DICKKOPF-1 Variant Antibodies and Methods of Use
Patent Information
- Application Number
- JP2024529460
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-14
- Filing Date
- 2022-11-17
- Publication Date
- 2025-11-25
AI Technical Summary
Current treatments for diseases associated with elevated Dickkopf WNT signaling pathway inhibitor 1 (DKK1) expression, such as cancer, are inadequate in effectively inhibiting its proliferation and growth.
Development of antibodies or antibody fragments, including monoclonal, polyclonal, bispecific, and humanized variants, with specific binding affinity to DKK1, designed to modulate its activity and inhibit its signaling pathway.
The developed antibodies demonstrate high affinity and specificity for DKK1, effectively blocking its signaling pathway and inhibiting cancer cell proliferation and growth, as shown in preclinical tumor models.
Abstract
Description
[Background technology]
[0001] (CROSS REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 280,840, filed November 18, 2021, U.S. Provisional Patent Application No. 63 / 286,522, filed December 6, 2021, U.S. Provisional Patent Application No. 63 / 374,497, filed September 2, 2022, and U.S. Provisional Patent Application No. 63 / 379,634, filed October 14, 2022, each of which is incorporated by reference in its entirety.
[0002] Dickkopf WNT signaling pathway inhibitor 1 (also known as dickkopf-1 or DKK1) is a secreted glycoprotein characterized by two cysteine-rich domains that mediate protein-protein interactions. DKK1 is involved in embryonic development of the heart, head, and forelimbs by inhibiting the WNT signaling pathway. In adults, elevated expression of this gene has been observed in many human cancers, and the protein may promote proliferation, invasion, and growth of cancer cell lines. Given the role of DKK1 in various diseases and disorders, improved therapeutic approaches are needed.
[0003] (Incorporated by reference) All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. Summary of the Invention
[0004] Provided herein is an antibody or antibody fragment comprising a variable domain heavy chain region (VH), the VH comprising complementarity determining regions CDRH1, CDRH2 and CDRH3, wherein (a) the amino acid sequence of CDRH1 is as set forth in any one of SEQ ID NOs: 1-98 or 919-1332, (b) the amino acid sequence of CDRH2 is as set forth in any one of SEQ ID NOs: 99-196 or 1333-1746, and (c) the amino acid sequence of CDRH3 is as set forth in any one of SEQ ID NOs: 197-294 or 1747-2160. Further provided herein is an antibody or antibody fragment, the antibody being a monoclonal antibody, a polyclonal antibody, a bispecific antibody, a multispecific antibody, a grafted antibody, a human antibody, a humanized antibody, a synthetic antibody, a chimeric antibody, a camelid antibody, a single-chain Fv (scFv), a single-chain antibody, a Fab fragment, a F(ab')2 fragment, a Fd fragment, a Fv fragment, a single domain antibody, an isolated complementarity determining region (CDR), a diabody, a fragment consisting of only a single monomeric variable domain, a disulfide-linked Fv (sdFv), an intrabody, an anti-idiotypic (anti-Id) antibody, or an ab antigen-binding fragment thereof. Further provided herein is an antibody or antibody fragment, the antibody being a single domain antibody. Further provided herein is an antibody or antibody fragment, the antibody or antibody fragment. Further provided herein is an antibody or antibody fragment, the antibody or antibody fragment having a K of less than 50 nM. D Further provided herein is an antibody or antibody fragment, the antibody or antibody fragment having a K of less than 25 nM. D Further provided herein is an antibody or antibody fragment, the antibody or antibody fragment having a K of less than 10 nM. D Further provided herein is an antibody or antibody fragment, the antibody or antibody fragment having a K of less than 5 nM. D Further provided herein is an antibody or antibody fragment, wherein the antibody or antibody fragment binds to DKK1.
[0005] Provided herein is an antibody or antibody fragment comprising a variable domain heavy chain region (VH) comprising an amino acid sequence at least about 90% identical to any one of SEQ ID NOs: 295-392, 394-712 or 2164-2258, wherein the VL comprises at least 90% sequence identity to any one of SEQ ID NOs: 713-918. Further provided herein is an antibody or antibody fragment, wherein the antibody or antibody fragment binds to spike glycoprotein. Further provided herein is an antibody or antibody fragment, wherein the antibody or antibody fragment binds to the receptor binding domain of spike glycoprotein. Further provided herein is an antibody or antibody fragment, wherein the antibody or antibody fragment has a K D Further provided herein is an antibody or antibody fragment, the antibody or antibody fragment having a K of less than 25 nM. D Further provided herein is an antibody or antibody fragment, the antibody or antibody fragment having a K of less than 10 nM. D Further provided herein is an antibody or antibody fragment, the antibody or antibody fragment having a K of less than 5 nM. D Further provided herein is an antibody or antibody fragment, wherein the antibody is a monoclonal antibody, a polyclonal antibody, a bispecific antibody, a multispecific antibody, a grafted antibody, a human antibody, a humanized antibody, a synthetic antibody, a chimeric antibody, a camelid antibody, a single chain Fv (scFv), a single chain antibody, a Fab fragment, a F(ab')2 fragment, a Fd fragment, a Fv fragment, a single domain antibody, an isolated complementarity determining region (CDR), a diabody, a fragment consisting of only a single monomeric variable domain, a disulfide-linked Fv (sdFv), an intrabody, an anti-idiotypic (anti-Id) antibody, or an ab antigen-binding fragment thereof. Further provided herein is an antibody or antibody fragment, wherein the antibody is a single domain antibody.
[0006] Provided herein is a nucleic acid composition comprising: a first nucleic acid encoding a variable domain heavy chain region (VH) comprising complementarity determining regions CDRH1, CDRH2, and CDRH3, wherein (a) the amino acid sequence of CDRH1 is as set forth in any one of SEQ ID NOs: 1-98 or 919-1332; (b) the amino acid sequence of CDRH2 is as set forth in any one of SEQ ID NOs: 99-196 or 1333-1746; and (c) the amino acid sequence of CDRH3 is as set forth in any one of SEQ ID NOs: 197-294 or 1747-2160; and an excipient.
[0007] Provided herein is a nucleic acid composition comprising: a) a first nucleic acid encoding a variable domain heavy chain region (VH) comprising an amino acid sequence that is at least about 90% identical to any one of SEQ ID NOs: 295 to 392, 394 to 712, or 2164 to 2258; and an excipient.
[0008] Provided herein is an antibody or antibody fragment comprising a variable domain light chain region (VL), the VL comprising complementarity determining regions CDRL1, CDRL2 and CDRL3, wherein (a) the amino acid sequence of CDRL1 is as set forth in any one of SEQ ID NOs: 2259 to 2464, (b) the amino acid sequence of CDRL2 is as set forth in any one of SEQ ID NOs: 2465 to 2521, and (c) the amino acid sequence of CDRL3 is as set forth in any one of SEQ ID NOs: 2522 to 2727. Further provided herein is an antibody or antibody fragment, the antibody being a monoclonal antibody, a polyclonal antibody, a bispecific antibody, a multispecific antibody, a grafted antibody, a human antibody, a humanized antibody, a synthetic antibody, a chimeric antibody, a camelid antibody, a single chain Fv (scFv), a single chain antibody, a Fab fragment, a F(ab')2 fragment, a Fd fragment, a Fv fragment, a single domain antibody, an isolated complementarity determining region (CDR), a diabody, a fragment consisting of only a single monomeric variable domain, a disulfide-linked Fv (sdFv), an intrabody, an anti-idiotypic (anti-Id) antibody, or an ab antigen-binding fragment thereof. Further provided herein is an antibody or antibody fragment, the antibody being a single domain antibody. Further provided herein is an antibody or antibody fragment, the antibody or antibody fragment. Further provided herein is an antibody or antibody fragment, the antibody or antibody fragment having a K of less than 50 nM. D Further provided herein is an antibody or antibody fragment, the antibody or antibody fragment having a K of less than 25 nM. D Further provided herein is an antibody or antibody fragment, the antibody or antibody fragment having a K of less than 10 nM. D Further provided herein is an antibody or antibody fragment, the antibody or antibody fragment having a K of less than 5 nM. D Further provided herein is an antibody or antibody fragment, wherein the antibody or antibody fragment binds to DKK1.
[0009] Provided herein is an antibody or antibody fragment comprising a variable domain light chain region (VL) comprising an amino acid sequence at least about 90% identical to a sequence set forth in any one of SEQ ID NOs: 713-918. Further provided herein is an antibody or antibody fragment, wherein the antibody or antibody fragment binds to spike glycoprotein. Further provided herein is an antibody or antibody fragment, wherein the antibody or antibody fragment binds to the receptor binding domain of spike glycoprotein. Further provided herein is an antibody or antibody fragment, wherein the antibody or antibody fragment has a K D Further provided herein is an antibody or antibody fragment, the antibody or antibody fragment having a K of less than 25 nM. D Further provided herein is an antibody or antibody fragment, the antibody or antibody fragment having a K of less than 10 nM. D Further provided herein is an antibody or antibody fragment, the antibody or antibody fragment having a K of less than 5 nM. D Further provided herein is an antibody or antibody fragment, wherein the antibody is a monoclonal antibody, a polyclonal antibody, a bispecific antibody, a multispecific antibody, a grafted antibody, a human antibody, a humanized antibody, a synthetic antibody, a chimeric antibody, a camelid antibody, a single chain Fv (scFv), a single chain antibody, a Fab fragment, a F(ab')2 fragment, a Fd fragment, a Fv fragment, a single domain antibody, an isolated complementarity determining region (CDR), a diabody, a fragment consisting of only a single monomeric variable domain, a disulfide-linked Fv (sdFv), an intrabody, an anti-idiotypic (anti-Id) antibody, or an ab antigen-binding fragment thereof. Further provided herein is an antibody or antibody fragment, wherein the antibody is a single domain antibody.
[0010] Provided herein is a nucleic acid composition comprising a first nucleic acid encoding a variable domain light chain region (VL) comprising complementarity determining regions CDRL1, CDRL2 and CDRL3, wherein (a) the amino acid sequence of CDRL1 is as set forth in any one of SEQ ID NOs: 2259-2464, (b) the amino acid sequence of CDRL2 is as set forth in any one of SEQ ID NOs: 2465-2521, and (c) the amino acid sequence of CDRL3 is as set forth in any one of SEQ ID NOs: 2522-2727, and an excipient.
[0011] Provided herein is a nucleic acid composition comprising: a) a first nucleic acid encoding a variable domain light chain region (VL) comprising an amino acid sequence at least about 90% identical to any one of SEQ ID NOs: 713-918; and an excipient. [Brief description of the drawings]
[0012] This patent or application contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee. [Figure 1A] 1 shows a first schematic diagram of an immunoglobulin. [Figure 1B] 1 shows a second schematic diagram of an immunoglobulin. [Diagram 2] FIG. 1 shows a schematic diagram of the motif for placement in immunoglobulins. [Diagram 3] FIG. 1 shows a diagram of steps illustrating an exemplary process workflow for gene synthesis as disclosed herein. [Figure 4] 1 illustrates an example computer system. [Diagram 5] FIG. 1 is a block diagram illustrating the architecture of a computer system. [Figure 6]FIG. 1 illustrates a network configured to incorporate multiple computer systems, multiple mobile phones and personal data assistants, and Network Attached Storage (NAS). [Figure 7] 1 is a block diagram of a multi-processor computer system using a shared virtual address memory space. [Figure 8A] 1 shows a schematic diagram of an immunoglobulin comprising a VH domain joined to a VL domain using a linker. [Figure 8B] FIG. 1 shows a schematic diagram of the complete domain structure of an immunoglobulin, including the VH domain connected to the VL domain using a linker, leader sequence and pill sequence. [Figure 8C] A schematic diagram of the four framework elements (FW1, FW2, FW3, FW4) and the three variable CDR (L1, L2, L3) elements of a VL or VH domain is shown. [Figure 9A] Long-read NGS sequencing of the eluted phage pool of antibody pool A. At the top of the figure, the cluster enrichment count (the number of instances an antibody appears in) is plotted against the cluster rank, and the antibody rank order of antibodies by cluster size is listed. At the bottom of the figure, a parallel histogram is shown showing the distribution of HCDR3 lengths among the top 95 antibody clusters. [Figure 9B] Long-read NGS sequencing of the eluted phage pool of antibody pool B. At the top of the figure, the cluster enrichment count (the number of instances an antibody appears in) is plotted against the cluster rank, and the antibody rank order of antibodies by cluster size is listed. At the bottom of the figure, a parallel histogram is shown showing the distribution of HCDR3 lengths among the top 95 antibody clusters. [Figure 9C]Long-read NGS sequencing of the eluted phage pool of antibody pool C. At the top of the figure, the cluster enrichment count (the number of instances an antibody appears in) is plotted against the cluster rank, and the antibody rank order of antibodies by cluster size is listed. At the bottom of the figure, a parallel histogram is shown showing the distribution of HCDR3 lengths among the top 95 antibody clusters. [Figure 10A] Distribution of antibody yields from 1.2 mL high-throughput antibody expression and purification among antibodies identified from three library pools. Dots are color-coded according to whether the antibody was identified by phage ELISA screening (blue) or NGS enrichment data (green). [Figure 10B] The distribution of antibody binding affinities for DKK1 as measured by SPR (Carterra) is shown. Dots are color-coded according to whether the antibody was identified by phage ELISA screening (blue) or by NGS enrichment data (green). [Figure 10C] The distribution of MFI ratios among antibodies identified from the three library pools is shown. The MFI ratio is defined as the measured MFI of antibodies binding to HEK293 cells overexpressing DKK1 divided by the measured MFI of antibodies binding to HEK293 cells. The points are color-coded according to whether the antibodies were identified by phage ELISA screening (blue) or NGS enrichment data (green). [Figure 11A] The relationship between MFI ratio as measured by SPR and binding affinity to DKK1 is shown. The size of each dot corresponds to the antibody yield from 1.2 mL of high-throughput antibody expression and purification. The dots are color-coded by the library pool used during panning. [Figure 11B]The relationship between MFI ratio as measured by SPR and binding affinity to DKK1 is shown. The size of each dot corresponds to the antibody yield from 1.2 mL of high-throughput antibody expression and purification. Dots are color-coded according to whether the antibody was identified by phage ELISA screening (blue) or NGS enrichment data (green). [Figure 12A] Carterra SPR kinetic graph showing VHH-Fc hits identified from NGS sequencing that bind with high affinity to DKK1. Antibody lawn (10ug / mL), 0-500nM antigen, HBSTE + 0.5mg / mL BSA pH 7.4. [Figure 12B] 1 shows Carterra SPR kinetic graphs showing VHH-Fc hits identified from ELISA screening that bind with high affinity to DKK1. [Figure 12C] 13 shows additional Carterra SPR kinetic graphs showing VHH-Fc hits identified from NGS sequencing that bind with high affinity to DKK1. [Figure 12D] 13 shows additional Carterra SPR kinetic graphs showing VHH-Fc hits identified from NGS sequencing that bind with high affinity to DKK1. [Figure 13] 1 shows the results of a TCF / LEF reporter (Wnt signaling) assay. Activation of Wnt signaling is plotted together with SPR binding affinity. [Figure 14A] Figure 14A shows the activation of primary immune cells in vitro, and Figure 14B shows an immune cell activation assay using peripheral blood mononuclear cells (PBMCs) and interferon gamma (IFN or IFN-γ). [Figure 14B]Figure 14B shows the activation of primary immune cells in vitro. Figure 14B shows an immune cell activation assay using PBMCs and granulocyte-macrophage colony-stimulating factor (GM-CSF), a marker for NK cell activation. [Figure 14C] Figure 14 shows activation of primary immune cells in vitro. Human PBMCs are treated with immune stimulants, mWnt3a, hDKK1 and Dkk1 leads from the ML synthetic library (Figure 14C) and MLs from the VHH library (Figure 14D). Cytokine release of GM-CSF is measured by ELISA. [Figure 14D] Figure 14 shows activation of primary immune cells in vitro. Human PBMCs are treated with immune stimulants, mWnt3a, hDKK1 and Dkk1 leads from the ML synthetic library (Figure 14C) and MLs from the VHH library (Figure 14D). Cytokine release of GM-CSF is measured by ELISA. [Figure 15A] Figure 1 shows the results of a tumor killing assay: Activated immune cells kill PC3 cells and hDKK1 treatment inhibits cytotoxicity. [Figure 15B] 1 shows a graph of the results of a tumor killing assay. [Figure 15C] Specific hits from the tumor killing assay that were also detected in the TCF / LEF reporter (Wnt signaling) assay are highlighted. [Figure 15D] We show that blocking hDKK1 interaction with the receptor by DKK1 leads restores the cytotoxic potency of MLs from the ML synthetic library and the VHH library. [Figure 15E] 1 shows the results of viability of PC3 tumor cells. [Figure 15F] The top clones in the PC3 cytotoxicity assay are shown. [Figure 15G] A subset of the top clones in the PC3 cytotoxicity assay is shown. [Figure 16] Antibody yield results from 1 mL of Expi293 cell culture are shown. [Figure 17A] Figure 1 shows anti-DKK1 binding to hDKK1 by SPR analysis, showing two epitope bindings evident between DKK1 reads (activation of Wnt signaling and immune response). [Figure 17B] Binding of anti-DKK1 to hDKK1 by SPR analysis. An example of hDDK1 protein with CRD1 and CRD2 annotated is shown. [Figure 17C] 1 shows anti-DKK1 binding to hDKK1 by SPR analysis, and DKK1 leads binding to hDKK1 CRD1 and / or hDKK1 CRD2 result in distinct activation pathways. [Figure 18A] Screening of Wnt TCF / LEF reporter assay. Wnt TCF / LEF signaling is blocked by DKK1 binding to LRP5 / 6. DKK1 leads were screened from VHH library (Figure 18A), ML synthetic library (Figure 18B), and ML from VHH library (Figure 18C). [Figure 18B] Screening of Wnt TCF / LEF reporter assay. Wnt TCF / LEF signaling is blocked by DKK1 binding to LRP5 / 6. DKK1 leads were screened from VHH library (Figure 18A), ML synthetic library (Figure 18B), and ML from VHH library (Figure 18C). [Figure 18C] Screening of Wnt TCF / LEF reporter assay. Wnt TCF / LEF signaling is blocked by DKK1 binding to LRP5 / 6. DKK1 leads were screened from VHH library (Figure 18A), ML synthetic library (Figure 18B), and ML from VHH library (Figure 18C). [Figure 19A]BsAb functional assay is shown. DKK1-99 binds to DKK1 CRD1 to activate immune response, and DKK1-100 binds to DKK1 CRD2 to activate Wnt signaling. DKK1-99 and DKK1-100 bispecific antibodies (Figure 19A) show potency in activating both Wnt (Figure 19B) and immune response (Figure 19C). Figure 19D shows another graph of immune response activation. [Figure 19B] BsAb functional assay is shown. DKK1-99 binds to DKK1 CRD1 to activate immune response, and DKK1-100 binds to DKK1 CRD2 to activate Wnt signaling. DKK1-99 and DKK1-100 bispecific antibodies (Figure 19A) show potency in activating both Wnt (Figure 19B) and immune response (Figure 19C). Figure 19D shows another graph of immune response activation. [Figure 19C] BsAb functional assay is shown. DKK1-99 binds to DKK1 CRD1 to activate immune response, and DKK1-100 binds to DKK1 CRD2 to activate Wnt signaling. DKK1-99 and DKK1-100 bispecific antibodies (Figure 19A) show potency in activating both Wnt (Figure 19B) and immune response (Figure 19C). Figure 19D shows another graph of immune response activation. [Figure 19D] BsAb functional assay is shown. DKK1-99 binds to DKK1 CRD1 to activate immune response, and DKK1-100 binds to DKK1 CRD2 to activate Wnt signaling. DKK1-99 and DKK1-100 bispecific antibodies (Figure 19A) show potency in activating both Wnt (Figure 19B) and immune response (Figure 19C). Figure 19D shows another graph of immune response activation. [Figure 20A] Figure 1 shows that DKK1 is involved in tumor regression. Schematic diagram of PC3 cells inoculated into mice. Treatment was started when the mean tumor volume was approximately 100 mm3, and 10 mg / kg was administered by intraperitoneal injection once every 3 days for 8 cycles. Tumor size was measured 3 times a week. [Figure 20B]Figure 1 shows that DKK1 is involved in tumor regression. Anti-DKK1 treatment inhibits tumor growth and is effective in tumor suppression. [Figure 20C] Figure 1 shows that DKK1 is involved in tumor regression. Anti-DKK1 treatment inhibited tumor growth, demonstrating efficacy in tumor inhibition from days 1 to 7 of the experiment. [Figure 20D] Figure 1 shows that DKK1 is involved in tumor regression.The mean tumor volumes from days 1 to 7 of the experiment are shown. [Figure 21] FIG. 1 shows a schematic diagram of panning rounds for DKK1 antibody production. [Figure 22A] 1 shows that the antagonism of DKK1 inhibition of WNT in the TCF / LEF assay is biphasic.The control DKN-01 antibody is shown. [Figure 22B] 1 shows that the antagonistic effect of WNT on DKK1 inhibition in the TCF / LEF assay is biphasic. Results for DKK1-28 are shown. [Figure 22C] 1 shows that the antagonistic effect of WNT inhibition of DKK1 in the TCF / LEF assay is biphasic. Results for DKK1-100 are shown. [Figure 23A] 1 shows concordance in ranking of transient and cell line TCF / LEF reporters in functional assays. [Figure 23B] A subset of the results from FIG. 23A is shown. [Figure 24] 1 shows the development of a DKK1 / LRP6 binding assay. [Figure 25A] 1 shows that the functional antagonist DKN-01 enhances binding of DKK1 to LRP6. [Figure 25B] 1 shows that the functional antagonist DKK1-100 enhances binding of DKK1 to LRP6. [Figure 25C] 1 shows that the functional antagonist DKK1-28 enhances binding of DKK1 to LRP6. [Figure 26A] 1 shows the results of a primary immune cell reactivation assay using an IFN-γ marker of immune cell activation. [Figure 26B]1 shows the results of a primary immune cell reactivation assay using GM-CSF marker for immune cell activation. [Figure 27A] The results of primary NK cell activation are shown. [Figure 27B] The results of immune cell activation assay of top clones are shown. [Figure 27C] A subset of the results from Figure 27B is shown. [Figure 28A] Identification of antagonist DKK1-473 by signaling titration assay. [Figure 28B] Identification of antagonist DKK1-478 by signaling titration assay. [Figure 28C] Identification of antagonist DKK1-477 (FIG. 28C) by signaling titration assay. [Figure 28D] Identification of antagonist DKK1-448 (FIG. 28D) by signaling titration assay. [Figure 29A] The results of the immunoassay are shown. [Figure 29B] A subset of the results from FIG. 29A is shown. [Diagram 30] Figure 2 shows immune cell-mediated killing of lung tumor organoids by DKK1 inhibition. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0013] The present disclosure employs, unless otherwise indicated, conventional molecular biology techniques within the skill of the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.
[0014] (definition) Throughout this disclosure, various embodiments are presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of any embodiment. Thus, the description of a range should be considered to have specifically disclosed all possible subranges within that range, and individual numerical values to the nearest tenth of the lower limit, unless the context clearly dictates otherwise. For example, the description of a range such as 1-6 is considered to have specifically disclosed subranges such as 1-3, 1-4, 1-5, 2-4, 2-6, 3-6, and individual values within that range, such as 1.1, 2, 2.3, 5, 5.9, etc. This applies regardless of the breadth of the range. The upper and lower limits of these intervening ranges may be independently included in the smaller ranges, and are also encompassed within the disclosure, subject to any specifically excluded limits in the stated range. Where the stated range includes upper and lower limits, or both, unless the context clearly indicates otherwise, ranges excluding one or both of those included limits are also included in the disclosure.
[0015] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit any embodiment. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms unless the context clearly indicates otherwise. Furthermore, as used herein, it will be understood that the terms "comprises" and / or "comprising" specify the presence of stated features, integers, steps, operations, elements, components, and / or groups, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0016] As used herein, unless otherwise specified or clear from the context, the term "about" in reference to numerical values or numerical ranges is understood to mean the stated numerical value and values + / - 10% thereof, or, for values stated as a range, 10% below the stated lower limit and 10% above the stated upper limit.
[0017] Unless otherwise indicated, the term "nucleic acid" as used herein encompasses single-stranded molecules as well as double- or triple-stranded nucleic acids. In double- or triple-stranded nucleic acids, the nucleic acid strands need not be coextensive (i.e., a double-stranded nucleic acid need not be double-stranded along the entire length of both strands). Nucleic acid sequences are written in a 5' to 3' direction unless otherwise indicated. The methods described herein provide for the production of isolated nucleic acids. The methods described herein further provide for the production of isolated and purified nucleic acids. A "nucleic acid" as referred to herein can comprise at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000 or more bases in length. Further provided herein are methods for the synthesis of any number of polypeptide segments encoding nucleotide sequences, including sequences encoding non-ribosomal peptides (NRPs), sequences encoding non-ribosomal peptide-synthetase (NRPS) modules and synthetic variants, polypeptide segments of other modular proteins such as antibodies, polypeptide segments of other protein families that contain non-coding DNA or RNA, such as regulatory sequences, e.g., promoters, transcription factors, enhancers, siRNAs, shRNAs, RNAi, miRNAs, small nucleolar RNAs derived from microRNAs, or any functional or structural DNA or RNA unit of interest.The following are non-limiting examples of polynucleotides: coding or non-coding regions of a gene or gene fragment, intergenic DNA, locus(s) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, short interfering RNA (siRNA), short-hairpin RNA (shRNA), micro-RNA (miRNA), small nucleolar RNA, ribozymes, complementary DNA (cDNA) (a DNA representation of messenger RNA (mRNA), usually obtained by reverse transcription or amplification of mRNA), DNA molecules produced by synthesis or amplification, genomic DNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. A cDNA encoding a gene or gene fragment as referred to herein may include at least one region encoding an exon sequence without intervening intronic sequences in the genomic equivalent sequence.
[0018] DKK1 Library Provided herein are methods and compositions related to WNT signaling pathway inhibitor 1 (DKK1) variant immunoglobulins (e.g., antibodies, VHHs) comprising nucleic acids encoding immunoglobulins comprising a DKK1 binding domain. The immunoglobulins described herein can stably support the DKK1 binding domain. The libraries described herein can be further diversified to provide variant libraries comprising nucleic acids each encoding a predefined variant of at least one predefined reference nucleic acid sequence. Also described herein are protein libraries that can be created when the nucleic acid libraries are translated. In some examples, the nucleic acid libraries described herein are transferred into cells to create cell libraries. Also provided herein are downstream applications of libraries synthesized using the methods described herein. Downstream applications include identification of variant nucleic acid or protein sequences with biologically relevant functions, such as improved stability, affinity, binding, functional activity, and enhanced functions for the treatment or prevention of disease conditions associated with DKK1.
[0019] Provided herein are libraries that include nucleic acids encoding immunoglobulins. In some examples, the immunoglobulins are antibodies. As used herein, the term antibody is understood to include proteins with the characteristic two-armed Y-shape of a typical antibody molecule, and one or more fragments of an antibody that retain the ability to specifically bind to an antigen. Exemplary antibodies include monoclonal antibodies, polyclonal antibodies, bispecific antibodies, multispecific antibodies, grafted antibodies, human antibodies, humanized antibodies, synthetic antibodies, chimeric antibodies, camelid antibodies, single chain Fvs (scFvs) (including fragments in which the VL and VH are linked using recombinant methods by a synthetic or natural linker, allowing the VL and VH regions to be paired together to produce a single protein chain that forms a monovalent molecule, including single chain Fabs and scFabs), single chain antibodies, Fab fragments (including monovalent fragments including the VL, VH, CL and CH1 domains), F(ab')2 fragments (including two Fab fragments linked by a disulfide bridge at the hinge region), and F(ab')2 fragments (including single chain Fab fragments that are linked by a disulfide bridge at the hinge region). The libraries include, but are not limited to, a bivalent fragment comprising a VL domain, a Fd fragment (including a fragment comprising a VH and a CH1 fragment), an Fv fragment (including a fragment comprising the VL and VH domains of a single arm of an antibody), a single domain antibody (dAb or sdAb) (including a fragment comprising a VH domain), an isolated complementarity determining region (CDR), a diabody (including a fragment comprising a bivalent dimer, such as two VL and VH domains that bind to each other and recognize two different antigens), a fragment consisting of only a single monomeric variable domain, a disulfide-linked Fv (sdFv), an intrabody, an anti-idiotypic (anti-Id) antibody, or an ab antigen-binding fragment thereof. In some examples, the libraries disclosed herein include nucleic acids encoding immunoglobulins, where the immunoglobulins are Fv antibodies, including Fv antibodies consisting of the smallest antibody fragment that contains a complete antigen recognition site and antigen binding site. In some embodiments, an Fv antibody consists of a dimer of one heavy- and one light-chain variable domain in tight, non-covalent association, where the three hypervariable regions of each variable domain interact to define an antigen-binding site on the surface of the VH-VL dimer, in some embodiments, the six hypervariable regions confer antigen-binding specificity to the antibody.In some embodiments, the single variable domain (or half of an Fv containing only three hypervariable regions specific for an antigen, including single domain antibodies isolated from camelids containing one heavy chain variable domain, such as VHH antibodies or nanobodies) has the ability to recognize and bind to the antigen. In some examples, the libraries disclosed herein include nucleic acids encoding immunoglobulins, where the immunoglobulins are single chain Fvs or scFvs, and include antibody fragments containing a VH, a VL, or both a VH and a VL domain, where both domains are present in a single polypeptide chain. In some embodiments, the Fv polypeptide further comprises a polypeptide linker between the VH and VL domains, allowing the scFv to form the desired structure for antigen binding. In some examples, the scFv is attached to an Fc fragment, or the VHH is attached to an Fc fragment (including a minibody). In some examples, the antibodies include immunoglobulin molecules and immunologically active fragments of immunoglobulin molecules, such as molecules containing an antigen binding site. Immunoglobulin molecules can be of any type (e.g., IgG, IgE, IgM, IgD, IgA, and IgY), class (e.g., IgG1, IgG2, IgG3, IgG4, IgA1, and IgA2), or subclass.
[0020] In some embodiments, the library comprises immunoglobulins adapted to the species of the intended therapeutic target. In general, these methods include "mammalization" and include methods to transfer donor antigen binding information to a less immunogenic mammalian antibody acceptor to create a useful therapeutic treatment. In some examples, the mammal is a mouse, rat, horse, sheep, cow, primate (e.g., chimpanzee, baboon, gorilla, orangutan, monkey), dog, cat, pig, donkey, rabbit, or human. In some examples, libraries and methods are provided herein for felineization and canineization of antibodies.
[0021] "Humanized" forms of non-human antibodies may be chimeric antibodies that contain minimal sequence derived from the non-human antibody. Humanized antibodies are generally human antibodies (recipient antibodies) in which one or more CDR residues have been replaced with one or more CDR residues of a non-human antibody (donor antibody). The donor antibody may be any suitable non-human antibody, such as a mouse, rat, rabbit, chicken, or non-human primate antibody with the desired specificity, affinity, or biological effect. In some instances, selected framework region residues of the recipient antibody are replaced with the corresponding framework region residues of the donor antibody. Humanized antibodies may also contain residues that are not found in either the recipient antibody or the donor antibody. In some instances, these modifications are made to further improve the performance of the antibody.
[0022] "Caninization" can include methods of transferring non-canine antigen binding information from a donor antibody to a less immunogenic canine antibody acceptor to create a therapeutic useful as a therapeutic in dogs. In some examples, the caninized forms of non-canine antibodies provided herein are chimeric antibodies, containing minimal sequence derived from the non-canine antibody. In some examples, the caninized antibody is a canine antibody sequence (the "acceptor" or "recipient" antibody) in which hypervariable region residues of the recipient have been replaced with hypervariable region residues of a non-canine species (the "donor" antibody), such as mouse, rat, rabbit, cat, dog, goat, chicken, cow, horse, llama, camel, dromedary, shark, non-human primate, human, humanized, recombinant, or engineered sequences with desired properties. In some examples, framework region (FR) residues of the canine antibody have been replaced with corresponding non-canine FR residues. In some examples, the caninized antibody includes residues that are not found in the recipient antibody or the donor antibody. In some instances, these modifications are made to further refine antibody performance. The caninized antibody may also contain at least a portion of the immunoglobulin constant region (Fc) of a canine antibody.
[0023] "Fenconization" may include methods of transferring non-feline antigen binding information from a donor antibody to a less immunogenic feline antibody receptor to create a therapeutic useful as a therapeutic in cats. In some examples, the felineized forms of non-feline antibodies provided herein are chimeric antibodies that contain minimal sequence derived from the non-feline antibody. In some examples, felineized antibodies are feline antibody sequences ("acceptor" or "recipient" antibody) in which hypervariable region residues of the recipient have been replaced with hypervariable region residues of a non-feline species ("donor" antibody), such as mouse, rat, rabbit, cat, dog, goat, chicken, cow, horse, llama, camel, dromedary, shark, non-human primate, human, humanized, recombinant, or engineered sequences with desired properties. In some examples, framework region (FR) residues of the feline antibody have been replaced with corresponding non-feline FR residues. In some examples, the felineized antibody contains residues that are not found in the recipient antibody or the donor antibody. In some examples, these modifications are made to further improve antibody performance. The fetinized antibody may also comprise at least a portion of the immunoglobulin constant region (Fc) of a feline antibody.
[0024] Provided herein is a library comprising nucleic acids encoding non-immunoglobulins. For example, the non-immunoglobulin is an antibody mimic. Exemplary antibody mimics include, but are not limited to, anticalins, affilins, affibody molecules, affimers, affitins, alphabodies, avimers, atrimers, DARPins, finomers, Kunitz domain-based proteins, monobodies, anticalins, knottins, armadillo repeat protein-based proteins, and bicyclic peptides.
[0025] The libraries described herein include nucleic acids encoding immunoglobulins that include mutations in at least one region of the immunoglobulin. Exemplary regions of an antibody for mutation include, but are not limited to, a complementarity determining region (CDR), a variable domain, or a constant domain. In some examples, the CDR is CDR1, CDR2, or CDR3. In some examples, the CDR is a heavy domain, including, but not limited to, CDRH1, CDRH2, and CDRH3. In some examples, the CDR is a light domain, including, but not limited to, CDRL1, CDRL2, and CDRL3. In some examples, the variable domain is a variable domain light chain (VL) or a variable domain heavy chain (VH). In some examples, the VL domain includes a kappa chain or a lambda chain. In some examples, the constant domain is a constant domain light chain (CL) or a constant domain heavy chain (CH).
[0026] The methods described herein provide for the synthesis of libraries comprising nucleic acids encoding immunoglobulins, each nucleic acid encoding a predetermined variant of at least one predetermined reference nucleic acid sequence. In some cases, the predetermined reference sequence is a nucleic acid sequence encoding a protein, and the variant library comprises sequences encoding at least a single codon mutation, such that multiple different variants of a single residue in the subsequent protein encoded by the synthesized nucleic acid are created by standard translation processes. In some examples, the variant library comprises diverse nucleic acids that collectively encode mutations at multiple positions. In some examples, the variant library comprises sequences encoding at least one codon mutation of a CDRH1, CDRH2, CDRH3, CDRL1, CDRL2, CDRL3, VL, or VH domain. In some examples, the variant library comprises sequences encoding multiple codon mutations of a CDRH1, CDRH2, CDRH3, CDRL1, CDRL2, CDRL3, VL, or VH domain. In some examples, the variant library includes sequences encoding mutations of multiple codons of framework element 1 (FW1), framework element 2 (FW2), framework element 3 (FW3), or framework element 4 (FW4). Exemplary numbers of codons for mutation include, but are not limited to, at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 225, 250, 275, 300, or more than 300 codons.
[0027] In some instances, at least one region of an immunoglobulin for mutation is from a heavy chain V gene family, a heavy chain D gene family, a heavy chain J gene family, a light chain V gene family, or a light chain J gene family. In some instances, the light chain V gene family comprises an immunoglobulin kappa (IGK) gene or an immunoglobulin lambda (IGL) gene. Exemplary genes include, but are not limited to, IGHV1-18, IGHV1-69, IGHV1-8, IGHV3-21, IGHV3-23, IGHV3-30 / 33rn, IGHV3-28, IGHV1-69, IGHV3-74, IGHV4-39, IGHV4-59 / 61, IGKV1-39, IGKV1-9, IGKV2-28, IGKV3-11, IGKV3-15, IGKV3-20, IGKV4-1, IGLV1-51, IGLV2-14, IGLV1-40, and IGLV3-1. In some examples, the gene is IGHV1-69, IGHV3-30, IGHV3-23, IGHV3, IGHV1-46, IGHV3-7, IGHV1, or IGHV1-8. In some examples, the gene is IGHV1-69 and IGHV3-30. In some examples, the gene is IGHJ3, IGHJ6, IGHJ, IGHJ4, IGHJ5, IGHJ2, or IGH1. In some examples, the gene is IGHJ3, IGHJ6, IGHJ, or IGHJ4.
[0028] Provided herein are libraries containing nucleic acids encoding immunoglobulins, which are synthesized with a variable number of fragments. In some examples, the fragments include CDRH1, CDRH2, CDRH3, CDRL1, CDRL2, CDRL3, VL, or VH domains. In some examples, the fragments include framework element 1 (FW1), framework element 2 (FW2), framework element 3 (FW3), or framework element 4 (FW4). In some examples, the immunoglobulin library is synthesized with at least or about 2 fragments, 3 fragments, 4 fragments, 5 fragments, or more than 5 fragments. The length of each of the nucleic acid fragments or the average length of the synthesized nucleic acids can be at least about 50, 75, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 525, 550, 575, 600, or more than 600 base pairs. In some examples, the length is about 50-600, 75-575, 100-550, 125-525, 150-500, 175-475, 200-450, 225-425, 250-400, 275-375, or 300-350 base pairs.
[0029] Libraries containing nucleic acids encoding immunoglobulins described herein contain amino acids of various lengths when translated. In some examples, the length or average length of each of the synthesized amino acid fragments can be at least or about 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, or more than 150 amino acids. In some examples, the length of the amino acids is about 15-150, 20-145, 25-140, 30-135, 35-130, 40-125, 45-120, 50-115, 55-110, 60-110, 65-105, 70-100, or 75-95 amino acids. In some examples, the length of amino acids is from about 22 amino acids to about 75 amino acids. In some examples, the immunoglobulin comprises at least or about 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, or more than 5000 amino acids.
[0030] A multiplicity of variant sequences for at least one region of an immunoglobulin for mutation are de novo synthesized using the methods described herein. In some examples, the multiplicity of variant sequences are de novo synthesized for CDRH1, CDRH2, CDRH3, CDRL1, CDRL2, CDRL3, VL, VH, or a combination thereof. In some examples, the multiplicity of variant sequences are de novo synthesized for framework element 1 (FW1), framework element 2 (FW2), framework element 3 (FW3), or framework element 4 (FW4). The number of variant sequences can be at least about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, or more than 500 sequences. In some examples, the number of variant sequences is at least about 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, or more than 8000 sequences. In some examples, the number of variant sequences is about 10-500, 25-475, 50-450, 75-425, 100-400, 125-375, 150-350, 175-325, 200-300, 225-375, 250-350, or 275-325 sequences.
[0031] The variant sequence of at least one region of an immunoglobulin differs in length or sequence in some examples. In some examples, the at least one region synthesized de novo is for CDRH1, CDRH2, CDRH3, CDRL1, CDRL2, CDRL3, VL, VH, or a combination thereof. In some examples, the at least one region synthesized de novo is for framework element 1 (FW1), framework element 2 (FW2), framework element 3 (FW3), or framework element 4 (FW4). In some examples, the variant sequence comprises at least or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, or more than 50 variant nucleotides or amino acids compared to the wild type. In some examples, the variant sequences contain at least or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 additional nucleotides or amino acids compared to the wild type. In some examples, the variant sequences contain at least or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 fewer nucleotides or amino acids compared to the wild type. In some examples, the library contains at least or about 10 1 , 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , or 10 10 Contains more than 100 variants.
[0032] After synthesis of the library described herein, the library can be used for screening and analysis. For example, the library is assayed for displayability and panning of the library. In some examples, displayability is assayed using a selectable tag. Exemplary tags include, but are not limited to, radioactive labels, fluorescent labels, enzymes, chemiluminescent tags, colorimetric tags, affinity tags, or other labels or tags known in the art. In some examples, the tag is histidine, polyhistidine, myc, hemagglutinin (HA), or FLAG. In some examples, libraries are assayed by sequencing using a variety of methods, including, but not limited to, single-molecule real-time (SMRT) sequencing, polony sequencing, sequencing by ligation, reversible terminator sequencing, proton detection sequencing, ion semiconductor sequencing, nanopore sequencing, electronic sequencing, pyrosequencing, Maxam-Gilbert sequencing, chain termination (e.g., Sanger) sequencing, +S sequencing, or sequencing by synthesis.
[0033] In some examples, the library is assayed for functional activity, structural stability (e.g., thermal or pH stability), expression, specificity, or a combination thereof. In some examples, the library is assayed for foldable immunoglobulins (e.g., antibodies). In some examples, regions of antibodies are assayed for functional activity, structural stability, expression, specificity, folding, or a combination thereof. For example, VH or VL regions are assayed for functional activity, structural stability, expression, specificity, folding, or a combination thereof.
[0034] DKK1 Library Provided herein are DKK1 variant immunoglobulins (e.g., antibodies, VHHs) that include nucleic acids encoding immunoglobulins (e.g., antibodies) that bind to DKK1. In some examples, the immunoglobulin sequence of the DKK1 binding domain is determined by the interaction between the DKK1 binding domain and DKK1.
[0035] The sequence of DKK1 binding domain based on surface interaction of DKK1 is analyzed using various methods. For example, several kinds of computer analysis are performed. In some examples, structural analysis is performed. In some examples, sequence analysis is performed. Sequence analysis can be performed using databases known in the art. Non-limiting examples of databases include, but are not limited to, NCBI BLAST (blast.ncbi.nlm.nih.gov / Blast.cgi), UCSC Genome Browser (genome.ucsc.edu / ), UniProt (www.uniprot.org / ) and IUPHAR / BPS Guide to PHARMACOLOGY (guidetopharmacology.org / ).
[0036] Herein, the DKK1 binding domain is designed based on sequence analysis between various organisms.For example, sequence analysis is carried out to identify the homologous sequence in different organisms.Exemplary organisms include, but are not limited to, mouse, rat, horse, sheep, cow, primate (e.g., chimpanzee, baboon, gorilla, orangutan, monkey), dog, cat, pig, donkey, rabbit, fish, fly, and human.
[0037] Following identification of the DKK1 binding domain, a library can be created that includes nucleic acids encoding the DKK1 binding domain. In some examples, the library of DKK1 binding domains includes sequences of DKK1 binding domains designed based on conformational ligand interactions, peptide ligand interactions, small molecule ligand interactions, the extracellular domain of DKK1, or antibodies targeting DKK1. In some examples, the library of DKK1 binding domains includes sequences of DKK1 binding domains designed based on peptide ligand interactions. The library of DKK1 binding domains can be translated to create a protein library. In some examples, the library of DKK1 binding domains is translated to create a peptide library, an immunoglobulin library, derivatives thereof, or combinations thereof. In some examples, the library of DKK1 binding domains is translated to create a protein library and further modified to create a peptidomimetic library. In some examples, the library of DKK1 binding domains is translated to create a protein library that is used to create small molecules.
[0038] The methods described herein provide for the synthesis of a library of DKK1 binding domains, each of which comprises nucleic acids encoding a predetermined variant of at least one predetermined reference nucleic acid sequence. In some cases, the predetermined reference sequence is a nucleic acid sequence encoding a protein, and the variant library comprises sequences encoding at least a single codon mutation, such that multiple different variants of a single residue in the subsequent protein encoded by the synthesized nucleic acid are created by standard translation processes. In some examples, the library of DKK1 binding domains comprises diverse nucleic acids that collectively encode mutations at multiple positions. In some examples, the variant library comprises sequences encoding at least one codon mutation of the DKK1 binding domain. In some examples, the variant library comprises sequences encoding multiple codon mutations of the DKK1 binding domain. Exemplary numbers of codons for mutation include, but are not limited to, at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 225, 250, 275, 300, or more than 300 codons.
[0039] The methods described herein provide for the synthesis of libraries comprising nucleic acids encoding DKK1 binding domains, the libraries comprising sequences encoding length variations of the DKK1 binding domains. In some examples, the libraries comprise sequences encoding length variations that are at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 225, 250, 275, 300, or more than 300 codons less than a given reference sequence. In some examples, the library includes sequences encoding mutations that are at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 275, 300, or more than 300 codons long compared to a given reference sequence.
[0040] Provided herein are DKK1 variant immunoglobulins (e.g., antibodies, VHHs) that include nucleic acids encoding immunoglobulins that include a DKK1 binding domain, including variations in domain type, domain length, or residue mutations. In some examples, the domain is a region of an immunoglobulin that includes a DKK1 binding domain. For example, the region is a VH, CDRH1, CDRH2, CDRH3, VL, CDRL1, CDRL2, or CDRL3 domain. In some examples, the domain is a DKK1 binding domain.
[0041] The methods described herein provide for the synthesis of a DKK1 binding library of nucleic acids, each encoding a predetermined variant of at least one predetermined reference nucleic acid sequence. In some cases, the predetermined reference sequence is a nucleic acid sequence encoding a protein, and the variant library includes sequences encoding at least a single codon mutation, such that multiple different variants of a single residue in the subsequent protein encoded by the synthesized nucleic acid are created by standard translation processes. In some examples, the DKK1 binding library includes diverse nucleic acids that collectively encode mutations at multiple positions. In some examples, the variant library includes sequences encoding at least one codon mutation of the VH, CDRH1, CDRH2, CDRH3, VL, CDRL1, CDRL2, or CDRL3 domain. In some examples, the variant library includes sequences encoding at least one codon mutation of the DKK1 binding domain. In some examples, the variant library includes sequences encoding multiple codon mutations of the VH, CDRH1, CDRH2, CDRH3, VL, CDRL1, CDRL2, or CDRL3 domain. In some examples, the variant library comprises sequences encoding mutations of multiple codons of the DKK1 binding domain. Exemplary numbers of codons for mutation include, but are not limited to, at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 225, 250, 275, 300, or more than 300 codons.
[0042] The methods described herein provide for the synthesis of a DKK1 binding library of nucleic acids, each encoding a predetermined variant of at least one predetermined reference nucleic acid sequence, the DKK1 binding library comprising sequences encoding domain length variations. In some examples, the domain is a VH, CDRH1, CDRH2, CDRH3, VL, CDRL1, CDRL2, or CDRL3 domain. In some examples, the domain is a DKK1 binding domain. In some examples, the library comprises sequences encoding length variations that are at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 225, 250, 275, 300, or more than 300 codons shorter than the predetermined reference sequence. In some examples, the library includes sequences encoding mutations that are at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 275, 300, or more than 300 codons long compared to a given reference sequence.
[0043] Provided herein are DKK1 variant immunoglobulins (e.g., antibodies, VHHs) that include nucleic acids encoding immunoglobulins that include a DKK1 binding domain, and DKK1 binding libraries are synthesized with a varying number of fragments. In some examples, the fragments include a VH, CDRH1, CDRH2, CDRH3, VL, CDRL1, CDRL2, or CDRL3 domain. In some examples, the DKK1 variant immunoglobulins (e.g., antibodies, VHHs) are synthesized with at least or about 2 fragments, 3 fragments, 4 fragments, 5 fragments, or more than 5 fragments. The length of each of the nucleic acid fragments or the average length of the synthesized nucleic acids can be at least about 50, 75, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 525, 550, 575, 600, or more than 600 base pairs. In some examples, the length is about 50-600, 75-575, 100-550, 125-525, 150-500, 175-475, 200-450, 225-425, 250-400, 275-375, or 300-350 base pairs.
[0044] DKK1 variant immunoglobulins (e.g., antibodies, VHHs), including nucleic acids encoding immunoglobulins comprising a DKK1 binding domain described herein, contain amino acids of various lengths when translated. In some examples, the length or average length of each of the amino acid fragments of the synthesized amino acids can be at least or about 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, or more than 150 amino acids. In some examples, the length of the amino acids is about 15 to 150, 20 to 145, 25 to 140, 30 to 135, 35 to 130, 40 to 125, 45 to 120, 50 to 115, 55 to 110, 60 to 110, 65 to 105, 70 to 100, or 75 to 95 amino acids. In some examples, the length of the amino acids is about 22 to about 75 amino acids.
[0045] The DKK1 variant immunoglobulin (e.g., antibody, VHH) comprising de novo synthesized variant sequences encoding an immunoglobulin comprising a DKK1 binding domain comprises multiple variant sequences. In some examples, multiple variant sequences are de novo synthesized for CDRH1, CDRH2, CDRH3, CDRL1, CDRL2, CDRL3, VL, VH, or combinations thereof. In some examples, multiple variant sequences are de novo synthesized for framework element 1 (FW1), framework element 2 (FW2), framework element 3 (FW3), or framework element 4 (FW4). In some examples, multiple variant sequences are de novo synthesized for a GPCR binding domain. For example, the number of variant sequences can range from about 1 to about 10 sequences for a VH domain, about 10 for a DKK1 binding domain, or about 10 for a DKK1 binding domain. 8 For VL domains, the number of variant sequences is about 1 to about 44 sequences. The number of variant sequences can be at least or about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, or more than 500 sequences. In some examples, the number of variant sequences is about 10 to 300, 25 to 275, 50 to 250, 75 to 225, 100 to 200, or 125 to 150 sequences.
[0046] Described herein are antibodies or antibody fragments thereof that bind to DKK1. In some embodiments, the antibody or antibody fragment thereof comprises a sequence set forth in Tables 4-8. In some embodiments, the antibody or antibody fragment thereof comprises a sequence that is at least about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to a sequence set forth in Tables 4-8.
[0047] In some examples, the antibodies or antibody fragments described herein comprise a CDRH1 sequence of any one of SEQ ID NOs: 1-98 or 919-1332. In some examples, the antibodies or antibody fragments described herein comprise a sequence that is at least 80% identical to a CDRH1 sequence of any one of SEQ ID NOs: 1-98 or 919-1332. In some examples, the antibodies or antibody fragments described herein comprise a sequence that is at least 85% identical to a CDRH1 sequence of any one of SEQ ID NOs: 1-98 or 919-1332. In some examples, the antibodies or antibody fragments described herein comprise a sequence that is at least 90% identical to a CDRH1 sequence of any one of SEQ ID NOs: 1-98 or 919-1332. In some examples, the antibodies or antibody fragments described herein comprise a sequence that is at least 95% identical to a CDRH1 sequence of any one of SEQ ID NOs: 1-98 or 919-1332.
[0048] In some examples, the antibodies or antibody fragments described herein comprise a CDRH2 sequence of any one of SEQ ID NOs: 99-196 or 1333-1746. In some examples, the antibodies or antibody fragments described herein comprise a sequence that is at least 80% identical to a CDRH2 sequence of any one of SEQ ID NOs: 99-196 or 1333-1746. In some examples, the antibodies or antibody fragments described herein comprise a sequence that is at least 85% identical to a CDRH2 sequence of any one of SEQ ID NOs: 99-196 or 1333-1746. In some examples, the antibodies or antibody fragments described herein comprise a sequence that is at least 90% identical to a CDRH2 sequence of any one of SEQ ID NOs: 99-196 or 1333-1746. In some examples, the antibodies or antibody fragments described herein comprise a sequence that is at least 95% identical to a CDRH2 sequence of any one of SEQ ID NOs: 99-196 or 1333-1746.
[0049] In some examples, the antibodies or antibody fragments described herein comprise a CDRH3 sequence of any one of SEQ ID NOs: 197-294 or 1747-2160. In some examples, the antibodies or antibody fragments described herein comprise a sequence that is at least 80% identical to a CDRH3 sequence of any one of SEQ ID NOs: 197-294 or 1747-2160. In some examples, the antibodies or antibody fragments described herein comprise a sequence that is at least 85% identical to a CDRH3 sequence of any one of SEQ ID NOs: 197-294 or 1747-2160. In some examples, the antibodies or antibody fragments described herein comprise a sequence that is at least 85% identical to a CDRH3 sequence of any one of SEQ ID NOs: 197-294 or 1747-2160. In some examples, the antibodies or antibody fragments described herein comprise a sequence that is at least 95% identical to a CDRH3 sequence of any one of SEQ ID NOs: 197-294 or 1747-2160.
[0050] In some examples, the antibodies or antibody fragments described herein comprise a CDRL1 sequence of any one of SEQ ID NOs: 2259-2464. In some examples, the antibodies or antibody fragments described herein comprise a sequence that is at least 80% identical to a CDRL1 sequence of any one of SEQ ID NOs: 2259-2464. In some examples, the antibodies or antibody fragments described herein comprise a sequence that is at least 85% identical to a CDRL1 sequence of any one of SEQ ID NOs: 2259-2464. In some examples, the antibodies or antibody fragments described herein comprise a sequence that is at least 90% identical to a CDRL1 sequence of any one of SEQ ID NOs: 2259-2464. In some examples, the antibodies or antibody fragments described herein comprise a sequence that is at least 95% identical to a CDRL1 sequence of any one of SEQ ID NOs: 2259-2464.
[0051] In some examples, the antibodies or antibody fragments described herein comprise a CDR2 sequence of any one of SEQ ID NOs: 2465-2521. In some examples, the antibodies or antibody fragments described herein comprise a sequence that is at least 80% identical to the CDRL2 sequence of any one of SEQ ID NOs: 2465-2521. In some examples, the antibodies or antibody fragments described herein comprise a sequence that is at least 85% identical to the CDRL2 sequence of any one of SEQ ID NOs: 2465-2521. In some examples, the antibodies or antibody fragments described herein comprise a sequence that is at least 90% identical to the CDRL2 sequence of any one of SEQ ID NOs: 2465-2521. In some examples, the antibodies or antibody fragments described herein comprise a sequence that is at least 95% identical to the CDRL2 sequence of any one of SEQ ID NOs: 2465-2521.
[0052] In some examples, the antibodies or antibody fragments described herein comprise a CDRL3 sequence of any one of SEQ ID NOs: 2522-2727. In some examples, the antibodies or antibody fragments described herein comprise a sequence that is at least 80% identical to a CDRL3 sequence of any one of SEQ ID NOs: 2522-2727. In some examples, the antibodies or antibody fragments described herein comprise a sequence that is at least 85% identical to a CDRL3 sequence of any one of SEQ ID NOs: 2522-2727. In some examples, the antibodies or antibody fragments described herein comprise a sequence that is at least 90% identical to a CDRL3 sequence of any one of SEQ ID NOs: 2522-2727. In some examples, the antibodies or antibody fragments described herein comprise a sequence that is at least 95% identical to a CDRL3 sequence of any one of SEQ ID NOs: 2522-2727.
[0053] Described herein, in some embodiments, is an antibody or antibody fragment comprising a variable domain heavy chain region (VH) and a variable domain light chain region (VL), wherein the VH comprises an amino acid sequence at least about 90% identical to a sequence set forth in any one of SEQ ID NOs: 295-392, 394-712, or 2164-2258, and the VL comprises an amino acid sequence at least about 90% identical to a sequence set forth in any one of SEQ ID NOs: 713-918. In some examples, the antibody or antibody fragment comprises a VH that comprises at least or about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one of SEQ ID NOs: 295-392, 394-712, or 2164-2258. In some examples, the antibody or antibody fragment comprises a VL that comprises at least or about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one of SEQ ID NOs:713-918.
[0054] The term "sequence identity" means that two polynucleotide sequences are identical (i.e., nucleotide-by-nucleotide) over a comparison window. The term "percentage of sequence identity" is calculated by comparing two optimally aligned sequences over a comparison window, determining the number of positions where identical nucleic acid bases (e.g., A, T, C, G, U, or I) occur in both sequences to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window (i.e., the window size), and multiplying the result by 100 to calculate the percentage of sequence identity. Alignment to determine percent amino acid sequence identity can be accomplished in a variety of ways within the skill of the art, for example, using publicly available computer software such as EMBOSS MATCHER, EMBOSS WATER, EMBOSS STRETCHER, EMBOSS NEEDLE, EMBOSS LALIGN, BLAST, BLAST-2, ALIGN, or Megalign (DNASTAR) software. Those skilled in the art can determine appropriate parameters for measuring alignment, including any algorithms needed to achieve maximal alignment over the full length of the sequences being compared.
[0055] In the context of using ALIGN-2 to compare amino acid sequences, the % amino acid sequence identity between or to a given amino acid sequence A and a given amino acid sequence B (which may also be expressed as a given amino acid sequence A having or containing a particular % amino acid sequence identity with or to a given amino acid sequence B) is calculated as follows: 100 x fraction X / Y, where X is the number of amino acid residues scored as identical matches by the sequence alignment program ALIGN-2 in the alignment of A and B in the program, and Y is the total number of amino acid residues in B. It should be understood that if the length of amino acid sequence A is not equal to the length of amino acid sequence B, then the % amino acid sequence identity between A and B is not equal to the % amino acid sequence identity between B and A. Unless otherwise stated, all % amino acid sequence identity values used herein are obtained using the ALIGN-2 computer program as described in the immediately preceding paragraph.
[0056] The term "homology" or "similarity" between two proteins is determined by comparing the amino acid sequence and its conserved amino acid substitutes of one protein sequence with a second protein sequence. Similarity can be determined by procedures well known in the art, such as, for example, the BLAST program (Basic Local Alignment Search Tool at the National Center for Biological Information).
[0057] The terms "complementarity determining region" and "CDR" are synonymous with "hypervariable region" or "HVR" and are known in the art to refer to non-contiguous sequences of amino acids in an antibody variable region that confer antigen specificity and / or binding affinity. Generally, there are three CDRs in each heavy chain variable region (CDRH1, CDRH2, CDRH3) and three CDRs in each light chain variable region (CDRL1, CDRL2, CDRL3). The terms "framework region" and "FR" are known in the art to refer to the non-CDR portions of the heavy and light chain variable regions. Generally, there are four FRs in each full-length heavy chain variable region (FR-H1, FR-H2, FR-H3, and FR-H4) and four FRs in each full-length light chain variable region (FR-L1, FR-L2, FR-L3, and FR-L4). The precise amino acid sequence boundaries of a given CDR or FR can be determined by Kabat et al. (1991), “Sequences of Proteins of Immunological Interest,” 5th Ed. Public Health Service, National Institutes of Health, Bethesda, MD ("Kabat" numbering scheme); Al-Lazikani et al., (1997) JMB273, 927-948 ("Chothia" numbering scheme); MacCallum et al., J. Mol. Biol. 262:732-745 (1996), “Antibody-antigen interactions: Contact analysis and binding site topography,” J. Mol. Biol. 262, 732-745.” ("Contact" numbering scheme); Lefranc MP et al.,“IMGT unique numbering for immunoglobulin and T cell receptor variable domains and Ig superfamily V-like domains,”Dev Comp Immunol,2003Jan;27(1):55-77(“IMGT”numbering scheme);Honegger A and Pluckthun A,“Yet another numbering scheme for immunoglobulin variable domains:an automatic modeling and analysis tool,”J Mol Biol,2001Jun8;309(3):657-70,(“Aho”numbering scheme); and Whitelegg NR and Rees AR,“WAM:an improved algorithm for modeling antibodies on the WEB,” Protein Eng. 2000Dec;13(12):819-24(“AbM”numbering The CDRs of the antibodies described herein can be readily determined using any of a number of well-known schemes, including those described by the method described herein. In certain embodiments, the CDRs of the antibodies described herein can be defined by a method selected from Kabat, Chothia, IMGT, Aho, AbM, or a combination thereof.
[0058] The boundaries of a given CDR or FR may differ depending on the scheme used for identification. For example, the Kabat scheme is based on structural alignment, whereas the Chothia scheme is based on structural information. The numbering in both the Kabat and Chothia schemes is based on the length of the most common antibody region sequences, with insertions addressed by an insertion letter (e.g., "30a") and deletions found in some antibodies. The two schemes place certain insertions and deletions ("indels") in different positions, resulting in different numbering. The Contact scheme is based on the analysis of complex crystal structures and is similar in many ways to the Chothia numbering scheme.
[0059] DKK1 variant immunoglobulins (e.g., antibodies, VHHs) that include de novo synthesized variant sequences encoding immunoglobulins that include a DKK1 binding domain include improved diversity. For example, variants are created by placing a DKK1 binding domain variant in an immunoglobulin that includes an N-terminal CDRH3 mutation and a C-terminal CDRH3 mutation. In some examples, the variants include affinity maturation variants. Alternatively or in combination, the variants include variants in other regions of the immunoglobulin, including but not limited to CDRH1 and CDRH2. In some examples, the number of variants in a DKK1 variant immunoglobulin (e.g., antibodies, VHHs) is at least or about 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 16 , 10 17 , 10 18 , 10 19 , 10 20 , or 10 20 There are more than two non-identical sequences.
[0060] In some examples, at least one region of the antibody for mutation is from a heavy chain V gene family, a heavy chain D gene family, a heavy chain J gene family, a light chain V gene family, or a light chain J gene family. In some examples, the light chain V gene family comprises an immunoglobulin kappa (IGK) gene or an immunoglobulin lambda (IGL) gene. Exemplary regions of an antibody for mutation include, but are not limited to, IGHV1-18, IGHV1-69, IGHV1-8, IGHV3-21, IGHV3-23, IGHV3-30 / 33rn, IGHV3-28, IGHV1-69, IGHV3-74, IGHV4-39, IGHV4-59 / 61, IGKV1-39, IGKV1-9, IGKV2-28, IGKV3-11, IGKV3-15, IGKV3-20, IGKV4-1, IGLV1-51, IGLV2-14, IGLV1-40, and IGLV3-1. In some examples, the gene is IGHV1-69, IGHV3-30, IGHV3-23, IGHV3, IGHV1-46, IGHV3-7, IGHV1, or IGHV1-8. In some examples, the gene is IGHV1-69 and IGHV3-30. In some examples, the region of the antibody for mutation is IGHJ3, IGHJ6, IGHJ, IGHJ4, IGHJ5, IGHJ2, or IGH1. In some examples, the region of the antibody for mutation is IGHJ3, IGHJ6, IGHJ, or IGHJ4. In some examples, at least one region of the antibody for mutation is IGHV1-69, IGHV3-23, IGKV3-20, IGKV1-39, or a combination thereof. In some examples, at least one region of the antibody for mutation is IGHV1-69 and IGKV3-20, and in some examples, at least one region of the antibody for mutation is IGHV1-69 and IGKV1-39. In some examples, at least one region of the antibody for mutation is IGHV3-23 and IGKV3-20. In some examples, at least one region of the antibody for mutation is IGHV3-23 and IGKV1-39.
[0061] Provided herein is a library that includes nucleic acids encoding DKK1 antibodies that include mutations in at least one region of the antibody, which region is a CDR region. In some examples, the DKK1 antibody is a single domain antibody that includes one heavy chain variable domain, such as a VHH antibody. In some examples, the VHH antibody includes mutations in one or more CDR regions. In some examples, the library described herein includes at least or about 1, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1200, 1400, 1600, 1800, 2000, 2400, 2600, 2800, 3000, or more than 3000 CDR1, CDR2, or CDR3 sequences. In some examples, the library described herein includes at least or about 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 16 , 10 17 , 10 18 , 10 19 , 10 20 , or 10 20 The library comprises more than one CDR1, CDR2, or CDR3 sequence. For example, the library comprises at least 2000 sequences of CDR1, at least 1200 sequences of CDR2, and at least 1600 sequences of CDR3. In some examples, each sequence is non-identical.
[0062] In some examples, the CDR1, CDR2, or CDR3 is a variable domain light chain (VL). The CDR1, CDR2, or CDR3 of a variable domain light chain (VL) may be referred to as CDRL1, CDRL2, or CDRL3, respectively. In some examples, the libraries described herein include at least or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1200, 1400, 1600, 1800, 2000, 2400, 2600, 2800, 3000, or more than 3000 VL CDR1, CDR2, or CDR3 sequences. In some examples, the libraries described herein comprise at least or about 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 16 , 10 17 , 10 18 , 10 19 , 10 20 , or 10 20 The library comprises more than 100 sequences of VL CDR1, CDR2, or CDR3. For example, the library comprises at least 20 sequences of VL CDR1, at least 4 sequences of VL CDR2, and at least 140 sequences of VL CDR3. In some examples, the library comprises at least 2 sequences of VL CDR1, at least 1 sequence of VL CDR2, and at least 3000 sequences of VL CDR3. In some examples, the VL is IGKV1-39, IGKV1-9, IGKV2-28, IGKV3-11, IGKV3-15, IGKV3-20, IGKV4-1, IGLV1-51, IGLV2-14, IGLV1-40, or IGLV3-1. In some examples, the VL is IGKV2-28. In some examples, the VL is IGLV1-51.
[0063] In some examples, the CDR1, CDR2, or CDR3 is a variable domain heavy chain (VH). The CDR1, CDR2, or CDR3 of a variable domain heavy chain (VH) may be referred to as CDRH1, CDRH2, or CDRH3, respectively. In some examples, the libraries described herein include at least or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1200, 1400, 1600, 1800, 2000, 2400, 2600, 2800, 3000, or more than 3000 VH CDR1, CDR2, or CDR3 sequences. In some examples, the libraries described herein comprise at least or about 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 16 , 10 17 , 10 18 , 10 19 , 10 20 , or 10 20 For example, the library may contain at least 30 sequences of VH CDR1, at least 570 sequences of VH CDR2, and at least 10 sequences of VH CDR3. 8 In some examples, the library includes at least 30 sequences of CDR1 of VH, at least 860 sequences of CDR2 of VH, and at least 10 sequences of CDR3 of VH. 7In some examples, the VH is IGHV1-18, IGHV1-69, IGHV1-8, IGHV3-21, IGHV3-23, IGHV3-30 / 33rn, IGHV3-28, IGHV3-74, IGHV4-39, or IGHV4-59 / 61. In some examples, the VH is IGHV1-69, IGHV3-30, IGHV3-23, IGHV3, IGHV1-46, IGHV3-7, IGHV1, or IGHV1-8. In some examples, the VH is IGHV1-69 or IGHV3-30. In some examples, the VH is IGHV3-23.
[0064] In some embodiments, the libraries described herein comprise CDRL1, CDRL2, CDRL3, CDRH1, CDRH2, or CDRH3 of varying lengths. In some examples, the length of CDRL1, CDRL2, CDRL3, CDRH1, CDRH2, or CDRH3 comprises a length of at least or about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, or more than 90 amino acids. For example, CDRH3 comprises a length of at least or about 12, 15, 16, 17, 20, 21, or 23 amino acids. In some examples, CDRL1, CDRL2, CDRL3, CDRH1, CDRH2, or CDRH3 comprises a range of lengths from about 1 to about 10, about 5 to about 15, about 10 to about 20, or about 15 to about 30 amino acids.
[0065] Libraries containing nucleic acids encoding antibodies with variant CDR sequences described herein contain a range of amino acid lengths when translated. In some examples, the length or average length of each of the synthesized amino acid fragments can be at least or about 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, or more than 150 amino acids. In some examples, the length is about 15-150, 20-145, 25-140, 30-135, 35-130, 40-125, 45-120, 50-115, 55-110, 60-110, 65-105, 70-100, or 75-95 amino acids. In some examples, the length is about 22 amino acids to about 75 amino acids. In some examples, the antibody comprises at least or about 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, or more than 5000 amino acids.
[0066] The length ratios of CDRL1, CDRL2, CDRL3, CDRH1, CDRH2, or CDRH3 may vary in the libraries described herein. In some examples, CDRL1, CDRL2, CDRL3, CDRH1, CDRH2, or CDRH3 that comprise at least about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, or more than 90 amino acids in length comprise about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or more than 90% of the library. For example, CDRH3s containing about 23 amino acids in length are present in 40% of the libraries, CDRH3s containing about 21 amino acids in length are present in 30% of the libraries, CDRH3s containing about 17 amino acids in length are present in 20% of the libraries, and CDRH3s containing about 12 amino acids in length are present in 10% of the libraries. In some examples, CDRH3s containing about 20 amino acids in length are present in 40% of the libraries, CDRH3s containing about 16 amino acids in length are present in 30% of the libraries, CDRH3s containing about 15 amino acids in length are present in 20% of the libraries, and CDRH3s containing about 12 amino acids in length are present in 10% of the libraries.
[0067] The libraries described herein encoding VHH antibodies include at least or about 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 16 , 10 17 , 10 18 , 10 19 , 10 20 , or 10 20 In some examples, the library comprises variant CDR sequences that have been shuffled to create a library with a theoretical diversity of sequences of at least or about 10 7 , 10 8 , 10 9 , 10 10, 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 16 , 10 17 , 10 18 , 10 19 , 10 20 , or 10 20 The final library diversity exceeds 100 sequences.
[0068] Provided herein are DKK1 variant immunoglobulins (e.g., antibodies, VHHs) encoding immunoglobulins. In some examples, the DKK1 immunoglobulin is an antibody. In some examples, the DKK1 immunoglobulin is a VHH antibody. In some examples, the DKK1 immunoglobulin comprises a binding affinity (e.g., kD) for DKK1 of less than 1 nM, less than 1.2 nM, less than 2 nM, less than 5 nM, less than 10 nM, less than 11 nM, less than 13.5 nM, less than 15 nM, less than 20 nM, less than 25 nM, or less than 30 nM. In some examples, the DKK1 immunoglobulin comprises a kD of less than 1 nM. In some examples, the DKK1 immunoglobulin comprises a kD of less than 1.2 nM. In some examples, the DKK1 immunoglobulin comprises a kD of less than 2 nM. In some examples, the DKK1 immunoglobulin comprises a kD of less than 5 nM. In some examples, the DKK1 immunoglobulin comprises a kD of less than 10 nM. In some examples, the DKK1 immunoglobulin comprises a kD of less than 13.5 nM. In some examples, the DKK1 immunoglobulin comprises a kD of less than 15 nM. In some examples, the DKK1 immunoglobulin comprises a kD of less than 20 nM. In some examples, the DKK1 immunoglobulin comprises a kD of less than 25 nM. In some examples, the DKK1 immunoglobulin comprises a kD of less than 30 nM.
[0069] Provided herein are DKK1 variant immunoglobulins (e.g., antibodies, VHHs) encoding immunoglobulins, wherein the immunoglobulin comprises a long half-life. In some examples, the half-life of the DKK1 immunoglobulin is at least or about 12 hours, 24 hours, 36 hours, 48 hours, 60 hours, 72 hours, 84 hours, 96 hours, 108 hours, 120 hours, 140 hours, 160 hours, 180 hours, 200 hours, or more than 200 hours. In some examples, the half-life of the DKK1 immunoglobulin ranges from about 12 hours to about 300 hours, from about 20 hours to about 280 hours, from about 40 hours to about 240 hours, or from about 60 hours to about 200 hours.
[0070] The DKK1 immunoglobulins described herein may include improved properties. In some examples, the DKK1 immunoglobulin is monomeric. In some examples, the DKK1 immunoglobulin is less prone to aggregation. In some examples, at least or about 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the DKK1 immunoglobulin is monomeric. In some examples, the DKK1 immunoglobulin is heat stable. In some examples, the DKK1 immunoglobulin exhibits reduced non-specific binding.
[0071] After synthesizing DKK1 variant immunoglobulins (e.g., antibodies, VHHs) that contain nucleic acids encoding immunoglobulins that contain DKK1 binding domains, the libraries can be used for screening and analysis. For example, the libraries are assayed for displayability and panning of the libraries. In some examples, displayability is assayed using a selectable tag. Exemplary tags include, but are not limited to, radioactive labels, fluorescent labels, enzymes, chemiluminescent tags, colorimetric tags, affinity tags, or other labels or tags known in the art. In some examples, the tag is histidine, polyhistidine, myc, hemagglutinin (HA), or FLAG. In some examples, the DKK1 variant immunoglobulins (e.g., antibodies, VHHs) contain multiple tags, such as GFP, FLAG, and Lucy, as well as nucleic acids encoding immunoglobulins with DNA barcodes. In some examples, libraries are assayed by sequencing using a variety of methods, including, but not limited to, single molecule real-time (SMRT) sequencing, polony sequencing, sequencing by ligation, reversible terminator sequencing, proton detection sequencing, ion semiconductor sequencing, nanopore sequencing, electronic sequencing, pyrosequencing, Maxam-Gilbert sequencing, chain termination (e.g., Sanger) sequencing, +S sequencing, or sequencing by synthesis.
[0072] Expression system Provided herein are libraries that include nucleic acids encoding immunoglobulins that include a DKK1-binding domain, the libraries having improved specificity, stability, expression, folding, or downstream activity. In some examples, the libraries described herein are used for screening and analysis.
[0073] Provided herein is a library comprising nucleic acids encoding immunoglobulins comprising a DKK1 binding domain, which nucleic acid library is used for screening and analysis. In some examples, the screening and analysis includes in vitro, in vivo, or ex vivo assays. The cells for screening include primary cells or cell lines taken from a living subject. The cells can be derived from prokaryotes (e.g., bacteria and fungi) or eukaryotes (e.g., animals and plants). Exemplary animal cells include, but are not limited to, those derived from mice, rabbits, primates, and insects. In some examples, the cells for screening include, but are not limited to, cell lines including Chinese hamster ovary (CHO) cell lines, human embryonic kidney (HEK) cell lines, or baby hamster kidney (BHK) cell lines. In some examples, the nucleic acid libraries described herein can also be delivered to multicellular organisms. Exemplary multicellular organisms include, but are not limited to, plants, mice, rabbits, primates, and insects.
[0074] The nucleic acid libraries described herein or the protein libraries encoded thereby can be screened for various pharmacological or pharmacokinetic properties. In some examples, the libraries are screened using in vitro, in vivo, or ex vivo assays. For example, in vitro pharmacological or pharmacokinetic properties that are screened include, but are not limited to, binding affinity, binding specificity, and binding activity. Exemplary in vivo pharmacological or pharmacokinetic properties of the libraries described herein that are screened include, but are not limited to, therapeutic efficacy, activity, preclinical toxicity profile, clinical efficacy profile, clinical toxicity profile, immunogenicity, potency, and clinical safety profile.
[0075] Pharmacological or pharmacokinetic properties that can be screened include, but are not limited to, cell binding affinity and cell activity.For example, cell binding affinity assay or cell activity assay is carried out to determine the agonistic effect, antagonistic effect, or allosteric effect of the library described herein.In some examples, the library described herein is compared with the cell binding or cell activity of a ligand of DKK1.
[0076] The libraries described herein can be screened in cell-based or non-cell-based assays, including, but not limited to, the use of viral particles, in vitro translated proteins, and proteoliposomes with DKK1.
[0077] The nucleic acid library described herein can be screened by sequencing. In some examples, next generation sequencing is used to determine the sequence enrichment of DKK1 binding variants. In some examples, V gene distribution, J gene distribution, V gene family, CDR3 number per length, or a combination thereof is determined. In some examples, clonal frequency, clonal accumulation, lineage accumulation, or a combination thereof is determined. In some examples, the number of sequences, sequences containing VH clones, clones, clones greater than 1, clonotypes, clonotypes greater than 1, lineages, Simpsons, or a combination thereof is determined. In some examples, the percentage of non-identical CDR3 is determined. For example, the percentage of non-identical CDR3 is calculated by dividing the number of non-identical CDR3s of a sample by the total number of sequences with CDR3 in the sample.
[0078] Provided herein is a nucleic acid library, which may be expressed in a vector. Expression vectors for inserting the nucleic acid library disclosed herein may include eukaryotic or prokaryotic expression vectors. Exemplary expression vectors include mammalian expression vectors: pSF-CMV-NEO-NH2-PPT-3XFLAG, pSF-CMV-NEO-COOH-3XFLAG, pSF-CMV-PURO-NH2-GST-TEV, pSF-OXB20-COOH-TEV-FLAG(R)-6His, pCEP4 Examples of vectors include, but are not limited to, pDEST27, pSF-CMV-Ub-KrYFP, pSF-CMV-FMDV-daGFP, pEF1a-mCherry-N1 vector, pEF1a-tdTomato vector, pSF-CMV-FMDV-Hygro, pSF-CMV-PGK-Puro, pMCP-tag(m), and pSF-CMV-PURO-NH2-CMYC; bacterial expression vectors: pSF-OXB20-BetaGal, pSF-OXB20-Fluc, pSF-OXB20, and pSF-Tac; plant expression vectors: pRI101-AN DNA and pCambia2301; and yeast expression vectors: pTYB21 and pKLAC2; and insect vectors: pAc5.1 / V5-His A and pDEST8. In some examples, the vector is pcDNA3 or pcDNA3.1.
[0079] The present specification describes a nucleic acid library expressed in a vector to generate a construct comprising an immunoglobulin comprising a sequence of a DKK1 binding domain.In some examples, the size of the construct varies.In some examples, the construct comprises at least or about 500, 600, 700, 800, 900, 1000, 1100, 1300, 1400, 1500, 1600, 1700, 1800, 2000, 2400, 2600, 2800, 3000, 3200, 3400, 3600, 3800, 4000, 4200, 4400, 4600, 4800, 5000, 6000, 7000, 8000, 9000, 10000, or more than 10000 bases. In some examples, the construct may be about 300-1,000, 300-2,000, 300-3,000, 300-4,000, 300-5,000, 300-6,000, 300-7,000, 300-8,000, 300-9,000, 300-10,000, 1,000-2,000, 1,000-3,000, 1,000-4,000, 1,000-5,000, 1 ,000~6,000, 1,000~7,000, 1,000~8,000, 1,000~9,000, 1,000~10,000, 2,000~3,000, 2,000~4,000, 2,000~5,000, 2,000~6,000, 2,000~7,000, 2,000~8,000, 2,000~9,000, 2,000~10,000, 3,000~4,000, 3 ,000~5,000, 3,000~6,000, 3,000~7,000, 3,000~8,000, 3,000~9,000, 3,000~10,000, 4,000~5,000, 4,000~6,000, 4,000~7,000, 4,000~8,000, 4,000~9,000, 4,000~10,000, 5,000~6,000, 5,000~7,000, 5 The ranges include, for example, 5,000 to 8,000, 5,000 to 9,000, 5,000 to 10,000, 6,000 to 7,000, 6,000 to 8,000, 6,000 to 9,000, 6,000 to 10,000, 7,000 to 8,000, 7,000 to 9,000, 7,000 to 10,000, 8,000 to 9,000, 8,000 to 10,000, or 9,000 to 10,000 bases.
[0080] Provided herein is a library comprising nucleic acids encoding immunoglobulins, the nucleic acid library being expressed in a cell. In some examples, the library is synthesized to express a reporter gene. Exemplary reporter genes include, but are not limited to, acetohydroxyacid synthase (AHAS), alkaline phosphatase (AP), β-galactosidase (LacZ), β-glucoronidase (GUS), chloramphenicol acetyltransferase (CAT), green fluorescent protein (GFP), red fluorescent protein (RFP), yellow fluorescent protein (YFP), cyan fluorescent protein (CFP), cerulean fluorescent protein, citrine fluorescent protein, orange fluorescent protein, cherry fluorescent protein, turquoise fluorescent protein, blue fluorescent protein, horseradish peroxidase (HRP), luciferase (Luc), nopaline synthase (NOS), octopine synthase (OCS), luciferase, and derivatives thereof. Methods for determining regulation of reporter genes are well known in the art and include, but are not limited to, fluorometric methods (e.g., fluorescence spectroscopy, fluorescence activated cell sorting (FACS), fluorescence microscopy), and antibiotic resistance determination.
[0081] Diseases and Disorders Provided herein are DKK1 variant immunoglobulins (e.g., antibodies, VHHs), including nucleic acids encoding immunoglobulins (e.g., antibodies) that include a DKK1 binding domain that can have a therapeutic effect. In some examples, the DKK1 variant immunoglobulins (e.g., antibodies, VHHs) are translated into proteins used to treat a disease or disorder. In some examples, the protein is an immunoglobulin. In some examples, the protein is a peptidomimetic.
[0082] Exemplary diseases include, but are not limited to, cancer (e.g., gastro-esophageal cancer, endometrial cancer, ovarian cancer, prostate cancer, liver cancer, etc.), inflammatory diseases or disorders, metabolic diseases or disorders, cardiovascular diseases or disorders, respiratory diseases or disorders, pain, digestive diseases or disorders, reproductive diseases or disorders, endocrine diseases or disorders, or nervous diseases or disorders. In some examples, the cancer is a solid cancer or a blood cancer. In some examples, the modulators of DKK1 described herein are used to treat weight gain (or induce weight loss), treat obesity, or treat type II diabetes. In some examples, the DKK1 modulators are used to treat hypoglycemia. In some examples, the DKK1 modulators are used to treat hypoglycemia following obesity. In some examples, the DKK1 modulators are used to treat severe hypoglycemia. In some examples, the DKK1 modulators are used to treat hyperinsulinism. In some examples, the DKK1 modulators are used to treat congenital hyperinsulinism.
[0083] DKK1 can be tumorigenic in cancer. DKK1 can also be immunosuppressive (e.g., via myeloid-derived suppressor cells (MDSC) or natural killer (NK) cells). DKK1 can cause immunosuppression through inactivation of T cells, accumulation of MDSC, or clearance of NK cells. DKK1 can inhibit Wnt binding to low-density lipoprotein (LDL) receptor-related protein 5 (LRP5). DKK1 can inhibit Wnt binding to LDL receptor-related protein 6 (LRP6). DKK1 can inhibit Wnt binding to the LRP5 / 6 complex. Mutations in Wnt-activating genes can lead to increased DKK1 expression.
[0084] Antagonist mAbs can activate the innate immune response, which has antiangiogenic and direct antitumor effects, and bind and remove DKK1 from the tumor microenvironment. Tumors with Wnt-activating mutations can respond to DKK1 antagonism. For example, high tumor DKK1 can be associated with longer progression-free survival in patients with esophageal and gastric cancer.
[0085] In some examples, the subject is a mammal. In some examples, the subject is a mouse, rabbit, dog, or human. The subject treated by the methods described herein may be an infant, an adult, or a child. Pharmaceutical compositions comprising the antibodies or antibody fragments described herein may be administered intravenously or subcutaneously.
[0086] Described herein are pharmaceutical compositions comprising an antibody or antibody fragment thereof that binds to DKK1. In some embodiments, the antibody or antibody fragment thereof comprises a sequence set forth in Tables 4-8. In some embodiments, the antibody or antibody fragment thereof comprises a sequence that is at least about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to a sequence set forth in Tables 4-8.
[0087] In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a CDRH1 sequence of any one of SEQ ID NOs: 1-98 or 919-1332. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a sequence that is at least 80% identical to a CDRH1 sequence of any one of SEQ ID NOs: 1-98 or 919-1332. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a sequence that is at least 85% identical to a CDRH1 sequence of any one of SEQ ID NOs: 1-98 or 919-1332. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a sequence that is at least 90% identical to a CDRH1 sequence of any one of SEQ ID NOs: 1-98 or 919-1332. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a sequence that is at least 95% identical to a CDRH1 sequence of any one of SEQ ID NOs: 1-98 or 919-1332.
[0088] In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a CDRH2 sequence of any one of SEQ ID NOs: 99-196 or 1333-1746. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a sequence that is at least 80% identical to a CDRH2 sequence of any one of SEQ ID NOs: 99-196 or 1333-1746. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a sequence that is at least 85% identical to a CDRH2 sequence of any one of SEQ ID NOs: 99-196 or 1333-1746. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a sequence that is at least 90% identical to a CDRH2 sequence of any one of SEQ ID NOs: 99-196 or 1333-1746. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a sequence that is at least 95% identical to a CDRH2 sequence of any one of SEQ ID NOs: 99-196 or 1333-1746.
[0089] In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a CDRH3 sequence of any one of SEQ ID NOs: 197-294 or 1747-2160. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a sequence that is at least 80% identical to a CDRH3 sequence of any one of SEQ ID NOs: 197-294 or 1747-2160. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a sequence that is at least 85% identical to a CDRH3 sequence of any one of SEQ ID NOs: 197-294 or 1747-2160. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a sequence that is at least 90% identical to a CDRH3 sequence of any one of SEQ ID NOs: 197-294 or 1747-2160. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a sequence that is at least 95% identical to a CDRH3 sequence of any one of SEQ ID NOs: 197-294 or 1747-2160.
[0090] In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a CDRL1 sequence of any one of SEQ ID NOs: 2259-2464. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a sequence that is at least 80% identical to a CDRL1 sequence of any one of SEQ ID NOs: 2259-2464. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a sequence that is at least 85% identical to a CDRL1 sequence of any one of SEQ ID NOs: 2259-2464. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a sequence that is at least 90% identical to a CDRL1 sequence of any one of SEQ ID NOs: 2259-2464. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a sequence that is at least 95% identical to a CDRL1 sequence of any one of SEQ ID NOs: 2259-2464.
[0091] In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a CDRL2 sequence of any one of SEQ ID NOs: 2465-2521. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a sequence that is at least 80% identical to a CDRL2 sequence of any one of SEQ ID NOs: 2465-2521. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a sequence that is at least 85% identical to a CDRL2 sequence of any one of SEQ ID NOs: 2465-2521. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a sequence that is at least 90% identical to a CDRL2 sequence of any one of SEQ ID NOs: 2465-2521. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a sequence that is at least 95% identical to a CDRL2 sequence of any one of SEQ ID NOs: 2465-2521.
[0092] In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a CDRL3 sequence of any one of SEQ ID NOs: 2522-2727. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a sequence that is at least 80% identical to a CDRL3 sequence of any one of SEQ ID NOs: 2522-2727. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a sequence that is at least 85% identical to a CDRL3 sequence of any one of SEQ ID NOs: 2522-2727. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a sequence that is at least 90% identical to a CDRL3 sequence of any one of SEQ ID NOs: 2522-2727. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a sequence that is at least 95% identical to a CDRL3 sequence of any one of SEQ ID NOs: 2522-2727.
[0093] In some embodiments, the antibody or antibody fragment comprises a variable domain heavy chain region (VH) and a variable domain light chain region (VL), wherein the VH comprises complementarity determining regions CDRH1, CDRH2, and CDRH3, and the VL comprises complementarity determining regions CDRL1, CDRL2, and CDRL3, and wherein (a) the amino acid sequence of CDRH1 is as set forth in any one of SEQ ID NOs: 1-98 or 919-1332, (b) the amino acid sequence of CDRH2 is as set forth in any one of SEQ ID NOs: 99-196 or 1333-1746, and (c) the amino acid sequence of CDRH3 is as set forth in any one of SEQ ID NOs: 197-294 or 1747-2160. In some embodiments, the antibody or antibody fragment comprises a variable domain heavy chain region (VH) and a variable domain light chain region (VL), wherein the VH comprises complementarity determining regions CDRH1, CDRH2, and CDRH3, and the VL comprises complementarity determining regions CDRL1, CDRL2, and CDRL3, and wherein (a) the amino acid sequence of CDRH1 is at least or about 80%, 85%, 90%, or 95% identical to any one of SEQ ID NOs: 1-98, (b) the amino acid sequence of CDRH2 is at least or about 80%, 85%, 90%, or 95% identical to any one of SEQ ID NOs: 99-196, and (c) the amino acid sequence of CDRH3 is at least or about 80%, 85%, 90%, or 95% identical to any one of SEQ ID NOs: 197-294.
[0094] In some embodiments, the antibody or antibody fragment comprises a variable domain heavy chain region (VH) and a variable domain light chain region (VL), wherein the VH comprises complementarity determining regions CDRH1, CDRH2, and CDRH3, and the VL comprises complementarity determining regions CDRL1, CDRL2, and CDRL3, and wherein (a) the amino acid sequence of CDRL1 is as set forth in any one of SEQ ID NOs: 2259-2464, (b) the amino acid sequence of CDRL2 is as set forth in any one of SEQ ID NOs: 2465-2521, and (c) the amino acid sequence of CDRL3 is as set forth in any one of SEQ ID NOs: 2522-2727. In some embodiments, the antibody or antibody fragment comprises a variable domain heavy chain region (VH) and a variable domain light chain region (VL), wherein the VH comprises complementarity determining regions CDRH1, CDRH2, and CDRH3, and the VL comprises complementarity determining regions CDRL1, CDRL2, and CDRL3, and wherein (a) the amino acid sequence of CDRH1 is at least or about 80%, 85%, 90%, or 95% identical to any one of SEQ ID NOs: 2259-2464, (b) the amino acid sequence of CDRL2 is at least or about 80%, 85%, 90%, or 95% identical to any one of SEQ ID NOs: 2465-2521, and (c) the amino acid sequence of CDRL3 is at least or about 80%, 85%, 90%, or 95% identical to any one of SEQ ID NOs: 2522-2727.
[0095] Described herein, in some embodiments, is an antibody or antibody fragment comprising a variable domain heavy chain region (VH), wherein the VH comprises an amino acid sequence at least about 90% identical to a sequence set forth in any one of SEQ ID NOs: 295-392, 394-712, or 2164-2258. In some examples, the antibody or antibody fragment comprises a VH that comprises at least or about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one of SEQ ID NOs: 295-392, 394-712, or 2164-2258.
[0096] Described herein, in some embodiments, is an antibody or antibody fragment comprising a variable domain light chain region (VL), wherein the VL comprises an amino acid sequence at least about 90% identical to a sequence set forth in any one of SEQ ID NOs: 713-918. In some examples, the antibody or antibody fragment comprises a VL that comprises at least or about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one of SEQ ID NOs: 713-918.
[0097] Variant Libraries Codon Mutation The variant nucleic acid library described herein may comprise a plurality of nucleic acids, each nucleic acid encoding a variant codon sequence compared to a reference nucleic acid sequence. In some examples, each nucleic acid of the first nucleic acid population contains a variant at a single variant site. In some examples, the first nucleic acid population contains multiple variants at a single variant site, and the first nucleic acid population contains multiple variants at the same variant site. The first nucleic acid population may comprise nucleic acids that collectively encode multiple codon variants at the same variant site. The first nucleic acid population may comprise nucleic acids that collectively encode up to 19 or more codons at the same position. The first nucleic acid population may comprise nucleic acids that collectively encode up to 60 variant triplets at the same position, or the first nucleic acid population may comprise nucleic acids that collectively encode up to 61 different triplets of a codon at the same position. Each variant may encode a codon that results in a different amino acid during translation.
[0098] A nucleic acid population may include a variety of nucleic acids that collectively encode up to 20 codon mutations at multiple positions. In such cases, each nucleic acid in the population includes codon mutations at multiple positions within the same nucleic acid. In some examples, each nucleic acid in the population includes codon mutations at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more codons within a single nucleic acid. In some examples, each variant length nucleic acid includes codon mutations at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more codons within a single length nucleic acid. In some examples, the variant nucleic acid population comprises codon mutations at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more codons within a single nucleic acid. In some examples, the variant nucleic acid population comprises codon mutations at at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more codons within a single length nucleic acid.
[0099] Highly parallel nucleic acid synthesis Provided herein is a platform approach that utilizes miniaturization, parallelization, and vertical integration of end-to-end processes from polynucleotide synthesis to gene assembly in nanowells on silicon to create an innovative synthesis platform. The device described herein provides a silicon synthesis platform that can generate up to about 1,000,000 or more polynucleotides or 10,000 or more genes in a single highly parallel run, increasing throughput by up to 1,000 times or more compared to traditional synthesis methods, in the same footprint as a 96-well plate.
[0100] With the advent of next-generation sequencing, high-resolution genomic data has become a key component of research that dissects the biological roles of various genes in both normal biology and the development of disease. At the core of this research is the central dogma of molecular biology and the concept of "sequential information transmission, residue by residue." Genomic information encoded in DNA is transcribed into messages that are then translated into proteins, which are the active products within specific biological pathways.
[0101] Another interesting research area is the discovery, development, and production of therapeutic molecules focused on very specific cellular targets. Highly diverse DNA sequence libraries are at the core of the development pipeline of targeted therapeutics. Gene variants are used to express proteins in a design, build, and test protein engineering cycle that ideally results in genes optimized for high expression of proteins with high affinity for their therapeutic target. As an example, consider the binding pocket of a receptor. All sequence permutations of all residues in the binding pocket can be tested simultaneously, allowing for an exhaustive search and increasing the chances of success. Saturation mutagenesis, in which researchers attempt to create all possible mutations at a specific site in a receptor, is one approach to this development challenge. Although costly, time-consuming, and laborious, it allows each variant to be introduced at every position. In contrast, in combinatorial mutagenesis, a few selected positions or short stretches of DNA can be extensively altered, but in that case biased representation creates an incomplete repertoire of variants.
[0102] To accelerate the drug development pipeline, libraries containing the desired variants available at the intended frequency in the right positions available for testing - i.e., precision libraries - can reduce the cost and time required for screening. Provided herein are methods for synthesizing nucleic acid synthetic variant libraries that provide precise introduction of each intended variant at the desired frequency. For the end user, this translates to the ability to not only exhaustively sample sequence space, but also to efficiently query these hypotheses, reducing costs and screening time. Genome-wide editing can uncover libraries where key pathways, individual variants and sequence permutations can be tested for optimal functionality, reconstructing pathways and entire genomes using thousands of genes to redesign biological systems for drug discovery.
[0103] In the first instance, the drug itself can be optimized using the methods described herein. For example, to improve a particular function of the antibody, a variant polynucleotide library encoding a portion of the antibody is designed and synthesized. A variant nucleic acid library of the antibody can then be created by the processes described herein (e.g., PCR mutagenesis followed by insertion into a vector). The antibody is then expressed in a production cell line and screened for enhanced activity. Examples of screening include testing binding affinity to the antigen, stability, or modulation of effector functions (e.g., ADCC, complement, or apoptosis). Exemplary regions for optimizing an antibody include the Fc region, the Fab region, the variable region of the Fab region, the constant region of the Fab region, the variable domains of the heavy or light chains (Vc, Vdc, Vs ... H Or V L ), and V H Or V L Examples of such complementarity determining regions (CDRs) include, but are not limited to, certain complementarity determining regions (CDRs) of the following amino acids:
[0104] The nucleic acid libraries synthesized by the methods described herein can be expressed in a variety of cells associated with a disease state. Cells associated with a disease state include cell lines, tissue samples, primary cells from a subject, cultured cells grown from a subject, or cells in a model system. Exemplary model systems include, but are not limited to, plant and animal models of a disease state.
[0105] To identify variant molecules associated with the prevention, alleviation or treatment of a disease state, the variant nucleic acid library described herein is expressed in cells associated with the disease state or cells capable of inducing the disease state. In some examples, an agent is used to induce the disease state in the cells. Exemplary tools for inducing a disease state include, but are not limited to, the Cre / Lox recombination system, LPS inflammation induction, and streptozotocin to induce hypoglycemia. The cells associated with the disease state can be cells from a model system or cultured cells as well as cells from a subject with a particular disease state. Exemplary disease states include bacterial, fungal, viral, autoimmune, or proliferative diseases (e.g., cancer). In some examples, the variant nucleic acid library is expressed in a model system, cell line, or primary cells obtained from a subject and screened for changes in at least one cellular activity. Exemplary cellular activities include, but are not limited to, proliferation, cycle progression, cell death, adhesion, migration, regeneration, cell signaling, energy production, oxygen utilization, metabolic activity, and aging, response to free radical damage, or any combination thereof.
[0106] substrate The device used as a surface for polynucleotide synthesis may be in the form of a substrate, including, but not limited to, a homogenous array surface, a patterned array surface, a channel, a bead, a gel, and the like. Provided herein is a substrate comprising a plurality of clusters, each cluster comprising a plurality of loci supporting attachment and synthesis of a polynucleotide. In some examples, the substrate comprises a homogenous array surface. For example, the homogenous array surface is a homogenous plate. As used herein, the term "locus" refers to a discrete region on a structure that provides support for a polynucleotide encoding a single predetermined sequence extending from the surface. In some examples, the locus is on a two-dimensional surface, e.g., a substantially flat surface. In some examples, the locus is on a three-dimensional surface, e.g., a well, a microwell, a channel, or a post. In some examples, the surface of the locus comprises a material that is actively functionalized to attach at least one nucleotide for polynucleotide synthesis, or preferably a population of identical nucleotides for synthesis of a population of polynucleotides. In some examples, the polynucleotide refers to a population of polynucleotides encoding the same nucleic acid sequence. In some examples, the surface of the substrate comprises one or more surfaces of the substrate. Using the systems and methods provided, the average error rate of polynucleotides synthesized in the libraries described herein, without error correction, is often less than 1 in 1000, less than about 1 in 2000, less than about 1 in 3000, or even lower.
[0107] Provided herein are surfaces that support parallel synthesis of multiple polynucleotides with different predetermined sequences at addressable locations on a common support. In some examples, the substrates are 50, 100, 200, 400, 600, 800, 1000, 1200, 1400, 1600, 1800, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, Support for the synthesis of 1,000,000, 1,200,000, 1,400,000, 1,600,000, 1,800,000, 2,000,000, 2,500,000, 3,000,000, 3,500,000, 4,000,000, 4,500,000, 5,000,000, 10,000,000 or more non-identical polynucleotides is provided. In some cases, the surface may be 50, 100, 200, 400, 600, 800, 1000, 1200, 1400, 1600, 1800, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,0 In some embodiments, the present invention provides support for the synthesis of polynucleotides encoding 00,000, 1,200,000, 1,400,000, 1,600,000, 1,800,000, 2,000,000, 2,500,000, 3,000,000, 3,500,000, 4,000,000, 4,500,000, 5,000,000, 10,000,000 or more different sequences, in which at least some of the polynucleotides have the same sequence or are configured to be synthesized with the same sequence. In some examples, the substrate provides a surface environment for the growth of polynucleotides having at least 80, 90, 100, 120, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500 or more bases.
[0108] Provided herein are methods for synthesizing polynucleotides on different loci of a substrate, where each locus supports the synthesis of a population of polynucleotides. In some cases, each locus supports the synthesis of a population of polynucleotides having a different sequence than a population of polynucleotides grown at another locus. In some examples, each polynucleotide sequence is synthesized with 1, 2, 3, 4, 5, 6, 7, 8, 9 or more redundancies across different loci within the same cluster of loci on a surface for polynucleotide synthesis. In some examples, the loci of the substrate are located within multiple clusters. In some examples, the substrate comprises at least 10, 500, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 11000, 12000, 13000, 14000, 15000, 20000, 30000, 40000, 50000 or more clusters. In some examples, the substrate is 2,000, 5,000, 10,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 1,100,000, 1,200,000, 1,300,000, 1,400,000, 1,500,000, 1,600,000, 1,700,000, 1,800,000, 1,900,000, 2,000,000, The substrate may comprise 0, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 1,200,000, 1,400,000, 1,600,000, 1,800,000, 2,000,000, 2,500,000, 3,000,000, 3,500,000, 4,000,000, 4,500,000, 5,000,000, or 10,000,000 or more distinct loci. In some examples, the substrate comprises about 10,000 distinct loci. The number of loci in a single cluster may vary in various cases.In some cases, each cluster comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 130, 150, 200, 300, 400, 500, or more loci. In some examples, each cluster comprises about 50-500 loci. In some examples, each cluster comprises about 100-200 loci. In some examples, each cluster comprises about 100-150 loci. In some examples, each cluster comprises about 109, 121, 130, or 137 loci. In some examples, each cluster comprises about 19, 20, 61, 64, or more loci. Alternatively, or in combination, polynucleotide synthesis is performed on a homogenous array surface.
[0109] In some instances, the number of different polynucleotides synthesized on a substrate depends on the number of different loci available within the substrate. In some instances, the density of loci within a cluster or surface of a substrate is greater than or equal to 1 mm 2 In some cases, the substrate is 10-500, 25-400, 50-500, 100-500, 150-500, 10-250, 50-250, 10-200, or 50-200 mm 2In some examples, the distance between the centers of two adjacent loci within a cluster or surface is about 10-500, about 10-200, or about 10-100 um. In some examples, the distance between the centers of two adjacent loci is greater than about 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 um. In some examples, the distance between the centers of two adjacent loci is less than about 200, 150, 100, 80, 70, 60, 50, 40, 30, 20, or 10 um. In some examples, each locus has a width of about 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 um. In some cases, each locus has a width of about 0.5-100, 0.5-50, 10-75, or 0.5-50 um.
[0110] In some instances, the density of clusters within the matrix is at least or about 100 mm 2 1 cluster per 10mm 2 1 cluster per 5mm 2 1 cluster per 4mm 2 1 cluster per 3mm 2 1 cluster per 2mm 2 1 cluster per 1mm 2 1 cluster per 1mm 2 2 clusters per 1mm 2 3 clusters per 1mm 2 4 clusters per 1mm 2 5 clusters per 1mm 2 10 clusters per 1mm 2 In some instances, the substrate has 50 or more clusters per 10 mm 2 Approximately 1 cluster ~ 1mm per 2In some examples, the distance between the centers of two adjacent clusters is at least or about 50, 100, 200, 500, 1000, 2000, or 5000 um. In some cases, the distance between the centers of two adjacent clusters is about 50-100, 50-200, 50-300, 50-500, and 100-2000 um. In some cases, the distance between the centers of two adjacent clusters is about 0.05-50, 0.05-10, 0.05-5, 0.05-4, 0.05-3, 0.05-2, 0.1-10, 0.2-10, 0.3-10, 0.4-10, 0.5-10, 0.5-5, or 0.5-2 mm. In some cases, each cluster has a cross section of about 0.5 to about 2, about 0.5 to about 1, or about 1 to about 2 mm. In some cases, each cluster has a cross section of about 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, or 2 mm. In some cases, each cluster has an internal cross section of about 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.15, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, or 2 mm.
[0111] In some examples, the substrate is as large as a standard 96-well plate, e.g., about 100 to about 200 mm, about 50 to about 150 mm. In some examples, the substrate has a diameter of about 1000, 500, 450, 400, 300, 250, 200, 150, 100, or 50 mm or less. In some examples, the substrate has a diameter of about 25 to 1000, 25 to 800, 25 to 600, 25 to 500, 25 to 400, 25 to 300, or 25 to 200 mm. In some examples, the substrate has a diameter of at least about 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 12,000, 15,000, 20,000, 30,000, 40,000, 50,000 mm. 2 In some examples, the thickness of the substrate is about 50-2000, 50-1000, 100-1000, 200-1000, or 250-1000 mm.
[0112] surface material The substrates, devices, and reactors provided herein may be fabricated from any type of material suitable for the methods, compositions, and systems described herein. In certain examples, the substrate material is fabricated to exhibit low levels of nucleotide binding. In some examples, the substrate material is modified to create different surfaces that exhibit high levels of nucleotide binding. In some examples, the substrate material is transparent to visible and / or UV light. In some examples, the substrate material is sufficiently conductive, e.g., capable of forming a uniform electric field across all or a portion of the substrate. In some examples, the conductive material is connected to an electrical ground. In some examples, the substrate is thermally conductive or insulating. In some examples, the material is chemically and thermally resistant to support chemical or biochemical reactions, e.g., polynucleotide synthesis reaction processes. In some examples, the substrate comprises a flexible material. For flexible materials, the material includes, but is not limited to, nylon (both modified and unmodified nylon), nitrocellulose, polypropylene, and the like. In some examples, the substrate comprises a rigid material. For rigid materials, the materials include, but are not limited to, glass, fused silica, silicon, plastics (e.g., polytetrafluoroethylene, polypropylene, polystyrene, polycarbonate, and mixtures thereof, etc.), and metals (e.g., gold, platinum, etc.). The substrate, solid support, or reactor may be made from a material selected from the group consisting of silicon, polystyrene, agarose, dextran, cellulose polymers, polyacrylamide, polydimethylsiloxane (PDMS), and glass. The substrate / solid support or microstructures / reactors therein may be made from a combination of the materials listed herein or other suitable materials known in the art.
[0113] surface structure Provided herein are substrates for the methods, compositions, and systems described herein, which have a surface structure suitable for the methods, compositions, and systems described herein. In some examples, the substrate includes raised and / or recessed features. One advantage of such features is an increased surface area to support synthesis of polynucleotides. In some examples, substrates with raised and / or recessed features are referred to as three-dimensional substrates. In some cases, the three-dimensional substrate includes one or more channels. In some cases, one or more loci include channels. In some cases, the channels are accessible for reagent deposition via a deposition device, such as a material deposition device. In some cases, the reagents and / or fluids are collected in a larger well that is in fluid communication with one or more channels. For example, the substrate includes multiple channels corresponding to multiple loci with a cluster, the multiple channels being in fluid communication with one well of the cluster. In some methods, a library of polynucleotides is synthesized at multiple loci of the cluster.
[0114] Provided herein are substrates for the methods, compositions, and systems described herein, which are configured for polynucleotide synthesis. In some examples, the structure is configured to allow controlled flow and mass transfer pathways for polynucleotide synthesis on the surface. In some examples, the substrate configuration allows controlled and uniform distribution of mass transfer pathways, chemical exposure times, and / or washing effects during polynucleotide synthesis. In some examples, the substrate configuration can increase sweep efficiency, for example, by providing sufficient volume for growing polynucleotides such that the volume displaced by the growing polynucleotides does not exceed 50, 45, 40, 35, 30, 25, 20, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1% or less of the initially available volume available or suitable for polynucleotide growth. In some examples, the three-dimensional structure allows for management of fluid flow and rapid exchange of chemical exposure.
[0115] Provided herein are substrates for the methods, compositions, and systems described herein, which have a structure suitable for the methods, compositions, and systems described herein. In some examples, separation is achieved by physical structure. In some examples, separation is achieved by differential functionalization of the surface, creating active and passive areas for polynucleotide synthesis. In some examples, differential functionalization is achieved by alternating hydrophobicity across the substrate surface, thereby creating water contact angle effects that cause beading or wetting of deposited reagents. Larger structures can be employed to reduce splashing and cross-contamination of reagents at different polynucleotide synthesis sites and adjacent sites. In some cases, devices such as material deposition devices are used to deposit reagents at different polynucleotide synthesis sites. Substrates with three-dimensional features are configured to allow synthesis of large numbers of polynucleotides (e.g., about 10,000 or more) with low error rates (e.g., less than about 1:500, 1:1000, 1:1500, 1:2,000, 1:3,000, 1:5,000, or 1:10,000). In some cases, the substrate is 1 mm 2 In some embodiments, the density of features may be greater than or equal to about 1, 5, 10, 20, 30, 40, 50, 60, 70, 80, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 300, 400, or 500 features per feature.
[0116] A well in the substrate can have the same or different width, height, and / or volume as another well in the substrate. A channel in the substrate can have the same or different width, height, and / or volume as another channel in the substrate. In some examples, the diameter of the cluster or the diameter of the well containing the cluster, or both, is about 0.05-50, 0.05-10, 0.05-5, 0.05-4, 0.05-3, 0.05-2, 0.05-1, 0.05-0.5, 0.05-0.1, 0.1-10, 0.2-10, 0.3-10, 0.4-10, 0.5-10, 0.5-5, or 0.5-2 mm. In some examples, the diameter of the cluster or well, or both, is about 5, 4, 3, 2, 1, 0.5, 0.1, 0.09, 0.08, 0.07, 0.06, or 0.05 mm or less. In some examples, the diameter of the cluster or well, or both, is about 1.0-1.3 mm. In some examples, the diameter of the cluster or well, or both, is about 1.150 mm. In some examples, the diameter of the cluster or well, or both, is about 0.08 mm. The diameter of the cluster refers to a cluster within a two-dimensional or three-dimensional matrix.
[0117] In some examples, the height of the wells is about 20-1000, 50-1000, 100-1000, 200-1000, 300-1000, 400-1000, or 500-1000 um. In some cases, the height of the wells is less than about 1000, 900, 800, 700, or 600 um.
[0118] In some instances, the substrate includes multiple channels corresponding to multiple loci in the cluster, and the channel height or depth is between 5-500, 5-400, 5-300, 5-200, 5-100, 5-50, or 10-50 um. In some instances, the channel height is less than 100, 80, 60, 40, or 20 um.
[0119] In some examples, the diameter of a channel, locus (e.g., in a substantially planar substrate), or both the channel and locus (e.g., in a three-dimensional substrate where the locus corresponds to a channel) is from about 1-1000, 1-500, 1-200, 1-100, 5-100, or 10-100 um, such as about 90, 80, 70, 60, 50, 40, 30, 20, or 10 um. In some examples, the diameter of a channel, locus, or both the channel and locus is less than about 100, 90, 80, 70, 60, 50, 40, 30, 20, or 10 um. In some examples, the distance between the centers of two adjacent channels, loci, or channels and locations is from about 1-500, 1-200, 1-100, 5-200, 5-100, 5-50, or 5-30, such as about 20 um.
[0120] Surface Modification Provided herein are methods for synthesizing polynucleotides on a surface, the surface comprising various surface modifications. In some examples, surface modifications are used to chemically and / or physically modify a surface by additive or subtractive processes to alter one or more chemical and / or physical properties of the substrate surface or selected sites or regions of the substrate surface. For example, surface modifications include, but are not limited to, (1) modifying the wettability of the surface, (2) functionalizing the surface, i.e., providing, modifying or replacing surface functional groups, (3) defunctionalizing the surface, i.e., removing surface functional groups, (4) changing the chemical composition of the surface, e.g., by etching, (5) increasing or decreasing the surface roughness, (6) providing a coating on the surface, e.g., a coating that exhibits a wettability different from that of the surface, and / or (7) depositing particulates on the surface.
[0121] In some cases, the addition of a chemical layer (referred to as an adhesion promoter) on the surface facilitates structured patterning of loci on the surface of the substrate. Exemplary surfaces for the application of adhesion promoters include, but are not limited to, glass, silicon, silicon dioxide, and silicon nitride. In some cases, the adhesion promoter is a chemical with high surface energy. In some examples, a second chemical layer is deposited on the surface of the substrate. In some cases, the second chemical layer has a low surface energy. In some cases, the surface energy of the chemical layer coated on the surface supports localization of the droplets on the surface. Depending on the selected pattern arrangement, the proximity of the loci and / or the fluid contact area at the loci can be altered.
[0122] In some examples, the substrate surface or resolved locus onto which nucleic acids or other moieties are deposited, for example for polynucleotide synthesis, is smooth or substantially flat (e.g., two-dimensional) or has irregularities, such as raised or depressed features (e.g., three-dimensional features). In some examples, the substrate surface is modified with one or more distinct layers of compounds. Such modified layers of interest include, but are not limited to, inorganic and organic layers, such as metals, metal oxides, polymers, small organic molecules, etc.
[0123] In some examples, the degraded loci of the substrate are functionalized with one or more moieties that increase and / or decrease the surface energy. In some cases, the moieties are chemically inert. In some cases, the moieties are configured to support one or more processes in a desired chemical reaction, e.g., a polynucleotide synthesis reaction. The surface energy, i.e., hydrophobicity, of the surface is a factor that determines the affinity with which nucleotides will attach onto the surface. In some examples, a method of functionalizing a substrate includes (a) providing a substrate having a surface that includes silicon dioxide, and (b) silanizing the surface using a suitable silanizing agent, e.g., an organofunctional alkoxysilane molecule, as described herein or known in the art. Methods and functionalizing agents are described in U.S. Pat. No. 5,474,796, which is incorporated herein by reference in its entirety.
[0124] In some instances, the substrate surface is functionalized by contacting with a derivatization composition containing a mixture of silanes, usually under reaction conditions effective to couple the silanes to the substrate surface via reactive hydrophilic moieties present on the substrate surface. Silanization generally involves covering the surface by self-assembly with organofunctional alkoxysilane molecules. Additionally, various siloxane functionalization reagents can be used, for example, to lower or increase surface energy, as currently known in the art. Organofunctional alkoxysilanes are classified according to their organic functionality.
[0125] Polynucleotide Synthesis The disclosed polynucleotide synthesis method may include a process involving phosphoramidite chemistry. In some examples, the polynucleotide synthesis includes coupling of a base with a phosphoramidite. The polynucleotide synthesis may include coupling of a base by deposition of a phosphoramidite under coupling conditions, where the same base is optionally deposited multiple times with a phosphoramidite (i.e., double coupling). The polynucleotide synthesis may include capping of unreacted sites. In some examples, capping is optional. The polynucleotide synthesis may also include oxidation or one or more oxidation steps. The polynucleotide synthesis may include deblocking, detritylation, and sulfurization. In some examples, the synthesis of a polynucleotide includes either oxidation or sulfurization. In some examples, between one or each step during the polynucleotide synthesis reaction, the device is washed, for example, using tetrazole or acetonitrile. The time frame of any one step in the phosphoramidite synthesis method may be less than about 2 minutes, 1 minute, 50 seconds, 40 seconds, 30 seconds, 20 seconds, and 10 seconds.
[0126] Polynucleotide synthesis using the phosphoramidite method can include adding phosphoramidite building blocks (e.g., nucleoside phosphoramidites) to a growing polynucleotide chain to form a phosphite triester bond. Synthesis of phosphoramidite polynucleotides proceeds in the 3' to 5' direction. Phosphoramidite polynucleotide synthesis allows for the controlled addition of one nucleotide to a growing nucleic acid chain per synthesis cycle. In some examples, each synthesis cycle includes a coupling step. Phosphoramidite coupling involves the formation of a phosphite triester bond between an activated nucleoside phosphoramidite and a nucleoside attached to a substrate, for example, via a linker. In some examples, the nucleoside phosphoramidite is provided to an activated device. In some examples, the nucleoside phosphoramidite is provided to a device together with an activator. In some examples, the nucleoside phosphoramidite is provided to the device in 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100 or more times excess of the nucleoside bound to the substrate. In some examples, the addition of the nucleoside phosphoramidite is performed in an anhydrous environment, for example, anhydrous acetonitrile. After the addition of the nucleoside phosphoramidite, the device is optionally washed. In some examples, the coupling step is repeated one or more times, optionally with a washing step between the addition of the nucleoside phosphoramidite to the substrate. In some examples, the polynucleotide synthesis method used herein comprises one, two, three or more consecutive coupling steps. Prior to coupling, the device-bound nucleoside is often deprotected by removing a protecting group, which serves to prevent polymerization. A common protecting group is 4,4'-dimethoxytrityl (DMT).
[0127] Following coupling, the phosphoramidite polynucleotide synthesis method optionally includes a capping step, in which the growing polynucleotide is treated with a capping agent. The capping step serves to block the unreacted substrate-bound 5'-OH group after coupling to prevent further chain elongation and to prevent the formation of polynucleotides with internal base deletions. Additionally, 1H-tetrazole-activated phosphoramidites may react to a small extent with the O6 position of guanosine. Without being bound by theory, this by-product may be depurinated, possibly via O6-N7 migration, upon oxidation with I2 / water. The apurinic site may be cleaved during the final deprotection of the polynucleotide, resulting in a reduced yield of the full-length product. The O6 modification may be removed by treatment with a capping reagent prior to oxidation with I2 / water. In some instances, the inclusion of a capping step during polynucleotide synthesis reduces the error rate compared to uncapping synthesis. In one example, the capping step comprises treating the substrate-bound polynucleotides with a mixture of acetic anhydride and 1-methylimidazole. After the capping step, the device is optionally washed.
[0128] In some examples, after addition of nucleoside phosphoramidites, and optionally after capping and one or more washing steps, the growing nucleic acid bound to the device is oxidized. The oxidation step involves a phosphite triester that is oxidized to a tetracoordinate phosphate triester, which is a protected precursor of the naturally occurring phosphodiester internucleoside linkage. In some examples, oxidation of the growing polynucleotide is accomplished by treatment with iodine and water, optionally in the presence of a weak base (e.g., pyridine, lutidine, collidine). Oxidation can be performed under anhydrous conditions, for example, using tert-butyl hydroperoxide or (1S)-(+)-(10-camphorsulfonyl)-oxaziridine (CSO). In some methods, a capping step is performed after oxidation. A second capping step allows the device to be dried, since residual water from possible oxidation may interfere with subsequent coupling. After oxidation, the device and growing polynucleotide are optionally washed. In some instances, the oxidation step is replaced by a sulfurization step to obtain polynucleotide phosphorothioates, and an optional capping step can be performed after sulfurization. Many reagents allow efficient sulfur transfer, including but not limited to 3-(dimethylaminomethylidene)amino)-3H-1,2,4-dithiazole-3-thione, DDTT, 3H-1,2-benzodithiol-3-one 1,1-dioxide (also called Beaucage reagent), and N,N,N'N'-tetraethylthiuram disulfide (TETD).
[0129] The protected 5' end of the growing polynucleotide bound to the device is removed so that the next cycle of nucleoside incorporation can occur by coupling, resulting in the primary hydroxyl group reacting with the next nucleoside phosphoramidite. In some examples, the protecting group is DMT, and deprotection occurs with trichloroacetic acid in dichloromethane. Longer detritylation times or use of stronger than recommended acid solutions can increase depurination of the solid support-bound polynucleotide, decreasing the yield of the desired full-length product. The disclosed methods and compositions described herein provide controlled deblocking conditions that limit undesired depurination reactions. In some examples, the polynucleotide bound to the device is washed after deblocking. In some examples, efficient washing after deblocking results in a low error rate of the synthesized polynucleotide.
[0130] Methods for synthesizing polynucleotides typically involve a repeated sequence of steps: applying a protected monomer to an active functionalized surface (e.g., a position), conjugating with either an activated surface, a linker, or a previously deprotected monomer, deprotecting the applied monomer so that it can react with a subsequently applied protected monomer, and applying another protected monomer for conjugation. One or more intermediate steps include oxidation or sulfurization. In some instances, one or more washing steps are performed before or after one or all of the steps.
[0131] Phosphoramidite-based polynucleotide synthesis methods include a series of chemical steps. In some examples, one or more steps of the synthesis method include reagent cycling, where one or more steps of the method include application of reagents useful for the step to a device. For example, the reagents are cycled through a series of liquid deposition and vacuum drying steps. In the case of substrates with three-dimensional structures such as wells, microwells, channels, etc., the reagents are optionally passed through one or more regions of the device via the wells and / or channels.
[0132] The methods and systems described herein relate to a polynucleotide synthesis device for synthesizing polynucleotides. The synthesis can be performed in parallel. For example, at least or about at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 1000, 10000, 50000, 75000, 100000 or more polynucleotides can be synthesized in parallel. The total number of polynucleotides that can be synthesized in parallel is 2 to 100,000, 3 to 50,000, 4 to 10,000, 5 to 1,000, 6 to 900, 7 to 850, 8 to 800, 9 to 750, 10 to 700, 11 to 650, 12 to 600, 13 to 550, 14 to 500, 15 to 450, 16 to 400, 17 to 350, 18 to 300, 19 to 250, 20 to 200, 21 to 150, 22 to 100, 23 to 50, 24 to 45, 25 to 40, 30 to 35. Those skilled in the art will understand that the total number of polynucleotides synthesized in parallel can fall within any range limited by any of these values, for example, 25 to 100. The total number of polynucleotides synthesized in parallel can fall within any range, defined by any value serving as the end point of the range. The total molar mass of the polynucleotides synthesized in the device, or the molar mass of each of the polynucleotides, can be at least or at least about 10, 20, 30, 40, 50, 100, 250, 500, 750, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 25000, 50000, 75000, 100000 picomoles, or more. The length of each of the polynucleotides in the device, or the average length of the polynucleotides, can be at least or at least about at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 150, 200, 300, 400, 500 nucleotides, or more.The length of each of the polynucleotides in the device or the average length of the polynucleotides may be at most or at most about 500, 400, 300, 200, 150, 100, 50, 45, 35, 30, 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10 nucleotides, or less. The length of each of the polynucleotides in the device or the average length of the polynucleotides may be in the range of 10-500, 9-400, 11-300, 12-200, 13-150, 14-100, 15-50, 16-45, 17-40, 18-35, 19-25. One of skill in the art will appreciate that the length of each of the polynucleotides in the device or the average length of the polynucleotides may fall within any range bounded by any of these values, for example, in the range of 100-300. The length of each of the polynucleotides or the average length of the polynucleotides in the device can fall within any range, defined by either of the values serving as the endpoints of the range.
[0133] The methods of polynucleotide synthesis on a surface provided herein allow for rapid synthesis. In one example, at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 125, 150, 175, 200 nucleotides or more are synthesized per hour. Nucleotides include adenine, guanine, thymine, cytosine, uridine building blocks, or analogs / variants thereof. In some examples, libraries of polynucleotides are synthesized in parallel on a substrate. For example, a device comprising about or at least about 100, 1,000, 10,000, 30,000, 75,000, 100,000, 1,000,000, 2,000,000, 3,000,000, 4,000,000, or 5,000,000 decomposed loci can support the synthesis of at least as many different polynucleotides, where polynucleotides encoding different sequences are synthesized on the decomposed loci. In some examples, a library of polynucleotides is synthesized on a low error rate device described herein in about 3 months, 2 months, 1 month, 3 weeks, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2 days, 24 hours, or less. In some examples, larger nucleic acids assembled from polynucleotide libraries synthesized with low error rates using the substrates and methods described herein are prepared in about 3 months, 2 months, 1 month, 3 weeks, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2 days, 24 hours, or less.
[0134] In some examples, the methods described herein provide for the creation of a library of nucleic acids comprising variant nucleic acids that differ at multiple codon sites, hi some examples, the nucleic acids can have 1 site, 2 sites, 3 sites, 4 sites, 5 sites, 6 sites, 7 sites, 8 sites, 9 sites, 10 sites, 11 sites, 12 sites, 13 sites, 14 sites, 15 sites, 16 sites, 17 sites, 18 sites, 19 sites, 20 sites, 30 sites, 40 sites, 50 sites, or more variant codon sites.
[0135] In some examples, one or more of the variant codon sites may be contiguous, hi some examples, one or more of the variant codon sites may be non-contiguous and separated by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more codons.
[0136] In some examples, a nucleic acid may contain multiple variant codon sites, where all variant codon sites are adjacent to each other to form a contiguous variant codon site. In some examples, a nucleic acid may contain multiple variant codon sites, where none of the variant codon sites are adjacent to each other. In some examples, a nucleic acid may contain multiple variant codon sites, where some variant codon sites are adjacent to each other to form a contiguous variant codon site, and some variant codon sites are not adjacent to each other.
[0137] Referring to the figures, Figure 3 shows an exemplary process workflow for synthesizing nucleic acids (e.g., genes) from short nucleic acids. The workflow is generally divided into phases: (1) de novo synthesis of single-stranded nucleic acid libraries, (2) joining nucleic acids to form larger fragments, (3) error correction, (4) quality control, and (5) shipping. Prior to de novo synthesis, the intended nucleic acid sequence or group of nucleic acid sequences is preselected. For example, a group of genes is preselected for creation.
[0138] Once the large nucleic acids to be created are selected, a given nucleic acid library is designed for de novo synthesis. Various suitable methods for generating high density polynucleotide arrays are known. In the example workflow, a device surface layer is provided. In this example, the surface chemistry is altered to improve the polynucleotide synthesis process. Areas of low surface energy are created to repel liquids, and areas of high surface energy are created to attract liquids. The surface itself can be a flat surface shape, or can include shape variations such as protrusions or microwells that increase the surface area. In the example workflow, the selected high surface energy molecules serve the dual function of supporting DNA chemistry, as disclosed in International Patent Application Publication No. WO / 2015 / 021080, which is hereby incorporated by reference in its entirety.
[0139] In situ preparation of polynucleotide arrays is created on a solid support and multiple oligomers are extended in parallel using a single nucleotide extension process. A deposition device, such as a material deposition device, is designed to release reagents in a stepwise manner so that multiple polynucleotides are extended in parallel, one residue at a time, to create oligomers having a predetermined nucleic acid sequence 302. In some instances, the polynucleotides are cleaved from the surface at this stage. Cleavage includes, for example, gas cleavage with ammonia or methylamine.
[0140] The created polynucleotide library is placed in a reaction chamber. In this exemplary workflow, the reaction chamber (also referred to as a "nanoreactor") is a silicon-coated well containing PCR reagents and lowered over the polynucleotide library 303. Reagents are added to release the polynucleotides from the substrate before or after sealing the polynucleotides 304. In the exemplary workflow, the polynucleotides are released after sealing the nanoreactor 305. Once released, the single-stranded polynucleotide fragments hybridize to span the entire long-range sequence of DNA. Partial hybridization 305 is possible because each synthesized polynucleotide is designed to have a small portion that overlaps with at least one other polynucleotide in the pool.
[0141] After hybridization, the PCA reaction is initiated. During polymerase cycles, polynucleotides anneal to complementary fragments and gaps are filled by the polymerase. Each cycle randomly increases the length of the various fragments depending on which polynucleotides find each other. Complementarity between the fragments allows the formation of complete long spans of double-stranded DNA306.
[0142] After PCA is complete, the nanoreactor is separated from the device 307 and placed to interact with a device with primers for PCR 308. After sealing, the nanoreactor undergoes PCR 309 to amplify larger nucleic acids. After PCR 310, the nanochamber is opened 311, error correction reagents are added 312, the chamber is sealed 313, and an error correction reaction occurs to remove mismatched base pairs and / or less complementary strands from the double stranded PCR amplification product 314. The nanoreactor is opened and separated 315. The error corrected product then undergoes additional processing steps such as PCR and molecular barcoding and is packaged 322 for shipping 323.
[0143] In some instances, quality control procedures are taken. After error correction, quality control steps include, for example, interacting the wafer with sequencing primers to amplify the error corrected product 316, sealing the wafer in a chamber containing the error corrected amplification product 317, and performing an additional round of amplification 318. The nanoreactors are opened 319 and the products are pooled 320 and sequenced 321. After acceptable quality control determinations are made, the packaged product 322 is approved for shipment 323.
[0144] In some examples, nucleic acids created by a workflow such as that of Figure 3 are mutagenized using overlapping primers as disclosed herein. In some examples, a library of primers is created by in situ preparation on a solid support and a single nucleotide extension process is utilized to extend multiple oligomers in parallel. A deposition device, such as a material deposition device, is designed to release reagents in a stepwise manner such that multiple polynucleotides are extended in parallel, one residue at a time, to create oligomers having a predetermined nucleic acid sequence 302.
[0145] Computer Systems Any of the systems described herein can be operably linked to a computer and can be automated locally or remotely via a computer. In various examples, the methods and systems of the present disclosure can further include software programs and the use of a computer system. Thus, computer control for synchronization of dispense / vacuum / replenish functions, such as coordinating and synchronizing the movement, dispense operation, and vacuum actuation of the material deposition device, is within the scope of the present disclosure. The computer system can be programmed to create an interface between the user-specified base sequence and the location of the material deposition device to deliver the correct reagent to the specified area of the substrate.
[0146] The computer system 400 shown in FIG. 4 can be understood as a logical device that can read instructions from a medium 411 and / or a network port 405, which can be optionally connected to a server 409 having a fixed medium 412. A system as shown in FIG. 4 can include a CPU 401, a disk drive 403, optional input devices such as a keyboard 415 and / or a mouse 416, and an optional monitor 407. Data communication can be made to a server at a nearby or remote location over a designated communication medium. A communication medium can include any means of transmitting and / or receiving data. For example, a communication medium can be a network connection, a wireless connection, or an Internet connection. Such a connection can provide communication over the World Wide Web. It is envisioned that data related to the present disclosure can be transmitted over such a network or connection for receipt and / or review by a party 422 as shown in FIG. 4.
[0147] FIG. 5 is a block diagram illustrating a first exemplary structure of a computer system 500 that can be used in connection with an example instance of the present disclosure. As shown in FIG. 5, the example computer system can include a processor 502 for processing instructions. Non-limiting examples of processors include Intel Xeon™ processors, AMD Opteron™ processors, Samsung 32-bit RISC ARM1176JZ(F)-S v1.0™ processors, ARM Cortex-A8 Samsung S5PC100™ processors, ARM Cortex-A8 Apple A4™ processors, Marvell PXA930™ processors, or functionally equivalent processors. Parallel processing can be performed using multiple execution threads. In some examples, multiple processors or processors with multiple cores can also be used within a single computer system, in a cluster, or distributed among multiple systems over a network, including multiple computers, mobile phones, and / or personal data assistant devices.
[0148] As shown in FIG. 5, a high speed cache 504 may be connected to or incorporated into the processor 502 to provide high speed memory for instructions or data recently or frequently used by the processor 502. The processor 502 is connected to a north bridge 506 by a processor bus 508. The north bridge 506 is connected to a random access memory (RAM) 510 by a memory bus 512 and manages access to the RAM 510 by the processor 502. The north bridge 506 is also connected to a south bridge 514 by a chipset bus 516. The south bridge 514 is further connected to a peripheral bus 518. The peripheral bus may be, for example, a PCI, PCI-X, PCI Express, or other peripheral bus. The north bridge and south bridge are often referred to as the processor chipset and manage data transfers between the processor, RAM, and peripheral components on the peripheral bus 518. In some alternative architectures, instead of using a separate north bridge chip, the functionality of the north bridge may be incorporated into the processor. In some examples, the system 500 can include an accelerator card 522 connected to the peripheral bus 518. An accelerator can include a field programmable gate array (FPGA) or other hardware for accelerating a particular process. For example, an accelerator can be used for adaptive data reconstruction or evaluation of algebraic expressions used in extended set processing.
[0149] Software and data may be stored on external storage 524 and loaded into RAM 510 and / or cache 504 for use by the processor. System 500 includes an operating system for managing system resources, non-limiting examples of which include Linux, Windows™, MACOS™, BlackBerry OS™, iOS™, and other functionally equivalent operating systems, as well as application software that runs on the operating system to manage data storage and optimization in accordance with example instances of the present disclosure. In this example, system 500 also includes network interface cards (NICs) 520 and 521 connected to the peripheral bus to provide a network interface to external storage, such as network attached storage (NAS) and other computer systems that can be used for distributed parallel processing.
[0150] FIG. 6 illustrates a network 600 with multiple computer systems 602a and 602b, multiple mobile phones and personal data assistants 602c, and network attached storage (NAS) 604a and 604b. In an example instance, systems 602a, 602b, and 602c can manage data storage and optimize data access to data stored in network attached storage (NAS) 604a and 604b. Mathematical models can be used for data and evaluated using distributed parallel processing across computer systems 602a and 602b and mobile phones and personal data assistant system 602c. Computer systems 602a and 602b and mobile phones and personal data assistant system 602c can also provide parallel processing for adaptive data reconstruction of data stored in network attached storage (NAS) 604a and 604b. FIG. 6 is merely an example, and a wide variety of other computer architectures and systems can be used in combination with various examples of the present disclosure. For example, blade servers can be used to provide parallel processing. The processor blades can be connected through a backplane to provide parallel processing. Storage can also be connected to the backplane or through a separate network interface as network attached storage (NAS). In some example instances, the processors can maintain separate memory spaces and send data through a network interface, backplane, or other connector for parallel processing by other processors. In other examples, some or all of the processors can use a shared virtual address memory space.
[0151] FIG. 7 is a block diagram of a multiprocessor computer system 700 using a shared virtual address memory space according to one example. The system includes multiple processors 702a-702f that can access a shared memory subsystem 704. The system incorporates multiple programmable hardware memory algorithm processors (MAPs) 706a-706f within the memory subsystem 704. Each MAP 706a-706f includes a memory 708a-708f and one or more field programmable gate arrays (FPGAs) 710a-710f. The MAPs provide configurable functional units that can provide specific algorithms or parts of algorithms to the FPGAs 710a-710f for processing in close cooperation with the respective processors. For example, the MAPs can be used to evaluate algebraic expressions over a data model or to perform adaptive data restructuring in an example instance. In this example, each MAP is globally accessible by all processors for these purposes. In one configuration, each MAP can use direct memory access (DMA) to access an associated memory 708a-708f and execute tasks asynchronously and independently from the respective microprocessor 702a-702f. In this configuration, a MAP can send results directly to another MAP for pipelined and parallel execution of algorithms.
[0152] The above computer architectures and systems are merely examples, and various other computer, mobile phone, and personal data assistant architectures and systems can be used in connection with the example instances, including systems that use any combination of general purpose processors, co-processors, FPGAs and other programmable logic devices, systems on a chip (SOC), application specific integrated circuits (ASICs), and other processing and logic elements. In some examples, the computer systems can be implemented in whole or in part in software or hardware. Various data storage media can be used in connection with the example instances, including random access memory, hard drives, flash memory, tape drives, disk arrays, network attached storage (NAS), and other local or distributed data storage devices and systems.
[0153] In an example instance, the computer system may be implemented using software modules executing on any of the above or other computer architectures and systems. In other examples, the functionality of the system may be implemented partially or fully in firmware, programmable logic devices such as field programmable gate arrays (FPGAs) as referenced in Figure 5, systems on chips (SOCs), application specific integrated circuits (ASICs), or other processing and logic elements. For example, the set processor and optimizer may be implemented using hardware acceleration through the use of a hardware accelerator card, such as accelerator card 522 shown in Figure 5.
[0154] The following examples are presented to more clearly illustrate to those skilled in the art the principles and practice of the embodiments disclosed herein, and should not be construed as limiting the scope of any claimed embodiments. All parts and percentages are by weight unless otherwise specified. EXAMPLES
[0155] The following examples are presented for the purpose of illustrating various embodiments of the present disclosure and are not intended to limit the disclosure in any manner. The examples, together with the methods described herein, are currently representative of preferred embodiments, are exemplary, and are not intended to limit the scope of the disclosure. Modifications therein and other uses encompassed within the spirit of the disclosure as defined by the claims will be apparent to those skilled in the art.
[0156] Example 1: Functionalization of the device surface The device was functionalized to support the attachment and synthesis of libraries of polynucleotides. First, the device surface was wet cleaned using a piranha solution containing 90% H2SO4 and 10% H2O2 for 20 min. The device was washed with DI water in multiple beakers, placed under a gooseneck tap of DI water for 5 min, and dried with N2. The device was then immersed in NH4OH (1:100, 3 mL:300 mL) for 5 min, rinsed with DI water using a hand gun, successively immersed in three beakers of DI water for 1 min each, and then rinsed again with DI water using a hand gun. The device was then plasma cleaned by exposing the device surface to O2. O2 was plasma etched for 1 min using a SAMCO PC-300 device at 250 watts in downstream mode.
[0157] The cleaned device surfaces were actively functionalized with a solution containing N-(3-triethoxysilylpropyl)-4-hydroxybutyramide using a YES-1224P deposition oven system with the following parameters: 0.5-1 Torr, 60 min, 70 °C, vaporizer at 135 °C. The device surfaces were resist-coated using a Brewer Science 200X spin coater. SPR™ 3612 photoresist was spin-coated onto the devices at 2500 rpm for 40 seconds. The devices were pre-baked at 90 °C for 30 minutes on a Brewer hotplate. The devices were photolithographically processed using a Karl Suss MA6 mask aligner device. The devices were exposed for 2.2 seconds and developed with MSF 26A for 1 minute. Residual developer was washed off with a hand gun and the devices were immersed in water for 5 minutes. The devices were baked in an oven at 100 °C for 30 minutes and then visually inspected for lithographic defects using a Nikon L200. A descum process was used to remove residual resist using a SAMCO PC-300 device with O2 plasma etching at 250 watts for 1 minute.
[0158] The device surface was passively functionalized with 100 μL of perfluorooctyltrichlorosilane solution mixed with 10 μL of light mineral oil. The device was placed in the chamber and pumped for 10 minutes, after which the valve to the pump was closed and left for 10 minutes. The chamber was vented. The device was resist stripped by immersing it twice in 500 mL of NMP at 70 °C for 5 minutes while sonicating at maximum power (9 for the Crest system). The device was then immersed in 500 mL of isopropanol at room temperature for 5 minutes and sonicated at maximum power. The device was then immersed in 300 mL of 200 proof ethanol and blown dry with N2. The functionalized surface was then activated to serve as a support for polynucleotide synthesis.
[0159] Example 2: Synthesis of 50-mer sequences on an oligonucleotide synthesis device The two-dimensional oligonucleotide synthesis device was assembled into a flow cell, which was then connected to a flow cell (Applied Biosystems (ABI394 DNA synthesizer)). The two-dimensional oligonucleotide synthesis device was uniformly functionalized with N-(3-triethoxysilylpropyl)-4-hydroxybutyramide (Gelest) and used to synthesize an exemplary 50 bp polynucleotide ("50-mer polynucleotide") using the polynucleotide synthesis methods described herein.
[0160] The sequence of the 50-mer is set forth in SEQ ID NO: 393. 5'AGACAATCAACCATTTGGGGTGGACAGCCTTGACCTCTAGACTTCGGCAT##TTTTTTTTTT3' (SEQ ID NO: 393), where # indicates thymidine-succinylhexamide CED phosphoramidite (CLP-2244 from ChemGenes), which is a cleavable linker that allows release of the oligo from the surface during deprotection.
[0161] The synthesis was carried out using standard DNA synthesis chemistry (coupling, capping, oxidation, and deblocking) according to the protocol in Table 1 and on an ABI synthesizer.
[0162] [Table 1-1]
[0163] [Table 1-2]
[0164] The phosphoramidite / activator combination was delivered similarly to the delivery of bulk reagents through a flow cell: no drying step was performed as the environment remained "wet" with reagents the entire time.
[0165] The flow restrictors were removed from the ABI394 synthesizer to allow for faster flows. Without the flow restrictors, the flow rates for amidites (0.1M in ACN), activator (0.25M benzoylthiotetrazole ("BTT", GlenResearch 30-3070-xx) in ACN, and Ox (0.02M I2 in 20% pyridine, 10% water and 70% THF) were approximately 100uL / sec, for acetonitrile ("ACN") and capping reagent (a 1:1 mix of CapA and CapB, where CapA is acetic anhydride in THF / pyridine and CapB is 16% 1-methylimidazole in THF), and for deblocking (Toluidine) were approximately 200uL / sec. The flow rate was approximately 300uL / sec for the oxidizer (3% dichloroacetic acid in ethylene) (compared to approximately 50uL / sec for all reagents when a flow restrictor was used). The time to fully pump out the oxidizer was observed and the timing of the chemical flow times was adjusted accordingly, and additional ACN washes were introduced between the different chemicals. After polynucleotide synthesis, the chip was deprotected in ammonia gas at 75 psi overnight. To recover the polynucleotides, five drops of water were applied to the surface. The recovered polynucleotides were then analyzed on a BioAnalyzer small RNA chip.
[0166] Example 3: Synthesis of 100-mer sequences on an oligonucleotide synthesis device Using the same process described in Example 2 for the synthesis of the 50-mer sequence, a 100-mer polynucleotide ("100-mer polynucleotide", 5'CGGGATCCTTATCGTCATCGTCGTACAGATCCCGACCCATTTGCTGTCCACCAGTCATGCTAGCCATACCATGATGATGATGATGATGAGAACCCCGCAT##TTTTTTTTTT3', where # indicates thymidine-succinylhexamide CED phosphoramidite (CLP-2244 from ChemGenes), SEQ ID NO: 2161) was synthesized on two different silicon chips, one functionalized homogeneously with N-(3-triethoxysilylpropyl)-4-hydroxybutyramide and the other functionalized with a 5 / 95 mixture of 11-acetoxyundecyltriethoxysilane and n-decyltriethoxysilane, and the polynucleotide extracted from the surface was analyzed with a bioanalyzer device.
[0167] All 10 samples from the two chips were run using forward (5'ATGCGGGGTTCTCATCATC3', SEQ ID NO: 2162) and reverse (5'CGGGATCCTTATCGTCATCG3', SEQ ID NO: 2163) primers in 50 uL PCR mix (25 uL NEB Q5 master mix, 2.5 uL 10 uM forward primer, 2.5 uL 10 uM reverse primer, 1 uL polynucleotide extracted from the surface, and up to 50 uL water) with the following thermal cycling program: 98℃, 30 seconds, 98℃, 10 sec; 63℃, 10 sec; 72℃, 10 sec; repeat 12 cycles. Further PCR amplification was carried out using 72°C for 2 minutes.
[0168] The PCR products were also run on a BioAnalyzer, which revealed a sharp peak at the 100-mer position. The PCR-amplified samples were then cloned and Sanger sequenced. Table 2 summarizes the results of Sanger sequencing for samples taken from spots 1–5 on chip 1 and spots 6–10 on chip 2.
[0169] [Table 2]
[0170] Thus, the high quality and uniformity of the synthesized polynucleotides was reproduced on the two chips with different surface chemistries. Overall, 89% of the sequenced 100-mers were perfect sequences without errors, which corresponds to 233 out of 262.
[0171] Table 3 summarizes the error characteristics of the sequences obtained from the polynucleotide samples of spots 1 to 10.
[0172] [Table 3]
[0173] Example 4: Exemplary Sequences
[0174] [Table 4-1]
[0175] [Table 4-2]
[0176] [Table 4-3]
[0177] [Table 4-4]
[0178] [Table 4-5]
[0179] [Table 4-6]
[0180]
Table 4-7
[0181]
Table 4-8
[0182]
Table 4-9
[0183]
Table 4-10
[0184]
Table 4-11
[0185]
Table 4-12
[0186]
Table 4-13
[0187]
Table 4-14
[0188]
Table 4-15
[0189]
Table 4-16
[0190]
Table 4-17
[0191]
Table 4-18
[0192]
Table 4-19
[0193]
Table 4-20
[0194]
Table 4-21
[0195]
Table 4-22
[0196]
Table 4-23
[0197]
Table 4-24
[0198]
Table 4-25
[0199]
Table 4-26
[0200]
Table 4-27
[0201]
Table 4-28
[0202]
Table 4-29
[0203]
Table 4-30
[0204]
Table 5-1
[0205]
Table 5-2
[0206]
Table 5-3
[0207]
Table 5-4
[0208]
Table 5-5
[0209]
Table 5-6
[0210]
Table 5-7
[0211]
Table 5-8
[0212]
Table 5-9
[0213]
Table 6-1
[0214]
Table 6-2
[0215]
Table 6-3
[0216]
Table 6-4
[0217]
Table 6-5
[0218]
Table 6-6
[0219]
Table 6-7
[0220]
Table 6-8
[0221]
Table 6-9
[0222]
Table 6-10
[0223]
Table 6-11
[0224]
Table 6-12
[0225]
Table 6-13
[0226]
Table 6-14
[0227]
Table 6-15
[0228]
Table 6-16
[0229]
Table 6-17
[0230]
Table 6-18
[0231]
Table 6-19
[0232]
Table 6-20
[0233]
Table 6-21
[0234]
Table 6-22
[0235]
Table 6-23
[0236]
Table 6-24
[0237]
Table 6-25
[0238]
Table 6-26
[0239]
Table 6-27
[0240]
Table 6-28
[0241]
Table 6-29
[0242]
Table 6-30
[0243]
Table 6-31
[0244]
Table 6-32
[0245]
Table 6-33
[0246]
Table 6-34
[0247]
Table 6-35
[0248]
Table 6-36
[0249]
Table 7-1
[0250]
Table 7-2
[0251]
Table 7-3
[0252]
Table 7-4
[0253]
Table 7-5
[0254]
Table 7-6
[0255]
Table 7-7
[0256]
Table 7-8
[0257]
Table 7-9
[0258]
Table 7-10
[0259]
Table 7-11
[0260]
Table 7-12
[0261]
Table 7-13
[0262]
Table 7-14
[0263]
Table 7-15
[0264]
Table 7-16
[0265]
Table 7-17
[0266]
Table 7-18
[0267]
Table 8-1
[0268]
Table 8-2
[0269]
Table 8-3
[0270]
Table 8-4
[0271]
Table 8-5
[0272] [Table 8-6]
[0273] [Table 8-7]
[0274] [Table 8-8]
[0275] [Table 8-9]
[0276] Example 5: DKK1 variants In this experiment, antibody yield, SPR affinity, and enrichment from eluted phages were tested (Tables 9-10).
[0277] The variable heavy and light domains of the anti-DKK1 antibody were reformatted into IgG2 or VHH-Fc based on IgG2 Fc for the nanobody leads. The reformatted leads were DNA reverse translated, synthesized and cloned into the mammalian expression vector pTwist CMV BG WPRE Neo. The light chain variable domains were reformatted into kappa and lambda frameworks accordingly. The cloned genes were delivered as purified plasmid DNA in preparation for transient transfection into HEK Expi293 cells (Thermo Fisher Scientific). A volume of 1.2 mL of culture was grown for 4 days, harvested and purified into 43 mM citrate, 148 mM HEPES, pH 6, 1.2 mL using Protein A resin (PhyNexus) on a Hamilton Microlab STAR platform. The yield was calculated by measuring the absorbance at 280 nm on a Lunatic instrument (UNCLE). The results are shown in Figure 10A.
[0278] SPR experiments were performed at 25°C in HBS-TE using a Carterra LSA SPR biosensor with a HC30M chip. Antibodies were diluted to 10 μg / mL and amine-coupled to the sensor chip by EDC / NHS activation followed by quenching with ethanolamine HCl. Increasing concentrations of analyte were flowed over the sensor chip in HBS-TE containing 0.5 mg / mL BSA and allowed to attach for 5 min and dissociate for 15 min. After each injection cycle, the surface was regenerated with two 30 s injections of IgG elution buffer (Thermo). Data were analyzed with Carterra's Kinetics Tool software using a 1:1 binding model. Results are shown in Figures 10B-C and 11A-B.
[0279] Long-read NGS sequencing was performed by sending PCR amplicons of DNA corresponding to the scFv or VHH of each clone to Loop Genomics for processing. The returned continuous FASTQ files were processed by the AIRR Python API to extract and annotate the antibody sequences. "NGS enrichment" refers to the number of instances where a particular antibody appeared in the fourth round of sequencing. "Cluster enrichment" refers to the number of instances where the exact antibody appeared in the fourth round or the number of instances where a variant within a Levenshtein distance of 3 appeared in the fourth round of sequencing. "Cluster rank" lists the antibody rank order of the antibodies belonging to the largest size cluster enrichment from lowest to highest. The results are shown in Figure 9A-C.
[0280] [Table 9-1]
[0281] [Table 9-2]
[0282] [Table 9-3]
[0283] [Table 10-1]
[0284] [Table 10-2]
[0285] [Table 10-3]
[0286] [Table 10-4]
[0287] [Table 10-5]
[0288] [Table 10-6]
[0289] [Table 10-7]
[0290] [Table 10-8]
[0291] [Table 10-9]
[0292] [Table 10-10]
[0293] Example 5: Panning and screening for identification of antibodies against DKK1 This example describes the identification of antibodies against DKK1. A phage display library was panned for binding to DKK1. Panning was performed as shown in FIG.
[0294] Carterra kinetics results showed that the VHH-Fc hits bound to DKK1 with high affinity (Figures 12A-D).
[0295] Figure 13 shows the results of a TCF / LEF reporter (Wnt signaling) assay. Wnt signaling activation was plotted together with SPR binding affinity.
[0296] 14A-14B show the results of an immune cell activation assay.
[0297] Tumor cell killing assays were performed, as shown in Figure 15A. The results showed that high affinity binders that are also Wnt signaling activators are not necessarily the same as potent immune cell activators and tumor killers (Figures 15B-G).
[0298] Figure 16 shows the antibody yield results from 1 mL of Expi293 cell culture. From DNA synthesis to antibody production, it took 31 days to generate 113 anti-DKK1 VHH-Fc.
[0299] Example 6: Testing of DKK1 antibodies This example describes the assays used to determine the validity of the anti-DKK1 leads identified in Example 5.
[0300] As shown in Figure 17A, two epitope bins were found among the top anti-DKK1 leads that bind to two different cysteine-rich domains (CRDs) (CRD1 or CRD2) within hDKK1, resulting in different activation pathways (Figures 17B-C).
[0301] Anti-DKK1 VHH lead was found to inhibit DKK1 binding to the receptor (Figure 18A-C). When DKK1 binds to LRP5 / 6, it inhibits Wnt TCF / LEF signaling, but anti-DKK1 lead inhibited DKK1 binding to the receptor, resulting in activation of TCF / LEF signaling.
[0302] The dual functional activities of DKK1-100 and DKK1-99 were tested in signal transduction assays, immune cell activation, and tumor cell killing (Figures 19A-C). Figures 22A-C show that the antagonism of WNT DKK1 inhibition in TCF / LEF assays is biphasic. Functional assays confirmed the concordant ranking of transient and cell line TCF / LEF reporters (Figures 23A-B).
[0303] DKK1 antibodies were tested for binding to LRP6 (Figure 24 and Figure 25A-C) and for immune cell activation (Figure 26A-B and Figure 27A-C). Antagonists were identified using signaling titration assays (Figure 28A-D). Additional immune assays were also performed (Figure 28A-B).
[0304] Example 7: In vivo efficacy of DKK1 antibodies Preclinical experiments on tumor regression using a mouse model are described in Figure 20A. In vivo efficacy results of PC3 tumor regression in SCID mice are shown in Figures 20B-D.
[0305] Killing of lung tumor organoids by immune cells through DKK1 inhibition is shown in Figure 30.
[0306] While preferred embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the present disclosure. It should be understood that various alternatives to the embodiments of the disclosure described herein may be employed in the practice of the present disclosure. The following claims define the scope of the present disclosure, and it is intended that methods and structures within the scope of these claims, and their equivalents, be covered thereby.
Claims
1. An antibody or antibody fragment that binds to dickkopf-1 (DKK1), wherein the antibody or antibody fragment comprises a VHH domain, and the VHH domain is (a) a heavy chain CDR1 comprising the amino acid sequence of SEQ ID NO: 923, a heavy chain CDR2 comprising the amino acid sequence of SEQ ID NO: 1337, and a heavy chain CDR3 comprising the amino acid sequence of SEQ ID NO: 1751; (b) a heavy chain CDR1 comprising the amino acid sequence of SEQ ID NO: 66, a heavy chain CDR2 comprising the amino acid sequence of SEQ ID NO: 164, and a heavy chain CDR3 comprising the amino acid sequence of SEQ ID NO: 262; (c) a heavy chain CDR1 comprising the amino acid sequence of SEQ ID NO: 69, a heavy chain CDR2 comprising the amino acid sequence of SEQ ID NO: 167, and a heavy chain CDR3 comprising the amino acid sequence of SEQ ID NO: 265; (d) a heavy chain CDR1 comprising the amino acid sequence of SEQ ID NO: 94, a heavy chain CDR2 comprising the amino acid sequence of SEQ ID NO: 192, and a heavy chain CDR3 comprising the amino acid sequence of SEQ ID NO: 290; (e) a heavy chain CDR1 comprising the amino acid sequence of SEQ ID NO: 98, a heavy chain CDR2 comprising the amino acid sequence of SEQ ID NO: 196, and a heavy chain CDR3 comprising the amino acid sequence of SEQ ID NO: 294; (f) a heavy chain CDR1 comprising the amino acid sequence of SEQ ID NO: 977, a heavy chain CDR2 comprising the amino acid sequence of SEQ ID NO: 1391, and a heavy chain CDR3 comprising the amino acid sequence of SEQ ID NO: 1805; (g) a heavy chain CDR1 comprising the amino acid sequence of SEQ ID NO: 1305, a heavy chain CDR2 comprising the amino acid sequence of SEQ ID NO: 1719, and a heavy chain CDR3 comprising the amino acid sequence of SEQ ID NO: 2133; (h) a heavy chain CDR1 comprising the amino acid sequence of SEQ ID NO: 10, a heavy chain CDR2 comprising the amino acid sequence of SEQ ID NO: 108, and a heavy chain CDR3 comprising the amino acid sequence of SEQ ID NO: 206; (i) a heavy chain CDR1 comprising the amino acid sequence of SEQ ID NO: 37, a heavy chain CDR2 comprising the amino acid sequence of SEQ ID NO: 135, and a heavy chain CDR3 comprising the amino acid sequence of SEQ ID NO: 233; (j) a heavy chain CDR1 comprising the amino acid sequence of SEQ ID NO: 919, a heavy chain CDR2 comprising the amino acid sequence of SEQ ID NO: 1333, and a heavy chain CDR3 comprising the amino acid sequence of SEQ ID NO: 1747; (k) a heavy chain CDR1 comprising the amino acid sequence of SEQ ID NO: 948, a heavy chain CDR2 comprising the amino acid sequence of SEQ ID NO: 1362, and a heavy chain CDR3 comprising the amino acid sequence of SEQ ID NO: 1776; (l) a heavy chain CDR1 comprising the amino acid sequence of SEQ ID NO: 951, a heavy chain CDR2 comprising the amino acid sequence of SEQ ID NO: 1365, and a heavy chain CDR3 comprising the amino acid sequence of SEQ ID NO: 1779; or (m) a heavy chain CDR1 comprising the amino acid sequence of SEQ ID NO: 1290, a heavy chain CDR2 comprising the amino acid sequence of SEQ ID NO: 1704, and a heavy chain CDR3 comprising the amino acid sequence of SEQ ID NO: 2118 An antibody or antibody fragment comprising:
2. The antibody or antibody fragment described in claim 1, wherein the VHH domain comprises an amino acid sequence that is at least 90% identical to the amino acid sequence set forth in any one of SEQ ID NOs: 360, 363, 388, 392, 452, 2231, 304, 331, 394, 398, 423, 426, and 2216.
3. The antibody or antibody fragment described in claim 1, wherein the VHH domain comprises an amino acid sequence set forth in any one of SEQ ID NOs: 360, 363, 388, 392, 452, 2231, 304, 331, 394, 398, 423, 426, and 2216.
4. The antibody or antibody fragment of claim 1, wherein the antibody is a monoclonal antibody, a bispecific antibody, a multispecific antibody, a grafted antibody, a humanized antibody, a synthetic antibody, a single domain antibody, an intrabody, or an antigen-binding fragment thereof.
5. An antibody or antibody fragment described in claim 1, wherein the antibody is a single domain antibody.
6. The antibody or antibody fragment described in claim 1, wherein the antibody or antibody fragment comprises an Fc region.
7. The antibody or antibody fragment of claim 6, wherein the Fc region is an IgG2 Fc region.
8. The antibody or antibody fragment of claim 1, wherein the antibody or antibody fragment comprises a K D of less than 50 nM.
9. A pharmaceutical composition comprising an antibody or antibody fragment described in any one of claims 1 to 8.
10. An isolated nucleic acid encoding an antibody or antibody fragment described in any one of claims 1 to 8.
11. An expression vector comprising the nucleic acid described in claim 10.
12. An isolated host cell containing the nucleic acid described in claim 10.
13. The pharmaceutical composition of claim 9, for use in a method for treating a disease or disorder, the method comprising administering the antibody or antibody fragment to a subject in need of treatment for the disease or disorder, wherein the disease or disorder is cancer, an inflammatory disease or disorder, a metabolic disease or disorder, a cardiovascular disease or disorder, a respiratory disease or disorder, pain, a digestive system disease or disorder, a reproductive system disease or disorder, an endocrine system disease or disorder, or a nervous system disease or disorder.
14. The pharmaceutical composition described in claim 13, wherein the cancer is gastroesophageal cancer, endometrial cancer, ovarian cancer, prostate cancer, or liver cancer.
15. The pharmaceutical composition of claim 13, wherein the metabolic disease or disorder or the endocrine disease or disorder is weight gain, obesity, type II diabetes, hypoglycemia, or hyperinsulinism.
16. The pharmaceutical composition of claim 9, for use in a method for activating immune cells, the method comprising contacting the immune cells with the antibody or antibody fragment.
17. The pharmaceutical composition described in claim 16, wherein the immune cells are natural killer (NK) cells.
18. The pharmaceutical composition of claim 9, for use in a method for increasing expression of interferon-gamma (IFNγ) in immune cells, the method comprising contacting the immune cells with the antibody or antibody fragment.
19. The pharmaceutical composition of claim 9, for use in a method for increasing expression of granulocyte-macrophage colony-stimulating factor (GM-CSF) in immune cells, the method comprising contacting the immune cells with the antibody or antibody fragment.
20. An antibody or antibody fragment that binds to dickkopf-1 (DKK1), said antibody or antibody fragment comprising a variable domain heavy chain region (VH) comprising complementarity determining regions CDRH1, CDRH2 and CDRH3, wherein (a) the amino acid sequence of said CDRH1 is as set forth in any one of SEQ ID NOs: 1 to 98, 919 to 1031 and 1238 to 1332, (b) the amino acid sequence of said CDRH2 is as set forth in any one of SEQ ID NOs: 99 to 196, 1333 to 1445 and 1652 to 1746, and (c) the amino acid sequence of said CDRH3 is as set forth in any one of SEQ ID NOs: 197 to 294, 1747 to 1859 and 2066 to 2160.
21. The antibody or antibody fragment described in claim 20, wherein the VH comprises an amino acid sequence that is at least 90% identical to an amino acid sequence set forth in any one of SEQ ID NOs: 295-392, 394-506, and 2164-2258.
22. An antibody or antibody fragment that binds to dickkopf-1 (DKK1), said antibody or antibody fragment comprising a variable domain heavy chain region (VH) comprising complementarity determining regions CDRH1, CDRH2, and CDRH3, wherein (a) the amino acid sequence of said CDRH1 is as set forth in any one of SEQ ID NOs: 1032 to 1237, (b) the amino acid sequence of said CDRH2 is as set forth in any one of SEQ ID NOs: 1446 to 1651, and (c) the amino acid sequence of said CDRH3 is as set forth in any one of SEQ ID NOs: 1860 to 2065. and a variable domain light chain region (VL) comprising complementarity determining regions CDRL1, CDRL2, and CDRL3, wherein (a) the amino acid sequence of CDRL1 is as set forth in any one of SEQ ID NOs: 2259 to 2464, (b) the amino acid sequence of CDRL2 is as set forth in any one of SEQ ID NOs: 2465 to 2521, and (c) the amino acid sequence of CDRL3 is as set forth in any one of SEQ ID NOs: 2522 to 2727.