Variant nucleic acid libraries for antibody optimization
By constructing a nucleic acid library containing multiple variant sequences, the problem of difficulty in balancing immunologic effects and other effects when designing therapeutic antibodies is solved, and the optimization of antibodies and the improvement of binding affinity is achieved.
Patent Information
- Application Number
- JP2025014222
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-04-05
- Filing Date
- 2025-01-30
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When designing therapeutic antibodies, it is difficult to balance immunologic effects with other effects, resulting in challenges in optimizing antibody properties.
By constructing a nucleic acid library containing multiple variant sequences, wherein each sequence corresponds to a predetermined antibody or antibody fragment, it is ensured that there are at least 5,000 variant sequences in the library, each sequence is expressed in the library no more than 50% of the other sequences, and at least one sequence has a higher binding affinity than the input sequence.
The optimization of the antibody is achieved, the binding affinity of the antibody is improved, and the function and stability of the antibody are expanded.
Smart Images

Figure 2025074080000001_ABST
Abstract
Description
[Technical field]
[0001] cross reference This application claims the benefit of U.S. Provisional Patent Application No. 62 / 810,379, filed February 26, 2019, and U.S. Provisional Patent Application No. 62 / 830,296, filed April 5, 2019, each of which is incorporated herein by reference in its entirety. [Background technology]
[0002] Antibodies have the ability to bind biological targets with high specificity and affinity. However, designing therapeutic antibodies is challenging due to the tradeoff between immunological action and efficacy. Therefore, there is a need to develop compositions and methods to optimize antibody properties.
[0003] Citation by reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. Summary of the Invention
[0004] Provided herein are methods, compositions, and systems for optimizing antibodies.
[0005] Provided herein is a nucleic acid library comprising a plurality of nucleic acids encoding an antibody or antibody fragment, each sequence of the plurality of sequences comprising a predetermined number of mutations relative to an input sequence, the library comprising at least 5,000 variant sequences, each sequence of the at least 5,000 variant sequences being represented in an amount no greater than 50% of the amount of other variant sequences in the library, and at least one sequence encoding an antibody or antibody fragment with a higher binding affinity than the input sequence. Further provided herein is a nucleic acid library, wherein the library comprises at least 50,000 variant sequences. Further provided herein is a nucleic acid library, wherein the library comprises at least 100,000 variant sequences. Further provided herein is a nucleic acid library, wherein at least a portion of the sequences encode an antibody light chain. Further provided herein is a nucleic acid library, wherein at least a portion of the sequences encode an antibody heavy chain. Further provided herein is a nucleic acid library, wherein each sequence of the plurality of sequences comprises at least one mutation in each CDR of the heavy or light chain relative to the input sequence. Further provided herein is a nucleic acid library, wherein each sequence of the plurality of sequences comprises at least two mutations in each CDR of the heavy or light chain relative to the input sequence. Further provided herein is a nucleic acid library, wherein at least one of the mutations is present in at least two individuals. Further provided herein is a nucleic acid library, wherein at least one of the mutations is present in at least three individuals. Further provided herein is a nucleic acid library, wherein each sequence of the plurality of sequences comprises at least one mutation in each CDR of the heavy or light chain relative to the germline sequence of the input sequence.
[0006] Provided herein is an antibody comprising a CDR-H3 comprising any one of the sequences of SEQ ID NOs: 1 to 35. Provided herein is an antibody comprising a CDR-H3 comprising any one of the sequences of SEQ ID NOs: 1 to 35, the antibody being a monoclonal antibody, a polyclonal antibody, a bispecific antibody, a multispecific antibody, a grafted antibody, a human antibody, a humanized antibody, a synthetic antibody, a chimeric antibody, a camelized antibody, a single-chain Fv (scFv), a single-chain antibody, a Fab fragment, a F(ab')2 fragment, a Fd fragment, a Fv fragment, a single domain antibody, an isolated complementarity determining region (CDR), a diabody, a fragment consisting of only a single monomeric variable domain, a disulfide-linked Fv (sdFv), an intrabody, an anti-idiotype (anti-Id) antibody, or an ab antigen-binding fragment thereof.
[0007] Provided herein is a method of inhibiting PD-1 activity, comprising administering an antibody as described herein. Provided herein is a method of treating a proliferative disorder, comprising administering an antibody as described herein to a subject. Provided herein is further a method of treating a proliferative disorder, wherein the proliferative disorder is cancer. Provided herein is further a method of treating a proliferative disorder, wherein the cancer is lung cancer, head and neck cancer, colon cancer, melanoma, liver cancer, classical Hodgkin's lymphoma, kidney cancer, gastric cancer, cervical cancer, Merkel cell, B cell lymphoma, or bladder cancer.
[0008] Provided herein is a nucleic acid library comprising a plurality of nucleic acids, each nucleic acid of the plurality of nucleic acids encoding a sequence that, upon translation, encodes an antibody, the antibody comprising a CDR-H3 loop comprising a PD-1-binding domain, and each nucleic acid of the plurality of nucleic acids comprising a sequence encoding a sequence variant of the PD-1-binding domain. Provided herein is a nucleic acid library, in which the length of the CDR-H3 loop upon translation is about 20 to about 80 amino acids. Provided herein is a nucleic acid library, in which the length of the CDR-H3 loop is about 80 to about 230 base pairs. Provided herein is a nucleic acid library, in which the antibody further comprises one or more domains selected from a light chain (VL) variable domain, a heavy chain (VH) variable domain, a light chain (CL) constant domain, and a heavy chain (CH) constant domain. Provided herein is a nucleic acid library, in which the length of the VH domain is about 90 to about 100 amino acids. Provided herein is a nucleic acid library, in which the length of the VL domain is about 90 to about 120 amino acids. Further provided herein is a nucleic acid library, wherein the VH domain is about 280 to about 300 base pairs in length. Further provided herein is a nucleic acid library, wherein the VL domain is about 300 to about 350 base pairs in length. Further provided herein is a nucleic acid library, wherein the library is at least 10 10 The present invention further provides a nucleic acid library, the library comprising at least 10 non-identical nucleic acids. 12 Further provided herein is a nucleic acid library, wherein the antibody comprises a single immunoglobulin domain. Further provided herein is a nucleic acid library, wherein the antibody comprises a peptide of at most 100 amino acids. Further provided herein is a nucleic acid library, wherein the PD-1 binding domain is a peptidomimetic or a small molecule mimetic.
[0009] Provided herein is a protein library comprising a plurality of proteins, each protein of the plurality of proteins comprising an antibody, the antibody comprising a CDR-H3 loop comprising a sequence variant of a PD-1 binding domain. Provided herein is a protein library comprising a plurality of proteins, the CDR-H3 loop being about 20 to about 80 amino acids in length. Provided herein is a protein library comprising a plurality of proteins, the antibody further comprising one or more domains selected from a light chain (VL) variable domain, a heavy chain (VH) variable domain, a light chain (CL) constant domain, and a heavy chain (CH) constant domain. Provided herein is a protein library comprising a plurality of proteins, the VH domain being about 90 to about 100 amino acids in length. Provided herein is a protein library comprising a plurality of proteins, the VL domain being about 90 to about 120 amino acids in length. Provided herein is a protein library comprising a plurality of proteins, the plurality of proteins being used to generate a peptidomimetic library. Provided herein is a protein library comprising a plurality of proteins, the library comprising an antibody.
[0010] Provided herein is a protein library comprising a plurality of proteins, the plurality of proteins comprising sequences encoding different PD-1-binding domains, each PD-1-binding domain being about 20 to about 80 amino acids in length. Further provided herein is a protein library comprising a plurality of proteins, wherein the protein library comprises peptides. Further provided herein is a protein library comprising a plurality of proteins, wherein the protein library comprises immunoglobulins. Further provided herein is a protein library comprising a plurality of proteins, wherein the protein library comprises antibodies. Further provided herein is a protein library comprising a plurality of proteins, wherein the plurality of proteins are used to generate a peptidomimetic library.
[0011] Provided herein is a vector library comprising a nucleic acid library as described herein. Provided herein is a cell library comprising a nucleic acid library as described herein. Provided herein is a cell library comprising a protein library as described herein.
[0012] Provided herein is a nucleic acid library comprising a plurality of nucleic acids encoding an antibody or antibody fragment, wherein each sequence of the plurality comprises a predetermined number of mutations relative to an input sequence, the library comprising at least 30,000 variant sequences, and at least some of the antibodies or antibody fragments have a K of less than 50 nM. D Further provided herein is a nucleic acid library, wherein the library comprises any one of CDR sequences of SEQ ID NOs: 1 to 35. Further provided herein is a nucleic acid library, wherein the library comprises any one of CDRH1, CDRH2, and CDRH3 sequences of SEQ ID NOs: 1 to 35. Further provided herein is a nucleic acid library, wherein the library comprises any one of CDRH1, CDRH2, and CDRH3 sequences of SEQ ID NOs: 1 to 35. Further provided herein is a nucleic acid library, wherein the library comprises any one of CDRH1, CDRH2, and CDRH3 sequences of SEQ ID NOs: 1 to 35, with a K D Provided herein is a nucleic acid library comprising at least one sequence encoding an antibody or antibody fragment that binds to PD-1 with a K of less than 5 nM. D Provided herein is a nucleic acid library comprising at least one sequence encoding an antibody or antibody fragment that binds to PD-1 with a K of less than 10 nM. D Provided herein is a nucleic acid library comprising at least five sequences encoding an antibody or antibody fragment that binds to PD-1 at a nucleotide sequence. Further provided herein is a nucleic acid library, wherein the library comprises at least 50,000 variant sequences. Further provided herein is a nucleic acid library, wherein the library comprises at least 100,000 variant sequences.
[0013] Provided herein is a computerized system for optimizing an antibody, the computerized system comprising: (a) a general purpose computer; and (b) a computer readable medium comprising a functional module comprising instructions for the general purpose computer, the computerized system configured to operate in the following manner: (i) receiving operational instructions, the operational instructions comprising a polynucleotide sequence encoding an antibody or antibody fragment; (ii) generating an antibody library, the antibody library comprising a plurality of variant sequences of the polynucleotide sequence; and (iii) synthesizing a plurality of variant sequences. Further provided herein is a computerized system for optimizing an antibody, the antibody library comprising at least 30,000 sequences. Further provided herein is a computerized system for optimizing an antibody, the antibody library comprising at least 50,000 sequences. Further provided herein is a computerized system for optimizing an antibody, the antibody library comprising at least 100,000 sequences. Further provided herein is a computerized system for optimizing an antibody, the computerized system further comprising enriching a subset of the variant sequences. Further provided herein is a computerized system for optimizing an antibody, further comprising expressing an antibody or antibody fragment corresponding to said variant sequence. Further provided herein is a computerized system for optimizing an antibody, wherein said polynucleotide sequence is a murine, human, or chimeric antibody sequence. Further provided herein is a computerized system for optimizing an antibody, wherein said antibody library comprises variant sequences, each of which is represented at 50% or less of the abundance of other variant sequences in said antibody library. Further provided herein is a computerized system for optimizing an antibody, wherein each sequence of said plurality of variant sequences comprises at least one mutation in each CDR of the heavy or light chain relative to said input sequence. Further provided herein is a computerized system for optimizing an antibody, wherein each sequence of said plurality of variant sequences comprises at least two mutations in each CDR of the heavy or light chain relative to said input sequence.Further provided herein is a computerized system for optimizing an antibody, wherein each sequence of the plurality of variant sequences comprises at least one mutation in each CDR of a heavy or light chain relative to a germline sequence of the input sequence. Further provided herein is a computerized system for optimizing an antibody, wherein the theoretical diversity of the antibody library is at least 10. 12 The present invention further provides a computerized system for optimizing an antibody, the antibody library having a theoretical diversity of at least 10 13 A computerized system for optimizing an antibody is provided, the sequence of which is:
[0014] Provided herein is a method of optimizing an antibody, comprising: (a) providing a polynucleotide sequence encoding an antibody or antibody fragment; (b) generating an antibody library, the antibody library comprising a plurality of variant sequences of a polynucleotide sequence; and (c) synthesizing a plurality of variant sequences. Provided herein is a method of optimizing an antibody, wherein the antibody library comprises at least 30,000 sequences. Provided herein is a method of optimizing an antibody, wherein the antibody library comprises at least 50,000 sequences. Provided herein is a method of optimizing an antibody, wherein the antibody library comprises at least 100,000 sequences. Provided herein is a method of optimizing an antibody, further comprising enriching a subset of the variant sequences. Provided herein is a method of optimizing an antibody, further comprising expressing an antibody or antibody fragment corresponding to the variant sequence. Provided herein is a method of optimizing an antibody, wherein the polynucleotide sequence is a murine, human, or chimeric antibody sequence. Further provided herein is a method of optimizing an antibody, wherein the antibody library comprises variant sequences, each of which is represented at 50% or less of the abundance of other variant sequences in the antibody library. Further provided herein is a method of optimizing an antibody, wherein each sequence of the plurality of variant sequences comprises at least one mutation in each CDR of the heavy or light chain relative to the input sequence. Further provided herein is a method of optimizing an antibody, wherein each sequence of the plurality of variant sequences comprises at least two mutations in each CDR of the heavy or light chain relative to the input sequence. Further provided herein is a method of optimizing an antibody, wherein each sequence comprises at least one mutation in each CDR of the heavy or light chain relative to the germline sequence of the input sequence. Further provided herein is a method of optimizing an antibody, wherein the theoretical diversity of the antibody library is at least 10 12 The present invention further provides a method for optimizing an antibody, the antibody library having a theoretical diversity of at least 10 13 The present invention provides a method for optimizing an antibody, the sequence of which is:
[0015] Provided herein is a nucleic acid library, the nucleic acid library comprising a plurality of sequences comprising a nucleic acid that upon translation encodes an antibody or antibody fragment, each of the plurality of sequences comprising a predetermined number of mutations in a CDR relative to an input sequence of an antibody, the library comprising at least 50,000 variant sequences, each sequence present in an amount within an average frequency of 1.5, and at least one sequence upon translation encodes an antibody or antibody fragment with a binding affinity at least 2.5 times higher than the binding infinity of the input sequence. Further provided herein is a nucleic acid library, the library comprising at least 100,000 variant sequences. Further provided herein is a nucleic acid library, at least a portion of the sequences encodes an antibody light chain. Further provided herein is a nucleic acid library, at least a portion of the sequences encodes an antibody heavy chain. Further provided herein is a nucleic acid library, each sequence of the plurality of sequences comprises at least one mutation in a CDR of the heavy or light chain relative to the input sequence. Further provided herein is a nucleic acid library, each sequence of the plurality of sequences comprises at least two mutations in a CDR of the heavy or light chain relative to the input sequence. Further provided herein is a nucleic acid library, wherein at least one of the mutations is present in at least two individuals. Further provided herein is a nucleic acid library, wherein at least one of the mutations is present in at least three individuals. Further provided herein is a nucleic acid library, wherein at least one sequence upon translation encodes an antibody or antibody fragment with a binding affinity at least 5 times higher than the binding infinity of the input sequence. Further provided herein is a nucleic acid library, wherein at least one sequence upon translation encodes an antibody or antibody fragment with a binding affinity at least 25 times higher than the binding infinity of the input sequence. Further provided herein is a nucleic acid library, wherein at least one sequence upon translation encodes an antibody or antibody fragment with a binding affinity at least 50 times higher than the binding infinity of the input sequence. Further provided herein is a nucleic acid library, wherein each sequence of the plurality of sequences comprises at least one mutation in a heavy or light chain CDR relative to the germline sequence of the input sequence.Further provided herein is a nucleic acid library, wherein the CDRs are CDR1, CDR2, and CDR3 on the heavy chain. Further provided herein is a nucleic acid library, wherein the CDRs are CDR1, CDR2, and CDR3 on the light chain. Further provided herein is a nucleic acid library, wherein at least one sequence upon translation encodes an antibody or antibody fragment with a binding affinity at least 70 times higher than the input sequence ... D The present invention further provides a nucleic acid library encoding an antibody or antibody fragment having a K of less than 50 nM. D The present invention further provides a nucleic acid library encoding an antibody or antibody fragment having a K of less than 25 nM. D The present invention further provides a nucleic acid library encoding an antibody or antibody fragment having a K of less than 10 nM. D Provided herein is a nucleic acid library encoding an antibody or antibody fragment having a K of less than 5 nM. Further provided herein is a nucleic acid library, wherein the library comprises a CDR sequence of any one of SEQ ID NOs: 1-6 or 9-70. Further provided herein is a nucleic acid library, wherein the library comprises a CDRH1, CDRH2, or CDRH3 sequence of any one of SEQ ID NOs: 1-6 or 9-70. Further provided herein is a nucleic acid library, wherein the library comprises an antibody or antibody fragment having a K of less than 10 nM. D Provided herein is a nucleic acid library comprising at least one sequence encoding an antibody or antibody fragment that binds to PD-1 with a K of less than 5 nM. D Provided herein is a nucleic acid library comprising at least one sequence encoding an antibody or antibody fragment that binds to PD-1 with a K of less than 10 nM. DProvided herein is a nucleic acid library comprising at least five sequences encoding an antibody or antibody fragment that binds to PD-1 at a target sequence. Further provided herein is a nucleic acid library, wherein the library comprises at least 100,000 variant sequences.
[0016] Provided herein is an antibody comprising any one of the sequences of SEQ ID NOs: 1 to 6 or 9 to 70. Provided herein is an antibody comprising any one of the sequences of SEQ ID NOs: 1 to 6 or 9 to 34, the antibody being a monoclonal antibody, a polyclonal antibody, a bispecific antibody, a multispecific antibody, a grafted antibody, a human antibody, a humanized antibody, a synthetic antibody, a chimeric antibody, a camelized antibody, a single chain Fv (scFv), a single chain antibody, a Fab fragment, a F(ab')2 fragment, an Fd fragment, an Fv fragment, a single domain antibody, an isolated complementarity determining region (CDR), a diabody, a fragment consisting of only a single monomeric variable domain, a disulfide-linked Fv (sdFv), an intrabody, an anti-idiotype (anti-Id) antibody, or an ab antigen-binding fragment thereof.
[0017] Provided herein is a method of inhibiting PD-1 activity, comprising administering an antibody described herein. Provided herein is a method of treating a proliferative disorder, comprising administering an antibody described herein to a subject. Provided herein further is a method, wherein the proliferative disorder is cancer. Provided herein further is a method, wherein the cancer is lung cancer, head and neck cancer, colon cancer, melanoma, liver cancer, classical Hodgkin's lymphoma, kidney cancer, gastric cancer, cervical cancer, Merkel cell, B-cell lymphoma, or bladder cancer.
[0018] Provided herein is a computerized system for optimizing an antibody, the computerized system comprising: (a) a general-purpose computer; and (b) a computer-readable medium comprising a functional module comprising instructions for the general-purpose computer, the computerized system configured to operate in the following manner: (i) receiving operational instructions, the operational instructions comprising a plurality of sequences encoding an antibody or antibody fragment; (ii) generating a nucleic acid library comprising a plurality of sequences comprising a nucleic acid that, when translated, encodes an antibody or antibody fragment, each of the plurality of sequences comprising a predetermined number of mutations in a CDR relative to an input sequence of an antibody, the library comprising at least 50,000 variant sequences, each sequence present at an amount within a 1.5-fold average frequency, and at least one sequence, when translated, encodes an antibody or antibody fragment with a binding affinity at least 2.5-fold higher than the binding affinity of the input sequence; and (iii) synthesizing at least 50,000 variant sequences. Further provided herein is a computerized system for optimizing an antibody, the nucleic acid library comprising at least 100,000 sequences. Further provided herein is a computerized system for optimizing an antibody, further comprising enriching a subset of said variant sequences. Further provided herein is a computerized system for optimizing an antibody, further comprising expressing an antibody or antibody fragment corresponding to said variant sequences. Further provided herein is a computerized system for optimizing an antibody, wherein said polynucleotide sequences are murine, human, or chimeric antibody sequences. Further provided herein is a computerized system for optimizing an antibody, wherein each sequence of said plurality of variant sequences comprises at least one mutation in a heavy or light chain CDR relative to said input sequence. Further provided herein is a computerized system for optimizing an antibody, wherein each sequence of said plurality of variant sequences comprises at least two mutations in a heavy or light chain CDR relative to said input sequence.Further provided herein is a computerized system for optimizing an antibody, wherein at least one sequence upon translation encodes an antibody or antibody fragment with a binding affinity at least 5 times higher than the binding infinity of the input sequence. Further provided herein is a computerized system for optimizing an antibody, wherein at least one sequence upon translation encodes an antibody or antibody fragment with a binding affinity at least 25 times higher than the binding infinity of the input sequence. Further provided herein is a computerized system for optimizing an antibody, wherein at least one sequence upon translation encodes an antibody or antibody fragment with a binding affinity at least 50 times higher than the binding infinity of the input sequence. Further provided herein is a computerized system for optimizing an antibody, wherein each sequence of the plurality of variant sequences comprises at least one mutation in a CDR of a heavy or light chain relative to a germline sequence of the input sequence. Further provided herein is a computerized system for optimizing an antibody, wherein the CDRs are CDR1, CDR2, and CDR3 on the heavy chain. Further provided herein is a computerized system for optimizing an antibody, wherein the CDRs are CDR1, CDR2, and CDR3 on the light chain. Further provided herein is that the theoretical diversity of the antibody library is at least 10. 12 The present invention further provides a computerized system for optimizing an antibody, the antibody library having a theoretical diversity of at least 10 13 A computerized system for optimizing an antibody is provided, the sequence of which is:
[0019] Provided herein is a method of optimizing an antibody, comprising: (a) providing a plurality of polynucleotide sequences encoding an antibody or antibody fragment; (b) generating a nucleic acid library comprising a plurality of sequences comprising nucleic acids that when translated encode an antibody or antibody fragment, each of the plurality of sequences comprising a predetermined number of mutations in CDRs relative to an input sequence of an antibody, the library comprising at least 50,000 variant sequences, each sequence present at an amount within a 1.5-fold average frequency, and at least one sequence when translated encodes an antibody or antibody fragment with a binding affinity at least 2.5-fold higher than the binding affinity of the input sequence; and (c) synthesizing at least 50,000 variant sequences. Provided herein is a method of optimizing an antibody, wherein the antibody library comprises at least 100,000 sequences. Provided herein is a method of optimizing an antibody, further comprising enriching a subset of the variant sequences. Provided herein is a method of optimizing an antibody, further comprising expressing an antibody or antibody fragment corresponding to the variant sequences. Further provided herein is a method of optimizing an antibody, wherein the polynucleotide sequence is a murine, human, or chimeric antibody sequence. Further provided herein is a method of optimizing an antibody, wherein each sequence of the plurality of variant sequences comprises at least one mutation in each CDR of the heavy or light chain relative to the input sequence. Further provided herein is a method of optimizing an antibody, wherein each sequence of the plurality of variant sequences comprises at least two mutations in each CDR of the heavy or light chain relative to the input sequence. Further provided herein is a method of optimizing an antibody, wherein at least one sequence upon translation encodes an antibody or antibody fragment with a binding affinity at least 5-fold higher than the binding infinity of the input sequence. Further provided herein is a method of optimizing an antibody, wherein at least one sequence upon translation encodes an antibody or antibody fragment with a binding affinity at least 25-fold higher than the binding infinity of the input sequence.Further provided herein is a method of optimizing an antibody, wherein at least one sequence upon translation encodes an antibody or antibody fragment with a binding affinity at least 50-fold higher than the binding affinity of the input sequence. Further provided herein is a method of optimizing an antibody, wherein each sequence comprises at least one mutation in each CDR of the heavy or light chain relative to the germline sequence of the input sequence. Further provided herein is a method of optimizing an antibody, wherein the theoretical diversity of the nucleic acid library is at least 10. 12 The present invention further provides a method for optimizing an antibody, the nucleic acid library having a theoretical diversity of at least 10 13 The present invention provides a method for optimizing an antibody, the sequence of which is: [Brief description of the drawings]
[0020] [Figure 1] 1 depicts the workflow for optimizing an antibody. [Figure 2A] FIG. 1 depicts a first schematic diagram of an immunoglobulin scaffold. [Figure 2B] FIG. 2 depicts a second schematic diagram of an immunoglobulin scaffold. [Figure 3A] 1 depicts exemplary sequences of optimized antibody input sequences in a library, with the number of mutations within regions of the sequence indicated. [Figure 3B] 1 shows the workflow for optimizing an antibody. [Figure 4A] The read length of the variable heavy chain after 1 to 5 rounds of panning is shown. [Figure 4B] The clonal frequencies of variable heavy chains after 1 to 5 rounds of panning are shown. [Figure 4C] The clonal accumulation of variable heavy chains after 1 to 5 rounds of panning is shown. [Figure 4D] 1 is a graph of the number of mutations for various panning conditions. [Figure 4E] Graph of sequence analysis of clones enriched for binding to PD-1. [Figure 5A] Plot of anti-scFv ELISA versus enrichment at the 5th panning time point. [Figure 5B] 13 is a graph of anti-scFv ELISA binding to PD-1. [Figure 6A-1] 1 shows the sequence of an optimized IgG binding to control PD-1. [Figure 6A-2] 1 shows the sequence of an optimized IgG binding to control PD-1. [Figure 6B] The increase in affinity of the optimized antibody (4.5 nM) compared to the parent antibody (330 nM) is shown. [Figure 6C] The sequence alignment of the CDRs is shown. [Figure 6D] The sequence alignment of the CDRs is shown. [Figure 6E] FIG. 1 depicts a dose-dependent phage PD-1 ELISA. [Figure 6F] Iso-affinity plots are presented. [Figure 7A] PD-1 / PDL-1 blocking assay of optimized IgG. [Figure 7B] Optimized IgG binding affinity and potency are shown. [Figure 7C] Graph of PD-1 / PDL-1 blockade IC50 (nM, on the Y-axis) versus SPR monovalent binding affinity (Kd (nM, on the X-axis) of optimized IgG. [Figure 7D] Graph of BVP scores (on the y-axis) for various IgGs (on the x-axis). [Figure 8] A process diagram is presented demonstrating a typical process workflow for gene synthesis as disclosed herein. [Figure 9] 1 illustrates an example computer system. [Figure 10] FIG. 1 is a block diagram illustrating the structure of a computer system. [Figure 11] 1 is a diagram demonstrating a network configured to incorporate multiple computer systems, multiple mobile phones, personal digital assistants, and network attached storage (NAS). [Figure 12]1 is a block diagram of a multiprocessor computer system using a shared virtual address memory space. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0021] This disclosure employs, unless otherwise specified, conventional molecular biology techniques within the skill of the art. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.
[0022] definition
[0023] Throughout this disclosure, various embodiments are presented in a range format. It should be understood that the description in range format is for convenience and brevity only and should not be construed as an inflexible limitation on the scope of any embodiment. Thus, the description of a range should be considered to have specifically disclosed all possible subranges, as well as each individual numerical value within that range, to the tenth of the unit of the lower limit, unless otherwise clear from the context. For example, the description of a range such as 1-6 should be considered to have specifically disclosed subranges such as 1-3, 1-4, 1-5, 2-4, 2-6, 3-6, etc., as well as each individual numerical value within that range, e.g., 1.1, 2, 2.3, 5, 5.9. This applies regardless of the breadth of the range. The upper and lower limits of these intervening ranges may be independently included in the smaller ranges, and are also encompassed within the disclosure, subject to any specifically excluded limit of the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure, unless the context makes clear otherwise.
[0024] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit any embodiment. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless otherwise clear from the context. It is further understood that, as used herein, the terms "comprises" and / or "comprising" specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0025] Unless otherwise specified or clear from the context, as used herein, the term "about" in connection with a number or range of numbers means the specified number and + / - 10% of that number, or, for any recited value in a range, less than 10% of the recited lower limit and more than 10% of the recited upper limit.
[0026] Unless otherwise specified, the term "nucleic acid" as used herein encompasses double-stranded or triple-stranded nucleic acids as well as single-stranded molecules. In double-stranded or triple-stranded nucleic acids, the strands of the nucleic acid need not be coextensive (a double-stranded nucleic acid need not be double-stranded along the entire length of both strands). Unless otherwise specified, nucleic acid sequences, as provided, are listed in a 5' to 3' orientation. The methods described herein provide for the production of isolated nucleic acids. The methods described herein further provide for the production of isolated and purified nucleic acids. As referred to herein, a "nucleic acid" may comprise at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000 or more bases in length. Further provided herein are methods for the synthesis of any number of polypeptide segments encoding nucleotide sequences, including sequences encoding non-ribosomal peptides (NRPs), sequences encoding non-ribosomal peptide synthetase (NRPS) modules and synthetic variants, polypeptide segments of other regulatory proteins such as antibodies, polypeptide segments from other protein families that contain non-coding DNA or RNA, such as regulatory sequences, e.g., promoters, transcription factors, enhancers, siRNAs, shRNAs, RNAi, miRNAs, small nucleolar RNAs derived from microRNAs, or any functional or structural DNA or RNA unit of interest.The following are non-limiting examples of polynucleotides: coding or non-coding regions of a gene or gene fragment, intergenic DNA, loci determined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, small interfering RNA (siRNA), small hairpin RNA (shRNA), microRNA (miRNA), small nucleolar RNA, ribozymes, complementary DNA (cDNA), which is a DNA representation of messenger RNA (mRNA), usually obtained by reverse transcription or amplification of mRNA, DNA molecules produced synthetically or by amplification, genomic DNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. The cDNA encoding the gene or gene fragment referred to herein may contain at least one region that encodes an exon sequence without the intron sequence intervening in the genomic equivalent sequence.
[0027] Antibody Optimization
[0028] Provided herein are methods, compositions, and systems for optimizing antibodies. In some examples, antibodies are optimized by designing an in silico library containing variant sequences of an input antibody sequence (FIG. 1). In some examples, the input sequence (100) is modified in silico (102) with one or more mutations that generate a library of optimized sequences (103). In some examples, such libraries are synthesized and cloned into expression vectors, and the translation products (antibodies) are evaluated for activity. In some examples, fragments of the sequences are synthesized and then assembled. In some examples, expression vectors are used to display and enrich for desired antibodies, such as phage display. In some examples, selection pressures used during enrichment include binding affinity, toxicity, immune tolerance, stability, or other factors. Such expression vectors allow for the selection of antibodies with specific properties ("panning"), and subsequent propagation or amplification of such sequences enriches the library with these sequences. Panning can be repeated any number of times, such as 1, 2, 3, 4, 5, 6, 7, or more than 7 times. In some instances, one or more rounds of sequencing are used to identify which sequences (105) are enriched in the library.
[0029] Described herein are methods and systems for designing in silico libraries. For example, antibody or antibody fragment sequences are used as input. In some examples, any antibody sequence is used as input for the methods and systems described herein. A database (102) containing known mutations from organisms is queried (101) to generate a library (103) of sequences containing combinations of these mutations. In some examples, the antibodies described herein contain CDR regions. In some examples, known mutations from the CDRs are used to construct the sequence library. In some examples, filters (104), i.e., exclusion criteria, are used to select for certain types of variants for members of the sequence library. For example, sequences with mutations are added if they are present in a minimum number of organisms in the database. In some examples, additional CDRs are identified for inclusion in the database. In some examples, certain mutations, or combinations of mutations, are excluded from the library (e.g., known immunogenic sites, structural sites, etc.). In some examples, certain sites in the input sequences are systematically replaced with histidine, aspartic acid, glutamic acid, or combinations thereof. In some examples, the maximum or minimum number of mutations allowed for each region of the antibody are specified. In some examples, the mutations are described relative to the input sequence or the corresponding germline sequence of the input sequence. For example, the sequence generated by optimization contains at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or more than 16 mutations from the input sequence. In some examples, the sequence generated by optimization contains no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or 18 mutations from the input sequence. In some examples, the sequence generated by optimization contains about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or 18 mutations from the input sequence. In some examples, the sequence generated by optimization contains about 1, 2, 3, 4, 5, 6, or 7 mutations from the input sequence in the first CDR region.In some examples, the sequences generated by optimization include about 1, 2, 3, 4, 5, 6, or 7 mutations from the input sequence in the second CDR region. In some examples, the sequences generated by optimization include about 1, 2, 3, 4, 5, 6, or 7 mutations from the input sequence in the third CDR region. In some examples, the sequences generated by optimization include about 1, 2, 3, 4, 5, 6, or 7 mutations from the input sequence in the first CDR region of the heavy chain. In some examples, the sequences generated by optimization include about 1, 2, 3, 4, 5, 6, or 7 mutations from the input sequence in the second CDR region of the heavy chain. In some examples, the sequences generated by optimization include about 1, 2, 3, 4, 5, 6, or 7 mutations from the input sequence in the third CDR region of the heavy chain. In some examples, the sequences generated by optimization include about 1, 2, 3, 4, 5, 6, or 7 mutations from the input sequence in the first CDR region of the light chain. In some examples, the sequences generated by optimization contain about 1, 2, 3, 4, 5, 6, or 7 mutations from the input sequence in the second CDR region of the light chain. In some examples, the sequences generated by optimization contain about 1, 2, 3, 4, 5, 6, or 7 mutations from the input sequence in the third CDR region of the light chain. In some examples, the first CDR region is CDR1. In some examples, the second CDR region is CDR2. In some examples, the third CDR region is CDR3. In silico antibody libraries are synthesized, assembled, and enriched for desired sequences in some examples.
[0030] The germline sequence corresponding to the input sequence may also be modified to generate sequences in the library. For example, the sequences generated by the optimization methods described herein contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or more than 16 mutations from the germline sequence. In some examples, the sequences generated by the optimization contain no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or 18 mutations from the germline sequence. In some examples, the sequences generated by the optimization contain about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or 18 mutations relative to the germline sequence.
[0031] Provided herein are methods, systems, and compositions for optimizing an antibody, where the input sequence includes mutations in an antibody region. Exemplary regions of an antibody include, but are not limited to, a complementarity determining region (CDR), a variable domain, or a constant domain. In some examples, the CDR is a CDR1, CDR2, or CDR3. In some examples, the CDR is a heavy domain, including, but not limited to, CDR-H1, CDR-H2, and CDR-H3. In some examples, the CDR is a light domain, including, but not limited to, CDR-L1, CDR-L2, and CDR-L3. In some examples, the variable domain is a light chain variable (VL) domain or a heavy chain variable (VH) domain. In some examples, the VL domain includes a kappa chain or a lambda chain. In some examples, the constant domain is a light chain constant (CL) domain or a heavy chain constant (CH) domain. In some examples, the sequence generated by optimization comprises about 1, 2, 3, 4, 5, 6, or 7 mutations from the germline sequence in the first CDR region. In some examples, the sequence generated by optimization comprises about 1, 2, 3, 4, 5, 6, or 7 mutations from the germline sequence in the second CDR region. In some examples, the sequence generated by optimization comprises about 1, 2, 3, 4, 5, 6, or 7 mutations from the germline sequence in the third CDR region. In some examples, the sequence generated by optimization comprises about 1, 2, 3, 4, 5, 6, or 7 mutations from the germline sequence in the first CDR region of the heavy chain. In some examples, the sequence generated by optimization comprises about 1, 2, 3, 4, 5, 6, or 7 mutations from the germline sequence in the second CDR region of the heavy chain. In some examples, the sequence generated by optimization comprises about 1, 2, 3, 4, 5, 6, or 7 mutations from the germline sequence in the third CDR region of the heavy chain. In some examples, the sequence generated by optimization contains about 1, 2, 3, 4, 5, 6, or 7 mutations from the germline sequence in the first CDR region of the light chain. In some examples, the sequence generated by optimization contains about 1, 2, 3, 4, 5, 6, or 7 mutations from the germline sequence in the second CDR region of the light chain.In some examples, the sequence generated by optimization contains about 1, 2, 3, 4, 5, 6, or 7 mutations from the germline sequence in the third CDR region of the light chain. In some examples, the first CDR region is CDR1. In some examples, the second CDR region is CDR2. In some examples, the third CDR region is CDR3.
[0032] Antibody Libraries
[0033] Provided herein are libraries generated from the antibody optimization methods described herein, which provide improved functional activity, structural stability, expression, specificity, or a combination thereof.
[0034] As used herein, the term antibody is understood to include proteins with the two Y-shaped configurations characteristic of typical antibody molecules, as well as one or more fragments of antibodies capable of specifically binding to an antigen. Exemplary antibodies include monoclonal antibodies, polyclonal antibodies, bispecific antibodies, multispecific antibodies, graft antibodies, human antibodies, humanized antibodies, synthetic antibodies, chimeric antibodies, camelized antibodies, single chain Fvs (scFv) (including fragments in which the VL and VH are linked using recombinant methods with synthetic or natural linkers, which allows the fragments to be produced as a single protein chain in which the VL and VH regions are paired to form monovalent molecules, including single chain Fab and scFab), single chain antibodies, Fab fragments (including monovalent fragments containing the VL, VH, CL, and CH1 domains), F(ab')2 fragments (including two Fab fragments linked by a disulfide bridge at the hinge region). Examples of antibody fragments include, but are not limited to, bivalent fragments including a VH fragment, a Fd fragment (including a fragment including a VH and a CH1 fragment), an Fv fragment (including a fragment including a VL and VH domain of a single group of antibodies), a single domain antibody (dAb or sdAb) (including a fragment including a VH domain), an isolated complementarity determining region (CDR)), a diabody (including a fragment including a bivalent dimer such as two VL and VH domains that are linked together and recognize two different antigens), a fragment consisting of only a single monomeric variable domain, a disulfide-linked Fv (sdFv), an intrabody, an anti-idiotypic (anti-Id) antibody, or an ab antigen-binding fragment thereof. In some examples, the libraries disclosed herein include nucleic acids encoding antibodies, where the antibodies are Fv antibodies, including Fv antibodies consisting of the smallest antibody fragment that includes a complete antigen recognition site and an antigen binding site. In some embodiments, said Fv antibody consists of a dimer of one heavy chain variable domain and one light chain variable domain in tight, non-covalent association, with the three hypervariable regions of each variable domain interacting to define an antigen-binding site on the surface of the VH-VL dimer.In some embodiments, the six hypervariable regions confer antigen-binding specificity to the antibody. In some embodiments, the single variable domain (or half of an Fv containing only three hypervariable regions specific for an antigen, including single domain antibodies isolated from camelids containing one heavy chain variable domain, such as VHH antibodies or nanobodies) is capable of recognizing and binding to an antigen. In some examples, the libraries disclosed herein include nucleic acids encoding antibodies, where the antibodies are single chain Fvs or scFvs, including antibody fragments containing VH, VL, or both VH and VL domains, where both domains are present in a single polypeptide chain. In some embodiments, the Fv polypeptide further comprises a polypeptide linker between the VH and VL domains, which allows the scFv to form the desired structure for antigen binding. In some examples, the scFv is joined to an Fc fragment, or the VHH is joined to an Fc fragment (including a minibody). In some examples, the antibodies include immunoglobulin molecules, and immunologically active fragments of immunoglobulin molecules, such as molecules that contain an antigen-binding site. Immunoglobulin molecules can be of any type (e.g., IgG, IgE, IgM, IgD, IgA, and IgY), class (e.g., IgG1, IgG2, IgG3, IgG4, IgA1, and IgA2), or subclass of molecule.
[0035] In some embodiments, the library contains immunoglobulins appropriate for the species of the intended therapeutic target. Overall, these methods include "mammalization" and include methods of transferring donor antigen binding information to less immunogenic mammalian antibody acceptors to generate useful therapeutic treatments. In some examples, the mammals are mice, rats, horses, sheep, cows, primates (e.g., chimpanzees, baboons, gorillas, orangutans, monkeys), dogs, cats, pigs, donkeys, rabbits, and humans. In some examples, methods are provided herein for felinization and caninization of antibodies.
[0036] "Humanized" forms of non-human antibodies may be chimeric antibodies that include minimal sequence derived from a non-human antibody. Humanized antibodies are generally human antibodies (recipient antibodies) in which residues from one or more CDRs are replaced with residues from one or more CDRs of a non-human antibody (donor antibody). The donor antibody may be any suitable non-human antibody, such as a mouse, rat, rabbit, chicken, or non-human primate antibody with the desired specificity, affinity, or biological effect. In some instances, selected framework region residues of the recipient antibody are replaced with the corresponding framework region residues of the donor antibody. Humanized antibodies may also include residues that are not found in either the recipient antibody or the donor antibody. In some instances, these modifications are made to further refine antibody performance.
[0037] "Caninization" may include methods of transferring non-canine antigen-binding information from a donor antibody to a less immunogenic canine antibody acceptor to generate a treatment useful as a therapeutic for canines. In some examples, the caninized forms of non-canine antibodies provided herein are chimeric antibodies that contain minimal sequence from the non-canine antibody. In some examples, the caninized antibody is a canine antibody sequence (the "acceptor" or "recipient" antibody) in which hypervariable region residues of the recipient are replaced with hypervariable region residues of a non-canine species (the "donor" antibody), such as mouse, rat, rabbit, cat, dog, goat, chicken, cow, horse, llama, camel, dromedary, shark, non-human primate, human, humanized, recombinant, or engineered sequences with desired properties. In some examples, framework region (FR) residues of the canine antibody are replaced with corresponding non-canine FR residues. In some examples, the caninized antibody includes residues that are not found in the recipient antibody or the donor antibody. In some instances, these modifications are made to further refine antibody performance. A caninized antibody may also contain at least a portion of the immunoglobulin constant region (Fc) of a canine antibody.
[0038] "Fenconization" may include methods of transferring non-feline antigen binding information from a donor antibody to a less immunogenic feline antibody acceptor to generate treatments useful as therapeutics for cats. In some examples, the felineized forms of non-feline antibodies provided herein are chimeric antibodies that contain minimal sequence from the non-feline antibody. In some examples, felineized antibodies are feline antibody sequences (the "acceptor" or "recipient" antibody) in which hypervariable region residues of the recipient are replaced with hypervariable region residues of a non-feline species (the "donor" antibody), such as mouse, rat, rabbit, cat, dog, goat, chicken, cow, horse, llama, camel, dromedary, shark, non-human primate, human, humanized, recombinant, or engineered sequences with desired properties. In some examples, framework region (FR) residues of the feline antibody are replaced with corresponding non-feline FR residues. In some examples, the felineized antibody contains residues that are not found in the recipient antibody or the donor antibody. In some examples, these modifications are made to further refine antibody performance. The fetinized antibody may also comprise at least a portion of the immunoglobulin constant region (Fc) of a feline antibody.
[0039] The methods as described herein may be used to optimize libraries encoding non-immunoglobulins. In some embodiments, the libraries include antibody mimetics. Exemplary antibody mimetics include, but are not limited to, anticalins, affilins, affibody molecules, affimers, affitins, alphabodies, avimers, atrimers, DARPins, finomers, Kunitz domain-based proteins, monobodies, anticalins, knottins, armadillo repeat protein-based proteins, and bicyclic peptides.
[0040] The libraries described herein, which include nucleic acids encoding antibodies, include mutations in at least one region of the antibody. Exemplary regions of an antibody for mutation include, but are not limited to, a complementarity determining region (CDR), a variable domain, or a constant domain. In some examples, the CDR is CDR1, CDR2, or CDR3. In some examples, the CDR is a heavy domain, including, but not limited to, CDR-H1, CDR-H2, and CDR-H3. In some examples, the CDR is a light domain, including, but not limited to, CDR-L1, CDR-L2, and CDR-L3. In some examples, the variable domain is a light chain variable (VL) domain or a heavy chain variable (VH) domain. In some examples, the VL domain includes a kappa chain or a lambda chain. In some examples, the constant domain is a light chain constant (CL) domain or a heavy chain constant (CH) domain.
[0041] The methods described herein synthesize libraries containing antibody-encoding nucleic acids, each nucleic acid encoding a predefined variant of at least one predefined reference nucleic acid sequence. Optionally, the predefined reference sequence is a protein-encoding nucleic acid sequence, and the variant library comprises sequences encoding at least one codon mutation, such that multiple different variants of a single residue in the subsequent protein encoded by the synthesized nucleic acid are generated by standard translation processes. In some examples, the antibody library comprises variant nucleic acids that collectively encode mutations at multiple positions. In some examples, the variant library comprises sequences encoding at least a single codon mutation of CDR-H1, CDR-H2, CDR-H3, CDR-L1, CDR-L2, CDR-L3, VL domain, or VH domain. In some examples, the variant library comprises sequences encoding multiple codon mutations of CDR-H1, CDR-H2, CDR-H3, CDR-L1, CDR-L2, CDR-L3, VL domain, or VH domain. In some examples, the position variant library includes sequences encoding mutations of multiple codons in framework element 1 (FW1), framework element 2 (FW2), framework element 3 (FW3), or framework element 4 (FW4). Exemplary numbers of codons for mutation include, but are not limited to, at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 225, 250, 275, 300, or more than 300 codons.
[0042] In some instances, at least one region of the antibody to which the mutation is directed is derived from a heavy chain V gene family, a heavy chain D gene family, a heavy chain J gene family, a light chain V gene family, or a light chain J gene family. In some instances, the light chain V gene family comprises an immunoglobulin kappa (IGK) gene or an immunoglobulin lambda (IGL).
[0043] Provided herein are libraries containing nucleic acids encoding antibodies, the libraries being composed of a variable number of fragments. In some examples, the fragments include CDR-H1, CDR-H2, CDR-H3, CDR-L1, CDR-L2, CDR-L3, a VL domain, or a VH domain. In some examples, the fragments include framework element 1 (FW1), framework element 2 (FW2), framework element 3 (FW3), or framework element 4 (FW4). In some examples, the antibody library is composed of at least or about 2 fragments, 3 fragments, 4 fragments, 5 fragments, or more than 5 fragments. The length of each nucleic acid fragment of the nucleic acid to be synthesized, or the average length of the nucleic acid, may be at least or about 50, 75, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 525, 550, 575, 600, or more than 600 base pairs. In some examples, the length is about 50-600, 75-575, 100-550, 125-525, 150-500, 175-475, 200-450, 225-425, 250-400, 275-375, or 300-350 base pairs.
[0044] The library comprising nucleic acids encoding antibodies as described herein comprises amino acids of various lengths upon translation. In some examples, the length of each amino acid fragment of the synthesized amino acids, or the average length of the amino acids, may be at least or about 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, or more than 150 amino acids. In some examples, the length is about 15-150, 20-145, 25-140, 30-135, 35-130, 40-125, 45-120, 50-115, 55-110, 60-110, 65-105, 70-100, or 75-95 amino acids. In some examples, the length is about 22 to about 75 amino acids. In some examples, the antibody comprises at least or about 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, or more than 5000 amino acids.
[0045] A number of variant sequences for at least one region of the antibody for mutation are de novo synthesized using the method as described herein. In some examples, the number of variant sequences are de novo synthesized for CDR-H1, CDR-H2, CDR-H3, CDR-L1, CDR-L2, CDR-L3, VL, VH, or a combination thereof. In some examples, the number of variant sequences are de novo synthesized for framework element 1 (FW1), framework element 2 (FW2), framework element 3 (FW3), or framework element 4 (FW4). See Figure 2A. The number of variant sequences may be at least or about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, or more than 500 sequences. In some examples, the number of variant sequences is at least or about 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, or more than 8000 sequences. In some examples, the number of variant sequences is about 10-500, 25-475, 50-450, 75-425, 100-400, 125-375, 150-350, 175-325, 200-300, 225-375, 250-350, or 275-325 sequences.
[0046] In some examples, the length or sequence of the variant sequence for at least one region of the antibody is varied. In some examples, the at least one region synthesized de novo is directed to CDR-H1, CDR-H2, CDR-H3, CDR-L1, CDR-L2, CDR-L3, VL, VH, or a combination thereof. In some examples, the at least one region synthesized de novo is directed to framework element 1 (FW1), framework element 2 (FW2), framework element 3 (FW3), or framework element 4 (FW4). In some examples, the variant sequence comprises at least or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 or more variant nucleotides or amino acids compared to the wild type. In some examples, the variant sequences contain at least or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 additional nucleotides or amino acids compared to the wild type. In some examples, the variant sequences contain at least or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 fewer nucleotides or amino acids compared to the wild type. In some examples, the library contains at least or about 10 1 , 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , or 10 10 Includes more variants.
[0047] After the antibody library is synthesized, the antibody library may be used for screening and analysis. For example, the antibody library is analyzed for the display ability and panning of the library. In some examples, the display ability is analyzed using a selectable tag. Exemplary tags include, but are not limited to, radioactive labels, fluorescent labels, enzymes, chemiluminescent tags, colorimetric tags, affinity tags, or other labels or tags known in the art. In some examples, the tag is histidine, polyhistidine, myc, hemagglutinin (HA), or FLAG. In some examples, the antibody library is analyzed by sequencing using a variety of methods, including, but not limited to, single molecule real-time (SMRT) sequencing, polony sequencing, sequencing by ligation, reversible terminator sequencing, proton detection sequencing, ion semiconductor sequencing, nanopore sequencing, electronic sequencing, pyrosequencing, Maxam-Gilbert sequencing, chain termination (e.g., Sanger) sequencing, +S sequencing, or sequencing by synthesis. In some examples, the antibody library is displayed on the surface of a cell or phage. In some examples, the antibody library is enriched for sequences with the desired activity using phage display.
[0048] In some examples, antibody libraries are analyzed for functional activity, structural stability (e.g., thermal or pH stability), expression, specificity, or a combination thereof. In some examples, antibody libraries are analyzed for foldable antibodies. In some examples, regions of antibodies are analyzed for functional activity, structural stability, expression, specificity, folding, or a combination thereof. For example, VH or VL regions are analyzed for functional activity, structural stability, expression, specificity, folding, or a combination thereof.
[0049] The antibodies or IgGs generated by the methods described herein include improved binding affinities. In some examples, the antibodies include binding affinities (e.g., kD) of less than 1 nM, less than 1.2 nM, less than 2 nM, less than 5 nM, less than 10 nM, less than 11 nM, less than 13.5 nM, less than 15 nM, less than 20 nM, less than 25 nM, or less than 30 nM. In some examples, the antibodies include a kD of less than 1 nM. In some examples, the antibodies include a kD of less than 1.2 nM. In some examples, the antibodies include a kD of less than 2 nM. In some examples, the antibodies include a kD of less than 5 nM. In some examples, the antibodies include a kD of less than 10 nM. In some examples, the antibodies include a kD of less than 13.5 nM. In some examples, the antibodies include a kD of less than 15 nM. In some examples, the antibodies include a kD of less than 20 nM. In some examples, the antibodies include a kD of less than 25 nM. In some examples, the antibodies include a kD of less than 30 nM.
[0050] In some examples, the affinity of the antibody or IgG produced by the methods described herein is at least or about 1.5-fold, 2.0-fold, 5-fold, 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold, or 200-fold or more improved binding affinity compared to a control drug antibody. In some examples, the affinity of the antibody or IgG produced by the methods described herein is at least or about 1.5-fold, 2.0-fold, 5-fold, 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold, or 200-fold or more improved function compared to a control drug antibody. In some examples, the control drug antibody is an antibody with a similar structure, sequence, or antigen target.
[0051] Expression system
[0052] Provided herein are libraries containing nucleic acids encoding antibodies comprising binding domains, where the libraries have improved specificity, stability, expression, folding, or downstream activity. In some examples, the libraries described herein are used for screening and analysis.
[0053] A library comprising nucleic acids encoding antibodies comprising binding domains is provided herein, and the nucleic acid library is used for screening and analysis. In some examples, the screening and analysis includes in vitro, in vivo, or ex vivo assays. The cells for screening include primary cells taken from a living subject or cell line. The cells may be derived from prokaryotic cells (e.g., bacteria and fungi) or eukaryotic cells (e.g., animals and plants). Exemplary animal cells include, but are not limited to, animal cells from mice, rabbits, primates, and insects. In some examples, the cells for screening include cell lines including, but are not limited to, Chinese hamster ovary (CHO) cell lines, human embryonic kidney (HEK) cell lines, or baby hamster kidney (BHK) cell lines. In some examples, the nucleic acid library described herein may also be delivered to a multicellular organism. Exemplary multicellular organisms include, but are not limited to, plants, mice, rabbits, primates, and insects.
[0054] The nucleic acid library described herein can be screened for various pharmaceutical or pharmacokinetic properties.In some examples, the library is screened using in vitro, in vivo, or ex vivo assays.For example, the in vitro pharmaceutical or pharmacokinetic properties that are screened include, but are not limited to, binding affinity, binding specificity, and binding activity.The typical in vivo pharmaceutical or pharmacokinetic properties that are screened for the library described herein include, but are not limited to, therapeutic effect, activity, preclinical toxicity properties, clinical efficacy properties, clinical toxicity properties, immunogenicity, efficacy, and clinical safety properties.
[0055] A nucleic acid library is provided herein, and the nucleic acid library may be expressed in a vector. The expression vector for inserting the nucleic acid library disclosed herein may include eukaryotic or prokaryotic expression vectors. Exemplary expression vectors include, but are not limited to, mammalian expression vectors: pSF-CMV-NEO-NH2-PPT-3XFLAG, pSF-CMV-NEO-COOH-3XFLAG, pSF-CMV-PURO-NH2-GST-TEV, pSF-OXB20-COOH-TEV-FLAG(R)-6His, pCEP4 pDEST27, pSF-CMV-Ub-KrYFP, pSF-CMV-FMDV-daGFP, pEF1a-mCherry-N1 Vector, pEF1a-tdTomato Vector, pSF-CMV-FMDV-Hygro, pSF-CMV-PGK-Puro, pMCP-tag(m), and pSF-CMV-PURO-NH2-CMYC, bacterial expression vectors: pSF-OXB20-BetaGal, pSF-OXB20-Fluc, pSF-OXB20, and pSF-Tac, plant expression vectors: pRI 101-AN DNA and pCambia2301, and yeast expression vectors: pTYB21 and pKLAC2, and insect vectors: pAc5.1 / V5-His A and pDEST8. In some examples, the vector is pcDNA3 or pcDNA3.1.
[0056] The nucleic acid library described herein is expressed in a vector to generate a construct comprising an antibody.In some examples, the size of the construct varies.In some examples, the construct comprises at least or about 500, 600, 700, 800, 900, 1000, 1100, 1300, 1400, 1500, 1600, 1700, 1800, 2000, 2400, 2600, 2800, 3000, 3200, 3400, 3600, 3800, 4000, 4200, 4400, 4600, 4800, 5000, 6000, 7000, 8000, 9000, 10000, or more than 10000 bases. In some examples, the construct is about 300-1,000, 300-2,000, 300-3,000, 300-4,000, 300-5,000, 300-6,000, 300-7,000, 300-8,000, 300-9,000, 300-10,000, 1,000-2,000, 1,000-3,000, 1,000-4,000, 1,000-5,000, 1,000 ~6,000, 1,000~7,000, 1,000~8,000, 1,000~9,000, 1,000~10,000, 2,000~3,000, 2,000~4,000, 2,000~5,000, 2,000~6,000, 2,000~7,000, 2,000~8,000, 2,000~9,000, 2,000~10,000, 3,000~4,000, 3,00 0~5,000, 3,000~6,000, 3,000~7,000, 3,000~8,000, 3,000~9,000, 3,000~10,000, 4,000~5,000, 4,000~6,000, 4,000~7,000, 4,000~8,000, 4,000~9,000, 4,000~10,000, 5,000~6,000, 5,000~7,000, 5,0 The ranges include 00-8,000, 5,000-9,000, 5,000-10,000, 6,000-7,000, 6,000-8,000, 6,000-9,000, 6,000-10,000, 7,000-8,000, 7,000-9,000, 7,000-10,000, 8,000-9,000, 8,000-10,000, or 9,000-10,000 bases.
[0057] A library is provided herein that includes nucleic acids encoding antibodies, wherein the nucleic acid library is expressed in a cell. In some examples, the library is synthesized to express a reporter gene. Exemplary reporter genes include, but are not limited to, acetohydroxyacid synthase (AHAS), alkaline phosphatase (AP), β-galactosidase (LacZ), β-glucoronidase (GUS), chloramphenicol acetyltransferase (CAT), green fluorescent protein (GFP), red fluorescent protein (RFP), yellow fluorescent protein (YFP), cyan fluorescent protein (CFP), deep blue fluorescent protein, light yellow fluorescent protein, orange fluorescent protein, cherry fluorescent protein, green-blue fluorescent protein, blue fluorescent protein, horseradish peroxidase (HRP), luciferase (Luc), nopaline synthase (NOS), octopine synthase (OCS), luciferase, and derivatives thereof. Methods for determining regulation of reporter genes are well known in the art and include, but are not limited to, fluorometric methods (e.g., fluorescence spectroscopy, fluorescence activated cell sorting (FACS), fluorescence microscopy), and antibiotic resistance determination.
[0058] PD-1 Library
[0059] Provided herein are methods and compositions related to programmed cell death protein 1 (PD-1) binding libraries that include nucleic acids encoding PD-1 antibodies. In some examples, such methods and compositions are generated by the antibody optimization methods and systems described herein. The antibodies described herein can stably support a PD-1 binding domain. The PD-1 binding domain may be designed based on the surface interaction of a PD-1 ligand and PD-1. The libraries described herein can be further varigated to provide variant libraries, each of which includes nucleic acids encoding a predetermined variant of at least one predetermined reference nucleic acid sequence. Further described herein are protein libraries that can be generated upon translation of the nucleic acid libraries. In some examples, the nucleic acid libraries described herein are transcribed into cells to generate cell libraries. Also provided herein are downstream applications of libraries synthesized using the methods described herein. Downstream applications include identification of variant nucleic acid or protein sequences with improved biologically relevant functions, such as stability, affinity, binding, functional activity, and for treating or preventing disease conditions associated with PD-1 signaling. In some examples, the antibodies described herein comprise a CDRH1 sequence of any one of SEQ ID NOs: 1-35. In some examples, the antibodies described herein comprise a sequence at least 80% identical to a CDRH1 sequence of any one of SEQ ID NOs: 1-35. In some examples, the antibodies described herein comprise a sequence at least 85% identical to a CDRH1 sequence of any one of SEQ ID NOs: 1-35. In some examples, the antibodies described herein comprise a sequence at least 90% identical to a CDRH1 sequence of any one of SEQ ID NOs: 1-35. In some examples, the antibodies described herein comprise a sequence at least 95% identical to a CDRH1 sequence of any one of SEQ ID NOs: 1-35. In some examples, the antibodies described herein comprise a CDRH2 sequence of any one of SEQ ID NOs: 1-35.In some examples, the antibodies described herein comprise a sequence at least 80% identical to the CDRH2 sequence of any one of SEQ ID NOs: 1-35. In some examples, the antibodies described herein comprise a sequence at least 85% identical to the CDRH2 sequence of any one of SEQ ID NOs: 1-35. In some examples, the antibodies described herein comprise a sequence at least 90% identical to the CDRH2 sequence of any one of SEQ ID NOs: 1-35. In some examples, the antibodies described herein comprise a sequence at least 95% identical to the CDRH2 sequence of any one of SEQ ID NOs: 1-35. In some examples, the antibodies described herein comprise a CDRH3 sequence of any one of SEQ ID NOs: 1-35. In some examples, the antibodies described herein comprise a sequence at least 80% identical to the CDRH3 sequence of any one of SEQ ID NOs: 1-35. In some examples, the antibodies described herein comprise a sequence at least 85% identical to the CDRH3 sequence of any one of SEQ ID NOs: 1-35. In some examples, the antibodies described herein comprise a sequence at least 90% identical to the CDRH3 sequence of any one of SEQ ID NOs: 1-35. In some examples, the antibodies described herein comprise a sequence that is at least 95% identical to the CDRH3 sequence of any one of SEQ ID NOs:1-35.
[0060] The term "sequence identity" means that two polynucleotide sequences are identical over a comparison window (i.e., nucleotide-by-nucleotide). The term "percentage of sequence identity" is calculated by comparing two optimally aligned sequences over a comparison window, determining the number of positions where the same nucleic acid base (e.g., A, T, C, G, U, or I) occurs in both sequences to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window (i.e., the window size), and multiplying the result by 100 to obtain the percentage of sequence identity.
[0061] The term "homology" or "similarity" between two proteins is determined by comparing the amino acid sequence and its conserved amino acid substitutes of one protein sequence with another protein sequence. Similarity may be determined by procedures well known in the art, for example, the BLAST program (Basic Local Alignment Search Tool of the National Center for Biological Information).
[0062] Provided herein are libraries containing nucleic acids encoding PD-1 antibodies. The antibodies described herein can improve the stability of a range of sequences encoding PD-1 binding domains. In some instances, the sequences encoding PD-1 binding domains are determined by the interaction between a PD-1 ligand and PD-1.
[0063] The sequence of the PD-1 binding domain based on the surface interaction between PD-1 ligand and PD-1 can be analyzed using various methods. For example, computational analysis of multiple organisms can be performed. In some examples, structural analysis can be performed. In some examples, sequence analysis can be performed. Sequence analysis can be performed using databases known in the art. Non-limiting examples of databases include, but are not limited to, NCBI BLAST (blast.ncbi.nlm.nih.gov / Blast.cgi), UCSC Genome Browser (genome.ucsc.edu / ), UniProt (www.uniprot.org / ), and IUPHAR / BPS Guide to PHARMACOLOGY (guidetopharmacology.org / ).
[0064] The PD-1 binding domain designed based on sequence analysis between various organisms is described herein.For example, sequence analysis is carried out to identify the homologous sequences of different organisms.Exemplary organisms include, but are not limited to, mouse, rat, horse, sheep, cow, primate (e.g., chimpanzee, baboon, gorilla, orangutan, monkey), dog, cat, pig, donkey, rabbit, fish, bird, and human.In some examples, homologous sequences are identified in the same organism between individuals.
[0065] After identification of a PD-1 binding domain, a library can be generated that includes nucleic acids encoding the PD-1 binding domain. In some examples, the library of PD-1 binding domains includes sequences of PD-1 binding domains that are designed based on conformational ligand interactions, peptide-ligand interactions, small molecule ligand interactions, the extracellular domain of PD-1, or antibodies targeting PD-1. The library of PD-1 binding domains is translated to generate a protein library. In some examples, the library of PD-1 binding domains is translated to generate a peptide library, an immunoglobulin library, derivatives thereof, or combinations thereof. In some examples, the library of PD-1 binding domains is translated to generate a protein library that is further modified to generate a peptidomimetic library. In some examples, the library of PD-1 binding domains is translated to generate a protein library that is used to generate small molecules.
[0066] The methods described herein provide for the synthesis of a library of PD-1 binding domains, each comprising nucleic acids encoding a predetermined variant of at least one predetermined reference nucleic acid sequence. In some cases, the predetermined reference sequence is a protein-encoding nucleic acid sequence, and the variant library comprises sequences encoding at least one codon mutation, such that multiple different variants of a single residue in the subsequent protein encoded by the synthesized nucleic acid are generated by standard translation processes. In some examples, the library of PD-1 binding proteins comprises mutant nucleic acids that collectively encode mutations at multiple positions. In some examples, the variant library comprises sequences encoding at least a single codon mutation of the PD-1 binding domain. In some examples, the variant library comprises sequences encoding multiple codon mutations of the PD-1 binding domain. Exemplary numbers of codons for mutation include, but are not limited to, at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 225, 250, 275, 300, or more than 300 codons.
[0067] The methods described herein provide for the synthesis of libraries comprising nucleic acids encoding PD-1 binding domains, the libraries comprising sequences encoding length variants of the PD-1 binding domains, in some examples, the libraries comprise sequences encoding length variants that are at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 225, 250, 275, or 300 or more fewer codons compared to a predetermined reference sequence. In some examples, the library includes sequences encoding mutations that are at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 275, or 300 or more codons in length compared to a predetermined reference sequence.
[0068] Following identification of the PD-1 binding domain, an antibody may be designed and synthesized to contain the PD-1 binding domain. Antibodies containing the PD-1 binding domain may be designed based on binding, specificity, stability, expression, folding, or downstream activity. In some examples, an antibody containing the PD-1 binding domain allows contact with PD-1. In some examples, an antibody containing the PD-1 binding domain allows high affinity binding to PD-1. Exemplary amino acid sequences of PD-1 binding domains include any one of SEQ ID NOs: 1-70.
[0069] In some examples, the PD-1 antibody comprises a binding affinity (e.g., kD) for PD-1 of less than 1 nM, less than 1.2 nM, less than 2 nM, less than 5 nM, less than 10 nM, less than 11 nM, less than 13.5 nM, less than 15 nM, less than 20 nM, less than 25 nM, or less than 30 nM. In some examples, the PD-1 antibody comprises a kD of less than 1 nM. In some examples, the PD-1 antibody comprises a kD of less than 1.2 nM. In some examples, the PD-1 antibody comprises a kD of less than 2 nM. In some examples, the PD-1 antibody comprises a kD of less than 5 nM. In some examples, the PD-1 antibody comprises a kD of less than 10 nM. In some examples, the PD-1 antibody comprises a kD of less than 13.5 nM. In some examples, the PD-1 antibody comprises a kD of less than 15 nM. In some examples, the PD-1 antibody comprises a kD of less than 20 nM. In some examples, the PD-1 antibody comprises a kD of less than 25 nM. In some examples, the PD-1 antibody comprises a kD of less than 30 nM.
[0070] In some examples, the affinity of PD-1 antibodies generated by the methods described herein is at least or about 1.5-fold, 2.0-fold, 5-fold, 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold, or 200-fold or more improved binding affinity compared to a reference drug antibody. In some examples, PD-1 antibodies generated by the methods described herein are at least or about 1.5-fold, 2.0-fold, 5-fold, 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold, or 200-fold or more improved function compared to a reference drug antibody. In some examples, the reference drug antibody is an antibody with a similar structure, sequence, or antigen target.
[0071] Provided herein are PD-1 binding libraries that include nucleic acids encoding antibodies that include a PD-1 binding domain, including mutations in domain type, domain length, or residue mutations. In some examples, the domain is a region in an antibody that includes a PD-1 binding domain. For example, the region is a VH, CDR-H3, or VL domain. In some examples, the domain is a PD-1 binding domain.
[0072] The methods described herein provide for the synthesis of a PD-1 binding library of nucleic acids, each encoding a predetermined variant of at least one predetermined reference nucleic acid sequence. In some cases, the predetermined reference sequence is a nucleic acid sequence encoding a protein, and the variant library includes sequences encoding at least one codon mutation such that multiple different variants of a single residue in the subsequent protein encoded by the synthesized nucleic acid are generated by standard translation processes. In some examples, the PD-1 binding library includes variant nucleic acids that collectively encode mutations at multiple positions. In some examples, the variant library includes sequences encoding at least one codon mutation of the VH, CDR-H3, or VL domain. In some examples, the variant library includes sequences encoding at least one codon mutation of the PD-1 binding domain. For example, at least one single codon of the PD-1 binding domain is altered. In some examples, the variant library includes sequences encoding multiple codon mutations of the VH, CDR-H3, or VL domain. In some examples, the variant library includes sequences encoding multiple codon mutations of the PD-1 binding domain. Exemplary numbers of codons for mutation include, but are not limited to, at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 225, 250, 275, 300, or more than 300 codons.
[0073] The methods described herein provide for the synthesis of a PD-1 binding library of nucleic acids, each encoding a predetermined variant of at least one predetermined reference nucleic acid sequence, the PD-1 binding library comprising sequences encoding domain length variants. In some examples, the domain is a VH, CDR-H3, or VL domain. In some examples, the domain is a PD-1 binding domain. In some examples, the library comprises sequences encoding at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 225, 250, 275, or 300 or more fewer codons in length variants compared to a given reference sequence. In some examples, the library includes sequences that encode variants with at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 275, or 300 or more codon lengths as compared to a given reference sequence.
[0074] Provided herein are PD-1 binding libraries comprising nucleic acids encoding antibodies comprising a PD-1 binding domain, where the PD-1 binding libraries are synthesized with a varying number of fragments. In some examples, the fragments comprise a VH, CDR-H3, or VL domain. In some examples, the PD-1 binding libraries are synthesized with at least or about 2 fragments, 3 fragments, 4 fragments, 5 fragments, or more than 5 fragments. The length of each nucleic acid fragment of the synthesized nucleic acid, or the average length of the nucleic acid, may be at least or about 50, 75, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 525, 550, 575, 600, or more than 600 base pairs. In some examples, the length is about 50 to 600, 75 to 575, 100 to 550, 125 to 525, 150 to 500, 175 to 475, 200 to 450, 225 to 425, 250 to 400, 275 to 375, or 300 to 350 base pairs.
[0075] A PD-1 binding library comprising nucleic acids encoding antibodies comprising a PD-1 binding domain described herein, when translated, comprises a range of amino acid lengths. In some examples, the length of each amino acid fragment of synthesized amino acids, or the average length of amino acids, may be at least or about 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, or more than 150 amino acids. In some examples, the length of the amino acids is about 15 to 150, 20 to 145, 25 to 140, 30 to 135, 35 to 130, 40 to 125, 45 to 120, 50 to 115, 55 to 110, 60 to 110, 65 to 105, 70 to 100, or 75 to 95 amino acids. In some examples, the length of the amino acids is about 22 to about 75 amino acids.
[0076] The PD-1 binding library, which includes de novo synthesized variant sequences encoding antibodies comprising a PD-1 binding domain, includes a multitude of variant sequences. In some examples, the multitude of variant sequences is synthesized de novo for CDR-H1, CDR-H2, CDR-H3, CDR-L1, CDR-L2, CDR-L3, VL, VH, or a combination thereof. In some examples, the multitude of variant sequences is synthesized de novo for framework element 1 (FW1), framework element 2 (FW2), framework element 3 (FW3), or framework element 4 (FW4). In some examples, the multitude of variant sequences is synthesized de novo for the PD-1 binding domain. The number of variant sequences may be at least or about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, or more than 500 sequences. In some examples, the number of variant sequences is about 10-300, 25-275, 50-250, 75-225, 100-200, or 125-150 sequences.
[0077] The PD-1 binding library, which includes de novo synthesized variant sequences encoding antibodies that include a PD-1 binding domain, contains improved diversity. In some examples, the variants include affinity matured variants. Alternatively, or in combination, the variants include variants in other regions of the antibody, including, but not limited to, CDR-H1, CDR-H2, CDR-L1, CDR-L2, and CDR-L3. In some examples, the number of variants in the PD-1 binding library is at least or about 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , or 10 14 are non-identical sequences.
[0078] After synthesis of the PD-1 binding library containing nucleic acids encoding antibodies comprising PD-1 binding domains, the library can be used for screening and analysis. For example, the library is analyzed for display ability and panning of the library. In some examples, display ability is analyzed using a selectable tag. Exemplary tags include, but are not limited to, radioactive labels, fluorescent labels, enzymes, chemiluminescent tags, colorimetric tags, affinity tags, or other labels or tags known in the art. In some examples, the tag is histidine, polyhistidine, myc, hemagglutinin (HA), or FLAG. For example, the PD-1 binding library contains multiple tags, such as GFP, FLAG, and Lucy, as well as nucleic acids encoding antibodies comprising PD-1 binding domains with DNA barcodes. In some examples, libraries are analyzed by sequencing using a variety of methods, including, but not limited to, single molecule real-time (SMRT) sequencing, polony sequencing, sequencing by ligation, reversible terminator sequencing, proton detection sequencing, ion semiconductor sequencing, nanopore sequencing, electronic sequencing, pyrosequencing, Maxam-Gilbert sequencing, chain termination (e.g., Sanger) sequencing, +S sequencing, or sequencing by synthesis.
[0079] Diseases and Disorders
[0080] Provided herein is a PD-1 binding library that includes nucleic acids encoding antibodies that include a PD-1 binding domain, which may have a therapeutic effect. In some examples, the PD-1 binding library translationally results in a protein that is used to treat a disease or disorder. In some examples, the protein is an immunoglobulin. In some examples, the protein is a peptidomimetic. Exemplary diseases include, but are not limited to, cancer, inflammatory disease or disorder, metabolic disease or disorder, cardiovascular disease or disorder, respiratory disease or disorder, pain, digestive disease or disorder, reproductive disease or disorder, endocrine disease or disorder, or nervous system disease or disorder. In some examples, the cancer is a solid cancer or hematological cancer. In some examples, the cancer is lung, head and neck squamous cell, colorectal, melanoma, liver, classical Hodgkin's lymphoma, kidney, stomach, cervical, Merkel cell, B cell lymphoma, or bladder cancer. In some examples, the cancer is an MSI H / dMMR cancer. In some examples, the inhibitors of PD-1 programmed cell death protein 1 described herein are used for the treatment of metabolic disorders. In some examples, the subject is a mammal. In some examples, the subject is a mouse, rabbit, dog, or human. The subject treated by the methods described herein may be an infant, an adult, or a child. The pharmaceutical composition comprising the antibody or antibody fragment described herein can be administered intravenously or subcutaneously. In some examples, the pharmaceutical composition comprises an antibody or antibody fragment described herein comprising a CDR-H3 comprising any one of the sequences of SEQ ID NOs: 1-70. In some examples, the sequence of any one of SEQ ID NOs: 1-70 is used to treat cancer. In some examples, the sequence of any one of SEQ ID NOs: 1-70 is used to treat lung cancer. In some examples, the sequence of any one of SEQ ID NOs: 1-70 is used to treat head and neck squamous cell carcinoma. In some examples, the sequence of any one of SEQ ID NOs: 1-70 is used to treat colon cancer. In some examples, the sequence of any one of SEQ ID NOs: 1-70 is used to treat melanoma.In some examples, any one of SEQ ID NOs: 1-70 is used to treat liver cancer. In some examples, any one of SEQ ID NOs: 1-70 is used to treat classical Hodgkin's lymphoma. In some examples, any one of SEQ ID NOs: 1-70 is used to treat kidney cancer. In some examples, any one of SEQ ID NOs: 1-70 is used to treat gastric cancer. In some examples, any one of SEQ ID NOs: 1-70 is used to treat cervical cancer. In some examples, any one of SEQ ID NOs: 1-70 is used to treat Merkel cell carcinoma. In some examples, any one of SEQ ID NOs: 1-70 is used to treat B-cell lymphoma. In some examples, any one of SEQ ID NOs: 1-70 is used to treat bladder cancer.
[0081] Mutant library
[0082] Codon Mutation
[0083] The variant nucleic acid library described herein comprises a plurality of nucleic acids, where each nucleic acid encodes a variant codon sequence compared to a reference nucleic acid sequence. In some examples, each nucleic acid of the first nucleic acid population includes one variant at a single variant site. In some examples, the first nucleic acid population includes multiple variants at a single variant site to include more than one variant at the same variant site. The first nucleic acid population may include nucleic acids that collectively encode multiple codon variants at the same variant site. The first nucleic acid population may include nucleic acids that collectively encode up to 19 or more codons at the same position. The first nucleic acid population may include nucleic acids that collectively encode up to 60 variant triplets at the same position, or may include nucleic acids that collectively encode up to 61 different triplets of codons at the same position. Each variant may encode a codon that results in a different amino acid during translation. Table 1 provides a list of each possible codon (and representative amino acid) at the variant site.
[0084] [Table 1-1]
[0085] [Table 1-2]
[0086] A nucleic acid population may include various nucleic acids that collectively encode up to 20 codon mutations at multiple positions. In such cases, each nucleic acid in the population includes a codon mutation at more than one position in the same nucleic acid. In some examples, each nucleic acid in the population includes a codon mutation at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more codons in a single nucleic acid. In some examples, each of the variant long nucleic acids includes a codon mutation at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more codons in a single long nucleic acid. In some examples, the variant nucleic acid population comprises codon mutations at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more codons in a single nucleic acid. In some examples, the variant nucleic acid population comprises codon mutations at at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more codons in a single nucleic acid.
[0087] Highly parallel nucleic acid synthesis
[0088] Provided herein is a platform approach that utilizes miniaturization, parallelization, and vertical integration of end-to-end processes from polynucleotide synthesis to gene assembly in nanowells on silicon to create an innovative synthesis platform. The device described herein has the same footprint as a 96-well plate, and the silicon synthesis platform can increase throughput by up to 1,000-fold or more compared to traditional synthesis methods, producing up to approximately 1,000,000 or more polynucleotides, or 10,000 or more genes, in a single highly parallel run.
[0089] With the advent of next-generation sequencing, high-resolution genomic data has become a key factor for research that delves deeply into the biological roles of various genes in both normal biology and pathogenesis. At the core of this study is the central dogma of molecular biology and the concept of "residue-by-residue transfer of sequence information". Genomic information encoded in DNA is transcribed into messages, which are then translated into proteins that are the active products in a given biological pathway.
[0090] Another exciting area of research concerns the discovery, development, and production of therapeutic molecules focused on highly specific cellular targets. Highly diverse DNA sequence libraries are at the core of the development pipeline for targeted therapeutics. Genetic mutants are used to express proteins in a protein engineering cycle of design, construction, and testing, ideally arriving at a gene optimized for high expression of a protein with high affinity for the therapeutic target. As an example, consider the binding pocket of a receptor. The ability to simultaneously test the entire sequence permutation of all residues within the binding pocket allows for a thorough exploration and increases the chances of success. Saturation mutagenesis, in which researchers attempt to generate all possible mutations at a particular site within the receptor, is one approach to this development challenge. This approach is expensive, time-consuming, and labor-intensive, but allows each mutant to be introduced at every position. In contrast, combinatorial mutagenesis, in which a few selected positions or short stretches of DNA may be extensively modified, generates an incomplete repertoire of mutants with biased representation.
[0091] To expedite the drug development pipeline, a library with the desired variants available at the correct location and at the intended frequency available for testing, in other words a precision library, allows for reduced screening turnaround time in addition to reduced costs. Provided herein is a method for synthesizing nucleic acid synthetic variant libraries that precisely introduce each intended variant at the desired frequency. For the end user, this means the ability to not only exhaustively sample sequence space but also to efficiently question these hypotheses, reducing costs and screening time. Genome-wide editing can uncover important pathways, libraries where each variant and sequence order can be tested for optimal functionality, and entire pathways and genomes can be reconstructed using thousands of genes to redesign biological systems for drug discovery.
[0092] In a first example, the drug itself may be optimized using the methods described herein. For example, to improve a specific function of an antibody, a variant nucleic acid library encoding a portion of the antibody is designed and synthesized. A variant nucleic acid library for the antibody can then be generated by the processes described herein (e.g., PCR mutagenesis followed by insertion into a vector). The antibody is then expressed in a production cell line and screened for enhanced activity. Examples of screening include testing binding affinity to the antigen, stability, or modulation of effector functions (e.g., ADCC, complement, or apoptosis). Exemplary regions for optimizing an antibody include the Fc region, the Fab region, the variable region of the Fab region, the constant region of the Fab region, the variable domains of the heavy or light chains (Vc, Vdc, Vs ... H or V L ), and V H or V L Examples of such sequences include, but are not limited to, specific complementarity determining regions (CDRs) of the following sequences:
[0093] The nucleic acid library synthesized by the method described herein can be expressed in various cells associated with disease state.Cells associated with disease state include cell lines, tissue samples, primary cells, cultured cells expanded from a subject, or cells of model systems.Typical model systems include, but are not limited to, plant and animal models of disease state.
[0094] To identify mutant molecules associated with the prevention, reduction, or treatment of a disease state, the mutant nucleic acid library described herein is expressed in cells associated with the disease state or cells in which the disease state can be induced. In some examples, an agent is used to induce the disease state in cells. Exemplary tools for inducing a disease state include, but are not limited to, the Cre / Lox recombination system, LPS inflammation induction, and streptozotocin to induce hypoglycemia. The cells associated with the disease state may be cells of a model system or cultured cells, as well as cells from a subject with a particular disease state. Exemplary disease states include bacterial, fungal, viral, autoimmune, and proliferative disorders (e.g., cancer). In some examples, the mutant nucleic acid library is expressed in a model system, cell line, or primary cells from a subject and screened for changes in at least one cellular activity. Exemplary cellular activities include, but are not limited to, proliferation, cycling progression, cell death, attachment, migration, reproduction, cell signaling, energy production, oxygen utilization, metabolic activity, and senescence, response to free radical damage, or any combination thereof.
[0095] substrate
[0096] The device used as a surface for polynucleotide synthesis may be in the form of a substrate, including, but not limited to, a homogenous array surface, a patterned array surface, a channel, a bead, a gel, and the like. Provided herein is a substrate comprising a plurality of clusters, where each cluster comprises a plurality of loci that support binding and synthesis of polynucleotides. In some examples, the substrate comprises a homogenous array surface. For example, the homogenous array surface is a homogenous plate. As used herein, the term "locus" refers to a structurally distinct region that supports a polynucleotide that encodes a single predefined sequence that extends from the surface. In some examples, the locus is on a two-dimensional surface, e.g., a substantially flat surface. In some examples, the locus is on a three-dimensional surface, e.g., a well, a microwell, a channel, or a post. In some examples, the surface of the locus comprises an actively functionalized material that binds at least one nucleotide for polynucleotide synthesis, or preferably a population of identical nucleotides for synthesis of a population of polynucleotides. In some examples, polynucleotide refers to a population of polynucleotides that encode the same nucleic acid sequence. Optionally, the surface of the substrate encompasses one or more surfaces of the substrate. Using the systems and methods provided, the average error rate for polynucleotides synthesized in the libraries described herein, without error correction, is often less than 1 in 1000, less than 1 in 2000, less than 1 in 3000, or much less.
[0097] Provided herein are surfaces that support parallel synthesis of multiple polynucleotides having different pre-defined sequences at addressable locations on a common support. In some examples, the substrate may be 50, 100, 200, 400, 600, 800, 1000, 1200, 1400, 1600, 1800, 2,000; 5,000; 10,000; 20,000; 50,000; 100,000; 200,000; 300,000; 400,000; 500,000; 600,000; 700,000; 800,000; 900, The present invention supports the synthesis of 000; 1,000,000; 1,200,000; 1,400,000; 1,600,000; 1,800,000; 2,000,000; 2,500,000; 3,000,000; 3,500,000; 4,000,000; 4,500,000; 5,000,000; 10,000,000 or more non-identical polynucleotides. Optionally, the surface may comprise 50, 100, 200, 400, 600, 800, 1000, 1200, 1400, 1600, 1800, 2,000; 5,000; 10,000; 20,000; 50,000; 100,000; 200,000; 300,000; 400,000; 500,000; 600,000; 700,000; 800,000; 90 Supports are provided for the synthesis of 0,000; 1,000,000; 1,200,000; 1,400,000; 1,600,000; 1,800,000: 2,000,000; 2,500,000; 3,000,000; 3,500,000; 4,000,000; 4,500,000; 5,000,000; 10,000,000 or more polynucleotides, in some instances, at least some of the polynucleotides have identical sequences or are configured to be synthesized with identical sequences. In some examples, the substrate provides a surface environment for growth of polynucleotides having at least 80, 90, 100, 120, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, or more bases.
[0098] Provided herein is a method for polynucleotide synthesis on separate loci of a substrate, where each locus supports the synthesis of a population of polynucleotides. Optionally, each locus supports the synthesis of a population of polynucleotides with a sequence different from the population of polynucleotides grown on another locus. In some examples, each polynucleotide sequence is synthesized with 1, 2, 3, 4, 5, 6, 7, 8, 9, or more redundancies across different loci within the same cluster of loci on a surface for polynucleotide synthesis. In some examples, the loci of the substrate are located within multiple clusters. In some examples, the substrate comprises at least 10, 500, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 11000, 12000, 13000, 14000, 15000, 20000, 30000, 40000, 50000, or more clusters. In some examples, the substrate may be 2,000; 5,000; 10,000; 100,000; 200,000; 300,000; 400,000; 500,000; 600,000; 700,000; 800,000; 900,000; 1,000,000; 1,100,000; 1,200,000; 1,300,000; 1,400,000; 1,500,000; 1,600,000; 1,700,000; 1,800,000; 1,900,000; 2,000, 000;300,000;400,000;500,000;600,000;700,000;800,000;900,000;1,000,000;1,200,000;1,400,000;1,600,000;1,800,000;2,000,000;2,500,000;3,000,000;3,500,000;4,000,000;4,500,000;5,000,000; or 10,000,000 or more distinct loci. In some examples, the substrate comprises about 10,000 distinct loci. The amount of loci within a single cluster varies in different examples. Optionally, each cluster includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 130, 150, 200, 300, 400, 500 or more loci.In some examples, each cluster comprises between about 50 and 500 loci. In some examples, each cluster comprises between about 100 and 200 loci. In some examples, each cluster comprises between about 100 and 150 loci. In some examples, each cluster comprises between about 109, 121, 130, or 137 loci. In some examples, each cluster comprises between about 19, 20, 61, 64, or more loci. Alternatively, or in combination, polynucleotide synthesis occurs on a homogenous array surface.
[0099] In some instances, the number of distinct polynucleotides synthesized on a substrate depends on the number of distinct loci available on the substrate. In some instances, the density of loci within a cluster or surface of a substrate is greater than or equal to 1 mm 2 Optionally, the substrate is 10-500, 25-400, 50-500, 100-500, 150-500, 10-250, 50-250, 10-200, or 50-200 mm. 2 In some examples, the distance between the centers of two adjacent loci within a cluster or surface is about 10-500, about 10-200, or about 10-100 um. In some examples, the distance between the centers of two adjacent loci is greater than about 10 μm, 20 μm, 30 μm, 40 μm, 50 μm, 60 μm, 70 μm, 80 μm, 90 μm, or 100 μm. In some examples, the distance between the centers of two adjacent loci is less than about 200 μm, 150 μm, 100 μm, 80 μm, 70 μm, 60 μm, 50 μm, 40 μm, 30 μm, 20 μm, or 10 μm. In some examples, the width of each locus is about 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 μm. In some cases, the width of each locus is about 0.5-100, 0.5-50, 10-75, or 0.5-50 μm.
[0100] In some cases, the density of clusters within a substrate is 2At least or about 1 cluster per 10 mm 2 1 cluster per 5mm 2 1 cluster per 4mm 2 1 cluster per 3mm 2 1 cluster per 2mm 2 1 cluster per 1mm 2 1 cluster per 1mm 2 2 clusters per 1mm 2 Clusters of 3 per 1mm 2 Clusters of 4 per 1mm 2 Clusters of 5 per 1mm 2 10 clusters per 1mm 2 In some examples, the substrate is 10 mm or thicker. 2 Approximately 1 cluster per ~1mm 2 In some examples, the distance between the centers of two adjacent clusters is at least or about 50, 100, 200, 500, 1000, 2000, or 5000 μm. In some cases, the distance between the centers of two adjacent clusters is between about 50-100, 50-200, 50-300, 50-500, or 100-2000 μm. In some cases, the distance between the centers of two adjacent clusters is between about 0.05-50, 0.05-10, 0.05-5, 0.05-4, 0.05-3, 0.05-2, 0.1-10, 0.2-10, 0.3-10, 0.4-10, 0.5-10, 0.5-5, or 0.5-2 mm. In some cases, each cluster has a cross section of about 0.5 to about 2, about 0.5 to about 1, or about 1 to about 2 mm. In some cases, each cluster has a cross section of about 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, or 2 mm. In some cases, each cluster has an internal cross section of about 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.15, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, or 2 mm.
[0101] In some examples, the substrate is the size of a standard 96-well plate, e.g., about 100 to about 200 mm by about 50 to about 150 mm. In some examples, the substrate has a diameter of about 1000, 500, 450, 400, 300, 250, 200, 150, 100, or 50 mm or less. In some examples, the substrate has a diameter of about 25 to 1000, 25 to 800, 25 to 600, 25 to 500, 25 to 400, 25 to 300, or 25 to 200 mm. In some examples, the substrate has a diameter of at least about 100; 200; 500; 1,000; 2,000; 5,000; 10,000; 12,000; 15,000; 20,000; 30,000; 40,000; 50,000 mm. 2 In some examples, the thickness of the substrate is between about 50-2000, 50-1000, 100-1000, 200-1000, or 250-1000 mm.
[0102] surface material
[0103] The substrates, devices, and reactors provided herein are manufactured from a variety of materials suitable for the methods, compositions, and systems described herein. In some examples, the substrate material is manufactured to exhibit low levels of nucleotide binding. In some examples, the substrate material is modified to generate discrete surfaces that exhibit high levels of nucleotide binding. In some examples, the substrate material is transparent to visible and / or UV light. In some examples, the substrate material is sufficiently conductive, e.g., capable of forming a uniform electric field across all or a portion of the substrate. In some examples, the conductive material is connected to an electric ground. In some examples, the substrate is thermally conductive or insulated. In some examples, the material is chemically and heat resistant to support chemical or biochemical reactions, e.g., polynucleotide synthesis reaction processes. In some examples, the substrate comprises a flexible material. Flexible materials may include, but are not limited to: modified and unmodified nylon, nitrocellulose, polypropylene, and the like. In some examples, the substrate comprises a rigid material. Rigid materials may include, but are not limited to: glass; quartz glass; silicon, plastics (e.g., polytetrafluoroethylene, polypropylene, polystyrene, polycarbonate, and mixtures thereof); metals (e.g., gold, platinum, etc.). The substrate, solid support, or reactor may be made of a material selected from the group consisting of silicon, polystyrene, agarose, dextran, cellulose acid polymers, polyacrylamide, polydimethylsiloxane (PDMS), and glass. The substrate / solid support or microstructure, the reactor may be made of a combination of materials listed herein, or other suitable materials known in the art.
[0104] Surface Architecture
[0105] Substrates for the methods, compositions, and systems described herein are provided herein, where the substrate has a surface structure suitable for the methods, compositions, and systems described herein. In some examples, the substrate includes raised features and / or recessed features. One advantage of having such features is an increase in surface area to support polynucleotide synthesis. In some examples, a substrate with raised and / or recessed features is referred to as a three-dimensional substrate. In some cases, the three-dimensional substrate includes one or more channels. In some cases, one or more loci include channels. In some cases, the channels are available for deposition of reagents by a deposition device, such as a material deposition device. In some cases, the reagents and / or fluids are collected in a larger well that is in fluid communication with one or more channels. For example, the substrate includes multiple channels corresponding to multiple loci with a cluster, and the multiple channels are in fluid communication with one well of the cluster. In some methods, a library of polynucleotides is synthesized at multiple loci of the cluster.
[0106] Substrates for the methods, compositions, and systems described herein are provided, where the substrate is configured for polynucleotide synthesis. In some examples, the structure is configured to allow controlled flow and mass transfer pathways for polynucleotide synthesis on the surface. In some examples, the substrate configuration allows controlled and even distribution of mass transfer pathways, chemical exposure times, and / or washing effects during polynucleotide synthesis. In some examples, the substrate configuration allows for enhanced sweep efficiency, for example, by providing sufficient volume for growing polynucleotides such that the volume displaced by the growing polynucleotides does not occupy more than 50, 45, 40, 35, 30, 25, 20, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1, or less of the initially available volume available or suitable for polynucleotide growth. In some examples, the three-dimensional structure allows for management of fluid flow to allow rapid exchange of chemical exposure.
[0107] Substrates for the methods, compositions, and systems described herein are provided herein, where the substrate comprises a structure suitable for the methods, compositions, and systems described herein. In some examples, isolation is achieved by physical structures. In some examples, isolation is achieved by differential functionalization of the surface to generate active and passive areas for polynucleotide synthesis. In some examples, differential functionalization is achieved by varying hydrophobicity across the substrate surface, thereby creating a water contact angle effect that causes beading or wetting of the deposited reagents. Larger structures can be utilized to reduce splashing and cross-contamination of separate polynucleotide synthesis locations with reagents from adjacent spots. In some cases, an apparatus such as a material deposition apparatus is used to deposit reagents at separate polynucleotide synthesis locations. Substrates with three-dimensional features are configured in a manner that allows for the synthesis of large numbers of polynucleotides (e.g., greater than about 10,000) with low error rates (e.g., less than about 1:500, 1:1000, 1:1500, 1:2,000, 1:3,000, 1:5,000, or 1:10,000). In some examples, the substrate is 1 mm 2 In some embodiments, the method comprises the step of: providing a plurality of features having a density of about 1, 5, 10, 20, 30, 40, 50, 60, 70, 80, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 300, 400, or 500 or more features per feature;
[0108] A well in a substrate may have the same or different width, height, and / or volume as another well in the substrate. A channel in a substrate may have the same or different width, height, and / or volume as another channel in the substrate. In some examples, the diameter of the cluster, or the diameter of the well containing the cluster, or both, is about 0.05-50, 0.05-10, 0.05-5, 0.05-4, 0.05-3, 0.05-2, 0.05-1, 0.05-0.5, 0.05-0.1, 0.1-10, 0.2-10, 0.3-10, 0.4-10, 0.5-10, 0.5-5, or 0.5-2 mm. In some examples, the diameter of the cluster, well, or both is about 5, 4, 3, 2, 1, 0.5, 0.1, 0.09, 0.08, 0.07, 0.06, or 0.05 mm or less. In some examples, the diameter of the cluster, well, or both is about 1.0-1.3 mm. In some examples, the diameter of the cluster, well, or both is about 1.150 mm. In some examples, the diameter of the cluster, well, or both is about 0.08 mm. The diameter of the cluster refers to a cluster within a two-dimensional or three-dimensional substrate.
[0109] In some examples, the height of the wells is about 20-1000, 50-1000, 100-1000, 200-1000, 300-1000, 400-1000, or 500-1000 um. In some cases, the height of the wells is less than about 1000, 900, 800, 700, or 600 um.
[0110] In some examples, the substrate includes a plurality of channels corresponding to a plurality of loci in the cluster, where the height or depth of the channels is between 5-500, 5-400, 5-300, 5-200, 5-100, 5-50, or 10-50 um. In some cases, the height of the channels is less than 100, 80, 60, 40, or 20 um.
[0111] In some examples, the diameter of the channel, locus (e.g., in a substantially planar substrate), or both the channel and locus (e.g., in a three-dimensional substrate where the locus corresponds to the channel) is about 1-1000, 1-500, 1-200, 1-100, 5-100, or 10-100 um, e.g., about 90, 80, 70, 60, 50, 40, 30, 20, or 10 um. In some examples, the diameter of the channel, locus, or both the channel and locus is less than about 100, 90, 80, 70, 60, 50, 40, 30, 20, or 10 μm. In some examples, the distance between the centers of two adjacent channels, loci, or channels and loci is about 1-500, 1-200, 1-100, 5-200, 5-100, 5-50, or 5-30, e.g., about 20 um.
[0112] Surface Modification
[0113] Methods for polynucleotide synthesis on a surface are provided herein, where the surface includes various surface modifications. In some examples, surface modifications are utilized to chemically and / or physically modify a surface by an additive or subtractive process to modify one or more chemical and / or physical properties of the substrate surface, or selected sites or regions of the substrate surface. For example, surface modifications include, but are not limited to, (1) modifying the wettability of the surface, (2) functionalizing the surface, i.e., providing, modifying, or replacing surface functional groups, (3) defunctionalizing the surface, i.e., removing surface functional groups, (4) otherwise modifying the chemical composition of the surface, e.g., by etching, (5) increasing or decreasing surface roughness, (6) providing a coating on the surface, e.g., a coating that exhibits a wettability different from that of the surface, and / or (7) depositing particles on the surface.
[0114] In some examples, the addition of a chemical layer (called an adhesion promoter) on the surface promotes structured patterning of sites on the surface of the substrate. Exemplary surfaces for the application of adhesion promoters include, but are not limited to, glass, silicon, silicon dioxide, and silicon nitride. In some cases, the adhesion promoter is a chemical with a high surface energy. In some examples, a second chemical layer is deposited on the surface of the substrate. In some cases, the second chemical layer has a low surface energy. In some cases, the surface energy of the chemical layer coated on the surface supports localization of the droplets on the surface. Depending on the patterning arrangement selected, the proximity of the sites and / or the area of fluid contact with the sites can be altered.
[0115] In some instances, the substrate surface or resolved locus on which nucleic acids or other moieties are deposited, for example for polynucleotide synthesis, is smooth or substantially planar (e.g., two-dimensional) or has irregularities, such as raised or recessed features (e.g., three-dimensional features). In some instances, the substrate surface is modified with one or more distinct layers of compounds. Such modified layers of interest include, but are not limited to, inorganic and organic layers, such as metals, metal oxides, polymers, small organic molecules, and the like.
[0116] In some examples, the decomposed sites of the substrate are functionalized with one or more moieties that increase and / or decrease the surface energy. In some cases, the moieties are chemically inert. In some cases, the moieties are configured to support one or more processes in a desired chemical reaction, such as a polynucleotide synthesis reaction. The surface energy, i.e., hydrophobicity, of the surface is a factor for determining the affinity of nucleotides to bind to the surface. In some examples, a method for functionalizing a substrate includes (a) providing a substrate having a surface comprising silicon dioxide, and (b) silanizing the surface using a suitable silanizing agent, such as an organofunctional alkoxysilane molecule, as described herein or otherwise known in the art. Methods and functionalizing agents are described in U.S. Pat. No. 5,474,796, which is incorporated herein by reference in its entirety.
[0117] In some examples, the substrate surface is functionalized by contact with a derivatization composition containing a mixture of silanes under reaction conditions effective to link the silanes to the substrate surface, typically via reactive hydrophilic moieties present on the substrate surface. The silane treatment generally covers the surface with organofunctional alkoxysilane molecules via self-assembly. As currently known in the art, various siloxane functionalization reagents can also be used, for example, to reduce or increase surface energy. Organofunctional alkoxysilanes are classified according to their organofunctional groups.
[0118] Polynucleotide Synthesis
[0119] The disclosed method for polynucleotide synthesis may include a process that includes phosphoramidite chemistry. In some examples, the polynucleotide synthesis includes coupling a base with a phosphoramidite. The polynucleotide synthesis may include coupling a base by deposition of a phosphoramidite under coupling conditions, where the same base is optionally deposited with the phosphoramidite more than once, i.e., in a double coupling. The polynucleotide synthesis may include capping of unreacted sites. In some examples, capping is optional. The polynucleotide synthesis may include oxidation, or an oxidation step or multiple oxidation steps. The polynucleotide synthesis may include deblocking, detritylation, and sulfurization. In some examples, the polynucleotide synthesis includes either oxidation or sulfurization. In some examples, between one or each step during the polynucleotide synthesis reaction, the equipment is washed, for example, using tetrazole or acetonitrile. The time frame for any one step in the phosphoramidite synthesis method can be less than about 2 minutes, 1 minute, 50 seconds, 40 seconds, 30 seconds, 20 seconds, and 10 seconds.
[0120] Polynucleotide synthesis using the phosphoramidite method may include the subsequent addition of phosphoramidite building blocks (e.g., nucleoside phosphoramidites) to a growing polynucleotide chain for the formation of a phosphite triester bond. Phosphoramidite polynucleotide synthesis proceeds in the 3' to 5' direction. Phosphoramidite polynucleotide synthesis allows for the controlled addition of one nucleotide to a growing nucleic acid chain per synthesis cycle. In some examples, each synthesis cycle includes a coupling step. Phosphoramidite coupling includes the formation of a phosphite triester bond between an activated nucleoside phosphoramidite and a nucleoside attached to a substrate, for example, via a linker. In some examples, the nucleoside phosphoramidite is provided to an activated device. In some examples, the nucleoside phosphoramidite is provided to an device along with an activator. In some examples, the nucleoside phosphoramidite is provided to the device in 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100 or more times excess over the nucleoside bound to the substrate. In some examples, the addition of the nucleoside phosphoramidite is carried out in an anhydrous environment, for example, in anhydrous acetonitrile. Following the addition of the nucleoside phosphoramidite, the device is optionally washed. In some examples, the coupling step is repeated one or more times, optionally with a washing step between the addition of the nucleoside phosphoramidite to the substrate. In some examples, the polynucleotide synthesis method used herein comprises one, two, three or more sequential coupling steps. Prior to coupling, the device-bound nucleoside is often deprotected by removal of a protecting group, which functions to prevent polymerization. A common protecting group is 4,4'-dimethoxytrityl (DMT).
[0121] After coupling, the phosphoramidite polynucleotide synthesis method optionally includes a capping step. In the capping step, the growing polynucleotide is treated with a capping agent. The capping step is useful to block the 5'-OH group attached to the substrate that remains unreacted after coupling from further chain extension, thereby preventing the formation of polynucleotides with internal base deletions. In addition, phosphoramidites activated with 1H-tetrazole can react to a small extent with the O6 position of guanosine. Without being bound by theory, upon oxidation with I2 / water, this by-product may also undergo depurination, possibly via O6-N7 migration. The apurinic site is eventually cleaved during the final deprotection of the polynucleotide, thus reducing the yield of the full-length product. The O6 modification can be removed by treatment with a capping reagent prior to oxidation with I2 / water. In some instances, the inclusion of a capping step during polynucleotide synthesis reduces the error rate compared to synthesis without capping. In one example, the capping step involves treating the substrate-bound polynucleotides with a mixture of acetic anhydride and 1-methylimidazole. After the capping step, the device is optionally washed.
[0122] In some examples, after addition of the nucleoside phosphoramidites, and optionally after capping and one or more washing steps, the growing nucleic acid bound to the device is oxidized. The oxidation step oxidizes the phosphite triester to a tetracoordinate phosphate triester, which is a protected precursor of the naturally occurring phosphodiester internucleoside linkage. In some examples, oxidation of the growing polynucleotide is accomplished by treatment with iodine and water, optionally in the presence of a weak base (e.g., pyridine, lutidine, collidine). Oxidation can be carried out under anhydrous conditions, for example, using tert-butyl hydroperoxide or (1S)-(+)-(10-camphorsulfonyl)-oxaziridine (CSO). In some methods, a capping step is carried out following oxidation. A second capping step allows for drying of the device, since residual water from possible continued oxidation can inhibit subsequent coupling. After oxidation, the device and growing polynucleotide are optionally washed. In some instances, the oxidation step is replaced by a sulfurization step to obtain polynucleotide phosphorothioates, where any capping step can be performed after sulfurization. Many reagents can perform efficient sulfur transfer, including but not limited to 3-(dimethylaminomethylidene)amino)-3H-1,2,4-dithiazole-3-thione, DDTT, 3H-1,2-benzodithiol-3-one 1,1-dioxide (also known as Beaucage reagent), and N,N,N'N'-tetraethylthiuram disulfide (TETD).
[0123] To allow subsequent cycles of nucleoside incorporation to occur via coupling, the protected 5' end of the growing polynucleotide bound to the device is removed and the primary hydroxyl group reacts with the next nucleoside phosphoramidite. In some examples, the protecting group is DMT and deblocking occurs with trichloroacetic acid in dichloromethane. Detritylation for a long period of time or with stronger than recommended acid solutions can increase depurination of the polynucleotide bound to the solid support and thus reduce the yield of the desired full-length product. The disclosed methods and compositions described herein provide controlled deblocking conditions that limit undesired depurination reactions. In some examples, the polynucleotide bound to the device is washed after deblocking. In some examples, efficient washing after deblocking contributes to synthesized polynucleotides with low error rates.
[0124] Methods for the synthesis of polynucleotides typically involve an iterating sequence of the following steps: application of a protected monomer to an activated surface, linker, or actively functionalized surface (e.g., site) for conjugation with a previously deprotected monomer, deprotection of the applied monomer to react with a subsequently applied protected monomer, and application of another protected monomer for conjugation. One or more intermediate steps include oxidation or sulfurization. In some instances, one or all of the steps are preceded or followed by one or more washing steps.
[0125] Methods for phosphoramidite-based polynucleotide synthesis include a series of chemical steps. In some examples, one or more steps of the synthesis method include reagent cycling, where one or more steps of the method include application of reagents useful for the process to a device. For example, the reagents are cycled through a series of liquid deposition and vacuum drying steps. For substrates that include three-dimensional features such as wells, microwells, channels, etc., the reagents are optionally passed through one or more regions of the device via the wells and / or channels.
[0126] The method and system described herein relate to a polynucleotide synthesis device for synthesizing polynucleotides.Synthesis can be performed in parallel.For example, at least or approximately at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 1000, 10000, 50000, 75000, 100000 or more polynucleotides can be synthesized in parallel. The total number of polynucleotides that can be synthesized in parallel can be between 2-100000, 3-50000, 4-10000, 5-1000, 6-900, 7-850, 8-800, 9-750, 10-700, 11-650, 12-600, 13-550, 14-500, 15-450, 16-400, 17-350, 18-300, 19-250, 20-200, 21-150, 22-100, 23-50, 24-45, 25-40, 30-35. Those skilled in the art will recognize that the total number of polynucleotides synthesized in parallel can be included within any range bounded by any of these values (e.g., 25-100). The total number of polynucleotides synthesized in parallel can be included within any range defined by any of the values that serve as the endpoints of the range. The total molar mass of the polynucleotides synthesized in the device or the molar mass of each of the polynucleotides can be at least or at least about 10, 20, 30, 40, 50, 100, 250, 500, 750, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 25000, 50000, 75000, 100000 picomoles, or more. The length of each of the polynucleotides in the device or the average length of the polynucleotides can be at least or at least about 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 150, 200, 300, 400, 500 nucleotides, or more.The length of each of the polynucleotides or the average length of the polynucleotides in the device can be at most or at most about 500, 400, 300, 200, 150, 100, 50, 45, 35, 30, 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10 nucleotides or less. The length of each of the polynucleotides or the average length of the polynucleotides in the device can be between 10-500, 9-400, 11-300, 12-200, 13-150, 14-100, 15-50, 16-45, 17-40, 18-35, 19-25. Those skilled in the art will recognize that the length of each of the polynucleotides or the average length of the polynucleotides in the device can fall within any range (e.g., 100-300) bounded by any of these values. The length of each of the polynucleotides in the device or the average length of the polynucleotides can fall within any range defined by any of the values that serve as the endpoints of the range.
[0127] The method for synthesizing polynucleotides on a surface provided herein allows for high-speed synthesis.In one example, at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45 50, 55, 60, 70, 80, 90, 100, 125, 150, 175, 200 nucleotides or more are synthesized per hour.Nucleotides include adenine, guanine, thymine, cytosine, uridine building blocks, or their analogs / modified versions.In some examples, a library of polynucleotides is synthesized in parallel on a substrate. For example, a device comprising about or at least about 100, 1,000, 10,000, 30,000, 75,000, 100,000, 1,000,000, 2,000,000, 3,000,000, 4,000,000, or 5,000,000 decomposed loci can support the synthesis of at least the same number of distinct polynucleotides, where polynucleotides encoding distinct sequences are synthesized at the decomposed loci. In some examples, a library of polynucleotides is synthesized on the device in less than about 3 months, 2 months, 1 month, 3 weeks, 15 days, 14 days, 13 days, 12 days, 11 days, 10 days, 9 days, 8 days, 7 days, 6 days, 5 days, 4 days, 3 days, 2 days, 24 hours with a low error rate as described herein. In some examples, large nucleic acids assembled from polynucleotide libraries synthesized with low error rates using the substrates and methods described herein are prepared in less than about 3 months, 2 months, 1 month, 3 weeks, 15 days, 14 days, 13 days, 12 days, 11 days, 10 days, 9 days, 8 days, 7 days, 6 days, 5 days, 4 days, 3 days, 2 days, or 24 hours.
[0128] In some examples, the methods described herein result in the generation of a library of nucleic acids that includes mutant nucleic acids that differ at multiple codon sites.In some examples, the nucleic acids can have one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen, sixteen, seventeen, eighteen, nineteen, twenty, thirty, forty, fifty, or more mutant codon sites.
[0129] In some instances, one or more of the variant codon sites may be contiguous. One or more of the variant codon sites may be non-contiguous, separated by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more codons.
[0130] In some instances, a nucleic acid may comprise multiple sites of a mutant codon site, where the mutant codon sites are adjacent to each other, forming a stretch of mutant codon sites. In some instances, a nucleic acid may comprise multiple sites of a mutant codon site, where the mutant codon sites are not adjacent to each other. In some instances, a nucleic acid may comprise multiple sites of a mutant codon site, where some mutant codon sites are adjacent to each other, forming a stretch of mutant codon sites, and some of the mutant codon sites are not adjacent to each other.
[0131] Referring to the drawings, Figure 8 shows the workflow of an exemplary process for the synthesis of nucleic acids (e.g., genes) from shorter nucleic acids. The workflow is generally divided into the following stages: (1) de novo synthesis of single-stranded nucleic acid libraries, (2) ligation of nucleic acids to form larger fragments, (3) error correction, (4) quality control, and (5) transport. Prior to de novo synthesis, the intended nucleic acid sequence or group of nucleic acid sequences is preselected. For example, a group of genes is preselected for production.
[0132] Once larger nucleic acids are selected for production, a predefined library of nucleic acids is designed for de novo synthesis. Various suitable methods for generating high density polynucleotide arrays are known. In an example workflow, a surface layer of a device is provided. In this example, the chemistry of the surface is modified to improve the polynucleotide synthesis process. Areas of low surface energy are generated to repel liquids, while areas of high surface energy are generated to attract liquids. The surface itself may be in the form of a flat surface, or may include geometric variations such as protrusions or microwells that increase the surface area. In an example workflow, the selected high surface energy molecules serve a dual function of supporting DNA chemistry, as disclosed in International Patent Application Publication WO / 2015 / 021080, which is incorporated herein by reference in its entirety.
[0133] In situ preparations of polynucleotide arrays are generated on solid supports utilizing a single nucleotide extension process to extend multiple oligomers in parallel. A deposition device, such as a material deposition device, is designed to release reagents in a stepwise manner, allowing multiple polynucleotides to be extended in parallel, one residue at a time, thereby generating oligomers having a predetermined nucleic acid sequence (802). In some instances, the polynucleotides are cleaved from the surface at this stage. Cleavage includes, for example, gas phase cleavage with ammonia or methylamine.
[0134] The generated polynucleotide library is placed in a reaction chamber. In this exemplary workflow, the reaction chamber (also called a "nanoreactor") is a silicon-coated well that contains PCR reagents and is lowered (803) onto the polynucleotide library. Reagents are added to release the polynucleotides from the substrate before or after sealing of the polynucleotides (804). In the exemplary workflow, the polynucleotides are released after sealing of the nanoreactor (805). Once released, the single-stranded polynucleotide fragments hybridize to span the full-length span of the DNA sequence. Partial hybridization (805) is possible because each synthesized polynucleotide is designed to have a small portion that overlaps with at least one other polynucleotide in the pool.
[0135] After hybridization, the PCA reaction begins. During polymerase cycles, polynucleotides are annealed to complementary fragments and gaps are filled by the polymerase. Each cycle randomly increases the length of the various fragments depending on which polynucleotides are found next to each other. Complementarity between the fragments allows the formation of complete large spans of double-stranded DNA (806).
[0136] After PCA is complete, the nanoreactors are separated from the device (807) and positioned for interaction with a device that has primers for PCR (808). After sealing, the nanoreactors are subjected to PCR (809) to amplify larger nucleic acids. After PCR (810), the nanochambers are opened (811), error correction reagents are added (812), the chambers are sealed (813), and an error correction reaction occurs to remove poorly complementary mismatched base pairs and / or strands from the double-stranded PCR amplification products (814). The nanoreactors are opened and separated (815). The error-corrected products are then subjected to additional processing steps, such as PCR and molecular barcoding, before being packaged (822) and shipped (823).
[0137] In some instances, quality control measures are taken. After error correction, the quality control steps may include, for example, interacting the wafer with sequencing primers for amplification of the error-corrected product (816), sealing the wafer in a chamber containing the error-corrected amplification product (817), and performing an additional round of amplification (818). The nanoreactors are opened (819) and the products are pooled (820) and sequenced (821). After an acceptable quality control determination is made, the packaged product (822) is approved for shipment (823).
[0138] In some examples, nucleic acids generated by a workflow such as that of FIG. 8 are subjected to mutagenesis using overlapping primers as disclosed herein. In some examples, a library of primers is generated by in situ preparation on a solid support, utilizing a single nucleotide extension process to extend multiple oligomers in parallel. A deposition device, such as a materials deposition device, is designed to release reagents in a stepwise manner, such that multiple polynucleotides are extended in parallel, one residue at a time, thereby generating oligomers having a predetermined nucleic acid sequence (802).
[0139] Computer Systems
[0140] Any of the systems described herein may be operably connected to a computer and automated locally or remotely via the computer. In various examples, the methods and systems of the present disclosure may further include a software program on a computer system, and the use thereof. Thus, computer control for synchronization of dispense / vacuum / refill functions, such as orchestrating and synchronizing the operation of the material deposition device, dispense actions, and vacuum actuation, is within the scope of the present disclosure. The computer system is programmed to interface between the base sequence specified by the user and the position of the material deposition device to deliver the correct reagent to the specified area of the substrate.
[0141] The computer system (900) illustrated in FIG. 9 may be understood as a logical device that can read instructions from a network port (905) that is optionally connectable to a server (909) having a medium (911) and / or a fixed medium (912). A system such as that shown in FIG. 9 may include a CPU (901), a disk drive (903), optional input devices such as a keyboard (915) and / or a mouse (916), and an optional monitor (907). Data communication may be accomplished over the indicated communication medium to a server at a local or remote location. The communication medium may include any means of transmitting and / or receiving data. For example, the communication medium may be a network connection, a wireless connection, or an Internet connection. Such a connection may provide for communication over the World Wide Web. It is envisioned that data relating to the present disclosure may be communicated over such networks or connections for receipt and / or consideration by the parties (922) as illustrated in FIG. 9.
[0142] FIG. 10 is a block diagram illustrating the architecture of a first example of a computer system (1000) that can be used in conjunction with examples of the present disclosure. As depicted in FIG. 10, the exemplary computer system may include a processor (1002) for processing instructions. Non-limiting examples of processors include Intel Xeon™ processors, AMD Opteron™ processors, Samsung 32-bit RISC ARM 1176JZ(F)-S v1.0™ processors, ARM Cortex-A8 Samsung S5PC100™ processors, ARM Cortex-A8 Apple A4™ processors, Marvell PXA 930™ processors, or functionally equivalent processors. Multiple threads of execution can be used for parallel processing. In some examples, multiple processors, or processors with multiple cores, can also be used, whether in a single computer system, in a cluster, or distributed across a network of systems including multiple computers, mobile phones, and / or personal digital assistant devices.
[0143] As shown in FIG. 10, a high speed cache (1004) may be connected to or incorporated into the processor (1002) to provide high speed memory for instructions or data that have been recently or frequently used by the processor (1002). The processor (1002) is connected to a north bridge (1006) by a processor bus (1008). The north bridge (1006) is connected to a random access memory (RAM) (1010) by a memory bus (1012) and manages access to the RAM (1010) by the processor (1002). The north bridge (1006) is also connected to a south bridge (1014) by a chipset bus (1016). The south bridge (1014) is in turn connected to a peripheral bus (1018). The peripheral bus may be, for example, a PCI, PCI-X, PCI Express, or other peripheral bus. The northbridge and southbridge are often referred to as the processor chipset and manage data transfers between the processor, RAM, and peripheral components over the peripheral bus (1018). In some alternative architectures, the functionality of the northbridge may be incorporated into the processor instead of using a separate northbridge chip. In some examples, the system (1000) may include an accelerator card (1022) attached to the peripheral bus (1018). The accelerator may include a field programmable gate array (FPGA) or other hardware to expedite specific processing. For example, the accelerator may be used for adaptive data reconstruction or to evaluate algebraic expressions used in extended configuration processing.
[0144] Software and data may be stored in external storage (1024) and loaded into RAM (1010) and / or cache (1004) for use by the processor. The system (1000) includes an operating system for management of system resources, non-limiting examples of which include Linux, Windows, MACOS, BlackBerry OS, iOS, and other functionally equivalent operating systems, as well as application software running on the operating system for managing data storage and optimization in accordance with examples of the present disclosure. In this example, the system (1000) further includes external storage, such as network attached storage (NAS), and network interface cards (NICs) (1020) and (1021) connected to the peripheral bus to provide a network interface to other computer systems that may be used for distributed parallel processing.
[0145] 11 is a schematic diagram showing a network (1100) with multiple computer systems (1102a) and (1102b), multiple mobile phones and personal digital assistants (1102c), and network attached storage (NAS) (1104a) and (1104b). In an example, systems (1102a), (1102b), and (1102c) can manage data storage and optimize data access to data stored in network attached storage (NAS) (1104a) and (1104b). Mathematical models can be used on this data and evaluated using distributed parallel processing across computer systems (1102a) and (1102b), and mobile phones and personal digital assistant systems (1102c). The computer systems (1102a) and (1102b) and the mobile phone and personal digital assistant system (1102c) can also provide parallel processing for adaptive data restructuring of data stored in network attached storage (NAS) (1104a) and (1104b). FIG. 11 shows only one example, and various other computer architectures and systems can be used with various examples of the present disclosure. For example, a blade server can be used to provide parallel processing. Processor blades can be connected through a backplane to provide parallel processing. Storage can also be connected to the backplane or can be connected as network attached storage (NAS) through another network interface. In some examples, the processors can maintain separate memory spaces and communicate data through a network interface, backplane, or other connector for parallel processing by other processors. In other examples, some or all of the processors can use a shared virtual address memory space.
[0146] FIG. 12 is a block diagram of a multiprocessor computer system using a shared virtual address memory space according to an example. The system includes multiple processors (1202a-f) that can access a shared memory subsystem (1204). The system incorporates multiple programmable hardware memory algorithm processors (MAPs) (1206a-f) in the memory subsystem (1204). Each MAP (1206a-f) may include a memory (1208a-f) and one or more field programmable gate arrays (FPGAs) (1210a-f). The MAPs provide configurable functional units, and specific algorithms or parts of algorithms can be provided to the FPGAs (1210a-f) for processing in close cooperation with each processor. For example, the MAPs can be used to evaluate algebraic expressions related to a data model and perform adaptive data reconstruction in the example. In this example, each MAP is globally accessible by all of the processors for such purposes. In one configuration, each MAP can use direct memory access (DMA) to access associated memory (1208a-f), allowing it to perform tasks independently and asynchronously from each microprocessor (1202a-f). In this configuration, a MAP can feed results directly to another MAP for pipelining and parallel execution of algorithms.
[0147] The above computer architectures and systems are merely examples, and various other computer, cell phone, and personal digital assistant architectures and systems can be used with the examples, including systems using general processors, co-processors, FPGAs, and other programmable logic devices, systems on chips (SOCs), application specific integrated circuits (ASICs), and any combination of other processing and logic elements. In some examples, all or part of the computer system can be implemented in software or hardware. Various data storage media can be used with the examples, including, for example, random access memory, hard drives, flash memory, tape drives, disk arrays, network attached storage (NAS), and other local or distributed data storage devices and systems.
[0148] In examples, the computer system may be implemented using software modules executing on any of the above or other computer architectures and systems. In other examples, the functionality of the system may be implemented partially or fully in firmware, programmable logic devices such as field programmable gate arrays (FPGAs) as referred to in FIG. 10, systems on chips (SOCs), application specific integrated circuits (ASICs), or other processing and logic elements. For example, the set processor and optimizer may be implemented with hardware acceleration through the use of a hardware accelerator card such as the accelerator card (1022) shown in FIG. 10.
[0149] The following examples are set forth to more clearly illustrate the principles and practice of the embodiments disclosed herein to those skilled in the art, and are not to be construed as limiting the scope of any claimed embodiments. Unless otherwise specified, all parts and percentages are by weight. EXAMPLES
[0150] The following examples are provided to illustrate various embodiments of the present disclosure and are not intended to limit the present disclosure in any manner. The examples, together with the methods described herein, illustrate preferred embodiments and are exemplary, and are not intended to limit the scope of the present disclosure. Variations herein and other uses encompassed within the spirit of the present disclosure, as defined by the scope of the claims, will occur to those skilled in the art.
[0151] Example 1: Functionalization of the device surface
[0152] The devices were functionalized to aid in the attachment and synthesis of libraries of polynucleotides. The device surface was first wet cleaned using a piranha solution containing 90% H2SO4 and 10% H2O2 for 20 minutes. The devices were rinsed in various beakers containing DI water, held under a gooseneck tap of DI water for 5 minutes, and dried with N2. The devices were then immersed in NH4OH (1:100; 3 mL: 300 mL) for 5 minutes, rinsed with DI water using a handgun, immersed in three successive beakers containing DI water for 1 minute each, and then rinsed again with DI water using a handgun. The devices were then plasma cleaned by exposing the device surface to O2. The devices were plasma etched with O2 at 250 watts for 1 minute in downstream mode using a SAMCO PC-300 instrument.
[0153] The cleaned device surfaces were actively functionalized with a solution containing N-(3-triethoxysilylpropyl)-4-hydroxybutyramide using a YES-1224P deposition oven system with the following parameters: 0.5-1 Torr, 60 min, 70 °C, vaporizer at 135 °C. The device surfaces were resist coated using a Brewer Science 200X spin coater. SPR™ 3612 photoresist was spin coated onto the devices at 2500 rpm for 40 seconds. The devices were prebaked at 90 °C for 30 minutes on a Brewer hotplate. The devices were exposed to photolithography using a Karl Suss MA6 mask aligner instrument. The devices were exposed for 2.2 seconds and developed in MSF 26A for 1 minute. The remaining developer was rinsed off with a hand gun and the devices were immersed in water for 5 minutes. The devices were baked in an oven at 100°C for 30 minutes and then visually inspected for lithographic defects using a Nikon L200. The remaining resist was removed using a descum process using a SAMCO PC-300 instrument and O2 plasma etched at 250 watts for 1 minute.
[0154] The device surface was passively functionalized with 100 μL of perfluorooctyltrichlorosilane solution mixed with 10 μL of diesel. The device was placed in the chamber and pumped for 10 minutes, after which the valve was closed, the pump was stopped, and the device was left for 10 minutes. The chamber was vented. The device was resist stripped by two immersions in 500 mL of NMP at 70° C. for 5 minutes with sonication at maximum power (9 on the Crest system). The device was then immersed in 500 mL of isopropanol at room temperature for 5 minutes with sonication at maximum power. The device was immersed in 300 mL of 200 proof ethanol and blown dry with N2. The functionalized surface was activated and served as a support for polynucleotide synthesis.
[0155] Example 2: Synthesis of a 50mer sequence on an oligonucleotide synthesizer
[0156] The two-dimensional oligonucleotide synthesizer was integrated into a flow cell, which was then connected to a flow cell (Applied Biosystems (ABI394 DNA Synthesizer)). The two-dimensional oligonucleotide synthesizer was uniformly functionalized with N-(3-triethoxysilylpropyl)-4-hydroxybutyramide (Gelest) and used to synthesize an exemplary 50 bp polynucleotide ("50-mer polynucleotide") using the polynucleotide synthesis method described herein.
[0157] The sequence of the 50-mer was as written: 5'AGACAATCAACCATTTGGGGTGGACAGCCTTGACCTCTAGACTTCGGCAT##TTTTTTTTTT3', where # represents thymidine-succinylhexamide CED phosphoramidite (CLP-2244 from ChemGenes), a cleavable linker that allows release of the oligo from the surface during deprotection.
[0158] The synthesis was carried out using standard DNA synthesis chemistry (coupling, capping, oxidation, and deblocking) according to the protocol in Table 2 and on an ABI synthesizer.
[0159] [Table 2-1]
[0160] [Table 2-2]
[0161] [Table 2-3]
[0162] The phosphoramidite / activator combination was delivered similarly to the delivery of bulk reagents through the flow cell, without any drying step as the environment remained "wet" with the reagents throughout the entire time.
[0163] The flow restrictor was removed from the ABI 394 synthesizer to allow faster flow. Without flow restrictors, the flow rates for amidite (0.1M in ACN), activator (0.25M benzoylthiotetrazole ("BTT"; GlenResearch 30-3070-xx) in ACN), and Ox (0.02M I2 in 20% pyridine, 10% water, and 70% THF) were approximately ~100uL / sec, for acetonitrile ("ACN") and capping reagent (a 1:1 mixture of CapA and CapB, where CapA is acetic anhydride in THF / pyridine and CapB is 16% 1-methylimidizole in THF), and for Deblock (3% dichloroacetic acid in toluene) were approximately ~300uL / sec (compared to ~50uL / sec for all reagents with flow restrictors). The time to completely push out the oxidizer was observed, the timing of the chemical flow time was adjusted accordingly, and an extra ACN wash was introduced between different chemicals. After polynucleotide synthesis, the chip was deprotected in gaseous ammonia at 75 psi overnight. Five drops of water were added to the surface to regenerate the polynucleotides. The recovered polynucleotides were then analyzed on a BioAnalyzer mini RNA chip.
[0164] Example 3: Synthesis of 100mer sequences on an oligonucleotide synthesizer
[0165] The same process described in Example 2 for the synthesis of the 50-mer sequence was used for the 100-mer polynucleotide ("100-mer polynucleotide"; 5'CGGGATCCTTATCGTCATCGTCGTACAGATCCCGACCCATTTGCTGTCCACCAGTCATGCTAGCCATACCATGATGATGATGATGATGAGAACCCCGCAT##TTTTTTTTTT3', where # represents thymidine-succinylhexamide CED phosphoramidite (CLP-2244 from ChemGenes) on two different silicon chips, the first silicon chip was functionalized uniformly with N-(3-triethoxysilylpropyl)-4-hydroxybutyramide and the second silicon chip was functionalized with a 5 / 95 mixture of 11-acetoxyundecyltriethoxysilane and n-decyltriethoxysilane) and the polynucleotide extracted from the surface was analyzed on a BioAnalyzer instrument.
[0166] All 10 samples from the two chips were further amplified using the forward primer (5'ATGCGGGGTTCTCATCATC3') and reverse primer (5'CGGGATCCTTATCGTCATCG3') in a 50uL PCR mixture (25uL NEB Q5 mastermix, 2.5uL 10uM forward primer, 2.5uL 10uM reverse primer, 1uL polynucleotide extracted from the surface, and up to 50uL water) using the following thermal cycling program: 98℃, 30 seconds 98°C, 10 sec; 63°C, 10 sec; 72°C, 10 sec; repeat 12 cycles 72℃, 2 minutes
[0167] The PCR products were also run on the BioAnalyzer and demonstrated a sharp peak at the 100mer position. The PCR amplified samples were then cloned and Sanger sequenced. Table 3 summarizes the results from Sanger sequencing for samples from spots 1-5 from chip 1 and spots 6-10 from chip 2.
[0168] [Table 3]
[0169] Therefore, synthetic polynucleotides of high quality and uniformity were replicated on two chips with different surface chemistries. Overall, 89% of the sequenced 100mers were error-free and had perfect sequences, corresponding to 233 out of 262.
[0170] Table 4 summarizes the error profiles for sequences obtained from polynucleotide samples from spots 1-10.
[0171] [Table 4]
[0172] Example 4: Antibody Optimization
[0173] Libraries generated from parent sequences
[0174] Antibody sequences targeting PD-1 were designed by in silico generating a library containing mutations from 12 individuals. The mutation space for the heavy and light chains was derived from the parental and closest germline sequences to generate an NGS database. The NGS database consists of the sequences of light chain CDR1-3 and heavy chain CDR1-3 containing mutations compared to the parental reference or germline sequences. All CDR sequences were represented in ≥2 individuals from the NGS database. The input sequences are shown in Figure 3A. The library consisted of 5.9x10 7 of various heavy chains, and 2.9x10 6 The light chains contained 100-fold more nucleotides than the IgG1 light chain (Fig. 3B).
[0175] Bead base selection
[0176] For five rounds of selection, C-terminally biotinylated PD-1 antigen was bound to streptavidin-coated magnetic beads. Bead-bound mutants were depleted between each round. The stringency of selection was increased in each subsequent round, and enrichment ratios track on-target binding.
[0177] ELISA and next-generation sequencing
[0178] Constructs expressing heavy and light chain combinations were synthesized and subjected to phage display to identify improved binders to PD-1. The pool was sequenced with 10 million reads, identifying 400,000 unique clones. The distribution of read lengths was highly uniform (Figure 4A). At each panning round, the frequency and accumulation of clones was measured (Figures 4B and 4C). Panning used four stringency conditions (Figure 4D). High stringency selection enriched the pool of binders for the same pool in repeated selections. Low stringency selection, with a broader range of low affinity binders, recovered 44 of 70 (63%) of the same clones as high stringency selection. At the fifth round, the majority of scFv binders were enriched (Figure 5A). In ELISA experiments measuring binding of scFv to PD-1, more than 90% (68 / 75) of the clones were present at 5-fold background levels and enriched to greater than 0.01%.
[0179] Five rounds of selection were completed under three different initial selection conditions. Clonal enrichment was tracked through each successive round by NGS. Sequences enriched for off-target or background binders were removed. Approximately 1000 clones represented in the fifth round were enriched to greater than 0.01% of the population. Sequence analysis showed that the majority (>95%) of clones enriched for binding to PD-1 were captured evenly across the different selection conditions (Figure 4E).
[0180] High-throughput IgG characterization
[0181] Clones were transiently transfected with Expi293 and purified by Kingfisher and Hamilton automation decks. Yield and purity were confirmed by Perkin Elmer Labchip and analytical HPLC. Binding affinity and epitope binning of over 170 IgG variants were assessed using the Carterra LSA system (data not shown).
[0182] The optimized IgG bound with comparable or improved affinity to control antibodies with sequences corresponding to pembrolizumab (Control 1) and nivolumab (Control 2) (data not shown). As seen in Figure 5B, scFV binding to PD-1 was also measured. The optimized antibody sequences had fewer germline mutations compared to control antibodies with sequences corresponding to pembrolizumab (Control 1) and nivolumab (Control 2) (Figure 6A and Table 5A). The light chains of the optimized antibodies were highly diverse, with over 90% of the clones containing unique light chains that were never repeated (Table 5B). The optimized IgG demonstrated a 100-fold improvement in monovalent binding affinity compared to the parental sequences. The PD1-1 clone bound to PD-1 with a KD of 4.52 nM, while several other clones showed binding affinities below 10 nM. Figure 6B shows the increase in affinity after optimization with the methods described herein. These high affinity binders each contained a unique CDRH3 and were not clustered by sequence lineage. Diversity between the various CDRs was also observed, as seen in Figures 6C-6D. PD-1 antibodies showed improved binding affinity compared to wild type (Figures 6E-6F and Table 5C).
[0183] [Table 5]
[0184] [Table 6-1]
[0185] [Table 6-2]
[0186] [Table 7]
[0187] Functional and developability assays
[0188] IgGs optimized by the methods described herein were tested for functional inhibition of PD-1 / PD-L1 interaction. Figure 7A shows that the high affinity variants showed improved IC50 compared to wild type and control anti-PD1 antibodies with sequences corresponding to nivolumab. IC50 and monovalent binding affinity were highly correlated. As seen in Figure 7B, the optimized IgGs showed improved binding affinity (up to 72-fold) and a 9.5-fold increase in function. Six antibodies were identified with higher binding affinity and function than the antibody with sequences corresponding to nivolumab. As seen in Figure 7C, addition of anti-PD-1 IgG inhibited PD-1 / PDL-1 interaction, releasing an inhibitory signal that resulted in TCR activation and NFAT-RE-mediated luminescence (RU). In addition, all binders retained binding to cyno PD-1. Several high affinity IgGs showed low polyspecificity scores measured by BVP binding ELISA (Figure 7D). In addition, IgG was tested on an Unchained UNCLE machine for Tm and Tagg, as well as analytical HPLC.
[0189] While preferred embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the present disclosure. It is understood that various alternatives to the embodiments of the present disclosure described herein may be utilized in implementing the present disclosure. The following claims define the scope of the disclosure, and it is intended that methods and structures within the scope of the claims and their equivalents be covered thereby.
Claims
1. 1. A nucleic acid library comprising: the nucleic acid library comprises a plurality of sequences comprising a nucleic acid that when translated encodes an antibody or antibody fragment, each of the plurality of sequences comprising a predetermined number of mutations in the CDRs relative to an input sequence of an antibody; the library comprises at least 50,000 variant sequences, each sequence being present at an amount within an average frequency of 1.5; A nucleic acid library, wherein at least one sequence upon translation encodes an antibody or antibody fragment having a binding affinity at least 2.5 times higher than the binding affinity of the input sequence.
2. The nucleic acid library of claim 1 , wherein the library comprises at least 100,000 variant sequences.
3. The nucleic acid library of claim 1 , wherein at least a portion of the sequences encode an antibody light chain.
4. The nucleic acid library of claim 1 , wherein at least a portion of the sequences encodes an antibody heavy chain.
5. 2. The nucleic acid library of claim 1, wherein each sequence of the plurality of sequences comprises at least one mutation in a heavy or light chain CDR relative to the input sequence.
6. 2. The nucleic acid library of claim 1, wherein each sequence of the plurality of sequences comprises at least two mutations in a heavy or light chain CDR relative to an input sequence.
7. The nucleic acid library of claim 1 , wherein at least one of the mutations is present in at least two individuals.
8. The nucleic acid library of claim 1 , wherein at least one of the mutations is present in at least three individuals.
9. The nucleic acid library of claim 1, wherein at least one sequence upon translation encodes an antibody or antibody fragment with a binding affinity at least 5-fold higher than the binding affinity of the input sequence.
10. 2. The nucleic acid library of claim 1, wherein upon translation at least one sequence encodes an antibody or antibody fragment with a binding affinity at least 25 times higher than the binding affinity of the input sequence.
11. The nucleic acid library of claim 1, wherein at least one sequence upon translation encodes an antibody or antibody fragment with a binding affinity at least 50 times higher than the binding affinity of the input sequence.
12. 2. The nucleic acid library of claim 1, wherein each sequence of the plurality of sequences comprises at least one mutation in a heavy or light chain CDR relative to a germline sequence of the input sequence.
13. 13. The nucleic acid library of claim 1, wherein the CDRs are CDR1, CDR2, and CDR3 on the heavy chain.
14. 14. The nucleic acid library of claim 1, wherein the CDRs are CDR1, CDR2, and CDR3 on a light chain.
15. 2. The nucleic acid library of claim 1, wherein at least one sequence upon translation encodes an antibody or antibody fragment with a binding affinity at least 70-fold higher than the input sequence.
16. During translation, at least one sequence is D The nucleic acid library of claim 1 , which encodes an antibody or antibody fragment having a molecular weight of less than 50 nM.
17. During translation, at least one sequence is D The nucleic acid library of claim 1 , which encodes an antibody or antibody fragment having a molecular weight of less than 25 nM.
18. During translation, at least one sequence is D The nucleic acid library of claim 1 , which encodes an antibody or antibody fragment having a molecular weight of less than 10 nM.
19. During translation, at least one sequence is D The nucleic acid library of claim 1 , which encodes an antibody or antibody fragment having a molecular weight of less than 5 nM.
20. The nucleic acid library of claim 1, wherein the library comprises any one of the CDR sequences of SEQ ID NOs: 1 to 6 or 9 to 70.
21. 2. The nucleic acid library of claim 1, wherein the library comprises a CDRH1, CDRH2, or CDRH3 sequence of any one of SEQ ID NOs: 1-6 or 9-70.
22. The library has a K of less than 10 nM D The nucleic acid library of claim 1, comprising at least one sequence encoding an antibody or antibody fragment that binds to PD-1 at a
23. The library has a K of less than 5 nM D The nucleic acid library of claim 1, comprising at least one sequence encoding an antibody or antibody fragment that binds to PD-1 at a
24. The library has a K of less than 10 nM D The nucleic acid library of claim 1, comprising at least five sequences encoding antibodies or antibody fragments that bind to PD-1.
25. The nucleic acid library of claim 1 , wherein the library comprises at least 100,000 variant sequences.
26. An antibody comprising any one of the sequences set forth in SEQ ID NOs: 1-6 or 9-70.
27. An antibody comprising any one of the sequences of SEQ ID NOs: 1-6 or 9-34, which is a monoclonal antibody, a polyclonal antibody, a bispecific antibody, a multispecific antibody, a grafted antibody, a human antibody, a humanized antibody, a synthetic antibody, a chimeric antibody, a camelized antibody, a single chain Fv (scFv), a single chain antibody, a Fab fragment, an F(ab')2 fragment, an Fd fragment, an Fv fragment, a single domain antibody, an isolated complementarity determining region (CDR), a diabody, a fragment consisting of only a single monomeric variable domain, a disulfide-linked Fv (sdFv), an intrabody, an anti-idiotypic (anti-Id) antibody, or an ab antigen-binding fragment thereof.
28. A method for inhibiting PD-1 activity comprising administering the antibody of claim 26 or 27.
29. 28. A method of treating a proliferative disorder comprising administering to a subject an antibody according to claim 26 or 27.
30. 30. The method of claim 29, wherein the proliferative disorder is cancer.
31. 30. The method of claim 29, wherein the cancer is lung cancer, head and neck cancer, colon cancer, melanoma, liver cancer, classical Hodgkin's lymphoma, renal cancer, gastric cancer, cervical cancer, Merkel cell, B-cell lymphoma, or bladder cancer.
32. 1. A computerized system for optimizing an antibody, comprising: (a) a general purpose computer; (b) a computer-readable medium containing functional modules containing instructions for the general-purpose computer; Including, The computerized system comprises: (i) receiving operational instructions, the operational instructions comprising a plurality of sequences encoding an antibody or antibody fragment; (ii) generating a nucleic acid library comprising a plurality of sequences comprising nucleic acids that, when translated, encode an antibody or antibody fragment, wherein each of the plurality of sequences comprises a predetermined number of mutations in a CDR relative to an input sequence of an antibody, the library comprising at least 50,000 variant sequences, each sequence being present at an amount that is within a 1.5-fold average frequency, and wherein, when translated, at least one sequence encodes an antibody or antibody fragment with a binding affinity that is at least 2.5-fold higher than the binding affinity of the input sequence; and (iii) Synthesis of at least 50,000 variant sequences. A computerized system configured to operate in the method of claim 1.
33. 33. The computerized system of claim 32, wherein the nucleic acid library comprises at least 100,000 sequences.
34. 33. The computerized system of claim 32, further comprising enriching the subset of variant sequences.
35. 33. The computerized system of claim 32, further comprising expressing an antibody or antibody fragment corresponding to said variant sequence.
36. 33. The computerized system of claim 32, wherein the polynucleotide sequence is a murine, human, or chimeric antibody sequence.
37. 33. The computerized system of claim 32, wherein each sequence of the plurality of variant sequences comprises at least one mutation in a heavy or light chain CDR relative to the input sequence.
38. 33. The computerized system of claim 32, wherein each sequence of the plurality of variant sequences comprises at least two mutations in a heavy or light chain CDR relative to the input sequence.
39. 33. The computerized system of claim 32, wherein upon translation, at least one sequence encodes an antibody or antibody fragment with a binding affinity at least 5-fold higher than the binding affinity of the input sequence.
40. 33. The computerized system of claim 32, wherein upon translation, at least one sequence encodes an antibody or antibody fragment with a binding affinity at least 25 times higher than the binding affinity of the input sequence.
41. 33. The computerized system of claim 32, wherein upon translation, at least one sequence encodes an antibody or antibody fragment with a binding affinity at least 50-fold higher than the binding affinity of the input sequence.
42. 33. The computerized system of claim 32, wherein each sequence of the plurality of variant sequences comprises at least one mutation in a heavy or light chain CDR relative to a germline sequence of the input sequence.
43. 43. The computerized system of any one of claims 32 to 42, wherein the CDRs are CDR1, CDR2, and CDR3 on the heavy chain.
44. 44. The computerized system of any one of claims 32 to 43, wherein the CDRs are CDR1, CDR2, and CDR3 on the light chain.
45. The theoretical diversity of the antibody library is at least 10 12 The computerized system of claim 32, wherein the sequence is:
46. The theoretical diversity of the antibody library is at least 10 13 The computerized system of claim 32, wherein the sequence is:
47. 1. A method for optimizing an antibody, comprising: (a) providing a plurality of polynucleotide sequences encoding antibodies or antibody fragments; (b) generating a nucleic acid library comprising a plurality of sequences comprising nucleic acids that when translated encode an antibody or antibody fragment, each of said plurality of sequences comprising a predetermined number of mutations in CDRs relative to an input sequence of an antibody, said library comprising at least 50,000 variant sequences, each sequence being present at an amount that is within a 1.5-fold average frequency, and at least one sequence when translated encodes an antibody or antibody fragment with a binding affinity at least 2.5-fold higher than the binding affinity of said input sequence; (c) synthesizing at least 50,000 variant sequences; A method comprising:
48. 48. The method of claim 47, wherein the antibody library comprises at least 100,000 sequences.
49. 48. The method of claim 47, further comprising enriching the subset of variant sequences.
50. 48. The method of claim 47, further comprising expressing an antibody or antibody fragment corresponding to said variant sequence.
51. 48. The method of claim 47, wherein the polynucleotide sequence is a murine, human, or chimeric antibody sequence.
52. 48. The method of claim 47, wherein each sequence of the plurality of variant sequences comprises at least one mutation in each CDR of the heavy or light chain relative to the input sequence.
53. 48. The method of claim 47, wherein each sequence of the plurality of variant sequences comprises at least two mutations in each CDR of the heavy or light chain relative to the input sequence.
54. 48. The method of claim 47, wherein upon translation, at least one sequence encodes an antibody or antibody fragment with a binding affinity at least 5-fold higher than the binding affinity of the input sequence.
55. 48. The method of claim 47, wherein upon translation, at least one sequence encodes an antibody or antibody fragment with a binding affinity at least 25 times higher than the binding affinity of the input sequence.
56. 48. The method of claim 47, wherein upon translation, at least one sequence encodes an antibody or antibody fragment with a binding affinity at least 50-fold higher than the binding affinity of the input sequence.
57. 48. The method of claim 47, wherein each sequence contains at least one mutation in each CDR of the heavy or light chain relative to the germline sequence of the input sequence.
58. The theoretical diversity of the antibody library is at least 10 12 The method of claim 47, wherein the sequence is:
59. The theoretical diversity of the antibody library is at least 10 13 The method of claim 47, wherein the sequence is:
Citation Information
Patent Citations
Express humanization of antibodies
US20170247473A1
PD-1 / tim-3 BI-specific antibodies, compositions thereof, and methods of making and using the same
WO2018156777A1