Methods, systems and compositions for generating and analyzing polypeptide libraries - Patents.com

JP2024525171A5Pending Publication Date: 2025-06-25PROTILLION BIOSCIENCES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023577679
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-06-15
Filing Date
2022-06-14
Publication Date
2025-06-25

AI Technical Summary

Technical Problem

Existing methods for identifying polypeptides of interest through directed evolution and screening techniques face challenges due to the complexity of sequence space and lack of sequence diversity, potentially leading to the loss of potentially useful polypeptides.

Method used

A high-throughput method involving polynucleotide and polypeptide libraries, along with display approaches, is used to analyze and optimize polypeptides by measuring characteristics such as equilibrium binding constants, kinetic binding constants, protein stability, and enzymatic activities, enabling the identification of optimized polypeptides.

Benefits of technology

This method allows for the rapid and efficient identification of polypeptides with specific characteristics, overcoming the limitations of existing techniques by analyzing large numbers of variants and generating quantitative data to guide iterative improvements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Methods, systems and compositions are disclosed for the analysis of polypeptides and the generation of polypeptide libraries. Analysis of polypeptide libraries can be used to generate polypeptides with specific characteristics. Antibodies with high affinity can be generated using the disclosed methods, systems and compositions.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] cross reference

[0001] This application claims priority to U.S. Provisional Application No. 63 / 210,905, filed June 15, 2021, the entirety of which is incorporated by reference herein. [Background technology]

[0002]

[0002] Polypeptides can be used for various purposes, such as therapeutics. Directed evolution or selection strategies can be used to identify polypeptides of interest. Methods of protein display can be used in conjunction with directed evolution. Directed evolution techniques can use protein display to screen for polypeptides of interest. Directed evolution and screening techniques can be effective in identifying polypeptides of interest, but can unintentionally lose potentially useful polypeptides due to the complexity of sequence space and lack of sequence diversity. Summary of the Invention [Means for solving the problem]

[0003]

[0003] Methods, systems and compositions are provided herein for the analysis of multiple polypeptides. The methods, systems and compositions allow for the generation of polypeptides with specific characteristics. The methods, systems and compositions may use polynucleotide and polypeptide libraries and polypeptide display approaches to develop polypeptides of interest.

[0004] In one aspect, the disclosure provides a high throughput method for identifying optimized polypeptides, the method comprising: (a) providing a first library of polynucleotides encoding a first library of variant polypeptides; (b) treating the first library of polynucleotides to produce the first library of variant polypeptides, wherein the variant polypeptides are attached to the first library of polynucleotides; and (c) measuring an equilibrium binding constant, kinetic binding constant, protein stability measurement, enzyme activity, fractional activity, nonspecific binding capacity, aggregation capacity, hydrophobicity, protein expression level, or maturation time of at least a portion of the first library of variant polypeptides. (d) providing a second library of polynucleotides encoding a second library of variant polypeptides selected based at least on the one or more characteristics identified in (c); (e) treating the second library of polynucleotides to produce a second library of variant polypeptides, wherein the variant polypeptides are attached to the second library of polynucleotides; and (f) analyzing the second library of variant polypeptides and generating optimized data.

[0005]

[0005] In another aspect, the disclosure provides a high-throughput method for measuring a characteristic of a polypeptide, the method comprising: (a) providing a first library of polynucleotides attached to a solid surface, where the library of polynucleotides encodes a library of variant polypeptides; (b) processing the library of polynucleotides to produce a library of variant polypeptides, where the variant polypeptides attach to the library of polynucleotides; and (c) identifying one or more characteristics of at least a portion of the library of variant polypeptides, including equilibrium binding constant, kinetic binding constant, protein stability measurement, enzymatic activity, fractional activity, nonspecific binding ability, aggregation ability, hydrophobicity, protein expression level, or maturation time.

[0006]

[0006] In another aspect, the disclosure provides a high-throughput method for screening a plurality of polypeptides, the method comprising: (a) providing a first library of polynucleotides encoding a library of variant polypeptides, wherein the first library of variant polypeptides comprises at least 90% of all single amino acid variants in which an amino acid residue is substituted with an amino acid selected from a set of 20 different amino acids; (b) processing the first library of polynucleotides to produce the first library of variant polypeptides, wherein the variant polypeptides are attached to the first library of polynucleotides; and (c) identifying one or more characteristics of the polypeptides in the first library of variant polypeptides.

[0007]

[0007] In another aspect, the disclosure provides a high-throughput method for screening a plurality of polypeptides, the method comprising the steps of: (a) providing a first library of polynucleotides encoding a first library of variant polypeptides, wherein the first library of variant polypeptides comprises single amino acid variant polypeptides corresponding to at least 90% of possible single nucleotide variants for a given reference sequence in a reference polypeptide, and for a given single amino acid variant, an amino acid residue is substituted with another amino acid selected from a set of 20 different amino acids; (b) processing the first library of polynucleotides to produce the first library of variant polypeptides, wherein the variant polypeptides are attached to the first library of polynucleotides; and (c) identifying one or more characteristics of the polypeptides of the first library of variant polypeptides.

[0008]

[0008] In some embodiments, the one or more characteristics include an equilibrium binding constant, a kinetic binding constant, a protein stability measurement, an enzymatic activity, a fractional activity, a non-specific binding capacity, an aggregation capacity, a hydrophobicity, a protein expression level, or a maturation time of at least a portion of the first library of variant polypeptides.

[0009]

[0009] In some embodiments, the method further comprises: (d) providing a second library of polynucleotides encoding a second library of variant polypeptides selected based at least on one or more characteristics identified in (c); (e) processing the second library of polynucleotides to produce a second library of variant polypeptides, wherein the variant polypeptides are attached to the second library of polynucleotides; and (f) analyzing the second library of variant polypeptides to generate optimized data. In some embodiments, the method further comprises (g) identifying an optimized polypeptide based on the optimized data. In some embodiments, the high throughput method does not comprise cells. In some embodiments, the first library of polynucleotides is a library of deoxyribonucleic acid molecules.

[0010] In some embodiments, the equilibrium binding constant is the dissociation constant (K d In some embodiments, the equilibrium binding constant is the association constant (K a In some embodiments, the kinetic binding constant is the association rate constant (k on In some embodiments, the kinetic binding constant is the dissociation rate constant (k off In some embodiments, the protein stability measurement is the protein melting temperature (T m In some embodiments, the protein stability measurement is the midpoint denaturation concentration (C m ).

[0011]

[0011] In some embodiments, the method further comprises in (d) identifying negative variants, positive variants and neutral variants from the first library of variant polypeptides. In some embodiments, the neutral variants have a dissociation constant that is greater than 0.25-fold and less than 2-fold that of the starting polypeptide. In some embodiments, the positive variants have a dissociation constant that is 0.25-fold or less than that of the starting polypeptide. In some embodiments, the negative variants have a dissociation constant that is 2-fold or more than that of the starting polypeptide.

[0012] In some embodiments, the first library of variant polypeptides comprises single amino acid variants in which an amino acid residue is replaced with one amino acid selected from the set of amino acids. In some embodiments, the set of amino acids comprises 10 different amino acids. In some embodiments, the set of amino acids comprises 20 different amino acids. In some embodiments, the set of amino acids comprises alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine and valine. In some embodiments, the first library of variant polypeptides consists of a variant of the starting polypeptide and the starting polypeptide. In some embodiments, the first library of variant polypeptides comprises double amino acid variants of interacting amino acid pairs. In some embodiments, the double amino acid variants of interacting amino acid pairs comprise variants in which the amino acid residues of the interacting amino acid pairs are replaced with all 20 amino acids. In some embodiments, the interacting amino acid pairs are identified through a crystal structure of the original polypeptide. In some embodiments, the interacting amino acid pairs include inter-polypeptide interactions and intra-polypeptide interactions. In some embodiments, the first library of variant polypeptides includes a single amino acid insertion at each position. In some embodiments, the first library of variant polypeptides includes a single amino acid deletion. In some embodiments, the first library of variant polypeptides includes a double amino acid deletion. In some embodiments, the first library of variant polypeptides includes a triple amino acid deletion. In some embodiments, the first library of variant polypeptides includes at least four amino acid deletions. In some embodiments, analyzing the first library of variant polypeptides includes transcribing and translating the polynucleotides of the first library of variant polynucleotides, wherein a polypeptide encoded by the polynucleotide is attached to the polynucleotide.In some embodiments, identifying the equilibrium binding constant, kinetic binding constant, protein stability measurement, enzymatic activity, fractional activity, non-specific binding capacity, aggregation capacity, hydrophobicity, protein expression level, or maturation time comprises performing a binding assay on the first library of variant polypeptides. In some embodiments, identifying the equilibrium binding constant, kinetic binding constant, protein stability measurement, enzymatic activity, fractional activity, non-specific binding capacity, aggregation capacity, hydrophobicity, protein expression level, or maturation time comprises sequencing the first library of polynucleotides and correlating the sequence of the first library of polynucleotides with a binding assay. In some embodiments, the binding assay comprises assaying binding of the first library of variant polypeptides to an antigen. In some embodiments, the binding assay comprises assaying binding of the first library of variant polypeptides to more than one antigen. In some embodiments, the binding assay comprises assaying binding of the first library of variant polypeptides to a plurality of antigens. In some embodiments, the method further comprises identifying variant polypeptides that bind to two or more antigens of the plurality of antigens. In some embodiments, the method further comprises identifying variant polypeptides that bind at least one antigen of the plurality of antigens and do not bind a different antigen of the plurality of antigens. In some embodiments, the method further comprises identifying variant polypeptides that do not bind to the plurality of antigens. In some embodiments, identifying the equilibrium binding constant, kinetic binding constant, protein stability measurement, enzyme activity, fractional activity, non-specific binding capacity, aggregation capacity, hydrophobicity, protein expression level, or maturation time comprises generating binding data for more than one target. In some embodiments, the second library is generated based at least on the binding data for more than one target.In some embodiments, processing the second library of variant polypeptides comprises transcribing and translating the polynucleotides of the second library of variant polynucleotides, and the polypeptides encoded by the polynucleotides are attached to the polynucleotides. In some embodiments, identifying the optimized polypeptide comprises performing a binding assay on the second library of variant polypeptides encoded by the second library of polynucleotides. In some embodiments, identifying the equilibrium binding constant, kinetic binding constant, protein stability measurement, enzyme activity, fractional activity, non-specific binding capacity, aggregation capacity, hydrophobicity, protein expression level, or maturation time comprises sequencing the second library of polynucleotides and correlating the sequences of the second library of polynucleotides with the binding assay. In some embodiments, the second library of variant polypeptides comprises at least 10. 4 In some embodiments, the first library of polynucleotides comprises at least 10 6 In some embodiments, the first library of variant polypeptides comprises at least 10 4 In some embodiments, the method is performed in less than 48 hours. In some embodiments, the first library of variant polypeptides comprises a library of distinct VHH antibodies. In some embodiments, the second library of variant polypeptides comprises a library of VHH antibody fusions. In some embodiments, the first library of variant polypeptides comprises a library of distinct single chain variable fragments (scFv). In some embodiments, the second library of variant polypeptides comprises a library of distinct single chain variable fragment (scFv) fusions.

[0013]

[0013] In another aspect, the disclosure provides a high throughput method for identifying an optimized polypeptide, the method comprising: (a) obtaining a dataset comprising binding data of an antigen to a first plurality of polypeptides, and providing a plurality of polynucleotides based at least in part on the dataset; (b) providing a plurality of polynucleotides attached to a solid surface; (c) processing the plurality of polynucleotides to produce a second plurality of polypeptides; (d) exposing an antigen to the second plurality of polypeptides and detecting an interaction of at least one polypeptide of the second plurality of polypeptides with the antigen; (e) detecting an interaction of at least one polypeptide of the second plurality of polypeptides with the antigen; (f) generating sequence data comprising (i) a sequence of a polypeptide, or (ii) a sequence of a corresponding polynucleotide encoding at least one of the polypeptides; (f) generating a plurality of fusion polypeptides based at least in part on the sequence data and the detection, wherein a fusion polypeptide of the plurality of fusion polypeptides comprises a polypeptide from each of the first plurality of polypeptides or the second plurality of polypeptides that is capable of binding to the antigen; and (g) repeating (a) through (e) to identify an optimized polypeptide, wherein the dataset comprises binding data of the antigen to the plurality of polypeptide fusions.

[0014]

[0014] In another aspect, the disclosure provides a method for identifying an optimized polypeptide, the method comprising: (a) providing a plurality of polynucleotides attached to a solid surface, the plurality of polynucleotides encoding a plurality of fusion polypeptides, a fusion polypeptide of the plurality of fusion polypeptides comprising two or more domains; (b) processing the plurality of polynucleotides to produce a plurality of fusion polypeptides; (c) exposing an antigen to the plurality of fusion polypeptides and detecting an interaction of at least one fusion polypeptide of the plurality of fusion polypeptides with the antigen; (d) generating sequence data comprising (i) a sequence of at least one fusion polypeptide, or (ii) a sequence of a corresponding polynucleotide encoding at least one fusion polypeptide; and (e) generating an optimized polypeptide capable of binding to the antigen based at least in part on a dataset comprising the sequence data, detection and binding data of the antigen to the plurality of single domain polypeptides. In some embodiments, the dataset is generated by identifying a polypeptide of a first plurality of polypeptides capable of interacting with the antigen. In some embodiments, the data set is generated by at least exposing an antigen to a first plurality of polypeptides and detecting an interaction between at least one polypeptide of the first plurality of polypeptides and the antigen. In some embodiments, the first plurality of polypeptides is generated by at least (i) providing a plurality of first polynucleotides encoding a plurality of first polypeptides; (ii) providing a plurality of first capture probes attached to a solid surface configured to anneal to the first plurality of polynucleotides, generating a plurality of captured polynucleotides; (iii) processing the plurality of captured polynucleotides to produce the first plurality of polypeptides. In some embodiments, the data related to the first plurality of polypeptides includes sequence data generated by sequencing the plurality of captured polynucleotides, and the plurality of captured polynucleotides are a plurality of VHH polynucleotides.

[0015] In some embodiments, the interaction of at least one polypeptide of the plurality of polypeptides with an antigen comprises identifying a quantitative characteristic of the polypeptide. In some embodiments, the identifying a quantitative characteristic of the polypeptide further comprises identifying the polypeptide as comprising one or more negative, neutral or positive mutations. In some embodiments, the plurality of fusion polypeptides comprises polypeptides for at least 50%, 60%, 70%, 80%, 90% or more of all possible fusion pair combinations or permutations of polypeptides of the first plurality of polypeptides. In some embodiments, the plurality of fusion polypeptides comprises polypeptides for all possible fusion pair combinations or permutations of polypeptides of the first plurality of polypeptides. In some embodiments, the dataset comprises data corresponding to single domain polypeptides corresponding to one or more domains of the fusion polypeptides. In some embodiments, the dataset is created by identifying a single domain polypeptide capable of interacting with the antigen. In some embodiments, the dataset is created by at least exposing an antigen to the plurality of single domain polypeptides and detecting an interaction of at least one single domain polypeptide of the plurality of single domain polypeptides with the antigen. In some embodiments, the multiple single domain polypeptides are generated by at least (i) providing multiple single domain polynucleotides encoding multiple single domain polypeptides, where the single domain polynucleotides are coupled to a solid surface; (iii) processing the multiple single domain polynucleotides to produce multiple single domain polynucleotide polypeptides. In some embodiments, the dataset comprises sequence data generated by sequencing the multiple single domain polynucleotides. In some embodiments, the single domain polypeptide comprises a VHH. In some embodiments, the fusion polypeptide comprises a VHH-VHH fusion. In some embodiments, the multiple fusion polypeptide comprises sequences corresponding to one or more polypeptides of the multiple single domain polypeptides.In some embodiments, a fusion polypeptide of the plurality of fusion peptides comprises sequences of two polypeptides of the plurality of single domain polypeptides. In some embodiments, the plurality of fusion polypeptides comprises polypeptides for at least 50%, 60%, 70%, 80%, 90% or more of all possible fusion pair combinations or permutations of single domain polypeptides of the plurality of single domain polypeptides. In some embodiments, the plurality of fusion polypeptides comprises polypeptides for all possible fusion pair combinations or permutations of single domain polypeptides of the plurality of single domain polypeptides. In some embodiments, the plurality of single domain polypeptides comprises a plurality of single domain polypeptides that differ by a single point mutation. In some embodiments, the plurality of single domain polypeptides comprises a plurality of single domain polypeptides that differ by a single point mutation at the binding interface. In some embodiments, the plurality of single domain polypeptides comprises a plurality of single domain antibody fragments that differ by a single point mutation in a CDR. In some embodiments, the plurality of single domain polypeptides comprises a plurality of 20 polypeptides in which a different amino acid is encoded at a given residue.

[0016] In some embodiments, detecting an interaction between at least one single domain polypeptide of the plurality of single domain polypeptides and an antigen comprises identifying a quantitative characteristic of the single domain polypeptide. In some embodiments, identifying a quantitative characteristic of the polypeptide further comprises identifying the single domain polypeptide as comprising one or more negative, neutral or positive mutations. In some embodiments, detecting an interaction between at least one fusion polypeptide of the plurality of fusion polypeptides and an antigen comprises identifying a quantitative characteristic of the fusion polypeptide. In some embodiments, identifying a quantitative characteristic of the polypeptide further comprises identifying the fusion polypeptide as comprising a bi-epitopic interaction. In some embodiments, identifying the fusion polypeptide as comprising an interaction with enhanced avidity comprises comparing the quantitative characteristic of the fusion polypeptide to a quantitative characteristic of the first single domain or the second single domain, wherein the sequence of the fusion polypeptide comprises the sequence of the first single domain and the sequence of the second single domain. In some embodiments, an interaction with enhanced avidity is identified when a quantitative characteristic of the fusion polypeptide exceeds a quantitative characteristic of the first single domain or the second single domain. In some embodiments, the optimized polypeptide comprises additional mutations of the fusion polypeptide identified as comprising an interaction with enhanced avidity, the mutations improving the binding affinity of the fusion polypeptide to the antigen. In some embodiments, the data comprising antigen binding data to the plurality of single domain polypeptides is obtained at the same time that (c) or (d) is performed. In some embodiments, the data comprising antigen binding data to the plurality of single domain polypeptides is obtained prior to (a), and the step of providing a plurality of polynucleotides attached to a solid support is based at least in part on the data set.

[0017] In some embodiments, the plurality of fusion polypeptides comprises a sequence of a single domain polypeptide having a moderate affinity to the antigen. In some embodiments, the plurality of fusion polypeptides comprises a sequence of a single domain polypeptide having minimal or no affinity to the antigen. In some embodiments, the sequence of a single domain polypeptide having minimal or no affinity has a substantially similar size or length as a single domain polypeptide capable of binding to the antigen. In some embodiments, the sequence of a single domain polypeptide having minimal or no affinity has a size or length difference of 10% or less from a single domain polypeptide capable of binding to the antigen. In some embodiments, one single domain polypeptide of the plurality of single domain polypeptides comprises an N-terminal linker or a C-terminal spacer. In some embodiments, one single domain polypeptide of the plurality of single domain polypeptides comprises an N-terminal linker and a C-terminal spacer. In some embodiments, the plurality of single domain polypeptides comprises a plurality of different N-terminal linker sequences and different C-terminal spacer sequences. In some embodiments, the data set is derived from data in a public database.

[0018] In some embodiments, the fusion polypeptide is a polypeptide-Fc fusion. In some embodiments, the polypeptide-Fc fusion comprises an antibody fragment crystallization region (Fc region) capable of binding to an antigen. In some embodiments, the fusion polypeptide comprises a chimeric antigen receptor. In some embodiments, the fusion polypeptide comprises a VHH nanobody. In some embodiments, the fusion polypeptide comprises a pair of bivalent VHH nanobodies. In some embodiments, the fusion polypeptide comprises a pair of bi-epitope VHH nanobodies. In some embodiments, the fusion polypeptide comprises a multivalent VHH nanobody. In some embodiments, the fusion polypeptide comprises a linker connecting a first domain of the fusion polypeptide and a second domain of the fusion polypeptide. In some embodiments, the first domain comprises a VHH. In some embodiments, the second domain comprises a VHH. In some embodiments, the first domain comprises a first VHH and the second domain comprises a second VHH. In some embodiments, the first VHH and the second VHH bind to the same antigen. In some embodiments, the same antigen comprises a polypeptide, a lipid or carbohydrate, or a cell. In some embodiments, the linker comprises at least 12 amino acids. In some embodiments, the linker comprises at least 20 amino acids. In some embodiments, the linker comprises at least 30 amino acids. In some embodiments, the linker has a net positive charge. In some embodiments, the linker has a net negative charge. In some embodiments, the linker has a net neutral charge.

[0019] In some embodiments, the plurality of polynucleotides is at least 10 4In some embodiments, the optimized polypeptide has an increased binding activity effect. In some embodiments, prior to (a), the solid surface comprises a plurality of capture oligonucleotides configured to anneal to the plurality of precursor polynucleotides, and the plurality of precursor polynucleotides anneal to the plurality of capture oligonucleotides, thereby creating a plurality of polynucleotides attached to the solid surface. In some embodiments, creating a plurality of polynucleotides attached to the solid surface comprises amplification or extension of the plurality of precursor polynucleotides. In some embodiments, the amplification comprises bridge amplification. In some embodiments, the solid support comprises beads. In some embodiments, the solid support comprises a sequencing flow cell.

[0020]

[0020] In some embodiments, (d) comprises sequencing the plurality of polynucleotides. In some embodiments, (e) comprises generating an optimized polypeptide based at least in part on sequence data generated from sequencing and detecting the plurality of polynucleotides. In some embodiments, a fusion polypeptide of the plurality of fusion polypeptides comprises an N-terminal linker or a C-terminal spacer. In some embodiments, a fusion polypeptide of the plurality of fusion polypeptides comprises an N-terminal linker and a C-terminal spacer. In some embodiments, a fusion polypeptide comprises a plurality of different N-terminal linker sequences and different C-terminal spacer sequences. In some embodiments, the optimized polypeptide comprises a bi-epitopic polypeptide. In some embodiments, the optimized polypeptide comprises a tri-epitopic polypeptide. In some embodiments, the optimized polypeptide comprises a tetra-epitopic polypeptide. In some embodiments, the optimized polypeptide comprises a multimeric polypeptide. In some embodiments, the optimized polypeptide comprises two or more domains capable of binding to an antigen, and at least two domains are identical. In some embodiments, an optimized polypeptide comprises two or more domains capable of binding to an antigen, and the two or more domains are distinct from one another.

[0021] In another aspect, the disclosure provides a method for identifying a bi-epitopic polypeptide, comprising the steps of: (a) providing a plurality of polynucleotides attached to a solid surface, wherein the plurality of polynucleotides encodes a plurality of VHH polypeptides; (b) processing the plurality of polynucleotides to produce a plurality of VHH polypeptides; (c) exposing an antigen to the plurality of polypeptides, and detecting an interaction of at least one VHH polypeptide of the plurality of VHH polypeptides with the antigen; (d) sequencing the plurality of polynucleotides; and (e) providing a second plurality of polynucleotides attached to a solid surface, wherein the second plurality of polynucleotides encodes a plurality of VHH polypeptides. The present invention provides a method comprising the steps of: (a) encoding a VHH fusion polypeptide; (b) processing a plurality of second polynucleotides to produce a plurality of VHH-VHH fusion polypeptides; (c) exposing an antigen to the plurality of VHH-VHH fusion polypeptides and detecting an interaction of at least one VHH-VHH fusion polypeptide of the plurality of VHH-VHH fusion polypeptides with the antigen; (d) sequencing the second plurality of polynucleotides; and (e) generating a bi-epitopic polypeptide capable of binding to the antigen based at least in part on sequence data generated from the sequencing steps of (d) and (e), and the detecting steps of (c) and (g).

[0022]

[0022] In another aspect, the disclosure provides a method for generating optimized polypeptides, the method comprising: (a) providing a plurality of polypeptides displayed on a solid substrate, wherein one polypeptide of the plurality of polypeptides comprises a binding domain and one or more of (i) an N-terminal spacer, (ii) a C-terminal spacer, and wherein the plurality of polypeptides comprises polypeptides comprising various combinations of N-terminal spacer sequences and C-terminal spacer sequences; (b) observing a signal of at least two polypeptides of the plurality of polypeptides, wherein the signal corresponds to (i) a binding interaction between the polypeptide and an antigen, or (ii) a physical characteristic of the polypeptide; (c) comparing the signals of the at least two polypeptides, and determining a combination of N-terminal spacer sequences and C-terminal spacer sequences that produces a target signal.

[0023]

[0023] In some embodiments, the N-terminal spacer or the C-terminal spacer does not bind to an antigen. In some embodiments, the target signal comprises a signal below a threshold level. In some embodiments, the target signal comprises a signal above a threshold level. In some embodiments, the target signal comprises the highest signal of the signals of a plurality of polypeptides. In some embodiments, the target signal comprises the lowest signal of the signals of a plurality of polypeptides.

[0024]

[0024] In some embodiments, the signal corresponds to an equilibrium binding constant, a kinetic binding constant, a protein stability measurement, an enzyme activity, a fractional activity, a non-specific binding capacity, an aggregation capacity, a hydrophobicity, a protein expression level, or a maturation time of a polypeptide.

[0025] In another aspect, the disclosure provides a method for discovering improved pairs of binders, comprising: (a) providing a comprehensive data set including: (i) measured quantitative binding characteristics for a plurality of polypeptides comprising two domains, the two domains being independently selected from a set of monomeric domains, the plurality of polypeptides including all possible pairs of monomeric polypeptides; and (ii) measured quantitative binding characteristics for each monomeric domain of the set of monomeric domains as individual monomeric polypeptides; (b) comparing the values ​​of (i) and (ii) to identify polypeptides including improved pairs of binders that exhibit a quantitative binding characteristic significantly greater than the binding characteristic of any of the individual monomeric polypeptide components. In some embodiments, the improved pairs of binders are bi-epitope binders. In some embodiments, the comprehensive data set includes measured quantitative binding characteristics of the set of individual monomeric polypeptides, and measured quantitative binding characteristics for at least 50%, 60%, 70%, 80%, 90% or more of all possible random pair combinations of the set of individual monomeric polypeptides. In some embodiments, the comprehensive dataset comprises measured quantitative binding characteristics for the set of individual monomeric polypeptides, and measured quantitative binding characteristics for all possible random pairwise combinations of the set of individual monomeric polypeptides.

[0026]

[0026] In another aspect, the disclosure provides a high-throughput method for identifying tandem polypeptides optimized for affinity and binding activity, comprising the steps of: (a) providing a first library of polynucleotides encoding a first library of monomeric variant polypeptides; (b) processing the first library of polynucleotides to produce a first library of variant polypeptides, wherein the variant polypeptides are attached to the first library of polynucleotides; (c) analyzing the first library of variant polypeptides and generating data; (d) identifying binding affinities of at least some of the first library of variant polypeptides based on the data; (e) providing a second library of second polynucleotides encoding a second library of monomeric variant polypeptides from the first library based on the binding data from the first library; and (f) comprising various combinations of monomeric variant polypeptides corresponding to the first library. The present invention provides a method comprising the steps of: (a) providing a third library of polynucleotides encoding a plurality of tandem polypeptides, wherein a tandem polypeptide of the plurality of tandem polypeptides comprises a first monomeric variant polypeptide and a second monomeric variant polypeptide; (g) treating the second and third libraries of polynucleotides to produce a second and third library of variant polypeptides, wherein the variant polypeptides are attached to the second and third libraries of polynucleotides; (h) analyzing the second and third libraries of variant polypeptides to identify affinity-enhanced monomeric polypeptide variants and avidity-enhanced tandem polypeptides; and (i) combining the enhanced avidity and affinity identified in the second and third libraries by substituting individually optimized monomers identified in the second library into corresponding positions in the avidity-enhanced tandem pairs discovered from the second library.In some embodiments, the third library comprises a plurality of polypeptides comprising different linkers between the first and second monomeric variant polypeptides. In some embodiments, the third library comprises monomeric variant polypeptides that have reduced affinity compared to the reference polypeptide based on the binding data from the first library.

[0027] In another aspect, the disclosure provides a composition comprising an array of polypeptides displayed on a solid surface, each polypeptide co-localized with a corresponding polynucleotide encoding the polypeptide, a polypeptide in the plurality of polypeptides comprises a first domain and a second domain, the first domain and the second domain are linked via a linker, the first domain binds a first epitope and the second domain binds a second epitope, and the first epitope and the second epitope are different. The composition may comprise an array of polypeptides, including a polypeptide library as described elsewhere herein.

[0028]

[0028] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, in which merely exemplary embodiments of the present disclosure are shown and described. As will be understood, the present disclosure is capable of other different embodiments, and its several details can be modified in various obvious aspects, all without departing from the present disclosure. Thus, the drawings and descriptions are to be regarded as illustrative in nature, and not as limiting.

[0029] Incorporation by Reference

[0029] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that a publication, patent, or patent application incorporated by reference conflicts with a disclosure contained in the specification, it is intended that the specification take precedence and / or supersede any such conflicting material.

[0030]

[0030] The novel features of the invention are set forth with particularity in the appended claims. "The patent or application file contains at least one figure executed in color. Copies of this patent or patent application publication containing the color figure(s) will be provided by the Patent Office upon request and payment of the necessary fee. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description and accompanying figures (also referred to herein as "figures" and "Figs") which set forth illustrative embodiments in which the principles of the invention are utilized: [Brief description of the drawings]

[0031] [Figure 1-1] FIG. 1A shows a schematic of nanobody sequences for initial display selection. [Figure 1-2] FIG. 1B shows a representative nanobody library displayed using ribosome display. [Diagram 2]

[0032] FIG. 2 shows a schematic diagram of the disclosed method by which DNA libraries are generated and quantified. [Diagram 3]

[0033] FIG. 3 shows a heat map of single mutations in the CDR regions. [Figure 4]

[0034] FIG. 4 shows a schematic diagram of the disclosed method in which a DNA library is generated and quantified, followed by the generation and quantification of a new library based on the analysis of the previous library. [Diagram 5]

[0035] FIG. 5 shows data regarding polypeptides produced by the methods of the disclosure. [Figure 6]

[0036] FIG. 6 presents data regarding a selection of polypeptides produced by the methods of the disclosure. [Figure 7]

[0037] FIG. 7 shows a schematic diagram of a polypeptide that can be produced using the methods of the present disclosure. [Figure 8]

[0038] FIG. 8 shows a schematic diagram of a multispecific or selective polypeptide. [Figure 9]

[0039] FIG. 9 shows a workflow schematic for the generation of bi-epitopic polypeptides. [Figure 10]

[0040] FIG. 10 shows a heat map of binding data for single mutants in the CDR regions of representative VHHs in the dataset. [Figure 11]

[0041] FIG. 11 shows a schematic of the design of a DNA library encoding tandem VHHs that can be expressed on a chip, assayed for binding, and analyzed to find enhanced binding activity using the methods of the present disclosure. [Figure 12-1]

[0042] FIG. 12A shows avidity enhancement data generated for a specific tandem VHH pair using the methods of the present disclosure. [Figure 12-2] FIG. 12B shows a heat map of the binding activity enhancement for all tandem VHH pairs in both orientations of the experiment. [Figure 13]

[0043] Figure 13A shows the distribution of the number of mutations in a VHH affinity-optimized library generated using the methods of the present disclosure. Figure 13B shows data for affinity-optimized VHHs generated against two different targets using the methods of the present disclosure. [Figure 14]

[0044] FIG. 14 shows a workflow schematic for the generation of affinity-optimized and avidity-enhanced multivalent tandem VHH pairs. [Figure 15-1]

[0045] 15A-15C show: (15A) sequential ("two-step") optimization using the methods of the present disclosure; (15B) a workflow diagram for discovery of tandem polypeptide pairs with enhanced binding activity; and (15C) a combinatorial workflow for discovery of affinity-optimized molecules formatted in a tandem configuration with high binding activity. [Figure 15-2]

[0045] Figures 15A-15C show (15A) sequential ("two-step") optimization using the methods of the present disclosure, (15B) a workflow diagram for discovery of tandem polypeptide pairs with enhanced binding activity, and (15C) a combinatorial workflow for discovery of affinity-optimized molecules formatted in a tandem configuration with high binding activity. [Figure 16]

[0046] 1 illustrates a computer control system programmed or otherwise configured to carry out the methods provided herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0032]

[0047] While various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be used.

[0033]

[0048] The present disclosure provides methods, systems and compositions for the generation of polypeptide libraries, as well as methods, systems and compositions for displaying libraries to identify or determine polypeptide characteristics. The approaches described herein can be effective for the optimization or generation of polypeptides with specific characteristics. Specifically, the approaches can be used to generate antibodies or antibody fragments that are capable of binding to antigens at low concentrations. The methods described herein allow for highly multiplexed quantitative assays that can result in the generation of data that is otherwise difficult to rapidly obtain. This data can be exploited and used to guide subsequent iterations of the methods described, or can be combined with other data generated to create polypeptides that can be optimized to have multiple characteristics. The methods can be performed iteratively to guide the construction of later iterations using data gathered by previous iterations to rapidly and efficiently identify polypeptides with extreme or rare functions. The generation of large data sets can be exploited to build polypeptides that other methods, such as directed evolution, are unable to identify. Due to the size of the sequence space that needs to be analyzed to identify polypeptides of interest, there is a need to analyze large amounts of promising polypeptides and generate quantitative data in a rapid, adjustable, and customizable manner.

[0034] Polypeptide library construction

[0049] In various aspects of the present disclosure, a polypeptide library is constructed. To identify and generate polypeptides with specific properties of interest, the polypeptide library can be constructed based on a set of parameters. Using the polypeptide library display methods described elsewhere herein, the polypeptide library can be subjected to analysis.

[0035]

[0050] In some embodiments, the polypeptide library comprises wild-type or reference polypeptides. In some embodiments, the polypeptide library may comprise variants of wild-type or reference polypeptides. The variants may comprise substitution mutations, insertions or deletions. The polypeptide library may comprise polypeptide variants having mutations at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more amino acids. The polypeptide library may comprise polypeptides corresponding to all possible single-point substitution variants for a single residue. A single-point mutation may comprise replacing an amino acid with another amino acid selected from a set of amino acids. The set of amino acids can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 or more amino acids. The set of amino acids can include alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, and valine. The set of amino acids may include alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, valine, or combinations thereof. For example, a polypeptide library may include 20 polypeptides (e.g., based on 20 standard amino acids) where the amino acid at the first residue is a different amino acid and all other amino acids are the same. In this manner, a polypeptide library may be analyzed to generate data regarding how an amino acid at a particular residue number may affect the properties of a polypeptide.A polypeptide library can include polypeptides that correspond to single-point substitutions at 20 different amino acids at all residues in a polypeptide.For example, for a 100 amino acid long polypeptide, 20 variants at each residue are generated for each standard amino acid, resulting in 2,000 (20x100) different polypeptides.Using this approach, a polypeptide library can be analyzed to generate data about how amino acids at specific residue numbers can affect the properties of a polypeptide over the entire length of the polypeptide.

[0036]

[0051] A polypeptide library may include polypeptides corresponding to single-point substitutions at 20 different amino acids at all residues in a region of a polypeptide. For example, a particular domain of a polypeptide may be associated with a function, such as binding to an antigen or other target. A polypeptide library may include polypeptides corresponding to single-point substitutions at 20 different amino acids at residues specific to a particular domain. For example, the polypeptide may be an antibody or antibody fragment, and the particular domain may be a complementarity determining region (CDR). A polypeptide library may include polypeptides corresponding to at least 80% of all single-point substitutions at 20 different amino acids at all residues in a region of a polypeptide. A polypeptide library may include polypeptides corresponding to at least 90% of all single-point substitutions at 20 different amino acids at all residues in a region of a polypeptide. A polypeptide library may include polypeptides corresponding to at least 95% of all single-point substitutions at 20 different amino acids at all residues in a region of a polypeptide. A polypeptide library may include polypeptides corresponding to at least 99% of all single-point substitutions at 20 different amino acids at all residues in a region of a polypeptide. The polypeptide library may include polypeptides corresponding to at least 80% of all single-point substitutions with the 20 amino acids at all residues in the polypeptide. The polypeptide library may include polypeptides corresponding to at least 90% of all single-point substitutions with the 20 amino acids at all residues in the polypeptide. The polypeptide library may include polypeptides corresponding to at least 95% of all single-point substitutions with the 20 amino acids at all residues in the polypeptide. The polypeptide library may include polypeptides corresponding to at least 99% of all single-point substitutions with the 20 amino acids at all residues in the polypeptide. The amino acids may include alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, and valine.

[0037]

[0052] The polypeptide library may be constructed at least based on structural data. The structure of the reference (or variant) polypeptide may be generated or may have already been generated. The structure may be generated based on a structure determination method, such as X-ray crystallography or nuclear magnetic resonance (NMR) spectroscopy, or other methods that reveal structural information. Using the structural data of the polypeptide, residues may be identified as interacting with other residues. The polypeptides of the polypeptide library may be generated based on information related to the interaction of residues by a structural model. For example, a reference polypeptide model may show an interaction between residue A and residue B. The polypeptide library may include double variants in which residue A and residue B are variants compared to the reference or wild-type polypeptide. This may be such that for each variant amino acid at residue A, all possible amino acid variants at residue B are generated, and vice versa. For a given residue A and residue B, 400 polypeptides (20 possible amino acids at residue A x 20 possible amino acids at residue B) may be generated. Using this approach, polypeptide libraries can be analyzed to generate data regarding how interacting amino acids at particular residue numbers can affect the properties of the polypeptide.

[0038]

[0053] The polypeptide of the polypeptide library can also correspond to the deletion of amino acids compared to wild type or reference polypeptide.Polypeptides can include deletion variants, where any single amino acid or group of amino acids is deleted.Polypeptides can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more amino acid deletions. A polypeptide may contain deletions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more contiguous amino acids. The deletions may be located in any part of the polypeptide chain.

[0039]

[0054] The polypeptides of the polypeptide library can also correspond to the insertion of amino acids compared to wild type or reference polypeptides.Polypeptides can include any single amino acid or group of amino acids inserted into the insertion variant.Polypeptides can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more amino acid insertions. A polypeptide may contain an insertion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more contiguous amino acids. The insertion may be located in any part of the polypeptide chain.

[0040]

[0055] The polypeptide library may include a combination of the polypeptide libraries described elsewhere herein. For example, the polypeptide library may include polypeptides that include insertion variants and polypeptides with single point substitution variants.

[0041]

[0056] A polypeptide library can be generated based on data generated from a polypeptide library described elsewhere herein. For example, a first polypeptide library can be generated corresponding to single-point substitutions for a particular domain of a polypeptide. The polypeptide library can be subjected to an assay in which binding to a particular antigen is analyzed. Data corresponding to binding of the polypeptides in the library can demonstrate that a particular single-point substitution variant can increase or decrease or maintain binding similarly compared to a reference or wild-type polypeptide. Using the data, a polypeptide containing multiple single-point substitution variants can be generated. For example, data for a polypeptide can show: (1) a single-point variant of residue A to amino acid X can increase binding; and (2) a single-point variant of residue B to amino acid Y can increase binding. Polypeptides can be generated and assayed for a polypeptide library containing a first single-point variant of residue A to amino acid X and a second single-point variant of residue B to amino acid Y. The synergistic effect of the variants can be analyzed, allowing for the generation of polypeptides with improved characteristics. A polypeptide library may include polypeptides that contain combinations of variants that are determined to improve or maintain the characteristics of the polypeptide. For example, 10 variants may be shown to have improved or neutral binding to antigen. A polypeptide library that contains combinations of 10 variants may be generated, where a first polypeptide may have any 2 variants of 10 possible variants, a second polypeptide may have any 3 variants of 10 possible variants, etc.

[0042]

[0057] These library construction approaches can be used iteratively to create a multi-step / multi-library approach to optimize or generate polypeptides with specific characteristics. A first library can be generated and assayed to determine the characteristics of the polypeptides of the first polypeptide library. Using the data generated, a second polypeptide library can be constructed taking into account the data, e.g., how variants affect the characteristics. A second library can be assayed and data can be generated to identify polypeptides with specific characteristics. This can be repeated, e.g., where a third library is generated based on data generated from the second library, or an n+1 library is generated from data generated from the n library (or other libraries). In addition, the data for the libraries can be analyzed by algorithms or used as a training set for predictive algorithms or machine learning, such as to identify variants of interest for use in the next library.

[0043]

[0058] The library may be constructed from sequences analyzed in a previously created library or from other data sources. For example, the library may be created by combining polypeptides analyzed in a previously generated library. A first library may be created that includes a plurality of polypeptides that bind to a given antigen. A second library may use one or more sequences of a plurality of polypeptides from the first library in combination with another sequence of a plurality of polypeptides from the first library. The first library may include a plurality of different scaffolds with characteristics. The second library may include a plurality of fusions of the various scaffolds analyzed in the first library. The first library may include a plurality of binding polypeptides with various structures or point mutations. The second library may include bivalent or biepitope polypeptides that include combinations of binding polypeptides from the first library. The second library may include bivalent or biepitope polypeptides that include all combinations of binding polypeptides from the first library. The second library may include bivalent or biepitope polypeptides that include all permutations of binding polypeptides from the first library.

[0044]

[0059] A library of polypeptides can be generated from a corresponding library of polynucleotides. 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 The library may contain 10 or more polynucleotides. 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 The library may contain at least 10 or more polypeptides. 3 , 10 4 , 10 5 , 106 , 10 7 , 10 8 , 10 9 A library may contain at least 10 or more polynucleotides on a single substrate, sequencing chip, or in a sample volume. 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 Or more than this number of polypeptides may be contained on a single substrate, sequencing chip or in a sample volume.

[0045]

[0060] A polypeptide can be any polymer composed of amino acids. A polypeptide can bind to another molecule, react (physically or chemically), transmit a signal, act as a structural component, transduce or cause other functions. A polypeptide can be an antibody or a fragment or fragments of an antibody. For example, a polypeptide can be a single chain variable fragment (scFv) or a nanobody (e.g., VHH).

[0046]

[0061] The methods described in this disclosure can be used to identify or generate polypeptides with specific or improved characteristics. The methods described can be performed on any reference or wild-type sequence to generate a library of polypeptides. The methods can consider any reference polypeptide with optimized functions to have improved functions. The specific characteristic can be the stability of the polypeptide. The specific characteristic can be the enzymatic rate or other reaction parameters. The specific characteristic can include at least a specific binding affinity or dissociation constant to a molecule. For example, the methods described can be used to generate antibodies or antibody fragments with high affinity to a target. The generated polypeptides can have a binding affinity to an antigen or target of less than 1 nM. The generated polypeptides can have a binding affinity to an antigen or target of less than or equal to 100 nM, 10 nM, 1 nM, 100 pM, 10 pM, 1 pM or less.

[0047]

[0062] The generated polypeptide may have an improved binding affinity measurement compared to a reference or wild-type polypeptide. For example, the binding affinity measurement may have a 10% improvement compared to a reference or wild-type polypeptide. For example, the binding affinity measurement may have a 25% improvement compared to a reference or wild-type polypeptide. For example, the binding affinity measurement may have a 50% improvement compared to a reference or wild-type polypeptide. For example, the binding affinity measurement may have a 75% improvement compared to a reference or wild-type polypeptide. For example, the binding affinity measurement may have a 100% improvement compared to a reference or wild-type polypeptide. For example, the binding affinity measurement may have a 200% improvement compared to a reference or wild-type polypeptide. For example, the binding affinity measurement may have a 300% improvement compared to a reference or wild-type polypeptide. For example, the binding affinity measurement may have a 400% improvement compared to a reference or wild-type polypeptide. For example, the binding affinity measurement may have a 500% improvement compared to a reference or wild-type polypeptide. For example, the binding affinity measurement may have a 1,000% improvement compared to a reference or wild-type polypeptide. For example, the binding affinity measurement may have a 100-fold improvement compared to the reference or wild-type polypeptide. For example, the binding affinity measurement may have a 1000-fold improvement compared to the reference or wild-type polypeptide. For example, the binding affinity measurement may have a 10,000-fold improvement compared to the reference or wild-type polypeptide. For example, the binding affinity measurement may have a 100,000-fold improvement compared to the reference or wild-type polypeptide. For example, the binding affinity measurement may have a 1,000,000-fold improvement compared to the reference or wild-type polypeptide. The generated polypeptide may be a polypeptide with enhanced binding activity.

[0048]

[0063] Avidity generally refers to the accumulation of the strength of multiple distinct non-covalent interactions between a binding molecule and an antigen, resulting in an increase in the binding affinity measurement. The avidity effect can result in an increase in the local concentration (of the antigen or binding molecule) by having multiple antigen binding sites that interact with the antigen. A single binding interaction may be disrupted, releasing the antigen and no longer interacting with the binding molecule, while a molecule with multiple binding sites (and multiple distinct non-covalent interactions) may maintain antigen binding even when individual binding interactions are disrupted. A polypeptide with enhanced avidity may have multiple distinct binding interactions, such as a bi-epitope binder that is capable of binding to two different epitopes. Similarly, a mono-epitope multimeric binder can maintain antigen binding by "swapping" the antigen between binding sites, effectively increasing the local concentration of binding sites, thereby increasing the binding affinity measurement.

[0049] Polypeptide Library Display

[0064] In various aspects of the present disclosure, polypeptides are generated and displayed as a library. Methods of displaying polypeptide libraries may incorporate methods that can associate genotypes and corresponding phenotypes. One such method for peptide display may include ribosome-based display methods. Display methods using ribosomes include those described in U.S. Patent Application Publication No. US2020 / 0048629 and U.S. Patent No. 10,011,830, which are incorporated herein by reference. Display methods may include a polypeptide displayed as a ribosomal translation product (e.g., a protein or peptide, biologically active fragments thereof or other ribosomally translated molecules) on a DNA template that encodes it. The DNA template may include a promoter operably linked to the open reading frame (ORF). The DNA template may further include a molecular roadblock that blocks the progression of RNA polymerase during transcription of the DNA template. The molecular roadblock may result in the stalling of RNA polymerase during transcription so that the DNA template and the transcribed mRNA remain associated. During translation of an RNA transcript, an RNA polymerase stalled at a molecular roadblock can block continued translation of the ribosome so that the ribosome remains associated with the RNA transcript and presents a nascent peptide chain (e.g., a protein or peptide, biologically active fragment thereof or other ribosomally translated molecule).Optionally, the single-stranded mRNA produced by transcription of a DNA template can be cleaved proximal to the ribosome after the ribosome reaches the molecular roadblock.

[0050]

[0065] A molecular obstacle may comprise one or more molecular configurations downstream of a transcribable region of DNA that are positioned such that when an RNA polymerase in the process of transcription encounters the obstacle, the polymerase stops and forms a stable complex that includes the RNA polymerase, the DNA template, and the nascent RNA transcript. An obstacle may be a molecular entity that associates with DNA covalently or non-covalently, or a chemical modification to DNA, such as a chemical crosslink between DNA strands that stops the RNA polymerase. An obstacle may be located at the 5' end of the antisense DNA strand or the 3' end of the sense DNA strand, or both. An obstacle may also comprise a molecule that selectively binds to a specific sequence of DNA at an appropriate position. In one embodiment, a molecular obstacle is formed by the binding of streptavidin following biotinylation of DNA at either the 3' end of the sense strand or the 5' end of the antisense strand, where the biotin-streptavidin complex acts as a molecular obstacle that blocks the RNA polymerase.

[0051]

[0066] In addition, the DNA template may code for an mRNA with a ribosome stop sequence. In certain embodiments, the ribosome stop sequence includes a stop codon (e.g., UAG (amber), UAA (ochre), or UGA (opal or umber) in the mRNA). In another embodiment, the ribosome stop sequence further includes a polyproline coding sequence adjacent to the stop codon. In one embodiment, the polyproline coding sequence includes a coding sequence for a triple proline motif, where the coding sequence for the triple proline motif is located before (i.e., 5') the stop codon. In another embodiment, the ribosome stop sequence further includes an arginine-histidine-arginine coding sequence adjacent to the polyproline coding sequence (e.g., a triple proline motif), where the arginine-histidine-arginine coding sequence is located before (i.e., 5') the polyproline coding sequence. The ribosome display method may also be performed under conditions that result in ribosome stalling. For example, amino acid starvation of the ribosome can be used, which can be achieved by limiting the amount of a particular amino acid (or tRNA or other related reagent) such that the ribosome cannot add the next amino acid to the elongating nascent peptide, thereby stalling the ribosome.

[0052]

[0067] The mRNA may further comprise a Shine-Dalgarno sequence, which may be optimized for a particular ORF of interest to promote efficient ribosome binding and translation initiation.

[0053]

[0068] The polynucleotides used in this disclosure may be derived from any nucleic acid of known or unknown sequence, for example, fragments of genomic DNA or cDNA. For example, the polynucleotides may be derived from a randomly fragmented primary nucleic acid sample. The polynucleotides may also be obtained by reverse transcription from a primary RNA sample to cDNA. Individual polynucleotides may include whole genes or parts of genes or cDNAs derived from mRNA that code for proteins or peptides or biologically active polypeptides or peptide fragments thereof. Additionally, the polynucleotides may include recombinantly engineered constructs. The polynucleotides may code for the polypeptides described throughout this disclosure. For example, the polynucleotides may code for nanobodies or scFvs.

[0054]

[0069] Protein translation can be carried out using in vitro cell-free expression systems. Translation can be carried out in vitro using crude lysates from any organism that provide all the components necessary for translation, including enzymes, tRNAs and cofactors (except for termination factors), amino acids and energy supply (e.g., GTP). Cell-free expression systems from Escherichia coli, wheat germ and rabbit reticulocytes are commonly used. Although E. coli-based systems provide high yields, eukaryotic-based systems are preferred for producing proteins that are post-translationally modified. Alternatively, artificially reconstituted cell-free systems can be used for protein production. For optimal protein production, the codon usage in the ORF of the DNA template can be optimized for expression in the particular cell-free expression system selected for protein translation. In addition, labels or tags can be added to the protein to facilitate high-throughput screening. See, e.g., Katzen et al. (2005) Trends Biotechnol. 23:150-156; Jermutus et al. (1998) Curr. Opin. Biotechnol. 9:534-548; Nakano et al. (1998) Biotechnol. Adv. 16:367-384; Spirin (2002) Cell-Free Translation Systems, Springer; Spirin and Swartz (2007) Cell-free Protein Synthesis, Wiley-VCH; Kudlicki (2002) Cell-Free Protein Expression, Landes Bioscience; each of which is incorporated by reference in its entirety.

[0055]

[0070] In certain embodiments, protein translation is carried out using an in vitro cell-free expression system that lacks one or more release factors so that ribosomes are not released from a stop codon on mRNA. One or more release factors may be absent, including release factor 1 (RF1), release factor 2 (RF2) and release factor 3 (RF3), or all release factors may be absent in the in vitro cell-free expression system. The release factor that is absent may depend on the stop codon selected for inclusion in the stop sequence. For example, RF1 normally mediates the release of ribosomes from RNA transcripts at amber codons. Thus, when an amber codon is included in the stop sequence, RF1 may be omitted from the in vitro cell-free expression system. On the other hand, RF2 normally mediates the release of ribosomes from RNA transcripts at either ochre or opal codons. Thus, when an ochre or opal codon is included in the stop sequence, RF2 may be omitted from the in vitro cell-free expression system. In some embodiments, protein translation is carried out using an in vitro cell-free expression system that lacks all termination factors. Additionally, ribosome recycling factors (RRFs) can also be omitted from the in vitro cell-free expression system to prevent the release of stalled ribosomes from the transcribed RNA molecules.

[0056]

[0071] In some embodiments, one or more non-standard amino acids, such as, but not limited to, D-amino acids, beta amino acids, or N-substituted glycines (peptoids), are incorporated into the ribosomal translation product. The non-standard amino acids can be introduced into a protein or peptide in either a residue-specific or site-specific manner. See, for example, Link et al., (2003) Curr. Opin. Biotechnol. 14(6):603-609; Johnson et al., (2010) Curr. Opin. Chem. Biol. 14(6):774-780; Zheng et al., (2012) Biotechnol J. 7(1):47-60; which are incorporated herein by reference.

[0057]

[0072] In some embodiments, the method of polypeptide display can include providing conditions on polynucleotide that allow only one RNA polymerase to start translation. For example, the DNA template can further include a stop sequence, where the first RNA polymerase that starts transcription stops at a position on the DNA template so that any other polymerase starts to start. Transcription is carried out under conditions of nucleotide starvation, where the RNA polymerase stops at a specific position on the DNA template because the nucleotides required for addition at that position are not provided (see, for example, Greenleaf and Block (2006) Science 313(5788):801; incorporated herein by reference). After the RNA polymerase stops, all unbound polymerases are removed, for example, by washing, and then the missing nucleotides required to resume transcription are added so that the one remaining RNA polymerase bound to the DNA template can continue transcription until it stops at a molecular obstacle. Alternatively, unbound RNA polymerase can be inactivated (eg, using heparin) rather than removed to ensure that only one RNA polymerase remains bound to the DNA template.

[0058]

[0073] In some embodiments, the method of polypeptide display can further comprise providing conditions that allow only one ribosome to start translation on the RNA transcript.For example, translation can be carried out under conditions of amino acid starvation, where ribosome stops at a specific position on the RNA transcript because the amino acid required for addition at that position is not provided.Then, all unbound ribosomes can be removed, for example, by washing, and the missing amino acid required to resume translation can be added to allow one bound ribosome to continue translation until it reaches the ribosome stop sequence.

[0059]

[0074] The ribosome translation product may contain one or more linkers or spacers, for example, to facilitate ribosome presentation, cloning, purification or detection, or to improve solubility. For example, short flexible linkers or spacers having 20 or fewer amino acids (i.e., 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2 or 1) are useful for separating domains in fusion constructs. Examples include short peptide sequences such as polyglycine linkers (Glyn, where n=2, 3, 4, 5, 6, 7, 8, 9, 10 or more), histidine tags (Hisn, where n=3, 4, 5, 6, 7, 8, 9, 10 or more), linkers composed of glycine and serine residues, soluble polypeptide linkers, GSAT, SEG and Z-EGFR linkers. Longer linkers with defined tertiary structures can be used to facilitate the presentation of proteins or peptides on the ribosome. Such linkers include, but are not limited to, a fragment of gene III from the filamentous phage M13mp192, a portion of the helical region of tolA, the extension region of tonB from E. coli, and a segment of protein D (pD) from the capsid of phage lambda (see, e.g., Yang et al., (2008) PLoS One 3(5):e2092; incorporated herein by reference). Other suitable linker amino acid sequences will be apparent to those of skill in the art (see, e.g., Argos (1990) J. Mol. Biol. 211(4):943-958; Crasto et al., (2000) Protein Eng. 13:309-312; George et al., (2002) Protein Eng. 15:871-879; Arai et al., (2001) Protein Eng. 14:529-532; and the Registry of Standard Biological Parts (partsregistry.org / Protein_domains / Linker). A polypeptide may comprise an N-terminal linker. An N-terminal linker may comprise an amino acid sequence at the N-terminus of the displayed polypeptide.The polypeptide may comprise a C-terminal spacer. The C-terminal spacer may comprise additional amino acids at the C-terminus of the polypeptide.

[0060]

[0075] Multiple polypeptides may be displayed simultaneously or on the same given substrate (e.g., solid surface such as a sequencing chip). For example, the method can be used to display a collection of proteins or peptides encoded by a genomic library for an organism or a cDNA library produced from RNA from an organism, or a selected subset of proteins or peptides of interest or engineered proteins or peptides expressed by an organism. The DNA library used for display may be wholly or partially synthetic and may contain sequences optimized for the expression of a particular set of polypeptides. Multiple DNA templates may be free in solution or fixed to a solid support. Polypeptide libraries and approaches for constructing polypeptide libraries are described elsewhere herein, and any number of polypeptides from such libraries may be displayed simultaneously or on the same surface.

[0061]

[0076] In some embodiments, the polynucleotides are immobilized on a solid support. The solid support may include, for example, glass, quartz, silica, metal, ceramic or plastic. Exemplary solid supports include slides, beads, plates, gels, membranes, or the inner surface of a flow cell or microchannel. Each DNA template may be located at a known, pre-determined position on the solid support, so that the identity of each protein produced from the DNA template may be determined from its position on the solid support. Alternatively, the DNA template may be randomly bound to the support, where the identity of the protein produced from each DNA template may be determined by sequencing the associated DNA template or characterizing the protein itself. Methods of immobilizing or coupling polynucleotides to beads and presenting polypeptides may be used, such as those described in WO2022026458A1, which is incorporated herein by reference.

[0062]

[0077] The nucleic acid may be covalently linked to a solid surface such as a polypeptide or a bead. Additionally, the polypeptide may be linked to a bead, for example, via direct conjugation to the bead or via conjugation to a nucleic acid attached to the bead. In some embodiments, the conjugation of the polypeptide to the nucleic acid molecule is catalyzed by a ligase. In some embodiments, the polypeptide is conjugated to the nucleic acid molecule by ligation of expressed proteins or by protein trans-splicing. In some embodiments, the polypeptide is conjugated to the nucleic acid molecule by the formation of a leucine zipper. In some embodiments, the bead or the nucleic acid molecule is conjugated to a capture component, and the polypeptide comprises a linking tag, where the capture component and the linking tag are conjugated, thereby conjugating the bead to the polypeptide or conjugating the nucleic acid molecule to the polypeptide. The ligase can be a sortase, butelase, trypsiligase, peptiligase, formylglycine generating enzyme, transglutaminase, tubulin tyrosine ligase, phosphopantetheinyl transferase, Spy Ligase, or Snoop Ligase.

[0063]

[0078] Nucleic acids can be coupled to solid supports by physical or chemical means using any method known in the art. Substrates can be added to the surface of solid supports to promote attachment of DNA templates. Methods for making DNA arrays are well known and include various photochemical-based methods, laser writing, electrospray deposition, inkjet and microjet deposition or spotting techniques, photolithographic oligonucleotide synthesis processes, and contact printing techniques including contact pin printing and microstamping. The combination of suitable robotics, microengineering-based systems, and microscopy techniques makes the ordered deposition of up to millions of nucleic acids per cm2 on solid supports technically feasible. See, e.g., Rehman et al. (1999) Nucleic Acids Research 27:649-655; Heller et al. (2002) Annu. Rev. Biomed. Eng. 4:129-153; Dufva (2009) Methods Mol. Biol. 529:1-22; Sethi et al. (2008) Bioconjug Chem. 19(11):2136-2143; Adessi et al. (2000) Nucleic Acids Res. 28(20):E87; Okamoto et al. (2000) Nat. Biotechnol. 18(4):438-441; Barbulovic-Nad et al. (2006) Crit. Rev. Biotechnol. 26(4):237-259; which are incorporated herein by reference.

[0064]

[0079] In one embodiment, the acrylamide modified nucleic acid is immobilized on a solid support (e.g., silanized glass or plastic) that contains exposed acryl groups. The acrylamide groups can be added to the nucleic acid during oligonucleotide synthesis using acrylamide phosphoramidites. The acrylamide modification is copolymerized with acrylamide monomers to allow the formation of stable polyacrylamide copolymers containing the immobilized nucleic acid. A layer containing immobilized DNA can be created on the support by polymerizing an acrylamide matrix on the surface of the support and adding the acrylamide modified nucleic acid. Polymerization is catalyzed using standard chemical or photochemical methods. See, for example, Rehman et al. (1999) Nucleic Acids Research 27:649-655; which is incorporated herein by reference in its entirety.

[0065]

[0080] Polynucleotides can be immobilized on solid support by hybridization to complementary capture oligonucleotides attached to the surface of solid support.Capture oligonucleotides can have unique sequences that are complementary to a single DNA template in a mixture of DNA templates, so as to allow selective capture of a specific DNA template.Additionally or alternatively, universal capture oligonucleotides can be used that bind to complementary adapter sequences added to DNA templates, so that a single type of capture oligonucleotide can be used to capture multiple DNA templates on solid support.DNA templates can be randomly arranged or ordered in an array on solid support, where each DNA template occupies a different position on solid support.

[0066]

[0081] The encoded polypeptide can be expressed and conjugated to beads (e.g., via conjugation to a nucleic acid conjugated to the beads), for example, by starting with nucleic acid-coated beads (e.g., DNA-coated beads) prepared using a method for presenting polynucleotides on beads. Conjugation of the polypeptide to the beads (e.g., directly or via attachment to the nucleic acid) can be performed in a microemulsion step. For example, DNA-coated beads are emulsified in a microemulsion with a mixture containing reagents for cell-free in vitro transcription and translation (IVTT) methods, resulting in transcription and translation of the DNA on the beads and production of the encoded polypeptide and / or protein. In some embodiments, the microemulsion contains reagents for IVTT and a catalytic enzyme or solution-phase DNA encoding a catalytic enzyme, which catalyzes the attachment of the polypeptide to the capture component on the nucleic acid. The components of the mixture can be adjusted as described herein to ensure an average of one DNA-coated bead and sufficient IVTT reagents.

[0067]

[0082] In some embodiments, the nucleic acid in each droplet is directly amplified on the surface of the bead through extension of the immobilized DNA oligo. In some embodiments, the nucleic acid can be amplified separately in a droplet that does not contain beads and then fused with another droplet that contains beads in a microfluidic channel. In some embodiments, in the generation of emulsion droplets, the nucleic acid in each droplet is amplified via polymerase chain reaction to create a clonal population of each nucleic acid variant. The physical fixation of the amplified nucleic acid in each microemulsion droplet can be achieved, for example, through ligation or extension of the immobilized DNA oligo to generate a nucleic acid-coated bead (e.g., a DNA-coated bead).

[0068]

[0083] In one embodiment, the method further comprises amplifying or extending at least one DNA template. Amplification or extension can be carried out using any known method, such as polymerase chain reaction (PCR) or other nucleic acid amplification process (e.g., ligase chain reaction (LGR), nucleic acid sequence-based amplification (NASBA), transcription-mediated amplification (TMA), Q-beta amplification, strand displacement amplification or target-mediated amplification). See, e.g., PCR Protocols, Vol. 226 (Methods in Molecular Biology, J. Bartlett and D. Stirling, eds., Humana Press; 2nd ed., 2003; Wiedmann et al., (1994) PCR Methods Appl. 3(4):551-64; Deiman et al., (2002) Mol. Biotechnol. 20(2):163-179; Guatelli et al., Proc. Natl. Acad. Sci. USA (1990) 87:1874-1878 and J. Compton, Nature (1991) 350:91-92 (1991); Hill (2001) Expert Rev. Mol. Diagn. 1:445-455; WO 89 / 1050; WO 88 / 10315; EPO Publication No. 2006-2011, all of which are incorporated by reference herein in their entireties. No. 408,295; EPO Application No. 8811394-8.9; WO 91 / 02818; U.S. Patent Nos. 5,399,491, 6,686,156 and 5,556,771; Walker et al., Clin. Chem. (1996) 42:9-13 and EPA 684,31.In particular, clonal amplification methods, such as, but not limited to, bridge amplification, emulsion PCR (ePCR), or rolling circle amplification, can be used to cluster amplified nucleic acids in separate regions (see, e.g., U.S. Pat. Nos. 7,790,418; 5,641,658; 7,264,934; 7,323,305; 8,293,502; 6,287,824; and International Application No. WO1998 / 044151 A1; Lizardi et al., (1998) Nature Genetics 19:225-232; Leamon et al., (2003) Electrophoresis 24:3769-3777; Dressman et al., (2003) Proc. Natl. Acad. Sci. USA, all of which are incorporated by reference herein). 100:8817-8822; Tawfik et al., (1998) Nature Biotechnol. 16:652-656; Nakano et al., (2003) J. Biotechnol. 102:117-124; see. For this purpose, the DNA template may contain at the 5' and 3' ends an adapter sequence suitable for high-throughput amplification (e.g., an adapter containing a sequence complementary to a universal amplification primer or a bridging PCR amplification primer). For example, a bridging PCR primer attached to a solid support may be used to capture the DNA template containing an adapter sequence complementary to the bridging PCR primer. The DNA templates may then be amplified, with the amplification products of each DNA template clustering in separate regions on the solid support. In one embodiment, the DNA templates are attached to a solid support, amplified, and sequenced prior to presenting the ribosomal translation products for functional screening.

[0069]

[0084] In various embodiments, microemulsion droplets may be used. Microemulsion droplets may be used to convert a bulk solution into multiple droplets. The droplets may contain reagents for reactions that may occur in the droplets and are separated from other microemulsion droplets or the bulk solution, allowing a microenvironment for the reactions to occur. For example, conjugation, transcription, translation or amplification reactions may occur in microemulsion droplets. Methods for making microemulsion droplets for chemical and biochemical reaction purposes are well known to those skilled in the art. Generally, microemulsion droplets contain an aqueous phase suspended in an oil phase (e.g., a water-in-oil emulsion). In one embodiment, the oil phase is composed of 95% mineral oil, 4.5% Span-80, 0.45% Tween-80 and 0.05% Triton X-100. In some embodiments, the microemulsion is formed via direct mixing and / or vortexing of the aqueous and oil phases. In some embodiments, the microemulsion is formed via a piezoelectric pump that pushes the aqueous phase into a microfluidic channel that contains the oil phase. In some embodiments, the microemulsion is formed via mechanical mixing of the aqueous and oil phases using a dispersing device or homogenizer. In one embodiment, each emulsion emulsion droplet contains, on average, a single primer-coated bead, one template DNA molecule, and multiple PCR primer molecules. Temperature cycling can be used to generate amplified clonal DNA from the template on the bead.

[0070] Identification of polypeptide library features

[0085] The polypeptide library can be generated and displayed as described elsewhere in this disclosure. The displayed polypeptide can be linked or associated with the corresponding polynucleotide by which the polypeptide is encoded. Sequencing reactions can be performed on the polynucleotides as described elsewhere herein. Any sequencing method can be used, including but not limited to Maxam-Gilbert sequencing, Sanger sequencing (i.e., chain termination), sequencing by synthesis (SBS), sequencing by ligation, pyrosequencing, ion torrent sequencing, nanopore sequencing, and single molecule real-time sequencing. In one embodiment, the multiple DNA templates are sequenced by a high-throughput DNA sequencing method. See, e.g., Pettersson et al. (2009) Genomics 93(2):105-111; Maxam & Gilbert (1977) Proc. Natl. Acad. Sci. USA 74(2):560-564; Sanger et al. (1977) Proc. Natl. Acad. Sci. USA 74(12):5463-5467; Ronaghi et al. (1996) Analytical Biochemistry 242(1):84-89; Brenner et al. (2000) Nature Biotechnology 18(6):630-634; Schuster (2008) Nat. Methods 5(1):16-18; Margulies et al. (2005) Nature 437:376-380; Shendure et al. (2005) Science 309:1728-1732; Thompson et al., (2012) Electrophoresis 33(23):3429-3436; Merriman et al., (2012) Electrophoresis. 33(23):3397-3417; and Pareek et al., (2011) Journal of applied genetics 52(4):413-435).

[0071]

[0086] The sequencing reaction can generate sequencing data for a polynucleotide. In some embodiments, the polynucleotide is attached to an array or solid support, or otherwise spatially distinctly separated. By sequencing the polynucleotide, a particular polynucleotide on the array or solid support can be identified as having a particular sequence. In this way, a particular point on the array can be identified as having a particular or known sequence. The polypeptide display technology described in this disclosure allows a polypeptide to be attached, linked, or otherwise associated with a polynucleotide that encodes the polypeptide. Since the sequencing reaction can identify a polynucleotide as having a particular sequence, the amino acid sequence of the corresponding polypeptide can be determined.

[0072]

[0087] Analysis of the polypeptides can be performed. Massively parallel high-throughput protein screening can be performed on the polypeptide library. For example, multiple assays can be performed where a library of polynucleotides can be immobilized on a solid support, such as beads in defined locations on a carrier (e.g., a capillary), or on the inner surface of a microchannel or flow chamber, or on the surface of a microscope slide. The surface can be a planar surface or a coated surface. Additionally, the surface can include a plurality of microfeatures arranged in spatially separated regions to create a structure on the surface, where the textured surface provides an increased surface area compared to an unstructured surface.

[0073]

[0088] The array may contain a plurality of presented ribosomal translation products, such as antigens, antibodies, enzymes, substrates, receptors, or regulatory molecules, or libraries thereof. Such arrays may be used, for example, in high-throughput genetic or pharmacological screening, epitope mapping, protein engineering, or proteomic profiling. For high-throughput screening, the array is preferably contained within a flow cell or microfluidic device. Potentially, tens of millions to billions of proteins, peptides, or small molecules translated by ribosomes can be screened quantitatively at the same time. Functional screening can be performed in a continuous flow or stop-flow system, where proteins are presented on immobilized polynucleotides as described herein, and various reagents and buffers are injected into the system at one end and exit the system at the other end. The reagents and buffers may be flowing continuously or may be held in place for a period of time to allow ligand binding or enzymatic reactions to proceed. Additionally, the ligand or substrate may be labeled to facilitate detection and quantitative analysis of binding interactions or enzymatic reactions.

[0074]

[0089] In some embodiments, protein characterization assay is carried out in high-throughput sequencing device.Ribosomal translation products (e.g., proteins or peptides, their biologically active fragments or other ribosome-translated molecules) can be presented on polynucleotide in sequencing device using the methods described herein, and then be directly functionally characterized simultaneously in sequencing flow cell.This allows high-throughput sequencing to be easily combined with protein screening, which can bring significant added value to high-throughput sequencing device.

[0075]

[0090] In some embodiments, sequencing the nucleic acid molecules and assaying one or more functions or properties of each polypeptide are performed (e.g., sequentially, in any order) on the same machine, device, or apparatus. In some embodiments, multiple assays are performed to determine more than one function or property of each polypeptide, or multiple assays are performed to determine a single function or property of each polypeptide under various conditions. Multiple assays can be performed simultaneously or sequentially on the same machine, device, or apparatus. For example, a single machine, device, or apparatus can be used to sequence the nucleic acid molecules conjugated to each bead to identify the polypeptide conjugated to the bead; and to perform one or more assays to characterize each polypeptide (e.g., binding affinity, binding specificity, enzyme activity, stability under various experimental conditions, including temperature and / or pH). In some embodiments, sequencing and one or more assays produce a fluorescent signature that is measured by a single machine, device, or apparatus.

[0076]

[0091] Polypeptide characterization may include generating a detectable signal based on the presence of a reaction or event. For example, a detectable signal may be generated upon binding of the polypeptide to an antigen. A detectable signal may be generated by a detectable label. The detectable label may be attached or coupled to the antigen (or target molecule), or may be attached to another reagent capable of detecting the antigen (or target molecule). For example, the antigen may be coupled to an enzyme capable of generating a signal. The polypeptide library may be contacted with the antigen or target molecule, and the polypeptide may bind to the antigen. After the excess antigen is removed, an enzyme substrate is added, and the enzyme allows a detectable signal to be generated. When the enzyme attached to the polypeptide-bound antigen is able to react with the enzyme substrate, a signal is generated, and the presence of a detectable signal may thereby indicate that the polypeptide is bound to the antigen. Similarly, the antigen may be coupled to a fluorophore, and a signal may be generated by excitation of the fluorophore. In another similar example, an antibody that binds to the antigen or target molecule may include an enzyme or fluorophore. The presented polypeptide library may be contacted with the antigen or target molecule. After removal of excess antigen, an antibody coupled to an enzyme or fluorophore is added and the excess is removed. Since the signal is generated by the antibody bound to the antigen bound to the polypeptide, the polypeptide bound to the antigen could be identified based on the generation of a signal.

[0077]

[0092] Detectable labels can be any composition detectable by spectroscopic, photochemical, biochemical, immunochemical, electrical, optical or chemical means. Detectable labels can include fluorescent dyes (e.g., phycoerythrin, YPet, fluorescein, tag RFP, Texas Red, rhodamine, green fluorescent protein, etc., see, e.g., Molecular Probes, Eugene, Oreg., USA), quantum dots, radiolabels (e.g., 3H, 125I, 35S, 14C or 32P), enzymes (e.g., horseradish peroxidase, alkaline phosphatase and others commonly used in ELISA), and colorimetric labels, such as colloidal gold (e.g., gold particles with a diameter size range of 40-80 nm that scatter green light with high efficiency) or colored glass or plastic (e.g., polystyrene, polypropylene, latex, etc.) beads. Patents which teach the use of such labels include, but are not limited to, U.S. Pat. Nos. 3,817,837; 3,850,752; 3,939,350; 3,996,345; 4,277,437; 4,275,149; 4,366,241; 7,416,854; 8,366,241; ,114,681; No. 7,229,769; No. 6,846,645; No. 7,232,659; No. 6,872,578; No. 7,897,257; No. 6,730,521; No. 5,972,721; No. 7,498,177; No. 7,235,361; and No. 6,306,610.

[0078]

[0093] Using the presence of a detectable signal, a multiplexed quantitative protein assay can be performed. A multiplexed quantitative protein assay can allow for the calculation, generation or identification of a quantitative characteristic of a polypeptide. The quantitative characteristic can be a kinetic or thermodynamic parameter associated with the polypeptide. For example, the quantitative characteristic can be a melting (or denaturation) temperature (T m ) or the midpoint concentration (C m) or a measure of polypeptide stability such as an equilibrium constant. The quantitative feature can be non-specific binding capacity, aggregation capacity, hydrophobicity, maturation time or protein expression level. The quantitative feature can be a rate constant or a kinetic parameter. The quantitative feature can be related to intra- or intermolecular interactions or reactions. For example, the quantitative feature can be enzyme reaction rate, enzyme activity, fractional activity or any relevant thermodynamic constant. In some cases, multiplexed quantitative protein binding assays can be performed. The quantitative feature can be binding affinity, association (K a ) or dissociation constant (K d ), kinetic constants of binding (e.g., k on Or k off The binding assay may be performed by observing a detectable signal generated in the presence of a binding event of a polypeptide of the library to a target molecule, and the intensity of the detectable signal may be used to quantify binding. By adding a series of known concentrations of the target molecule, binding of the target molecule to the polypeptide library is allowed, intensity data for each polypeptide is obtained, and a binding curve may be generated for every polypeptide in the polypeptide library. This concentration-dependent binding curve may be fitted and the binding affinity for each polypeptide in the library may be calculated. For polypeptides presented on an array, each polypeptide may be observed as a point on the array, and the intensity of each point on the array at a given concentration of the target molecule may be observed. In this way, multiple polypeptides may be analyzed in the same assay, and quantitative characteristics may be obtained for multiple polypeptides in the assay.

[0079]

[0094] Binding data or other data from multiplexed quantitative protein assays can be used to characterize the polypeptides in a polypeptide library. A polypeptide library includes variants of a reference or wild-type sequence, and these assays can characterize the variants as having a neutral effect, a positive effect, or a negative effect on the characteristics of the polypeptide. For example, to characterize binding affinity, a polypeptide variant can be characterized as having an increased binding affinity, a decreased binding affinity, or a slightly altered binding affinity to an antigen. For example, a neutral variant can have a dissociation constant that is greater than 0.25 times and less than 2 times the dissociation constant of the reference or starting polypeptide. A positive variant can have a dissociation constant that is 0.25 times or less than the dissociation constant of the reference or starting polypeptide. A negative variant can have a dissociation constant that is 2 times or more than the dissociation constant of the reference or starting polypeptide. By using this data of quantitative characteristics, novel polypeptide libraries can be constructed, such as polypeptides that include a combination of multiple variants with increased binding affinity. In addition, using quantitative measurements, the strength or magnitude of a feature may be used to guide future library construction, data that may otherwise be lost in a typical enrichment or selection assay. Additionally, observations of variants with negative or neutral effects may be positively observed, as opposed to potentially being lost in a typical selection or enrichment assay that only enriches for variants with positive effects.

[0080]

[0095] The multiple quantitative protein assays described herein allow for the observation of a large number of proteins in a given assay. 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9Or more than one characteristic of the polypeptide may be observed in a single assay, or simultaneously (or substantially simultaneously). The assay may be performed in a short period of time. The assay may be performed for up to or less than 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 9 hours, 10 hours, 11 hours, 12 hours, 13 hours, 14 hours, 15 hours, 16 hours, 17 hours, 18 hours, 19 hours, 20 hours, 21 hours, 22 hours, 23 hours, 24 hours, 25 hours, 26 hours, 27 hours, 28 hours, 29 hours, 30 hours, 31 hours, 32 hours, 33 hours, 34 hours, 35 hours, 36 hours, 37 hours, 38 hours, 39 hours, 40 hours, 41 hours, 42 hours, 43 hours, 44 hours, 45 hours, 46 hours, 47 hours, 48 ​​hours, 49 hours, 50 hours, 55 hours, 60 hours, 65 hours, 70 hours.

[0081]

[0096] Multiple quantitative protein binding assays can be performed on the polypeptide library using different antigens or under different conditions. For example, a first binding assay can be performed using a first antigen to identify polypeptides that bind to the first antigen. A second binding assay can be performed using a second antigen to identify polypeptides that bind to the second antigen. Using data generated from the two binding assays, polypeptides that bind to both the first and second antigens can be identified. The polypeptide library construction can be repeated as described elsewhere, and synergistic combinations of variants can be identified as binding to both the first and second antigens. Additionally, binding assays can be performed on a third antigen, a fourth antigen, or an nth antigen, and polypeptides that bind (or do not bind) to a particular set or subset of antigens. Based on the data generated and the iterative library design, polynucleotides that are specific for an antigen(s) and do not bind (or do not bind very well) to other antigens can be generated. For example, polypeptides that bind to the first antigen and the second antigen, but do not bind to the third antigen can be generated. In another example, a polypeptide can be generated that binds to a first and second antigen and also binds to a third antigen. Figure 8 shows an example Venn diagram relating to the different types of polypeptides that can be made in relation to three antigens. A polypeptide can fall anywhere within this diagram, such that it binds or does not bind (or binds very little or only slightly) to the respective antigen.

[0082]

[0097] Identification of a polypeptide with a particular characteristic can be used to generate additional protein constructs or polypeptide conjugates. The polypeptides in the polypeptide library can represent functional domains or fragments of a full-length protein. Based on the sequence of the polypeptide (or corresponding polynucleotide), a polypeptide can be expressed that has a particular characteristic and includes a polypeptide that includes the polypeptide sequence of another protein, domain, or fragment. For example, a polypeptide-chimeric antigen receptor fusion can be generated. A polypeptide drug conjugate (e.g., an antibody drug conjugate) can be generated. For example, the polypeptides in the library can be heavy chain fragments, light chain fragments, nanobodies, or scFvs. Once a fragment is identified as having a particular characteristic, a new full-length polypeptide that includes the sequence of the fragment can be generated. For example, a full-length antibody can be generated by expressing a polynucleotide that includes the coding region of the fragment together with the coding sequence of the Fc region. For example, a CDR sequence can be identified based on the method of the present disclosure, and a full-length IgG antibody can be generated based on the CDR sequence and the sequence of the IgG backbone. For example, a bivalent nanobody can be generated based on the sequence of the polypeptide analyzed by the method of the present disclosure. In this way, it may be possible to identify and generate full-length antibodies (or other functional proteins) based on data generated from libraries that do not use full-length proteins. This may be advantageous since it allows the construction of a protein of interest to be performed modularly, allowing each domain of the protein to be characterized individually. For example, a library may be generated corresponding to a first CDR of an antibody, and a characterization method may be performed on the library. A second library may be generated corresponding to a second CDR of an antibody, and a characterization method may be performed on the second library. These libraries may be analyzed on the same sequencing chip or substrate, or at the same time or at different times. CDR libraries may be subjected to different antigens or the same antigen, so that multispecific, multi-epitope, or highly specific antibodies may be generated. Additionally, smaller fragments may be more easily characterized or expressed on a given polypeptide display array.

[0083]

[0098] Identification of polypeptides with particular characteristics can be used to generate additional polypeptide libraries. Polypeptides in a polypeptide library can exhibit functional domains with various characteristics. For example, polypeptides in a polypeptide library can have different binding affinities to an antigen. Based at least on the characteristics of a given polypeptide, additional libraries can be generated to optimize or improve the characteristics. For example, a polypeptide in a library can exhibit a medium or low affinity to an antigen. Subsequent libraries can use a polypeptide with a medium affinity and generate multiple polypeptides that include point mutants of the polypeptide or fusions that include the polypeptide. Since the original polypeptide exhibits a medium to low affinity, point mutants or fusions with improved affinity can be more easily identified compared to using the original polypeptide that already has a high affinity to the antigen. Data obtained on constructs with improved affinity (or other characteristics) can be used to generate further improved constructs. For example, a fusion protein that includes a first domain with medium binding and a second domain with medium binding can exhibit an avidity effect. The first domain can be "swapped" for a domain with a high affinity to generate a polypeptide construct with increased binding, avidity, or a combination of both. The library can also include fusion polypeptides or constructs with a domain that does not bind to an antigen or has a low affinity to bind to an antigen. For example, a fusion polypeptide can have a first domain that binds and a second domain that does not bind. The presence of a domain or monomer that does not bind allows the characteristics of the polypeptide to be compared to another polypeptide with more similar physical characteristics. In the example of a polypeptide with a first domain that binds and a second domain that does not bind, this can be directly compared to a polypeptide with the same first domain but with a second domain that binds. These polypeptides can be of more similar size, length, shape compared to a polypeptide with only one domain. In this way, the comparison can lead to more accurate results.A domain or polypeptide region that does not bind (or has minimal or no affinity for an antigen) may have the same length, size, shape, net charge as a domain that binds or has affinity for an antigen. A domain or polypeptide region that does not bind (or has minimal or no affinity for an antigen) may have substantially the same length, size, shape, net charge as a domain that binds or has affinity for an antigen. A domain or polypeptide region that does not bind (or has minimal or no affinity for an antigen) may have a length, size, shape, net charge that differs by no more than 10% from a domain that binds or has affinity for an antigen.

[0084]

[0099] Polypeptides generated from the methods of the present disclosure may use quantitative features analyzed in various libraries to generate optimized polypeptides. For example, a first library may generate data related to binding affinity for multiple point mutations of a first scaffold. A second library may generate data related to binding affinity for multiple different scaffolds including the first scaffold. A third library may include data related to binding affinity from a combination of any two scaffolds of the second library. A polypeptide may be generated that includes two scaffolds that include point mutations analyzed in the first library. In this way, optimized polypeptides may be generated that utilize information collected at a first level of detail (e.g., point mutations for a given scaffold) and information collected at a second level of detail (e.g., bivalent or bi-epitope scaffolds) to generate polypeptides that are not necessarily present in their entirety in a given library.

[0085]

[0100] For example, the first library may include a plurality of single domains that bind to an antigen. The second library may include point mutations of one or more of the single domains in the plurality of single domains in the first library. The first library may allow for the identification of a first scaffold that binds to an antigen. The second library may generate variants of the first scaffold with different binding characteristics. Determining the binding characteristics (or other quantitative characteristics) may be used to generate a new library, or separate libraries may be assayed simultaneously without using data generated from a previously generated library. The generated second library may identify mutations that result in desirable or targeted binding characteristics. For example, the binding characteristics may be an improvement in binding. The third library may be generated by combining the single domains with fusion polypeptides that include pairs of single domains. The third library may include all possible combinations of single domain pairs. The third library may include all possible permutations of single domain pairs. The third library may include single domain pairs, where the single domains include reduced binding characteristics compared to a reference or wild type single domain. A third library may be used to identify bi-epitope binders, and the use of a single domain with reduced binding may allow bi-epitope binders to be more easily identified. The use of two strong binders in a construct may result in an increase in binding that is difficult to unravel or identify, since bi-epitope binders may significantly increase binding characteristics based on avidity effects. By using a weaker binder that still binds to the epitope, the avidity effect obtained in the bi-epitope construct may become more easily evident and assayable using a given binding assay. The information generated by each library may be combined to generate optimized polypeptides, where the optimized polypeptides were not necessarily analyzed in either library. For example, a library containing constructs containing two or more domains may be used to determine and identify domains or scaffolds that bind in tandem or bi-epitopes.Data obtained using libraries containing point mutations of the scaffold can identify mutations that result in high or highest binding affinity to the antigen. Mutations can then be introduced by substitution into the biepitope construct to generate biepitope (or multiepitope) constructs in which each domain has an optimized binding affinity or binding characteristics.

[0086]

[0101] The fragments analyzed using the methods of the present disclosure can be used to generate larger polypeptides, such as fusion proteins. Libraries can be generated to code and generate larger polypeptides. For example, libraries can be generated to code fusion proteins. Larger polypeptides can be generated without generating libraries. For example, data related to scFvs or CDRs can be generated using the methods and systems disclosed elsewhere herein, and full-length antibodies can be generated using this data without using libraries that code for full-length antibodies.

[0087]

[0102] The polypeptide may include a linker or spacer domain. The linker may link two domains to form a fusion protein. The linker may be a polypeptide linker. The linker or spacer domain may include at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 40, 45, 50, 60, 70, 80, 90, 100 or more amino acids. The linker or spacer domain may comprise 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 40, 45, 50, 60, 70, 80, 90, 100 or less amino acids. The spacer domain may be a polypeptide spacer domain. The spacer domain may be an N-terminal spacer domain. The spacer domain may be a C-terminal spacer domain. The spacer domain or linker may have a positive, negative or neutral charge. The spacer domain or linker may have a net positive, net negative or net neutral charge. The spacer domain or linker may be hydrophobic, hydrophilic or partially hydrophobic or hydrophilic. For example, a first VHH may be analyzed using the methods described and a library corresponding to the first VHH (e.g., a library of single point mutations). Once the analysis of the first VHH is performed, a certain VHH with a particular characteristic (such as binding to a target or epitope) may be used to create a second library that includes another combination of VHHs separated by a linker sequence. The other VHH may be analyzed by creating a library such that both VHHs are analyzed and selected independently, prior to the generation of a subsequent library that includes a construct that includes multiple VHHs.A library containing constructs containing two or more VHHs separated by a linker sequence(s) can then be subjected to the analysis described elsewhere herein. In this way, bi-epitope constructs can be generated, where each binding unit is analyzed individually or simultaneously to identify constructs with desired parameters or certain characteristics. Libraries can also be analyzed or generated independently, and can be assayed simultaneously or sequentially. For example, a library containing constructs of two or more VHHs can be generated and tested together with a library containing constructs of a single VHH, without using data from a single VHH library guide, or used to determine the polypeptides of a library containing constructs of two or more VHHs.

[0088]

[0103] The library may include the generation of polypeptides with different linker or spacer domains. The library includes polypeptides comprising a scaffold or domain and an N-terminal spacer, where the polypeptides have different N-terminal spacers. The N-terminal spacer may alter the presentation or other characteristics of the polypeptide, and a library of different N-terminal spacers allows for the determination of the optimal or preferred N-terminal spacer for a given polypeptide or scaffold. Similarly, libraries may be generated and assayed for N-terminal spacers, C-terminal spacers, linkers, or combinations thereof. The N-terminal spacers, C-terminal spacers, or linkers may include various lengths, charges, mobilities, steric bulk, hydrophobicity, or other characteristics that may affect the characteristics of the polypeptide. The library allows for the selection of appropriate spacers and linkers for the polypeptide construct. In the context of bi-epitope (or multi-epitope) binders, variation in the length of the linker may affect the binding characteristics. Since epitopes for an antigen may be separated by a certain distance, the spatial characteristics of the binder may be suitable for optimizing binding. For example, a linker that separates two binding domains that is too short may prevent the binding agent from simultaneously binding to both binding domains on the antigen, thereby affecting the overall binding capacity. Thus, a library that contains the same two scaffolds or binding domains with different linkers can be used to identify optimal or suitable linkers.

[0089]

[0104] In various aspects, data that can be used to generate a polypeptide can be generated or obtained. For example, data related to the binding characteristics of multiple polypeptides can be generated or obtained. This data can be used to guide the design of a library. For example, a first library of different scaffolds can be generated and data related to the binding characteristics of the scaffolds can be generated. Scaffolds that did not bind to the antigen can be omitted from future libraries. Scaffolds that bind to the antigen can be used as reference scaffolds or polypeptides to generate a library of point mutants of that scaffold. Data can be obtained from publicly available databases. For example, publicly available data on polypeptides that bind to antigens can be used to determine a reference polypeptide or scaffold. Multiple data sets can be used and compared. For example, data on a polypeptide that includes a single domain can be compared to data on a polypeptide that includes a fusion of the single domain. By comparing the data of the single domains that correspond to polypeptides that include the same single domain, improvements to binding based on the addition of another domain (e.g., a bi-epitope construct) can be determined.

[0090]

[0105] 15A-15C show an example schematic workflow that can be used to generate a library and that uses library-derived data to generate a polypeptide of interest. FIG. 15A shows a schematic workflow that allows for the generation of affinity-optimized variants. An initial library 1501 is generated that contains mutations of a polypeptide. The library can be a systematic mutation scan library in which single point mutations are made at every residue from a region of the polypeptide, substituting each of all 20 standard amino acids. Analysis of library 1501 results in information about the mutation landscape of the polypeptide, where the impact of individual mutations can be analyzed. A second library 1505 is generated using analysis of the data, with "targeting" based on the information discovered in library 1501. For example, library 1505 can contain mutations to multiple residues identified in library 1501 that can result in improved binding. Initial library 1501 can, for example, identify single point mutations that improve binding affinity. Library 1505 can contain a polypeptide that contains multiple single point mutations identified in library 1501. The initial library 1501 can, for example, identify residues suitable for mutation where, for example, some or all single point mutations result in a neutral or positive increase in binding. The library 1505 can include polypeptides with all combinations of mutations at the residues identified as potentially suitable for mutation. Screening of the library 1505 allows for the creation of a large data set of different polypeptides that differ by multiple mutations from the initial reference or wild type polypeptide. Data analysis 1515 is performed on this data set, allowing for the identification of affinity-optimized variants.

[0091]

[0106] FIG. 15B shows an example schematic of identifying tandem pairs that lead to increased binding activity. A first library 1520 of monomeric polypeptides capable of binding to an antigen is generated, and data for the different individual monomeric polypeptides is generated. A second library 1525 is also generated, which includes polypeptides made by creating fusion tandem polypeptides that include the polypeptide sequences of the two monomeric polypeptides. The second library 1525 can have all possible permutations of the two monomeric polypeptides. The libraries 1520 and 1525 can also include polypeptides with different N- and / or C-terminal spacers that can affect the binding and presentation of the polypeptides. Additionally, the second library 1525 can also include various linkers between the two monomeric polypeptides. For example, the second library 1525 can include a polypeptide that includes two monomeric polypeptides with a linker, and a second polypeptide that includes the same two monomeric polypeptides with a different linker. Additionally, the library 1525 can include polypeptides with one monomeric polypeptide capable of binding to an antigen and another monomer that does not bind to the antigen. This can generate polypeptides that act as a baseline for comparison against other tandem polypeptides, since it creates "pseudomonomers" of similar size but with only one binding domain. Data analysis 1530 is performed by comparing data from the monomer polypeptide library 1520 with data from the tandem library 1525 (and pseudomonomers) to find pairs in the tandem library that result in an increase in binding affinity compared to its constituent individual monomers (and pseudomonomers).

[0092]

[0107] FIG. 15C shows a schematic diagram of an example workflow combining analysis and libraries as described and illustrated in FIG. 15A and 15B. A set of libraries and data 1540 is generated for multiple reference and wild-type molecules. For each of these polypeptides, an initial systematic mutation scan library, e.g., library 1501, is generated. Analysis of library 1540 creates information about the mutation landscape of the polypeptide, where the impact of individual mutations can be analyzed. The information about the mutation landscape can then be used to generate three different libraries. Similar to what was described for library 1505, a targeting library is generated for each of the reference or wild-type polypeptides. Another set of libraries 1545 with "targeting" based on the information found in library 1540 is generated using analysis of the data. For example, library 1545 can include mutations to multiple residues identified in library 1540 that can result in improved binding. The set of libraries 1540 can, for example, identify single point mutations that improve binding affinity. Library 1545 may include polypeptides that include multiple single point mutations identified in library 1540. Library 1540 may, for example, identify residues suitable for mutation where some or all single point mutations result in a neutral or positive increase in binding. Library 1545 may include polypeptides with all combinations of mutations at residues identified as potentially suitable for mutation. Screening library 1545 allows for the generation of a large data set of different polypeptides that differ from an initial reference or wild type polypeptide by multiple mutations. Data analysis 1550 is performed on this data set, allowing for the identification of affinity-optimized variants. A second library 1560 is generated that includes multiple monomers that exhibited moderate to low affinity as determined by the set of library 1540. A third library 1565 is also generated that includes polypeptides created by creating a fusion tandem polypeptide that includes the polypeptide sequences of two monomer polypeptides.The second library 1565 may have all possible permutations of the two monomeric polypeptides. The libraries 1560 and 1565 may also include polypeptides with different N-terminal spacers and / or C-terminal spacers, which may affect the binding and presentation of the polypeptides. Additionally, the second library 1565 may include different linkers between the two monomeric polypeptides. For example, the second library 1565 may include a polypeptide that includes two monomeric polypeptides with a linker, and a second polypeptide that includes the same two monomeric polypeptides with a different linker. Additionally, the library 1565 may include a polypeptide with one monomeric polypeptide that can bind to an antigen and another monomer that does not bind to the antigen. This may generate a polypeptide that acts as a baseline for comparison against other tandem polypeptides, since it creates a "pseudomonomer" of similar size but with only one binding domain. Data analysis 1570 is performed by comparing data from the monomer polypeptide library 1560 with data from the tandem library 1565 (and pseudomonomers) to find pairs in the tandem library that result in an increase in binding affinity compared to its constituent individual monomers (and pseudomonomers). Data analysis 1580 is then performed to identify high affinity tandem binders based on data analysis 1550 and data analysis 1570. Data analysis 1570 identifies monomers that bind in tandem, however each monomer generated may not have high affinity by itself. Data analysis 1550 determines mutations that result in increased affinity in a given monomer construct. By combining the data and adding mutations to each of the monomers of the tandem pairs found in data analysis 1570, tandem binders in which each monomer has high affinity can be generated.

[0093]

[0108] Fiducial markers can be used because multiple protein assays can be performed and imaged on a protein array. Fiducial markers allow for alignment of multiple images from a given array. Because a multiplexed protein assay involves multiple polypeptides on a given array, this can be advantageous to prevent one polypeptide from being confused with another. By imaging one or more fiducial markers with the polypeptide, a location on the array can be identified as the location of the fiducial marker. The signal for the polypeptide on the array is referenced to one or more fiducial markers, so that the location of each polypeptide can be precisely mapped. For binding assays, multiple images of the polypeptide array can be generated. These images can be aligned based on the location of one or more fiducial markers.

[0094]

[0109] A reference marker can be generated by capturing a reference polynucleotide on an array. A polynucleotide complementary to the reference polynucleotide can then be added, where the polynucleotide complementary to the reference polynucleotide comprises a detectable label. This detectable label can act as the reference marker.

[0095]

[0110] In various embodiments, the polypeptide library is bound to the antigen, and the binding data is derived from the polypeptide library. The antigen can be a small molecule, a protein or polypeptide, a receptor, a hormone, or any molecule. The antigen can be from an animal, a plant, a fungus, a microorganism, a virus, or other biological organism. The antigen can be an inorganic or organic compound. The antigen can be from or generated from a pathogen. For example, the antigen can be from or generated from SARS-CoV-2. The antigen can be the SARS-CoV-2 receptor binding domain (RBD).

[0096]

[0111] Polypeptides produced using the methods, compositions and systems described in this disclosure can be used to produce antibodies or antibody fragments. Antibodies and antibody fragments can be used as therapeutic or diagnostic agents, and antibodies with high affinity and / or high specificity can be very useful. The methods, compositions and systems provided elsewhere herein are capable of producing antibodies with high affinity and / or high specificity. Additionally, by multiplexing the capabilities of the methods described, antibodies of specific characteristics can be assayed and designed in a highly efficient manner.

[0097] Computer Control System

[0112] The present disclosure provides a computer control system programmed to carry out the method of the present disclosure. Figure 16 shows a computer system 1601 programmed or otherwise configured to carry out a part of the method, such as image processing corresponding to a polypeptide library, or calculation of binding affinity. The computer system 1601 can control various aspects of the method of the present disclosure, such as image reception, image processing for intensity, output of binding curves, etc. The computer system 1601 can be a user's electronic device or a computer system located remotely to the electronic device. The electronic device can be a mobile electronic device.

[0098]

[0113] The computer system 1601 includes a central processing unit (CPU, also referred to herein as "processor" and "computer processor") 1605, which may be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 1601 also includes memory or memory locations 1610 (e.g., random access memory, read-only memory, flash memory), electronic storage 1615 (e.g., hard disk), communication interface 1620 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 1625, such as cache, other memory, data storage devices, and / or electronic display adapters. The memory 1610, storage 1615, interface 1620, and peripheral devices 1625 communicate with the CPU 1605 through a communication bus (solid line), such as a motherboard. The storage 1615 may be a data storage device (or data storage location) for storing data. The computer system 1601 may be operatively coupled to a computer network ("network") 1630 with the aid of the communication interface 1620. The network 1630 may be the Internet, an Internet and / or an extranet, or an intranet and / or an extranet in communication with the Internet. In some cases, the network 1630 is a telecommunications and / or data network. The network 1630 may include one or more computer servers that enable distributed computing, such as cloud computing. The network 1630 may implement a peer-to-peer network, in some cases with the help of the computer system 1601, that allows devices coupled to the computer system 1601 to function as clients or servers.

[0099]

[0114] The CPU 1605 can execute a series of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 1610. The instructions may be directed to the CPU 1605, which can then be programmed or otherwise configured to execute the methods of the present disclosure. Examples of operations performed by the CPU 1605 include fetch, decode, execute, and writeback.

[0100]

[0115] The CPU 1605 may be part of a circuit, such as an integrated circuit. One or more other components of the system 1601 may be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).

[0101]

[0116] The storage device 1615 can store files such as drivers, libraries, and saved programs. The storage device 1615 can store user data, such as user settings and user programs. In some cases, the computer system 1601 may include one or more additional data storage devices external to the computer system 1601, such as located on a remote server in communication with the computer system 1601 over an intranet or the Internet.

[0102]

[0117] The computer system 1601 can communicate with one or more remote computer systems through the network 1630. For example, the computer system 1601 can communicate with a remote computer system of a user. Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., Apple® iPad®, Samsung® Galaxy Tab), a telephone, a smartphone (e.g., Apple® iPhone®, Android-enabled devices, Blackberry®), or a personal digital assistant. A user can connect to the computer system 1601 through the network 1630.

[0103]

[0118] The methods described herein may be performed by machine (e.g., a computer processor) executable code stored in an electronic storage location, such as, for example, memory 1610 or electronic storage device 1615, of the computer system 1601. The machine executable code, or machine readable code, may be provided in the form of software. In use, the code may be executed by the processor 1605. In some cases, the code may be read from storage device 1615 and stored in memory 1610 for immediate access by the processor 1605. In some cases, the electronic storage device 1615 may be omitted and the machine executable instructions are stored in memory 1610.

[0104]

[0119] The code may be pre-compiled and configured for use with a machine having a processor adapted to execute the code, or may be compiled at run-time. The code may be provided in a programming language that may be selected to allow the code to execute in a pre-compiled or compiled fashion.

[0105]

[0120] Aspects of the systems and methods provided herein, such as the computer system 1601, may be embodied in programming. Various aspects of the technology are typically thought of as "products" or "articles of manufacture" in the form of machine (or processor) executable code and / or associated data executed or embodied in some type of machine-readable medium. The machine executable code may be stored in electronic storage, such as memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. A "storage" type medium may include any or all tangible memory of a computer, processor, etc., or their associated modules, such as various semiconductor memories, tape drives, disk drives, etc., that can provide persistent storage at any time for software programming. At times, all or part of the software may be communicated over the Internet or various other telecommunications networks. Such communication may, for example, enable loading of the software from one computer or processor to another, such as from a management server or host computer to the computer platform of an application server. Thus, other types of media that can carry software elements include light waves, radio waves, and electromagnetic waves used across physical interfaces between local devices through wired and optical landline networks, and through various air links. Physical elements that transmit such waves, such as wired or wireless links, optical links, etc., can also be considered media that carry software. As used herein, without being limited to persistent, tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.

[0106]

[0121] Thus, machine-readable media such as computer executable code may take many forms, including but not limited to tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media include optical or magnetic disks, such as any storage device in any computer(s) shown in the figures, which may be used to run a database. Volatile storage media include dynamic memory, such as the main memory of a computer platform. Tangible transmission media include coaxial cables, including the wires that comprise the bus in a computer system; copper wire and optical fiber. Carrier wave transmission media may take the form of electric or electromagnetic signals, or sound or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer readable media include, for example: floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punch cards paper tape, any other physical storage media using a pattern of holes, RAM, ROM, PROMs and EPROMs, FLASH-EPROMs, any other memory chips or cartridges, carrier waves transmitting data or instructions, cables or links transmitting such carrier waves, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.

[0107]

[0122] The computer system 1601 includes or is in communication with an electronic display 1635 that includes a user interface (UI) 1640 for providing, for example, the sequence of a polypeptide or antigen concentration for each image. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.

[0108]

[0123] The methods and systems of the present disclosure may be implemented by one or more algorithms. The algorithms may be implemented by way of software upon execution by the central processing unit 1605. The algorithms may, for example, generate sequences for polypeptides, calculate binding coefficients, or fit curves. EXAMPLES

[0109] Example 1 Generation of nanobodies

[0124] Nanobodies (or VHHs) are a class of single-domain antibodies found in camelids, including camels, llamas and alpacas. Nanobodies, composed of a single variable heavy chain, exhibit high specificity and affinity for their antigenic targets and often have favorable immunogenicity and toxicity profiles. Due to their small size (approximately 15 kDa), they are easily produced and can be more stable than conventional antibodies. These properties make nanobodies exciting targets for the development of novel therapeutics. Indeed, since their discovery in the 1990s, nanobodies have increasingly entered clinical trials as drug candidates to combat a variety of diseases, including a number of cancers, thrombotic thrombocytopenic purpura, inflammation and Alzheimer's, among others.

[0110]

[0125] Since the end of 2019, the global pandemic caused by the SARS-CoV novel coronavirus has infected more than 80 million people worldwide, resulting in the deaths of approximately 2 million people. The viral envelope is studded with multiple copies of a spike protein that binds to the angiotensin-converting enzyme 2 (ACE2) receptor on human epithelial cells, thereby initiating viral entry. Many groups have therefore focused on developing affinity reagents that can bind to this spike protein, and several V HHThe sequences were reported to have shown both high affinity binding to the spike protein and high levels of neutralization of viral entry in vitro. Moreover, pharmaceutical companies have already begun clinical trials to test the efficacy of spike-binding nanobodies.

[0111]

[0126] Sy62 is an anti-SARS-CoV-2 VHH previously described in the literature. Sy62 exhibits high signal-to-noise ratio and excellent binding affinity (apparent K of approximately 3.4 nM). D ) and was used as a reference sequence to generate variants. An initial optimization of the display was performed by generating a polypeptide library with different spacer and linker regions. Various C-terminal spacers and N-terminal linkers were screened. Screening of successful displays is analyzed by observing proper folding and function of the VHHs on the display chip. Figure 1A shows a schematic diagram for display screening, where about 1,200 to about 30,000 combinations are displayed and analyzed for binding. Figure 1B shows a schematic example of a library of polypeptides displayed using ribosome display, where different shapes show different N-terminal linkers and C-terminal spacers that can be displayed.

[0112]

[0127] Individual amino acids within the complementarity determining region (CDR) regions of Sy62 that contribute to binding were then analyzed by creating large targeted mutation libraries and measuring the effect of each mutation on binding as well as characterizing cooperative interactions between mutations.

[0113]

[0128] Such an analysis would yield a comprehensive list of functional mutations within the Sy62 CDRs, providing a means for affinity tuning and improvement. To generate these data sets, a multi-pronged approach was used. In the first experiment, the mutant affinity landscape of the Sy62 CDR, containing approximately 90,000 different variants, was split into three separate sub-libraries. The first sub-library contained an exhaustive set of single mutants in which each CDR residue was mutated to all 20 possible amino acids using a degenerate NNK codon. In the second sub-library, compensatory mutations between interacting residues in the Sy62 CDRs were identified. Candidate intra- and inter-CDR interacting residues were identified by analyzing the crystal structure of the parent nanobody from which Sy62 was derived, and pairs of residues were then mutated to all possible double mutation combinations. The third and final sub-library investigated the dependence of Sy62 binding affinity on the length of the CDR3, including single residue insertions at each position in addition to all possible deletions within the length range of 1-17 amino acids. These three CDR sub-libraries were assembled into six different framework scaffolds consisting of the wild-type (WT) Sy62 framework (FR) with some diversity introduced at four key residues in the FR2 framework region. The libraries were constructed by generating multiple polynucleotides encoding the polypeptide variants and then using ribosome display on a sequencing chip.

[0114]

[0129] FIG. 2 is a schematic diagram of the general workflow for the first sub-library, where a DNA library is generated for every single point mutation, and then quantitative analysis can be performed. Specifically, analysis of the first sub-library was performed by displaying the polypeptides of the sub-library on a sequencing chip. First, a library of polynucleotides encoding the polypeptides was added and captured on the sequencing chip. The polynucleotides were sequenced to determine the location on the chip of each polynucleotide and the corresponding polypeptides that were subsequently displayed. Reagents for ribosome display were added (e.g., RNA polymerase, dNTPs, ribosomes, tRNA) to display the corresponding VHH polypeptides from each polynucleotide. To analyze binding, various concentrations of labeled SARS-CoV-2 RBD were added to the sequencing chip and allowed to bind to the displayed VHH polypeptides, and excess SARS-CoV-2 RBD was removed. Fluorescent signals from the labeled SARS-CoV-2 RBD were generated, and the intensity of each polypeptide was collected by imaging the sequencing chip. By imaging the chip against various concentrations of labeled SARS-CoV-2 RBD, binding curves were generated for each polypeptide on the chip, which can then be fitted to determine binding coefficients or other quantitative binding measurements.

[0115]

[0130] Protein presentation on massively parallel arrays (Prot-MaP) analysis of the first sublibrary revealed strong binding signals and diverse binding constants as well as a complex dependence of the CDRs on both amino acid position and identity. Certain residues were observed to be mutagenized without affecting binding, while others were only amenable to mutation to certain other amino acids. Furthermore, some amino acids showed increased binding when mutated. Indeed, residue CDR2.6 showed improved activity when mutated from WT to any of approximately 15 different amino acids. Furthermore, the second sublibrary not only confirmed that target-interacting residues were highly sensitive to mutation, but also allowed the identification of compensatory mutations that restored function in otherwise ineffective single mutants, validating the structure-guided approach and providing a potential method to optimize even highly sensitive residues. Figure 3 shows the apparent Kd (K d app ) are shown. In detail, single mutant CDR variants for each VHH were first grouped and binned by the sequence of their specific parent CDR. The binding data for each set of CDR variants were then constructed as individual heatmaps with the residues constituting the CDR on the x-axis and the identity of the 20 individual amino acids (at which each position was mutated) on the y-axis. The WT amino acid identity at each position is marked by a black square on the heatmap. The binding affinity of the variants in the heatmap is colored from light red (weak affinity) to dark red (high affinity). Variants for which no binding was observed even at the highest tested concentration are shown in white, while the highest affinity variants are colored purple. The variants can be grouped as neutral (Kd=1.5-7 nM), negative (Kd>7 nM) or positive (Kd≦1.5 nM) based on the wild-type Kd of 3.4 nM.

[0116]

[0131] In the second step of the process, we found variants of Sy62 capable of maintaining high affinity binding across diverse mutational landscapes through single mutant analysis, selecting 21 mutations at 13 positions from a total of 34 residues in the CDRs that showed comparable or improved signal and binding affinity compared to the wild type. This second library explored all possible combinations at any of positions 1 to 13 where all possible combinations of these neutral to beneficial (when considered individually) mutations were simultaneously accepted, resulting in a library containing approximately 200,000 Sy62 variants. Figure 4 shows a corresponding schematic diagram of the general workflow in which a first DNA library is generated and then quantitative analysis is performed. Using the data from the first DNA library, a second DNA library can be generated and quantitative analysis can be performed to generate optimized variants.

[0117]

[0132] Sequencing and Prot-MaP analysis of a library containing approximately 200,000 Sy62 variants identified variants that were highly distant in sequence space (differences from wild type (WT) by 13 mutations) that performed as well as or better than their parental sequences. Figure 5 shows results from analysis of the initial sublibrary ("first experiment"), and from a library generated based on variants identified in the initial sublibrary ("second experiment"). Figure 5A shows Sy62 CDR variants from each of the two experiments plotted as frequency histograms binned by the number of mutations observed in each experiment. In the first experiment (blue bars), the majority of variants differed from the WT sequence by one to three mutations. Neutral and beneficial mutations from this library were then combined with the second experiment (black bars) in multiple different permutations to generate a diverse combinatorial library of variants that differed from the WT sequence by between 3 and 17 mutations. Most members of the second library contained between 6 and 8 mutations from the WT. Figure 5B shows the apparent binding affinities (y-axis) of the variants from each of the two experiments (first experiment shown by the blue line; second experiment shown by the black line), ranked from highest affinity to lowest affinity, plotted as a function of rank (x-axis). In each experiment, the rank of the WT sequence is shown by a red dashed line. In the first experiment, less than 9% of the variants had improved affinity over the WT. The affinity maturation step produced an approximately 9-fold increase (from about 8.7% to about 77%) in the number of variants with greater affinity to the ligand than the WT between the two experiments. Figure 5C shows the apparent binding affinities of Sy62 variants from the first experiment (left panel, blue) and the second experiment (right panel, black) plotted individually on a 3-dimensional scatter plot as a function of mutation distance of each CDR from the Sy62 WT sequence. The apparent binding affinity of the variants is colored from light (weak affinity) to dark (high affinity).

[0118]

[0133] Some of the highest affinity variants identified differed by 7-11 mutations from the WT. Figure 4 shows selected high affinity variants (arrows) and highly mutated (gray) variants that were superior to the WT Sy62 nanobody (black). Fluorescence binding data of variants from the combinatorial library (second experiment) were fitted to a 1:1 equilibrium binding model. Figure 6 shows ligand binding (y-axis) as a function of ligand concentration (x-axis) with shaded areas indicating ± standard deviation at each fit parameter. The left panel shows selected variants (left curve) with binding affinity 17-28 times higher than WT Sy62 (right curve). These variants contained 7-11 mutations from the WT sequence. The right panel shows improved binding of variants (light gray line) that differed by 13 mutations from the WT sequence (dark gray line). Overall, approximately 75,000 variants were identified as having stronger binding affinity than the starting sequence, while the strongest binding variants showed a significant increase in apparent affinity compared to the WT, as shown in Figure 5B. ( K d app ) This shows an improvement of about 100 times.

[0119] Example 2

[0134] Generation of polypeptide fusions, multi-epitope or specific polypeptides.

[0135] Using a method similar to that described in Example 1, further composite polypeptides can be generated based on quantitative analysis of the polypeptide library. A first library containing scFv variants or VHH variants is generated. The first library contains the sub-libraries described in Example 1, for example, a sub-library containing 20 variants for each residue corresponding to a single amino acid substitution to each standard amino acid at each residue number. As in Example 1, the library is then subjected to a quantitative binding assay that allows a labeled antigen of interest to interact with the polypeptide library. Labeled antigen is added at various concentrations and the intensity of the label is imaged to determine the interaction at each concentration. Binding curves for each polypeptide are generated and fitted to determine quantitative binding characteristics. Once data related to the library is generated, a second library is constructed using information about the variants. For example, variants containing multiple mutations corresponding to combinations of variants with neutral or positive effects can be constructed for the second library. The second library is assayed to identify polypeptides with optimized or improved binding characteristics. These optimized polypeptides can be used as cores or domains for novel polypeptide constructs. Although the library is created using scFv or VHH, larger polypeptides or polypeptide fusions can be generated. Figure 7 shows a schematic diagram of the polypeptide fusions that can be generated. Based on the identification of the optimized scFv, a complete IgG antibody can be generated using the sequence information of the optimized scFv to encode an IgG antibody that includes the structure or sequence of the optimized scFv. A similar method can be used for the VHH library. As shown in Figure 7, the sequence of the optimized VHH can be used to construct VHH-Fc fusions, can be combined with other VHHs to generate multispecific or multi-epitope polypeptides, can be conjugated with drugs to generate antibody drug conjugates, and can be combined with chimeric antigen receptors to generate VHH-CARs. For multispecific or multi-epitope constructs, Figure 8 shows a Venn diagram of binding to various antigens.VHHs can be assayed individually for a specific antigen and then combined to allow for multispecificity.

[0120] Example 3

[0136] Generation of biepitope polypeptides

[0137] Bi-epitope polypeptide is a class of antibody or antibody fragment that can bind to two different epitopes on the same antigen.Bi-epitope antibody can have many different advantages over the antibody that targets a single epitope, including increased avidity for target antigen and reduced susceptibility to antigenic mutation that escapes antibody.For example, the bi-epitope VHH developed by Janssen / Johnson & Johnson has been approved by FDA for use as CAR-T cell therapy that targets BCMA for the treatment of relapsed / refractory multiple myeloma.

[0121]

[0138] Traditional approaches to develop biepitope antibodies have relied on prior knowledge of antibodies or antibody fragments that bind different epitopes on the target antigen, or have utilized low-throughput epitope binning methods to individually screen and discover pairs of antibody fragments that bind different epitopes on the same antigen. The Prot-MaP platform enables a systematic, high-throughput approach to screen large libraries of tandemly arranged VHHs to identify and characterize biepitope tandem VHHs (Figure 9). Input of VHHs into these libraries can be done in several ways, including but not limited to DNA synthesis, immunization of animals (alpaca, llama, rat, mouse, among others) and utilization of human immune repertoire sequences.

[0122]

[0139] Using publicly available sources, we identified a large set of VHHs targeting the SARS-CoV-2 spike and RBD proteins. To validate the RBD-binding activity of these VHHs, we first constructed a survey library in which all VHHs in the set were placed in the context of N-terminal linker and C-terminal spacer polypeptide diversity to optimize initial presentation. From this library, we identified several VHHs (and their associated presentation contexts) that bind with moderate to high affinity to the SARS-CoV-2 RBD. Next, to optimize the affinity of the selected VHHs, we generated a library containing single mutant variants of the 14 highest affinity VHHs identified in the previous step, as in Example 1. The library was sequenced and the affinity of these variants was quantitatively characterized in a Prot-MaP experiment. A series of fluorescently labeled SARS-CoV-2 RBD solutions at various concentrations were added sequentially to the sequencing chip, allowing them to bind to the displayed VHHs, and imaged. The fluorescent signal from the bound RBD was quantified and fitted to the binding curves used to derive the binding affinity of each presented VHH to its RBD target, thereby generating single mutant binding affinity landscapes that quantitatively describe the effect of specific amino acid changes to every residue in the CDR of each of these VHHs. Figure 10 shows the resulting heatmap of binding data for all single mutants from the subset of 14 VHHs.

[0123]

[0140] In a next step, the single mutant binding data was used to construct two additional libraries. First, a tandem VHH library was generated to explore the enhanced binding activity achieved through tandem display of pairs of VHHs. Single mutant variants with moderate affinity (Kd in the range of 5-30 nM) were selected from 12 of the 14 VHHs. To this set were added three positive control VHHs predicted to bind SARS-CoV2-RBD and two negative control VHHs not predicted to bind SARS-CoV-2 RBD. Then, all possible pairwise combinations of 17 VHHs connected to each other by a flexible protein linker were generated. Fourteen unique linker sequences, varying in length (12-30 amino acids), charge and predicted secondary structure, were used to connect each pair of VHHs. Finally, each pair was incorporated in a variety of different C-spacer contexts, as described in Example 1 and shown in schematic form in Figure 11, resulting in a library containing >80,000 variants. To identify the large increase in avidity expected from simultaneous bi-epitope binding of two high affinity VHHs, it is necessary to compare the affinity measured for the tandem pair (tandem data set) with the affinity of each constituent VHH as an individual monomer (monomer data set). In principle, it would be more efficient to generate both tandem and monomer data sets together on the same chip (instead of two separate experiments), but one of the challenges of doing so is that clustering and sequencing libraries of significantly different lengths together at the same time often results in large and unpredictable distortions in relative representation. To minimize such distortions, it is beneficial for library members to be sequenced together at similar lengths, and for this purpose we included pseudomonomer VHHs (consisting of a given VHH and a negative control "no effect" VHH arranged in both orientations (ab and ba)) that were used as a substitute for the individual, monomeric VHHs. The libraries were sequenced and assayed for binding to the SARS-CoV-2 RBD as described above.Tandem VHH pairs in a given orientation were thereby identified that bound to the RBD with an affinity significantly greater than the average affinity of the pseudomonomer VHHs in the pair (Figure 12).

[0124]

[0141] Using the single mutant binding data (Figure 10), a second library was constructed to optimize the affinity of the individual VHHs forming the di-epitope tandem pairs. As described in Example 1, an affinity optimized library was generated based on the data from the single mutant library and subjected to binding assays to identify individual VHHs with improved affinity over the starting variant (Figure 13).

[0125]

[0142] To generate the final affinity and avidity enhanced molecules, tandem VHH pairs showing significant avidity enhancement were reconstructed by replacing the moderate affinity single mutant VHHs in the tandem VHH pairs with the optimized tightest binding affinity variants of each VHH (Figure 14).

[0126]

[0143] Although preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the present invention be limited by the specific examples provided in the specification. Although the present invention has been described with reference to the foregoing specification, the description and illustration of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the present invention. Furthermore, it is understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions described herein, depending upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the present invention described herein may be used in carrying out the present invention. It is therefore contemplated that the present invention will cover any such alternatives, modifications, variations, or equivalents. It is intended that the following claims define the scope of the present invention, and that methods and structures within the scope of these claims and their equivalents are covered thereby.

Claims

1. A high-throughput method for identifying an optimized polypeptide, comprising: (a) providing a first library of polynucleotides encoding a first library of variant polypeptides; (b) processing the first library of polynucleotides to produce the first library of variant polypeptides, wherein the variant polypeptides are attached to the first library of polynucleotides; (c) identifying one or more characteristics including at least the equilibrium binding constant, kinetic binding constant, protein stability measurement, enzyme activity, fractionation activity, non-specific binding ability, aggregation ability, hydrophobicity, protein expression level, or maturation time of at least a portion of the first library of variant polypeptides; (d) providing a second library of polynucleotides encoding a second library of variant polypeptides selected based at least on the one or more characteristics identified in (c); (e) processing the second library of polynucleotides to produce the second library of variant polypeptides, wherein the variant polypeptides are attached to the second library of polynucleotides; and (f) analyzing the second library of variant polypeptides to generate optimized data A method comprising the steps of:

2. The method according to claim 1, further comprising (g) identifying an optimized polypeptide based on the optimized data.

3. wherein the equilibrium binding constant includes a dissociation constant (K d ), or an association constant (Ka), the method according to claim 1.

4. wherein the kinetic binding constant is an association rate constant (k on ), or a dissociation rate constant (koff), the method according to claim 1.

5. wherein the protein stability measurement value is the protein melting temperature (T m ), or the midpoint concentration (C m) of the chemical denaturant in the formulation, according to the method of claim 1.

6. The method according to claim 1, further comprising, in (d), identifying negative variants, positive variants, and neutral variants from the first library of variant polypeptides.

7. The neutral variant includes a dissociation constant greater than 0.25 times and less than 2 times the dissociation constant of the starting polypeptide; The positive variant includes a dissociation constant that is 0.25 times or less the dissociation constant of the starting polypeptide; or The negative variant includes a dissociation constant that is 2 times or more the dissociation constant of the starting polypeptide. The method according to claim 6.

8. The method according to claim 1, wherein the first library of variant polypeptides comprises single amino acid variants in which an amino acid residue is substituted with one amino acid selected from a set of amino acids.

9. The method according to claim 1, wherein the first library of variant polypeptides comprises single amino acid insertions, single amino acid deletions, double amino acid deletions, triple amino acid deletions, or deletions of at least 4 amino acids.

10. The method according to claim 1, wherein the first library of variant polypeptides comprises single amino acid variant polypeptides corresponding to at least 90% of the possible single nucleotide variants for a given reference sequence in a reference polypeptide.

11. The method according to claim 1, wherein the step of analyzing the first library of variant polypeptides comprises the steps of transcribing and translating the polynucleotides of the first library of variant polynucleotides, and the polypeptide encoded by the polynucleotide attaches to the polynucleotide.

12. The step of identifying the equilibrium binding constant, kinetic binding constant, protein stability measurement, enzyme activity, fractionation activity, non-specific binding ability, aggregation ability, hydrophobicity, protein expression level, or maturation time is performing a binding assay on the first library of variant polypeptides, or sequencing the first library of polynucleotides and correlating the sequence of the first library of polynucleotides with the binding assay The method according to claim 1, comprising.

13. The method according to claim 12, wherein the binding assay comprises assaying the binding of the first library of variant polypeptides to one antigen, more than one antigen, or a plurality of antigens.

14. Binding to two or more of the plurality of antigens, Binding to at least one of the plurality of antigens and not binding to a different antigen among the plurality of antigens, or Not binding to the plurality of antigens The method according to claim 13, further comprising the step of identifying a variant polypeptide.

15. The method according to claim 1, wherein the step of identifying the equilibrium binding constant, kinetic binding constant, protein stability measurement value, enzyme activity, fractionation activity, non-specific binding ability, aggregation ability, hydrophobicity, protein expression level, or maturation time includes the step of generating binding data for more than one target.

16. The method according to claim 15, wherein the second library is generated based at least on binding data for more than one target.

17. The method according to claim 1, wherein the step of identifying the optimized polypeptide includes the step of performing a binding assay on the second library of variant polypeptides encoded by the second library of polynucleotides.

18. The step of identifying the equilibrium binding constant, kinetic binding constant, protein stability measurement value, enzyme activity, fractionation activity, non-specific binding ability, aggregation ability, hydrophobicity, protein expression level, or maturation time is performing a binding assay on the second library of variant polypeptides encoded by the second library of polynucleotides, or sequencing the second library of polynucleotides and associating the sequence of the second library of polynucleotides with the binding assay The method according to claim 17, comprising.

19. The second library of variant polypeptides comprises at least 10 4 polypeptides, the method according to claim 1.

20. The first library of polynucleotides contains at least 10 6 The method according to claim 1, comprising polynucleotides.

21. The method according to claim 1, wherein the first library of variant polypeptides comprises a library of individual VHH antibodies.

22. The method according to claim 21, wherein the second library of variant polypeptides comprises a library of VHH antibody fusions.

23. The method according to claim 1, wherein the first library of variant polypeptides comprises a library of individual single-chain variable fragments (scFvs).

24. The method according to claim 23, wherein the second library of variant polypeptides comprises a library of individual single-chain variable fragment (scFv) fusions.

25. The method according to any one of claims 1 to 24, wherein the second library of variant polypeptides is generated based on an algorithm or machine learning configured to identify one or more variant polypeptides from the first library of variant polypeptides.