Methods and compositions for protein detection
Patent Information
- Application Number
- JP2024527085
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-11-03
- Filing Date
- 2022-11-02
- Publication Date
- 2025-11-11
AI Technical Summary
Current methods for evaluating binding agents, such as phages and antibodies, are slow, costly, and limited in throughput, especially when assessing multiple agents in complex environments like in vivo systems, and existing DNA barcoding techniques suffer from instability and immunogenicity.
The use of peptide barcodes associated with binding agents, such as phages, allows for high-precision detection and measurement of multiple agents in complex systems by sequencing nucleic acids without covalent modification, enabling simultaneous evaluation of multiple characteristics of therapeutic agents.
This approach provides accurate, high-throughput evaluation of proteins and therapeutic agents with minimal impact on their functionality, suitable for complex environments, and allows for rapid generation and characterization of binders and barcodes.
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Application No. 63 / 275,172, filed November 3, 2021, and U.S. Application No. 17 / 518,221, filed November 3, 2021, the disclosures of which are incorporated herein by reference in their entireties. [Background technology]
[0002] The evaluation of binding agents (eg, phage, polypeptide binders, antibodies, etc.) is central to much of molecular and pharmaceutical biology.
[0003] Sequence Listing This specification references a Sequence Listing (submitted electronically as an .xml file named ST26_2013703-0005). The .xml file was generated on October 18, 2022 and is 12,323,567 bytes in size. The entire contents of the Sequence Listing are incorporated herein by reference. Summary of the Invention
[0004] The present disclosure provides insights and techniques to achieve improved or otherwise desirable evaluation of agents (e.g., binding agents, therapeutic agents, and / or polypeptide agents, such as, in some embodiments, antibody agents).
[0005] Among other things, the present disclosure recognizes that many current techniques for assessing, and particularly determining, the presence and / or abundance of one or more agents of interest typically rely on mass spectrometry and / or affinity (e.g., immuno) detection. The present disclosure recognizes that many available affinity detection techniques are slow and / or costly to perform or implement, many such techniques must be performed one at a time, and many are limited, for example, by the availability of fluorogenic substrates (which can be assessed, for example, by related techniques such as light microscopy).
[0006] The present disclosure further recognizes that certain other techniques, such as DNA barcoding techniques, sometimes utilized to evaluate agents of interest may also suffer from drawbacks. DNA barcodes, for example, may lack stability and / or exhibit undesirable immunogenicity when utilized in vivo. Thus, the present disclosure recognizes that such techniques may encounter problems, particularly with respect to the evaluation of agents (e.g., protein agents) in complex environments (e.g., in vivo).
[0007] Among other things, the present disclosure encompasses the recognition of the source of certain problems with available techniques typically utilized to evaluate agents of interest, particularly protein agents and / or binding agents (i.e., agents involved in one or more binding interactions of interest). In particular, the present disclosure identifies the source of certain problems encountered with such techniques with respect to the evaluation (e.g., detection and / or measurement of quantities such as concentrations) of multiple agents, particularly when such agents are present in complex systems (e.g., in complex solutions and / or in vivo).
[0008] Moreover, the present disclosure provides, in some embodiments, certain techniques for achieving such assessments with a surprisingly high degree of accuracy. Those skilled in the art will appreciate that there are many contexts in which detection and / or measurement (e.g., precise amounts) of multiple agents within a complex system is desirable. Moreover, those skilled in the art will appreciate the benefits of high accuracy in many such contexts.
[0009] Among other things, the present disclosure provides techniques for achieving detection and / or measurement (e.g., high precision and / or otherwise accurate measurement) of one or more, and in some embodiments, multiple, agents (e.g., protein agents) comprising complex systems (e.g., in vivo). In some embodiments, the detected agent(s) may be or may include proteins (i.e., polypeptides) and / or forms thereof (e.g., aggregated, complexed, covalently modified by disulfide bond formation, glycosylation, pegylation, phosphorylation, etc., cleaved by proteolytic cleavage, etc.).
[0010] In some embodiments, the provided techniques are particularly useful or effective for evaluating therapeutic agents. For example, in some embodiments, the provided techniques may be particularly useful for evaluating one or more characteristics (e.g., properties (e.g., concentration, localization, persistence, affinity, etc.)) of an agent(s) of interest, and in some such embodiments, the relevant agent(s) may be characterized by one or more attributes that are appropriate or desirable for therapeutic use. For example, in some embodiments, the provided techniques can be used to screen potential therapeutic agents (e.g., polypeptide entities) for one or more characteristics (e.g., properties, attributes) that are suitable for therapeutic use. In some embodiments, each characteristic of the potential therapeutic agent(s) may be measured one at a time. In some embodiments, two or more characteristics of the potential therapeutic agent(s) may be measured simultaneously. For example, in some embodiments, one or more therapeutic agents may be screened for affinity for a target protein, while other desirable characteristics, such as molecular stability in a physiologically relevant environment, are not yet known. In some embodiments, for example, one or more therapeutic agents may be screened for affinity to a target protein along with other desirable properties, such as molecular stability in a physiologically relevant environment.
[0011] The present disclosure recognizes that many current methods of protein measurement rely on determining the abundance or overall emission of light at certain wavelengths, such as Western blots or ELISAs (Towbin 1979, Engvall 1972). Due to the constraints of visible light wavelengths, these methods are only able to measure a small number of different proteins at a time, often less than four, within a single reaction (Elshal, 2006). The present disclosure recognizes that many applications, including drug discovery applications, benefit from (and in some cases require) dramatically higher throughput.
[0012] The present disclosure further recognizes that nucleic acid sequencing technologies (e.g., DNA sequencing technologies) have been developed that are capable of analyzing billions of individual DNA molecules in a single experiment (Shendure, 2005). In an effort to apply this massive throughput achievable with nucleic acid sequencing techniques to protein detection and measurement, various strategies have been developed, particularly by tagging proteins with DNA fragments (typically referred to as "DNA barcodes") attached to proteins, which can then be sequenced to indirectly detect the protein (Trads, 2017) or one or more features of the protein.
[0013] While this disclosure recognizes the power of applying high-throughput nucleic acid sequencing technology to the evaluation of other agents, particularly protein agents, it also identifies the source of certain problems associated with many approaches utilized to study proteins through the attachment of DNA barcodes. For example, this disclosure recognizes that modifying a protein through the attachment of a DNA barcode can often alter its functionality (Trads, 2017), which may defeat the purpose of using DNA barcodes to evaluate proteins.
[0014] Known techniques for quantifying barcoded proteins include those presented in Egloff et al. (2019), which use mass spectrometry to determine the presence or absence of protein sequences in a mixture (Egloff 2019). However, the present disclosure identifies sources of problems with such approaches and further provides certain advantages over them, including, for example, by using nucleic acids (e.g., DNA) to amplify the original signal; approaches such as those described in Egloff et al. cannot include (or benefit from) such features. Furthermore, as will be understood by those skilled in the art upon reading this disclosure, methods using mass spectrometry are limited in their overall throughput because mass spectrometry only reads the mass-to-charge ratio of related sequences, and therefore, different sequences may have the same mass-to-charge ratio. In comparison, the present invention is not limited by such drawbacks, as nucleic acid sequences associated with one or more binding agents associated with each barcode are sequenced and measured to determine and quantify the barcoded protein.
[0015] Other techniques available in the art use antibodies displayed on phages (Fab-phages) to determine the presence of endogenous proteins expressed on cell surfaces (Pollock, 2018). In such methods, one Fab-phage is generated per endogenous protein (i.e., the target protein being evaluated), and barcodes are not utilized. In contrast, the present technology contemplates the use of generalizable engineered barcode sequences that can be used to mark any protein, whether endogenous or exogenous to the context in which it is applied, and then measured using one or more binding agents with which each barcode, and thus each barcoded protein, uniquely associates (i.e., "barcode fingerprints," as described elsewhere in this disclosure). Such complex association of one or more binding agents with the barcode is then measured, and accurate quantification of the associated proteins is achieved, for example, using a complex algorithm (i.e., "decoding," as described elsewhere in this disclosure).
[0016] The present disclosure recognizes the ability of antigens displayed on phage to determine the epitopes of antibodies in blood to which the phage can bind (Mohan, 2018). However, this method provides limited information about any antibodies that specifically bind to the antigen displayed on the phage because it cannot determine the sequence of the antibody to which the antigen binds. However, the present disclosure provides systems, compositions, and methods that offer the advantage of using a generic barcode with known affinity for one or more binders or binding agents, which can be used to tag any target(s) of interest in a complex mixture, including but not limited to blood, and to determine and quantify the target(s).
[0017] The present disclosure provides, among other things, techniques by which the evaluation (e.g., detection and / or quantification) of multiple agents (e.g., multiple protein agents) within a pool of such agents can be accomplished using DNA sequencing without requiring covalent association (direct or indirect) between the DNA and the agents being evaluated or otherwise limiting the agents being evaluated.
[0018] Described herein are peptide barcodes (also known as "barcodes") and techniques for creating and / or utilizing them. In some embodiments, barcodes are utilized to mark payloads. Notably, such an approach can achieve pooled measurement of protein payloads without modifying non-protein identifiers. In some embodiments, the peptide barcode is an amino acid polypeptide sequence. In some embodiments, the peptide barcode is contained within a protein (e.g., the protein being measured (e.g., an antibody); i.e., endogenous to the protein being measured). In some embodiments, the peptide barcode is not contained with the protein (e.g., the protein being measured (e.g., an antibody); e.g., exogenous to the protein being measured). In some embodiments, the barcode is a sequence (e.g., a designed sequence) contained within, for example, the protein (e.g., the protein being measured). In some embodiments, the barcode is associated with (e.g., attached (e.g., covalently linked)) the N-terminus of the protein (e.g., the protein being measured). In some embodiments, the barcode is associated with (e.g., attached (e.g., covalently linked)) the C-terminus of the protein (e.g., the protein being measured). In some embodiments, the barcode is associated (e.g., attached (e.g., covalently linked)) proximal to the N-terminus of the protein (e.g., the protein to be measured) (e.g., within the protein (e.g., in a loop region proximal to the N-terminus). In some embodiments, the barcode is associated (e.g., attached (e.g., covalently linked)) proximal to the C-terminus of the protein (e.g., the protein to be measured), e.g., within the protein (e.g., in a loop region proximal to the C-terminus).
[0019] The methods disclosed herein may use peptide barcodes designed to have various lengths. In some embodiments, peptide barcodes may have lengths ranging from 1 to 100, 5 to 50, 8 to 25, 9 to 25, or 9 to 15 amino acids. In some embodiments, peptide barcodes may have a length of at least 25 amino acids. In some embodiments, peptide barcodes may have a length of up to 8 amino acids. In some embodiments, peptide barcodes may have a length of 10 amino acids.
[0020] The barcode sequences described herein can be reused to quantify different payloads (e.g., proteins of interest) or mixtures of payloads (mixtures of proteins of interest being measured). In some embodiments, the barcode is generated so that it can be easily reused across several different payload (e.g., protein to be measured) molecules across different experiments.
[0021] Notably, the barcodes described herein are designed to be distinct / unique. In some embodiments, the barcodes are designed to have distinct (e.g., distinct from other barcodes) sequences. For example, each barcode is designed to be different (e.g., unique) from all other barcodes used in the experiment, each payload (e.g., protein to be measured) is attached to at least one barcode, and each barcode (e.g., barcode with a particular sequence) is attached to only one payload. As will be understood by those skilled in the art, the diversity of barcodes contained within a pool is limited only by the possible diversity of amino acid sequences for a given barcode length. For example, for a barcode length "N," 20 of the length N may be used. N There are distinct amino acid barcode sequences.
[0022] The methods described herein relate to the detection of one or more barcodes using a binder. In some embodiments, the barcode is contacted with a binder that is associated with or comprises a detectable nucleic acid. For example, in some embodiments, the binder may be or comprise a phage, ribosome, mRNA, DNA, etc. In some embodiments, the binder is a phage (e.g., a polypeptide binder as described herein) having a binding motif on its surface. In some embodiments, the binder comprises a detectable nucleic acid. In some embodiments, the binder expresses a detectable nucleic acid. In some embodiments, the binder expresses a detectable nucleic acid on (e.g., on) the binder (e.g., binder). In some embodiments, the binder is a polypeptide. In some embodiments, the binder is associated with (e.g., has known specificity and affinity for) the barcode. In some embodiments, the binder is associated with (e.g., has known specificity and affinity for) one or more barcodes. In some embodiments, the binder is an antibody (e.g., expressed on the surface of the binder). In some embodiments, for example, to detect the presence of a specific (e.g., distinct) barcode, the present disclosure contemplates the association of a distinct detectable nucleic acid (e.g., a DNA sequence, an RNA sequence, etc.) to the specific barcode. This is accomplished by contacting a binder, which may be expressed on (e.g., on the surface of) a binding agent that includes the distinct detectable nucleic acid.
[0023] Described herein are binders. In some embodiments, the binder is a polypeptide. In some embodiments, for example, the binder is generated to have a known specificity and affinity for a given barcode. In some embodiments, the binder is generated to have a known specificity and affinity for one barcode. In some embodiments, the binder is generated to have a known specificity and affinity for multiple (e.g., two or more, three or more, etc.) barcodes. In some embodiments, the binder is generated to have a known specificity and affinity for at least one barcode. In some embodiments, the binder is expressed on the surface of a binding agent (e.g., a phage, a ribosome, etc.), for example, using methods known to those of skill in the art.
[0024] Among other things, for example, the systems and methods described herein identify the advantages of nucleic acid sequencing techniques and effectively apply them to protein detection and measurement methods. For example, the methods described herein can use several binders with known specificity and affinity for different barcodes, which can be expressed on a binder and mixed together in a single pool. When mixed with a pool of barcoded proteins (i.e., proteins that each associate with a barcode as described herein), the binders expressed on the binders bind to any given barcode in the pool with known but varying affinities. This spectrum of binder affinities for various barcodes is referred to herein as a "binder fingerprint." Conversely, a barcode can bind to any given binder in a pool of binders with known but varying affinities. This spectrum of barcode affinities for various binders is referred to herein as a "barcode fingerprint." Thus, the presence of a particular barcoded protein can be detected, for example, by extracting and sequencing associated nucleic acids (e.g., detectable nucleic acids (e.g., DNA sequences, RNA sequences, etc.)) of a population of binding agents (e.g., phages) bound to barcodes associated with the protein in a complex solution.
[0025] Other methods using binders to identify protein sequences have been developed. However, these methods encounter several challenges, including the difficulty of generating and characterizing binders and effectively decoding their binding to specifically identify proteins. Another limitation of previously developed binders is their nonspecific binding, which results in poor signal-to-noise ratios and thereby adversely affects the accuracy of detection. In contrast, the present technology rapidly generates large numbers of binders (e.g., in about 1 week, about 2 weeks, about 3 weeks, about 4 weeks, about 1 month, about 2 months, about 3 months, about 4 months, about 5 months, about 6 months, or about 1 year). In some embodiments, for example, about 100 to about 1,000 binders can be rapidly generated. In some embodiments, about 10 to about 1,000 binders can be rapidly generated. In some embodiments, about 10 to about 10,000 binders can be rapidly generated. In some embodiments, at least about 10,000 binders can be rapidly generated.
[0026] The binders described herein are robust. The binders can bind to barcodes as described herein (e.g., with robust affinity for one or more barcodes) in a variety of conditions and / or environments. For example, the binders described herein can bind to barcodes (e.g., with robust affinity for one or more barcodes) in a variety of complex environments (e.g., blood, tissue, serum, plasma, etc.). Thus, the binders of the present disclosure can be used to detect targets (e.g., proteins of interest) in a variety of conditions (e.g., physiological conditions).
[0027] Similarly, the barcodes described herein can be generated in a rapid and robust manner. In some embodiments, the barcodes described herein are specific to the binders described herein. In some embodiments, for example, about 100 to about 2000 barcodes can be generated rapidly (e.g., in about 1 week, about 2 weeks, about 3 weeks, about 4 weeks, about 1 month, about 2 months, about 3 months, about 4 months, about 5 months, about 6 months, or about 1 year). In some embodiments, about 10 to about 1000 barcodes can be generated rapidly. In some embodiments, about 10 to about 10,000 barcodes can be generated rapidly. In some embodiments, at least about 10,000 barcodes can be generated rapidly.
[0028] The barcodes described herein are robust. The barcodes can bind to binders (e.g., with robust affinity for one or more binders) as described herein in a variety of conditions and / or environments. For example, the barcodes described herein can bind to binders (e.g., with robust affinity for one or more binders) in a variety of complex environments (e.g., blood, tissue, serum, plasma, etc.). Thus, the barcodes of the present disclosure can be used to detect targets (e.g., proteins of interest) in a variety of conditions (e.g., physiological conditions). Thus, the present disclosure corrects the shortcomings and deficiencies of existing methods (e.g., nonspecific binding, variable binding in different environments, etc.) by rapidly generating large numbers of robust binders and barcodes that can be used in combination with the computational methods described herein (e.g., deconvolution methods) to enable specific, well-characterized binder-barcode binding / association and accurate detection methods.
[0029] The present disclosure also contemplates the ability to modify the sequence(s) of one or more peptide barcode sequences so that they are readily distinguishable from one another and / or from potential background protein sequences. Similarly, the present disclosure also contemplates the ability to modify the sequence(s) of one or more polypeptide binder sequences so that they are readily distinguishable from one another and / or from potential background protein sequences.
[0030] Among other things, the invention described herein provides methods for testing "n" distinct protein candidates, where n > 1, in a single assay or animal model. In some embodiments, the protein candidates are therapeutic protein candidates. In some embodiments, multiple protein candidates are designed, with each distinct protein candidate associated with a unique peptide barcode as described herein. Such barcoding has many advantages, including, but not limited to, cost- and time-efficiently injecting all protein candidates into assays and / or animals in a single injection. A sample (e.g., tissue sample, serum sample, blood sample, extracellular sample, single-cell sample, etc.) from the injected animal can then be obtained and barcodes extracted. In some embodiments, such extracted barcodes provide a measure of the relative abundance of the originally injected protein candidate. For example, one or more extracted barcodes can be identified by contacting them with a pool of binders (e.g., expressed on binders) known to bind to the barcode originally bound to the protein candidate. After binding of the barcodes and binders, the bound binders (e.g., phage) are selected and their detectable nucleic acids (e.g., DNA sequences, RNA sequences, etc.) are extracted. In some embodiments, the extracted nucleic acids are subjected to sequencing (e.g., next-generation sequencing). The sequenced nucleic acids may then be used to identify one or more barcodes to which they are designed to bind, which, together with previously established information about the binding affinities between various binder-barcode pairs, may be used to identify and determine the relative abundance of each protein originally injected.
[0031] Also described herein are methods used to convert nucleic acid counts, e.g., from a sequencing experiment, to relative or absolute protein quantification. In some embodiments, nucleic acid sequences are counted and translated in silico into protein sequences. As described herein, the nucleic acid sequences correspond to binder sequences with established and characterized affinities for all barcodes provided in a pool. In some embodiments, binder counts are compared to a database of known propensities for binding to a single barcode. In some embodiments, binder counts are compared to a database of known propensities for binding to multiple barcodes (e.g., two or more, three or more, etc.). In some embodiments, the relative proportions of binder counts are compared directly to determine the relative proportions of barcodes and / or proteins associated with the barcodes, e.g., in a sequencing experiment. In some embodiments, as would be known to one of skill in the art, sequences (e.g., control or accessory sequences) of known abundance (e.g., count, quantitation, concentration, etc.) can be utilized (e.g., added to a sequencing experiment) to determine the absolute abundance (e.g., count, quantitation, concentration, etc.) of a given binder(s), which can then be used to estimate the absolute abundance (e.g., count, quantitation, concentration, etc.) of the barcode(s) and / or protein(s) associated with the barcode(s) using either direct counting or linear models as described herein.
[0032] In some embodiments, the payload is or comprises a protein. In some embodiments, the payload is or comprises a therapeutic protein. In some embodiments, the payload is associated with (e.g., linked to) a barcode as described herein.
[0033] Among other things, the present disclosure provides methods for evaluating the barcodes, binders (e.g., binding agents (e.g., having binders expressed on their surfaces)), and payloads (e.g., barcoded payloads (e.g., barcoded proteins)) described herein. In some embodiments, the methods include subjecting a population of barcoded payloads (e.g., barcoded proteins) to evaluation, separating members of the population that satisfy the evaluation from those that do not so that either a positive population or a negative population, or both, are identified, contacting the positive population or the negative population, or each population separately from the other, with a set of binders comprising at least one specific binder specific for each barcode in the population, and determining which binders bind to the separated members, thereby determining which barcoded payloads (e.g., barcoded proteins) are present in the contacted population(s).
[0034] The present disclosure provides methods including contacting a set of binders with either a first population, a second population, or each of the first and second populations separately of barcoded payloads (e.g., barcoded proteins); and determining which binders of the set bind to members of the first population, the second population, or both, thereby determining which barcoded payloads (e.g., barcoded proteins) are present in the contacted population(s). In some embodiments, each binder specifically binds (e.g., with a known affinity) to one or more barcodes. In some embodiments, the set of binders collectively comprises at least one binder specific for each of the barcodes in the first population and the second population. In some embodiments, the first population and the second population are separated from each other based on performance in an evaluation.
[0035] In some embodiments, the method further comprises determining a difference between the first population and the second population to determine a functional effect of the performance assessment. In some embodiments, the method comprises isolating binders that bind to at least one payload (e.g., a barcoded payload (e.g., a barcoded protein)).
[0036] In some embodiments, the determining step comprises quantifying the number of binders that bind to the barcoded payload (e.g., barcoded protein). In some embodiments, the quantifying may be performed by decoding the nucleotide sequence of each binder that binds to the barcoded payload (e.g., barcoded protein). In some embodiments, quantifying the number of binders that bind to the payload (e.g., barcoded protein) provides a measure of the payload (e.g., protein) in the population.
[0037] In some embodiments, the determining step comprises amplifying nucleic acid of the bound phage particles. In some embodiments, the determining step comprises determining a nucleotide sequence of the amplified nucleic acid. In some embodiments, one or more of the determined nucleotide sequences corresponds to a coding sequence of the binder. In some embodiments, the determining step comprises detecting one or more payloads (e.g., proteins) from a population of barcoded payloads (e.g., barcoded proteins) using the determined sequence(s) of the coding sequence of the binder. In some embodiments, the determining step comprises identifying one or more barcoded payloads (e.g., barcoded proteins) as therapeutics or targets for treating a disease, disorder, or condition.
[0038] In some embodiments, the determining step includes performing one or more of amplification, amplification, and sequencing (e.g., nucleic acid (e.g., DNA, RNA) amplification, amplification, and / or sequencing). In some embodiments, the amplification may be performed using one or more of polymerase chain reaction (PCR), loop-mediated isothermal amplification (LAMP), rolling circle amplification (RCA), or similar known techniques. In some embodiments, the sequencing may be performed using one or more of Illumina, next-generation sequencing (NGS), nanopore sequencing, Pac Bio long-read sequencing, or similar known techniques.
[0039] In some embodiments, the separating step comprises purifying one or more barcoded payloads (e.g., barcoded proteins) from the sample. In some embodiments, the barcoded payloads (e.g., barcoded proteins) are purified from a complex sample. In some embodiments, the barcoded payloads (e.g., barcoded proteins) are purified from a complex mixture. In some embodiments, the barcoded payloads (e.g., barcoded proteins) are purified using affinity purification methods (e.g., FLAG IP, protein G / A) or protein precipitation methods.
[0040] In some embodiments, the method further comprises injecting the population of barcoded payloads into the animal. In some embodiments, the method further comprises injecting the population of barcoded payloads (e.g., barcoded proteins) into the animal. In some embodiments, each barcode is bound to a specific binder expressed on a phage. In some embodiments, the method further comprises obtaining a sample from the animal to be evaluated.
[0041] In some embodiments, the methods described herein involve determining the relative amount of each binder present in the sample, thereby identifying a subset of the injected population of barcoded payloads (e.g., barcoded proteins) present in the sample. In some embodiments, the methods described herein involve comparing the relative amounts to standards of known concentration to determine the absolute amount of each binder present in the sample.
[0042] In some embodiments, the methods described herein optionally include repeating one or more of the method steps described herein using the identified subset of payloads (e.g., proteins).
[0043] In some embodiments, the methods described herein involve identifying one or more payloads (e.g., proteins) as therapeutic agents or targets for treating a disease, disorder, or condition.
[0044] In some embodiments, the methods described herein include removing any unassociated (e.g., unbound) binders. In some embodiments, removal can be done by washing.
[0045] In some embodiments, the barcoded payload (e.g., barcoded protein) is in a sample. In some embodiments, the barcoded payload (e.g., barcoded protein) is in a complex sample. In some embodiments, the barcoded payload (e.g., barcoded protein) is in a complex mixture. In some embodiments, the barcoded payload (e.g., barcoded protein) is in a purified sample.
[0046] In some embodiments, the sample is or comprises one or more of serum, blood, tissue, or tumor. In some embodiments, the sample is a control (e.g., a positive control or a negative control).
[0047] In some embodiments, the sample is a complex sample. In some embodiments, the complex sample is or comprises tissue. In some embodiments, the complex sample is or comprises blood. In some embodiments, the complex sample is a complex mixture. In some embodiments, the complex sample is or comprises one or more of serum, blood, or tissue.
[0048] In some embodiments, the barcode is or comprises one or more amino acids. In some embodiments, the barcode is comprised in a complementarity determining region (CDR) of a payload (e.g., a protein). In some embodiments, the barcode is synthetic. In some embodiments, the barcode is 1-100, 5-50, 8-25, 9-25, or 9-15 amino acids in length. In some embodiments, the barcode is 10 amino acids in length. In some embodiments, the barcode has relatively little effect on payload (e.g., protein) function. In some embodiments, the barcode does not tamper with the immune response. In some embodiments, the barcodes are orthogonal to one another. In some embodiments, at least one barcode is linked to a polypeptide of interest (e.g., a polypeptide binder, a payload).
[0049] In some embodiments, the barcode is attached to a payload (e.g., a protein). In some embodiments, the barcode is attached to a suitable position on the payload (e.g., a protein). In some embodiments, the suitable position is the N-terminus or C-terminus.
[0050] In some embodiments, the binders are or comprise binding moieties displayed on a phage. In some embodiments, each binder of a set of binders is expressed on a phage. In some embodiments, the binders are expressed on the surface of a phage particle.
[0051] In some embodiments, the phage is selected from the group consisting of M13, T4, T7, lambda, and filamentous phage, hi some embodiments, the phage is M13.
[0052] The present disclosure provides, inter alia, nucleic acids whose nucleotide sequence is or includes a sequence encoding a peptide barcode. In some embodiments, the peptide barcode has a length in the range of 1-100, 5-50, 8-25, 9-25, or 9-15 amino acids. In some embodiments, the peptide barcode has a length of 8-25 amino acids. In some embodiments, the peptide barcode has a length of 10 amino acids. In some embodiments, the peptide barcode has been determined to specifically bind to a particular group of polypeptide binders within a set of binders.
[0053] In some embodiments, the peptide barcode has an amino acid sequence selected from the group consisting of SEQ ID NOs: 5347 to 8398. In some embodiments, the coding sequence is selected from the group consisting of SEQ ID NOs: 1148 to 4199.
[0054] The present disclosure provides a library comprising a plurality of nucleic acids. In some embodiments, the plurality of nucleic acids together encode a collection of peptide barcodes. In some embodiments, each nucleic acid comprises, in 5' to 3' or 3' to 5' order, one or more of: a) a first invariant sequence (e.g., a linker sequence or a payload sequence), b) a variant sequence at least 9 nucleotides in length, and c) a second invariant sequence (e.g., a linker sequence, a stop codon, or a payload sequence).
[0055] In some embodiments, the variant sequence is at least 15, 24, 27, 45, 150, or 300 nucleotides in length.
[0056] In some embodiments, the library further comprises one or more of: d) sequences containing short helical motifs; e) sequences containing disordered motifs; and f) invariant sequences linking the sequences to a protein of interest.
[0057] In some embodiments, each peptide barcode of the collection specifically binds to a particular group of polypeptide binders within the set of binders, hi some embodiments, each peptide barcode of the collection specifically binds to one or more polypeptide binders within the set of binders.
[0058] The present disclosure provides nucleic acids, the nucleotide sequence of which is or comprises a sequence encoding a polypeptide binder moiety. In some embodiments, the polypeptide binder moiety has a length in the range of 10 to 400 amino acids. In some embodiments, the polypeptide binder moiety has been determined to specifically bind to a particular group of peptide barcodes within a collection of barcodes.
[0059] In some embodiments, the polypeptide binder portion comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 4200 to 5346. In some embodiments, the coding sequence is selected from the group consisting of SEQ ID NOs: 1 to 1147.
[0060] The present disclosure provides a library comprising a plurality of nucleic acids. In some embodiments, the plurality of nucleic acids together encode a set of polypeptide binder moieties. In some embodiments, each nucleic acid comprises, in 5' to 3' or 3' to 5' order: a) a first constant sequence (e.g., an antibody germline sequence (e.g., IGHV / IGKV)), b) a first variant sequence (e.g., a CDR (e.g., CDR3) sequence) of at least 10 nucleotides in length, and c) a second constant sequence (e.g., an antibody germline sequence (e.g., IDHJ / IGKJ)).
[0061] In some embodiments, each nucleic acid further comprises one or more of: d) a stop codon (e.g., after the second constant sequence); e) a linker sequence; f) a third constant sequence (e.g., an antibody germline sequence (e.g., IGHV / IGKV)); g) a second variant sequence (e.g., a CDR (e.g., CDR3) sequence) at least 10 nucleotides in length; and h) a fourth constant sequence (e.g., an antibody germline sequence (e.g., IDHJ / IGKJ)).
[0062] The present disclosure provides, among other things, a library of phage particles, each phage particle comprising one or more nucleic acids described herein.
[0063] In some embodiments, the phage is selected from the group consisting of M13, T4, T7, lambda, and filamentous phage, hi some embodiments, the phage is M13.
[0064] The present disclosure provides sets of barcodes and binders. In some embodiments, each barcode is a peptide 1-100, 5-50, 8-25, 9-25, or 9-15 amino acids in length that specifically binds to a particular group of binders in the set. In some embodiments, each binder is a polypeptide that specifically binds to at least one barcode among the barcodes in the set.
[0065] In some embodiments, specific binding is observed when the binders are expressed on a phage that contacts the barcode. In some embodiments, each binder is expressed on a phage.
[0066] The present disclosure provides kits including a set of binders, each of which is a polypeptide that specifically binds to at least a particular peptide barcode within a collection of barcodes. In some embodiments, each binder is provided as a polypeptide, a nucleic acid encoding the polypeptide, or both. In some embodiments, one or more of the binders are provided as a phage particle or collection thereof engineered to express the binder. In some embodiments, one or more of the binders are provided as a nucleic acid within a phagemid vector or as an insert suitable for cloning into a phage vector.
[0067] In some embodiments, the kit further comprises information designating a peptide barcode for each binder, hi some embodiments, each binder has been determined to specifically bind to at least a particular peptide barcode in a collection of barcodes, each of which specifically binds to at least one binder in the set.
[0068] In some embodiments, the kit further comprises a set of instructions for sequencing one or more phage particles bound to one or more barcodes. In some embodiments, the kit further comprises a computer readable program for decoding the sequencing data. In some embodiments, the kit further comprises reagents for expressing binders on the phage particles.
[0069] In some embodiments, the kit comprises a nucleic acid encoding one or more barcodes. In some embodiments, the kit comprises a nucleic acid encoding one or more binders.
[0070] The present disclosure provides methods of pharmacokinetic screening. In some embodiments, the methods include injecting an animal with a set of barcoded therapeutic candidate proteins. In some embodiments, each barcoded therapeutic candidate protein comprises a specific peptide barcode. In some embodiments, the methods include obtaining a sample from the animal, purifying one or more barcoded therapeutic candidate proteins from the sample, contacting the sample with a set of binders (e.g., binding agents having the binders expressed thereon) comprising at least one specific binder specific for each barcode in the sample, and determining the relative amount of each binder present in the sample to determine the pharmacokinetic properties or biodistribution of each barcoded therapeutic candidate protein.
[0071] In some embodiments, the purified proteins can be a subset of the barcoded therapeutic candidate proteins that are injected into the animal.
[0072] In some embodiments, multiple samples can be obtained from an animal.
[0073] In some embodiments, the animal is a mammal. In some embodiments, the animal is a human. In some embodiments, the animal is genetically modified to express a barcoded payload (e.g., a barcoded protein).
[0074] In some embodiments, the animal is a model for a disease, disorder, or condition, hi some embodiments, the disease, disorder, or condition is cancer, an autoimmune disease, a neurodegenerative disease, or a pathogenic (e.g., viral / bacterial) disease, disorder, or condition.
[0075] In some embodiments, the determining step comprises (i) sequencing nucleic acids from binder-expressing binders; (ii) decoding the relative amount of each barcode present, thereby determining the relative amount of each therapeutic candidate protein; and / or (iii) performing one or more of FACS, or MACS (magnetic activated cell sorting), or affinity-based purification.
[0076] In some embodiments, the determining step comprises quantifying the number of binders that bind to the barcoded payload (e.g., the barcoded protein (e.g., the barcoded therapeutic protein)). In some embodiments, the quantifying is performed by decoding the nucleotide sequence of each binder that binds to the barcoded payload (e.g., the barcoded protein).
[0077] In some embodiments, some nucleotide sequences provide a measure of a payload (e.g., a target protein) in a population of barcoded payloads (e.g., barcoded proteins).
[0078] In some embodiments, the step of injecting comprises administering the barcoded payload (e.g., a barcoded protein, a barcoded therapeutic candidate protein, etc.) orally or intravenously. In some embodiments, the barcoded payload (e.g., a barcoded protein) is injected (e.g., delivered) by viral delivery or mRNA delivery.
[0079] The present disclosure provides a method for characterizing a collection of peptide barcodes, the method including providing (i) a library of phage particles, each phage particle designed to express a polypeptide binder, each binder binding to one or more peptide barcodes; (ii) a collection of peptide barcodes; contacting each phage particle with each barcode to form bound phage-barcode particles; determining the amount of binding between each phage particle and the barcode; and identifying phage-barcode pairs that specifically bind to each other between the barcodes in the collection and the phages in the library.
[0080] The present disclosure provides a method for characterizing a collection of peptide barcodes, the method including providing (i) a set of binders, each binder being a polypeptide that binds to one or more peptide barcodes, and (ii) a collection of peptide barcodes; contacting each binder with each barcode to form a bound binder-barcode particle; determining the relative amount of binding between each polypeptide binder and the peptide barcode; and identifying binder-barcode pairs between barcodes in the collection and binders in the set that specifically bind to each other.
[0081] The present disclosure provides a database of amino acid or encoding nucleic acid sequences for a collection of peptide barcodes, the database being embodied in a computer-readable format. In some embodiments, each barcode sequence has a length ranging from 1 to 100, 5 to 50, 8 to 25, 9 to 25, or 9 to 15 amino acids. In some embodiments, each barcode sequence has been determined to specifically bind to one or more polypeptide binders in a set of binders, each of which specifically binds to one or more of the barcodes in the collection.
[0082] In some embodiments, the binding pattern of one or more polypeptide binders to the barcode is used to identify the peptide barcode.
[0083] The present disclosure provides a database of amino acid or encoding nucleic acid sequences for a set of polypeptide binders. In some embodiments, the database is embodied in a computer-readable format. In some embodiments, each binder sequence has a length in the range of 10 to 400 amino acids. In some embodiments, each binder sequence has been determined to specifically bind to one or more peptide barcodes in a collection of barcodes, each of which specifically binds to one or more of the binders in the set.
[0084] The present disclosure provides, among other things, a database of amino acid or coding nucleic acid sequences for a set of barcode-binder associations embodied in a computer-readable format. In some embodiments, each barcode is a peptide 1-100, 5-50, 8-25, 9-25, or 9-15 amino acids in length. In some embodiments, each binder is a polypeptide that specifically binds to one or more of the barcodes in the set.
[0085] The present disclosure provides a set of barcode-binder association designations embodied in a computer-readable format. In some embodiments, each barcode is a peptide between 1 and 100, 5 and 50, 8 and 25, 9 and 25, or 9 and 15 amino acids in length. In some embodiments, each binder is a polypeptide that specifically binds to one or more of the barcodes in the set.
[0086] In some embodiments, specific binding is observed when the binder is expressed on a phage particle that is then contacted with the barcode.
[0087] The present disclosure provides methods of treatment using the techniques described herein. In some embodiments, the method includes administering a therapeutic protein that has been determined to satisfy the assessment. In some embodiments, satisfying the assessment may be performed by a process including: a) subjecting a population of barcoded proteins to the assessment; b) separating members of the population that satisfy the assessment from those that do not to identify either a positive population or a negative population, or both; c) contacting the positive population or the negative population, or each population separately from the other, with a set of binders comprising at least one specific binder specific for each barcode in the population; d) determining which binders bind to the separated members, thereby determining which barcoded proteins are present in the contacted population(s); and e) identifying the therapeutic protein from the barcoded proteins determined to be present in the contacted population(s).
[0088] The present disclosure provides methods of treatment that include administering a therapeutic protein determined to satisfy an assessment by a process that includes: a) separately contacting a set of binders with a first population, a second population, or each of the first and second populations of barcoded proteins; b) determining which binders of the set bind to members of the first population, the second population, or both, thereby determining which barcoded proteins are present in the contacted population(s); and c) identifying the therapeutic protein from the barcoded proteins determined to be present in the contacted population(s). In some embodiments, each binder specifically binds to one or more barcodes relative to other barcodes. In some embodiments, the set of binders collectively includes binders specific for each barcode in the first and second populations. In some embodiments, the first and second populations are separated from each other based on their performance in the assessment.
[0089] These and other aspects encompassed by the present disclosure are described in more detail below and in the claims. [Brief explanation of the drawings]
[0090] [Figure 1A] 1 is a schematic diagram of a barcoded payload described herein, according to an exemplary embodiment. It shows the barcoded payload and the corresponding DNA encoding the barcoded payload. LN refers to the "linker N-terminus" and LC refers to the "linker C-terminus." In some embodiments, the LN and LC sequences are constant and encode the amino acids that connect the payload to the barcode. In some embodiments, the LN and LC sequences are constant and are nucleic acid sequences used for modular cloning of barcodes with different payloads. In some embodiments, the LN and LC sequences are flanked by type IIS restriction site sequences. [Figure 1B] 1 is a schematic diagram of a barcoded payload described herein, according to an exemplary embodiment. It shows the nucleic acid sequences encoding the barcode and / or barcoded payload. LN refers to the "linker N-terminus" and LC refers to the "linker C-terminus." In some embodiments, the LN and LC sequences are constant and encode amino acids that connect the payload to the barcode. In some embodiments, the LN and LC sequences are constant and are nucleic acid sequences used for modular cloning of barcodes with different payloads. In some embodiments, the LN and LC sequences are flanked by type IIS restriction site sequences. [Figure 2]1 is a schematic diagram of a method for detecting and / or quantifying and / or characterizing payloads (e.g., proteins) in a pool using barcodes and binders described herein, according to an exemplary embodiment. A library of barcoded payload proteins is contacted with a library of binders containing identifying DNA. A wash step is applied to remove binders that do not associate (e.g., link (e.g., form strong links)) with any of the barcoded payload proteins, leaving only binders that associate with the barcode. After washing, a DNA sequencing process is applied to the associated binders. In some embodiments, sequencing can be performed using next-generation sequencing (NGS) (e.g., operated by an Illumina sequencer). The relative abundance of DNA sequences is reported as a computer file (e.g., .fastq data). A computer algorithm is applied to the .fastq data combined with prior biophysical characterization of the binders to infer the abundance of each of the barcoded payload proteins in the pool. [Figure 3A]
[0023] Figure 1 is a schematic diagram for capturing a barcode described herein so that it can be contacted by a binding agent described herein, according to an exemplary embodiment. This shows a capture scaffold that can have a barcode associated with it (e.g., immobilized on its surface), and a binding agent (e.g., a phage with a binder expressed on its surface (e.g., a phage with binder DNA within the phage)) contacted to characterize the biophysical interaction. In some embodiments, the biophysical characterization is a measure of the dissociation constant (Kd) between the binding agent and the protein barcode. [Figure 3B]
[0023] Figure 1 is a schematic diagram of an exemplary embodiment for capturing a barcode described herein so that it can be contacted by a binder described herein. This is a schematic diagram of the barcode-binder platform described herein, according to an exemplary embodiment. The schematic shows magnetic beads with bead-binding domains conjugated to a universally tagged (e.g., HALO, Chitin BD, Avitag (Strep), etc.) barcoded payload. To detect the captured barcoded payload, a binder with known affinity for the barcode (e.g., a phage expressing a binder on its surface (e.g., a phage with binder DNA / lib)) is allowed to bind to the immobilized payload. The DNA in the phage encoding the binder is then amplified and subjected to NGS to detect the payload. [Figure 3C]
[0023] Figure 1 is a schematic diagram of an exemplary embodiment for capturing a barcode described herein so that it can be contacted by a binder described herein. This is a schematic diagram of the barcode-binder platform described herein, according to an exemplary embodiment. The schematic shows magnetic beads with Fc / Protein A conjugated barcoded payloads. To detect the captured barcoded payload, a binder with known affinity for the barcode (e.g., a phage expressing a binder on its surface (e.g., a phage with binder DNA / lib)) is allowed to bind to the immobilized payload. The DNA in the phage encoding the binder is then amplified and subjected to NGS to detect the payload. [Figure 4]
[0023] Figure 1 is a schematic diagram of a method for learning a barcode fingerprint for a given barcode described herein, according to an exemplary embodiment. Peptide barcodes displayed on a capture scaffold are contacted with a library of binders containing identified DNA. A wash step is applied to remove binders that do not associate with (e.g., do not ligate (e.g., do not form strong bonds)) any of the barcodes, while leaving only those that associate with the barcode. After washing, a DNA sequencing process is applied to the associated binders. In some embodiments, sequencing can be performed using next-generation sequencing (NGS) (e.g., operated by an Illumina sequencer). The relative abundance of DNA sequences is reported as a computer file (e.g., in .fastq format). A computer algorithm is applied to the .fastq data to compute the barcode fingerprint, which is a vector of relative counts of members of the binder library. The method for learning a barcode fingerprint can be repeated for any barcode to identify a unique fingerprint. In some embodiments, steps 1-4 of Figure 4 can be repeated for barcodes with existing fingerprints or for new barcodes, each time starting with a focused binder library to improve the fingerprint. In some embodiments, the focused binder library is generated by oligonucleotide library synthesis. [Figure 5]
[0023] Figure 1 is a schematic diagram of a method for using a fingerprint matrix of a set of barcodes to determine the relative abundance of a mixture of barcodes, according to an exemplary embodiment. The set of barcodes, whose individual fingerprints have been determined, are combined together in known ratios and displayed (e.g., on a scaffold) for subsequent contact with a binder library. The binder library is contacted with the set of barcodes, and nonspecific binders are washed away. Specific binders are quantified by NGS and reported as a mixture measurement computer file (e.g., in .fastq format). This data is provided to a computer algorithm that uses the mixture measurements to learn the relative scaling of reads to the original fingerprints and assembles the scaled fingerprints together into a scaled matrix. This scaled fingerprint matrix can then be used in further applications, where the computer algorithm quantifies the relative abundance of barcoded payloads using these barcodes applied to NGS reads from contacting a sample with the binder library. [Figure 6A] Figure 1 shows the results of quantifying a complex mixture of barcodes. Up to six barcodes were pooled and then measured using the decoding method described herein. A shows the actual relative proportion of a given barcode (left panel) and the measured relative proportion of a given barcode (right panel). Rows are individual experimental conditions, columns are barcodes, and colors are measurements (100% barcode = white, 0% barcode = black). [Figure 6B] Figure 1 shows the results of quantifying complex mixtures of barcodes. Up to six barcodes were pooled and then measured using the decoding method described herein. Figure 2 shows a plot of the measured concentration of a barcode against its actual concentration for all experiments compared across all barcodes. A 0.95 Pearson correlation was calculated between the measured and actual proportions across all experiments and mixtures. [Figure 6C]Figure 1 shows the results of quantifying a complex mixture of barcodes. Up to six barcodes were pooled and then measured using the decoding method described herein. Figure 2 shows a plot of NGS count values normalized to counts per million for each single barcode measurement, as well as the mixture used to predict the relative abundance of each barcode within the mixture. Rows are experiments; therefore, all values in a row are generated from a single .fastq file; columns are binders. Figure 3 discloses SEQ ID NOs: 8400-8413, respectively, in the order they appear. [Figure 7] Figure 1A shows a schematic representation of data obtained using the method and decoding method for payload proteins with barcodes contained within an internal region of the protein sequence (i.e., endogenous barcodes). The schematic shows the results of a synthetic pooled barcode measurement assay. Figure 1B shows that two barcoded payloads (BC1 and BC2) were combined together at various known concentrations in different wells of a 96-well plate. Each mixture was contacted with the same binder pool and decoded as described herein. Each mixture was quantified and then compared to the known value of barcoded payload. Figure 1B shows a schematic representation of data obtained using the method and decoding method for payload proteins with barcodes contained within an internal region of the protein sequence (i.e., endogenous barcodes). The schematic shows the results of a synthetic pooled barcode measurement assay. Figure 1B shows the relative observed ratios (X-axis) of each barcode correlated to the relative observed ratios (Y-axis) with a Pearson of 0.96. [Figure 8A]
[0023] Figure 1 shows a schematic diagram of a method for detecting protein payloads in serum using the barcode-binder platform described herein, according to an exemplary embodiment. The payload protein has a barcode contained within an internal region of the protein sequence (i.e., an intrinsic barcode). A indicates that barcoded therapeutic antibody agents of interest (barcoded mAbs) were mixed at known concentrations and then added to serum. The barcoded payloads were then purified, contacted with a binder, and subjected to decoding. [Figure 8B]
[0023] Figure 1 shows a schematic diagram of a method for detecting protein payloads in serum using the barcode-binder platform described herein, according to an exemplary embodiment. The payload protein has a barcode contained within an internal region of the protein sequence (i.e., an endogenous barcode). Figure 2 shows the relative observed barcoded antibody percentages (left) and the relative measured antibody percentages (right) for three experimental conditions, each with three replicates. Rows correspond to experimental conditions, columns correspond to barcodes, and the color of the heatmap cells is a measure of the percentage of barcoded antibody present. [Figure 8C] (A) shows a schematic diagram of a method for detecting protein payloads in serum using the barcode-binder platform described herein, according to an exemplary embodiment. The payload protein has a barcode contained within an internal region of the protein sequence (i.e., an intrinsic barcode). (B) shows a scatter plot of all data across all experimental conditions for all barcodes, with a Spearman correlation of 0.926 across all experimental measurements. [Figure 9A]
[0023] Figure 1 shows a schematic diagram of the experimental description provided in Examples 1, 2, and 8. Six unique barcodes (BC1, BC2, BC3, BC4, BC5, and BC6) were mixed in known proportions, contacted with a binder, and subjected to decoding as described herein. Two barcodes were retained as experimental negative controls, but prediction of these barcodes allowed for the determination of background prediction. [Figure 9B] Figure 1 shows data on the accuracy of the decoding procedure over a 10-fold concentration range for six unique barcodes. A shows a plot of the observed (input) and measured data obtained after decoding one mixture of known barcode concentrations. The input known concentrations (left bars) are shown next to the predicted / measured data (right bars) for each barcode across three replicates. [Figure 9C]Figure 1 shows data on the accuracy of the decoding procedure over a 10-fold concentration range for six unique barcodes. Figure 2 shows a plot of the observed (input) and measured data obtained after decoding for five different mixtures of known barcode concentrations (i.e., pools 1-5). The input known concentrations (left bars) are shown next to the predicted / measured data (right bars) for each barcode across three replicates. [Figure 10A]
[0023] Figure 1 shows the method and data for determining the absolute concentration of a single test barcode described herein. A shows a schematic diagram of the experiment. A single test barcode was assayed at several concentrations, and a "spike-in" barcode (i.e., reference barcode) was added to each assay mixture at a known concentration. The test barcodes at various concentrations were contacted with a binding agent and decoded as described herein. The prediction of the "spike-in" barcode was used to determine the absolute amount of the test barcode being measured. [Figure 10B] (B) shows the method and data for determining the absolute concentration of a single test barcode described herein. (C) shows a plot of the measured absolute amount of the test barcode (right bar) compared to the known input concentration of the test barcode (left bar) for each titration of the test barcode. The Y-axis is the logarithm of the test barcode concentration in nanograms per milliliter (ng / mL). [Figure 10C] FIG. 1 shows the method and data for determining the absolute concentration of a single test barcode described herein. C shows the results of determining the absolute concentration for six different barcodes. The plot shows the known input concentration (left bar) and the measured concentration (right bar) for the six different barcodes. [Figure 11]
[0023] Figure 1 shows a method for determining the relative abundance of two proteins after in vivo injection using the binder-barcode system described herein, according to an exemplary embodiment. The figure shows a graphic depiction of the experimental setup. In group 1 (top), mice were injected with one barcoded payload. In group 2 (middle), mice were injected with two barcoded payloads. In group 3 (bottom), mice were injected with no barcoded payload. For each of the three groups, serum samples were collected at 24 hours, and the barcoded payload(s) captured using the binders described herein were subjected to decoding. The measured barcoded payload concentration (right bar) is shown compared to the known input concentration (left bar) for each group. [Figure 12A] Figure A shows a graphical depiction of the experiment, showing the determination of 24 barcodes contained within a single mixture. Of the 24 total barcodes, the algorithm was able to predict, with 10 present at equal concentrations in the mixture. The remainder were excluded from the pool, but the prediction was computationally acceptable. Three separate pools covering all possible barcodes were measured in replicates. [Figure 12B] A shows the determination of 24 barcodes contained within a single mixture. B shows the prediction of the first pool. Input concentrations (left bar) and measured concentrations (right bar) are displayed. [Figure 12C] C shows the determination of 24 barcodes contained within a single mixture. C shows the prediction across all three pools. As in B, input concentrations are the left bars and measurements are the right bars. [Figure 12D]Figure 1 shows the determination of 24 barcodes contained within a single mixture. Figure 2 shows the barcode fingerprints of the 24 barcodes used to computationally determine the relative abundance of the barcodes in the three pools. Columns represent barcode fingerprints, and rows represent binder fingerprints. Figure 3 shows SEQ ID NOs: 8414, 8415, 8414, 8416, 8414, 8413, 8414, 8417, 8414, 8418, 8414, 8419, 8414, 8420-8425, 8422, 8426, 8427, 8426, 8428-8431, 8430, 8432, 8433. 3, 8432, 8434, 8432, 8435, 8432, 8430, 8432, 8436~8453, 8413, 8453, 8454, 8453, 8455~8475, 8474, 8476~8480, 8479, 8481~8484, 8483, 8484~8493, 8472 , 8494, 8472, 8495, 8472, 8496, 8472, 8497, 8472, 8498-8502, 8501, 8503-8505, 8504, 8506, 8504, 8507, 8504, 8508-8516, 8515, 8517, 8518, 8417, 8519, 8520, 8519, 8521-8532, 8403, 8533-8542, 8541, 8543-8545, 8544, 8546, 8544, 8547, 8544, 8548-8552, 8551, 8553, 8551, 8554-8562 are disclosed in the order shown. [Figure 12E]Figure 1 shows the determination of 24 barcodes contained within a single mixture. E shows binder counts from three pools used to mathematically determine pool percentages. Rows are binder counts, columns are pools, and cells are binder counts within a particular pool. E represents SEQ ID NOs: 8414, 8415, 8414, 8416, 8414, 8413, 8414, 8417, 8414, 8418, 8414, 8419, 8414, 8420-8425, 8422, 8426, 8427, 8426, 8428-8431, 8430, 8432, 8433 3, 8432, 8434, 8432, 8435, 8432, 8430, 8432, 8436~8453, 8413, 8453, 8454, 8453, 8455~8475, 8474, 8476~8480, 8479, 8481~8484, 8483, 8484~8493, 8472 , 8494, 8472, 8495, 8472, 8496, 8472, 8497, 8472, 8498-8502, 8501, 8503-8505, 8504, 8506, 8504, 8507, 8504, 8508-8516, 8515, 8517, 8518, 8417, 8519, 8520, 8519, 8521-8532, 8403, 8533-8542, 8541, 8543-8545, 8544, 8546, 8544, 8547, 8544, 8548-8552, 8551, 8553, 8551, 8554-8562 are disclosed in the order shown. [Figure 13A]
[0023] Figure 1 is a schematic diagram of a method for detecting and / or quantifying and / or characterizing 14 exemplary payloads (e.g., proteins) in a pool using the binder-barcode platform described herein. A library of barcoded payload proteins was contacted with a library of binders containing identifying DNA ("binder-barcode particles"). The binder-barcode particles were injected in vivo as a pooled library into wild-type (wt) BALB / c mice (n=3 / time point). Blood was collected from individual mice at 30 minutes, 6 hours, 24 hours, and 48 hours (n=3 / time point), and serum was extracted. The binder-barcode particles were captured and subjected to the decoding procedure described herein. [Figure 13B]Figure 13B shows plots illustrating the clearance of 14 exemplary binder-barcoded particles injected in vivo into wild-type (wt) BALB / c mice (n=3 / time point). Data were collected at 30 minutes, 6 hours, 24 hours, and 48 hours, as measured by the decoding procedure described herein. The Y-axis is normalized to 100% of the injected volume for each exemplary binder-barcoded particle. The plots shown in Figure 13B were measured simultaneously. Each plot includes exemplary binder-barcoded particles characterized as having a specific measurable phenotype. The left plot shows the clearance (% injected) of a clinical control with known characteristics. The middle plot shows the clearance (% injected) of an exemplary binder-barcoded particle characterized as having slow clearance characteristics. The right plot shows the clearance (% injected) of an exemplary binder-barcoded particle characterized as having fast clearance characteristics. [Figure 14A] Figure 1 shows a schematic diagram of a method for detecting and / or quantifying and / or characterizing 36 payloads (e.g., proteins) in a pool using the binder-barcode platform described herein. A library of barcoded payload proteins was contacted with a library of binders containing identifying DNA ("binder-barcode particles"). The binder-barcode particles were injected as a pooled library into tumor-bearing NSG mice previously implanted in vivo with two tumor cell lines ("Tumor 1" and "Tumor 2") (n = 2-4 / time point). Blood and tumor tissue were collected from individual mice at 30 minutes, 6 hours, 24 hours, and 48 hours (n = 3 / time point). Tissues were lysed using a standard lysis buffer, and serum was separated from the blood. The binder-barcode particles were captured and subjected to the decoding procedure described herein. [Figure 14B]
[0023] Figure 1 is a heat map of data collected from 36 exemplary binder-barcoded particles using the decoding procedure described herein. Rows identify each exemplary binder-barcoded particle tested in this example. Columns represent data for mice across serum, tumor 1, or tumor 2 time points. Color intensity indicates relative units of drug measured via the decoding procedure described herein. Color intensity represents a normalized readout of relative concentrations measured via next-generation sequencing (NGS). [Figure 14C] 14B shows plots of the binder-barcoded particles described by FIG. 14B using the decoding procedure described herein. Variation in properties was measured simultaneously. For example, binder-barcoded particle P14_A5 was rapidly cleared from serum with minimal accumulation in tumor 1 or tumor 2, while binder-barcoded particle P17_A10 was cleared more slowly and persisted in tumor 1 over time. [Figure 15A] A plot representing ELISA quantification of two groups of payloads (Group 1: protein payload without a barcode; Group 2: a pool of eight binder-barcode particles, where each particle contains the same protein payload used in Group 1 and each particle is barcoded with a different barcode) is shown. [Figure 15B] Quantification of group 2 using the decoding procedure described herein is shown. [Figure 15C] Shown is a comparison of half-life measurements for Groups 1 and 2, quantified using ELISA and the decoding procedure described herein, respectively. [Figure 16A] FIG. 1 is a schematic diagram of a method for detecting and / or quantifying and / or characterizing 35 payloads (e.g., proteins) in separate pools with different numbers of barcoded payloads at different concentrations using the binder-barcode platform described herein. [Figure 16B]Figure 16B shows a plot of measured barcode levels (arbitrary units) versus expected barcoded payload levels (ng) generated by sequencing 96 separate mixtures containing 10-35 barcoded proteins, each at a known concentration ranging from 1 pg to 1 μg. Each data point in Figure 16B represents a comparison of the known concentration of binder-barcode particles from one of the 96 separate mixtures to the concentration determined by the decoding procedure described herein. DETAILED DESCRIPTION OF THE INVENTION
[0091] definition About: When used herein in connection with a value, the term "about" refers to a value that is similar to the reference value in the context. Generally, a person skilled in the art who is familiar with the context will understand the reasonable degree of difference that "about" encompasses in the context. For example, in some embodiments, the term "about" may encompass a range of values that are within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1% or less of the reference value.
[0092] Affinity: As known in the art, "affinity" is a measure of the tightness with which two or more binding partners associate with one another. Those skilled in the art will be knowledgeable about various assays that can be used to assess affinity and will also recognize appropriate controls for such assays. In some embodiments, affinity is assessed in a quantitative assay. In some embodiments, affinity (e.g., of one binding partner at a time) is assessed across multiple concentrations. In some embodiments, affinity is assessed in the presence of one or more potentially competing entities (which may be present in a relevant (e.g., physiological) context). In some embodiments, affinity is assessed relative to a reference (e.g., having a known affinity above a certain threshold [see "positive control"] or having a known affinity below a certain threshold [see "negative control"]). In some embodiments, affinity may be assessed relative to a concurrent reference, and in some embodiments, affinity may be assessed relative to a background reference. Typically, when affinity is assessed relative to a reference, it is assessed under comparable conditions.
[0093] Agent: Generally, as used herein, the term "agent" is used to refer to an entity (e.g., a lipid, metal, nucleic acid, polypeptide, polysaccharide, small molecule, etc., or a complex, combination, mixture, or system thereof (e.g., a cell, tissue, organism)) or phenomenon (e.g., heat, an electric current or electric field, a magnetic force or magnetic field, etc.). Under appropriate circumstances, as will be clear from the context to one of skill in the art, the term may be used to refer to an entity that is or includes a cell or organism, or a fraction, extract, or component thereof. Alternatively or additionally, as will be clear from the context, the term may be used to refer to a natural product that exists in and / or is taken from nature. In some cases, again as will be clear from the context, the term may be used to refer to one or more artificial and / or non-natural entities that have been designed, engineered, and / or created by the act of the hand of man. In some embodiments, an agent may be utilized in isolated or pure form, and in some embodiments, an agent may be utilized in crude form. In some embodiments, potential agents may be provided as a collection or library that may be screened, for example, to identify or characterize active agents therein. In some cases, the term "agent" may refer to a compound or entity that is or includes a polymer, and in some cases, the term may refer to a compound or entity that includes one or more polymer moieties. In some embodiments, the term "agent" may refer to a compound or entity that is not a polymer and / or is substantially free of any polymer and / or one or more specific polymer moieties. In some embodiments, the term may refer to a compound or entity that lacks or is substantially free of any polymer moieties.
[0094] Amino acid: As used herein, in its broadest sense, refers to any compound and / or substance that can be incorporated into a polypeptide chain, for example, by the formation of one or more peptide bonds. In some embodiments, an amino acid has the general structure HN-C(H)(R)-COOH. In some embodiments, an amino acid is a naturally occurring amino acid. In some embodiments, an amino acid is an unnatural amino acid, in some embodiments, an amino acid is a D-amino acid, and in some embodiments, an amino acid is an L-amino acid. A "standard amino acid" refers to any of the 20 standard L-amino acids commonly found in naturally occurring peptides. A "non-standard amino acid" refers to any amino acid other than the standard amino acids, whether it is synthetically prepared or obtained from a natural source. In some embodiments, an amino acid, including the carboxy- and / or amino-terminal amino acids in a polypeptide, may contain structural modifications compared to the general structure above. For example, in some embodiments, an amino acid may be modified compared to the general structure by methylation, amidation, acetylation, pegylation, glycosylation, phosphorylation, and / or substitution (e.g., of the amino group, the carboxylic acid group, one or more protons, and / or the hydroxyl group). In some embodiments, such modifications may, for example, alter the circulating half-life of a polypeptide comprising the modified amino acid compared to one comprising the otherwise identical amino acid. In some embodiments, such modifications do not significantly alter the relevant activity of a polypeptide comprising the modified amino acid compared to one comprising the otherwise identical amino acid. As will be clear from the context, in some embodiments, the term "amino acid" may be used to refer to a free amino acid, and in some embodiments, the term may be used to refer to an amino acid residue of a polypeptide.
[0095] Animal: As used herein, refers to any member of the animal kingdom. In some embodiments, "animal" refers to humans of either sex and at any stage of development. In some embodiments, "animal" refers to non-human animals at any stage of development. In certain embodiments, the non-human animal is a mammal (e.g., a rodent, mouse, rat, rabbit, monkey, dog, cat, sheep, cow, primate, and / or pig). In some embodiments, animals include, but are not limited to, mammals, birds, reptiles, amphibians, fish, insects, and / or parasites. In some embodiments, the animal may be a transgenic animal, a genetically engineered animal, and / or a clone.
[0096] Antibody: As used herein, the term "antibody" refers to a polypeptide containing canonical immunoglobulin sequence elements sufficient to confer specific binding to a particular target antigen. As is known in the art, intact antibodies, as produced in nature, are approximately 150 kD tetrameric agents composed of two identical heavy chain polypeptides (about 50 kD each) and two identical light chain polypeptides (about 25 kD each) that associate with each other into what is commonly referred to as a "Y-shaped" structure. Each heavy chain consists of at least four domains (each about 110 amino acids long): an amino-terminal variable (VH) domain (located at the tip of the Y structure) followed by three constant domains: CH1, CH2, and a carboxy-terminal CH3 domain (located at the base of the stem of the Y). A short region known as the "switch" connects the heavy chain variable and constant regions. A "hinge" connects the CH2 and CH3 domains to the rest of the antibody. Two disulfide bonds in this hinge region connect two heavy chain polypeptides to each other in intact antibodies. Each light chain consists of two domains: an amino-terminal variable (VL) domain followed by a carboxy-terminal constant (CL) domain, which are separated from each other by another "switch." An intact antibody tetramer consists of two heavy-light chain dimers, in which the heavy and light chains are linked to each other by a single disulfide bond, and two other disulfide bonds connect the heavy chain hinge regions to each other, forming a tetramer. Naturally produced antibodies are typically glycosylated in the CH2 domain. Each domain in a natural antibody has a structure characterized by an "immunoglobulin fold" formed by two beta sheets (e.g., a three-, four-, or five-stranded sheet) packed against each other in a compressed antiparallel beta barrel. Each variable domain contains three hypervariable loops known as "complementarity determining regions" (CDR1, CDR2, and CDR3), and four somewhat invariant "framework" regions (FR1, FR2, FR3, and FR4).When a natural antibody folds, the FR regions form beta sheets to provide the structural framework for the domain, and the CDR loop regions of both the heavy and light chains join in three-dimensional space to create a single hypervariable antigen-binding site located at the tip of a Y-structure. The Fc region of a naturally occurring antibody binds to elements of the complement system and also to receptors on effector cells, including, for example, effector cells that mediate cytotoxicity. As is known in the art, the affinity and / or other binding properties of the Fc region for the Fc receptor can be modulated through glycosylation or other modifications. In some embodiments, antibodies produced and / or utilized in accordance with the present invention comprise a glycosylated Fc domain, an Fc domain that has been modified or engineered, such as by glycosylation. For purposes of the present invention, in certain embodiments, any polypeptide or complex of polypeptides that comprises a sufficient immunoglobulin domain sequence as found in a natural antibody may be referred to and / or used as an "antibody," regardless of whether such polypeptide is naturally produced (e.g., produced by an organism in response to an antigen) or produced by recombinant engineering, chemical synthesis, or other artificial systems or methodologies. In some embodiments, an antibody is polyclonal; in some embodiments, an antibody is monoclonal. In some embodiments, an antibody has constant region sequences characteristic of mouse, rabbit, primate, or human antibodies. In some embodiments, antibody sequence elements are humanized, primatized, chimeric, etc., as known in the art. Furthermore, as used herein, the term "antibody" may refer, in appropriate embodiments (unless otherwise stated or apparent from the context), to any of the constructs or formats known or developed in the art for utilizing the structural and functional characteristics of antibodies in alternative presentations.For example, in embodiments, antibodies utilized in accordance with the present invention include, but are not limited to, intact IgA, IgG, IgE, or IgM antibodies, bispecific or multispecific antibodies (e.g., Zybodies®, etc.); antibody fragments, such as Fab fragments, Fab' fragments, F(ab')2 fragments, Fd' fragments, Fd fragments, and isolated CDRs or sets thereof; single chain Fv, polypeptide-Fc fusions; single domain antibodies (e.g., shark single domain antibodies such as IgNAR or fragments thereof); camelid antibodies; masked antibodies (e.g., Probodies®); Small Modular The format is selected from ImmunoPharmaceuticals ("SMIP™"); single-chain diabodies or tandem diabodies (TandAb®); VHH; Anticalins®; Nanobodies® minibodies; BiTE®; ankyrin repeat proteins or DARPIN®; Avimers®; DART; TCR-like antibodies; Adnectins®; Affilins®; Trans-bodies®; Affibodies®; TrimerX®; MicroProteins; Fynomers®; Centyrins®, and KALBITOR®. In some embodiments, the antibody may lack a covalent modification (e.g., glycan attachment) that it may have when naturally produced. In some embodiments, the antibody may include a covalent modification (e.g., glycan attachment), a payload (e.g., a detectable moiety, a therapeutic moiety, a catalytic moiety, etc.), or other pendant group (e.g., polyethylene glycol, etc.).
[0097] Antibody agent: As used herein, the term "antibody agent" refers to an agent that specifically binds to a particular antigen. In some embodiments, the term encompasses any polypeptide or polypeptide complex that contains sufficient immunoglobulin structural elements to confer specific binding. Exemplary antibody agents include, but are not limited to, monoclonal or polyclonal antibodies. In some embodiments, an antibody agent may contain one or more constant region sequences characteristic of murine, rabbit, primate, or human antibodies. In some embodiments, an antibody agent may contain one or more sequence elements that are humanized, primatized, chimeric, etc., as known in the art. In many embodiments, the term "antibody agent" is used to refer to one or more of the constructs or formats known or developed in the art for utilizing the structural and functional characteristics of antibodies in alternative presentations.For example, in embodiments, antibody agents utilized in accordance with the present invention include, but are not limited to, intact IgA, IgG, IgE, or IgM antibodies, bispecific or multispecific antibodies (e.g., Zybodies®, etc.); antibody fragments, such as Fab fragments, Fab' fragments, F(ab')2 fragments, Fd' fragments, Fd fragments, and isolated CDRs or sets thereof; single chain Fv, polypeptide-Fc fusions; single domain antibodies (e.g., shark single domain antibodies such as IgNAR or fragments thereof); camelid antibodies; masked antibodies (e.g., Probodies®); Small Modular The format is selected from ImmunoPharmaceuticals ("SMIP™"); single-chain diabodies or tandem diabodies (TandAb®); VHH; Anticalins®; Nanobodies® minibodies; BiTE®; ankyrin repeat proteins or DARPIN®; Avimers®; DART; TCR-like antibodies; Adnectins®; Affilins®; Trans-bodies®; Affibodies®; TrimerX®; MicroProteins; Fynomers®; Centyrins®, and KALBITOR®. In some embodiments, the antibody may lack a covalent modification (e.g., glycan attachment) that it may have when naturally produced. In some embodiments, the antibody may include a covalent modification (e.g., glycan attachment), a payload (e.g., a detectable moiety, a therapeutic moiety, a catalytic moiety, etc.), or other pendant group (e.g., polyethylene glycol, etc.). In many embodiments, an antibody agent is or comprises a polypeptide whose amino acid sequence comprises one or more structural elements recognized by those skilled in the art as complementarity determining regions (CDRs).In some embodiments, an antibody agent is a polypeptide comprising at least one CDR (e.g., at least one heavy chain CDR and / or at least one light chain CDR) whose amino acid sequence is substantially identical to that found in a reference antibody, or comprises a polypeptide comprising at least one CDR (e.g., at least one heavy chain CDR and / or at least one light chain CDR) whose amino acid sequence is substantially identical to that found in a reference antibody. In some embodiments, the included CDR is substantially identical to the reference CDR in that it is sequence identical or contains one to five amino acid substitutions compared to the reference CDR. In some embodiments, the included CDR is substantially identical to the reference CDR in that it exhibits at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the reference CDR. In some embodiments, the included CDRs are substantially identical to the reference CDRs in that they exhibit at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the reference CDRs. In some embodiments, the included CDRs are substantially identical to the reference CDRs in that at least one amino acid within the included CDRs is deleted, added, or substituted when compared to the reference CDRs, but the included CDRs have an amino acid sequence that is otherwise identical to that of the reference CDRs. In some embodiments, the included CDRs are substantially identical to the reference CDRs in that one to five amino acids within the included CDRs are deleted, added, or substituted when compared to the reference CDRs, but the included CDRs have an amino acid sequence that is otherwise identical to that of the reference CDRs. In some embodiments, the included CDRs are substantially identical to the reference CDRs in that at least one amino acid within the included CDRs is substituted when compared to the reference CDRs, but the included CDRs have an amino acid sequence that is otherwise identical to that of the reference CDRs.In some embodiments, the included CDR is substantially identical to the reference CDR in that one to five amino acids within the included CDR are deleted, added, or substituted when compared to the reference CDR, but the included CDR has an amino acid sequence that is otherwise identical to the reference CDR. In some embodiments, the antibody agent is a polypeptide whose amino acid sequence comprises structural elements recognized by those skilled in the art as an immunoglobulin variable domain, or comprises a polypeptide whose amino acid sequence comprises structural elements recognized by those skilled in the art as an immunoglobulin variable domain. In some embodiments, the antibody agent is a polypeptide protein having a binding domain that is homologous or largely homologous to an immunoglobulin binding domain.
[0098] Associated: As used herein, two events or entities are "associated" with one another when the presence, level, degree, type, and / or form of one correlates with that of the other. For example, a particular entity (e.g., a polypeptide, gene signature, metabolite, microorganism, etc.) is considered to be associated with a particular disease, disorder, or condition if its presence, level, and / or form correlates with the occurrence and / or susceptibility of the disease, disorder, or condition (e.g., across a relevant population). In some embodiments, two or more entities are physically "associated" with one another when they interact directly or indirectly to be in physical proximity and / or remain in close proximity to one another. In some embodiments, two or more entities that are physically associated with one another are covalently linked to one another; in some embodiments, two or more entities that are physically associated with one another are not covalently linked to one another but are non-covalently associated, e.g., by hydrogen bonding, van der Waals interactions, hydrophobic interactions, magnetism, and combinations thereof.
[0099] Barcode: As used herein, the term "barcode" refers to a peptide sequence that associates with a binder of known specificity and affinity. In some embodiments, the barcode binds to a specific antibody agent. In some embodiments, the barcode may be contained within a particular payload of interest. In some embodiments, the barcode may be at the end of a particular payload of interest. In some embodiments, the barcode may be synthetic. In some embodiments, the barcode may be designed. For example, a barcode sequence may be ordered as a DNA polynucleotide and cloned into a payload of interest using methods of molecular cloning known to those of skill in the art.
[0100] Binder: As used herein, the term "binder" or "binder moiety" refers to a polypeptide sequence that associates with a barcode with known specificity and affinity. In some embodiments, the binder is or includes an antibody agent. In some embodiments, the binder is expressed on the surface of the binding agent. In some embodiments, the binder can bind to one or more barcodes.
[0101] Binding: As used herein, the terms "binding" or "binding" are understood to typically refer to a non-covalent association between two or more entities. "Direct" binding involves physical contact between the entities or moieties, while indirect binding involves a physical interaction due to physical contact through one or more intermediate entities. Binding between two or more entities can typically be assessed in any of a variety of contexts, including when the interacting entities or moieties are studied alone or in the context of a more complex system (e.g., covalently or otherwise associated with a carrier entity and / or in a biological system or cell).
[0102] Binder: In general, the term "binder" is used herein to refer to any entity that binds to a target of interest (e.g., a barcode, a barcoded target, etc.) as described herein. In many embodiments, a binder of interest is one that specifically binds to its target in that it distinguishes the target from other potential binding partners in the context of a particular interaction. Generally, a binder may be or include an entity of any chemical class (e.g., polymer, non-polymer, small molecule, polypeptide, carbohydrate, lipid, nucleic acid, etc.) or biological class (e.g., bacteria, phage, ribosome, mRNA, DNA, etc.). In some embodiments, a binder is a single chemical entity. In some embodiments, a binder is a complex of two or more distinct chemical entities that associate with each other under relevant conditions through non-covalent interactions. For example, one of skill in the art will understand that in some embodiments, a binding agent can include a "general" binding moiety (e.g., one of biotin / avidin / streptavidin and / or a class-specific antibody) as well as a "specific" binding moiety (e.g., an antibody or aptamer with a specific molecular target) linked to the general binding moiety partner. In some embodiments, such an approach can allow for modular assembly of multiple binding agents through linkage of different specific binding moieties to the same general binding moiety partner. In some embodiments, a binding agent is or comprises a phage. In some embodiments, a binding agent is or comprises a polypeptide (e.g., including an antibody or antibody fragment). In some embodiments, a binding agent is or comprises a small molecule. In some embodiments, a binding agent is or comprises a nucleic acid. In some embodiments, a binding agent is or comprises an aptamer. In some embodiments, a binding agent is a polymer, and in some embodiments, a binding agent is not a polymer. In some embodiments, a binding agent is non-polymeric in that it lacks a polymeric portion. In some embodiments, a binding agent is or comprises a carbohydrate. In some embodiments, the binding agent is or comprises a lectin.In some embodiments, the binding agent is or comprises a peptidomimetic. In some embodiments, the binding agent is or comprises a scaffold protein. In some embodiments, the binding agent is or comprises a mimeotope. In some embodiments, the binding agent is or comprises a staple peptide. In certain embodiments, the binding agent is or comprises a nucleic acid, such as DNA or RNA.
[0103] Biological sample: As used herein, the term "biological sample" typically refers to a sample obtained or derived from a biological source of interest (e.g., a tissue or organism or cell culture) as described herein. In some embodiments, the source of interest includes an organism, such as an animal or a human. In some embodiments, the biological sample is or includes a biological tissue or fluid. In some embodiments, the biological sample can be or include bone marrow, blood, blood cells, ascites, tissue or fine needle biopsy samples, cell-containing body fluids, suspended nucleic acids, sputum, saliva, urine, cerebrospinal fluid, peritoneal fluid, pleural effusion, feces, lymph, gynecological fluids, skin swabs, vaginal swabs, oral swabs, nasal swabs, washings or lavage fluids such as ductal lavage or bronchoalveolar lavage, aspirates, scrapings, bone marrow specimens, tissue biopsy specimens, surgical specimens, feces, other body fluids, secretions, and / or excretions, and / or cells therefrom, and / or the like. In some embodiments, a biological sample is or comprises cells obtained from an individual. In some embodiments, the obtained cells are or comprise cells derived from the individual from whom the sample is obtained. In some embodiments, a sample is a "primary sample" obtained directly from the source of interest by any suitable means. For example, in some embodiments, a primary biological sample is obtained by a method selected from the group consisting of biopsy (e.g., fine needle aspirate or tissue biopsy), surgery, collection of a bodily fluid (e.g., blood, lymph, stool, etc.), and the like. In some embodiments, as will be clear from the context, the term "sample" refers to a preparation obtained by processing a primary sample (e.g., by removing one or more components and / or adding one or more agents), such as filtration using a semipermeable membrane. Such a "processed sample" can include, for example, nucleic acids or proteins extracted from the sample or obtained by subjecting the primary sample to techniques such as amplification or reverse transcription of mRNA, isolation and / or purification of certain components, etc.
[0104] CDR: As used herein, "CDR" refers to a complementarity-determining region within an antibody variable region. There are three CDRs in each of the heavy and light chain variable regions, designated CDR1, CDR2, and CDR3 for each variable region. A "set of CDRs" or "CDR set" refers to a group of three or six CDRs that occur in either a single variable region capable of binding antigen or the CDRs of cognate heavy and light chain variable regions capable of binding antigen. Certain systems have been established in the art for defining CDR boundaries (e.g., Kabat, Chothia, etc.), and those of skill in the art will appreciate the differences between these systems and will understand CDR boundaries to the extent necessary to understand and practice the claimed invention.
[0105] Comparable: As used herein, the term "comparable" refers to a set of two or more agents, entities, circumstances, or conditions that may not be identical to one another, but that are sufficiently similar to permit a comparison between them, where one of ordinary skill in the art would understand that conclusions can be reasonably drawn based on observed differences or similarities. In some embodiments, comparable sets of conditions, circumstances, individuals, or populations are characterized by multiple substantially identical characteristics and one or a few altered characteristics. One of ordinary skill in the art will understand what level of identity is required for two or more such sets of agents, entities, circumstances, conditions, etc. to be considered comparable in any given situation in context. For example, one of ordinary skill in the art will understand that sets of circumstances, individuals, or populations are comparable to one another if they are characterized by a sufficient number and type of substantially identical characteristics to warrant a reasonable conclusion that differences in results obtained or phenomena observed under or with different sets of circumstances, individuals, or populations are attributable to or indicative of differences in these various characteristics.
[0106] Comprising: Compositions or methods described herein as "comprising" one or more named elements or steps are open-ended, meaning that the named elements or steps are required, but that other elements or steps may be added within the scope of the composition or method. To avoid redundancy, it should also be understood that any composition or method described as "comprising" (or "comprises") one or more named elements or steps also represents a corresponding, more limited composition or method "consisting essentially of" (or "consists essentially of") the same named elements or steps, meaning that the composition or method includes the named essential elements or steps, and may include additional elements or steps that do not materially affect the basic and novel property(ies) of the composition or method. It should also be understood that any composition or method described herein as "comprising" or "consisting essentially of" one or more named elements or steps also represents a corresponding, more limited, closed-ended composition or method "consisting of" (or "consists of") the named elements or steps, excluding any other elements or steps not named. In any composition or method disclosed herein, known or disclosed equivalents of any named essential element or step may be substituted for that element or step.
[0107] Decoding: As used herein, the term "decoding" refers to the laboratory and / or bioinformatics process of identifying and quantifying unique sets of amino acids within a barcode. In some embodiments, such identification and quantification is achieved by using nucleic acid (e.g., DNA) counts from a sequencing experiment and measuring binder counts. In some embodiments, previously measured fingerprints (e.g., binder fingerprints or barcode fingerprints) are used to determine relationships between an unknown barcode mixture, which is decoded, for example, by comparing the binder counts of a previously known mixture across binders with known and varying affinities for several barcodes within the pool.
[0108] Designed: As used herein, the term "designed" refers to (i) an agent whose structure is selected or chosen by the hand of man, (ii) an agent produced by a process requiring human intervention, and / or (iii) an agent that is distinct from natural substances and other known agents.
[0109] Determining: Many methodologies described herein include a "determining" step. Those skilled in the art will understand, upon reading this specification, that such "determining" may utilize or be accomplished by use of any of a variety of techniques available to those skilled in the art, including, for example, certain techniques explicitly mentioned herein. In some embodiments, determining involves manipulation of the physical sample. In some embodiments, determining involves consideration and / or manipulation of data or information, for example, with the aid of a computer or other processing unit adapted to perform the relevant analysis. In some embodiments, determining involves receiving relevant information and / or materials from a source. In some embodiments, determining involves comparing one or more characteristics of the sample or entity to a comparable reference.
[0110] Engineered: In general, the term "engineered" refers to aspects that have been manipulated by the hand of man. For example, in some embodiments, a small molecule can be considered engineered if its structure and / or production are designed and / or implemented by the hand of man. Similarly, in some embodiments, a polynucleotide can be considered "engineered" if two or more sequences that are not inherently linked together in that order are directly linked to each other within the engineered polynucleotide. For example, in some embodiments of the invention, an engineered polynucleotide includes a regulatory sequence that is naturally found in operative association with a first sequence (e.g., a coding sequence), but is not operatively associated with a second sequence (e.g., a coding sequence), and is ligated by the hand of man into operative association with the second sequence. Comparably, a cell or organism is considered "engineered" if it has been manipulated so that its genetic information is altered (e.g., new genetic material not previously present has been introduced, e.g., by transformation, mating, somatic cell hybridization, transfection, transduction, or other mechanisms, or previously present genetic material has been altered or removed, e.g., by substitution or deletion mutations or by mating protocols). As is common practice and understood by those of skill in the art, the expression products of engineered polynucleotides, and / or the progeny of engineered polynucleotides or cells, will typically still be referred to as "engineered," regardless of the actual manipulation of the prior entity.
[0111] Expression: As used herein, "expression" of a nucleic acid sequence refers to one or more of the following events: (1) generation of an RNA template from a DNA sequence (e.g., by transcription), (2) processing of the RNA transcript (e.g., by splicing, editing, 5' capping, and / or 3' end formation), (3) translation of the RNA into a polypeptide or protein, and / or (4) post-translational modification of the polypeptide or protein.
[0112] Fingerprint: As used herein, the term "fingerprint" refers to a count of one or more unknown agents to which a known agent may bind or associate. In some embodiments, a fingerprint may be for a known barcode or barcode mixture. In some embodiments, a fingerprint may be for a known binder or binder mixture. For example, in some embodiments, a fingerprint (e.g., a barcode fingerprint) refers to the count of one or more binders (e.g., determined by sequencing analysis) that can specifically bind to a known barcode or barcode mixture. That is, in some embodiments, a fingerprint of a barcode refers to the count of one or more binders, some of which may have a high affinity for the barcode and some of which may have a low affinity for the barcode. In some embodiments, a fingerprint may be used in a decoding process, which is used to determine the relative or absolute abundance of a given barcode within a pool of barcodes. As will be understood by one of skill in the art, a fingerprint may be determined for a known barcode or barcode mixture, or a known binder or binder mixture. For example, in some embodiments, a fingerprint (e.g., a binder fingerprint) may refer to the count of one or more barcodes (e.g., determined by sequencing analysis) that specifically bind to a known binder or mixture of binders. That is, in some embodiments, a binder fingerprint refers to the count of one or more barcodes, some of which may have a high affinity for the binder and some of which may have a low affinity for the binder. Thus, fingerprints may also, in some embodiments, be used in the decoding process to determine the relative or absolute abundance of a given binder within a pool of binders.
[0113] Fragment: A "fragment" of a substance or entity as described herein comprises a distinct portion of the whole, but has a structure that lacks one or more portions found in the whole. In some embodiments, the fragment consists of such a distinct portion. In some embodiments, the fragment consists of or comprises characteristic structural elements or portions found in the whole. In some embodiments, a polymer fragment comprises or consists of at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500 or more monomer units (e.g., residues) as found throughout the polymer. In some embodiments, a polymer fragment comprises or consists of at least about 5%, 10%, 15%, 20%, 25%, 30%, 25%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more of the monomer units (e.g., residues) found in the whole polymer. The whole substance or entity may, in some embodiments, be referred to as the "parent" of the whole.
[0114] Human: In some embodiments, the human is an embryo, fetus, infant, child, teenager, adult, or elderly.
[0115] "Improve," "Increase," "Inhibit," or "Reduce": As used herein, the terms "improve," "increase," "inhibit," "reduce," or their grammatical equivalents refer to a value relative to a baseline or other reference measurement. In some embodiments, a suitable reference measurement may be or include a measurement in a particular system (e.g., in a single individual) under otherwise comparable conditions in the absence (e.g., before and / or after) of a particular agent or treatment, or in the presence of a suitable comparable reference agent. In some embodiments, a suitable reference measurement may be or include a measurement in a comparable system known or expected to respond in a particular manner in the presence of the relevant agent or treatment.
[0116] In vitro: As used herein, the term "in vitro" refers to events that take place in an artificial environment, e.g., in a test tube or reaction vessel, in cell culture, etc., rather than within a multicellular organism.
[0117] In vivo: As used herein, refers to events that occur within multicellular organisms, such as humans and non-human animals. In the context of cell-based systems, the term can be used to refer to events that occur within living cells (as opposed to, for example, in vitro systems).
[0118] Library: The term "library," as used herein, refers to a mixture of one or more distinct molecules. In some embodiments, all elements of a library share one or more common components. In some embodiments, all elements of a library do not share a common component. In some embodiments, one or more elements of a library are distinguished by one or more unique components. In some embodiments, as may be clear from the context, a library may refer to a mixture of binders. In some embodiments, a library may be a phage library. In some embodiments, for example, a phage library may consist of phage with distinct binders displayed (e.g., on their surface) and encapsulating DNA encoding the binders within the phage. In some embodiments, a library may refer to a mixture of barcoded payload proteins. In some embodiments, a library may refer to a mixture of barcodes (e.g., peptide barcodes).
[0119] Linker: As used herein, the term "linker" refers to a portion of a multi-element agent that connects different elements to one another. For example, those skilled in the art will understand that polypeptides whose structure includes two or more functional or organizational domains often contain a stretch of amino acids between them that connects such domains. In some embodiments, polypeptides containing linker elements have an overall structure of the general form S1-L-S2, where S1 and S2, which may be the same or different, represent two domains associated with each other by the linker. In some embodiments, the polypeptide linker is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, or more amino acids in length. In some embodiments, the linker is characterized by not tending to adopt a rigid three-dimensional structure, but rather providing flexibility to the polypeptide. A variety of different linker elements that can be appropriately used in engineering polypeptides (e.g., fusion polypeptides) are known in the art (see, e.g., Holliger, P., et al. (1993) Proc. Natl. Acad. Sci. USA 90:6444-6448; Poljak, RJ, et al. (1994) Structure 2:1 121-1123).
[0120] Nucleic Acid: As used herein, in its broadest sense, refers to any compound and / or substance that is or can be incorporated into an oligonucleotide chain. In some embodiments, nucleic acids are compounds and / or substances that are or can be incorporated into an oligonucleotide chain via a phosphodiester linkage. As will be clear from the context, in some embodiments, "nucleic acid" refers to individual nucleic acid residues (e.g., nucleotides and / or nucleosides), and in some embodiments, "nucleic acid" refers to an oligonucleotide chain comprising individual nucleic acid residues. In some embodiments, "nucleic acid" is or comprises RNA, and in some embodiments, "nucleic acid" is or comprises DNA. In some embodiments, a nucleic acid is, comprises, or consists of one or more naturally occurring nucleic acid residues. In some embodiments, a nucleic acid is, comprises, or consists of one or more nucleic acid analogs. In some embodiments, a nucleic acid analog differs from a nucleic acid in that it does not utilize a phosphodiester backbone. For example, in some embodiments, the nucleic acid is, comprises, or consists of one or more "peptide nucleic acids," which are known in the art and have peptide bonds instead of phosphodiester bonds in the backbone, and are considered within the scope of the present invention. Alternatively or additionally, in some embodiments, the nucleic acid has one or more phosphorothioate and / or 5'-N-phosphoramidite linkages rather than phosphodiester linkages. In some embodiments, the nucleic acid is, comprises, or consists of one or more naturally occurring nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine).In some embodiments, the nucleic acid is, comprises, or consists of one or more nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, C-5 propynyl-cytidine, C-5 propynyl-uridine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, 2-thiocytidine, methylated bases, intercalating bases, and combinations thereof). In some embodiments, the nucleic acid comprises one or more modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose) compared to those of naturally occurring nucleic acids. In some embodiments, the nucleic acid has a nucleotide sequence that encodes a functional gene product such as RNA or a protein. In some embodiments, the nucleic acid comprises one or more introns. In some embodiments, the nucleic acid is prepared by one or more of isolation from a natural source, enzymatic synthesis by polymerization based on a complementary template (in vivo or in vitro), replication in a recombinant cell or system, and chemical synthesis. In some embodiments, the nucleic acid is at least 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 20, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000 or more residues in length. In some embodiments, the nucleic acid is partially or entirely single-stranded, and in some embodiments, the nucleic acid is partially or entirely double-stranded.In some embodiments, the nucleic acid has a nucleotide sequence that includes at least one element that encodes a polypeptide or is the complement of a sequence that encodes a polypeptide. In some embodiments, the nucleic acid has enzymatic activity.
[0121] Payload: As used herein, the term "payload" refers to a protein sequence that can be associated with a peptide barcode via at least one covalent bond. In some embodiments, the payload is a protein that is detected in a pool of proteins. In some embodiments, the payload is an unmodified protein that is detected in a pool of proteins without a peptide barcode attached. In some embodiments, the payload is a modified protein that is detected in a pool of proteins. In some embodiments, the payload can be associated with a barcode (e.g., a peptide barcode). In some embodiments, the payload may not be associated with a barcode (e.g., a peptide barcode).
[0122] Peptide: As used herein, the term "peptide" refers to a polypeptide that is typically relatively short, e.g., having a length of less than about 100 amino acids, less than about 50 amino acids, less than about 40 amino acids, less than about 30 amino acids, less than about 25 amino acids, less than about 20 amino acids, less than about 15 amino acids, or less than 10 amino acids.
[0123] Polypeptide: As used herein, refers to any polymeric chain of residues (e.g., amino acids) typically linked by peptide bonds. In some embodiments, a polypeptide has a naturally occurring amino acid sequence. In some embodiments, a polypeptide has a non-naturally occurring amino acid sequence. In some embodiments, a polypeptide has an engineered amino acid sequence, in that it has been designed and / or created by the hand of man. In some embodiments, a polypeptide can comprise or consist of natural amino acids, unnatural amino acids, or both. In some embodiments, a polypeptide can comprise or consist of only natural amino acids or only unnatural amino acids. In some embodiments, a polypeptide can comprise D-amino acids, L-amino acids, or both. In some embodiments, a polypeptide can comprise only D-amino acids. In some embodiments, a polypeptide can comprise only L-amino acids. In some embodiments, a polypeptide can include one or more pendant groups or other modifications, e.g., modification of or attachment to one or more amino acid side chains, at the N-terminus of the polypeptide, the C-terminus of the polypeptide, or any combination thereof. In some embodiments, such pendant groups or modifications may be selected from the group consisting of acetylation, amidation, lipidation, methylation, pegylation, and the like (including combinations thereof). In some embodiments, a polypeptide may be cyclic and / or include a cyclic moiety. In some embodiments, a polypeptide is not cyclic and / or does not include a cyclic moiety. In some embodiments, a polypeptide is linear. In some embodiments, a polypeptide may be or include a stapled polypeptide. In some embodiments, the term "polypeptide" may be appended to the name of a reference polypeptide, activity, or structure, and in such cases, it is used herein to refer to polypeptides that share a related activity or structure and therefore can be considered members of the same class or family of polypeptides.For each such class, exemplary polypeptides within the class are provided herein, and / or those of skill in the art will be aware of, whose amino acid sequences and / or functions are known. In some embodiments, such exemplary polypeptides are reference polypeptides of a class or family of polypeptides. In some embodiments, members of a polypeptide class or family exhibit significant sequence homology or identity with the reference polypeptide of the class (and in some embodiments, with all polypeptides within the class), share common sequence motifs (e.g., characteristic sequence elements), and / or share a common activity (in some embodiments, at a similar level or within a specified range). For example, in some embodiments, a member polypeptide exhibits an overall degree of sequence homology or identity with a reference polypeptide that is at least about 30-40%, and often greater than about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more, and / or contains at least one region (e.g., a conserved region that, in some embodiments, is or may contain a distinctive sequence element) that exhibits very high sequence identity, often greater than 90%, or even 95%, 96%, 97%, 98%, or 99%. Such a conserved region typically encompasses at least 3-4, and often up to 20 or more, amino acids, and in some embodiments, the conserved region encompasses at least one stretch of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more contiguous amino acids. In some embodiments, a useful polypeptide may comprise or consist of a fragment of a parent polypeptide, hi some embodiments, a useful polypeptide may comprise or consist of multiple fragments, each of which is found in the same parent polypeptide in a different spatial arrangement relative to each other than that found in the polypeptide of interest (e.g., a fragment directly linked to the parent may be spatially separated in the polypeptide of interest, or vice versa, and / or the fragments may be present in a different order in the polypeptide of interest than in the parent), and thus the polypeptide of interest is a derivative of that parent polypeptide.
[0124] Protein: As used herein, the term "protein" refers to a polypeptide (i.e., a series of at least two amino acids linked together by a peptide bond). A protein may contain moieties other than amino acids (e.g., it may be a glycoprotein, proteoglycan, etc.) and / or may be otherwise processed or modified. Those of skill in the art will understand that a "protein" may be an entire polypeptide chain (with or without a signal sequence) produced by a cell, or a characteristic portion thereof. Those of skill in the art will understand that a protein may in some cases comprise two or more polypeptide chains linked by one or more disulfide bonds or associated by other means, for example. Polypeptides may contain l-amino acids, d-amino acids, or both, and may contain any of a variety of amino acid modifications or analogs known in the art. Useful modifications include, for example, terminal acetylation, amidation, methylation, etc. In some embodiments, a protein may comprise natural amino acids, unnatural amino acids, synthetic amino acids, and combinations thereof. In some embodiments, a protein is an antibody, an antibody fragment, a biologically active portion thereof, and / or a distinctive portion thereof.
[0125] Reference: As used herein, refers to a standard or control against which a comparison is made. For example, in some embodiments, an agent, animal, individual, population, sample, sequence, or value of interest is compared to a reference or control agent, animal, individual, population, sample, sequence, or value. In some embodiments, the reference or control is tested and / or determined substantially contemporaneously with the test or determination of interest. In some embodiments, the reference or control is a historical reference or control, optionally embodied in a tangible medium. Typically, as will be understood by those of skill in the art, a reference or control is determined or characterized under conditions or circumstances comparable to those being evaluated. Those of skill in the art will understand when there is sufficient similarity to justify reliance on and / or comparison to a particular reference or control considered.
[0126] Sample: As used herein, the term "sample" typically refers to an aliquot of material obtained or derived from a source of interest. In some embodiments, the term "sample" may be used interchangeably with terms such as "mixture," or "complex mixture," or "complex sample," as understood from the context by one of skill in the art. In some embodiments, the source of interest is a biological or environmental source. In some embodiments, the source of interest may be or include cells or organisms, such as microorganisms, plants, animals (e.g., humans), etc. In some embodiments, the source of interest is or includes biological tissue or fluid. In some embodiments, the biological tissue or fluid may be or include cells, serum, extracellular matrix, CSF, and / or combinations or component(s) thereof. In some embodiments, the biological tissue or fluid may be or include amniotic fluid, aqueous humor, peritoneal fluid, bile, bone marrow, blood, breast milk, cerebrospinal fluid, uterine cavity, chyle, earwax, semen, intestinal lymph, exudate, feces, gastric acid, gastric juice, lymph, mucus, pericardial fluid, perilymph, peritoneal fluid, pleural fluid, pus, catarrhal secretions, saliva, sebum, sperm, serum, smegma, sputum, synovial fluid, sweat, tears, urine, vaginal fluid, vitreous humor, vomit, and / or combinations or component(s) thereof. In some embodiments, the biological fluid may be or include intracellular fluid, extracellular fluid, intravascular fluid (plasma), interstitial fluid, lymph, and / or transcellular fluid. In some embodiments, the biological fluid may be or include plant exudates. In some embodiments, the biological tissue or sample may be obtained by, for example, aspiration, biopsy (e.g., fine needle or tissue biopsy), swab (e.g., oral, nasal, skin, or vaginal swab), scraping, surgery, lavage, or irrigation (e.g., brocheoalvealar, ductal, nasal, ocular, oral, uterine, vaginal, or other lavage or irrigation). In some embodiments, the biological sample is or comprises cells obtained from an individual. In some embodiments, the sample is a "primary sample" obtained directly from the source of interest by any suitable means.In some embodiments, as will be clear from the context, the term "sample" refers to a preparation obtained by processing a primary sample (e.g., by removing one or more components and / or adding one or more agents), such as filtration using a semi-permeable membrane. Such a "processed sample" can include, for example, nucleic acids or proteins extracted from the sample or obtained by subjecting the primary sample to one or more techniques, such as amplification or reverse transcription of nucleic acids, isolation and / or purification of certain components, etc.
[0127] Specific: The term "specific," as used herein with respect to an active agent, will be understood by those skilled in the art to mean that the agent distinguishes between potential target entities or aspects. For example, in some embodiments, an agent is said to bind "specifically" to a target if it preferentially binds to that target in the presence of one or more competing alternative targets. In many embodiments, the specific interaction depends on the presence of a particular structural feature of the target entity (e.g., an epitope, cleft, binding site). It should be understood that specificity need not be absolute. In some embodiments, specificity can be assessed relative to the specificity of a binding agent for one or more other potential target entities (e.g., competitors). In some embodiments, specificity is assessed relative to that of a reference specific binding agent. In some embodiments, specificity is assessed relative to that of a reference nonspecific binding agent. In some embodiments, an agent or entity does not directly bind to a competing alternative target under conditions in which it binds to its target entity. In some embodiments, a binding agent binds to its target entity with a higher on-rate, a lower off-rate, increased affinity, decreased dissociation, and / or increased stability when compared to competing surrogate target(s).
[0128] Subject: As used herein, the term "subject" refers to an organism, typically a mammal (e.g., in some embodiments, a human, including prenatal human forms). In some embodiments, the subject is afflicted with the relevant disease, disorder, or condition. In some embodiments, the subject is susceptible to the disease, disorder, or condition. In some embodiments, the subject exhibits one or more symptoms or characteristics of the disease, disorder, or condition. In some embodiments, the subject does not exhibit any symptoms or characteristics of the disease, disorder, or condition. In some embodiments, the subject is one who possesses one or more characteristics characteristic of susceptibility to or risk for a disease, disorder, or condition. In some embodiments, the subject is a patient. In some embodiments, the subject is an individual to whom and / or to whom diagnosis and / or therapy is administered.
[0129] Substantially: As used herein, the term "substantially" refers to the qualitative state of exhibiting the entire or nearly entire extent or degree of a desired characteristic or property. Those skilled in the art of biology will understand that biological and chemical phenomena rarely, if ever, proceed to completion and / or perfection or achieve or avoid absolute results. Thus, the term "substantially" is used herein to capture the potential lack of completeness inherent in many biological and chemical phenomena.
[0130] Therapeutic Agent: As used herein, the phrase "therapeutic agent" generally refers to any agent that induces a desired pharmacological effect when administered to an organism. In some embodiments, an agent is considered to be a therapeutic agent if it shows a statistically significant effect in an appropriate population. In some embodiments, the appropriate population may be a population of model organisms. In some embodiments, the appropriate population may be defined by various criteria, such as a particular age group, sex, genetic background, pre-existing clinical conditions, etc. In some embodiments, a therapeutic agent is a substance that can be used to alleviate, ameliorate, relieve, inhibit, prevent, delay the onset, reduce the severity, and / or reduce the incidence of one or more symptoms or characteristics of a disease, disorder, and / or condition. In some embodiments, a "therapeutic agent" is an agent that has been, or needs to be, approved by a government agency before it can be sold for administration to humans. In some embodiments, a "therapeutic agent" is an agent that requires a medical prescription for administration to humans. In some embodiments, a therapeutic agent is a therapeutic protein.
[0131] I. Barcoded Payload Methods and systems for generating and using barcodes and barcoded payloads are described herein. In some embodiments, the methods disclosed herein are used to detect and / or characterize a payload. In some embodiments, the methods disclosed herein are used to detect and / or characterize a protein. In some embodiments, the methods disclosed herein are used to detect and / or characterize a therapeutic protein. In some embodiments, the methods disclosed herein are used to detect and / or characterize a non-therapeutic protein. In some embodiments, the methods disclosed herein are used to detect and / or characterize a protein by tagging the protein with a barcode (e.g., a barcoded protein). In some embodiments, the methods disclosed herein are used to detect and / or characterize a protein in vitro. In some embodiments, the methods disclosed herein are used to detect and / or characterize a protein in vivo. In some embodiments, the methods disclosed herein are used to detect and / or characterize a protein. In some embodiments, the methods disclosed herein are used to detect and / or characterize a plurality (e.g., two or more, three or more, four or more, etc.) of proteins.
[0132] Barcode: In some embodiments, the barcode is or comprises an amino acid sequence. In some embodiments, the barcode is or comprises a naturally occurring amino acid sequence. In some embodiments, the barcode is or comprises a non-naturally occurring amino acid sequence. In some embodiments, the barcode is or comprises an amino acid sequence that is synthetic. In some embodiments, the barcode comprises a naturally occurring amino acid. In some embodiments, the barcode comprises a non-naturally occurring amino acid (e.g., a modified amino acid). In some embodiments, the barcode is or comprises a peptide barcode.
[0133] Barcodes of the present disclosure may vary in length. For example, in some embodiments, barcodes may have lengths ranging from 1 to 100 amino acids. In some embodiments, barcodes may have lengths ranging from 5 to 50 amino acids. In some embodiments, barcodes may have lengths ranging from 8 to 25 amino acids. In some embodiments, barcodes may have lengths ranging from 9 to 25 amino acids. In some embodiments, barcodes may have lengths ranging from 9 to 15 amino acids. In some embodiments, barcodes may have lengths of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids. In some embodiments, barcodes may have lengths of at least 5 amino acids. In some embodiments, barcodes may have lengths of up to 100 amino acids.
[0134] The barcodes described herein may be available in libraries of different formats. For example, in some embodiments, the barcodes described herein may be described as nucleic acid sequences. In other examples, the barcodes described herein may be described as amino acid sequences. Those skilled in the art will understand that barcodes described in one format can be converted to another format using basic biological principles. Thus, barcodes described as nucleic acid sequences may be translated into proteins, which may be used to detect the presence or absence of payloads (e.g., proteins) in a mixture. Such translated barcodes are referred to herein as peptide barcodes.
[0135] Thus, when written using nucleic acids, barcodes of the present disclosure may have lengths that differ from the amino acid sequence lengths disclosed in the paragraph above. For example, in some embodiments, barcodes may have lengths ranging from 3 to 300 nucleotides. In some embodiments, barcodes may have lengths ranging from 15 to 150 nucleotides. In some embodiments, barcodes may have lengths ranging from 24 to 75 nucleotides. In some embodiments, barcodes may have lengths ranging from 27 to 75 nucleotides. In some embodiments, barcodes may have lengths ranging from 27 to 45 nucleotides. In some embodiments, the barcode may have a length of 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, or 75 nucleotides. In some embodiments, the barcode may have a length of at least 15 nucleotides. In some embodiments, the barcode may have a length of up to 300 nucleotides.
[0136] Barcodes of the present disclosure may have one or more properties. In some embodiments, the barcode may be naturally occurring. In some embodiments, the barcode may not be naturally occurring (e.g., synthetic). In some embodiments, the barcode may have relatively no effect on payload function. For example, in some embodiments, tagging a payload (e.g., a protein) with a barcode as described herein does not relatively alter or change the function of the tagged payload. In some embodiments, the barcode may have an effect (e.g., positive or negative) on payload function. For example, in some embodiments, tagging a payload (e.g., a protein) with a barcode described herein may relatively alter or change the function of the tagged payload (e.g., half-life (e.g., longer half-life), enhanced targeting to a particular tissue, etc.). In some embodiments, the barcode may not tamper with an immune response (e.g., IgG response, complement response, etc.). In some embodiments, the barcodes are orthogonal to one another. In some embodiments, the barcodes are not orthogonal to one another.
[0137] The barcodes of the present disclosure may be attached to various positions on a payload. For example, in some embodiments, the barcode may be attached to a preferred position on a payload. In some embodiments, the barcode may be attached to an unfavorable position on a payload. In some embodiments, the barcode may be attached to a preferred position on a protein. For example, in some embodiments, the barcode may be attached to the N-terminus of a protein. In some embodiments, the barcode may be attached to the C-terminus of a protein. In some embodiments, the barcode may be attached to a non-terminal position on a protein (e.g., a side chain). In some embodiments, the barcode may be attached to an unfavorable position on a protein.
[0138] Among other things, barcodes of the present disclosure (e.g., peptide barcodes or nucleic acids encoding peptide barcodes) may be flanked by additional sequences (e.g., nucleic acid sequences, amino acid sequences, etc.). In some embodiments, a barcode may be flanked by additional sequences on the 5' end of the barcode. In some embodiments, a barcode may be flanked by additional sequences on the 3' end of the barcode. In some embodiments, a barcode may be flanked by additional sequences on the 3' and 5' ends of the barcode. In some embodiments, the additional sequence may be a primer binding site, a restriction endonuclease recognition sequence, a restriction enzyme site (e.g., cleavage site), a sequence encoding an amino acid sequence, a sequence that does not encode an amino acid sequence, an amino acid sequence, or a nucleic acid sequence. For example, in some embodiments, a barcode may be flanked by a nucleic acid sequence that encodes an amino acid sequence. In some embodiments, a barcode may be flanked by a nucleic acid sequence that does not encode an amino acid sequence. In some embodiments, a barcode may be flanked by an amino acid sequence. In some embodiments, a peptide barcode may be flanked by an amino acid sequence (e.g., glycine-serine (GS)). Similarly, in some embodiments, a nucleic acid encoding a peptide barcode of the present disclosure may be flanked by additional sequences (e.g., nucleic acid sequences, amino acid sequences, etc.). In some embodiments, a nucleic acid encoding a peptide barcode may be flanked by a nucleic acid sequence on the 5' end. In some embodiments, a nucleic acid encoding a peptide barcode may be flanked by a nucleic acid sequence on the 3' end. In some embodiments, a nucleic acid encoding a peptide barcode may be flanked by nucleic acid sequences on the 3' and 5' ends. In some embodiments, a nucleic acid encoding a peptide barcode may be flanked by a nucleic acid sequence encoding an amino acid sequence comprising glycine-serine (GS).
[0139] In some embodiments, the barcode may be flanked by restriction endonuclease recognition sequences. In some embodiments, the barcode may be flanked by restriction endonuclease recognition sequences on the 5' end of the barcode. In some embodiments, the barcode may be flanked by restriction endonuclease recognition sequences on the 3' end of the barcode. In some embodiments, the barcode may be flanked by restriction endonuclease recognition sequences on the 3' and 5' ends of the barcode. In some embodiments, the nucleic acid encoding the peptide barcode may be flanked by a restriction endonuclease recognition sequence. In some embodiments, the nucleic acid encoding the peptide barcode may be flanked by a restriction endonuclease recognition sequence on the 5' end. In some embodiments, the nucleic acid encoding the peptide barcode may be flanked by a restriction endonuclease recognition sequence on the 3' end. In some embodiments, the nucleic acid encoding the peptide barcode may be flanked by restriction endonuclease recognition sequences on the 3' and 5' ends. In some embodiments, the restriction endonuclease recognition sequence may be recognized by one or more restriction enzymes (e.g., BsaI, BsmBI, BbsI, SapI, etc.). In some embodiments, the restriction endonuclease recognition sequence is a type I, type II, or type IIs restriction endonuclease recognition sequence. Such a recognition sequence can be used, for example, to generate a universal overhang that can be used to clone peptide barcodes into different positions of various payloads. Such flexibility allows the barcodes to be used to detect different protein payloads in different experiments.
[0140] In some embodiments, a nucleic acid sequence encoding a barcode (e.g., a peptide barcode) can be associated (e.g., attached, linked) with a second nucleic acid sequence encoding a payload (e.g., a protein payload of interest). Such a nucleic acid sequence can be translated, for example, to form a barcoded payload (e.g., a barcoded protein). In some embodiments, the nucleic acid sequence encoding the barcode (e.g., a peptide barcode) is separate from the second nucleic acid sequence encoding the payload (e.g., a protein payload of interest). For such nucleic acid sequences, for example, the nucleic acid sequence encoding the barcode can be translated separately from the second nucleic acid sequence encoding the payload and then attached using one or more methods known in the art to join the separate amino acid sequences (e.g., using a linker).
[0141] Barcodes of the present disclosure may be associated with (e.g., directly or indirectly attached to) a payload to form a barcoded payload. For example, in some embodiments, each barcode sequence (e.g., a peptide barcode sequence) may be associated with only one payload of interest (e.g., a protein of interest) in a mixture. In some embodiments, each barcode sequence may be associated with two or more payloads of interest (e.g., payloads having different sequences) in a mixture. In some embodiments, multiple (e.g., two or more, three or more, four or more, etc.) barcode sequences may be associated with one payload of interest in a mixture. For example, in some embodiments, one or more barcode sequences may be associated with various different positions on a given payload; such configurations may be useful, for example, in studying and identifying the stability and / or cleavage of such barcoded payloads. In some embodiments, each payload in a mixture is a unique sequence (e.g., each payload has a sequence that is different from all other payloads in the mixture). In some embodiments, each payload in a mixture is a non-unique sequence.
[0142] Various methods and parameters can be used to select a suitable barcode for a given payload. For example, the stability of the barcoded payload is key in determining whether the payload can be tagged with that barcode. In some embodiments, a barcode can be tagged to a particular payload across different experiments. In some embodiments, a barcode can be tagged to different payloads in different experiments. For example, in some embodiments, a barcode can be tagged to two or more, three or more, four or more, ten or more, one hundred or more, one thousand or more, or ten thousand or more different payloads across different experiments.
[0143] In some embodiments, a barcode may be associated with only one payload in a given experiment. In some embodiments, a barcode may be associated with multiple payloads in a given experiment. For example, in some embodiments, one or more barcodes are associated with multiple payloads in a mixture (i.e., a barcode is tagged to multiple payloads), such that each payload is associated with a unique set of barcodes within the mixture. That is, each payload may be associated with a unique "pattern" of barcodes in the mixture. Similarly, in some embodiments, several payloads may be associated with the same barcode.
[0144] Notably, the barcodes described herein are designed to have distinct (i.e., unique) sequences. In some embodiments, the barcodes are designed to have distinct (e.g., separate from other barcodes) sequences. For example, each barcode is designed to be different (e.g., unique) from all other barcodes used in the experiment, each payload (e.g., protein to be measured) is attached to at least one barcode, and each barcode (e.g., barcode with a particular sequence) is attached to only one payload. As will be understood by those skilled in the art, the diversity of barcodes contained within a pool is limited only by the possible diversity of amino acid sequences for a given barcode length. For example, for a barcode length "N," 20 of the length N may be used. N There are distinct amino acid barcode sequences (if only unmodified / naturally occurring amino acids are used), i.e., for a barcode length of 15, the theoretical limit is 20 15 , i.e. 3.2768×10 19 is.
[0145] Examples of barcodes according to various embodiments of the present disclosure are listed in Tables 1 and 2. In some embodiments, the barcode (e.g., peptide barcode) is or includes an amino acid sequence selected from SEQ ID NOs: 5347-8398. In some embodiments, the barcode (e.g., peptide barcode) is encoded by a sequence that is or includes a nucleic acid sequence selected from SEQ ID NOs: 1148-4199.
[0146] payload: The methods and systems disclosed herein can be used to detect one or more payloads described herein. In some embodiments, the payload is a protein of interest. In some embodiments, the payload is a protein with a therapeutic function. In some embodiments, the payload is a protein without a therapeutic function (e.g., it may be complementary to another payload with a therapeutic function). For example, possible payloads are proteins that one might want to screen as drugs, such as monoclonal antibodies, single domain antibodies, enzymes, bispecific antibodies, or any other protein that may have a therapeutic function.
[0147] In one aspect, the systems and methods disclosed herein can be used to detect payloads (e.g., proteins) in a mixture. Specifically, barcodes disclosed herein are tagged to payloads (e.g., barcoded payloads (e.g., barcoded proteins)) in the mixture and used to detect the payloads in the mixture. In some embodiments, each payload differs from all other payloads in the mixture. In some embodiments, each payload in the mixture differs from all other payloads in the mixture by at least one amino acid. In some embodiments, each payload in the mixture differs from all other payloads in the mixture by two or more amino acids. In some embodiments, payloads (e.g., in a mixture) can be tagged with barcodes. In some embodiments, each payload (e.g., in a mixture) can be tagged with the same barcode. In some embodiments, each payload (e.g., in a mixture) can be tagged with a different barcode. In some embodiments, a payload (e.g., in a mixture) can be tagged with a barcode that differs from all other barcodes (e.g., associated with other payloads) in the mixture by at least one amino acid. In some embodiments, a payload (e.g., in a mixture) can be tagged with a barcode that differs from all other barcodes (e.g., associated with other payloads) in the mixture by two or more amino acids.
[0148] As discussed elsewhere herein, payloads may be tagged with different barcodes (e.g., in different mixtures, different experiments, etc.). For example, as described above, in some embodiments, each barcode sequence may be attached to only one payload of interest within a mixture. In some embodiments, each barcode sequence may be attached to two or more payloads of interest within a mixture (e.g., payloads having different sequences). In some embodiments, multiple (e.g., two or more, three or more, four or more, etc.) barcode sequences may be attached to one payload of interest within a mixture. For example, in some embodiments, one or more barcode sequences may be attached to various different locations on a given payload; such a configuration may be useful, for example, in studying the stability and identification of such barcoded payloads. In some embodiments, each payload in a mixture is a unique sequence (e.g., each payload has a sequence that is different from all other payloads in the mixture). In some embodiments, each payload in a mixture is a non-unique sequence.
[0149] In some embodiments, a payload may be tagged to a particular barcode across different experiments. In some embodiments, a payload may be tagged to different barcodes in different experiments. For example, in some embodiments, a payload may be tagged to two or more, three or more, four or more, ten or more, one hundred or more, one thousand or more, or ten thousand or more different barcodes across different experiments.
[0150] In some embodiments, a payload may be associated with only one barcode in a given experiment. In some embodiments, a payload may be associated with multiple barcodes in a given experiment. For example, in some embodiments, one or more barcodes (e.g., in a mixture) are associated with multiple payloads in the mixture (i.e., a barcode is tagged to multiple payloads), such that each payload is associated with a unique set of barcodes within the mixture. That is, each payload may be associated with a unique "pattern" of barcodes in the mixture. In some embodiments, several payloads may be associated with the same barcode.
[0151] Linker: In particular, the systems and methods described herein may use a linker. In some embodiments, a payload described herein and a barcode described herein are separated by a linker. In some embodiments, the linker (L) provides distance between the payload (P) and the barcode (b). That is, structurally, the barcoded payload may have the sequence of PLb in some embodiments. This may contribute, for example, to folding properties, payload functionality, and / or payload stability.
[0152] In some embodiments, the linker can be a nucleic acid. In some embodiments, the linker can be an amino acid. The linkers described herein can have a variety of lengths. For example, in some embodiments, the linker can be at least 3 amino acids long. In some embodiments, the linker can be 1 to 50 amino acids long (e.g., 1 to 30 amino acids long).
[0153] In some embodiments, the linkers of the invention can be cleaved upon treatment. For example, in some embodiments, the linker can include one or more motifs that can be cleaved upon treatment.
[0154] In some embodiments, the linkers of the present invention may be resistant to cleavage. In some embodiments, the linkers of the present invention may be resistant to cleavage in an assay. In some embodiments, the linkers of the present invention may be resistant to cleavage in vivo.
[0155] In one aspect, linkers can be used to tag barcodes. In some embodiments, each linker sequence is associated with a distinct barcode sequence. For example, in some embodiments, linker sequences can be used as unique tags associated with distinct barcode sequences (e.g., nucleic acid sequences) in a mixture. That is, in some embodiments, such linkers can be used to amplify associated barcode sequences. For example, in some embodiments, such linkers can be used as primers to amplify associated barcode sequences. In some embodiments, the amplified linkers can then be used to isolate the associated barcode sequences, allowing for the removal of the barcode sequence (e.g., nucleic acid sequence) from a given linker-barcode pair. In some embodiments, the linker-barcode pair can be subjected to DNA sequencing for identification of the barcode sequence.
[0156] In some embodiments, a nucleic acid sequence encoding the linker-barcode pair may be used to associate (e.g., link) the linker-barcode pair to a new payload.
[0157] II. Binders and Binding Agents binder: In some embodiments, a Binder (i.e., a Binder moiety) is or comprises a nucleic acid sequence. In some embodiments, a Binder is or comprises a naturally occurring nucleic acid sequence. In some embodiments, a Binder is or comprises a non-naturally occurring nucleic acid sequence. In some embodiments, a Binder is or comprises a nucleic acid sequence that is synthetic. In some embodiments, a Binder comprises a naturally occurring nucleic acid. In some embodiments, a Binder comprises a non-naturally occurring nucleic acid (e.g., a modified nucleic acid).
[0158] In some embodiments, the binder nucleic acid sequence is or comprises a sequence encoding a polypeptide sequence. For example, in some embodiments, the binder nucleic acid sequence may contain a region encoding a polypeptide sequence that confers high affinity and / or specificity for a given barcode (e.g., a peptide barcode). In some embodiments, the binder nucleic acid sequence is or comprises a sequence encoding an antibody. In some embodiments, the binder nucleic acid sequence is or comprises a sequence encoding a fragment of an antibody. In some embodiments, the binder nucleic acid sequence is or comprises a sequence encoding a single chain variable fragment (scFv). As may be known to one of skill in the art, an scFv is a fragment of an immunoglobulin heavy chain (V H ) and light chain (V L In some embodiments, the V H and V L The chains may be connected by a short linker peptide (eg, a linker of about 5 to 50 amino acids in length, 10 to 25 amino acids in length, etc.).
[0159] In some embodiments, for example, binders are generated with known specificity and affinity for a given barcode. In some embodiments, binders are generated with known specificity and affinity for one barcode. In some embodiments, binders are generated with known specificity and affinity for multiple (e.g., two or more, three or more, etc.) barcodes. In some embodiments, binders are generated with known specificity and affinity for at least one barcode. In some embodiments, binders are expressed on the surface of a binding agent (e.g., a phage, a ribosome, etc.), for example, using methods known to those of skill in the art.
[0160] In some embodiments, the binder associates with (eg, has known specificity and affinity for) the barcode.
[0161] In some embodiments, the binder is or comprises a naturally occurring polypeptide sequence. In some embodiments, the binder is or comprises a non-naturally occurring polypeptide sequence. In some embodiments, the binder is or comprises a synthetic polypeptide sequence. In some embodiments, the binder comprises naturally occurring amino acids. In some embodiments, the binder comprises non-naturally occurring amino acids (e.g., modified amino acids).
[0162] Binders of the present invention can vary in length. For example, in some embodiments, binders can have lengths ranging from 5 to 1000 amino acids. In some embodiments, binders can have lengths ranging from 5 to 800 amino acids. In some embodiments, binders can have lengths ranging from 6 to 500 amino acids. In some embodiments, binders can have lengths ranging from 10 to 400 amino acids. In some embodiments, binders can have lengths ranging from 5 to 500 amino acids. In some embodiments, binders can have lengths ranging from 5 to 1000 amino acids. In some embodiments, binders can have lengths of 10 amino acids. In some embodiments, binders can have lengths of at least 5 amino acids. In some embodiments, binders can have lengths of up to 1000 amino acids.
[0163] The binders described herein may be available in libraries of different formats. For example, in some embodiments, the binders described herein may be described as nucleic acid sequences. In other examples, the binders described herein may be described as amino acid sequences. Those skilled in the art will understand that binders described in one format can be converted to another format using basic biological principles. Thus, a binder described as a nucleic acid sequence may be translated into a protein, and the protein may be used to detect the presence or absence of a payload (e.g., a barcoded payload (e.g., a barcoded protein)) in a mixture. Such translated binders are referred to herein as polypeptide binders or polypeptide binder moieties.
[0164] Thus, when described using nucleic acids, binders of the present disclosure can have lengths that differ from the amino acid sequence lengths disclosed in the paragraph above. For example, in some embodiments, binders can have lengths ranging from 15 to 3000 nucleotides. In some embodiments, binders can have lengths ranging from 15 to 2400 nucleotides. In some embodiments, binders can have lengths ranging from 24 to 1500 nucleotides. In some embodiments, binders can have lengths ranging from 30 to 1200 nucleotides. In some embodiments, binders can have lengths of 30 nucleotides. In some embodiments, binders can have lengths of at least 15 nucleotides. In some embodiments, binders can have lengths of up to 3000 nucleotides.
[0165] The binders of the present disclosure may have one or more particular properties. In some embodiments, the binders may be naturally occurring. In some embodiments, the binders may not be naturally occurring (e.g., synthetic). In some embodiments, the binders may not tamper with the immune response (e.g., IgG response, complement response, etc.).
[0166] In particular, binders of the present disclosure (e.g., polypeptide binders, nucleic acids encoding binders) may be flanked by additional sequences (e.g., nucleic acid sequences, amino acid sequences, etc.), such as the barcodes described above. In some embodiments, a binder may be flanked by additional sequences on the 5' end of the binder. In some embodiments, a binder may be flanked by additional sequences on the 3' end of the binder. In some embodiments, a binder may be flanked by additional sequences on the 3' and 5' ends of the binder. In some embodiments, the additional sequence may be a primer binding site, a restriction endonuclease recognition sequence, a restriction enzyme site (e.g., a cleavage site), a sequence encoding an amino acid sequence, a sequence that does not encode an amino acid sequence, an amino acid sequence, or a nucleic acid sequence.
[0167] In one aspect of the invention, a binder nucleic acid sequence can be associated with (e.g., attached or linked to) another nucleic acid sequence. For example, in some embodiments, a binder nucleic acid sequence can be associated with a nucleic acid sequence encoding one or more genes. In some embodiments, a binder nucleic acid sequence can be associated with a nucleic acid sequence encoding one or more genes of a phage (e.g., m13). In some embodiments, a binder nucleic acid sequence can be associated with a nucleic acid sequence encoding a polypeptide. In some embodiments, a binder nucleic acid sequence can be associated with a nucleic acid sequence encoding a phage polypeptide (e.g., m13 gene 3 protein). Binder-gene 3 protein fusions can be expressed and incorporated into m13 phage.
[0168] In particular, the binders described herein are designed to have distinct (i.e., unique) sequences. In some embodiments, the binders are designed to have distinct (e.g., distinct from other binders) sequences. For example, each binder is designed to be distinct (e.g., unique) from all other binders used in the experiment.
[0169] In one aspect of the invention, as described herein, binders bind to barcodes or barcoded payloads with high specificity and high affinity. In some embodiments, a barcode or barcoded payload (e.g., a barcoded protein to be measured) binds to one binder, and each binder (e.g., a binder with a specific sequence) binds to one barcode or barcoded payload. In some embodiments, a barcode or barcoded payload (e.g., a barcoded protein to be measured) binds to at least one binder. In some embodiments, each binder (e.g., a binder with a specific sequence) binds to at least one barcode or barcoded payload. In some embodiments, multiple binders (e.g., with different sequences (e.g., polypeptide sequences)) bind to a single barcode. In some embodiments, multiple barcodes (e.g., with different sequences (e.g., peptide sequences)) bind to a single binder.
[0170] Examples of binders according to various embodiments of the present disclosure are listed in Tables 1 and 2. In some embodiments, the binder (e.g., a polypeptide binder) is or comprises an amino acid sequence selected from SEQ ID NOs: 4200-5346. In some embodiments, the binder (e.g., a polypeptide binder) is or is encoded by a sequence that comprises a nucleic acid sequence selected from SEQ ID NOs: 1-1147.
[0171] Binder: The methods described herein relate to the detection of one or more barcodes using a binding agent. In some embodiments, the binding agent is associated with or includes a detectable nucleic acid. In some embodiments, the binding agent expresses a detectable nucleic acid. In some embodiments, the binding agent expresses a detectable nucleic acid on its surface (e.g., a binder). In some embodiments, the binding agent expresses an antibody on its surface.
[0172] In some embodiments, for example, to detect the presence of a specific (e.g., distinct) barcode, the present invention contemplates the association of a distinct, detectable nucleic acid (e.g., a DNA sequence, an RNA sequence, etc.) to the specific barcode. This is accomplished by contacting the barcode with a binding agent. In some embodiments, one or more barcodes may be contacted with a binding agent. In some embodiments, one or more binding agents may be contacted with a barcode.
[0173] In some embodiments, the binding agent may be or include a phage, a ribosome, mRNA, DNA, etc. In some embodiments, the binding agent is a phage. In some embodiments, the binding agent may be an M13 phage, a T4 phage, a T7 phage, a lambda phage, or a filamentous phage. In some embodiments, the binding agent may be an M13 phage.
[0174] The binders disclosed herein may be expressed on a binding agent using methods known in the art. For example, one of skill in the art may be able to express nucleic acid encoding a polypeptide binder on (e.g., on the surface of) a phage using techniques and methods available in the art.
[0175] III. Generation Generate barcode: Disclosed herein are methods and systems for generating barcodes for use in the disclosed systems and methods. In some embodiments, the barcodes described herein may be generated rapidly (e.g., in about 1 week, about 2 weeks, about 3 weeks, about 4 weeks, about 1 month, about 2 months, about 3 months, about 4 months, about 5 months, about 6 months, or about 1 year). In some embodiments, for example, about 100 to about 1,000 barcodes may be generated rapidly. In some embodiments, about 10 to about 1,000 barcodes may be generated rapidly. In some embodiments, about 10 to about 10,000 barcodes may be generated rapidly. As described herein, large numbers of barcodes can be generated rapidly, but such barcodes are also robust, and barcodes generated using the methods disclosed herein may bind specifically and with distinct affinities to a set of known binders.
[0176] According to various embodiments, the barcodes described herein can be synthesized using nucleic acid (e.g., oligonucleotide) arrays. In some embodiments, the barcodes described herein can be synthesized using DNA arrays. In some embodiments, the nucleic acids (e.g., oligonucleotides) of the nucleic acid array are expressed into barcodes. In some embodiments, the barcodes described herein can be synthesized using nucleic acid libraries. In some embodiments, the nucleic acid libraries are synthesized using nucleic acid arrays. In some embodiments, the nucleic acids (e.g., oligonucleotides) of the nucleic acid library are expressed into barcodes.
[0177] In some embodiments, a barcode nucleic acid library contains about 1 or more, about 2 or more, about 3 or more, about 4 or more, about 5 or more, about 10 or more, about 50 or more, about 100 or more, about 200 or more, about 300 or more, about 400 or more, about 500 or more, about 600 or more, about 700 or more, about 800 or more, about 900 or more, about 1000 or more, about 2000 or more, about 3000 or more, about 4000 or more, or about 5000 or more potential barcodes. In some embodiments, a nucleic acid library contains one or more potential barcode sequences. Such potential barcode sequences can be screened for functionality as peptide barcodes (i.e., after translation of the potential barcode nucleic acid sequences) using one or more methods described herein.
[0178] Barcodes of the present disclosure can be screened for one or more particular properties. In some embodiments, barcodes can be screened for specific binding (e.g., specificity, binding affinity) to a binder. In some embodiments, barcodes can be screened for specific binding to one or more binders. In some embodiments, barcodes can be screened for specific binding to at least a binder. In some embodiments, barcodes can be screened for specific binding to up to a binder. In some embodiments, barcodes can be screened for specific binding to multiple binders.
[0179] As can be appreciated by one of skill in the art, barcodes are designed to be distinct (i.e., unique (e.g., have a unique sequence)) within a pool of barcodes. Such distinction, in some embodiments, can be achieved by altering one or more amino acids in the barcode. In some embodiments, a barcode differs from other barcodes in the pool of barcodes by one amino acid. In some embodiments, a barcode differs from other barcodes in the pool of barcodes by at least one amino acid. In some embodiments, a barcode differs from other barcodes in the pool of barcodes by at most one amino acid. In some embodiments, a barcode differs from other barcodes in the pool of barcodes by 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids from other barcodes in the barcode pool. In some embodiments, a barcode differs from other barcodes in the pool of barcodes by at least two amino acids. In some embodiments, the barcode differs from other barcodes in the pool of barcodes by a maximum of 50 amino acids.
[0180] Generate a barcoded payload: Barcoded payloads according to the present invention can be generated in a variety of ways. In some embodiments, payload-barcode nucleic acid sequence pairs can be inserted into a plasmid to allow for expression in different expression systems (e.g., protein expression systems). In some embodiments, at least one payload-barcode nucleic acid sequence pair is inserted into a plasmid. In some embodiments, at least two payload-barcode nucleic acid sequence pairs are inserted into a plasmid. In some embodiments, at least three payload-barcode nucleic acid sequence pairs are inserted into a plasmid. In some embodiments, one or more payload-barcode nucleic acid sequence pairs are inserted into a plasmid.
[0181] In some embodiments, the payload-barcode nucleic acid sequence may comprise additional sequences. In some embodiments, the payload-barcode nucleic acid sequence may comprise additional nucleic acid sequences. In some embodiments, the payload-barcode nucleic acid sequence may comprise a universal motif sequence. In some embodiments, the payload-barcode nucleic acid sequence may comprise at least one universal motif sequence. In some embodiments, the payload-barcode nucleic acid sequence may comprise at least two universal motif sequences. In some embodiments, the payload-barcode nucleic acid sequence may comprise more than one universal motif sequence.
[0182] In some embodiments, at least one payload-barcode nucleic acid sequence in the pool of payload-barcode nucleic acid sequences may comprise a universal motif sequence, hi some embodiments, all payload-barcode nucleic acid sequences in the pool of payload-barcode nucleic acid sequences may comprise a universal motif sequence.
[0183] Different plasmids can be used to generate the techniques described herein. In some embodiments, the plasmid is a DNA plasmid. In some embodiments, the plasmid is an RNA plasmid. In some embodiments, the plasmid is a fertility F plasmid. In some embodiments, the plasmid is a resistance plasmid. In some embodiments, the plasmid is a toxic plasmid. In some embodiments, the plasmid is a degradative plasmid. In some embodiments, the plasmid is a Col plasmid.
[0184] Different hosts (e.g., host cells, host cell lines, etc.) can be used to generate the technology described herein. In some embodiments, the host is a mammalian host. In some embodiments, the host is a non-mammalian host. In some embodiments, the host is an insect. In some embodiments, the host is a bacterium. In some embodiments, the host is E. coli.
[0185] In some embodiments, the payload-barcode pair is expressed in vitro. In some embodiments, the payload-barcode pair is expressed in vivo. In some embodiments, the payload-barcode pair is expressed from RNA. In some embodiments, the payload-barcode pair is expressed from transcribed RNA. In some embodiments, the payload-barcode pair is expressed from DNA. In some embodiments, the payload-barcode pair is expressed using a protein component (e.g., required for protein translation).
[0186] After expression of the barcoded payload construct, the construct can be purified from the pool. In some embodiments, purification can be performed using a universal motif. In some embodiments, purification can be performed using a HIS tag, a FLAG tag, a HALO tag, a SNAP tag, an Avitag, a Twin Strep tag, or any other tag-based method for protein purification known in the art.
[0187] Generate the binder: Disclosed herein are methods and systems for generating binders for use in the disclosed systems and methods. In some embodiments, the binders described herein may be generated rapidly (e.g., in about 1 week, about 2 weeks, about 3 weeks, about 4 weeks, about 1 month, about 2 months, about 3 months, about 4 months, about 5 months, about 6 months, or about 1 year). In some embodiments, for example, about 100 to about 1,000 binders may be generated rapidly. In some embodiments, about 10 to about 1,000 binders may be generated rapidly. In some embodiments, about 10 to about 10,000 binders may be generated rapidly. As described herein, while large numbers of binders can be generated rapidly, such binders are also robust, and binders generated using the methods disclosed herein may bind to a set of known barcodes with specificity and distinct affinity.
[0188] In some embodiments, a binder nucleic acid library contains about 1 or more, about 2 or more, about 3 or more, about 4 or more, about 5 or more, about 10 or more, about 50 or more, about 100 or more, about 200 or more, about 300 or more, about 400 or more, about 500 or more, about 600 or more, about 700 or more, about 800 or more, about 900 or more, about 1000 or more, about 2000 or more, about 3000 or more, about 4000 or more, or about 5000 or more potential binders. In some embodiments, a nucleic acid library contains one or more potential binder sequences. Such potential binder sequences can be screened for functionality as polypeptide binders (i.e., after translation of the potential nucleic acid binder sequences) using one or more methods described herein.
[0189] Binders according to the present invention can be produced in a variety of ways. In some embodiments, binder nucleic acid sequences may be inserted into a plasmid to allow expression in different expression systems. In some embodiments, at least one binder nucleic acid sequence is inserted into a plasmid. In some embodiments, at least two binder nucleic acid sequences are inserted into a plasmid. In some embodiments, at least three binder nucleic acid sequences are inserted into a plasmid. In some embodiments, one or more binder nucleic acid sequences are inserted into a plasmid.
[0190] In some embodiments, the binder nucleic acid sequence is attached to one or more genes. In some embodiments, the binder nucleic acid sequence is attached to one or more genes before insertion into the plasmid. In some embodiments, the binder nucleic acid sequence is attached to one or more genes after insertion into the plasmid. In some embodiments, the binder nucleic acid sequence is attached to a bacteriophage gene. In some embodiments, the binder nucleic acid sequence is attached to an m13 bacteriophage gene. In some embodiments, the binder nucleic acid sequence is attached to gene 3 (i.e., encoding the gene 3 protein) of m13 bacteriophage.
[0191] In some embodiments, a plasmid (e.g., containing a binder sequence, containing a binder and a bacteriophage sequence, etc.) can be transformed into a host. In some embodiments, the plasmid can be transformed into a host and expressed. In some embodiments, the plasmid is transformed into bacteria. In some embodiments, the plasmid is transformed into E. coli.
[0192] In some embodiments, expression of the plasmid results in phage production. In some embodiments, expression of the plasmid results in display of a binder on the surface of the phage. In some embodiments, expression of the plasmid results in display of two binders on the surface of the phage. In some embodiments, expression of the plasmid results in display of at least one binder on the surface of the phage. In some embodiments, expression of the plasmid results in display of one or more binders on the surface of the phage. In some embodiments, expression of the plasmid results in display of one or more binders on one or more surfaces of the phage. In some embodiments, expression of the plasmid results in display of at least one binder on one or more surfaces of the phage.
[0193] After phage generation, the resulting pool can be purified to determine the presence of one or more polypeptide binders. In some embodiments, purification can be performed using a universal motif. In some embodiments, purification can be performed using a HIS tag, a FLAG tag, a HALO tag, a SNAP tag, an Avitag, a Twin Strep tag, or any other tag-based method of protein purification known in the art.
[0194] In some embodiments, the purified binder pool may be highly diverse. In some embodiments, the purified binder pool may not be highly diverse. In some embodiments, the purified binder pool is subjected to a screening method to select binders of interest.
[0195] Binders of the present disclosure may be screened for one or more particular properties. In some embodiments, binders may be screened for specific binding to a barcode. In some embodiments, binders may be screened for specific binding to one or more barcodes. In some embodiments, binders may be screened for specific binding to at least a barcode. In some embodiments, binders may be screened for specific binding to at most a barcode. In some embodiments, binders may be screened for specific binding to multiple barcodes.
[0196] As can be appreciated by one of skill in the art, binders are designed to be distinct (i.e., unique (e.g., have a unique sequence)) within a pool of binders. Such distinction may, in some embodiments, be achieved by altering one or more amino acids in the binder. In some embodiments, a binder differs from other binders in the pool of binders by one amino acid. In some embodiments, a binder differs from other binders in the pool of binders by at least one amino acid. In some embodiments, a binder differs from other binders in the pool of binders by at most one amino acid. In some embodiments, a binder differs from other binders in the binder pool by 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids. In some embodiments, a binder differs from other binders in the pool of binders by at least two amino acids. In some embodiments, a binder differs from other binders in the pool of binders by at most 1000 amino acids.
[0197] IV. Characterization sample: As described elsewhere in this disclosure, the sample may be a biological sample. In some embodiments, the sample may contain one or more barcoded payloads. In some embodiments, the sample may contain one or more barcoded proteins.
[0198] In some embodiments, the sample is from an organism. In some embodiments, the sample is from an animal. In some embodiments, the sample is from an animal model of a disease. In some embodiments, the sample is from a non-mammal. In some embodiments, the sample is from a mammal (e.g., a rodent, mouse, rat, rabbit, monkey, dog, cat, sheep, cow, primate, and / or pig). In some embodiments, the sample is from a mouse. In some embodiments, the sample is from a human. In some embodiments, the sample is from a cell (e.g., in vitro). In some embodiments, the sample is a human cell line.
[0199] In some embodiments, the sample may be purified. In some embodiments, the sample may not be purified.
[0200] In some embodiments, the sample is obtained from cells treated with a barcoded payload. In some embodiments, the sample is obtained from cells that were not treated with a barcoded payload. In some embodiments, the sample is obtained from an animal that was treated with a barcoded payload. In some embodiments, the sample is obtained from an animal that was not treated with a barcoded payload. For example, in some embodiments, the sample is obtained from a human that was treated with a barcoded protein.
[0201] In some embodiments, the sample is obtained from genetically modified cells. In some embodiments, the sample is obtained from cells modified by gene therapy. In some embodiments, the sample is obtained from cells genetically modified to include one or more barcoded payloads. In some embodiments, the sample is obtained from cells genetically modified to express a barcoded payload. In some embodiments, the sample is obtained from cells genetically modified to include one or more barcodes. In some embodiments, the sample is obtained from cells genetically modified to express a barcode. In some embodiments, the sample is obtained from cells genetically modified to include one or more binders. In some embodiments, the sample is obtained from cells genetically modified to express a binder.
[0202] In some embodiments, the sample is obtained from a genetically modified animal. In some embodiments, the sample is obtained from an animal modified by gene therapy. In some embodiments, the sample is obtained from an animal genetically modified to include one or more barcoded payloads. In some embodiments, the sample is obtained from an animal genetically modified to express a barcoded payload. In some embodiments, the sample is obtained from an animal genetically modified to include one or more barcodes. In some embodiments, the sample is obtained from an animal genetically modified to express a barcode. In some embodiments, the sample is obtained from an animal genetically modified to include one or more binders. In some embodiments, the sample is obtained from an animal genetically modified to express a binder.
[0203] Fingerprint: In particular, the systems and methods described herein identify the advantages of nucleic acid sequencing techniques and effectively apply them to protein detection and measurement methods. For example, the methods described herein can use several binders with known binding specificities and affinities for different barcodes, which can be expressed on binders and mixed together in a single pool. When mixed with a pool of barcoded proteins (i.e., proteins each associated with a barcode as described herein), each binder expressed on the binder binds to one or more barcodes in the pool with known but varying affinities. Such a spectrum of affinities of a given barcode for one or more binders can be determined through NGS, resulting in a distinct distribution of binder counts for a given barcode, referred to herein as a "barcode fingerprint." In some embodiments, the collective barcode fingerprint of a set of barcodes is referred to herein as a "fingerprint matrix." Similarly, a spectrum of binder affinities for various (e.g., one or more) barcodes is referred to herein as a "binder fingerprint." In some embodiments, using the provided technology, the presence of a barcoded protein(s) can be detected, for example, by extracting and sequencing the associated nucleic acids (e.g., detectable nucleic acids (e.g., DNA sequences, RNA sequences, etc.)) of a population of binders (e.g., phage) that bind to the barcoded protein(s) in the complex solution. That is, for example, in some embodiments, the presence of a protein in a complex solution is determined not by a single binder, but by a specific combination of multiple binders that bind to the barcodes associated with that protein in a fixed, known proportion.
[0204] As disclosed herein, fingerprinting has many advantages. In some embodiments, the fingerprinting approach to detection allows for noise reduction. For example, the use of multiple binders to detect barcodes in a complex solution introduces redundancy into the detection method, thereby reducing signal noise. Additionally, another advantage of the "fingerprinting" approach is that partial non-specificity within the binders (e.g., to barcodes other than the barcode of interest being detected) can be tolerated and compensated for by computational prediction methods.
[0205] In some embodiments, binder sequences can be modified to alter the fingerprint. In some embodiments, binder sequences can be modified to improve the fingerprint.
[0206] The barcode fingerprints described herein for a given barcode may include affinity information for the given barcode to one or more binders. In some embodiments, the barcode fingerprint may include affinity information for the given barcode to one binder. In some embodiments, the barcode fingerprint may include affinity information for the given barcode to at least one binder. In some embodiments, the barcode fingerprint may include affinity information for the given barcode to 2, 3, 4, 5, 10, 20, 25, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10,000 or more binders. In some embodiments, a barcode fingerprint may contain affinity information for up to 10,000 binders of a given barcode.
[0207] The binder fingerprints described herein for a given binder may include affinity information for the given binder to one or more barcodes. In some embodiments, the binder fingerprint may include affinity information for a given binder to one barcode. In some embodiments, the binder fingerprint may include affinity information for a given binder to at least one barcode. In some embodiments, the binder fingerprint may include affinity information for a given binder to 2, 3, 4, 5, 10, 20, 25, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10,000 or more barcodes. In some embodiments, a binder fingerprint may contain affinity information for up to 10,000 barcodes of a given binder.
[0208] As discussed herein, in some embodiments, multiple barcode fingerprints for a set of barcodes may be grouped together, referred to herein as a "fingerprint matrix." In some embodiments, a fingerprint matrix may include one barcode fingerprint. In some embodiments, a fingerprint matrix may include at least one barcode fingerprint. In some embodiments, a fingerprint matrix may include 2, 3, 4, 5, 10, 20, 25, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10,000 or more barcode fingerprints. In some embodiments, a fingerprint matrix may include up to 10,000 barcode fingerprints.
[0209] The techniques described herein enable the generation and characterization of a unique fingerprint for each barcode. This enables, for example, the availability of methods for payload (i.e., target (e.g., protein)) detection that may not require orthogonality between barcode-binder pairs. In some embodiments, the barcode-binder pairs may be orthogonal. In some embodiments, the barcode-binder pairs may not be orthogonal. As will be apparent to one of skill in the art, the barcode-binder pairs described herein offer the advantage of being more robust because the availability of a unique fingerprint makes non-specific binding less of a concern, a major advantage in complex environments (e.g., serum, blood, etc.).
[0210] Decoding: analysis A key component of the present invention is the method used to estimate relative or absolute protein concentrations from DNA sequencing of binders. In the present invention, DNA sequences are translated in silico into amino acid sequences corresponding to each binder and tabulated to generate a table of binder counts. The binder count table measured singly for any given barcode is hereinafter known as the barcode's "fingerprint." When applying the present invention to an unknown mixture of barcoded payloads, the relative or absolute abundance of each barcode is determined by comparing the binder count table to the predetermined fingerprints of each barcode and applying the computational prediction method described below. In some embodiments, the binder count table for a mixture of m unknown barcodes is assumed to be a linear combination of their respective fingerprints, and the coefficients of the linear combination are estimated by least-squares fitting of the equation Ax = b, where A is an n x m matrix of fingerprints, b is a length-n vector of binder counts, and x is an unknown length-m vector of the abundance of each barcode. In some embodiments, the abundance of each of the barcodes is estimated using a Bayesian method, whereby a suitable prior probability distribution for barcode abundances is assumed, a likelihood ratio for the observed count table given the barcode abundances is calculated from a model of uncertainty in the experimental system, and a posterior probability distribution is estimated using the product of the prior and the likelihood ratio. In some embodiments, the posterior distribution is estimated using a Monte Carlo sampling method. In some embodiments, the maximum value of the posterior distribution is determined by a computational optimization procedure. In some embodiments, the binder count table is assumed to be a nonlinear function of the abundances of the various barcodes to account for saturation of specific barcode-binder interactions or competition between distinct barcodes or distinct binders.
[0211] In some embodiments, the relative proportions of binder counts are compared directly to determine the relative proportions of barcodes. In some embodiments, sequences of known abundance are used to determine the absolute abundance of a given binder that is mixed into the experiment and used to estimate the absolute concentration of the barcode.
[0212] V. Use Evaluate the payload: The techniques described herein can be used to detect, evaluate, and / or characterize payloads (e.g., proteins). In some embodiments, the provided techniques can be used, for example, to assay payloads in complex environments (e.g., serum, blood, tissue, etc.). In some embodiments, the payload can be a protein. In some embodiments, the payload can be a therapeutic protein.
[0213] As described herein, a payload can be associated with a barcode (i.e., a barcoded payload). In some embodiments, the barcoded payload can be assayed using a binding agent (e.g., phage with binders expressed thereon) using methods described herein. In some embodiments, the barcoded payload can be captured (e.g., using affinity reagents) onto a surface (e.g., a bead or plate). In some embodiments, the barcoded payload can be immobilized for barcode assay. In some embodiments, the barcoded payload is contacted with one or more binders and subjected to decoding as described herein.
[0214] In some embodiments, the payload may be detected, assessed, and / or characterized in vitro, hi some embodiments, the payload may be detected, assessed, and / or characterized in vivo.
[0215] In some embodiments, the methods described herein determine simultaneous in vivo assessment of a payload's phenotype in multiple tissues. In some embodiments, the phenotype includes biodistribution information associated with the payload. In some embodiments, the phenotype includes pharmacokinetic (clearance) information for the payload. In some embodiments, the phenotype includes half-life information for the payload. In some embodiments, the phenotype includes tissue-mediated drug disposition (TMDD) for the payload. In some embodiments, the phenotype includes characteristics related to the in vivo stability of the payload. In some embodiments, multiple phenotypes of a payload may be determined simultaneously.
[0216] kit: The technology described herein may be provided in the form of a composition. For example, in some embodiments, a composition may include one or more elements (e.g., nucleic acids, amino acids, etc.) for producing or generating one or more barcodes and / or binders as described herein. In some embodiments, a composition may include one or more elements for producing or generating a set of barcodes. In some embodiments, a composition may include one or more elements for producing or generating a set of binders. In some embodiments, a composition may include one or more elements for producing or generating a pool of barcode-binder pairs. In some embodiments, a composition may include one or more elements for producing or generating binders (e.g., phage-expressed binders). In some embodiments, a composition may be a barcode composition. In some embodiments, a composition may be a binder composition. In some embodiments, a composition may be a barcode-binder composition. In some embodiments, a composition may be a binder composition. In some embodiments, a composition may include one or more of a barcode, a binder, a binder, and / or components thereof. In some embodiments, the composition may include a barcode, a binder, a binding agent, and / or one or more sets / pools of components thereof.
[0217] Provided herein are compositions comprising barcodes, binders, binding agents, or components thereof. In some embodiments, the compositions comprise barcodes, binders, binding agents, components thereof, and / or combinations thereof that have been assessed, identified, characterized, or assayed using the methods described herein. In some embodiments, the compositions provided herein comprise one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more barcodes, binders, binding agents, components thereof, and / or combinations thereof that have been assessed, identified, characterized, or assayed using the methods described herein.
[0218] In some embodiments, the compositions provided herein comprise two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more nucleic acid or amino acid sequences listed in Table 1 or Table 2.
[0219] The compositions described herein can be formulated in various forms.For example, in some embodiments, the compositions described herein can be formulated in powder form (e.g., lyophilized).In some embodiments, the compositions described herein can be formulated in liquid form.
[0220] In some embodiments, a composition for use according to the present disclosure is a pharmaceutical composition, e.g., for administration (e.g., topical, oral, subcutaneous, intravenous, intramuscular, intracerebral, intrathecal, rectal (e.g., rectal intubation), intraocular, intravitreal, or suprachoroidal administration) to a subject (e.g., a mammal (e.g., a human)). In some embodiments, such a composition is administered to a subject to detect, characterize, and / or evaluate one or more attributes of one or more payloads administered or to be administered to the subject. Pharmaceutical compositions typically include an agent to be administered (e.g., a barcode, binder, binding agent, and / or components thereof) and a pharmaceutically acceptable carrier. Certain exemplary pharmaceutically acceptable carriers include, for example, saline, solvents, dispersion media, coatings, antibacterial and antifungal agents, isotonic and absorption delaying agents, etc., compatible with pharmaceutical administration. Pharmaceutical compositions are typically formulated to be compatible with their intended route of administration. Examples of routes of administration include topical, oral, subcutaneous, intravenous, intramuscular, intracerebral, intrathecal, rectal (eg, rectal intubation), intraocular, intravitreal, or suprachoroidal administration.
[0221] In some embodiments, pharmaceutically compatible binding agents, and / or adjuvant materials can be included as part of the pharmaceutical composition. In some particular embodiments, the pharmaceutical composition can contain, for example, any one or more of the following inactive ingredients, or compounds of a similar nature: binders, excipients, lubricants, glidants, or some similar such compound.
[0222] The compositions may be included in a kit, container, pack, or dispenser with instructions for administration (e.g., to a subject) or for use in the methods described herein. In some embodiments, the instructions may include how to reconstitute a powder form composition into a liquid form composition for further use. In some embodiments, the kit may include instructions that allow a user to generate a new set of binders for a new set of barcodes. In some embodiments, the kit includes a set of instructions for sequencing one or more phage particles bound to one or more barcodes.
[0223] In some embodiments, the kit includes information indicating the peptide barcode of each binder, hi some embodiments, the kit includes a computer readable program for decoding the sequencing data.
[0224] In some embodiments, the kit includes reagents for expressing binders on phage particles. In some embodiments, the kit includes nucleic acids encoding one or more barcodes. In some embodiments, the kit includes nucleic acids encoding one or more binders.
[0225] Those skilled in the art will understand, upon reading this disclosure, that in some embodiments, the compositions described herein (e.g., binder compositions, barcode compositions, binding agent compositions, etc.) can be or can include one or more cells, tissues, or organisms (e.g., plant or microbial cells, tissues, viruses, or organisms) that produce (e.g., produced and / or are producing) the associated binder, barcode, and / or binding agent described herein.
[0226] Those skilled in the art will understand that in some embodiments, techniques for preparing compositions and / or preparations (and particularly techniques for preparing pharmaceutical compositions) may include one or more steps of evaluating or characterizing the compound, preparation, or composition, e.g., as part of quality control. In some embodiments, if assayed material does not meet predetermined specifications for the relevant evaluation, it is discarded. In some embodiments, if such assayed material meets predetermined specifications, it continues to be processed as described herein.
[0227] In some embodiments, the composition is tailored to a particular subject (e.g., a particular mammal, e.g., a patient). In some embodiments, the composition is specific to the payload being evaluated for an individual subject (e.g., a mammal (e.g., a human, a mouse, etc.)). In some embodiments, the composition is specific to the payload being evaluated for an individual subject (e.g., a mammal (e.g., a human, a mouse)). In some embodiments, the composition is specific to the payload for a population of subjects (e.g., a mammal (e.g., a human, a mouse, etc.)). A population of subjects can include, but is not limited to, a family, subjects in the same geographic location (e.g., a neighborhood, city, state, or country), subjects with the same disease or condition, subjects of a particular age or age range, subjects consuming a particular diet (e.g., a food, food source, or caloric intake). [Example]
[0228] Example 1: Identification of barcodes, corresponding binders, and fingerprint determination This example demonstrates how to identify barcodes, corresponding binding agents (e.g., binders expressed on the binding agents), determine fingerprints (e.g., barcode fingerprints), and use the information to determine the proportion of barcodes in a given mixture. The resulting material can then be used to measure and quantify different payloads.
[0229] Barcode library design and synthesis: Barcode sequences were designed containing specific sequence motifs that were thought to fold into a given helix or loop structure. All sequences from the Protein Data Bank (PDB) were downloaded along with their corresponding secondary structure predictions. Sequences were selected and subset from the full sequence if they met the criteria of being a continuous helix or loop sequence between 8 and 25 amino acids in length. 100,000 random subsets of peptide sequences matching this criteria were then ordered as oligo pools containing constant overhangs and type IIS sites for cloning into vectors (see Figure 1B).
[0230] Cloning the barcode library into an expression plasmid: The designed barcode pool was cloned into a pET expression vector to obtain barcodes attached to payload proteins. A plasmid containing 6xHIS-HALO-TEV-LN-IIS-LC was constructed to enable direct cloning of the oligo pool via Golden Gate assembly. LN and LC represent constant overhangs within the oligo pool used for ligation (Figure 1B). 1 μg of vector was predigested with BsaI at 37°C and purified. The oligo pool was amplified by polymerase chain reaction (PCR). A 1:10 molar ratio of purified vector and insert was added to the Golden Gate assembly reaction using NEB Golden Gate Assembly Mix (catalog E1601S), incubated at 37°C for 1 hour, and then heat-killed at 70°C for 5 minutes. This material was then purified, drop-dialyzed into pure H2O for 60 minutes, and electroporated into electrocompetent BL21 bacteria (lucigen). Serial dilutions were performed, and individual colonies were recovered. Colonies were then picked, grown in medium containing 1% glycerol and 100 μg / ml carbenicillin, and stocked in 20% glycerol at −80° C. After cloning, linkers are used to generate constructs containing barcodes attached to proteins (FIG. 1A).
[0231] Expression of individual barcodes: Expression was performed using either in vitro transcription / translation (IVTT) or BL21 induction. For IVTT, PCR was performed directly from the glycerol stock by adding primers specific to the T7 and T7 terminator sequences of BL21. The resulting amplifier contained the T7 and T7 terminator for expression, making the protein 8xHIS-HALO-TEV-LN-barcode-LC. Approximately 1 μL of PCR product containing 50 ng of DNA was added to a 10 μL IVTT reaction using a NEBPure (catalog number: E6800S) assembled according to the manufacturer's instructions. The reaction was then incubated at 37°C for 4 hours. For expression in Escherichia coli (E. coli), cultures were grown to an OD of 0.5 at 37°C, then induced using isopropyl β-d-1-thiogalactopyranoside (IPTG) and grown overnight at 25°C. The next day, cells were lysed using sonication in lysis buffer, the lysed material was separated from the inclusion bodies by centrifugation, and the protein-containing supernatant was collected and purified using Ni-NTA affinity chromatography and stored for future use.
[0232] Capture of individual barcodes on HALO magnetic beads: 10 μL of IVTT was diluted to 50 μL in PBS supplemented with 1 mg / ml BSA. 30 μL of Halo-tag magnetic beads (catalog: G7281) were added to this mixture, which was then shaken at 400 rpm for 2 hours and then incubated overnight at 4°C with shaking. The beads were captured on a magnetic stand, and the supernatant was removed. The beads were then washed twice with PBS-T containing 0.1% Tween 20 (PBS-T). A schematic diagram of the captured barcode is shown in Figure 3B.
[0233] Construction of a phage display library containing binders with different affinities for the barcode: Binders with strong affinity for at least one barcode were generated via methods known to those skilled in the art (e.g., phage display, hybridomas, etc.). These binders were then displayed on phage as scFv fragments fused to the m13 gene 3 protein (g3). Briefly, oligos containing the scFv binding sequences were generated by DNA synthesis. The oligos were cloned into a plasmid containing the constant region of the scFv connected to G3 via a G4S linker (SEQ ID NO: 8399) and a myc tag. 30 μg of the library was electroporated into TG1 (lucigen) and plated onto several 25 mm plates containing 100 μg / ml carbenecrine and 1% glucose. Dilutions of the electroporation were plated for diversity analysis. The Q-trays were scraped and stocked with glycerol. To generate phage, a 2 L culture was inoculated at an OD of approximately 0.05 and grown at 37°C with 100 μg / ml carbenicillin and 1% glycerol to an OD of 0.5. At OD of 0.5, helper phage was added at a phage:cell ratio of 10:1 and incubated with shaking at 250 rpm. After 1 h, shaking was reduced to 150 rpm, and the temperature was increased to 30°C for overnight incubation. The next day, phage were prepared via PEG precipitation (Barbas et al. 2001), resuspended in 10 mL, and titrated. Phage were stored at 4°C until use.
[0234] Assessment of phage-binder-barcode interactions: Ten microliters of the phage library prepared using the method described above was added to the captured barcodes and incubated at room temperature for 2 hours to allow phage binding to the barcodes (Figure 3C). After incubation, the mixture was washed three times by successively transferring the magnetic beads to fresh PBS-T. The phage were resuspended in PBS containing TEV protease + 0.1% DTT and eluted from the beads by incubation at 37°C for 30 minutes. The beads were collected on a magnetic stand, and the supernatant was collected. Trypsin was added to the supernatant and incubated for an additional 30 minutes. To amplify the phage, the supernatant was added to 50 mL of TG1 E. coli grown at 37°C in 2xYT to an OD of 0.5 and incubated at 37°C for 1 hour with shaking. 100 μg / ml carbenicillin and 1% glucose were then added for 1 hour of incubation under the same conditions. Helper phage (catalog: PH050L) was added and further incubated for 1 hour at 37° C. The culture was centrifuged, placed in fresh medium containing 100 μg / ml carbenicillin and 50 μg / ml kanamycin, and incubated overnight at 30° C. Subsequently, the culture was centrifuged at 4000 g for 20 minutes, and the phage-containing supernatant was collected.
[0235] Analysis and establishment of a fingerprint for a given barcode: After individual selections from the original phage pool for each barcode, the selectivity of phage-scfv (i.e., phage-binders) was analyzed via next-generation sequencing (NGS). Phages were lysed by heating at 98°C for 10 min, and the resulting genomes were PCR-enhanced using primers flanking the CDR regions of both the heavy chain CDR3 (5-prime) and light chain CDR3 (3-prime). A second round of PCR was performed to add the necessary Illumina sequences (i5 / i7, sequencing primer binding regions) for next-generation sequencing (NGS). The resulting DNA was pooled, quantified, and subjected to next-generation sequencing using an Illumina instrument. This process is illustrated in Figure 2 and Figures 3A-3C.
[0236] NGS reads are demultiplexed using the Illumina software bcl-convert so that each final .fastq contains DNA sequences from a given phage CDR3 pair corresponding to the output from a given barcode well. The corresponding CDR3 sequences are then counted using a computer program, revealing the distribution of binders present in a given barcode. The barcode fingerprint corresponds to the vector of counts for each scFv binder within a given pool. Each fingerprint is the median of n=3 individual barcode replicates. The process and resulting fingerprint for a single barcode using a pool of phage binders are shown in Figure 4.
[0237] Determining the proportion of barcodes in a given mixture using the fingerprint matrix: Once the fingerprints for each barcode in the set of barcodes were determined, the proportion of barcodes in the unknown sample was determined in the following manner. Binder-barcode interactions were assessed as described above, and the resulting NGS readouts were fitted to a linear combination of the known fingerprints via least squares fitting. That is, the coefficients of the linear combination were selected by minimizing the sum of squares of the differences between the measured NGS counts and the expected NGS counts for all bound species, as described in Example 8. The expected NGS counts were given by the matrix product of the fingerprint matrix and the set of barcode abundance coefficients. Once the coefficients were obtained, they were normalized to sum to 1 to obtain a proportion.
[0238] Decoding an equal proportion barcode mixture to assess fingerprint scaling: To determine scaling issues that may arise due to varying affinities between binders and barcodes, a scaling factor was generated by measuring barcodes mixed in equal proportions. Briefly, all validated barcodes were mixed at a uniform concentration after production. Phage-binder interactions were evaluated as described above, and the resulting phages were subjected to NGS. Using the fingerprints determined for each individual barcode above, the proportion of barcodes in the mixture was estimated by least-squares regression, as described in Example 8. The proportion predicted using least squares forms the basis of a scaling factor (sf), where sf = 1 / p (p = predicted proportion). The process is described in Figure 5.
[0239] Evaluation of a mixture of barcodes of known proportions: The barcodes generated above were mixed in known proportions and subjected to evaluation to determine their accuracy. Analysis of phage-binder (i.e., binding agent) interactions was then performed as described above for barcodes mixed in different proportions (Figure 6A). NGS counts for each binder within the pool were counted. Least squares was used to determine the proportion of barcodes that gave the fingerprint constructed above using a single barcode. Predictions from the least squares analysis were then rescaled using the established scaling factor p' = p * sf, where p' is the new prediction, p is the original prediction, and sf is the established scaling factor. The new predictions were then renormalized to sum to 1.
[0240] Figure 6B shows the accuracy of this method's proportion measurement. Six different barcodes were measured and normalized via a scaling factor to estimate relative barcode proportions with a global Pearson correlation of 0.95 across all measurements. Measurements spanned a 100-fold gradient of barcodes. Figure 6C shows a plot of NGS count values normalized to counts per million for each single barcode measurement, as well as the mixture used to predict the relative abundance of each barcode within the mixture. Rows are experiments; therefore, all values in a row are generated from a single .fastq file, and columns are binders.
[0241] Example 2: In vitro detection of payloads in mixtures using the binder-barcode platform This example demonstrates how the binder-barcode platform described herein can be used to determine the presence or absence of a given payload in a mixture.
[0242] The barcode generated in Example 1 was transferred to a new payload using DNA cloning. Briefly, the barcode was amplified from pET 6xHIS-HALO-TEV-LN-barcode-LC so that the LN-barcode-LC portion was amplified. The barcode insert was cloned using Gibson into a new pET vector containing 6xHIS-payload-LN-barcode-LC, whose payload was the novel protein of interest. The payload-barcode was generated in E. coli as described above. The barcoded payload protein was then purified via affinity chromatography using NiNTA, washed in 500 mM NaCl, 50 mM Tris-HCl, 50 mM imidazole, and eluted using 500 mM imidazole. The barcoded payload protein was then subjected to decoding using the phage binder library described in Example 1. Figures 6, 9, and 12 show experimental setups for detecting payloads using previously generated barcodes in different contexts; therefore, the results of such experiments show that these barcodes contain detection properties that generalize across various numbers of barcodes in the pool.
[0243] In the experiment depicted in Figure 9, six unique barcodes (BC1, BC2, BC3, BC4, BC5, and BC6) were mixed in known proportions, contacted with a binder, and subjected to decoding as described herein (Figure 9A). Two barcodes were retained as experimental negative controls, allowing for prediction of these barcodes, thereby enabling the determination of background prediction. Figures 9B and 9C show data on the accuracy of the decoding procedure over a 10-fold concentration range for six unique barcodes. Figure 9B shows a plot of the observed (input) and measured data obtained after decoding one mixture of known barcode concentrations. The input known concentrations (left bar) are shown next to the predicted / measured data (right bar) for each barcode across three replicates. Figure 9C shows a plot of the observed (input) and measured data obtained after decoding five different mixtures of known barcode concentrations (i.e., pools 1–5). The input known concentrations (left bar) are shown next to the predicted / measured data (right bar) for each barcode across three replicates.
[0244] In the experiment described in Figure 12, the quantification of 24 barcodes contained within a single mixture was determined. Figure 12A shows a graphical depiction of the experiment. Of the 24 total barcodes, the algorithm was able to predict, with 10 present at equal concentrations in the mixture. The remainder were excluded from the pool, but the prediction was computationally acceptable. Three separate pools covering all possible barcodes were measured in duplicate. These pools were mixed and then captured on HALO beads as described in Example 1. The immobilized barcoded payloads were then contacted with a pool of binders (i.e., binders with binders expressed on them) and decoded into CDR3 sequence counts as described in Example 1. The CDR3 sequence counts determined via NGS were then used to determine the presence and total concentration of each barcoded payload in the sample via decoding (see Example 8). Figure 12B shows the prediction for the first pool, where the input concentration (left bar) and measured concentration (right bar) are plotted for each barcode in the pool. Figure 12C plots the predictions for each of the three pools. As in Figure 12B, the bar graph plots input concentration on the left and measured concentration on the right. Figure 12D shows the barcode fingerprints for each of the 24 barcodes used to computationally determine the relative abundance of the barcodes in each of the three pools. The columns represent barcode fingerprints, and the rows represent binder fingerprints. Figure 12E shows binder counts from the three pools used to computationally determine pool proportions. The rows are binder counts, the columns are pools, and each cell is the binder count within a particular pool.
[0245] Example 3: In vitro assessment of payload stability within a pool of payloads using the binder-barcode platform This example demonstrates how barcode decoding can be used to determine the general aggregation tendency of several payloads within a pool.
[0246] A purified pool of barcoded payloads is generated using the method described in Example 2. The purified pool is then subjected to size exclusion chromatography using standard methods. Different fractions are collected (corresponding to monomeric vs. aggregated payloads). The general presence or absence of a given barcoded payload within the purified pool is unknown. Separated fractions containing unknown abundances of each barcoded payload are then immobilized on beads or immunosorbent assay plates, contacted with a pool of binders (i.e., binding agents with binders expressed thereon), and decoded into CDR3 sequence counts as described in Example 1. The CDR3 sequence counts determined via NGS are then used to determine the presence and total concentration of barcoded payloads in each fraction. The concentrations of barcoded payloads in different fractions are then compared to determine the percentage of monomeric vs. aggregated barcoded payloads within the purified pool.
[0247] Example 4: In vivo evaluation of the pharmacokinetics of payloads within a pool of payloads using the binder-barcode platform This example demonstrates how to determine the overall residence time and clearance time of a given payload contained within a pool of payloads using a mouse model, as demonstrated in FIG.
[0248] Three pools of barcoded payloads were injected into three different groups of mice (n = 3): Pool 1 in Group 1, Pool 2 in Group 2, and Pool 3 in Group 3. Pool 1 contained a single barcoded antibody at 10 mg / kg. Pool 2 contained two barcoded antibodies pooled at equal concentrations and injected at 20 mg / kg. Pool 3 contained PBS only. The injection volume was kept constant at 100 μL per pool. 24 hours after injection, blood was collected from each mouse, and serum was separated. 10 μL of serum was diluted 1:10 in PBS and captured using anti-human IgG magnetic beads (Ray biotech catalog #801-101-1) by incubating overnight at 4°C with mixing at 700 rpm. The immobilized barcoded payload was washed three times with PBS-T to remove all serum proteins not associated with the affinity reagent. The immobilized barcoded payloads were then contacted with a pool of binders (i.e., binders with binders expressed thereon) and decoded to CDR3 sequence counts as described in Example 1. The CDR3 sequence counts determined via NGS were then utilized to determine the presence and total concentration of each barcoded payload in the sample via decoding (see Example 8). The percentage of barcoded payload measured at 24 hours was compared to the injected concentration to determine the relative rate of clearance for each barcoded payload from the organism. In each of the groups, only the injected barcoded antibody was detected with high accuracy via decoding, as evidenced by the graph plotted in Figure 11. Group 2 mice injected with Pool 2, which contains both antibodies at equal concentrations, measured slightly different amounts via decoding in serum at 24 hours. This difference is hypothesized to be due to differences in clearance rates between the two barcoded antibodies. As expected, the control group showed very little antibody.
[0249] Example 5: In vivo assessment of payload biodistribution using the binder-barcode platform This example demonstrates how a mouse model can be used to determine the overall distribution of a barcoded payload across a diverse set of tissues.
[0250] The purified payload pool is intravenously injected into BALB-6 mice. At least 24 hours later, different tissue samples, such as liver, lung, and brain, are harvested from the organism. The tissues are then processed into a single-cell suspension via vigorous shaking with beads. The suspension is then lysed using a lysis buffer to release the barcoded payload (e.g., barcoded protein) contained within the tissue. The lysed suspension is then purified using a universal tag affinity reagent contained within the payload to isolate the barcoded payload. The purified barcoded payload is then immobilized, and barcode decoding is performed according to the method described in Example 1. The CDR3 sequence counts determined via NGS are then used to determine the presence and total concentration of each barcoded payload in each sample. Payload abundance across different tissue samples is then compared to determine the percentage of each payload contained in each tissue. This data can then be used to select the best payload with specific biodistribution characteristics.
[0251] Example 6: In vitro demonstration of recovery of a known mixture of unmodified antibodies This example demonstrates how the protein quantification invention described herein was used to quantify a known mixture of antibody proteins that did not have barcodes attached.
[0252] Briefly, scFv binders to antibodies were generated using methods known to those skilled in the art. The binders were then cloned and displayed on phage as described in Example 1. Two antibodies of interest were expressed in CHO cells and purified from the culture medium using protein A affinity chromatography. The antibodies were mixed together in known ratios (FIG. 7A). The antibodies were then captured using 50 μL of anti-human Fc magnetic beads and incubated in PBS. The antibodies were then subjected to phage evaluation as described in Example 1, and the relative abundance of each antibody was estimated using the algorithm described in Example 8. A Pearson accuracy of 0.96 was calculated for determining the relative concentrations of these two antibodies in mixtures of various ratios (FIG. 7B).
[0253] Example 7: In vitro demonstration of recovery of a known mixture of antibodies in the presence of serum This example demonstrates how a known mixture of antibody proteins with barcodes contained within internal regions of the protein sequence was quantified in mouse serum using the protein quantification techniques described herein.
[0254] Briefly, antibodies were produced as described in Example 6. A similar experiment was performed, except that after production, the antibodies were mixed with mouse serum, incubated at 37°C for 30 minutes, and then captured using anti-Fc magnetic beads (Figure 8A). The immobilized barcoded antibodies were then contacted with a pool of binders (i.e., binders with binders expressed thereon) and decoded into CDR3 sequence counts as described in Example 1. The CDR3 sequence counts determined via NGS were then used to determine the presence and total concentration of each barcoded antibody in the sample via decoding (Example 8). The relative concentrations of the three antibodies were estimated with a Spearman's error of 0.926 across a mixture of three antibody concentrations, as shown in Figures 8B and 8C. These results demonstrate the ability to use the technology described herein to accurately rank the antibodies present in a sample after incubation with serum, i.e., in a complex environment.
[0255] Example 8: Detailed description of the decoding algorithm used to infer barcode abundance How to estimate barcode volume: Before decoding unknown samples, the set of barcodes and their interactions with the binder pool (i.e., phage binder pool) must first be characterized. This is done by decoding a set of known samples under known conditions. The binder pool and experimental conditions are fixed across all samples.
[0256] To characterize a set of barcodes, a set of fingerprints is measured, one for each barcode. A fingerprint represents the ideal readout of an individual barcode. Roughly speaking, it is the spectrum of affinities between a given barcode and all binder species in the pool. A fingerprint can be estimated by decoding multiple identical samples containing purely one barcode, averaging the replicate readouts together, and rescaling accordingly. Alternatively, fingerprints can be learned by decoding samples containing a mixture of known barcodes and appropriately deconvolving to isolate individual fingerprints. Together, the fingerprints of a set of barcodes are known as a "fingerprint matrix."
[0257] Once the fingerprint matrix for a set of barcodes has been determined, it can be used to infer the barcode composition in an unknown sample. The decoding algorithm accomplishes this by matching the readout of the unknown sample to a linear combination of the fingerprints, as described in more detail in the algorithm section herein. A key assumption of the algorithm is that the decoding process is linear; if a sample contains two barcodes mixed in equal proportions, the readout is assumed to be equal to the sum (plus noise) of the fingerprints of the two barcodes. More generally, the readout of a mixture of barcodes is assumed to be the sum of the fingerprints of each barcode, appropriately weighted by its prevalence in the mixture. This assumption has been empirically found to be true.
[0258] Barcode quantification tasks come in a variety of difficulty levels. These, from easiest to most difficult, include the following: Binary classification: Detects the presence or absence of a barcode in a sample Rank order quantification: Rank barcodes from most common to least common within a sample Relative quantification: determining the ratio between barcodes in a sample Absolute quantification: Determine the absolute amount of each barcode in a sample
[0259] In this example, absolute quantification is discussed in more detail.
[0260] Mathematical model of decoding: The decoding process can be represented by the following mathematical model: x is a length-n vector representing the input sample, where each entry is the amount of barcode species in units of ng, y is a length-m vector representing the bound binder fraction, where each entry is the number of particles of a binder species in units of pfu, and z is a length-m vector representing the NGS readout, where each entry is the number of counts of a binder species.
[0261] The bound binder fraction is modeled as a linear combination of the fingerprints, and the NGS readout is modeled as the bound binder fraction multiplied by a conversion factor,
number
[0262] This model assumes that the binding between barcodes and binders is linear. In other words, if a sample contains a mixture of barcodes, the readout is assumed to be equal to the sum of the fingerprints of the individual barcodes, weighted by their relative barcode abundance. In Appendix A, we provide a detailed biophysical model that justifies the linear assumption under one important condition: the amount of available binder cannot be significantly depleted by binding to barcodes in the sample. Thus, as in a typical immunoassay, the binder must be in excess and not be the limiting reagent.
[0263] Fingerprint matrix A ji : Each column of the mn fingerprint matrix is a fingerprint. Each fingerprint represents an ideal, properly normalized readout of a pure barcode. Entry A ji represents the contribution of the jth binder to the fingerprint of barcode i.
[0264] The fingerprint of a barcode depends on the binding affinity to all binders in the pool, as well as the relative abundance of each binder species in the pool. Furthermore, the fingerprint is sensitive to the binding, equilibration, and elution steps. In the simplest case, the fingerprint matrix is given by: A ji =d j / K ji In the formula, d j is the concentration of binder j in the binder pool, and K ji is the dissociation constant of the complex between barcode i and binder j (see Appendix A). In more complex cases, A ji can also include effects such as adhesion to surfaces, debonding during cleaning steps, etc.
[0265] The matrix product of A and the barcode mixture x gives the composition of the ideal bound binder fraction (i.e., in the absence of noise), in units of number of phage particles.
[0266] The fingerprint matrix can be determined from measuring the readouts of multiple known samples. Multiple replicates are performed and averaged over noise. Furthermore, the fingerprints are appropriately scaled relative to each other or to an absolute reference (see normalization section).
[0267] Conversion factor s j : The post-binding step introduces a conversion factor between the number of bound phage particles and the number of NGS reads. j In the simplest case, s j is the same for all binder species and represents the global normalization,
number
[0268] If the conversion factor is the same for all binder species, it is a single number that must be determined for each sample, which can be done by using a DNA sequence spiked in at some point in the process (as described herein).
[0269] Noise ε: The noise sources are denoted by the terms ε1 and ε2. In the absolute simplest case, ε1 does not exist and ε2 is a Gaussian noise with fixed variance, in which case prediction can be made by ordinary least-squares regression. In practice, noise arises from multiple non-Gaussian sources, as detailed in the sections above. These include lognormal noise involved in exponential steps such as phage propagation and PCR amplification, Poisson noise due to finite sequencing depth (and possibly stochasticity in binding / elution at low concentrations), and Gaussian noise from all sorts of other processes such as sample degradation.
[0270] Converting read counts to phage counts: One feature of NGS is that the readout is a relative measurement, which gives the ratio of abundance between different binder species, but not necessarily the absolute concentration of the binder species. To obtain an absolute readout, the raw readout must be divided by a conversion factor between the number of bound phage particles and the NGS read count.
[0271] Without knowing the conversion factor, it is only possible to determine the relative abundance (i.e., percentage) of the barcode in the sample. Absolute quantification requires a standard of known concentration (either barcode or binder) spiked into the process.
[0272] Absolute quantification methods: Spiking the eluate with phage ladder: One method of normalization is to use a known concentration, y スパイクイン The key to success is to add a unique binder species to the eluate at 100 kJ / ml. This reference species must be distinct from any existing binders in the pool. In subsequent steps (growing the eluted phage, extracting DNA, PCR, and sequencing), the reference phage is (ideally) amplified by the same factor as the other phages in the pool, i.e., by 1 / 2 the factor of 1.
number
[0273] The conversion factor can be estimated by dividing the number of reads corresponding to the reference phage by the (known) concentration at which it was added, i.e.,
number
[0274] A generalization is to spike multiple binder species into the eluate. Each reference species can be spiked at a different concentration. If the concentrations are evenly spaced, this forms a "phage ladder," similar to the ladders used in gel electrophoresis. To estimate the conversion factor, the read counts of each reference sequence can be compared to the (known) concentration at which it was added to the eluate. Averaging across species then yields a more accurate estimate of the conversion factor.
[0275] Spiking barcodes into samples: Alternatively, a reference barcode of known concentration can be added to the sample at the beginning of the decoding process. This reference barcode must be distinct from any existing barcodes in the sample. The decoding algorithm can use the raw readout to determine the proportion of all barcodes in the sample, including the reference barcode and the sample barcode. A barcode conversion factor can be determined by dividing the reference barcode concentration by its predicted proportion. Multiplying all predicted proportions by this factor gives the absolute abundance of the barcode.
[0276] Note that this method is only applicable to decoding unknown samples after a properly normalized set of fingerprints has been determined.
[0277] Fingerprint scaling: An important subtlety is the need to scale readouts even for relative quantification. Specifically, fingerprints for a set of barcodes must be appropriately scaled relative to each other. Intuitively, this is because raw, unscaled fingerprints cannot be directly compared between barcodes; binder read counts have different meanings in the context of fingerprints for different barcodes because each barcode has a different conversion factor between combined binder counts and read counts. To ensure that the fingerprints are measured in the same units, each unscaled fingerprint must be scaled by (the inverse of) its conversion factor.
[0278] To illustrate this, consider the case of two barcodes. Due to differences in binding affinity across the binder pool, assume that the total amount of binder bound to 100 ng of barcode A is 10 times greater than the total amount of binder bound to 100 ng of barcode B. However, due to the nature of the method, the raw readouts will have the same number of reads. In this example, one NGS read in the raw fingerprint of barcode A corresponds to 10 reads in the raw fingerprint of barcode B. The conversion factors are different. Next, consider a sample containing a 1:1 mixture of the two barcodes. Due to the difference in affinity, the bound binder fraction in this sample is 10:1, and therefore the number of NGS reads corresponding to A and B is also in a 10:1 ratio. In other words, the readouts are proportional to 10a + b, where a is the raw readout of A and b is the raw readout of B. Based on this, one would erroneously conclude that A and B are in a 10:1 ratio. To correct for this, the raw fingerprint of barcode A needs to be multiplied by a factor of 10 compared to B to get the correctly scaled fingerprints, a' = 10a and b' = b. Then the correct result is determined when the two barcodes are mixed in equal proportions, and the readout is an equally weighted mixture of the two correctly scaled fingerprints, a' + b'.
[0279] This example shows that the relative scaling factors between barcode raw fingerprints can be determined by measuring the readout of a known mixture of barcodes. If the barcodes are mixed in equal proportions, the composition of the mixture readout is each raw fingerprint, weighted together by their scaling. (Note that this only gives the relative conversion factors between barcodes, not the absolute conversion factors for absolute barcode amounts.)
[0280] Decoding Algorithm: The decoding algorithm has two phases. "Training phase": By measuring the readouts of several known samples, a fingerprint matrix A is created. ji Learn "Test Phase": Measure the readout and A ji Predict the amount of barcode in an unknown sample by comparing
[0281] Training phase: learning the fingerprint matrix In the training phase, the fingerprints of a set of barcodes are determined by measuring the readouts of a set of samples with known composition. The fingerprints must be correctly scaled relative to each other. One method for measuring the fingerprint matrix is outlined below.
[0282] First, a set of samples is prepared, each containing a purely single barcode. Each sample is decoded. The fingerprint for each barcode is estimated by taking multiple replicates of the barcode and averaging the readouts together. The error is reduced when more replicates are averaged. This produces an unscaled fingerprint for each barcode.
[0283] Next, the fingerprints are scaled correctly relative to each other. This is done by multiplying each unscaled fingerprint by a scaling factor. To determine the scaling factor for each barcode, a sample consisting of all barcodes mixed in equal proportions is decoded. In theory, this readout of this sample should be the sum of the normalized fingerprints of all barcodes, weighted equally. However, when fitting this readout of the mixture to the set of unnormalized fingerprints determined from the previous step, the barcode weights are not equal. The coefficient to the linear fit is precisely the coefficient by which each barcode's fingerprint should be multiplied to obtain an accurately normalized fingerprint. By averaging multiple iterations of this together, a more accurate estimate of the scaling factor can be determined.
[0284] This method is described in further detail below.
number
number
number
number
number
number
number
number
number
[0285] Note that there are other methods for measuring the fingerprint matrix. This method only uses a single barcode sample to determine the (unnormalized) fingerprint and a mixture sample to determine the scaling factor. More sophisticated methods could use heterogeneous mixtures to determine the scaling factor and / or use information in these samples to better estimate the fingerprint (rather than just learning the scaling factor).
[0286] Testing phase: Predicting the amount of barcodes in unknown samples In the testing phase, we are provided with a readout y of an unknown sample x and aim to infer its composition.
number
number
[0287] The fitting is done by the coefficient x j This is done by selecting a set of σ, which minimizes a loss function. The loss function measures the deviation between the expected readout and the measured readout. The expected readout is calculated based on the determined fingerprint and the proposed mixture coefficients, using the matrix product A x In the simplest case, the loss function is the sum of squared errors,
number
number
number
number
[0288] Note that the L2 loss function above is the simplest case: it is the negative log-likelihood (proportional) when ε1 is absent (no noise in the combination) and ε2 is Gaussian noise with fixed variance. Other loss functions can be chosen to model more realistic forms of noise.
[0289] Appendix A: Biophysical models of binding processes Let n be the number of barcoded protein species in the sample,
number
number
number
number
[0290] For any given decoder species, the amount that binds to and is subsequently eluted from the sample is
number
[0291] In the simplest case, we assume that each species is an ideal solute in dilute solution, allowing the binding between the barcode and the decoder to approach thermodynamic equilibrium. At equilibrium, some of the decoder binds to the barcode, but the concentration
number
number
[0292] The equilibrium state is given by minimizing the overall free energy of the system, which has been shown to be equivalent to solving the following set of equations:
number
[0293] The first equation ensures conservation of mass: the total amount of barcode species is the unbound amount plus the amount bound in complexes with all possible decoder partners. The second equation is a similar description for the decoder. The final equation is the definition of the equilibrium constant for the barcode-decoder pair.
[0294] The known value is the equilibrium constant K ij and x i and d j and represent the total concentration of each species of barcode and decoder added to the well, respectively. The unknown variables are
number
[0295] linear approximation In the decoding process, K ij and dj A decoder pool with a fixed value of x is used to generate the unknown barcode composition, x i The observable output of the binding process is the amount of each decoder species that binds to the sample.
number
[0296] In general, the above system of equations is nonlinear, but if certain conditions are met, the combining process can be well approximated by a set of linear equations. If the underlying equations are linear, the decoding process is greatly simplified. A linear system implies at least the following: 1) When the density of a barcode is doubled, all decoders associated with that barcode are correspondingly doubled (no saturation). 2) When two barcodes are mixed together, the decoders bound to the mixture are the sum of the decoders bound to each barcode alone (no conflicts). These two criteria can be thought of as the linearity of single and multiple barcode detection, respectively. They are necessary but not sufficient conditions.
[0297] To see how these criteria work in practice, consider the following binding situation where one or more of the criteria are violated: Assume that the decoder is the limiting reagent in a 1:1 stoichiometry of barcode and decoder. Above a certain barcode concentration, all available decoders become saturated and x i The increase in y j Therefore, to avoid saturation, the barcode concentration should be kept below the Kd of the interaction (or the decoder concentration, whichever is greater (see below)).
[0298] As another example, consider a situation in which two barcodes, A and B, both have affinity for a particular decoder, D, but barcode A has a much stronger affinity than barcode B. In the absence of A, the binding of barcode B has a particular binding curve. However, if enough A is present to deplete most of the available D, the binding of barcode B to the remaining D changes significantly. This is a situation in which A and B compete for the available decoder, D. However, if the binding of A did not deplete the amount of D remaining in the pool, the binding between B and D would not be altered by the presence of A. Competitive behavior only occurs in situations in which the available D is significantly consumed by binding to A (when A is highly abundant and / or A-D affinity is strong). In both examples of nonlinear situations, one or more decoders in the pool are significantly depleted by binding to the barcode.
[0299] These examples show that the coupling process is linear when a small fraction of the decoding pool is coupled to a sample. Indeed, if this condition is met, the above equations can be simplified to a simple linear system. Under this assumption, the amount of uncoupled decoders is
number
number
number
[0300] Example 9: Evaluation of absolute payload abundance of a single test barcode using a reference barcode (or spike-in barcode) This example demonstrates how to determine absolute payload abundance of payloads in a pool using reference barcodes and barcode decoding.
[0301] Figure 10A shows a schematic of the experiment. A single test barcode attached to a payload was assayed at several concentrations ranging from 0 ng / mL to 1250 ng / mL. In the same sample, a "spiked-in" barcode (i.e., reference barcode) attached to the payload was added to each assay mixture at 250 ng / mL. Varying concentrations of the test barcoded payload were contacted with binding agents (i.e., binding agents with binders expressed thereon), and decoding was performed as described herein. Predictions of the reference or "spiked-in" barcode were used to determine the absolute amount of the test barcode, and by extension, the barcoded payload, being measured (see Example 8). Figure 10B shows a plot of the measured absolute amount of the test barcode (right bar) compared to the known input concentration of the test barcode (left bar) for each titration of the test barcode. The Y-axis is the logarithm of the test barcode concentration in nanograms per milliliter (ng / mL). Figure 10C shows the results of absolute concentration determinations for six different barcoded payloads. The plot shows the known input concentrations (left bars) and measured concentrations (right bars) for six different barcoded payloads.
[0302] Example 10: Simultaneous in vivo assessment of payload phenotypes using the binder-barcode platform This example demonstrates how the binder-barcode platform described herein can be used to determine simultaneous in vivo assessment of a payload's phenotype. In some embodiments, the phenotype includes pharmacokinetic (or clearance) data as described herein.
[0303] Figure 13A is a schematic diagram of a method for detecting and / or quantifying and / or characterizing 14 exemplary payloads (e.g., proteins) in a pool using the binder-barcode platform described herein. Fourteen exemplary binder molecules were generated using different barcodes ("binder-barcode particles") described herein. The binder-barcode particles were injected as a pool into wild-type (wt) BALB / c mice. Blood was collected from individual mice (n=3 / time point) at 30 minutes, 6 hours, 24 hours, and 48 hours, and serum was extracted. The binder-barcode particles were captured and subjected to the decoding procedure described.
[0304] Figure 13B shows simultaneous in vivo assessment of payload clearance phenotypes using the binder-barcode platform described herein. Specifically, Figure 13B shows plots of clearance of multiple payloads measured simultaneously and grouped by rate of clearance phenotype (e.g., slow vs. fast clearance). Figure 13B (left) shows controls with known properties measured. Figure 13B (center) shows payloads identified as having slow clearance properties over time. Figure 13B (right) shows payloads identified as having fast clearance properties over time. Data was normalized to 100% of injected volume for each binder-barcode particle.
[0305] Thus, this example confirms that the binder-barcode platform described herein can be used for simultaneous characterization of phenotypes in vivo.
[0306] Example 11: Simultaneous in vivo assessment of payload phenotypes in multiple tissues using the binder-barcode platform Among other things, this disclosure provides insight that on-target, off-tumor toxicity is a biodistribution challenge. This example demonstrates how the binder-barcode platform described herein can be used to determine simultaneous in vivo assessment of payload phenotypes in multiple tissues. In some embodiments, the phenotype comprises biodistribution data described herein. In some embodiments, the phenotype comprises pharmacokinetic (clearance) data described herein. In some embodiments, the phenotype comprises pharmacokinetic and biodistribution data described herein.
[0307] Figure 14A is a schematic diagram of a method for detecting and / or quantifying and / or characterizing 36 exemplary payloads (e.g., proteins) in a pool using the binder-barcode platform described herein. Thirty-six binder molecules were generated with different barcodes ("binder-barcode particles"). The binder-barcode particles were injected as a pool into tumor-bearing NSG mice previously implanted with two tumor cell lines (Tumor 1, Tumor 2). Blood and tumor tissues were collected from individual mice at 30 minutes, 6 hours, 24 hours, and 48 hours (n = 2-4 per time point). In some embodiments, other tissues, including lung, liver, brain, etc., can be collected. Tissues were lysed using a standard lysis buffer. Serum was separated from the blood. Binder-barcode particles were captured from each tissue and subjected to the decoding procedure described herein.
[0308] Figure 14B is a heat map of all binder-barcode particle data collected as described herein. Rows indicate different binder construct identifiers (IDs) correlating to binder-barcode particles tested by this example. Columns indicate data for mice across serum, tumor 1, or tumor 2 time points. Color intensity indicates relative units of drug measured via the decoding procedure described herein. Color intensity indicates a normalized readout of relative concentrations measured via next-generation sequencing (NGS).
[0309] Figure 14C shows a plot of the binder-barcoded particles described by Figure 14B using the decoding procedure described herein. Variation in properties was measured simultaneously. For example, binder-barcoded particle P14_A5 was rapidly cleared from serum with minimal accumulation in tumor 1 or tumor 2, while binder-barcoded particle P17_A10 was cleared more slowly and persisted in tumor 1 over time.
[0310] Thus, this example confirms that the binder-barcode platform described herein can be used for simultaneous phenotypic characterization of payloads across multiple tissue types in vivo.
[0311] Example 12: Simultaneous in vivo assessment of payload phenotypes using the binder-barcode platform This example demonstrates how the binder-barcode platform described herein can be used to determine a simultaneous in vivo assessment of the phenotype of a payload (e.g., a protein). In some embodiments, the phenotype is a half-life measurement as described herein.
[0312] Figure 15A shows plots demonstrating ELISA quantification of two payload groups (Group 1: protein payload without a barcode; Group 2: a pool of eight binder-barcode particles, each containing the same protein payload used in Group 1 but barcoded with a different barcode). Group 1 was injected into a cohort of wild-type (wt) BALB / c mice. Group 2 was also injected into a cohort of wild-type (wt) BALB / c mice. Blood was collected from individual mice at 6, 24, and 48 hours (n = 3 / time point), and serum was extracted. ELISA quantification showed similar measurements between Group 1 and Group 2, confirming that the barcode does not affect payload function. Figure 15B shows plots demonstrating simultaneous and individual evaluation of eight distinct entities within Group 2 using the decoding procedure described herein. Figure 15C shows a comparison of half-life measurements for Group 1 and Group 2. Half-life measurements for Group 1 were quantified using ELISA (dashed line). Half-life measurements for Group 2 were quantified using the decoding procedure described herein (bars). As shown in Figure 15C, the variation in half-life measurements across different binder-barcoded particles within a pool can be resolved using the binder-barcode platform described herein. Furthermore, ELISA does not allow for simultaneous measurement of phenotypes in one experiment.
[0313] Thus, this example confirms that the binder-barcode platform described herein can be used for simultaneous and accurate characterization of phenotypes in vivo. Furthermore, this example confirms that attaching barcodes to payloads using the methods described herein does not destroy the in vivo properties of the payloads.
[0314] Example 13: Sensitivity and Dynamic Range of the Binder-Barcode Platform This example demonstrates how to determine the sensitivity and dynamic range of the binder-barcode platform described herein.
[0315] An array of 96 mixtures containing 10-35 barcoded payloads (e.g., proteins), each with a known concentration between 1 picogram (pg) and 1 microgram (μg), was designed to determine the sensitivity and dynamic range of the binder-barcode platform described herein (Figures 16A-16B). Each data point in Figure 16B represents a comparison of the known concentration of a barcoded payload particle from one of the 96 distinct mixtures to the concentration determined by the decoding procedure described herein. As shown in Figure 16B, payloads were quantified over a 10,000-fold range of concentrations. Figure 16B also shows that payloads down to 0.1 nanograms (ng) were quantified.
[0316] Thus, this example confirms that the binder-barcode platform described herein can be used for simultaneous, sensitive characterization of barcoded payloads over a wide dynamic range of concentrations in diverse mixtures.
[0317] Example 14: Comparison of in vitro and in vivo assessment of payload phenotypes using the binder-barcode platform This example provides insight that in vitro models may not accurately model the in vivo environment. Among other things, the present disclosure provides insight that certain in vitro systems may be limited (e.g., with respect to their modeling of in vivo performance) by one or more of the following: the two-dimensional monolayer of the in vitro model versus the complex three-dimensional architecture of the in vivo environment, differences in target post-translational modifications or accessibility, differences in gene expression, the absence of stromal cells and / or extracellular matrix, lack of circulation and diffusion from the vasculature, or off-target and antigen sink effects that are not modeled in the in vitro model.
[0318] The binder-barcoded particles tested in Example 11 were also tested in an in vitro T cell activation assay (data not shown). Binder-barcoded particles that demonstrated the highest T cell activation in vitro exhibited poor tumor tissue accumulation and rapid clearance in vivo. Furthermore, binder-barcoded particles that demonstrated moderate T cell activation in vitro exhibited high tumor accumulation and slow clearance in vivo. Thus, this example confirms that the binder-barcode platform described herein can be used to identify binders with unexpected in vivo properties. Furthermore, this example confirms that the binder-barcoded platform described herein can be used to select therapeutic candidates that exhibit desirable phenotypes in vivo.
[0319] Example 15: Barcodes do not significantly affect payload characteristics This example confirms that barcodes have relatively little effect on payload (e.g., protein) properties. In some embodiments, the properties include payload affinity as described herein. In some embodiments, the properties include payload production (or yield) as described herein.
[0320] The affinity and production (or yield) of payloads (e.g., proteins) with and without barcodes were evaluated using biolayer interferometry (BLI). Ten different barcodes were tested. The affinity of the barcoded payloads, as measured by BLI, showed similar affinity to payloads without barcodes (data not shown). Furthermore, the payload production (or yield) of the barcoded payloads, as measured by BLI, showed similar payload production (or yield) to payloads without barcodes (data not shown).
[0321] Therefore, this example confirms that payload performance is not significantly affected by the barcode.Furthermore, this example confirms that payload production is not significantly affected by the barcode.
[0322] References Towbin H, Staehelin T, Gordon J. Electrophoretic transfer of proteins from polyacrylamide gels to nitrocellulose sheets: procedure and some applications. Proc Natl Acad Sci U S A. 1979;76(9):4350 - 4354. doi:10.1073 / pnas.76.9.4350 Engvall E, Perlmann P. Enzyme - linked immunosorbent assay, Elisa. 3. Quantitation of specific antibodies by enzyme - labeled anti - immunoglobulin in antigen - coated tubes. J Immunol. 1972 Jul;109(1):129 - 35. PMID:4113792. Elshal MF, McCoy JP. Multiplex bead array assays: performance evaluation and comparison of sensitivity to ELISA. Methods. 2006 Apr;38(4):317 - 23. doi:10.1016 / j.ymeth.2005.11.010. PMID:16481199; PMCID:PMC1534009. Shendure J, Porreca GJ, Reppas NB, Lin X, McCutcheon JP, Rosenbaum AM, Wang MD, Zhang K, Mitra RD, Church GM. Accurate multiplex polony sequencing of an evolved bacterial genome. Science. 2005 Sep 9;309(5741):1728 - 32. doi:10.1126 / science.1117389. Epub 2005 Aug 4. PMID:16081699. Trads JB,Torring T,Gothelf KV.Site-Selective Conjugation of Native Proteins with DNA.Acc Chem Res.2017 Jun 20;50(6):1367-1374.doi:10.1021 / acs.accounts.6b00618.Epub 2017 May 9.PMID:28485577. Egloff P,Zimmermann I,Arnold FM,Hutter CAJ,Morger D,Opitz L,Poveda L,Keserue HA,Panse C,Roschitzki B,Seeger MA.Engineered peptide barcodes for in-depth analyses of binding protein libraries.Nat Methods.2019 May;16(5):421-428.doi:10.1038 / s41592-019-0389-8.Epub 2019 Apr 22.PMID:31011184;PMCID:PMC7116144. Pollock SB,Hu A,Mou Y,Martinko AJ,Julien O,Hornsby M,Ploder L,Adams JJ,Geng H,Muschen M,Sidhu SS,Moffat J,Wells JA.Highly multiplexed and quantitative cell-surface protein profiling using genetically barcoded antibodies.Proc Natl Acad Sci U S A.2018 Mar 13;115(11):2836-2841.doi:10.1073 / pnas.1721899115.Epub 2018 Feb 23.PMID:29476010;PMCID:PMC5856557. Mohan D,Wansley DL,Sie BM,Noon MS,Baer AN,Laserson U,Larman HB.PhIP antibodies characterization of serum using oligonucleotide-encoded peptidomes.Nat Protoc.2018 Sep;13(9):1958-1978.doi:10.1038 / s41596-018-0025-6.Erratum in:Nat Protoc.2018 Oct 25;:PMID:30190553;PMCID:PMC6568263. Barbas, CF, Burton, DR, Scott, JK, & Silverman, GJ Phage Display: A Laboratory Manual. 2001. Cold Spring Harbor Laboratory Press.
[0323] Other embodiments It should be understood that various changes, modifications, and improvements to this disclosure will readily occur to those skilled in the art. Such changes, modifications, and improvements are intended to be part of this disclosure and are intended to be within the spirit and scope of the invention. Accordingly, the foregoing description and drawings are by way of example only, and any inventions described in this disclosure are further described in detail by the following claims.
[0324] Those of ordinary skill in the art will understand the typical basis for deviation or error attributable to values obtained in the assays or other processes described herein. Publications, websites, and other reference materials referred to herein to describe the background of the invention and to provide additional details regarding its practice are hereby incorporated by reference in their entireties.
[0325] While embodiments of the present invention have been described in conjunction with the detailed description thereof, it should be understood that the foregoing description is intended to be illustrative and not limiting of the scope of the invention as defined by the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.
[0326] equivalent Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the invention described herein. The scope of the present invention is not intended to be limited to the above Description, but rather is as set forth in the following claims.
[0327] [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 1-5] [Table 1-6] [Table 1-7] [Table 1-8] [Table 1-9] [Table 1-10]
[0328] [Table 1-11] [Table 1-12] [Table 1-13]
Table 1-14
Table 1-15
Table 1-16
Table 1-17
Table 1-18
Table 1-19
[0329]
Table 1-20
Table 1-21
Table 1-22
Table 1-23
Table 1-24
Table 1-25
Table 1-26
Table 1-27
Table 1-28
Table 1-29
[0330]
Table 1-30
Table 1-31
Table 1-32
Table 1-33
Table 1-34
Table 1-35
Table 1-36
Table 1-37
Table 1-38
Table 1-39
[0331]
Table 1-40
Table 1-41
Table 1-42
Table 1-43
Table 1-44
Table 1-45
Table 1-46
Table 1-47
Table 1-48
Table 1-49
[0332]
Table 1-50
Table 1-51
Table 1-52
Table 1-53
Table 1-54
Table 1-55
Table 1-56
Table 1-57
Table 1-58
Table 1-59
[0333]
Table 1-60
Table 1-61
Table 1-62
Table 1-63
Table 1-64
Table 1-65
Table 1-66
Table 1-67
Table 1-68
Table 1-69
[0334]
Table 1-70
Table 1-71
Table 1-72
Table 1-73
Table 1-74
Table 1-75
Table 1-76
Table 1-77
Table 1-79
[0335]
Table 1-80
Table 1-81
Table 1-82
Table 1-83
Table 1-84
Table 1-85
Table 1-86
Table 1-87
Table 1-88
Table 1-89
[0336]
Table 1-90
Table 1-91
Table 1-92
Table 1-93
Table 1-94
Table 1-95
Table 1-96
Table 1-97
Table 1-98
Table 1-99
[0337]
Table 1-100
Table 1-101
Table 1-103
Table 1-104
Table 1-105
Table 1-106
Table 1-108
Table 1-109
[0338]
Table 1-110
Table 1-111
Table 1-112
Table 1-113
Table 1-114
Table 1-115
Table 1-116
Table 1-117
Table 1-118
Table 1-119
[0339]
Table 1-120
Table 1-121
Table 1-123
Table 1-124
Table 1-125
Table 1-126
Table 1-127
Table 1-129
[0340] Table 1-130
Table 1-131
Table 1-133
Table 1-136
Table 1-137
Table 1-138
Table 1-139
[0341] Table 1-140
Table 1-141
Table 1-145
Table 1-147
Table 1-149
Table 1-150
[0342]
Table 1-151
Table 1-152
Table 1-153
Table 1-155
Table 1-156
Table 1-157
Table 1-158
Table 1-159
[0343] Table 1-160
Table 1-161
Table 1-166
Table 1-167
[0344]
Table 1-171
Table 1-176
Table 1-177
Table 1-178
Table 1-179
[0345] Table 1-181 Table 1-182 Table 1-183 Table 1-184 Table 1-185
Table 1-186
[0346] Table 1-190
Table 1-191
Table 1-193
Table 1-196
Table 1-197
Table 1-199
[0347]
Table 1-200
Table 1-201
Table 1-202
Table 1-203
Table 1-205
Table 1-207
Table 1-209
[0348]
Table 1-211
Table 1-212
Table 1-217
Table 1-219
[0349]
Table 1-220
Table 1-221
Table 1-222
Table 1-223
Table 1-224
Table 1-225
Table 1-226
Table 1-227
Table 1-228
Table 1-229
[0350] Table 1-230
Table 1-231
Table 1-233
Table 1-236
Table 1-237
Table 1-239
[0351] Table 1-241 Table 1-242 Table 1-243 Table 1-244 Table 1-245
Table 1-246
Table 1-247
Table 1-249
Table 1-250
[0352]
Table 1-251
Table 1-252
Table 1-253
Table 1-254
Table 1-256
Table 1-257
Table 1-258
Table 1-259
[0353] Table 1-261 Table 1-262
Table 1-263
Table 1-269
[0354]
Table 1-271
Table 1-276
[0355] Table 1-281 Table 1-282
Table 1-283
Table 1-288
Table 1-289
[0356]
Table 1-291
Table 1-297
Table 1-299
[0357]
Table 1-300
Table 1-303
Table 1-304
Table 1-305
Table 1-306
Table 1-307
Table 1-309
[0358]
Table 1-310
Table 1-311
Table 1-313
Table 1-315
Table 1-316
Table 1-317
Table 1-319
[0359]
Table 1-321
Table 1-323
Table 1-325
Table 1-326
Table 1-327
Table 1-329
Table 1-331
[0360] Table 1-332
Table 1-333
Table 1-334
Table 1-335
Table 1-336
Table 1-337
Table 1-338
Table 1-339
Table 1-341
[0361] Table 1-342
Table 1-343
Table 1-344
Table 1-345
Table 1-346
Table 1-347
Table 1-349
[0362]
Table 1-350
Table 1-351
Table 1-352
Table 1-353
Table 1-354
Table 1-355
Table 1-356
Table 1-357
Table 1-359
[0363]
Table 1-361
Table 1-363
Table 1-366
Table 1-367
Table 1-369
[0364]
Table 1-371
Table 1-373
Table 1-374
Table 1-377
Table 1-378
Table 1-379
[0365] Table 1-381 Table 1-382 Table 1-383 Table 1-384 Table 1-385 Table 1-386 Table 1-387 Table 1-388 Table 1-389
[0366] Table 1-390
Table 1-391
Table 1-396
Table 1-397
Table 1-399
[0367] Table 1-402 Table 1-403 Table 1-404 Table 1-405 Table 1-406 Table 1-407 Table 1-408 Table 1-409
Table 1-410
[0368] Table 1-411 Table 1-412
Table 1-413
Table 1-414
Table 1-415
Table 1-416
Table 1-417
Table 1-419
[0369]
Table 1-421
Table 1-422
Table 1-423
Table 1-425
Table 1-426
Table 1-427
Table 1-428
Table 1-429
Table 1-430
[0370]
Table 1-431
Table 1-432
Table 1-433
Table 1-434
Table 1-436
Table 1-437
Table 1-438
Table 1-439
[0371]
Table 1-440
Table 1-441
Table 1-442
Table 1-443
Table 1-444
Table 1-445
Table 1-446
Table 1-447
Table 1-448
Table 1-449
[0372]
Table 1-450
Table 1-451
Table 1-452
Table 1-453
Table 1-454
Table 1-455
Table 1-456
Table 1-457
Table 1-458
Table 1-459
[0373]
Table 1-461
Table 1-462
Table 1-463
Table 1-465
Table 1-466
Table 1-467
Table 1-468
Table 1-469
[0374]
Table 1-470
Table 1-471
Table 1-472
Table 1-473
Table 1-474
Table 1-475
Table 1-476
Table 1-477
Table 1-478
Table 1-479
[0375]
Table 1-481
Table 1-483
Table 1-484
Table 1-486
Table 1-487
Table 1-488
Table 1-489
Table 1-490
[0376]
Table 1-491
Table 1-492
Table 1-493
Table 1-495
Table 1-496
Table 1-497
Table 1-498
Table 1-499
Table 1-500
[0377]
Table 1-501
Table 1-503
Table 1-506
Table 1-507
Table 1-509
Table 1-510
[0378]
Table 1-511
Table 1-514
Table 1-515
Table 1-516
Table 1-517
Table 1-519
[0379]
Table 1-520
Table 1-521
Table 1-522
Table 1-523
Table 1-524
Table 1-525
Table 1-526
Table 1-527
Table 1-529
[0380] Table 1-530
Table 1-531
Table 1-532
Table 1-533
Table 1-534
Table 1-535
Table 1-536
Table 1-537
Table 1-538
Table 1-539
Table 1-540
[0381]
Table 1-541
Table 1-542
Table 1-543
Table 1-545
Table 1-546
Table 1-547
Table 1-548
Table 1-549
Table 1-550
[0382]
Table 1-551
Table 1-552
Table 1-553
Table 1-554
Table 1-555
Table 1-556
Table 1-557
Table 1-558
Table 1-559
Table 1-560
[0383]
Table 1-561
Table 1-562
Table 1-563
Table 1-564
Table 1-565
Table 1-566
Table 1-567
Table 1-568
Table 1-569
Table 1-570
[0384]
Table 1-571
Table 1-572
Table 1-573
Table 1-574
Table 1-575
Table 1-576
Table 1-577
Table 1-578
Table 1-579
Table 1-581
[0385]
Table 1-582
Table 1-583
Table 1-584
Table 1-585
Table 1-586
Table 1-587
Table 1-588
Table 1-589
Table 1-590
Table 1-591
[0386]
Table 1-592
Table 1-593
Table 1-594
Table 1-595
Table 1-596
Table 1-597
Table 1-598
Table 1-599
Table 1-600
Table 1-603
[0387]
Table 1-604
Table 1-605
Table 1-606
Table 1-607
Table 1-608
Table 1-609
Table 1-610
Table 1-611
Table 1-612
Table 1-613
[0388]
Table 1-614
Table 1-615
Table 1-616
Table 1-617
Table 1-618
Table 1-619
Table 1-620
Table 1-621
Table 1-622
Table 1-623
Table 1-624
[0389]
Table 1-625
Table 1-626
Table 1-627
Table 1-628
Table 1-629
Table 1-630
Table 1-631
Table 1-632
Table 1-633
Table 1-634
Table 1-635
Table 1-636
[0390]
Table 1-637
Table 1-638
Table 1-639
Table 1-641
Table 1-642
Table 1-643
Table 1-644
Table 1-645
[0391]
Table 1-646
Table 1-647
Table 1-648
Table 1-649
Table 1-650
Table 1-651
Table 1-652
Table 1-653
Table 1-654
Table 1-655
Table 1-656
Table 1-657
Table 1-658
Table 1-659
[0392]
Table 1-660
Table 1-661
Table 1-662
Table 1-663
Table 1-664
Table 1-665
Table 1-666
Table 1-667
Table 1-668
Table 1-669
Table 1-670
Table 1-671
Table 1-672
Table 1-673
Table 1-674
[0393]
Table 1-675
Table 1-676
Table 1-677
Table 1-678
Table 1-679
Table 1-680
Table 1-681
Table 1-682
Table 1-683
Table 1-684
Table 1-685
Table 1-686
Table 1-687
Table 1-688
Table 1-689
[0394]
Table 1-690
Table 1-691
Table 1-692
Table 1-693
Table 1-694
Table 1-695
Table 1-696
Table 1-697
Table 1-698
Table 1-699
Table 1-700
Table 1-701
[0395] Table 1-702
Table 1-703
Table 1-704
Table 1-705
Table 1-706
Table 1-707
Table 1-708
Table 1-709
Table 1-710
[0396]
Table 2-1
Table 2-2
Table 2-3
Table 2-4
Table 2-5
Table 2-6
Table 2-7
Table 2-8
Table 2-9
Table 2-10
Table 2-11
Table 2-12
Table 2-13
Table 2-14
[0397]
Table 2-15
Table 2-16
Table 2-17
Table 2-18
Table 2-19
Table 2-20
Table 2-21
Table 2-22
Table 2-23
Table 2-24
Table 2-25
Table 2-26
Table 2-27
Table 2-28
Table 2-29
[0398]
Table 2-30
Table 2-31
Table 2-32
Table 2-33
Table 2-34
Table 2-35
Table 2-36
Table 2-37
Table 2-38
Table 2-39
Table 2-40
Table 2-41
Table 2-42
Table 2-43
Table 2-44
[0399]
Table 2-45
Table 2-46
Table 2-47
Table 2-48
Table 2-49
Table 2-50
Table 2-51
Table 2-52
Table 2-53
Table 2-54
Table 2-55
Table 2-56
Table 2-57
Table 2-58
Table 2-59
[0400]
Table 2-60
Table 2-61
Table 2-62
Table 2-63
Table 2-64
Table 2-65
Table 2-66
Table 2-67
Table 2-68
Table 2-69
Table 2-70
Table 2-71
Table 2-72
Table 2-73
Table 2-74
[0401]
Table 2-75
Table 2-76
Table 2-77
Table 2-78
Table 2-79
Table 2-80
Table 2-81
Table 2-82
Table 2-83
Table 2-84
Table 2-85
Table 2-86
Table 2-87
Table 2-88
Table 2-89
[0402]
Table 2-90
Table 2-91
Table 2-92
Table 2-93
Table 2-94
Table 2-95
Table 2-96
Table 2-97
Table 2-98
Table 2-99
Table 2-100
[0403]
Table 2-101
Table 2-102
Table 2-103
Table 2-104
Table 2-105
Table 2-106
Table 2-107
Table 2-108
Table 2-109
Table 2-110
Table 2-111
Table 2-112
Table 2-113
Table 2-114
Table 2-115
[0404]
Table 2-116
Table 2-117
Table 2-118
Table 2-119
Table 2-120
Table 2-121
Table 2-122
Table 2-123
Table 2-124
Table 2-125
Table 2-126
Table 2-127
Table 2-128
Table 2-129
Table 2-131
Table 2-132
Table 2-133
Table 2-134
[0405]
Table 2-135
Table 2-136
Table 2-137
Table 2-138
Table 2-139
Table 2-141
Table 2-142
Table 2-143
Table 2-144
Table 2-145
Table 2-146
Table 2-147
Table 2-148
Table 2-149
Table 2-150
Table 2-151
Table 2-152
Table 2-153
Table 2-154
[0406]
Table 2-155
Table 2-156
Table 2-157
Table 2-158
Table 2-159
Table 2-160
Table 2-161
Table 2-162
Table 2-163
Table 2-164
Table 2-165
Table 2-166
Table 2-167
Table 2-169
Table 2-170
Table 2-171
Table 2-172
Table 2-173
Table 2-174
Table 2-175
Table 2-176
Table 2-177
Table 2-178
Table 2-179
[0407]
Table 2-181
Table 2-183
Table 2-184
Table 2-186
Table 2-187
Table 2-189
Table 2-190
Table 2-191
Table 2-192
Table 2-193
Table 2-194
Table 2-195
Table 2-196
Table 2-197
Table 2-199
Table 2-200
[0408]
Table 2-201
Table 2-202
Table 2-203
Table 2-204
Table 2-205
Table 2-206
Table 2-207
Table 2-208
Table 2-209
Table 2-210
Table 2-211
Table 2-212
Table 2-213
Table 2-214
Table 2-215
Table 2-216
Table 2-217
Table 2-218
Table 2-219
[0409]
Table 2-220
Table 2-221
Table 2-222
Table 2-223
Table 2-224
Table 2-225
Table 2-226
Table 2-227
Table 2-228
Table 2-229
Table 2-230
Table 2-231
Table 2-232
Table 2-233
Table 2-234
Table 2-235
Table 2-236
Table 2-237
Table 2-238
Table 2-239
[0410]
Table 2-240
Table 2-241
Table 2-242
Table 2-243
Table 2-244
Table 2-245
Table 2-246
Table 2-247
Table 2-248
Table 2-249
Table 2-250
Table 2-251
Table 2-252
Table 2-253
Table 2-254
Table 2-255
Table 2-256
Table 2-257
Table 2-258
Table 2-259
Table 2-260
[0411]
Table 2-261
Table 2-262
Table 2-263
Table 2-264
Table 2-265
Table 2-266
Table 2-267
Table 2-268
Table 2-269
Table 2-270
Table 2-271
Table 2-272
Table 2-273
Table 2-274
Table 2-275
Table 2-276
Table 2-277
Table 2-278
Table 2-279
[0412]
Table 2-280
Table 2-281
Table 2-282
Table 2-283
Table 2-284
Table 2-285
Table 2-286
Table 2-287
Table 2-288
Table 2-289
Table 2-290
Table 2-291
Table 2-292
Table 2-293
Table 2-294
Table 2-295
Table 2-297
Table 2-299
Table 2-300
Table 2-301
Table 2-302
Table 2-303
[0413]
Table 2-304
Table 2-305
Table 2-306
Table 2-307
Table 2-308
Table 2-309
Table 2-310
Table 2-311
Table 2-312
Table 2-313
Table 2-314
Table 2-315
Table 2-316
Table 2-317
Table 2-319
Table 2-320
Table 2-321
[0414]
Table 2-322
Table 2-323
Table 2-324
Table 2-325
Table 2-326
Table 2-327
Table 2-328
Table 2-329
Table 2-330
Table 2-331
Table 2-332
Table 2-333
Table 2-334
Table 2-335
Table 2-336
Table 2-337
[0415]
Table 2-338
Table 2-339
Table 2-340
Table 2-341
Table 2-342
Table 2-343
Table 2-344
Table 2-345
Table 2-346
Table 2-347
Table 2-348
Table 2-349
Table 2-350
Table 2-351
Table 2-352
Table 2-353
Table 2-354
Table 2-355
Table 2-356
Table 2-357
Table 2-358
Table 2-359
Table 2-361
Table 2-362
Table 2-363
Table 2-364
Table 2-365
Table 2-366
[0416]
Table 2-367
Table 2-368
Table 2-369
Table 2-370
Table 2-371
Table 2-372
Table 2-373
Table 2-374
Table 2-375
Table 2-376
Table 2-377
Table 2-378
Table 2-379
Table 2-381
Table 2-382
Table 2-383
Table 2-385
Table 2-386
Claims
1. 1. A library comprising a plurality of nucleic acids, the plurality of nucleic acids together encoding a set of peptide barcodes, each peptide barcode comprising: (a) having a length in the range of 5 to 50, 8 to 25, 9 to 25, 9 to 25, or 9 to 15 amino acids; (b) determined to specifically bind to a particular group of polypeptide binders within the set of binders, wherein: The nucleic acids are arranged in 5' to 3' or 3' to 5' order: i) a first immutable sequence; ii) a variant sequence that is at least 9 nucleotides in length, and iii) a second immutable sequence and the collection of peptide barcodes comprises at least one peptide barcode having an amino acid sequence selected from the group consisting of SEQ ID NOs: 5347-8398; and the plurality of nucleic acids comprises at least one nucleic acid having a coding sequence selected from the group consisting of SEQ ID NOs: 1148-4199; Library.
2. c) a sequence containing a short helical motif; d) sequences containing disordered motifs; e) an invariant sequence linking the sequence to the protein of interest The library of claim 1 , further comprising one or more of:
3. 10. The library of claim 1, wherein each peptide barcode of the collection specifically binds to one or more polypeptide binders in a set of binders.
4. 2. The library of claim 1, wherein the first constant sequence comprises or is a linker sequence or a payload sequence.
5. 2. The library of claim 1, wherein the second invariant sequence comprises or is a linker sequence, a stop codon, or a payload sequence.
6. 2. The library of claim 1, wherein the first constant sequence comprises or is a first linker sequence, the second constant sequence comprises or is a second linker sequence, and the first linker sequence is the same as the second linker sequence.
7. 2. The library of claim 1, wherein the first constant sequence comprises or is a first linker sequence, the second constant sequence comprises or is a second linker sequence, and the first linker sequence is different from the second linker sequence.
8. A library comprising a plurality of nucleic acids, said plurality of nucleic acids together encoding a set of polypeptide binder moieties, each polypeptide binder moiety comprising: a) having a length in the range of 10 to 400 amino acids; b) determined to specifically bind to a particular group of peptide barcodes within a collection of barcodes, wherein: The nucleic acids are arranged in 5' to 3' or 3' to 5' order: i) a first immutable sequence; ii) a first variant sequence that is at least 10 nucleotides in length; and iii) a second immutable sequence Including, the set of polypeptide binder moieties comprises at least one polypeptide binder moiety having an amino acid sequence selected from the group consisting of SEQ ID NOs: 4200-5346; and the plurality of nucleic acids comprises at least one nucleic acid having a coding sequence selected from the group consisting of SEQ ID NOs: 1-1147; Library.
9. A library of phage particles, each phage particle comprising one or more nucleic acids, wherein each nucleic acid comprises a sequence encoding a polypeptide binder moiety, said polypeptide binder moiety comprising: a) having a length in the range of 10 to 400 amino acids; b) determined to specifically bind to a particular group of peptide barcodes within a collection of barcodes, wherein: The nucleic acids are arranged in 5' to 3' or 3' to 5' order: i) a first immutable sequence; ii) a first variant sequence that is at least 10 nucleotides in length; and iii) a second immutable sequence Including, the polypeptide binder portion has an amino acid sequence selected from the group consisting of SEQ ID NOs: 4200-5346; and the one or more nucleic acids include at least one nucleic acid having a coding sequence selected from the group consisting of SEQ ID NOs: 1-1147; Library.
10. Each nucleic acid is d) stop codon e) a linker sequence; f) a third immutable array; g) a second variant sequence that is at least 10 nucleotides in length, and h) a fourth immutable array 10. The library of claim 8 or 9, further comprising one or more of:
11. 10. The library of claim 9, wherein the phage particles are selected from the group consisting of M13, T4, T7, lambda, and filamentous phage.
12. The library of claim 9 , wherein the phage particles are M13.
13. 10. The library of claim 8 or 9, wherein the first constant sequence or the second constant sequence comprises an antibody germline sequence.
14. 14. The library of claim 13, wherein the antibody germline sequences comprise an amino acid sequence of IGHV, IGKV, IDHJ, or IGKJ.
15. 10. The library of claim 8 or 9, wherein the first variant sequence comprises a CDR sequence.
16. 16. The library of claim 15, wherein the CDR sequences are CDR3 sequences.
17. The library of claim 10, wherein a stop codon is located after the second invariant sequence.
18. 11. The library of claim 10, wherein the third constant sequence or the fourth constant sequence comprises an antibody germline sequence.
19. 19. The library of claim 18, wherein the antibody germline sequences comprise an amino acid sequence of IGHV, IGKV, IDHJ, or IGKJ.
20. The library of claim 10 , wherein the second variant sequence comprises a CDR sequence.
21. 21. The library of claim 20, wherein the CDR sequences are CDR3 sequences.
22. 10. The library of claim 9, wherein the one or more nucleic acids encode at least one polypeptide binder moiety.