Peptide libraries having enhanced subsequence diversity and methods for use thereof
The engineered peptide library with enhanced subsequence diversity addresses the limitation of traditional peptide synthesis by representing a larger number of unique sequences on a single array, effectively identifying peptide binders through an algorithmic tiling method.
Patent Information
- Application Number
- JP2025107097
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-06-27
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-25
AI Technical Summary
Existing peptide synthesis techniques are limited by the number of reaction sites on a single array, bead, or chip, constraining the diversity of peptide sequences that can be explored, particularly for longer peptides.
An engineered peptide library is developed with enhanced subsequence diversity, where each peptide feature comprises a composite region representing multiple distinct elements, allowing for a larger number of unique peptide sequences to be represented on a single array through an algorithmic approach that tiles shorter peptide elements within longer sequences.
This approach significantly increases the number of unique peptide sequences that can be represented on a single array, enhancing the ability to identify peptide binders by achieving nearly 100% representation of possible sequences, surpassing the limitations of traditional methods.
Smart Images

Figure 2025138750000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Patent Application Nos. 62 / 867,765 and 62 / 867,666, filed June 27, 2019, the contents of which are incorporated herein by reference in their entirety for any and all purposes.
[0002] FIELD OF THE DISCLOSURE This disclosure relates generally to the design and selection of synthetic peptides for biomarker discovery, and more particularly to peptide libraries with enhanced subsequence diversity and methods for their use. [Background technology]
[0003] Peptides are biological polymers assembled in part by the formation of amide bonds between amino acid monomer units. In general, peptides can be distinguished from their protein counterparts based on factors such as size (e.g., number of monomer units or molecular weight), complexity (e.g., number of peptides, presence of coenzymes, cofactors, or other ligands), and the like. Experimental approaches to identify binding motifs, epitopes, mimotopes, disease markers, and the like can be successful in using peptides in place of larger or more complex proteins that may be more difficult to obtain or manipulate. As a result, the study of peptides and the possibility of synthesizing them are of great interest in biological science and medicine.
[0004] Several methods exist for the synthesis of peptides, including both in vivo and in vitro translation systems, as well as organic synthesis routes, such as solid-phase peptide synthesis. Solid-phase peptide synthesis is a technique in which the first amino acid is attached to a solid surface, such as a bead, microscope slide, or another similar surface. Subsequent amino acids are then added stepwise to the first amino acid to form a peptide chain. Because the peptide chain is attached to the solid surface, manipulations such as washing steps, side chain modifications, cyclization, or other processing steps can be performed while maintaining the peptide chain in place.
[0005] Recent advances in solid-phase peptide synthesis have led to automated synthesis platforms for the parallel assembly of millions of unique peptide features into arrays on a single surface (e.g., an approximately 75 mm x approximately 25 mm microscope slide). The utility of such peptide arrays relies, at least in part, on their ability to simultaneously explore the diversity of peptide sequences. While existing techniques can enable the exploration of millions of different sequences, the number of unique sequences that can be explored is inherently limited, for example, by the number of reaction sites on a single array, bead, chip, or other solid support. Summary of the Invention
[0006] The present disclosure provides a series of peptide binders to biologically relevant proteins that have been identified by a method that includes identifying overlapping binding of a target protein to small peptides among a comprehensive collection of peptides immobilized on a microarray, then performing one or more rounds of maturation of the isolated core hit peptides, followed by one or more rounds of N- and C-terminal extension of the mature peptides.
[0007] In one aspect, the technology relates to an engineered peptide library comprising a plurality of peptide features, each of the peptide features comprising at least one peptide, the at least one peptide comprising a composite region having a defined sequence of amino acids of length N, the composite region representing k distinct elements, each distinct element having a defined sequence of amino acids of length x, where x, N, and k are integers, x is less than N, k is at least 2, and the total number of distinct elements represented by the engineered peptide library is greater than or equal to K. Eng where F is the number of peptide features contained in the engineered peptide library, and K Eng is greater than F. In some embodiments, k=N−x+1.
[0008] In some embodiments, the plurality of peptides represents at least about 90% of the target proteome. The engineered peptide libraries described herein can have enhanced subsequence diversity.
[0009] In one aspect, an engineered peptide library may include a plurality of peptide features, each of the peptide features comprising at least one peptide, wherein the at least one peptide comprises a composite region having a defined sequence of amino acids of length N, wherein the composite region represents k distinct elements, each distinct element having a defined sequence of amino acids of length x, where x, N, and k are integers, x is less than N, k is at least 2, and the total number of distinct elements represented by the engineered peptide library is greater than or equal to K. Eng where the number of peptide features contained in the engineered peptide library is F, K is greater than F, the ratio of K to F is a measure of subsequence diversity, each of the k distinct elements in each composite region has a unique sequence compared to each of the other distinct elements in the same composite region, the defined sequences of each composite region are selected for maximum subsequence diversity relative to the average subsequence diversity for the total number of random elements K, the random elements have a sequence of amino acids of length x represented by a random peptide library having F peptide features, each peptide feature of the random peptide library includes at least one random peptide, and the at least one random peptide has a random sequence of amino acids of length N.
[0010] The technology also provides an engineered peptide library comprising a plurality of peptide features, each of the peptide features comprising at least one peptide, the at least one peptide comprising a composite region having a defined sequence of at least 15 amino acids, the composite region representing 10 distinct elements, each distinct element having a defined sequence of 6 amino acids, wherein the total number of distinct elements represented by the engineered peptide library is K Engwhere F is the number of peptide features contained in the engineered peptide library, and K Eng is at least 9.5*F.
[0011] In another aspect, the present technology relates to a method for identifying peptide binders, the method comprising contacting a first sample with an engineered peptide library described herein and selecting at least one of a plurality of peptides from a first subset of peptides. [The present invention 1001] a plurality of peptide features, each of said peptide features comprising at least one peptide, said at least one peptide comprising a composite region having a defined sequence of amino acids of length N, said composite region representing k distinct elements, each of said distinct elements having a defined sequence of amino acids of length x; 1. An engineered peptide library comprising: x, N, and k are integers; x is less than N, k is at least 2, The total number of distinct elements represented by the engineered peptide library is K Eng and the number of peptide features contained in the engineered peptide library is F; K Eng is greater than F, The engineered peptide library. [The present invention 1002] 1001. An engineered peptide library of the present invention, wherein k is defined by the equation: k=N-x+1. [The present invention 1003] K Eng The engineered peptide library of the present invention 1001 or 1002, wherein k is at least 0.8*k*F. [The present invention 1004] K Eng The engineered peptide library of any of claims 1001 to 1003, wherein k is at least 0.9*k*F. [The present invention 1005] K Eng 1005. The engineered peptide library of any of claims 1001 to 1004, wherein k is at least 0.95*k*F. [The present invention 1006] K Eng 1006. The engineered peptide library of any of claims 1001 to 1005, wherein k is at least 0.99*k*F. [The present invention 1007] K Eng 1006. The engineered peptide library of any of claims 1001 to 1006, wherein F is at least 10*F. [The present invention 1008] K Eng 1008. The engineered peptide library of any of claims 1001 to 1007, wherein F is at least 20*F. [The present invention 1009] K Eng is the total number K of random elements with sequences of amino acids of length x represented by a random peptide library with F peptide features. Rnd 1009. The engineered peptide library of any of claims 1001 to 1008, wherein each of said peptide features of said random peptide library comprises at least one random peptide, said at least one random peptide having a random sequence of amino acids of length N. [The present invention 1010] 1009. The engineered peptide library of any of claims 1001 to 1009, wherein each of said elements in a selected one of said composite regions overlaps with each of its adjacent elements in said composite region by at least one amino acid. [The present invention 1011] 1010. The engineered peptide library of any of claims 1001 to 1010, wherein said plurality of peptides represents at least 90% of the target proteome. [The present invention 1012] 1012. The engineered peptide library of any of claims 1001 to 1011, wherein each of said peptides in said plurality of peptides is 12 amino acids to 16 amino acids in length. [The present invention 1013] 13. The engineered peptide library of any of claims 1001 to 1012, wherein N is at least 7 amino acids. [The present invention 1014] 10. The engineered peptide library of any of claims 1001 to 1013, wherein N is at least 10 amino acids. [The present invention 1015] 15. The engineered peptide library of any of claims 1001 to 1014, wherein N is at least 15 amino acids. [The present invention 1016] 10. The engineered peptide library of any of claims 1001 to 1015, wherein x is at least 5. [The present invention 1017] The engineered peptide library of any of claims 1001 to 1016, wherein x is 6. [The present invention 1018] 1. A method for identifying peptide binders, comprising: contacting a first sample with any one of the engineered peptide libraries of the present inventions 1001 to 1017; selecting at least one of said plurality of peptides from a first subset of peptides; A method for identifying said peptide binders comprising: [The present invention 1019] detecting a first signal output characteristic of said first interaction with said engineered peptide library; selecting the at least one peptide based on the first signal output; The method of the present invention 1018 further comprises: [The present invention 1020] The first signal output is a fluorescence intensity obtained through fluorophore excitation emission, and the fluorescence intensity is i) the abundance of a component in one of the first sample and the second sample associated with the first plurality of peptides; and ii) the binding affinity of the components of one of the first sample and the second sample to the first plurality of peptides. The method of the present invention 1019 reflects at least one of the above. [The present invention 1021] a plurality of peptide features, each of said peptide features comprising at least one peptide, said at least one peptide comprising a composite region having a defined sequence of at least 15 amino acids, said composite region representing 10 distinct elements, each of said distinct elements having a defined sequence of 6 amino acids; 1. An engineered peptide library comprising: The total number of distinct elements represented by the engineered peptide library is K Eng and the number of peptide features contained in the engineered peptide library is F; K Eng is at least 9.5*F, The engineered peptide library. [The present invention 1022] 1021. The engineered peptide library of any of claims 1001 to 1017 or 1021, wherein each of said peptides is attached to a solid support. [The present invention 1023] 1022. The engineered peptide library of the present invention, wherein the solid support is one of a bead and a chip. [The present invention 1024] 1. An engineered peptide library with enhanced subsequence diversity, the engineered peptide library comprising: a plurality of peptide features, each of said peptide features comprising at least one peptide, said at least one peptide comprising a composite region having a defined sequence of amino acids of length N, said composite region representing k distinct elements, each of said distinct elements having a defined sequence of amino acids of length x; Including, x, N, and k are integers; x is less than N, k is at least 2, The total number of distinct elements represented by the engineered peptide library is K Eng and the number of peptide features contained in the engineered peptide library is F; KEng is greater than F, The ratio of KEng to F is a measure of subsequence diversity, each of the k distinct elements within each composite region has a unique arrangement compared to each of the other distinct elements of the same composite region; The defined sequences of each of the composite regions are selected for their maximum subsequence diversity relative to the average subsequence diversity of the total number of random elements, KRnd; the random element has a sequence of amino acids of length x represented by a random peptide library having F peptide features, each of the peptide features of the random peptide library comprising at least one random peptide, the at least one random peptide having a random sequence of amino acids of length N; The engineered peptide library having enhanced subsequence diversity. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a schematic diagram showing a peptide array for peptide binder discovery. [Figure 2] FIG. 2 is an example of a method for identifying peptide binders according to the present disclosure. [Figure 3] FIG. 3 is a schematic diagram depicting one embodiment of a maturation array comprising a population of peptides immobilized on a solid support, where each peptide comprises a mature core hit peptide sequence. [Figure 4] FIG. 4 is a schematic diagram depicting one embodiment of a method for identifying peptide binders. [Figure 5]Figure 5A is a schematic diagram of an embodiment of a peptide array comprising a population of peptide features for identification and characterization of control peptides. Figure 5B is a schematic diagram of an embodiment of the peptide array of Figure 5A after exposure of the peptide features to a plurality of receptor molecules. Figure 5C is a schematic diagram of an embodiment of the peptide array of Figure 5B after attachment of a detectable tag to the receptor molecule. [Figure 6] FIG. 6 is a schematic diagram representing 16-mer peptides tiled at either one-amino acid resolution or four-amino acid resolution, and includes a table showing the tiling of a portion of the exemplary protein sequence EGVKLTALNDSSLDLSMDSDNSMSV (SEQ ID NO: 69), represented as 16-mer peptides tiled at one-amino acid resolution (SEQ ID NOs: 71-80, respectively, in order of appearance). [Figure 7] FIG. 7 is a schematic diagram depicting 15-mer peptides tiled at one amino acid resolution, showing how a composite 15-mer peptide can represent ten 6-mer elements or subsequences. [Figure 8] Figure 8 shows a single substitution plot of the DFWHGDTCKVTQFDQ peptide, DsbA_1_WT (SEQ ID NO: 70). The position of each peptide is represented by 21 bars (one bar for each of the 20 amino acids, plus the deletion (last bar)). The height of each bar indicates the median signal intensity. [Figure 9] FIG. 9 shows binding of DsbA_l_WT, DFWHGDTCKVTQFDQ-NH2 (SEQ ID NO: 70) to biotin-labeled DsbA immobilized on a (Biacore) SA chip. [Figure 10] FIG. 10 shows a comparison of the fluorescent signal intensity and binding affinity.
[0013] Throughout the following detailed description, like numbers are used to describe like parts from figure to figure. DETAILED DESCRIPTION OF THE INVENTION
[0014] I. Overview As described above, in various situations, it may be useful to provide a collection of peptides prepared by solid-phase peptide synthesis. In the case of peptide arrays, the ability to simultaneously probe diverse peptide sequences is desirable in various applications. While existing techniques can enable the probing of millions of different sequences, the number of unique sequences that can be probed is inherently limited, for example, by the number of available reaction sites or features on a single array, bead, chip, or other solid support. These and other challenges may be overcome by peptide libraries with enhanced subsequence diversity according to the present disclosure.
[0015] The following terms are used throughout, as defined below:
[0016] As used herein and in the appended claims, in the context of describing elements (particularly in the context of the claims that follow), singular articles such as "a," "an," and "the" and similar referents are to be construed to cover both the singular and the plural unless otherwise stated herein or clearly contradicted by context. Recitation of ranges of values herein is merely intended to serve as a shorthand method of individually referring to each separate value falling within the range, unless otherwise indicated herein, and each separate value is incorporated herein as if it were individually listed herein. All methods described herein can be performed in any suitable order unless otherwise stated herein or clearly contradicted by context. The use of any and all examples or exemplary language (e.g., "e.g., "etc.") provided herein is intended merely to better clarify embodiments and does not pose a limitation on the scope of the claims unless otherwise stated. No language herein should be construed as indicating any non-claimed element as essential.
[0017] As used herein, "about" will be understood by one of ordinary skill in the art and will vary to some extent depending on the context in which it is used. If there are uses of the term that are not clear to persons of ordinary skill in the art, given the context in which it is used, "about" will mean plus or minus 10% of the particular term.
[0018] As used herein, an "engineered peptide library" is a library of peptide sequences that has been designed and synthesized to allow for the exploration and testing of significantly more motifs than would otherwise be available in a given fixed library format. Libraries of peptides prepared using known synthetic techniques are defined by parameters including the number of peptides or peptide features (e.g., in the case of microarrays) and the total length of the peptide in amino acids. However, if it is desired to screen a larger number of peptides than is provided in a given library format, multiple libraries must be generated to provide the required library size. For example, a library of all possible 6-mers can be generated in 20 6 or would require a peptide library size of 64 million unique peptides. To allow a single peptide library to explore a larger portion of this sequence space, a design approach has been developed in which multiple x-mers are embedded into an N-mer peptide sequence, where N and x are integers and N is greater than x. In one aspect, this approach provides for the representation of multiple unique x-mer peptides in a single N-mer peptide feature. In one example, the synthesis of over 30 million unique 6-mer peptide motifs was achieved in a peptide feature space of approximately 3 million. This approach has been validated by screening the aforementioned library against the antibacterial target DsbA.
[0019] First, consider an exemplary array with approximately 3 million features available for peptide synthesis. This exemplary array can accommodate the synthesis of up to 3 million unique peptides. As used herein, the term "unique peptide" means that each peptide in a fixed collection of peptides has a unique amino acid sequence compared to each of the other peptides in the collection. For example, two peptides are unique if they differ from each other by at least one amino acid.
[0020] Continuing with the above example, considering the use of all 20 standard amino acids, the number of unique 5-mer peptides that can be prepared is 20 5 , i.e., 3.2 million unique 5-mer peptide sequences. Thus, an array with approximately 3 million features can accommodate most, if not all, of the 5-mer peptide sequences prepared from the 20 standard amino acid building blocks. Such comprehensive 5-mer peptides have demonstrated utility for identifying peptide binders to various targets (see, e.g., U.S. Patent Application No. 15 / 132,951, entitled "Specific Peptide Binders to Proteins Identified via Systematic Discovery, Maturation, and Extension Process"). However, in some cases, it may be useful to provide core binder sequences greater than five amino acids in length. In such cases, it may be desirable to provide an array of unique 6-, 7-, 8-, 9-, 10-, 11-, 12-, 13-, 14-, 15-, 16-, 17-, 18-, 19-, 20-mer, or longer, peptides.
[0021] When considering the use of longer peptides, the number of unique amino acid sequences that can be represented on a single array is largely constrained by the number of available features. For 6-mer peptides prepared from all 20 standard amino acids, the number of unique 6-mer peptides that can be prepared is 20 6, or 64 million unique 6-mer peptide sequences. Given the constraint of 3 million features, a single array can represent at most about 4.7% of all possible unique 6-mer peptide sequences. Alternatively, to represent all 64 million possible 6-mer sequences, 22 separate arrays, each with 3 million unique features, would be required. While this approach may be feasible under selection conditions, it becomes infeasible when moving to peptides with lengths of 7 amino acids or more.
[0022] In one embodiment, it may be possible to select a subset of peptides. For example, it may be possible to only consider the 6-mer peptide sequences present in the human genome, but it has previously been shown that there are sequences not found in the human genome that are relevant to binder discovery for human targets (see at least Patel A, Dong JC, Trost B, Richardson JS, Tohme S, et al. (2012) Pentamers Not Found in the Universal Proteome Can Enhance Antigen Specific Immune Responses and Adjuvant Vaccines. PLoS ONE 7(8):e43802.doi:10.1371 / journal.pone.0043802). Therefore, new approaches are needed to increase the representation of unique peptide sequences without needing to increase the feature capacity of a given platform.
[0023] Toward this goal, the inventors have made the surprising discovery that it is possible to increase the effective number of x-mer peptide sequences represented on a single array by preparing an array of peptides, each having a total length N, where N is greater than x. Turning to Figure 8, an example of a 15-mer peptide is shown as a series of 15 blocks, each representing a single amino acid. The 15-mer peptides define a composite sequence that can be decomposed into a series of overlapping 6-mer elements with a one-amino acid tiling resolution. Effectively, the 15-mer peptide sequence provides up to 10 unique 6-mer peptide sequences. For an array with 3 million features, up to 3 million 15-mer peptides can be prepared, representing up to 30 million unique 6-mer peptide sequences (i.e., 10 6-mer peptides for each of the 30 million 15-mer peptide features). This approach can be generalized to any composite peptide of length N representing multiple x-mer elements. Furthermore, the x-mer elements need not overlap with a one-amino acid tiling resolution. The tiling decomposition can be varied to result in overlaps of 2 or more amino acids, or the x-mer elements may not overlap at all.
[0024] In one embodiment, the engineered peptide library comprises a plurality of peptide features. Each peptide feature comprises at least one peptide, and at least one peptide comprises a complex region having a defined sequence of amino acids of length N. The complex region represents k distinct elements, each distinct element (k) having a defined sequence of amino acids of length x. In particular, x, N, and k are integers, and x is necessarily less than N. In some cases, x is at least 1 less than N (i.e., x≦N−1), and k is at least 2. In some embodiments, k is at least 3, 4, or 5. The total number of distinct elements represented by the engineered peptide library is K Eng and the number of peptide features contained in the engineered peptide library can be defined as F, where K Eng is greater than F. In the above example, K Engis at least 0.8*k*F, indicating that at least 80% of the x-mer peptide elements collectively represented by the 15-mer composite sequence are unique. Eng may be at least 0.85*k*F. In any of the embodiments, K Eng may be at least 0.9*k*F. In any of the embodiments, K Eng may be at least 0.95*k*F. In any of the embodiments, K Eng may be at least 0.99*k*F. In any of the embodiments, K Eng can be at least 0.999*k*F. Depending on the technique used, K Eng may be at least 0.8*k*F, 0.85*k*F, 0.95*k*F, 0.9*k*F, 0.99*k*F, 0.999*k*F, or more.
[0025] Although identifying at least 3 million N-mer sequences that represent only unique x-mer elements (depending on the values selected for N and x) can be computationally challenging, the inventors have further discovered an efficient algorithm that allows for the selection of N-mer peptides whose represented x-mers approach a population of completely unique peptide sequences. According to the present disclosure, an algorithmic approach can be used to prepare a set of N-mer composite sequences in a relatively short time. This algorithm was developed using the general-purpose scripting language Perl. The algorithm generates peptides by randomly selecting an amino acid at each position of the N-mer peptide from a list of available amino acids. The algorithm then tiled the newly generated peptides to identify the presence of all possible x-mer elements, which were added to a list of elements encountered by the algorithm. Next, a new N-mer peptide was generated, performing the same tasks as above, except that if it encountered an x-mer element already present in the list of encountered elements, the newly generated N-mer peptide was discarded. This process was repeated until a user-specified number of N-mer peptides was reached. Additionally, the algorithm tracks the number of times each x-mer element is recognized and allows user control to define the number of allowed repetitions of a given element.
[0026] In particular, this algorithm is very versatile and can be used for any N-mer peptide and x-mer element as long as x < N. In the present disclosure, non-limiting examples of 15-mer composite peptides and 6-mer elements were investigated. Using this approach, it was possible to generate approximately 3 million 15-mer peptides representing over 30 million unique 6-mer peptide sequences, which represents less than half of all possible unique 6-mer sequences prepared from all 20 standard amino acids. Then, for each of the 3 million + 15-mer peptides identified using the algorithm described, a single peptide array was synthesized and that array was used effectively to identify binders to the target DsbA. It should be understood that using an array of 3 million features with each feature having a different 5-mer peptide synthesized (i.e., a 5-mer array in contrast to a 15-mer array) was insufficient to identify binders with the desired properties to DsbA as compared to the use of the 15-mer arrays according to the present disclosure.
[0027] It should be further appreciated that this approach is effective for preparing peptide libraries with enhanced subsequence diversity. That is, many approaches are available for preparing sets of unique N-mer peptide sequences. For example, a list of unique 15-mer peptides can be easily generated where each peptide differs from the next without considering the subsequence (i.e., the x-mer elements represented by the overall sequence of the N-mer). Referring to Table 1, six such sets of unique 15-mer peptides were prepared, regardless of the composition of the subsequence. The resulting peptides were then analyzed to determine the number of unique 6-mer peptides represented therein. The maximum number of possible 6-mer peptides indicates the maximum number of unique 6-mer peptides that can be represented by approximately 3 million 15-mer peptides. The actual number of unique 6-mer peptides then accurately indicates the number of unique 6-mers ultimately represented by the lists of unique 15-mer peptide sequences randomly prepared for each of the six sets. The final column indicates the percentage of represented 6-mers compared to the maximum possible number of 6-mers. In all six cases, it was determined that 15-mer peptides represented approximately 73.6% of the maximum possible number.
[0028] Table 1. Unique 6-mer peptide sequences represented by randomly unique 15-mer peptide sequences prepared by the comparative method. TIFF2025138750000002.tif47169
[0029] In contrast to the data presented in Table 1, by using the disclosed approach, it was possible to increase the diversity of 6-mer peptide sequences represented by a population of 15-mer peptides to nearly 100%. In particular, the disclosed algorithm was able to obtain an equal number of 15-mer peptides representing 30,325,760 unique 6-mer peptide sequences, or 99.6% of the maximum number of unique 6-mers possible with this approach. Without being limited by theory, it is hypothesized that by increasing the local diversity within each N-mer peptide, it is possible to increase the overall sequence diversity represented on a given peptide array, thereby enhancing the ability to effectively identify peptide binders for a given target. Table 2 further illustrates the x-mer representation of a series of N-mer arrays.
[0030] Table 2: Unique X-mers represented by randomly unique N-mer peptide sequences prepared by the method of the present invention TIFF2025138750000003.tif81159
[0031] In Table 2, the first row (5-mer array) represents a 5-mer design that includes all 5-mer sequences prepared from all 20 standard amino acids except methionine. This approach provides a total of 3,035,196 unique peptides, representing 94.9% of all possible 5-mer sequences prepared from all 20 standard amino acids. Because the length of each peptide is limited to 5 amino acids, the array does not necessarily represent 6-mer peptide sequences. The second array contains 16-mer peptides tiled across the entire human proteome. This design represents 73.3% of all possible 5-mer peptides prepared from all 20 standard amino acids. Notably, this number is far less than 100%, because the human proteome does not include all possible 5-mer peptide sequences prepared from all 20 standard amino acids. Each 16-mer peptide in this design can represent up to 11 unique 6-mer elements, but because the design is not optimized for 6-mers and is merely representative of the human proteome, only 12.4% of all possible 6-mer peptide sequences are represented.
[0032] Here, an array of 15-mer sequences selected for both uniqueness and subsequence diversity, taking into account the "pseudo-6-mer" design in line 3, was prepared according to the methods disclosed herein. This design further excluded the use of methionine. The resulting library represented 77.4% of all possible 5-mer sequences prepared from all 20 standard amino acids and 47.4% of all possible 6-mer sequences prepared from all 20 standard amino acids, using a single array with approximately 3 million features. Notably, this final approach significantly expands the subsequence diversity of the 6-mer elements of comprehensive 15-mer composite peptides.
[0033] In some embodiments, N is at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids. In some embodiments, N is at least 7, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids. In some embodiments, N is 6-20 amino acids. In some embodiments, N is 7-16 amino acids.
[0034] In some embodiments, x is at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19. In some embodiments, x is at least 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14. In some embodiments, x is 5-19 amino acids. In some embodiments, x is 5-12 amino acids. In some embodiments, x is 6-14 amino acids. In some embodiments, x is 6-10 amino acids. In some embodiments, x is 6-9 amino acids.
[0035] In some embodiments, the plurality of peptides represents at least about 80%, 85%, 90%, or 95%, hi some embodiments, the plurality of peptides represents about 80-100%, 85-100%, 90-100%, 95-100%, 80-99.9%, 85-99.9%, 90-99.9%, 95-99.9%, 80-99%, 85-99%, 90-99%, or 95-99% of the target proteome.
[0036] For applications of peptide arrays according to the present disclosure, it is generally useful to evaluate a population of peptide features by probing the population for the presence of receptors that have affinity for multiple binder sequences. Receptors include any peptide, protein, antibody, small molecule, or other similar structure that can specifically bind a given peptide sequence or feature. Generally, to determine whether a receptor binds to a particular peptide or peptide feature, a side of the receptor should be detectable. For example, the receptor itself may contain a fluorophore that is detectable by fluorescence microscopy. Alternatively (or additionally), the receptor may be bound by a secondary molecule, such as a fluorescent antibody. Additional techniques would also be within the scope of the present disclosure.
[0037] As described above, the receptor can bind to or otherwise interact with a known binder or affinity sequence. One example of a binder sequence is a defined amino acid sequence or motif. The defined amino acid sequence can represent at least a portion of a full-length peptide within a synthetic peptide population. However, the binder sequence can itself be a full-length peptide. For example, the eight amino acid peptide sequence Trp-Ser-His-Pro-Gln-Phe-Glu-Lys (SEQ ID NO: 1), known as the "Strep-tag," exhibits unique affinity for an engineered form of the protein streptavidin. According to the present disclosure, the Strep-tag can be incorporated into either the N- or C-terminus of a given peptide, or even into an intermediate point within the peptide. A peptide population containing peptides consisting of (or including) the Strep-tag binder sequence can then be bound by the streptavidin receptor. Binding of streptavidin to the Strep-tag sequence can then be detected using various techniques. Further examples of binder sequences include hexahistidine tags (His tags) (SEQ ID NO: 2), FLAG tags, calmodulin-binding peptides, covalently linked but separable peptides, the heavy chain of a protein C-tag, etc. Alternative (or additional) binder sequence-receptor pairs are also within the scope of this disclosure.
[0038] With continued reference to binder sequences as disclosed herein, each binder sequence will have a specific or defined amino acid sequence. A binder sequence can contain at least three amino acids. Exemplary binder sequences disclosed herein contain from about 5 amino acids to about 12 amino acids. However, binder sequences having fewer than 5 or more than 12 amino acids can also be used. The position of each amino acid in a particular binder sequence can define a start at either the N-terminus ([N]) or C-terminus ([C]). For example, the position of an amino acid in the aforementioned Strep-tag binder sequence can be defined as [N]-Trp-Ser-His-Pro-Gln-Phe-Glu-Lys-[C] (SEQ ID NO: 1). Thus, the position of the amino acid histidine (His) is defined as the third amino acid from the N-terminus of the Strep-tag binder sequence. Notably, as noted above, the Strep-tag binder sequence can be flanked by one or more additional amino acids at either or both the N-terminus and C-terminus.
[0039] The method according to the present disclosure further includes detecting a signal output characteristic of the interaction between the receptor and the first control peptide feature. Detecting a signal output may include any method of monitoring or otherwise observing a measurable aspect of one or more peptides or peptide features in a population of peptides in the presence or absence of a receptor. Examples of signal outputs include optical output (e.g., luminescence), electrical output, chemical output, etc., and combinations thereof. Consequently, detecting a signal output may include measuring, recording, or otherwise observing the signal output using any suitable device. Examples of devices include optical and digital detection devices such as fluorescence microscopes and digital cameras. In some embodiments, detecting a signal output further includes perturbations such as excitation with one or more wavelengths of light, thermal manipulation, introduction of one or more chemical reagents, the like, and combinations thereof. In particular, a synthetic peptide population may include a population of peptide features that are synthesized to all include alternative building blocks, such as unnatural amino acids, amino acid derivatives, or other monomer units.
[0040] II. Peptides According to various embodiments of the present disclosure, peptides (e.g., control peptides, peptide binder sequences) are disclosed. Each peptide comprises two or more natural or unnatural amino acids, as described herein. In the examples provided herein, linear forms of the peptides are shown. However, one of skill in the art will readily understand that peptides can be converted to cyclic forms by reacting the N-terminus with the C-terminus, as disclosed, for example, in U.S. Patent Publication No. 2015 / 0185216, filed December 19, 2014, to Albert et al. Thus, embodiments of this technology include both cyclic and linear peptides.
[0041] As used herein, the terms "peptide," "oligopeptide," and "peptide binder" refer to organic compounds composed of amino acids, which may be arranged in a linear chain (joined together by peptide bonds between the carboxyl and amino groups of adjacent amino acid residues), in a cyclic form (cyclized using internal sites), or in a constrained form (e.g., a "macrocycle" in a head-to-tail cyclized form). The term "peptide" or "oligopeptide" also refers to shorter polypeptides, i.e., organic compounds composed of fewer than 50 amino acid residues. As used herein, macrocycle (or constrained peptide) is used in its conventional sense to describe cyclic small molecules, such as peptides of about 500 daltons to about 2000 daltons.
[0042] The term "natural amino acid" or "standard amino acid" refers to one of the 20 amino acids commonly found in proteins and used in protein biosynthesis, as well as other amino acids (including pyrrolysine and selenocysteine) that can be incorporated into proteins during translation. The 20 naturally occurring amino acids include the L-stereoisomers of histidine (His; H), alanine (Ala; A), valine (Val; V), glycine (Gly; G), leucine (Leu; L), isoleucine (Ile; I), aspartic acid (Asp; D), glutamic acid (Glu; E), serine (Ser; S), glutamine (Gln; Q), asparagine (Asn; N), threonine (Thr; T), arginine (Arg; R), proline (Pro; P), phenylalanine (Phe; F), tyrosine (Tyr; Y), tryptophan (Trp; W), cysteine (Cys; C), methionine (Met; M), and lysine (Lys; K). The term "all 20 amino acids" refers to the 20 naturally occurring amino acids.
[0043] The term "unnatural amino acid" refers to an organic compound that is not encoded by the standard genetic code or incorporated into proteins during translation. Thus, unnatural amino acids include amino acids or amino acid analogs, including, but not limited to, the D-stereoisomers of all 20 amino acids, the beta-amino analogs of all 20 amino acids, citrulline, homocitrulline, homoarginine, hydroxyproline, homoproline, ornithine, 4-amino-phenylalanine, cyclohexylalanine, α-aminoisobutyric acid, N-methyl-alanine, N-methyl-glycine, norleucine, N-methyl-glutamic acid, tert-butylglycine, α-aminobutyric acid, tert-butylalanine, 2-aminoisobutyric acid, α-aminoisobutyric acid, 2-aminoindan-2-carboxylic acid, selenomethionine, dehydroalanine, lanthionine, γ-aminobutyric acid, and derivatives thereof in which the amine nitrogen is mono- or di-alkylated.
[0044] According to embodiments of the present disclosure, peptides are presented immobilized on a support surface (e.g., a microarray, beads, etc.). In some embodiments, peptides selected for use as control peptides can optionally undergo one or more rounds of extension and maturation processes to obtain the control peptides disclosed herein.
[0045] III. Microarray The peptides disclosed herein can be generated using oligopeptide microarrays. As used herein, the term "microarray" refers to a two-dimensional arrangement of features on the surface of a solid or semi-solid support. A single microarray, or in some cases multiple microarrays (e.g., 3, 4, 5, or more microarrays), can be arranged on a single solid support. For fixed-dimensional solid supports, the size of the microarray depends on the number of microarrays on the solid support. That is, the more microarrays per solid support, the smaller the array must be to fit on the solid support. The array can be designed in any shape, but is preferably designed as a square or rectangle. The ready-to-use product is an oligopeptide microarray on a solid or semi-solid support (microarray slide).
[0046] The term "peptide microarray" or "oligopeptide microarray" or "peptide chip" or "peptide epitope microarray" refers to a microarray, i.e., a population or collection of peptides arranged on a solid surface, such as a glass, carbon composite, or plastic array, slide, or chip.
[0047] The term "feature" refers to a defined region on the surface of a microarray. Features include biomolecules such as peptides (i.e., peptide features), nucleic acids, carbohydrates, etc. One feature may contain biomolecules with different properties, such as a different sequence or orientation, compared to other features. The size of a feature is determined by two factors: i) the number of features on the array (the more features on the array, the smaller each individual feature will be); and ii) the number of individually addressable aluminum mirror elements used to illuminate a feature. The more mirror elements used to illuminate a feature, the larger each individual feature will be. The number of features on an array may be limited by the number of mirror elements (pixels) present on the micromirror device. For example, a state-of-the-art micromirror device from Texas Instruments, Inc. (Dallas, Tex.) currently contains 4.2 million mirror elements (pixels); therefore, the number of features in such an exemplary microarray is limited by this number. However, other micromirror devices allow for higher density arrays.
[0048] The term "solid or semi-solid support" refers to any solid material having a surface area to which organic molecules can be attached by bond formation or absorbed by electronic or static interactions, such as covalent bonds or complexation through specific functional groups. The support can be a combination of materials such as plastic on glass, carbon on glass, etc. Functional surfaces can be simple organic molecules, but can also include copolymers, dendrimers, molecular brushes, etc.
[0049] The term "plastic" refers to synthetic materials such as homo- or hetero-copolymers of organic building blocks (monomers) with functionalized surfaces so that organic molecules can be attached by covalent bond formation or absorbed by electronic or static interactions, such as bond formation via functional groups. Preferably, the term "plastic" refers to polyolefins, which are polymers derived from the polymerization of olefins (e.g., ethylene propylene diene monomer polymer, polyisobutylene). Most preferably, the plastic is a polyolefin with defined optical properties, such as TOPAS® or ZEONOR / EX®.
[0050] The term "functional group" refers to any of a number of combinations of atoms that form part of a chemical molecule and that undergo characteristic reactions themselves and affect the reactivity of the remainder of the molecule. Typical functional groups include, but are not limited to, hydroxyl, carboxyl, aldehyde, carbonyl, amino, azide, alkynyl, thiol, and nitrile. Potentially reactive functional groups include, for example, amines, carboxylic acids, alcohols, double bonds, and the like. Preferred functional groups are the potentially reactive functional groups of amino acids, such as amino or carboxyl groups.
[0051] Various methods for producing oligopeptide microarrays are known in the art. For example, spotting of pre-made peptides or in situ synthesis by spotting reagents (e.g., on a membrane) are examples of known methods. Another known method used to generate higher density peptide arrays is the so-called photolithography technique, in which the synthetic design of the desired biopolymer is controlled by suitable photolabile protecting groups (PLPGs), which release binding sites for each subsequent component (amino acid, oligonucleotide) upon exposure to electromagnetic radiation such as light (Fodor et al., (1993) Nature 364:555-556; Fodor et al., (1991) Science 251:767-773). Currently, two different photolithography techniques are known. The first is a photolithography mask, which is used to direct light to specific regions of the synthesis surface to locally deprotect the PLPGs. "Masked" methods involve polymer synthesis using a mount (e.g., a "mask") that engages with the substrate and provides a reaction space between the substrate and the mount. Exemplary embodiments of such "masked" array synthesis are described, for example, in U.S. Pat. Nos. 5,143,854 and 5,445,934, the disclosures of which are incorporated herein by reference. However, potential drawbacks of this technique include the need for multiple masking steps, which result in relatively low overall yields and high costs; for example, the synthesis of a peptide only six amino acids in length may require more than 100 masks. A second photolithography technique is so-called maskless photolithography, in which light is directed at specific regions of the synthesis surface, resulting in localized deprotection of PLPG by digital projection techniques such as micromirror devices (Singh-Gasson et al., Nature Biotechn. 17 (1999) 974-978). Thus, such "maskless" array synthesis eliminates the need for time-consuming and expensive fabrication of exposure masks. It should be understood that embodiments of the systems and methods disclosed herein may include or utilize any of the various array synthesis techniques described above.
[0052] The use of PLPG (photolabile protecting group) to provide the basis for photolithography-based synthesis of oligopeptide microarrays is well known in the art. PLPGs commonly used in photolithography-based biopolymer synthesis include, for example, α-methyl-6-nitropiperonyloxycarbonyl (MeNPOC) (Pease et al., Proc. Natl. Acad. Sci. USA (1994) 91:5022-5026), 2-(2-nitrophenyl)-propoxycarbonyl (NPPOC) (Hasan et al. (1997) Tetrahedron 53:4247-4264), nitroveratryloxycarbonyl (NVOC) (Fodor et al. (1991) Science 251:767-773), and 2-nitrobenzyloxycarbonyl (NBOC).
[0053] Amino acids were introduced into the photolithographic solid-phase peptide synthesis method for oligopeptide microarrays. They were protected with NPPOC as a photocleavable amino-protecting group, and glass slides were used as supports (US Patent Publication No. 20050101763). The method using NPPOC-protected amino acids has the disadvantage that all (except one) protected amino acids exhibit half-lives upon irradiation with light within the range of approximately 2–3 minutes under certain conditions. In contrast, under the same conditions, NPPOC-protected tyrosine exhibits a half-life of approximately 10 minutes. Because the overall rate of the synthesis process depends on the slowest subprocess, this phenomenon increases the synthesis process time by 3–4 times. Concomitantly, the extent of damage to growing oligomers by photogenerated radical ions increases with increasing excess light dose requirements.
[0054] As will be understood by one of skill in the art, a peptide microarray comprises thousands (or in the case of the present disclosure, millions) of peptides (presented in multiple copies in some embodiments) bound or immobilized to a solid support (which in some embodiments includes a glass, carbon composite, or plastic chip or slide).
[0055] In some embodiments, the peptide microarray is exposed to a sample of interest, such as a receptor, antibody, enzyme, peptide, oligonucleotide, etc. The peptide microarray exposed to the sample of interest undergoes one or more washing steps and then undergoes a detection process. In some embodiments, the array is exposed to an antibody targeting the sample of interest (e.g., anti-IgG human / mouse, anti-phosphotyrosine, or anti-myc). Typically, the secondary antibody is tagged with a fluorescent label that can be detected with a fluorescent scanner. Other detection methods are chemiluminescence, colorimetry, or autoradiography. In other embodiments, the sample of interest is biotinylated and then detected with streptavidin conjugated to a fluorophore. In yet other embodiments, the protein of interest is tagged with a specific tag, such as a His tag, a FLAG tag, or a Myc tag, and detected with a fluorophore-conjugated antibody specific to the tag.
[0056] After scanning the microarray slide, the scanner records a 20-bit, 16-bit, or 8-bit numerical image in tagged image file format (*.tif). The tif image allows the interpretation and quantification of each fluorescent spot on the scanned microarray slide. This quantitative data is the basis for performing statistical analyses on the binding events or peptide modifications measured on the microarray slide. For the evaluation and interpretation of the detected signals, it is necessary to perform an assignment of peptide spots (shown in the image) and the corresponding peptide sequences.
[0057] Peptide microarrays are slides onto which peptides are spotted or assembled directly on the surface by in situ synthesis. Ideally, peptides are covalently attached via chemoselective bonding, directing peptides in the same orientation as interaction profiling. Alternative procedures include nonspecific covalent bonding and adhesive immobilization.
[0058] According to a specific embodiment of the present disclosure, specific peptide binders are identified using maskless array synthesis to fabricate peptide binder probes on a substrate. According to such an embodiment, the maskless array synthesis used allows for ultra-high-density peptide synthesis of up to 2.9 million unique peptides. Each of the 2.9 million features / regions has up to 107 reaction sites with the potential to generate full-length peptides. Smaller arrays can also be designed. For example, an array representing a comprehensive list of all possible 5-mer peptides using the 19 natural amino acids excluding cysteine would have 2,476,099 peptides. In other examples, arrays can include unnatural and natural amino acids. Arrays of 5-mer peptides using all combinations of the 18 natural amino acids excluding cysteine and methionine can also be used. Furthermore, arrays can exclude other amino acids or amino acid dimers. In some embodiments, arrays can be designed to generate a library of 1,360,732 unique peptides, excluding any dimers or longer repeats of the same amino acid, as well as any peptides containing the HR, RH, HK, KH, RK, KR, HP, and PQ sequences. Smaller arrays can have replicates of each peptide on the same array to increase the reliability of conclusions drawn from the array data.
[0059] In various embodiments, the peptide arrays described herein comprise at least 1.0×10 mAb bound to a solid support of the peptide array. 5 , 1.2x10 5 , 1.4x10 5 , 1.6x10 5 , 1.8x10 5 , 2.0x10 5 , 1.6x10 6 , 1.8x10 6 , 2.0x10 6 of peptides, and / or up to about 1.0x10 7 , 5.0x10 7 , 8.0x10 7 , 1.0x10 8As described herein, a peptide array comprising a particular number of peptides can refer to a single peptide array on a single solid support, or the peptides can be resolved and attached to multiple solid supports to achieve the number of peptides described herein.
[0060] Arrays synthesized according to such embodiments can be designed for the discovery of peptide binders, in linear or circular form (as described herein), with or without modifications such as N-methyl or other post-translational modifications. Arrays can also be designed to further extend potential binders using a block approach by performing iterative screening at the N- and C-termini of potential hits (as described in further detail herein). Once ideal affinity hits are identified, they can be further matured using a combination of maturation arrays (described further herein), which allow for insertion, deletion, and substitution analysis of various amino acid combinations, both natural and unnatural.
[0061] The peptide arrays of the present disclosure are used to identify specific binders or binder sequences of the present technology, as well as for maturation and extension of binder sequences for use in the design and selection of control peptides.
[0062] IV. Search for peptide binders In one aspect, the present disclosure provides for the discovery of novel binders. Referring now to FIG. 1 , according to one embodiment of the present disclosure, a peptide array 100 can be designed to include a population of hundreds, thousands, tens of thousands, hundreds of thousands, or even millions of peptides 102. In some embodiments, the population of peptides 102 can be configured so that the peptides 102 collectively represent an entire protein, gene, chromosome, or even an entire genome of interest (e.g., the human proteome). Furthermore, the peptides 102 can be configured according to specific criteria, such that certain amino acids or motifs are excluded. Furthermore, the peptides 102 can be configured so that each of the peptides 102 includes the same length. For example, in some embodiments, the population of peptides 102 immobilized on the array substrate 104 can all include 3-, 4-, 5-, 6-, 7-, 8-, 9-, 10-, 11-, or even 12-mers or more. In some embodiments, the peptides 102 may also each comprise an N-terminal sequence (N-terminus 106) or a C-terminal sequence (C-terminus 108), where each peptide 102 comprises both an N-terminal sequence and a C-terminal peptide sequence of a specific and identical length (e.g., a 3-, 4-, 5-, 6-, 7-, or even 8-mer, or longer). In particular, the sequence of a peptide at a specific location on the array is known.
[0063] According to some embodiments, the peptide array 100 is designed to contain a population of up to 2.9 million peptides 102, configured to represent a comprehensive list of all possible 5-mer probe peptides 110 in the genome immobilized on the array substrate 104. In some such embodiments, the 5-mer probe peptides 110 (comprising the 2.9 million peptides of the array) may exclude one or more of the 20 amino acids. For example, Cys may be excluded to help control abnormal peptide folding. The amino acid Met may be excluded as it is a rare amino acid in the proteome. Other selective exclusions are repeats of two or more amino acids of the same amino acid (to help control nonspecific interactions such as charge and hydrophobic interactions) or a specific amino acid motif (e.g., in the case of streptavidin binders), such as the His-Pro-Gln sequence, where His-Pro-Gln is a known streptavidin-binding motif. Continuing with reference to FIG. 1 , in some exemplary embodiments, the 5-mer probe peptides 110 may exclude one or more of the above amino acids or amino acid motifs. One embodiment of this technology includes a peptide array 100 comprising a population of up to 2.9 million peptides 102, where the 5-mer probe peptide 110 portion of the peptides 102 represents the entire human genome. In one example, the 5-mer probe peptides 110 do not include the amino acids Cys and Met, do not include amino acid repeats of two or more amino acids, and do not include the amino acid motif His-Pro-Gln. Another embodiment of this technology includes a peptide array comprising up to 2.9 million peptides 102 comprising 5-mer probe peptides 110 representing the protein content encoded by the entire human genome, where the 5-mer probe peptides 110 do not include the amino acids Cys and Met, and do not include amino acid repeats of two or more amino acids.
[0064] According to further embodiments, as shown in FIG. 1 , each 5-mer probe peptide 110 comprising a population of up to 2.9 million peptides 102 in a peptide array 100 can be synthesized in five cycles of wobble synthesis at each of the N-terminus 106 and C-terminus 108. As used herein, "wobble synthesis" refers to the synthesis (by any means disclosed herein) of a sequence of peptides (either fixed or random) located at the N-terminus or C-terminus of a 5-mer probe peptide 110 of interest. As shown in FIG. 1 , the particular amino acid comprising wobble synthesis at either the N-terminus 106 or C-terminus 108 is represented by "Z." According to various embodiments, wobble synthesis can include any number of amino acids or other monomer units at the N-terminus 106 or C-terminus 1-8. For example, each of the N-terminus 106 and C-terminus 108 can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more (e.g., 15-20) amino acids. Additionally, the wobble synthesis can include N-terminus and C-terminus with the same or different numbers of wobble synthesis amino acids.
[0065] According to various embodiments, the N-terminal 106 and C-terminal 108 wobble oligopeptide compositions are flexible with respect to amino acid composition and with respect to amino acid ratios or concentrations. For example, a wobble oligopeptide composition can include a mixture of two or more amino acids. An exemplary embodiment of a flexible wobble mix includes a wobble oligopeptide composition of Gly and Ser in a ratio of 3:1 (Gly:Ser). Other examples of flexible wobble mixes include equal concentrations (e.g., equal ratios) of the amino acids Gly, Ser, Ala, Val, Asp, Pro, Glu, Leu, and Thr, equal concentrations (e.g., equal ratios) of the amino acids Leu, Ala, Asp, Lys, Thr, Gln, Pro, Phe, Val, and Tyr, and combinations thereof. Other examples include N-terminal 106 and C-terminal 108 wobble oligopeptide compositions that include equal concentrations of any of the 20 standard amino acids.
[0066] As disclosed herein, wobble oligopeptide synthesis in various embodiments allows for the generation of peptides on an array with a combination of random and directed synthesized amino acids. For example, oligopeptide probes on an array may include combined 15-mer peptides with peptide sequences in the following format: ZZZZZ-[5-mer]-ZZZZZ, where Z is an amino acid from a particular wobble amino acid mixture. In another aspect, ZZZZZ can be abbreviated as 5Z, where nZ corresponds to n consecutive amino acids selected from the set of amino acids comprising the wobble amino acid mixture.
[0067] In some embodiments, the features are about 10 7 In some such embodiments, the complexity of the population of each feature can vary depending on the complexity of the wobble mixture. As disclosed herein, by using wobble synthesis in semi-directed synthesis to create such complexity, up to 10 12 This allows for screening of binders on arrays using peptides with unique sequence diversity. An example of binder screening for streptavidin is shown below. However, additional protein targets such as prostate-specific antigen, urokinase, or tumor necrosis factor are also possible according to the described methods and systems.
[0068] It was further discovered that the linkers (e.g., N-terminal 106 and C-terminal 108) can be varied in length and are optional. In some embodiments, a 3Z or 1Z linker can be used instead of a 5Z linker. In such embodiments, Z can be synthesized using a random mixture of all 20 amino acids. It was discovered that the same target can generate additional 5-mer binder sequences when using a 1Z linker or no linker. It was discovered that varying the length of the linker or deleting the linker identified additional peptide binders that were not found, for example, when using the original 5Z linker.
[0069] Indeed, with reference to FIG. 1 , a peptide array 100 includes an array substrate 104 including a solid support 112 having a reactive surface 114 (e.g., a reactive amine layer) with a population of peptides 102 (e.g., a population of 5-mers representing the entire human proteome, etc.) immobilized thereon. Exemplary 5-mer peptides comprising the population of peptides 102 according to such embodiments include no amino acids Cys or Met, no amino acid repeats of more than one amino acid, and no amino acid motif His-Pro-Gln. According to the embodiment shown in FIG. 1 , a population of peptides 102 representing the entire human proteome would include 1,360,732 individual peptides comprising the population of peptides 102. In some embodiments, replicates or repeats may be located on the same array. For example, a population of peptides 102 comprising a single replicate would include 2,721,464 individual features. Additionally, peptide 102 includes N- and C-terminal wobble synthetic oligopeptides (i.e., N-terminus 106 and C-terminus 108), respectively. In one example, N-terminus 106 and C-terminus 108 each have five amino acids, each randomly selected from a mixture of Gly and Ser in a ratio of 3:1 (Gly:Ser). The wobble oligopeptides forming N-terminus 106 and C-terminus 108 can be omitted or replaced with a single amino acid selected from any of the 20 standard amino acids, an unnatural amino acid (e.g., 6-aminohexanoic acid), or a random mixture of combinations thereof. Some embodiments can include a non-amino acid moiety (e.g., polyethylene glycol).
[0070] Referring now generally to FIG. 2, a process 200 for preparing a peptide array (e.g., peptide array 100 shown in FIG. 1) includes step 202 of peptide binder discovery. In one example of step 202, the peptide array is exposed to an enriched and purified protein of interest (as in standard microarray practice), whereby the protein of interest may bind to or otherwise interact with one or more of a population of peptides (e.g., a population of peptides 102 as shown in FIG. 1). In one embodiment, the protein of interest may bind to a selected one of the population of peptides independently of another one of the peptides that make up the population. After exposure to the protein of interest, binding of the protein of interest to the peptide binder is assayed, for example, by exposing the array to an antibody (specific for the protein) to which a reportable label (e.g., peroxidase) is attached. Because the peptide sequence of each 5-mer at each location on the array is known, the sequence (and binding strength) of protein binding to specific 5-mer peptides can be graphed, quantified, or compared. One such method for comparing protein binding to peptides comprising a population is to examine binding with principled analytical distribution-based clustering, such as that described by White et al. (Standardizing and Simplifying Analysis of Peptide Library Data, Chem. Inf. Model., 2013, 53(2), pp. 493-499) and presented herein. As exemplified herein, clustering of protein-5-mer binding (also known as "hits," as depicted in principled analytical distribution-based clustering) reveals 5-mers with overlapping peptide sequences. As shown in more detail below, from the overlapping peptide sequences (in each cluster), "core hit" peptide sequences (e.g., peptide sequences shared by prominent protein-peptide binding events in the array) can be identified, or at least hypothesized and constructed for further evaluation. In one aspect, an array such as that exemplified herein can identify multiple core hit peptide sequences.Furthermore, the core hit peptide sequences may contain more amino acids than the 5-mer peptide binders comprising the population of peptides, for example, due to the potential for identifying overlapping and adjacent sequences during principled analytical distribution-based clustering.
[0071] V. Peptide Maturation Continuing with reference to Figure 2, upon identification of core hit peptide sequences (through the peptide binder discovery 202 process disclosed, described, and exemplified herein), step 204 of process 200 involves peptide maturation, whereby the core hit peptide sequences are modified in various ways (via amino acid substitutions, deletions, and insertions) at each core hit peptide position to further optimize or validate suitable core hit sequences. For example, according to some embodiments (e.g., when the core hit peptide sequences comprise a given number of amino acids), a maturation array is generated. According to the present disclosure, a maturation array has immobilized thereon a population of core hit peptides, whereby each amino acid in the core hit peptide undergoes an amino acid substitution at each position.
[0072] To further illustrate the process of hit maturation or peptide maturation 204, an exemplary or hypothetical core hit peptide is described as consisting of a 5-mer peptide having the amino acid sequence -M1M2M3M4M5- (SEQ ID NO: 3). In accordance with the present disclosure, hit maturation 204 can include any or all combinations of amino acid substitutions, deletions, and insertions at positions 1, 2, 3, 4, and 5. For example, with respect to the hypothetical core hit peptide -M1M2M3M4M5- (SEQ ID NO: 3), embodiments of the present disclosure can include the amino acid M at position 1 being substituted with each of the other 19 amino acids (e.g., A1M2M3M4M5- (SEQ ID NO: 4), P1M2M3M4M5- (SEQ ID NO: 5), V1M2M3M4M5- (SEQ ID NO: 6), Q1M2M3M4M5- (SEQ ID NO: 7), etc.). Each position (2, 3, 4, and 5) can also have the amino acid M substituted with each of the other 19 amino acids (e.g., a substitution at position 2 would be similar to M1A2M3M4M5- (SEQ ID NO: 8), M1Q2M3M4M5- (SEQ ID NO: 9), M1P2M3M4M5- (SEQ ID NO: 10), M1N2M3M4M5- (SEQ ID NO: 11), etc.). It is understood that the peptides (immobilized on the array) are generated to include core hit peptides containing one or more substitutions, deletions, insertions, or combinations thereof.
[0073] In some embodiments of process 200, step 204 of peptide maturation involves preparing a double amino acid substitution library. A double amino acid substitution involves changing an amino acid at a first position in combination with substituting the amino acid at the second position with each of the other 19 amino acids. This process is repeated until all possible combinations of first and second positions have been combined. As an example, referring back to a hypothetical core hit peptide having a 5-mer peptide of amino acid sequence -M1M2M3M4M5- (SEQ ID NO: 3), double amino acid substitutions for positions 1 and 2 could include, for example, an M→P substitution at position 1, followed by substitutions of all 20 amino acids at position 2 (e.g., -P1A2M3M4M5- (SEQ ID NO: 12), -P1F2M3M4M5- (SEQ ID NO: 13), -P1V2M3M4M5- (SEQ ID NO: 14), -P1E2M3M4M5- (SEQ ID NO: 15), etc.), an M→V ... and substitutions of all 20 amino acids at position 2 (e.g., -V1A2M3M4M5- (SEQ ID NO: 16), -V1F2M3M4M5- (SEQ ID NO: 17), -V1V2M3M4M5- (SEQ ID NO: 18), -V1E2M3M4M5- (SEQ ID NO: 19), etc.), an M to A substitution at position 1 followed by substitutions of all 20 amino acids at position 2 (e.g., -A1A2M3M4M5- (SEQ ID NO: 20), -A1F2M3M4M5- (SEQ ID NO: 21), -A1V2M3M4M5- (SEQ ID NO: 22), -A1E2M3M4M5- (SEQ ID NO: 23), etc.).
[0074] In some embodiments of step 204 of peptide maturation according to the present disclosure, amino acid deletion can be performed at each amino acid position of the core hit peptide. Amino acid deletion involves preparing peptides that contain the core hit peptide sequence, but by deleting a single amino acid from the core hit peptide sequence (thus creating peptides with an amino acid deleted at each position). As an example, referring back to a hypothetical core hit peptide having a 5-mer peptide with the amino acid sequence -M1M2M3M4M5- (SEQ ID NO: 3), amino acid deletion would involve preparing a series of peptides with the following sequences: -M2M3M4M5- (SEQ ID NO: 24), M1M3M4M5- (SEQ ID NO: 24), -M1M2M4M5- (SEQ ID NO: 24), -M1M2M3M5- (SEQ ID NO: 24), and -M1M2M3M4- (SEQ ID NO: 24). It should be noted that following amino acid deletion of the hypothetical 5-mer, five new 4-mers are created. According to some embodiments of the present disclosure, an amino acid substitution or double amino acid substitution scan can be performed on each new 4-mer generated.
[0075]
[0001] Similar to the amino acid deletion scan described above, some embodiments of the peptide maturation step 204 disclosed herein may include an amino acid insertion scan, whereby each of the 20 amino acids is inserted before and after every position of the core hit peptide. As an example, referring back to a hypothetical core hit peptide having a 5-mer peptide of amino acid sequence -M1M2M3M4M5- (SEQ ID NO: 3), an amino acid insertion scan would look like this: -XM1M2M3M4M5- (SEQ ID NO: 25), -M1XM2M3M4M5- (SEQ ID NO: 26), -M1M2XM3M4M5- (SEQ ID NO: 27), -M1M2M3XM4M5- (SEQ ID NO: 28), -M1M2M3M4XM5- (SEQ ID NO: 29), and -M1M2M3M4M5X- (SEQ ID NO: 30) (where X represents an individual amino acid selected from the 20 naturally occurring amino acids or a specific defined subset of amino acids, whereby peptide copies are made for each of the 20 amino acids or defined subset of amino acids).
[0076] It should also be understood that the above-described amino acid substitution peptides, double amino acid substitution peptides, amino acid deletion scan peptides, and amino acid insertion scan peptides can also include one or both of N-terminal and C-terminal wobble amino acid sequences (similar to that described for N-terminal 106 and C-terminal 108 in FIG. 1 ). Similar to the N-terminal and C-terminal wobble amino acid sequences described in FIG. 1 (N-terminal 106 and C-terminal 108), the N-terminal and C-terminal wobble amino acid sequences can include as few as one amino acid or as many as 15 or 20 amino acids, and the N-terminal wobble amino acid sequence can be the same length as, longer than, or shorter than the C-terminal wobble amino acid sequence. In another embodiment, either or both of the N-terminal and C-terminal wobble sequences can be omitted entirely. Furthermore, the N-terminal and C-terminal wobble amino acid sequences can include any defined group of amino acids in any given ratio. For example, a wobble amino acid sequence can contain glycine and serine in a ratio of 3:1 (Gly:Ser), or a random mixture of all 20 standard amino acids.
[0077] In one embodiment of step 204, a core hit peptide having 7 amino acids undergoes exhaustive single and double amino acid screening to include wobble amino acid sequences at both the N- and C-termini. In this example, the N- and C-terminal sequences each contain three amino acids (all glycines). In other embodiments, different terminal sequences can be added by using different mixtures of amino acids during the maturation process. Any single amino acid or any mixture of two or more amino acids can be used. In yet another embodiment, a mixture of Gly and Ser in a ratio of 3:1 (Gly:Ser) is used. In other embodiments, a "random mixture" consisting of a random mixture of all 20 amino acids is used. In some embodiments, an unnatural amino acid (e.g., 6-aminohexanoic acid) is used. Additionally, some embodiments include a non-amino acid moiety (e.g., polyethylene glycol).
[0078] Once various substitution, deletion, and insertion variants of the core hit peptide have been prepared (e.g., immobilized on a solid support such as a microarray), the strength of binding of the purified and enriched target protein is assayed. As shown in the examples provided below, the process of hit maturation allows the core hit peptide to be refined into amino acid sequences that demonstrate the most favorable amino acid sequence for binding to the target protein with the highest affinity.
[0079] VI. Peptide Elongation (N-Terminal and C-Terminal) Motifs identified in 5-mer array experiments may represent only short versions of optimal protein binders. In one aspect, the present invention includes a strategy for identifying longer motifs by extending sequences selected from 5-mer array experiments by one or more amino acids from either or both the N- and C-termini. Starting with a selected peptide, one can add one or more amino acids to each of the N- and C-termini to create an extension library for further selection. For example, starting with a single peptide, one can create an extension library of 160,000 unique peptides using all 20 naturally occurring amino acids. In some embodiments, each extended peptide is synthesized in replicates.
[0080] Referring now to step 206 of process 200 of Figure 2, once the core hit peptide has been matured in step 204 (such that a more optimal amino acid sequence of the core hit peptide is identified for binding to the target protein), either or both of the N-terminal and C-terminal positions undergo an extension step, whereby the length of the mature core hit peptide from step 204 is further extended to increase its specificity and affinity for the target peptide.
[0081] An example of a C-terminal extension according to the present disclosure is shown in Figure 3. Peptide extension or maturation array 300 includes a first population of peptides 302a and a second population of peptides 302b. Each of peptides 302a and 302b includes a mature core hit peptide 304 identified through the maturation process of step 204 of process 200 (Figure 2). A specific peptide probe (e.g., 5-mer probe peptide 110 of Figure 1) selected from the population of probe peptides from step 202 of peptide binder discovery is added to (or synthesized onto) the C-terminus of the mature core hit peptide 304 of the first population of peptides 302a. In this way, the most N-terminal amino acid of each peptide sequence is positioned immediately adjacent to the most C-terminal amino acid of the mature core hit peptide 304.
[0082] Similarly, according to various embodiments of N-terminal extension of the present disclosure, with reference to Figure 3, once the sequences of mature core hit peptides 304 have been identified by the maturation process (step 204, Figure 2), a particular one of each of the 5-mer probe peptides 110 (5-mer probe peptides 110, Figure 1) of the population 102 from step 202 of the peptide binder search is added to the N-terminus of the mature core hit peptides 304 in the second population of peptides 302b. In this way, the most C-terminal amino acid of each peptide sequence (5-mer probe peptide 110, Figure 1) is added immediately adjacent to the most N-terminal amino acid of the mature core hit peptides 304.
[0083] According to some embodiments of the present disclosure (FIG. 3), one or both of the mature core hit peptides 304 used in the C-terminal extension and the N-terminal extension can also include one or both of an N-terminal wobble sequence (N-terminus 306) and a C-terminal wobble sequence (C-terminus 308). Similar to N-terminus 106 and C-terminus 108 in FIG. 1, N-terminus 306 and C-terminus 308 can contain as few as one amino acid or as many as 15-20 amino acids (or more), and N-terminus 306 can be the same length as, longer than, or shorter than C-terminus 308. Furthermore, N-terminus 306 and C-terminus 308 can be added by using different mixtures of amino acids during the maturation process. Any single amino acid or any "wobble mix" consisting of two or more amino acids can be used. In yet other embodiments, a "flexible wobble mix" consisting of a mixture of Gly and Ser in a 3:1 (Gly:Ser) ratio is used. In other embodiments, a "random wobble mix" consisting of a random mixture of all 20 amino acids is used. In some embodiments, unnatural amino acids (e.g., 6-aminohexanoic acid) can also be used. Some embodiments may include non-amino acid moieties (e.g., polyethylene glycol).
[0084] 3 shows a peptide maturation array 300 having a population of peptides with C-terminal extensions 302a and a population of peptides with N-terminal extensions 302b. In the illustrated embodiment, the peptide maturation array 300 includes an array substrate 310 including a solid support 312 having a reactive surface 314 (e.g., a reactive amine layer) to which a first population of peptides 302a and a second population of peptides 302b are immobilized. Each of the first population of peptides 302a and the second population of peptides 302b can include the entire complement of 5-mer probe peptides 110 from the peptide array 100 (e.g., used in peptide binder searching step 204). As further shown, each peptide in both the first population of peptides 302a and the second population of peptides 302b can include the same mature core hit peptide 304, each with a different 5-mer probe peptide 110 (from the population of 5-mer probe peptides 110 (FIG. 1) from peptide binder searching step 102). Also shown in FIG. 3, each peptide in the first population of peptides 302a and the second population of peptides 302b includes a wobble amino acid sequence at the N-terminus 306 and C-terminus 308.
[0085] In some embodiments, mature array 300 (comprising peptides 302a and 302b) is exposed to an enriched and purified protein of interest or another similar receptor (similar to the peptide binder search of step 202 of process 200), whereby the protein may bind to any peptide in either the first population of peptides 302a or the second population of peptides 302b, independent of other peptides comprising the first population of peptides 302a and the second population of peptides 302b. After exposure to the protein of interest, binding of the protein of interest to the peptides in the first population of peptides 302a and the second population of peptides 302b is assayed, for example, by exposing the individual peptides and protein complexes in the first population of peptides 302a and the second population of peptides 302b to an antibody (specific for the protein) having a reportable label (e.g., peroxidase) attached. In another embodiment, the protein of interest may be directly labeled with a reporter molecule. Because the sequence of each of the 5-mer probe peptides 110 for each location on the array is known, it is possible to chart, quantify, compare, contrast, or a combination thereof, the sequence (and binding strength) of protein binding to a particular probe, including the mature core hit peptide 304 with each one of the 5-mer probe peptides 110.
[0086] An exemplary method for comparing proteins (of interest) that bind to a combination of mature core hit peptides 304 and 5-mer probe peptides 110 (including either a first population of peptides 302a or a second population of peptides 302b) is to consider binding strength in a principled analytical distribution-based clustering method, such as that described by White et al. (Standardizing and Simplifying Analysis of Peptide Library Data, J Chem Inf Model, 2013, 53(2), pp. 493-499). As exemplified herein, the clustering of proteins that bind to each probe (from the first population of peptides 302a and the second population of peptides 302b) shown in the principled analytical distribution-based clustering method reveals 5-mer probe peptides 110 with overlapping peptide sequences. From the overlapping peptide sequences (for each cluster), sequences of mature core hit peptides 304 can be identified, or at least hypothesized, and constructed for further evaluation, as described in more detail below. In some embodiments of the present application, the extended mature core hit peptide 304 undergoes a maturation process (as described and exemplified herein and shown in step 204 of Figure 2).
[0087] Additional rounds of optimization of extended peptide binders are also possible. For example, a third round of binder optimization may involve extending sequences identified in extension array experiments with Gly amino acids. Further optimization may involve creating double substitution or deletion libraries containing all possible single and double substitution or deletion variants of the reference sequence (i.e., peptide binders optimized and selected in any of the previous steps).
[0088] VII. Specificity analysis of extended mature core hit peptide binders Following identification of the extended mature core hit peptides, specificity analysis can be performed by any method available in the art for measuring peptide affinity and specificity. One example of specificity analysis includes the "BIACORE™" system analysis, which is used to characterize molecules in terms of target-directed molecular interaction, kinetic rates ("on," binding, and "off," dissociation), and affinity (binding strength). BIACORE™ is a trademark of General Electric Company and is available from the company's website.
[0089] FIG. 4 is a simplified schematic diagram of a method 400 for identifying novel peptide binders (e.g., process 200 of FIG. 2). As shown, an array 402 for peptide binder discovery is prepared by synthesizing a population of peptides (e.g., via maskless array synthesis) on an array substrate 404. As shown, each peptide 406 (or peptide feature) in the array 402 includes five cycles of wobble synthesis at the N-terminus (N-terminus 408) and five cycles of wobble synthesis at the C-terminus (C-terminus 410), such that each of the N-terminus (N-terminus 408) and C-terminus (C-terminus 410) contains five amino acids. It should be understood that the wobble synthesis at the N-terminus 408 and C-terminus 410 can include any composition, as described above. For example, the wobble synthesis can include only the amino acids Gly and Ser in a ratio of 3:1 (Gly:Ser), or a random mixture of all 20 amino acids. Each peptide 406 is also shown as including a 5-mer peptide binder or probe peptide 412, which, as described above, can include up to 2.9 million different peptide sequences, such that the entire human proteome is represented. Furthermore, it should be noted that different probe peptides 412 can be synthesized according to specific "rules." Non-limiting exemplary rules include excluding one or more amino acids (e.g., Cys, Met, or combinations thereof), excluding repeats of the same amino acids in consecutive order, excluding motifs already known to bind to target proteins (e.g., the His-Pro-Gln amino acid motif for streptavidin), and combinations thereof. As described above, protein targets of interest (e.g., in purified and enriched form) are exposed to the 5-mer probe peptides 412, and binding is scored (e.g., by principled clustering analysis), whereby "core hit peptide" sequences are identified based on overlapping binding motifs.
[0090] In some embodiments, upon identification of core hit peptide sequences, an exhaustive maturation process can be performed, as shown for maturation or maturation array 414. Maturation array 414 comprises a population of peptides 416 immobilized on an array substrate 418. In some embodiments, core hit peptides (illustrated as 5-mer core hit peptides 420) are synthesized on array substrate 418 with both an N-terminal wobble sequence (N-terminus 422) and a C-terminal wobble sequence (C-terminus 424). In the example shown in Figure 4, each of peptides 416 comprises three cycles of N-terminal and C-terminal wobble synthesis of only the amino acid Gly, although the wobble amino acid can be varied as described above. In some embodiments of exhaustive maturation, core hit peptides 416 are synthesized on an array substrate 418, where every amino acid position of the core hit peptide 416 is substituted with each of the other 19 amino acids, or double amino acid substitutions (as described above) are synthesized on the array substrate 418, or an amino acid deletion scan is synthesized on the array substrate 418, or an amino acid insertion scan is synthesized on the array substrate 418. In some cases, all of the above maturation processes are performed (and optionally repeated as described above for new peptides generated as a result of the amino acid deletion and insertion scans). Upon synthesis of a maturation array 414 containing various peptides (including substitutions, deletions, and insertions as described herein), the target protein is exposed to the modified core hit peptides 420 on the maturation array 414 and the strength of binding is assayed, thereby identifying a "mature core hit peptide" sequence.
[0091] In a further embodiment, after identification of the "mature core hit peptide" sequences, one or both of N-terminal and C-terminal extensions can be performed, as shown for extension array 426. Extension array 426 includes a first population of peptides 428a and a second population of peptides 428b, each immobilized on array substrate 430. As shown for selected peptides 432 of second population of peptides 428b, each of first population of peptides 428a and second population of peptides 428b includes a mature core hit peptide 434 (MC hit) attached at either the N-terminus (in the case of second population of peptides 428b) or C-terminus (in the case of first population of peptides 428a) to extension sequence 436. N-terminal and C-terminal extensions include synthesis of mature core hit peptides 434 adjacent to a population of probe peptides 412 (5-mers in this example). Probe peptides 416 are synthesized at either the N-terminus or C-terminus of mature core hit peptides 434. As shown for the first population of peptides 428a, the C-terminal extension involves five rounds of wobble synthesis to provide a C-terminal wobble sequence (C-terminus 438) and extension sequence 436 synthesized at the C-terminus of mature core hit peptide 434, followed by five more cycles of wobble synthesis to provide an N-terminal wobble sequence (N-terminus 440). Similarly, as shown for the second population of peptides 428b, the N-terminal extension involves five rounds of wobble synthesis (as described above) to generate C-terminus 438 synthesized at the C-terminus of mature core hit peptide 434, followed by another five cycles of wobble synthesis to provide extension sequence 436 and N-terminus 440. Upon synthesis of an extension array 426 comprising various C-terminal and N-terminal extension peptides (i.e., a first population of peptides 428a and a second population of peptides 428b), the target protein is exposed to the extension array 426 and binding is scored (e.g., by principle-based clustering analysis), thereby identifying the sequences of mature core hit peptides 434 extended at the C-terminus or N-terminus.As represented by the arrow indicated at 442, according to some embodiments, after an extended mature core hit peptide (e.g., peptide 432) is identified, the maturation process of the extended mature core hit peptide may be repeated, and then the extension process may be repeated on the resulting altered peptide sequence.
[0092] VIII. Identification of binder peptides for specific targets According to embodiments of the present disclosure, peptide microarrays are incubated with samples containing target proteins to obtain binders specific to various receptors. Examples of receptors include streptavidin, Taq polymerase, human proteins such as prostate-specific antigen, thrombin, tumor necrosis factor alpha, urokinase-type plasminogen activator, etc. Methods and exemplary peptide binders for the aforementioned receptors are described by Albert et al. (U.S. Patent Application No. 2015 / 0185216 and U.S. Provisional Patent No. 62 / 150,202).
[0093] The identified peptide binders can be used for various binder-specific purposes, although some uses are common to all binders. For example, for each of the targets described herein, the peptide binders of the present technology can be used as quality control peptides for inclusion in the synthesis of a wider collection of peptides (e.g., for use on peptide arrays to search for new peptide binder sequences).
[0094] 6A-6C, a peptide array 600 includes a population of peptide features 602 immobilized on an array substrate 604. Each of the peptide features 602 includes multiple co-localized peptides that share the same amino acid sequence. Depending on the synthesis method used, the peptide features may have different footprints or feature densities. In one example, a peptide feature has a square footprint of approximately 10 μm×10 μm, with a maximum density of approximately 10 75. However, as will be appreciated by those of ordinary skill in the art, other footprints and feature densities are possible. In this example, peptide feature 606 comprises multiple peptides, each having the same amino acid sequence. Peptide array 600 further comprises peptide feature 608, which has multiple peptide sequences that differ from the sequence comprising peptide feature 606. In particular, peptide array 600 can include a larger number of peptide features than the number of features shown in the embodiment depicted in FIG. 5.
[0095] In one aspect, the peptide features on peptide array 600 can collectively define at least one naturally occurring amino acid sequence. For example, the peptides can be tiled at one-amino acid resolution along the entire length of a partial or full-length protein sequence of interest (see FIGS. 6 and 7). In another example, the peptides can be tiled at four-amino acid resolution along the entire length of a partial or full-length protein sequence of interest (see FIG. 6). In some embodiments, the peptides can have amino acid sequences that collectively represent the entire human proteome or another proteome of interest. In this example, peptide array 600 includes peptide feature 610, peptide feature 612, and peptide feature 614, where each of peptide feature 610, peptide feature 612, and peptide feature 614 includes a plurality of peptides, each having a different amino acid sequence compared to each of the other peptide features.
[0096] Once the peptide array 600 is synthesized as shown in FIG. 5A, multiple receptor molecules known to interact with selected peptide binder sequences can be contacted with the peptide array 600 to probe the population of peptide features 602 in the presence of the receptor molecules (FIG. 5B). Multiple receptor molecules 616 are shown interacting with the peptide feature 606. Interactions between the receptor molecules 616 and the peptide feature 606 can include binding, catalysis (or participation) of a reaction involving the peptide in the peptide feature 606, digestion of the peptide in the feature 606, or the like, and combinations thereof. In the present example shown in FIGS. 5A-5C, the receptor 616 was used to identify the peptide binder sequence represented by the peptide in feature 606. Therefore, a strong degree of interaction between the peptide in the peptide feature 606 and the receptor molecule 616 would be expected to be represented by multiple receptor molecules 616 associated with the feature 606. In one aspect, interaction of receptor molecules 616 with a population of peptide features 602 on peptide array 600 can be detected, for example, by labeling receptor molecules 616 with a detectable tag 618 (FIG. 5C). As shown in the illustrated embodiment, detectable tag 618 is a labeled antibody specific for targeting receptor molecules 616. However, other detection schemes are within the scope of this disclosure.
[0097] While multiple receptor molecules 616 are associated with feature 606 in FIG. 5B , relatively few or no receptor molecules 616 are associated with any one of peptide feature 608, peptide feature 610, peptide feature 612, and peptide feature 614. In one embodiment, the sequence of the peptide in peptide feature 608 resulted in little or no interaction of receptor molecule 616 with peptide feature 608. In another embodiment, the sequence of the peptide in peptide feature 610, peptide feature 612, and peptide feature 614 resulted in little or no interaction of receptor molecule 616 with the aforementioned peptide features. As a result, the preference of receptor molecules for the peptide sequence represented in feature 606 can be inferred. Similarly, the degree of interaction or relative changes in the degree of interaction of receptor molecule 616 with any of the peptide features on peptide array 600 can be investigated.
[0098] Examples are provided herein to demonstrate the advantages of the present technology and to further assist those skilled in the art in preparing or using the compounds or salts, pharmaceutical compositions, derivatives, solvates, metabolites, prodrugs, racemic mixtures, or tautomers thereof of the present technology. Examples are also provided herein to more fully illustrate preferred embodiments of the present technology. The examples should not be construed in any way as limiting the scope of the present technology, as defined by the appended claims. The examples may include or incorporate any of the variations, aspects, or aspect(s) of the present technology described above. The variations, aspects, or aspect(s) above may also each include or incorporate variations of any or all of the other variations, aspects, or aspect(s) of the present technology. [Example]
[0099] Example 1: Naive search for novel DsbA-binding peptides Novel DsbA-binding peptides were discovered using the enhanced and improved peptide library as previously described. Briefly, peptide arrays were synthesized by light-directed array synthesis on a Roche NimbleGen Maskless Array Synthesizer (MAS) using amino-functionalized substrates, as previously reported (Forsstrom, et al., "Proteome-wide Epitope Mapping of Antibodies Using Ultra-dense Peptide Arrays," Molecular & Cellular Proteomics 13:1585-1597 (2014) and Lyamichev, et al., "Stepwise Evolution Improves Identification of Diverse Peptides Binding to a Protein Target," Nature Scientific Reports 7:12116 (2017), both of which are incorporated herein by reference in their entireties). L-amino acids were synthesized by Orgentis Chemicals GmbH. Custom amino acids were purchased from Lifetein. Cy5™-streptavidin was purchased from GE Healthcare, Blocker™ BSA (10%) in PBS from ThermoFisher, and SecureSeal™ hybridization chambers from GraceBio-Labs. Final side-chain deprotection was performed by incubating the microarray in 25 mM TIPS in 60 mM EDT and 95% TFA (v / v) for 30 minutes at room temperature. The microarray was then washed twice for 30 seconds with methanol, once for 1 minute with 1xTBST, and twice for 30 seconds with TBS, and then spun dry in a microcentrifuge with the array holder. To determine which of these peptides were able to bind to DsbA and to what extent, DsbA was incubated on the array.Here, 2.6 μL of biotin-labeled DsbA (1.5 mg / mL, Abcam) was incubated overnight at 4°C in a hybridization chamber on the array in 49 μL of binding buffer containing 100 mM HEPES (pH 7.3), 1% BSA, 250 mM NaCl, 20 mM L-gluathione-reduced, and 0.2 mM L-gluathione-oxidized. After incubation, the array was washed with water for 15 seconds and placed directly in a streptavidin-Cy5 detection bath. Positive binding to DsbA was determined by streptavidin-Cy5 detection. Streptavidin-Cy5 binding to all arrays was performed in a 30 mL pap jar at room temperature for 1 hour using 20 μL of streptavidin-Cy5 (1 μg / mL) in 30 mL of binding buffer containing 10 mM Tris-HCl (pH 7.4), 1% casein, and 0.05% Tween-20. After incubation, the arrays were washed for 30 seconds with 20 mM Tris-HCl (pH 7.8), 0.2 M NaCl, and 1% SDS, followed by 30 seconds with water. The arrays were then dried by spinning in a microcentrifuge with the array holder. Data were analyzed by measuring the Cy5 fluorescence intensity of the arrays and extracting the data, as previously described in Lyamichev, et al. (2017).
[0100] To generate an improved unique peptide array library, two peptide libraries of approximately 3 million unique 15-mer linear peptides were synthesized. One library contained peptides that could be found in the human proteome, and the other library contained 15-mer peptides composed of L-amino acids. To narrow the number of possible 15-mer peptides in both libraries to approximately 3 million peptides, calculations were performed to select the maximum diversity of 6-mer peptide sequences. Each 15-mer sequence was replicated twice. Table 3 shows the peptide sequences of the 20 best DsbA-binding peptides identified from the two unique 15-mer peptide libraries, human and nonhuman, that bind and specifically interact with biotin-labeled DsbA.
[0101] Table 3. The 20 best DsbA-binding peptides discovered from two unique 15-mer peptide libraries. TIFF2025138750000004.tif137128
[0102] The peptide sequences shown in Table 3 represent peptides that bind to DsbA with high affinity and specificity. As an example, Figure 8 shows the Cy5 fluorescence intensity of an array of DsbA-binding peptides with the sequence DFWGDTCKVTQFDQ (SEQ ID NO: 70). Data was analyzed and extracted for all other DsbA-binding peptides from Table 3 (data not shown), but only DFWGDTCKVTQFDQ (SEQ ID NO: 70) is shown here. Figure 8 shows a single substitution plot of the DFWGDTCKVTQFDQ peptide, DsbA_1_WT (SEQ ID NO: 70). The position of each peptide is represented by 21 bars (one bar for each of the 20 amino acids and one bar for the deletion). The height of each bar indicates the median signal intensity.
[0103] Example 2: Validation of discovered and stepwise optimized novel DsbA-binding peptides by kinetic characterization studies The novel DsbA-binding peptides discovered and stepwise optimized were subjected to kinetic characterization. Kinetic analysis of the interaction between the discovered peptides and DsbA was performed using a Biacore XI00. Biotin-labeled DsbA (1 μM) was immobilized on a streptavidin-coated chip (GE Healthcare) in HBS-P+ at a target level of 1,000 response units. A typical immobilization level was 1,400. Fourteen dilutions of the peptides, ranging from 5,000 nM to 0.6 nM in HBS-P+, were run in triplicate over the DsbA-coated chip with a 120-second contact time and a 600-second separation time. Table 4 shows the kinetic parameters for binding of the discovered peptides to immobilized biotin-labeled DsbA on the Biacore XI00. All peptides were amidated at the C-terminus.
[0104] Table 4. Kinetic parameters for binding of probed peptides to immobilized biotin-labeled DsbA on BiocoreX100. TIFF2025138750000005.tif207169
[0105] Figure 9 shows the binding of DsbA_l_WT, DFWHGDTCKVTQFDQ-NH2 (SEQ ID NO: 70) to biotin-labeled DsbA immobilized on a (Biacore) SA chip. The peptide concentrations range from 1.2 nM to 1250 nM.
[0106] Figure 10 shows a comparison of fluorescent signal intensity and binding affinity. Fluorescent signal intensity was measured by peptide microarray, and KD values were determined by Biacore or ITC (see Duprez, et al. (2015)). The values in the graph were found in Table 10 above. Peptides tested via Biacore were amidated at the C-terminus.
[0107] In summary, the data show that the discovered and stepwise optimized peptides bind to DsbA with high affinity and specificity.
[0108] equivalent The schematic flowcharts depicted in the figures are generally described as logical flowchart diagrams. Thus, the depicted order and labeled steps are indicative of one embodiment of the presented method. Other steps and methods may be conceived that are equivalent in function, logic, or effect to one or more steps, or portions thereof, of the depicted method. Furthermore, it is understood that the format and symbols used in the figures are provided to describe the logical steps of the method and do not limit the scope of the method. While various arrow and line types may be used, it is understood that they do not limit the scope of the corresponding method. In fact, some arrows or other connectors may be used to indicate only the logical flow of the method. For example, arrows may indicate waiting or monitoring periods of unspecified duration between listed steps of the depicted method. Furthermore, the order in which a particular method occurs may or may not strictly follow the order of the corresponding steps shown.
[0109] The present invention is presented in several various embodiments in the following description with reference to the figures, wherein like numbers represent the same or similar elements. Reference throughout this specification to "one embodiment," "an embodiment," or similar language means that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of the invention. Thus, appearances of the phrases "in one embodiment," "in an embodiment," and similar language throughout this specification can, but do not necessarily, all refer to the same embodiment.
[0110] The described features, structures, or characteristics of the invention may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are set forth to provide a thorough understanding of embodiments of the present system. However, those skilled in the relevant art will recognize that both the system and method may be practiced without one or more specific details or with other methods, components, materials, etc. In other instances, well-known structures, materials, or operations have not been shown or described in detail to avoid obscuring aspects of the invention. Therefore, the foregoing description is intended to be illustrative and not to limit the scope of the inventive concepts.
[0111] The present technology is also not limited in terms of the specific embodiments described herein, which are intended as single illustrations of individual embodiments of the technology. As will be apparent to those skilled in the art, many modifications and variations of the present technology can be made without departing from its spirit and scope. Functionally equivalent methods within the scope of the present technology, in addition to those recited herein, will be apparent to those skilled in the art from the foregoing description. Such modifications and variations are intended to fall within the scope of the appended claims. It is understood that the present technology is not limited to particular methods, reagents, compounds, compositions, labeled compounds, or biological systems, which may, of course, vary. It is also understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. Therefore, it is intended that the specification be considered exemplary only, with the breadth, scope, and spirit of the present technology indicated solely by the appended claims, the definitions therein, and any equivalents thereof.
[0112] The embodiments illustratively described herein may suitably be practiced in the absence of any element(s), limitation(ies) not specifically disclosed herein. Thus, for example, terms such as "comprising," "including," and "containing" shall be read expansively and without limitation. Furthermore, the terms and expressions used herein are used as terms of description, not of limitation, and the use of such terms and expressions is not intended to exclude any equivalents of the shown and described features or portions thereof, recognizing that various modifications are possible within the scope of the claimed technology. Furthermore, the phrase "consisting essentially of" shall be understood to include the specifically recited elements and additional elements that do not materially affect the basic and novel characteristics of the claimed technology. The phrase "consisting of" excludes any elements not specified.
[0113] Furthermore, where features or aspects of the disclosure are described in terms of a Markush group, one of ordinary skill in the art will recognize that the disclosure also describes in terms of any individual member or subgroup of members of the Markush group. Each of the narrower species and subgeneric groupings falling within the generic disclosure also form part of the invention. This includes the generic description of the invention with a condition or negative limitation removing any subject matter from the genus, regardless of whether the deleted material is specifically recited herein.
[0114] As will be understood by those skilled in the art, for any and all purposes, particularly in terms of providing a written description, all ranges disclosed herein encompass any and all possible subranges and combinations thereof. Any recited range can be readily recognized as fully descriptive and allowing for the same range to be broken down into at least equal halves, 1 / 3, 1 / 4, 1 / 5, 1 / 10, etc. As a non-limiting example, each range discussed herein can be readily broken down into a lower 1 / 3, a middle 1 / 3, an upper 1 / 3, etc. Also, as will be understood by those skilled in the art, language such as "up to," "at least," "greater than," "less than," and the like, all refer to ranges that are inclusive of the recited numbers and that can be subsequently broken down into the subranges discussed above. Finally, as will be understood by those skilled in the art, a range includes each individual member.
[0115] All publications, patent applications, issued patents, and other documents (e.g., journals, articles, and / or textbooks) referenced herein are incorporated by reference herein to the same extent as if each individual publication, patent application, issued patent, or other document were specifically and individually indicated to be incorporated by reference in its entirety. Definitions contained in the text incorporated by reference are excluded to the extent they conflict with definitions in this disclosure.
[0116] Other embodiments are set forth in the following claims, along with the full scope of equivalents to which such claims are entitled.
[0117] Array information SEQUENCE LISTING <110> ROCHE SEQUENCING SOLUTIONS, INC. <120> PEPTIDE LIBRARIES HAVING ENHANCED SUBSEQUENCE DIVERSITY AND METHODS FOR USE THEREOF <150> US 62 / 867,765 <151> 2019-06-27 <150> US 62 / 867,666 <151> 2019-06-27 <160> 80 <170> PatentIn version 3.5 <210> 1 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 1 Trp Ser His Pro Gln Phe Glu Lys 1 5 <210> 2 <211> 6 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic 6xHis tag" <400> 2 His His His His His His 1 5 <210> 3 <211> 5 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 3 Met Met Met Met Met 1 5 <210> 4 <211> 5 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 4 Ala Met Met Met Met 1 5 <210> 5 <211> 5 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 5 Pro Met Met Met Met 1 5 <210> 6 <211> 5 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 6 Val Met Met Met Met 1 5 <210> 7 <211> 5 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 7 Gln Met Met Met Met 1 5 <210> 8 <211> 5 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 8 Met Ala Met Met Met 1 5 <210> 9 <211> 5 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 9 Met Gln Met Met Met 1 5 <210> 10 <211> 5 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 10 Met Pro Met Met Met 1 5 <210> 11 <211> 5 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 11 Met Asn Met Met Met 1 5 <210> 12 <211> 5 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 12 Pro Ala Met Met Met 1 5 <210> 13 <211> 5 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 13 Pro Phe Met Met Met 1 5 <210> 14 <211> 5 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 14 Pro Val Met Met Met 1 5 <210> 15 <211> 5 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 15 Pro Glu Met Met Met 1 5 <210> 16 <211> 5 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 16 Val Ala Met Met Met 1 5 <210> 17 <211> 5 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 17 Val Phe Met Met Met 1 5 <210> 18 <211> 5 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 18 Val Val Met Met Met 1 5 <210> 19 <211> 5 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 19 Val Glu Met Met Met 1 5 <210> 20 <211> 5 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 20 Ala Ala Met Met Met 1 5 <210> 21 <211> 5 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 21 Ala Phe Met Met Met 1 5 <210> 22 <211> 5 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 22 Ala Val Met Met Met 1 5 <210> 23 <211> 5 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 23 Ala Glu Met Met Met 1 5 <210> 24 <211> 4 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 24 Met Met Met Met 1 <210> 25 <211> 6 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <220> <221> MOD_RES <222> (1)..(1) <223> Any of the 20 natural amino acids <400> 25 Xaa Met Met Met Met Met 1 5 <210> 26 <211> 6 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <220> <221> MOD_RES <222> (2)..(2) <223> Any of the 20 natural amino acids <400> 26 Met Xaa Met Met Met Met 1 5 <210> 27 <211> 6 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <220> <221> MOD_RES <222> (3)..(3) <223> Any of the 20 natural amino acids <400> 27 Met Met Xaa Met Met Met 1 5 <210> 28 <211> 6 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <220> <221> MOD_RES <222> (4)..(4) <223> Any of the 20 natural amino acids <400> 28 Met Met Met Xaa Met Met 1 5 <210> 29 <211> 6 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <220> <221> MOD_RES <222> (5)..(5) <223> Any of the 20 natural amino acids <400> 29 Met Met Met Met Xaa Met 1 5 <210> 30 <211> 6 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <220> <221> MOD_RES <222> (6)..(6) <223> Any of the 20 natural amino acids <400> 30 Met Met Met Met Met Xaa 1 5 <210> 31 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 31 Asn Asp Tyr Gln Tyr Lys Gly Gln Gln Cys Trp Asp Pro Arg Tyr 1 5 10 15 <210> 32 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 32 His Asn Trp Ala Ala Gln Cys Thr Asn Pro Asp Gln Ala Lys Leu 1 5 10 15 <210> 33 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 33 Trp Ser Leu Gly Gln Cys Val Ala Asp Gln Pro Asn Gln Thr Trp 1 5 10 15 <210> 34 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 34 Ser Ile Ile Gly Gln Cys Tyr Ser Lys Asp Pro Ser Gln Ala Thr 1 5 10 15 <210> 35 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 35 Gln Asn Phe Tyr Gly Glu Thr Arg Phe Thr Pro Ile Ser Ser Trp 1 5 10 15 <210> 36 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 36 Leu Ser Tyr Asn Leu Phe Tyr Gly Lys Cys Trp Ala Pro Trp Phe 1 5 10 15 <210> 37 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 37 Cys Gln Asn Pro Ala Tyr Gln Lys Lys Tyr Asn Trp Glu Pro Val 1 5 10 15 <210> 38 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 38 Ile Trp Leu Gly Glu Cys Cys Lys Asn Asn Tyr Ser Arg Ala Arg 1 5 10 15 <210> 39 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 39 Val Ile Val Gly Gln Cys Leu Lys Lys Tyr Asn Asp Thr Arg Val 1 5 10 15 <210> 40 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 40 Lys Pro Ile Tyr Leu Gly Gln Cys Cys Asn Thr Ser His Ala Arg 1 5 10 15 <210> 41 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 41 Arg Ser Asn Asn Asp Ile Cys Phe Trp Thr Gly Gln Gln Cys Cys 1 5 10 15 <210> 42 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 42 Tyr Val Tyr Gly Thr Cys Trp Lys Asp Tyr Thr Asp Thr Asp Lys 1 5 10 15 <210> 43 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 43 Gln Ile Gly Gln Cys Leu Lys Asn Val Gly Tyr Arg Glu Val Pro 1 5 10 15 <210> 44 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 44 Phe Gln Val Gly Gln Cys Ile Asn Val Gln Tyr Asp Thr Asn Pro 1 5 10 15 <210> 45 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 45 Lys Pro Asn Val Asp Phe Trp His Gly Gly Gln Cys Asn Phe Thr 1 5 10 15 <210> 46 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 46 Phe Gln Ile Gly Gln Cys Tyr Gly Thr Gln His Trp Ala Ser Phe 1 5 10 15 <210> 47 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 47 Trp Gln Tyr Glu Asn Trp Ala Ile Thr Lys Trp Gln Cys Val Thr 1 5 10 15 <210> 48 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 48 Ser Asp Gly Trp Ala Lys Gln Cys Trp Ala Pro Thr His Asn Thr 1 5 10 15 <210> 49 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 49 Gly Gln Asp Trp His Gly Gln Gln Cys Cys Ala Glu Tyr Ala Glu 1 5 10 15 <210> 50 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 50 Ile Pro His Gln Tyr Ala Pro Gly Gln Cys Cys Lys Ile Ala Tyr 1 5 10 15 <210> 51 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 51 Thr Leu Gly Gln Cys Cys Ala Phe Pro Val Leu Asp Trp Lys Asn 1 5 10 15 <210> 52 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 52 Ala Leu Gly Glu Cys Val Lys Pro Leu Ala Phe Glu Lys Gln Ser 1 5 10 15 <210> 53 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 53 Val Ile Gly Phe Cys Ala Lys Pro Gln Asp Asn Ser Ser Ala Pro 1 5 10 15 <210> 54 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 54 Pro Ser Pro Trp Ala Thr Cys Asp Phe 1 5 <210> 55 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 55 Pro Ser Pro Phe Ala Thr Cys Asp Phe 1 5 <210> 56 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 56 Asp Phe Trp His Gly Glu Thr Cys Lys Val Thr Gln Phe Asp Gln 1 5 10 15 <210> 57 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 57 Asp Phe Trp His Gly Glu Gln Cys Lys Val Thr Gln Phe Asp Gln 1 5 10 15 <210> 58 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 58 Asn Asp Tyr Gln Tyr Lys Gly Gln Gln Cys Leu Asp Pro Lys Tyr 1 5 10 15 <210> 59 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 59 His Asn Trp Ala Ala Gln Cys Leu Asn Pro Asp Gln Ala Lys Leu 1 5 10 15 <210> 60 <211> 14 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 60 His Asn Trp Ala Ala Gln Cys Leu Lys Asp Gln Ala Lys Leu 1 5 10 <210> 61 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 61 Ser Ile Ile Gly Gln Cys Tyr Lys Lys Asp Pro Ser Gln Ala Thr 1 5 10 15 <210> 62 <211> 12 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 62 Ser Ile Ile Gly Met Cys Tyr Lys Lys Asp Pro Ser 1 5 10 <210> 63 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 63 Val Ile Gly Gln Cys Leu Lys Asn 1 5 <210> 64 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 64 Trp Gln Tyr Glu Asn Trp Ala Ile Thr Lys Trp Gln Cys Val Lys 1 5 10 15 <210> 65 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 65 Val Ile Val Gly Gln Cys Leu Lys Lys 1 5 <210> 66 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 66 Val Ile Leu Gly Gln Cys Leu Lys Gln 1 5 <210> 67 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 67 Tyr Val Tyr Gly Met Cys Trp Lys Asp Tyr Thr Asp Thr Asp Lys 1 5 10 15 <210> 68 <211> 14 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 68 Tyr Val Tyr Gly Met Cys Trp Lys Asp Tyr Thr Thr Asp Lys 1 5 10 <210> 69 <211> 25 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 69 Glu Gly Val Lys Leu Thr Ala Leu Asn Asp Ser Ser Leu Asp Leu Ser 1 5 10 15 Met Asp Ser Asp Asn Ser Met Ser Val 20 25 <210> 70 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 70 Asp Phe Trp His Gly Asp Thr Cys Lys Val Thr Gln Phe Asp Gln 1 5 10 15 <210> 71 <211> 16 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 71 Glu Gly Val Lys Leu Thr Ala Leu Asn Asp Ser Ser Leu Asp Leu Ser 1 5 10 15 <210> 72 <211> 16 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 72 Gly Val Lys Leu Thr Ala Leu Asn Asp Ser Ser Leu Asp Leu Ser Met 1 5 10 15 <210> 73 <211> 16 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 73 Val Lys Leu Thr Ala Leu Asn Asp Ser Ser Leu Asp Leu Ser Met Asp 1 5 10 15 <210> 74 <211> 16 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 74 Lys Leu Thr Ala Leu Asn Asp Ser Ser Leu Asp Leu Ser Met Asp Ser 1 5 10 15 <210> 75 <211> 16 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 75 Leu Thr Ala Leu Asn Asp Ser Ser Leu Asp Leu Ser Met Asp Ser Asp 1 5 10 15 <210> 76 <211> 16 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 76 Thr Ala Leu Asn Asp Ser Ser Leu Asp Leu Ser Met Asp Ser Asp Asn 1 5 10 15 <210> 77 <211> 16 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 77 Ala Leu Asn Asp Ser Ser Leu Asp Leu Ser Met Asp Ser Asp Asn Ser 1 5 10 15 <210> 78 <211> 16 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 78 Leu Asn Asp Ser Ser Leu Asp Leu Ser Met Asp Ser Asp Asn Ser Met 1 5 10 15 <210> 79 <211> 16 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 79 Asn Asp Ser Ser Leu Asp Leu Ser Met Asp Ser Asp Asn Ser Met Ser 1 5 10 15 <210> 80 <211> 16 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic peptide" <400> 80 Asp Ser Ser Leu Asp Leu Ser Met Asp Ser Asp Asn Ser Met Ser Val 1 5 10 15
Claims
1. a plurality of peptide features, each of said peptide features comprising at least one peptide, said at least one peptide comprising a composite region having a defined sequence of amino acids of length N, said composite region representing k distinct elements, each of said distinct elements having a defined sequence of amino acids of length x; 1. An engineered peptide library comprising: x, N, and k are integers; x is less than N; k is at least 2; The total number of distinct elements represented by the engineered peptide library is K Eng and the number of peptide features contained in the engineered peptide library is F; K Eng is greater than F, The engineered peptide library.
2. 2. The engineered peptide library of claim 1, wherein k is defined by the equation: k=N-x+1.
3. K Eng 3. The engineered peptide library of claim 1 or 2, wherein k is at least 0.8*k*F.
4. K Eng 4. The engineered peptide library of claim 1, wherein k is at least 0.9*k*F.
5. K Eng 5. The engineered peptide library of any one of claims 1 to 4, wherein k is at least 0.95*k*F.
6. K Eng 6. The engineered peptide library of any one of claims 1 to 5, wherein k is at least 0.99*k*F.
7. K Eng 7. The engineered peptide library of any one of claims 1 to 6, wherein F is at least 10*F.
8. K Eng 8. The engineered peptide library of any one of claims 1 to 7, wherein F is at least 20*F.
9. K Eng is the total number K of random elements having sequences of amino acids of length x represented by a random peptide library having F peptide features. Rnd 9. The engineered peptide library of any one of claims 1-8, wherein each of the peptide features of the random peptide library comprises at least one random peptide, the at least one random peptide having a random sequence of amino acids of length N.
10. 10. The engineered peptide library of any one of claims 1 to 9, wherein each of the elements in a selected one of the composite regions overlaps with each adjacent element in the composite region by at least one amino acid.
11. 11. The engineered peptide library of any one of claims 1 to 10, wherein the plurality of peptides represents at least 90% of the target proteome.
12. 12. The engineered peptide library of any one of claims 1 to 11, wherein each of the peptides in the plurality of peptides is between 12 and 16 amino acids in length.
13. 13. The engineered peptide library of any one of claims 1 to 12, wherein N is at least 7 amino acids.
14. 14. The engineered peptide library of any one of claims 1 to 13, wherein N is at least 10 amino acids.
15. 15. The engineered peptide library of any one of claims 1 to 14, wherein N is at least 15 amino acids.
16. 16. The engineered peptide library of any one of claims 1 to 15, wherein x is at least 5.
17. 17. The engineered peptide library of any one of claims 1 to 16, wherein x is 6.
18. 1. A method for identifying peptide binders, comprising: contacting a first sample with the engineered peptide library of any one of claims 1 to 17; selecting at least one of the plurality of peptides from a first subset of peptides; A method for identifying said peptide binders comprising:
19. detecting a first signal output characteristic of said first interaction with said engineered peptide library; selecting the at least one peptide based on the first signal output; 20. The method of claim 18, further comprising:
20. The first signal output is a fluorescence intensity obtained through fluorophore excitation emission, and the fluorescence intensity is i) the abundance of a component in one of the first sample and the second sample associated with the first plurality of peptides; and ii) the binding affinity of the components of one of the first sample and the second sample to the first plurality of peptides.
20. The method of claim 19, wherein the method reflects at least one of:
21. a plurality of peptide features, each of said peptide features comprising at least one peptide, said at least one peptide comprising a composite region having a defined sequence of at least 15 amino acids, said composite region representing 10 distinct elements, each of said distinct elements having a defined sequence of 6 amino acids; 1. An engineered peptide library comprising: The total number of distinct elements represented by the engineered peptide library is K Eng and the number of peptide features contained in the engineered peptide library is F; K Eng is at least 9.5*F, The engineered peptide library.
22. 22. The engineered peptide library of any one of claims 1 to 17 or 21, wherein each of the peptides is attached to a solid support.
23. 23. The engineered peptide library of claim 22, wherein the solid support is one of a bead and a chip.
24. 1. An engineered peptide library with enhanced subsequence diversity, the engineered peptide library comprising: a plurality of peptide features, each of said peptide features comprising at least one peptide, said at least one peptide comprising a composite region having a defined sequence of amino acids of length N, said composite region representing k distinct elements, each of said distinct elements having a defined sequence of amino acids of length x; Including, x, N, and k are integers; x is less than N; k is at least 2; The total number of distinct elements represented by the engineered peptide library is K Eng and the number of peptide features contained in the engineered peptide library is F; KEng is larger than F; The ratio of KEng to F is a measure of subsequence diversity, each of the k distinct elements within each composite region has a unique arrangement compared to each of the other distinct elements of the same composite region; The defined sequences of each of the composite regions are selected for their maximum subsequence diversity relative to the average subsequence diversity of the total number of random elements, KRnd; the random element has a sequence of amino acids of length x represented by a random peptide library having F peptide features, each of the peptide features of the random peptide library comprising at least one random peptide, the at least one random peptide having a random sequence of amino acids of length N; The engineered peptide library having enhanced subsequence diversity.