Methods and compositions for protein and peptide sequencing

By utilizing aptamers that specifically bind to N-terminal amino acids and the SELEX method, combined with artificial intelligence optimization, we have achieved highly efficient protein and peptide sequencing, solving the problem of low efficiency in existing protein sequencing technologies and providing a high-throughput sequencing solution for the dynamic range of proteins.

CN114555810BActive Publication Date: 2025-11-07GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080061216.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-13
Filing Date
2020-09-11
Publication Date
2025-11-07
Estimated Expiration
2040-09-11

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently sequence proteins, especially low-expression proteins and non-target proteins, and lack high-throughput sequencing strategies that span the dynamic range of protein expression, making it difficult to infer disease state information from the genome.

Method used

Protein sequencing was performed using aptamers that specifically bind to N-terminal amino acids via the SELEX method. The experimental seed conjugates were optimized using artificial intelligence and deep learning. The protein sequences were recorded using DNA barcoding, and efficient sequencing was achieved using PCR amplification and sequencing technologies.

Benefits of technology

It enables high-throughput, low-cost sequencing of proteins and peptides, allowing for the identification and recording of protein amino acid sequences, inference of the correlation between protein levels and enzymatic activity, and support for disease state monitoring and treatment efficacy evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114555810B_ABST
    Figure CN114555810B_ABST
Patent Text Reader

Abstract

The present disclosure describes methods and compositions for protein and peptide sequencing.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates generally to methods and compositions for protein and peptide sequencing. BACKGROUND

[0002] Over the past decade, rapid improvements in DNA sequencing technology have generated vast amounts of molecular information. And while the ability to read genomes has revolutionized biological research, vast amounts of phenotypic and disease state information cannot be inferred from the genome. RNA sequencing provides a deeper understanding of the functional elements of the genome and their expression levels. However, there remain significant challenges around efforts to correlate proteins with mRNA expression levels (de Sousa, Abreu, Penalva, Marcotte, & Vogel, 2009) (Vogel & Marcotte, 2012), leading to a loss of cell state information in understanding precise protein quantification, modification, or even sequence. In assessing proteins in serum, RNA analysis cannot predict the presence of a protein, as the protein can be expelled from the cell and circulate throughout the blood system, leading to a loss of spatial connection between the RNA sequence and its translated target. Furthermore, protein sequencing can reveal many unknown proteins, i.e., proteins from other organisms (e.g., viruses, bacteria, etc.) that are present in the host’s bloodstream and affect the host organism.

[0003] RNA and DNA sequencing have limited knowledge of antibody sequences, as the diversity of the antibody repertoire is generated by somatic hypermutation events. To capture information that arises after DNA processing and secretion, such as post-translational protein modifications, translational fidelity, protein folding integrity, etc., scientists must be able to sequence proteins directly from a sample of interest (i.e., read their amino acid sequence) to infer correlations between protein levels and their enzymatic action. De novo protein sequencing can lead to the discovery of rare and novel proteins from any organism (e.g., various different tissues, pathogens, mutated cancer cells) or from any protein-containing sample (e.g., blood, skin, cerebrospinal fluid, fecal matter). Protein sequencing can also serve as a measure of therapeutic efficacy by allowing for extensive physiological monitoring during disease treatment. However, there currently exists no cost- and time-effective strategy for large-scale, high-throughput sequencing of proteins and proteomes that spans the entire dynamic range of protein expression. There also exists no reliable method for sequencing non-target, low-expression proteins. Thus, there remain obstacles to sequencing antibodies and low-expression proteins using current technology, and it is practically infeasible outside of the most specialized research efforts. SUMMARY

[0004] The present disclosure describes a series of methods and compositions that form a pipeline for developing and using a protein sequencing platform that utilizes aptamers that specifically bind to N-terminal amino acids.Figure 1 ). Amino acid specific aptamers can be generated using the novel methods described herein (RCHT-SELEX and NTAA-SELEX). Such amino acid specific aptamers can be used to recognize, identify and translate each amino acid of a protein or peptide into a DNA sequence (PROSEQ) or such amino acid specific aptamers can be used to recognize or identify each amino acid of a protein or peptide based on a visual signal (PROSEQ-VIS). Furthermore, many different target specific aptamers can be generated simultaneously and can be used to generate and screen a large number of binders (MULTIPLEX). Simultaneous and specific aptamer selection relies on the reliable identification of the target. The generation of targets with nucleic acid barcodes can be achieved in vivo by non-covalent bonds between peptides or proteins that utilize RNA binding proteins and their corresponding recognition sequences (TURDUCKEN). Finally, a successful SELEX experiment requires the inclusion of aptamers with a certain specific binding preference and affinity for the molecular target in the original pool of 10 14 -10 15 candidate sequences out of a pool of all possible DNA sequences. Artificial intelligence (AI), deep learning (DL) and machine learning (ML) can optimize the experimental seed binders, so that unlike conventional SELEX experiments, the optimal binders do not have to be present in the initial starting library, but can be generated from the characteristics of the binders found experimentally. The ability to construct computationally derived, customizable DNA libraries to perform SELEX screening with controlled input pools can significantly increase the exploration space by systematically analyzing aptamer candidates containing sequences with known binding properties (LEGO).

[0005] In one aspect, a method of obtaining an aptamer with affinity and specificity for a target is provided. Such a method generally comprises: (a) providing a plurality of aptamers; (b) optionally subjecting the plurality of aptamers to negative selection; (c) optionally incorporating a control oligonucleotide into the plurality of aptamers prior to PCR amplification; (d) optionally amplifying the plurality of aptamers; (e) incubating the plurality of aptamers with a plurality of potential targets under conditions that allow the plurality of aptamers to bind to the plurality of potential targets; (f) optionally, for a parallel experiment, incubating the plurality of amplified aptamers with a plurality of potential targets or null targets in different reactions under conditions that allow the plurality of amplified aptamers to bind to the plurality of potential targets; (g) removing unbound aptamers; (h) sequencing target-bound aptamers; and (i) repeating steps (a)-(h) a plurality of times, thereby obtaining an aptamer with affinity and specificity for the target.

[0006] In certain embodiments, the potential target is a polypeptide, an amino acid, a nucleic acid, a small molecule, a whole protein or protein complex, or a cell.

[0007] In certain embodiments, the method further comprises amplifying the plurality of aptamer candidates in the initial random library or ML designed library in a single incubation amplification step or a double incubation amplification step to generate an input pool for SELEX containing multiple copies of the aptamer candidates.

[0008] In certain embodiments, the same incubation is assayed for multiple targets, in parallel experiments, or a combination thereof.

[0009] In certain embodiments, the method optionally further comprises introducing a known amount of a known oligonucleotide in the sample prior to the step of amplifying the plurality of aptamers.

[0010] In certain embodiments, the method optionally further comprises introducing a known amount of a known oligonucleotide in the sample prior to the step of sequencing.

[0011] In certain embodiments, the sequencing data of the incorporated known oligonucleotide is observed to detect experimental errors.

[0012] In certain embodiments, the method further comprises amplifying a standardized amount of target-bound aptamer from each sample at each repetition of the steps.

[0013] In certain embodiments, the method further comprises amplifying the plurality of aptamers under conditions optimized for the particular primers used to obtain maximal amplification and minimal bias.

[0014] In certain embodiments, the method further comprises digesting the post-PCR dsDNA to ssDNA in order to preserve the desired strand.

[0015] In certain embodiments, the method further comprises amplifying the plurality of aptamers in the presence of primers that enrich for the desired ssDNA.

[0016] In certain embodiments, the method further comprises performing a unit test prior to each dsDNA digestion to determine the optimal digestion conditions for each sample.

[0017] In certain embodiments, the method further comprises changing the primer sequence associated with each member of the plurality of aptamers prior to repeating the step of incubating the plurality of aptamers with potential targets multiple times to identify strong binders that are not dependent on the primer region.

[0018] In certain embodiments, for experiments in which the desired aptamer is one that specifically binds to a smaller portion of a molecule rather than the entire molecule, the method further comprises alternating the target with a different local environment binding region between each repetition of steps (a)-(h).

[0019] In certain embodiments, the method further comprises performing the same PCR reaction on a small sample of the pool of aptamers prior to step (e) of Method 1 but without analysis for beads or targets to assess the effectiveness of SELEX using the selected selection components.

[0020] In certain embodiments, the method further comprises: (a) incubating the plurality of aptamers with a plurality of different targets in the same reaction under conditions that allow the plurality of aptamers to bind to the plurality of potential targets; (b) removing unbound aptamers; (c) amplifying target-bound aptamers; (d) sequencing target-bound aptamers; (e) repeating steps (a)-(d) a plurality of times; (f) incubating the plurality of aptamers with a plurality of single targets in each experiment for each different target; (g) repeating steps (b)-(d); thereby identifying aptamer binders that bind to a plurality of targets.

[0021] In certain embodiments, step (e) of claim 1 is repeated a plurality of times in separate reactions, each reaction containing a potential target.

[0022] The SELEX methods described herein, referred to herein as RCHT SELEX, are designed to have a flow that is ideal for integrating machine learning (ML) into the SELEX process (e.g. prioritizing computational demands). Additional SELEX methods are also described, referred to herein as N-terminal SELEX (or N-terminal amino acid (NTAA) SELEX), and are designed to have a flow that is ideal for small and / or difficult targets (e.g. prioritizing experimental demands).

[0023] While both SELEX methods can be modified as needed, differences between the two methods can include:

[0024] (a) Incubation, referred to as SELEX-RCHT, is typically performed at the beginning of the SELEX portion of the method, reducing the initial input pool to 10A12 molecules. On the other hand, incubation is typically not included in the method referred to as SELEX-NTAA, and in some cases no incubation effect is better. Thus, the SELEX-NTAA method typically starts with a pool of 10A14-10A15 random aptamers.

[0025] (b) Reactions using the SELEX-RCHT method typically use parallel samples that are run in parallel (e.g. 2 to 3 parallel reactions); parallel reactions are typically not necessary when using the SELEX-NTAA method so that as many experiments as possible can be run in parallel.

[0026] (c) Control reactions using the SELEX-RCHT method are typically run in parallel reactions, which take up 3 of the 12 possible incubation inputs; parallel samples for control reactions are not necessary when using the SELEX-NTAA method, although it is recommended to run one target in each experiment to determine overall experimental failure or contamination.

[0027] (d) The SELEX-NTAA method uses a target-switching step, which allows for the pursuit of small or difficult targets (or sub-regions of larger targets); the SELEX-RCHT method does not typically use this additional step.

[0028] (e) The SELEX-NTAA method incorporates an additional step of counter-selection, particularly when pursuing sub-regions of a target, to isolate the best experimental binder. The SELEX-RCHT method does not include a counter-selection step, however a counter-selection step can be used with the SELEX-RCHT method, provided that such counter-selection step is used carefully, either alone or in combination with other steps of the method, to avoid biasing the results.

[0029] In one aspect, a method of sequencing a protein or peptide is provided. Such a method generally comprises: (a) incubating the protein or peptide with a library of DNA aptamers exhibiting binding specificity to at least one N-terminal amino acid under conditions in which one or more of the aptamers specifically bind to at least one N-terminal amino acid of the protein or peptide, wherein each aptamer within the library comprises a peptide-binding ssDNA region and a unique barcode sequence indicative of the first sequencing round and the associated peptide-binding ssDNA region; (b) ligating the DNA aptamer bound to the N-terminus of the protein or peptide to a DNA barcode construct proximal thereto; (c) removing the peptide-binding sequence from the DNA aptamer, leaving only the barcode of the DNA aptamer and a short, consensus sequence for subsequent ligation covalently attached to the DNA barcode construct, so as to record the identity of the binder and thus the putative amino acid at the N-terminus of the peptide; (d) removing the N-terminal amino acid from the protein or peptide, to yield an N-terminal amino acid shortened protein or peptide; (e) incubating the N-terminal amino acid shortened protein or peptide with a library of aptamers exhibiting binding specificity to at least one amino acid under conditions in which one or more of the aptamers specifically bind to at least one N-terminal amino acid of the N-terminal amino acid shortened protein or peptide, wherein each aptamer within the library comprises a peptide-binding ssDNA region and a unique barcode sequence indicative of the second sequencing round and the associated peptide-binding ssDNA region; (f) ligating the DNA aptamer bound to the N-terminus of the protein or peptide to a DNA barcode construct proximal thereto; (g) removing the peptide-binding sequence from the DNA aptamer, leaving only the barcode of the DNA aptamer and a short, consensus sequence for subsequent ligation covalently attached to the DNA barcode construct, so as to record the identity of the binder and thus the putative amino acid at the N-terminus of the peptide; (h) removing the N-terminal amino acid from the N-terminal amino acid shortened protein or peptide; (i) repeating steps (a)-(d) a plurality of times to construct a chain of position barcodes corresponding to consecutive N-terminal amino acids in the protein or peptide; and (j) sequencing the chain of position barcodes, thereby obtaining the sequence of the protein or peptide.

[0030] In certain embodiments, the protein or peptide is from a synthetic sample, a biological sample, or a combination thereof. In certain embodiments, the biological sample is selected from blood, urine, saliva, a tissue biopsy sample, sputum, fecal matter, a single cell, an environmental sample, a bacterial swab, or any sample containing a peptide or protein.

[0031] In certain embodiments, the protein or peptide is a full-length protein, a peptide fragment, or a protein or peptide contained within a complex. In certain embodiments, the method further comprises fragmenting the protein or peptide prior to step (a). In certain embodiments, the fragmenting step comprises fragmenting the protein or peptide with trypsin, Lys-C, another fragmenting enzyme, an optional protein fragmentation or degradation method, or a combination thereof.

[0032] In certain embodiments, the C-terminal end of the protein or peptide is attached to a solid support. In certain embodiments, the C-terminal end of the protein or peptide is attached to an oligonucleotide tail. In certain embodiments, removing the aptamer comprises cleaving the aptamer at a restriction site using a restriction enzyme. In certain embodiments, the aptamer is attached to the barcode using a hydrostatic method, and removal of the peptide binding sequence is mediated by hydrogen bond disruption (rather than by DNA cleavage by a restriction enzyme).

[0033] In certain embodiments, the step of removing the N-terminal amino acid comprises Edman degradation of the protein or peptide, cleavage of the protein or peptide with one or more aminopeptidases, heat, pH, or a combination thereof. In certain embodiments, the sequencing step uses a next generation sequencing (NGS) platform. In certain embodiments, the number of sequencing reads associated with the amino acid sequence of a known protein is analyzed to determine the relative amount of protein in the sample.

[0034] In certain cases, methods of identifying new biomarkers are provided. Such methods generally comprise: (a) providing protein samples from a biological sample of interest and a control or comparison biological sample according to the methods described herein; (b) optionally removing very high concentrations of known proteins; (c) performing steps (a)-(j) of the methods described herein; (d) removing high concentrations of DNA barcode construct sequences associated with proteins that are typically highly expressed or contaminants, so as to increase the ratio of DNA barcode constructs associated with lowly expressed proteins compared to highly expressed proteins, thereby generating ratio-adjusted DNA barcodes; (e) PCR amplifying the sup-diffed DNA barcode constructs; and (f) comparing the number of sequencing reads associated with each lowly expressed protein from the control sample to the sample of interest, thereby identifying putative biomarkers that have significantly different relative expression levels between the control sample and the sample of interest.

[0035] In certain embodiments, methods of assessing a disease state, assessing a response to a treatment, predicting a response to a treatment, or a combination thereof using the protein sequencing methods described herein are provided, wherein one or more signs of the disease is an abnormal expression level of a known protein biomarker. Such methods generally comprise: (a) providing a protein sample from a patient sample according to the methods described herein; (b) optionally stripping very high concentrations of known proteins; (c) performing steps (a)-(j) of the methods described herein; (d) removing high concentrations of DNA barcode construct sequences associated with generally highly expressed proteins or contaminants in order to increase the ratio of DNA barcode constructs associated with low expressed proteins compared to high expressed proteins, thereby generating ratio-adjusted DNA barcodes; (e) PCR amplifying the sup-diffed DNA barcode constructs; (f) determining the relative amount of the known biomarker by analyzing the number of sequencing reads associated with the known protein biomarker; and (g) determining the presence or absence of a deviation in the expression level of the known biomarker from a standard value, thereby assessing a disease state, assessing a response to a treatment, predicting a response to a treatment, or a combination thereof.

[0036] In certain embodiments, the aptamer library is generated using the RCHT-SELEX methods described herein. In certain embodiments, the aptamer exhibits binding specificity for one N-terminal amino acid. In certain embodiments, the aptamer exhibits binding specificity for two or more N-terminal amino acids. In certain embodiments, the unique barcode sequence indicative of the relevant peptide-binding ssDNA region of the aptamer and sequencing round comprises about 6 to about 20 nucleotides. In certain embodiments, the BCS-compatible portion of the aptamer construct can comprise one or more complementary DNA sequences that hybridize to the aptamer described herein. In certain embodiments, the proximal DNA barcode base contains a unique barcode indicative of the relevant protein or peptide (if known) or the sample from which the protein or peptide is derived.

[0037] In another aspect, articles for protein or peptide sequencing are provided. Such articles generally comprise a library of DNA aptamers, wherein each member of the library exhibits binding specificity for at least one N-terminal amino acid.

[0038] In certain embodiments, each member of the library comprises a common sequence indicative of the cycle number (e.g., first, second, third, etc.) and a unique barcode sequence. In certain embodiments, each member of the library comprises a restriction site. In certain embodiments, each member of the library further comprises at least one sequence for ligation, annealing, or a combination thereof.

[0039] The methods described herein can also be used to sequence full-length proteins.

[0040] The methods described herein can also be used to sequence proteins within a protein complex.

[0041] The methods described herein can also be used to sequence proteins within a complex pool of proteins.

[0042] Other methods to overcome the difficulty of functional P5 adapters being removed from the sequencing chip surface as a result of Edman degradation are also described herein. Loss of functional P5 adapters on the sequencing chip surface prevents the clustering of DNA barcode constructs and thus prevents the ability to sequence directly on the same chip.

[0043] In certain embodiments, after the DNA barcode construct containing the strands of DNA barcodes indicative of the binding order of the aptamers to the peptide is established, the construct can be amplified on the chip or excised from the chip and amplified in solution. The amplification methods used can include, but are not limited to, PCR, loop-mediated isothermal amplification, nucleic acid sequence-based amplification, strand displacement amplification, and multiple displacement amplification. In addition, the original DNA barcode construct can be transcribed on the chip into a large number of RNA constructs, which can then be converted into a cDNA library containing many copies of the original DNA barcode. The amplified products that are copies of the original DNA barcode construct can be removed from the microfluidic chamber and sequenced using standard DNA sequencing methods including, but not limited to, Sanger sequencing, NGS, Ion Torrent sequencing, SOLiD technology, cPAS, etc. The number of reads can be normalized to the number of PCR cycles to estimate the amount of each protein or peptide sequenced from the original sample.

[0044] In certain embodiments, the methods described herein can utilize the empty P7 adapters available on the chip for cluster generation. After the DNA barcode construct is established, a second sequencing primer adapter containing at least (a) the antisense restriction site and (b) the reverse complement of the P7 adapter on the chip can be ligated to the 3' end of the barcode construct. After bridge amplification of the barcode construct, the reverse strand can be selectively cleaved to allow accurate base calling in each individual cluster.

[0045] In another aspect, methods are provided for recording one or more binding events between a plurality of putative binders and a plurality of targets (BCS binding assay). Such methods generally comprise: (a) incubating known putative binding partners with a library of binders of unknown binding affinity and specificity bearing DNA barcodes, wherein each binder within the library comprises a target binder and a unique barcode sequence indicative of the associated binder; (b) ligating the DNA barcode of the target binder to its proximal DNA barcode construct, which itself can contain a unique barcode; (c) optionally removing the target binder, leaving only the barcode of the target binder and a short, consensus sequence covalently attached to the DNA barcode construct for subsequent ligation, in order to record the identity of the binder and thus the putative identity of the bound target; (d) optionally repeating steps (b)-(c) for multiple rounds of validation; (e) optionally, if the binder is an aptamer, not removing the target binder in step (c) but instead ligating sequencing adapters, so that sequencing will be performed directly on the nucleic acid sequence of the binder; and (f) ligating appropriate sequencing adapters; and (g) sequencing the substrate and binder barcodes, thereby identifying a plurality of targets and their binding partners.

[0046] Representative binders include, but are not limited to, aptamers, antibodies, and other small molecule binders. Representative targets include, but are not limited to, peptides, proteins and protein complexes, lipid molecules, viruses, ultramicrobacteria, and inorganic molecules.

[0047] In certain embodiments, the putative binders are attached to a solid substrate, and the targets are modified with DNA barcode tails and in solution.

[0048] In one aspect, methods for sequencing a protein or peptide using fluorescently tagged aptamers are provided. Such methods generally comprise: (a) providing a solid support having attached thereto at least one protein or peptide, wherein the at least one protein or peptide is attached to the solid support via a nucleic acid linker, wherein the nucleic acid linker comprises a sequencing adaptor sequence; (b) incubating the protein or peptide with a library of aptamers that exhibit binding specificity for at least one N-terminal amino acid under conditions in which one or more aptamers within the library specifically bind to the at least one N-terminal amino acid of the protein or peptide, wherein each aptamer within the library comprises a unique optical signature; (c) detecting the unique optical signature and the location of the unique optical signature; (d) removing the aptamer from the protein or peptide and removing the N-terminal amino acid to yield an N-terminal amino acid shortened protein or peptide; (e) incubating the N-terminal amino acid shortened protein or peptide with a library of DNA aptamers that exhibit binding specificity for at least one N-terminal amino acid under conditions in which one or more aptamers within the library specifically bind to the at least one N-terminal amino acid of the protein or peptide, wherein each aptamer within the library comprises a peptide-binding ssDNA region and a unique barcode sequence, the barcode sequence comprising a single DNA barcode indicative of the first probe iteration and the associated peptide-binding ssDNA region; (f) detecting the unique optical signature and the location of the unique optical signature; (g) removing the aptamer from the protein or peptide and removing the N-terminal amino acid to yield an N-terminal amino acid shortened protein or peptide; (h) repeating steps (b)-(g) a plurality of times to construct a chain of locations of optical barcodes; thereby obtaining the sequence of the protein or peptide.

[0049] In another aspect, methods are provided for sequencing a protein or peptide using aptamers complementary to fluorescently tagged probes. Such methods generally comprise: (a) providing a solid support having attached thereto at least one protein or peptide, wherein the at least one protein or peptide is attached to the solid support via a nucleic acid linker, wherein the nucleic acid linker comprises a sequencing adaptor sequence; (b) incubating the protein or peptide with a library of DNA aptamers exhibiting binding specificity for at least one N-terminal amino acid under conditions in which one or more of the aptamers specifically bind to the at least one N-terminal amino acid of the protein or peptide, wherein each aptamer within the library comprises a series of one or more sequences complementary to an optical labeled nucleic acid probe indicative of the sequencing round and associated peptide-binding ssDNA region, and wherein the probe hybridization region is hybridized to a protective complementary oligonucleotide; (c) denaturing and washing away the protective complementary oligonucleotide; (d) incubating the bound aptamer with a fluorescently tagged oligonucleotide probe complementary to a specific region of the aptamer barcode tail; (e) detecting the unique optical signature and the location of the unique optical signature; (f) denaturing and washing away the bound probe; (g) repeating steps (d)-(f) for a desired number of iterations; (h) removing the aptamer from the protein or peptide and removing the N-terminal amino acid to yield an N-terminal amino acid shortened protein or peptide; (i) repeating steps (b)-(h) a plurality of times to construct a chain of locations of optical barcodes; thereby obtaining a sequence of the protein or peptide.

[0050] In another aspect, methods are provided for identifying new biomarkers using any of the protein sequencing methods described herein. Such methods generally comprise: (a) providing a protein sample from a biological sample of interest and a control or comparison biological sample; (b) optionally removing very high concentrations of known proteins; (c) performing steps (a)-(h) of the method of claim 1 or steps (a)-(i) of claim 2; (d) comparing the number of optical barcode reads associated with each lowly expressed protein from the control sample to the sample of interest; thereby identifying putative biomarkers having significantly different relative expression levels between the control sample and the sample of interest.

[0051] In another aspect, methods of using the protein sequencing methods described herein to assess a disease state, to assess a response to a treatment, to predict a therapeutic response, or a combination thereof are provided, wherein one or more signs of the disease is an abnormal expression level of a known protein marker. Such methods generally comprise: (a) providing a protein sample from a patient sample; (b) optionally stripping very high concentrations of known proteins; (c) performing steps (a)-(h) of the method of claim 1 or steps (a)-(i) of the method of claim 2; (d) determining the relative amount of the known biomarker by analyzing the number of optical barcode reads associated with the known protein biomarker; thereby determining the presence or absence of a deviation in the expression level of the known biomarker from a standard value.

[0052] In another aspect, methods of using the protein sequencing methods described herein to screen for potential antibodies are provided. Such methods generally comprise: (a) providing a plasma sample from immunized and non-immunized biological samples; (b) optionally stripping very high concentrations of known proteins; (c) optionally isolating immunoglobulins; (d) performing steps (a)-(h) of the methods described herein or steps (a)-(i) of the methods described herein; (e) comparing the number of optical barcode reads associated with each polypeptide from the non- immunized sample to the immunized sample of interest; thereby identifying putative antibodies having significantly different relative expression levels between the non- immunized sample and the immunized sample of interest.

[0053] In certain embodiments, the method further comprises fragmenting the protein or peptide between step (a). In certain embodiments, the fragmenting step comprises fragmenting the protein or peptide using trypsin, another fragmenting enzyme, or a combination thereof.

[0054] In certain embodiments, the C-terminal end of the protein or peptide is attached to a solid support. In certain embodiments, the C-terminal end of the protein or peptide is attached to an oligonucleotide tail. In certain embodiments, the aptamer is optionally crosslinked to the N-terminal amino acid after step (b) and before step (c).

[0055] In certain embodiments, the protein or peptide is from a biological sample. In certain embodiments, the biological sample is selected from the group consisting of blood, urine, saliva, a tissue biopsy sample, sputum, fecal matter, a single cell, an environmental sample, a bacterial swab, or any sample containing a peptide or protein. In certain embodiments, the protein or peptide is a full-length protein, a peptide fragment, or a protein or peptide contained within a complex.

[0056] In certain embodiments, the unique label is selected from the group consisting of a fluorophore, a dye, a nanolanthanide, and a quantum dot. In certain embodiments, the optically labeled probe is an oligonucleotide complementary to a barcode sequence. In certain embodiments, one or more oligonucleotide probes of one or more colors are hybridized to the aptamer barcode tail in the same iteration of probe incubation. In certain embodiments, the detecting step is performed using optical imaging, total internal reflection fluorescence (TIRF), super-resolution microscopy, structured optical microscopy, wide-field microscopy, or confocal microscopy.

[0057] In certain embodiments, the aptamer library comprises aptamers that are partially dsDNA in regions unrelated to aptamer binding. In certain embodiments, the dsDNA is denatured and the protective complementary oligonucleotide is washed away. In certain embodiments, the bound aptamer is crosslinked to the N-terminal amino acid with PFA. In certain embodiments, the step of removing aptamer comprises cleaving the aptamer with a restriction enzyme. In certain embodiments, the step of removing the N-terminal amino acid comprises Edman degradation of the protein or peptide, cleavage of the protein or peptide with one or more aminopeptidases, heat, pH, or a combination thereof.

[0058] In certain embodiments, the amino acid recognized by a member of the aptamer library is a natural amino acid, an unmodified amino acid, and a modified amino acid. In certain embodiments, the aptamer library is generated using the RCHT-SELEX method described herein. In certain embodiments, the aptamer exhibits binding specificity for one N-terminal amino acid. In certain embodiments, the aptamer exhibits binding specificity for two or more N-terminal amino acids.

[0059] In one aspect, methods of screening a library of DNA aptamers for protein or peptide binding partners are provided. Such methods generally comprise: (a) incubating a plurality of proteins or peptides with a library of DNA aptamer candidates that can exhibit binding specificity to the proteins or peptides under conditions in which the aptamer specifically binds to a protein or peptide of the plurality of proteins or peptides, wherein each protein or peptide of the plurality of proteins or peptides comprises a DNA bridge annealing sequence and a unique DNA barcode, wherein each aptamer within the library comprises a DNA bridge annealing sequence; (b) incubating the barcoded proteins or peptides and DNA aptamer candidates with short oligonucleotide bridges, wherein a portion of the short oligonucleotide bridge is complementary to the bridge annealing sequence at the 3' end of the aptamer, and wherein an additional portion of the short oligonucleotide bridge is complementary to the bridge annealing sequence coupled to the 5' peptide tail; (c) ligating the bridge annealing portion of each element of the aptamer library that specifically binds to a polypeptide to those bridge annealing portions of polypeptides that are connected by the oligonucleotide bridge; (d) amplifying the aptamers within the library that specifically bind to the proteins or peptides; (e) repeating steps (a)-(d) a plurality of times to identify aptamers that exhibit binding specificity to each protein or peptide; and (f) sequencing the annealed aptamers and DNA barcodes; thereby identifying a plurality of polypeptides and their aptamer binding partners.

[0060] In certain embodiments, the amplifying step comprises performing nested PCR. In certain embodiments, the method further optionally comprises separating the proteins or peptides from the aptamers to which they specifically bind and purifying the aptamers prior to step (d). In certain embodiments, the sequencing step uses a next generation sequencing (NGS) platform.

[0061] In one aspect, methods of producing barcoded polypeptides are provided. Such methods generally comprise: transforming an expression construct into a microbial cell under conditions in which about one construct is introduced per cell, wherein the expression construct comprises a nucleic acid encoding: (a) a fusion protein comprising the polypeptide, a purification tag, and a nucleic acid binding protein (naBP); and (b) a nucleic acid sequence recognized by the naBP and a unique nucleic acid barcode; and culturing the microorganism under conditions in which the construct is expressed and the naBP portion of the fusion protein binds to the naBP recognition sequence, thereby producing barcoded polypeptides.

[0062] In certain embodiments, the microbial cell is selected from a eukaryotic or prokaryotic cell. In certain embodiments, the method further comprises purifying the barcoded polypeptide. In certain embodiments, the expression construct comprises any combination of constitutive, inducible, or repressible promoters compatible with the host organism, in any copy number. In certain embodiments, components of the system are expressed using different promoters. In certain embodiments, components of the system are expressed using the same promoter present at different locations within the expression construct. In certain embodiments, the components are expressed using a Gal 1,10-bidirectional promoter, ADH1, GDS, TEF, CMV, EF1a, SV40, T7, lac, or any other promoter and promoter combinations compatible with the host organism.

[0063] In certain embodiments, the purification step comprises pull down of the barcoded polypeptide using a pull down method corresponding to the encoded purification tag. In certain embodiments, the immunoprecipitation step comprises pull down of the barcoded polypeptide with protein purification magnetic beads (e.g. anti-His antibody, agarose, nickel, etc.). In certain embodiments, the method further comprises elution of the barcoded polypeptide from the beads using a mild elution buffer such as glycine to release the fusion peptide without denaturing the RNA-protein / peptide binding.

[0064] In certain embodiments, the polypeptide comprises one or more site-specific protease cleavage sites for release of the barcoded polypeptide from the anti-affinity tag beads using a site-specific protease (e.g. enterokinase, factor Xa, tobacco etch virus protease, thrombin). In certain embodiments, the nucleic acid sequence comprises a restriction enzyme cleavage site for release of the barcoded polypeptide from the beads using a restriction endonuclease.

[0065] In certain embodiments, the nucleic acid sequence recognized by the nucleic acid binding protein and the nucleic acid binding protein are an MS2 RNA hairpin or variant thereof and an MS2 bacteriophage coat protein or mutant thereof. In certain embodiments, the nucleic acid recognized by the nucleic acid binding protein and the nucleic acid binding protein are a boxB sequence or variant thereof and bacteriophage anti-terminator protein N (λN).

[0066] In certain embodiments, the cells are irradiated with UV radiation prior to purification of the barcoded polypeptide. In certain embodiments, the purified complex is irradiated with UV radiation.

[0067] In another aspect, there is provided a DNA-barcoded polypeptide or protein made by the methods described herein.

[0068] In another aspect, methods of producing dsDNA oligonucleotides with high control over sequence content are provided. Such methods generally comprise: (a) ligating a first LEGO block of dsDNA having a 5' phosphorylated single nucleotide overhang in the direction of sequence extension to a second LEGO block of dsDNA having a 5' phosphorylated single nucleotide overhang at each end, one overhang of the second LEGO block being complementary to the overhang of the first LEGO block and the other overhang being non-complementary, using a dsDNA ligase, thereby leaving one 5' phosphorylated single nucleotide overhang on the second LEGO block in the direction of sequence extension; (b) ligating the second LEGO block of dsDNA to a third LEGO block of dsDNA having a 5' phosphorylated single nucleotide overhang at each end, one overhang of the third LEGO block being complementary to the overhang of the second LEGO block and the other overhang being non-complementary, using a dsDNA ligase, thereby leaving one 5' phosphorylated single nucleotide overhang on the third LEGO block in the direction of sequence extension; (c) repeating steps (a)-(b) a plurality of times until the sequence construct is one LEGO block shorter than the desired length; and (d) ligating the sequence construct to a last LEGO block of dsDNA having a 5' phosphorylated single nucleotide overhang in the opposite direction of sequence extension.

[0069] In certain embodiments, the 3' or 5' modification of the LEGO blocks is compatible with the dsDNA ligase used. In certain embodiments, to produce a random library, a heterogeneous pool of LEGO blocks is used at specific positions where diversity is desired. In certain embodiments, the double stranded LEGO blocks are enzymatically ligated using T4 DNA ligase or any other dsDNA ligase compatible with the 3' or 5' end modification utilized by the ligase of choice. In certain embodiments, the ligation reaction is performed in solution, on beads, on a solid support, in a gel, etc. In certain embodiments, the first dsDNA LEGO block is a PCR primer. In certain embodiments, the last dsDNA LEGO block is a PCR primer. In certain embodiments, the dsDNA product is PCR amplified to produce a library with parallel samples. In certain embodiments, the dsDNA product after PCR amplification is digested to produce a ssDNA library.

[0070] In another aspect, methods of generating ssDNA oligos with high control over sequence content are provided. Such methods generally comprise: (a) ligating the 3' end of a first ssDNA LEGO block to the 5' end of a second ssDNA LEGO block, wherein one of the ends involved in the ligation is phosphorylated; (b) ligating the 3' end of the second ssDNA LEGO block to the 5' end of a third ssDNA LEGO block, wherein one of the ends involved in the ligation is phosphorylated; (c) repeating steps (a)-(b) a plurality of times until a sequence construct is one LEGO block shorter than desired; and (d) ligating the sequence construct to a final LEGO block.

[0071] In certain embodiments, the 3' or 5' modifications of the LEGO blocks are compatible with the ssDNA or RNA ligase used. In certain embodiments, single stranded LEGO blocks are enzymatically ligated using RtcB ssRNA ligase, CircLigase, or any other ssDNA or RNA ligase compatible with the 3' or 5' end modifications required by the ligase of choice. In one embodiment, the ligation reaction is performed in solution, on beads, on a solid support, in a gel, etc. In certain embodiments, the first ssDNA LEGO block is a PCR primer. In certain embodiments, the final ssDNA LEGO block is a PCR primer. In certain embodiments, the ssDNA product is PCR amplified to generate a library of double stranded parallel samples. In certain embodiments, the dsDNA product after PCR amplification is digested to generate a library of ssDNA.

[0072] In another aspect, methods of generating RNA oligos with high control over sequence content are provided. Such methods generally comprise: (a) ligating the 3' end of a first RNA LEGO block to the 5' end of a second RNA LEGO block, wherein one of the ends involved in the ligation is phosphorylated; (b) ligating the 3' end of the second RNA LEGO block to the 5' end of a third RNA LEGO block, wherein one of the ends involved in the ligation is phosphorylated; (c) repeating steps (a)-(b) a plurality of times until a sequence construct is one LEGO block shorter than desired; and (d) ligating the sequence construct to a final LEGO block.

[0073] In certain embodiments, the 3’ or 5’ modification of the LEGO block is compatible with the RNA ligase used. In certain embodiments, the RNA LEGO blocks are enzymatically ligated using any RNA ligase compatible with the 3’ or 5’ end modification required by the chosen ligase. In certain embodiments, the ligation reaction is performed in solution, on beads, on a solid support, in a gel, etc. In certain embodiments, the first RNA LEGO block is a PCR primer. In certain embodiments, the last RNA LEGO block is a PCR primer. In certain embodiments, to generate a ssDNA library, the RNA product is reverse transcribed into cDNA, the second strand is synthesized with a DNA polymerase, the dsDNA product is PCR amplified, and the antisense strand is digested.

[0074] In another aspect, there is provided a pool of oligonucleotides manufactured by any of the methods described herein.

[0075] Definitions

[0076] A nucleic acid can be single-stranded or double-stranded, which generally depends on its intended use. As used herein, an “isolated” nucleic acid molecule is a nucleic acid molecule that does not contain sequences flanking one or both ends of the nucleic acid molecule in the genome of the organism from which the isolated nucleic acid molecule is derived (e.g., a cDNA or genomic DNA fragment produced by PCR or restriction endonuclease digestion). Such an isolated nucleic acid molecule is typically introduced into a vector (e.g., a cloning vector or an expression vector) to facilitate manipulation or to produce a fusion nucleic acid molecule, which is discussed in more detail below. In addition, an isolated nucleic acid molecule can be an engineered nucleic acid molecule, such as a recombinant or synthetic nucleic acid molecule.

[0077] An aptamer is a single-stranded nucleic acid sequence, which can be composed of RNA, DNA, XNA such as TNA, modified nucleic acids (e.g., replacing natural DNA nucleotides with alternative functional groups (Chelsea et al., 2019 and Pfeiffer et al., 2017)) or other synthetic nucleic acid analogs. Aptamers are typically identified using the SELEX assay, which relies primarily on the evolution of a diversified pool of sequences using PCR round-by-round amplification. Aptamer sequences typically have 20-45 base pairs (bp) plus additional flanking primer regions (typically 20-23 bp in length for both forward and reverse primers). Capillary electrophoresis SELEX (CE-SELEX) does not rely on aptamers with primer regions, however CE-SELEX is limited to nL scale working volumes, thereby reducing the initial sequence starting pool from 10 14 -10 16 limited to 10 8 -10 9 .

[0078] As used herein, a "purified" polypeptide is one that has been separated or purified away from cellular components naturally accompanying it. Generally, a polypeptide is considered "purified" when at least 70% (e.g., at least 75%, 80%, 85%, 90%, 95%, or 99%) by dry weight of the polypeptide is not a polypeptide and a naturally occurring molecule naturally accompanying it. A chemically synthesized polypeptide is "purified" by nature of being separated from the components naturally accompanying it.

[0079] Nucleic acids can be isolated using techniques routine in the art. For example, nucleic acids can be isolated using any method including, but not limited to, recombinant nucleic acid techniques and / or polymerase chain reaction (PCR). Ordinary PCR techniques are described in, for example, PCR Primer: A Laboratory Manual, Dieffenbach & Dveksler, eds., Cold Spring Harbor Laboratory Press, 1995. Recombinant nucleic acid techniques include, for example, restriction enzyme digestion and ligation, which can be used to isolate nucleic acids. Isolated nucleic acids can also be chemically synthesized, either as a single nucleic acid molecule or as a series of oligonucleotides, by traditional methods such as bead purification, enzyme digestion, column purification, and the like.

[0080] Polypeptides can be purified from natural sources (e.g., biological samples) by known methods such as DEAE ion exchange, gel filtration, HIS-tag bead pull-down, affinity chromatography, and hydroxyapatite chromatography. Polypeptides can also be purified, for example, by expressing nucleic acids in expression vectors. In addition, purified polypeptides can be obtained by chemical synthesis. The purity of a polypeptide can be measured using any suitable method, such as column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis.

[0081] Vectors containing nucleic acids (e.g., nucleic acids encoding polypeptides) are also provided. Vectors, including expression vectors, are commercially available or can be produced by routine recombinant DNA techniques in the art. Vectors containing nucleic acids can have expression elements operably linked to such nucleic acids, and can also include sequences encoding selectable markers (e.g., antibiotic resistance genes). Vectors containing nucleic acids can encode chimeric or fusion polypeptides (e.g., polypeptides operably linked to heterologous polypeptides, which can be at the N-terminus or C-terminus of the polypeptide). Representative heterologous polypeptides are polypeptides useful for purification of the encoded polypeptide (e.g., 6xHis tag, glutathione S-transferase (GST)).

[0082] Expression elements include nucleic acid sequences that direct and regulate expression of nucleic acid coding sequences. One example of an expression element is a promoter sequence. Expression elements can also include introns, enhancer sequences, response elements, or inducible elements that modulate nucleic acid expression. Expression elements can be of bacterial, yeast, insect, mammalian, or viral origin, and vectors can contain combinations of elements from different sources. As used herein, operably linked means that the promoter or other expression element is placed in the vector in a manner so as to direct or regulate expression of the nucleic acid.

[0083] The vectors described herein can be introduced into a host cell. As used herein, "host cell" refers to the particular cell into which a nucleic acid is introduced, and also includes the progeny of such a cell that has the vector. Host cells can be any prokaryotic or eukaryotic cell. For example, the nucleic acid can be expressed in a bacterial cell, such as E. coli, or in an insect cell, yeast, or mammalian cell, such as Chinese hamster ovary cells (CHO) or COS cells. Other suitable host cells will be known to those of skill in the art. Numerous methods for introducing nucleic acids into host cells in vitro and in vivo are known to those of skill in the art, and include, but are not limited to, electroporation, calcium phosphate precipitation, polyethylene glycol (PEG) transformation, heat shock, lipofection, microinjection, and viral-mediated nucleic acid transfer.

[0084] As used herein, "specific" recognition or "specific" binding means that a molecule exhibits high substrate specificity for a given target within a known operating concentration range, and very low or no substrate specificity for any other target.

[0085] As used herein, "semi-specific" recognition or "semi-specific" binding means that a molecule exhibits high substrate specificity for a known target, and intermediate to low binding specificity for a subset of other targets.

[0086] As used herein, "prefix" means at least the N-terminal amino acid, and can also include the penultimate N-terminal amino acid at the N-terminus of a protein or peptide.

[0087] As used herein, "suffix" means one or more amino acids in a peptide C-terminal to the "prefix" amino acid as defined above.

[0088] As used herein, "DNA barcode" means an oligonucleotide sequence bearing information indicative of the identity of at least one molecule. Although the barcode is referred to throughout this text as a construct of "DNA", the barcode molecule can in fact comprise DNA, RNA, XNA, modified nucleic acids, or combinations thereof.

[0089] As used herein, "DNA barcode construct" refers to a DNA strand comprising at least two DNA barcodes.

[0090] As used herein, "barcode sequencing (BCS) compatible" aptamer refers to a partially double-stranded aptamer wherein one or more regions not involved in target binding can hybridize to a complementary oligonucleotide and can or can not contain overhangs.

[0091] As used herein, "blocked aptamer" refers to a partially double-stranded aptamer wherein at least the primer region of the aptamer, but not the aptamer region itself, can hybridize to a protective complementary oligonucleotide.

[0092] As used herein, "sup-diff" refers to a method of removing DNA barcode constructs of highly expressed proteins.

[0093] As used herein, "optical barcode" or "optical signature" refers to the detection of a fluorescently tagged molecule either directly integrated into an oligonucleotide or attached via one or more conjugates.

[0094] As used herein, "optical barcode" refers to an ordered combination of optical signatures.

[0095] As used herein, "dsDNA LEGO block" refers to a DNA oligonucleotide 5 or more base pairs in length with a 5' nucleotide overhang (e.g., one or more nucleotides) at one or both ends, wherein the most 5'-end nucleotide on at least one strand is phosphorylated.

[0096] As used herein, "ssDNA LEGO block" refers to a DNA oligonucleotide 5 or more nucleotides in length with a phosphorylated 3' or 5' end.

[0097] As used herein, "RNA LEGO block" refers to a RNA oligonucleotide 5 or more nucleotides in length with a phosphorylated 3' or 5' end.

[0098] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the subject methods and compositions belong. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the subject methods and compositions, suitable methods and materials are described below. Also, the materials, methods, and examples are illustrative only and not intended to be limiting. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. BRIEF DESCRIPTION OF DRAWINGS

[0099] Figure 1is a schematic depicting how all of the individual inventions described herein contribute to the development of the PROSEQ platform.

[0100] Figure 2 is a schematic showing the two-amino acid identity redundancy scheme, where each diabody binding event provides the putative identity of two N-terminal amino acids, and each round of degradation removes only one amino acid, allowing each amino acid except the original N-terminal amino acid to be exposed to two rounds of aptamer binding.

[0101] Figure 3A is a schematic showing the steps in a representative conventional SELEX method.

[0102] Figure 3B is a schematic showing the steps in one embodiment of the ML-SELEX method described herein.

[0103] Figure 4A The schematic of FIG. 1 shows that conventional SELEX methods can undesirably enrich for aptamers that bind to components of the selection process ("non-specific high-affinity binders") in addition to aptamers that bind to the desired target ("specific high-affinity binders").

[0104] Figure 4B The schematic of FIG. 2 shows that the addition of a negative selection step in the SELEX methods described herein can reduce the eventual enrichment of aptamers that non-specifically bind to selection components by first removing aptamers that bind to beads, biotin, oligonucleotides, or other selection components before incubation amplification or input into SELEX.

[0105] Figure 5A is a schematic demonstrating the various different steps in the RCHT-SELEX procedure (from Figure 2 ) that can incorporate single incubation experiments, double incubation experiments, and / or intra-experimental replicates.

[0106] Figure 5B is a schematic demonstrating the single incubation experiments, double incubation experiments, intra-experimental replicates, and full-bead control experiments that can be used in parallel or sequentially during the RCHT-SELEX methods described herein.

[0107] Figure 6 is a schematic showing a bead-based multiplexed version of RCHT-SELEX that allows for the selection of aptamers against multiple targets in each experiment. Aptamers identified in the bead-based multiplexed version of RCHT-SELEX can be de-multiplexed in the final round by incubating those aptamers with beads that are coupled to only one of the initial targets, respectively.

[0108] Figure 7Figure 1 is a schematic showing the identification method of aptamers that specifically bind to N-terminal amino acid prefixes independent of the composition of the peptide suffix tail, the method comprising assaying aptamers in iterative rounds, wherein the peptide suffix is changed in different rounds while the desired N-terminal amino acid prefix remains the same. Four types of iterations are shown: dipeptide switch (column 1), wherein the N-terminal amino acid remains the same while the suffix is switched; single amino acid switch (column 2); consistent peptide target (column 3); complete switch or null (column 4), wherein the peptide target is completely different between alternating rounds.

[0109] Figure 8 Figure 2 is a schematic showing how double-stranded (ds) DNA can be converted to single-stranded (ss) DNA using lambda exonuclease. Lambda exonuclease has a bias towards degrading targets that are phosphorylated on the 5' end at a ratio of approximately 20: 1. In order to fold and bind to peptides, aptamers must be single-stranded, so the bound aptamer is PCR amplified with a specific protection / phosphorylation primer, which produces dsDNA, which is then digested with lambda exonuclease to convert the amplified product so that the forward ssDNA aptamer survives.

[0110] Figure 9 Figure 3 A-9C are electropherograms showing the extent of lambda exonuclease digestion of random aptamer libraries monitored on an Agilent Bioanalyzer Chip system using the Small RNA kit. Representative bioanalyzer curves are shown corresponding to (A) dsDNA, (B) partially digested DNA, and (C) ssDNA aptamer. Data are shown in the gel-like images to the right of each electropherogram, with the green line representing the RNA marker. Confirmation of complete conversion to ssDNA is performed prior to introducing each aptamer library to each new RCHT-SELEX round.

[0111] Figure 10A Figure 4-10C is a schematic showing control and sham experiments that can be used in the SELEX method described herein. Positional incorporation added in specific wells of a 96-well plate can be used to determine local contamination between wells (A). Different incorporations are added at different stages of SELEX (i.e. prior to incubation, prior to PCR amplification after each round of incubation, and in each NGS sample) to determine PCR bias steps at each step (B). In sham SELEX, the removed incubation is incubated in the absence of beads and target and PCR amplified (C).

[0112] Figure 11A Figure 5 is a schematic showing threshold PCR, wherein DNA from different samples at different concentrations is PCR amplified to near equal concentrations to ensure near equal input is introduced in each reaction in subsequent SELEX rounds.

[0113] Figure 11B Figure 3 is a plot showing the expression intensity of each 8mer combination from sequencing runs of DNA pools before (top panel) and after (bottom panel) threshold PCR. The X and Y axes are each possible 4mer DNA sequences. The expression intensity is very similar between pools, with a log variance of 0.132.

[0114] Figure 11C Figure 4 is a heatmap reporting the log ratio of the quotient of expression intensity of each 8mer combination from sequencing runs of DNA pools after and before threshold PCR in Figure 11B Figure 4 is a heatmap reporting the log ratio of the quotient of expression intensity of each 8mer combination from sequencing runs of DNA pools after and before threshold PCR in

[0115] Figure 12 Figure 5 is a schematic showing that primer switching can be used to select aptamers with binding affinity that are not dependent on the primer region.

[0116] Figure 13 Figure 6 is a schematic showing the peptide sequencing method described herein. Step 0 comprises establishing a substrate consisting of a 5' phosphorylated barcode backbone, forward and reverse co-localized linkers, and a protein or peptide target (PT) bearing a C-terminal oligonucleotide sequence tag oriented with the 3' end attached to the protein or peptide and a free phosphorylated 5' end; Step 1 comprises attaching the peptide-substrate complex to a solid substrate; Step 2 comprises incubating the bound protein or peptide with a barcoded aptamer library under conditions that allow specific aptamer binding to the appropriate N-terminal amino acid; Step 3 comprises attaching the aptamer tail to a second oligonucleotide bound to the substrate; and Step 4 comprises cleaving off the aptamer, leaving a DNA barcode bound to the specific amino acid bound to the second oligonucleotide. After removal of the N-terminal amino acid from the protein or peptide using Edman degradation and / or aminopeptidase, steps 2-5 are repeated, generating a strand of DNA barcodes that can be used to identify each subsequent N-terminal amino acid.

[0117] Figure 14 is a schematic showing the construct of an aptamer tail and bridge oligonucleotide. Figure 14A Figure 15 is a schematic depicting a barcode-specific bridge, wherein the bridge is fully complementary to the aptamer tail including the barcode region except for a 3' single-stranded overhang region. Figure 14B Figure 16 is a schematic depicting a universal bridge, wherein the bridge is only complementary to both the restriction site spacer and consensus sequence that are conserved in all aptamers and flanking the barcode.

[0118] Figure 15Ais a schematic showing the peptide or protein sequencing method described herein, wherein the peptide or protein sequence is determined on the basis of a DNA sequence. In this embodiment, step 1 comprises attaching the C-terminal end of a protein or peptide to a DNA primer oligonucleotide bound to a substrate; step 2 comprises incubating the bound protein or peptide with a library of barcoded aptamers under conditions that allow for specific aptamer binding to the appropriate N-terminal amino acid; step 3 comprises ligating an aptamer tail to a second oligonucleotide bound to the substrate; and step 4 comprises cleaving off the aptamer, leaving behind the DNA barcode bound to the specific amino acid bound to the second oligonucleotide. After removal of the N-terminal amino acid from the protein or peptide using Edman degradation and / or aminopeptidase, steps 1-4 are repeated, generating a strand of DNA barcodes that can be used to identify each subsequent N-terminal amino acid.

[0119] Figure 15B is a schematic showing an example of the correlation between a single amino acid and the corresponding aptamer barcode.

[0120] Figure 16 is a schematic showing prior and non-prior sup-diff strategies to remove DNA constructs bound to a known target or unknown but high concentration of DNA constructs.

[0121] Figure 17 shows an example of a variation of the steps in the PROSEQ platform.

[0122] Figure 18is a heatmap showing the estimated percentage of the human proteome that is potentially identifiable for each library of up to 100 binders, each of which has up to 400 different dipeptide binding on the ProSeq platform, where a protein is digested at each lysine, resulting in 12mer or smaller peptides. The simulation to obtain the details of the percentage proteome coverage for a hypothetical binder set was performed as follows: (a) the protein is digested into fragments with LysC, (b) one of the fragments of a protein is uniquely identified when it has a matching barcode that is unique in the proteome, at which point the protein is identified, (c) a set of dipeptides (pairing amino acids) is randomly selected for the binder to have affinity for out of 400 possible dipeptides, (d) a set of 20 binders is randomly selected, (e) given the set of binders and the dipeptides each binder has affinity for, the barcodes read for each protein fragment are determined and the number of uniquely identified proteins is determined, (f) 12 rounds of Edman degradation, binding, and barcode determination are performed for each fragment. The simulation does not simulate noise (binders not binding when they should or binding when they should not). In a real system, some noise can be reduced by redundancy in dipeptide reads and by reading multiple copies of the same protein. Furthermore, only 20 possible sets were evaluated to obtain the matching percentage, so the curves are expected to be more smooth for lower specificity binder sets.

[0123] Figure 19 is a schematic showing the binding confirmation method described herein. Step 0 includes establishing a substrate consisting of a 5’ phosphorylated barcode substrate, forward and reverse co-localization linkers, and a target with a C-terminal oligonucleotide sequence tag oriented with the 3’ end attached to a protein or peptide and a free phosphorylated 5’ end; Step 1 includes attaching the target-substrate complex to a solid substrate; Step 2 includes incubating the target with a hypothetical library of barcoded binders under conditions that allow the hypothetical binders to bind to the target; Step 3 includes attaching the oligonucleotide barcode tail to the second proximal substrate oligonucleotide barcode attached to the substrate; and Step 4 includes cleaving off the binder barcode tail, leaving the barcode that was bound to the specific hypothetical binder attached to the substrate oligonucleotide barcode. Optionally, after the hypothetical binders are removed from the attached target, steps 2-5 are repeated, resulting in a strand of DNA barcodes that can be used to identify multiple binding events. Note that the binding event is not limited to the N-terminal amino acid or the free end of the attached target, and can occur at any exposed region of the target.

[0124] Figure 20is an overview of the peptide sequencing method described herein, where the peptide or protein sequence is determined using fluorescence and microscopy. Peptides are attached to a known linker on a chip (A). A library of fluorescent dye conjugated aptamers selected for binding to a specific N-terminal amino acid is flowed over the peptide, incubated with the target, and unbound aptamers are washed off the chip (B). The optical barcodes of bound aptamers are imaged. For each round, a z-axis multi-layer scan of the image is taken to generate a spectral signature of the N-terminal amino acid (C). The N-terminal amino acid is removed from the immobilized peptide, the sample is washed, and the same pool of aptamers is flowed over it to interrogate the newly exposed N-terminal amino acid (D). After repeating this series of steps on the slide, the identity of each successive N-terminal amino acid for each round can be calculated by comparing the optical barcode of each peptide to the organism's proteome.

[0125] Figure 21 is a schematic showing one embodiment of the methods described herein, where proteins are isolated from cells and processed before the proteins are attached to a solid substrate. For example, cells can be lysed and proteins isolated (A), then denatured and digested (B). The side chains and N-terminus of the peptides can be protected (C), the C-terminal amino acid modified with an oligonucleotide or linker (D), and attached to a solid substrate (E). Optically labeled aptamers can be flowed over the complex (F), the image captured, and the process repeated.

[0126] Figure 22 is a schematic showing the construction of an aptamer with regions that bind to complementary fluorescently tagged oligonucleotides. The aptamer comprises (a) an active binding region, (b) an optional spacer, and (c) one or more barcode tails of one or more combinations of barcode units (BCs) that indicate the number of probing iterations and fluorescent tags, where each BC is complementary to a fluorescently tagged oligonucleotide. The barcode tails are designed in two variations: (1) the BCs are spatially separated and can anneal to one or up to all unique complementary probes at the same time, and (2) the BCs are designed such that the BC sequences overlap and can only anneal to probes complementary to non-overlapping BCs at the same time. Note that since the BC sequences themselves contain information of the number of probing iterations, the BCs do not have to be spatially oriented in the order of the time of probing iteration incubation (as shown in the figure).

[0127] Figure 23is a schematic showing the peptide sequencing method described herein. Step 1 includes immobilizing a peptide-oligonucleotide target on a solid substrate; Step 2 includes incubating the bound protein or peptide with a library of barcoded aptamer probes under conditions that allow specific aptamers to bind specifically to the appropriate N-terminal amino acid; Step 3 includes removing the protective complementary oligonucleotide, exposing the barcode region for probe annealing; Step 4 includes incubation with a library of probes that hybridize to the barcode region indicating probe iteration 1; Step 5 includes washing away unbound probes and imaging the bound probes; Step 6 includes denaturing the bound probes from the aptamer and washing the probes off the substrate; Step 7 includes repeating steps 4-6 for all probe iterations required to identify the aptamer. Steps 2-8 are repeated after removal of the N-terminal amino acid from the protein or peptide using Edman degradation and / or aminopeptidase, resulting in a series of optical barcodes that can be used to identify each subsequent N-terminal amino acid.

[0128] Figure 24 is a schematic depicting the method described herein for PROSEQ VIS when the library of aptamer probes consists of high affinity binders that specifically bind to unique N-terminal amino acid prefixes. Individual binding events that indicate the hypothesized identity of the N-terminal amino acid prefix being probed are observed by detecting a unique combination of aptamers directly coupled to dyes or dye-coupled oligonucleotides hybridized to the aptamers. In Step 1, a peptide is positioned to the sequencing platform and incubated with aptamers that recognize specific N-terminal dipeptides. In Step 2, each aptamer has multiple binding sites for binders coupled to dyes. These strong binders can be hybridized to the aptamer simultaneously and still remain bound. The identity of the aptamer is determined by assessing the combination of colors detected at each position and, by extension, to the identity of the N-terminal amino acid (SEQ ID NO: 121). In Step 3, the aptamer is washed away and the new N-terminal amino acid is exposed by degradation. The cycle is repeated for the remaining amino acids (SEQ ID NO: 122).

[0129] Figure 25is a schematic depicting the method for PROSEQVIS described herein when the library of aptamer probes consists of non-specific binders to a set of N-terminal amino acid prefixes and have variable probability distributions for each unique binding pair. Multiple binding events indicating the hypothesized identity of the N-terminal amino acid prefix probed are observed by detecting dye-modified aptamers in multiple cycles of incubation and wash-off. In step 1, peptides are positioned to the sequencing platform and incubated with aptamers recognizing a set of N-terminal dipeptides. In step 2, dye-coupled binders hybridize to the single-stranded portion of the aptamer, but as they are "weak" binders, they lack the specificity of stronger binders. As the cycle progresses, dye-coupled binders that fluoresce at each peptide position are tracked to determine the accuracy of amino acid visitation rates. Individual colors or optical barcodes can be used. In step 3, the identity of the N-terminal amino acid at each round is calculated by comparing the combination of observed fluorescent signals to the probability distribution of binding events for each aptamer to each N-terminal amino acid prefix (SEQ ID NO: 123).

[0130] Figure 26 A-26C is a schematic showing the multiplexing method described herein. A library of aptamers (A) is incubated with unbound DNA barcode-bearing protein or peptide targets (B). After the aptamer binds to the barcoded target, the 3' end of the single-stranded aptamer is ligated to the ssDNA barcode specific for the target identity (C) through a ssDNA bridge that is semi-complementary to the 3' end of the aptamer and semi-complementary to the 5' end of the ssDNA barcode. The gap between the aptamer and the ssDNA peptide barcode can be ligated and run through to obtain the aptamer sequence and the peptide barcode, which in turn provides the target to which the aptamer bound. Figure 26 B) is a schematic showing the multiplexing method described herein. A library of aptamers (A) is incubated with unbound DNA barcode-bearing protein or peptide targets (B). After the aptamer binds to the barcoded target, the 3' end of the single-stranded aptamer is ligated to the ssDNA barcode specific for the target identity (C) through a ssDNA bridge that is semi-complementary to the 3' end of the aptamer and semi-complementary to the 5' end of the ssDNA barcode. The gap between the aptamer and the ssDNA peptide barcode can be ligated and run through to obtain the aptamer sequence and the peptide barcode, which in turn provides the target to which the aptamer bound.

[0131] Figure 26 D is a schematic indicating the steps of the SELEX procedure (from FIG. 3) into which the multiplexing technique can be incorporated.

[0132] Figure 27 is a schematic of a peptide-oligonucleotide conjugate (POC) comprising a single-stranded (ss) DNA tail (a) covalently linked to the C-terminus of a peptide or protein target (b) at its 3' end. The ssDNA tail (a) comprises a 3' primer region (c), a unique DNA barcode (d), and a 5' bridge binding sequence (e). An aptamer (f) comprises a 3' bridge binding sequence (g). A short oligonucleotide bridge (h) that is half-complementary to the 3' bridge binding sequence (g) at the 3' end of the aptamer (f) and half-complementary to the 5' bridge binding sequence (e) of the ssDNA tail (a) can be used to ligate the aptamer (f) to the peptide (b).

[0133] Figure 28 is a schematic of the nested PCR technique in multiplexing.

[0134] Figure 29 is a schematic showing the in vivo production of barcoded (D) complexes of protein of interest (POI) (A) in the TURDUCKEN method described herein. This method takes advantage of the non-covalent interaction between RNA binding proteins (B) and their corresponding binding sites (C).

[0135] Figure 30 A-30C is a schematic showing one embodiment of the TURDUCKEN method described herein. Pools of plasmids containing various different protein of interest (POI)-RNA binding protein (RBP) fusion genes and their corresponding RNA barcode sequences are transformed into cells at an approximate dilution of 1 plasmid per cell (A), the POI-RBP fusions are expressed and bind their corresponding RNA barcodes (B), which are then purified (C).

[0136] Figure 30 D is a schematic indicating the steps of the SELEX procedure (from Figure 3) into which the TURDUCKEN method can be incorporated.

[0137] Figure 31A-31B is a schematic showing one embodiment of the LEGO method described herein for dsDNA (A) ligation and ssDNA and RNA ligation (B).

[0138] Figure 32 A-32C is a schematic showing one embodiment of the LEGO method described herein. Pools of LEGO blocks of the first, second, third, etc. position (A) are sequentially ligated (B) and PCR amplified to create parallel samples. The resulting dsDNA is then digested into ssDNA to form a library of folded aptamers (C).

[0139] Figure 32 D is a schematic indicating the steps of the SELEX procedure (from Figure 3) into which the LEGO method can be incorporated.

[0140] Figure 33 is a schematic of the universal workflow for all SELEX (RCHT-SELEX and NTAA-SELEX) experiments.

[0141] Figure 34 A is a schematic depicting the 400 possible amino acid prefixes that the SELEX method described herein uses to find aptamers for PROSEQ and PROSEQ VIS.

[0142] Figure 34B is a schematic depicting how the 400 possible amino acid prefixes are organized into 16 chunks.

[0143] Figure 34 C is a schematic depicting how the suffixes paired with 2-mer prefixes alternate between odd and even rounds, with only the 2-mer prefixes having a constant peptide composition across all 4 rounds.

[0144] Figure 34 D is a specific example of how the suffixes ("backbones") are switched in alternating rounds while the prefixes remain the same to find aptamers specific for DD and DC prefixes regardless of the suffix (DD / DD, SEQ ID NOs: 124-127; DC / DC, SEQ ID NOs: 128-131; DD / DC, SEQ ID NOs: 132-135). The same incubations were also used to assay targets with alternating backbones and prefixes similar to the aptamers picked that were not specific for DD and DX.

[0145] Figure 35 is a comparison of the implementation of three variations of SELEX aptamer incubations with peptides (Variations 1-3) to the BCS conditions (BCS).

[0146] Figure 36 is a plot showing the log of the ratio of the expression level of each 12-mer combination from sequencing runs of the post-incubation DNA pools divided by the expression level pre-incubation for 96 conditions, where two conditions failed (the two plots in the lower right). The X and Y axes of each plot are the possible each 6-mer DNA sequences. Plots with high red or blue scales show increased variance from the Gaussian distribution, indicating that the experimental conditions perturbed the random input pool further from its state at input.

[0147] Figure 37 is two tables showing the sequences and read counts of the top 20 common sequences randomly sampled from 100,000 reads in the post-SELEX and post-pseudo-SELEX aptamer pools. The sequences from the pseudo-SELEX (SEQ ID NOs: 136-155) are all different from the sequences from SELEX (SEQ ID NOs: 156-175), indicating that the aptamers pulled down by the peptide targets exhibit higher affinity than random sequences.

[0148] Figure 38Table 1 is a table showing the counts of parallel sample sequences between 9 experiments (3 targets each 3 parallel experiments) performed using the same incubation pool. All parallel experiments of each round were pooled together and non-specific aptamers were filtered out by bead control subtraction. Counts highlighted in red are counts of identical sequences found in experiments of different targets. Bradykinin 1 r5 means that the target is Bradykinin, position 1, parallel 1 and SELEX round 5. GNRH 4 r5 means that the target is GnRH, position 4, parallel 1 and SELEX round 5. Sequence contamination occurs between the closest parallel indicated in red area, which is significantly reduced after changing the automation process and the position of the targets on the plate.

[0149] Figure 39 Table 2 are two examples of aptamers selected against small peptides using the RCHT-SELEX method herein: one against vasopressin (SEQ ID NOs: 176-179) and the other against bradykinin (SEQ ID NOs: 180-183). Aptamer structures are the lowest Gibbs free energy structures obtained by IDT's proprietary UNAFold software.

[0150] Figure 40 Table 3 reports the top 5 aptamer sequences specifically continuously enriched in the presence of peptides prefixed with an N-terminal lysine (SEQ ID NOs: 184-188) or N-terminal cysteine (SEQ ID NOs: 189-193) identified in the peptide switch ML-SELEX experiment. These results illustrate the ability of ML-SELEX to discover unique aptamers against each amino acid.

[0151] Figure 41 A is a schematic of the N-terminal amino acid SELEX experiment strategy of Example 2. 12 selection parallels containing each target mixture were run for 5 rounds. The workflow started with the initial pool of ssDNA being negatively selected against streptavidin beads and split into 12 random pools. 2 parallel selections were performed for each control reference target and 3 parallel selections were performed for the target (proline-proline) with and without backbone switching (C and D backbone) in alternating rounds. Representative pools of ssDNA from each selection of each round were sequenced and sequence enrichment across rounds was analyzed.

[0152] Figure 41 B reports the target composition and amino acid sequences (SEQ ID NOs: 194-203) in non-switched and switched SELEX.

[0153] Figure 42The sequencing counts of the top 10 most enriched sequences of each round are reported. The X-axis is the round of SELEX and the Y-axis is the number of counts seen during sequencing of the 10 sequences. The 10 sequences shown were selected for their calculated enrichment values.

[0154] Figure 43 A is a box plot summarizing the enrichment of the top ranked aptamers for each target. Specifically, enrichment was calculated from round 2 to round 5. Each box plot shows a summary (minimum, first quartile, median, third quartile, and maximum) of the top ten aptamers from each selection performed for a given target. (Total number of sequences = 20 for scaffold, bradykinin, beads, and 30 for PPC and PPCD). The X-axis is on a log scale and shows enrichment. The Y-axis is the target of each selection. The enrichment of the PPCD switch is higher than the negative control (beads) but lower than the positive control (bradykinin).

[0155] Figure 43 B is a categorical scatter plot reporting the difference in enrichment between the top ranked enriched sequences of each selection for each target. There were 2 selections for scaffold, beads, and bradykinin. There were 3 selections for PPC and PPCD. (Total number of sequences = 20 for scaffold, bradykinin, beads, and 30 for PPC and PPCD). The Y-axis is the target and the x-axis is enrichment (penalized growth). For some selections / parallel samples, higher enrichment was observed for the same target. For example, 3 unique sequences in parallel sample 2 observed high enrichment (>3, equivalent to 1000 fold) while only 1 unique sequence in parallel sample 1 observed high enrichment in the selection performed on the target scaffold.

[0156] Figure 44 is a confusion matrix of the top 10 enriched sequences of each parallel sample for each target (scaffold, beads, bradykinin, PPC-C, PPC-CD). 0 indicates no sequence overlap between two selections, 1 indicates one sequence overlap, etc. -1 indicates the same selection. In these selections, there is some observed sequence overlap (1-2 sequences). This information can be incorporated into the final candidate selection. The candidate aptamers for PPC-CD can be selected to have no overlap with the other control targets (scaffold, beads, bradykinin) but allow for the selection of candidates that can recognize the PPC and PPCD switch as they can recognize the PPC on the N-terminus.

[0157] Figure 45is the result of single point binding assay of 10 potential aptamer candidates. The binding indicated by fluorescence signal (y-axis) was measured for 100 nM of 10 aptamers. Aptamer 4 showed higher binding to target PP-C compared to control (non-aptamer and buffer). Aptamer 1, 2, 3, 4, 7, 8, 9 showed higher binding to PP-D compared to control. Data was normalized to positive control (FAM directly conjugated to beads).

[0158] Figure 46A and 46B are binding curves for aptamer 1 and 4, respectively. Aptamer 1 (panel A) shows an elevated signal for PP-D, much higher than for PP-C. It appears saturated for PP-C, but not for PP-D, indicating non-specific binding. Aptamer 4 (panel B) shows saturated binding for PP-D and no binding for PP-C.

[0159] Figure 47 is an example of an electropherogram from an Agilent bioanalyzer assay with the desired peak shape at 60 seconds, indicating that the PCR product was properly digested into ssDNA.

[0160] Figure 48 is an example of an electropherogram from an Agilent bioanalyzer assay with the desired peak shape indicating that most of the product has the desired length (86 nt for the example described herein).

[0161] Figure 49 is a schematic of the BCS core sequencing unit.

[0162] Figure 50A is a heatmap reporting the read counts of barcodes added in each cycle at each position of the barcode construct for 12 cycles of barcoding, each barcode having an expected position on the barcode construct. In an ideal case, the barcode added in the nth cycle should be at the nth position on the barcode construct. In case of x failed ligation or no aptamer binding events, the barcode will be observed in the (n-x)th position. The results confirm that it is possible to achieve the tandem ligation of 12 barcodes in the expected positions. Note that the barcodes used in cycles 1-6 are repeated in the same order in cycles 7-12, and the results are not demultiplexed; therefore, the small fraction of counts from each boxed number of the expected cycles 1-6 can be attributed to the grid to the right of the five squares (annotated with *), meaning that for these sequences, at least after cycle 6, no barcodes were not ligated.

[0163] Figure 50Bis an arrow diagram depicting that in 3 rounds of ligation mediated by a universal bridge, 3 barcodes were successfully ligated into a row, demonstrating that it is possible to achieve serial ligation using a universal bridge.

[0164] Figure 51 is a heatmap reporting the sequencing of each target substrate using the aptamer barcodes ligated to it. Figure 51A Total counts (SEQ ID NOs: 204-243) are reported, while Figure 51B Normalized percentages (SEQ ID NOs: 244-279) are reported. Arginine vasopressin aptamers identified by RCTH-SELEX (highlighted in red) showed specificity over the bradykinin target and the peptide target with DD N-terminus (DD target) as their barcodes were ligated to all types of arginine vasopressin substrates but rarely or not to the empty control, bradykinin, and DD target substrates.

[0165] Figure 52 is a fluorescent image of a flow cell with bradykinin surface attached before Edman degradation and after two rounds of Edman degradation. The flow cell was probed with fluorescent bradykinin antibody and imaged through the 555 channel. The reduced but not eliminated signal indicates reduced antibody binding, which can indicate that the peptide was partially degraded but still remained attached to the flow cell surface.

[0166] Figure 53A is a 100% stacked histogram depicting the distribution of RNA decoys complementary to 5 different sequences (9, 13, 11, 12, 19) produced from the original pool of sequences 9 at 0.000125% by weight, 13 at 0.01%, 11 at 0.1%, 12 at 10%, and 10 at 89% and various concentrations of in vitro transcriptase (IVT). The change in frequency of RNA decoy sequences indicates that treatment with different concentrations of IVT can produce different ratios of RNA decoy sequences.

[0167] Figure 53B is a table reporting the percentage of counts of each RNA decoy sequence produced using various concentrations of IVT.

[0168] Figure 54 is an image of an electrophoretic mobility shift assay (EMSA) gel demonstrating that Spot-tag nanobodies were coupled to oligonucleotides (VHH-oligo). The first 4 lanes of the gel show the electrophoretic mobility of the Spot-tag nanobody itself uncoupled. In the subsequent lanes, multiple bands of higher molecular weight are observed on the gel, presumably corresponding to multiple oligonucleotides coupled to a single nanobody.

[0169] Figure 55is a schematic of the complete core sequencing unit construct for each target and their corresponding structure on the sequencing chip after ligation and formamide wash. DNA targets serve as positive controls. 5' Phos. Ol control is used for noise associated with the complete oligonucleotide tail ligated to all peptide targets, while CLR. Null. Block. Br control is used for noise associated with sequencing chip components.

[0170] Figure 56 is a heatmap reporting the case for each target substrate sequenced with binder barcodes ligated to the substrate when Spot-tag nanobodies are coupled to oligonucleotides. Control runs were performed in triplicate with each parallel having a different associated barcode, and DNA and Spot-tag experiments were run using 6 experimental replicates. DNA controls (Kd pM) bind and label complementary oligonucleotides with high fidelity (in terms of sequencing counts), and Spot-tag nanobodies bind and label Spot-tag peptides (Kd 6 nM) with strong fidelity. Differences in sequencing counts between experimental replicates are attributed to differences in barcodes used for each replicate. Effects on barcode sequence were screened and analyzed to derive a set of barcodes for use in downstream experiments. No known variables (GC content, sequential base pairs, etc.) were found to correlate with barcode effects on sequencing noise beyond target type (DNA vs. nanobody, etc.). Experiments were repeated and confirmed, demonstrating that this protocol can be used for DNA:DNA binding systems and peptide:nanobody binding systems.

[0171] Figure 57 is a heatmap reporting the case for each target substrate sequenced with binder barcodes ligated to the substrate when Spot-tag nanobodies are not coupled to oligonucleotides. Experiments were run in triplicate with each parallel having a different associated barcode. Differences in sequencing counts between experimental replicates are attributed to differences in barcodes used for each replicate. Effects on barcode sequence were screened and analyzed to derive a set of barcodes for use in downstream experiments. No known variables (GC content, sequential base pairs, etc.) were found to correlate with barcode effects on sequencing noise beyond target type (DNA vs. nanobody, etc.). For this experiment, only DNA binder AV.B4.U2.SA4.2 and its corresponding target (SP9) had high sequencing counts. Experiments were repeated and confirmed, demonstrating that this protocol can be used for DNA:DNA binding systems and peptide:nanobody binding systems.

[0172] Figure 58 is an implementation of the results and computational deconvolution process for single molecule peptides from imaging to peptide identification. Figure 58 A is an implementation of a series of images produced by 4 iterations of probe incubation for a single peptide molecule at a location (X,Y) on a chip.Figure 58 B is a table reporting the fluorescence signal observed through each channel (350, 433, 532, 555, 647) reflecting the results of A. Color regions indicate signals above the noise threshold, which together make up the optical signature of the bound aptamer. Figure 58 B is a table reporting the fluorescence signal observed through each channel (350, 433, 532, 555, 647) reflecting the results of A. Color regions indicate signals above the noise threshold, which together make up the optical signature of the bound aptamer. Figure 58 C is an implementation of a lookup table matching each aptamer identity to the optical signature observed through multiple iterations. Figure 58 D is an implementation of a series of aptamers observed at a location (X, Y) on the chip calculated from 8 rounds of aptamer incubation. Overlapping N-terminal amino acid calls from the double amino acid identity redundancy scheme are indicated in black, while contentious calls are indicated in red. Figure 58 E is a schematic of the sequence calling strategy, where the calculated sequences produced by the peptide sequencing method described herein are matched to a database of known peptides or a reference proteome.

[0173] Figure 59 are magnified 20x, 60x, and lOOx images of fluorescent beads-streptavidin conjugates on a glass slide (single molecule control) and a single oligonucleotide bound to a sequencing chip. The similarity in the size of the observed spots between the fluorescent beads on the chip and the sequencing chip indicates that the observed spots on the sequencing chip are single molecules.

[0174] Figure 60A are fluorescence images of fluorescent beads-streptavidin conjugates on a sequencing chip and intensity measurements after background subtraction using a local threshold. The threshold is the median intensity of the local neighborhood (30 x 30 pixels) of a pixel.

[0175] Figure 60B is the thresholded intensity distribution of all fluorescent spots in Figure 60A

[0176] Figure 61 is a heatmap reporting the selectivity performance of multiplexing. In 5 target (GNRH, NC2, NC3, Tl, Angiopep) assays, the aptamers were first subjected to an abundance filter (at least 12 reads) and the top 5 sequences for each target were ranked on the basis of selectivity (number of reads for the desired target / number of reads for all targets). Off-target hits are shown and selectivity is emphasized with a color gradient from red (low specificity) to blue (greater specificity). For each target, the top 5 target-specific aptamers exhibited a selectivity of 0.500 to 0.923, indicating that at least half of each aptamer's reads were bound to its target of interest. In comparison, no more than 25.0% of the same aptamer's reads were bound to any single non-target of interest.

[0177] Figure 62 ​Peptide target sequences used in the multiplexing experiment (SEQ ID NOs: 280-285) are shown.

[0178] Figure 63 is an image of a SDS-PAGE gel showing denatured peptides purified using an anti-His antibody affinity pull-down assay have the expected size of dMS-EmGFP and dMS2, indicating that both dMS-EmGFP and dMS2 were expressed. BSA was included as a standard.

[0179] Figure 64 is an image of an electrophoretic mobility shift assay (EMSA) gel confirming that the dMS2-EmGFP fusion protein binds to 2 nM RNA containing MS2 coat protein binding sites (protein concentration used in the binding reaction is labeled at the top, nM).

[0180] Figure 65 is an image of an EMSA gel showing that the dMS2 protein (without EmGFP) binds to ~2 nM RNA containing MS2 coat protein binding sites (protein concentration used in the binding reaction is labeled at the top, nM), verifying the identity of the dMS2 protein.

[0181] Figure 66 is a violin plot showing the percent of sequences from each experiment that were desired full-length constructs obtained using 10mer dsDNA blocks with 1 base pair overhang, one of which achieved 78.9% efficiency.

[0182] Figure 67 Reports the percent of unique sequences produced in LEGO experiment 87P from Figure 66 LEGO, where 78.9% of the constructs were sequences of LEGO blocks with the desired length, order, and orientation. DETAILED DESCRIPTION

[0183] The present disclosure describes methods and compositions forming a pipeline for developing and using a protein sequencing platform that utilizes aptamers that specifically bind to N-terminal amino acids Figure 1). The protein sequencing methods described herein rely primarily on aptamers with various different characteristics depending on the specific application. For example, amino acid specific aptamers can be generated using the new methods described herein (RCHT-SELEX or NTAA-SELEX). Such amino acid specific aptamers can be used to identify, characterize and convert 1-2 amino acid residues of a protein or peptide into a DNA sequence (PROSEQ) by a region with a nucleic acid barcode, or such amino acid specific aptamers can be generated and used to identify and characterize each amino acid of a protein or peptide on the basis of a visual signal (PROSEQ-VIS). Furthermore, many target specific aptamers can be generated simultaneously and used to generate and screen a large number of binders (multiplexing). Simultaneous and specific aptamer selection relies on reliable identification of the target. Generation of targets with nucleic acid barcodes can be achieved in vivo by exploiting non-covalent bonds between peptides or proteins and their corresponding recognition sequences using RNA binding proteins (TURDUCKEN). Finally, successful SELEX experiments require the inclusion of aptamers with certain specific binding preferences and affinities for the molecular target in the original pool of 10 14 -10 15 candidate sequences that is only a small fraction of all possible DNA sequences. Machine learning (ML) can help to optimize experimental seed binders, so that unlike conventional SELEX experiments, the optimal binder does not necessarily exist in the experimental dataset. The ability to construct computationally derived, customizable DNA libraries to perform SELEX screening using controlled input pools can significantly customize the exploration space by systematically analyzing aptamer candidates containing sequences with known effective binding properties (LEGO).

[0184] Aptamer

[0185] Aptamers are short single-stranded nucleic acid strands, which can be composed of RNA, DNA, modified nucleic acids, or other synthetic nucleic acid analogs, that fold into unique conformations that allow for the acquisition of binding specificity to biological targets such as proteins and peptides (Mckeague & Derosa, 2012). Aptamers are used to interrogate binding interactions involving molecular targets in a large number of research areas including drug development, diagnostics, imaging, and basic science. Specifically, aptamers bind to targets with high specificity and affinity, can be produced and modified more rapidly and at lower cost than antibodies, have a wider potential target range than antibodies (Zhou & Rossi, 2016), and are less likely to elicit immunological side effects than antibodies (Bouchard, Hutabarat, & Thompson, 2010). However, aptamers have not achieved widespread success in clinical or industrial use, primarily due to the difficulty of discovering and identifying aptamers with the desired binding properties (Zhou & Rossi, 2016). Furthermore, aptamers discovered in isolation (i.e., selected against a purified target) exhibit high binding affinity under experimental conditions, but fail to bind to its target of interest under in vivo conditions (Chen et al., 2016). The present disclosure provides methods of manufacturing and using aptamers with very specific binding properties for amino acid residues at the N-terminal end of a peptide chain.

[0186] Aptamers with high peptide binding affinity have a higher chance of binding and generating a record of a binding event than aptamers with lower binding affinity. Specific aptamers bind to only a small number of possible peptides and thus generate a record of information about which molecules are present. Therefore, aptamers with high affinity (K d Aptamers with high affinity (K dand selective aptamers enable us to relatively easily accurately quantify mixtures of known proteins. For non-de novo applications, PROSEQ and PROSEQ-VIZ technology can use proteomic maps to resolve any resolution gaps in the data. Furthermore, subsequent cycles can be repeated before removal of amino acids to allow for additional bits of information to be obtained before cleavage. Finally, high specificity aptamers are not necessary even for de novo sequencing if PROSEQ and PROSEQ-VIZ are limited to aptamers that selectively bind to N-terminal dipeptide prefixes. Noise from reduced specificity is compensated by the additional observed binding events that result from the double amino acid identity redundancy scheme, as it allows for each amino acid to be observed twice (except for the first N-terminal amino acid) to confirm its identity Figure 2 ) Each dipeptide aptamer binding event provides interpretation of the identity of two N-terminal amino acids, while each round of degradation removes only one amino acid, allowing each amino acid to be exposed to two rounds of aptamer binding except for the original N-terminal and C-terminal amino acids, which are only read once. In the case of amino acid misidentification, downstream computational algorithms can be used to correct or detect inaccurate read bit results with a certain level of confidence.

[0187] Robust and compressed high-throughput - Ligand System Evolution by Exponential Enrichment (RCHT-SELEX) and N-terminal amino acid SELEX (NTAA-SELEX) Figure 3A

[0188] Systematic evolution of ligands by exponential enrichment (SELEX) is a known high-throughput screening (HTS) process that has been used to identify aptamers that bind to specific target ligands in vitro selection (Tuerk & Gold, 1990). The conventional SELEX protocol generally involves screening a diverse and random oligonucleotide library against a single peptide or protein target, which includes flowing the aptamers over the bead-bound target and eliminating weakly binding aptamers through multiple rounds of selection in which weakly and non-binding aptamers are washed away (Blind & Blank, 2015).

[0189] The conventional SELEX method starts with the synthesis of about 1010 14 -10 15 unique sequences for an oligonucleotide library, followed by 10-20 iterative rounds of: a) incubating a single target with a random pool of candidate aptamer sequences to promote aptamer / target binding; b) separating the target-bound oligonucleotides from unbound sequences; and c) amplification and characterization of the bound aptamers Figure 3A Several variations of the original SELEX method have been developed, such as capillary electrophoresis SELEX (CE-SELEX), microfluidic SELEX, and CELL SELEX, to meet different research needs.

[0190] The goal of the conventional SELEX method is to improve the binding affinity of aptamers identified through experimental screening. The conventional SELEX method for identifying aptamers has two major issues that hinder large-scale screening:

[0191] • The conventional SELEX method relies on a repetitive screening process, in which experimental errors can compound in each subsequent round of screening. For example, in each round, aptamers undergo PCR amplification, DNA clean-up, and conversion from double-stranded to single-stranded DNA by separation or enzymatic digestion. Variability in one or more of these processes between iterations and / or experiments can contribute to biased selection of the pool of aptamers engineered to withstand the selection process for a particular experimental setup.

[0192] • The lack of parallel selection using the same input library for controls and parallel samples prevents (a) inter- and intra-experimental comparisons, respectively, (b) signal-to-noise ratio analysis, and (c) background truth measurement, all of which complicate the application of downstream computational analysis, data clean-up, and predictive modeling, such as ML. Models are attracted to the strongest signal, regardless of origin. In the case of biological experiments, there is usually operator error / noise, instrument noise, biological process noise, and noise caused by the handling of physical reagents (i.e., contamination), and the combination of all these different noise elements can often drown out the experimental signal. Therefore, models often make predictions based on very noisy signals, unless they are pre-trained for different noise elements. To this end, several different features (incubation, parallel samples, spiking controls, mock SELEX, etc.) were designed to calculate and remove noise during data processing before the model or train the model for the noise elements in the prediction phase. Furthermore, there are several classes of models that have limited predictive power outside of the linear range, and in biology, processes are often non-linear (e.g., PCR). The advantage of linear models is that they are well-studied, computationally inexpensive, and often provide reliable predictions. However, when applied to non-linear data sets, linear models often give incorrect predictions. On the other hand, non-linear modeling methods can be more computationally expensive and are also prone to overfitting (e.g., polynomial modeling of sparse data), but are often needed when linear models do not accurately describe the data set. Therefore, a large number of unit tests were run to calculate the regions of linear and non-linear processes in order to best determine which type of modeling method can be applied.

[0193] • The conventional SELEX method allows for screening of a pool of aptamers against only one peptide or protein target at a time. That is, each protein or peptide target must be screened in isolation in order to be able to identify the target. Therefore, screening against 1,000 peptide targets using the conventional SELEX method would require 1,000 independent SELEX experiments, each comprising multiple rounds of screening.

[0194] Furthermore, for example for 40-mer ssDNA oligonucleotides, 1010 24 possible oligonucleotides can be generated and the exploration of the total possible experimental space of 1010 12-15 can lead to difficulties in finding unique aptamers for a target. Currently, there are a number of obstacles to efficiently screening such a large number of candidates:

[0195] • Low hit rate: A successful SELEX experiment requires that an aptamer with high affinity for the molecular target is contained in the original pool of 10 14 - 10 15 candidate sequences. For 10 14 samples, only 8.27 x 10 -11 % of the experimental space of possible DNA sequences is explored, making even the most optimized experiments in practice have a high probability of failure.

[0196] • Excessively time consuming: It often takes more than 6 months to a year to identify specific aptamer candidates.

[0197] • Non-specific: Traditional SELEX experiments incubate candidates with one target at a time, which only demonstrates the relative affinity of the candidate aptamer and not their specificity in a competitive environment.

[0198] • Cannot be used for direct comparison: Since most experiments start with a new random pool of input oligonucleotides, direct comparison across experiments is not possible.

[0199] • Difficult to transfer to different environments from discovery conditions: The transfer of discovered aptamers can also be fraught with difficulties due to their sensitive structural properties related to the discovery environment. Since structure determines function, aptamers selected in a particular environment can not fold and bind to their target in the same way when conditions are different from the experimental conditions.

[0200] There are two significant gaps in the current SELEX protocol. None of the existing methods are tailored to accommodate large-scale computational analysis of multiple targets between each round with the goal of using experimental data to complement computationally derived aptamers. If a working protocol exists, the empirical dataset could be seamlessly integrated with a machine learning analysis and prediction pipeline, allowing for the use of computer predictions of aptamers for targets. Computationally predicted aptamers allow for the exploration of a wider range of sequences for the best aptamer target and also save resources and time in the aptamer search query. Furthermore, the SELEX protocol lacks the precision and resolution to find binders with high resolution for a small subset of larger targets and can be used as N-terminal amino acid binders. The methods developed to address these two gaps are described in detail below. A new SELEX method (referred to herein as RCHT-SELEX) is provided in Section A that optimizes the selection of high affinity and specific aptamers in a time-efficient manner through the creative combination of existing and new technologies to address the gap in developing an ML-compatible empirical dataset. In addition, another new SELEX method is provided in Section B that was developed to preferentially find N-terminal amino acid specific binders (referred to herein as NTAA-SELEX).

[0201] Section A: RCHT-SELEX

[0202] Figure 3 is a schematic showing how the conventional SELEX method (left) is modified to produce RCHT-SELEX (right). The main differences between the two technologies are highlighted below: Figure 3B Figure 5B • Step #1 of conventional SELEX does not amplify the input pool; RCHT-SELEX amplifies the input pool after the negative selection step and addition of the incorporator, such that: (a) there are approximately 100 copies of each aptamer binder present, and (b) the same input pool is used in each experiment.

[0203] • Step #2 of conventional SELEX is a single experiment of a single target incubated with the aptamer library; in contrast, RCHT-SELEX

[0204] • The incubated pool is split into triplicate samples run in parallel (including 3 experimental controls with beads only) in many experiments of several targets, and

[0205] • The aptamers are assayed against targets with alternating regions in different rounds, such that only the constant region that drives the selection process is the region for which the user wishes to find a specific binder, regardless of the adjacent region of the target.

[0206] • The incubated pool is split into triplicate samples run in parallel (including 3 experimental controls with beads only) in many experiments of several targets, and

[0207] ​• Step #4 of traditional SELEX sequences the evolved pool of aptamers after 8-20 rounds of repeating steps #2-#5, whereas step #4 of RCHT-SELEX includes sequencing after each round of selection as well as various techniques for maximizing and normalizing the amount of DNA input into the next round in each experiment.

[0208] • Step #5 of traditional SELEX includes obtaining aptamers that bound to the target protein in the previous round, so that by repeating steps #2, #3, #4, and #5 for 8-20 rounds, those aptamers can continue the selection process; RCHT-SELEX can be performed in just 4 rounds and the aptamers are again assayed against the target after the primer region is replaced with alternative primer sequences.

[0209] Since several experiments are run in parallel in RCHT-SELEX, and the goal is to reduce experimental bias between each experiment, several additional steps are added to the RCHT-SELEX protocol to support running >36 experiments simultaneously. RCHT-SELEX can include techniques such as:

[0210] • Thresholding the same amount of DNA as input for subsequent rounds to reduce PCR bias ("Threshold PCR")

[0211] • Optimizing PCR conditions for a specific pool of candidates ("PCR Optimization")

[0212] • Performing unit tests before each digestion to determine the optimal digestion conditions for each sample ("dsDNA Digestion").

[0213] Other changes to RCHT-SELEX can include:

[0214] • Using the same pool of aptamer candidates assayed against multiple targets pooled together in early rounds, and de-multiplexing by independently incubating the aptamers against each target in the last round ("Bead-based Multiplexing-SELEX")

[0215] • Alternating between targets with variable local environment binding regions between alternating rounds of RCHT-SELEX for experiments where the desired aptamer is one that specifically binds to a smaller portion of the molecule rather than the entire molecule ("Switching")

[0216] • Switching primers in the middle of the experiment to identify aptamers that are strong binders independent of the primer region ("Primer Switching").

[0217] Negative SELEX:

[0218] One technique that can be used to reduce the enrichment of aptamers against unwanted targets is to screen the initial pool of aptamer candidates for aptamers that bind to the selection component (e.g. beads, streptavidin) used in the SELEX experiment. Aptamers that exhibit binding affinity to the selection component are non-specific to the target and can be removed from the pool of candidates, such that only aptamers that do not bind to the selection component will be part of the pool of aptamer candidates analyzed against the target. See, e.g., Figure 4. Single incubations, double incubations, and intra-experimental replicates:

[0219] For example, a pool of 10 15 DNA aptamers is selected from an original pool of 10 12 DNA aptamers and amplified by 13 cycles of PCR using unmodified primers, resulting in approximately 2000 copies of each aptamer. Amplification is dependent on primer sequence and PCR conditions, and the incubation PCR protocol can be adjusted for each individual library. The goal is to have at least 100 copies of the majority of sequences present in each experiment, and at least 30 copies of each aptamer sequence present. Libraries are sequenced in the protocol optimization phase to help achieve approximately even amplification copy numbers between sequences.

[0220] After amplification, approximately 2000 copies of each aptamer are distributed into 12 samples, resulting in approximately 166 copies of each aptamer in each initial starting library pool. The process of having multiple copies of the same aptamer present before starting selection allows for direct comparison of results from the same initial incubation. Computationally, this feature allows for direct experimental replicates to be run in parallel, and also provides the ability for the trained model to move towards a particular target and away from another. Since determining accurate amplification of 10 12 sequences would take many sequencing runs, a single NextSeq run of 400 million reads can be performed as an approximation of the library amplification feature across the entire pool. Single incubations stop at this step.

[0221] For double incubations, a second incubation is performed by taking approximately 75 copies of each aptamer from the first incubation and amplifying it through 6 cycles using primers that are protected from phosphorylation, which allows for comparison of results from the same initial incubation between approximately 300 experiments (approximately 2000 copies of each aptamer from single incubations, selecting 75 aptamers gives 26 possible draws; each group of 75 aptamers will produce a double incubation pool for 12 experiments, so 12*26 = 312 total experiments; note that there can be some loss in purification, digestion, and other processes, and amplification yield is highly dependent on the nature of the primers and PCR master mix components). Amplification of aptamer candidates from each incubation also increases the likelihood that strong and moderate binders will make it through early rounds. See, e.g., Figure 5Awhich schematically demonstrates single and double incubations and experimental replicates described herein, and Figure 6 which schematically demonstrates that single and double incubations and experimental replicates can be used in the RCHT-SELEX method.

[0222] Bead-based multiplexing-SELEX

[0223] After, for example, four rounds of RCHT-SELEX using multiple bead-bound targets pooled together, the aptamers can be de-multiplexed in round 5 by incubating the pool of amplified aptamers with beads coupled to only one of the original targets, respectively (see, for example, Figure 7 ). Bead-based multiplexing-SELEX adds a competitive target environment and changes the number of targets that can be explored in the same experiment.

[0224] Peptide switching

[0225] In designing binders for protein sequencing, four goals must be achieved: (1) target a specific amino acid, (2) target the specific amino acid in the N-terminal position, (3) not bind to the same amino acid in non-N-terminal positions, and (4) robustly bind to the targeted N-terminal amino acid regardless of the neighboring amino acids. The rationale behind goal #4 is that the local biochemical environment (e.g. neighboring amino acids) can affect the binding activity of aptamers, decreasing their effective K d . Since the goal of protein sequencing is to establish binders that can be utilized in peptide strings across the entire proteome, the binder design must account for the influence of the local environment. To achieve goal #4, a changing local environment is introduced during binder selection to develop binders that are independent of the neighboring amino acids. This is done by fixing 1-2 amino acids in precise positions within the peptide string (typically the N-terminal position) and changing the amino acids connected or surrounding in different rounds. Figure 8 A method of identifying aptamers that specifically bind to the N-terminal amino acid prefix independent of the composition of the peptide suffix is shown. This technique, labeled "peptide switching", evolves aptamers in iterative rounds in which only the peptide suffix is changed while the desired N-terminal amino acid prefix remains the same, removing negative binders. Peptide switching experiments can also include empty, scrambled, or "dummy" targets to define promiscuous binders to eliminate false positives.

[0226] PCR optimization

[0227] PCR conditions can be optimized to maximize DNA output while minimizing unwanted products such as concatemers. PCR optimization must be performed for each library. In SELEX experiments, the initial library primer must be replaced frequently between experiments to prevent PCR contamination in experiments. For each library, master mix and PCR optimization unit testing is performed after each change in library primer, which includes adjusting as many parameters as possible (buffer conditions, cycle number, enzyme, primer concentration, number of protected base pairs, etc.), then the SELEX library can be used in experiments. Results are analyzed using sequencing, Qubit, TapeStation, Bioanalyzer, and digestion unit testing in order to select the ideal optimization settings for each library. For example, amplification can be performed in a 50 μΐ, reaction volume consisting of 38.49 μΐ, nuclease-free water, 0.30 μΐ, 1 mM forward primer complementary to the first 6 nucleotides (referred to as 6XP), 0.30 μΐ, 1 mM phosphorylated reverse primer (referred to as RP04), 0.50 μΐ, 2X Phusion® HF PCR Master Mix (New England Biolabs®), 0.01 μΐ, 100 mM MgCl2, 0.01 μΐ, 10 mM dNTP, 0.01 μΐ, 50 mM Betaine, 0.01 μΐ, 20 U / μΐ, Phusion® DNA Polymerase (New England Biolabs®), and 0.01 μΐ, 100 μΜ template. PCR can be performed using an Eppendorf Mastercycler nexus eco PCR machine. The thermocycling can be programmed for an initial denaturation at 95 °C for 5 minutes, followed by 13 cycles of 95 °C denaturation for 30 seconds, 55 °C annealing for 30 seconds, and 72 °C extension for 30 seconds, and finally 72 °C extension for 5 minutes. Annealing conditions are primer dependent and can be re-optimized for different primer sets used. II Fusion DNA Polymerase, 10 μΐ, Herc Buffer, 0.40 μΐ, 25 mM dNTP, and 0.01 μΐ, template. PCR can be performed using an Eppendorf Mastercycler nexus eco PCR machine. The thermocycling can be programmed for an initial denaturation at 95 °C for 5 minutes, followed by 13 cycles of 95 °C denaturation for 30 seconds, 55 °C annealing for 30 seconds, and 72 °C extension for 30 seconds, and finally 72 °C extension for 5 minutes. Annealing conditions are primer dependent and can be re-optimized for different primer sets used.

[0228] Digestion of dsDNA

[0229] Lambda exonuclease is a highly active exonucleic deoxyribonuclease that preferentially digests the 5-phosphorylated strand of dsDNA and has significantly lower activity on ssDNA and non-phosphorylated DNA (Little, 1967) (Mitsis & Kwagh, 1999). Lambda exonuclease can be used to efficiently digest PCR-amplified dsDNA into ssDNA in three steps: a) unit test for optimal digestion conditions, b) splitting the pre-digested library into three, and c) bioanalyzer quality control (QC) assay to test the amount of ssDNA versus dsDNA. Single stranded PCR products can be generated by first performing PCR using two different primers (e.g. a 3’-thiophosphate protected primer complementary to the unwanted reverse strand and a 5’-phosphorylated primer complementary to the desired forward strand) and then performing PCR amplification where the phosphorylated strand of the PCR product can then be removed by digestion with lambda exonuclease. The RNA kit of the bioanalyzer system can be repurposed to quantify ssDNA as the dye in the RNA kit also binds to ssDNA. Although the measurement output is not calibrated for ssDNA, inferences can be made from the bands and peaks. See, e.g. Figure 9 The RNA kit of the bioanalyzer system can be modified to quantify the amount of ssDNA versus dsDNA in a sample as the dye in the RNA kit binds to both ssDNA and dsDNA. When a sample containing both ssDNA and dsDNA is processed on the bioanalyzer by capillary electrophoresis, unique non-overlapping peaks are generated for ssDNA and dsDNA, where the relative area under each curve depicts the percentage of the sample attributed to ssDNA and dsDNA. The purpose of analyzing the RNA bioanalyzer kit in the digestion assay is to confirm that all dsDNA has been converted to ssDNA without over-digestion of the ssDNA library. Although the measurement output is not calibrated for ssDNA, inferences can be made from the bands and peaks about the nature of the DNA mixture.

[0230] During the experiment, the data indicated that the quality and quantity of the PCR output affected the ability to predict the lambda exonuclease digestion behavior. Libraries containing extra concatemer products digested very slowly or very rapidly depending on the fraction of protected or phosphorylated base pairs present in the concatemer sequence. Therefore, unit tests can be performed when evaluating new libraries to prevent complete digestion of the sample. Unit tests can be performed to determine the optimal reaction time for high efficient ssDNA production for each sample before digesting all PCR products. A small sample of purified PCR product can be incubated at 37°C for e.g. 2, 5, 10, 15 or 20 minutes, incubated at 75°C for 10 minutes, and kept at 4°C afterwards, and subjected to a time course analysis of lambda exonuclease digestion. An RNA Bioanalyzer can be run for each sample to evaluate the digestion and determine the optimal digestion conditions to be applied to the rest of the PCR product samples.

[0231] Lambda exonuclease digestion of the whole sample can be performed as follows: incubate the optimal time determined by the time course analysis at 37°C, then heat inactivate the enzyme at 75°C for 10 min, and keep at 4°C.

[0232] A representative sample of the final lambda exonuclease digestion mixture can be run on another RNA Bioanalyzer chip to ensure that the PCR products were sufficiently digested to ssDNA Figure 10A ) before the next round of RCHT-SELEX. If the digestion is not complete, more lambda nuclease and ATP can be incorporated.

[0233] Additional controls: Bead controls, spiking and pseudo-SELEX

[0234] • Spiking oligos: A small amount of spiking of known aptamer mimics can be added at various steps throughout the RCHT-SELEX as a control to detect experimental errors. For example, a mixture of 9 oligos with 3 representative sequences and 3 different GC content levels (e.g. 40%, 50%, 60%) of known sequence (i.e. a “9-oligo mixture”) can be added prior to PCR to provide relevant information on sample variability related to PCR differences. See e.g. Figure 10B . Alternatively or in addition, a known sequence (e.g. a positional spike) can be added to each well to provide information on the spatial position on e.g. a 96-well plate. See e.g. Figure 10B .

[0235] • Full bead control: The full bead control includes parallel and sequential controls. See e.g. Figure 10CFor parallel experiments, a full-bead control (e.g. bead sample with no coupled peptide) can be run in triplicate with the experiment to determine the amount of aptamer that binds to beads only from the incubation pool. In addition, these controls can be used to determine the level of inter-well contamination or noise from each experiment. Sequential bead controls can be used after each round of RCHT-SELEX, where aptamer bound to beads coupled with peptide is incubated with beads that are not coupled to peptide. If desired, aptamer bound to empty beads can be sequenced to identify common sequences in the aptamer bound to empty beads.

[0236] • False SELEX: Prior to each round of RCHT-SELEX, a small sample of the original input can be removed and kept at room temperature as a control to determine the effect of PCR bias, as there is no target present. See, e.g., Figure 11A .

[0237] Threshold PCR

[0238] Bound aptamer from a bead-based RCHT-SELEX experiment can be amplified directly on the magnetic beads. Therefore, there is no need to denature the aptamer from the beads prior to running the PCR, limiting the number of steps of handling, manipulation, and potential library loss in the sensitive phase of the SELEX assay (Hoon, Zhou, Janda, Brenner, & Scolnick, 2011). However, the PCR reaction can reach a saturation point, where reagents become limited or concentrations become too high for uniform replication to continue. Since the concentration of bound aptamer prior to PCR amplification is unknown and can only be estimated, it is not possible to precisely determine how many amplification cycles are needed before amplification saturation occurs. In addition, PCR amplification can be affected by some of the magnetic beads coated with bovine serum albumin (BSA), where if the concentration of BSA is too high, the total product produced by PCR is reduced. Furthermore, internal experiments have shown that there is an uneven distribution of aptamer between beads, such that if the aptamer library on the beads is physically split into separate solutions prior to amplification, one would observe different endpoint amounts and variances of unwanted PCR product between the splits, resulting in an unknown introduced inter-sample variance. To (a) address the complexity of introducing unquantifiable inter-sample bias, (b) amplify each library to the same concentration endpoint, and (c) mitigate the problems caused by PCR saturation and the presence of BSA, PCR amplification is performed in two stages: (1) PCR on the beads and (2) threshold PCR. If problems occur with digestion, performing PCR amplification in two stages provides the benefit of library redundancy.

[0239] When running many experiments in parallel from the same incubation pool, the PCR reaction can produce a mixture of aptamers with different end concentrations (e.g. low, medium, and high) depending on the amount of DNA pulled down in each experiment (Figure 11). To allow for computational comparisons between many experiments, and to balance the experimental requirements of a minimum amount of material that can be manipulated automatically (e.g. minimum pipetting volume for magnetic beads), the amount of input library is normalized prior to the second amplification step. Differences in the amount of input DNA template can affect the effect of PCR bias. The DNA concentration of each library can be measured after the PCR on beads, and the post-PCR library with the lowest DNA concentration or a standard amount can be used as a threshold amount standard. The remaining samples are then adjusted to the threshold amount and subjected to PCR for the subsequent round, which then produces the input for the subsequent RCHT-SELEX round. See, e.g. Figure 11B A large number of control experiments confirmed that using this threshold PCR method, the shape of the sequence distribution does not change Figure 12 and 11C .

[0240] Primer Switching

[0241] The construct of aptamer candidates can include a) random sequences of DNA that participate in or facilitate binding to the target, and b) one or more regions to which DNA primers can hybridize so that the aptamer sequence can be PCR amplified. The primer regions can contribute to the structure of the aptamer and binding affinity to the target molecule. The primer regions can be alternated with different primer sequences or removed altogether, and the aptamer can be assayed again to isolate aptamers that have high affinity to the target molecule independent of the primer regions. See, e.g. Protein or peptide sequencing (PROSEQ) .

[0242] Sequencing the aptamer pool after each round

[0243] A representative aliquot of the dsDNA from each round of selection is sequenced prior to threshold PCR and the sequence enrichment is analyzed round to round. Unit tests of sequencing before and after threshold PCR confirm that the distribution of sequences does not change during threshold PCR. Since the sequence distribution does not change, and a direct comparison point at each stage of SELEX is ideal for computational analysis, the stage prior to threshold PCR is chosen to: (1) reduce the additional steps at the end of the SELEX experiment, and (2) allow for the storage of DNA samples at higher concentrations and reduced volumes without additional manipulations (i.e. SpeedVac, etc.).

[0244] As discussed herein, the RCHT-SELEX method incorporates several new improvements: (1) simultaneous screening of up to 300 different targets, (2) maintenance of high DNA concentrations between selection rounds and reduced PCR bias, (3) additional features for advanced post-hoc computational analysis including comparisons between every possible experiment regardless of day of performance, and (4) improved binding specificity to small molecule targets such as small peptides or amino acid targets. These capabilities can accelerate the large-scale identification of aptamers to biological targets with potential uses in diagnostics, therapeutics, and basic scientific research. The new features of the RCHT-SELEX method described herein include, but are not limited to:

[0245] • Single or double incubations allow for direct comparisons between results from targets, experiments, and / or parallel samples from the same initial incubation;

[0246] • Analysis of within-experiment parallel samples reinforces positive signals and saves time and money testing unwanted aptamer candidates;

[0247] • Threshold PCR produces robust aptamer library input for multiple parallel experiments with minimized PCR bias, provides an earlier recovery point if there are experimental issues associated with converting the post-PCR dsDNA library to a ssDNA library, and reduces library loss from concatemers;

[0248] • Switching allows detection of aptamers specific for desired sequences at particular locations on the target (e.g. small fragments of larger molecules);

[0249] • Bead-based multiplexed SELEX increases the number of targets within the same experiment and reveals the binding capabilities of aptamers in a competitive environment;

[0250] • Incorporation of a control concentration can be used to detect experimental errors and PCR bias;

[0251] • The combination of next-generation sequencing (NGS) performed at each round with a sensitivity analysis can: (a) locate binders earlier, and (b) produce input data for machine learning (ML) models. ML models can predict high-specificity new aptamers with fewer SELEX rounds and explore a larger DNA input space than is possible with experiments. The use of ML in aptamer prediction can improve the efficacy of the SELEX method described herein while saving valuable research dollars and time.

[0252] The RCHT-SELEX method described herein reduces labor and reagent costs, while more importantly improving data quality, downstream analysis, and broadening the screening capabilities. Furthermore, the multiplexing method described herein can yield aptamers that specifically bind to targets in environments with multiple available targets (e.g., cell surfaces, human blood), thus greatly increasing the discovery of the application pipeline of aptamers.

[0253] The RCHT-SELEX method described herein can be used to examine binding of substances other than DNA:peptide interactions. For example, binding between a large number of biological targets can be examined, so long as both targets include oligonucleotides that can be linked to one another. For example, similar techniques can be used to screen for RNA aptamers that bind small molecule targets or protein complexes.

[0254] Additionally, many procedural modifications can be made to adapt this method to different applications. For example, but not limited to, other "input" nucleic acids such as RNA or modified nucleic acid bases can be screened for binding affinity to a molecular target of interest, or aptamers can be screened for binding to targets other than proteins or peptides (e.g., small molecules, intact proteins, other nucleic acids, specific cell lines). Another example of a modification is to replace lambda exonuclease dsDNA digestion with asymmetric PCR to generate ssDNA input into subsequent rounds of SELEX.

[0255] The RCHT-SELEX method described herein can be used to screen for aptamers that have selective binding to a specific peptide target in a competitive, multi-peptide environment. Like selective antibodies, the resulting aptamers can be used individually or in combination with two or more aptamers to create a complex that exhibits a multi-target binding profile. For example, two aptamers that each have high selectivity for different targets can be used sequentially, in tandem, or linked together in order to create a single construct that binds to two independent targets. Alternatively, two aptamers that have the same primary target but different off-target binding profiles can be linked together to increase the selectivity of binding to their common target through avidity, and simultaneously reduce off-target effects.

[0256] In addition to being used to measure binding between an aptamer and a target, the RCHT-SELEX method described herein can also be used to measure binding between different mixtures of any of the previously described classes of molecules (e.g., by replacing the aptamer with a molecule that has a DNA barcode and a 3' C-overhang arm), enabling bidirectional multiplexed competitive measurements of any of the combinations of molecular classes including, but not limited to, peptide to protein, protein-protein, antibody-protein, small molecule-protein, peptide-cell surface marker, antibody-cell surface marker, etc. In certain embodiments, two binding molecules (e.g., a binder and a target) can be drawn from a mixture of molecules from any of the above classes, allowing for the measurement of cross-binding in a complex, competitive environment.

[0257] Part B: NTAA-SELEX

[0258] We have developed a new SELEX method through the creative combination of existing and new technologies to optimize the selection of high affinity and specific aptamers in a time-efficient manner:

[0259] 1) Negative selection

[0260] One common technique to reduce the enrichment of aptamers to unwanted targets (e.g. magnetic beads, PEG, reagents in the binding buffer (e.g. BSA, etc.)) is to screen the initial pool of aptamer candidates for binding to the selection component used in the SELEX experiment, which in our case is streptavidin beads in SELEX buffer (1x PBS, 0.025% Tween-20, 0.1 mg / mL BSA, 1 mM MgCl2). Aptamers that exhibit binding affinity to the selection component are non-specific to the target and are removed from the pool of candidates, such that only aptamers that do not bind the selection component are part of the pool of aptamer candidates for the assay against the target. The library can be subjected to a single or multiple rounds of negative selection prior to the start of the SELEX round. In selecting the library size (e.g. 10 14 For negative selection, a larger library needs to be used to ensure that the supernatant contains enough molecules for downstream SELEX experiments.

[0261] 2) Peptide backbone switching

[0262] During each parallel selection, each parallel sample of the target of interest can be subjected to peptide switching. Specifically, a "switched" target can be developed that has a different backbone sequence, e.g. the amino acid sequence of the peptide target differs except for, e.g., two amino acids at the N-terminus. By switching between at least two different backbones in alternating rounds, the chance of enriching aptamers that bind to any species that is not the dipeptide of interest is reduced.

[0263] 3) Screening of multiple parallel targets

[0264] In this technology, parallel selection of DNA aptamers can be used for closely related as well as unrelated targets. The following indicators can be used between targets: 1) the count of each aptamer in each round determined by NGS sequencing, 2) the enrichment of each aptamer round to round, and 3) the enrichment from the first round of sequencing to the last round of sequencing. By comparing these indicators between different target selections, one is able to determine what binding signal looks like for "true binders" that bind to known targets that have been shown to be "aptamerogenic" previously, and what binding signal looks like for "non-specific binders" that non-specifically bind to the surface (e.g. beads) on which the targets are immobilized. These indicators between parallel target selections allow for tracking of the specificity of the aptamers and protection from unknown contamination.

[0265] 4) Parallel sample target selection

[0266] In this technology, parallel selection of DNA aptamers can be used for the same target. Unique random DNA libraries can be used to perform 2 or 3 SELEXs on the same target at the same time. This allows the experimenter to establish confidence in the indicators for each of the aptamers above, especially if they are of the same order of magnitude. In addition, it allows the experimenter to observe if there are outliers in the pool of aptamers. For example, if one random library has significantly lower enrichment than another random library when looking for final aptamer candidates, the experimenter can choose to work with only the lead aptamer candidates from the library that showed higher enrichment.

[0267] 5) Reverse SELEX

[0268] Reverse SELEX is a technique similar to negative selection, with the difference that the aptamer library is incubated with a molecule similar to the desired target on beads, the beads are pulled down with a magnet, and the resulting supernatant contains the library of aptamers that did not bind to the similar target. The supernatant can then be used in downstream experiments to help enrich for N-terminal binders. Reverse SELEX can be performed in parallel or sequentially with negative selection at the beginning of the experiment, and can be run in a single or multiple cycles. Reverse SELEX can be run between rounds of regular SELEX or after the last round of SELEX to enhance the signal of N-terminal aptamer binders in the library pool.

[0269] Many types of molecules can be used during reverse SELEX. Reverse SELEX can be used for targets that are similar in nature to the target but have slight modifications (e.g. to distinguish post-translationally modified N-terminal amino acids from unmodified N-terminal amino acids), the peptide backbone (or suffix) used during peptide switch, or large pools of targets representing the proteome to ensure specific N-terminal aptamer binders to unique target of interest.

[0270] If multiple backbones are used in the peptide switch experiments, multiple peptide suffixes can be used sequentially during the reverse SELEX experiments. For example, if two different backbones are used for the peptide switch, parallel reverse SELEX can be run on a mixture of targets between SELEX rounds, where the "target" pool for reverse SELEX consists of one half of the backbone bound to the beads and one half of the other backbone bound to the beads. Other embodiments can vary the stringency and / or introduce combinations of other molecules such as random peptide libraries, various different backbone designs, backbones with other N-terminal dipeptide suffixes.

[0271] 6) PCR and Digestion Techniques

[0272] PCR optimization, threshold PCR, and dsDNA digestion techniques can be used for NTAA-SELEX and are described in Section A: RCHT-SELEX.

[0273] New features of the NTAA-SELEX method described herein include, but are not limited to:

[0274] 1) This protocol provides a pathway to discover aptamer binders to N-terminal amino acids, which can revolutionize the method to enable high resolution identification of protein sequences and high throughput protein sequencing analysis. The stability and flexibility of nucleic acids make aptamers a versatile tool for protein sequencing and quantification techniques, including imaging and the DNA barcode method described herein;

[0275] 2) Multiple parallel SELEX experiments can allow for scaled aptamer discovery and removal of aptamers that are non-specific binders to multiple peptide targets;

[0276] 3) Sequencing of reverse SELEX experiments can facilitate discovery of N-terminal binders and removal of aptamer binders to other regions along the target;

[0277] 4) Control targets can be run in each SELEX experiment to allow for assessment of inter-experimental comparison metrics;

[0278] 5) Peptide backbone switching allows for detection of aptamers specific to N-terminal amino acids of larger peptides, or if desired, aptamers to amino acid sequences or modified amino acids internal to the peptide string.

[0279] Figure 13

[0280] The PROSEQ method described herein uses barcoded amino acid specific aptamers to convert protein sequences into DNA signals that are readable on next generation sequencing (NGS) platforms. Mass spectrometry (MS) is one of the commonly used tools in the identification and quantification of proteins, however the technique lacks the ability to cover a wide dynamic range required to detect low expressing proteins in complex samples (Schiess, Wollscheid, & Aebersold, 2008). Other existing specific protein quantification assays include antibody or aptamer binding assays, where a detectable antibody, aptamer, or other small molecule binder binds specifically to a known protein, thus not enabling de novo sequencing or measuring proteins for which no specific binder has been discovered. The PROSEQ protein sequencing method described herein can be used for small input of samples, including single cells or small volumes of blood, to identify the entire proteome, including low expressing proteins and single amino acid mutations, to better understand diseases caused by aberrant or mutated proteins. Furthermore, the PROSEQ method described herein is capable of sequencing heterogeneous samples or multiple samples in parallel, as proteins can be barcoded with unique DNA tags that can be incorporated into the DNA sequences that encode the protein sequence information. Furthermore, the method described herein is capable of sequencing significantly deeper than existing methods such as mass spectrometry, as DNA sequences are derived from single peptides, amplified, and read out from a sequencer (DR 100-10 9 ), which does not suffer from the same dynamic range limitations as mass spectrometry (DR > 10 5 )(Yates, Ruse, & Nakorchevsky, 2009). Furthermore, samples can be processed to remove reads associated with high abundance proteins in the sample by either 1) removing high abundance proteins in the raw input pool that enters PROSEQ, or 2) isolating DNA barcodes associated with high abundance proteins to increase the NGS readout count of DNA sequences associated with low abundance proteins.

[0281] The PROSEQ method described herein can be used in a clinical setting to quantify protein expression levels or identify new protein fusions or mutations associated with disease from individual patient samples to aid in patient diagnosis and disease onset. Furthermore, the method described herein can be widely used in the research fields of molecular and cellular biology and protein engineering, for example: sequencing proteins, discovering new biomarkers, profiling the entire proteome or metaproteome, evaluating mechanisms related to protein abundance, etc.

[0282] 1) Aptamers provide the ability to perform de novo sequencing.

[0283] The methods described herein rely on a library of aptamers specific to a unique combination of one or two N-terminal amino acids, where each residue or pair of residues has at least one or more possible aptamer binders. The ssDNA aptamers are designed to contain a 5' phosphate for ligation, a unique DNA barcode that indicates the identity of the particular aptamer and the corresponding cycle number, a spacer / consensus region for subsequent barcode ligation (e.g. ligation consensus sequence), a restriction enzyme site with spacer, and an amino acid recognition sequence (e.g. single stranded DNA aptamer sequence). See, e.g. Figure 13 These aptamers can be incubated with the peptide target with or without complementary DNA strands that cover some or all of the barcode sequences, ligation consensus sequence, and restriction enzyme site with spacer. In cases where these regions are not covered, DNA complementary to the ligation and restriction site can hybridize after incubation to facilitate ligation and restriction, respectively.

[0284] The aptamers described herein can be used to sequence proteins or peptides in any of the following ways:

[0285] (A) Peptide fragments from proteins processed in solution or on a solid substrate

[0286] Proteins can be obtained from a sample (e.g. a blood sample, cell lysate, or single cell), denatured, coupled to oligonucleotides, and digested into peptide fragments. It will be appreciated that there are multiple ways to obtain and digest proteins and couple peptide fragments to oligonucleotides prior to sequencing steps. One such strategy includes denaturing proteins using a mild surfactant and reducing and alkylating the denatured proteins to protect cysteine side chains. For example, the amino group on lysine amino acid side chains is reacted with aldehyde-modified oligonucleotides through a reductive amination reaction using sodium cyanoborohydride. Proteins can be digested with Lys-C, which cleaves proteins on the C-terminal side of lysines. By using this method, each digested peptide has a lysine residue attached to the tail of the oligonucleotide. The reductive amination reaction can also occur between the side chain of lysine and an alkyne with an aldehyde functional group, which is prepared for a click chemistry reaction with an azide-modified DNA oligonucleotide. In another method, the side chains of proteins can be protected, modified with oligonucleotides or click chemistry linkers, and then cleaved into peptide fragments, such as using a conventional trypsin method that cleaves at lysines and arginines and / or other fragmenting enzymes that cleave at random amino acid sites Figure 13 , step 2), or they can be processed in solution (see modifications below). At this point, the DNA-coupled protein fragments can be ligated to DNA oligonucleotides on the sequencing substrate surface, where they will remain ligated throughout the DNA barcode encoding process and removed prior to DNA sequencing.

[0287] Aptamers can be obtained directly from SELEX experiments and applied to BCS assays through the generation of a BCS-compatible aptamer pool in which one of the SELEX primer regions is converted to a BCS handle. The aptamer region of the conjugate is sequenced and serves as the "barcode" for the conjugate. To generate a BCS-compatible aptamer pool, a pool of single-stranded aptamers is incubated with a bridge oligonucleotide that is partially complementary to the aptamer tail and partially complementary to a ligation region on the barcode sequence on the barcode base (single-stranded overhang shown in FIG. 14) prior to incubation of the peptide target with the aptamer to (a) promote binding of the aptamer tail to the barcode sequence, and (b) block the ssDNA region of the aptamer that is not involved in target binding from interfering with proper aptamer folding. The library of BCS-compatible aptamers hybridized to the bridge can be flowed over the peptide and incubated, allowing the appropriate aptamer to specifically bind to the N-terminal amino acid residue Figure 13 , step 3).

[0288] After aptamer binding, unbound aptamer is washed away, and the tail of the bound aptamer can be ligated to a second glass-immobilized DNA oligonucleotide co-localized with the peptide Figure 13 , step 4). A restriction enzyme site included at the distal end of the aptamer barcode can be used to cleave the remainder of the aptamer, leaving the DNA barcode attached to the adjacent oligonucleotide Figure 15A , step 5). The N-terminal amino acid can then be removed from the immobilized peptide using Edman degradation and / or aminopeptidase. In Edman degradation, once the new N-terminal amino acid is exposed, another pool of aptamers with unique DNA barcodes indicating the target recognition sequence and cycle number can be introduced, and another cycle of DNA barcode ligation can be performed. After repeating this series of steps several times, a chain of DNA barcodes can be established that indicates the aptamer binding order of the peptide, which can be read using conventional NGS techniques. Using this information, the amino acid sequence of the bound peptide can be obtained. In the case of aminopeptidase, more than one N-terminal amino acid can be cleaved at a time in a less controllable manner, which, while not conducive to de novo sequencing, can reveal information for non-de novo sequencing methods.

[0289] (B) Full-length proteins processed in solution

[0290] For full-length proteins, the protocol is similar to the above, but with some important differences. The following steps can be performed: (a) lysing cells (if the protein is obtained from cells), isolating or purifying, denaturing and protecting the protein, (b) protecting the reactive side chains of the amino acid residues (e.g., thiol, carboxyl, and amine groups), (c) coupling ssDNA oligos to the C-terminus of the protein, where the ssDNA oligos contain a primer region, a unique barcode, and an initial ligation region, (d) deprotecting all of the side chain protecting groups, (e) incubating the protein with a pool of aptamers, where the aptamers can contain a tail including a 5’ phosphate for ligation, a unique DNA barcode (which provides information about the aptamer binding sequence plus the sequencing round), a spacer / common region for subsequent barcode ligation (e.g., a ligation common sequence), a restriction enzyme site with a spacer, and an N-terminal amino acid recognition sequence (e.g., a single-stranded DNA aptamer sequence), (f) ligating the bound aptamer to the DNA tail of the protein, (g) pulling down the protein / aptamer complex using a biotinylated reagent with complementarity to the primer region of the protein / DNA conjugate molecule, (h) washing away unbound aptamer pool, (i) cleaving off the binding region of the aptamer, leaving its DNA barcode attached to the DNA tail of the protein, (j) cleaving off the N-terminal amino acid, (k) denaturing the protein from its biotinylated oligo, (1) collecting the supernatant of the protein with the DNA barcode, (m) repeating steps (c)-(1) until the entire protein has been converted to a DNA strand, followed by PCR amplification and sequencing of the DNA barcode. If the conjugate remains bound and is destroyed during the protein-aptamer complex pull-down, then step (g) can also be performed prior to ligation of the bound aptamer to the DNA tail of the protein [bind, pull down, wash, ligate] (step f). It should also be understood that a biotinylated reagent with complementarity to the primer region of the protein / DNA conjugate molecule can be added during the aptamer incubation (step e) to prevent the aptamer from binding to the DNA region of the peptide target instead of the N-terminal prefix.

[0291] The length of the barcode, including the overhang, can be about 8 to about 26 nucleotides (nt) (e.g., about 9, 10, 12, 15, 16, 18, 20, 21, 22, 23, 24, or 26 nt in length). NGS technologies are currently optimized for short read sequences or up to about 300-600 cycles. For many proteins, long sequencing experiments can be performed (e.g., by PacBio), or the DNA strand can be fragmented into smaller regions and re-aligned after sequencing.

[0292] (C) Protein complex processed in solution and subsequent solid substrate step

[0293] For protein complexes, the proteins within the protein complex can be tagged with DNA oligonucleotide tags through amino acid side chains, and the proximal side chains can be ligated together, then the proteins denatured, then the above outlined protocol (e.g. under part (B)) performed in the absence of fragmentation. The protocol can be optimized so that only immediately adjacent proteins (e.g. bound complex) bear oligonucleotide tags that can be ligated to one another. The protein complex can be pulled down and attached to a solid substrate, which can have DNA adapters placed in such a way that the protein complex can be processed locally. The DNA adapters on the chip can have unique DNA starting barcodes, which when isolated and sequenced, can reveal information about the identity of the adjacent sequenced peptide fragments, and thus the protein complex.

[0294] The PROSEQ method described herein does not rely on prior knowledge of the protein or protein complex (as is required when using e.g. mass spectrometry), and provides a route for de novo sequencing. Once the protein or peptide molecule has been converted into a DNA molecule, the sequence can be amplified, augmented, and modified using conventional tools such as PCR amplification, biotin pull-down, and / or digestion to allow pooling of many samples or determination of low expressing molecules in a sample. Many new biological information can also be obtained using PROSEQ for many non-de novo applications, such as high resolution protein quantification, which is not possible with current conventional protein sequencing techniques.

[0295] Figure 15B is another schematic showing an example of the aptamer-based peptide sequencing method described herein, in which the C-terminal end of a peptide is coupled to an amine-modified oligonucleotide bound to a substrate, or the peptide is covalently bound to an oligonucleotide (1) using other strategies such as click chemistry or SMCC linker (4-(N-maleimidomethyl)cyclohexane-l-carboxylate), the bound peptide is incubated with a library of aptamers bearing DNA barcodes (2), the aptamer bound to the peptide is ligated to a second oligonucleotide immobilized on a solid substrate (3), and the aptamer is cleaved, leaving the DNA barcode attached to the second oligonucleotide (4). Figure 16 is a schematic showing representative aptamers for different amino acids and the corresponding aptamer barcodes, the sequence of which identifies the specific amino acid at that position.

[0296] 2) The protein sequencing method described herein overcomes the processivity limitations of Edman degradation

[0297] The methods described herein overcome the processivity limitations of Edman degradation. For example, after cleavage by Edman degradation, liquid chromatography (LC) is typically used to identify the terminal amino acid. A hypothetical drawback in standard Edman degradation is that there is a maximum number of cycles (~10 cycles) for the accurate degradation and detection of the N-terminal amino acid physically. Since the methods of the present invention do not measure the cleaved-off amino acid, the detection limitation of the cleaved amino acid is not an obstacle. Furthermore, any processivity limitations in the PROSEQ method described herein can be overcome by alternating between cleaving the terminal amino acid using Edman degradation and an aminopeptidase (e.g., trypsin and pepsin). After, for example, about 30 cycles, the methods described herein can cleave the peptide at a particular amino acid site using an exopeptidase, which allows sequencing to start again from a new region of the peptide.

[0298] 3) The protein sequencing methods described herein allow for sequencing of a non-homogenous pool of proteins

[0299] One of the important features of the PROSEQ method described herein is the ability to sequence large pools of proteins in which one or more proteins of interest (e.g., target proteins) are expressed at low or very low levels (e.g., proteins present at parts per trillion; possibly even lower when using the "Sup-Diff" method described herein). This is particularly useful when working with samples such as plasma, which (a) is readily accessible from patients, (b) allows for longitudinal studies, and (c) can provide information for difficult to study diseases, such as neurodegenerative diseases, due to the presence of biomarkers in the bloodstream. In plasma, 13 proteins plus albumin account for 96% of the protein sample, with some of the most interesting molecules, such as tissue leak products and cytokines, accounting for the last 4% of the sample, and finding them completely below the instrument detection resolution limits of MS (Schiess, Wollscheid, & Aebersold, 2008). Thus, it can be extremely difficult to identify biomarkers or new proteins in a plasma sample using MS. Unlike HPLC and MS, identifying amino acids on the basis of aptamer binding is not limited by the detection limits of a single protein at high concentration in the sample. Since the end product of the actual sequencing is DNA rather than protein, there are well developed tools for amplifying, annealing, and pulling down specific populations of DNA of interest. After DNA barcode strand formation, DNA sequencer platforms can clonally amplify the sequence (e.g., using bridge amplification). The thousands of clusters of each DNA sequence produce a much larger readable signal than the initial input signal from the low expressing protein, while bypassing single molecule techniques. This ability to sequence large non-homogenous pools allows for sequencing of thousands of antigens across the entire proteome of an organism.

[0300] For samples with large dynamic ranges, a method called "sup-diff" can be used to remove the DNA barcode constructs of highly expressed proteins, resulting in an increased proportion of DNA barcode constructs of low expressed peptides or protein clusters remaining in the oligonucleotide pool to be sequenced. For example, there are two methods used to increase the proportion of desired or low expressed peptides: a priori and a posteriori methods. The overall strategy is to develop a pool of ssDNA baits that contain biotinylated RNA sequences complementary to certain sequences in the initial pool of ssDNA (Diatchenko et al., 1996) (Gnirke et al., 2009). The RNA bait pool is used to capture ssDNA targets by hybridization in solution and subsequent pull-down on streptavidin-coated magnetic beads.

[0301] The main difference between the a priori and a posteriori methods is that the a priori method pulls out only known sequences, while the a posteriori method pulls out high abundance sequences in a pool whose distribution and composition are unknown. In the a priori method, the pool of ssDNA is first sequenced, and then the user can design baits specific to the sequences that the user wants to pull out from the pool, which can include very high concentrations of sequences that can be contaminants. The a priori method enriches for sequences that are not pulled down by the designed baits, thus reducing the NGS sequencing reads for the targets that one initially hopes to pull out from the pool. In the a posteriori method, the initial pool of ssDNA is used directly to generate the RNA bait pool. The RNA bait pool can have the same fraction distribution as the original target pool or a distribution slightly biased towards the initial high abundance sequences. Assuming that higher abundance target sequences are more likely to find their RNA bait partners under optimized time, temperature, and overall bait to target ratios, high concentration sequences are more likely to be pulled out when the RNA baits are hybridized to the initial pool of ssDNA. See, e.g., Figure 17 .

[0302] 4) The protein sequencing methods described herein allow for sequencing of DNA barcodes using a variety of DNA sequencing technologies

[0303] The methods for sequencing proteins described herein can be performed in conjunction with any existing DNA sequencing technology. Using custom flow cells with DNA printed on glass in a specified manner and automated fluidics systems, barcoding can be established as described in the previous section without reprogramming or changing the use of existing DNA sequencing platforms. These DNA barcodes representing protein / peptide sequences can then be sequenced on any existing DNA sequencing platform or technology.

[0304] 5) The protein sequencing methods described herein include strategies to ensure robust protein and DNA sequencing capabilities despite the demanding chemistry of Edman degradation

[0305] The ProSeq method described herein converts protein sequences into DNA signals readable on next generation sequencing (NGS) platforms using barcoded amino acid specific aptamers. The methods described herein overcome the distortion of protein sequencing platform components caused by Edman degradation that prevents clustering of DNA barcode constructs and thus sequencing directly on the same chip. Trifluoroacetic acid (TFA) and the pH fluctuations that occur during Edman degradation cause two major problems: (1) loss of DNA cluster generation through removal or modification of the P5 and P7 DNA adapters on the chip, (2) modification of the constructed DNA barcode causing loss of sequence information and amplification ability.

[0306] (A) Off-chip sequencing of DNA barcodes

[0307] After establishing the DNA barcode construct containing the strands of DNA barcodes indicative of the aptamer binding sequence of the peptide, the construct is amplified on the chip or excised from the chip and amplified in solution. Amplification methods used include but are not limited to PCR, loop-mediated isothermal amplification, nucleic acid sequence-based amplification, strand displacement amplification, and multiple displacement amplification. In addition, the original DNA barcode construct can be transcribed on the chip into a large number of RNA constructs, which can then be converted into a cDNA library consisting of many copies of the original DNA barcode. The amplified product, i.e., copies of the original DNA barcode construct, can be removed from the microfluidic chamber and sequenced using standard DNA sequencing methods including but not limited to Sanger sequencing, NGS, Ion semiconductor sequencing, SOLiD technology, cPAS, etc. The number of reads is normalized to the number of PCR cycles used to estimate the amount of each protein or peptide sequenced from the initial sample.

[0308] (B) XNA or modified DNA / RNA adapters, substrates, and barcodes

[0309] The methods described herein are a single-chip strategy that overcomes the degradation of DNA components on the BCS platform by utilizing XNA or modified DNA / RNA with the following properties: (a) resistance to transformations caused by Edman degradation or highly acidic conditions, (b) ability to be manufactured as chimeras with conventional DNA nucleotides, and (c) compatibility with existing polymerases that can amplify these non-natural nucleic acids or convert the modified sequences into conventional DNA bp. Such modified nucleic acids can include modifications to the 2' carbon of ribose or to the purine bases themselves that enhance their hydrolytic stability (Watt et al., 2009). Examples include but are not limited to 2'-0-methylated RNA, 2'-fluoro deoxyadenosine, 7-deaza-2'-deoxyadenosine, and 7-deaza-8-aza-deoxyguanosine.

[0310] • Adding XNA or modified DNA / RNA adapters to degraded P7: The methods herein can utilize the degraded P7 adapters available on the chip as a basis for custom XNA or modified DNA / RNA adapters. After subjecting the P7 and P5 adapters to acidic conditions, the P5 adapter is at least partially removed and the P7 is degraded. Two methods of adding new adapters for ligation and barcode generation handles after barcode cluster generation are:

[0311] • Method 1: Several rounds of Edman degradation can be performed to remove P5 and deprotonate P7, and XNA or modified DNA / RNA adapters can be ligated to the remaining region of P7. One method of XNA adapter ligation is to ligate XNA or modified DNA / RNA adapters with a phosphorylated 5' end to the 3' end of P7. If the modified nucleic acid analogs reduce ligase efficiency, the adapter sequence can be a chimeric XNA or modified DNA / RNA molecule with one or more standard cytosine or thymine nucleotides at its 5' end.

[0312] • Method 2: Several rounds of Edman degradation are performed to remove P5 and deprotonate P7, and XNA or modified DNA / RNA adapters are attached to the remaining region of P7 using click chemistry. Another strategy for adding XNA adapters is to attach XNA or modified DNA / RNA adapters to P7 by ligation of an oligonucleotide linker with a reactive group at the 3' end onto the 3' end of P7. A functional XNA or modified DNA / RNA adapter that can optionally contain a cleavage site and have a corresponding reactive group at its 5' end upon ligation to the oligonucleotide linker can be attached to P7 by a chemical reaction. Examples of pairs of reactive groups include, but are not limited to, NHS ester and amine (azide reaction), azide and alkyne (triazole reaction), maleimide and thiol (thioether reaction), and tetrazine and alkene. P7 and the linker can be blocked from unwanted annealing of oligonucleotides that are partially complementary to both P7 and the extender oligonucleotide during aptamer incubation.

[0313] • XNA or modified DNA / RNA bases and barcodes: The base segments of the methods herein, the binding regions of the aptamers, the BCS cassette components, the aptamer barcode regions, or combinations thereof can comprise XNA or modified DNA / RNA.

[0314] Once the P5 adapter is no longer detected, the Illumina sequencing protocol ends the sequencing run, so in embodiments where P5 is removed from the sequencing platform, additional steps can be required to prevent premature sequencing stoppage. These steps can include, individually or in combination:

[0315] • Adding multiple P5 to the chip after the last round of Edman degradation by enzymatic or chemical means

[0316] • Adapt sequencing instrument protocol code to continue sequencing run in the absence of P5

[0317] • Attach custom primer sequence to the cleavage site of the altered P7 strand by enzymatic or chemical means and adapt the sequencing protocol code to detect the custom primer sequence instead of P5 to determine whether to terminate the sequencing run.

[0318] 6) Exemplary variations of the protein sequencing method described herein include, but are not limited to, Figure 18 ) :

[0319] • Multiple rounds of aptamer binding: In some cases (e.g. if there are aptamer-specific binding issues), several rounds of aptamer binding / DNA barcode encoding / aptamer denaturation can be performed before proceeding with N-terminal amino acid degradation for error correction. The additional data acquisition allows downstream computational analysis to reduce noise for each measurement.

[0320] • Aptamer for two amino acids: In some cases (e.g. if the aptamer for a single amino acid does not have high enough affinity or is not specific enough for the method), aptamers for two or more consecutive amino acids can be generated Figure 14A ). The additional benefit of aptamers that bind and encode two amino acids is that the signal-to-noise ratio is improved because each amino acid (except the N- and C-termini) is read twice.

[0321] • Substrate: This barcode sequencing method can also be performed on a glass or quartz substrate with DNA oligos printed or chemically attached in random or patterned events. Such chips can be custom made or purchased; for example, academic labs fabricate chips using clean rooms and DNA spotting machines, Agilent prints microarrays with known oligo sequences on glass in a spot pattern, and Illumina’s next-generation sequencing chips are glass slides with a random distribution of DNA adapters for P5 and P7 sequence binding sites attached to the solid surface. In the case of custom glass slides or substrates, the DNA oligos can have a specific pattern to reduce off-target ligation noise.

[0322] • Different oligo orientation: The protein sequencing method described herein orients the DNA barcode sequence such that the 5’ end is attached to the DNA adapter on the chip. Using alternative or custom chips, it is possible to instead attach the 3’ end of the barcode sequence to the chip surface.

[0323] • In solution: The need for a solid substrate can be completely eliminated by directly ligating DNA barcodes to the C-terminus of the peptide. The C-terminus of the peptide can initially contain a short oligonucleotide sequence that allows ligation between the aptamer end and the peptide tail that is bridged by, for example, a 5-mer oligonucleotide. After Edman degradation, a subsequent DNA barcode can be ligated to the free end of the peptide tail. The resulting barcode sequence can then be PCR amplified and sequenced using standard NGS technology.

[0324] • Beads in solution: Peptides and oligonucleotides can be attached to beads (magnetic, glass, glass-coated magnetic beads, or other beads coated with acid-resistant material), and the successive peptide sequencing steps (e.g., aptamer binding, barcode incorporation, and peptide degradation) can be performed by soaking and isolating the beads in solution. After the desired number of sequencing cycles, the DNA barcodes that provide the sequence of the peptide can be PCR amplified directly on the beads and sequenced using standard NGS technology (Hoon, Zhou, Janda, Brenner, & Scolnick, 2011).

[0325] • Different binders: In addition to aptamers, barcoded binders such as RNA, peptides, proteins, nanobodies, or other small molecules can be used to recognize amino acids.

[0326] • Different proteases: In processing protein samples, different proteases such as Lys-C, trypsin, or a combination of multiple proteases as described above can be used. In addition, the sample can be divided into multiple samples and treated with multiple proteolysis strategies to establish different proteomic profiles.

[0327] • Single platform versus separation of steps: Edman degradation of the peptide and generation of the DNA barcode can be performed outside of the sequencer platform, or a complete end-to-end automated single platform can be established. The DNA barcode strand can be immobilized and sequenced in separate steps.

[0328] • Bridge design: Bridges are oligonucleotides that are complementary to the aptamer tail portion with a 3’ single-stranded overhang that anneal to the restriction site spacer and barcode (Figure 14). The bridges can be designed such that they can be (a) barcode-specific bridges, where the bridge is fully complementary to the aptamer tail including the barcode region except for the 3’ single-stranded overhang region, such that each unique aptamer has a unique bridge associated with it ( Figure 14B ), or (b) universal bridges, where the bridge is only complementary to the restriction site spacer and the consensus sequence, which are conserved in all aptamers and located on the aptamer tail flanking the barcode, such that all unique aptamers share the same bridge oligonucleotide ( Figure 19). For universal bridges, the region that forms a duplex with the barcode on the aptamer tail can consist of (a) a sequence of universal base analogs such as 5-nitroindole, 3-nitroindole, and 4-nitrobenzimidazole, or (b) a gap of no bases, such that the universal bridge consists of two separate oligonucleotides that anneal to regions flanking the barcode.

[0329] • Ligation method: DNA barcodes can be chemically ligated rather than enzymatically ligated together.

[0330] • Different readouts: Instead of using DNA barcodes, one can use fluorescent dyes, beads, nanoparticles, etc. (see also the PROSEQ-VIS method described herein) to identify the amino acid conjugates.

[0331] • Sequential amino acid degradation: Cleaving off individual amino acids between rounds can be done enzymatically or chemically, for example by Edman degradation.

[0332] • Sequencing directionality: Individual amino acids can be cleaved off from either the N-terminal or C-terminal end (Casagranda and Wilshire, 1994) (Cederlund et al., 2001). Protein sequencing starting from the N-terminal end is described in detail herein. On the basis of the present disclosure, it should be recognized that similar methods can be combined with aptamers designed to specifically recognize and bind one or more C-terminal amino acids, for application to protein sequencing starting from the C-terminal end. Methods for removing C-terminal amino acids and producing C-terminal amino acid-shortened proteins or peptides (instead of using, for example, Edman degradation to produce N-terminal amino acid-shortened proteins or peptides) for C-terminal sequencing are known in the art and can be used, including but not limited to Bergman et al. (2001, Anal. Biochem., 290(1): 74-82) and Casagranda and Wilshire (1994, Methods Mol. Biol., 32:335-49).

[0333] It should be understood that the PROSEQ method described herein can also serve as a large scale high throughput binding specificity assay to characterize interactions in different material binding contexts (BCS binding assays). A key advantage of this assay is that it allows one to record one or more binding events between a plurality of hypothetical binders and a plurality of targets in one experiment. Once the desired targets are coupled to the co-localization substrate, the substrate can be immobilized onto a glass substrate or processed in solution. A library of DNA barcoded hypothetical binders (PBL) is then incubated with the desired targets and non-targets, allowing binding to occur. Each DNA barcoded hypothetical binder comprises a binder molecule coupled to a DNA sequence that contains at least a) a restriction site, b) a ligation site (e.g., a first ligation site), c) a unique DNA barcode that indicates the identity of the hypothetical binder and the binding round, and d) another ligation site (e.g., a second ligation site). When a hypothetical binder binds to an immobilized target, its DNA barcode tail ligates to the proximal target-barcoded DNA substrate that is co-localized with the target. The ligated barcode is cleaved with a restriction enzyme, exposing a DNA barcode construct that is ligated to another binder barcode in the next round. After repeating this series of steps on the chip, the DNA barcode strand that contains information about the identity of the binder and the target and the order of binding events can be read out using conventional DNA NGS technology Protein or peptide sequencing with visualization (PROSEQ-VIS)

[0334] The PROSEQ method described herein produces a number of advantages, including but not limited to the ability to:

[0335] • produce a probability distribution of binding events in one mixture by interrogating the same target multiple times;

[0336] • separate binding events from unbound binder molecules through wash steps for solid state methods. Separation of binding and ligation events reduces off-target ligation events.

[0337] • assay large libraries of hypothetical binders in a variety of different contexts (e.g., in the presence of non-targets, other targets of interest, etc.). This is particularly important for identifying binders that will be used in applications where a variety of different targets are present through selection processes where binders are selected in isolation from other hypothetical targets;

[0338] • detect rare binding events in a high noise environment (due to high resolution data in NGS);

[0339] • determine the dynamic range of functional buffer conditions for binders;

[0340] ​• If the reaction does not proceed in solution, the process of separating bound and unbound ligand is simplified by simply flowing a wash buffer over the surface.

[0341] Figure 20

[0342] The PROSEQ-VIS method described herein converts amino acid sequences into optical barcodes. In the PROSEQ-VIS method described herein, aptamers conjugated to fluorophores can be used to deconvolute amino acid sequences, allowing de novo protein sequencing. The PROSEQ-VIS method described herein is capable of sequencing a wide variety of samples, particularly samples in which one or more proteins of interest (e.g., target proteins) are present at low or very low concentrations (e.g., proteins present at parts per trillion). The PROSEQ-VIS method described herein also provides computational tools to determine the identity of the N-terminal amino acid based on the unique spectral signature of the observed binding events.

[0343] The PROSEQ-VIS method described herein uses aptamer binding to convert protein sequences into a series of fluorescent images or "optical barcodes" that can be read by microscope imaging. The optical fluorophores can be assigned to their aptamer, revealing the underlying protein sequence. See, e.g., Figure 21 This protein sequencing method can be used on small amounts of sample (including single cells or small volumes of blood) to identify the entire expressed proteome, lowly expressed proteins, and single amino acid mutations to better understand complex disease phenotypes. In addition, the PROSEQ-VIS method described herein can be performed on intact cells and tissues to visualize not only the sequence of the protein, but also the location in the sample. See, e.g., Table 1.

[0344] Table 1

[0345]

[0346] The PROSEQ-VIS method described herein can be used in a clinical setting to identify new protein fusions or disease-related mutations from individual patient samples, develop diagnostics or prognostics, evaluate patient response to treatment, or predict the likelihood of a possible response to certain treatments. In addition, the methods described herein can be used broadly to characterize proteins, discover new biomarkers, analyze the whole or macro proteome, establish cell lines, and evaluate mechanisms related to protein abundance, sequence, or function.

[0347] 1) Aptamers provide the ability to perform de novo sequencing

[0348] The PROSEQ-VIS method described herein uses a library of aptamers specific for unique combinations of one or two N-terminal amino acids described herein, where each residue pair has at least one (e.g., more than one, e.g., multiple) aptamer binders. The ssDNA aptamers are designed to contain a region that includes a fluorophore or a region for annealing a short dye-coupled ssDNA probe so that the N-terminal amino acid can be identified by the unique spectral signature of the binding event between the N-terminal amino acid and its corresponding aptamer.

[0349] Proteins can be obtained from a sample (e.g., a blood sample, cell lysate, or single cell), denatured, blocked, and cleaved into peptide fragments. While intact proteins that are denatured can be analyzed without cleavage, proteins cleaved into smaller peptide fragments are optimal because: (1) the rounds of Edman increase the background noise in the imaging, so fewer sequencing rounds can be used to determine the sequence of the peptide fragment, and (2) certain imaging modes (like TIRF) have a narrow focus window (nms of 10s-100s) and signal detection is highly dependent on the sample being sufficiently contained within the optimal imaging window. Proteins can be cleaved into peptide fragments using, for example, a conventional trypsin method that cleaves at lysine and arginine and / or other fragmenting enzymes that cleave at random amino acid sites. A combination of both methods can help reduce errors in computational alignment after sequencing. Once the proteins are converted into short peptides, the free and unblocked C-terminal end can be coupled to a DNA primer oligonucleotide on a glass substrate or directly to a glass Figure 20 ) substrate. Aptamer libraries can then be flowed over the peptides for incubation, allowing the aptamers to specifically bind to the N-terminal amino acid residues. There are many ways to fluorescently label the aptamer tails. Two possible imaging options are that the aptamer tails can have: (a) tails with optical barcodes for imaging, or (b) regions where one or more short fluorescently tagged DNA probes can anneal to the aptamer: amino acid complex.

[0350] 1.1 Direct aptamer-dye coupling

[0351] After the aptamer binds to the N-terminal prefix, the optical signature (a) of the aptamer can be imaged by a multi-channel single molecule epi-fluorescence or total internal reflection fluorescence (TIRF) imaging setup. For each N-terminal prefix readout (“round”), unbound aptamers are washed away, and z-axis multi-layer scans of the images can be obtained during incubation to confirm the spectral signature of the N-terminal amino acid. The next round is then started by removing the N-terminal amino acid on the immobilized peptide using Edman degradation and / or aminopeptidase. The same pool of aptamers can then be used to interrogate the newly exposed N-terminal amino acid Figure 20A-20D). After repeating this series of steps, the identity of each N-terminal amino acid can be calculated at each round by comparing the binding events observed for each peptide against the probability distribution of binding events for each aptamer-amino acid complex. Using this information, the amino acid sequence of each peptide can be deduced on the basis of the series of amino acid signatures obtained in successive rounds of imaging and degradation. See, e.g., Figure 22 E.

[0352] 1.2 Hybridization of oligonucleotide-conjugated dyes to aptamers

[0353] In the case of using aptamers with regions that bind to complementary fluorescently-tagged oligonucleotides, each "round" of N-terminal prefix readout for the assay comprises multiple "iterations" of probe incubation and imaging. The aptamers comprise 3 regions: (a) an active binding region, (b) an optional spacer, and (c) a barcode tail with one or more combinations of barcode elements that indicate the number of probing iterations and fluorescent tags, where each barcode is complementary to a fluorescently-tagged oligonucleotide Figure 23 ). To prevent the barcode region from affecting the folding of the binding region of the aptamer, the region of the oligonucleotide that does not bind to the N-terminal prefix can be partially or completely protected by hybridization to a complementary oligonucleotide to form an aptamer with partial duplexing as the aptamer library flows through. The aptamer:amino acid complex can be incubated with a library of probes that hybridize to the barcode region indicating probe iteration 1. The number of unique fluorescent tags that can be used per iteration depends on the number of channels in the imaging device, the nature of the fluorescent dyes and emission filters, and the sensitivity of the detector. During each iteration, each aptamer can hybridize to one or more oligonucleotide-bound probes for multiplexing, as long as the complementary barcode elements on the aptamer do not overlap for that iteration. The unbound probes can then be washed away, and the bound probes can be imaged to obtain the first segment of optical barcodes. Subsequently, the bound aptamers can be incubated with the next set of probes that hybridize to the barcode region indicating probe iteration 2. The iterations of probe incubation, imaging, and washing can be repeated until a complete optical barcode is obtained. Finally, Edman degradation can be performed to remove the N-terminal amino acid and the aptamer bound to it to expose the next N-terminal amino acid for the next round of sequencing Figure 24

[0354] It should be appreciated that procedural modifications can be made, particularly to the imaging and downstream signal deconvolution strategies, to account for the affinity and specificity of the aptamer used to probe the N-terminal amino acid. In the case of utilizing a high-specificity binder, a library of aptamers specific to unique N-terminal amino acid prefixes and with low K d (tight binding) is flowed through, unbound aptamers are washed away, and the optical barcode is observed as described above Figure 25 ​). In the case of aptamers with medium to low specificity, a library of fluorophore conjugated aptamers can be flowed over the peptide, allowing the aptamers to semi-specifically bind to a set of N-terminal amino acid residues. Such aptamers preferentially bind to a given target and can also bind to a subset of known N-terminal amino acids, with each binding pair having a known probability distribution. For each round of sequencing, images can be acquired before (to obtain background), after (to obtain specific binding), or during (to obtain kinetics on off determination) the aptamer incubation, in order to generate a spectral signature of the N-terminal amino acid prefix consisting of multiple binding events, before the N-terminal amino acid is removed to expose the next amino acid to be probed. Several rounds of incubation and detection can be performed before the N-terminal amino acid is removed by Edman, in order to increase the confidence of the detected signal. After multiple rounds of aptamer binding are repeated, the identity of the N-terminal amino acid can be calculated at each round by comparing the binding events observed for each peptide to the known probability distribution of binding events for each aptamer amino acid prefix, as each unique N-terminal amino acid is expected to have its own unique binding signature for a given pool of medium to strong binders Simultaneous screening of multiple targets (multiplexing) In addition to or as an alternative to aptamers, binders such as RNA or small molecules can be used to identify amino acids.

[0355] The methods described herein do not rely on prior knowledge of the protein (e.g. a database of peptides as required in mass spectrometry) and provide a route to de novo sequencing. However, if a database of proteins is available, it can only be necessary to identify a portion of the amino acids in order to accurately map the peptide fragments back to the full-length protein. Furthermore, if the protein is purified or selected (e.g. by molecular weight, charge, or affinity for a known molecule) prior to sequencing, it will further focus the list of candidates on the basis of the portion of the amino acid sequence of the full-length protein that is identified.

[0356] The PROSEQ-VIS method described herein has a number of advantages and applications, including but not limited to the ability to:

[0357] 1) sequence peptides regardless of peptide concentration;

[0358] 2) convert protein sequences into optical sequences, which allows the separation of signals from lowly expressed proteins;

[0359] 3) perform de novo protein sequencing (to allow, for example, the direct discovery of sequences in molecules such as cytokines);

[0360] 4) handle small volume samples, down to single cell protein sequencing; and

[0361] 5) sequence peptides in situ to obtain localization data for proteins in intact tissues.​

[0362] Instead of using fluorophore-conjugated aptamers or oligonucleotide probes to identify amino acids, other optical methods such as quantum dots and dye-conjugated nanoparticles can be used. Other microscopy techniques, offering varying degrees of resolution quality, can be used for imaging instead of TIRF. Finally, in the PROSEQ-VIS method described in this paper, another type of N-terminal amino acid-binding small molecule, barcoded with optical barcodes, is used instead of aptamers, which also allows for protein sequencing on the PROSEQ-VIS platform.

[0363] Figure 26

[0364] Attempts by others to screen multiple targets using SELEX have successfully multiplexed up to 30 biologically similar targets in a single SELEX experiment (e.g., BasePair's VENN multiplexed SELEX). While the specific method for achieving this is not yet known, it likely involves binding targets to beads with different spectral contents and incubating them with aptamer candidates, followed by sorting via fluorescence-activated cell sorting (FACS). This method limits the number of targets that can be multiplexed at once due to the optical limitations of the machine.

[0365] The multiplexing method described herein allows for the simultaneous screening of conjugates to multiple peptide or protein targets. Furthermore, this method enables the detection of rare binding events in noisy environments, improves target specificity, and allows for specificity assays for multi-target cross-validation matrix analysis and machine learning analysis. The multiplexing method described herein can be used to identify interactions between virtually any two biological molecules (e.g., two DNA or RNA barcoded molecules such as oligonucleotides and molecular targets, proteins and antibodies, small molecules and barcoded proteins), provided that both targets can be coupled to the oligonucleotide, which can then be linked to each other.

[0366] The multiplexing method described in this paper involves multiplexing aptamer candidates ( Figure 26 A) A diverse pool of unbound DNA-barcoded peptide targets ( Figure 26 B) Incubation. After aptamer binding, the 3' end of the single-stranded aptamer is linked to the ssDNA barcode of the peptide ( Figure 26 C) The DNA portion is then amplified by PCR. Sequencing the aptamer and its covalently attached DNA barcode provides the aptamer sequence and a unique identifier indicating the target to which the aptamer binds, thereby eliminating the barrier to identifying which aptamers bind to which targets. Figure 27 D is a schematic diagram indicating the steps (from Figure 3) in which multiplexed SELEX procedures can be incorporated.

[0367] The multiplexing methods described herein can reduce labor and reagent costs while improving data quality and broadening screening capabilities. Furthermore, the multiplexing methods described herein can produce aptamers that specifically bind to their unique targets in environments with a large number of available targets (e.g., cell surfaces, human blood), greatly increasing the aptamer discovery-to-application pipeline.

[0368] 1) Identification of peptide or protein targets using DNA barcodes

[0369] As described above, in the multiplexing methods described herein, the targets are peptide-oligonucleotide conjugates (POCs), with reference to Figure 27 which are single-stranded (ss) DNA tails (a) covalently attached to the C-terminus of a peptide or protein target (b) at their 3’ end. The ssDNA tail (a) includes a 3’ primer region (c), a unique DNA barcode (d), and a 5’ bridge binding sequence (e). The aptamer (f) includes a 3’ bridge binding sequence (g). After the POC-aptamer binds in solution, a short oligonucleotide bridge (h) can be introduced, with one half of the short oligonucleotide bridge (h) complementary to the 3’ bridge binding sequence (g) at the 3’ end of the aptamer (f) and the other half complementary to the 5’ bridge binding sequence (e) of the ssDNA tail (a). After the bridge oligonucleotide binds to both the aptamer and the peptide tail, a ligase can be added to close the gap, unused bridge oligonucleotides can be degraded and / or removed, and the ligase can be inactivated. This results in a covalent linkage of the aptamer (f) to the peptide (b).

[0370] After ligation, bead-bound POC targets can be obtained (e.g., pulled down with complementarity to biotinylated oligonucleotides), and unbound aptamers can be removed (e.g., washed away). PCR can be performed on the beads through the ssDNA tail and aptamer, and the resulting DNA construct can be sequenced to obtain the aptamer sequence and its protein binding partner’s barcode identifier (boxed region in Figure 28 ).

[0371] 2) Identification of local aptamer binding events from global noise using proximity-dependent DNA ligation

[0372] One difficulty encountered in the multiplexing methods disclosed herein is constraining the assay in a way that favors ligation of bound partners over ligation of randomly available species in solution, as physically proximal peptide tails and aptamers are more likely to ligate to each other than to free-floating DNA. Thus, ligation reaction conditions can be developed and optimized to maximize local signal by optimizing several experimentally tested parameters including, but not limited to, reaction time, substrate concentration, temperature, and reaction solution. Furthermore, tail lengths and bridge regions of different lengths can be designed and characterized to optimize local interactions in high-noise environments.

[0373] 3) Nested PCR for additional rounds of multiplexing-SELEX

[0374] To implement multiple rounds in the multiplexing methods described herein, the aptamer segments of the ligated aptamer-barcode products can be re-amplified (e.g., using a primer pair flanking the aptamer sequence to perform nested PCR on the ligated complex) and processed (e.g., purified by automated electrophoresis gel separation) and then converted to ssDNA (e.g., using enzymatic digestion). See Target protein and RNA binding protein fusion (TURDUCKEN) .

[0375] 4) Alternative forms and variations of the multiplexing methods

[0376] Numerous procedural modifications can be made to adapt the multiplexing methods described herein to different applications.

[0377] The multiplexing methods described herein can be used to examine interactions in different material binding scenarios; for example, but not limited to: a) DNA-peptide binding, where the interaction region comprises an aptamer that binds to a peptide target; b) DNA-DNA binding, where the interaction region comprises a region of base complementarity between two DNA strands. For DNA-DNA interactions, the ability to identify local signals has been demonstrated at as low as 0.001% of the total pool in solution at a concentration of 500 nM of the binding partner, demonstrating the sensitivity of the multiplexing methods described herein

[0378] In addition, the multiplexing methods described herein can be used to examine material binding other than DNA-DNA or DNA-peptide interactions. For example, the multiplexing methods described herein can be used to examine binding between any number of biological targets, so long as the two targets can bind to each other (e.g., through ligation of oligonucleotides). For example, a similar multiplexing method as described herein can be used to screen RNA aptamers that bind small molecule targets or protein complexes.

[0379] The ssDNA tail can be attached to the C-terminus of a peptide or protein using any number of different techniques, including but not limited to chemical linkers (e.g., click chemistry, SMCC linkers, EMCS linkers, etc.), biological linkers (e.g., biotin-streptavidin systems), cross-linking (e.g., using formaldehyde or UV), etc.

[0380] In addition, it can be recognized that the ssDNA tail can be attached to different regions of a protein or peptide (i.e., other than the C-terminus). For example, the ssDNA tail can be attached to the N-terminus, specific functional groups, amino acid side chains, etc. Additionally or alternatively, multiple ssDNA tails can be attached to a single peptide or protein.

[0381] The ligation between DNA ends can occur in a variety of ways. Enzymatic ligation in aqueous solution can be used, but chemical ligation of DNA ends can also be used. In certain embodiments, optional ends of the bridge can be used for ligation. The overhangs and / or bridges can also be modified to contain base pair mismatches to introduce a gradient of binding interactions, such that the binding interaction between the binder and the target is favored over the binding interaction of the bridge.

[0382] It should be understood that the multiplexing methods described herein can be performed in aqueous solution, or they can be tailored for use in different systems, e.g., on a fixed surface, on a bead, in vivo, in a gel, etc.

[0383] The multiplexing methods described herein have been used to identify aptamers that have selective binding to a peptide target in a competitive, multiplicity of peptides environment. Similar to selective antibodies, the resulting aptamers are suitable for use individually or in combination of two or more to create constructs that control their multi-target binding profile. For example, two aptamers each having high selectivity for different targets can be linked together to create a construct that binds two independent targets; or two aptamers having the same primary target but different off-target binding profiles can be added to a pool in parallel or sequentially to improve the binding readout to their common target by analysis of the overlapping profile area.

[0384] Substitution of molecules that have been DNA-barcoded and have 3' C-overhang arms for aptamers in the multiplexing methods described herein allows for the measurement of binding between different mixtures of any of the foregoing classes of molecules, enabling bidirectional multiplexed competitive measurements of combinations of any class of molecules, including peptides to proteins, protein-protein, antibody-protein, small molecule-protein, peptide-cell surface marker, antibody-cell surface marker, etc. In certain embodiments, both binder and target molecules can be drawn from any mixture of molecules of any of the foregoing classes, allowing for the measurement of cross-binding in a complex competitive environment.

[0385] The multiplexing methods described herein provide a highly sensitive tool for detecting low level binding events in large pools of substances. The multiplexing methods described herein reduce the need for a large number of SELEX rounds (e.g., 8 to 20 rounds) and at the same time allow multiplexing of several peptide targets in one solution. As a result of the reduction in rounds, the multiplexing methods described herein minimize the number of PCR amplifications that must be performed on the aptamer pool and thus minimize the bias introduced by each round of amplification. Increased specificity and reduced off-target binding are additional benefits in the multiplexing methods described herein. For example, if a unique aptamer is isolated that binds to peptide target number 1 in a mixture containing targets 1-10, it is also known that the aptamer does not bind to targets 2-10 (under those same conditions) in addition to target number 1. This reduces the likelihood of selecting a non-specific aptamer that can bind to other targets in addition to the target of interest.

[0386] Figure 29

[0387] Classification of binding interactions is highly desirable in a large number of research areas including drug development, diagnostics, and basic research. Protein and peptide libraries contain a library of biological targets of interest to which binders (e.g., aptamers, small molecules, antibodies, etc.) can be screened. Currently, screening is typically performed in individual reactions in which the identity of the protein or peptide target is known, making large-scale screening of unknown targets prohibitively expensive and laborious. Pooling and screening several targets at once allows for scale-up and higher binding specificity, however, no method is currently available that can produce a target library from which the identity of each target in the pool can be easily deduced.

[0388] Biological methods for producing protein or peptide libraries rely on cloning each protein individually into a model system such as yeast or E. coli and purification (Jia & Jeon, 2016). To produce a library of 1,000 unique proteins, a researcher must perform 1,000 independent transformation reactions, protein purification, and QC processes, finally pooling the proteins together. Chemical synthesis can reliably produce peptide pools and can quickly become prohibitively expensive and technically challenging for larger proteins and protein complexes.

[0389] Importantly, existing methods for generating libraries do not allow scientists to easily identify individual elements after component pooling. Commonly used techniques for protein identification include mass spectrometry, antibody binding assays, and affinity tag binding assays (Miteva, Busaeva, & Cristea, 2012). Concentration thresholds of unique elements within a protein pool limit the use of mass spectrometry for identifying low-expressed individual proteins from large pools; antibodies are often inconsistent, absent, or prohibitively expensive for novel targets; and affinity tagging methods limit the diversity of the pool to the number of unique affinity tags available.

[0390] The TURDUCKEN method described in this paper allows for the creation of mixtures of thousands of unique proteins, their tagging, and screening and identification within a single pool. The TURDUCKEN method also allows for the generation and screening of pools of diverse proteins.

[0391] 1) Protein expression

[0392] An in vivo system in *Saccharomyces cerevisiae* and *Escherichia coli* is described, wherein each transformed cell is engineered to produce a different protein of interest (POI), which can be non-covalently linked to an RNA barcode of its sequence that can be used to identify the POI; the non-covalent link depends on the natural interaction between the RNA binding site and its corresponding RNA-binding protein (RBP). See, for example... Figure 29 Representative RNA binding sites and their corresponding RBPs that can be used for such constructs include, but are not limited to, MS2 RNA hairpins bound to MS2 phage coat proteins and boxB sequences bound to phage anti-termination protein N (λN). Each POI ( Figure 29 A) can be expressed as an RNA-binding protein ( Figure 29 A fusion protein of part B), wherein the POI can be non-covalently linked to a specific RNA binding site recognized by an RNA-binding protein. Figure 29 Part C) and unique barcode ( Figure 30 Part D). Each construct in the pool typically contains a POI fused to the RBP, a DNA sequence encoding an RNA sequence recognized by the RBP, a unique RNA barcode, and a promoter to drive expression. Representative promoters include, for example, the Gal 1,10 bidirectional promoter, ADH1, GDS, TEF, CMV, EF1a, SV40, T7, lac, or any other promoter and promoter combination compatible with the host organism. Pools containing plasmids of various different POI-RBP fusion genes and their corresponding RNA barcode sequences can be transformed into Saccharomyces cerevisiae at an approximate dilution of one plasmid per cell. Figure 30A). POI fusions made in vivo then bind their corresponding RNA barcodes Figure 30 B) which can then be purified Figure 30 C). Generation of large, diverse, and controlled DNA libraries by ligation (LEGO) D is a schematic demonstrating where the product of the TURDUCKEN described herein can be used relative to the SELEX method (Figure 3).

[0393] 2) Protein purification

[0394] POI-RNA complexes can be obtained using any of a variety of methods, resulting in the collection of only complexes containing both POI fusion proteins and RNA barcodes. By way of simple example, the complexes can be pulled down from cell lysate by a His tag or other purification tag that can be included in the protein fusion component of the POI. The POIs can then be washed and released from anti-His beads or other pull-down method compatible with the purification tag used, and further purified using streptavidin-coated beads and biotinylated oligonucleotides reverse complementary to the sequences in the RNA barcodes. After this pull-down step, a mixture of biotinylated oligonucleotides annealed to random RNA sequences bound to POI-RNA complexes, or beads without binding, is obtained. The POI-RNA complexes can be released from streptavidin-coated beads and purified by heating and washing the mixture to denature the RNA and biotinylated oligonucleotides or by using a restriction endonuclease to release the complexes.

[0395] 3) Protein pools for aptamer binding assays

[0396] The final product from this method is a diverse pool of proteins, each of which can be identified by the attached RNA barcode. This design allows for the use of this protein pool in multiplexed aptamer screening assays. For example, a pool of potential aptamers also containing their own unique nucleic acid barcodes can be incubated with the protein pool and allowed to bind their targets from the pool of potential aptamers. Through controlled enzymatic ligation (see, e.g., the multiplexing methods described herein), the barcodes of the non-covalently bound aptamers can be ligated (e.g., covalently) to the POI-RNA complex barcodes. Through sequencing of the ligation products, the aptamer sequences can be obtained, which provide the identity of its target.

[0397] The TURDUCKEN methods described herein allow for:

[0398] a) the in vivo labeling of proteins using nucleic acid barcodes;

[0399] b) the production of large, diverse pools of proteins in a single transformation reaction;

[0400] c) identifying each component of the pool of proteins using NGS sequencing; and

[0401] d) screening for multiple targets in one pooled reaction.

[0402] Other methods of producing DNA-barcoded proteins, such as chemical synthesis, cannot be scaled up and must be performed in single samples or wells. The TURDUCKEN method described herein provides the ability to express thousands to millions of different proteins in the same pool in vivo and add barcodes to them with low protein mislabeling rates. This method saves significant time and money. In addition, the TURDUCKEN method described herein provides the advantage of being able to screen many targets simultaneously at one time.

[0403] It should be understood that procedural modifications can be made to adapt the TURDUCKEN method described herein to different applications. For example:

[0404] • Any organism other than yeast (e.g. E. coli, mammalian CHO cells) can be engineered to produce POI-NA complexes.

[0405] • The nucleic acids used in the TURDUCKEN method described herein can be expressed from a variety of different constructs or vectors (e.g. circular plasmids, linear inserts, or chromosomally integrated DNA).

[0406] • Alternative strategies for linking two entities in vivo to produce POI-NA complexes (e.g. different RNA binding proteins such as MS2 or BoxB / lambda N systems, HUH-endonuclease domains, CRISPR-associated proteins).

[0407] • DNA barcodes can be used instead of RNA barcodes using linker systems such as Spycatcher / Spytag, TALEs, etc.

[0408] There are many potential uses for the in vivo protein labeling provided by the TURDUCKEN method described herein. For example, the TURDUCKEN method described herein can be used to study interactions between molecular targets (e.g. aptamers, small molecules, etc.) for basic or translational research. For example, fluorescent probes that hybridize to POI-DNA complexes can be used to visualize proteins in vivo as a screening tool for drug discovery applications. For example, the TURDUCKEN method described herein can be used to mine aptamers that can then be used as a replacement for antibodies (e.g. as molecular probes, for targeted drug delivery, etc.).

[0409] Figure 31A

[0410] Systematic evolution of exponentially enriched ligands (SELEX) is a biomolecular technique traditionally used to identify aptamers. It is designed to isolate strong binders from large pools of random aptamer candidates because synthesizing such large pools of specific sequences is extremely difficult and expensive. However, if one can generate their own initial SELEX starting aptamer pool, the scenario for SELEX experiments can be specifically adapted, for example, using sequences predicted by ML as targets as the starting aptamer pool. To achieve the generation of such large, diverse, and still controlled or known libraries, a scheme called LEGO has been developed. For 40-merssDNA oligonucleotides, there are 10 24 There are 10 possible oligonucleotides that can be explored, but each SELEX experiment only measures 10 of the total possible experimental space. 8 -10 14 This represents only a small fraction of all possible DNA sequences, making it difficult to find the best aptamer for a specific target in practice, even with optimized experiments. Studies have confirmed the existence of specific two-dimensional or secondary structures, such as G-quadruplexes, often seen in aptamers (Tucker, Shum, & Tanner, 2012), and hypothesized that these secondary structures enhance aptamer binding affinity. The ability to generate an initial input library, rather than being limited to random libraries biased towards popular secondary structures over unstructured aptamers, will accelerate binder discovery. Furthermore, since artificial intelligence prediction algorithms such as ML improve their predictive power, ML-guided input libraries for aptamer experiments will significantly increase the relative proportion of potential aptamer candidates to non-candidates in the initial pool and may reduce the number of rounds required to discover aptamers with equally high affinity. As a result, aptamer candidates can be discovered faster with fewer SELEX rounds, requiring lower discovery costs, and the discovered candidates are less affected by experimental noise such as PCR bias. In other words, fewer downstream quality control assays are needed to confirm that top-ranked aptamer candidates are true binders rather than aptamer candidates that perform excellent PCR but lack specificity for the target of interest. Furthermore, one could consider iterative approaches, in which several rounds of SELEX are performed from a random library, the library is sequenced, and the resulting data is fed into an ML model that predicts what the next initial starting pool should look like (e.g., secondary structure or GC content characteristics or direct sequences), generating new libraries and initiating new, more targeted SELEX experiments.

[0411] Although randomized libraries can be synthesized inexpensively, there are currently no cost-effective methods for generating large pools whose parameters (e.g., GC content, recurring motifs, fixed regions, length, etc.) can be easily determined and manipulated. Current methods for synthesizing short (>200 bp) DNA pools offer the following:

[0412] a) High diversity but little control over sequence content: Random DNA libraries with customizable primer regions can be synthesized at low cost (e.g. under $300, TriLink Biotech). However, it is prohibitively expensive to produce 10 14 thousand specified sequences by conventional microarray synthesis (e.g. Integrated DNA Technologies: $2000 for 1 thousand 200bp long sequences; Agilent: $13,000 for 244 thousand sequences of maximum 90-bp; Twist Biosciences: $46k for 1 million sequences).

[0413] b) High control over sequence content but limited sequence diversity: Research groups have developed methods to construct DNA libraries by piecing together building blocks in a one-pot reaction using 12 base fragments (Fujishima et al., 2015) or sequentially on an immobilized system using 8 base fragments (Horspool et al., 2010). Both of these methods have limitations that limit their use for aptamer library construction.

[0414] The LEGO method described herein allows the construction of computationally derived, customizable DNA libraries, allowing scientists to perform SELEX screening with controlled input pools at reasonable cost. It utilizes commercially available ligases to assemble random 40-mer libraries from sequential ligation of 5-mer or longer DNA LEGO blocks. This can be done in at least two ways: by double-stranded ligation using a dsDNA ligase such as T4 DNA ligase ( Figure 31B ), or by template-independent single-stranded ligation using a ssDNA or ssRNA ligase such as RNA ligase RtcB ( Figure 32 ). In both strategies, ligation starts with the forward PCR primer ligating to the first LEGO block and continues by adding one LEGO block at a time. The final ligation reaction occurs between the last LEGO block and the reverse PCR primer ( Figure 32 A - 32B). After producing the said primer-bearing 40-mer, an amplification method can be performed, such as PCR using a protected forward primer and a phosphorylated reverse primer. The PCR product can be cleaned using any preferred method, and the product with the correct base pair length can be selected using a size selection method such as the automated Pippin HT program. The library can then be converted from double-stranded to single-stranded DNA, for example using lambda exonuclease digestion, and the single-stranded product can be cleaned and concentrated ( Figure 32 C). Relevant information for both RCHT and N-terminal amino acid SELEX experiments D is a schematic demonstrating where the product of the LEGO described herein can be used relative to the SELEX method (Figure 3).

[0415] The method described herein has several unique features that make it best suited for the production of aptamer libraries:

[0416] 1) Unique overhang design allows position control over dsDNA ligation

[0417] Successful ligation between two double stranded DNA fragments requires having complementary single base overhangs on both fragments. A pair of DNA building blocks with compatible overhangs (e.g. A and T, G and C) ligate together preferentially. Building blocks with incompatible overhangs (e.g. A and C, G and T, etc.) ligate together significantly less. By using building blocks with different combinations of A, T, C, and G overhangs, one can control building block positioning. For example, by designing the building blocks such that the overhangs of building blocks 1 and 2 are compatible while the overhangs of building blocks 1 and 3 are incompatible, one can facilitate assembly of the building blocks in the order 1-2-3 but not 2-1-3, 3-1-2, etc.

[0418] 2) Short building blocks allow exploration of the entire DNA space including sequences that are difficult to synthesize

[0419] Using shorter LEGO blocks one can produce libraries that are several orders of magnitude more diverse than libraries produced by other ligation methods. Using a library of 1,024 5-mers one can produce the entire space of 40-mer DNA libraries (10 24 Using a single 1536 well plate one can assemble any 40-mer aptamer or feature interval library that the experiment requires. Furthermore, certain sequences (e.g. long stretches of G) are difficult to synthesize accurately by conventional methods. Splicing together many shorter building blocks provides a useful way to obtain these sequences.

[0420] It should be understood that many modifications can be made to the method described herein. For example:

[0421] • Library design: While the method described herein uses 5-mers to build 40-mers, one can build libraries of different lengths / multiple lengths from building blocks of different lengths / multiple lengths. During DNA synthesis, the 5' phosphorylation rate of short (i.e. <6 nt) oligos is low due to steric interactions from the glass substrate. Increasing the length of the building used will increase the percentage of phosphorylated oligo reagents. However, increasing the length of the oligo blocks requires using larger amounts of different oligo blocks to assemble a library with the desired statistical distribution of sequences.

[0422] • Block design: The methods described herein that use dsDNA use blocks with phosphate group modifications on the 5' end of both strands to facilitate block ligation to the growing strand and to the next block in the sequence. In contrast, blocks with only one 5' phosphorylation can be used to reduce the possibility of flipped DNA blocks being integrated / ligated to the growing sequence. Alternatively, a modification that inhibits ligation can be added on the 5' or 3' strand to prevent ligation of flipped blocks. For ssDNA ligation, the methods described herein use blocks with 3' phosphorylation modifications required by the RtcB enzyme to facilitate this reaction.

[0423] • Starting material: XNA, RNA, modified RNA, single stranded DNA, or modified DNA instead of unmodified double stranded DNA can be used to construct libraries using compatible ligases.

[0424] • Ligation method: There are multiple ways to ligate strands of DNA together. The methods described herein use T4 DNA ligase or RtcB ssRNA ligase to enzymatically ligate DNA construction blocks together. Different ligases (e.g. E. coli DNA ligase, CircLigase, thermostable ligases, etc.) can be used or the construction blocks can be ligated by chemical methods (e.g. click chemistry).

[0425] • Ligation method: Instead of performing a one-pot sequential ligation reaction, several smaller ligation reactions can be performed to produce large blocks and then the products can be combined to ligate the large blocks together. This can improve control over block placement.

[0426] • Medium: Instead of performing library construction in solution, reactions can be performed on beads, on solid supports, in gels, etc.

[0427] • Size selection: In ligating these small DNA blocks together, the ligation products will often not be the desired length. To purify full-length products, manual and automated size selection methods such as the PippinHT automated DNA size selection system can be used.

[0428] In addition, while the methods described herein can be used to produce random libraries for SELEX aptamer screening, the methods described herein can also be used to produce DNA libraries for different applications, for example:

[0429] • Establishing ML derived DNA libraries for peptide / protein production by translation. The priority in the SELEX aptamer selection described herein is to find aptamers specific for their amino acid targets. To this end, the same pool of random aptamers can be incubated with different sequences of peptides. Given that often many different variations of the same sequence need to be tested, it can be quite expensive to obtain all the different peptide sequences that can be needed from a vendor. To expand the space of random peptides available for SELEX, it would be helpful to be able to produce these peptides in-house. The random DNA library production methods described herein can produce these peptide libraries by either cell-free translation kits or conventional DNA plasmid transformation experiments in cells. Promoter sequences can be included in the design of the linker region blocks, or attached after library production, and peptides can be produced from these sequences in vivo or in vitro.

[0430] • Expanding the sequence of the DNA barcodes. A key to protein sequencing is the ability to encode and subsequently read out the amino acid sequence. In many of the protein sequencing methods described herein, DNA barcodes can be used to encode the identified regions of the amino acid sequence. In these methods, when the aptamer binds to the portion of the protein or peptide being sequenced, the DNA barcode region on the aptamer is attached to the growing barcode chain by any suitable ligation method. The enzymatic ligation methods described herein can be used to ligate the barcodes together to form the barcode chain, or to attach the barcodes to a universal linker.

[0431] • Modifying the PROSEQ reagents. In many of the protein sequencing methods described herein, the functional aptamer and processed peptide contain DNA regions such as spacers, barcodes, and ligation consensus regions. For peptides to be sequenced, shorter oligonucleotide adapters (e.g., >6 nt) can be coupled to the amino acid residues to increase the rate of reaction, and then the rest of the DNA elements can be attached in a LEGO-like fashion. For aptamers discovered in SELEX, DNA tails including unique barcodes indicating the aptamer identity, cycle number, and restriction sites, etc. can be directly attached to the 5’ end of the aptamer using a single-stranded ligase such as RtcB. In addition, the binders discovered in SELEX can be modified for direct use on the PROSEQ platform using asymmetric PCR.

[0432] The LEGO methods described herein allow for the production of oligonucleotide libraries that can be tailored to have certain properties (e.g., GC content, recurring motifs, etc.). The diversity of these libraries is orders of magnitude higher than libraries produced by other ligation methods, and can be assembled at a reasonable cost.

[0433] According to the present application, conventional molecular biology, microbiology, biochemistry and recombinant DNA techniques are used. These techniques are within the skill of the art. The present application is further described in the following examples, which do not limit the scope of the subject methods and compositions described in the claims.

[0434] Examples

[0435] Figure 33

[0436] The following will be described:

[0437] A. General methods for all SELEX experiments

[0438] B. RCHT-SELEX experiments

[0439] B.1 General RCHT-SELEX experiment Part I

[0440] B.2 RCHT-SELEX incubation variation

[0441] B.3 General RCHT-SELEX experiment Part II

[0442] B.4 Other components for RCHT-SELEX

[0443] C. RCHT-SELEX results

[0444] D. N-terminal amino acid SELEX experiments

[0445] E. N-terminal amino acid SELEX results

[0446] F. General SELEX protocol

[0447] The general workflow for all SELEX (RCHT-SELEX and N-terminal amino acid SELEX) experiments is shown in Figure 34

[0448] Reagents

[0449] Aptamer libraries were purchased from TriLink Biotechnologies and IDT, all other oligonucleotides were purchased from IDT or synthesized by K&A ​H-8 DNA & RNA synthesizer internal synthesis. All oligonucleotides were purified by HPLC (IDT internal system or internal Agilent 1290 Infinity II). All automation procedures were performed in Agilent Bravo NGS workstation or Opentrons OT-2. All SPRI purifications utilized Mag-Bind TotalPure NGS beads from Omega Biotek. All DNA quantification was obtained using dsDNA and / or ssDNA High Sensitivity Qubit fluorometric quantitation (Thermofisher A9932). All water used was Ambion TM Nuclease-free water.

[0450] Library

[0451] The single-stranded N40 aptamer library consists of 40 random bases, flanked by custom primer regions. To mitigate contamination from overabundant aptamers from past experiments, the primers on the N40 library are switched every 2-3 months. The initial N40 library (TAGGGAAGAGAAGGACATATGATNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNTTGACTAGTACATGACCACTTGA (SEQ ID NO: 1)) was ordered directly from TriLink Technologies. Subsequent custom primers were designed using the Random Sequence Generator tool to generate putative sequences, cross-validated against the internal primer set to avoid overly similar sequences, and then checked for melting temperature and self-dimer and heterodimer using the IDT Oligo Analyzer. Custom primers were also quality checked using a short SELEX cycle before being used for the full SELEX process.

[0452] N40 library used:

[0453] • SELEX N40 library 1 (also referred to as TriLink library): TAGGGAAGAGAAGGACATATGATNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNTTGACTAGTACATGACCACTTGA (SEQ ID NO: 2)

[0454] • SELEX N40 library 2 (also referred to as OMB63): (TTGACTAGTACATGACCACTTGANNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNCACATCAGACTGGACGACAGAA (SEQ ID NO: 3))

[0455] • SELEX N40 library 3 (also known as OMB105 or Wolverine2):

[0456] TGATGCTATGCGACTTATTGTACNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNTACTTGGCGTTCTTACCACCA (SEQ ID NO: 4)

[0457] Peptides

[0458] Biotinylated peptides were synthesized by Genscript. To facilitate the attachment of the peptide to the biotin, all C-terminal residues were lysines. The construct for each peptide was as follows: N-terminal - (2-mer prefix) - (8-mer suffix) - C-terminal - biotin.

[0459] 2-mer prefix: The 20 naturally occurring amino acids were divided into 4 groups of 5 amino acids each. The 2-mer prefix was determined by pairing the amino acids within a block to each other and to amino acids from other groups. Thus each 2-mer prefix belongs to one of 16 blocks (each block has 25 possible 2-mers). There are 400 possible 2-mer prefixes in total. As a reference, the 400 possible prefixes are depicted in Figure 34 A. The 16 blocks are depicted in Figure 34 B.

[0460] 8-mer suffix: For the dipeptide switch experiment, each 2-mer prefix is associated with 2 of 4 possible suffixes. In addition, whether there is a K or C on the end depends on whether the peptide is biotinylated (no DNA oligo attached) or fabricated with an attached DNA oligo (PoC). These suffixes are:

[0461] • A' suffix: ADRWADR(K or C) (SEQ ID NO: 5)

[0462] • B' suffix: MSQPLQP(K or C) (SEQ ID NO: 6)

[0463] • C' suffix: NHFENEI(K or C) (SEQ ID NO: 7)

[0464] • D' suffix: TKYVGTG(K or C) (SEQ ID NO: 8)

[0465] • E' suffix: TAYVETE(K or C) (SEQ ID NO: 9)

[0466] • F' suffix: QGHSIDN(K or C) (SEQ ID NO: 10)

[0467] The two suffixes assigned to each 2-mer prefix are chosen to avoid similarity to the 2-mer prefix. For example, a 2-mer prefix from an AB block should be associated with C' and D' suffixes, but not A' and B' suffixes.

[0468] The suffixes paired with the 2-mer prefixes are alternated between odd and even rounds, only the 2-mer prefix constant peptide combination exerts selection pressure on the aptamer in all 4 rounds Figure 34 C) Examples of suffix and prefix combinations for DD and DC prefix experiments are depicted in Part B. RCHT-SELEX experiment D.

[0469] Example 1 - RCHT-SELEX experiment

[0470] B.1 RCHT-SELEX General Experiment Part I

[0471] Figure 35

[0472] Methods

[0473] SELEX pre-cycling method:

[0474] Incubation

[0475] Depending on the needs of the experiment, incubation is performed by one of three variations. All incubations are performed using 50 microliter PCR reactions using Herculase II fusion DNA polymerase (Agilent Technologies). PCR uses Mag-Bind TotalPure NGS beads (Omega-Biotek) and 100% ethanol added on a Bravo automated liquid handling platform (Agilent) for 0.6X ratio SPRI purification. Amplification conditions for this and all subsequent PCR reactions (except NGS preparation) are as follows: initial denaturation at 95°C for 5 minutes, followed by 13 cycles of 95°C denaturation for 30 seconds, 55°C annealing for 30 seconds, and 72°C extension for 30 seconds, and a final extension at 72°C for 5 minutes.

[0476] To facilitate regeneration of the ssDNA library for aptamer incubation (detailed in the section on digestion), protected and phosphorylated primers are used. For the following primer constructs, * indicates that the nucleotide is modified such that the oxygen atom in the phosphate backbone is replaced with a sulfur atom to obtain a phosphorothioate, which makes the sequence more resistant to nuclease digestion.

[0477] • SELEX N40 library 1 (also known as TriLink library):

[0478] o Forward primer: 5'-T*A*G*G*G*A*AGAGAAGGACATATGAT-3' (SEQ ID NO: 11)

[0479] o Reverse primer: / 5Phos / -TCAAGTGGTCATGTACTAGTCAA-3' (SEQ ID NO: 12)

[0480] • SELEX N40 library 2 (also known as OMB63):

[0481] o Forward primer: 5'-T*T*G*A*C*T*AGTACATGACCACTTGA-3' (SEQ ID NO: 13)

[0482] o Reverse primer: / 5Phos / -TTCTGTCGTCCAGTCTGATGTG-3' (SEQ ID NO: 14)

[0483] • SELEX N40 library 3 (also known as OMB105 or Wolverine2):

[0484] o Forward primer: 5'-T*G*A*T*G*C*TAT GCG ACT TAT TGT AC-3' (SEQ ID NO: 15)

[0485] o Reverse primer: / 5phos / -TGG TGG TAA GAACGCCAAGTA-3' (SEQ ID NO: 16)

[0486] Incubation variations

[0487] Option 1 (mainly used):

[0488] Samples of 10 12 sequences from the single-stranded N40 library (-48 ng) were amplified in 288 reactions of 50 microliters each. The SPRI-purified products of all 288 reactions were combined, giving us a final incubation of 10 12 sequences and with approximately 1200 copies, which was split into 12 SELEX reactions. This method was used to identify aptamers against the biological controls bradykinin, arginine vasopressin, and GnRH, as well as a portion of the dipeptide switch experiment.

[0489] Option 2:

[0490] Samples of two 10 12Samples of the sequence (~48 ng each, ~96 ng total) were amplified in 576 reactions of 50 microliters each. The SPRI purified product of all 576 reactions was combined, giving us a final pool of 2 x 10 12 sequences, which was split into 36 SELEX reactions. This approach provides input pools for most dipeptide switch experiments.

[0491] Option 3: Double incubation:

[0492] Incubation was performed in the manner of Variation 1, but with unmodified primers instead of the protected and phosphorylated versions. An aliquot of the purified incubation (with a diversity of 10 12 sequences) was used as a dsDNA input library for a second incubation using modified primers (Variation 1 or 2). A total of ~48 ng of each dsDNA aliquot was amplified in 288 reactions. Double incubation allows the use of the same input of 10 12 sequences in multiple sets of experiments, far exceeding the usual 12-18 SELEX reactions limited by their distribution.

[0493] Incubation: Incorporation

[0494] Depending on experimental needs, N40 constructs with known sequences were incorporated into the incubation and completed subsequent rounds of SELEX. These sequences are:

[0495] • A6: high_gc_5: TAGGGAAGAGAAGGACATATGATCACCGCATCCTGAGGCCGGTGTGGAGGGCACGAAGTCTGGTTGACTAGTACATGACCACTTGA (SEQ ID NO: 17)

[0496] • C2: high_gc_5: TAGGGAAGAGAAGGACATATGATCTAGCATGGTGCCCTTACCCTCAGAGCGGAAGTACCTGATTTGACTAGTACATGACCACTTGA (SEQ ID NO: 18)

[0497] During the initial incubation, there were ~5.39 million molecules of each incorporator present in each 50 ul reaction, making each incorporator 53,947 times more abundant than the average of random N40 sequences.

[0498] Refolding

[0499] The aptamer library was heated to 95°C for 5 minutes, then cooled on ice for 30 minutes to refold the DNA secondary structure into their lowest energy state.

[0500] Negative selection

[0501] To remove aptamers that would otherwise bind to reagents that are present in the sample throughout the assay, the oligonucleotide library is subjected to negative selection prior to being used as input for SELEX. 166.62 pmol (4650 ng) of refolded ssDNA library is added to 500 ug of streptavidin-coated beads (C1, T1, M270 or M280 depending on the needs of the experiment) and the final volume is made to 400 ul in IX PBS, 0.025% Tween and 10 mg / ml BSA. The reaction is incubated at room temperature (RT) with rotation for 30 minutes, after which the supernatant is collected.

[0502] When using peptide-oligonucleotide conjugates, selection is performed for oligonucleotide tails only. Oligonucleotide tails are incubated with full-length 5’ biotinylated oligonucleotides complementary to the oligonucleotide tails at a 1:2 ratio of tail:complement. A sample containing 1.67 pmol of oligonucleotide tails and 3.34 pmol of complement is then added to 166.62 pmol of refolded ssDNA library that has been previously subjected to negative selection on beads. The reaction is incubated at room temperature (RT) with rotation for 30 minutes, after which 200 ug of streptavidin-coated beads are added and incubation is continued for 30 minutes. The supernatant from this incubation is then collected as the final negatively selected input.

[0503] Digestion

[0504] Amplified libraries are converted to single-stranded DNA (ssDNA) by enzymatic digestion using lambda exonuclease (New England BioLabs) and SPRI purified by automated bead cleanup. ssDNA digestion completion is qualitatively assessed on a Bioanalyzer 2100 (Agilent) using the Small RNA kit (Agilent) and quantitatively assessed for concentration by ssDNA Qubit assay (Thermofisher) after cleanup.

[0505] SELEX cycle method:

[0506] Refolding

[0507] Prior to each SELEX incubation, the aptamer library is heated to 95°C for 5 minutes and then cooled on ice for 30 minutes to refold the DNA secondary structure into their lowest energy state prior to each SELEX incubation.

[0508] B.2 RCHT-SELEX incubation variation

[0509] SELEX incubation:

[0510] There are three variations on how to incubate the peptide with the ssDNA aptamer. With variation 1, the initial SELEX incubation occurs in the presence of streptavidin beads (Variation 1: SsDNA incubated with peptide-bead conjugate); with variation 2, the streptavidin beads are added after most of the incubation is complete (Variation 2: SsDNA incubated with peptide-oligonucleotide target, followed by bead pull-down). With variation 3, the peptide-oligonucleotide target is incubated with a biotinylated primer, followed by the addition of partially double-stranded aptamer (Variation 3: (5') blocked aptamer incubated with peptide-oligonucleotide conjugate, bead pull-down used). See Part C: RCHT-SELEX results .

[0511] In all cases, the ssDNA pool is heated to 95°C for 5 minutes prior to incubation, then quickly cooled on ice. For each reaction, up to 166.62 pmol (4650 ng) of refolded aptamer is added to the peptide or peptide-bead conjugate, and the total volume is brought to 400 ul with IX PBS and 0.025% TWEEN 20 at a final concentration. The final incubation buffer for Variation 3 also incorporates BSA at a final concentration of 10 mg / ml. These buffer conditions can be distinguished as:

[0512] • SELEX Buffer V.1 (also referred to as SELEX Buffer): IX PBS and 0.025% TWEEN 20

[0513] • SELEX Buffer V.2 (also referred to as SELEX Buffer enriched with BSA): IX PBS, 0.025% TWEEN 20, 10 mg / ml BSA

[0514] These buffers are prepared from 10X PBS (Sigma-Aldrich), TWEEN 20 (Sigma Aldrich), and powdered bovine serum albumin (Sigma Aldrich).

[0515] Variation 1: SsDNA incubated with peptide-bead conjugate

[0516] Conjugation of peptide to beads

[0517] After deciding on the concentration gradient for the SELEX experiment, the peptide target on beads can be manufactured in advance in a large batch to avoid round-to-round error from multiple conjugations. Beads can be frozen and thawed once without any experimental defects. Aliquots for each round are manufactured and stored at -20°C in Eppendorf LoBind or Nunc plates until thawed for use. To ensure similar properties, unit tests are performed on freshly conjugated beads and frozen beads and compared, and no differences are found. The amount of target produced should be based on the number of rounds, the starting concentration of the first round, and buffer stock in case of experimental disaster. In this example, a starting ratio of 1 : 10 target:DNA aptamer is used. Using a Bravo automated liquid handling platform (Agilent), 18.5 pmol of peptide is mixed with 87.2 ug (8.72 ul of 10 mg / ml stock) of MyOne streptavidin Cl beads (ThermoFisher) and incubated for 30 minutes. After an additional 2 washes with SELEX buffer, the initial mix of 18.5 pmol of peptide and 87.2 ug of beads is resuspended in 50 ul of SELEX buffer. These numbers are scaled up proportionally to produce a large volume bead conjugate stock that can be aliquoted and frozen at the beginning of each experiment. For a 1 : 10 target: ssDNA stringency experiment, 50 ul of this stock can be added to 4650 ng of input ssDNA, and scaled down to smaller volumes directly for experiments using less than 4650 ng of input ssDNA. For experiments using a higher stringency of 1 : 25, the volume of added peptide-bead conjugate is scaled down further using a factor of 0.6X.

[0518] M280 or T1 beads blocked with BSA or M270 or Cl beads unblocked are used depending on the needs of the experiment. M280 and M270 beads have a diameter of 2.7 um and Cl and T1 beads have a diameter of 1 um. Unit tests confirmed that the Cl beads, which the manufacturer indicated were optimal for automation, pulled down different aptamer sequences from the incubation than the M280, M270, and T1 beads. The mechanism of this result is unknown. As a result of the unit tests, M280 beads are chosen for the next step of the experiment because BSA blocking is preferred to prevent selection of aptamers that bind to the bead surface, and a larger surface area target can provide a platform to place individual peptides further apart, reducing selection of aptamers that prefer dimerization of peptides.

[0519] Blank bead "conjugate" was produced by placing a mixture of beads and water in the same automated Bravo protocol, with a total of 30 minute incubation and 2-3 wash cycles. The initial input of 87.2 ug of beads was also resuspended in 50 ul SELEX buffer and added to the ssDNA later at a ratio of 87.2 ug beads per 4650 ng ssDNA (for a 1 : 10 stringency reaction) or 34.88 ug beads per 4650 ng ssDNA (1 :25 stringency reaction).

[0520] SELEX incubation

[0521] Up to 50 ul of bead conjugate was added to 166.62 pmol (4650 ng) of folded aptamer and incubated at RT with rotation for 2 hours.

[0522] Streptavidin-biotin pull down

[0523] Streptavidin M280 beads (Invitrogen) were added to the SELEX incubation at a ratio of 83.33 ug of beads per 51.02 pmol of peptide present, for 30 minutes with rotation.

[0524] Variation 2: SsDNA incubation with peptide-oligo and aptamer incubation followed by bead pull down

[0525] Peptide conjugation

[0526] For this variation, no conjugation is required prior to incubation. The target is a peptide-oligo.

[0527] SELEX incubation

[0528] The amount of target added depends on the stringency gradient desired. Typically, for small molecule targets, stringency conditions ranging from 1 : 1 to 1 : 10 (target: ssDNA) are used as starting conditions, with the ratio between target and DNA increased in subsequent rounds until sequencing data confirms enrichment of the aptamer. Here, the method used for a protocol starting with a 1 : 10 target: ssDNA is described. For the 1st and 2nd rounds, 166.62 pmol (4650 ng) of folded aptamer was added directly to 18.51 pmol of peptide-oligo construct for a 1 : 10 target: ssDNA stringency. Considering the decreased stringency of 1 : 25 in the 3rd and 4th rounds, 166.62 pmol (4650 ng) of aptamer was added directly to 7.40 pmol of peptide. The peptide and ssDNA were incubated at RT with rotation for 2 hours.

[0529] Streptavidin-biotin pull down

[0530] In the case where the target has a DNA oligonucleotide tail, a biotinylated primer that anneals to a portion of the oligonucleotide tail (5' Biotin TAGGGAAGAGAAGGACATATGAT 3' (SEQ ID NO: 19)) is added to the SELEX incubation at a 1 :2 peptide: biotinylated oligonucleotide ratio for every 51.02 pmol of peptide present for 30 minutes with rotation. The primer serves two functions: (1) to prevent the aptamer from binding to the DNA oligonucleotide tail, and (2) to allow for the pull down of the target by performing a biotin-streptavidin reaction after incubation.

[0531] Streptavidin M280 beads (Invitrogen) are then added to the SELEX incubation at 83.33 ug for every 51.02 pmol of peptide present for 30 minutes with rotation. After the incubation with the beads, which allows the biotin-streptavidin reaction to complete, the beads are pulled down (either manually or automatically) using a magnet, washed, and prepared for PCR.

[0532] Variation 3: Incubation of (5') blocked aptamer with peptide-oligonucleotide conjugate, pull down using beads

[0533] Incubation solution preparation (POC and biotinylated primer incubation)

[0534] In addition to blocking the region of the tail of the peptide-oligonucleotide conjugate (POC), a portion of the aptamer can also be blocked to prevent unwanted binding between the primer region of the aptamer and the region of the DNA tail on the POC. The POC is added to the 5' biotinylated primer that is complementary to the length of the oligonucleotide tail at a 1 :2 POC: biotinylated primer ratio. 10X PBS, TWEEN-20, BSA, and water are added to make a final 265 ul solution and IX PBS, 0.025% TWEEN-20, and 0.1509 mg / ml BSA per reaction. The entire solution is incubated at RT with rotation for 30 minutes.

[0535] The amount of POC input for each reaction is determined by the expected aptamer input. An exemplary method for a 1 : 10 target: ssDNA stringency round is set forth below. For rounds 1 and 2, 18.5 pmol POC is prepared for an input of 166.62 pmol (4650 ng) of aptamer, resulting in a 1 : 10 target: ssDNA stringency. In this particular gradient, after two rounds of 1 : 10 stringency, the next two rounds are accelerated to a 1 : 25 stringency to increase the signal of the enriched aptamer. It should be noted that too rapid of a stringency increase or too high of a starting stringency can result in the loss or disappearance of the true aptamer signal. However, too slow of a stringency increase or a starting stringency that does not create competition between binders can result in a loss of time and resources due to the additional SELEX rounds required before enrichment can be observed. In this example, given the target reduction required for a 1 : 25 stringency in rounds 3 and 4, the amount of POC prepared for an input of 166.62 pmol (4650 ng) of aptamer is reduced to 7.40 pmol.

[0536] SELEX incubation

[0537] The peptide and ssDNA are incubated with rotation at RT for 2 hours. The final incubation buffer for a 400 ul reaction is IX PBS, 0.025% TWEEN 20, and BSA at a concentration matching the hybridization buffer used in the BCS experiment (see Example 3 - ProSeq experiment and Example 4 - BCS binding assay experiment, varied in the range of 0.10 mg / ml - 10 mg / ml).

[0538] POC control

[0539] For the negative control for SELEX variation 3, the aptamer is incubated with only the oligonucleotide tail of the POC and not the peptide.

[0540] Possible oligonucleotide tails for this purpose are as follows:

[0541] • / 5phos / cttagatgcacgtggataATCATATGTCCTTCTCTTCCCTA (SEQ ID NO: 20)

[0542] • / 5phos / cttagatgcacgcagcatATCATATGTCCTTCTCTTCCCTA (SEQ ID NO: 21)

[0543] Streptavidin-biotin pull down

[0544] M280 beads (Invitrogen) were added to the SELEX incubations in an amount of 83.33 ug per 51.02 pmol of peptide present for 30 minutes under rotation.

[0545] B.3 RCHT-SELEX General Experiment Part II

[0546] Post SELEX Cycling Method:

[0547] Post Incubation Wash (Applicable to all transformations)

[0548] Bead-peptide-aptamer conjugates were collected on Bravo using an automated wash protocol. Each SELEX reaction was incubated on a magnetic plate for 2 minutes. The supernatant containing unbound aptamer was aspirated and the beads were washed twice with SELEX buffer and finally with IX PBS. The IX PBS was aspirated at the end of the protocol.

[0549] PCR on Beads

[0550] Immediately after the automated wash protocol was complete, 50 ul of PCR solution was added to each well containing beads. An unmodified version of the primers used to incubate the aptamers were used to amplify the 86 nt construct, except for the Wolverine2 library where the construct was 84 nt long (full library construct provided in the description of the library above).

[0551] NGS Preparation

[0552] Following PCR amplification on the beads, the DNA concentration was measured by Qubit dsDNA assay and 10 ng of the SPRI purified PCR sample on the beads was taken for NGS preparation. Each aptamer identified from sequencing of these samples had a 6 bp barcode for the peptide they were assumed to bind to in solution. The P5 and P7 adapters required for Illumina sequencing were incorporated by PCR using custom NGS primers (5'-CAAGCAGAAGACGGCATACGAGATNNNNNNNN- (forward primer)-3') (SEQ ID NO: 22) and 5'-AATGATACGGCGACCACCGAGATCTACACNNNNNN- (reverse primer)-3') (SEQ ID NO: 23). The forward and reverse primer regions were variable depending on the N40 library used for SELEX. The amplification conditions for these PCR reactions were as follows: initial denaturation at 95 °C for 5 minutes followed by 10 cycles of denaturation at 95 °C for 30 seconds, annealing at 65 °C for 30 seconds and extension at 72 °C for 30 seconds, and finally extension at 72 °C for 5 minutes. The final NGS library was SPRI purified, pooled and size selected for the 177 bp construct by PippinHT (Sage Science).

[0553] Threshold PCR

[0554] For each SELEX reaction, 4.08 ng of SPRI purified product from the bead PCR was amplified in 24 50ul PCR reactions using modified primers tailored to each library (sequences provided in the incubation section). The SPRI purified dsDNA product of this library is an 86-bp (or 84-bp for the Wolverine2 library) amplicon with the same construct as the original N40 library with protected and phosphorylated ends to facilitate enzymatic digestion of the reverse strand. The regenerated ssDNA library serves as input for the next round of SELEX.

[0555] SELEX cycle

[0556] The protocol steps between aptamer refolding, target selection, aptamer incubation, unbinders separation, wash, amplification, NGS sample extraction, threshold amplification, ssDNA library generation, and refolding can be repeated as a “SELEX round” until enriched aptamers are found in the NGS sequencing data. Incubation and initial negative selection are not repeated between rounds.

[0557] B.4 Additional components of RCHT-SELEX

[0558] False SELEX

[0559] During the first 2 hours of SELEX variation 2, the negative control is incubated with water and SELEX buffer only. After each round of SELEX, samples from the false SELEX are sequenced to determine the impact of PCR bias (as no enrichment should occur due to the lack of target). False SELEX can be used for computational analysis and ML modeling of aptamers to train the model to focus on enrichment signals of aptamer counts rather than operator error, contamination, PCR bias, or other experimental or instrument noise.

[0560] BCS compatible aptamer preparation

[0561] Application of BCS or DNA aptamers in ProSeq requires modification of the primer region of the aptamer to include the correct ligation, restriction enzyme, and spacer sequences to facilitate binding and recording events in BCS. However, unique barcodes are not required because sequencing can be performed across the entire aptamer sequence in order to record which aptamer bound to which target on the BCS chip. There are several ways to convert an aptamer library into a BCS compatible library, however the fastest, cheapest, and highest throughput method is to use PCR to modify the primer region of the aptamer. To do this, add the ssDNA pool (up to 166.62 pmol per reaction) to 23nt oligonucleotides “bridge mimics” that are complementary to the forward primer region of the aptamer at a 1 : 10 aptamer: bridge mimic ratio. Bring the solution up to 135ul solution with IX PBS and 0.25% TWEEN 20. Heat the mixture to 95°C for 5 minutes, quick chill on ice, then add to the incubation solution.

[0562] For the SELEX N40 library 3 (also known as OMB105, Wolverine2), the library has the following construct

[0563] • 5’ TGATGCTATGCGACTTATTGTAC NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNTACTTGGCGTTCTTACCACCA 3’ (SEQ ID NO: 24)

[0564] And the following forward primer

[0565] • 5’ TGATGCTATGCGACTTATTGTAC 3’ (SEQ ID NO: 25)

[0566] The bridge mimic used was 5’ GTACAATAAGTCGCATAGCATCA 3’ (SEQ ID NO: 26).

[0567] Bead-based multiplexed SELEX

[0568] This assay is almost identical to SELEX, with the difference that multiple peptides are added to each reaction. Peptides are coupled to beads individually at the beginning of the experiment and aliquoted into individual stocks, which are mixed in equimolar proportions at the beginning of the SELEX incubation. The first four rounds are processed by the usual incubation / threshold PCR, digestion, incubation, automated wash and PCR cycle on beads. To de-multiplex in the last round, N*4.08 ng of each reaction resulting from PCR on beads is amplified in N*24 reactions, where N is the number of peptides incubated simultaneously with the aptamer pool. The SsDNA from this reaction is incubated in each SELEX reaction at a stringency of 1 :50, with only one peptide present in each reaction.

[0569] After washing away unbound aptamers using Bravo's automated wash protocol, 50ul of PCR solution is added to each de-multiplexed well. The SPRI purified product of each of these PCR reactions is barcoded and sequenced during NGS preparation to reveal the aptamer associated with each peptide isolated.

[0570] Primer Switching

[0571] Custom primers flanking the N40 region are switched between different rounds and replaced with alternative primer sequences. The purpose of this primer switching is to reduce contamination by overabundant aptamers from experiments using the same N40 library.

[0572] The current primer switching design is designed for TriLink N40 libraries. By amplifying the initial N40 construct (5' TAGGGAAGAGAAGGACATATGAT NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNTTGACTAGTACATGACCACTTGA (SEQ ID NO: 27)) with primers TriLinkFwd_Fokl (5' TAGGGAAGAGGGATGAAGGACATATGAT (SEQ ID NO: 28)) and TriLinkRev_Fokl (5' TCAAGTGGTCGGATGATGTACTAGTCAA (SEQ ID NO: 29)), a Fokl restriction site is introduced to create a new full-length construct (5' TAGGGAAGAGGGATGAAGGACATATGAT NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNTTGACTAGTACATCATCCGACCACTTGA (SEQ ID NO: 30)).

[0573] This altered PCR product was digested with the nuclease Fokl (NEB) cutting 9 and 13 bp downstream of its restriction sites (5'...GGATG(N)9 / 3'...CCTAC(N) 13 (SEQ ID NO: 31)) we cut off (5' TAGGGAAGAGGGATGAAGGACATA (SEQ ID NO: 32) and 5' TTGACTAGTACATCATCCGACCACTTGA (SEQ ID NO: 33)) leaving sticky ends. End filling of this construct with Klenow fragment (NEB) resulted in a blunt end. Incubation of this blunt ended double stranded library with new double stranded primers and ligase completed the protocol leaving us with the original N40 library with the new primer set swapped in.

[0574] Plate layout

[0575] To minimize the effect of local contamination between adjacent wells, technical replicates (3 per experimental condition) were spatially randomized across different rows and / or different plates. For the dipeptide switch experiments, no technical replicates were adjacent to each other. This allowed for computational filtering of noise during post-sequencing analysis.

[0576] Figure 36

[0577] Incubation

[0578] For incubation, 96 unit tests were performed to determine the optimal incubation conditions for each library, defined as the conditions that introduced the least bias or variance in the expression levels of all possible 6-mers after incubation was performed. The expression intensity of each combination of possible 6-mers from sequencing runs of the DNA pool after incubation was divided by the expression intensity before incubation. For the OMB63 library, the optimal conditions that caused the least variance in the expression levels of each combination of 6-mers were 11 PCR amplification cycles, using Herculase II fusion DNA polymerase and 0% DMSA, and an input of 10 10 DNA molecules Figure 37 .

[0579] Pseudo-SELEX

[0580] The top 20 sequences from the 100,000 sequences randomly sampled from the pseudo-SELEX samples and real SELEX rounds were confirmed to be different, indicating that the DNA pool after SELEX incubation was altered as a result of bead-coupled targets rather than pulling down random sequences Figure 9). The pseudo-SELEX analysis can be used to determine PCR bias elements during the SELEX experiment and also to train the model towards the background truth values of positive aptamer signals.

[0581] digestion

[0582] The bioanalyzer small RNA kit trace shows a single clear peak at approximately 75 nt after the digestion process, which correlates with the desired ssDNA product size (86 bp for most SELEX libraries) considering the measurement error in the technique Figure 11B C). The confirmation of the full conversion of the dsDNA PCR product to ssDNA occurs before each aptamer library is introduced into each new round of SELEX.

[0583] threshold PCR

[0584] Unit tests showed that the threshold PCR introduced minimal bias. Comparison of the sequencing data of the DNA before and after the threshold PCR run showed that the threshold PCR introduced low variance in the sequence distribution between the pools before and after the threshold PCR (variance of the log ratio was 0.132) Figure 11C and Figure 38 ).

[0585] parallel experiments

[0586] Parallel experiments were performed between experiments of the same target with aptamer sequences from the same incubation until round 5, providing higher confidence in the identified aptamers. The wells in which the chymotrypsin and GNRH experiments were performed were physically adjacent on the same plate. Significant bleed between the targets chymotrypsin and GNRH was detected in the biological control SELEX experiment, allowing the detection of spatial contamination Biological controls ). Therefore, randomization of the sample placement was performed on each plate, with different targets placed on the same row and without intervals between each experiment, and parallel samples of the same target were placed with a distance of 2 columns between each parallel sample to reduce contamination. After significance evaluation, it was found that the observed contamination was a result of reagents being carried out from the automation.

[0587] aptamer

[0588] Figure 39

[0589] As a proof-of-concept for the RCHT-SELEX method, DNA aptamers for arginine vasopressin (peptide sequence: CYFQNCPRG{LYS(biotin)}(SEQ ID NO: 34)) and bradykinin (peptide sequence: RPPGFSPFR{LYS(biotin)}(SEQ ID NO: 35)) were identified as having high binding affinity. The equilibrium dissociation constant (K0) was estimated based on the experimental conditions of SELEX incubation. d The value is 45nM ( Peptide switching The aptamers can be further characterized to determine K with and without primers. d The N40 binding region sequence of the aptamer identified for each target is:

[0590] Arginine vasopressin: 5'-ATATTCTAGGTTGGTAGGGAAGGCATGTATCTAATTCCTG-3' (SEQ ID NO: 36)

[0591] ·Bradykinin: 5'-CAAATCGGTGCCGGCCGGGAAGGGGCAAAAACAGTGCAAC-3' (SEQ ID NO: 37)

[0592] During RCHT-SELEX, the two aptamers have the following primers on both sides:

[0593] • Forward primer: TAGGGAAGAGAAGGACATATGAT (SEQ ID NO: 38)

[0594] • Reverse primer and reverse complement: TTGACTAGTACATGACCACTTGA (SEQ ID NO: 39)

[0595] In three parallel experiments for each target, the same culture was assayed against arginine vasopressin and bradykinin; the identified sequences yielded consistent results in experiments using the same target, but no consistent results were obtained in experiments using different targets. These findings suggest that these aptamers may be specific aptamers for arginine vasopressin and bradykinin, and could be used to detect these targets in samples.

[0596] Figure 40

[0597] In the A-peptide switching assay, the sequence was sequentially enriched for specific N-terminal amino acids. Part D: N-terminal amino acid SELEX experiment The report identifies the top-ranking aptamers for lysine and cysteine, defined as the aptamer with the highest sequence count after noise filtering. During RCHT-SELEX, both sets of aptamers were flanked by the following primers:

[0598] • Forward primer: TAGGGAAGAGAAGGACATATGAT (SEQ ID NO: 40)

[0599] • Reverse primer reverse complement: TTGACTAGTACATGACCACTTGA (SEQ ID NO: 41)

[0600] Further experiments can be performed to characterize and validate the identified aptamers for use in protein sequencing.

[0601] Example 2 - N-terminal amino acid SELEX

[0602] Figure 41

[0603] Reagents

[0604] DNA libraries were purchased from TriLink Biotechnologies, all DNA primers were purchased from Integrated DNA Technologies and HPLC purified. All peptides were purchased from Genscript. 10X PBS and Tween-20 were purchased from Sigma-Aldrich. Lambda exonuclease and buffer were purchased from New England Biolabs. Mag-Bind Total Pure NGS beads were purchased from Omega-Biotek. Bioanalyzer and all reagents, Bravo liquid handler, and Herculase II Phusion polymerase and buffers were purchased from Agilent. Test tubes, plates, and thermocycler were purchased from Eppendorf. Nunc plates were purchased from VWR. 70% and anhydrous ethanol were purchased from Fisher Scientific. Nuclease-free water, MgCl2, bovine serum albumin, dNTP mix, Dynabeads M280 streptavidin, and QuBit reagents were purchased from Thermo Scientific.

[0605] Methods

[0606] In this example, aptamers specific for the dipeptide proline-proline (PP) were isolated using an N-terminal amino acid SELEX method (see Example 1). Round). Twelve selections were run in parallel for a total of 5 targets (2 targets of interest and 3 control targets). 3 selections were run for each target of interest and 2 selections were run for each control target. All positive selection rounds were sequenced and used for enrichment analysis across rounds and targets. In addition, automation was used in several steps to ensure minimal potential error across samples and to facilitate running parallel selections. For this experiment, the dipeptide PP was chosen as the N-terminal dipeptide of interest because its bulky cyclic side chain allows for multiple potential binding sites. The PP target is a 10-mer peptide with a region of two prolines at the N-terminus and 8 other amino acids (“scaffold”), followed by a C-terminal coupled biotin (biotinylated target) or DNA tail (PoC target). To improve the chances of isolating aptamers specific for the N-terminal PP dipeptide, both “switched” and “unswitched” targets were utilized, with multiple selections for each. For PP targets with a C-scaffold (“unswitched”), the target is referred to as PP-C, or for PP targets with a D-scaffold (“unswitched”), referred to as PP-D. If both targets are used in a selection (“switched”), they are referred to as PPCD.

[0607] Target-bead coupling

[0608] Target-bead coupling was performed fresh before each round of incubation. Biotinylated peptide targets were coupled to M280 streptavidin beads using an Agilent Bravo liquid handling platform. For each coupling reaction, beads were vortexed to homogeneity, then 25 uL of beads were added to an appropriate volume of 75 ng of peptide target. Beads and target were incubated on a chilled plate (4°C) for 2 minutes to allow biotin and streptavidin to interact and form a tight bond, then the beads were washed several times with SELEX buffer (1x PBS, 0.025% Tween-20, 0.1 mg / mL BSA, 1 mM MgCl2). The final product of the bead coupling reaction was resuspended in 50 uL of SELEX buffer.

[0609] Negative SELEX

[0610] DNA aptamer generation was performed using a protocol involving aptamers in solution and biotinylated targets coupled to streptavidin beads. An initial library of 10 15 aptamers was taken from the library stock and subjected to a 30 minute negative selection against 10 mg / mL streptavidin beads in 50 ul of SELEX buffer. The supernatant was retained and placed directly into the positive selection against the peptide target. This positive selection was the first step of 5 rounds of SELEX using the following workflow: selection, amplification (small-scale PCR and large-scale PCR), and single-stranded generation.

[0611] Positive SELEX

[0612] Prior to each selection step, aptamers were annealed in refolding buffer (lx PBS, 0.025% Tween-20, ImM MgCl2) at 95°C for 5 minutes and incubated at room temperature (RT) at 22-24°C for at least 30 minutes. Selections were performed in SELEX buffer with rotation for 30 minutes (negative selection) or 1 hour (positive selection). Stringency reports for each round for "switched" and "no switch" are reported in Table 2.1.

[0613] Table 2.1 Stringency as a function of round and target type

[0614] "Tight" stringency "no switch" "Tight" stringency "switch" "Switch" backbone Figure 42 1 1:1 1:1 C 2 1:2 1:1 D 3 1:5 1:2 C 4 1:10 1:2 D 5 1:25 1:5 C

[0615] Amplification was performed in two steps: small scale PCR and large scale PCR. After washing away non-binders, the remaining target-aptamer conjugate was directly placed in a small scale PCR reaction, 1 reaction (50 uL) per sample. PCR reaction conditions included all DNA retained from the washing step, 3 uM forward primer, 3 uM reverse primer, Herculase buffer, 0.2 mM DNTP, 0.0.5 units / L Herculase polymerase, in a final volume of 50 uL.

[0616] After cleaning the PCR reaction, an aliquot of the product was placed in a large scale PCR, 24 reactions of 50 uL each. The purpose of this large scale PCR was to amplify as much DNA as possible without introducing too much PCR bias. PCR reaction conditions included 0.17 ng DNA, 6 uM forward primer, 6 uM reverse primer, lx Herculase buffer, 0.2 mM DNTP, 0.5 units / uL Herculase polymerase, in a final volume of 50 uL.

[0617] Both small and large scale PCR were performed using a Mastercycler Nexus using the following conditions: 95°C 5 min, 13 cycles of 95°C 30 seconds, 55°C 30 seconds, and 72°C 30 seconds, and 72°C 5 minutes. PCR reactions used dNTPs from Omega Bio-Tek TotalPure NGS beads purification and was performed using an Agilent Bravo liquid handling platform. ssDNA and TotalPure NGS beads were incubated at a 3:5 ratio and cleaned with 70% ethanol.

[0618] To generate single stranded DNA from the large scale PCR product, a lambda exonuclease digestion was performed for an optimized time. Digestion was qualitatively followed using a bioanalyzer. The cleaned digest was quantified and used as input for the next selection.

[0619] NGS preparation and sequencing

[0620] Samples after SELEX rounds were prepared for sequencing. The samples were normalized to a concentration of 10 ng / ul. A 50 ul PCR reaction was set up for each sample (2 ul 6.25 uM forward and reverse primers, 10 ul 10 ng / ul DNA sample, 36 ul master mix) to amplify the DNA and the reaction was performed using a Mastercycler Nexus (PCR conditions: 98°C 5 min, 10 cycles of 98°C 30 sec, 65°C 30 sec, 72°C 30 sec, and 72°C 5 min). After the reaction, the PCR product was cleaned (Agilent Bravo Liquid Handling Platform). The size of the PCR product was then quantified using a Tapestation to determine if the PCR reaction was successful. The samples should have a DNA size of 170-190 bp. The concentration of the PCR product was determined using a qubit dsDNA assay. The PCR products were then pooled in a tube according to the concentration of each product. The concentration of the pooled product was determined using a qubit dsDNA assay. The PCR product was purified (Pippin Prep system, Sage Science) by selecting a DNA size of 177 bp. The concentration of the purified product was determined using a qubit dsDNA assay. After purification, 10 uL of the purified product was finally sent for NGS sequencing.

[0621] Analysis

[0622] A rapid increase in enrichment of all targets was observed from round 2 to round 3, and a plateau was reached from round 3 to round 5 ( Figure 43 ). In addition, for aptamers binding to Bradykinin, PP-C and PP-CD targets, a logarithmic enrichment value of around 3.5, 3.2 and 3.0 respectively was observed, indicating that these targets have putative binders ( Figure 43 A). To further examine these binders, the top 10 binders obtained from enrichment were drawn for each parallel of each target ( Figure 44 B). The enrichment clusters of binders for each target were clustered between experimental parallels, indicating that the selection on these targets is isolating binders of interest. Further analysis of experimental parallels of binders for a target indicated that the overlap between binders in different parallels was overall low ( Part E: N-terminal SELEX results ). Due to the size of the initial random pool, the chance of having the same sequence in different experimental parallels or for different targets is low, indicating that they are rather contaminant sequences, allowing to filter out these possible contaminant sequences before testing. These candidates were further filtered down to a short list of candidates to test binding properties in vitro.

[0623] To identify the final aptamer sequences for full characterization, two filtering steps were performed. Candidates from the PP-CD target with high enrichment (greater than 2, which correlates to at least 100-fold improvement from R2 to R5) and selective binding to PP-CD (binders that do not bind to other targets) were selected. Filtering the candidate sequences resulted in 26 candidates, of which 10 were selected for final testing. The 10 final candidates were selected on the basis of highest enrichment ratio, total sequencing count, representation within each selection replicate, and zero sequence contamination in selection replicates.

[0624] Enrichment calculation (formula defining growth and penalty growth):

[0625] The number of times a given aptamer sequence appears in the sequencing data set is the aptamer count. The "pre" and "post" rounds of SELEX were defined as subsets of the sequencing data to track unique aptamer sequences. "Pre" is the subset from round 2, and "post" is the subset from round 5. A log scaling factor was applied to each aptamer count to accommodate the wide range of aptamer counts from 0 to 10 5

[0626] Pre = log 10 (pre ct + 1)

[0627] Post = log 10 (post ct + 1)

[0628] Growth was defined as the enrichment of a given aptamer between the "pre" round (round 2) and the "post" round (round 5).

[0629] Growth = Post - Pre = log 10 [(pre ct + 1) / (post ct + 1)]

[0630] Penalty was calculated for sequences with low count numbers in both round 2 and round 5, multiplying the original penalty value by a factor γ, and applying it to the growth factor by subtracting the product of γ and the original penalty.

[0631]

[0632] γ = 1.26

[0633] Penalty Growth = Growth - γ x Original Penalty

[0634] Technical details: If pre < c, c can be substituted in the formula, where:

[0635]

[0636] K d Measurement

[0637] 200 pmol peptides (PP-C, PP-D) were coupled to 100 μL of Dynabeads according to the manufacturer's instructions. TM M-280 streptavidin (Thermo Scientific) was resuspended to its original concentration in SELEX buffer. 5 mg of fluorescein biotin (Biotinium, #80019) was resuspended in DMSO. 650 pmol of fluorescein biotin was conjugated to 100 μL of Dynabeads according to the manufacturer's instructions. TM M-280 streptavidin (Thermo Scientific) was used as a positive control and resuspended to the original concentration. FAM-tagged aptamer candidates #1-10 at the 5' end were purchased from IDT. Aptamers were synthesized using complementary forward and reverse primers and tested to be of full length. The full sequence of each aptamer is as follows: 5'-TTGACTAGTACATGACCACTTGA-N40-TTCTGTCGTCCAGTCTGATGTG-3' (SEQ ID NO: 42). The N40 sequences of the tested aptamers are reported in Table 2.2.

[0638] Table 2.2 Tested aptamer candidate sequences

[0639]

[0640] Dilute the peptide-conjugated beads to 0.03 mg / mL or 1:320 of the original concentration for use in the binding assay. Aliquot 100 μL of diluted peptide-conjugated or fluorescein-conjugated beads into each well of a 96-well plate. Place the plate on a magnetic rack for 2 minutes and remove the supernatant. Add 100 μL of 5'-terminal FAM-labeled aptamer candidates diluted in SELEX buffer at different concentrations (0, 100 nM, 250 nM, 500 nM, 750 nM, 1 μM, 2.5 μM, 5 μM, 10 μM, 20 μM) to the appropriate wells. Seal the plate with plate sealant (AB 0558 adhesive PCR film, ThermoFisher) and incubate in the dark at room temperature for 1 hour. After incubation, remove the sealant, wash the beads three times with 100 μL of SELEX buffer, and resuspend them in 100 μL of SELEX buffer. The beads were transferred to a black plate and a single-endpoint fluorescence reading was measured using a Biotek plate reader.

[0641] Note that this is a combination assay to measure K. dOne method of producing even more accurate measurements is microscale thermophoresis, bio-layer interferometry, flow cytometry, and surface plasmon resonance.

[0642] Figure 45

[0643] By the above plate-based K d The aptamers were tested for their measurement methods. At a single concentration (100 nM), 7 aptamers showed higher fluorescent signals for the target PP-D compared to the control (no aptamer, just buffer). One aptamer showed higher fluorescent signals for the target PP-C compared to the control ( Figure 46A ). Two aptamers, aptamer 1 and aptamer 4, were selected for further testing. Aptamer 1 showed possible saturation binding for PP-C but no specific binding for PP-D ( Figure 46B ). Aptamer 4 showed saturation binding for PP-D but no binding for PP-C ( Part: General ).

[0644] F Protocol SELEX Negative selection

[0645] The above lists various methods that were used, optimized, and utilized in order to obtain aptamer binders from SELEX results, however, for each application of the SELEX described herein: (1) RCHT-SELEX for ML analysis or (2) N-terminal binder aptamer using NTAA-SELEX, there are different combinations of methods used. Below is a template protocol that can be used to decipher the combination of methods needed.

[0646] Overall workflow:

[0647] 1. Bead coupling

[0648] 2. Amplification

[0649] 3. Single-stranded generation of the eluate / antisense digestion

[0650] 4. Incubation

[0651] 5. Amplification from the incubated beads

[0652] 6. Threshold amplification

[0653] 7. Single-stranded generation of the threshold amplification / antisense digestion

[0654] 8. Counterselection

[0655] 9. Gradient 1

[0656] Device protocol:

[0657] 1. Qubit: Qubit was used to measure DNA concentration according to manufacturer's protocol.

[0658] 2. Bravo: Three types of protocols were run on the Bravo liquid handler: (1) PCR cleanup ("bulk" and "variable volume"), (2) bead coupling ("bead coupling"), and (3) bead washing ("post-SELEX wash but not elution"). For PCR cleanup, the Bravo was programmed to follow the manufacturer's guidelines for using Mag-Bind TotalPure NGS. For bead coupling, the Bravo was programmed to follow the manufacturer's guidelines for using Dynabeads TM M-280 streptavidin. Incubation time and buffer were optimized for coupled peptides. For bead washing, the Bravo was programmed to perform 3 washes of the peptide beads (after incubation with aptamer). The plate was incubated on a magnet for 2 minutes. The first two washes were performed using SELEX buffer, and the last wash was performed using lx PBS. After the last wash, the beads were not resuspended but left in the plate for the next step of the SELEX protocol.

[0659] 3. Bioanalyzer: Two types of protocols were run on the Agilent 2100 Bioanalyzer with 2100 Expert software. Library quality check and post-PCR quality check were performed using the High Sensitivity DNA chip according to the manufacturer's instructions using the High Sensitivity DNA protocol. Post-digestion / single-stranded generation quality check was performed using the Small RNA chip according to the manufacturer's instructions using the Small RNA II Series protocol.

[0660] Table 2.3 SELEX stringency gradient:

[0661] R1 R2 R3 R4 R5 n / a 1:10 1:10 1:25 1:25 Gradient 2 Target 1:5 1:10 1:25 1:50 1:100

[0662] SELEX buffer: IX PBS, 0.025% tween-20, 1 mM MgCl2, 0.1 mg / mL BSA, nuclease free H2O

[0663] Technical terms:

[0664] Fwd RC: Reverse complement of the 5' end of the forward aptamer. This is a mimic of the bridge used in the BCS because it makes the 5' end of the aptamer double-stranded.

[0665] POC: Peptide-oligonucleotide conjugate: This is the SELEX Peptide primer, biotinylated primer, DNA tail complement, blockerTarget: The substance we are looking for the aptamer binder to bind to. The POC produces a 10-mer peptide and a 41 nt ssDNA tail.

[0666] bt Peptide Oligonucleotide Complement: Also known as Segment Suffix The fragment is the complement of the ssDNA "tail" region of the peptide- oligonucleotide conjugate (POC). This fragment has biotin on the 3' side to bind to streptavidin beads and is a complete "blocker" of the oligonucleotide tail of the POC. It is incubated with the POC in a 2: 1 ratio, then the target is incubated with the aptamer.

[0667] Tail: Refers to the DNA tail of the peptide conjugated into the PoC (but can be used alone without the peptide attached).

[0668] Backbone: Also known as Figure 47 This is the 8-mer region on the dipeptide target (both the biotinylated target and the PoC) between the N-terminal dipeptide and the C-terminal conjugated biotin (biotinylated target) or DNA tail (PoC target). The backbone is named according to the following convention: [letter]' (e.g. C' or D').

[0669] Stringency: This corresponds to the ratio of target: aptamer. For example, a stringency of 1:10 means there is 10 aptamer sequences for every 1 target, conversely, a stringency of 10:1 means there is 10 targets for every 1 aptamer. 10:1 is not very stringent, while 1:100 is very stringent.

[0670] Positive selection: Selection where aptamers are incubated with their target, pulled down, and the supernatant (containing non-binders) is discarded.

[0671] Negative selection: Selection where aptamers are incubated against a random surface (test tube sidewall, beads, etc.) and the supernatant (containing sequences that do not bind to the random surface) is retained.

[0672] Counterselection: Selection where aptamers are incubated against a substance that is closely similar to the target (e.g. a different dipeptide or just the backbone) and the supernatant is retained.

[0673] Workflow

[0674] Negative selection (just beads or beads + tail)

[0675] Purpose: To eliminate aptamers from the library that have high binding affinity to the beads.

[0676] 1. Dilute input ssDNA (10 15 molecules) in refolding solution (1X PBS, 0.025% Tween-20, 1 mM MgCl2, NF H2O). Total volume is 150 uL.

[0677] 2. Anneal (refold aptamer): Heat to 95°C for 5 minutes and cool on bench for 30 minutes.

[0678] 3. Wash 55 uL of 10 mg / mL M280 beads in 500 uL SELEX buffer 3 times. Resuspend in 55 uL SELEX buffer.

[0679] 4. In a 1.5 mL low-binding tube, spin incubate 50 uL of washed M280 beads in 200 uL of modified SELEX buffer (1X PBS, 0.025% Tween-20, 1 mM MgCl2, 0.16 mg / mL BSA, NF H2O) with cooled annealed library solution (150 uL) for 30 minutes.

[0680] 5. Place the tube on a magnetic stand and wait for 1 minute for the beads to fully concentrate near the magnet.

[0681] 6. Remove the supernatant (~200 uL) and transfer to a new tube.

[0682] 7. Measure the DNA concentration using the Qubit ssDNA kit. A typical expected concentration is in the range of 8-20 ng / uL.

[0683] Bead coupling

[0684] Purpose: To couple biotinylated peptide targets to streptavidin beads, which magnetically pull down the aptamer binders during incubation.

[0685] Note: The peptide-bead couplings can be made in advance and aliquoted in a 96-well eppendorf plate for freezing (up to 1 freeze-thaw cycle), or made fresh before each incubation for fresh use.

[0686] 1. Dilute the stock peptide to an appropriate concentration so that the peptide and beads can be combined at a ratio of 200 pmol of peptide target to 1 mg of Dynabeads M280 beads (per manufacturer's protocol).

[0687] 2. Pipette the appropriate amount of peptide and water into each well of a 96-well eppendorf plate to a volume of 50 uL.

[0688] 3. Pipette the appropriate amount of M280 streptavidin beads into the NUNC plate, filling only the wells that will be used.

[0689] 4. Run the "Bead Coupling" protocol using the liquid handler. This performs the incubation, mixing, and washing steps as defined by the manufacturer.

[0690] 5. Dilute the peptide beads to the appropriate stringency, divide into aliquots and store at -20°C.

[0691] Amplification (incubation)

[0692] Purpose: To generate multiple copies of each aptamer of the negatively selected library.

[0693] 1. Prepare master mix using 50mL conical tubes. Master mix: 3uM forward primer, 3uM reverse primer, Herculase buffer, 0.2mM dNTPs, 0.5 units / uL Herculase polymerase in a 16000uL final volume (this is a total of 320 reactions, 50uL per reaction). Each 50uL reaction should have 0.17ng DNA.

[0694] 2. Aliquot the master mix into 3 x 96 well plates, 50uL per reaction.

[0695] 3. Seal the 96 well plates and place in a thermocycler using the following PCR protocol: 95°C 5min, (95°C 30sec, 55°C 30sec, 72°C 30sec) x 13 cycles, 72°C 5min, hold at 4°C.

[0696] 4. Combine the 3 plates into 1 plate of 150uL reactions.

[0697] 5. Clean up using "large volume" protocol on a liquid handler. This uses the manufacturer's protocol for Mag-Bind TotalPure NGS beads.

[0698] 6. Combine the incubations into 1 x 5mL eppendorf low bind tube.

[0699] 7. Measure the concentration of double stranded DNA using QuBit dsDNA kit to check the concentration. Typically, the concentration is in the range of 40-90ng / uL.

[0700] Single strand generation (digestion of incubation)

[0701] Purpose: To digest the antisense strand of the double stranded DNA using lambda exonuclease. In order for the aptamer to be able to bind to the target, ssDNA must be generated.

[0702] 1. Set up single strand generation reaction according to lambda exonuclease (M0262, NEB) manufacturer's instructions (for a 50uL reaction, use up to 5ug DNA, 5uL 10x reaction buffer, 1uL lambda exonuclease and up to 50uL H2O). First add 10x reaction buffer to the DNA, vortex mix. Next add lambda, pipette mix.

[0703] 2. Depending on the DNA input concentration, incubate the reaction for 10-20 minutes at 37°C.

[0704] 3. Heat inactivate the exonuclease by incubating at 72°C for 10 minutes, hold at 4°C.

[0705] 4. Check the quality of the DNA after digestion by running the DNA product on a bioanalyzer small RNA kit following the manufacturer's protocol. If the trace shows that there is still double stranded product, add the same amount of lambda exonuclease as the original reaction and incubate at 37°C for an additional 5-10 minutes. Check the quality again.

[0706] 5. Pool the DNA in one plate and clean on a liquid handler using the "variable volume" protocol. This uses Mag-Bind TotalPure NGS beads following the manufacturer's protocol.

[0707] 6. Check the DNA concentration using the QuBit ssDNA kit. The concentration is usually around 30 ng / ul or higher.

[0708] PoC target incubation - no bead conjugation

[0709] Purpose: Incubate the aptamer library with the target to see which aptamers bind to the target.

[0710] This incubation is only for PoC targets where the PoC is exposed to the aptamer before the beads are introduced and pulled down. For any protocol using bead conjugation, use biotinylated targets for the incubation.

[0711] 1. Dilute the input ssDNA (10 15 molecules) and FWD RC / bridge if used in refolding solution (1X PBS, 0.025% Tween-20, 1 mM MgCl2, NF H2O). Total volume is 150 uL.

[0712] 2. Anneal (refold aptamer): heat to 95°C for 5 minutes and cool on the bench for 30 minutes.

[0713] 3. Target tail blocking incubation: Incubate the target with the bt peptide oligo complement primer at a 1 :2 ratio in a total volume of 250 uL of modified SELEX buffer (1X PBS, 0.025% Tween-20, 1 mM MgCl2, 0.16 mg / mL BSA, NF H2O) in a sealed NUNC plate for 30 minutes with rotation. The target concentration will vary with the stringency gradient.

[0714] 4. Selection Incubation: Combine 150 uL of cooled ssDNA in refolding solution with 250 uL of target and annealed biotinylated peptide oligonucleotide complement in modified SELEX buffer for a total volume of 400 uL and incubate with rotation for 1 hour in a sealed NUNC plate.

[0715] 5. Isolation / Pull-down Incubation: Pre-wash M280 beads 3 times in SELEX buffer and resuspend at original concentration in SELEX buffer. Add beads to 400 uL selection incubation reaction and incubate for 30 minutes.

[0716] 6. Use liquid handler to wash off non-specifically bound or unbound DNA from target beads (protocol: "wash but not elute").

[0717] Biotinylated target incubation - using bead coupling

[0718] Purpose: Incubate our aptamer library with target to see which aptamers bind to target.

[0719] This incubation protocol should be used for any target (biotinylated or PoC) that is coupled to beads prior to the start of SELEX. Note that in this protocol the aptamer is exposed to target and beads, as opposed to the "PoC target incubation" protocol, in which the PoC is exposed to aptamer, then beads are introduced and pulled down.

[0720] 1. Dilute input ssDNA (10 15 molecules) in refolding solution (1X PBS, 0.025% Tween-20, 1 mM MgCl2, NF H2O). Total volume is 150 uL.

[0721] 2. Anneal (refold aptamer): Heat to 95°C for 5 minutes and cool on bench for 30 minutes.

[0722] 3. Thaw frozen bead coupling plate and add modified SELEX buffer (1X PBS, 0.025% Tween-20, 1 mM MgCl2, 0.16 mg / mL BSA, NF H2O) to a total volume of 250 uL.

[0723] 4. Combine 150 uL of cooled ssDNA in refolding solution with 250 uL of bead target coupling in modified SELEX buffer for a total volume of 400 uL and incubate with rotation for 1 hour in a sealed NUNC plate.

[0724] 5. Use liquid handler to wash off non-specifically bound or unbound DNA from target beads (protocol: "wash but not elute").

[0725] Amplification (PCR from beads [PoB])

[0726] Objective: Amplify the aptamer bound to the target using PCR. At this point, the aptamer is still bound to the target and all non-specific DNA has been washed away.

[0727] 1. Immediately after the wash protocol, add master mix (3 uM forward primer, 3 uM reverse primer, Herculase buffer, 0.2 mM DNTP, 0.5 units / uL Herculase polymerase, final volume of 50 uL) to the wells to avoid bead drying.

[0728] 2. Transfer to Eppendorf low binding 96 well plate, seal and place in thermal cycler using the following PCR protocol: 95 °C 5 min, (95 °C 30 sec, 55 °C 30 sec, 72 °C 30 sec) x 13 cycles, 72 °C 5 min, hold at 4 °C.

[0729] 3. Clean using "variable volume" protocol on liquid handler.

[0730] 4. Measure concentration of double stranded DNA using QuBit dsDNA kit on plate reader to check concentration. Typical concentration is in the range of 4-20 ng / uL.

[0731] Threshold PCR

[0732] Objective, amplify the aptamer library using protected primers (forward primer with 6 thiol sulfates, reverse primer with 5' phosphate).

[0733] 1. Prepare master mix using 50 mL conical. Master mix: 3 uM forward primer, 3 uM reverse primer, Herculase buffer, 0.2 mM dNTP, 0.5 units / uL Herculase polymerase, final volume of 16000 uL (this is 320 reactions total, 50 uL per reaction).

[0734] 2. Make 1:10 dilution of PoB DNA and normalize input concentration by pipetting 0.17 ng dsDNA per 50 uL reaction. Prepare stock solutions for each sample by adding 4.3 ng dsDNA, 300 uL H2O, and 954 uL master mix to each well. Aliquot each sample stock solution into 50 uL per reaction.

[0735] 3. Seal plate and place in thermal cycler using the following PCR protocol: 95 °C 5 min, (95 °C 30 sec, 55 °C 30 sec, 72 °C 30 sec) x 13 cycles, 72 °C 5 min, hold at 4 °C.

[0736] 4. Clean DNA using "large volume" protocol on liquid handler.

[0737] 5. Measure concentration of double stranded DNA using QuBit dsDNA kit on plate reader to check concentration. Typically, concentration is in the range of 30-90 ng / uL.

[0738] Single stranded regeneration (digestion of threshold PCR product)

[0739] Purpose: To generate ssDNA for next round of SELEX. This needs to be done as multiple reactions, as each selection has a different DNA concentration.

[0740] 1. Set up single stranded generation reactions according to lambda exonuclease (M0262, NEB) manufacturer's instructions (for 50 uL reactions, use up to 5 ug DNA, 5 uL 10x reaction buffer, 1 uL lambda exonuclease, and up to 50 uL H2O). First add 10x reaction buffer to DNA, vortex mix. Next add lambda, pipette mix.

[0741] 2. Incubate reactions at 37°C for 10-20 minutes, depending on DNA input concentration. Group reactions on different plates depending on reaction time.

[0742] 3. Heat inactivate the exonuclease by incubating at 72°C for 10 minutes, hold at 4°C.

[0743] 4. Check quality of DNA after digestion by running DNA product on bioanalyzer small RNA kit according to manufacturer's protocol. If trace shows that double stranded product is still present, add same amount of lambda exonuclease as original reaction and extend 37°C incubation by 5-10 minutes. Check quality again.

[0744] 5. Pool DNA on one plate and clean using "variable volume" protocol. This uses Mag-Bind TotalPure NGS beads according to manufacturer's protocol.

[0745] 6. Check DNA concentration using QuBit ssDNA kit. Typically, the concentration is around 30 ng / uL or higher.

[0746] Counter selection

[0747] Purpose: Incubate target against other targets that are closely similar to the target in one or more aspects to ensure that the aptamer being enriched is specific and actually binds to the target itself. This is very similar to positive selection, with the difference being that the target is different and there is no "wash but not elute" step.

[0748] 1. Depending on the experiment, the aptamer is refolded and incubation is set up according to the PoC or biotinylation incubation steps listed above.

[0749] 2. After incubation, the plate is placed on a magnet for 2 minutes to allow all beads to be magnetically concentrated.

[0750] 3. The supernatant is removed from each well and stored in a clean eppendorf 96 well PCR plate.

[0751] 4. The DNA concentration is measured using the Qubit ssDNA kit.

[0752] NGS preparation

[0753] The PoB DNA from round 2 onwards is sequenced. The samples are prepared using the NextSeq protocol (NGS preparation).

[0754] Additional protocols:

[0755] Bioanalyzer check after digestion (small RNA kit):

[0756] The purpose of the bioanalyzer test is to verify that the dsDNA from the incubation / threshold PCR has been effectively digested to ssDNA by the lambda exonuclease. The small RNA kit is used according to the manufacturer's instructions.

[0757] To analyze the results of the bioanalyzer assay, look for the position of the ssDNA and dsDNA peaks. The ssDNA peak is at 60 seconds and the dsDNA peak is at 40 seconds. If there are concatemers, they are observed at 55-65 seconds (a broad, uneven peak). Digestion is complete when a sharp peak is seen at 60 seconds. See Figure 48 .

[0758] dsDNA bioanalyzer check

[0759] The purpose of this bioanalyzer test is to assess the quality of the dsDNA after PCR / after incubation + clean up according to size (base size). We use the high sensitivity DNA kit according to the manufacturer's instructions.

[0760] To analyze the results of this assay, look for the lower marker at 35 bp and the upper marker at 10380 bp. Check the match of the aptamer length to the expected library length (in this example 86 bp). See ​ .

[0761] Example 3 - PROSEQ experiment

[0762] The following will be described below:

[0763] Part A: ProSeq experimental methods

[0764] Part B: ProSeq results

[0765] Part C: General ProSeq protocol

[0766] Part A: ProSeq experimental methods

[0767] Reagents

[0768] Aptamers and substrate oligonucleotides were purchased from IDT or synthesized by K&A TE H-8 DNA & RNA synthesizer was used for internal synthesis and purification by HPLC (Agilent 1290 Infinity II). Peptide-oligonucleotide constructs bradykinin, arginine vasopressin and GNRH were purchased from Genscript. Aptamer incubation and later DNA barcode sequencing was performed on NextSeq or MiSeq kits supplemented with PhiX Control v3 and sequenced on MiSeq500 (Illumina). Bound aptamers were ligated to barcode substrates using T4 ligase (Blunt / TA Master Mix preparation) and cut with EcoRI in buffer and all reagents were purchased from New England Biolabs. Excess aptamer and hybridization buffer was washed away with For Edman degradation, peptides were coupled with phenylisothiocyanate (PITC) in coupling buffer (0.4 M dimethylallylamine in pyridine:water 3:2 (v / v), pH 9.5), cleaved in trifluoroacetic acid (TFA) and dried under a stream of nitrogen. All reagents for Edman degradation were purchased from Sigma-Aldrich. All buffers were diluted with Ambion TM Nuclease-free water. Analysis of NGS data was done using a custom analysis pipeline running on the Colaboratory notebook environment.

[0769] Methods

[0770] Protein sequencing

[0771] Construction of substrates and immobilization to solid substrates

[0772] The core sequencing unit is composed of 4 independent DNA fragments: a 5' phosphorylated barcode base (BF), forward and reverse co-localization adapters (FC and RC) and a protein or peptide target (PT) with a C-terminal oligonucleotide sequence tag attached to the protein or peptide at the 3' end with a free phosphorylated 5' end orientation. The 5' end of the BF sequence is complementary to the 5' end of FC to allow hybridization, while the 3' end of BF contains a unique barcode (for sample multiplexing or related PT identification) and a short consensus sequence complementary to the bridge sequence to facilitate aptamer attachment to BF. The FC contains a BF-complementary region at the 5' end, followed by a sequence complementary to the glass-bound oligonucleotide, followed by a flexible T-spacer, and a short high GC content sequence complementary to RC at the 3' end. In turn, the 3' end of RC is complementary to the 3' end of FC, followed by a long T-spacer, followed by a sequence complementary to the glass-bound oligonucleotide, followed by a sequence complementary to the PT-binding oligonucleotide. Likewise, the 5' end of the PT oligonucleotide is complementary to the 5' end of RC, followed by a spacer Figure 49

[0773] The four fragments are then combined and hybridized in solution, such that the PT is attached to the unique BF through the FC and RC, allowing for PT identification (in the case of confirmatory and spike-in controls) or sample demultiplexing (in the case of multiple pools of peptides being sequenced simultaneously). Following hybridization, the four-component complex is incubated on an oligonucleotide-seeded glass substrate. The FC and RC hybridize to the glass-bound oligonucleotide, and through the addition of DNA ligase, the BF and PT oligonucleotides are covalently attached to the glass-bound oligonucleotide through ligation (in this case, "nick repair" ligation). In this way, the BF-PT pair is co-localized and spatially separated from all other BF-PT pairs to ensure that a given PT's binding event is confined to a single BF. Furthermore, the covalent attachment of BF and PT to the glass facilitates the maintenance of BF and PT co-localization after multiple rounds of PT sequencing, despite the harsh reagents required for PT degradation. Once the BF and PT are covalently attached to the glass-bound oligonucleotide, formamide is used to wash away the forward and reverse co-localization adapters that were annealed to the BF and PT.

[0774] Aptamer incubation

[0775] ​After the BF and PT are covalently attached to the substrate, the sequencing process begins as follows: a pool of first BCS-compatible aptamers is incubated, then unbound aptamers are washed away and ligase is added to covalently attach the aptamers to the BF. This cycle of incubation and ligation is performed multiple times, with ligation occurring after each incubation or after all aptamer pools have been introduced. Prior to incubating the peptide target with the aptamers, a pool of single-stranded aptamers is incubated with the bridge oligonucleotide to form a library of BCS-compatible aptamers. It should be noted that only a single barcode is recorded between cycles of restriction digestion (as described below). After ligation, a restriction enzyme (along with an excess of sequence complementary to the restriction site and spacer) is introduced to cut the peptide-binding sequence of the aptamer off the aptamer barcode on the 5' end, leaving only the aptamer barcode and a short consensus sequence for subsequent ligation attached to the BF. After restriction, the PT is degraded from the N-terminus using Edman degradation, aminopeptidase, or any other processive degradation process. Clearly, the technique of building a barcode sequence encoding an aptamer can be applied equally to C- to N-terminal peptide or protein sequencing, as the barcode sequence synthesis process is independent of the orientation of the PT on its oligonucleotide tether. Furthermore, multiple rounds of aptamer incubation, ligation, and restriction can be used prior to PT degradation to interrogate the same N-terminal amino acid sequence multiple times, allowing for more accurate identification of the N-terminal composition.

[0776] After degradation, another pool of aptamers is incubated and the process is repeated. The aptamers in each round contain unique barcodes (even if the peptide-binding sequence is the same) so that missed incorporation events (e.g. obvious deletions) can be easily identified and accounted for in subsequent data analysis steps.

[0777] DNA barcode construct sequencing

[0778] The final step in the sequencing process is the addition of PCR or next-generation sequencing (NGS) adapters. Using the same consensus and bridge sequences, adapters are ligated to the 3' end of the aptamer barcode sequence, which represents a series of aptamer binding events, in turn used to determine the sequence of the PT. Using the glass-bound oligonucleotide sequence and / or the 5' sequence of the BF as one primer and the PCR / NGS adapter as the other primer, the barcode construct is amplified from the chip and sequenced using standard NGS techniques, or, in the case where the NGS sequencing flow cell serves as the PT sequencing platform and the NGS adapters are suitably designed, the barcode construct is amplified and sequenced directly on the NGS flow cell without further processing.

[0779] Sup-Diff

[0780] A priori Sup-Diff

[0781] Biotinylated RNA bait production

[0782] Prior Sup-Diff was performed on pools of BCS barcode constructs. Initial NGS datasets revealed sequences with high read counts as targets to be stripped by Sup-Diff. The targets were manufactured separately from IDT or by in-house K&A H8 DNA synthesizer with other pool constituents. The target sequences were PCR'd using a standard forward primer and a reverse primer containing a T7 RNA polymerase promoter sequence. The PCR product was cleaned (~1-2ug) according to the automated Bravo clean-up protocol and then used as a template to generate complementary biotinylated RNA baits by in vitro transcription in a 20ul TranscriptAid T7 High Yield Transcription Kit (Thermo Scientific) reaction containing 10mM ATP, CTP and GTP, 7.5mM UTP and 2.5mM Biotin-16-UTP (Roche). After 4-6 hours at 37°C, the DNA template and unincorporated nucleotides were removed by DNase I (NEB) treatment and RNeasy Mini Kit column filtration (Qiagen).

[0783] Hybridization in solution and bead pull-down

[0784] The mixture containing the target pool and nuclease-free water was heated at 95°C for 5 minutes, cooled on ice for 2 min, then mixed with the biotinylated RNA bait and SUPERase In RNase inhibitor (Invitrogen) in pre-warmed (65°C) 2X hybridization buffer (10X SSPE, 10X Denhardt's, 10mM EDTA and 0.2% SDS). After 16 hours at 65°C, the hybridization mixture was added to MyOne C1 streptavidin Dynabeads (Invitrogen) that were washed three times and resuspended in 2X B&W buffer (10mM Tris-HCl (pH 7.5), 1 mM EDTA, 2M NaCl). After 30 minutes at RT, the beads were pulled down and the supernatant was retained.

[0785] “Soup” processing and sequencing

[0786] The supernatant (“soup”) was treated with a cocktail of two RNases, RNase H (NEB) and RNase A (Zymo) for 30 minutes at 37°C. The treated ssDNA was then amplified for 18 or more cycles. Initial denaturation was 95°C for 5 min. Each cycle was 95°C for 30 seconds, 55°C for 30s and 72°C for 30s. Final extension was 72°C for 5 min. The Bravo cleaned PCR product was then prepared for NGS to sequence on an Illumina Miseq using custom primers.

[0787] Non-A priori Sup-Diff

[0788] There can also be cases where a non-priori version of Sup-Diff can be desirable. In this case, samples of the target pool can be used as templates for in vitro transcription (IVT). As a proof of concept, IVT optimization was performed to make the baits in the RNA bait pool representative of a bias towards high abundance species.

[0789] RNA bait pool generation

[0790] A gradient of SELEX incorporated sequences (mass %) was generated: sequence 9 (0.000125%), sequence 13 (0.01%), sequence 11 (1%), sequence 12 (10%), sequence 10 (88.98%). This ssDNA gradient pool was used as a template in a 20ul TranscriptAid T7 High Yield Transcription Kit (Thermo Scientific) reaction containing 0.1 mM, 0.25 mM, 1 mM, 2.5 mM, or 10 mM rNTPs (no biotinylated UTP). After 4-6 hours at 37°C, the DNA template and unincorporated nucleotides were removed by DNase I (NEB) treatment and RNeasy Mini Kit column filtration (Qiagen).

[0791] Reverse transcription

[0792] The purified RNA bait pool was then reverse transcribed into cDNA using the Maxima Reverse Transcription Kit (Thermo Fisher). A 28ul initial reaction contained 500ng RNA bait pool, 15-20pmol TriLink forward primer, 0.5mM dNTP equimolar mix, and nuclease-free water, which was incubated at 65°C for 5min. 8ul 5X Reverse Transcriptase Buffer, 2ul SUPERase In RNase Inhibitor (Invitrogen), and 2ul Maxima Reverse Transcriptase were then added, and the reaction was incubated at 50°C for 30min, followed by heat inactivation at 85°C for 5min. The resulting cDNA pool was treated with a cocktail of two RNases, RNase H (NEB) and RNase A (Zymo) for 30min at 37°C.

[0793] Amplification and sequencing

[0794] The treated ssDNA was then amplified for 13 or more cycles. Initial denaturation was 95°C for 5 min. Each cycle was 95°C for 30 seconds, 55°C for 30 seconds, and 72°C for 30 seconds. Final extension was 72°C for 5 min. The Bravo cleaned PCR product was then prepared for NGS using custom primers for sequencing on an Illumina Miseq. A 41x8x6 readout was performed using the Miseq V2 Nano kit.

[0795] Part B: ProSeq Results

[0796] Results - Proof of concept for barcode sequence synthesis

[0797] As a proof of concept for the synthesis of DNA barcodes representing a series of binding events, and in turn, a hypothetical amino acid sequence of a protein or peptide to be sequenced, the barcode synthesis process was performed using a "mock aptamer" DNA-DNA binding (e.g., hybridization) system. In this way, the uncertainties in binding kinetics and binder-target specificity were reduced to create an "ideal" binder-target system in which to demonstrate the serial barcode addition strategy. Furthermore, these DNA-DNA binders can be used as internal controls in future experiments to assess overall run quality.

[0798] Using this idealized platform with barcode-specific bridges, up to 12 cycles of aptamer barcode ligation and restriction were performed with an ...

Claims

1. A method of producing a barcoded polypeptide, the method comprising: transforming an expression construct into a microbial cell under conditions in which each cell incorporates one construct, wherein the expression construct comprises nucleic acid encoding: (a) a fusion protein comprising the polypeptide, a purification tag, and a nucleic acid binding protein; and (b) a nucleic acid sequence recognized by the nucleic acid binding protein and a unique nucleic acid barcode; and culturing the microorganism under conditions in which the construct is expressed and the nucleic acid binding protein portion of the fusion protein binds to the nucleic acid binding protein recognized sequence, thereby producing the barcoded polypeptide in vivo.

2. The method of claim 1, wherein the microbial cell is selected from a eukaryotic or prokaryotic cell.

3. The method of claim 1, further comprising purifying the barcoded polypeptide.

4. The method of claim 1, wherein the expression construct comprises any copy number of an origin of replication compatible with the host organism.

5. The method of claim 1, wherein expression is driven by any combination of constitutive, inducible, or repressible promoters compatible with the host organism.

6. The method of claim 1, wherein components (a) and (b) are expressed using different promoters.

7. The method of claim 1, wherein components (a) and (b) are expressed using the same promoter present at different locations within the expression construct.

8. The method of claim 1, wherein components (a) and (b) are expressed using Gal 1,10-bidirectional promoter, ADH1, GDS, TEF, CMV, EF1a, SV40, T7, lac, or any other promoter and promoter combination compatible with the host organism.

9. The method of claim 3, wherein the purification step comprises pull down of the barcoded polypeptide using a pull down method corresponding to the encoded purification tag.

10. The method of claim 9, wherein the pull down step comprises pull down of the barcoded polypeptide with protein purification magnetic beads.

11. The method of claim 10, wherein the protein purification magnetic beads comprise at least one of an anti-His antibody, agarose, and nickel.

12. The method of claim 10, further comprising elution of the barcoded polypeptide from the protein purification magnetic beads using a mild elution buffer to release the fusion peptide without denaturing the RNA-protein / peptide binding.

13. The method of claim 12, wherein the mild elution buffer is glycine.

14. The method of claim 1, wherein the polypeptide comprises one or more site-specific protease cleavage sites for release of the barcoded polypeptide from the anti-affinity tag beads using a site-specific protease.

15. The method of claim 14, wherein the site-specific protease comprises at least one of enterokinase, factor Xa, tobacco etch virus protease, and thrombin.

16. The method of claim 1, wherein the nucleic acid sequence comprises a restriction enzyme cleavage site so as to release the barcoded polypeptide from the anti-affinity tag bead using a restriction endonuclease.

17. The method of claim 1, wherein the nucleic acid sequence recognized by the nucleic acid binding protein and the nucleic acid binding protein are an MS2 RNA hairpin or a variant thereof and an MS2 phage coat protein or a mutant thereof.

18. The method of claim 1, wherein the nucleic acid recognized by the nucleic acid binding protein and the nucleic acid binding protein are a boxB sequence or a variant thereof and a phage antitermination protein N (λN).

19. The method of claim 3, wherein the cells are irradiated with UV radiation prior to purification of the barcoded polypeptides.

20. The method of claim 3, wherein the purified complex is irradiated with UV radiation.

21. A DNA barcoded polypeptide or protein made by the method of any one of claims 1-20.

Citation Information

Patent Citations

  • Single-cell proteomic assay using aptamers

    US20180320224A1