Methods for generating and screening compartmentalized peptide libraries

The co-compartmentalization of cyclic peptides with their encoding polynucleotides using microfluidic techniques addresses the limitations of existing methods, enabling efficient screening and selection of cyclic peptides for pharmaceutical assays, particularly functional assays, suitable for drug discovery.

JP7748797B2Active Publication Date: 2025-10-03UNIV OF SOUTHAMPTON
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2019546242
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-02-23
Filing Date
2018-02-22
Publication Date
2025-10-03
Estimated Expiration
2038-02-22

AI Technical Summary

Technical Problem

Existing methods for generating and screening cyclic peptide libraries face challenges in compatibility with pharmaceutical assays, particularly functional assays, and require complex operational skills, limiting their applicability and efficiency.

Method used

A method for co-compartmentalizing cyclic peptides with their encoding polynucleotides using microfluidic techniques, allowing for passive or active cyclization within compartments, followed by screening and sorting based on activity, enabling high-throughput identification and selection of cyclic peptides suitable for pharmaceutical assays.

Benefits of technology

Enables easy and unique identification of cyclic peptides compatible with pharmaceutical assays, facilitating high-throughput screening and selection of promising compounds for drug discovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007748797000001
    Figure 0007748797000001
  • Figure 0007748797000002
    Figure 0007748797000002
  • Figure 0007748797000003
    Figure 0007748797000003
Patent Text Reader

Abstract

A method for co-compartmentalizing a cyclic polypeptide and a polynucleotide encoding the cyclic polypeptide, comprising the steps of: a) forming a compartment comprising a polynucleotide encoding the cyclic polypeptide; b) expressing the polypeptide from the polynucleotide; and c) cyclizing the polypeptide. Co-compartmentalized cyclic polypeptides and encoding polynucleotides. Libraries of co-compartmentalized cyclic polypeptides and encoding polynucleotides. Methods for screening libraries of co-compartmentalized cyclic polypeptides and encoding polynucleotides. Incorporation of non-standard nucleic acids into such libraries.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] 1. Technical Field The present invention relates to the co-compartmentalization of cyclic peptides with polynucleotides encoding them. In particular, the present invention relates to libraries of such compartmentalized cyclic peptides and methods for screening, selecting, and sorting compartmentalized cyclic peptides from such libraries. [Background technology]

[0002] 2. Background of the Invention Compartmentalizing individual samples within aqueous droplets dispersed in an oil phase is a powerful method for high-throughput assays in chemistry and biology. In this case, the droplets are equivalent to test tubes. They contain everything needed to evaluate and interpret a particular experiment or profile of library members. Droplets generated by bulk emulsion techniques are not uniform in size, presenting complex challenges for experiments requiring quantitative readouts. Microfluidic devices can generate highly monodisperse aqueous droplets at rates up to thousands (tens of thousands) per second. They typically range in diameter from 10 to 200 microns, corresponding to volumes of 0.5 pL to 4 nL. In addition to droplet formation, microfluidic formats enable numerous other unit operations, such as droplet fission, fusion, incubation, analysis, and sorting.

[0003] Droplet microfluidics involves the formation of nanoliter to femtoliter sized oil-in-water droplets that provide well-defined, individualized environments particularly suited to the compartmentalization of single genes, organisms, and cells. Miniaturization in particular offers many advantages, including reduced sample consumption, improved analytical performance, rapid mixing of contents, and laminar (streamlined) flow.

[0004] Aqueous droplets link genotype to phenotype because the compartmentalization they provide resembles that of natural cells. In each artificial compartment, a single gene is transcribed and translated by cell-free means, resulting in multiple copies of its encoded protein. This in vitro compartmentalization (IVC) generates "monoclonal units" that are well suited for high-throughput library selection and the identification of novel peptide-based biologics. Because the link between genotype and phenotype is co-compartmentalized rather than chemically linked, a variety of functions (enzymatic, regulatory, inhibitory, binding, structural) can be selected for. This offers an advantage over existing display selection strategies (e.g., bacterial, viral, yeast, and phage display), which select solely based on binding interactions. Because binding does not correlate with inhibition or catalysis, these strategies can inadvertently select weak inhibitors that bind strongly to the enzyme surface but lie outside the enzyme's active site.

[0005] Studies on in vitro protein expression in emulsified water-in-oil (w / o) droplets often use green fluorescent protein (GFP) as proof of concept. The realization of a microstructured device for cell-free protein synthesis compartmentalization and GFP expression was first demonstrated by Dittrich et al. Subsequently, Courtois et al. described the utility of an integrated chip system for the storage and online detection of droplets containing measurable amounts of in vitro-expressed GFP. However, isolation of only a single template may not produce enough protein to reach the detection threshold. As in the case of a single lacZ gene, β-galactosidase expression was undetectable (kJ / k). cat =187s -1, KM = 150 μM). This problem can be overcome by co-encapsulating DNA with an isothermal amplification system. On-chip isothermal amplification with Φ29 DNA polymerase easily generates μg quantities of double-stranded DNA. This pre-amplification step involves the coordinated fusion of IVTT-containing droplets with droplets containing amplified DNA. However, the optimal conditions required for DNA amplification, IVTT, and subsequent enzymatic assays often differ; for example, components required for successful DNA amplification may inhibit IVTT. Similarly, it may be necessary to perform each process at different pH values ​​or buffer compositions.

[0006] Compatibility between different biochemical reactions is an issue when working with water-in-oil droplets. Most biochemical protocols are multistep processes, involving the addition of solutions, centrifuging the sample, removing the supernatant, washing the pellet, and so on. Transferring such protocols to droplets is challenging. Even if droplet-to-droplet fusion for adding and / or diluting components can be performed using specialized microfluidic devices, it is practically challenging when multiple droplets are involved and additional steps (e.g., washing, buffer exchange, adding or removing reagents) are required. Furthermore, while droplet fusion is important for the controlled coalescence of droplet compartments and the initiation of chemical and biological reactions, the process itself is challenging due to, for example, the need for both temporal and spatial synchronization of droplets for fusion to occur, the need to overcome the stabilizing effects of surfactants used, and the need for specialized equipment such as electrodes to actively induce droplets to coalesce by subjecting them to an electric field (electrofusion). This makes it difficult or impossible to achieve droplet fusion in a continuous workflow.

[0007] An alternative approach for performing multi-step compartmentalized reactions in vitro is the use of gel beads in microfluidic droplets, which can be particularly useful for biochemical processes such as high-throughput screening, directed evolution, and genotyping.

[0008] For example, PCT Application No. PCT / GB2012 / 051106 describes the following method for screening a plurality of polypeptides for activity in converting a reporter substrate into a detectable reporter: 1) emulsifying an aqueous reporter solution comprising a population of polynucleotides, a reporter substrate, and a gel-forming agent into microdroplets, wherein each polynucleotide encodes a product that converts the reporter substrate into a detectable reporter within the microdroplets; 2) solidifying the gel-forming agent within the microdroplets to produce gel beads containing the polynucleotide and the product-generated detectable reporter; 3) demulsifying the aqueous microdroplets and resuspending the beads in an aqueous detection solution; 4) Detecting, determining, or measuring the detectable reporter within one or more beads in the population.

[0009] One or more beads can be identified and / or isolated from an aqueous solution containing or not containing a detectable reporter. Polynucleotides from one or more identified beads can be identified, amplified, cloned, sequenced, or otherwise examined. The polynucleotides or nucleic acids may be isolated, for example, in a plasmid, virus, or PCR product, or may be contained in a cell or virus particle. The polynucleotides are retained within the gel beads.

[0010] In contrast to simple water-in-oil droplets, water-in-oil-in-water (w / o / w) double emulsions provide an internal aqueous compartment for isolating biological components, along with an aqueous carrier for flow cytometry analysis. Tawfick and Griffiths previously utilized w / o / w double emulsions as miniaturized microreactors for the formation of genotype-phenotype linkages in directed protein evolution to identify potential catalytic "hits" based on the appearance of product fluorescence. Library members exhibiting the desired activity were then selected by flow cytometry. However, although numerous microfluidic systems have been developed for creating and sorting droplets, the required operational skills preclude easy implementation for non-experts. Therefore, we employ a two-step monodisperse double emulsion droplet method, as described by Zinchenko et al., which uses two separate microfluidic devices with different surface properties for single (hydrophobic device surface coating) and double emulsion droplet formation (hydrophilic device surface coating). The resulting double emulsion droplets are suitable for quantitative analysis and sorting by fluorescence activated cell sorting (FACS).

[0011] International Patent Application WO 2000 / 036093 A2 discloses a method for generating cyclic peptides and splicing peptide intermediates within a loop structure. This method utilizes the trans-splicing ability of a split intein to catalyze the cyclization of a peptide from a precursor peptide with the target peptide sandwiched between two portions of the split intein. The interaction of the two portions of the split intein forms a catalytically active intein, which forces the target peptide into a loop configuration that stabilizes ester isomers of amino acids at the junction between one of the intein portions and the target peptide. A heteroatom from the other intein portion then reacts with the ester to form a cyclic ester intermediate. The active intein then catalyzes the formation of an aminosuccinimide, liberating the target peptide in cyclized form, which then undergoes spontaneous rearrangement to form the thermodynamically favorable backbone cyclic peptide product.

[0012] International Patent Application WO 2012 / 156744 A2 discloses the use of gel beads in microfluidic droplets to perform multi-step compartmentalized reactions in vitro. The method can include emulsifying an aqueous reporter solution containing polynucleotides, a reporter substrate, and a gel-forming agent into microdroplets. Each polynucleotide encodes a product that converts the reporter substrate into a detectable reporter within the microdroplet. The gel-forming agent is then solidified within the microdroplet to produce gel beads containing both the polynucleotide and the detectable reporter generated by the product. The aqueous microdroplets are then demulsified and resuspended in an aqueous detection solution, and the reporter is detected within one or more beads in the population.

[0013] Cui, Weitz, Chong, et al. (Scientific Reports 6, Article number: 22575) developed a streamlined mix-and-read drop-IVT2H method for screening random DNA libraries by incorporating an in vitro two-hybrid system (IVT2H) into microfluidic droplets. Drop-IVT2H is based on the correlation between the binding affinity of two interacting protein domains and the transcriptional activation of a fluorescent reporter. A DNA library encoding potential peptide binders was encapsulated by IVT2H so that single DNA molecules were distributed in individual droplets.

[0014] (PNAS 2009, 106:34, pp. 14195-14200) presented a droplet-based microfluidic technology that enables high-throughput screening of single mammalian cells and developed an optically encoded droplet library that allows for identification of droplet composition during assay readout. Using this integrated droplet technology, a drug library was screened for cytotoxic effects on U937 cells.

[0015] Thiele et al. (Lab on a chip, 2014, pages 2651-2652) reported the use of hyaluronic acid hydrogel beads for membrane-free in vitro transcription / translation. However, the insoluble hyaluronic acid hydrogel may reduce protein yield.

[0016] Small linear peptides exhibit a wide range of biological activities and can be easily synthesized in an almost infinite variety of sequences using conventional techniques in solid-phase synthesis and combinatorial chemistry, making them useful for probing a variety of physiological phenomena. These properties also make them particularly useful for identifying and developing new drugs. For example, large libraries of numerous different small linear peptides can be synthetically prepared and screened for specific properties in various biological assays. (See, e.g., Scott, JK and GP Smith, Science 249:386, 1990; Devlin, JJ, et al., Science 24:404, 1990; Furka, A. et al., Int. J. Pept. Protein Res. 37:487, 1991; Lam, KS, et al., Nature 354:82, 1991.) Peptides within the library that exhibit specific properties can then be isolated and serve as candidates for further study. Sequencing can be used to characterize selected peptides, for example, by their associated polynucleotides.

[0017] Despite these advantages, only a few small linear peptides have been developed into widely used pharmaceuticals, in part because they are typically cleared rapidly from the body, thereby limiting their therapeutic value.

[0018] Ring closure, or cyclization, can reduce the rate at which peptides are degraded in vivo, thereby dramatically improving their pharmacokinetic properties. Synthetic methods that generate large numbers of different peptides with infinitely variable amino acid sequences greatly facilitate the identification of specific cyclic peptides as potential new drugs.

[0019] Various methods for producing cyclic peptides have been described. For example, chemical reaction protocols have been devised to produce cyclic peptides, such as those described in U.S. Patent Nos. 4,033,940 and 4,102,877. Other approaches combine biological and chemical methods to produce cyclic peptides. These latter methods involve first expressing a linear precursor of the cyclic peptide and then chemically converting the linear precursor into a cyclic peptide by adding an exogenous substance, such as a protease or a nucleophile. See, for example, Camerero, JA, and Muir, TW, J. Am. Chem. Society. 121:5597 (1999); Wu, H. et al., Proc. Natl. Acad. Sci. USA, 95:9226 (1998).

[0020] Once produced, the cyclic peptides can be screened for pharmacological activity. For example, a library containing a large number of different cyclic peptides can be prepared and screened for a particular property, such as the ability to bind to a particular target ligand. The library is mixed with the target ligand, and library members that bind to the target ligand can be isolated and identified by sequencing the associated polynucleotide. Similarly, libraries of cyclic peptides can be added to assays for specific biological activities. Cyclic peptides that modulate biological activity can then be isolated and identified by sequencing.

[0021] Cyclic peptide libraries are increasingly being used in high-throughput screening to identify inhibitors of various challenging targets. Methods for generating such libraries can be divided into two categories: genetically encoded approaches, such as phage display or SICLOPPS (Split-Intein Circular Ligation of Peptides and Proteins), and chemically synthesized libraries that require tagging with a deconvolutional code, typically using a unique nucleic acid (e.g., RNA or DNA) code to identify each member of the library.

[0022] Recent advances in molecular biology allow some molecules to be co-selected for properties along with the nucleic acids that encode them. The selected nucleic acids can then be cloned for further analysis or use, or subjected to additional rounds of mutation and selection.

[0023] Common to these methods is the establishment of large nucleic acid libraries, and molecules with desired properties (activities) can be isolated through selection regimes that select for the desired activity of the encoded polypeptides, such as a desired biochemical or biological activity, e.g., binding activity.

[0024] Phage display technology has been highly successful in providing a vehicle that allows for the selection of displayed proteins by providing an essential link between the nucleic acid and the activity of the encoded polypeptide (Smith, 1985; Bass et al., 1990; McCafferty et al., 1990; for a review, see Clackson and Wells, 1994). Filamentous phage particles function as genetic display packages, containing proteins on the outside and the polynucleotides that encode them on the inside. This tight link between nucleic acid and the activity of the encoded polypeptide is a result of phage assembly within bacteria. Because individual bacteria are rarely infected with multiple genes, most often all phages produced from an individual bacterium carry the same polynucleotide and display the same protein.

[0025] However, phage display relies on the in vivo generation of nucleic acid libraries in bacteria. Therefore, the practical limit to the library size possible with phage display technology is 10, even when using λ phage vectors with excisable filamentous phage replicons. 7 〜10 11 This approach has been applied primarily to the selection of molecules with binding activity. A few catalytically active proteins have also been isolated using this technique; however, selection was not directly directed at the desired catalytic activity, but rather at binding to transition-state analogs (Widersten and Mannervik, 1995) or reaction with suicide inhibitors (Soumillion et al., 1994; Janda et al., 1997). Recently, there have been several examples of enzymes selected using phage display for product formation (Atwell and Wells, 1999; Demartis et al., 1999; Jestin et al., 1999; Pederson et al., 1998), but in all these cases selection was not directed at multiple turnover events.

[0026] Although mRNA display allows for the generation of large libraries, such libraries are limited to affinity-based screening and suffer from the same drawbacks as phage display.

[0027] mRNA display is a method for linking genotype and phenotype by covalently linking mRNA as the genotype and a peptide molecule as the phenotype using a cell-free translation system (in vitro transcription / translation system). It is applied by linking a synthesized peptide molecule to the mRNA encoding it via puromycin, an analogue of the 3' end of tyrosyl-tRNA.

[0028] In mRNA display (see Szostak, JW and Roberts, RW, U.S. Pat. No. 6,258,558, the entire disclosure of which is incorporated herein by reference), each mRNA molecule in a library is modified at its 3' end by the covalent attachment of a puromycin-like moiety. The puromycin-like moiety is an aminoacyl-tRNA acceptor stem analog that functions as a peptidyl acceptor and can be added to a growing polypeptide chain by the peptidyl transferase activity of the ribosome translating the mRNA. During in vitro translation, the mRNA and the encoded polypeptide are covalently linked via the puromycin-like moiety to form an RNA-polypeptide fusion. After selection of fusion molecules by binding of their polypeptide component to a target (i.e., by screening), the RNA component of the selected fusion molecule can be amplified by PCR and subsequently characterized.Several other methods have been developed to create a physical linkage between a polypeptide and its encoding nucleic acid to facilitate selection and amplification (Yanagawa, H., Nemoto, N., Miyamoto, E., and Husimi, Y., U.S. Pat. No. 6,361,943; Nemoto, H., Miyamoto-Sato, E., Husimi, H., and Yanagawa, H. (1997). FEBS Lett. 414:405-408; Gold, L., Tuerk, C., Pribnow, D., and Smith, JD, U.S. Pat. Nos. 5,843,701 and 6,194,550; Williams, RB, U.S. Pat. No. 6,962,781; Baskerville, S. and Bartel, DP (2002). Proc. Natl. Acad. Sci. 2002, 1:101-104, each of which is incorporated herein by reference in its entirety). Sci. USA 99:9154-9159; Baskerville, DS and Bartel, DP, US Pat. No. 6,716,973; Sergeeva, A. et al. (2006). Adv. Drug Deliv. Rev. 58:1622-1654).

[0029] mRNA display represents a particularly useful method for generating large peptide libraries.

[0030] In mRNA display, mRNA containing puromycin already attached to its 3' end via an appropriate linker is introduced into a cell-free translation system to synthesize peptides from the mRNA. Puromycin is fused to the C-terminus of the growing peptide chain as a substrate for peptidyl transfer reactions on the ribosome. The translated peptide molecule is fused to the mRNA via the puromycin moiety. Unlike the 3' end of aminoacyl-tRNA, puromycin forms an amide bond with the nascent peptide rather than an ester bond. Therefore, the puromycin-peptide conjugate fused to each other on the ribosome is resistant to hydrolysis and stable.

[0031] In mRNA display, puromycin must be attached to the 3' end of the mRNA. This attachment can be achieved by first preparing a puromycin-binding linker with a linear polymer spacer and then fusing the linker to the 3' end of the mRNA. Attachment can also be achieved by first attaching a spacer to the 3' end of the mRNA and then fusing puromycin to the conjugate. In either method, the linear polymer spacer generally contains a phosphate group or a nucleotide at its end, and the bond between the 3' end of the mRNA and the 5' end of the linker is a covalent bond via the m phosphate group. This covalent bond can be formed using RNA ligase or DNA ligase or by standard organic chemistry reactions.

[0032] EP2492344A1 discloses a modified mRNA display method known as "RAPID." In the RAPID display method of the present invention, a linker is attached to the 3'-terminus of the mRNA at one end and to the C-terminus of the nascent peptide at the other end, thereby connecting the mRNA and the peptide translated therefrom, as in known mRNA display methods.

[0033] However, the linker used in the RAPID display method has a different structure at both ends from that used in the known mRNA display method. In this specification, the linker used in the RAPID display method may be referred to as a "RAPID linker."

[0034] The region at one end of the linker that binds to the C-terminus of the peptide may be referred to herein as a "peptidyl acceptor" or simply as an "acceptor." Accordingly, the term "peptidyl acceptor" refers to a molecule or moiety having a structure that can bind to a growing peptide (peptidyl-tRNA) by peptidyl transfer reaction on a ribosome. The peptidyl acceptor may refer to a region located at the end of the linker or to the entire structure including the linker. For example, the peptidyl acceptor in known mRNA display methods is puromycin located at one end of the linker, or a puromycin-binding linker as the entire structure including the linker.

[0035] RAPID linkers are characterized by the structure and preparation process of the peptidyl acceptor.

[0036] In the RAPID display method, a linker having a sequence consisting of the four-residue ribonucleotide ACCA is synthesized at the 3' end, and then an amino acid is attached to the adenosine at the 3' end, thereby conferring a peptidyl acceptor structure to the linker. During the peptide elongation reaction in the ribosome, the amino acid attached to the end of the linker accepts the C-terminus of the peptide from the peptidyl-tRNA and binds to the peptide. The structure in which the amino acid is attached to the RNA sequence ACCA via an ester bond is referred to herein as the "peptidyl acceptor region."

[0037] The peptidyl acceptor in known mRNA display methods is puromycin, which has an aminonucleoside structure in which the ribose in the adenosine-like moiety is linked to an amino acid via an amide bond. However, in RAPID display, the amino acid is linked to the 3'-O of ribose via an ester bond. In other words, the peptidyl acceptor in RAPID display has a nucleoside structure similar to that of natural aminoacyl-tRNA. By utilizing a structure similar to that of the natural acceptor, the peptidyl acceptor exhibits incorporation efficiency equal to or greater than that of puromycin.

[0038] The bond formation between the peptidyl acceptor and the C-terminus of the peptide is thought to occur when the amino group of the peptidyl acceptor incorporated into the A-site comes into close proximity with an ester bond at the C-terminus of the attached peptide of the peptidyl-tRNA in the P-site, similar to the normal peptidyl transfer reaction in the ribosome. Therefore, the covalent bond formed by the C-terminus of the peptide chain is generally an amide bond, as in mRNA display. It should be noted that linkers containing non-natural (non-standard / unnatural / non-proteinogenic) amino acids such as D-amino acids or β (beta)-amino acids can also be used in the RAPID display of the present invention by using an artificial RNA catalyst (flexizyme) for linker synthesis.

[0039] In the RAPID display method, the 5' end of the linker and the 3' end of the mRNA molecule form a complex through hybridization based on base pairing. This region of the RAPID linker is referred to herein as the "single-stranded structural region." Specific examples of single-stranded structures with nucleic acid bases in their side chains include single-stranded DNA, single-stranded RNA, and single-stranded PNA (peptide nucleic acid). The resulting complex must remain stable during peptide selection from an mRNA library. Increasing the complementarity between the nucleotide sequence of the single-stranded structural region of the linker and the sequence at the 3' end of the mRNA molecule increases the efficiency of duplex formation and also increases stability. Stability also depends on the GC content, the salt concentration of the reaction solution, and the reaction temperature. In particular, this region desirably has a high GC content, specifically a GC content of 80% or more, preferably 85% or more.

[0040] The remainder of the linker, excluding both ends, is designed to have a simple, flexible, hydrophilic, linear structure with few side chains overall, similar to the structure of linkers used in known mRNA display methods. Therefore, for example, linear polymers including oligonucleotides such as single-stranded or double-stranded DNA or RNA, polyalkylenes such as polyethylene, polyalkylene glycols such as polyethylene glycol, polystyrene, polysaccharides, or combinations thereof can be appropriately selected and used. The linker preferably has a length of 100 angstroms or more, more preferably about 100 to 1,000 angstroms.

[0041] An advantage of the RAPID system is that the association of the peptidyl acceptor region with the library mRNA can occur in situ by the IVTT system, i.e., no preliminary ligation step is required, for example, between puromycin and the linker or between the linker and the mRNA.

[0042] Furthermore, unnatural (non-standard) amino acids can be incorporated into the same translation system. For example, the acylation reaction to charge tRNA with non-proteinogenic amino acids or hydroxy acids, the building blocks of unusual peptides, can be mediated by artificial RNA catalysts (ribozymes), such as flexizymes.

[0043] Specific peptide ligands have been selected for receptor binding by affinity selection using large peptide libraries linked to the C-terminus of the lac repressor, LacI (Cull et al., 1992). When expressed in E. coli, the repressor protein physically links the ligand to the encoding plasmid by binding to the lac operator sequence on the plasmid.

[0044] A completely in vitro polysome display system has also been reported in which nascent peptides are physically attached to the RNAs that encode them via ribosomes (Mattheakis et al., 1994; Hanes and Pluckthun, 1997). Alternative completely in vitro systems that link genotype to phenotype by creating RNA-peptide fusions have also been described (Roberts and Szostak, 1997; Nemoto et al., 1997).

[0045] However, within the scope of the above systems, activities other than binding, such as catalytic or regulatory activity, cannot be directly selected. Most of these approaches are only compatible with affinity-based assays. However, most pharmaceutical assays are functional assays, and in any case, functional assays are superior; simply because a compound binds to a target protein does not mean that it has a biological function.

[0046] One exception is the SICLOPPS method, which is coupled to both affinity-based and functional assays, however, to date SICLOPPS has only been demonstrated in cell-based in vivo systems.

[0047] In the SICLOPPS system, a split (or trans) intein (C-intein or I-intein) is inserted at one end of a nucleotide sequence encoding a peptide to be cyclized. C ) at the other end, a split intein (N-intein or I N The nucleic acid molecule can be constructed such that the two split intein components (i.e., I) are flanked by nucleotide sequences encoding the amino-terminal portions of the two split intein components (i.e., I). Expression of the construct results in the production of a fusion protein. The two split intein components (i.e., I) of the fusion protein are then fused together. C and I N) associate to form an active enzyme that splices together the amino and carboxy termini of the intervening sequence to generate the backbone cyclic peptide. The chemical reaction is shown in Figure 14. This method can be adapted to facilitate the selection or screening of cyclic peptides with desired properties.

[0048] Thus, the invention can feature a non-naturally occurring nucleic acid molecule that encodes a polypeptide having a first portion of a split intein, a second portion of the split intein, and a target peptide sandwiched between the first portion of the split intein and the second portion of the split intein. Expression of the nucleic acid molecule produces a polypeptide that spontaneously splices to produce a cyclized form of the target peptide or a splicing intermediate of the cyclized form of the target peptide, such as an active intein intermediate, a thioester intermediate, or a lariat intermediate.

[0049] Both the first portion of the split intein and the second portion of the split intein are derived from a naturally occurring split intein, such as Ssp DnaE. In other variations, one or both of the split intein portions may be derived from a non-naturally occurring split intein, such as those derived from RecA, DnaB, Psp Pol-1, and Pfu inteins.

[0050] Tawfik and Griffiths (1998) and International Patent Application PCT / GB98 / 01889 describe a system for in vitro evolution that overcomes many of the above limitations by using microcapsule compartmentalization to link genotype and phenotype at the molecular level.

[0051] In Tawfik and Griffiths (1998) and in some embodiments of International Patent Application PCT / GB98 / 01889, the desired activity of a polypeptide results in the modification of the polynucleotide that encodes it (and is present in the same microcapsule). The modified polynucleotide can be selected in a subsequent step. However, these approaches do not consider cyclic polypeptides.

[0052] Thus, there is a need for methods of generating and tagging cyclic peptide libraries that are compatible with any pharmaceutical assay. The present invention addresses this need by providing cyclic polypeptides co-compartmentalized with their encoding polynucleotides in a format that is compatible with pharmaceutical assays and peptide library generation. [Prior art documents] [Patent documents]

[0053] [Patent Document 1] WO2012 / 156744A2 [Patent Document 2] WO2000 / 036093A2 [Non-patent literature]

[0054] [Non-Patent Document 1] Townend and Tavassoli, 2016, ACS Chemical Biology, 1624-1630 [Non-patent document 2] Tavassoli and Benkovic, 2007, Nature Protocols, 1126-1133 Summary of the Invention

[0055] 3. Summary of the invention

[0056] The present invention has the advantage of providing cyclic polypeptides in a format that can be coupled to any pharmaceutical assay, particularly a functional assay, allowing for easy and unique identification of each polypeptide.

[0057] In a first aspect, there is provided a method for producing a cyclic polypeptide and a polynucleotide encoding the cyclic polypeptide in a co-compartmentalized manner, the method comprising: a) forming a compartment comprising a polynucleotide encoding a cyclic polypeptide; b) expressing a polypeptide from the polynucleotide; c) cyclizing the polypeptide.

[0058] In embodiments, cyclization of a polypeptide is a passive process. In other words, the polypeptide may self-cyclize or autocyclize. A passive process may be an "autocatalytic" process. For example, the polypeptide may be a SICLOPPS polypeptide. In such cases, the step "(c) cyclizing the polypeptide" comprises simply cyclizing the polypeptide, e.g., by allowing the reaction to proceed for a suitable length of time. In embodiments, a suitable length of time for cyclization may be several seconds. In other embodiments, a suitable length of time for cyclization may be several minutes, several hours, or several days. The time required for cyclization may vary greatly for any given polypeptide. However, determining a suitable length of time for cyclization to the desired degree is within the skill of one of ordinary skill in the art.

[0059] Alternatively, polypeptides may be cyclized by other methods known in the art, such as chemical or enzymatic methods. In these "non-passive" cyclization methods, the components necessary for cyclization are contacted with the polypeptide and allowed to react under the necessary conditions for a suitable period of time to effect the desired degree of cyclization. In embodiments, the length of time suitable for cyclization may be several seconds. In other embodiments, the length of time suitable for cyclization may be several minutes, hours, or days. The time required for cyclization may vary greatly for any given polypeptide. However, determining the length of time suitable for cyclization to the desired degree is within the skill of one of ordinary skill in the art.

[0060] Since expression of the cyclic polypeptide occurs within the compartment, the polynucleotide encoding the cyclic peptide is contained within the same compartment. Thus, the cyclic polypeptide can be uniquely identified by isolating and sequencing the co-compartmentalized polynucleotide. Furthermore, the compartments form a microreactor that contains all components of the expression system and / or other reaction components.

[0061] A polynucleotide may be non-covalently associated with its encoded product, e.g., the polynucleotide and encoded product may be contained within the same bead, compartment, cell, or viral particle. The polynucleotide and its encoded product may be linked directly or indirectly via a non-covalent bond.

[0062] Alternatively or additionally, a polynucleotide may be covalently associated with its encoded product, for example, via a puromycin moiety in an mRNA display system.

[0063] In a second aspect, there is provided a method for sorting cyclic polypeptides, comprising: a) forming a compartment comprising a polynucleotide encoding a cyclic polypeptide; b) expressing a polypeptide from the polynucleotide; c) cyclizing the polypeptide; d) screening the cyclic polypeptides for activity; e) selecting a cyclic polypeptide that exhibits a desired activity.

[0064] For example, cyclic polypeptides can be used to induce or inhibit fluorescence. Beads and / or compartments that exhibit or do not exhibit fluorescence can then be selected and / or sorted accordingly, such as by Fluorescence Activated Droplet Sorting (FADS). Cyclic peptides that exhibit desired properties can then be identified by sequencing the associated polynucleotides.

[0065] In embodiments, the method further comprises identifying the selected cyclic polypeptide, for example, by sequencing a polynucleotide associated with the polypeptide according to the invention.

[0066] In some embodiments, the compartments may be water-in-oil (w / o) or water-in-oil-in-water (w / o / w) droplets obtained by microfluidic manipulation of a solution containing the polynucleotide.

[0067] In other embodiments, the compartments may be microcapsules obtained by bioelectrospraying or jetting of a suitable solution of polynucleotides in a polyelectrolyte, such as alginate compartments.

[0068] In yet other embodiments, the compartments may be vesicles, such as lipid vesicles.

[0069] The method may further comprise amplifying the polynucleotide. Increased copy number of the polynucleotide may ultimately result in higher expression of the cyclic polypeptide, which may facilitate detection of the cyclic polypeptide in any subsequent assays, thereby increasing sensitivity.

[0070] The compartments may contain a gel-forming agent that solidifies (or forms) into gel beads after polynucleotide amplification. The amplified polynucleotides are captured (or immobilized) in the gel matrix, thereby retaining identical polynucleotide copies within a single bead. This prevents the amplified copies of the polynucleotide from being lost from the beads if the compartments are later destroyed, thereby maintaining the ability to uniquely identify the beads. The gel-forming agent can be made to form a gel, for example, by cooling to a temperature at which a gel forms.

[0071] An external heat source may be utilized during the amplification reaction. This may be achieved by applying heat to the reaction vessel or reaction container, e.g., a glass syringe, in which the amplification is occurring. This may stimulate the amplification reaction. Alternatively or additionally, in embodiments utilizing a gel-forming agent, the application of heat maintains the gel in a liquid phase.

[0072] Preferably, heat is applied to the container uniformly and continuously at a constant temperature. By "uniformly," we mean that substantially all surfaces of the container receive the same degree of heating, i.e., any point on the surface of the container receives the same amount of thermal energy as any other point on the surface of the container. In embodiments, "substantially all" of the container includes only one surface of the container that laterally surrounds or encircles the container, or multiple surfaces that together laterally surround or encircle the container. For example, the container may be a syringe with a circular cross-section, thus having an overall cylindrical shape capped with two circular surfaces (top and bottom) on a single outer curved surface (side) that surrounds the syringe. In this example, heat is applied uniformly to the surface of the container if any point on the curved surface (side) receives the same amount of thermal energy as any other point on the curved surface (side), regardless of the thermal energy received by either of the cap surfaces (top and bottom).

[0073] The external heat source may comprise a flexible heating element or filament. The heat source may be electronically powered. In embodiments, the heat source is a commercially available electronic heating pad. In other embodiments, the heat source may be a fluid-filled jacket into which heated fluid, e.g., water, is continuously supplied from a heated fluid source. In embodiments, the external heat source may be integrally formed with the vessel.

[0074] After the formation of the gel beads, the compartments can be disrupted, allowing the conditions under which the beads are placed to be changed. For example, the buffer can be changed by a buffer exchange procedure. The new buffer can penetrate the gel beads and interact with the polynucleotide copies or other components held within the beads. This may lead to the activation of new processes, such as gene expression.

[0075] Each gel bead represents a mechanically stable unit and can be considered a reservoir of polynucleotides with identical sequences. Each bead can be individually manipulated. For example, the beads can be fed into a microfluidic device for emulsification.

[0076] In some embodiments, the gel beads can be subjected to conditions for expressing the cyclic polypeptide, for example, the gel beads can be contacted with an in vitro transcription and translation (IVTT) system.

[0077] After and / or during exposure to conditions for expressing the cyclic polypeptide, compartments may form around the gel beads. In other words, the beads are re-compartmentalized, meaning that new compartments are formed around the non-compartmentalized beads. This second compartmentalization step may follow the same procedure as the first compartmentalization described above. Alternatively, the second compartmentalization step may follow a different procedure. For example, the first compartmentalization may be via microfluidics to form droplets of an emulsion containing the gel beads, and the second compartmentalization may be via a polyelectrolyte jetting procedure to form compartments containing the gel beads. In another example, both the first and second compartmentalization procedures may be via microfluidics to form droplets of an emulsion containing the beads.

[0078] After translation, the beads can be assayed. For example, cyclic peptides can be evaluated for their potential to inhibit a target enzyme by optical assays, such as colorimetric or fluorometric assays. If the enzyme catalyzes a reaction that produces a colored or fluorescent product, beads containing inhibitory polypeptides will exhibit reduced color or fluorescence intensity compared to other beads containing non-inhibitors. The beads can be sorted in a high-throughput manner, for example, by fluorescence-activated cell sorting (FACS) or fluorescence-activated droplet sorting (FADS). In embodiments, the beads are within droplets of an emulsion. To be compatible with FACS / FADS, the continuous phase of the emulsion should be aqueous.

[0079] The polynucleotide can contain a sequence encoding an N-terminal intein fragment, followed by a sequence encoding a cyclic polypeptide, followed by a sequence encoding a C-terminal intein fragment. When expressed as a polypeptide, the N-terminal intein fragment binds to the C-terminal intein fragment. This results in an intervening polypeptide containing the desired cyclic polypeptide forming a polypeptide loop. The intein loop structure then undergoes a spontaneous splicing reaction, thereby generating a free intein and the desired cyclic polypeptide. The desired polynucleotide can be obtained by conventional molecular cloning techniques.

[0080] In any aspect or embodiment, the compartments may be water-in-oil-in-water (w / o / w) emulsions, vesicles, or compartments.

[0081] In a third aspect, a) a polynucleotide encoding an N-terminal intein fragment, followed by a sequence encoding a cyclic peptide, followed by a sequence encoding a C-terminal intein fragment; b) a cyclic polypeptide.

[0082] Expression of the polynucleotide produces a linear polypeptide containing an N-terminal intein fragment, an intervening cyclic polypeptide sequence, and a C-terminal intein fragment. In the linear peptide, spontaneous splicing reactions can occur at the junction between the intein fragment and the cyclic polypeptide sequence, thereby producing a cyclic polypeptide and a free intein portion. This is one way in which compartments can be created that contain both the cyclic polypeptide and the polynucleotide encoding it. Such compartmentalized cyclic polypeptides can be assayed. In particular, compartmentalized cyclic polypeptides are compatible with pharmaceutical assays, including functional assays. This allows for high-throughput screening and selection of promising cyclic polypeptide candidate compounds, for example, in drug discovery. This is particularly true when using libraries according to the present invention.

[0083] Thus, the invention can feature a non-naturally occurring nucleic acid molecule that encodes a polypeptide having a first portion of a split intein, a second portion of a split intein, and a target peptide sandwiched between the first portion of the split intein and the second portion of the split intein. Expression of the nucleic acid molecule produces a polypeptide that spontaneously splices to produce a cyclized form of the target peptide or a splicing intermediate of the cyclized form of the target peptide, such as an active intein intermediate, a thioester intermediate, or a lariat intermediate.

[0084] The present invention may be used to encode cyclic polypeptides containing one or more unnatural amino acids by using reassigned codon sets in combination with a custom IVTT mixture containing specific tRNAs loaded with unnatural amino acids, and tRNAs charged with unnatural amino acids can be easily generated using previously reported methods. For example, the acylation reaction to charge tRNAs with nonproteinogenic amino acids or hydroxy acids, which are building blocks of unusual peptides, can be mediated by artificial RNA catalysts (ribozymes), such as flexizymes.

[0085] The first portion of the split intein and the second portion of the split intein are both derived from a naturally occurring split intein, such as Npu or Ssp DnaE. In other variations, one or both of the split intein portions can be derived from a naturally occurring, artificial, or non-naturally occurring split intein, such as those derived from RecA, DnaB, Psp Pol-1, and Pfu inteins.

[0086] In a fourth aspect, a library is provided in which a cyclic polypeptide is co-compartmentalized with a polynucleotide encoding the same. The library contains a plurality of such compartments, and the polynucleotide in at least one compartment contains a sequence that is different (i.e., not identical) to the sequence of the polynucleotide in at least one other compartment (i.e., a different compartment). The library can be screened, for example, by performing a fluorescent or colorimetric assay to determine whether the compartments exhibit a desired signal, as selected, for example, by FADS.

[0087] In another aspect, a) a microfluidic device; b) a polynucleotide encoding an N-terminal intein fragment; and and c) a polynucleotide encoding a C-terminal intein fragment.

[0088] In embodiments, the kit comprises an encapsulating material, which may be an oil, a lipid, or a polyelectrolyte.

[0089] In some embodiments, the kit further comprises a gel-forming agent.

[0090] Polynucleotides encoding intein fragments can be used to generate polynucleotides encoding cyclic polypeptide sequences flanked by an N-terminal intein fragment at one end and a C-terminal intein fragment at the other end, which can be accomplished by conventional molecular cloning techniques, e.g., ligating an intein-encoding sequence to a cyclic polypeptide-encoding nucleotide sequence.

[0091] 4. Brief explanation of the figure [Brief explanation of the drawings]

[0092] [Figure 1] Microfluidic device used for the generation and manipulation of femtoliter-sized droplets. (Left) A PDMS device bonded to a glass slide for the controlled generation of femtodroplets. (Right) Photograph of device operation - droplet formation occurs at the nozzle (10 μm wide x 5 μm deep) of the microfluidic chip, and the newly generated droplets travel through the channel (100 μm wide x 25 μm deep). Four inlets are provided for the injection and introduction of fluids into the device; two outer ports are used for oil and the two central ports for aqueous solutions. Scale bar = 50 μm. [Figure 2] Analysis of the size of single water-in-oil emulsion droplets. The ability to finely control the diameter of the generated droplets was determined by maintaining a constant aqueous solution flow rate of 10 μL / h while increasing the oil / surfactant flow rate from 10 to 60 μL / h (A to F, respectively). The bright-field photographs above each diameter distribution histogram represent the resulting droplet generation stream during collection under each condition. For each experiment, three independent images were analyzed, and the resulting droplet diameters were plotted to obtain the diameter histograms. In each case, the y-axis (number of droplets) was normalized, and the results are presented as decimals. A Gaussian normal distribution was fitted in each case with a constraint amplitude of 1. The coefficient of variation (CV) for each photograph is presented next to each set. Bin size = 0.5 μm. [Figure 3] Analysis of single emulsion droplet diameter, volume, and production rate. (Left) Droplet diameter and oil flow rate for single emulsion droplet production. (Right) Droplet production rate in kHz and average volume plotted against oil flow rates from 10 to 60 μl / h. An increase in droplet volume corresponds to a consequent decrease in production rate. Values ​​plotted as the mean (standard deviation) of three independent photographs of three separate samples analyzed. [Figure 4]A two-chip microfluidic setup for generating double emulsion droplets. Single emulsion droplets for re-emulsification are contained within a glass syringe in an upright position. After settling, the emulsion is driven through a second microfluidic device (enlarged crop) with larger channel dimensions, allowing for double emulsion droplet formation. [Figure 5] Diagram showing double emulsion generation at the junction of a microfluidic device. Droplets move from left to right. (A) Diagram showing the formation of single and double emulsions at the chip. (B) Bright-field image of double emulsion droplet formation at the orifice of the device. [Figure 6] Controllable and flexible generation of highly monodisperse double emulsion (w / o / w) droplets. Water-in-oil-in-water (w / o / w) double emulsion droplets containing a 100 μM fluorescein inner core and a 1% Tris / Tween 80 outer phase within an FC-40 oil shell. Droplets were generated using the following flow rates: emulsion: 4 μl / h, FC-40 spacer oil: 15 μl / h, and 1% Tris / Tween: 60 μl / h. The pronounced spherical structures in the 10x magnification photograph represent mineral oil droplets. Arrows highlight doubly filled double emulsion droplets. (Top left) Inset shows a magnified image of an individual double emulsion droplet, indicating the diameter of the inner aqueous core and the overall double emulsion diameter (ID and OD). [Figure 7]Flow cytometry analysis of FITC-containing double emulsion droplets. (A) Log plot of side scatter vs. forward scatter for a double emulsion sample. (B) Log plot of green fluorescence intensity vs. forward scatter. Histograms of forward scatter (top) and green fluorescence intensity (right) are also shown. (C) A single oil-in-water emulsion sample prepared in the absence of primary emulsion, using identical flow rates and conditions as the previous sample, to determine the level of background fluorescence. (D) Fluorescence vs. forward scatter for single oil-in-water emulsion droplets. Single emulsions were generated using a JUS device, while double emulsions were formed using a 15 × 16 μm (h × w) device. The flow rates and device design used for oil-in-water generation were identical to those for FITC-based w / o / w emulsions to allow for comparison. Insets in (A) and (C) show enlarged gated regions for the desired droplet populations. [Figure 8] Diagram showing the formation of triple emulsion IVTT-containing droplets using DNA-loaded agarose beads and a three-chip microfluidic system. (A) Fluorinated oil, agarose / DNA suspension, and Phi29 DNA polymerase are injected into a hydrophobic microfluidic flow-focusing device and collected upon formation of a steady stream of monodisperse agarose droplets from the flow-focusing junction. (B) The solidified agarose beads are re-injected into a second hydrophobic microfluidic device along with the IVTT mixture. (C) The IVTT / DNA-containing droplets are re-emulsified a third and final time, resulting in an aqueous external phase that allows for flow cytometry analysis. [Figure 9] Microfluidic setup for generating monodisperse agarose beads. The microfluidic device was mounted on an inverted optical microscope. A syringe pump was elevated to the level of the device on top of a laser-cut PMMA stand. Tubing was used to connect the glass syringe to the microfabricated microfluidic system. A microwave-heated commercial heat pad was placed over the agarose syringe to prevent the agarose from solidifying during device operation. A constant temperature of approximately 40 °C was then maintained using a USB-connected electric heating pad. [Figure 10]Analysis of agarose beads containing Phi29 DNA amplified plasmid DNA. (A) The process of agarose droplet formation on the chip; the inset shows a diagram of solidified agarose beads containing amplified DNA. (B) Diameter size distribution of four individual agarose beads after separation from the emulsion. (C) Bright-field image of unwashed agarose beads in 1x TAE buffer, and fluorescent image of agarose beads incubated with a fluorescent double-stranded DNA-binding dye. [Figure 11] Flow cytometry analysis of monodisperse 1% agarose beads containing Phi29 DNA polymerase-amplified plasmid DNA. After incubation at 30°C, isolated agarose beads were stained for fluorescent ds-DNA binding to enable analysis and visualization. For each condition, the starting DNA solution was statistically diluted to ensure an average of (A) 100, (B) 10, (C) 1, or (D) 0.1 DNA copies per droplet (λ value) using Poisson statistics. [Figure 12] In vitro protein expression from agarose beads containing Phi29 pre-amplified SICLOPPS library plasmid DNA in a polydisperse bulk emulsion. After DNA amplification, the agarose beads were isolated and washed by centrifugation at 6,500 rpm three times. After washing, the beads were resuspended in a solution containing PURExpress IVTT components and QX200 oil / surfactant, and finally vortexed for 10-15 seconds to generate a polydisperse bulk emulsion. In contrast to monodisperse droplets generated by microfluidic devices, the generation of polydisperse droplets allows for the rapid determination of conditions suitable for the desired biochemical reaction. Thus, in the presence of pre-amplification, GFP-mediated fluorescence from the SICLOPPS plasmid (encoding the SICLOPPS intein and GFP) is clearly visible. In contrast, no GFP fluorescence is observed in the absence of DNA amplification, highlighting the importance of amplification from a single copy of the IVTT in the droplets. Emulsion samples were incubated at 37°C for 2 hours before imaging. [Figure 13]IVTT of GFP from Phi29 DNA polymerase-preamplified SICLOPPS library plasmid DNA within agarose beads. (A) Microfluidically generated agarose beads containing SICLOPPS plasmid DNA with an average of 100 starting DNA copies and Phi29 DNA amplification components were incubated at 30°C for 16 hours and analyzed using flow cytometry. When stained with a DNA-intercalating dye, a clear increase in green fluorescent intensity signal was observed compared to the corresponding negative control, indicating successful amplification. (B) Prior to IVTT, the solidified DNA-encapsulated agarose beads were washed by centrifugation at 6,500 rpm to remove the amplification buffer and then injected into a 15 x 16 μm hydrophobic device with PURExpress IVTT components. After 2 hours of incubation at 37°C, the IVTT droplets were re-emulsified into a triple emulsion format for flow cytometry analysis in a hydrophilic device. In the forward vs. side scatter plot, two distinct droplet populations are observed, the higher one corresponding to the agarose-in-IVTT-in-oil-in-aqueous droplets. Compared to the control in the absence of amplified DNA and IVTT, a distinct GFP-mediated green fluorescence is observed, approximately 10 times greater than the background fluorescence of the IVTT mixture. [Figure 14]Spontaneous polypeptide cyclization catalyzed by split inteins of SICLOPPS polypeptides. (A) Formation of an active intein from amino- and carboxy-terminal intein fragments (B) stabilizes an ester isomer of an amino acid at the junction between the N-intein and the peptide to be cyclized. (C) A heteroatom from the C-intein is ready to attack the ester, generating a cyclic ester intermediate. (D) Intein-catalyzed aminosuccinimide formation liberates a cyclic peptide (lactone form, not shown) that spontaneously rearranges to form a thermodynamically favorable backbone (lactam form) cyclic peptide product. (E) One or more of the cysteines shown may be substituted with threonine, i.e., the thiol group may be replaced with an alcohol (OH) group. Other heteroatoms may be substituted for the sulfur atoms of the cysteine ​​residues shown, as long as they do not interfere with the formation of the ester intermediate, nucleophilic attack, or backbone rearrangement. [Figure 15] Diagram showing construction of a pETDuet-1-based SICLOPPS vector for generating a cyclic peptide library. Diagram of the pDuetNpu-GFP expression vector, where MCS1 encodes SICLOPPS with an Npu intein and MCS2 encodes GFP. "Library" indicates the location of the variable polynucleotide sequence of the library. This sequence encodes a cyclic polypeptide of the invention that contains a nucleophilic amino acid (e.g., cysteine, serine, or threonine) at the first position. Because it is located between two intein fragments, it is sometimes referred to as the "extein" sequence, as it encodes an extein that is spliced ​​out of the SICLOPPS polypeptide after translation. [Figure 16]pDuetNpu encodes the in vitro cyclized polypeptide CLLFVY. Mass spectrometry analysis of the vector CSpDuetNpuHisCLLFVY after IVTT in bulk format and after FACS sorting. (Left) Standard PURExpress IVTT reaction assembled in the presence of vector CSpDuetNpuHisCLLFVY. (Right) Preamplification of vector CSpDuetNpuHisCLLFVY in agarose beads. Washed agarose beads containing monoclonal amplified DNA were reencapsulated with IVTT to express CLLFVY before the third emulsification step, generating double-emulsion FACS-compatible droplets. FACS-sorted samples were separated from the emulsion and subjected to mass spectrometry analysis. In both cases, a peak representing cycloCLLFVY is evident. [Figure 17] In vitro expression of AB42-GFP fusion with cycloTAFDR (Matis et al.) in microfluidic droplets. TAFDR was cloned into MCS1 using the vector CSpDuetNpuHisAB42-GFP. Plasmid DNA was then encapsulated in agarose droplets along with isothermal DNA amplification reaction components and incubated overnight. Agarose beads were prepared as previously described and reencapsulated with IVTT. A final emulsification step was performed to generate a FACS-compatible double emulsion. A negative control vector in the presence of CA5 was similarly constructed to verify inhibition of aggregation of cycloTAFDR AB42-GFP. A positive shift in green fluorescence was observed upon cycloTAFDR expression compared to buffer alone, IVTT alone, and cycloCA5 data. [Figure 18]In vitro compartmentalization and FACS screening of the TX4 SICLOPPS library in double emulsion droplets. Plasmid construction of the pETDuet-1 vector containing NpuHis TX4 in MCS1 and AB42-GFP fusion in MCS2. To enable monoclonal DNA amplification prior to IVTT, plasmid DNA was encapsulated in agarose femtodroplets with TempliPhi isothermal DNA amplification components. The beads were separated from the emulsion, and the aqueous phase was extracted and washed to remove the amplification buffer. The beads were then encapsulated into highly monodisperse droplets with the PURExpress IVTT system and incubated at 37°C for 2 hours. After protein expression, the sample was converted to a double emulsion format (agarose beads in oil-in-water IVTT) to enable FACS screening. The double emulsion population was identified and gated to enable library sorting (B). Only fluorescence measurements above the IVTT background were gated from the double emulsion population and sorted by FL1-H (GFP, A). The percentage +ve (potential positive candidate peptides) and -ve gated particles are shown for the TX4 library sample only (orange line). DETAILED DESCRIPTION OF THE INVENTION

[0093] Detailed Description of the Preferred Embodiments

[0094] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art, including those in the fields of peptide chemistry, microfluidics, nucleic acid chemistry, molecular genetics and cloning, and biochemistry. Standard procedures are used for molecular biology, genetics, and biochemistry (see Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd ed., 2001, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY; Ausubel et al., Short Protocols in Molecular Biology (1999) 4th ed., John Wiley & Sons, Inc.), which are incorporated herein by reference.

[0095] A "linear polypeptide" is a peptide that is not cyclic and generally has both a carboxy terminal amino acid with a free carboxy terminus and an amino terminal amino acid with a free amino terminus.

[0096] As used herein, the term "intein" refers to a naturally occurring or artificially engineered polypeptide sequence embedded within a precursor protein that is capable of catalyzing a splicing reaction during post-translational processing of the protein. A partial list of known inteins is available at http: / / www.inteins.com.

[0097] A "split intein" is an intein that has two or more separate components that are not fused to each other.

[0098] As used herein, the terms "interposed" or "intervening" mean positioned between. Thus, in a polypeptide having a first sequence inserted between a second and a third sequence (or a second and a third sequence with an intervening first sequence), the stretch of amino acids making up the first sequence is physically located between the stretch of amino acids making up the second sequence and the stretch of amino acids making up the third sequence.

[0099] On the other hand, a "cyclic polypeptide" is a polypeptide that has been "cyclized." The term "cyclic" means having constituent atoms that form a ring. With respect to a peptide, the term "cyclization" means rendering the peptide in a circular or "cyclized" form. Thus, for example, a linear carboxypeptide is "cyclized" when its free amino terminus is covalently linked to a free carboxy terminus (i.e., in a head-to-tail manner) so that no free carboxy or amino termini remain in the peptide.

[0100] As used herein, the term "spontaneously" means that the action being described occurs without the addition of exogenous materials. For example, a precursor polypeptide will spontaneously splice to produce a cyclic peptide when nothing is added to a host system other than the precursor polypeptide or a nucleic acid molecule encoding the precursor polypeptide. On the other hand, a precursor polypeptide within a host system will not spontaneously splice within the host system when an agent external to the host system is required to produce the cyclic peptide.

[0101] As used herein, the phrase "expression vector" refers to a vehicle that facilitates the transcription and / or translation of a nucleic acid molecule in an appropriate in vitro or in vivo system. An expression vector is "inducible" if the vector can be expressed (e.g., a nucleic acid molecule within the vector is transcribed into mRNA) by the addition of an exogenous material to a host system containing the expression vector.

[0102] As used herein, the phrase "regulatory sequence" refers to a nucleotide sequence that regulates the expression (e.g., transcription) of a nucleic acid molecule. For example, promoters and enhancers are regulatory sequences.

[0103] The present invention provides new features and attendant advantages, which can be further described in detail, related to the generation of identifiably tagged (or labeled) polypeptides, particularly cyclic polypeptides, by co-compartmentalization with the polypeptides they encode. In particular, the present invention provides new means for performing functional assays of gene libraries that are unattainable with known gene libraries. Prior art libraries, such as phage display libraries, are compatible with affinity-based assays but do not effectively interface with functional assays. The majority of pharmaceutical assays are functional assays, in part because functional assays provide a more accurate indication of biological (or physiological) activity than affinity assays. Thus, there is an unmet need for a technology that combines the diversity and high throughput of gene libraries with the flexibility of interface with functional assays.

[0104] 4.1 Plots

[0105] In a first aspect, the present invention relates to a method for expressing a cyclic polypeptide in a compartment, the method comprising forming a compartment comprising a polynucleotide comprising a sequence encoding the cyclic polypeptide.

[0106] A compartment consists of a physical boundary that defines an internal volume separated from the external environment. In embodiments, the compartment is substantially spherical. A compartment may be a particle in a solution, e.g., a particle suspended in a solution such as a colloid, or a droplet in an emulsion.

[0107] The interior volume of the compartment should be sufficient to contain at least one polynucleotide molecule and one circular polypeptide.

[0108] To ensure that polynucleotides and polypeptides cannot diffuse between compartments, the contents of each compartment are preferably isolated from the contents of surrounding compartments, so that there is little or no movement of polynucleotides and polypeptides between compartments over the time scale of the experiment.

[0109] The physical boundaries of a compartment can be formed from any material capable of preventing the exit of the encapsulated polynucleotide and any associated polypeptides from the compartment. A compartment can be semipermeable, forming a barrier to the movement of one component, such as a polynucleotide, while allowing the movement of other components, such as buffer components or nucleotide phosphates (e.g., nucleotide triphosphates). Thus, the interior of a semipermeable compartment is a thermodynamically open system, allowing both matter and energy to cross the boundary, albeit in a selective manner. Alternatively, a compartment can be impermeable, forming a barrier to the movement of all components, species, and moieties, including media such as aqueous solutions, water, oil, etc., while still allowing energy exchange with the external environment.

[0110] Compartmentalization is the process by which compartments are formed. When an entity is described as being compartmentalized, this means that the entity is contained within a compartment.

[0111] As used herein, the term "compartmentalization" is synonymous with "encapsulation." Thus, the term "compartment" is synonymous with "capsule."

[0112] The compartments of the present invention require appropriate physical properties that allow the practice of the invention. The formation and composition of the compartments has the advantage of not eliminating the function of the machinery for expression of polynucleotides or activity of polypeptides. As will be apparent to those skilled in the art, the appropriate system(s) may vary depending on the exact nature of the requirements for each application of the invention.

[0113] Suitable compartments and compartment-forming materials include emulsion droplets, lipid vesicles, and microcapsules.

[0114] 4.1.1 Partition Size

[0115] The preferred compartment size may vary depending on the exact requirements of the particular selection process to be performed by the present invention. In all cases, there is an optimal balance between the size of the gene library, the required enrichment, and the required component concentrations in the individual compartments to achieve efficient expression and reactivity of the polypeptides.

[0116] The process of expression occurs within each individual compartment provided by the present invention. At sub-nanomolar DNA concentrations, the efficiency of both in vitro transcription and coupled transcription-translation decreases. The requirement that only a limited number of DNA molecules should be present in each compartment sets a practical upper limit to the possible compartment size. Preferably, the average volume of a compartment is less than 5.2 x 10 -16 m 3 less than 10 μm in diameter, more preferably less than 6.5 × 10 -17 m 3 (diameter 5 μm), more preferably less than about 4.2 × 10 -18 m 3 (2 μm diameter), ideally about 9 × 10 -18 m 3 (corresponding to a spherical compartment with a diameter of 2.6 μm).

[0117] The available DNA or RNA concentration within a compartment can be artificially increased by a variety of methods well known to those skilled in the art. These include, for example, the addition of volume-excluding chemicals such as polyethylene glycol (PEG) and various gene amplification techniques, including transcription using RNA polymerases from bacteria such as E. coli (Roberts, 1969; Blattner and Dahlberg, 1972; Roberts et al., 1975; Rosenberg et al., 1975), eukaryotes (Weil et al., 1979; Manley et al., 1983), and bacteriophages such as T7, T3, and SP6 (Melton et al., 1984), polymerase chain reaction (PCR) (Saiki et al., 1988), and Qb replicase amplification (Miele et al., 1983; Cahill et al., 1991; Chetverin and Spirin, 1995; Katanaev et al., 2000). These include the ligase chain reaction (LCR) (Landegren et al., 1988; Barany, 1991), as well as self-sustained sequence replication systems (Fahy et al., 1991) and strand displacement amplification (Walker et al., 1992). If the emulsion and in vitro transcription or coupled transcription-translation system are thermostable (e.g., if a coupled transcription-translation system can be generated from a thermotolerant organism such as Thermus aquaticus), gene amplification techniques that require thermal cycling, such as PCR and LCR, can be used.

[0118] Increasing the effective local nucleic acid concentration allows for the effective use of larger compartments, which brings the preferred practical upper limit to approximately 5.2×10 -16 m 3 (equivalent to a sphere with a diameter of 10 μm)

[0119] The size of the compartment is preferably large enough to accommodate all the components required for the biochemical reactions that need to occur within the compartment, e.g., in vitro, both transcription and coupled transcription-translation reactions require a total nucleoside triphosphate concentration of approximately 2 mM.

[0120] For example, to transcribe a gene into a single short RNA molecule 500 bases in length, a minimum of 500 nucleoside triphosphates (8.33 × 10 -22 To make a 2 mM solution, the number of molecules required is 4.17 x 10 -19 Volume in liters (4.17 x 10 -22 m 3 , and in the case of a sphere, it is contained within a compartment of 93 nm in diameter.

[0121] Furthermore, particularly for reactions involving translation, it should be noted that the ribosomes required for translation to occur are themselves approximately 20 nm in diameter, so a suitable lower limit for the compartment is approximately 0.1 μm (100 nm) in diameter.

[0122] Therefore, the compartment volume is preferably 5.2×10 corresponding to a sphere with a diameter of 0.1 μm to 10 μm. -22 m 3 ~5.2×10 -16 m 3 , more preferably about 5.2×10 -19 m 3 ~6.5×10 -17 m 3 (1 μm and 5 μm). A spherical diameter of about 2.6 μm is most advantageous.

[0123] It is not coincidental that the preferred dimensions of the compartments (droplets with an average diameter of 2.6 μm) are very similar to those of bacteria; for example, Escherichia coli is a rod-shaped cell measuring 1.1–1.5 × 2.0–6.0 μm, and Azotobacter is an oval cell measuring 1.5–2.0 μm in diameter. In its simplest form, Darwinian evolution is based on a "one genotype, one phenotype" mechanism. The concentration of a single compartmentalized gene or genome decreases from 0.4 nM in a 2 μm-diameter compartment to 25 pM in a 5 μm-diameter compartment. Prokaryotic transcription / translation machinery has evolved to operate in compartments approximately 1–2 μm in diameter, where single genes are present at approximately nanomolar concentrations. In a 2.6 μm-diameter compartment, the concentration of a single gene is 0.2 nM. This gene concentration is sufficiently high for efficient translation. Compartmentalization in such volumes ensures that even if only a single molecule of a polypeptide is formed, it is present at approximately 0.2 nM. Thus, the volume of the compartment is selected with the requirements of transcription and translation of the polynucleotide in mind.

[0124] The size of the emulsion compartments can be varied depending on the requirements of the selection system simply by adjusting the emulsion conditions used to form the emulsion. The ultimate limiting factor is the size of the compartments, and therefore the number of compartments possible per unit volume; therefore, the larger the compartment size, the larger the volume required to compartmentalize a given polynucleotide library.

[0125] The size of the compartments is selected taking into account not only the requirements of the transcription / translation system, but also the requirements of the selection system to be utilized for the polynucleotides. That is, components of the selection system, such as chemical modification systems, may require reaction volumes and / or reagent concentrations that are not optimal for transcription / translation. Such requirements can be addressed by a secondary reencapsulation step. Preferably, optimal compartment volumes and reagent concentrations are determined empirically.

[0126] The compartments are preferably obtained on a microfluidic scale. For example, the maximum diameter of the compartment (e.g., the outer diameter (O) in the upper left inset of Figure 6) d)") may be 100 μm or less. In embodiments, the maximum diameter of a compartment may be 0.1 μm to 100 μm in diameter, preferably 0.5 μm to 10 μm, for example 0.5 μm to 5 μm, or 1.0 μm to 4 μm. In embodiments, the maximum diameter of a compartment may be 0.8 μm. When multiple compartments are involved, these measurements apply to the average of the maximum diameters of the compartments. This can be determined by taking a micrograph of a sample containing compartments, measuring the maximum diameter of each compartment in the micrograph multiplied by the appropriate magnification, and determining the average of the resulting measurements.

[0127] A variety of compartmentalization procedures are available (see Benita, 1996) and can be used to create the compartments used in accordance with the present invention. Indeed, over 200 compartmentalization methods have been identified in the literature (Finch, 1993).

[0128] These include membrane-bound aqueous vesicles such as lipid vesicles (liposomes) (New, 1990) and nonionic surfactant vesicles (van Hal et al., 1996). These are closed membranous capsules of single or multiple bilayers of noncovalently associated molecules, each separated from the next by an aqueous compartment. In the case of liposomes, the membrane is composed of lipid molecules, usually phospholipids, although sterols such as cholesterol may also be incorporated into the membrane (New, 1990). Various enzyme-catalyzed biochemical reactions, including RNA and DNA polymerization, can be carried out within liposomes (Chakrabarti et al., 1994; Oberholzer et al., 1995a; Oberholzer et al., 1995b; Walde et al., 1994; Wick & Luisi, 1996).

[0129] In membrane-bound vesicular systems, much of the aqueous phase is outside the vesicles and is therefore not compartmentalized. This continuous aqueous phase must be removed or the internal biological system inhibited or destroyed (e.g., by digestion of nucleic acids by DNase or RNase) so that reactions are confined to the compartment (Luisi et al., 1987).

[0130] Enzyme-catalyzed biochemical reactions have also been demonstrated in compartments generated by a variety of other methods: many enzymes are active in reverse micellar solutions, such as the AOT-isooctane-water system (Menger & Yamada, 1979) (Bru & Walde, 1991; Bru & Walde, 1993; Creagh et al., 1993; Haber et al., 1993; Kumar et al., 1989; Luisi & B., 1987; Mao & Walde, 1991; Mao et al., 1992; Perez et al., 1992; Walde et al., 1994; Walde et al., 1993; Walde et al., 1988).

[0131] Compartments can also be generated by interfacial polymerization and interfacial complex formation (Whateley, 1996). These compartments can have rigid, impermeable or semipermeable membranes. Semipermeable compartments surrounded by nitrocellulose, polyamide, and lipid-polyamide membranes can all support biochemical reactions, including multienzyme systems (Chang, 1987; Chang, 1992; Lim, 1984). Alginate / polylysine compartments (Lim & Sun, 1980) can be formed under very mild conditions and have also proven to be highly biocompatible, providing an effective method for encapsulating, for example, living cells and tissues (Chang, 1992; Sun et al., 1992).

[0132] 4.1.2 Emulsions

[0133] Non-membranous compartmentalized systems based on phase partitioning of aqueous environments in colloidal systems such as emulsions may also be used.

[0134] Preferably, the compartments of the present invention are formed from emulsions, i.e., heterogeneous systems of two immiscible liquid phases, one dispersed in the other, resulting in microscopic or colloidal droplets (Becher, 1957; Sherman, 1968; Lissant, 1974; Lissant, 1984).

[0135] Emulsions can be produced by any suitable combination of immiscible liquids. Preferably, the emulsions of the present invention have water (containing the biochemical components) as a phase (dispersed, internal, or discontinuous phase) present in the form of finely divided droplets, and a hydrophobic immiscible liquid (oil) as a matrix (non-dispersed, continuous, or external phase) in which these droplets are suspended. Such emulsions are called "water-in-oil" (W / O). This has the advantage that the entire aqueous phase containing the biochemical components is compartmentalized into individual droplets (internal phase). The external phase is a hydrophobic oil and is generally biochemical-free and therefore inert.

[0136] Emulsions can be stabilized by the addition of one or more surface-active agents (surfactants). These surfactants, called emulsifiers, act at the water / oil interface, preventing (or at least delaying) phase separation. Many oils and emulsifiers can be used to create water-in-oil emulsions; a recent compilation listed over 16,000 surfactants, many of which are used as emulsifiers (Ash and Ash, 1993). Suitable oils include light white mineral oils and nonionic surfactants (Schick, 1966), such as sorbitan monooleate (Span™ 80, ICI) and polyoxyethylene sorbitan monooleate (Tween™ 80, ICI).

[0137] The use of anionic detergents can also be beneficial. Suitable detergents include sodium cholate and sodium taurocholate. Sodium deoxycholate is particularly preferred, preferably at a concentration of 0.5% w / v or less. Inclusion of such detergents can, in some cases, increase polynucleotide expression and / or polypeptide activity. Addition of some anionic detergents to non-emulsified reaction mixtures completely abolishes translation. However, during emulsification, the detergent is transported from the aqueous phase to the interface, restoring activity. Addition of an anionic detergent to the mixture to be emulsified ensures that the reaction proceeds only after compartmentalization.

[0138] Creating an emulsion generally requires the application of mechanical energy to force the phases together. This can be accomplished in a variety of ways, utilizing a variety of mechanical devices, including agitators (magnetic stir bars, propeller and turbine agitators, paddle devices, whisks, etc.), homogenizers (rotor-stator homogenizers, high-pressure valve homogenizers, jet homogenizers, etc.), colloid mills, ultrasound, and "membrane emulsification" devices (Becher, 1957; Dickinson, 1994). In a preferred embodiment, the emulsion is created by a microfluidic process. Most preferably, the microfluidic emulsion is obtained droplet by droplet, producing one droplet at a time.

[0139] The aqueous compartments formed in water-in-oil emulsions are generally stable, with little, if any, trafficking of polynucleotides or polypeptides between the compartments. Furthermore, biochemical reactions proceed in the emulsion compartments. Furthermore, complex biochemical processes, particularly gene transcription and translation, are also active in the emulsion compartments. This technology exists to create emulsions in quantities up to industrial scales of thousands of liters (Becher, 1957; Sherman, 1968; Lissant, 1974; Lissant, 1984).

[0140] In embodiments, the compartments are droplets of an emulsion, such as droplets of a primary emulsion (e.g., water-in-oil (w / o)). In embodiments involving emulsions, the interface between the outermost layer of the dispersed (or discontinuous) phase and the continuous phase constitutes the boundary surface of the compartment. In embodiments, the droplets are microfluidic droplets. Microfluidic droplets can be obtained using a microfluidic device, such as a microfluidic device on a chip. Microfluidic droplets of water-in-oil emulsions can be obtained in a microfluidic device, preferably a hydrophobic microfluidic device. This can be achieved by using an aqueous phase as the first phase (also called the "dispersed" or "internal" phase) and a non-aqueous phase (e.g., lipophilic) as the second phase (also called the "continuous" or "external" phase).

[0141] "Double" emulsion droplets can also be obtained in microfluidic devices. For example, water-in-oil-in-water (w / o / w) emulsion droplets can be obtained in microfluidic devices, preferably hydrophilic microfluidic devices. This can be achieved by using a water-in-oil emulsion as the discontinuous phase and an aqueous fluid as the continuous phase.

[0142] Thus, double emulsions can be produced in two steps. First, a single emulsion is obtained (e.g., water-in-oil), and second, the single emulsion is self-emulsified to obtain a double emulsion. The first and second steps can be performed in tandem, for example, by connecting the output from a first microfluidic device to the input of a second microfluidic device. This example is particularly advantageous because the microfluidic devices are commercially available and the process does not require specialized equipment. Alternatively, the single emulsion can be collected and, optionally, stored after the first and second steps have been performed. Three, four, or more microfluidic devices operating in tandem in the manner described can be used to obtain any number of emulsion layers.

[0143] 4.2 Polynucleotides

[0144] In the present invention, it is desirable for compartments to be formed around polynucleotides or for polynucleotides to be inserted into pre-formed compartments. The methods of the present invention require only a limited number of polynucleotides per compartment. This ensures that the polypeptides of each polynucleotide are separated from other polypeptides. Thus, the binding between the polynucleotide and the polypeptide is highly specific. The enrichment factor is maximized with an average of one or fewer polynucleotides per compartment, separating each polynucleotide from the products of all other polynucleotides, ensuring the tightest possible link between the nucleic acid and the activity of the encoded polypeptide. However, even if the theoretically optimal situation of an average of one or fewer polynucleotides per compartment is not used, a ratio of 5, 10, 50, 100, or even 1000 or more polynucleotides per compartment can be beneficial for sorting large libraries. Subsequent sorting iterations, including new encapsulations with different polynucleotide distributions, allow for more stringent polynucleotide sorting. Preferably, there is a single polynucleotide or fewer per compartment.

[0145] A compartment may contain a single polynucleotide (i.e., a single molecule of a polynucleotide). Alternatively, a compartment may contain multiple polynucleotides (i.e., two or more molecules of a polynucleotide, e.g., two molecules of a polynucleotide). Preferably, when a compartment contains multiple polynucleotides, all of the polynucleotides have the same or substantially identical polynucleotide sequence. When multiple polynucleotides within a compartment have substantially identical sequences, this means that the sequence of one polynucleotide within the compartment does not differ from the sequences of other polynucleotides within the same compartment by more than determined by the cumulative error rate of the method used to synthesize the polynucleotides.

[0146] The number of polynucleotides in a compartment can be statistically controlled by appropriate dilution of the bulk medium containing the polynucleotides prior to the compartmentalization step.

[0147] When the compartments are emulsion droplets, the polynucleotides are contained in the discontinuous phase fluid before emulsification. The number of polynucleotides in the compartments is statistically controlled by appropriate dilution of the discontinuous phase fluid. For example, the discontinuous phase fluid can be diluted so that the average number of polynucleotides in a single droplet is 1 or less. After emulsification, the thus diluted discontinuous phase forms the interior volume of the emulsion droplets, and each droplet thus produced contains 0 or 1 polynucleotide.

[0148] The polynucleotides of the present invention encode cyclic polypeptides. Upon expression, a linear polypeptide can be produced from the polynucleotide. The linear polypeptide can then be subjected to a cyclization process to obtain the desired cyclic polypeptide. Polypeptide cyclization can occur spontaneously, i.e., the polypeptide self-cyclizes. The amino acid sequence of the final cyclic polypeptide may not be the same as the amino acid sequence of the linear peptide from which it is derived. For example, the cyclization process may remove one or more N-terminal residues from the linear peptide. Alternatively, the cyclization process may remove one or more C-terminal residues from the linear polypeptide. In some embodiments, the cyclization process may remove one or more residues from both the N-terminus and C-terminus of the linear polypeptide. A linear polypeptide is sometimes referred to as a "pre-cyclic" polypeptide. The sequence of a cyclic polypeptide may refer to the sequence of amino acids that make up the cyclic polypeptide or the sequence of amino acids contained in the linear polypeptide that will be included in the cyclic polypeptide after cyclization. The sequence of the cyclic polypeptide begins with the amino acid residue that was the N-terminal residue of the cyclic polypeptide sequence contained in the linear polypeptide. In some embodiments, the sequence of the cyclic polypeptide is the same as the sequence of the linear polypeptide.

[0149] A sequence encoding a cyclic polypeptide refers to a sequence of nucleotides in a polynucleotide that encodes a portion of a linear polypeptide that will be included in the cyclic polypeptide after cyclization.

[0150] A polynucleotide is a molecule or construct selected from the group consisting of a DNA molecule, an RNA molecule, a partially or completely artificial nucleic acid molecule consisting of only synthetic bases or a mixture of natural and synthetic bases, any of the foregoing linked to a polypeptide, and any of the foregoing with any other molecule or construct. The other molecule or construct may advantageously be selected from the group consisting of nucleic acids, polymeric materials, particularly beads, e.g., polystyrene beads, and magnetic or paramagnetic materials, such as magnetic or paramagnetic beads. Polynucleotides may also be referred to herein as "nucleic acids" or "nucleic acid molecules."

[0151] The nucleic acid portion of the polynucleotide may include appropriate regulatory sequences, such as promoters, enhancers, translation initiation sequences, polyadenylation sequences, splice sites, etc., necessary for efficient expression of the polypeptide.

[0152] Nucleic acid molecules within the invention include those that encode a polypeptide having a first portion of a split intein, a second portion of the split intein, and a target peptide located between the first portion of the split intein and the second portion of the split intein. In one embodiment of the invention, expression of the nucleic acid molecule results in a polypeptide that spontaneously splices to produce a cyclized form of the target peptide.

[0153] In other embodiments of the invention, expression of the nucleic acid molecule results in a polypeptide that is a splicing intermediate of a cyclized form of the target peptide.

[0154] Nucleic acids of the invention can be prepared by methods for preparing and manipulating nucleic acid molecules generally known in the art (see, e.g., Ausubel et al., Current Protocols in Molecular Biology, New York: John Wiley & Sons, 1997; Sambrook et al., Molecular Cloning: A Laboratory Manual (2nd Edition), Cold Spring Harbor Press, 1989). For example, a nucleic acid molecule of the invention can be made by separately preparing a polynucleotide encoding a first portion of the split intein, a polynucleotide encoding a second portion of the split intein, and a polynucleotide encoding a target peptide. The three polynucleotides can be ligated together to form a nucleic acid molecule encoding a polypeptide having the target peptide sandwiched between the first portion of the split intein and the second portion of the split intein.

[0155] Inteins that are not naturally split (i.e., those that exist as a single continuous chain of amino acids) can be artificially split using known techniques. For example, two or more nucleic acid molecules encoding different portions of such an intein can be created, and their expression can result in two or more artificially split intein components. See, for example, Evans et al., J. Biol. Chem. 274:18359, 1999; Mills et al., Proc. Natl. Acad. Sci. USA 95:3543, 1998. Nucleic acids encoding such non-natural intein components (portions) can be used in the present invention. Nucleic acid molecules encoding non-natural split intein portions that efficiently interact with the same precursor polypeptide to produce a cyclic peptide or splicing intermediate are preferred.

[0156] Examples of non-natural split inteins that can provide such nucleic acid molecules include Psp Pol-1 (Southworth, MW, et al., The EMBO J. 17:918, 1998), Mycobacterium tuberculosis RecA intein (Lew, BM, et al., J. Biol. Chem. 273:15887, 1998; Shingledecker, K., et al., Gene 207:187, 1998; Mills, KV, et al., Proc. Natl. Acad. Sci. USA 95:3543, 1998), Ssp DnaB / Mxe GyrA (Evans, TC et al., J. Biol. Chem. 274:18359, 1999), and Pfu (Otomo et al., Biochemistry 38:16040, 1999; Yamazaki, M., et al., J. Biol. Chem. 274:18359, 1999). et al, J. Am. Chem. Soc. 120:5591, 1998).

[0157] In embodiments, a polynucleotide can be associated with a polypeptide encoded by the polynucleotide by a bond, such as a covalent or non-covalent bond. Examples of non-covalent bonds include ionic bonds, hydrogen bonds, and induced dipole interactions (also known as van der Waals forces). In a preferred embodiment, the bond is a covalent bond. This can be achieved by the mRNA display method described above. In this embodiment, the polynucleotide includes a 3' peptidyl acceptor region, such as a puromycin moiety or an amino acid moiety. When such a polynucleotide is translated, a terminal aminoacyltransferase reaction creates a covalent bond between the nascent polypeptide and the 3' peptidyl acceptor region. In a preferred embodiment, a linker is provided that includes a 5' end that specifically hybridizes to the 3' end of the polynucleotide and a 3' end that includes a peptidyl acceptor region (or moiety).

[0158] Polynucleotides of the invention can also be modified or engineered so that particular codons are located at desired positions, either in the polynucleotide itself or in the mRNA encoded by the polynucleotide. This can be useful in embodiments in which some tRNAs (those carrying the corresponding anticodons) are loaded with non-natural, e.g., non-proteinogenic, residues ("reassigned codon sets"), as described in more detail below. However, engineering codons in polynucleotides of the invention is not required for embodiments that use such non-standard acyl-tRNAs to function.

[0159] 4.3 Expression vectors

[0160] The expression vectors of the present invention can be prepared by inserting a polynucleotide encoding a target peptide into any suitable expression vector capable of promoting expression of the polynucleotide. Such suitable vectors include plasmids, bacteriophage, and viral vectors. Many of these are known in the art, and many are commercially available or available from the scientific community. One of skill in the art can select an appropriate vector for use in a particular application based, for example, on the type of system selected (e.g., in vitro systems, prokaryotic cells such as bacteria, and eukaryotic cells such as yeast or mammalian cells) and the expression conditions selected.

[0161] An expression vector within the present invention can include a series of nucleotides that encodes a target polypeptide and a series of nucleotides that act as a regulatory domain that regulates or controls the expression (e.g., transcription) of the nucleotide sequence within the vector. For example, the regulatory domain can be a promoter or an enhancer.

[0162] Expression vectors within the present invention can include nucleotide sequences encoding peptides that facilitate screening of the cyclized form of the target peptide or splicing intermediate for a particular property (e.g., a DNA-binding domain, a chitin-binding domain or an affinity tag such as a biotin tag, a colored or luminescent label, a radioactive tag, etc.) or purification of the cyclized form of the target peptide or splicing intermediate (e.g., a chitin-binding domain, an affinity tag such as a biotin tag, a colored or luminescent label, a radioactive tag, etc.).

[0163] In preferred embodiments, expression vectors of the present invention are generated with restriction sites between and within the nucleic acid sequences encoding the split intein portions, allowing for the cloning of various circularization targets or splicing intermediates. In some embodiments, expression vectors of the present invention can be inducible expression vectors, such as arabinose-inducible vectors. The inducer can be permeable to the compartment material of the present invention.

[0164] 4.4 Polypeptides

[0165] Several methods of polypeptide cyclization are known in the art. Polypeptide cyclization can be performed between two side chains, between a side chain and a terminal group (i.e., the N-terminus or C-terminus), or between two terminal groups (i.e., "head-to-tail" or "backbone" cyclization). One such side chain-to-side chain cyclization method can be performed enzymatically. Numerous enzymes exist that cyclize peptide sequences. For example, transglutaminase can catalyze an aminotransferase reaction between a glutamine side chain and a lysine side chain, resulting in a covalent isopeptide bond between the two side chains. When glutamine and lysine are present on the same polypeptide, the polypeptide is cyclized by this reaction. Backbone cyclization can also be induced enzymatically, for example, by treatment with subtilisin. Other non-limiting examples include ProcM and PatG. Generally, these methods rely on the presence of a "leader" sequence of amino acids within the polypeptide. The leader sequence recruits and directs the action of the cognate enzyme by interacting specifically with the enzyme. Each enzyme may be specific for a particular leader sequence. The leader sequence of a particular enzyme may be added to a polypeptide of the invention through manipulation of the polynucleotide encoding it, for example, by conventional molecular cloning techniques that are within the skill of the art.

[0166] Macrocyclization methods have been described in detail by Bashiruddin and Suga, Curr. Op. in Chem. Bio., vol. 24, pp. 131-138, particularly in relation to mRNA display, but those skilled in the art will be familiar with the use of these methods outside of mRNA display technology. The most basic approach to synthesizing macrocyclic peptides using the translational machinery relies on disulfide bonds between cysteine ​​residues. However, this is susceptible to reduction in the intracellular environment, making it undesirable for some applications. Therefore, methods have been devised to form non-reducible covalent bonds for cyclization by simple chemical post-translational modifications. Macrocyclic peptides have been successfully generated using dibromoxylene to bridge two primary amines between the N-terminus and lysine side chains using disuccinimidyl glutarate or the sulfhydryl groups of two cysteine ​​residues. A similar method for generating bicyclic peptides via thioether bridges of three cysteine ​​residues has also been reported (Heinis & Winter, Nat Chem Biol, 5 (2009), pp. 502-507). The advantage of these methods is their applicability to standard proteinogenic amino acids. However, the appearance of more than three reactive residues in the random regions of these libraries can scramble the cross-linking patterns, making it difficult to deconvolute the results of selections based on these cyclization methods.

[0167] Although technically more challenging than those mentioned above, a much more reliable method for constructing macrocyclic peptides is based on the concept of genetic code manipulation, known as genetic code reprogramming, in which designated codons are emptied and then reassigned to non-proteinogenic amino acids. Two main methodologies have been reported so far, both of which utilize custom-made reconstituted translation systems.

[0168] One method utilizes the mischarge property of aminoacyl-tRNA synthetases in the presence of excess non-proteinogenic amino acids to generate the corresponding aminoacyl-tRNA. Szostak et al. reported a method for generating peptides containing 4-selenalysine in the peptide chain, followed by selective oxidation and simultaneous removal of the seleno group to generate dehydroalanine residues. The dehydroalanine then reacts with the sulfhydryl group of cysteine ​​via a Michael addition to form a thioether bond, resulting in a lanthionine-like macrocyclic peptide.

[0169] The other approach involves a "flexible" tRNA acylation ribozyme, known as flexizyme, developed by Suga et al., which facilitates the preparation of a wide range of nonproteinogenic aminoacyl-tRNAs with nearly unlimited options. Combining flexizyme with a custom-made in vitro translation system, called the FIT (flexible in vitro translation) system, enables the ribosomal synthesis of macrocyclic peptides using nonproteinogenic amino acids that can be crosslinked with other proteinogenic or nonproteinogenic residues. The FIT system allows for various cyclization methods; for example, methyllanthionine-like macrocyclic peptides can be synthesized by incorporating vinylglycine, which is thermally isomerized to dehydrobutyrine, which can form a thioether bond with a cysteine ​​residue (Y. Goto, K. Iwasaki, K. Torikai, H. Murakami, H. Suga; Chem Commun (Camb), 23 (2009), pp. 3419-3421). Translation of peptides containing a benzylamine group designated by the N-terminal initiating amino acid and a downstream 5-hydroxytryptophan provides a unique method for mild oxidative macrocyclization to form fluorescent indole bonds (Y. Yamagishi, H. Ashigai, Y. Goto, H. Murakami, H. Suga; ChemBioChem, 10 (2009), pp. 1469-1472). Furthermore, ribosomal synthesis of head-to-tail linked peptides can also be achieved using programmed peptidyl-tRNA drop-off containing a C-terminal Cys-Pro-HOG (glycolic acid) sequence or a C-terminal Cys-Pro sequence (T. Kawakami, Nat Chem Biol, 5 (2009), pp. 888-890; Y. Ohshiro, ChemBioChem, 12 (2011), pp. 1183-1187; TJ Kang, Angew Chem Int Ed Engl, 50 (2011), pp. 2159-2161). In both cases, the C-terminal ester bond accelerates the self-rearrangement of N→S migration to form a C-terminal diketopiperazine thioester, ultimately generating a backbone-cyclized monocyclic or disulfide-bridged bicyclic peptide.

[0170] The most convenient and reliable cyclization method is based on the translation of peptides with an N-chloroacetyl-amino acid initiator capable of reacting with downstream cysteines (Y. Goto, A. Ohta, Y. Sako, Y. Yamagishi, H. Murakami, H. Suga; ACS Chem Biol, 3 (2008), pp. 120-129). The advantage of this method is the spontaneous and selective formation of a thioether bond between the N-terminal chloroacetyl group and the sulfhydryl group of the nearest cysteine ​​residue. The only exception is that the cysteine ​​residue adjacent to the N-chloroacetyl amino acid cannot react with the chloroacetyl group due to ring constraints, leaving a free sulfhydryl group at this position. However, this selectivity has been shown to provide a convenient method for translating fused bicyclic peptides with a thioether (sulfide) bond between the N-terminus and the second cysteine ​​and a disulfide bond between the first and third cysteine ​​residues. Importantly, the FIT system has been demonstrated to promote translation of peptides containing D-amino acids, N-methyl amino acids, N-alkylglycines, and those with non-canonical side chains.

[0171] In other methods, cyclization of a polypeptide can be achieved spontaneously by intramolecular interactions within the polypeptide, for example, by fusing sequences derived from an "intein" to the desired cyclic polypeptide.

[0172] Numerous methods for generating nucleic acids encoding peptides of known or random sequence are known in the art. For example, polynucleotides having predetermined or random sequences can be prepared chemically by solid-phase synthesis using commercially available equipment and reagents. The polymerase chain reaction can also be used to prepare polynucleotides of known or random sequence. See, e.g., Ausubel et al., supra. As another example, restriction endonucleases can be used to enzymatically digest larger nucleic acid molecules or entire chromosomal DNA into multiple smaller polynucleotide fragments that can be used to prepare the nucleic acid molecules of the present invention.

[0173] A polynucleotide encoding a peptide sequence to be cyclized is preferably prepared so that one end of the polynucleotide encodes an asparagine, serine, cysteine, or threonine residue that promotes the cyclization reaction. For the same reason, a polynucleotide encoding a peptide sequence for generating a splicing intermediate is preferably prepared so that the end encodes an amino acid other than an asparagine, serine, cysteine, or threonine residue that prevents the cyclization reaction.

[0174] Once generated, the nucleic acid molecule encoding the intein portion can be ligated to a nucleic acid molecule encoding the target peptide (or a peptide within a splice intermediate) using conventional methods to form a larger nucleic acid molecule encoding a polypeptide having the order first intein portion-target peptide-second intein portion (see, e.g., Ausubel et al., supra).

[0175] 4.4.1 SICLOPPS

[0176] The trans-splicing ability of split inteins has been exploited to develop a general method for generating cyclic peptides and splicing intermediates that display the peptide in a loop structure (PCT / US1999 / 030162). In this method, a target peptide is inserted between two portions of a split intein in a precursor polypeptide. The two portions of the split intein physically join together to form an active intein, a structure that further forces the target peptide into a loop configuration. In this configuration, one of the intein portions (e.g., I N ) and the target peptide, ester isomers of the amino acids at the junction between the intein and the target peptide are stabilized, and other parts of the intein (e.g., I C The heteroatom from the intein is then able to react with the ester to form a cyclic ester intermediate. The active intein then catalyzes the formation of an aminosuccinimide, which liberates the cyclized form of the target peptide (i.e., the lactone form), which then spontaneously rearranges to form the thermodynamically favored backbone cyclic peptide product (i.e., the lactam form).

[0177] By stopping the reaction at a predetermined point before the release of a cyclic peptide, a splice intermediate can be generated that retains the target peptide in a loop configuration. To generate such peptides, nucleic acid molecules can be constructed that encode polypeptides in which the target peptide sequence is sandwiched between two intein moieties. Introduction of these constructs into expression vectors provides a method for producing the polypeptide in an appropriate expression system, where the polypeptide can be spliced ​​into a cyclic peptide or splice intermediate. Using this method, several different cyclic peptides or splice intermediates can be prepared, generating a library of cyclic or partially cyclic peptides that can be screened for specific properties.

[0178] An intein is a protein segment that can cleave itself and join the remaining portion (an extein) via a peptide bond in a process called "protein splicing." Intein-mediated protein splicing occurs after an intein-containing mRNA is translated into a protein. The precursor protein contains three segments: an N-extein, followed by an intein, followed by a C-extein. After splicing occurs, the resulting protein contains the N-extein linked to the C-extein. In some cases, the inteins in the precursor protein are derived from two genes. In this case, the intein is called a split intein. The intein portion (or fragment) encoded by one gene interacts with the intein fragment encoded by the other gene to generate a catalytically active intein, which then cleaves itself and splices the exteins from the two genes together.

[0179] If two split intein fragments are instead located at opposite ends of an intervening polypeptide sequence, a splicing and cleavage process will produce a circular polypeptide with the sequence of the original intervening polypeptide. This can be achieved using traditional molecular genetic techniques, for example, by cloning sequences encoding two complementary split intein fragments into a vector containing a sequence encoding the desired circular polypeptide. The three polynucleotide sequences are positioned in the vector so that upon expression, the resulting polypeptide will have an N-terminal intein fragment followed by a circular polypeptide sequence followed by a C-terminal intein fragment, and the N- and C-terminal intein fragments can associate to form a functional intein that subsequently catalyzes a splicing reaction to produce the circular polypeptide. In other words, the polynucleotides contain a sequence encoding the N-terminal intein fragment followed by a sequence encoding the circular polypeptide sequence followed by a sequence encoding the C-terminal intein fragment. This process is known as split intein circular ligation of peptides and proteins (SICLOPPS).

[0180] Expression from a polynucleotide can be expression of a polypeptide from the polynucleotide. Expression from a polynucleotide can also be expression of a second polynucleotide from a first polynucleotide. For example, an RNA polynucleotide can be expressed from a DNA polynucleotide by the process of transcription by an appropriate RNA polymerase. A polypeptide can be expressed from an RNA polynucleotide by the process of translation by an appropriate ribosome. In some usages, expression from a polynucleotide refers to the ultimate expression of a polypeptide from the genetic material encoding it, i.e., both transcription into RNA and translation into a polypeptide. The intended usage will be clear from the context.

[0181] One suitable split intein is the Npu split intein from dnaE. An example of a polypeptide according to the invention comprising an Npu split intein may have the following sequence: HHHHHHGENLYFKLQAMGMIKIATRKYLGKQNVYDIGVERYHNFALKNGFIASNX ~~~~~ CLSYDTEILTVEYGILPIGKIVEKRIECTVYSVDNNGNIYTQPVAQWHDRGEQEVFEYCLEDGCLIRATKDHKFMTVDGQMMPIDEIFERELDLMRVDNLPNGTAANDENYALAA where X ~~~~~ is the cyclic peptide to be produced, X is C, S, T or any other amino acid, and ~ " indicates an amino acid in the cyclic peptide sequence. It will be apparent to one skilled in the art that any sequence may be inserted after the "X" in the above sequences. The sequence may be one or more amino acids in length, in embodiments at least three or more amino acids in length, and preferably at least six amino acids in length.

[0182] The above array contains the following components: 1: HHHHHH An optional hexa-histidine tag aids purification, for example on a nickel-NTA column. Other purification systems are contemplated, such as "FLAG-TAG," in which the hexa-histidine is replaced with a suitable tag sequence. 2: GENLYFKLQAMGMIKIATRKYLGKQNVYDIGVERYHNFALKNGFIASN Contains a C-terminal intein fragment. 3: X ~~~~~ Cyclic polypeptide sequences. 4: CLSYDTEILTVEYGILPIGKIVEKRIECTVYSVDNNGNIYTQPVAQWHDRGEQEVFEYCLEDGCLIRATKDHKFMTVDGQMMPIDEIFERELDLMRVDNLPNGTAANDENYALAA Contains the N-terminal intein fragment.

[0183] 4.5 Decorating the content of regions

[0184] 4.5.1 In vitro transcription and translation

[0185] To co-compartmentalize a cyclic polypeptide and a polynucleotide encoding it, a means for expressing the polypeptide from the polynucleotide can also be enclosed within the compartment. Such a means can include an in vitro transcription and translation (IVTT) system. The IVTT system can include, for example, RNA polymerase, ribosomes, phosphate nucleotides, tRNAs loaded with amino acids, and translation factors such as initiation and elongation factors. Suitable in vitro transcription / translation reagents are well known in the art (e.g., Isalan, M. et al. (2005) PLoS Biol. 3 e64). Thus, expression of the polynucleotide within the compartment results in the co-location of the polypeptide and polynucleotide. When the compartments are droplets of an emulsion, this can be achieved by including IVTT components in the discontinuous (or dispersed or internal) phase fluid prior to emulsification.

[0186] Furthermore, acylated tRNAs precharged with a desired non-proteinogenic amino acid (or hydroxy acid) (i.e., with an activated amino acid bound thereto) can be added to a reconstituted cell-free translation (IVTT) system containing only limited natural amino acids. By correlating the codon for the excluded natural amino acid with the anticodon of a tRNA acylated with a non-proteinogenic amino acid (or hydroxy acid), a peptide containing the non-proteinogenic amino acid (or hydroxy acid) can be synthesized by ribosome translation based on the genetic information of mRNA encoded by or constituting the polynucleotide of the present invention. Alternatively, a peptide free of a natural amino acid can be synthesized by translation by adding only acylated tRNAs charged with a non-proteinogenic amino acid (or hydroxy acid) to a reconstituted cell-free translation system free of natural amino acids.

[0187] Acylated tRNAs charged with nonproteinogenic amino acids (or hydroxy acids) can be prepared using the artificial RNA catalyst "Flexizyme," which can catalyze aminoacyl-tRNA synthesis. As described above, these artificial RNA catalysts can charge amino acids with any side chain and recognize only the consensus sequence 5'-RCC-3' (R = A or G) at the 3' end of tRNA to acylate the 3' end of the tRNA. Therefore, they can act on any tRNA with a different anticodon. Furthermore, Flexizyme can charge tRNAs with not only L-amino acids but also hydroxy acids (having a hydroxyl group at the α-position), N-methyl amino acids (having an N-methyl amino acid at the α-position), N-acyl amino acids (having an N-acyl amino group at the α-position), D-amino acids, and other amino acids. Detailed explanations are provided in Y. Goto and H. Suga (2009) "Translation initiation with initiator tRNA charged with exotic peptides," Journal of the American Chemical Society, Vol. 131, No. 14, pp. 5040-5041; WO2008 / 059823 "Translation and synthesis of polypeptides having non-natural structures at the N-terminus and their applications," Goto et al., ACS Chem. Biol., 2008, 3, 120-129; and WO2008 / 117833 "Method for synthesizing cyclic peptide compounds," the disclosures of which are incorporated herein by reference in their entireties.

[0188] Other unnatural amino acids that can be ligated to tRNA using Flexizyme include amino acids with various side chains, β (beta) amino acids, γ (gamma) amino acids, and δ (delta) amino acids, D amino acids, and derivatives with substituted amino or carboxyl groups on the amino acid backbone. Furthermore, abnormal peptides obtained by incorporating unnatural amino acids can have backbone structures other than the usual amide bond. For example, abnormal peptides include depsipeptides composed of amino acids and hydroxy acids, polyesters produced by sequential condensation of hydroxy acids, peptides methylated at the nitrogen atom of the amide bond by introducing N-methyl amino acids, and peptides with various acyl groups (acetyl, pyroglutamic acid, fatty acids, etc.) at the N-terminus. Furthermore, cyclic peptides can also be synthesized by cyclizing a non-cyclic peptide composed of an amino acid sequence having a functional group pair capable of forming a bond at the opposite two ends (or, when N-methyl peptides are used, cyclic N-methyl peptides can be obtained). Cyclization can occur under the conditions of a cell-free translation (IVTT) system with any functional group pair, as exemplified by a cyclic peptide cyclized via a thioether bond obtained by translation / synthesis of a peptide sequence with a chloroacetyl group and a cysteine ​​group at the opposite two termini.

[0189] 4.5.2 Polynucleotide Amplification

[0190] It can be difficult to generate enough polypeptides from a single polynucleotide to provide a detectable signal in a subsequent assay. In embodiments, multiple copies of a polynucleotide are generated within a compartment. This can be achieved by an amplification process, such as polymerase chain reaction (PCR) or the use of Phi29 polymerase. If the components of the amplification process are contained within the compartment along with the polynucleotide, amplification can be carried out by subjecting the compartment to the necessary amplification conditions, such as thermal cycling or incubation under heat. The compartment thus becomes a self-contained microreactor, similar to that described above for transcription and translation of polynucleotides. Because the compartment forms a barrier to polynucleotide movement, all copies of the amplified polynucleotide are contained and separated within the compartment. Thus, the compartment becomes a monoclonal unit, i.e., a co-localized unit of identical or substantially identical copies of a polynucleotide.

[0191] Where the compartments are emulsion droplets, this is achieved by including the components of the amplification reaction in the discontinuous phase fluid prior to emulsification. The emulsion can then be subjected to the conditions necessary to amplify to the desired copy number.

[0192] Primers for PCR can be selected or designed to amplify an entire desired sequence of a polynucleotide (an "amplicon"). For example, one primer can be designed to anneal to the beginning of the desired sequence on the template strand, and a second primer can be designed to anneal to the beginning of the desired sequence on the coding strand.

[0193] If the polynucleotide encodes a SICLOPPS self-splicing polypeptide, for example, a first primer can anneal to a sequence encoding an N-terminal intein fragment on the template (or antisense) strand, and a second primer can anneal to a sequence encoding a C-terminal intein fragment on the coding (or sense) strand.

[0194] Alternatively, vectors can be selected or designed to contain specific sequences for primer annealing. For example, a sequence complementary to a first primer can be inserted after the desired sequence on the template strand of the vector (i.e., after the 3' end of the desired sequence on the template strand), and a sequence complementary to a second primer can be inserted after the desired sequence on the coding strand of the vector (i.e., after the 3' end of the desired sequence on the coding strand). This option is particularly suitable when the desired sequence to be amplified is unknown, such as when inserting randomized polynucleotides into a vector to construct a gene library. In this case, primers can be selected or designed so that, upon translation of the amplified polynucleotide, a polypeptide sequence with specific properties is generated. For example, one primer can be designed to encode the N-terminal glutamine donor sequence Ala-Leu-Gln, and a second primer can be designed to encode the C-terminal region lysine as a substrate for transglutaminase.

[0195] Preferably, the amplification reagent is an isothermal amplification reagent. The method of isothermal amplification in agarose gel is well known in the art, and includes multiply-primed RCA using Phi29 DNA polymerase, which generates high molecular weight (more than 40 kb) highly branched products containing copies of amplified polynucleotides (Michikawa, Y. et al. (2008). Anal. Biochem., 383, 151-158). Amplified DNA, especially highly branched amplified DNA, cannot diffuse out of the bead matrix.

[0196] If the polynucleotide and encoded product remain co-localized without compartmentalization, the beads may be contacted with an aqueous expression solution without emulsification. The ability of the polynucleotide and encoded product to remain co-localized depends on the concentration and biophysical properties (e.g., size) of the encoded product.

[0197] In embodiments, polynucleotides are amplified using Phi29 DNA polymerase. In such embodiments, the primers may be random primers, such as polynucleotide hexamers having a random sequence of six nucleotides. Alternatively, the primers may be designed or selected in a manner similar to that described for PCR above. An advantage of Phi29-based amplification is that no thermal cycling is required; amplification can be performed with Phi29 by simple incubation of the reaction mixture.

[0198] In embodiments, Phi29 amplification is carried out by incubating the compartments for 1 hour or more. Preferably, amplification is carried out for 8 hours or more. Most preferably, amplification is carried out for 16 hours or more. Incubation of the Phi29 amplification reaction may be carried out for 1, 2, or 3 days or more.

[0199] In an embodiment, incubation for Phi29 amplification may be performed at room temperature, or Phi29 amplification may be incubated under heating at 20°C to 50°C, most preferably 30°C to 40°C, for example 37°C.

[0200] However, the conditions required to perform an amplification reaction may not be compatible with IVTT. Therefore, in order to perform IVTT on the amplified polynucleotides in the compartment, it may be necessary to manipulate the compartment after amplification is complete to change the internal environment. Methods in the art, such as the method of Brouzes et al. (PNAS, 2009, 106:34, pp. 14195-14200), change the conditions within microfluidic droplets through the process of droplet merging. However, droplet merging is a technically challenging process and requires specialized equipment and expertise to be performed reliably.

[0201] 4.5.3 Gel transition

[0202] In embodiments, the present invention avoids the drawbacks associated with droplet merging by utilizing a gel-forming material within the compartments, as in WO 2012 / 156744 A2. The gel-forming material can be engineered to undergo a reversible transition from a liquid phase to a solid or gel phase to immobilize the amplified polynucleotides. Once solidified within the compartments, the gel can form beads containing the immobilized polynucleotides. The compartments can then be disrupted without affecting the integrity of the single clones of amplified polynucleotides captured in the gel.

[0203] The process for breaking down the compartments varies depending on the compartment material used. In some embodiments, a demulsifier, such as a weak surfactant, can be added to the emulsion to separate the phases and remove the aqueous phase containing the beads. The demulsifier competes with the surfactant at the oil / water interface, disrupting it, also known as "breaking the emulsion." Suitable weak surfactants include perfluorooctanol (PFO) and other fluorous compounds with small hydrophilic groups when using fluorinated oils, or buffers containing SDS and Triton and other compounds with carbon chains on one side and small hydrophilic groups on the other when using mineral oils. The demulsifier can be added to the emulsion and the mixture can be stirred, for example, with a pipette.

[0204] In other embodiments, the emulsion may be centrifuged to separate the phases and remove the aqueous phase containing the beads. Suitable techniques for re-emulsifying gel beads are known in the art (Abate, AR et al (2009) Lab Chip, 9, 2628-2631).

[0205] After demulsification, the beads may be isolated and / or washed to remove buffer and other reagents. The beads may be isolated and / or washed by centrifugation or filtration using standard techniques.

[0206] After washing, the beads may be immediately subjected to other steps of the methods described herein or may be stored, for example, at room temperature or refrigerated or frozen (preferably in the presence of glycerol). In embodiments in which the beads contain viable cells, the beads may be treated with a preservative such as glycerol prior to freezing, according to known techniques.

[0207] Decompartmentalized gel beads are porous, and the internal conditions of the gel can be altered by suspending them in different media. Thus, the beads can be removed from conditions suitable for performing PCR and transferred to conditions suitable for performing IVTT without losing monoclonality. The environment within the beads is then suitable for IVTT of polynucleotides.

[0208] A gel-forming agent is an agent, eg, a polymer such as a polysaccharide or polypeptide, that can be solidified from a liquid into a gel upon a change in conditions, such as heating, cooling, or a change in pH.

[0209] Suitable gel-forming agents include alginate, gelatin, and agarose, as well as other gels with a sol phase sufficiently fluid to move through the channels of a microfluidic device. Preferably, agarose, a linear polymer composed of repeating units of disaccharides (D-galactose and 3,6-anhydro-L-galactopyranose), is used. The gel-forming agent can be solidified into beads by any convenient method, for example, by changing the conditions. Preferably, the hydrogel-forming agent is solidified by changing the temperature, for example, by cooling. Hyaluronic acid is another gel-forming agent that can be used in the present invention.

[0210] In a preferred embodiment, the gel can be induced to change phase from a liquid form to a solid form by cooling below the phase change (or transition) temperature. For example, if the agent is agarose, it can be solidified by lowering the temperature, e.g., below 25°C, below 20°C, or below 15°C. The exact gelling temperature depends on the type and concentration of agarose and can be readily determined by one of skill in the art. For example, the gelling point of 0.5% to 2% ultra-low melting point agarose Type IX-A (Sigma) is approximately 17°C.

[0211] Solidification of the agent within the compartments causes the solidified gel to assume the shape of a bead. The population of solidified beads may be monodisperse. The encoded product may be retained within the beads by any convenient method. For example, the product may be retained within the beads by entrapment within the gel matrix, or by covalent or non-covalent bonding to a retention agent or the gel matrix itself.

[0212] Molecules such as polypeptides and polynucleotides can be retained within the beads due to their size. For example, molecules larger than a threshold size may be unable to diffuse out of the beads through the pores of the gel and thus be trapped within the gel matrix of the beads. For example, polynucleotides such as plasmids and amplified copies of polynucleotides may be retained within the beads. In some embodiments, the gel may retain particles with diameters of 50 nm or greater, although the exact threshold depends on several factors, including the type and concentration of the gel.

[0213] In some embodiments, the aqueous solution in which the gel beads solidify after emulsification, e.g., a solution containing a hydrogel-forming agent, such as the aqueous expression and / or amplification solution described above, may further comprise one or more retention agents that, e.g., bind to the polynucleotide or encoded product, thereby reducing or preventing diffusion of the polynucleotide or encoded product from the bead.

[0214] In other embodiments, compartment components such as substrates and encoded products may be held by direct binding to the hydrogel scaffold, for example, the scaffold may be engineered to contain one or more binding sites that bind to droplet components and hold them within the beads.

[0215] In other embodiments, the polynucleotide and / or encoded product may be sufficiently retained within the bead without the need for binding to a retention agent or gel scaffold.

[0216] The gel-forming agent may be incorporated into the compartment at any stage before the compartment is disrupted.

[0217] 4.6 Libraries

[0218] A library of genetically tagged cyclic peptides within the compartment can be generated. Randomized polynucleotides of desired length can be cloned into a vector for amplification and expression, for example, by the methods described above, prior to compartmentalization of the vector within a single copy. For example, the vector shown in Figure 15 has a compartment labeled "Library," which is the location of the randomized polynucleotide in this embodiment.

[0219] The length of the randomized polynucleotide inserted into the vector will depend on various factors that can be determined by one of skill in the art. A primary consideration is the size of the final polypeptide to be expressed. In a preferred embodiment, the polypeptide is six amino acids long. Therefore, a suitable randomized polynucleotide is 18 nucleic acids long. When forming a cyclic peptide, consideration must be given to whether the length of the polypeptide is sufficient to allow the cyclization reaction to proceed, i.e., whether the length allows the formation of a closed peptide cycle. In embodiments, peptides are cyclized with linkers of any length. Thus, a cyclic polypeptide may be achieved by encoding only two amino acids, in which case the randomized polynucleotide would be at least six nucleic acids long. Another consideration is the maximum insert size allowed by the vector and corresponding replication system. In embodiments, the randomized sequence may be even longer, e.g., at least 9, 30, 60, 90, 180, 300, 600, 900, 1800, 3,000, or more nucleic acids long. In preferred embodiments, the randomized nucleotide sequence has a length of 6, 9, 12, 15, 18, 21, 24, 27, or 30 nucleotides. While the randomized sequence encodes a polypeptide, its length is not necessarily a multiple of three. For example, it can be 7, 8, 10, 11, 13, 14, 16, 17, 19, 20, 22, 23, 25, 26, 28, or 29 nucleotides in length. A randomized polynucleotide sequence may be referred to herein as a variable sequence. In embodiments, one or more positions of a "random" or "variable" sequence may be fixed in nature. For example, in embodiments where cyclization is achieved by the SICLOPPS method, the first position may be occupied by an invariant cysteine, serine, or threonine residue, followed by a variable or random amino acid sequence.

[0220] It will also be understood that the libraries of the present invention may contain components with randomized sequences of different lengths. For example, a certain percentage of library components may contain randomized sequences with a length of 9 nucleotides, while another percentage of the same library components may contain randomized sequences with a length of 19 nucleotides. Any number of different lengths may be present in the same library.

[0221] Each individual compartment then functions as a microreactor for amplification, e.g., by including components for Phi29 amplification in the vector medium prior to compartmentalization. In embodiments utilizing gels, the compartments are subjected to conditions that solidify the gel, at which point the compartments may be disrupted and the medium conditions for IVTT changed. The gel beads are optionally recompartmentalized, with each bead now functioning as a microreactor for IVTT of the immobilized polynucleotides, resulting in a library. Concurrent with or following IVTT, some or all of the components necessary to carry out the selection protocol may be introduced into the beads and / or capsules.

[0222] The integrity of the association between the polypeptide and the encoding polynucleotide can be further enhanced by utilizing the mRNA display technology described above in the compartmentalization and gel transfer protocols of the present invention. In this embodiment, the polynucleotide is structured so that it will be linked to the resulting polypeptide by the process of translation. For example, the 3' end of the polynucleotide can contain a peptidyl acceptor region, such as a puromycin or amino acid moiety.

[0223] It will be understood that the display technologies described herein (e.g., mRNA display, phage display, compartmentalized randomized polynucleotides) are not necessarily mutually exclusive and may be implemented together within the same embodiment. For example, a polynucleotide of the invention may encode a SICLOPPS polypeptide while including a peptidyl acceptor region that forms a covalent bond to the nascent peptide upon completion of translation of the polynucleotide.

[0224] Furthermore, compartmentalization of an mRNA library by any of the methods described herein allows functional assays to be performed on the mRNA library, thereby overcoming one of the major drawbacks of mRNA display technology.

[0225] 4.7 Selection

[0226] Libraries according to the invention may be screened by subjecting them to assay conditions, with each bead or compartment optionally possessing the characteristics of a positive or negative assay result. Beads or compartments may then be selected based on this characteristic. For example, fluorescent signals may be selected by FACS / FADS.

[0227] All reporters, labels, and tags disclosed herein may be used in any of the embodiments disclosed herein.

[0228] 4.7.1 Affinity Selection

[0229] When selecting a polypeptide with affinity for a specific ligand, the polynucleotide can be linked to the polypeptide within the microcapsule via the ligand. For example, the ligand can be covalently attached to the polynucleotide via a reaction between the ligand and the 3'-OH or 5'-phosphate of the polynucleotide, etc. As used herein, "ligand" can refer to any entity that binds to or can be bound by another entity. For example, the ligand can be another polypeptide containing a receptor binding site. In this format, the peptide of the invention is selected based on the strength of its interaction with the ligand polypeptide, ideally via the receptor binding site. Alternatively, the ligand can be a small molecule, another polynucleotide (e.g., an aptamer), or a macroscale physical structure such as a polystyrene bead or a magnetic bead. These examples are not intended to be limiting.

[0230] Only polypeptides with affinity for the ligand will bind to the polynucleotide, and only polynucleotides having a polypeptide bound via a ligand will acquire altered optical properties that allow them to be retained in the selection step. Thus, in this embodiment, the polynucleotide comprises a nucleic acid encoding a polypeptide linked to a ligand for the polypeptide.

[0231] The change in the optical properties of the polynucleotide upon binding of the polypeptide with a ligand can be induced in a variety of ways, including: (1) The polypeptide itself may have unique optical properties, for example, it is fluorescent (e.g., green fluorescent protein (Lorenz et al., 1991)). (2) The optical properties of a polypeptide may be modified upon binding to a ligand; for example, the fluorescence of the polypeptide may be quenched or enhanced upon binding (Guixe et al., 1998; Qi and Grabowski, 1998). (3) The optical properties of the ligand may change upon binding to the polypeptide; for example, the fluorescence of the ligand may be quenched or enhanced upon binding (Voss, 1993; Masui and Kuramitsu, 1998). (4) The optical properties of both the ligand and the polypeptide may be modified upon binding, e.g., fluorescence resonance energy transfer (FRET) from the ligand to the polypeptide (or vice versa), resulting in emission at the "acceptor" emission wavelength when excitation is at the "donor" absorption wavelength (Heim & Tsien, 1996; Mahajan et al., 1998; Miyawaki et al., 1997).

[0232] In this embodiment, the binding of the polypeptide to the polynucleotide via the ligand does not need to directly induce a change in optical properties. All polypeptides to be selected can contain a putative binding domain to be selected and a tag as a common feature. The polynucleotide within each microcapsule is physically linked to the ligand. If the polypeptide generated from the polynucleotide has affinity for the ligand, it will bind to the ligand and become physically linked to the same polynucleotide that encoded the polypeptide, leaving the polynucleotide "tagged."

[0233] At the end of the reaction, all microcapsules may be combined, pooling all polynucleotides and polypeptides together in one environment. Polynucleotides encoding polypeptides exhibiting the desired binding can be selected by adding a reagent that specifically binds to or reacts specifically with the "tag," thereby inducing a change in the optical properties of the polynucleotides and allowing sorting. For example, a fluorescently labeled anti-"tag" antibody can be used, or an anti-"tag" antibody followed by a second fluorescently labeled antibody that binds to the first can be used.

[0234] In other embodiments, polynucleotides can be sorted based on the fact that polypeptides that bind to a ligand simply mask the ligand from other binding partners that would otherwise modify, for example, the optical properties of the polynucleotide, in which case polynucleotides with unmodified optical properties are selected.

[0235] In another embodiment, the present invention provides a method in which a polypeptide binds to a polynucleotide encoding it. The polypeptide, along with the attached polynucleotide, is then sorted as a result of ligand binding to a polypeptide with a desired binding activity. For example, all polypeptides can include a constant region that binds covalently or noncovalently to a polynucleotide and a second region that is diversified to generate the desired binding activity.

[0236] In other embodiments, the ligand of the polypeptide is itself encoded by a polynucleotide and binds to the polynucleotide. In other words, the polynucleotide encodes two (or indeed more) polypeptides, at least one of which binds to the polynucleotide and potentially each other. Only if the polypeptides interact within the compartment is the polynucleotide modified in a manner that ultimately results in a change in its optical properties that allows its sorting. This embodiment is used, for example, to search a gene library for pairs of genes that encode pairs of proteins that bind to each other. Each polypeptide may be encoded by an individual polynucleotide.

[0237] Tyramide Signal Amplification(TSA) TM ) amplification may be used to enhance fluorescence and render the polynucleotide fluorescent. This involves peroxidase catalyzing the conversion of fluorescein-tyramine to a free radical form that binds to the polynucleotide (links to another protein) and reacts (locally) with the polynucleotide. Methods for performing TSA are known in the art, and kits are commercially available from NEN.

[0238] The TSA may be configured to result in a direct increase in the fluorescence of the polynucleotide, or a ligand may be attached to the polynucleotide that is bound to a second fluorescent molecule, or an array of molecules, one or more of which are fluorescent.

[0239] 4.7.2 Catalyst Selection

[0240] If the selection is for catalysis, the polynucleotide in each microcapsule can contain a substrate for the reaction. If the polynucleotide encodes a polypeptide capable of acting as a catalyst, the polypeptide catalyzes the conversion of the substrate to a product. Thus, at the end of the reaction, the polynucleotide is physically linked to the product of the catalyzed reaction.

[0241] In some cases, it may be desirable for the substrate not to be a component of a polynucleotide. In this case, the substrate contains an inert "tag" that requires further activation, such as photoactivation (e.g., of a "caged" biotin analog (Sundberg et al., 1995; Pirrung and Huang, 1996)). A catalyst to be selected then converts the substrate to product. The "tag" is then activated, and the "tagged" substrate and / or product is bound to a tag-binding molecule (e.g., avidin or streptavidin) complexed with the nucleic acid. Thus, the ratio of substrate to product attached to the nucleic acid via the "tag" reflects the ratio of substrate and product in solution.

[0242] The optical properties of the polynucleotide to which the product is attached and which encodes a polypeptide having the desired catalytic activity can be modified by any of the following: (1) A product-polynucleotide complex that has distinctive optical properties not found in the substrate-polynucleotide complex, for example, because of: (a) substrates and products with different optical properties (many fluorogenic enzyme substrates are commercially available (see, e.g., Haugland, 1996), including substrates for glycosidases, phosphatases, peptidases, and proteases (Craig et al., 1995; Huang et al., 1992; Brynes et al., 1982; Jones et al., 1997; Matayoshi et al., 1990; Wang et al., 1990)), or (b) Substrates and products that have similar optical properties, but only the products, and not the substrates, bind to or react with the polynucleotide. (2) Addition of reagents that specifically bind or react with the products, thereby inducing a change in the optical properties of the polynucleotides that allows for sorting (these reagents can be added before or after disrupting the compartments and pooling the polynucleotides). (a) specifically binds to or reacts specifically with the product but not the substrate when both are attached to a polynucleotide; or (b) Both substrate and product are optionally bound, where only the product, but not the substrate, binds or reacts with the polynucleotide.

[0243] The pooled polynucleotides encoding catalytic molecules can then be enriched by selecting for polynucleotides with modified optical properties.

[0244] Another option is to bind the nucleic acid to a product-specific antibody (or other product-specific molecule). In this format, the substrate (or one of the substrates) is present in each compartment, not linked to the polynucleotide, but bearing a molecular "tag" (e.g., biotin, DIG or DNP, or a fluorescent group). When the catalyst to be selected converts the substrate to product, the product retains the "tag" and is then captured in a microcapsule by a product-specific antibody. In this way, the polynucleotide becomes associated with the "tag" only if it encodes or produces an enzyme capable of converting the substrate to product. Once all reactions have stopped, the polynucleotide encoding the active enzyme becomes "tagged" and may already have altered optical properties, for example, if the "tag" is a fluorescent group. Alternatively, a change in the optical properties of the "tagged" gene can be induced by adding a fluorescently labeled ligand that binds to the "tag" (e.g., fluorescently labeled avidin / streptavidin, a fluorescent anti-"tag" antibody, or a non-fluorescent anti-"tag" antibody that can be detected by a second fluorescently labeled antibody).

[0245] Alternatively, selection may be performed indirectly by coupling a first reaction to a subsequent reaction occurring in the same compartment. There are two general ways this can be done. In a first embodiment, the product of the first reaction reacts or binds to a molecule that does not react with the substrate of the first reaction. The second, coupled reaction proceeds only in the presence of the product of the first reaction. The properties of the product of the second reaction can then be used to identify a polynucleotide encoding a polypeptide with a desired activity and induce a change in the optical properties of the polynucleotide as described above.

[0246] Alternatively, the product of the selected reaction may be a substrate or cofactor for a second enzyme-catalyzed reaction. The enzyme catalyzing the second reaction may be translated in situ within the microcapsules or incorporated into the reaction mixture prior to compartmentalization. Only if the first reaction proceeds will the bound enzyme produce a product that can be used to induce a change in the optical properties of the polynucleotide, as described above.

[0247] This coupling concept can be adapted to incorporate multiple enzymes, each of which uses the product of the previous reaction as a substrate. This allows for the selection of enzymes that do not react with the immobilized substrate. It can also be designed to enhance sensitivity through signal amplification, where the product of one reaction serves as a catalyst or cofactor for a second reaction or series of reactions leading to a selectable product (see, e.g., Johannsson and Bates, 1988; Johannsson, 1991). Furthermore, enzyme cascade systems can be based on the generation of an enzyme activator or the destruction of an enzyme inhibitor (see Mize et al., 1989). Coupling also has the advantage of allowing a common selection system to be used across a group of enzymes that produce the same product, enabling the selection of complex chemical transformations that cannot be performed in a single step.

[0248] This method of coupling thus allows novel "metabolic pathways" to be evolved stepwise in vitro, selecting and improving first one step and then the next. Because the selection strategy is based on the final product of the pathway, all previous steps can be evolved independently or sequentially, without the need to set up a new selection system for each step of the reaction.

[0249] Alternatively stated, there is provided a method for isolating one or more polynucleotides encoding a polypeptide having a desired catalytic activity, the method comprising: (1) expressing the polynucleotides to obtain the respective polypeptides; (2) allowing the polypeptide to catalyze the conversion of the substrate to a product that may or may not be directly selectable, depending on the desired activity; (3) optionally, coupling the initial reaction to one or more subsequent reactions, each conditioned by the product of the previous reaction, leading to the creation of a final selectable product; (4) ligating the selectable product of the catalysis to a polynucleotide; a) attaching a substrate to a polynucleotide in such a way that the product is associated with the polynucleotide; or b) reacting or binding the selectable product with a polynucleotide by means of an appropriate molecular "tag" attached to the substrate that remains on the product; or c) binding the selectable product (but not the substrate) to the polynucleotide by a product-specific reaction or interaction with the product; (5) selecting the product of the catalytic reaction together with its bound polynucleotide by its characteristic optical properties or by adding a reagent that specifically binds to or reacts specifically with the product, thereby inducing a change in the optical properties of the polynucleotide, wherein in steps (1) through (4), each polynucleotide and each polypeptide is contained within a microcapsule.

[0250] All catalytic strategies described herein can also be implemented with the roles of polypeptide and substrate / inhibitor / product / reagent reversed. For example, library polypeptides of the invention can be screened for substrate activity or regulatory activity by providing an enzyme as the target. Selection by regulatory activity is also described below.

[0251] 4.7.3 Substrate specificity / selectivity

[0252] Polynucleotides encoding enzymes with substrate specificity or selectivity can be specifically enriched by positively selecting for reactions with one substrate and negatively selecting for reactions with another substrate. This combination of positive and negative selection pressures could be crucial for separating regioselective and stereoselective enzymes (e.g., enzymes that can distinguish between two enantiomers of the same substrate). For example, two substrates (e.g., two different enantiomers) could be labeled with different tags (e.g., two different fluorophores), which then become attached to the polynucleotide through an enzyme-catalyzed reaction. If the two tags confer different optical properties to the polynucleotide, the substrate specificity of the enzyme can be determined by the optical properties of the polynucleotide, allowing polynucleotides encoding polypeptides with the wrong specificity (or no specificity) to be rejected. Tags that do not alter optical activity can also be used if tag-specific ligands with different optical properties are added (e.g., tag-specific antibodies labeled with different fluorophores).

[0253] 4.7.4 Adjustment

[0254] Similar systems can be used to select for regulatory properties of cyclic polypeptides.

[0255] When selecting regulator molecules that act as activators or inhibitors of a biochemical process, components of the biochemical process can be translated in situ in each compartment or incorporated into the reaction mixture prior to compartmentalization, except that components that are permeable to the compartment material may be contacted with the compartment after compartmentalization. If the selected polynucleotide encodes an activator, selection can be for the product of the regulated reaction, as described above in connection with catalysis. If an inhibitor is desired, selection can be for a chemical property specific to the substrate of the regulated reaction, or for the lack of a chemical property specific to the product of the regulated reaction.

[0256] Accordingly, there is provided a method for sorting one or more polynucleotides encoding a polypeptide that exhibits a desired modulating activity, the method comprising: (1) expressing the polynucleotides to obtain the respective polypeptides; (2) causing the polypeptide to activate or inhibit a biochemical reaction or series of coupled reactions depending on the desired activity, in a manner that allows for the production or survival of a selectable molecule; (3) ligating the selectable molecule to the polynucleotide; a) attaching a selectable molecule, or the substrate from which it is derived, to a polynucleotide; or b) reacting or binding the selectable product to a polynucleotide by means of an appropriate molecular "tag" attached to the substrate that remains on the product; or c) binding of the product of catalysis (but not the substrate) to the polynucleotide by a product-specific reaction or interaction with the product; or d) immobilizing the components of the reaction and any products within a gel (in this mode, a gel-forming agent must be included in the encapsulated reaction components at the start of the process, and the gel can be formed, for example, by cooling the reaction to the phase change temperature of the gel); or e) maintaining the reaction components and any products in microcapsules; (4) selecting the selectable product together with its bound polynucleotide by its characteristic optical properties or by adding a reagent that specifically binds to or reacts specifically with the product, thereby inducing a change in the optical properties of the polynucleotide, wherein in steps (1) through (3), each polynucleotide and each polypeptide is contained within a microcapsule.

[0257] In modes involving option (e) of step (3) and the addition of reagents in step (4), a semipermeable compartment material should be selected that is permeable to any such reagents. In embodiments, the compartments are not disrupted prior to sorting.

[0258] Generally, when selecting a regulator or modulator of a biochemical activity, such as an activator or inhibitor of an enzyme, a candidate regulator (in this case, a cyclic polypeptide) is contacted with the target of modulation. The target and candidate are then contacted with all other reaction conditions and components normally required for activity to occur (i.e., in the absence of any regulators that may interfere with the assay). The mixture is incubated for a period of time, after which the activity or absence of the target is assessed.

[0259] A reporter may be included in the reaction to facilitate detection of any activity. A detectable reporter is a molecule, atom, ion, or group that can be detected by standard detection methods. For example, the reporter may be capable of producing a detectable signal in response to a stimulus, such as contact with a chromogenic substrate or light of an appropriate excitation wavelength.

[0260] The presence or amount of a detectable reporter can be determined by detecting or measuring a signal generated by the reporter. Suitable detectable labels can include fluorescent reporters, chromogenic reporters, Raman-active reporters such as mercaptopyridine, thiophenol (TP), mercaptobenzoic acid (MBA), and dithiobissuccinimidyl nitrobenzoate (DNBA), mass spectrometry reporters, and particles identifiable by shape by image analysis.

[0261] Suitable fluorescent reporters include fluorescein and fluorescein derivatives such as O-methyl-fluorescein or fluorescein isothiocyanate (FITC), phycoerythrin, europium, TruRed, allophycocyanin (APC), PerCP, Lissamine, rhodamine, B X-rhodamine, TRITC, BODIPY-FL, FluorX, Red 613, R-phycoerythrin (PE), NBD, Lucifer Yellow, Cascade Blue, methoxycoumarin, aminocoumarin, Texas Red, hydroxycoumarin, Alexa Fluor dyes (Molecular Probes), sulfonate cyanine dyes such as Cy2, Cy3, Cy3.5, Cy5, Cy5.5, and Cy7 (AP Biotech), IRD41 IRD700 (Li-Cor, Inc.), NIR-1 (Dejindom, Japan), La Jolla Blue (Diatron), DyLight TM 405, 488, 549, 633, 649, 680, and 800 reactive dyes (Pierce / Thermo Fisher Scientific Inc.), or IRDye TM LI-COR etc. TM Dye (LI-COR TM Biosciences).

[0262] The reporter can be a substrate, product, or intermediate of a biochemical process. The reporter can be a direct substrate or product of the target of regulation. In other embodiments, the reporter may not be a direct substrate or product of the target of regulation. For example, the reporter can act on or be produced at an early step in a biochemical activity. Or, the reporter can act on or be produced at a late step in a biochemical activity. In a simple enzyme cascade, accumulation of the reporter (relative to a control lacking the cyclic polypeptide) indicates activation at any step preceding reporter production (an "upstream" step) and / or inhibition of any step following reporter production (a "downstream" step). Conversely, absence of the reporter indicates inhibition of any step preceding reporter production and / or activation of any step following reporter production. Alternatively, the reporter is a substrate or product of a separate biochemical process coupled to the biochemical process involving the target of regulation.

[0263] Suitable reporter substrates that can be converted into detectable reporters by the action of specific enzymes are well known in the art. For example, the reporter substrate fluorescein disulfate can be converted to the detectable reporter fluorescein by arylsulfatase. Similarly, fluorescein phosphate and fluorescein acetate can be used with phosphatases and acetylases, respectively.

[0264] Many reporter substrates are commercially available (eg, Molecular Probes Inc) or may be synthesized by standard procedures.

[0265] Preferably, the detectable reporter is carried within a bead.

[0266] In other embodiments, the reporter may be a ligand that binds to a component of a biochemical process. The candidate regulator may thereby modulate the binding activity of the reporter. In this case, a suitable reporter undergoes a detectable physicochemical change upon binding. The reporter may bind directly to the target of regulation. In other embodiments, the reporter may not directly bind to the target of regulation, but may instead bind to another component upstream or downstream of the target of regulation in biochemical activity. In either case, however, reporter binding, whether direct or indirect, must depend on the activity of the target of regulation. If the candidate regulator inhibits reporter binding, the signal associated with the unbound reporter will dominate (i.e., be increased relative to a positive control). Conversely, if the candidate regulator promotes binding, the signal associated with the bound reporter will dominate (i.e., be increased relative to a negative control).

[0267] Other formats will be readily contemplated by one of skill in the art. For example, the target of regulation may be an enzyme that catalyzes a reaction that produces an inhibitor that prevents another component of the biochemical activity from binding to the reporter. In either case, the relationship between the presence of the regulatory component and its effect on the reporter will be readily ascertained by one of skill in the art.

[0268] In embodiments, components of the reporter assay are incorporated into the compartment simultaneously with the IVTT system, and some components of the reporter assay may be transcribed from an additional polynucleotide by the IVTT system within the capsule concomitantly with expression of the cyclic polypeptide.

[0269] In some embodiments, compartmentalized / decompartmentalized gel beads containing co-immobilized polynucleotides and cyclic peptides may be contacted with media containing the components of a reporter assay.

[0270] In some embodiments, some components of the reporter assay and / or polynucleotides encoding components of the reporter assay are incorporated into the compartments simultaneously with the IVTT system, in which the remaining components of the reporter assay are contacted with the compartments or decompartmentalized gel beads at a later stage.

[0271] In embodiments in which a reporter assay component is contacted with a compartment, the compartment material should be selected so that the reporter assay component, which is initially present only in the external environment of the compartment, can penetrate across the compartment boundary into the interior volume of the compartment.

[0272] 4.7.5 Protein-Protein Interactions

[0273] Protein-protein interactions (PPIs) can be monitored by assays (e.g., FRET or HTRF assays) that are encapsulated with the library or later merged using droplet merging, or incorporated by reconstituting decompartmentalized gel beads in assay medium.

[0274] To conduct the assay, a FRET donor moiety can be attached to one of a pair of interacting proteins, and a FRET acceptor moiety can be attached to the other of the interacting pair. Positioning the donor / acceptor moieties to ensure sufficient proximity for resonance energy transfer when the two proteins interact is within the skill of one of ordinary skill in the art. The FRET-functionalized protein is then contacted with a polypeptide from the library.

[0275] The presence of a polypeptide according to the invention that inhibits the interaction between the two proteins will not result in FRET, whereas the presence of an inactive molecule will display a FRET signal. These two species can be easily separated by FACS.

[0276] The assay may be performed in the opposite direction to identify polypeptides that form or enhance an interaction between two proteins, where the presence or increase in FRET indicates a polypeptide that promotes the interaction.

[0277] 4.7.6 Optical properties of polypeptides

[0278] When the polypeptide binds back to the polynucleotide within the compartment, for example, via a common element of the polypeptide that binds to a ligand that is part of the polynucleotide, it becomes possible to select for the unique optical properties of the polypeptide. The polynucleotides can then be sorted using the optical properties of the bound polypeptide. This embodiment can be used, for example, to select variants of green fluorescent protein (GFP) with improved fluorescence and / or novel absorption and emission spectra (Cormack et al., 1996; Delagrave et al., 1995; Ehrig et al., 1995).

[0279] 4.7.7 Polynucleotide / Polypeptide Flow Sorting

[0280] In a preferred embodiment of the invention, beads, capsules, polynucleotides, and / or polypeptides are sorted by flow cytometry. Sorting can be performed using a variety of optical properties, including light scattering (Kerker, 1983) and fluorescence polarization (Rolland et al., 1985). In a highly preferred embodiment, the difference in the optical property of the beads, capsules, polynucleotides, and / or polypeptides is a difference in fluorescence, and the beads, capsules, polynucleotides, and / or polypeptides are sorted using a fluorescence-activated cell sorter (Norman, 1980; Mackenzie and Pinder, 1986) or similar device. In a particularly preferred embodiment, the beads, capsules, polynucleotides, and / or polypeptides comprise non-fluorescent, non-magnetic (e.g., polystyrene) or paramagnetic microbeads (see Fornusek and Vetvicka, 1986), optimally 0.6-1.0 μm in diameter, to which are attached both genes and groups responsible for generating a fluorescent signal: (1) Commercially available fluorescence-activated cell sorting equipment from reputable manufacturers (e.g., Becton-Dickinson, Coulter) can sort up to 10 cells per hour. 8 individual beads, capsules, polynucleotides, and / or polypeptides (events) can be sorted; (2) The fluorescent signal from each bead corresponds closely to the number of fluorescent molecules attached to the bead, and currently only a few hundred fluorescent molecules per particle can be quantitatively detected. (3) The wide dynamic range of the fluorescence detector (typically 4 log units) allows for easy setting of the stringency of the sorting procedure, thereby allowing for the recovery of an optimal number of beads, capsules, polynucleotides, and / or polypeptides from the starting pool (gates can be set to separate beads with small differences in fluorescence or only those with large differences in fluorescence, depending on the selection being performed); (4) Commercially available fluorescence-activated cell sorting equipment is capable of simultaneous excitation at up to two different wavelengths and detection of fluorescence at up to four different wavelengths (Shapiro, 1983), and positive and negative selections can be performed simultaneously by monitoring the labeling of beads, capsules, polynucleotides, and / or polypeptides with two (or more) different fluorescent markers. For example, if two alternative substrates (e.g., two different enantiomers) of an enzyme are labeled with different fluorescent tags, the beads, capsules, polynucleotides, and polypeptides can be labeled with different fluorophores depending on the substrate used, and only genes encoding enzymes with enantioselectivity are selected. (5) Highly uniform derivatized / underivatized nonmagnetic / paramagnetic microparticles (beads) are commercially available from many sources (Sigma, Molecular Probes, etc.) (Fornusek and Vetvicka, 1986).

[0281] 4.8 Polynucleotide Isolation

[0282] The polynucleotides from one or more beads may be isolated, amplified, cloned, sequenced, or otherwise manipulated.

[0283] The methods described herein can further include identifying and / or isolating polynucleotides from one or more compartments or beads identified as containing the candidate cyclic polypeptide.

[0284] Polynucleotides from one or more beads may be isolated, amplified, sequenced, cloned, and / or otherwise interrogated. For example, the polynucleotides may be extracted using conventional techniques such as gel extraction columns or agarose, and / or sequentially amplified by isothermal amplification (e.g., multi-prime RCA), or PCR, or both.

[0285] The methods described herein can also be adapted for systems in which the compartments are, for example, vesicles or alginate microcapsules. These systems can also be engineered to provide microreactors for PCR, agarose gelation, IVTT, and subsequent assays and high-throughput sorting. [Example]

[0286] 5. Working Example

[0287] 5.1 Introduction to the Examples

[0288] In one embodiment, the present invention overcomes limitations of the art by implanting SICLOPPS into microfluidic droplets. Because it is possible to express individual SICLOPPS library components in each droplet, achieved by in vitro transcription / translation, the DNA code for each cyclic peptide is also contained in the same droplet, greatly simplifying hit identification by PCR amplification and DNA sequencing. This library can be coupled with any pharmaceutical assay, including fluorescent or colorimetric outputs (e.g., FRET, HTRF). This method further enables the generation of libraries of hundreds of millions of cyclic peptides, which can include unnatural amino acids.

[0289] To overcome the limitations of droplet merging, we utilize an agarose-in-oil droplet-based method for highly parallel and highly efficient single-molecule DNA amplification prior to in vitro protein expression. This method utilizes the thermoresponsive sol-gel switching properties of agarose to capture single DNA molecules prior to amplification with Phi29 DNA polymerase. Following DNA amplification, the agarose droplets are gelled (or solidified) to form agarose beads, capturing all amplicons within each microreactor and preserving the monoclonality of each droplet.

[0290] The use of monodisperse agarose-in-oil microdroplets has previously been used to monitor biochemical reactions within gelled microbeads. As previously described, the transition from agarose droplets to solidified agarose particles enhances stability and facilitates manipulation. Following solidification, we describe the isolation of isothermally amplified DNA solely within the agarose beads (i.e., pre-amplified DNA does not diffuse outward). Therefore, the resulting particles can be recovered by demulsifying the emulsion, washed, and resuspended in an IVTT-compatible buffer, effectively circumventing the limitations associated with droplet merging. After resuspension in an IVTT-compatible buffer, other components of the IVTT system may be added, and in embodiments, the IVTT-competent agarose beads may be re-emulsified to form an IVTT microreactor.

[0291] The developed methodology allows for high-throughput generation of uniform droplets with efficient single-molecule DNA amplification, providing a promising platform for many single-copy gene-based studies.

[0292] Overall, we demonstrate a novel, ultra-high-throughput screening platform that can be used for the rapid selection of protein-protein interaction inhibitors via in vitro compartmentalization of SICLOPPS-derived cyclic peptide libraries in femtoliter-sized aqueous compartments, and use FACS to selectively recover droplets displaying desired phenotypic traits.

[0293] We describe successful single-molecule DNA amplification in monodisperse agarose droplets, following solidification and subsequent demulsification (via the addition of PFO), which allows the isolation of DNA-enriched agarose particles. The amplified DNA is detected by staining with a fluorescent dsDNA-intercalating dye, allowing simultaneous detection by flow cytometry. A single, distinct population of agarose particles is observed in side-to-forward scatter plots, displaying a significant increase in green fluorescence intensity compared to a control without Phi29 DNA polymerase.

[0294] Prior to introducing the agarose particles into the IVTT-containing droplets, the agarose beads are washed to remove unbound DNA along with the removal of incompatible DNA amplification buffer; if washing is not performed, IVTT is inhibited and no GFP expression is observed.

[0295] For agarose-in-oil IVTT (double emulsion) formation, the agarose particles, along with the IVTT components, are reintroduced into the hydrophobic microfluidic device. After incubation at 37 °C, the emulsion is placed on ice to stop protein expression. For triple emulsion formation (agarose-in-oil-in-water IVTT), droplets of the double emulsion are injected into a third, final hydrophilic microfluidic device and re-emulsified into a configuration compatible with flow cytometry.

[0296] Finally, triple emulsions containing GFP expressed in vitro from a GFP-encoding plasmid can be sorted using FACS: droplets are injected and sorted at a rate of 1,000 events per second, or over 60,000 events per minute, or 3.2 million events per hour.

[0297] Overall, we demonstrate the successful implementation of agarose particles in IVTT for in vitro expression of polypeptides in microfluidically generated droplets.

[0298] Example 1: Generation of monodisperse single water-in-oil droplets at the femtoliter scale

[0299] The generation of aqueous femtodroplets using a 10 × 5 μm (width × depth) flow-focusing microfluidic device on a chip, capable of achieving droplet volumes of 5–50 fL, is shown in Figure 1, whereby two immiscible liquids enter the device through four parallel microchannels, with a continuous HFE-7500-based phase entering through the outer two channels and a dispersed aqueous phase entering through the inner two channels. Downstream, the aqueous microchannels merge before entering the orifice. The fixed, elongated water streams initially observed at the nozzle orifice during droplet formation, as reported in Shim et al., represent the "tip streaming" mechanism of droplet generation in the drop break-off regime, i.e., the formation of a conical droplet shape with a very sharp tip from which smaller droplets (approximately 0.5 μm in diameter) are released. First reported in 1934, surfactant-mediated microscale tip streaming represents a hydrodynamic phenomenon capable of generating submicron-sized droplets due to surfactant concentration gradients at interfaces resulting from extensional flows generated within flow-focusing geometries.

[0300] To demonstrate well-controlled generation of monodisperse water-in-oil (w / o) femtodroplets for SICLOPPS library encapsulation, droplets containing 100 μM fluorescein (FITC dye) in 1× TAE (pH ∼7.5, since FITC fluorescence emission is strongly dependent on the ambient pH) were formed in hydrofluoroether HFE-7500, premixed with 5% Krytox 157 FSL Jeffamine ED-600 di-salt surfactant (JUS, not commercially available) to reduce interfacial tension at the oil-water interface. To probe the frequency of femtodroplet formation and thus determine the theoretical number of encapsulated SICLOPPS library variants achievable over time (assuming single-molecule encapsulation), the effect of volumetric flow rate (defined by the oil / water flow rate ratio, R1 / R2) on droplet size was investigated by increasing the oil-phase flow rate in stages (from 10 to 60 μl / h) while maintaining a constant water flow of 10 μl / h. The frequency of droplet formation is expected to increase with decreasing droplet size and has been theoretically calculated as: f = Q aq / D v , where f is the generation frequency (Hz), Q aq is the aqueous phase flow rate (μl / s), V d is the final drop volume (μl).

[0301] To determine the diameter of microfluidically generated emulsion droplets, samples were pipetted onto disposable Fast-Read Counting Slides and imaged using an Olympus CKX41 inverted microscope or a Zeiss Axio II. A QIClick CCD camera was attached for image capture and controlled using open-source microscope software (Micro-Manager). For analysis, bright-field and fluorescent images were saved as 16-bit files and processed using ImageJ software. Statistics of droplet diameter and volume were generated using the "Particle Analysis" function. Accordingly, circularity measurements above a user-defined threshold were performed for all objects greater than the specified minimum cutoff (0.8) and tabulated for data processing. The resulting droplet size distributions under each flow rate condition are shown in Figure 2.

[0302] Increasing the oil / surfactant flow rate subsequently decreases the overall droplet size. Similarly, as shown in Figure 3, the decrease in droplet volume coincides with an increase in their generation frequency.

[0303] Example 2: Generation of water-in-oil-in-water double emulsion droplets for flow cytometry analysis

[0304] For flow cytometry analysis of SICLOPPS-encapsulated emulsion samples, the compatibility of the water-in-oil-in-water double emulsion configuration generated by two-step re-emulsification was essential. Therefore, primary FITC-containing water-in-oil emulsions generated using a 10 × 5 μm microfluidic device were converted into double emulsion droplets using a second hydrophilic flow-focusing chip with microchannel dimensions of 15 × 16 μm. To enable re-emulsification, the primary emulsions were stored upright in glass syringes and flaked with a fluorinated mineral oil solution. During encapsulation, a second FC-40-filled syringe was added to facilitate controlled manipulation and encapsulation of single primary droplets per double emulsion. Simultaneously, a continuous phase containing 1% Tris / Tween 80 surfactant was used to stabilize the interface. The on-chip two-step emulsification process is shown in Figure 4.

[0305] Additionally, the workflow for single and double emulsion droplet generation using flow-focusing microfluidic geometry is shown in Figure 5.

[0306] The resulting droplet inner diameter averaged about 6.69 μm, while the double emulsion outer diameter was about 14.98 μm. Therefore, these 1.76 pL of double emulsion resulted in a droplet size of about 5.68 × 10 8 Droplets / mL were obtained. The majority (98.21% on average of 10 images) were observed to contain a single aqueous core, while a small proportion (1.79%) was found to be doubly occupied (arrows in Figure 6). Droplet stability and structural integrity were maintained after 1 month of storage at 4 °C.

[0307] The double emulsion droplets in Figure 6 were then analyzed using a BD Accuri C6 flow cytometer. Density plots of the on-chip generated double emulsion droplets show two distinct clusters in forward and side cytometric scatter (FSC and SSC, Figure 7). While FSC distinguishes droplets based on size (diameter), SSC is used to distinguish droplets based on internal complexity (or granularity). As a result, double emulsion droplets containing an internal aqueous core exhibit greater side scatter than empty oil-in-water droplets and therefore occupy a higher position in the SSC vs. FSC density / scatter plot. Monodisperse droplet populations exhibit similar forward scatter. Singly occupied (droplets containing an internal fluorescein core) and unoccupied droplets, i.e., those that failed to encapsulate the primary water-in-oil droplets during on-chip re-emulsification, are shown in Figure 7A. This distinction is further supported by the corresponding green fluorescence intensity profiles, as shown in Figure 6B, where the highly fluorescent droplet population, representing singly occupied FITC-containing droplets, has a fluorescence signal that is approximately three orders of magnitude greater than that of the second, lower-lying population (unoccupied droplets), representing the baseline level of fluorescence.

[0308] Background fluorescence levels were confirmed by preparing simple oil-in-water droplets using the exact same reagents and conditions used previously. A single fluorescence peak and droplet population were observed in the resulting fluorescence histogram and SSC vs. FSC plot, respectively, which correspond directly to those in Figures 7A and 7B.

[0309] Example 3: Creation of a library of 3.2 million cyclic peptides in droplets

[0310] To facilitate identification of droplets containing cyclic peptides, we designed an expression system encoding both the SICLOPPS intein and a fluorescent protein (GFP). Droplets containing the SICLOPPS plasmid can be easily identified and sorted / isolated based on their fluorescence using FACS. A pETDuet-based plasmid encoding the Npu split intein in the SICLOPPS format (IC-extein-IN) was created to construct the CX5 SICLOPPS library from the vector pARNpuHisSsrA-CX5 (Townend and Tavassoli, 2016, ACS Chemical Biology, 1624-1630). The C- and N-intein-encoding regions, along with the peptide-encoding hexamer, were PCR amplified using forward and reverse primers encoding EcoRI and HindIII restriction enzymes, respectively. The reverse primer was designed to omit the SsrA degradation tag. This amplified product was then digested and subsequently cloned into the first multiple cloning site (MCS1) of the vector pETDuet-1. GFP was PCR-amplified from the coding gene block using forward and reverse primers encoding NdeI and EcoRV restriction enzyme sites, respectively, and this construct was similarly cloned into MCS2 of vector pETDuet-1 to generate a vector encoding both SICLOPPS and GFP proteins. The resulting plasmid was subsequently used as a template in the construction of the SICLOPPS cyclic peptide library.

[0311] Contains 5 variable amino acids, thus 3.2 × 10 6A hexamer library containing the library components (X = any amino acid residue) was constructed as previously described (Tavassoli and Benkovic, 2007, Nature Protocols, pp. 1126-1133). The library was designed to encode cyclic peptide hexamers within the first multiple cloning site (MCS) of pETDuet-1, with the first cysteine ​​residue followed by five randomized residues, because the nucleophile at the first position is required for the transesterification step of intein processing. (Other residues suitable for providing the nucleophile at the first position include serine (S) and threonine (T).) The second MCS was constructed to express GFP, thus allowing simultaneous expression of both the cyclic peptide and the fluorescent protein, enabling rapid and systematic identification of cyclic peptide-containing droplets.

[0312] The resulting plasmid is shown in FIG.

[0313] As previously detailed (Tavassoli and Benkovic, 2007, Nature Protocols, 1126-1133), the random oligonucleotides in the library were C and the 3' end of I N A two-step PCR-based approach was used, in which a forward primer was incorporated between the region that binds to the 5' end of the target peptide. The variable segments were coded in the format NNS, where N represents any of the four DNA bases (A, C, G, or T) and S represents C or G. The NNS sequence generated 32 codons, encoding all 20 natural amino acids, although the ochre (UAA) and opal (UGA) stop codons were excluded from the library. Note that there is no limit to the number of amino acids in the target peptide, allowing for the generation of cyclic peptides of various sizes.

[0314] The generation of the first linear PCR product resulted in mismatches in the random nucleotide regions due to the sequence complexity of the library. Therefore, a second PCR using a "zipper" primer corresponding to the 3' end of the C-intein was used to ensure that all DNA sequences annealed to complementary strands. The resulting DNA library was then incorporated into the SICLOPPS plasmid using standard molecular biology techniques, generating the desired CX5 library in the vector pDuetNpuHis-GFP.

[0315] After ligation, the SICLOPPS CX5 peptide library was transformed into electrocompetent DH5α E. coli cells, plated onto LB agar plates, and grown overnight at 37°C. After growth, the resulting colonies were harvested by colony scraping and miniprepped to yield a plasmid library ready for encapsulation in agarose-in-oil droplets and Phi29-mediated preamplification as described above.

[0316] Example 4: Miniaturization of biochemical manipulations using agarose-based droplet microfluidics

[0317] The above-described plasmid library (generated in Example 3) was used in the following experiments. The application of agarose-in-oil emulsion droplets for amplicon capture and DNA amplification in gelled agarose beads has been realized as a powerful approach to address the challenges of traditional PCR and primer-functionalized microbead methods. Here, we describe a highly parallel Phi29 DNA polymerase-mediated protocol for amplifying single DNA plasmid molecules in femtoliter-sized agarose beads with a diameter of 6 to 7 μm, leveraging the unique thermoresponsive sol-gel switching properties of agarose. We utilize a two-aqueous inlet emulsion droplet generator with co-encapsulated isothermal amplification reagents, along with agarose as the capture matrix. The monoclonality of each product is consequently preserved within the robust and inert biochemical environment of each reservoir / bead. The solidified beads encapsulating the captured amplicons are then co-encapsulated in an in vitro transcription / translation (IVTT) system without droplet merging to study protein expression in a defined reaction volume, as shown in Figure 8. Consequently, we demonstrate the confinement of gene transcription and translation to membrane-free agarose particles, thereby demonstrating the potential of this emerging technology in enabling novel quantitative studies of complex biological systems.

[0318] Figure 9 shows the agarose droplet microfluidic setup used for encapsulation and amplification of single DNA plasmid molecules using Phi29 DNA polymerase within a microfabricated hydrophobic device for the controlled generation of highly uniform, monodisperse femtoliter emulsion droplets containing 1% agarose solution. DNA template molecules are introduced into the droplets along with the agarose solution at a statistically diluted concentration so that the average number of molecules in a single droplet is approximately 1 or less. Ultra-low gelling temperature agarose, which remains fluid at 37 °C and has a transition gel point between 8 and 17 °C, is used for encapsulation, allowing for the facile generation of agarose droplets during device operation. Following off-chip incubation and DNA amplification, cooling the solution below its critical gel point allows the agarose matrix to switch to a solid gel phase. Once solidified, the beads remain in a solid state unless the temperature is increased. As a result, the amplified DNA remains entrapped within the solidified agarose matrix, preserving the monoclonality of the droplets after removal of the external oil phase.

[0319] For droplet generation, the resulting plasmid library was statistically diluted according to Poisson statistics to ensure that either (i) a mean (λ) of 0.1, representing single-molecule encapsulation, or (ii) an average of 100 library copies (as proof-of-principle) were encapsulated in 110 fL agarose-in-oil droplets. The resulting plasmid library, along with Phi29 DNA polymerase-mediated isothermal amplification reagents, was injected into a JUS device and encapsulated using aqueous flow rates of 10 μl / h for each aqueous phase and 30 μl / h for the oil / surfactant mixture QX200 (droplet generation oil). To generate uniform agarose emulsion droplets, agarose was loaded into the microfluidic device together with 1x TAE buffer, with QX200 oil / surfactant as the continuous phase. To maintain the agarose in a liquid state for droplet generation, the syringe containing the agarose was preheated with a commercially available microwave heating pack before agarose loading and during device operation (Figure 9A). Alternatively, a commercially available 5V DC-powered 5 x 10 cm heating pad (Figure 9B) was purchased and manually connected to a standard 5V USB cable to create a continuous-use heating element capable of generating a temperature of approximately 40 °C during 10 minutes of operation.

[0320] After collection (Figure 10A), the agarose droplets containing DNA with Phi29 DNA polymerase amplification components were incubated overnight at 30°C. The desired agarose bead population was isolated by solidification on ice and subsequent bead extraction using PFO. As a result, the beads were stained with fluorescent double-stranded DNA that binds to the DNA, allowing visualization of the internal DNA (Figure 10C).

[0321] Flow cytometry analysis of stained agarose beads containing an average of 100, 10, 1, or 0.1 starting DNA copies (λ value) (0.1 represents the encapsulation of a single molecule) prior to amplification (Figure 11) is shown in Figure 13. A 16-hour overnight incubation (corresponding to approximately 12,000 final DNA copies per drop) is preferable to a 4-hour incubation in generating larger amounts of DNA. In all cases, clear Phi29 DNA polymerase-mediated DNA amplification is observed in each experimental condition compared to negative controls containing DNA only or empty (no DNA) agarose beads.

[0322] Because the buffer system used for DNA amplification is known to be incompatible with droplet-based IVTT, agarose beads containing monoclonal amplified plasmid DNA are solidified, then washed by centrifugation and resuspended in a solution compatible with IVTT. To further investigate the compatibility of our plasmid-containing agarose beads with IVTT using the PURExpress system, we resuspended the washed agarose beads in IVTT solution with QX200 oil / surfactant and vortexed to quickly generate bulk polydisperse droplets, the brightfield and fluorescence images of which are shown in Figure 12. Interestingly, GFP-mediated fluorescence is observed in agarose beads containing in vitro protein expression components from samples containing Phi29-amplified DNA, whereas no fluorescence is observed in the absence of amplification. Therefore, the addition of a pre-amplification step following single-copy DNA encapsulation is fundamental to successful in vitro-mediated protein expression of protein constructs. This demonstrates the generation of a SICLOPPS library and GFP in the droplets described above using our dual expression vector.

[0323] To demonstrate in vitro-mediated protein expression of GFP in microfluidically generated monodisperse droplets using the PURExpress system, agarose droplets containing approximately 100 starting plasmid DNA copies were preamplified by Phi29 DNA amplification (Figure 13A) and washed as previously described. The agarose beads were then loaded into a glass syringe and injected into a hydrophobic 15 × 16 μm microfluidic device along with IVTT components and a 1% Tris / Tween 80 continuous phase. After collection, the droplets were incubated at 37 °C for 2 h and then emulsified through a 15 × 25 μm hydrophilic device to form triple emulsion droplets with an external aqueous phase for flow cytometry analysis. As a result, in vitro GFP expression was achieved in highly monodisperse microfluidically generated droplets, as shown in Figure 13B. Although the IVTT PURExpress system itself was observed to exhibit high levels of background fluorescence, clear GFP-mediated fluorescence was observed at approximately 10-fold above background levels, allowing for easy selection of droplets containing cyclic peptides using, for example, fluorescence-activated cell sorting (FACS) to recover desired phenotypic traits. Thus, a robust agarose and IVTT platform for in vitro protein expression in microfluidic droplets is provided without complex droplet merging procedures.

[0324] Example 5: Formation and sorting of cyclic peptides

[0325] Mass spectrometry demonstrated the formation of cyclic peptides via intein splicing within the emulsion, which persisted after FACS sorting of emulsion droplets.

[0326] The vector from Example 3 was used to encode the intein-linked hexapeptide CLLFVY. This vector was used to generate circular CLLFVY as follows.

[0327] A standard PURExpress IVTT reaction was assembled in the presence of the vector CSpDuetNpuHisCLLFVY. The reaction was incubated at 37° C. for 2 hours before mass spectrometry analysis. The results are shown in the left panel of FIG. 16.

[0328] In the second experiment, the vector CSpDuetNpuHisCLLFVY was preamplified overnight in highly monodisperse, approximately 6 μm 1% agarose beads using the TempliPhi amplification system at 30°C. After heat inactivation of the enzyme, the beads were separated from the emulsion and washed by centrifugation three times at approximately 6,000 rpm in diH2O to efficiently remove the amplification buffer (which could inadvertently interfere with the subsequent IVTT procedure). The washed agarose beads containing the monoclonal amplified DNA were then reencapsulated with IVTT to express CLLFVY and incubated at 37°C for 2 hours before a third emulsification step to generate double-emulsion FACS-compatible droplets.

[0329] A double emulsion population (MCS2 clone GFP) showing increased GFP fluorescence above background was selected for sorting. Sorted samples were separated from the emulsion and subjected to mass spectrometry analysis. See Figure 16, right panel.

[0330] In both cases, the peak representing cycloCLLFVY is evident.

[0331] Example 6: In vitro expression of AB42-GFP fusions with cycloTAFDR in microfluidic droplets

[0332] The cyclic peptide TAFDR is a cyclic peptide that has been shown to inhibit the aggregation of the Alzheimer's disease protein Aβ42 fused to GFP. When GFP is produced as an aggregation fusion, its fluorescence is lost; however, when aggregation is disrupted by TAFDR, GFP can fluoresce. See Mathis et al., Nature Biomedical Engineering volume 1, pages 838-852 (2017).

[0333] The performance of our system was evaluated using an Aβ42 aggregation assay. Because this assay has been combined with the SICLOPPS screen (Mathis et al.), cycloTAFDR serves as a validated positive cyclic peptide control. Figure 17 shows that expression of this control peptide in droplets, together with an Aβ GFP fusion protein, causes a significant increase in fluorescence associated with disruption of Aβ aggregation.

[0334] TAFDR was cloned into MCS1 using the vector CSpDuetNpuHisAB42-GFP. Plasmid DNA was then encapsulated into highly monodisperse 1% agarose droplets along with isothermal DNA amplification reaction components and incubated overnight at 30°C. Agarose beads were prepared as described above, reencapsulated with IVTT, and incubated at 37°C for 2 hours. A final emulsification step was performed to generate a FACS-compatible double emulsion. A negative control vector in the presence of CA5 was similarly constructed to verify inhibition of cycloTAFDR AB42-GFP aggregation. A positive shift in green fluorescence was observed upon cycloTAFDR expression compared to buffer alone, IVTT alone, and cycloCA5 data.

[0335] Example 7: In vitro compartmentalization and FACS screening of the TX4 SICLOPPS library in double emulsion droplets.

[0336] A full screen was performed on the TX4 library using the methods of Example 6. Figure 18 shows a clear pool of potentially active compounds (area marked +ve 13.1%).

[0337] As in Example 6, a plasmid was constructed in the pETDuet-1 vector containing NpuHis TX4 in MCS1 and AB42-GFP fusion in MCS2. To enable monoclonal DNA amplification prior to IVTT, the plasmid DNA was encapsulated into highly monodisperse, approximately 6 μm agarose femtodroplets with TempliPhi isothermal DNA amplification components. Samples were incubated overnight (16 h) for maximum amplification at 30°C. Phi29 DNA polymerase was heat-inactivated at 65°C for 10 min. The beads were incubated on ice to promote gel phase conversion and capture of the isothermally amplified DNA products. The beads were separated from the emulsion, and the aqueous phase was extracted and washed three times in diH2O (6,000 rpm) to remove the amplification buffer. The beads were then encapsulated into highly monodisperse droplets with the PURExpress IVTT system and incubated at 37°C for 2 h. After protein expression, samples were converted to a double emulsion format (agarose beads in oil-in-water IVTT) to allow for FACS screening. The double emulsion population was identified and gated to allow for library sorting (B). Only fluorescence measurements above IVTT background were gated from the double emulsion population and sorted by FL1-H (GFP, A). The percentage +ve (potential positive candidate peptide) and -ve gated particles are shown for the TX4 library sample only (orange line).

[0338] ---

[0339] All publications mentioned in the above specification are incorporated herein by reference. Various modifications and variations of the described aspects and embodiments of the present invention will be apparent to those skilled in the art without departing from the scope of the present invention. Although the present invention has been described in connection with specific preferred embodiments, it should be understood that the invention as claimed should not be unduly limited to such specific embodiments. Indeed, various modifications of modes for carrying out the invention that are apparent to those skilled in the art are intended to be within the scope of the following claims. [Explanation of symbols]

[0340] [Figure 1] Aqueous: Aqueous solution Oid / surfactant: oil / surfactant Nozzel (5μm deep): Nozzle (5μm deep) Flow channel (25μm deep): Flow channel (25μm deep) Femto-droplets: Femto droplets [Figures 2A to 2D] Oid / surfactant: oil / surfactant Femtodroplets Nozzle, 5μm deep: Nozzle (depth 5μm) Flow channel, 25μm deep: Flow channel (depth 25μm) Normalized droplet count: Normalized droplet count μL / h oil: μL / h oil Droptet diameter: Droplet diameter [Figure 3] Droplet diameter: Oil flow rate: Oil flow rate Volume: Volume Droplet generation rate: Droplet generation rate [Figure 4] Emulsion filled syringe: Emulsion filled syringe Syringe needle: Syringe needle Microfluidic device: Microfluidic device [Figure 5A] Water-in-oil droplet generation: Water-in-oil droplet generation Aq: Aqueous solution Oil: Oil 44,000 droplets per second: 44,000 droplets per second Collection: Agarose bead-in-water-in-oil droplet generation: Agarose bead-in-water-in-oil droplet generation Aqueout phase: aqueous phase [Figure 5B] Primary emulsion (water-in-oil) droplets: Primary emulsion (water-in-oil) droplets Aqueous inner phase: External aqueous phase with surfactant: Hydrophilic channels: Organic middle phase: organic middle phase Inlet: Inlet Outlet: Outlet Aqueous inlet: Aqueous inlet Debris filter: Debris filter Oil inlet: oil inlet Emulsion inlet: Emulsion inlet Flow-focusing junction: Flow-focusing junction Oil: Oil Aq: Aqueous solution Emulsion: Emulsion O D =Outline I D =Outline [Figures 7A to 7D] Side scatter: Side scatter Detection threshold: Detection threshold Monodisperse droplets: Monodisperse droplets Occupied: occupied Unoccupied: Unoccupied Forward scatter: forward scatter Green fluorescence intensity: Green fluorescence intensity Monodisperse oil-in-water droplets: Monodisperse oil-in-water droplets [Figure 8A] Agarose-in-oil droplet generation: Agarose-in-oil droplet generation 2% agarose & DNA: 2% agarose and DNA Phi29 DNA polymerase: Phi2 DNA polymerase Oil: Oil 44,000 droplets per second: 44,000 droplets per second External Organic phase: External organic phase 30 ℃ incubation: Incubation at 30℃ Break emulsion: break emulsion Amplified DNA: Amplified DNA Agarose beads: Agarose beads with Phi29 DNA polymerase amplified DNA: Agarose beads with Phi2 DNA polymerase amplified DNA [Figures 8B and 8C] Agarose-in-IVTT-in-oil droplet generation: Agarose-type droplet generation in IVTT in oil Agarose beads: Oil: Oil 4,000 droplets per second: 4,000 droplets per second IVTT phase: IVTT phase Surfactant layer: surfactant layer External oil phase: external oil phase Agarose-in-IVTT-in-oil-in-water droplet generation: Agarose-type droplet generation in IVTT in oil-in-water Aq: Aqueous solution Agarose inner phase with amplified DNA: Agarose inner phase with amplified DNA Read out: Read Flow cytometry Microscopy: Microscope Oil shell: Oil shell External aqueous phase: external aqueous phase [Figure 9A] Microfluidic device: Microfluidic device Fluidic tubing: Fluid tubing Glass syringes: Glass syringes Heating pad: Heating pad Syringe pumps: Syringe pumps [Figure 9B] Phi29 + hexamers: Phi29+hexamers Surfactant + oil: Surfactant + oil 2% agarose + DNA: 2% agarose + DNA Electric heater: Electric heater [Figure 10A] Phi29 + hexamers: Phi29+hexamers 2% agarose + DNA: 2% agarose + DNA Nozzel (5μm deep): Nozzle (5μm deep) Femto-droplets: Femto droplets Flow channel (25μm deep): Flow channel (25μm deep) Agarose beads: [Figure 10B] Normalized bead count: Normalized bead count Bead diameter: [Figure 10C] Unwashed: Unwashed Fluorescence: Fluorescence Merge: fusion [Figures 11A to 11D] Count: Count Control: Control DNA alone: ​​DNA only Green fluorescence intensity: Green fluorescence intensity Single-molecule encapsulation: Single-molecule encapsulation [Figure 13A] Phi29 DNA amplification of plasmid DNA in agarose beads Side scatter: Side scatter Agarose beads: Forward scatter: forward scatter Count: Count Negative control: Negative control Green fluorescence intensity: Green fluorescence intensity [Figure 13B] IVTT of GFP from Phi29 amplified DNA in droplets Side scatter: Side scatter Singly occupied: Single occupancy Unoccupied droplets: unoccupied droplets Forward scatter: forward scatter Agarose beads: TAE alone: ​​TAE alone Green fluorescence intensity: Green fluorescence intensity [Figure 14] SICLOPPS intein: SICLOPPS intein active intein: active intein cyclic peptide lariat: lariat thioester: thioester [Figure 15] LacI coding sequence: LacI coding sequence f1 origin: f1 starting point bla (Ap) coding sequence: bla (Ap) coding sequence pBR322 origin: pBR322 origin C-intein: C-intein Library: Library N-intein: N-intein [Figure 16] Intensity: Intensity CLLFVY in IVTT bulk: CLLFVY in IVTT bulk CLLFVY IVTT post FACS sorting: CLLFVY IVTT post FACS sorting [Figure 17] Mormalized to Mode: Normalized to mode Buffer alone: ​​Buffer only IVTT alone: ​​IVTT only Green fluorescence intensity: Green fluorescence intensity [Figure 18A] Normalized to Mode: Normalized to Mode Single emulsion: Single emulsion Buffer alone: ​​Buffer only IVTT alone: ​​IVTT only TX4 library: TX4 library Green fluorescence intensity: Green fluorescence intensity [Figure 18B] Side scatter: Side scatter Double emulsions: Double emulsions Forward scatter:

Claims

1. 1. A method for co-compartmentalizing a cyclic polypeptide and a polynucleotide encoding said cyclic polypeptide, comprising: a) forming a compartment comprising a polynucleotide encoding said cyclic polypeptide; b) expressing a linear polypeptide from said polynucleotide; c) cyclizing the polypeptide; the compartments are substantially spherical with an outer diameter of 0.1 μm to 100 μm; the polynucleotide comprises a sequence encoding an N-terminal intein fragment, followed by a sequence encoding the cyclic polypeptide, followed by a sequence encoding a C-terminal intein fragment; a method wherein the compartment is: Agarose droplets in an oil-in-water in vitro transcription / translation (IVTT) system.

2. A method for sorting cyclic polypeptides, comprising the steps of claim 1, further comprising: c) screening the cyclic polypeptide for activity; d) selecting said cyclic polypeptides that exhibit a desired activity.

3. The method of claim 1 or 2, further comprising amplifying the polynucleotide within the compartment.

4. 4. The method of claim 3, wherein the compartment further comprises a gel-forming agent that solidifies into gel beads after amplifying the polynucleotides.

5. 5. The method of claim 4, wherein the compartments are destroyed after the gel-forming agent has solidified into gel beads.

6. The method of claim 5, wherein the gel beads are subjected to conditions for expression of the cyclic polypeptide.

7. 7. The method of claim 5 or 6, wherein new compartments are formed around the gel beads.

8. 8. The method of any of claims 1-7, wherein expressing the polypeptide from the polynucleotide comprises contacting the polynucleotide with an in vitro transcription and translation (IVTT) mixture comprising one or more tRNAs charged with an unnatural amino acid.

9. 9. The method of claim 4, wherein amplification is carried out in a vessel and heat is applied uniformly and continuously over the entire surface area surrounding said vessel.

10. a) a polynucleotide comprising a sequence encoding an N-terminal intein fragment, followed by a sequence encoding a cyclic polypeptide, followed by a sequence encoding a C-terminal intein fragment; b) a compartment comprising the cyclic peptide. wherein the compartments are substantially spherical with an outer diameter of 0.1 μm to 100 μm; The compartments are: Agarose-type droplets in an oil-in-water in vitro transcription / translation (IVTT) system

11. A library of cyclic polypeptides comprising a plurality of compartments according to claim 10.

Citation Information

Patent Citations

  • Cyclic peptide

    JP2003534768A

  • Method for producing peptide

    JP2006514543A

  • Intein-modified proteases, their production and industrial applications

    US20150232827A1

  • Cyclic peptides

    WO2000036093A2

  • Gel beads in microfluidic droplets

    WO2012156744A2