High-throughput methods for polypeptide synthesis
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- MASSACHUSETTS INST OF TECH
- Filing Date
- 2026-02-02
- Publication Date
- 2026-08-06
Smart Images

Figure 00000064_0000 
Figure 00000065_0000 
Figure 00000066_0000
Abstract
Description
Attorney Docket: 114203-3092HIGH-THROUGHPUT METHODS FOR POLYPEPTIDE SYNTHESISCROSS-REFERENCE TO RELATED PATENT APPLICATION
[0001] This application claims the benefit under 35 U.S.C. § 119(e) of U.S. Provisional Application No. 63 / 753,240, filed February 3, 2025, the entire contents of which are incorporated herein by reference.TECHNICAL FIELD
[0002] The subject matter disclosed herein is generally directed to high-throughput, cell-free polypeptide synthesis systems and methods of using the same for producing in vitro polypeptides.BACKGROUND
[0003] Cell-free protein expression systems have emerged as an important tool for polypeptide production, but several significant limitations hinder their broader adoption and utility. Current systems typically require multiple days for polypeptide production, creating substantial delays in research and development pipelines. These traditional methods are further constrained by their limited scalability, as they often rely on labor-intensive column-based purification steps that employ protein-specific tags and buffers. Such requirements make it particularly challenging to implement these systems in high-throughput applications or industrial-scale production.
[0004] The flexibility of existing cell-free systems is also restricted by their inability to express diverse protein types simultaneously efficiently. Existing cell-free protein expression platforms are poorly suited for high-throughput screening applications, which are increasingly critical in modern biotechnology and pharmaceutical research. The limited capacity to efficiently, reliably, and robustly perform parallel expression, purification, and quantification in discrete volumes hinders cell response profiling, variant screening, disease driver identification, and therapeutic target discovery. This limitation also impacts protein engineering applications and drug response testing, where rapid and parallel protein production is essential.
[0005] The purification steps in cell-free protein expression systems present a particular bottleneck. They require time-consuming and labor-intensive column-based methods that areAttorney Docket: 114203-3092difficult to automate or scale. These purification challenges significantly impact the overall efficiency of protein production and limit the process’s throughput. The lack of streamlined, automation-compatible purification methods has remained a persistent obstacle in advancing cell-free protein expression technology.
[0006] There is a need in the art for improved cell-free protein expression systems to address these limitations and provide rapid, scalable, and versatile protein production capabilities. Such improvements would significantly advance fields ranging from basic research to industrial-scale protein production and therapeutic development.
[0007] Citation or identification of any document in this application is not an admission that such a document is available as prior art to the present invention.SUMMARY
[0008] Disclosed herein is an expression library comprising a set of expression templates, each expression template encoding a different polypeptide and each expression template further comprising polynucleotide sequences encoding one or more of an N-terminal translation enhancer sequence; a translation initiation sequence, a first purification tag; a solubility tag; an open reading frame (ORF) encoding the polypeptide and further comprising a 3’ untranslated region (3'-UTR) sequence, an engineered poly-A tail, or both; an optional detection tag; or an optional second purification tag.
[0009] Also disclosed herein is a high-throughput in vitro polypeptide synthesis system comprising the expression library of the invention; a plurality of individual discrete volumes, each volume containing a different expression template from the expression library; a cell lysate; an RNA polymerase; a reaction buffer; a capture agent for binding the first capture tag and an optional second capture agent for binding the second capture tag; and one or more purification buffers.
[0010] Also disclosed herein are methods of highly parallel in vitro polypeptide synthesis comprising:a. allocating one or more expression templates from an expression library into individual discrete volumes, each discrete volume comprising a different transcription template, each transcription template comprising a polynucleotideAttorney Docket: 114203-3092sequence encoding a different polypeptide and further comprising an N-terminal translation enhancer sequence and further comprising one or more of;i. a translation initiation sequence;ii. a first purification tag;iii. a solubility tag;iv. an ORF encoding the polypeptide;v. a 3'-UTR sequence, an engineered poly-Atail, or both; and vi. an optional detection tag; andb. expressing the polypeptide from the expression template in the individual discrete volume in the presence of a cell lysate and a reaction buffer wherein the expressed proteins are captured in the individual discrete volume by the first capture agent;c. washing the captured polypeptide with a wash bufferd. for proteins that are not surface polypeptide, releasing the proteins into solution via a protease that cleaves a protease cleavage site adjacent to the first purification tag.
[0011] These and other aspects, objects, features, and advantages of the example embodiments will become apparent to those with ordinary skill in the art upon considering the following detailed description of example embodiments.BRIEF DESCRIPTION OF THE DRAWINGS
[0012] An understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention may be utilized, and the accompanying drawings of which:
[0013] FIG. 1 is a uniform manifold approximation and projection (UMAP) 2D visualization of cellular phenotypes.
[0014] FIGs. 2A-2F show a schematic featuring an exemplary embodiment of the systems and methods of this invention. FIG. 2A shows an arrayed plate containing DNA plasmids that encode signaling proteins. FIG. 2B shows cell-free polypeptide synthesis using a human-based in vitro transcription and translation (IVTT) system. FIG. 2C shows molecular clean-up to produceAttorney Docket: 114203-3092arrayed signaling polypeptides, featuring surface presentation via a terminal tag to capture and release soluble signaling polypeptides through a protease cleavage site. FIG.2D shows cell culture on arrayed signaling proteins. FIG. 2E shows cellular phenotypic responses induced by the signaling polypeptides. FIG. 2F shows a schematic demonstrating the iterative approaches described herein for measuring complex responses to signaling proteins where the polypeptide synthesis of the present disclosure is paired with Next-Gen sequencing, microscopy, flow cytometry and computational approaches.
[0015] FIGS. 3A and 3B are schematics showing an exemplary embodiment comprising computational methods for mapping primary human cell responses across all secreted signaling polypeptides, detecting intercellular signals in a disease phenotype, and selecting a therapeutic target. FIG. 3A shows experimental platforms that define signal-to-phenotype maps of primary human cells’ responses across all secreted signaling polypeptides. FIG. 3B shows computational methods for detecting intercellular signals in a disease phenotype and selecting a therapeutic target.
[0016] FIG. 4 is a gel showing exemplary features that may be included in the system of this disclosure that may synthesize polypeptides.
[0017] FIG.5 is a bar graph showing polypeptides. The systems and methods of this disclosure successfully synthesize polypeptides across diverse classes and with various structural and biochemical characteristics.
[0018] FIGS. 6A-6B show gene expression of interferon target genes IRF1 and OAS2 (measured via qPCR) in response to in-house synthesized interferon as compared to commercial interferon in A549 cells. FIG. 6A shows that interferon gamma produced by the present platform (marked as “TAPAS”) causes downstream effects indicative of strong biological activity. FIG.6B shows that TNFSF10 produced by the present platform (marked as “TNFSF10 TAPAS”) promotes apoptosis as expected.
[0019] FIG. 7 shows cell death induced by synthesized TRAIL protein compared to commercial TRAIL protein in MDA-MB-231 breast cancer cells.
[0020] FIG. 8 shows the phosphorylation cascade of IL-2 protein, comparing commercially available IL-2 to IL-2 synthesized using example methods disclosed herein with and without solubility tags. IL2 produced by the present platform (marked as “IL2 TAPAS”) causes downstream ERK and STAT5 phosphorylation, indicating that it is biologically active.Attorney Docket: 114203-3092
[0021] FTG. 9 shows exemplary polypeptides that may be synthesized by the systems and methods of the disclosure.
[0022] FIG. 10 provides an example of an expression template map.
[0023] The figures herein are for illustrative purposes only and are not necessarily drawn to scale.DETAILED DESCRIPTION OF THE EXAMPLE EMBODIMENTSGeneral Definitions
[0024] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. Definitions of common terms and techniques in molecular biology may be found in Molecular Cloning: A Laboratory Manual, 2ndedition (1989) (Sambrook, Fritsch, and Maniatis); Molecular Cloning: A Laboratory Manual, 4thedition (2012) (Green and Sambrook); Current Protocols in Molecular Biology (1987) (F.M. Ausubel et al. eds.); the series Methods in Enzymology (Academic Press, Inc.): PCR2: APractical Approach (1995) (M.J. MacPherson, B.D. Hames, and G.R. Taylor eds.): Antibodies, A Laboratory Manual (1988) (Harlow and Lane, eds.): Antibodies A Laboratory Manual, 2ndedition 2013 (E.A. Greenfield ed.); Animal Cell Culture (1987) (R.I. Freshney, ed.); Benjamin Lewin, Genes IX, published by Jones and Bartlet, 2008 (ISBN 0763752223); Kendrew et al. (eds.), The Encyclopedia of Molecular Biology, published by Blackwell Science Ltd., 1994 (ISBN 0632021829); Robert A. Meyers (ed.), Molecular Biology and Biotechnology: a Comprehensive Desk Reference, published by VCH Publishers, Inc., 1995 (ISBN 9780471185710); Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, N.Y. 1994), March, Advanced Organic Chemistry Reactions, Mechanisms and Structure 4th ed., John Wiley & Sons (New York, N.Y. 1992); and Marten H. Hofker and Jan van Deursen, Transgenic Mouse Methods and Protocols, 2ndedition (2011).
[0025] As used herein, the singular forms “a,” “an,” and “the” include both singular and plural referents unless the context dictates otherwise.
[0026] The term “optional” or “optionally” means that the subsequent described event, circumstance or substituent may or may not occur and that the description includes instances where the event or circumstance occurs and instances where it does not.Attorney Docket: 114203-3092
[0027] The recitation of numerical ranges by endpoints includes all numbers and fractions subsumed within the respective ranges and the recited endpoints.
[0028] The terms “about” or “approximately,” as used herein when referring to a measurable value, such as a parameter, an amount, a temporal duration, and the like, are meant to encompass variations of and from the specified value, such as variations of + / - 10% or less, + / -5% or less,+ / -1% or less, and + / -0.1% or less of and from the specified value, insofar such variations are appropriate to perform in the disclosed disclosure. It is understood that the value to which the modifier “about” or “approximately” refers is also specifically and preferably disclosed.
[0029] As used herein, a “biological sample” may contain whole cells and / or live cells and / or cell debris. The biological sample may include (or be derived from) a “bodily fluid.” The present disclosure encompasses embodiments wherein the bodily fluid is selected from amniotic fluid, aqueous humor, vitreous humor, bile, blood serum, breast milk, cerebrospinal fluid, cerumen (earwax), chyle, chyme, endolymph, perilymph, exudates, feces, female ejaculate, gastric acid, gastric juice, lymph, mucus (including nasal drainage and phlegm), pericardial fluid, peritoneal fluid, pleural fluid, pus, rheum, saliva, sebum (skin oil), semen, sputum, synovial fluid, sweat, tears, urine, vaginal secretion, vomit and mixtures of one or more thereof. Biological samples include cell cultures, bodily fluids, and cell cultures from bodily fluids. Bodily fluids may be obtained from a mammal organism, for example, by puncture or other collecting or sampling procedures.
[0030] As used herein, the term “polypeptide” refers to a sequence of linked single amino acids (monopeptides) of any length, including naturally occurring amino acids, modified amino acids, and non-natural amino acids. Modified amino acids may include, but are not limited to, post-translationally modified amino acids, chemically modified amino acids, and amino acid analogs. Non-natural amino acids include synthetic amino acids and amino acid mimetics that can be incorporated into a peptide chain. The terms 'peptide,' 'oligopeptide,' or 'protein' may be used interchangeably herein with the term 'polypeptide. The monopeptides may be linked via peptide bonds or by other peptide coupling methods. Reagents such as HATU / DIEA, EDC / HOBt, PyBOP / DIEA, or HBTU / DIEA may couple natural and non-natural amino acids containing amine and carboxyl groups. Click chemistry may also be used, including azide-alkyne cycloaddition (CuAAC) between azide-modified and alkyne-modified amino acids, strain-promoted azide-Attorney Docket: 114203-3092alkyne cycloaddition (SPAAC) for copper-free conditions, and thiol-ene click reactions. Thiol-based conjugation methods provide additional flexibility through maleimide-thiol coupling, disulfide bond formation, and thioether linkages. Selective side chain chemistry can be utilized when working with specific residues: lysine-like side chains can be conjugated using NHS esters, tyrosine-like residues can undergo diazonium coupling, and tryptophan-like residues can be modified through Mannich reactions. Native Chemical Ligation (NCL) offers another approach, using N-terminal cysteine or thiol derivatives and C-terminal thioesters, which can be performed under aqueous conditions. The amino acids may also be linked through polypeptide-polypeptide binding, such as encoded tags like SpyTag / SpyCatcher polypeptide ligation systems and their derivatives. Keeble and Horvath. Chem. Sci, 2020, 11, 7281-8291.
[0031] As used herein, unless otherwise stated, the term “transcription” refers to the synthesis of RNA from a DNA template; the term “translation” refers to the synthesis of a polypeptide from an mRNA template. Translation may be regulated by the sequence and structure of the 5’ UTR and may include further 5’ modifications such as 7-methyl guanosine (m7G) caps, additional 2’-O-methylation at the +1 position, and 5-phosphorothioate modifications. The translation may also be enhanced by incorporating nucleobase modifications such as pseudouridine and Nl-methylpseudouridine, 5, methylcytidine. The translation may also include the incorporation of chemically modified poly-A tails as well as the incorporation of untranslated regions (UTRs) from highly expressed human genes such as alpha- and beta-globins.
[0032] As used herein, “z>z vitro" refers to systems outside a cell or organism and may sometimes be referred to cell free system. In vivo systems relate to essentially intact cells whether in suspension or attached to or in contact with other cells or a solid. In vitro systems have an advantage of being more readily manipulated. For example, delivering components to a cell interior is not a concern; manipulations incompatible with continued cell function are also possible. However, in vitro systems involve disrupted cells or the use of various components to provide the desired function and thus spatial relationships of the cell are lost. When an in vitro system is prepared, components, possibly critical to the desired activity can be lost with discarded cell debris. Thus in vitro systems are more manipulatable and can function differently from in vivo systems..
[0033] The terms “z / z vitro transcription” (IVT) and “cell-free transcription” are used interchangeably herein and are intended to refer to any method for cell-free synthesis of RNA fromAttorney Docket: 114203-3092DNA without synthesis of protein from the RNA. In some embodiments, the RNA is messenger RNA (mRNA), which encodes proteins. The terms “zzz vitro transcription-translation” (IVTT), “cell-free transcription-translation,” “DNA template-driven in vitro protein synthesis,” and “DNA template-driven cell -free protein synthesis” are used interchangeably herein and are intended to refer to any method for cell-free synthesis of mRNA from DNA (transcription) and of protein from mRNA (translation). The cell-free system may be a synthetic mixture such as PURE, PURExpress, TXTL Toolbox, or a customized synthetics system. The cell-free system may include eukaryotic cell-free systems, including but not limited to, Bombyx mori, CHO, HeLa, Leishmcinia tarentolae, Nicotiana tabacum, Pichia pastoris, Rabbit reticulocyte, Saccharomyces cerevisiae, Spodoptera frugiperda, Tichophisia ni, and Wheat germ based cell free systems. The terms “z z vitro protein synthesis” (IVPS), “in vitro translation,” “cell-free translation,” “RNA template-driven in vitro protein synthesis,” “RNA template-driven cell-free protein synthesis,” and “cell-free protein synthesis” are used interchangeably herein and are intended to refer to any method for cell-free synthesis of a protein. IVTT, including coupled transcription and transcription, is one non-limiting example of IVPS.
[0034] As used herein, the terms “nucleic acid,” “polynucleotide,” and “oligonucleotide” are used interchangeably and refer to naturally occurring or synthetic polymeric forms of nucleotides. The oligonucleotides and nucleic acid molecules of the present disclosure may be formed from naturally occurring nucleotides, for example, deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) molecules. Alternatively, the naturally occurring oligonucleotides may include structural modifications to alter their properties, such as in peptide nucleic acids (PNA) or locked nucleic acids (LNA). The solid phase synthesis of oligonucleotides and nucleic acid molecules with naturally occurring or artificial bases is well known in the art. The terms should be understood to include equivalents, analogs of either RNA or DNA made from nucleotide analogs, and singlestranded or double-stranded polynucleotides as applicable to the described embodiment. Nucleotides useful in the disclosure include, for example, naturally occurring nucleotides (for example, ribonucleotides or deoxyribonucleotides), natural or synthetic modifications of nucleotides, or artificial bases.
[0035] As used herein, the term monomer refers to a member of a set of small molecules that can be joined together to form an oligomer, a polymer, or a compound composed of two or moreAttorney Docket: 114203-3092members. The particular ordering of monomers within a polymer is referred to herein as the “sequence” of the polymer. The set of monomers includes but is not limited to, for example, the set of common L-amino acids, the set of D-amino acids, the set of synthetic and / or natural amino acids, the set of nucleotides and the set of pentoses and hexoses.
[0036] As used herein, the term “predefined sequence” means that the sequence of the polymer is designed or known and chosen before synthesis or assembly of the polymer. In particular, aspects of the disclosure are described herein primarily concerning the preparation of nucleic acid molecules, the sequence of the oligonucleotides or polynucleotides being known and chosen before the synthesis or assembly of the nucleic acid molecules. In some embodiments of the technology provided herein, immobilized oligonucleotides or polynucleotides are used as a source of material. In various embodiments, the methods described herein use oligonucleotides. Their sequence is determined based on the sequence of the final polynucleotide constructs to be synthesized. In one embodiment, oligonucleotides are short nucleic acid molecules. For example, oligonucleotides may be from 10 to about 300 nucleotides, from 20 to about 400 nucleotides, from 30 to about 500 nucleotides, from 40 to about 600 nucleotides, or more than about 600 nucleotides long. However, shorter or longer oligonucleotides may be used. Oligonucleotides may be designed to have different lengths. In some embodiments, the polynucleotide construct's sequence may be divided into a plurality of shorter sequences that can be synthesized in parallel and assembled into a single or plurality of desired polynucleotide constructs using the methods described herein.
[0037] The terms “subject,” “individual,” and “patient” are used interchangeably herein to refer to a vertebrate, e g. a mammal, or a human. Mammals include, but are not limited to, mice, non-human primates, humans, farm animals, sport animals, and pets. Tissues, cells, and the progeny of a biological entity obtained in vivo or cultured in vitro are also encompassed.
[0038] The term “bead” or “microbead” as used herein refers to a substantially spherical particle having a diameter in the range of 0.1 to 1000 micrometers, wherein said diameter is more in the range of 1 to 500 micrometers, in the range of 10 to 300 micrometers, or in the range of 20 to 100 micrometers, wherein said particle is composed of a material selected from the group consisting of polystyrene, polymethylmethacrylate (PMMA), silica, magnetic iron oxidecontaining polymers, agarose, dextran, and combinations thereof. The microbead may optionally comprise surface modifications including, but not limited to, carboxyl groups, amine groups,Attorney Docket: 114203-3092streptavidin, antibodies, oligonucleotides, or fluorescent moieties. Tn certain embodiments, the microbead is substantially monodisperse, having a coefficient of variation in diameter of less than 15%, less than 10%, less than 5%, and less than 2%. The microbead is capable of suspension in aqueous media and is suitable for use in droplet-based or microwell-based assays, including but not limited to immunoassays, nucleic acid amplification reactions, and cell-based assays.
[0039] As used herein, the term "chaperone" refers to a protein that assists other proteins in achieving proper folding, assembly, and stability during synthesis or under stress conditions. Chaperones prevent misfolding and aggregation, which can compromise protein function and lead to cellular stress. They are essential in both in vivo and in vitro systems to maintain protein integrity during expression. Chaperones may also facilitate refolding of denatured proteins and support complex assembly processes. In some embodiments, a chaperone comprises heat shock protein 70 (Hsp70), heat shock protein 90 (Hsp90), Growth Arrest And DNA-Damage-Inducible 34 (GADD34), vaccinia virus K3L, or a disulfide isomerase.
[0040] As used herein, the phrase "crowding agent" refers to a compound that simulates the dense molecular environment of a cell, thereby promoting proper protein folding and / or stability in vitro. Crowding agents create conditions that mimic intracellular macromolecular crowding, which influences protein folding kinetics and thermodynamics. These agents can enhance reaction efficiency and improve yield in cell-free systems. They are commonly used in biochemical reactions to replicate physiological conditions. In some embodiments, crowding agent comprises polyethylene glycol, an alcohol, a sugar, an amino acid, or a polyol.
[0041] As used herein, the phrase "energy regeneration enzyme" refers to an enzyme that replenishes energy molecules such as ATP during biochemical reactions, ensuring sustained activity in cell-free protein synthesis systems. These enzymes catalyze reactions that regenerate high-energy phosphate compounds from precursors, maintaining the energy balance required for transcription and translation. Without energy regeneration, protein synthesis reactions would quickly stall due to depletion of ATP and GTP. Energy regeneration systems are critical for long-duration or high-yield protein expression. In some embodiments, energy regeneration enzyme comprises creatine kinase or pyruvate kinase.
[0042] As used herein, the phrase "reducing agent" refers to a compound that maintains a reducing environment in biochemical reactions, preventing unwanted disulfide bond formationAttorney Docket: 114203-3092and oxidative damage. Reducing agents stabilize thiol groups in proteins, ensuring correct folding and activity. They are particularly important in cell-free systems where oxidative stress can impair protein integrity. Reducing conditions also support the activity of enzymes sensitive to oxidation. In some embodiments, reducing agent comprises dithiothreitol, beta mercaptoethanol, or a solution of reduced and oxidized glutathione.
[0043] As used herein, the phrase "divalent cation" refers to a positively charged ion with two valence electrons that plays a critical role in stabilizing nucleic acids and supporting enzymatic reactions. Divalent cations are essential cofactors for polymerases and ribosomes in transcription and translation processes. They influence the structural integrity of macromolecules and reaction efficiency. Common examples include magnesium and calcium ions. In some embodiments, divalent cation comprises magnesium or calcium.
[0044] As used herein, the phrase “halide salt” refers to a compound formed by the combination of a halide anion (such as fluoride, chloride, bromide, or iodide) with a suitable cation, typically a metal or ammonium ion. In some embodiments, a halide salt comprises sodium chloride, potassium chloride, lithium bromide, calcium chloride, magnesium chloride, sodium iodide, potassium iodide, or ammonium chloride.
[0045] As used herein, “U / pL” stands for units per microliter, which expresses how many enzyme units, cells, or molecules are present in one microliter (pL) of solution. As used herein, one unit refers to the amount of an enzyme that catalyzes the conversion of 1 micromole of substrate per minute under defined conditions (e g., at 30 °C).
[0046] Various embodiments are described hereinafter. It should be noted that the specific embodiments are not intended as an exhaustive description or as a limitation to the broader aspects discussed herein. One aspect described in conjunction with a particular embodiment is not necessarily limited to that embodiment and can be practiced with any other embodiment(s). Reference throughout this specification to “one embodiment,” “an embodiment,” and “an example embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. Thus, appearances of the phrases “in one embodiment,” “in an embodiment,” or “an example embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment but may.Attorney Docket: 114203-3092
[0047] Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner, as would be apparent to a person skilled in the art from this disclosure in one or more embodiments. Furthermore, while some embodiments described herein include some but not other features included in other embodiments, combinations of features of different embodiments are meant to be within the scope of the disclosure. For example, in the appended claims, any of the claimed embodiments can be used in any combination.
[0048] All publications, published patent documents, and patent applications cited herein are hereby incorporated by reference to the same extent as though each publication, published patent document, or patent application was specifically and individually indicated as being incorporated by reference.OVERVIEW
[0049] Existing cell-free protein expression systems suffer from several technical limitations that hinder their utility in high-throughput applications. Traditional approaches typically require multiple days for polypeptide production and rely on labor-intensive, column-based purification methods that employ protein-specific tags and buffers. These systems lack the flexibility to express diverse protein types simultaneously and are poorly suited for parallel expression, purification, and quantification in discrete volumes. The purification steps present a particular bottleneck, as they are difficult to automate or scale and significantly impact overall throughput.
[0050] The technical problem addressed by the present disclosure is to provide a rapid, scalable, and versatile cell-free protein expression system capable of high-throughput parallel polypeptide synthesis and purification without requiring protein-specific optimization or labor-intensive column-based workflows.
[0051] The disclosed solution comprises an expression library of templates, each encoding a different polypeptide and incorporating standardized elements including an N-terminal translation enhancer sequence, a translation initiation sequence, a first purification tag, an optional solubility tag, an open reading frame encoding the polypeptide with a 3' untranslated region sequence or engineered poly-A tail or both, and optionally a detection tag or second purification tag. These templates are allocated into individual discrete volumes where polypeptides are expressed in the presence of cell lysate and reaction buffer, captured by a first capture agent, washed with purification buffers, and released via protease cleavage of a site adjacent to the purification tag.Attorney Docket: 114203-3092The combination of standardized expression elements enables robust expression across diverse polypeptide classes, while the capture-and-release purification approach eliminates the need for column-based methods and enables automation-compatible workflows.
[0052] The technical effect achieved is the production of substantially pure polypeptides within an hourly timescale instead of days, with simultaneous expression of diverse polypeptide classes using automatable workflows compatible with high-throughput formats. The system further enables direct phenotypic screening by culturing cells in the individual discrete volumes containing expressed polypeptides and measuring cellular responses through techniques including RNA sequencing, single-cell RNA sequencing, spatial transcriptomics, flow cytometry, microscopy imaging, transwell migration assays, immunoassays, proliferation assays, or coculture systems.
[0053] Disclosed embodiments provide an advanced polypeptide synthesis platform designed for the rapid, on-demand production of hundreds or more polypeptides, tailored for specific experiments or assays and accommodating diverse polypeptide sequences. This solution utilizes an optimized cell-free polypeptide synthesis platform, where expression templates are efficiently translated (RNA-based expression templates) or transcribed into RNA and translated into polypeptides (DNA-based expression templates) within hours, significantly faster than traditional cell-based methods, which typically require days. Expression templates may comprise synthetic or naturally made dsDNA. The platform also features streamlined polypeptide purification using optimized pulldowns and washes, which are compatible with high-throughput formats and automation, eliminating the labor-intensive steps of traditional column-based purification. These systems and methods can be successfully applied to a broad range of polypeptide classes, including, for example, interleukins, chemokines, growth factors, nanobodies, and fluorescent proteins, offering a scalable and efficient alternative to conventional protein production methods. The systems and methods can also be used to make intracellular proteins (with or without attached transmembrane domains, transmembrane delivery peptides, cell-penetrating peptides, etc.), genetic variants of the same protein for screening, artificially designed polypeptides, and more.
[0054] Unlike traditional cell-based polypeptide production, which often requires several days, the protein synthesis platform disclosed herein provides reliable and robust polypeptides on an hourly timescale instead of days. In one embodiment, the methods disclosed herein may beAttorney Docket: 114203-3092completed within 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, or 36 hours. This accelerated process significantly reduces turnaround time. The embodiments disclosed herein combine high-purity, high-throughput, and / or high-flexibility polypeptide purification. The platform employs a method compatible with the simultaneous expression of diverse polypeptide classes, utilizing automatable workflows. In contrast, existing solutions often rely on polypeptide-specific tags and buffers, requiring labor-intensive, low-throughput column-based purification steps that are difficult to scale. This streamlined, scalable platform enables faster, more flexible polypeptide production than conventional methods, making it ideal for high-demand research and industrial applications.
[0055] The polypeptides are expressed in individual discrete volumes, which enables a number of further screening techniques. For example, cells may be cultured in individual discrete volumes, and changes in phenotype to the variously expressed polypeptides measured. Accordingly, the platform disclosed herein may be used to experimentally profile cell responses to a set of polypeptides. This may, in turn, be used to identify drivers of a particular disease and or identify new therapeutic targets. The cells may be further cultured in the presence of an agent, such as a therapeutic agent, to determine the effects of polypeptide expression on treatment efficacy. The expression library may comprise different variants of the same or a set of polypeptides, and the system is used to screen for variant-specific effects on cell phenotype. The ability to screen multiple polypeptides, including variants thereof, may also guide polypeptide engineering design for many functions, including but not limited to screening and design of de novo binders and effectors for different therapeutic or diagnostic applications.IN VITRO HIGH-THROUGHPUT POLYPEPTIDE SYNTHESIS AND SCREENING SYSTEM
[0056] In one aspect, embodiments disclosed herein provide a high-throughput in vitro or cell-free polypeptide synthesis system. In an embodiment, the system comprises an expression library, a plurality of individual discrete volumes, and a set of reagents for polypeptide expression and purification. Individual expression templates from the library are placed in discrete volumes where they can be expressed and purified using reagents, which will be described in further detail below. In an embodiment, more than one expression template may be added to an individual discreteAttorney Docket: 114203-3092volume. Cells may then be added to the individual discrete volume to measure, for example, changes in phenotype in the presence of the different proteins expressed.Expression Library
[0057] The library comprises a set of expression templates (or cassettes), each encoding a polypeptide. The expression template may be RNA- or DNA-based. In an embodiment, the expression template is an mRNA, a circular double-stranded DNA, linear dsDNA, branched / mixed ssDNA and dsDNA, or a genomic template. The RNA and DNA-based transcription templates may be synthetic or naturally made.
[0058] A expression template may also be referred to as an expression vector. Various vectors are publicly available. The vector may, for example, be in the form of a plasmid, cosmid, viral particle, or phage. Both expression and cloning vectors contain a nucleic acid sequence that enables the vector to replicate in one or more selected host cells. Such vector sequences are well-known for various bacteria, yeast, and viruses. Useful expression templates include for example, segments of chromosomal, non-chromosomal, and synthetic DNA sequences. Suitable vectors include, but are not limited to, derivatives of SV40 and pcDNA and known bacterial plasmids such as col El, pCRl, pBR322, pMal-C2, pET, pGEX as described by Smith et al., Gene 57:31-40 (1988), pMB9 and derivatives thereof, plasmids such as RP4, phage DNAs such as the numerous derivatives of phage I such as NM989, as well as other phage DNA such as Ml 3 and filamentous single-stranded phage DNA; yeast plasmids such as the 2-micron plasmid or derivatives of the 2 pm plasmid, as well as centromeric and integrative yeast shuttle vectors; vectors useful in eukaryotic cells such as vectors useful in insect or mammalian cells; vectors derived from combinations of plasmids and phage DNAs, such as plasmids that have been modified to employ phage DNA or the expression control sequences; and the like. The requirements are that the vectors are replicable and viable in the host cell of choice. Low- or high-copy number vectors may be used as desired. The expression template may be an RNA molecule (e.g., mRNA) or a nucleic acid that encodes an mRNA (e.g., RNA, DNA) and be in any form (e.g., linear, circular, supercoiled, single-stranded, doublestranded, etc.).
[0059] A map of an example expression template is shown in FIG. 10. Each vector may encode a different polypeptide or variant of a given polypeptide. A different variant may comprise a polypeptide with a particular mutation or set of mutations relevant to a reference polypeptide. ItAttorney Docket: 114203-3092may also include insertions or deletions of several amino acids. The insertions or deletions may add or remove all or a portion of different polypeptide domains.
[0060] The expression templates can include appropriate promoter and translation sequences for in vitro protein synthesis or in vitro transcription / translation. Any suitable promoter, such as the area B, tac promoter, T7, T3, or SP6 promoters, or one that is artificially designed, can be used. The promoter is placed so that it is operably linked to the DNA sequences of the disclosure such that such sequences are expressed. Each promoter sequence may be native or foreign to the polynucleotide sequence to which it is operably linked. Each promoter sequence may be any nucleic acid sequence that shows transcriptional activity. A variety of promoters can be utilized. For example, the different promoter sequences may have different promoter strengths. In one embodiment, the library of promoter sequences comprises promoter variant sequences. In one embodiment, the promoter variants cover a wide range of promoter activities, from the weak promoter to the strong promoter. Putative promoter sequences may be identified using computerized algorithms such as the Neural Network of Promoter Prediction software (Demeter et al. (Nucl. Acids. Res. 1991, 19:1593-1599). Putative promoters may also be identified by examination of the family of genomes and homology analysis.
[0061] In one embodiment, the disclosure provides libraries of expression templates encoding at least about 5 different polypeptides or polypeptides variants, at least about 10 different polypeptides or polypeptides variants, at least about 20 different polypeptides or polypeptides variants, at least about 30 different polypeptides or polypeptides variants, at least about 40 different polypeptides or polypeptides variants, at least about 50 different polypeptides or polypeptides variants, at least about 60 different polypeptides or polypeptides variants, at least about 70 different polypeptides or polypeptides variants, at least about 80 different polypeptides or polypeptides variants, at least about 90 different polypeptides or polypeptides variants, at least about 100 different polypeptides or polypeptide variants, at least about 200 different polypeptides or polypeptides variants, at least about 300 different polypeptides or polypeptide variants, at least about 400 different polypeptides or polypeptide variants, at least about 500 different polypeptides or polypeptide variants, at least about 600 different polypeptides or polypeptide variants, at least about 700 different polypeptides or polypeptide variants, at least about 800 different polypeptides or polypeptide variants, at least about 900 different polypeptides or polypeptide variants, at leastAttorney Docket: 114203-3092about 1,000 different polypeptides or polypeptide variants, at least about 2,000 different polypeptides or polypeptide variants, at least about 3,000 different polypeptides or polypeptide variants, at least about 4,000 different polypeptides or polypeptide variants, at least about 5,000 different polypeptides or polypeptide variants, at least about 6,000 different polypeptides or polypeptide variants, at least about 7,000 different polypeptides or polypeptide variants, at least about 8,000 different polypeptides or polypeptide variants, at least about 9,000 different polypeptides or polypeptide variants, or at least about 10,000 different polypeptides or polypeptide variants or more different polypeptides or polypeptide variants. In one embodiment, the library comprises about 6 different polypeptides or polypeptide variants. In another embodiment, the library provides about 12 different polypeptides or polypeptide variants. In another embodiment, the library provides about 24 different polypeptides or polypeptide variants. In another embodiment, the library provides about 96 different polypeptides or polypeptide variants. In another embodiment, the library provides about 384 different polypeptides or polypeptide variants. In another embodiment, the library provides about 1,536 different polypeptides or polypeptide variants. In another embodiment, the library comprises about 3,456 different polypeptides or polypeptide variants. In another embodiment, the library comprises about 9,600 different polypeptides or polypeptide variants.
[0062] The expression templates comprise a set of features that enable robust and reliable expression across a diverse number of polypeptides in contrast to current approaches, which typically must tailor a set of expression elements to a particular protein. In an embodiment, the expression template comprises one or more expression elements selected from the group of an N-terminal translation enhancer sequence, a translation initiation sequence, a first purification tag, a solubility tag, a 3’ untranslated region (3'-UTR) sequence appended to an open reading frame (ORF) encoding a polypeptide, an engineered poly-A tail appended to the 3 '-end of the ORF, a detection tag, and a second purification tag. The ORF may be codon optimized. In one embodiment, the ORF is codon optimized for eukaryotic cells, especially mammalian cells, e.g., a human cell. In another embodiment, the ORF is human codon optimized. The first and second purification tags may further comprise a protease cleavage sequence, allowing the purification tags to be released from the expressed polypeptide.Attorney Docket: 114203-3092
[0063] In an embodiment, the expression template comprises a first purification tag and an ORF. In another embodiment, the expression template further comprises a 3'-UTR sequence appended to the ORF. In another embodiment, the expression template further comprises an engineered poly-A tail appended to a 3'-end of the ORF. In another embodiment, the expression template further comprises a 3'-UTR and an engineered poly-A tail.
[0064] In an embodiment, the expression template comprises an N-terminal translation enhancer sequence, a first purification tag, and an ORF. In another embodiment, the expression template further comprises a 3'-UTR sequence appended to the ORF. In another embodiment, the expression template further comprises an engineered poly-A tail appended to a 3' end of the ORF. In another embodiment, the expression template further comprises a 3'-UTR and an engineered poly-A tail.
[0065] In an embodiment, the expression template comprises a translation initiation sequence, a first purification tag, and an ORF. In another embodiment, the expression template further comprises a 3’UTR sequence appended to the ORF. In another embodiment, the expression template further comprises an engineered poly-A tail appended to a 3' end of the ORF. In another embodiment, the expression template further comprises a 3'-UTR and an engineered poly-A tail.
[0066] In an embodiment, the expression template comprises a solubility tag, a first purification tag, and an ORF. In another embodiment, the expression template further comprises a 3'-UTR sequence appended to the ORF. In another embodiment, the expression template further comprises an engineered poly-A tail appended to a 3' end of the ORF. In another embodiment, the expression template further comprises a 3'-UTR and an engineered poly-A tail.
[0067] In an embodiment, the expression template comprises a first purification tag, an ORF, and a detection tag. In another embodiment, the expression template further comprises a 3'-UTR sequence appended to the ORF. In another embodiment, the expression template further comprises an engineered poly-A tail appended to a 3' end of the ORF. In another embodiment, the expression template further comprises a 3'-UTR and an engineered poly-A tail.
[0068] In an embodiment, the expression template comprises a first purification tag, an ORF, and a second purification tag. In another embodiment, the expression template further comprises a 3'-UTR sequence appended to the ORF. In another embodiment, the expression template furtherAttorney Docket: 114203-3092comprises an engineered poly-A tail appended to a 3' end of the ORF. In another embodiment, the expression template further comprises a 3' UTR and an engineered poly-A tail.
[0069] In an embodiment, the expression template comprises an N-terminal translation enhancer sequence, a first purification tag, and an ORF. In another embodiment, the expression template further comprises a 3'-UTR sequence appended to the ORF. In another embodiment, the expression template further comprises an engineered poly-A tail appended to a 3' end of the ORF. In another embodiment, the expression template further comprises a 3' UTR and an engineered poly-A tail.
[0070] In another embodiment, the expression template comprises an N-terminal translation enhancer, a translation initiation sequence, a first purification tag, and an ORF. In another embodiment, the expression template further comprises a 3'-UTR sequence appended to the ORF. In another embodiment, the expression template further comprises an engineered poly-A tail appended to a 3' end of the ORF. In another embodiment, the expression template further comprises a 3'-UTR and engineered poly-A tail.
[0071] In another embodiment, the expression template comprises an N-terminal translation enhancer, a translation initiation sequence, a first purification tag, a solubility tag, and an ORF. In another embodiment, the expression template further comprises a 3'-UTR sequence appended to the ORF. In another embodiment, the expression template further comprises an engineered poly-A tail appended to a 3' end of the ORF. In another embodiment, the expression template further comprises a 3' UTR and engineered poly-A tail.
[0072] In another embodiment, the expression template comprises an N-terminal translation enhancer, a translation initiation sequence, a first purification tag, a solubility tag, an ORF, and a detection tag. In another embodiment, the expression template further comprises a 3'-UTR sequence appended to the ORF. In another embodiment, the expression template further comprises an engineered poly-A tail appended to a 3' end of the ORF. In another embodiment, the expression template further comprises a 3'-UTR and engineered poly-A tail.
[0073] In another embodiment, the expression template comprises an N-terminal translation enhancer, a translation initiation sequence, a first purification tag, a solubility tag, an ORF, a detection tag, and a second purification tag. In another embodiment, the expression template further comprises a 3'-UTR sequence appended to the ORF. In another embodiment, the expressionAttorney Docket: 114203-3092template further comprises an engineered poly-A tail appended to a 3' end of the ORF. In another embodiment, the expression template further comprises a 3'-UTR and engineered poly-A tail.
[0074] In an embodiment, the expression template comprises an N-terminal translation enhancer sequence, a first purification tag, a solubility tag, and ORF encoding of the polypeptide. In another embodiment, the expression template further comprises a 3'-UTR sequence appended to the ORF. In another embodiment, the expression template further comprises an engineered poly-A tail appended to a 3' end of the ORF. In another embodiment, the expression template further comprises a 3' UTR and engineered poly-A tail.
[0075] In an embodiment, the expression template comprises an N-terminal translation enhancer sequence, a first purification tag, a solubility tag, ORF encoding the polypeptide, and a detection tag. In another embodiment, the expression template further comprises a 3'-UTR sequence appended to the ORF. In another embodiment, the expression template further comprises an engineered poly-A tail appended to a 3' end of the ORF. In another embodiment, the expression template further comprises a 3'-UTR and engineered poly-A tail.
[0076] In an embodiment, the expression template comprises an N-terminal translation enhancer sequence, a first purification tag, a solubility tag, ORF encoding of the polypeptide, a detection tag, and a second purification tag. In another embodiment, the expression template further comprises a 3'-UTR sequence appended to the ORF. In another embodiment, the expression template further comprises an engineered poly-A tail appended to a 3' end of the ORF. In another embodiment, the expression template further comprises a 3' UTR and engineered poly-A tail.
[0077] In an embodiment, the expression template comprises a translation initiation sequence, a first purification tag, and an ORF. In another embodiment, the expression template further comprises a 3'-UTR sequence appended to the ORF. In another embodiment, the expression template further comprises an engineered poly-A tail appended to a 3' end of the ORF. In another embodiment, the expression template further comprises a 3' UTR and engineered poly-A tail.
[0078] In an embodiment, the expression template comprises a translation initiation sequence, a first purification tag, a solubility tag, and an ORF. In another embodiment, the expression template further comprises a 3'-UTR sequence appended to the ORF. In another embodiment, the expression template further comprises an engineered poly-A tail appended to a 3' end of the ORF.Attorney Docket: 114203-3092In another embodiment, the expression template further comprises a 3'-UTR and engineered poly-A tail.
[0079] In an embodiment, the expression template comprises a translation initiation sequence, a first purification tag, a solubility tag, an ORF, and a detection tag. In another embodiment, the expression template further comprises a 3'-UTR sequence appended to the ORF. In another embodiment, the expression template further comprises an engineered poly-A tail appended to a 3' end of the ORF. In another embodiment, the expression template further comprises a 3'-UTR and engineered poly-A tail.
[0080] In an embodiment, the expression template comprises a translation initiation sequence, a first purification tag, a solubility tag, an ORF, a detection tag, and a second purification tag. In another embodiment, the expression template further comprises a 3’UTR sequence appended to the ORF. In another embodiment, the expression template further comprises an engineered poly-A tail appended to a 3' end of the ORF. In another embodiment, the expression template further comprises a 3’ UTR and engineered poly-A tail.
[0081] In an embodiment, the expression template comprises a first purification tag, a solubility tag, and an ORF. In another embodiment, the expression template further comprises a 3'-UTR sequence appended to the ORF. In another embodiment, the expression template further comprises an engineered poly-A tail appended to a 3' end of the ORF. In another embodiment, the expression template further comprises a 3 '-UTR and engineered poly-A tail.
[0082] In an embodiment, the expression template comprises a first purification tag, a solubility tag, an ORF, and a detection tag. In another embodiment, the expression template further comprises a 3'-UTR sequence appended to the ORF. In another embodiment, the expression template further comprises an engineered poly-A tail appended to a 3' end of the ORF. In another embodiment, the expression template further comprises a 3'-UTR and engineered poly-A tail.
[0083] In an embodiment, the expression template comprises a first purification tag, a solubility tag, an ORF, a detection tag, and a second purification tag. In another embodiment, the expression template further comprises a 3'-UTR sequence appended to the ORF. In another embodiment, the expression template further comprises an engineered poly-A tail appended to a 3' end of the ORF. In another embodiment, the expression template further comprises a 3'-UTR and engineered poly-A tail.Attorney Docket: 114203-3092
[0084] In an embodiment, the expression template comprises a first purification tag, an ORF, and a detection tag. In another embodiment, the expression template further comprises a 3'-UTR sequence appended to the ORF. In another embodiment, the expression template further comprises an engineered poly-A tail appended to a 3' end of the ORF. In another embodiment, the expression template further comprises a 3'-UTR and engineered poly-A tail.
[0085] In an embodiment, the expression template comprises a first purification tag, an ORF, a detection tag, and a second purification tag. In another embodiment, the expression template further comprises a 3'-UTR sequence appended to the ORF. In another embodiment, the expression template further comprises an engineered poly-A tail appended to a 3' end of the ORF. In another embodiment, the expression template further comprises a 3' UTR and engineered poly-A tail.
[0086] In an embodiment, the expression template comprises a first purification tag, an ORF, and a second purification tag. In another embodiment, the expression template further comprises a 3 '-UTR sequence appended to the ORF. In another embodiment, the expression template further comprises an engineered poly-A tail appended to a 3' end of the ORF. In another embodiment, the expression template further comprises a 3' UTR and engineered poly-A tail.
[0087] The following N-terminal translation enhancer sequences, translation initiation sequences, first and second purification tags, solubility tags, 3' UTR sequences, engineered poly-A tail sequences, and detection tags may be used with any of the above embodiments.N-terminal Translation Enhancer Sequence
[0088] Translational elements in N-terminal regions can modulate polypeptide production. An N-terminal translation enhancer sequence may be a prokaryotic sequence, such as a Shine-Dalgamo consensus sequence. In an embodiment, the translational enhancer sequence is a viral nucleotide sequence, such as a plant virus, e.g., Tobacco Mosaic Virus (TMV), Alfalfa Mosaic Virus (AMV); Tobacco Etch Virus (TEV) (SEQ ID NO: 1); Potato Virus Y (PVY); Turnip Mosaic (poty) Virus and Pea Seed Borne Mosaic Virus. In one embodiment, the N-terminal translation enhancer sequence is selected from the chloramphenicol acetyltransferase (CAT) sequence, ubiquitin-activating enzyme (El) sequence, activity -regulated cytoskeleton-associated protein 1 (ARC-1) sequence, or an internal ribosome entry site (IRES) sequence. In one embodiment, the N-terminal translational enhancer sequence is from a 5' UTR of the aforementioned sequences.Attorney Docket: 114203-3092
[0089] In some embodiments, the TEV enhancer sequence comprises:TAAATAACAAATCTCAACACAACATATACAAAACAAACGAATCTCAAGCAATCAAG CATTCTACTTCTATTGCAGCAATTTAAATCATTTCTTTTAAAGCAAAAGCAATTTTCT GAAAATTTTCACCATTTACGAACGAT AGCA (SEQ ID NO: 1).
[0090] In one embodiment, the N-terminal translation enhancer sequence is a 5'-UTR sequence comprises a translational enhancer sequence. A 5'-UTR comprises may be homologous to another nucleotide sequence in the 5'-UTR. In one embodiment, the 5'-UTR is exogenous to another sequence in the 5' UTR. A translational enhancer sequence sometimes is 10 or more nucleotides in length, 11 or more nucleotides in length, and sometimes 12 or more, 13 or more, 14 or more, 15 or more, 20 or more, 25 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more or 100 or more nucleotides in length. An N-terminal translation enhancer sequence often is located between the 3' end of a promoter and the 5' end of a translated nucleotide sequence. The distance between the 3' end of the translational enhancer sequence and the 5' end of a translated nucleotide sequence sometimes is about 5 to about 50 nucleotides, about 10 to about 20 nucleotides, or about 15 nucleotides in length.
[0091] In one embodiment, the N-terminal translation enhancer sequence is ARC-1 or an ARC-1 like sequence, such as, for example, GCCGGCGGAG (SEQ ID NO: 6), CUCAUAAGGU, GACUUUGAUU (SEQ ID NO: 7), CGGAACCCAA (SEQ ID NO: 8), AUACUCCCCC (SEQ ID NO: 9), or CCUUGCGACC (SEQ ID NO: 10), or a substantially identical sequence thereof.
[0092] In one embodiment, the N-terminal translation enhancer sequence is an internal IRES sequence, such as one or more of EMBL nucleotide sequences J04513, X87949, M95825, M12783, AF025841, AF013263, AF006822, M17169, M13440, M22427, D14838 and M17446, or a substantially identical nucleotide sequence thereof. An IRES sequence may be a type I IRES (e.g., from enterovirus (e.g., poliovirus), rhinovirus (e.g., human rhinovirus)), a type II IRES (e.g., from cardiovirus (e.g., encephalomyocraditis virus), aphthovirus (e.g., foot-and-mouth disease virus)), a type III IRES (e.g., from Hepatitis A virus) or other picornavirus sequences (e.g., Paulos et al. supra, and Jackson et al., RNA 1: 985-1000 (1995)).
[0093] In an embodiment, the N-terminal translation enhancer sequence is the six-base Shine-Dalgarno consensus sequence AGGAGG (SEQ ID NO: 11). In an embodiment, the Shine-Dalgarno consensus sequence is AGGAGGU (SEQ ID NO: 12).Attorney Docket: 114203-3092
[0094] In one embodiment, the N-terminal translation enhancer sequence is Tobacco Mosaic Virus (TEV) or a substantially identical sequence. In one embodiment, the N-terminal terminal enhancer sequence is GGGAAAGCUU (SEQ ID NO: 13).
[0095] In one embodiment, the N-terminal translation enhancer sequence is a Tobacco Etch Virus (SEQ ID NO: 1) or a substantially identical sequence.
[0096] In another embodiment, the N-terminal translation enhancer sequence is a chloramphenicol acetyltransferase (CAT) sequence or a substantially identical sequence thereof.
[0097] In another embodiment, the N-terminal translation enhancer sequence is a ubiquitin-activating enzyme (El) or substantially identical sequence.Purification Tags
[0098] Introducing a first purification tag can lead to selective pull-down of full-length polypeptides and, thus, high-purity polypeptides without needing poorly scalable, highly manual column-based purifications. In one embodiment, wherein the first purification tag is selected from an haloalkane dehalogenase-derived tag, Albumin-binding protein (ABP), Alkaline Phosphatase (AP), AU1 epitope, AU5 epitope, Bacteriophage T7 epitope (T7-tag), Bacteriophage V5 epitope (V5-tag), Biotin-carboxy carrier protein (BCCP), Bluetongue virus tag (B-tag), Calmodulin binding peptide (CBP), Chloramphenicol Acetyl Transferase (CAT), Cellulose binding domain (CBP), Chitin binding domain (CBD), Choline-binding domain (CBD), Dihydrofolate reductase (DHFR), E2 epitope, FLAG epitope, Galactose-binding protein (GBP), Green fluorescent protein (GFP), Glu-Glu (EE-tag), Glutathione S-transferase (GST), Human influenza hemagglutinin (HA), Histidine affinity tag (HAT), Horseradish Peroxidase (HRP), HSV epitope, Ketosteroid isomerase (KSI), KT3 epitope, LacZ, Luciferase, Maltose-binding protein (MBP), Myc epitope, NusA, PDZ domain, PDZ ligand, Polyarginine (Arg-tag), Polyaspartate (Asp-tag), Polycysteine (Cys-tag), Polyhistidine (His-tag), Polyphenylalanine (Phe-tag), Profinity eXact, Protein C, SI -tag, S-tag, Streptavidin-binding peptide (SBP), Staphylococcal protein A (Protein A), Staphylococcal protein G (Protein G), Strep-tag, Streptavidin, Small Ubiquitin-like Modifier (SUMO), Tandem Affinity Purification (TAP), T7 epitope, Thioredoxin (Trx), TrpE, Ubiquitin, Universal, and VSV-G.
[0099] The introduction of an optional second purification tag can increase purity. In one embodiment, the optional second purification tag is a histidine (His) tag, which enables anAttorney Docket: 114203-3092additional round of purification. It may also be any of the tags mentioned above, as long as the primary and secondary tags are different. Proteins with anN-terminal His tag can undergo a second series of washes to achieve greater purity. His-tagged proteins may then be eluted using either imidazole — retaining their attachment to solubility tags — or via cleavage with a second orthogonal protease, which simultaneously releases the protein and removes the solubility tag.Solubility Tags
[0100] The introduction of an N-terminal solubility tag can enhance protein production. In one embodiment the solubility tag is selected from a small ubiquitin-like modifier (SUMO) tag, SUMOstar protease, a glutathione S-transferase (GST)-tag, a maltose-binding protein (MBP)-tag, a thioredoxin (Trx)-tag, RRGRRGRRGRRGRRG (SEQ ID NO: 4) tag, a NusA tag, an N-terminal extension (NEXT) tag, and Bl domain of Streptococcal protein G (GB1). In one embodiment, the template comprises a protease cleavage site after the solubility tag, wherein the solubility tag is selected from the cleavage site motifs of SUMOstar protease, HRV3C protease, and HaloTEV.3’ UTR Sequences
[0101] Sequences found in 3'-UTR to the open reading frame to be translated can also affect translation. 3'-UTRs are known to have stretches of Adenosines and Uridines embedded in them. These AU-rich signatures are particularly prevalent in genes with high rates of turnover. Based on their sequence features and functional properties, the AU-rich elements (AREs) can be separated into three classes (Chen et al., 1995). Class I AREs contain several dispersed copies of an AUUUA motif within U-rich regions. C-Myc and MyoD contain class I AREs. Class II AREs possess two or more overlapping UUAUUUA(U / A)(U / A) nonamers. Molecules containing this type of AREs include GM-CSF and TNF-a. Class III ARES are less well defined. These U-rich regions do not contain an AUUUA motif. Jun and Myogenin are two well-studied examples from this class. Most proteins binding to the AREs are known to destabilize the messenger, whereas members of the ELAV family, most notably HuR, have been documented to increase mRNA stability. HuR binds to the AREs of all three classes. Engineering the HuR-specific binding sites into the 3'-UTR of nucleic acid molecules will lead to HuR binding and, thus, stabilization of the message in vivo.
[0102] Introduction, removal, or modification of 3'UTRAU-rich elements (AREs) can be used to modulate stability. One or more copies of an ARE can be introduced to make an ORF less stable, curtailing translation and decreasing the production of the resultant protein. Likewise, AREs canAttorney Docket: 114203-3092be identified and removed or mutated to increase stability and thus increase translation and production of the resultant protein. Cell-free polypeptide experiments can be conducted in relevant reaction mixes using polynucleotides disclosed herein, and polypeptide production can be assayed at various time points. For example, cell-free polypeptide synthesis reactions can be initiated with different ARE-engineering molecules, using an ELISA kit to the relevant protein or assays relevant to detection tags described below and assaying protein produced at 1 hr, 2 hr, 3 hr, 4 hr, 5 hr, and 6 hr after reaction initiation. .Engineered poly-A tail sequences
[0103] During RNA processing, a poly-A tail — a long chain of adenine nucleotides — is often added to a polynucleotide such as an mRNA molecule to enhance its stability. Following transcription, the 3' end of the RNA transcript is cleaved, exposing a 3' hydroxyl group. Poly-A polymerase then catalyzes the addition of adenine nucleotides, forming a poly-A tail typically ranging from approximately 100 to 250 residues in length. This process, known as polyadenylation, plays a critical role in RNA stability and function.
[0104] Poly-A tails may also be added after the RNA construct is exported from the nucleus. Modifications to the terminal groups of the poly-A tail can further stabilize the molecule. Polynucleotides, which may include structural moieties such as 2'-O-methyl modifications or dess' hydroxyl tails, can be engineered to include alternative polyadenylation structures. For example, constructs encoding histone mRNA may incorporate stem-loop structures instead of poly-A tails, as these structures, in combination with stem-loop binding proteins (SLBP), perform similar stability functions to poly-A binding proteins (PABP). This mechanism is described in works by Norbury (“Cytoplasmic RNA: A Case of the Tail Wagging the Dog,” Nature Reviews Molecular Cell Biology, 2013) and Junjie Li et al. (Current Biology, Vol. 15, 1501-1507, August 23, 2005).
[0105] Generally, the length of a poly-A tail, when present, is greater than 30 nucleotides in length. In another case, the poly-A tail is greater than 35 nucleotides in length (e.g., at least or greater than about 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, 2,000, 2,500, and 3,000 nucleotides).Attorney Docket: 114203-30925'-UTR Elements
[0106] Natural 5'-UTRs bear features which play roles in fortranslation initiation. They harbor signatures like Kozak sequences that are commonly known to be involved in the process by which the ribosome initiates translation of many genes. Kozak sequences have the consensus CCR(A / G)CCAUGG (SEQ ID NO: 2), where R is a purine (adenine or guanine) three bases upstream of the start codon (AUG), which is followed by another 'G'. 5'-UTR also have been known to form secondary structures that are involved in elongation factor binding.
[0107] By engineering the features typically found in abundantly expressed genes of specific target organs, one can enhance the stability and protein production. For example, introduction of 5'-UTR of liver-expressed mRNA, such as albumin, serum amyloid A, Apolipoprotein A / B / E, transferrin, alpha fetoprotein, erythropoietin, or Factor VIII, could be used to enhance expression of a nucleic acid molecule in hepatic cell lines or liver. Likewise, use of 5-UTR from other tissuespecific mRNA to improve expression in that tissue is possible - for muscle (MyoD, Myosin, Myoglobin, Myogenin, Herculin), for endothelial cells (Tie-1, CD36), for myeloid cells (C / EBP, AMLI, G-CSF, GM-CSF, CDIlb, MSR, Fr-1, i-NOS), for leukocytes (CD45, CD18), for adipose tissue (CD36, GLUT4, ACRP30, adiponectin) and for lung epithelial cells (SP-A / B / C / D).
[0108] Other non-UTR sequences maybe incorporated into the 5'-UTRs. For example, introns or portions of introns sequences may be incorporated into the flanking regions of the polynucleotides disclosed herein. Incorporation of intronic sequences may increase protein production as well as mRNA levels.
[0109] The 5'-UTR may selected for use in the present disclosure may be a structured UTR such as, but not limited to, 5'-UTRs to control translation. In one embodiment, the translational element in the 5'-UTR comprises internal ribosome entry site (IRES) elements or cap-independent translation enhancer sequences to initiate translation. Such sequence elements can circumvent the need to utilize 5 '-capped mRNA templates for efficient protein translation in the CFPS platforms.. In an embodiment, IRES is an EMCV IRES. In another embodiment, the IRES is an HCV IRES.Translation Initiation Sequence
[0110] The vectors disclosed herein may include a translation initiation sequence to assist in the initiation of translation. In one embodiment, the translation initiation sequence is a Kozak sequence. In an embodiment, the Kozak sequence is eukaryotic. In one embodiment, the eukaryoticAttorney Docket: 114203-3092Kozak sequence is a yeast Kozak sequence, a Drosophila species Kozak sequence, or a vertebrate Kozak sequence. In an embodiment, the vertebrate Kozak sequence may be a human Kozak sequence. In an embodiment, the Kozak sequence may be engineered. See e.g., Li et al. “Optimization of extended Kozak elements enhances recombinant protein expression in CHO cells” J. Biotech. 2024, 392:96-102; Sample et al. “Human 5’ UTR-mediated regulation of protein abundance in Yeast using nucleotide sequence activity relationships” ACS Synth. Biol. 2018, 7:2709-2714. In one embodiment, the Kozak sequence is selected from GATGATAATATGG (SEQ ID NO: 3) and GCCGCCACCATGG (SEQ ID NO: 4).Detection Tags[OHl] Introduction of an optional detection tag allows for measuring protein concentrations, whether in solution or bound to a surface. As used herein, the term “detection tag,” refers to a covalently linked chemical moiety that may be selectively bound and isolated. In some embodiments, the detection tag is an affinity tag in which the chemical moiety has a specific binding partner. Exemplary detection tags include, but are not limited to, Green Fluorescent Protein, biotin, streptavidin, FLAG (DYKDDDDK) (SEQ ID NO: 5), cyclodextrin, adamantane, and combinations thereof. In one embodiment, the detection tag is HiBiT-RR. Lin et al. bioRxiv 2024.05.14.594249.Solid Supports
[0112] The systems provided herein may use polypeptide that are immobilized on a support or substrate. As used herein the term “support” and “substrate” are used interchangeably and refers to a porous or non-porous solvent insoluble material on which polymers such as polypeptides are synthesized or immobilized. As used herein “porous” means that the material contains pores having substantially uniform diameters (for example in the nm range). Porous materials include paper, synthetic filters and the like. In such porous materials, the reaction may take place within the pores. The support can have any one of a number of shapes, such as pin, strip, plate, disk, rod, bends, cylindrical structure, particle, including bead, nanoparticle and the like. The support can have variable widths.
[0113] The support can be hydrophilic or capable of being rendered hydrophilic and includes inorganic powders such as silica, magnesium sulfate, and alumina; natural polymeric materials, particularly cellulosic materials and materials derived from cellulose, such as fiber containingAttorney Docket: 114203-3092papers, e.g, filter paper, chromatographic paper, etc.; synthetic or modified naturally occurring polymers, such as nitrocellulose, cellulose acetate, poly (vinyl chloride), polyacrylamide, cross linked dextran, agarose, polyacrylate, polyethylene, polypropylene, poly (4-methylbutene), polystyrene, polymethacrylate, poly(ethylene terephthalate), nylon, poly(vinyl butyrate), polyvinylidene difluoride (PVDF) membrane, glass, controlled pore glass, magnetic controlled pore glass, ceramics, metals, and the like etc.; either used by themselves or in conjunction with other materials.
[0114] In one embodiment, the solid support is a bead. Each bead may serve as a localized platform for transcription and translation reactions, containing immobilized DNA templates encoding the desired polypeptide. Following one or multiple rounds of synthesis, each bead generates a unique polypeptide or combinations of polypeptides, corresponding to its immobilized template(s). To ensure efficient identification and tracking, beads can be tagged with a unique sequence or identifier linked to the synthesized polypeptide. The beads may be made from a variety of materials including polystyrene, polyethylene, polyeththylene terephthalate, nylon, polypropylene, polylactic acid, or agarose. In an embodiment, the beads may be magnetic beads.Individual Discrete Volumes
[0115] The present disclosure relates to high throughput expression of proteins using in vitro transcription / translation. In one embodiment, the systems and methods disclosed herein use multiple individual discrete volumes.
[0116] In one embodiment, the individual discrete volumes are droplets, or the individual discrete volumes are defined on a solid support, or the individual discrete volumes are microwells, or the individual discrete volumes are spots defined on a solid support, such as a bead (e.g., a magnetic bead).
[0117] In one embodiment, the individual discrete volumes are minimized sample volumes, such as microvolumes, nanovolumes, picovolumes, or sub-picovolumes. Accordingly, in one embodiment, the systems and methods relate to amplification and / or assembly of polynucleotide sequences and of expression of proteins in small volume droplets on separate and addressable features of a support. In one embodiment, predefined reaction microvolumes of between about 0.5 pL and about 100 nL, may be used. However, smaller or larger volumes may be used. One would appreciate that the minimized sample volume increases the number of samples that can beAttorney Docket: 114203-3092processed in an efficient and parallel manner. Tn one embodiment, the transcription / translation reactions can be generated or performed within a vessel of minimized proportions. In one embodiment, the transcription / translation reaction can take place on a solid surface or support, such as an array.
[0118] An “individual discrete volume” is a discrete volume or discrete space, such as a container, receptacle, or other defined volume or space that can be defined by properties that prevent and / or inhibit migration of nucleic acids and reagents necessary to carry out the methods disclosed herein, for example a volume or space defined by physical properties such as walls, for example the walls of a well, tube, or a surface of a droplet, which may be impermeable or semipermeable, or as defined by other means such as chemical, diffusion rate limited, electromagnetic, or light illumination, or any combination thereof. By “diffusion rate limited” (for example diffusion defined volumes) is meant spaces that are only accessible to certain molecules or reactions because diffusion constraints effectively defining a space or volume as would be the case for two parallel laminar streams where diffusion will limit the migration of a molecule from one stream to the other. By “chemical” defined volume or space is meant spaces where only certain molecules can exist because of their chemical or molecular properties, such as size, where for example gel beads may exclude certain species from entering the beads but not others, such as by surface charge, matrix size or other physical property of the bead that can allow selection of species that may enter the interior of the bead. By “electro-magnetically” defined volume or space is meant spaces where the electro-magnetic properties of molecules or their supports such as charge or magnetic properties can be used to define certain regions in a space such as capturing magnetic particles within a magnetic field or directly on magnets. By “optically” defined volume is meant any region of space that may be defined by illuminating it with visible, ultraviolet, infrared, or other wavelengths of light such that only molecules within the defined space or volume may be labeled. Typically, a discrete volume will include a fluid medium, (for example, an aqueous solution, an oil, a buffer, and / or a media capable of supporting cell growth) suitable for labeling of the molecule with the indexable nucleic acid identifier under conditions that permit labeling.
[0119] Exemplary discrete volumes or spaces useful in the disclosed methods include, but are not limited to, droplets (for example, microfluidic droplets and / or emulsion droplets), hydrogel beads or other polymer structures (for example poly-ethylene glycol di-acrylate beads or agaroseAttorney Docket: 114203-3092beads), tissue slides (for example, fixed formalin paraffin embedded tissue slides with particular regions, volumes, or spaces defined by chemical, optical, or physical means), microscope slides with regions defined by depositing reagents in ordered arrays or random patterns, tubes (such as, centrifuge tubes, microcentrifuge tubes, test tubes, cuvettes, conical tubes, and the like), bottles (such as glass bottles, plastic bottles, ceramic bottles, Erlenmeyer flasks, scintillation vials and the like), wells (such as wells in a plate), plates, pipettes, or pipette tips among others. In certain example embodiments, the individual discrete volumes are the wells of a microplate. In certain example embodiments, the microplate is a 96 well, a 384 well, or a 1536 well microplate.Expression Reagents
[0120] The IVPS system may further comprise expression reagents to facilitate the expression of the polypeptide expression template. In general, the expression reagents include cell extracts that support the synthesis of proteins in vitro from purified mRNA transcripts or from mRNA transcribed from DNA during the in vitro synthesis reaction. Nucleic acid templates guide the production of the desired protein. The reagents may further comprise RNA polymerase, accessory, proteins, buffers, ribonucleotides, crowding agents (e g. PEG) to stabilize RNA / increase macromolecular interactions, regulatory elements (e g. IPTG, cAMP), amino acids, tRNAs, polyamines to stabilize RNA and regulate gene expression (e.g. spermidine and putrescine), folinic acid, reducing agents to stabilize proteins with free sulfhydryls (e.g., DTT), RNase inhibitors, protease inhibitors. The expression reagents may further include energy sources to fuel metabolism, such as acetyl phosphate, phosphoenolpyruvate (PEP), creatine phosphate, pyruvate, glucose-6-phosphate, fructose-l,6-bisphosphate, 3 -phosphoglycerate, glucose, maltose, maltodextrin, starch, and glutamate. The expression reagents may include cofactors supporting metabolic and enzymatic functions, such as NAD and CoA. Buffers may include HEPES or BisTris and contain ions to maintain ionic strength required for polypeptide activity such as KGlu, NEEClu, MgClu, oxalic acid, NH4C2H3O2, Mg (C2H3O2), and K2HPO4. Accessory proteins may include chaperones like Hsp70 and Hsp90, energy regeneration systems such as creatine kinase or pyruvate kinase to regenerate ATP, tRNA synthetases to improve translation efficiency, co-factors and post-translational modifiers (e.g., ferredoxin reductase, disulfide isomerases, peptidyl-prolyl isomerases, glycosyltransferases, hydroxylases, acetylases, myristoylases, methyltransferases,Attorney Docket: 114203-3092nitrosylases, and kinases), and ribosome stabilizing factors like ribosome recycling factor, or any combination thereof.
[0121] In some embodiments, the methods, and systems of the present disclosure utilize an optimized reaction buffer comprising: a buffering agent selected from HEPES or Bis-Tris present at 2-500 mM; ATP, GTP, CTP, and UTP each present at 0.01-20 mM; polyamines selected from spermidine and putrescine present at 0.05-30 mM total; an organic anion selected from glutamate and acetate salts present at 1-2000 mM; a divalent cation present at 0.4-90 mM; glycerol present at 0.5-50 % v / v; an RNase inhibitor and a protease inhibitor.
[0122] In some embodiments, the cell free protein synthesis module utilizes an optimized reaction buffer comprising: a buffering agent selected from HEPES or Bis-Tris present at 20- 50 mM; ATP, GTP, CTP, and UTP each present at 0.1-2mM; polyamines selected from spermidine and putrescine present at 0.5-3 mM total; an organic anion selected from glutamate and acetate salts present at 10-200 mM; a divalent cation present at 4-9 mM; glycerol present at 5-20 % v / v; an RNase inhibitor and a protease inhibitor.
[0123] In some embodiments, the buffer further comprises at least one of a halide salt, an amino acid, a tRNA, a tRNA synthetase, a chaperone, a crowding agent, an energy regeneration enzyme, or a reducing agent.
[0124] In some embodiments, the buffer operates at a reaction temperature between 1 °C and 95 °C. In some embodiments, the buffer operates at a reaction temperature between 4 °C and 42 °C. In some embodiments, the buffer operates at a reaction temperature between 25 °C and 40 °C.
[0125] In some embodiments, the buffer further comprises a halide salt in a concentration of 1-2000 mM. In some embodiments, the buffer further comprises a halide salt in a concentration of 10-200 mM. In some embodiments, the buffer further comprises a halide salt in a concentration of 20-150 mM. In some embodiments, the buffer further comprises a halide salt in a concentration of 40- 100 mM.
[0126] In some embodiments, the buffer further comprises a chaperone protein. In some embodiments, the chaperone protein is selected from Hsp70, Hsp90, GADD34, vaccinia virus K3L, and a disulfide isomerase.Attorney Docket: 114203-3092
[0127] In some embodiments, the buffer further comprises a crowding agent. In some embodiments, the crowding agent comprises polyethylene glycol, an alcohol, a sugar, an amino acid, or a polyol.
[0128] In some embodiments, the buffer further comprises a reducing agent. In some embodiments, the reducing agent comprises dithiothreitol, a solution of reduced and oxidized glutathione, or beta mercaptoethanol.
[0129] In some embodiments, the buffer further comprises at least one energy regeneration component. In some embodiments, the energy regeneration component comprises phosphoenolpyruvate, creatine phosphate, acetyl phosphate or glucose 6 phosphate.
[0130] In some embodiments, the RNase inhibitor is present at a final concentration between 0.01 U / pL and 0.5 U / pL and the protease inhibitor is present at a final concentration between 0.1 pg / mL and 3 pg / mL. Exemplary RNase inhibitors include, but are not limited to, RNase A, RNase B, and RNase T2. Exemplary protease inhibitors include, but are not limited to, aprotinin and leupeptin, 4-(2-Aminoethyl)benzene-l-sulfonyl fluoride (AEBSF), bestatin, N-[N-( -3-trans-carboxyirane-2-carbonyl)-L-leucyl]-agmatine (E-64), 2,2',2'',2'"-(Ethane-l,2-diyldinitrilo)tetraacetic acid (EDTA), pepstatin A, and Phenylmethanesulfonyl fluoride (PMSF). In one embodiment the protease inhibitor is aprotinin an / or leupeptin.
[0131] In one embodiment, polymerases used in the combined transcription / translation IVPS platform include any polymerase capable of supporting in vitro transcription within the IVPS platform extract and reaction. Examples of suitable polymerases include E. coli RNA Polymerase, T3 RNA Polymerase, T7 RNA Polymerase, and SP6 RNA Polymerase, among others.
[0132] The expression reagents, IVPS extract and translation buffer can be added separately, or two or more of these solutions can be combined before their addition, or added contemporaneously. To synthesize a protein of interest in vitro, an IVPS extract generally at some point comprises a mRNA molecule that encodes the protein of interest. In early IVPS experiments, mRNA was added exogenously after being purified from natural sources or prepared synthetically in vitro from cloned DNA using bacteriophage RNA polymerases. In other systems, the mRNA is produced in vitro from a template DNA; transcription and translation occur in this IVPS reaction. Techniques using coupled or complementary transcription and translation systems, which synthesize both RNA and protein in the same reaction, have been developed. In such in vitroAttorney Docket: 114203-3092transcription and translation (IVTT) systems, the IVPS extracts contain all the components necessary both for transcription (to produce mRNA) and for translation (to synthesize protein) in a single system. An early IVTT system was based on a bacterial extract (Lederman and Zubay, Biochim, Biophys. Acta, 1-49: 253, 1967). In IVTT systems, the input nucleic acid is DNA, which is usually much easier to obtain than mRNA and more readily manipulated (e.g., by cloning, sitespecific recombination, and the like).
[0133] The reaction buffer may include NTPs, spermidine, putrescine, glutamate salt, magnesium salt, and glycerol. In one embodiment, the reaction buffer comprises at least one component selected from NTPs: a polyamine, an organic anion, a divalent cation, an alcohol, and combinations thereof. In one embodiment, the polyamine is selected from spermidine and putrescine; the organic anion is selected from glutamate and acetate; the divalent cation is selected from magnesium, calcium, and manganese; and the alcohol includes glycerol.
[0134] In one embodiment, the reaction reagents can be tailored to the specific type of protein being expressed. For example, polypeptides may be expressed using bacterial cell lysates, such as E. coli, Leishmania tarentolae, Vibrio natriegens, Pseudomonas putida, Streptomyces, Bacillus subtillis, and archaeal cell lysates. Eukaryotic proteins can also be expressed with compatible systems, including wheat germ, tobacco, CHO, insect, Hela, yeast, reticulocyte, and erythrocyte cell lysate-based expression systems. Incubation of the reaction under appropriate reaction conditions, such as 37° C., can be followed by an analysis of protein expression.
[0135] In one embodiment, the presence of polypeptides of interest can be assessed by measuring protein activity. A variety of protein activities can be assayed. Non-limiting examples include binding activity (e.g., specificity, affinity, saturation, competition), enzyme activity (kinetics, substrate specificity, product, inhibition), etc. Exemplary methods include but are not limited to, spectrophotometric, fluorometric, calorimetric, chemiluminescent, light scattering, radiometric, and chromatographic methods.Purification Reagents
[0136] The system leverages a high-purity, high-throughput, and / or highly versatile process for polypeptide purification. By comparison, traditional methods for polypeptide purification are constrained using specialized polypeptide-specific tags and buffers, as well as inefficient, manual, column-based workflows that lack scalability.Attorney Docket: 114203-3092
[0137] As described above, the vector may encode a first purification tag and, optionally, a second purification tag. Accordingly, the system may further comprise purification reagents to facilitate the expressed polypeptide's purification and / or immobilization. The purification reagents may comprise capture agents that bind the first or second purification tag and a set of wash buffers.
[0138] For example, capture agent are the cognate binding parts of the first and second purification tags described above. For example, if the purification tag is strep-tag, then the capture reagent could be streptavidin; if the purification tag is streptavidin, the capture reagent could be biotin; if the purification tag is albumin binding protein, the capture reagent could be serum albumin; if the purification tag is Halo-tag the capture reagent could be haloalkanes; if the purification tag was a His-tag then the capture reagent could be a nickel / cobalt covered substrate. In an embodiment, the One of ordinary skill in the art can select the appropriate capture tag based on the initial selection(s) of the purification tag. In an embodiment, the first capture agent, the second capture agent, or both, is attached to a bead.
[0139] In one embodiment, the first purification tag is cleaved with a protease to yield a substantially pure polypeptide or polypeptide variant. In one embodiment, the protease is human rhinovirus (HRV) 3C protease, enteropeptidase, thrombin, factor Xa, protease tobacco etch virus (TEV), human rhinovirus 3C (HRV3C) protease, carboxypeptidase A, carboxypeptidase B, trypsin, chymotrypsin A4, thermolysin, dipeptidyl amino peptidase (DAPase), endoproteinase arg-C, endoproteinase glu-C (V8), endoproteinase lys-C, and endoproteinase asp-N.
[0140] In an embodiment, the first purification tag is cleaved by a protease that targets a cleavage site adjacent to the first purification tag, wherein the cleavage site adjacent to the first purification tag comprises one or more non-canonical amino acids comprising photocleavable bonds, chemically cleavable bonds, or environmentally cleavable bonds.
[0141] In another embodiment, an optional, second purification tag is introduced into the vector, which allows for higher purity polypeptide once cleaved. In one embodiment, the optional second purification tag is a histidine (His) tag. His-tagged polypeptides can be eluted in one of two ways: using imidazole, which retains the solubility tag on the polypeptides, or by employing a second orthogonal protease to simultaneously release the substantial pure polypeptide and remove the solubility tag.Attorney Docket: 114203-3092
[0142] As used herein, the term “substantially pure polypeptide” as used herein, refers to a composition that is at least about 75% by weight (including, for example, at least about 76% by weight, at least about 77% by weight, at least about 78% by weight, at least about 79% by weight, at least about 80% by weight, based on the total weight of the composition), at least about 85% by weight, at least about 86% by weight, at least about 87% by weight, at least about 88% by weight, at least about 89% by weight, least about 90% by weight (including, for example, at least about 91% by weight, at least about 92% by weight, at least about 93% by weight, at least about 94% by weight, at least about 90% by weight, based on the total weight of the composition), at least about 95% by weight, at least about 96% by weight, at least about 97% by weight, at least about 98% by weight, or at least about 99% by weight) of the pure polypeptide.
[0143] In one embodiment, the expression template comprises a second purification tag, and the method further includes capturing the proteins with a second capture agent that binds the second purification tag and washing the captured proteins with a second wash buffer. The term “wash buffer” is used herein to refer to the buffer passed over the solid support (with immobilized polypeptide) following loading and before elution of the polypeptide of interest. Wash buffers can remove impurities, such as host cell impurities, but not polypeptides of interest. The conductivity and / or pH of the wash buffer is / are such that the impurities are eluted from the polypeptide, but not any significant amounts of the polypeptide of interest. In one embodiment, the salt concentration of the wash buffers is high salt concentration wash buffers. In another embodiment, the salt concentration of the wash buffer is low salt concentration wash buffers. In another embodiment, the salt concentration of the wash buffers is an alternative between high salt concentration wash buffers and low salt concentration wash buffers.
[0144] In one embodiment, the wash buffer may include a detergent to increase protein yield. As used herein, the term “detergent” refers to amphipathic molecules containing both a non-polar “tail” with aliphatic or aromatic characteristics and a polar “head.” Exemplary detergents include, but are not limited to, octylphenoxypolyethoxy ethanol, sodium dodecyl sulfate (SDS), 2-[4-(2,4,4-trimethylpentan-2-yl)phenoxy]ethanol, nonyl phenoxypolyethoxylethanol, polyoxyethylene (20), polyoxyethylene (80), 3-((3-cholamidopropyl)dimethylammonio)-l -propanesulfonate, and 3-{dimethyl[3-(3a,7u,12a-trihydroxy-5P-cholan-24-amido)propyl]azaniumyl}propane-l -sulfonate.Attorney Docket: 114203-3092
[0145] In another embodiment, the wash buffer may include a blocking reagent to prevent protein losses due to non-specific adhesion. As used herein, the term “blocking reagent” refers to a highly concentrated protein that can be added to the wash buffer. These proteins function by outcompeting non-specific binding interactions, thereby preventing unintended interactions between proteins or other molecules and the surface of interest. Exemplary blocking reagents include, but are not limited to, bovine serum albumin (BSA), casein, non-fat dry milk, or commercially available formulations specifically designed to minimize background noise and enhance assay specificity.
[0146] In one embodiment, the second purification tag can be eluted via a protease that cleaves a protease cleavage site adjacent to the solubility tag.
[0147] The present disclosure also utilizes one or more capture agents to bind one or more capture tags. In one embodiment, the capture agents are each coated on an individual discrete volume.
[0148] In one embodiment, the capture agents are each coated on a support, e.g., a bead e.g., a magnetic bead, or an agarose bead. In one embodiment, the capture reagents are immobilized on a surface of the individual discrete volume. In another embodiment, the individual discrete volumes may be in the format of a filter plate, e.g., HIS-Select® Filter Plate from Millipore Sigma.
[0149] In one embodiment, a first capture agent can bind a first capture tag. In another embodiment, an optional second capture agent can bind an optional second capture tag.
[0150] In one embodiment, the first capture agent is attached to a microwell surface. In another, a photocleavable linker attaches it to the microwell surface. Alternatively, a different photocleavable linker may be used in different wells to allow the protein's spatiotemporal release.
[0151] In one embodiment, the optional second capture agent is attached to a solid support. In one embodiment, the solid support is ahead. In one embodiment, the bead is magnetic or agarose.
[0152] In one embodiment, the systems disclosed herein further comprise an RNase inhibitor and / or a protease inhibitor. Exemplary RNase inhibitors include, but are not limited to, RNase A, RNase B, and RNase T2. Exemplary protease inhibitors include, but are not limited to, aprotinin and leupeptin, 4-(2-Aminoethyl)benzene-l-sulfonyl fluoride (AEBSF), bestatin, N-[N-( -3-trans-carboxyirane-2-carbonyl)-L-leucyl]-agmatine (E-64), 2,2',2'',2'"-(Ethane-l,2-Attorney Docket: 114203-3092diyldinitrilo)tetraacetic acid (EDTA), pepstatin A, and Phenylmethanesulfonyl fluoride (PMSF). In one embodiment the protease inhibitor is aprotinin an / or leupeptin.
[0153] In another embodiment, the systems disclosed herein comprise one or more proteases capable of cleaving one or more protease cleavage site(s). Exemplary proteases include but are not limited to, enteropeptidase, thrombin, factor Xa, protease tobacco etch virus (TEV), human rhinovirus 3C (HRV3C) protease, carboxypeptidase A, carboxypeptidase B, trypsin, chymotrypsin A4, thermolysin, dipeptidyl amino peptidase (DAPase), endoproteinase arg-C, endoproteinase glu-C (V8), endoproteinase lys-C, and endoproteinase asp-N, aprotinin, and leupeptin. In one embodiment, the protease inhibitor is aprotinin. In another embodiment, the protease inhibitor is leupeptin.METHODS OF POLYPEPTIDE SYNTHESIS AND SCREENING
[0154] Also disclosed herein are methods of highly parallel in vitro polypeptide synthesis. In one embodiment, the method comprises allocating expression templates from the expression library described above into individual discrete volumes, wherein each discrete volume includes a different expression template. In one embodiment, the individual discrete volumes comprise minimized sample volumes, such as micro, nano, pico, or sub-pico volumes. In an embodiment, predefined reaction micro volumes between about 0.1 pL and about 100 nL may be used. However, smaller or larger volumes may be used. One would appreciate that the minimized sample volume increases the number of samples that can be processed efficiently and in parallel.
[0155] In one embodiment, the IVPS reaction utilizes a microwell format to facilitate parallel polypeptide expression. Each microwell contains a distinct expression template encoding a polypeptide with defined characteristics, including N-terminal translation enhancers, solubility tags, and purification tags. The reaction environment is optimized using a cell lysate supplemented with RNA polymerase and protease inhibitors to ensure high-yield polypeptide synthesis.
[0156] In an embodiment, the expression templates from the expression library are allocated to different wells of a 96-well, 384-well, or 1536-well microplate. In another embodiment, the individual expression templates are encapsulated in droplets on a microfluidic device.Polypeptide Expression and Purification
[0157] The polypeptides encoded by the expression templates in each individual discrete volume are then expressed in parallel in the presence of expression reagents. In general, expressionAttorney Docket: 114203-3092is carried out in the presence of a cell extract that provides the translation machinery and an expression buffer (e.g., having appropriate salts, detergents, and pH), an RNA polymerase that recognizes the promoter(s) to which the polypeptide is operably linked and, optionally, one or more transcription factors directed to an optional regulatory sequence to which the template nucleic acid is operably linked; ribonucleotide triphosphates (rNTPs); ribosomes; transfer RNA (tRNA); optionally, other transcription factors and co-factors therefor; amino acids (optionally comprising one or more detectably labeled amino acids); one or more energy sources, (e.g., ATP, GTP); and other or optional translation factors (e.g., translation initiation, elongation and termination factors) and co-factors therefor.
[0158] The temperature may be any temperature suitable for IVTT. Temperature may be in the general range from about 10° C to about 40° C, including intermediate specific ranges within this general range, include from about 15° C. to about 35° C., form about 15° C. to about 30° C., form about 15° C to about 25° C. In one embodiment, the reaction temperature can be about 15° C, about 16° C, about 17° C, about 18° C, about 19° C, about 20° C, about 21° C, about 22° C, about 23° C, about 24° C, about 25° C. In one embodiment, the reaction temperature can be about 21° C.
[0159] The IVPS reaction can include any organic anion suitable for IVTT. In one embodiment, the organic anions can be glutamate, acetate, among others. In one embodiment, the concentration for the organic anions is independently in the general range from about 0 mM to about 200 mM, including intermediate specific values within this general range, such as about 0 mM, about 10 mM, about 20 mM, about 30 mM, about 40 mM, about 50 mM, about 60 mM, about 70 mM, about 80 mM, about 90 mM, about 100 mM, about 110 mM, about 120 mM, about 130 mM, about 140 mM, about 150 mM, about 160 mM, about 170 mM, about 180 mM, about 190 mM and about 200 mM, among others.
[0160] The IVPS reaction can also include any halide anion suitable for IVTT. In one embodiment, the halide anion can be chloride, bromide, iodide, among others. In some embodiments, the halide anion is chloride. Generally, the concentration of halide anions, if present in the reaction, is within the general range from about 0 mM to about 200 mM, including intermediate specific values within this general range, such as those disclosed for organic anions generally herein.Attorney Docket: 114203-3092
[0161] The IVPS reaction may also include any organic cation suitable for IVTT. In one embodiment, the organic cation can be a polyamine, such as spermidine or putrescine, among others. In some embodiments, polyamines are present in the CFPS reaction. In one embodiment, the concentration of organic cations in the reaction can be in the general about 0 mM to about 3 mM, about 0.5 mM to about 2.5 mM, about 1 mM to about 2 mM. In one embodiment, more than one organic cation can be present.
[0162] The IVPS reaction can include any inorganic cation suitable for IVTT. For example, suitable inorganic cations can include monovalent cations, such as sodium, potassium, lithium, among others; and divalent cations, such as magnesium, calcium, manganese, among others. In one embodiment, the inorganic cation is magnesium. In one embodiment, the magnesium concentration can be within the general range from about 1 mM to about 50 mM, including intermediate specific values within this general range, such as about 1 mM, about 2 mM, about 3 mM, about 5 mM, about 6 mM, about 7 mM, about 8 mM, about 9 mM, about 10 mM, among others. In some aspects, the concentration of inorganic cations can be within the specific range from about 4 mM to about 9 mM or alternatively, within the range from about 5 mM to about 7 mM.
[0163] The IVPS reaction includes NTPs. In one embodiment, the reaction uses ATP, GTP, CTP, and UTP. In one embodiment, the concentration of individual NTPs is within the range from about 0.1 mM to about 2 mM.
[0164] The IVPS reaction can also include any alcohol suitable for IVTT. In one embodiment, the alcohol may be a polyol, and more specifically glycerol. In one embodiment, the alcohol is between the general range from about 0% (v / v) to about 25% (v / v), including specific intermediate values of about 5% (v / v), about 10% (v / v) and about 15% (v / v), and about 20% (v / v), among other.
[0165] In some embodiments, the IVPS reaction includes glutamate salts, NTPs, spermidine, putrescine, glycerol, and magnesium.
[0166] The polypeptide can be expression of from the expression template in the individual discrete volume in the presence of a cell lysate and a reaction buffer. In one embodiment, the reaction reagents can be tailored to the specific type of protein being expressed. For example, prokaryotic proteins can be expressed using bacterial reagents, such as an A. cell lysate based expression systems, e.g., the PureExpress® system from New England Biolabs. EukaryoticAttorney Docket: 114203-3092proteins can also be expressed with compatible systems, including wheat germ (e g., T7 coupled reticulocyte lysate TNT™ system (Promega, Madison, Wis.)) and erythrocyte lysate based expression systems. Incubation of the reaction under appropriate reaction conditions, such as 37° C, can be followed by an analysis of protein expression. In one embodiment, synthesized protein products can be separated using gel electrophoresis and visualized or imaged by staining or Western blot.
[0167] In one embodiment, the expressed polypeptide is captured within an individual discrete volume using a first capture agent. The capture process is designed to ensure high purity by incorporating multiple sequential washing steps that effectively remove non-specific binding interactions. The solid support, such as magnetic beads, undergoes at least two, e.g., five, sequential washes. These washes alternate between high-salt and low-salt buffer conditions, as detailed in Example 1. This alternating buffer strategy disrupts and eliminates weakly bound contaminants, including non-specific proteins and residual impurities, while maintaining the structural integrity and functional activity of the target polypeptide. By employing this rigorous washing protocol, the purity and yield of the captured polypeptide are significantly enhanced, making it suitable for downstream applications requiring high-quality biomolecules. In an embodiment, the target polypeptide may be captured with a second capture agent that binds the second purification tag to achieve greater purity. Capture of the target polypeptide with the second capture agent comprises washing with a second wash buffer. For example, in one embodiment, the target polypeptide has an N-terminal His-tag and then washed with a second wash buffer.
[0168] The term “pure,” as used herein, refers to materials that have only the target polypeptide such that the presence of irrelevant polypeptides ( / .<?., impurities or contaminants, including polypeptide fragments) is reduced or eliminated. For example, the purified polypeptide sample comprises one or more target polypeptides but is substantially free of other polypeptides detectable by the method. As used herein, the term “substantially free” is operatively used in the context of analytical testing of a material. In some embodiments, the purification material is substantially free of one or more impurities or contaminants including polypeptide fragments, and has a purity of at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, or 97%. In some embodiments, the purification material is substantially free of one or more impurities or contaminants including polypeptideAttorney Docket: 114203-3092fragments, and has a purity of at least 98%, 99%, or more. In an embodiment, the pure polypeptide sample consists of 100% target polypeptide, and no other polypeptides are included.
[0169] Following the washing steps, a protease is introduced to cleave a protease cleavage site adjacent site to the solubility tag. Exemplary proteases include, but are not limited to, enteropeptidase, thrombin, factor Xa, protease tobacco etch virus (TEV), human rhinovirus 3C (HRV3C) protease, carboxypeptidase A, carboxypeptidase B, trypsin, chymotrypsin A4, thermolysin, dipeptidyl amino peptidase (DAPase), endoproteinase arg-C, endoproteinase glu-C (V8), endoproteinase lys-C, and endoproteinase asp-N.
[0170] For spatiotemporal polypeptide release, a photocleavable linker may be attached to capture agents in individual wells. These linkers enable the controlled release of polypeptides at desired time intervals, allowing for temporal studies of polypeptides’ function or sequential reactions within the same microwell array. For example, polypeptides synthesized and immobilized on the microwell surface can be selectively released into solution for further analysis or interaction assays.
[0171] In one embodiment, the ORF encodes only a mature protein or polypeptide form. For example, the mature protein or polypeptide form may exclude one or more of signal peptides; intracellular domains of transmembrane proteins, receptors, intra-protein linkers for proteins; polypeptides, or any amino acids associated with post-translational proteolytic processing.Screening of Polypeptide Expression
[0172] In an embodiment, the method may further comprising screening the expressed polypeptides for different functions. In one embodiment, the method further comprises culturing cells in individual discrete volumes and measuring a phenotypic response of the cell to each expressed polypeptide. The cell may be a bacterial cell, a fungal cell, or a eukaryotic cell. The eukaryotic cell may be any eukaryotic cell type including protozoal cells, plant cells, mammalian cells, and human cells. In one embodiment a virus, or a cell already inoculated with a virus, may be further cultured in the individual discrete well. In one embodiment, the cell may be further cultured in the presence of a therapeutic agent.
[0173] A phenotypic response of the cell to the expressed protein, and any other perturbation added to the culture, may be measured using, for example, RNA-seq, single-cell RNA seq, spatialAttorney Docket: 114203-3092transcriptomics, flow cytometry, microscopy imaging, transwell migration, immunoassay, proliferation assay, co-culture, or a combination thereof.
[0174] In one embodiment, RNA sequencing (RNA-seq) methodology is employed to quantify transcriptome-wide changes resulting from protein expression. The method comprises preparing RNA libraries from cells exposed to the synthesized polypeptides and performing high-throughput sequencing, and analyzing differential gene expression patterns. This approach enables comprehensive characterization of transcriptional responses to protein expression, including identification of downstream effectors and regulatory networks. Detection tags may be incorporated to support the integration of phenotypic assays within the microwells, enabling direct observation of cellular responses to expressed polypeptides in the same modular platform. A further embodiment utilizes single-cell RNA sequencing to resolve cellular heterogeneity in response to polypeptide exposure or screening. The method comprises isolating single cells from the individual discrete volumes, generating barcoded cDNA libraries, and performing massively parallel sequencing. This technique reveals cell-specific transcriptional responses and identifies distinct cellular subpopulations with varying polypeptide expression responses.
[0175] In another embodiment, spatial transcriptomics may be used to preserve spatial information while analyzing gene expression changes. The method comprises preparing tissue sections or cell cultures, capturing spatially-resolved RNA, and performing sequencing with spatial barcoding. This approach enables the mapping of transcriptional responses relative to polypeptide exposure / screening localities, revealing the spatial organization of cellular responses.
[0176] A further embodiment employs flow cytometry for high-throughput single-cell analysis. The method comprises labeling cells with fluorescent markers, analyzing individual cells through a fluidic system, and quantifying protein expression levels and cellular parameters. This technique enables rapid screening of large cell populations and quantifying polypeptide expression effects on cellular phenotypes. The detection tags disclosed above may be the fluorescent marker for flow cytometry.
[0177] In another embodiment, microscopy imaging techniques are utilized for visual analysis of cellular responses. The method comprises preparing cells for imaging, capturing multi-channel images, and performing automated image analysis. This approach provides detailed informationAttorney Docket: 114203-3092about morphological changes, protein localization, and cellular organization following polypeptide expression.
[0178] A further embodiment employs trans-well migration or invasion assays to assess effects on cellular motility. The method comprises seeding cells in an upper chamber of the migration assay device, allowing migration through permeable membranes and quantifying migrated cells. This technique evaluates how polypeptide expression influences cell migration and chemotactic responses. Accordingly, in an embodiment employing a trans-well migration assay, the individual discrete volumes may be arranged on a microwell plate with cell migration or cell invasion insert that can be placed within each well, providing the aforementioned upper chamber and permeable membrane.
[0179] In another embodiment, immunoassay methodology enables specific protein quantification. The method comprises preparing cellular samples, performing antibody-based detection, and measuring protein levels or modifications. This approach provides precise quantification of polypeptide expression that, when combined with an RNA or other cellular, provides a readout of downstream molecular effects. In one embodiment, the immunoassay is an ELISA-based assay.
[0180] A further embodiment utilizes proliferation assays to measure effects on cell growth. The method comprises monitoring cell numbers over time through various techniques, including metabolic assays and direct counting. This approach quantifies how polypeptide expression influences cell proliferation and survival.
[0181] In another embodiment, co-culture systems enable analysis of intercellular effects. The method combines two or more cell types in the individual discrete volume and measures how polypeptide expression influences cell-cell interactions or communication. In one embodiment, the individual discrete volume comprises an immune cell and a diseased cell type and measures the amount of cell killing by the immune cell of the diseased cell type or measures changes in cell signaling by the immune cell in the presence of the expressed polypeptide and the diseased cell type. In one embodiment, the diseased cell type is a cancer cell.
[0182] These embodiments may be employed individually or in combination to characterize cellular responses to protein expression comprehensively. The methods described herein enable aAttorney Docket: 114203-3092multi -parameter analysis of protein expression effects across various biological scales and cellular processes.Exemplary Applications
[0183] Combining the in vitro polypeptide expression systems and phenotypic readouts provides versatile applications across multiple biotechnology and pharmaceutical development domains. The following examples illustrate representative applications but should not be construed as limiting the scope of the invention.
[0184] In therapeutic target identification, the invention enables systematic characterization of potential drug targets. For example, when investigating a polypeptide implicated in cancer progression, the RNA-seq embodiment can reveal genome-wide transcriptional changes upon polypeptide expression, identifying previously unknown causal drivers of disease outcomes or downstream pathways in either the cancer cell or in a co-cultured cell, such as an immune cell. The single-cell RNA sequencing embodiment further resolves cell-type specific responses, potentially showing that the polypeptide drives proliferation specifically, for example, in cancer stem cells but not in differentiated tumor cells. The spatial transcriptomics embodiment might demonstrate that polypeptide exposure / screening in one cell type induces protective responses in neighboring cells, suggesting potential resistance mechanisms.
[0185] For drug screening applications, the present disclosure provides a comprehensive assessment of compound effects. Consider a screen of polypeptides targeting a specific receptor or enzyme. The flow cytometry embodiment enables rapid quantification of target engagement across thousands of cells, while the microscopy embodiment reveals the subcellular distribution of the expressed polypeptide and morphological changes induced by inhibitor binding. The transwell migration embodiment can demonstrate that specific inhibitors block cancer cell invasion, while co-culture experiments might reveal that only some compounds effectively prevent cancer cell-induced changes in, for example, stromal cells.
[0186] In protein engineering applications, the present disclosure facilitates iterative optimization of protein properties. For instance, when engineering chimeric antigen receptors (CARs) for immunotherapy, the immunoassay embodiment can quantify expression levels of different CAR variants, while the proliferation assay embodiment measures their effects on T cell expansion. Single-cell RNA sequencing might reveal that certain CAR designs induce optimal T-Attorney Docket: 114203-3092cell activation while others trigger exhaustion programs. Co-culture experiments with target cells can confirm CAR-specific cytotoxicity while revealing potential off-target effects.
[0187] The present disclosure also enables the discovery of novel protein-protein interactions. For example, when expressing a library of modified polypeptides, the microscopy embodiment can identify variants that co-localize with targets of interest. RNA-seq can reveal transcriptional signatures indicating successful pathway modulation, while immunoassays confirm physical interaction. Flow cytometry enables rapid screening of libraries for desired binding properties.
[0188] In the context of cellular differentiation protocols, the present disclosure can optimize polypeptide expression timing and levels. Spatial transcriptomics might reveal that certain morphogen combinations create optimal gradient patterns fortissue organization. Single-cell RNA sequencing can track differentiation trajectories, while immunoassays monitor key marker proteins. Co-culture systems can evaluate the capacity of engineered cells to integrate into tissuelike structures.
[0189] For studying disease mechanisms, the present disclosure enables systematic perturbation of disease-associated polypeptides. Expression of disease variants can be assessed through RNA-seq to identify causal drivers and dysregulated pathways, while single-cell analysis reveals effects on specific cell populations. Migration assays might uncover novel aspects of pathological cell behavior, while co-culture systems model disease effects on tissue organization.
[0190] The present disclosure further facilitates the development of protein-based therapeutics. Flow cytometry can rapidly screen for variants with desired binding properties when optimizing therapeutic antibodies. Microscopy can confirm appropriate subcellular targeting, while RNA-seq reveals downstream pathway modulation. Co-culture systems can evaluate therapeutic efficacy in complex cellular environments.
[0191] These examples demonstrate the present disclosure's versatility across multiple biological research and therapeutic development applications. Combining different embodiments provides insight into protein function and cellular responses, enabling more effective development of therapeutic strategies and protein-based technologies.
[0192] Further embodiments are illustrated in the following Examples which are given for illustrative purposes only and are not intended to limit the scope of the present disclosure.Attorney Docket: 114203-3092EXAMPLESExample 1 -Protein Production, HaloTag Bead Pulldown, and Elution
[0193] Protein production begins by preparing HeLa IVTT reactions. HeLa lysate aliquots, which must not exceed three freeze-thaw cycles, are combined with accessory proteins, reaction mix, and inhibitors to a final reaction volume of 5 pL. Specifically, the reaction consists of 2.5 pL HeLa lysate, 0.5 pL accessory proteins, and 1 pL reaction mix. Protector RNase inhibitor is added to a final concentration of 0.2 U / pL, and aprotinin and leupeptin protease inhibitors are each included at a final concentration of 1 pg / mL. For a no-DNA control, 1 pL of water is used, while 1 pL of plasmid DNA (from colonies yielding maximum DNA) is added for test proteins. All reagents are thawed on ice, with unused portions immediately returned to -80°C storage. HeLa lysate and accessory proteins are combined in a protein-free tube, placed on ice for 10 minutes, and mixed with the remaining reagents. The reaction is incubated at 30°C for 6 hours. The reactions are stored on ice for immediate use; long-term storage is frozen at -20°C or colder.
[0194] HaloTag bead pulldown is then performed. Magnetic beads are equilibrated using PBS containing 0.005% IGEPAL and 0.025% BSA as the buffer. Beads are resuspended thoroughly before use, with 7.5 pL of beads allocated per 50 pL of IVTT reaction. The beads are washed four times by capturing them on a magnetic stand for 30 seconds, removing the supernatant, adding equilibration buffer, and mixing for 5 minutes. After the final wash, the beads are resuspended in an equilibration buffer volume equal to the IVTT reaction volume. The beads are then added to the IVTT reaction containing the bait protein in a 1:1 volume ratio and mixed thoroughly. The mixture is incubated overnight at 4°C using an end-over-end rotator set to 1-2 rpm.
[0195] Next, the beads undergo five sequential washes to remove non-specific bindings. The washing buffers alternate between high-salt (50 mM HEPES, 500 mM NaCl, 0.005% IGEPAL, 0.025% BSA) and low-salt (50 mM HEPES, 150 mM NaCl, 0.005% IGEPAL, 0.025% BSA) conditions, with each wash lasting 3 minutes at room temperature. 50mM HEPES may also be replaced with 20 mM Tris-HCl (pH 7.4). Following the washes, HRV3C protease is added in a buffer containing 20 mM HEPES (pH 7.5), 500 mM NaCl, 0.005% IGEPAL, and 0.025% BSA. The mixture is incubated for 8 hours at 4°C to cleave the target protein from the beads.Attorney Docket: 114203-3092
[0196] Finally, on Day 2 (Evening), the cleaved proteins are eluted into new, pre-blocked wells. These eluted samples are now ready for downstream analysis or storage, completing the protein production and purification workflow. This method ensures high efficiency, reproducibility, and scalability for protein expression and isolation.Additional Embodiments
[0197] (Al) An expression library comprising a set of expression templates, each expression template encoding a different polypeptide and each vector further comprising polynucleotide sequences encoding one or more of:a. an N-terminal translation enhancer sequence;b. a translation initiation sequence;c. a first purification tag;d. a solubility tag;e. an open reading frame (ORF) encoding a polypeptide and further comprising a 3' UTR sequence, an engineered poly-A tail, or bothf. an optional detection tag; org. an optional second purification tag.
[0198] (A2) For the expression library denoted as (Al), wherein the N-terminal translation enhancer sequence is selected from chloramphenicol acetyltransferase (CAT), E01, UTR 17, ARC-1, and omega sequence of tobacco mosaic virus (SEQ ID NO: 1).
[0199] (A3) For the expression library denoted as (Al) or (A2), wherein the translation initiation sequence is a Kozak sequence.
[0200] (A4) For the expression library denoted as (A3), wherein the Kozak sequence is selected from GATGATAATATGG (SEQ ID NO: 2) and GCCGCCACCATGG (SEQ ID NO: 3).
[0201] (A5) For the expression library denoted as any one of (A1)-(A4), wherein the first purification tag is selected from an haloalkane dehalogenase-derived tag, Albumin-binding protein (ABP), Alkaline Phosphatase (AP), AU1 epitope, AU5 epitope, Bacteriophage T7 epitope (T7-tag), Bacteriophage V5 epitope (V5-tag), Biotin-carboxy carrier protein (BCCP), Bluetongue virus tag (B-tag), Calmodulin binding peptide (CBP), Chloramphenicol Acetyl Transferase (CAT), Cellulose binding domain (CBP), Chitin binding domain (CBD), Choline-binding domain (CBD), Dihydrofolate reductase (DHFR), E2 epitope, FLAG epitope, Galactose-binding protein (GBP),Attorney Docket: 114203-3092Green fluorescent protein (GFP), Glu-Glu (EE-tag), Glutathione S-transferase (GST), Human influenza hemagglutinin (HA), Histidine affinity tag (HAT), Horseradish Peroxidase (HRP), HSV epitope, Ketosteroid isomerase (KSI), KT3 epitope, LacZ, Luciferase, Maltose-binding protein (MBP), Myc epitope, NusA, PDZ domain, PDZ ligand, Polyarginine (Arg-tag), Polyaspartate (Asp-tag), Polycysteine (Cys-tag), Polyhistidine (His-tag), Polyphenylalanine (Phe-tag), Profmity eXact, Protein C, SI -tag, S-tag, Streptavadin-binding peptide (SBP), Staphylococcal protein A (Protein A), Staphylococcal protein G (Protein G), Strep-tag, Streptavadin, Small Ubiquitin-like Modifier (SUMO), Tandem Affinity Purification (TAP), T7 epitope, Thioredoxin (Trx), TrpE, Ubiquitin, Universal, and VSV-G.
[0202] (A6) For the expression library denoted as any one of (A1)-(A5), wherein the solubility tag is selected from a small ubiquitin-like modifier (SUMO) tag, SUMOstar protease, a glutathione S-transferase (GST)-tag, a maltose-binding protein (MBP)-tag, a thioredoxin (Trx)-tag, RRGRRGRRGRRGRRG (SEQ ID NO: 4) tag, a NusA tag, an N-terminal extension (NEXT) tag, and Bl domain of Streptococcal protein G (GB1).
[0203] (A7) For the expression library denoted as any one of (A1)-(A6), wherein the vector further comprises a protease cleavage site after the solubility tag.
[0204] (A8) For the expression library denoted as (A7), wherein the solubility tag is selected from SUMOstar protease, HRV3C protease, and HaloTEV.
[0205] (A9) For the expression library denoted as any one of (A1)-(A8), wherein the detection tag is a fluorescent tag.
[0206] (A10) For the expression library denoted as any one of (A1)-(A9), wherein the second purification is a histidine tag.
[0207] (All) For the expression library denoted as any one of (Al)-(A10), wherein each template further comprises a protease cleavage site adjacent to the second purification tag.
[0208] (A 12) For the expression library denoted as any one of (Al)-(Al 1), wherein the ORF encodes only domains included in a mature polypeptide form.
[0209] (A13) For the expression library denoted as (A12), wherein the mature polypeptide form excludes one or more of signal peptides; intracellular domains of transmembrane polypeptides or receptors; intra-polypeptides linkers for polypeptides; or any amino acids associated with post-translational proteolytic processing.Attorney Docket: 114203-3092
[0210] (Al 4) For the expression library denoted as any one of (A1)-(A13), wherein the ORF is codon optimized.
[0211] (Bl) A high-throughput in vitro polypeptide synthesis system comprising:the expression library denoted as any one of (Al) to (A 14);a plurality of individual discrete volumes, each volume comprising a different expression template from the expression library;a cell lysate;an RNA polymerase;a reaction buffer;a first capture agent for binding a first capture tag and an optional second capture agent for binding an optional second capture tag; andone or more purification buffers.
[0212] (B2) For the system denoted as (Bl), wherein the plurality of individual discrete volumes is defined in a microwell format.
[0213] (B3) For the system denoted as (Bl), wherein the plurality of individual discrete volumes are droplets.
[0214] (B4) For the system denoted as (B3), wherein the droplets are formed on a microfluidic device.
[0215] (B5) For the system denoted as (B3), wherein the droplets are formed on a glass surface.
[0216] (B6) For the system denoted as any one of (B2)-(B5), wherein the first capture agent is attached to a microwell surface.
[0217] (B7) For the system denoted as (B6), wherein the first capture agent is attached to the microwell surface by a photocleavable linker.
[0218] (B8) For the system denoted as any one of (B2)-(B7), wherein a different photocleavable linker may be used in different wells to allow for spatiotemporal release of the polypeptide.
[0219] (B9) For the system denoted as any one of (B 1)-(B8), wherein the first capture agent is attached to a bead.
[0220] (B10) For the system denoted as any one of (B1)-(B9), wherein the optional second capture agent is attached to beads.Attorney Docket: 114203-3092
[0221] (Bl 1) For the system denoted as any one of (Bl) to (BIO), further comprising a RNase inhibitor and a protease inhibitor.
[0222] (B 12) For the system denoted as (B 11), wherein the RNase inhibitor inactivates one or more of RNase A, RNase B, or RNase T2.
[0223] (B13) For the system denoted as (Bll), wherein the protease inhibitor is aprotinin and / or leupeptin.
[0224] (B14) For the system denoted as any one of (Bl) to (B13), further comprising one or more proteases capable of cleaving one or more protease cleavage site(s).
[0225] (Cl) A method of highly parallel in vitro polypeptide synthesis comprising:a. allocating expression templates from an expression library into individual discrete volumes, each individual discrete volume comprising a different transcription template, each transcription template comprising a polynucleotide sequence encoding a different polypeptide and further comprisingi. an N-terminal translation enhancer sequence;ii. a translation initiation sequence;iii. a first purification tag;iv. a solubility tagv. an ORF encoding the polypeptide;vi. a 3' UTR sequence, an engineered poly-A tail, or both; andvii. an optional detection tag;viii. an optional second purification tag; andb. expressing the polypeptide from the expression template in the individual discrete volume in the presence of a cell lysate and a reaction buffer wherein the expressed polypeptide is captured in the individual discrete volume by a first capture agent;c. washing the captured polypeptide with a wash buffer;d. for polypeptides that are not surface polypeptides, releasing the polypeptide into solution via a protease that cleaves a protease cleavage site adjacent to the first purification tag.
[0226] (C2) For the method denoted as (Cl), wherein the N-terminal translation enhancer sequence is selected from chloramphenicol acetyltransferase (CAT), E01, UTR 17, ARC-1, and omega sequence of tobacco mosaic virus (SEQ ID NO: 1).Attorney Docket: 114203-3092
[0227] (C3) For the method denoted as (Cl) or (C2), wherein the translation initiation sequence is a Kozak sequence.
[0228] (C4) For the method denoted as (C3), wherein the Kozak sequence is selected from GATGATAATATGG (SEQ ID NO: 2) and GCCGCCACCATGG (SEQ ID NO: 3).
[0229] (C5) For the method denoted as any one of (C1)-(C4), wherein the first purification tag is an haloalkane dehalogenase-derived tag.
[0230] (C6) For the method denoted as any one of (C1)-(C5), wherein the solubility tag is a small ubiquitin-like modifier (SUMO) tag, SUMOstar protease, a glutathione S-transferase (GST)-tag, a maltose-binding protein (MBP)-tag, a thioredoxin (Trx)-tag, or a NusA tag, , an N-terminal extension (NEXT) tag, and Bl domain of Streptococcal protein G (GB1).
[0231] (C7) For the method denoted as any one of (C1)-(C6), wherein the template further comprises a protease cleavage site after the solubility tag.
[0232] (C8) For the method denoted as (C7), wherein the solubility tag is selected from SUMOstar protease, HRV3C protease, and HaloTEV.
[0233] (C9) For the method denoted as any one of (C1)-(C8), wherein the detection tag is a fluorescent tag.
[0234] (CIO) For the method denoted as any one of (C1)-(C9), wherein the second purification tag is a histidine tag.
[0235] (C 11 ) For the method denoted as any one of (C 1 )-(C 10), wherein each template further comprises a protease cleavage site adjacent to the second purification tag.
[0236] (Cl 2) For the method denoted as any one of (Cl)-(Cl 1), wherein the ORF encodes only domains included in a mature polypeptide form.
[0237] (C 13 ) For the method denoted as (C 12), wherein the mature polypepti de form excludes one or more of signal peptides; intracellular domains of transmembrane polypeptides or receptors; intra-polypeptide linkers for polypeptide; or any amino acids associated with post-translational proteolytic processing.
[0238] (Cl 4) For the method denoted as any one of (C1)-(C13), wherein the ORF is codon optimized.
[0239] (Cl 5) For the method denoted as any one of (C1)-(C14), wherein the expression template comprises a second purification tag, and the method further comprises after step (c)Attorney Docket: 114203-3092capturing the polypeptide with a second capture agent that binds the second purification tag and washing the captured polypeptide with a second wash buffer.
[0240] (C16) For the method denoted as any one of (C 1)-(C 15), further comprising removing the second purification tag via a protease that cleaves a protease cleavage site adjacent to the solubility tag.
[0241] (Cl 7) For the method denoted as any one of (Cl) to (Cl 5), further comprising culturing cells in each individual discrete volume and measuring a phenotypic response to the polypeptide expressed in each individual discrete volume.
[0242] (Cl 8) For the method denoted as (Cl 7), wherein the cells are selected from bacterial cells, yeast cells, fungi cells, and eukaryotic cells.
[0243] (C19) For the method denoted as (C 17), wherein the phenotypic response is measured by RNA-seq, single-cell RNA seq, spatial transcriptomics, flow cytometry, microscopy imaging, transwell migration, immunoassay, proliferation assay, co-culture or a combination thereof.
[0244] (C20) For the method denoted as any one of (Cl) to (Cl 9), wherein a plurality of individual discrete volumes is defined in a microwell format.
[0245] (DI) An expression template for cell-free polypeptide synthesis comprising:a promoter operably linked to a polynucleotide sequence encoding a polypeptide;an N-terminal translation enhancer sequence positioned upstream of an open reading frame (ORF) encoding the polypeptide;a first purification tag positioned N-terminal to the ORF; anda 3' untranslated region (3'-UTR) sequence, an engineered poly-A tail, or both, positioned downstream of the ORF.
[0246] (D2) For the expression template denoted as (DI), wherein the expression template is DNA-based and selected from circular double-stranded DNA, linear double-stranded DNA, or branched DNA.
[0247] (D3) For the expression template denoted as (DI), wherein the expression template is RNA-based and comprises messenger RNA (mRNA).
[0248] (D4) For the expression template denoted as any one of (D1)-(D3), wherein the promoter is selected from T7, T3, SP6, and tac promoters.Attorney Docket: 114203-3092
[0249] (D5) For the expression template denoted as any one of (DI )-(D4), further comprising a solubility tag positioned between the first purification tag and the ORF.
[0250] (D6) For the expression template denoted as (D5), further comprising a protease cleavage site positioned between the solubility tag and the ORF.
[0251] (D7) For the expression template denoted as any one of (D1)-(D6), wherein the ORF is codon optimized for expression in mammalian cells.
[0252] (El) A kit for high-throughput cell-free polypeptide synthesis comprising:an expression library according to any one of (A1)-(A14);a cell lysate suitable for in vitro transcription and translation;a reaction buffer comprising ribonucleotide triphosphates (rNTPs), polyamines, and divalent cations;capture agents for binding purification tags;one or more wash buffers; andone or more proteases for cleaving protease cleavage sites.
[0253] (E2) For the kit denoted as (El), further comprising RNase inhibitors and protease inhibitors.
[0254] (E3) For the kit denoted as (El) or (E2), wherein the capture agents are attached to magnetic beads or agarose beads.
[0255] (E4) For the kit denoted as any one of (E1)-(E3), wherein the cell lysate is selected from HeLa cell lysate, E. coli cell lysate, wheat germ lysate, and reticulocyte lysate.* * *
[0256] Various modifications and variations of the described methods, pharmaceutical compositions, and kits of the invention will be apparent to those skilled in the art without departing from the scope and spirit of the invention. Although the invention has been described in connection with specific embodiments, it will be understood that it is capable of further modifications and that the invention as claimed should not be unduly limited to such specific embodiments. Indeed, various modifications of the described modes for carrying out the invention that are obvious to those skilled in the art are intended to be within the scope of the invention. This application is intended to cover any variations, uses, or adaptations of the invention following, in general, theAttorney Docket: 114203-3092principles of the invention and including such departures from the present disclosure come within known customary practice within the art to which the invention pertains and may be applied to the essential features herein before set forth.
Claims
Attorney Docket: 114203-3092CLAIMSWhat is claimed is:
1. An expression library comprising a set of expression templates, each expression template encoding a different polypeptide and each vector further comprising polynucleotide sequences encoding one or more of:a. an N-terminal translation enhancer sequence;b. a translation initiation sequence;c. a first purification tag;d. a solubility tag;e. an open reading frame (ORF) encoding a polypeptide and further comprising a 3' UTR sequence, an engineered poly-Atail, or bothf. an optional detection tag; org. an optional second purification tag.
2. The expression library of claim 1, wherein the N-terminal translation enhancer sequence is selected from chloramphenicol acetyltransferase (CAT), E01, UTR 17, ARC-1, and omega sequence of tobacco mosaic virus (SEQ ID NO: 1).
3. The expression library of claim 1 or claim 2, wherein the translation initiation sequence is a Kozak sequence.
4. The expression library of claim 3, wherein the Kozak sequence is selected from GATGATAATATGG (SEQ ID NO: 2) and GCCGCCACCATGG (SEQ ID NO: 3).
5. The expression library of any one of claims 1-4, wherein the first purification tag is selected from an haloalkane dehalogenase-derived tag, Albumin-binding protein (ABP), Alkaline Phosphatase (AP), AU1 epitope, AU5 epitope, Bacteriophage T7 epitope (T7-tag), Bacteriophage V5 epitope (V5-tag), Biotin-carboxy carrier protein (BCCP), Bluetongue virus tag (B-tag), Calmodulin binding peptide (CBP), Chloramphenicol Acetyl Transferase (CAT), Cellulose binding domain (CBP), Chitin binding domain (CBD), Choline-bindingAttorney Docket: 114203-3092domain (CBD), Dihydrofolate reductase (DHFR), E2 epitope, FLAG epitope, Galactose- binding protein (GBP), Green fluorescent protein (GFP), Glu-Glu (EE-tag), Glutathione S- transferase (GST), Human influenza hemagglutinin (HA), Histidine affinity tag (HAT), Horseradish Peroxidase (HRP), HSV epitope, Ketosteroid isomerase (KSI), KT3 epitope, LacZ, Luciferase, Maltose-binding protein (MBP), Myc epitope, NusA, PDZ domain, PDZ ligand, Polyarginine (Arg-tag), Polyaspartate (Asp-tag), Polycysteine (Cys-tag), Polyhistidine (His-tag), Polyphenylalanine (Phe-tag), Profinity eXact, Protein C, SI -tag, S-tag, Streptavadin-binding peptide (SBP), Staphylococcal protein A (Protein A), Staphylococcal protein G (Protein G), Strep-tag, Streptavadin, Small Ubiquitin-like Modifier (SUMO), Tandem Affinity Purification (TAP), T7 epitope, Thioredoxin (Trx), TrpE, Ubiquitin, Universal, and VSV-G.
6. The expression library of any one of claims 1-5, wherein the solubility tag is selected from a small ubiquitin-like modifier (SUMO) tag, SUMOstar protease, a glutathione S- transferase (GST)-tag, a maltose-binding protein (MBP)-tag, a thioredoxin (Trx)-tag, RRGRRGRRGRRGRRG (SEQ ID NO: 4) tag, a NusA tag, an N-terminal extension (NEXT) tag, and Bl domain of Streptococcal protein G (GB1).
7. The expression library of any one of claims 1-6, wherein the vector further comprises a protease cleavage site after the solubility tag.
8. The expression library of claim 7, wherein the solubility tag is selected from SUMOstar protease, HRV3C protease, and HaloTEV.
9. The expression library of any one of claims 1-8, wherein the detection tag is a fluorescent tag.
10. The expression library of any one of claims 1-9, wherein the second purification is a histidine tag.
11. The expression library of any one of claims 1-10, wherein each template further comprises a protease cleavage site adjacent to the second purification tag.Attorney Docket: 114203-309212. The expression library of any one of claims 1-11, wherein the ORF encodes only domains included in a mature polypeptide form.
13. The expression library of claim 12, wherein the mature polypeptide form excludes one or more of signal peptides; intracellular domains of transmembrane polypeptides or receptors; intra-polypeptides linkers for polypeptides; or any amino acids associated with post- translational proteolytic processing.
14. The expression library of any one of claims 1-13, wherein the ORF is codon optimized.
15. A high-throughput in vitro polypeptide synthesis system comprising:the expression library of any one of claims 1 to 14;a plurality of individual discrete volumes, each volume comprising a different expression template from the expression library;a cell lysate;an RNA polymerase;a reaction buffer;a first capture agent for binding a first capture tag and an optional second capture agent for binding an optional second capture tag; andone or more purification buffers.
16. The system of claim 15, wherein the plurality of individual discrete volumes is defined in a microwell format.
17. The system of claim 15, wherein the plurality of individual discrete volumes are droplets.
18. The system of claim 17, wherein the droplets are formed on a microfluidic device.
19. The system of claim 17, wherein the droplets are formed on a glass surface.Attorney Docket: 114203-309220. The system of any one of claims 16-19, wherein the first capture agent is attached to a microwell surface.
21. The system of claim 20, wherein the first capture agent is attached to the microwell surface by a photocleavable linker.
22. The system of any one of claims 16-21, wherein a different photocleavable linker may be used in different wells to allow for spatiotemporal release of the polypeptide.
23. The system of any one of claims 15-22, wherein the first capture agent is attached to a bead.
24. The system of any one of claims 15-23, wherein the optional second capture agent is attached to beads.
25. The system of any one of claims 15 to 24, further comprising a RNase inhibitor and a protease inhibitor.
26. The system of claim 25, wherein the RNase inhibitor inactivates one or more of RNase A, RNase B, or RNase T2.
27. The system of claim 25, wherein the protease inhibitor is aprotinin and / or leupeptin.
28. The system of any one of claims 15 to 27, further comprising one or more proteases capable of cleaving one or more protease cleavage site(s).
29. A method of highly parallel in vitro polypeptide synthesis comprising:a. allocating expression templates from an expression library into individual discrete volumes, each individual discrete volume comprising a different transcription template, each transcription template comprising a polynucleotide sequence encoding a different polypeptide and further comprisingi. an N-terminal translation enhancer sequence;ii. a translation initiation sequence;Attorney Docket: 114203-3092i i i . a fi r st puri fi cati on tag;iv. a solubility tagv. an ORF encoding the polypeptide;vi. a 3’ UTR sequence, an engineered poly-Atail, or both; and vii. an optional detection tag;viii. an optional second purification tag; andb. expressing the polypeptide from the expression template in the individual discrete volume in the presence of a cell lysate and a reaction buffer wherein the expressed polypeptide is captured in the individual discrete volume by a first capture agent; c. washing the captured polypeptide with a wash buffer;d. for polypeptides that are not surface polypeptides, releasing the polypeptide into solution via a protease that cleaves a protease cleavage site adjacent to the first purification tag.
30. The method of claim 29, wherein the N-terminal translation enhancer sequence is selected from chloramphenicol acetyltransferase (CAT), E01, UTR 17, ARC-1, and omega sequence of tobacco mosaic virus (SEQ ID NO: 1).
31. The method of claim 29 or 30, wherein the translation initiation sequence is a Kozak sequence.
32. The method of claim 31, wherein the Kozak sequence is selected from GATGATAATATGG (SEQ ID NO: 2) and GCCGCCACCATGG (SEQ ID NO: 3).
33. The method of any one of claims 29, wherein the first purification tag is an haloalkane dehalogenase-derived tag.
34. The method of any one of claims 29, wherein the solubility tag is a small ubiquitin-like modifier (SUMO) tag, SUMOstar protease, a glutathione S-transferase (GST)-tag, aAttorney Docket: 114203-3092maltose-binding protein (MBP)-tag, a thioredoxin (Trx)-tag, or aNusAtag, , an N-terminal extension (NEXT) tag, and Bl domain of Streptococcal protein G (GB1).
35. The method of any one of claims 29, wherein the template further comprises a protease cleavage site after the solubility tag.
36. The method of claim 35, wherein the solubility tag is selected from SUMOstar protease, HRV3C protease, and HaloTEV.
37. The method of any one of claims 29, wherein the detection tag is a fluorescent tag.
38. The method of any one of claims 29-37, wherein the second purification tag is a histidine tag.
39. The method of any one of claims 29-38, wherein each template further comprises a protease cleavage site adjacent to the second purification tag.
40. The method of any one of claims 29, wherein the ORF encodes only domains included in a mature polypeptide form.
41. The method of claim 40, wherein the mature polypeptide form excludes one or more of signal peptides; intracellular domains of transmembrane polypeptides or receptors; intrapolypeptide linkers for polypeptide; or any amino acids associated with post-translational proteolytic processing.
42. The method of any one of claims 29, wherein the ORF is codon optimized.
43. The method of any one of claims 29, wherein the expression template comprises a second purification tag, and the method further comprises after step (c) capturing the polypeptide with a second capture agent that binds the second purification tag and washing the captured polypeptide with a second wash buffer.Attorney Docket: 114203-309244. The method of any one of claims 29, further comprising removing the second purification tag via a protease that cleaves a protease cleavage site adjacent to the solubility tag.
45. The method of any one of claims 29 to 44, further comprising culturing cells in each individual discrete volume and measuring a phenotypic response to the polypeptide expressed in each individual discrete volume.
46. The method of claim 45, wherein the cells are selected from bacterial cells, yeast cells, fungi cells, and eukaryotic cells.
47. The method of claim 45, wherein the phenotypic response is measured by RNA-seq, singlecell RNA seq, spatial transcriptomics, flow cytometry, microscopy imaging, transwell migration, immunoassay, proliferation assay, co-culture or a combination thereof.
48. The method of any one of claims 29 to 47, wherein a plurality of individual discrete volumes is defined in a microwell format.
49. An expression template for cell-free polypeptide synthesis comprising:a promoter operably linked to a polynucleotide sequence encoding a polypeptide;an N-terminal translation enhancer sequence positioned upstream of an open reading frame (ORF) encoding the polypeptide;a first purification tag positioned N-terminal to the ORF; anda 3' untranslated region (3'-UTR) sequence, an engineered poly-Atail, or both, positioned downstream of the ORF.
50. The expression template of claim 49, wherein the expression template is DNA-based and selected from circular double-stranded DNA, linear double-stranded DNA, or branched DNA.
51. The expression template of claim 49, wherein the expression template is RNA-based and comprises messenger RNA (mRNA).
52. The expression template of any one of claims 49-51, wherein the promoter is selected from T7, T3, SP6, and tac promoters.Attorney Docket: 114203-309253. The expression template of any one of claims 49-52, further comprising a solubility tag positioned between the first purification tag and the ORF.
54. The expression template of claim 53, further comprising a protease cleavage site positioned between the solubility tag and the ORF.
55. The expression template of any one of claims 49-54, wherein the ORF is codon optimized for expression in mammalian cells.
56. A kit for high-throughput cell-free polypeptide synthesis comprising:an expression library according to any one of claims 1-14;a cell lysate suitable for in vitro transcription and translation;a reaction buffer comprising ribonucleotide triphosphates (rNTPs), polyamines, and divalent cations;capture agents for binding purification tags;one or more wash buffers; andone or more proteases for cleaving protease cleavage sites.
57. The kit of claim 56, further comprising RNase inhibitors and protease inhibitors.
58. The kit of claim 56 or 57, wherein the capture agents are attached to magnetic beads or agarose beads.
59. The kit of any one of claims 56-58, wherein the cell lysate is selected from HeLa cell lysate, E. coli cell lysate, wheat germ lysate, and reticulocyte lysate.