Microcompartmentalized ultra-high-throughput screening from single-copy gene libraries

The integration of IVTT with a nucleic acid replication system in microcompartments addresses the inefficiencies of current uHTS methods by enabling direct encapsulation and amplification of single DNA molecules, ensuring clonality and high protein expression for efficient screening of protein variants.

JP2025539740APending Publication Date: 2025-12-09PARIS SCI & LETTRES +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025526834
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-10
Filing Date
2023-11-10
Publication Date
2025-12-09

Smart Images

  • Figure 2025539740000013
    Figure 2025539740000013
  • Figure 2025539740000014
    Figure 2025539740000014
  • Figure 2025539740000015
    Figure 2025539740000015
Patent Text Reader

Abstract

The present invention provides a novel method that is specifically designed to statistically start from a single copy of a nucleic acid molecule and requires only a single encapsulation step to provide a clonal microcompartment library that represents a high concentration of encoded polypeptides. The method of the present invention makes it possible to perform in vitro uHTS by one-step encapsulation of linear DNA constructs containing candidate sequences to be tested in a molecular mixture, allowing for simultaneous specific gene amplification and protein expression.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the fields of directed protein evolution, protein engineering, and screening. [Background technology]

[0002] Directed evolution of proteins Directed evolution (DE) of proteins (e.g., enzymes) is an engineering approach inspired by natural evolution that allows the manipulation of proteins and enzymes without requiring precise knowledge of the protein's structural, functional, or mechanistic features. It is carried out through cycles of genetic diversification and enrichment. Starting with a known gene or a small set of known genes with a given sequence, in a first step, some diversity is introduced to obtain a collection (library) of genes that differ in nucleotide sequence but are related to the parent nucleotide sequence. These genes are called variants. Various methods are then used to enrich this library for variants with desired properties. This screening or selection step, performed based on the properties of the encoded protein variants, ultimately allows the recovery of genes encoding proteins with the most interesting properties (e.g., improved catalytic rate, or greater stability, unnatural specificity, etc.), which can then be used for further diversification and enrichment. By combining this approach with DNA sequencing, it is also possible to identify beneficial or deleterious mutations with respect to the target property. To date, directed evolution has been applied to improve or modify the properties of a wide variety of enzymes. New methodologies are being introduced to manipulate ever-larger libraries and screen for various types of activities (e.g., enzymatic activity, fluorescence, binding to specific receptors / ligands, etc.). In particular, the use of microcompartments has enabled ultra-high-throughput screening methods. Microcompartmentalized ultra-high-throughput screening (uHTS) methods use random encapsulation of variants in individual microcompartments to manipulate large libraries of variants or candidate genes. This random encapsulation is performed at such low concentrations that many microcompartments contain zero or one member of the library, with only a small fraction of these microcompartments containing more than one member (e.g., two or more).This process, called Poisson partitioning, allows for clonality (i.e., the fact that signals observed in most microcompartments are associated with unique variants in the library). For this reason, the microcompartments resulting from Poisson partitioning of variants are sometimes referred to as "monoclonal" microcompartments (Holstein, JM, Gylstorff, C. & Hollfelder, F. Cell-free Directed Evolution of a Protease in Microdroplets at Ultrahigh Throughput. ACS Synthetic Biology acssynbio.0c00538 (2021) doi:10.1021 / acssynbio.0c00538). By recovering the microcompartments that exhibit the most interesting signals, the library can be enriched for interesting variants. These microcompartments can be, for example, water-in-oil microdroplets or liposomes generated and manipulated using microfluidic devices. The volumes of these microcompartments range from femtoliters to microliters.

[0003] In vitro methods for directed evolution Traditionally, organisms are used to express proteins of interest from their encoding genes. The expressed recombinant protein is typically confined to the cytoplasm or periplasm of the transformed cell, or displayed on the cell surface. In all of these cases, the cell itself encapsulates both the encoding gene (e.g., as a multicopy number plasmid) and the encoded protein, thereby maintaining a strong phenotype-genotype link. However, there are cases where it is beneficial to avoid the use of cells.

[0004] In vitro protein evolution has been developed as a subset of directed evolution techniques, and the DE process is carried out in a cell-free environment. The advantage of performing DE in vitro is the ability to design cytotoxic proteins, proteins whose function requires unusual, unnatural, or cell-impermeable substrates, or the use of lethal selection conditions. Additionally, the introduction of genetic variants into living hosts (a process called transformation) is often the throughput-limiting step in workflows using in vivo protein expression, which limits the size of the libraries that can be engineered. Removing this constraint paves the way for higher-throughput DE protocols compared to those currently available for in vivo approaches.

[0005] Genotype-phenotype correlations In the absence of an organism, genotype-phenotype linkages must be maintained artificially. Additionally, to efficiently screen a protein library, as explained above, it is necessary that most microcompartments do not contain more than one of the gene library members.

[0006] One approach to genotype-phenotype linkage in the absence of cells is to create a direct physical connection between the encoding gene and its gene product (so-called "display system"). To accomplish this function, various display systems, such as phage display, ribosome display, mRNA display, CIS display, and SNAP display, have been developed over the years. However, these "display systems" only provide one or a few copies of the protein of interest attached to each gene variant, and in some cases (e.g., phage display) still require in vivo steps performed in living transformed cells. While such "display systems" can be used for affinity-based selection (panning), they are often insufficient for assessing the catalytic properties of enzymes or other functional protein properties (e.g., fluorescence). In fact, it is generally difficult to detect enzymatic activity or other protein properties from a single molecule or only a few molecules.

[0007] Amplification and expression in droplets Another option for maintaining genotype-phenotype linkages rather than establishing physical interactions between genes and their protein products is to confine both within the same microcompartment. In the case of emulsion microdroplet, liposome, or microfabricated chamber uHTS technologies, the microcompartment can perform this function in addition to separating genetic variants from each other.

[0008] It is possible to express proteins from genes without using live cells using cell extracts or in vitro expression platforms such as the PURE system (Shimizu, Y et al. Cell-free translation reconstituted with purified components. Nat. Biotechnol. 19, pp. 751-755 (2001)). This reaction is called IVTT ("in vitro transcription and translation"). In this case, no live cells are involved. To use IVTT reactions in microcompartments for screening applications, it is necessary to have only one variant gene present in most compartments. In an in vitro setting, this is achieved by diluting the gene library before encapsulation, so that compartments typically contain only one gene molecule (Poisson partitioning). However, the use of high dilutions required to ensure Poisson partitioning, where many compartments contain no more than one variant, also reduces the yield of protein expression from the IVTT reaction and therefore the screening efficiency.

[0009] A common approach for in vitro uHTS is to improve the performance of the IVTT reaction in each compartment by generating more protein from unique copies of the encapsulated gene, thereby obtaining more signal associated with the activity of this protein in the sorting device. Indeed, if sufficient amounts of protein are generated within a compartment, it becomes easier to distinguish between positive and negative microcompartments (e.g., those that exhibit the desired enzymatic activity and those that do not), thus facilitating the recovery of only those macrocompartments containing genes encoding active protein variants. However, current in vitro methods for obtaining multiple copies of the same gene variant in each microcompartment while maintaining clonality are cumbersome.

[0010] To mitigate the low expression levels resulting from the use of single molecules of DNA template, several methods have been developed, which are based on the clonal pre-amplification of gene variants.

[0011] The first approach uses isothermal amplification to create concatemers. For example, rolling circle amplification (RCA) starts with a single circular DNA molecule containing a gene and creates multiple concatemerized copies of the same gene. Thus, a single molecule contains multiple copies of the genetic information, increasing gene concentration even with a single encapsulation event. However, in this case, concatemer generation requires a preprocessing step, and in the case of libraries, it is not guaranteed that a single DNA molecule contains a copy of a single variant of the library.

[0012] The second approach uses droplet PCR amplification from a single encapsulated gene followed by pico-injection of the IVTT (in vitro transcription and translation) mix (Mazutis, L et al. Droplet-Based Microfluidic Systems for High-Throughput Single DNA Molecule Isothermal Amplification and Analysis. Anal Chem 81, pp. 4813-4821 (2009)).

[0013] However, this method requires multiple steps, which increases the complexity and requires costly droplet-based microfluidic handling. For example, a 2021 study described a method for ultra-high-throughput cell-free directed evolution using such droplet-based microfluidic devices. In this method, a circular library of gene variants is dispersed into microdroplets after Poisson partitioning (the concentration of this library is adjusted so that approximately 10% of the droplets contain a single variant and most droplets contain no variants) and first amplified by rolling circle amplification. An IVTT mixture for protein expression is then pico-injected into the droplets. A substrate for the enzyme of interest is then pico-injected into the droplets again. Finally, after incubation, the resulting droplets can be sorted according to the fluorescent signal associated with the enzymatic conversion of the substrate (Holstein, JM, Gylstorff, C. & Hollfelder, F. Cell-free Directed Evolution of a Protease in Microdroplets at Ultrahigh Throughput. ACS Synthetic Biology acssynbio.0c00538 (2021) doi:10.1021 / acssynbio.0c00538). This cumbersome protocol, involving multiple microfluidic steps, is technically challenging and limits the throughput of this method.

[0014] A third approach avoids the tedious multiple microfluidic manipulation of droplets by using so-called bead display (or microbead display; see A Sepp, DS Tawfik, AD Griffiths. Microbead display by in vitro compartmentalization: selection for binding using flow cytometry. FEBS letters 532 (3), pp. 455-458). In this case, gene amplification (e.g., by PCR) and expression reactions are performed clonal-wise on beads. The resulting microbeads carry both multiple clonal copies of the gene and multiple copies of the corresponding protein. Thus, they physically maintain the phenotype-genotype linkage and can be manipulated in bulk. The beads are then encapsulated in droplets containing a fluorescent substrate, the activity of the retained protein is assayed, and the droplets are sorted according to the detected activity level to recover beads carrying the most active protein variants and, ultimately, the genes encoding these variants. However, this bead display protocol also involves several preparatory steps before screening can be performed.

[0015] For example, in 2013, a bead display technology based on the SNAP tag system was developed, which can display up to several million copies of a protein of interest along with its gene (Diamante, L., Gatti-Lafranconi, P., Schaerli, Y. & Hollfelder, F. In vitro affinity screening of protein and peptide binders by megavalent bead surface display. Protein engineering, design & selection: PEDS 26, pp. 713-724 (2013)). This approach requires complex, multi-step protocols for bead preparation, encapsulation, and selection, which increases the workload and the likelihood of failure.

[0016] A fourth option, particularly desirable for its simplicity, would be a "one-pot" process that combines amplification / replication of a single compartmentalized nucleic acid molecule with protein expression, all in the same mixture, without external intervention. Combining DNA replication and expression systems (IVTT, such as the PURE system) in a single compartment has been attempted in the past. However, significant problems arose due to the poor compatibility of expression systems (including the PURE system) with DNA replication. For example, PCR cannot be used in this situation because thermal cycling would disrupt the protein expression machinery. While this compatibility could be improved by selecting a specific amplification scheme and adjusting reaction conditions (e.g., by reducing the ribonucleotide triphosphate concentration), existing approaches still require high DNA concentrations for efficiency. In the case of the microcompartments required for ultra-high throughput screening applications, this required high concentration necessitates many initial DNA copies in each compartment, which is incompatible with random Poisson partitioning and clonal expression (Han et al., Applied microbiology and biotechnology, 106(24), 2022; Libicher et al., Nature Communication, 11(1), 2020).

[0017] It would therefore be advantageous to have a simple method of in vitro uHTS screening that works by direct encapsulation of single gene copies (especially in linear form), such as PCR products, and allows immediate post-incubation screening without additional steps, avoiding complex microfluidic fusion or pico-injection protocols.

[0018] Therefore, known methods for performing microcompartmentalized in vitro uHTS of functional proteins (e.g., enzymes, fluorescent proteins, binders) require tedious protocols to create microcompartments containing both a single gene variant (one or multiple copies) and sufficient amounts of the corresponding protein / enzyme so that activity is detectable by a sorting device within the microcompartment, allowing for high-frequency sorting. In particular, current methods for obtaining multiple copies of the same gene variant in each microcompartment while maintaining clonality are cumbersome and prone to error, as they require multiple complex steps before and after compartmentalization. As detailed above, current strategies rely on pre-amplification steps on solid supports (bead display) or complex microfluidic protocols using multiple fusion / pico-injection steps. [Prior art documents] [Non-patent literature]

[0019] [Non-Patent Document 1] Holstein, JM, Gylstorff, C. & Hollfelder, F. Cell-free Directed Evolution of a Protease in Microdroplets at Ultrahigh Throughput. ACS Synthetic Biology acssynbio.0c00538 (2021) doi:10.1021 / acssynbio.0c00538 [Non-patent document 2] Shimizu, Y et al. Cell-free translation reconstituted with purified components. Nat. Biotechnol. 19, pp. 751-755 (2001) [Non-patent document 3] Mazutis, L et al. Droplet-Based Microfluidic Systems for High-Throughput Single DNA Molecule Isothermal Amplification and Analysis. Anal Chem 81, pp. 4813~4821 (2009) [Non-patent document 4] A Sepp, DS Tawfik, AD Griffiths. Microbead display by in vitro compartmentalization: selection for binding using flow cytometry. FEBS letters 532 (3), pp. 455-458 [Non-Patent Document 5] Diamante, L., Gatti-Lafranconi, P., Schaerli, Y. & Hollfelder, F. In vitro affinity screening of protein and peptide binders by megavalent bead surface display. Protein engineering, design & selection: PEDS 26, pp. 713-724 (2013) [Non-patent document 6] Han et al., Applied microbiology and biotechnology, 106(24), 2022 [Non-Patent Document 7] Libicher et al., Nature Communication, 11(1), 2020 [Non-patent document 8] "Self-replication of DNA by its encoded proteins in liposome-based synthetic cells". NATURE COMMUNICATIONS (2018)9:1583 page [Non-Patent Document 9] Uhlmann, E. & Peyman, A. (1990) Chemical Reviews, 90, pp. 543~584 [Non-licensed Document 10] Schomburg D., Schomburg I., Springer Handbook of Enzymes. 2 edn. Heidelberg: Springer; pp. 2001~2009 [Non-licensed Document 11] Liebecq C., IUPAC-IUBMB Joint Commission on Biochemical Nomenclature (JCBN) and Nomenclature Committee of IUBMB (NC-IUBMB) Biochem. Mol. Biol. Int. 1997;43:1151~1156 pages; IUBMB (1992), Enzyme Nomenclature 1992, Academic Press, San Diego [Non-licensed Document 12] Schomburg D, Schomburg I. Methods Mol Biol. 2010;609:113~28 pages [Non-licensed Document 13] Chang A, Schomburg I, Placzek S, Jeske L, Ulbrich M, Xiao M, Sensen CW, Schomburg D, Nucleic Acids Res. 2015 Jan;43. Epub 2014 Nov 5. BRENDA in 2015: exciting developments in its 25th year of existence [Non-licensed Document 14] Quantitative assessment of fluorescent proteins. Nat Methods. 2016 Jul;13(7):557~62 pages [Non-licensed Document 15] Yusuf, Dimas, et al. "The transcription factor encyclopedia." Genome biology 13.3 (2012): pages 1-25 [Non-Patent Document 16] Bernath, Kalia, Shlomo Magdassi, and Dan S. Tawfik. "Directed evolution of protein inhibitors of DNA-nucleases by in vitro compartmentalization (IVC) and nano-droplet delivery." Journal of molecular biology 345.5 (2005): pp. 1015-1026 [Non-Patent Document 17] Ueno H et al. "Amplification of over 100 kbp DNA from Single Template Molecules in Femtoliter Droplets" ACS Synth Biol. 2021 Sep 17;10(9):2179~2186 [Non-Patent Document 18] Su'etsugu, Masayuki et al. "Exponential propagation of large circular DNA by reconstitution of a chromosome-replication cycle." Nucleic acids research 45.20 (2017): pp. 11525-11534 [Non-Patent Document 19] Kulczyk AW et al. "The Replication System of Bacteriophage T7" Enzymes. 2016;39:89~136 pages [Non-Patent Document 20] Shimizu, Ya. Cell-free translation reconstituted with purified components. Nat. Biotechnol. 19, pp. 751~755 (2001) [Non-licensed Document 21] Schwartz, JJ, Lee, C. & Shendure, J. Accurate gene synthesis with tag-directed retrieval of sequence-verified DNA molecules. Nature methods 9, pp. 913~915 (2012) [Non-licensed Document 22] Christopher, GF & Anna, SL Microfluidic methods for generating continuous droplet streams. Journal Of Physics D-Applied Physics 40, pp. R319~R336 (2007) [Non-licensed Document 23] Sam Duwe, Peter Dedecker, Optimizing the fluorescent protein toolbox and its use, Current Opinion in Biotechnology, Volume 58, 2019, pages 183~191, ISSN 0958-1669 [Non-licensed Document 24] Koveal, D., Rosen, PC, Meyer, DJ. A high-throughput multiparameter screen for accelerated development and optimization of soluble genetically encoded fluorescent biosensors. Nat Commun 13, p. 2919 (2022) [Non-licensed Document 25] Baret, J.-C. et al. Fluorescence-activated droplet sorting (FADS): efficient microfluidic cell sorting based on enzymatic activity. Lab on a chip 9, pp. 1850~1858 (2009) [Non-Patent Document 26] D. Blanken, D. Foschepoth, A. Calaca Serrao, and C. Danelon. Genetically controlled membrane synthesis in liposomes. Nat. Commun. 2020, 11(1):4317 pages [Non-Patent Document 27] Mencia et al., "Terminal protein-primed amplification of heterologous DNA with a minimal replication system based on phage Φ29". PNAS 108, pp. 18655~18660 (2011) [Non-patent document 28] Van Nies et al., "Self-replication of DNA by its encoded proteins in liposome-based synthetic cells". Nature Communications 9, p. 1583 (2018) [Non-Patent Document 29] Conrad et al., "Maximizing transcription of nucleic acids with efficient T7 promoters". Commun Biol 3, page 439 (2020) Summary of the Invention [Problem to be solved by the invention]

[0020] Therefore, there is an urgent need to provide a novel microcompartmentalized uHTS method for screening proteins with desired activities that is free from the above-mentioned cumbersome constraints and works by directly Poisson partitioning single DNA molecules into microcompartments. [Means for solving the problem]

[0021] The present invention fulfills this need. Indeed, the inventors have developed a novel method specifically designed to amplify a single copy of a protein-encoding nucleic acid molecule (e.g., a standard linear nucleic acid (e.g., a PCR product)), simultaneously express the encoded protein, and assay its activity in microcompartments. This method is based on the combined use of an IVTT system and a system capable of replicating and amplifying nucleic acids in vitro (i.e., in a non-cellular environment). In the provided example, a replicator linear DNA is used that contains, in addition to the encoded gene and its regulatory sequences, an origin of replication 5' of each DNA strand (e.g., the origin of replication of phage Phi29 or its derivatives), and minimal replication machinery (e.g., phage Phi29 p2, p3, p5, and p6 proteins) is added either as purified protein or as DNA encoding this protein.

[0022] Van Nies et al. ("Self-replication of DNA by its encoded proteins in liposome-based synthetic cells", NATURE COMMUNICATIONS (2018) 9:1583) recently described an in vitro reaction called IVTTR (i.e., IVTT coupled with an in vitro replication system), in which a reconstituted phage Phi29 replication machinery allows replication of linear DNA fragments simultaneously with the expression of encoded proteins in the IVTT mixture. However, this coupled replication-expression reaction has only been demonstrated for the expression of specific proteins from phage Phi29 (replication proteins p2, p3, p5, and p6), not for other non-Phi29 proteins. Furthermore, it requires a high initial concentration of replicating DNA, which is incompatible with Poisson partitioning in microcompartments such as microdroplets or liposomes.

[0023] In accordance with the present invention, the inventors have surprisingly discovered that: i) separate replication-expression can be used to express any protein of interest, or any library of DNA fragments (e.g., those encoding variants of a protein of interest); ii) this modified reaction can be performed (preferably using Poisson partitioning into microcompartments of single nucleic acid molecules) starting from very low concentrations of carrying replicated DNA containing the gene of interest; iii) replication and transcription / translation can be coupled; iv) this reaction is compatible with enzymatic assays (particularly fluorescent assays); v) the reaction can be performed (preferably using Poisson partitioning) starting from very low concentrations of carrying replicated DNA containing the gene of interest; When performed in microcompartments starting with a sufficiently low initial concentration of a library of DNA fragments (using the uHTS technique), a population of clonal microcompartments containing high copy numbers of one of the DNA fragments from the library and high copy numbers of its encoded protein is obtained, generating a signal related to the activity of the encapsulated protein that can be used in an ultra-high-throughput sorting device; and vi) overall, this approach enables ultra-high-throughput in vitro sorting of protein or enzyme libraries directly by encapsulation of the gene library, without any intermediate steps. Poisson partitioning enables clonality, i.e., ensuring that only one copy of a nucleic acid molecule (i.e., the library-specific variant of the library being tested) is incorporated into the majority of microcompartments. Thus, the present invention provides a unique, reliable, rapid, and efficient method for performing clonal in vitro microcompartmentalized uHTS by one-step encapsulation of nucleic acid molecules (e.g., linear DNA constructs) containing the gene variants (or candidate sequences) to be tested into a molecular mixture that allows for simultaneous specific gene amplification and protein expression, and optionally activity / enzyme assays.

[0024] SUMMARY OF THE INVENTION To this end, the present invention provides a novel method for screening a library of nucleic acid molecules, each encoding a candidate polypeptide, for a polypeptide having a desired activity by direct encapsulation. In this method, (i) a first composition comprising an in vitro platform including a nucleic acid replication mechanism, an optional nucleic acid transcription mechanism, and a nucleic acid translation mechanism, and (ii) a second composition comprising a library of nucleic acid molecules, each containing at least one candidate sequence encoding at least one candidate polypeptide, are prepared and mixed. The resulting mixture is partitioned (preferably using Poisson partitioning) into microcompartments, at least one of which contains a single copy of the nucleic acid molecule. Poisson partitioning allows for clonality, i.e., the fact that most microcompartments contain unique variants of the library (i.e., only one copy of a nucleic acid molecule from the library being tested is incorporated into most of the microcompartments). Then, within the microcompartments, the candidate sequences are amplified by the nucleic acid replication mechanism in the microcompartments, optionally transcribed by the nucleic acid transcription mechanism, and translated into at least one candidate polypeptide by the nucleic acid translation mechanism.

[0025] The activity of the candidate polypeptides thus obtained is then detected, and a selection device is used to ultimately recover candidate sequences that encode candidate polypeptides having the desired activity.

[0026] The present invention also relates to a kit comprising an in vitro platform as defined above, said kit also comprising at least one component for producing microcompartments. DETAILED DESCRIPTION OF THE INVENTION

[0027] definition Unless expressly defined herein, all technical and scientific terms used herein have the same meaning as commonly understood by one skilled in the art of chemistry, biochemistry, cell biology, molecular biology, protein engineering, and medicine.

[0028] As used herein, when used to define products, compositions, cell lines, uses, and methods, the words "comprising" (and all forms of "comprising," such as "comprise" and "comprises"), "having" (and all forms of "having," such as "have" and "has"), "including" (and all forms of "including," such as "includes" and "include"), or "containing" (and all forms of "containing," such as "contains" and "contain") are open-ended and do not exclude additional, unrecited elements or method steps. "Consisting of" means the exclusion of any other component or step. "Consisting essentially of" means the exclusion of any other essentially significant component or step (but does not exclude other minor / insignificant components or steps). In this disclosure, the terms "comprising," "consisting," and "consisting essentially of" may be interchanged where appropriate.

[0029] As used herein, "directed evolution" refers to an engineering approach inspired by natural evolution, which allows for the manipulation of proteins and enzymes without requiring precise knowledge of protein structure, function, or mechanism. This is accomplished through cycles of genetic diversification and enrichment. Starting with a known gene or a small set of known genes with a given sequence, in a first step, some diversity is introduced to obtain a collection (library) of genes that differ in nucleotide sequence but are related to the parent nucleotide sequence. Various methods are then used to enrich this library for variants with desired properties. This screening or selection step, performed based on the properties of the encoded protein variants, ultimately allows for the recovery of genes encoding proteins with the most interesting properties, which can then be used for further diversification and enrichment. This approach can also be combined with DNA sequencing to identify beneficial or deleterious mutations with respect to the target property.

[0030] As used herein, "bioprospection" and "functional metagenomics" are used interchangeably and refer to alternative approaches for identifying proteins with desired activities. In bioprospection, the screened library can be a library of natural microorganisms or cells, or isolated nucleic acid sequences. These libraries of nucleic acid strands can be obtained by extraction and cloning of nucleic acids obtained from microorganisms, plants, or animals, or multiple of these.

[0031] As used herein, the terms "screening" or "screening" or "selection" (all of which are considered synonymous herein) refer to the process of testing and selecting compounds / active agents (particularly proteins and / or nucleic acid molecules encoding such proteins) for a particular effect / activity.

[0032] "Candidate sequence," as used herein, refers to a genetic sequence (i.e., a nucleotide sequence) encoding a candidate polypeptide whose activity is to be tested in the methods of the invention. With respect to directed evolution, a candidate sequence (also referred to as a genetic variant or mutant) generally encodes a polypeptide variant or mutant (i.e., the original polypeptide having one or more mutations selected from substitutions, insertions, deletions, and permutations) of a known polypeptide having a desired activity, with the goal of identifying a variant that has an increased desired activity or improved properties compared to the known polypeptide.

[0033] As used herein, a "library" of nucleic acid molecules refers to a mixture of several, preferably many (eg, hundreds to billions, or more) different nucleic acid molecules.

[0034] As used herein, the term "partitioning" refers to the separation of a sample into multiple portions, also called "zones" or "compartments." Zones are generally physical such that the sample in one zone does not mix, or does not substantially mix, with the sample in an adjacent zone.

[0035] The term "micro-compartment" or "microcompartment" as used herein refers to a compartment having an internal volume between femtoliters and microliters (i.e., a characteristic size of less than 1 mm). Microcompartments include, but are not limited to, microdroplets (e.g., water-in-oil microdroplets), vesicles (e.g., liposomes), microchambers, etc. While the compartments used in the screening methods of the present invention may advantageously be identical in shape and volume, collections of microcompartments of different sizes (e.g., obtained by liposome formation protocols or bulk emulsification protocols) may also be used. Methods for obtaining microcompartments are known to those skilled in the art. In the case of microdroplets and emulsions, such methods include various emulsification techniques, such as membrane emulsification, mechanical emulsification, microfluidic emulsification, high homogenization, microfluidization, and sonication, among others. Emulsions are generally stabilized by suitable surfactants known to those skilled in the art, which may confer long-term stability to the emulsions. For example, microdroplets can be obtained as monodisperse emulsions using a microfluidic approach or some other approach.

[0036] As used herein, the term "microcompartmentalization" refers to the separation of a reaction into multiple microreactors (eg, multiple microcompartments).

[0037] The terms "polynucleotide," "nucleic acid molecule," and "nucleic acid" are used interchangeably herein and refer to a polymeric or oligomeric macromolecule made from nucleotide monomers (also called nucleotide residues). A nucleotide monomer is composed of a nucleobase, a sugar, or a sugar analog (e.g., but not limited to, ribose or 2'-deoxyribose in the case of RNA and DNA; however, other sugars or sugar analogs may be used in the case of heterologous nucleic acids called XNAs, such as 1,5-anhydrohexitol in the case of HNA, cyclohexene in the case of CeNA, threose in the case of TNA, glycol in the case of GNA, ribose modified with an additional bridge connecting the 2' oxygen and 4' carbon in the case of locked nucleic acids (LNA), and N-(2-aminoethyl)-glycine units in the case of peptide nucleic acids (PNA)), and one to three phosphate groups. Typically, polynucleotides are formed by phosphodiester bonds linking ribose or deoxyribose groups in natural nucleic acids, but other artificial bonds (e.g., peptide bonds in PNAs) may be used between individual nucleotide monomers in XNAs. Nucleic acid molecules include, but are not limited to, ribonucleic acid (RNA), deoxyribonucleic acid (DNA), and mixtures thereof, including, for example, RNA-DNA hybrids (mixed polyribo-polydeoxyribonucleotides). These terms encompass single- or double-stranded, linear or circular, natural or synthetic, unmodified or modified versions (e.g., genetically modified polynucleotides, optimized polynucleotides), sense or antisense polynucleotides, and chimeric mixtures (e.g., RNA-DNA hybrids). Furthermore, polynucleotides may contain non-natural nucleotides and be interrupted by non-nucleotide components.Exemplary DNA nucleic acids include, but are not limited to, complementary DNA (cDNA), genomic DNA, plasmid DNA, DNA vectors, viral DNA (e.g., viral genomes, viral vectors), oligonucleotides, probes, primers, satellite DNA, microsatellite DNA, coding DNA, non-coding DNA, antisense DNA, and any mixture thereof. Exemplary RNA nucleic acids include, but are not limited to, messenger RNA (mRNA), precursor messenger RNA (pre-mRNA), small interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), RNA vectors, viral RNA, guide RNA (gRNA), antisense RNA, coding RNA, non-coding RNA, antisense RNA, satellite RNA, small cytoplasmic RNA, small nuclear RNA, etc. The polynucleotides described herein can be synthesized by standard methods known in the art, for example, by using an automated DNA synthesizer (e.g., commercially available from Biosearch, Applied Biosystems, etc.), or by DNA assembly and gene synthesis methods, or by mutagenesis methods, or can be obtained from natural sources (e.g., genomes, cDNA, etc.) or artificial sources (e.g., commercially available libraries, plasmids, etc.) using molecular biology techniques known in the art (e.g., cloning, PCR, etc.). Nucleic acids can also be chemically synthesized, for example, according to the phosphate triester method (see, for example, Uhlmann, E. & Peyman, A. (1990) Chemical Reviews, 90, 543-584).

[0038] In the context of the present invention, a "replicator" refers to a nucleic acid molecule ("nucleic acid replication machinery") that can replicate under specific conditions in the presence of a specific compound. The replicator serves as a template for amplification, so that the newly created nucleic acid molecule is its copy or reverse-complement copy. This replication results in an increase in the concentration of the replicator. For example, if starting with a single replicator molecule, after a while there may be two, then four, eight, etc. molecules, i.e., an exponential increase in the replicator concentration. Replicators can be any type of nucleic acid (DNA, RNA, or hybrid or XNA) and may contain specific subsequences or non-standard modifications. For example, replicators preferably contain subsequences selected from an origin of replication (OR), a primer binding site, one or more genes with appropriate regulatory sequences (e.g., promoters), a ribosome binding site (RBS), a barcode, etc.

[0039] The terms "protein" and "polypeptide" are used interchangeably herein and refer to any peptide-linked polymer of amino acids, regardless of length or post-translational modification. These terms preferably refer to a polymer of amino acid residues containing at least six amino acids covalently linked by peptide bonds. The polymer may be linear, branched, or cyclic. The polymer may contain naturally occurring and / or amino acid analogs and may be interrupted by non-amino acids. The maximum number of amino acids contained in a polypeptide is not limited. As a general indication, the term refers to both short polymers (typically referred to in the art as peptide or protein fragments) and longer polymers (typically referred to in the art as polypeptides or proteins). The term encompasses, among other things, native polypeptides, modified polypeptides (also referred to as derivatives, analogs, variants, or mutants), polypeptide fragments, polypeptide multimers (e.g., dimers), mutant polypeptides, artificial polypeptides, and fusion polypeptides. A polypeptide is understood to be any translational product of a polynucleotide, regardless of size and whether or not glycosylated, and includes peptides and proteins. The polypeptides / proteins usable herein (including protein derivatives, protein variants, protein fragments, protein domains, protein epitopes, and protein domains) may be further modified by chemical or enzymatic modification. This means that such chemically or enzymatically modified polypeptides contain chemical groups other than the 20 naturally occurring amino acids. Examples of such chemical or enzymatic modifications include post-translational modifications. Chemical or enzymatic modification of a polypeptide may confer advantageous properties compared to the parent polypeptide, such as one or more of improved stability, increased biological half-life, increased water solubility, increased activity, improved properties, labeling, etc.

[0040] As a general guideline, and not by way of limitation, if an amino acid polymer contains more than 50 amino acid residues, the amino acid polymer is preferably referred to as a polypeptide or protein, whereas if the polymer consists of 50 or fewer amino acids, the polymer is preferably referred to as a "peptide." The sense of writing and writing amino acid sequences of polypeptides, proteins, and peptides as used herein is the conventional sense of writing and writing. By convention, the amino terminus is placed on the left, and the sequence is then written and read from the amino terminus (N-terminus) to the carboxyl terminus (C-terminus) from left to right.

[0041] The term "functional protein" or "functional polypeptide" refers to a protein or polypeptide that has a desired activity, as opposed to a non-functional protein or peptide that does not have the desired activity (e.g., a variant of a functional polypeptide that has lost the desired activity). In the case of directed evolution of proteins, a library of protein variants, some of which are functional and some of which are non-functional, is constructed and then screened using the methods of the present invention to recover functional variants.

[0042] The terms "activity," "activity of interest," or "function," when referring to a protein, refer to any biological activity of the protein that is screened for using the methods of the present invention. Activities of interest include, among others, enzymatic activity, reporter activity (e.g., fluorescent activity), regulatory activity (e.g., transcription factor activity), and biological activity (e.g., drug or antibiotic activity).

[0043] "Enzyme activity" refers to the specific catalytic activity of an enzyme. As used herein, "enzyme" refers to a protein with catalytic properties (enzymatic properties). While most biomolecules capable of catalyzing chemical reactions within cells are enzymes, some catalytic biomolecules are made of RNA and are therefore ribozymes, distinct from protein enzymes. Enzymes act by lowering the activation energy of a chemical reaction, thereby increasing the reaction rate. Enzymes either remain unchanged during the reaction or are regenerated unchanged at the end of the catalytic reaction. The initial molecule is the enzyme's "substrate," and the molecule formed from this substrate is the product of the reaction. Enzymes are often characterized by their extremely high specificity. Furthermore, most enzymes are characterized by being reusable, meaning they can catalyze many substrate-to-product conversions.

[0044] Enzymes are generally globular proteins that act alone or as a complex of several enzymes or subunits. Like all proteins, enzymes consist of one or more polypeptide chains that fold to form a three-dimensional structure corresponding to the native state. However, some enzymes can be completely or partially disordered proteins. Many enzymes are composed of multiple peptide chains or assemble as multimers.

[0045] The size can vary from about 50 residues to over 2000 residues. Only a small portion of the enzyme (most often 2-4 residues, but sometimes more) is directly involved in catalysis (the so-called catalytic site, located within the catalytic domain). This catalytic site may be located near one or more binding sites where a substrate is bound and oriented to catalyze a chemical reaction. The catalytic site and the binding site form the active site of the enzyme.

[0046] Enzymes perform numerous functions in living organisms. For example, enzymes may be involved in signal transduction and regulation of cellular processes, generation of movement, active membrane transport, digestion, metabolism, the immune system, nucleic acid digestion or cleavage mechanisms, or nucleic acid production (referred to herein as "nucleic acid-acting enzymes"), prodrug conversion mechanisms (conversion of prodrugs to drugs). The enzyme is preferably a prokaryotic, eukaryotic, or viral enzyme, and is preferably an enzyme derived from an animal, plant, algae, microalgae, insect, microorganism, archaea, bacteria, parasite, yeast, fungus, or virus. Within animals, the enzyme may be a mammalian enzyme, such as a human enzyme. Various categories of enzymes are known to those skilled in the art, who may refer in particular to references in the field (e.g., specialized databases such as those described in Schomburg D., Schomburg I., Springer Handbook of Enzymes. 2 edn. Heidelberg: Springer; 2001-2009; Liebecq C., IUPAC-IUBMB Joint Commission on Biochemical Nomenclature (JCBN) and Nomenclature Committee of IUBMB (NC-IUBMB) Biochem. Mol. Biol. Int. 1997; 43: 1151-1156; IUBMB (1992), Enzyme Nomenclature 1992, Academic Press, San Diego; and Schomburg D, Schomburg I. Methods Mol. Biol. 2010; 609: 113-28). Enzyme databases; in particular the BRENDA database (particularly available at brenda-enzymes.org) as described, for example, in Chang A, Schomburg I, Placzek S, Jeske L, Ulbrich M, Xiao M, Sensen CW, Schomburg D, Nucleic Acids Res. 2015 Jan;43. Epub 2014 Nov 5. BRENDA in 2015: exciting developments in its 25th year of existence.Enzymes are also classified based on the type of activity in the Enzyme Commission number (EC number) classification, and any enzyme in any subcategory of EC numbers 1 to 7 is of interest in the present invention. The term enzyme is intended to encompass naturally occurring enzymes and their derivatives (e.g., mutant enzymes and / or artificial enzymes), provided that such derivatives may have enzymatic activity. As used herein, an "enzyme fragment" refers to any portion of an enzyme, provided that preferably such fragment / portion may have enzymatic activity. In the case of a protein enzyme, the enzyme fragment preferably comprises at least 6 consecutive amino acid residues of the enzyme (preferably the enzyme catalytic site) (preferably at least 8 consecutive amino acid residues of the enzyme, preferably at least 10, preferably at least 15, preferably at least 20, preferably at least 30 amino acid residues of the enzyme).

[0047] Some enzymes use "cofactors," a term that refers to non-protein molecules that form stable or transient complexes with the enzyme and whose presence is essential for the activity of the enzyme. Cofactors can be inorganic ions (especially metal ions) or other organic molecules (e.g., FAD or NADH).

[0048] Some enzymes use "coenzymes," a term that refers to nonprotein organic molecules that form complexes with the enzyme and whose presence is essential for the enzyme's activity. Coenzymes can be divided into two types. The first, called "prosthetic groups," consist of a coenzyme that is tightly (or covalently) and permanently bound to the protein. The second type of coenzyme, called "cosubstrates," is temporarily bound to the protein. The cosubstrate can be released from the protein at some point and then re-bound.

[0049] "Reporter activity" refers to the specific activity of a reporter protein. "Reporter protein" or "reporter polypeptide" refers to a protein that is directly or indirectly detectable in the screening step of the method of the invention, preferably by optical means. Reporter proteins that are directly detectable in the screening step include, in particular, colored or fluorescent proteins. "Fluorescent protein" refers to a protein that emits light when exposed to light or other radiation, the emission being at a longer wavelength than the light to which the protein is exposed. Non-limiting examples of fluorescent proteins include blue (e.g., EBFP2, mTagBFP2), blue-green (e.g., mTurquoise, mTurquoise2, mCerulean, mCerulean3, mCerulean3), UV-excited green (e.g., mT-Sapphire), green (e.g., EGFP, mEGFP, Emerald, mEmerald, sfGFP), yellow-green (e.g., YFP, mPapaya, YPet, Citrine, mCitrine, Venus, mVenus), , Topaz, mTopaz, Clover, mClover, mNeonGreen), orange (e.g., mOrange, mOrange2, mKO, mKO2), orange-red (e.g., tdTomato, TagRFP, TagRFP-T, DsRed2), red (e.g., mRuby, mRuby2, mApple, mRFP1, mCherry, FusionRed), and far-red (e.g., mKate2, mNeptune, mCardinal, mPlum) fluorescent proteins (Cranfill PJ, Sell BR, Baird MA, Allen JR, Lavagnino Z, de Gruiter HM, Kremers GJ, Davidson MW, Ustione A, Piston DW. Quantitative assessment of fluorescent proteins. Nat Methods. 2016 Jul;13(7):557-62).Indirectly detectable reporter proteins in the screening process include, inter alia, enzymes that produce a product (e.g., a colored, fluorescent, or luminescent product) from a precursor molecule (e.g., a chromogenic compound, a fluorescent compound, or a luciferin) that can be easily detected, for example, by spectroscopic means. A "chromogenic compound" is a compound that produces color when modified by an enzyme, a "fluorescent compound" is a compound that produces fluorescence when modified by an enzyme, and a "luciferin" is a compound that produces light when modified by an enzyme.

[0050] The terms "regulatory activity" and "transcription factor activity" are used interchangeably and refer to the specific transcription-inducing activity of a transcription factor. The term "transcription factor" (or "TF"), as used herein, refers to a protein that controls the rate of transcription from DNA to mRNA through a specific binding mechanism. All TFs share the essential characteristic of containing at least one DNA-binding domain (DBD). Examples of TFs include, but are not limited to, AP-1, CREB, C / EBP, c-Myc, NF-1, and TFII factors. Various categories of TFs are known to those skilled in the art, and those skilled in the art may refer, in particular, to references in the field (e.g., Yusuf, Dimas, et al. "The transcription factor encyclopedia." Genome biology 13.3 (2012): pp. 1-25). The term transcription factor is intended to encompass natural TFs and their derivatives (e.g., mutant TFs and / or artificial TFs), provided that such derivatives are capable of controlling the rate of transcription from DNA to mRNA through a specific binding mechanism. As used herein, a "TF fragment" refers to any portion of a TF, provided that preferably such a fragment / portion is capable of controlling the rate of transcription from DNA to mRNA via a specific binding mechanism (e.g., DBD, etc.). In the case of a protein TF, a TF fragment preferably comprises at least 6 consecutive amino acid residues of the TF (and preferably is the DBD) (preferably comprising at least 8 consecutive amino acid residues of the TF, preferably at least 10, preferably at least 15, preferably at least 20, preferably at least 30 amino acid residues of the TF).

[0051] The terms "peptide" or "protein fragment" or "portion of a peptide or protein" or "protein domain" as used herein refer to a portion of a peptide or protein, i.e., a portion of the contiguous amino acid sequence that constitutes said peptide or protein (referred to as the peptide or protein from which the fragment is derived). When the fragment is a peptide or protein, the fragment preferably comprises at least 6 contiguous amino acids of the peptide or protein from which the fragment is derived, more preferably at least 8 contiguous amino acids, more preferably at least 10 contiguous amino acids, more preferably at least 12 contiguous amino acids, more preferably at least 15 contiguous amino acids, more preferably at least 20 contiguous amino acids, more preferably at least 30 contiguous amino acids. When the fragment is a peptide or protein fragment, it preferably has a three-dimensional structure under non-denaturing conditions (e.g., conditions that are normally non-denaturing for proteins, in particular in the absence of denaturants and / or chaotropic agents). When the fragment is a peptide or protein fragment, it is preferably a functional fragment. By "functional fragment" is meant any peptide or protein fragment that retains at least one of the original functions of the peptide or protein from which the fragment is derived.Preferably, a functional fragment performs its function with an efficiency at least equal to 30% of that of the peptide, protein, or molecule, and preferably at least 40%, preferably at least 45%, preferably at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, preferably at least 85%, preferably at least 90%, preferably at least 91%, preferably at least 92%, preferably at least 93%, preferably at least 94%, preferably at least 95%, preferably at least 96%, preferably at least 97%, preferably at least 98%, preferably at least 99%, and preferably at least 100% of that of the peptide or protein. Examples of protein fragments (particularly functional fragments) include protein domains, protein epitopes, and the like. Fragments usable herein (including protein fragments, peptide fragments, protein domains, protein epitopes, and protein domains) can be further modified by chemical or enzymatic modification (e.g., post-translational modification). If the fragment is a peptide or protein fragment, the chemically / enzymatically modified fragment contains chemical groups other than the 20 naturally occurring amino acids (e.g., may contain post-translational modifications).

[0052] The term "protein derivative" or "derivative" as used herein refers to mutant proteins and / or artificial proteins and / or mimetics. Derivatives are preferably functional derivatives. "Functional derivative" refers to any protein derivative that retains at least one of the original functions of the protein from which it is derived. Preferably, a functional derivative performs said function with an efficiency equal to at least 30% of that of said protein, preferably at least 40%, preferably at least 45%, preferably at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, preferably at least 85%, preferably at least 90%, preferably at least 91%, preferably at least 92%, preferably at least 93%, preferably at least 94%, preferably at least 95%, preferably at least 96%, preferably at least 97%, preferably at least 98%, preferably at least 99%, preferably at least 100% of the effectiveness of said protein.

[0053] "Three-dimensional structure" or "tertiary structure," as used herein, refers to the unique folding of a molecule in space. A molecule with a three-dimensional structure is one that has a stable spatial arrangement (commonly called a fold) that is unique to the molecule and generally closely related to the function of the molecule. When this structure is dissociated by the use of denaturants or chaotropic agents, the molecule is said to be denatured and its function is lost. In particular, a three-dimensional structure is one with limited flexibility. In contrast, a non-three-dimensional structure can adopt a set of dynamic configurations that are constantly changing over time. In the case of proteins, peptides, or fragments thereof, the three-dimensional structure refers to the folding of a polypeptide chain in space. In this case, the three-dimensional structure is not a linear chain of amino acids (not a linear continuum of amino acids without any spatial organization) that can adopt a set of dynamic configurations that are constantly changing over time. The three-dimensional structure of a protein, peptide, mixed molecule containing a protein or peptide, or fragment thereof, is maintained by various interactions that may include: - covalent interactions (disulfide bridges between cysteines); - electrostatic interactions (ionic bonds, hydrogen bonds); - Van der Waals interactions; - Interactions with the solvent and environment (ions, lipids, etc.).

[0054] As used herein, "post-translationally modified" refers to a chemical or enzymatic modification that occurs naturally or artificially on a protein or protein fragment after or simultaneously with protein translation (e.g., biological or biochemical synthesis using the cellular machinery) or after or simultaneously with protein synthesis (e.g., artificial and / or chemical synthesis). This means that at least one of the naturally occurring amino acids of a protein or protein fragment has been modified by the addition of at least one chemical group and / or the modification (e.g., without limitation, removal) of at least one chemical group of this naturally occurring amino acid. Examples of such chemical or enzymatic modifications include, but are not limited to, glycosylation, phosphorylation, acylation, carboxylation, acetylation, biotinylation, hydroxylation, lipoylation, amidation, ubiquitination, sumoylation, deamination, etc. A "post-translationally modified protein," as used herein, refers to a protein having at least one (i.e., one or more) post-translational modification. A "post-translationally modified protein fragment," as used herein, refers to a protein fragment that has at least one (ie, one or more) post-translational modifications.

[0055] The terms "platform" and "system" are considered synonymous herein. They refer to an in vitro material apparatus made up of one or more elements operatively linked to each other to achieve a specific purpose. The platform of the present invention is preferably a cell-free system. Such a cell-free system may be selected from a cell extract, a mixture of purified cell components, a reconstituted in vitro transcription / translation (IVTT) system, a DNA amplification and replication system, an enzyme and protein activity assay, and any combination thereof. In the context of the present invention, the platform is designed to achieve at least two biological processes (i.e., isothermal gene replication and protein expression).

[0056] In particular, the platform of the present invention combines an IVTT system with a system capable of performing in vitro replication under the highly specific circumstances of a microcompartment containing only one nucleic acid molecule comprising a candidate sequence encoding a candidate polypeptide, resulting in an in vitro transcription, translation, and replication system (referred to herein as the "IVTTR" system).

[0057] As used herein, the term "mechanism" refers to a biological device made of one or more substances that functionally cooperate to realize a biological mechanism.

[0058] The methods of the present invention provide a "nucleic acid replication mechanism." As such, the nucleic acid replication mechanism comprises, consists essentially of, or consists of all materials necessary to complete nucleic acid replication. Specifically, the nucleic acid replication mechanism comprises one or more materials selected from the following: a DNA polymerase, an RNA polymerase, a DNA termination protein, a single-stranded DNA binding protein (SSBP), a double-stranded DNA binding protein (DSBP), a nucleic acid sequence encoding a DNA polymerase, a nucleic acid sequence encoding an RNA polymerase, a nucleic acid sequence encoding a DNA termination protein, a nucleic acid sequence encoding an SSBP, a nucleic acid sequence encoding a DSBP, necessary cofactors and substrates (e.g., nucleoside triphosphates, ammonium sulfate), auxiliary nucleic acids (e.g., primers), and any combination thereof. Other protein enzyme activities may also be included. For example, enzymes involved in DNA replication in organisms or viruses, such as recombinases, topoisomerases, transposases, reverse transcriptases, helicases, nuclease processivity factors, and enzymes associated with nucleic acid metabolism, among others, may also be included in the in vitro platform of the first composition.

[0059] The methods of the present invention optionally provide a "nucleic acid transcription machinery." As such, the nucleic acid transcription machinery comprises, consists essentially of, or consists of all materials necessary to complete nucleic acid transcription starting from DNA and ending with mRNA. Specifically, the nucleic acid transcription machinery includes one or more materials selected from an RNA polymerase, a nucleic acid sequence encoding the RNA polymerase, and any combination thereof (as well as substrates, cofactors, auxiliary enzymes, proteins such as transcription factors (TFs), etc.).

[0060] In the methods of the invention, "nucleic acid translation machinery" is provided, which thus comprises, consists essentially of, or consists of all materials necessary to complete nucleic acid translation beginning with mRNA and ending with a polypeptide.

[0061] Specifically, the nucleic acid translation machinery includes one or more ribosomes, translation factors (eg, initiation factors, elongation factors, release factors, NTP recycling enzymes), aminoacyl-tRNA synthetases, tRNAs, and the like.

[0062] Any one of the mechanisms defined above may be selected from, or derived from, a native or wild-type mechanism from a bacteriophage (e.g., Phi29, T7, Qβ, MS2, or T4), a bacterium (e.g., Escherichia coli), a yeast (e.g., Saccharomyces cerevisiae), a virus (e.g., Venezuelan equine encephalitis virus, adeno-associated virus, norovirus, or influenza virus), another eukaryotic cell (e.g., rabbit reticulocyte), or any combination thereof (e.g., obtained or engineered after one or more modifications (e.g., genetic modifications such as gene mutations)).

[0063] Any one of the mechanisms defined above may be provided as a cell extract or a reconstituted system. "Cell extract" refers to an extract of cellular components, which may be used in crude form (i.e., without any prior purification) or in a more or less purified form (at least one purification step that enriches the extract for the desired compound). "Reconstituted system" refers to a reconstituted mixture of previously purified or recombinantly expressed components (e.g., the PURE system as a reconstituted transcription and translation system, or a mixture of Phi29 bacteriophage proteins p2, p3, p5, and p6 as a reconstituted replication system, both of which are described in more detail below).

[0064] The term "parameter" specifically refers to a characteristic that is a value used to adequately describe, define, or classify an entity or situation at a particular given time, and specifically refers to physical, quantifiable properties (e.g., concentration, temperature, pH, viscosity, and the like).

[0065] The term "substrate," as used herein, refers to a starting material that is subjected to a chemical or biological reaction (specifically, an enzymatic reaction in the presence of an enzyme). The substrate may be advantageously conjugated and / or fused to a compound or moiety that has a detectable optical property. Typically, a reaction (e.g., an enzymatic or chemical reaction) begins with one or more substrates and ends with one or more products. Thus, in this context, a "product" refers to a substance obtained by a reaction that begins with a substrate.

[0066] As used herein, the terms "optical technique," "optical method," or "optical property" refer to a technique, method, or property that allows optical detection and / or measurement. Optical techniques may be selected from absorption-based detection techniques (e.g., measurement of single or multiple wavelengths, etc.), fluorescence-based detection techniques (e.g., fluorescence intensity, fluorescence lifetime, fluorescence wavelength, FRET, etc.), and any combination thereof. Optical properties may be, inter alia, electromagnetic and / or color and / or fluorescent properties occurring at visible or invisible wavelengths. A "fluorescent compound" refers to a compound that emits light when exposed to light or other radiation, the emitted light having a longer wavelength than the light to which the compound is exposed. This emission may be advantageously detected and / or measured by optical techniques as defined above.

[0067] The term "compound," as used herein, refers to a chemical or biological entity such as a molecule. Examples of compounds include, but are not limited to, reagents, substrates, co-substrates, cofactors, coenzymes, DNA, RNA, hormones, antigens, epitopes, ligands, antibodies, receptors, toxins, and the like.

[0068] The expression "at least one microcompartment contains a single copy of a nucleic acid molecule" means herein that the average number of nucleic acid molecules per microcompartment (corresponding to the ratio of the sum of the individual nucleic acid molecule contents of the microcompartments to the total number of microcompartments) is 0 to 10, preferably 0 to 5, more preferably 0 to 2, and most preferably 0.01 to 1.

[0069] For example, if the average number of nucleic acid molecules per microcompartment is 1 after random distribution of DNA molecules in microcompartments of the same internal volume, this means that approximately 37% of the microcompartments are empty, approximately 37% of the microcompartments contain one nucleic acid molecule, and approximately 26% of the microcompartments contain two or more nucleic acid molecules. These values ​​are calculated from the known "Poisson distribution" formula P(k) = λk.ek / k!, where P(k) is the probability of observing a compartment containing k DNA molecules, and λ is the average number of DNA molecules per droplet. λ is easily calculated when the DNA concentration in the encapsulated mixture and the volume of each compartment are known. Such a division of DNA molecules into microcompartments, in which many microcompartments contain only a single DNA molecule (necessarily many compartments do not contain any DNA molecules), is called "Poisson division." This Poisson partitioning is crucial in all microcompartmentalized uHTS methods to allow for the preservation of clonality and genotype-phenotype associations at a statistical level. To do so, the concentration of the encapsulated entities (here, DNA molecules, but in other approaches, it could be bacteria or beads) is adjusted so that only a small fraction of the microcompartments contain multiple entities. The result is a collection of microcompartments, most of which contain zero or one entity. This approach therefore allows for the separation of most entities into individual microcompartments without relying on deterministic partitioning.

[0070] As used herein, the phrase "at least one nucleic acid molecule comprising at least one candidate sequence encoding at least one polypeptide" means: - the nucleic acid molecule may comprise a plurality of candidate sequences; and / or - each candidate sequence may encode multiple polypeptides; and / or - Each candidate sequence may be present as a single copy in the nucleic acid molecule.

[0071] Screening Method While known in vitro uHTS methods for screening functional proteins (e.g., enzymes, fluorescent proteins, binders, etc.) require cumbersome protocols, the present inventors have developed a novel system that is specifically designed to statistically start with a single copy of a nucleic acid molecule (e.g., a standard linear form of the nucleic acid (e.g., a PCR product)) and requires only a single encapsulation step to provide a clonal microcompartment library (that can be sent directly to the selection stage) containing multiple copies of a given nucleic acid molecule, including one or more genes and a high concentration of the encoded protein. Thus, the present invention provides a unique, reliable, and efficient method for performing in vitro uHTS by one-step encapsulation of nucleic acid molecules (e.g., linear DNA constructs) containing the gene variants (or candidate sequences) to be tested in a molecular mixture that allows for simultaneous specific gene amplification and protein expression.

[0072] Indeed, the inventors have unexpectedly demonstrated that it is possible to obtain sufficient levels of protein activity for uHTS screening from the encapsulation of single copies of linear nucleic acid molecules according to a Poisson distribution in microcompartments containing a platform that combines the two activities of isothermal gene replication and protein expression. Poisson partitioning allows for clonality, i.e., the fact that most microcompartments contain unique variants of the library. Therefore, the platform developed by the inventors combines an IVTT system with a system capable of replicating these unique variants in each microcompartment, resulting in an IVTTR system. This innovative platform relies on the association of (i) a nucleic acid replication mechanism, (ii) optionally a nucleic acid transcription mechanism (depending on the type of starting nucleic acid molecule used in the screening method), and (iii) a nucleic acid translation mechanism. The data obtained by the inventors demonstrates that by coupling both replication and translation machinery in a one-pot reaction, even when starting with a single copy of the gene-containing nucleic acid molecule in a microcompartment, simultaneous nucleic acid amplification and protein production of any gene can occur in quantities sufficient for automated detection of the protein activity of interest. The inventors have successfully applied this novel method to mutation screening or bioprospecting of any protein of interest.

[0073] Because this innovative approach relies on direct encapsulation of the mixture and does not involve complex, multi-step, or throughput-limiting procedures (e.g., microfluidic droplet manipulation, bacterial transformation, pre-amplification, or microbead manipulation), the data support that this approach can be successfully used to manipulate large libraries and enrich them for genes encoding proteins with targeted properties. Importantly, because this novel method is a compartmentalized approach, it is compatible with the enrichment of any protein library, including enzyme libraries and libraries of proteins with other activities, such as fluorescent proteins or transcription factors (TFs). This is in contrast to uHTS approaches, which are only compatible with binder selection.

[0074] More specifically, the inventors have shown that this system can be used to efficiently and reliably screen proteins with desired activities (e.g., enzymes, proteins in protein complexes, nucleic acid-binding proteins, transcription factors, and any combination thereof) at very high speeds (e.g., millions of variants per hour).

[0075] This novel method offers the following advantages, as supported by experimental data: (i) Nucleic acid molecules, even if initially present as single copies in microcompartments, undergo cycles of isothermal exponential amplification driven by the replication machinery, with the accumulated gene copies simultaneously serving as templates for protein translation, synthesizing high and easily detectable concentrations of gene products; (ii) This approach is advantageous for screening applications because the nucleic acid molecules are replicated exponentially in the microcompartments, starting with at least one single copy upon encapsulation and reaching nanomolar concentrations in the microcompartments upon entering the selection phase. All selected positive microcompartments are then pooled, facilitating the recovery of genetic material at the end of the screening round if their contents are analyzed or used as the starting library for the next directed evolution cycle. This amplification increases the amount of nucleic acid molecules recovered after screening, even if only a few compartments are sorted into positive bins, simplifying further processing of the selected library (i.e., recovery of selected genes); (iii) this approach is compatible with various microencapsulation strategies, such as liposomes or water-in-oil microdroplets; (iv) As starting materials, any type of nucleic acid molecule may be used, such as DNA (e.g., cDNA, single-stranded DNA, double-stranded DNA, partially double-stranded DNA, linear DNA, circular DNA, capped DNA, uncapped DNA, etc.), RNA (e.g., mRNA, capped RNA, uncapped RNA, DNA / RNA hybrid), or vector (e.g., plasmid or viral vector). For example, in vitro replication may be achieved using terminal P3-capped DNA containing a specific replication origin. Alternatively, in vitro replication may be initiated with 5'-phosphate-capped DNA, which is easily obtained (using a 5'-P primer during PCR, or by limiting an adapter, or by enzymatic phosphorylation of DNA). Therefore, a library composed of standard linear 5'-P DNA molecules, which can be easily obtained from PCR performed with a 5'-P primer or by enzymatic phosphorylation of 5'-OH DNA, may be directly used; (v) this gene amplification process can be adapted to long DNA or RNA constructs (up to multi-kilobase constructs), thus allowing the inclusion of additional functionality in the replicator, which can be used to improve the protein screening process; (vi) Experimental conditions can be changed at any stage (e.g., between the IVTTR stage and the protein characterization assay stage) without complex microfluidic manipulations (e.g., mixing of pico-injected droplets). For example, temperature can be easily adjusted during the replication and IVTT reactions, and then during evaluation of target protein activity. Concentrations in microcompartments can also be changed at any stage using evaporation, osmotic exchange, or diffusion of compounds from a continuous phase without disrupting the droplets or liposomes. When liposomes are used, they may also contain pores through which small molecules can diffuse from a continuous solution. Finally, several methods have been described for using nanoemulsions to deliver compounds with microemulsion compartments (Bernath, Kalia, Shlomo Magdassi, and Dan S. Tawfik. "Directed evolution of protein inhibitors of DNA-nucleases by in vitro compartmentalization (IVC) and nano-droplet delivery." Journal of molecular biology 345.5 (2005): pp. 1015-1026). In other words, activity assay conditions are not limited to those required for the IVTTR reaction.

[0076] Thus, the present invention provides a novel method for screening a library of nucleic acid molecules, each encoding a candidate polypeptide, for one or more nucleic acid molecules encoding a polypeptide having a desired activity, comprising the steps of: a)See below: (i) Nucleic acid replication mechanism, (ii) optional nucleic acid transcription machinery, and (iii) nucleic acid translation mechanism providing a first composition comprising an in vitro platform comprising: b) providing a second composition comprising a library of nucleic acid molecules, each comprising at least one candidate sequence encoding at least one candidate polypeptide; c) mixing the first composition of step a) with the second composition of step b); d) dividing the mixture thus obtained into microcompartments, at least one of said microcompartments containing a single copy of said nucleic acid molecule; e) amplifying said candidate sequences in the microcompartments by nucleic acid replication machinery; f) optionally transcribing the amplified candidate sequences of step e) in the microcompartments by a nucleic acid transcription machinery; g) translating the amplified candidate sequence of step e) or the transcribed candidate sequence of step f) into at least one candidate polypeptide by a nucleic acid translation machinery in the microcompartment; h) optionally adding at least one compound to the mixture of step c) or to said microcompartments of step g) and / or modifying at least one parameter in the microcompartments of step d) or g), wherein said at least one compound is preferably selected from a reagent, substrate, cofactor, coenzyme, and any combination thereof, of the candidate polypeptide encoded by the candidate sequence, and / or said at least one parameter is preferably selected from a concentration, temperature, pH, viscosity, and any combination thereof, i) detecting, preferably using optical methods, in the microcompartment, the activity of the candidate polypeptide of step g) under conditions that are optionally altered in step h); and j) using a selection device to recover candidate sequences encoding candidate polypeptides having a desired activity from the microcompartments that exhibit a specific activity signal; The present invention relates to a method comprising, consisting essentially of, or consisting of:

[0077] The method of the present invention is particularly suitable for high-throughput screening (HTS), preferably for ultra-high-throughput screening (uHTS). Specifically, the method of the present invention is carried out by using a method for screening at least 10 5 nucleic acid molecules, preferably at least 10 containing at least one candidate sequence 6 This allows for high-throughput or ultra-high-throughput screening of individual nucleic acid molecules.

[0078] The screening method of the present invention can be applied to any activity of interest, but the activity of interest is particularly enzymatic polymer degradation. The polymer may be a plastic polymer, particularly a plastic polymer selected from PLA (polylactic acid), PET (polyethylene terephthalate), PE (polyethylene), PP (polypropylene), PCL (polycaprolactone), and PS (polystyrene). In one embodiment, the polypeptide having the activity of interest is selected from the group consisting of an enzyme, a polypeptide of a protein complex, a nucleic acid-binding protein, a transcription factor, and any combination thereof (i.e., the method is preferably for screening a library of nucleic acid molecules each encoding a candidate polypeptide selected from the group consisting of a candidate enzyme, a candidate polypeptide of a protein complex, a candidate nucleic acid-binding protein, a transcription factor, and any combination thereof).

[0079] Step a) - Providing a first composition comprising an in vitro platform In step a), the following is carried out: (i) Nucleic acid replication mechanism, (ii) optional nucleic acid transcription machinery, and (iii) nucleic acid translation mechanism A first composition is provided that includes an in vitro platform comprising:

[0080] As described above, the nucleic acid replication machinery comprises, or consists essentially of, or consists of all materials necessary to complete nucleic acid replication. Specifically, the nucleic acid replication machinery comprises, or consists essentially of, or consists of one or more materials selected from DNA polymerase, RNA polymerase, DNA termination protein (TP), single-stranded DNA binding protein (SSBP), double-stranded DNA binding protein (DSBP), a nucleic acid sequence encoding a DNA polymerase, a nucleic acid sequence encoding an RNA polymerase, a nucleic acid sequence encoding a DNA termination protein, a nucleic acid sequence encoding an SSBP, a nucleic acid sequence encoding a DSBP, necessary cofactors and substrates (e.g., nucleoside triphosphates, etc.), auxiliary nucleic acids (e.g., primers), and any combination thereof.

[0081] The combination of materials required to complete nucleic acid replication will vary depending on the specific type of replication machinery selected.

[0082] For example, when the minimal replication machinery of phage Phi29 is used, the nucleic acid replication machinery preferably comprises (or consists essentially of, or consists of) a DNA polymerase, a DNA terminal protein (TP), a single-stranded DNA binding protein (SSBP), and a double-stranded DNA binding protein (DSBP). Since the replication reaction occurs in the IVTT mixture, one or more of these enzymes may be replaced with their encoding DNA having appropriate regulatory sequences (promoter, ribosome binding site, and terminator). In this case, the enzymes required for replication are produced in situ by in vitro transcription and translation before the replication reaction occurs. Even more preferably, when the minimal replication machinery of phage Phi29 is used, it is DNA polymerase, DNA terminal protein (TP), single-stranded DNA binding protein (SSBP), and double-stranded DNA binding protein (DSBP); or Nucleic acid sequences encoding DNA polymerases, DNA TPs, single-stranded DNA binding proteins (SSBPs), and double-stranded DNA binding proteins (DSBPs) Comprising (or consisting essentially of, or consisting of).

[0083] However, other viruses or organisms may have other replication mechanisms, some of which do not use DNA terminal proteins (TP).

[0084] As a result, a general minimal replication machinery preferably comprises (or consists essentially of, or consists of) a DNA or RNA polymerase, a single-stranded DNA binding protein (SSBP), and a double-stranded DNA binding protein (DSBP). This minimal replication machinery can then be completed by other substances, depending on the particular type of replication machinery selected.

[0085] Proteins included in the replication machinery (or proteins encoded by nucleic acid molecules included in the replication machinery) may be selected from, or derived from, native or wild-type machinery derived from a bacteriophage (e.g., Phi29, T7, Qβ, MS2, or T4), a bacterium (e.g., Escherichia coli), a yeast (e.g., Saccharomyces cerevisiae), a virus (e.g., Venezuelan equine encephalitis virus, adeno-associated virus, norovirus, or influenza virus), other eukaryotic cells (e.g., rabbit reticulocytes), or any combination thereof (e.g., obtained after one or more modifications (e.g., genetic modifications such as gene mutations) or engineered).

[0086] Preferably, the proteins comprised in the replication machinery (or the proteins encoded by the nucleic acid molecules comprised in the replication machinery) are selected or derived (e.g., obtained or engineered after one or more modifications (e.g., genetic modifications such as gene mutations)) from a natural or wild-type machinery from a bacteriophage, and more preferably from bacteriophage Phi29 (hence referred to as a "Phi29-based replication machinery"). In this latter case, the replication machinery preferably comprises (or consists essentially of, or consists of) a DNA polymerase, a DNA terminal protein (TP), a single-stranded DNA binding protein (SSBP), and a double-stranded DNA binding protein (DSBP), wherein the DNA polymerase corresponds to bacteriophage Phi29 protein p2, the DNA terminal protein (TP) corresponds to bacteriophage Phi29 protein p3, the single-stranded DNA binding protein (SSBP) corresponds to bacteriophage Phi29 protein p5, and the double-stranded DNA binding protein (DSBP) corresponds to bacteriophage Phi29 protein p6. In a particularly preferred embodiment, the replication machinery comprises: purified Phi29 proteins p2, p3, p5, and p6 (preferably purified recombinant Phi29 proteins p2, p3, p5, and p6 of the amino acid sequences of SEQ ID NOs: 45 to 48); or Nucleic acid sequences encoding Phi29 proteins p2 and p3 (preferably pUC57_OriLR_p2p3 of the sequence SEQ ID NO: 34), as well as purified Phi29 proteins p5 and p6 (preferably purified recombinant Phi29 proteins p5 and p6 of the amino acid sequences SEQ ID NO: 47 and 48). Comprising (or consisting essentially of, or consisting of).

[0087] Alternatively, other replication mechanisms may be used, such as a reconstituted plasmid replication mechanism (Ueno H et al. "Amplification of over 100 kbp DNA from Single Template Molecules in Femtoliter Droplets" ACS Synth Biol. 2021 Sep 17;10(9):2179-2186), a reconstituted chromosomal replication mechanism (Su'etsugu, Masayuki et al. "Exponential propagation of large circular DNA by reconstitution of a chromosome-replication cycle." Nucleic acids research 45.20 (2017):11525-11534), or a reconstituted viral replication mechanism (Kulczyk AW et al. "The Replication System of Bacteriophage T7" Enzymes. 2016;39:89-136).

[0088] The above-described proteins of this replication machinery may be provided as a cell extract or a reconstituted system. Preferably, a reconstituted system is used which contains the protein mixture defined above.

[0089] Other proteins with enzymatic activity may also be included. For example, nucleic acid replication machinery may include, among others, recombinases, topoisomerases, transposases, reverse transcriptases, helicases, nuclease processivity factors, enzymes involved in nucleic acid metabolism, and the like, which are involved in DNA replication in organisms or viruses. These may also be selected from, or derived from, native or wild-type machinery (e.g., obtained after one or more modifications (e.g., genetic modifications such as gene mutations) or engineered) from bacteriophages (e.g., Phi29, T7, Qβ, MS2, or T4), bacteria (e.g., Escherichia coli), yeast (e.g., Saccharomyces cerevisiae), viruses (e.g., Venezuelan equine encephalitis virus, adeno-associated virus, norovirus, or influenza virus), other eukaryotic cells (e.g., rabbit reticulocytes), or any combination thereof.

[0090] The replication machinery should also contain deoxyribonucleotide triphosphates (dNTPs).

[0091] The in vitro platform in the first composition optionally includes a nucleic acid transcription mechanism depending on the type of nucleic acid present in the second composition provided in step b).If the second composition in step b) includes mRNA, the nucleic acid transcription mechanism is not required.In other cases, particularly when the second composition in step b) includes DNA, the in vitro platform in the first composition also includes a nucleic acid transcription mechanism.

[0092] As described above, the nucleic acid transcription machinery comprises, or consists essentially of, or consists of all materials necessary to complete nucleic acid transcription starting from DNA and ending with mRNA. Specifically, the nucleic acid transcription machinery comprises one or more materials selected from an RNA polymerase, a nucleic acid sequence encoding the RNA polymerase, and any combination thereof (and substrates, cofactors, auxiliary enzymes, proteins such as transcription factors (TFs), etc.). Preferably, the nucleic acid transcription machinery comprises an RNA polymerase, for example, T7 RNA polymerase.

[0093] The in vitro platform of the first composition also includes nucleic acid translation machinery, which comprises, consists essentially of, or consists of all materials necessary to complete nucleic acid translation starting from mRNA and ending in a polypeptide. Specifically, the nucleic acid translation machinery includes one or more ribosomes, translation factors (e.g., initiation factors, elongation factors, release factors, NTP recycling enzymes), aminoacyl-tRNA synthetases, tRNAs, and the like.

[0094] Proteins included in the nucleic acid translation machinery or encoded by nucleic acid molecules included in this nucleic acid translation machinery, and proteins included in the optional nucleic acid transcription machinery or encoded by nucleic acid molecules included in this nucleic acid transcription machinery may be selected from or derived (e.g., obtained after one or more modifications (e.g., genetic modifications such as gene mutations) or engineered) from native or wild-type machinery from bacteriophage (e.g., Phi29, T7, Qβ, MS2, or T4), bacteria (e.g., Escherichia coli), yeast (e.g., Saccharomyces cerevisiae), viruses (e.g., Venezuelan equine encephalitis virus, adeno-associated virus, norovirus, or influenza virus), other eukaryotic cells (e.g., rabbit reticulocytes), or any combination thereof. These proteins may more particularly be selected from natural or wild-type mechanisms from bacteria (particularly Escherichia coli) or may be derived (e.g., obtained or engineered after one or more modifications (e.g., genetic modifications such as gene mutations)).

[0095] The in vitro platform contained in the first composition is preferably a cell-free system, which may be provided as a cell extract, a reconstituted system, or any combination thereof. Preferably, a reconstituted system is used.

[0096] When a nucleic acid transcription machinery is provided, the nucleic acid translation machinery and the nucleic acid transcription machinery are preferably provided in a single reconstituted system, preferably a purified gene expression machinery derived from Escherichia coli (E. coli), such as the PURE system ("Protein synthesis Using Recombinant Elements"), which is commercially available from various suppliers (e.g., GeneFrontier's "PUREfrex" system), which is a reconstituted nucleic acid transcription and translation machinery capable of isothermally expressing any open reading frame under a T7 promoter and an E. coli ribosome binding site, and contains 32 individually purified components: initiation factors IF1, IF2, and IF3; elongation factors EF-G, EF-Tu, and EF-T; release factors RF1 and RF3; ribosome recycling factor (RRF); 20 aminoacyl-tRNA synthetases (ARSs), methionyl-tRNA transformylase (MTF), T7 RNA polymerase, and ribosomes. In addition, the PURE system contains 46 tRNAs, nucleoside triphosphates (NTPs), creatine phosphate, 10-formyl-5,6,7,8-tetrahydrofolate, 20 amino acids, creatine kinase, myokinase, nucleoside diphosphate kinase, and pyrophosphatase (Shimizu, Y et al. Cell-free translation reconstituted with purified components. Nat. Biotechnol. 19, pp. 751-755 (2001)). There are three PUREfrex versions: PUREfrex 2.0 is an upgraded version of PUREfrex 1.0 for protein production, and PUREfrex 2.1 is a version of PUREfrex 2.0 that does not contain any reducing agents and allows for the adjustment of the redox environment of the solution.

[0097] The first composition prepared in step a) may further comprise other components useful in the method of the present invention. For example, when the method is applied to screening nucleic acid molecules encoding candidate enzymes, the substrate, cofactor, coenzyme (prosthetic group or cosubstrate) of the enzyme, or any combination thereof, may be added to the first composition prepared in step a).

[0098] Step b) - Providing a second composition comprising a library of nucleic acid molecules, each comprising at least one candidate sequence encoding at least one candidate polypeptide. In step b), a second composition is provided that comprises a library of nucleic acid molecules, each comprising at least one candidate sequence that encodes at least one polypeptide.

[0099] The nucleic acid molecule is preferably selected from the group consisting of DNA (e.g., gDNA, cDNA), RNA (e.g., mRNA), and any combination thereof (e.g., DNA / RNA hybrid). The nucleic acid molecule can be single-stranded, double-stranded, or partially double-stranded. The nucleic acid molecule can be linear or circular. The nucleic acid molecule can be derived from an in vitro preparation (e.g., chemical synthesis, PCR, isothermal amplification) or an in vivo preparation (e.g., plasmid, viral vector). The nucleic acid molecule can include (in addition to the candidate sequence encoding the aforementioned polypeptide) specific subsequences and specific caps such as terminal proteins. The nucleic acid molecule can further include (non-standard) modifications such as non-standard bases, epigenetic marks, internal modifications, or terminal modifications.

[0100] Each nucleic acid molecule in this library is preferably a replicator, i.e., a nucleic acid molecule capable of replicating under specific conditions in the presence of a specific compound (the "nucleic acid replication machinery"). The replicator serves as a template for amplification, so that newly generated nucleic acid molecules are copies or reverse-complement copies of this replicator. This replication results in an increase in the concentration of replicators. For example, if starting with a single replicator molecule, after a while there may be two, then four, eight, etc. molecules, i.e., an exponential increase in the concentration of replicators. Replicators can be any type of nucleic acid (DNA, RNA, or hybrid) and may contain specific subsequences or non-standard modifications (see above). For example, replicators preferably contain one or more subsequences selected from one or more genes with an origin of replication (OR), a primer binding site, and appropriate regulatory sequences (e.g., a promoter, a terminator, and a ribosome binding site (RBS)). More preferably, the replicator comprises one or more genes having an origin of replication (OR), a primer binding site, appropriate regulatory sequences (eg, a promoter), and a ribosome binding site (RBS).

[0101] When using a Phi29-based replication mechanism (comprising Phi29 p2, p3, optionally p5, and optionally p6 proteins, or nucleic acid sequences encoding them), the replicator is preferably a linear, double-stranded DNA containing at least a Phi29 origin of replication or a derivative thereof (e.g., a minimal origin) at each end and one or more sequences encoding a polypeptide with appropriate regulatory sequences (e.g., a promoter, a terminator, and a ribosome binding site (RBS)), each strand of which is further capped at the 5' end by phage Phi29 P3 protein or a phosphate group. Phosphorylated 5' ends (i.e., 5'-capped with a phosphate group) can be readily obtained, for example, by using a 5' phosphate primer during PCR, by limiting adapters, or by enzymatic phosphorylation of 5' OH DNA.

[0102] In one embodiment, the replicator contains only one sequence encoding a candidate polypeptide.

[0103] However, the present inventors have shown that a nucleic acid molecule functioning as a replicator can contain multiple sequences encoding polypeptides. For example, the present inventors have started with a nucleic acid molecule into which a reporter gene (e.g., a gene encoding a fluorescent or chromogenic gene) and a gene encoding a polypeptide of interest have been inserted as a fusion. In this case, the colored or fluorescent signal associated with the reporter gene indicates the expression level of the polypeptide of interest and can be used during the selection step to normalize activity with respect to the expression level. In this case, the selection process selects for nucleic acid sequences encoding polypeptides with high specific activity rather than polypeptides with high apparent activity (apparent activity is the product of specific activity and polypeptide concentration).

[0104] In another embodiment, the replicator comprises several sequences each encoding a different polypeptide, where at least one of the sequences encodes the candidate polypeptide to be screened, while other sequences may encode other candidate polypeptides to be screened, reporter proteins (e.g., colored or fluorescent proteins) that can be used to normalize the expression levels of the polypeptides to be screened, or substrate proteins, or functional variants or fragments thereof (if the candidate polypeptide is an enzyme with a protein substrate).

[0105] When a replicator contains several sequences (e.g., two) each encoding a different polypeptide, these different sequences can be present as one or more fusions, as independent transcription units, or as a multicistronic (e.g., bicistronic) construct. Fusions are particularly suitable for normalization because they ensure that the expression levels of the fused polypeptides are exactly the same.

[0106] Preferably, the candidate polypeptide encoded by the candidate sequence is selected from the group consisting of a putative enzyme, a putative transcription factor, a putative functional derivative thereof, and a putative functional fragment thereof.

[0107] In one embodiment, the candidate sequences are pre-selected based on at least one of the sequence, a portion of the sequence, a structural property, a functional property, and any combination thereof.

[0108] When the PURE system is used as the transcription and translation machinery, the nucleic acid sequence encoding the candidate (or other additional) polypeptide is preferably under the control of a T7 promoter.

[0109] The replicator may further comprise additional subsequences, particularly selected from a ribosome binding site (RBS), a barcode, a restriction enzyme binding site, an aptamer.

[0110] When the PURE system is used as the transcription and translation machinery, the nucleic acid sequence encoding the candidate (or other additional) polypeptide preferably contains an E. coli ribosome binding site.

[0111] When a Phi29-based replication mechanism and a PURE transcription and translation mechanism are used, a particularly preferred replicator has a 5' phosphate capped end and comprises, on one strand, the following nucleic acid sequences from 5' to 3': oriL (preferably Phi29 oriL), a T7 promoter, an RBS (ribosome binding site, preferably an E. coli RBS), a gene encoding a candidate polypeptide, a terminator, and oriR (preferably Phi29 oriR).

[0112] Additionally, we have applied this novel method in combination with DNA barcoding technology by including randomized DNA subsequences called barcodes or unique molecular identifiers (UMIs) in the replicator. These barcodes or UMIs can be used in combination with sequencing to, for example, calculate the change in frequency of a given variant during screening cycles, correct for amplification bias, calculate a high-quality consensus from multiple reads, assemble short reads into longer complete sequences, or confirm that recombination does not occur during the amplification process or library preparation, or "call" one specific variant (e.g., by performing PCR using the barcode sequence as a primer) (Schwartz, JJ, Lee, C. & Shendure, J. Accurate gene synthesis with tag-directed retrieval of sequence-verified DNA molecules. Nature Methods 9, 913-915 (2012)). All of these technologies based on random DNA barcodes or UMIs are known.

[0113] Preferably, the method is a high-throughput screening (HTS) method, more preferably an ultra-high-throughput screening (uHTS) method. In a particularly preferred embodiment, the method produces at least 10 clones containing at least one candidate sequence (preferably a replicator as defined herein). 5 nucleic acid molecules, preferably at least 10 containing at least one candidate sequence (preferably a replicator as defined herein) 6 This allows for high-throughput or ultra-high-throughput screening of individual nucleic acid molecules.

[0114] The second composition prepared in step b) may further contain other components useful in the method of the present invention. For example, when this method is applied to screening nucleic acid molecules encoding candidate enzymes, the second composition prepared in step b) may contain an enzyme substrate, a cofactor, a coenzyme (prosthetic group or cosubstrate), or any combination thereof.

[0115] Step c) - Mixing the first composition of step a) with the second composition of step b). In step c), the first composition prepared in step a) is mixed with the second composition prepared in step b), which is preferably carried out at low temperature and / or immediately before step d) to avoid premature initiation of the replication or expression reaction.

[0116] Step d) - dividing the mixture thus obtained into microcompartments In step d), the mixture obtained in step c) is divided into microcompartments, at least one of said microcompartments containing a single copy of said nucleic acid molecule.

[0117] In a preferred embodiment, a substantial proportion of the microcompartments contain a single copy of the nucleic acid molecule, more preferably essentially all non-empty microcompartments contain a single copy of the nucleic acid molecule.

[0118] Several techniques for performing such partitioning are known in the art, resulting in different types of microcompartments, preferably selected from the group consisting of microdroplets (e.g., water-in-oil microdroplets), microvesicles (e.g., liposomes), microchambers, and any combination thereof.

[0119] The desired type of microcompartment can be generated using standard techniques known to those skilled in the art. For example, microdroplets can be obtained by microfluidic or bulk emulsification techniques. Advantageously, monodisperse water-in-oil microdroplets can be obtained using microfluidic systems (Christopher, GF & Anna, SL. Microfluidic methods for generating continuous droplet streams. Journal of Physics D - Applied Physics 40, pp. R319-R336 (2007)).

[0120] Advantageously, in the method of the present invention, the average number of nucleic acid molecules per microcompartment after division (corresponding to the ratio between the sum of the individual nucleic acid molecule contents of the microcompartments and the total number of microcompartments) is 0 to 10, preferably 0 to 5, more preferably 0 to 2, and most preferably 0.1 to 1.

[0121] This can be achieved by using Poisson partitioning, i.e., by adjusting the concentration of nucleic acid molecules containing at least one candidate sequence encoding at least one candidate polypeptide of the library in the composition prepared in step b) so that a small number of microcompartments contain multiple molecules. In this setting, many microcompartments contain no nucleic acid molecules (these are referred to as empty microcompartments), and most of the non-empty microcompartments contain a single nucleic acid molecule. For example, if the concentration of the nucleic acid molecules (preferably replicators) of the library in the composition prepared in step b) is set to 0.63 pM and droplets with a diameter of 10 μm are generated, the average number of nucleic acid molecules (preferably replicators) per droplet will be 0.2. Under such conditions, Poisson partitioning is predicted, with 82% of the droplets being empty, 16% containing one nucleic acid molecule (preferably replicator), and approximately 2% containing two or more nucleic acid molecules (preferably replicators). These conditions are suitable for uHTS screening experiments because most (16 / (16+2)=88%) of the non-empty droplets harbor only one nucleic acid molecule (preferably a replicator) and are therefore clonal in the sense that they express many copies of the polypeptide originally encoded by a single nucleic acid variant (and thus make many copies of the nucleic acid variant).

[0122] In a preferred embodiment of step d), the mixture obtained in step c) is partitioned into microcompartments using Poisson partitioning (i.e., Poisson distribution), such that preferably a substantial proportion of the microcompartments contain a single copy of the nucleic acid molecule, more preferably, most of the microcompartments contain a single copy of the nucleic acid molecule, and even more preferably, essentially all non-empty microcompartments contain a single copy of the nucleic acid molecule. Poisson partitioning allows for clonality, i.e., ensures that only one copy of the nucleic acid molecule (i.e., unique variants of the library) is incorporated into the majority of microcompartments.

[0123] Depending on the type of compartment and the method of manufacture, the partitioning of nucleic acid molecules may deviate from a Poisson distribution (e.g., a power-law distribution). In such cases, experiments can be designed to obtain empirical validation that a small fraction of microcompartments contain multiple nucleic acid variant molecules upon production.

[0124] Steps e), f), and g) - amplifying the candidate sequences by a nucleic acid replication mechanism, optionally transcribing the amplified candidate sequences of step e) by a nucleic acid transcription mechanism, and translating the amplified candidate sequences of step e) or the transcribed candidate sequences of step f) into at least one polypeptide by a nucleic acid translation mechanism. In steps e), optional step f), and step g), the microcompartments are incubated under conditions suitable for the replication, optional transcription, and translation machinery to perform their functions, resulting in the amplification of both the candidate nucleic acid sequence encoding the candidate polypeptide and the candidate polypeptide itself.

[0125] Suitable conditions will depend on the particular replication, optional transcription, and translation machinery selected and are known in the art.

[0126] For example, when bacteriophage Phi29 p2, p3, p5, and p6 proteins are used as the replication mechanism, the conditions used in step e) are in the presence of 20 mM ammonium sulfate for 1 hour to 16 hours, preferably at a temperature in the range of 20°C to 50°C, more preferably 25°C to 37°C.

[0127] In steps f) and g), when the PURE system is used as the transcription-translation system, the conditions are preferably those recommended by the manufacturer of the PURE system. Specifically, the reaction time may be 1 to 16 hours, and the reaction temperature may be less than 40°C, more preferably in the range of 25 to 37°C.

[0128] Step h) - Optionally, adding at least one compound to the mixture of step c) or to the microcompartment of step g) and / or modifying at least one parameter in the microcompartment of step d) or g). In step h), at least one compound may be added to the mixture of step c) or to the microcompartment of step g) and / or at least one parameter may be changed in the microcompartment of step d) or g).

[0129] For example, if the candidate polypeptide is a candidate enzyme, the at least one compound that may be added to the mixture in step c) or to the microcompartment in step g) may be selected from a reagent, substrate, cofactor, coenzyme (including prosthetic group and cosubstrate) of the candidate polypeptide encoded by the candidate sequence, and any combination thereof.

[0130] In certain embodiments, the activity of interest is polymer enzymatic degradation, and the at least one compound added to the mixture in step c) or to the microcompartment in step g) is a fluorescent compound (e.g., fluorescein acetate substrate (FDA), resorufin di-β-D-galactopyranoside (FDG), etc.), a chromogenic compound (e.g., 3,3'-diaminobenzidine tetrahydrochloride), a solvatochromic compound (Nile Red), or an environmentally sensitive compound (e.g., Bromocresol Green) embedded in one or more polymer particles, such that degradation of the polymer particles catalyzed by the candidate polypeptide releases the compound and converts it into a fluorescent or chromogenic product or changes its spectral properties in the microcompartment. Sorting can then be performed depending on whether the amount of fluorescent or chromogenic product detected in step i) is higher or lower than a user-defined threshold.

[0131] The polymer may in particular be a plastic polymer, in particular selected from PLA (polylactic acid), PET (polyethylene terephthalate), PE (polyethylene), PP (polypropylene), PCL (polycaprolactone) and PS (polystyrene).

[0132] If the candidate polypeptide is a candidate transcription factor, at least one compound that may be added to the mixture in step c) or to the microcompartment in step g) may be a nucleic acid molecule comprising a nucleic acid sequence encoding a reporter polypeptide (preferably a fluorescent or chromogenic polypeptide) under the control of said transcription factor.

[0133] Alternatively, or in addition, at least one parameter may be altered within the microcompartment of step d) to place this microcompartment under optimal conditions for replication in step e), optional transcription in step f), and translation in step g).

[0134] Alternatively, or in addition, at least one parameter may be altered within the microcompartment in step g) to place this microcompartment under optimal conditions for the detection of candidate polypeptide activity in the subsequent step i).

[0135] Examples of optimal conditions for steps e), f), g), and i) are disclosed in the sections related to the corresponding steps, so that one or more parameters may be adapted in step h) to meet these optimal conditions.

[0136] In particular, the at least one parameter that may be altered in the microcompartments of step d) or g) may be selected from concentration, temperature, pH, viscosity, and any combination thereof.

[0137] Step i) - detecting the activity of the candidate polypeptide of step g) in the microcompartment under conditions that are optionally modified in step h). In step i), the activity of the candidate polypeptide of step g) is detected in the microcompartment.

[0138] Any suitable method known in the art may be used for this detection.

[0139] In a preferred embodiment, the activity of the candidate polypeptide encoded by the candidate sequence is determined in step i) by measuring the intensity of one or more signals, which may be performed using one or more suitable techniques selected from optical techniques, electrical techniques (e.g., capacitance, resistance, impedance measurements, and the like), and magnetic techniques.

[0140] This detection can be a direct detection assay (preferably a fluorescent or colorimetric assay), i.e., the activity to be measured directly produces a detectable optical, electrical, and / or magnetic change in the microcompartment. As an example, if the candidate polypeptide is a candidate enzyme, a substrate containing an optical moiety can be used, and if the candidate enzyme has enzymatic activity, the reaction of the candidate enzyme encoded by the candidate sequence on the substrate will produce a detectable optical change in the microcompartment. Alternatively, a substrate conjugated / fused to a compound with detectable optical properties can be used. More specifically, the activity of the candidate enzyme encoded by the candidate sequence can be determined by detecting a product obtained from the substrate (e.g., a product converted from the substrate by the candidate enzyme or released from the reaction of the candidate enzyme with the substrate). A direct detection assay can also be used when the candidate polypeptide is a candidate reporter protein, such as a fluorescent or chromogenic protein.

[0141] Alternatively, an indirect activity assay (preferably a fluorescent or colorimetric assay) can be used, and the activity of the polypeptide encoded by the candidate sequence can be determined by detecting the product of the indirect activity assay. As an example, if the candidate polypeptide is a candidate transcription factor, a nucleic acid molecule containing a nucleic acid sequence encoding a reporter polypeptide (preferably a fluorescent or colorimetric polypeptide) having optical, electrical, and / or magnetic properties under the control of the transcription factor can be present in the microcompartment due to its presence in the first or second composition prepared in step a) or b) or by its addition in step h). In this case, the activity of the candidate transcription factor can be indirectly detected by detecting the reporter polypeptide, which is expressed only when the candidate transcription factor is active. As another example, the product released by the activity of the candidate enzyme can be used as a substrate by a second enzyme contained in the reaction mixture. The reaction of this second enzyme with the product of the candidate enzyme generates a secondary product that can be easily detected, for example, by spectroscopic methods, optionally in the presence of a specific cosubstrate. Many indirect enzyme detection assays (e.g., by fluorescent readout) have been described and are known.

[0142] As another example, the product released by the activity of a candidate enzyme may be conjugated to one or more macromolecules to enhance its fluorescence by irreversible photoactivation / photoconversion or reversible photoswitching (e.g., but not limited to, fluorescent biosensor proteins or aptamers) (Sam Duwe, Peter Dedecker, Optimizing the fluorescent protein toolbox and its use, Current Opinion in Biotechnology, Volume 58, 2019, pp. 183-191, ISSN 0958-1669; Koveal, D., Rosen, PC, Meyer, DJ et al. A high-throughput multiparameter screen for accelerated development and optimization of soluble genetically encoded fluorescent biosensors. Nat Commun 13, pp. 2919 (2022)).

[0143] It is preferred to use at least one optical method in step i), which is more preferably selected from absorption-based detection methods (e.g., single or multiple wavelength measurement, turbidity measurement, etc.), fluorescence-based detection methods (e.g., fluorescence intensity, fluorescence lifetime, fluorescence wavelength, fluorescence quenching / unquenching, BRET, FRET, relaxation, etc.), scattering-based detection methods (e.g., Rayleigh scattering or Raman scattering, nephelometry, laser diffraction, static light scattering, dynamic light scattering, etc.), or spectroscopy-based methods (e.g., infrared spectroscopy or Raman spectroscopy). More specifically, the optical method in step i) is preferably selected from the group consisting of absorption-based detection techniques (e.g., single or multiple wavelength measurement, etc.), fluorescence-based detection techniques (e.g., fluorescence intensity, fluorescence lifetime, fluorescence wavelength, FRET, etc.), and any combination thereof.

[0144] If the candidate polypeptide is a candidate reporter protein (e.g., a fluorescent or chromogenic protein), its activity can be detected directly using fluorescence-based or absorbance-based detection techniques (see Example 1 below).

[0145] If the candidate polypeptide is a candidate enzyme, fluorescence-based or absorbance-based detection techniques may still be used.

[0146] For example, fluorescence-based detection techniques may be used where a fluorescent substrate is used (see Example 2 below), or may be used where the candidate enzyme is fused to a fluorescent polypeptide (i.e., the candidate nucleic acid molecule comprises at least one candidate sequence encoding a fusion of the candidate polypeptide with a fluorescent polypeptide, see Examples 3-5 below).

[0147] Similarly, absorption-based detection techniques can still be used when a chromogenic substrate is used, or when the candidate enzyme is fused to a chromogenic polypeptide (i.e., the candidate nucleic acid molecule comprises at least one candidate sequence encoding a fusion of the candidate polypeptide with a chromogenic polypeptide).

[0148] For example, fluorescence-based detection techniques can also be used to detect the product of an enzyme, e.g., using a fluorescent probe (or a probe conjugated to a fluorescent compound), which is specific for the product of the enzyme (see, e.g., Example 8 below).

[0149] Similarly, absorption-based detection techniques can still be used if a chromogenic probe (specific for the enzyme's product) is used, or if the probe (specific for the enzyme's product) is fused to a chromogenic compound.

[0150] If the candidate polypeptide is a candidate regulatory protein, such as a transcription factor, fluorescence-based or absorption-based detection techniques may still be used. For example, a nucleic acid molecule comprising a nucleic acid sequence encoding a reporter polypeptide (preferably a fluorescent or chromogenic polypeptide) under the control of said transcription factor may be present in the microcompartment due to its presence in the first or second composition provided in step a) or b), or by addition in step h).

[0151] The method of the present invention can be a flow-based method (e.g., sequential detection in microcompartments) or a non-flow-based method (e.g., parallel detection in microcompartments). Flow-based methods are preferably performed using a microfluidic device (e.g., a fluorescence-activated droplet sorting (FADS) device (Baret, J.-C. et al., Fluorescence-activated droplet sorting (FADS): efficient microfluidic cell sorting based on enzymatic activity. Lab on a chip 9, pp. 1850-1858 (2009)) or a fluorescence-activated cell sorting (FACS) device. Non-flow-based methods include immobilizing liposomes on a flat surface, each on a small electrode, so that the activity of the protein to be detected changes the electrical properties of the droplets, thereby inducing the release of vesicles into solution.

[0152] Step j) - using a selection device to recover candidate sequences encoding polypeptides having the desired activity from the microcompartments that show specific activity signals; In step j), candidate sequences encoding polypeptides with the desired activity (detected in step i)) are recovered from the microcompartments that exhibit a specific activity signal using a sorting device.

[0153] The detection device used in step i) to detect the activity of the candidate polypeptides and the sorting device used in step j) are generally contained in a single device containing several modules (hence the term detection / sorting device). In this case, the detection / sorting device combines a sensor module that detects the activity of interest in each microcompartment and a sorting module that separates the contents of the microcompartments according to the activity level measured by the sensor module based on a user-defined threshold. The sorting module includes software capable of comparing the activity level measured by the sensor module with the user-defined threshold, and means for separating the contents of the microcompartments according to whether the activity level measured by the sensor module is above or below the user-defined threshold.

[0154] This threshold value is preferably at least equal to, and preferably higher than, the activity level measured for a control polypeptide. Depending on the desired stringency of the screen and the type of polypeptide being screened, this threshold value may be, for example, at least 1.2, at least 1.3, at least 1.4, at least 1.5, at least 1.6, at least 1.7, at least 1.8, at least 1.9, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 10 3 , at least 10 4 , at least 10 5 , or at least 10 6 Twice as expensive.

[0155] For directed evolution, this threshold value is preferably at least equal to, and preferably higher than, the activity level measured for the original polypeptide from which the candidate polypeptide is derived (any of the values ​​described above).

[0156] For bioprospecting, this threshold value is preferably at least equal to, and preferably higher than, the activity level measured for a known polypeptide (any of the values ​​described above), and more preferably higher than the activity level measured for the polypeptide with the best known activity level.

[0157] Any suitable sorting device may be used, depending on the particular type of microcompartment being used.

[0158] For example, when microdroplets (e.g., water-in-oil microdroplets) are used as the microcompartments, a fluorescence-activated droplet sorting (FADS) device can be preferably used. When liposomes are used as the microcompartments, a fluorescence-activated cell sorting (FACS) device can be preferably used.

[0159] Such sorting equipment is commercially available and should be used according to the manufacturer's recommendations.

[0160] Preferably, the sorting device pools the contents of all microcompartments where the activity level measured by the sensor module is higher than a user-defined threshold, whereupon the nucleic acid molecules in this pool can be analyzed or, for directed evolution, used as the starting library for the next directed evolution cycle.

[0161] Use of the screening method according to the present invention Two particularly interesting uses of the screening method according to the invention are directed evolution and bioprospecting (also called functional metagenomics).

[0162] In the case of directed evolution, the initial library (consisting of the second composition that is mixed with the first composition and then partitioned into all microcompartments) is typically derived from one or a small number of candidate sequences (e.g., natural or artificial sequences encoding proteins with known properties), which are further diversified by mutagenesis (in particular, by shuffling, random mutation, targeted mutation, combinatorial mutation, or any combination thereof).

[0163] Therefore, the present invention also relates to the use of the screening method according to the invention in a directed evolution method, which preferably comprises multiple cycles of mutagenesis (in particular obtained by shuffling, random mutation, targeted mutation, combinatorial mutagenesis, or any combination thereof), followed by screening using the screening method according to the invention.

[0164] For example, the present invention also provides a method of directed evolution comprising: a) generating a library of nucleic acid molecules each encoding a candidate polypeptide by mutagenesis (in particular by shuffling, random mutation, targeted mutation, combinatorial mutation, or any combination thereof) of one or more nucleic acid molecules each encoding a known polypeptide having a desired activity; b) screening the library generated in step a) for one or more nucleic acid molecules encoding a polypeptide having the desired activity using a screening method according to the invention; c) mutagenesis of one or more nucleic acid molecules selected in step b) (in particular obtained by shuffling, random mutation, targeted mutation, combinatorial mutation, or any combination thereof) to generate a library of nucleic acid molecules each encoding a candidate polypeptide; d) screening the library generated in step c) for one or more nucleic acid molecules encoding a polypeptide having the desired activity using a screening method according to the invention; and e) repeating steps c) and d) until one or more nucleic acid molecules encoding polypeptides having the desired activity of interest are obtained. The present invention also relates to a method comprising:

[0165] In the specific case of directed evolution of enzymes capable of degrading polymers, the directed evolution method can begin with one or more nucleic acid molecules, each encoding a known enzyme with a desired activity. Such enzymes are known in the art. For example, known enzymes capable of degrading PLA (polylactic acid) include cutinases (e.g., Humicola insolens cutinase (HiC)). Known enzymes capable of degrading PET (polyethylene terephthalate) include several proteases and cutinases, or IsPETase from Ideonella sakaiensis. Known enzymes capable of degrading PE (polyethylene) and PP (polypropylene) include cytochrome P450. Known enzymes capable of degrading PCL (polycaprolactone) include Humicola insolens cutinase (HiC) and proteases. A known enzyme capable of degrading PS (polystyrene) is hydroquinone peroxidase HM121 from Azotobacter beijerinckii.

[0166] In functional metagenomics applications, the initial library is typically a set of natural sequences obtained from an environment where genes encoding polypeptides with desired properties are likely to be present. For example, if one is looking for biomass-degrading enzymes, it may be interesting to use a metagenomic library obtained from the intestine of a ruminant.

[0167] Therefore, the present invention also relates to the use of the screening method according to the invention in functional metagenomics methods.

[0168] The present invention also provides a) generating a library of naturally occurring nucleic acid molecules each encoding a candidate polypeptide obtained from an environment likely to contain a gene encoding a polypeptide having a desired activity; and b) screening the library generated in step a) for one or more nucleic acid molecules encoding a polypeptide having the activity of interest using a screening method according to the invention. The present invention also relates to a functional metagenomics method comprising:

[0169] The screening method of the present invention can also be used in a method that starts with a first step of functional metagenomics and uses selected nucleic acid molecules in a second step of directed evolution, in which case the screening method of the present invention can be used to screen for the functional metagenomics step, the directed evolution step, or both steps.

[0170] kit The present invention also provides a kit (preferably a kit for carrying out the method described above), comprising: a)See below: (i) Nucleic acid replication mechanism, (ii) optional nucleic acid transcription machinery, and (iii) nucleic acid translation mechanism In vitro systems / platforms, including b) at least one substance for creating microcompartments, and c) Optionally, instructions for use The present invention also relates to a kit comprising, consisting essentially of, or consisting of:

[0171] In this kit according to the invention, the in vitro platform may be selected from any of the general or preferred embodiments disclosed above for the first composition provided in step a) of the method according to the invention.

[0172] Specifically, a preferred in vitro platform included in a kit of the present invention comprises both a bacteriophage Phi29-based replication mechanism (described herein) and a PURE transcription and translation system (described herein).

[0173] The kit according to the invention further comprises at least one material for producing the microcompartments.

[0174] The materials for creating the microcompartments may be selected from reagents for creating the microcompartments and devices for creating the microcompartments.

[0175] As a reagent for creating microcompartments, the kit according to the invention may comprise at least one dividing agent, which may in particular be selected from continuous phases and surfactants suitable for generating emulsions of microdroplets.

[0176] As a device for producing microcompartments, the kit according to the present invention may include a microfluidic microarray.

[0177] The kit may include both the reagents for creating the microcompartments and the device for creating the microcompartments. [Brief explanation of the drawings]

[0178] [Figure 1] Demonstration of the IVTTR reaction using PUREfrex 1.0 and the reconstituted replication machinery of phage Phi29 (P2 P3 P5 P6). Endpoint GFP fluorescence (gray bars) of 10 μL of IVTTR and final GFP fluorescence (gray bars) after 5 hours of incubation at 30°C. <gfp>A, B and C, D: Quantification by qPCR of low template concentrations (3 pM or 12 pM, respectively) <gfp>Replication is required to initiate high protein production from the ribosomal DNA (gene). E, F, G: Purified P2 and P3 proteins enable high replication and protein production while maintaining a balance between protein production and replication. [Figure 2] Endpoint GFP fluorescence (gray bar) of 10 μL of IVTTR 1.0 and final GFP fluorescence after 5 h of IVTTR at 30 °C. <gfp>Quantification by qPCR of the concentration (short black line). A-D: Increasing the amount of pUC57_OriLR_p2p3 plasmid has a non-monotonic effect on GFP yield, which is maximal at 200 pM of plasmid (B). <gfp>The effect on concentration is positive and saturates when the plasmid concentration exceeds 300 pM (C-G). The optimal pUC57_OriLR_p2p3 concentration is approximately 200 pM. [Figure 3] at a concentration of less than one copy per compartment <gfp>Brightfield images (A and C) and green fluorescent images (B and D) of an IVTTR 1.0 experiment initiated with a replicator. The dNTP-supplemented conditions (A and C) show numerous, discrete fluorescent droplets. <gfp>Replication and expression in droplets is shown for droplets harboring one or more copies of , while the absence of dNTPs (C and D) does not show any fluorescence. [Figure 4] After 8 hours of incubation at 33°C, <rfp>Endpoint mCherry fluorescence for 5 μL of IVTTR 2.0 containing replicators and varying amounts of PLA-FDL microparticles. Adjacent bars represent results from two replicates. Little difference in mCherry yield was observed, demonstrating the low toxicity of PLA-FDL microparticles to IVTTR 2.0. [Figure 5] After 8 hours of incubation at 33°C, <hic>Replicator (A-D) or <rfp>Steep slopes of green fluorescence appearance for 5 μL of IVTTR 2.0 containing replicators (E-H) and various amounts of PLA-FDL microparticles. Adjacent bars represent the results of two replicates. <rfp>Replicators show low fluorescein hydrolysis, but <hic>Replicators show that fluorescein hydrolysis increases with substrate concentration. [Figure 6] 1 pM in a 40 μm droplet <rfp-hic>Bright-field images (A, D, and G) at 4x magnification, as well as green (B, E, and H) and red (C, F, and I) fluorescent images, are shown for IVTTR 2.0 experiments (A-I) or PUREfrex 2.0-only experiments (G-I) initiated with a replicator. The dNTP-supplemented conditions (A-C) show numerous, discontinuous fluorescent droplets. <rfp-hic>Replication and expression in droplets is shown for droplets harboring one or more copies of . Colocalization of green and red fluorescence indicates that cutinase activity is associated with protein production. Conditions without dNTPs (D-F) do not show any fluorescence. Conditions with PUREfrex show a high background of green fluorescence (H) compared to IVTTR 2.0 without dNTPs (E), but not as high as IVTTR 2.0 with fluorescence. [Figure 7] Four-fold magnified green (A), red (B), far-red (C) fluorescence (corresponding to Hic esterase activity, mCherry fluorescence, and fluorescent dyes uniformly dispersed in the initial master mix for droplet tracking, respectively), and bright-field image (D) of an IVTTR 2.0 simulated library experiment after overnight incubation and before sorting. <hic>Replicator and <rfp>Starting with a 10:90 molar mixture of replicators, the majority of non-fluorescent droplets in the green and red channels exhibit a Poisson distribution with parameter λ << 1. The independence of green and red fluorescence confirms the Poisson distribution and implies clonality of the DNA after encapsulation. [Figure 8] Scatter plot of screened droplets according to far-red fluorescence (x-axis) and green fluorescence (y-axis). The horizontal line represents the threshold for green fluorescence above which droplets are sorted (positive hints). [Figure 9] qPCR quantification of the percentage of hic in gene libraries during mock library screening experiments. Standard deviations are calculated from technical triplicate qPCR quantifications. [Figure 10] A. Microscopic images of polyesterase activity in microdroplets containing wild-type HiC replicators fused to mCherry from IVTTR reactions performed at less than one copy per compartment (top: red fluorescence, bottom: green fluorescence). B. Histogram of green fluorescence values ​​(polyesterase activity) in non-empty droplets. C. Scatter plot of green (polyesterase activity) and red (protein concentration) values ​​for each droplet, showing a strong correlation. D. Histogram of the ratio of fluorescent signals for all non-empty droplets, showing that normalizing enzyme activity by the red expression signal resulted in a sharp peak of specific activity. [Figure 11] A) Flow cytometry analysis of YFP-expressing liposomes in a mock enrichment experiment. The initial library contained 1 nM of <yfp>Except for the control sample, the yfp gene and the minD gene were contained in a ratio of 1:10. <yfp>Starting from the IVTT sample, IVTTR exhibits higher YFP expression levels compared to IVTT samples. B) Gel electrophoresis analysis of DNA recovered by PCR from FACS-sorted liposomes. Full-length DNA was successfully recovered from IVTTR samples, but not from IVTT samples, using purified or expressed DNAP and TP proteins. C) Quantitative PCR data showing that both the yfp and minD genes are amplified in IVTTR samples, enabling DNA recovery. D) Percentages of yfp and minD genes in the library before and after FACS sorting using IVTTR in liposomes, as measured by qPCR. Two gating stringencies (sorts 1 and 2) were used, and two IVTTR constructs (containing P2 and P3 purified or expressed from the plasmid pUC57_OriLR_p2p3) were tested. The results show that liposomes sorted based on high YFP fluorescence contain up to 90% excess yfp gene content. The results demonstrate that functional gene variants can be enriched in liposomes after IVTTR when starting with picomolar amounts of DNA. [Figure 12] IVTTR is a non-catalytic and non-fluorescent protein. (A) Schematic of intraliposomal ori-p3 DNA amplification and expression by IVTTR. (B) Absolute quantification of ori-p3 DNA by PCR in lysed liposomes. A total input template DNA concentration of 10 pM was used, which was reduced due to exogenously supplied DNase I. (C) Population variation of dsGreen fluorescence in IVTTR liposomes measured by flow cytometry. (D) Quantification of the percentage of liposomes exhibiting dsGreen fluorescence above background, estimated by the diagonal gate in C. [Figure 13] Application of IVTTR to improve lipid synthesis by membrane-associated enzymes. (A) Schematic of CDP-DAG conversion to PS by PssA. (B) Schematic of IVTTR liposomes expressing the PssA enzyme and detection of PS-positive liposomes by LactC2-mCherry binding. (C) Absolute DNA quantification by qPCR of dissolved CADGE liposome samples. **P≦0.01; ***P≦0.001. (D) Box plot and (E) quantification of PS-positive CADGE liposomes expressing PssA as assessed by flow cytometry. [Figure 14] In the presence (A, C, E) or absence (B, D, F) of the DNK gene <gfp>GFP fluorescence (gray bars) and final replicator concentration (short black lines) in IVTTR 2.0 with various replicator and dNTP compositions. Conditions A and B contain equal concentrations of dATP, dCTP, dGTP, and dTTP. Conditions C and D replace dTTP with an equal concentration of dTMP, and conditions E and F contain no dNTPs. [Figure 15] Brightfield image (left) and fluorescent image (right) showing detection of DNK activity from a single copy of an encapsulated DNA fragment containing the DNK gene and a fluorescent protein. [Example]

[0179] Although the invention herein has been described with reference to particular embodiments, it is to be understood that these embodiments are merely illustrative of the principles and applications of the present invention. It is therefore to be understood that numerous modifications can be made to the exemplary embodiments and other arrangements can be devised without departing from the spirit and scope of the invention as defined by the appended claims.

[0180] Example 1 Coupled in vitro replication-expression reactions achieve high expression levels even starting from picomolar gene concentrations We have shown for the first time that a coupled replication-expression in vitro reaction can achieve high levels of protein expression, even when starting from very low concentrations of the encoding genes, but that a simple in vitro expression reaction cannot achieve these levels starting from the same low concentrations. Two different methods for introducing the replication machinery were tested: with purified P2 and P3 proteins, or by in situ generation of P2 and P3 from the encoding genes on the plasmid pUC57_OriLR_p2p3. In situ generation utilizes the protein expression machinery in the mixture to express some of the proteins of the replication machinery from their genes.

[0181] A replicator encoding GFP was constructed by combining, in order, the following nucleic acid sequences: Phi29 oriL, T7 promoter, E. coli ribosome binding site, gfp gene, terminator, and Phi29 oriR. The DNA was capped with a 5'-phosphate introduced by PCR using a 5'-phosphate primer. This molecule was then <gfp>(Hereafter, brackets <...> indicate DNA molecules that can replicate via the reconstituted Phi29 machinery.)

[0182] A typical coupled replication-expression reaction was performed using PUREfrex 1.0 (GeneFrontier) with an initial replicator DNA concentration of 3 pM or 12 pM in a 10 μL reaction volume. A 1 pmol replicator concentration corresponds to a concentration approaching one molecule per microcompartment, assuming the internal volume of the microcompartment is approximately 1.5 picoliters (e.g., the volume of a liquid with a diameter of 15-20 μm, a size typically used in droplet screening applications). The concentration of pUC57_OriLR_p2p3 was fixed at 200 pM under in situ protein production conditions. Under purified protein conditions, P2 protein (NEB) was fixed at 1 U / μL, and P3 (homemade) was varied from 0.8 mg / mL to 3.2 mg / mL.

[0183] Green fluorescence was monitored over 5 hours at 30°C in a standard 96-well plate fluorescence spectrometer (here we use a Biorad cfx thermocycler operated at constant temperature). The final green fluorescence values ​​(mature GFP production, grey bars) and the final <gfp>The concentration values ​​(short black lines) are reported in Figure 1. When no dNTPs were present in the mixture, <gfp>Replication is not possible and GFP fluorescence is not detected (A and C). In the presence of dNTPs, GFP fluorescence is readily detected, and the inventors <gfp>We measure an order-of-magnitude increase in concentration. This is true whether in situ-produced or purified P2 and P3 are used (B and D, and E, F, and G, respectively). In the case of purified P3 (F), the existence of an optimum concentration at which maximal protein production is 1.6 mg / mL P3 suggests a competitive effect between replication and protein production.

[0184] Next, we investigated the optimal concentration of the plasmid pUC57_OriLR_p2p3 for GFP production in the IVTTR system. Because the in situ synthesis of P2 and P3 also uses the PURE instrument, excessively high concentrations would adversely affect the final yield of the target protein (GFP). A typical IVTTR reaction was performed using PUREfrex 1.0 (GeneFrontier) and an initial concentration of 12 pM of GFP-encoding replicator DNA in a 10 μL reaction volume. The concentration of pUC57_OriLR_p2p3 was varied from 100 pM to 700 pM in 100 pM increments.

[0185] Green fluorescence was monitored over 5 hours at 30°C in CFX. The final green fluorescence values ​​(gray bars) and the final <gfp>The concentration values ​​(short black lines) are reported in Figure 2. Figure 2 clearly shows that GFP production increases and then decreases with increasing pUC57_OriLR_p2p3 concentration, with fluorescence maximal at 200 pM pUC57_OriLR_p2p3. Replication yield also increases with increasing pUC57_OriLR_p2p3 concentration and remains high even at the highest concentration (700 pM).

[0186] Using these optimal concentrations for the IVTTR system, we verified the detection of GFP fluorescence in microliter droplets, where the reaction was initiated from a unique copy of the DNA replicator. A typical IVTTR reaction was performed using PUREfrex 1.0 (GeneFrontier) and a fixed pUC57_OriLR_p2p3 concentration of 200 pM. Green fluorescence was monitored for 5 hours at 30°C in CFX. <gfp>The replicator density was set at 0.63 pM, generating droplets with a diameter of 10 μm, resulting in an average number of 0.2 replicators per droplet. Under these conditions, Poisson partitioning is expected, with 82% of droplets being empty, 16% containing one replicator, and approximately 2% containing two or more replicators. These conditions are suitable for uHTS screening experiments, as most non-empty droplets (16 / (16 + 2) = 88%) contain only one replicator and are therefore clonal. In this experiment, the medium was supplemented with yeast total RNA (Sigma-Aldrich, reference 10109223001) and Pluronic F127 (Sigma-Aldrich, reference P2443-250G) at final concentrations of 1 ng / μL and 0.4% (w / v), respectively. This master mix was divided into two aliquots: one supplemented with 0.3 mM dNTPs and the other with MilliQ water only. A flow-focusing device, adjusted to generate 10 μm diameter droplets, was used with fluorinated oil (HFE7500) supplemented with 64 mg / mL (4%) Fluosurf (Emulseo). The emulsion was incubated in a PCR machine at 30°C for 4 hours and then held at 4°C. The emulsion was spread between a hydrophobized slide and a cover glass, which was sealed with epoxy adhesive to prevent evaporation. Both samples were photographed at 10x magnification using the same parameters with an epifluorescence microscope (Nikon) equipped with a GFP filter. Fluorescence images were inverted for easier interpretation. Brightfield and fluorescence images are summarized in Figure 3. This figure shows that in the presence of dNTPs, only a portion of the droplets exhibited GFP fluorescence (C), whereas in the absence of dNTPs, no fluorescence was observed (D). This demonstrates that the replication machinery of the IVTTR system is necessary to generate detectable protein levels starting from a single DNA template in a microdroplet, whereas simple IVTT systems cannot.

[0187] Example 2 One-pot in vitro coupled replication-expression and fluorescence monitoring of enzymatic protein activity starting from pM amounts of replicator-form genes Humicola insolens cutinase (HiC) is an enzyme of the esterase family, more precisely, a cutinase, with hydrolytic activity toward various esters or polyesters. It is known to degrade cutin and polyesters. Herein, we use the solid polyester polylactic acid (PLA) as a model substrate. To evaluate the ability of the IVTTR system to replicate and express the enzyme to detectable activity and the possibility of detecting activity in situ with a fluorogenic substrate, we used the fluorogenic compound fluorescein dilaurate (FDL) embedded in PLA microparticles. Degradation of PLA by HiC cutinase releases FDL into solution, followed by its hydrolysis to fluorescein, which is also catalyzed by the esterase activity of HiC. Therefore, the appearance of green fluorescence indicates polyesterase activity and indirectly indicates the presence of HiC. Fluorescein has absorption and emission spectra similar to those of GFP, and we used mCherry protein as a reporter and negative control. The mCherry replicator is a "red fluorescent protein" <rfp>and HiC replicators are represented by <hic>It is written as:

[0188] A typical IVTTR reaction was performed using PUREfrex 2.0 (GeneFrontier) with an initial replicator DNA concentration of 10 pM. The reaction volume was set to 5 μL. The concentration of pUC57_OriLR_p2p3 was fixed at 200 pM. Green and red fluorescence were monitored over 8 hours at 33°C in CFX. IVTTR experiments were performed using the following methods: <hic>or <rfp>The experiments were carried out using a replicator. Plastic microparticles were supplemented to final concentrations of 4.3 mM, 2.1 mM, 1.1 mM, or 0 mM. Pipetting of each plastic concentration was performed twice. <rfp>Red fluorescence produced by the replicators was used as a control for plastic toxicity to the replication and expression reactions. <rfp>The baseline-corrected red fluorescence of the samples is shown in Figure 4. Figure 4 shows that the final mCherry fluorescence did not decrease with increasing amounts of PLA-FDL microparticles, thus indicating that the substrate was not significantly toxic to the IVTTR.

[0189] <hic>In tubes where replication and expression of was performed, green fluorescence increased strongly, indicating that the activity of the replicator-encoded enzyme could be detected, i.e., replication, expression, and enzyme assays could be combined in a one-pot condition. This fluorescence still increased after 8 hours of incubation, and samples with high substrate concentrations saturated the CFX sensor, so we evaluated plastic degradation activity by Vmax measurements. The steepest slope of the green fluorescence time trace is shown in Figure 5. Figure 5 shows the <rfp>Whereas the replicators showed low fluorescein hydrolysis, <hic>Figure 5 shows that replicators exhibit an increase in fluorescein hydrolysis with substrate concentration, demonstrating that hydrolysis is carried out by HiC cutinase synthesized in situ from the replicators. Figure 5 also shows that doubling the PLA-FDL concentration nearly doubles Vmax, consistent with the nonsaturating nature of the HiC enzyme.

[0190] Example 3 Unlike simple IVTT reactions, IVTTR reactions start with a single copy of the enzyme gene and produce detectable amounts of enzyme activity in microcompartments. This example demonstrates the use of a single IVTTR microcompartment containing a fluorescent substrate. <hic>We demonstrate that gene molecules can produce observable amounts of fluorescence, but not under IVTT conditions (i.e., no gene replication). The microcompartmentalization step used monodisperse water-in-oil microdroplets with internal volumes on the order of picoliters. In other words, this example demonstrates that gene replication is necessary to be able to perform uHTS screening in microcompartments directly from a library of linear genes.

[0191] To demonstrate this, we took advantage of the fact that the IVTTR system can amplify and express longer DNA constructs (e.g., containing multiple genes). Herein, we designed a DNA construct encoding two proteins, mCherry and HiC cutinase, and expressed it as a single fusion peptide. The linker between these two proteins contains the thrombin cleavage site LVPRGS flanked by the flexible peptide sequence GGGGS, although many other linkers would be suitable. Co-localization of protein production (red fluorescence) and cutinase activity (green fluorescence) demonstrated that cutinase activity was not detected. <rfp-hic>It will be shown that it depends on the existence of

[0192] A typical IVTTR reaction was performed using PUREfrex 2.0 (GeneFrontier) with an initial replicator DNA concentration of 1 pM in 40-micrometer diameter droplets. The concentration of pUC57_OriLR_p2p3 was set to 200 pM droplets. This master mix was split into two aliquots: one supplemented with 0.3 mM dNTPs and the other with MilliQ water only. The three solutions from the PUREfrex 2.0 kit and the 1 pM concentration of pUC57_OriLR_p2p3 were used. <rfp-hic>A third sample containing a replicator was generated. A flow focusing device adjusted to generate 42 μm (39 pL) droplets was used with fluorinated oil (HFE7500) supplemented with 64 mg / mL (4%) Fluosurf (Emulseo). The emulsion was incubated in the PCR machine at 33 °C for 16 h and then held at room temperature. The emulsion was spread between a hydrophobized slide and a cover glass, which was sealed with epoxy adhesive to prevent evaporation. Both samples were photographed at 4x magnification using the same parameters with an epifluorescence microscope (Nikon) equipped with a GFP filter. Fluorescence images were inverted for ease of interpretation. Bright-field images (A, D, G), green fluorescence images (B, E, H), and red fluorescence images (C, F, I) are summarized in Figure 6.

[0193] Figure 6 shows that in the presence of dNTPs (A–C), some but not all droplets exhibit green and red fluorescence indicative of a partial Poisson distribution (i.e., the initial distribution of replicator DNA molecules in the droplets is such that some droplets contain no replicator molecules, some contain only one, and some contain multiple), and that the reaction is initiated from a single DNA molecule. Most importantly, colocalization of green and red fluorescence can be observed (meaning that cutinase activity correlates with the mCherry protein), demonstrating the presence of the fusion protein in the droplets. In the absence of dNTPs and use of PURE 2.0 (D–F), no fluorescence is observed, demonstrating that replication is necessary in the IVTTR process to observe high protein production and enzyme activity. In the absence of dNTPs and use of PUREfrex (G–I), a slightly higher background green fluorescence (H) is observed compared to the absence of dNTPs and use of PURE 2.0 (E). However, this background was present in all droplets and its level was much lower compared to the positive droplets in IVTTR under PURE 2.0 conditions, which can be interpreted as a higher background hydrolysis of polyester particles under PUREfrex conditions.

[0194] This demonstrates that the replication machinery of the IVTTR system is necessary to generate detectable enzymatic activity starting from a single DNA template in a microdroplet, whereas simple IVTT systems cannot.

[0195] Example 4 Clonal expression in microdroplets followed by selection and enrichment of genes encoding active enzymes using IVTTR of single gene molecules in microdroplets containing fluorescent substrates We then tested the possibility of obtaining clonal expression in microcompartments when starting from a mixture (i.e., a library) of genes with different replicator forms. First, we prepared a mock library consisting of a mixture of two DNA sequences: one encoding an active polyesterase (cutinase HiC) and the other encoding a catalytically inactive protein (the fluorescent protein mCherry). Both sequences were inserted between the origin of replication and the expression elements (promoter, terminator, and RBS) as described above, so that they could be replicated and express the proteins at high levels. This initial library contained 10% <hic>Replicators and 90% <rfp>A replicator was included. The reaction mixture was prepared using the PUREfrex 2.0 kit and fluorogenic particles for detecting esterase activity. The reaction mixture was encapsulated in 47 μm droplets (54 pL volume) in fluorinated oil containing surfactant. A far-infrared fluorescent dye was added to the solution to allow detection of all droplets.

[0196] After overnight incubation, emulsion droplets were spread between a hydrophobized slide and a cover glass and sealed with epoxy adhesive to prevent evaporation. Images were taken at 4x magnification using an epifluorescence microscope (Nikon) equipped with a GFP filter. Fluorescence images were inverted for easier interpretation. Green (A), red (B), and far-infrared (C) fluorescence images, as well as a bright-field image (D), are summarized in Figure 7. The majority of non-fluorescent droplets in the green and red channels exhibit a Poisson distribution with parameter λ<<1 (λ corresponds to the average number of replicators per droplet). In addition, we observed a strong green signal in some droplets and a strong red fluorescent signal in others, suggesting high levels of pf protein expression. As expected, the green and red images do not overlap. The statistical independence of green and red fluorescence supports a Poisson distribution and clonal expression of genes encoded by the replicators of the library.

[0197] We then subjected this emulsion to a microfluidic droplet sorting method to explore the feasibility of recovering the gene encoding the active enzyme (and depleting this fraction from the gene encoding catalytically inactive mCherry). Droplets were sorted using a microfluidic droplet sorting device (also known as a FADS). A sorting gate was set according to green fluorescence, such that only droplets exhibiting significant green fluorescence (indicative of Hic esterase activity) were sent to the sorting gate. Figure 8 shows a scatter plot of green droplet fluorescence (y-axis) versus far-red droplet fluorescence (x-axis). The horizontal line represents the green fluorescence threshold above which droplets are sorted. 1 × 10 6 900 positive droplets were successfully sorted ("sorted" population) against 100 negative droplets ("discard" population).

[0198] Finally, to evaluate sorting efficiency, we used a PCR assay to measure the relative proportions of HiC and mCherry genes in the mock library, sorted bottles, and discard bottles. Sorted, discarded, and unscreened (mock) droplets were separately extracted in pure water using 1H,1H,2H,2H-perfluoro-1-octanol to break the emulsion. Gene-specific qPCR primers were used to quantify the amount of mCherry and HiC coding sequences in each sample. To ensure high fidelity in the HiC-to-mCherry ratio measurement, a fusion of the mCherry and HiC coding sequences was used as a unique standard. Figure 9 shows a summary of the qPCR data for the proportion of HiC genes in each sample. The initial HiC proportion remained unchanged between the mock library DNA mixture and the unscreened droplets, approximately 10%. The HiC proportion in the sorted population increased up to 54%, but decreased to 4% in the discard population. These results confirm that high levels of clonal expression demonstrate the ability of the method to distinguish between active and inactive variants and demonstrate the potential to enrich DNA populations with active variants directly from the encapsulation of linear DNA constructs.

[0199] Example 5 Quantification of specific activity levels in single DNA microdroplet-IVTTR assays initiated using hic-rfp fusion constructs The above example demonstrates that the IVTTR system can amplify and express longer DNA constructs, including those containing multiple genes, and can be used to significantly improve the determination of enzyme activity in microdroplets. Here, we used the rfp-hic fusion gene (rfp is the fluorescent protein mCherry, and hic is the enzyme) to measure specific enzyme activity rather than apparent enzyme activity as in previous examples. When IVTTR was performed in droplets from a single DNA template, two fluorescent signals were observed: one corresponding to enzyme activity and one (red signal) indicating the expression level of the construct. This allowed us to measure specific activity in the droplets. Indeed, total enzyme activity can be calculated as the product of specific activity and enzyme concentration. Because the concentration of expressed protein varies between microcompartments due to stochastic effects in the kinetics of the encapsulation and expression mechanisms, we observed a distribution of activity values ​​even when all droplets contained the same wild-type gene (Figure 10). In addition, expression levels may vary between variants, potentially obscuring differences in specific activity.

[0200] To assess specific activity from measuring enzyme activity (A) in a sample, the concentration of enzyme present in the sample (C) was evaluated, and the ratio A / C was calculated by computation. To directly measure enzyme concentration in microdroplets, a fluorescent tag was attached as an N-terminal fusion. This allows activity to be normalized to protein concentration, reducing droplet-to-droplet heterogeneity. A construct was created by genetically fusing mCherry fluorescent protein with HiC cutinase, connected by a flexible, cleavable linker. Both proteins are covalently linked, and each cutinase enzyme carries an mCherry tag, allowing us to measure cutinase concentration by the level of red fluorescence.

[0201] Fluorescent particles and fusion <rfp-hic>Esterase activity assays were performed in droplets using mCherry as the replicating DNA. IVTTR mixtures were constructed using the PURE 2.0 kit and fluorescent microparticles. The IVTTR solution was encapsulated in 42 μm (39 pL) droplets in fluorinated oil containing surfactant. The emulsions were incubated overnight at 33°C. The next day, the emulsions were imaged on microscope slides using appropriate filters for mCherry and fluorescein detection. <hic>and <rfp>A control experiment was performed under identical conditions using a mock library of replicator DNA. Microscope slides were incubated and imaged after 3 days to confirm that the green fluorescence was still increasing and therefore the reaction had not terminated during the first measurement.

[0202] Image analysis allowed extraction of two signals from each microdroplet: one in the red channel related to the concentration of the expressed fusion protein (Fig. 10A, bottom), and one in the green channel reporting polyesterase activity (Fig. 10A, top). The results for the green channel indicate that, although all active droplets contain the same gene, the levels of activity are not identical and show a large distribution (Fig. 10B). This heterogeneity is due to differences in enzyme expression levels in each compartment. However, the green signal was found to be strongly correlated with the red signal, indicating that most of the activity variation can be explained by droplet-to-droplet differences in enzyme concentration (Fig. 10C). When specific activity (i.e., the ratio of green to red) was calculated by computer, a narrow-peaked histogram was obtained, as expected since all droplets express the same wild-type gene (Fig. 10D).

[0203] Example 6 Enrichment of genes encoding fluorescent proteins using IVTTR in liposomes and FACS The possibility of selecting for active variants after direct encapsulation of linearized constructs was tested in another enrichment experiment using liposomes as an alternative compartmentalization strategy. In this experiment, we used a commercially available fluorescence-activated cell sorter (FACS) to sort active variants based on the fluorescence of expressed YFP in the presence of excess irrelevant template encoding a non-fluorescent protein. <mind>From DNA template <yfp>(yellow fluorescent protein). <yfp>The DNA template was added in a 10-fold excess <mind>The IVTTR mixture was mixed with the template and encapsulated in liposomes at a total DNA concentration of 10 pM (λ = 0.2 in this case). At such a low template DNA concentration, YFP expression was very low compared to the higher DNA concentrations typically used in cell-free reactions, resulting in a low signal-to-noise ratio. In contrast, liposomes in the IVTTR condition exhibited higher levels of YFP fluorescence (Figure 11A). Two stringent conditions were tested for this sorting gate: "Gate 1" encompassing the top 1% of all liposomes (applied to both IVTT and IVTTR samples), and "Gate 2" including only the top (0.2%) of high-intensity liposomes (applied to IVTTR samples only). While it was difficult to reproduce full-length DNA recovery by PCR from unamplified liposome samples, full-length DNA was readily recovered from IVTTR-treated liposomes (Figure 11B). This finding may be explained by the higher DNA titer in liposomes sorted from IVTTR samples. Indeed, when assayed by qPCR, the IVTTR liposomes <yfp> / <mind>The mixtures were amplified to a similar extent (>100-fold) and uniformly (Figure 11C). Furthermore, qPCR quantification of sorted liposome samples suggests that using the more stringent conditions of "Gate 2" in IVTTR samples improved the purity of YFP sorting compared to "Gate 1" in both IVTTR configurations (Figure 11D). These findings demonstrate that IVTTR enables efficient enrichment of genes encoding functional proteins and DNA recovery in a single FACS sort of liposomes.

[0204] Example 7 Compartmentalized IVTTR (in liposomes) for catalyzed, non-fluorescent proteins This example demonstrates that the indirect fluorescent readout method can be used to detect non-catalytic, non-fluorescent proteins separated from a single gene. Here, we demonstrate the terminal protein (TP) of phage phi29, encoded by the p3 gene. The p3 gene was inserted between the replication origins downstream of a T7 promoter. This template was introduced into the IVTTR mix at a concentration of 10 pM, supplemented with an excess amount of a plasmid encoding a DNA polymerase, and encapsulated in liposomes with a Poisson parameter of λ = 0.2 (Figure 12A). After incubation, quantitative PCR showed that the p3 gene was amplified by three orders of magnitude in the liposomes in the presence of dNTPs compared to the -dNTP control (Figure 12B). DNA amplification in single vesicles was assessed by flow cytometry using the DNA-intercalating dye dsGreen as a fluorescent marker. In the presence of dNTPs, a fraction of liposomes were detected with increased dsGreen fluorescence compared to background, corresponding to liposomes initially containing copies of the ori-p3 replicator from which DNA amplification had occurred (Fig. 12C,D). Liposomes with background levels of fluorescence corresponded to liposomes that initially did not harbor copies of the p3 gene.

[0205] Example 8 Compartmentalized IVTTR (in liposomes) for lipid synthesis by membrane-associated enzymes In this example, we demonstrate detection of IVTTR and encoded enzyme activity in liposomes containing an average of less than one replicator copy per compartment. We selected PssA from the E. coli Kennedy phospholipid biosynthesis pathway. This enzyme conjugates cytidine diphosphate-diacylglycerol (CDP-DAG) with L-serine to produce cytidine monophosphate and phosphatidylserine (PS), a precursor to phosphatidylethanolamine (Figure 13A). We assayed the activity of PssA enzyme synthesized intravesicularly from the pssA gene in the PURE system by preparing liposomes containing 5 mol% CDP-DAG (Figure 13B). PS lipid production was detected by external staining of liposomes with a PS-specific probe consisting of the C2 domain of lactadherin protein (LactC2) fused to a fluorescent protein such as mCherry or eGFP (Figure 13B) (Blanken et al., 2019). First, using qPCR, we confirmed 10- to 100-fold amplification of the pssA gene in both IVTTR configurations (i.e., using purified or expressed replication proteins) compared to the -dNTP control, where the input ori-pssA concentration was 10 pM (Poisson parameter λ = 0.2) (Figure 13C). By flow cytometry analysis of LactC2-eGFP-stained liposomes, we observed that the average intensity of the fluorescent signal (derived from recruited LactC2-eGFP) increased with functional IVTTR (+dNTP), even though low levels of PS synthesis were detectable in the -dNTP sample (Figure 13D). When IVTTR was applied, i.e., when gene expression was coupled to clonal expansion, the number of liposomes exhibiting a PS-positive phenotype was also higher (Figure 13E).

[0206] Reference: D. Blanken, D. Foschepoth, A. Calaca Serrao, and C. Danelon. Genetically controlled membrane synthesis in liposomes. Nat. Commun. 2020, 11(1):4317.

[0207] Example 9 Compartmentalized IVTTR (in microdroplets) for deoxyribonucleoside monophosphate kinase enzymes This example demonstrates the feasibility of applying the present invention to the screening of additional enzyme activities, demonstrating that this screening can be performed in microcompartments (picoliter scale) that are much larger than liposomes. This picoliter scale is typical of the water-in-oil droplets often used in directed evolution protocols based on microfluidic droplet sorting chips. Here, we selected deoxyribonucleoside monophosphate kinase from T5 phage (Uniprot:Q6QGP4). Its activity is the phosphorylation of deoxythymidine monophosphate (dTMP). In the IVTTR reaction, in which a nucleoside triphosphate (here, dTTP) is replaced by its monophosphate counterpart (dTMP), the DNK enzyme can transfer the terminal phosphate of an ATP molecule to dTMP to generate deoxynucleoside diphosphate (dTDP). This dTMP is converted to dTTP by nucleoside diphosphate kinase (NDK) and adenylate kinase (ADK, myokinase) using ATP as the phosphate donor. Both enzymes are present in the PURE system, so their addition is not necessary. The newly generated dTTP can ultimately be used by Phi29 polymerase to amplify DNA flanked by replication origins. The replicating DNA contains a gene for a fluorescent reporter protein, which is expressed in high yield to provide a fluorescent readout.

[0208] The bulk experiment reported in Figure 14 shows that in situ expression of the DNK gene can induce DNA amplification in IVTTR when dTTP is replaced by dTMP, whereas DNA amplification does not occur in the absence of the DNK gene. The T5 phage wild-type DNK protein coding sequence was inserted downstream of a T7 promoter, synthesized as a double-stranded DNA fragment (GeneStrand, Eurofins), diluted in ultrapure water (Merck Milli-Q), and used without further purification. This fragment is not flanked by replication origins and therefore is not recognized by the Phi29 replication machinery. PUREfrex 2.0 (GeneFrontier) was used and the only replicates were analyzed. <gfp>IVTTR reactions were prepared with a starting concentration of 12 mM. P2 and p3 were introduced via expression plasmids (the concentration of the pUC57_oriLR_p2p3 plasmid was fixed at 100 pM), and p5 and p6 purified proteins were maintained at final concentrations of 0.375 mg / mL and 0.105 mg / mL, respectively, and supplemented with 20 mM ammonium sulfate. Six conditions were generated by crossing the presence or absence of 25 pM of the DNK gene fragment and the composition of the dNTP mixture. The dNTP mixture consisted of either 0.3 mM of each of the four dNTPs, or 0.3 mM dATP, dGTP, dCTP, and dTMP, or no dNTPs. Fluorescence of 5 μL of the reaction mixture was monitored over 20 hours at 33°C in a CFX thermocycler. The Cal Gold 540 channel was used because the green fluorescence of the positive condition saturated the detector for the FAM channel (>65,000 RFU). The final fluorescence of the reaction mixture containing all dNTPs, regardless of the presence of the DNK gene (Figure 14A and B), and the condition in which dTTP was replaced with dTMP in the presence of the DNK gene (Figure 14C), was approximately 8.10 3 In the absence of all dNTPs and the DNK gene, when dTTP was replaced with dTMP (Fig. 14D), the final fluorescence was about 100 RFU. This result suggests that when dTTP is replaced with dTMP in IVTTR, <gfp>It is concluded that the presence of the DNK gene is required for replication and high fluorescence production.

[0209] The results shown in Figure 15 demonstrate that a single DNK gene dispersed in a microcompartment (here, a 15-μm-diameter droplet) can induce replication of the encoded DNA and express readily detectable levels of fluorescent protein. To ensure co-encapsulation of the DNK gene and the fluorescent protein gene, the gene encoding the DNK enzyme was fused at its C-terminus to the mCherry red fluorescent protein, a downstream T7 promoter, followed by a T7 terminator, and flanked by a replication origin. This structure was obtained by isolating a clonal plasmid and verifying its sequence by Sanger sequencing. The replicator was then amplified from the plasmid using PCR with phosphorylated primers, resulting in a replicable linear fragment. A typical IVTTR reaction was prepared using PUREfrex 2.0 (GeneFrontier). The concentration of pUC57_oriLR_p2p3 plasmid was fixed at 200 pM, and p5 and p6 purified proteins were maintained at final concentrations of 0.375 mg / mL and 0.105 mg / mL, respectively. The solution was supplemented with 20 mM ammonium sulfate. 2 μM Dextran, Alexa Fluor™ 647; 10000 MW (Invitrogen), and 2 μM Dextran Alexa Fluor™ 680; 10000 MW (Invitrogen) were added for pipetting control purposes and do not interfere with the reaction. The dNTP mixture consisted of dATP, dGTP, dCTP, and dTMP, each at a final concentration of 0.3 mM.

[0210] A replicator containing a single copy of the DNK gene and a single copy of the mCherry gene was added to a final concentration of 1 pM, resulting in a theoretical Poisson parameter λ of approximately 1 for 15 μm diameter droplets. This means that, on average, each droplet contains one copy of the replicator. A flow focusing device tuned to generate 15 μm diameter droplets was used with fluorinated oil (HFE7500) supplemented with 64 mg / mL (4%) Fluosurf (Emulseo). The emulsion was incubated overnight at 33 °C in a PCR machine. The emulsion was spread between a hydrophobized slide and a cover glass, which was sealed with epoxy adhesive to prevent evaporation. Images were taken at 4x magnification using an epifluorescence microscope (Nikon) equipped with the mCherry configuration (625 nm emission filter). Fluorescence images were inverted for easier interpretation. Bright-field and fluorescence images are summarized in Figure 15. The observation of individual fluorescent droplets indicates that the presence of a single copy of the DNK gene in picoliter droplets can induce gene replication and expression of a fluorescent protein in sufficient amounts to produce bright fluorescence.

[0211] Example 10 Materials and Methods for Examples 1-9 10.1. Bacterial strains and enzymes for molecular biology Enzymes and bacterial strains were purchased from New England Biolabs (NEB) unless otherwise specified. The strain for plasmid isolation was E. coli NEB® 5-alpha.

[0212] 10.2.Chemical products Chemicals such as ammonium sulfate (reference A4418-1KG) were purchased from Sigma-Aldrich unless otherwise specified.

[0213] 10.3. qPCR Equipment and Consumables All bulk fluorescence kinetics were monitored on a CFX96 instrument (BioRad) using white "Low-Profile 0.2 ml 8-Tube Strips" and "Optical Flat 8-Cap Strips" (BioRad #TLS0851 and #TCS0803, respectively).

[0214] Oligonucleotides All oligonucleotides except T22 and T23 were purchased from Eurofins in high-purity, salt-free grade purification, which were purified using HPLC grade purification. The sequences of the oligonucleotides are shown in Section 7.14 below (Table 1).

[0215] 10.5. Double-stranded DNA The starting dsDNA materials (OriLR, T7 promoter, and T7 terminator) listed in Table 2 below were submitted to Eurofins (Germany) for standard GeneStrand synthesis. Plasmid pUC57_OriLR_p2p3 was isolated from E. coui by intermediate preparation using the Plasmid Midi Kit (Qiagen). pR045-pIVEX-mCherry, pUC57_OriLR_gfp, pUC57_OriLR_rfp-hic_AC, and pEX_A128_pANobar were isolated from E. coli using the Plasmid Mini Kit (Macherey Nagel) and eluted with MilliQ water. All plasmids were Sanger sequenced for the regions of interest. These are reported in Table 4 below.

[0216] 10.6. Preparation of DNA constructs 10.6.1. PCR of phosphorylated DNA replicators All replicators were linear and phosphorylated to enable the P2P3 replication system. To obtain replicators, all DNA constructs were amplified by PCR using Q5® HotStart polymerase and the T22 / T23 phosphorylated primers listed in Table 1 below. Typical PCR conditions were as follows: In a PCR microtube, a fixed amount of DNA template (typically 1 fmol), 2 units of Q5® HotStart DNA polymerase (NEB), 0.2 mM dNTPs, and 500 nM forward and reverse primers were mixed in a total volume of 50 μL of 1× Q5 High Fidelity buffer. After an initial denaturation at 98°C for 30 seconds, 25 PCR reaction cycles were performed, each consisting of a denaturation step at 98°C for 10 seconds, a 69°C annealing step for 10 seconds, and an extension step at 72°C for 30 seconds per 1 kb of replicator DNA. The cycle was followed by a finishing step at 72°C for 2 minutes.

[0217] 10.6.2. <gfp>construct Cloned DNA encoding GFP protein was obtained by the typical PCR described above using 0.2 fmol of pUC57_OriLR_gfp and an extension time of 37 seconds, resulting in a single 1229 bp product visible on an agarose gel. The replicator was purified using a PCR cleanup kit (Macherey Nagel) according to the manufacturer's protocol, and the DNA was eluted in 30 μL of MilliQ water. The nucleic acid sequence of the gfp gene used is shown in Table 5 below.

[0218] 10.6.3. <hic>Constructs and <rfp>construct Replicative DNA encoding the HiC and mCherry proteins was constructed using Golden Gate Assembly (GGA). The nucleic acid sequences of the HiC (cutinase from the soft-rot fungus Humicola insolens (HiC) (AM Ronkvist, 2009)), rfp, and mCherry genes used, which are known to have esterase and polyesterase activity, are shown in Table 5 below. The GGA fragment was obtained by PCR using the primers listed in Table 1 below. A typical PCR reaction was performed in a total volume of 50 μL with 1 fmol of DNA template, 2 units of Q5® Hot Start DNA Polymerase (NEB), 0.2 mM dNTPs, and 500 nM of forward and reverse primers. After initial denaturation at 98°C for 30 seconds, the PCR reaction consisted of two cycles of a first step at a low annealing temperature, followed by 20 cycles of increasing the annealing temperature to 72°C. The cycle consisted of a 10-second denaturation step at 98°C, followed by a 10-second annealing step, followed by a 30-second extension step at 72°C per 1 kb of amplicon product. The low annealing temperature and extension times are reported in Table 2 below. After cycling, a 2-minute completion step at 72°C was performed. The reaction products were purified with a PCR cleanup kit (Macherey Nagel) according to the manufacturer's protocol, and DNA was eluted with 30 μL of MilliQ water. Golden Gate assemblies were performed with 50 fmol of each purified PCR product in a final volume of 20 μL using the BsaI-HF® v2 Golden Gate Assembly Kit (NEB). GGA was incubated for 15 minutes at 37°C, followed by denaturation at 60°C for 5 minutes, resulting in a circular assembly product designated pLR-mCherry or pLR-HiC. One microliter of the GGA product was used directly in PCR amplification of the DNA construct using the phosphorylated primers T22 and T23 described above, using a 45 second extension time. <hic>and <rfp>Single products of 1408 bp and 1507 bp were obtained, respectively. The PCR products were purified with a PCR clean-up kit (Macherey Nagel) according to the manufacturer's protocol, and the DNA was eluted in 30 μL of MilliQ water.

[0219] 10.6.4. Barcoded XOriLbar and XOriRbar DNA Fragments Using barcoded OriL and OriR DNA fragments, <rfp-hic>A replicator was constructed. First, 1 fmol of the plasmid pEX_A128_pAnobar was linearly amplified by PCR using the T93 and T360 primers. In a PCR microtube, 1 fmol of plasmid DNA, 2 units of Q5® Hot Start DNA Polymerase (NEB), 0.2 mM dNTPs, and 500 nM of forward and reverse primers were mixed in a total volume of 50 μL of 1X Q5 High Fidelity Buffer. After an initial denaturation at 98°C for 30 seconds, 18 PCR cycles were performed, consisting of a denaturation step at 98°C for 10 seconds, followed by an annealing step at 68°C for 10 seconds, and an extension step at 72°C for 60 seconds. A completion step at 72°C for 2 minutes resulted in a single product of 2072 bp visible on an agarose gel. The PCR products were purified using a Zymo-5 kit (Zymo Research) according to the manufacturer's protocol, and the DNA was eluted in 30 μL of water. A second round of PCR amplification was then performed using primers T22 / M75 or M74 / T23 for OriL and OriR, respectively. Primers M74 and M75 were barcoded with the degenerate sequence NNNNNWNNNNNWNNNNN, where N represents equal amounts of ACGT and W represents equal amounts of AT in the chemical synthesis (Table 1 below for the sequences of M74 and M75). In a PCR microtube, 0.3 fmol of plasmid DNA, 2 units of Q5® Hot Start DNA Polymerase (NEB), 0.2 mM dNTPs, and 500 nM of forward and reverse primers were mixed in a total volume of 50 μL of 1X Q5 High Fidelity buffer. After an initial denaturation at 98° C. for 30 seconds, 17 PCR cycles were performed consisting of a denaturation step at 98° C. for 10 seconds, followed by an annealing step at 55° C. for 10 seconds, and an extension step at 72° C. for 20 seconds. The cycles were followed by a completion step at 72° C. for 2 minutes, yielding single products of 346 bp and 343 bp for OriL and OriR, respectively.The PCR products were purified using a Zymo-5 kit (Zymo Research) according to the manufacturer's protocol, and the DNA was eluted in 30 μL of water. At this stage, half of the DNA molecules were labeled with mismatched barcodes. We performed only one round of PCR using primers T22 / M81 or T23 / M80 for OriL or OriL, respectively. Primers M80 and M81 are at the beginning of the degenerate sequence, so the polymerase completes the matched sequence during this cycle. 16 fmol (10 10 100 μL of 1× Q5 High Fidelity buffer was mixed with 100 μL of 1× Q5 High Fidelity buffer, 2 units of Q5® Hot Start DNA Polymerase (NEB), 0.2 mM dNTPs, and 500 nM of forward and reverse primers. The thermocycler parameters were 98°C for 40 seconds, 70°C for 10 seconds, and 72°C for 2 minutes and 10 seconds. After cycling, a 2-minute completion step at 72°C was performed, yielding single products of 346 bp and 343 bp for OriL and OriR, respectively. The PCR products were purified using a Zymo-5 kit (Zymo research) according to the manufacturer's protocol, and the DNA was eluted in 30 μL of water. These barcoded products were designated XOriLbar and XoriRbar.

[0220] 10.6.5. <rfp-hic>fusion construct Barcoded fusion protein replicators <rfp-hic>was obtained by Golden Gate Assembly from four different DNA fragments, designated XOriLbar, XOriRbar, prom_rfp, and hic_ter. The construction of XoriLbar and XoriRbar has been described above. promrfp and hic_ter were obtained by PCR on the plasmid pUC57_OriLR_rfp-hic_AC using primers M105 / M106 or M104 / M84 for prom_rfp and hic_ter, respectively. In a PCR microtube, 1 fmol of pUC57_OriLR_rfp-hic_AC plasmid, 2 units of Q5® Hot Start DNA Polymerase (NEB), 0.2 mM dNTPs, and 500 nM of forward and reverse primers were mixed in a total volume of 50 μL of 1X Q5 High Fidelity buffer. After an initial denaturation at 98°C for 30 seconds, the PCR reaction consisted of four cycles of an annealing temperature of 59°C in the first stage, followed by 18 cycles of increasing the annealing temperature to 72°C. A cycle consisted of a 10-second denaturation step at 98°C, followed by a 10-second annealing step, followed by a 30-second extension step at 72°C. A 2-minute completion step at 72°C was performed after cycling, resulting in single amplicon products of 813 bp and 756 bp for prom_rfp and hic_ter, respectively. Golden Gate Assembly (GGA) was performed as follows: 2.5 fmol of each of the four aliquots was mixed with 0.5 μL of BsmBI-v2 Golden Gate Enzyme Mix (NEB) in 10 μL of 1× ligase buffer in a PCR microtube. The GGA mix was incubated in a thermocycler for 30 cycles of 1 minute at 42°C and 1 minute at 16°C, followed by denaturation at 65°C for 5 minutes. One microliter of the GGA mix was directly used in a typical PCR amplification using the phosphorylated T22 / T23 primers described above, with an extension time of 66 seconds, resulting in a single product of 2199 bp. The PCR product was purified using a Zymo-5 kit (Zymo research) according to the manufacturer's protocol, and the DNS was eluted with 30 μL of water.

[0221] 10.7. Purified P5 and P6 Proteins DNA-binding proteins P5 and P6 were prepared according to (Soengas, Gutierrez, and Salas, 1995) and (Mencia et al., "Terminal protein-primed amplification of heterologous DNA with a minimal replication system based on phage Φ29", PNAS 108, pp. 18655-18660 (2011)), respectively, with the following stock concentrations and storage buffers: P5 (10 mg / mL in 50 mM Tris pH 7.5, 60 mM ammonium sulfate, 1 mM EDTA, 7 mM BME, 50% glycerol) and P6 (10 mg / mL in 50 mM Tris pH 7.5, 0.1 M ammonium sulfate, 1 mM EDTA, 7 mM BME, 50% glycerol). Proteins were aliquoted and stored at -80°C.

[0222] 10.8. In vitro transcription and translation (IVTT) of genes The in vitro transcription and translation systems (IVTT) PUREfrex 1.0 and PUREfrex 2.0 were purchased from Euromedex (France). The three main components of this kit are energy solution, enzyme solution, and ribosome solution. These were added to the final mixture as specified by the manufacturer: for PUREfrex 1.0, the ratio was 1 / 2 energy solution, 1 / 20 enzyme solution, and 1 / 20 ribosome solution. For PUREfrex 2.0, the ratio was 1 / 2 energy solution, 1 / 20 enzyme solution, and 1 / 10 ribosome solution. To ensure high protein production, mimicking successful replication, linear DNA was added at a final concentration in the nanomolar range (typically 2 nM).

[0223] 10.9. In vitro transcription, translation, and replication (IVTTR) The basics of IVTTR were the same as for IVTT. The same kit was used in the same proportions as above. To allow replication to occur, the solution was typically supplemented with 10 mM ammonium sulfate, 100 pM pUC57_OriLR_p2p3, 38 μg / mL purified P5 protein, 112 μg / mL purified P6 protein, and 300 μM dNTP solution (New England Biolabs). The remainder of the reaction mixture consisted of replicator DNA and various concentrations of substrate particles. In some cases, purified P2 and P3 were used instead of pUC57_OriLR_p2p3.

[0224] Any piece of DNA that can be replicated in the IVTTR system is referred to herein as a "replicator" and is labeled with a lowercase letter between two outward-facing chevrons (e.g., in the case of a red fluorescent protein reporter, <rfp>) To be recognized by the Phi29 replication machinery, the DNA construct must be double-stranded, linear, and flanked by active origins of replication. These origins, designated OriL and OriR, are typically located on the left and right sides of the construct, respectively. For expression in the PURE system, the coding sequence must be placed downstream of a T7 promoter. A T7 terminator is typically placed downstream of the gene. For Golden Gate Assembly, we used a promoter and terminator previously described in IVTTR experiments (Van Nies et al., "Self-replication of DNA by its encoded proteins in liposome-based synthetic cells", Nature Communications 9, 1583 (2018)) with some modifications, deleting the BsaI recognition site downstream of the T7 promoter sequence. The present inventors selected the +1 to +8 sequence GGGATAAT, which has a high transcription rate, according to (Conrad et al., "Maximizing transcription of nucleic acids with efficient T7 promoters". Commun Biol 3, 439 (2020)). <rfp> 、 <hic>, and <rfp-hic>Three replicators were designed, representing the red fluorescent protein mCherry, cutinase from Humicola insolens, and a fusion of these two proteins. The peptide bond between the two proteins consisted of the thrombin cleavage site LVPRGS (SEQ ID NO: 49) flanked by the flexible peptide GGGGS (SEQ ID NO: 50).

[0225] 10.10. Preparation of fluorescent PLA-FDL particles 100 mg of polylactic acid (PLA) (high molecular weight PLA LX175, Total Carbion) and 4 mg of fluorescein diacetate substrate (FDA) (ThermoFisher Scientific) were dissolved in 1 mL of dichloromethane (DCM, Sigma-Aldrich). This solution was added to 9 mL of water containing 5% (w / v) polyvinyl acid (PVA) (MW 13,000; Sigma-Aldrich). The mixture was then emulsified by sonication for 20 seconds in a 15 mL glass vial on ice. Solid PCL particles were formed by overnight evaporation of the DCM solvent at room temperature with magnetic stirring at 400 rpm. After evaporation of the DCM, the solid PLA-FDL (4%) particles were washed with three cycles of pelleting the particles by centrifugation at 10,000 g for 20 minutes, removing the particles, and adding 10 mL of clean MilliQ water to remove residual PVA. Optionally, an ethanol wash can be added to remove unencapsulated FDA. The particle suspension was then filtered through a 5 μm filter to remove potential aggregates.

[0226] 10.11. Quantitation of PLA-FDL Plastic Preparations Polyester concentration was quantified by the fluorescence signal obtained after complete degradation with proteinase K. Serial dilutions of the plastic solution diluted in activity buffer (100 mM Tris-HCl, pH 7.8) were mixed with 0.06 U Proteinase K (New England Biolabs) in a total volume of 10 μL in Low-Profile 0.2 ml 8-Tube Strips (BioRad). Fluorescence was monitored in a CFX instrument at 33 °C (lid temperature 50 °C) until a plateau was reached. The final fluorescein concentration was determined using a calibration curve of known concentrations of Dextran FITC 2 MDa (Invitrogen D7137) in the same buffer. Considering the complete degradation of available PLA and fluorescein dilaurate by Proteinase K and a 4% w / w ratio of fluorescein dilaurate to PLA, the polyester microparticle concentration was calculated according to the lactic acid monomer equivalent: C3H4O2 (MW 72.06 g / mol).

[0227] 10.12. Liposome Sorting <yfp>construct (encoding yellow fluorescent protein YFP) and <mind>The construct (encoding the non-fluorescent protein MinD) was mixed at a 1:10 molar ratio, and the IVTTR reaction mixture was assembled in either expression buffer (PUREfrex 2.0: 50% V / V Solution I, 5% V / V Solution II, and 10% V / V Solution III, with 0.6 units / μl Superase·In RNase inhibitor) or replication buffer (PUREfrex 2.0 supplemented with 20 mM ammonium sulfate, 300 μM dNTPs, 375 μg / ml purified P5 protein, 52.5 μg / ml purified P6 protein, 3 ng / μl purified P2 protein, 3 ng / μl purified P3 protein, and 0.6 units / μl Superase·In RNase inhibitor) at a final and total DNA concentration of 10 pM. The thoroughly mixed solution was encapsulated into liposomes by adding 10 mg of lipid-coated beads and rotating for 30 minutes at 4°C on an automated tube rotator (VWR). The mixture was then subjected to four freeze / freeze cycles. Ten microliters of the bead-free liposome suspension was then transferred to a PCR tube, mixed with 0.5 units of Proteinase K (Thermo Scientific), and incubated at 30°C for 16 hours. Three microliters of the liposome suspension was mixed with 497 μl of buffer and filtered through a 35 μm nylon mesh in a cell filter cap from a 5 ml round-bottom polystyrene test tube (Falcon).

[0228] Fluorescence-activated cell sorting was performed on a FACSMelody (BD Biosciences). Lasers were PE-CF594(YG) and FITC-BB515, a nozzle diameter of 100 microns, a pressure of 23.14 PSI, and a drop frequency of 34.2 kHz. The applied photon multiplier voltages were 320 V for forward scatter, 455 V for side scatter, 337 V for Texas Red, and 673 V for GFP, with a threshold of 359 V for side scatter. Liposomes with 1% of the highest YFP signal were sorted from liposomes prepared in expression buffer (Gate 1), and the same gate was applied to liposomes prepared in replication buffer (Sort 1) or a trimming gate containing only 0.2% of the highest YFP signal (Sort 2). Approximately 50,000 (sort 1) or 10,000 (sort 2) liposomes were sorted into 1.5 ml Eppendorf tubes. Liposomes in gate 1 were further concentrated by centrifugation at 12,000 g for 3 min and removal of ¾ of the supernatant volume. Proteinase K was heat inactivated at 95°C for 5 min.

[0229] The liposome suspension was then used as a template for PCR amplification using phosphorylated primers (ChD 491 / ChD 492). Reactions were set up in a 100 μl volume, containing 300 nM of each primer in Xtreme buffer, 400 μM dNTPs, 10 μl of diluted liposome suspension, and 2 units of KOD Xtreme Hotstart DNA Polymerase, and thermally cycled as follows: 94°C for 2 minutes to activate the polymerase, followed by 30 cycles of (98°C for 10 seconds, 65°C for 20 seconds, 68°C for 1.5 minutes). The amplified PCR fragment was purified using QIAquick PCR purification buffer (Qiagen) and RNeasy MinElute Cleanup columns (Qiagen) according to the manufacturer's guidelines for QIAquick PCR purification, except for a longer pre-elution drying step (10,000 g for 4 min in an open column) and a final elution step with 14 μl of ultrapure water (Merck Milli-Q). Purified DNA was quantified using a Nanodrop 2000c spectrophotometer (Isogen Life Sciences).

[0230] Arrays All sequences used in the above examples are listed in the table below.

[0231] [Table 1A]

[0232] [Table 1B]

[0233] [Table 2]

[0234] [Table 3]

[0235] Table 4

[0236] Table 5A

[0237] Table 5B

[0238]

Table 5C

[0239] Table 5D

[0240]

Table 5E

[0241] Table 6

[0242] Table 7 < / mind> < / yfp> < / hic> < / rfp> < / rfp> < / rfp> < / hic> < / rfp> < / hic> < / gfp> < / gfp> < / gfp> < / mind> < / yfp> < / mind> < / yfp> < / yfp> < / mind> < / rfp> < / hic> < / rfp> < / hic> < / hic> < / hic> < / rfp> < / hic> < / rfp> < / rfp> < / rfp> < / hic> < / hic> < / rfp> < / gfp> < / gfp> < / gfp> < / gfp> < / gfp> < / gfp> < / gfp> < / yfp> < / yfp> < / rfp> < / hic> < / hic> < / rfp> < / rfp> < / hic> < / rfp> < / gfp> < / gfp> < / gfp> < / gfp> < / gfp> < / gfp>

Claims

1. 1. A method for screening a library of nucleic acid molecules, each encoding a candidate polypeptide, for one or more nucleic acid molecules encoding a polypeptide having a desired activity, comprising: a)See below: (i) Nucleic acid replication mechanism, (ii) optional nucleic acid transcription machinery, and (iii) nucleic acid translation mechanism providing a first composition comprising an in vitro platform comprising: b) providing a second composition comprising a library of nucleic acid molecules, each comprising at least one candidate sequence encoding at least one candidate polypeptide; c) mixing the first composition of step a) with the second composition of step b); d) dividing the mixture so obtained into microcompartments, at least one of said microcompartments containing a single copy of said nucleic acid molecule; e) amplifying said candidate sequences in said microcompartments by said nucleic acid replication machinery; f) optionally transcribing the amplified candidate sequences of step e) in the microcompartments by the nucleic acid transcription machinery; g) translating the amplified candidate sequence of step e) or the transcribed candidate sequence of step f) into at least one candidate polypeptide by the nucleic acid translation machinery in the microcompartment; h) optionally adding at least one compound to the mixture of step c) or to the microcompartment of step g) and / or modifying at least one parameter in the microcompartment of step d) or g), wherein the at least one compound is preferably selected from a reagent, a substrate, a cofactor, a coenzyme, and any combination thereof, of the candidate polypeptide encoded by the candidate sequence, and / or the at least one parameter is preferably selected from a concentration, a temperature, a pH, a viscosity, and any combination thereof; i) detecting, preferably using optical methods, in said microcompartments, the activity of said candidate polypeptide of step g) under conditions that are optionally altered in step h); and j) using a selection device to recover the candidate sequences encoding the candidate polypeptides having the desired activity from the microcompartments that exhibit a specific activity signal; A method comprising:

2. 10. The method of claim 1, wherein the microcompartment is selected from the group consisting of a microdroplet, a microvesicle, a microchamber, and any combination thereof.

3. 3. The method of claim 1, wherein the nucleic acid molecule is selected from the group consisting of DNA, RNA, and any combination thereof.

4. The method of any one of claims 1 to 3, wherein the nucleic acid replication machinery comprises one or more substances selected from a DNA polymerase, an RNA polymerase, a DNA termination protein, a single-stranded DNA binding protein (SSBP), a double-stranded DNA binding protein (DSBP), a nucleic acid sequence encoding a DNA polymerase, a nucleic acid sequence encoding an RNA polymerase, a nucleic acid sequence encoding a DNA termination protein, a nucleic acid sequence encoding an SSBP, a nucleic acid sequence encoding a DSBP, and any combination thereof.

5. The method of any one of claims 1 to 4, wherein the nucleic acid transcription machinery comprises an RNA polymerase.

6. The method of any one of claims 1 to 5, wherein the nucleic acid translation machinery contains all the materials necessary to complete nucleic acid translation, including, for example, ribosomes, translation factors, aminoacyl-tRNA synthetases, and tRNAs.

7. The method according to any one of claims 1 to 6, wherein the in vitro platform is a cell-free system, such as a cell extract, a reconstituted in vitro system, and any combination thereof.

8. 8. The method of any one of claims 1 to 7, wherein the proteins comprised in or encoded by the nucleic acid molecules comprised in the nucleic acid replication machinery, optionally the nucleic acid transcription machinery, the nucleic acid translation machinery, or any combination thereof, are selected from or derived from a bacteriophage, a bacterium, a yeast, a virus, or any combination thereof.

9. 9. The method of any one of claims 1 to 8, wherein the activity of interest is polymer enzymatic degradation and the at least one compound added to the mixture in step c) or to the microcompartments in step g) is a fluorescent or chromogenic compound embedded in one or more polymer particles, and degradation of the polymer particles catalyzed by the candidate polypeptide releases and converts the compound into a fluorescent or chromogenic product.

10. 10. The method of any one of claims 1 to 9, wherein the activity of the candidate polypeptide encoded by the candidate sequence is determined by measuring the intensity of one or more signals, preferably using an optical technique selected from the group consisting of absorbance-based detection techniques, fluorescence-based detection techniques, and any combination thereof.

11. The method according to any one of claims 1 to 10, wherein the method is a high throughput screening (HTS) method, preferably an ultra high throughput screening (uHTS) method.

12. By the method, at least 10 5 This allows for high-throughput or ultra-high-throughput screening of at least 10 nucleic acid molecules, preferably at least 10 containing at least one candidate sequence. 6 12. The method according to any one of claims 1 to 11, which allows high-throughput or ultra-high-throughput screening of nucleic acid molecules.

13. The method of any one of claims 1 to 12, wherein the candidate polypeptide encoded by the candidate sequence is selected from the group consisting of a putative enzyme, a putative transcription factor, a putative functional derivative thereof, and a putative functional fragment thereof.

14. 14. The method of any one of claims 1 to 13, wherein the compound in step h) is selected from substrates, cofactors, and coenzymes (including prosthetic groups and co-substrates) of the candidate polypeptide encoded by the candidate sequence, or wherein the nucleic acid molecule further comprises a sequence encoding a protein substrate, and the substrate is selected from the group consisting of a substrate for an enzyme, a substrate for an enzyme cascade, functional derivatives thereof, and functional fragments thereof.

15. 15. The method according to any one of claims 1 to 14, wherein in step d), the mixture is preferably divided into microcompartments using Poisson partitioning, such that a substantial proportion of the microcompartments contain a single copy of the nucleic acid molecule.

16. a)See below: (i) Nucleic acid replication mechanism, (ii) optional nucleic acid transcription machinery, and (iii) nucleic acid translation mechanism In vitro systems / platforms, including b) at least one substance for creating microcompartments, preferably at least one reagent or device for creating microcompartments, more preferably at least one partitioning agent or at least one microfluidic microarray; A kit for carrying out the method according to any one of claims 1 to 15, comprising:

17. 16. Use of the method according to any one of claims 1 to 15 in a directed evolution method, said directed evolution method preferably comprising multiple cycles of mutagenesis followed by screening by the method according to any one of claims 1 to 15.