Methods for ultra-high-throughput profiling of nucleic acid binding or modifying proteins
The barcode and print approach allows for high-throughput profiling of nucleic acid binding or modifying proteins by linking pooled sequence libraries to specific proteins, overcoming the scalability and cost issues of traditional methods.
Patent Information
- Application Number
- PCT/US2024/060944
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-12-19
- Publication Date
- 2025-06-26
AI Technical Summary
Current methods for profiling nucleic acid binding or modifying proteins are costly and labor-intensive due to the need to express, purify, and assay each new protein variant individually, limiting the scalability to many proteins.
A barcode and print approach that links specific members of pooled sequence libraries to specific proteins, allowing for simultaneous expression, purification, and assessment of interactions between 10^2-10^3 different proteins and thousands of molecules, enabling high-throughput profiling of protein-sequence interactions.
Enables the measurement of 100,000+ protein-sequence interactions in a day with high-quality kinetic data, significantly reducing the cost and labor required compared to traditional methods.
Smart Images

Figure US2024060944_26062025_PF_FP_ABST
Abstract
Description
PATENT Attorney Docket No.: 110221-1475359-010610WO Client Reference No.: CZB-291S-PC METHODS FOR ULTRA-HIGH-THROUGHPUT PROFILING OF NUCLEIC ACID BINDING OR MODIFYING PROTEINS CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority benefit of U.S. Provisional Application No. 63 / 611,965, filed December 19, 2023, which is incorporated by reference in its entirety for all purposes. STATEMENT AS TO RIGHTS TO INVENTIONS MADE UNDER FEDERALLY SPONSORED RESEARCH AND DEVELOPMENT
[0002] This invention was made with Government support under contract 2142336 awarded by the National Science Foundation. The Government has certain rights in the invention. BACKGROUND
[0003] While next-generation sequencing has made it possible to assay nucleic acid binding or modification for a single protein interacting with thousands-to-millions of sequences simultaneously, scaling this up to many proteins has been prohibitively costly and labor- intensive due to a need to laboriously recombinantly express, purify, and assay each new protein variant.
[0004] Previous microfluidic platforms have been developed that can assess binding affinities for either a single transcription factor (TF) interacting with up to 106different sequences or for hundreds of recombinantly expressed and purified TFs interacting with ~10 different sequences. Similar microfluidic devices have also been shown to return accurate information about binding kinetics for a single TF interacting with hundreds of DNA sequences. BRIEF SUMMARY
[0005] It is understood that this summary features various aspects of the disclosure and is not provided as a comprehensive summary of all embodiments encompassed by the disclosure.1 30840863V.1
[0006] The present disclosure provides a barcode and print approach that links particular members of pooled sequence libraries to specific proteins that are able to interact with the particular library members; and devices that allow such library-on-library measurements. Accordingly, in some embodiments, the disclosure features a device and methods that provide the ability to simultaneously express, purify, and assess, e.g., quantify, interactions, e.g., between 10^2-10^3 different proteins, and thousands of molecules, e.g., tens of thousands of molecules, evaluated for interaction with such proteins, e.g., nucleic acids, which can produce measurements of 100,000+ protein-sequence interactions in a day. Moreover, the present disclosure makes it possible to obtain high-quality kinetic data about these interactions.
[0007] In one aspect, the disclosure thus provides a method of evaluating interactions of protein variants and members of a population of barcoded molecules, the method comprising: (a) providing a device comprising: a solid surface with an imprinted array comprising (i) a library of protein variants, e.g., homologs, mutagenized proteins, structurally related proteins, or a domain thereof; wherein variants are imprinted at discrete known locations on the solid surface to provide a localized site for each protein variant, wherein the position of the localized site identifies the variant protein; and (ii) a multiplicity of barcoded individual pools of a population of molecules comprising different molecules to be evaluated for interactions with variant proteins, wherein each individual pool comprises the population of molecules barcoded with a barcode specific for the pool; which barcode differs from the barcodes of other individual pools; wherein each individual pool is imprinted at a discrete known location on the solid surface; and a valved microfluidic device comprising a first set of imprint chambers, a set of reaction chambers wherein each reaction chamber has a valve controlling flow into the chamber and a second set of imprint chambers, wherein the set of reaction chambers is fluidly couplable to both the first and second set of imprint chambers, the first set of imprint chambers is physically aligned with the localized sites for each variant protein to provide a variant protein chamber for each variant; and the second set of imprint chambers is physically aligned with each individual pool of barcoded molecules to provide a barcoded pool chamber for each individual pool; (b) introducing the protein from each variant protein chamber into its connected reaction chamber;2 30840863V.1(c) introducing solubilized barcoded molecules into the reaction chamber from the connected barcoded pool chamber to allow formation of complexes between barcoded molecules and variant proteins; (d) incubating complexes formed in the reaction chambers with a capture agent that binds to the variant proteins; and (e) collecting barcoded molecules and determining the sequence of barcodes of captured barcoded molecules, thereby identifying molecules that bind to variant proteins. In some embodiments, captured molecules are eluted from the device and sequencing is performed on a pool comprising eluted captured molecules.
[0008] In some embodiments, each of the first set of imprint chambers comprises an expression construct and step (b) further comprises introducing a cell-free expression mixture into each of the variant protein chambers. In other embodiments, each of the first set of imprint chambers comprises a lyophilized variant protein and step (b) comprises solubilizing protein in each of the variant protein chambers.
[0009] In some embodiments, the method further comprises detecting the amount of barcoded molecules in individual pools that bind to proteins produced by the library of variant proteins. In some embodiments, the capture agent is immobilized to a surface in the reaction chamber. In some embodiments, the method further comprises washing the reaction chambers to remove unbound barcoded molecules. In additional embodiments, the method may further comprise collecting barcoded molecules present in complexes bound by the capture agent and determining the sequence of the barcodes. In further embodiments, the method further comprises collecting unbound barcoded molecules and determining the sequence of the barcodes of the barcoded molecules in the unbound fraction. For example, in some embodiments, collecting barcoded molecules comprises separately collecting fractions of unbound barcoded molecules that have dissociated over time from the immobilized protein at desired time points; determining the sequence of the barcodes in the unbound fraction at each time point; and quantifying the amount of barcoded molecule from each pool present in the unbound fraction at each time point. In some embodiments, the method further comprises collecting unbound barcoded molecules and determining the sequence of the barcodes of the barcoded molecules in the unbound fraction. In some embodiments, the method comprises quantifying the amount of protein captured in each3 30840863V.1reaction chamber. In some embodiments, the expressed proteins comprise a tag region, for examples comprising a fluorescent or luminescent protein that binds to the capture reagent, which can be, for example, an antibody that binds to the fluorescent or luminescent protein.
[0010] In some embodiments, protein variants comprise a nucleic acid binding protein and variants thereof. In some embodiments, the population of barcoded molecules comprises DNA oligonucleotides. In other embodiments, the population of barcoded molecules comprises RNA oligonucleotides. In addition, the population of barcoded molecules can comprise single- stranded oligonucleotides, double stranded oligonucleotides, or oligonucleotides comprising single and double stranded regions.
[0011] In some embodiments, protein variants comprise a library of nucleic acid modification enzymes and the pooled populations are substrate variants. In some embodiments, the modified and unmodified substrates are collected and sequenced in a pool. Thus, in some instances, protein variants need not be immobilized in the reaction chambers.
[0012] In some embodiments, variant proteins evaluated in accordance with the methods of the invention are transcription factor variants. In some embodiments, proteins evaluated in accordance with the invention are homologs or orthologs. In some embodiments, the proteins evaluated are synthetically designed proteins. In some embodiments, the population of barcoded molecules comprises candidate binding partners of a protein of interest. For example, in certain embodiments, the protein of interest comprises a binding domain of a receptor and the barcoded molecules comprise candidate ligands for binding to the receptor. In some instances, the variant proteins comprise a population of targeting binding sites for an antigen and the barcoded molecules comprise candidate antibodies. Thus, in some embodiments, barcoded molecules comprise proteins linked to a DNA or RNA sequence.
[0013] In a further aspect, the method employs a microfluidic device comprising a set of reaction chambers aligned in parallel with the first set of imprint chambers and aligned in parallel with the second set of imprint chambers, thereby providing an array of chambers in which each protein variant chamber is connected by a channel to a reaction chamber, which is connected by a second channel to a barcoded pool chamber. In some embodiments, the reaction chambers are fluidly connected via a through channel, each fluid channel divided by an actuatable through-channel valve, wherein each reaction chamber is isolatable from the other4 30840863V.1reaction chambers via actuation of through-channel valves in the through channel connecting each reaction chamber to the other reaction chambers, wherein introducing the protein from each protein variant chamber into its connected reaction chamber comprises actuating the through- channel valves to isolate the connected reaction chamber from the other reaction chambers. In some embodiments in which capture agents are employed, the method further comprises a step of washing each reaction chamber to remove uncaptured protein, wherein washing comprises actuating the through-channel valves to allow wash fluid to pass through the reaction chamber via the through channel. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] FIG.1 provides a schematic that illustrates a 2-layer microfluidic devices with 2 “encoding” / “print” chambers, each connected to a reaction chamber for library-on-library measurements.
[0015] FIG.2A-C illustrates DNA library design for STAMMP-Seq. A. Schematic of ordered ssDNA sequences. B. Addition of protein variant barcodes, printed on-chip. C. Schematic of STAMMP-Seq barcoded DNA libraries after library preparation, submitted for sequencing.
[0016] FIG.3A-B illustrates a workflow for assessing interactions of barcoded oligonucleotide pools with arrayed protein variants.
[0017] FIG.4 provides a illustrative STAMMP-Seq binding data.
[0018] FIG.5A-C illustrates STAMMP-Seq kinetics data.
[0019] FIG.6A-C. A: Schematic illustrating DNA library design for STAMMP-seq applied to large serine recombinase (LSR) enzymes and reaction scheme. Two DNA sequences were used to evaluate activity, a single "attP" sequence and a single "attB" sequence, both with illustrated adaptor sequences. AttP and attB are the recombination substrates for LSR enzymes. A=adaptor sequence, BC=barcode, attP & attB = attachment sites, recombination substrates, R2 = read2 adaptor sequence, bp=base pairs. B: Cell-free expression of LSR and variant on-chip quantified as mean intensity per-device chamber ± standard deviation. Significance markers indicate independent t-test comparing chamber identity to blank (no plasmid) chambers(***p < 0.001; **p < 0.01; *p < 0.05). C: Percent recombined quantified as total attL divided by5 30840863V.1total sampled attL and attP counts. Plot shows mean ± standard deviation percent recombined values across indicated device and / or barcode replicates. DETAILED DESCRIPTION OF THE INVENTION Terminology
[0020] A "polynucleotide" or “nucleic acid” or “oligonucleotide” can include any form of RNA or DNA, and includes nucleic acid molecules produced synthetically or by amplification. Polynucleotides may include chimeric molecules and nucleic acids comprising non-standard bases (e.g., inosine) or nucleotide analogs. For example, an oligonucleotide may contain naturally occurring nucleotides and / or analogs thereof. Polynucleotides may be single-stranded or double-stranded. Thus, an “oligonucleotide” as used herein can refer to duplexed oligonucleotides.
[0021] A “barcode” as used herein is a nucleic acid sequence that identifies a pool of molecules, such that one pool of molecules can be pooled with other pools of molecules, each pool having different barcodes, to identify the particular pool that is the source of a molecule following an interaction with a protein variant. A barcode need not be contiguous, but comprises one or more regions of identical sequence. A region of identical sequence is typically at least 4 nucleotides, preferably 5 or 6 nucleotides in length, and often greater, e.g., 10 or 18 nucleotides in length.
[0022] A “variant” of a protein as used herein refers to any population of proteins comprising different members to be studied for the ability to interact with a pool of barcoded molecules. Such members can include a protein and variants of the protein obtained from mutagenesis; homologs, orthologs, or paralogs; synthetically designed proteins; or any population of different proteins to be assessed for interaction with the barcoded pool molecules. Configuration
[0023] The present disclosure provides a microfluidic device that provides the ability to perform ultra-high throughput profiling of interactions of variants of a protein of interest with a library of barcoded molecules comprising variants of a molecule that interacts with the protein of6 30840863V.1interest. Pools of barcoded molecules can be assessed for the ability to interact with a protein of interest and its variants for quantification of binding interactions.
[0024] With reference now to FIG.1, one embodiment of the microfluidic device 100 is shown. The microfluidic device 100 includes one or several sets of chambers 102 that are fluidly connected to each other via channels 104. Three such sets of chambers 102 are shown in detail view, 106, which is a magnified view of a portion of FIG.1.
[0025] As used herein, the term “imprint” chamber refers to a chamber of the device that contains one of the components to be analyzed for interactions with another component. In one embodiment, a protein present in a first imprint chamber to be assessed with interaction with the barcoded population, e.g., a variant of a protein of interest, is encoded by an expression construct. In some embodiments an array of such constructs, each encoding a different variant protein, can be printed onto the base of a device, e.g., a coated glass slide, or the device itself, in a known configuration, before the device is assembled, such that the position of each variant is known. The imprinted nucleic acid constructs can subsequently be solubilized after assembly in a solution comprising components for in vitro expression added to the imprint chamber. Alternatively, proteins can be individually expressed and lyophilized onto the base surface of the device or the device itself in an arrayed pattern.
[0026] The second imprint chamber, also referred to herein as the “pool chamber”, contains a pool of barcoded molecules to be tested for interaction with proteins from the first chamber. Each individual pool chamber contains a population of molecules to be evaluated for interactions with protein variants, each molecule having the same barcode. The barcode differs in sequence, however, for each of the pool chambers, thereby allowing for the identification of the pool chamber that is the source of each barcode that is sequenced. Thus, the barcode information and positional information for the imprint chambers can be used to determine which protein variant(s) interacts with any given molecule from the pool chamber. In some embodiments, pool chambers contain populations comprising the same molecules. In other embodiments, individual pool chambers contain populations comprising different molecules. In some embodiments, a device may contain some pool chambers that each contain a population of the same molecules as well as pool chambers that contains populations of different molecules.7 30840863V.1
[0027] In one aspect, the microfluidic device 100 comprises an array of sets 101 of three chambers 102, also referred to herein as compartments, in which each set of three chambers 102 is configured such that two “imprint” chambers 102-B, 102-C are each connected to a single reaction chamber 102-A. The first “imprint” chamber 102-B contains a variant of a protein of interest and the second “imprint” chamber 102-C contains a pool of barcoded variants of a molecule, e.g., a polynucleotide or protein that interacts with the protein of interest. Each molecule contained within a pool of barcoded variants is barcoded with the same barcode sequence.
[0028] In some embodiments, and as seen in FIG.1, the reaction chambers 102, and specifically the set of reaction chambers 102-A is aligned in parallel with the first set of imprint chambers 102-B and is aligned in parallel with the second set of imprint chambers 102-C. In some embodiments, the reaction chambers 102-A and the imprint chambers 102-B, 102-C thereby provide an array of chambers 102 in which each first imprint chamber 102-B that can, in some embodiments, be a protein variant chamber, is connected by a channel to a reaction chamber 102-A, which reaction chamber 102-A is connected by a second channel to a second imprint chamber 102-C that can, in some embodiments, be a barcoded pool chamber.
[0029] These chambers 102-A, 102-B, 102-C are connected to each other by channels 104. Specifically, the first imprint chamber 102-B is connected to the reaction chamber 102-A via a first channel 104-A, and the second imprint chamber 102-C is connected to the reaction chamber 102-A via a second channel 104-B.
[0030] Further, each set 101 of chambers can be connected to one or several other sets of chambers. Specifically, and as shown in FIG.1, the first set 101-A is connected to the second set 101-B, and the second set 101-B is connected to the third set 101-C. As further seen, the sets 101 are connected via a through channel 108, and specifically are fluidly connected via the through channel 108.
[0031] The microfluidic device 100 further includes a plurality of actuatable components. These components can be actuated to affect flow through all or portions of the microfluidic device 100 and / or to achieve a desired outcome in all or portions of the microfluidic device 100. Each of these actuatable components are actuated by applying pressure to an inlet, set of inlets, and / or potentially an outlet. In some embodiments, the actuatable components can be8 30840863V.1pneumatically controlled via air pressure applied at the inlet and / or an outlet connected with the actuatable component.
[0032] These actuatable components can include, in some embodiments, three sets of valves to control fluid between chambers: ‘‘neck’’ valves, which separate imprint (e.g., plasmid) and reaction (e.g., binding) compartments; ‘‘sandwich’’ valves, which physically sequester reaction chambers from one another to prevent cross-contamination, and ‘‘button’’ valves in the binding compartment to enable selective surface patterning within the device and trap macromolecular binding interactions, e.g., at equilibrium for quantitative affinity measurements. In some embodiments, some or all of the actuatable components can have a distinct inlet and outlet via which the associated actuatable components can be controlled. In some embodiments, for example, the button valves can be controlled via one or several inlets fluidly coupled to the button valves. Likewise, the sandwich valves can be controlled via one or several inlets fluidly coupled to the sandwich valves, and / or the neck valves can be controlled via one or several inlets fluidly coupled to the neck valves.
[0033] These valves include a plurality of “sandwich” valves 110 located along the through channel 108, which sandwich valves 110 can divide the through channel 108 into one or several segments. In some embodiments, the through channel 108 can be configured to flow one or several liquids through the microfluidic device 100 and / or through one or several reaction chambers 102-A of the microfluidic device 100. In some embodiments, the through channel 108 can be fluidly coupled with an inlet via which fluid can be flowed into the through channel 108 and an outlet via which fluid can exit the through channel 108.
[0034] Each of the sandwich valves 110 has an open position in which flow can pass the sandwich valve 110 via the through channel 108 and a closed position in which flow cannot pass the sandwich valve 110 via the through channel 108. In some embodiments, the sandwich valves 110 are individually controllable, and in some embodiments, the sandwich valves 110 are controlled as a group.
[0035] These sandwich valves 110 can include a sandwich valve positioned along the through channel 108 before and after each set 101 of the chambers 102 such that that set 101 of chambers 102 can be isolated via the closing of the sandwich valves 110. Thus, as seen in FIG.1, the second set 101-B is isolatable from the first set 101-A via a first sandwich valve 110-A and the9 30840863V.1second set 101-B is isolatable from the third set 101-C via the second sandwich valve 110-B. Thus, in some embodiments, the sandwich valves 110 can divide the through channel 108 and isolate each reaction chamber 102-A from other reaction chambers 102-A. In some embodiments, this isolation of the reaction chambers 102-A can be performed as part of a process, for example, in some embodiments, introducing a protein from the first imprint chamber 102-B, which can be a protein variant chamber, into its connected reaction chamber 102-A can include actuating one or several sandwich valves 110 to isolate the connected reaction chamber 102-A from the other reaction chambers 102-A.
[0036] Within each set 101, a pair of valves allows isolation of each of the imprint chambers 102-B, 102-C from the reaction chamber 102-A. Specifically, a “neck” valve 112 can be positioned along the channel 104 extending between one of the imprint chambers 102-B, 102-C and the reaction chamber 102-A. As seen in FIG.1, this includes a first neck valve 112-A positioned along the first channel 104-A connecting the first imprint chamber 102-B and the reaction chamber 102-A, and a second neck valve 112-B positioned along the second chamber 104-B connecting the second imprint chamber 102-C and the reaction chamber 102-A. In some embodiments, the neck valves 112 can be independently controlled, or can be controlled in groups of neck valves 112. In one embodiment, for example, some or all of the first neck valves 112-A can be controlled as a group, and some or all of the second neck valves 112-B can be controlled as a group.
[0037] As seen in FIG.1, each set 101 of the microfluidic device 100 further includes a button valve 114. The button valve 114 can be located in the reaction chamber 102-A of the set 101. The button valve 114 can be moved from a first, open position to a second, closed position. In contrast to the sandwich valves 110 and the neck valves 112, the button valve 114 does not isolate any portion of the set 101 or isolate one set 101 from another set 101. Rather, in the closed position, the button valve 114 moves portions of the reaction chamber 102-A to a position proximate to each other and / or contacting each other. Specifically, in the closed position, a top of the reaction chamber 102-A is moved adjacent to a bottom portion of the reaction chamber 102- A. In some embodiments, the top of the reaction chamber 102-A can contact the bottom of the reaction chamber 102-A when the button valve 114 is in the closed position.10 30840863V.1
[0038] Upon pressurization, the configuration of the microfluidic device allows mixing of the contents of the two imprint chambers only with the contents of the reaction chamber to which they are connected. The position of each chamber in the array is known, as is the barcode content of each of the second imprint chambers. Thus, the use of the two compartments each connected to the single reaction chamber provides a 1:1 linkage between a protein variant contained in the first compartment to the particular barcode sequence of the contents of the second compartment. The combination of the positional information of the array components with the barcode information provides the ability to collect the contents of the reaction chambers following incubation of the protein variant of the first imprint chamber with the contents of the second imprint chamber for processing in bulk for sequence analysis.
[0039] In some embodiments, the reaction chamber contains a capture agent attached to the surface of the chamber that binds to the protein of interest, or variant thereof, thereby allowing the protein contents of the first chamber to be immobilized in the reaction chamber during binding interactions with the contents of the second chamber attached to the reaction chamber. Bound fractions can then be collected and combined to be processed for sequencing to identify the barcodes. Unbound fractions can also be processed for sequencing. A “capture” agent as used herein refers to an agent that comprises a binding agent, such as an antibody, that has an affinity for the desired target, e.g., a protein of interest and variants thereof. The capture agent thus allows for separation of the protein of interest, and, for example, any molecule bound to the protein of interest, from other components of a reaction, e.g., by a wash step.
[0040] Thus, in some embodiments, the button valve is closed and the reaction chamber is washed to remove uncaptured protein. In some embodiments, this further includes actuating the sandwich valves, and specifically opening the sandwich valves to allow wash fluid to pass through the reaction chamber via the through channel. Arrays
[0041] An array of the present disclosure contains any type of compartments to separate the contents of one compartment from another when valves of connecting channels are closed. In some embodiments, an array can comprise 500 or more compartments, or 1,000 or more compartments. In some embodiments, a microfluidic channel has hydraulic diameters ranging from 1 to 10,000 m, typically 5 to 1,000 m .11 30840863V.1
[0042] A chamber can be square or rectangular, or an alternative shape, e.g., a cylindrical shape. In some embodiments, a chamber can be from 1 μm-2 mm per diameter or per side for square or rectangular wells. In some embodiments, a chamber has a volume of from 1 pL to 10 L. In some embodiments, the chamber has a volume from about 50 nL to 5 L.
[0043] A microfluidic device can be of any suitable material. For arrays, a variety of suitable fabrication methods, such as wet etching, reactive ion etching, machining, photolithography, soft lithography (for example, multi-layer soft lithography), hot embossing, injection molding, laser ablation, in situ construction, or plasma etching may be employed. A non-limiting example of a suitable fabrication method uses multi-layer soft lithography, e.g., a described in the “Techniques Section”. Selection of a suitable fabrication method depends at least in part on the material to be used in the fabrication. Materials that may be in the fabrication include, but are not limited to, silicon, glass, quartz, polydimethylsiloxane (PDMS), polymethylmethacrylate (PMMA), thermoset polyester (TPE), polycarbonate (PC), cyclic olefin copolymer (COC), polystyrene (PS), polyvinylchloride (PVC), and polyethyleneterephthalate glycol (PETG). For example, in some embodiments, an array is produced using standard soft-lithography. In some embodiments, PDMS is employed for formulating the microarray.
[0044] One of skill in the art understands that prior to use, a microarray device may be treated so that the surface is suitable for imprinting and bonding.
[0045] Chambers can be arranged in a variety of configurations in which the first and second imprint chambers are connected by a valved channel to a single reaction chamber, such that the two imprint chambers are connected to only one reaction chamber. Thus, the chambers need not be arranged in parallel. Similarly, in some embodiments, the imprint chambers may be arranged serially with one imprint chamber connected to the second imprint chamber, which is in turn connected to the reaction chamber. Proteins to be evaluated for interactions with members of barcoded pools
[0046] Any protein that interacts with a substrate can be evaluated in accordance with the invention. In some embodiments, proteins to be assessed can be a protein of interest and mutagenized variants of the protein of interest. In some embodiments, such proteins can be randomly mutagenized in a desired region. In other embodiments, mutations can be introduced12 30840863V.1into specific sites. In some embodiments, proteins to be assess can be structurally or evolutionarily related, e.g., homologs, orthologs, paralogs; or synthetically designed protein sequences.
[0047] In some embodiments the protein is a nucleic acid binding protein. Examples of such proteins include double-stranded DNA binding proteins, single-stranded DNA binding proteins, RNA binding proteins, or any other protein that binds to a nucleic acid. In some embodiments, the protein is a transcription factor.
[0048] In some embodiments the protein is a nucleic acid modifying enzyme. For example, such a protein can a nuclease, a ligase, a polymerase, a recombinase, a DNA repair protein, a methylase, or other enzyme that acts on a nucleic acid substrate. Pooled populations Oligonucleotides
[0049] In some embodiments, a protein of interest interacts with nucleic acid molecules, e.g., DNA and / or RNA molecules. Accordingly, variants of such a protein can be analyzed for their ability to interact with a library of oligonucleotides, e.g., duplexed oligonucleotides, where portions of the library are distributed to an array such that second imprint chambers contain a population, or pool, from the library Each pool of an imprint chamber is barcoded with a sequence representative of the individual imprint chamber, which differs from barcodes of the pools in other imprint chambers.
[0050] For example, a library of oligonucleotides may comprise sequence variants of a nucleic acid binding site of a nucleic acid binding protein of interest. In the present disclosure, relative binding energies for one nucleic acid binding protein, e.g., a transcription factor, can be assessed, for example, for at least 105and preferably greater than 106, oligonucleotide sequences in parallel.
[0051] Similarly, in some embodiments, a library of oligonucleotides may comprise sequence variants of nucleic acid sequences that are modified by a nucleic acid modification enzyme. Again, many iterations of the nucleic acid sequence surrounding a modification site can be assessed for each variant enzyme.13 30840863V.1
[0052] In some embodiments, the population of barcoded nucleic acids can comprise candidate aptamers to be assessed for the ability of bind to a protein of interest and variants thereof. Barcodes
[0053] Unique barcodes for each population present in the “pool” imprint chambers can be generated using known methods, for example rational design using Hamming distances for maximal sequencing discrimination of barcodes. Common algorithms used for barcode design known (e.g., Bioconductor, DNA barcodes; Glenn et al., Plos ONE, 2012: doi:10.1371 / journal.pone.0042543).
[0054] In some embodiments, the number of permutations nucleotides is equal to or more than the number of protein variants of interest being evaluated in an assay. For example, in some instances, a typical assay will measure anywhere from 50-1000 variants and use barcodes 3-5 bases in length at minimum (yielding 43-45 or 64-1024 unique codes). Barcodes longer than the length minimally necessary for the number of variants are preferably used to provide more robust assignment relative to sequencing noise to reduce inadvertent mis-assigning a barcode. In some instances, barcodes that contain known binding motifs Thus, for typical experiments profiling 50-1000 protein variants, in the 8-12 base range can be used. In some embodiments, barcodes form 8-20 nucleotide in length can be used.
[0055] Barcodes can be incorporated in oligonucleotide members of a pooled population using well known methods, including for example, methods based on extension reactions, e.g., amplification reactions, ligation reactions and the like.
[0056] Barcoded oligonucleotides can comprise additional sequences, such as Unique Molecular Identifiers (UMI) that can be used to determine whether multiple reads with the same barcodes sequence originate from the same template molecule.
[0057] Additional components of barcoded oligonucleotides include adaptors, PCR primer sites, or other regions to facilitate sequencing.
[0058] Each population of barcoded oligonucleotides can be printed onto the bottom surface, e.g., glass slide, bottom surface of a device, or the device itself in a specific position in an array.14 30840863V.1Once the device is assembled, the oligonucleotides in each pooled chamber can then be solubilized. Polypeptide pools for the analysis of protein-protein interactions
[0059] In some embodiments, the barcoded pool can comprise a population of polypeptides, wherein each member of the pool is barcoded directly, e.g., via conjugation, or indirectly, via interaction of a moiety or sequence common to the population that interacts with a barcoded affinity (i.e., binding) agent. For example, in some embodiments, polypeptide members of a pool may comprise a tag to which a barcoded binding agent, e.g., an antibody, can bind thereby labeling members of a pool with the same barcode.
[0060] Thus, such a pool barcoded population can be a portion from a library of variant polypeptide to be assessed for the ability to bind to a protein of interest, and variants thereof, encoded by expression constructs in the first imprint chamber. For example, in some instances, a protein of interest may be a receptor and the pool population comprise peptide ligands to be assessed for the ability to bind to the protein of interest, or variant thereof. In some embodiments, the pool of barcoded polypeptides can comprise variants of a substrate region of an enzyme, such as a kinase or phosphorylase
[0061] The pool of polypeptides can be encoded by expression constructs printed onto a localized site and subsequently lyophilized for analysis after assembly of the device. In other embodiments, the pools are lyophilized and printed onto the surface of the device.
[0062] In some embodiments, antibody (including any configuration of antibody, e.g., nanobody, affibody, scFV and the like) binding interactions with a target can be evaluated to characterize antibody binding activity. In some embodiments, protein-peptide interactions can be evaluated, e.g., SLiM-containing peptides binding to a folded protein. In some embodiments, various protein domains, e.g., SH2, SH3, PDZ, WW domains and the like) can be evaluated for binding specificity for various target molecules.
[0063] In some embodiments, protein binding molecules, e.g., antibodies, can be assessed for binding to a pooled populations expressed by a display library, e.g., cDNA or mRNA display library.15 30840863V.1Illustrative Reaction Capture agent immobilized in reaction chamber
[0064] In some embodiments, microfluidic devices are aligned to printed arrays comprising discrete sites to which expression constructs encoding variants of a protein of interest are localized to provide a position for each variant in the array; and discrete sites to which barcoded pooled populations are localized to provide a position, relative to a variant of a protein of interest, for a pooled population, and hence the barcode that is characteristic of the pooled population.
[0065] As explained above, in some embodiments, the reaction chamber connected to the two imprint chambers contains a capture agent attached to the surface of the chamber that binds to the protein of interest, or variant thereof, thereby allowing the protein contents of the first chamber to be immobilized in the reaction chamber during binding interactions with the contents of the second chamber attached to the reaction chamber. Bound fractions can then be collected, e.g., eluted, and combined for processing for sequencing to identify the barcodes. Unbound fractions can also be processed for sequencing.
[0066] In some embodiments, reaction chambers are surface patterned for protein immobilization following assembly of the microfluidic device. In embodiments, in which expression constructs are printed to provide a set for first imprint chambers, constructs can be solubilized for producing protein by introducing an expression solution comprising reagents for in vitro expression. In some instances, such a solution can be introduced to simultaneously express protein variants by flowing the expression solution through the channels with neck and button valves closed. The outlet valve is then closed and button channels opened. The neck valve controlling access to the first set of imprint channels is then opened for a time sufficient to provide a sufficient volume, after which sandwich valves can be closed, followed by closing the button valves to mix the expression solution with solubilized plasmid. After a time sufficient to express the proteins, e.g., 30-120 minutes at 37ºC, the button valves are opened and the expressed protein can be captured by antibodies in the reaction chamber that bind to a site, e.g., a tag region incorporated into the expressed proteins. The captured expressed protein can then be sequestered in the reactions chambers by closing the button valves to wash the device to remove16 30840863V.1nonspecifically bound expressed protein. In some instances, e.g., where a fluorescent protein serves as a tag, the device can be imaged to quantify protein levels in each reaction chamber.
[0067] In order to solubilize DNA libraries, buffer, e.g., binding buffer is flowed through the device with all “neck” and “button” valves closed to completely fill all channels. The outlet valve is then closed and the button valves open to allow buffer to fill each reaction chamber. With the button and sandwich valves closed and neck valves for the pooled imprint chamber open, each pool is allowed to solubilize. The button valves can then be opened to allow the binding reaction to proceed.
[0068] After binding, unbound oligonucleotides can be eluted by flowing elution buffer through the device with buttons closed to collect a pool of unbound oligonucleotides. Any remaining unbound oligonucleotides in the microfluidic device can then be washed. Thus, in some embodiments, the button valve is closed and the reaction chamber is washed to remove uncaptured protein. In some embodiments, this further includes actuating the sandwich valves, and specifically opening the sandwich valves to allow wash fluid to pass through the reaction chamber via the through channel. The bound oligonucleotide fraction can then be eluted by opening the button valves and the collected fraction processed for sequencing. Alternative embodiment
[0069] In alternative embodiments, however, the reaction chamber need not have a capture agent attached to the surface. For example variants of a DNA-modifying enzyme that can modify oligos in solution can be tested. The enzymes themselves (and / or DNA encoding the enzymes) are lyophilized in the imprint chambers and then solubilized and released into the reaction chamber containing DNA oligonucleotide substrates already in solution. In this example, because detecting the modification via sequencing downstream doesn’t require separation into bound and unbound fractions, the protein does not need to be immobilized in the reaction chambers. Technical section
[0070] This technical section highlights certain technical features of the invention.17 30840863V.1
[0071] FIG.3A-B illustrates a workflow for assessing interactions of barcoded oligonucleotide pools with arrayed protein variants. The example of the methods described herein can be employed to recombinantly express and purify 900 transcription factors (TFs) and measure relative binding affinities for each to >10,000 different oligonucleotides. A. Experimental pipeline. Libraries of TF variants and barcoded dsDNA pools are printed on glass slides. Valved microfluidic devices can be aligned to these arrays, allowing high-throughput expression and purification of TF variants. Surface-immobilized TFs can then interact with solubilized barcoded dsDNA libraries. Pneumatic valves mechanically ‘trap’ bound DNA to allow washing out unbound material without loss of weak interactions. Sequencing bound and unbound materialallows counting molecules and quantifying relative binding energies ( Gs). B. Detailedschematic of TF expression, surface immobilization, and DNA binding.
[0072] FIG.4 provides illustrative STAMMP-Seq binding data obtained using methods described herein. Far left panel: Photograph of STAMMP-Seq microfluidic device. Second Panel: MAX crystal structure (PDB ID: 1HLO), motif, and details for initial TF and DNA libraries. Third Panel: Preliminary STAMMP-Seq data demonstrating ability to measure relative Gs for 1 TF interacting with 256 DNA sequences (third panel, left) and 4 TF variants interacting with a consensus motif sequence (third panel, right)
[0073] FIG.5A-C illustrates STAMMP-Seq kinetics data obtained using methods and devices as described herein. A. Experimental pipeline for quantifying dissociation of bound dsDNA from surface-immobilized TFs. B. Example dissociation data and exponential fits for dsDNA bearing cognate E- box sites C. Example dissociation data and exponential fits for dsDNA bearing cognate E-box sites interacting with MAX and Pho4.
[0074] FIG.6A-C provides data illustrating application of the compositions and methods of the current disclosure to the analysis of variant enzymes. FIG.6A is a schematic illustrating DNA library design for STAMMP-seq applied to large serine recombinase (LSR) enzymes. FIG.6B shows expression of LSR variants on-chip quantified as mean intensity per devicechamber. Significance indicators (***p < 0.001; **p < 0.01; *p < 0.05) reflect independent t-testcomparing chamber identity to blank (no plasmid) chambers. FIG.6C shows percent recombined quantified as total attL divided by total sampled attL and attP counts. The graph18 30840863V.1shows mean ± standard deviation percent recombined values across indicated device and / or barcode replicates.
[0075] Illustrative methods are further described below. Microfluidic device fabrication
[0076] STAMMP-Seq devices were designed in AutoCAD2023 (Autodesk, Inc.) and each layer was reproduced as a film photomask at 32,000 dpi (Fineline Imaging). Photolithography molds (Markin et al. Science, 2021, doi.org / 10.1126 / science.abf87612021) and microfluidic devices (Aditham et al. Cell Systems, 2020, doi.org / 10.1016 / j.cels.2020.11.012; Le et al., PNAS, 2018, doi.org / 10.1073 / pnas.1715888115; Fordyce et al., Lab on a Chip, 2012 doi.org / 10.1039 / c2lc40414a) ) were produced as described previously, with the following minor modifications to device fabrication protocols.
[0077] Two-layer microfluidic devices were then cast from these molds using polydimethylsiloxane (PDMS) polymer (RS Hughes, RTV615). Control layers of the microfluidic device were generated by pouring 60 grams of PDMS (1:5 ratio of cross-linker to base) onto the molds, degassing to remove all air in a vacuum chamber under vacuum for 45 minutes, and baking for 60 minutes at 80°C in a convection oven. After this step, control layers for each device were cut out and removed from the wafer and the fluid line inlets were punched using a catheter hole punch (SYNEO, CR0350255N20R4) mounted onto a drill press (Technical Innovations). The flow layer was generated by spin-casting PDMS (1:20 ratio of cross-linker to polymer) onto the molds at 500 rpm with an acceleration of 133 rpm / second for 10 seconds, followed by 1750-1850 rpm with an acceleration of 266 rpm / second for 75 seconds. Layers were relaxed on a flat surface for 10 minutes at room temperature prior to baking at 80°C for 40 minutes in an oven. Cut and punched device control layers were then aligned to flow layers remaining on master molds manually using a stereoscope. Aligned devices were then baked for 50 minutes at 80°C in an oven, excised from the molds using a scalpel, and the remaining flow- layer fluidics inlets made using the same catheter punch as above. TF mutant library preparation
[0078] TF variant plasmids compatible with IVTT reactions were cloned and validated as described previously (Aditham et al 2020).19 30840863V.1Preparation of barcoded oligonucleotide pools
[0079] Degenerate nucleotide libraries for TF binding assays were ordered as ssDNA from IDT (Integrated DNA Technologies) at the 100 nmole synthesis scale with standard desalting purification, normalized to 100 M in IDTE pH 8.0. Sequences were as follows:Generation of barcoding primer plates
[0080] Barcoding primer sets for STAMMP-Seq library generation were designed using NEXTflex-HT™ Barcode Indices (Bioo Scientific) barcodes as a basis set, excluding barcodes with a Hamming distance of 2 or less to the Max and Pho4 consensus motif CACGTG. Primers composed of unique barcodes + sequencing adaptor handles (Figure 2) from IDT (Integrated DNA Technologies) at the 0.5 nmol synthesis scale in a 384-well plate format with standard desalting purification. These primers were normalized at 100 M per well. Generation of barcoded DNA libraries
[0081] Using a 96-channel manual pipette (Liquidator, Rainin), we added 5 L of solubilized primers to 45 L of Milli-Q water to dilute to a working concentration of 10 M and mixed well by pipetting up and down. Finally, we transferred 2.5 L of these diluted primers to a new 96- well plate before adding 47.5 L PCR reaction Master Mix to all wells, composed of the following components (per 50 L reaction): 25 L of NEBNext Ultra II Q52x Master Mix (M0544X) 5L of 50 ng / L degenerate nucleotide library template (Figure S2)0.25 L of 10 M reverse primer (5’-GTGACTGGAGTTCAGACGTGTGCTCT-3’) 17.25 L of DNase- / RNase-free water The solution was then incubated at the following thermocycler settings:20 30840863V.11. Initial Denaturation: 98ºC for 30 sec 2. 2. 12 cycles: (a) Cycle melt: 98ºC for 5 sec (b) Cycle anneal and extend: 72ºC for 30 sec 3. Final extension: 72ºC for 1 min 4. Hold at 4ºC Quantification of barcoded DNA libraries:
[0082] An aliquot (1.5 L) of the resulting PCR reactions of select barcoded DNA libraries were spot-checked through gel electrophoresis on a 2.5% agarose gel stained with 1x Gel Green (Biotium, #41005) and a 50 bp ladder (New England Biolabs, N3236S) as reference to confirm uniform amplification. The amplified products were then purified as follows.
[0083] Two microliters Speed Beads (Cytiva, 65152105050250) and 78 l 20% PEG- 8000 / 2.5 M NaCl (12% PEG-8000 final concentration) were added to each 50 l PCR reaction and mixed by inversion. This mixture was incubated at room temperature for 10 minutes, and the beads were collected on a 96-well magnet plate for 3-5 min. The supernatant was removed with a multichannel pipette, then 180 μl of 70% ethanol was added to each reaction. Each reaction was washed by moving tubes to opposite sides of the magnet bars, moving the beads back-and-forth 20 times, and inverting tubes briefly, then the washed reactions were briefly spun down for subsequent collection on the magnet. This procedure was repeated once more for a total of 2 ethanol washes. After washing, the supernatant was removed and the beads were allowed to dry on the magnet (~10-20 minutes, or until cracks in the beads appeared) before elution into a buffer containing 10 mM Tris pH 8.0 and 0.05% Tween-20. The resulting yields of each barcoding PCR reaction were subsequently quantified by measuring A260 with a DeNovix instrument. DNA Array Printing Plate preparation Prior to printing, we transferred mini-prepped plasmid into 384-well plates. To standardize volumes of plasmids, the wells were evaporated until completely dry. We resuspended each plasmid with 20 uL of print buffer at a final concentration of plasmid of 100 ng / μL. Plasmid print buffer was formulated as described below and sterile filtered prior to use:21 30840863V.11% (10mg / mL) Bovine Serum Albumin (Sigma Life Science, B4287-25G) 200 mM (11.65 mg / mL) NaCl (Sigma Life Science, 71376-1KG) 12 mg / mL trehalose dihydrate (Sigma Life Science, T9531-25G)
[0084] Prior to printing, we transferred cleaned, barcoded oligonucleotide pools into the same 384-well plates. To standardize concentrations of oligonucleotide pools, the wells were evaporated until completely dry using a plate drier. We resuspended each oligonucleotide pool in variable volumes of print buffer to achieve a final concentration of 4 μM. Volumes of print buffer to be used for each well were calculated using the prior Denovix measurements of yield. Oligonucleotide print buffer was formulated as described below: 1% (10mg / mL) Bovine Serum Albumin (Sigma Life Science, B4287-25G) 1X SSC buffer (15mM sodium citrate, 150mM sodium chloride - Invitrogen, 15557044) 12 mg / mL trehalose dihydrate (Sigma Life Science, T9531-25G)
[0085] When not in use, plasmid plates were sealed with foil covers and stored them at -20ºC. Prior to printing, plates were defrosted overnight at 4ºC and centrifuged at 2000 RPM for 2 minutes. Printing & device alignment:
[0086] Plasmids were printed using a SciFlex S3 Arrayer (SCIENION AG) using the PDC70 nozzle (Type 2 coating). A “field file” was generated to map each well on a 384-well plate to positions within the printed plasmid array using custom Python scripts. To prevent cross- contamination between plasmids, the glass nozzle was washed stringently with room temperature Milli-Q water in between spotting different plasmid samples. We printed plasmid arrays on 25 mm x 75 mm epoxysilane-coated glass slides (ArrayIt SME2, SuperChip C50-5588-M20, or self-coated as previously described (Volpetti et al, Plos ONE, 2015 doi.org / 10.1371 / journal.pone.0117744). After drying arrays overnight at room temperature, we aligned microfluidic devices to “program” each chamber with its own printed spot. These devices were then baked for 4 hours at 95ºC on a hotplate. Microscopy and instrumentation:22 30840863V.1
[0087] Devices were imaged as previously described (Aditham et al, et al. Cell Systems, 2020 doi.org / 10.1016 / j.cels.2020.11.012); Markin et al. Science, 2021, doi.org / 10.1126 / science.abf8761) using a Nikon Ti-S microscope. Devices were controlled using a pneumatics manifold (Brower et al. HardwareX, 2018 doi.org / 10.1016 / j.ohx.2017.10.001). Custom scripting and automation enabled integrated control of both the microscope and the pneumatics manifold (github.com / FordyceLab / RunPack). Microfluidic device operation: On-Chip Surface Patterning
[0088] Microfluidic devices aligned to printed libraries consisting of plasmids and barcoded DNA were subsequently surface patterned for protein immobilization. First, to prevent premature solubilization of the DNA spots from osmotic transfer of water from the control layer, the control lines of the device (neck1, neck2, butR, butL, sandR, and sandL) were filled with a 0.55 M NaCl solution. Next, device priming and surface functionalization was carried out largely as described previously (Markin et al., 2021, supra). All reagents were introduced to the device as previously described (Aditham et al, 2020, supra). On-chip expression & purification of TF mutants
[0089] PURExpress (NEB E6800L) was employed as previously described (Aditham et al., 2020, Markin et al., 2021, both supra) with some modifications to express all TF variants simultaneously. Briefly, Parts A and B of PURExpress were equilibrated on ice until defrosted. For one device using 25 μL total of PURExpress, we first incubated 10 μL of Part A with 7.5 μL of Part B on ice for 45 minutes. Then, we added 1.5 μL of recombinant RNAsin (Promega N2515) and 6μL of nuclease-free water (Promega P1193) and mixed by pipette until no phase separation was visible.
[0090] PURExpress was introduced onto the device by flowing PURExpress through flow channels for 10 minutes with all “neck” and “button” valves closed and the “sandwich” valves open to completely fill all channels. After this period, we closed the outlet valve was closed and opened the “button” valves for 30 seconds. Next, we opened the “neck2” valve gating access to TF expression plasmids and monitored flow into the imprint chambers by eye. When imprint chambers became approximately 80% full, we closed the “sandwich” valves and paused for 3023 30840863V.1seconds before additionally closing the “button” valves to mix PURExpress with solubilized plasmid. Devices were then placed on a pre-heated hotplate at 37ºC for 45 minutes to express all proteins. We then placed devices on the scope and allowed the GFP to mature to a fluorescent state over the course of 45-60 minutes with the button valves on the device closed. After this was completed, we opened the button valves and recruited GFP-tagged protein to the antibodies for 20-30 minutes, imaging to monitor progression. We then closed the buttons to shield trapped TFs while we washed the device with PBS (ThermoFisher 12604-013) to remove nonspecifically bound TFs from the device walls. Finally, the device was imaged in the GFP channel (using an exposure time of 500 ms) when complete to quantify TF expression levels in each chamber and brightfield channel to monitor solubilization of oligonucleotide pools. Solubilization and binding of barcoded oligonucleotide libraries
[0091] In order to solubilize DNA libraries, we prepared and flowed Binding Buffer (225 mM NaCl, 20 mM Tris pH 7.5, 2 mM DTT) through the device for 10 minutes with all “neck” and “button” valves closed to completely fill all channels. We then closed the “out” valve and opened the “button” valves for 30 seconds to allow Binding Buffer to fill the entire volume of each reaction chamber.
[0092] We then opened the “neck1” valve and monitored flow into the imprint chambers by eye. When imprint chambers became approximately 80% full, we closed the “sandwich” valves and paused for 2 minutes. Then, we toggled the “button” valves ~5 times to mix the Binding Buffer with DNA libraries. With “button” and “sandwich” valves closed and “neck1” valves open, we allowed DNA libraries to solubilize and diffuse into chambers for 20 minutes. To initiate binding of on-chip expressed TF variants and barcoded DNA pools, we opened the “button” valves and allowed the binding reaction to proceed at room temperature for 90 minutes. At this stage, subsequent fraction collection steps were conducted as follows depending on whether binding or kinetic data was being acquired. Fraction collection for binding experiments
[0093] After binding, we closed the “button” valves to trap the equilibrium-bound DNA. We eluted the unbound DNA fraction by flowing Elution Buffer (20 mM Tris HCl pH 8.0, 0.2 mM EDTA, 0.05% Tween-20, 0.25% BSA) through the device with buttons closed for 5 minutes to24 30840863V.1collect pooled DNA in a 100 μl gel loading pipette tip (USA Scientific 10220810) plugged into the outlet hole. This gel loading pipette tip with the eluted fraction was removed and the liquid fraction was then subsequently transferred to a DNA LoBind Eppendorf tube (Eppendorf022431021) and stored at -20 C. Any remaining unbound DNA in the microfluidic device wasthen removed by washing with Elution Buffer for 10 minutes at 3.5 psi through tubing attached to the outlet hole.
[0094] Thus, in some embodiments, the button valve is closed and the reaction chamber is washed to remove uncaptured protein. In some embodiments, this further includes actuating the sandwich valved, and specifically opening the sandwich valves to allow wash fluid to pass through the reaction chamber via the through channel.
[0095] The bound DNA fraction was eluted by flowing Proteinase K mixture (1 mg / ml Proteinase K, 1x NEB Buffer 2, 0.05% Tween-20) through the device with buttons closed for 10 minutes at 3 psi. After 10 minutes, flow through the device was stopped by closing the outlet control valve and the “button” valves were opened, exposing immobilized TFs with equilibrium - bound DNA to the Proteinase K mixture. This mixture was incubated at RT on the device for 20 minutes to ensure full digestion and elution of bound DNA. Finally, bound DNA fractions were collected by flowing Elution Buffer through the device with “button” valves open for 5 minutes to collect pooled DNA in a 100 μl gel loading pipette tip plugged into the outlet hole. This gel loading pipette tip with the eluted fraction was removed and the liquid fraction was thensubsequently transferred to a DNA LoBind Eppendorf tube and stored at -20 C.Fraction collection for kinetic experiments
[0096] After binding, unbound DNA fractions were first eluted and collected from the device as described previously. To elute and collect kinetic fractions – oligonucleotide pools that spontaneously dissociate from immobilized TFs during a short time interval (on the order of seconds) – Kinetic Buffer (115 mM NaCl, 10 mM Tris pH 7.5, 1 mM DTT, 50 μg / mL BSA, 1 uM “dark competitor” DNA) was flowed through the device with buttons closed for 5 minutes. The inclusion of competitor DNA at high concentrations containing high affinity motifs (but not compatible with downstream library preparation steps and therefore invisible during sequencing readouts, hence the term “dark”) competitor during dissociation is critical to prevent rebinding of barcoded DNA pools, which would lead to systematic underestimation of dissociation rates.25 30840863V.1“Dark competitor” DNA (5’-GTCAATATTTCCTCCCACGTGACTGTGATTCCTTCC-3’) for kinetic assays was ordered as duplexed dsDNA from IDT (Integrated DNA Technologies) at the 100 nmole synthesis scale with standard desalting purification, normalized to 100 M in IDTE pH 8.0.
[0097] After flowing Kinetic Buffer through the device for 5 minutes, flow through the device was stopped by pressurizing the “sandwich” valves and pausing briefly. The “button” valves were then opened for 2 seconds, briefly exposing the immobilized TFs with equilibrium-bound DNA to the buffer-exchanged solution before closing the “button” valves again, pausing briefly before resuming flow to ensure “buttons” were fully closed. Spontaneously dissociated DNA species were eluted from the device by flowing Kinetic Buffer with “buttons” closed for 5 minutes and pooled DNA was collected in a 100 μl gel loading pipette tip plugged into the outlet hole. This gel loading pipette tip with the eluted fraction was removed and the liquid fraction was then subsequently transferred to a DNA LoBind Eppendorf tube for long-term storage.
[0098] This process was then iteratively repeated to collect a total of 12 kinetic fractions, sampling dissociation over 24 seconds in 2 second intervals. Bound DNA remaining on the device after this time period was then collected as described previously. Sequencing library preparation: Addition of “spike-in” oligonucleotide to kinetic fractions:
[0099] To normalize between different read counts of separately eluted kinetic fractions and convert counts to concentration, we added a “spike-in” oligonucleotide to collected kinetic fractions prior to subsequent library preparation. The “spike-in” oligonucleotide (5’-GTC ATA CCG CCG GAC AAG ACT TCG ATA CGT GCG CTC GAC TTG GGA CTG GCT TTC TAC AGC CTA TTC CTG GAG GAT AGG ATA CAC ATA CTC CGA GAT CGG AAG AGC ACA CGT CTG AAC TCC AGT CAC-3’) for kinetic assays was ordered as duplexed dsDNA from IDT (Integrated DNA Technologies) at the 100 nmole synthesis scale with standard desalting purification, normalized to 100 M in IDTE pH 8.0.
[0100] Concentrations of DNA present in a subset of kinetic fractions was quantified via qPCR. For the 1st, 6th, and 12th kinetic timepoint fractions, 1 μl of eluted fraction was added to 5 μl of 2x iTaq Universal SYBR Green Supermix (Bio-Rad 1725120) and 0.5 μM each of26 30840863V.1common forward and reverse qPCR primers, and reaction volumes were brought to 10 μl total with ultrapure DNase-free water. A 10-fold dilution series of input library was prepared by measuring library concentration with the Qubit Flex fluorometer (Invitrogen Q33327) and making 10-fold dilutions from 1 nM to 0.1 fM in buffer containing 10 mM Tris pH 8.0, 0.05% Tween-20.10 μl qPCR reactions were prepared with 1 μl of each of these input library dilutions as described above. All qPCR measurements were performed in triplicate along with no sample controls to determine the noise floor of the assay. qPCR reactions were carried out with a Bio- Rad qPCR machine using the following cycling parameters: 1. Initial Denaturation: 95 C for 30 sec2.39 cycles (18 cycles for kinetic fractions): (a) Cycle melt: 95 C for 5 sec(b) Cycle anneal / extend / read: 65.5 for 25 sec3. Melting Curve: 65 - 95 C, 5 seconds / step with 0.5 C steps.
[0101] Raw qPCR intensity data was analyzed by fitting a sigmoid to each SYBR fluorescence curve using the fit inflection points as measures of concentration or cycle thresholds. A standard curve was built using the inflection points of the input library samples of known concentration, and the original concentrations of kinetic fraction elutions were determined by comparing fit inflection points to this standard curve. Finally, a bulk dissociation curve was fit to the calculated kinetic fraction concentrations to determine a desired concentration of the “spike-in” oligonucleotide to normalize concentrations between kinetic fractions estimated to result in 15% of the total reads. A working concentration of 500 fM “spike-in” oligonucleotide was diluted into each kinetic fraction, normalized to the total volume of the eluted fraction, to a final concentration of 27 fM.Addition of unique molecular identifiers to bound and unbound DNA molecules:27 30840863V.1
[0102] In order to add unique molecular identifiers (UMI) and Illumina “Read1” primer landing sites to the eluted bound and unbound fractions, primer extension reactions were performed as follows using an oligonucleotide containing degenerate nucleotides (hereafter UMI-oligo) (5’- ACACTCTTTCCCTACACGACGCTCTTCCGATCNNNNNNNNNNNNNNNNNNNNNNNN NNGTCATACCGCCGGACAAGAC-3’). Unbound fractions were diluted at a 1:100 ratio in TT buffer (10 mM Tris pH 8.0, 0.05% Tween-20), and 1 μl of the dilution was used for the library preparation. Entire bound fractions were used for library preparation. First, elutions were annealed to the UMI-oligo in a 12 μl reaction consisting of 167 nM UMI-oligo and 50 mM NaClbrought up to 12 μl with TT and placed in a thermocycler for a 10 minute incubation at 95 C(both to denature DNA and deactivate residual proteinase K) and 57 cycles of 30 secondincubations starting at 94 C and decreasing by 1 C each cycle until 37 C is reached. Then, foreach sample, 8 μl of 2x NEBNext Ultra II Q5 PCR master mix (NEB M0544X) was heated to98 C for 50 seconds to activate the polymerase, brought back to room temperature, and added tothe annealed DNA for a total volume of 20 μl. The extension reaction was then carried out byincubating at 55 C for 1 hour and bringing samples back to 4 C in a thermocycler. To removeunextended single-stranded UMI-oligo, 1 μl (or 20 units) of Exo1 (NEB M0293L) was added toeach extension reaction and incubated at 37 C for 60 minutes, 80 C for 20 minutes to heatinactivate the ExoI, and then brought back to 4 C.Amplification and addition of sequencing indexes:
[0103] To add Illumina sequencing adapters and indexes for demultiplexing, we prepared 50 μl PCR reactions by adding 17 μl of 2x NEBNext Ultra II Q5 PCR master mix, 0.5 μM final of P5-adapter-fwd oligo, and 0.5 μM final of P7-adapter-rev oligo to the 21 μl of each sample, bringing the volume up to 50 μl with Ultrapure Water (Invitrogen 10977015), and amplifying with the following thermocycling program: 1. Initial Denaturation: 98 C for 30 sec2.12 cycles (18 cycles for kinetic fractions): (a) Cycle melt: 98 C for 5 sec(b) Cycle anneal: 63 for 10 sec(c) Cycle extend: 72 C for 20 sec28 30840863V.13. Final extension: 72 C for 1 min4. Hold at 4 C.Quantification of STAMMP-seq libraries:
[0104] Each PCR reaction was cleaned up using PEG-8000 / NaCl / speedbeads as above with the following alterations.2 μl of speedbeads and 42.5 μl of 20% PEG-8000 / 2.5 M NaCl were added to each 50 μl PCR reaction for a final concentration of 8.5% PEG-8000 to avoid precipitating adapter-dimers. Libraries were eluted in 20 μl of TT. Library concentrations were quantified using the Qubit-Flex fluorometer (Invitrogen Q33327) and the high-sensitivity dsDNA detection kit (Invitrogen Q32854).
[0105] To assess clean-up efficiency, 2 L of each library were analyzed by gel electrophoresis on a 2.5% agarose gel stained with 1x Gel Green (Biotium, #41005) and a 50 bp ladder (New England Biolabs, N3236S) as reference. Sequencing
[0106] STAMMP-seq libraries were sequenced to an average depth of 20 million reads on an Illumina NextSeq 500 or MiSeq. Data analysis: Illumina sequencing data processing
[0107] UMIs were extracted from 5’ ends of reads using UMI-tools (Smith et al, Genome Res., 2017, doi.org / 10.1101 / gr.209601.1162017) using the “extract” function with the first 2629 30840863V.1basepairs of the read specified as UMIs . Constant sequences on both 5’ and 3’ ends of reads were then trimmed from reads using the Cutadapt tool (Martin, EMBnet, 2011, doi.org / 10.14806 / ej.17.1.200). Trimmed reads were aligned using bowtie2 (Langmead and Salzberg, Nature Methods, 2012, doi.org / 10.1038 / nmeth.1923) to bowtie indexes built from a fasta file containing all enumerated library sequences (as described above) using default settings. Resulting SAM files were converted to BAM files, sorted, and indexed using SAMtools (Danecek et al, GigaScience, 2021, doi.org / 10.1093 / gigascience / giab008). Reads from each sample were then deduplicated using the “dedup” function of UMI-tools with the method option set to “unique” and using the sorted and indexed BAM files as inputs. Finally, deduplicated samples were converted to SAM files for further processing using custom Python scripts. Calculation of relative free energies of binding
[0108] Using custom Python scripts and deduplicated SAM files as input, counts of every well-barcode:library sequence combination were tabulated for bound and unbound fractions, and combinations that were unobserved in both bound and unbound fractions were removed from further analysis. A pseudocount of 2*1,000,000 / total reads for each sample was added to the count of each well-barcode:library sequence combination to prevent log(0) errors in later steps, and probabilities of observing each combination calculated using these resultant counts. Probabilities were then used to calculate relative free-energies of binding using the following equation (Le et al., PNAS, 2018 doi.org / 10.1073 / pnas.1715888115) where R = 1.987094E-3 kcal / (K*mol) and T = 298 K:Calculation of dissociation rate constants
[0109] Using custom Python scripts and raw counts of every well-barcode:library sequence combination in the bound fraction, in the unbound fraction, and in each kinetic fraction as input, dissociation rates for each protein-DNA interaction were calculated as follows. First, counts of each well-barcode:library sequence were converted into concentration by normalizing to raw30 30840863V.1counts of the “spike-in” oligonucleotide with a known concentration in each kinetic fraction. Next, the concentration of each well-barcode:library sequence combination in each kinetic fraction (a collection of all dissociated sequences in a 2 second interval) were fit to the derivative of a single exponential function for dissociation to determine koff. LSR variant analysis
[0110] Analysis of LSR variant recombination activity using STAMMP seq was performed as described above for the analysis of TF binding activity with the following adaptations to evaluate enzymatic activity rather than binding activity.
[0111] In brief, a different buffer containing 300 mM NaCl 40 mM Tris pH 8.0, 2mM DTT, 2 mM EDTA, and 2 mM spermidine was employed for the step of solubilization and binding of barcoded oligonucleotide libraries for the enzymatic reaction. To initiate mixing of the on-chip expressed LSR variants and barcoded DNA pools, the “button” valves were opened as described above and the reaction was allowed to proceed for 16 hours at room temperature. Captured agents were not employed for this analysis. Thus, after mixing, DNA fractions were eluted and collected from the device and stored as described above for the “unbound fraction”.
[0112] Illumina sequencing data processing for enzymatic experiments was performed as follows. UMIs (first 26 basepairs of read1) were extracted from 5’ ends of read1 and constant sequences on both 5’ and 3’ ends of reads were trimmed from paired-end reads using the Cutadapt tool (Martin, EMBnet, 2011, doi.org / 10.14806 / ej.17.1.200). Both trimmed reads were aligned using bowtie2 (Langmead and Salzberg, Nature Methods, 2012, doi.org / 10.1038 / nmeth.1923) to bowtie indexes built from a fasta file containing all enumerated library sequences (as described above) using default settings. Resulting SAM files were converted to BAM files, sorted, and indexed using SAMtools (Danecek et al, GigaScience, 2021, doi.org / 10.1093 / gigascience / giab008). Reads from each sample were then deduplicated using the “dedup” function of UMI-tools with the method option set to “unique” and using the sorted and indexed BAM files as inputs. Finally, deduplicated samples were converted to SAM files for further processing using custom Python scripts.
[0113] Calculation of endpoint turnover was performed as follows. Using custom Python scripts and deduplicated SAM files as input, counts of every well-barcode:library sequence31 30840863V.1combination were tabulated and only exact matches to enumerated library sequences were considered for further analysis. Fraction recombined was then calculated for each unique device, barcode, and library combination as the ratio of observed attL sequences over the sum of observed attL and attP sequences.
[0114] It is understood that examples and embodiments described herein are for illustrative purposes only and that various modifications or changes in light thereof will be suggested to persons skilled in the art and are to be included within the spirit and purview of this application and scope of the appended claims.
[0115] All publications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication or patent application was specifically and individually indicated to be incorporated by reference for the material for which it is cited.32 30840863V.1
Claims
1. A method of evaluating interactions of protein variants and members of a population of barcoded molecules, the method comprising: (a) providing a device comprising: a solid surface with an imprinted array comprising (i) a library of protein variants, wherein variants are imprinted at discrete known location on the solid surface to provide a localized site for each variant protein, wherein the position of the localized site identifies the protein variant; and (ii) a multiplicity of barcoded individual pools of a population of molecules comprising different molecules to be evaluated for interactions with variant proteins, wherein each individual pool comprises the population of molecules barcoded with a barcode specific for the pool; which barcode differs from the barcodes of other individual pools; wherein each individual pool is imprinted at a discrete known location on the solid surface; and a valved microfluidic device comprising a first set of imprint chambers, a set of reaction chambers wherein each reaction chamber has a valve controlling flow into the chamber and a second set of imprint chambers, wherein the set of reaction chambers is fluidly couplable to both the first and second set of imprint chambers, the first set of imprint chambers is physically aligned with the localized sites for each protein variant to provide a protein variant chamber for each variant protein; and the second set of imprint chambers is physically aligned with each individual pool of barcoded molecules to provide a barcoded pool chamber for each individual pool; (b) introducing the protein from each protein variant chamber into its connected reaction chamber; (c) introducing solubilized barcoded molecules into the reaction chamber from the connected barcoded pool chamber to allow formation of complexes between barcoded molecules and variant proteins; (d) incubating complexes formed in the reaction chambers with a capture agent that that binds to the protein of interest and variants thereof; and (e) determining the sequence of barcodes of captured barcoded molecules, thereby identifying molecules that bind to the protein of interest or variants thereof.33 30840863V.
12. The method of claim 1, wherein each of the first set of imprint chambers comprises an expression construct and step (b) further comprises introducing a cell-free expression mixture into each of the variant protein chambers.
3. The method of claim 1, wherein each of the first set of imprint chambers comprises a lyophilized protein variant and step (b) comprises solubilizing protein in each of the variant protein chambers.
4. The method of claim 1, 2, or 3, further comprising detecting the amount of barcoded molecules in individual pools that interact with proteins produced by the library of variant proteins.
5. The method of any one of claims 1-4, wherein a capture agent is immobilized to a surface in the reaction chamber.
6. The method of claim 5, further comprising washing the reaction chambers to remove unbound barcoded molecules.
7. The method of claim 5 or 6, further comprising collecting barcoded molecules present in complexes bound by the capture agent and determining the sequence of the barcodes.
8. The method of any one of claims 1-7, further comprising collecting unbound barcoded molecules and determining the sequence of the barcodes of the barcoded molecules in the unbound fraction.
9. The method of claim 7, wherein collecting barcoded molecules comprises separately collecting fractions of unbound barcoded molecules that have dissociated over time from the immobilized protein at desired time points; determining the sequence of the barcodes in the unbound fraction at each time point; and quantifying the amount of barcoded molecule from each pool present in the unbound fraction at each time point.
10. The method of any one of claims claim 1-9, comprising quantifying the amount of protein captured in each reaction chamber.34 30840863V.
111. The method of any one of claims claim 1-10, wherein the expressed proteins comprise a tag region that binds to the capture reagent and allows for quantification of the amount of bound protein in each reaction chamber.
12. The method of claim 11, wherein the tag region comprises a fluorescent or luminescent protein and the capture reagent is an antibody that binds to the fluorescent or luminescent protein.
13. The method of any one of claims 1-12, wherein the variant proteins comprise nucleic acid binding proteins.
14. The method of claim 13, wherein the population of barcoded molecules comprises DNA oligonucleotides.
15. The method of claim 13, wherein the population of barcoded molecules comprises RNA oligonucleotides.
16. The method of claim 14 or 15, wherein the population of barcoded molecules comprises single-stranded oligonucleotides.
17. The method of claim 14 or 15, wherein the population of barcoded molecules comprises double-stranded oligonucleotides.
18. The method of claim 13, wherein the protein is a transcription factor.
19. The method of any one of claims 1-12, wherein the population of barcoded molecules comprises candidate protein binding partners of a protein of interest.
20. The method of claim 19, wherein the population of barcoded molecules comprises a display library that expressed candidate binding partners.
21. The method of claim 19 or 20, wherein the protein of interest comprises an binding domain of a receptor and the barcoded molecules comprise variants of a ligand that binds to the binding domain of the receptor.35 30840863V.
122. The method of any one of claims 1-4, wherein the variant proteins comprise nucleic acid modification proteins.
23. The method of any one of claims 1-22, wherein the set of reaction chambers is aligned in parallel with the first set of imprint chambers and is aligned in parallel with the second set of imprint chambers, thereby providing an array of chambers in which each protein variant chamber is connected by a channel to a reaction chamber, which is connected by a second channel to a barcoded pool chamber.
24. The method of any one of claims 1-23, wherein the reaction chambers are fluidly connected via a through channel, each fluid channel divided by an actuable through- channel valve, wherein each reaction chamber is isolatable from the other reaction chambers via actuation of through-channel valves in the through channel connecting each reaction chamber to the other reaction chambers, wherein introducing the protein from each protein variant chamber into its connected reaction chamber comprises actuating the through-channel valves to isolate the connected reaction chamber from the other reaction chambers.
25. The method of claim 24, wherein the method comprises a capture agent immobilized to a surface in the reaction chamber and further comprises a step of washing each reaction chamber to remove uncaptured protein, wherein washing comprises actuating the through-channel valves to allow wash fluid to pass through the reaction chamber via the through channel. .36 30840863V.1
Citation Information
Patent Citations
Manipulation of microparticles in microfluidic systems
US20160040226A1
Barcoded Protein Array for Multiplex Single-Molecule Interaction Profiling
US20230295704A1
Sample tracking using molecular barcodes
WO2003052101A1