Methods for displaying biomolecules
Patent Information
- Application Number
- JP2024515532
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-09-10
- Filing Date
- 2022-09-06
- Publication Date
- 2025-09-30
AI Technical Summary
Existing methods for displaying RNA and polypeptides on flow cells face challenges such as instability due to non-covalent binding, degradation over time, and limitations in withstanding denaturing conditions, which affect high-throughput analysis and drug discovery processes.
A method involving the immobilization of a first nucleic acid on a substrate, followed by polymerization with a nucleic acid polymerase to create a complementary non-DNA nucleic acid, such as RNA or XNA, using a DNA primer to form a bridge, and subsequent cleavage and linearization to enhance stability and enable high-throughput sequencing and screening.
Enables the generation of high-resolution sequence-function pair information for biomolecule libraries, facilitating rapid screening and drug discovery by allowing up to 10^8 sequencing and screening of variants in 2-3 days, with improved stability and efficiency under denaturing conditions.
Smart Images

Figure 00000072_0000 
Figure 00000073_0000 
Figure 00000073_0001
Abstract
Description
[Technical field]
[0001] FIELD OF THEINVENTION The present invention relates to a method for presenting biomolecules on a substrate, for example on a flow cell surface. The present invention relates to a method for the upstream, downstream or direct presentation of xenonucleic acid (XNA) molecules, RNA molecules and / or polypeptides on a substrate. The present invention further relates to a substrate presenting biomolecules obtained or obtainable by the method of the present invention.
[0002] BACKGROUND OF THE PRESENTINVENTION Platforms that allow high-throughput analysis of biological molecules, such as analysis of binding affinities or other properties, are important for drug discovery. Flow cells are commonly used as substrates for displaying DNA molecules that can then be interrogated to obtain both positional and sequence information.
[0003] Attempts have been made to utilize flow cells to present RNA molecules or polypeptides. For example, there are several methods that use DNA clusters immobilized on a flow cell to generate RNA that is non-covalently bound to the DNA clusters via stalled RNA polymerase. The RNA can then be translated. Such methods include those disclosed in WO2014 / 189768, Layton et al. (Layton et al., 2019, Molecular Cell 73, 1075 1082), and US2019 / 0112730. However, these approaches have known drawbacks. For example, these complexes are not covalently bound to the flow cell, can degrade over time, and assays require the use of signal loss normalization techniques. Furthermore, analyses that require conditions that can denature these complexes cannot be performed. These complexes dissociate, for example, at high temperatures, chemical denaturants, and low or high concentrations of magnesium. Low concentrations of magnesium dissociate the ribosome from the complex. High concentrations of magnesium can dissociate the RNA polymerase from the complex. Thus, these presentation techniques have limitations.
[0004] Svensen et al. report a method in which flow-cell-bound clusters of identical DNA strands, generated by Illumina DNA sequencing technology, are converted into clusters of complementary RNA and then into peptide clusters (Chembiochem. 2016 September 02;17(17): 1628-1635. doi:10.1002 / cbic.201600298). This method requires that flow-cell-bound primers be modified with ribonucleotides so that they can be used by poliovirus 3Dpol polymerase. The yield of RNA produced is suboptimal, so the yield of the resulting polypeptide could be increased.
[0005] Thus, there is a need in the art for a method for displaying biomolecules that overcomes the aforementioned problems.
[0006] Moriizumi et al. have provided insights into in vitro translation (Moriizumi, Yoshiki et al., "Osmolyte-enhanced protein synthesis activity of a reconstituted translation system," ACS synthetic biology 8.3(2019): 557-567). Summary of the Invention [Problem to be solved by the invention]
[0007] (Summary of the invention) We provide herein methods that allow the creation of high throughput drug discovery platforms. We generate substrate-bound libraries of biological molecules, including XNA, RNA, and polypeptides, and show that these libraries can be interrogated, for example, by measuring binding affinity or enzymatic activity.
[0008] Specifically, the present inventors have 8A library of variants can be sequenced and screened in 2-3 days, including replication measurements. Sequence-feature pair information can be generated for all library variants.
[0009] In this manner, the platform disclosed herein generates an unprecedented amount of high-resolution data, which can be used in combination with machine learning, for example, in engineering therapeutic agents.
[0010] In an embodiment of the invention, there is provided a method for displaying non-DNA nucleic acid molecules on a substrate, comprising: i) providing a first nucleic acid immobilized on a substrate, wherein the first nucleic acid is oriented such that its 5' end is proximal and its 3' end is distal to its point of immobilization; ii) generating a second nucleic acid complementary to the first nucleic acid, the generating step of the second nucleic acid comprising: a) contacting the first nucleic acid with a nucleic acid polymerase under conditions suitable for polymerization, wherein the primer for polymerization is a DNA primer immobilized on the substrate such that a bridge is formed during polymerization; the product of the polymerization is a non-DNA nucleotide strand immobilized on the substrate via the primer, and The nucleic acid polymerase is a polymerase capable of acting on a DNA primer to synthesize a non-DNA nucleic acid molecule complementary to a single-stranded nucleic acid template; and and iii) removing said first nucleic acid resulting in the presentation of said second nucleic acid on said substrate.
[0011] The second nucleic acid may be an RNA molecule. The nucleic acid polymerase may comprise an amino acid sequence having at least 36% identity to the amino acid sequence of SEQ ID NO:1, and the amino acid sequence may comprise a Y409 and an E664 mutation with respect to the amino acid sequence of SEQ ID NO:1. The Y409 mutation may be Y409N or Y409G, and the E664 mutation may be E664K or E664Q. In one particular embodiment, the Y409 mutation is Y409G, and the E664 mutation is E664K. The amino acid sequence of the nucleic acid polymerase may comprise SEQ ID NO:3.
[0012] The second nucleic acid may be an XNA molecule. For example, the XNA molecule may include arabino nucleotides, arabino nucleic acid (ANA) nucleotides, 2'-fluoroarabino nucleic acid (FANA) nucleotides, 2'-O-methyl ribonucleic acid (2'OMe) nucleotides, 2'-O-methoxyethyl (MOE) nucleic acid nucleotides, phosphorothioate 2'-O-methoxyethyl (PS-MOE) nucleotides, phosphorodiamidate morpholino nucleotides, locked nucleic acid (LNA) nucleotides, P-alkylphosphonate nucleic acid (phNA) nucleotides, threose nucleic acid (TNA) nucleotides, hexitol nucleic acid (HNA) nucleotides, 2'-hydroxyhexitol (AtNA) nucleotides, cyclohexene nucleic acid (CeNA) nucleotides, or 3'-deoxy-DNA (2'-5') nucleotides.
[0013] The nucleic acid polymerase may comprise an amino acid sequence having at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO: 1, and further comprising a mutation that allows polymerization of at least one XNA nucleotide or RNA nucleotide. The amino acid sequence of the nucleic acid polymerase may comprise one or more or all of the following mutations: V93Q, D141A, E143A, and A485L. In particular, the polymerase may be TGK, TGLLK, 2M, Bst, RT521, 6G12, 6G12521, C7, PGLVV, PGLVVWA, D4K, or a variant thereof.
[0014] The method may include, after step ii)a), cleaving the first nucleic acid to linearize the bridge. The method may further include re-contacting the linearized product with a nucleic acid polymerase under conditions suitable for polymerization.
[0015] In another embodiment of the invention, there is provided a method for displaying non-DNA nucleic acid molecules on a substrate, comprising: i) providing a first nucleic acid immobilized on a substrate, wherein the first nucleic acid is oriented such that its 5' end is proximal and its 3' end is distal to its point of immobilization; ii) generating a second nucleic acid complementary to the first nucleic acid, the generating step of the second nucleic acid comprising: a) contacting the first nucleic acid with a nucleic acid polymerase under conditions suitable for polymerization, wherein a primer for the polymerization is immobilized on the substrate such that a bridge is formed during polymerization; and The product of the polymerization is a non-DNA nucleotide strand immobilized on the substrate via the primer; b) cleaving the first nucleic acid to linearize the bridge; and c) contacting the linearized product of step b) with a nucleic acid polymerase under conditions suitable for polymerization; and and iii) removing said first nucleic acid resulting in the presentation of said second nucleic acid on said substrate.
[0016] The second nucleic acid may be an RNA molecule. The bridge may be denatured by temperature.
[0017] The first nucleic acid may be cleaved at the 8-oxoguanine site with formamidopyrimidine DNA glycosylase (Fpg). The third nucleic acid may be annealed to the first nucleic acid at the 8-oxoguanine site prior to cleavage with Fpg.
[0018] Step ii)a) may comprise at least 5, 10, 12, 15, 20 or 25 cycles of bridged polymerisation. The first nucleic acid may be removed in step iii) by contacting the first nucleic acid with a denaturing agent. The denaturing agent may be a buffer comprising 1-500 mM NaOH and 0-20 mM EDTA; or 100 mM NaOH and 5 mM EDTA.
[0019] In one embodiment, the second nucleic acid is an RNA molecule and encodes a polypeptide, and the method further comprises the step of iv) contacting the second nucleic acid with a ribosome under conditions suitable for translation of the encoded polypeptide. The conditions of step iv) may include trimethylamine N-oxide (TMAO). TMAO may be at a concentration of 0.05-1.5 M, 0.05-1.2 M, or 4 M.
[0020] In an embodiment of the invention, there is provided a method for displaying a polypeptide on a substrate, comprising: i) providing a first nucleic acid comprising an antisense sequence encoding a single chain variable fragment (scFv), wherein the first nucleic acid is immobilized on a substrate and oriented such that the 5' end is proximal and the 3' end is distal to the point of immobilization; ii) generating a second nucleic acid complementary to the first nucleic acid, the generating step of the second nucleic acid comprising: contacting the first nucleic acid with a nucleic acid polymerase under conditions suitable for RNA polymerization, wherein a primer for the polymerization is immobilized on the substrate such that a bridge is formed during polymerization; and The product of the polymerization is an RNA nucleotide strand immobilized on the substrate via the primer; iii) removing the first nucleic acid resulting in the presentation of the second nucleic acid on the substrate; and and iv) contacting the second nucleic acid with ribosomes under conditions suitable for translation of the encoded scFv, wherein the conditions in step iv) comprise trimethylamine N-oxide (TMAO).
[0021] The ribosome-polypeptide complex may be stabilized by application of a ribosome presentation buffer, which may contain a magnesium concentration greater than 7 mM MgCl2; or equivalent to 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, or 100 mM MgCl2 or MgAc; or equivalent to 8-100 mM, 10-90 mM, 15-85 mM, 20-80 mM, 25-75 mM, 30-70 mM, 35-65 mM, 40-60 mM, 45-55 mM MgCl2; or equivalent to 8-100 mM, 10-90 mM, 15-85 mM, 20-80 mM, 25-75 mM, 30-70 mM, 35-65 mM, 40-60 mM, or 45-55 mM MgAc.
[0022] In one embodiment, the second nucleic acid is an RNA molecule and a plurality of first nucleic acids encoding a plurality of polypeptides is provided in step i) such that a display library is generated. The encoded polypeptide may be an antibody fragment or an enzyme. The encoded polypeptide may be a single chain variable fragment (scFv), a peptide, a fibronectin type III domain (FN3 domain), a single domain antibody (sdAb, also known as a nanobody), an affibody, a darpin, a fynomer, an OBody, or an avimer.
[0023] The first nucleic acid immobilized on the substrate provided in step i) comprises: 1) providing a template nucleic acid; 2) hybridizing the template nucleic acid to a primer immobilized on a substrate; 3) contacting the hybridized template nucleic acid with a polymerase under conditions suitable for extension of the immobilized primer to synthesize the first nucleic acid, which is a nucleotide strand complementary to the template; 4) performing bridge amplification of the first nucleic acid to generate clusters of the first nucleic acid; and 5) sequencing at least a portion of the first nucleic acid.
[0024] In one embodiment, the bridge amplification comprises: 32-35 amplification cycles, has an extension time of 60-120 seconds per cycle, and includes the use of an amplification buffer with a Mg concentration equivalent to 2-6 mM MgSO4, and / or includes the use of a denaturation buffer comprising 95-99.9% formamide, optionally 1-10 mM NaOH, and optionally 1-5 mM EDTA. In one particular embodiment, the bridge amplification comprises 32 amplification cycles, has an extension time of 60 seconds per cycle, and includes the use of an amplification buffer with a Mg concentration equivalent to 6 mM MgSO4, and / or includes the use of a denaturation buffer comprising 98% formamide, 10 mM NaOH, and 1 mM EDTA.
[0025] In an embodiment of the invention, there is provided a method for preparing clusters of substrate-bound nucleic acids, comprising: 1) providing a template nucleic acid; 2) hybridizing the template nucleic acid to a primer immobilized on a substrate; 3) contacting the hybridized template nucleic acid with a polymerase under conditions suitable for extension of the immobilized primer to synthesize the first nucleic acid, which is a nucleotide strand complementary to the template; and 4) performing bridge amplification of the first nucleic acid to generate clusters of the first nucleic acid, the bridge amplification being performed with 32-35 amplification cycles, having an extension time of 60-120 seconds per cycle, comprising the use of an amplification buffer having a Mg concentration equivalent to 2-6 mM MgSO4, and comprising the use of a denaturing buffer comprising 95-99.9% formamide, optionally 1-10 mM NaOH, and optionally 1-5 mM EDTA.
[0026] In one embodiment, the bridge amplification comprises 32 amplification cycles, has an extension time of 60 seconds per cycle, and comprises the use of an amplification buffer having a Mg concentration equivalent to 6 mM MgSO4, and / or comprises the use of a denaturing buffer comprising 98% formamide, 10 mM NaOH, and 1 mM EDTA.
[0027] In an embodiment of the present invention, a substrate is provided that displays a non-DNA nucleic acid molecule obtained or obtainable by the method disclosed herein. In an embodiment of the present invention, a substrate is provided that displays an RNA molecule obtained or obtainable by the method disclosed herein. In an embodiment of the present invention, a substrate is provided that displays an XNA molecule obtained or obtainable by the method disclosed herein. In an embodiment of the present invention, a substrate is provided that displays a polypeptide molecule obtained or obtainable by the method disclosed herein.
[0028] In an embodiment of the present invention, there is provided a use of a nucleic acid polymerase for extending a DNA primer immobilized on a substrate to synthesize a non-DNA nucleic acid molecule complementary to a single-stranded nucleic acid template. The nucleic acid polymerase may comprise an amino acid sequence having at least 36% similarity or identity to the amino acid sequence of SEQ ID NO:1 and may comprise Y409 and E664 mutations, and may polymerize an RNA molecule complementary to the nucleic acid template. The nucleic acid polymerase may comprise a sequence having at least 80%, 90%, 95%, 99%, or 100% identity to the amino acid sequence of SEQ ID NO:3, and residues 93, 141, 143, 409, 485, and 664 are invariant. The nucleic acid polymerase may comprise an amino acid sequence having at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO: 1, and may further comprise a mutation that allows polymerization of at least one XNA nucleotide or RNA nucleotide. The amino acid sequence of the nucleic acid polymerase may comprise one or more or all of the following mutations: V93Q, D141A, E143A, and A485L. In particular, the polymerase may be TGK, TGLLK, 2M, Bst, RT521, 6G12, 6G12521, C7, PGLVV, PGLVVWA, D4K, or a variant thereof.
[0029] In an aspect of the invention, there is provided a method of screening a substrate displaying a plurality of biomolecules, the substrate being any of those disclosed herein, and the biomolecules forming a library.
[0030] In another aspect of the invention, there is provided a method of displaying a non-DNA nucleic acid molecule or a polypeptide as disclosed herein, said method further comprising the step of screening the displayed non-DNA nucleic acid molecule or polypeptide molecule.
[0031] The screens disclosed herein may involve measuring the affinity of the displayed biomolecules, non-DNA nucleic acid molecules, or polypeptide molecules for a ligand or target molecule, or measuring enzymatic function. [Brief description of the drawings]
[0032] BRIEF DESCRIPTION OF THE DRAWINGS [Figure 1] Figure 1 shows a representative process of displaying a cluster of polypeptides on a substrate according to the present invention. This diagram covers the following processes: A) Generating clusters of DNA molecules and obtaining their positional and sequence information. In one embodiment, the method of the present invention relates to an improved method of generating clusters as disclosed herein, which is particularly suitable for use in the display of long polypeptides of more than 300 amino acids. B) Converting the clusters of DNA into clusters of RNA by engineering polymerases and a linearization step. This linearization releases the torque accumulated during the synthesis of the bridged RNA, allowing full-length synthesis of long RNA molecules (>1.2 kb). C) In vitro translation and ribosome display of polypeptides encoded within the RNA clusters according to the described structure design, and the process of measuring binding affinity to the displayed polypeptides. D) Downstream analysis of sequencing and binding data.
[0033] [Diagram 2] Figure 2 shows an example of the generation of a library of RNA molecules encoding single-stranded variable fragments (scFv) on a flow cell. The top panel, "without torque release," shows 12 cycles of RNA synthesis with TGK polymerase, followed by linearization of the DNA template with Fpg. The bottom panel, "with torque release," shows 12 cycles of RNA synthesis with TGK polymerase, followed by 2 cycles of DNA template linearization with Fpg, followed by further RNA synthesis with TGK polymerase. The images show the levels of binding of fluorescent oligonucleotide probes to the RNA 3' ends. The method for generating the top panel achieved incomplete amounts of RNA synthesis, whereas the method for the bottom panel achieved much higher levels.
[0034] [Diagram 3]Figure 3 shows an example of measuring the effect of magnesium concentration in the ribosome presentation buffer. Some prior art methods are limited to magnesium concentrations equivalent to 7 mM MgCl2 or less because higher concentrations can denature the stalled RNAP:RNA complex. The methods disclosed herein can use much higher concentrations of magnesium. Figure 3 shows the levels of Her2 binding (100 nM Her2-biotin and 100 nM AF532-strepavidin) in low versus high magnesium conditions. High magnesium conditions improved binding and presentation efficiency by 5-fold.
[0035] [Figure 4]Figure 4. Deep screening workflow. 1) Deep screening begins with library preparation, followed by the addition of 5' and 3' untranslated regions (UTRs) to both ends of the protein coding region of the library. The assembled library is then clustered and the N28 UMIs are sequenced on a HiSeq 2500, resulting in the reporting of the UMI sequences and their physical xy coordinates on the flow cell. 2) The sequenced flow cell is then used for deep screening by converting the clusters of DNA to clusters of RNA and removing the DNA template. The RNA clusters are labeled with Atto647N-labeled oligos, then translated into proteins and bound to the RNA via ribosome presentation. Following presentation, an on-chip binding assay is performed with equilibrium binding at high concentrations of biotinylated antigen and AF532-labeled streptavidin, followed by kinetic dissociation. 3) If the binding assay reports a hit in that library, a sequencing run is performed on a new flow cell by sequencing the UMIs and CDRs with internal sequencing primers. The CDRs are then paired with the binding data using the common UMI between the two experiments. 4) The paired CDR:binding data is analyzed for hits and / or machine learning models are trained to predict hits, which can be used to generate either a library for subsequent deep screening rounds, or 5) a shortlist of potential hits for characterization in an appropriate format.
[0036] [Diagram 5]Figure 5. A) Selection workflow. B) Library statistics. C) Deep-screen integrated binding strength versus abundance for R3 MACS and R3 FACS libraries at 300 nM HEL. The library mean intensity is shown as a grey dashed line, and the green solid line indicates a hit threshold of 2x library background. Spearman rank correlation constants of 0.361 and 0.442, respectively, indicate low correlation between abundance and deep-screen binding strength. D) Characterization hit candidates show library construct organization and ID, where M1-23 are from the R3 MACS library, C1-8 were identified by colony picking of the R3 MACS library, and F1-10 are from the R3 FACS library. Abundance, CDR sequences, deep-screen fitted equilibrium binding KD, and Octet fitted kinetic KD are also shown. Sequences shown are SEQ ID NOs: 41-150 in order of appearance from beginning to end. E) Deep screening equilibrium binding and kinetic dissociation curves for clones M5, M6, M14, and M15. F) Octet kinetics for the same four clones at 50 nM on a streptavidin chip loaded with HEL-biotin. G) Octet KD plotted against deep screening binding intensity at 300 nM HEL for all characterized clones showed a Spearman rank correlation constant of -0.697.
[0037] [Figure 6]Figure 6. A) Summary of direct affinity maturation experiment. B) Library statistics for L1L3 library not selected in deep screen. C) Deep screen intensity of huIL-7 at 0, 100 pM and 1 nM sorted by intensity (rank) and background normalized. D) IL70001 and top 19 clones showing VL1 and VL3 sequences, raw deep screen intensity of huIL-7 at 333 pM, Octet fitted KD and IL7R IC50. Sequences shown are SEQ ID NOs: 151-176 in order of appearance from the beginning. E) Octet kinetics of IL70001, IL70100, IL70102 and IL70105 Fabs at 50 nM against a streptavidin chip loaded with huIL-7. F) Deep screen average intensity of top 19 clones of huIL-7 at 333 pM plotted against fitted Octet KD. Error bars are standard error of the mean (SEM). Grey vertical line shows average library intensity of huIL-7 at 333 pM. G) TF-1 STAT5 IL7Rα+γ luciferase inhibition assay with IL70001, IL70100, IL70102, and IL70105 shown as representative ranges for the assay. All inhibition assay curves are shown in FIG. 15. Error bars are standard deviation, n=2. H) Plotting BLI / Octet fitted KD versus IC50 revealed a strong non-linear correlation between affinity and inhibition (ρ=0.956, R2=0.901, fitting log(y)=m×log(x)+c (grey line)).
[0038] [Figure 7]FIG. 7 presents a library of characterized anti-Her2 single chain antibodies (scFv) and shows experiments in which equilibrium binding affinity and kinetic dissociation rates were measured. A) Structure design showing CDR sequences (VH3, VL1, and VL3) and binding affinity of clones G98A, C6.5, ML3-9, H3B1, and B1D2+A1. Sequences shown are SEQ ID NOs: 177-184 in order of appearance from the beginning. B) Flow cell images of equilibrium binding and kinetic dissociation. Shown are raw images of flow cells with increasing concentrations of Her2-biotin and 100 nM AF532 streptavidin, as well as raw images of flow cells after time lapse after injecting binding buffer into the flow cell at 100 μl / min. The min / max thresholds of the images were set identically at 100 / 1000. C) Curves were fitted to the equilibrium binding and kinetic dissociation data, showing the clones from A) plus Herceptin (Trastuzumab). This shows the median integrated binding signal with processed image data plotted against the concentration of Her2-biotin. Equilibrium binding curves were fitted to the data, and error bars are shown as standard error of the mean (SEM). Given the large number of clusters observed per clone, the SEM is not visible in the plot. The second graph shows the median integrated binding signal with processed image data plotted against the wash time course. Two-phase heterogeneous dissociation kinetics were fitted to the data, and error bars are shown as SEM.
[0039] [Figure 8]Figure 8. Affinity maturation of G98A. A) Schematic of the structure of G98A showing its CDR H3 sequences and how the six scan window NNS sub-libraries are structured (sequence is SEQ ID NO: 177). B) Experimental statistics showing 159.8M clusters in the deep screening component, yielding 297k unique barcodes with 12 replicates and 236k unique CDR VH3 protein sequences. Sequences shown are SEQ ID NOs: 177, 180, 185-187. C) PCA plot of the two-dimensional projection of all 236k VH3 protein sequences, colored by mean fluorescence intensity at 100 nM Her2. Red dots indicate the position of G98A wild type with respect to the library. D) CDR H3 sequences of G98A, ML3-9, and the three top scoring clones identified in the deep screening. Next to the sequences are the deep screening (DS) fitted equilibrium binding KD and Octet determined binding KD. E) Deep screening equilibrium binding and kinetic dissociation curves for G98A, ML3-9 and the three top scoring clones. Error bars are SEM. F) Octet kinetics of G98A, ML3-9 and the three top scoring clones at 20 nM of each clone on an octet chip loaded with Her2.
[0040] [Figure 9] Figure 9. Machine learning assisted antibody engineering. A) Workflow used for in silico mutagenesis of three anti-Her2 seed sequences and selection of 13,121 random and 11,916 ML-assisted variants prior to the second round of deep screening. Sequences shown are SEQ ID NOs: 188-192. B) Evaluation of selected ML and random variants at 5 min wash conditions revealed that the binding distribution of ML variants was substantially shifted 5-fold compared to random variant generation. C) Equilibrium binding and kinetic dissociation curves on flow cells for G98A, Her20006, 13 and 19 as scFv. Error bars are SEM. D) Octet binding kinetics of purified Fab at 20 nM concentration on Her2-loaded chips.
[0041] [Figure 10] Figure 10. Equilibrium binding and kinetic dissociation curves from deep screening for anti-HEL nanobody clones selected from the MACS library for characterization. Each concentration condition in the curves represents at least 12 measurements from deep screening experiments. Error bars are SEM. The equilibrium KD, area under the curve (AUC) for equilibrium binding, and two dissociation rates and AUC for the two-phase dissociation model are reported.
[0042] [Figure 11] Figure 11. Equilibrium binding and kinetic dissociation curves from deep screening for anti-HEL nanobody clones selected by picking 96 colonies from the R3 MACS output and clones selected by R3 FACS library screening. Each concentration condition within the curve represents at least 12 measurements from the deep screening experiment. Error bars are SEM. The equilibrium KD, area under the curve (AUC) for equilibrium binding, and the two dissociation rates and AUC for the two-phase dissociation model are reported.
[0043] [Figure 12] Figure 12. Octet-measured binding and dissociation kinetics for anti-HEL nanobody clones (M1–M23) selected from the MACS library. Nanobody clones were bound at 50 nM to a streptavidin chip loaded with HEL-biotin.
[0044] [Figure 13] Figure 13. Octet-measured binding and dissociation kinetics for anti-HEL nanobody clones (C1–C8) selected by 96 colony picking and anti-HEL nanobody clones (M1–M10) selected by MACS library screening. Nanobody clones were bound at 50 nM to a streptavidin chip loaded with HEL-biotin.
[0045] [Figure 14] Figure 14. Octet measured binding and dissociation kinetics for anti-IL7 scFv clones selected for characterization, where each clone was converted from scFv to Fab, expressed, purified, and normalized to 50 nM. The Fab was then bound to a streptavidin chip preloaded with huIL7-biotin. A 1:1 model was fitted to all clones except IL70001.
[0046] [Figure 15] Figure 15. TF-1 STAT5 IL7 receptor (IL7R) α+γ luciferase inhibition assay. IL7R signaling luminescence was plotted against the log molar concentration of all characterized clones individually. Error bars are standard deviation, n=2.
[0047] [Figure 16] Figure 16. TF-1 STAT5 IL7 receptor (IL7R) α+γ luciferase inhibition assay. IL7R signaling luminescence was plotted against the log molar concentration of all characterized clones. Error bars are standard deviation, n=2.
[0048] [Figure 17] Figure 17. Equilibrium binding and kinetic dissociation curves from deep screening for anti-Her2 scFvs selected for characterization. Each concentration condition in the curves represents at least 12 measurements from either "Her2affmat" (G98A-HER20011) or "Her2 ML vs. Random" (HER20012-HER20026) deep screening experiments. Error bars are SEM. The equilibrium KD, area under the curve (AUC) for equilibrium binding, and two dissociation rates and AUC for the two-phase dissociation model are reported.
[0049] [Figure 18]Figure 18. Octet measured binding and dissociation kinetics for anti-Her2 scFv clones selected for characterization. Each clone was converted from scFv to Fab, expressed, purified, and normalized to 20 nM. The Fab was then bound to a streptavidin chip preloaded with Her2-biotin. A 1:1 model was fitted to all clones except G98A.
[0050] [Figure 19] FIG. 19 demonstrates successful display and functional fluorescent assay of FANA polymers and 2'OMe-RNA polymers on a substrate.
[0051] [Figure 20] FIG. 20 demonstrates successful display and functional fluorescent assay of peptides, fibronectin type III (FN3) scaffold, nanobodies, and scFv on substrates.
[0052] [Figure 21] FIG. 21 shows representative XNAs that can be represented by the methods of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0053] (Detailed Description) Techniques that allow the display of biomolecules on a substrate are important for enabling downstream analysis, such as high-throughput screening. The present inventors herein provide techniques that allow the display of non-DNA nucleic acids on a substrate.
[0054] In one embodiment, the invention uses a polymerase capable of synthesizing a non-DNA nucleic acid from a DNA primer to generate a biomolecule immobilized on a substrate.
[0055] Thus, in one aspect, the present invention provides a method for displaying a non-DNA nucleic acid molecule on a substrate, comprising: i) providing a first nucleic acid immobilized on a substrate, wherein the first nucleic acid is oriented such that its 5' end is proximal and its 3' end is distal to its point of immobilization; ii) generating a second nucleic acid complementary to the first nucleic acid, the generating step of the second nucleic acid comprising: a) contacting the first nucleic acid with a nucleic acid polymerase under conditions suitable for polymerization, wherein the primer for polymerization is a DNA primer immobilized on the substrate such that a bridge is formed during polymerization; the product of the polymerization is a non-DNA nucleotide strand immobilized on the substrate via the primer, and The nucleic acid polymerase is a polymerase capable of acting on a DNA primer to synthesize a non-DNA nucleic acid molecule complementary to a single-stranded nucleic acid template; and and iii) removing said first nucleic acid resulting in the presentation of said second nucleic acid on said substrate.
[0056] The resulting second nucleic acid is a single-stranded, non-DNA nucleic acid molecule displayed on the substrate. Conditions suitable for polymerization of non-DNA nucleic acids are known in the art and include the provision of appropriate nucleotides, such as RNA nucleotides or XNA nucleotides.
[0057] In some embodiments, the non-DNA nucleic acid is an XNA molecule. The XNA molecule includes a nucleotide chain having a non-natural sugar backbone, a non-natural nucleobase, a non-natural phosphodiester bond, a non-natural bond, or any combination thereof. The XNA may be polymerized by a polymerase capable of acting on a DNA primer to synthesize the XNA molecule. In particular, the XNA may be either a naturally modified nucleic acid or a non-natural nucleic acid, for which a natural or engineered polymerase can synthesize a polynucleotide from a DNA template using a DNA primer. Suitable polymerases are described herein.
[0058] For example, XNA molecules may include arabinonucleotides, which are structural analogs of deoxynucleotides and differ only by the presence of a β-hydroxyl at the 2' position of the sugar moiety. The arabinonucleotide molecule may be an arabinonucleic acid (ANA) molecule, or a 2'-fluoroarabinonucleic acid (FANA) molecule. In other embodiments, the XNA molecule may be a 2'-O-methylribonucleic acid (2'OMe) molecule, a 2'-O-methoxyethyl (MOE) nucleotide, a phosphorothioate 2'-O-methoxyethyl (PS-MOE) nucleotide, a phosphorodiamidate morpholino oligonucleotide (PMO), or a combination thereof. Alternatively, the XNA may be a P-alkylphosphonate nucleic acid (phNA). In phNAs, the non-bridging oxygen of the normal phosphodiester bond is replaced with an uncharged alkyl substituent, in particular a methyl (Met) or ethyl (Et) group. In other embodiments, the XNA molecule may be a threose nucleic acid (TNA), a hexitol nucleic acid (HNA), a 2'hydroxy-hexitol (AtNA), a cyclohexene nucleic acid (CeNA), a locked nucleic acid (LNA), or a 3'deoxy-DNA (2'-5').
[0059] In some embodiments, the non-DNA nucleic acid is an RNA molecule. The RNA molecule may include natural and non-natural modifications, such as m6A, 5-ethynyl-U, diaminopurine, phosphorothioate, 2'fluoro, 2'N3, 2'NH2, 3'O-methyl, and non-natural base pair derivatives. In some embodiments, the RNA molecule is unmodified.
[0060] The table below lists examples of RNAs and XNAs that may be presented in the methods of the invention. [Table 1] *Includes RNA containing modifications such as m6A, 5-ethynyl-U, diaminopurine, and unnatural base pairs.
[0061] The above table lists representative polymerases for the preparation of nucleic acid polymers. Representative polymerases are further described herein.
[0062] In embodiments in which RNA is displayed on a substrate, the first nucleic acid may encode a polypeptide, for example, the first nucleic acid may include an antisense sequence that can serve as a template for an RNA molecule that can be translated into a protein.
[0063] The first nucleic acid in step i) may be a nucleic acid that is part of a cluster generated on a substrate. For example, the nucleic acid may be a DNA molecule with a first adaptor at one end and a second adaptor at the other end, attached to the substrate via an immobilized primer capable of hybridizing to one of the adaptors. This nucleic acid may then be amplified into a cluster via bridge amplification, for example by using said primer and a second immobilized primer capable of hybridizing to the other adaptor. Methods for generating such clusters of DNA molecules are known in the art, and the present invention encompasses the use of all such methods. Particularly preferred methods are disclosed herein.
[0064] "Nucleic acid cluster" is a term of the art that refers to a group of nucleic acid molecules immobilized in close proximity. Close proximity is usually due to the fact that the cluster is generated by amplification from a single parent molecule, and therefore the cluster is made up of nucleic acids that contain identical sequences. Examples of clustering techniques include bridge amplification and binding equilibrium exclusion exponential amplification.
[0065] The substrate may be a solid surface such as the surface of a flow cell, a bead, a slide, or a membrane. In particular, the substrate may be a flow cell. The flow cell may be a patterned or non-patterned flow cell. The substrate may comprise glass, quartz, silica, metal, ceramic, or plastic. The substrate surface may comprise a polyacrylamide matrix or coating.
[0066] As used herein, the term "flow cell" is intended to have its usual meaning in the art, particularly in the field of sequencing by synthesis. Exemplary flow cells include, but are not limited to, those used in nucleic acid sequencing devices, such as flow cells for the Genome Analyzer®, MiSeq®, NextSeq®, HiSeq®, or NovaSeq® platforms available from Illumina, Inc. (San Diego, Calif.), or flow cells for the SOLiD™ or Ion Torrent™ sequencing platforms available from Life Technologies, Inc. (Carlsbad, Calif.). Exemplary flow cells, as well as methods for their manufacture and use, are also described, for example, in WO2014 / 142841A1, U.S. Patent Application Publication No. 2010 / 0111768A1, and U.S. Patent No. 8,951,781.
[0067] At least a part of the first nucleic acid may be sequenced before step i) of the method of presenting a non-DNA nucleic acid molecule on a substrate.For example, at least one adapter may include a barcode sequence, and the barcode may be sequenced.The location of each such barcode sequence on the substrate may be known.These techniques are known in the art.
[0068] Immobilization to a substrate means that the nucleic acid is bound to the substrate even under conditions that may denature the double-stranded nucleic acid. For example, the nucleic acid may be covalently attached to the substrate. The nucleic acid may be immobilized on a polyacrylamide-coated substrate.
[0069] The first nucleic acid is oriented with its 5' end proximal and its 3' end distal to its immobilization point. This configuration can allow for bridge amplification in combination with an immobilized primer. These configurations are standard in the art.
[0070] "Nucleic acid bridge" is a term of the art that refers to a nucleic acid attached at both ends to a substrate. Typically, one end (e.g., the 5' end) is immobilized to a substrate and the other end is attached by hybridization to a complementary nucleic acid that is itself immobilized to a substrate. Bridge amplification occurs when the template is a bridge.
[0071] In an embodiment of the present invention, the immobilized first nucleic acid is contacted with a nucleic acid polymerase under conditions suitable for polymerization. The primer for this polymerization is a DNA primer that is also immobilized on the substrate, so that a bridge is formed during polymerization. When a polymerase is used, it can act on the DNA primer to synthesize a non-DNA molecule complementary to the first nucleic acid. Thus, the polymerization product is a non-DNA nucleotide chain immobilized on the substrate via the primer.
[0072] The DNA primer may contain modified or non-DNA nucleotides. However, the DNA primer does not contain RNA nucleotides at its 3' end, and therefore the polymerase is a polymerase that does not require an RNA primer. In particular, the method of the present invention is suitable for use with commercially available adapter / primers such as Illumina's P5 and P7 adapters, and the method does not require the primer to be modified with ribonucleotides. For example, 3Dpol from poliovirus is an RNA-dependent RNA polymerase that cannot act on DNA primers.
[0073] Polymerases capable of acting on DNA primers to synthesize XNA polymers are disclosed in documents such as Arangundy-Franklin et al. (Nature Chemistry, Vol. 11, pp. 533-542 (2019)), WO2011 / 135280, and WO2013 / 156786. These documents disclose that mutations in the backbone of a polB family polymerase, excluding viral polymerases, can render the polymerase capable of synthesizing XNA polymers. In particular, the backbone may be a polymerase from the archaeal genera Thermococcus and / or Pyrococcus.
[0074] The polymerase may be a polymerase variant from T. goronarius that has been mutated to enable it to polymerize XNA molecules. The polymerase may comprise an amino acid sequence having at least at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO: 1. The polymerase may comprise an amino acid sequence having at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO: 1 and including a mutation that enables it to polymerize RNA or XNA.
[0075] The nucleic acid polymerase may be a polymerase capable of acting on a DNA primer to synthesize an RNA or XNA molecule, such as 2'F-RNA, 2'N3-RNA, 2'NH2-RNA, or PS-RNA, complementary to a single-stranded nucleic acid template. Such polymerases include any polymerase capable of synthesizing an RNA or XNA molecule, as disclosed in WO2011 / 135280 or Cozens et al. (Cozens, Pinheiro, Vaisman, Woodgate, and Holliger, "A short adaptive path from DNA to RNA polymerases," PNAS, May 22, 2012, 109(21)8067-8072; https: / / doi.org / 10.1073 / pnas.1120964109), each of which is incorporated herein by reference. For example, the polymerase may be D4N, TNQ, TNK, or TGK, or a variant thereof, as disclosed in these documents. In particular, the polymerase may be TGK or a variant thereof.
[0076] The polymerase may comprise mutations corresponding to Y409N or Y409G and E664K or E664Q (as described with respect to SEQ ID NO: 1) in the backbone of a polymerase from the polB family. In a particular embodiment, the backbone is any polB polymerase, except viral polymerases. The backbone may be that of a polymerase from the archaeal genera Thermococcus and / or Pyrococcus.
[0077] The polymerase may be a variant of the polymerase from T. goronarius (Tgo) that contains additional mutations that activate the RNA polymerase. The sequence of wild-type Tgo is shown below. [ka]
[0078] The polymerase may have an amino acid sequence having at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99% or 100% similarity or identity to the amino acid sequence of SEQ ID NO: 1, the amino acid sequence comprising Y409 and E664 mutations with respect to the amino acid sequence of SEQ ID NO: 1. In an embodiment, the Y409 mutation is Y409N or Y409G and the E664 mutation is E664K or E664Q. In a particular embodiment, the Y409 and E664 mutations are a combination of i) Y409N and E664Q, ii) Y409N and E664k, or iii) Y409G and E664K. The amino acid sequence of the nucleic acid polymerase may further comprise one or more or all of the following mutations: A485L, V93Q, D141A, and E143A.
[0079] V93Q is a mutation known to relieve uracil stop, D141A and E143A reduce 3'-5' exonuclease function, and the "Therminator" mutation (A485L) is known to promote incorporation of unnatural substrates. The sequence of Tgo polymerase containing these mutations (hereafter referred to as "TgoT") is shown below. [ka]
[0080] The polymerase may be of SEQ ID NO:2 containing the following mutations: i) Y409N and E664Q (TNQ), ii) Y409N and E664K (TNK), or iii) Y409G E664K (TGK).
[0081] In one preferred embodiment, the amino acid sequence of the nucleic acid polymerase comprises SEQ ID NO:1 as follows and the mutations V93Q, D141A, E143A, Y409G, A485L, and E664K (TGK). [ka]
[0082] The polymerase may have an amino acid sequence that has at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO:3, and in which residues 93, 141, 143, 409, 485, and 664 are unchanged (i.e., the mutations V93Q, D141A, E143A, Y409G, A485L, and E664K are maintained).
[0083] In a particular embodiment, there is provided a method for displaying RNA molecules on a substrate, comprising: i) providing a first nucleic acid immobilized on a substrate, wherein the first nucleic acid is oriented such that its 5' end is proximal and its 3' end is distal to its point of immobilization; ii) generating a second nucleic acid complementary to the first nucleic acid, the generating step of the second nucleic acid comprising: a) contacting the first nucleic acid with a nucleic acid polymerase under conditions suitable for RNA polymerization, wherein the primer for polymerization is a DNA primer immobilized on the substrate such that a bridge is formed during polymerization; the product of the polymerization is an RNA nucleotide strand immobilized on the substrate via the primer, and The nucleic acid polymerase is a polymerase capable of acting on a DNA primer to synthesize an RNA molecule complementary to a single-stranded nucleic acid template, and has at least 80%, 90%, 95%, 99%, or 100% identity to the amino acid sequence of SEQ ID NO:3, and includes a sequence in which residues 93, 141, 143, 409, 485, and 664 are invariant; and and iii) removing said first nucleic acid resulting in the presentation of said second nucleic acid on said substrate. In this embodiment, the second nucleic acid is an RNA polymer.
[0084] The nucleic acid polymerase may be a polymerase capable of acting on a DNA primer to synthesize an arabinonucleotide polymer, such as an XNA, e.g., an ANA molecule or a FANA molecule, that is complementary to a single-stranded nucleic acid template. Such polymerases include any polymerase capable of synthesizing an arabinonucleotide polymer molecule, such as those disclosed in WO2013 / 156786A1, which is incorporated herein by reference. In one particular embodiment, the polymerase may be the D4YK polymerase disclosed in WO2013 / 156786A1. Such polymerases include any polymerase capable of synthesizing a polymer, such as those disclosed in Pinheiro et al., "Synthetic genetic polymers capable of heredity and evolution," Science. 2012 Apr 20;336(6079):341-344. For example, the polymerase may be D4K or a variant thereof as disclosed in the literature.
[0085] The polymerase may comprise mutations corresponding to P657T, E658Q, K659H, Y663H, E664K, D669A, K671N, and T676I (as described with respect to SEQ ID NO: 1) in the backbone of a polymerase from the polB family. The polymerase may further comprise L403P. In a particular embodiment, the backbone is any polB polymerase except viral polymerases. The backbone may be of a polymerase from the archaeal genera Thermococcus and / or Pyrococcus. The polymerase may be a variant of the polymerase from T. gorgonarius (Tgo) (SEQ ID NO: 1).
[0086] The L403P mutation is an additional useful mutation within the A motif of the polymerase. It has the advantage of facilitating polymerization and can help produce longer polymers. It can improve the polymerization of arabinonucleotides by 3-4 fold or more. Depending on the application, this improvement can be 10-fold.
[0087] The polymerase may be an amino acid sequence having at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO:1, which may include the mutations P657T, E658Q, K659H, Y663H, E664K, D669A, K671N, and T676I, and optionally L403P, relative to the amino acid sequence of SEQ ID NO:1.
[0088] The amino acid sequence of the nucleic acid polymerase may further comprise one or more or all of the following mutations: V93Q, D141A, E143A, and A485L, which are described elsewhere herein.
[0089] In one particular embodiment, the nucleic acid polymerase capable of acting on a DNA primer to synthesize an arabinonucleotide polymer has an amino acid sequence having at least 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO:1, and optionally includes the mutations P657T, E658Q, K659H, Y663H, E664K, D669A, K671N, T676I, V93Q, D141A, E143A, L403P, and A485L with respect to the amino acid sequence of SEQ ID NO:1.
[0090] In one particular embodiment, the nucleic acid polymerase capable of acting on a DNA primer to synthesize an arabinonucleotide polymer has the following amino acid sequence: [ka] (SEQ ID NO: 4; also known as D4YK or D4K). It may include or be the above.
[0091] The polymerase may have an amino acid sequence that has at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO:4, and where residues 93, 141, 143, 403, 485, 657, 658, 659, 663, 664, 669, 671, and 676 are unchanged (i.e., the mutations V93Q, D141A, E143A, L403P, A485L, P657T, E658Q, K659H, Y663H, E664K, D669A, K671N, and T676I are maintained).
[0092] In certain embodiments, a method for displaying arabinonucleotide polymers on a substrate comprises: i) providing a first nucleic acid immobilized on a substrate, wherein the first nucleic acid is oriented such that its 5' end is proximal and its 3' end is distal to its point of immobilization; ii) generating a second nucleic acid complementary to the first nucleic acid, the generating step of the second nucleic acid comprising: a) contacting the first nucleic acid with a nucleic acid polymerase under conditions suitable for arabinonucleotide polymerization, wherein the primer for polymerization is a DNA primer immobilized on the substrate such that a bridge is formed during polymerization; the product of the polymerization is an arabinonucleotide chain immobilized on the substrate via the primer, and the nucleic acid polymerase is capable of acting on a DNA primer to synthesize an arabinonucleotide polymer complementary to a single-stranded nucleic acid template; and having at least 80%, 90%, 95%, 99%, or 100% identity to the amino acid sequence of SEQ ID NO:4, with residues 93, 141, 143, 403, 485, 657, 658, 659, 663, 664, 669, 671, and 676 being unchanged; and and iii) removing said first nucleic acid resulting in the presentation of said second nucleic acid on said substrate. In this embodiment, the second nucleic acid is a single-stranded arabinonucleotide polymer presented on a substrate. In one particular embodiment, the arabinonucleotide polymer presented on a substrate is an ANA molecule or a FANA molecule.
[0093] The nucleic acid polymerase may be a polymerase capable of acting on a DNA primer to synthesize an XNA molecule, such as 2'OMe, MOE, PS-MOE, or LNA polymer, complementary to a single-stranded nucleic acid template. Such polymerases include polymerases that contain mutations corresponding to Y409G, I521L, T541G, F545L, K592A, and E664K (described with reference to SEQ ID NO: 1) in the backbone of a polymerase from the polB family. In a particular embodiment, the backbone is any polB polymerase, except viral polymerases. The backbone may be of a polymerase from the archaeal genera Thermococcus and / or Pyrococcus. The polymerase may be a variant of a polymerase from T. goronarius (Tgo).
[0094] The polymerase may have an amino acid sequence having at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO:1, and the amino acid sequence may include the mutations Y409G, I521L, T541G, F545L, K592A, and E664K with respect to the amino acid sequence of SEQ ID NO:1.
[0095] The amino acid sequence of the nucleic acid polymerase may further comprise one or more or all of the following mutations: V93Q, D141A, E143A, and A485L, which are described elsewhere herein.
[0096] In one particular embodiment, the nucleic acid polymerase is capable of acting on a DNA primer to synthesize a 2'OMe, MOE, PS-MOE or LNA polymer, and has an amino acid sequence having at least 80%, 90%, 95%, 99% or 100% similarity or identity to the amino acid sequence of SEQ ID NO:1, which may include the mutations V93Q, D141A, E143A, Y409G, A485L, I521L, T541G, F545L, K592A and E664K with respect to the amino acid sequence of SEQ ID NO:1.
[0097] In one particular embodiment, the nucleic acid polymerase is capable of acting on a DNA primer to synthesize a 2'OMe, MOE, PS-MOE, or LNA polymer having the following amino acid sequence: [ka] (SEQ ID NO:20; also known as 2M polymerase.) It may include or be the above.
[0098] The polymerase may have an amino acid sequence that has at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO:20, and wherein residues 93, 141, 143, 409, 485, 521, 541, 545, 592, and 664 are unchanged (i.e., the mutations V93Q, D141A, E143A, Y409G, A485L, I521L, T541G, F545L, K592A, and E664K are maintained).
[0099] In a particular embodiment, a method for presenting 2'-O-methyl ribonucleotide polymers or 2'-O-methoxyethyl nucleotide polymers on a substrate comprises: i) providing a first nucleic acid immobilized on a substrate, wherein the first nucleic acid is oriented such that its 5' end is proximal and its 3' end is distal to its point of immobilization; ii) generating a second nucleic acid complementary to the first nucleic acid, the generating step of the second nucleic acid comprising: a) contacting the first nucleic acid with a nucleic acid polymerase under conditions suitable for 2'-O-methylribonucleotides or 2'-O-methoxyethyl nucleotide polymerization, wherein the primer for polymerization is a DNA primer immobilized on the substrate such that a bridge is formed during polymerization; the product of the polymerization is a 2'-O-methylribonucleotide or 2'-O-methoxyethyl nucleotide chain immobilized on the substrate via the primer; and the nucleic acid polymerase is capable of acting on a DNA primer to synthesize 2'-O-methylribonucleotides or 2'-O-methoxyethyl nucleotides complementary to a single-stranded nucleic acid template; and having at least 80%, 90%, 95%, 99%, or 100% identity to the amino acid sequence of SEQ ID NO:20, with residues 93, 141, 143, 409, 485, 521, 541, 545, 592, and 664 remaining unchanged; and and iii) removing said first nucleic acid resulting in the presentation of said second nucleic acid on said substrate. In this embodiment, the second nucleic acid is a single stranded 2'-O-methyl ribonucleotide polymer or a 2'-O-methoxyethyl nucleotide polymer presented on a substrate.
[0100] The nucleic acid polymerase may be a polymerase capable of acting on a DNA primer to synthesize 2'NH2-RNA, 2'O-methyl-RNA, 3'deoxy-DNA (2'-5'), or 3'O-methyl-RNA polymers complementary to a single-stranded nucleic acid template. Such polymerases include any polymerase capable of synthesizing a polymer as disclosed in Cozens et al. (Cozens, Mutschler, Nelson, Houlihan, Taylor, and Holliger, "Enzymatic Synthesis of Nucleic Acids with Defined Regioisomeric 2'-5' Linkages," Angew Chem Int Ed Engl. 2015 Dec 14;54(51): 15570-15573), which is incorporated herein by reference. For example, the polymerase may be TGLLK or a variant thereof as disclosed in the literature.
[0101] The polymerase may comprise mutations corresponding to Y409G, I521L, F545L, and E664K (as described with respect to SEQ ID NO: 1) in the backbone of a polymerase from the polB family. In a particular embodiment, the backbone is any polB polymerase except viral polymerases. The backbone may be of a polymerase from the archaeal genera Thermococcus and / or Pyrococcus. The polymerase may be a variant of the polymerase from T. gorgonarius (Tgo) (SEQ ID NO: 1).
[0102] The polymerase may have an amino acid sequence having at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99% or 100% similarity or identity to the amino acid sequence of SEQ ID NO: 1, and the amino acid sequence may include the mutations Y409G, I521L, F545L, and E664K with respect to the amino acid sequence of SEQ ID NO: 1. The amino acid sequence of the nucleic acid polymerase may further include one or more or all of the following mutations V93Q, D141A, E143A, and A485L, which mutations are described elsewhere herein.
[0103] In one particular embodiment, the nucleic acid polymerase has an amino acid sequence having at least 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO:1, and optionally includes the mutations V93Q, D141A, E143A, Y409G, A485L, I521L, F545L, and E664K relative to the amino acid sequence of SEQ ID NO:1 (the polymerase with 100% identity is sometimes referred to as TGLLK).
[0104] The nucleic acid polymerase may be a polymerase capable of acting on a DNA primer to synthesize an XNA polymer, such as a TNA polymer, complementary to a single-stranded nucleic acid template. Such polymerases include any polymerase capable of synthesizing a polymer, such as those disclosed in Chen and Romesberg (FEBS Lett. 2014 Jan 21;588(2): 219-229) or Pinheiro et al. ("Synthetic genetic polymers capable of heredity and evolution" Science. 2012 Apr 20;336(6079): 341-344), each of which is incorporated herein by reference. For example, the polymerase may be RT521 or a variant thereof as disclosed in the literature. As disclosed in these literature, this polymerase is capable of synthesizing XNA polymers other than TNA.
[0105] The polymerase may comprise mutations corresponding to E429G, I521L, and K726R (as described with respect to SEQ ID NO: 1) in the backbone of a polymerase from the polB family. In a particular embodiment, the backbone is any polB polymerase except viral polymerases. The backbone may be of a polymerase from the archaeal genera Thermococcus and / or Pyrococcus. The polymerase may be a variant of the polymerase from T. gorgonarius (Tgo) (SEQ ID NO: 1).
[0106] The polymerase may have an amino acid sequence having at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99% or 100% similarity or identity to the amino acid sequence of SEQ ID NO: 1, which may include the mutations E429G, I521L, and K726R with respect to the amino acid sequence of SEQ ID NO: 1. The amino acid sequence of the nucleic acid polymerase may further include one or more or all of the following mutations V93Q, D141A, E143A, and A485L, which mutations are described elsewhere herein.
[0107] In one particular embodiment, the nucleic acid polymerase has an amino acid sequence having at least 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO:1, and optionally includes the mutations V93Q, D141A, E143A, E429G, A485L, I521L, and K726R relative to the amino acid sequence of SEQ ID NO:1 (the polymerase with 100% identity is sometimes referred to as RT521).
[0108] The nucleic acid polymerase may be a polymerase capable of acting on a DNA primer to synthesize an XNA polymer, such as an HNA polymer, complementary to a single-stranded nucleic acid template. Such polymerases include any polymerase capable of synthesizing such polymers, as disclosed in Taylor et al. ("Catalysts from synthetic genetic polymers" Nature. 2015 Feb 19;518(7539): 427-430) or Pinheiro et al. ("Synthetic genetic polymers capable of heredity and evolution" Science. 2012 Apr 20;336(6079): 341-344), each of which is incorporated herein by reference. For example, the polymerase may be 6G12 or a variant thereof as disclosed in the literature. As disclosed in these publications, this polymerase is capable of synthesizing XNA polymers other than HNA.
[0109] The polymerase may comprise mutations corresponding to V589A, E609K, I610M, K659Q, E664Q, Q665P, R668K, D669Q, K671H, K674R, T676R, A681S, L704P, and E730G (as described with reference to SEQ ID NO: 1) in the backbone of a polymerase from the polB family. In a particular embodiment, the backbone is any polB polymerase except viral polymerases. The backbone may be of a polymerase from the archaeal genera Thermococcus and / or Pyrococcus. The polymerase may be a variant of the polymerase from T. gorgonarius (Tgo) (SEQ ID NO: 1).
[0110] The polymerase may have an amino acid sequence having at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99% or 100% similarity or identity to the amino acid sequence of SEQ ID NO: 1, the amino acid sequence including the mutations V589A, E609K, I610M, K659Q, E664Q, Q665P, R668K, D669Q, K671H, K674R, T676R, A681S, L704P, and E730G with respect to the amino acid sequence of SEQ ID NO: 1. The amino acid sequence of the nucleic acid polymerase may further include one or more or all of the following mutations V93Q, D141A, E143A, and A485L, which are described elsewhere herein.
[0111] In one particular embodiment, the nucleic acid polymerase has an amino acid sequence having at least 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO:1, and optionally includes the mutations V93Q, D141A, E143A, A485L, V589A, E609K, I610M, K659Q, E664Q, Q665P, R668K, D669Q, K671H, K674R, T676R, A681S, L704P, and E730G relative to the amino acid sequence of SEQ ID NO:1 (the polymerase with 100% identity is sometimes referred to as 6G12).
[0112] The nucleic acid polymerase may be a polymerase capable of acting on a DNA primer to synthesize an XNA polymer, such as an HNA, AtNA, CeNA, or LNA polymer, complementary to a single-stranded nucleic acid template. Such polymerases include any polymerase capable of synthesizing polymers, such as those disclosed in Taylor et al., ("Catalysts from synthetic genetic polymers" Nature. 2015 Feb 19;518(7539): 427-430) or Mutschler et al., ("Random-sequence genetic oligomer pools display an innate potential for ligation and recombination; eLife 2018;7:e43022 DOI: 10.7554 / eLife.4302), each of which is incorporated herein by reference. For example, the polymerase can be the 6G12 I521L variant ("6G12521") or variants thereof disclosed therein.
[0113] The polymerase may comprise mutations corresponding to I521L, V589A, E609K, I610M, K659Q, E664Q, Q665P, R668K, D669Q, K671H, K674R, T676R, A681S, L704P, and E730G (as described with reference to SEQ ID NO: 1) in the backbone of a polymerase from the polB family. In a particular embodiment, the backbone is any polB polymerase except viral polymerases. The backbone may be of a polymerase from the archaeal genera Thermococcus and / or Pyrococcus. The polymerase may be a variant of the polymerase from T. gorgonarius (Tgo) (SEQ ID NO: 1).
[0114] The polymerase may have an amino acid sequence having at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99% or 100% similarity or identity to the amino acid sequence of SEQ ID NO: 1, the amino acid sequence including the mutations I521L, V589A, E609K, I610M, K659Q, E664Q, Q665P, R668K, D669Q, K671H, K674R, T676R, A681S, L704P, and E730G with respect to the amino acid sequence of SEQ ID NO: 1. The amino acid sequence of the nucleic acid polymerase may further include one or more or all of the following mutations V93Q, D141A, E143A, and A485L, which are described elsewhere herein.
[0115] In one particular embodiment, the nucleic acid polymerase has an amino acid sequence having at least 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO:1, and optionally includes the mutations V93Q, D141A, E143A, A485L, I521L, V589A, E609K, I610M, K659Q, E664Q, Q665P, R668K, D669Q, K671H, K674R, T676R, A681S, L704P, and E730G relative to the amino acid sequence of SEQ ID NO:1 (the polymerase with 100% identity is sometimes referred to as 6G12521).
[0116] The nucleic acid polymerase may be a polymerase capable of acting on a DNA primer to synthesize an XNA polymer, such as CeNa or LNA polymer, complementary to a single-stranded nucleic acid template. Such polymerases include any polymerase capable of synthesizing such polymers, as disclosed in Pinheiro et al. ("Synthetic genetic polymers capable of heredity and evolution" Science. 2012 Apr 20;336(6079):341-344). For example, the polymerase may be PolC7 (also known as "C7") or a variant thereof, as disclosed in the literature.
[0117] The polymerase may comprise mutations corresponding to E654Q, E658Q, K659Q, V661A, E664Q, Q665P, D669A, K671Q, T676K, and R709K (as described with reference to SEQ ID NO: 1) in the backbone of a polymerase from the polB family. In a particular embodiment, the backbone is any polB polymerase except viral polymerases. The backbone may be of a polymerase from the archaeal genera Thermococcus and / or Pyrococcus. The polymerase may be a variant of the polymerase from T. gorgonarius (Tgo) (SEQ ID NO: 1).
[0118] The polymerase may have an amino acid sequence having at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99% or 100% similarity or identity to the amino acid sequence of SEQ ID NO: 1, and the amino acid sequence may include the mutations E654Q, E658Q, K659Q, V661A, E664Q, Q665P, D669A, K671Q, T676K, and R709K with respect to the amino acid sequence of SEQ ID NO: 1. The amino acid sequence of the nucleic acid polymerase may further include one or more or all of the following mutations V93Q, D141A, E143A, and A485L, which mutations have been described elsewhere herein.
[0119] In one particular embodiment, the nucleic acid polymerase has an amino acid sequence having at least 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO:1, and optionally includes the mutations V93Q, D141A, E143A, A485L, E654Q, E658Q, K659Q, V661A, E664Q, Q665P, D669A, K671Q, T676K, and R709K relative to the amino acid sequence of SEQ ID NO:1 (the polymerase with 100% identity is sometimes referred to as C7).
[0120] The nucleic acid polymerase may be a polymerase capable of acting on a DNA primer to synthesize an XNA molecule, such as a phNA, PMO, or P-alkyl-moNA molecule, complementary to a single-stranded nucleic acid template. Such polymerases include any polymerase capable of synthesizing a phNA molecule, such as those disclosed in Arangundy-Franklin et al. (Nature Chemistry, Vol. 11, pp. 533-542 (2019)), which is incorporated herein by reference. In particular, the polymerase may be "GV", "GV2", or "PGV2" (also known as "PGLVV") or a variant thereof, as disclosed therein.
[0121] The polymerase may comprise mutations corresponding to E429G, D455P, K487G, I521L, R606V, R613V, and K726R (as described with respect to SEQ ID NO: 1) in the backbone of a polymerase from the polB family. In a particular embodiment, the backbone is any polB polymerase except viral polymerases. The backbone may be of a polymerase from the archaeal genera Thermococcus and / or Pyrococcus. The polymerase may be a variant of the polymerase from T. gorgonarius (Tgo) (SEQ ID NO: 1).
[0122] The polymerase may have an amino acid sequence having at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99% or 100% similarity or identity to the amino acid sequence of SEQ ID NO: 1, and the amino acid sequence may include the mutations E429G, D455P, K487G, I521L, R606V, R613V, and K726R with respect to the amino acid sequence of SEQ ID NO: 1. The amino acid sequence of the nucleic acid polymerase may further include one or more or all of the following mutations V93Q, D141A, E143A, and A485L, which mutations are described elsewhere herein.
[0123] In one particular embodiment, the nucleic acid polymerase has an amino acid sequence having at least 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO:1, and optionally includes the mutations V93Q, D141A, E143A, E429G, D455P, A485L, K487G, I521L, R606V, R613V, and K726R relative to the amino acid sequence of SEQ ID NO:1 (the 100% identical polymerase is sometimes referred to as PGV2 or PGLVV).
[0124] The nucleic acid polymerase may be a polymerase capable of acting on a DNA primer to synthesize an XNA molecule, such as a phNA, PMO, or P-alkyl-moNA molecule, complementary to a single-stranded nucleic acid template. In one embodiment, the polymerase may include mutations corresponding to N269W, E429G, D455P, K487G, I521L, V589A, R606V, R613V, and K726R (as described with reference to SEQ ID NO: 1) in the backbone of a polymerase from the polB family. In a particular embodiment, the backbone is any polB polymerase, except viral polymerases. The backbone may be of a polymerase from the archaeal genera Thermococcus and / or Pyrococcus. The polymerase may be a variant of the polymerase from T. gorgonarius (Tgo) (SEQ ID NO: 1).
[0125] The polymerase may have an amino acid sequence having at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99% or 100% similarity or identity to the amino acid sequence of SEQ ID NO: 1, and the amino acid sequence may include the mutations N269W, E429G, D455P, K487G, I521L, V589A, R606V, R613V, and K726R with respect to the amino acid sequence of SEQ ID NO: 1. The amino acid sequence of the nucleic acid polymerase may further include one or more or all of the following mutations V93Q, D141A, E143A, and A485L, which mutations are described elsewhere herein.
[0126] In one particular embodiment, the nucleic acid polymerase has an amino acid sequence having at least 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO:1, and optionally includes the mutations V93Q, D141A, E143A, N269W, E429G, D455P, A485L, K487G, I521L, V589A, R606V, R613V, and K726R with respect to the amino acid sequence of SEQ ID NO:1 (the polymerase with 100% identity is sometimes referred to as PGLVVWA).
[0127] When using alternative polymerase backbones, mutations are introduced at equivalent positions known in the art. For example, for the representative polymerase 6G12, the table below shows how mutations can be introduced into alternative backbones. The table shows the Pol6G12 mutations and the structurally equivalent positions in other PolBs. The mutations found in Pol6G12 are shown relative to the underlying wild-type Tgo sequence. The structurally equivalent residues in other well-studied B-family polymerases are shown. Residues that did not map to an equivalent position are shown as ND. [Table 2]
[0128] A mutation may refer to a substitution or a truncation or deletion of the referred to residue, motif or domain. In one particular embodiment, a mutation is a substitution of one type of amino acid residue for another type of amino acid residue.
[0129] The polymerase may be a polymerase fragment that maintains polymerase function.
[0130] The conditions suitable for polymerization in step ii)a) may be a cycle comprising a denaturation step, an annealing step, and an amplification step. The denaturation step may be the application of a denaturation buffer, for example a buffer comprising 98% formamide and / or NaOH. The NaOH may be at a concentration of 1 mM NaOH, preferably 10 mM NaOH or higher. The denaturation buffer may also comprise EDTA, for example 1 mM EDTA. The annealing step may be the application of a premixed buffer, which may comprise the same components as the amplification buffer without NTPs or polymerase. For example, the premixed buffer may comprise 2 M betaine, 20 mM Tris, 10 mM ammonium sulfate, 6 mM MgSO4, 0.1% Triton-X, 1.3% DMSO, and 18.1 U / ml RNAse inhibitor at pH 8.8. The amplification step may comprise contacting the substrate-bound nucleic acid with a polymerase, RNA nucleotide triphosphates, and a suitable amplification buffer. For example, the amplification buffer may include 2M betaine, 20 mM Tris, 10 mM ammonium sulfate, 6 mM MgSO4, 0.1% Triton-X, 1.3% DMSO, 625 uM NTPs, 10 nM TGK polymerase, and 18.1 U / ml RNAse inhibitor at pH 8.8. In another example, the amplification buffer may include 20 mM Tris-HCL, 10 mM (NH4)2SO4, 10 mM KCl, 2 mM MgSO4, 0.1% Triton X-100, 200 uM faNTPs, 10 nM D4YK pol at pH 8.8. In a third example, the amplification buffer may include 200 uM 2'OMe NTP, 10 nM 2M polymerase, 2 M betaine, 20 mM Tris, 10 mM ammonium sulfate, 6 mM MgSO4, 0.1% Triton-X, 1.3% DMSO at pH 8.8. In some embodiments, at least 5, 10, 12, 15, 20, or 25 cycles of bridge amplification are performed. In some embodiments, the RNAse inhibitor may be or include SuperaseIn, RNAseOUT, RNasein, RiboSafe, or other commercially available products that do not inhibit polymerase activity.
[0131] In addition to the above disclosure, the inventors provide further steps to improve the synthesis of non-DNA polymers in the method of the invention. These further steps are particularly relevant for long structures. Without being bound by theory, the inventors believe that during the synthesis of non-DNA or RNA in the bridge, the dsDNA:RNA complex (or other non-DNA nucleic acid complex) starts to accumulate a significant amount of torque, which causes the polymerase to slow down and eventually stop. The inventors have overcome this problem. See for example FIG. 2, which shows the improvements related to this aspect of the invention.
[0132] Thus, in one embodiment, the method further comprises a step performed after the first polymerization step or cycle, in which the first nucleic acid is cleaved. This cleavage allows the bridge to be linearized and the torque to be released, while the first nucleic acid is maintained. In an embodiment in which a polypeptide is presented, the cleavage site should be after the open reading frame encoding the polypeptide, to avoid interference with further rounds of polymerization. Thus, the cleavage site in the first nucleic acid is located 5' to the sequence in the first nucleic acid that corresponds to the encoded polypeptide. In a particular embodiment, the cleavage site is in an immobilized adaptor / primer that couples the first nucleic acid to the substrate.
[0133] In some embodiments, this step applies to methods involving a first nucleic acid that is greater than 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, or 1500 nucleotides in length. In one particular embodiment, this step applies to methods involving a first nucleic acid that is greater than 800 or 900 nucleotides in length. This is particularly relevant to embodiments in which the second nucleic acid is an RNA molecule, since RNA molecules may encode polypeptides and therefore are generally longer.
[0134] The cleavage may be any that allows for targeted cleavage of the first nucleic acid in a manner that does not alter other components, such as the newly formed nucleic acid strand. In a particular embodiment, the cleavage site may be incorporated into the adaptor / primer that couples the first nucleic acid to the substrate. For example, the cleavage site may be a 2-deoxyuridine that is cleavable with the Uracil Specific Removal Reagent (USER) enzyme. Alternatively, the cleavage site may be an 8-oxoguanine that is cleavable with formamidopyrimidine DNA glycosylase (Fpg). When using Fpg to cleave the first nucleic acid, the inventors have found that cleavage efficiency increases when the cleavage site is double-stranded DNA. Thus, in one particular embodiment, a third nucleic acid hybridizes to the cleavage site. The third nucleic acid may be a DNA oligo that is complementary to a sequence that spans the cleavage site, such as a DNA oligo that is hybridizable to the 8-oxoguanine site of an Illumina P7 adaptor. The 3' end of the third nucleic acid may be modified to prevent extension of the third nucleic acid during the method. For example, the 3' end of the third nucleic acid may be phosphorylated.
[0135] After cleavage, the bridges exist in a linear form. The temperature may be increased until the bridges are denatured.
[0136] The linearized product is then recontacted with the polymerase under conditions suitable for polymerization. The inventors have discovered that these steps increase the efficiency of polymerization, leading to high yields of intact non-DNA molecules.
[0137] In one particular embodiment, the first nucleic acid is contacted with a nucleic acid polymerase under conditions suitable for polymerization, where a primer for the polymerization is immobilized on the substrate such that a bridge is formed during polymerization, and at least 5, 10, 12, 15, 20, or 25 cycles of bridge amplification are performed. In an example embodiment, 12 cycles are performed, but this number may be increased. After the bridging amplification cycle, the first nucleic acid is cleaved and the bridge is linearized. Polymerization is then performed again. The polymerization after the linearization step does not need to include a denaturation step, the absence of which avoids, for example, dissociation of the DNA:RNA duplex.
[0138] The steps of cleavage, linearization, and continued polymerization can be cycled, for example, for two cycles, and in other embodiments, for three, four, five, or more cycles.
[0139] Thus, in one particular embodiment, there is provided a method for displaying RNA molecules on a substrate, comprising: i) providing a first nucleic acid immobilized on a substrate, wherein the first nucleic acid is oriented such that its 5' end is proximal and its 3' end is distal to its point of immobilization; ii) generating a second nucleic acid complementary to the first nucleic acid, the generating step of the second nucleic acid comprising: a) contacting the first nucleic acid with a nucleic acid polymerase under conditions suitable for RNA polymerization, wherein the primer for polymerization is a DNA primer immobilized on the substrate such that a bridge is formed during polymerization; The product of the polymerization is RNA immobilized on the substrate via the primer; The nucleic acid polymerase is a polymerase capable of acting on a DNA primer to synthesize an RNA molecule complementary to a single-stranded nucleic acid template, including those disclosed herein; and At least 5, 10, 12, 15, 20, or 25 cycles of bridge amplification are performed; b) cleaving the first nucleic acid to linearize the bridge; and c) contacting the linearized product of step b) with a polymerase under conditions suitable for RNA polymerization; d) at least one further cycle of step c) after step b) has been performed; and and iii) removing the first nucleic acid resulting in the presentation of the second nucleic acid on the substrate. The first nucleic acid may comprise an antisense sequence encoding a polypeptide. The cleavage in step ii)b) may be at a site 5' to the encoded polypeptide sequence.
[0140] The surprisingly efficient process for polymerization using substrate-bound templates is also applicable when the polymerase used cannot act on DNA primers, e.g., the steps of polymerization, linearization, and then additional polymerization are also applicable to other methods that utilize polymerases that act on, for example, RNA or non-DNA primers.
[0141] Thus, in an embodiment of the invention there is provided a method for displaying non-DNA nucleic acid molecules on a substrate comprising: i) providing a first nucleic acid immobilized on a substrate, wherein the first nucleic acid is oriented such that its 5' end is proximal and its 3' end is distal to its point of immobilization; ii) generating a second nucleic acid complementary to the first nucleic acid, the generating step of the second nucleic acid comprising: a) contacting the first nucleic acid with a nucleic acid polymerase under conditions suitable for polymerization, wherein a primer for the polymerization is immobilized on the substrate such that a bridge is formed during polymerization; and The product of the polymerization is a nucleotide chain immobilized on the substrate via the primer; b) cleaving the first nucleic acid to linearize the bridge; and c) contacting the linearized product of step b) with a polymerase under conditions suitable for polymerization; and and iii) removing the first nucleic acid resulting in the presentation of the second nucleic acid on the substrate. In embodiments where the second nucleic acid molecule is RNA, the first nucleic acid may comprise an antisense sequence encoding a polypeptide. The cleavage in step ii)b) may be at a site 5' to the encoded polypeptide sequence.
[0142] Further details of the above method may be as disclosed herein. For example, the number of cycles of bridge amplification and / or the number of cycles of linearization and repolymerization may be as described in the preceding paragraphs. The buffer may be as disclosed in the preceding paragraphs. However, as stated above, the polymerase may be any of those disclosed herein, although the method is not limited thereto, and the polymerase may be, for example, a 3D pol It may be a polymerase.
[0143] In step iii) of the method for presenting a non-DNA nucleic acid molecule on a substrate, the first nucleic acid is removed and the newly synthesized nucleic acid molecule is presented on the substrate as a single stranded nucleic acid molecule. Various techniques are available for removing the first nucleic acid, and in one particular embodiment, the first nucleic acid is removed by use of a denaturing agent. For example, the denaturing agent may be a buffer comprising 1-500 mM, 10-400 mM, 25-300 mM, 50-200 mM, or 75-125 mM NaOH. In one embodiment, the denaturing agent comprises 100 mM NaOH. The denaturing agent may comprise 0-20 mM EDTA. In one embodiment, the denaturing agent comprises 5 mM EDTA. The denaturing agent may comprise 100 mM NaOH and 5 mM EDTA, and the substrate-nucleic acid complex may be contacted with the buffer. In a particular embodiment, step iii) does not involve the use of DNase1.
[0144] The method of presenting non-DNA nucleic acid molecules thus results in a substrate with immobilized nucleic acid molecules on its surface. As discussed herein, the nucleic acid molecules may be present in clusters, and sequence and position information may be obtained. The presented nucleic acid molecules may form a library. For example, a library of aptamers, RNA, XNA, FANA, ANA, or 2'-OMe aptamers. The library may be of XNAzymes, for example, XNAzyme enzymes composed of FANA polymers or other XNA polymers. Alternatively, the nucleic acid molecules themselves may be presented for analysis. For example, the binding of the molecules to non-DNA nucleic acid molecules may be evaluated.
[0145] In embodiments involving the presentation of an XNAzyme, a nucleic acid oligo, such as a DNA oligo, 5' and 3' adaptors may be annealed to ensure that the XNAzyme is not blocked by the adaptors.
[0146] Some embodiments result in RNA molecules displayed on a substrate, where the RNA molecules may encode polypeptides. As discussed herein, the RNA molecules may be present in clusters, and sequence and position information may be obtained. The RNA clusters may form a library of encoded polypeptides. In some embodiments, the RNA molecules encode peptides or proteins of 1-25 kDa size. In other embodiments, libraries of peptides or proteins of 1-25 kDa size are displayed. For example, the libraries may be of scFvs, peptides, fibronectin type III domains (FN3 domains), or single domain antibodies (sdAbs, also known as nanobodies). Other displayable scaffolds include affibodies, darpins, finomers, OBodys, or avimers.
[0147] The method of displaying RNA molecules on a substrate to obtain a library can start with a substrate whose immobilized first nucleic acids are a plurality of first nucleic acids encoding a plurality of polypeptides, the first nucleic acids being present as a cluster, at least a portion of which may be sequenced.
[0148] When an immobilized single-stranded RNA molecule is obtained, a probe can optionally be annealed to the single-stranded RNA molecule. For example, a nucleic acid probe complementary to the 3' end of the second nucleic acid can be hybridized to the second nucleic acid. The hybridization site is preferably not within the open reading frame of the encoded polypeptide. The hybridization site can be located away from the stop codon of the open reading frame to avoid steric clashes between the probe and the ribosome. For example, the hybridization site can be at least 10, 15, 20, 25, 30, 35, or 40 nucleotides from the stop codon. In one particular embodiment, the hybridization site is at least 30 nucleotides from the stop codon. The probe can be, for example, fluorescently labeled so that RNA synthesis can be verified, visualized, and quantified.
[0149] In a further embodiment, the inventors use such polymerases to generate clusters of RNA molecules immobilized on a substrate, such as a flow cell, and then demonstrate the surprisingly efficient display of polypeptides translated from the RNA clusters.
[0150] Thus, the method may further comprise contacting the second nucleic acid, which is a newly formed RNA molecule, with ribosomes under conditions suitable for translation of the encoded polypeptide, which allows in vitro translation of the RNA sequence to form the polypeptide itself.
[0151] The displayed polypeptides can comprise or consist of standard amino acids. The displayed polypeptides can comprise non-standard amino acids. The displayed polypeptides can comprise unnatural amino acids. In one embodiment, the displayed polypeptides comprise any combination of standard amino acids, non-standard amino acids, and / or unnatural amino acids.
[0152] The second nucleic acid can include a ribosome binding site 5' to the open reading frame, e.g., the second nucleic acid can include a Shine-Dalgarno sequence.
[0153] The prior art has attempted to provide a method for displaying a large number of polypeptides on a surface in a manner suitable for high-throughput screening and analysis.However, these methods have drawbacks, particularly inefficient translation of polypeptides or instability of the displayed polypeptides.The present inventors herein provide a further technique for translation of RNA molecules displayed on a substrate, overcoming the drawbacks of the prior art.
[0154] The inventors have demonstrated that in vitro translation and folding of certain polypeptides may be inefficient. This is particularly relevant for large folded polypeptides such as scFvs. To improve translation and folding, the inventors have determined that trimethylamine N-oxide (TMAO) can be included in the in vitro translation buffer. In particular, the inventors have determined that a TMAO concentration of 0.05M to 1.5M improves yields when performing in vitro translation at 37°C. Furthermore, in vitro translation should be performed in a buffer with minimal or no RNAse activity.
[0155] Thus, in one embodiment, the method includes contacting the second nucleic acid with ribosomes under conditions suitable for translation of the encoded polypeptide, the conditions including trimethylamine N-oxide (TMAO). TMAO may be at a concentration of 0.05 to 1.5 M, or 0.05 to 1.2 M. The TMAO concentration may be 0.05 M to 1.5 M, 0.1 M to 1.2 M, 0.15 M to 1 M, 0.2 M to 0.8 M, 0.25 M to 0.6 M, 0.3 M to 0.5 M, or 0.35 M to 0.45 M. In one embodiment, the TMAO concentration is about 0.4 M.
[0156] Alternatively, the inventors have found that in vitro translation can be improved when dimethylsulfoxide (DMSO) is present in the translation buffer. For example, 10% DMSO can be included in the translation buffer. The inventors have found improvement when DMSO is included during translation of scFv, but have not observed improvement of the entire encoded protein.
[0157] Surprisingly, the efficient process for translation of immobilized RNA molecules is also applicable to other methods. Thus, in an embodiment of the invention, a method for displaying a polypeptide on a substrate is provided, comprising: i) providing a first nucleic acid comprising an antisense sequence encoding a polypeptide, wherein the first nucleic acid is immobilized on a substrate and oriented such that the 5' end is proximal and the 3' end is distal to the point of immobilization; ii) generating a second nucleic acid complementary to the first nucleic acid, the generating step of the second nucleic acid comprising: contacting the first nucleic acid with a nucleic acid polymerase under conditions suitable for RNA polymerization, wherein a primer for the polymerization is immobilized on the substrate such that a bridge is formed during polymerization; and The product of the polymerization is an RNA nucleotide strand immobilized on the substrate via the primer; iii) removing the first nucleic acid resulting in the presentation of the second nucleic acid on the substrate; and and iv) contacting the second nucleic acid with ribosomes under conditions suitable for translation of the encoded polypeptide, wherein the conditions of step iv) comprise TMAO. Optionally, TMAO is at a concentration of 0.05 M to 1.5 M or 0.05 M to 1.2 M. In one embodiment, the TMAO concentration is 0.4 M.
[0158] Further details of the above methods may be any as disclosed herein. Alternatively, TMAO may be replaced with DMSO, for example 10% DMSO.
[0159] The encoded polypeptide may be present as an open reading frame that ends with a stop codon. Once translation has stopped at the stop codon, the ribosome will then be stabilized. The ribosome may be stabilized by contacting the complex with a stabilizing buffer, for example, a buffer having a Mg concentration at least equal to or greater than 7 mM MgCl2.
[0160] Ribosome stabilization buffers containing more than 7 mM MgCl2 are not suitable for use in prior art methods that rely on DNA-RNAP-RNA complexes that cannot be denatured. However, the inventors have found that high Mg concentrations are associated with high display and stabilization efficiency and are suitable for use in the present method (see, e.g., FIG. 3). For example, the inventors have observed a 30-fold increase in ribosome display efficiency when comparing 7 mM MgCl2 with 50 mM MgAc in the present system. Thus, the stabilization buffer may contain 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, or 100 mM MgCl2 or MgAc. In some embodiments, the buffer has a magnesium concentration that is equal to or exceeds 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, or 100 mM MgCl2 or MgAc. The buffer may have a magnesium concentration greater than that provided by 7 mM MgCl2, hi some embodiments, the buffer has a magnesium concentration that is or is equivalent to 8-100 mM, 10-90 mM, 15-85 mM, 20-80 mM, 25-75 mM, 30-70 mM, 35-65 mM, 40-60 mM, or 45-55 mM MgCl2 or MgAc.
[0161] The ribosome stabilizing buffer may be phosphate buffered saline having the magnesium concentration mentioned above. The buffer may further comprise Tween 20 or Triton X-100.
[0162] In one particular embodiment, the ribosome presentation buffer may contain 50 mM TrisAc (tris(hydroxymethyl)aminomethane acetate), 150 mM NaCl, 0.1% Tween 20, 0.1% BSA, 20 U / ml RNase inhibitor, magnesium concentrations as disclosed herein, and may be at pH 7.5. The magnesium concentration may be provided by 50 mM MgAc (magnesium acetate).
[0163] In this manner, a polypeptide is generated that is displayed on the substrate surface. As discussed herein, a library of polypeptides, such as a library of scFv molecules, may be displayed on this surface.
[0164] The displayed polypeptide may be between 5-25 kDa, 10-25 kDa, 15-25 kDa, or 20-25 kDa. In some embodiments, the displayed polypeptide does not exceed 25 kDa. In particular embodiments, the polypeptide may be greater than 15 kDa.
[0165] The substrate surface with the displayed polypeptides may be washed and blocked. Suitable blocking reagents include bovine serum albumin, casein, recombinant bovine serum albumin, and the like.
[0166] Substrate surfaces that display polypeptides can be used for further studies. For example, if a surface displays a library of target-binding proteins or potential target-binding proteins, candidate targets, antigens, peptides, or proteins can be contacted with the surface to determine the binding characteristics of the displayed target-binding fragments. Candidates can be fluorescently labeled or can be detected in other ways. In this way, the displayed library can be used to analyze binding characteristics.
[0167] The present invention is also not limited to the measurement of binding properties, as the invention may be used to analyze any other properties, for example, a library encoding variants of an enzyme may be prepared and the library used to analyze the enzyme activity.
[0168] In one particular embodiment, there is provided a method for displaying a polypeptide on a substrate, comprising: i) providing a first nucleic acid comprising an antisense sequence encoding a polypeptide, e.g., an scFv, wherein the first nucleic acid is immobilized on a substrate and oriented such that the 5' end is proximal and the 3' end is distal to the point of immobilization; ii) generating a second nucleic acid complementary to the first nucleic acid, the generating step of the second nucleic acid comprising: contacting the first nucleic acid with a nucleic acid polymerase under conditions suitable for RNA polymerization, wherein the primer for polymerization is a DNA primer immobilized on the substrate such that a bridge is formed during polymerization; The product of the polymerization is an RNA nucleotide strand immobilized on the substrate via the primer; and The nucleic acid polymerase is a polymerase capable of acting on a DNA primer to synthesize an RNA molecule complementary to a single-stranded nucleic acid template, e.g., TGK; b) cleaving the first nucleic acid at a site 5' to the encoded polypeptide sequence to linearize the bridge; and c) contacting the linearized product of step b) with a polymerase under conditions suitable for RNA polymerization; iii) removing the first nucleic acid resulting in the presentation of the second nucleic acid on the substrate; and iv) contacting the second nucleic acid with ribosomes under conditions suitable for translation of the encoded polypeptide, said conditions in step iv) optionally including TMAO at a concentration of 0.05 to 1.5 M; and v) the ribosome-polypeptide complex is stabilized with a ribosome presentation buffer, optionally having a magnesium concentration of greater than 7 mM MgCl2. In one embodiment, the TMAO concentration is 0.05 M to 1.2 M. In one particular embodiment, the TMAO concentration is 0.4 M.
[0169] As discussed herein, a method for displaying a biomolecule on a substrate includes providing a first nucleic acid immobilized on a substrate. The first nucleic acid may be present as part of a cluster of clones, and at least some sequence and position information may be available. Methods for obtaining immobilized nucleic acids in this manner, and for obtaining the aforementioned information, are known in the art. However, the inventors herein provide a particularly improved method that is optimized for the downstream method disclosed herein. In particular, it is desirable to be able to generate longer immobilized nucleic acid sequences, for example 1.2 Kbp in length or more. The inventors provide an improved method for the production of this long structure.
[0170] Thus, in one embodiment, the first nucleic acid immobilized on the substrate provided in step i) comprises: 1) providing a template nucleic acid encoding a polypeptide sequence; 2) hybridizing the template nucleic acid to a primer immobilized on a substrate; 3) contacting the hybridized template nucleic acid with a polymerase under conditions suitable for extension of the immobilized primer to synthesize the first nucleic acid, which is a nucleotide strand complementary to the template; 4) performing bridge amplification of the first nucleic acid to generate clusters of the first nucleic acid; and 5) sequencing at least a portion of the first nucleic acid.
[0171] The template nucleic acid may have adapter oligonucleotides at its 5' and 3' ends. For example, if the substrate is an Illumina flow cell, the adapters may be P5 and P7 adapters. The primers immobilized on the substrate may be complementary to at least a portion of the template nucleic acid, such as the adapters.
[0172] The bound template nucleic acid is then contacted with a polymerase under conditions suitable for extension of the immobilized primer to synthesize the first nucleic acid, which is a strand of nucleotides complementary to the template. Thus, the first nucleic acid is an extension of the immobilized primer. The first nucleic acid and template nucleic acid can then be denatured to produce a single-stranded first nucleic acid immobilized on the substrate.
[0173] Bridge amplification can then be used to generate clonal clusters of the first nucleic acid. Bridge amplification can include cycles of annealing, amplifying, and denaturing steps. For example, the amplification step can include the following characteristics: 28-35 amplification cycles, an extension time of 1-120 seconds, an amplification buffer having a Mg concentration equivalent to 2-6 mM MgSO4, and a denaturing buffer containing 95-99.9% formamide, with or without the addition of 1-10 mM NaOH and 1-5 mM EDTA.
[0174] The inventors have discovered that the following characteristics can be used to specifically optimize this step for downstream RNA / polypeptide display: 32-35 amplification cycles, an extension time of 60-120 seconds, an amplification buffer with a Mg concentration equivalent to 2-6 mM MgSO4, and a denaturing buffer containing 95-99.9% formamide (with or without the addition of 1-10 mM NaOH and 1-5 mM EDTA).
[0175] Thus, in one embodiment, the first nucleic acid immobilized on the substrate provided in step i) comprises: 1) providing a template nucleic acid encoding a polypeptide sequence; 2) hybridizing the template nucleic acid to a primer immobilized on a substrate; 3) contacting the hybridized template nucleic acid with a polymerase under conditions suitable for extension of the immobilized primer to synthesize the first nucleic acid, which is a nucleotide strand complementary to the template; 4) performing bridge amplification of the first nucleic acid to generate clusters of the first nucleic acid, the bridge amplification comprising: 32-35 amplification cycles, having an extension time of 60-120 seconds per cycle, comprising the use of an amplification buffer having a Mg concentration equivalent to 2-6 mM MgSO4, and comprising the use of a denaturing buffer comprising 95-99.9% formamide, optionally 1-10 mM NaOH, and optionally 1-5 mM EDTA; and 5) sequencing at least a portion of the first nucleic acid.
[0176] In one particular embodiment, the bridge amplification comprises 32 cycles. The extension time may be 60 seconds. The amplification buffer may comprise a Mg concentration equivalent to 6 mM MgSO4. The denaturation buffer may comprise 98% formamide, 10 mM NaOH, and 1 mM EDTA.
[0177] The amplification buffer may be 2 M betaine, 20 mM Tris, 10 mM ammonium sulfate, 6 mM MgSO4, 0.1% Triton-X, 1.3% DMSO, 200 uM dNTPs, 80 U / ml Bst 2.0, pH 8.8.
[0178] The polymerase may be Bst large fragment, Bst 2.0 polymerase, or Bst 3.0 polymerase (New England Biolabs).
[0179] After cluster generation, the double-stranded bridges may be linearized and denatured according to techniques known in the art. At least a portion of the first nucleic acid may then be sequenced in a standard manner. For example, the first nucleic acid may include a primer binding site followed by a unique molecular identifier or barcode sequence, which may be sequenced. The barcode sequence may be a random barcode of 15-30 nucleotides.
[0180] After sequencing, the sequencing product can be removed. The 3' phosphate group of immobilized phosphate can be deprotected to make the method of the present invention more applicable. For example, when using Illumina flow cell and reagent, the 3' phosphate group of P5 primer can be deprotected. The enzyme T4 PNK can be used for deprotection.
[0181] As mentioned above, the inventors provide an optimized method for generating clusters of nucleic acid molecules immobilized on a substrate, which is particularly useful for certain downstream applications. Thus, in an aspect of the invention, a method for preparing clusters of nucleic acids bound to a substrate, comprising: 1) providing a template nucleic acid encoding a polypeptide sequence; 2) hybridizing the template nucleic acid to a primer immobilized on a substrate; 3) contacting the hybridized template nucleic acid with a polymerase under conditions suitable for extension of the immobilized primer to synthesize the first nucleic acid, which is a nucleotide strand complementary to the template; and 4) performing bridge amplification of the first nucleic acid to generate clusters of the first nucleic acid, the bridge amplification being performed with 32-35 amplification cycles, having an extension time of 60-120 seconds per cycle, comprising the use of an amplification buffer having a Mg concentration equivalent to 2-6 mM MgSO4, and comprising the use of a denaturing buffer comprising 95-99.9% formamide and optionally 1-10 mM NaOH, and optionally 1-5 mM EDTA.
[0182] In one particular embodiment, the bridge amplification comprises 32 cycles. The extension time may be 60 seconds. The amplification buffer may comprise a Mg concentration equivalent to 6 mM MgSO4. The denaturation buffer may comprise 98% formamide, 10 mM NaOH, and 1 mM EDTA.
[0183] The methods disclosed herein for preparing clusters of substrate-bound nucleic acids may be used to present nucleic acids of at least 0.5, 1, 1.2 or 1.5 Kbp in length. The methods may be used to present nucleic acids of 1-1.5 Kbp, 1.1-1.3 Kbp, or 1.2 Kbp in length.
[0184] In an aspect of the invention, a substrate is provided that displays a non-DNA nucleic acid molecule, such as an XNA, FANA, 2'OMe, or RNA molecule, obtained or obtainable by any of the methods disclosed herein.
[0185] In an aspect of the invention, there is provided a substrate presenting a polypeptide obtained or obtainable by any of the methods disclosed herein.
[0186] In another aspect of the invention, there is provided the use of a nucleic acid polymerase to extend a DNA primer immobilized on a substrate to synthesize a non-DNA nucleic acid molecule complementary to a single-stranded nucleic acid template.
[0187] Features disclosed in relation to the methods of the invention may also be applied to this aspect of the invention, for example any of the features relate to an RNA polymerase or an XNA polymerase.
[0188] For example, the nucleic acid polymerase may comprise an amino acid sequence having at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO: 1 and further comprising a mutation that allows polymerization of at least one XNA nucleotide or RNA nucleotide. The nucleic acid polymerase may comprise one or more or all of the following mutations: V93Q, D141A, E143A, and A485L.
[0189] In particular, the polymerase may have an amino acid sequence that has at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO:3, and in which residues 93, 141, 143, 409, 485, and 664 are unchanged (i.e., the mutations V93Q, D141A, E143A, Y409G, A485L, and E664K are maintained). The polymerase may have an amino acid sequence that has at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO:4, and where residues 93, 141, 143, 403, 485, 657, 658, 659, 663, 664, 669, 671, and 676 are unchanged (i.e., the mutations V93Q, D141A, E143A, L403P, A485L, P657T, E658Q, K659H, Y663H, E664K, D669A, K671N, and T676I are maintained). The polymerase may have an amino acid sequence that has at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO:20, and wherein residues 93, 141, 143, 409, 485, 521, 541, 545, 592, and 664 are unchanged (i.e., the mutations V93Q, D141A, E143A, Y409G, A485L, I521L, T541G, F545L, K592A, and E664K are maintained). In one particular embodiment, the nucleic acid polymerase is an amino acid sequence having at least 80%, 90%, 95%, 99% or 100% similarity or identity to the amino acid sequence of SEQ ID NO: 1, which may include the mutations V93Q, D141A, E143A, Y409G, A485L, I521L, F545L and E664K with respect to the amino acid sequence of SEQ ID NO: 1. In one particular embodiment, the nucleic acid polymerase is an amino acid sequence having at least 80%, 90%, 95%, 99% or 100% similarity or identity to the amino acid sequence of SEQ ID NO: 1, which may include the mutations V93Q, D141A, E143A, E429G, A485L, I521L and K726R with respect to the amino acid sequence of SEQ ID NO: 1.In one particular embodiment, the nucleic acid polymerase has an amino acid sequence having at least 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO:1, and optionally the amino acid sequence includes the mutations V93Q, D141A, E143A, A485L, V589A, E609K, I610M, K659Q, E664Q, Q665P, R668K, D669Q, K671H, K674R, T676R, A681S, L704P, and E730G relative to the amino acid sequence of SEQ ID NO:1. In one particular embodiment, the nucleic acid polymerase has an amino acid sequence having at least 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO:1, and optionally the amino acid sequence includes the mutations V93Q, D141A, E143A, A485L, I521L, V589A, E609K, I610M, K659Q, E664Q, Q665P, R668K, D669Q, K671H, K674R, T676R, A681S, L704P, and E730G relative to the amino acid sequence of SEQ ID NO:1. In one particular embodiment, the nucleic acid polymerase has an amino acid sequence having at least 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO:1, and optionally the amino acid sequence includes the mutations V93Q, D141A, E143A, A485L, E654Q, E658Q, K659Q, V661A, E664Q, Q665P, D669A, K671Q, T676K, and R709K relative to the amino acid sequence of SEQ ID NO:1. In one particular embodiment, the nucleic acid polymerase is an amino acid sequence having at least 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO: 1, and may include the mutations V93Q, D141A, E143A, E429G, D455P, A485L, K487G, I521L, R606V, R613V, and K726R with respect to the amino acid sequence of SEQ ID NO: 1. The polymerase may be Bst. The polymerase may be PGLVVWA.In one particular embodiment, the nucleic acid polymerase has an amino acid sequence having at least 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO:1, and optionally the amino acid sequence includes the mutations V93Q, D141A, E143A, N269W, E429G, D455P, A485L, K487G, I521L, V589A, R606V, R613V, and K726R with respect to the amino acid sequence of SEQ ID NO:1.
[0190] In a further aspect of the invention, a nucleic acid polymerase is provided that comprises mutations corresponding to N269W, E429G, D455P, K487G, I521L, V589A, R606V, R613V, and K726R (as described with reference to SEQ ID NO: 1) in the backbone of a polymerase from the polB family. In a particular embodiment, the backbone is any polB polymerase except viral polymerases. The backbone may be of a polymerase from the archaeal genera Thermococcus and / or Pyrococcus. The polymerase may be a variant of the polymerase from T. gorgonarius (Tgo) (SEQ ID NO: 1). The polymerase of this aspect of the invention may be associated with efficient polymerization of XNA molecules, such as phNA, PMO, or P-alkyl-moNA polymers. The polymerase of this aspect of the invention may be capable of synthesizing the polymer as a strand complementary to a nucleic acid template, such as a DNA template.
[0191] The polymerase may have an amino acid sequence having at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99% or 100% similarity or identity to the amino acid sequence of SEQ ID NO: 1, and the amino acid sequence may include the mutations N269W, E429G, D455P, K487G, I521L, V589A, R606V, R613V, and K726R with respect to the amino acid sequence of SEQ ID NO: 1. The amino acid sequence of the nucleic acid polymerase may further include one or more of the following mutations: V93Q, D141A, E143A, and A485L. These mutations are described elsewhere herein.
[0192] In one particular embodiment, the nucleic acid polymerase has an amino acid sequence having at least 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO:1, and optionally the amino acid sequence includes the mutations V93Q, D141A, E143A, N269W, E429G, D455P, A485L, K487G, I521L, V589A, R606V, R613V, and K726R with respect to the amino acid sequence of SEQ ID NO:1.
[0193] In an aspect of the invention, there is provided a method of screening a substrate displaying a plurality of biomolecules, the substrate being as disclosed herein or obtainable by a method disclosed herein, and the biomolecules forming a library. The library may be any of those disclosed herein. For example, the library may comprise a plurality of variants of a parent nucleic acid or polypeptide sequence.
[0194] The screens disclosed herein may involve measuring the affinity of the displayed biomolecule for a ligand or target molecule, or measuring enzymatic function. For example, the screen may involve measuring the affinity of the displayed variants of the parent scFv, or other binding polypeptide, for a target ligand. Alternatively, the screen may involve measuring the enzymatic function of the displayed variants of the parent molecule, such as activity against a substrate.
[0195] Sequence comparisons can be performed using readily available sequence comparison programs. These publicly available and commercially available computer programs can calculate the sequence identity between two or more sequences.
[0196] Those skilled in the art will understand how to calculate the percent identity between two nucleic acid sequences. To calculate the percent identity between two nucleic acid sequences, the alignment of the two sequences must first be prepared, and then the sequence identity value must be calculated. The percent identity of two sequences can take different values depending on i) the method used to align sequences, such as the Needleman-Wunsch algorithm (e.g., applied in Needle (EMBOSS) or Stretcher (EMBOSS)), the Smith-Waterman algorithm (e.g., applied in Water (EMBOSS)), or the LALIGN application (e.g., applied in Matcher (EMBOSS)); and (ii) the parameters used in the alignment method, such as local alignment vs. global alignment, the matrix used, and the parameters applied to gaps.
[0197] There are many different ways to calculate the percent identity between two sequences after alignment. For example, the number of identities may be divided by: (i) the length of the shortest sequence; (ii) the length of the alignment; (iii) the average length of the sequences; (iv) the number of non-gap positions; or (iv) the number of equivalent positions excluding overhangs. Furthermore, it will be recognized that the percent identity is also strongly length-dependent. Thus, the shorter a pair of sequences is, the higher the sequence identity can be expected to occur by chance.
[0198] The percent identity between two nucleic acid sequences can then be calculated such that the alignment is (N / T) x 100, where N is the number of positions where the sequences share identical residues and T is the total number of compared positions, including gaps but excluding overhangs.
[0199] The sequence alignment may be a pairwise sequence alignment. Suitable services include Needle (EMBOSS), Stretcher (EMBOSS), Water (EMBOSS), Matcher (EMBOSS), LALIGN, or GeneWise. As an example, the similarity or identity between two amino acid sequences can be calculated using the Needle (EMBOSS) service, set to default parameters, such as matrix (BLOSUM62), gap open (10), gap extend (0.5), end gap penalty (false), end gap open (10), end gap extend (0.5). In another example, the similarity or identity between two amino acid sequences can be calculated using the Matcher (EMBOSS) service, set to default parameters, such as matrix (BLOSUM62), gap open (14), gap extend (4), alternative match (1). As an example, identity between two nucleic acid sequences can be calculated using the Needle (EMBOSS) service with default parameters set to, for example, matrix (DNAfull), gap open (10), gap extend (0.5), end gap penalty (false), end gap open (10), end gap extend (0.5). In another example, identity between two nucleic acid sequences can be calculated using the Matcher (EMBOSS) service with default parameters set to, for example, matrix (DNAfull), gap open (16), gap extend (4), alternative match (1).
[0200] Every feature described in this specification (including the accompanying claims, abstract and drawings), and / or every step of any method or process so disclosed, may be combined with any of the above aspects in any combination, except combinations in which at least some of such features and / or steps are mutually exclusive.
[0201] For a better understanding of the present invention and to show how embodiments of the same may be practiced, reference is now made to the following examples which are not intended to be limiting of the invention in any way. EXAMPLES
[0202] (Example) Therapeutic antibodies have a transformative impact on the clinic, especially in inflammatory diseases and cancer, but their development remains time- and cost-intensive. Herein, we present a deep screening, 10-fold increase in antibody yields using the Illumina HiSeq platform. 8 We report an ultra-high throughput screening approach for parallel sequencing, display, and rapid affinity screening at levels exceeding 1000 of individual antibody-antigen interactions. Deep screening allows the discovery of tens to hundreds of low nanomolar to high picomolar distinct nanobody (VHH) and single-chain Fv (scFv) antibody variants, both from yeast-display enriched VHH libraries and directly from unselected synthetic scFv repertoires. The large antibody-antigen interaction datasets generated by combining deep screening with machine learning models allow the in silico prediction of novel high affinity scFv antibody sequences not present in the original repertoire. Deep screening promises to greatly accelerate the discovery of high affinity antibodies against a broad range of targets.
[0203] Massively parallel assays offer the potential to greatly increase both throughput and data generation rates in biomedical sciences. Repertoire selection approaches and directed biomolecular screening strategies have proven important in the discovery of enzyme catalysts and therapeutic antibodies, peptides and small molecule drugs, the inclusion of which could target many diseases.
[0204] Diversification methods at the level of high-throughput DNA oligonucleotide synthesis are highly developed, and various selection strategies (e.g., phage, yeast, and ribosome display) have been used to generate large-scale (10 10) combinatorial (poly)peptide repertoires can be processed and fractionated, but these still only sample a fraction of the possible sequence space. Moreover, all selection methods (to varying degrees) have inherent and unavoidable biases due to different levels of protein expression, display, folding efficiency, and potential toxicity to the host organism. Finally, these selections are typically performed "blind," with little or no global information on the output, until diversity has been sufficiently reduced to allow the occurrence, abundance, and enrichment of genotypes to be determined by next generation sequencing (NGS).
[0205] Furthermore, even if NGS of selection repertoires could provide information on genotypic distribution and enrichment, several studies suggest that both genotypic abundance and enrichment are only weakly associated with function (due to the biases mentioned above). Thus, genotypic distributions obtained from sequencing data are only an imperfect proxy for the global phenotypic and functional map of a particular biomolecular repertoire and therefore do not significantly improve the discovery of highly functional but low-abundance clones during selection experiments.
[0206] Due to these shortcomings and the need to obtain a more reliable picture of genotype-phenotype correlations, a number of high-throughput screening methods have been developed. However, most screening approaches are limited in their scope, scale, and information output. Scaling isolated screens (one clone / compound / drug per well) is not easy, even with robotics, while for biologics, the sequence composition of each well is particularly challenging and is often only performed for specific hits. DNA, peptide, and protein microarrays, where known sequences are printed or synthesized at defined locations on a surface, allow for linked measurements of sequence and function, but scales beyond 500k spots tend to be limited by prohibitive costs for many laboratories.
[0207] A potentially transformative approach aims to directly merge sequencing with functional screening. NGS technology is based on extremely high parallelization by sequencing clonal DNA from randomly sequenced DNA clusters on the polony and Illumina platforms. Both platforms are utilized for the characterization of DNA, RNA and polypeptides presented on a flow cell or captured in a polyacrylamide matrix after sequencing. This allows the detection of up to 2 × 10 6 It is now possible to simultaneously collate DNA and RNA:protein, as well as RNA:RNA and protein:protein interactions.
[0208] Herein, this concept is applied to a large number of antibody candidates, up to 2 × 10 9 We seek to extend the power of the Illumina HiSeq platform with its cluster / interaction diversity capabilities. We demonstrate highly diverse display and screening of both preselected and unselected synthetic nanobody (VHH) and single-chain Fv (scFv) antibody libraries to discover high affinity (low nM to mid pM) binders directly from global equilibrium antigen binding data. Our approach, which we term deep screening, expedites high affinity antibody discovery from months to days. We also demonstrate the utility of large deep screening datasets for machine learning of antibody-antigen interaction parameters and in silico prediction of high affinity antibody hit sequences directly from antibody repertoire deep screening data.
[0209] Example 1 - Implementation of ribosome display and deep screening on HiSeq 2500 Our goal was to validate an ultra-high-throughput antibody screening method on the Illumina HiSeq sequencing platform, an approach we call “deep screening.” To achieve our goal, we had to overcome several technical challenges, as described below.
[0210] Illumina next-generation sequencing can generate up to 2 billion (2 × 10 9 The system operates on a highly integrated instrument, a HiSeq 2500 equipped with a flow cell containing clonal DNA clusters of 1000 ssDNA. These are generated in situ from individual single-stranded (ss) DNA template molecules by a process called bridge amplification. Individual clusters typically contain an array of about 1,000 DNA molecules in a spot of about 1 μm in diameter. The clusters are aligned and then sequenced in parallel using Illumina sequencing by synthesis (SBS) technology, resulting in a large number of sequences and their physical xy coordinates as output. To perform protein interaction screening, it was necessary to develop a methodology to quantitatively convert DNA clusters first into RNA clusters and then into protein clusters. For this purpose, we exploited the engineered polymerase TGK, which has an efficient primer-dependent DNA-templated RNA polymerase activity, to convert the sequenced DNA clusters into RNA clusters. Specifically, we exploited a paired-end turnaround process to perform DNA bridge-templated RNA synthesis (Figure 4), whereby a surface-bound P5 primer is repeatedly extended for RNA synthesis on the DNA template. Once DTR is complete, the template is removed by restriction enzyme digestion and DNA-RNA duplex denaturation at alkaline pH, resulting in the formation of ssRNA clusters covalently attached to the flow cell surface by the P5 primer (Figure 4). These can be either directly interrogated or converted to peptide and protein clusters by in vitro translation (IVT).
[0211] We then developed a robust workflow to translate RNA clusters into polypeptides and stably display the resulting peptides or proteins on the flow cell surface. Because 5'-linked RNA clusters are susceptible to nuclease degradation, we used a reconstituted PURExpress IVT system rather than the more standard S30 IVT extract, which can contain significant amounts of endonucleases and exonucleases. Specifically, we used PURExpress ΔRF123, -T7 RNAP, which not only lacks T7 RNA polymerase but also all release factors (RF-1, RF-2, RF-3), and associates an RNA structure in which the desired open reading frame (ORF) is preceded by an N28 unique molecular identifier (UMI / barcode), a 5'-UTR containing a translation initiation signal, followed by a 3'-extension sequence (to space the ORF coding domain from the ribosome exit tunnel), and two stop codons to stall the ribosome (Figure 4). The stalled mRNA:ribosome:nascent polypeptide complexes are allowed to stabilize in high magnesium buffer at ambient temperature for several days, during which time flow cell arrays of up to hundreds of millions to billions of protein clusters with known sequence (or known unique molecular identifiers (UMIs)) can be interrogated for various functional assays, such as antigen binding.
[0212] Another technical challenge was raised because the HiSeq instrument is not designed for quantitative measurements, but rather its imaging system is designed to determine base calls during sequencing by thresholding the fluorescence intensity signals among the four color channels. This is a challenge to measure the quantification of binding interactions, which we solved algorithmically and experimentally by integrating the equilibrium binding signal intensity at different concentrations with the redundancy of each UMI. Furthermore, the HiSeq 2500 imaging platform used an epifluorescence line-scanning microscope with 532 nm and 660 nm lasers. The line-scanning process of imaging the flow cell requires the instrument to detect a significant amount of illuminated signal in one of the 660 nm channels (as would be expected during a sequencing run) to first locate on the flow cell surface and maintain focus during subsequent scans. This imaging mode is less suitable for screening binding interactions when clusters showing high signals are rare and sufficient signal is not obtained for focus. We solved this problem by labeling all RNA clusters by hybridization to their 3' ends with fluorescently labeled DNA oligos in the 660 nm channel, allowing focused imaging of the entire flow cell even when cluster signals are sporadic or absent in the 532 nm channel. Furthermore, this signal can be used as a diagnostic for RNA synthesis efficiency / cluster size and as a normalization factor for functional / protein binding signals from the same cluster. Finally, all steps (including sequencing, RNA and protein synthesis, and imaging) can be performed within the same instrument, streamlining the experimental, imaging, and data processing pipeline and avoiding image alignment issues. Indeed, the excellent xy reproducibility of the HiSeq optical stage allows efficient correlation of flow cell binding data with sequencing coordinates prior to fluorescence quantification of each cluster (Figure 4).Thus, using a custom image and data analysis pipeline, we can map barcode:phenotype pairs with barcode:genotype pairs and recover genotype:phenotype associations for millions of clones.
[0213] Finally, cluster size and protein expression levels are variables that, along with other possible artifacts, can introduce noise into genotype:phenotype binding data sets from deep screens. To correct for this inherent variability, a redundant measure of binding signals from multiple clusters of the same barcode is utilized along with statistical outlier rejection to obtain reliable data. In the implementation described herein, a two-lane HiSeq 2500 rapid run flow cell is used to measure up to 3 × 10 8 We use this with the proposed cluster of 12-fold redundancy, aiming for 12-fold redundancy. To achieve redundancy on the flow cell, we bottleneck the library after addition of UMIs to between 0.1 and 1 fmol. This results in a theoretical maximum diversity of 2.5 × 10 7 It becomes UMI.
[0214] From the above, the deep screening workflow proceeds in two stages. During the first stage, the N28 UMI barcodes are sequenced for cost and time reasons. Then, RNA synthesis is performed on the flow cell after sequencing, followed by in vitro translation (IVT) of the RNA clusters into protein clusters, which are matched for target binding in equilibrium binding and kinetic dissociation assays. Binding and kinetic data are generated in the form of raw images of the flow cell and processed by our data analysis pipeline that groups UMI and equilibrium binding data, allowing rapid validation of functionality within the library. If binding is observed, a second sequencing run is performed to sequence the library members (complete or various fragments thereof) and associate them with the N28 UMI barcodes, and thus the binding data. Depending on the number and length of the variable regions to be sequenced, the deep screening experiment can be completed in as little as three days, and data processing can usually be completed in a few hours.
[0215] Example 2 - Identification of rare, high affinity nanobodies for lysozyme The technical issues of RNA cluster generation, protein display and imaging of HiSeq flow cells after sequencing were overcome, and we first explored deep screening of nanobody libraries. Nanobodies (VHHs) are important tools in molecular and structural biology. Developed by Kruse lab, 10 8 We obtained a commercially available yeast-displayed VHH library with a reported diversity of 1000 and performed several (2-3) rounds of positive and negative magnetic-activated cell sorting (MACS) and fluorescence-activated cell sorting (FACS) on the library to select for binding to a model antigen (hen egg lysozyme (HEL)), after which deep screening for HEL binding was performed on the output on a flow cell (Figure 5A). In processing the barcoded binding data, we identified 1,479 (MACS) and 3,687 (FACS) barcodes with integrated cluster mean fluorescence intensity values that exceeded the background binding signal by at least 2-fold (for 300 nM HEL), indicating the presence of true HEL-binding VHH clones. The apparent HEL binding affinity (K D_app Equilibrium binding affinity titrations involving increasing concentrations of HEL (1, 10, 100 and 300 nM) were performed to determine the apparent dissociation constant k. Binding at 300 nM HEL was followed by dissociation rate measurements (the rate of dimming of the cluster signal corresponds to the apparent dissociation constant k). off_app provide.
[0216] Library sequencing was then performed to identify the three CDR sequences (nanobody genotypes) and their equilibrium binding signals and dissociation rates (K D_app , k off_app ) (nanobody phenotype), yielding 379,300 (MACS) and 39,900 (FACS) unique CDR combinations (Figure 5B). Binding data grouped by unique CDR combinations yielded a total of 47 (MACS) / 53 (FACS) putative VHH hits (Figure 5C).
[0217] The deep screening dataset allows a global analysis of the antibody discovery process. For both MACS and FACS selections of yeast-displayed VHHs, we observe a low correlation between CDR abundance and high equilibrium binding signal (as a surrogate for affinity) (Spearman rank correlation constant ρ = 0.361 (MACS), ρ = 0.442 (FACS) at 300 nM HEL) (Figure 5C). This suggests that high affinity clones that are inefficiently enriched in both MACS and (to a lesser extent) FACS selection setups are likely due to strong biases in VHH yeast display (including toxicity, expression, folding or display efficiency variability) or biased amplification / transcription at the DNA / RNA level. Thus, from both R3 selections by deep screening, rare binders could be isolated (estimated hit frequency 0.049% (MACS) or 0.834% (FACS) (see above).
[0218] However, this conclusion is supported by the high equilibrium binding signal (or equilibrium binding K D :K D_app ) is the “true” high affinity binding (K D ) would correlate with the binding kinetics of the CDRs. To evaluate this hypothesis, 20 clones (M1-M19 and M23) and 10 clones (F1-F10) from the R3 MACS / FACS (respectively) screens that showed a wide range of fluorescence intensity, equilibrium binding signal, and abundance were selected for characterization (Figure 5D). At the same time, 96 random colonies from the R3 MACS selection were plated and picked for colony PCR and Sanger sequencing. Four of the identified 28 unique CDR sequences had already been selected for characterization from the MACS / FACS library. Another 8 clones were selected from the remaining 24 (C1-C8) for a total of 38 clones, expressed, and characterized for measurement of binding kinetics by Biolayer Interferometry (BLI) (Figures 5D-F, Figures 10-13). This characterization identified the dissociation constant (KD ) is 9-20 nM (M5(1.9×10 -8 M), M6(1.42×10 -8 M), and M15 (9.81 x 10 -9 M)) three VHH hits with low K in the range of 20–100 nM D The K obtained from the BLI measurements were 9 clones with the K1 and K2 sequences, including two clones (C1 and C2) from randomly picked colonies (Figures 5D, 5E). D Values were plotted against the integrated mean intensity obtained from the deep screen, showing a Spearman rank correlation coefficient ρ = -0.697 at 300 nM HEL (Figure 5G).
[0219] While there are many possible factors that could explain the differences between deep screening and BLI, these results suggest that both nanobody abundance and enrichment can only weakly correlate with affinity, at least in some selection experiments. Despite the use of standard nanobody selection libraries and protocols, several more rounds would be necessary to further enrich for the highest affinity library clones. Our results suggest that deep screening could shorten this process, even in cases where enrichment is still low (2.9 × 10 6 This suggests that it is possible to find high affinity binders (M5, M6, M15) even at 3, 11, and 145 UMI in the screen. Indeed, identifying the same clones using standard procedures would require laborious and time-consuming microplate expression and screening of tens of thousands of colonies.
[0220] Example 3 - Affinity maturation of scFv antibodies directly without selection Having demonstrated the ability of deep screening to identify low nanomolar binders from preselected libraries, we next explored whether high affinity antibodies could be discovered without a selection step, i.e., directly from a diverse repertoire of low affinity parental clones.
[0221] Starting with the parent antibody IL70001, isolated by phage display from a human scFv library, we developed IC50001 against human interleukin-7 (huIL-7), a promising pharmaceutical target relevant to multiple autoimmune and allergic inflammatory diseases. 50 was measured to be approximately 7 μM (FIGS. 15-16). An affinity maturation library was prepared by diversifying both the Vk light chain CDRs L1 and L3, and this unselected library was then directly subjected to ab initio screening (FIG. 6A).
[0222] Deep screening and CDR L1 and L3 sequencing yielded 2.4 × 10 6 1.7 x 10 including unique barcode 8 measured, and 1.9 × 10 in protein space. 5 We obtained unique CDR combinations of 1000 huIL-7 clones (Figure 6B). Due to aggregation issues, we were only able to collect huIL-7 binding data up to 1 nM, but these revealed 173 promising unique hits (Figure 6C). Despite the high diversity of this input library, an overall convergence of CDR L3 loop sequences was observed, with significant diversity maintained in the central region of CDR L1, presumably reflecting a larger contribution of CDR L3 to the IL-7 paratope. Thus, we selected a subset (top 19 clones as determined by equilibrium binding signal at 1 nM huIL-7) and IL70001 for in-depth characterization (Figure 6D). These were recloned as Fab fragments to avoid possible pitfalls in affinity measurements by scFv multimerization, Fabs were expressed and purified in CHO cells, and binding kinetics were measured by BLI at 50 nM for each Fab, revealing that all 19 anti-IL-7 Fabs had K values ranging from 3 nM to 429 pM. D It was revealed that IL70001 had the K value (Fig. 6D, 3E, S6). D Since is significantly weaker than 50 nM, the measured maximum response and the ratio of the rates of association and dissociation are estimated by the dissociation constant (K D ) is insufficient for accurate fitting.
[0223] The role of IL-7 in autoimmune and allergic inflammatory diseases is dependent on its binding to the interleukin-7 receptor (IL7R). Therefore, we sought to assess whether our high affinity hits have the capacity to inhibit IL7 receptor (IL7R) signaling by sequestration of IL7 using a TF-1 STAT5 IL7Rα+γ luciferase cell-based reporter assay. Indeed, the inhibitory potential (IC 50 ) was observed to be increased by an average of 10,000-fold over IL70001 and 37,000-fold over IL70105 (Figure 6G, Figure 15, Figure 16). A strong correlation was also observed between affinity and inhibition (ρ = 0.956, R for fitting log(y) = m × log(x) + c). 2 = 0.901) (Figure 6H).
[0224] The data demonstrate that deep screening can rapidly identify multiple picomolar affinity antibodies against therapeutically relevant pharmaceutical targets directly from unselected VL1 / VL3 libraries. Furthermore, a strong correlation between flow cell signal and BLI-measured affinity was observed even when switching the antibody format from scFv to Fab (Figure 6D, Figure 6F, ρ = -0.788). Moreover, the affinity increase from the parent antibody resulted in a four-order of magnitude increase in the inhibitory potency of the target ligand. Avoiding selection in affinity maturation by deep screening results in a significant increase in the speed of discovery and provides a direct route from low affinity leads to high affinity (pM) affinity antibodies without the need for intermediate selection and screening processes. Furthermore, the isolated Fab clones showed desirable general properties and indicators of developability, such as good expression profiles (0.25-0.59 mg / ml culture) and excellent monomericity (12 / 19 clones showed ≥ 98% monomeric fraction) by HP-SEC (high performance size exclusion chromatography).
[0225] Example 4 - Affinity maturation of anti-Her2 scFv Having demonstrated the ability to rapidly screen and identify high affinity nanobodies and scFvs from both selected and unselected libraries, we sought to further explore whether a large, internally consistent deep screening dataset could be leveraged in a supervised machine learning approach to enable more efficient exploration of CDR sequence space and discovery of high affinity antibodies.
[0226] The target we chose was HER2 (ERBB2), a cell surface protein tyrosine kinase that is overexpressed in 30% of breast cancers, as well as ovarian, gastric, and lung cancers. Her2 is the target of the highly therapeutic antibody trastuzumab (Herceptin), with a reported binding affinity of approximately 1 nM. We used the Herceptin scFv and a well-characterized affinity panel of five scFvs (G98A, C6.5, ML3-9, H3B1, and B1D2+A1), which have reported binding affinities (K D ) ranged from 320 nM to 15 pM, which is the benchmark for our experiments (Figure 8A). First, we validated the display of Herceptin and anti-HER2 scFv affinity panels on a flow cell (Figure 8B) and determined the equilibrium binding K D These clones were ranked overall by peak intensity (except for B1D2+A1 and H3B1, which were found to have the opposite ranking) (Figure 8C). Interestingly, Herceptin scFv showed significantly higher peak intensity compared to the affinity panel clones (Figure 8C; equilibrium K D app 6.35 nM), which is presumably due to the well-known favorable expression, folding, and stability of Herceptin scFv.
[0227] Having demonstrated effective scFv display, the lowest affinity scFv G98A(K D= 320 nM, from the affinity panel with barely detectable binding above the Her2 background binding signal at 100 nM. Thereby, six G98A CDR H3 scanning libraries with 4 NNS codons per window were constructed (Figure 8A). Deep screening and subsequent CDR sequencing yielded 2.98 × 10 5 Unique barcodes were measured, which is (6.2 × 10 possible 6 of protein space. 5 100 nM Her2) (Figure 8B). Although this samples only a small fraction of the potential CDR H3 library diversity, a two-dimensional projection of VH3 sequence space by principal component analysis (PCA) revealed that function was highly localized to three matching peaks in close local proximity to each other, and that the vast majority of sampled mutations had no detectable binding at the highest concentration tested (100 nM Her2) (Figure 8C). Examination of the three highest scoring clones (peak intensity at 100 nM Her2 with wash steps) revealed no detectable binding at the known K D is 1.0 x 10 -9 Binding curves were obtained that closely matched ML 3-9 from the affinity panel with M (Figures 8D, 8E, 17). Based on the established correlation between flow cell data and affinity, these clones (HER20003, HER20004, and HER20005) likely have low nanomolar affinity for Her2, suggesting a 100-fold improvement in affinity through deep screening alone. These three clones were then converted to Fab, expressed in CHO cells, purified, and characterized for binding kinetics using Octet. Fitting a 1:1 model to the BLI data, the kinetic binding affinities were measured to be 2.8 nM, 3.4 nM, and 1.8 nM for HER20003, HER20004, and HER20005, respectively, supporting the deep screening observations (Figures 8D, 8F, 18). G98A and ML3-9 were expressed and purified, and the 1:1 model fit of G98A at 20 nM gave a K D(46.9 nM), which is 6.8-fold higher than that in the published literature. The maximum response signal in BLI was rather low, suggesting that specific binding was not well captured at this Fab concentration. Therefore, the K D Value (3.2×10 -7 I will refer to M).
[0228] Thus, deep screening was able to reproduce the phage display affinity maturation of anti-Her2 G98A scFv in a single 3-day experiment. However, the motivation for this experiment was not affinity maturation, but rather the matching of CDR H3 sequences (genotype) to binding affinities (phenotype) (here, 2.4 × 10 5 The objective of this study was to generate a large dataset ("HER2affmat") that combines multiple histograms of HER2-specific markers with multiple histograms of HER2-specific markers (including CDR H3 sequences) and in silico prediction of higher affinity Her2 binders.
[0229] For this purpose, we built a machine learning model to predict Her2 binding by formulating a classification problem and partitioning the predictors into three categories (no hits, low hits, high hits) according to the fluorescence intensity threshold from a 5 min wash step (Figure 9A). We selected the threshold such that the parent clone G98A was approximately centered in the low hit category with G98A intensity of 190.05, and chose a no hit threshold of 150.0 and a high hit threshold of 250 to balance the selection of sufficiently low nanomolar binders. These thresholds yielded 232,693 no hit VH sequences, 1,284 low hit VH sequences, and 111 high hit VH sequences from the "HER2affmat" data. Our best model yielded acceptable performance metrics of F1 scores of 0.993, 0.329, and 0.480 for the no hit, low hit, and high hit categories, respectively (Tables S2, S3). Although the F1 scores for low and high hits were lower than ideal, they were dominated by high false positive rates, likely due to the difficulty of defining class boundaries across the continuous space of measurements.
[0230] With the model trained, we explored whether it could be used to generate better and more diverse anti-Her2 binding sequences than those observed in the "HER2affmat" dataset, compared to random mutagenesis. To this end, we took the three top-scoring clones (seeds) from the "HER2affmat" dataset (HER20003, HER20004, and HER20005) and performed in silico mutagenesis of 1.98 × 10 for each seed. 6 We generated 10 mutant VH3 sequences (Figure 9A). Specifically, we generated all single, double, and triple mutants, as well as up to 10 mutants. 8 We randomly generated fourth and fifth order variants of all 594 million mutations, which were then scored by this model before a selection was made for the next round of deep screening.
[0231] To compare this model against random mutagenesis, we devised a selection scheme to compile a random mutation set from all single mutants and up to 1000 mutants with edit distances between 2 and 5 for each seed sequence, resulting in a pool of 13,121 mutations ("random / mut"). We then assembled a pool of mutation-only sequences generated by machine learning by removing all sequences with a high hit score <0.9, randomly selecting up to 1000 mutants with edit distances between 2 and 5, and rejecting those already selected in the "random / mut" set. This resulted in the assembly of a pool of 11,916 mutations ("ml / mut") (Figure 9A).
[0232] This resulted in a total of 25,042 CDR VH3 sequences synthesized as oligonucleotide pools, including clones G98A, ML3-9, HER20003, HER20004, and HER20005. Subsequent deep screening (under identical conditions to the "HER2affmat" library) yielded a total of 199,737 unique VH3 sequences in protein space, including 24,968 (99.72% coverage) of the 25,037 clones from the designed library, plus 174,700 extra variants due to errors during array synthesis and cloning. The ML-generated VH CDR3 sequences ("ml / mut") showed a significant enhancement in fluorescence intensity with a significant upward shift in the distribution of high-intensity clones under 5 min washing conditions compared to random mutagenesis ("random / mut") (Figure 9B), indicating that our machine learning model was able to extract salient features of high affinity Her2 binding from the "HER2affmat" dataset and use them to accurately predict a large number of novel Her2 binders.
[0233] Since our goal was to leverage machine learning to discover antibodies with higher affinity than the parent G98A, HER20003, HER20004, and HER20005 clones, the resulting deep screening data was explored as a binary classification problem, with G98A centered in the no-hit category, and clones with 1.5-fold intensity of G98A classified as hits (Figure 9B). The resulting classification threshold showed an overall hit performance of 13.23% for ML clones ("ml / mut") versus 2.31% for randomly selected clones ("random / mut") (Figure 9B, Table S4). Examination of the number of hits per edit distance from the seed showed that the machine learning model improved sequence space sampling by 2.6- to 23.3-fold over random mutagenesis, with an average improvement of 5.7-fold (Table S4).
[0234] For affinity determination, 21 new anti-Her2 scFv clones (6 from the "HER2affmat" library, 9 from the ML set ("ml / mut"), and 6 from the random set ("random / mut")) were selected and converted to monovalent Fabs, expressed in CHO cells, purified, and characterized (Figure 17). These clones were selected based on a variety of criteria (peak fluorescence intensity at all concentrations, shape of equilibrium binding curve, fluorescence intensity at washing conditions) to identify patterns that better correlate with affinity, expression rate, and monomericity.
[0235] All selected clones obtained from screening of the “HER2affmat” library containing three seeds (HER20003, HER20004, and HER20005) had a total of 8.58 × 10 -10 M and 5.25×10 -9 K between M D The clones in the ML / random library showed a 300-fold improvement in affinity (by intensity values and binding curves) with the top clone from the ML set (HER20013) showing a 5,220-fold improvement in affinity (K D =6.07×10 -11 M), and another four clones from the ML set (HER20015, 20, 21, 22) showed a >1,000-fold improvement in affinity to G98A (Figures 9C, 9D, Figure 18).
[0236] Although high-potency clones were 5-fold rarer in the random set, we still identified two clones (HER20024 and 25) that showed a >1,000-fold increase in affinity for G98A (1.65 × 10 -10 M and 2.83 x 10 -10M). In addition to the affinity increase, an overall improvement in monomericity of ML and random clones was observed, ranging from 93.5% to 98.1% relative to G98A monomericity. However, no strong correlation between deep screening data and monomericity could be identified. Collectively, these results demonstrate the extraordinary power of combining deep screening with state-of-the-art machine learning models to discover high affinity antibody binders.
[0237] Example 5 - Presentation of anti-Her2 scFv affinity panel In this example, 3.2×10 -7 ~1.5×10 -11 A library of anti-Her2 scFvs with known affinity ranges of M were clustered. 28 nucleotides were then sequenced, revealing the known unique barcode of each clone and the spatial location of each cluster on the flow cell surface. RNA synthesis was then performed and verified as described in the method before performing in vitro translation and ribosome display. Ribosomes were stabilized by adding a buffer containing 50 mM MgAc and then blocked with 0.1% BSA. After a short incubation of 100 nM AF532 streptavidin and a buffer wash, a control image of 0 nM was taken. Equilibrium binding affinity titrations were then performed in a stepwise fashion with 0.03 nM to 100 nM Her2-biotin and 100 nM AF532 streptavidin, and images of the flow cell at each concentration were saved before measuring the kinetic off-rate as described in the method.
[0238] Processing the raw flow cell images through our data analysis pipeline allowed us to report the mean, median, and SEM (standard error of the mean) for each clone at each concentration. Flow cell images from the entire experiment are shown in Figure 7B, and the median intensities for each clone are reported in Figure 7C for equilibrium binding affinity and off-rate measurements. Curves were fitted to these data as described in the method, and the fitting results are shown in Table 4, along with the SPR-validated binding affinities for Her2.
[0239] (Table 4) [Table 3]
[0240] Although the fitting curves in this example do not completely match the results previously characterized, likely due to incomplete saturation and differences in the methods employed, there are some interesting observations to note in the data. In particular, the binding curve for Herceptin shows a slower equilibrium binding rate than H3B1 and B1D2+A1, but with a significantly higher Rmax. Since Herceptin is known in the literature to be a very well-behaved scFv, in that it is well expressed and well folded, the significantly higher Rmax is likely due to a combination of binding affinity and expression / folding. In any case, there is a clear ranking of clones that is observable in the data, and low and high nanomolar binding affinities can be identified. Thus, this data demonstrates the ability to display and measure the equilibrium binding and off-rates of single chain antibodies on an Illumina flow cell.
[0241] Example 6 - Presentation of XNA-2'-Fluoroarabinonucleic Acid (FANA) Following sequencing, the flow cell was imaged and offsets were measured to allow for correction of chromatic distortion between the different optical paths of the instrument. The sequencing products were then denatured in a formamide wash at 65°C, followed by Illumina's "End Deblock" protocol using the reagents "Cleavage Reagent Mix (CRM)" and "Cleavage Wash Mix (CWM)" to remove any dye-terminal nucleotides still present on the flow cell surface. Next, with the single-stranded DNA template present on the flow cell, the 3' phosphate group must be "deprotected" or removed from the P5 primer. This is done using the "Fast Resynthesis Mix (FRM)" or T4 Polynucleotide Kinase (T4 PNK) and Illumina's deprotection protocol.
[0242] With the free 3' hydroxyl group on the P5 grafted primer, a cyclic RNA primer extension is performed using D4YK polymerase, reusing the paired-end turnaround process. Here, D4YK takes the DNA primer (grafted P5) annealed to the DNA template and extends it with FANA ribonucleotides (faNTPs). This is done by heating the flow cell to 55°C and performing 12 cycles of injection of denaturation mix, annealing, and extension (incubation time for each extension step is 900 seconds) using 1x Thermopol buffer (20 mM Tris-HCl, 10 mM (NH4)2SO4, 10 mM KCl, 2 mM MgSO4, 0.1% Triton X-100, 200 uM faNTPs, 10 nM D4YK pol, pH 8.8).
[0243] Following 12 cycles of FANA extension, oligos are annealed onto the 8-oxoG site of the grafted P7 primer and cleavage is performed (using Illumina's FLM2 reagent or 200 U / ml Fpg, 100 ul / ml BSA, and 1x NE buffer 1).
[0244] Following DNA cleavage and final extension, the DNA:FANA duplex is denatured and the DNA template is washed away using a mixture of 100 mM NaOH and 5 mM EDTA, and the flow cell is then washed with 2 ml of 6 M GuHCl, 10 mM Tris, pH 7.4, and 2 ml of 5X SSC, 0.1% Tween 20. With the clusters of single-stranded FANA present on the flow cell, 100 nM R2_atto647N and P7'_surface_hyb are annealed to the P7 adaptor at the 3' end of each molecule of FANA.
[0245] Example 7: Presentation of XNA-2'-O-methyl ribonucleic acid (2'OMe) Following sequencing, the flow cell was imaged and offsets were measured to allow for correction of chromatic distortion between the different optical paths of the instrument. The sequencing products were then denatured in a formamide wash at 65°C, followed by Illumina's "End Deblock" protocol using the reagents "Cleavage Reagent Mix (CRM)" and "Cleavage Wash Mix (CWM)" to remove any dye-terminal nucleotides still present on the flow cell surface. Next, with the single-stranded DNA template present on the flow cell, the 3' phosphate group must be "deprotected" or removed from the P5 primer. This is done using the "Fast Resynthesis Mix (FRM)" or T4 Polynucleotide Kinase (T4 PNK) and Illumina's deprotection protocol.
[0246] With the free 3' hydroxyl group on the P5 grafted primer, a cyclic RNA primer extension with 2M polymerase is performed, reusing the paired-end turnaround process. Here, 2M takes the DNA primer (grafted P5) annealed to the DNA template and extends it with 2'O-methyl ribonucleotides (2'OMe NTPs). This is done by heating the flow cell to 55°C and performing 12 cycles of injection of denaturation mix, annealing, and extension (incubation time for each extension step is 3600 seconds) with TAM (TGK amplification mix; 200 uM 2'OMe NTPs, 10 nM 2M pol, 2 M betaine, 20 mM Tris, 10 mM ammonium sulfate, 6 mM MgSO4, 0.1% Triton-X, 1.3% DMSO, pH 8.8).
[0247] Following 12 cycles of 2'OMe extension, an oligo is annealed onto the 8-oxoG site of the grafted P7 primer and cleavage is performed (using Illumina's FLM2 reagent or 200 U / ml Fpg, 100 ul / ml BSA, and 1x NE buffer 1).
[0248] Following DNA cleavage and final extension, the DNA:2'OMe duplex is denatured and the DNA template is washed away using a mixture of 100 mM NaOH and 5 mM EDTA, and the flow cell is then washed with 2 ml of 6 M GuHCl, 10 mM Tris, pH 7.4, and 2 ml of 5X SSC, 0.1% Tween 20. With the clusters of single-stranded 2'OMe present on the flow cell, 100 nM R2_atto647N and P7'_surface_hyb are annealed to the P7 adaptor at the 3' end of each molecule of 2'OMe.
[0249] Example 8 - Discussion It has long been recognized that a broader and more diverse antibody (and, in general, biomolecule) repertoire will contain a higher probability of high affinity binders, thus covering the geometric space of possible epitopes in a more complete manner. The experiments disclosed herein build on pioneering work from many groups that repurposed the Illumina sequencing platform for high-throughput screening, by demonstrating efficient display of single domain and single chain antibodies, and also extending screening depth to 3×10 on a two-lane rapid run flow cell. 8 (And 2 × 10 on an 8-lane flow cell 9 This was achieved by extending the range of available binding sites to 100,000,000 unique high affinity binders directly from unselected repertoires against two different human therapeutic targets (IL-7 and Her2), with affinity improvements of 2–3 orders of magnitude over the usual range.
[0250] This screening depth can reveal salient features of antigen-binding paratopes. For example, in the case of the IL-7 affinity maturation library, high affinity binders showed a high degree of convergence in CDR L3 sequences, while CDR L1 remained more diverse despite the emergence of a consensus sequence signature for the highest affinity clones. Similarly, in the case of the Her2 affinity maturation library, high affinity binders showed some convergence around three core motifs.
[0251] An important observation from our experiments is that the "true" binding affinities of individual purified monovalent antibody Fab fragments (after conversion from scFv) determined by state-of-the-art biophysical measurements (biolayer interferometry, BLI) correlate well with the ranking and relative affinity estimated by equilibrium antigen binding on the flow cell (ρ=-0.788 for IL-7 clones), despite confounding factors such as possible avidity effects, differences in clustering and presentation efficiency, diffusion-related flow effects, and heterogeneous specificity of the flow cell. Antibody rankings can be further improved and differentiated by utilizing antigen dissociation kinetics, which are not generally utilized here. In the future, the combination of equilibrium binding and off-rate measurements will enable the collection of apparent overall affinity measurements across the entire displayed antibody library, which in turn will provide a large and internally consistent data set for machine learning-assisted sampling of CDR sequence space associated with high affinity antigen binding with desirable binding kinetics.
[0252] We illustrated the utility of machine learning to predict anti-Her2 binders using an experimental affinity maturation dataset generated by deep screening by training a state-of-the-art machine learning model (Figure 9A). We evaluated the ML model in a head-to-head comparison with random mutagenesis to computationally generate a set of 11,916 ML-guided anti-Her2 mutants and a set of 13,121 random mutants, which were arrayed as oligopools and experimentally measured for binding by deep screening. Here, we observed that the number of high-strength binders obtained from the ML set increased by 5-fold overall compared to random mutagenesis (Figure 9B), and the ML model produced a greater improvement over random as the number of mutations increased (23-fold for 5 mutations, Table S4). We characterized 15 high-strength clones (9 from the ML set and 6 from the random set), yielding 2.26 × 10 -9 and 6.07 x 10 -11 The binding affinity between M was observed, with the ML clone HER200013 showing a 5,200-fold improved binding affinity relative to the parent clone G98A (Figures 9C and 9D).
[0253] While we are not the first to train machine learning models and predict function on protein sequences, the majority of the literature attempts to predict refinements to function that have been embedded in the vast evolutionary history of natural protein classes. Engineering binding to specific antigens is substantially more challenging because the information is not present in large sequencing datasets, and all ML models rely on antigen-specific sequences to function datasets that can be easily generated for deep screening.
[0254] Interestingly, antibodies isolated by deep screening not only exhibit high affinity antigen binding properties, but also other desirable "developability" features important for therapeutics, such as retention of affinity upon conversion to Fab or IgG, high degree of monomericity, and high expression yields in CHO cells. We hypothesize that these features arise in part through pre-selection for desirable physicochemical properties, expressed using a minimal (chaperone-deficient) translational machinery and allowed to express and fold at 37°C for 1 h. Misfolded or aggregation-prone scFvs will be deselected due to low equilibrium binding signal strength.
[0255] The human immune system has approximately 10 9 The immune repertoire contains many B cells, each of which displays a different antibody, and should be ready to respond to any antigenic challenge. In rodents, the immune repertoire is much smaller (10 7), whose antibodies can respond to virtually any non-self antigen. If the naive repertoire could be faithfully represented by deep screening, one repertoire could in principle yield binders to any desired target. However, such initial binders in the immune system are usually of moderate affinity (low micromolar to high nanomolar) with slow on-rate, fast off-rate kinetics, which is a problem to capture by current implementations of deep screening. This is because imaging of flow cells (two-lane flow cells take 4 min, eight-lane flow cells take 16 min) is often slower than the half-life of low-micromolar binders (dissociation half-life of 1 μM affinity binders is about 4-30 s). Nevertheless, further improvements in detection sensitivity (experimental and hardware) will enable screening of naive libraries in the future.
[0256] Deep screening is currently implemented on the HiSeq 2500 platform, but there are no obvious obstacles to extending it to the more advanced HiSeq 4000 and NovaSeq platforms, where the principles of clustering and imaging are similar, but where patterned flow cells are used instead of random clustering. Also, currently both sequencing and flow cell binding and imaging are performed on the same instrument, but external imaging is also possible, as demonstrated on the MiSeq platform, with advantages such as a wide range of color channels and fluorescent imaging modes, potentially allowing measurement of protein expression, non-specific binding and competitive binding in the same assay.
[0257] In conclusion, deep screening expands the post-sequencing screening capabilities of the HiSeq platform into the realm of hundreds of millions to billions of measurements. With methodological advances, this allows for the presentation and direct screening of selected VHH and unselected scFv antibody libraries and the discovery of picomolar affinity binders from such libraries in days rather than weeks or months. Furthermore, the large genotype-phenotype correlation datasets generated by deep screening allow for efficient machine learning and the sampling of antibody CDR sequences and antigen binding space resulting in novel high affinity antibody sequences not present in the starting library. Many applications of deep screening platforms are anticipated, especially in the discovery and development of therapeutic antibody drugs.
[0258] (method) (Structure design) To transcribe and translate the sequenced DNA clusters on an Illumina flow cell, our DNA construct contains the following elements: a P5 adaptor, followed by a 28nt unique barcode, a 27nt unstructured spacer (5p UNS v2), a ribosome binding site, a start codon, a protein coding region, a TolAK short linker, 2x stop codons, a 27nt unstructured spacer (3p UNS v2), and a P7 adaptor.
[0259] (Table 1) [Table 4]
[0260] (Preparation of anti-Her2 scFv clones) Anti-Her2 scFv clones including Her2_G98A, Her2_C6.5, Her2_ML3-9, Her2_H3B1, Her2_B1D2+A1 and Herceptin are disclosed in US8580263B2 and US5772997A. These clones were found to be capable of producing up to 3.2×10 -7 ~1.5×10 -11The constructs were designed to cover the affinity range of M and contain the above structural elements. The designed constructs were ordered as gBlocks from IDT (Integrated DNA Technologies), cloned into E. coli, single colonies were picked, verified by Sanger sequencing, and extracted by PCR.
[0261] Linear double-stranded DNA was diluted to 10 nM and the concentration was quantified by qPCR using the KAPA Quant kit (KK4824, Roche).
[0262] (Table 2) [Table 5] TIFF2024534985000011.tif248170TIFF2024534985000012.tif124170
[0263] (Cluster generation and barcode sequencing) Libraries containing 5% of each of the above clones were clustered on an Illumina HiSeq 2500 using paired-end rapid run flow cells (PE-402-4002, HiSeq PE Rapid Cluster Kit v2, Illumina) at 6 pM, which typically yields about 200 m reads. Although these flow cells can be fully clustered to yield more than 400 m reads in the downstream RNA synthesis and ribosome presentation steps, we chose to hybridize fluorescent Atto 647N oligos to the P7 adaptors of each cluster to allow normalization of the binding assay. At densities higher than 200 m reads, our HiSeq 2500 cannot reliably focus and image a flow cell with all the labeled RNA clusters. Using fiducials would allow higher cluster densities, but would not provide information on RNA synthesis efficiency.
[0264] The standard Illumina DNA cluster generation protocol was modified by increasing the number of bridge amplification cycles from 28 to 32 and adding a 60 second wait time after each amplification cycle. Further modifications included an amplification mix containing 2 M betaine, 20 mM Tris, 10 mM ammonium sulfate, 6 mM MgSO4, 0.1% Triton-X, 1.3% DMSO, 200 uM dNTPs, 80 U / ml Bst 2.0, pH 8.8, and a denaturation mix containing 98% formamide, 10 mM NaOH, and 1 mM EDTA. The combination of these modifications was shown to greatly improve the signal of clusters grown from long templates such as single-chain antibodies, which can be up to 1.2 kb in length.
[0265] Clustering and sequencing were performed in paired-end, single-read runs with 28 cycles on read 1 and 0 cycles on read 2 without indexing and were accomplished using HiSeq Control Software (HCS v. 2.2.68, Illumina). Flow cells and clustering reagents were obtained from HiSeq PE Rapid Cluster Kit v2 (PE-402-4002, Illumina) and sequencing reagents were obtained from HiSeq Rapid SBS Kit v2 (FC-402-4023, Illumina).
[0266] (RNA synthesis) Imaging sequencing of the flow cell allows offsets to be measured and chromatic distortions between the different optical paths of the instrument to be normalized. Exit HCS and start HiSeq engineering software (Archimedes Test Software v. 3.8.317.0, Illumina), initialize the instrument, home the stage, set the chemistry module run mode to "RapidRun", and set the flow cell temperature to 20°C. Next, 120 ul of Illumina's Universal Sequencing Buffer (USB) is pumped into the flow cell, followed by auto-tilting, aligning, and imaging the flow cell using the "Bruno Scan" module, with the surface set to "Dual Lane", the scan speed set to 2.0 mm / sec, and the swath set to "Dual Swath". The flow cell image is saved, allowing offsets and chromatic distortions between the different optical paths of the instrument to be measured.
[0267] The sequencing products are then denatured with a formamide wash (e.g., FDR-Illumina's "Fast Denaturant") at 65°C, followed by Illumina's "End Deblock" protocol, the reagents "Cleavage Reagent Mix (CRM)" and "Cleavage Wash Mix (CWM)" to remove any remaining dye-terminal nucleotides on the flow cell surface. The single-stranded DNA template present on the flow cell then requires "deprotection," or removal of the 3' phosphate group from the P5 primer. This is done using the "Fast Resynthesis Mix (FRM)" or T4 Polynucleotide Kinase (T4 PNK) and Illumina's deprotection protocol.
[0268] With the free 3' hydroxyl group on the P5-grafted primer, a cyclic RNA primer extension is performed using TGK polymerase, reusing the paired-end turnaround process. Here, TGK takes the DNA primer (grafted P5) annealed to the DNA template (cluster strand) and extends it with ribonucleotides (NTPs). This is done by heating the flow cell to 55°C and performing 12 cycles of injection of denaturation mix (FDR), annealing, and extension (incubation time for each extension step is 1800 seconds) using TAM (TGK amplification mix; 625 uM NTPs, 10 nM TGK, 18 U / ml Superase In (AM2696, Thermo), 2 M betaine, 20 mM Tris, 10 mM ammonium sulfate, 6 mM MgSO4, 0.1% Triton-X, 1.3% DMSO, pH 8.8).
[0269] Following 12 cycles of RNA extension, we observed that for long templates (>800 nt or >900 nt), TGK was unable to complete the synthesis of the strand. This was likely due to torque build-up within the DNA:RNA duplex, which is covalently attached to the surface via its respective 5' end. To release this torque, we annealed the oligos spanning the grafted P7 primer Ueno 8-oxoG site and performed two cycles of cleavage (using Illumina's "Fast Linearisation Mix 2" (FLM2) reagent or 200 U / ml Fpg, 100 ul / ml BSA and 1x NE buffer 1) and extension (with TAM) for 30 min at 37°C and 1 h at 55°C, respectively.
[0270] Following DNA cleavage and final extension, the DNA:RNA duplex is denatured and the DNA template is washed away using a mixture of 100 mM NaOH and 5 mM EDTA (or Illumina's FDR mix). The flow cell is then cleaned with 2 ml of 6 M GuHCl, 10 mM Tris, pH 7.4, and 2 ml of 5X SSC, 0.1% Tween 20. With the single-stranded RNA clusters present on the flow cell, 100 nM R2_atto647N and P7'_surface_hyb are annealed to the P7 adaptor at the 3' end of each molecule of RNA.
[0271] (Table 3) [Table 6]
[0272] (Ribosome display on Illumina flow cell) Ribosome display was performed using a custom PURExpress kit from New England Biolabs (NEB) (lacking release factors 1, 2, and 3, and lacking T7 RNA polymerase). Specifically, 200 ul of master mix was prepared containing 80 ul of solution A, 60 ul of solution B, 4 ul of disulfide enhancer 1 and 2 (E6820S, NEB) (if required), 4 ul of Superase In (AM2696, Thermo), 10 ul of 10 mM Tris, pH 7.0, 4 M trimethylamine N-oxide, and 10 ul of Millipore water (if required). 90 ul of master mix was then injected into each lane of the flow cell using a custom designed low dead volume manifold, taking care to avoid introducing bubbles, and the flow cell was then incubated on the HiSeq at 37°C for 60 minutes. After the incubation period, the flow cell is cooled to 20°C, followed by washing and stabilization of the ribosomes with 1 ml of ribosome presentation buffer per lane (50 mM TrisAc (tris(hydroxymethyl)aminomethane acetate), 150 mM NaCl, 50 mM MgAc (magnesium acetate), 0.1% Tween 20, 1 U / ml Superase In (AM2696, ThermoFisher), pH 7.5).
[0273] Next, with the ribosomes stabilized by the presentation buffer, the flow cell was blocked with binding buffer (ribosome presentation buffer and 0.1% bovine serum albumin (BSA) (A9647, Sigma-Aldrich)) by 6× 250 ul injections with 10 min between each injection. Next, 100 nM AF532 streptavidin (S11224, ThermoFisher Scientific) in binding buffer was injected and incubated for 30 min at 20° C., then the flow cell was washed with 250 ul of binding buffer per lane and imaged. These images serve as a baseline for background fluorescence, non-specific binding, and fluorescence remaining from sequencing.
[0274] (CDR sequencing using internal sequencing primers) After a successful deep screening presentation experiment, a subsequent sequencing experiment was set up using the same library to elucidate the CDR sequences with internal sequencing primers. The CDR sequencing experiment was performed in HCS with a custom recipe that first sequenced the N28 UMI with Illumina's lead 1 sequencing primer for 28 cycles, followed by FDR, where the sequencing product was denatured at 65°C, annealed with the appropriate internal sequencing primer, and sequenced for enough cycles to cover the variable region. All internal sequencing primers used in this experiment were obtained from IDT, purified by HPLC, and resuspended in IDTE at 100 μM.
[0275] (Oligos used for internal sequencing of CDRs) [Table 7]
[0276] (Her2 binding affinity panel to anti-Her2 scFv) Equilibrium binding assays are performed by preparing a dilution series of Her2-biotin (HE2-H822R, Acro biosystems) from 0.03 nM to 100 nM in binding buffer, and 100 nM AF532 streptavidin in binding buffer. The steps of the binding assay are: 1) inject 100 ul of Her2-biotin per lane, 2) incubate for 40 minutes, 3) wash with 100 ul of binding buffer per lane, 4) inject 100 ul of 100 nM AF532 streptavidin per lane, 5) incubate for 10 minutes, 6) wash with 150 ul of binding buffer, and 7) image the flow cell. This process is performed stepwise from the lowest concentration to the highest concentration.
[0277] Following equilibrium binding experiments, pseudokinetic off-rates can be measured by injecting binding buffer into the flow cell at a constant rate and imaging at fixed time points, in this case the flow cell was imaged at 5, 10, 20, 60, 120, and 240 minute intervals.
[0278] (Image Processing) A single scan of a two-lane rapid run flow cell generates 8 x 2048 x 160,000 pixel 16-bit tif images in four color channels, resulting in a total of 32 images. The HiSeq 2500 uses 532 nm and 660 nm lasers with emission filters set (558-32 nm, 610-60 nm, 687-20 nm, and 740-60 nm) that pass out to a 4x time-delay integration (TDI) line-scan CCD detector. These dyes with signal from the Alexa / Atto 647 "A" and "C" channels, and the Alexa 532 "G" and "T" channels, were detectable, with the greatest signal-to-noise ratios observed for the "C" and "T" channels. Therefore, only analyses using the "C" and "T" color channels were performed.
[0279] Our image processing pipeline works by dividing each image of 2048 x 160,000 pixels into 16 tiles that are processed independently and in parallel. For a given tile image, we first perform a non-uniform illumination correction by applying a disk-shaped structure element with a radius of 25 pixels to the morphological aperture, and then remove the morphological aperture from the tile image. We then detect the centroids of clusters present in the tile image using a peak local maximum function that works first by performing a morphological dilation of the tile image with a square kernel of 3 x 3 pixels. The algorithm then goes through each pixel of the tile image and checks whether the pixel image of the tile image is equal to the value of the same dilated pixel and whether its pixel intensity exceeds a set threshold. If a given pixel meets these conditions, it is considered as a centroid and added to the centroid map. In this case, we use a pixel intensity threshold of 400 or 600 (this value was manually adjusted for our device). This cluster detection method is simple and computationally fast, and is generally sufficient for our needs.
[0280] We use the detected cluster coordinates on the "C" and "T" images to align them against known sequencing coordinates using the DFT (Discrete Fourier Transform) phase correlation function from the OpenCV package. Because of some variability in the repeatability of the microscope stage and optical distortions within the HiSeq, we perform fine alignment by subdividing the tile image into non-overlapping sub-images of 128 × 128 pixels and store the fine offsets in an offset map.
[0281] We use the refined offset map to quantify the intensity of all known clusters from the sequencing data by extracting 9 × 9 pixel sub-images with their centroids at the offset-corrected cluster coordinates. We then perform an element-wise multiplication of the 9 × 9 pixel sub-images with a 9 × 9 pixel array constructed from a 2D Gaussian point spread function (PSF) with sigma 0.5. The following equation is used to describe our 2D Gaussian PSF:
number
[0282] The sum of the multiplied pixel values per element is defined as the cluster intensity. The image processing pipeline reports the cluster intensities of all sequenced clusters on the "C" and "T" channels from all scans of the flow cell and stores this on disk or inserts it into a database.
[0283] (Data Analysis) The barcode sequences and combined cluster intensities are matched and grouped by unique barcode sequences through our custom data processing pipeline. This then performs outlier rejection using median absolute deviation and a cutoff of 2.0. Global statistics are reported for each unique barcode; mean, median, standard deviation, standard error of the mean, minimum and maximum intensity, etc. This is done for both the "C" and "T" channels, allowing for normalization of protein binding signals to the RNA probe signal.
[0284] More specifically, data analysis begins by first grouping all cluster data by their common N28 UMI. If there are at least 12 replicates that are not rejected because the cluster is outside the imaged region, the UMI is retained. Next, UMIs and combined data that includes UMI and CDR sequencing data, with at least 3 CDR reads per UMI, are grouped. Following grouping, median absolute deviation outlier rejection is performed to correct for CDR read consensus errors (and eliminate UMIs where no consensus exists) before calculating the mean, median, standard deviation, and standard error of the mean for each UMI in both the T (532 nm; protein) and C (660 nm; RNA) color channels.
[0285] (Equilibrium binding and dissociation rate fitting) With binding data grouped by unique barcode and outliers removed, median binding data were plotted for each member of the anti-Her2 scFv library and equilibrium binding curves were fitted using the following equation and a least squares fit with a trust region reflection algorithm as the solver.
number
[0286] To fit the kinetic off-rates, the following equation was fitted to the data using least squares and a trust region reflectance solver:
number
[0287] The equilibrium binding and dissociation rate fitting may alternatively be written as follows:
[0288] Flow cell-based equilibrium binding curves are fitted to the integrated mean intensity of a given UMI using the least squares method, as implemented in the curve_fit() function of the python Package SciPy using the following equation:
number
[0289] Flow cell-based kinetic dissociation curves are fitted using the least squares method as implemented in the curve_fit() function of the python Package SciPy using the following two-phase dissociation equation:
number
[0290] (Nanobody Yeast Surface Display Selection) Nanobody yeast display libraries contain >2.5 × 10 9 Frozen stocks of cells (EF0014-FP, Kerafast) were obtained from the Kruse laboratory. An aliquot of the library was first thawed at 30°C and then harvested in 1 L of "Yglc4.5 -Trp" (3.8 g / L -Trp Yeast Dropout Medium Supplement (Y1876, Merck), 6.7 g / L Yeast Nitrogen Base (Y0626, Merck), 10 mL / L Pen-Strep (P4333, Merck)) by shaking at 230 RPM overnight at 30°C. The harvested culture was then expanded into 3 L of medium and grown to stationary phase (OD ) over 48 hours. 600 Cultures were centrifuged at 3,500 × g for 5 min and resuspended in fresh Yglc4.5 -Trp supplemented with 10% DMSO to a final density of 10 per mL. 10 After cells were allowed to settle, 2 mL aliquots were made and frozen at -80°C.
[0291] To prepare the naïve library for the first selection round, one aliquot was thawed at 30°C and used to inoculate 1 L of Yglc4.5-Trp supplemented with 2% galactose. This culture was then grown for 72 h at 24°C. Expression was confirmed by flow cytometry using a FITC-labeled anti-HA antibody (GG8-1F3.3.1, Miltenyi Biotech) before the first selection round. Cells showing more than 10-fold library diversity were first deselected against streptavidin microbeads (Miltenyi Biotech) for 1 h at 4°C in PBS-T-BSA (0.1% Tween-20, 0.1% BSA) and then detached from the beads on a Miltenyi MACS magnet. Deselected cells were then incubated in the presence of 500 nM HEL-biotin (GTX82960-pro, GeneTex) for 1 h at 4°C. Streptavidin beads were added and incubated for a further 15 min before selection and washing on a Miltenyi MACS magnet. Beads and bound cells were eluted, pelleted, and resuspended in 1 L of Yglc4.5-Trp supplemented with 2% galactose before growing for 72 h at 24° C. The second round was performed similarly to the first, except that no deselected cells were present and HEL-biotin was reduced to 300 nM, followed by addition of streptavidin microbeads, panning on a MACS column, washing, and cell recovery.
[0292] After the second round, the recovered cells were split into half the volume and each aliquot was used for a third round of MACS (magnetic activated cell sorting) and FACS (fluorescence activated cell sorting). The third round of MACS was performed as in the second round, with the HEL-biotin further reduced to 200 nM, followed by harvesting, centrifugation and cell acquisition by miniprep of plasmid DNA (D2004, Zymo Research). Prior to cell acquisition, 100 μL of cells were serially diluted and plated on YPD agar plates to allow picking of 96 colonies for colony PCR and Sanger sequencing. For the third round of FACS, cells were incubated with 200 nM HEL-biotin for 1 h at 4°C, pelleted, resuspended in fresh PBS-T-BSA, combined with 100 μg Neutravidin PE (A2660, ThermoFisher Scientific) and a 1:1000 dilution of anti-HA-FITC antibody for 15 min, then sorted on a Synergy 3 cell sorter (Sony Biotechnology) and gated on double-labeled (FITC / PE) events, yielding 50,135 cells. Sorted cells were harvested and miniprepped for the third round of MACS.
[0293] (Nanobody library preparation and deep screening) Third-round minipreps for MACS and FACS were subjected to 20 cycles of PCR amplification (Q5 polymerase; M0492, NEB) with primers that anneal to the N-terminal framework region, the C-terminal HA tag, and introduce a 20-nucleotide overhang at the 5' end of each primer with homology to the 5' flow cell adaptor (RBS+ATG; KF_olap.fwd) and the 3' flow cell adaptor (TolAk linker; KF_olap.rev).
[0294] [Table 8]
[0295] The current nanobody library with homology to the adapters was run on a 1% agarose gel and a band of approximately 449 bp (approximate as the library contains CDR loops of variable size) was gel extracted, purified and quantified with nanodrop. The library was then assembled into a deep screening presentation construct by Gibson assembly with 0.2 pmol of 5' adapter, nanobody library fragments and 3' adapter, and HiFi DNA Assembly Master Mix (E2621, NEB) and incubated at 50°C for 30 min. The library was then bottlenecked by removing 300 amol of material from the Gibson assembly reaction (estimated efficiency 100%) and PCR amplified for 25 cycles using Q5 polymerase and outnest P5 and P7 primers.
[0296] [Table 9]
[0297] The PCR products were run on a 1% agarose gel and an approximately 800 bp band was gel extracted, purified and quantified first by nanodrop and then by qPCR (NEBNext library quant kit, E7630, NEB).
[0298] The quantified library was ready for deep screening and was diluted to 2 nM, denatured (10 μL of library was mixed with 10 μL of 100 mM NaOH and incubated at room temperature for 5 min) and rapidly diluted to 20 pM with HT1 buffer from the rapid PE flow cell clustering kit (PE-402-4002, Illumina). The library was then diluted to a concentration of 6 pM and then loaded into the template slot of the HiSeq 2500 and the deep screening experiment was set up as described above.
[0299] Following acquisition of a baseline flow cell image, equilibrium binding assays were performed with successively increasing concentrations of HEL-biotin. Each condition specifically involved the injection of 120 μL of HEL-biotin (GTX82960-pro, GeneTex) precomplexed with AF532 streptavidin (S11224, ThermoFisher) in a 1:1 ratio in presentation buffer at 20°C, incubation for 45 min at 20°C, washing with 200 μL of presentation buffer, followed by complete imaging of the flow cell. This was performed for 1 nM, 10 nM, 100 nM, and 300 nM HEL and a 1:1 amount of AF532 streptavidin. The highest concentration of HEL was followed to measure the kinetic dissociation rate. This was accomplished by pumping presentation buffer over the flow cell and imaging at 5, 10, 15, 20, 30, 60, and 120 minutes. Raw images were then processed as described above.
[0300] (Nanobody expression and periplasmic extraction) Nanobody hits were computationally constructed assuming no mutations outside the sequenced CDR regions (including 3 nucleotides before and after the actual variable region). The constructed hits were then codon optimized and ordered as gBlocks from IDT, and subsequently cloned by FX cloning into the E. coli periplasmic expression vector pSBinit (Addgene plasmid #110100; http: / / n2t.net / addgene:110100; RRID:Addgene_110100), a gift from Markus Seeger. Single colonies were picked and verified to be correct by Sanger sequencing. After verification, single colonies were grown overnight at 37°C in 5 mL TB + 25 μg / mL chloramphenicol and then subcultured 1:100 into 5 mL TB (with chloramphenicol). Cultures were grown at 37°C to an OD of approximately 0.6–0.9 in 0.05% w / v L-arabinose. 600Cultures were grown for a further 3.5 h and then harvested by centrifugation at 2,500xg for 20 min at 4°C and discarding the supernatant. The pellet was resuspended (1 / 20 of the original culture volume) in 250 μL TES buffer (50 mM Tris-HCl, pH 7.2, 0.1 mM EDTA, 20% sucrose) and incubated on ice for 60 min for periplasmic extraction. Supernatants were harvested by centrifugation at 20,000xg for 30 min at 4°C and protein yields were quantified by SDS page. All clones were normalized to a concentration of 500 nM in SuperBlock PBS (37515, ThermoFisher Scientific) before BLI kinetic measurements.
[0301] (Nanobody kinetics measurement) Periplasmic extracted nanobodies standardized to 500 nM in SuperBlock PBS were further diluted to 50 nM. Biolayer Interferometry (BLI) kinetics were performed with Octet Red384 (Sartorius) with a reference subtraction performed for each nanobody clone using a non-loaded streptavidin chip (18-5136, Sartorius). Kinetics were measured using the following steps: 1) sensor check for 30 seconds, 2) loading with HEL-biotin 25 μg / mL for 400 seconds, 3) baseline measurement for 240 seconds, 4) binding kinetics at 50 nM of each nanobody for 400 seconds or 500 seconds, 5) dissociation kinetics for 600 seconds. SuperBlock PBS was used as buffer for all steps.
[0302] (Biolayer Interferometry Data Fitting) BLI kinetic data were collected on an Octet Red384 instrument as described above and in the Kinetics Measurement section below. In all cases, a streptavidin chip (18-5136, Sartorius) was loaded with biotinylated target antigen and bound with a fixed concentration of each VHH or Fab clone after washing to baseline signal. After obtaining on-rate kinetics, the chip was bathed in fresh buffer to measure off-rate kinetics. Data measured for each clone was referenced to a streptavidin-only chip to remove non-specific binding to streptavidin.
[0303] A 1:1 model was fitted to all data by least squares using a custom python script.
[0304] The binding rate is given by the following equation:
number
[0305] The dissociation rate is given by the following equation:
number
[0306] K D The value is calculated as follows:
number
[0307] (IL-7 library preparation and deep screening) An unselected IL-7 Vkappa light chain CDR L1 and L3 scFv library was prepared and provided to us by AstraZeneca in the pCANTAB6 plasmid. The scFv library was extracted by 20 cycle PCR using Q5 polymerase and primers that provide 25 nucleotides of homology to the 5' and 3' presented adapters. The PCR products were run on a 1% agarose gel and the band of approximately 778 bp was gel extracted and purified. As with the nanobody library assembly, 0.2 pmol of the 5' adapter, scFv library fragment, and 3' adapter, as well as HiFi DNA Assembly Master Mix (E2621, NEB) were combined and incubated at 50°C for 30 minutes. The library was then bottlenecked by taking 500 amol of material from the Gibson assembly reaction (estimated efficiency 100%) and PCR amplifying for 25 cycles using Q5 polymerase and outnest P5 and P7 primers. The PCR product was run on a 1% agarose gel and the 1.2 kb band was gel extracted, purified and quantified first by nanodrop and then by qPCR (NEBNext library quant kit, E7630, NEB).
[0308] The quantified library was ready for deep screening and was diluted to 2 nM, denatured (10 μL of library was mixed with 10 μL of 100 mM NaOH and incubated at room temperature for 5 min) and rapidly diluted to 20 pM with HT1 buffer from the rapid PE flow cell clustering kit (PE-402-4002, Illumina). The library was then diluted to a concentration of 6 pM and then loaded into the template slot of the HiSeq 2500 and the deep screening experiment was set up as described above.
[0309] Following acquisition of baseline flow cell images, equilibrium binding assays were performed with successively increasing concentrations of hu-IL7-biotin precomplexed at a 1:1 ratio (100 pM, 333 pM, and 1 nM) with AF532 streptavidin (S11224, ThermoFisher). In this experiment, substantial aggregation of hu-IL-7 at the flow cell surface was observed, which resulted in the inability to image above a concentration of 1 nM hu-IL-7, and therefore no kinetic dissociation measurements could be obtained. Images were processed as described above to reveal CDR sequences, which were used to identify putative hits.
[0310] (Expression and purification of anti-IL-7 and anti-Her2 Fabs) The top 19 putative anti-IL7 hits (and IL70001) and all 26 anti-Her2 hits (including G98A, and ML3-9) were converted from scFv to Fab format and the heavy and light chain variable regions were synthesized separately and cloned into mammalian expression vectors pEU10.1 and pEU4.4, respectively. Vectors were transiently transfected into CHO (Chinese Hamster Ovary) cells using PEI and proprietary media. Expressed Fabs were purified by loading clarified culture supernatants onto a CaptureSelect™ CH1-XL column (Life Technologies, ThermoFisher, The Netherlands), running in DPBS, eluting with 25 mM acetic acid pH 3.6, and buffer exchanging into DPBS pH 7.4 using a PD-10 desalting column (Cytiva). Concentrations were determined spectrophotometrically using extinction coefficients based on the amino acid sequence. Protein purity was verified by SDS-PAGE and verification of correct molecular weight (MW) was achieved by LC-MS analysis. HP-SEC analysis was performed after purification by loading 70 μl of each protein onto a TSKgel G3000SWXL; 5 μm, 7.8 mm x 300 mm column at a flow rate of 1 ml / min using 0.1 M anhydrous disodium phosphate (anhydrous) + 0.1 M sodium sulfate, pH 6.8 as running buffer. For comparison purposes, a gel filtration standard (BIORAD, Cat. No. 151-1901) was also used.
[0311] (IL-7 kinetics measurement) Binding kinetics of the top 19 hits and IL70001 were measured using an Octet BLI and streptavidin coated chip (18-5136, Sartorius). In all cases, the buffer used was DPBS (14190-169, Gibco) + 0.1% BSA + 0.02% Tween-20. Purified Fabs were diluted to a final concentration of 50 nM. Kinetics were measured using the following steps: 1) sensor check for 60 seconds, 2) loading with hu-IL7-biotin 5 μg / mL for 30 seconds, 3) baseline measurement for 60 seconds, 4) binding kinetics at 50 nM of each Fab for 300 seconds, 5) dissociation kinetics for 600 seconds.
[0312] (TF-1 STAT5 IL7Rα+γ cell-based reporter assay) 10 of 1 ml 7 Two vials containing 100 µg / ml TF-1 STAT5 IL7α+γ luciferase cG3 cells were removed from liquid nitrogen, thawed, and transferred to 1 x 50 ml Falcon tubes (2 vials per tube) containing 40 ml complete medium and centrifuged at 1,200 rpm for 5 min. The supernatant was aspirated and the cell pellet resuspended in 40 ml RPMI (11875093 ThermoFisher) + 10% FBS + 1% sodium pyruvate, followed by a further centrifugation at 1,200 rpm for 5 min, then the supernatant was aspirated as before. Finally, the cells were resuspended in 40 ml RPMI + 10% FBS + 1% sodium pyruvate, placed in a T175 flask, and incubated at 37°C for 24 h in a 5% CO2 atmosphere.
[0313] Hu-IL7 (CHO expressed) was prepared up to 0.12 nM in RPMI + 10% FCS + sodium pyruvate, which was then diluted 1:100 to a final volume of 20 mL and added to a 384-well plate. Purified Fab was added undiluted to the 384-well plate and 11-point duplicate 3-fold serial dilutions were performed in complete RPMI using a Bravo liquid handling platform. After 24 hours of incubation, cells were removed, pelleted by centrifugation at 1,200 rpm for 5 minutes, and resuspended in 10 mL of RPMI + 10% FCS + 1% sodium pyruvate. Cells were counted and diluted with complete RPMI to a concentration of 10,000 cells / 20 μL. Cells (20 μL) were then added to 3× 384-well clear assay plates. 10 μL of titrated Fab was added to the cells, followed by 10 μL of 120 pM Hu-IL7. Plates were then kept in a tissue culture incubator for 6 hours at 37°C under 5% CO2 atmosphere. 100 mL of Steady-Glo Reagent (E2520, Promega) was thawed before use and 40 μL was added to each well of the 384-well plate. Plates were sealed and incubated for 10 minutes on a plate shaker before measurement. Luminescence readings were measured using an EnVision plate reader with a 1 second pulse time. Each Fab was measured in duplicate.
[0314] Data was exported and processed using a custom Python script, and the average data was fitted using least squares to a log inhibitor response curve defined as:
number
[0315] (Deep screening of anti-Her2 affinity panel) The anti-Her2 scFv affinity panel and Herceptin protein sequences were reverse translated, codon optimized, and constructed into deep screening display constructs with known 28 nucleotide UMIs. DNA constructs were ordered as gBlocks from IDT and clustered onto rapid PE flow cells at 1% per construct, with the remaining clusters on the flow cell containing PhiX controls (FC-110-3001, Illumina). The flow cells were sequenced for 28 cycles and deep screening display was performed as described above.
[0316] The nucleic acid sequences of the anti-Her2 scFv clones are shown in Table 2.
[0317] Following successful presentation, equilibrium binding assays were performed using biotinylated human Her2 (HE2-H822R-25ug, Acro biosystems) and AF532 streptavidin (S11224, ThermoFisher). In this case, a binding assay cycle was performed with injection of 120 μL of Her2-biotin, incubation for 45 minutes at 20° C., washing with 200 μL of presentation buffer, injection of 120 μL of 100 nM AF532-streptavidin, incubation for 10 minutes at 20° C., followed by washing with 200 μL of presentation buffer, and imaging. Equilibrium binding assays were performed with 100 pM, 333 pM, 1 nM, 3.33 nM, 10 nM, 33.3 nM, and 100 nM Her2-biotin before initiating the kinetic dissociation assay. Dissociation assays were performed by pumping wash buffer over the flow cell and imaging at 5, 10, 20, 60, 240, and 420 min. Data collected from this experiment was processed as previously described and summary statistics were calculated by grouping by known UMIs.
[0318] (Anti-Her2 scFv affinity maturation library preparation and deep screening) A CDR VH3 affinity maturation library was constructed using G98A as the parental starting clone. This was accomplished by TOPO cloning (450245, ThermoFisher) the G98A gBlock from the previous section into Top 10 chemically competent cells (C404010, ThermoFisher), picking 6 colonies, growing them overnight in 5 mL TB + 50 μg / mL kanamycin, and miniprepping 2 mL of culture. Plasmids were sent for Sanger sequencing using M13 forward and reverse primers. One of the correct colonies was carried forward.
[0319] To construct the VH3 affinity maturation library, it was first necessary to extract the upstream and downstream regions of VH3. This was done by PCR amplification of plasmid DNA using Q5 polymerase with primer set 1 (G98A_olap.fwd and G98A_5p_VH3.rev) and primer set 2 (G98A_3p_VH3.fwd and G98A_olap.rev) for 25 cycles in two reactions. Both PCR products were then treated with DpnI (R0176L, NEB) for 1 h at 37°C and then purified with a PCR clean-up kit (T1030S, NEB). This process yielded the upstream and downstream fragments of the G98A clones with homology to the deep screening display construct and excluded the contamination of wild-type plasmid DNA.
[0320] [Table 10]
[0321] The Her2 affinity maturation library was then assembled using 20 cycles of PCR with Q5 polymerase, the upstream and downstream fragments of G98A, equimolar amounts of VH3 NNS oligos generating a scanning window of 4 NNS codons across the CDR VH3, and the G98A olap forward and reverse primers. The products were then column purified using a PCR cleanup kit (T1030S, NEB). Deep screening 5' and 3' adapters were then added using Gibson assembly with 0.2 pmol of each fragment and NEB HiFi Assembly Master Mix (E2621, NEB) at 50°C for 60 minutes. The library was then bottlenecked by removing 300 amol of material from the Gibson assembly reaction (estimated efficiency 100%) and PCR amplified for 25 cycles using Q5 polymerase and outnest P5 and P7 primers. The PCR product was run on a 1% agarose gel and the 1.2 kb band was gel extracted, purified and quantified first by nanodrop and then by qPCR (NEBNext library quant kit, E7630, NEB).
[0322] The quantified library was ready for deep screening and was diluted to 2 nM, denatured (10 μL of library was mixed with 10 μL of 100 mM NaOH and incubated at room temperature for 5 min) and rapidly diluted to 20 pM with HT1 buffer from the rapid PE flow cell clustering kit (PE-402-4002, Illumina). The library was then diluted to a concentration of 6 pM and then loaded into the template slot of the HiSeq 2500 and the deep screening experiment was set up as described above.
[0323] Following acquisition of a baseline flow cell image, equilibrium binding assays were performed with successively increasing concentrations of human Her2-biotin (HE2-H822R-25ug, Acro biosystems) precomplexed with AF532 streptavidin (S11224, ThermoFisher) in a 1:1 ratio (100 pM, 333 pM, 1 nM, 3.33 nM, 10 nM, 33.3 nM, and 100 nM). In this case, a binding assay cycle was performed by injecting 120 μL of Her2-biotin:AF532 streptavidin precomplex, incubating for 45 minutes at 20°C, washing with 200 μL of presentation buffer, and then imaging the flow cell. Kinetic dissociation assays were performed following the 100 nM maximum condition by pumping the presentation buffer over the flow cell and imaging at 5, 10, 20, 60, 120, and 240 min. Images were then processed and CDR sequences elucidated by internal primer sequencing as described above and used to assemble a CDR:binding dataset called "HER2affmat".
[0324] (ML vs. Random library preparation and deep screening) For each seed sequence, we devised a selection scheme to compile a random mutation set from all single mutants and up to 1000 mutants with edit distances between 2 and 5, resulting in a pool of 13,121 mutations ("random / mut"). We then assembled a pool of mutation-only sequences generated by machine learning by removing all sequences with a high hit score <0.9, randomly selecting up to 1000 mutants with edit distances between 2 and 5, and rejecting those already selected in the "random / mut" set. This resulted in the assembly of a pool of 11,916 mutations ("ml / mut"). The sequences were assembled into an oligo pool of 25,042 CDR VH3 sequences and ordered from Twist Bioscience. The "Her2 ML vs. Random" library was assembled for deep screening in the same way as the "HER2affmat" library, and PCR was performed for 20 cycles using Q5 polymerase, the upstream and downstream fragments of G98A, combined with the oligo pool and the G98A olap forward and reverse primers. The products were then column purified using a PCR cleanup kit (T1030S, NEB). Deep screening 5' and 3' adapters were then added using Gibson assembly with 0.2 pmol of each fragment and NEB HiFi Assembly Master Mix (E2621, NEB) at 50°C for 60 minutes. The library was then bottlenecked by removing 300 amol of material from the Gibson assembly reaction (estimated efficiency 100%) and PCR amplified for 25 cycles using Q5 polymerase and outnest P5 and P7 primers. The PCR product was run on a 1% agarose gel and the 1.2 kb band was gel extracted, purified and quantified first by nanodrop and then by qPCR (NEBNext library quant kit, E7630, NEB).
[0325] The quantified library was ready for deep screening and was diluted to 2 nM, denatured (10 μL of library was mixed with 10 μL of 100 mM NaOH and incubated at room temperature for 5 min) and rapidly diluted to 20 pM with HT1 buffer from the rapid PE flow cell clustering kit (PE-402-4002, Illumina). The library was then diluted to a concentration of 6 pM and then loaded into the template slot of the HiSeq 2500 and the deep screening experiment was set up as described above.
[0326] Following acquisition of a baseline flow cell image, equilibrium binding assays were performed with successively increasing concentrations of human Her2-biotin (HE2-H822R-25ug, Acro biosystems) precomplexed with AF532 streptavidin (S11224, ThermoFisher) in a 1:1 ratio (100 pM, 333 pM, 1 nM, 3.33 nM, 10 nM, 33.3 nM, and 100 nM). In this case, a binding assay cycle was performed by injecting 120 μL of Her2-biotin:AF532 streptavidin precomplex, incubating for 45 minutes at 20°C, washing with 200 μL of presentation buffer, and then imaging the flow cell. Kinetic dissociation assays were performed following the 100 nM maximum condition by pumping the display buffer over the flow cell and imaging at 5, 10, 20, 60, 120, and 240 min. Images were then processed and CDR sequences elucidated by internal primer sequencing as described above and used to assemble a CDR:binding dataset called "Her2 ML vs. Random."
[0327] (Anti-Her2 hit kinetics measurement) Binding kinetics of all anti-Her2 Fabs were measured using an Octet BLI and a streptavidin-coated chip (18-5136, Sartorius). In all cases, the buffer used was DPBS (14190-169, Gibco) + 0.1% BSA + 0.02% Tween-20. Purified Fabs were diluted to a final concentration of 20 nM. Kinetics were measured using the following steps: 1) sensor check for 60 seconds, 2) loading human Her2-biotin (HE2-H822R-25ug, Acro biosystems) at 5 μg / mL for 30 seconds, 3) baseline measurement for 60 seconds, 4) binding kinetics at 20 nM of each Fab for 300 seconds, 5) dissociation kinetics (in buffer) for 600 seconds.
[0328] (Table S2. ML model training:test confusion matrix) [Table 11]
[0329] (Table S3. ML model training: Test precision, recall and F1 score) [Table 12] *Precision rate is defined as TP / (TP+FP). **Recall is defined as TP / (TP+FN). ***F1 score is defined as the harmonic mean of precision and recall.
[0330] (Table S4. ML vs. Random selected clones; hit performance) [Table 13] *This total excludes single point mutations from the ML set.
[0331] The present invention may be described with reference to the following non-limiting clauses, which set forth specific aspects and embodiments of the invention. 1. A method for displaying non-DNA nucleic acid molecules on a substrate comprising: i) providing a first nucleic acid immobilized on a substrate, wherein the first nucleic acid is oriented such that its 5' end is proximal and its 3' end is distal to its point of immobilization; ii) generating a second nucleic acid complementary to the first nucleic acid, the generating step of the second nucleic acid comprising: a) contacting the first nucleic acid with a nucleic acid polymerase under conditions suitable for polymerization, wherein the primer for polymerization is a DNA primer immobilized on the substrate such that a bridge is formed during polymerization; the product of the polymerization is a non-DNA nucleotide strand immobilized on the substrate via the primer, and The nucleic acid polymerase is a polymerase capable of acting on a DNA primer to synthesize a non-DNA nucleic acid molecule complementary to a single-stranded nucleic acid template; and iii) removing said first nucleic acid resulting in the presentation of said second nucleic acid on said substrate. 2. The method of claim 1, wherein the second nucleic acid is an RNA molecule. 3. The method of clause 2, wherein the nucleic acid polymerase comprises an amino acid sequence having at least 36% identity to the amino acid sequence of SEQ ID NO:1, and the amino acid sequence comprises a Y409 and an E664 mutation relative to the amino acid sequence of SEQ ID NO:1; optionally, the Y409 mutation is Y409G and the E664 mutation is E664K; and optionally, the amino acid sequence of the nucleic acid polymerase comprises SEQ ID NO:3. 4. The method of claim 1, wherein the second nucleic acid is an XNA molecule. 5. The method of claim 4, wherein the XNA molecule comprises an arabino nucleotide, an arabino nucleic acid (ANA) nucleotide, a 2'-fluoroarabino nucleic acid (FANA) nucleotide, a 2'-O-methyl ribonucleic acid (2'OMe) nucleotide, a 2'-O-methoxyethyl (MOE) nucleic acid nucleotide, a phosphorothioate 2'-O-methoxyethyl (PS-MOE) nucleotide, a phosphorodiamidate morpholino nucleotide, a locked nucleic acid (LNA) nucleotide, a P-alkylphosphonate nucleic acid (phNA) nucleotide, a threose nucleic acid (TNA) nucleotide, a hexitol nucleic acid (HNA) nucleotide, a 2'-hydroxyhexitol (AtNA) nucleotide, a cyclohexene nucleic acid (CeNA) nucleotide, or a 3'-deoxy-DNA (2'-5') nucleotide. 6. The method of any one of clauses 1 to 5, wherein the nucleic acid polymerase comprises an amino acid sequence having at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO:1, and further comprises a mutation that enables polymerization of at least one XNA nucleotide or RNA nucleotide. 7. The method of claim 6, wherein the amino acid sequence of the nucleic acid polymerase comprises one or more or all of the following mutations: V93Q, D141A, E143A, and A485L. 8. The method of any one of clauses 1 to 7, wherein the polymerase is TGK, TGLLK, 2M, Bst, RT521, 6G12, 6G12521, C7, PGLVV, PGLVVWA, D4K, or a variant thereof. 9. The method of any one of clauses 1 to 8, wherein the method comprises, after step ii)a), cleaving the first nucleic acid to linearize the bridge. 10. The method of clause 9, further comprising recontacting the linearized product with the nucleic acid polymerase under conditions suitable for polymerization. 11. A method for displaying a non-DNA nucleic acid molecule on a substrate, comprising: i) providing a first nucleic acid immobilized on a substrate, wherein the first nucleic acid is oriented such that its 5' end is proximal and its 3' end is distal to its point of immobilization; ii) generating a second nucleic acid complementary to the first nucleic acid, the generating step of the second nucleic acid comprising: a) contacting the first nucleic acid with a nucleic acid polymerase under conditions suitable for polymerization, wherein a primer for the polymerization is immobilized on the substrate such that a bridge is formed during polymerization; and The product of the polymerization is a non-DNA nucleotide strand immobilized on the substrate via the primer; b) cleaving the first nucleic acid to linearize the bridge; and c) contacting the linearized product of step b) with a polymerase under conditions suitable for polymerization; and iii) removing said first nucleic acid resulting in the presentation of said second nucleic acid on said substrate. 12. The method of any one of clauses 1 to 11, wherein step ii)a) comprises at least 5, 10, 12, 15, 20, or 25 cycles of bridge amplification. 13. The method of any one of clauses 1 to 12, wherein in step iii) the first nucleic acid is removed by contacting the first nucleic acid with a denaturant, wherein the denaturant is a buffer comprising 100 mM NaOH and 5 mM EDTA. 14. The method of any one of clauses 1-13, wherein the second nucleic acid is an RNA molecule and encodes a polypeptide, and further comprising the step of iv) contacting the second nucleic acid with ribosomes under conditions suitable for translation of the encoded polypeptide, and optionally the conditions in step iv) include trimethylamine N-oxide (TMAO). 15. The method of clause 14, wherein the encoded polypeptide is a single chain variable fragment (scFv), a peptide, a fibronectin type III domain (FN3 domain), a single domain antibody (sdAb, also known as a nanobody), an affibody, a darpin, a finomer, an OBody, or an avimer. 16. A method for displaying a polypeptide on a substrate, comprising: i) providing a first nucleic acid comprising an antisense sequence encoding a single chain variable fragment (scFv), wherein the first nucleic acid is immobilized on a substrate and oriented such that the 5' end is proximal and the 3' end is distal to the point of immobilization; ii) generating a second nucleic acid complementary to the first nucleic acid, the generating step of the second nucleic acid comprising: contacting the first nucleic acid with a nucleic acid polymerase under conditions suitable for RNA polymerization, wherein a primer for the polymerization is immobilized on the substrate such that a bridge is formed during polymerization; and The product of the polymerization is an RNA nucleotide strand immobilized on the substrate via the primer; iii) removing the first nucleic acid resulting in the presentation of the second nucleic acid on the substrate; and iv) contacting the second nucleic acid with ribosomes under conditions suitable for translation of the encoded scFv, wherein the conditions in step iv) comprise trimethylamine N-oxide (TMAO). 17. The method according to any one of clauses 14 to 16, wherein the TMAO is at a concentration of 0.05 to 1.5 M, 0.05 to 1.2 M, or 4 M. 18. The method of any one of clauses 14 to 17, wherein the ribosome-polypeptide complex is stabilized by application of a ribosome presentation buffer, The ribosome display buffer comprises: greater than 7 mM MgCl2; or equivalent to 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, or 100 mM MgCl2 or MgAc; or is equivalent to 8-100 mM, 10-90 mM, 15-85 mM, 20-80 mM, 25-75 mM, 30-70 mM, 35-65 mM, 40-60 mM, 45-55 mM MgCl2; or The method according to any one of claims 1 to 5, wherein the magnesium concentration is equivalent to 8 to 100 mM, 10 to 90 mM, 15 to 85 mM, 20 to 80 mM, 25 to 75 mM, 30 to 70 mM, 35 to 65 mM, 40 to 60 mM, or 45 to 55 mM MgAc. 19. The method of any one of clauses 1 to 18, wherein the second nucleic acid is an RNA molecule, and a plurality of first nucleic acids encoding a plurality of polypeptides is provided in step i), such that a display library is generated. 20. The first nucleic acid immobilized on the substrate provided in step i) comprises: 1) providing a template nucleic acid; 2) hybridizing the template nucleic acid to a primer immobilized on a substrate; 3) contacting the hybridized template nucleic acid with a polymerase under conditions suitable for extension of the immobilized primer to synthesize the first nucleic acid, which is a nucleotide strand complementary to the template; 4) performing bridge amplification of the first nucleic acid to generate clusters of the first nucleic acid; and 5) sequencing at least a portion of the first nucleic acid; Optionally, the bridge amplification comprises: Contains 32-35 amplification cycles, With an extension time of 60 to 120 seconds per cycle, and / or using an amplification buffer with a Mg concentration equivalent to 2-6 mM MgSO4. 20. The method of any one of clauses 1-19, comprising use of a denaturing buffer comprising 95-99.9% formamide, optionally 1-10 mM NaOH, and optionally 1-5 mM EDTA. 21. A method for preparing clusters of substrate-bound nucleic acids, comprising: 1) providing a template nucleic acid; 2) hybridizing the template nucleic acid to a primer immobilized on a substrate; 3) contacting the hybridized template nucleic acid with a polymerase under conditions suitable for extension of the immobilized primer to synthesize the first nucleic acid, which is a nucleotide strand complementary to the template; and 4) performing bridge amplification of the first nucleic acid to generate clusters of the first nucleic acid, the bridge amplification being performed with 32-35 amplification cycles, having an extension time of 60-120 seconds per cycle, comprising the use of an amplification buffer having a Mg concentration equivalent to 2-6 mM MgSO4, and comprising the use of a denaturing buffer comprising 95-99.9% formamide, optionally 1-10 mM NaOH, and optionally 1-5 mM EDTA. 22. i) An RNA molecule obtained or obtainable by the method according to any one of clauses 1 to 3, 6 to 13 or 19 to 20; ii) an XNA molecule obtained or obtainable by the method according to any one of clauses 1, 4 to 13 or 20; or iii) A substrate presenting a polypeptide molecule obtained or obtainable by the method according to any one of clauses 14 to 20. 23. The use of a nucleic acid polymerase to extend a DNA primer immobilized on a substrate to synthesize a non-DNA nucleic acid molecule complementary to a single-stranded nucleic acid template. 24. The use of clause 23, wherein the nucleic acid polymerase comprises an amino acid sequence having at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO:1, and further comprises a mutation that enables polymerization of at least one XNA nucleotide or RNA nucleotide. 25. The use of clause 23 or clause 24, wherein the nucleic acid polymerase comprises a sequence having at least 80%, 90%, 95%, 99%, or 100% identity to the amino acid sequence of SEQ ID NO:3, and wherein residues 93, 141, 143, 409, 485, and 664 are unchanged.
Claims
1. 1. A method for displaying non-DNA nucleic acid molecules on a substrate, comprising: i) providing a first nucleic acid immobilized on a substrate, the first nucleic acid being oriented such that its 5' end is proximal and its 3' end is distal to the point of immobilization; ii) generating a second nucleic acid complementary to said first nucleic acid: a) contacting the first nucleic acid with a nucleic acid polymerase under conditions suitable for polymerization, the primer for polymerization is a DNA primer immobilized on the substrate such that a bridge is formed during polymerization; The product of the polymerization is a non-DNA nucleotide chain immobilized on the substrate via the primer; and the nucleic acid polymerase is a polymerase capable of acting on a DNA primer to synthesize a non-DNA nucleic acid molecule complementary to the single-stranded nucleic acid template; and b) cleaving the first nucleic acid to linearize the bridge; and iii) removing said first nucleic acid to result in the display of said second nucleic acid on said substrate.
2. (i) the second nucleic acid is an RNA molecule or an XNA molecule, optionally wherein the XNA molecule comprises an arabinonucleotide, an arabinonucleic acid (ANA) nucleotide, a 2'-fluoroarabinonucleic acid (FANA) nucleotide, a 2'-O-methylribonucleic acid (2'OMe) nucleotide, a 2'-O-methoxyethyl (MOE) nucleic acid nucleotide, a phosphorothioate 2'-O-methoxyethyl (PS-MOE) nucleotide, a phosphorodiamidate morpholino nucleotide, a locked nucleic acid (LNA) nucleotide, a P-alkylphosphonate nucleic acid (phNA) nucleotide, a threose nucleic acid (TNA) nucleotide, a hexitol nucleic acid (HNA) nucleotide, a 2'-hydroxyhexitol (AtNA) nucleotide, a cyclohexene nucleic acid (CeNA) nucleotide, or a 3'-deoxy-DNA (2'-5') nucleotide; and / or (ii) the nucleic acid polymerase comprises an amino acid sequence having at least 36% identity to the amino acid sequence of SEQ ID NO:1, wherein the amino acid sequence comprises a Y409 and an E664 mutation with respect to the amino acid sequence of SEQ ID NO:1; and optionally (a) the Y409 mutation is Y409N or Y409G, and the E664 mutation is E664K or E664Q; (b) the Y409 mutation is Y409G and the E664 mutation is E664K, or (c) the amino acid sequence of the nucleic acid polymerase comprises SEQ ID NO:
3.
3. (a) the nucleic acid polymerase comprises an amino acid sequence having at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO: 1, and further comprises a mutation that allows for the polymerization of at least one XNA nucleotide or RNA nucleotide, and optionally the amino acid sequence of the nucleic acid polymerase comprises one or more or all of the following mutations: V93Q, D141A, E143A, and A485L; and / or 2. The method of claim 1, wherein (b) the polymerase is TGK, TGLLK, 2M, Bst, RT521, 6G12, 6G12521, C7, PGLVV, PGLVVWA, D4K, or a variant thereof.
4. 10. The method of claim 1, further comprising the step of re-contacting the linearized product with the nucleic acid polymerase under conditions suitable for polymerization.
5. 1. A method for displaying non-DNA nucleic acid molecules on a substrate, comprising: i) providing a first nucleic acid immobilized on a substrate, the first nucleic acid being oriented such that its 5' end is proximal and its 3' end is distal to the point of immobilization; ii) generating a second nucleic acid complementary to said first nucleic acid: a) contacting the first nucleic acid with a nucleic acid polymerase under conditions suitable for polymerization, a primer for the polymerization is immobilized on the substrate such that a bridge is formed during polymerization; and the contacting, wherein the product of the polymerization is a non-DNA nucleotide chain immobilized on the substrate via the primer; b) cleaving the first nucleic acid to linearize the bridge; and c) contacting the linearized product of step b) with a nucleic acid polymerase under conditions suitable for polymerization; and iii) removing the first nucleic acid to result in the display of the second nucleic acid on the substrate, wherein optionally the second nucleic acid is an RNA molecule.
6. (i) the bridge is temperature denaturable, and / or (ii) the first nucleic acid comprises an 8-oxoguanine site and is cleaved at the 8-oxoguanine site with formamidopyrimidine DNA glycosylase (Fpg), and optionally a third nucleic acid is annealed to the first nucleic acid at the 8-oxoguanine site prior to cleavage with Fpg; and / or 2. The method of claim 1, wherein (iii) the first nucleic acid comprises a 2-deoxyuridine site and is cleaved with a Uracil-Specific Removal Reagent (USER) enzyme.
7. (I) step ii)a) comprises at least 5, 10, 12, 15, 20, or 25 cycles of bridge amplification; and / or (II) the first nucleic acid is removed in step iii) by contacting the first nucleic acid with a denaturing agent, optionally the denaturing agent being: 1-500 mM NaOH and 0-20 mM EDTA, or 100 mM NaOH and 5 mM EDTA, and / or (III)(a) the second nucleic acid is an RNA molecule, encodes a polypeptide, and further comprises: iv) contacting the second nucleic acid with ribosomes under conditions suitable for translation of the encoded polypeptide, optionally the conditions in step iv) including trimethylamine N-oxide (TMAO); and / or (b) the encoded polypeptide is an antibody fragment or an enzyme, and / or (c) the encoded polypeptide is a single-chain variable fragment (scFv), a peptide, a fibronectin type III domain (FN3 domain), a single-domain antibody (sdAb, also known as a nanobody), an affibody, a darpin, a fynomer, an OBody, or an avimer; The method according to any one of claims 1 to 6.
8. 1. A method for displaying a polypeptide on a substrate, comprising: i) providing a first nucleic acid comprising an antisense sequence encoding a single-chain variable fragment (scFv), wherein the first nucleic acid is immobilized on a substrate and oriented such that the 5' end is proximal and the 3' end is distal to the point of immobilization; ii) generating a second nucleic acid complementary to the first nucleic acid; a) contacting the first nucleic acid with a nucleic acid polymerase under conditions suitable for RNA polymerization, a primer for the polymerization is immobilized on the substrate such that a bridge is formed during polymerization; and the contacting, wherein the product of the polymerization is an RNA nucleotide strand immobilized on the substrate via the primer; and b) cleaving the first nucleic acid to linearize the bridge; iii) removing the first nucleic acid to result in the presentation of the second nucleic acid on the substrate; and iv) contacting the second nucleic acid with a ribosome under conditions suitable for translation of the encoded scFv, wherein the conditions in step iv) include trimethylamine N-oxide (TMAO).
9. (i) the TMAO is at a concentration of 0.05 to 1.5 M, 0.05 to 1.2 M, or 4 M; and / or (ii) the ribosome-polypeptide complex is stabilized by application of a ribosome presentation buffer, optionally comprising: 7 mM MgCl 2 exceeding; or 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, or 100 mM MgCl 2 or equivalent to MgAc; or 8-100 mM, 10-90 mM, 15-85 mM, 20-80 mM, 25-75 mM, 30-70 mM, 35-65 mM, 40-60 mM, or 45-55 mM MgCl 2 is equivalent to; or equivalent to 8-100 mM, 10-90 mM, 15-85 mM, 20-80 mM, 25-75 mM, 30-70 mM, 35-65 mM, 40-60 mM, or 45-55 mM MgAc; 8. The method of claim 7, wherein the magnesium concentration is
10. The method according to any one of claims 1 to 6, (i) the second nucleic acid is an RNA molecule, and a plurality of first nucleic acids encoding a plurality of polypeptides is provided in step i), such that the method generates a display library; and / or (ii) the first nucleic acid immobilized on the substrate provided in step i) comprises: 1) providing a template nucleic acid; 2) hybridizing the template nucleic acid to a primer immobilized on a substrate; 3) contacting the hybridized template nucleic acid with a polymerase under conditions suitable for extension of the immobilized primer to synthesize the first nucleic acid, which is a nucleotide strand complementary to the template; 4) performing bridge amplification of the first nucleic acid to generate clusters of the first nucleic acid; and 5) sequencing at least a portion of the first nucleic acid, and optionally the bridge amplification is generated by: containing 32-35 amplification cycles, with an extension time of 60 to 120 seconds per cycle, 2–6 mM MgSO 4 and / or using an amplification buffer with a Mg concentration equivalent to Preferably, the bridge amplification comprises the use of a denaturing buffer comprising 95-99.9% formamide, optionally 1-10 mM NaOH, and optionally 1-5 mM EDTA, and Contains 32 amplification cycles, with an extension time of 60 seconds per cycle, 6 mM MgSO 4 and / or, The method comprises using a denaturing buffer comprising 98% formamide, 10 mM NaOH, and 1 mM EDTA.
11. 1. A method for preparing clusters of substrate-bound nucleic acids, comprising: 1) providing a template nucleic acid; 2) hybridizing the template nucleic acid to a primer immobilized on a substrate; 3) contacting the hybridized template nucleic acid with a polymerase under conditions suitable for extension of the immobilized primer to synthesize the first nucleic acid, which is a nucleotide strand complementary to the template; and 4) performing bridge amplification of the first nucleic acid to generate clusters of the first nucleic acid; The bridge amplification was carried out for 32-35 amplification cycles with an extension time of 60-120 seconds per cycle, and with 2-6 mM MgSO 4 and using a denaturing buffer comprising 95-99.9% formamide, optionally 1-10 mM NaOH, and optionally 1-5 mM EDTA; and optionally The bridge amplification included 32 amplification cycles with an extension time of 60 seconds per cycle, and was performed in 6 mM MgSO 4 and / or the method comprises using a denaturing buffer comprising 98% formamide, 10 mM NaOH, and 1 mM EDTA.
12. (i) a non-DNA nucleic acid molecule obtained or obtainable by the method of claim 1; or (ii) an RNA molecule obtained or obtainable by the method of claim 1; or (iii) an XNA molecule obtained or obtainable by the method of claim 1; or (iv) A polypeptide molecule obtained or obtainable by the method of claim 7. Substrate presenting.
13. 1. Use of a nucleic acid polymerase to extend a DNA primer immobilized on a substrate to synthesize a non-DNA nucleic acid molecule complementary to a single-stranded nucleic acid template, optionally comprising: (i) the nucleic acid polymerase comprises an amino acid sequence having at least 36% similarity or identity to the amino acid sequence of SEQ ID NO: 1 and comprises Y409 and E664 mutations, and polymerizes an RNA molecule complementary to the nucleic acid template, preferably the nucleic acid polymerase comprises a sequence having at least 80%, 90%, 95%, 99%, or 100% identity to the amino acid sequence of SEQ ID NO: 3, with residues 93, 141, 143, 409, 485, and 664 remaining unchanged; and / or (ii) the nucleic acid polymerase comprises an amino acid sequence that is at least 36%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, or 100% similarity or identity to the amino acid sequence of SEQ ID NO: 1 and further comprises a mutation that allows for the polymerization of at least one XNA nucleotide or RNA nucleotide; and / or (iii) the amino acid sequence of the nucleic acid polymerase comprises one or more or all of the following mutations: V93Q, D141A, E143A, and A485L; and / or (iv) The aforementioned use, wherein the polymerase is TGK, TGLLK, 2M, Bst, RT521, 6G12, 6G12521, C7, PGLVV, PGLVVWA, D4K, or a variant thereof.
14. A method for screening a substrate displaying a plurality of biomolecules, wherein the substrate is the substrate described in claim 12 and the biomolecules form a library, and optionally the screening comprises measuring the affinity of the displayed biomolecules, non-DNA nucleic acid molecules, or polypeptide molecules for a ligand or target molecule, or measuring enzymatic function.
15. The method of any one of claims 1 to 6, further comprising a step of screening the displayed non-DNA nucleic acid molecule or polypeptide molecule, optionally wherein the screening comprises measuring the affinity of the displayed biomolecule, non-DNA nucleic acid molecule, or polypeptide molecule for a ligand or target molecule, or measuring enzymatic function.