Transmembrane beta barrels
The 12-stranded transmembrane beta barrel polypeptide design addresses the limitations of existing TMB nanopores by providing precise control over nanopore characteristics, enabling advanced molecular sensing and sequencing applications.
Patent Information
- Application Number
- PCT/EP2024/087538
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-12-19
- Publication Date
- 2025-06-26
AI Technical Summary
Existing transmembrane beta barrel (TMB) nanopores have limited tunability and desirable pore properties, making them unsuitable for advanced applications such as single-molecule enzymology, protein fingerprinting, and detection of small molecules and biomarkers.
Design and production of a 12-stranded transmembrane beta barrel polypeptide with a specific amino acid sequence and structure, allowing for precise control over nanopore characteristics such as channel size, conductivity, and stability, enabling custom engineering for various industrial applications.
The 12-stranded TMB polypeptide design achieves improved control over nanopore properties, resulting in stable, conductive, and tuneable nanopores suitable for advanced molecular sensing and sequencing technologies.
Smart Images

Figure IMGF000045_0001 
Figure IMGF000031_0001 
Figure IMGF000038_0001
Abstract
Description
[0001] TRANSMEMBRANE BETA BARRELS
[0002] Field of the invention
[0003] The present invention relates to the field of de novo (trans)membrane protein design, in particular the field of transmembrane beta barrel polypeptides and the creation, generation and uses thereof. Furthermore, the invention relates to a scaffold structure designed to provide for nanopores comprising a transmembrane beta barrel polypeptide, and uses thereof, in particular the use in molecule sensing. The invention also relates to a method for obtaining transmembrane beta barrel polypeptides and a method for producing the transmembrane beta barrel polypeptides.
[0004] Background
[0005] Transmembrane beta barrel (TMBs) nanopores provide robust scaffolds for monitoring and regulating transmembrane transport, and are of great interest for single-molecule analytical technologies. TMBs nanopores can be formed by a single, curved, p-sheet, the p-sheet being also known as transmembrane beta barrel. Some naturally occurring transmembrane beta barrels can spontaneously fold and insert into membranes and form stable pores, but the range of pore properties that can be achieved by repurposing the natural TMBs is limited. Thus, there is a need for industrially applicable pores of improved and / or tuneable characteristics. A de novo computational design coupled with a "hypothesis, design, and test" approach was taken previously in a study (Vorobieva et al., 2021, Science 371, eabc8182) to determine TMB design principles. The study focused on simplest and smallest transmembrane beta barrel architecture, resulting in de novo design of the transmembrane beta barrel comprising eight beta strands. The downside of the resulting barrel structure is that such smallest transmembrane beta barrels do not form a pore of a desirable diameter. In particular, said 8-stranded TMBs were too small to contain a central conducting channel. Similar efforts in designing nanopores are known for example from WO2022 / 051457.
[0006] There is a need for design of transmembrane beta barrels which would assemble the nanopore in lipid membranes and allow custom engineering of new sequencing and sensing technologies. There is a need for better control of the channel size of nanopores and other characteristics, such as physicochemical properties and / or conductivity thereof. There is an increased need for provision of transmembrane beta barrels and / or nanopores for applications such as single molecule enzymology, protein fingerprinting and the detection of small molecules and biomarkers and sequencing of biological and synthetic polymers.
[0007] The present invention aims to provide an alternative transmembrane beta barrel polypeptides and nanopores comprising said transmembrane beta barrel polypeptides. The features of said beta barrel polypeptides and nanopores would preferably be better controlled and allow for various industrial applications, for example small molecule and / or biomarker detection and / or sensing.
[0008] Summary
[0009] In a first aspect, the application provides a transmembrane beta barrel (TMB) polypeptide comprising the formula X1-X2-X3-X4-X5-X6-X7-X8-X9-X10-X11-X12-X13-X14-X15-X16-X17-X18-X19-X20-X21-X22- X23, wherein:
[0010] • XI comprises a beta strand, preferably said beta strand comprising or consisting of a sequence Z1-Z2-Z3-G-Z3-Z2-Z3-Z2-Y;
[0011] • X3 comprises a beta strand comprising or consisting of a sequence Z1-Z2-Z3-G-Z3-Y-Z3-Z2- Y-Z2-Z1;
[0012] • X5 comprises a beta strand comprising or consisting of the sequence Z1-Z2-Z3-Z2-Z3-Z2-Y- Z2-W;
[0013] • X7 comprises a beta strand comprising or consisting of the sequence (N / Q)-Z2-Z3-Z2-Z3-G- Z3-Z2-Y-Z2-Z1;
[0014] • X9 comprises a beta strand comprising or consisting of the sequence Z1-Z2-Z3-Z2-Z3-Z2-Z3- Z2-Y;
[0015] • Xll comprises a beta strand comprising or consisting of the sequence Z1-Z2-Z3-Z2-Z3-Z2-Z3- Z2-Z3-Z2-Y-P-Z1;
[0016] • X13 comprises a beta strand comprising or consisting of the sequence Z1-Z2-Z3-Z2-P-G-Z3- Z2-Y-Z2-W;
[0017] • X15 comprises a beta strand comprising or consisting of the sequence (N / Q)-Z4-Z3-Z2-P-Y- Z3-Z2-Z3-Z2-Y-Z2-Z1;
[0018] • X17 comprises a beta strand comprising or consisting of the sequence Z1-Z2-Z3-Z2-Z3-Z2-Z3- Z2-Y;
[0019] • X19 comprises a beta strand comprising or consisting of the sequence Z1-Z2-Z3-Z2-Z3-G-Z3- Z2-Y-Z2-Z1;
[0020] • X21 comprises a beta strand comprising or consisting of the sequence Z1-Z2-Z3-Z2-Z3-Z2-Y- G-W;
[0021] • X23 comprises a beta strand comprising or consisting of the sequence (N / Q)-Z2-Z3-Z2-Z3- Z2-Z3-Z2-Y-Z2-Z1-Z5;
[0022] • X2, X6, X10, X14, X18 and X22 each comprise a flexible loop, preferably the flexible loop comprising at least 3 and / or at most 10 amino acids; • X4, X8, X12, X16 and X20 each are a beta turn, said beta turn preferably comprising or consisting of 3 to 5 amino acids; wherein each one of XI, X3, X5, X7, X9, Xll, X13, X15, W17, X19 and X21 comprises a sequence of amino acids that alternate in their relative position within the TMB between the internal side of the lumen formed by the TMB and the surface region of the TMB, and said sequence starting and / or ending with an amino acid located in the surface region of said TMB polypeptide; and wherein X23 comprises a sequence of amino acids that alternate in their relative position within the TMB between the internal side of the lumen formed by the TMB and the surface region of the TMB, and said sequence starting and / or ending with the amino acid in the surface region of said polypeptide, except for the two amino acids on the C-terminus indicated as Z1 and Z5 which are located in the surface region; wherein said Z1 amino acid is chosen from the group consisting of V, L, A and F, and said amino acid is located in the surface region of the TMB polypeptide; wherein said Z2 amino acid is chosen from the group A, D, E, F, H, K, L, M, N, Q, R, S, T, V and W, and said amino acid is located in the pore region; wherein said Z3 amino acid is chosen from the group consisting of V, L, A, G, S, T and F, and said amino acid is located in the surface region; and wherein said Z4 amino acid is chosen from the group consisting of A, D, E, H, K, L, M, N, Q, R, S, T and V, and said amino acid is located in the pore region; wherein Z5 is a single amino acid for a bulge in the TMB structure, and wherein X represents an element of the polypeptide defined by an amino acid sequence, said element comprising at least two amino acids and / or amino acid residues and / or said element having a secondary structure, wherein each of the "X" elements is connected or linked to other elements through peptide bonds.
[0023] The term "Z" as used herein, relates to an amino acid and / or a residue thereof. Said amino acid and / or residue are preferably being chosen from a group, and are connected or linked to other amino acids through peptide bonds.
[0024] The transmembrane beta barrel polypeptide according to the first aspect of the invention relates to new and improved transmembrane beta barrel construct comprising or consisting of twelve beta strands and allowing for an expression in a host cell, such as Escherichia coli. Furthermore, the polypeptide according to the first aspect of the invention may be seen as better tuneable structure which can allow various applications. The TMB polypeptide is a synthetic, non-natural man-made TMB, which can allow for better prediction of its structure and characteristics. In a preferred embodiment, the invention relates to a nucleic acid encoding the TMB polypeptide according to any embodiment of the first aspect. Said nucleic acid would allow for easier production of the transmembrane beta barrel polypeptide of any embodiment according to the first aspect.
[0025] In a preferred embodiment, the invention concerns an expression vector comprising the nucleic acid encoding the TMB polypeptide according to any embodiment of the first aspect, said expression vector preferably being operably linked to a control sequence, such as a chimeric gene construct or promoter sequence for expression of the coding sequence of the TMB. Said expression vector allows for easier production of the transmembrane beta barrel polypeptide of any embodiment according to the first aspect.
[0026] In a particularly preferred embodiment, the recombinant host cell is provided, said cell comprising the nucleic acid encoding the TMB polypeptide according any embodiment of the first aspect or the expression vector comprising said nucleic acid. Said cell is preferably a recombinant cell, more preferably a recombinant microorganism, most preferably a recombinant bacterium. In a particularly preferred embodiment, said cell is a recombinant E. coli cell. Such recombinant host cell may allow for more efficient 12-stranded TMB polypeptide production which may be scaled up to the industrial quantities. Said host cell may allow for more efficient and better targeted TMB polypeptide production, which TMB polypeptide can be designed for a specific use.
[0027] According to a second aspect, the invention relates to a nanopore comprising a lumen and a wall, wherein said wall comprises the TMB polypeptide according to the any embodiment of first aspect of the invention. Said nanopore is can have desired and / or adjustable lumen diameter, and can be preferably of a square cross shape. Said nanopore can be used in various molecular sensing applications. Moreover, said nanopore may be of an improved electrophysiology in comparison to the pores known in the art, as it allows for a precision design of a nanopore wall, said wall being preferably more "silent" and / or with minimal intermediate structures blocking the lumen. The nanopore according to the second aspect can be characterized as a more stable pore with rigid scaffold which allows an improved conductivity. Said pore can allow for a better transport of desired molecules, preferably of smallmolecules. Furthermore, said pore may be further engineered to allow for selected and / or controllable detection of molecules, also known as molecular sensing. In yet another aspect, the present invention relates to a use of the TMB polypeptide according any one of the embodiments according to the first aspect, the nucleic acid encoding said TMB polypeptide, the expression vector comprising said nucleic acid, the recombinant host cell, said cell comprising said nucleic acid or said expression vector, and / or nanopore according to any one of embodiments according to the second aspect for molecular sensing and / or receiving and / or capturing of small-molecules or (poly)peptides or antigens, and / or drug delivery and / or as ion, water or small-molecule permeable transmembrane channels. The TMB polypeptide and the nanopore of the present invention can be designed to achieve any desirable structural property, which property, for example, is suited for a particular technological use. Said use would allow broader possibilities and / or alternative to known use, or an improvement for an existing use. Naturally occurring channels such as CsgG, which have suitable to carry out limited functions very different from for example, molecular sensing, and are difficult to be repurposed for other needs in the art. The TMB polypeptide and / or nanopore of the invention are of a design which can allow for repurposing and modifying the structure thereof, in order to adjust to a specific use.
[0028] Another aspect of the present invention relates to a pharmaceutical composition, said composition comprising:
[0029] • the TMB polypeptide according to any one of any one of the embodiments of the first aspect, the nucleic acid encoding said polypeptide, the expression vector comprising said nucleic acid, the recombinant host cell comprising said expression vector or said nucleic acid and / or the nanopore according to any one of embodiments according to the second aspect; and
[0030] • a carrier, preferably pharmaceutically acceptable carrier, or a diluent or excipient.
[0031] Said method allows for a design and / or engineering of the TMB polypeptides tailored to a particular need, for example metabolite and / or small-molecule sensing. The model has been surprisingly shown to generate TMB polypeptides which could be expressed in a host cell, for example a bacterial cell such as E. coli.
[0032] In yet another aspect, the present invention relates to a computer-implemented method for obtaining the TMB polypeptide according to any one of embodiments of the first aspect, said method comprising:
[0033] • Generating backbones corresponding to the TMB polypeptide according to any one of embodiments of the first aspect;
[0034] • Assembling the generated backbones using the Rosetta BlueprintBDR.
[0035] Said method allows to obtain the TMB polypeptide and / or the nanopore according to the present invention, which are of such design and advantages to the known TMB and TMB pores, and can be suited for various uses. In yet another aspect, the present invention relates to a method for producing the TMB polypeptide according to any one of embodiments of the first aspect, the method comprising:
[0036] • culturing a host cell under conditions conducive to the expression of the polypeptide;
[0037] • recovering the expressed polypeptide; and
[0038] • optionally, reconstitute the polypeptide in vitro in detergent micelles or lipids.
[0039] The TMB polypeptides and / or nanopores of the invention are designed in terms of structural and physicochemical characteristics which allows expression, preferably high expression thereof in a suited host cell, for example in E. coli cell.
[0040] Brief description of the Figures
[0041] Figure 1. Sculpting p-barrel geometry. A. Pore diameter can be controlled through the number of p- strands in the p-barrel blueprint structure. B. p-barrel 2D interaction map. Strong bends in the p-strands (< 90° bend, right enlarged square) can be achieved by stacking several glycine kink residues (grey spheres in enlarged squares) along the p-barrel axis, as opposed to placing one kink (>90° bend, left enlarged square). C-D. Cross-sections of explicitly assembled p-barrel backbone polypeptides without (cylinder, C) and with (D) glycine kinks. The cp atoms of the residues facing the pore are shown as spheres and shaded based on their respective repulsion energy. Glycine kinks positions are shown with arrows; placement at the corners of the embedded rectangular, oval and triangular shapes (dashed lines in D) generates the desired backbone geometries. E. Polar threonine residues are tolerated on the membrane- exposed surface of TMBs (right) as they can form a hydrogen bond to the backbone, mimicking the interactions with water molecules observed in similarly curved areas of water-exposed p-strands (left).
[0042] Figure 2. TMB Blueprint of the invention. A. TMB blueprint corresponding to a beta barrel of 12 strands (N=12) and a shear number of 14 (S=14) and a square cross shape. The residues facing the beta barrel lumen and surface are shown as grey and white circles, respectively. Glycine kinks are shown as smaller dark grey circles and are facing the lumen. The tyrosine residues belonging to the Tyr-Gly-Asp / Glu folding motif are indicated by grey hexagons. B. The p-strands of de novo designed TMB polypeptides are connected with short p-turns on both sides of the barrel: cis-hairpins (N- and C-termini side) are connected with canonical type I p-turns preceded by a p-bulge; trans-hairpins are connected with type I P-turns directly followed by a G-bulge. By comparison, naturally-occurring TMBs (exemplified here by OmpG, right) feature mostly long, disordered, loops on the trans side.
[0043] Figure 3. Characterization of TMB polypeptide consisting of SEQ IN No. 3. A. Coomassie stained SDS- PAGE gels showing expression bands for the TMB polypeptide consisting of SEQ ID No. 3 from insoluble fractions of corresponding lysed cell pellets. B. Folding kinetics of the TMB polypeptide consisting of SEQ ID No. 3 (A) in DUPC (upper set of lines) and DMPC (lower set of lines) LUVs at 30°C and monitored by intrinsic tryptophan fluorescence. The difference in folding rates associated with the length of the lipid chain is consistent with intramembrane folding. C. Secondary chemical shifts of TMB polypeptide represented by SEQ ID No. 3 from sequence-specific resonance assignments. Consecutive stretches of large negative values indicate the presence of p-strand secondary structure. The positions of the 12 p- strands are indicated by straight lines on the top.
[0044] Figure 4. Biophysical characterization of a designed nanopore. The 12-stranded TMB polypeptide (consisting of SEQ ID No. 3) formed a nanopore with a square cross-section. The design elutes predominantly as monodispersed species with retention times consistent with monomeric protein in complex with DPC detergent (A), show cooperative and reversible folding / unfolding transitions in DUPC liposomes (B), dispersed NMR 1H-15N HSQC spectra (C) and distinct negative maxima in far UV CD spectra at 215 nm in DUPC LUVs, but not in 8 M urea or in the absence of lipid (D).
[0045] Figure 5. A. Raw unfiltered current traces of one example for a nanopore comprising 12-stranded TMB polypeptide consisting of sequence SEQ ID No. 3 recorded at 5kHz sampling rate. Applied voltage is 100 mV and the cis and trans buffer for all conditions is 500mM NaCI. Low noise levels are a result of the different bilayer capacitances at the time of recording and noise from adjacent cavities in the MECA recording chip from Nanion. 10s reads show characteristics of a stable non-gating pore in the membrane. B. Current vs Voltage plots for a nanopore design comprising a TMB polypeptide consisting of SEQ ID No. 3. All measurements were carried out in 500mM NaCI solution (symmetric across bilayer). C. MOLE 2.5 pore size calculations (left), charge and hydrophobicity profiles (right) for a nanopore comprising the TMB polypeptide of SEQ ID No. 3.
[0046] Figure 6. Experimentally determined nanopore structures adopt structures that closely align with the computational design models. A-C. TMB polypeptide (SEQ ID No. 3) structure in LDAO micelles. A. 2D [^N Hj-TROSY NMR spectrum of [U-2H,15N]- TMB polypeptide (SEQ ID No. 3) with sequence-specific resonance assignments. B. Long-range NMR NOE contacts mapped to the expected TMB12_3 hydrogen bonds (dashed black lines). Residues with amide assignment are shown in white and light gray, unassigned residues are shown in dark grey. Residues with p-sheet secondary structure are shown as squares, all others as circles. Bold outlines indicate available methyl assignments. NOE contacts are shown as thick grey lines (long-range amide-amide, dashes indicating diagonal overlap) and double black lines (contacts involving side chain methyl groups). C. Ensemble of the 20 lowest energy solution NMR structures (P-sheets shown in brown). D. Strips from the 3D [1H,1H]-NOESY-15N-TROSY experiment of the TMB polypeptide consisting of SEQ ID No. 3 (i.e. TMB12_3) in LDAO micelles. Strips were taken for the residue pairs involved in the antiparallel pi-pi2 pairing. The NOE cross peaks are connected to the diagonal peaks by blue dashed lines. D. Strips from the 3D [1H,1H]-NOESY-15N-TROSY experiment of the TMB polypeptide represented by SEQ IN No. 3 in LDAO micelles. Strips were taken for the residue pairs involved in the antiparallel pi-pi2 pairing. The NOE cross peaks are connected to the diagonal peaks by dashed lines.
[0047] Figure 7. Conductance of the nanopore comprising a TMB polypeptide (SEQ ID No. 3). A. i) Top view cartoon representation, ii) Vertical cross sections of the pore, iii) single channel conductance (smallest observed conductance jump), iv) sequential insertions of designed pore in planar lipid bilayer membrane from detergent solubilized sample at low concentrations, v) histogram of smallest measured current jumps for each design up to 50 pA. The applied voltage across the bilayer was lOOmV and experiments were performed in a buffer containing 500 mM NaCI. A gaussian mixture model was used to fit the histogram peaks observed. B. i) Respective cartoons indicating top view of the designs, ii) Size Exclusion Chromatography plots for all designs carried out in buffer containing 0.1% DPC detergent, iii) Corresponding Circular Dichroism (CD) plots in the near-UV range, iv) CD melt plots from 25°C to 95°C.
[0048] Figure 8. Circular dichroism (CD) spectra in DPC micelles. The samples are collected at 25°C (black line) and 95°C (grey line) for designs corresponding to the TMB polypeptides of three sequences SEQ ID No. 3 (A), SEQ ID No. 5 (B) and SEQ ID No. 8 (C). All three polypeptides show the 218 nm minimum characteristic of proteins composed mainly of p-sheet secondary structure.
[0049] Figure 9. The nanopores comprising TMB polypeptides of SEQ ID. 5 (first row A-D) and SEQ ID 9 (second raw A-D). Both nanopores feature monodispersed SEC elution profiles consistent with a monomeric 12- stranded TMP polypeptide (A) and CD spectra characteristic of p-sheet proteins (D) in DUPC LUVs but not in denaturing conditions (8 M urea) or in the absence of lipid. TMB polypeptide of SEQ ID 9 (second row) cooperatively and reversibly folded in an urea titration experiment (B) and had a dispersed NMR HSQC spectra (C). (E) Stable nanopore activity experiments showing the nanopore activity of pores comprising a TMB polypeptide of SEQ ID No. 3 (E left) and SEQ ID No. 9 (E right).
[0050] Detailed description
[0051] Definitions
[0052] The present invention will be described with respect to particular embodiments and with reference to certain drawings but the invention is not limited thereto but only by the claims. Any reference signs in the claims shall not be construed as limiting the scope. It is to be understood that not necessarily all aspects or advantages may be achieved in accordance with any particular embodiment of the invention. Thus, for example those skilled in the art will recognize that the invention may be embodied or carried out in a manner that achieves or optimizes one advantage or group of advantages as taught herein without necessarily achieving other aspects or advantages as may be taught or suggested herein. The drawings described are only schematic and are non-limiting. In the drawings, the size of some of the elements may be exaggerated and not drawn on scale for illustrative purposes.
[0053] The invention, both as to organization and method of operation, together with features and advantages thereof, may best be understood by reference to the following detailed description when read in conjunction with the accompanying drawings. The aspects and advantages of the invention will be apparent from and elucidated with reference to the embodiment(s) described hereinafter. Reference throughout this specification to "one embodiment" or "an embodiment" means that a particular feature, structure or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, appearances of the phrases "in one embodiment" or "in an embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment, but may. Similarly, it should be appreciated that in the description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof for the purpose of streamlining the disclosure and aiding in the understanding of one or more of the various inventive aspects. This method of disclosure, however, is not to be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single foregoing disclosed embodiment.
[0054] Where an indefinite or definite article is used when referring to a singular noun e.g. "a" or "an", "the", this includes a plural of that noun unless something else is specifically stated. Where the term "comprising" is used in the present description and claims, it does not exclude other elements or steps. Furthermore, the terms first, second, third and the like in the description and in the claims, are used for distinguishing between similar elements and not necessarily for describing a sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances and that the embodiments, of the invention described herein are capable of operation in other sequences than described or illustrated herein. The following terms or definitions are provided solely to aid in the understanding of the invention. Unless specifically defined herein, all terms used herein have the same meaning as they would to one skilled in the art of the present invention. Practitioners are particularly directed to Sambrook et al., Molecular Cloning: A Laboratory Manual, 4th ed., Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016), for definitions and terms of the art. The definitions provided herein should not be construed to have a scope less than understood by a person of ordinary skill in the art.
[0055] The terms "protein" and "polypeptide" are interchangeably used further herein to refer to a polymer of amino acid residues and to variants and synthetic analogues of the same. Thus, these terms apply to amino acid polymers in which one or more amino acid residues is a synthetic non-naturally occurring amino acid, such as a chemical analogue of a corresponding naturally occurring amino acid, as well as to naturally-occurring amino acid polymers. This term also includes posttranslational modifications of the polypeptide, such as glycosylation, phosphorylation and acetylation. Based on the amino acid sequence and the modifications, the atomic or molecular mass or weight of a polypeptide is expressed in (kilo)dalton (kDa). By "recombinant polypeptide" is meant a polypeptide made using recombinant techniques, i.e., through the expression of a recombinant or synthetic polynucleotide. When the chimeric polypeptide or biologically active portion thereof is recombinantly produced, it is also preferably substantially free of culture medium, i.e., culture medium represents less than about 20 %, more preferably less than about 10 %, and most preferably less than about 5 % of the volume of the protein preparation. By "isolated" is meant material that is substantially or essentially free from components that normally accompany it in its native state. The expression "heterologous protein" may mean that the protein is not derived from the same species or strain that is used to display or express the protein.
[0056] "Amino acids" and "amino acid residues" are interchangeably used further herein to refer to molecules which are building units of proteins.
[0057] "Homologue", "Homologues" of a protein encompass peptides, oligopeptides, polypeptides, proteins and enzymes having amino acid substitutions, deletions and / or insertions relative to the unmodified protein in question and having similar biological and functional activity as the unmodified protein from which they are derived. The term "amino acid identity" as used herein refers to the extent that sequences are identical on an amino acid-by-amino acid basis over a window of comparison. Thus, a "percentage of sequence identity" is calculated by comparing two optimally aligned sequences over the window of comparison, determining the number of positions at which the identical amino acid residue occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison (i.e., the window size), and multiplying the result by 100 to yield the percentage of sequence identity. A "substitution", or "mutation" as used herein, results from the replacement of one or more amino acids or nucleotides by different amino acids or nucleotides, respectively as compared to an amino acid sequence or nucleotide sequence of a parental protein or a fragment thereof. It is understood that a protein or a fragment thereof may have conservative amino acid substitutions which have substantially no effect on the protein's activity.
[0058] A "protein domain" is a distinct functional and / or structural unit in a protein. Usually a protein domain is responsible for a particular function or interaction, contributing to the overall role of a protein. Domains may exist in a variety of biological contexts, where similar domains can be found in proteins with different functions. Protein secondary structure elements (SSEs) typically spontaneously form as an intermediate before the protein folds into its three dimensional tertiary structure. The two most common secondary structural elements of proteins are alpha helices and beta (P) sheets, though -turns and omega loops occur as well. Beta sheets consist of beta strands (also p-strand) connected laterally by at least two or three back-bone hydrogen bonds, forming a generally twisted, pleated sheet. A p-strand is a stretch of polypeptide chain typically 3 to 10 amino acids long with backbone in an extended conformation. Thus, p-sheets are an assembly of p-strands, often bringing together parts of the protein which are separated along the backbone. Two arrangements of the p-sheets are possible within a strand: Parallel (pP) and antiparallel (p ) arrangements. Said arrangements of strands in the sheet differ in the hydrogen bond pattern between strands and in the type of chain connectivity they allow: short reverse turn connections are characteristic for pAand longer crossover connections for pP.
[0059] A beta (P) turn is a type of non-regular secondary structure in proteins that causes a change in direction of the polypeptide chain. Beta turns (P turns, p-turns, p-bends, tight turns, reverse turns) are very common motifs in proteins and polypeptides, which mainly serve to connect p-strands. A beta bulge is a region between two consecutive beta-type hydrogen bonds which includes two residues (positions 1 and 2) on one strand opposite a single residue (position x) on the other strand. In other words, the beta bulge can be described as a localized disruption of the regular hydrogen bonding of beta sheet by inserting extra residues into one or both hydrogen bonded p-strands.
[0060] "Loops" are irregular structures which connect two secondary structure elements in proteins. They are generally located on the protein's surface in solvent exposed areas in the proteins (Choi Y. et al., "How long is a piece of loop?", PeerJ., 2013, l:el). Generally, loops are longer in amino acid number than turns, see, e.g., Milner-White and Poet, "Loops, bulges, turns and hairpins in proteins", Trends in Biochemical Sciences, 1987, 12:189-192. For instance, loops that have only 4 or 5 amino acid residues, when they have internal hydrogen bonds, can also be referred to as turns. The term "flexible loop" is known to a person in the art (for example disclosed in Kempner, 1993, FEBS letters: 1,2,3, pp. 1-10), and it refers to a part of the polypeptide or a protein, which part is prone to small stretches of amino acid residues leading to structural deformations of the polypeptide or protein. Said flexible loops may have, for example, at least two stable conformations, preferably three, at least four, at least five, at least six stable conformations. The loops in proteins often move in response to small molecules, covering small surface areas of the polypeptide and / or protein; ligands may be trapped and solvent molecules may be excluded. The flexible loops are found in many naturally occurring (for example streptavidin) and synthetic polypeptides and / or proteins and have profound biochemical consequences, such as permitting reactions which would otherwise be prevented. Sequence-based structure prediction tools are known to the skilled person to determine whether a peptidic sequence as part of the TMB proposed in the present invention is defined as a flexible loop or loop within said TMB 3D structure.
[0061] Amino acids are presented herein by their 3- or 1-lettercode nomenclature as defined and provided also in the IUPAC-IUB Joint Commission on Biochemical Nomenclature (Nomenclature and Symbolism for Amino Acids and Peptides. Eur. J. Biochem. 138: 9-37 (1984)); as follows: Alanine (A or Ala), Cysteine (C or Cys), Aspartic acid (D or Asp), Glutamic acid (E or Glu), Phenylalanine (F or Phe), Glycine (G or Gly), Histidine (H or His), Isoleucine (I or He), Lysine (K or Lys), Leucine (L or Leu), Methionine (M or Met), Asparagine (N or Asn), Proline (P or Pro), Glutamine (Q. or Gin), Arginine (R or Arg), Serine (S or Ser), Threonine (T or Thr), Valine (V or Vai), Tryptophan (W or Trp), and Tyrosine (Y or Tyr).
[0062] As described herein each position of an amino acid sequence as present in the Xn elements is indicated by a generic amino acid symbol (e.g. Zl, Zx, ...) or by the 1-letter code of the amino acid as generally known. For instance, the element 'X15' may be defined as comprising a beta strand comprising or consisting of the sequence '(N / Q)-Z5-Z3-Z2-P-Y-Z3-Z2-Z3-Z2-Y-Z2-Z1', wherein the (N / Q) means that the first amino acid residue of X15 is N (Asn) or Q(Gln), followed by the amino acid defined as Z5, which is an amino acid chosen from the group consisting of A, D, E, H, K, L, M, N, Q, R, S, T and V; Z3 is the following amino acid chosen from the group of V, L, A, G, S, T and F,; Z2 is the following amino acid chosen from the group of A, D, E, F, H, K, L, M, N, Q, R, S, T, V and W; P (Pro) is the following amino acid, then Y(Tyr) is the following amino acid, etc.
[0063] The term "circular permutation of a polypeptide", "a protein which is circularly permuted", "a circularly permutated variant of a polypeptide" or "circularly permuted polypeptide", as interchangeably used herein, refers to a molecule which in its linear form has the termini joined together, either directly or through a linker, to produce a circular molecule (as an intermediate), followed by opening or cleaving of the circular molecule at another location or position to produce a new molecule which in its linear form is a molecule with termini (XI and X2) different from the termini in the original molecule. The opening or cleaving of the circular molecule at another location may comprise the removal of one or more nucleotides / amino acids of the original sequence. For instance, at least one such as one, two, three, four, five or more residues may be removed when opening or cleaving the circular molecule. Circularly permuted molecules include those molecules whose structure is equivalent to a molecule that has been circularized and then opened, and / or with regards to proteins include those wherein the amino and carboxy ends are joined together, directly or through a linker, and new amino and carboxy terminal ends are formed at a different location within the protein sequence. Alternatively, a circularly permuted molecule may also be synthesized de novo starting from a new linear form of the molecule (as compared to the original molecule) and never go through a circularization and opening step. Circularly permuted molecules provide for a rearrangement in the molecule as compared to the original wild type molecule, though without impact on activity or functionality, since the folding or appearance of the final (folded) molecule is similar or the same as the original molecule, with the only difference that the beginning and end point is at a different location. As stated above, in some cases, one or more nucleotides / amino acids of the original molecule are removed from the original molecule. So circularly permuted molecules, which may be nucleic acid molecules, or proteins, have their normal termini fused, often with a linker, and contain new termini at another position. See Goldenberg, et al. J. Mol. Biol., 165: 407-413 (1983) and Pan et al. Gene 125: 111-114 (1993), both incorporated by reference herein. Circular permutation is functionally equivalent to taking a straight-chain molecule, fusing the ends to form a circular molecule, and then cutting the circular molecule at a different location to form a new straight chain molecule with different termini. Circular permutation thus has the effect of essentially preserving the sequence and identity of the amino acids of a protein while generating new termini at different locations (also see Pastan et al. - EP 0754 192 Bl). Hence, in the context of the present invention, the terms "circular permutation", "circularly permuted", or "circularly permutated", or "circular permutated variant", refer to the process of taking a TMP polypeptide sequence, or its cognate nucleic acid sequence, and fusing the N- and C-termini (directly or through a linker, e.g., using protein or recombinant DNA methodologies) to form a circular molecule, and then cutting (opening) the circular molecule at a different location, preferably in the region of a loop or a turn of said TMB polypeptide, to form a new protein, or cognate nucleic acid molecule, with termini different from the termini in the original molecule. Circular permutation thus preserves the overall sequence (besides the linkers, if introduced, and the one or more amino acids removed, if any), structure, and function of a protein, while generating new C- and / V-termini at different locations that results in an improved orientation for fusing a desired polypeptide fusion partner as compared to the original molecule. As stated above, a circularly permuted molecule may be synthesized de novo as a linear molecule and never go through a circularization and opening step. In addition, in the context of the present invention, the fusion of the N- and C-termini of the molecule (protein) may take place between the original / V- and C-termini of the protein or may take place between the N- and C-termini created after deletion of one or more residues, such as one, two, three, four, five or more residues from the original / V-terminus, between the N- and C-termini created after deletion of one or more residues, such as one, two, three, four, five or more residues from the original C-terminus, or between the N- and C-termini created after deletion of one or more residues, such as one, two, three, four, five or more residues from the both the original N- and C-termini. In the context of the present invention, the opening of the circular molecule is preferably performed at a R-turn or loop of said TMB protein, so that the folding (3D structure) of the circularly permuted TMB protein is retained or similar as compared to the folding of the wild-type protein. Therefore, in the context of the present invention, the term "circular permutation of a protein" or "circularly permuted protein variant" refers to a protein which has a changed order of amino acids in its amino acid sequence, as compared to the wild-type protein sequence, with as a result a protein structure with different connectivity, but overall similar three-dimensional (3D) shape. A circular permutation of a protein is analogous to the mathematical notion of a cyclic permutation, in the sense that the sequence of the first portion of the wild-type protein (adjacent to the / V-terminus) is related to the sequence of the second portion of the resulting circularly permuted protein (near its C-terminus), as described for instance in Bliven and Prlic (2012) (Circular permutation in proteins. PLOS Comput. Biol. 8(3):el002445). A circular permutation of a protein as compared to its wild protein is obtained through genetic or artificial engineering of the protein sequence, whereby the N- and C-terminus of the wild-type protein are 'connected' (directly, by means of a linker and / or with one or more amino acids having been removed, as explained above) and the protein sequence is interrupted at another site (where one or more amino acids can be removed, as explained above), to create a novel N- and C-terminus of said protein. The circularly permuted proteins of the invention (cytokines) are the result of a connected N- and C-terminus of the wild-type cytokine sequence, and a cleavage or interrupted sequence at an accessible or exposed site (preferentially a p-turn or loop) of said cytokine, whereby the folding (3D structure) of the circularly permuted cytokine is retained or similar as compared to the folding of the wild-type protein. Said connection of the N- and C-terminus in said circularly permuted cytokine may be the result of a peptide bond linkage, or of introducing a peptide linker, or of a deletion of a peptide stretch near the original N- and C-terminus in the wild-type protein, followed by a peptide bond or the remaining amino acids. The terms "circularly permuted" and "circular permutation" are well known in the art, see, e.g., "CPSARST: Circular Permutation Search Aided by Ramachandran Sequential Transformation"
[0064] (http: / / 140.113.120.231 / ~lab / iSARST 2019 / srv / index.php?c=m2), Lo WC, Lyu PC. CPSARST: an efficient circular permutation search tool applied to the detection of novel protein structural relationships. Genome Biol. 2008 Jan 18;9(1):R11, or "CPDB - the Circular Permutation Database" (http: / / 10.life.nctu.edu.tw / cpdb / ). Circular permutation is performed to obtain a circularly permuted protein.
[0065] The term "fused to", as used herein, and interchangeably used herein as "connected to", "conjugated to", "ligated to", "linked to" or "bound to" refers, in particular, to "genetic fusion", e.g., by recombinant DNA technology, as well as to "chemical and / or enzymatic conjugation" resulting in a stable covalent link. As used herein, "a molecular sensor", also known referred to as a probe, is a molecular or supramolecular system that is able to transform probe-analyte interactions into a signal which allows analyte sensing through optical and / or electrochemical changes.
[0066] "Expression vector", as used herein, includes vectors that operatively link a nucleic acid coding region or gene to any control sequences capable of effecting expression of the gene product. "Control sequences" operably linked to the nucleic acid sequences of the disclosure are nucleic acid sequences capable of effecting the expression of the nucleic acid molecules. The control sequences need not be contiguous with the nucleic acid sequences, so long as they function to direct the expression thereof. Thus, for example, intervening untranslated yet transcribed sequences can be present between a promoter sequence and the nucleic acid sequences and the promoter sequence can still be considered "operably linked" to the coding sequence. Other such control sequences include, but are not limited to, polyadenylation signals, termination signals, and ribosome binding sites. Such expression vectors can be of any type, including but not limited plasmid and viral -based expression vectors. The control sequence used to drive expression of the disclosed nucleic acid sequences in a mammalian system may be constitutive (driven by any of a variety of promoters, including but not limited to, CMV, SV40, RSV, actin, EF) or inducible (driven by any of a number of inducible promoters including, but not limited to, tetracycline, ecdysone, steroid-responsive). The expression vector is replicable in the host organisms either as an episome or by integration into host chromosomal DNA. In various embodiments, the expression vector may comprise a plasmid, viral-based vector, or any other suitable expression vector. In another aspect, the disclosure provides host cells that comprise the nucleic acids or expression vectors (i.e.: episomal or chromosomally integrated) disclosed herein, wherein the host cells can be either prokaryotic or eukaryotic. The cells can be transiently or stably engineered to incorporate the expression vector of the disclosure, using techniques including but not limited to bacterial transformations, calcium phosphate co-precipitation, electroporation, or liposome mediated-, DEAE dextran mediated-, polycationic mediated-, or viral mediated transfection.
[0067] "Nucleotide sequence", "DNA sequence" or "nucleic acid molecule(s)" as used herein refers to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. This term refers only to the primary structure of the molecule. Thus, this term includes double- and single-stranded DNA, and RNA. It also includes known types of modifications, for example, methylation, "caps" substitution of one or more of the naturally occurring nucleotides with an analog. By "nucleic acid construct" it is meant a nucleic acid sequence that has been constructed to comprise one or more functional units not found together in nature. Examples include circular, linear, double-stranded, extrachromosomal DNA molecules (plasmids), cosmids (plasmids containing COS sequences from lambda phage), viral genomes comprising non-native nucleic acid sequences, and the like. "Coding sequence" is a nucleotide sequence, which is transcribed into mRNA and / or translated into a polypeptide when placed under the control of appropriate regulatory sequences. The boundaries of the coding sequence are determined by a translation start codon at the 5'-terminus and a translation stop codon at the 3'-terminus. A coding sequence can include, but is not limited to mRNA, cDNA, recombinant nucleotide sequences or genomic DNA, while introns may be present as well under certain circumstances.
[0068] "Promoter region of a gene" as used here refers to a functional DNA sequence unit that, when operably linked to a coding sequence and possibly placed in the appropriate inducing conditions, is sufficient to promote transcription of said coding sequence. "Operably linked" refers to a juxtaposition wherein the components so described are in a relationship permitting them to function in their intended manner. A promoter sequence "operably linked" to a coding sequence is ligated in such a way that expression of the coding sequence is achieved under conditions compatible with the promoter sequence. "Gene" as used here includes both the promoter region of the gene as well as the coding sequence. It refers both to the genomic sequence (including possible introns) as well as to the cDNA derived from the spliced messenger, operably linked to a promoter sequence. The term "terminator" or "transcription termination signal" encompasses a control sequence which is a DNA sequence at the end of a transcriptional unit which signals 3' processing and polyadenylation of a primary transcript and termination of transcription. The terminator can be derived from the natural gene, from a variety of other plant genes, or from T-DNA. The terminator to be added may be derived from, for example, the nopaline synthase or octopine synthase genes, or alternatively from another plant gene, or less preferably from any other eukaryotic gene.
[0069] With a "chimeric gene" or "chimeric construct" or "chimeric gene construct" is meant a recombinant nucleic acid sequence in which a promoter or regulatory nucleic acid sequence is operatively linked to, or associated with, a nucleic acid sequence that codes for an mRNA, such that the regulatory nucleic acid sequence is able to regulate transcription or expression of the associated nucleic acid coding sequence. The regulatory nucleic acid sequence of the chimeric gene is not operatively linked to the associated nucleic acid sequence as found in nature.
[0070] An "expression cassette" comprises any nucleic acid construct capable of directing the expression of a gene / coding sequence of interest, which is operably linked to a promoter of the expression cassette. Expression cassettes are generally DNA constructs preferably including (5' to 3' in the direction of transcription): a promoter region, a polynucleotide sequence, homologue, variant or fragment thereof operably linked with the transcription initiation region, and a termination sequence including a stop signal for RNA polymerase and a polyadenylation signal. It is understood that all of these regions should be capable of operating in biological cells, such as prokaryotic or eukaryotic cells, to be transformed. The promoter region comprising the transcription initiation region, which preferably includes the RNA polymerase binding site, and the polyadenylation signal may be native to the biological cell to be transformed or may be derived from an alternative source, where the region is functional in the biological cell. Such cassettes can be constructed into a "vector".
[0071] The term "vector", "vector construct," "expression vector," or "gene transfer vector," as used herein, is intended to refer to a nucleic acid molecule capable of transporting another nucleic acid molecule to which it has been linked, and includes any vector known to the skilled person, including any suitable type, including, but not limited to, plasmid vectors, cosmid vectors, phage vectors, such as lambda phage, viral vectors, such as adenoviral, AAV or baculoviral vectors, or artificial chromosome vectors such as bacterial artificial chromosomes (BAC), yeast artificial chromosomes (YAC), or Pl artificial chromosomes (PAC). Expression vectors comprise plasmids as well as viral vectors and generally contain a desired coding sequence and appropriate DNA sequences necessary for the expression of the operably linked coding sequence in a particular host organism (e.g., bacteria, yeast, plant, insect, or mammal) or in in vitro expression systems. Expression vectors are capable of autonomous replication in a host cell into which they are introduced (e.g., vectors having an origin of replication which functions in the host cell). Other vectors can be integrated into the genome of a host cell upon introduction into the host cell, and are thereby replicated along with the host genome. Suitable vectors have regulatory sequences, such as promoters, enhancers, terminator sequences, and the like as desired and according to a particular host organism (e.g. bacterial cell, yeast cell). Cloning vectors are generally used to engineer and amplify a certain desired DNA fragment and may lack functional sequences needed for expression of the desired DNA fragments. The construction of expression vectors for use in transfecting prokaryotic cells is also well known in the art, and thus can be accomplished via standard techniques (see, for example, Sambrook, et al. Molecular Cloning: A Laboratory Manual, 4th ed., Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016), for definitions and terms of the art.
[0072] "Host cells" can be either prokaryotic or eukaryotic. The cells can be transiently or stably transfected. Such transfection of expression vectors into prokaryotic and eukaryotic cells can be accomplished via any technique known in the art, including but not limited to standard bacterial transformations, calcium phosphate co-precipitation, electroporation, or liposome mediated-, DEAE dextran mediated-, polycationic mediated-, or viral mediated transfection. For all standard techniques see, for example, Sambrook et al., Molecular Cloning: A Laboratory Manual, 4th ed., Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016). Recombinant host cells, in the present context, are those which have been genetically modified to contain an isolated DNA molecule, nucleic acid molecule or expression construct or vector of the invention. The DNA can be introduced by any means known to the art which are appropriate for the particular type of cell, including without limitation, transformation, lipofection, electroporation or viral mediated transduction. A DNA construct capable of enabling the expression of the chimeric protein of the invention can be easily prepared by the art-known techniques such as cloning, hybridization screening and Polymerase Chain Reaction (PCR). Standard techniques for cloning, DNA isolation, amplification and purification, for enzymatic reactions involving DNA ligase, DNA polymerase, restriction endonucleases and the like, and various separation techniques are those known and commonly employed by those skilled in the art. A number of standard techniques are described in Sambrook et al. (2012), Wu (ed.) (1993) and Ausubel et al. (2016). Representative host cells that may be used with the invention include, but are not limited to, bacterial cells, yeast cells, plant cells and animal cells. Bacterial host cells suitable for use with the invention include Escherichia spp. cells, Bacillus spp. cells, Streptomyces spp. cells, Erwinia spp. cells, Klebsiella spp. cells, Serratia spp. cells, Pseudomonas spp. cells, and Salmonella spp. cells. Animal host cells suitable for use with the invention include insect cells and mammalian cells (most particularly derived from Chinese hamster (e.g. CHO), and human cell lines, such as HeLa. Yeast host cells suitable for use with the invention include species within Saccharomyces, Schizosaccharomyces, Kluyveromyces, Pichia (e.g. Pichia pastoris), Hansenula (e.g. Hansenula polymorpha), Yarrowia, Schwaniomyces, Schizosaccharomyces, Zygosaccharomyces and the like. Saccharomyces cerevisiae, S. carlsbergensis and K. lactis are the most commonly used yeast hosts, and are convenient fungal hosts. The host cells may be provided in suspension or flask cultures, tissue cultures, organ cultures and the like. Alternatively, the host cells may also be transgenic animals.
[0073] "Higher" or "increased" as used herein refers to at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 1.5 fold, at least 2 fold, at least 3 fold, at least 5 fold or at least 10 fold higher quantity or an effect. In some embodiments, "higher" or "increased" refers to a statistically significant difference. "Predominantly" as used herein, means at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95% of quantity or an effect.
[0074] "Lower" or "decreased" as used herein is defined herein as a statistically significantly decreased, more particularly an at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 1.5 fold, at least 2 fold, at least 3 fold, at least 5 fold or at least 10 fold lower quantity or an effect. In some embodiments, "lower" or "decreased" refers to a statistically significant difference.
[0075] The term "moiety" as used herein, means a group of atoms within a molecule or a part of molecule that is responsible for characteristic chemical reactions of that molecule. In some embodiments, the term moiety may refer to a functional group. The term "pharmaceutically acceptable carrier" as used herein, means a carrier or excipient or vehiculum or diluent that at the dosages and concentrations employed are suitable and / or safe to be administered to a patient.
[0076] The term "small molecule", as used herein, refers to wide class of inorganic and organic compounds, which are of a lower molecular weight compared to biological macromolecules, such as proteins, polysaccharides, lipides or polynucleotides. Inorganic small molecule drugs can comprise, for example, 20 to 100 atoms and / or have a molecular mass of less than 1000 g / mol or 1 kilodalton (kDa). Organic small molecule compounds can be, for example, of the molecular mass up to 2.5 kDa. It should be understood that any class of compounds known to a skilled person as a small molecule, featuring any kind of physicochemical characteristic can be considered as a small molecule, without departing from the scope of the present invention.
[0077] Detailed description of embodiments
[0078] The present invention aims to solve the need to precise design of TMB polypeptides and nanopores comprising said TMP polypeptides. Some natural TMB polypeptides are known in the art but lacking the described for DNA and RNA sequencing. Current engineering approaches for nanopore sensors are limited to the known naturally occurring nanopores, i.e. channels such as CsgG, were shown to have only limited number of applications and / or are not suitable for molecular sensing techniques and thus provide sub-optimal starting points. In contrast, novel protein design which is described in various aspects of the present invention, was shown to be able to create a number new nanopores with any desired structural properties. Particularly, said TMB polypeptide design model allows for provision of said TMB polypeptides by an expression in a suitable host organism, such as, but not limited to a microorganism, preferably bacterial cell such as E. coli. Provided NMR and crystallographic characterization shows that the designs are stably folded with structures of the design models suitable for various applications, for example molecular sensing or a moiety detection.
[0079] According to the first aspect, the present invention concerns de novo designed TMB polypeptides with a 12 stranded p-barrel, the transmembrane beta barrel (TMB) polypeptide comprising the formula X1-X2- X3-X4-X5-X6-X7-X8-X9-X10-X11-X12-X13-X14-X15-X16-X17-X18-X19-X20-X21-X22-X23, wherein:
[0080] • XI comprises a beta strand, preferably said beta strand comprising or consisting of a sequence Z1-Z2-Z3-G-Z3-Z2-Z3-Z2-Y;
[0081] • X3 comprises a beta strand comprising or consisting of a sequence Z1-Z2-Z3-G-Z3-Y-Z3-Z2- Y-Z2-Z1; • X5 comprises a beta strand comprising or consisting of the sequence Z1-Z2-Z3-Z2-Z3-Z2-Y- Z2-W;
[0082] • X7 comprises a beta strand comprising or consisting of the sequence (N / Q)-Z2-Z3-Z2-Z3-G- Z3-Z2-Y-Z2-Z1;
[0083] • X9 comprises a beta strand comprising or consisting of the sequence Z1-Z2-Z3-Z2-Z3-Z2-Z3- Z2-Y;
[0084] • Xll comprises a beta strand comprising or consisting of the sequence Z1-Z2-Z3-Z2-Z3-Z2-Z3- Z2-Z3-Z2-Y-P-Z1;
[0085] • X13 comprises a beta strand comprising or consisting of the sequence Z1-Z2-Z3-Z2-P-G-Z3- Z2-Y-Z2-W;
[0086] • X15 comprises a beta strand comprising or consisting of the sequence (N / Q)-Z4-Z3-Z2-P-Y- Z3-Z2-Z3-Z2-Y-Z2-Z1;
[0087] • X17 comprises a beta strand comprising or consisting of the sequence Z1-Z2-Z3-Z2-Z3-Z2-Z3- Z2-Y;
[0088] • X19 comprises a beta strand comprising or consisting of the sequence Z1-Z2-Z3-Z2-Z3-G-Z3- Z2-Y-Z2-Z1;
[0089] • X21 comprises a beta strand comprising or consisting of the sequence Z1-Z2-Z3-Z2-Z3-Z2-Y- G-W;
[0090] • X23 comprises a beta strand comprising or consisting of the sequence (N / Q)-Z2-Z3-Z2-Z3- Z2-Z3-Z2-Y-Z2-Z1-Z5;
[0091] • X2, X6, X10, X14, X18 and X22 each comprise a flexible loop, preferably the flexible loop comprising at least 3 and / or at most 10 amino acids;
[0092] • X4, X8, X12, X16 and X20 each are a beta turn, said beta turn preferably comprising or consisting of 3 to 5 amino acids; wherein each one of XI, X3, X5, X7, X9, Xll, X13, X15, W17, X19 and X21 comprises a sequence of amino acids that alternate in their relative position within the TMB between the internal side of the lumen formed by the TMB and the surface region of the TMB, and said sequence starting and / or ending with an amino acid located in the surface region of said TMB polypeptide; and wherein X23 comprises a sequence of amino acids that alternate in their relative position within the TMB between the internal side of the lumen formed by the TMB and the surface region of the TMB, and said sequence starting and / or ending with the amino acid in the surface region of said polypeptide, except for the two amino acids on the C-terminus indicated as Z1 and Z5 which are located in the surface region; wherein said Z1 amino acid is chosen from the group consisting of V, L, A and F, and said amino acid is located in the surface region of the TMB polypeptide; wherein said Z2 amino acid is chosen from the group A, D, E, F, H, K, L, M, N, Q, R, S, T, V and W, and said amino acid is located in the pore region; wherein said Z3 amino acid is chosen from the group consisting of V, L, A, G, S, T and F, and said amino acid is located in the surface region; and wherein said Z4 amino acid is chosen from the group consisting of A, D, E, H, K, L, M, N, Q, R, S, T and V, and said amino acid is located in the pore region; wherein Z5 is a single amino acid for a bulge in the TMB structure.
[0093] The X elements defines in the TMB polypeptide according to the first aspect and any embodiment thereof represent an element of the polypeptide defined by an amino acid sequence, said element comprising at least two amino acids and / or amino acid residues and / or said element having a secondary structure, wherein each of the "X" elements is connected or linked to other elements through peptide bonds.
[0094] The term "Z" as used herein, relates to an amino acid and / or a residue thereof. Said amino acid and / or residue can be, preferably are connected or linked to other amino acids through peptide bonds.
[0095] The TMB polypeptide according to the first aspect can be expressed in a host cell, which can allow for a more economic and / or easier industrial production.
[0096] Said TMB polypeptide is of an amino acid sequence allowing to be recombinantly expressed in a host cell to fold into a single chain TMB with a stability to successfully integrate and fold into a lipid membranes and / or micelles. The newly designed amino acid blueprint scaffold provides any of said amino acid structures of 12-stranded TMB polypeptides according to the first aspect of the invention. The new TMB designs provided in the first aspect are 12 stranded TMB polypeptides of high complexity, which polypeptides still can be expressed in a host cell, preferably a bacterial host cell such as E. coli and are capable of assembling onto lipid-lipid interfaces, such as lipid bilayer membranes, or detergent micelles. Said TMB polypeptide can be seen as a new complex beta-barrel structure of a particular and / or favoured design suitable for provision of nanopore channels of high stability and / or rigidity and / or conducting properties. The TMB polypeptide according the first aspect of the invention shows preferably less "noise" and / or "noisy" recordings than naturally occurring TMB polypeptides known in the art. The "noisy" electrophysiology recording can be described as undesired interference in the signal readout which can lead to challenging data interpretation when the TMB polypeptides and the pores thereof are used for sensing applications. For example, the TMB polypeptide can be designed in a such way that bulky and / or large elements of said polypeptide are preferably flexible and / or moved or placed onto a position which allows minimal interference of the lumen and / or surface of the TMB polypeptide with for example surrounding structures and / or molecules.
[0097] The TMB polypeptide according to the first aspect, thus, can allow forming a nanopore comprising a circularly closed beta sheet, providing a rigid scaffolds for the transport of molecules across various membranes, such as cellular or organelle membranes (of bacterial, mitochondria and chloroplast).
[0098] The TMB polypeptide according to the first aspect of the present invention is preferably of the exterior which is nonpolar to some extent, preferably largely nonpolar, more preferably predominantly non-polar for membrane insertion, while the interior is polar to some extents, preferably largely polar, more preferably predominantly polar to support a solvated conducting channel. Furthermore, the structure of TMB polypeptide is preferably specified in the vast majority by short-range interactions between residues located on adjacent strands since there is no close-packed core. Finally, the amphipathic p- strands are highly aggregation prone prior to p-barrel assembly, and hence the design must strongly favour intra-chain rather than inter-chain interactions during folding. Said structural characteristics of the TMB polypeptides can allow design of stable monomeric channels with tuneable pore characteristics, size, and / or channel conductance.
[0099] In a preferred embodiment, Z5 is a single amino acid for a bulge structural feature being chosen from the group consisting of S, T or D. Said bulge can contribute to the better expression and / or more efficient folding of the TMB polypeptides of the invention and / or stability of the structure.
[0100] In an embodiment, X4, X8, X12, X16 and X20 are beta turns comprising a bulge and two to four amino acid turn residues. The beta turns in the TMB polypeptide of the invention are designed to be of a structure allowing for translocation though the bilayer. Thus said beta turns preferably are destabilized to allow folding and assembly into the membrane. Thus said beta turns allow for desired folding and assembly in the lipid membrane bilayers and / or detergent micelles, thereby forming of the nanopore comprising said TMB polypeptide.
[0101] In a preferred embodiment, any one of said beta turn X4, X8, X12, X16 or X20 is a beta turn comprising or consisting of Z5-Zx-Zx+i, Z5-ZX-ZX+I-ZX+2or Z5-ZX-ZX+I-ZX+2-ZX+3 wherein Zx, Zx+i, Zx+2, Zx+3are single turn amino acids.
[0102] In a preferred embodiment the TMB polypeptide comprises the beta turns consisting of Z5-Zx-Zx+ii.e. said element consists of a bulge and two amino acid residues. Said design allows for improved structure flexibility and / or local "frustrations" or "destabilization" allowing for more efficient expression and / or folding.
[0103] In a preferred embodiment, any one of said beta turn X4, X8, X12, X16 or X20 is a beta turn comprising or consisting of Z5-Zx-Zx+i, wherein Zxis preferably P and Zx+iis preferably chosen from the group consisting of E, D, N or Y. In a preferred embodiment, any one of X4, X8, X12, X16 or X20 elements, said elements being beta turns consisting of the sequence (S / T / D)-P-(E / D / N / Y). Said beta turns contribute to better and / or desired folding and expression of the TMB polypeptide of the invention.
[0104] The TMB polypeptide of the invention comprises six flexible loop, i.e. comprises elements X2, X6, X10, X14, X18 and X22 each comprising a flexible loop. Each one of said loops preferably comprises at least 3 and / or at most 10 amino acids, preferably at most 5 amino acids. Said flexible loops allow for obtaining more "silent" nanopores, thus providing for pore structures with less "noise" as compared to pores with less defined loops or with different structural features positioned between the sequential beta-strands of the TMB, as the flexibility can contribute to moving of said loop away from the lumen of the pore. Thus, said flexible loops preferably reduce the interference with nanopore conductivity and / or the transport through the lumen of said nanopore, compared to pores known in the art.
[0105] In a specific embodiment said TMB polypeptide and / or pore are provided as an engineered TMB polypeptide, preferably wherein said engineering is present as a circular permutated TMB polypeptide variant, for which the design and construction is described herein. More preferably, said engineered TMB polypeptide relates to a circular permutated TMB polypeptide wherein the N- and C-terminus of the engineered TMB polypeptide are derived from a position in a loop (X2, X6, X10, X14, X18, X22) or a turn (X4, X8, X12, X16, X20) of the original TMB polypeptide provided herein. Alternatively, said engineered TMB polypeptide and / or pore comprise a circular permutated TMB polypeptide allowing for specific uses in a modified form.
[0106] Preferably the TMB polypeptide described here is designed with the shortest loops that, do not compromise protein folding. Such short flexible loop provides quiet TMB platforms and / or designs which can be readily used for sensor engineering. The particularly preferred TMB polypeptide designs according to the first aspect can for example, allow up to 2 hours of recording which would enable the detection and quantification of molecules at low concentration.
[0107] In a particular embodiment X2 is the loop comprising or consisting of the sequence N-T-D-N-T.
[0108] In a particular embodiments, any one of X6, X14 or X22 is the loop comprising or consisting of the sequence N-N-S-S-L.
[0109] In a particular embodiment, any one of X10 or X18 is the loop comprising or consisting of the sequence N-T-D-N-T.
[0110] In particular embodiments, the TMB polypeptide has a shear number (S) of 14. The shear number is a discrete parameter which is a characteristic of a beta barrel (Vorobieva et al., 2021, Science 371, eabc8182), which characterizes the transmembrane span and the connectivity between p-strands. In this particular embodiments, the shear number of 14 allows for an optimal transmembrane span and the connectivity between p-strands of the TMB polypeptide. In a preferred embodiment, the connection between the beta strands XI and X23 comprises a canonical antiparallel seam and / or antiparallel arrangement between the strands. Said antiparallel arrangements of the beta sheets allows for better folding characteristics and / or stability of the TMB polypeptides of the invention.
[0111] In an embodiment, the TMB polypeptide further comprises a linker XO at the N-terminal domain of said peptide. Preferably, said linker comprises or consists of the sequence Z6-Z7-Z8-G-(S / T / D / N), wherein underlined amino acid is invariant, and wherein Z6, Z7 and Z8 is any amino acid except C.
[0112] In a particularly preferred embodiment, the TMB polypeptide comprises the amino acid sequence of at least 50 % of sequence identity to SEQ ID No. 3 or SEQ ID NO. 9, or the amino acid sequence of at least 60 % of sequence identity, the amino acid sequence of at least 70 % of sequence identity, the amino acid sequence of at least 80 % of sequence identity, the amino acid sequence of at least 90 % of sequence identity, the amino acid sequence of at least 95 % of sequence identity, or the amino acid sequence of at least 99 % of sequence identity to SEQ ID No. 3 or SEQ ID NO. 9. In further preferred embodiments, said TMB comprises or consists of the polypeptide according to any one of SEQ ID NOs:l-9. More particularly, any of the said TMB polypeptides allows for preferred expression and / or folding characteristics as described herein.
[0113] In a particularly preferred embodiment, said polypeptide consists of SEQ ID No. 3 or SEQ ID NO. 9. Said TMB polypeptide allows for preferred expression and / or folding characteristics. Furthermore, the TMB polypeptide consisting of said sequences showed particularly good pore forming characteristics. This particularly preferred embodiment of TMB polypeptide can be folded into a nanopore which provides, for example up to 0.5, 1, 1.5, 2, 2.5 or 3 hours of recording, preferably at least 2 hours of recording. Such long stability and / or recording time would enable the detection and quantification of molecules at low concentration. Thus said TMB polypeptide may be particularly suitable molecular sensing applications where high sensitivity is desired.
[0114] According to one embodiment of the first aspect, the TMP polypeptide comprises a structure corresponding to a blueprint or structural scaffold shown in Table 1 in Example 1.
[0115] In a particularly preferred embodiment, the invention relates to a nucleic acid encoding the TMB polypeptide according to any one of the embodiments of the first aspect.
[0116] In a particularly preferred embodiment, the invention relates to an expression vector comprising the nucleic acid nucleic acid encoding the TMB polypeptide according to any one of the embodiments of the first aspect, said expression vector being preferably operably linked to a control sequence.
[0117] In a preferred embodiment, the present invention relates to a cell, preferably a recombinant host cell comprising the nucleic acid encoding the TMB polypeptide according to any embodiment of the first aspect, or the expression vector comprising said nucleic acid. Some non-limiting examples are a bacterial or a yeast or a fungal cell, preferably the bacterial cell. In a particularly preferred embodiment said host cell is of E. coli.
[0118] The invention according to the first aspect provides said the nucleic acid nucleic acid encoding the TMB polypeptide of the invention, the expression vector comprising the nucleic acid nucleic acid encoding the TMB polypeptide and / or the host cell comprising said expression vector and / or said nucleic acid, which all allow an easier and / or more efficient and / or up-scaled and / or industrial production of said TMB polypeptide of the invention. The specific secondary and tertiary structure of the TMB polypeptide of the invention, for example a balance between the optimization of tertiary structure energy and negative design (introduction of locally frustrated residues) to disfavor premature p-strand formation before membrane insertion, allows for the expression of the 12-stranded, complex TMB nanopores according to any embodiment of the first aspect in a suitable host cell, or even more preferably, in a specific body and / or part of the host cell. In one embodiment, the TMB polypeptide according to the first aspect is expressed in E coli inclusion bodies
[0119] In a particularly preferred embodiment, present invention relates to a micelle and / or a lipid membrane and / or an organelle membrane comprising the TMB polypeptide according to any embodiment of the first aspect, the nucleic acid encoding said polypeptide, the expression vector comprising said nucleic acid or a host cell comprising said nucleic acid or said vector. Said micelle and / or lipid membrane or an organelle can be used as a useful tool and / or a carrier for sensing and / or metabolite detecting and / or protein fingerprinting and / or custom engineering.
[0120] TMB polypeptide and / or the design thereof according to the first aspect and any embodiments thereto, enables with atomic level precision, as highlighted by the close agreement between the experimentally determined crystal and NMR structures a new stable beta barrel structures, which can be specified by strategic placement of glycine residues at which bending takes place to reduce strain. The obtained TMB polypeptide designs are stable and of the favourable folding, despite the inversion of the hydrophobicpolar core compared to globular proteins, and the almost entirely local nature of the side-chain interactions.
[0121] In a second aspect, the present invention relates to a nanopore comprising a lumen and a wall, wherein said wall comprises the TMB polypeptide, or the nanopore is obtainable by nucleic acid encoding the TMB polypeptide, the expression vector comprising said nucleic acid or the host cell according to any of the embodiments of the first aspect of the invention. The TMB polypeptide preferably allows for assembly in detergent and / or lipid vesicles, thereby forming a nanopore preferably comprising a single beta sheet. In comparison with previously known oligomeric protein nanopores, the nanopores according to the second aspect have the advantage that they can be built from a single chain. Such a strategy can enable controlled assembly of, for example, monodisperse nanopores without alternative oligomeric states and much greater control over the shape and specific surface properties of the transmembrane channel. While nanopore designs from oligomeric p-hairpins require lipid nanoparticles for solubility and assemble under electric current at lipid-lipid interfaces, the monomeric TMB design fold efficiently into detergent micelles and lipid vesicles. Thus, the nanopores according to the second aspect are a stable and / or more rigid and / or more silent and / or better controlled alternative to the known TMB pores. The stability of the TMB polypeptide designs of the invention enables their spontaneous insertion into planar lipid membranes following dilution out of detergent micelles, opening the door to the use of synthetic transmembrane nanopores in commercial flowcells developed for singlemolecule sensing and sequencing. Known monomeric integral TMBs such as OmpG have been applied for the sending of a range of metabolically and therapeutically-relevant proteins by displaying analyterecognition motifs or biotin-bound antibodies in the solvent-exposed loops. However, the long disordered loops of naturally-occurring TMBs result in noisy reading and the engineering of quiet pores with shorter or mutated loops has been a long-standing problem. The TMBs described here were designed with the preferably short loops that preferably do not compromise protein folding. They provide quiet TMB polypeptides and / or nanopores comprising said TMB polypeptides readily for sensor engineering. The nanopores according to the second aspect thus provide for a new and / or more efficient and / or sensitive option for development of various metabolite development techniques, compared to the known technologies.
[0122] In a preferred embodiment, the nanopore allows for at least 1, 5, 10, 15, 20, 25, 30 min and at most 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 min of recording of a current, preferably of an electrical current or any other electro- type of measurements. This can be particularly advantageous where the concentration of the analyte is low and the high sensitivity is requires.
[0123] In a particularly preferred embodiment, the nanopore comprising the TMB polypeptide of SEQ ID No. 3 provides up to 10 min, 30 min, 1, 1.5, 2, 2.5 or 3 hours of recording which would enable the detection and quantification of molecules at low concentration. Preferably the nanopore comprising the TMB polypeptide of SEQ ID No. 3 can allow for at least 1.5 h and at most 2.5 h of the measurement recording. In a preferred embodiment, the nanopore has the lumen is of a diameter of 21.0 to 24.6 A, preferably it has the lumen of 21.1, 21.2, 21.3, 21.4, 21.5, 21.6, 21.7, 21.8, 21.9, 22.0, 22.1, 22.2, 22.3, 22.4, 22.5, 22.6, 22.7, 22.8, 22.9, 23.0, 23.1, 23.2, 23.3, 23.4, 23.5, 23.6, 23.7, 23.8, 23.9, 24.0, 24.1, 24.2, 24.3, 24.4? 24.5 or 24.6 A. In a particularly preferred embodiment, said nanopore has a lumen of 22.0 to 23.6 A. The nanopores of the invention can have any distinct conductance's that correlate with their pore diameter, for example, ranging from 210 pS (~0.5 nm pore diameter) to 430 pS ("'1.1 nm pore diameter). The capability to generate on demand p-barrel pores of defined geometry opens up fundamentally new opportunities for custom engineering of sequencing and sensing technologies.
[0124] In a preferred embodiment, the nanopore has the lumen is of a substantially square shape. This lumen was favourable because of improved conductivity.
[0125] In a particularly preferred embodiment, said nanopore comprises the lumen further comprises a detectable moiety and / or a detector polypeptide. Said moiety and / or detector polypeptide could allow specifical application in the field of analytics, preferably by allowing the adjustability of the pore to a particular analyte. The nanopore of the invention can allow understanding chemical properties of a nanopore and link it to directed detection of an analyte in the pore lumen. The nanopores of the invention are custom pores, of the design which could be adjusted, which opens up exciting new directions for sequencing and sensing since the pore size and the sidechains lining the pore can be designed specifically for the desired application. Unlike native pores, which are of a defined structure, there is no limit on the number of distinct designed pores that can be generated by using the TMB polypeptide design of the invention. Such design can lead to pore geometry and chemistry to be custom built to be optimal for applications ranging from detection and selective transport of a wide range of molecules of interest to biopolymer sequencing.
[0126] In a particular embodiment, said detectable moiety is selected from the group consisting of a pH- detectable moiety, electrical conductivity moiety, chromophore, fluorophore and chemically detectable moiety.
[0127] In a preferred embodiment, the nanopore further comprises the detector polypeptides, wherein said polypeptide is an ion-binding polypeptide, a pH responsive polypeptide and / or a nucleic acid binding polypeptide.
[0128] In another aspect, the present invention relates to use of the TMB polypeptide according to any embodiment of the first aspect, the nucleic acid encoding said peptide, the expression vector comprising said nucleic acid, the host cell comprising said vector and / or the nucleic acid, and / or nanopore according to the second aspect for molecular sensing and / or receiving and / or capturing of small-molecules or of proteins, and / or drug delivery and / or as ion, water or small-molecule permeable transmembrane channels.
[0129] In yet another independent aspect, the present invention relates to a pharmaceutical composition, comprising:
[0130] • The TMB polypeptide according to any embodiment of the first aspect, the nucleic acid encoding said polypeptide, the expression vector comprising said nucleic acid, the host cell according comprising said vector and / or said nucleic acid, and / or the nanopore according to any embodiment of the second aspect;
[0131] • A carrier, preferably pharmaceutically acceptable carrier.
[0132] The present invention provides pharmaceutical compositions, comprising one or more TMB polypeptides, nucleic acids, expression vectors, and / or host cells according to any embodiment of the first aspect and / or the nanopores according to any embodiment of the second aspect of the disclosure and a pharmaceutically acceptable carrier. The pharmaceutical compositions of the disclosure can be used, for example, for the same purposes as explained for the TMB polypeptide, the nucleic acid and / or a vector and / or the host cell, or the nanopore of the invention. The pharmaceutical composition may comprise additional ingredients, such as (a) a lyoprotectant; (b) a surfactant; (c) a bulking agent; (d) a tonicity adjusting agent; (e) a stabilizer; (f) a preservative and / or (g) a buffer.
[0133] In another aspect, the present invention relates to a computer-implemented method for obtaining the TMB polypeptide according to any embodiment of the first aspect:
[0134] • Generating backbones corresponding to the TMB polypeptide according to any one embodiment of the first aspect;
[0135] • Assembling the generated backbones using the Rosetta BlueprintBDR.
[0136] In another independent aspect, the present invention relates to a method for producing the TMB polypeptide according to any embodiment of the first aspect, the method comprising:
[0137] • culturing a host cell under conditions conducive to the expression of the polypeptide;
[0138] • recovering the expressed polypeptide; and
[0139] • optionally, reconstitute the polypeptide in vitro in detergent micelles or lipids.
[0140] EXAMPLES
[0141] Example 1: Computational design of TMB polypeptide backbones
[0142] TMB backbones accommodating water-accessible pores were built starting from the principles elucidated during the design of 8-stranded TMBs lacking pores (Vorobieva et al., 2021, Science, 371, eabc8182). To modulate the size of the pore, the number of p-strands was increased to 12 strands, while keeping the transmembrane span and the connectivity between p-strands the shear number constant (Liu, 1998, J. Mol. Biol. 275, 541-545, Principles determining the structures of p-sheet barrels in proteins, 1994, J. Mol. Biol. 236, 1369-1381)). This resulted in an increase of the average p-barrel diameter to about 22.8 A corresponding to 12-stranded barrel (Figure 1A). By comparison to 8-stranded TMBs, the diameters of the 12-stranded p-barrels do not allow long-range sidechain contacts across the pores and the structural properties of the pores (P-strand pairing, p-barrel shape) must be locally encoded. Naturally-occurring TMBs typically feature long, disordered, loops on one side of the barrel (Franklin et al., 2018, J. Mol. Biol. 430, 3251-3256), which results in noisy electrophysiology recording and challenging data interpretation when the pores are used for sensing applications. To design quiet pores and reduce the noise, we connected the p-strands on both sides of the barrels with 2- and 3- residues (Figure S2B) loops, also referred to as -hairpins, the shortest loops we have previously found to support TMB folding ( Vorobieva et al., 2021, Science, 371, eabc8182). The first-generation backbones corresponding to these designs were assembled with the Rosetta BlueprintBDR application (Koga et al., 2012, Nature. 491, 222-227) and had similar cylindrical shapes. Such cylindrical p-sheet configurations are strained (Dou et al., 2018, Nature. 561, 485-491, Salemme, 1981, J. Mol. Biol. 146, 143-156) due to repulsion between side-chains packing the barrel lumen (Figure 1C). Glycine kinks (glycine residues in extended positive-phi conformation) were introduced into the blueprint to relieve the strain and to bend the p-strands to form corners in the p-barrel cross-section. We generated four blueprints with the same topology but different glycine kink distributions to design 12-stranded p-barrel backbones with square, cross-sections (Figure 2A). A single glycine kink was used in corners of an angle of >= 90° and several adjacent and / or stacked kinks were placed to form corners of < 90° (Figure IB). Sequence-agnostic TMB backbones incorporating these constraints were assembled in silica and had the shapes expected based on the placement of the glycine kinks (Figure ID).
[0143] A challenge for TMB design is to balance the optimization of the folded p-barrel state in the membrane with delayed folding in water to reduce misfolding and aggregation that would prevent successful integration into a membrane bilayer (Vorobieva et al., 2021, Science, 371, eabc8182, Danoff and Fleming, 2017, Biochemistry. 56, 47-60). For the 8-stranded TMBs, this was achieved by incorporating local secondary-structure frustration to reduce premature formation of aggregation-prone p-strands prior to full barrel assembly: hydrophobic amino acids were designed into the water-accessible pore to disrupt the hydrophobic-polar amino acid alternation pattern characteristic of amphipathic p-sheets. To test whether such balancing is necessary for larger p-barrel designs that need to have water-accessible (and hence more polar) channels, we first designed "optimal" 12-strand TMBs with only polar and charged amino acids facing the pore, but all of such designs failed to express in E. coli. We therefore set out to design larger TMB nanopores incorporating local secondary structure frustration. In the water- accessible pore, networks of polar residues were designed around the canonical TMB folding motif Tyr- Gly-Asp / Glu (Vorobieva et al., 2021, Science, 371, eabc8182, Michalik et al., 2017, PLoS One. 12, e0182016) to optimize strong local p-register defining interactions while alternating with patches of hydrophobic and small, disorder promoting, residues. To compensate for the more hydrophilic pores in the larger TMBs, we further reduced the p-sheet propensity by replacing a small number of p-branched residues with Ser and Thr amino acids on the lipid-exposed surface. Despite the fact that it seems perhaps counterintuitive to expose hydroxyl groups to the lipid environment, a model was built which indicated that Ser and Thr could form a hydrogen bond to the p-strand backbone when placed in close proximity with a glycine kink, effectively mimicking the backbone-water hydrogen bonds observed in strongly bent p-strands of water-soluble p-barrels (Figure IE).
[0144] During combinatorial design of sequences for p-barrels of different size, it was found that the frequency of incorporation of each amino acid type strongly depended on the curvature of the p-sheet. During our design, the Rosetta solvation and reference energies were adjusted ( Alford et al., 2017, J. Chem. Theory Comput. 13, 3031-3048) to achieve the desired balance of frustrated and energetically-favorable contacts. Following several iterations of combinatorial sequence design and structure relaxation, designs were selected based on hydrogen bond network descriptors, secondary structure and aggregation propensities (Fernandez-Escamilla et al., 2004, Nat. Biotechnol. 22, 1302-1306).
[0145] Table 1. Blueprint of the 12-stranded TMB polypeptides.
[0146] I A NOTAA C
[0147] 3 A NOTAA C
[0148] 4 A NOTAA C
[0149] 5 A PIKAA G
[0150] 6 A PIKAA STDN
[0151] 7 A PIKAA VLAF #surface
[0152] 8 A PIKAA ADEFHKLMNQRSTVW #pore
[0153] 9 A PIKAA VLAGSTF #surface
[0154] 10 A PIKAA G #pore
[0155] II A PIKAA VLAGSTF #surface
[0156] 12 A PIKAA ADEFHKLMNQRSTVW #pore
[0157] 13 A PIKAA VLAGSTF #surface
[0158] 14 A PIKAA ADEFHKLMNQRSTVW #pore
[0159] 15 A PIKAA Y #surface 16 A PIKAA N #loop
[0160] 17 A PIKAA T #loop
[0161] 18 A PIKAA D #loop
[0162] 19 A PIKAA N #loop
[0163] 20 A PIKAA T #loop
[0164] 21 A PIKAA VLAF #surface
[0165] 22 A PIKAA ADEFHKLMNQRSTVW #pore
[0166] 23 A PIKAA VLAGSTF #surface
[0167] 24 A PIKAA G #pore
[0168] 25 A PIKAA VLAGSTF #surface
[0169] 26 A PIKAA Y #pore 1 A PIKAA VLAGSTF #surface
[0170] 28 A PIKAA ADEFHKLMNQRSTVW #pore
[0171] 29 A PIKAA Y #surface
[0172] 30 A PIKAA ADEFHKLMNQRSTVW #pore
[0173] 31 A PIKAA VLAF #surface
[0174] 32 A PIKAA STD #bulge
[0175] 33 A PIKAA P #turn
[0176] 34 A PIKAA EDNY #turn
[0177] 35 A PIKAA VLAF ftsurface
[0178] 36 A PIKAA ADEFHKLMNQRSTVW #pore
[0179] 37 A PIKAA VLAGSTF #surface
[0180] 38 A PIKAA ADEFHKLMNQRSTVW #pore
[0181] 39 A PIKAA VLAGSTF #surface
[0182] 40 A PIKAA ADEFHKLMNQRSTVW #pore
[0183] 41 A PIKAA Y #surface
[0184] 42 A PIKAA ADEFHKLMNQRSTVW #pore 43 A PIKAA W#surface
[0185] 44 A PIKAA N #loop
[0186] 45 A PIKAA N #loop
[0187] 46 APIKAAS#loop
[0188] 47 APIKAAS#loop
[0189] 48 APIKAAL#loop
[0190] 49 A PIKAA NQ#surface
[0191] 50 A PIKAA ADEFHKLMNQRSTVW #pore
[0192] 51 A PIKAA VLAGSTF #surface
[0193] 52 A PIKAA ADEFHKLMNQRSTVW #pore
[0194] 53 A PIKAA VLAGSTF #surface
[0195] 54 A PIKAA G #pore
[0196] 55 A PIKAA VLAGSTF #surface
[0197] 56 A PIKAA ADEFHKLMNQRSTVW #pore
[0198] 57 A PIKAA Y #surface
[0199] 58 A PIKAA ADEFHKLMNQRSTVW #pore
[0200] 59 APIKAAVLAF #surface
[0201] 60 A PIKAA STD #bulge
[0202] 61 A PIKAA P #turn
[0203] 62 APIKAAEDNY #turn
[0204] 63 APIKAAVLAF #surface
[0205] 64 A PIKAA ADEFHKLMNQRSTVW #pore
[0206] 65 A PIKAA VLAGSTF #surface
[0207] 66 A PIKAA ADEFHKLMNQRSTVW #pore
[0208] 67 A PIKAA VLAGSTF #surface
[0209] 68 A PIKAA ADEFHKLMNQRSTVW #pore 69 A PIKAA VLAGSTF #surface
[0210] 70 A PIKAA ADEFHKLMNQRSTVW #pore
[0211] 71 A PIKAA Y #surface
[0212] 72 A PIKAA N #loop
[0213] 73 APIKAAT#loop
[0214] 74 A PIKAA D #loop
[0215] 75 A PIKAA N #loop
[0216] 76 APIKAAT#loop
[0217] 77 APIKAAVLAF #surface
[0218] 78 A PIKAA ADEFHKLMNQRSTVW #pore
[0219] 79 A PIKAA VLAGSTF #surface
[0220] 80 A PIKAA ADEFHKLMNQRSTVW #pore
[0221] 81 A PIKAA VLAGSTF #surface
[0222] 82 A PIKAA ADEFHKLMNQRSTVW #pore
[0223] 83 A PIKAA VLAGSTF #surface
[0224] 84 A PIKAA ADEFHKLMNQRSTVW #pore
[0225] 85 A PIKAA VLAGSTF #surface
[0226] 86 A PIKAA ADEFHKLMNQRSTVW #pore
[0227] 87 A PIKAA Y #surface
[0228] 88 APIKAAP#pore
[0229] 89 APIKAAVLAF #surface
[0230] 90 A PIKAA STD #bulge
[0231] 91 APIKAAP#turn
[0232] 92 APIKAAEDNY #turn
[0233] 93 A PIKAA VLAF #surface
[0234] 94 A PIKAA ADEFHKLMNQRSTVW #pore 95 A PIKAA VLAGSTF #surface
[0235] 96 A PIKAA ADEFHKLMNQRSTVW #pore
[0236] 97 A PIKAA P #surface
[0237] 98 A PIKAA G #pore
[0238] 99 A PIKAA VLAGSTF #surface
[0239] 100 A PIKAA ADEFHKLMNQRSTVW #pore
[0240] 101 A PIKAA Y #surface
[0241] 102 A PIKAA ADEFHKLMNQRSTVW #pore
[0242] 103 A PIKAA W #surface
[0243] 104 A PIKAA N #loop
[0244] 105 A PIKAA N #loop
[0245] 106 A PIKAA S #loop
[0246] 107 A PIKAA S #loop
[0247] 108 A PIKAA L #loop
[0248] 109 A PIKAA NQ #surface
[0249] 110 A PIKAA ADEHKLMNQRSTV #pore
[0250] 111 A PIKAA VLAGSTF #surface
[0251] 112 A PIKAA ADEFHKLMNQRSTVW #pore
[0252] 113 A PIKAA P #surface
[0253] 114 A PIKAA Y #pore
[0254] 115 A PIKAA VLAGSTF #surface
[0255] 116 A PIKAA ADEFHKLMNQRSTVW #pore
[0256] 117 A PIKAA VLAGSTF #surface
[0257] 118 A PIKAA ADEFHKLMNQRSTVW #pore
[0258] 119 A PIKAA Y #surface
[0259] 120 A PIKAA ADEFHKLMNQRSTVW #pore
[0260] 121 A PIKAA VLAF #surface 122 A PIKAA STD #bulge
[0261] 123 A PIKAA P #turn
[0262] 124 A PIKAA EDNY #turn
[0263] 125 A PIKAA VLAF #surface
[0264] 126 A PIKAA ADEFHKLMNQRSTVW #pore
[0265] 127 A PIKAA VLAGSTF #surface
[0266] 128 A PIKAA ADEFHKLMNQRSTVW #pore
[0267] 129 A PIKAA VLAGSTF #surface
[0268] 130 A PIKAA ADEFHKLMNQRSTVW #pore
[0269] 131 A PIKAA VLAGSTF #surface
[0270] 132 A PIKAA ADEFHKLMNQRSTVW #pore
[0271] 133 A PIKAA Y #surface
[0272] 134 A PIKAA N #loop
[0273] 135 A PIKAA T #loop
[0274] 136 A PIKAA D #loop
[0275] 137 A PIKAA N #loop
[0276] 138 A PIKAA T #loop
[0277] 139 A PIKAA VLAF #surface
[0278] 140 A PIKAA ADEFHKLMNQRSTVW #pore
[0279] 141 A PIKAA VLAGSTF #surface
[0280] 142 A PIKAA ADEFHKLMNQRSTVW #pore
[0281] 143 A PIKAA VLAGSTF #surface
[0282] 144 A PIKAA G #pore
[0283] 145 A PIKAA VLAGSTF #surface
[0284] 146 A PIKAA ADEFHKLMNQRSTVW #pore
[0285] 147 A PIKAA Y #surface
[0286] 148 A PIKAA ADEFHKLMNQRSTVW #pore 149 A PIKAA VLAF #surface
[0287] 150 A PIKAA STD #bulge
[0288] 151 A PIKAA P #turn
[0289] 152 A PIKAA EDNY #turn
[0290] 153 A PIKAA VLAF #surface
[0291] 154 A PIKAA ADEFHKLMNQRSTVW #pore
[0292] 155 A PIKAA VLAGSTF #surface
[0293] 156 A PIKAA ADEFHKLMNQRSTVW #pore
[0294] 157 A PIKAA VLAGSTF #surface
[0295] 158 A PIKAA ADEFHKLMNQRSTVW #pore
[0296] 159 A PIKAA Y #surface
[0297] 160 A PIKAA G #pore
[0298] 161 A PIKAA W #surface
[0299] 162 A PIKAA N #loop
[0300] 163 A PIKAA N #loop
[0301] 164 A PIKAA S ttloop
[0302] 165 A PIKAA S #loop
[0303] 166 A PIKAA L #loop
[0304] 167 A PIKAA NQ #surface
[0305] 168 A PIKAA ADEFHKLMNQRSTVW #pore
[0306] 169 A PIKAA VLAGSTF #surface
[0307] 170 A PIKAA ADEFHKLMNQRSTVW #pore
[0308] 171 A PIKAA VLAGSTF #surface
[0309] 172 A PIKAA ADEFHKLMNQRSTVW #pore
[0310] 173 A PIKAA VLAGSTF #surface
[0311] 174 A PIKAA ADEFHKLMNQRSTVW #pore 175 A PIKAA Y #surface
[0312] 176 A PIKAA ADEFHKLMNQRSTVW #pore
[0313] 177 A PIKAA VLAF #surface
[0314] 178 A PIKAA STD #surface Example 2. Recombinant production of 12-stranded TMB polypeptides.
[0315] It was previously found that AlphaFold2 (Jumper et al., 2021, Nature. 596, 583-589) could accurately predict the structures of designed TM Bs even in the absence of evolution information (from a single sequence input and without a multiple sequence alignment (Mirdita et al., 2022, Nat. Methods. 19, 679- 682) when weak 3D contacts contained in the sequence were amplified by 48 rounds of molecular model recycling through the prediction network, and that the confidence assigned to the model (pIDDT) was a good discriminator of the sequences with higher probability of experimentally folding (Hermosilla et al., 2023, BioRxiv p. 2023.06.06543955). Therefore, 9 sequence designs (shown in Table 2) were selected from the blueprint (Table 1), for which AlphaFold2 predicted high-confidence structures closely matching the design models. Table 2. Sequences of the selected designs for TMB polypeptide based on a blueprint.
[0316] In order to verify that the 'de novo designed' sequences according to the proposed scaffold of Table 1 are novel in view of known protein sequences, we performed a BLASTP analysis on databases containing all non-redundant GenBank CDS translations+PDB+SwissProt+PIR+PRF excluding environmental samples from WGS projects. This resulted in a list of proteins with the highest percent identity as compared to any one of SEQ ID Nos: 1-9 within a range of 44-51%, leading to the conclusion that the designed TMB blueprint only provides for novel polypeptide sequences.
[0317] Moreover, the percent identity matrix was build comparing the 9 sequences as used further herein, using
[0318] Clustal 12.1, with the matrix shown here (Table 3), indicating that the blueprint results in protein sequence with at least 50 % identity
[0319] Table 3. Percent identity matrix.
[0320] 1: TMB12sol09_8 100.00 59.44 53.33 57 .78 61.11 58.89 57.78 61.67 54.44
[0321] 2: TMB12sol09_7 59.44 100.00 54.44 60 .56 58.89 53.33 61.67 58.33 61.67
[0322] 3: TMB12sol09_3 53.33 54.44 100.00 60 .56 60.00 55.56 55.00 61.67 60.00
[0323] 4 : TMB12sol09_0 57.78 60.56 60.56 1 00 .00 63.89 58.89 57.78 61.11 62.22
[0324] 5: TMB12sol09_9 61.11 58.89 60.00 63 .89 100.00 58.89 56.11 62.78 61.67
[0325] 6: TMB12sol09_4 58.89 53.33 55.56 58 .89 58.89 100.00 62.22 61.67 58.89
[0326] 7 : TMB12sol09_6 57.78 61.67 55.00 57 .78 56.11 62.22 100.00 63.33 58.33
[0327] 8: TMB12sol09_2 61.67 58.33 61.67 61 .11 62.78 61.67 63.33 100.00 63.33
[0328] 9: TMB12sol09 5 54.44 61.67 60.00 62 .22 61.67 58.89 58.33 63.33 100.00 The synthesized 12 p-strand designs (nine TMB polypeptides Table 2) predicted to have a square crosssection were tested. Genes were synthesized and the proteins were expressed as inclusion bodies in E. coli to avoid the complexity of targeting the outer membrane (Konovalova et al., 2017, Annu. Rev. Microbiol. 71, 539-556) (Figure 4). All tested proteins with a design based on the blueprint of Table 1 showed expression, except for TMB12_4, , and the majority even at high-levels (Figure 3A). The purified proteins were solubilized in guanidine hydrochloride and refolded by slow dilution into a buffer containing either detergent (fos-choline 12 (DPC) at a concentration double the critical micellar concentration (CMC)) or synthetic lipid vesicles (Material & Methods). As previously observed for the 8- stranded TMB designs, the standard band-shift assay on cold SDS-PAGE used to assess folding of natural TMBs was not informative to identify properly folded synthetic TMBs. Instead, the proteins were characterized by size exclusion chromatography (SEC), far UV circular dichroism (CD), urea titration (tryptophan fluorescence) and NMR in the presence of DPC detergent or DUPC (C11:OPC) large unilamellar vesicles (LUVs). For instance, the selected 12-strand TMB consisting of SEQ ID No. 3 showed a particularly favourable monodisperse SEC profile in DPC micelles (Figure 4A), sharp and reversible folding / unfolding transitions in the presence of DUPC LUVs (Figure 4B) (mid-point urea concentrations for folding (CmF): 4.5 M and 5.7 M, respectively) and NMR1H-15N HSQC chemical shift dispersion in DPC micelles indicative of folding (Figure 4C). The equilibrium unfolding curves were fitted to a two-states transition, with the calculated unfolding free energies (AGOUF) of -63.1 kJ / mol (for said TMB polypeptide design consisting of SEQ ID No. 3) in the range of natural (AGOUF -10 to -140 kJ / mol) and previously designed 8-stranded TMBs (-38 and -56 kJ / mol (Vorobieva et al., 2021, Science, 371, eabc8182)). The far UV CD spectra of the design folded into DUPC LUVs (Figure 4D) or DPC micelles (Figure 8A) showed the expected enrichment in p-sheet secondary structure by comparison with conditions in which the TMBs are not expected to fold (in 8 M urea, in the absence of LUVs or after thermal denaturation). To confirm that the design folded by integration into the bilayer rather than partial folding on its surface, the kinetics of folding were recorded in thinner (DUPC, C11:OPC) and thicker (DMPC, C14:0PC) membranes. Dramatically decreased folding rates were observed with DMPC LUVs compared with DUPC membranes (Figure 3B), consistent with integral membrane folding. The same experimental evidence of polypeptide folding was shown for TMB12_9, i.e. the TMB polypeptide consisting of SEQ ID No. 9.
[0329] Example 3. 12-stranded TMB functional characteristics.
[0330] Encouraged by these results, we assessed the nanopore activity of the TMB polypeptide design (SEQ ID No 3), based on the capacity of the TMB peptide to integrate spontaneously into planar Dipalmitoylphosphatidylcholine (DPhPC) membranes after dilution out of DPC micelles, and to conduct Na+and Cl’ ions. The TMB polypeptide design inserted successfully into the membrane, producing distinct jumps of current of reproducible intensities (Figure 5A), and had stable nanopore conductance. Unlike some native pores which can fall off from the membrane stochastically, the pore obtained by folding of the TMB polypeptide designs according to the invention remain stably inserted over long periods of time with the longest recording acquired being 2 hours for the design consisting of SEQ ID No. 3. Recording of the current-to-voltage response showed an incremental increase in observed conductance under positive and negative changes of voltage, indicative of stable transmembrane channels (1 / V curves in Figure 5B). The NMR 1H-15N HSQC spectra and smooth, folding / unfolding transitions of other TMB polypeptide designs is shown in Figure 9. The results suggest that membrane integration and stable nanopore activity correlates with stable and cooperative TMB folding in vitro, and that far UV CD and SEC, can be used to support evaluation of the folding of water-soluble de novo designed proteins, but may be less useful parameters in assessment of integral membrane folding of TMBs.
[0331] As a next step, the structures of the designs to assess the accuracy of the computational design methods. The structure of the TMB polypeptide consisting of SEQ ID No. 3 is elucidated by NMR spectroscopy. Optimization of the in vitro folding conditions showed that the protein was structured in aqueous solution in LDAO detergent micelles, as indicated by well-dispersed amide and side chain methyl spectra (Figure 6A). Secondary chemical shifts indicated the presence of twelve p-strands arranged into polypeptide segments expected from the design (Figure 3C). Amide and side chain methyl NOEs spanned a dense network of experimental connectivities that reached around the barrel circumference and thus confirmed the correct arrangement of the strands into the predicted barrel structure (Figure 6B). The TMB12_3 features the designed p-strands connectivity (shear number of 14) with the barrel closed by the canonical antiparallel pi-pi2 seam (Figure 3C, Figure 3D, Table 4). These experimental structures demonstrate that the computational design method can design TMB nanopores with precisely controlled shear, channel width and shape.
[0332] Table 4. NMR data and (refinement) statistics for TMB12_3 (TMB polypeptide consisting of SEQ ID No. 3) in LDAO micelles
[0333] * Calculated among 20 refined structures
[0334] Example 4: Electrophysiology characterization of obtained 12-stranded nanopores comprising the TMB polypeptides The TMB polypeptide designs were screened for nanopore activity directly after validating their folding into monodispersed -sheet structures in DPC micelles. The TMB polypeptide designs were evaluated for their capacity to insert into planar membranes from dilute detergent solution and form conducting pores (Figure 7B). The 12-stranded TMB polypeptide designs had similar conductances to each other (210-230 pS), based on measured conductances for the nanopores consisting of TMB polypeptides of SEQ. ID No. 3 and TMB polypeptide of SEQ ID No. 9. The conductances were consistent with a cylindrical nanopore of around 5 A. The predicted diameter is close to the average pore diameters of 9.4 A calculated from square 12-stranded design models using MOLE 2.5 (Figure 5C). In comparison to naturally-occurring pores used for sensing, such as OmpG which undergoes several transient and complete occlusion events by its solvent-exposed loops over a time span of 1 sec (Chen et al., 2008, Proc. Natl. Acad. Sci, USA. 105, 6272-6277), our TMB designs show remarkably quiet conductance. No occlusion events were detected over 10 sec measurements (Figure 5A). The obtained results anticipate that modulation of the nanopore shape and chemical lining could allow control over the permeability of the pores to larger and more complex solutes in the future.
[0335] Materials and Methods
[0336] General purification of all designs
[0337] All designs were purified from E. coli following a similar protocol as described previously in Vorobieva, Anastassia A., et al. "De novo design of transmembrane p barrels." Science 371.6531 (2021): eabc8182. Custom genes in a pET29b vector containing the kanamycin resistance gene were ordered from IDT and chemically transformed into BL21DE3 cells. All proteins were purified from inclusion body fractions following complete denaturation in 6M GuCI (Guanidine Hydrochloride) buffer. Briefly, inclusion pellets were washed several times with buffers containing 1% w / v of Triton X-100 and Brij-35 alternatively. A typically washing step involved resuspension of the insoluble pellet in the appropriate buffer, brief sonication and subsequent incubation for one hour at room temperature or overnight at 4°C. After solubilisation of the pellet in GuCI, the protein was diluted to 80-100 pM and refolded in a buffer containing 25mM Tris-CI at pH 8.0, 150 mM NaCI and 0.1% DPC (Dodecyl-Phosphatidyl-Choline) either using a dropwise dilution method or by spontaneous dilution to achieve a final GuCI concentration of 0.3 M. The diluted buffer was concentrated after overnight incubation with shaking at 4°C and run on a S200 Cytiva superdex 200 column. Fractions at expected volume were concentrated with a 10 kDa cutoff filter and used for subsequent analysis.
[0338] Conductance measurement in planar lipid bilayers
[0339] All ion-conductance measurements were carried out using the Nanion Orbit 16TC instrument (https: / / www.nanion.de / products / orbit-16-tc / ) on MECA chips. Lipid stock solutions were freshly made in dodecane at a final concentration of 5mg / mL. DPhPC (Di-Phytanoyl-Phosphatidyl-Choline) lipids were used for all experiments. Designed proteins were diluted in a buffer containing 0.05% DPC (~ 1 CMC), 25 mM Tris-CI pH 8.0 and 150 mM NaCI to a final concentration of ~100 nM. Subsequently, 0.5 pL or less of this stock was added to the cis chamber of the chip containing 200 pL of buffer while simultaneously making lipid bilayers using the in-built rotating stir-bar setup. All measurements were carried out at 25°C. Spontaneous insertions were recorded over multiple rounds of bilayer formation. All chips were washed with multiple rounds of ethanol and water and completely dried before testing subsequent designs. A 500 mM NaCI buffer was used on both sides of the membrane for all current recordings. Raw signals were recorded at a sampling frequency of 5 kHz. Only current recordings from bilayers whose capacitances were in the range 15-25 pF were used for subsequent analysis. The raw signals at 5 kHz were downsampled to 100 Hz using an 8-pole bessel filter. Estimation of current jumps were carried out using a custom script with appropriate thresholds. Current jumps larger than 2 times the smallest observed jump were discarded for single channel histogram calculations for each design.
[0340] TMB12sol9_3 (TMB polypeptide consisting of SEQ ID 3) Expression for NMR
[0341] BL21(DE3) Lemo cells were transformed with a pET29b-derived expression plasmid for TMB12sol9_3. Cells were grown in M9 minimal medium and expression was induced with 0.5 mM IPTG for 22h at 24 °C. For expression of [U-99%2H,15N,13C]-labeled samples, M9 was prepared with D2O,15NH4CI and deuterated-13C-glucose. For the expression of the [U-99%-2H,15N] labeled sample, M9 was prepared with D2O and15NH4CI. For the selectively labeled [15N-Lys] and [15N-Phe] samples, the desired15N labeled amino acid was added to the culture 45 min before induction.
[0342] Purification and Refolding of TMB polypeptides
[0343] Cells were lysed using a M110L from Microfluidics. Inclusion bodies were isolated and dissolved in denaturing buffer (20 mM Tris / HCI pH 8, 150 mM NaCI, 6M GdnHCI), then dialyzed against H2O in a 10,000 MWCO dialysis membrane for 2h, followed by centrifugation at 30,000 g to precipitate the protein. Precipitated TMB polypeptides were dissolved in 10 mM Tris / HCI pH 8, 7 M urea. Refolding was done at 4°C by dropwise rapid dilution into a stirred refolding buffer (20 mM Tris / HCI, 2 mM EDTA, 0.6 M L-Arg, 15 mM LDAO, pH 10). The dilution ratio was set to 1:20 and after overnight stirring, the refolded protein was dialyzed against 20 mM NaPi pH 6.8, 1 mM EDTA for 2h. The refolded protein was concentrated with MWCO 10,000 and the sample was loaded on an S200 size exclusion column, preequilibrated with 20 mM NaPi, 1 mM EDTA, 15 mM LDAO. The fractions containing protein were pooled and concentrated using MWCO 10,000.
[0344] Isotope labelling samples
[0345] The following samples were made: [U-99%-15N]-TMB12sol9_3 in LDAO, [U-99%-2H,15N,13C]- TMB12sol9_3 in LDAO, [U-99%-2H,15N; OO / o-^ / V-A; OO / o^H ^-M; OO / o^H61,13^1-!; 99° / o-1Hyl,13Cyl- V]-TMB12sol9_3 in [U-99%-2H]-LDAO, [15N-Lys]-TMB12sol9_3 in LDAO, [15N-Phe]-TMB12sol9_3 in LDAO. The final sample conditions were 20 mM Na-PO4, 1 mM EDTA, pH 6.8, 300-500 mM LDAO, 0.2-1 mM TMB12sol9_3. NMR experiments
[0346] All experiments were carried out at 25°C on Bruker spectrometers operating at field strengths of 700, 800 and 900 MHz. All spectrometers were equipped with a cryogenic triple-resonance probe. The following experiments were recorded: 2D [15N, -BEST-TROSY7, 3D BEST-TROSY-HNCACB with2H decoupling7, 3D H / HJ-NOESY-^N-TROSY8, 2D13C-Methyl-SOFAST7, 3D H / HPNOESY-^C-HMQC9.
[0347] Structure calculation
[0348] Structure calculation was performed with CYANA 3.98.15 (Guntert and Buchner, 2015, J. Biomol. NMR 62, 453-471). All spectra were processed and analyzed with NMRPipe11and ccpNMR version 312. Dihedral constraints were derived by TALOS-N13from the experimentally determined Ca, Cp, N and HN chemical shifts. Only TALOS-N predictions classified as "strong" were used. Tolerances were set to one standard deviation, capped at a maximum of 20°. The experimental NOEs were obtained from 3D [1H,1H]- NOESY-15N-TROSY and 3D [1H,1H]-NOESY-13C-HMQC spectra. Hydrogen bond constraints were inferred from the measured NOE data with upper limits of 2.0 / 3.0 A and lower limits of 1.8 / 2.7 A for the HN...0 and N...O, respectively. In regions with sparse assignment, these constraints were inferred indirectly from the experimentally established -strand topology (dashed black lines in Figure 3F). A total of 200 structures were calculated and the ensemble of 20 lowest energy structures was selected. Ramachandran statistics of this ensemble showed 73.5% of residues in most favored regions, 22.4% in allowed regions, 3.0% in generously allowed regions and 1.1% in disallowed regions.
[0349] Calculating the theoretical diameters of nanopores
[0350] The diameters of the nanopores were inferred from the observed single-channel conductance (G (nS)) using the access resistance model (eq. 1), where R is the resistance, S the conductivity of the solution (S / m), d is the diameter of the pore (nm) and L is the pore length (nm). The purely geometric model approximates the properties of a cylindrical nanopore and assumes homogeneous solution, pore and membrane neutrality and constant potential at the pore mouth.
[0351] A pore length of 3.5 nm was used for all the calculations based on the total transmembrane span of the TMB designs. The conductivity of a solution of 0.5 M NaCI was estimated to be 40.5 mS / cm based on the previously reported relationship between NaCI molarity and conductivity of the solution. Folding kinetics measured by Tryptophan fluorescence
[0352] Protein samples were buffer exchanged into unfolding buffer (50 mM Glycine-NaOH pH 9.5, 8 M / 10 M urea) using a 0.5 mL ZebaSpin 7K MWCO desalting column. The concentration was determined by nanodrop. For kinetic experiments, 15 pL of unfolded protein was added to 485 pL of pre-warmed (25 °C) pre-fold buffer of LUVs in 50 mM glycine-NaOH, pH 9.5, in a QS quartz cuvette to give final concentrations of 0.4 pM OMP, 600-3200 LPR (mol:mol), 0.24-4 M urea, 50 mM glycine-NaOH pH 9.5. Immediately after mixing, a time based fluorescence scan was carried out on a PTI QuantaMaster™ spectrofluorometer (Photon Technology International), controlled by FelixGX v4.3 software, with excitation at 280 nm, and emission measured at 335 nm. The slit settings were 0.5 nm for excitation, and 5 nm for emission, to minimize photobleaching. Integration was set at 1 s between time points, and the temperature was maintained at 25°C throughout. Fluorescence emission spectra were measured by exciting tryptophan at 280 nm, and measuring fluorescence emission between 300-400 nm, using the same slit-width settings as above, with samples in urea concentrations between 0.24-9.9 M urea.
[0353] Circular dichroism in liposomes
[0354] Protein samples were prepared in a similar manner as for the tryptophan fluorescence samples. Samples were made to a 600:1 (mol:mol) LPR, with final concentrations of 4 pM TMB, 1.2 mM lipid-LUV, 0.24-8 M urea, 50 mM glycine-NaOH pH 9.5 in a final reaction volume of 300 pL. The reaction was allowed to proceed overnight at 25°C to maximize the fraction of protein folded into the lipid-LUV bilayer. Controls were made where the volume of substrate was replaced with 50 mM glycine-NaOH pH 9.5 and an appropriate volume of urea to match the protein samples. These were used to normalize the data by subtracting their CD signals from the CD signal from the protein containing samples. Measurements were taken using 300 pL of sample in a 1 mm QS quartz cuvette, using a Chirascan plus CD Spectrometer (Applied Photophysics). The bandwidth was set at 2.5 nm, and used adaptive sampling to adjust the integration time for the optimal signaknoise. Four scans were averaged between 260 nm to the lowest useable wavelength for each respective sample, which was the point where the voltage reached its upper limit of 1000 V, after which the data became unusable. During temperature ramp experiments, only single scans were taken as the temperature ranged between 25 °C and 87 °C.
[0355] Equilibrium denaturation analysis
[0356] To determine the urea dependence of TMB folding, urea denatured TMB polypeptides in 50 mM glycine- NaOH pH 9.5, 10 M urea were diluted into DUPC or DLPC LUVs at an Lipid-to-Protein ratio (LPR) of 600:1 (mol / mol) to give a final concentration of 0.4 pM TMB in 50 mM glycine-NaOH pH 9.5 containing 2-9.9 M urea, and folding was allowed to proceed overnight at 25°C. For urea dependence of unfolding, TMBs were folded in DUPC or DMPC LUVs (LPR 600:1 (mol / mol)) in 50 mM glycine-NaOH pH 9.5, 2 M urea overnight at 25°C. Pre-folded TMBs were then unfolded by dilution into 50 mM glycine-NaOH pH 9.5 containing 2-9 M urea to a final TMB concentration of 0.4 pM and incubated overnight at 25 °C. Tryptophan fluorescence emission spectra were obtained using a PTI QuantaMaster spectrofluorometer (Photon Technology International) in QS quartz cuvettes with excitation slits set to 1 nm and emission slits set to 5 nm. Fluorescence was excited at 280 nm and emission spectra were acquired between 300- 400 nm using a step size of 1 nm and an integration time 0.5 seconds. Average wavelength between 325- 375 nm was calculated using equation 2, where <X> is the average wavelength, lx is the fluorescence intensity at a given wavelength, X is the wavelength, and 1 is the sum of the intensity of the entire emission spectra.
[0357] < X > = ^~ (eq. 2)
[0358] The experimental data were fitted to a 2-state transition model ("Determination and analysis of urea and guanidine hydrochloride denaturation curves” in Methods in Enzymology, 1986, vol. 131, 266-280) to extract AGo (the Gibb's free energy for unfolding in the absence of denaturant), the m-value (muF,the global dependence of AGo on the concentration of denaturant and Cm(the transition midpoint) based on ObsFand Obsu (the observed <X> for the folded and unfolded states in the absence of denaturant ([D]=0)), rrif and mu(the linear dependence of ObsFand Obsu to [D]). The observed <X> was corrected to account for the difference in quantum yield between the folded and unfolded states based on the Cifactor (QF), which is calculated by taking the ratio of the summed fluorescence intensities at the folded and unfolded states. R is the universal gas constant and T is the absolute temperature.
[0359] De novo protein design
[0360] The blueprint representations of the beta-barrel backbones were generated using a custom python script. The backbones were assembled based on such a blueprint and constraints file using the Rosetta BluePrintBDR application (Koga, N. et al., 2012, Nature. 491, 222-227). The highest-scoring protein backbones were used as templates for several rounds of combinatorial sequence design, alternating between the water-accessible pore and the lipid-exposed surface. To fine-tune the amino acid propensities achieved by combinatorial design on each layer, different energy function weights were tested on a representative protein backbone input. To fine-tune the water-accessible pore-lining residues, the weight of the Rosetta full-atom solvation energy (fasol from 0.8 to 1.5) and of the reference energies of small disorder-promoting amino acids (ALA and SER) were systematically varied. The scoring function used for subsequent sequence design was selected based on the closest match between the resulting designed sequences and naturally-occurring TMB sequences at the level of the overall hydropathy of the pore and the frequency of ALA and SER amino acids. To fine-tune the lipid-exposed surface residues, the weight of the Rosetta full-atom solvation energy and of the reference energies of large hydrophobic (PHE) and of small disorder-promoting amino acids (GLY and ALA) were systematically varied. The scoring function used for subsequent combinatorial sequence design was selected based on the closest match between the resulting designed sequences and naturally-occurring TMB sequences at the level of the overall hydropathy of the surface and the frequency of ALA and GLY amino acids. The designed sequences were continuously selected based on the stability of the mortise-tenon folding motifs and the properties of the networks of hydrogen bonds lining the pore. The final sequences were filtered based on the predicted secondary structure and aggregation propensities (with RaptorX (Wang et al., 2011, Proteomics. 11, 3786-3792) and Tango (Fernandez-Escamilla et al., 2004, Nat. Biotechnol.
[0361] 22, 1302-1306) predictors) and based on the capacity of AlphaFold2 (Jumper et al., 2021, Nature. 596, 583-589) to fold the sequence into the designed TMB structure without an input MSA.
Claims
Claims1. A transmembrane beta barrel (TMB) polypeptide comprising the formula X1-X2-X3-X4-X5-X6-X7-X8-X9-X10-X11-X12-X13-X14-X15-X16-X17-X18-X19-X20-X21-X22-X23, wherein:• XI comprises a beta strand, preferably said beta strand comprising or consisting of a sequence Z1-Z2-Z3-G-Z3-Z2-Z3-Z2-Y;• X3 comprises a beta strand comprising or consisting of a sequence Z1-Z2-Z3-G-Z3-Y-Z3-Z2- Y-Z2-Z1;• X5 comprises a beta strand comprising or consisting of the sequence Z1-Z2-Z3-Z2-Z3-Z2-Y- Z2-W;• X7 comprises a beta strand comprising or consisting of the sequence (N / Q)-Z2-Z3-Z2-Z3-G- Z3-Z2-Y-Z2-Z1;• X9 comprises a beta strand comprising or consisting of the sequence Z1-Z2-Z3-Z2-Z3-Z2-Z3- Z2-Y;• Xll comprises a beta strand comprising or consisting of the sequence Z1-Z2-Z3-Z2-Z3-Z2-Z3- Z2-Z3-Z2-Y-P-Z1;• X13 comprises a beta strand comprising or consisting of the sequence Z1-Z2-Z3-Z2-P-G-Z3- Z2-Y-Z2-W;• X15 comprises a beta strand comprising or consisting of the sequence (N / Q)-Z4-Z3-Z2-P-Y- Z3-Z2-Z3-Z2-Y-Z2-Z1;• X17 comprises a beta strand comprising or consisting of the sequence Z1-Z2-Z3-Z2-Z3-Z2-Z3- Z2-Y;• X19 comprises a beta strand comprising or consisting of the sequence Z1-Z2-Z3-Z2-Z3-G-Z3- Z2-Y-Z2-Z1;• X21 comprises a beta strand comprising or consisting of the sequence Z1-Z2-Z3-Z2-Z3-Z2-Y- G-W;• X23 comprises a beta strand comprising or consisting of the sequence (N / Q)-Z2-Z3-Z2-Z3- Z2-Z3-Z2-Y-Z2-Z1-Z5;• X2, X6, X10, X14, X18 and X22 each comprise a flexible loop, preferably the flexible loop comprising at least 3 and / or at most 10 amino acids;• X4, X8, X12, X16 and X20 each are a beta turn, said beta turn preferably comprising or consisting of 3 to 5 amino acids; wherein each one of XI, X3, X5, X7, X9, Xll, X13, X15, W17, X19 and X21 comprises a sequence of amino acids that alternate in their relative position within the TMB between the internal sideof the lumen formed by the TMB and the surface region of the TMB, and said sequence starting and / or ending with an amino acid located in the surface region of said TMB polypeptide; and wherein X23 comprises a sequence of amino acids that alternate in their relative position within the TMB between the internal side of the lumen formed by the TMB and the surface region of the TMB, and said sequence starting and / or ending with the amino acid in the surface region of said polypeptide, except for the two amino acids on the C-terminus indicated as Z1 and Z5 which are located in the surface region; wherein said Z1 amino acid is chosen from the group consisting of V, L, A and F, and said amino acid is located in the surface region of the TMB polypeptide; wherein said Z2 amino acid is chosen from the group A, D, E, F, H, K, L, M, N, Q, R, S, T, V and W, and said amino acid is located in the pore region; wherein said Z3 amino acid is chosen from the group consisting of V, L, A, G, S, T and F, and said amino acid is located in the surface region; and wherein said Z4 amino acid is chosen from the group consisting of A, D, E, H, K, L, M, N, Q, R, S, T and V, and said amino acid is located in the pore region; wherein Z5 is a single amino acid for a bulge in the TMB structure.
2. The TMB polypeptide according to claim 1 wherein Z5 is a single amino acid for a bulge being chosen from the group consisting of S, T or D.
3. The TMB polypeptide according to any one of claims 1 or 2, wherein any one of said beta turn X4, X8, X12, X16 or X20 is a beta turn comprising or consisting of Z5-Zx-Zx+i, Z5-ZX-ZX+I-ZX+2 or Z5- Zx-Zx+i-Zx+2-Zx+3 wherein Zx, Zx+i, Zx+2, Zx+3 are single turn amino acids.
4. The TMB polypeptide according to claim 3, wherein any one of said beta turn X4, X8, X12, X16 or X20 is a beta turn comprising or consisting of Z5-Zx-Zx+i, wherein Zx is preferably P and Z x+1 is preferably chosen from the group consisting of E, D, N or Y.
5. The TMB polypeptide according to any one of the preceding claims, wherein any one of X4, X8, X12, X16 or X20 consists of the sequence (S / T / D)-P-(E / D / N / Y).
6. The TMB polypeptide according to any one of the preceding claims, wherein X2 is the flexible loop comprising or consisting of the sequence N-T-D-N-T.
7. The TMB polypeptide according to any one of the preceding claims, wherein any one of X6, X14 or X22 is the flexible loop comprising or consisting of the sequence N-N-S-S-L.
8. The TMB polypeptide according to any one of the preceding claims, wherein any one of X10 or X18 is the flexible loop comprising or consisting of the sequence N-T-D-N-T.
9. The TMB polypeptide according to any one of the preceding claims wherein the connection between the beta strands XI and X23 comprises a canonical antiparallel seam.
10. The TMB polypeptide according to any one of the preceding claims, wherein said polypeptide further comprises a linker X0 preceding XI at the N-terminal domain of said peptide.
11. The TMB polypeptide according to claim 10, wherein said linker comprises or consists of the sequence Z6-Z7-Z8-G-(S / T / D / N), wherein Z6, Z7 and Z8 can be any amino acid except C.
12. The TMB polypeptide according to any one of the preceding claims, wherein said polypeptide comprises an amino acid sequence with at least 50 % sequence identity to SEQ ID NO: 3 or SEQ ID NO. 9.
13. The TMB polypeptide according to any one of the preceding claims wherein said polypeptide consists of SEQ ID NO: 3 or SEQ ID NO: 9.
14. The TMB of any one of claims 1 to 13, which is an engineered TMB polypeptide, preferably a circular permutated variant TMB polypeptide.
15. A nucleic acid encoding the TMB polypeptide according to any one of the preceding claims.
16. An expression vector comprising the nucleic acid of claim 15, preferably operably linked to a promoter sequence.
17. A recombinant host cell comprising the nucleic acid of claim 15 or the expression vector according to claim 16.
18. A nanopore comprising a lumen and a wall, wherein said wall comprises the TMB polypeptide according to any one of claims 1 to 14, the TMB encoded by the nucleic acid according to claim 15, or the TMB obtainable by the expression vector according to claim 16 or from the recombinant host cell according to claim 17.
19. The nanopore according to claim 18, wherein the lumen has a diameter of 21.0 to 24.6 A, preferably 22.0 to 23.6 A.
20. The nanopore according to any one of claims 18 or 19, wherein said lumen is of a substantially square shape.
21. The nanopore according to any one of claims 18 to 20, wherein the lumen further comprises a detectable moiety and / or a detector polypeptide.
22. The nanopore according to any one of claims 18 to 21 wherein said detectable moiety is selected from the group consisting of a pH-detectable moiety, electrical conductivity moiety, chromophore, fluorophore and chemically detectable moiety.
23. The nanopore according to any one of claims 18 to 22, wherein said peptide is an antigen-binding polypeptide, an ion-binding polypeptide, a pH responsive polypeptide and / or a nucleic acid binding polypeptide.
24. Use of the TMB polypeptide according any one of claims 1 to 14, the nucleic acid according to claim 15, the expression vector according to claim 16, the recombinant host cell according toclaim 17 and / or nanopore according to any one of claims 18 to 23 for molecular sensing and / or receiving and / or capturing of small-molecules or (poly)peptides, and / or drug delivery and / or as ion, water or small-molecule permeable transmembrane channels.
25. A pharmaceutical composition, comprising: • The TMB polypeptide according to any one of the claim 1 to 14, the nucleic acid according to claim 15, the expression vector according to claim 16, the recombinant host cell according to claim 17 and / or the nanopore according to any one of claims 18 to 23;• A carrier, preferably pharmaceutically acceptable carrier.
26. A computer-implemented method for obtaining the TMB polypeptide according to any one of claims 1 to 14, said method comprising:• Generating backbones corresponding to the TMB polypeptide according to any one of claims 1 to 14;• Assembling the generated backbones using the Rosetta BlueprintBDR.
27. A method for producing the TMB polypeptide according to any one of the claim 1 to 14, the method comprising: culturing a host cell under conditions conducive to the expression of the polypeptide; recovering the expressed polypeptide; and optionally, reconstitute the polypeptide in vitro in detergent micelles or lipids.
Citation Information
Patent Citations
Circularly permuted ligands and circularly permuted chimeric molecules
EP0754192B1
Transmembrane beta barrel proteins
WO2022051457A1
Beta barrel polypeptides and methods for their use
US20210047373A1
Transmembrane beta barrel proteins
US20230295230A1