Mutant CSGG pore
By modifying the CsgG nanopore structure for nucleic acid sequencing and combining it with nucleic acid processing enzymes for current signal measurement, the problems of slow speed and high cost in existing technologies have been solved, achieving rapid and inexpensive nucleic acid sequencing.
Patent Information
- Application Number
- CN202311730120.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2015-08-27
- Filing Date
- 2015-09-01
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2035-09-01
AI Technical Summary
Current nucleic acid sequencing technologies are slow and costly, relying on amplification techniques and large amounts of fluorescent chemicals, making it difficult to achieve rapid and inexpensive nucleic acid sequencing.
Molecular sensing is achieved using CsgG nanopores. By modifying CsgG monomers to optimize their structure, biopores suitable for nucleic acid sequencing are formed. Combined with nucleic acid processing enzymes, current signals are measured to identify nucleotides and simplify the system.
It improves the speed and reduces the cost of nucleic acid sequencing, enhances the ability to identify electrical signals by reducing the number of nucleotides and signal changes, and simplifies the sequencing system.
Smart Images

Figure BDA0004609908080000441 
Figure BDA0004609908080000442 
Figure BDA0004609908080000451
Abstract
Description
[0001] This patent application is a divisional application of the patent application with application number 2015800595058, application date September 1, 2015, and invention title "Mutant CSGG Pore". Technical Field
[0002] This invention relates to a novel protein pore and its uses. In particular, it relates to bio-nanopores in nucleic acid sequencing applications and molecular sensing.
[0003] This invention relates to mutant forms of CsgG. It also relates to the detection and characterization of analytes using CsgG. Background Technology
[0004] Protein pores are transmembrane polypeptides and complexes that form channels in a membrane, through which ions and certain molecules can pass. The minimum diameter of these channels is typically in the nanometer range (10⁻⁶). -9 (meters), so some of these peptides were named "nanopores".
[0005] Nanopores possess great potential as biosensors. When a potential is applied to a nanopore bound to a membrane, ions flow through the channel. This ion flow can be measured as an electric current. Suitable electrical detection techniques using single-channel recording devices are described, for example, in WO2000 / 28312 and D. Stoddart et al., Proc. Natl. Acad. Sci., 2010, 106, 7702-7. Multi-channel recording techniques are described, for example, in WO2009 / 077734.
[0006] Molecules transcribed through, or bound to, or near the pore, transcribe to impede and thus reduce the ion flow through the channel. The degree of reduction in ion flow indicates the size of the blockage within or near the pore, which is measured by a decrease in current. Therefore, the measured current can be used to measure the size or extent of channel blockage. Changes in current can be used to identify molecules or portions of molecules already bound in or near the pore (molecular sensing) or in a particular system, and can be used to determine the identity of molecules located within the pore based on molecular size (nucleic acid sequencing).
[0007] "Strand sequencing" is a known method for sequencing nucleic acids using biological nanopores. Single-stranded polynucleotides are passed through nanopores, and the bases on each nucleotide are determined by detecting changes in electrical current as they transiently pass through the nanopore channels. This method saves significant time and cost compared to traditional nucleic acid sequencing methods.
[0008] Previously reported protein nanopores, such as the mutant MspA (Manrao et al., Nature Biotechnology, 2012, 30(4), 349-353) and the α-hemolysin nanopore (Nat. Nanotechnol., 2009, 4(4), 265-70), have been used for nucleic acid sequencing using the "strand sequencing" method. Similarly, other pores for protein sensing, such as α-hemolysin (J Am Chem Soc, 2012, 134(5), 2781-7) and ClyA (Am. Chem. Soc. Nano. 2014, 8(12), 12826-35) (J. Am. Chem. Soc, 2013, 135(36), 13456-63), have also been used.
[0009] The need for novel nanopores that overcome the limitations of existing technologies remains, especially for optimizing the dimensions and characteristics of pores for molecular sensing applications and, for example, nucleic acid sequencing applications.
[0010] Nanopore sensing is a sensing method that relies on observing the binding of individual analyte molecules or the interaction between them and acceptors. Nanopore sensors are generated by placing nanoscale single pores in an insulating membrane and measuring voltage-driven ion transport through these pores in the presence of analyte molecules. Analyte identification is indicated by its unique current signal, significant duration and range of current blocks, and variations in current level.
[0011] There is a current demand for rapid, inexpensive, and widely applicable nucleic acid (e.g., DNA or RNA) sequencing technologies. Existing technologies are slow and expensive primarily because they rely on amplification techniques to generate large quantities of nucleic acids and require significant amounts of specialized fluorescent chemicals for signal detection. Nanopore sensing holds the potential to provide a rapid and inexpensive nucleic acid sequencing method by reducing the amount of nucleotides and reagents required.
[0012] Two essential components of nucleic acid sequencing using nanopore sensing are (1) controlling the movement of nucleic acids through the pore; and (2) recognizing nucleotides as nucleic acid polymers move through the pore. In the past, to achieve nucleotide recognition, nucleic acids were passed through a mutant of hemolysin. This provided an electrical signal, which was sequence-dependent. This also suggests that when using a hemolysin pore, a large number of nucleotides contribute to the observation of the current, making the direct relationship between the observed current and polynucleotides intriguing.
[0013] While mutations in the hemolysin wells have improved the range of currents used to identify nucleotides, the sequencing system would perform even better if the current differences between nucleotides were further increased. Furthermore, some current states exhibited high variability as nucleic acids moved through the wells. It was also shown that certain mutant hemolysin wells exhibited higher variability than others. While these state variations can contain sequence-specific information, it is desirable to generate wells with low variability to simplify the system. It is also desirable to reduce the number of nucleotides that contribute to the observation of the current. Summary of the Invention
[0014] The inventors have identified the structure of the bacterial amyloid secretory channel CsgG. The CsgG channel is a transmembrane oligomeric protein that forms channels with a minimum diameter of approximately 0.9 nm. The structure of the CsgG nanopores makes it suitable for protein sensing applications, particularly for nucleic acid sequencing. Modified variants of the CsgG peptide can be used to further enhance the channel's suitability for these specific applications.
[0015] Because its structure is better suited for DNA sequencing applications, the CsgG pore offers advantages over existing protein pores such as ClyA or α-hemolysin. The CsgG pore has a more favorable aspect ratio, including a shorter transmembrane channel than ClyA. The CsgG pore has a wider channel opening compared to the α-hemolysin pore. For some applications, this can facilitate enzyme attachment, such as for nucleic acid sequencing applications. In these embodiments, it can also minimize the length of the nucleic acid strand located between the enzyme and the read head (defined as the narrowest part of the pore), resulting in an improved readout signal. In embodiments of the invention relating to nucleic acid sequencing, the narrow internal constriction within the CsgG pore channel also facilitates the translocation of single-stranded DNA. The constriction consists of two loops formed by the juxtaposition of a tyrosine residue at position 51 (Tyr 51) and phenylalanine and asparagine residues at positions 56 and 55, respectively, of adjacent protein monomers. The size of the constriction can be modified. ClyA has a wider internal contraction structure, preventing the passage of currently unused double-stranded DNA for sequencing. The α-hemolysin well not only has a 1.3 nm wide internal contraction structure but also a 2 nm wide β-barrel structure with additional readhead features.
[0016] In a first aspect, the present invention relates to a method for molecular sensing, comprising:
[0017] a) Provides a CsgG biopore formed from at least one CsgG monomer within an insulating layer;
[0018] b) Apply an electric potential to the insulating layer to establish a current through the biopore;
[0019] c) Contact the CsgG biowell with the test substrate; and
[0020] d) Measure the current flowing through the biopore.
[0021] Typically, the insulating layer is a membrane, such as a phospholipid bilayer. In one embodiment, the current through the pores is carried by a flow of soluble ions flowing from a first side of the insulating layer to a second side of the insulating layer.
[0022] In one embodiment of the invention, the molecular sensing is analyte detection. In a particular embodiment, the method for analyte detection includes an additional step after step (d): determining the presence of a test substrate by a decrease in the current through the biopore compared to the current through the biopore when the test substrate is absent.
[0023] In an alternative embodiment of the invention, the molecular sensing is nucleic acid sequencing. Typically, the nucleic acid sequenced by this method is DNA or RNA. In a particular embodiment of the invention, the CsgG biopore is adapted to accommodate additional accessory proteins. Typically, the additional accessory proteins are nucleic acid processing enzymes selected from: DNA or RNA polymerases; isomerases; topoisomerases; helicases; telomerases; exonucleases; and helicases.
[0024] In embodiments of the present invention, the CsgG biopore is a modified CsgG pore, wherein the modified CsgG pore has at least one modification on the monomeric wild-type *E. coli* CsgG polypeptide sequence in at least one CsgG monomer of the CsgG pore. Typically, all CsgG monomers forming the CsgG pore undergo the same modification. In a specific embodiment of the present invention, the modified CsgG monomer has a polypeptide sequence from position 38 to 63 according to SEQ ID NO: 4 to 388.
[0025] In a second aspect, the present invention relates to modified CsgG biopores comprising at least one CsgG monomer, wherein the modified CsgG biopore has no more than one channel contraction structure having a diameter ranging from 0.5 nm to 1.5 nm. Typically, the modification is located between positions 38 and 63 of the CsgG monomer polypeptide sequence. Suitablely, the modification is located at positions selected from Tyr51; Asn55; and Phe56. In a specific embodiment, the modification is at the Tyr51 position, or at both Asn55 and Phe56 positions.
[0026] In embodiments of the invention, the modification of the CsgG monomer is selected from the following: substitution of naturally occurring amino acids; deletion of naturally occurring amino acids; and modification of the side chains of naturally occurring amino acids. Suitably, the modification reduces or removes steric hindrance of the unmodified amino acids. In a specific embodiment, at least one CsgG monomer of the pore has a polypeptide sequence from position 38 to 63 according to SEQ ID NO 4 to 388.
[0027] In a third aspect, the present invention relates to isolated polypeptides encoding at least one CsgG monomer of the modified CsgG biopore of the second aspect of the present invention.
[0028] In a fourth aspect, the present invention relates to isolated nucleic acids that encode the isolated polypeptides of the third aspect of the present invention.
[0029] In a fifth aspect, the present invention relates to a biosensor comprising:
[0030] a) Insulating layer;
[0031] b) CsgG biopores in the insulating layer; and
[0032] c) A device for measuring the current passing through the biopore.
[0033] In a specific embodiment, the CsgG biopore in the biosensor is a modified CsgG biopore according to the second aspect of the present invention.
[0034] In a sixth aspect, the present invention relates to the use of CsgG biopores for biosensing applications, wherein the biosensing applications are analyte detection or nucleic acid sequencing.
[0035] In one embodiment of the sixth aspect of the present invention, the nucleic acid sequencing is DNA sequencing or RNA sequencing.
[0036] The inventors have unexpectedly demonstrated that CsgG and its novel mutants can be used to characterize analytes, such as polynucleotides. The present invention relates to mutant CsgG monomers, wherein one or more modifications are made to enhance the ability of said monomers to interact with analytes, such as polynucleotides. The inventors have also unexpectedly demonstrated that pores containing novel mutant monomers have enhanced ability to interact with analytes, such as polynucleotides, and thus exhibit improved performance for evaluating analyte properties (e.g., polynucleotide sequences). The mutant pores unexpectedly exhibit improved identification of nucleotides. In particular, the mutant pores unexpectedly exhibit an increased current range, which makes it easier to distinguish different nucleotides and reduces state changes that increase the signal-to-noise ratio. Additionally, the number of nucleotides contributing to the current as polynucleotides move through the pore is reduced. This makes it easier to determine the direct relationship between the observed current and the polynucleotide as it moves through the pore. Furthermore, the mutant pores can exhibit increased flux, e.g., a greater likelihood of interaction with analytes, such as polynucleotides. This makes it easier to characterize analytes using said pores. The mutant pores can be more easily inserted into membranes.
[0037] Therefore, the present invention provides a mutant CsgG monomer comprising the sequence variant shown in SEQ ID NO: 390, wherein the variant comprises a mutation at one or more positions in Y51, N55 and F56.
[0038] Therefore, the present invention provides a mutant CsgG monomer comprising the sequence variant shown in SEQ ID NO: 390, wherein the variant comprises one or more of the following: (i) one or more mutations at the following positions (i.e., mutations at one or more of the following positions) N40, D43, E44, S54, S57, Q62, R97, E101, E124, E131, R142, T150, and R192; (ii) mutations at Y51 / N55, Y51 / F56, N55 / F56, or Y51 / N55 / F56; (iii) Q42R or Q42K; (iv) K49R (v) N102R, N102F, N102Y or N102W; (vi) D149N, D149Q or D149R; (vii) E185N, E185Q or E185R; (viii) D195N, D195Q or D195R; (ix) E201N, E201Q or E201R; (x) E203N, E203Q or E203R; and (xi) missing one or more of the following positions: F48, K49, P50, Y51, P52, A53, S54, N55, F56 and S57.
[0039] This invention also provides:
[0040] - A construct comprising two or more covalently linked CsgG monomers, wherein at least one monomer is a mutant monomer as described in this invention;
[0041] - A polynucleotide encoding the mutant monomer of this invention or the construct of this invention;
[0042] - A homologous oligomer pore derived from CsgG containing the same mutant monomer of the present invention or the same construct of the present invention;
[0043] - A heterooligoporous pore derived from CsgG containing at least one mutant monomer of the present invention or at least one construct of the present invention;
[0044] - Methods for determining the presence, absence, or one or more properties of a target analyte, including:
[0045] a) Contacting the target analyte with a CsgG well or a mutant thereof to move the target analyte relative to the well; and
[0046] b) Acquire one or more measurements as the analyte moves relative to the orifice and thereby determine the presence, absence, or one or more characteristics of the analyte;
[0047] - A method for forming a sensor for characterizing a target polynucleotide, comprising forming a complex between a CsgG pore or a mutant thereof and a polynucleotide-binding protein, thereby forming a sensor for characterizing the target polynucleotide;
[0048] - Sensors for characterizing target polynucleotides, including complexes between CsgG pores or mutants thereof and polynucleotide-binding proteins;
[0049] - Application of CsgG wells or their mutants in determining the presence, absence, or one or more properties of a target analyte;
[0050] - A kit for characterizing a target analyte, comprising (a) a Csg well or a mutant thereof, and
[0051] (b) Membrane composition;
[0052] - An apparatus for characterizing a target analyte in a sample, comprising (a) a plurality of CsgG wells or mutants thereof and (b) a plurality of membranes;
[0053] - Methods for characterizing target polynucleotides include:
[0054] a) Contacting the polynucleotide with a CsgG pore or a mutant thereof, a polymerase, and a labeled nucleoside to sequentially add a phosphate-labeled specie to the target polynucleotide via the polymerase, wherein the phosphate-labeled specie contains a label specific to each nucleotide; and
[0055] b) Using the wells to detect the phosphate markers and thereby characterize the polynucleotides; and
[0056] A method for producing the mutant monomer or construct described in this invention comprises expressing the polynucleotide described in this invention in a suitable host cell and thereby producing the mutant monomer or construct described in this invention. Attached Figure Description
[0057] Figure 1 This represents a side cross-sectional view of the CsgG nonamer in its channel conformation using ribbon and surface representation.
[0058] Figure 2 A cross-sectional view of the CsgG channel contraction structure (i.e., the hole readhead in a nanopore sensing application environment) and related diameter measurements are shown.
[0059] Figure 3 The structural motif representing the pore-shrinking structure consists of three stacked concentric side chain layers: Tyr 51, Asn55, and Phe 56.
[0060] Figure 4 Sequence homology among CsgG homologs is represented, including multiple sequence alignments of CsgG-like proteins (SEQ ID NO: 442 to SEQ ID NO: 448). Selected sequences are chosen from a monophyletic clade across the CsgG-like sequence phylogenetic tree (not shown) to provide a representative view of sequence diversity. Secondary structure elements are shown, with arrows or bars indicating β-chains and α-helices, respectively, and these secondary structure elements are based on the *E. coli* CsgG crystal structure. Importantly, residues corresponding to *E. coli* Tyr 51, Asn 55, and Phe 56 are highlighted by arrows. These residues form the internal contraction structure of the pore, i.e., the pore readout head in the aforementioned nanopore sensing applications.
[0061] Figure 5 Representative single-channel current records (a) and conductance histograms (b) of CsgG, recombined in a planar phospholipid bilayer and measured under an electric field of +50 mV (n = 33) or -50 mV (n = 13).
[0062] Figure 6This represents a single-channel current recording of CsgG recombined with PPB, performed at +50mV or -50mV with CsgE concentration increments. The horizontal scale bar is at 0pA.
[0063] Figure 7 a. Original negative-stained EM image of CsgG dissolved in C8E4 / LDAO. Arrows indicate different particle groups marked in the elution curves of the size shown in Figure g, which are aggregates of (I) CsgG nonamers, (II) CsgG octamericals, and (III) CsgG nonamers. Scale bar, 20 nm. b. Schematic diagram of the top and side views of the indicated oligomers, c. Rotational autocorrelation function plot of LDAO-dissolved CsgG in the top view, showing ninefold symmetry, d. CsgG C1S The original negative-stained EM image - arrows indicate hexammer (IV) and octamer (V) particles observed by size exclusion chromatography in Figure g. e,CsgG C1S A schematic diagram of the average side view of the oligomer. No top view of the construct was observed, f, CsgG observed by size exclusion chromatography as shown in Figure g. C1S Elution volume (EV) of CsgG particles, calculated molecular weight (MWcalc), and corresponding CsgG oligomer state (CsgG n ) expected molecular weight (MW) CsgG (and a table showing the symmetry of the particles as observed by negative staining electron microscopy (EM) and X-ray crystallography. g,CsgG) C1S Size exclusion chromatograms (black) and C8E4 / LDAO-dissolved CsgG (gray) on a Superdex200 10 / 300GL (GE Healthcare), h, i, schematic diagram of bands showing the top and side views of the crystalline oligomers, which display CsgG C1S (h) D8 hexadecomer and membrane-extracted D9 octadecomer in CsgG(i). A propolymer is colored with iridescent colors from the N-terminus (blue) to the C-terminus (red). Two C8 octamers (CsgG) are captured in the crystal to form a tail-to-tail dimer. C1S The C9 nonamer (CsgG) is colored blue and brown. r and θ represent the radius and interprotomer rotation, respectively.
[0064] Figure 8 Indicates CsgG C1S exist The electron density map (map) was calculated using experimental SAD phase averaging and density modification at 1.5g, and contours were plotted at 1.5g. The map shows the channel-constructed regions (CL; marking individual protopolymers) and is overlaid on the final refined model. (CsgG) C1S It is a mutant CsgG in which the N-terminal Cys, or Cys 1, of the mature CsgG sequence is replaced by Ser, resulting in the absence of lipid modification via the E. coli LOL pathway. This yields a soluble homooctameric oligomer existing in a pre-pore conformation (see [link to original text]). Figure 42 The conformation is the opposite of that formed by targeting homononomer pores through natural, ester-modified CsgG. Figure 43 ).
[0065] Figure 9 A top view showing a CsgG contractile structure modeled with polyalanine chains in an extended conformation shown from C-terminus to N-terminus passing through the channel. Figure 9 a) and side view ( Figure 9 b). Modelled solvation of polyacrylic acid chains, i.e. Figure 9 The position in b, such as Figure 9 As shown in c, the C-ring structure has been removed for clarity (the solvent molecules shown are those relative to the entire polyalanine chain). (molecules within).
[0066] Figure 10 : indicates CsgG in Escherichia coli.
[0067] Figure 11 : Indicates the size of CsgG.
[0068] Figure 12 : Explanation in The translocation of a single G. In *E. coli* CsgG (CsgG-Eco), there is a significant barrier to guanine entering the F56 ring. *=G enters the F56 ring. A=G ceases to interact with ring 56. B=G ceases to interact with ring 55. C=G ceases to interact with ring 51.
[0069] Figure 13 : Explanation in The translocation of ssDNA. A significant force is required to pull the DNA through the contractile structure of CsgG in the E. coli.
[0070] Figure 14 : Explanation in The translocation of ssDNA. Both the CsgG-F56A-N55S and CsG-F56A-N55S-Y51A mutants have low resistance to ssDNA translocation.
[0071] Figure 15-17 The mutant pores showed an increased range compared to the wild type (WT).
[0072] Figure 18-19 The mutant pores exhibited increased flux compared to the wild-type (WT).
[0073] Figure 20-21 The mutant pores exhibit enhanced insertion compared to the wild type (WT).
[0074] Figure 22 : indicates DNA construct X used in Example 18. The region labeled 1 corresponds to 30 SpC3 spacer regions. The region labeled 2 corresponds to SEQ ID NO:415. The region labeled 3 corresponds to 4 iSp18 spacer regions. The region labeled 4 corresponds to SEQ ID NO:416. The segment labeled 5 corresponds to 4 5-nitroindole. The region labeled 6 corresponds to SEQ ID NO:417. The region labeled 7 corresponds to SEQ ID NO:418. The region labeled 8 corresponds to SEQ ID NO:419, which has 4 iSp18 spacer regions (region labeled 9) connected to the 3' end of SEQ ID NO:419. At the other end of the iSp18 spacer regions is a 3' cholesterol chain (labeled 10). The region labeled 11 corresponds to 4 SpC3 spacer regions.
[0075] Figure 23 Example chromatographic trace of CsgG protein purified by a streptococcal trap (GE Healthcare) (x-axis label = elution volume (mL), y-axis label = absorbance (mAu)). The sample was loaded into 25 mM Tris, 150 mM NaCl, 2 mM EDTA, and 0.01% DDM, and eluted with 10 mM dethiobiotin. The elution peak of the eluted CsgG protein was labeled E1.
[0076] Figure 24 This section shows an example of typical SDS-PAGE visualization of CsgG protein after initial strep purification. 4-20% TGX gels (Bio Rad) were run at 300V for 22 minutes in 1X TGS buffer. The gels were stained with Sypro Ruby staining agent. Lanes 1-3 indicate the major elution peaks containing CsgG protein, as indicated by the arrows. Figure 23 The middle label is E1). Lanes 4-6 correspond to the main elution peaks containing contaminants. Figure 23 The eluted fraction at the tail end (labeled E1). M indicates the molecular weight marker used, which is unstained Novex Sharp (unit = kD).
[0077] Figure 25 Example of CsgG protein size exclusion chromatography (SEC) (120 mL S200 GE healthcare, x-axis label = elution volume (mL), y-axis label = absorbance (mAu)). The SEC was performed after strep purification and heating of the protein sample. The electrophoresis buffer for the SEC was 25 mM Tris, 150 mM NaCl, 2 mM EDTA, 0.01% DDM, 0.1% SDS, pH 8.0, and the column was run at a rate of 1 mL / min. The trace labeled X has absorbance at 220 nm, and the trace labeled Y has absorbance at 280 nm. Peaks labeled with an asterisk were collected.
[0078] Figure 26 Lane 1 shows a typical SDS-PAGE visualization of CsgG protein after SEC. 4-20% TGX gels (BioRad) were run at 300V for 22 minutes in 1X TGS buffer, and the gels were stained with Sypro Ruby staining agent. Lane 1 represents a CsgG protein sample after strep purification and heating but before SEC. Lanes 2-8 represent… Figure 25 The fraction collected at approximately 48-60 mL (intermediate peak = 55 mL) of the peak was run through the system and then... Figure 25 An asterisk is used to indicate the molecular weight marker used, which is unstained Novex Sharp (unit = kD). The bands corresponding to the E. coli CsgG wells are indicated by arrows.
[0079] Figures 27 to 33 The mutant pores showed an increased range compared to the wild type (WT).
[0080] Figures 34-39 The mutant pores showed an increased range compared to the wild type (WT).
[0081] Figure 40 This represents a snapshot (run) of the enzyme (T4 Dda-(E94C / C109A / C136A / A360C)(with mutant E94C / C109A / C136A / A360C and (△M1)G1G2) SEQ ID NO: 412) located at the top of well (CsgG-Eco-(Y51T / F56Q)-Strepll(C))9 (SEQ ID NO: 390 with mutant Y51 T / F56Q, where Stepll(C) is SEQ ID NO: 435 and is linked to well mutant No. 20 at the C-terminus)) at 0 and 20 ns. Figures 1 to 3 ).
[0082] Figure 41 The snapshots (runs) taken at 30 and 40 ns show the enzyme (T4 Dda-(E94C / C109A / C136A / A360C)(with mutant Y51 T / F56Q SEQ ID NO:390, where Stepll(C) is SEQ ID NO:435 and is linked to well mutant No.20 at the C-terminus) located at the top of well (CsgG-Eco-(Y51T / F56Q)-Strepll(C)9 (SEQ ID NO:390 with mutant Y51 T / F56Q, where Stepll(C) is SEQ ID NO:435 and is linked to well mutant No.20 at the C-terminus)) at the top of well (CsgG-Eco-(Y51T / F56Q)-Strepll(C)9 (SEQ ID NO:390 with mutant Y51 T / F56Q, where Stepll(C) is SEQ ID NO:435 and is linked to well mutant No.20 at the C-terminus)). Figures 1 to 3 ).
[0083] Figure 42 CsgG represents the form of the pre-pore conformation. C1S X-ray structure of a,CsgG C1S The monomer's banding diagram is colored a rainbow from blue to red from the N-terminus to the C-terminus. Secondary structure elements are labeled according to the ABD-like folding pattern, with additional N- and C-terminal α-helices and extended loops connecting β1 and α1 labeled as αN, αC, and C-rings (CL), respectively. C1S A side view of the C8 octamer, which has subunits distinguished by color and a subunit oriented and colored as shown in Figure a.
[0084] Figure 43 This represents the structure of CsgG in its channel conformation. a, CsgG C1S (Blue) and membrane-extracted CsgG (red) ATR-FTIR spectra in the amide I region (1,700–1,600 cm⁻¹). -1 b, TM1 and TM2 sequences (SEQ ID NO:449 and SEQ ID NO:450) (residues facing the bilayer, in blue) and the Congo red binding region of *E. coli* BW25141ΔcsgG complementary to wild-type csgG (WT), empty vector, or csgG lacking the underlined fragment of TM1 or TM2. Data are presented using 3 sets of biological replicates. c, Cover view of CsgG monomers in the prepore conformation (light blue; TM1 pink, TM2 purple) and channel conformation (brown; TM1 green; TM2 orange). CL, C ring. d, e, Side view (d) and cross-sectional view (e) of the CsgG nonmer in band-faceted representation; helix 2, core domain, and TM hairpin structure are shown in blue, light blue, and brown, respectively. Single protomers are shown as... Figure 42 a is colored. The magenta sphere indicates the location of Leu 2. OM, outer membrane.
[0085] Figure 44The contraction structure of the CsgG channel is shown. a, Cross-sectional view of the contraction structure of the CsgG channel and its diameter for solvent exclusion. b, The contraction structure consists of three stacked concentric side chain layers: Tyr 51, Asn 55, and Phe 56, preceded by Phe 48 from the periplasmic side. c, Topology of the CsgG channel. d, Congo red binding region of E. coli BW25141ΔcsgG with csgG (WT), empty vector, or csgG with a mutation showing the contraction structure. Data are presented using six sets of biological replicates. e, f, Representative single-channel current records (e) and conductance histograms (f) of CsgG recombined in a planar phospholipid bilayer and measured at electric fields of +50 mV (n = 33) or -50 mV (n = 13).
[0086] Figure 45 Model representing the CsgG transport mechanism. a, Non-denaturing gel electrophoresis (NativePAGE) of CsgE(E), CsgG(G), and CsgG supplemented with excess CsgG(E+G), showing the formation of the CsgG-CsgE complex (EG*). Data are presented as seven experiments including four different batches of protein. b, SDS-PAGE of CsgE(E), CsgG(G), and the EG* complex recovered from non-denaturing gel electrophoresis (NativePAGE). Data are presented as two replicate experiments. M, Molecular weight marker. c, Average plot of selected species of CsgG-CsgE particles. From left to right: Top and side views seen by cryo-electron microscopy (cryo-EM), and a comparison plot of the negatively stained side view with the CsgG nonamic. d, Cryo-electron microscopy (cryo-EM) averages of top and oblique side views of CsgE particles. Rotational autocorrelation shows ninefold symmetry. e. A 3D reconstruction of CsgG-CsgE (24 Å resolution, 1221 single particles) showing the density of nonamic particles including CsgG (blue) and additionally represented as CsgE nonamic particles (orange). f. Single-channel current recording of PPB-recombined CsgG at +50 mV or -50 mV, supplemented with increased CsgE concentration. The horizontal scale bar is at 0 pA. g. A tentative model of CsgG-mediated protein secretion. CsgG and CagE were proposed for the formation of a secretory complex that captures CsgA (in... Figure 54 (Discussed in the previous section), an entropic potential is generated in the channel. After CsgA is captured in the channel's contraction structure, DS-rectified Brownian diffusion causes the polypeptide to gradually translocate across the outer membrane.
[0087] Figure 46This diagram illustrates the Curli biosynthetic pathway in *E. coli*. The larger curli subunit, CsgA (light green), is secreted from the cell as a soluble monomeric protein. The smaller curli subunit, CsgB (dark green), is attached to the outer membrane (OM) and acts as a nucleating agent, converting CsgA from a soluble protein into amyloid precipitate. CsgG (orange) assembles into oligomeric curli-specific transport channels in the outer membrane. CsgE (purple) and CsgF (light blue) form soluble accessory proteins required for the transport and deposition of abundant CsgA and CsgB. CsgC forms a putative oxidoreductase with unknown function. All curli proteins possess a putative Sec signaling sequence for transcytoplasmic (inner) membrane (IM) transport.
[0088] Figure 47 This indicates the CsgG and CsgG analyzed by size exclusion chromatography and negative staining electron microscopy. C1S Oligomeric states in solution. a, Original negative-stained EM image of CsgG dissolved in C8E4 / LDAO. Arrows indicate different particle groups marked in the size elution curves as shown in g, which are aggregates of (I) CsgG nonamers, (II) CsgG octamerics, and (III) CsgG nonamers. Scale bar, 20 nm. b, Schematic diagram of class-averaged top and side views of the indicated oligomeric states. c, Rotational autocorrelation function plot of LDAO-dissolved CsgG in the top view, showing ninefold symmetry. d, CsgG C1S The original negative-stained EM image, with arrows indicating hexammer (IV) and octamer (V) particles observed by size exclusion chromatography in Figure g. e, CsgG C1S A schematic diagram of the oligomer's side view. A top view of the construct not observed. f, CsgG observed by size exclusion chromatography as shown in Figure g. C1S Elution volume (EV) of CsgG particles, calculated molecular weight (MWcalc), and corresponding CsgG oligomer (CsgG n A table showing the expected molecular weight (MWCsgG) of g,CsgG and the symmetry of the particles as observed by negative staining EM and X-ray crystallography. C1S Size exclusion chromatograms of (black) and C8E4 / LDAO-dissolved CsgG (gray) on a Superdex 200 10 / 300GL (GE Healthcare). Schematic diagrams of the banded crystalline oligomers in h, i, top and side views, showing CsgG. C1S(h) D8 hexadecomer and membrane-extracted CsgG(i) D9 octadecomer. A propolymer is colored iridescently from the N-terminus (blue) to the C-terminus (red). Two C8 octamers (CsgG) are captured in the crystal to form tail-to-tail dimers. C1S The C9 nonamer (CsgG) is stained blue and brown. r and h represent the radius and the angle between protomers, respectively.
[0089] Figure 48 This section shows a comparison of CsgG with its structural homologs and the contacts between protopolymers in contact with CsgG. a, b, CsgG C1S Banding diagrams of monomers (e.g., CsgG in the pre-pore conformation) (a) and the nucleotide-binding domain-like domain of TolB (b) (PDB 2hqs), both colored in iridescent hues from the N-terminus (blue) to the C-terminus (red). Ordinary secondary structure elements are labeled in the same manner. c, CsgG C1S (Gray), from left to right, stacked with Xanthomonas campestris rare lipoprotein B (PDB 2r76, colored pink), Shewanella oneidensis hypothetical lipoprotein DUF330 (PDB 2iqi, colored pink), and Escherichia coli TolB (PDB 2hqs, whose N-terminus and b-propeller domains are colored pink and yellow, respectively). CsgG-specific structural elements are labeled and colored, as shown in the upper left image. d, e, band diagrams of two adjacent protomers found in the CsgG structure, viewed along the bilayer plane from the outside (c) or inside (d) of the oligomer. The first protomer is represented by iridescent colors (dark blue to red) from the N-terminus to the C-terminus; the second protomer is represented by light blue (core domain), blue (helix 2), and brown (TM domain). Four main oligomer interfaces are evident: the interaction of the b6-b39 backbone within the b-barrel structure, the contraction loop (CL), the side chain stacking of helix 1 (α1) abutting against b1-b3-b4-b5, and the helix-helix stacking of helix 2 (α2). It was also observed that an N-terminal loop of 18 residues connecting the lipid anchor (magenta spheres indicate Ca positions in Leu 2) and the N-terminal helix (αN) encloses two adjacent protomers. The projected position of the lipid anchor is expected to abut against the TM1 and TM2 hairpin structures of the +2 protomer (not shown for clarity).
[0090] Figure 49Cys accessibility test of selected surface residues in CsgG oligomers. Schematic diagram of CsgG nonamers showing bands in ac, periplasmic (a), lateral (b), and extracellular (c) views. A protomer is colored iridescently from the N-terminus (blue) to the C-terminus (red). Cysteine substituents are labeled, and the equivalent position of the S atom is shown as a sphere, colored according to the accessibility of MAL-PEG (5,000 Da) labeling in the outer membrane of *E. coli*. d, Western blot analysis of the sample reacted with MAL-PEG on SDS-PAGE, showing an increase of 5 kDa in binding of the introduced cysteine to MAL-PEG. Accessible (11 and 111), moderately accessible (1), and inaccessible (2) sites are colored green, orange, and red in ae, respectively. For Arg 97 and Arg 110, a second substance at 44 kDa is present, corresponding to a portion of the protein where both introduced and native cysteine residues are labeled. Data are presented using four sets of independent biological repeatability experiments. e, a side view of the dimer interface in the D9 octameric structure in X-ray diffraction. Introduced cysteine residues are labeled at the dimer interface or within the lumen of the D9 particle. These residues are accessible to MAL-PEG in membrane-bound CsgG, demonstrating that the D9 particle is a product of membrane-extracted CsgG concentrate and that the C9 complex forms physiologically relevant substances. Residues in the C-terminal helix (aC; Lys 242, Asp 248, and His 255) were found to be inaccessible to difficult to access, suggesting that aC can form additional contact with the *E. coli* cell membrane, possibly the peptidoglycan layer.
[0091] Figure 50 The images show molecular dynamics simulations of the CsgG contractile structure using a polyalanine chain model. a, b, top (a) and side (b) views of the simulated CsgG contractile structure with the polyalanine chain extending its conformation through the channel, indicated here with the C-terminus to N-terminus orientation. The substrate channel via the CsgG transporter is not sequence-specific. 参考文献16,23For clarity, a polyalanine chain is used to simulate the hypothetical interactions of a forward-transitioning polypeptide chain. The simulated region consists of nine concentric CsgG C-ring structures, each containing residues 47-58. The side chains connecting the contracted structures are schematically represented using rods, where Phe 51 is colored slate blue, Asn 55 (amide clip) is colored cyan, and Phe 48 and Phe 56 (Φ clip) are colored light orange and dark orange, respectively. N, O, and H atoms (only hydroxyl groups or amide H atoms on the side chains are shown) are colored blue, red, and white, respectively. The C, N, O, and H atoms in the polyalanine chain are colored green, blue, red, and white, respectively. The polyalanine residues in the contracted structure... Solvent molecules (water) within the range are represented by red dots. c, The simulated solvation of the polyalanine is shown in the position shown in Figure b, and the C-ring structure has been clearly removed (the solvent molecules shown are those within the entire polyalanine). Solvent molecules within the range). At the height of the amide clamp and Φ clamp, the solvation of the polyalanine chain weakens into a single water shell bridging the peptide backbone and the amide clamp side chains. Most side chains in the Tyre 51 ring, along with their CsgG (and CsgG) C1S The inward-pointing, center-oriented position observed in the X-ray structure rotates towards the solvent. This model is based on GROMACS. 参考文献53 The results of a 40 ns all-atom explicit solvent molecular dynamics simulation were obtained using an AMBER99SB-ILDN54 force field, with the positions of the Cα atoms of the residues (Gin 47 and Thr 58) at the C ring ends restricted.
[0092] Figure 51The diagram illustrates sequence conservation in CsgG homologs. a, Schematic diagram of the surface of the CsgG nonamic, colored according to sequence similarity (conservation is scored from low to high based on the coloring from yellow to blue), and viewed from the periplasm (far left), side (middle left), extracellular environment (middle right), or as a cross-sectional view (far right). The accompanying figure shows the regions of highest sequence conservation mapping on the entrance of the periplasmic vestibule, the vestibular side of the contractile loop, and the luminal surface of the TM domain. b, Multiple sequence alignment of CsgG-like lipoproteins. Selected sequences were chosen from monophyletic clades (not shown) of three CsgG-like sequences across phylogeny to provide a representative view of sequence diversity. The β-chains and α-helices of secondary structural elements are indicated by arrows or bars, and these secondary structural elements are based on the E. coli CsgG crystal structure. c, d, Schematic diagram of the secondary structure of the CsgG protomer (c) and cross-sectional view of the surface of the CsgG nonamer (d), both are gray, and three consecutive highly sequence-conserved regions are colored red (HCR1), blue (HCR2), and yellow (HCR3). HCR1 and HCR2 form the vestibular side of the contractile ring; HCR3 corresponds to helix 2 and is located at the entrance of the periplasmic vestibule. Inside the contractile structure, Phe 56 is 100% conserved, while Asn 55 can be substituted by Ser or Thr conserved substitution, for example, by small polar side chains that can act as hydrogen bond donors / acceptors. The concentric side chain ring (Tyr51) at the exit of the contractile structure is not conserved. The presence of the Phe-ring at the entrance of the contractile structure is topologically similar to the Phe427 ring (called Φ clip) in the anthrax protective antigen PA63, which is shown to catalyze the capture and passage of peptides. 参考文献20 The MST of the toxB superfamily proteins showed a conserved motif D(D / Q)(F)(S / N)S at the height of the Phe ring. This is similar to the motif S(Q / NT)(F)ST seen in curli-like transporters. Although the atomically resolved PA63 structure in the pore conformation could not be obtained, the obtained structure suggests that the Phe ring could similarly be subsequently formed by a conserved hydrogen bond donor / acceptor (Ser / Asn428) as a concentric ring in the subsequent translocation channel (it should be noted that the orientation of the elements is reversed in both transporters).
[0093] Figure 52Single-channel current analysis of CsgG and CsgG:CsgE orifices. a) Under negative field potential, the CsgG orifice exhibits two conductance states. The upper left and upper right plots show representative single-channel current trajectories for the normal (measured at +50, 0, and -50 mV) and low-conductance forms (measured at 0, +50, and -50 mV), respectively. No transition between the two states was observed over the total observation time (n = 22), indicating a long lifetime (timescale from seconds to minutes) for the conductance states. The lower left plot shows the current histograms for the normal and low-conductance forms of the CsgG orifice obtained at +50 and -50 mV (n = 33). The IV curves of the CsgG orifice under normal and low-conductance conditions are represented by the lower right plot. The data represent the mean and standard deviation over at least four independent records. The natural or physiological presence of the low-conductance form is unknown. b) Electrophysiological properties of the CsgG channel titrated with the cofactor CsgE. This curve shows the proportion of channels in open, intermediate, and closed states as a function of CsgE concentration. The open and closed states of CsgG are shown as follows: Figure 45 As shown in Figure f, increasing the CsgE concentration to more than 10 nM caused the CsgG pores to close. This effect occurred at +50 mV (left panel) and -50 mV (right panel), ruling out the possibility that CsgE (calculated pI 4.7) electrophoresis into the CsgG pores caused by pore blockage. A low-probability (5%) intermediate state had approximately half the open-channel conductance. This could represent incomplete closure of the CsgG channel caused by CsgE; or it could represent temporary formation of CsgG dimers caused by residual CsgG monomers from the electrolyte solution binding to the pores embedded in the membrane. The proportions of these three states can be obtained from full-point histogram analysis of the single-channel current traces. The peak regions of up to three states produced by the above histogram, and the proportion of a given state, can be obtained by dividing the corresponding peak area by the sum of all other states recorded. At a negative field potential, two open-conductance states were identified, similar to the observations for CsgG (see Figure a). Because changes in both open channels are impeded by higher CagE concentrations, the “open” trace in b combines both conductance forms. The data in the graph represent the mean and standard deviation from three independent records. c, Crystal structure, size exclusion chromatography, and EM show that detergent-extracted CsgG pores form non-natural tail-tail stacked dimers (e.g., two nonamers such as D9 particles) at higher protein concentrations; Figure 47These dimers can also be observed in single-channel recordings. The top panel shows the single-channel current traces of stacked CsgG orifices under +50, 0, and -50 mV conditions (from left to right). The bottom left panel shows the current histograms of the dimerized CsgG orifices recorded under +50 and -50 mV conditions. The experimental conductivities of +16.2 ± 1.8 and -16.0 ± 3.0 pA (n = 15) at +50 and -50 mV conditions, respectively, are close to the theoretically calculated value of 23 pA. The bottom right panel shows the IV curves of the stacked CsgG orifices. Data represent the mean and standard deviation of six independent recordings. d, The ability of CsgE to bind to and block stacked CsgG orifices was tested electrophysiologically. It shows the single-channel current traces of stacked CsgG orifices in the presence of 10 or 100 nMCsgE under +50 mV (top panel) and -50 mV (bottom panel) conditions. The current traces indicate that the pre-existing saturation concentration of CsgE does not lead to pore closure of the stacked CsgG dimers. These observations are consistent with mappings from the CsgG-CagE contact region to the helix 2 and CsgG periplasmic cavity entrances, as identified by EM and site-directed mutagenesis. Figure 45 and Figure 52 ).
[0094] Figure 53 Representing CsgE oligomers and CsgG-CsgE complexes, a) size exclusion chromatography of CsgE (Superose 6, 16 / 600; electrophoresis buffer 20 mM Tris-HCl pH 8, 100 mM NaCl, 2.5% glycerol) shows an equilibrium of two oligomeric states, 1 and 2, with an apparent molecular weight ratio of 9.16:1. Negative-stain EM observation of peak 1 shows discrete CsgE particles (inset: average image of five representative species, ordered by increasing tilt angle) comparable in size to nine CsgE replicas. b) Average image of selected species of CsgE oligomers observed from a top-view perspective by cryo-electron microscopy (cryo-EM), and its rotational autocorrelation plot shows the presence of C9 symmetry. c) FSC analysis of the CsgG-CsgE cryo-electron microscopy (cryo-EM) model. Using 125 categories corresponding to 1221 particles, and with a correlation threshold of 0.5, FSC analysis determined that the 3D reconstruction was achieved. d, Overlay of the CsgG nonammer as observed in CsgG-CsgE cryo-electron microscopy density and X-ray structure. The overlay is shown as a translucent density map (left) or cross-sectional view from a side view. e, Congo red binding region of E. coli BW25141ΔcsgG complementary to wild-type csgG (WT), empty vector (ΔcsgG), or csgGhelix 2 mutant (replaced by a single amino acid with a single letter coding label). Data are presented as 4 sets of biological replicates. f, Effect of bile salt toxicity on E. coli LSR12 complementary to csgG (WT) or csgG carrying different helix 2 mutations (complementary to csgE (1) or non-complementary (2)). 10-fold serial dilutions starting from 107 bacteria were spotted on MacConkey agar plates. Expression of CsgG pores in the outer membrane leads to increased bile salt sensitivity, which can be blocked by co-expression of CsgE (n=6, three biological replicates, twice per replicate). g, Cross-sectional view of the CsgG X-ray structure in the molecular surface schematic diagram. CsgG mutants that have no effect on the Congo red binding region or toxin are shown in blue; mutants that interfere with the rescue of CsgE-mediated bile salt sensitivity are shown in red.
[0095] Figure 54 This describes the assembly and substrate recruitment of the CsgG secretion complex. The Curli transporter CsgG forms a secretion complex with a stoichiometric ratio of 9:9 with the soluble secretion cofactor CsgE, which encloses a... The chamber, which is believed to be used to capture CsgA substrates and facilitate their transmembrane (OM); see above and Figure 45Entropy-driven diffusion. Theoretically, three hypothetical pathways (ac) can be envisioned for the substrate recruitment and assembly of the secretion complex. a, The “catch-and-cap” mechanism requires CsgA to bind to the apolipoprotein CsgG translocation channel (1), causing a conformational change in the latter, exposing a high-affinity binding platform for CsgE binding (2). The binding of CsgE results in capping the substrate cage structure. Upon secretion of CsgA, CsgG will return to its low-affinity conformation, leading to the dissociation of CsgE and the release of the secretion channel for a new secretion cycle. b, In the “dock-and-trap” mechanism, CagE first captures CsgA in the periplasm (1), causing CagE to adopt a high-affinity complex that docks to the CsgG translocation pore (2), surrounding CsgA in the secretion complex. CsgA can bind directly to either the CsgE oligomer or the CagE monomer, the latter leading to subsequent oligomerization and binding to CsgG. CsgA secretion causes CagE to revert to its low-affinity conformation and dissociate from the secretion channel. c, CsgG and CagE form a constitutive complex, in which the conformational kinetics of CagE cycle between open and closed forms during CsgA recruitment and secretion. Currently published or existing data do not allow us to distinguish these putative recruitment modes or their derivative modes, or to propose one of the aforementioned modes.
[0096] Figure 55 Indicates CsgG C1S Data acquisition statistics and electron density mapping of CsgG. a,CsgG C1S Statistical chart of data acquisition for CsgG X-ray structure. b, in Download CsgG C1S The electron density maps, calculated using NCS averaging and density-modified experimental SAD phases, are plotted as contour lines at 1.5σ. The maps show the channel-constructed regions (CL; labeled with individual protopolymers) and are overlaid on the final refined model. Electron density maps in the c,CsgG TM domains (with resolutions of 3.6, 3.7, and 3.7 along the reciprocal vectors a*, b*, and c*, respectively) are also shown. The value is calculated from NCS-averaged and density-modified molecular substitution phases (TM rings are not present in the input model); B-factor sharpening. Contours were plotted at 1.0σ. The figure shows the TM1 (Lys 135-Leu 154) and TM2 (Leu182-Asn 209) regions of individual CsgG promers, overlaid on the final refined model.
[0097] Figure 56 The image shows a single-channel current trace (left) and its magnified region (right) of a CsgG WT protein interacting with a DNA hairpin structure carrying single-stranded DNA overhangs. The trace shows the changing current (as indicated by arrows) in response to potentials measured at +50 mV or -50 mV intervals. The stagnation of the descending current in the final +50 mV segment indicates that the hairpin double strands simultaneously enter the pore lumen and spiral the single-stranded hairpin ends into the pore's internal structure, resulting in near-complete current stagnation. Reversing the electric field to -50 mV causes electrophoretic unblocking of the pore. The new +50 mV phase again leads to the entry / coiling of the DNA hairpin structure and pore blockage. Uncoiling of the hairpin structure in the +50 mV segment can cause cessation of the current stagnation, as indicated by the reversal of the current stagnation. The hairpin structure—which has the sequence: 3'GCGGGGA GCGTATT AGAGTTG GATCGGATGCA GCTGGCTACTGACGTCATGACGTCAGTAGCCAGCATGCATCCGATC-5'—is added to the cis side of the chamber at a final concentration of 10 nM.
[0098] Figure 57 The purification and channel properties of CsgG-ΔPYPA are described, wherein CsgG-ΔPYPA is a mutated CsgG pore in which the PYPA sequence (residues 50-53) at the contracted residue Y51 is mutated to GG.
[0099] Sequence List Description
[0100] SEQ ID NO:1 represents the amino acid sequence of wild-type Escherichia coli CsgG containing the signal sequence (Uniprot access number P0AEA2).
[0101] SEQ ID NO:2 represents the polynucleotide sequence of wild-type Escherichia coli CsgG containing the signal sequence (gene ID: 12932538).
[0102] SEQ ID NO:3 represents the amino acid sequence of wild-type Escherichia coli CsgG at positions 53 to 77 in SEQ ID NO:2. This corresponds to the amino acid sequence at positions 38 to 63 of the mature wild-type Escherichia coli CsgG monomer (i.e., the signal sequence is missing).
[0103] SEQ ID NO:4 to 388 represent the amino acid sequences from position 38 to 63 of the CsgG modified monomer that lacks the signal sequence.
[0104] SEQ ID NO:389 represents a codon-optimized polynucleotide sequence encoding a wild-type CsgG monomer from Escherichia coli Str.K-12substr.MC4100. This monomer lacks the aforementioned signal sequence.
[0105] SEQ ID NO:390 represents the amino acid sequence of a mature wild-type CsgG monomer from *Escherichia coli* Str.K-12substr.MC4100. This monomer lacks the aforementioned signal sequence. The abbreviation used here, CsgG, means CsgG-Eco.
[0106] SEQ ID NO:391 represents the amino acid sequence of the hypothetical protein CKO_02032, YP_001453594.1:1-248 [Citrobacter koseri ATCC BAA-895], which is 99% identical to the sequence of SEQ ID NO:390.
[0107] SEQ ID NO:392 represents the amino acid sequence of WP_001787128.1:16-238 of the curli production assembly / transport component CsgG, partially [Salmonella enterica], which is 98% identical to the sequence of SEQ ID NO:390.
[0108] SEQ ID NO:393 represents the amino acid sequence of KEY44978.1|:16-277 of the curli production assembly / transport protein CsgG [Citrobacter amalonaticus], which is 98% identical to the sequence of SEQ ID NO:390.
[0109] SEQ ID NO:394 represents the amino acid sequence of YP_003364699.1:16-277 [Citrobacter rodentium ICC168] of the curli production assembly / transport component, which is 97% identical to the sequence of SEQ ID NO:390.
[0110] SEQ ID NO:395 represents the amino acid sequence of YP_004828099.1:16-277 of the curli production assembly / transport component CsgG [Enterobacter asburiae LF7a], which is 94% identical to the sequence of SEQ ID NO:390.
[0111] SEQ ID NO:396 represents a polynucleotide sequence encoding Phi29 DNA polymerase.
[0112] SEQ ID NO:397 represents the amino acid sequence of Phi29 DNA polymerase.
[0113] SEQ ID NO:398 represents a codon-optimized polynucleotide sequence derived from the *E. coli* sbcB gene. It encodes an exonuclease I (EcoExo I) from *E. coli*.
[0114] SEQ ID NO:399 represents the amino acid sequence of an exonuclease I (EcoExo I) from Escherichia coli.
[0115] SEQ ID NO:400 represents a codon-optimized polynucleotide sequence of the xthA gene from Escherichia coli. It encodes an exonuclease III enzyme derived from E. coli.
[0116] SEQ ID NO:401 represents the amino acid sequence of an exonuclease III enzyme from *Escherichia coli*. This enzyme performs distributed digestion of one strand of double-stranded DNA (dsDNA) along the 3' to 5' direction using 5' monophosphate nucleosides. Initiation of the enzyme on the strand requires a 5' overhang of approximately 4 nucleotides.
[0117] SEQ ID NO:402 represents a codon-optimized polynucleotide sequence of the recJ gene derived from *T. thermophilus*. It encodes the RecJ enzyme (TthRecJ-cd) from *T. thermophilus*.
[0118] SEQ ID NO:403 represents the amino acid sequence (TthRecJ-cd) of the RecJ enzyme from *Streptococcus thermophilus*. This enzyme progressively digests 5' monophosphate nucleotides of single-stranded DNA (ssDNA) from the 5' to 3' direction. Initiation of the enzyme on the strand requires at least 4 nucleotides.
[0119] SEQ ID NO:404 represents a codon-optimized polynucleotide sequence derived from the bacterial phage λ exonuclease (redX). It encodes the bacterial phage λ exonuclease.
[0120] SEQ ID NO:405 represents the amino acid sequence of a bacterial bacteriophage λ exonuclease. This sequence is one of the three identical subunits that assemble into a trimer. This enzyme performs highly progressive nucleotide digestion of one strand of double-stranded DNA (dsDNA) from the 5' to 3' direction (http: / / www.neb.com / nebecomm / products / productM0262.asp). Initiation of the enzyme on the strand preferably requires a protrusion of approximately four nucleotides with a 5' phosphate.
[0121] SEQ ID NO:406 represents the amino acid sequence of Hel308 Mbu.
[0122] SEQ ID NO:407 represents the amino acid sequence of Hel308 Csy.
[0123] SEQ ID NO:408 represents the amino acid sequence of Hel308 Tga.
[0124] SEQ ID NO:409 represents the amino acid sequence of Hel308 Mhu.
[0125] SEQ ID NO:410 represents the amino acid sequence of Tral Eco.
[0126] SEQ ID NO:411 represents the amino acid sequence of XPD Mbu.
[0127] SEQ ID NO:412 represents the amino acid sequence of Dda 1993.
[0128] SEQ ID NO:413 represents the amino acid sequence of Trwc Cba.
[0129] SEQ ID NO:414 represents the amino acid sequence of the transporter WP_006819418.1:19-280 [Yokenella regensburgei], which is 91% identical to the sequence of SEQ ID NO:390.
[0130] SEQ ID NO:415 represents the amino acid sequence of WP_024556654.1:16-277 [Cronobacter pulveris] of curli production assembly / transport protein CsgG, which is 89% identical to the sequence of SEQ ID NO:390.
[0131] SEQ ID NO:416 represents the amino acid sequence of YP_005400916.1:16-277 [Rahnella aquatilis HX2] of curli for producing the assembly / transport protein CsgG, which is 84% identical to the sequence of SEQ ID NO:390.
[0132] SEQ ID NO:417 represents the amino acid sequence of KFC99297.1:20-278 [Kluyvera ascorbata ATCC 33433] of the CsgG family curli production assembly / transport component, which is 82% identical to the sequence of SEQ ID NO:390.
[0133] SEQ ID NO:418 represents the amino acid sequence of KFC86716.11:16-274 [Hafnia alvei ATCC 13337] of the CsgG family curli for producing assembly / transport components, which is 81% identical to the sequence of SEQ ID NO:390.
[0134] SEQ ID NO:419 represents the amino acid sequence of YP_007340845.1|:16-270 of the uncharacterized protein involved in the formation of curli polymer [Enterobacteriaceae bacterium strain FGI 57], which is 76% identical to the sequence of SEQ ID NO:390.
[0135] SEQ ID NO:420 represents the amino acid sequence of WP_010861740.1:17-274 of curli production assembly / transport protein CsgG [Plesiomonas shigelloides], which is 70% identical to the sequence of SEQ ID NO:390.
[0136] SEQ ID NO:421 represents the amino acid sequence of YP_205788.1:23-270 of the curli production assembly / transport of outer membrane lipoprotein component CsgG [Vibrio fischeri ES1 14], which is 60% identical to the sequence of SEQ ID NO:390.
[0137] SEQ ID NO:422 represents the amino acid sequence of WP_017023479.1:23-270 of the curli-produced assembly protein CsgG [Aliivibrio logei], which is 59% identical to the sequence of SEQ ID NO:390.
[0138] SEQ ID NO:423 represents the amino acid sequence of WP_007470398.1:22-275 of Curli's production assembly / transport component CsgG [Photobacterium sp. AK15], which is 57% identical to the sequence of SEQ ID NO:390.
[0139] SEQ ID NO:424 represents the amino acid sequence of WP_021231638.1:17-277 of the curli-produced assembly protein CsgG [Aeromonas veronii], which is 56% identical to the sequence of SEQ ID NO:390.
[0140] SEQ ID NO:425 represents the amino acid sequence of Curli producing the assembly / transport protein CsgG WP_033538267.1:27-265 [Shewanella sp. ECSMB14101], which is 56% identical to the sequence of SEQ ID NO:390.
[0141] SEQ ID NO:426 represents the amino acid sequence of WP_003247972.1:30-262 of curli producing the assembled protein CsgG [Pseudomonas putida], which is 54% identical to the sequence of SEQ ID NO:390.
[0142] SEQ ID NO:427 represents the amino acid sequence of YP_003557438.1:1-234 of the curli production assembly / transport component CsgG [Shewanella violacea DSS12], which is 53% identical to the sequence of SEQ ID NO:390.
[0143] SEQ ID NO:428 represents the amino acid sequence of WP_027859066.1:36-280 of curli production of assembly / transport protein CsgG [Marinobacterium jannaschii], which is 53% identical to the sequence of SEQ ID NO:390.
[0144] SEQ ID NO:429 represents the amino acid sequence of CEJ70222.1:29-262 of Curli production assembly / transport component CsgG [Chryseobacterium oranimense) G311], which is 50% identical to the sequence of SEQ ID NO:390.
[0145] SEQ ID NO:430 represents the polynucleotide sequence used in Example 18.
[0146] SEQ ID NO:431 represents the polynucleotide sequence used in Example 18.
[0147] SEQ ID NO:432 represents the polynucleotide sequence used in Example 18.
[0148] SEQ ID NO:433 represents the polynucleotide sequence used in Example 18.
[0149] SEQ ID NO:434 represents the polynucleotide sequence used in Example 18. Attached to the 3' end of SEQ ID NO:434 are six iSp18 spacer regions, and the other end is attached to two thymines and a 3' cholesterol TEG.
[0150] SEQ ID NO:435 represents the amino acid sequence of Step II (C).
[0151] SEQ ID NO:436 represents the amino acid sequence of Pro.
[0152] SEQ ID NO:437 to 440 represent primers from Example 1.
[0153] SEQ ID NO:441 represents the hair clip structure from Example 21.
[0154] SEQ ID NO:442 to 448 indicate origin Figure 4 sequence.
[0155] SEQ ID NO:449 to 450 indicate origin Figure 43 sequence. Detailed Implementation
[0156] Unless otherwise specified, the present invention is carried out using conventional techniques in chemistry, molecular biology, microbiology, recombinant DNA technology, and chemical methods, which are within the capabilities of those skilled in the art. These techniques are also explained in the literature, for example, MR Green, J. Sambrook, 2012, Molecular Cloning: A Laboratory Manual, 4th Edition, Volumes 1-3, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY; Ausubel, FMef al. (1995 and periodically supplemented; Current Protocols in Molecular Biology, Chapters 9, 13 and 16, John Wiley & Sons, New York, NY); B. Roe, J. Crabtree, and A. Kahn, 1996, DNA Isolation and Sequencing: Fundamental Techniques, John Wiley & Sons; J. M. Polak and James O'D. McGee, 1990, In Situ Hybridization: Principles and Practice, Oxford University Press; M. J. Gait (ed.), 1984, Oligonucleotide Synthesis: Practical Methods, IRL Press; and DM. J. Lilley and J. E. Dahlberg, 1992, Enzymatic Methods: DNA Structure Part A: Enzymatic Methods for DNA Synthesis and Physical Analysis, Academic Press. The full contents of the aforementioned documents are included in this paper through citation.
[0157] Before describing this invention, several definitions are provided to aid in understanding it. All references cited herein are incorporated herein by reference in their entirety. Unless otherwise stated, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0158] As used herein, the term "comprising" means that any of the listed elements must be included, and other elements may optionally be included. "consistently composed of" means that any of the listed elements must be included, excluding elements that have a significant impact on the essential and novel characteristics of the listed elements, and other elements may optionally be included. "composed of" means that all unlisted elements are excluded. Embodiments defined by each of these terms are within the scope of the invention.
[0159] As used herein, the term "nucleic acid" refers to a single-stranded or double-stranded covalently linked sequence of nucleotides, wherein the 3' and 5' ends of each nucleotide are linked by a phosphodiester bond. The polynucleotide may consist of deoxyribonucleotide or ribonucleotide bases. Nucleic acids may include DNA and RNA and may be synthesized in vitro or isolated from natural resources. Nucleic acids may further include modified DNA or RNA, such as methylated DNA or RNA, or post-translational modified RNA, such as 5'-capping with 7-methylguanosine, 3'-end processing, such as cleavage and polyadenylation, and splicing. Nucleic acids may also include synthetic nucleic acids (XNAs), such as hexetol nucleic acids (HNAs), cyclohexene nucleic acids (CeNAs), threonine nucleic acids (TNAs), glycerol nucleic acids (GNAs), locked nucleic acids (LNAs), and peptide nucleic acids (PNAs). The size of a nucleic acid (also referred to herein as a "polynucleotide") is typically expressed in terms of the number of base pairs (bp) of a double-stranded polynucleotide, or in the case of a single-stranded polynucleotide, in terms of the number of nucleotides (nts). One thousand bp or nt equals one kilobase pair (kb). Polynucleotides shorter than about 40 nucleotides are often called "oligonucleotides" and can contain primers used in DNA manipulations, such as by polymerase chain reaction (PCR).
[0160] The term "amino acid" as used in this invention is used in its broadest sense, encompassing naturally occurring Lα-amino acids or residues. Naturally occurring amino acids as used herein are generally abbreviated using single or three letters: A = Ala; C = Cys; D = Asp; E = Glu; F = Phe; G = Gly; H = His; I = I1e; K = Lys; L = Leu; M = Met; N = Asn; P = Pro; Q = Gln; R = Arg; S = Ser; T = Thr; V = Val; W = Trp; and Y = Tyr (Lehninger, AL, (1975) Biochemistry, 2nd ed., pp. 71-92, Worth Publishers, New York). The general term "amino acid" also includes D-amino acids, retro-inverso amino acids, and chemically modified amino acids, such as amino acid analogs, naturally occurring amino acids not typically found in proteins, such as ortholeucine, and chemically synthesized compounds possessing amino acid characteristics known in the art, such as β-amino acids. For example, analogues or simulants of phenylalanine or proline—which allow the peptide compound to have the same conformational restrictions as natural Phe or Pro—are included in the definition of an amino acid. Such analogues and simulants are referred to herein as “functional equivalents” of the corresponding amino acid. Other examples of amino acids are cited in: Roberts and Vellaccio, Peptides: Analysis, Synthesis, Biology, Gross and Meiehofer, eds., Vol. 5, p. 341, Academic Press, Inc., NY 1983, which is incorporated herein by reference.
[0161] A "polypeptide" is a polymer of amino acid residues linked by peptide bonds, whether naturally occurring or synthesized in vitro. Polypeptides shorter than about 12 amino acid residues are generally called "peptides," while those between about 12 and 30 amino acid residues are called "oligopeptides." The term "polypeptide" as used herein refers to naturally occurring polypeptides, precursor forms, or preprotein products. Polypeptides can also undergo ripening or post-translational modifications, including but not limited to: glycosylation, protein hydrolysis, lipolysis, signal peptide cleavage, propeptide cleavage, phosphorylation, etc. The term "protein" as used herein refers to a macromolecule containing one or more polypeptide chains.
[0162] "Biopores" refer to transmembrane protein structures that define channels or pores, allowing molecules and ions to migrate from one side of the membrane to the other. The migration of ionic substances through a pore can be driven by applying a potential difference to either side of the pore. Nanopores are a type of biopore in which the minimum diameter of the channel through which molecules or ions pass is on the nanometer scale (10⁻⁶). -9 rice).
[0163] In all embodiments or aspects of the present invention, the polynucleotide may comprise a polynucleotide having at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% complete sequence identity with wild-type *E. coli* CsgG, as shown in SEQ ID NO: 2. Similarly, the polypeptide may comprise a polypeptide having at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% complete sequence identity with wild-type *E. coli* CsgG, as shown in SEQ ID NO: 1. The polypeptide may comprise a polypeptide containing the PFAM domain PF03783, which is characteristic of CsgG-like proteins. A list of currently known CsgG homologs and CsgG architectures can be found at http: / / pfam.xfam.org / / family / PF03783. Sequence identity can therefore be a full-length polynucleotide or a fragment or portion of a polypeptide. Therefore, a sequence may have only 50% overall sequence identity with the sequence of the present invention, but specific regions, domains, or subunits may share 80%, 90%, or up to 99% sequence identity with the sequence of the present invention. According to the present invention, homology to the nucleic acid sequence of SEQ ID NO:2 is not limited to sequence identity. Many nucleic acid sequences can exhibit significant biological homology to each other despite having very low sequence identity. In the present invention, homologous nucleic acid sequences are considered to be heterozygous sequences under low stringency conditions (MR Green, J. Sambrook, 2012, Molecular Cloning: A Laboratory Manual, 4th Edition, Volumes 1-3, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY).
[0164] The term "vector" is used to refer to a linear or circular DNA molecule into which a suitable fragment of another nucleic acid (typically DNA) sequence can be integrated. This DNA fragment may include additional segments prepared for transcription of a gene encoded by the DNA sequence fragment. These additional segments may include, but are not limited to, promoters, transcription terminators, enhancers, internal ribosome entry sites, untranslated regions, polyadenylation signals, selectable markers, origins of replication, etc. A variety of promoters suitable for prokaryotic (e.g., β-lactamase and lactose promoter systems, alkaline phosphatase, tryptophan (trp) promoter systems, *E. coli* lac, tac, T3, T7 promoters) and eukaryotic (e.g., simian virus 40 early or late promoters, Rous sarcoma virus long terminal repeat promoters, cytomegalovirus promoters, adenovirus late promoters, EG-1a promoters) hosts can be used. Expression vectors are typically derived from plasmids, colloids, viral vectors, and yeast artificial chromosomes; vectors are usually recombinant molecules containing DNA sequences from several sources. Specific embodiments of the present invention provide expression vectors encoding wild-type or modified CsgG polypeptides as described herein. The term "operably linked," when applied to DNA sequences, such as in nucleic acid vectors as described above, indicates that the sequences are arranged such that they work synergistically to achieve their intended purpose, i.e., the promoter sequence allows transcription to proceed from the associated coding sequence to the termination sequence.
[0165] The transmembrane protein structure of a biopore can be either a monomer or an oligomer. Typically, the pore comprises multiple polypeptide subunits arranged around a central axis, forming a channel that extends substantially perpendicular to the membrane containing the nanopore residues. The number of polypeptide subunits is not limited. Typically, the number of subunits ranges from 5 to a maximum of 30, with 6 to 10 being suitable. Alternatively, in the case of perfringolysin or related large membrane pores, the number of subunits is not limited. The protein subunit portions within the nanopore forming the protein liner channel typically include secondary structural motifs, which may include one or more transmembrane β-barrel structures and / or α-helical portions.
[0166] It should be understood that different applications of the disclosed products and methods can be adapted to the specific needs of the art. It should also be understood that the terminology used herein is for the purpose of describing specific embodiments of the invention only and is not restrictive.
[0167] Furthermore, as used in this specification and the appended claims, unless otherwise expressly stated, the singular forms “a,” “an,” and “the” include plural references. Thus, for example, reference to “a polynucleotide” includes two or more polynucleotides; reference to “a polynucleotide-binding protein” includes two or more such proteins; reference to “a helicase” includes two or more helicases; reference to “a monomer” refers to two or more monomers; reference to “a pore” includes two or more pores, etc.
[0168] All publications, patents, and patent applications cited in this article, whether above or below, are included in this document by reference in their entirety.
[0169] This invention relates in part to a bacterial amyloid secretion channel (CsgG), its preparation method, and its use in nucleic acid sequencing and molecular sensing.
[0170] CsgG is a membrane lipoprotein present in the outer membrane of *E. coli* (Uniprot accession number P0AEA2; gene ID: 12932538). In this outer lipid membrane, CsgG forms nanopores containing oligomeric complexes of nine CsgG monomeric subunits. The pre-CsgG protein translocates across the SEC translocase via a type II (lipoprotein) signal sequence, followed by triacylation at the N-terminal Cys residue of mature CsgG (i.e., CsgG with a cleaved type II signal sequence). The triacylated or “lipomeric” CsgG is transported to the outer membrane of the Gram-negative host, where it inserts into the bilayer as a nonameric pore. Non-lipomeric forms of CsgG, such as CsgG… C1S It exists in the periplasm as a soluble protein in the prepore conformation. Figure 42 ).
[0171] X-ray structure of wild-type CsgG nanopores (Goyal et al., Nature, 2014, 516(7530), 250-3) shows that they possess... width and The height (in the following text, the term "width" of a nanopore refers to its dimension parallel to the membrane surface, while the term "height" of a nanopore refers to its dimension perpendicular to the membrane). The CsgG pore composite passes through the membrane via 36 strands of β-barrel-like structures to provide... Inner diameter channel ( Figure 1 The assembly monomers of the CsgG channels each have a conserved 12-residue ring structure (C-ring, "CL"); Figure 2 Their combined action creates a channel with a diameter of approximately contraction structure ( Figure 1 and 2The contractile structure of wild-type CsgG nanopores consists of three stacked concentric rings, formed by the side chains of amino acid residues Tyr51, Asn55, and Phe56 of each CsgG monomer present in the CsgG oligomer. Figure 3 The numbering of this residue is based on the mature protein with the natural 15-amino acid signal sequence missing at the N-terminus. Therefore, the mature protein corresponds to residues 16 to 277 of SEQ ID NO:1. Tyre56 is located at position 66 of SEQ ID NO:1, Asn at position 70 of SEQ ID NO:1, and Phe56 at position 71 of SEQ ID NO:1.
[0172] The function of this contractile structure is to restrict the path of ions and other molecules through the CsgG channel. Single-channel current recordings of reconstituted CsgG in a planar phospholipid bilayer under standard electrolyte conditions and potentials of +50 mV or -50 mV resulted in steady-state currents of 43.1 ± 4.5 pA (n = 33) or -45.1 ± 4.0 pA (n = 13), respectively. Figure 5 ).
[0173] Adding a stoichiometric periplasmic factor, CsgE (Uniprot accession number POAE95; Examples 10-12), can effectively block current flow through the CsgG channels. Not wishing to be bound by theory, the evidence from the current points to a mechanism in which CagE forms a complex with the CsgG pores, effectively sealing one end of the channel. A significant reduction in ion flux through the CsgG channels can be detected using standard single-channel recording techniques (Examples 12 and 13). Figure 6 The inventors have discovered that, according to one embodiment of the invention, the measurement parameters of the current (maximum current and the ability to monitor changes in current) make nanopores suitable for applications in nucleic acid sequencing and molecular sensing.
[0174] Therefore, a part of this invention relates to methods and uses of the CsgG nanopore protein complex in nucleic acid sequencing based on changes in electrical measurements of current flowing through nanopores.
[0175] Nucleic acids are particularly well-suited for nanopore sequencing. Naturally occurring nucleic acid bases in DNA and RNA can be distinguished by their physical size. When a nucleic acid molecule or a single base passes through the channel of a nanopore, the size differences between the bases result in a directly correlated reduction in the ion flow through the channel. This change in ion flow can be recorded. Suitable electrical measurement techniques for recording changes in ion flow are described, for example, in WO 2000 / 28312 and D. Stoddart et al., Proc. Natl. Acad. Sci., 2010, 106, pp 7702-7 (single-channel recording devices); and, for example, in WO 2009 / 077734 (multi-channel recording techniques). With appropriate calibration, the characteristic reduction in ion flow can be used to identify specific nucleotides and associated bases passing through the channel in real time.
[0176] The size of the narrowest contraction structure in a transmembrane channel is often a key factor determining the suitability of a nanopore for nucleic acid sequencing applications. If the contraction structure is too small, the molecule to be sequenced will not be able to pass through. However, to maximize the influence on the ion flow through the channel, the channel should not be too large at its narrowest point (i.e., at the contraction structure). Ideally, any contraction structure should be as close as possible in diameter to the size of the bases passing through. For sequencing nucleic acids and nucleic acid bases, suitable contraction structure diameters are in the nanometer range (10⁻⁶). -9 (Meter range). Suitablely, the diameter should be in the range of 0.5 to 1.5 nm, typically in the range of 0.7 to 1.2 nm. The contractile structure in wild-type Escherichia coli CsgG has approximately The diameter is 0.9 nm. The inventors infer that the size and conformation of the contraction structure in the CsgG channel are suitable for nucleic acid sequencing.
[0177] For applications related to nucleic acid sequencing, the CsgG nanopores can be used in wild-type form or can be further modified, for example, through targeted mutations of specific amino acid residues, to further enhance the desired properties of the nanopores used. For example, in embodiments of the invention, mutations are considered to alter the number, size, shape, position, or orientation of contractile structures within the channel. Modified mutant CsgG nanopore complexes can be prepared using known genetic engineering techniques that result in the insertion, substitution, and / or deletion of specific target amino acid residues in the polypeptide sequence. In the case of oligomeric CsgG nanopores, mutations can be performed on each monomeric polypeptide subunit, or on any one or all monomers. Suitably, in one embodiment of the invention, the aforementioned mutations are performed on all monomeric polypeptides within the oligomeric protein structure.
[0178] According to one embodiment of the present invention, a modified mutant CsgG nanopore is provided, wherein the number of channel contraction structures within the pore is reduced.
[0179] Wild-type Escherichia coli CsgG pores include two channel contraction structures (see...) Figure 1 These contractile structures are formed by (i) amino acid residues Phe56 and Asn55, and (ii) amino acid residue Tyr51, as part of a wider structure that includes additional amino acids from position 54 to 53, as well as a C-ring motif. Figure 2 and Figure 3 ).
[0180] In typical nanopore nucleic acid sequencing, as individual nucleotides of the nucleic acid sequence of interest sequentially pass through the channel of the nanopore, the ion flow through the open channel is reduced due to the partial blockage of the channel by the nucleotides. This reduction in ion flow is measured using the appropriate recording technique described above. The reduction in ion flow can be calibrated to the reduction in ion flow measured when a known nucleotide passes through the channel, resulting in a method for determining which nucleotide passes through the channel, and thus a method for determining the nucleotide sequence of the nucleic acid passing through the nanopore when the nucleic acid passes through the channel sequentially. For accurate determination of individual nucleotides, it is generally required that the reduction in ion flow through the channel be directly correlated with the size of the individual nucleotide passing through the single contractile structure (or “read head”). It should be understood that the sequencing can be performed, for example, by the action of a related polymerase, targeting an intact nucleic acid polymer that is “threaded” through the pore. Alternatively, the sequence can be determined by the passage of nucleotide triphosphates that have been sequentially removed from the target nucleic acid near the pore (see, for example, WO 2014 / 187924).
[0181] When two or more contractile structures are present and spaced apart, each contractile structure can simultaneously read or interact with isolated nucleotides in the nucleic acid chain. In this case, the reduction in ion flow through the channel will be a result of the combined limitation of the flows from all the nucleotide-containing contractile structures. Therefore, in some cases, dual contractile structures may lead to complex current signals. In some situations, when two such readouts are present, the current reading for one contractile structure or "readout" may not be determined individually.
[0182] The wild-type CsgG pore structure can be redesigned using recombinant gene technology to enlarge, alter, or remove one of the two contractile structures, leaving a single contractile structure within the channel, thereby defining a single read head. The contractile motif in the CsgG oligomeric pore is located at amino acid residues 38 to 63 of the wild-type *E. coli* CsgG polypeptide. The wild-type amino acid sequence of this region is shown in SEQ ID NO:3. When considering this region, any mutations at positions 50 to 53, 54 to 56, and 58 to 59 are considered to be within the scope of this invention. Based on sequence similarity to CsgG homologs (… Figure 4 The amino acid residues at positions 38 to 49, 53, 57, and 61 to 63 are considered highly conserved and therefore may not be well-suited for substitution or other modifications. Mutations at these positions may be advantageous to modify or alter the characteristics of the readout head due to the critical positioning of the Tyr51, Asn55, and Phe56 side chains within the channel of the wild-type CsgG structure.
[0183] Mutations at a designated position in the monomeric CsgG protein can result in the wild-type amino acid at that position being replaced by any other natural or non-natural amino acid. In one embodiment of the invention, it is desirable to expand or remove a contractile structure; suitably, the amino acid side chains in the modified CsgG protein are selected such that they have less steric hindrance compared to the amino acid side chains in the wild-type structure they replace. The substituted amino acid residues at the designated position can have similar electrostatic properties, or they can have different electrostatic properties. Suitably, the substituted amino acid side chains will have electrostatic charges similar to those in the wild-type structure they replace, to minimize disruption to secondary structure or channel properties.
[0184] The selection of substituted amino acids can be based on the BLOSUM62 matrix, which provides a standard method for calculating the probability of an amino acid being substituted by another based on a large number of multiple sequence alignments. Examples of BLOSUM62 matrices are freely available to those skilled in the art on the Internet; see, for example, the website of the National Center for Biotechnology Information (http: / / ...).
[0185] / / www.ncbi.nlm.nih.gov / Class / Structure / aa / aa_explorer.cgi).
[0186] For Tyr51 in the wild-type CsgG structure, it forms a first contractile structure in the CsgG channel, providing substitution of any amino acid. In particular, in some embodiments of the invention, Tyr51 can be substituted with alanine, glycine, valine, leucine, isoleucine, asparagine, glutamine, and phenylalanine (SEQ ID: 39-318). In embodiments of the invention, substitution of Tyr51 with alanine or glycine is particularly suitable (SEQ ID: 39-108). In embodiments, residues 50 to 53 (PYPA in the wild-type sequence) can be substituted with glycine-glycine (GG) (SEQ ID: 354-388).
[0187] For Asn55—which contributes to the second contractile structure in the CsgG channel—any amino acid substitution is provided. Specifically, in certain embodiments of the invention, Asn55 may be substituted with alanine, glycine, valine, serine, or threonine (SEQ IDs: 9-33, 44-68, 79-103, 14-14, 138, 149-173, 184-208, 219-243, 254-278, 289-313, and 324-348).
[0188] For Phe56—which forms part of the second contractile structure in the CsgG channel—any amino acid substitution is provided. In particular, in certain embodiments of the invention, Phe56 may be substituted with alanine, glycine, valine, leucine, isoleucine, asparagine, and glutamine (SEQ ID NO: 1). ID: 5-13,15-18,20-23,25-28,30-33,40-43,45-48,50-53,55-58,60-63,65-68,75-78,80-83,85-88,90-93,95-98,100-103,110- 113,115-118,120-123,125-128,130-133,135-138,145-148,150-153,155-158,160-163,165-168,170-173,180-183,185-188,190 -193, 195-198, 200-203, 205-208, 215-218, 220-223, 225-228, 230-233, 235-238, 240-243, 250-253, 255-258, 260-263, 265-268, 270-273, 275-578, 285-288, 290-293, 295-298, 300-303, 305-308, 310-313, 320-323, 325-328, 330-333, 335-338, 340-343 and 345-348).In embodiments of the present invention, the substitution of Phe56 with alanine and glycine is particularly suitable (SEQ ID: 5,10,15,20,25,30,35,40,45,50,55,60,65,70,75,80,85,90,95,100,105,1 10,1). 15,120,125,130,135,140,145,150,155,160,165,170,175,180,185,190,195,200,205,210,215,220,225,230,235,240,245,250,255,260,265,270,275,280,285,290,295,300,305,310,315,320,325,330,335,340,345,350,355,360,365,370,375,380,385,6,11,16,21,26,31 ,36,41,46,51,56,61,66,71,76,81,86,91,96,101,106,111,116,121,126,131,136,141,146,151,156,161,166,171,176,181,186,191,196,201,206,211,216,221,226,231,236,241,246,251,256,261,266,271,276,281,286,291,296,301,306,311,316,321,326,331,336,341 346, 351, 356, 361, 366, 371, 376, 381 and 386).
[0189] In a given mutant CsgG protein, the substitution of Tyr51 can occur simultaneously with either position 55 or 56 (SEQ ID: 44,49,54,59,64,79,84,89,94,99,114,119,124,129,134,149,154,159,164,169,184,189,194,199,204,219,224,229,234,239,254,259,264,269,274,289,294,299,304,309,359,364,369,374,379,40,41,42,75,76). 77, 110, 111, 112, 145, 146, 147, 180, 181, 182, 215, 216, 217, 250, 251, 252, 285, 286 and 287), wherein at least one appropriately sized contractile structure is maintained within the channel. Alternatively, the substitution of Tyr51 is mutually exclusive with the substitution at positions Asn55 and Phe56 (SEQ ID: 39, 74, 109, 144, 179, 214, 249, 284, 354, 10, 11, 12, 15, 16, 17, 20, 21, 22, 25, 26, 27, 30, 31 and 32).
[0190] Alternatively, one or more of Tyr51, Asn55, or Phe56 in the wild-type CsgG protein may be missing (SEQ ID: 319-353,34-38,69-73,104-108,139-143,174-178,209-213,244-248,279-283,314-318,384-388,8,13,18,23,28,33,43,48,53,58,63,68,78,83,88,93,98,103,1 13,1). 18,123,128,133,138,148,153,158,163,168,173,183,188,193,198,203,208,218,223,228,233,238,243,253,258,263,268,273,278,288,293,298,303,308 and 313). To maintain at least one contractile structure in the channel, in the given embodiments, the deletion of amino acid residue Tyr51 is mutually exclusive with the deletion of amino acid residues Asn55 and Phe56 (SEQ ID: 319-322, 324-327, 329-332, 334-337, 339-342, 344-347, 38, 73, 108, 143, 178, 213, 248, 283, and 318). Certain adjacent amino acid residues at positions 53 and 54, and 48 and 49, may also be deleted.
[0191] It should be understood that the embodiments provided by the present invention may include modifications made individually or in any combination.
[0192] Removing the contraction structure at Tyr51 or at Asn55 / Tyr56 results in a single contraction structure within the CsgG channel. Not wanting to be bound by theory, it is assumed that the Asn55 / Tyr56 contraction structure has higher conformational stability compared to the likely desired contraction structure at Tyr51. However, the Asn55 / Tyr56 contraction structure may be too high (measured along the axis of the central pore) compared to nucleotides. This could lead to poor resolution of individual base pairs during DNA strand shifts.
[0193] The opposite is likely true for the Tyr51 contraction structure. After removing the Asn55 / Tyr56 contraction structure, the remaining Tyr51 residues in the oligomer may be conformably less stable compared to the native structure. However, the Tyr51 contraction structure is shorter (when measured along the axis of the central pore) and may be more capable of providing a contraction structure that distinguishes individual bases within the channel.
[0194] In any embodiment, when the well is used in nucleic acid sequencing applications, the presence of a single narrow constriction structure (when measured along the central well axis) may reduce the complexity of current reading. The modulation of the observation current generated as nucleic acids shift through the well will therefore reflect only the passage of isolated nucleotides through the single constriction structure or readout head.
[0195] Effective removal of a shrinkage structure can also increase the open channel current of the pore mutant. The increased open channel current will be advantageous because the higher background conductivity leads to higher resolution current hindrance levels for signals targeting different nucleic acid base pairs. In this way, modifications to the readhead can improve the suitability of biopores for nucleic acid sequencing and other molecular sensing applications.
[0196] As an alternative implementation, or in addition to the sequence modifications described above, the Asn55 / Phe56 contractile structure may be further adapted to adjust its height (measured along the central pore axis). This further adjustment of the Asn55 / Phe56 contractile structure may or may not be accompanied by mutations at other locations within the Tyr51 or CsgG channel. Suitably, further adjustments to the Asn55 / Phe56 contractile structure are considered as part of a mutation that expands or removes the contractile structure formed by the Tyr51 residues.
[0197] In the wild-type form, the Asn55 / Phe56 channel contraction structure consists of two amino acid loops arranged perpendicularly to each other. Therefore, the contraction structure has a length greater than 1 nm. A contraction structure of 1 nm in length may not allow for the resolution of the electrical signal generated by the ion flow that separates bases during nucleic acid chain translocation. Typically, the contraction structure of known nanopores used for nucleic acid sequencing is generally less than 1 nm in length. For example, the MspA nanopore used for DNA sequencing has a contraction structure height of 0.6 nm (measured along the central pore axis; Manrao et al., Nature Biotechnology, 2012, 30(4), 349-353).
[0198] To reduce the height of the Asn55 / Phe56 contractile structure (measured along the central hole axis), either of the two residues can be substituted or deleted, resulting in an expansion of the top or bottom of the contractile structure.
[0199] For Asn55, which forms part of the second contractile structure in the CsgG channel, substitution with any amino acid can be considered. In particular, substitution with alanine, glycine, valine, serine, or threonine (SEQ ID: 9-33, 44-68, 79-103, 14-14, 138, 149-173, 184-208, 219-243, 254-278, 289-313, and 324-348).
[0200] For Phe56, which forms part of the second contractile structure in the CsgG channel, substitution with any amino acid can be considered. Specifically, substitution with alanine, glycine, valine, leucine, isoleucine, asparagine, and glutamine (SEQ ID NO: 1) is possible. ID: 5-13,15-18,20-23,25-28,30-33,40-43,45-48,50-53,55-58,60-63,65-68,75-78,80-83,85-88,90-93,95-98,100-103,110- 113,115-118,120-123,125-128,130-133,135-138,145-148,150-153,155-158,160-163,165-168,170-173,180-183,185-188,190 -193, 195-198, 200-203, 205-208, 215-218, 220-223, 225-228, 230-233, 235-238, 240-243, 250-253, 255-258, 260-263, 265-268, 270-273, 275-578, 285-288, 290-293, 295-298, 300-303, 305-308, 310-313, 320-323, 325-328, 330-333, 335-338, 340-343 and 345-348).In embodiments of the present invention, the substitution of Phe56 with alanine and glycine is particularly suitable (SEQ ID: 5,10,15,20,25,30,35,40,45,50,55,60,65,70,75,80,85,90,95,100,105,1 10,1). 15,120,125,130,135,140,145,150,155,160,165,170,175,180,185,190,195,200,205,210,215,220,225,230,235,240,245,250,255,260,265,270,275,280,285,290,295,300,305,310,315,320,325,330,335,340,345,350,355,360,365,370,375,380,385,6,11,16,21,26,31 ,36,41,46,51,56,61,66,71,76,81,86,91,96,101,106,111,116,121,126,131,136,141,146,151,156,161,166,171,176,181,186,191,196,201,206,211,216,221,226,231,236,241,246,251,256,261,266,271,276,281,286,291,296,301,306,311,316,321,326,331,336,341 346, 351, 356, 361, 366, 371, 376, 381 and 386).
[0201] Modification to adjust the minimum diameter of the contraction structure in wild-type CsgG pores is also conceivable. The minimum diameter of the two contraction structures in a CsgG pore is approximately 0.9 nm. Its diameter is smaller than the 1.2 nm diameter of the contractile structure in the MspA nanopores known to be effective for DNA sequencing (Manrao et al., Nature Biotechnology, 2012, 30(4), 349-353). Any of the above mutations that can provide the remaining contractile structure in modified CsgG pores with a minimum diameter of 0.5 to 1.5 nm is suitable.
[0202] Any of the modifications listed above can advantageously alter the hydrophilicity and charge distribution of the amino acids at the contractile structure to improve the passage of the translocated nucleic acid chain and its non-covalent interactions, thereby improving current readout. Any of the mutations listed above can also advantageously alter the hydrophilicity and charge distribution near the channel contractile structure to optimize the flow of electrolyte ions through the contractile structure and achieve better differentiation of nucleotides passing through the translocated nucleic acid chain.
[0203] Further modifications to the CsgG protein are thought to potentially alter the surface charge distribution within the channel lumen. In one embodiment of the invention, these modifications prevent unwanted electrostatic adsorption of translocated nucleic acids onto the channel walls. This is because nucleic acids carry a negative charge, and the CsgG channel lumen contains some positive charge (…). Figure 1 It is assumed that electrostatic interactions may interfere with the helical or translocation of nucleic acids during sequencing. Suitable positively charged amino acid residues, such as lysine, histidine, and arginine, can be replaced by neutral or negatively charged side chains to further improve the efficiency of nucleic acid translocation through the pore, thereby improving the clarity of the current reading.
[0204] Embodiments of the present invention also provide modifications to the channel lumen of wild-type Escherichia coli CsgG or mutant CsgG to facilitate the translocation or helical entry of nucleic acid chains (or individual nucleotides) into the pore contraction structure. The transmembrane portion of CsgG with an internal contraction structure resembles a barrel structure with a cap-like structure characterized by a central pore. Helical entry can be facilitated by adding additional loops within the pore lumen closest to the barrel-like and cap-like structures.
[0205] The present invention provides that the CsgG pore can be composed of one or more covalently linked monomers, dimers, or oligomers. As a non-limiting example, the monomers can be fused in any conformation, for example, through their terminal amino acids. In this case, the amino terminus of one monomer can be fused to the carboxyl terminus of another monomer.
[0206] According to one embodiment of the invention, the CsgG pore is also adapted to accommodate additional accessory proteins that may possess advantageous properties as molecules pass through the pore. Adaptation to the pore may facilitate the anchoring of nucleic acid processing enzymes. Nucleic acid processing enzymes may include DNA or RNA polymerases; isomerases; topoisomerases; helicases; telomerases; and helicases. Associated with one or more of these enzymes to the nanopore may be beneficial in enhancing the entry of nucleic acid helices into the pore and controlling the rate at which nucleic acid chains translocate through the pore (Manrao et al., Nature Biotechnology, 2012, 30(4), 349-353). Controlling the translocation rate of nucleic acid chains through the pore has the advantage of providing an improved response in current measurements of more suitable and consistent ion currents.
[0207] In embodiments of the invention, it is envisioned that modification of the extracellular region of the nanopore membrane may facilitate docking of suitable nucleic acid processing enzymes (e.g., DNA polymerases) within or near the channel lumen by providing one or more binding / anchoring sites. Suitable anchoring sites may include electrostatic patches for enzyme electrostatic binding; one or more cysteine residues allowing covalent coupling; and / or altered internal width of the transmembrane channel portion for providing a spatial anchor.
[0208] The present invention provides further adaptive variations of the CsgG wild-type well for nucleic acid sequencing in the specific embodiments listed below in more detail.
[0209] In an embodiment of the present invention, the extramembrane region of the CsgG pore (e.g. Figure 1 The bottom region (shown) can be truncated or removed to facilitate the departure of nucleic acid chains on the other side of the channel lumen. Truncating or removing the extramembrane region can also improve the current resolution of the electrical signal from the ion flow within the channel. The latter benefit arises from the reduced resistance to the ion flow caused by the extramembrane region. In the CsgG pore of this invention, the transmembrane channel, the internal constriction structure, and the cap region represent three consecutive resistance regions. Removing or reducing the contribution of one of these regions increases the open channel current, thereby improving current resolution. This modification may include the deletion of α-helix 2 (α2), C-terminal α-helix (αC) (… Figure 1 ), and / or combinations thereof.
[0210] In a further embodiment of the invention, the membrane facing the amino acids on the outer surface of the wild-type E. coli CsgG pores can be modified to facilitate insertion of the pores into the membrane. In some embodiments, it is provided that a single amino acid substitution can replace the wild-type residue with a suitable more hydrophobic analog. For example, one or more of residues Ser136, Gly138, Gly140, Ala148, Ala188, or Gly202 can be changed to Ala, Val, Leu, or IIe. Furthermore, aromatic residues such as tyrosine or tryptophan can replace suitable wild-type amino acids such that they are located at the interface between the hydrophobic membrane and the hydrophilic solvent. For example, one or more of residues Leu 154 or Leu182 can be substituted with Tyr, Phe, or Trp.
[0211] In embodiments of the invention, the thermal stability of the protein pores can be increased. This results in a favorable increase in the shelf life of the nanopores in the sequencing device. In embodiments, the increase in the thermal stability of the protein pores is achieved by modifying the β-turn sequences or improving the electrostatic interactions on the protein surface. In one embodiment, β-hairpin structures in the transmembrane region can be stabilized by forming a covalent disulfide across two adjacent β-chains in or between adjacent β-hairpin structures. Examples of such cross-chain cysteine pairs can be Val139-Asp203; Gly139-Gly205; Lys135-Thr207; Glu201-Ala141; Gly147-Gly189; Asp149-Gln187; Gln151-Glu185;
[0212] Thr207-Glu185; Gly205-Gln187; Asp203-Gly189; Ala153-Lys135; Gly137-Gln151;
[0213] Val139-Asp149 or Ala141-Gly147.
[0214] In a further embodiment of the invention, according to the method described in Biotechnol. J., 2011, 6(6), 650-659, the codons for the polypeptide sequence can be modified to allow expression of the CsgG protein at high levels and with a low error rate. The modification can also target any secondary structure of the mRNA.
[0215] This invention also provides protein sequence alterations to improve their stability against proteases. This can be achieved by removing the flexible loop region, for example, by deleting α-helix 2 (α2) or the C-terminal α-helix (αC). Figure 1 ), and / or combinations thereof.
[0216] In embodiments of the present invention, the CsgG polypeptide sequence / expression system is modified to avoid potential protein aggregation.
[0217] This invention also provides for altering polypeptide sequences to improve protein folding efficiency. Suitable techniques are provided in Biotechnol.J., 2011, 6(6), 650-659.
[0218] One embodiment of the invention further provides the replacement or addition of a bioaffinity tag to facilitate the purification of the CsgG protein. The disclosed structure of the CsgG pore includes a Strep1 tag (Goyal et al., Nature, 2014, 516(7530), 250-3). Embodiments of the invention include other bioaffinity tags, such as histidine tags for facilitating purification by metal chelate affinity chromatography. In alternative embodiments of the invention, the tag may include a FLAG tag or an epitope tag, such as a Myc-tag or a HA-tag. In a further embodiment of the invention, the nanopore may be modified by biotinylation with biotin or an analogue of biotin (e.g., dethiobiotin), thereby facilitating purification via interaction with streptavidin.
[0219] In embodiments of the invention, a negative charge can be added to the ends of the protein to increase the net charge of the polypeptide and promote protein migration in polyacrylamide gel electrophoresis. This may lead to improved separation of CsgG heterooligomers in cases of interest to these species (see Howorka et al. Proc. Nat. Acad. Sci., 2001, 98(23), 12996-13001). Heterooligomers can be used to introduce a single cysteine residue into each well, which can facilitate the ligation of suitable nucleic acid processing enzymes, such as DNA polymerases as described above.
[0220] Part of this invention relates to the use of wild-type or modified CsgG nanopores in molecular sensing applications based on changes in electrical measurements of current flowing through the nanopores.
[0221] The binding of molecules within or near any opening of the CsgG pore affects the open-channel ion flow through the pore. Similar to the nucleic acid sequencing applications described above, changes in open-channel ion flow can be measured by current variations using suitable measurement techniques (e.g., WO 2000 / 28312 and D. Stoddart et al, Proc. Natl. Acad. Sci., 2010, 106, 7702-7 or WO 2009 / 077734). The degree of reduction in ion flow (measured by a decrease in current) is related to the size of the blockage within or near the pore. Therefore, the binding of molecules of interest, also known as analytes, within or near the pore provides a detectable and measurable event, forming the basis of a biosensor. Suitable molecules for nanopore sensing include nucleic acids; proteins; peptides; and small molecules such as drugs, toxins, or cytokines.
[0222] The detection of the presence of biomolecules can be applied to personalized medicine development, pharmaceuticals, diagnostics, life science research, environmental monitoring, and the safety and / or defense industries.
[0223] In embodiments of the present invention, the wild-type or modified *E. coli* CsgG nanopores or homologs disclosed in this application can be used as molecular sensors. The analyte detection process is described in Howorka et al. *Nature Biotechnology* (2012) Jun 7; 30(6) 506-7. The analyte molecule to be detected can bind to either side of the channel or within the lumen of the channel itself. The binding site can be determined by the size of the molecule to be sensed. The wild-type CsgG pore can be used as a sensor, or, in embodiments of the present invention, the CsgG pore can be modified by recombination or chemical methods to increase binding strength, binding site, or binding specificity of the molecule to be sensed. Typical modifications include the addition of a specific binding moiety that is structurally complementary to the molecule to be sensed. When the analyte molecule contains nucleic acids, this binding moiety can contain cyclodextrins or oligonucleotides; for small molecules, this can be a known complementary binding region, such as the antigen-binding portion of an antibody or non-antibody molecule, including a single-stranded variable fragment (scFv) region or antigen recognition domain from a T-cell receptor (TCR); or for proteins, it can be a known ligand of the target protein. In this way, wild-type or modified E. coli CsgG nanopores or their homologs can be used as molecular sensors to detect the presence of suitable antigens (including epitopes) in a sample, including cell surface antigens, such as receptors, markers of solid tumors or blood cancer cells (e.g., lymphoma or leukemia), viral antigens, bacterial antigens, protozoan antigens, allergens, allergy-related molecules, albumins (e.g., human, rodent, or bovine), fluorescent molecules (including luciferin), blood group antigens, small molecules, drugs, enzymes, catalytic sites of enzymes or enzyme substrates, and transition state analogs of enzyme substrates.
[0224] Modification can be achieved using known genetic engineering and recombinant DNA techniques. The localization of any adaptation will depend on the properties of the molecule to be sensed, such as its size, three-dimensional structure, and biochemical properties. The selection of the adaptive structure can be achieved using computer-aided structural design. A series of customized CsgG nanopores have been envisioned, each specifically suited to its intended sensing application. The determination and optimization of protein-protein interactions or protein-small molecule interactions can be achieved using, for example... The study utilizes techniques such as surface plasmon resonance to detect molecular interactions (BIAcore, Inc., Piscataway, NJ; see also www.biacore.com).
[0225] In one embodiment, the CsgG pore can be in the form of a water-soluble octamer, wherein the N-terminal Cys residue is replaced by another amino acid to prevent lipidation of the protein's N-terminus. In an alternative embodiment, the protein can be expressed in the cytoplasm by removing the N-terminal leader sequence to avoid processing via the bacterial lipidation pathway.
[0226] Goyal et al. (Nature, 2014; 516(7530)250-3) described methods for preparing CsgG monomeric soluble protein, octamer soluble protein and oligolipidized CsgG pores, which are incorporated herein by reference and in Examples 1 and 2.
[0227] Mutant CsgG monomer
[0228] This invention provides a mutant CsgG monomer. The mutant CsgG monomer can be used to form the pores described in this invention. A mutant CsgG monomer is a monomer whose sequence differs from that of a wild-type CsgG monomer and retains the ability to form pores. Methods for confirming the pore-forming ability of mutant monomers are well known in the art and will be discussed in more detail below.
[0229] The mutant monomers exhibit enhanced polynucleotide readout properties, demonstrating improved polynucleotide capture and nucleotide recognition. Specifically, pores constructed from mutant monomers capture nucleotides and polynucleotides more readily than those constructed from wild-type monomers. Furthermore, pores constructed from mutant monomers exhibit an increased current range (making it easier to distinguish different nucleotides) and reduced state changes (increasing the signal-to-noise ratio). Additionally, the number of nucleotides contributing to the current is reduced as polynucleotides move through the pores constructed from mutant monomers. This makes it easier to identify a direct relationship between the observed current and the polynucleotide sequence when polynucleotides move through the pores. Moreover, pores constructed from mutant monomers can exhibit increased flux, meaning they are more likely to interact with analytes such as polynucleotides. This makes it easier to characterize analytes using the pores. Pores constructed from mutant monomers can be more easily inserted into membranes.
[0230] The mutant monomers of this invention comprise variants of the sequence shown in SEQ ID NO:390. SEQ ID NO:390 is a wild-type CsgG monomer from the Escherichia coli Str.K-12 substrain MC4100. Variants of SEQ ID NO:390 are polypeptides having an amino acid sequence different from that of SEQ ID NO:390 and retaining their ability to form pores. The ability of a variant to form pores can be determined using any method known in the art. For example, the variant can be inserted into an amphiphilic molecular layer along with other suitable subunits, and the ability of the variant to oligomerize and form pores can be determined. Methods for inserting subunits into membranes, such as amphiphilic molecular layers, are known in the art. For example, the subunits can be suspended in a purified form in a solution containing a triblock copolymer membrane, allowing them to diffuse into the membrane and be inserted by binding to the membrane and assembling into a functional state.
[0231] Throughout the discussion of this application, standard single-letter codes are used for amino acids. These codes are as follows: alanine (A), arginine (R), asparagine (N), aspartic acid (D), cysteine (C), glutamic acid (E), glutamine (Q), glycine (G), leucine (L), lysine (K), methionine (M), phenylalanine (F), proline (P), serine (S), threonine (T), tryptophan (W), and valine (V). Standard substitution symbols are also used, i.e., Q42R indicates that the Q at position 42 is replaced by R.
[0232] In one embodiment, the mutant monomer of the present invention comprises a mutant of SEQ ID NO:390, which includes one or more of the following: (i) one or more mutations at the following positions (i.e., mutations at one or more of the following positions): N40, D43, E44, S54, S57, Q62, R97, E101, E124, E131, R142, T150, and R192, such as one or more mutations at the following positions (i.e., mutations at one or more of the following positions): N40, D43, E44, S54, S57, Q62, E101, E131, and T150 or N4 (ii) mutations in Y51N55, Y51F56, N55F56, or Y51N55F56; (iii) Q42R or Q42K; (iv) K49R; (v) N102R, N102F, N102Y, or N102W; (vi) D149N, D149Q, or D149R; (vii) E185N, E185Q, or E185R; (viii) D195N, D195Q, or D195R; (ix) E201 N, E201 Q or E201 R; (x) E203N, E203Q or E203R; and (xi) missing one or more of the following positions: F48, K49, P50, Y51, P52, A53, S54, N55, F56 and S57. The variants may contain any combination of (i) to (xi). In particular, the variants may include {i}{ii}{iii}{iv}{v}{vi}
[0233] {vii}{viii}{ix}{x}{xi}{i,ii}{i,iii}{i,iv}{i,v}{i,vi}{i,vii}{i,viii}{i,ix}{i,x}{i,xi}{ii,iii}
[0234] {ii,iv}{ii,v}{ii,vi}{ii,vii}{ii,viii}{ii,ix}{ii,x}{ii,xi}{iii,iv}{iii,v}{iii,vi}{iii,vii}{iii,viii}
[0235] {iii,ix}{iii,x}{iii,xi}{iv,v}{iv,vi}{iv,vi}{iv,viii}{iv,ix}{iv,x}{iv,xi}{v,vi}{v,vii}{v,vii}{v,ix}{v,x}{v,xi}{vi,vi}{vi,vi {vi,xi}{vii,viii}{vii,xi}{vii,x}{vii,xi}{viii,ix}{viii,x}{viii,xi}{ix,x}{ix,xi}{x,xi}{i,ii,iii}{i,ii,iv}{i,ii,v}{i,ii,vi}{i,ii,vii}{i,ii,viii}{i,ii,viii}
[0236] {i,ii,ix}{i,ii,x}{i,ii,xi}{i,iii,v}{i,ii,v}{i,iii,vi}{i,iii,vii}{i,iii,viii}{i,iii,ix}{i,iii,x}{i,iii ,xi}{i,v,v}{i,v,v}{i,v,vi}{i,v,vii}{i,v,ix}{i,v,x}{i,v,xi}{i,v,vi}{i,v,vii}{i,v,viii}{i,v,ix}
[0237] {i,v,x}{i,v,xi}{i,v,vi}{i,v,viii}{i,vi,ix}{i,v,x]{i,v,xi}{i,vi,vii}{i,vi,ix}{i,vi,x}{i,vi,xi}{i,viii,ix}{i,viii,x}{i,viii,xi}{i,iii,x}{i,viii,xi}{i,xi,x}{i,iii,xi}{i,x,xi}{ii,iii,iv}{ii,iii,v}{ii,iii,vi]{ii,iii,vii}{ii,iii,vii}
[0238] {ii,iii,vii}{ii,iii,ix}{ii,iii,x}{ii,iii,xi}{ii,iv,v}{ii,ii,vi}{ii,iv,vii} {ii,v,viii}{ii,v,ix}{ii,v,x}{ii,iv,xi}{ii,v,v]{ii,v,vii}{ii,v,viii}{ii, v,ix}{ii,v,x}{ii,v,xi}{ii,v,vii}{ii,v,viii}{ii,v,ix}{ii,v,x}{ii,v,xi}{ ii,vi,vii}{ii,vi,ix}{ii,vi,x]{ii,vi,xi}{ii,vii,ix}{ii,vii,x}{ii,viii, xi}{ii,ix,x}{ii,ix,xi}{ii,x,xi}{iii,iv,v}{iii,iv,vi}{iii,iv,vii}{iii,iv,vii i}{iii,iv,ix}{iii,iv,x}{iii,v,xi}{iii,v,v}{iii,v,vii}{iii,v,viii}{iii,v,ix}{iii,v,x}{iii,v,xi}{iii,v,vii}{iii,v,vii}{iii,v,ix}{iii,v,x}{iii,v,x i}{iii,vi,vii}{iii,vi,ix}{iii,vi,x}{iii,vi,xi}{iii,viii,ix}{iii,viii,x}{iii,vii,x}
[0239] {i,ii,iii,iv}{i,ii,iii,v}{i,ii,iii,vi}{i,ii,iii,vi}{i,ii,iii,viii]{i,ii,iii,ix}{i,ii,iii,x}{i,xi,iii i}{i,i,v,v}{i,i,v,v}{i,i,i,v}{i,i,i,v,vii}{i,i,i,iv,ix}{i,i,v,x}{i,i,iv,xi}{i,i,v,v}{i,i,i, v,vii}{i,ii,v,vii}{i,ii,v,ix}{i,i,v,x}{i,i,v,xi}{i,i,vi,vii}{i,ii,vi,vii}{i,i,vi, ix}{i,ii,vi,x}{i ,ii,vi,xi}{i,ii,vi,viii}{i,ii,vii,xi}{i,i,vi,x}{i,ii,vi,xi}{i,ii,viii,ix}{i,ii,viii,x}{i,ii,viii,xi}{i,ii,vii,xi}
[0240] {i,ii,ix,x}{i,ii,ix,xi}{i,ii,x,xi}{i,ii,ii,v,v}{i,iii,iv,vi}{i,iii,iii,vii]{i,iii,iii,iv,viii}{i,iii,iv,ix}
[0241] {i,iii,v,x}{i,iii,iv,xi}{i,iii,v,v}{i,iii,v,vii}{i,iii,v,viii}{i,iii,v,ix}{i,iii,v,x}{i,iii,v,xi}
[0242] {i,iii,vi,vi]{i,iii,vi,viii}{i,iii,vi,ix}{i,iii,vi,x}{i,iii,vi,xi}{i,iii,vii,vii}{i,iii ,vii,ix}{i,iii,vi,x}{i,iii,vi,xi}{i,iii,viii,ix}{i,iii,viii,x}{i,iii,viii,xi}{i,iii,ix,x}{i,iii,ix,xi}{i,iii,x,xi}{i,iv,v,vi}{i,iv,v,vii}{i,iv,v,viii}{i,iv,v,ix}{i,iv,v,x}{i,iv,v ,xi}{i,v,v,vi}{i,v,v,vii}{i,v,v,ix}{i,v,v,x}{i,v,v,xi}{i,v,v,viii}{i,v,vii ,ix}{i,iv,vii,x}{i,iv,vii,xi}{i,iv,viii,ix}{i,iv,viii,x}{i,iv,viii,xi}{i,iv,ix,x}{i,iv,ix ,xi}{i,v,x,xi}{i,v,v,v}{i,v,v,viii}{i,v,v,ix}{i,v,v,x}{i,v,v,xi}{i,v,vii,viii}{i,v ,vii,ix}{i,v,vii,x}{i,v,vii,xi}{i,v,viii,ix}{i,v,viii,x}{i,v,viii,xi}{i,v,ix,x}{i,v,ix,xi} {i,v,x,xi}{i,v,vi,vi,vii}{i,v,vi,ix}{i,v,vi,x]{i,v,vi,xi}{i,v,viii,ix}{i,v,viii,x}
[0243] {i,vi,viii,xi}{i,vi,x,x}{i,vi,xi,xi}{i,vi,x,xi}{i,vii,viii,ix}{i,vii,viii,x}{i,vii,viii,xi]{i,vii,ix,x}
[0244] {i,vii,x,xi}{i,vii,x,xi}{i,viii,ix,x}{i,viii,ix,xi}{i,viii,x,xi}{i,ix,x,xi}{ii,i ii,f,v}{ii,iii,iii,v,v}{ii,iii,iv,vii]{ii,iii,iv,viii}{ii,iii,iv,ix}{ii,iii,iv,x}{i i,iii,iv,xi}{ii,iii,v,v}{ii,iii,v,vii}{ii,iii,v,viii}{ii,iii,v,ix}{ii,iii,v,x}{ii ,iii,v,xi}{ii,iii,vi,vi}{ii,iii,vi,viii}{ii,iii,vi,ix}{ii,iii,vi,x}{ii,iii,vi,xi}
[0245] {ii,iii,vi,vii}{ii,iii,vi,ix}{ii,iii,vi,x}{ii,iii,vii,xi}{ii,iii,viii,ix]{ii,iii,viii,x}{ii,iii,viii,xi}
[0246] {ii,iii,x,x}{ii,iii,ix,xi}{ii,iii,x,xi}{ii,iv,v,vi}{ii,iv,v,vii}{ii,iv,v,viii}{ii,iv,v,ix}{ii,iv,v,x}
[0247] {ii,v,v,xi}{ii,v,v,v}{ii,v,v,v}{ii,v,v,ix}{ii,v,v,x}{ii,v,v,xi}{ii,v,v,v,v}{ii,v,v,ix}
[0248] {ii,iv,vi,x}{ii,iv,vii,xi}{ii,iv,viii,ix}{ii,iv,viii,x}{ii,iv,viii,xi}{ii ,iv,ix,x}{ii,iv,ix,xi}{ii,iv,x,xi}{ii,v,v,vii}{ii,v,v,viii}{ii,v,vi,ix}{i i,v,v,x}{ii,v,v,xi}{ii,v,vi,vii}{ii,v,vi,ix}{ii,v,vii,x}{ii,v,vi,xi} {ii,v,viii,ix}{ii,v,viii,x}{ii,v,viii,xi}{ii,v,ix,x}{ii,v,ix,xi}{ii,v,x,xi}
[0249] {ii,vi,vi,vii}{ii,vi,vii,ix}{ii,vi,vi,x}{ii,vi,vi,xi}{ii,vi,viii,ix}{ii,vi,viii,x}{ii,vi,vii,xi}{ii,vi,vii,xi}
[0250] {ii,vi,x,x}{ii,vi,ix,xi}{ii,vi,x,xi}{ii,vi,vii,ix}{ii,vii,viii,x}{ii,vi,vii,xi}{ii,vii,ix,x}
[0251] {ii,vii,x,xi}{ii,vii,x,xi}{ii,viii,ix,x}{ii,vii,ix,xi}{ii,vii,x,xi}{ii,ix,x,xi}{iii,iv,v,vi}{iii,iv,v,vii}{iii,iv,v,viii}{iii,iv,v,ix}{iii,iv,v,x}{iii,iv,v,xi}{ iii,iv,vi,vi}{iii,iv,v,viii]{iii,iv,v,ix}{iii,iv,v,x}{iii,iv,v,xi}{iii,iv,vi i,viii}{iii,v,v,vi,ix}{iii,iv,vii,x}{iii,iv,vi,xi}{iii,iv,viii,ix}{iii,iv,viii,x}
[0252] {iii,v,viii,xi}{iii,iv,ix,x}{iii,iv,ix,xi}{iii,iv,x,xi}{iii,v,v,vii}{iii,v,v,viii}{iii,v,v,ix}{iii,v ,vi,x}{iii,v,v,xi}{iii,v,vii,viii]{iii,v,vii,ix}{iii,v,vii,x}{iii,v,vii,xi}{iii,v,viii,ix}{iii,v,vii,x}{iii,v,vii,x}
[0253] {iii,v,vii,xi}{iii,v,ix,x}{iii,v,ix,xi}{iii,v,x,xi}{iii,vi,vii,viii}{iii,vi,vii,ix}{iii,vi,vii,x}{iii,vi,vi,x}
[0254] {iii,vi,vi,xi}{iii,vi,viii,xi}{iii,vi,viii,x}{iii,vi,viii,xi}{iii,vi,ix,x}{iii,vi,xi,xi}{iii,vi,x,xi]{iii,vi,x,xi]
[0255] {iii,vii,viii,x}{iii,vii,vii,x}{iii,vii,viii,xi}{iii,vii,ix,x}{iii,vii,ixi,xi}{iii,vii,x,xi}{iii,viii,ix,x}
[0256] {iii,viii,ix,xi}{iii,viii,x,xi]{iii,ix,xi}{iv,v,vi,vii}{iv,v,v,viii}{iv,v,vi,ix}{iv,v,vi,x}{iv,v, vi,xi}{iv,v,vii,viii}{iv,v,vii,ix}{iv,v,vii,x}{iv,v,vii,xi}{iv,v,viii,ix}{iv,v,viii,x}{iv,v,viii,xi}{iv,v,viii,xi}
[0257] {iv,v,ix,x}{iv,v,ix,xi}{iv,v,x,xi}{iv,v,vii,viii}{iv,v,vii,ix}{iv,v,vii,x]{iv,vi,vii,xi}{iv,vi,vii i,ix}{iv,vi,viii,x}{iv,vi,viii,xi}{iv,vi,ix,x}{iv,vi,xi,xi}{iv,vi,x,xi}{iv,vii,viii,ix}{iv,vii,viii,x]
[0258] {iv,vii,viii,xi}{iv,vii,x,x}{iv,vii,ix,xi}{iv,vii,x,xi}{iv,viii,ix,x}{iv,viii,xi,xi}{iv,viii,x,xi}
[0259] {iv,ix,x,xi}{v,vi,vi,vii,viii]{v,vi,vii,ix}{v,vi,vii,x}{v,vi,vii,xi}{v,vi,viii,ix}{v,vi,viii,x}
[0260] {v,vi,viii,xi}{v,vi,ix,x}{v,vi,ix,xi}{v,vi,x,xi}{v,vii,viii,ix]{v,vii,viii,x}{v,vii,viii,x i}{v,vii,ix,x}{v,vii,ix,xi}{v,vii,x,xi}{v,viii,ix,x}{v,viii,ix,xi}{v,viii,x,xi}{v,ix,x,xi}{v,ix,x,xi}
[0261] {vi,vii,viii,ix}{vi,vii,viii,x}{vi,vii,vii,xi}{vi,vii,ix,x}{vi,vii,xi,xi}{vi,vii,x,xi}{vi,viii,ix,x}
[0262] {vi,viii,x,xi}{vi,viii,x,xi}{vi,ix,x,xi}{vii,viii,ix,x}{vii,viii,ix,xi}{vii,viii,x,xi}{vii,ix,x,xi}{vii,ix,x,xi}
[0263] {viii,ix,x,xi}{i,ii,iii,v,v}{i,ii,ii,iii,v,v}{i,ii,iii,iii,iii}{i,ii,iii,iv,viii}{i,ii,iii,iii,iii,viii}{i,ii,ii,iii,iii, ix}{i,ii,xiii,iii,iv
[0264] {i,ii,iii,v,xi}{i,ii,iii,v,v}{i,ii,iii,v,vii}{i,ii,iii,v,viii}{i,ii,iii,v,ix}{i,ii,iii,v,x]{i,ii,xi}iii
[0265] {i,ii,iii,v,vii}{i,ii,iii,vi,vii}{i,ii,iii,vi,ix}{i,ii,iii,vi,x}{i,ii,iii,vi,vi,xi}{i,ii,iii,vi,viii}{iii,vii,ii
[0266] {i,ii,iii,vi,x}{i,ii,iii,vi,xi}{i,ii,iii,vii,ix}{i,ii,iii,viii,x}{i,ii,ii,iii,viii,xi}{i,ii,iii, ixi,xi,xi}iiiiiii
[0267] {i,ii,iii,x,xi}{i,ii,v,v,v}{i,ii,v,v}{i,ii,iv,v,viii}{i,ii,v,v,ix}{i,ii,v,v,x}{i,ii,v,v,xi}{i,ii,v,v,xi}
[0268] {i,ii,v,v,v}{i,i,v,v,v}{i,i,v,v,v}{i,i,v,v,ix}{i,i,v,v,x}{i,i,v,v,xi}{i,i,i,v,v,v,v}{i,ii,v,v,ix}
[0269] {i,ii,iv,vii,x}{i,ii,iv,vii,xi}{i,ii,iv,viii,ix}{i,ii,iv,viii,x}{i,ii,iv,viii,xi}{i,ii,iv,ix,x}{i,ii,iv,ix,xi]
[0270] {i,ii,v,x,xi}{i,i,v,v,vi}{i,i,v,v,vii}{i,i,v,v,ix}{i,i,v,v,x}{i,i,v,v,vi,xi}{i,i,v,vi,vii}{i,i,v,vi,vii}
[0271] {i,ii,v,vi,ix}{i,ii,v,vi,x}{i,ii,v,vi,xi}{i,ii,v,viii,ix}{i,ii,v,viii,x}{i,ii,v,viii,xi}{i,ii,v,ix,x}
[0272] {i,ii,v,ix,xi}{i,ii,v,x,xi}{i,ii,vi,vi,viii}{i,ii,vi,vii,ix}{i,ii,vi,vii,x}{i,ii,vi,vi,xi}{i,ii,vi,viii,ix}
[0273] {i,ii,vi,viii,x}{i,ii,vi,viii,xi}{i,ii,vi,x,x}{i,ii,vi,ii,xi}{i,ii,vi,x,xi}{i,ii,vii,viii,ix]{i,ii,vii,viii,x}{i,ii,vii,viii,x}
[0274] {i,ii,vii,viii,xi}{i,ii,vii,ix,x}{i,ii,vii,xi,xi}{i,ii,vii,x,xi}{i,ii,viii,ix,x}{i,ii,viii,xi,xi}{i,ii,viii,x,xi}{i,ii,vii,x,xi}
[0275] {i,ii,ix,x,xi}{i,iii,v,v,vi}{i,iii,v,v,vii}{i,iii,iv,v,viii}{i,iii,iv,v,ix}{i,iii,iv,v,x}{i,iii,iv,v,v,xi}{i,iii,iv,v,xi}
[0276] {i,iii,v,v,v}{i,iii,v,v,vii}{i,iii,v,v,ix}{i,iii,v,v,x}{i,iii,v,v,xi}{i,iii,iv,v,viii}
[0277] {i,iii,v,vi,x}{i,iii,v,vi,x}{i,iii,v,vi,xi}{i,iii,iv,viii,ix}{i,iii,iv,vii,x}{i,iii,iv,vii,xi}{i,iii,iv,vii,xi}
[0278] {i,iii,iv,ix,x}{i,iii,v,ix,xi}{i,iii,iv,x,xi}{i,iii,v,v,vii}{i,iii,v,vi,viii}{i,iii,v,vi,ix}{i,iii,v,vi,x}
[0279] {i,iii,v,vi,xi]{i,iii,v,vi,vii}{i,iii,v,vi,ix}{i,iii,v,vii,x}{i,iii,v,vi,xi}{i,iii,v,viii,ix}{i,iii,iii,}v,vi
[0280] {i,iii,v,viii,xi}{i,iii,v,ix,x}{i,iii,v,ix,xi}{i,iii,v,x,xi}{i,iii,vi,vii,viii}{i,iii,vi,vi, ix}{i,iii,vi,vi,vii
[0281] {i,iii,vi,vii,xi}{i,iii,vi,viii,xi}{i,iii,vi,viii,x}{i,iii,vi,viii,xi}{i,iii,vi, ix,x}{i,iii,vi,ixi,xi}{i,iii,vi ,x,xi}{i,iii,vii,viii,ix}{i,iii,vii,viii,x}{i,iii,vii,viii,xi}{i,iii,vii,ix,x}{i,iii,vii,xi,xi}{i,iii,vii,x,xi}
[0282] {i,iii,viii,ix,x}{i,iii,viii,ix,xi}{i,iii,viii,x,xi}{i,iii,iii,x,xi}{i,iv,v,vi,vi}{i,iv,v,v,viii}{i,iv,v,vi ,ix}{i,v,v,v,x}{i,v,v,v,xi}{i,v,v,v,vii}{i,v,v,vii,ix}{i,v,v,vii,x}{i,v,v,v,vi,xi}{i,v,v,viii,ix}
[0283] {i,v,v,viii,x}{i,iv,v,vii,xi]{i,iv,v,v,vii,xi]{i,iv,v,ix,x}{i,iv,v,ix,xi}{i,iv,v,x,xi}{i,iv,vi,vii,viii}{i,iv,v,vii,xi}{i,iv,v,vii,xi]
[0284] {i,v,v,vi,x}{i,v,v,v,xi}{i,v,v,viii,ix}{i,v,v,viii,x}{i,v,v,viii,xi}{i,v,v,ix,x}{i,v,v, ix,xi}{i,iv,vi,x,xi}{i,iv,vii,viii,ix}{i,iv,vii,viii,x}{i,iv,vii,viii,xi}{i,iv,vii,ix,x}{i,iv,vii,ix,xi}{i,iv,vii,vii,xi}
[0285] {i,iv,vii,x,xi}{i,iv,viii,ix,x}{i,iv,viii,xi,xi}{i,iv,viii,x,xi}{i,iv,ix,x,xi}{i,v,vi,vii,viii}{i,v,vi,vii ,ix}{i,v,vi,vi,vi,x}{i,v,vi,vi,xi}{i,v,vi,viii,ix}{i,v,vi,viii,x}{i,v,vi,viii,xi}{i,v,vi,ix,x}{i,v,vi,ix,xi}{i,v,vi,ix,xi}
[0286] {i,v,vi,x,xi}{i,v,vii,viii,ix}{i,v,vii,viii,x}{i,v,vii,viii,xi}{i,v,vii,ix,x}{i,v,vii,xi,xi}{i,v,vii,x,xi}{i,v,vii,x,xi}
[0287] {i,v,viii,ix,x}{i,v,viii,ix,xi}{i,v,viii,x,xi}{i,v,ix,x,xi}{i,vi,vii,viii,ix}{i, vi,vi,vii,x}{i,vi,vi,viii,xi}{i,vi,vii,ix,x}{i,vi,vii,ix,xi}{i,vi,vii,x,xi}{i ,vi,viii,ix,x}{i,vi,viii,ix,xi}{i,vi,viii,x,xi}{i,vi,ix,x,xi}{i,vii,viii,ix,x}{i ,vii,viii,x,xi}{i,vii,viii,x,xi}{i,vii,ix,x,xi}{i,viii,ix,x,xi}{ii,iii,iv,v,vi}{ii,iii,iv,v,vi}
[0288] {ii,iii,v,v,vii}{ii,iii,v,v,viii}{ii,iii,v,v,ix}{ii,iii,v,v,x}{ii,iii,v,v,xi}{ii,iii,v,v,vii}{ii,iii,v,v,vii}
[0289] {ii,iii,v,v,viii}{ii,iii,v,v,ix}{ii,iii,v,v,x}{ii,iii,v,v,x}{ii,iii,iii,v,xi}{ii,iii,v,vi,vii}{ii,iii,iv,vi,ix}
[0290] {ii,iii,v,vi,x}{ii,iii,v,vi,xi}{ii,iii,iv,viii,ix}{ii,iii,iii,iv,viii,x}{ii,iii,iv,viii,xi}{ii,iii,iv,ix,x}
[0291] {ii,iii,v,x,xi}{ii,iii,v,x,xi}{ii,iii,v,v,vi}{ii,iii,v,v,viii}{ii,iii,v,vi,ix}{ii,iii,v,v,x}{ii,,iii,v vi,xi}{ii,iii,v,vi,vii}{ii,iii,v,vi,iii}{ii,iii,v,vii,x}{ii,iii,v,vii,xi}{ii,iii,v,viii,ix}{ii,iii,v},x,viii
[0292] {ii,iii,v,vii,xi}{ii,iii,v,ix,x}{ii,iii,v,ix,xi}{ii,iii,v,x,xi}{ii,iii,vi,vii,vii}{ii,iii,vi,vi,ix}{ii,iii,vi,vi,ix}
[0293] {ii,iii,vi,vi,x}{ii,iii,vi,vi,xi}{ii,iii,vi,viii,ix}{ii,ii,iii,vi,viii,x}{ii,iii,vi,viii,xi}{ii,iii,vi,ix,x}{ii,iii,vi,ix,x}
[0294] {ii,iii,vi,x,xi}{ii,iii,vi,x,xi}{ii,iii,vi,viii,ix}{ii,ii,iii,vii,viii,x}{ii,iii,vii,viii,xi}{ii,iii,vii,ix,x}
[0295] {ii,iii,vii,x,xi}{ii,iii,vi,x,xi}{ii,iii,viii,ix,x}{ii,ii,iii,viii,xii,xi}{ii,iii,viii,x,xi}{ii,iii,ix,x,xi}
[0296] {ii,v,v,v,v}{ii,v,v,v,v}{ii,v,v,v,ix}{ii,v,v,v,x}{ii,v,v,v,xi}{ii,v,v,v,v,v}{ii,v,v ,vii,ix}{ii,iv,v,vii,xi}{ii,iv,v,viii,ix}{ii,iv,v,v,viii,x}{ii,iv,v,viii,xi}{ii,iv,v,ix,x}
[0297] {ii,iv,v,ix,xi}{ii,v,v,x,xi}{ii,iv,v,v,viii}{ii,iv,v,vii,ix}{ii,iv,v,vii,x}{ii,iv,v,vi,xi}{ii,iv,v,vi,xi}
[0298] {ii,iv,v,viii,ix}{ii,iv,v,viii,x}{ii,iv,v,viii,xi}{ii,iv,v,ix,x}{ii,iv,vi,ix,xi}{ii,iv,v,x,xi}{ii,iv,v,x,xi}
[0299] {ii,iv,vii,viii,ix}{ii,iv,vii,viii,x}{ii,iv,vii,viii,xi}{ii,iv,vii,ix,x}{ii,iv,vii,ix,xi}{ii,iv,vii,x,xi}{ii,iv,vii,x,xi}
[0300] {ii,iv,viii,ix,x}{ii,v,viii,ix,xi}{ii,iv,viii,x,xi}{ii,iv,ix,x,xi}{ii,v,vi,vii,viii}{ii,v,v,vii,ix}{ii, v,vi,vi,x}{ii,v,v,vi,vi,xi}{ii,v,vi,viii,ix}{ii,v,vi,viii,x}{ii,v,vi,viii,xi}{ii,v,vi,vi,x}{ii,v,vi,ix,xi}{ii,v,vi,ix,xi}
[0301] {ii,v,v,x,xi}{ii,v,vii,viii,ix}{ii,v,vii,viii,x}{ii,v,vii,viii,xi}{ii,v,vii,ix ,x}{ii,v,vii,x,xi}{ii,v,vii,x,xi}{ii,v,viii,ix,x}{ii,v,viii,ix,xi}{ii,v,viii,x ,xi}{ii,v,ix,x,xi}{ii,vi,vi,viii,ix}{ii,vi,vii,viii,x}{ii,vi,vii,viii,xi}{ii,v i,vii,ix,x}{ii,vi,vii,ix,xi}{ii,vi,vii,x,xi}{ii,vi,viii,ix,x}{ii,vi,viii,ix,xi}{ii,vi,viii,ix,xi}
[0302] {ii,vi,vii,x,xi}{ii,vi,ix,x,xi}{ii,vii,viii,ix,x}{ii,vii,vii,vii,ii,xi,xi}{ii,vii,vii,x,xi}{ii,vii,ix,x,xi}{ii,vii,ix,x,xi}
[0303] {ii,viii,ix,x,xi}{iii,iv,v,v,vii}{iii,iv,v,v,vii}{iii,iv,v,v,vi,vii}{iii,iv,v,v,vi,ix}{iii,iv,v,vi,x}{iii,iv,v,v,vi,xi}
[0304] {iii,iv,v,vii,viii}{iii,iv,v,vii,ix}{iii,iv,v,vii,x}{iii,iv,v,vii,xi}{iii,iv,v,viii,ix}{iii,iv,v,viii,x}{iii,iv,v,vii,x}
[0305] {iii,iv,v,viii,xi}{iii,iv,v,v,ix,x}{iii,iv,v,ix,xi}{iii,iv,v,x,xi}{iii,iv,vi,vii,viii}{iii,iv,v,vii,ix}
[0306] {iii,iv,v,vi,x}{iii,iv,v,vi,xi}{iii,iv,v,viii,ix}{iii,iv,v,viii,x}{iii,iv,vi,viii,xi}{iii,iv,v,ix,x}
[0307] {iii,iv,vi,x,xi}{iii,iv,vi,x,xi}{iii,iv,vi,vii,xi}{iii,iv,vii,vii,x}{iii,iv,vii,viii,xi}{iii,iv,vii,ix,x}{iii,iv,vii,ix,x}
[0308] {iii,iv,vii,x,xi}{iii,iv,vii,x,xi}{iii,iv,viii,x,x}{iii,iv,viii,ix,xi}{iii,iv,viii,x,xi}{iii,iv,ix,x,xi}
[0309] {iii,v,v,vii,viii}{iii,v,vi,vii,ix}{iii,v,vi,vii,x}{iii,v,v,vii,xi}{iii,v,vi,viii,ix}{iii,v,vi,viii,x}{iii,v,vi,viii,x}
[0310] {iii,v,vi,viii,xi}{iii,v,vi,viii,x}{iii,v,vi,ixi,xi}{iii,v,vi,x,xi}{iii,v,vii,viii,ix}{iii,v,vii,viii,x}{iii,v,vii,viii,x}
[0311] {iii,v,vii,viii,xi}{iii,v,vii,xi,x}{iii,v,vi,ii,xi}{iii,v,vii,xi,xi}{iii,v,vii,x,xi}{iii,v,viii,ix,x}{iii,v,vii,ix,xi}
[0312] {iii,v,vii,x,xi}{iii,v,ix,x,xi}{iii,vi,vi,vii,ix}{iii,vi,vi,vii,vii,x}{iii,vi,vii,vii,vii,xi}{iii,vi,vii,ix,x}{iii,vi,vii,ix,x}
[0313] {iii,vi,vii,x,xi}{iii,vi,vii,x,xi}{iii,vi,viii, ix,x}{iii,vi,viii,xi,xi}{iii,vi,viii,x,xi}{iii,vi,vii,x,xi}{iii,vi,vii,x,xi}
[0314] {iii,vii,viii,ix,x}{iii,vii,viii,ix,xi}{iii,vii,viii,x,xi}{iii,vii,ii,x,xi}{iii,viii,ix,x,xi}{iv,v,vi,vii,viii}
[0315] {iv,v,vi,vii,ix}{iv,v,v,vii,x}{iv,v,vi,vii,xi}{iv,v,vi,viii,ix}{iv,v,v,viii,x}{iv,v,vi,viii,xi}{iv,v,vi,viii,xi}
[0316] {iv,v,vi,ix,x}{iv,v,vi,ix,xi}{iv,v,vi,x,xi}{iv,v,vii,viii,ix}{iv,v,v,vii,viii,x}{iv,v,vii,viii,xi}{iv,v,vii,viii,xi}
[0317] {iv,v,vii,ix,x}{iv,v,vii,ix,xi}{iv,v,vii,x,xi}{iv,v,viii,ix,x}{iv,v,viii,ix,xi}{iv,v,viii,x,xi}
[0318] {iv,v,ix,x,xi}{iv,vi,vi,viii,ix}{iv,vi,vii,viii,x}{iv,vi,vii,viii,x{iv,v,vii,ix,x}{iv,vi,vi,ix,xi}{
[0319] {iv,vi,vii,x,xi}{iv,vi,viii,ix,x}{iv,vi,viii,ix,xi}{iv,vi,viii,x,xi}{iv,vi,ix,x,xi}{iv,vii,viii,ix,x}{iv,vii,viii,ix,x}
[0320] {iv,vii,viii,ix,xi}{iv,vii,viii,x,xi}{iv,vii,ix,x,xi}{iv,viii,ix,x,xi}{v,vi,vii,viii,ix}{v,vi,vii,viii,x}{v,vi,vi,vii,vii,x}{
[0321] {v,vi,vii,viii,xi}{v,vi,vii,ix,x}{v,vi,vii,xi,xi}{v,vi,vii,x,xi}{v,vi,viii,ix,x}{v,vi,viii,ix,xi}{v,vi,viii,ix,xi}
[0322] {v,vi,viii,x,xi}{v,vi,ix,x,xi}{v,vii,viii,ix,x}{v,vii,viii,ix,xi}{v,vii,viii,x,xi}{v,vii,ix,x,xi}{v,vii,ix,x,xi}
[0323] {v,viii,x,x,xi}{vi,vii,viii,xi,x}{vi,vii,viii,ix,xi}{vi,vii,vii,x,xi}{vi,vii,ix,x,xi}{vi,viii,ix,x,xi}{vi,viii,ix,x,xi}
[0324] {vii,viii,ix,x,xi}{i,ii,iii,iv,v,vi}{i,ii,iii,iv,v,vii}{i,ii,iii,iv,v,viii}{i,ii,ii,iii,iv,v,ix}{i,ii,iii},x,v,v,
[0325] {i,ii,iii,iv,v,xi}{i,ii,iii,v,v,vi}{i,ii,iii,iv,v,viii}{i,ii,iii,iv,v,ix}{i,ii,ii,iii,iv,v,x}{i,ii,iii},iii,v,v
[0326] {i,ii,iii,iv,vi,viii}{i,ii,iii,iv,vii,ii}{i,ii,iii,iv,vii,x}{i,ii,iii,iv,vii,xi}{i,ii,iii,iii,vii,ix}{i,ii,ii,iii i,x}{i,ii,iii,iv,viii,xi}{i,ii,iii,iv,iii,xi}{i,ii,iii,iv,ix,xi}{i,ii,iii,iv,x,xi}{i,ii,iii,v,vi,vii}{i,ii},ii,iii
[0327] {i,ii,iii,v,v,ix}{i,ii,iii,v,v,x}{i,ii,iii,v,v,xi}{i,ii,iii,v,vi,viii}{i,ii,iii,iii,v,vii,ix}{i,ii,iiiiiii},v,v
[0328] {i,ii,iii,v,vi,xi}{i,ii,iii,v,viii,ix}{i,ii,iii,v,viii,x}{i,ii,iii,v,vii,xi}{i,ii,iii,v,ix,x}{i,ii,ii,iiiii
[0329] {i,ii,iii,v,x,xi}{i,ii,iii,v,vi,vii}{i,ii,iii,vi,vi,ix}{i,ii,iii,vi,vi,x}{i,ii,iiii,vi,vi,xi}{i,iii},ii,viiii
[0330] {i,ii,iii,vi,viii,x}{i,ii,iii,vi,viii,xi}{i,ii,iii,vi,ix,x}{i,ii,iii,vi, ix,xi}{i,ii,iii,vi,x,xi}{i,iii,iii,iii,vii
[0331] {i,ii,iii,vi,vi,vii,x}{i,ii,iii,vii,viii,xi}{i,ii,iii,vii,i,x}{i,ii,iii,vi, ix,xi}{i,ii,iii,vi,x,xi}{ii,iii,iii x,x}{i,ii,iii,viii,ix,xi}{i,ii,ii,iii,viii,x,xi}{i,ii,iii,iii,x,xi}{i,ii,v,v,vi,vii}{i,ii,iv,v,vi,vii}{i,ii,ii,iii},v
[0332] {i,ii,v,v,v,x}{i,ii,v,v,v,xi}{i,ii,v,v,viii}{i,ii,v,v,vii,ix}{i,ii,iv,v,v,vii,x}{i,ii,v,v,vi,xi}{i,ii,v,v,vi,xi}
[0333] {i,ii,v,v,viii,ix}{i,ii,v,v,viii,x}{i,ii,v,v,viii,xi}{i,ii,iv,v,x,x}{i,ii,iv,v,v,ix,xi}{i,ii,iv,v,x,xi}{i,ii,v,v,x,xi}
[0334] {i,ii,f,v,v,vii}{i,ii,v,v,v,ii}{i,ii,v,v,x}{i,ii,v,v,vi,x}{i,ii,v,v,viii,ix}{i,ii,v,v,viii, x}{i,ii,v,v,viii,x}{i,ii,v,v,x,x}{i,ii,v,v,x,xi}{i,ii,v,v,x,xi}{i,ii,v,vii,viii,x}{i,ii,iv,vi,viii,x}
[0335] {i,ii,v,vii,viii,xi}{i,ii,iv,vii,x}{i,ii,vii,xi,xi}{i,ii,ii,vii,x,xi}{i,ii,iv,viii,x,x}{i,ii,iv,viii,ix, xi}{i,ii,v,vii,x,xi}{i,ii,v,ix,x,xi}{i,ii,v,v,vi,vii}{i,ii,v,v,vii,ix}{i,ii,v,v,vi,vi,x}{i,ii,v,vi,vi,xi}{i,ii,v,vi,vi,vi,xi}
[0336] {i,ii,v,vi,viii,ix}{i,ii,v,vi,viii,x}{i,ii,v,vi,viii,xi}{i,ii,v,vi,x,x}{i,ii,v,vi,vi,xi,xi}{i,ii,v,vi,x,xi}{i,ii,v,vi,x,xi}
[0337] {i,ii,v,vii,viii,ix}{i,ii,v,vii,viii,x}{i,ii,v,vii,viii,xi}{i,ii,v,vii,ix,x}{i,ii,v,vi,i,xi,xi}{i,ii,v,vii,x,xi}{i,ii,v,vii,x,xi}
[0338] {i,ii,v,viii,ix,x}{i,ii,v,viii,ix,xi}{i,ii,v,viii,x,xi}{i,ii,v,x,xi}{i,ii,vi,vi,vii,viii,ix}{i,ii,vi,vii,viii,x}
[0339] {i,ii,v,vii,vii,xi}{i,ii,vi,vii,x,x}{i,ii,vi,vii,ix,xi}{i,ii,vi,vii,x,xi}{i,ii,vi, viii,ix,x}{i,ii,v,viii,x,xi}{i,ii,vi,viii,x,xi}{i,ii,vi,vi,x,x,xi}{i,ii,vii,viii,ix,x}{i,ii,vii,viii,x,xi}{i,ii,vii,vii,x,xi}{i,ii,vii,ix,x,xi}{i,ii,viii,ii,x,xi}{i,iii ,iv,v,v,vii}{i,iii,v,v,v,viii}{i,iii,v,v,v,ix}{i,iii,iv,v,v,x}{i,iii,v,v,v,xi}{i,iii,v,v,v,xi}
[0340] {i,iii,v,v,vi,vi,vi}{i,iii,v,v,vi,ix}{i,iii,v,v,vi,x}{i,iii,v,v,vi,xi}{i,iii,iv,v,viii,ix}{i,iii,iv,x,v,viii}{i,iii,v,v,viii,xi}{i,iii,iv,v,v,x}{i,iii,v,v,ix,xi}{i,iii,v,v,x,xi}{i,iii,iv,v,vi,vi,viii}{i,iii,iv,iv}vi,vi,vi
[0341] {i,iii,v,v,vi,x}{i,iii,v,v,vi,xi}{i,iii,v,v,viii,ix}{i,iii,iv,v,viii,x}{i,iii,iv,v,viii,xi}
[0342] {i,iii,v,v,ix,x}{i,iii,v,v,x,xi}{i,iii,v,v,x,xi}{i,iii,iv,vii,viii,x}{i,iii,iv,vii,viii,x}{i,iii,x}
[0343] {i,iii,iv,vii,vii,xi}{i,iii,iv,vii,x,x}{i,iii,iv,vii,ix,xi}{i,iii,iv,vii,x,xi}{i,iii,iv,viii,x,x}
[0344] {i,iii,iv,viii,x,xi}{i,iii,iv,viii,x,xi}{i,iii,iv,ix,x,xi}{i,iii,v,vi,vii,viii}{i,iii,v,vi,vi, ix}{i,iii,v,vi,vii x}{i,iii,v,v,vi,xi}{i,iii,v,vi,viii,ix}{i,iii,v,vi,viii,x}{i,iii,v,vi,viii,xi}{i,iii,v,vi, ix,x}{i,iii,xiv},vi,ix
[0345] {i,iii,v,v,x,xi}{i,iii,v,vi,viii,ix}{i,iii,v,vii,viii,x}{i,iii,v,vii,viii,xi}{i,i ii,v,vii,x,x}{i,iii,v,vii,x,xi}{i,iii,v,vii,x,xi}{i,iii,v,viii, ix,x}{i,iii,v,viii ,ix,xi}{i,iii,v,viii,x,xi}{i,iii,v,ix,x,xi}{i,iii,vi,vii,viii,ix}{i,iii,vi,vii,viii ,x}{i,iii,vi,vi,vii,xi}{i,iii,vi,vi,i,x}{i,iii,vi,vii,ix,xi}{i,iii,vi,vii,x,xi}
[0346] {i,iii,vi,viii,x,x}{i,iii,vi,viii,ix,xi}{i,iii,vi,viii,x,xi}{i,iii,vi, ix,x,xi}{i,iii,vii,viii, ix,x}
[0347] {i,iii,vii,viii,ix,xi}{i,iii,vii,viii,x,xi}{i,iii,vii,ix,x,xi}{i,iii,viii, ix,x,xi}{i,iv,v,vi,vii,viii}{i,iv,v,vi,vii,viii}
[0348] {i,v,v,v,vi,x}{i,v,v,v,vi,x}{i,v,v,v,vi,xi}{i,v,v,v,viii,ix}{i,v,v,v,viii,x}{i,v,v,v,vii,xi}
[0349] {i,iv,v,vi,ix,x}{i,iv,v,vi,ix,xi}{i,iv,v,v,vi,xi,xi}{i,iv,v,v,x,xi}{i,iv,v,vii,viii,ix}{i,iv,v,vii,viii,x}{i,iv,v,vii,viii,xi}{i,iv,v,vii,vii,xi}{
[0350] {i,iv,v,vii,ix,x}{i,iv,v,vii,ix,xi}{i,iv,v,vii,x,xi}{i,iv,v,viii,ix,x}{i,iv,v,viii,ix,xi}{i,iv,v,vii,x,xi}{i,iv,v,vii,x,xi}
[0351] {i,iv,v,ix,x,xi}{i,iv,vi,vi,viii,ix}{i,iv,vi,vii,viii,x}{i,iv,vi,vi,vii,viii,xi}{i,iv,vi,vi,vi,ix,x}
[0352] {i,iv,vi,vi,x,xi}{i,iv,vi,vi,x,xi}{i,iv,vi,viii,ix,x}{i,iv,vi,viii,ix,xi}{i,iv,vi,viii,x,xi}{i,iv,vi,i x,x,xi}{i,iv,vii,viii,ix,x}{i,iv,vii,viii,xi,xi}{i,iv,vii,viii,x,xi}{i,iv,vii,ix,x,xi}{i,iv,viii,ix,x,xi}{i,iv,viii,ix,x,xi}
[0353] {i,v,vi,vi,vii,ix}{i,v,vi,vi,vii,x}{i,v,vi,vi,vii,xi}{i,v,vi,vi,ix,x}{i,v,vi,vi,vi,xi,xi}{i,v,vi,vi,x,xi}{i,v,vi,viii,ix,x}{i,v,vi,viii,ix,xi}{i,v,vi,viii,x,xi}{i,v,vi,x,xi}{i,v,vii,viii,ix,x}{i,v,vii,viii,ix,xr
[0354] {i,v,vii,vii,x,xi}{i,v,vii,ix,x,xi}{i,v,viii,ix,x,xi}{i,vi,vii,viii,ix,x}{i,vi,vii,viii,ix,xi}{i,vi,vii,viii,ix,xi}
[0355] {i,vi,vii,viii,x,xi}{i,vi,vii,xi,x,xi}{i,vi,viii,ix,x,xi}{i,vii,viii,ix,x,xi}{ii,iii,iv,v,vi,vii}
[0356] {ii,iii,v,v,v,viii}{ii,iii,v,v,v,ix}{ii,iii,v,v,v,x}{ii,iii,v,v,vi,xi}{ii,iii,iv,v,vi,vii}{ii,iii,iv,v,vi,viii}
[0357] {ii,iii,v,v,vi,ix}{ii,iii,v,v,vi,x}{ii,iii,v,v,vi,xi}{ii,iii,iv,v,viii,ix}{ii,iii,iv,v,viii,x}{ii,iii,iv,v,viii,x}
[0358] {ii,iii,iv,v,viii,xi}{ii,iii,iv,v,ix,x}{ii,iii,iv,v,ix,xi}{ii,iii,iv,v,x,xi}{ii,iii,iv,vi,vi,viii}
[0359] {ii,iii,v,vi,vi,ix}{ii,iii,iv,vi,vi,x}{ii,iii,iv,vi,vi,xi}{ii,iii,iv,vi,viii,ix}{ii,iii,iv,vi,vii,x}{ii,iii,iv,vi,viii,x}
[0360] {ii,iii,iv,vi,viii,xi}{ii,iii,iv,vi,x,x}{ii,iii,iv,vi,ix,xi}{ii,iii,iv,vi,x,xi}{ii,iii,iv,vi,viii,ix}
[0361] {ii,iii,v,vi,vii,x}{ii,iii,iv,vii,viii,xi}{ii,iii,iv,vii,x}{ii,iii,iv,vii,ix,xi}{ii,iii,iv,vii,x,xi}{ii,iii,iv,vii,x,xi}
[0362] {ii,iii,iv,viii,ix,x}{ii,iii,iv,viii,ix,xi}{ii,iii,iv,viii,x,xi}{ii,iii,iv,ix,x,xi}{ii,iii,v,vi,vi,viii}{ii,iii,v,vi,vi,viii}
[0363] {ii,iii,v,vi,vi,ix}{ii,iii,v,vi,vi,x}{ii,iii,v,vi,vi,xi}{ii,iii,v,vi,viii,ix}{ii,iii,v,vi,vii,x}{ii,iii,v,vi,viii,x}
[0364] {ii,iii,v,vi,vii,xi}{ii,iii,v,vi,x}{ii,iii,v,vi, ix,xi}{ii,iii,v,vi,x,xi}{ii,iii,v,vii,viii,ix}
[0365] {ii,iii,v,vi,vii,x}{ii,iii,v,vii,viii,xi}{ii,iii,v,vii,x}{ii,iii,v,vii,ix,xi}{ii,iii,v,vii,x,xi}{ii,iii,v,vii,x,xi}
[0366] {ii,iii,v,viii,ix,x}{ii,iii,v,viii,ix,xi}{ii,iii,v,viii,x,xi}{ii,iii,v,ix,x,xi}{ii,iii,vi,vii,viii,ix}
[0367] {ii,iii,vi,vi,vi,x}{ii,iii,vi,vi,vii,xi}{ii,iii,vi,vi,i,x}{ii,iii,vi,vi, ix,xi}{ii,iii,vi,vi,x,xi}{ii,iii,vi,vi,x,xi}
[0368] {ii,iii,viii,viii,x,x}{ii,iii,vi,viii,ix,xi}{ii,iii,vi,viii,x,xi}{ii,iii,vi, ix,x,xi}{ii,iii,vi,viii, ix,x}
[0369] {ii,iii,vi,vii,ix,xi}{ii,iii,vii,vii,x,xi}{ii,iii,vii,ix,x,xi}{ii,iii,vii, ix,x,xi}{ii,iv,v,vi,vii,viii}{ii,iv,v,vi,vii,viii}
[0370] {ii,v,v,v,vi,ix}{ii,v,v,v,vi,x}{ii,v,v,v,vi,xi}{ii,iv,v,v,v,viii,x}{ii,v,v,v,viii,x}
[0371] {ii,iv,v,v,viii,xi}{ii,iv,v,vi,x}{ii,iv,v,v,vi,x}{ii,iv,v,v,ix,xi}{ii,iv,v,v,x,xi}{ii,iv,v,vii,viii,x}{ii,iv,v,vi i,viii,x}{ii,iv,v,vii,viii,xi}{ii,iv,v,v,vii,ix,x}{ii,iv,v,vii,ix,xi}{ii,iv,v,vii,x,xi}{ii,iv,v,viii,x,x}{ii,iv,v,viii,ix,x}
[0372] {ii,iv,v,viii,ix,xi}{ii,iv,v,viii,x,xi}{ii,iv,v,ix,x,xi}{ii,iv,v,vii,viii,ix}{ii,iv,vi,vi,vii,viii,x}{ii,iv,vi,vi,vii,x}
[0373] {ii,iv,vi,vii,viii,xi}{ii,iv,vi,vii,x,x}{ii,iv,vi,vii,xi,xi}{ii,iv,vi,vii,x,xi}{ii,iv,vi,viii,ix,x}
[0374] {ii,iv,vi,viii,ix,xi}{ii,iv,vi,vii,x,xi}{ii,iv,vi,ix,x,xi}{ii,iv,vii,viii,ix,x}{ii,iv,vii,viii,ix,xi}{ii,iv,vii,viii,ix,xi}
[0375] {ii,v,vii,vii,x,xi}{ii,iv,vii,x,x,xi}{ii,iv,viii,ix,x,xi}{ii,v,vi,vii,viii,ix}{ii,v,vi,vi,vii,x}{ii,v,vi,vii,vii,x}
[0376] {ii,v,vi,vi,vii,xi}{ii,v,vi,vii,x,x}{ii,v,vi,vi,vi,xi,xi}{ii,v,vi,vi,x,xi}{ii,v,vi,viii,ix,x}
[0377] {ii,v,vi,viii,ix,xi}{ii,v,vi,viii,x,xi}{ii,v,vi,ix,x,xi}{ii,v,vii,viii,ix,x}{ii,v,vii,viii,ix,xi}{ii,v,vii,viii,ix,xi}
[0378] {ii,v,vii,vii,x,xi}{ii,v,vii,x,x,xi}{ii,v,viii,ix,x,xi}{ii,vi,vii,viii, ix,x}{ii,vi,vii,viii,ii,xi}{ii,vi,vii,viii,ix,xi}
[0379] {ii,vi,vii,viii,x,xi}{ii,vi,vii,ii,x,xi}{ii,vi,viii, ix,x,xi}{ii,vii,viii, ix,x,xi}{iii,iv,v,vi,vii,vii}{iii,iv,v,vi,vii,viii}
[0380] {iii,iv,v,vi,vi,ix}{iii,iv,v,vi,vi,x}{iii,iv,v,vi,vi,xi}{iii,iv,v,vi,viii,ix}{iii,iv,v,v,viii,x}{iii,iv,v,v,viii,x}
[0381] {iii,iv,v,vi,viii,xi}{iii,iv,v,vi,x}{iii,iv,v,vi,ix,xi}{iii,iv,v,v,x,xi}{iii,iv,v,v,vii,viii,ix}{iii,iv,v,vii,viii,ix}
[0382] {iii,iv,v,vii,viii,x}{iii,iv,v,vii,viii,xi}{iii,iv,v,vii,ix,x}{iii,iv,v,vii,ix,xi}{iii,iv,v,vii,x,xi}{iii,iv,v,vii,x,xi}{iii,iv,v,vii,x,xi}
[0383] {iii,iv,v,viii,x,x}{iii,v,v,viii,ix,xi}{iii,iv,v,viii,x,xi}{iii,iv,v,x,x,xi}{iii,iv,v,vii,viii,ix}
[0384] {iii,iv,vi,vi,vii,x}{iii,iv,vi,vi,viii,xi}{iii,iv,vi,vii,ix,x}{iii,iv,vi,vii,ix,xi}{iii,iv,vi,vi,x,xi}{iii,iv,vi,vi,x,xi}
[0385] {iii,iv,v,viii,x,x}{iii,v,v,viii,ix,xi}{iii,iv,vi,viii,x,xi}{iii,iv,v,ix,x,xi}{iii,iv,vii,viii,x,x}{iii,iv,vii,viii,ix,x}
[0386] {iii,iv,vii,viii,ix,xi}{iii,iv,vii,vii,x,xi}{iii,iv,vii,ix,x,xi}{iii,iv,viii,ix,x,xi}{iii,v,vi,vii,vii,xi}{iii,v,vi,vii,viii,ix}
[0387] {iii,v,vi,vi,viii,x}{iii,v,vi,vi,vii,xi}{iii,v,vi,vii,ix,x}{iii,v,vi,vi, ix,xi}{iii,v,vi,vi,x,xi}{iii,v,vi,vi,x,xi}
[0388] {iii,v,vi,viii,x,x}{iii,v,vi,viii,ix,xi}{iii,v,vi,viii,x,xi}{iii,v,vi,x,x,xi}{iii,v,vii,viii,x,x}{iii,v,vii,viii,ix,x}
[0389] {iii,v,vii,vii,ix,xi}{iii,v,vii,vii,x,xi}{iii,v,vii,ix,x,xi}{iii,v,viii,ix,x,xi}{iii,vi,vii,vii,viii,x}{iii,vi,vii,viii,ix,x}
[0390] {iii,vi,vii,viii,xi,xi}{iii,vi,vii,vii,x,xi}{iii,vi,vii, ix,x,xi}{iii,vi,viii, ix,x,xi}{iii,vii,viii, ix,x,xi}{iii,vii,viii, ix,x,xi}{iii,vii,viii, ix,x,xi}
[0391] {iv,v,vi,vii,viii,ix}{iv,v,vi,vii,viii,x}{iv,v,vi,vii,viii,xi}{iv,v,vi,vii,ix,x}{iv,v,vi,vii,ix,xi}{iv,v,vi,vii,ix,xi}
[0392] {iv,v,vi,vii,x,xi}{iv,v,vi,viii,ix,x^iv,v,vi,viii,ix,xi}{iv,v,vi,viii,x,xi}{iv,v,vi,ix,x,xi}{iv,v,vi,ix,x,xi}
[0393] {iv,v,vii,viii,ix,x}{iv,v,vii,viii,ix,xi}{iv,v,vii,viii,x,xi iv,v,vii,ix,x,xi}{iv,v,viii,ix,x,xi}{iv,v,viii,ix,x,xi}
[0394] {iv,vi,vii,viii,ix,x}{iv,vi,vii,viii,ix,xi}{iv,vi,vii,viii,x,xi}{iv,vi,vii,ix,x,xi iv,vi,viii,ix,x,xi}
[0395] {iv,vii,viii,ix,x,xi}{v,vi,vii,viii,ix,x}{v,vi,vii,viii,ix,xi}{v,vi,vii,viii,x,xi}{v,vi,vii,ix,x,xi v,
[0396] vi,viii,ix,x,xi}{v,vi,viii,ix,x,xi}{vi,vi,viii,ix,x,xi}{i,ii,iii,iv,v,vi,vii}{i,ii,iii,iv,v,vi,x,vii}{
[0397] {i,ii,iii,iv,v,v,ix i,v,v,vi,x}{i,ii,iii,iv,v,vi,xi}{i,ii,iii,iv,v,vii,viii}{i,ii,iii,iv,v,vii,ix}
[0398] {i,ii,iii,iv,v,vi,x}{i,ii,iii,iv,v,vi,xi i,iv,v,viii,ix}{i,ii,iii,iv,v,viii,x}{i,ii,iii,iv,v,viii,xi}{
[0399] {i,ii,iii,iv,v,ix,x}{i,ii,iii,iv,v,ix,xi}{i,ii,iii,iv,v,x,xi i,iv,vi,vii,vii}{i,ii,iii,iv,vi,vii,ix}
[0400] {i,ii,iii,iv,vi,vi,x}{i,ii,iii,iv,vi,vi,xi}{i,ii,iii,iv,vi,viii,ix}{i,ii,iii,iv,vi,viii,xi,iv,vi,viii,xi}
[0401] {i,ii,iii,iv,v,ix,x}{i,ii,iii,v,v,ix,xi}{i,ii,iii,v,v,x,xi}{i,ii,iii,iv,vii,viii,ix}{i,ii,iii,iv,vii,vii,x^^
[0402] i,iv,vii,viii,xi}{i,ii,iii,iv,vii,ix,x}{i,ii,iii,iv,vii,ix,xi}{i,ii,iii,iv,vii,x,xi}{i,ii,iii,iv,viii,ix,x}{
[0403] {i,ii,iii,iv,viii,ix,xi i,iv,viii,x,xi}{i,ii,iii,iv,ix,x,xi}{i,ii,iii,v,vi,vii,vii}{i,ii,iii,v,vi,vii,ix}
[0404] {i,ii,iii,v,v,vi,x}{i,ii,iii,v,v,vi,xi i,v,vi,viii,ix}{i,ii,iii,v,vi,viii,x}{i,ii,iii,v,vi,viii,xi}{
[0405] {i,ii,iii,v,v,ix,x}{i,ii,iii,v,v,ix,xi}{i,ii,iii,v,vi,x,xi}{i,ii,iii,v,vii,viii,ix}{i,ii,iii,v,vii,viii,x}
[0406] {i,ii,iii,v,vii,viii,xi}{i,ii,iii,v,vii,ix,x}{i,ii,iii,v,vii,xi,xi}{i,ii,iii,v,vii,x,xi i,v,viii,ix,x}
[0407] {i,ii,iii,v,viii,ix,xi}{i,ii,iii,v,viii,x,xi}{i,ii,iii,v,ixi,x,xi}{i,ii,iii,vi,vi,vii,iii}{i,ii,iii,vi,vii,vii i,xi,ii,vii,viii,xi}{i,ii,iii,vi,vii,ix,x}{i,ii,iii,vi,vii,ix,xi}{i,ii,iii,vi,vii,x,xi}{i,ii,iii,vi,viii,ix,x}{i,ii,iii,vi,vii,ix,x}
[0408] {i,ii,iii,vi,viii,ix,xi i,vi,viii,x,xi}{i,ii,iii,vi, ix,x,xi}{i,ii,iii,vii,viii, ix,x}{i,ii,iii,vii,viii,ix,xi}{
[0409] {i,ii,iii,vii,vii,x,xi}{i,ii,iii,vii,ii,x,xi i,viii, iii,x,xi}{i,ii,iv,v,vi,vii,vii}{i,ii,iv,v,vi,vii,ix}
[0410] {i,ii,iv,v,v,vi,x}{i,ii,v,v,v,vi,xi}{i,ii,v,v,v,viii,ix iv,v,vi,viii,x}{i,ii,iv,v,v,viii,xi}{
[0411] {i,ii,v,v,v,ix,x}{i,ii,v,v,v,x}{i,ii,v,v,v,x,xi}{i,ii,v,v,vii,viii,ix}{i,ii,iv,v,vii,viii,x}
[0412] {i,ii,v,v,vii,vii,xi}{i,ii,iv,v,vii,x,x}{i,ii,iv,v,vii,xi,xi}{i,ii,iv,v,vii,x,xi}{i,ii,iv,v,vii,ix, xiv,v,viii,x,xi}{i,ii,iv,v,viii,x,xi}{i,ii,iv,v,ix,x,xi}{i,ii,iv,v,vii,viii,x}{i,ii,iv,vi,vii,viii,x}{i,ii,iv,vi,vii,vii,x}
[0413] {i,ii,iv,vi,vi,vii,xi iv,vi,vii,x,x}{i,ii,iv,vi,vii,xi,xi}{i,ii,iv,vi,vii,x,xi}{i,ii,iv,vi,vi,vii,x,xi}{i,ii,iv,vi,viii,ix,x}
[0414] {i,ii,iv,vi,viii,ix,xi}{i,ii,iv,vi,viii,x,xi iv,vi,ix,x,xi}{i,ii,iv,vii,viii,ix,x}{i,ii,iv,vii,viii,ix,xi}{i,ii,iv,vii,viii,ix,xi}
[0415] {i,ii,v,vii,viii,x,xi}{i,ii,iv,vii,x,x,xi}{i,ii,iv,viii,ix,x,xi}{i,ii,v,vii,viii,ix}{i,ii,v, vi,vi,vii,x}{i,ii,v,vi,vi,vii,xi}{i,i,v,vi,vi,vi,x}{i,i,v,vi,vi,vii,xi}{i,i,v,vi,vi,vi,x,xi v,vi,viii,x,x}
[0416] {i,ii,v,vi,viii,xi,xi}{i,ii,v,vi,viii,x,xi}{i,ii,v,vi,ix,x,xi}{i,ii,v,vii,viii,ix,x}{i,ii,v,vii,viii,ix,xi v,
[0417] vii,viii,x,xi}{i,ii,v,vii,ix,x,xi}{i,ii,v,viii,ix,x,xi}{i,ii,vi,vii,viii,x,x}{i,ii,vi,vi,vii,viii,ix,xi}{i,ii,vi,vi,viii,ix,xi}
[0418] {i,ii,vi,vii,viii,x,xi vi,vii,ii,x,xi}{i,ii,vi,viii, ix,x,xi}{i,ii,vii,viii, ix,x,xi}{i,iii,iv,v,vii,vii,viii}{
[0419] {i,iii,iv,v,v,vi,ix}{i,iii,v,v,v,vi,x}{i,iii,iv,v,v,vi,xi}{i,iii,iv,v,vi,viii,ix}{i,iii,iv,v,vi,vi ii,x}{i,iii,iv,v,v,v,vii,xi}{i,iii,iv,v,v,vi,x,x}{i,iii,iv,v,v,x,xi,iv,v,v,x}{i,iii,iv,v,vii,viii,ix}
[0420] {i,iii,iv,v,vii,viii,x}{i,iii,iv,v,vii,viii,xi}{i,iii,iv,v,v,vii,ix,x}{i,iii,iv,v,vii,xi,xi,iv,v,vii,x,xi}
[0421] {i,iii,iv,v,viii,ix,x}{i,iii,iv,v,viii,ix,xi}{i,iii,iv,v,viii,x,xi}{i,iii,iv,v,ix,x,xi}
[0422] {i,iii,v,vi,vi,vii,vii, iv,v,vi,vi,viii,x}{i,iii,iv,vi,vi,viii,xi}{i,iii,iv,vi,vii,ix,x}{i,iii,iv,vi,vi,ix,xi}
[0423] {i,iii,iv,vi,vi,x,xi}{i,iii,iv,vi,viii,ix,x}{i,iii,iv,vi,viii,ix,xi}{i,iii,iv,vi,viii,x,xi}{i,iii,iv,vi,ix,x,xi}{i,iii,iv,vi,ix,x,xi}
[0424] {i,iii,iv,vii,viii,ix,x}{i,iii,iv,vii,viii,ix,xi}{i,iii,iv,vii,viii,x,xi}{i,iii,iv,vii,ix,x,xi}
[0425] {i,iii,v,viii,ix,x,xi}{i,iii,v,vi,vi,viii,ix}{i,iii,v,vi,vii,viii,x}{i,iii,v,vi,vii,viii,xi}{i,iii,v,vi}vii,x,ix
[0426] {i,iii,v,vi,vii,x,xi}{i,iii,v,vi,vi,x,xi}{i,iii,v,vi,viii,x,x}{i,iii,v,vi,viii,ix,xi}{i,iii,v,vi,vii,x,xi}{i,iii,v,vi,vii,x,xi}
[0427] {i,iii,v,vi,x,x,xi,v,vii,viii,x,x}{i,iii,v,vii,viii,ix,xi}{i,iii,v,vii,viii,x,xi}{i,iii,v,vii,vii,x,xi}{i,iii,v,vii,iii,x,xi}{i,iii,v,vii,ii,x,xi}
[0428] {i,iii,v,viii,ix,x,xi}{i,iii,vi,vi,viii,ix,x}{i,iii,vi,vii,viii,ix,xi}{i,iii,vi,vii ,viii,x,xi}{i,iii,vi,vii,ix,x,xi}{i,iii,vi,viii,ix,x,xi}{i,iii,vii,viii,ix,x,xi}{i,i v,v,v,vii,viii,ix}{i,v,v,vi,vii,viii,x}{i,iv,v,vi,vii,viii,xi}{i,iv,v,vi,vii,ix,x}{i,iv,v,vi,vii,ix,x} {i,iv,v,vi,vii,xi,xi}{i,iv,v,vi,vii,x,xi}{i,iv,v,vi,viii,ix,x}{i,iv,v,vi,viii,ix,xi}{i,iv,v,vi,viii,ix,xi}
[0429] {i,iv,v,v,viii,x,xi}{i,v,v,vi,x,x,xi}{i,iv,v,vii,viii,ix,x}{i,iv,v,vii,viii,ix,xi}{i,iv,v,vii,viii,x,xi}{i,iv,v,vii,viii,x,xi}{i,
[0430] iv,v,vii,x,x,xi}{i,iv,v,viii,x,x,xi}{i,iv,v,vii,viii,ix,x}{i,iv,vi,vii,viii,ix,xi}{i,iv,vi,vi,vii,viii,x,xi}{i,iv,vi,vi,vii,x,xi}{
[0431] {i,iv,vi,vii,x,x,xi iv,vi,viii,ix,x,xi}{i,iv,vii,viii,ix,x,xi}{i,v,vi,vii,viii,ix,x}{i,v,vi,vi,vii,viii,ix,xi}{i,v,vi,vi,vii,vii,ix,xi}
[0432] {i,v,vi,vi,vii,x,xi}{i,v,vi,vii,ix,x,xi}{i,v,vi,viii,ix,x,xi}{i,v,vii,viii,ix,x,xi}{i,vi,vi,vii,vii,x,xi}{i,vi,vi,vii,ix,x,xi}{i,vi,vi,vii,viii,x,xi}
[0433] {ii,iii,v,v,v,v,vi,vii}{ii,iii,iv,v,v,vi,vi,ix}{ii,iii,iv,v,v,vii,x}{ii,iii,iv,v,vi,vi,xi}{ii,iii,v,v,v,v,vi,viii ,ix}{ii,iii,iv,v,v,vi,vii,x}{ii,iii,v,v,v,viii,xi}{ii,iii,iv,v,v,vi,x,x}{ii,iii,iv,v,v,v,ix,x}{ii,iii,iv,v},v,x,x
[0434] {ii,iii,iv,v,vi,v,vii,ix}{ii,iii,iv,v,v,vii,x}{ii,iii,iv,v,vii,viii,xi}{ii,iii,iv,v,v,vii,ix,x}{ii,iii,iv,iv,v,vii ,x{ii,iii,v,v,vi,x,xi}{ii,iii,iv,v,v,viii,ix,x}{ii,iii,iv,v,viii,ix,xi}{ii,iii,iv,v,v,viii,x,xi}{ii,iii,iv,x},xiv,ix,ix
[0435] {ii,iii,iv,vi,vi,viii,ix}{ii,iii,iv,vi,vii,viii,x}{ii,iii,iv,vi,vii,viii,xi}{ii,iii,iv,vi,vii,ix,x}
[0436] {ii,iii,iv,v,vi,vi,x,xi}{ii,iii,iv,vi,vi,x,xi}{ii,iii,iv,vi,viii,ix,x}{ii,iii,iv,vi,viii,ix,xi}
[0437] {ii,iii,iv,vi,vii,x,xi}{ii,iii,iv,vi,x,x,xi}{ii,iii,iv,vii,vii,iii,ix,x}{ii,iii,iv,vii,viii,ix,xi}
[0438] {ii,iii,iv,vi,vi,x,x,x{ii,iii,iv,vii,x,x,xi}{ii,iii,iv,viii,ix,x,xi}{ii,iii,v,vi,vi,vii,ii, ix}{ii,iii,v,vi,vi,vi,vi,i x}{ii,iii,v,v,vi,vi,vii,xi}{ii,iii,v,vi,vi,vi,x,x}{ii,iii,v,vi,vi,i, ix,xi}{ii,iii,v,vi,vi,vi,x,xi}{ii,iii,v,vi,vi,x}viii,i
[0439] {ii,iii,v,vi,viii,x,xi}{ii,iii,v,vi,viii,x,xi}{ii,iii,v,v,vi,x,x,x}{ii,iii,v,vii,viii,ix,x}{ii,iii,v,vii,xii,xi}viii
[0440] {ii,iii,v,vi,vii,x,xi}{ii,iii,v,vii,ix,x,xi}{ii,iii,v,viii,ix,x,xi}{ii,iii,vi,vi,vii,viii,ix,x}
[0441] {ii,iii,vi,vi,viii,ix,xi}{ii,iii,vi,vii,viii,x,xi}{ii,iii,vi,vii, ix,x,xi}{ii,iii,vi,viii,ix,x,xi}
[0442] {ii,iii,vi,vii,ix,x,xi}{ii,iv,v,v,vi,vii,viii,ix}{ii,iv,v,vi,vi,viii,x}{ii,iv,v,vi,vi,vii,xi}{
[0443] {ii,iv,v,v,vi,vi,x,x}{ii,v,v,vi,vi,xi,xi}{ii,iv,v,vi,vi,x,x}{ii,iv,v,vi,viii,ix,x}{ii,iv,v,vi,viii,ix,xi}{ii,iv,v,vi,viii,ix,xi}
[0444] {ii,iv,v,v,viii,x,xi}{ii,v,v,vi,x,x,xi}{ii,iv,v,vii,viii,ix,x}{ii,iv,v,vii,viii,ix,xi}{ii,iv,v,vii,viii,x,x}{ii,iv,v,vii,viii,x,x}
[0445] {ii,iv,v,vii,ix,x,xi}{ii,iv,v,viii,ix,x,xi}{ii,iv,vi,vii,viii,ix,x}{ii,iv,vi,vii,vii,ix,xi}{ii,iv,vi,vii,viii,ix,xi}
[0446] {ii,v,vi,vi,vii,x,xi}{ii,vi,vi,vi,vi,x,x,x}{ii,iv,vi,vii,ix,x,xi}{ii,iv,vi,vii,ix,x,xi}{ii,v,vi,vi,vii,ix, x}{ii,v,vi,vi,vii,x,xi}{ii,v,vi,vi,vii,vii,x,xi}{ii,v,vi,vi,vii,ix,x,x{ii,v,viii,ix,x,xi}{ii,v,vi,vii,ix,x,xi}{ii,v,vi,vii,ix,x,xi}
[0447] {ii,vi,vi,viii,ix,x,xi}{iii,v,v,vi,vi,viii,ix}{iii,v,v,vi,vi,viii,x}{iii,iv,v,v,vi,vii,viii,xi}
[0448] {iii,iv,v,v,vi,vi,x,x}{iii,v,v,v,vi,vi,ix,xi}{iii,iv,v,vi,vii,x,xi}{iii,iv,v,vi,viii,ix,x}{iii,iv,v,vi, viii,ix,x}{iii,iv,v,v,viii,x,xi}{iii,iv,v,v,vi,ix,x,xi}{iii,iv,v,vii,viii,x,x}{iii,iv,v,vii,viii,ix,x}{iii,iv,v,vii,viii,ix,xi}
[0449] {iii,iv,v,vii,viii,x,xi}{iii,iv,v,vii,x,x,x iii,iv,v,viii,ix,x,xi}{iii,iv,vi,vii,viii,ix,x}{iii,iv,vii,viii,ix,x}
[0450] {iii,iv,vi,vii,viii,x,xi}{iii,iv,vi,vii,viii,x,xi}{iii,iv,vi,vi,vii,x,x,xr iii,iv,vi,vii,ix,x,xi}
[0451] {iii,iv,vii,viii,ix,x,xi}{iii,v,vi,vii,viii,ix,x}{iii,v,vi,vii,viii,ix,xi}{iii,v,vi,vi,vii,viii,x,xi}
[0452] {iii,v,vi,vii,ix,x,xi}{iii,v,vi,viii,ix,x,xi}{iii,v,vii,viii,ix,x,xi}{iii,vi,vii,viii,ix,x,xi}
[0453] {iv,v,vi,vii,viii,ix,x}{iv,v,vi,vii,viii,ix,xi}{iv,v,vi,vii,viii,x,xi}{iv,v,vi,vii,ix,x,xi}{
[0454] {iv,v,vi,vii,ix,x,xi}{iv,v,vii,viii,ix,x,xi}{iv,vi,vii,viii,ix,x,xi}{v,vi,vii,viii,ix,x,xi}{v,vi,vii,viii,ix,x,xi}
[0455] {i,ii,iii,iv,v,v,v,viii}{i,ii,iii,iv,v,vi,vii,ix}{i,ii,iii,iv,v,vi,vii,x}{i,ii,iii,iv,v,vi,vii,xi}{i,ii,xi}
[0456] {i,ii,iii,iv,v,v,viii,ix}{i,ii,iii,iv,v,vi,viii,x}{i,ii,iii,iv,v,vi,viii,xi}{i,ii,iii,iv,v,vi,ix,x}
[0457] {i,ii,iii,iv,v,v,x,xi}{i,ii,iii,iv,v,v,x,xi}{i,ii,iii,iv,v,vii,viii,ix}{i,ii,iii,iv,v,vii,vii,x}
[0458] {i,ii,iii,iv,v,v,v,viii,x}{i,ii,iii,iv,v,vii,x,x}{i,ii,iii,iv,v,vii,ix,xi}{i,ii,iii,iv,v,vii,x,xi}
[0459] {i,ii,iii,iv,v,viii,ix,x}{i,ii,iii,iv,v,viii,ix,xi,ii,iii,iv,v,viii,x,xi}{i,ii,iii,iv,v,ix,x,xi}
[0460] {i,ii,iii,iv,v,v,v,vii,ix}{i,ii,iii,v,v,v,viii,x}{i,ii,iii,v,v,vii,vii,x}{i,ii,iii,v,vi,vi,ix,x}
[0461] {i,ii,iii,iv,v,v,i,x,xi}{i,ii,iii,iv,v,vi,x,xi}{i,ii,iii,iv,v,viii,x,x}{i,ii,iii,iv,vi,viii,ix,x}
[0462] {i,ii,iii,iv,vi,viii,x,xi}{i,ii,iii,iv,vi,i,x,xi}{i,ii,iii,iv,vii,viii,ixi,x}{i,ii,iii,iv,vii,viii,ix,xi}
[0463] {i,ii,iii,iv,vi,vii,x,x}{i,ii,iii,iv,vii,x,x,xi}{i,ii,ii,iii,iv,viii,ix,x,xi}{i,ii,iii,v,vi,vi,vii,ix}
[0464] {i,ii,iii,v,v,v,v,vii,x}{i,ii,iii,v,v,vi,viii,x}{i,ii,iii,v,vi,vii,x}{i,ii,iii,v,vi,vii,ix,xi}
[0465] {i,ii,iii,v,vi,vi,x,xi}{i,ii,iii,v,vi,viii,ix,x}{i,ii,iii,v,vi,viii,ix,x}{i,ii,iii,v,vi,viii,x,xi}
[0466] {i,ii,iii,v,vi,x,x,xi}{i,ii,iii,v,vii,viii,iii,x,x}{i,ii,iii,v,vii,viii,ix,xi}{i,ii,iii,v,vii,viii,x,x}
[0467] {i,ii,iii,v,vii,ix,x,xi}{i,ii,iii,v,viii,ix,x,xi}{i,ii,ii,iii,vi,vii,viii,ixi,x}{i,ii,iii,vi,vii,viii,ix,xi}
[0468] {i,ii,iii,vi,vi,vii,x,x}{i,ii,iii,vi,vii,x,x,xi}{i,ii,ii,iii,vi,viii, ix,x,xi}{i,ii,iii,ii,viii, ix,x,xi}
[0469] {i,ii,f,v,v,v,vi,vii,ix}{i,ii,v,v,v,v,vii,x}{i,ii,v,v,v,vi,viii,xi}{i,ii,iv,v,v,vi,vi, ix,x}
[0470] {i,ii,v,v,v,vi,vi,x,xi}{i,ii,v,v,v,vi,x,xi}{i,ii,iv,v,v,viii,x,x}{i,ii,v,v,vi,viii,ix,xi}
[0471] {i,ii,iv,v,v,viii,x,xi}{i,ii,iv,v,v,vi,x,xi}{i,ii,iv,v,vii,viii,ix,x}{i,ii,iv,v,vii,viii,ix,x}
[0472] {i,ii,iv,v,vi,vii,x,xi}{i,ii,iv,v,vii,ix,x,xi}{i,ii,iv,v,viii,ix,x,xi}{i,ii,iv,vi,vi,vii,viii,ix,x}
[0473] {i,ii,iv,vi,vi,viii,ix,x}{i,ii,iv,vi,vii,viii,x,xi}{i,ii,iv,vi,vii,ix,x,xi}{i,ii,iv,vi,viii,ix,x,xi}{i,ii,vi,vi,viii,ix,x,xi}
[0474] {i,ii,iv,vi,viii,ix,x,xi}{i,ii,v,vi,vi,vii,viii,ix,x}{i,ii,v,vi,vi,viii,viii,ix,xi}{i,ii,v,vi,vi,vii,viii,x,xi}
[0475] {i,ii,v,vi,vi,ix,x,xi}{i,ii,v,vi,viii,ix,x,xi}{i,ii,v,vii,viii,ix,x,xi}{i,ii,vi,vi,vii,viii,ix,x,xi}
[0476] {i,iii,v,v,v,v,v,viii,ix}{i,iii,v,v,v,v,vii,x}{i,iii,v,v,v,vii,vii,xi}{i,iii,v,v,v,vii,x,x}
[0477] {i,iii,iv,v,v,vi,vi,x,xi}{i,iii,iv,v,v,vi,vii,x,xi}{i,iii,iv,v,v,viii,viii,ix,x}{i,iii,iv,v,v,vi,viii,ix,xi}{i,iii,xi}
[0478] {i,iii,iv,v,vi,viii,x,xi}{i,iii,iv,v,v,vi,x,x,xi}{i,iii,iv,v,vii,viii,ix,x}{i,iii,iv,v,v,vii,viii,ix,xi}{i,iii,xi}
[0479] {i,iii,iv,v,vi,vi,vii,x,xi}{i,iii,iv,v,vii,ix,x,xi}{i,iii,iv,v,viii,ix,x,xi}{i,iii,iv,v,vi,vii,viii,ix,x}
[0480] {i,iii,iv,vi,vi,vii,vii,iii,x,xi}{i,iii,iv,vi,vii,viii,x,xi}{i,iii,iv,vi,vii,ix,x,xi}{i,iii,iv,vi,viii,ix,x,xi}{i,iii,iv,vi,viii,ix,x,xi}
[0481] {i,iii,iv,vi,viii,ix,x,xi}{i,iii,v,vi,vi,viii,ix,x}{i,iii,v,vi,vii,viii, ix,xi}{i,iii,v,vi,vii,viii,x,x i}{i,iii,v,vi,vi,vi,x,x,xi}{i,iii,v,vi,viii,ix,x,xi}{i,iii,v,vii,viii,ix,x,xi}{i,iii,vi,vi,vii,viii,ix,x,xi}{i,x,xi}
[0482] {i,iv,v,vi,vi,viii,ix,x}{i,iv,v,vi,vii,viii,ix,xi}{i,iv,v,vi,vi,vii,viii,x,xi}{i,iv,v,vi,vi,vii,ix,x,xi}{i,iv,v,vi,vi,vii,ix,x,xi}
[0483] {i,iv,v,vi,viii,ix,x,xi}{i,iv,v,vii,viii,ix,x,xi}{i,iv,vi,vii,viii,ix,x,xi}{i,v,vi,vii,viii,ix,x,xi}{i,v,vi,vii,viii,ix,x,xi}
[0484] {ii,iii,v,v,v,v,v,vii,ix}{ii,iii,v,v,v,v,viii,x}{ii,iii,v,v,v,vi,vii,xi}{ii,iii,v,v,vi,vi,ix,x}
[0485] {ii,iii,iv,v,v,v,v,i,x,xi}{ii,iii,iv,v,v,vi,vi,x,xi}{ii,iii,iv,v,v,viii,x,x}{ii,iii,iv,v,v,viii,x,xi} {ii,iii,iv,v,v,v,v,v,x,xi}{ii,iii,iv,v,v,vi,x,x,xi}{ii,iii,iv,v,v,vii,viii,x,x}{ii,iii,iv,v,v,vii,viii,ix,xi}
[0486] {ii,iii,iv,v,vi,vi,vii,x,xi}{ii,iii,iv,v,vii,ix,x,xi}{ii,iii,iv,v,viii,ix,x,xi}{ii,iii,iv,vi,vii,x,xi}{ii,iii,iv,vi,vii,viii,ix,x}
[0487] {ii,iii,iv,vi,vi,vii,vii,viii,x,xi}{ii,iii,iii,vi,vi,vii, ix,x,xi}{ii,iii,iv,vi,vii, ix,x,xi}{ii,iii,iv,vi,viii, ix,x,xi}
[0488] {ii,iii,iv,vi,viii,vii,vii,xi,xi}{ii,iii,v,vi,vii,viii,ix,x}{ii,iii,v,vi,vii,viii,ix,xi i,v,vi,vi,viii,x,xi}
[0489] {ii,iii,v,vi,vi,vi,x,x,xi}{ii,iii,v,vi,viii,ix,x,xi}{ii,iii,v,vii,viii,iii, ix,x,xi}{ii,ii,iii,vi,vii,viii, ix,x,xi v,
[0490] v,vi,vii,viii,ix,x}{ii,iv,v,vi,vi,viii,ix,xi}{ii,iv,v,vi,vii,viii,x,xi}{ii,iv,v,vi,vi,ix,x,xi}
[0491] {ii,iv,v,vi,viii,ix,x,xi v,v,vii,viii,ix,x,xi}{ii,iv,vi,vii,viii,ix,x,xi}{ii,v,vi,vii,viii,ix,x,xi}{ii,v,vi,vii,viii,ix,x,xi}
[0492] {iii,iv,v,v,vi,v,v,vi,vii,x,xi}{iii,iv,v,v,vi,vii,viii,x,xi}{iii,v,v,v,vii,x,x,xi}{iii,v,v,v,vii,x,x,xi}
[0493] {iii,iv,v,vi,viii,ix,x,xi}{iii,iv,v,vii,viii,viii,ix,x,xi}{iii,iv,vi,vii,viii,viii,ix,x,xi}{iii,v,vi,vi,vii,viii,ix,x,xi}{iii,x,xi}
[0494] {i,v,v,vi,viii,ix,x,xi}{i,ii,iii,iv,v,vi,vii,viii,ix}{i,ii,iii,iv,v,vi,vii,viii,x}{i,ii,iii,iv,v,v,vi,viii,xi}{i,ii,iii,iv,v,vi,vi,x}{i,ii,iii,iv,v,vi,vi,i,x,x i}{i,ii,iii,iv,v,v,v,x,xi}{i,ii,iii,v,v,v,viii,x,x}{i,ii,iii,v,v,vi,viii,ix,x i}{i,ii,iii,iv,v,v,viii,x,xi}{i,ii,iii,iv,v,vi,x,x,xi}{i,ii,iii,iv,v,vii,viii,ix,x}
[0495] {i,ii,iii,iv,v,vii,viii,ix,xi}{i,ii,iii,ii,v,vii,viii,x,xi}{i,ii,iii,iv,v,vii,ix,x,xi}{i,ii,iii,iv,v,viii},xi,x,x
[0496] {i,ii,iii,v,v,vi,vi,vii,ix,x}{i,ii,iii,ii,vi,vi,viii, ix,xi}{i,ii,iii,iv,vi,vi,vii,x,xi}{i,ii,iii, iv,vi,vi,xi,x,x}{i,ii,iii,iv,v,viii,ix,x,xi}{i,ii,iii,iv,vii,viii,ix,x,xi}{i,ii,iii,v,vi,vii,viii,iii,x,x}{i,ii,iii,v,vi,vii,xii}viii
[0497] {i,ii,iii,v,vi,vi,vii,x,xi}{i,ii,iii,v,vii,vi,x,x,xi}{i,ii,iii,v,vi,viii, ix,x,xi}{i,ii,iii,v,vii,viii}{i,xi,x,
[0498] ii,iii,vi,vi,viii,ix,x,xi}{i,ii,iv,v,v,vi,vii,vii,ix,x}{i,ii,iv,v,v,vii,viii,ix,xi}{i,ii,iv,v,vi,vii,vii,x,xi}
[0499] {i,ii,v,v,v,vi,vi,x,x,xi}{i,ii,v,v,v,vi,viii,x,x,xi}{i,ii,iv,v,v,vii,viii,ix,x,xi}{i,ii,iv,vi,vii,viii,ix,x,xi}
[0500] {i,ii,v,v,vii,viii,x,xi}{i,iii,iv,v,v,vii,viii,ix,x}{i,iii,iv,v,vi,vii,viii,ix,xi}{i,iii,iv,v,v,vi,vi,viii,x,xi}{i,iii,iv,v,vi,vii,x,x,xi}{i,iii,iv,v,vi,viii,ix,x,xi}{i ,iii,iv,v,vii,viii,ix,x,xi}{i,iii,iv,vi,vi,viii,ix,x,xi}{i,iii,v,vi,vii,viii,ix,x,xi}{ i,iv,v,vi,vii,viii,ix,x,xi}{ii,iii,iv,v,vi,vi,viii,ix,x}{ii,iii,iv,v,vi,vii,viii,ix,xi}
[0501] {ii,iii,iv,v,v,vi,vi,viii,x,xi}{ii,iii,iv,v,vi,vii,x,x,xi}{ii,iii,iv,v,vi,viii,ix,x,xi}{ii,iii,
[0502] iv,v,vii,viii,ix,x,xi}{ii,iii,iv,vi,vii,viii,ix,x,xi}{ii,iii,v,vi,vii,viii,ix,x,xi}{ii,iv,v,vi,vii,viii,ix,x,xi}
[0503] {iii,iv,v,vi,vii,viii,ix,x,xi}{i,ii,iii,iv,v,vi,vii,viii,ix,x}{i,ii,iii,iv,v,vi,vii,viii,ix,xi}
[0504] {i,ii,iii,iv,v,vi,vii,viii,x,xi}{i,ii,iii,iv,v,vi,vii,ix,x,xi}{i,ii,iii,iv,v,vi,viii,ix,x,xi}
[0505] {i,ii,iii,iv,v,vii,viii,ix,x,xi}{i,ii,iii,iv,vi,vii,viii,ix,x,xi}{i,ii,iii,v,vi,vii,viii,ix,x,xi}{i,ii,iv,v,vi,vii ,viii,ix,x,xi}{i,iii,iv,v,vi,vii,viii,ix,x,xi}{ii,iii,iv,v,vi,vii,viii,ix,x,xi} or {i,ii,iii,iv,v,vi,viiviii,ix,x,xi}.
[0506] If the variant includes any one of (i) and (iii) to (xi), it may further include mutations at one or more of the following positions: Y51, N55, and F56, such as Y51N55, Y51 / F56, N55 / F56, or Y51 / N55 / F56.
[0507] In (i), the variant may include any number and combination of mutations from N40, D43, E44, S54, S57, Q62, R97, E101, E124, E131, R142, T150, and R192. In (i), the variant preferably includes one or more mutations located at the following positions (i.e., mutations occurring at one or more of the following positions): N40, D43, E44, S54, S57, Q62, E101, E131, and T150. In (i), the variant preferably includes one or more mutations located at the following positions (i.e., mutations occurring at one or more of the following positions): N40, D43, E44, E101, and E131. In (i), the variant preferably includes mutations located at S54 and / or S57. In (i), the variant more preferably includes one or more of (a) S54 and / or S57 and (b) Y51, N55 and F56, such as mutations in Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56. If S54 and / or S57 are missing in (xi), then they cannot mutate in (i), and vice versa. In (i), the variant preferably includes a mutation in T150, such as T150I. Alternatively, the variant preferably includes one or more of (a) T150 and (b) Y51, N55 and F56, such as mutations in Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56. In (i), the variant preferably includes a mutation located at Q62, such as Q62R or Q62K. Alternatively, the variant preferably includes one or more of (a) Q62 and (b) Y51, N55, and F56, such as mutations located at Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56, or Y51 / N55 / F56. The variant may include mutations located at D43, E44, Q62, or any combination thereof, such as D43, E44, Q62, D43 / E44, D43 / Q62, E44 / Q62, or D43 / E44 / Q62. Alternatively, the variants preferably include one or more of (a) D43, E44, Q62, D43 / E44, D43 / Q62, E44 / Q62 or D43 / E44 / Q62 and (b) Y51, N55 and F56, such as mutations in Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56.
[0508] In (ii) and elsewhere in this application, the symbol / denotes "and," for example, Y51 / N55 refers to Y51 and N55. In (ii), the variant preferably includes a mutation at Y51 / N55. It has been proposed that the contractile structure in CsgG consists of three stacked concentric loops formed by the side chains of residues Y51, N55, and F56 (Goyal et al, 2014, Nature, 516, 250-253). In (ii), mutations in these residues can thus reduce the number of nucleotides contributing to the current as the polynucleotide moves through the pore, thereby making it easier to identify the direct relationship between the observed current (when the polynucleotide moves through the pore) and the polynucleotide. Y56 can be mutated in any manner discussed below with respect to the variants and pores available in the method of the invention.
[0509] In (v), the variant may include N102R, N102F, N102Y, or N102W. The variant preferably includes (a) N102R, N102F, N102Y, or N102W and (b) mutations at one or more locations in Y51, N55, and F56, such as those at Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56, or Y51 / N55 / F56.
[0510] In (xi), any number and combination of the following positions: K49, P50, Y51, P52, A53, S54, N55, F56, and S57 may be omitted. Preferably, one or more of K49, P50, Y51, P52, A53, S54, N55, and S57 may be omitted. If any of Y51, N55, and F56 is omitted in (xi), then it / they cannot mutate in (ii), and vice versa.
[0511] In (i), the variant preferably includes one or more of the following substitutions: N40R, N40K, D43N, D43Q, D43R, D43K, E44N, E44Q, E44R, E44K, S54P, S57P, Q62R, Q62K, R97N, R97G, R97L, E101N, E101Q, E101R, E101K, E101F, E101Y, E101W. E124N, E124Q, E124R, E124K, E124F, E124Y, E124W, E131D, R142E, R142N, T150I, R192E and R192N, e.g., N40R, N40K, D43N, D43Q, D43R, D43K, E44N, E44Q, E44R, E44K, S54P, S57P, Q62R, Q62K, E101 One or more of N, E101 Q, E101 R, E101K, E101F, E101Y, E101W, E131D, and T150I, or one or more of N40R, N40K, D43N, D43Q, D43R, D43K, E44N, E44Q, E44R, E44K, E101N, E101Q, E101R, E101K, E101F, E101Y, E101W, and E131D. The variants may include any number of these substitutions and any combination of these substitutions. In (i), the variants preferably include S54P and / or S57P. In (i), the mutant preferably includes (a) S54P and / or S57P and (b) mutations located at one or more positions in Y51, N55, and F56, such as Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56, or Y51 / N55 / F56. The mutation located at one or more positions in Y51, N55, and F56 can be any of the mutations described below. In (i), the variant preferably includes F56A / S57P or S54P / F56A. The variant preferably includes T150I. Alternatively, the variants preferably include one or more of (a) T150I and (b) Y51, N55 and F56, such as mutations in Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56.
[0512] In (i), the variant preferably includes Q62R or Q62K. Alternatively, the variant preferably includes (a) Q62R or Q62K and (b) mutations at one or more of the following positions: Y51, N55, and F56, such as Y51 / N55, Y51 / F56, N55 / F56, or Y51 / N55 / F56. The variant may include D43N, E44N, Q62R, or Q62K or any combination thereof, such as D43N, E44N, Q62R, Q62K, D43N / E44N, D43N / Q62R, D43N / Q62K, E44N / Q62R, E44N / Q62K, D43N / E44N / Q62R, or D43N / E44N / Q62K. Alternatively, the variants preferably include (a) D43N, E44N, Q62R, Q62K, D43N / E44N, D43N / Q62R, D43N / Q62K, E44N / Q62R, E44N / Q62K, D43N / E44N / Q62R or D43N / E44N / Q62K and (b) mutations at one or more of the following locations: Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56.
[0513] In (i), the variant preferably includes D43N.
[0514] In (i), the variants preferably include E101R, E101S, E101F or E101N.
[0515] In (i), the variants preferably include E124N, E124Q, E124R, E124K, E124F, E124Y, E124W or E124D, such as E124N.
[0516] In (i), the variants preferably include R142E and R142N.
[0517] In (i), the variants preferably include R97N, R97G or R97L.
[0518] In (i), the variants preferably include R192E and R192N.
[0519] In (ii), the variants preferably include F56N / N55Q, F56N / N55R, F56N / N55K, and F56N / N55S.
[0520] F56N / N55G,F56N / N55A,F56N / N55T,F56Q / N55Q,F56Q / N55R,F56Q / N55K,F56Q / N55SF56Q / N55G,F56Q / N55A,F56Q / N55T,F56R / N55Q,F56R / N55R,F56R / N55K,F56R / N55SF56R / N55G,F56R / N55A,F56R / N55T,F56S / N55Q,F56S / N55R,F56S / N55K,F56S / N55S,F56S / N55G F56S / N55A,F56S / N55T,F56G / N55Q,F56G / N55R,F56G / N55K,F56G / N55S,F56G / N55G,F56G / N55A,F56G / N55T,F56A / N55Q,F56A / N55R,F56A / N55K,F56A / N55S,F56A / N55G,F56A / N55A F56A / N55T,F56K / N55Q,F56K / N55R,F56K / N55K,F56K / N55S,F56K / N55G,F56K / N55A,F56K / N55T F56N / Y51L,F56N / Y51V,F56N / Y51A,F56N / Y51N,F56N / Y51Q,F56N / Y51S,F56N / Y51G,F56Q / Y51L F56Q / Y51V,F56Q / Y51A,F56Q / Y51N,F56Q / Y51Q,F56Q / Y51S,F56Q / Y51G,F56R / Y51L,F56R / Y51V,F56R / Y51A,F56R / Y51N,F56R / Y51Q,F56R / Y51S,F56R / Y51G,F56S / Y51L,F56S / Y51V,F56S / Y51A F56S / Y51N,F56S / Y51Q,F56S / Y51S,F56S / Y51G,F56G / Y51L,F56G / Y51V,F56G / Y51A,F56G / Y51NF56G / Y51Q,F56G / Y51S,F56G / Y51G,F56A / Y51L,F56A / Y51V,F56A / Y51A,F56A / Y51N,F56A / Y51Q F56A / Y51S,F56A / Y51G,F56K / Y51L,F56K / Y51V,F56K / Y51A,F56K / Y51N,F56K / Y51Q,F56K / Y51S / F56K / Y51G,N55Q / Y51L,N55Q / Y51V,N55Q / Y51A,N55Q / Y51N,N55Q / Y51Q,N55Q / Y51S / N55Q / Y51G,N55R / Y51L,N55R / Y51V,N55R / Y51A,N55R / Y51N,N55R / Y51Q,N55R / Y51S / N55R / Y51G,N55K / Y51L,N55K / Y51V,N55K / Y51A,N55K / Y51N,N55K / Y51Q,N55K / Y51S / N55K / Y51G,N55S / Y51L,N55S / Y51V,N55S / Y51A,N55S / Y51N,N55S / Y51Q,N55S / Y51S / N55S / Y51G,N55G / Y51L,N55G / Y51V,N55G / Y51A,N55G / Y51N,N55G / Y51Q,N55G / Y51S N55G / Y51G,N55A / Y51L,N55A / Y51V,N55A / T51A,N55A / T51N,N55A / Y51Q,N55AY51S N55A / Y51G,N55T / Y51L,N55T / Y51V,N55T / T51A,N55T / T51N,N55T / Y51Q,N55T / Y51S,N55T / Y51G,F56N / N55Q / Y51L,F56N / N55Q / Y51V,F56N / N55Q / Y51A,F56N / N55Q / Y51N,F56N / N55Q / Y51Q,F56N / N55Q / Y51S,F56N / N55Q / Y51G,F56N / N55R / Y51L,F56N / N55R / Y51V,F56N / N55R / Y51A,F56N / N55R / Y51N,F56N / N55R / Y51Q,F56N / N55R / Y51S,F56N / N55R / Y51G,F56N / N55K / Y51L F56N / N55K / Y51V,F56N / N55K / Y51 A,F56N / N55K / Y51N,F56N / N55K / Y51Q,F56N / N55K / Y51S,F56N / N55K / Y51G,F56N / N55S / Y51L,F56N / N55S / Y51V,F56N / N55S / Y51A,F56N / N55S / Y51N,F56N / N55S / Y51Q,F56N / N55S / Y51S,F56N / N55S / Y51G,F56N / N55G / Y51L,F56N / N55G / Y51V,F56N / N55G / Y51A,F56N / N55G / Y51N,F56N / N55G / Y51Q,F56N / N55G / Y51S,F56N / N55G / Y51G,F56N / N55A / Y51L,F56N / N55A / Y51V,F56N / N55A / Y51A,F56N / N55A / Y51N,F56N / N55A / Y51Q,F56N / N55A / Y51S,F56N / N55A / Y51G,F56N / N55T / Y51L,F56N / N55T / Y51V,F56N / N55T / Y51A,F56N / N55T / Y51N,F56N / N55T / Y51Q,F56N / N55T / Y51S,F56N / N55T / Y51G,F56Q / N55Q / Y51L,F56Q / N55Q / Y51V,F56Q / N55Q / Y51A,F56Q / N55Q / Y51N,
[0521] F56Q / N55Q / Y51Q,F56Q / N55Q / Y51S,F56Q / N55Q / Y51G,F56Q / N55R / Y51L,F56Q / N55R / Y51V,
[0522] F56Q / N55R / Y51A,F56Q / N55R / Y51N,F56Q / N55R / Y51Q,F56Q / N55R / Y51S,F56Q / N55R / Y51G,
[0523] F56Q / N55K / Y51L,F56Q / N55K / Y51V,F56Q / N55K / Y51A,F56Q / N55K / T51N,
[0524] F56Q / N55K / Y51Q,F56Q / N55K / Y51S,F56Q / N55K / Y51G,F56Q / N55S / Y51L,
[0525] F56Q / N55S / Y51V,F56Q / N55S / Y51A,F56Q / N55S / Y51N,F56Q / N55S / Y51Q,F56Q / N55S / Y51S,
[0526] F56Q / N55S / Y51G,F56Q / N55G / Y51L,F56Q / N55G / Y51V,F56Q / N55G / Y51A,F56Q / N55G / Y51N,
[0527] F56Q / N55G / Y51Q,F56Q / N55G / Y51S,F56Q / N55G / Y51G,F56Q / N55A / Y51L,F56Q / N55A / Y51V,
[0528] F56Q / N55A / Y51A,F56Q / N55A / Y51N,F56Q / N55A / Y51Q,F56Q / N55A / Y51S,
[0529] F56Q / N55A / Y51G,F56Q / N55T / Y51L,F56Q / N55T / Y51V,F56Q / N55T / Y51A,F56Q / N55T / Y51N,
[0530] F56Q / N55T / Y51Q,F56Q / N55T / Y51S,F56Q / N55T / Y51G,F56R / N55Q / Y51L,F56R / N55Q / Y51V,
[0531] F56R / N55Q / Y51A,F56R / N55Q / Y51N,F56R / N55Q / Y51Q,F56R / N55Q / Y51S,F56R / N55Q / Y51G,
[0532] F56R / N55R / Y51L,F56R / N55R / Y51V,F56R / N55R / Y51A,F56R / N55R / Y51N,F56R / N55R / Y51Q,
[0533] F56R / N55R / Y51S,F56R / N55R / Y51G,F56R / N55K / Y51L,F56R / N55K / Y51V,F56R / N55K / Y51A,
[0534] F56R / N55K / Y51N,F56R / N55K / Y51Q,F56R / N55K / Y51S,F56R / N55K / Y51G,F56R / N55S / Y51L,
[0535] F56R / N55S / Y51V,F56R / N55S / Y51A,F56R / N55S / Y51N,F56R / N55S / Y51Q,F56R / N55S / Y51S,
[0536] F56R / N55S / Y51G,F56R / N55G / Y51L,F56R / N55G / Y51V,F56R / N55G / Y51A,F56R / N55G / Y51N,
[0537] F56R / N55G / Y51Q,F56R / N55G / Y51S,F56R / N55G / Y51G;F56R / N55A / Y51L,F56R / N55A / Y51V,
[0538] F56R / N55A / Y51A,F56R / N55A / Y51N,F56R / N55A / Y51Q,F56R / N55A / Y51S,F56R / N55A / Y51G
[0539] F56R / N55T / Y51L,F56R / N55T / Y51V,F56R / N55T / Y51A,F56R / N55T / Y51N,F56R / N55T / Y51Q
[0540] F56R / N55T / Y51S,F56R / N55T / Y51G,F56S / N55Q / Y51L,F56S / N55Q / Y51V,F56S / N55Q / Y51A
[0541] F56S / N55Q / Y51N,F56S / N55Q / Y51Q,F56S / N55Q / Y51S,F56S / N55Q / Y51G,F56S / N55R / Y51L
[0542] F56S / N55R / Y51V,F56S / N55R / Y51A,F56S / N55R / Y51N,F56S / N55R / Y51Q,F56S / N55R / Y51S
[0543] F56S / N55R / Y51G,F56S / N55K / Y51L,F56S / N55K / Y51V,F56S / N55K / Y51A,F56S / N55K / Y51N
[0544] F56S / N55K / Y51Q,F56S / N55K / Y51S,F56S / N55K / Y51G,F56S / N55S / Y51L,F56S / N55S / Y51V
[0545] F56S / N55S / Y51A,F56S / N55S / Y51N,F56S / N55S / Y51Q,F56S / N55S / Y51S,F56S / N55S / Y51G
[0546] F56S / N55G / Y51L,F56S / N55G / Y51V,F56S / N55G / Y51A,F56S / N55G / Y51N,F56S / N55G / Y51Q
[0547] F56S / N55G / Y51S,F56S / N55G / Y51G,F56S / N55A / Y51L,F56S / N55A / Y51V,F56S / N55A / Y51A
[0548] F56S / N55A / Y51N,F56S / N55A / Y51Q,F56S / N55A / Y51S,F56S / N55A / Y51G,F56S / N55T / Y51L
[0549] F56S / N55T / Y51V,F56S / N55T / Y51A,F56S / N55T / Y51N,F56S / N55T / Y51Q,F56S / N55T / Y51S
[0550] F56S / N55T / Y51G,F56G / N55Q / Y51L,F56G / N55Q / Y51V,F56G / N55Q / Y51A,F56G / N55Q / Y51N
[0551] F56G / N55Q / Y51Q,F56G / N55Q / Y51S,F56G / N55Q / Y51G;F56G / N55R / Y51L,F56G / N55R / Y51V
[0552] F56G / N55R / Y51A,F56G / N55R / Y51N,F56G / N55R / Y51Q,F56G / N55R / Y51S,F56G / N55R / Y51G
[0553] F56G / N55K / Y51L,F56G / N55K / Y51V,F56G / N55K / Y51A,F56G / N55K / Y51N,
[0554] F56G / N55K / Y51Q F56G / N55K / Y51S,F56G / N55K / Y51G,F56G / N55S / Y51L,F56G / N55S / Y51V,
[0555] F56G / N55S / Y51A F56G / N55S / Y51N,F56G / N55S / Y51Q,F56G / N55S / Y51S,F56G / N55S / Y51G,
[0556] F56G / N55G / Y51L F56G / N55G / Y51V,F56G / N55G / Y51A,F56G / N55G / Y51N,F56G / N55G / Y51Q,
[0557] F56G / N55G / Y51S F56G / N55G / Y51G,F56G / N55A / Y51L,F56G / N55A / Y51V,F56G / N55A / Y51A,
[0558] F56G / N55A / Y51N F56G / N55A / Y51Q,F56G / N55A / Y51S,F56G / N55A / Y51G,F56G / N55T / Y51L,
[0559] F56G / N55T / Y51V,F56G / N55T / Y51A,F56G / N55T / Y51N,F56G / N55T / Y51Q,F56G / N55T / Y51S,F56G / N55T / Y51G,F56A / N55Q / Y51L,F56A / N55Q / Y51V,F56A / N55Q / Y51A,F56A / N55Q / Y51N,F56A / N55Q / Y51Q,F56A / N55Q / Y51S,F56A / N55Q / Y51G,F56A / N55R / Y51L,F56A / N55R / Y51 V,F56A / N55R / Y51A,F56A / N55R / Y51N,F56A / N55R / Y51Q,F56A / N55R / Y51S,F56A / N55R / Y51G,F56A / N55K / Y51L,F56A / N55K / Y51V,F56A / N55K / Y51A,F56A / N55K / Y51N,F56A / N55K / Y51Q,F56A / N55K / Y51S,F56A / N55K / Y51G,F56A / N55S / Y51L,F56A / N55S / Y51V,F56A / N55S / Y51A,F56A / N55S / Y51N,F56A / N55S / Y51Q,F56A / N55S / Y51S,F56A / N55S / Y51G,F56A / N55G / Y51L,F56A / N55G / Y51V,F56A / N55G / Y51A,F56A / N55G / Y51N,F56A / N55G / Y51Q,F56A / N55G / Y51S,F56A / N55G / Y51G,F56A / N55A / Y51L,F56A / N55A / Y51V,F56A / N55A / Y51A,F56A / N55A / Y51N,F56A / N55A / Y51Q,F56A / N55A / Y51S,F56A / N55A / Y51G,F56A / N55T / Y51L,F56A / N55T / Y51V,F56A / N55T / Y51A,F56A / N55T / Y51N,F56A / N55T / Y51Q,F56A / N55T / Y51S,F56A / N55T / Y51G,F56K / N55Q / Y51L,F56K / N55Q / Y51V,F56K / N55Q / Y51A,F56K / N55Q / Y51N,F56K / N55Q / Y51Q,F56K / N55Q / Y51S,F56K / N55Q / Y51G,F56K / N55R / Y51L,F56K / N55R / Y51V,F56K / N55R / Y51A,F56K / N55R / Y51N,F56K / N55R / Y51Q,F56K / N55R / Y51S,F56K / N55R / Y51G,F56K / N55K / Y51L,F56K / N55K / Y51V,F56K / N 55K / Y51A,F56K / N55K / Y51N,F56K / N55K / Y51Q,F56K / N55K / Y51S,F56K / N55K / Y51G,F56K / N55S / Y5 1L,F56K / N55S / Y51V,F56K / N55S / Y51A,F56K / N55S / Y51N,F56K / N55S / Y51Q,F56K / N55S / Y51S,F56 K / N55S / Y51G,F56K / N55G / Y51L,F56K / N55G / Y51V,F56K / N55G / Y51A,F56K / N55G / Y51N,F56K / N55G / Y51Q,F56K / N55G / Y51S,F56K / N55G / Y51G,F56K / N55A / Y51L,F56K / N55A / Y51V,F56K / N55A / Y51A,F 56K / N55A / Y51N,F56K / N55A / Y51Q,F56K / N55A / Y51S,F56K / N55A / Y51G,F56K / N55T / Y51L,F56K / N5 5T / Y51V,F56K / N55T / Y51A,F56K / N55T / Y51N,F56K / N55T / Y51Q,F56K / N55T / Y51S,F56K / N55T / Y51 G, F56E / N55R, F56E / N55K, F56D / N55R, F56D / N55K, F56R / N55E, F56R / N55D, F56K / N55E or F56K / N55D. ,
[0560] In (ii), the variants preferably include Y51R / F56Q, Y51N / F56N, Y51M / F56Q, Y51L / F56Q, Y51I / F56Q, Y51V / F56Q, Y51A / F56Q, Y51P / F56Q, Y51G / F56Q, Y51C / F56Q, Y51Q / F56Q, Y51N / F56Q, Y51S / F56Q, Y51E / F56Q, Y51D / F56Q, Y51K / F56Q, or Y51H / F56Q.
[0561] In (ii), the variants preferably include Y51T / F56Q, Y51Q / F56Q or Y51A / F56Q.
[0562] In (ii), the variants preferably include Y51T / F56F, Y51T / F56M, Y51T / F56L, Y51T / F56I, Y51T / F56V, Y51T / F56A, Y51T / F56P, Y51T / F56G, Y51T / F56C, Y51T / F56Q, Y51T / F56N, Y51T / F56T, Y51T / F56S, Y51T / F56E, Y51T / F56D, Y51T / F56K, Y51T / F56H or Y51T / F56R.
[0563] In (ii), the variants preferably include Y51T / N55Q, Y51T / N55S or Y51T / N55A.
[0564] In (ii), the variants preferably include Y51A / F56F, Y51A / F56L, Y51A / F56I, Y51A / F56V, Y51A / F56A, Y51A / F56P, Y51A / F56G, Y51A / F56C, Y51A / F56Q, Y51A / F56N, Y51A / F56T, Y51A / F56S, Y51A / F56E, Y51A / F56D, Y51A / F56K, Y51A / F56H, or Y51A / F56R.
[0565] In (ii), the variants preferably include Y51C / F56A, Y51E / F56A, Y51D / F56A, Y51K / F56A, Y51H / F56A, Y51Q / F56A, Y51N / F56A, Y51S / F56A, Y51P / F56A or Y51V / F56A.
[0566] In (xi), the variant preferably includes the deletion of Y51 / P52, Y51 / P52 / A53, P50 to P52, P50 to A53, K49 to Y51, K49 to A53 and substitutions with monoproline (P), K49 to S54 and substitutions with monop, Y51 to A53, Y51 to S54, N55 / F56, N55 to S57, N55 / F56, substitutions with monop, N55 / F56 and substitutions with monoglycine (G), N55 / F56 and substitutions with monoalanine (A), N55 / F56 and substitutions with P and Y51. Substitutions of N, N55 / F56 with monop and Y51Q, substitutions of N55 / F56 with monop and Y51S, substitutions of N55 / F56 with monoglycine (G) and Y51N, substitutions of N55 / F56 with monoglycine (G) and Y51Q, substitutions of N55 / F56 with monoglycine (G) and Y51S, substitutions of N55 / F56 with monoalanine (A) and Y51N, substitutions of N55 / F56 with monoalanine (A), substitutions of Y51Q or N55 / F56 with monoalanine (A) and Y51S.
[0567] The variants more preferably include D195N / E203N, D195Q / E203N, and D195N / E203Q.
[0568] D195Q / E203Q, E201N / E203N, E201Q / E203N, E201N / E203Q, E201Q / E203Q, E185N / E203Q,
[0569] E185Q / E203Q, E185N / E203N, E185Q / E203N, D195N / E201N / E203N, D195Q / E201N / E203N,
[0570] D195N / E201Q / E203N, D195N / E201N / E203Q, D195Q / E201Q / E203N, D195Q / E201N / E203Q,
[0571] D195N / E201Q / E203Q, D195Q / E201Q / E203Q, D149N / E201N, D149Q / E201N, D149N / E201Q,
[0572] D149Q / E201Q, D149N / E201N / D195N, D149Q / E201N / D195N, D149N / E201Q / D195N,
[0573] D149N / E201N / D195Q,D149Q / E201Q / D195N,D149Q / E201N / D195Q,D149N / E201Q / D195Q,
[0574] D149Q / E201Q / D195Q, D149N / E203N, D149Q / E203N, D149N / E203Q, D149Q / E203Q,
[0575] D149N / E185N / E201N,D149Q / E185N / E201N,D149N / E185Q / E201N,D149N / E185N / E201Q,
[0576] D149Q / E185Q / E201N,D149Q / E185N / E201Q,D149N / E185Q / E201Q,D149Q / E185Q / E201Q,
[0577] D149N / E185N / E203N, D149Q / E185N / E203N, D149N / E185Q / E203N, D149N / E185N / E203Q,
[0578] D149Q / E185Q / E203N,D149Q / E185N / E203Q,D149N / E185Q / E203Q,D149Q / E185Q / E203Q,
[0579] D149N / E185N / E201N / E203N,D149Q / E185N / E201N / E203N,D149N / E185Q / E201N / E203N,
[0580] D149N / E185N / E201Q / E203N,D149N / E185N / E201N / E203Q,D149Q / E185Q / E201N / E203N,
[0581] D149Q / E185N / E201Q / E203N, D149Q / E185N / E201N / E203Q, D149N / E185Q / E201Q / E203N,
[0582] D149N / E185Q / E201N / E203Q,D149N / E185N / E201Q / E203Q,D149Q / E185Q / E201Q / E203Q,
[0583] D149Q / E185Q / E201N / E203Q,D149Q / E185N / E201Q / E203Q,D149N / E185Q / E201Q / E203Q,
[0584] D149Q / E185Q / E201Q / E203N,D149N / E185N / D195N / E201N / E203N,
[0585] D149Q / E185N / D195N / E201N / E203N,D149N / E185Q / D195N / E201N / E203N,
[0586] D149N / E185N / D195Q / E201N / E203N,D149N / E185N / D195N / E201Q / E203N,
[0587] D149N / E185N / D195N / E201N / E203Q,D149Q / E185Q / D195N / E201N / E203N,
[0588] D149Q / E185N / D195Q / E201 N / E203N, D149Q / E185N / D195N / E201Q / E203N,
[0589] D149Q / E185N / D195N / E201N / E203Q,D149N / E185Q / D195Q / E201N / E203N,
[0590] D149N / E185Q / D195N / E201Q / E203N,D149N / E185Q / D195N / E201N / E203Q,
[0591] D149N / E185N / D195Q / E201Q / E203N,D149N / E185N / D195Q / E201N / E203Q,
[0592] D149N / E185N / D195N / E201Q / E203Q,D149Q / E185Q / D195Q / E201N / E203N,
[0593] D149Q / E185Q / D195N / E201Q / E203N,D149Q / E185Q / D195N / E201N / E203Q,
[0594] D149Q / E185N / D195Q / E201Q / E203N,D149Q / E185N / D195Q / E201N / E203Q,
[0595] D149Q / E185N / D195N / E201Q / E203Q,D149N / E185Q / D195Q / E201Q / E203N,
[0596] D149N / E185Q / D195Q / E201N / E203Q,D149N / E185Q / D195N / E201Q / E203Q,
[0597] D149N / E185N / D195Q / E201Q / E203Q,D149Q / E185Q / D195Q / E201Q / E203N,
[0598] D149Q / E185Q / D195Q / E201N / E203Q,D149Q / E185Q / D195N / E201Q / E203Q,
[0599] D149Q / E185N / D195Q / E201Q / E203Q,D149N / E185Q / D195Q / E201Q / E203Q,
[0600] D149Q / E185Q / D195Q / E201Q / E203Q,D149N / E185R / E201N / E203N,D149Q / E185R / E201N / E203N,
[0601] D149N / E185R / E201Q / E203N,D149N / E185R / E201N / E203Q,D149Q / E185R / E201Q / E203N,
[0602] D149Q / E185R / E201N / E203Q,D149N / E185R / E201Q / E203Q,D149Q / E185R / E201Q / E203Q,
[0603] D149R / E185N / E201N / E203N,D149R / E185Q / E201N / E203N,D149R / E185N / E201Q / E203N,
[0604] D149R / E185N / E201N / E203Q,D149R / E185Q / E201Q / E203N,D149R / E185Q / E201N / E203Q,
[0605] D149R / E185N / E201Q / E203Q,D149R / E185Q / E201Q / E203Q,D149R / E185N / D195N / E201N / E203N,
[0606] D149R / E185Q / D195N / E201N / E203N,D149R / E185N / D195Q / E201N / E203N,
[0607] D149R / E185N / D195N / E201Q / E203N,D149R / E185Q / D195N / E201N / E203Q,
[0608] D149R / E185Q / D195Q / E201N / E203N,D149R / E185Q / D195N / E201Q / E203N,
[0609] D149R / E185Q / D195N / E201N / E203Q,D149R / E185N / D195Q / E201Q / E203N,
[0610] D149R / E185N / D195Q / E201N / E203Q,D149R / E185N / D195N / E201Q / E203Q,
[0611] D149R / E185Q / D195Q / E201Q / E203N,D149R / E185Q / D195Q / E201N / E203Q,
[0612] D149R / E185Q / D195N / E201Q / E203Q,D149R / E185N / D195Q / E201Q / E203Q,
[0613] D149R / E185Q / D195Q / E201Q / E203Q,D149N / E185R / D195N / E201N / E203N,
[0614] D149Q / E185R / D195N / E201N / E203N,D149N / E185R / D195Q / E201N / E203N,
[0615] D149N / E185R / D195N / E201Q / E203N,D149N / E185R / D195N / E201N / E203Q,
[0616] D149Q / E185R / D195Q / E201N / E203N,D149Q / E185R / D195N / E201Q / E203N,
[0617] D149Q / E185R / D195N / E201N / E203Q,D149N / E185R / D195Q / E201Q / E203N,
[0618] D149N / E185R / D195Q / E201N / E203Q,D149N / E185R / D195N / E201Q / E203Q,
[0619] D149Q / E185R / D195Q / E201Q / E203N,D149Q / E185R / D195Q / E201N / E203Q,
[0620] D149Q / E185R / D195N / E201Q / E203Q,D149N / E185R / D195Q / E201Q / E203Q,
[0621] D149Q / E185R / D195Q / E201Q / E203Q,D149N / E185R / D195N / E201R / E203N,
[0622] D149Q / E185R / D195N / E201R / E203N,D149N / E185R / D195Q / E201R / E203N,
[0623] D149N / E185R / D195N / E201R / E203Q,D149Q / E185R / D195Q / E201R / E203N,
[0624] D149Q / E185R / D195N / E201R / E203Q,D149N / E185R / D195Q / E201R / E203Q,
[0625] D149Q / E185R / D195Q / E201R / E203Q,E131D / K49R,E101N / N102F,E101N / N102Y,E101N / N102W,
[0626] E101F / N102F,E101F / N102Y,E101F / N102W,E101Y / N102F,E101Y / N102Y,E101Y / N102W,
[0627] E101W / N102F, E101W / N102Y, E101W / N102W, E101N / N102R, E101F / N102R, E101Y / N102R or E101W / N102F.
[0628] Preferred variants of the pores formed in this invention—where fewer nucleotides contribute to the current as the polynucleotide moves through the pore—include Y51A / F56A, Y51A / F56N, Y51I / F56A, Y51L / F56A, Y51T / F56A, Y51I / F56N, Y51L / F56N, or Y51T / F56N, or more preferably Y51I / F56A, Y51L / F56A, or Y51T / F56A. As described above, this makes it easier to identify a direct relationship between the observed current (when the polynucleotide moves through the pore) and the polynucleotide.
[0629] Preferred variants of forming pores that exhibit an increased range include abrupt changes at the following locations:
[0630] Y51, F56, D149, E185, E201 and E203;
[0631] N55 and F56;
[0632] Y51 and F56;
[0633] Y51, N55, and F56; or
[0634] F56 and N102.
[0635] Preferred variations that form an increased range of apertures include:
[0636] Y51 N, F56A, D149N, E185R, E201 N and E203N;
[0637] N55S and F56Q;
[0638] Y51A and F56A;
[0639] Y51A and F56N;
[0640] Y51I and F56A;
[0641] Y51L and F56A;
[0642] Y51T and F56A;
[0643] Y51I and F56N;
[0644] Y51L and F56N;
[0645] Y51T and F56N;
[0646] Y51T and F56Q;
[0647] Y51A, N55S and F56A;
[0648] Y51A, N55S and F56N;
[0649] Y51T, N55S, and F56Q; or
[0650] F56Q and N102R.
[0651] Preferred variants of the pore—where fewer nucleotides contribute to the current as the polynucleotide moves through the pore—include mutations at the following locations:
[0652] N55 and F56, such as N55X and F56Q, where X is any amino acid; or
[0653] Y51 and F56, for example Y51X and F56Q, where X is any amino acid.
[0654] Preferred variants of the orifice that exhibit increased flux include abrupt changes at the following locations:
[0655] D149, E185 and E203;
[0656] D149, E185, E201, and E203; or
[0657] D149, E185, D195, E201 and E203.
[0658] Preferred variants of the orifice that exhibit increased flux include:
[0659] D149N, E185N and E203N;
[0660] D149N, E185N, E201N and E203N;
[0661] D149N, E185R, D195N, E201N, and E203N; or
[0662] D149N, E185R, D195N, E201R and E203N.
[0663] Preferred variants forming the pore—in which polynucleotide capture is increased—include the following mutations:
[0664] D43N / Y51T / F56Q;
[0665] E44N / Y51T / F56Q;
[0666] D43N / E44N / Y51T / F56Q;
[0667] Y51T / F56Q / Q62R;
[0668] D43N / Y51T / F56Q / Q62R;
[0669] E44N / Y51T / F56Q / Q62R; or
[0670] D43N / E44N / Y51T / F56Q / Q62R.
[0671] Preferred variants include the following mutations:
[0672] D149R / E185R / E201R / E203R or Y51T / F56Q / D149R / E185R / E201R / E203R;
[0673] D149N / E185N / E201N / E203N or Y51T / F56Q / D149N / E185N / E201N / E203N;
[0674] E201R / E203R or Y51T / F56Q / E201R / E203R
[0675] E201N / E203R or Y51T / F56Q / E201N / E203R;
[0676] E203R or Y51T / F56Q / E203R;
[0677] E203N or Y51T / F56Q / E203N;
[0678] E201R or Y51TF56Q / E201 R;
[0679] E201N or Y51T / F56Q / E201 N;
[0680] E185R or Y51 T / F56Q / E185R;
[0681] E185N or Y51T / F56Q / E185N;
[0682] D149R or Y51T / F56Q / D149R;
[0683] D149N or Y51T / F56Q / D149N;
[0684] R142E or Y51T / F56Q / R142E;
[0685] R142N or Y51T / F56Q / R142N;
[0686] R192E or Y51T / F56Q / R192E; or
[0687] R192N or Y51T / F56Q / R192N.
[0688] Preferred variants include the following mutations:
[0689] Y51A / F56Q / E101N / N102R;
[0690] Y51A / F56Q / R97N / N102G;
[0691] Y51A / F56Q / R97N / N102R;
[0692] Y51A / F56Q / R97N;
[0693] Y51A / F56Q / R97G;
[0694] Y51A / F56Q / R97L;
[0695] Y51A / F56Q / N102R;
[0696] Y51A / F56Q / N102F;
[0697] Y51A / F56Q / N102G;
[0698] Y51A / F56Q / E101R;
[0699] Y51A / F56Q / E101F;
[0700] Y51A / F56Q / E101N; or
[0701] Y51A / F56Q / E101G.
[0702] The present invention also provides mutant CsgG monomers comprising variants of the sequence shown in SEQ ID NO:390, wherein said variants include a mutation at the T150 position. Preferred variants forming pores exhibiting increased insertion include T150I. A mutation at the T150 position, such as T150I, can be combined with any of the above-described mutations or combinations thereof.
[0703] The present invention also provides mutant CsgG monomers comprising variants of the sequence shown in SEQ ID NO:390, said variants comprising combinations of mutations present in the variants disclosed in the embodiments.
[0704] Methods for introducing or replacing naturally occurring amino acids are well known in the art. For example, at the relevant position of a polynucleotide encoding a mutant monomer, methionine (M) can be substituted for arginine (R) by replacing the codon of methionine (ATG) with the codon of arginine (CGT). The polynucleotide can then be expressed as described below.
[0705] Methods for introducing or replacing non-naturally occurring amino acids are also well known in the art. For example, non-naturally occurring amino acids can be introduced by including synthetic aminoacyl-tRNA in an IVTT system used to express mutant monomers. Alternatively, they can be introduced by expressing the mutant monomer in *E. coli* that is auxotrophic for the specific amino acid in the presence of synthetic (i.e., non-naturally occurring) analogs of those specific amino acids. If partial peptide synthesis is used to generate mutant monomers, they can also be prepared by naked linking.
[0706] Other monomers of the present invention
[0707] In another embodiment, the present invention provides a mutant CsgG monomer comprising a variant of the sequence shown in SEQ ID NO:390, wherein the variant comprises a mutation at one or more of positions Y51, N55, and F56. The variant may comprise a mutation at positions Y51; N55; F56; Y51 / N55; Y51 / F56; N55 / F56; or Y51 / N55 / F56. The variant preferably comprises a mutation at position Y51, N55, or F56. The variant may comprise any specific mutation at one or more of positions Y51, N55, and F56 discussed above, and any combination thereof. One or more of Y51, N55, and F56 may be substituted with any amino acid. Y51 may be substituted with F, M, L, I, V, A, P, G, C, Q, N, T, S, E, D, K, H, or R, for example, A, S, T, N, or Q. N55 can be replaced by F, M, L, I, V, A, P, G, C, Q, T, S, E, D, K, H, or R, for example, A, S, T, or Q. F56 can be replaced by M, L, I, V, A, P, G, C, Q, N, T, S, E, D, K, H, or R, for example, A, S, T, N, or Q.
[0708] The variant may further include one or more of the following: (i) one or more mutations at the following positions (i.e., mutations at one or more of the following positions): (i) N40, D43, E44, S54, S57, Q62, R97, E101, E124, E131, R142, T150 and R192; (iii) Q42R or Q42K; (iv) K49R; (v) N102R, N102F, N102Y or N102W; (vi) D149N, D149Q or D149R; (vii) E185N, E185Q or E185R; (viii) D195N, D195Q or D195R; (ix) E201N, E201Q or E201 R; (x) E203N, E203Q, or E203R; and (xi) missing one or more of the following positions: F48, K49, P50, Y51, P52, A53, S54, N55, F56, and S57. The variant may include any combination of (i) and (iii) through (xi) discussed above. The variant may include any implementation of (i) and (iii) through (xi) discussed above.
[0709] variants
[0710] In addition to the specific mutations described above, variants may include other mutations. Based on amino acid identity, the variant is preferably at least 50% homologous to the amino acid sequence of SEQ ID NO:390 over its entire length. More preferably, based on amino acid identity, the variant may be at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, and more preferably at least 95%, 97%, or 99% homologous to the amino acid sequence of SEQ ID NO:390 over its entire length. Over a continuous amino acid extension of 100 or more, such as 125, 150, 175, or 200 or more, there may be at least 80%, such as at least 85%, 90%, or 95% amino acid identity (“strict hard homology”).
[0711] Standard methods in the prior art can be used to determine homology. For example, the UWGCG software package provides the BESTFIT program for calculating homology, for example using its default settings (Devereux et al. (1984) Nucleic Acids Research 12, pp. 387-395). The PILEUP and BLAST algorithms can be used to calculate homology or sequence ordering (e.g., identifying equivalent residues or corresponding sequences (usually with their default settings)), for example as described in Altschul SF (1993) J Mol Evol 36290-300; Altschul, S. Fet al. (1990) J Mol Biol 215403-10. Software for performing BLAST analysis is publicly available from the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih.gov).
[0712] SEQ ID NO:390 is a wild-type CsgG monomer from *Escherichia coli* Str.K-12substr.MC4100. Variants of sequence SEQ ID NO:390 may include any substitutions present in another CsgG homolog. Preferred CsgG homologs are shown in SEQ ID: NO:391 to 395 and 414 to 429. These variants may include combinations of one or more substitutions present in SEQ ID: NO:391 to 395 and 414 to 429 compared to SEQ ID NO:390.
[0713] In addition to the amino acid sequences discussed above, amino acid substitutions can be made to the amino acid sequence of SEQ ID NO:390, for example, up to 1, 2, 3, 4, 5, 10, 20, or 30 substitutions. Conservative substitution involves replacing an amino acid with another amino acid having a similar chemical structure, similar chemical properties, or similar side chain volume. The introduced amino acid may have similar polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality, or charge compared to the amino acid it replaces. Alternatively, the conservative substitution may introduce another aromatic or aliphatic amino acid to replace a previously present aromatic or aliphatic amino acid. Conservative amino acid changes are well known in the art and can be selected based on the properties of the 20 major amino acids defined in Table 1 below. When amino acids have similar polarity, this can also be determined with reference to the hydrophilicity level of the amino acid side chains in Table 2.
[0714] Table 1 - Chemical properties of amino acids
[0715]
[0716] Table 2 - Hydrophilicity Level
[0717]
[0718]
[0719] One or more amino acid residues of the amino acid sequence of SEQ ID NO:390 may be additionally deleted from the above-described polypeptide. Up to 1, 2, 3, 4, 5, 10, 20, or 30 residues may be deleted.
[0720] The variant may contain the fragment of SEQ ID NO:390. The fragment retains the pore-forming activity. The length of the fragment may be at least 50, at least 100, at least 150, at least 200, or at least 250 amino acids. The fragment can be used to generate the pore. The fragment preferably includes the transmembrane domains of SEQ ID NO:390, namely K135-Q153 and S183-S208.
[0721] One or more amino acids may be added alternatively or additionally to the above-described polypeptide. An extension may be provided at the amino or carboxyl terminus of the amino acid sequence of SEQ ID NO:390 or its polypeptide variants or fragments thereof. The extension may be relatively short, for example, 1 to 10 amino acids in length. Alternatively, the extension may be longer, for example, reaching 50 or 100 amino acids. The carrier protein may be fused with the amino acid sequence according to the invention. Other fusion proteins are discussed in more detail below.
[0722] As described above, a variant is a polypeptide having an amino acid sequence derived from SEQ ID NO:390 and retaining its pore-forming ability. Variants typically include the regions responsible for pore formation in SEQ ID NO:390. The pore-forming ability of CsgG, which contains a β-barrel structure, is provided by β-sheets in each subunit. Variants of SEQ ID NO:390 typically include the regions in SEQ ID NO:390 that form β-variants, namely K135-Q153 and S183-S208. One or more modifications can be made to the regions in SEQ ID NO:390 that form β-sheets, as long as the resulting variant retains its pore-forming ability. Variants of SEQ ID NO:390 preferably include one or more modifications, such as substitution, addition, or deletion, in their α-helical and / or loop structural regions.
[0723] Monomers derived from CsgG can be modified to aid in their identification or purification, for example, by adding a streptavidin tag or by adding a signaling sequence to promote their secretion from cells, wherein the monomer does not naturally contain such a sequence. Other suitable tags are discussed in more detail below. The monomers can be labeled with display tags. The display tag can be any suitable tag that allows the monomer to be detected. Suitable tags are described below.
[0724] Monomers derived from CsgG can also be prepared using D-amino acids. For example, monomers derived from CsgG may include a mixture of L-amino acids and D-amino acids. This is common practice in the field of preparing the aforementioned proteins or peptides.
[0725] Monomers derived from CsgG contain one or more specific modifications to facilitate nucleotide differentiation. Monomers derived from CsgG may also contain other non-specific modifications, provided they do not interfere with pore formation. Many non-specific side-chain modifications are known in the art and can be performed on the side chains of monomers derived from CsgG. These modifications include, for example, amino acid reductive alkylation via reaction with an aldehyde followed by reduction with NaBH4, amidoylation with methylacetimidate, or acylation with acetic anhydride to obtain the amino acid.
[0726] The CsgG-derived monomer can be produced using methods known in the art. The CsgG-derived monomer can be prepared synthetically or by recombination methods. For example, the monomer can be synthesized by in vitro translation and transcription (IVTT). Suitable methods for producing the pores and monomers are discussed in international applications PCT / GB09 / 001690 (published as WO 2010 / 004273), PCT / GB09 / 001679 (published as WO 2010 / 004265), or PCT / GB10 / 000133 (published as WO 2010 / 086603). The method for inserting the pores into the membrane is described below.
[0727] In some embodiments, the mutant monomer is chemically modified. The mutant monomer can be chemically modified in any manner and at any site. Preferably, the mutant monomer is chemically modified by linking the molecule to one or more cysteine residues (cysteine linkage), linking the molecule to one or more lysine residues, linking the molecule to one or more non-natural amino acids, enzymatic modification of the antigenic epitope, or terminal modification. Suitable methods for implementing the above modifications are known in the art. The mutant monomer can be chemically modified by linking any molecule. For example, the mutant monomer can be chemically modified by linking a dye or fluorophore.
[0728] In some embodiments, the mutant monomer is chemically modified with a molecular adapter that promotes the interaction between the pore containing the monomer and the target nucleotide or polynucleotide sequence. The presence of the adapter improves the host-guest chemistry between the pore and the nucleotide or polynucleotide, thereby enhancing the sequencing capability of the pore formed from the mutant monomer. The principles of host-guest chemistry are well known in the art. The adapter affects the physical or chemical properties of the pore, improving its interaction with the nucleotide or polynucleotide sequence. The adapter can alter the charge of the barrel or channel structure of the pore, or specifically interact with or bind to the nucleotide or polynucleotide sequence to promote its interaction with the pore.
[0729] The molecular adapter is preferably a cyclic molecule, cyclodextrin, a substance capable of hybridization, a DNA binder or chelator, a peptide or peptide analog, a synthetic polymer, an aromatic planar molecule, a positively charged small molecule, or a small molecule capable of hydrogen bonding.
[0730] The adapter can be annular. The annular adapter preferably has the same symmetry as the hole. Since CsgG typically has 8 or 9 subunits around a central axis, the adapter preferably has 8- or 9-fold symmetry. This will be discussed in more detail below.
[0731] The adapter typically interacts with the nucleotide or polynucleotide sequence via host-guest chemistry. The adapter is generally capable of interacting with the nucleotide or polynucleotide sequence. The adapter includes one or more chemical groups capable of interacting with the nucleotide or polynucleotide sequence. These one or more chemical groups preferably interact with the nucleotide or polynucleotide sequence via non-covalent interactions, such as hydrophobic interactions, hydrogen bonds, van der Waals forces, π-cation interactions, and / or electrostatic forces. The one or more chemical groups capable of interacting with the nucleotide or polynucleotide sequence are preferably positively charged. More preferably, the one or more chemical groups capable of interacting with the nucleotide or polynucleotide sequence include amino groups. The amino groups may be attached to primary, secondary, or tertiary carbon atoms. The adapter even more preferably includes a ring of amino groups, such as a ring with 6, 7, or 8 amino groups. Most preferably, the adapter includes a ring with 8 amino groups. The protonated ring of amino groups can interact with negatively charged phosphate groups in the nucleotide or polynucleotide sequence.
[0732] The correct positioning of the adapter within the pore can be facilitated by host-guest chemistry between the adapter and the pore containing the mutant monomer. The adapter preferably comprises one or more chemical groups capable of interacting with one or more amino acids in the pore. More preferably, the adapter comprises one or more chemical groups capable of interacting with one or more amino acids in the pore through non-covalent interactions, such as hydrophobic interactions, hydrogen bonds, van der Waals forces, π-cation interactions, and / or electrostatic forces. The chemical groups capable of interacting with one or more amino acids in the pore are typically hydroxyl or amine groups. The hydroxyl group may be attached to a primary, secondary, or tertiary carbon atom. The hydroxyl group may form hydrogen bonds with uncharged amino acids in the pore. Any adapter that facilitates the interaction between the pore and the nucleotide or polynucleotide sequence can be used.
[0733] Suitable adapters include, but are not limited to, cyclodextrins, cyclic peptides, and cucurbitacins. The adapter is preferably a cyclodextrin or a derivative thereof. The cyclodextrin or its derivative may be any one disclosed in Elisev, AV, and Schneider, HJ. (1994) J. Am. Chem. Soc. 116, 6081-6088. The adapter is more preferably hepta-6-amino-β-cyclodextrin.
[0734] (am7-βCD), 6-monodeoxy-6-monoamino-β-cyclodextrin (am1-βCD) or hepta-(6-deoxy-6-guanidinyl)-cyclodextrin (gu7-βCD). The guanidinyl group in gu7-βCD has a higher pKa value than the primary amine in am7-βCD, and therefore it is more likely to carry a positive charge. The gu7-βCD adapter can be used to increase the residence time of nucleotides in the well to improve the accuracy of the measured residual current and increase the baseline detection rate at high temperatures or low data acquisition rates.
[0735] If the succinimide 3-(2-pyridyl dithio)propionate crosslinking agent is used as described in more detail below, the adapter is preferably hepta(6-deoxy-6-amino)-6-N-mono(2-pyridyl)dithiopropionyl-β-cyclodextrin (am6amPDP1-βCD).
[0736] A more suitable adapter includes γ-cyclodextrin comprising nine sugar units (and thus having ninefold symmetry). γ-cyclodextrin may contain linker molecules, or may be modified to contain all or more of the modified sugar units used in the β-cyclodextrin examples described above.
[0737] The molecular adapter is preferably covalently linked to the mutant monomer. The adapter can be covalently linked to the pore by any method known in the art. The adapter is typically linked by chemical linkage. If the molecular adapter is linked via a cysteine linker, it is preferable to introduce one or more cysteine residues into the mutant by substitution, for example, into the barrel structure. The mutant monomer can be chemically modified by linking the molecular adapter to one or more cysteine residues in the mutant monomer. The one or more cysteine residues can be naturally occurring, i.e., located at positions 1 and / or 215 in SEQ ID NO:390. Alternatively, the mutant monomer can be chemically modified by linking the molecule to one or more cysteine residues introduced at other positions. The cysteine at position 215 can be removed, for example by substitution, to ensure that the molecular adapter does not link to positions other than the cysteine at position 1 or to a cysteine introduced at another position.
[0738] The reactivity of cysteine residues can be enhanced by modifying neighboring residues. For example, basic groups on the side chains of arginine, histidine, or lysine residues can alter the pKa of the cysteine thiol group to a more reactive S. - The pKa of the group. The reactivity of the cysteine residue can be protected by a thiol protecting group such as dTNB. These protecting groups can react with one or more cysteine residues of the mutant monomer prior to linking the linker.
[0739] The molecule can be directly linked to the mutant monomer. Preferably, the molecule is linked to the mutant monomer via a linker, such as a chemical cross-linking agent or a peptide linker.
[0740] Suitable chemical crosslinking agents are well known in the art. Preferred crosslinking agents include 2,5-dioxopyrrolidone-1-yl 3-(pyridin-2-ylthioalkyl)propionate, ethyl 2,5-dioxopyrrolidone-1-yl 4-(pyridin-2-ylthioalkyl)butyrate, and 2,5-dioxopyrrolidone-1-yl 8-(pyridin-2-yldithioalkyl)octanoate. The most preferred crosslinking agent is succinimide 3-(2-pyridyldithio)propionate (SPDP). Typically, the molecule is covalently linked to the bifunctional crosslinking agent before the molecule / crosslinking agent complex is covalently linked to the mutant monomer, but it is also possible that the bifunctional crosslinking agent is covalently linked to the monomer before the bifunctional crosslinking agent / monomer complex is linked to the molecule.
[0741] The linker is preferably resistant to dithiothreitol (DTT). Suitable linkers include, but are not limited to, iodoacetamide-based and maleimide-based linkers.
[0742] In another embodiment, the monomer can be linked to a polynucleotide-binding protein. This forms a modular sequencing system that can be used in the sequencing methods of the present invention. Polynucleotide-binding proteins are discussed below.
[0743] The polynucleotide-binding protein is preferably covalently linked to the mutant monomer. The protein can be covalently linked to the monomer using any method known in the art. The monomer and protein can be chemically or genetically fused. If the entire construct is expressed by a single polynucleotide sequence, the monomer and protein can be genetically fused. Genetic fusion of the monomer with the polynucleotide-binding protein is discussed in International Application PCT / GB09 / 001679 (published as WO 2010 / 004265).
[0744] If the polynucleotide-binding protein is linked via a cysteine linker, the one or more cysteine residues are preferably introduced into the mutant through substitution. The one or more cysteine residues are preferably introduced into a cyclic region that exhibits low conservation among homologs, indicating tolerance to mutations or insertions. These are therefore suitable for linking polynucleotide-binding proteins. In such embodiments, the naturally occurring cysteine residue at position 251 can be removed. The reactivity of the cysteine residues can be enhanced by the modifications described above.
[0745] The polynucleotide-binding protein can be linked to the mutant monomer directly or via one or more linkers. The molecule can be linked to the mutant monomer using a heterozygous linker as described in International Application No. PCT / GB10 / 000132 (published as WO 2010 / 086602). Alternatively, peptide linkers can be used. Peptide linkers are amino acid sequences. The length, flexibility, and hydrophilicity of the peptide linkers are generally designed so as not to disrupt the function of the monomer and molecule. Preferred flexible peptide linkers have a chain length of 2 to 20, for example, 4, 6, 8, 10, or 16 serine and / or glycine. More preferred flexible linkers include (SG)1, (SG)2, (SG)3, (SG)4, (SG)5, and (SG)8, where S is serine and G is glycine. Preferred rigid linkers have a chain length of 2 to 30, for example, 4, 6, 8, 16, or 24 proline. More preferred rigid linkers include (P) 12 , where P is proline.
[0746] The mutant monomer can be chemically modified using molecular adapters and polynucleotide binding proteins.
[0747] The molecule (used to chemically modify the monomer) can be directly linked to the monomer or linked via a linker as disclosed in international applications PCT / GB09 / 001690 (publication number WO 2010 / 004273), PCT / GB09 / 001679 (publication number WO 2010 / 004265) or PCT / GB10 / 000133 (publication number WO 2010 / 086603).
[0748] Any protein described herein, such as the mutant monomers and pores described herein, can be modified to facilitate their identification or purification, for example by adding histidine residues (his tags), aspartic acid residues (asp tags), streptavidin tags, flag tags, SUMO tags, GST tags, or MBP tags, or by adding signaling sequences that promote their secretion from cells, wherein the polypeptide does not naturally contain such sequences. An alternative approach to introducing genetic tags is to chemically react the tags to native or modified sites on the protein. One such example could be reacting a gel translocation reagent with cysteine residues modified externally to the protein. This method has been proven to be a method for isolating hemolysin heterooligomers (Chem Biol. 1997 Jul; 4(7) 497-505).
[0749] Any protein described herein, such as the mutant monomers and wells described in this invention, can be labeled with a display label. The display label can be any suitable label that allows the protein to be detected. Suitable labels include, but are not limited to, fluorescent molecules, radioactive isotopes, etc. 125 I, 35 S, enzymes, antibodies, antigens, polynucleotides, and ligands such as biotin.
[0750] Any protein described herein, such as the monomers and pores described in this invention, can be prepared synthetically or by recombinant methods. For example, the protein can be synthesized by in vitro translation and transcription (IVTT). The amino acid sequence of the protein can be modified to include non-naturally occurring amino acids or to increase the stability of the protein. When the protein is prepared synthetically, such amino acids can be introduced during the preparation process. The protein may also be altered after synthetic or recombinant preparation.
[0751] Proteins can also be prepared using D-amino acids. For example, the protein may comprise a mixture of L-amino acids and D-amino acids. This is conventional in the field of preparing the aforementioned proteins or peptides.
[0752] The protein may also contain other non-specific modifications, as long as they do not interfere with the protein's function. Many non-specific side-chain modifications are known in the art and can be performed on the side chains of the protein. These modifications include, for example, amino acid reductive alkylation performed by reacting with an aldehyde followed by reduction with NaBH4, amidoylation with methyl iminoacetate, or acylation with acetic anhydride to obtain the amino acid.
[0753] Any protein described herein, including the monomers and pores of this invention, can be prepared using standard methods known in the art. The polynucleotide sequence encoding the protein can be derived and replicated using standard methods in the art. The polynucleotide sequence encoding the protein is expressed in bacterial host cells using standard techniques in the art. The protein can be prepared from a recombinant expression vector by in situ expression of the polypeptide in cells. The expression vector optionally carries an inducible promoter to control the expression of the polypeptide. These methods are described in Sambrook, J. and Russell, D. (2001). Molecular Cloning: A Laboratory Manual, 3rd Edition. Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY.
[0754] Proteins can be mass-produced after purification by any protein liquid chromatography system from the organism that produces them or after recombinant expression. Typical protein liquid chromatography systems include FPLC, AKTA system, Bio-Cad system, Bio-RadBioLogic system, and Gilson HPLC system.
[0755] Construct
[0756] This invention also provides a construct comprising two or more covalently linked CsgG monomers, wherein at least one of the monomers is a mutant monomer as described in this invention. The construct of this invention retains its ability to form wells. This can be determined as described above. One or more constructs of this invention can be used to form wells for characterization, such as multinucleotide sequencing. The construct may comprise at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 monomers. The construct preferably comprises two monomers. The two or more monomers may be the same or different.
[0757] At least one monomer in the construct is a mutant monomer as described in this invention. Two, three, four, five, six, seven, eight, nine, or ten monomers in the construct may be mutant monomers as described in this invention. All monomers in the construct are preferably mutant monomers as described in this invention. The mutant monomers may be the same or different. In a preferred embodiment, the construct includes two mutant monomers as described in this invention.
[0758] The mutant monomers of the present invention in the constructed organism preferably have approximately the same length or the same length. The barrel-shaped structure of the mutant monomers of the present invention in the constructed organism preferably has approximately the same length or the same length. The length can be measured by the number of amino acids and / or unit length.
[0759] The construct may include one or more monomers that are not mutant monomers of the present invention. CsgG mutant monomers that are not mutant monomers of the present invention include monomers comprising SEQ ID NO: 390, 391, 392, 393, 394, 395, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428, or 429, or control variants comprising SEQ ID NO: 390, 391, 392, 393, 394, 395, 414, 415, 416, 417, 418, 419, 414, 421, 422, 423, 424, 425, 426, 427, 428, or 429, wherein none of the aforementioned amino acids / positions have been mutated. At least one monomer in the construct may include a control variant of the sequence shown in SEQ ID NO:390,391,392,393,394,395,414,415,416,417,418,419,420,421,422,423,424,425,426,427,428 or 429, or SEQ ID NO:390,391,392,393,394,395,414,415,416,417,418,419,420,421,422,423,424,425,426,427,428 or 429. The control variants of SEQ ID NO:390,391,392,393,394,395,414,415,416,417,418,419,420,421,422,423,424,425,426,427,428 or 429 are at least 50% homologous in their entire sequence based on amino acid identity. More preferably, the control variant is at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 75%, at least 80%, at least 85%, at least 90%, and more preferably at least 95%, 97%, or 99% homologous throughout its length based on amino acid identity.
[0760] The monomers in the construct are preferably genetically fused. If the entire construct is expressed by a single polynucleotide sequence, the monomers can be genetically fused. The coding sequences of the monomers can be recombinated in any manner to form a single polynucleotide sequence encoding the construct.
[0761] The monomers can be genetically fused in any conformation. The monomers can be fused via their terminal amino acids. For example, the amino terminus of one monomer can be fused to the carboxyl terminus of another monomer. The second and subsequent monomers in the construct (from amino to carboxyl) can contain a methionine at their amino terminus (each fused to the carboxyl terminus of a previous monomer). For example, if M is a monomer (without an amino-terminal methionine) and mM is a monomer with an amino-terminal methionine, the construct can include the sequence M-mM, M-mM-mM, or M-mM-mM-mM. The presence of these methionines typically arises from the expression of the start codon (i.e., ATG) at the 5' end of the polynucleotide encoding the second or subsequent monomer in the polynucleotide encoding the entire construct. The first monomer in the construct (from amino to carboxyl) can also include methionines (e.g., mM-mM, mM-mM-mM, or mM-mM-mM-mM).
[0762] The two or more monomers can be directly genetically fused together. Preferably, the monomers are genetically fused using a linker. The linker can be designed to restrict the movement of the monomers. A preferred linker is an amino acid sequence (i.e., a peptide linker). Any peptide linker described above can be used.
[0763] In another preferred embodiment, the monomers are chemically fused. If two monomers are chemically linked, for example by a chemical crosslinking agent, then the two monomers are chemically fused. Any chemical crosslinking agent described below may be used. The linker may be attached to one or more cysteine residues introduced into the mutant monomers of the present invention. Alternatively, the linker may be attached to the end of one of the monomers in the construct.
[0764] If the construct contains different monomers, cross-linking between the monomers can be prevented by maintaining the concentration of the linkers in a large excess of said monomers. Alternatively, a "lock and key" setup can be used when using two linkers. Only one end of each linker can react together to form a longer linker, and the other end of each linker reacts with a different monomer. Such a linker is described in International Application No. PCT / GB10 / 000132 (Publication No.: WO 2010 / 086602).
[0765] Polynucleotides
[0766] The present invention also provides a polynucleotide sequence encoding the mutant monomer described herein. The mutant monomer can be any of those described above. The polynucleotide sequence preferably comprises a sequence that is at least 50%, 60%, 70%, 80%, 90%, or 95% homologous over the entire sequence to the sequence SEQ ID NO:389, based on nucleotide identity. Over a continuous nucleotide extension of 300 or more, for example 375, 450, 525, or 600 or more, at least 80%, for example at least 85%, 90%, or 95% nucleotide identity (“strict homology”) may be present. Homology can be calculated as described above. The polynucleotide sequence may contain sequences that differ from SEQ ID NO:389 based on the degeneracy of the genetic code.
[0767] The present invention also provides a polynucleotide sequence encoding any of the genetic fusion constructs described herein. The polynucleotide preferably comprises two or more variants of the sequence shown in SEQ ID NO:389. The polynucleotide sequence preferably comprises two or more sequences having at least 50%, 60%, 70%, 80%, 90%, or 95% homology over the entire sequence based on nucleotide identity. On a continuous nucleotide extension of more than 600, for example, 750, 900, 1050, or 1200 nucleotides, at least 80%, for example, at least 85%, 90%, or 95% nucleotide identity (“strict homology”) may be present. Homology can be calculated as described above.
[0768] The polynucleotide sequence can be derived and replicated using standard methods in the art. Chromosomal DNA encoding wild-type CsgG can be extracted from pore-producing organisms such as *E. coli*. The gene encoding the pore subunit can be amplified using PCR including specific primers. The amplified sequence can then be subjected to site-directed mutagenesis. Suitable methods for site-directed mutagenesis are known in the art and include, for example, combinatorial chain reactions. The polynucleotide encoding the constructions of this invention can be prepared using techniques known in the art, such as those described in *Sambrook, J. and Russell, D. (2001). *Molecular Cloning: A Laboratory Manual*, 3rd Edition. Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY*.
[0769] The obtained polynucleotide sequence can then be integrated into a recombinant reproducible vector, such as a cloning vector. The vector can then be used to replicate the polynucleotide in a compatible host cell. Thus, a polynucleotide sequence can be prepared by introducing the polynucleotide into a reproducible vector, introducing the vector into a compatible host cell, and allowing the host cell to grow under conditions that induce vector replication. The vector can be recovered from the host cell. Suitable host cells for polynucleotide cloning are known in the art and are described in more detail below.
[0770] The polynucleotide sequence can be cloned into a suitable expression vector. In the expression vector, the polynucleotide sequence is typically operatively linked to a control sequence that enables the host cell to express the coding sequence. The expression vector can be used to express the pore subunit.
[0771] The term "operably ligated" refers to a juxtaposition where the described components are in a relationship that allows them to function in their intended manner. "Operably ligated" control sequence to coding sequence means ligated in a manner that enables expression of the coding sequence to be compatible with the control sequence. Multiple copies of the same or different polynucleotide sequences can be introduced into the vector.
[0772] The expression vector can then be introduced into a suitable host cell. Therefore, the mutant monomers or constructs of the present invention can be prepared by inserting a polynucleotide sequence into an expression vector, introducing the vector into a compatible bacterial host cell, and growing the host cell under conditions that induce expression of the polynucleotide sequence. The recombinant expressed monomer or construct can self-assemble into a pore in the host cell membrane. Alternatively, the recombinant pores prepared in this manner can be removed from the host cell and inserted into another membrane. When pores containing at least two different monomers or constructs are generated, the different monomers or constructs can be expressed separately in different host cells as described above, removed from the host cells, and assembled into pores in separate membranes, such as rabbit cell membranes or synthetic membranes.
[0773] The vector may be, for example, a plasmid, viral, or phage vector providing an origin of replication, optionally a promoter for expressing the polynucleotide sequence, and optionally a regulator of the promoter. The vector may contain one or more optional marker genes, such as a tetracycline resistance gene. The promoter and other expression regulatory signals may be selected to be compatible with the host cell targeted by the designed expression vector. Typically, T7, trc, lac, ara, or λ are used. L Promoter.
[0774] The host cells typically express the monomer or construct at high levels. Host cells transformed with a polynucleotide sequence are selected to be compatible with the expression vector used to transform the cells. The host cells are typically bacteria, with *E. coli* being preferred. Any cell line containing λDE3 lysogens, such as C41(DE3), BL21(DE3), JM109(DE3), B834(DE3), TUNER, Origami, and Origami B, can express vectors containing the T7 promoter. In addition to the above conditions, any method cited in Cao et al., 2014, PNAS, Structure of the nonameric bacterial amyloid secretion channel, doi-1411942111 and Goyal et al., 2014, Nature, 516, 250-253, structural and mechanical insights into the bacterial amyloid secretion channel CsgG, can also be used to express the CsgG protein.
[0775] The present invention also includes a method for generating mutant monomers or constructs of the present invention. The method includes expressing the polynucleotides of the present invention in suitable host cells. The polynucleotides are preferably part of a vector and are preferably operably linked to a promoter.
[0776] hole
[0777] This invention also provides a variety of wells. The wells described herein are ideal for characterizing polynucleotide sequences, such as sequencing polynucleotide sequences, because they can distinguish different nucleotides with high sensitivity. The wells can unexpectedly distinguish the four nucleotides in DNA and RNA. The wells described herein can even distinguish between methylated and unmethylated nucleotides. The base resolution of the wells described herein is unexpectedly high. The wells exhibit near-complete separation of all four DNA nucleotides. The wells further distinguish deoxycytosine monophosphate (dCMP) and methyl-dCMP based on residence time in the well and current flowing through the well.
[0778] The pores described in this invention can also distinguish different nucleotides under a range of conditions. Specifically, the pores distinguish nucleotides under conditions favorable for nucleic acid characterization such as sequencing. The degree to which the pores of this invention distinguish different nucleotides can be controlled by varying the applied potential, salt concentration, buffer solution, temperature, and the presence of additives such as urea, betaine, and DTT. This allows for fine-tuning of the pore's function, particularly during sequencing. This will be discussed in more detail below. The pores described in this invention can also be used to identify polynucleotide polymers by interaction with one or more monomers rather than on nucleotide-based nucleotides.
[0779] The pores described in this invention can be isolated, substantially isolated, purified, or substantially purified. The pores described in this invention are isolated or purified if they contain absolutely no other components, such as liposomes or other pores. The pores are substantially isolated if they are mixed with a carrier or diluent that will not interfere with their intended use. For example, the pores are substantially isolated or substantially purified if they contain less than 10%, less than 5%, less than 2%, or less than 1% of other components such as triblock copolymers, liposomes, or other pores. Alternatively, the pores described in this invention can be present in a membrane. Suitable membranes are described below.
[0780] The pores of the present invention can exist as a single or individual pore. Alternatively, the pores of the present invention can exist in a homologous or heterologous group of two or more pores.
[0781] Homologous oligomerization pores
[0782] This invention also provides a homologous oligomer well derived from CsgG containing the same mutant monomer of this invention. The homologous oligomer well may include any mutant of this invention. The homologous oligomer well described in this invention is ideally suited for polynucleotide characterization, such as sequencing. The homologous oligomer well described in this invention may possess any of the advantages described above.
[0783] The homologous oligomer pores may contain any number of mutant monomers. The pores generally contain at least 7, at least 8, at least 9, or at least 10 identical mutant monomers, for example, 7, 8, 9, or 10 mutant monomers. The pores preferably contain 8 or 9 identical mutant monomers. One or more of the mutant monomers, for example 2, 3, 4, 5, 6, 7, 8, 9, or 10, are preferably chemically modified as discussed above.
[0784] The method for preparing the pores is discussed in more detail below.
[0785] Heterogeneous oligomerization pores
[0786] This invention also provides heterooligomeric pores derived from CsgG containing at least one mutant monomer of this invention. The heterooligomeric pores described herein are ideal for polynucleotide characterization, such as sequencing. The heterooligomeric pores can be prepared using methods known in the art (e.g., Protein Sci. 2002 Jul; 11(7) 1813-24).
[0787] The heterooligomeric pores contain sufficient monomers to form the pores. The monomers can be of any type. The pores generally contain at least 7, at least 8, at least 9, or at least 10 monomers, for example, 7, 8, 9, or 10 monomers. The pores preferably contain 8 or 9 monomers.
[0788] In a preferred embodiment, all of the monomers (e.g., 10, 9, 8, or 7 monomers) are mutant monomers of the present invention, and at least one of the monomers is different from the other monomers. In a more preferred embodiment, the pores comprise eight or nine mutant monomers of the present invention, and at least one of the monomers is different from the other monomers. They may also all be different from each other.
[0789] The mutant monomers of the present invention preferably have approximately the same length or the same length in the wells. The barrel-shaped structure of the mutant monomers of the present invention in the wells preferably has approximately the same length or the same length. The length can be measured using the number of amino acids and / or length units.
[0790] In another preferred embodiment, at least one of the mutant monomers is not a mutant monomer of the present invention. In this embodiment, the remaining monomers are preferably mutant monomers of the present invention. Therefore, the wells may include 9, 8, 7, 6, 5, 4, 3, 2, or 1 mutant monomer of the present invention. Any number of monomers in the wells may not be mutant monomers of the present invention. The wells preferably include seven or eight mutant monomers of the present invention and one monomer that is not a monomer of the present invention. The mutant monomers of the present invention may be the same or different.
[0791] The mutant monomers of the present invention in the constructed organism preferably have approximately the same length or have the same length. The barrel-shaped structure of the mutant monomers of the present invention in the constructed organism preferably has approximately the same length or has the same length. The length can be measured using the number of amino acids and / or length units.
[0792] The pores may include one or more monomers that are not mutant monomers of the present invention. CsgG monomers that are not mutant monomers of the present invention include monomers comprising SEQ ID NO: 390, 391, 392, 393, 394, 395, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428, or 429, or control variants comprising SEQ ID NO: 390, 391, 392, 393, 394, 395, 414, 415, 416, 417, 418, 419, 414, 421, 422, 423, 424, 425, 426, 427, 428, or 429, wherein all amino acids / positions related to the present invention as described above have not been mutated / substituted. The control variants of SEQ ID NO:390,391,392,393,394,395,414,415,416,417,418,419,420,421,422,423,424,425,426,427,428 or 429 are generally at least 50% homologous to SEQ ID NO:390,391,392,393,394,395,414,415,416,417,418,419,420,421,422,423,424,425,426,427,428 or 429 based on amino acid identity throughout their entire sequence. More preferably, the control variant can be based on amino acid identity and be at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, and more preferably at least 95%, 97%, or 99% homologous throughout the entire sequence to the amino acid sequence of SEQ ID NO: 390, 391, 392, 393, 394, 395, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428, or 429.
[0793] In all the embodiments described above, one or more, such as 2, 3, 4, 5, 6, 7, 8, 9 or 10, the mutant monomers are preferably chemically modified as described above.
[0794] The method for preparing the pores is discussed in more detail below.
[0795] Holes containing the building blocks
[0796] The present invention also provides a pore comprising at least one construct of the present invention. The construct of the present invention comprises two or more covalently linked monomers derived from CsgG, wherein at least one of the monomers is a mutant monomer of the present invention. That is, the construct must contain more than one monomer. The pore contains sufficient constructs and, if necessary, sufficient monomers to form the pore. For example, an octamer pore may comprise (a) four constructs, each containing two monomers, (b) two constructs, each containing four monomers, or (b) a construct containing two monomers and six monomers not forming part of the construct. For example, a nonamer pore may comprise (a) four constructs each containing two monomers and one monomer not forming part of the construct, (b) two constructs each containing four monomers and one monomer not forming part of the construct, or (b) a construct containing two monomers and seven monomers not forming part of the construct. Other combinations of constructs and monomers can be conceived by those skilled in the art.
[0797] At least two of the monomers in the wells are in the form of the construct of the present invention. The construct, and therefore the wells, include at least one mutant monomer of the present invention. The wells generally contain at least 7, at least 8, at least 9, or at least 10 monomers, for example, 7, 8, 9, or 10 monomers (at least two of which must be in the construct). The wells preferably contain eight or nine monomers (at least two of which must be in the construct).
[0798] The pores containing the constructs can be homologous oligomers (i.e., containing the same constructs) or heterologous oligomers (i.e., at least one of the constructs is different from the others).
[0799] The well typically contains (a) a construct comprising two monomers and (b) 5, 6, 7, or 8 monomers. The construct can be any of those described above. The monomers can be any of those described above, including the mutant monomers of the present invention, containing SEQ ID NO:
[0800] Monomers of SEQ ID NO: 390, 391, 392, 393, 394, 395, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428, or 429, and mutant monomers containing control variants of SEQ ID NO: 390, 391, 392, 393, 394, 395, 414, 415, 416, 417, 418, 419, 414, 421, 422, 423, 424, 425, 426, 427, 428, or 429 as described above.
[0801] Another typical aperture includes more than one of the inventive constructs, such as two, three, or four inventive constructs. If necessary, the aperture further includes sufficient additional monomers or constructs to form the aperture. The additional monomer can be any of the above-described types, including the mutant monomers of the present invention, monomers comprising SEQ ID NO:390,391,392,393,394,395,414,415,416,417,418,419,420,421,422,423,424,425,426,427,428 or 429, and mutant monomers comprising control variants comprising SEQ ID NO:390,391,392,393,394,395,414,415,416,417,418,419,414,421,422,423,424,425,426,427,428 or 429 as described above. The additional construct may be any of the types described above or may be a construct comprising two or more covalently linked CsgG monomers, each of which comprises a monomer containing SEQ ID NO: 390, 391, 392, 393, 394, 395, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428 or 429, and a monomer containing a control variant of SEQ ID NO: 390, 391, 392, 393, 394, 395, 414, 415, 416, 417, 418, 419, 414, 421, 422, 423, 424, 425, 426, 427, 428 or 429 as described above.
[0802] Further, the pores of this invention include only constructs containing two monomers; for example, the pores may include 4, 5, 6, 7, or 8 constructs containing two monomers. At least one construct is a construct of this invention, that is, at least one monomer in at least one construct, and preferably each monomer in at least one construct is a mutant monomer of this invention. All constructs containing two monomers can be constructs of this invention.
[0803] A specific pore of the present invention comprises four constructs of the present invention, each construct containing two monomers, wherein at least one monomer in each construct, and preferably each monomer in each construct, is a mutant monomer of the present invention. The constructs can be oligomerized into pores having a structure in which only one monomer in each construct contributes to the pore channel. Typically, the other monomers in the constructs are located outside the channel of the pore. For example, the pore of the present invention may comprise 7, 8, 9, or 10 constructs containing two monomers, wherein the channel comprises 7, 8, 9, or 10 monomers.
[0804] Mutations can be introduced into the constructs described above. The mutations can be alternating, meaning the mutation is different for each monomer in the two monomer constructs, and the constructs are assembled into homologous oligomers resulting in alternating modifications. That is, monomers containing MutA and MutB are fused and assembled to form AB:AB:AB:AB pores. Alternatively, the mutations can be adjacent, meaning the same mutation is introduced into two monomers in a construct and then oligomerized with different mutated monomers or constructs. That is, a monomer containing MutA is fused and then oligomerized with a monomer containing MutB to form AA:B:B:B:B:B:B.
[0805] One or more monomers of the present invention can be chemically modified in the pores containing the constructs as described above.
[0806] Characterization of analytes
[0807] This invention provides a method for determining the presence, absence, or one or more properties of a target analyte. The method involves contacting the target analyte with a CsgG well or a mutant thereof (e.g., the well of this invention), such that the target analyte moves relative to, for example through, the well, and acquiring one or more measurements as the analyte moves relative to the well, thereby determining the presence, absence, or one or more properties of the analyte. The target analyte may also be referred to as a template analyte or an analyte of interest.
[0808] The method involves contacting the target analyte with a CsgG well or a mutant thereof (e.g., the well of the present invention), causing the target analyte to move through the well. The well typically contains at least 7, at least 8, at least 9, or at least 10 monomers, for example, 7, 8, 9, or 10 monomers. The well preferably contains 8 or 9 identical monomers. One or more of the monomers, for example 2, 3, 4, 5, 6, 7, 8, 9, or 10, are preferably chemically modified as discussed above.
[0809] The CsgG pores can be derived from any organism. The CsgG pores can contain monomers including sequences shown in SEQ ID NO: 390, 391, 392, 393, 394, 395, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428, or 429. The CsgG pores may contain at least 7, at least 8, at least 9, or at least 10 monomers, for example, 7, 8, 9, or 10 monomers, each monomer comprising the sequence shown in SEQ ID NO: 390, 391, 392, 393, 394, 395, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428, or 429 (i.e., the pores are homologous oligomers containing the same monomers from the same organism). The CsgG well may contain any combination of monomers, each of which includes the sequence shown in SEQ ID NO:390,391,392,393,394,395,414,415,416,417,418,419,420,421,422,423,424,425,426,427,428, or 429. For example, the well may contain seven monomers including the sequence shown in SEQ ID NO:390 and two monomers including the sequence shown in SEQ ID NO:391.
[0810] The CsgG mutant may contain any number of mutant monomers, for example, at least 7, at least 8, at least 9, or at least 10 monomers, such as 7, 8, 9, or 10 monomers. The mutant monomers may include comparative variants of the sequences shown in SEQ ID NO: 390, 391, 392, 393, 394, 395, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428, or 429. Comparative variants have been discussed above. Comparative variants must be able to form wells and may have any percentage of the homology discussed above compared to the wells described in this invention.
[0811] The CsgG mutant preferably comprises nine monomers, and at least one of the monomers is a variant of the sequence shown in SEQ ID NO:390, the variant comprising: (a) mutations at one or more of the following positions: N40, Q42, D43, E44, K49, Y51, S54, N55, F56, S57, Q62, E101, N102, E124, E131, R142, D149, T150, E185, R192, D195, E201, and E203, for example, mutations at one or more of the following positions: N40, Q42, D43, E44, K49, Y51, S54, N55, F56, S57, Q62 E101, N102, E131, D149, T150, E185, D195, E201, and E203, or one or more mutations at one or more of the following positions: N40, Q42, D43, E44, K49, Y51, N55, F56, E101, N102, E131, D149, T150, E185, D195, E201, and E203; and / or (b) deletions at one or more of the following positions: F48, K49, P50, Y51, P52, A53, S54, N55, F56, and S57. The variants may contain (a); (b); or (a) and (b). In (a), any number and combination of the following positions: N40, Q42, D43, E44, K49, Y51, S54, N55, F56, S57, Q62, E101, N102, R124, E131, R142, D149, E185, R192, D195, E201, and E203 can be mutated. As discussed above, mutations in one or more of Y51, N55, and F56 can reduce the number of nucleotides contributing to the current when the polynucleotide moves through the pore, thereby making it easier to identify the direct relationship between the observed current and the polynucleotide when the polynucleotide moves through the pore.
[0812] In (b), any number and combination of the following positions: F48, K49, P50, Y51, P52, A53, S54, N55, F56, and S57 may be omitted. The variant may contain any specific mutation / substitution or combination thereof discussed above with respect to the mutant monomers of this invention.
[0813] Preferred variants of the method described in this invention include one or more of the following substitutions: (a) F56N, F56Q, F56R, F56S, F56G, F56A or F56K or F56A, F56P, F56R, F56H, F56S, F56Q, F56I, F56L, F56T or F56G; (b) N55Q, N55R, N55K, N55S, N55G, N55A or N55T; (c) Y51L, Y51V, Y51A, Y51N, Y51Q, Y51S or Y51G; (d) T150I; (e) S54P; and (f) S57P. The variants may include any number and combination of (a) to (f).
[0814] Preferred variants of the method described in this invention include Q62R or Q62K.
[0815] Preferred variants of the method described in this invention include mutations in D43, E44, Q62 or any combination thereof, such as D43, E44, Q62, D43 / E44, D43 / Q62, E44 / Q62 or D43 / E44 / Q62.
[0816] The variant may contain mutations at the following locations:
[0817] Y51, F56, D149, E185, E201 and E203, such as Y51 N, F56A, D149N, E185R, E201 N and E203N;
[0818] N55, such as N55A or N55S;
[0819] Y51, for example, Y51N or Y51T;
[0820] S54, for example, S54P;
[0821] S57, for example, S57P;
[0822] F56, such as F56N, F56Q, F56R, F56S, F56G, F56A or F56K or F56A, F56P, F56R, F56H, F56S, F56Q, F56I, F56L, F56T or F56G.
[0823] Y51 and F56, such as Y51A and F56A, Y51A and F56N, Y51I and F56A, Y51L and F56A, Y51T and F56A, Y51T and F56Q, Y51I and F56N, Y51L and F56N or Y51T and F56N, preferably Y51I and F56A, Y51L and F56A or Y51T and F56A, more preferably Y51T and F56Q or even more preferably Y51X and F56Q, wherein X is any amino acid;
[0824] N55 and F56, such as N55X and F56Q, where X is any amino acid;
[0825] Y51, N55 and F56, for example Y51 A, N55S and F56A, Y51 A, N55S and F56N or Y51 T, N55S and F56Q;
[0826] S54 and F56, such as S54P and F56A or S54P and F56N;
[0827] F56 and S57, such as F56A and S57P or F56N and S57P;
[0828] D149, E185 and E203, such as D149N, E185N and E203N;
[0829] D149, E185, E201 and E203, such as D149N, E185N, E201N and E203N;
[0830] D149, E185, D195, E201 and E203, such as D149N, E185R, D195N, E201 N and E203N or D149N, E185R, D195N, E201 R and E203N;
[0831] F56 and N102, such as F56Q and N102R;
[0832] (a) Q62, such as Q62R or Q62K, and (b) one or more of Y51, N55 and F56, such as in Y51, N55, F56, Y51N55, Y51F56, N55F56 or Y51N55F56, such as Y51TF56QQ62R;
[0833] (i) D43, E44, Q62 or any combination thereof, such as D43, E44, Q62, D43 / E44, D43 / Q62, E44 / Q62 or D43 / E44 / Q62 and (ii) one or more of Y51, N55 and F56, such as Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56 Or Y51 / N55 / F56, for example D43N / Y51T / F56Q, E44N / Y51T / F56Q, D43N / E44N / Y51T / F56Q, D43N / Y51T / F56Q / Q62R, E44N / Y51T / F56Q / Q62R or D43N / E44N / Y51T / F56Q / Q62R; or
[0834] T150, for example, T150I.
[0835] The preferred aperture for use in the method of the present invention comprises at least 7, at least 8, at least 9, or at least 10 monomers, for example, 7, 8, 9, or 10 monomers, each of said monomers comprising a variant of SEQ ID NO:390, said variant comprising one or more of the following substitutions: (a) F56N, F56Q, F56R, F56S, F56G, F56A or F56K or F56A, F56P, F56R, F56H, F56S, F56Q, F56I, F56L, F56T or F56G; (b) N55Q, N55R, N55K, N55S, N55G, N55A or N55T; (c) Y51L, Y51V, Y51A, Y51N, Y51Q, Y51S or Y51G; (d) T150I; (e) S54P; and (f) S57P. The variants may include any number and combination of (a) to (f). The monomers are preferably identical in these preferred pores.
[0836] A preferred aperture for use in the method of the present invention comprises at least 7, at least 8, at least 9, or at least 10 monomers, for example, 7, 8, 9, or 10 monomers, each of said monomers comprising a variant of SEQ ID NO:390, said variant comprising:
[0837] Y51N, F56A, D149N, E185, E201N and E203N;
[0838] N55A;
[0839] N55S;
[0840] Y51N;
[0841] S54P;
[0842] S57P;
[0843] F56N, F56Q, F56R, F56S, F56G, F56A or F56K;
[0844] F56A, F56P, F56R, F56H, F56S, F56Q, F56I, F56L, F56T or F56G;
[0845] Y51A and F56A;
[0846] Y51A and F56N;
[0847] Y51I and F56A;
[0848] Y51L and F56A;
[0849] Y51T and F56A;
[0850] Y51T and F56Q;
[0851] Y51I and F56N;
[0852] Y51L and F56N;
[0853] Y51T and F56N;
[0854] N55S and F56Q;
[0855] Y51A, N55S and F56A;
[0856] Y51A, N55S and F56N;
[0857] Y51T, N55S and F56Q;
[0858] S54P and F56A;
[0859] S54P and F56N;
[0860] F56A and S57P;
[0861] F56N and S57P;
[0862] D149N, E185N and E203N;
[0863] D149N, E185N, E201N and E203N;
[0864] D149N, E185R, D195N, E201N and E203N;
[0865] D149N, E185R, D195N, E201R and E203N;
[0866] T150I;
[0867] F56Q and N102R;
[0868] F56 and N102, such as F56Q and N102R;
[0869] Y51T / F56Q / Q62R;
[0870] D43N / Y51T / F56Q;
[0871] E44N / Y51T / F56Q;
[0872] D43N / E44N / Y51T / F56Q;
[0873] D43N / Y51T / F56Q / Q62R;
[0874] E44N / Y51T / F56Q / Q62R; or
[0875] D43N / E44N / Y51T / F56Q / Q62R.
[0876] The CsgG mutant used in the method of the present invention preferably comprises nine monomers, and at least one of the monomers is a variant of the sequence shown in SEQ ID NO:390, the variant including mutations at one or more of positions Y51, N55, and F56. Preferred wells for the method of the present invention comprise at least 7, at least 8, at least 9, or at least 10 monomers, for example, 7, 8, 9, or 10 monomers, each of the monomers comprising a variant of SEQ ID NO:390, the variant including mutations at one or more of positions Y51, N55, and F56. The monomers are preferably identical in these preferred wells. The variant may include mutations at positions Y51; N55; F56; Y51 / N55; Y51 / F56; N55 / F56; or Y51 / N55 / F56. The variant may include any specific mutations and any combination thereof at one or more of the positions Y51, N55, and F56 discussed above. One or more of Y51, N55, and F56 may be substituted with any amino acid. Y51 can be replaced by F, M, L, I, V, A, P, G, C, Q, N, T, S, E, D, K, H, or R, for example, A, S, T, N, or Q. N55 can be replaced by F, M, L, I, V, A, P, G, C, Q, T, S, E, D, K, H, or R, for example, A, S, T, or Q. F56 can be replaced by M, L, I, V, A, P, G, C, Q, N, T, S, E, D, K, H, or R, for example, A, S, T, N, or Q. The variant may further include one or more of the following: (i) one or more mutations at the following positions (i) N40, D43, E44, S54, S57, Q62, R97, E101, E124, E131, R142, T150 and R192; (iii) Q42Ror Q42K; (iv) K49R; (v) N102R, N102F, N102Y, or N102W; (vi) D149N, D149Q, or D149R; (vii) E185N, E185Q, or E185R; (viii) D195N, D195Q, or D195R; (ix) E201N, E201Q, or E201R; (x) E203N, E203Q, or E203R; and (ix) missing one or more of the following positions: F48, K49, P50, Y51, P52, A53, S54, N55, F56, and S57. The variants may include any combination of (i) and (iii) through (xi) discussed above. The variants may include any implementation discussed above for (i) and (iii) through (xi).
[0877] (1) A preferred variant of the method of the present invention or (2) a preferred aperture of the method of the present invention, comprising at least 7, at least 8, at least 9 or at least 10 monomers, for example 7, 8, 9 or 10 monomers, each of the monomers comprising a variant of SEQ ID NO:390, the variant comprising: Y51R / F56Q, Y51N / F56N, Y51M / F56Q, Y51L / F56Q, Y51I / F56Q, Y51V / F56Q, Y51A / F56Q, Y51P / F56Q, Y51G / F56Q, Y51C / F56Q, Y51Q / F56Q, Y51N / F56Q, Y51S / F56Q, Y51E / F56Q, Y51D / F56Q, Y51K / F56Q or Y51H / F56Q.
[0878] (1) A preferred mutant for use in the method of the present invention or (2) a preferred well for use in the method of the present invention, comprising at least 7, at least 8, at least 9 or at least 10 monomers, for example 7, 8, 9 or 10 monomers, each of the monomers comprising a variant of SEQ ID NO:390, the variant comprising: Y51TF56Q, Y51QF56Q or Y51A / F56Q.
[0879] (1) a preferred mutant for use in the method of the present invention or (2) a preferred well for use in the method of the present invention, comprising at least 7, at least 8, at least 9 or at least 10 monomers, for example 7, 8, 9 or 10 monomers, each of said monomers comprising SEQ Variants of IDNO:390, including: Y51T / F56F, Y51T / F56M, Y51T / F56L, Y51T / F56I, Y51T / F56V, Y51T / F56A, Y51T / F56P, Y51T / F56G, Y51T / F56C, Y51T / F56Q, Y51T / F56N, Y51T / F56T, Y51T / F56S, Y51T / F56E, Y51T / F56D, Y51T / F56K, Y51T / F56H, or Y51T / F56R.
[0880] (1) A preferred variant of the method of the present invention or (2) a preferred aperture of the method of the present invention, comprising at least 7, at least 8, at least 9 or at least 10 monomers, such as 7, 8, 9 or 10 monomers, each of the monomers comprising a variant of SEQ ID NO:390, the variant comprising: Y51T / N55Q, Y51T / N55S or Y51T / N55A.
[0881] (1) A preferred variant of the method of the present invention comprises or (2) a preferred aperture for the method of the present invention comprises at least 7, at least 8, at least 9 or at least 10 monomers, for example 7, 8, 9 or 10 monomers, each of the monomers comprising a variant of SEQ ID NO: 390, the variant comprising: Y51A / F56F, Y51A / F56L, Y51A / F56I, Y51A / F56V, Y51A / F56A, Y51A / F56P, Y51A / F56G, Y51A / F56C, Y51A / F56Q, Y51A / F56N, Y51A / F56T, Y51A / F56S, Y51A / F56E, Y51A / F56D, Y51A / F56K, Y51A / F56H or Y51A / F56R.
[0882] (1) A preferred variant of the method of the present invention or (2) a preferred aperture for the method of the present invention, comprising at least 7, at least 8, at least 9 or at least 10 monomers, for example 7, 8, 9 or 10 monomers, each of the monomers comprising a variant of SEQ ID NO:390, the variant comprising: Y51C / F56A, Y51E / F56A, Y51D / F56A, Y51K / F56A, Y51H / F56A, Y51Q / F56A, Y51N / F56A, Y51S / F56A, Y51P / F56A or Y51V / F56A.
[0883] (1) a preferred variant of the method of the present invention or (2) a preferred aperture for the method of the present invention, comprising at least 7, at least 8, at least 9 or at least 10 monomers, for example 7, 8, 9 or 10 monomers, each of said monomers comprising a variant of SEQ ID NO:390, said variant comprising:
[0884] D149R / E185R / E201R / E203R or Y51T / F56Q / D149R / E185R / E201R / E203R;
[0885] D149N / E185N / E201N / E203N or Y51T / F56Q / D149N / E185N / E201N / E203N;
[0886] E201R / E203R or Y51T / F56Q / E201R / E203R
[0887] E201N / E203R or Y51T / F56Q / E201N / E203R;
[0888] E203R or Y51T / F56Q / E203R;
[0889] E203N or Y51T / F56Q / E203N;
[0890] E201 R or Y51T / F56Q / E201R;
[0891] E201 N or Y51T / F56Q / E201 N;
[0892] E 185R or Y51 T / F56Q / E 185R;
[0893] E185N or Y51T / F56Q / E185N;
[0894] D149R or Y51T / F56Q / D149R;
[0895] D149N or Y51TF56Q / D149N;
[0896] R142E or Y51T / F56Q / R142E;
[0897] R142N or Y51T / F56Q / R142N;
[0898] R192E or Y51T / F56Q / R192E; or
[0899] R192N or Y51T / F56Q / R192N;
[0900] (1) a preferred variant of the method of the present invention or (2) a preferred aperture for the method of the present invention, comprising at least 7, at least 8, at least 9 or at least 10 monomers, for example 7, 8, 9 or 10 monomers, each of said monomers comprising a variant of SEQ ID NO:390, said variant comprising:
[0901] Y51A / F56Q / E101 N / N102R;
[0902] Y51 A / F56Q / R97N / N 102G;
[0903] Y51A / F56Q / R97N / N102R;
[0904] Y51A / F56Q / R97N;
[0905] Y51A / F56Q / R97G;
[0906] Y51A / F56Q / R97L;
[0907] Y51A / F56Q / N102R;
[0908] Y51A / F56Q / N102F;
[0909] Y51A / F56Q / N102G;
[0910] Y51A / F56Q / E101R;
[0911] Y51A / F56Q / E101F;
[0912] Y51A / F56Q / E101N; or
[0913] Y51A / F56Q / E101G.
[0914] The monomers are preferably the same in these preferred pores.
[0915] The preferred apertures used in the method of the present invention comprise at least 7, at least 8, at least 9, or at least 10 monomers, for example, 7, 8, 9, or 10 monomers, each of said monomers comprising a variant of SEQ ID NO:390, said variant comprising F56A, F56P, F56R, F56H, F56S, F56Q, F56I, F56L, F56T, or F56G. The monomers are preferably identical in these preferred apertures.
[0916] The CsgG mutant may include any variant of SEQ ID NO:390 disclosed in the examples or may be any well disclosed in the examples.
[0917] The CsgG mutant is most preferably found in the well of the present invention.
[0918] Steps (a) and (b) are preferably performed with a potential applied across the pore. As discussed in more detail below, the applied potential generally results in the formation of a complex between the pore and the polynucleotide-binding protein. The applied potential can be a voltage potential. Alternatively, the applied potential can be a chemielectric potential. One example is the use of a salt gradient across the amphiphilic molecular layer. Salt gradients are disclosed in Holden et al., J Am Chem Soc. 2007 Jul 11; 129(27) 8650-5.
[0919] The method is used to determine the presence, absence, or one or more characteristics of a target analyte. The method can be used to determine the presence, absence, or one or more characteristics of at least one target analyte. The method can involve determining the presence, absence, or one or more characteristics of two or more target analytes. The method can include determining the presence, absence, or one or more characteristics of any number of target analytes, such as 2, 5, 10, 15, 20, 30, 40, 50, 100, or more target analytes. Any number of characteristics of one or more target analytes can be determined, such as 1, 2, 3, 4, 5, 10, or more characteristics.
[0920] The target analytes are preferably metal ions, inorganic salts, polymers, amino acids, peptides, polypeptides, proteins, nucleotides, oligonucleotides, polynucleotides, dyes, bleaching agents, pharmaceuticals, diagnostic agents, recreational drugs, explosives, or environmental pollutants. The method may involve determining the presence, absence, or one or more characteristics of two or more target analytes of the same category, such as two or more proteins, two or more nucleotides, or two or more pharmaceuticals. Alternatively, the method may involve determining the presence, absence, or one or more characteristics of two or more target analytes of different categories, such as one or more proteins, one or more nucleotides, and one or more pharmaceuticals.
[0921] The target analyte can be secreted by cells. Alternatively, the target analyte can be an intracellular analyte such that it must be extracted from the cells prior to the present invention.
[0922] The analyte is preferably an amino acid, peptide, polypeptide, and / or protein. The amino acid, peptide, polypeptide, or protein may be naturally occurring or non-natural. The polypeptide or protein may contain synthetic or modified amino acids. A wide variety of modifications to amino acids are known in the art. Suitable amino acids and their modifications are described above. For the purposes of this invention, it is understood that the target analyte can be modified by any method known in the art.
[0923] The protein may be an enzyme, antibody, hormone, growth factor, or growth regulatory protein, such as a cytokine. The cytokine may be selected from interleukins, preferably IFN-1, IL-1, IL-2, IL-4, IL-5, IL-6, IL-10, IL-12, and IL-13; interferon, preferably IL-γ; and other cytokines, such as TNF-α. The protein may be a bacterial protein, fungal protein, viral protein, or parasitic protein.
[0924] The target analyte is preferably a nucleotide, oligonucleotide, or polynucleotide. Nucleotides and polynucleotides are discussed below. Oligonucleotides are short-chain nucleotide polymers that generally have 50 or fewer nucleotides, such as 40 or fewer, 30 or fewer, 20 or fewer, 10 or fewer, or 5 or fewer nucleotides. The oligonucleotide may include any nucleotide discussed below, including baseless and modified nucleotides.
[0925] The target analyte, such as the target polynucleotide, can be present in any suitable sample discussed below.
[0926] As described below, the pores are generally present in the membrane. Target analytes can be coupled to or transferred to the membrane using methods discussed below.
[0927] Any measurement method discussed below can be used to determine the presence, absence, or one or more characteristics of a target analyte. The method preferably includes contacting the target analyte with the orifice such that the analyte moves relative to (e.g., through) the orifice, and measuring the current passing through the orifice as the analyte moves relative to the orifice, thereby determining the presence, absence, or one or more characteristics of the analyte.
[0928] If the current flows through the well in an analyte-specific manner (i.e., if a specific current associated with the analyte is detected flowing through the well), the target analyte is present. If the current does not flow through the well in a nucleotide-specific manner, the analyte is not present. Control experiments can be performed in the presence of the analyte to determine whether they affect the manner in which the current flows through the well.
[0929] This invention can be used to distinguish analytes with similar structures—based on their different effects on the current passing through the pore. Individual analytes can be identified at the single-molecule level from the magnitude of the current they generate when interacting with the pore. This invention can also be used to determine the presence of a specific analyte in a sample. This invention can also be used to measure the concentration of a specific analyte in a sample. The use of pores other than CsgG for analyte characterization is known in the art.
[0930] Characterization of polynucleotides
[0931] This invention provides a method for characterizing target polynucleotides, such as polynucleotide sequencing. There are two main strategies for characterizing or sequencing polynucleotides using nanopores: strand characterization / sequencing and exonuclease characterization / sequencing. The method described in this invention can relate to either of these approaches.
[0932] In strand sequencing, DNA is displaced through a nanopore with or against an applied potential. Exonucleases that act progressively or processively on double-stranded DNA can be used on the cis side of the pore to feed the remaining single strand through the pore under the applied potential, or on the trans side of the pore under the reverse potential. Similarly, helicases that untangle double-stranded DNA can be used in a similar manner. Polymerases can also be used. Sequencing applications requiring strand displacement against the applied potential may also be used, but the DNA must first be "captured" by the enzyme under the reverse potential or no potential. Then, as the potential switches back after binding, the strand will pass through the pore from cis to trans and remain in an extended conformation by current. Single-stranded DNA exonucleases or single-stranded DNA-dependent polymerases can act as molecular motors, pulling the just-displaced single strand back through the pore in a controlled, stepwise manner against the applied potential, from trans to cis.
[0933] In one embodiment, the method for characterizing the target polynucleotide includes contacting the target sequence with a pore and a helicase. Any helicase can be used in this method. Suitable helicases are discussed below. The helicase can operate relative to the pore in two modes. First, the method is preferably performed using a helicase that controls the movement of the target sequence through the pore by a field generated by an applied voltage. In this mode, the 5' end of the DNA is first captured into the pore, and the enzyme controls the movement of the DNA into the pore, such that the target sequence passes through the pore under the electric field until it is finally displaced to the trans side of the bilayer. Alternatively, the method is preferably performed such that the helicase controls the movement of the target sequence through the pore against the field generated by the applied voltage. In this mode, the 3' end of the DNA is first captured into the pore, and the enzyme controls the movement of the DNA through the pore, such that the target sequence is pulled out of the pore against the applied field until it finally springs back to the cis side of the bilayer.
[0934] In exonuclease sequencing, an exonuclease releases individual nucleotides from one end of a target polynucleotide, and these individual nucleotides are identified as described below. In another embodiment, the method for characterizing the target polynucleotide includes contacting the target sequence with a well and an exonuclease. Any exonuclease described below can be used in this method. The enzyme can be covalently linked to the well as described below.
[0935] An exonuclease is an enzyme that typically locks onto one end of a polynucleotide and digests one nucleotide of the sequence at a time from that end. The exonuclease can digest the polynucleotide in a 5' to 3' orientation or a 3' to 5' orientation. The end of the polynucleotide to which the exonuclease is attached is typically determined by selecting the enzyme used and / or employing methods known in the art. A hydroxyl group or cap structure at either end of the polynucleotide can often be used to prevent or facilitate the binding of the exonuclease to a specific end of the polynucleotide.
[0936] The method includes contacting the polynucleotide with the exonuclease such that the nucleotide is digested from the ends of the polynucleotide at a rate that allows characterization or identification of a subset of the nucleotides as described above. Methods for performing these steps are well known in the art. For example, Edman degradation is used to sequentially digest individual amino acids from the ends of a polypeptide, allowing them to be identified using high-performance liquid chromatography (HPLC). Homology methods can be used in this invention.
[0937] The rate of action of the exonuclease is generally lower than the optimal rate of the wild-type exonuclease. Suitable rates of exonuclease activity in the method of this invention include digestion of 0.5 to 1000 nucleotides per second, 0.6 to 500 nucleotides per second, 0.7 to 200 nucleotides per second, 0.8 to 100 nucleotides per second, 0.9 to 50 nucleotides per second, or 1 to 20 or 10 nucleotides per second. The preferred rates are 1, 10, 100, 500, or 1000 nucleotides per second. Suitable rates of exonuclease activity can be achieved in various ways. For example, according to this invention, different exonucleases with reduced optimal activity rates can be used.
[0938] In an embodiment of chain characterization, the method includes contacting the polynucleotide with a CsgG pore or a mutant thereof, such as the pore of the present invention, such that the polynucleotide moves relative to the pore, for example through the pore, and acquiring one or more measurements as the polynucleotide moves relative to the pore, wherein the measurements indicate one or more characteristics of the polynucleotide, thereby characterizing the target polynucleotide.
[0939] In an embodiment of exonuclease characterization, the method includes contacting the polynucleotide with a CsgG pore or a mutant thereof, such as the pore of the present invention, and an exonuclease, such that the exonuclease digests each nucleotide from one end of the target polynucleotide and the each nucleotide moves relative to the pore, for example through the pore, and acquiring one or more measurements as the each nucleotide moves relative to the pore, wherein the measurements indicate one or more characteristics of the each nucleotide, thereby characterizing the target polynucleotide.
[0940] Each nucleotide is a single nucleotide. Each nucleotide is a nucleotide that is not bound to another nucleotide or polynucleotide via a nucleotide bond. A nucleotide bond includes one of the phosphate groups of the nucleotide that is bound to the sugar moiety of another nucleotide. Each nucleotide is typically not bound to another polynucleotide having at least 5, at least 10, at least 20, at least 50, at least 100, at least 200, at least 500, at least 1000, or at least 5000 nucleotides via a nucleotide bond. For example, the individual nucleotides have been digested from a target polynucleotide sequence, such as a DNA or RNA strand. The nucleotides can be any of those described below.
[0941] The individual nucleotides can interact with the pore in any manner at any site. The nucleotides preferably bind reversibly to the pore via or connected to an adapter, as described above. Most preferably, the nucleotides bind reversibly to the pore via or connected to an adapter as they cross the membrane. The nucleotides can also reversibly bind to the barrel structure or channel of the pore via or connected to the adapter as they cross the membrane.
[0942] In the interaction between the various nucleotides and the pore, the nucleotides typically affect the current flowing through the pore in a nucleotide-specific manner. For example, a particular nucleotide will reduce the current flowing through the pore by a specific average time period and a specific degree. That is, the current flowing through the pore is specific to a particular nucleotide. Control experiments can be performed to determine the effect of a specific nucleotide on the current flowing through the pore. The results obtained by applying the method of the present invention to the test sample are then compared with the results from the control experiments described above to identify specific nucleotides in the sample or to determine the presence of a specific nucleotide in the sample. The frequency at which the current flowing through the pore is affected in a manner indicating a specific nucleotide can be used to determine the concentration of that nucleotide in the sample. The proportions of different nucleotides in the sample can also be calculated. For example, the proportion of dCMP to methyl-dCMP can be calculated.
[0943] The method includes measuring one or more properties of the target polynucleotide. The target polynucleotide may also be referred to as a template polynucleotide or a polynucleotide of interest.
[0944] This implementation can also use CsgG wells or mutants thereof, such as the wells of this invention. Any of the wells and implementation methods described above for the target analyte can be used.
[0945] Polynucleotides
[0946] Polynucleotides, such as nucleic acids, are macromolecules containing two or more nucleotides. The polynucleotide or nucleic acid can contain any combination of any nucleotides. The nucleotides can be naturally occurring or artificially synthesized. One or more nucleotides in the polynucleotide can be oxidized or methylated. One or more nucleotides in the polynucleotide can be damaged. For example, the polynucleotide can contain pyrimidine dimers. Such dimers are commonly associated with damage caused by ultraviolet radiation and are a major cause of skin melanoma. One or more nucleotides in the polynucleotide can be modified, for example, by labeling or tagging. Suitable labeling is described below. The polynucleotide can contain one or more spacer regions.
[0947] Nucleotides typically contain a nucleobase, a sugar, and / or at least one phosphate group. The nucleobase and sugar form a nucleoside.
[0948] The nucleobases are typically heterocyclic. Nucleobases include, but are not limited to, purines and pyrimidines, more specifically adenine (A), guanine (G), thymine (T), uracil (U), and cytosine (C).
[0949] The sugar is typically a pentose. Nucleoside sugars include, but are not limited to, ribose and deoxyribose. The sugar is preferably deoxyribose.
[0950] The polynucleotide preferably includes the following nucleosides: deoxyadenosine (dA), deoxyuridine (dU) and / or deoxythymidine (dT), deoxyguanosine (dG) and deoxycytidine (dC).
[0951] The nucleotide is typically a ribonucleotide or a deoxyribonucleotide. The nucleotide typically contains a monophosphate, diphosphate, or triphosphate. The nucleotide may contain more than three phosphates, for example, four or five phosphates. The phosphate may be attached to the 3' or 5' side of the nucleotide. Nucleotides include, but are not limited to, adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), 5-methylcytidine monophosphate, 5-hydroxymethylcytidine monophosphate, cytidine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythymidine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), deoxycytidine monophosphate (dCMP), and deoxymethylcytidine monophosphate. The nucleotides are preferably selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, dCMP, and dUMP.
[0952] Nucleotides can be baseless (i.e., lacking a nucleobase). Nucleotides can also lack both a nucleobase and a sugar (i.e., a C3 spacer).
[0953] The nucleotides in the polynucleotide can be linked to each other in any manner. The nucleotides are typically linked by their sugar and phosphate groups, as in nucleic acids. The nucleotides can also be linked by their nucleobases, as in pyrimidine dimers.
[0954] The polynucleotide can be single-stranded or double-stranded. At least a portion of the polynucleotide is preferably double-stranded.
[0955] The polynucleotide can be a nucleic acid, such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The polynucleotide may contain an RNA strand hybridized to a DNA strand. The polynucleotide can be any synthetic nucleic acid known in the art, such as peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threonine nucleic acid (TNA), locked nucleic acid (LNA), or other synthetic polymers with nucleotide side chains. The PNA backbone consists of repeating N-(2-aminoethyl)-glycine units linked by peptide bonds. The GNA backbone consists of repeating ethylene glycol units linked by phosphodiester bonds. The TNA backbone consists of repeating threoyl groups linked together by phosphodiester bonds. LNA is formed from the above-described ribonucleic acids, with additional bridging structures linking the 2' oxygen and 4' carbon in the ribose moiety. Bridged nucleic acids (BNA) are modified RNA nucleotides. They may also be referred to as restricted or inaccessible RNA. BNA monomers can contain 5-, 6-, or even 7-membered bridging structures with “fixed” C3'-endo sugar puckering. These bridging structures are synthesized to introduce the 2',4'-position of the ribose to produce 2',4'-BNA monomers.
[0956] The polynucleotide is most preferably ribonucleic acid (RNA) or deoxyribonucleic acid (DNA).
[0957] The polynucleotide can be of any length. For example, the length of the polynucleotide can be at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, or at least 500 nucleotides or nucleotide pairs. The length of the polynucleotide can be 1000 or more nucleotides or nucleotide pairs, 5000 or more nucleotides or nucleotide pairs, or 100000 or more nucleotides or nucleotide pairs.
[0958] Any number of polynucleotides can be studied. For example, the method of the present invention can involve characterizing 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 50, 100 or more polynucleotides. If two or more polynucleotides are characterized, they can be either different polynucleotides or the same polynucleotide.
[0959] The polynucleotide can be naturally occurring or artificially synthesized. For example, the method can be used to verify the sequence of the prepared oligonucleotide. The method is typically performed in vitro.
[0960] sample
[0961] The polynucleotide is typically present in any suitable sample. This invention is generally performed on samples known to contain or suspected of containing the polynucleotide. Alternatively, this invention can be performed on a specific sample to confirm the identity of the polynucleotide known to be present or expected to be present in the sample.
[0962] The sample may be a biological sample. This invention can be performed in vitro on samples obtained or extracted from any organism or microorganism. The organism or microorganism is typically archaeal, prokaryotic, or eukaryotic, and generally belongs to one of the following five kingdoms: plant, animal, fungi, prokaryotes, and protists. This invention is performed in vitro on samples obtained or extracted from any virus. The sample is preferably a fluid sample. The sample typically includes bodily fluids from a patient. The sample may be urine, lymph, saliva, mucus, or amniotic fluid, but blood, plasma, or serum are preferred.
[0963] Typically, the sample is derived from a human, but alternatively it may be derived from other mammals, such as commercially raised animals like horses, cattle, sheep, fish, chickens, or pigs, or alternatively from pets such as cats or dogs. Alternatively, the sample may be derived from plants, such as samples obtained from commercial crops like grains, legumes, fruits, or vegetables, such as wheat, barley, oats, rapeseed, corn, soybeans, rice, rhubarb, bananas, apples, tomatoes, potatoes, grapes, tobacco, kidney beans, lentils, sugarcane, cocoa, or cotton.
[0964] The sample may be a non-biological sample. The non-biological sample is preferably a fluid sample. Examples of non-biological samples include surgical fluids, water such as drinking water, seawater, or river water, and reagents used for laboratory testing.
[0965] The samples are typically pretreated in this invention, for example by centrifugation or membrane filtration to remove unwanted molecules or cells, such as red blood cells. The samples can be measured immediately after collection. Samples may also be stored prior to measurement, preferably at temperatures below -70°C.
[0966] Characterization
[0967] The method may involve measuring two, three, four, or five or more features of a polynucleotide. The one or more features are preferably selected from (i) the length of the polynucleotide, (ii) the identity of the polynucleotide, (iii) the sequence of the polynucleotide, (iv) the secondary structure of the polynucleotide, and (v) whether the polynucleotide is modified. According to the invention, any combination of (i) to (v) can be measured, for example, {i},{ii},{iii},{iv},{v},{i,ii},{i,iii},{i,iv},{i,v},{ii,iii},{ii,iv},{ii,v},{iii,iv},{iii,v},{iv,v},{i,ii,iii},{i,ii,iv},{i,ii,v},{i,iii,iv},{i,iii,v},{i,iv,v},{ii,iii,iv},{ii,iii,v},{ii,iii,iv},{ii,iii,v},{ii,iii,iv},{ii,iii,v},{ii,iii,v},{ii,iii,iv},{ii,iii,v},{ii,iii,v},{ii,iii,v},{ii,iii,v},{ii,iii,v},{ii,iii,v},{ii,iii,v},{ii,iii,v},{ii,iii,v},{ii,iii,v},{ii,iii,v},{ii,iii,v},{ii,iii,v},{ii,iii,v},{ii,iii,v},{ii,iii,v},{ii,iii,v},
[0968] {ii,iv,v},{iii,iv,v},{i,ii,iii,iv},{i,ii,iii,v},{i,ii,iv,v},{i,iii,iv,v},{ii,iii,iv,v} or {i,ii,iii,iv,v}. Different combinations of (i) to (v) can measure the first polynucleotide, compared to the second polynucleotide, including any of the combinations listed above.
[0969] For (i), the length of the polynucleotide can be measured, for example, by determining the amount of interaction between the polynucleotide and the pore or the duration of the interaction between the polynucleotide and the pore.
[0970] Regarding (ii), the identity of a polynucleotide can be measured in several ways, either in conjunction with or without measuring the polynucleotide sequence. The former is simpler; the polynucleotide is sequenced and then identified. The latter can be accomplished in several different ways. For example, the presence of a specific motif in the polynucleotide can be measured (without measuring the rest of the polynucleotide's sequence). Alternatively, the measurement of specific electrical and / or optical signals in the method can identify that the polynucleotide originates from a specific source.
[0971] For (iii), the sequence of the polynucleotide can be determined as previously described. Suitable sequencing methods, particularly those using electrical measurement methods, are described in Stoddart D et al., Proc Natl Acad Sci, 12; 106(19)7702-7, Lieberman KR et al., J Am Chem Soc. 2010; 132(50)17961-72, and international application WO2000 / 28312.
[0972] Regarding (iv), the secondary structure can be measured using various methods. For example, if the method involves an electrical measurement, the secondary structure can be measured using variations in residence time or changes in the current flowing through the orifice. This allows for the differentiation of regions of single-stranded and double-stranded polynucleotides.
[0973] For (v), the presence of any modification can be measured. The method preferably includes determining whether the polynucleotide has been modified by methylation, oxidation, damage, with one or more proteins, or with one or more labels, tags, or spacer regions. Specific modifications will result in a specific interaction with the pore, which can be measured using methods described below. For example, methylcytosine can be distinguished from cytosine based on the current flowing through the pore during its interaction with each nucleotide.
[0974] The target polynucleotide contacts a CsgG pore or a mutant thereof, such as the pore of this invention. The pore is typically present in a membrane. Suitable membranes are described below. The method can be performed using any apparatus suitable for studying membrane / pore systems—where pores are present in the membrane. The method can be performed using any apparatus suitable for use on the transmembrane pore-sensing side. For example, the apparatus includes a chamber containing an aqueous solution and a barrier dividing the chamber into two parts. The barrier typically has pores, forming a pore-containing membrane within the pores. Alternatively, the barrier forms a membrane in which pores are present.
[0975] The method can be performed using the apparatus described in International Application No. PCT / GB08 / 000562 (WO 2008 / 102120).
[0976] Various types of measurements can be performed. These include, but are not limited to, electrical and optical measurements. Possible electrical measurements include current measurements, impedance measurements, tunneling measurements (Ivanov AP et al., Nano Lett. 2011 Jan 12; 11(1):279-85), and FET measurements (International Application WO 2005 / 124888). Optical measurements can be combined with electrical measurements (Soni GV et al., Rev Sci Instrum. 2010 Jan; 81(1)014301). The measurements can be transmembrane current measurements, such as measurements of the ion current flowing through the pore.
[0977] Electrical measurements can be performed using standard single-channel recording equipment as described in Stoddart D et al., Proc Natl Acad Sci, 12; 106(19)7702-7, Lieberman KR et al., J Am Chem Soc. 2010; 132(50)17961-72 and International Application WO2000 / 28312. Alternatively, electrical measurements can be performed using multi-channel systems, such as those described in International Applications WO 2009 / 077734 and WO 2011 / 067559.
[0978] The method preferably employs a transmembrane applied potential. The applied potential can be a voltage potential. Alternatively, the applied potential can be a chemical potential. One example is the use of a transmembrane salt gradient, such as that of an amphiphilic molecular layer. Salt gradients are disclosed in Holden et al., J Am Chem Soc. 2007 Jul 11; 129(27): 8650-5. In some cases, the current flowing through the pore as the polynucleotide moves relative to it is used to estimate or determine the sequence of the polynucleotide. This is strand sequencing.
[0979] The method may include measuring the current flowing through the pore as the polynucleotide moves relative to the pore. Therefore, the apparatus used in the method may also include circuitry capable of applying a potential and measuring the electrical signal passing through the membrane and the pore. The method may be performed using patch clamp or voltage clamp. The method preferably involves the use of voltage clamp.
[0980] The method of the present invention may include measuring the current flowing through the pore as a polynucleotide moves relative to the pore. Suitable conditions for measuring the ion current through the transmembrane protein pore are known in the art and disclosed in the embodiments. The method is generally carried out by applying a voltage to the membrane and the pore. The voltage used is typically from +5V to -5V, for example from +4V to -4V, from +3V to -3V, or from +2V to -2V. The voltage used is typically from -600mV to +600V or from -400mV to +400mV. The voltage used is preferably selected from -400mV, -300mV, -200mV, -150mV, -100mV, -50mV, ...
[0981] The lower limits of -20mV and 0mV are independently selected from the ranges of +10mV, +20mV, +50mV, +100mV, +150mV, +200mV, +300mV, and +400mV. The voltage used is more preferably in the range of 100mV to 240mV, and most preferably in the range of 120mV to 220mV. By using an increased applied potential, the pore's recognition of different nucleotides can be increased.
[0982] The method is typically carried out in the presence of any charge carrier, such as metal salts (e.g., alkali metal salts), halide salts (e.g., chloride salts), or alkali metal chloride salts. The charge carrier may include ionic liquids or organic salts, such as tetramethylammonium chloride, trimethylphenylammonium chloride, phenyltrimethylammonium chloride, or 1-ethyl-3-methylimidazolium chloride. In the exemplary apparatus described above, the salt is present in an aqueous solution within the chamber. Potassium chloride (KCl), sodium chloride (NaCl), cesium chloride (CsCl), or mixtures of potassium ferrocyanide and potassium ferrocyanide are commonly used. Mixtures of KCl, NaCl, and potassium ferrocyanide and potassium ferrocyanide are preferred. The charge carrier on the membrane may be asymmetric. For example, the type and / or concentration of the charge carrier may differ on each side of the membrane.
[0983] The salt concentration can be saturated. The salt concentration can be 3M or lower, and is typically 0.1 to 2.5M, 0.3 to 1.9M, 0.5 to 1.8M, 0.7 to 1.7M, 0.9 to 1.6M, or 1M to 1.4M. The salt concentration is preferably 150mM to 1M. The method is preferably performed using a salt concentration of at least 0.3M, for example at least 0.4M, at least 0.5M, at least 0.6M, at least 0.8M, at least 1.0M, at least 1.5M, at least 2.0M, at least 2.5M, or at least 3.0M. High salt concentrations provide a high signal-to-noise ratio and allow the presence of the nucleotide to be identified to be indicated by current under normal current fluctuations.
[0984] The method is typically carried out in the presence of a buffer solution. In the exemplary apparatus described above, the buffer solution is present in an aqueous solution within the chamber. Any buffer solution can be used in the method of the present invention. Typically, the buffer solution is a phosphate buffer. Other suitable buffer solutions are HEPES or Tris-HCl buffer. The method is typically carried out at pH values of 4.0 to 12.0, 4.5 to 10.0, 5.0 to 9.0, 5.5 to 8.8, 6.0 to 8.7, or 7.0 to 8.8 or 7.5 to 8.5. A pH of about 7.5 is preferably used.
[0985] The method can be performed at temperatures ranging from 0°C to 100°C, 15°C to 95°C, 16°C to 90°C, 17°C to 85°C, 18°C to 80°C, 19°C to 70°C, or 20°C to 60°C. The method is typically performed at room temperature. Optionally, the method is performed at a temperature that supports enzyme function, such as approximately 37°C.
[0986] Polynucleotide binding protein
[0987] The chain characterization method preferably involves contacting the polynucleotide with a polynucleotide-binding protein, such that the protein controls the movement of the polynucleotide relative to a pore, for example, through the pore.
[0988] More preferably, the method includes (a) contacting a polynucleotide with a CsgG pore or a mutant thereof (e.g., the pore of the present invention) and a polynucleotide-binding protein, such that the protein controls the movement of the polynucleotide relative to the pore, for example, through the pore, and (b) acquiring one or more measurements as the polynucleotide moves relative to the pore, wherein the measurements indicate one or more characteristics of the polynucleotide, thereby characterizing the polynucleotide.
[0989] More preferably, the method includes (a) contacting a polynucleotide with a CsgG pore or a mutant thereof (e.g., the pore of the present invention) and a polynucleotide-binding protein, such that the protein controls the movement of the polynucleotide relative to the pore, for example, through the pore, and (b) measuring a current through the pore as the polynucleotide moves relative to the pore, wherein the current indicates one or more characteristics of the polynucleotide, thereby characterizing the polynucleotide.
[0990] Polynucleotide-binding proteins can be any protein capable of binding polynucleotides and controlling their movement through pores. Determining whether a protein is bound to a polynucleotide is relatively straightforward in the art. Proteins typically interact with polynucleotides and modify at least one property of the polynucleotide. Proteins can modify polynucleotides by cleaving them to form individual nucleotides or short chains of nucleotides (e.g., dinucleotides or trinucleotides). Proteins can also modify polynucleotides by orienting them or moving them to a specific location, i.e., controlling their movement.
[0991] Polynucleotide-binding proteins are preferably derived from polynucleotide-processing enzymes. Polynucleotide-processing enzymes are polypeptides capable of interacting with polynucleotides and modifying at least one property of the polynucleotides. The enzyme can modify the polynucleotide by cleaving it to form individual nucleotides or short chains of nucleotides (e.g., dinucleotides or trinucleotides). The enzyme can also modify the polynucleotide by orienting it or moving it to a specific location. Polynucleotide-processing enzymes do not need to exhibit enzymatic activity, as long as they can bind to polynucleotides and control their movement through pores. For example, the enzyme can be modified to remove its enzymatic activity, or it can be used under conditions that prevent it from being used as an enzyme. These conditions are discussed in more detail below.
[0992] The polynucleotide processing enzyme is preferably derived from a nucleolytic enzyme. More preferably, the polynucleotide processing enzyme used in the enzyme construct is derived from any one of the enzyme classification (EC) groups 3.1.11, 3.1.13, 3.1.14, 3.1.15, 3.1.16, 3.1.21, 3.1.22, 3.1.25, 3.1.26, 3.1.27, 3.1.30, and 3.1.31. The enzyme may be any of those enzymes disclosed in the international application PCT / GB10 / 000133 (published as WO2010 / 086603).
[0993] Preferably, the enzyme is a polymerase, exonuclease, helicase, and topoisomerase, such as a gyrase. Suitable enzymes include, but are not limited to, exonuclease I (SEQ ID NO: 399) from *Escherichia coli*, exonuclease III (SEQ ID NO: 401) from *Escherichia coli*, RecJ (SEQ ID NO: 403) from *Streptococcus thermophilus*, and bacterial phage λ exonuclease (SEQ ID NO: 405), TatD exonuclease, and variants thereof. Three subunits comprising the sequence shown in SEQ ID NO: 403 or a variant thereof interact to form a trimer exonuclease. These exonucleases can also be used in the exonuclease method of the present invention. The polymerase may be... 3173 DNA polymerase (which can be obtained from...) (purchased by the company), SD polymerase (available from...) (obtained) or a variant thereof. The enzyme is preferably Phi29 DNA polymerase (SEQ ID NO: 397) or a variant thereof. The topoisomerase is preferably any one of enzyme classification (EC) groups 5.99.1.2 and 5.99.1.3.
[0994] The enzyme is most preferably derived from a helicase, such as Hel308Mbu (SEQ ID NO: 406), Hel308 Csy (SEQ ID NO: 407), Hel308Tga (SEQ ID NO: 408), Hel308Mhu (SEQ ID NO: 409), Tral Eco (SEQ ID NO: 410), XPD Mbu (SEQ ID NO: 411), or variants thereof. Any helicase can be used in this invention. The helicase can be or is derived from Hel308 helicase, RecD helicase, such as Tral helicase or TrwC helicase, XPD helicase, or Dda helicase. The helicase may be any one of the helicases, modified helicases, or helicase constructs disclosed in the following international applications: PCT / GB2012 / 052579 (published as WO2013 / 057495); PCT / GB2012 / 053274 (published as WO2013 / 098562); PCT / GB2012 / 053273 (published as WO2013098561); PCT / GB2013 / 051925 (published as WO2013 / 057495); PCT / GB2 012 / 053274 (published as WO2013 / 098562); PCT / GB2012 / 053273 (published as WO2013098561); PCT / GB2013 / 051925 (published as WO2014 / 013260); PCT / GB2013 / 051924 (published as WO2014 / 013259); PCT / GB2013 / 051928 (published as WO2014 / 013262) and PCT / GB2014 / 052736.
[0995] The helicase preferably comprises the sequence shown in SEQ ID NO: 413 (Trwc Cba) or a variant thereof, the sequence shown in SEQ ID NO: 406 (Hel308Mbu) or a variant thereof, or the sequence shown in SEQ ID NO: 412 (Dda) or a variant thereof. Variants may differ from the native sequence in any of the ways discussed below regarding transmembrane pores. Preferred variants of SEQ ID NO: 412 include (a) E94C and A360C or (b) E94C, A360C, C109A and C136A, then optionally (ΔM1)G1G2 (i.e., M1 is deleted, then G1 and G2 are added).
[0996] According to the present invention, any number of helicases can be used. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more helicases can be used. In some embodiments, different numbers of helicases can be used.
[0997] The method of the present invention preferably involves contacting the polynucleotide with two or more helicases. The two or more helicases are typically the same helicase. Alternatively, the two or more helicases can be different helicases.
[0998] The two or more helicases can be any combination of the helicases described above. The two or more helicases can be two or more Dda helicases. The two or more helicases can be one or more Dda helicases and one or more TrwC helicases. The two or more helicases can be different variants of the same helicase.
[0999] The two or more helicases are preferably linked to each other. More preferably, the two or more helicases are covalently linked to each other. The helicases can be linked in any order and using any method. Preferred helicase constructs used in this invention are described in the following international applications: PCT / GB2013 / 051925 (published as WO2014 / 013260); PCT / GB2013 / 051924 (published as WO2014 / 013259); PCT / GB2013 / 051928 (published as WO2014 / 013262) and PCT / GB2014 / 052736.
[1000] Variants of SEQ ID NO: 397, 399, 401, 403, 405, 406, 407, 408, 409, 410, 411, 412, or 413 are enzymes having an amino acid sequence derived from the amino acid sequence of SEQ ID NO: 397, 399, 401, 403, 405, 406, 407, 408, 409, 410, 411, 412, or 413 and retaining polynucleotide binding capacity. This can be measured using any method known in the art. For example, the variant can be contacted with a polynucleotide, and its ability to bind to and move along the polynucleotide can be measured. Variants may include modifications that promote polynucleotide binding and / or enhance its activity at high salt concentrations and / or room temperature. Variants can be modified to bind polynucleotides (i.e., retain polynucleotide-binding ability) but not act as helicases (i.e., do not move along polynucleotides when provided with all the necessary components to promote movement, such as ATP and Mg). 2+ This modification is known in the art. For example, Mg in helicases 2+ Modification of the binding domain usually renders the variants unusable as helicases. These types of variants can act as molecular brakes (see below).
[1001] Based on amino acid identity, the variant peptide is preferably at least 50% homologous to the amino acid sequence of SEQ ID NO: 397,339,401,403,405,406,407,408,409,410,411,412, or 413 over its entire length. More preferably, based on amino acid identity, the variant peptide may be at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, more preferably, at least 95%, 97%, or 99% homologous to the amino acid sequence of SEQ ID NO: 397,399,401,403,405,406,407,408,409,410,411,412, or 413 over its entire length. Over a length of 200 or more, for example, 230, 250, 270, 280, 300, 400, 500, 600, 700, 800, 900, or 1000 or more consecutive amino acids, there may be at least 80%, for example, at least 85%, 90%, or 95% amino acid identity (“strict homology”). Homology is determined as described above. Variants may differ from the wild-type sequence in any of the manner discussed above with respect to SEQ ID NO: 390. The enzyme may be covalently linked to the pore. Any method may be used to covalently link the enzyme to the pore.
[1002] The preferred molecular brake is TrwC Cba-Q594A (SEQ ID NO: 413 with the mutant Q594A). This variant does not act as a helicase (i.e., it binds to polynucleotides when all the necessary components to promote movement are provided, but does not move along them, e.g., ATP and Mg). 2+ ).
[1003] During strand sequencing, polynucleotides shift through a well with or against an applied potential. Exonucleases that act gradually or progressively on double-stranded polynucleotides can be used on the cis side of the well to feed the remaining single strand through the well with an applied potential, or on the trans side with a reverse potential. Similarly, helicases that untangle double-stranded DNA can be used in a similar manner. Polymerases can also be used. Sequencing applications requiring strand shift with a reverse potential may also be used, but the DNA must first be "captured" by the enzyme in the presence of a reverse potential or no potential at all. Then, as the potential switches back after binding, the strand will pass through the well from cis to trans and remain in an extended conformation by current. Single-stranded DNA exonucleases or single-stranded DNA-dependent polymerases can act as molecular motors, pulling the just-shifted single strand back through the well in a controlled, stepwise manner against an applied potential, from trans to cis.
[1004] Any helicase can be used in the method. The helicase can operate relative to the pore in two modes. First, the method is preferably carried out using a helicase such that the helicase moves the polynucleotide through the pore by a field generated by an applied voltage. In this mode, the 5' end of the polynucleotide is first captured into the pore, and the helicase moves the polynucleotide into the pore such that it passes through the pore under the electric field until it is finally displaced to the trans side of the membrane. Alternatively, the method is preferably carried out such that the helicase moves the polynucleotide through the pore against the field generated by the applied voltage. In this mode, the 3' end of the polynucleotide is first captured into the pore, and the helicase moves the polynucleotide through the pore such that it is pulled out of the pore against the applied field until it finally springs back to the cis side of the membrane.
[1005] The method can also be performed in the reverse direction. The 3' end of the polynucleotide can be captured in the pore first, and the helicase can move the polynucleotide into the pore, allowing it to utilize the field through the pore until it is finally translocated to the trans side of the membrane.
[1006] When the helicase does not provide the necessary components to promote movement, or when the helicase is modified to hinder or prevent its movement, it can bind to polynucleotides and act as a brake to slow down polynucleotide movement as they are pulled into the pore by the applied field. In the non-active mode, whether the 3' or 5' end of the polynucleotide is captured is irrelevant; the applied field pulls the polynucleotide toward the trans side into the pore by the enzyme acting as a brake. The helicase's control over polynucleotide movement in the non-active mode can be described in several ways, including ratcheting, sliding, and braking. Helicase variants lacking helicase activity can also be used in this way.
[1007] The polynucleotide can be contacted with the polynucleotide-binding protein and the pore in any order. Preferably, when the polynucleotide is contacted with the polynucleotide-binding protein (e.g., a helicase) and the pore, the polynucleotide first forms a complex with the protein. When a voltage is applied to the pore, the polynucleotide / protein complex then forms a complex with the pore and controls the movement of the polynucleotide through the pore.
[1008] Any step in a method employing polynucleotide-binding proteins is typically performed in the presence of free nucleotides or free nucleotide analogs and enzyme cofactors that facilitate the action of polynucleotide-binding proteins. Free nucleotides can be any one or more individual nucleotides discussed above. Free nucleotides include, but are not limited to, adenosine monophosphate (AMP), adenosine diphosphate (ADP), adenosine triphosphate (ATP), guanosine monophosphate (GMP), guanosine diphosphate (GDP), guanosine triphosphate (GTP), thymidine monophosphate (TMP), thymidine diphosphate (TDP), thymidine triphosphate (TTP), uridine monophosphate (UMP), uridine diphosphate (UDP), uridine triphosphate (UTP), cytidine monophosphate (CMP), cytidine diphosphate (CDP), cytidine triphosphate (CTP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), and deoxyadenosine monophosphate (cGMP). Phosphate (dAMP), deoxyadenosine diphosphate (dADP), deoxyadenosine triphosphate (dATP), deoxyguanosine monophosphate (dGMP), deoxyguanosine diphosphate (dGDP), deoxyguanosine triphosphate (dGTP), deoxythymidine monophosphate (dTMP), deoxythymidine diphosphate (dTDP), deoxythymidine triphosphate (dTTP), deoxyuridine monophosphate (dUMP), deoxyuridine diphosphate (dUDP), deoxyuridine triphosphate (dUTP), deoxycytidine monophosphate (dCMP), deoxycytidine diphosphate (dCDP), and deoxycytidine triphosphate (dCTP). The free nucleotide is preferably selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, or dCMP. The free nucleotide is preferably adenosine triphosphate (ATP). The enzyme cofactor is a factor that enables the construct to function. The enzyme cofactor is preferably a divalent metal cation. The divalent metal cation is preferably Mg. 2+ Mn 2+ Ca 2+ , or Co 2+ The optimal enzyme cofactor is Mg. 2+ .
[1009] helicases and molecular brakes
[1010] In a preferred embodiment, the method includes:
[1011] (a) Providing a polynucleotide with one or more helicases and one or more molecular brakes linked to the polynucleotide;
[1012] (b) Contacting the polynucleotide with a CsgG pore or a mutant thereof, such as the pore of the present invention, and applying a potential to the pore to cause one or more helicases and one or more molecular brakes to aggregate together, both of which control the movement of the polynucleotide relative to the pore, for example, through the pore;
[1013] (c) As the polynucleotide moves relative to the pore, one or more measurements are acquired, wherein the measurements indicate one or more characteristics of the polynucleotide, thereby characterizing the polynucleotide.
[1014] This type of method is discussed in detail in international application PCT / GB2014 / 052737.
[1015] The one or more helicases may be any of those discussed above. The one or more molecular brakes may be any compound or molecule that binds to polynucleotides and slows their movement through the pore. The one or more molecular brakes preferably comprise one or more compounds that bind to polynucleotides. The one or more compounds are preferably one or more macrocyclic compounds. Suitable macrocyclic compounds include, but are not limited to, cyclodextrins, calixarenes, cyclic peptides, crown ethers, cucurbitacins, columnar aromatics, their derivatives, or combinations thereof. Cyclodextrins or their derivatives may be any of those disclosed in Elisev, AV and Schneider, HJ. (1994) J. Am. Chem. Soc 116, 6081-6088. The reagent is more preferably hepta-6-amino-β-cyclodextrin (am7-βCD), 6-monodeoxy-6-monoamino-β-cyclodextrin (am1-βCD), or hepta-(6-deoxy-6-guanidinyl)-cyclodextrin (gu7-βCD).
[1016] The one or more molecular brakes are preferably one or more single-chain binding proteins (SSBs). More preferably, the one or more molecular brakes are single-chain binding proteins (SSBs) comprising a carboxyl-terminal (C-terminal) region without a net negative charge, or (ii) modified SSBs having one or more modifications in their C-terminal region, said modifications reducing the net negative charge of the C-terminal region. Most preferably, the one or more molecular brakes are one of the SSBs disclosed in International Application PCT / GB2013 / 051924 (published as WO2014 / 013259).
[1017] The one or more molecular brakes are preferably one or more polynucleotide-binding proteins. A polynucleotide-binding protein can be any protein capable of binding to a polynucleotide and controlling its movement through a pore. It is readily apparent in the art whether a protein is bound to a polynucleotide. Proteins typically interact with polynucleotides and modify at least one property of the polynucleotide. Proteins can be modified by cleaving the polynucleotide to form individual nucleotides or short chains of nucleotides (e.g., dinucleotides or trinucleotides). This modification can also be achieved by orienting the polynucleotide or moving it to a specific location, i.e., controlling its movement.
[1018] The polynucleotide-binding protein is preferably derived from a polynucleotide processing enzyme. The one or more molecular brakes can be derived from any of the aforementioned polynucleotide processing enzymes. A modified version of the Phi29 polymerase (SEQ ID NO: 396) as a molecular brake is disclosed in U.S. Patent 5,576,204. The one or more molecular brakes are preferably derived from helicases.
[1019] Any number of molecular brakes derived from helicases can be used. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more helicases can be used as molecular brakes. If two or more helicases are used as molecular brakes, the two or more helicases are usually the same helicase. The two or more helicases can be different helicases.
[1020] The two or more helicases can be any combination of the helicases described above. The two or more helicases can be two or more Dda helicases. The two or more helicases can be one or more Dda helicases and one or more TrwC helicases. The two or more helicases can be different variants of the same helicase.
[1021] The two or more helicases are preferably linked to each other. More preferably, the two or more helicases are covalently linked to each other. The helicases can be linked in any order and using any method. Preferably, the one or more molecular brakes derived from the helicases are modified to reduce the size of the opening in the polynucleotide-binding domain, wherein the polynucleotide unwinds from the helicase through the polynucleotide-binding domain in at least one conformational state. This is disclosed in WO2014 / 013260.
[1022] The preferred helicase constructs used in this invention are described in the following international applications: PCT / GB2013 / 051925 (published as WO2014 / 013260); PCT / GB2013 / 051924 (published as WO2014 / 013259); PCT / GB2013 / 051928 (published as WO2014 / 013262) and PCT / GB2014 / 052736.
[1023] If the one or more helicases are used in active mode (i.e., when the one or more helicases provide all the necessary components for easy movement, such as ATP and Mg), 2+The one or more molecular brakes are preferably used in (a) a non-active mode (i.e., in the absence of necessary components to promote movement or in the absence of active movement); (b) an active mode, wherein the one or more molecular brakes move in the opposite direction to the one or more helicases; or (c) an active mode, wherein the one or more molecular brakes move in the same direction as the one or more helicases and move more slowly than the one or more helicases.
[1024] If the one or more helicases are in non-active mode (i.e., when the one or more helicases do not provide all the necessary components to promote movement, such as ATP and Mg), 2+ The one or more molecular brakes are preferably used (a) in an inactive mode (i.e., when there are no necessary components to promote movement or when active movement is not possible) or (b) in an active mode, wherein the one or more molecular brakes move along the polynucleotide in the same direction as the direction in which the polynucleotide passes through the pore.
[1025] The one or more helicases and one or more molecular brakes can be linked to polynucleotides at any position, causing them to aggregate together, and both controlling the movement of the polynucleotides through the pore. The one or more helicases and one or more molecular brakes are at least one separate nucleotide, for example, at least 5, at least 10, at least 50, at least 100, at least 500, at least 1000, at least 5000, at least 10,000, at least 50,000 or more separate nucleotides. If the method involves characterizing a double-stranded polynucleotide with a Y-adapter at one end and a hairpin loop adapter at the other end, the one or more helicases are preferably linked to the Y-adapter, and the one or more molecular brakes are preferably linked to the hairpin loop adapter. In this embodiment, the one or more molecular brakes are preferably one or more helicases modified to bind to polynucleotides but not used as helicases. The one or more helicases linked to the Y-adapter preferably stop at the spacer region, which will be discussed in more detail below. The one or more molecular brakes linked to the hairpin loop adapter preferably do not stop at the spacer region. When the one or more helicases reach the hairpin loop, the one or more helicases and one or more molecular brakes are preferably aggregated together. The one or more helicases can be connected to the Y adapter before or after the Y adapter is connected to the polynucleotide. The one or more molecular brakes can be connected to the hairpin loop adapter before or after the hairpin loop adapter is connected to the polynucleotide.
[1026] The one or more helicases and the one or more molecular brakes are preferably not interconnected. More preferably, the one or more helicases and the one or more molecular brakes are not covalently linked. As described in international applications PCT / GB2013 / 051925 (published as WO2014 / 013260); PCT / GB2013 / 051924 (published as WO2014 / 013259); PCT / GB2013 / 051928 (published as WO2014 / 013262) and PCT / GB2014 / 052736, the one or more helicases and the one or more molecular brakes are preferably not interconnected.
[1027] Interval
[1028] The one or more helicases may stop at one or more spacer regions as described in international application PCT / GB2014 / 050175. Any conformation of the one or more helicases and one or more spacer regions disclosed in that international application may be used in this invention.
[1029] As a portio...
Claims
1. A method for determining one or more characteristics of a target polynucleotide, comprising: (a) The target polynucleotide is brought into contact with a mutant CsgG pore within a membrane under an applied potential, such that the target polynucleotide binds to or moves through the mutant CsgG pore, wherein the membrane is a synthetic or isolated, naturally occurring lipid bilayer; and (b) One or more electrical measurements are performed while the polynucleotide is bound to or moves through the mutant CsgG well to determine one or more characteristics of the polynucleotide, wherein the one or more characteristics of the target polynucleotide are selected from (i) the length of the polynucleotide, (ii) the identity of the polynucleotide, (iii) the sequence of the polynucleotide, (iv) the secondary structure of the polynucleotide, and (v) whether the polynucleotide is modified, wherein the mutant CsgG well comprises at least one mutant CsgG monomer selected from: CsgG-Eco-(F56A), CsgG-Eco-(F56A-N55S) and CsgG-Eco-(F56A-N55S-Y51A); CsgG-Eco-(Y51N-F56A-D149N-E185R-E201N-E203N-Strepll(C)), CsgG-Eco-(N55A-Strepll(C)), CsgG-Eco-(N55S-Strepll(C)), CsgG-Eco-(Y51N-Str epll(C)), CsgG-Eco-(Y51A-F56A-Strepll(C)), CsgG-Eco-(Y51A-F56N-Strepll(C)), CsgG-Eco-(Y51A-N55S-F56A-Strepll(C)), CsgG-Eco-(Y51A-N55S-F56N-Strepll(C)), CsgG-Eco-(F56H-Strepll(C)), CsgG-Eco -(F56Q-Strepll(C)), CsgG-Eco-(F56T-Strepll(C)), CsgG-Eco-(S54P / F56A-Strepll(C)), CsgG-Eco-(Y51T / F56A-Strepll(C)), CsgG-Eco-(F56P-Strepll(C)), CsgG-Eco-(F56A-Strepll(C)), CsgG-Eco-(Y51T / F56Q-Str epll(C)), CsgG-Eco-(N55S / F56Q-Strepll(C)), CsgG-Eco-(Y51T / N55S / F56Q-Strepll(C)), CsgG-Eco-(F56Q / N102R-Strepll(C)), CsgG-Eco-(Y51Q / F56Q-Strepll(C)) and CsgG-Eco-(Y51A / F56Q-Strepll(C)), wherein Strepll(C) is SEQ ID NO:435 and is attached to the end of C; as well as CsgG-ΔPYPA, in which the sequence PYPA (residues 50-53) in the contractile ring is replaced by GG.
2. A method for determining the presence of a target polynucleotide, comprising: (a) The target polynucleotide is brought into contact with a mutant CsgG pore within a membrane at an applied potential, such that the target polynucleotide binds in or near the mutant CsgG pore, thereby influencing the ion flow through or moving through the mutant CsgG pore, wherein the membrane is a synthetic or isolated, naturally occurring lipid bilayer; and (b) The presence of a polynucleotide is determined by one or more electrical measurements performed when the polynucleotide is bound to, near, or moves through a mutant CsgG pore, wherein the mutant CsgG pore comprises at least one mutant CsgG monomer selected from: CsgG-Eco-(F56A), CsgG-Eco-(F56A-N55S) and CsgG-Eco-(F56A-N55S-Y51A); CsgG-Eco-(Y51N-F56A-D149N-E185R-E201N-E203N-Strepll(C)), CsgG-Eco-(N55A-Strepll(C)), CsgG-Eco-(N55S-Strepll(C)), CsgG-Eco-(Y51N-Str epll(C)), CsgG-Eco-(Y51A-F56A-Strepll(C)), CsgG-Eco-(Y51A-F56N-Strepll(C)), CsgG-Eco-(Y51A-N55S-F56A-Strepll(C)), CsgG-Eco-(Y51A-N55S-F56N-Strepll(C)), CsgG-Eco-(F56H-Strepll(C)), CsgG-Eco -(F56Q-Strepll(C)), CsgG-Eco-(F56T-Strepll(C)), CsgG-Eco-(S54P / F56A-Strepll(C)), CsgG-Eco-(Y51T / F56A-Strepll(C)), CsgG-Eco-(F56P-Strepll(C)), CsgG-Eco-(F56A-Strepll(C)), CsgG-Eco-(Y51T / F56Q-Str epll(C)), CsgG-Eco-(N55S / F56Q-Strepll(C)), CsgG-Eco-(Y51T / N55S / F56Q-Strepll(C)), CsgG-Eco-(F56Q / N102R-Strepll(C)), CsgG-Eco-(Y51Q / F56Q-Strepll(C)) and CsgG-Eco-(Y51A / F56Q-Strepll(C)), wherein Strepll(C) is SEQ ID NO:435 and is attached to the end of C; as well as CsgG-ΔPYPA, in which the sequence PYPA (residues 50-53) in the contractile ring is replaced by GG.
3. The method according to claim 1 or 2, wherein: (i) The electrical measurement is a current measurement; (ii) the combination of electrical measurements and optical measurements; and / or (ii) Step (a) further includes contacting the polynucleotide with a polynucleotide-binding protein such that the protein controls the movement of the polynucleotide through the mutant CsgG pore.
4. The method according to claim 3, wherein the polynucleotide binding protein is a helicase or a helicase derived from Hel308, RecD, XPD, or Dda.
5. A sensor suitable for characterizing a target polynucleotide according to the method of claim 1, the sensor comprising a mutant CsgG pore and a polynucleotide-binding protein within a membrane, wherein the mutant CsgG pore and the polynucleotide-binding protein form a complex in the presence of the target polynucleotide under an applied potential, and the membrane is a synthetic or isolated, naturally occurring lipid bilayer, wherein the mutant CsgG pore comprises at least one mutant CsgG monomer selected from: CsgG-Eco-(F56A), CsgG-Eco-(F56A-N55S) and CsgG-Eco-(F56A-N55S-Y51A); CsgG-Eco-(Y51N-F56A-D149N-E185R-E201N-E203N-Strepll(C)), CsgG-Eco-(N55A-Strepll(C)), CsgG-Eco-(N55S-Strepll(C)), CsgG-Eco-(Y51N-Str epll(C)), CsgG-Eco-(Y51A-F56A-Strepll(C)), CsgG-Eco-(Y51A-F56N-Strepll(C)), CsgG-Eco-(Y51A-N55S-F56A-Strepll(C)), CsgG-Eco-(Y51A-N55S-F56N-Strepll(C)), CsgG-Eco-(F56H-Strepll(C)), CsgG-Eco -(F56Q-Strepll(C)), CsgG-Eco-(F56T-Strepll(C)), CsgG-Eco-(S54P / F56A-Strepll(C)), CsgG-Eco-(Y51T / F56A-Strepll(C)), CsgG-Eco-(F56P-Strepll(C)), CsgG-Eco-(F56A-Strepll(C)), CsgG-Eco-(Y51T / F56Q-Str epll(C)), CsgG-Eco-(N55S / F56Q-Strepll(C)), CsgG-Eco-(Y51T / N55S / F56Q-Strepll(C)), CsgG-Eco-(F56Q / N102R-Strepll(C)), CsgG-Eco-(Y51Q / F56Q-Strepll(C)), and CsgG-Eco-(Y51A / F56Q-Strepll(C)), wherein Strepll(C) is SEQ ID NO:435 and is attached to the C-terminus; and CsgG-ΔPYPA, in which the sequence PYPA (residues 50-53) in the contractile ring is replaced by GG.
6. A kit suitable for characterizing a target polynucleotide according to the method of claim 1, comprising (a) a mutant CsgG well, (b) components of a synthetic membrane, and (c) a voltage device, wherein the mutant CsgG well comprises at least one mutant CsgG monomer selected from: CsgG-Eco-(F56A), CsgG-Eco-(F56A-N55S) and CsgG-Eco-(F56A-N55S-Y51A); CsgG-Eco-(Y51N-F56A-D149N-E185R-E201N-E203N-Strepll(C)), CsgG-Eco-(N55A-Strepll(C)), CsgG-Eco-(N55S-Strepll(C)), CsgG-Eco-(Y51N-Str epll(C)), CsgG-Eco-(Y51A-F56A-Strepll(C)), CsgG-Eco-(Y51A-F56N-Strepll(C)), CsgG-Eco-(Y51A-N55S-F56A-Strepll(C)), CsgG-Eco-(Y51A-N55S-F56N-Strepll(C)), CsgG-Eco-(F56H-Strepll(C)), CsgG-Eco -(F56Q-Strepll(C)), CsgG-Eco-(F56T-Strepll(C)), CsgG-Eco-(S54P / F56A-Strepll(C)), CsgG-Eco-(Y51T / F56A-Strepll(C)), CsgG-Eco-(F56P-Strepll(C)), CsgG-Eco-(F56A-Strepll(C)), CsgG-Eco-(Y51T / F56Q-Str epll(C)), CsgG-Eco-(N55S / F56Q-Strepll(C)), CsgG-Eco-(Y51T / N55S / F56Q-Strepll(C)), CsgG-Eco-(F56Q / N102R-Strepll(C)), CsgG-Eco-(Y51Q / F56Q-Strepll(C)), and CsgG-Eco-(Y51A / F56Q-Strepll(C)), wherein Strepll(C) is SEQ ID NO:435 and is attached to the C-terminus; and CsgG-ΔPYPA, in which the sequence PYPA (residues 50-53) in the contractile ring is replaced by GG.
7. A pore-forming CsgG monomer, wherein the CsgG monomer is intramembrane and is a mutant of the SEQ ID NO:390 sequence with an amino acid mutation, and the mutant is selected from: CsgG-Eco-(F56A), CsgG-Eco-(F56A-N55S) and CsgG-Eco-(F56A-N55S-Y51A); CsgG-Eco-(Y51N-F56A-D149N-E185R-E201N-E203N-Strepll(C)), CsgG-Eco-(N55A-Strepll(C)), CsgG-Eco-(N55S-Strepll(C)), CsgG-Eco-(Y51N-Str epll(C)), CsgG-Eco-(Y51A-F56A-Strepll(C)), CsgG-Eco-(Y51A-F56N-Strepll(C)), CsgG-Eco-(Y51A-N55S-F56A-Strepll(C)), CsgG-Eco-(Y51A-N55S-F56N-Strepll(C)), CsgG-Eco-(F56H-Strepll(C)), CsgG-Eco -(F56Q-Strepll(C)), CsgG-Eco-(F56T-Strepll(C)), CsgG-Eco-(S54P / F56A-Strepll(C)), CsgG-Eco-(Y51T / F56A-Strepll(C)), CsgG-Eco-(F56P-Strepll(C)), CsgG-Eco-(F56A-Strepll(C)), CsgG-Eco-(Y51T / F56Q-Str epll(C)), CsgG-Eco-(N55S / F56Q-Strepll(C)), CsgG-Eco-(Y51T / N55S / F56Q-Strepll(C)), CsgG-Eco-(F56Q / N102R-Strepll(C)), CsgG-Eco-(Y51Q / F56Q-Strepll(C)), and CsgG-Eco-(Y51A / F56Q-Strepll(C)), wherein Strepll(C) is SEQ ID NO:435 and is attached to the C-terminus; and CsgG-ΔPYPA, in which the sequence PYPA (residues 50-53) in the contractile ring is replaced by GG.
8. A construct comprising two or more covalently linked CsgG monomers, wherein at least one of the monomers is a monomer according to claim 7.
9. A modified CsgG pore, which is as follows: (a) A modified CsgG pore comprising at least one CsgG monomer, wherein: The pore comprises the monomer according to claim 7; or (b) Homologous oligomeric pores comprising the same monomer as claimed in claim 7 or the same construct as claimed in claim 8; or (c) Includes at least one monomer according to claim 7 or at least one construct according to claim 8 with heterogeneous oligopores.
10. A polynucleotide encoding a CsgG monomer, wherein the monomer is a mutant selected from: CsgG-Eco-(F56A), CsgG-Eco-(F56A-N55S) and CsgG-Eco-(F56A-N55S-Y51A); CsgG-Eco-(Y51N-F56A-D149N-E185R-E201N-E203N-Strepll(C)), CsgG-Eco-(N55A-Strepll(C)), CsgG-Eco-(N55S-Strepll(C)), CsgG-Eco-(Y51N-Str epll(C)), CsgG-Eco-(Y51A-F56A-Strepll(C)), CsgG-Eco-(Y51A-F56N-Strepll(C)), CsgG-Eco-(Y51A-N55S-F56A-Strepll(C)), CsgG-Eco-(Y51A-N55S-F56N-Strepll(C)), CsgG-Eco-(F56H-Strepll(C)), CsgG-Eco -(F56Q-Strepll(C)), CsgG-Eco-(F56T-Strepll(C)), CsgG-Eco-(S54P / F56A-Strepll(C)), CsgG-Eco-(Y51T / F56A-Strepll(C)), CsgG-Eco-(F56P-Strepll(C)), CsgG-Eco-(F56A-Strepll(C)), CsgG-Eco-(Y51T / F56Q-Str epll(C)), CsgG-Eco-(N55S / F56Q-Strepll(C)), CsgG-Eco-(Y51T / N55S / F56Q-Strepll(C)), CsgG-Eco-(F56Q / N102R-Strepll(C)), CsgG-Eco-(Y51Q / F56Q-Strepll(C)), and CsgG-Eco-(Y51A / F56Q-Strepll(C)), wherein Strepll(C) is SEQ ID NO:435 and is attached to the C-terminus; and CsgG-ΔPYPA, in which the sequence PYPA (residues 50-53) in the contractile ring is replaced by GG.
Citation Information
Patent Citations
Enzyme-pore constructs
EP2682460A1
phi 29 DNA polymerase
US5576204A
A miniature support for thin films containing single channels or nanopores and methods for using same
WO2000028312A1
Suspended carbon nanotube field effect transistor
WO2005124888A1
Deliver of molecules to a li id bila
WO2006100484A2