Improved separated HaloTag
By introducing specific amino acid substitutions and linker sequences into the HaloTag system, the stability and affinity of HaloTag were improved, solving the problems of instability and insufficient affinity of the existing system at high temperatures, and achieving more efficient biological process research.
Patent Information
- Application Number
- CN202480014434.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-22
- Filing Date
- 2024-02-22
- Publication Date
- 2025-10-03
AI Technical Summary
The existing HaloTag system is not stable enough under high temperature conditions and has insufficient affinity, resulting in high background signals and low signal-to-noise ratios in biological process research, making it difficult to achieve rapid self-complementary labeling reactions.
A new modular peptide complex was designed. By introducing specific amino acid substitutions and linker sequences into the first and second effector sequences of HaloTag, the thermal stability and affinity of cpHaloΔ were improved, and the binding ability of the halogenated peptide to cpHaloΔ was enhanced.
The stability and activity of the HaloTag system at 37°C are improved, the background signal is reduced, and the signal-to-noise ratio is enhanced, making it suitable for a wider range of bioprocess research applications.
Smart Images

Figure BDA0005561260450000241 
Figure BDA0005561260450000242 
Figure BDA0005561260450000251
Abstract
Description
[0001] This application claims the benefit of priority from European patent application 23158085.3 filed on February 22, 2023, the contents of which are incorporated herein by reference. Technical Field
[0002] The present invention relates to improved variants of the split HaloTag system. HaloTag is a self-labeling protein tag derived from a bacterial enzyme and capable of covalently binding to synthetic ligands, including reactive chloroalkanes. HaloTag can be split into the cpHaloΔ protein and a complementary peptide, which is active only when brought into close proximity. The present invention relates to improved sequences of this split HaloTag system that are more stable and exhibit faster kinetics. Background Art
[0003] Methods for integrating biochemical processes over time are of great scientific interest. These processes include, for example, protein-protein interactions, changes in metabolite concentrations, or protein subcellular (re)localization. Currently, these processes are primarily studied using real-time fluorescence microscopy using specialized biosensors. However, fluorescence microscopy is limited to a limited field of view and tissue depth. Despite significant improvements in instrumentation, image processing, and biosensor design, it remains impossible to study fundamental events in intact rodent organs in vivo using fluorescence microscopy.
[0004] An alternative approach involves recording biochemical processes that can be read out at a later time point. This recording process is known as integration. It results in the accumulation of irreversible markers over time in response to the biochemical process being studied.
[0005] Post-experimental evaluation is valuable not only for in vivo studies, but also for cell-based analysis. Especially in the case of rare biochemical processes, signal integration over time can improve readout due to the enhanced signal-to-noise ratio. In high-throughput screening, signal integration can provide a snapshot of the phenomenon under study, thereby increasing multiplexing and reducing costs by replacing lengthy recordings with real-time microscopy. Finally, using fluorescent substrates, cell populations can be identified and ultimately classified based on their metabolic / signaling profiles for downstream analysis and / or processing.
[0006] Based on the self-labeling protein HaloTag, a family of chemical genetic integrins has been produced that have been engineered to react irreversibly with chloroalkane substrates in response to a given biochemical process. These integrons are composed of split ring arrangements of HaloTag, which are complemented as a result of conformational changes in response to biochemical processes as sensors, and cause labeling activity as a result of this complementation. The circularly arranged HaloTag is shown in SEQ ID NO: 001. The original system is disclosed in WO2020212537 and U.S. application 17 / 604,417. It includes a first part effector (SEQ ID NO: 367; cpHaloΔ protein) consisting of two components of the HaloTag protein (SEQ ID NO: 002 and 003) connected by an internal linker, and is complementary to a short second part effector peptide (SEQ ID NO: 004 or SEQ ID NO: 005). However, the original separated HaloTag system has several disadvantages.
[0007] When the split HaloTag system is connected to two relatively far apart (e.g. When the end of the interacting protein of cpHaloΔ was fused, no complementation / labeling was observed. d In addition, the low thermal stability of cpHaloΔ at 30°C limits its application in many cell or in vivo systems.
[0008] Based on the above-described prior art, the object of the present invention is to provide means and methods for designing novel halogenated peptides with increased affinity for cpHaloΔ, as well as for designing cpHaloΔ variants with increased thermostability and high activity at 37°C. This object is achieved by the subject matter of the independent claims and by the further advantageous embodiments described in the dependent claims, examples, figures, and general description of the present description. Summary of the Invention
[0009] The first aspect of the present invention relates to a modular polypeptide complex comprising a first partial effector sequence constituting the cpHaloΔ protein and a second partial effector sequence (or essentially consisting thereof), characterized in that the second partial effector sequence consists of a sequence selected from the group consisting of SEQ ID NO: 6 to SEQ ID NO: 343.
[0010] A second aspect of the present invention relates to a modular polypeptide complex comprising a first part effector sequence and a second part effector sequence (or consisting essentially of the same), characterized in that
[0011] a The N-terminal first effector sequence portion comprises at least one N-effector amino acid substitution; and / or
[0012] b. The C-terminal first effector sequence portion comprises at least one C-effector amino acid substitution; and / or
[0013] c. The internal cpHalo linker is characterized by the novel sequences provided herein, in particular the sequences selected from SEQ ID NOs: 344 to 366, as listed below.
[0014] The third aspect of the present invention relates to a modular polypeptide complex, comprising the first partial effector sequence as described in the second aspect, and the second partial effector sequence as described in the first aspect.
[0015] These alternatives involve modular polypeptide complexes comprising a variant of the first portion, the effector sequence, consisting of:
[0016] The N-terminal first effector sequence portion is characterized by SEQ ID NO: 002,
[0017] a C-terminal first effector sequence portion characterized by SEQ ID NO: 003, and
[0018] an internal cpHalo linker, consisting of 10 to 35 amino acids, connecting the C-terminus of the N-terminal first effector sequence portion to the N-terminus of the C-terminal first effector sequence portion;
[0019] and a second partial effector sequence consisting of a sequence selected from the group consisting of SEQ ID NO:004 to SEQ ID NO:343, wherein the first and second partial effector sequences together constitute an active, circularly permuted, variant of the self-labeling protein identified as GenBank-ID: AQS79242.1 and are capable of achieving covalent attachment of a haloalkyl moiety when brought into close proximity with one another. According to this alternative aspect of the invention, the modular polypeptide complex is characterized by at least one substitution relative to the sequence of GenBank-ID: AQS79242.1 selected from the group consisting of E20S; V184E; N119H; V197K; K117R; N217D; F205W; F80T; R30P; and G7D.
[0020] Other aspects of the present invention relate to nucleic acid sequences or nucleic acid expression systems encoding modular polypeptide complexes as described herein, cells or non-human transgenic animals or plants comprising modular polypeptide complexes, and kits and methods for detecting molecular interaction events.
[0021] The inventors surprisingly discovered that the previously developed circularly arrayed HaloTag system can be fine-tuned through a combination of phylogenetics and in silico force field calculations to generate a toolbox that provides a wide range of affinity and activity values resulting from mutations or variations in the original sequence.
[0022] Low affinity between components of the cpHaloTag system is advantageous in the study of biological processes, resulting in reduced background signal, while high affinity is advantageous for rapid self-complementation. As shown in the Examples, mutations interact in a synergistic manner rather than delivering cumulative increases in activity or stability.
[0023] Of the cpHaloΔ variant positions analyzed (for interpretation of variant / mutation designations, see below), E20S was a particularly useful variant position because of its contribution to stability and activity in the experiments provided herein.
[0024] Terms and Definitions
[0025] For purposes of interpreting this specification, the following definitions will apply, and whenever appropriate, terms used in the singular will also include the plural, and vice versa. In the event of a conflict between any definition set forth below and any document incorporated herein by reference, the set forth definition will control.
[0026] As used herein, the terms "comprising," "having," "containing," and "including," and other similar forms, and grammatical equivalents thereof, are intended to be equivalent in meaning and open-ended, in that one or more items following any of these terms is not meant to be an exhaustive list of such one or more items, or to be limited to only the listed one or more items. For example, something that "comprising" components A, B, and C may consist of (i.e., contain only) components A, B, and C, or may contain not only components A, B, and C, but may also include one or more other components. Thus, it is intended and understood that "comprising" and its similar forms and grammatical equivalents thereof include disclosure of embodiments that "consist essentially of" or "consist of."
[0027] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit, between the upper and lower limits of that range, and any other stated or intervening value in that range, is encompassed within the invention, subject to any specifically excluded limit in that range, unless the context clearly dictates otherwise. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also encompassed within the disclosure.
[0028] Reference herein to "about" a value or parameter includes (and describes) variations involving the value or parameter itself. For example, a description referring to "about X" includes a description of "X."
[0029] As used herein, including in the appended claims, the singular forms "a," "an," "or," and "the" include plural referents unless the context clearly dictates otherwise.
[0030] “And / or” is used herein as a specific recitation of each of two specified features or components, with or without the other features or components. Thus, the term “and / or” used in phrases such as “A and / or B” is intended to include “A and B,” “A or B,” “A” (alone), and “B” (alone). Similarly, the term “and / or” used in phrases such as “A, B, and / or C” is intended to cover each of the following: A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone); B (alone); and C (alone).
[0031] Unless defined otherwise, all technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art (e.g., in cell culture, molecular genetics, nucleic acid chemistry, hybridization techniques and biochemistry, organic synthesis). Standard techniques are used for molecular, genetic, and biochemical methods (see generally Sambrook et al., Molecular Cloning: A Laboratory Manual, 4 th ed. (2012) Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY and Ausubelet al., Short Protocols in Molecular Biology (2002) 5 th Ed, John Wiley & Sons, Inc.) and chemical methods.
[0032] In the context of this specification, the term close proximity to each other relates to situations where the first and second partial effector sequences are able to interact non-covalently. In certain embodiments, the sensor polypeptide is necessary to bring the first and second partial effector sequences into close enough proximity to interact. In certain embodiments, the first and second partial effector sequences have a sufficiently high mutual affinity that they can interact without triggering by the sensor polypeptide. In this case, the first and second partial effector sequences interact even if both are expressed only in the same cell or cell-free expression system. In certain embodiments, "close proximity" refers to a distance In certain embodiments, "close proximity" refers to a distance In certain embodiments, "close proximity" refers to a distance In certain embodiments, "close proximity" refers to a distance It is understood that the distance between partners is affected by the heat flux and that the distance is not static. Therefore, the distance ” means that the partners are at a distance apart for a time interval sufficient for a reaction or interaction between the partners to occur.
[0033] In the context of this specification, the term coupling relates to a covalent bond connecting two partners that are coupled to each other. This covalent bond can be directly between the two partners or there can be another moiety in between.
[0034] In the context of this specification, the term deletion relates to the event whereby one amino acid is deleted from a sequence, resulting in a sequence that has one less amino acid than the parent sequence.
[0035] In the context of this specification, the term insertion relates to the event of inserting an amino acid into a sequence resulting in a sequence having one more amino acid than the parent sequence.
[0036] In the context of this specification, the term substitution refers to the event whereby a "wild-type" amino acid is replaced by another mutant or substituted amino acid at the same position in the sequence, resulting in a sequence having an amino acid that differs from the amino acid at the same position within the parent sequence (a substitution compared to the parent sequence). Variable substitutions refer to substitutions that do not significantly reduce the function and activity of the parent sequence. Variable substitutions are substitutions that retain all the important characteristics of the parent sequence. N-effector and C-effector amino acid substitutions are substitutions that improve certain characteristics of the parent sequence, as described in the specification. N-effector and C-effector amino acid substitutions are selected from a defined list of possible substitutions.
[0037] In the context of this specification, the term detection moiety refers to a chemical moiety that can be detected in solution, in a suspension of organic particles, or within a cell. Detection can be by, but is not limited to, light emission, radioactivity, or certain specific binding properties of the detection moiety. In certain embodiments, the detection moiety comprises a fluorophore. In certain specific embodiments, the detection moiety comprises a fluorescent dye molecule.
[0038] In the context of this specification, the term purification tag refers to a chemical moiety that has a high affinity for and specifically binds to a receptor binding moiety. The binding of the purification tag to the receptor moiety is specific and can be used to purify the moiety covalently coupled to the purification tag, for example, by affinity chromatography. Non-limiting examples of purification tags include a T7 tag, a calmodulin binding peptide, a FLAG epitope, a GST tag, an HA tag, a polyhistidine tag, a Myc tag, biotin, and a Strep tag.
[0039] Terms used in this manual Relates to self-labeling protein tags. In its most commonly used form, the conventional HaloTag is a 297 residue protein (33 kDa) derived from the bacterial haloalkane dehalogenase, designed for covalent binding to a synthetic ligand. The bacterial enzyme can be fused to a variety of target proteins [Los et al., ACS Chemical Biology. 3(6): 373–82]. Depending on the type of experiment to be performed, the synthetic ligand is selected from a number of available ligands. The enzyme is designed to facilitate visualization of the subcellular localization of the target protein, immobilization of the target protein, or capture of the target protein's binding partner in its biochemical environment [Giepmans et al., Science. 312(5771): 217–24]. The HaloTag system consists of two covalently bound fragments, including the haloalkane dehalogenase and a selected synthetic ligand. These synthetic ligands consist of a reactive chloroalkane linker bound to a functional group. HaloTag is a trademark of Promega Corporation.
[0040] In the context of this specification, the term haloalkyl moiety or HaloTag substrate is synonymous with the term haloalkyl dehalogenase substrate and relates to an ω-haloalkyl moiety that can be covalently attached to a HaloTag protein. In certain embodiments, the HaloTag substrate is a moiety of the formula:
[0041] (6-chlorohexyl),
[0042] Wherein R can be any moiety as further defined in the specification. In certain embodiments, R is a detection moiety. In certain specific embodiments, R comprises a fluorescent dye. In certain embodiments, R also comprises a linker. The HaloTag substrate can also be an ω-bromoalkane, while aryl-ω-haloalkyl substrates are known and can be substituted (see Shields et al., Thousandfold Cell-Specific Pharmacology of Neurotransmission; BioRxiv Oct. 21, 2022).
[0043] In certain embodiments, non-covalent HaloTag substrates are encompassed by the term haloalkane moiety or HaloTag substrate. Non-covalent HaloTag substrates are disclosed in WO2022229466A1, which is incorporated herein by reference.
[0044] When used in relation to a HaloTag protein or variants thereof, the term "activity" refers to the ability of the protein to self-tag, i.e. covalently attach to a HaloTag substrate, or - if reference is made in the context of a HaloTag system, where non-covalent attachment of a HaloTag substrate is contemplated as disclosed in WO2022229466 - non-covalently bind and release a non-covalent HaloTag substrate.
[0045] The term "hpep" or "hpeps" (plural) refers to the short complementary peptide of the cp HaloTag system, also referred to as the second partial effector sequence when describing the present invention.
[0046] In the context of this specification, the term in-frame insertion of a coding sequence relates to inserting an open reading frame at a position relative to another open reading frame such that both open reading frames can be transcribed and translated into one continuous peptide or polypeptide.
[0047] In the context of this specification, the term multiple cloning site relates to a short DNA fragment containing several restriction sites, each of which is uniquely cut by a specific restriction endonuclease.
[0048] Any patent documents cited herein should be deemed to be incorporated by reference in their entirety.
[0049] sequence
[0050] Sequences similar or homologous to the sequences disclosed herein are also part of the present invention (always assuming that any mutations characterizing the present invention are present and that the sequences are functional in the sense indicated herein). In some embodiments, the sequence identity at the amino acid level may be about 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more. At the nucleic acid level, the sequence identity may be about 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more. Alternatively, substantial identity exists when the nucleic acid segment hybridizes to the complement of the strand under selective hybridization conditions (e.g., very high stringency hybridization conditions). The nucleic acid may be present in whole cell, cell lysate, or partially purified or substantially pure form.
[0051] In the context of this specification, the terms sequence identity and sequence identity percentage refer to a single quantitative parameter representing a sequence comparison result determined by comparing two aligned sequences position by position. Methods for aligning sequences for comparison are well known in the art. Sequence alignments for comparison can be performed by the local homology algorithm of Smith and Waterman, Adv.Appl.Math.2:482 (1981), by the global alignment algorithm of Needleman and Wunsch, J.Mol.Biol.48:443 (1970), by the similarity search method of Pearson and Lipman, Proc.Nat.Acad.Sci.85:2444 (1988), or by computerized implementations of these algorithms, including but not limited to: CLUSTAL, GAP, BESTFIT, BLAST, FASTA, and TFASTA. Software for performing BLAST analysis is publicly available, for example, through the National Center for Biotechnology Information ( http: / / blast.ncbi.nlm.nih.gov / ).
[0052] An example of an algorithm for amino acid sequence comparison is the BLASTP algorithm, which uses the default settings: Desire threshold: 10; Word length: 3; Maximum match in query: 0; Matrix: BLOSUM62; Gap cost: Existence 11, Extension 1; Composition correction: Conditional combination score matrix adjustment. An example of an algorithm for nucleic acid sequence comparison is the BLASTN algorithm, which uses the default settings: Desire threshold: 10; Word length: 28; Maximum match in query: 0; Match / mismatch score: 1.-2; Gap cost: Linear. Unless otherwise indicated, sequence identity values provided herein refer to the values obtained using the BLAST suite of programs (Altschule et al., J. Mol. Biol. 215:403–410 (1990)) using the default parameters specified above for protein and nucleic acid alignments, respectively.
[0053] Reference to identical sequences without specifying a percentage value implies 100% identical sequences (ie, identical sequences).
[0054] General biochemistry: peptides, amino acid sequences
[0055] In the context of this specification, the term polypeptide refers to a molecule consisting of 50 or more amino acids forming a linear chain in which the amino acids are linked by peptide bonds. The amino acid sequence of a polypeptide may represent the amino acid sequence of an entire (as found physiologically) protein or a fragment thereof. The terms "polypeptide" and "protein" are used interchangeably herein and include proteins and fragments thereof. A polypeptide is disclosed herein as a sequence of amino acid residues.
[0056] In the context of the present specification, the term peptide relates to a molecule consisting of up to 50 amino acids, particularly 8 to 30 amino acids, more particularly 8 to 15 amino acids, forming a linear chain wherein the amino acids are linked by peptide bonds.
[0057] The amino acid residue sequences are given from the amino terminus to the carboxyl terminus. The capital letters at the sequence positions refer to the L-amino acids in the single-letter code (Stryer, Biochemistry, 3 rd ed.p.21). Lowercase letters at the amino acid sequence position refer to the corresponding D- or (2R)-amino acid. The sequence is written from left to right in the direction from the amino terminus to the carboxyl terminus. According to standard nomenclature, the amino acid residue sequence is named using the three-letter or one-letter code as follows: alanine (Ala, A), arginine (Arg, R), asparagine (Asn, N), aspartic acid (Asp, D), cysteine (Cys, C), glutamine (Gln, Q), glutamic acid (Glu, E), glycine (Gly, G), histidine (His, H), isoleucine (Ile, I), leucine (Leu, L), lysine (Lys, K), methionine (Met, M), phenylalanine (Phe, F), proline (Pro, P), serine (Ser, S), threonine (Thr, T), tryptophan (Trp, W), tyrosine (Tyr, Y), and valine (Val, V).
[0058] In the context of this specification, the term having substantially the same activity relates to the activity of the effector polypeptide pair, i.e., the haloalkyl transferase activity. A polypeptide having substantially the same activity as a reference does not necessarily show the same amount of activity as the reference polypeptide; in the specific case of the present invention, for certain applications, a reduction in (self-)labeling activity relative to the reference peptide (SEQ ID NO: 004 or 005) may indeed be desirable. As described below, in order to distinguish polypeptides encompassed by the present invention from polypeptides not encompassed, the inventors set a 10% scavenger concentration in a standard assay as described in Example 3 using halo-CPY (T5-CPY HaloTag substrate) as substrate. 2 s -1 Activity threshold of M-1. When polypeptides or oligopeptides are identified herein as belonging to the present invention as part of a modular peptide complex, it is understood that they have haloalkyl transferase activity.
[0059] For purposes where the above definition of activity is not applicable, 3 standard deviations above background of haloalkyl transferase activity should be considered a reference threshold for having substantially the same activity. In certain embodiments, at least 5 standard deviations are used as a reference threshold. In certain specific embodiments, at least 10 standard deviations are used as a reference threshold.
[0060] In the context of this specification, the term amino acid linker refers to a polypeptide of variable length for connecting two polypeptides to produce a single-chain polypeptide. Unless otherwise indicated, exemplary embodiments of linkers for practicing the invention specified herein are oligopeptide chains consisting of 1, 2, 3, 4, 5, 10, 20, 30, 40 or 50 amino acids.
[0061] In the context of this specification, the term "circular permutation" or "circular permutation," as commonly used in the field of protein engineering, refers to the process of reorganizing the primary sequence of a protein by joining the original N-terminus (start) and C-terminus (stop) of the protein together and then introducing new N-termini and C-termini at different positions within the protein. This technique does not change the sequence of amino acids, but changes the order in which they are linked, potentially affecting the structure, stability, and function of the protein. Circular permutation can be used to study protein folding and function, as well as to engineer proteins with new properties or functions.
[0062] General molecular biology: nucleic acid sequences, expression
[0063] In the context of this specification, the term nucleic acid expression vector refers to a plasmid, viral genome, or RNA that is used to transfect (in the case of a plasmid or RNA) or transduce (in the case of a viral genome) a target cell with a gene of interest, or—in the case of a transfected RNA construct—translate the corresponding protein of interest from the transfected mRNA. For vectors that operate at the level of transcription and subsequent translation, the gene of interest is under the control of a promoter sequence that is operable within the target cell, so that the gene of interest is transcribed constitutively, in response to a stimulus, or depending on the state of the cell. In certain embodiments, the viral genome is packaged into a capsid to become a viral vector capable of transducing a target cell.
[0064] Labels, ligands
[0065] In the context of this specification, the term fluorescent dye or fluorophore refers to a small molecule that can fluoresce in the visible or near-infrared spectrum. Examples of fluorescent labels or labels that exhibit visible colors include, but are not limited to, fluorescein, rhodamine and silicon-rhodamine-based dyes, allophycocyanin (APC), peridinin chlorophyll (PerCP), phycoerythrin (PE), Alexa Fluors (Life Technologies, Carlsbad, California, USA), DyLight Fluors (Thermo Fisher Scientific, Waltham, Massachusetts, USA), ATTO dyes (ATTO-TEC GmbH, Siegen, Germany), BODY dyes (based on 4,4-difluoro-4-borane-3a,4a-diaza-s-dicyclopentacarbonyl acene dyes), etc. In the context of this specification, the term fluorescent dye or fluorophore also refers to the dyes described in WO2020115286 and WO2019122269A1, or US16956596, which are incorporated herein by reference. DETAILED DESCRIPTION
[0066] Improved marking speed:
[0067] The first aspect of the present invention relates to a modular polypeptide complex comprising a first partial effector sequence and a second partial effector sequence (or essentially consisting thereof), characterized in that the second partial effector sequence consists of a sequence selected from the group consisting of SEQ ID NO: 6 to SEQ ID NO: 343.
[0068] The improved sequences provided by the present invention can modulate the affinity between two different, separate parts of the HaloTag protein, allowing for a wider range of uses for the tool, given the many different peptides (sensors, see below) that can be attached. The peptides provided by the present invention cover a wide range of affinities, enabling a variety of different applications of this technology.
[0069] Optionally, the second partial effector sequence can include one or two amino acid variations independently selected from the group consisting of a deletion, an insertion, and a variable substitution.
[0070] The first and second partial effector sequences together comprise an active, circularly arranged, haloalkane dehalogenase-derived, self-labeling protein and are capable of covalently linking the haloalkane moiety to the first partial effector sequence when brought into close proximity with one another.
[0071] The first part of the effector sequence includes
[0072] o an N-terminal first effector sequence portion characterized by SEQ ID NO: 002 (original N-terminal cpHaloΔ protein) or a sequence at least (≥) 90% identical (particularly ≥93%, 95%, 97% or ≥98% identical) to SEQ ID NO: 002,
[0073] o a C-terminal first effector sequence portion characterized by SEQ ID NO: 003 (the original C-terminal cpHaloΔ protein) or a sequence at least (≥) 90% identical (in particular ≥93%, 95%, 97% or ≥98% identical) to SEQ ID NO: 003,
[0074] An internal cpHalo linker consisting of 10 to 35 amino acids, wherein the internal cpHalo linker connects the C-terminus of the N-terminal first effector sequence portion to the N-terminus of the C-terminal first effector sequence portion.
[0075] In a specific embodiment, the C-terminal effector sequence portion is characterized by a variant of SEQ ID NO: 003, wherein the D (aspartic acid) residue in the active center of the protein has been exchanged for, for example, an A (see SEQ ID NO: 370), to generate a class of non-covalent HaloTag mutants that bind substrates but do not covalently link them, which provides applications in advanced (super-resolution) microscopy techniques. The inventors have been able to determine that these non-covalent HaloTag systems benefit from the invention described herein in any of its aspects.
[0076] In certain embodiments, the internal cpHalo linker consists of 12 to 20 amino acids. In certain embodiments, the internal cpHalo linker consists of ~15 amino acids.
[0077] An alternative to this first aspect of the invention provides an isolated second partial effector sequence described by any one of SEQ ID NO: 6 to SEQ ID NO: 343, which can be combined with the first partial effector sequence alone as a kit, or in combination, as part of two separate fusion proteins comprising the first and second partial effector sequences, respectively, or as part of a larger fusion protein.
[0078] In certain embodiments, the second partial effector sequence consists of a sequence selected from the group consisting of SEQ ID NOs: 6–10, 12, 13, 15–26, 28, 30, 31, 34, 40, 46, 48, 50, 75, 85, 98, 103, 112, 154, 156, 160, 167, 173, 175, 182, 189, 195, 212, 214, 233, and 337–342, wherein optionally, the second partial effector sequence may include one or two amino acid variations independently selected from deletions, insertions, and variable substitutions.
[0079] In certain embodiments, the second portion effector sequence consists of a sequence selected from the group consisting of
[0080] -WREEVRKAFKLFRQ (SEQ ID NO: 25)
[0081] -WREEVRKAFKLFRS (SEQ ID NO: 24)
[0082] -WRETFQLFRT (SEQ ID NO: 26)
[0083] -WREMFRLFRTGRVQ (SEQ ID NO: 27)
[0084] -WREMFQAFRT (SEQ ID NO: 28)
[0085] -WREMFRLFRTGQRS (SEQ ID NO: 29)
[0086] -WREMFQLFRT (SEQ ID NO: 30)
[0087] -WREMFRLFRT (SEQ ID NO: 20)
[0088] -SKRDWREMFRLFRT (SEQ ID NO: 31)
[0089] -RVMSWREMFRLFRT (SEQ ID NO: 32)
[0090] -RMWSWREMFRLFRT (SEQ ID NO: 33)
[0091] -WKRDWREMFRLFRT (SEQ ID NO: 34)
[0092] -RQWTWREMFRLFRT (SEQ ID NO: 35)
[0093] -RGWTWREMFRLFRT (SEQ ID NO: 36)
[0094] -RMWTWREMFRLFRT (SEQ ID NO: 37)
[0095] -RQWSWREMFRLFRT (SEQ ID NO: 38)
[0096] -RGWSWREMFRLFRT (SEQ ID NO: 39)
[0097] Optionally, the second partial effector sequence may include one or two amino acid variations independently selected from deletion, insertion, and variable substitution.
[0098] In certain embodiments, the second partial effector sequence consists of WREEVRKAFKLFRQ (SEQ ID NO: 25). In certain embodiments, the second partial effector sequence consists of WREEVRKAFKLFRS (SEQ ID NO: 24). In certain embodiments, the second partial effector sequence consists of WRETFQ LFRT (SEQ ID NO: 26). In certain embodiments, the second partial effector sequence consists of WREMFRLFRTGRVQ (SEQ ID NO: 27). In certain embodiments, the second partial effector sequence consists of WREMFQAFRT (SEQ ID NO: 28). In certain embodiments, the second partial effector sequence consists of WREMFRLFRTGQRS (SEQ ID NO: 29). In certain embodiments, the second partial effector sequence consists of WREMFQLFRT (SEQ ID NO: 30). In certain embodiments, the second partial effector sequence consists of WREMFRLFRT (SEQ ID NO: 20). In certain embodiments, the second partial effector sequence consists of SKRDWREMFRLFRT (SEQ ID NO: 31). In certain embodiments, the second partial effector sequence consists of RVMSWREMFRLFRT (SEQ ID NO:32). In certain embodiments, the second partial effector sequence consists of RMWSWREMFRLFRT (SEQ ID NO:33). In certain embodiments, the second partial effector sequence consists of WKRDWREMFRLFRT (SEQ ID NO:34). In certain embodiments, the second partial effector sequence consists of RQWTWREMFRLFRT (SEQ ID NO:35). In certain embodiments, the second partial effector sequence consists of RGWTWREMFRLFRT (SEQ ID NO:36). In certain embodiments, the second partial effector sequence consists of RMWTWREMFR LFRT (SEQ ID NO:37). In certain embodiments, the second partial effector sequence consists of RQWSWREMFRLFRT (SEQ ID NO:38). In certain embodiments, the second partial effector sequence consists of RGWSWREMFRLFRT (SEQ ID NO:39).
[0099] In certain embodiments, the second portion effector sequence consists of a sequence selected from the group consisting of
[0100] -SKRDAREMFQAFRT(SEQ ID NO:6);
[0101] -ARLFFQLFRT (SEQ ID NO: 7);
[0102] -YFQGARETFQAFRT(SEQ ID NO:8);
[0103] -WIETFKLYRE (SEQ ID NO: 9);
[0104] -RKKEARETFQAFRT(SEQ ID NO:10);
[0105] -VIETFKLFRS (SEQ ID NO: 11);
[0106] -AREMFQLFRT (SEQ ID NO: 12);
[0107] -AYKIFQLFRT (SEQ ID NO: 13);
[0108] -AIRMFQLFRT (SEQ ID NO: 14);
[0109] -WGDEARETFQAFRT(SEQ ID NO:15);
[0110] -AREMFQAFRT (SEQ ID NO: 16);
[0111] -WKDEVIDAFRKFRE(SEQ ID NO:17);
[0112] -ERETWQAFRT (SEQ ID NO: 18);
[0113] -WREEVRKTFKLFRQ(SEQ ID NO:19);
[0114] -WREMFRLFRT (SEQ ID NO: 20);
[0115] -YFQGAREMFQAFRT(SEQ ID NO:21);
[0116] -WKEEVIKAFKLFRD(SEQ ID NO:22);
[0117] -AVNMFQLFRT (SEQ ID NO: 23);
[0118] -WREEVRKAFKLFRS(SEQ ID NO:24);
[0119] Optionally, the second partial effector sequence may include one or two amino acid variations independently selected from deletion, insertion, and variable substitution.
[0120] In certain embodiments, the second partial effector sequence consists of SKRDAREMFQAFRT (SEQ ID NO: 6). In certain embodiments, the second partial effector sequence consists of ARLFFQLFRT (SEQ ID NO: 7). In certain embodiments, the second partial effector sequence consists of YFQGARETFQAFRT (SEQ ID NO: 8). In certain embodiments, the second partial effector sequence consists of WIETFKLYRE (SEQ ID NO: 9). In certain embodiments, the second partial effector sequence consists of RKKEARETFQAFRT (SEQ ID NO: 10). In certain embodiments, the second partial effector sequence consists of VIETFKLFRS (SEQ ID NO: 11). In certain embodiments, the second partial effector sequence consists of AREMFQLFRT (SEQ ID NO: 12). In certain embodiments, the second partial effector sequence consists of AYKIFQLFRT (SEQ ID NO: 13). In certain embodiments, the second partial effector sequence consists of AIRMFQLFRT (SEQ ID NO: 14). In certain embodiments, the second partial effector sequence consists of WGDEARETFQAFRT (SEQ ID NO: 15). In certain embodiments, the second partial effector sequence consists of AREMFQAFRT (SEQ ID NO: 16). In certain embodiments, the second partial effector sequence consists of WKDEVIDAFRKFRE (SEQ ID NO: 17). In certain embodiments, the second partial effector sequence consists of ERETWQAFRT (SEQ ID NO: 18). In certain embodiments, the second partial effector sequence consists of WREEVRKTFKLFRQ (SEQ ID NO: 19). In certain embodiments, the second partial effector sequence consists of WREMFRLFRT (SEQ ID NO: 20). In certain embodiments, the second partial effector sequence consists of YFQGAREMFQAFRT (SEQ ID NO: 21). In certain embodiments, the second partial effector sequence consists of WKEEVIKAFKLFRD (SEQ ID NO: 22). In certain embodiments, the second partial effector sequence consists of AVNMFQLFRT (SEQ ID NO: 23). In certain embodiments, the second partial effector sequence consists of WREEVRKAFKLFRS (SEQ ID NO: 24).
[0121] The first partial effector sequences as given in SEQ ID NOs: 367 and 371 may be used together with the improved second partial effector peptides as claimed in this first aspect of the invention, such as any of the cpHaloΔ variants claimed in the second aspect of the invention described below:
[0122] Improved cpHaloΔ with higher stability and / or activity:
[0123] A second aspect of the present invention relates to a modular polypeptide complex comprising a first part effector sequence and a second part effector sequence (or consisting essentially of the same), characterized in that
[0124] a The N-terminal first effector sequence portion comprises at least one N-effector amino acid substitution; and / or
[0125] b. The C-terminal first effector sequence portion comprises at least one C-effector amino acid substitution; and / or
[0126] c. The internal cpHalo linker is characterized by a sequence selected from SEQ ID NOs: 344 to 366.
[0127] The first and second partial effector sequences together comprise an active, circularly arranged, self-labeling protein and enable covalent attachment of the haloalkane moiety to the first partial effector sequence when brought into close proximity with one another.
[0128] The second effector sequence consists of a sequence selected from the following
[0129] - SEQ ID NO: 004 and SEQ ID NO: 005 (original cpHalo peptide), or
[0130] - Any one of SEQ ID NO: 6 to SEQ ID NO: 343, or a subset thereof, as described according to the first aspect of the present invention.
[0131] In a specific embodiment, the inventors contemplate that the second partial effector sequence may include one or two amino acid variations independently selected from deletion, insertion, and variable substitution. In a more specific embodiment, the second partial effector sequence is identical to the sequence disclosed in SEQ ID NO: 004 to 343.
[0132] The first partial effector sequence is a variant of SEQ ID NO: 367 and 371, in other words, different from these known sequences, and consists of
[0133] o an N-terminal first effector sequence portion characterized by SEQ ID NO: 002 or a sequence that is at least (≥) 90% identical (particularly ≥93%, 95%, 97% or ≥98% identical) to SEQ ID NO: 002,
[0134] o a C-terminal first effector sequence portion characterized by SEQ ID NO: 003, or a sequence at least (≥) 90% identical (particularly ≥ 93%, 95%, 97% or ≥ 98% identical) to SEQ ID NO: 003, or a non-covalent variant of SEQ ID NO: 003, such as embodied by SEQ ID NO: 370;
[0135] o An internal cpHalo linker consisting of 10 to 35 amino acids, wherein the internal cpHalo linker connects the C-terminus of the N-terminal first effector sequence portion to the N-terminus of the C-terminal first effector sequence portion.
[0136] In certain embodiments, the internal cpHalo linker consists of 20 to 26 amino acids. In certain embodiments, the internal cpHalo linker consists of 22 to 24 amino acids.
[0137] The first partial effector sequence as given in SEQ ID NO: 367 and 371 is not claimed as part of this second aspect of the invention, which relates to the improved cpHaloΔ mutation.
[0138] In certain embodiments, only the N-terminal portion of the first effector sequence comprises at least one N-effector amino acid substitution. In certain embodiments, only the C-terminal portion of the first effector sequence comprises at least one C-effector amino acid substitution. In certain embodiments, both the N-terminal and C-terminal first effector sequences are as disclosed, and only the internal cpHalo linker differs, characterized by the sequences newly disclosed herein and listed below.
[0139] In certain embodiments, the N-terminal first effector sequence portion comprises at least one N-effector amino acid substitution and the C-terminal first effector sequence portion comprises at least one C-effector amino acid substitution. In certain embodiments, the N-terminal first effector sequence portion comprises at least one N-effector amino acid substitution and the internal cpHalo linker is characterized by certain sequences listed below. In certain embodiments, the C-terminal first effector sequence portion comprises at least one C-effector amino acid substitution and the internal cpHalo linker is characterized by certain sequences listed below.
[0140] In certain embodiments, the N-terminal first effector sequence portion comprises at least one N-effector amino acid substitution and the C-terminal first effector sequence portion comprises at least one C-effector amino acid substitution and the internal cpHalo linker is characterized by certain sequences listed below.
[0141] N effector amino acid substitutions (mutations) are selected from the group consisting of
[0142] οN217D;
[0143] οF205W;
[0144] οV184E;
[0145] οV197K;
[0146] οR254W.
[0147] In certain embodiments, the N-effect amino acid substitution is N217D. In certain embodiments, the N-effect amino acid substitution is F205W. In certain embodiments, the N-effect amino acid substitution is V184E. In certain embodiments, the N-effect amino acid substitution is V197K. In certain embodiments, the N-effect amino acid substitution is R254W.
[0148] C effector amino acid substitutions (mutations) are selected from the group consisting of
[0149] οF80T;
[0150] οN119H;
[0151] οK117R;
[0152] οR30P;
[0153] οG7D;
[0154] οE20S.
[0155] In certain embodiments, the C effector amino acid substitution is F80T. In certain embodiments, the C effector amino acid substitution is N119H. In certain embodiments, the C effector amino acid substitution is K117R. In certain embodiments, the C effector amino acid substitution is R30P. In certain embodiments, the C effector amino acid substitution is G7D. In certain embodiments, the C effector amino acid substitution is E20S.
[0156] The numbering of the N-effector and C-effector amino acid substitutions refers to the sequence position numbering of the original Halo7 sequence with GenBank-ID: AQS79242.1 (the original halo7 sequence without rearrangements).
[0157] With respect to SEQ ID NOs: 002 and 003, the mutations given in the preceding paragraphs translate to:
[0158] -N217D: position 62 of SEQ ID NO: 002;
[0159] -F205W: position 50 of SEQ ID NO: 002;
[0160] - V184E: position 29 of SEQ ID NO: 002;
[0161] - V197K: position 42 of SEQ ID NO: 002;
[0162] -R254W: position 99 of SEQ ID NO: 002;
[0163] -F80T: position 77 of SEQ ID NO: 003;
[0164] -N119H: position 116 of SEQ ID NO: 003;
[0165] -K117R: position 114 of SEQ ID NO: 003;
[0166] -R30P: position 27 of SEQ ID NO: 003;
[0167] -G7D: position 4 of SEQ ID NO: 003;
[0168] -E20S: Position 17 of SEQ ID NO: 003.
[0169] In certain embodiments, the internal cpHalo linker is characterized by a sequence selected from the group consisting of
[0170] οRSDDPRKTQTIASKISRDLNGS(SEQ ID NO:344);
[0171] οKGGTKRDADKAVRDTLLSLNGQ(SEQ ID NO:345);
[0172] οGGAPRDEALKKIEKAKRDTGDQ (SEQ ID NO:346);
[0173] οKSKYDRDQILKIIAELEKKTGGS (SEQ ID NO:347);
[0174] οQSKYPPEWLEKVIRELLKRKNGR(SEQ ID NO:348);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0175] <h2 style=";text-align:left;direction:ltr"> οKSKYDKRQIRDIADKIAKDNNHQ(SEQ ID NO:349);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0176] <h2 style=";text-align:left;direction:ltr"> οGADDKTKIEKILEEIKRRWQGR(SEQ ID NO:350);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0177] <h2 style=";text-align:left;direction:ltr"> οGTSDPRNQEIAKKLARDASTVP(SEQ ID NO:351);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0178] <h2 style=";text-align:left;direction:ltr"> οNGADKEQIDRAIEKAKRDLNNQ(SEQ ID NO:352);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0179] <h2 style=";text-align:left;direction:ltr"> οKGASDRDEAKKLADDIRKKKGDQ(SEQ ID NO:353);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0180] <h2 style=";text-align:left;direction:ltr"> οNSNGHRDELEKILQTIRKQNNDI(SEQ ID NO:354);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0181] <h2 style=";text-align:left;direction:ltr"> οLKDERQRDKALEIADRADKYPTS(SEQ ID NO:355);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0182] <h2 style=";text-align:left;direction:ltr"> οKGAEDAKERLERGDIEKKKKEQP(SEQ ID NO:356);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0183] <h2 style=";text-align:left;direction:ltr"> οKGAEDAKERLERGDLDRWSQEWR(SEQ ID NO:357);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0184] <h2 style=";text-align:left;direction:ltr"> οKGAEDAKERLERRDLDKINSRNS(SEQ ID NO:358);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0185] <h2 style=";text-align:left;direction:ltr"> οKGAEDAKERLEKGYIDDKASKQQ(SEQ ID NO:359);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0186] <h2 style=";text-align:left;direction:ltr"> οKGAEDAKERLEKGELEKKWKDHP(SEQ ID NO:360);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0187] <h2 style=";text-align:left;direction:ltr"> οKGAEDAKERLERGEMEKAVKHGS(SEQ ID NO:361);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0188] <h2 style=";text-align:left;direction:ltr"> οKGAEDAKERLERDELTRDIKTYPY(SEQ ID NO:362);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0189] <h2 style=";text-align:left;direction:ltr"> οKGAEDAKERLEQGMLEEIKKKYPE(SEQ ID NO:363);<h2 style=";text-align:left;direction:ltr">
[0190] οKGAEDAKERLERDELTKIAKNLGG (SEQ ID NO:364);
[0191] οKGAEDAKERLEKNALDKIAKSKGD(SEQ ID NO:365);
[0192] οKGAEDAKERLEKNDETLKKAKDKP(SEQ ID NO:366);
[0193] Optionally, the internal cpHalo linker sequence may include one, two, or three amino acid variations independently selected from deletions, insertions, and variable substitutions.
[0194] In certain embodiments, the internal cpHalo linker is characterized by the sequence RSDDPRKTQTIASKISRDLNGS (SEQ ID NO: 344). In certain embodiments, the internal cpHalo linker is characterized by the sequence KGGTKRDADKAVRDTLLSLNGQ (SEQ ID NO: 345). In certain embodiments, the internal cpHalo linker is characterized by the sequence GGAPRDEALKKIEKAKRDTGDQ (SEQ ID NO: 346). In certain embodiments, the internal cpHalo linker is characterized by the sequence KSKYDRDQILKIIAELEKKTGGS (SEQ ID NO: 347). In certain embodiments, the internal cpHalo linker is characterized by the sequence QSKYPPEWLEKVIRELLKRKNGR (SEQ ID NO: 348). In certain embodiments, the internal cpHalo linker is characterized by the sequence KSKYDKRQIRDIADKIAKDNNHQ (SEQ ID NO: 349). In certain embodiments, the internal cpHalo linker is characterized by the sequence GADDKTKIEKIL EEIKRRWQGR (SEQ ID NO: 350). In certain embodiments, the internal cpHalo linker is characterized by the sequence GTSDPRNQEIAKKLARDASTVP (SEQ ID NO: 351). In certain embodiments, the internal cpHalo linker is characterized by the sequence NGADKEQIDRAIEKAKRDLNNQ (SEQ ID NO: 352). In certain embodiments, the internal cpHalo linker is characterized by the sequence KGASDRDEAKKLADDIRKKKGDQ (SEQ ID NO: 353). In certain embodiments, the internal cpHalo linker is characterized by the sequence NSNGHRDELEKILQTIRKQNNDI (SEQ ID NO: 354). In certain embodiments, the internal cpHalo linker is characterized by the sequence LKDERQRDKALEIADRADKYPTS (SEQ ID NO: 355). In certain embodiments, the internal cpHalo linker is characterized by the sequence KGAEDAKERLERGDIEKKKKEQP (SEQ ID NO: 356). In certain embodiments, the internal cpHalo linker is characterized by the sequence KGAEDAKERLERGDLDRWSQEWR (SEQ ID NO: 357). In certain embodiments, the internal cpHalo linker is characterized by the sequence KGAEDAKERLERRDLDKINSRNS (SEQ ID NO: 358).In certain embodiments, the internal cpHalo linker is characterized by the sequence KGAEDAKERLEKGYIDDKASKQQ (SEQ ID NO: 359). In certain embodiments, the internal cpHalo linker is characterized by the sequence KGAEDAKERLEKGELEKKWK DHP (SEQ ID NO: 360). In certain embodiments, the internal cpHalo linker is characterized by the sequence KGAEDAKERLERGEMEKAVKHGS (SEQ ID NO: 361). In certain embodiments, the internal cpHalo linker is characterized by the sequence KGAEDAKERLERDELTRDIKTYPY (SEQ ID NO: 362). In certain embodiments, the internal cpHalo linker is characterized by the sequence KGAEDAKERLEQGMLEEIKKKYPE (SEQ ID NO: 363). In certain embodiments, the internal cpHalo linker is characterized by the sequence KGAEDAKERLERDELTKIAKNLGG (SEQ ID NO: 364). In certain embodiments, the internal cpHalo linker is characterized by the sequence KGAEDAKERLEKNALDKIAKSKGD (SEQ ID NO: 365). In certain embodiments, the internal cpHalo linker is characterized by the sequence KGAEDAKERLEKNDETLKKAKDKP (SEQ ID NO: 366).
[0195] In certain embodiments, the first partial effector sequence comprises an N-effector amino acid substitution and / or a C-effector amino acid substitution and / or an internal cpHalo linker sequence selected from the group consisting of:
[0196] -E20S and V184E;
[0197] -E20S, N119H, and V184E;
[0198] -E20S, N119H, V184E, and V197K;
[0199] - E20S, N119H, and V184E, and linker sequence KSKYDRDQILKIIAELEKKTGGS (SEQ ID NO: 347);
[0200] -E20S, N119H, V184E, and V197K, and linker sequence KSKYDRDQILKIIAELEKKTGGS (SEQ ID NO: 347).
[0201] It will be understood that the variants of the previously disclosed HaloTag system provided herein also encompass any analogous variant of the non-covalent HaloTag system in which the nucleophilic D residue has been substituted, for example, by A (see SEQ ID NO: 371).
[0202] In certain embodiments, the first portion of the effector sequence includes E20S, N119H, V184E, and V197K, and the linker sequence KSKYDRDQILKII AELEKKTGGS (SEQ ID NO: 347).
[0203] Combination of peptides with improved labeling speed and improved cpHaloΔ:
[0204] The third aspect of the present invention relates to a modular polypeptide complex, comprising the first partial effector sequence as described in the second aspect, and the second partial effector sequence as described in the first aspect.
[0205] The second effector sequence consists of a sequence selected from the following
[0206] - SEQ ID NO: 004 and SEQ ID NO: 005 (original cpHalo peptide), or
[0207] - Any one of SEQ ID NO: 6 to SEQ ID NO: 343, or a subset thereof, as described according to the first aspect of the present invention.
[0208] In certain embodiments, the second partial effector sequence consists of a sequence selected from the group consisting of any one of SEQ ID NO: 6 to SEQ ID NO: 343, and the N-terminal first effector sequence portion includes at least one N-effector amino acid substitution selected from the group consisting of N217D, F205W, V184E, V197K, and R254W.
[0209] In certain embodiments, the second partial effector sequence consists of a sequence selected from the group consisting of any one of SEQ ID NO: 6 to SEQ ID NO: 343, and the C-terminal first effector sequence portion comprises at least one C-effector amino acid substitution selected from the group consisting of F80T, N119H, K117R, R30P, G7D, and E20S.
[0210] In certain embodiments, the second portion effector sequence consists of a sequence selected from the group consisting of any one of SEQ ID NO: 6 to SEQ ID NO: 343, and the internal cpHalo linker is characterized by a sequence selected from the group consisting of SEQ ID NO: 344 to SEQ ID NO: 366.
[0211] In certain embodiments, the second partial effector sequence consists of a sequence selected from the group consisting of any one of SEQ ID NO: 6 to SEQ ID NO: 343, and
[0212] the N-terminal first effector sequence portion comprises at least one N-effector amino acid substitution selected from the group consisting of N217D, F205W, V184E, V197K, R254W; and
[0213] • The C-terminal first effector sequence portion comprises at least one C-effector amino acid substitution selected from the group consisting of F80T, N119H, K117R, R30P, G7D, E20S.
[0214] In certain embodiments, the second partial effector sequence consists of a sequence selected from the group consisting of any one of SEQ ID NO: 6 to SEQ ID NO: 343, and
[0215] the N-terminal first effector sequence portion comprises at least one N-effector amino acid substitution selected from the group consisting of N217D, F205W, V184E, V197K, R254W; and
[0216] • The internal cpHalo linker is characterized by a sequence selected from the group consisting of SEQ ID NO: 344 to SEQ ID NO: 366.
[0217] In certain embodiments, the second partial effector sequence consists of a sequence selected from the group consisting of any one of SEQ ID NO: 6 to SEQ ID NO: 343, and
[0218] the C-terminal first effector sequence portion comprises at least one C-effector amino acid substitution selected from the group consisting of F80T, N119H, K117R, R30P, G7D, E20S, and
[0219] • The internal cpHalo linker is characterized by a sequence selected from the group consisting of SEQ ID NO: 344 to SEQ ID NO: 366.
[0220] In certain embodiments, the second partial effector sequence consists of a sequence selected from the group consisting of any one of SEQ ID NO: 6 to SEQ ID NO: 343, and
[0221] the N-terminal first effector sequence portion comprises at least one N-effector amino acid substitution selected from the group consisting of N217D, F205W, V184E, V197K, R254W; and
[0222] the C-terminal first effector sequence portion comprises at least one C-effector amino acid substitution selected from the group consisting of F80T, N119H, K117R, R30P, G7D, E20S, and
[0223] • The internal cpHalo linker is characterized by a sequence selected from the group consisting of SEQ ID NO: 344 to SEQ ID NO: 366.
[0224] In certain embodiments, the second partial effector sequence consists of a sequence selected from the group consisting of any one of SEQ ID NOs: 006-010, 012, 013, 015-026, 028, 030, 031, 034, 040, 046, 048, 050, 075, 085, 098, 103, 112, 154, 156, 160, 167, 173, 175, 182, 189, 195, 212, 214, 233, and 337-342, and the first partial effector sequence comprises at least one feature selected from the group consisting of an N-effector amino acid substitution, a C-effector amino acid substitution, a novel internal cpHalo linker as provided herein (particularly a sequence selected from SEQ ID NOs: 344 to 366).
[0225] In certain embodiments, the second partial effector sequence consists of a sequence selected from the group consisting of any one of SEQ ID NOs: 020 or 024-039, and the first partial effector sequence comprises at least one feature selected from N-effector amino acid substitutions, C-effector amino acid substitutions, and a novel internal cpHalo linker as provided herein (particularly a sequence selected from SEQ ID NOs: 344 to 366).
[0226] Optionally, the second portion effector sequences described in this section can include one or two amino acid variations independently selected from deletions, insertions, and variable substitutions.
[0227] Replacement rules:
[0228] Optionally, variable substitutions may be selected from substitutions according to the following rules:
[0229] a. Glycine (G) and alanine (A) are interchangeable; valine (V), leucine (L) and isoleucine (I) are interchangeable, and A and V are interchangeable;
[0230] b. Tryptophan (W), phenylalanine (F), and tyrosine (Y) are interchangeable;
[0231] c. Serine (S) and threonine (T) are interchangeable;
[0232] d. Aspartic acid (D) and glutamic acid (E) are interchangeable;
[0233] e. Asparagine (N) and glutamine (Q) are interchangeable; N and S are interchangeable; N and D are interchangeable; E and Q are interchangeable;
[0234] f. Methionine (M) and Q are interchangeable;
[0235] g. Cysteine (C), A and S are interchangeable;
[0236] h. Proline (P), G and A are interchangeable;
[0237] i. Arginine (R), lysine (K) and Q are interchangeable;
[0238] j. Histidine (H) and Y are interchangeable, and H and N are interchangeable;
[0239] kL and M are interchangeable, I and M are interchangeable, V and M are interchangeable;
[0240] lE and K are interchangeable;
[0241] mA and C are interchangeable, A and G are interchangeable; A and T are interchangeable, A and V are interchangeable;
[0242] nR and N are interchangeable, R and E are interchangeable, R and H are interchangeable;
[0243] oN and Q are interchangeable, N and E are interchangeable, N and G are interchangeable, N and K are interchangeable, N and T are interchangeable;
[0244] pD and Q are interchangeable, D and S are interchangeable;
[0245] qQ and H are interchangeable, Q and M are interchangeable, Q and S are interchangeable;
[0246] rE and H are interchangeable, E and S are interchangeable;
[0247] sG and S are interchangeable;
[0248] tF and I are interchangeable, F and L are interchangeable, F and M are interchangeable;
[0249] uK and S are interchangeable;
[0250] vT and V are interchangeable;
[0251] In certain embodiments, variable substitutions are selected from substitutions according to rules a to l (rather than according to m to v).
[0252] Spontaneous complementarity
[0253] In certain embodiments of the modular peptide complexes according to the first, second, or third aspects of the present invention, the first and second partial effector sequences are capable of forming an active, circularly arranged, self-labeling protein without the need for a sensor module polypeptide. These self-complementary first and second partial effector sequences need only be expressed in the same cell or expression system to construct the protein complex, and if a corresponding substrate is present, the haloalkane moiety is attached to the first partial effector sequence.
[0254] In certain embodiments of this subaspect, the second partial effector sequence is selected from the group consisting of SEQ ID NO: 20 and SEQ ID NOs: 24-39. The inventors have demonstrated that these sequences are capable of complementing cpHaloΔ (a first partial effector sequence; one previously disclosed or a cpHaloΔ disclosed herein) in the absence of direct linkage, such as established by a sensor polypeptide. The inventors do not exclude the possibility that any of the sequences in SEQ ID NOs: 006-343 could perform this function; however, to date, this complementation capability in the absence of a linker has only been demonstrated for SEQ ID NOs: 20 and 24-39. The second partial effector sequences disclosed in the original cpHaloΔ disclosure (SEQ ID NOs: 004 and 005) were unable to perform this function.
[0255] Sensor Module Peptide:
[0256] In certain embodiments of the modular peptide complex according to the first, second or third aspect of the present invention, the first partial effector sequence and the second partial effector sequence are linked via a sensor module polypeptide.
[0257] In certain embodiments, the sensor moiety polypeptide is a single sensor polypeptide capable of undergoing a conformational change from a first conformation to a second conformation in response to the presence or concentration of an analyte compound.
[0258] In certain embodiments, the sensor moiety polypeptide is a single sensor polypeptide that is capable of undergoing a conformational change from a first conformation to a second conformation in response to the presence or concentration of an external stimulus, such as light radiation.
[0259] In the first conformation, the first and second partial effector sequences are in close proximity (resulting in the first and second partial effector sequences forming an active entity).
[0260] In the second conformation, the first and second partial effector sequences are not in close proximity (resulting in the first and second partial effector sequences constituting an inactive (non-self-labeling) entity).
[0261] If the first and second effector portion sequences are unable to form an active, circularly arranged, self-labeling protein in the absence of an agent that induces their spatial proximity, they must be brought into close proximity to form an active protein. The sensor module polypeptide is coupled to the first and second effector portion sequences, and the sensor module polypeptide—upon sensing a stimulus—brings the first effector portion sequence and the second effector portion sequence into close proximity such that they can interact and form an active, circularly arranged, self-labeling protein.
[0262] Optionally,
[0263] One of the first and second effector sequences can be linked to the C-terminus of the sensor module sequence (single sensor polypeptide); and
[0264] The other of the first effector sequence and the second effector sequence can be linked to the N-terminus of the sensor module sequence (single sensor polypeptide)
[0265] • Alternatively, the first partial effector sequence and / or the second partial effector sequence are inserted into a single sensor polypeptide.
[0266] In certain embodiments, the sensor module polypeptide is a sensor polypeptide pair comprising a first sensor polypeptide and a second sensor polypeptide, wherein the first sensor polypeptide is covalently linked to a first effector portion sequence via a peptide bond and the second sensor polypeptide is covalently linked to a second effector portion sequence. The first sensor polypeptide and the second sensor polypeptide are capable of specific molecular interaction (protein-protein binding), optionally wherein the specific molecular interaction can be dependent on analyte concentration, and the first and second sensor polypeptides are part of separate polypeptide chains.
[0267] Using sensor module polypeptides, it is also possible to detect the co-localization of two polypeptides or peptides within or outside a cell, for example, when expressed in a cell-free expression system. The two (poly)peptides of interest constitute a sensor polypeptide pair, and if they co-localize, the first and second effector sequences are in close proximity and thus can form an active protein that can attach the HaloTag substrate to the first effector sequence.
[0268] Non-limiting examples of sensor polypeptides are
[0269] a) Calmodulin. The protein calmodulin is an example of a single sensor polypeptide. Calmodulin is flanked by a first partial effector sequence and a second partial effector sequence (one at the N-terminus and one at the C-terminus). If calmodulin senses calcium, it undergoes a conformational change, bringing the first and second partial effector sequences into close proximity and enabling interaction.
[0270] b) FKBP / FRP. The FKBP / FRP polypeptide is an example of a sensor polypeptide pair. One of the proteins, FKBP and FRB, is coupled to a first effector sequence, while the other is coupled to a second effector sequence. When rapamycin is added to the expression system, FKBP and FRB form a protein complex, bringing the first and second effector sequences into close proximity and enabling interaction.
[0271] c) Glutamate intensity sensor (iGlusnFR). iGluSnFR is derived from the bacterial periplasmic glutamate-binding protein Glt1. The second effector sequence is linked to the N-terminus of iGlusSnFR, and the first effector sequence is linked to the C-terminus. When glutamate is added, iGlusSnFR undergoes a conformational change, resulting in close proximity and interaction between the first and second effector sequences.
[0272] Further entities:
[0273] Another aspect of the present invention relates to a nucleic acid sequence, or multiple nucleic acid sequences, encoding a modular polypeptide complex according to any of the aforementioned aspects. In certain embodiments, the nucleic acid sequence is integrated into the cellular genome. If the modular polypeptide complex comprises a single sensor polypeptide linked to a first and second part effector component, the system is encoded by a single nucleic acid sequence; however, if the sensor is a two-part sensor, it may be necessary or convenient to use two separate nucleic acid sequences.
[0274] Another aspect of the present invention relates to a nucleic acid expression system comprising a nucleic acid sequence or a plurality of nucleic acid sequences according to the aforementioned aspects, each nucleic acid sequence being under the control of a promoter sequence.
[0275] In certain embodiments, the nucleic acid expression system is a plasmid.
[0276] In certain embodiments of the nucleic acid sequence or nucleic acid expression system, the sequence encoding the modular polypeptide complex comprises a promoter 5' to each open reading frame encoding the modular polypeptide complex. In certain embodiments of the nucleic acid sequence or nucleic acid expression system, each open reading frame encoding the modular polypeptide complex is adjacent to a multiple cloning site (in the 3' or 5' direction) that allows for in-frame insertion of the coding sequence of interest and the first and / or second partial effector sequence.
[0277] Another aspect of the present invention relates to a cell comprising the nucleic acid expression system according to the preceding aspects, in particular wherein the promoter is operable in the cell.
[0278] Another aspect of the present invention relates to a non-human transgenic animal or plant comprising the nucleic acid sequence according to the aforementioned aspect or the nucleic acid expression system according to the aforementioned aspect.
[0279] Another aspect of the present invention relates to a kit comprising a nucleic acid sequence or nucleic acid expression system according to the aforementioned aspects, and a HaloTag substrate. In certain embodiments, the HaloTag substrate comprises a detection moiety. In certain specific embodiments, the HaloTag substrate comprises a fluorescent dye moiety as the detection moiety.
[0280] In certain embodiments, the HaloTag substrate comprises a purification tag. In certain specific embodiments, the purification tag is selected from the group consisting of a T7 tag, a calmodulin binding peptide, a FLAG epitope, a GST tag, an HA tag, a polyhistidine tag, a Myc tag, biotin, and a Strep tag. Similarly, reactive groups, such as those known in the context of "click" chemistry, may be part of the HaloTag substrate. Non-limiting examples thereof are alkynyl (particularly ethynyl) or azide moieties, trans-cyclooctene moieties, tetrazine moieties, N-hydroxysuccinimide (NHS) moieties, maleimide moieties.
[0281] Another aspect of the present invention relates to a nucleic acid sequence comprising a peptide-encoding sequence element encoding a peptide sequence selected from the group consisting of SEQ ID NOs: 006–110, 112–194, and 196–343. The peptide-encoding sequence element is flanked by cloning sites located directly 5′ or 3′ of the peptide-encoding sequence element, allowing for in-frame insertion of a coding sequence of interest into the peptide-encoding portion. This aspect relates to a nucleic acid sequence capable of accepting the insertion of a protein or peptide of interest cloned in-frame with a second effector sequence (meaning a peptide that forms a circular, self-labeling protein when bound to a cpHaloΔ protein). This fusion polypeptide of the protein or peptide of interest and the second effector sequence can be detected by adding the cpHaloΔ protein and a HaloTag substrate comprising a detection moiety. In this case, the first and second effector sequences form a circular, self-labeling protein that is then labeled with the HaloTag substrate. This labeling can (and subsequently also is) be read out by detecting the detection moiety attached to the HaloTag substrate.
[0282] In certain embodiments of this aspect, the peptide sequence encoded by the peptide coding sequence element is selected from the group of SEQ ID NOs: 6-10, 12, 13, 15-26, 28, 30, 31, 34, 40, 46, 48, 50, 75, 85, 98, 103, 112, 154, 156, 160, 167, 173, 175, 182, 189, 212, 214, 233, and 337-342.
[0283] In certain embodiments of this aspect, the peptide sequence encoded by the peptide encoding sequence element is selected from the group of SEQ ID NOs: 6-24.
[0284] In certain embodiments of this aspect, the sequence encoded by the peptide encoding sequence element is selected from the group consisting of SEQ ID NO: 20 and SEQ ID NOs: 24-39.
[0285] Methods for detecting molecular interaction events:
[0286] Another aspect of the present invention relates to a method for detecting a molecular interaction event, comprising the steps of:
[0287] a) providing an expression system capable of expressing the modular polypeptide complex according to any one of the first, second and third aspects;
[0288] b) adding a haloalkane moiety (HaloTag substrate) to the expression system under conditions that result in expression of the modular polypeptide, wherein the haloalkane moiety is coupled to the detection moiety;
[0289] c) In a detection step, detecting the detection moiety coupled to the first portion effector sequence.
[0290] In certain embodiments, during the detecting step, a molecular interaction event between the first sensor polypeptide and the second sensor polypeptide is detected.
[0291] In certain embodiments, during the detecting step, an internal molecular interaction event (conformational change) of a single sensor polypeptide is detected.
[0292] In certain embodiments, during the detecting step, co-expression of a first partial effector sequence and a second partial effector sequence is detected, wherein the first partial effector sequence and the second partial effector sequence are capable of spontaneously interacting. This method can be applied to quantify the activity of a specific promoter that regulates expression of a peptide (second partial effector sequence) or cpHaloΔ (first partial effector sequence).
[0293] In certain embodiments, during the detecting step, the presence of the protein of interest is detected. The protein of interest can be fused to a first partial effector sequence. Alternatively, the protein of interest can be fused to a second partial effector sequence.
[0294] When using this method to detect the presence of a protein of interest, the first effector sequence and the second effector sequence must be able to interact spontaneously.
[0295] In certain embodiments, the expression system is a prokaryotic or eukaryotic cell, and the protein of interest is an endogenous protein of the cell.
[0296] Using this method, it is possible to determine, among other things, how much of the protein of interest is present in the cell or expression system, and where the protein of interest is located. This can be analyzed at a later stage, since the HaloTag substrate is covalently coupled to the first effector portion when the first and second effector portions interact.
[0297] Since the interaction (of the first and second part effector sequences) can be analyzed at a later time point, the method can also be performed in living cells or even in animals or plants. First, the expression and / or localization of the protein of interest can be recorded, and at a later time point, the recorded expression and / or localization can be analyzed. This analysis can be performed by light emission from a fluorescent dye that is part of the detection moiety coupled to the HaloTag substrate.
[0298] In certain embodiments, the nucleic acid sequence encoding the second partial effector sequence is inserted at the 3' end or the 5' end of the endogenous gene encoding the protein of interest.
[0299] In certain embodiments, the nucleic acid sequence encoding the second partial effector sequence is inserted by the CRISPR / Cas9 complex.
[0300] In certain embodiments, the first partial effector sequence is expressed and encoded by a plasmid.
[0301] In certain embodiments, the first portion of the effector sequence is introduced into the cell via a viral vector.
[0302] In certain embodiments, the second partial effector sequence is selected from the group consisting of SEQ ID NO: 20 and SEQ ID NOs: 24-39. The inventors have demonstrated that these sequences are capable of complementing cpHaloΔ (a first partial effector sequence; one previously disclosed or a cpHaloΔ disclosed herein) in the absence of direct linkage, such as established by a sensor polypeptide. The inventors do not exclude the possibility that any of the sequences in SEQ ID NOs: 006-343 could perform this function; however, to date, this complementation capability in the absence of a linker has only been demonstrated for SEQ ID NOs: 20 and 24-39. The second partial effector sequences disclosed in the original cpHaloΔ disclosure (SEQ ID NOs: 004 and 005) were unable to perform this function.
[0303] In certain embodiments, the detection moiety comprises a purification tag.
[0304] In certain embodiments, the detection moiety comprises a fluorophore.
[0305] If the expression and / or localization of an endogenous protein of interest (POI) is to be studied in living cells (by microscopy), this POI must be tagged with a living cell-compatible fluorescent tag on the genome (by genome engineering techniques such as CRISPR-Cas9). This tag can be a fluorescent protein such as GFP or a self-labeling protein such as HaloTag. The smaller the genome modification, the more efficient the genome engineering process. This allows the introduction of only small peptides in the split system, which, after complementation with the larger split part, can be used as fluorescent tags, greatly increasing throughput and the chances of success compared to the introduction of large tags.
[0306] In certain embodiments, the peptide (second part effector sequence) is inserted into the genome at the locus of the POI and the other split part (first part effector sequence) can be expressed from a vector or from another genomic locus. When there is no sensor module polypeptide, the split parts (first and second part effector sequences) must be characterized by very high affinity for each other so that they spontaneously complement each other and remain together as a stable complex. Even if the label is not directly coupled to the peptide (second part effector sequence) but to the cpHaloΔ part (first part effector sequence), the complex remains intact and the label will therefore remain on the POI. The advantage of the currently claimed split HaloTag system is that it allows the use of synthetic dyes instead of GFP, which is highly preferred for (super-resolution) microscopy for the detection of low-abundance proteins.
[0307] Wherever alternatives to a single separable feature, such as an isoform protein or an effector or sensor sequence, are listed herein as "embodiments," it is to be understood that such alternatives can be freely combined to form discrete embodiments of the invention disclosed herein. Thus, any alternative embodiment of a detectable label can be combined with any alternative embodiment of an isoform protein, and these combinations can be combined with any effector or sensor sequence mentioned herein.
[0308] This manual also includes the following items:
[0309] item:
[0310] 1. A modular polypeptide complex comprising
[0311] - The first part of the effector sequence consists of the following
[0312] o an N-terminal first effector sequence portion characterized by SEQ ID NO: 002, or a sequence at least (≥) 90% identical (particularly ≥93%, 95%, 97% or ≥98% identical) to SEQ ID NO: 002,
[0313] o a C-terminal first effector sequence portion characterized by SEQ ID NO: 003 or SEQ ID NO: 370, or a sequence that is at least (≥) 90% identical (particularly ≥93%, 95%, 97% or ≥98% identical) to SEQ ID NO: 003 or SEQ ID NO: 370;
[0314] o internal cpHalo linker, consisting of 10 to 35 amino acids,
[0315] Specifically, it is composed of 20 to 25 amino acids.
[0316] More particularly 22 to 24 amino acids,
[0317] wherein an internal cpHalo linker connects the C-terminus of the N-terminal first effector sequence portion to the N-terminus of the C-terminal first effector sequence portion;
[0318] - The second part of the effector sequence,
[0319] in
[0320] The first and second effector sequences together constitute an active, circularly arranged, self-labeling protein and, when brought into close proximity with one another, enable covalent attachment of the haloalkyl moiety, and
[0321] The modular polypeptide complex is characterized in that the second part effector sequence consists of a sequence selected from the group consisting of SEQ ID NO: 006-343 (any one of 338 variants).
[0322] 2. A modular polypeptide complex according to claim 1, wherein the first partial effector sequence is flanked at the N-terminus and / or C-terminus by a tetrapeptide selected from LKPG or EKKG (SEQ ID NO372, 373) at the N-terminus and PDYE, GDVE, PDSN or PDPQ (SEQ ID NO374–377) at the C-terminus.
[0323] 3. The modular polypeptide complex according to item 1 or 2, wherein the second partial effector sequence may comprise one or two amino acid variations relative to the sequence given in SEQ ID NO: 006–343, said variations being independently selected from deletions, insertions and variable substitutions.
[0324] 4. A modular polypeptide complex according to any of the preceding items, wherein the second portion effector sequence consists of a sequence selected from the group consisting of SEQ ID NO: 6–10, 12, 13, 15-26, 28, 30, 31, 34, 40, 46, 48, 50, 75, 85, 98, 103, 112, 154, 156, 160, 167, 173, 175, 182, 189, 195, 212, 214, 233 and 337–342.
[0325] 5. The modular polypeptide complex according to item 1 or 2, wherein the second partial effector sequence consists of a sequence selected from the group consisting of
[0326] -WREEVRKAFKLFRQ (SEQ ID NO: 25)
[0327] -WREEVRKAFKLFRS (SEQ ID NO: 24)
[0328] -WRETFQLFRT (SEQ ID NO: 26)
[0329] -WREMFRLFRTGRVQ (SEQ ID NO: 27)
[0330] -WREMFQAFRT (SEQ ID NO: 28)
[0331] -WREMFRLFRTGQRS (SEQ ID NO: 29)
[0332] -WREMFQLFRT (SEQ ID NO: 30)
[0333] -WREMFRLFRT (SEQ ID NO: 20)
[0334] -SKRDWREMFRLFRT (SEQ ID NO: 31)
[0335] -RVMSWREMFRLFRT (SEQ ID NO: 32)
[0336] -RMWSWREMFRLFRT (SEQ ID NO: 33)
[0337] -WKRDWREMFRLFRT (SEQ ID NO: 34)
[0338] -RQWTWREMFRLFRT (SEQ ID NO: 35)
[0339] -RGWTWREMFRLFRT (SEQ ID NO: 36)
[0340] -RMWTWREMFRLFRT (SEQ ID NO: 37)
[0341] -RQWSWREMFRLFRT (SEQ ID NO: 38)
[0342] -RGWSWREMFRLFRT (SEQ ID NO:39).
[0343] 6. The modular polypeptide complex according to item 1, wherein the second effector sequence consists of a sequence selected from the group consisting of
[0344] -SKRDAREMFQAFRT(SEQ ID NO:6);
[0345] -ARLFFQLFRT (SEQ ID NO: 7);
[0346] -YFQGARETFQAFRT(SEQ ID NO:8);
[0347] -WIETFKLYRE (SEQ ID NO: 9);
[0348] -RKKEARETFQAFRT(SEQ ID NO:10);
[0349] -VIETFKLFRS (SEQ ID NO: 11);
[0350] -AREMFQLFRT (SEQ ID NO: 12);
[0351] -AYKIFQLFRT (SEQ ID NO: 13);
[0352] -AIRMFQLFRT (SEQ ID NO: 14);
[0353] -WGDEARETFQAFRT(SEQ ID NO:15);
[0354] -AREMFQAFRT (SEQ ID NO: 16);
[0355] -WKDEVIDAFRKFRE(SEQ ID NO:17);
[0356] -ERETWQAFRT (SEQ ID NO: 18);
[0357] -WREEVRKTFKLFRQ(SEQ ID NO:19);
[0358] -WREMFRLFRT (SEQ ID NO: 20);
[0359] -YFQGAREMFQAFRT(SEQ ID NO:21);
[0360] -WKEEVIKAFKLFRD(SEQ ID NO:22);
[0361] -AVNMFQLFRT (SEQ ID NO: 23);
[0362] -WREEVRKAFKLFRS (SEQ ID NO:24).
[0363] 7. A modular polypeptide complex comprising
[0364] - a first partial effector sequence; and
[0365] - The second effector sequence consists of a sequence selected from the following
[0366] oSEQ ID NO: 004 and SEQ ID NO: 005, or
[0367] o the group of SEQ ID NOs: 006–343,
[0368] Optionally, the second partial effector sequence may include one or two amino acid variations independently selected from deletion, insertion, and variable substitution,
[0369] in
[0370] The first and second effector sequences together constitute an active, circularly arranged, self-labeling protein and, when brought into close proximity with one another, enable covalent attachment of the haloalkyl moiety, and
[0371] -The first effector sequence consists of the following
[0372] o an N-terminal first effector sequence portion characterized by SEQ ID NO: 002 or a sequence that is at least (≥) 90% identical (particularly ≥93%, 95%, 97% or ≥98% identical) to SEQ ID NO: 002,
[0373] o a C-terminal first effector sequence portion characterized by SEQ ID NO: 003 or SEQ ID NO: 370, or a sequence at least (≥) 90% identical (particularly ≥93%, 95%, 97% or ≥98% identical) to SEQ ID NO: 003;
[0374] o an internal cpHalo linker consisting of 10 to 35 amino acids, particularly consisting of 20 to 26 amino acids, more particularly 22 to 24 amino acids, wherein the internal cpHalo linker connects the C-terminus of the N-terminal first effector sequence portion to the N-terminus of the C-terminal first effector sequence portion;
[0375] The modular polypeptide complex is characterized by
[0376] The N-terminal first effector sequence portion comprises at least one N-effector amino acid substitution selected from the group consisting of
[0377] οN217D;
[0378] οF205W;
[0379] οV184E;
[0380] οV197K;
[0381] οR254W;
[0382] and / or
[0383] The C-terminal first effector sequence portion comprises at least one C-effector amino acid substitution selected from the group consisting of
[0384] οF80T;
[0385] οN119H;
[0386] οK117R;
[0387] οR30P;
[0388] οG7D;
[0389] οE20S;
[0390] and / or
[0391] The internal cpHalo linker is characterized by a sequence selected from the group consisting of
[0392] οRSDDPRKTQTIASKISRDLNGS(SEQ ID NO:344);
[0393] οKGGTKRDADKAVRDTLLSLNGQ(SEQ ID NO:345);
[0394] οGGAPRDEALKKIEKAKRDTGDQ (SEQ ID NO:346);
[0395] οKSKYDRDQILKIIAELEKKTGGS (SEQ ID NO:347);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0396] <h2 style=";text-align:left;direction:ltr"> οQSKYPPEWLEKVIRELLKRKNGR(SEQ ID NO:348);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0397] <h2 style=";text-align:left;direction:ltr"> οKSKYDKRQIRDIADKIAKDNNHQ(SEQ ID NO:349);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0398] <h2 style=";text-align:left;direction:ltr"> οGADDKTKIEKILEEIKRRWQGR(SEQ ID NO:350);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0399] <h2 style=";text-align:left;direction:ltr"> οGTSDPRNQEIAKKLARDASTVP(SEQ ID NO:351);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0400] <h2 style=";text-align:left;direction:ltr"> οNGADKEQIDRAIEKAKRDLNNQ(SEQ ID NO:352);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0401] <h2 style=";text-align:left;direction:ltr"> οKGASDRDEAKKLADDIRKKKGDQ(SEQ ID NO:353);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0402] <h2 style=";text-align:left;direction:ltr"> οNSNGHRDELEKILQTIRKQNNDI(SEQ ID NO:354);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0403] <h2 style=";text-align:left;direction:ltr"> οLKDERQRDKALEIADRADKYPTS(SEQ ID NO:355);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0404] <h2 style=";text-align:left;direction:ltr"> οKGAEDAKERLERGDIEKKKKEQP(SEQ ID NO:356);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0405] <h2 style=";text-align:left;direction:ltr"> οKGAEDAKERLERGDLDRWSQEWR(SEQ ID NO:357);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0406] <h2 style=";text-align:left;direction:ltr"> οKGAEDAKERLERRDLDKINSRNS(SEQ ID NO:358);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0407] <h2 style=";text-align:left;direction:ltr"> οKGAEDAKERLEKGYIDDKASKQQ(SEQ ID NO:359);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0408] <h2 style=";text-align:left;direction:ltr"> οKGAEDAKERLEKGELEKKWKDHP(SEQ ID NO:360);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0409] <h2 style=";text-align:left;direction:ltr"> οKGAEDAKERLERGEMEKAVKHGS(SEQ ID NO:361);<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0410] <h2 style=";text-align:left;direction:ltr"> οKGAEDAKERLERDELTRDIKTYPY(SEQ ID NO:362);<h2 style=";text-align:left;direction:ltr">
[0411] οKGAEDAKERLEQGMLEEIKKKYPE(SEQ ID NO:363);
[0412] οKGAEDAKERLERDELTKIAKNLGG (SEQ ID NO:364);
[0413] οKGAEDAKERLEKNALDKIAKSKGD(SEQ ID NO:365);
[0414] οKGAEDAKERLEKNDETLKKAKDKP(SEQ ID NO:366);
[0415] Optionally, the internal cpHalo linker sequence may include one, two, or three amino acid variations independently selected from deletions, insertions, and variable substitutions.
[0416] 7A) A modular polypeptide according to item 7, wherein the second partial effector sequence is selected from the group consisting of SEQ ID NO: 004 and SEQ ID NO: 005.
[0417] 7B) A modular polypeptide according to item 7, wherein the second partial effector sequence is selected from the group of SEQ ID NOs: 006-343.
[0418] 7C) A modular polypeptide according to item 7, wherein the second partial effector sequence is selected from the group of SEQ ID NO: 6-10, 12, 13, 15-26, 28, 30, 31, 34, 40, 46, 48, 50, 75, 85, 98, 103, 112, 154, 156, 160, 167, 173, 175, 182, 189, 195, 212, 214, 233 and 337-342.
[0419] 7D) A modular polypeptide according to item 7, wherein the second partial effector sequence is selected from the group of SEQ ID NO: 020 or 024-039.
[0420] 7E) A modular polypeptide according to item 7, wherein the second partial effector sequence is selected from the group of SEQ ID NO: 020 or 006-024.
[0421] 7F) A modular polypeptide according to item 7, wherein the second partial effector sequence is SEQ ID NO: 006.
[0422] 7G) A modular polypeptide according to item 7, wherein the second partial effector sequence is SEQ ID NO: 037.
[0423] 7H) A modular polypeptide according to item 7, wherein the second partial effector sequence is SEQ ID NO:339.
[0424] 7I) Modular polypeptide according to any one of items 7 to 7H, wherein the modular polypeptide complex is characterized by an E20S substitution relative to the sequence of GenBank-ID: AQS79242.1.
[0425] 7J) Modular polypeptide according to any one of items 7 to 7I, wherein the modular polypeptide complex is characterized by the substitution V184E relative to the sequence of GenBank-ID: AQS79242.1.
[0426] 7K) Modular polypeptide according to any one of items 7 to 7J, wherein the modular polypeptide complex is characterized by the substitution N119H relative to the sequence of GenBank-ID: AQS79242.1.
[0427] 7L) Modular polypeptide according to any one of items 7 to 7K, wherein the modular polypeptide complex is characterized by a V197K substitution relative to the sequence of GenBank-ID: AQS79242.1.
[0428] 7M) Modular polypeptide according to any one of items 7 to 7L, wherein the modular polypeptide complex is characterized by a K117R substitution relative to the sequence of GenBank-ID: AQS79242.1.
[0429] 8. The modular polypeptide complex according to any one of items 7 to 7M, wherein the first partial effector sequence comprises an N-effector amino acid substitution and / or a C-effector amino acid substitution and / or an internal cpHalo linker sequence selected from the group consisting of:
[0430] -E20S and V184E;
[0431] -E20S, N119H, and V184E;
[0432] -E20S, N119H, V184E, and V197K;
[0433] - E20S, N119H, and V184E, and linker sequence KSKYDRDQILKIIAELEKKTGGS (SEQ ID NO: 347);
[0434] - E20S, N119H, V184E and V197K, and linker sequence KSKYDRDQILKIIAELEKKTGGS (SEQ ID NO: 347);
[0435] 9. The modular polypeptide complex of item 8, wherein the first partial effector sequence comprises E20S, N119H, V184E, and V197K, and a linker sequence KSKYDRDQILKIIAELEKKTGGS (SEQ ID NO: 347).
[0436] 10. A modular polypeptide complex comprising the first partial effector sequence according to items 7 to 9, and the second partial effector sequence according to items 1 to 6.
[0437] 11. The modular polypeptide complex according to any one of items 3, or 7 to 10, wherein the variable substitution is selected from substitutions according to the following rules:
[0438] a. Glycine (G) and alanine (A) are interchangeable; valine (V), leucine (L) and isoleucine (I) are interchangeable, and A and V are interchangeable;
[0439] b. Tryptophan (W), phenylalanine (F), and tyrosine (Y) are interchangeable;
[0440] c. Serine (S) and threonine (T) are interchangeable;
[0441] d. Aspartic acid (D) and glutamic acid (E) are interchangeable;
[0442] e. Asparagine (N) and glutamine (Q) are interchangeable; N and S are interchangeable; N and D are interchangeable; E and Q are interchangeable;
[0443] f. Methionine (M) and Q are interchangeable;
[0444] g. Cysteine (C), A and S are interchangeable;
[0445] h. Proline (P), G and A are interchangeable;
[0446] i. Arginine (R), lysine (K) and Q are interchangeable;
[0447] j. Histidine (H) and Y are interchangeable, and H and N are interchangeable;
[0448] kL and M are interchangeable, I and M are interchangeable, V and M are interchangeable;
[0449] lE and K are interchangeable;
[0450] mA and C are interchangeable, A and G are interchangeable; A and T are interchangeable, A and V are interchangeable;
[0451] nR and N are interchangeable, R and E are interchangeable, R and H are interchangeable;
[0452] oN and Q are interchangeable, N and E are interchangeable, N and G are interchangeable, N and K are interchangeable, N and T are interchangeable;
[0453] pD and Q are interchangeable, D and S are interchangeable;
[0454] qQ and H are interchangeable, Q and M are interchangeable, Q and S are interchangeable;
[0455] rE and H are interchangeable, E and S are interchangeable;
[0456] sG and S are interchangeable;
[0457] tF and I are interchangeable, F and L are interchangeable, F and M are interchangeable;
[0458] uK and S are interchangeable;
[0459] vT and V are interchangeable;
[0460] In particular wherein the variable substitution is selected from substitutions according to rules A to L.
[0461] 12. The modular polypeptide complex according to any one of the preceding items, wherein the first effector sequence and the second effector sequence are connected via a sensor module polypeptide, wherein
[0462] - The sensor module polypeptide is selected from
[0463] a) a single sensor polypeptide capable of undergoing a conformational change from a first conformation to a second conformation in response to an external stimulus, in particular the presence or concentration of an analyte compound, or light radiation, wherein
[0464] In the first conformation, the first and second partial effector sequences are in close proximity, and
[0465] In the second conformation, the first and second partial effector sequences are not in close proximity,
[0466] and
[0467] b) a sensor polypeptide pair comprising a first sensor polypeptide and a second sensor polypeptide, wherein the first sensor polypeptide is covalently linked to a first partial effector sequence via a peptide bond and the second sensor polypeptide is covalently linked to a second partial effector sequence,
[0468] The first sensor polypeptide and the second sensor polypeptide are capable of specific molecular interactions,
[0469] And the first and second sensor polypeptides are part of separate polypeptide chains.
[0470] 13. A modular polypeptide complex according to any of the preceding items, wherein optionally, the first partial effector sequence is flanked at the N-terminus and / or C-terminus by a tetrapeptide selected from LKPG or EKKG (SEQ ID NOs 372, 373) at the N-terminus and PDYE, GDVE, PDSN or PDPQ (SEQ ID NOs 374-377) at the C-terminus.
[0471] 14. A nucleic acid sequence, or multiple nucleic acid sequences, encoding the modular polypeptide complex according to any one of the preceding items.
[0472] 15. A nucleic acid expression system comprising the nucleic acid sequence or the plurality of nucleic acid sequences according to item 14, each nucleic acid sequence being under the control of a promoter sequence.
[0473] 16. A cell comprising the nucleic acid expression system according to item 15, wherein the promoter is operable in the cell.
[0474] 17. A non-human transgenic animal or plant comprising the nucleic acid sequence according to item 14 or the nucleic acid expression system according to item 15.
[0475] 18. A kit comprising the nucleic acid sequence according to item 14 or the nucleic acid expression system according to item 15, and a HaloTag substrate.
[0476] 19. A method for detecting a molecular interaction event, comprising the steps of:
[0477] a) providing an expression system capable of expressing the modular polypeptide complex according to any one of items 1 to 13;
[0478] b) adding a haloalkyl moiety to the expression system under conditions that result in expression of the modular polypeptide, wherein the haloalkyl moiety is coupled to the detection moiety;
[0479] c) In a detection step, detecting the detection moiety coupled to the first portion effector sequence.
[0480] 20. The method of claim 19, wherein in the detecting step,
[0481] - a molecular interaction event between the first sensor polypeptide and the second sensor polypeptide, or
[0482] -Internal molecular interaction events of individual sensor peptides
[0483] Be detected.
[0484] 21. The method according to item 19, wherein in the detecting step, co-expression of the first partial effector sequence and the second partial effector sequence is detected.
[0485] 22. The method of claim 19, wherein in the detecting step, the presence of a protein of interest is detected, wherein the protein of interest is fused to
[0486] a) the first effector sequence, or
[0487] b) The second part of the effector sequence.
[0488] 23. The method according to item 22, wherein the expression system is a cell, and the protein of interest is an endogenous protein of the cell.
[0489] 24. The method according to item 23, wherein the nucleic acid sequence encoding the second partial effector sequence is inserted at the 3' end or the 5' end of the endogenous gene encoding the target protein.
[0490] 25. A method according to claim 24, wherein the nucleic acid sequence encoding the second partial effector sequence is inserted by the CRISPR / Cas9 complex.
[0491] 26. The method according to any one of items 23 to 26, wherein the first partial effector sequence is expressed and encoded by a plasmid.
[0492] 27. The method according to any one of items 23 to 25, wherein the first part of the effector sequence is introduced into the cell via a viral vector.
[0493] 28. The method according to any one of items 22 to 27, wherein the second partial effector sequence is selected from the group consisting of SEQ ID NO: 20 and SEQ ID NO: 24-39.
[0494] 29. A method according to any one of items 19 to 28, wherein the detection moiety comprises a purification tag.
[0495] 30. A method according to any one of items 19 to 28, wherein the detection moiety comprises a fluorophore.
[0496] 31. A method according to any one of item 30, wherein in the detection step, the spatial positioning of the detection part is determined.
[0497] 32. A nucleic acid sequence comprising a peptide coding portion encoding a peptide sequence selected from the group consisting of SEQ ID NOs: 6-110, 112-194, 196-343, wherein the peptide coding portion comprises a cloning site,
[0498] - is located at the 5' or 3' end of the peptide encoding portion, and
[0499] - allows for in-frame insertion of the coding sequence of interest into the peptide coding portion.
[0500] 33. A nucleic acid sequence according to item 32, wherein the peptide sequence is selected from the group consisting of SEQ ID NO: 6–10, 12, 13, 15–26, 28, 30, 31, 34, 40, 46, 48, 50, 75, 85, 98, 103, 112, 154, 156, 160, 167, 173, 175, 182, 189, 212, 214, 233, and 337–342.
[0501] 34. The nucleic acid sequence according to item 32, wherein the peptide sequence is selected from the group of SEQ ID NOs: 6-24.
[0502] 35. The nucleic acid sequence according to item 32, wherein the sequence is selected from the group consisting of SEQ ID NO: 20 and SEQ ID NO: 24-39.
[0503] The present invention is further illustrated by the following examples and figures, from which further embodiments and advantages can be derived. These examples are intended to illustrate the present invention but are not intended to limit its scope. BRIEF DESCRIPTION OF THE DRAWINGS
[0504] Figure 1 : shows the concept of the split HaloTag, the labeling reaction of cpHaloΔ in the presence or absence of Hpep, and the concept of a split HaloTag-based recorder for biological activity.
[0505] Figure 2 : It shows that the newly identified Hpep has higher activity than the original Hpep.
[0506] Figure 3 : A) illustrates the identification of Hpep with increased initial labeling speed. The sequence in this figure relates to the SEQ ID NO:5–10,12–13,15–26,28,30–31,34,40,46,48,50,75,85,98,103,112,154,156,160,167,173,175,182,189,195,212,214,233,337–342 in the appended ST26 sequence table, and is also listed in Table 1.
[0507] B) Shows selected Hpeps with increased initial labeling speed. The sequences in this figure relate to SEQ ID NOs: 27, 29, 32-33, 35-39, 343.
[0508] Figure 4 : Shows the relative initial labeling speed of selected Hpep and the EC for identifying Hpep 50Values range from 124 nM to 3.0 mM. The sequences in this figure relate to SEQ ID NOs: 5-6, 12, 16, 20-22, 24-39, 47, 50, 341-342.
[0509] Figure 5 : Shows the FKBP / FRB / RAPA model system.
[0510] Figure 6 : Shows peptides with higher activity in the FKBP / FRB / RAPA model system. The sequences in this figure relate to SEQ ID NOs: 5-23 and 369 in the attached ST26 sequence listing and are also listed in Table 1.
[0511] Figure 7 : shows sequence mutations that increase labeling rate and shows an increase in melting temperature.
[0512] Figure 8 : cpHaloΔCP linker showing increased labeling speed and melting temperature.
[0513] Figure 9 : shows the cpHaloΔ variant with increased labeling kinetics and increased melting temperature.
[0514] Figure 10 : Shows the relative initial labeling speed of the improved cpHaloΔ variant (E20S, N119H, V184E, and V197K, and the linker sequence KSKYDRDQILKIIAELEKKTGGS; SEQ ID NO: 347) with selected high affinity Hpep. The sequences in this figure are related to SEQ ID NOs: 37–39; see "EC of Hpep for Split HaloTag Labeling Reactions" in Example 1 below. 50 value".
[0515] Figure 11 : Shows the results of testing cpHaloΔ variants with increased affinity for Hpep. Using sequences identified by yeast display screening, improved cpHaloΔ variants (E20S, N119H, V184E, and V197K, and linker sequence KSKYDRDQILKIIAELEKKTGGS; SEQ ID NO: 347) were extended by 4 amino acids at the N-terminus and / or C-terminus (flanking sequences are SEQ ID NOs 372 to 377).
[0516] Figure 12Figure 2: Results of testing certain high-affinity split HaloTag variants using the non-covalent HaloTag substrate T5-CPY. cpHaloΔ variants are based on a modified version (E20S, N119H, V184E, and V197K, and linker sequence KSKYDRDQILKIIAELEKKTGGS; SEQ ID NO: 347) plus extensions (SEQ ID NOs: 372, 373, 374, and 375), as shown. Peptides correspond to SEQ ID NOs: 31 and 34.
[0517] Example
[0518] Example 1: Improved Hpep
[0519] We recently developed a break-apart system for the self-labeling protein HaloTag to generate transient cellular event recorders ( Figure 1 ), which is described in patent US20220275350A1 (also PCT / EP2020 / 060785, Hiblot et al.). We circularly permuted HaloTag and discovered new termini at positions 154 / 156 or 141 / 145 that preserved the protein's overall folding and tagging activity. Removal of residues 142–155 to generate cpHaloΔ preserved the overall folding of HaloTag but essentially abolished its activity. However, activity could be restored by reversible binding of the peptide Hpep1 (residues 145–154, SEQ ID NO: 5).
[0520] We then tested whether this split HaloTag system could detect the rapamycin-dependent interaction between FKBP and FRB. If the split HaloTag fragment was fused to the C-termini of FKBP and FRB, labeling was highly dependent on the presence of rapamycin. When fused to the more distal N-termini of FKBP and FRB (47 Å between the N-termini and 10 Å between the C-termini), the labeling was highly dependent on the presence of rapamycin. ), no labeling was observed. We hypothesize that this is due to the low affinity (K D =4.6 mM), which does not allow for complementation of the split fragments over longer distances. Therefore, we computationally designed peptides with higher affinity for cpHaloΔ by building a custom RosettaScripts protocol. 40,000 peptide sequences were generated and ranked by Rosetta total score and peptide binding free energy. 384 selected peptide sequences were synthesized and tested for their ability to activate purified cpHaloΔ protein. 80% of these peptides resulted in cpHaloΔ labeling faster than Hpep1 ( Figure 2 ).
[0521] Promising peptides were purified, and additional peptides with combinations of beneficial features (mutations or extensions) were synthesized and purified. Other candidate peptides identified by yeast surface display and phage display were also synthesized and purified. Sixty of these peptides showed that cpHaloΔ labeled faster than Hpepl, with some increase in the initial labeling rate, reaching 30,000 times ( Figure 3 To evaluate the affinity of these modified peptides for cpHaloΔ, we determined the EC 50 Many of these peptides showed efficient labeling even at low μM or nM concentrations, under conditions where Hpepl had no detectable activity, and were characterized by EC 50 The values ranged from 124 nM to 3.0 mM ( Figure 4 We then fused selected modified Hpep and cpHaloΔ to the distal N-termini of FKBP and FRB and observed efficient labeling of several variants by the complexes in the presence of rapamycin (>100-fold increase compared to Hpep1), whereas no labeling was detected in the absence of rapamycin ( Figure 5 、 6 Thus, access to different Hpeps with different affinities significantly facilitates the application of the split HaloTag system and greatly increases its scope by making it compatible with protein-protein interactions that are not characterized by adjacent termini.
[0522] To also test the usability of the isolated HaloTag in living mammalian cells, we expressed membrane-localized FKBP-cpHaloΔ and cytosolic Hpep1-FRB in cultured HeLa cells. The membrane-localized FKBP-cpHaloΔ and CPY-CA were highly dependent on the addition of rapamycin, and CPY-labeled cells could be identified by fluorescence microscopy or flow cytometry analysis.
[0523] Example 2: Improved cpHaloΔ
[0524] Even when using high concentrations of optimized Hpep, the labeling speed of split HaloTag was still significantly slower than that of HaloTag. To investigate this difference, we measured the melting temperature (Tm) of the cpHaloΔ protein and found that it had greatly reduced thermal stability (Tm: 30.4°C) compared to HaloTag (Tm: 62.0°C). This suggests that at the physiological temperature of 37°C, most of the protein may not be in its active fold. To overcome this problem, we aimed to stabilize cpHaloΔ through computational protein design.
[0525] We use the PROSS server ( https: / / pross.weizmann.ac.il / step / pross-terms / ) to predict stabilizing mutations within cpHaloΔ and select 10 mutation candidates that were tested as single point mutants. Seven mutants had higher melting temperatures, of which six were characterized by increased labeling speed in the presence of the Hpep variant SKRDAREMFQAFRT (SEQ ID NO: 6) ( Figure 7 ).
[0526] Next, we decided to use the Rosetta software suite ( https: / / www.rosettacommons.org ) redesigned the CP linker connecting the original N- and C-termini of HaloTag in cpHaloΔ. We hypothesized that replacing the flexible (GGTGGSGGTGGS GGS, SEQ ID NO: 368) linker with a well-folded helical linker that binds to the protein surface would increase stability. Characterization of 23 cpHaloΔ variants featuring different redesigned linkers in the presence of the Hpep variant SKRDAREMFQAFRT revealed that all 23 linkers increased labeling speed ( Figure 8 ) and 22 linkers increase the melting temperature of the protein ( Figure 8 ), among which linker 04 (KSKYDRDQILKIIAELEKKTGGS, SEQ ID NO: 34) showed the best performance.
[0527] We then combined beneficial point mutations with linker 04 to determine whether their effect was additive to further improve the stability and activity of cpHaloΔ. Indeed, several combinations showed increased Tm and labeling speed. The best variant (linker 04 / E20S / N119H / V184E / V197K) was characterized by a 96-fold increase in labeling kinetics and a 13.6°C increase in melting temperature (Tm: 45.4°C). Figure 9 ). This variant is referred to hereinafter as cpHaloΔ2.
[0528] We expected that cpHaloΔ2 would not only be characterized by higher stability and activity, but also have a higher affinity for Hpep. To test this hypothesis, we determined the EC values of three high-affinity Hpep 50 Values (RMWTWREMFRLFRT, SEQ ID NO: 37, RQWSWREMFRLFRT, SEQ ID NO: 38, RGWSWREMFRLFRT, SEQ ID NO: 39). On average, EC 50 The value was 12 times lower than that of the original cpHaloΔ, reaching 2.5 nM ( Figure 10 This high-affinity cpHaloΔ-Hpep pair can be classified as a self-complementary split-type system and opens up new application cases for the split-type HaloTag system.
[0529] Example 3: Improved cpHaloΔ variants for increased affinity for Hpep
[0530] A particularly promising application of the self-complementing split HaloTag is the labeling of endogenous proteins expressed in living cells. The endogenous protein can be tagged with Hpep through genome engineering, and cpHaloΔ2 can be expressed from a vector or another genomic site. The split HaloTag will then spontaneously complement, and the endogenous protein can be labeled with the HaloTag ligand. Due to the relative ease of introducing a short DNA sequence encoding the small Hpep into an endogenous locus compared to introducing larger tags (such as the entire HaloTag or fluorescent protein), this can greatly facilitate the labeling of endogenous proteins with bright and stable fluorescent HaloTag substrates.
[0531] To further improve the spontaneous complementation between cpHaloΔ2 and Hpep:RGWSWREMFRLFRT (SEQ ID NO:34), we extended the N-terminus or C-terminus of cpHaloΔ2 by 4 amino acids and screened for high-affinity binders by yeast surface display. Promising N-terminal and C-terminal sequences were combined, and the EC values of these extended cpHaloΔ variants with Hpep:RGWSWREMFRLFRT (SEQ ID NO:39) were determined. 50 The best variant (LKPG-cpHaloΔ2-PDYE) (flanking sequences SEQ ID 372 and 374) was characterized by an EC value 6.6 times lower than that of cpHaloΔ2. 50 Value (382pM)( Figure 11 This subnanomolar affinity should improve the performance of the isolated HaloTag as a tool for labeling endogenous proteins in living cells.
[0532] Example 4: Testing high affinity split HaloTag variants with non-covalent HaloTag substrates
[0533] The main advantage of using split HaloTag to tag endogenous proteins over other split systems such as split GFP is that fluorescent HaloTag ligands are superior for challenging imaging techniques such as super-resolution microscopy (e.g. STED, PAINT, MINLUX). This is especially true when using exchangeable HaloTag ligands (xHTLs) that reduce photobleaching effects. Therefore, we tested whether self-complementing split HaloTag variants are compatible with xHTL by measuring xHTL (T5-CPY) affinity in the presence or absence of Hpep: WKRDWREMFRLFRT (SEQ ID NO: 34) and SKRDWREMFRLFRT (SEQ ID NO: 31). All cpHaloΔ variants tested showed significantly stronger xHTL binding in the presence of either Hpep. The optimal split HaloTag pair is characterized by a K of 100 in the presence of Hpep. D The affinity value was as low as 238 nM, and the affinity difference in the presence and absence of Hpep was as high as 315-fold ( Figure 12 ).
[0534] Materials and methods:
[0535] Melting temperature of nanoDSF
[0536] The fluorescence intensity ratio at 350 nm and 330 nm was monitored by measuring the change in the fluorescence intensity ratio at 350 nm and 330 nm in 20 μM activity buffer (50 mM HEPES, 50 mM NaCl, pH 7.3) at 1 °C min on a Prometheus NT 48 nanometer differential scanning fluorimeter (NanoTemper). -1 The thermal stability of proteins was measured in the temperature range of 20°C to 95°C at a heating rate of 1.5°C. The melting temperatures shown (average of 2 samples) correspond to the inflection points (maximum of the first derivative).
[0537] General Methods for Fluorescence Polarization Analysis
[0538] For different analyses, the concentrations of protein, dye, and peptide used were different and are given in the following sections. An exemplary reaction condition for determining labeling speed is shown below: 100 μL of a solution containing 400 nM protein and 10 μM peptide was mixed with 100 μL of 100 nM fluorescent HaloTag substrate (i.e., halo-CPY) in a buffer (50 mM NaCl, 50 mM HEPES, pH 7.3, 0.5 mg / ml BSA) in a 96-well plate (black - no binding - flat bottom). Alternatively, 20 μL of a solution containing 400 nM protein and 10 μM peptide was mixed with 20 μL of 100 nM fluorescent HaloTag substrate (i.e., halo-CPY) in a buffer (50 mM NaCl, 50 mM HEPES, pH 7.3, 0.5 mg / ml BSA) in a 385-well plate (black - no binding - flat bottom). Labeling kinetics were measured by recording fluorescence polarization over time at 37°C in a microplate reader (TECAN Spark20M). If the labeling reaction reached a plateau, the data were fit to a second-order reaction model to obtain an estimate of the apparent second-order rate constant, kapp, and the initial slope. If the reaction did not reach a plateau, a linear model was fitted to obtain an estimate of the initial slope.
[0539] Initial screening of 384 Hpep libraries
[0540] The labeling kinetics of cpHaloΔ (2 μM) with tetramethylrhodamine HaloTag ligand (TMR-HTL, 100 nM) were measured by fluorescence polarization reading in the presence of each peptide. Peptides were ranked by the initial slope of the labeling reaction compared to native cpHaloΔ and native Hpep (10-mer).
[0541] Characterization of purified peptides
[0542] The labeling kinetics of cpHaloΔ (500 nM) and TMR-HTL (100 nM) were measured using fluorescence polarization readings in the presence of 10 μM of each peptide. Peptides were ranked by the initial slope of the labeling reaction compared to native cpHaloΔ and native Hpep (10-mer).
[0543] EC of Hpep for split HaloTag labeling reaction 50 value
[0544] The labeling kinetics of cpHaloΔ (20 nM or 500 nM) and TMR-HTL (4 nM or 100 nM) were measured over a range of Hpep concentrations (lower cpHaloΔ and TMR-HTL concentrations were used for high affinity peptides). The initial slope of the labeling reaction was determined. The slope was plotted against Hpep concentration and fitted with a sigmoidal model to determine the EC50 Value. EC 50 : Peptide concentration at which half-maximal response velocity is observed.
[0545] Testing the ability of the new Hpep to label distal protein complexes (FKBP / FRB)
[0546] The labeling kinetics of purified recombinant proteins (FKBP / FRB split HaloTag fusion, 250 nM) with TMR-HTL (50 nM) were measured in the presence or absence of the dimerization agent rapamycin (500 nM). The initial slope of the labeling reaction was determined using native cpHaloΔ and native Hpep (10-mer) and is presented in comparison with the corresponding constructs.
[0547] cpHaloΔ point mutation screening
[0548] The labeling kinetics of the cpHaloΔ variant (100 nM) and TMR-HTL (20 nM) were measured in the presence of the Hpep variant SKRDAREMFQAFRT (SEQ ID NO: 6) (6.25 μM). The initial slope of the labeling reaction was determined and compared to the parent protein. The melting temperature was measured using tryptophan fluorescence (ie, NanoDSF).
[0549] cpHaloΔCP linker screening
[0550] The labeling kinetics of cpHaloΔ variants (100 nM) and TMR-HTL (20 nM) were measured in the presence of Hpep variant SKRDAREMFQAFRT (SEQ ID NO: 6) (12.5 μM). The apparent second-order rate constant (k app The melting temperature was measured using tryptophan fluorescence (ie, NanoDSF).
[0551] Characterization of cpHaloΔ variants with combinations of beneficial point mutations and CP linkers
[0552] The labeling kinetics of cpHaloΔ variants (100 nM) and TMR-HTL (20 nM) were measured in the presence of Hpep variant SKRDAREMFQAFRT (SEQ ID NO: 6) (12.5 μM). The apparent second-order rate constant (k app The melting temperature was measured using tryptophan fluorescence (ie, NanoDSF).
[0553] Hpep for split HaloTag tagging reactions with N-terminally and / or C-terminally extended cpHaloΔ variants: RGWSWREMFRLFRT, EC 50 value
[0554] The labeling kinetics of cpHaloΔ (4 nM) and TMR-HTL (0.8 nM) were measured over a range of Hpep:RGWSWREMFRLFRT (SEQ ID NO: 39) concentrations. The initial slope of the labeling reaction was determined. The slope was plotted against Hpep concentration and fitted with a sigmoidal model to determine the EC 50 Value. EC 50 : Peptide concentration at which half-maximal response velocity is observed.
[0555] Testing high-affinity split-type HaloTag variants with non-covalent HaloTag substrate (xHTL)
[0556] Binding of xHTL T5-CPY (10 nM) was measured by fluorescence polarization at different concentrations (0.00128 to 100 μM) of cpHaloΔ protein alone or in complex with Hpep. A sigmoidal model was fitted to the data to obtain K D value.
[0557] T5-CPY HaloTag Substrate:
[0558]
[0559] Example 4: Sequence
[0560] Table 1 provides the peptide sequences, which are also described in the accompanying ST26 sequence listing.
[0561]
[0562]
[0563]
[0564]
Claims
1. A modular polypeptide complex comprising - A variant of the first effector sequence consisting of: o an N-terminal first effector sequence portion characterized by SEQ ID NO: 002, o a C-terminal first effector sequence portion characterized by SEQ ID NO: 003, and o an internal cpHalo linker consisting of 10 to 35 amino acids, connecting the C-terminus of the N-terminal first effector sequence portion to the N-terminus of the C-terminal first effector sequence portion; and - a second partial effector sequence consisting of a sequence selected from the group consisting of SEQ ID NO: 004 to SEQ ID NO: 343; in The first and second partial effector sequences together constitute an active, circularly permuted, variant of the self-labeling protein identified as GenBank-ID: AQS79242.1, and The modular polypeptide complex is characterized by at least one substitution relative to the sequence of GenBank-ID: AQS79242.1 selected from the group consisting of E20S; V184E; N119H; V197K; K117R; N217D; F205W; F80T; R30P; G7D.
2. The modular polypeptide complex according to claim 1, wherein the first effector sequence and the second effector sequence are connected by a sensor module polypeptide, wherein - The sensor module polypeptide is selected from a) a single sensor polypeptide capable of undergoing a conformational change from a first conformation to a second conformation in response to an external stimulus, in particular the presence or concentration of an analyte compound, or light radiation, wherein In the first conformation, the first and second partial effector sequences are in close proximity, and In the second conformation, the first and second partial effector sequences are not in close proximity, and b) a sensor polypeptide pair comprising a first sensor polypeptide and a second sensor polypeptide, wherein the first sensor polypeptide is covalently linked to a first partial effector sequence via a peptide bond and the second sensor polypeptide is covalently linked to a second partial effector sequence, The first sensor polypeptide and the second sensor polypeptide are capable of specific molecular interactions, And the first and second sensor polypeptides are part of separate polypeptide chains.
3. The modular polypeptide complex according to claim 1 or 2, wherein the modular polypeptide complex is characterized by an E20S substitution.
4. The modular polypeptide complex according to any one of the preceding claims, wherein the modular polypeptide complex is characterized by a V184E substitution.
5. The modular polypeptide complex according to any one of the preceding claims, wherein the modular polypeptide complex is characterized by an N119H substitution.
6. The modular polypeptide complex according to any one of the preceding claims, wherein the modular polypeptide complex is characterized by a V197K substitution.
7. The modular polypeptide complex according to any one of the preceding claims, wherein the modular polypeptide complex is characterized by a K117R substitution.
8. The modular polypeptide complex according to any one of the preceding claims, wherein the internal cpHalo linker is characterized by a sequence selected from the group consisting of οRSDDPRKTQTIASKISRDLNGS(SEQ ID NO:344); οKGGTKRDADKAVRDTLLSLNGQ(SEQ ID NO:345); οGGAPRDEALKKIEKAKRDTGDQ (SEQ ID NO:346); οKSKYDRDQILKIIAELEKKTGGS (SEQ ID NO:347); οQSKYPPEWLEKVIRELLKRKNGR(SEQ ID NO:348); οKSKYDKRQIRDIADKIAKDNNHQ(SEQ ID NO:349); οGADDKTKIEKILEEIKRRWQGR(SEQ ID NO:350); οGTSDPRNQEIAKKLARDASTVP(SEQ ID NO:351); οNGADKEQIDRAIEKAKRDLNNQ(SEQ ID NO:352); οKGASDRDEAKKLADDIRKKKGDQ(SEQ ID NO:353); οNSNGHRDELEKILQTIRKQNNDI(SEQ ID NO:354); οLKDERQRDKALEIADRADKYPTS(SEQ ID NO:355); οKGAEDAKERLERGDIEKKKKEQP(SEQ ID NO:356); οKGAEDAKERLERGDLDRWSQEWR(SEQ ID NO:357); οKGAEDAKERLERRDLDKINSRNS(SEQ ID NO:358); οKGAEDAKERLEKGYIDDKASKQQ(SEQ ID NO:359); οKGAEDAKERLEKGELEKKWKDHP(SEQ ID NO:360); οKGAEDAKERLERGEMEKAVKHGS(SEQ ID NO:361); οKGAEDAKERLERDELTRDIKTYPY(SEQ ID NO:362); οKGAEDAKERLEQGMLEEIKKKYPE(SEQ ID NO:363); οKGAEDAKERLERDELTKIAKNLGG (SEQ ID NO:364); οKGAEDAKERLEKNALDKIAKSKGD(SEQ ID NO:365); οKGAEDAKERLEKNDETLKKAKDKP (SEQ ID NO: 366).
9. The modular polypeptide complex according to any one of the preceding claims, wherein the internal cpHalo linker is KSKYDRDQILKIIAELEKKTGGS (SEQ ID NO: 347).
10. The modular polypeptide complex according to any of the preceding claims, wherein the second portion effector sequence consists of a sequence selected from the group consisting of SEQ ID NO: 6–10, 12, 13, 15–26, 28, 30, 31, 34, 40, 46, 48, 50, 75, 85, 98, 103, 112, 154, 156, 160, 167, 173, 175, 182, 189, 195, 212, 214, 233, and 337–342.
11. The modular polypeptide complex according to any one of the preceding claims, wherein the second part effector sequence is selected from the group consisting of SEQ ID NO: 006, SEQ ID NO: 037 and SEQ ID NO:
339.
12. A nucleic acid sequence, or a plurality of nucleic acid sequences, encoding a modular polypeptide complex according to any one of the preceding claims.
13. A nucleic acid expression system comprising the nucleic acid sequence or the plurality of nucleic acid sequences according to claim 12, each nucleic acid sequence being under the control of a promoter sequence. A cell comprising the nucleic acid expression system according to claim 13 , wherein the promoter is operable in the cell.
15. A method for detecting a molecular interaction event, comprising the steps of: a) providing an expression system capable of expressing the modular polypeptide complex according to any one of claims 1 to 11; b) adding a haloalkyl moiety to the expression system under conditions that result in expression of the modular polypeptide, wherein the haloalkyl moiety is coupled to the detection moiety; c) In a detection step, detecting the detection moiety coupled to the first portion effector sequence.
16. The method according to claim 15, wherein in the detecting step, - a molecular interaction event between the first sensor polypeptide and the second sensor polypeptide is detected, or - the internal molecular interaction events of a single sensor polypeptide are detected, or - co-expression of the first partial effector sequence and the second partial effector sequence is detected, or - detecting the presence of a protein of interest, wherein the protein of interest is fused to the first partial effector sequence or the second partial effector sequence.
Citation Information
Patent Citations
Circularly permutated haloalkane transferase fusion molecules
US12371676B2
Circularly permutated haloalkane transferase fusion molecules
US20220275350A1
Novel tunable photoactivatable silicon rhodamine fluorophores
WO2019122269A1
Cell-permeable fluorogenic fluorophores
WO2020115286A2
Circularly permutated haloalkane transferase fusion molecules
WO2020212537A1