calibration protein
An ancestral citrate synthase protein forms multiple complex states for stable and cost-effective calibration of mass photometry, addressing the limitations of existing standards by providing precise and reproducible mass calibration.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- MAX PLANCK GESELLSCHAFT ZUR FOERDERUNG DER WISSENSCHAFTEN EV
- Filing Date
- 2024-05-13
- Publication Date
- 2026-05-19
AI Technical Summary
Current calibration standards for mass photometry are costly and not optimized for the technique, requiring frequent recalibration due to fluctuating interference contrast, and existing proteins form monomers rather than a variety of stable multimers.
A single ancestral citrate synthase protein that forms multiple distinct complex formation states, such as dimers, trimers, and tetramers, even at low concentrations, providing stable and reproducible mass calibration.
The ancestral citrate synthase protein offers precise and cost-effective calibration with a reproducible mass error of approximately 1%, maintaining stability over time and through freeze-thaw cycles, suitable for mass photometry and other techniques like native polyacrylamide gel electrophoresis.
Smart Images

Figure 2026516134000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the provision of molecular weight standards for various instruments and techniques for determining (measuring) molecular weight, among other things. Generally, molecular weight standards contain a diverse range of unknown proteins with molecular weights in the range of thousands to hundreds of thousands of Daltons. The present invention simplifies the provision of molecular weight standards by providing a single protein species that has the ability to form multiple distinct (different) complex formation states (i.e., dimers, trimers, and / or tetramers) even at relatively low concentrations in solution. Such a simplified provision enables the calibration of instruments to be carried out, reduces costs, and provides appropriate and robust standards for specific purposes for newer techniques such as mass photometry.
[0002] The present invention relates to an ancestral citrate synthase protein that can simultaneously form at least two, at least three, at least four, or at least five different complex formation states. The present invention further relates to a composition comprising the ancestral citrate synthase protein, the use of the ancestral citrate synthase protein or the composition, a nucleic acid encoding the ancestral citrate synthase protein, and a method for purifying the ancestral citrate synthase protein.
Background Art
[0003] Mass photometry is a novel technique that has only recently been commercially established, and is an example of an interferometric scattering microscopy technique (WO2018 / 011591) that can be used to determine the molecular mass of proteins, protein complexes, and other biomolecules in solution in a single-molecule approach. The signal measured is the interference contrast created by the light scattered by the biomolecule of interest and the light reflected by the measurement surface. The measured contrast directly correlates with the molecular mass (Young, G. et al.). However, translating that signal to the molecular weight of the biomolecule requires a calibration step that must be performed using samples of various known masses. These calibration measurements need to be performed regularly, usually at least once a day, or sometimes more frequently, after each measurement session, because the interference contrast fluctuates with temperature, laser run time, and other changing parameters. Thus, calibration samples, also called calibration standards, are consumables necessary to operate this technique.
[0004] The calibration standards currently in use are not specifically designed or optimized for mass photometry, but rather reuse commercially available protein mixtures used for gel electrophoresis, such as NativeMark® Unstained Protein Standard (Invitrogen® LC0725). These standards contain mixtures of multiple undisclosed proteins at different sizes and concentrations, and their use is quite costly.
[0005] This invention relates to a single protein that forms multiple distinct complex-forming states in solution. The resulting protein population (population of complex-forming states) is stable and has distinct molecular weight differences, which can be detected and identified by mass photometry. This invention enables reproducible and convenient mass calibration using a single protein that can be mass-produced in heterologous host bacteria. The protein itself is a citrate synthase, but its amino acid sequence is not found in naturally occurring or extant organisms. It is an ancient representative example of this enzyme family, and its amino acid sequence is inferred by ancestral sequence reconstruction, a method of reviving ancestral proteins using the relevant amino acid sequence.
[0006] Natural citrate synthases are known to form complexes; for example, eukaryotes (in their mitochondria) and Gram-positive bacteria use type I citrate synthase, which forms dimers, while Gram-negative bacteria use type II citrate synthase, which forms hexamers. However, in extant citrate synthase proteins, monomers form only specific multimers, rather than a variety of stable multimers. In a study by Schmidtmann et al., citrate synthases from the genus Arabidopsis were analyzed. When all mitochondrial proteins were studied together, Blue-Native PAGE showed the presence of dimers (100 kDa) and high molecular weight complexes up to 1000 kDa. Considering the presence of other proteins, the authors concluded that these could be "multimeric aggregates." Mitochondrial extracts were then further analyzed by gel filtration chromatography, followed by Western blotting. Alternatively, recombinant CS4 was analyzed. Although the protein concentration was not confirmed, the amounts used were 100 μg (recombinant CS4) and 180 μg (mitochondrial extract), and the maximum loading volume of the column used was 500 μL; therefore, the concentration far exceeded the amount used for mass standard (reference) purposes. Based on the isolated proteins, the authors concluded that the recombinant proteins were able to form dimers (notably the most abundant form) and oligomeric complexes.
[0007] In contrast, the citrate synthase protein of the present invention produces a defined and stable series of complex-forming states even at low protein concentrations such as 10 nM. Each individual complex-forming state occurs at appropriate abundances, is sufficiently different in size to be detectable, and can be identified by mass photometry. The mass difference may be at least 25 kDa, and this difference is measured between different complex-forming states for each monomeric unit (i.e., monomer versus dimer).
[0008] The resulting mass calibration yields a reproducible mass error of approximately 1%, which is at least as good as commercially available protein standards. The citrate synthase protein provided herein is also relatively more stable over long periods at -20°C and after repeated freeze-thaw cycles. It has also been shown to be sufficiently stable even after storage at 4°C for several weeks, which is not the case with commercially available protein mixtures. The abundance of species of different masses, and the differences in these masses, are significantly better optimized for mass photometer calibration than commercially available protein mixtures. The protein can also be used as a molecular weight standard for other non-denaturing techniques such as native polyacrylamide gel electrophoresis (Native PAGE).
[0009] In addition, the present invention includes a two-step protein purification protocol, which involves heterologous production in Escherichia coli (E. coli), resulting in high yield, purity, and batch-to-batch consistency. [Overview of the project]
[0010] The present invention provides a protein that can simultaneously form multiple distinct complex-forming states, and once formed, maintains these complex-forming states, providing a heterogeneous mixture of complexes, each having a distinct molecular mass. Such a protein is ideal for the calibration of instruments, devices, and techniques for determining the mass of biomolecules and the like.
[0011] The present invention can be described in several ways. The invention provided herein is a citrate synthase protein or a functional fragment thereof encoded by a nucleic acid sequence, wherein the nucleic acid is an ancestral gene, and the citrate synthase protein or a functional fragment thereof can simultaneously form at least two, at least three, at least four, or at least five complex-forming states.
[0012] The inventions provided herein are citrate synthase proteins or functional fragments thereof having an amino acid sequence that is at least 75%, 80%, 85%, 90%, or 95% identical to a sequence selected from Sequence IDs 1 to 7, wherein the citrate synthase proteins or functional fragments thereof can simultaneously form at least two, at least three, at least four, or at least five complex-forming states.
[0013] The invention provided herein is a citrate synthase protein or a functional fragment thereof having an amino acid sequence having at least 75%, 80%, 85%, 90%, or 95% sequence similarity to a sequence selected from Sequence IDs 1 to 7, wherein the citrate synthase protein or a functional fragment thereof can simultaneously form at least two, at least three, at least four, or at least five complex-forming states.
[0014] The invention provided herein is a citrate synthase protein or a functional fragment thereof comprising a sequence selected from any of SEQ ID NOs: 1 to 7. The invention provided herein is a citrate synthase protein or a functional fragment thereof comprising a sequence selected from SEQ ID NOs: 15-21 and a sequence having more than 75%, 80%, 85%, 90%, or 95% sequence identity, wherein the citrate synthase protein or a functional fragment thereof can simultaneously form at least two, at least three, at least four, or at least five complex-forming states.
[0015] The invention provided herein is a citrate synthase protein or a functional fragment thereof comprising a sequence having more than 75%, 80%, 85%, 90%, or 95% sequence similarity to a sequence selected from SEQ ID NOs: 15-21, wherein the citrate synthase protein or a functional fragment thereof can simultaneously form at least two, at least three, at least four, or at least five complex-forming states.
[0016] The invention provided herein is a citrate synthase protein or a functional fragment thereof comprising a sequence selected from any of SEQ ID NOs: 15 to 21. The invention provided herein is a citrate synthase protein or a functional fragment thereof containing at least one ancestral (sudden) mutation, and the citrate synthase protein or a functional fragment thereof can simultaneously form at least two, at least three, at least four, or at least five complex-forming states.
[0017] The invention provided herein is a composition comprising a citrate synthase protein or a functional fragment thereof according to any of the above embodiments, wherein the concentration of the citrate synthase protein or a functional fragment thereof is at least 5 nM, preferably 5 nM to 100 μM.
[0018] It may be preferable that citrate synthase according to any definition of the present invention can simultaneously form at least three complex-forming states. Each of these three complex-forming states is present in solution in an amount sufficient to be detectable. Notably, each of the three complex-forming states may constitute at least 15% of the total citrate synthase protein complex present.
[0019] It may be preferable that citrate synthase according to any definition of the present invention can simultaneously form at least four complex-forming states. Each of these four complex-forming states is present in solution in an amount sufficient to be detectable. Notably, each of the four complex-forming states may constitute at least 5% of the total citrate synthase protein complex present.
[0020] The most abundant complex-forming state may preferably constitute less than 50% of the total citrate synthase present. Optionally, the most abundant (most common, dominant) complex-forming state may preferably constitute less than 45%, less than 40%, or less than 35% of the total citrate synthase present.
[0021] It is preferable that at least three distinct complex-forming states are formed in solution and are sufficiently specific (different) to enable the use of the protein as a mass standard. This can be achieved by having a difference of at least 25 kDa between each of the complex-forming states. Optionally, there can be a difference of at least 30 kDa, 35 kDa, 40 kDa, 45 kDa, or 50 kDa between each of the complex-forming states. Thus, for example, monomers may be 50 kDa, dimers 100 kDa, trimers 150 kDa, and so on. The difference can be measured between different complex-forming states for each monomeric unit (i.e., monomer vs. dimer, tetramer vs. pentamer, etc.).
[0022] The invention provided herein is the use of citrate synthase protein or a functional fragment thereof, or a composition thereof, as defined in any of the foregoing definitions of the present invention as molecular weight standards.
[0023] The invention provided herein is the use of any of the compositions defined herein in the calibration or standardization of biochemical techniques. The invention provided herein is a method for calibrating a mass photometry device, the method comprising detecting the mass of particles in a composition according to any of the above definitions of the invention using a mass photometry device.
[0024] The invention provided herein is a nucleic acid sequence encoding any of the above-defined citrate synthases or a functional fragment thereof of the invention. The invention provided herein is a vector comprising the nucleic acid according to any of the above definitions of the invention.
[0025] The invention provided herein is a cell transformed with a vector comprising the nucleic acid according to the above definition of the invention. The invention provided herein is a method for producing a citrate synthase protein or a functional fragment thereof according to the above definition of the invention, the method comprising culturing a cell transformed with a vector according to the above definition of the invention, and performing at least one purification step.
Brief Description of Drawings
[0026] [Figure 1] FIG. 1 shows the overall profiles of phylogenetic trees 1 and 2 having general taxonomic descriptors common to each branch. <( [Figure 2a] FIGS. 2a - f show enlarged views of tree 1 used for the reconstruction of ancestral sequences. The amino acid sequences of 418 existing citrate synthase genes (from bacteria, archaea, and eukaryotes) were collected from the NCBI Reference Sequence Database and aligned using MUSCLE v3.8.31 28. The maximum likelihood phylogeny was inferred from the multiple sequence alignment using raxML v8.2.10 29. The nodes corresponding to the resurrected proteins are labeled CS1, CS5, CS6, and CS7. Labels a - j indicate how the phylogenetic branches connect across the figure. [Figure 2b]Figures 2a–f show enlarged views of Tree 1 used for ancestral sequence reconstruction. Amino acid sequences of 418 extant citrate synthase genes (from bacteria, archaea, and eukaryotes) were collected from the NCBI Reference Sequence Database and aligned using MUSCLE v3.8.31 28. The maximum likelihood phylogeny was inferred from multiple sequence alignments using raxML v8.2.10 29. The sections corresponding to the reconstructed proteins are labeled CS1, CS5, CS6, and CS7. Labels a–j show how phylogenetic branching connects across the figure. [Figure 2c] Figures 2a–f show enlarged views of Tree 1 used for ancestral sequence reconstruction. Amino acid sequences of 418 extant citrate synthase genes (from bacteria, archaea, and eukaryotes) were collected from the NCBI Reference Sequence Database and aligned using MUSCLE v3.8.31 28. The maximum likelihood phylogeny was inferred from multiple sequence alignments using raxML v8.2.10 29. The sections corresponding to the reconstructed proteins are labeled CS1, CS5, CS6, and CS7. Labels a–j show how phylogenetic branching connects across the figure. [Figure 2d] Figures 2a–f show enlarged views of Tree 1 used for ancestral sequence reconstruction. Amino acid sequences of 418 extant citrate synthase genes (from bacteria, archaea, and eukaryotes) were collected from the NCBI Reference Sequence Database and aligned using MUSCLE v3.8.31 28. The maximum likelihood phylogeny was inferred from multiple sequence alignments using raxML v8.2.10 29. The sections corresponding to the reconstructed proteins are labeled CS1, CS5, CS6, and CS7. Labels a–j show how phylogenetic branching connects across the figure. [Figure 2e]Figures 2a–f show enlarged views of Tree 1 used for ancestral sequence reconstruction. Amino acid sequences of 418 extant citrate synthase genes (from bacteria, archaea, and eukaryotes) were collected from the NCBI Reference Sequence Database and aligned using MUSCLE v3.8.31 28. The maximum likelihood phylogeny was inferred from multiple sequence alignments using raxML v8.2.10 29. The sections corresponding to the reconstructed proteins are labeled CS1, CS5, CS6, and CS7. Labels a–j show how phylogenetic branching connects across the figure. [Figure 2f] Figures 2a–f show enlarged views of Tree 1 used for ancestral sequence reconstruction. Amino acid sequences of 418 extant citrate synthase genes (from bacteria, archaea, and eukaryotes) were collected from the NCBI Reference Sequence Database and aligned using MUSCLE v3.8.31 28. The maximum likelihood phylogeny was inferred from multiple sequence alignments using raxML v8.2.10 29. The sections corresponding to the reconstructed proteins are labeled CS1, CS5, CS6, and CS7. Labels a–j show how phylogenetic branching connects across the figure. [Figure 3a] Figures 3a–g are enlarged views of the generated tree 2 used for the reconstruction of the ancestral sequence. Amino acid sequences of 418 extant citrate synthase genes (from bacteria, archaea, and eukaryotes) were collected from the NCBI Reference Sequence Database and aligned using MUSCLE v3.8.31 28. Maximum likelihood lineages were inferred from multiple sequence alignments using raxML v8.2.10 29. The sections corresponding to the reconstructed proteins are labeled CS2, CS3, and CS4. Labels a–i show how phylogenetic branching connects across the figure. [Figure 3b]Figures 3a–g are enlarged views of the generated tree 2 used for ancestral sequence reconstruction. Amino acid sequences of 418 extant citrate synthase genes (from bacteria, archaea, and eukaryotes) were collected from the NCBI Reference Sequence Database and aligned using MUSCLE v3.8.31 28. Maximum likelihood lineages were inferred from multiple sequence alignments using raxML v8.2.10 29. The sections corresponding to the reconstructed proteins are labeled CS2, CS3, and CS4. Labels a–i show how phylogenetic branching connects across the figure. [Figure 3c] Figures 3a–g are enlarged views of the generated tree 2 used for ancestral sequence reconstruction. Amino acid sequences of 418 extant citrate synthase genes (from bacteria, archaea, and eukaryotes) were collected from the NCBI Reference Sequence Database and aligned using MUSCLE v3.8.31 28. Maximum likelihood lineages were inferred from multiple sequence alignments using raxML v8.2.10 29. The sections corresponding to the reconstructed proteins are labeled CS2, CS3, and CS4. Labels a–i show how phylogenetic branching connects across the figure. [Figure 3d] Figures 3a–g are enlarged views of the generated tree 2 used for ancestral sequence reconstruction. Amino acid sequences of 418 extant citrate synthase genes (from bacteria, archaea, and eukaryotes) were collected from the NCBI Reference Sequence Database and aligned using MUSCLE v3.8.31 28. Maximum likelihood lineages were inferred from multiple sequence alignments using raxML v8.2.10 29. The sections corresponding to the reconstructed proteins are labeled CS2, CS3, and CS4. Labels a–i show how phylogenetic branching connects across the figure. [Figure 3e]Figures 3a–g are enlarged views of the generated tree 2 used for ancestral sequence reconstruction. Amino acid sequences of 418 extant citrate synthase genes (from bacteria, archaea, and eukaryotes) were collected from the NCBI Reference Sequence Database and aligned using MUSCLE v3.8.31 28. Maximum likelihood lineages were inferred from multiple sequence alignments using raxML v8.2.10 29. The sections corresponding to the reconstructed proteins are labeled CS2, CS3, and CS4. Labels a–i show how phylogenetic branching connects across the figure. [Figure 3f] Figures 3a–g are enlarged views of the generated tree 2 used for ancestral sequence reconstruction. Amino acid sequences of 418 extant citrate synthase genes (from bacteria, archaea, and eukaryotes) were collected from the NCBI Reference Sequence Database and aligned using MUSCLE v3.8.31 28. Maximum likelihood lineages were inferred from multiple sequence alignments using raxML v8.2.10 29. The sections corresponding to the reconstructed proteins are labeled CS2, CS3, and CS4. Labels a–i show how phylogenetic branching connects across the figure. [Figure 3g] Figures 3a–g are enlarged views of the generated tree 2 used for ancestral sequence reconstruction. Amino acid sequences of 418 extant citrate synthase genes (from bacteria, archaea, and eukaryotes) were collected from the NCBI Reference Sequence Database and aligned using MUSCLE v3.8.31 28. Maximum likelihood lineages were inferred from multiple sequence alignments using raxML v8.2.10 29. The sections corresponding to the reconstructed proteins are labeled CS2, CS3, and CS4. Labels a–i show how phylogenetic branching connects across the figure. [Figure 4] Figure 4 shows images of the SDS-PAGE gel from CS1-His protein to CS7-His protein. [Figure 5]Figure 5 shows images of the Native PAGE gels for CS1-His protein, CS2-His protein, and CS3-His protein. [Figure 6] Figure 6a shows mass photometry histograms taken from CS1-His samples obtained from four different purification methods, demonstrating batch-by-batch reproducibility. Figure 6b shows mass photometry histograms taken from CS1-His samples obtained from four different purification methods, demonstrating batch-by-batch reproducibility. Figure 6c shows mass photometry histograms taken from CS1-His samples obtained from four different purification methods, demonstrating batch-by-batch reproducibility. Figure 6d shows mass photometry histograms taken from CS1-His samples obtained from four different purification methods, demonstrating batch-by-batch reproducibility. [Figure 7] Figure 7 shows a comparison of the mass of a protein with a known molecular weight with the mass obtained by mass photometry using CS1-His as a calibrator. [Figure 8] Figure 8a shows a comparison of mass photometry histograms using NativeMark® versus CS1-His HMW. Figure 8b shows a comparison of mass photometry histograms using NativeMark® versus CS1-His HMW. [Modes for carrying out the invention]
[0027] Citrate synthases are ubiquitous enzymes involved in the citrate cycle, catalyzing the formation of citrate through the condensation of acetate (derived from acetyl-CoA) and oxaloacetate. While all citrate synthases likely originate from the same common ancestral protein, eukaryotes (in their mitochondria) and Gram-positive bacteria utilize type I citrate synthase, which forms dimers, while Gram-negative bacteria utilize type II citrate synthase, which forms hexamers. To better understand the factors contributing to this structural diversity, ancestral protein sequences were inferred using phylogenetic methods, and these proteins were expressed and purified in E. coli. Quite surprisingly, rather than forming only dimers or hexamers, the resulting "ancestral" citrate synthase proteins simultaneously exhibited a range of complex formation states in solution. Different ancestral citrate proteins exhibit complex formation states with varying distributions, most commonly dimers, tetramers, hexamers, octamers, and decamers, providing a detection target set with regularly spaced molecular weights across the molecular weight range of approximately 85 kDa to 430 kDa. Ancestral citrate synthase proteins are also highly stable, surviving multiple freeze-and-thaw cycles and prolonged incubation under refrigeration with minimal changes in their complex formation properties. As a result of these characteristics, ancestral citrate synthase proteins offer an ideal solution to problems in mass photometry calibration.
[0028] According to one definition of the present invention, a citrate synthase protein or a functional fragment thereof encoded by a nucleic acid sequence is provided herein, the nucleic acid being an ancestral gene, and the citrate synthase protein or its functional fragment can simultaneously form at least two, at least three, at least four, or at least five complex-forming states.
[0029] The present invention may simply relate to the citrate synthase protein as defined herein. As used herein, the term “functional fragment” means a portion or part of a protein sequence that retains the ability to form multiple complex-forming states. Functional fragments can be produced by any suitable method, but may involve deletions of sequences that are not important for interaction with other protein sequences. At least 3%, 5%, 7%, 10%, 12%, 15%, or 20%, or more, of the protein sequence may be removed in a functional fragment. To avoid doubt with respect to the present invention, the term “functional fragment” in this example does not refer to any particular enzyme function.
[0030] As used herein, “ancestor gene” refers to a common gene proposed to originate from a family of genes. Ancestor genes may originate from phylogenetic techniques such as ancestor gene reconstruction, in which an ancestral protein sequence is inferred as described above (Harms & Thornton 2010, Harms & Thornton 2013, Selberg et al., 2021), and the DNA molecule encoding that protein is synthesized. Thus, an ancestor gene is typically an artificial DNA sequence inferred from an artificial amino acid sequence. That amino acid sequence may be inferred using ancestor sequence reconstruction as described below.
[0031] Citrate synthase proteins that originate from ancestral genes can be simply referred to as "ancestral citrate synthase." This invention utilizes ancestral sequence reconstruction as a starting point for engineering operations on citrate synthase protein, which can be used as a mass calibration standard. Thus, the ancestral gene can be inferred by ancestral sequence reconstruction. To the extent relevant to this invention, ancestral sequence reconstruction uses the vast and ever-expanding amount of sequence data available in sequence databases to produce the current amino acid sequence alignment of citrate synthase protein. Then, using phylogenetic and statistical analysis with a suitable evolutionary model, the amino acid sequences at the branching points or "nodes" of the phylogenetic tree generated by ancestral sequence reconstruction are defined. The amino acid sequences at the nodes of the tree are candidates for the ancestral amino acid sequence of citrate synthase, which led to the amino acid sequence of the extant citrate synthase protein. Briefly, a phylogenetic tree, also known as a phylogenetic tree, is a diagram that depicts the lines of evolutionary descent of different genes that originate from a common ancestor.
[0032] The citrate synthase proteins produced by the generation of ancestral sequence reconstructions according to the present invention are found to possess unique properties not observed in existing citrate synthase proteins. Therefore, these ancestral sequences benefit from inherently different self-interaction specificities. It has been found that ancestral sequences can simultaneously form different multimeric morphologies in solution. Thus, ancestral sequences can simultaneously form two or more, preferably three or more, higher-order multimeric morphologies including dimers, trimers, tetramers, pentamers, hexamers, and octamers and decamers.
[0033] This ancestral sequence reconstruction technique has been applied to citrate synthase proteins from bacteria, archaea, and eukaryotes. Specifically, as described in the examples, extant citrate synthase proteins were used to construct phylogenetic trees 1 and 2 shown in Figures 1-3. Briefly, the extant protein sequences are aligned. Comparison is performed by aligning the sequences, thereby allowing for scoring of amino acid differences. This is not easy when the sequences are not relatively similar, and / or when they diverge due to the accumulation of insertions and deletions, and also due to point mutations. Therefore, for these reasons, and to align multiple sequences, it is more common to rely on relevant computer programs to perform the alignment. In the examples, MUSCLE (Multiple Sequence Comparison by Log-Expectation, described in Edgar, RCBMC Bioinformatics 5, 113 (2004).doi.org / 10.1186 / 1471-2105-5-113) was used for sequence alignment. Other available alignment software includes Clustal, T-Coffee, Bali-Phy, or Phylo.
[0034] Once alignment is complete, one of several suitable models can be used to reconstruct the tree. There are several widely used methods for estimating phylogenetic trees (neighbor-joining method, maximum parsimony, Bayesian estimation, and maximum likelihood estimation). In this example, RAxML (Randomized Axelerated Maximum Likelihood, described in Stamatakis A., Bioinformatics) was used, which is a program for sequential and parallel maximum likelihood-based estimation of large phylogenetic trees. Other available phylogenetic tree generation software includes PhyML (Guindon S. & Gascuel O.Syst. Biol). Other available software for estimating phylogenetic trees includes PhyML, MrBayes, FastTree, or IQ-Tree.
[0035] As used herein, “complex-forming state” refers to a protein complex containing a specific number of monomeric units of the same protein. When two or more monomeric units are present, the complex-forming state may also be described as a macromolecule or multimer, which are generally formed by monomers that are non-covalently bonded to one another. Generally, non-covalent interactions between hydrophobic and hydrophilic regions in monomeric units help stabilize the quaternary structure of the complex-forming state. Therefore, in multimers, non-covalent interactions between monomeric units typically position the monomers into a specific structure. Thus, monomeric units are usually assembled into a regular and predictable structure. The number of monomeric units in a complex-forming state can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10. Therefore, the complex-forming state can be monomer (1), dimer (2), trimer (3), tetramer (4), pentamer (5), hexamer (6), and larger polymers (usually 7 or more). Thus, it may be considered that this term includes both monomers and homooligomers, where a monomer is a single protein capable of forming a complex, and a homooligomer is a complex containing a specific number of identical monomers.
[0036] Monomeric units exist in solution, forming various complex-forming states. These are distinct and stable, and as a result, environmental conditions such as freezing / thawing do not affect the proportion of each polymer present. The solution can be any suitable solution for the protein. The solution may contain, for example, any suitable buffer and / or salt. The concentration of the protein in the solution can be 5 nM to 1000 nM (1 μM).
[0037] In contrast, agglomeration refers to the formation of protein clusters that are not constructed uniformly but are based on random interactions. The ancestral citrate synthase sequence identified by the inventors possesses a remarkable ability to form two or more complex-forming states when present in solution. In fact, the inventors observed that the ancestral citrate synthase protein was able to simultaneously form more than three distinct complex-forming states in solution. Indeed, in Figures 6a–d, where each citrate synthase protein complex was counted by mass photometry, at least five distinct and identifiable complex-forming states were identifiable. For example, Figure 6d has seven identifiable complex-forming states. Furthermore, monomeric units appear to be able to form multiple complex-forming states without one complex-forming state being excessively dominant (exceeding 50% of all citrate synthase protein complexes present in terms of number). For example, in Figures 6a–6d, the maximum amount of the most common complex-forming state in the total solution was 33% (172 kDa) of all measured citrate synthase protein complexes. In contrast, the least common complex-forming state present in Figures 6a and 6b constituted 2% of the total. However, many complex formation states exist in a relatively good proportion, ranging from 10 to 35% (in terms of number) of all citrate synthase protein complexes.
[0038] Mass photometry has the ability to detect the mass of a single particle in a solution without the need for labeling. Generally, it is common to detect the mass of all particles present (e.g., on a coverslip) and count the number of particles assigned to each mass; see, for example, Figures 6a-6d, where the "count" of each particle is shown. This gives us the total number of particles present, as well as the count of each type of particle (in this invention, each protein complex). Since monomers are one type that can be formed by citrate synthase, they are understood to be included in the term "protein complex." Therefore, in order to determine (measure) the percentage of each citrate synthase protein complex, the number of such complexes present is determined together with the total number. Thus, the percentages listed herein are based on numbers.
[0039] Therefore, it may be preferable that the ancestral citrate synthase can simultaneously form at least three complex-forming states, and that each of these three complex-forming states contributes at least 10%, 15%, 20%, 21%, 22%, 23%, 24%, or 25% (in terms of number) of the total citrate synthase protein complexes in solution. Alternatively, the three major (or most common) complex-forming states may each be present in 15–40% (in terms of number), or optionally, 18–35% (in terms of number), of the total citrate synthase protein complexes in solution.
[0040] As used herein, “simultaneously” means at the same time. Therefore, in a solution of the citrate synthase of the present invention, different complex-forming states exist that are formed at the same time, such as monomers, dimers, trimers, and tetramers. Once formed, these complex-forming states are stable and separate, i.e., dissociation and reassociation occur only minimally, if any.
[0041] As with most standardization techniques, it is theoretically possible to provide measurements of just two known points that define a gradient, and in the case of mass photometry, this is the molecular weight gradient against radiometric contrast. However, in practice, at least three known points are generally required, as this allows for confirmation of the gradient and definition of the endpoint of the detection range of the instrument used. Thus, it is preferable that at least three distinct complex-forming states are formed by the citrate synthase of the present invention. This allows for the use of citrate synthase as a molecular mass marker in mass photometry. As described above, the at least three complex-forming states of the citrate synthase of the present invention are individually present in at least 10%, 15%, 20%, 21%, 22%, 23%, 24%, or 25% (in terms of number) of the total citrate synthase protein complex present in solution. Thus, each of the three most common / most abundant complex-forming states is present in a concentration sufficient to allow for accurate detection, and therefore calibration, of the instrument.
[0042] The citrate synthase protein or its functional fragment of the present invention may preferably have an amino acid sequence that is at least 75%, 80%, 85%, 90%, or 95% identical to a sequence selected from SEQ ID NOs: 1 to 7, and the citrate synthase protein or its functional fragment may also simultaneously form at least three complex-forming states.
[0043] The citrate synthase protein or its functional fragment of the present invention may preferably have an amino acid sequence having at least 75%, 80%, 85%, 90%, or 95% sequence similarity to a sequence selected from SEQ ID NOs: 1 to 7, and the citrate synthase protein or its functional fragment may also simultaneously form at least three complex-forming states.
[0044] Sequence identity refers to the comparison of aligned sequences. If both sequences contain the same amino acid at the same position after alignment, the sequences are considered identical at that position. Tools for calculating sequence identity are readily available and widely used in many fields of life science. For the purposes of this application, sequence identity refers to the identity output from the publicly available Emboss Needle program https: / / www.ebi.ac.uk / Tools / psa / emboss_needle / (Needleman and Wunsch(https: / / www.sciencedirect.com / science / article / abs / pii / 0022283670900574?via%3Dihub)).
[0045] Sequence similarity is related to sequence identity but allows for similar properties of some amino acids. For example, both glutamic acid and aspartic acid have similar chemical, structural, and electrostatic properties. Therefore, if one sequence contains glutamic acid at the position where the other contains aspartic acid after alignment, these sequences are considered similar but not identical at this position. Similar relationships are considered between other amino acids. For the purposes of this application, sequence similarity refers to the similarity output from the publicly available Emboss Needle program https: / / www.ebi.ac.uk / Tools / psa / emboss_needle / (Needleman and Wunsch(https: / / www.sciencedirect.com / science / article / abs / pii / 0022283670900574?via%3Dihub)).
[0046] The present invention may more preferably relate to a citrate synthase protein or a functional fragment thereof containing a sequence selected from any of SEQ ID NOs: 1 to 7. Furthermore, the citrate synthase protein may contain, or consist of, the sequence of SEQ ID NO: 1.
[0047] The citrate synthase protein or its functional fragment of the present invention may contain at least one ancestral mutation. The citrate synthase protein or its functional fragment of the present invention may contain at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 ancestral mutations.
[0048] As used herein, “ancestral mutation” refers to a change (alteration) in the first amino acid sequence of a protein such as citrate synthase, where the change to the first amino acid sequence increases the similarity and / or identity of the first amino acid sequence to a second amino acid sequence of the same protein, where the second amino acid sequence is an ancestral sequence generated by phylogenetic methods such as ancestral sequence reconstruction. In such techniques, the ancestral sequence is inferred by comparing known sequences of extant organisms with the knowledge of their evolutionary history described above.
[0049] As a simple example, the sequence of citrate synthase from E. coli (accession number WP_166726827) is shown below, aligned with sequence number 1 (sequence number 29). Sequence number 1 is an ancestral sequence inferred by comparing known sequences of existing organisms (including E. coli, but not a requirement) with knowledge of their evolutionary history. These sequences have 37.1% identity and 52.7% similarity, as determined by Emboss Needle.
[0050] [ka] Examples of ancestral mutations are as follows:
[0051] This is a T50V mutation in SEQ ID NO: 29, which increases its identity with SEQ ID NO: 1. This is the W396I mutation in SEQ ID NO: 29, which increases the similarity to SEQ ID NO: 1.
[0052] Removing any / all of residues 297-299 of sequence number 29 increases the overlap with sequence number 1. The addition of suitable residues between N349 and D350 in SEQ ID NO: 29 increases duplication with SEQ ID NO: 1, and optionally increases identity or similarity.
[0053] The citrate synthase protein or its functional fragment of the present invention may preferably have at least 2, at least 3, at least 4, at least 5, at least 10, at least 15, or at least 20 ancestral mutations. Those skilled in the art can easily determine, by trial and error, the type, location, and number of ancestral mutations required to obtain a citrate synthase protein that can simultaneously form at least 2, at least 3, at least 4, or at least 5 complex-forming states.
[0054] Any of the citrate synthase proteins or functional fragments referenced above may be conjugated with another amino acid sequence to form part of a fusion protein. For example, any of the citrate synthase proteins provided herein may further include a purification tag, which may be a His tag, a Strep tag, or any other suitable affinity tag. The purification tag may be located at the N-terminus or C-terminus of the citrate synthase protein. Such a purification tag facilitates the purification of the target protein. A spacer may be present between the citrate synthase protein and the purification tag. The spacer may consist of 1 to 5 amino acids. The spacer may increase the flexibility of the purification tag in solution and thus increase its ability to bind to its affinity partner. The spacer may include a target for cleavage so that the purification tag can be removed after the processing step.
[0055] Additionally or alternatively, citrate synthase proteins or functional fragments thereof are provided herein, comprising sequences selected from SEQ ID NOs: 15-21 and sequences having more than 75%, 80%, 85%, 90%, or 95% sequence identity, and the citrate synthase proteins or functional fragments thereof can simultaneously form at least two, at least three, at least four, or at least five complex-forming states.
[0056] Alternatively, citrate synthase proteins or functional fragments thereof are provided herein, comprising sequences having more than 75%, 80%, 85%, 90%, or 95% sequence similarity to sequences selected from SEQ ID NOs: 15-21, and the citrate synthase proteins or functional fragments thereof can simultaneously form at least two, at least three, at least four, or at least five complex-forming states.
[0057] Alternatively, a citrate synthase protein or a functional fragment thereof comprising a sequence selected from any of SEQ ID NOs: 15 to 21 is provided herein. In one embodiment, the citrate synthase protein or a functional fragment thereof comprises the sequence of SEQ ID NO: 15.
[0058] According to any description of the present invention, the citrate synthase protein or its functional fragment is not expressed by any living organism. “Living organism” means any currently living organism, including any known living organism, but not including organisms whose genes have been recombined / transformed / edited for the purpose of expressing the citrate synthase protein or its functional fragment described herein. The citrate synthase protein described herein is a hypothetical ancestral protein and is not thought to be produced by any living organism. Therefore, the citrate synthase protein provided herein has an artificially constructed amino acid sequence and is not known to exist naturally.
[0059] In any embodiment provided herein, the citrate synthase protein or its functional fragment may be an artificial citrate synthase protein or its functional fragment. In this context, "artificial" means produced from a non-natural gene sequence rather than a gene sequence found in existing organisms.
[0060] The citrate synthase protein or its functional fragments according to the present invention can be readily screened for their ability to form multiple complex-forming states using mass photometry or other native mass spectrometry techniques.
[0061] In addition, compositions comprising citrate synthase protein or a functional fragment thereof as described herein are provided herein, wherein the concentration of citrate synthase protein or a functional fragment thereof is at least 5 nM, preferably 5 nM to 100 μM. Thus, the concentration of the protein in the solution may be 5 nM to 1000 nM (1 μM) or 1 μM to 100 μM, and any value between these ranges. The concentration of the protein in the solution may vary depending on the parameters in the instrument or technique used for calibration.
[0062] composition, A buffering agent suitable for maintaining pH at 5-10 during regular use, and / or It may further include any suitable salt that can be optionally selected from a list containing sodium chloride and potassium chloride, The total salt concentration is at least 50 mM, and preferably in the range of 50 to 500 mM.
[0063] The use of citrate synthase proteins or functional fragments thereof, or compositions, as described in any of the foregoing descriptions of the present invention as molecular weight standards is further provided herein. Several biochemical techniques rely on the use of standard proteins having known molecular weights. The citrate synthase proteins of the present invention may be used in any non-denaturing technique, including but not limited to mass photometry, native PAGE, and gel filtration.
[0064] Mass photometry requires periodic calibration, usually at least once a day, or sometimes more frequently, during measurement sessions, because interference contrast varies with temperature, laser run time, and other changing parameters.
[0065] The complex formation state of citrate synthase proteins or their functional fragments depends on the protein structure that determines protein-protein interactions. Therefore, citrate synthase proteins are not suitable as molecular weight standards in denaturation techniques that reduce proteins to their primary structure, such as sodium dodecyl sulfate PAGE (SDS-PAGE). If the complex formation state of proteins is fixed, for example, by chemical crosslinking, they become resistant to dissociation and may have broader applications in more techniques.
[0066] In another embodiment, the use of compositions according to any description of the present invention in the calibration or standardization of instruments and / or biochemical techniques is provided herein. Preferably, the compositions may be used for the calibration of mass photometry devices. Optionally, the compositions may be used as molecular weight standards in Native PAGE.
[0067] In another embodiment, a method for calibrating a mass photometry device is further provided herein, comprising the step of detecting the mass of particles in a composition described herein.
[0068] For calibration purposes, the detected contrast values are converted to molecular mass using the instrument's software, and the instrument's software is calibrated using molecular weight standards (as described herein). Calibration should be periodically confirmed by measuring the compositions of the present invention and verifying the molecular masses of known protein complexes. Thus, the molecular mass of the calibrator / molecular weight standard is known, and the instrument's software can be adjusted based on the detected mass of the calibrator / molecular weight standard. As good laboratory practice, it is recommended to include calibration measurements in each set of experiments.
[0069] In addition, nucleic acid sequences encoding citrate synthase according to any of the above embodiments are provided herein. Citrate synthase protein can be encoded by ancestral genes.
[0070] Selectively, nucleic acids are codon-optimized to produce proteins in living organisms, and preferably, this sequence is codon-optimized for expression in E. coli. Genetic codes have some redundancy, meaning that different codons code for the same amino acid. However, organisms typically do not produce tRNA at the same level for each codon, and may have preferred codons that produce more tRNA. Codon optimization refers to modifying the redundancy of protein-coding nucleic acids to take advantage of the known tRNA distribution of an organism, thereby increasing protein expression.
[0071] In addition, vectors containing nucleic acids described herein are provided herein. As used herein, a vector refers to any particle used as a vehicle to artificially transport an exogenous nucleic acid sequence into a cell in which the exogenous nucleic acid sequence can be replicated and / or expressed. Vectors include, but are not limited to, plasmids, cosmids, viruses, and phages.
[0072] Optionally, the vector is a plasmid, and preferably, the vector is a pET plasmid. Furthermore, cells transformed with a vector containing the nucleic acid described herein are provided herein. Preferably, the cells are E. coli cells. Optionally, the transformed E. coli cells express the citrate synthase described herein.
[0073] Furthermore, the present invention includes a method for producing the citrate synthase protein described herein, comprising the steps of culturing the cells described herein and performing at least one purification step.
[0074] Optionally, the purification step may include affinity chromatography, or the purification step may include size exclusion chromatography (SEC). In the embodiments described above, the term "citric acid synthase protein" is used to describe the proteins of the present invention, but it is not required that the proteins have any enzymatic activity to satisfy this definition. The proteins simply need to contain a sequence that is originally derived from or similar to the citrate synthase protein provided herein, and may be modified by any of the methods described above. The proteins provided herein can be isolated or purified.
[0075] It should be noted that ancestral proteins and genes are products of algorithmic analysis that determine the probability that each amino acid is at any given position in the protein sequence at a specific node of the phylogenetic tree. The ancestral sequences disclosed herein are high-probability solutions for ancestral sequence reconstruction algorithms. It is highly likely that the actual sequence of the ancestral protein or encoding gene cannot be reliably known because the biological sample is too old to survive intact. This does not diminish the usefulness of the present invention in any way. [Examples]
[0076] Example 1 - Ancestral protein / gene rearrangement of citrate synthase protein Two sets of amino acid sequences (Table 2) of 418 extant citrate synthase genes from bacteria, archaea, and eukaryotes were collected from the NCBI Reference Sequence Database, and each set was aligned using MUSCLE v3.8.31. Maximum likelihood phylogenetic trees were inferred from multiple sequence alignments using raxML v8.2.10. The LG substitution matrix was determined and used by selecting an automated best fit model, as well as by a gamma model of fixed base frequency and rate heterogeneity. The two trees are generally similar, differing mainly in the position of peroxisome citrate synthases. The position of this family in tree 2 better reflects its established evolutionary history.
[0077] Based on a citrate synthase tree and multiple sequence alignments, ancestral sequences (CS1-CS7) were inferred using the codeML package in PAML v4.9. To account for gaps and different N-terminus lengths of the citrate synthase sequences, these ancestral states were determined using parsimony inference by PAUP 4.0a based on binary versions of multiple sequence alignments (1=amino acid, 0=gap, no residue). The state assignments (amino acids or gaps) for each node in the tree were then applied to the inferred ancestral sequences. CS1 Sequence ID 1 MVAKGLEGVVAAESSISYIDGQEGRLYYRGYPIEELAEHSSFEEVAYLLWHGRLPTREELDEFKEELAENRAIPEEIIDLLRTLPKSAHPMAALRTAVSALGMFDPDADDVQSPEANYRKAIRLIAKIPTIVAAFHRIRQGQEPVAPRPDLSHAANFLYMLNGEEPSPVQAKVMDVALILHAEHEMN ASTFAARVVASTLSDMYSAITAAIGALKGPLHGGANEQVMKMLQEIGSPDKAEPWVQEKLANKRRIMGFGHRVYKTYDPRAKILKKMARQLAEKHGDTKLYEIAEAVEKVVVERLGPKGIYPNVDFYSGLVYHALGIPTDLFTPIFAMARVSGWTAHVLEQLEDNRLIRPRAVYVGPTDRKYVPIDQR CS2 Sequence ID 2 MATMEIKKGLEGVVVAETKISYIDGQEGRLYYRGYPIQELAEHSTFEEVAYLLLYGRLPTRDELDEFKEELAEHRALPEQIIDLLKNLPKDAHPMAALRTAVSALGMFDPDADDTSPEARYRKAIRLIAKIPTIVAAFHRIRQGQDPVAPRPDLSHAANFLYMLNGEEPSPVEAKVFDVALILHADHEM NASTFAALVVASTLSDMYSAITAAIGALKGPLHGGANEVMKMLQEIGSPDKAEPWVQEKLANKERIMGFGHRVYKTYDPRARILKKYAKQLAEKHGDSKLYEIAEAVEKVVVERLGPKGIYPNVDFYSGIVYYSMGIPTDLFTPIFAMARIAGWTAHILEYLKDNRLIRPRAVYVGPTDRKYVPIDQR CS3 Sequence ID 3 MSTTEIAKGLEGVVFTETKLSFIDGQEGRLYYLGYPIQELAEHSTFEEVSFLLLHGRLPTREELEAFKEELAANRALPEELIDALRAYPKDAHPMSALRTAVSELGMFDPDAEDTSPEGRYQKSVRLIAKFATIVAAIKRIREGQDPVAPRPDLSHAANFLYMLNGEEPSPEQAKLFDVALILHADHGM NASTFTALAVASTLSDMYSSITAAIGALKGPLHGGANEAVMKMLQEIGSPDKAEAWVQEKLANKERIMGMGHRVYKAFDPRARILKKYAEQVAEKHGKSKYYEILETVEKEVVKRLGPKGIYPNVDFYSGVVYSDLGIPTEFFTPIFAVARISGWTAHILEYTRDNRLLRPKAVYVGELDRKYVPIDQR CS4 Sequence ID 4 MASMEYTPGLAGVVAAESSISYIDGQEGILRYRGYPIEELAEHSTFEEVAYLLLFGELPTRDELEEFDHELKHHRALPERIIDLLKNLPKSAHPMAALQSAVAALGMFYPDADVTDPEGNYEAAVRLIAKLPTIVAAFHRIRRGQDPIAPRDDLGHAANFLYMLNGEEPDPLAARVFDVCLILHAEHSM NASTFTARVVGSTLADPYSAIAAAIGSLSGPLHGGANEEVLQMLEEIGSPDNVEPWLEEKLARKEKIMGFGHRVYKVKDPRATILQKMAEQLFEKHGSTPLYDIALELEKVAAERLGPKGIYPNVDFYSGIVYQKMGIPTDLFTPIFAIARVAGWTAHWLEQLEDNRIFRPSQIYVGPTDRSYVPIDER CS5 Sequence ID 5 MYTPGLAGVVAAESSISYIDGQEGILRYRGYPIEELAEHSSFEEVAYLLLYGRLPTREELEEFDHELKHHRAIPEGIIDLLKTLPKSAHPMAALQSAVAALGMFYPEADDVEDPEANYKKAVRLIAKLPTIVAAFHRIRRGKEPVAPRSDLSHAANFLYMLNGEEPDPLAARVMDVCLILHAEHEMN ASTFTARVVGSTLADPYSAIAAAIGALSGPLHGGANEEVLKMLQEIGSADNVEPWLEEKLATKQKIMGFGHRVYKTYDPRAKILKKMAEQLFEKHGSTPLYDIAVELEKVAAERLGPKGIYPNVDFYSGLVYQALGIPTDLFTPIFAIARVAGWTAHWLEQLEDNRIFRPRQVYVGPTDRKYVPIDQR CS6 Sequence ID 6 MIAKGLEGVVIAETSISYIDGQEGRLYYRGYPIEELAEHSSFEEVAYLLWHGRLPTREELEEFKEELAKNRAIPEEIIDLLRTLPKSAHPMAALRTAVSALGMFDPDADDTQSPEARYRKAIRLIAKIPTIVAAFHRIRQGQEPVAPRPDLSHAANFLYMLNGEEPSPVQAKVMDVALILHAEHEMN ASTFAALVVASTLSDMYSAITAAIGALKGPLHGGANEQVMKMLQEIGSPDKAEPWVQEKLANKRRIMGFGHRVYKTYDPRAKILKKYARQLAEKQGDTTLYEIAEAVEKVVVERLGPKGIYPNVDFYSGLVYHALGIPTELFTPIFAMARVSGWTAHVLEYLKDNRLIRPRAVYVGPTDRKYVPIDQR CS7 Sequence ID 7 MYTPGLAGVVAAESSISYIDGQEGILRYRGYPIEELAEHSSFEEVAYLLLFGKLPTREELEEFDHELKHHRAIPEGIIDLLKTLPKSAHPMALQSAVAALGMFYPADDDVEDPEANYEAAVRLIAKLPTIVAAFHRIRRGKDPIAPRSDLGHAANFLYMLNGEEPDPLAARVMDVCLILHAEHSMNA STFTARVVGSTLADPYSAIAAAIGSLSGPLHGGANEEVLQMLQEIGSAENVEPWLEEKLATKQKIMGFGHRVYKVKDPRATILQKMAEQLFEKHGSTPLYDIAVELEKVAAERLGPKGIYPNVDFYSGLVYQKLGIPTDLFTPIFAIARVAGWTAHWLEQLEDNRIFRPTQIYVGPTDRSYVPIDQR CS1-His sequence number 15 MVAKGLEGVVAAESSISYIDGQEGRLYYRGYPIEELAEHSSFEEVAYLLWHGRLPTREELDEFKEELAENRAIPEEIIDLRTLPKSAHPMALRTAVSALGMDFDPDADDVQSPEANYRKAIRLIAKIPTIVAAFHRIRQGQEPVAPRPDLSHAANFLYMNLNGEESPVQAKVMDVALILHAEHEMNASTF AARVVASTLSDMYSAITAAIGALKGPLHGGANEQVMKMLQEIGSPDKAEPWVQEKLANKRRIMGFGHRVYKTYDPRAKILKKMARQLAEKHGDTKLYEIAEAVEKVVVERLGPKGIYPNVDFYSGLVYHALGIPTDLFTPIFAMARVSGWTAHVLEQLEDNRLIRPRAVYVGPTDRKYVPIDQRLEHHHHHH CS2-His sequence number 16 MATMEIKKGLEGVVVAETKISYIDGQEGRLYYRGYPIQELAEHSTFEEVAYLLLYGRLPTRDELDEFKEELAEHRALPEQIIDLLKNLPKDAHPMAALRTAVSALGMFDPDADDTSPEARYRKAIRLIAKIPTIVAAFHRIRQGQDPVAPRPDLSHAANFLYMNLNGEESPVEAKVFDVALILHADHEMNAST FAALVVASTLSDMYSAITAAIGALKGPLHGGANEEVMKMLQEIGSPDKAEPWVQEKLANKERIMGFGHRVYKTYDPRARILKKYAKQLAEKHGDSKLYEIAEAVEKVVVERLGPKGIYPNVDFYSGIVYYSMGIPTDLFTPIFAMARIAGWTAHILEYLKDNRLIRPRAVYVGPTDRKYVPIDQRLEHHHHHH CS3-His sequence number 17 MSTTEIAKGLEGVVFTETKLSFIDGQEGRLYYLGYPIQELAEHSTFEEVSFLLLLHGRLPTREELEAFKEEALANRALPEELIDALRAYPKDAHPMSALRTAVSELGMFDPDAEDTSPEGRYQKSVRLIAKFATIVAAIKRIREGQDPVAPRPDLSHAANFLYMNLNGEEPSPEQAKLFDVALILHADHGMNAST FTALAVASTLSDMYSSITAAIGALKGPLHGGANEAVMKMLQEIGSPDKAEAWVQEKLANKERIMGMGHRVYKAFDPRARILKKYAEQVAEKHGKSKYYEILETVEVKRLGPKGIYPNVDFYSGVVYSDLGIPTEFFTPIVARISGWTAHILEYTRDNRLLRPKAVYVGELDRKYVPIDQRLEHHHHHH CS4-His sequence number 18 MASMEYTPGLAGVVAAESSISYIDGQEGILRYRGYPIEELAEHSTFEEVAYLLLFGELPTRDELEEFDHELKHHRALPERIIDLLKNLPKSAHPMALQSAVAALGMFYPDADVTDPEGNYEAAVRLIAKLPTIVAAFHRIRRGQDPIAPRDDLGHAANFLYMLNGEEDPPLAARVFDVCLILHAEHSMNAST FTARVVGSTLADPYSAIAAAIGSLSGPLHGGANEEVLQMLEEIGSPDNVEPWLEEKLARKEKIMGFGHRVYKVKDPRATILQKMAEQLFEKHGSTPLYDIALELEKVAAERLGPKGIYPNVDFYSGIVYQKMGIPDTLFTPIFAIARVAGWTAHWLEQLEDNRIFRPSQIYVGPTDRSYVPIDERLEHHHHHH CS5-His sequence number 19 MYTPGLAGVVAAESSISYIDGQEGILRYRGYPIEELAEHSSFEEVAYLLLYGRLPTREELEEFDHELKHHRAIPEGIIDLLKTLPKSAHPMALQSAVAALGMFYPEADDVEDPEANYKKAVRLIAKLPTIVAAFHRIRRGKEPVAPRSDLSHAANFLYMNLNGEEDPPLAARVMDVCLILHAEHEMNASTF TARVVGSTLADPYSAIAAAIGALSGPLHGGANEEVLKMLQEIGSADNVEPWLEEKLATKQKIMGGFHRVYKTYDPRAKILKKMAEQLFEKHGSTPLYDIAVELEKVAAERLGPKGIYPNVDFYSGLVYQALGIPTDLFTPIFAIARWAGWTAHWLEQLEDNRIFRPRQVYVGPTDRKYVPIDQRLEHHHHHH CS6-His sequence number 20 MIAKGLEGVVIAETSISYIDGQEGRLYYRGYPIEELAEHSSFEEVAYLLWHGRLPTREELEEFKEELAKNRAIPEIIDLRTLPKSAHPMALRTAVSALGMFDPDADDTQSPEARYRKAIRLIAKIPTIVAAFHRIRQGQEPVAPRPDLSHAANFLYMNLNGEESPVQAKVMDVALILHAEHEMNASTF AALVVASTLSDMYSAITAAIGALKPLHGGANEQVMKMLQEIGSPDKAEPWVQEKLANKRRIMGFGHRVYKTYDPRAKILKKYARQLAEKQGDTTTLYEIAEAVEKVVVERLGPKGIYPNVDFYSGLVYHALGIPTELFTPIFAMARVSGWTAHVLEYLKDNRLIRPRAVYVGPTDRKYVPIDQRLEHHHHHH CS7-His sequence number 21 MYTPGLAGVVAAESSISYIDGQEGILRYRGYPIEELAEHSSFEEVAYLLLFGKLPTREELEEFDHELKHHRAIPEGIIDLLKTLPKSAHPMALQSAVAALGMFYPADVEDPEANYEAAVRLIAKLPTIVAAFHRIRRGKDPIAPRSDLGHAANFLYMLNGEEPDPLAARVMDVCLILHAEHSMNASTFT ARVVGSTLADPYSAIAAAIGSLSGPLHGGANEEVLQMLQEIGSAENVEPWLEEKLATKQKIMGFGHRVYKVKDPRATILQKMAEQLFEKHGSTPLYDIAVELEKVAAERLGPKGIYPNVDFYSGLVYQKLGIPTDLFTPIFAIARWAGWTAHWLEQLEDNRIFRPTQIYVGPTDRSYVPIDQRLEHHHHHH
[0078]
table 1
[0079]
table 2-1
[0080] [Table 2-2]
[0081] [Table 2-3]
[0082] [Table 3] Example 2 - Expression and purification of citrate synthase protein BL21(DE3) cells were transformed with pET vectors containing sequence numbers 22-28, and these cells were cultured at 30°C with shaking in a 2L baffled flask containing 500mL of LB medium supplemented with 6.25g of lactose and 100μg / mL of carbenicillin. The cells were cultured for approximately 16 hours, after which they were harvested by centrifugation at 4500×g for 15 minutes.
[0083] The cells were resuspended in 20 mM Tris:HCl, 300 mM NaCl, and 20 mM imidazole pH 8.0 buffer (Buffer A) at a ratio of 30 mL per liter of the original cell culture. The cells were lysed using a microfluidizer at 15,000 psi for 3 cycles. The lysed cells were ultracentrifuged at 30,000 × g for 30 minutes to remove membranes and cell fragments. The clarified lysate was then filtered through a 0.45 micrometer (micron) syringe filter.
[0084] The lysate was loaded onto a nickel NTA pre-packed column pre-equilibrated with buffer A. The binding protein was washed with 7 column volumes of buffer A, and 7 column volumes of 20 mM Tris:HCl, 300 mM NaCl, and 68 mM imidazole pH 8.0 (buffer B). The binding protein was then eluted with 20 mM Tris:HCl, 300 mM NaCl, and 500 mM imidazole pH 8.0 (buffer C). The eluted protein was transferred to 20 mM Tris:HCl, 200 mM NaCl, pH 7.5 (buffer X) or phosphate-buffered saline (PBS) using a PD-10 column.
[0085] The proteins are pure enough to be used as molecular weight standards / mass photometry calibrators, can be frozen in liquid nitrogen, and can be stored for extended periods at -20°C. SDS-PAGE (Figure 4) and Native PAGE (Figure 5) were performed on the samples.
[0086] The obtained protein can be enriched to a high molecular weight complex-forming state using SEC. This step can be performed on the protein after elution from the nickel NTA column or after the step of transferring to buffer X.
[0087] The analytical SEC column Enrich 650 (Biorad) was equilibrated with buffer X. Protein samples were concentrated to 5-25 mg / mL protein and 250 μL was loaded onto the column. The samples were eluted at 1 mL / min. The samples eluted as a first broad peak encompassing high molecular weight (HMW) complex formation states (trimers or more), and a second sharp peak belonging to low molecular weight (LMW) complex formation states (dimers or less), but the two peaks could not be completely separated from each other.
[0088] Example 3 - Mass photometry measurement of citrate synthase protein A reusable silicone gasket (CultureWell®, CW-50R-1.0, 50-3mm diameter x 1mm depth) was placed on a clean microscope coverslip (1.5H, 24 x 60mm, Carl Roth) and mounted on the stage of a mass photometer (Refeyn Ltd., UK) using immersion oil (Immersol 518F, Zeiss). The gasket was packed with 19 μL of buffer (PBS or buffer C) and the instrument was focused. The protein was pre-diluted to a concentration of approximately 400 nM in the same buffer. Then, 1 μL of the pre-diluted protein solution was added to a drop of buffer and thoroughly mixed. The final protein concentration at the time of measurement was 20 nM. Data was acquired for 60 seconds at 100 frames per second.
[0089] [Table 4] Example 4 - Batch-to-batch reproducibility To ensure batch-to-batch consistency of mass photometry data, several further purifications of CS1-His were prepared and tested as described above. Mass photometry (MP) experiments were conducted using Acquire MP Two using software MP The procedure was performed using a mass photometer (Refeyn Ltd, UK). Figures 6a–d show mass photometry plots for 20 nM samples of CS1-His purified by two separate purifications derived from two cell lysates. All of these samples show the same peaks, with only small differences in the ratio of peak areas between batches. All of these samples are suitable for use as mass photometry standards.
[0090] Example 5 - Use of CS1-His as a mass photometry standard for proteins of known molecular weight Using CS1-His, the Two mentioned above MPThe mass photometer was calibrated, and proteins of known molecular weight were analyzed by mass photometry. Sample concentrations ranged from 10 to 40 nM depending on the protein being analyzed. Figure 7 shows that the masses of proteins determined by mass photometry, when calibrated using CS1-His, are shown next to the masses predicted from the protein's primary sequence. The results demonstrate that CS1-His calibration yields calculated masses within 2% of the predicted masses, indicating the validity of the calibration.
[0091] Example 6 - Comparison with NativeMark (trademark) as a calibration standard. Prior to the development of the citrate synthase protein of the present invention, the calibration standard commonly used in mass photometry was NativeMark® unstained protein standard (Invitrogen®). NativeMark®, designed for use as a standard for native PAGE, is advertised as containing eight proteins (1236, 1048, 720, 480, 242, 146, 66, and 20 kDa).
[0092] NativeMark® was compared to CS1-His HMW for calibration usability and accuracy. NativeMark® was diluted 250-fold in PBS, while CS1-His HMW was diluted to 40 nM (monomer concentration). After setting up the mass photometry device, 10 μL of the diluted protein was added to the well, and a 1-minute video was recorded. The data was then used in Discover MP The analysis was performed using software (Refeyn®). A comparison between NativeMark® and CS1-His HMW was conducted on the same day.
[0093] Figures 8a and 8b show a comparison of results from the two standards. Clearly, the peaks obtained with NativeMark™ are heavily biased towards a single species, accounting for 67% of the binding events. As a result, the low-intensity peaks have a low signal-to-noise ratio and are less accurate. The error in the mass / radiometry contrast gradient for NativeMark™ is 5.8%, while for the dataset shown for CS1-His HMW, the error is only 0.7%. Higher intensity peaks can be achieved using more concentrated samples with NativeMark™, and thus more accurate data can be achieved for higher molecular weight species, but this adds complexity, time, and cost.
[0094] These results are not particularly surprising, because NativeMark™ is not designed and formulated for use as a mass photometry standard, but rather to achieve nearly identical staining intensity when stained after performing native PAGE.
[0095] References Guindon S. & Gascuel O. PhyML : "A simple, fast, and accurate algorithm to estimate large phylogenies by maximum likelihood." Syst. Biol. 52(5):696-704 (2003) DOI: 10.1080 / 10635150390235520). Harms, M., Thornton, J. Evolutionary biochemistry: revealing the historical and physical causes of protein properties. Nat Rev Genet 14:559-571 (2013). https: / / doi.org / 10.1038 / nrg3540 Harms, M.J., Thornton, J.W. Analyzing protein structure and function using ancestral gene reconstruction. Curr. Opin. Struct. Biol.:20, 360-366. (2010). https: / / doi.org / 10.1016 / j.sbi.2010.03.005. Needleman, S.B., Wunsch, C.D. A general method applicable to the search for similarities in the amino acid sequence of two proteins. JMB48:443-453 (1970). https: / / doi.org / 10.1016 / 0022-2836(70)90057-4 Selberg, A.G.A., Gaucher, E.A. & Liberles, D.A. Ancestral Sequence Reconstruction: From Chemical Paleogenetics to Maximum Likelihood Algorithms and Beyond. J Mol Evol 89:157-164 (2021). https: / / doi.org / 10.1007 / s00239-021-09993-1 Schmidtmann et al., Redox Regulation of Arabidopsis Mitochondrial Citrate Synthase. Mol. Plant 7:156-159 (2014) doi.org / 10.1093 / mp / sst144. Stamatakis A. Randomized Axelerated Maximum Likelihood Bioinformatics. 30(9):1312-3 (2014). doi: 10.1093 / bioinformatics / btu033. Young G, Hundt N, Cole D, et al., Quantitative mass imaging of single biological macromolecules. Science. 360:423-427 (2018). doi:10.1126 / science.aar5839 The above description provides several methods, proteins, and nucleic acids of the present invention. The present invention is modifiable in terms of methods and materials. Such modifications will be obvious to those skilled in the art by considering this application or practice of the invention provided herein. Accordingly, the present invention is not intended to be limited to the specific embodiments provided herein, but is intended to encompass all modifications and substitutions that fall within the true scope and spirit of the invention, to be embodied in the appended claims. All patents, applications, and other references cited herein are incorporated herein by reference in their entirety.
[0096] Clauses: A. A citrate synthase protein or a functional fragment thereof comprising at least one ancestral mutation, wherein the protein can simultaneously form at least two, at least three, at least four, or at least five complex-forming states.
[0097] B. The citrate synthase protein described in Clause A, comprising at least two, at least three, at least four, at least five, at least ten, at least fifteen, or at least twenty ancestral mutations.
[0098] C. A citrate synthase protein or a functional fragment thereof having an amino acid sequence that is at least 75%, 80%, 85%, 90%, or 95% identical to a sequence selected from Sequence ID No. 1 to 7, wherein the protein can simultaneously form at least two, at least three, at least four, or at least five complex-forming states.
[0099] D. A citrate synthase protein or a functional fragment thereof having an amino acid sequence having at least 75%, 80%, 85%, 90%, or 95% sequence similarity to a sequence selected from Sequence ID No. 1 to 7, wherein the protein can simultaneously form at least two, at least three, at least four, or at least five complex-forming states.
[0100] E. A citrate synthase protein or a functional fragment thereof, comprising or consisting of an amino acid sequence selected from any of Sequence IDs 1-7. F. The citrate synthase protein described in clause E, comprising the sequence of sequence number 1.
[0101] G. Citrate synthase protein as described in any of clauses A through F, further including tags to facilitate purification. H. The citrate synthase protein described in clause G, wherein the purified tag is a His tag.
[0102] I. A citrate synthase protein described in any of clauses A through H, which is not expressed by any existing organism. J. A nucleic acid sequence encoding the citrate synthase protein described in clauses A to I.
[0103] A nucleic acid sequence encoding a K. citrate synthase protein or a functional fragment thereof, wherein the nucleic acid sequence is an ancestral gene, and the citrate synthase protein encoded by the nucleic acid sequence can simultaneously form at least two, at least three, at least four, or at least five complex-forming states.
[0104] A citrate synthase protein or a functional fragment thereof, encoded by the nucleic acid sequence described in Clause L, K. N. A nucleic acid sequence according to clause J or K, which is codon-optimized for expression in a living organism, and preferably codon-optimized for expression in Escherichia coli.
[0105] A vector comprising a nucleic acid sequence as described in any of clauses J, K, or N. The vector described in clause O, which is a P plasmid, preferably a pET plasmid.
[0106] Q. A cell comprising a nucleic acid sequence or vector as described in any of clauses K, N, or O. R. Escherichia coli cells, as described in clause Q.
[0107] S. Escherichia coli cells as described in clause R, expressing the citrate synthase protein described in any of clauses A to I or L. A method for producing a citrate synthase protein as described in either clause A to I or clause L, comprising the steps of culturing a bacterium as described in clause R, and performing at least one purification step.
[0108] A composition comprising the citrate synthase protein described in any of clauses A to I or L, wherein the concentration of the citrate synthase protein is at least 5 nM, preferably 5 nM to 100 μM.
[0109] V. The composition according to clause U, further comprising at least one component selected from buffer reagents and salts. Use of citrate synthase proteins described in clauses A to I or L, or compositions described in clause U or V, as molecular weight standards.
[0110] Use of the citrate synthase protein or composition described in clause W as a molecular weight standard, and as a calibrator for mass photometry. Y. Use of the citrate synthase protein or composition described in Clause W as a molecular weight standard, for use as a molecular weight standard in native polyacrylamide gel electrophoresis.
[0111] Z. A method for calibrating a mass photometry device, comprising the step of assaying a composition described in clause W or clause X.
Claims
1. A citrate synthase protein or a functional fragment thereof encoded by a nucleic acid sequence, wherein the nucleic acid is an ancestral gene, and the citrate synthase protein or its functional fragment can simultaneously form at least three complex formation states.
2. At least three complex formation states can be formed simultaneously. i) An amino acid sequence that is at least 75%, 80%, 85%, 90%, or 95% identical to a sequence selected from Sequence IDs 1 to 7, ii) An amino acid sequence having at least 75%, 80%, 85%, 90%, or 95% sequence similarity to a sequence selected from Sequence IDs 1 to 7, iii) An amino acid sequence selected from any of SEQ ID NOs: 1 to 7, and / or iv) Amino acid sequence containing at least one ancestral mutation A citrate synthase protein or a functional fragment thereof having an amino acid sequence selected from the above.
3. A citrate synthase protein or a functional fragment thereof according to claim 1 or 2, which can simultaneously form at least four complex-forming states.
4. The citrate synthase protein or a functional fragment thereof according to any one of claims 1 to 3, wherein each of the three most abundant complex-forming states constitutes at least 15% of the total number of citrate synthase protein complexes present.
5. The citrate synthase protein or a functional fragment thereof according to any one of claims 1 to 4, further comprising a tag, preferably a His tag, for the purpose of facilitating purification.
6. A citrate synthase protein or a functional fragment thereof according to any one of claims 1 to 5, which is not expressed by existing organisms.
7. A composition comprising a citrate synthase protein or a functional fragment thereof according to any one of claims 1 to 6, wherein the concentration of the citrate synthase protein or a functional fragment thereof is at least 5 nM, and optionally the concentration is 5 nM to 100 μM.
8. Use of the citrate synthase protein or a functional fragment thereof according to any one of claims 1 to 6, or the composition according to claim 7, as a molecular weight standard.
9. Use of the citrate synthase protein, a functional fragment thereof, or a composition according to claim 8 as a molecular weight standard, for use as a calibrator for mass photometry.
10. A method for calibrating a mass photometry device, comprising the step of detecting the mass of particles in the composition according to claim 7 using the mass photometry device.
11. A nucleic acid sequence encoding a citrate synthase protein or a functional fragment thereof according to any one of claims 1 to 6.
12. The nucleic acid sequence according to claim 11, which is codon-optimized for expression in a living organism, and preferably codon-optimized for expression in Escherichia coli.
13. A vector comprising the nucleic acid sequence described in claim 11 or claim 12, preferably a pET plasmid.
14. A cell, preferably an Escherichia coli cell, containing the nucleic acid sequence described in claim 11 or claim 12 or the vector described in claim 13.
15. The cell according to claim 14, expressing the citrate synthase protein or a functional fragment thereof according to any one of claims 1 to 6.
16. A method for producing a citrate synthase protein or a functional fragment thereof according to any one of claims 1 to 6, comprising the steps of culturing cells according to claim 14 or 15, and performing at least one purification step.