PROTEIN CONTAINING gfasCP MUTANT SEQUENCE
Novel gfasCP mutants with altered amino acids at positions 61 and 62 provide a shorter absorption wavelength, improving color-based markers and FRET efficiency for biological applications.
Patent Information
- Application Number
- JP2024094243
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-11
- Publication Date
- 2025-12-23
AI Technical Summary
There is a lack of gfasCP mutants with a shorter maximum absorption wavelength compared to the wild-type, limiting their applications as color-based markers and fluorescence resonance energy transfer (FRET) acceptors.
Development of gfasCP mutant sequences with specific amino acid modifications at positions 61 and 62, such as replacing serine and glutamine with hydrophobic or polar uncharged amino acids, resulting in a maximum absorption wavelength of 578 nm or less.
The novel gfasCP mutants exhibit a shorter absorption wavelength, enhancing color variation for cell labeling and increasing FRET efficiency, suitable for various biological applications.
Smart Images

Figure 2025185823000002 
Figure 2025185823000003 
Figure 2025185823000004
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to proteins comprising gfasCP mutant sequences. [Background technology]
[0002] gfasCP is a chromoprotein derived from the coral Galaxea fascicularis. It has a molar extinction coefficient of over 200,000 [L / mol cm], which is larger than that of other chromoproteins. Its maximum absorption wavelength is 577 nm, and it appears purple to the naked eye (Non-Patent Document 1).
[0003] To date, several gfasCP mutants with different absorption spectral shapes have been reported, primarily as part of basic research into the relationship between amino acid sequence and absorption wavelength. For example, Non-Patent Document 1 discloses that mutants with S61C, I154L, and / or S175T mutations introduced into gfasCP have a longer maximum absorption wavelength and exhibit a bluish color when visually observed. Non-Patent Document 2 also discloses gfasCP mutants with various mutations introduced that have absorption spectral shapes different from those of wild-type gfasCP. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] International Publication No. 2023 / 182169 [Non-patent literature]
[0005] [Non-Patent Document 1] Naila O. Alieva et al., "Diversity and Evolution of Coral Fluorescent Proteins", PLoS ONE 3(7): e2680 (2008). [Non-patent document 2] F Hafna Ahmed et al., "Overthe rainbow: structural characterization of the chromoproteins gfasPurple, amilCP, spisPink and eforRed", Acta Crystallogr D Struct Biol. 78(Pt5):599-612 (2022). Summary of the Invention [Problem to be solved by the invention]
[0006] Because gfasCP and its mutants have a large molar extinction coefficient, they are expected to be used as markers that use coloration when expressed in cells as an indicator, or as acceptors for fluorescence resonance energy transfer (FRET) (Non-Patent Document 2). When using gfasCP mutants for these applications, the greater the color variation available for color-based markers, the more proteins and cells can be distinguishably labeled. Furthermore, when used as FRET acceptors, FRET efficiency is increased when gfasCP mutants with an absorption wavelength appropriate for the donor's fluorescence wavelength are used. Therefore, if gfasCP mutants with a variety of absorption wavelengths can be developed, the breadth and quality of applications for gfasCP mutants are expected to be further enhanced.
[0007] Many of the gfasCP mutants reported to date have a longer maximum absorption wavelength compared to the wild-type, resulting in a visually shifted color toward blue. In contrast, there are few reported examples of gfasCP mutants with a shorter maximum absorption wavelength compared to the wild-type, resulting in a visually shifted color toward red. Furthermore, the few reported examples of gfasCP mutants (Y116H, Y116H+E73D) that exhibit a reddish color, as disclosed in Non-Patent Document 2, have an extremely pale color.
[0008] In view of these circumstances, an object of the present disclosure is to provide a protein comprising a novel gfasCP mutant sequence and a protein comprising a novel gfasCP mutant sequence whose maximum absorption wavelength is shortened compared to that of wild-type gfasCP. [Means for solving the problem]
[0009] The present disclosure relates, for example, to the following: [1] A gfasCP mutant sequence having an amino acid sequence having 90% or more sequence identity with the amino acid sequence shown in SEQ ID NO: 1; In the gfasCP mutant sequence, the amino acid corresponding to the tyrosine at position 63 in the amino acid sequence shown in SEQ ID NO: 1 is tyrosine, The gfasCP mutant sequence satisfies at least one of the following (A) to (C): (A) the amino acid corresponding to the serine at position 61 in the amino acid sequence shown in SEQ ID NO: 1 is a hydrophobic amino acid, and the amino acid corresponding to the glutamine at position 62 in the amino acid sequence shown in SEQ ID NO: 1 is a polar, uncharged amino acid; (B) the amino acid corresponding to the serine at position 61 in the amino acid sequence shown in SEQ ID NO: 1 is a hydrophobic amino acid, and the amino acid corresponding to the glutamine at position 62 in the amino acid sequence shown in SEQ ID NO: 1 is a hydrophobic amino acid; (C) The amino acid corresponding to the serine at position 61 in the amino acid sequence shown in SEQ ID NO: 1 is a polar, uncharged amino acid, and the amino acid corresponding to the glutamine at position 62 in the amino acid sequence shown in SEQ ID NO: 1 is a polar, uncharged amino acid other than glutamine. [2] The hydrophobic amino acid is alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, glycine, or cysteine; The protein according to [1], wherein the polar uncharged amino acid is asparagine, glutamine, serine, threonine, or tyrosine. [3] The protein according to [1] or [2], wherein in the gfasCP mutant sequence, the amino acids corresponding to the 61st serine and the 62nd glutamine in the amino acid sequence shown in SEQ ID NO: 1 are, in the order of 61st to 62nd, leucine-serine, phenylalanine-threonine, isoleucine-serine, leucine-leucine, phenylalanine-cysteine, valine-glycine, serine-serine, alanine-glycine, leucine-glycine, or tyrosine-serine. [4] A protein described in any one of [1] to [3], wherein the gfasCP mutant sequence comprises an amino acid sequence having 90% or more sequence identity with the amino acid sequence shown at positions 1 to 60 of SEQ ID NO: 1, and an amino acid sequence having 90% or more sequence identity with the amino acid sequence shown at positions 64 to 221 of SEQ ID NO: 1. [5] The protein according to any one of [1] to [4], which has a β-barrel containing 11 β-strands as a higher-order structure. [6] The protein according to any one of [1] to [5], which forms an oligomer in an aqueous solution at pH 8.0. [7] The protein according to any one of [1] to [6], which has a maximum absorption wavelength of 578 nm or less in an aqueous solution at pH 8.0. [8] The molar absorption coefficient at the maximum absorption wavelength in an aqueous solution of pH 8.0 is 5.0 × 10 4 The protein according to any one of [1] to [7], wherein the protein has a molecular weight of [L / mol·cm] or more. [9] The protein according to any one of [1] to [8], wherein the gfasCP mutant sequence is the amino acid sequence shown in SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10 or 11.
[10] The protein according to any one of [1] to [9], which consists of the gfasCP mutant sequence.
[11] A nucleic acid comprising a sequence encoding the protein according to any one of [1] to
[10] .
[12] A vector for expressing the protein according to any one of [1] to
[10] or a fusion protein thereof in a cell.
[13] A cell expressing the protein according to any one of [1] to
[10] or a fusion protein thereof.
[14] A method for evaluating a cell, comprising detecting a signal derived from the protein or a fusion protein thereof according to any one of [1] to
[10] in a cell expressing the protein or a fusion protein thereof. [Effects of the Invention]
[0010] The present disclosure provides proteins comprising novel gfasCP mutant sequences, and also provides proteins comprising novel gfasCP mutant sequences that have a shorter maximum absorption wavelength than wild-type gfasCP. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 shows a photograph of the wild-type gfasCP solution obtained in Example 1, taken under room light. [Figure 2] FIG. 1 shows a photograph taken under room light of a 10 cm diameter dish in which transformants were inoculated onto LB agar medium in Example 2. [Figure 3] FIG. 1 shows the base sequences of pMGFE003, pMGFE005, and pMGFE006 in Example 2 near the sites where random mutations were introduced, and the underlined parts are the positions where random mutations were introduced. [Figure 4] FIG. 4 shows the results of identifying the amino acid sequences near the sites where random mutations were introduced in pMGFE003, pMGFE005, and pMGFE006 from the results of FIG. 3 in Example 2. [Figure 5] FIG. 1 shows a photograph taken under room light of a 10 cm diameter dish in which transformants were inoculated onto LB agar medium in Example 3. [Figure 6] FIG. 1 shows the base sequences near the sites where random mutations were introduced in pMGFE014, pMGFE015, pMGFE016, pMGFE020, pMGFE022, pMGFE025, and pMGFE027 in Example 3, where the underlined parts are the positions where random mutations were introduced. [Figure 7] FIG. 7 shows the results of identifying the amino acid sequences near the sites where random mutations were introduced in pMGFE014, pMGFE015, pMGFE016, pMGFE020, pMGFE022, pMGFE025, and pMGFE027 from the results of FIG. 6 in Example 3. [Figure 8] FIG. 1 shows photographs taken under room light of a solution of wild-type gfasCP and solutions of gfasCP mutants obtained from pMGFE003, pMGFE005, pMGFE006, pMGFE014, pMGFE015, pMGFE016, pMGFE020, pMGFE022, pMGFE025, and pMGFE027 in Example 4. [Figure 9] FIG. 1 shows the results of polyacrylamide gel electrophoresis of the soluble fraction or gfasCP mutant solution (eluted fraction) obtained by centrifugation after bacterial cell culture in Example 4. [Figure 10] 1 shows the absorption spectra of wild-type gfasCP and gfasCP mutants. [Figure 11] Absorption spectrum of the gfasCP mutant. DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, embodiments for carrying out the present disclosure will be described, but the present disclosure is not limited to the following embodiments.
[0013] In the present disclosure, when a protein or nucleic acid comprises an amino acid sequence or a nucleotide sequence that has 90% or more sequence identity with a given amino acid sequence or a nucleotide sequence, the protein or nucleic acid may have 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity with the given sequence, or in a preferred embodiment, 95% or more sequence identity, and in a most preferred embodiment, 100% sequence identity.
[0014] In the present disclosure, when a sequence contained in a certain protein or nucleic acid has a mutation (i.e., sequence identity is not 100%) with respect to a predetermined amino acid sequence or nucleotide sequence, the mutation may be a mutation selected from substitution, deletion, insertion, and addition in 1 to 20 consecutive or dispersed residues or 1 to 60 bases. In a preferred embodiment, the mutation may be a mutation selected from substitution, deletion, insertion, and addition in 1 to 10 consecutive or dispersed residues or 1 to 30 bases. In a more preferred embodiment, the mutation may be a mutation selected from substitution, deletion, insertion, and addition in 1 to 3 residues or 1 to 10 bases. In an even more preferred embodiment, the mutation may be a mutation selected from substitution, deletion, insertion, and addition in 1 residue or 1 to 3 bases.
[0015] In the present disclosure, peptides and proteins are described from the N-terminus to the C-terminus unless otherwise specified. In the present disclosure, polynucleotides are described from the 5'-terminus to the 3'-terminus unless otherwise specified. In the present disclosure, natural amino acids are described in the L-form unless otherwise specified. For example, in the present disclosure, the terms "alanine" and "L-alanine" refer to L-alanine, and the term "D-alanine" refers to D-alanine. In the present disclosure, amino acids may be natural or unnatural amino acids, and in one aspect, may be natural amino acids.
[0016] In the present disclosure, a "protein consisting of the amino acid sequence shown in <SEQ ID NO>" refers to a protein whose amino acid residues, from the N-terminus, are the amino acid sequence shown in <SEQ ID NO>, and includes not only molecules having each of those amino acids themselves as residues, but also molecules that are obtained by spontaneous molecular structural changes (e.g., ring formation) that occur in a living body after being biosynthesized as a polypeptide having that amino acid sequence.
[0017] A first embodiment of the present disclosure is a protein comprising a gfasCP mutant sequence. In a preferred aspect, the protein according to the first embodiment is a protein consisting of L-amino acid residues, in a more preferred aspect, a protein consisting of L-α-amino acid residues, and in an even more preferred aspect, a protein consisting of natural amino acid residues. In the present disclosure, proline is treated as an α-amino acid.
[0018] gfasCP (also known as gfasPurple) is a chromoprotein derived from the coral Galaxea fascicularis. The amino acid sequence of gfasCP is shown in SEQ ID NO: 1, and its accession number in Genpept, a protein database maintained by the National Center for Biotechnology Information (NCBI) of the National Institutes of Health (NIH), is ABB17967. The nucleotide sequence of the mRNA encoding gfasCP in Galaxea fascicularis is shown in SEQ ID NO: 12, and its accession number in GenBank, a genome database maintained by the NIH's NCBI, is DQ206394. SEQ ID NO: 1: MSVIAKQMTYKVYMSGTVNGHYFEVEGDGKGKPYEGEQTVKLTVTKGGPLPFAWDILSPQSQYGSIPFTKYPEDIPDYVKQSFPEGYTWERIMNFEDGAVCTVSNDSSIQGNCFIYHVKFSGLNFPPNGPVMQKKTQGWEPNTERLFARDGMLIGNNFMALKLEGGGHYLCEFKSTYKAKKPVKMPGYHYVDRKLDVTNHNKDYTSVEQCEISIARKSVVA
[0019] The gfasCP mutant sequence according to the present disclosure is an amino acid sequence having 90% or more sequence identity with the amino acid sequence shown in SEQ ID NO: 1. That is, the gfasCP mutant sequence according to the present disclosure has 90% or more sequence identity with the amino acid sequence of gfasCP.
[0020] In the gfasCP mutant sequence of the present disclosure, the amino acid corresponding to the tyrosine at position 63 in the amino acid sequence shown in SEQ ID NO: 1 is tyrosine. That is, in the gfasCP mutant sequence, the tyrosine at position 63 in the amino acid sequence of gfasCP is not mutated. The present inventors have found that when the amino acid corresponding to the tyrosine at position 63 in the amino acid sequence shown in SEQ ID NO: 1 is tyrosine, the visible light absorption of the gfasCP mutant sequence is likely to be maintained and the maximum absorption wavelength in the visible light region may be shorter than that of wild-type gfasCP. In one embodiment, the gfasCP mutant sequence of the present disclosure may further include glycine at the amino acid corresponding to the glycine at position 64 in the amino acid sequence shown in SEQ ID NO: 1. In one aspect, the gfasCP mutant sequence of the present disclosure may further comprise amino acids at positions 57 to 60, 58 to 60, 59 to 60 or 60 and / or 63 to 67, 63 to 66, 63 to 65, 63 to 64 or 63 of the amino acid sequence shown in SEQ ID NO: 1, which are identical to the corresponding positions in the amino acid sequence shown in SEQ ID NO: 1.
[0021] The gfasCP mutant sequence according to the present disclosure satisfies at least one of the following conditions (A) to (C). The present inventors have found that when a gfasCP mutant sequence satisfies at least one of conditions (A) to (C), the maximum absorption wavelength in the visible light region becomes shorter than that of wild-type gfasCP. The gfasCP mutant sequence according to one embodiment of the present disclosure satisfies one of conditions (A) to (C). Hereinafter, "serine at position 61 in the amino acid sequence shown in SEQ ID NO: 1" will also be referred to as S61, and "glutamine at position 62 in the amino acid sequence shown in SEQ ID NO: 1" will also be referred to as Q62. (A) The amino acid corresponding to the serine at position 61 in the amino acid sequence shown in SEQ ID NO: 1 is a hydrophobic amino acid, and the amino acid corresponding to the glutamine at position 62 in the amino acid sequence shown in SEQ ID NO: 1 is a polar, uncharged amino acid. (B) The amino acid corresponding to the serine at position 61 in the amino acid sequence shown in SEQ ID NO: 1 is a hydrophobic amino acid, and the amino acid corresponding to the glutamine at position 62 in the amino acid sequence shown in SEQ ID NO: 1 is a hydrophobic amino acid. (C) The amino acid corresponding to the serine at position 61 in the amino acid sequence shown in SEQ ID NO: 1 is a polar, uncharged amino acid, and the amino acid corresponding to the glutamine at position 62 in the amino acid sequence shown in SEQ ID NO: 1 is a polar, uncharged amino acid other than glutamine.
[0022] In the present disclosure, a polar group refers to a group that has an electron-withdrawing structure in part or in its entirety. Such a group causes a charge imbalance (polarization) in the side chain. Examples of polar groups include hydroxyl, carboxyl, amino, carbonyl, amido, oxycarbonyl, sulfo, and phospho. In one embodiment, when the amino acid is a natural amino acid, polar groups that can be contained in its side chain are hydroxyl, carboxyl, amino, carbonyl, and amido.
[0023] In the present disclosure, a hydrophobic amino acid refers to an amino acid that does not have a polar group in its side chain and that has an uncharged side chain at pH 7.4 (physiological pH). In one embodiment, a hydrophobic amino acid may refer to an amino acid that does not contain an atom with an electronegativity of 3.10 or higher (i.e., a chlorine atom, an oxygen atom, or a fluorine atom) in its side chain and that has an uncharged side chain at pH 7.4. In one embodiment, when the hydrophobic amino acid is a natural amino acid, the hydrophobic amino acid may refer to an amino acid that does not contain a hydroxyl, carboxyl, amino, carbonyl, or amide in its side chain and that has an uncharged side chain under physiological conditions. In one embodiment, the hydrophobic amino acid may be alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, glycine, or cysteine.
[0024] In the present disclosure, an acidic amino acid refers to an amino acid whose side chain can be negatively charged at pH 7.4. The fact that the side chain can be negatively charged at pH 7.4 is due to the pK aBased on this, it can be determined that at pH 7.4, 1.0% or more of the molecules have a negative total charge on their side chains. In one embodiment, the acidic amino acid may be an amino acid containing a carboxyl group in the side chain. In one embodiment, the acidic amino acid may be aspartic acid or glutamic acid.
[0025] In the present disclosure, a basic amino acid refers to an amino acid whose side chain can be positively charged at pH 7.4. The fact that the side chain can be positively charged at pH 7.4 is due to the pK a Based on this, it can be determined that at pH 7.4, 1.0% or more of the molecules have a positive total charge on their side chains. In one embodiment, the basic amino acid may be an amino acid containing an amino or imidazolyl group in the side chain. In one embodiment, the acidic amino acid may be arginine, histidine, or lysine.
[0026] In the present disclosure, a polar uncharged amino acid refers to an amino acid that has a polar group in its side chain and that has an uncharged side chain at pH 7.4. In one embodiment, a polar uncharged amino acid may be an amino acid that does not belong to any of the categories of hydrophobic amino acids, acidic amino acids, and basic amino acids. In one embodiment, a polar uncharged amino acid may refer to an amino acid that contains an atom in its side chain with an electronegativity of 3.10 or more (i.e., a chlorine atom, an oxygen atom, or a fluorine atom) and that has an uncharged side chain at pH 7.4. In one embodiment, a polar uncharged amino acid may be asparagine, glutamine, serine, threonine, or tyrosine.
[0027] In one aspect, the gfasCP mutant sequence may satisfy one of the following conditions (A') to (C'): Examples of such aspects include when the protein of the first embodiment is a protein consisting of L-amino acid residues, a protein consisting of L-α-amino acid residues, or a protein consisting of naturally occurring amino acid residues. (A') The amino acid corresponding to the 61st serine in the amino acid sequence shown in SEQ ID NO: 1 is alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, glycine, or cysteine, and the amino acid corresponding to the 62nd glutamine in the amino acid sequence shown in SEQ ID NO: 1 is asparagine, glutamine, serine, threonine, or tyrosine. (B') The amino acid corresponding to the 61st serine in the amino acid sequence shown in SEQ ID NO: 1 is alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, glycine, or cysteine, and the amino acid corresponding to the 62nd glutamine in the amino acid sequence shown in SEQ ID NO: 1 is alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, glycine, or cysteine. (C') The amino acid corresponding to the 61st serine in the amino acid sequence shown in SEQ ID NO: 1 is asparagine, glutamine, serine, threonine or tyrosine, and the amino acid corresponding to the 62nd glutamine in the amino acid sequence shown in SEQ ID NO: 1 is asparagine, serine, threonine or tyrosine.
[0028] In one embodiment of the gfasCP mutant sequence, the amino acids corresponding to the serine at position 61 and the glutamine at position 62 in the amino acid sequence shown in SEQ ID NO: 1 may be, in order from position 61 to position 62, leucine-serine, phenylalanine-threonine, isoleucine-serine, leucine-leucine, phenylalanine-cysteine, valine-glycine, serine-serine, alanine-glycine, leucine-glycine, or tyrosine-serine.
[0029] In one embodiment, the gfasCP mutant sequence may comprise an amino acid sequence having 90% or more sequence identity with the amino acid sequence of positions 1 to 60 of SEQ ID NO: 1. In one embodiment, the gfasCP mutant sequence may have amino acids 1 to 60 of the amino acid sequence of SEQ ID NO: 1 substituted with an amino acid sequence having 90% or more sequence identity with the amino acid sequence of positions 1 to 60 of the amino acid sequence of SEQ ID NO: 1.
[0030] In one embodiment, the gfasCP mutant sequence may comprise an amino acid sequence having 90% or more sequence identity with the amino acid sequence of positions 64 to 221 of SEQ ID NO: 1. In one embodiment, the gfasCP mutant sequence may have amino acids 64 to 221 of the amino acid sequence of SEQ ID NO: 1 substituted with an amino acid sequence having 90% or more sequence identity with the amino acid sequence of positions 64 to 221 of the amino acid sequence of SEQ ID NO: 1.
[0031] In one embodiment, the gfasCP mutant sequence may comprise an amino acid sequence having 90% or more sequence identity to the amino acid sequence set forth at positions 1 to 60 of SEQ ID NO: 1, and an amino acid sequence having 90% or more sequence identity to the amino acid sequence set forth at positions 64 to 221 of SEQ ID NO: 1. In one embodiment, the gfasCP mutant sequence may have the amino acids at positions 1 to 60 of the amino acid sequence set forth at positions SEQ ID NO: 1 substituted with an amino acid sequence having 90% or more sequence identity to the amino acid sequence set forth at positions 1 to 60 of the amino acid sequence set forth at positions SEQ ID NO: 1, and the amino acids at positions 64 to 221 of the amino acid sequence set forth at positions SEQ ID NO: 1 substituted with an amino acid sequence having 90% or more sequence identity to the amino acid sequence set forth at positions 64 to 221 of the amino acid sequence set forth at positions SEQ ID NO: 1.
[0032] In one embodiment, the gfasCP mutant sequence may be the amino acid sequence set forth in SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11. In the amino acid sequence set forth in SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11, the amino acids corresponding to the serine at position 61 and the glutamine at position 62 in the amino acid sequence set forth in SEQ ID NO: 1 are, in order from position 61 to position 62, leucine-serine, phenylalanine-threonine, isoleucine-serine, leucine-leucine, phenylalanine-cysteine, valine-glycine, serine-serine, alanine-glycine, leucine-glycine, or tyrosine-serine, respectively, and the amino acid sequence corresponding to amino acids at positions 1 to 60 in the amino acid sequence set forth in SEQ ID NO: 1 and the amino acid sequence corresponding to amino acids at positions 64 to 221 in the amino acid sequence set forth in SEQ ID NO: 1 share 100% sequence identity with the identical positions in the amino acid sequence set forth in SEQ ID NO: 1.
[0033] The protein according to one aspect of the first embodiment may have a higher-order structure of a β-barrel containing 11 β-strands, the β-barrel containing 11 β-strands being a structure derived from a gfasCP mutant sequence. When a protein according to one aspect has a β-barrel containing 11 β-strands, it is likely to have high absorbance in the visible light range, making it suitable for use in cell labeling. Those skilled in the art can easily predict that a protein according to one aspect of the first embodiment has a β-barrel containing 11 β-strands using known gfasCP 3D structures (e.g., Non-Patent Document 2) as a template using known software such as AlphaFold, PyMOL, Rosseta, HHpred, RaptorX, CRNPRED, Spanner, and SFAS. Furthermore, the protein can be easily evaluated by comparing known gfasCP 3D structures with measured values using 3D structure analysis techniques commonly used by those skilled in the art, such as cryo-electron microscopy, X-ray crystallography, or nuclear magnetic resonance spectroscopy. Therefore, a protein having a β-barrel containing 11 β-strands according to one aspect of the first embodiment can be obtained without excessive trial and error based on techniques commonly used by those skilled in the art.
[0034] The protein according to one aspect of the first embodiment may form oligomers in a physiological environment, such as an intracellular environment. For example, the protein may form oligomers in an aqueous solution at pH 8.0 (e.g., 50 mM Tris-HCl, 300 mM NaCl, 500 mM imidazole), and the oligomers are formed between structures derived from the gfasCP mutant sequence. When the protein according to one aspect forms an oligomer, it tends to exhibit increased absorbance in the visible light range, making it suitable for use in cell labeling. The oligomer of the protein according to one aspect is, for example, a dimer or tetramer, preferably a dimer. The formation of oligomers by the protein according to one aspect of the first embodiment can be easily evaluated by subjecting a solution containing the protein to a method commonly used by those skilled in the art to evaluate the physical size or molecular weight of proteins, such as size exclusion chromatography or non-denaturing electrophoresis. Therefore, the oligomer-forming protein according to one aspect of the first embodiment can be obtained without undue trial and error by following techniques commonly used by those skilled in the art.
[0035] In a preferred embodiment, the oligomer-forming protein may form oligomers via glutamic acid at position 96 (E96), glutamic acid at position 140 (E140), arginine at position 149 (R149), and phenylalanine at position 158 (F158) in the amino acid sequence set forth in SEQ ID NO: 1 in a structure derived from the gfasCP mutant sequence. When the oligomer-forming protein forms oligomers via these residues, it tends to exhibit increased absorbance in the visible light range, making it suitable for use in cell labeling (Non-Patent Document 2). Whether the oligomer-forming protein forms oligomers via these residues can be easily determined by assessing whether introducing mutations into these residues changes the nature of oligomer formation using methods commonly used by those skilled in the art, such as size exclusion chromatography or non-denaturing electrophoresis, to evaluate the physical size or molecular weight of proteins. Thus, in one preferred embodiment, the gfasCP mutant sequence may have the amino acid corresponding to E96 as glutamic acid, the amino acid corresponding to E140 as glutamic acid, the amino acid corresponding to R149 as arginine, and the amino acid corresponding to F158 as phenylalanine.
[0036] The protein according to one aspect of the first embodiment may have a maximum absorption wavelength that is equal to or lower than a predetermined upper limit and / or equal to or higher than a predetermined lower limit in a physiological environment such as an intracellular environment, for example, in an aqueous solution at pH 8.0 (e.g., 50 mM Tris-HCl, 300 mM NaCl, 500 mM imidazole) that is equal to or lower than a predetermined upper limit and / or equal to or higher than a predetermined lower limit. Furthermore, the protein according to one aspect of the first embodiment may have a maximum absorption wavelength that is equal to or lower than a predetermined upper limit and / or equal to or higher than a predetermined lower limit in a wavelength region of 400 to 800 nm in an aqueous solution at pH 8.0 (e.g., 50 mM Tris-HCl, 300 mM NaCl, 500 mM imidazole) that is equal to or lower than a predetermined upper limit and / or equal to or higher than a predetermined lower limit. In these cases, the predetermined upper limit may be, for example, 578 nm, 577 nm, 576 nm, 575 nm, 570 nm, 550 nm, 530 nm, or 515 nm. The predetermined lower limit may be, for example, 440 nm or more, 470 nm or more, 500 nm or more, or 510 nm or more. The predetermined upper and lower limits can be freely combined. The maximum absorption wavelength or the local maximum absorption wavelength is, for example, 400 nm or more and 577 nm or more, 440 nm or more and 570 nm or less, or 510 nm or more and 570 nm or less, and examples thereof include 441 nm, 513 nm, 565 nm, 566 nm, 567 nm, 569 nm, and 576 nm, with 441 nm and 513 nm being preferred. If the absorption maximum or the absorption maximum wavelength satisfies these conditions, the absorption maximum or the absorption maximum wavelength will be smaller than that of wild-type gfasCP, and therefore the protein may be applicable to observations and devices suitable for shorter wavelength regions, and may also be usable in combination with gfasCP through signal discrimination. The absorption maximum or the absorption maximum wavelength of the protein according to one aspect of the first embodiment can be easily measured by one skilled in the art using an absorbance meter, a microwell plate reader, or the like.
[0037] The protein according to one aspect of the first embodiment may have a molar extinction coefficient at its maximum absorption wavelength in a physiological environment such as an intracellular environment that is equal to or greater than a predetermined lower limit. For example, the molar extinction coefficient at its maximum absorption wavelength in an aqueous solution at pH 8.0 (e.g., 50 mM Tris-HCl, 300 mM NaCl, 500 mM imidazole) may be equal to or greater than a predetermined lower limit. Furthermore, the protein according to one aspect of the first embodiment may have a molar extinction coefficient at its maximum absorption wavelength in a wavelength range of 400 to 800 nm in a physiological environment such as an intracellular environment that is equal to or greater than a predetermined lower limit. For example, the molar extinction coefficient at its maximum absorption wavelength in a wavelength range of 400 to 800 nm in an aqueous solution at pH 8.0 (e.g., 50 mM Tris-HCl, 300 mM NaCl, 500 mM imidazole) may be equal to or greater than a predetermined lower limit. In these cases, the predetermined lower limit is, for example, 5.0 x 10 4 , 8.0×10 4 , 1.0×10 5 , 1.2 × 10 5 , 1.4×10 5 , 1.6×10 5 , 1.8×10 5 or 2.0 x 10 5 [L / mol cm]. When the maximum absorption wavelength or the local maximum absorption wavelength is equal to or greater than the lower limit, the signal intensity derived from the absorbance of the protein according to one aspect is increased, making it suitable for use in cell labeling. Since gfasCP has a large molar extinction coefficient compared to other chromoproteins (Non-Patent Document 1), the protein having a gfasCP mutant sequence according to one aspect also has a large molar extinction coefficient. The molar extinction coefficient at the maximum absorption wavelength or the local maximum absorption wavelength of the protein according to one aspect of the first embodiment can be easily determined by those skilled in the art by measuring absorbance using an absorption spectrometer, a microwell plate reader, or the like.
[0038] The protein according to one aspect of the first embodiment may have a fluorescence quantum yield in a physiological environment such as an intracellular environment of 30% or less, 10% or less, 5% or less, 4% or less, 3% or less, 2% or less, 1% or less, 0.3% or less, 0.1% or less, 0.03% or less, 0.01% or less, or 0.001% or less, and for example, the fluorescence quantum yield in an aqueous solution of pH 8.0 (e.g., 50 mM Tris-HCl, 300 mM NaCl, 500 mM imidazole) may be 30% or less, 10% or less, 5% or less, 4% or less, 3% or less, 2% or less, 1% or less, 0.3% or less, 0.1% or less, 0.03% or less, 0.01% or less, or 0.001% or less. Furthermore, the protein according to one aspect of the first embodiment may have a phosphorescence quantum yield under a physiological environment such as an intracellular environment of 30% or less, 10% or less, 5% or less, 4% or less, 3% or less, 2% or less, 1% or less, 0.3% or less, 0.1% or less, 0.03% or less, 0.01% or less, or 0.001% or less, and for example, the phosphorescence quantum yield in an aqueous solution of pH 8.0 (e.g., 50 mM Tris-HCl, 300 mM NaCl, 500 mM imidazole) may be 30% or less, 10% or less, 5% or less, 4% or less, 3% or less, 2% or less, 1% or less, 0.3% or less, 0.1% or less, 0.03% or less, 0.01% or less, or 0.001% or less. When the fluorescence quantum yield and / or phosphorescence quantum yield of the protein according to one aspect of the first embodiment is not more than the above upper limit, noise derived from fluorescence and / or phosphorescence can be suppressed in detection of a signal derived from the protein. Furthermore, when the fluorescence quantum yield and / or phosphorescence quantum yield of the protein according to one aspect of the first embodiment is equal to or less than the upper limit, photobleaching of the protein due to fluorescence and / or phosphorescence can be suppressed.
[0039] The protein according to one aspect of the first embodiment may be a protein comprising or consisting of a gfasCP mutant sequence (i.e., a gfasCP mutant). Examples of proteins comprising a gfasCP mutant sequence include a fusion protein of a gfasCP mutant with another protein, a tagged gfasCP mutant, or a signal peptide-added gfasCP mutant.
[0040] The other protein in the fusion protein of a gfasCP mutant and another protein is not particularly limited and may be, for example, a protein endogenous to mammalian cells or a mutant protein thereof, preferably a protein endogenous to mammalian cells or a mutant protein thereof whose intracellular localization is known, or a protein listed in Genpept, an NIH-funded NCBI database. Such a fusion protein can be used, for example, to label a specific organelle in a mammalian cell or to label the other protein in a mammalian cell.
[0041] Furthermore, a fusion protein of a gfasCP mutant with another protein can be used as an embodiment of a protein complex containing the fusion protein. An example of such a protein complex is an antibody labeled with a gfasCP mutant, which can be used to label an antigen to which the antibody binds.
[0042] Examples of tags in tagged gfasCP mutants include peptide tags (e.g., FLAG tags) or protein tags (e.g., HALO protein or SNAP) that bind to specific ligands. Such tagged mutants can be used to label specific proteins in mammalian cells by utilizing the affinity of the tag, and can be easily purified by affinity chromatography or other methods.
[0043] Examples of signal peptides in the signal peptide adducts of gfasCP mutants include nuclear localization signals (NLS), nuclear export signals (NES), and mitochondrial localization signals (MTS), and these signal peptides can be known. These signal peptide adducts can be used to label specific organelles within cells (e.g., mammalian cells).
[0044] A protein consisting of a gfasCP mutant sequence can be used for whole cell labeling, for example, by expressing it in a cell. In a specific embodiment, the protein consisting of a gfasCP mutant sequence may be a protein consisting of the amino acid sequence set forth in SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11.
[0045] The protein according to the first embodiment can be obtained according to methods commonly used by those skilled in the art. For example, the protein according to the first embodiment can be obtained by introducing a plasmid vector encoding the protein into bacteria such as Escherichia coli, culturing the bacteria with activated RNA polymerase to express the protein, and then isolating and purifying the protein from the bacteria. Introduction of the plasmid vector into bacteria can be carried out according to conventional methods, such as heat shock, lipofection, or electroporation. After vector introduction, bacteria into which the vector has been introduced can be selected using a selection marker inserted into the vector and its corresponding antibiotic. Activation of RNA polymerase can be carried out according to conventional methods, such as by adding isopropyl β-D-1-thiogalactopyranoside (IPTG) to the culture medium. Protein isolation and purification from bacteria can be carried out according to conventional methods, such as disrupting the cell membrane by sonication or exposure to a detergent, and separating the eluted protein by affinity chromatography, electrophoresis, or the like.
[0046] A second embodiment of the present disclosure is a nucleic acid comprising a sequence encoding a protein according to one aspect of the first embodiment of the present disclosure. That is, the second embodiment of the present disclosure is a nucleic acid comprising a base sequence consisting of codons corresponding to the amino acid sequence of the protein according to one aspect of the first embodiment of the present disclosure. A sequence encoding a given protein can be easily optimized by a person skilled in the art using known methods with respect to the amino acid sequence of the protein, depending on the intended use (e.g., the origin of the cell into which it is introduced). The nucleic acid according to the second embodiment may be, for example, DNA or RNA.
[0047] In one aspect, the nucleic acid of the second embodiment may be a nucleic acid consisting of a sequence encoding a protein according to one aspect of the first embodiment. Such a nucleic acid can be used as a nucleic acid to be ligated in the preparation of a nucleic acid for expressing the protein encoded by the nucleic acid or a fusion protein thereof in a cell (e.g., a mammalian cell or E. coli).
[0048] In one aspect, the nucleic acid of the second embodiment may be a nucleic acid that causes a cell to express the protein according to one aspect of the first embodiment or a fusion protein thereof. The cell in which the nucleic acid is expressed is not particularly limited and may be, for example, an animal cell or Escherichia coli, and the animal cell may be, for example, a mammalian cell. The type and structure of such a nucleic acid can be appropriately selected by those skilled in the art depending on the cell in which the nucleic acid is expressed.
[0049] In one aspect, the nucleic acid of the second embodiment may be mRNA or a precursor thereof comprising a sequence encoding the protein according to one aspect of the first embodiment or a fusion protein thereof. Such mRNA or a precursor thereof can be used to express the protein or a fusion protein thereof in mammalian cells. The mRNA precursor is not particularly limited as long as it can provide an mRNA comprising a sequence encoding the protein according to one aspect of the first embodiment or a fusion protein thereof through splicing in mammalian cells. Such mRNA and precursors can be appropriately designed by those skilled in the art depending on the amino acid sequence of the protein to be expressed. Such mRNA and precursors can be introduced into cells by, for example, lipofection or electroporation.
[0050] In one aspect, the nucleic acid of the second embodiment may be a vector that causes a cell to express the protein according to one aspect of the first embodiment or a fusion protein thereof. Such a vector may be, for example, a viral vector or a plasmid vector. Viral vectors include, for example, retroviral vectors, adenoviral vectors, adeno-associated viral (AAV) vectors, and lentiviral vectors, which can introduce genes into mammalian cells via the infectivity of the virus and cause the target protein to be expressed in the animal cells (e.g., mammalian cells). Plasmid vectors can be introduced into animal cells (e.g., mammalian cells) to express the target protein in the mammalian cells, and can also be introduced into bacteria such as Escherichia coli to synthesize the target protein. Plasmid vectors can be introduced into cells by, for example, lipofection or electroporation.
[0051] The nucleic acid according to the second embodiment can be designed and obtained according to methods commonly used by those skilled in the art. For example, it can be obtained by outsourcing synthesis to an external company that performs custom synthesis of nucleic acids with designated sequences, or by ligating a nucleic acid having a sequence encoding a protein according to one aspect of the first embodiment into the cloning site of a commercially available backbone vector via restriction enzyme treatment. Furthermore, such nucleic acids can be amplified according to methods commonly used by those skilled in the art, such as by introducing them into bacteria such as Escherichia coli and culturing them.
[0052] The protein according to the first embodiment and the nucleic acid according to the second embodiment can be used to express the protein according to the first embodiment in a cell. That is, a third embodiment of the present disclosure is a cell expressing the protein according to the first embodiment or a fusion protein thereof. The cell may be, for example, an animal cell, such as a mammalian cell or a human cell. Such a cell can be evaluated by detecting a signal derived from the protein according to the first embodiment or a fusion protein thereof. That is, a fourth embodiment of the present disclosure is a method for evaluating a cell, comprising detecting a signal derived from the protein according to the first embodiment or a fusion protein thereof in a cell expressing the protein according to the first embodiment or a fusion protein thereof. Such an evaluation method utilizes a signal derived from the absorbance of a structure derived from a gfasCP mutant sequence to evaluate a cell, and the method may be an in vitro method.
[0053] The method of the fourth embodiment uses, for example, coloration based on the absorbance of a structure derived from the gfasCP mutant sequence as an index, and evaluates cells by using the coloration of cells expressing the protein of the first embodiment or a fusion protein thereof. Such a method can be used, for example, to evaluate the condition dependence of protein expression (e.g., environmental dependence or concentration dependence of a specific substance). Cell coloration can be evaluated, for example, visually, or the color tone can be evaluated by using image analysis software (e.g., ImageJ) on a digital photograph of the cells.
[0054] In the method of the fourth embodiment, for example, a protein in which a structure derived from the gfasCP mutant sequence acts as a fluorescence resonance energy transfer (FRET) acceptor is used to evaluate cells co-expressing a donor protein or expressing a fusion protein with the donor protein, using the fluorescence intensity of the donor protein as an indicator. Such a method can be used to detect intracellular environments or events detected by fluorescent probes that operate on FRET, such as protein-protein interactions, ion concentrations, membrane potentials, or pH changes.
[0055] The method of the fourth embodiment evaluates the distribution of the protein of the first embodiment or a fusion protein thereof in cells based on, for example, a refractive index change due to the absorbance of a structure derived from a gfasCP mutant sequence. Patent Document 1 discloses a cell evaluation method including: a labeling step of labeling a specific site of a cell with a labeling substance having different refractive indices at a first wavelength and a second wavelength; a refractive index distribution acquisition step of acquiring the refractive index distribution of the cell whose specific site has been labeled in the labeling step at the first wavelength and the second wavelength; and an analysis step of evaluating the distribution of the specific site in the cell by comparing the refractive index distributions at the first wavelength and the second wavelength. A chromoprotein is cited as an example of the labeling substance. This method evaluates cells by utilizing the difference in refractive index between the two wavelengths of the labeling substance. On the other hand, pigments such as chromoproteins are known to exhibit a large refractive index change near their maximum absorption wavelength. Because the absorbance of a structure derived from a gfasCP mutant sequence is higher than that of other chromoproteins, a large refractive index change near their maximum absorption wavelength is expected. Therefore, one aspect of the method of the fourth embodiment is considered to be suitable as one aspect of a cell evaluation method that uses a refractive index as an index, such as the method described in Patent Document 1. [Example]
[0056] The present disclosure will be described in more detail below using examples, but the present disclosure is not limited to the following examples.
[0057] Example 1: Obtaining wild-type gfasCP As a nucleic acid encoding the amino acid sequence of wild-type (WT) gfasCP (SEQ ID NO: 1), we commissioned Eurofins Genomics to synthesize a DNA (gfasCP-DNA) consisting of the nucleotide sequence shown in SEQ ID NO: 13, with codons optimized for expression in E. coli K12. The gfasCP-DNA was amplified by PCR using primers shown in SEQ ID NOs: 14 and 15 under the PCR conditions described below. The amplified gfasCP-DNA was subcloned into the multicloning site of the pCold® I vector (Takara Bio, 3361, SEQ ID NO: 16) using 2× Ligation Mix (Nippon Gene Co., Ltd.) and the restriction enzymes NdeI and SalI to obtain the WT gfasCP expression vector (pKT005). The pCold® I vector used was purchased from Takara Bio and replicated in E. coli. pKT005 was introduced into competent cells BL21 (Takara Bio, 9126), a type of E. coli strain. The pKT005-introduced BL21 cells were plated on LB agar medium containing 100 μg / mL ampicillin as a selection marker and cultured overnight at 37°C. All of the resulting colonies were inoculated into 20 mL of LB medium containing 100 μg / mL ampicillin. Isopropyl β-D-1-thiogalactopyranoside (IPTG, Fujifilm Wako Pure Chemical Industries) was added to a final concentration of 1 mM, and the LB medium was cultured at 15°C for 20 hours. SEQ ID NO: 14: ATGCATATGAGTGTGATTGCCAAACAG (forward) SEQ ID NO: 15: ATGGTCGACTCATTAGGCTACAACCGATTTAC (reverse) PCR conditions: 95°C for 2 minutes → (95°C for 30 seconds → 55°C for 30 seconds → 72°C for 60 seconds) x 30 cycles → 72°C for 10 minutes → 4°C infinity
[0058] After cultivation, the LB medium was centrifuged at 5,000 rpm for 10 minutes at 4°C to collect a pellet of wild-type gfasCP-expressing E. coli. A photograph taken under room light is shown in Figure 1. As previously reported, the pellet of wild-type gfasCP-expressing E. coli was purple.
[0059] <Example 2: Obtaining gfasCP mutant 1> Random mutations were introduced into the amino acids corresponding to the 61st serine (S61), the 62nd glutamine (Q62), and the 63rd tyrosine (Y63) in the amino acid sequence shown in SEQ ID NO: 1, which are amino acids near the chromophore of gfasCP, in an attempt to obtain gfasCP mutants with shorter absorption wavelengths.
[0060] The following primer (SEQ ID NO: 17, N randomly represents A, C, G, or T; hereinafter referred to as "random primer") was designed and synthesis was outsourced to Eurofins Genomics. While the random primer has random base sequences corresponding to S61, Q62, and Y63 of gfasCP, the 13 bases before and after it are complementary to the base sequence of gfasCP-DNA, allowing it to bind to single-stranded gfasCP-DNA. SEQ ID NO: 17: TCTTAGTCCGCAANNNNNNNNNGGCTCAATCCCGT
[0061] The random primers were incubated at 37°C for 30 minutes in a reaction system containing the following reagents (50 μL in 1x T4 Polynucleotide Kinase Buffer (70 mM Tris-HCl, 10 mM MgCl2, 5 mM DTT, pH 7.6)), to phosphorylate their 5' ends and obtain 5'-phosphorylated random primers. Random primer (final concentration 6 μM) T4 Polynucleotide Kinase (10 units, New England Biolabs) T4 Polynucleotide Kinase Buffer (final concentration 1x, New England Biolabs) ATP (final concentration 1 mM, Takara Bio)
[0062] PCR was performed in a reaction system (50 μL) containing the following reagents. The PCR reaction was first incubated at 65°C for 5 minutes, then at 95°C for 5 minutes. 25 cycles of incubation at 95°C for 10 seconds, 50°C for 30 seconds, and 65°C for 14 minutes were then performed. Finally, the PCR reaction was completed by incubating at 75°C for 7 minutes. pKT005 (150 ng) 5' phosphorylated primer (final concentration 0.3 pmol / μL) dNTP mixture (final concentration 0.05 mM each, Takara Bio) Pfu Ultra High-Fidelity DNA Polymerase (2.5 Units, Agilent Technologies) Reaction Buffer (final concentration 1x, Agilent Technologies) Taq DNA Ligase (20 Units, New England Biolabs) Taq DNA Ligase Buffer (final concentration 1x, New England Biolabs)
[0063] This PCR reaction system contains a 5'-phosphorylated random primer and pKT005 obtained in Example 1 as template DNA. Therefore, during the PCR reaction, the 5'-phosphorylated random primer binds to a site in pKT005 centered on the base sequence corresponding to S61, Q62, and Y63 of gfasCP-DNA. DNA polymerase then replicates the DNA in the portion of pKT005 other than the site where the 5'-phosphorylated random primer binds, going almost completely around pKT005. This results in a vector consisting of pKT005 template DNA and circular DNA in which random sequences have been introduced into the base sequences corresponding to S61, Q62, and Y63 of gfasCP-DNA.
[0064] Next, 20 units of DpnI (Takara Bio) was added as a restriction enzyme to the solution after the PCR reaction, and the mixture was incubated at 37°C for 30 minutes to degrade the template DNA. m ATC( m As mentioned above, the DNA derived from pKT005 was amplified from the purchased pCold® I vector using E. coli with DNA adenine methyltransferase (Dam) activity. m ATC sequence. Therefore, the DNA derived from pKT005 is m Since it has an ATC sequence, it is digested with DpnI. On the other hand, the circular DNA synthesized by PCR and containing random sequences is digested with G m Since it does not have an ATC sequence, it is not digested by DpnI.
[0065] Next, the solution after template DNA digestion was added to the pCold I vector. The primer shown in SEQ ID NO: 18, which was synthesized and phosphorylated in the same manner as the 5'-phosphorylated random primer, was phosphorylated at the 5' end. Pfu Ultra High-Fidelity DNA Polymerase (0.625 Units, Agilent Technologies) was then added to the primer. The circular DNA containing the random sequence was amplified by PCR. The PCR reaction was first incubated at 95°C for 30 seconds. Two cycles were then performed: 95°C for 30 seconds, 55°C for 1 minute, and 70°C for 15 minutes. Finally, the PCR was completed by incubating at 75°C for 7 minutes. SEQ ID NO: 18: GGCAGGGATCTTAGATTCTG
[0066] A 1 μL aliquot of the PCR solution was added to competent Ecos SonicCompetent E. coli BL21(DE3) cells (Nippon Gene) to introduce the circular DNA containing the random sequence. The BL21(DE3) cells were plated on LB agar medium containing 100 μg / mL ampicillin as a selection marker and coated with 0.5 μL of IPTG per 10 cm dish, and cultured overnight at 37°C. The sample was then cooled to 15°C and then warmed to room temperature.
[0067] Figure 2 shows a photograph taken under room light of a 10-cm diameter dish inoculated with transformants on LB agar medium. As shown in Figure 2, numerous colonies of transformants were obtained that were visually visible. Among them, colonies 3, 5, and 6 in the enlarged image on the right exhibited characteristic red, magenta, and orange colors, respectively. These three colonies were then collected and inoculated into LB medium containing 100 μg / mL ampicillin for overnight culture. After cultivation, the plasmids were extracted using NucleoSpin Plasmid Easy Pure (Macherey-Nagel), a kit for purifying plasmid DNA from E. coli cultures, according to the instructions. Hereafter, the plasmids derived from colonies 3, 5, and 6 in the photograph in Figure 2 are referred to as pMGFE003, pMGFE005, and pMGFE006, respectively. DNA sequence analysis of pMGFE003, pMGFE005, and pMGFE006 using the primer shown in SEQ ID NO: 19 was outsourced to Eurofins Genomics. SEQ ID NO: 19: ACGCCATATCGCCGAAAGG
[0068] Figure 3 shows the nucleotide sequences near the sites of random mutagenesis in pMGFE003, pMGFE005, and pMGFE006, with the underlined regions indicating the sites of random mutagenesis. Figure 4 shows the results of identifying the amino acid sequences near the sites of random mutagenesis in pMGFE003, pMGFE005, and pMGFE006 from the results of Figure 3. As shown in Figures 3 and 4, the nucleotide sequences encoding amino acids corresponding to S61, Q62, and Y63 of gfasCP were a sequence encoding leucine, serine, and tyrosine (S61L, Q62S mutation) in pMGFE003, a sequence encoding phenylalanine, threonine, and tyrosine (S61F, Q62T mutation) in pMGFE005, and a sequence encoding isoleucine, serine, and tyrosine (S61I, Q62S mutation) in pMGFE006. Thus, the gfasCP mutants expressed by these plasmids all had S61 substituted with a hydrophobic amino acid and Q62 substituted with a polar, uncharged amino acid other than glutamine, while Y63 was unmutated. These results suggest that gfasCP mutants without a Y63 mutation are advantageous for pigmentation. The amino acid sequences of the proteins expressed by pMGFE003, pMGFE005, and pMGFE006 are shown in SEQ ID NOs: 2, 3, and 4, respectively, and the amino acids other than the mutated sites were identical to those of the wild-type.
[0069] <Example 3: Obtaining gfasCP mutant 2> Based on the results of Example 2, a gfasCP mutant was obtained in the same manner as in Example 2 using random primers whose nucleotide sequence is shown in SEQ ID NO: 20, which do not introduce a mutation into Y63 but introduce random mutations into S61 and Q62. In the nucleotide sequence of SEQ ID NO: 20, N randomly represents A, C, G, or T. SEQ ID NO: 20: ATTCTTAGTCCGCAANNNNNNTATGGCTCAATCCCG
[0070] Figure 5 shows a photograph taken under room light of a 10 cm diameter dish inoculated with transformants on LB agar medium. As shown in Figure 5, many colonies of transformants were obtained that were visually visible. Among them, 16 colonies indicated in white letters as 012 to 027 in Figure 5 were reddish compared to gfasCP and exhibited shorter absorption wavelengths. These 16 colonies were then collected, treated as in Example 2, and subjected to DNA sequence analysis. Hereinafter, the plasmids derived from the colonies indicated as 012 to 027 in the photograph in Figure 5 will be referred to as pMGFE012 to pMGFE027, respectively.
[0071] DNA sequence analysis revealed that the nucleotide sequences encoding the amino acids corresponding to S61 and Q62 in the plasmids derived from 9 of the 16 colonies were either the same as the wild-type or contained a mixture of multiple mutants due to contamination. On the other hand, the plasmids derived from 7 of the 16 colonies (pMGFE014, pMGFE015, pMGFE016, pMGFE020, pMGFE022, pMGFE025, and pMGFE027) expressed novel mutants in which mutations had been introduced into the amino acids corresponding to S61 and Q62. Figure 6 shows the nucleotide sequences of the regions near the random mutagenesis sites in pMGFE014, pMGFE015, pMGFE016, pMGFE020, pMGFE022, pMGFE025, and pMGFE027. The underlined regions indicate the random mutagenesis sites. FIG. 7 shows the results of identifying the amino acid sequences near the sites where random mutations were introduced in pMGFE014, pMGFE015, pMGFE016, pMGFE020, pMGFE022, pMGFE025, and pMGFE027, based on the results of FIG. 6. As shown in Figures 6 and 7, the base sequences encoding the amino acids corresponding to S61 and Q62 of gfasCP were a sequence encoding leucine, leucine in pMGFE014 (S61L, Q62L mutation), a sequence encoding phenylalanine, cysteine in pMGFE015 (S61F, Q62C mutation), a sequence encoding valine, glycine in pMGFE016 (S61V, Q62G mutation), a sequence encoding serine, serine in pMGFE020 (Q62S mutation), a sequence encoding alanine, glycine in pMGFE022 (S61A, Q62G mutation), a sequence encoding leucine, glycine in pMGFE025 (S61L, Q62G mutation), and a sequence encoding tyrosine, serine in pMGFE027 (S61Y, Q62S mutation). The amino acid sequences of the proteins expressed by pMGFE014, pMGFE015, pMGFE016, pMGFE020, pMGFE022, pMGFE025, and pMGFE027 are shown in SEQ ID NOs: 5, 6, 7, 8, 9, 10, and 11, respectively, and the amino acids other than the mutated sites were identical to those of the wild type.
[0072] The gfasCP mutants obtained in Examples 2 and 3 are shown in Table 1. As shown in Table 1, the obtained gfasCP mutants did not have a mutation at Y63, and furthermore, either (A) the amino acid corresponding to S61 was a hydrophobic amino acid and the amino acid corresponding to Q62 was a polar, uncharged amino acid, (B) the amino acid corresponding to S61 was a hydrophobic amino acid and the amino acid corresponding to Q62 was a hydrophobic amino acid, or (C) the amino acid corresponding to S61 was a polar, uncharged amino acid and the amino acid corresponding to Q62 was a polar, uncharged amino acid other than glutamine.
[0073] [Table 1]
[0074] Example 4: Evaluation of gfasCP mutants The gfasCP mutants found in Examples 2 and 3 were expressed in bacterial cells, purified, and their optical properties were evaluated. Note that the plasmids used in Examples 2 and 3 were designed so that a His-tag was added to the N-terminus of the protein, and therefore the gfasCP mutants were purified using Ni-NTA agarose beads.
[0075] pMGFE003, pMGFE005, pMGFE006, pMGFE014, pMGFE015, pMGFE016, pMGFE020, pMGFE022, pMGFE025 and pMGFE027 were introduced into competent cells BL21 (Takara Bio, 9126) and cultured in the same manner as in Example 1. After cultivation, the cells are collected by centrifugation.To this solution, 1 mL of CellLytic B (Sigma-Aldrich) was added per 100 mg of bacterial cells and incubated at room temperature for at least 10 minutes with end-over-end mixing. To the incubated solution, 50 units / mL of Benzonase (Sigma-Aldrich) and 0.2 mg / mL of Lysozyme (Fujifilm Wako Pure Chemical Industries) were added and further incubated at room temperature for at least 10 minutes with end-over-end mixing. After incubation, the mixture was centrifuged to obtain the soluble fraction. The obtained soluble fraction was mixed with 0.25 mL of Ni-NTA agarose and incubated at room temperature for at least 30 minutes with end-over-end mixing. The Ni-NTA agarose was recovered and washed with Buffer B shown below. Subsequently, Ni-NTA agarose was added to Buffer C to elute the protein bound to the Ni-NTA agarose, yielding a gfasCP mutant solution. [Buffer B] 50mM Tris-HCl pH 8.0 300mM NaCl 50mM Imidazole [Buffer C] 50mM Tris-HCl pH 8.0 300mM NaCl 500mM Imidazole
[0076] Figure 8 shows photographs taken under room light of the wild-type gfasCP solution obtained in Example 1 and the gfasCP mutant solutions obtained from pMGFE003, pMGFE005, pMGFE006, pMGFE014, pMGFE015, pMGFE016, pMGFE020, pMGFE022, pMGFE025, and pMGFE027. As shown in Figure 8, the color of each gfasCP mutant solution was confirmed visually, indicating that a chromoprotein had been obtained.
[0077] Next, the purified gfasCP mutant was confirmed. The soluble fraction or gfasCP mutant solution (eluted fraction) obtained by centrifugation after bacterial cell culture was mixed with SDS-PAGE Sample Buffer (Tokyo Chemical Industry Co., Ltd.) and incubated at 100°C for 5 minutes. The solution was applied to a SuperSep Ace 12.5% acrylamide gel (Fujifilm Wako Pure Chemical Industries, Ltd.) and electrophoresed in SDS-PAGE 1× Running Buffer (Nippon Gene Co., Ltd.). The electrophoresed gel was stained with CBB Stain One Super (Nacalai Tesque) and destained with ultrapure water. Figure 9 shows the results of polyacrylamide gel electrophoresis of the soluble fraction or gfasCP mutant solution (eluted fraction) obtained by centrifugation after bacterial cell culture. In Figure 9, M indicates a molecular weight marker, S indicates the soluble fraction, and E indicates the eluted fraction. As shown in FIG. 9, for both wild-type gfasCP and the gfasCP mutants, a strong band was observed only around 29 kDa, the molecular weight of gfasCP, in the eluted fractions, confirming high purity.
[0078] Next, absorption spectra of the obtained gfasCP mutants were obtained. Absorption spectra were measured using a spectrophotometer U-4100 (Hitachi, Ltd.) in the above-mentioned Buffer C (50 mM Tris-HCl, 300 mM NaCl, 500 mM imidazole). Figures 10 and 11 show the absorption spectra of wild-type gfasCP and gfasCP mutants. As shown in Figures 10 and 11, the maximum absorption wavelengths of all 10 obtained gfasCP mutants were shorter than that of wild-type gfasCP (579 nm). As a result, new gfasCP mutants with maximum absorption wavelengths shorter than that of wild-type gfasCP (579 nm) were obtained as gfasCP mutants having the mutations described in Table 1 at S61 and / or Q62.
Claims
1. The gfasCP mutant sequence has an amino acid sequence having 90% or more sequence identity with the amino acid sequence shown in SEQ ID NO: 1, In the gfasCP mutant sequence, the amino acid corresponding to tyrosine at position 63 in the amino acid sequence shown in SEQ ID NO: 1 is tyrosine; The gfasCP mutant sequence satisfies at least one of the following (A) to (C): (A) the amino acid corresponding to the serine at position 61 in the amino acid sequence shown in SEQ ID NO: 1 is a hydrophobic amino acid, and the amino acid corresponding to the glutamine at position 62 in the amino acid sequence shown in SEQ ID NO: 1 is a polar, uncharged amino acid; (B) the amino acid corresponding to the serine at position 61 in the amino acid sequence shown in SEQ ID NO: 1 is a hydrophobic amino acid, and the amino acid corresponding to the glutamine at position 62 in the amino acid sequence shown in SEQ ID NO: 1 is a hydrophobic amino acid; (C) The amino acid corresponding to the serine at position 61 in the amino acid sequence shown in SEQ ID NO: 1 is a polar, uncharged amino acid, and the amino acid corresponding to the glutamine at position 62 in the amino acid sequence shown in SEQ ID NO: 1 is a polar, uncharged amino acid other than glutamine.
2. the hydrophobic amino acid is alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, glycine, or cysteine; The protein of claim 1 , wherein the polar uncharged amino acid is asparagine, glutamine, serine, threonine, or tyrosine.
3. The protein according to claim 1, wherein in the gfasCP mutant sequence, the amino acids corresponding to the serine at position 61 and the glutamine at position 62 in the amino acid sequence shown in SEQ ID NO: 1 are, in order from position 61 to position 62, leucine-serine, phenylalanine-threonine, isoleucine-serine, leucine-leucine, phenylalanine-cysteine, valine-glycine, serine-serine, alanine-glycine, leucine-glycine, or tyrosine-serine.
4. The protein of claim 1, wherein the gfasCP mutant sequence comprises an amino acid sequence having 90% or more sequence identity with the amino acid sequence shown at positions 1 to 60 of SEQ ID NO: 1, and an amino acid sequence having 90% or more sequence identity with the amino acid sequence shown at positions 64 to 221 of SEQ ID NO:
1.
5. The protein according to claim 1, which has a β-barrel containing 11 β-strands as its higher-order structure.
6. The protein of claim 1, which forms oligomers in an aqueous solution at pH 8.
0.
7. The protein of claim 1, which has a maximum absorption wavelength of 578 nm or less in an aqueous solution at pH 8.
0.
8. The molar absorption coefficient at the maximum absorption wavelength in an aqueous solution of pH 8.0 is 5.0 × 10 4 The protein according to claim 1, wherein the solubility is [L / mol·cm] or more.
9. The protein of claim 1 , wherein the gfasCP mutant sequence is the amino acid sequence shown in SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10 or 11.
10. The protein of claim 1 , consisting of the gfasCP mutant sequence.
11. A nucleic acid comprising a sequence encoding a protein according to any one of claims 1 to 10.
12. A vector that causes a cell to express the protein or a fusion protein thereof according to any one of claims 1 to 10.
13. A cell expressing the protein or a fusion protein thereof according to any one of claims 1 to 10.
14. A method for evaluating a cell, comprising detecting a signal derived from the protein or a fusion protein thereof according to any one of claims 1 to 10 in a cell expressing the protein or a fusion protein thereof.
Citation Information
Patent Citations
Cell evaluation method and cell evaluation device
WO2023182169A1