Protein for addition to target protein

A novel protein with a specific tertiary structure enhances the stabilization and structural analysis of target proteins by forming stable attachments, addressing the limitations of existing methods.

WO2025225731A1PCT designated stage Publication Date: 2025-10-30THE UNIV OF TOKYO +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/016100
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-26
Filing Date
2025-04-25
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Existing methods for stabilizing and facilitating structural analysis of target proteins, such as membrane proteins, are inadequate, particularly those involving bRIL protein or T4 lysozyme, as they do not provide sufficient stabilization or structural analysis effectiveness.

Method used

A novel protein with a predetermined tertiary structure comprising a central domain and two peripheral domains, each formed by helices, where the peripheral domains have specific hydrophilic and hydrophobic amino acid distributions, allowing for stable attachment to target proteins like G protein-coupled receptors, thereby enhancing structural analysis.

Benefits of technology

The novel protein stabilizes target proteins and facilitates effective structural analysis by preventing orientation bias and allowing precise visualization using electron microscopy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JPOXMLDOC01-APPB-T000001
    Figure JPOXMLDOC01-APPB-T000001
  • Figure JPOXMLDOC01-APPB-T000002
    Figure JPOXMLDOC01-APPB-T000002
  • Figure JPOXMLDOC01-APPB-T000003
    Figure JPOXMLDOC01-APPB-T000003
Patent Text Reader

Abstract

Provided is a novel protein for facilitating the stabilization or structural analysis of a target protein. The protein is for addition to a target protein 200 and forms a predetermined tertiary structure including a central domain 130, a first peripheral domain 110, and a second peripheral domain 120. The central domain 130, the first peripheral domain 110, and the second peripheral domain 120 each include a hydrophobic core formed by a plurality of helices. The hydrophobic core of the central domain is formed by a plurality of helices including at least one helix included in the first peripheral domain 110 and at least one helix included in the second peripheral domain 120.
Need to check novelty before this filing date? Find Prior Art

Description

Proteins to be added to target proteins CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application is based on Japanese Patent Application No. 2024-72890, filed on April 26, 2024, the contents of which are incorporated herein by reference.

[0002] The present disclosure relates to proteins for attachment to target proteins.

[0003] Conventionally, techniques for adding a predetermined protein to a target protein have been known in order to stabilize the target protein or facilitate structural analysis of the target protein.

[0004] Japanese Patent Application Laid-Open No. 2021-127333

[0005] However, the protein structure for membrane protein structural analysis described in Patent Document 1 is a general-purpose structure that can be commonly used in the purification, crystallization, and structural analysis of membrane proteins, and therefore may not be suitable for stabilizing a target protein or facilitating structural analysis. Furthermore, while a technique for stabilizing a target protein or facilitating structural analysis is known in which bRIL protein or T4 lysozyme is added to the target protein, this technique has only limited effect on stabilizing the target protein or facilitating structural analysis.

[0006] Therefore, an object of the present disclosure is to provide a novel protein for stabilizing a target protein or facilitating structural analysis.

[0007] The present disclosure is as follows: [1] A protein to be added to a target protein, the protein forming a predetermined tertiary structure including a central domain, a first peripheral domain, and a second peripheral domain, the central domain, the first peripheral domain, and the second peripheral domain each including a hydrophobic core formed by a plurality of helices, the hydrophobic core of the central domain being formed by a plurality of helices including at least one helix included in the first peripheral domain and at least one helix included in the second peripheral domain. [2] The protein according to [1], wherein at least 90% of the amino acid residues on the surface of the first peripheral domain are hydrophilic amino acids, and at least 90% of the amino acid residues within the hydrophobic core of the first peripheral domain are hydrophobic amino acids. [3] The protein according to [1] or [2], wherein 90% or more of the amino acid residues on the surface of the second peripheral domain are hydrophilic amino acids, and 90% or more of the amino acid residues within the hydrophobic core of the second peripheral domain are hydrophobic amino acids. [4] The protein according to any one of [1] to [3], wherein the helices forming the first peripheral domain and the helices forming the second peripheral domain consist of 10 to 40 amino acid residues. [5] The protein according to any one of [1] to [4], wherein the helices forming the central domain include at least one helix having fewer amino acids than the helices forming the first peripheral domain and the second peripheral domain. [6] The protein according to any one of [1] to [5], wherein the first peripheral domain is formed by 6 to 12 helices, and the second peripheral domain is formed by 6 to 12 helices that are all different from the 6 to 12 helices forming the first peripheral domain. [7] The protein according to any one of [1] to [6], wherein the number of helices forming the first peripheral domain is the same as the number of helices forming the second peripheral domain.[8] The protein according to any one of [1] to [7], wherein the predetermined tertiary structure comprises, in order from the N-terminus, a first peripheral domain, a central domain, and a second peripheral domain, wherein the helix forming the first peripheral domain comprises a first connecting helix that connects to a first target helix contained in the target protein to form a series of helices, and the helix forming the second peripheral domain comprises a second connecting helix that connects to a second target helix contained in the target protein to form a series of helices. [9] The protein according to any one of [1] to [8], wherein the target protein is a G protein-coupled receptor, wherein the first connecting helix connects to the first target helix, which is the fifth transmembrane domain of the G protein-coupled receptor, to form a series of helices, and the second connecting helix connects to the second target helix, which is the sixth transmembrane domain of the G protein-coupled receptor, to form a series of helices.

[0008] According to the present disclosure, it is possible to provide a novel protein for stabilizing a target protein or facilitating structural analysis.

[0009] FIG. 1 is a diagram showing the structure of an additional protein 100. FIG. 2 is a diagram showing the structure of an additional protein 100. FIG. 3 is a diagram showing the structure of an additional protein 100. FIG. 4 is a diagram showing the structure of an additional protein 100. FIG. 5 is a diagram showing the structure of an additional protein 100. FIG. 6 is a diagram showing a schematic representation of the positional relationship of helices contained in an additional protein. FIG. 7 is a diagram showing a schematic representation of a helix forming a first peripheral domain. FIG. 8 is a diagram showing a schematic representation of a helix forming the first peripheral domain. FIG. 9 is a diagram showing a schematic representation of a helix forming the first peripheral domain. FIG. 10 is a diagram showing a schematic representation of a helix forming the first peripheral domain. FIG. 11 is a diagram showing a schematic representation of a helix forming the second peripheral domain. FIG. 12 is a diagram showing a schematic representation of a helix forming the second peripheral domain. FIG. 13 is a diagram showing an example of a method for designing an additional protein 100.

[0010] Preferred embodiments of the present disclosure will be described with reference to the accompanying drawings.

[0011] In this embodiment, the protein to be added to the target protein is referred to as an addition protein. The addition protein and the target protein are connected at a connection region. In this embodiment, the protein in which the addition protein is added to the target protein is also referred to as a whole protein.

[0012] In recent years, in research on target proteins, particularly in structural analysis of target proteins, additional proteins may be added to the target protein. Examples of the purpose of adding additional proteins include stabilizing the target protein and increasing its apparent molecular size to facilitate structural analysis. However, the purpose of adding additional proteins is not limited to these.

[0013] The addition of an additional protein to a target protein may be performed, for example, by inserting the additional protein into a predetermined region of the target protein. More specifically, a DNA sequence corresponding to the additional protein is inserted into a predetermined region of a DNA sequence corresponding to the target protein, and the entire protein is synthesized using the inserted DNA sequence. When inserting the additional protein, a portion of the additional protein, for example, the amino acid residue at the insertion site, may be appropriately edited, or a portion of the additional protein, for example, the amino acid residue at the insertion site, may be appropriately edited. When inserting the additional protein, a portion of the target protein, for example, the amino acid residue at the insertion site, may be appropriately edited, or a portion of the target protein, for example, the amino acid residue at the insertion site, may be appropriately edited.

[0014] The additional protein in this embodiment is a novel protein that is added to a target protein for, for example, stabilization or structural analysis of the target protein.

[0015] The additional protein forms a tertiary structure including a first peripheral domain, a second peripheral domain, and a central domain. The additional protein forms a tertiary structure in the order of, for example, the first peripheral domain, the central domain, and the second peripheral domain from the N-terminus.

[0016] The first peripheral domain, the second peripheral domain, and the central domain are each formed, for example, by a plurality of helices.

[0017] Additionally, the first peripheral domain, the second peripheral domain, and the central domain each include a first hydrophobic core, a second hydrophobic core, and a third hydrophobic core, each formed, for example, by a plurality of helices.

[0018] For example, the first peripheral domain may have at least 90% of the amino acid residues on its surface that are hydrophilic, and at least 90% of the amino acid residues within the first hydrophobic core that are hydrophobic. The surface of the first peripheral domain may have, for example, amino acid residues that are not within the first hydrophobic core and that form hydrogen bonds with a solvent (e.g., water). This allows the attached protein to form a stable tertiary structure.

[0019] Each of the multiple helices forming the first peripheral domain may be a helix consisting of, for example, 10 to 40 amino acid residues, which can prevent the occurrence of structural variations in the attached protein (so-called orientation bias) and allow for effective structural analysis of the entire protein using, for example, an electron microscope.

[0020] Specifically, each of the multiple helices forming the first peripheral domain may be a helix consisting of, for example, 11 to 15, 18 to 22, 25 to 29, or 32 to 36 amino acid residues. More specifically, when the additional protein is composed of 14 helices, the first peripheral domain is formed by six helices (helices 1 to 6), the central domain is formed by at least two helices (helix 7 and helix 8), and the second peripheral domain is formed by six helices (helices 9 to 14), the suitable numbers of amino acid residues for helices 1 to 6 are, for example, as follows: Helix 1: 15, 22, 29, or 36 residues Helix 2: 13, 20, 27, or 34 residues Helix 3: 14, 21, 28, or 35 residues Helix 4: 11, 18, 25, or 32 residues Helix 5: 12, 19, 26, or 33 residues Helix 6: 13, 20, 27, or 34 residues

[0021] The combination of the number of residues in each helix may be any of the following combinations, in the order of helices 1 to 6: 15, 13, 14, 11, 12, and 13; 22, 20, 21, 18, 19, and 20; 29, 27, 28, 25, 26, and 27; and 36, 34, 35, 32, 33, and 34. This allows the helices to be connected by a linker with an appropriate number of residues.

[0022] The first peripheral domain may be formed, for example, by 6 to 12 helices. The number of amino acid residues forming each of the 6 to 12 helices may vary from one helix to another. This allows for the production of an additional protein of an appropriate size, which allows for effective structural analysis of the entire protein using, for example, an electron microscope.

[0023] Specifically, the number of helices forming the first peripheral domain is preferably a multiple of 2, particularly, for example, 6, 8, 10, or 12. This allows the N-terminus and C-terminus of the attached protein to face in the same direction when each helix forms a hairpin-type helix-loop-helix structure.

[0024] A predetermined linker may be formed between two of the helices forming the first peripheral domain. Here, the linker may be, for example, a loop structure formed by 1 to 10 amino acid residues, preferably a loop structure formed by 2 to 4 amino acid residues. The predetermined linker may also be, for example, another protein domain.

[0025] The first peripheral domain also includes a first connecting helix that connects to the first target helix contained in the target protein to form a series of helices. That is, the first target helix and the first connecting helix are connected at a connecting region, for example, a first connecting region. When the additional protein forms a tertiary structure consisting of the first peripheral domain, the central domain, and the second peripheral domain in this order from the N-terminus, the first connecting helix may be located on the N-terminus side of the first peripheral domain. The first connecting helix may be located contiguous with the helix most N-terminal of the helices that make up the first peripheral domain.

[0026] The second peripheral domain may have, for example, at least 90% of the amino acid residues on the surface of the second peripheral domain being hydrophilic amino acids, and at least 90% of the amino acid residues within the second hydrophobic core being hydrophobic amino acids. Note that the surface of the second peripheral domain may have, for example, amino acid residues that are not within the second hydrophobic core but that form hydrogen bonds with a solvent, such as water.

[0027] Each of the multiple helices forming the second peripheral domain may be a helix consisting of, for example, 10 to 40 amino acid residues, which can prevent the structural variation of the attached protein, i.e., the occurrence of so-called orientation bias, and can effectively perform structural analysis of the entire protein using, for example, an electron microscope.

[0028] Specifically, each of the multiple helices forming the second peripheral domain may be a helix consisting of, for example, 11 to 15, 18 to 22, 25 to 29, or 32 to 36 amino acid residues. More specifically, when the additional protein is composed of 14 helices, the first peripheral domain is formed by six helices (helices 1 to 6), the central domain is formed by at least two helices (helices 7 and 8), and the second peripheral domain is formed by six helices (helices 9 to 14), the suitable numbers of amino acid residues for helices 9 to 14 are, for example, as follows: Helix 9: 18, 23, 31, 39 residues Helix 10: 11, 18, 25, or 32 residues Helix 11: 16, 23, 31, or 39 residues Helix 12: 12, 19, 26, or 33 residues Helix 13: 13, 20, 27, or 34 residues Helix 14: 16, 23, 30, or 37 residues

[0029] The combination of the number of residues in each helix may be any of the following combinations, in the order of helices 9 to 14: 18, 11, 16, 12, 13, and 16; 23, 18, 23, 19, 20, and 23; 31, 25, 31, 26, 27, and 30; and 39, 32, 39, 33, 34, and 37. This allows the helices to be connected by a linker with an appropriate number of residues.

[0030] The second peripheral domain may be formed, for example, by 6 to 12 helices. The number of amino acid residues forming each of the 6 to 12 helices may vary from one helix to another. This allows for the production of an additional protein of an appropriate size, which allows for effective structural analysis of the entire protein using, for example, an electron microscope.

[0031] Specifically, the number of helices forming the second peripheral domain is preferably a multiple of 2, particularly, for example, 6, 8, 10, or 12. This allows the N-terminus and C-terminus of the additional protein to face in the same direction when each helix forms a hairpin-type helix-loop-helix structure.

[0032] A predetermined linker may be formed between two of the helices forming the second peripheral domain. Here, the linker may be, for example, a loop structure formed by 1 to 10 amino acid residues, preferably a loop structure formed by 2 to 4 amino acid residues. The predetermined linker may also be, for example, another protein domain.

[0033] The second peripheral domain also includes a second connecting helix that connects to a second target helix contained in the target protein to form a series of helices, i.e., the second target helix and the second connecting helix are connected at a connecting region, e.g., a second connecting region.

[0034] The number of helices forming the first peripheral domain and the number of helices forming the second peripheral domain may be the same or different, and the helices forming the first peripheral domain and the helices forming the second peripheral domain may all be different helices.

[0035] The third hydrophobic core may be formed by multiple helices, including at least one helix in the first peripheral domain and at least one helix in the second peripheral domain, which allows the additional protein to form a stable tertiary structure.

[0036] The helices forming the central domain are preferably formed, for example, by at least one helix forming the first peripheral domain, at least one helix forming the second peripheral domain, and at least two helices that do not form either the first or second peripheral domain.

[0037] The helices forming the central domain may, for example, include at least one helix that has fewer amino acids than the helices forming the first and second peripheral domains, allowing the helices forming the central domain to fit inside the additional protein, allowing the additional protein to form a stable tertiary structure.

[0038] Each of the multiple helices forming the central domain may be a helix consisting of, for example, 10 to 40 amino acid residues, which can prevent the structure of the attached protein from varying, i.e., prevent orientation bias, and can effectively perform structural analysis of the entire protein using, for example, an electron microscope.

[0039] Specifically, each of the multiple helices forming the central domain may be a helix consisting of, for example, 11 to 15, 18 to 22, 25 to 29, or 32 to 36 amino acid residues. More specifically, if the additional protein consists of 14 helices, the number of multiple helices forming the first peripheral domain is 6 (helices 1 to 6), the central domain is formed by at least two helices (helices 7 and 8), and the number of multiple helices forming the second peripheral domain is 6 (helices 9 to 14), the suitable numbers of amino acid residues for helices 7 and 8 are, for example, Helix 7: 15, 22, 29, or 36 residues Helix 8: 18, 25, or 32 residues

[0040] The combination of the number of residues in each helix may be any of the following, in the order of helices 7 and 8: 15 and 18, 22 and 18, 22 and 25, 29 and 25, 29 and 32, and 36. This allows the helices to be connected by a linker with an appropriate number of residues.

[0041] A predetermined linker may be formed between two of the helices forming the central domain. Here, the linker may be, for example, a loop structure formed by 1 to 10 amino acid residues, preferably a loop structure formed by 2 to 4 amino acid residues. The predetermined linker may also be, for example, another protein domain.

[0042] The amino acid composition of the multiple helices of the additional protein will be described using an example in which the additional protein is composed of 14 helices, the number of multiple helices forming the first peripheral domain is six (helices 1 to 6), the central domain is formed by at least two helices (helices 7 and 8), and the number of multiple helices forming the second peripheral domain is six (helices 9 to 14).

[0043] The amino acid composition of the helices located at the four corners of the first peripheral domain is preferably similar to that of the helices located at the four corners of the second peripheral domain, e.g., helix 1, helix 3, helix 4, and helix 6, and the helices located at the four corners of the second peripheral domain are preferably helix 9, helix 11, helix 12, and helix 14.

[0044] Preferably, the amino acid composition of the central helix of the first peripheral domain is similar to that of the central helix of the second peripheral domain, e.g., helix 2 and helix 5, and the central helix of the second peripheral domain is e.g., helix 10 and helix 13.

[0045] The amino acid composition of each helix is ​​specifically described below. If the amino acid residue closest to the N-terminus of helix 1 is designated as 1, then the amino acid residues at positions 2, 5, 6, 13, 14, 16, 24, and 28 are preferably hydrophobic amino acid residues. Furthermore, in this case, the amino acid residues at positions 1, 3, 4, 7, 8, 11, 12, 15, 18, 19, 21, 22, 23, 25, and 26 are preferably charged amino acid residues. The other amino acid residues constituting helix 1 are preferably hydrophobic amino acids or polar, uncharged amino acids.

[0046] If the amino acid residue closest to the N-terminus of helix 2 is designated as 1, then the amino acid residues at positions 1, 2, 4, 12, 13, 16, 17, 21, 23, 26, and 27 are preferably hydrophobic amino acid residues. In this case, the amino acid residues at positions 3, 7, 8, 10, 11, 14, 18, 22, and 25 are preferably charged amino acid residues. The other amino acid residues constituting helix 2 are preferably hydrophobic amino acids or polar, uncharged amino acids.

[0047] If the amino acid residue closest to the N-terminus of helix 3 is designated as 1, then the amino acid residues at positions 3, 4, 7, 11, 14, 15, 22, and 26 are preferably hydrophobic amino acid residues. In this case, the amino acid residues at positions 1, 2, 5, 6, 8, 9, 10, 12, 13, 16, 17, 19, 20, 21, 23, 24, 27, and 28 are preferably charged amino acid residues. The other amino acid residues constituting helix 3 are preferably hydrophobic amino acids or polar, uncharged amino acids.

[0048] If the amino acid residue closest to the N-terminus of helix 4 is designated as 1, then the amino acid residues at positions 1, 2, 4, 5, 12, 19, and 20 are preferably hydrophobic amino acid residues. In this case, the amino acid residues at positions 3, 6, 7, 8, 10, 11, 13, 14, 15, 17, 18, 21, and 24 are preferably charged amino acid residues. The other amino acid residues constituting helix 4 are preferably hydrophobic amino acids or polar, uncharged amino acids.

[0049] If the amino acid residue closest to the N-terminus of helix 4 is designated as 1, then the amino acid residues at positions 1, 2, 4, 5, 12, 19, and 20 are preferably hydrophobic amino acid residues. In this case, the amino acid residues at positions 3, 6, 7, 8, 10, 11, 13, 14, 15, 17, 18, 21, and 24 are preferably charged amino acid residues. The other amino acid residues constituting helix 4 are preferably hydrophobic amino acids or polar, uncharged amino acids.

[0050] If the amino acid residue closest to the N-terminus of helix 5 is designated as 1, then the amino acid residues at positions 4, 7, 10, 11, 13, 14, 18, 19, 21, 22, and 25 are preferably hydrophobic amino acid residues. In this case, the amino acid residues at positions 2, 5, 6, 9, 12, 15, 16, 20, 23, and 26 are preferably charged amino acid residues. The other amino acid residues constituting helix 5 are preferably hydrophobic amino acids or polar, uncharged amino acids.

[0051] If the amino acid residue closest to the N-terminus of helix 6 is designated as 1, then the amino acid residues at positions 1, 4, 5, 6, 11, 12, 16, 17, and 18 are preferably hydrophobic amino acid residues. In this case, the amino acid residues at positions 2, 3, 8, and 13 are preferably charged amino acid residues. The other amino acid residues constituting helix 6 are preferably hydrophobic amino acids or polar, uncharged amino acids.

[0052] If the amino acid residue closest to the N-terminus of helix 7 is designated as 1, then the amino acid residues at positions 5, 8, 9, and 12 are preferably hydrophobic amino acid residues. In this case, the amino acid residues at positions 2, 3, 6, 7, 13, and 14 are preferably charged amino acid residues. The other amino acid residues constituting helix 7 are preferably hydrophobic amino acids or polar, uncharged amino acids.

[0053] If the amino acid residue closest to the N-terminus of helix 8 is designated as 1, then the amino acid residues at positions 3, 4, 7, 8, 10, 11, 14, 15, and 17 are preferably hydrophobic amino acid residues. In this case, the amino acid residues at positions 2, 5, 6, 9, 12, 13, 16, and 18 are preferably charged amino acid residues. The other amino acid residues constituting helix 7 are preferably hydrophobic amino acids or polar, uncharged amino acids.

[0054] If the amino acid residue closest to the N-terminus of helix 9 is designated as 1, then the amino acid residues at positions 3, 4, 7, 10, 11, 14, 17, and 21 are preferably hydrophobic amino acid residues. In this case, the amino acid residues at positions 1, 2, 5, 8, 9, 12, 15, 16, 18, 19, and 23 are preferably charged amino acid residues. The other amino acid residues constituting helix 9 are preferably hydrophobic amino acids or polar, uncharged amino acids.

[0055] If the amino acid residue closest to the N-terminus of helix 10 is designated as 1, then the amino acid residues at positions 1, 3, 4, 5, 7, 11, 12, 15, 18, 21, and 22 are preferably hydrophobic amino acid residues. In this case, the amino acid residues at positions 6, 9, 10, 13, 17, 20, 24, and 25 are preferably charged amino acid residues. The other amino acid residues constituting helix 10 are preferably hydrophobic amino acids or polar, uncharged amino acids.

[0056] If the amino acid residue closest to the N-terminus of helix 11 is designated as 1, then the amino acid residues at positions 1, 3, 4, 7, 8, 14, 15, 18, 19, and 23 are preferably hydrophobic amino acid residues. In this case, the amino acid residues at positions 2, 5, 6, 9, 13, 16, 17, 20, 21, and 23 are preferably charged amino acid residues. The other amino acid residues constituting helix 11 are preferably hydrophobic amino acids or polar, uncharged amino acids.

[0057] If the amino acid residue closest to the N-terminus of helix 12 is designated as 1, then the amino acid residues at positions 1, 8, 12, 15, and 26 are preferably hydrophobic amino acid residues. In this case, the amino acid residues at positions 2, 3, 4, 5, 6, 7, 9, 10, 11, 13, 14, 16, 17, 18, 20, 21, 22, and 24 are preferably charged amino acid residues. The other amino acid residues constituting helix 12 are preferably hydrophobic amino acids or polar, uncharged amino acids.

[0058] If the amino acid residue closest to the N-terminus of helix 13 is designated as 1, then the amino acid residues at positions 1, 3, 4, 7, 8, 11, 14, 15, 18, 19, 22, and 25 are preferably hydrophobic amino acid residues. In this case, the amino acid residues at positions 2, 5, 9, 17, 20, 23, 24, 26, and 27 are preferably charged amino acid residues. The other amino acid residues constituting helix 13 are preferably hydrophobic amino acids or polar, uncharged amino acids.

[0059] If the amino acid residue closest to the N-terminus of helix 14 is designated as 1, then the amino acid residues at positions 2, 4, 5, 8, 9, 10, 14, 16, 17, 19, and 20 are preferably hydrophobic amino acid residues. In this case, the amino acid residues at positions 3, 7, 18, 21, 24, 25, 26, 27, 28, 29, and 30 are preferably charged amino acid residues. The other amino acid residues constituting helix 14 are preferably hydrophobic amino acids or polar, uncharged amino acids.

[0060] The target protein may be, for example, a G protein-coupled receptor, in which case the first target helix may be, for example, the fifth transmembrane domain of the G protein-coupled receptor, and the second target helix may be, for example, the sixth transmembrane domain of the G protein-coupled receptor.

[0061] Although not particularly limited, the present invention will be described with reference to Fig. 1 for better understanding. Fig. 1A and Fig. 1B are diagrams showing the structure of an additional protein 100 according to one embodiment of the present disclosure. Fig. 1B is a diagram showing the additional protein 100 shown in Fig. 1A as viewed from below.

[0062] The additional protein 100 is a protein to be added to the target protein 200. The additional protein 100 and the target protein 200 are connected at a first connection region 310a and a second connection region 310b. The additional protein 100 in this embodiment is a novel protein that is added to the target protein 200, for example, for stabilization or structural analysis of the target protein 200.

[0063] 2A and 2B are diagrams showing the structure of the additional protein 100. The structure shown in Fig. 2A is shown in a format that displays the side chains of the amino acids of the additional protein 100.

[0064] Figure 2B shows the structure shown in Figure 2A, viewed from below the additional protein 100 in Figure 2A. The additional protein 100 forms a tertiary structure including a first peripheral domain 110, a second peripheral domain 120, and a central domain 130.

[0065] The additional protein 100 forms a tertiary structure consisting of, for example, a first peripheral domain 110, a central domain 130, and a second peripheral domain 120 in this order from the N-terminus.

[0066] The first peripheral domain 110, the second peripheral domain 120, and the central domain 130 are each formed, for example, by a plurality of helices.

[0067] Furthermore, the first peripheral domain 110, the second peripheral domain 120, and the central domain 130 each include a first hydrophobic core 111, a second hydrophobic core 121, and a third hydrophobic core 131, each formed by, for example, a plurality of helices.

[0068] The first target helix and the first connecting helix connect at a first connecting region 310a, and the second target helix and the second connecting helix connect at a second connecting region 310b.

[0069] Figure 3 is a diagram showing the positional relationship of helices contained in an additional protein. The diagram shown in Figure 3 is an example in which the additional protein consists of 14 helices. In Figure 3, the numbers on the helices indicate the helix numbers from the N-terminus.

[0070] Helices 1 to 6 form the first peripheral domain. Helices 9 to 14 form the second peripheral domain. Helices 6, 7, 8, and 14 form the central domain.

[0071] 4A to 4D are schematic diagrams showing the helices that form the first peripheral domain, in which the numbers of the helices indicate the number of the helices from the N-terminus.

[0072] Figure 4A is a schematic diagram of a first peripheral domain consisting of six helices. In this case, the helices located at the four corners of the first peripheral domain are, for example, helix 1, helix 3, helix 4, and helix 6. Furthermore, in this case, the helices located in the center of the first peripheral domain are helix 2 and helix 5. The helices located at the four corners of the first peripheral domain in Figure 4A (i.e., helix 1, helix 3, helix 4, and helix 6 in Figure 4A) preferably have the same amino acid composition and number of residues as helix 1, helix 3, helix 4, and helix 6 of the additional protein 100 shown in the amino acid sequence of SEQ ID NO: 1, which will be described later. Furthermore, the helices located in the center of the first peripheral domain in Figure 4A (i.e., helix 2 and helix 5 in Figure 4A) preferably have the same amino acid composition and number of residues as helix 2 and helix 5 of the additional protein 100 shown in the amino acid sequence of SEQ ID NO: 1, which will be described later.

[0073] Figure 4B is a schematic diagram of a first peripheral domain consisting of eight helices. In this case, the helices located at the four corners of the first peripheral domain are, for example, helix 1, helix 4, helix 5, and helix 8. In addition, in this case, the helices located in the center of the first peripheral domain are helix 2, helix 3, helix 6, and helix 7. The helices located at the four corners of the first peripheral domain in Figure 4B (i.e., helix 1, helix 4, helix 5, and helix 8 in Figure 4B) preferably have the same amino acid composition and number of residues as helix 1, helix 3, helix 4, and helix 6 of additional protein 100 shown in the amino acid sequence of SEQ ID NO: 1, which will be described later. Furthermore, the helix located in the center of the first peripheral domain in FIG. 4B (i.e., helix 2, helix 3, helix 6, and helix 7 in FIG. 4B) preferably has the same amino acid composition and number of residues as helix 2 or helix 5 of additional protein 100 shown in the amino acid sequence of SEQ ID NO: 1 described below.

[0074] Figure 4C is a schematic diagram of a first peripheral domain consisting of 10 helices. In this case, the helices located at the four corners of the first peripheral domain are, for example, helix 1, helix 5, helix 6, and helix 10. In addition, in this case, the helices located in the center of the first peripheral domain are helix 2 to helix 4 and helix 7 to helix 9. The helices located at the four corners of the first peripheral domain in Figure 4C (i.e., helix 1, helix 5, helix 6, and helix 10 in Figure 4C) preferably have the same amino acid composition and number of residues as helix 1, helix 3, helix 4, and helix 6 of additional protein 100 shown in the amino acid sequence of SEQ ID NO: 1, which will be described later. Furthermore, the helices located in the center of the first peripheral domain in FIG. 4C (i.e., helices 2 to 4 and helices 7 to 9 in FIG. 4C) preferably have the same amino acid composition and number of residues as helix 2 or helix 5 of additional protein 100 shown in the amino acid sequence of SEQ ID NO: 1, which will be described later.

[0075] Figure 4D is a schematic diagram of a first peripheral domain consisting of 12 helices. In this case, the helices located at the four corners of the first peripheral domain are, for example, helix 1, helix 6, helix 7, and helix 12. In addition, in this case, the helices located in the center of the first peripheral domain are helix 2 to helix 5 and helix 8 to helix 11. The helices located at the four corners of the first peripheral domain in Figure 4D (i.e., helix 1, helix 6, helix 7, and helix 12 in Figure 4D) preferably have the same amino acid composition and number of residues as helix 1, helix 3, helix 4, and helix 6 of additional protein 100 shown in the amino acid sequence of SEQ ID NO: 1, which will be described later. Furthermore, the helices located in the center of the first peripheral domain in FIG. 4D (i.e., helices 2 to 5 and helices 8 to 11 in FIG. 4D) preferably have the same amino acid composition and number of residues as helix 2 or helix 5 of additional protein 100 shown in the amino acid sequence of SEQ ID NO: 1, which will be described later.

[0076] 5A to 5D are schematic diagrams showing the helices that form the second peripheral domain, in which the numbers next to the helices indicate the helix numbers from the N-terminus.

[0077] Figure 5A is a schematic diagram of a second peripheral domain consisting of six helices. In this case, the helices located at the four corners of the second peripheral domain are, for example, helix 9, helix 11, helix 12, and helix 14. Furthermore, in this case, the helices located in the center of the second peripheral domain are helix 10 and helix 13. The helices located at the four corners of the second peripheral domain in Figure 5A (i.e., helix 9, helix 11, helix 12, and helix 14 in Figure 5A) preferably have the same amino acid composition and number of residues as helix 9, helix 11, helix 12, and helix 14 of the additional protein 100 shown in the amino acid sequence of SEQ ID NO: 1, which will be described later. Furthermore, the helices located in the center of the second peripheral domain in Figure 5A (i.e., helix 10 and helix 13 in Figure 5A) preferably have the same amino acid composition and number of residues as helix 2 and helix 5 of the additional protein 100 shown in the amino acid sequence of SEQ ID NO: 1, which will be described later.

[0078] Figure 5B is a schematic diagram of a second peripheral domain consisting of eight helices. In this case, the helices located at the four corners of the second peripheral domain are, for example, helix 9, helix 12, helix 13, and helix 16. In this case, the helices located in the center of the second peripheral domain are helix 10, helix 11, helix 14, and helix 15. The helices located at the four corners of the second peripheral domain in Figure 5B (i.e., helix 9, helix 12, helix 13, and helix 16 in Figure 5B) preferably have the same amino acid composition and number of residues as helix 9, helix 11, helix 12, and helix 14 of additional protein 100 shown in the amino acid sequence of SEQ ID NO: 1, which will be described later. Furthermore, the helices located in the center of the second peripheral domain in FIG. 5B (i.e., helix 10, helix 11, helix 14, and helix 15 in FIG. 5B) preferably have the same amino acid composition and number of residues as helix 10 and helix 13 of additional protein 100 shown in the amino acid sequence of SEQ ID NO: 1, which will be described later.

[0079] Figure 5C is a schematic diagram of a second peripheral domain consisting of 10 helices. In this case, the helices located at the four corners of the second peripheral domain are, for example, helix 9, helix 13, helix 14, and helix 18. In addition, in this case, the helices located in the center of the second peripheral domain are helices 10 to 12 and helices 15 to 17. The helices located at the four corners of the second peripheral domain in Figure 5C (i.e., helix 9, helix 13, helix 14, and helix 18 in Figure 5C) preferably have the same amino acid composition and number of residues as helix 9, helix 11, helix 12, and helix 14 of additional protein 100 shown in the amino acid sequence of SEQ ID NO: 1, which will be described later. Furthermore, the helices located in the center of the second peripheral domain in FIG. 5C (i.e., helices 10 to 12 and helices 15 to 17 in FIG. 5C) preferably have the same amino acid composition and number of residues as helix 10 and helix 13 of additional protein 100 shown in the amino acid sequence of SEQ ID NO: 1, which will be described later.

[0080] Figure 5D is a schematic diagram of a second peripheral domain consisting of 12 helices. In this case, the helices located at the four corners of the second peripheral domain are, for example, helix 9, helix 14, helix 15, and helix 20. In addition, in this case, the helices located in the center of the second peripheral domain are helices 10 to 13 and helices 16 to 19. The helices located at the four corners of the second peripheral domain in Figure 5D (i.e., helix 9, helix 14, helix 15, and helix 20 in Figure 5D) preferably have the same amino acid composition and number of residues as helix 9, helix 11, helix 12, and helix 14 of additional protein 100 shown in the amino acid sequence of SEQ ID NO: 1, which will be described later. Furthermore, the helices located in the center of the second peripheral domain in FIG. 5D (i.e., helices 10 to 13 and helices 16 to 19 in FIG. 5D) preferably have the same amino acid composition and number of residues as helix 10 and helix 13 of additional protein 100 shown in the amino acid sequence of SEQ ID NO: 1, which will be described later.

[0081] A method for designing the amino acid sequence of the added protein 100 on a computer will now be described. Figure 6 is a diagram showing an example of a method for designing the added protein 100.

[0082] First, the first connection region 310a and the second connection region 310b are identified in the target protein 200. When the target protein 200 is a G protein-coupled receptor, the first connection region 310a and the second connection region 310b in the target protein 200 may be, for example, a predetermined amino acid residue in the fifth transmembrane domain of the G protein-coupled receptor and a predetermined amino acid residue in the sixth transmembrane domain of the G protein-coupled receptor, respectively.

[0083] Next, a predetermined first program is used to design the main chain structure of the attached protein 100 (S301). Here, the predetermined first program may be, for example, a program that can virtually design the main chain structure of a protein on a computer, in particular, for example, Foldit Standalone.

[0084] Specifically, using a predetermined first program, multiple helices are first designed starting from the first connecting region 310a. The multiple helices to be designed may be, for example, multiple helices consisting of 10 to 40 amino acid residues connected by amino acid residues that form a predetermined loop structure. Multiple helices starting from the second connecting region 310b are similarly designed. The main chain structure of the added protein 100 is then obtained by connecting the end (e.g., C-terminus) of the helix starting from the first connecting region 310a with the end (e.g., N-terminus) of the helix starting from the second connecting region 310b. At this time, the main chain structure of the added protein 100 is designed so as to form a tertiary structure including the first peripheral domain 110, the second peripheral domain 120, and the central domain 130. The number of helices in the added protein 100 may be, for example, 14 to 26.

[0085] Next, a predetermined second program is used to add side chains to the main chain structure of the added protein 100 (S302). The predetermined second program may be, for example, a program that predicts appropriate side chains that are thought to form the main chain structure based on the main chain structure, particularly, for example, ProteinMPNN. This allows the amino acid sequence of the added protein 100 to be predicted.

[0086] Next, at least one of the hydrophobic amino acid residues on the surface of the first peripheral domain 110 and the second peripheral domain 120 is replaced with a hydrophilic amino acid residue, and at least one of the hydrophilic amino acid residues predicted to exist within the first hydrophobic core 111, the second hydrophobic core 121, and the third hydrophobic core 131 is replaced with a hydrophobic amino acid residue (S303).

[0087] After the substitution, a predetermined third program is used to predict the structure of the substituted added protein 100. The predetermined third program may be, for example, a program that inputs a predicted structure of a protein and outputs a more likely predicted structure of the protein, particularly, for example, ColabFold.

[0088] Substitution and prediction are again performed on the predicted structure of the added protein 100 (S304). This operation is repeated multiple times to obtain a protein that is defined as the added protein 100 (S305).

[0089] A method for synthesizing the additional protein 100 will now be described.

[0090] First, a whole protein in which the additional protein 100 is inserted into the target protein 200 can be synthesized using, for example, a known biochemical method using a nucleic acid sequence corresponding to the amino acid sequence and a predetermined expression system. The expressed whole protein can be purified, for example, using a known protein purification technique, to obtain the whole protein.

[0091] An example will be described.

[0092] In this Example 1, the vasopressin V2 receptor, which is one of the G protein-coupled receptors, was used as the target protein 200. Figure 1 shows an additional protein 100 and target protein 200 corresponding to this Example 1.

[0093] Based on the design method for additional protein 100, the N-terminus of additional protein 100 containing 14 helices was designed to be connected to S235 of target protein 200 (in this case, vasopressin V2 receptor), and the C-terminus of additional protein 100 was designed to be connected to V266 of target protein 200 (in this case, vasopressin V2 receptor).

[0094] The first peripheral domain 110 of the additional protein 100 in this Example 1 was designed to be formed by six helices, and the second peripheral domain 120 of the additional protein 100 in this Example 1 was designed to be formed by six helices. Furthermore, the central domain of the additional protein 100 in this Example 1 was designed to be formed by two helices contained in the first peripheral domain 110, two helices contained in the second peripheral domain 120, and two more helices.

[0095] In this Example 1, of the amino acid residues on the surface of the first peripheral domain 110 of the additional protein 100, 90% or more were hydrophilic amino acids, and of the amino acid residues within the first hydrophobic core 111 of the first peripheral domain 110, 90% or more were hydrophobic amino acids.Furthermore, of the amino acid residues on the surface of the second peripheral domain 120 of the additional protein 100 in this Example 1, 90% or more were hydrophilic amino acids, and of the amino acid residues within the second hydrophobic core 121 of the second peripheral domain 120, 90% or more were hydrophobic amino acids.

[0096] An additional protein 100 has been described as one embodiment of the present disclosure. The additional protein 100 is a protein that forms a predetermined tertiary structure including, for example, a central domain, a first peripheral domain, and a second peripheral domain, and the central domain, the first peripheral domain, and the second peripheral domain each include a hydrophobic core formed by multiple helices, and the hydrophobic core of the central domain comprises a structure formed by multiple helices including at least one helix included in the first peripheral domain and at least one helix included in the second peripheral domain. In other words, the additional protein 100 may be any protein that has this structure, and is not limited to the amino acid sequence shown in Table 1. The inventors have discovered that the additional protein 100 having this structure is effective for addition to a target protein 200, for example, in stabilizing or structural analysis of the target protein 200.

[0097] In this regard, it has been known in the technical field of the present disclosure that proteins with similar structures may exist even if they have different amino acid sequences. In recent years, advances in information processing technology have made it possible to artificially design multiple proteins with similar structures even if they have different amino acid sequences. In light of this state of technological development, the significance of the present disclosure lies not only in providing novel proteins with the amino acid sequences shown in Table 1, but also in providing novel additional proteins 100 with the structure described in this embodiment.

[0098] Furthermore, in conventional techniques for attaching other proteins (e.g., bRIL or T4 lysozyme) to a target protein, antibodies may be used for the purpose of higher-resolution three-dimensional structural analysis. However, because the attached protein 100 has an appropriate molecular weight, there is no need to use an additional antibody.

[0099] Furthermore, since the added protein 100 has high thermal stability, the expression level of the entire protein in which the added protein 100 is added to the target protein can be increased, and the three-dimensional structural analysis of the target protein can be carried out effectively.

[0100] Furthermore, since the structure of the added protein 100 is less flexible, higher resolution structural analysis can be performed.

[0101] It should be noted that the "protein to be added to a target protein" in the present disclosure includes, for example, a protein that has been added to a target protein or a protein that is being added to a target protein.

[0102] Furthermore, among the amino acids in this embodiment, hydrophobic amino acids include, for example, glycine, alanine, valine, leucine, isoleucine, methionine, proline, tyrosine, phenylalanine, and tryptophan.

[0103] Furthermore, among the amino acids in this embodiment, hydrophilic amino acids include, for example, polar uncharged amino acids such as serine, threonine, asparagine, glutamine, and histidine, and also include charged amino acids such as arginine, lysine, aspartic acid, glutamic acid, and cysteine.

[0104] The amino acid sequence of additional protein 100 in this Example 1 is as shown in Table 1. In Table 1, amino acid residues are represented by single-letter codes. Here, additional protein 100 represented by the amino acid sequence of SEQ ID NO: 1 is composed of 14 helices, with the first peripheral domain being formed by helices 1 to 6, the second peripheral domain being formed by helices 9 to 14, and the central domain being formed by helices 6, 7, 8, and 14.

[0105]

[0106] The coordinates of each atom of the added protein 100 in this Example 1 are as shown in Table 2. In Table 2, amino acid residues are shown in three-letter notation. The X coordinate, Y coordinate, and Z coordinate are shown as relative coordinates from a predetermined origin. Each of the atom types listed in Table 2 will be clear to those skilled in the art (see, for example, http: / / www.schmieder.fmp-berlin.info / teaching / educational_scripts / pdf / aminoacids.pdf).

[0107]

[0108]

[0109]

[0110]

[0111]

[0112]

[0113]

[0114]

[0115]

[0116]

[0117]

[0118]

[0119]

[0120]

[0121]

[0122]

[0123]

[0124]

[0125]

[0126]

[0127]

[0128]

[0129]

[0130]

[0131]

[0132]

[0133]

[0134]

[0135]

[0136]

[0137]

[0138]

[0139]

[0140]

[0141]

[0142]

[0143]

[0144]

[0145]

[0146]

[0147]

[0148]

[0149]

[0150]

[0151]

[0152]

[0153]

[0154]

[0155]

[0156]

[0157]

[0158]

[0159]

[0160]

[0161]

[0162]

[0163]

[0164]

[0165]

[0166]

[0167]

[0168]

[0169]

Claims

1. A protein for attachment to a target protein, said protein forming a predetermined tertiary structure comprising a central domain, a first peripheral domain, and a second peripheral domain, said central domain, said first peripheral domain, and said second peripheral domain each comprising a hydrophobic core formed by a plurality of helices, said hydrophobic core of said central domain being formed by a plurality of helices including at least one helix comprised in said first peripheral domain and at least one helix comprised in said second peripheral domain.

2. The protein of claim 1, wherein at least 90% of the amino acid residues on the surface of the first peripheral domain are hydrophilic amino acids, and at least 90% of the amino acid residues within the hydrophobic core of the first peripheral domain are hydrophobic amino acids.

3. The protein of claim 1, wherein at least 90% of the amino acid residues on the surface of the second peripheral domain are hydrophilic amino acids, and at least 90% of the amino acid residues within the hydrophobic core of the second peripheral domain are hydrophobic amino acids.

4. A protein according to any one of claims 1 to 3, wherein the helix forming the first peripheral domain and the helix forming the second peripheral domain consist of 10 to 40 amino acid residues.

5. A protein according to any one of claims 1 to 3, wherein the helices forming the central domain comprise at least one helix having fewer amino acids than the helices forming the first peripheral domain and the second peripheral domain.

6. A protein according to any one of claims 1 to 3, wherein the first peripheral domain is formed by 6 to 12 helices, and the second peripheral domain is formed by 6 to 12 helices that are all different from the 6 to 12 helices that form the first peripheral domain.

7. The protein of claim 6, wherein the number of helices forming the first peripheral domain is the same as the number of helices forming the second peripheral domain.

8. The protein according to any one of claims 1 to 3, wherein the predetermined tertiary structure comprises, in order from the N-terminus, the first peripheral domain, the central domain, and the second peripheral domain; the helix forming the first peripheral domain comprises a first connecting helix that connects with a first target helix contained in the target protein to form a series of helices; and the helix forming the second peripheral domain comprises a second connecting helix that connects with a second target helix contained in the target protein to form a series of helices.

9. The protein of claim 8, wherein the target protein is a G protein-coupled receptor, the first connecting helix connects to the first target helix, which is the fifth transmembrane domain of the G protein-coupled receptor, to form a series of helices, and the second connecting helix connects to the second target helix, which is the sixth transmembrane domain of the G protein-coupled receptor, to form a series of helices.

Citation Information

Patent Citations

  • GPCR: Binding domains created for G protein complexes and their derived uses

    JP2014525735A

  • Fusion proteins comprising cytokines and scaffold proteins

    JP2022515150A

  • Anti-BRIL antibody and stabilization method of BRIL fusion protein using said antibody

    WO2019013308A1

  • Sequence extraction system, sequence extraction method, and sequence extraction program

    WO2024143531A1