Split Intein and its Use
Intein-degron combinations address the accumulation of protein intermediates in gene therapy by enhancing splicing yields and reducing unwanted by-products, thereby improving the efficiency of protein production.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- スプライスバイオ エセエレ
- Filing Date
- 2021-03-26
- Publication Date
- 2026-06-01
AI Technical Summary
The accumulation of protein intermediates and cleaved split inteins during gene therapy using adeno-associated viruses (AAVs) hinders the efficient production of target proteins, particularly for genes that cannot be encapsulated within AAVs.
Combining specific inteins with degradation signals, known as degrons, to reduce the accumulation of protein intermediates and improve splicing yields by using intein-degron combinations.
The use of intein-degron combinations significantly reduces protein intermediate accumulation and enhances the yield of reconstituted proteins, improving the efficiency of gene therapy by eliminating unwanted by-products.
Smart Images

Figure 0007867703000016 
Figure 0007867703000017 
Figure 0007867703000018
Abstract
Description
[Technical Field]
[0001] This invention falls under the field of biotechnology and specifically relates to split-intein and its use. [Background technology]
[0002] It is known that genes that cannot be encapsulated within adeno-associated viruses (AAVs) for gene therapy can be split into two fragments. Therefore, each fragment must be fused to a split intein and delivered within its respective AAV. Infecting target cells with the intein allows for the production of the desired target protein (see Figure 1).
[0003] Nevertheless, one of the key limitations of this approach is the accumulation of protein intermediates (N-extine-ItnN and IntC-C-extine) and cleaved split inteins. To address this challenge, in this application, the inventors have shown that the accumulation of such protein intermediates can be significantly reduced by combining certain inteins with certain degradation signals. This invention focuses on these intein-degron combinations. The inventors have also identified certain intein-degron combinations that contribute to improved splicing yields in vivo. [Overview of the Initiative]
[0004] In a preferred first embodiment, the present invention is a composition, a. A first polynucleotide encoding a polypeptide containing a split-intene N fragment, wherein the split-intene N fragment is optionally directly linked via a peptide bond to the N-terminal fragment of the reconstituted protein, and is selected from the group consisting of any functionally equivalent variants such as CfaN of SEQ ID NO: 27, CatN of SEQ ID NO: 30, and Gp41N of SEQ ID NO: 38, or ConN of SEQ ID NO: 39, and b. A second polynucleotide encoding a polypeptide containing a split-intane C fragment, wherein the split-intane C fragment is optionally directly bound to the C-terminal fragment of the reconstituted protein via a peptide bond by a peptide linker, and the second polynucleotide is selected from the group consisting of CfaC (SEQ ID NO: 28), CatC (SEQ ID NO: 31), Gp41C (SEQ ID NO: 104), or any functionally equivalent variant thereof such as CfaCmut (SEQ ID NO: 29) or ConC (SEQ ID NO: 105), Includes, Both polynucleotides of the composition may be filled together in a single formulation, or they may be filled separately in different formulations. The first and second polynucleotides respectively encode the N-terminal and C-terminal fragments of the reconstituted protein, such that when both fragments are combined, the N-terminal fragment of the protein is linked to the C-terminal fragment of the protein to generate the entire protein. The protein being reconstituted is over 25 kDa, The treatment method is further, The split intein N fragment is further directly linked to degron via peptide bonds, the degron is linked to the intein N fragment via the C-terminus of intein with or without a linker between the intein N fragment and degron, the N-terminus of the split intein N fragment is directly linked to the N-terminal fragment of the reconstituted protein via peptide bonds, and / or The split intein C fragment is further directly linked to degron via a peptide bond, and degron is linked to the intein C fragment via the N-terminus of intein, with or without a linker between the intein C fragment and degron, and the C-terminus of the split intein C fragment is directly linked to the C-terminal fragment of the reconstituted protein via a peptide bond. A composition characterized by the following features.
[0005] In a preferred embodiment of the first aspect, the deglon directly linked to the split-intane N fragment and / or split-intane C fragment is selected from the group consisting of CL1, Deg1, PEST, DD1, DD2, DD3, M1, M2, SopE, SopE-1-78, SopE-15-78, SopE-15-50, L2, L6, L9, L10, L11, L12, L15, L16, M3, M4, M5, V12, DD4, DD5, DD6, and DD7. Preferably, the deglon directly linked to the split-intane N fragment and / or split-intane C fragment is selected from the group consisting of DD1, DD3, PEST, SopE, L2, L9, M4, or V12. More preferably, the deglon directly linked to the split-intei N fragment and / or split-intei C fragment is selected from the group consisting of SopE, L2, L9, M4, or V12.
[0006] In another preferred embodiment of the first aspect, the split-intane N fragment is CfaN of SEQ ID NO: 27, directly linked to the N-terminal fragment of the reconstituted protein via a peptide bond, optionally by a peptide linker; the split-intane C fragment is CfaC of SEQ ID NO: 28, directly linked to the C-terminal fragment of the reconstituted protein via a peptide bond, optionally by a peptide linker; preferably, the degrons directly linked to the split-intane N fragment and / or split-intane C fragment are CL1, Deg1, PEST, DD1, DD2, DD3, M1, M2, SopE, S The deglon directly linked to the split-intane N fragment and / or split-intane C fragment is selected from the group consisting of opE-1-78, SopE-15-78, SopE-15-50, L2, L6, L9, L10, L11, L12, L15, L16, M3, M4, M5, V12, DD4, DD5, DD6, and DD7, and more preferably, the deglon directly linked to the split-intane N fragment and / or split-intane C fragment is selected from the group consisting of DD1, DD3, PEST, SopE, L2, L9, M4, or V12, and even more preferably, the deglon directly linked to the split-intane N fragment and / or split-intane C fragment is selected from the group consisting of SopE, L2, L9, M4, or V12.
[0007] In another preferred embodiment of the first aspect, the split-intane N fragment is CfaN of SEQ ID NO: 27, directly linked to the N-terminal fragment of the reconstituted protein via a peptide bond, optionally by a peptide linker; the split-intane C fragment is CfaCmut of SEQ ID NO: 29, directly linked to the C-terminal fragment of the reconstituted protein via a peptide bond, optionally by a peptide linker; preferably, the degron directly linked to the split-intane N fragment and / or the split-intane C fragment is CL1, Deg1, PEST, DD1, DD2, DD3, M1, M2, SopE, The deglon directly linked to the split-intane N fragment and / or split-intane C fragment is selected from the group consisting of SopE-1-78, SopE-15-78, SopE-15-50, L2, L6, L9, L10, L11, L12, L15, L16, M3, M4, M5, V12, DD4, DD5, DD6, and DD7, and more preferably, the deglon directly linked to the split-intane N fragment and / or split-intane C fragment is selected from the group consisting of DD1, DD3, PEST, SopE, L2, L9, M4, or V12, and even more preferably, the deglon directly linked to the split-intane N fragment and / or split-intane C fragment is selected from the group consisting of SopE, L2, L9, M4, or V12.
[0008] In another preferred embodiment of the first aspect, the split-intane N fragment is NpuN of SEQ ID NO: 32, directly linked to the N-terminal fragment of the reconstituted protein via a peptide bond, optionally by a peptide linker; the split-intane C fragment is CfaCmut of SEQ ID NO: 29 or NpuCmut of SEQ ID NO: 36, directly linked to the C-terminal fragment of the reconstituted protein via a peptide bond, optionally by a peptide linker; preferably, the degrons directly linked to the split-intane N fragment and / or split-intane C fragment are CL1, Deg1, PEST, DD1, DD2, DD3, M1, The deglon directly linked to the split-intane N fragment and / or split-intane C fragment is selected from the group consisting of M2, SopE, SopE-1-78, SopE-15-78, SopE-15-50, L2, L6, L9, L10, L11, L12, L15, L16, M3, M4, M5, V12, DD4, DD5, DD6, and DD7, and more preferably, the deglon directly linked to the split-intane N fragment and / or split-intane C fragment is selected from the group consisting of DD1, DD3, PEST, SopE, L2, L9, M4, or V12, and even more preferably, the deglon directly linked to the split-intane N fragment and / or split-intane C fragment is selected from the group consisting of SopE, L2, L9, M4, or V12.
[0009] In another preferred embodiment of the first aspect, the composition is characterized by comprising two degrons, and in particular, the composition is The split intein N fragment is further directly linked to degron via a peptide bond, and degron is linked to the intein N fragment via the C-terminus of intein, with or without a linker between the intein N fragment and degron, and the N-terminus of the split intein N fragment is directly linked to the N-terminal fragment of the reconstituted protein via a peptide bond. The split intein C fragment is further directly linked to degron via a peptide bond, and degron is linked to the intein C fragment via the N-terminus of intein, with or without a linker between the intein C fragment and degron, and the C-terminus of the split intein C fragment is directly linked to the C-terminal fragment of the reconstituted protein via a peptide bond. It is characterized by the following:
[0010] In the first embodiment or any of its preferred embodiments, the composition is a. Sequence ID No. 27, which is directly linked via peptide bonds to the N-terminal fragment of the reconstituted protein, optionally by a peptide linker, and optionally further linked via peptide bonds to a deglon selected from the group consisting of CL1, Deg1, PEST, DD1, DD2, DD3, M1, M2, SopE, SopE-1-78, SopE-15-78, SopE-15-50, L2, L6, L9, L10, L11, L12, L15, L16, M3, M4, M5, V12, DD4, DD5, DD6, and DD7. A first polynucleotide encoding a polypeptide comprising a split-intene N fragment selected from the group consisting of CfaN or any functionally equivalent variant thereof, wherein a degron is linked to the intein N fragment via the C-terminus of intein, with or without a linker between the intein N fragment and the degron, and the N-terminus of the split-intene N fragment is directly linked to the N-terminal fragment of the reconstituted protein via a peptide bond; b. CfaC of SEQ ID NO: 28 or CfaC of SEQ ID NO: 29, which is directly linked to the C-terminal fragment of the reconstituted protein via peptide bonds, optionally by a peptide linker, and optionally further linked via peptide bonds to a deglon selected from the group consisting of CL1, Deg1, PEST, DD1, DD2, DD3, M1, M2, SopE, SopE-1-78, SopE-15-78, SopE-15-50, L2, L6, L9, L10, L11, L12, L15, L16, M3, M4, M5, V12, DD4, DD5, DD6 and DD7. A second polynucleotide encoding a polypeptide comprising a split intein C fragment selected from the group consisting of mut or NpuCmut of SEQ ID NO: 36, or any functionally equivalent variant thereof, wherein a degron is linked to the intein C fragment via the N-terminus of intein, with or without a linker between the intein N fragment and the degron, and the C-terminus of the split intein N fragment is directly linked to the N-terminal fragment of the reconstituted protein via a peptide bond; It is characterized by including.
[0011] In the first embodiment or any of its preferred embodiments, the composition is a. A first polynucleotide encoding a polypeptide comprising a split intein N fragment selected from the group consisting of CfaN of SEQ ID NO: 27 or any functionally equivalent variant thereof, wherein the degron is linked to the intein N fragment via a peptide bond, optionally by a peptide linker, and optionally further linked via a peptide bond, to a degron selected from the group consisting of DD1, DD3, PEST, SopE, L2, L9, M4, or V12, wherein the degron is linked to the intein N fragment via the C-terminus of the intein, with or without a linker between the intein N fragment and the degron, and the N-terminus of the split intein N fragment is linked to the N-terminal fragment of the protein to be reconstituted via a peptide bond. b. A second polynucleotide encoding a polypeptide comprising a split intein C fragment selected from the group consisting of CfaC of SEQ ID NO: 28, CfaCmut of SEQ ID NO: 29, or NpuCmut of SEQ ID NO: 36, or any functionally equivalent variant thereof, which is directly linked to the C-terminal fragment of the protein to be reconstituted via a peptide bond, optionally by a peptide linker, and optionally further linked via a peptide bond to a degron selected from the group consisting of DD1, DD3, PEST, SopE, L2, L9, M4, or V12, wherein the degron is linked to the intein C fragment via the N-terminus of the intein, with or without a linker between the intein N fragment and the degron, and the C-terminus of the split intein N fragment is directly linked via a peptide bond to the N-terminal fragment of the protein to be reconstituted. Includes.
[0012] In the first embodiment or another preferred embodiment thereof, the composition is a. A first polynucleotide encoding a polypeptide comprising a split intein N fragment selected from the group consisting of CfaN of SEQ ID NO: 27 or any functionally equivalent variant thereof, wherein the degron is linked to the intein N fragment via a peptide bond, optionally by a peptide linker, and optionally further linked via a peptide bond, to a degron selected from the group consisting of SopE, L2, L9, M4, or V12, wherein the degron is linked to the intein N fragment via the C-terminus of the intein, with or without a linker between the intein N fragment and the degron, and the N-terminus of the split intein N fragment is linked to the N-terminal fragment of the reconstituted protein via a peptide bond. b. A second polynucleotide encoding a polypeptide comprising a split intein C fragment selected from the group consisting of CfaC of SEQ ID NO: 28 or CfaCmut of SEQ ID NO: 29 or any functionally equivalent variant thereof, which is directly linked to the C-terminal fragment of the reconstituted protein via peptide bonds, optionally by a peptide linker, and optionally further linked via peptide bonds to a degron selected from the group consisting of SopE, L2, L9, M4, or V12, wherein the degron is linked to the intein C fragment via the N-terminus of the intein, with or without a linker between the intein N fragment and the degron, and the C-terminus of the split intein N fragment is directly linked via peptide bonds to the N-terminal fragment of the reconstituted protein, Includes.
[0013] In the first embodiment or another preferred embodiment thereof, the composition is a. Directly linked via peptide bonds, optionally by a peptide linker, to the N-terminal fragment of the reconstituted protein, and optionally further linked via peptide bonds, to a deglon selected from the group consisting of CL1, Deg1, PEST, DD1, DD2, DD3, M1, M2, SopE, SopE-1-78, SopE-15-78, SopE-15-50, L2, L6, L9, L10, L11, L12, L15, L16, M3, M4, M5, V12, DD4, DD5, DD6, or DD7, as of SEQ ID NO: 38 A first polynucleotide encoding a polypeptide comprising a split intein N fragment selected from the group consisting of Gp41N or any functionally equivalent variant thereof, wherein a degron is linked to the intein N fragment via the C-terminus of the intein, with or without a linker between the intein N fragment and the degron, and the N-terminus of the split intein N fragment is directly linked to the N-terminal fragment of the reconstituted protein via a peptide bond; b. Linked directly via a peptide bond, optionally via a peptide linker, to the C-terminal fragment of the protein to be reconstituted, and optionally further via a peptide bond, to a degron selected from the group consisting of CL1, Deg1, PEST, DD1, DD2, DD3, M1, M2, SopE, SopE-1-78, SopE-15-78, SopE-15-50, L2, L6, L9, L10, L11, L12, L15, L16, M3, M4, M5, V12, DD4, DD5, DD6 or DD7, a second polynucleotide encoding a polypeptide comprising a split intein C fragment selected from the group consisting of Gp41C of SEQ ID NO: 104 or any functionally equivalent variant thereof, wherein the degron is linked to the intein C fragment via the N-terminus of the intein, with or without a linker between the intein N fragment and the degron, and the C-terminus of the split intein N fragment is linked directly via a peptide bond to the N-terminal fragment of the protein to be reconstituted. comprising.
[0014] In another preferred embodiment of either the first aspect or its preferred embodiments, the first polynucleotide and the second polynucleotide encode the N-terminal fragment and the C-terminal fragment of the ABCA4 protein, respectively, such that when both polynucleotides are translated into their respective protein complexes and combined according to the method of the invention, the N-terminal fragment of the ABCA4 protein is linked to the C-terminal fragment of the ABCA4 protein, thereby generating the entire ABCA4 protein.
[0015] In connection with the embodiments described above, where the first polynucleotide and the second polynucleotide encode the N-terminal fragment and the C-terminal fragment of the ABCA4 protein, respectively, the present invention provides the following alternatives.
[0016] First alternative: a. The first polynucleotide is directly linked via a peptide bond, optionally via a peptide linker, to the N-terminal fragment of the protein to be reconstituted and is optionally further directly linked via a peptide bond to a degron selected from the group consisting of CL1, Deg1, PEST, DD1, DD2, DD3, M1, M2, SopE, SopE-1-78, SopE-15-78, SopE-15-50, L2, L6, L9, L10, L11, L12, L15, L16, M3, M4, M5, V12, DD4, DD5, DD6 and DD7, the degron being linked to the intein N fragment via the C-terminus of the intein with or without a linker between the intein N fragment and the degron, and the N-terminus of the split intein N fragment being directly linked via a peptide bond to the N-terminal fragment of the protein to be reconstituted, encoding a polypeptide comprising a split intein N fragment selected from the group consisting of CfaN of SEQ ID NO: 27 or any functionally equivalent variant thereof. b. The second polynucleotide is directly linked via a peptide bond, optionally via a peptide linker, to the C-terminal fragment of the protein to be reconstituted and is optionally further directly linked via a peptide bond to a degron selected from the group consisting of CL1, Deg1, PEST, DD1, DD2, DD3, M1, M2, SopE, SopE-1-78, SopE-15-78, SopE-15-50, L2, L6, L9, L10, L11, L12, L15, L16, M3, M4, M5, V12, DD4, DD5, DD6 and DD7, the degron being linked to the intein C fragment via the N-terminus of the intein with or without a linker between the intein N fragment and the degron, and the C-terminus of the split intein N fragment being directly linked via a peptide bond to the N-terminal fragment of the protein to be reconstituted, encoding a polypeptide comprising a split intein C fragment selected from the group consisting of CfaCmut of SEQ ID NO: 29 or NpuCmut of SEQ ID NO: 36 or any functionally equivalent variant thereof. If the first polynucleotide codes for positions 1-1149, 1-1139, 1-1176, or 1-1178 of the N-terminal fragment of the ABCA4 protein, and the second polynucleotide codes for positions 1150-2273, 1140-2273, 1177-2273, or 1179-2273 of the C-terminal fragment of the ABCA4 protein, and the first polynucleotide codes for positions 1-1149, then the second poly If the nucleotide codes from position 1150 to 2273 and the first polynucleotide codes from position 1 to 1139, if the second polynucleotide codes from position 1140 to 2273 and the first polynucleotide codes from position 1 to 1176, if the second polynucleotide codes from position 1177 to 2273 and the first polynucleotide codes from position 1 to 1178, then the second polynucleotide codes from position 1179 to 2273.
[0017] Second alternative: a. Encoding a polypeptide comprising a split intein N fragment selected from the group consisting of CfaN of SEQ ID NO: 27 or any functionally equivalent variant thereof, wherein the first polynucleotide is directly linked via a peptide bond, optionally by a peptide linker, to the N-terminal fragment of the protein to be reconstituted, and optionally further linked via a peptide bond, to a degron selected from the group consisting of DD1, DD3, PEST, SopE, L2, L9, M4, or V12, wherein the degron is linked to the intein N fragment via the C-terminus of intein, with or without a linker between the intein N fragment and the degron, and the N-terminus of the split intein N fragment is directly linked via a peptide bond to the N-terminal fragment of the protein to be reconstituted. b. Encoding a polypeptide comprising a split intein C fragment selected from the group consisting of CfaCmut of SEQ ID NO: 29 or NpuCmut of SEQ ID NO: 36, or any functionally equivalent variant thereof, wherein the second polynucleotide is directly linked via a peptide bond, optionally by a peptide linker, to the C-terminal fragment of the protein to be reconstituted, and optionally further linked via a peptide bond, to a degron selected from the group consisting of DD1, DD3, PEST, SopE, L2, L9, M4, or V12, wherein the degron is linked to the intein C fragment via the N-terminus of the intein, with or without a linker between the intein N fragment and the degron, and the C-terminus of the split intein N fragment is directly linked via a peptide bond to the N-terminal fragment of the protein to be reconstituted. If the first polynucleotide codes for positions 1-1149, 1-1139, 1-1176, or 1-1178 of the N-terminal fragment of the ABCA4 protein, and the second polynucleotide codes for positions 1150-2273, 1140-2273, 1177-2273, or 1179-2273 of the C-terminal fragment of the ABCA4 protein, and the first polynucleotide codes for positions 1-1149, then the second poly If the nucleotide codes from position 1150 to 2273 and the first polynucleotide codes from position 1 to 1139, if the second polynucleotide codes from position 1140 to 2273 and the first polynucleotide codes from position 1 to 1176, if the second polynucleotide codes from position 1177 to 2273 and the first polynucleotide codes from position 1 to 1178, then the second polynucleotide codes from position 1179 to 2273.
[0018] A third alternative: a. Encoding a polypeptide comprising a split intein N fragment selected from the group consisting of CfaN of SEQ ID NO: 27 or any functionally equivalent variant thereof, wherein the first polynucleotide is directly linked via a peptide bond, optionally by a peptide linker, to the N-terminal fragment of the protein to be reconstituted, and optionally further linked via a peptide bond, to a degron selected from the group consisting of SopE, L2, L9, M4, or V12, wherein the degron is linked to the intein N fragment via the C-terminus of intein, with or without a linker between the intein N fragment and the degron, and the N-terminus of the split intein N fragment is directly linked via a peptide bond to the N-terminal fragment of the protein to be reconstituted. b. Encoding a polypeptide comprising a split intein C fragment selected from the group consisting of SEQ ID NO: 29 CfaCmut or SEQ ID NO: 36 NpuCmut, or any functionally equivalent variant thereof, wherein the second polynucleotide is directly linked to the C-terminal fragment of the reconstituted protein via a peptide bond, optionally by a peptide linker, and optionally further linked via a peptide bond to a degron selected from the group consisting of SopE, L2, L9, M4, or V12, wherein the degron is linked to the intein C fragment via the N-terminus of the intein, with or without a linker between the intein N fragment and the degron, and the C-terminus of the split intein N fragment is directly linked to the N-terminal fragment of the reconstituted protein via a peptide bond. If the first polynucleotide codes for positions 1-1149, 1-1139, 1-1176, or 1-1178 of the N-terminal fragment of the ABCA4 protein, and the second polynucleotide codes for positions 1150-2273, 1140-2273, 1177-2273, or 1179-2273 of the C-terminal fragment of the ABCA4 protein, and the first polynucleotide codes for positions 1-1149, then the second poly If the nucleotide codes from position 1150 to 2273 and the first polynucleotide codes from position 1 to 1139, if the second polynucleotide codes from position 1140 to 2273 and the first polynucleotide codes from position 1 to 1176, if the second polynucleotide codes from position 1177 to 2273 and the first polynucleotide codes from position 1 to 1178, then the second polynucleotide codes from position 1179 to 2273.
[0019] Fourth alternative: a. Encoding a polypeptide comprising a split intein N fragment selected from the group consisting of CfaN of SEQ ID NO: 27 or any functionally equivalent variant thereof, wherein the first polynucleotide is directly linked via a peptide bond, optionally by a peptide linker, to the N-terminal fragment of the protein to be reconstituted, and optionally further linked via a peptide bond, to a degron selected from the group consisting of DD1, DD3, PEST, SopE, L2, L9, M4, or V12, wherein the degron is linked to the intein N fragment via the C-terminus of intein, with or without a linker between the intein N fragment and the degron, and the N-terminus of the split intein N fragment is directly linked via a peptide bond to the N-terminal fragment of the protein to be reconstituted. b. Encoding a polypeptide comprising a split intein C fragment selected from the group consisting of CfaC of SEQ ID NO: 28, CfaCmut of SEQ ID NO: 29, or NpuCmut of SEQ ID NO: 36, or any functionally equivalent variant thereof, wherein the second polynucleotide is directly linked to the C-terminal fragment of the reconstituted protein via a peptide bond, optionally by a peptide linker, and optionally further linked via a peptide bond to a degron selected from the group consisting of DD1, DD3, PEST, SopE, L2, L9, M4, or V12, wherein the degron is linked to the intein C fragment via the N-terminus of the intein, with or without a linker between the intein N fragment and the degron, and the C-terminus of the split intein N fragment is directly linked to the N-terminal fragment of the reconstituted protein via a peptide bond. The first polynucleotide codes for positions 1 to 1149 of the N-terminal fragment of the ABCA4 protein, and the second polynucleotide codes for positions 1150 to 2273 of the C-terminal fragment of the ABCA4 protein. If the first polynucleotide codes for positions 1 to 1149, then the second polynucleotide codes for positions 1150 to 2273.
[0020] Fifth alternative: a. The first polynucleotide is optionally directly linked to the N-terminal fragment of the reconstituted protein via a peptide linker, and further linked via a peptide linker, selected from the group consisting of CL1, Deg1, PEST, DD1, DD2, DD3, M1, M2, SopE, SopE-1-78, SopE-15-78, SopE-15-50, L2, L6, L9, L10, L11, L12, L15, L16, M3, M4, M5, V12, DD4, DD5, DD6, or DD7. Encoding a polypeptide comprising a split intein N fragment selected from the group consisting of Gp41N of SEQ ID NO: 38 or any functionally equivalent variant thereof, directly linked to degron, wherein degron is linked to the intein N fragment via the C-terminus of intein, with or without a linker between the intein N fragment and degron, and the N-terminus of the split intein N fragment is directly linked via a peptide bond to the N-terminal fragment of the reconstituted protein. b. A second polynucleotide is optionally directly linked to the C-terminal fragment of the reconstituted protein via a peptide bond using a peptide linker, and further linked via a peptide bond to a selected DNA from the group consisting of CL1, Deg1, PEST, DD1, DD2, DD3, M1, M2, SopE, SopE-1-78, SopE-15-78, SopE-15-50, L2, L6, L9, L10, L11, L12, L15, L16, M3, M4, M5, V12, DD4, DD5, DD6, or DD7. The polypeptide comprises a split intein C fragment selected from the group consisting of Gp41C of SEQ ID NO: 104, or any functionally equivalent variant thereof, directly linked to degron, wherein degron is linked to the intein C fragment via the N-terminus of intein, with or without a linker between the intein N fragment and degron, and the C-terminus of the split intein N fragment is directly linked via a peptide bond to the N-terminal fragment of the reconstituted protein. If the first polynucleotide codes for positions 1 to 1095 or 1 to 1185 of the N-terminal fragment of the ABCA4 protein, and the second polynucleotide codes for positions 1096 to 2273 or 1186 to 2273 of the C-terminal fragment of the ABCA4 protein, and the first polynucleotide codes for positions 1 to 1095, then the second polynucleotide codes for positions 1096 to 2273, and if the first polynucleotide codes for positions 1 to 1185, then the second polynucleotide codes for positions 1186 to 2273.
[0021] In the first embodiment or any of its preferred embodiments or alternatives, both polynucleotides are contained within a vector that enables the proliferation or insertion of the polynucleotides in a suitable host cell. Preferably, the vector is an adeno-associated virus (AAV), and more preferably, the vector is an AAV of serotype 1, 2, 3, 4, 5, 6, 7, 8, or 9.
[0022] In the first embodiment or any other further preferred embodiment thereof or any alternative, the compositions described herein are used for the treatment of any of the diseases specified in Table 1.
[0023] Further aspects of the present invention include in vitro or in vivo methods for expressing a target gene in a cell, which are as follows: (i) cells, Contacting the first polynucleotide and the second polynucleotide as defined in any of the preceding embodiments, examples, or alternatives, Preferably, both polynucleotides are contained within the adeno-associated virus (AAV), and at least one of the polynucleotides encodes a split-intene fragment directly linked to degron via a peptide bond; (ii) Expressing the first polynucleotide and the second polynucleotide so that the first fusion protein and the second fusion protein are produced, (iii) The split intein N fragment binds to the split intein C fragment to form an intein intermediate, and the intein intermediate reacts to bring the first protein and the second protein into contact so that the C-terminus of the first target polypeptide is covalently bonded to the N-terminus of the second target polypeptide. This includes mentioning methods. [Brief explanation of the drawing]
[0024] [Figure 1] This is an overview of the strategy for reconstituting large proteins for gene therapy using intein, and intein in combination with degron. Top diagram: (1) The target gene is recombinated at a cleverly selected site, and the resulting 5' end is recombinated with IntN and the 3' end with IntC so that during protein expression, the N-terminal fragment of the protein is expressed fused with IntN to its C-terminus, and the C-terminal fragment of the protein is expressed fused with IntC to its N-terminus (2). Each construct is encapsulated in separate AAVs. These two AAVs are administered to the patient to simultaneously transduce into specific cells (3). The DNA delivered via the AAV is transcribed into RNA by simultaneous transduction, translated into protein, and intein performs a protein trans-splicing reaction at the protein level to reconstitute the desired splicing product (4). Bottom diagram: Intein is combined with degradation signals to prevent the accumulation of starting materials and to eliminate cleaved intein. Deglon is fused to the C-terminus and / or N-terminus of the N-intine. [Figure 2] Top figure: Reaction scheme between the EGFPN-IntN-H6 and IntC-EGFPC-H6 constructs, resulting in the formation of full-length EGFP with an H6 tag and excision of the intein. Bottom figure: Fluorescence microscope images of cells transfected with N-terminal and C-terminal fragments, or cells simultaneously transfected with both fragments using either Cfa or Npu intein. [Figure 3]This figure shows the Western blot analysis of lysates of HEK293 cells transfected with EGFP-intane polynucleotide. Left figure: Western blot analysis of cells transfected with EGFP-intane plasmid. Cells were transfected with either full-length EGFP plasmid, EGFPN-IntN, IntC-EGFPC plasmid, or both simultaneously. Plasmids containing Cfa or Npu intein were tested. Right figure: Quantification of the increase ratio of splicing products relative to Npu. Product yield was determined by densitometry using β-tubulin as a loading control. The graph shows the increase ratio of EGFP products relative to Npu when cells are transfected with full-length EGFP plasmid (EGFP), or when simultaneously transfected with EGFPN-CfaN and CfaC-EGFPC (Cfa), or EGFPN-NpuN and NpuC-EGFPC (Npu). [Figure 4] Transfected HEK293 cells were analyzed by flow cytometry. Left figure: Flow cytometry data of cells transfected with different constructs or combinations of constructs. The black curve corresponds to controls transfected with either the N-terminal or C-terminal fragment of EGFP. The blue line represents four replicas of cells transfected with plasmids encoding full-length EGFP, the green curve corresponds to cells co-transfected with EGFPN-CfaN and CfaC-EGFPC (EGFP split at position 71), and the yellow curve corresponds to cells co-transfected with EGFPN-NpuN and NpuC-EGFPC (EGFP split at position 71). Center figure: Plot of mean fluorescence intensity (MFI) obtained for each sample set, showing that the use of Cfa intein can restore more than 90% of the signal corresponding to full-length EGFP. Right figure: Plot showing the number of positive cells in each sample, showing that the transfection efficiency was similar for all three sets. [Figure 5]This figure compares the splicing yield at position 1150 of ABCA4 using Cfa and Npu. Cells were co-transfected with ABCA4N1150-IntN and IntC-ABCA4C1150, lysed, and analyzed by Western blotting (WB). ABCA4N1150 corresponds to the N-terminal fragment of ABCA4 from residues 1 to 1149, and ABCA4C1150 corresponds to the C-terminal fragment of ABCA4 from residues 1150 to 2273. Constructs containing either Cfa or Npu were tested. The left figure shows a blot using an anti-ABCA4 mAb. The right figure corresponds to densitometry quantification in Western blotting. [Figure 6] This figure compares the splicing yield at position 1140 of ABCA4 using Cfa and Npu. Cells were co-transfected with ABCA4N1140-IntN and IntC-ABCA4C1140, lysed, and analyzed by Western blotting. ABCA4N1140 corresponds to the N-terminal fragment of ABCA4 from residues 1 to 1139, and ABCA4C1140 corresponds to the C-terminal fragment of ABCA4 from residues 1140 to 2273. Constructs containing either Cfa or Npu were tested. The left panel shows the results with anti-FLAG tags, and the center panel shows the results with anti-ABCA4 mAb blotting. The right panel corresponds to densitometry quantification in Western blotting. [Figure 7] This figure shows a comparison of splicing yields at positions 1140, 1150, 1179, and 1188 of ABCA4. Cells were co-transfected with the corresponding ABCA4N-CfaN and CfaC (or CfaCmut)-ABCA4C constructs, lysed, and analyzed in Western blotting (WB). The graph corresponds to the densitometry quantification of reconstituted ABCA4 protein by WB and a comparison of ABCA4 levels obtained by transfection with plasmids encoding the full-length protein. [Figure 8]This figure shows the reaction scheme of PTS including degron. The N-terminal fragment of the target protein is fused to IntN, a linker (His6 tag in this case), and degron. The C-terminus of the protein is fused to IntC, and its N-terminus is linked to degron. The splicing reaction results in the formation of the desired full-length protein, which is not fused to degron, and cleaved inteins, each carrying a degradation signal. [Figure 9] This figure shows combinations of intein and degron for the reconstitution of the reporter protein EGFP. HEK293 cells were transfected with the specified constructs, lysed, and analyzed by Western blotting with an anti-His6 tag mAb. The results show that the use of constructs combining CfaN or C-intine with a degradation signal eliminated all signals originating from the starting material (EGFP-CfaN-6H) as well as from the intein (CfaN-6H). [Figure 10] This figure shows combinations of intein and degron for reconstituting the large protein ABCA4. HEK293 cells were transfected with ABCA4N1150-CfaN and CfaC-ABCA4C1150, with and without degron, lysed, and analyzed by Western blotting using anti-FLAG tag mAbs. Cells were harvested at two different time points (24 hours and 48 hours). Top figure: Western blot analysis. Bottom figure: Densitometry quantification of Western blot. The results show that the use of constructs combining CfaN or C-intine with degradation signals eliminated the signals that would have been produced from the starting materials (ABCA4-CfaN and CfaC-ABCA4). Interestingly, the results also demonstrate that the combination of Cfa-intine with degradation signals resulted in an increased yield of reconstituted full-length ABCA4. [Figure 11]The effect of the intein-degron combination is mediated by proteasomes. Cells were co-transfected with and without degron using the constructs ABCA4N1150-CfaN and CfaC-ABCA4C1150, and incubated for specified times (30 minutes, 3 hours, 6 hours, and 24 hours) in or out of the proteosome inhibitor MG132 after 24 hours. After the specified time, cells were lysed and analyzed by Western blotting with anti-FLAG mAb. The addition of the proteosome inhibitor MG132 increased the intensity of the band corresponding to the fragment, suggesting inhibition of degradation. [Figure 12] This figure shows a comparison of reconstitution yields using different inteins, or combinations of inteins and degron. Left figure: Western blot of PTS-mediated reconstitution of ABCA4 divided at position 1150 using Cfa, Cfa-SopE, or Npu. Cells were simultaneously transfected with ABCA4N1150-CfaN and CfaC-ABCA4C1150(Cfa), ABCA4N1150-CfaN-SopE and SopE-CfaC-ABCA4C1150(Cfa-SopE), or ABCA4N1150-NpuN and NpuC-ABCA4C1150(Npu). After 48 hours, cells were collected, lysed, and analyzed by WB. Cfa constructs were analyzed in double strips, while CfaSopE and Npu were analyzed in triple strips. Right figure: Densitometry quantification of the products obtained with each combination of intein and degron, and the residual starting material. The results show how much higher the product reconstruction becomes when Cfa and degron (SopE) are used together. Interestingly, the data show that reconstruction with Npu is the lowest, and the largest amount of unreacted starting material remains. [Figure 13]This figure shows the rearrangement of ABCA4 at position 1140 using CfaCmut, with and without degradation signals. Left figure: Cells were simultaneously transfected with ABCA4N1140-CfaN and CfaCmut-ABCA4C1140, ABCA4N1140-CfaN-SopE and SopE-CfaC-ABCA4C1140 (Cfa-SopE), or full-length ABCA4 (ABCA). Cells were collected after 48 hours, lysed, and analyzed by WB using anti-FLAG tag mAb. Right figure: Densimetry quantification of the level of ABCA4 rearrangement compared to transfection with full-length ABCA4. [Figure 14] This figure shows the reconstitution of the reporter protein EGFP using a combination of intein and degron with DD1 degron. HEK293 cells were transfected with the specified construct, lysed, and analyzed by Western blotting with an anti-His6 tag mAb. The combination of Cfa intein and DD1 degron eliminates unwanted starting material and cleaved intein. [Figure 15] This figure shows the results of mass spectrometry. Cells were transfected with ABCA4N1150-CfaN-SopE and SopE-CfaC-ABCA4C1150, collected after 48 hours, and lysed. Cell lysates were run on an SDSPAGE gel, and the band around 300 kDa was excised, proteolytic (see Materials and Methods), and analyzed by LC-MS / MS. The results are summarized in the figure. Peptides identified by MS / MS with a false positive rate (FDR) of less than 1% are shown in green. Peptides identified with an FDR of less than 5% are shown in yellow. The red underline indicates the sequence of the cleavage site. [Figure 16] This figure shows a three-piece ligation using orthogonal consensus intent. [Figure 17]This figure shows three-piece ligation using orthogonal consensus intein. The N, M, and C constructs shown in Figure 16 were transfected into HEK293 either individually, in pairs, or all three fragments were transfected together. The cells were lysed and analyzed by Western blotting using antibodies against the 3FT tags present in the C-terminus of the N-terminal, M-terminal, and C-terminal fragments (left figure) or antibodies against the POIN fragment (right figure). [Figure 18] A) This figure shows the analysis of the effect of degron on the expression level of the split-intane N fragment of the present invention. HEK293 cells were transfected with constructs containing degron and constructs without degron, and the expression levels were analyzed by Western blotting. Adding degron to the N-intane fragment resulted in a decrease in the detection level of the expressed protein. B) This figure shows the analysis of the effect of degron on the expression level of the split-intane C fragment of the present invention. HEK293 cells were transfected with constructs containing degron and constructs without degron, and the expression levels were analyzed by Western blotting. Adding degron to the C-intane fragment resulted in a decrease in the detection level of the expressed protein. C) This figure shows the analysis of cells transfected with either the N-terminal or C-terminal fragment, or cells simultaneously transfected with both fragments, by Western blotting. Results are shown for constructs without degron (no degron) or those containing representative degron (SopE, L2, and L9). The figure on the right provides a quantitative comparison of the amount of PTS product obtained with each specified degron compared to that obtained without degron. It can be seen that some degrons provide higher yields than the construct without degron. Interestingly, all degrons are functional, and all of them can reconstitute at least 25% of the product obtained without degron. Some degrons enable reconstitution of more than 50% of the level obtained without degron, and some degrons ultimately result in 100% (or higher) reconstitution. [Figure 19] This figure shows the results of PTS reactions using different deglons, analyzed in Western blotting using different types of gels to detect intein. As can be seen, some deglons reduce the amount of intein cleaved compared to others (e.g., SpoE and PEST), while others, such as L2, L9, M2, M4, or DD3, completely eliminate the intein band. [Figure 20] This figure shows the Western blot analysis of the PTS reaction at separation position 1150, using different deglon combinations and without deglon. In this case, the deglon for the N fragment was fixed (indicated as N in L9 on the left and L2 on the right), while different deglons were used for the C fragment (indicated as C in L2, L9 M4, M2, V12, DD3, and DD1). [Figure 21] This figure shows the WB and quantification of the PTS reaction at sites 1140 and 1179 in the N-terminal and C-terminal intein ABCA4 fragment, with and without the use of degron L9. [Figure 22] This figure shows the analysis of PTS reactions using gp41 inteins by Western blotting. Left figure: Western blotting of PTS reconstructions of ABCA4 separated using Cfa intein at position 1150 and gp41 intein at position 1185. After 48 hours, cells were collected, lysed, and analyzed by WB. Right figure: Densitometry quantification of products obtained with each intein and full-length proteins. [Figure 23] This figure shows in vivo retinal EGFP: (A) Fundus autofluorescence (FAF) of mice injected with AAV8 encoding full-length EGFP, or two AAVs encoding the N-terminal fragment (N,EGFPN-CfaN) and C-terminal fragment (C,CfaC-EGFPC) of the present invention. (B) Immunohistochemical analysis (IHC) of saline or treated retinas. The results show native EGFP fluorescence at the photoreceptor and EGFP reconstitution via protein splicing. All constructs were controlled by the GRK1 promoter, and two different doses of the N+C construct were injected into Balanced Sterile Saline (BSS). [Modes for carrying out the invention]
[0025] This invention relates to methods for using engineered or naturally derived split inteins and combinations thereof with degradation signals (deglons, destabilization domains) to reconstitute large proteins for gene therapy.
[0026] The method of the present invention allows for the reconstitution of large proteins in vitro, ex vivo, and in vivo. When used in vivo, it becomes possible to reconstitute desired large target proteins in any desired tissue, including, but not limited to, the central nervous system (CNS), peripheral nervous system (PNS), muscle, liver, eyeball, pancreas, retina, kidney, inner ear, heart, lung, blood, spleen, and skin.
[0027] As used herein, the term "intene" refers to a naturally occurring or artificially constructed polypeptide sequence capable of catalyzing a protein splicing reaction in which an intein sequence is cleaved from a precursor protein and its neighboring sequences (N-extene and C-extene) are linked by peptide bonds. Inteins are typically 150–550 amino acids in size and may contain a homing endonuclease domain. A list of known inteins is available at http: / / www.inteins.com and https: / / inteins.biocenter.helsinki.fi / index.php.
[0028] As used herein, the term “split intein” means any intein whose N-terminal and C-terminal amino acid sequences are not directly linked via a peptide bond, such that the N-terminal and C-terminal sequences form separate fragments that can be reassociated non-covalently into a functional intein for trans-splicing reactions.
[0029] The term "peptide bond" refers to a covalent chemical bond (-CO-NH-) formed between two molecules when the carboxyl portion of one molecule, called the carboxyl component, reacts with the amino portion of another molecule, called the amino component, releasing a molecule. For example, proteogenic L-amino acids can form peptide bonds in conjunction with the release of water molecules. Therefore, proteins and peptides can be considered as chains of amino acid residues linked by peptide bonds. A peptide bond is also known as an "amide bond" or "amide linkage."
[0030] In this specification, the terms “polypeptide,” “peptide,” or “protein” are used interchangeably to refer to polymers of amino acids.
[0031] The term "amino acid" refers to naturally occurring and synthetic amino acids, as well as amino acid analogs and amino acid mimics that function similarly to naturally occurring amino acids. Furthermore, the term "amino acid" includes both D-amino acids and L-amino acids (stereoisomers).
[0032] The term “natural amino acids” or “naturally derived amino acids” includes 20 naturally derived amino acids, such as hydroxyproline, phosphoserine, and phosphothreonine, which are commonly post-translationally modified in vivo, as well as other less common amino acids, including, but not limited to, 2-aminoadipic acid, hydroxylysine, isodesmosine, norvaline, norleucine, and ornithine.
[0033] As used herein, the terms “unnatural amino acid” or “synthetic amino acid” refer to carboxylic acids or derivatives thereof that are structurally related to natural amino acids, in which the α (alpha) position is substituted with an amine group. Exemplary, non-limiting examples of modified or uncommon amino acids include 2-aminoadipic acid, 3-aminoadipic acid, beta-alanine, 2-aminobutyric acid, 4-aminobutyric acid, 6-aminocaproic acid, 2-aminoheptanoic acid, 2-aminoisobutyric acid, 3-aminoisobutyric acid, 2-aminopimelic acid, 2,4-diaminobutyric acid, desmosine, 2,2'-diaminopimelic acid, 2,3-diaminopropionic acid, N-ethylglycine, N-ethylasparagine, and hydroxypropyl Examples include roxylysine, ariohydroxylysine, 3-hydroxyproline, 4-hydroxyproline, isodesmosine, alloisoleucine, N-methylglycine, N-methylisoleucine, 6-N-methyllysine, N-methylvaline, norvaline, norleucine, ornithine, p-acetylphenylalanine, p-halophenylalanine, p-propargyloxyphenylalanine, p-azidophenylalanine, and p-benzoylphenylalanine. This group also includes D-isomers of "natural amino acids."
[0034] As used herein, the terms “split-intane N fragment,” “N-terminal split-intane,” “N-terminal intein fragment,” or “N-terminal intein sequence” (abbreviated as “IntN”) refer to any intein sequence that includes a functional N-terminal amino acid sequence for trans-splicing reactions, i.e., it can associate with a functional split-intane C fragment to form a complete intein that can cleave itself from a host protein, or it can catalyze ligation via peptide bonds of an extein or neighboring sequence, or it catalyzes “N-terminal cleavage” when associated with a split-intane C fragment, i.e., it nucleophilically attacks the peptide bond between the extein and the N-terminus of the split-intane N fragment, resulting in the breakdown of the peptide bond. Thus, IntN also includes sequences that are spliced when trans-splicing occurs. IntN may include sequences that are modified from the N-terminal portion of naturally occurring intein sequences. For example, IntN may contain additional amino acid residues and / or mutated residues, provided that such additional residues and / or mutated residues do not render IntN non-functional in trans-splicing. Preferably, the inclusion of additional residues and / or mutated residues improves or enhances the trans-splicing activity of IntN.
[0035] As used herein, the term “degron” means a naturally occurring or artificially constructed polypeptide sequence that, when recombinantly fused to another polypeptide, promotes its protein degradation via the proteasomal degradation pathway or any other cellular degradation mechanism.
[0036] As used without distinction herein, the terms “split-intane C fragment,” “C-terminal split-intane,” “C-terminal intein fragment,” and “C-terminal intein sequence” (abbreviated as “IntC”) refer to any intein sequence containing a functional C-terminal amino acid sequence for trans-splicing reactions, that is, a sequence that, upon association, can associate with a functional split-intane N fragment to form a complete intein capable of cleaving itself from a host protein, or can catalyze ligation via peptide bonds of an extein or neighboring sequence, or, upon association with a split N-intane, catalyzes “C-terminal cleavage,” i.e., nucleophilically attacks the peptide bond between the extein and the C-terminus of the split-intane C fragment, resulting in the breakdown of the peptide bond. Thus, IntC also includes sequences that are spliced when trans-splicing occurs. IntC may include sequences that are modified versions of the C-terminal portion of naturally occurring intein sequences. For example, IntC may contain additional amino acid residues and / or mutated residues, provided that such additional residues and / or mutated residues do not render IntC non-functional in trans-splicing. Preferably, the inclusion of additional residues and / or mutated residues improves or enhances the trans-splicing activity of IntC.
[0037] In this specification, in the context of the present invention, “N-terminal fragment of the reconstituted protein” refers to the N-terminal fragment of a protein, more specifically, a protein fragment greater than 25 kDa, greater than 50 kDa, or greater than 100 kDa. As used herein, the term “N-terminal fragment of a protein” therefore refers to a fragment of variable length that includes the N-terminus (mature or immature) of a protein. In certain embodiments, the N-terminal fragment is a fragment comprising less than 100%, less than 90%, less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, or less than 5% of the total length of the protein.
[0038] In this specification, in the context of the present invention, “C-terminal fragment of the protein to be reconstituted” refers to the C-terminal fragment of a protein, more specifically, a protein fragment greater than 25 kDa, greater than 50 kDa, or greater than 100 kDa. Therefore, as used herein, the term “C-terminal fragment of a protein” refers to a fragment of variable length that includes the C-terminus of a protein. In certain embodiments, the C-terminal fragment is a fragment comprising less than 100%, less than 90%, less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, or less than 5% of the total length of the protein.
[0039] As used herein, the term "polynucleotide" refers to a polymer composed of multiple nucleotide units (deoxyribonucleotides or ribonucleotides, or their related structural variants or synthetic analogs) linked via phosphodiester bonds (or related structural variants on their synthetic analogs). The term polynucleotide includes double-stranded or single-stranded genomes and cDNA, RNA, any synthetic and genetically engineered polynucleotides, and both sense and antisense polynucleotides (however, only sense strands are disclosed in this invention). This includes single-stranded and double-stranded molecules, i.e., DNA-DNA, DNA-RNA, and RNA-RNA hybrids.
[0040] Split-Intein N Fragment This invention refers to a split-intene N-terminal fragment directly linked to the N-terminal fragment of the protein to be reconstituted via a peptide bond. The construct will have a general structure from the N-terminus to the C-terminus: (N-terminal fragment of the protein to be reconstituted)-inteinN
[0041] The protein to be reconstituted could be any large gene whose reconstitution could potentially yield a positive therapeutic effect. Non-limiting examples of proteins that can be reconstituted in this invention include the ABCA4 protein and proteins encoded by the genes listed in Table 1 below. The purpose of reconstituting these proteins is to treat diseases associated with mutations in their coding genes by providing the correct version of such genes. Furthermore, the protein to be reconstituted could be a protein or enzyme for treating a disease, but not necessarily a replacement for a mutated one. For example, the protein to be reconstituted could be a gene, a base, or a CRISPR / Cas9 system for prime editing.
[0042] [Table 1] TIFF0007867703000002.tif254170TIFF0007867703000003.tif253170TIFF00078677030 00004.tif253170TIFF0007867703000005.tif254170TIFF0007867703000006.tif138170
[0043] Accordingly, a first aspect of the present invention refers to a split-intane N fragment directly linked to the N-terminal fragment of a protein to be reconstituted via a peptide bond, wherein, as reflected above, IntN is directly or via a linker linked to the C-terminus of the N-terminal fragment (hereinafter referred to as "the split-intane N fragment of the present invention").
[0044] A preferred embodiment of the first aspect of the present invention refers to a split-intane N fragment directly linked to a degron via a peptide bond (hereinafter referred to as "the split-intane N fragment degron of the present invention"), wherein the degron is linked to the intein N fragment via the C-terminus of intein, with or without a linker between the intein N fragment and the degron, and the N-terminus of the split-intane N fragment is directly linked via a peptide bond to the N-terminal fragment of the reconstituted protein.
[0045] Therefore, the structure of the split-intei N-fragment degron construct of the present invention, from the N-terminus to the C-terminus, is as follows: (N-terminal portion of the protein to be reconstituted)-(intene-N fragment)-(degron).
[0046] If a linker is introduced between IntN and degron, the structure of the N-terminus to C-terminus of the split-intene N-fragment degron of the present invention is as follows: (N-terminal portion of the protein to be reconstituted)-(intene-N fragment)-(linker)-(degron).
[0047] Preferably, the intein N in either the split-intine N fragment of the present invention or the split-intine N fragment degron of the present invention can be selected from any of those listed in Table 2 below.
[0048] [Table 2]
[0049] Here, preferably, intein N is any variant thereof such as CfaN intein of SEQ ID NO: 27 or ConN of SEQ ID NO: 39; or CatN intein of SEQ ID NO: 30 or any variant thereof; or NpuN intein of SEQ ID NO: 32 or any variant thereof; or Gp41N intein of SEQ ID NO: 38 or any variant thereof.
[0050] As used herein, the term “mutant” refers to a polypeptide molecule that is substantially similar to a particular polypeptide sequence. In this invention, the inventors refer to mutants of the N-intane or C-intane of Cfa, Npu, Cat, or Gp41 as polypeptides substantially similar to any of those polypeptide sequences; for example, Con-intane is understood herein as a mutant of Cfa. Thus, a mutant may be similar in structure and biological activity to the polypeptide from which it is derived. Thus, a mutant may also refer to a mutant of a polypeptide sequence. The term “mutant” refers to a polypeptide molecule in which one or more amino acids in its sequence are added, deleted, substituted, or otherwise chemically modified compared to the polypeptide molecule from which it is derived. A mutant may retain substantially the same properties as the polypeptide molecule from which it is derived, or it may lack the biological activity of the sequence described in the claims. In certain embodiments, an intein mutant includes a mutant in which a non-catalyzed Cys residue, i.e., the first residue of N-intane, is mutated to serine or alanine.
[0051] Preferably, the term “mutant” as commonly understood in this invention refers to a mutant with increased promiscuity, where promiscuity is understood as the ability of the intein to undergo protein trans-splicing (PTS) reactions without depending on the identity of the residue immediately adjacent to the splitting site. More specifically, the promiscuity of the intein refers to its ability to undergo PTS reactions without depending on the identity of the amino acid immediately adjacent to the splitting site, i.e., the site in which the intein is inserted into the target protein. A schematic diagram of the splitting site is shown below: -3-2-1 +1+2+3 XXX-InteinN InteinCXXX
[0052] Typically, inteins strongly prefer certain amino acids at their positions. For example, Cat intein (SEQ ID NO: 30) prefers Cys at position +1 and Glu at position -1 (Stevens, AJ, Sekar, G., Gramespacher, JA, Cowburn, D., & Muir, TW (2018)). An Atypical Mechanism of Split Intein Molecular Recognition and Folding. Journal of the American Chemical Society, 140(37), 11791-11799. http: / / doi.org / 10.1021 / jacs.8b07334). More ambiguous mutants can perform protein trans-splicing reactions in good yield even in the absence of such preferred residues.
[0053] In certain embodiments, intein mutants include mutants in which the catalytic residue and second-shell accelerator residues of intein are maintained, while only non-catalytic residues or residues outside the second shell are mutated. The second-shell accelerator residues are residues adjacent to the active site of intein and play a crucial role in regulating the splicing activity of intein. For example, the catalytic and accelerator residues of Cfa intein and inteins of the DnaE family are well known and described in the art (Stevens, AJ, Brown, ZZ, Shah, NH, Sekar, G., Cowburn, D., & Muir, TW (2016). Journal of the American Chemical Society. http: / / doi.org / 10.1021 / jacs.5b13528). Similarly, catalytic residues of several other intein families, including DnaE, GyrA, GyrB, DnaB, TerL, gp41, and IMPDH, have been studied and characterized (Shah, NH, & Muir, TW (2014). Inteins: Nature's Gift to Protein Chemists. Chemical Science (Royal Society of Chemistry: 2010), 5(1), 446-461. http: / / doi.org / 10.1039 / C3SC52951G). A methodology for identifying the second shell promoter residue has also been described (Stevens et al. 2016 Journal of the American Chemical Society. http: / / doi.org / 10.1021 / jacs.5b13528), which can be used to create intein variants (both N-intines and C-intines) with properties suitable for use in the present invention.
[0054] In certain embodiments, the intein variants include mutants in which the catalytic residue of intein is maintained, and the variants retain the key functional characteristics of CfaN, NpuN, CatN, or gp41N intein, including a fast splicing rate (half-life less than 5 minutes) and high activity in the presence of chaotropic agents or at high temperatures. In more detailed embodiments, the variants of CfaN, CatN, NpuN, or gp41N intein are functionally equivalent variants of any of these sequences.
[0055] As used herein, the term “functionally equivalent variant” is understood to mean all such proteins obtained from a sequence by modification, insertion and / or deletion, or by one or more amino acids, provided that the function is substantially maintained or improved, and in particular, in the case of functionally equivalent variants of the split-intane N fragment, it refers to the maintenance of its activity. As used herein, the term “activity” refers to the ability of the split-intane N fragment to perform protein trans-splicing reactions by binding to the split-intane C fragment.
[0056] Examples of functionally equivalent variants of CfaN, NpuN, CatN, or gp41N intein are shown below.
[0057] Functionally equivalent variants of CfaN can be selected from one of the following lists: A CfaN sequence (SEQ ID NO: 27) in which Cys28 and / or Cys59 are mutated to Ser, Thr, or Ala. The CfaN sequence (SEQ ID NO: 27) contains additional amino acids at the N-terminus of Met1, such as a linker, degron, or detection tag. A CfaN sequence (SEQ ID NO: 27) in which one of the residues has been mutated to a residue with similar physicochemical properties, except for the following residues that must be maintained: Cys1, Lys70, His72, Met75, and Met81. The CfaN sequence (SEQ ID NO: 27) is a sequence in which any of the following residues must be maintained: Cys1, Asp5, Phe15, Glu24, Thr32, Lys35, Phe38, Val39, Ile44, Asn49, Ile65, Thr69, Lys70, His72, Met75, Thr77, Met81, Gly91, Lys95, Gln96, and Gly99, all of which have been mutated to residues with similar physicochemical properties. The following residues: Cys1, Lys70, His72, Met75, Met81 are maintained, and one or a combination of the following mutations are introduced: Asp5Glu, Phe15Leu, Glu24Lys, Thr32Ser, Lys35Asn, Phe38Asn, Val39Ile, Ile44Val, Asn49Asp, Ile65Leu, Thr77Val, Gly91Glu, Lys95Met, Gln96Arg, Gly99Asn, in the CfaN sequence (SEQ ID NO: 27). The following residues: Cys1, Lys70, His72, Met75, Met81 are maintained, and one to six of the following mutations (1 and 6 are within the range) are introduced: Asp5Glu, Phe15Leu, Glu24Lys, Thr32Ser, Lys35Asn, Phe38Asn, Val39Ile, Ile44Val, Asn49Asp, Ile65Leu, Thr77Val, Gly91Glu, Lys95Met, Gln96Arg, Gly99Asn, CfaN sequence (SEQ ID NO: 27). The following residues are maintained: Cys1, Asp5, Glu24, Phe38, Val39, Ile44, Lys70, His72, Met75, Met81, Gly91, Lys95, Gln96, Gly99, and one or a combination of the following mutations, Phe15Ala, Thr32Ser, Lys35Glu, Asn49Asp, Ile65Thr, Thr77Glu, is introduced into the CfaN sequence (SEQ ID NO: 27).
[0058] Functionally equivalent variants of NpuN can be selected from one of the following lists: An NpuN sequence (SEQ ID NO: 32) in which Cys28 and / or Cys59 are mutated to Ser, Thr, or Ala. The NpuN sequence (SEQ ID NO: 32) allows any residue of NpuN to be mutated to a residue with similar physicochemical properties, except for the following residues that must be maintained: Cys1, Thr69, Lys70, His72, Met75, and Met81. The NpuN sequence (SEQ ID NO: 32) allows any residue in NpuN to be mutated to the corresponding amino acid in the CfaN sequence.
[0059] Functionally equivalent variants of Gp41N can be selected from any of the following lists: A Gp41N sequence (SEQ ID NO: 38) in which Cys59 and / or Cys83 are mutated to Ser, Thr, or Ala. The Gp41N sequence (SEQ ID NO: 38) can be mutated to any residue of Gp41N with similar physicochemical properties, except for the following residues that must be maintained: Cys1 and His63, and residue 60 is Ser or Thr.
[0060] Functionally equivalent variants of CatN can be selected from one of the following lists: A CatN sequence (SEQ ID NO: 30) that allows any residue in CatN to be mutated to a residue with similar physicochemical properties, except for the Cys1 residue which must be maintained.
[0061] In certain embodiments, intein N comprises or consists of a variant of the amino acid sequence of SEQ ID NO: 27, SEQ ID NO: 30, SEQ ID NO: 32, or SEQ ID NO: 38, having at least 90% sequence identity with SEQ ID NO: 27, SEQ ID NO: 30, SEQ ID NO: 32, or SEQ ID NO: 38 across the entire sequence. In certain embodiments, a variant of intein N of SEQ ID NO: 27, SEQ ID NO: 30, SEQ ID NO: 32, or SEQ ID NO: 38 has at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 27, SEQ ID NO: 30, SEQ ID NO: 32, or SEQ ID NO: 38 across the entire sequence.
[0062] In relation to two or more amino acid or nucleotide sequences, the terms “identity,” “identity,” “identity percentage,” or “sequence identity” refer to two or more sequences or subsequences that are identical or have a specific percentage of the same amino acid residues when compared and aligned (with gaps introduced as necessary) for maximum correspondence, without considering any conserved amino acid substitutions as part of sequence identity. The identity percentage can be measured by sequence comparison software or algorithms, or by visual inspection. Various algorithms and software that can be used to obtain amino acid sequence alignment are known in the art. One such non-restrictive example of a sequence alignment algorithm is the algorithm described in Karlin et al., 1990, Proc. Natl. Acad. Sci., 87:2264-8, modified in Karlin et al., 1993, Proc. Natl. Acad. Sci., 90:5873-7, and is incorporated into the N BLAST and X BLAST programs (Altschul et al., 1991, Nucleic Acids Res., 25:3389-402). In certain embodiments, Gapped BLAST can be used, as described in Altschul et al., 1997, Nucleic Acids Res. 25:3389-402. BLAST-2, WU-BLAST-2 (Altschul et al., 1996, Methods in Enzymology, 266:460-80), ALIGN, ALIGN-2 (Genentech, South San Francisco, California), or Megalign (DNASTAR) are additional publicly available software programs that can be used to align sequences.In certain alternative embodiments, the percentage of identity between two amino acid sequences can be determined using the GAP program of the GCG software package incorporating the algorithm of Needleman and Wunsch (J. Mol. Biol. 48:444-53 (1970)) (e.g., using either the Blossum 62 matrix or the PAM250 matrix, with gap weights of 16, 14, 12, 10, 8, 6, or 4 and length weights of 1, 2, 3, 4, or 5). Alternatively, in certain embodiments, the percentage of identity between amino acid sequences is determined using the algorithm of Myers and Miller (CABIOS, 4:1 1-7 (1989)). For example, the percentage of identity can be determined using the ALIGN program (version 2.0) with a residue table for PAM120, a gap length penalty of 12, and a gap penalty of 4. Appropriate parameters for maximum alignment by specific alignment software can be determined by those skilled in the art. In certain embodiments, the default parameters of the alignment software are used. In a particular embodiment, the identity percentage "X" of the first amino acid sequence to the second amino acid sequence is calculated as 100 × (Y / Z), where Y is the number of amino acid residues scored as identical in the alignment of the first and second sequences (aligned visually or by a specific sequence alignment program), and Z is the total number of residues in the second sequence. If the second sequence is longer than the first sequence, a global alignment considering the entirety of both sequences is used, so all letters and nulls in each sequence must be aligned. In this case, the same formula as above can be used, but the length of the overlapping region of the first and second sequences is used as the Z value, and this region is assumed to be substantially the same length as the length of the first sequence.
[0063] As a non-restrictive example, whether any particular polypeptide has a certain percentage of sequence identity with respect to a reference sequence (e.g., at least 80% identical, at least 85% identical, at least 90% identical, and in some embodiments, at least 95%, 96%, 97%, 98%, or 99% identical) can, in certain embodiments, be determined using the Bestfit program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, 575 Science Drive, Madison, Wl 5371 1). Bestfit uses the local homology algorithm from Smith and Waterman, Advances in Applied Mathematics 2:482-9 (1981) to find the optimal segment of homology between two sequences. When using Bestfit or any other sequence alignment program to determine whether a particular sequence is 95% identical to, for example, a reference sequence according to the present invention, the percentage of identity is calculated over the entire length of the reference amino acid sequence, and the parameters are set to allow a homology gap of up to 5% of the total number of nucleotides in the reference sequence.
[0064] The term "linker" between the split intein N fragment and degron refers to the amino acid sequence that connects the C-terminus of the intein N fragment to the N-terminus of the degron sequence. The linker is preferably a polypeptide of 1 to 100 amino acids, or 1 to 5 amino acids, or 1 to 10 amino acids, or 1 to 50 amino acids, or 1 to 25 amino acids. The linker may be a glycy-rich peptide or a glycy-ser-rich peptide. Examples of glycy-ser-rich peptides include the sequence GGS, or polymers of this linker with the general formula (GGS). n Examples include those with (n can be 1 to 10). The general formula for linker is ((G) n S) yHere, n is 1 to 5 and y is 1 to 10. Furthermore, epitope tags can be used as linkers, for example, hexahistidine (His6, sequence: GHHHHHHG (sequence number 84) tag) or triple flag tag (3FT, sequence: DYKDHDGDYKDHDIDYKDDDDK (sequence number 103)).
[0065] In another specific embodiment, the degron directly linked to the split-intei N fragment via a peptide bond is selected from any of those listed in Table 3 below.
[0066] [Table 3]
[0067] Other non-limiting examples of degron include polypeptides associated with DHFR (dihydrofolate reductase), FKBP (FK506-binding protein), FRB (FKBP-rapamycin-binding protein), and PDE5 (phosphodiesterase type 5). Degradation signals in the endoplasmic reticulum-associated degradation (ERAD) pathway are also included.
[0068] In yet another preferred embodiment, degron is a polypeptide with fewer than 75 amino acids, which, upon fusion with the intein N fragment of the present invention, induces its degradation.
[0069] In yet another preferred embodiment, degron is a polypeptide that, when fused with the N-terminal or C-terminal fragment of the split-intene of the present invention, results in a reduction of more than 10%, 20%, 30%, 50%, 75%, or 90% in the expression level of the fragment. Importantly, the incorporation of degron into the fragment of the present invention does not significantly affect the yield of the reconstituted protein obtained from protein trans-splicing between the N-terminus and C-terminus of the split-intene of the present invention and the degron fragment. For the purposes of this preferred embodiment, the expression levels are measured as shown in Figure 18, i.e., the expression level of the fragment with degron is compared to the expression level of the fragment without degron under conditions similar to those shown in Figure 18, and the former must be at least 10%, at least 20%, at least 30%, at least 50%, at least 75%, or at least 90% lower than the latter.
[0070] In yet another preferred embodiment, preferred combinations of N-intane and degron of the split-intane N-fragment degron of the present invention are selected from the group consisting of: The following: CfaN or functionally equivalent variants thereof combined with any of the degrons CL1, Deg1, PEST, DD1, DD2, DD3, M1, M2, SopE, SopE-1-78, SopE-15-78, SopE-15-50, L2, L6, L9, L10, L11, L12, L15, L16, M3, M4, M5, V12, DD4, DD5, DD6, and DD7. Preferably: CfaN or functionally equivalent variants thereof combined with any of the degrons CL1, Deg1, DD1, DD2, DD3, SopE, L2, L9, M4, V12, or DD4, or The following: Gp41N or a functionally equivalent variant thereof in combination with any of the degrons CL1, Deg1, PEST, DD1, DD2, DD3, M1, M2, SopE, SopE-1-78, SopE-15-78, SopE-15-50, L2, L6, L9, L10, L11, L12, L15, L16, M3, M4, M5, V12, DD4, DD5, DD6, or DD7. Preferably: Gp41N or a functionally equivalent variant thereof in combination with any of the degrons CL1, Deg1, DD1, DD2, DD3, SopE, L2, L9, M4, V12, or DD4.
[0071] Figure 18A shows the effect of several representative degrons when fused to split N-intane fragments, particularly CfaN-intane. Addition of degrons (SopE, L2, L9, M4, V12) significantly reduces the detectable expression level of N-intane-containing protein fusions fused with degron. Interestingly, as described in the examples and in Figures 18, 20, and 21, the inclusion of several preferred degrons (SopE, L2, L9, M4, or V12) does not adversely affect the yield of splicing product formation. This result was not predicted from the state of the art, as a decrease in the amount of starting material in the protein trans-splicing reaction, as achieved by degron addition, would be expected to reduce the protein splicing yield.
[0072] In yet another preferred embodiment, a preferred combination of split-intane N fragment and degron is selected from the group consisting of combinations of CfaN intein with the following degrons DD1, DD2, DD3, SopE, L2, L9, M4, V12, DD4, DD5, DD6, or DD7.
[0073] Thus, the inventors observed that when CfaN intein was combined with any of the above degrons, particularly DD1, DD2, DD3, SopE, L2, L9, M4, V12, DD4, DD5, DD6, or DD7, a decrease in the expression level of the N-terminal intein fragment was detected (Figures 9 and 18A). Importantly, this did not affect the yield of the protein trans-splicing reaction (Figures 9, 18C, 20, or 21). Typically, a decrease in the yield, expression, or stability of one of the two intein fragments involved in a protein trans-splicing reaction leads to an overall decrease in the amount of splicing product formed, so this result was unexpected from the prior art.
[0074] The inventors have observed this effect not only with membrane proteins such as ABCA4 (Figures 18, 20, and 21) but also with cytoplasmic soluble proteins such as EGFP, demonstrating the universality of the effect. In Figure 9, it can be observed that when degron is added to EGFPN-CfaN, the band corresponding to that protein completely disappears. However, when CfaC-EGFPC fragments (with or without degron) are added, the full-length EGFP product is formed in the same yield as when degron was not used. This example shows that adding one or two degron molecules (one for each fragment) does not adversely affect the yield of the target protein.
[0075] A complex containing the N-fragment of split-intane In another embodiment, the present invention relates to any of the following complexes (which may be more than one), namely the split-intine N fragment of the present invention or the split-intine N fragment degron of the present invention (hereinafter referred to as the first complex (which may be more than one) of the present invention), each of which complexes is as follows: (i) The N-terminal fragment of the protein to be reconstituted, (ii) As defined in the above section, a split-intine N fragment, or a split-intine N fragment directly linked to degron via a peptide bond or optionally by a linker (split-intine N fragment degron of the present invention), Includes, The complex (or there may be multiple complexes) optionally includes a linker between (i) and (ii), The N-terminal fragment of the target protein is linked to the split-intane N-terminal fragment by an amide bond, or If the complex(s) contains a linker, the N-terminal fragment of the protein of interest is bound to the linker by an amide bond, and / or the linker is bound to the N-terminus of the split-intane N-fragment by an amide bond.
[0076] Non-limiting examples of target proteins useful in any of the above complexes are shown in Table 1 above. Further target proteins can be selected from antibodies, antibody fragments containing Fc domains or scFvs, nanobodies, bispecific antibodies, proteins, preferably any protein greater than 25 kDa, 50 kDa, or 100 kDa.
[0077] As already shown, the target protein and the split-intane N-fragment may optionally be linked by a linker, so the linker will be located between the target protein and the N-intane. The properties of the linker depend on the properties of the target protein. In certain embodiments, the linker is a peptide. In certain embodiments, the linker is a peptide having a length of 1, 2, 3, 4, 5, 10, 20, 50, 100 or more amino acid residues, and specifically, it may be 1 to 3 amino acid residues. Preferably, by peptide bond, the N-terminus of the linker is linked to the C-terminus of the target protein, and the C-terminus of the linker is linked to the N-terminus of the N-intane.
[0078] In another specific embodiment, the complex does not contain a linker between the N-terminal fragment of the protein of interest and the split-intane N-fragment. In this specific embodiment, the protein of interest is linked to the N-terminus of the split-intane N-fragment by an amide bond.
[0079] In another specific embodiment, if the complex contains a linker, the complex is a fusion protein. The term “fusion protein” is well known in the art and refers to a single polypeptide chain that is artificially designed and contains two or more sequences from different natural and / or artificial origins. A fusion protein is, by definition, not found in nature on its own.
[0080] Split Intein C Fragment In another embodiment, the present invention relates to a split-intane C fragment (hereinafter referred to as "the split-intane C fragment of the present invention") that is directly linked to the C-terminal fragment of a protein to be reconstituted via a peptide bond.
[0081] More preferably, the present invention refers to a split-intane C fragment directly linked to a degron via a peptide bond (hereinafter referred to as "the split-intane C fragment degron of the present invention"), wherein the degron is linked to the intein C fragment via the N-terminus of intein, with or without a linker between the intein C fragment and the degron, and the C-terminus of the split-intane C fragment is directly linked via a peptide bond to the N-terminus of the C-terminal fragment of the reconstituted protein.
[0082] Therefore, the structure of the split-intei C-fragment degron construct of the present invention, from the N-terminus to the C-terminus, is as follows: (Degron)-(Intein-C fragment)-(C-terminal portion of the reconstituted protein).
[0083] If a linker is present between degron and intein C, the N-terminus to C-terminus structure of the split intein C fragment degron construct of the present invention is as follows: (Degron)-(linker)-(intene-C fragment)-(C-terminal portion of the reconstituted protein).
[0084] Preferably, the intein C in either the split intein C fragment of the present invention or the split intein C fragment degron of the present invention is selected from any of those listed in Table 4 below.
[0085] [Table 4]
[0086] Here, preferably, intein C is any variant thereof such as CfaC intein of SEQ ID NO: 28, CfaCmut (SEQ ID NO: 29), or ConC (SEQ ID NO: 105); or CatC intein of SEQ ID NO: 31 or any variant thereof, or NpuC intein (SEQ ID NO: 33), or NpuCmut (SEQ ID NO: 36), or any variant thereof such as GP41C intein of SEQ ID NO: 104; or any variant thereof with increased ambiguity as defined above.
[0087] The term "mutant" is defined as it is for the N fragment of split intein. Examples of intein C mutants include SEQ ID NOs. 28, 29, 31, 33, 36, 104, and 105, which lack the N-terminal methionine residue.
[0088] In certain embodiments, variants of CfaC, NpuC, or ConC with increased ambiguity contain the amino acids Gly, Glu, and Pro at positions 21, 22, and 23, respectively, as described in Stevens et al. Proc. Natl. Acad. Sci. 2017.
[0089] In a particular embodiment, examples of functionally equivalent variants of CfaC, NpuC, CatC, or gp41c intein are shown below:
[0090] Functionally equivalent variants of CfaC can be selected from one of the following lists: The CfaC (sequence number 28) sequence with Met1 removed. The CfaC (SEQ ID NO: 28) sequence includes additional amino acids in the N-terminal Met1, such as a linker, degron, or detection tag. A CfaC (SEQ ID NO: 28) sequence in which any residue of CfaC is mutated to a residue with similar physicochemical properties, except for the following residues that should be maintained: Asp17, His24, and Asn36. A CfaC (SEQ ID NO: 28) sequence in which any residue of CfaC is mutated to a residue with similar physicochemical properties, except for the following residues that should be maintained: Asp17, Asp23, His24, Ser35, and Asn36. A CfaC (SEQ ID NO: 28) sequence in which any residue of CfaC is mutated to a residue with similar physicochemical properties, except for the following residues that should be maintained: Asp17, Asp23, His24, Ser35, and Asn36. The CfaC (SEQ ID NO: 28) sequence is characterized by the following residues: Asp17, Asp23, His24, Ser35, Asn36 being maintained, and one of the following amino acids, Val2, Ile5, Ser6, Ser9, Lys22, Leu27, Leu32, or Val33, being mutated. The following residues are mutated in the CfaC (SEQ ID NO: 28) sequence: Glu21Gly, Lys22Glu, and Asp23Pro. This variant corresponds to CfaCmut intein (SEQ ID NO: 29). A CfaC (SEQ ID NO: 28) sequence in which any residue is mutated to a residue with similar physicochemical properties, except for the following residues that should be in the specified positions: Asp17, Gly21, Glu22, Pro23, His24, and Asn36. The CfaC (SEQ ID NO: 28) sequence is one in which the following residues should be in the specified positions: Asp17, Gly21, Glu22, Pro23, His24, Ser35, Asn36, and one of the following amino acids Val2, Ile5, Ser9, Leu27, Leu32, or Val33 is mutated. A CfaC (SEQ ID NO: 28) or CfaCmut (SEQ ID NO: 29) sequence containing one to five of the following mutations (1 and 5 are included in the range): Val2Ile, Ile5Ala, Ser6Thr, Ser9Tyr, Thr12Lys, Leu27Ala, Leu32Phe, Val33Ile.
[0091] Functionally equivalent variants of NpuC can be selected from one of the following lists: The NpuC (sequence number 33) sequence with Met1 removed. The NpuC (SEQ ID NO: 33) sequence contains additional amino acids at the N-terminus of Met1, such as a linker, degron, or detection tag. The NpuC (SEQ ID NO: 33) sequence allows any residue of NpuC to mutate into a residue with similar physicochemical properties, except for the following residues that should be maintained: Asp17, His24, and Asn36. The NpuC (SEQ ID NO: 33) sequence allows any residue in NpuC to mutate into a residue with similar physicochemical properties, except for the following residues that need to be maintained: Asp17, His24, Ser35, and Asn36. The following residues: Asp17, His24, Ser35, Asn36 are maintained, and one of the following amino acids, Val2, Ile5, Ser9, Lys22, Leu32, or Val33, is mutated in the NpuC (SEQ ID NO: 33) sequence. The following residues: Asp17, His24, Ser35, Asn36 are maintained, and one of the following mutations is incorporated: Ile2Val, Ala5Ile, Tyr9Ser, Ala27Leu, Phe32Leu, or Ile33Val, in the NpuC (Sequence ID 33) sequence. The NpuC (SEQ ID NO: 33) sequence has mutations in the following residues, Glu21Gly, Arg22Glu, and Asp23Pro, to increase the ambiguity of Npu.
[0092] Functionally equivalent variants of Gp41C can be selected from one of the following lists: The Gp41C (sequence number 104) sequence with Met1 removed. The Gp41C (SEQ ID NO: 104) sequence contains additional amino acids at the N-terminus of Met1, such as a linker, degron, or detection tag. The sequence of Gp41C (SEQ ID NO: 104) is such that any residue of Gp41C can be mutated to a residue with similar physicochemical properties, except for the following residues that need to be maintained: His26, His36, and Asn37.
[0093] In certain embodiments, Intein C comprises or consists of a variant of the amino acid sequence of SEQ ID NO: 28, SEQ ID NO: 31, SEQ ID NO: 33, or SEQ ID NO: 104, having at least 90% sequence identity with SEQ ID NO: 28, SEQ ID NO: 31, SEQ ID NO: 33, or SEQ ID NO: 104 across the entire sequence. In certain embodiments, a variant of Intein C of SEQ ID NO: 28, SEQ ID NO: 31, SEQ ID NO: 33, or SEQ ID NO: 104 has at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 28, SEQ ID NO: 31, SEQ ID NO: 33, or SEQ ID NO: 104 across the entire sequence.
[0094] In another specific embodiment, the deglon directly linked to the split-intei C fragment via a peptide bond is selected from any of those listed in Table 3. Non-limiting examples of the target protein useful in this section are shown in Table 1 above.
[0095] In yet another preferred embodiment, preferred combinations of C-intane and degron in the split-intane C-fragment degron of the present invention are selected from the group consisting of: The following: Any functionally equivalent mutant of CfaC or CfaCmut, etc., combined with any of the degrons CL1, Deg1, PEST, DD1, DD2, DD3, M1, M2, SopE, SopE-1-78, SopE-15-78, SopE-15-50, L2, L6, L9, L10, L11, L12, L15, L16, M3, M4, M5, V12, DD4, DD5, DD6, and DD7. Preferably: Any functionally equivalent mutant of CfaC or CfaCmut, etc., combined with any of the degrons CL1, Deg1, DD1, DD2, DD3, SopE, L2, L9, M4, V12, DD4, DD5, DD6, or DD7, or The following: Gp41C or any functionally equivalent variant thereof, combined with any of the degrons CL1, Deg1, PEST, DD1, DD2, DD3, M1, M2, SopE, SopE-1-78, SopE-15-78, SopE-15-50, L2, L6, L9, L10, L11, L12, L15, L16, M3, M4, M5, V12, DD4, DD5, DD6, or DD7. Preferably: Gp41C or any functionally equivalent variant thereof, combined with any of the degrons CL1, Deg1, DD1, DD2, DD3, SopE, L2, L9, M4, V12, or DD4.
[0096] Figure 18B shows the effect of fusing degron to the C-terminal construct. As observed with the N-terminal degron, fusing degron to the C-terminal construct reduces the construct's expression level, but no effect on splicing yield is observed when a preferred combination is used. The following degrons are particularly compatible with CfaC: DD1, DD2, DD3, SopE, L2, L9, M4, V12, DD4, DD5, DD6, or DD7. As observed with the N-terminal fragment, incorporating degron into the C-terminal fragment of the present invention reduced its expression (detected in WB), but interestingly, no adverse effect on PTS yield, i.e., the formation of the target product, was observed. This result has been observed with several proteins, including the examples of EGFP and ABCA4 shown herein.
[0097] complex containing split-intei C fragment In another embodiment, the present invention relates to any (or more) of the following complexes, namely the split-intine C fragment of the present invention or the split-intine C fragment degron of the present invention (hereinafter referred to as the second (or more) complex of the present invention), each of these complexes being: (i) The C-terminal fragment of the target protein and... (ii) A split-intine C fragment or split-intine C fragment (split-intine C fragment degron of the present invention) that is directly linked to degron via a peptide bond, optionally by a peptide linker, as defined in the above section, Includes, The complex (or there may be multiple complexes) optionally includes a linker between (i) and (ii), The C-terminal fragment of the target protein is linked to the C-terminus of the split-intene C fragment by an amide bond, or If the complex(s) contains a linker, the C-terminal fragment of the protein of interest is bound to the linker by an amide bond, and / or the linker is bound to the C-terminus of the split-intane C fragment by an amide bond.
[0098] Non-limiting examples of target proteins useful in the present invention are listed in Table 1 above. Further examples of target proteins can be selected from antibodies, antibody fragments containing Fc domains or scFv, nanobodies, bispecific antibodies, proteins, preferably any protein greater than 25 kDa, 50 kDa, or 100 kDa.
[0099] The terms “target protein” and “linker” have already been defined in relation to the first complex of the present invention. All specific embodiments of the target protein and linker of the first complex of the present invention are also fully applicable to the second complex of the present invention.
[0100] In certain embodiments, the complex does not contain a linker between the compound of interest and the split-intane C fragment. In this particular embodiment, the compound of interest is linked to the C-terminus of the split-intane C fragment by an amide bond.
[0101] In certain embodiments, the complex includes a linker between the compound of interest and the split-intane C fragment. In this particular embodiment, the compound of interest may be bonded to the linker by any suitable means, depending on the chemical properties of the compound of interest and the linker. In this particular embodiment, the linker is bonded to the C-terminus of the split-intane C fragment by an amide bond. In another particular embodiment, the compound of interest is bonded to the linker by an amide bond, in which case the linker may be bonded to the C-terminus of the split-intane C fragment by any suitable means. In yet another particular embodiment, the compound of interest is bonded to the linker by an amide bond, and the linker is bonded to the C-terminus of the split-intane C fragment by an amide bond.
[0102] In a particular embodiment, if the complex includes a linker, the linker is a peptide linker. In this particular embodiment, the complex is a fusion protein.
[0103] Compositions containing the composite of the present invention In another embodiment, the present invention refers to a composition comprising a first(or more) complex and / or a second(or more) complex (hereinafter referred to as the first composition of the present invention).
[0104] The term “composition” is intended to encompass not only products containing specific components, but also products that result directly or indirectly from the combination of specific components in specific amounts. The components of a composition may be filled together in a single formulation, or they may be filled separately in different formulations. Thus, in one embodiment, the first complex of the present invention is filled together with the second complex of the present invention in a single formulation. In another embodiment, the first complex of the present invention and the second complex of the present invention are filled separately.
[0105] In a particularly preferred embodiment, the first and second complexes each contain the same N-terminal and C-terminal fragments of the protein, respectively, such that when both complexes are combined according to the method of the present invention, the N-terminal fragment of the protein is linked to the C-terminal fragment of the protein to generate the whole protein.
[0106] Polynucleotides, vectors, and host cells of the present invention In another embodiment, the present invention relates to polynucleotides encoding a first(or more) complex or a second(or more) complex of the present invention.
[0107] Preferably, in preferred embodiments, the present invention refers to two polynucleotides, one of which encodes a first complex (or more) of the present invention (hereinafter referred to as the first polynucleotide of the present invention), and the other which encodes a second complex (or more) of the present invention (hereinafter referred to as the second polynucleotide of the present invention), wherein the first and second polynucleotides of the present invention each encode the same N-terminal and C-terminal fragments of the protein, respectively, such that when both polynucleotides are translated into their respective protein complexes and / or combined according to the methods of the present invention, the N-terminal fragment of the protein is ligated to the C-terminal fragment of the protein, thus producing the entire reconstituted protein. It is important to note that the first and second polynucleotides of the present invention each preferably encode split inteins of the same intein, i.e., the N-intine and C-intine are a homogeneous pair (see Shah et al. Journal of the American Chemical Society 2012 https: / / doi.org / 10.1021 / ja303226x). Examples of congeneral intein pairs include CfaN / CfaC (or any of their variants, including CfaCmut), NpuN / NpuC (or any of their variants), CatN / CatC (or any of their variants), and Gp41N / Gp41C (or any of their variants). However, regardless of whether both polynucleotides encode a split intein involved in the same intein, when translated into protein, the N-terminal and C-terminal sequences of each split intein must become separate fragments that can be non-covalently reassembled, i.e., reconstructed, into a functional intein for trans-splicing reactions.
[0108] In a preferred embodiment, the present invention refers to the two polynucleotides mentioned above, one of which encodes the split-intane N-fragment degron of the present invention (hereinafter referred to as the first polynucleotide encoding the degron of the present invention), and the other which encodes the split-intane C-fragment degron of the present invention (hereinafter referred to as the second polynucleotide encoding the degron of the present invention), wherein the first and second polynucleotides encoding the degron of the present invention each encode the same N-terminal and C-terminal fragments of the protein, respectively, such that when both polynucleotides are translated into their respective protein complexes and combined according to the method of the present invention, the N-terminal fragment of the protein is ligated to the C-terminal fragment of the protein, thus producing the entire protein. As in the previous embodiment, it is important to note that the first and second polynucleotides encoding the degron of the present invention each preferably encode the same intein split-intane. In any case, and as reflected above, when used herein, each polynucleotide must encode a split intein such that, when translated into a protein, the N-terminal and C-terminal sequences become separate fragments that can be non-covalently reassembled or reconfigured into a functional intein for trans-splicing reactions.
[0109] In yet another preferred embodiment, the present invention refers to a three-piece ligation strategy using orthogonal split-intane pairs. This methodology has three polynucleotides, the first of which is POI N -CfaN(POI N The first part codes for the "target protein" (which is understood as the N fragment), and the second polynucleotide is CfaC-POI M -CatN(in this case, POI M (This is an intermediate fragment of the target protein), and the third polynucleotide is CatC-POI C (POIC encodes what is understood as the C - fragment of the target protein. Alternatively, by swapping the positions of the Cfa intein and the Cat intein, a construct with the following structure: POI N - CatN, CatC - POI M - CfaN, and CfaC - POI C can also be generated. CfaN is either the CfaN intein fragment SEQ ID NO: 27 or a variant thereof. CatN is either the CatN intein fragment SEQ ID NO: 30 or a variant thereof. CfaC is either the CfaC intein fragment SEQ ID NO: 28 or a variant thereof including SEQ ID NO: 29. CatC is either the CatC intein fragment SEQ ID NO: 31 or a variant thereof.
[0110] It is noted that, in the present specification, for those embodiments referring to two polynucleotides, the first polynucleotide (or specifically, the first polynucleotide encoding the degron of the present invention) can be any split - intein N - fragment of the present invention or any split - intein N - degron fragment of the present invention. In particular, the first polynucleotide (or specifically, the first polynucleotide encoding the degron of the present invention) of the present invention can be selected from any of the following list: SEQ ID NO: 1, SEQ ID NO: 3, SEQ ID NO: 5, SEQ ID NO: 8, SEQ ID NO: 10, SEQ ID NO: 12, SEQ ID NO: 14, SEQ ID NO: 17, SEQ ID NO: 19, SEQ ID NO: 21, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 66, SEQ ID NO: 68, SEQ ID NO: 70, SEQ ID NO: 72, SEQ ID NO: 74, SEQ ID NO: 76 and SEQ ID NO: 78. More preferably, the first polynucleotide encoding the degron of the present invention encodes any of the following specific combinations of intein and degron: The following: Functionally equivalent variants of CfaN or ConN, etc., combined with any of the following degrons: CL1, Deg1, PEST, DD1, DD2, DD3, M1, M2, SopE, SopE-1-78, SopE-15-78, SopE-15-50, L2, L6, L9, L10, L11, L12, L15, L16, M3, M4, M5, V12, DD4, DD5, DD6, and DD7. Preferably: Functionally equivalent variants of CfaN or ConN, etc., combined with any of the following degrons: CL1, Deg1, DD1, DD2, DD3, SopE, L2, L9, M4, V12, or DD4, or The following: Gp41N or a functionally equivalent variant thereof in combination with any of the degrons CL1, Deg1, PEST, DD1, DD2, DD3, M1, M2, SopE, SopE-1-78, SopE-15-78, SopE-15-50, L2, L6, L9, L10, L11, L12, L15, L16, M3, M4, M5, V12, DD4, DD5, DD6, or DD7. Preferably: Gp41N or a functionally equivalent variant thereof in combination with any of the degrons CL1, Deg1, DD1, DD2, DD3, SopE, L2, L9, M4, V12, or DD4.
[0111] In this specification, with respect to embodiments referring to two polynucleotides, it should be further noted that the second polynucleotide (or specifically the second polynucleotide encoding the degron of the present invention) may be any split-intine C fragment of the present invention or any split-intine C degron fragment of the present invention. In particular, the second polynucleotide of the present invention (or specifically the second polynucleotide encoding the degron of the present invention) may be selected from any of the following lists: SEQ ID NOs: 2, 4, 6, 9, 11, 13, 15, 18, 20, 22, 26, 67, 69, 71, 73, 75, 77, and 79. More preferably, the second polynucleotide encoding the degron of the present invention may encode any of the following specific combinations of intein and degron: The following: Functionally equivalent variants of CfaC, CfaCmut, or ConC, etc., combined with any of the following degrons: CL1, Deg1, PESTt, DD1, DD2, DD3, M1, M2, SopE, SopE-1-78, SopE-15-78, SopE-15-50, L2, L6, L9, L10, L11, L12, L15, L16, M3, M4, M5, V12, DD4, DD5, DD6, and DD7. Preferably, the following: Functionally equivalent variants of CfaC, CfaCmut, etc., combined with any of the following degrons: CL1, Deg1, DD1, DD2, DD3, SopE, L2, L9, M4, V12, DD4, DD5, DD6, or DD7, or The following: Gp41C or a functionally equivalent variant thereof in combination with any of the degrons CL1, Deg1, PEST, DD1, DD2, DD3, M1, M2, SopE, SopE-1-78, SopE-15-78, SopE-15-50, L2, L6, L9, L10, L11, L12, L15, L16, M3, M4, M5, V12, DD4, DD5, DD6, or DD7. Preferably: Gp41C or a functionally equivalent variant thereof in combination with any of the degrons CL1, Deg1, DD1, DD2, DD3, SopE, L2, L9, M4, V12, or DD4.
[0112] In a preferred embodiment, the first and second polynucleotides encoding the degron of the present invention encode the N-terminal and C-terminal fragments of the ABCA4 protein, respectively, such that when both polynucleotides are translated into their respective protein complexes and combined according to the method of the present invention, the N-terminal fragment of the ABCA4 protein is ligated to the C-terminal fragment of the ABCA4 protein to produce the entire ABCA4 protein. More preferably, the first polynucleotide encoding the degron of the present invention encodes positions 1-1149, 1-1139, 1-1178, or 1-1187 of the N-terminal fragment of the ABCA4 protein with any of the following specific combinations of intein and degron: The following: Functionally equivalent variants of CfaN or ConN, etc., combined with any of the following degrons: CL1, Deg1, PEST, DD1, DD2, DD3, M1, M2, SopE, SopE-1-78, SopE-15-78, SopE-15-50, L2, L6, L9, L10, L11, L12, L15, L16, M3, M4, M5, V12, DD4, DD5, DD6, and DD7. Preferably: Functionally equivalent variants of CfaN or ConN, etc., combined with any of the following degrons: CL1, Deg1, DD1, DD2, DD3, SopE, L2, L9, M4, V12, or DD4; The second polynucleotide encoding degron of the present invention encodes positions 1150-2273, 1140-2273, 1179-2273, or 1188-2273 of the C-terminal fragment of the ABCA4 protein, and any of the following specific combinations of intein and degron: The following: CfaC combined with any of the degrons CL1, Deg1, PEST, DD1, DD2, DD3, M1, M2, SopE, SopE-1-78, SopE-15-78, SopE-15-50, L2, L6, L9, L10, L11, L12, L15, L16, M3, M4, M5, V12, DD4, DD5, DD6, and DD7, or functionally equivalent variants thereof such as CfaCmut or ConC. Preferably: CfaC combined with any of the degrons CL1, Deg1, DD1, DD2, DD3, SopE, L2, L9, M4, V12, DD4, DD5, DD6, or DD7, or functionally equivalent variants thereof such as CfaCmut or CfaCmut.
[0113] It is important to note that if the first polynucleotide encoding degron in the present invention encodes positions 1 to 1149, the N-terminal fragment of the ABCA4 protein must be ligated with the C-terminal fragment of the ABCA4 protein, thereby generating the entire ABCA4 protein. Therefore, the second polynucleotide encoding degron in the present invention encodes positions 1150 to 2273 of the C-terminal fragment of the ABCA4 protein. The same applies to other positions.
[0114] Using the preferred combination described above, it has been confirmed that efficient reconstitution of the target protein ABCA4, in these examples, can be obtained while significantly reducing the levels of the N-terminal and C-terminal starting fragments. If one construct contains at least one degron, or both, the degradation of the N-terminal and / or C-terminal fragments is promoted, resulting in reduced expression. Based on the state of the art, such a reduction in the levels of starting materials would be expected to result in a decrease in the splice product. Nevertheless, the inventors observed the opposite behavior, in which the yield of the full-length splicing product was unaffected and, in some cases, even improved. Furthermore, the inventors observed that by reconstituting the protein using the method described herein, they were able to produce a functional protein, as determined by an ATPase activity assay. The inventors also confirmed efficient reconstitution of the target protein in vivo in the mouse retina.
[0115] In another preferred embodiment, the first polynucleotide encoding the degron of the present invention encodes positions 1-1095, or 1-1185, of the N-terminal fragment of the ABCA4 protein, and any of the following specific combinations of intein and degron: The following: Gp41N or a functionally equivalent variant thereof in combination with any of the following degrons: CL1, Deg1, PEST, DD1, DD2, DD3, M1, M2, SopE, SopE-1-78, SopE-15-78, SopE-15-50, L2, L6, L9, L10, L11, L12, L15, L16, M3, M4, M5, V12, DD4, DD5, DD6, or DD7. Preferably: Gp41N or a functionally equivalent variant thereof in combination with any of the following degrons: CL1, Deg1, DD1, DD2, DD3, SopE, L2, L9, M4, V12, or DD4; The second polynucleotide encoding degron of the present invention encodes positions 1096-2273, or 1186-2273, of the C-terminal fragment of the ABCA4 protein, and any of the following specific combinations of intein and degron: The following: Gp41C or a functionally equivalent variant thereof in combination with any of the following degrons: CL1, Deg1, PEST, DD1, DD2, DD3, M1, M2, SopE, SopE-1-78, SopE-15-78, SopE-15-50, L2, L6, L9, L10, L11, L12, L15, L16, M3, M4, M5, V12, DD4, DD5, DD6, or DD7. Preferably: Gp41C or a functionally equivalent variant thereof in combination with any of the following degrons: CL1, Deg1, DD1, DD2, DD3, SopE, L2, L9, M4, V12, or DD4.
[0116] Similar to the previous example, if the first polynucleotide encoding degron in the present invention encodes positions 1 to 1095, the N-terminal fragment of the ABCA4 protein must be ligated with the C-terminal fragment of the ABCA4 protein, thereby generating the entire ABCA4 protein. Therefore, it is important to note that the second polynucleotide encoding degron in the present invention encodes positions 1096 to 2273 of the C-terminal fragment of the ABCA4 protein. The same applies to other positions.
[0117] A more preferred combination of nucleotide sequences encoding the above sequence (the first and second polynucleotides of the present invention, or the first and second polynucleotides encoding degron of the present invention, respectively) is selected from any of the following pairs: 1 and 2; 3 and 4; 5 and 6; 66 and 67; 68 and 69; 70 and 71; 72 and 73; 74 and 75; 76 and 77; and 78 and 79.
[0118] The terms “the first polynucleotide and the second polynucleotide of the present invention” include, in particular, “the first polynucleotide and the second polynucleotide encoding degron of the present invention,” and therefore, the inventors should note that hereafter, “the first polynucleotide and the second polynucleotide of the present invention” are referred to only as terms that already encompass all of the above-mentioned substitutes for polynucleotides.
[0119] Preferably, the present invention relates to a composition comprising the first polynucleotide and / or second polynucleotide of the present invention, or the first polynucleotide and / or second polynucleotide encoding the degron of the present invention, and / or each of the POI-intane fragments of a three-piece ligation system using an orthogonal split-intane pair, separately or jointly, hereinafter referred to as the second composition of the present invention. The term “composition” is intended in this particular context to encompass products containing the specified components. The components of the composition may be filled together in a single formulation or separately in different formulations. Thus, in an embodiment, the first polynucleotide of the present invention is filled together with the second polynucleotide of the present invention in a single formulation. In another embodiment, the first polynucleotide and the second polynucleotide of the present invention are filled separately. More preferably, the first polynucleotide and the second polynucleotide of the present invention encode the N-terminal and C-terminal fragments of a reconstituted protein, respectively, such that when the two complexes are combined according to the method of the present invention, the N-terminal fragment of the protein is ligated to the C-terminal fragment of the protein, generating the entire reconstituted protein.
[0120] All of the polynucleotides described above in the present invention may be found to be isolated on their own or to form part of a vector that enables the delivery and / or propagation of the polynucleotides in a suitable host cell. In another embodiment, the present invention relates to a vector comprising any of the polynucleotides described above in the present invention.
[0121] Vectors suitable for the insertion of the above polynucleotides include vectors derived from prokaryotic expression vectors such as pUC18, pUC19, Bluescript and its derivatives, mpl8, mpl9, pBR322, pMB9, ColEl, pCRl, RP4, and phages, as well as "shuttle" vectors such as pSA3 and pAT28; yeast expression vectors such as 2-micron plasmids, integrated plasmids, YEP vectors, and centromere plasmid type vectors; insect cell expression vectors such as the pAC series and pVL series vectors; plant expression vectors such as the pIBI, pEarleyGate, pAVA, pCAMBIA, pGSA, pGWB, pMDC, pMY, and pORE series; and eukaryotic cell expression vectors containing baculoviruses suitable for transfection of insect cells using any commercially available baculovirus system. Preferred vectors for eukaryotic cells include viral vectors (adenovirus, adeno-associated virus (AAV), adenovirus-related viruses such as retroviruses, particularly lentiviruses), and non-viral vectors (pSilencer 4.1-CMV (Ambion), pcDNA3, pcDNA3.1 / hyg, pHMCV / Zeo, pCR3.1, pEFI / His, pIND / GS, pRc / HCMV2, pSV40 / Zeo2, pTRACER-HCMV, pUB6 / V5-His, pVAXl, pZeoSV2, pCI, pSVL, and PKSV-10, pBPV-1, pML2d, and pTDTl, etc.). Preferably, the vector is adeno-associated virus (AAV). Preferably, the vector is AAV serotype 1, 2, 3, 4, 5, 6, 7, 8, or 9. Dimers or self-complementary AAV vectors (scAAVs) can also be used for the insertion of the polynucleotides described above.
[0122] Therefore, the present invention is preferably directed toward the development of AAV as a gene therapy vector. Preferably, these vectors have lost their ability to integrate by removing rep and cap from the vector DNA. The desired gene (polynucleotide), along with a promoter that drives gene transcription, is inserted between reverse terminal repeats (ITRs) that facilitate concatemer formation in the nucleus after the single-stranded vector DNA has been converted to double-stranded DNA by the host cell's DNA polymerase complex. Gene therapy vectors using AAV form episomal concatemers in the nucleus of the host cell. In non-dividing cells, these concatemers remain until the end of the host cell's lifespan. In dividing cells, the episomal DNA is not replicated along with the host cell's DNA, so the AAV DNA is lost during cell division. Random integration of AAV DNA into the host genome is detectable but occurs at a very low frequency.
[0123] A desired gene (polynucleotide) can be combined with a regulatory element that functions in the host cell in which the construct is to be expressed. Those skilled in the art can select a suitable regulatory element for use in a host cell, e.g., a mammalian or human host cell. Examples of regulatory elements include promoters, transcription termination sequences, translation termination sequences, enhancers, signal peptides, and polyadenylation elements. The polynucleotides of the present invention can be operably ligated to promoter sequences. Promoter contemplated for use in the subject matter of the invention includes, but is not limited to, native gene promoters, cytomegalovirus (CMV) promoters (KF853603.1, bp149~735), chimeric CMV / chicken beta-actin promoters (CBA), and cleaved CBA (smCBA) promoters (U.S. Patent No. 8,298818, and Light-Driven Cone Arrestin Translocation in Cones of Postnatal Guanylate Cyclase-1 Knockout Mouse Retina Treated with AAVGC). 1) Examples include the rhodopsin promoter (NG009115, bp4205~5010), the photoreceptor-retinoid-binding protein (IRBP) promoter (NG_029718.1, bp4777~5011), the vitiligo macular degeneration 2 (VMD2) promoter (NG009033.1, bp4870~5470), the PR-specific human G protein-coupled receptor kinase 1 (hGRKl; AY327580.1 bp1793~2087 or bp1793~1991) (Haire et al. 2006; U.S. Patent No. 8,298,818), and the proximal mouse rhodopsin promoter (MOPS). However, any suitable promoter known in the art can be used. In specific embodiments, the promoter is the CMV promoter or the hGRKl promoter. In one embodiment, the promoter is a tissue-specific promoter that exhibits selective activity in one tissue or a group of tissues, but has low or no activity in other tissues.In one embodiment, the promoter is a photoreceptor-specific promoter. In a further embodiment, the promoter is a cone cell-specific and / or rod cell-specific promoter.
[0124] Preferred promoters are CMV, GRK1, CBA, and IRBP promoters. More preferred promoters are hybrid promoters that combine regulatory elements from various promoters (for example, a chimeric CBA promoter combining an enhancer from the CMV promoter, the CBA promoter, and the Sv40 chimeric intron, referred to herein as the CBA hybrid promoter).
[0125] Furthermore, AAV exhibits very low immunogenicity, seemingly limited to the production of neutralizing antibodies, and does not induce clearly distinguishable cytotoxic reactions. This characteristic, along with its ability to infect quiescent cells, gives it an advantage over adenoviruses as a vector for human gene therapy.
[0126] AAV Genome, Transcriptome, and Proteome: The AAV genome consists of positive-sense or negative-sense single-stranded deoxyribonucleic acid (ssDNA) approximately 4.7 kilobases long. The genome contains reverse terminal repeat sequences (ITRs) at both ends of the DNA strand and two open reading frames (ORFs): rep and cap. The former consists of four duplicated genes encoding the Rep protein, which is essential for the AAV life cycle, while the latter contains duplicated nucleotide sequences of capsid proteins: VP1, VP2, and VP3, which interact to form an icosahedral symmetric capsid.
[0127] Reverse terminal repeat (ITR) sequences each contain 145 base pairs. They have been shown to be necessary for the efficient proliferation of the AAV genome and are named after their symmetry. Another characteristic of these sequences is their ability to form hairpins, which contributes to so-called self-priming, enabling the synthesis of a primer-independent second DNA strand. Furthermore, ITRs have been shown to be necessary for both the integration of AAV DNA into the host cell genome (chromosome 19 in humans) and its rescue from there, and also for the efficient encapsulation of AAV DNA in combination with the generation of fully assembled deoxyribonuclease-resistant AAV particles.
[0128] Regarding gene therapy, it appears that the only sequence required in cis orientation next to the therapeutic gene is the ITR, and structural genes (cap) and packaging genes (rep) can be introduced in trans orientation. Based on this premise, numerous methods have been established for efficiently producing recombinant AAV (rAAV) vectors containing reporter genes or therapeutic genes. However, it has also been reported that the ITR is not the only element required in cis orientation for effective replication and encapsulation. Several research groups have identified a sequence called a cis-acting Rep-dependent element (CARE) within the coding sequence of the rep gene. It has been shown that the presence of CARE in cis orientation enhances replication and encapsulation.
[0129] By 2006, 11 AAV serotypes had already been reported. All known serotypes can infect cells of multiple diverse tissue types. Tissue specificity is determined by the serotype of the capsid, and altering its tropism range by pseudotyping of the AAV vector is considered important for therapeutic use. In this invention, ITRs of AVV serotypes 1, 2, 3, 4, 5, 6, 7, 8, or 9 are preferred.
[0130] Serotype 2 (AAV2) has been the most extensively studied so far. AAV2 exhibits innate tropism towards skeletal muscle, nerve cells, vascular smooth muscle cells, and hepatocytes.
[0131] Furthermore, the vector may include a reporter or marker gene that can identify cells that have taken up the vector after contact.
[0132] Reporter genes useful in relation to the present invention include lacZ, luciferase, thymidine kinase, GFP, etc. Marker genes useful in relation to the present invention include, for example, a neomycin resistance gene that confers resistance to aminoglycoside G418; a hygromycin phosphotransferase gene that confers resistance to hygromycin; an ODC gene that confers resistance to ornithine decarboxylase inhibitor (2-(difluoromethyl)-DL-ornithine (DFMO)); a dihydrofolate reductase gene that confers resistance to methotrexate; a puromycin-N-acetyltransferase gene that confers resistance to puromycin; and resistance to zeosin. Examples include the ble gene which confers sex; the adenosine deaminase gene which confers resistance to 9-beta-D-xylofuranose adenine; the cytosine deaminase gene which enables cell proliferation in the presence of N-(phosphonoacetyl)-L-aspartate; the thymidine kinase which enables cell proliferation in the presence of aminopterin; the xanthine-guanine phosphoribosyltransferase gene which enables cell proliferation in the presence of xanthine and in the absence of guanine; the trpB gene of Escherichia coli (E. coli) which enables cell proliferation in the presence of indole instead of tryptophan; and the hisD gene of Escherichia coli which enables cells to use histidinol instead of histidine. The selected gene is incorporated into a plasmid that may further contain a promoter suitable for the expression of the above gene in eukaryotic cells (e.g., CMV or SV40 promoter), an optimized translation initiation site (e.g., a site following the so-called Kozak rule or IRES), polyadenylation sites such as the SV40 polyadenylation site or phosphoglycerate kinase site, and introns such as the betaglobulin gene intron. Alternatively, a combination of reporter and marker genes can be used simultaneously in the same vector.
[0133] On the other hand, as is known to those skilled in the art, the choice of vector depends on the host cell into which it is subsequently introduced. For example, the vector into which the above polynucleotides are introduced may also be a yeast artificial chromosome (YAC), a bacterial artificial chromosome (BAC), or a PI-derived artificial chromosome (PAC). The properties of YACs, BACs, and PACs are known to those skilled in the art. Detailed information on the above types of vectors is provided, for example, by Giraldo and Montoliu (Giraldo, P. & Montoliu L., 2001 Size matters: use of YACs, BACs and PACs in transgenic animals, Transgenic Research 10(2):83-110). The vectors of the present invention can be obtained by conventional methods known to those skilled in the art (Sambrook J. et al., 2000 "Molecular cloning, a Laboratory Manual", 3rd ed., Cold Spring Harbor Laboratory Press, NY Vol 1-3).
[0134] The polynucleotides of the present invention can be introduced into host cells in vivo as naked DNA plasmids, but they can also be introduced using vectors by methods well known in the art, such as transfection, electroporation (e.g., percutaneous electroporation), microinjection, transduction, cell fusion, DEAE dextran, calcium phosphate precipitation, the use of gene guns, or the use of DNA vector transporters. Methods for formulating naked DNA and administering it to mammalian muscle tissue are also known. See Feigner P, et al., U.S. Patents 5,580,859 and 5,589,466. Other molecules that promote nucleic acid transfection in vivo, such as cationic oligopeptides, peptides derived from DNA-binding proteins, or cationic polymers, are also useful. See Bazile D, et al., International Publication No. 1995021931 and Byk G, et al., International Publication No. 1996025508.
[0135] Another well-known method that can be used to introduce polynucleotides into host cells is particle beam irradiation (also known as bioristic transformation). Bioristic transformation is generally carried out in one of several ways. A common method involves propelling inert or biologically active particles toward the cells. See Sanford J, et al., U.S. Patents 4,945,050, 5,036,006, and 5,100,792.
[0136] Alternatively, vectors can be introduced in vivo by lipofection. Cationic lipids can be used to promote the encapsulation of negatively charged nucleic acids and their fusion with negatively charged cell membranes. See Feigner P, Ringold G, Science 1989; 337:387-388. Lipid compounds and compositions particularly useful for nucleic acid transcription are described. See Feigner P, et al., U.S. Patent No. 5,459,127, Behr J, et al., International Publication No. 1995018863, and Byk G, International Publication No. 1996017823.
[0137] Finally, and especially preferably, the vector can be introduced in vivo by viral delivery systems, including but not limited to adenovirus vectors, adeno-associated virus (AAV) vectors, pseudotyped AAV vectors, herpesvirus vectors, retrovirus vectors, lentivirus vectors, and baculovirus vectors. A pseudotyped AAV vector contains the genome of a certain AAV serotype in the capsid of a second AAV serotype; for example, the AAV2 / 8 vector contains the AAV8 capsid and the AAV2 genome (Auricchio et al. (2001) Hum. Mol. Genet. 10(26):3075-81). Such vectors are also known as chimeric vectors. Examples of other delivery systems include, but are not limited to, ex vivo delivery systems, and DNA transfection methods such as electroporation, DNA bioristics, lipid-mediated transfection, and compactified DNA-mediated transfection.
[0138] The construction of AAV vectors can be carried out according to procedures and techniques known to those skilled in the art. The theory and practice of constructing and using adeno-associated virus vectors in therapy are described in several scientific publications and patent publications (the following references constitute part of this specification by reference: Flotte TR. Adeno-associated virus-based gene therapy for inherited disorders. Pediatr Res. 2005 Dec; 58(6):1143-7; Goncalves MA. Adeno-associated virus: from defective virus to effective vector, Virol J. 2005 May 6;2:43; Surace EM, Auricchio A. Adeno-associated viral vectors for retinal gene transfer. Prog Retin Eye Res. 2003 Nov; 22(6):705-19; Mandel RJ, Manfredsson FP, Foust KD, Rising A, Reimsnider S, Nash K, Burger C. Recombinant adeno-associated viral vectors as therapeutic agents to treat neurological disorders. Mol Ther. 2006 Mar; 13(3):463-83).
[0139] Preferred forms of administration for pharmaceutical compositions containing AAV vectors include, but are not limited to, injectable solutions or suspensions, eye lotions, and ophthalmic ointments. Therefore, in another embodiment, the present invention relates to host cells containing the polynucleotide or vector of the present invention. These cells can be obtained by conventional methods known to those skilled in the art (see, for example, Sambrook et al.).
[0140] As used herein, the term “host cell” refers to a cell into which nucleic acids of the present invention, such as polynucleotides or vectors according to the present invention, have been introduced and which can express the split-intane N fragment of the present invention or a fusion protein containing the split-intane N fragment. In this specification, the terms “host cell” and “recombinant host cell” are used interchangeably. It should be understood that such terms refer not only to specific target cells but also to the offspring or potential offspring of such cells. Because certain changes may occur in the progeny due to mutation or environmental influences, such offspring may not actually be identical to the parent cells, but they are still included within the scope of the terms as used herein. This term also includes cultureable cells that can be modified by the introduction of heterologous DNA. Preferably, the host cell is capable of stably expressing the polynucleotide of the present invention, undergoing post-translational modification, localizing to an appropriate intracellular compartment, and enabling it to participate in an appropriate transcriptional mechanism. The selection of an appropriate host cell also influences the selection of the detection signal. For example, the reporter construct can provide a selectable or screenable trait for activation or inhibition of gene transcription in response to a transcription regulatory protein, as described above, and the phenotype of the host cell is considered to achieve optimal selection or screening. Examples of host cells of the present invention include prokaryotic cells and eukaryotic cells. Examples of prokaryotes include Gram-negative or Gram-positive organisms, such as Escherichia coli or Bacillus. Preferably, prokaryotic cells are used for the proliferation of transcription regulatory sequences containing polynucleotides or vectors of the present invention. Examples of prokaryotic host cells suitable for transformation include Escherichia coli, Bacillus subtilis, Salmonella typhimurium, and various other species within the genera Pseudomonas, Streptomyces, and Staphylococcus. Eukaryotic cells include, but are not limited to, yeast cells, plant cells, fungal cells, insect cells (e.g., baculoviruses), mammalian cells, and parasitic cells (e.g., trypanosomes).As used herein, yeast includes not only yeast in the strict taxonomic sense, i.e., unicellular organisms, but also yeast-like multicellular fungi of the filamentous fungi. Exemplary species include Kluyverei lactis, Schizosaccharomyces pombe, and Ustilaqo maydis, with Saccharomyces cerevisiae being preferred. Other yeasts that can be used in carrying out the present invention include Neurospora crassa, Aspergillus niger, Aspergillus nidulans, Pichia pastoris, Candida tropicalis, and Hansenula polymorpha. Established cell lines such as COS cells, L cells, 3T3 cells, Chinese hamster ovary (CHO) cells, and embryonic stem cells are used as mammalian host cell culture systems, with BHK cells, HeK cells, or HeLa cells being preferred. Eukaryotic cells are preferred for recombinant gene expression.
[0141] In a preferred embodiment, the second composition of the present invention is for therapeutic use, and more specifically for use in any of the diseases specified in Table 1 depending on the type of gene encoded in the composition (the correlation between diseases and genes is clearly shown in Table 1, and therefore identifying the correct combination will be obvious to those skilled in the art).
[0142] A method for expressing a gene that codes for a target protein within a cell. In another embodiment, the present invention relates to a method for expressing a target gene in a cell in vitro or in vivo (hereinafter referred to as the first method for expressing the target gene), and this method is (i) cells, (a) The first polynucleotide of the present invention, or the first polynucleotide encoding the degron of the present invention, (b) Contacting the second polynucleotide of the present invention, or the second polynucleotide encoding degron of the present invention, (ii) Expressing the first polynucleotide and the second polynucleotide so that the first fusion protein and the second fusion protein are produced, (iii) The split intein N fragment binds to the split intein C fragment to form an intein intermediate, and the intein intermediate reacts to bring the first protein and the second protein into contact so that the C-terminus of the first target polypeptide is covalently bonded to the N-terminus of the second target polypeptide. Includes.
[0143] In certain embodiments, the first polynucleotide is a first polynucleotide encoding degron, the second polynucleotide is a second polynucleotide encoding degron, preferably the target protein having a magnitude greater than 25 kDa, greater than 50 kDa, or greater than 100 kDa, and the entire protein is obtained by covalently bonding the C-terminus of the target first polypeptide to the N-terminus of the target second polypeptide.
[0144] Contact between the first polynucleotide and / or the second polynucleotide of the present invention and cells can be made in vitro or in vivo by any suitable means for introducing the polynucleotide of interest into cells, such as transfection, electroporation, microinjection, transduction, lipofection, cell fusion, DEAE dextran, calcium phosphate precipitation, use of a gene gun, or use of a DNA vector transporter. Preferably, the vector is adeno-associated virus (AAV).
[0145] In a preferred embodiment, the present invention relates to a method for expressing a gene encoding a target protein in a cell (hereinafter referred to as the first method for expressing the target gene), and this method is (i) cells, (a) A first AAV comprising the first polynucleotide of the present invention, or the first polynucleotide encoding the degron of the present invention, and (b) Contacting or transducing a second AAV containing the second polynucleotide of the present invention, or the second polynucleotide encoding degron of the present invention, (ii) Expressing the first polynucleotide and the second polynucleotide so that the first fusion protein and the second fusion protein are produced, (iii) The split intein N fragment binds to the split intein C fragment to form an intein intermediate, and the intein intermediate reacts to bring the first fusion protein and the second fusion protein into contact so that the C-terminus of the first target polypeptide is covalently bonded to the N-terminus of the second target polypeptide. Includes, Preferably, the first polynucleotide is a first polynucleotide encoding degron, and the second polynucleotide is a second polynucleotide encoding degron. More preferably, the target protein has a molecular weight greater than 25 kDa, greater than 50 kDa, or greater than 100 kDa. The entire protein is obtained by covalently bonding the C-terminus of the target first polypeptide to the N-terminus of the target second polypeptide.
[0146] In this sense, and as shown in the examples, the presence of degron causes rapid degradation of the starting material (i.e., the N-terminal and / or C-terminal fragment degron of the split-intei of the present invention). The selected degron has a degradation rate suitable for protein splicing, and at steady state, the amount of starting material is reduced compared to that observed in the absence of degron. However, the level of the splicing product is maintained or even increased compared to that observed in the absence of degron.
[0147] In the method for expressing the target gene of the present invention, cells are intended to be brought into contact with the first polynucleotide and the second polynucleotide simultaneously, or to be brought into contact with the first polynucleotide and the second polynucleotide sequentially in any order. That is, cells may be brought into contact with the first polynucleotide first, then the second polynucleotide, or with the second polynucleotide first, then the first polynucleotide. The same applies to the vector encoding the polynucleotide.
[0148] These methods can use any cell that has been predefined as a host cell.
[0149] The present invention will be described by the following examples, but these are merely illustrative and are not intended to limit the scope of the invention. [Examples]
[0150] In the following table, the inventors establish a nomenclature for sequences formed by adding sequences already enumerated in this book. Such sequences are represented in the following table as a list of sequences identified by sequence numbers, from the N-terminus to the C-terminus. For example, for a sequence in which amino acids 1-1149 of sequence number 106 are directly linked to sequence number 27 (CfaN), which is then directly linked to sequence number 103 (3FT), the nomenclature used is as follows: sequence number 106(1-1149)-sequence number 27-sequence number 103. This specific example should be interpreted as a sequence that begins with amino acids 1-1149 of sequence number 106 from the N-terminus, followed immediately by the polypeptide corresponding to sequence number 27, which is directly linked via a peptide bond, and then immediately by the polypeptide corresponding to sequence number 103, which is also linked via a peptide bond. Therefore, the sequence may have the following structure from the N-terminus to the C-terminus: [Sequence ID 106 (residues 1-1149)]-[Sequence ID 27]-[Sequence ID 103]
[0151] Here, the sequence enclosed in brackets "[ ]" represents all the amino acids within that sequence, and "-" represents a peptide bond connecting the two sequences within the brackets. The residues from x to y in a particular sequence mean all amino acid residues from position x to position y (including both positions x and y). For example, sequence number 106 (residues 1 to 1149) means the sequence containing residues 1 to 1149 of sequence number 106 (including both residues 1 and 1149).
[0152] [Table 5] TIFF0007867703000011.tif250170TIFF0007867703000012.tif92170TIFF0007867703000013.tif151170
[0153] Materials and methods material: Oligonucleotides were purchased from Eurofins Genomics. Synthetic genes were purchased from Genewiz. Pfu Ultra fusion polymerase and all restriction enzymes for cloning were purchased from Thermofisher Scientific. High-competency cells used for cloning were prepared from XL10-Gold chemically competent E. coli. HEK293T cells were purchased from ATCC. A DNA purification kit was purchased from Qiagen. All plasmids were sequenced using Macrogen. Luria Bertani (LB) medium and all buffer salts were purchased from Thermofisher Scientific. Coomassie Brilliant Blue, MG-132 proteasome inhibitor, phenylmethanesulfonyl fluoride, iodoacetamide, NH4HCO3, DTT, formic acid, fetal bovine serum, and soybean-derived asolectin were purchased from Sigma-Aldrich. Acetonitrile (ACN) was purchased from Carlo-Erba. EDTA-free complete protease inhibitors were purchased from Roche. Lipofectamine 2000 transfection reagent, DMEM high glucose GlutaMAX supplement, RPMI1640 medium GlutaMAX supplement, RIPA lysis and extraction buffer, BCA protein assay kit, MES-SDS running buffer, pre-stained protein ladder, and SDS-PAGE (Bis-tris gel and Tris-acetic acid gel) were purchased from Thermofisher Scientific. Primary anti-6×His tag mouse monoclonal antibody, anti-flag tag mouse monoclonal antibody, anti-ABCA4 rabbit polyclonal antibody, and anti-tubulin rabbit polyclonal antibody were purchased from Invitrogen. Secondary antibodies, goat anti-mouse IgG(H+L) highly cross-adsorbed antibody alexa fluor plus 680 and goat anti-rabbit IgG(H+L) secondary antibody dylight 800 4×PEG, were purchased from Invitrogen. Solutions of dodecyl maltoside (D310) and cholesteryl hemisuccinate (CH210) were purchased from Anatrace. I bought trypsin from Promega.
[0154] device: Electrospray ionization mass spectrometry (ESI-MS) was performed using a no-acquity liquid chromatograph (Waters) connected to an LTQ-Orbitrap Velos (Thermo Scientific) mass spectrometer. Gel and Western blot images were acquired using a LI-COR Odyssey Infrared Imager. Cell lysis was performed using an SFX550 Branson sonifier. FACS measurements were performed using a Gallios Beckman Coulter.
[0155] Cloning of recombinant DNA Synthetic genes for constructs ABCA4-1150-CfaN (SEQ ID NO: 1) and ABCA4-1150-CfaC (SEQ ID NO: 2) were purchased and introduced into a pEGFP-N1 expression vector using KpnI and NotI restriction enzymes. Synthetic genes for NpuN and NpuC were purchased and introduced into ABCA4-1150-CfaN and ABCA4-1150-CfaC by restriction enzyme-free cloning to obtain constructs ABCA4-1150-NpuN (SEQ ID NO: 8) and ABCA4-1150-NpuC (SEQ ID NO: 9).
[0156] Constructs ABCA4-3FT (SEQ ID NO: 7), ABCA4-1140-CfaN (SEQ ID NO: 3), ABCA4-1140-CfaCmut (SEQ ID NO: 4), ABCA4-1188-CfaN (SEQ ID NO: 5), ABCA4-1188-CfaCmut (SEQ ID NO: 6), ABCA4-1179-CfaN (SEQ ID NO: 93), and ABCA4-1179-CfaCmut (SEQ ID NO: 94) were prepared by restriction enzyme-free cloning. Mutations in CfaCmut were introduced using inverse PCR with Pfu Ultra II HF polymerase. ABCA4-1140-NpuN (SEQ ID NO: 10), ABCA4-1140-NpuC (SEQ ID NO: 11), ABCA4-1188-NpuN (SEQ ID NO: 12), and ABCA4-1188-NpuC (SEQ ID NO: 13) were prepared by restriction enzyme-free cloning.
[0157] The constructs EGFP-71-CfaN (SEQ ID NO: 14), EGFP-71-CfaC (SEQ ID NO: 15), EGFP (SEQ ID NO: 16), EGFP-71-NpuN (SEQ ID NO: 17), and EGFP-71-NpuC (SEQ ID NO: 18) were prepared by restriction enzyme-free cloning.
[0158] Synthetic genes for SopE were purchased and introduced into ABCA4-1150-CfaN, ABCA4-1150-CfaC, ABCA4-1140-CfaN, ABCA4-1140-CfaC, EGFP-71-CfaN, and EGFP-71-CfaC via restriction enzyme-free cloning, yielding constructs ABCA4-1150-CfaN-SopE (SEQ ID NO: 19), ABCA4-1150-CfaC-SopE (SEQ ID NO: 20), ABCA4-1140-CfaN-SopE (SEQ ID NO: 21), ABCA4-1140-CfaC-SopE (SEQ ID NO: 22), EGFP-71-CfaN-SopE (SEQ ID NO: 23), and EGFP-71-CfaC-SopE (SEQ ID NO: 24). The gene for DD1 was prepared by overlap extension PCR and introduced into EGFP-71-CfaN and EGFP-71-CfaC by restriction enzyme-free cloning to obtain the constructs EGFP-71-CfaN-DD1 (SEQ ID NO: 25) and EGFP-71-CfaC-DD1 (SEQ ID NO: 26). The identity of all recombinant plasmids was confirmed by sequencing, and the sequences of the corresponding proteins are reported in Table 1. Constructs containing degron were similarly prepared by cloning different elements, including the split protein gene, split intein, and degron, into expression plasmids.
[0159] AAV production AAVs encoding the N fragment of the present invention, which encodes either an EGFP fragment or an ABCA4 fragment containing a CfaN intein, were prepared using a standard triple transfection strategy. An rAAV8 vector containing a wild-type AAV2 ITR was prepared by polyethyleneimine (PEI)-mediated simultaneous transfection of the plasmid pAAV-Rep2-Cap8, pHelper, and AAV-transgene plasmid into HEK293 cells.
[0160] Cells and culture medium were collected 48 hours after transfection. After treatment with 0.1% Triton, the virus was released from the cells, and AAV particles were obtained from the culture medium by PEG precipitation. Purification of the crude lysate and removal of empty capsids were performed by iodixanol gradient according to the method of Zolotukhin et al. [Zolotukhin S, Byrne BJ, Mason E, Zolotukhin I, Potter M, Chestnut K, Summerford C, Samulski, RJ, Mucyczka N. (1999) Recombinant adeno-associated virus purification using novel methods improves infectious titer and yield. Gene Ther; 6(6):973-85.].
[0161] The purified batch was determined by quantitative polymerase chain reaction (qPCR) using Vivaspin 20 Centrifugal Concentrator 100,000 MWCO (Sartorius, catalog number VS0641), resulting in a final concentration of 1 × 10⁻¹⁶. 13 The solution was concentrated to vg / ml and then formulated. The formulation buffer used was characterized by BSS (Alcon) + Pluronic.
[0162] After concentration, the virus batch is placed in an Acrodisc® syringe filter (Acrodisc PP, PES, 0.2 μM, 1 cm). 2The sample was filtered using ) and stored at -80°C. Viral titer was determined in terms of genome copy number per milliliter by qPCR using ITR-specific PCR primers. Furthermore, capsid titer was obtained using AAV8 ELISA (Progen) according to the manufacturer's instructions. 1 × 10⁶ titers were present in the final product. 10 The protein corresponding to vg was isolated by 10% SDS-PAGE using the Pierce silver staining kit (catalog number 24612) for protein visualization. VP1, VP2, and VP3 proteins were detectable with the correct 1:1:10 stoichiometry, indicating a purity of over 95% in the AAV preparation.
[0163] Transfection of intein plasmids in HEK293 cells: HEK293T cells were maintained in DMEM containing 10% FBS and antibiotics at 37°C in a 6% CO2 atmosphere. Lipofectamine 2000 and 1.25 μg of each plasmid were used to co-transfect cells in a 6-well plate format at approximately 80% confluence. In experiments using plasmids encoding full-length genes, scrambled plasmids were co-transfected with full-length plasmids so that the amount of transfected DNA was the same when using two intein plasmids. Cells were harvested 48 hours after transfection, and EGFP was analyzed by Western blotting.
[0164] Transfection of degron-inteiin plasmid in HEK293 cells: Cells were co-transfected in a 6-well plate format using Lipofectamine 2000 and 1.25 μg of each plasmid, at approximately 80% confluence. Cells transfected with the intein-degron plasmid were collected 48 hours after transfection and analyzed by Western blotting. In the time-course experiment, cells were collected 24 and 48 hours after transfection and analyzed by Western blotting.
[0165] Proteosome inhibitor experiments were conducted using the MG-132 inhibitor. Cells were co-transfected, and DMSO or MG-132 (dissolved in DMSO) was added 24 hours after transfection to a final concentration of 50□M. Cells were collected at 30 minutes, 3 hours, 6 hours, and 24 hours and analyzed by Western blotting.
[0166] Western blot analysis: EGFP-transfected cells (HEK293 cells and WERI-RB1 cells) were dissolved in RIPA buffer supplemented with a protease inhibitor and 1 mM phenylmethylsulfonyl. After lysis, the samples were quantified using a BCA protein assay kit. Samples with 10 μg of total protein were denatured in 1 × Laemmli sample buffer at 95°C for 10 minutes. The lysates were separated on a 12% Bis-tris SDS-PAGE gel at 165 V for 40 minutes. Antibodies used for immunoblotting were anti-6 × His tag for EGFP detection and anti-β-tubulin as a loading control. The EGFP band detected by Western blotting was quantified using a LI-COR Odyssey Infrared Imager.
[0167] Cells transfected with ABCA4 (HEK293) were dissolved in a solution of dodecyl maltoside (D310) and cholesteryl hemisuccinate (CH210) in PBS (1:1) supplemented with a protease inhibitor and 1 mM phenylmethylsulfonyl. After lysis, the ABCA4 sample was quantified using a BCA protein assay kit. Samples with a total protein content of 25 μg were denatured at 37°C for 15 minutes in 1× Laemmli sample buffer containing 2.5 mg / ml asolectin. The lysates were separated on a 3%-8% Tris acetate SDS-PAGE gel at 150 V for 1.5 hours. Antibodies used for immunoblotting were either anti-flag tag or anti-ABCA4 to detect the ABCA4 protein, and anti-β-tubulin as a loading control. Quantification of the ABCA4 band detected by Western blotting was performed using a LI-COR Odyssey Infrared Imager.
[0168] Fluorescence-activated cell sorting measurement: HEK293T cells were co-transfected with 2 μg of each EGFP-intane plasmid. HEK293T cells were cultured as a negative control, but were not transfected. In experiments using plasmids encoding the full-length gene, scrambled plasmids were co-transfected with full-length plasmids so that the amount of transfected DNA was the same when using two intein plasmids. Cells were analyzed by flow cytometry 48 hours after transfection. Transfected and untransfected cells were resuspended in FACS analysis buffer (PBS, 2% FSA, 2 mM EDTA), respectively. The percentage of EGFP+ cells was assessed by comparing different transfected and untransfected cells using Kaluza flow cytometry software on a Beckman Coulter Gallios flow cytometer. The number of dead cells was detected using Dapi staining.
[0169] Analysis by mass spectrometry: The gel bands were washed with ammonium bicarbonate (50 mM NH4HCO3) and acetonitrile (ACN). The samples were reduced with 20 mM DTT at 60°C for 60 minutes, and then alkylated with 55 mM iodoacetamide at 25°C for 30 minutes in the dark. Subsequently, the samples were double-digested with trypsin (sequence-grade modified trypsin) at 37°C for 2 hours and overnight. Finally, the resulting peptide mixture was extracted from the gel matrix with 5% formic acid (FA) in 50% ACN and 100% ACN, and dried in a SpeedVac vacuum system. The resulting peptide mixture was cleaned up on a C18 tip (PolyLC Inc.) according to the manufacturer's protocol. Finally, the cleaned-up peptide solution was dried. The trypsin-digested mixture was resuspended in 1% FA solution, and aliquots were injected for each sample for chromatographic separation. Peptides were captured on a Symmetry C18 capture column (5 μm, 180 μm × 20 mm; Waters) and separated using a C18 reversed-phase capillary column (ACQUITY UPLC M-class peptide BEH column; 130 Å, 1.7 μm, 75 μm × 250 mm, Waters). The gradient used for peptide elution was 1% → 40% B for 25 minutes, followed by 40% → 60% for 5 minutes at a flow rate of 250 nL / min (A: 0.1% FA; B: 100% ACN, 0.1% FA). The eluted peptides were subjected to electrospray ionization at an applied voltage of 2000 V using an emitter needle (PicoTip®, New Objective). Peptide masses (m / z 300-1700) were analyzed in data-dependent mode, and full-scan MS was acquired in Orbitrap at 400 m / z with a resolution of 60,000 FWHM. The 15th most abundant peptide (minimum intensity 500 counts) was selected from each MS scan, and then fragmented using a linear ion trap with CID (38% normalized collision energy) using helium as the collision gas. Database searches were performed using Thermo Proteome Discover with the Sequest HT search engine.
[0170] in vivo experiment The experiment was conducted using 4-week-old wild-type mice of the 129Sv strain. The test sample was injected subretinically into the mice.
[0171] Subretinal injection was performed using a 5 ml Hamilton syringe with a 33G needle after pupil dilation. EGFP expression in the fundus was evaluated 4 weeks post-injection using a Canon UVI retinal camera connected to a digital imaging system. Retinal structure and the outer granular layer (ONL) were quantitatively evaluated by SD-OCT (Optical Coherence Tomography) using a Spectralis imaging system (Heidelberg Engineering Inc.). After imaging, the animals were euthanized and enucleated for further analysis by immunohistochemistry and Western blotting. Furthermore, experiments were also conducted in which the vector was injected into the inner ear of 6-week-old C57BL / 6 mice.
[0172] Results and Discussion Recently, several novel inteins have been designed based on consensus design (Stevens et al., 2016; Stevens, Sekar, Gramespacher, Cowburn, & Muir, 2018), and have been shown to possess superior properties compared to naturally occurring inteins. One of these inteins, called Cfa, was generated by consensus design from DnaE intein alignment and has been shown to have superior properties to some of the best DnaE family inteins, such as Npu. Cfa has been reported to exhibit faster expression rates, higher expression levels, and high tolerance to extreme conditions such as high temperatures and high concentrations of denaturants. Interestingly, Cfa variants with a degree of homology of over 90% also exhibit similar properties. Applying the same consensus design strategy to the TerL-AceL intein family resulted in a consensus sequence called Cat, which also showed improved properties compared to the rest of the TerL-AceL family.
[0173] To demonstrate the usefulness of consensus inteins for protein reconstitution via protein transsplicing in gene therapy applications, direct comparisons were performed. These comparisons were conducted by transfecting cells with purified plasmids.
[0174] The initial experiment used EGFP as the reporter gene. In short, EGFP was split at position 71, and the N-terminal fragment (residues 1-70, EGFPN) was recombinately fused to the N-intane, while the C-terminal fragment (residues 71-239, EGFPC) was recombinately fused to the C-intane. To compare the reconstitution efficiency of the Cfa consensus sequence with Npu, four constructs were generated: EGFPN-CfaN, EGFPN-NpuN, CfaC-EGFPC, and NpuC-EGFPC. Cultured HEK293 cells were co-transfected with equimolar plasmids encoding the N-terminal and C-terminal fragments, and splicing efficiency was monitored by fluorescence microscopy, flow cytometry, and Western blotting.
[0175] Our results (see Figures 2, 3, and 4) show that consensus Cfa intein provides a 2.5-fold increase in EGFP reconstitution compared to Npu intein, and that inteins obtained by the consensus design provide a higher yield of target protein upon co-transfection. It was known that Cfa divides faster than Npu in vitro, and that its N-terminal fragment is better expressed when transformed or transfected into cells individually, but it was important to note that it had not been shown whether there was any advantage to co-transfecting both fragments into the same cells. Through these experiments, we demonstrated that when both IntN and IntC fragments are expressed in the same cell, using Cfa provides a clear advantage in target protein reconstitution yield compared to known ultrafast split inteins such as Npu. Furthermore, for in vivo delivery, the N and C fragments of EGFP, and the genes encoding the N-intine and C-intine, were incorporated into recombinant AAV. The results demonstrated that EGFP is reconstituted in several organs, including the retina and inner ear, by Cfa intein-mediated protein splicing (Figure 23).
[0176] To confirm that the positive features of this Cfa consensus sequence are also observed in other proteins, the inventors repeated the experiment using the ABCA4 protein. ABCA4 is a large protein that mutates in Stargardt disease, and its reconstruction has been proposed as a viable strategy for treating the disease. Currently, several approaches based on AAV gene therapy are being explored to reconstruct the ABCA4 protein. Due to its large size, ABCA4 cannot be encapsulated in a single AAV, and therefore, various strategies have been proposed to deliver ABCA4 in two fragments and reconstruct it in target cells. To confirm that manipulated consensus sequences such as Cfa offer advantages over naturally derived inteins such as Npu, the inventors split ABCA4 at position 1150 and cloned the N-intine and C-intine into the N-terminal and C-terminal fragments of the protein, respectively.
[0177] ABCA4 N (1~1149)-IntN and IntC-ABCA4 C (1150~2273) and Cfa or Npu inteins were cloned into expression plasmids and co-transfected into HEK293 cells. The cells were lysed, and the protein reconstitution yield was determined by Western blotting. Full-length control ABCA4 was also co-transfected.
[0178] As observed with EGFP, we observed a higher reconstitution yield with Cfa than with Npu (see Figure 5).
[0179] Furthermore, a Cfa variant called CfaCmut (Stevens et al., 2017), which contains a specific mutation in the CfaC fragment, has recently been reported, adding value to the advantages of Cfa. In particular, the CfaCmut variant has been shown to efficiently splice the protein fragment even without the original Phe residue at the +2 position of the C-extinne. The +2 position defined above refers to the position +2 relative to the last amino acid of IntC. When ABCA4 is split at position 1150, Phe remains at the +2 position of the C-extinne, which, like Cfa, is an optimal residue for Npu.
[0180] Although ABCA4 can be reconstructed by splitting at position 1150, this may not be the optimal location for therapeutic purposes. To test the ability to split and reconstruct ABCA4 at a site other than 1150, the inventors used CfaCmut, not limited to a position that provides Phe at position +2 relative to intein.
[0181] Several ABCA4 constructs were cloned by splitting the ABCA4 protein at a Phe-free site at the +2 position of the splitting site. Two additional sites were tested: position 1140, corresponding to the Ser at the +2 position, and position 1188, corresponding to the Pro at the +2 position. The constructs were cloned and tested as described above. Specifically, splicing of ABCA4 at position 1140 was evaluated using the ABCA4(1~1139)-IntN and IntC-ABCA4(1140~2273) constructs. Splitting of ABCA4 at position 1188 was evaluated using ABCA4(1~1187)-IntN and IntC-ABCA4(1188~2273). Sites were selected based on the topological structure of ABCA4, taking into account the presence of folded domains. Sites outside the well-folded intracellular domain of ABCA4 were selected. When different sites were tested, it was observed that Cfa and Cfa variants all provided higher reconstruction yields than Npu, and that the efficiency differed depending on the splitting site (Figures 5, 6, and 7). The inventors also demonstrated that this approach can reconstruct ABCA4 splitting at positions 1179 and 1177. Similar constructs to those defined above were cloned for 1140 and 1188 and used to test the efficiency of splicing at positions 1179 and 1177, with the results at position 1179 shown in Figure 21.
[0182] The inventors also demonstrated that ABCA4 can be reconstituted from the split fragments using other efficient inteins such as gp41. ABCA4 constructs split at positions 1185 and 1095 were designed, and fusions corresponding to the N-terminal and C-terminal inteins were cloned and their splicing activity analyzed as described above. Western blotting analysis of the PTS reaction at split position 1096 and the gp41 intein is shown in Figure 22.
[0183] Table 7 below shows the sequences of the ABCA4 fragments obtained when the protein is split at different sites (1177, 1179, 1096, and 1185), and the sequences of the corresponding ABCA4-inteinin N or intein C-ABCA4 fusions.
[0184] The use of split inteins in gene therapy is subject to significant limitations beyond reconstitution yield, including the presence of unreacted starting materials and / or inteins. To prevent the accumulation of undesirable starting materials and eliminate cleaved intein fragments, the inventors added degron to a construct having the general structure shown in Figure 8.
[0185] The inventors have identified several degrons that can be used in combination with intein to develop gene therapy approaches for ABCA4, which can also be repurposed for diseases caused by mutations on other large genes, and for protein reconstitution to treat those diseases, such as the reconstitution of the CRISPR / Cas9 system.
[0186] To identify appropriate degron-intine combinations, EGFP-Int constructs containing selected degrons were cloned. Degrons were cloned at the C-terminus of the N-intine and the N-terminus of the C-intine. For detection purposes, a His6 tag was included as a linker between the two elements. As proof of principle, a degron consisting of the SopE destabilization domain (1-100) and the peptide sequence DD1 (see Table 3) was used.
[0187] The results show that including degron effectively removes any detectable amount of starting material and intein (see Figures 9 and 14). Importantly, and unexpectedly, it was demonstrated that including both degrons simultaneously removes any unreacted starting material without reducing the yield of the protein splicing reaction.
[0188] Interestingly, when the inventors applied the degron strategy to the large gene ABCA4, they observed similar results, but with increased levels of spliced ABCA4 product, demonstrating an unexpected synergistic effect between consensus intein and degron (see Figure 10).
[0189] The effects of proteosome inhibitors on these results were also investigated (Figure 11). The inventors showed that proteosome inhibitors reduced the effect of degron, thus confirming that SopE degradation is mediated by proteasomes. Interestingly, the inventors confirmed that by combining intein and degron, they were not only able to eliminate the presence of the starting material but also maintain and even enhance the reconstitution level of ABCA4 (Figure 10).
[0190] The inventors also conducted experiments to investigate the effects of combining the consensus intein Cfa with degron, comparing the results with those without degron or with other inteins. Interestingly, the inventors observed that using Cfa with degron resulted in a significant decrease in the starting material level and an increase in the amount of product (Figure 12). The inventors also confirmed that this effect was observed at positions 1140 (Figures 13 and 21) and 1179 (Figure 21) of ABCA4. Importantly, the inventors showed that the observed effects were maintained by combining several different degrons with Cfa (and Cfamut) inteins that shared certain characteristics (see Example 3).
[0191] Based on the results obtained in ABCA4, other diseases, genes, and proteins to which this approach may be applicable are listed in Table 1.
[0192] Example 2. Three-piece ligation strategy using orthogonal split-intane pairs Materials and methods: Transfection of a 3-piece plasmid in HEK293 cells: Cells were co-transfected in a 6-well plate format with Lipofectamine 2000 and 1.25 μg of each plasmid at approximately 80% confluence. Cells transfected with 3-piece plasmids were harvested 48 hours after transfection and analyzed by Western blotting. As a control, cells were co-transfected with two plasmids (POIN-CfaN+CfaC-POIM-CatN, CfaC-POIM-CatN+CatC-POIC, and POIN-CfaN+CatC-POIC).
[0193] Western blot analysis: Cells (HEK293) transfected using a 3-piece technique were dissolved in RIPA buffer supplemented with a protease inhibitor and 1 mM phenylmethylsulfonyl. After lysis, the samples were quantified using a BCA protein assay kit. Samples with a total protein content of 10 μg were denatured in 1 × Laemmli sample buffer at 95°C for 10 minutes. The lysates were separated on a 12% Bis-tris SDS-PAGE gel at 165 V for 40 minutes. Antibodies used for immunoblotting were an anti-flag tag for detecting POIs and anti-β-tubulin as a loading control.
[0194] result The inventors have shown that combining an ultrafast consensus intein with degron is a viable strategy for reconstituting large proteins while reducing the levels of starting material and cleaved intein. This strategy, when combined with a suitable delivery vector such as those disclosed herein, may be used in gene therapy to reconstitute large proteins using gene replacement strategies and to generate therapeutic agents in vivo. However, for even larger proteins, this strategy may be insufficient. For example, for proteins with coding regions larger than 7kb–8kb, splitting the coding gene into two fragments is not sufficient to obtain fragments that can be encapsulated in an AAV vector. For such proteins, it may be necessary to split the coding gene into three pieces. Assembling a protein from three separate fragments using inteins requires an orthogonal intein pair. To achieve maximum reconstitution yield, a highly efficient and orthogonal intein is required. The inventors decided to use Cfa intein in combination with Cat intein. These two inteins were obtained through consensus design and share several common characteristics, including high expression yield, heat resistance, resistance to chaotropic agents, and a fast splicing rate. The inventors designed a series of constructs: POIN-CfaN, CfaC-POIM-CatN, and CatC-POIC, according to the structures described in Figure 16. Alternatively, the positions of the Cfa intein and Cat intein can be swapped to generate constructs with the following structures: POIN-CatN, CatC-POIM-CfaN, and CfaC-POIC.
[0195] When the inventors tested their orthogonality, they demonstrated that both inteins were indeed orthogonal and did not react with each other; that is, the N fragment of either the Cfa or Cat intein reacted only with its respective homologous pair. When the three fragments were simultaneously transfected in equimolar amounts, the main product detected was that which resulted from the assembly of the three fragments into the desired full-length product. Importantly, the inventors detected only low levels of unreacted material, and also did not detect any intermediate products (i.e., products resulting from the reaction of only two fragments). This result indicates that using highly active pairs of orthogonal consensus inteins is a suitable strategy for assembling large proteins from three individual fragments.
[0196] Example 3. To carry out the present invention, several degrons were tested. For application to AAV gene therapy, degrons smaller than 75 amino acids are desirable to minimize the size of the resulting transgene; therefore, degrons were selected from the literature based on their size. In some cases, designs were created by combining known N-terminal or C-terminal degrons.
[0197] Degrons were cloned into the C-terminus of the N-fragment and the N-terminus of the C-fragment, both of which were fragments intended for the reconstruction of full-length EGFP protein as representative examples of soluble cytoplasmic proteins. Figure 9 shows an overview of the results, demonstrating that including one or two degrons (one for each EGFP fragment) can eliminate undesirable starting material fragments without negatively impacting splicing yield. Similar results were obtained with different degrons (Figure 14), demonstrating the versatility of the approach.
[0198] Degrons were cloned into the C-terminus of the N-fragment and the N-terminus of the C-fragment, both fragments intended to reconstruct the full-length ABCA4 protein at different splitting locations, including representative examples 1140, 1150, and 1179. Table 6 below shows all the degrons tested, including their size, sequence, origin, and the ligases involved in ubiquitination and eventual degradation.
[0199] [Table 6]
[0200] Degron was tested to determine whether it could induce degradation of the starting material and the cleaved intein, and to examine its effect on the yield of the protein trans-splicing reaction, i.e., the ability to reconstitute the full-length ABCA4 protein. Cells were co-transfected with N-terminal and C-terminal constructs containing the different degrons shown above. Cells were lysed and analyzed by Western blotting to detect the presence of PTS products with and without degron. Figure 18 shows the results using the same degron with the N and C fragments. The gel in the left panel (Figure 18C) shows the PTS products (ABCA4) and starting materials (ABCAN-IntN-DD and DD-IntC-ABCAC) for each reaction. PTS corresponds to the results without degron, and the labels in the other lanes indicate which degron was used in each case. The right panel (Figure 18C) provides quantitative analysis of the amount of PTS product in each case. Surprisingly, some degrons show higher yields than the constructs without degron. Interestingly, all degrons functioned, and all of them could reconstitute at least 25% of the product obtained without the degrons. Some degrons could reconstitute more than 50% of the level obtained without the degrons, and some degrons ultimately yielded 100% (or more) reconstitution. In particular, the following degrons yielded 100% (or more) reconstitution: DD1, V12, M4, L2, L9, DD3, and SopE100.
[0201] Furthermore, the effect of deglon on cleaved intein was also investigated (Figure 19). Intein was detected by analyzing the PTS reaction with different deglons using Western blotting (WB) with a high-efficiency acrylamide gel. As can be seen, some deglons reduce the amount of cleaved intein compared to others (e.g., SpoE and PEST), while others, such as L2 (SEQ ID NO: 52), L9 (SEQ ID NO: 54), M2 (SEQ ID NO: 47), M4 (SEQ ID NO: 35), or DD3 (SEQ ID NO: 45), completely eliminate the intein band.
[0202] Furthermore, the inventors also conducted experiments to investigate the effects of using different degrons at the N-terminus or C-terminus. As an example, Figure 20 shows Western blot analysis in which the inventors immobilized degron on the N-fragment and varied the degron on the C-fragment. The results shown in Figure 20 demonstrate that by combining several preferred degrons, it is possible to achieve further reduction of starting materials without reducing the yield of the desired splicing product, thereby further improving the results. These combinations are considered useful in certain applications, particularly when it is important to completely eliminate all starting materials.
[0203] Based on the results obtained, the inventors were able to classify degron into three categories: Degron enables the reconstitution of target proteins with at least 25% of the yield obtained without Degron. Degron enables the reconstitution of target proteins with a yield of over 50% compared to when Degron is not present. Degron enables the reconstitution of target proteins with yields equal to or better than those achieved without Degron.
[0204] When using a combination of CfaN and CfaC or CfaCmut intein, several sites in ABCA4 were shown to be suitable for use in the invention described herein. Specifically, these are the sites between the C-terminus of nucleotide-binding domain 1 (NBD1) and the seventh transmembrane domain (TMD), including positions 1140, 1150, 1177, 1179, and 1188. The yield of ABCA4 reconstitution was highest at positions 1140, 1150, and 1179, and although position 1188 functioned with very low efficiency, it still resulted in a higher level of ABCA4 reconstitution than that obtained using non-consensus Npu intein.
[0205] The combination of CfaN and CfaCmut inteins provides increased ambiguity. In fact, we have now revealed for the first time that this increased ambiguity is observed in transmembrane proteins such as ABCA4, and that high reconstitution yields are obtained when using CfaN-CfaC mutant pairs at different splicing sites. Furthermore, such splicing reactions are shown to be highly effective, as less than 10% of each starting material is observed when using these inteins (CfaN and CfaCmut inteins) (estimated as shown in Figure 12). This was unexpected from previously reported results using Npu intein and ABCA4 (Auricchio et al. 2019), which suggested that a significant portion of the starting material remains unreacted and that protein-transsplicing is inherently difficult in such membrane proteins.
[0206] Based on the results reported herein, the inventors conclude that when using mutant pairs of consensus inteins CfaN and CfaC, any site having the following characteristics will function efficiently: (1) a site outside the enzyme domain or transmembrane domain, and (2) a site where the +1 position of the splitting site (as described above) corresponds to a nucleophilic residue such as Cys or Ser, and the +2 position is any site other than Pro.
[0207] [Table 7]
[0208] item 1. A composition, The polynucleotide comprises a first polynucleotide encoding a polynucleotide containing a split-intane N fragment directly linked to the N-terminal fragment of the protein to be reconstituted via a peptide bond, optionally by a peptide linker, and a second polynucleotide encoding a polynucleotide containing a split-intane C fragment directly linked to the C-terminal fragment of the protein to be reconstituted via a peptide bond, optionally by a peptide linker. Both components of the composition may be filled together in a single formulation, or they may be filled separately in different formulations. The first and second polynucleotides each encode the N-terminal and C-terminal fragments of the reconstituted protein, respectively, such that when both fragments are combined, the N-terminal fragment of the protein is ligated to the C-terminal fragment of the protein to form the entire protein. Each polynucleotide must encode a split intein such that, upon translation into a protein, its N-terminal and C-terminal sequences become separate fragments that can be reassembled or rearranged non-covalently into a functional intein for trans-splicing reactions. A composition in which the reconstituted protein has a value greater than 25 kDa. 2. The composition according to item 2, wherein a first polynucleotide encodes a split-intene N fragment directly linked to degron via a peptide bond, and degron is linked to the intein N fragment via the C-terminus of intein, with or without a linker between intein N fragment and degron, and the N-terminus of the split-intene N fragment is directly linked to the N-terminal fragment of the reconstituted protein via a peptide bond; and a second polynucleotide encodes a split-intene C fragment directly linked to degron via a peptide bond, and degron is linked to the intein C fragment via the N-terminus of intein, with or without a linker between intein C fragment and degron, and the C-terminus of the split-intene C fragment is directly linked to the C-terminal fragment of the reconstituted protein via a peptide bond. 3. The composition according to claim 1 or 2, wherein the first polynucleotide encodes either the CfaN intein of SEQ ID NO: 27 or a variant thereof, and the second polynucleotide encodes either the CfaC intein of SEQ ID NO: 28 or a variant thereof, and the variant is understood to be a split intein N fragment or C fragment of SEQ ID NO: 27 or SEQ ID NO: 28 having at least 90% sequence identity with any of these sequences, such as CfaCmut (SEQ ID NO: 29). 4. The composition according to claim 2 or 3, wherein degron is selected from the group consisting of SEQ ID NOs: 40 to 59, SEQ ID NOs: 34, SEQ ID NOs: 35 and SEQ ID NOs: 37. 5. The composition according to claim 2, wherein split-intine is as defined in claim 3, and degron is selected from the group consisting of SEQ ID NOs: 42, 43, 49, 50, 51, 54, and 34. 6. The composition according to any one of claims 1 to 5, wherein both polynucleotides are contained within a vector that enables the propagation of the polynucleotides in a suitable host cell. 7. The composition described in item 6, which is an adeno-associated virus (AAV). 8. The composition according to item 7, wherein the vector is AAV of serotype 1, 2, 3, 4, 5, 6, 7, 8, or 9. 9. The composition according to any one of items 1 to 8, wherein the gene encoding the whole protein is selected from the group consisting of any of the proteins listed in Table 1. 10. The composition according to any one of items 1 to 9, which is used for treatment. 11. A method for expressing a target gene in a cell, comprising: (i) contacting the cell with (a) the first polynucleotide defined in item 1, and (b) the second polynucleotide defined in item 1; (ii) expressing the first polynucleotide and the second polynucleotide so that the first fusion protein and the second fusion protein are produced; (iii) contacting the first protein and the second protein so that the split intein N fragment binds to the split intein C fragment to form an intein intermediate, and the intein intermediate reacts to covalently bond the C-terminus of the target first polypeptide to the N-terminus of the target second polypeptide. A method comprising the above steps. 12. A method for expressing a target gene in a cell, comprising: (i) contacting the cell with (a) the first polynucleotide defined in item 2, and (b) the second polynucleotide defined in item 2; (ii) expressing the first polynucleotide and the second polynucleotide so that the first fusion protein and the second fusion protein are produced; (iii) contacting the first protein and the second protein so that the split intein N fragment binds to the split intein C fragment to form an intein intermediate, and the intein intermediate reacts to covalently bond the C-terminus of the target first polypeptide to the N-terminus of the target second polypeptide. A method comprising the above steps. 13. The composition according to claim 11 or 12, wherein the first polynucleotide encodes either the CfaN intein of SEQ ID NO: 27 or a variant thereof, and the second polynucleotide encodes either the CfaC intein of SEQ ID NO: 28 or a variant thereof, and the variant is understood to be a split intein N fragment or C fragment of SEQ ID NO: 27 or SEQ ID NO: 28 having at least 90% sequence identity with any of these sequences, such as CfaCmut (SEQ ID NO: 29). 14. The method according to claim 11 or 12, wherein split-intane is as defined in claim 3, and degron is selected from the group consisting of SEQ ID NOs: 42, 43, 49, 50, 51, 54, and 34. 15. The method described in any of sections 11-14, wherein both polynucleotides are contained within the adeno-associated virus (AAV).
Claims
1. A combination of a composition or formulation, a. A first polynucleotide encoding a polypeptide containing a split-intane N fragment, wherein the split-intane N fragment is directly linked to the N-terminal fragment of the reconstituted protein via a peptide bond or a peptide linker, and is CfaN of SEQ ID NO: 27, or any functionally equivalent variant thereof, wherein the variant has at least 98% sequence identity with SEQ ID NO: 27, and the ability of the split-intane N fragment to perform protein trans-splicing reactions by binding to the split-intane C fragment is maintained or improved, and the first polynucleotide and, b. A second polynucleotide encoding a polypeptide containing a split-intane C fragment, wherein the split-intane C fragment is CfaC of SEQ ID NO: 28, or CfaCmut of SEQ ID NO: 29, or any functionally equivalent variant thereof, wherein the variant has at least 98% sequence identity with SEQ ID NO: 28 or 29, and the ability of the split-intane C fragment to perform protein trans-splicing reactions by binding to a split-intane N fragment is maintained or improved, and the second polynucleotide and, Includes, Both of the aforementioned polynucleotides may be filled together in a single formulation, or they may be filled separately in different formulations. The first polynucleotide and the second polynucleotide respectively encode the N-terminal and C-terminal fragments of the reconstructed protein, such that when both fragments are combined, the N-terminal fragment of the protein is linked to the C-terminal fragment of the protein to generate the entire protein. The protein to be reconstituted is greater than 25 kDa, The protein to be reconstituted is ABCA4, and The first polynucleotide codes for positions 1 to 1139 of the N-terminal fragment of the ABCA4 protein, and the second polynucleotide codes for positions 1140 to 2273 of the C-terminal fragment of the ABCA4 protein; or The first polynucleotide codes for positions 1 to 1178 of the N-terminal fragment of the ABCA4 protein, and the second polynucleotide codes for positions 1179 to 2273 of the C-terminal fragment of the ABCA4 protein; composition.
2. The composition or combination according to claim 1, wherein the split-intane N fragment is CfaN of SEQ ID NO: 27, directly linked to the N-terminal fragment of the reconstituted protein via a peptide bond or by a peptide linker, and the split-intane C fragment is CfaCmut of SEQ ID NO: 29, directly linked to the C-terminal fragment of the reconstituted protein via a peptide bond or by a peptide linker.
3. The composition or combination according to claim 1, wherein the first polynucleotide and the second polynucleotide are selected from the following combinations: a. Sequence IDs 3 and 4; b. Sequence IDs 68 and 69; c. Sequence IDs 74 and 75; and d. Sequence IDs 78 and 79.
4. The first polynucleotide encoding the split-intene N fragment is further directly linked to a degron via a peptide bond, the degron is linked to the split-intene N fragment via its C-terminus with or without a linker between the split-intene N fragment and the degron, the N-terminus of the split-intene N fragment is directly linked via a peptide bond to the N-terminal fragment of the reconstituted protein, and The second polynucleotide encoding the split-intane C fragment is further directly linked to a degron via a peptide bond, the degron is linked to the split-intane C fragment via its N-terminus, with or without a linker between the split-intane C fragment and the degron, and the C-terminus of the split-intane C fragment is directly linked via a peptide bond to the C-terminal fragment of the reconstituted protein. The composition or combination according to any one of claims 1 to 3.
5. The composition or combination according to claim 4, wherein the degron is selected from the group consisting of SEQ ID NOs: 40-59, 34, 35, and 37.
6. The composition or combination according to claim 5, wherein the degron is selected from the group consisting of Sequence ID Nos. 42, 43, 49, 50, 51, 54, and 34.
7. The composition or combination according to any one of claims 1 to 6, wherein both polynucleotides are contained within a vector that enables the propagation of the polynucleotides in a suitable host cell.
8. The composition or combination according to claim 7, wherein the vector is adeno-associated virus (AAV).
9. The composition or combination according to claim 8, wherein the vector is an AAV of serotype 1, 2, 3, 4, 5, 6, 7, 8, or 9.
10. A pharmaceutical composition for the treatment of Stargardt disease, comprising the composition or combination described in any one of claims 1 to 9.
11. An in vitro method for expressing a target gene within a cell, (i) the cells Contacting the first polynucleotide and the second polynucleotide as defined in any one of claims 1 to 3, (ii) Expressing the first polynucleotide and the second polynucleotide as defined in any one of claims 1 to 3 so as to produce the first fusion protein and the second fusion protein, (iii) The split intein N fragment binds to the split intein C fragment to form an intein intermediate, and the first fusion protein and the second fusion protein are brought into contact such that the intein intermediate reacts and covalently bonds the C-terminus of the N-terminal fragment of the reconstructed protein to the N-terminus of the reconstructed C-terminal fragment. Methods that include...