Fusion protein for purifying nucleic acids
Patent Information
- Application Number
- JP2026512093
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-16
- Filing Date
- 2024-08-23
- Publication Date
- 2026-09-08
AI Technical Summary
【0018】 本明細書には、一本鎖核酸を精製する方法が提供され、これは、以下:(i)一本鎖核酸と少なくとも1つの汚染物質を含む組成物を、請求項1~28のいずれか一項に記載の融合タンパク質と接触させるステップであって、融合タンパク質が一本鎖核酸と結合して、複合体を形成するステップ;(ii)複合体を含む組成物に第1の環境因子を添加し、それによって複合体のサイズを増加させるステップ;(iii)少なくとも1つの汚染物質から複合体を分離するステップ;並びに(iv)複合体を第2の環境因子と接触させることにより、融合タンパク質から一本鎖核酸を分離し、それによって一本鎖核酸を含む生成物を形成するステップを含む。実施形態では、一本鎖核酸は、ssRNAを含む。実施形態では、複合体は、サイズに基づいて、少なくとも1つの汚染物質から分離される。実施形態では、サイズに基づく分離は、接線流ろ過、膜クロマトグラフィー、分析用超遠心分離、高速液体クロマトグラフィー、全量ろ過、音波分離、遠心分離、向流遠心分離、及び高速タンパク質液体クロマトグラフィーのいずれか1つから選択される方法を用いて実施される。実施形態では、サイズに基づく分離は、遠心分離を含む方法を用いて行われる。実施形態では、第1の環境因子は、以下:(a)温度、pH、塩濃度、精製マトリックスの濃度、ウイルス粒子の濃度、若しくは圧力のうちの1つ以上の変化;(b)1つ若しくは複数の界面活性剤、補因子、ビタミン、分子密集剤、還元剤、酸化剤、酵素、若しくは変性剤の添加;又は(c)電磁波若しくは音波の適用のうちの1つ若しくは複数を含む。実施形態では、第2の環境因子は、以下:(a)温度、pH、塩濃度、精製マトリックスの濃度、ウイルス粒子の濃度、若しくは圧力のいずれか1つ以上の変化;(b)1若しくは複数の界面活性剤、補因子、ビタミン、分子密集剤、還元剤、酸化剤、酵素、若しくは変性剤の添加;又は(c)電磁波若しくは音波の適用のうちの1つ若しくは複数を含む。実施形態では、少なくとも1つの汚染物質は、以下:溶媒、タンパク質、ペプチド、炭水化物、核酸、ウイルス、細胞(例えば、細菌、酵母、若しくは哺乳動物細胞)、炭水化物、脂質、又はリポ多糖から選択される。実施形態では、少なくとも1つの汚染物質は、二本鎖核酸である。実施形態では、二本鎖核酸は、dsRNAである。実施形態では、塩は、約0.5M~約3Mを含む濃度で添加される。実施形態では、塩は、約1.2M~約1.7Mの濃度で添加される。実施形態では、生成物は、ステップ(iv)の生成物中の総核酸含有量に基づいて、約70%から約100%の一本鎖核酸を含む。実施形態では、一本鎖核酸は、ssRNAを含む。実施形態では、生成物は、生成物中の総核酸含有量に基づいて、約10%以下の少なくとも1つの汚染物質を含有する。実施形態では、少なくとも1つの汚染物質は、dsRNAを含む。実施形態では、精製方法は、ステップ(i)の組成物中の汚染物質の量と比較して、少なくとも約1-logの汚染物質の除去を含む。実施形態では、精製方法は、ステップ(i)の組成物中の汚染物質の量と比較して、1~10-logの汚染物質の除去を含む。実施形態では、ステップ(i)において、融合タンパク質は、約1μM~約200μMの濃度で存在する。実施形態では、ステップ(i)において、融合タンパク質は、約30μM~約60μMの濃度で存在する。
Smart Images

Figure 2026530421000001_ABST
Abstract
Description
Technical Field
[0001] Cross-reference to Related Applications This application claims priority to U.S. Patent Application Serial No. 63 / 578,551 filed on August 24, 2023 and U.S. Patent Application Serial No. 63 / 554,715 filed on February 16, 2024, the entire contents of which are incorporated herein by reference for all purposes.
[0002] Reference to Electronic Sequence Listing The content of the electronic sequence listing (ISOL_011_02WO_SeqList_ST26.xml; size: 1,566,780 bytes; date of creation: August 21, 2024) is incorporated herein by reference in its entirety.
[0003] The present disclosure generally relates to fusion proteins, compositions comprising the fusion proteins, and methods of using the fusion proteins for nucleic acid purification. Background Art
[0004] As highlighted by mRNA-based vaccines against SARS-CoV-2, the use of nucleic acids as therapeutic agents is increasing rapidly. Purification of nucleic acids such as mRNA is difficult because mRNA is a much larger molecule than proteins. Due to its large size, mRNA does not interact sufficiently with conventional affinity chromatography resins. In addition, the contaminant profile of nucleic acid compositions is often complex, and it is difficult to find a purification platform that can remove all contaminants from the composition. Summary of the Invention Problem to be Solved by the Invention
[0005] There is a need in the art for improved compositions and methods for rapidly and cost-effectively purifying nucleic acids such as mRNA. Means for Solving the Problem
[0006] This disclosure provides a fusion protein comprising a nucleic acid-binding protein, a method for using the fusion protein to purify nucleic acids, and a method for using the fusion protein.
[0007] This specification provides compositions comprising a fusion protein. In embodiments, the fusion protein comprises a nucleic acid-binding protein (NBP) and a polypeptide having phase behavior. In this embodiment, NBP is: RNA-specific adenosine deaminase 1 (ADAR1), ADAR1 double-stranded RNA-binding domain 3 (dsRBD3), Bacillus subtilis cold shock protein B (Bs-CspB), cold shock domain Y-box protein (CSD-Ybox), eukaryotic translation initiation factor 4E (eIF4e), Fox-1 protein (FOX1), heteronuclear ribonucleoprotein Q1 (hnRNPQ1), human (Homo sapiens) zinc finger CCCH-type 14 (HsZC3H14), poly(A)-binding protein (PABP), poly(A)-binding protein nucleus 1 (PABPN1), pentatricopeptide repeat protein A (PPRpA), pumilio-like repeat protein A (PUFpA), and Staufen, 12-O-tetradecanoylphorbol-13-acetate-inducible sequence 11 The NBP is selected from one of D(TIS11D), Z-type DNA / RNA binding protein 1 (ZBP1), and zinc finger nuclease (ZNF). In embodiments, the NBP contains an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one polypeptide of SEQ ID NOs. In the embodiment, the NBP is selected from one of the following: ADAR1 double-stranded RNA-binding domain 3 (dsRBD3), heteronuclear ribonucleoprotein Q1 (hnRNPQ1), human (Homo sapiens) zinc finger CCCH-type 14 (HsZC3H14), poly(A)-binding protein nucleus 1 (PABPN1), pentatricopeptide repeat protein A (PPRpA), and pumilio-like repeat protein A (PUFpA).In this embodiment, the NBP contains an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one polypeptide of SEQ ID NOs. In this embodiment, the polypeptide having phase behavior has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one polypeptide of SEQ ID NOs: 1 to 60 or 217. In this embodiment, the polypeptide having phase behavior has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with respect to either polypeptide SEQ ID NO: 57 or 60.
[0008] In the embodiment, the fusion protein includes a linker. In the embodiment, the linker includes an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to any one polypeptide of SEQ ID NOs.
[0009] In embodiments, NBP binds to RNA, DNA, or both. In embodiments, NBP binds to RNA selected from double-stranded RNA (dsRNA), single-stranded RNA (ssRNA), mRNA, premRNA, polyadenosine (polyA)RNA, Z-conformational RNA (Z-type RNA), or a combination thereof. In embodiments, NBP binds to the 3' end of mRNA, the 3' untranslated region (UTR) of mRNA, the polyA tail of mRNA, or the AU-rich element of mRNA, or a combination thereof. In embodiments, NBP binds to premRNA. In embodiments, NBP binds to the introns, exons, polyA tail of premRNA, or a combination thereof. In embodiments, NBP binds to DNA. In embodiments, NBP binds to single-stranded DNA, double-stranded DNA, polyadenosine (polyA)DNA, Z-conformational DNA (Z-type DNA), or a combination thereof.
[0010] In one embodiment, the fusion protein is encoded by a nucleic acid. In one embodiment, the nucleic acid contains a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to any one of the nucleic acids of sequence numbers 119-133. In one embodiment, the fusion protein is encoded by a vector. In one embodiment, the vector comprises a nucleic acid sequence, wherein the nucleic acid comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one of the nucleic acids of sequence numbers 119 to 133. In one embodiment, the vector has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with respect to any one nucleic acid of sequence numbers 104-118, 134, and 135.
[0011] This specification provides a method for purifying nucleic acids, which includes: (i) contacting a fusion protein of the Disclosure herein with a composition comprising nucleic acids and at least one contaminant, wherein the fusion protein binds to the nucleic acids to form a complex; (ii) contacting the complex with a first environmental factor to increase the size of the complex; (iii) separating the complex from at least one contaminant; and (iv) separating the nucleic acid from the fusion protein by contacting the complex with a second environmental factor. In embodiments, the complex is separated from at least one contaminant based on size. In embodiments, the size-based separation is carried out using one of the following methods selected from tangential flow filtration, membrane chromatography, analytical ultracentrifugation, high-performance liquid chromatography, membrane chromatography, total filtration, ultrasonic separation, centrifugation, countercurrent centrifugation, and high-performance protein liquid chromatography. In embodiments, the first environmental factor includes: (a) a change in one or more of the following: temperature, pH, salt concentration, concentration of the purified matrix, concentration of virus particles, or pressure; (b) the addition of one or more surfactants, cofactors, vitamins, molecular tungsten agents, reducing agents, oxidizing agents, enzymes, or denaturants; or (c) the application of electromagnetic waves or sound waves. In embodiments, the second environmental factor includes: (a) a change in one or more of the following: temperature, pH, salt concentration, concentration of the purified matrix, concentration of virus particles, or pressure; (b) the addition of one or more surfactants, cofactors, vitamins, molecular tungsten agents, reducing agents, oxidizing agents, enzymes, or denaturants; or (c) the application of electromagnetic waves or sound waves. In embodiments, at least one contaminant is selected from: solvents, proteins, peptides, carbohydrates, nucleic acids, viruses, cells (e.g., bacteria, yeast, or mammalian cells), carbohydrates, lipids, or lipopolysaccharides.
[0012] This specification provides a fusion protein comprising Bacillus subtilis cold shock protein B (Bs-CspB) and a polypeptide having phase behavior. In embodiments, Bs-CspB comprises an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polypeptide of SEQ ID NO: 90. In embodiments, the polypeptide having phase behavior comprises P and G motifs comprising a plurality of proline residues and a plurality of glycine residues. In embodiments, the P and G motifs comprise at least about 10% proline residues and at least about 20% glycine residues. In embodiments, the polypeptide exhibiting phase behavior includes a pentapeptide repeat having the sequence (Val-Pro-Gly-Xaa-Gly)n (SEQ ID NO: 217), or its randomized, scrambled analogues; where Xaa may be any amino acid other than proline. In embodiments, n is an integer from 1 to 360 including the endpoint.In this embodiment, polypeptides having phase behavior are as follows: a(GRGDSPY)n(SEQ ID NO: 1); b.(GRGDSPH)n(SEQ ID NO: 2); c.(GRGDSPV)n(SEQ ID NO: 3); d.(GRGDSPYG)n(SEQ ID NO: 4); e.(RPLGYDS)n(SEQ ID NO: 5); f.(RPAGYDS)n(SEQ ID NO: 6); g.(GRGDSYP)n(SEQ ID NO: 7); h.(GRGDSPYQ)n(SEQ ID NO: 8); i.(GRGNSPYG)n(SEQ ID NO: 9); j.(GVGVP)n(SEQ ID NO: 10); k.(GVGVPGLGVPGVGVPGLGVPGVGVP)m(SEQ ID NO: 11); l.(GVGVPGVGVPGAGVPGVGVPGVGVP)m( An amino acid sequence selected from SEQ ID NO: 12); m.(GVGVPGWGVPGVGVPGWGVPGVGVP)m(SEQ ID NO: 13); n.(GVGVPGVGVPGVGVPGVGVGVPGVGVGVPGEGVPGFGVPGVGVP)m(SEQ ID NO: 14); o.(GVGVPGVGVPGVGVPGVGVPGVGVGVPGKGVPGFGVPGVGVP)m(SEQ ID NO: 15); and p.(GAGVPGVGVPGAGVPGVGVPGAGVP)m(SEQ ID NO: 16); or including randomized and scrambled analogues thereof; where n is an integer in the range of 20 to 360 including the endpoint; and m is an integer in the range of 4 to 25 including the endpoint. In the embodiment, the polypeptide having phase behavior comprises an amino acid sequence selected from the following: a. (GVGVP)m (SEQ ID NO: 22); b. (ZZPXXXXGZ)m (SEQ ID NO: 23); c. (ZZPXGZ)m (SEQ ID NO: 24); d. (ZZPXXGZ)m (SEQ ID NO: 25); or e. (ZZPXXXGZ)m (SEQ ID NO: 26), where m is an integer from 10 to 160 including the endpoint, X is any amino acid except proline or glycine if present, and Z is any amino acid if present.
[0013] In the embodiment, the polypeptide having phase behavior comprises an amino acid sequence selected from the following: a. (GVGVPGVGVPGAGVPGVGVPGVGVP)m (SEQ ID NO: 17); or b. (GVGVPGVGVPGLGVPGVGVPGVGVP)m (SEQ ID NO: 18); where m is an integer from 2 to 32 including the endpoint. In the embodiment, the polypeptide having phase behavior comprises an amino acid sequence selected from the following: a. (GVGVPGVGVPGAGVPGVGVPGVGVP)m (SEQ ID NO: 19) (wherein m is 8 or 16); b. (GVGVPGAGVP)m (SEQ ID NO: 20) (wherein m is an integer from 5 to 80 including the endpoint); or c. (GXGVP)m (SEQ ID NO: 21) (wherein m is an integer from 10 to 160 including the endpoint); wherein each repeat, X is independently selected from the group consisting of glycine, alanine, valine, isoleucine, leucine, phenylalanine, tyrosine, tryptophan, lysine, arginine, aspartic acid, glutamic acid, and serine. In this embodiment, the polypeptide having phase behavior has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one polypeptide of SEQ ID NOs: 1-60, 217, or 262-264. In this embodiment, the polypeptide having phase behavior has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one polypeptide of SEQ ID NO: 56 or 262-264.In one embodiment, the polypeptide having phase behavior has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polypeptide of SEQ ID NO: 90. In another embodiment, the fusion protein has at least 90% identity with the polypeptide of SEQ ID NO: 90 and at least 90% identity with any one of the polypeptides of SEQ ID NO: 56 and 262-264. In yet another embodiment, the fusion protein has at least 90% identity with the polypeptide of SEQ ID NO: 90 and at least 90% identity with the polypeptide of SEQ ID NO: 56.
[0014] In the embodiment, the fusion protein includes a linker. In the embodiment, the linker includes an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to any one polypeptide of SEQ ID NOs. In one embodiment, the linker contains an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polypeptide of SEQ ID NO: 261. In one embodiment, the fusion protein contains an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with either one polypeptide of SEQ ID NO: 225 or 226.
[0015] In the embodiment, the fusion protein has at least 90% identity with the polypeptide of SEQ ID NO: 225 or 226. In the embodiment, Bs-CspB binds to RNA, DNA, or both. In the embodiment, Bs-CspB binds to RNA, DNA, or both. In the embodiment, Bs-CspB binds to ssRNA. In the embodiment, Bs-CspB binds to the 3'-end of mRNA, the 5'-end of mRNA, the coding region of mRNA, the non-coding region of mRNA, or a combination thereof. In the embodiment, Bs-CspB binds to premRNA. In the embodiment, Bs-CspB binds to the intron, exon, 5'UTR, 3'UTR, or a combination thereof of premRNA. In the embodiment, Bs-CspB binds to DNA. In the embodiment, Bs-CspB binds to single-stranded DNA, double-stranded DNA, polyadenosine (polyA)DNA, Z-conformation DNA (Z-type DNA), or a combination thereof.
[0016] This specification provides nucleic acids encoding fusion proteins. In embodiments, the nucleic acid comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to any one of the nucleic acids of SEQ ID NOs.
[0017] This specification provides vectors encoding fusion proteins. This specification provides vectors containing the nucleic acids described above. This specification provides vectors having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to any one of the nucleic acids of SEQ ID NOs.
[0018] This specification provides a method for purifying single-stranded nucleic acids, which includes: (i) contacting a composition comprising single-stranded nucleic acids and at least one contaminant with a fusion protein according to any one of claims 1 to 28, wherein the fusion protein binds to the single-stranded nucleic acids to form a complex; (ii) adding a first environmental factor to the composition comprising the complex to increase the size of the complex; (iii) separating the complex from at least one contaminant; and (iv) separating the single-stranded nucleic acid from the fusion protein by contacting the complex with a second environmental factor to form a product comprising single-stranded nucleic acids. In embodiments, the single-stranded nucleic acid comprises ssRNA. In embodiments, the complex is separated from at least one contaminant based on size. In embodiments, the size-based separation is carried out using a method selected from one of tangential flow filtration, membrane chromatography, analytical ultracentrifugation, high-performance liquid chromatography, volumetric filtration, ultrasonic separation, centrifugation, countercurrent centrifugation, and high-performance protein liquid chromatography. In embodiments, the size-based separation is carried out using a method comprising centrifugation. In embodiments, the first environmental factor includes: (a) a change in one or more of the following: temperature, pH, salt concentration, concentration of the purified matrix, concentration of virus particles, or pressure; (b) the addition of one or more surfactants, cofactors, vitamins, molecular tungsten agents, reducing agents, oxidizing agents, enzymes, or denaturants; or (c) the application of electromagnetic waves or sound waves. In embodiments, the second environmental factor includes: (a) a change in one or more of the following: temperature, pH, salt concentration, concentration of the purified matrix, concentration of virus particles, or pressure; (b) the addition of one or more surfactants, cofactors, vitamins, molecular tungsten agents, reducing agents, oxidizing agents, enzymes, or denaturants; or (c) the application of electromagnetic waves or sound waves. In embodiments, at least one contaminant is selected from: solvents, proteins, peptides, carbohydrates, nucleic acids, viruses, cells (e.g., bacteria, yeast, or mammalian cells), carbohydrates, lipids, or lipopolysaccharides. In embodiments, at least one contaminant is a double-stranded nucleic acid. In this embodiment, the double-stranded nucleic acid is dsRNA.In the embodiment, the salt is added at a concentration of about 0.5 M to about 3 M. In the embodiment, the salt is added at a concentration of about 1.2 M to about 1.7 M. In the embodiment, the product contains about 70% to about 100% single-stranded nucleic acids, based on the total nucleic acid content in the product of step (iv). In the embodiment, the single-stranded nucleic acids include ssRNA. In the embodiment, the product contains at least one contaminant at a concentration of about 10% or less, based on the total nucleic acid content in the product. In the embodiment, the at least one contaminant includes dsRNA. In the embodiment, the purification method includes the removal of at least about 1-log of contaminants compared to the amount of contaminants in the composition of step (i). In the embodiment, the purification method includes the removal of 1 to 10-log of contaminants compared to the amount of contaminants in the composition of step (i). In the embodiment, in step (i), the fusion protein is present at a concentration of about 1 μM to about 200 μM. In the embodiment, in step (i), the fusion protein is present at a concentration of approximately 30 μM to approximately 60 μM.
[0019] The accompanying drawings are incorporated herein and constitute part thereof, illustrating several, but not exclusive, exemplary embodiments and / or features. The embodiments and drawings disclosed herein are intended to be considered illustrative, not restrictive. [Brief explanation of the drawing]
[0020] [Figure 1] This study compares the nonspecific binding of various polypeptides exhibiting phase behavior to RNA templates and linearized plasmid DNA of varying lengths and chain states using agarose gel electrophoresis. Successful binding of nucleic acid species by biomacromolecules is demonstrated by retention within the wells. Polypeptide 40L80 (SEQ ID NO: 264), exhibiting phase behavior, showed nonspecific binding to a diverse range of nucleic acid species. [Figure 2]This shows the elution rate of total RNA directly captured from in vitro transcription (IVT) reaction products that individually transcribe mRNA templates of various sizes, using an ssRNA purification reagent containing a fusion protein (SEQ ID NO: 225) that contains a polypeptide (called 100V80) (SEQ ID NO: 56) having phase behavior with Bs-CspB (SEQ ID NO: 90). The elution rate represents the final yield after elution compared to the total RNA present in the initial IVT reaction product. [Figure 3A] , [Figure 3B] This demonstrates that ssRNA purification reagents containing Bs-CspB and 100V80 are highly selective for ssRNA. Figure 3A shows the Log10 removal value (LRV) of dsRNA during RNA capture of purified RNA with added dsRNA across various ssRNA purification reagents:RNA moles. The reagent maintained selectivity for ssRNA and was able to remove dsRNA exceeding 4-log. Figure 3B shows the elution rate of total dsRNA in samples with various percentages of added dsRNA. Selectivity was maintained in the capture reaction performed on purified RNA despite the addition of a high value of 10% of total RNA for dsRNA. [Figure 4] This study demonstrates the screening of ssRNA binding candidates in independent capture reaction products using agarose gel electrophoresis, assessing their ability to capture diverse ssRNA templates and dsRNA targets. Fusion proteins containing Bs-CspB and 100V80 selectively captured a variety of ssRNA templates. The presence of bands indicates nucleic acids that were not captured. [Modes for carrying out the invention]
[0021] definition As used herein and in the appended claims, the singular forms “a,” “an,” and “the” refer to multiple subjects unless the context clearly indicates otherwise. Thus, for example, “protein” refers to one protein or a mixture of such proteins, and “method” refers to equivalent steps and / or methods known to those skilled in the art, etc.
[0022] Where used herein, the terms “about” or “approximately” preceding a number indicate a value that is within a 10% range of plus or minus that value. For example, “about 100” includes 90 and 110.
[0023] Furthermore, as used herein, “and / or” means and encompasses all possible combinations of one or more of the related enumerated items, and the absence of any combination as interpreted alternatively ("or").
[0024] Unless otherwise indicated by the context, the various features described herein are explicitly intended to be used in any combination.
[0025] Furthermore, this disclosure intends that in some embodiments, any feature or combination of features described herein may be excluded or omitted. To elaborate further, for example, where the specification indicates that a particular amino acid can be selected from A, G, I, L, and / or V, this wording also indicates that an amino acid can be selected from any subset of these amino acids, e.g., A, G, I, or L; A, G, I, or V; A or G; L only, as each such partial combination is expressly described herein. Furthermore, such wording also indicates that one or more of the specified amino acids can be excluded. For example, in certain embodiments, an amino acid may be neither A, G, nor I; not A; not G, nor V; and so on, as if each such possible exclusion were expressly described herein.
[0026] As used herein, the term “environmental factor” refers to any factor that, when applied to a composition containing a fusion protein, alters one or more properties of that composition. Non-limiting examples of environmental factors include changes in one or more of the following: temperature, pH, salt concentration, fusion protein concentration, nucleic acid concentration, or pressure; the addition of one or more surfactants, molecular condensers, denaturants, reducing agents, or oxidizing agents; or the application of electromagnetic waves or sound waves.
[0027] As used herein, the term “contaminant” may refer to any substance undesirable in the purified composition. In embodiments, the contaminant is any substance other than nucleic acids that are desirable to be purified. Non-limiting examples of contaminants include, but are not limited to, solvents, proteins, peptides, carbohydrates, nucleic acids, viruses, cells (e.g., bacteria, yeast, or mammalian cells), lipids, or lipopolysaccharides. In embodiments, the contaminant is an endotoxin or mycotoxin. In embodiments, the cells are immune cells. In embodiments, the immune cells are T cells, B cells, NK cells, peripheral blood mononuclear cells, monocytes, macrophages, or neutrophils. In embodiments, the cells are T cells expressing chimeric antigen receptors (CARs). In embodiments, the contaminant is a double-stranded nucleic acid.
[0028] As used herein, the terms “peptide,” “polypeptide,” and “protein” are interchangeable and refer to compounds consisting of amino acid residues covalently linked by peptide bonds. A protein must contain at least two amino acids, and there is no limit to the maximum number of amino acids that can constitute a protein sequence. The term “peptide” may refer to short-chain amino acids, including, for example, native peptides, recombinant peptides, synthetic peptides, or combinations thereof. Proteins and peptides may include, for example, biologically active fragments, substantially homologous polypeptides, oligopeptides, homodimers, heterodimers, polypeptide variants, modified polypeptides, derivatives, analogs, and fusion proteins.
[0029] A "polynucleotide" is a sequence of nucleotide bases and may be RNA, DNA, or a DNA-RNA hybrid sequence (including both naturally occurring and / or non-naturally occurring nucleotides). In embodiments, a polynucleotide is a single-stranded or double-stranded DNA sequence.
[0030] As used herein, “isolating” or “purifying” (or grammatically equivalent) a nucleic acid means that the nucleic acid is separated at least partially from at least some of the other components in the starting material containing the nucleic acid (e.g., a cell lysate). In embodiments, the “isolated” or “purified” nucleic acid may be concentrated at least about 10 times, about 100 times, about 1,000 times, about 10,000 times, or more compared to the starting material.
[0031] As used herein, the term “polypeptide having phase behavior” refers to any polypeptide that can undergo a phase transition. In some embodiments, polypeptides undergo a phase transition upon application of environmental factors. Examples of polypeptides having phase behavior include elastin-like polypeptides (ELPs) and resilin-like polypeptides (RLPs).
[0032] As used herein, the term “fusion protein” refers to a polypeptide produced when two heterogeneous nucleotide sequences or fragments thereof, which encode two (or more) different polypeptides that are not found to be fused in nature, fuse with each other in a correct translational reading frame.
[0033] As used herein, the term “nucleic acid-binding protein” is also known as NBP and may refer to any amino acid sequence (protein, peptide, etc.) that binds to a target nucleic acid. In embodiments, an NBP may include the full-length, cleaved, or modified form of a receptor for the target nucleic acid. In embodiments, an NBP may be the antigen-binding portion of a monoclonal antibody (e.g., Fab), a monoclonal antibody-derived single-chain variable region fragment (scFv); a native ligand of the target nucleic acid; a peptide with sufficient affinity for the target nucleic acid; a single-domain binder such as a camelid binder; an artificial binder such as Darpin; or a single chain derived from a T cell receptor.
[0034] As used herein, the term “fragment” referring to a protein or polypeptide includes cleaved forms of the protein or polypeptide. For example, a fragment of NBP may contain about 50% to about 99.9% of the full-length NBP. In embodiments, a fragment of NBP may contain about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 97%, or about 99% of the amino acids of the full-length NBP.
[0035] As used herein, the term “capture efficiency” in relation to the fusion protein described herein refers to the amount of nucleic acid captured by the fusion protein relative to the amount of nucleic acid present in the starting composition. The capture efficiency is calculated using the following formula: 100 × (amount of nucleic acid captured by the fusion protein / amount of nucleic acid in the composition before purification).
[0036] In relation to two or more nucleic acid or polypeptide sequences, the term “identity percentage” refers to two or more sequences or subsequences that, when compared, have a predetermined percentage of the same nucleotide or amino acid residues. Unless otherwise noted, sequence identity is determined using the National Center for Biotechnology Information (NCBI) Basic Local Alignment Search Tool (BLAST®) (blast.ncbi.nlm.nih.gov / Blast.cgi). In embodiments, sequence identity is calculated over the entire length of the sequences being compared. In embodiments, sequence identity is calculated over fragments of approximately 20, 50, 75, 100, 250, 500, 750, or 1000 amino acids from each of the sequences being compared.
[0037] Fusion protein This disclosure provides a fusion protein and a method for using the fusion protein to purify nucleic acids. In embodiments, the fusion protein comprises an NBP (i.e., a nucleic acid-binding protein) that binds to a target nucleic acid and a polypeptide having phase behavior, where the NBP is bound to the polypeptide having phase behavior.
[0038] nucleic acid-binding proteins In embodiments, the fusion protein includes a nucleic acid-binding protein (NBP). In embodiments, the NBP binds to one or more nucleic acids. In embodiments, the NBP binds to single-stranded nucleic acids (e.g., single-stranded DNA or RNA). In embodiments, the NBP binds to double-stranded nucleic acids (e.g., double-stranded DNA or RNA). In embodiments, the NBP binds to DNA. In embodiments, the DNA is single-stranded or double-stranded. In embodiments, the NBP binds to RNA. In embodiments, the RNA is single-stranded or double-stranded. In embodiments, the NBP binds to both DNA and RNA, where DNA and RNA are single-stranded, double-stranded, or a combination of both. In embodiments, the NBP binds to polyadenosine (polyA) DNA. In embodiments, the NBP binds to polyadenosine (polyA) RNA. In embodiments, the NBP binds to polythymidine (polyT) DNA. In embodiments, the NBP binds to polyuridine (PolyU) RNA. In the embodiment, NBP binds to polycytidine (polyC)RNA. In the embodiment, NBP binds to polyguanidine (polyG)RNA. In the embodiment, NBP binds to polycytidine (polyC)DNA. In the embodiment, NBP binds to polyguanidine (polyG)DNA.
[0039] In the embodiment, NBP binds to DNA, microRNA, capped RNA, DNA, double-stranded RNA, transfer RNA, ribosomal RNA, nuclear small RNA, regulatory RNA, ribozyme, transfer RNA, or messenger RNA.
[0040] In embodiments, NBPs bind to AU-rich RNA elements (AREs). AREs refer to adenylate-uridilate-rich elements located in the 5' or 3' untranslated region of mRNA. AREs contain the core sequence AUUUA. AREs are determinants of RNA stability and often occur in the mRNAs of proto-oncogenes, nuclear transcription factors, and cytokines. Proteins that bind to AREs are called ARE-binding proteins (ARE-BPs). In embodiments, ARE-BPs stabilize mRNA. Non-exclusive examples of ARE-BPs include human antigen R (huR, also known as "ELAV"), tristetrapolin (TTP), AU-rich element RNA-binding protein (AUF), and fragile X mental retardation syndrome-related protein 1 (FXR1).The following papers describe ARE-BP, and their entirety is incorporated herein by reference: Otsuka et al. Front. Genet., 02 May 2019; Brennan and Steitz. Cell Mol Life Sci. 2001 Feb; 58(2): 266-77; Carballo et al. 1998. Science, 281, 1001-1005; Mazan-Mamczarz et al. Oncogene volume 27, pages 6151-6163 (2008); Vasudevan and Steitz. Cell. 2007 Mar 23; 128(6): 1105-18; Curr Cancer Drug Targets. 2019; 19(5): 382-399; J Biol Chem. 2017 Apr 28;292(17):6869-6881.doi:10.1074 / jbc.M116.772947.Epub 2017 Mar 16;Wiley Interdiscip Rev RNA.Jul-Aug 2014;5(4):549-64.doi:10.1002 / wrna.1230;and Elife.2017 Aug 2;6:e26129.doi:10.7554 / eLife.26129;Mazan-Mamczarz et al.2008.Nucleic Acids Research,37,204-214.In embodiments, the NBP binds to ARE.In embodiments, the NBP that binds to ARE incorporates a binding element for huR, TTP, AUF, or FXR1. The entirety of ARE is disclosed in the references incorporated herein: Barreau C, et al. AU-rich elements and associated factors: are there unifying principles? Nucleic Acids Res. 2006 Jan 3;33(22):7138-50. In embodiments, NBP binds to Z-conformation DNA (Z-type DNA).
[0041] In this embodiment, NBP binds to Z-conformation RNA (Z-type RNA). Generally, Z-type DNA and R-DNA are left-handed structures of DNA and RNA, respectively, as described in the references incorporated herein as a whole: Barreau C, et al. AU-rich elements and associated factors: are there unifying principles? Nucleic Acids Res. 2006 Jan 3;33(22):7138-50.
[0042] In this embodiment, NBP binds to microRNA.
[0043] In the embodiment, NBP binds to nuclear small RNA (snRNA). In the embodiment, NBP binds to nucleolar small RNA (snoRNA).
[0044] In this embodiment, NBP binds to regulatory RNA.
[0045] In this embodiment, NBP is bound to the ribozyme.
[0046] In the embodiment, NBP binds to transfer RNA (tRNA). In the embodiment, NBP binds to long non-coding RNA (lncRNA).
[0047] In the embodiment, NBP binds to mRNA. In the embodiment, NBP binds to premRNA. In the embodiment, NBP binds to premRNA introns. In the embodiment, NBP binds to premRNA exons. In the embodiment, NBP binds to the 3' end of mRNA. In the embodiment, NBP binds to the 5' end of mRNA. In the embodiment, NBP binds to the 5' cap of mRNA. In the embodiment, NBP binds to positively charged intrinsically disordered regions (IDRs) or nucleic acids. In the embodiment, NBP binds to nucleic acid sequences. In the embodiment, NBP binds to nucleic acid structures. In the embodiment, NBP binds to nucleic acid secondary structures. In the embodiment, NBP binds to nucleic acid tertiary structures. In the embodiment, NBP binds to naturally occurring nucleic acids. In the embodiment, NBP binds to synthetically produced nucleic acids. In the embodiment, the nucleic acids are naturally occurring but modified by synthetic means.
[0048] In the embodiment, the NBP is bound to a nucleic acid sequence, where the nucleic acid sequence is approximately 30 nucleotides to approximately 10,000 nucleotides in length. In the embodiment, the nucleic acid is at least approximately 30 nucleotides in length. In an embodiment, the nucleic acid is at least about 35 nucleotides long, for example, at least about 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, 2,000, 2,500, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, or at least about 10,000 nucleotides long. In the embodiment, the nucleic acid is at least about 100 nucleotides long. In the embodiment, the nucleic acid is at least about 250 nucleotides long. In the embodiment, the nucleic acid is at least about 500 nucleotides long. In the embodiment, the nucleic acid is at least about 1,000 nucleotides long. In the embodiment, the nucleic acid is at least about 1,500 nucleotides long. In the embodiment, the nucleic acid is at least about 2,000 nucleotides long. In the embodiment, the nucleic acid is at least about 2,500 nucleotides long. In the embodiment, the nucleic acid is at least about 3,000 nucleotides long. In the embodiment, the nucleic acid is at least about 4,000 nucleotides long. In the embodiment, the nucleic acid is at least about 5,000 nucleotides long. In the embodiment, the nucleic acid is at least about 6,000 nucleotides long. In the embodiment, the nucleic acid is at least about 7,000 nucleotides long. In the embodiment, the nucleic acid is at least about 8,000 nucleotides long. In the embodiment, the nucleic acid is at least about 9,000 nucleotides long. In this embodiment, the nucleic acid is at least about 10,000 nucleotides long.
[0049] In some embodiments, the nucleic acid has a diameter or length of approximately 0.001 μm to approximately 500 μm. In some embodiments, the nucleic acid has a diameter of 1 nm to 100 μm including the endpoint. In some embodiments, the nucleic acid has a diameter of 1 nm to 100 nm including the endpoint. In some embodiments, the nucleic acid has a diameter of 100 nm to 1 μm including the endpoint. In some embodiments, the nucleic acid has a diameter of 1 μm to 50 μm including the endpoint. In some embodiments, the nucleic acid has a diameter of 50 μm to 100 μm including the endpoint.
[0050] In the embodiment, the nucleic acid size (i.e., diameter or length) is approximately 0.001 μm, approximately 0.002 μm, approximately 0.003 μm, approximately 0.004 μm, approximately 0.005 μm, approximately 0.006 μm, approximately 0.007 μm, approximately 0.008 μm, approximately 0.009 μm, approximately 0.010 μm, approximately 0.020 μm, approximately 0.030 μm, approximately 0.040 μm, approximately 0.050 μm, approximately 0.060 μm, approximately 0.070 μm, approximately 0. 080μm, about 0.090μm, about 0.1μm, about 0.2μm, about 0.3μm, about 0.4μm, about 0.5μm, about 0.6μm, about 0.7μm, about 0.8μm, about 0.9μm, about 1μm, about 2μm , about 3μm, about 4μm, about 5μm, about 6μm, about 7μm, about 8μm, about 9μm, about 10μm, about 11μm, about 12μm, about 13μm, about 14μm, about 15μm, about 16μm, about 17μm, about 18μ m, about 19μm, about 20μm, about 21μm, about 22μm, about 22μm, about 23μm, about 24μm, about 25μm, about 26μm, about 27μm, about 28μm, about 29μm, about 30μm, about 31μm, about 32μm, about 33μm, about 34μm, about 35μm, about 36μm, about 37μm, about 38μm, about 39μm, about 40μm, about 41μm, about 42μm, about 43μm, about 44μm, about 45μm, about 46μm , approximately 47 μm, approximately 48 μm, approximately 49 μm, approximately 50 μm, approximately 55 μm, approximately 60 μm, approximately 65 μm, approximately 70 μm, approximately 75 μm, approximately 80 μm, approximately 85 μm, approximately 90 μm, approximately 95 μm, approximately 100 μm, approximately 150 μm, approximately 200 μm, approximately 250 μm, approximately 300 μm, approximately 350 μm, approximately 400 μm, approximately 450 μm, or approximately 500 μm or larger (including all values and ranges in between). In the embodiment, the nucleic acid has a size of 10 μm or larger. In the embodiment, the nucleic acid has a size of 25 μm or larger. In the embodiment, the nucleic acid has a size of 50 μm or larger. In the embodiment, the nucleic acid has a size of 100 μm or larger.
[0051] In some embodiments, the nucleic acids have a size (i.e., molecular weight) of about 2 kDa to about 1000 MDa. In embodiments, the nucleic acids are about 2 kDa, about 5 kDa, about 15 kDa, about 20 kDa, about 20 kDa, about 25 kDa, about 30 kDa, about 35 kDa, about 40 kDa, about 45 kDa, about 50 kDa, about 55 kDa, about 60 kDa, about 65 kDa, about 70 kDa, about 75 kDa, about 80 kDa, about 85 kDa, about 90 kDa, about 95 kDa Da, about 100kDa, about 150kDa, about 200kDa, about 250kDa, about 300kDa, about 350kDa, about 400kDa, about 450kDa, about 500kD a, about 550kDa, about 600kDa, about 650kDa, about 700kDa, about 750kDa, about 800kDa, about 850kDa, about 900kDa, about 950kDa, It has a molecular weight of approximately 1000 kDa, approximately 1 MDa, approximately 5 MDa, approximately 10 MDa, approximately 15 MDa, approximately 20 MDa, approximately 25 MDa, approximately 50 MDa, approximately 75 MDa, approximately 100 MDa, approximately 125 MDa, approximately 150 MDa, approximately 175 MDa, approximately 200 MDa, approximately 225 MDa, approximately 250 MDa, approximately 275 MDa, approximately 300 MDa, approximately 325 MDa, approximately 350 MDa, approximately 400 MDa, approximately 425 MDa, approximately 450 MDa, approximately 500 MDa, approximately 550 MDa, approximately 600 MDa, approximately 650 MDa, approximately 700 MDa, approximately 750 MDa, approximately 800 MDa, approximately 850 MDa, approximately 900 MDa, approximately 950 MDa, or approximately 1000 MDa (including all values and ranges between them).
[0052] In some embodiments, the fusion protein contains 1 to about 100 NBPs, 1 to about 75 NBPs, 1 to about 50 NBPs, 1 to about 40 NBPs, 1 to about 30 NBPs, 1 to about 20 NBPs, 1 to about 15 NBPs, 1 to about 10 NBPs, or 1 to about 5 NBPs. In embodiments, the fusion protein contains about 1, about 5, about 10, about 20, about 30, about 40, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, or about 100 NBPs. In embodiments, a single polypeptide having phase behavior may bind to multiple NBPs, such as about 1 to 100 NBPs.
[0053] In the embodiment, the fusion protein may have two or more NBPs that bind to different nucleic acids, or that bind to the same nucleic acid.
[0054] In the embodiment, the affinity of NBP to nucleic acids is adjusted to facilitate the separation of nucleic acids from the fusion protein.
[0055] In this embodiment, NBP is: RNA-specific adenosine deaminase 1 (ADAR1), ADAR1 double-stranded RNA-binding domain 3 (dsRBD3), Bacillus subtilis cold shock protein B (Bs-CspB), cold shock domain Y-box protein (CSD-Ybox), eukaryotic translation initiation factor 4E (eIF4e), Fox-1 protein (FOX1), heteronuclear ribonucleoprotein Q1 (hnRNPQ1), human zinc finger CCCH-type 14 (HsZC3H14), poly(A)-binding protein (PABP), poly(A)-binding protein nucleus 1 (PABPN1), pentatricopeptide repeat protein A (PPRpA), Pumilio homologous domain (PUM-HD), pumilio-like repeat protein A (PUFpA), and Staufen, 12-O-tetradecanoylphorbol-13-acetate-inducible sequence 11 The NBP is selected from one of the following: D(TIS11D), Z-type DNA / RNA binding protein 1 (ZBP1), and zinc finger nuclease (ZNF). In the embodiment, the NBP is ADAR1. In the embodiment, the NBP is ADAR1 dsRBD3. In the embodiment, the NBP is Bs-CspB. In the embodiment, the NBP is CSD-Ybox. In the embodiment, the NBP is eIF4e. In the embodiment, the NBP is FOX1. In the embodiment, the NBP is hnRNPQ1. In the embodiment, the NBP is HsZC3H14. In the embodiment, the NBP is PABP. In the embodiment, the NBP is PABPN1. In the embodiment, the NBP is PPRpA. In the embodiment, the NBP is PUFpA. In the embodiment, the NBP is PUM-HD. In the embodiment, the NBP is Staufen. In the embodiment, NBP is TIS11D. In the embodiment, NBP is ZBP. In the embodiment, NBP is ZNF.
[0056] In embodiments, NBP is RNA-specific adenosine deaminase 1 (ADAR1) or a fragment, subunit, or domain thereof. ADAR1 is a polypeptide that catalyzes the post-transcriptional deamination of adenosine, thereby converting it to inosine. ADAR1 binds to and catalyzes double-stranded RNA. ADAR1 is described in the following reference (which is incorporated herein by reference in its entirety): Song B, et al. The role of RNA editing enzyme ADAR1 in human disease. Wiley Interdiscip Rev RNA. 2022 Jan;13(1):e1665. ADAR1 comprises one or more Z-type DNA-binding domains, one or more dsRNA-binding domains, and a deaminase domain. In embodiments, NBP is ADAR1 double-stranded RNA-binding domain 3 (dsRBD3).
[0057] In one embodiment, ADAR1 dsRBD3 has an amino acid sequence that is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the polypeptide of Sequence ID No. 89.
[0058] In embodiments, the NBP includes a cold shock domain (CSD). The CSD contains five antiparallel β-chains that form a β-barrel structure known as oligosaccharide / oligonucleotide-linked folding. The CSD binds to single-stranded RNA and single-stranded DNA. In embodiments, the CSD consists of 60 to 80 amino acids, for example, about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, or about 80 amino acids. In embodiments, the CSD contains about 70 amino acids. In embodiments, the CSD is a bacterial CSD. In embodiments, the bacterial CSD prefers ssDNA to ssRNA by up to 10 times.
[0059] In embodiments, NBP is Bacillus subtilis cold shock protein B (Bs-CspB) or a fragment, subunit, or domain thereof. Bs-CspB is also known as Bacillus subtilis-derived CspB (Bscscp). Bs-CspB binds to single-stranded DNA (ssDNA) and single-stranded RNA (ssRNA). Bs-CspB is a polypeptide that functions as an RNA chaperone and transcriptional antiterminator. Bs-CspB is described in the following reference (the entire reference is incorporated herein by reference): Sachs R, et al. RNA single strands bind to a conserved surface of the major cold shock protein in crystals and solution. RNA. 2012 Jan;18(1):65-76. Bs-CspB contains one or more nucleic acid binding motifs.
[0060] In the embodiments, the fusion protein includes Bs-CspB. In the embodiments, Bs-CspB binds to RNA, DNA, or both. In the embodiments, Bs-CspB binds to RNA selected from one of the following: double-stranded RNA (dsRNA), single-stranded RNA (ssRNA), mRNA, premRNA, polyadenosine (polyA)RNA, Z-conformation RNA (Z-type RNA), or a combination thereof. In the embodiments, Bs-CspB binds to ssRNA. In the embodiments, Bs-CspB binds to the 3' end of mRNA, the 3' untranslated region (UTR) of mRNA, the polyA tail of mRNA, or the AU-rich element of mRNA, or a combination thereof. In the embodiments, Bs-CspB binds to the intron, exon, polyA tail of premRNA, or a combination thereof. In the embodiments, Bs-CspB binds to DNA. In the embodiment, Bs-CspB binds to single-stranded DNA, double-stranded DNA, polyadenosine (polyA) DNA, Z-conformation DNA (Z-type DNA), or a combination thereof.
[0061] In this embodiment, Bs-CspB has an amino acid sequence that is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to that of the polypeptide of SEQ ID NO: 90.
[0062] In embodiments, NBP is a cold shock domain Y-box protein (CSD-Ybox) or a fragment, subunit, or domain thereof. CSD-Ybox is also known as Y-box protein 1 (YB1) cold shock domain (CSD) (YB1-CSD). CSD-Ybox binds to single-stranded DNA (ssDNA) and single-stranded RNA (ssRNA). CSD-Ybox is a polypeptide that regulates nucleic acid metabolism, for example, through DNA repair, premRNA transcription and splicing, mRNA packaging, and regulation of mRNA stability and translation. CSD-Ybox is described in the following reference (the entire reference is incorporated herein by reference): Heinemann U. and Roske Y. Cold-Shock Domains-Abundance, Structure, Properties, and Nucleic-Acid Binding. Cancers (Basel) 2021 Jan 7;13(2):190. CSD-Ybox contains one or more nucleic acid binding motifs.
[0063] In one embodiment, CSD-Ybox has an amino acid sequence that is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the polypeptide of Sequence ID No. 91.
[0064] In embodiments, NBP is eukaryotic translation initiation factor 4E (eIF4e) or a fragment, subunit, or domain thereof. eIF4e binds to the 5'-cap of mRNA. eIF4e is a polypeptide that, in synergy with other proteins, binds to mRNA and enables the recruitment of ribosomes for translation initiation. eIF4e is described in the following reference (which is incorporated herein by reference in its entirety): Davis, MR, et al. Nuclear eIF4E Stimulates 3'-end Cleavage of Target RNAs. Cell Reports, 27, 1397-1408. eIF4e contains a cap-binding pocket.
[0065] In one embodiment, eIF4e has an amino acid sequence that is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to any one polypeptide of SEQ ID NOs.
[0066] In one embodiment, eIF4e has an amino acid sequence that is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the polypeptide of either SEQ ID NO: 86 or 87.
[0067] In embodiments, NBP is the Fox-1 protein (FOX1) or a fragment, subunit, or domain thereof. FOX1 binds to premRNA introns and exons. FOX1 is a polypeptide that promotes or represses exon expression by regulating alternative splicing. FOX1 is described in the following reference (which is incorporated herein by reference in its entirety): Kuroyanagi H. Fox-1 family of RNA-binding proteins. Cell Mol Life Sci. 2009 Dec;66(24):3895-907. FOX1 contains an RNA recognition motif (RRM).
[0068] In one embodiment, FOX1 has an amino acid sequence that is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the polypeptide of Sequence ID No. 92.
[0069] In embodiments, NBP is heteronuclear ribonucleoprotein Q1 (hnRNPQ1) or a fragment, subunit, or domain thereof. hnRNPQ1 binds to mRNA. hnRNPQ1 is a polypeptide that regulates mRNA processing events, such as premRNA splicing, mRNA transport, and translational regulation. hnRNPQ1 is described in the following reference (which is incorporated herein by reference in its entirety): Xing, L. et al. Negative regulation of RhoA translation and signaling by hnRNP-Q1 affects cellular morphogenesis. Molecular Biology of the Cell 2012 23:8, 1500-1509. hnRNPQ1 comprises one or more RNA recognition motifs (RRMs), an acidic domain, and an Arg-Gly-Gly (RGG) box domain.
[0070] In one embodiment, hnRNPQ1 has an amino acid sequence that is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the polypeptide of SEQ ID NO: 93.
[0071] In embodiments, NBP is human (Homo sapiens) zinc finger CCCH-type 14 (HsZC3H14) or a fragment, subunit, or domain thereof. HsZC3H14 binds to poly(A) RNA. HsZC3H14 is a polypeptide that regulates the length of the 3'-polyadenosine (poly(A)) tail. HsZC3H14 is described in the following reference (the entire reference is incorporated herein by reference): Rha J, et al. The RNA-binding protein, ZC3H14, is required for proper poly(A) tail length control, expression of synaptic proteins, and brain function in mice. Hum Mol Genet. 2017 Oct 1;26(19):3663-3681. HsZC3H14 contains an N-terminal proline-tryptophan-isoleucine (PWI)-like domain and a C-terminal tandem CysCysCysHis (CCCH) zinc finger (ZF) domain.
[0072] In one embodiment, HsZC3H14 has an amino acid sequence that is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the polypeptide of either SEQ ID NO: 94 or 95.
[0073] In embodiments, NBP is a poly(A)-binding protein (PABP) or a fragment, subunit, or domain thereof. PABP binds to poly(A) RNA. PABP is a polypeptide that mediates the circularization of mRNA. PABP is described in the following reference (which is incorporated herein by reference in its entirety): Mangus DA et al. Poly(A)-binding proteins: multifunctional scaffolds for the post-transcriptional control of gene expression. Genome Biol 4, 223 (2003). PABP contains one or more RNA recognition motifs (RRMs).
[0074] In this embodiment, PABP has an amino acid sequence that is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to any one polypeptide of SEQ ID NOs.
[0075] In this embodiment, PABP has an amino acid sequence that is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the polypeptide of Sequence ID No. 77.
[0076] In embodiments, NBP is poly(A)-binding protein nucleus 1 (PABPN1) or a fragment, subunit, or domain thereof. PABPN1 binds to the poly(A) RNA tail. PABPN1 is a polypeptide that regulates RNA processing, for example, by inhibiting the export of unspliced RNA from the nucleus and regulating the length of the poly(A) tail. PABPN1 is described in the following reference (the entire reference is incorporated herein by reference): Mangus DA et al. Poly(A)-binding proteins: multifunctional scaffolds for the post-transcriptional control of gene expression. Genome Biol 4, 223 (2003). PABPN1 contains an RNA recognition motif (RRM).
[0077] In this embodiment, PABPN1 has an amino acid sequence that is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the polypeptide of Sequence ID No. 97.
[0078] In embodiments, NBP is a pentatricopeptide repeat (PPR) or a fragment, subunit, or domain thereof. The PPR binds to the 3' or 5' end of an RNA transcript. The PPR is a polypeptide that regulates RNA stabilization and translational activation. In embodiments, NBP is PPR protein a (PPRpA). The PPR contains 20 to 50 amino acids, for example, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, or about 50 amino acids. In embodiments, the PPR contains about 35 amino acids. In embodiments, the NBP contains PPRs that are repeated about 2 to 30 times within the NBP sequence. In embodiments, the NBP contains PPRs that are repeated about 10 to 30 times within the NBP sequence. In embodiments, the PPRs may be repeated about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 times. In embodiments, the PPRs are repeated at least 10 times. The PPR repeats may be contiguous or separated by one or more amino acids. The PPR repeats form two antiparallel α-helices. In embodiments, the NBP containing the PPRs forms a solenoid structure. In embodiments, the NBP containing PPR binds to single-stranded RNA, single-stranded DNA, or mRNA. In embodiments, the NBP containing PPR binds to the 5' cap of mRNA. In embodiments, the NBP containing PPR binds to the 3' poly-A tail of mRNA. PPR is described in the following reference (which is incorporated herein by reference in its entirety): Manna, S. An overview of pentatricopeptide repeat proteins and their applications, Biochimie, 113, 2015 93-99. The PPR comprises one or more PPR motifs and one or more helix-turn-helix motifs.
[0079] In this embodiment, PPRpA has an amino acid sequence that is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the polypeptide of Sequence ID No. 98.
[0080] In embodiments, NBP is a Pumilio-like repeat (PUF) or a fragment, subunit, or domain thereof. PUFs bind to the 3'-UTR of mRNA in the cytosol and to the precursor of rRNA in the nucleolus. PUFs are polypeptides that function as post-transcriptional and translational regulators. In embodiments, NBP is PUF protein A (PUFpA). The PUF domain contains eight α-helix repeats consisting of a conserved 36-amino acid sequence, which form a concave RNA-binding surface. In embodiments, NBP contains one to eight α-helix repeats of the PUF domain, e.g., 1, 2, 3, 4, 5, 6, 7, or 8 α-helix repeats. In embodiments, NBP containing a PUF binds to the poly-A tail. In embodiments, NBP containing a PUF binds to mRNA. PUFs are described in the following reference (which is incorporated herein by reference in its entirety): Wang M, et al. The PUF Protein Family: Overview on PUF RNA Targets, Biological Functions, and Post Transcriptional Regulation. Int J Mol Sci. 2018 Jan 30;19(2):410. PUFs contain one or more Pumilio homologous domains (PUM-HD).
[0081] In one embodiment, PUFpA has an amino acid sequence that is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the polypeptide of Sequence ID No. 99.
[0082] In embodiments, NBP is a Pumilio homologous domain (PUM-HD) or a fragment, subunit, or domain thereof. PUM-HD binds to the 3'-UTR of mRNA and represses its translation. PUM-HD is a polypeptide that functions as a post-transcriptional and translational regulator. PUFs are described in the following reference (which is incorporated herein by reference in its entirety): Wang M, et al. The PUF Protein Family: Overview on PUF RNA Targets, Biological Functions, and Post Transcriptional Regulation. Int J Mol Sci. 2018 Jan 30;19(2):410. PUM-HD contains an RNA recognition motif.
[0083] In one embodiment, PUM-HD has an amino acid sequence that is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the polypeptide of either SEQ ID NO: 102 or 103.
[0084] In embodiments, NBP is the Staufen protein or a fragment, subunit, or domain thereof. Staufen binds to double-stranded RNA (dsRNA). Staufen is a polypeptide that functions in RNA transport, degradation, and translational repression. Staufen is described in the following reference (which is incorporated herein by reference in its entirety): Visentin S, et al. A multipronged approach to understanding the form and function of hStaufen protein. RNA. 2020 Mar;26(3):265-277. Staufen comprises one or more dsRNA-binding domains (RBDs) and one or more tubulin-binding domains (TBDs).
[0085] In one embodiment, Staufen has an amino acid sequence that is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the polypeptide of Sequence ID No. 100.
[0086] In embodiments, NBP is 12-O-tetradecanoylphorbol-13-acetate (TPA)-derived sequence 11D (TIS11D) or a fragment, subunit, or domain thereof. TIS11D binds to AU-rich elements in mRNA. TIS11D is a polypeptide that functions in RNA metabolism and secondary hematopoiesis. TIS11D is described in the following reference (which is incorporated herein by reference in its entirety): Morgan, B.R. et al. Probing the Structural and Dynamical Effects of the Charged Residues of the TZF Domain of TIS11d, 2015 Biophysical Journal, 108:6, 1503-1515. TIS11D comprises one or more CCCH-type tandem zinc finger domains and one or more (R / K)YKTEL motifs.
[0087] In one embodiment, TIS11D has an amino acid sequence that is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to any one polypeptide of sequence numbers 61 to 66.
[0088] In one embodiment, TIS11D has an amino acid sequence that is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the polypeptide of Sequence ID No. 65.
[0089] In embodiments, NBP is Z-type DNA / RNA binding protein 1 (ZBP1) or a fragment, subunit, or domain thereof. ZBP1 binds to Z-conformation DNA and RNA. ZBP1 is a polypeptide that functions in the innate immune response by binding to foreign nucleic acids and inducing type I interferon production. ZBP1 is described in the following reference (the entire reference is incorporated herein by reference): Maelfait J, et al. Sensing of viral and endogenous RNA by ZBP1 / DAI induces necroptosis. EMBO J. 2017 Sep 1;36(17):2529-2543. ZBP1 has one or more Z-binding domains (ZBDs).
[0090] In one embodiment, ZBP1 has an amino acid sequence that is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to any one polypeptide of sequence numbers 67 to 73.
[0091] In one embodiment, ZBP1 has an amino acid sequence that is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the polypeptide of Sequence ID No. 65.
[0092] In embodiments, NBP is a zinc finger nuclease (ZNF) or a fragment, subunit, or domain thereof. ZNF binds to double-stranded DNA and RNA. ZNF is a polypeptide that regulates gene expression at the transcriptional and translational levels. ZNF is described in the following reference (which is incorporated herein by reference in its entirety): Chaves-Arquero B, et al. The distinct RNA-interaction modes of a small ZnF domain underlay TUT4(7) diverse action in miRNA regulation. RNA Biol. 2021 Nov 12;18(sup2):770-781. ZNF comprises one or more C2H2 domains, one or more CCHC domains, and a catalytic domain.
[0093] In one embodiment, the ZNF has an amino acid sequence that is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to any one of the polypeptides of SEQ ID NO: 101.
[0094] NBPs containing Bs-CspB are described herein. In embodiments, the NBP binds to double-stranded RNA. In embodiments, the NBP comprises a dsRNA-binding protein (dsRBD) or a fragment thereof. The dsRBD is described in its entirety in the following reference, which is incorporated herein by reference: Banerjee et al. RNA Biol. 2014 Oct;11(10):1226-1232.
[0095] In the embodiment, the NBP binds to capped mRNA. In the embodiment, the NBP that binds to capped mRNA includes eukaryotic translation initiation factor 4E (eIF4E), eukaryotic translation initiation factor 3 subunit D (eIF3D), or a combination thereof.
[0096] In embodiments, NBPs bind to the groove of DNA or RNA. Non-limiting examples of nucleic acid-binding proteins that bind to the groove of DNA or RNA include the transactivator (Tat) protein of human immunodeficiency virus-1 (HIV-1), the REV protein of HIV-1, and the RSG-1.2 peptide. The RSG-1.2 peptide is a synthetic peptide that binds to the Rev response element located within the env gene of the HIV-1 genome. The RSG-1.2 peptide is described in the following paper, which is incorporated herein by reference in its entirety: Kumar et al. PLoS One. 2011;6(8):e23300.
[0097] In the embodiments, NBP binds to mRNA. In the embodiments, the NBP that binds to mRNA is a ribosomal protein. In the embodiments, the ribosomal protein is a 70S ribosome or an 80S ribosome. In the embodiments, the ribosomal protein is derived from the 40S or 60S subunit of an 80S ribosome. In the embodiments, the ribosomal protein is derived from the 30S or 50S subunit of a 70S ribosome. In the embodiments, the ribosomal protein is selected from the group consisting of L3 ribosomal protein, L4 ribosomal protein, L13 ribosomal protein, L20 ribosomal protein, L22 ribosomal protein, L24 ribosomal protein, L24e ribosomal protein, S12 ribosomal protein, S14 ribosomal protein, and eukaryotic initiation factor 4E-binding protein 1 (4EBP1).
[0098] In the embodiment, the NBP that binds to mRNA is part of the spliceosome. In the embodiment, the NBP that is part of the spliceosome is a splicing factor. In the embodiment, the splicing factor is selected from ASF / SF2 splicing factor, serine / arginine-rich splicing factor 4 (SRp75), and serine and arginine-rich splicing factor 1 (SRSF1).
[0099] In the embodiment, the NBP that binds to mRNA is a protein localized to P granules. In the embodiment, the protein localized to P granules is selected from the group consisting of LAF-1, MEG-1, and MEG-3. LAF-1, MEG-1, and MEG-3 are described in the following references, which are incorporated herein by reference in their entirety: Leacock et al. Genetics, Volume 178, Issue 1, 1 January 2008, Pages 295-306; Wu et al. Mol Biol Cell. 2019 Feb 1;30(3):333-345; Elbaum-Garfinkle et al. Proc Natl Acad Sci USA. 2015 Jun 9;112(23):7189-94.
[0100] In embodiments, the NBP that binds to mRNA is a protein that removes or facilitates the removal of the mRNA's 5' cap, and is referred to herein as a “decapping protein.” In embodiments, the protein that removes or facilitates the removal of the mRNA's 5' cap is Dcp1, Dcp2, or a combination thereof. Dcp1 and Dcp2 are described in their entirety in the following reference, which is incorporated herein by reference: Valkov et al. Nature Structural & Molecular Biology volume 23, pages 574-579 (2016).
[0101] In the embodiment, the NBP that binds to mRNA is a component of the processing body (p-body). In the embodiment, the components of the p-body are Edc3, DHX9, or Xrn1. The components of the p-body are described in their entirety in the following reference, which is incorporated herein by reference: Luo et al. Biochemistry 2018, 57, 17, 2424-2431.
[0102] In the embodiments, the mRNA-binding NBP is a stem-loop binding protein (SLBP). SLBP binds to the histone 3' untranslated region (UTR) stem-loop structure of replication-dependent histone mRNA. In the embodiments, the mRNA-binding NBP is a heteronuclear ribonucleoprotein (hnRNP). hnRNP is described in its entirety in the following reference, which is incorporated herein by reference: Geuens et al. Hum Genet. 2016;135:851-867.
[0103] In this embodiment, the NBP that binds to mRNA is GroEL.
[0104] In embodiments, NBPs are proteins involved in in vitro transcription. Non-limiting examples of NBPs involved in in vitro transcription include T7 RNA polymerase, RNase inhibitors, 2'-O-methyltransferase, inorganic pyrophosphatase, poly(A) polymerase, DNase I, calf intestinal phosphatase, Antarctic phosphatase, the D1 subunit of vaccinia virus mRNA capping enzyme, guanine-7-methyltransferase (located in the D1 subunit of vaccinia virus mRNA capping enzyme), guanylyltransferase (located in the D1 subunit of vaccinia virus mRNA capping enzyme), RNA triphosphatase (located in the D1 subunit of vaccinia virus mRNA capping enzyme), and the D12 subunit of vaccinia virus mRNA capping enzyme. The following references describe the aforementioned protein, and their entirety is incorporated herein by reference: Dickson et al. Prog Nucleic Acid Res Mol Biol. 2005;80:349-374; Shuman et al. J Biol Chem. 1980 Dec 10;255(23):11588-11598; Luo et al. J Virol. 1995 Jun;69(6):3852-3856; Kobori et al. PNAS November 1, 1984 81(21)6691-6695.
[0105] In this embodiment, NBP includes poly(A)-binding protein (PABP), eukaryotic translation initiation factor 4E (eIF4E), eukaryotic translation initiation factor 3 subunit D (eIF3D), heteronuclear ribonucleoprotein (hnRNP), RNA-specific adenosine deaminase 1 (ADAR1), RNA-specific adenosine deaminase 2 (ADAR2), Bacillus subtilis-derived CspB (Bscscp), Y-box protein 1 cold shock domain (YB1-CSD), Fox-1 protein (FOX1), poly(A)-binding protein (PABP), Staufen protein, TIS11d, zinc finger protein (ZNF), Z-DNA-binding protein 1 (ZBP1), retinoic acid-inducible gene-I (RIG-I)-like protein, Toll-like receptor 7 (TLR7), Toll-like receptor 8 (TL The following are selected from the group consisting of R3), Toll-like receptor 8 (TLR8), retinoic acid-inducible gene I (RIG-I), melanoma differentiation-associated protein 5 (MDA5), interferon-inducible protein with tetratricopeptide repeat 1 (IFIT1), protein kinase R (PKR), 2'-5'-oligoadenylate synthetase, oligoadenylate synthase-like (OASL) proteins (e.g., OAS1, OAS2, OAS3, or OASL), ribonuclease E (RNASE E), gamma interferon-inducible protein Ifi-16 (IF116), and cyclic GMP-AMP synthase (cGAS).The following references describe the selected proteins mentioned above, and their entirety is incorporated herein by reference: Kuroyanagi. Cell Mol Life Sci. 2009;66(24):3895-3907; Baou et al. J Biomed Biotechnol. 2009;2009:634520; and Brisse and Ly. Front.Immunol., 17 July 2019;10:1586; Rehwinkel et al. Nature Reviews Immunology volume 20, pages 537-551 (2020); and Brisse et al. Front Immunol. 2019;10:1586; Luo et al. Cell. 2011 Oct 14;147(2):409-422.
[0106] In embodiments, the NBP comprises one or more RNA-binding domains (RBDs) and one or more intrinsically disordered regions (IDRs). In embodiments, the IDR comprises an RG[G] repeat, an RS / RG-rich domain, a K / R patch, a molecular recognition function, a low complexity sequence, a pentatricopeptide domain, or a combination thereof.
[0107] In embodiments, NBP includes one or more of the following domains: short linear motif (SLiM), RG repeat, RGG repeat, RS / RG rich domain, K / R basicity patch, molecular recognition function, low complexity sequence, RNA recognition motif, double-stranded RNA binding domain, K homology domain, zinc finger domain (e.g., CCHH ZF domain, CCCC (Ran-BP2) domain, CCCH ZF domain), RGG domain, Pumillo family domain, pentatricopeptide domain, cold shock domain, helicase domain, La motif, Piwi-Argonaute-Zwille (PAZ) domain, PIWI (P-element induced wimpy testis), pseudouridine synthase, and PUA (archaeosine). Domains with YT521-B homology include transglycosylate, Pumillo-like repeats (PUM), ribosomal S1-like (S1), Sm and Like-Sm (Sm / Lsm) repeats, thiridine synthase and RNA methylase and pseudouridine synthase (THUMP), and YT521-B homology. The following references describe many of these domains, and their entirety is incorporated herein by reference: Balcerak et al. Open Biol. 2019 Jun;9(6)190096; Jarvelin et al. Cell Commun Signal, 2016:14,9; Corley et al. Mol.Cell. 2020 Apr 2;78(1):9-29; De Franco et al. Sci Rep:2019:9,2484; Shotwell et al. 2020. Wiley Interdiscip Rev RNA,11,e1573; Simon et al. 2019. Molecular Cell,75,66-75.e5; Varadi et al, 2015, PLoS One,10,e0139731; Zeke et al, 2020, WIREs RNA,n / a,e1714.
[0108] In embodiments, NBPs include short-chain motifs (SLiMs). SLiMs consist of up to 10 amino acid residue motifs located primarily outside the protein domain. SLiMs bind to RNA nonspecifically with low affinity. SLiMs are often repeated multiple times throughout the protein.
[0109] In the embodiment, the NBP includes pseudouridine synthase and archaeosine transglycosylase (PUA) domains. The PUA domain is in the range of 67 to 94 amino acids in length and has a β1α1β2β3β4β5α2β6 structure that forms a pseudo-barrel enclosed by two α-helices. In the embodiment, the NBP containing the PUA binds to double-stranded RNA.
[0110] In the embodiment, the NBP includes an S1 RNA-binding domain. In the embodiment, the NBP including the S1 RNA-binding domain interacts with single-stranded RNA, double-stranded RNA, or mRNA. In the embodiment, the S1 RNA-binding domain contains about 60 to about 80 amino acids, for example, about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, or about 80 amino acids. In the embodiment, the S1 RNA-binding domain contains about 70 amino acids.
[0111] In the embodiments, the NBP contains an Sm RNA-binding motif. The Sm RNA-binding motif is present in eukaryotic and archaeal Sm and Like-sm (Lsm) proteins, as well as in prokaryotic Hfq proteins. The Sm motif consists of approximately 70 residues having an α1β1β2β3β4β5 topology that forms a curved antiparallel β-sheet. Sm-containing proteins readily multimerize via interactions between β4 and β5 chains in two Sm motifs. In the embodiments, the NBP contains 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 Sm motifs. In the embodiments, the NBP contains two Sm motifs. In the embodiments, the Sm-binding motif binds to RNA via hydrogen bonding and base stacking interactions.
[0112] In this embodiment, the NBP contains thiouridine synthase, RNA methylase, and pseudouridine synthase (THUMP) domains. The THUMP domain is present in many tRNA modifying enzymes. The THUMP domain is located near the RNA modifying domain and, in some cases, near the N-terminal ferredoxin-like domain. The THUMP domain presents an α1α2β1α3β2β2 topology, forming parallel α-helices adjacent to the β-sheet. In this embodiment, the NBP containing the THUMP domain binds to tRNA.
[0113] In one embodiment, NBP contains a YT521-B homologous domain. In one embodiment, the YT521-B homologous domain contains 100 to 150 amino acids, for example, about 100, about 101, about 102, about 103, about 104, about 105, about 106, about 107, about 108, about 109, about 110, about 111, about 112, about 113, about 114, about 115, about 116, about 117, about 118, about 119, about 120, about 121, about 1 It contains approximately 22, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, or 150 amino acids. In embodiments, the NBP containing the YT521-B homologous domain binds to methylated adenosine.
[0114] In the embodiment, the NBP includes a Piwi-Argonaut-Zwille (PAZ) domain. In the embodiment, the PAZ domain facilitates the binding of small interfering mRNAs and / or microRNA guides to mRNA targets. In the embodiment, the PAZ domain is derived from a Dicer protein or an Argonaute protein. The PAZ domain has two α-helices at the top and presents a six-stranded β-barrel adjacent to a special appendage containing a β-hairpin and a short α-helix on the opposite side.
[0115] In the embodiment, NBP includes a P-element induced Wimpy Testis (PIWI) domain. In the embodiment, the PIWI domain facilitates the binding of small interfering mRNAs and / or microRNA guides to mRNA targets. In the embodiment, the PIWI domain is located on the Argonaute protein. The tertiary structure of the PIWI domain forms an RNase H-like folding consisting of a quintuple β-sheet with α-helices adjacent on both sides.
[0116] In this embodiment, NBP includes a PAZ domain and a PIWI domain.
[0117] In embodiments, the NBP includes an RS / RG-rich domain. The RS / RG-rich domain includes arginine-serine repeats (RS), arginine-glycine repeats (RG), or a combination thereof. The RS / RG-rich domain mediates specific or nonspecific interactions with RNA. Examples of proteins containing RS / RG-rich domains include SR proteins and SR-like protein-like serine / arginine-rich splicing factor 1 (SRSF1) and RNA helicase DDX23.
[0118] In embodiments, the NBP includes a helicase domain. The helicase comprises six superfamilies (SF) including SF1, SF2, SF3, SF4, SF5, and SF6. In embodiments, the helicase domain is a eukaryotic RNA and DNA helicase derived from the SF1 or SF2 superfamily. Non-restrictive examples of families within the SF1 and SF2 superfamilies include the Upf1-like family, the DEAD-box family, the DEAH family, the RIG-I-like family, the Ski2-like family, and the NS3 family. In embodiments, the helicase domain is a bacterial or viral helicase derived from the SF3, SF4, SF5, or SF6 superfamily. When ATP binds to the helicase, it promotes a higher affinity of the helicase domain for RNA. ATP hydrolysis promotes a structural change that causes substrate unwinding and / or single nucleotide transfer of the helicase.
[0119] In embodiments, the NBP includes a La motif. The La motif consists of five α-helices and three β-chains that form a small antiparallel β-sheet relative to the modified “winged helix” folding. In embodiments, the La motif binds to the 3'-terminal UUU-OH element on polymerase III transcription small RNA. The La motif contains 80 to 100 amino acids, for example, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, or about 100 amino acids. In embodiments, the La motif contains about 90 amino acids.
[0120] In embodiments, NBP includes RG[G] repeats. RG[G] repeats are known to have extensive degenerate binding. An RG[G] repeat is an arginine and glycine-rich motif consisting of at least three RG / RGG repeats (e.g., 3-500) separated by 10 amino acid residues. An RG / RGG motif includes RGG and / or RG repeats of varying lengths interspersed with spacers of different amino acids. In embodiments, NBP includes a di-RGG motif. A di-RGG motif includes two repeating RGG sequences separated by 0-4 amino acids. In embodiments, NBP includes a di-RG motif. A di-RG motif includes two repeating RG sequences separated by 0-4 amino acids. In embodiments, NBP includes a tri-RGG motif. A tri-RGG motif includes three repeating RGG sequences separated by 0-4 amino acids. In embodiments, NBP includes a tri-RG motif. The tri-RG motif contains three repeating RG sequences separated by 0 to 4 amino acids. These motifs are described in their entirety in the following paper, which is incorporated herein by reference: Thandapani et al. (2013). Molecular Cell, 50, 613-623.
[0121] In embodiments, the amino acid sequence of the NBP includes one or more RG, RGG, RGGR, or RGGGR, or a combination thereof. In embodiments, the NBP containing RG, RGG, RGGR, RGGGR, or a combination thereof mediates hydrogen bonding and base stacking with DNA and RNA via the arginine moiety. In embodiments, the NBP containing RG, RGG, RGGR, RGGGR, or a combination thereof binds to the DNA G quadruplex. An exemplary protein containing repeats of RGG, RGGR, or RGGGR is the RNA-binding protein FUS. In embodiments, the NBP includes FUS. In embodiments, the NBP sequence includes consecutive repeats of RGG, RGGR, RGGGR, or a combination thereof. An exemplary NBP containing a combination of RGG, RGGR, or RGGGR repeats may include the sequence RGGRGGRGGRRGGRRGGRRGGGRRGG. In embodiments, the NBP may contain one or more RGG, RGGR, or RGGGR scattered throughout its sequence. In the embodiment, the NBP comprises 1 to 100 RG, RGG, RGGR, or RGGGR sequences. The RGG, RGGR, and RGGGR sequences may be scattered throughout the sequence (separated by one or more amino acids) or may be contiguous. In the embodiment, the NBP is approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, Includes approximately 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or approximately 100 RG, RGG, RGGR, or RGGGR repeats.The RG, RGG, RGGR, and RGGGR repeats may be consecutive or scattered throughout the sequence. The following paper describes an exemplary RGG sequence, which is incorporated herein by reference in its entirety: Simon et al. Molecular Cell (2019), 75, 66-75.e5.
[0122] In embodiments, the NBP includes an RG domain. The RG domain contains about 2 to about 500 RG (arginine-glycine) repeats. In embodiments, the NBP includes an RGG domain. The RGG domain contains about 2 to about 500 RGG (arginine-glycine-glycine) repeats. In embodiments, the NBP includes an RGGR domain. The RGGR domain contains about 2 to about 500 RGGR (arginine-glycine-glycine-arginine) repeats. In embodiments, the NBP includes an RGGGR domain. The RGGGR domain contains about 2 to about 500 RGG (arginine-glycine-glycine-glycine-arginine) repeats. In embodiments, the NBP includes an RG mixed domain. The RG mixed domain contains 2 to 500 simultaneous repeats of RG, RGG, RGGR, and / or RGGGR. For example, the RG mixed domain may include RGG, followed by RG, followed by RGGR, followed by RG, followed by RGGGR.
[0123] In embodiments, the NBP includes a K / R basic patch. The K / R basic patch contains 4 to 8 consecutive lysine, arginine, or a combination thereof. The K / R basic patch forms a strong positive exposed interface that binds to RNA. The K / R basic patch is often contained in multiple clusters on the same protein.
[0124] In the embodiment, NBP includes a molecular recognition function (MoRF). In the embodiment, the MoRF is up to 25 amino acids long, 50 or more amino acids long, or 25 to 50 amino acids long. The MoRF dynamically transitions from irregular to regular upon ligand binding.
[0125] In embodiments, NBP includes a low-complexity (LC) sequence. In embodiments, the LC sequence contains up to 100 amino acids and consists of the same amino acid or numerous repeats of several amino acids. The LC sequence can polymerize into amyloid-like fibrils and undergo a reversible phase transition to a hydrogel-like state. Examples of proteins containing LC sequences include FUS and hnRNPA2.
[0126] In embodiments, the NBP includes an RNA recognition motif (RRM). The RRM binds to RNA. Typically, the binding is sequence-specific. In embodiments, the RRM has a length of about 75 to about 125 amino acids, for example, about 75, about 80, about 85, about 90, about 95, about 100, about 105, about 110, about 115, about 120, or about 125 amino acids. In embodiments, the RRM contains about 85 amino acids. The RRM typically adopts a β1α1β2β3α2β4 topology, forming two α helices relative to an antiparallel β sheet, and accommodating the RNA-binding RNP1 and RNP2 motifs conserved in the central β1 and β3 chains.
[0127] In some embodiments, the NBP includes a double-stranded RNA-binding domain (dsRBD). In some embodiments, the dsRBD contains about 55–80 amino acids or about 65–70 amino acids. In some embodiments, the dsRBD contains 68 amino acids. The dsRBD typically takes on an αβββα structure. In some embodiments, the dsRBD exists as a tandem repeat or in combination with other RNA-binding domains. There are two subclasses of dsRBD: type B and type A. Type A binds to dsRNA better than type B. The dsRBD typically binds in a shape-dependent manner rather than in a sequence-specific manner. However, ADAR2 is a rare example of a dsRBD that exhibits sequence-specific binding.
[0128] In some embodiments, the NBP contains a K homologous domain. In some embodiments, the K homologous domain contains 60 to 80 amino acids. In some embodiments, the K homologous domain contains 70 amino acids. There are two types of K homologous domains: type I and reverse type II. The type I K homologous domain adopts a β1α1α2β2β'α' topology. The reverse type II K homologous domain adopts an α'β'β1α1α2β2 topology. The K homologous domain does not use aromatic amino acids for bonding, but instead uses hydrogen bonds. NBPs containing a K homologous domain are difficult to design due to their strict sequence specificity.
[0129] In embodiments, NBP comprises one or more zinc finger (ZF) domains. In embodiments, NBP comprises 1 to 100 ZF domains, for example, about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, Includes approximately 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 ZF domains. In embodiments, the zinc finger domain is selected from one of the following subtypes: CCHC (zinc knuckle), CCCH, CCCC (RanBP2), and CCHH. C and H refer to scattered cysteine and histidine residues that coordinate to the zinc atom. In embodiments, the zinc finger domain contains about 20 to about 40 amino acids, e.g., about 20, about 22, about 24, about 26, about 28, about 30, about 32, about 34, about 36, about 38, or about 40 amino acids. The CCHH ZF domain contains two conserved cysteine residues and two conserved histidine residues. The CCHH ZF domain recognizes both structure-specific and sequence-specific elements. To date, no engineered versions of the CCHH ZF domain exist. The CCHH ZF domain binds to both single-stranded and double-stranded DNA and RNA. CCCC ZF may not require a specific RNA structure for binding. Typically, CCCC ZF recognizes short three-nucleotide repeats.An engineered version of the CCCC ZF is described in the following document, which is incorporated herein by reference in its entirety: De Franco et al. Sci Rep:2019:9,2484.
[0130] In this embodiment, the nucleic acid-binding protein (NBP) is a polypeptide that has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one of the polypeptides in Table 1. In one embodiment, the NBP contains an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polypeptide of SEQ ID NO: 90.
[0131] [Table 1]
[0132] [Table 2]
[0133] [Table 3]
[0134] [Table 4]
[0135] [Table 5]
[0136] Polypeptides exhibiting phase behavior In embodiments, the fusion protein described herein comprises one or more polypeptides having phase behavior. In embodiments, the fusion protein described herein comprises polypeptides having phase behavior, namely 1-2, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8, 1-9, 1-10, 1-15, 1-20, 1-25, 1-30, 1-35, 1-40, 1-45, or 1-50. In the embodiment, the fusion protein comprises one polypeptide having phase behavior of about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, or about 50.
[0137] In embodiments, the polypeptide having phase behavior is a resilin-like polypeptide (RLP). A resilin-like polypeptide is an elastic polypeptide having desirable mechanical properties including elasticity, compressive modulus, tensile modulus, shear modulus, elongation to fracture, maximum tensile strength, hardness, rebound, and compression set. In embodiments, the resilin-like polypeptide described herein is a polymer comprising one or more repeats. In embodiments, the polymer repeats may have an amino acid sequence selected from any one of SEQ ID NOs: 1 to 9.
[0138] In this embodiment, the resilin-like polypeptide comprises two or more repeats, for example, a repeat of SEQ ID NO: 1 and a repeat of SEQ ID NO: 3.
[0139] In embodiments, the resilin-like polypeptide described herein includes repetitions that appear up to 500 times within a given RLP. In embodiments, the repetitions include about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, and It appears approximately 190 times, 200 times, 210 times, 220 times, 230 times, 240 times, 250 times, 260 times, 270 times, 280 times, 290 times, 300 times, 310 times, 320 times, 330 times, 340 times, 350 times, 360 times, 370 times, 380 times, 390 times, 400 times, 450 times, or 500 times.
[0140] In embodiments, RLP comprises one or more partial repeats. In embodiments, the length of the partial repeat is one, two, three, four, five, six, seven, eight, nine, ten, or more amino acids. In embodiments, RLP comprises one or more additional amino acids at the N-terminus or C-terminus of RLP that are not part of the repeat.
[0141] In embodiments, one or more RLP repeats are scrambled, i.e., they contain different amino acid sequences but retain the same amino acid composition. For example, repeats may have a different amino acid sequence than SEQ ID NO: 8 but retain the same amino acid composition. In embodiments, the polypeptide having phase behavior is an elastin-like polypeptide. Elastin-like polypeptides (ELPs) are biomacromolecules derived from tropoelastin. In embodiments, the polypeptide having phase behavior contains a P and G motif comprising multiple P residues and multiple G residues. In embodiments, the P and G motifs contain at least about 10% proline and at least about 20% glycine. In embodiments, the elastin-like polypeptide described herein has the sequence (Val-Pro-Gly-Xaa-Gly) nThe polymer contains a pentapeptide repeat having (SEQ ID NO: 217). In the embodiment, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108 ,109,110,111,112,113,114,115,116,117,118,119,120,121,122,123,124,125,126,127,128,129,130,131,132,133,134,135,136,137,138,139,140,141,142,143,144,145,146,147,148,149,15 0, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, or 500 (including all values and ranges between them). In embodiments, n is an integer from 1 to 360 including the endpoint. In embodiments, the polypeptide having phase behavior is as follows: a. (GRGDSPY) n (Sequence ID 1) b. (GRGDSPH) n (Sequence 2) c.(GRGDSPV) n (Sequence ID 3) d.(GRGDSPYG) n (Sequence No. 4) e.(RPLGYDS) n (Sequence ID 5) f.(RPAGYDS)n (SEQ ID NO: 6) g. (GRGDSYP) n (SEQ ID NO: 7) h. (GRGDSPYQ) n (SEQ ID NO: 8) i. (GRGNSPYG) n (SEQ ID NO: 9) j. (GVGVP) n (SEQ ID NO: 10); k. (GVGVPGLGVPGVGVPGLGVPGVGVP) m (SEQ ID NO: 11); l. (GVGVPGVGVPGAGVPGVGVPGVGVP) m (SEQ ID NO: 12); m. (GVGVPGWGVPGVGVPGWGVPGVGVP) m (SEQ ID NO: 13); n. (GVGVPGVGVPGVGVPGVGVPGVGVPGVGVPGEGVPGFGVPGVGVP) m (SEQ ID NO: 14); o. (GVGVPGVGVPGVGVPGVGVPGVGVPGVGVPGKGVPGFGVPGVGVP) m (SEQ ID NO: 15); and p. (GAGVPGVGVPGAGVPGVGVPGAGVP) m (SEQ ID NO: 16); comprises an amino acid sequence selected from, or a randomized, scrambled analog thereof. In embodiments, n is an integer ranging from 1 to 500 (including 1, 20, 40, 60, 80, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, 300, 320, 340, 360, 380, 400, 420, 440, 460, 480, or 500, including any value or range subsumed therein). In embodiments, n is an integer in the range of 20 to 360, inclusive of the endpoints. In one embodiment, m is an integer ranging from 1 to 100 (including 1, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, including all values and ranges therebetween). In embodiments, m is an integer in the range of 4 to 25, inclusive of the endpoints.
[0142] In the embodiment, the polypeptide having phase behavior is as follows: (GVGVP) m (Sequence code 22); (ZZPXXXXGZ) m (Sequence number 23); (ZZPXGZ) m (Sequence number 24); (ZZPXXGZ) m (Sequence number 25); or (ZZPXXXGZ) m The polypeptide comprises an amino acid sequence selected from (SEQ ID NO: 26), where m is an integer from 10 to 160 including the endpoint, X is any amino acid other than proline or glycine, if present, and Z is any amino acid, if present. In the embodiment, the polypeptide having phase behavior is as follows: (GVGVPGVGVPGAGVPGVGVPGVGVP) m (Sequence ID 17); or (GVGVPGVGVPGLGVPGVGVPGVGVP) m The polypeptide having a phase behavior is as follows: (GVGVPGVGVPGAGVPGVGVPGVGVP) m (Sequence ID 19) (m is 8 or 16); (GVGVPGAGVP) m (Sequence code 20) (m is an integer between 5 and 80, including the endpoint); or (GXGVP) m The sequence comprises an amino acid sequence selected from (SEQ ID NO: 21) (where m is an integer from 10 to 160 including the endpoint), where X in each repeat is independently selected from the group consisting of glycine, alanine, valine, isoleucine, leucine, phenylalanine, tyrosine, tryptophan, lysine, arginine, aspartic acid, glutamic acid, and serine. In embodiments, m is an integer from 1 to 100 (1, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 (including all values and ranges between them)).
[0143] In the embodiment, the pentapeptide repeat is scrambled and, for example, contains different amino acid sequences but maintains the same amino acid composition. For example, ELP may contain a different amino acid sequence than SEQ ID NO: 217 but maintains the same amino acid composition, for example, 40% of the sequence is glycine, 20% is Xaa, 20% is proline, and 20% is valine.
[0144] In embodiments, the ELP comprises one or more partial repeats. In embodiments, the length of the partial repeat is one, two, three, or four amino acids. In embodiments, the ELP comprises one or more additional amino acids at its N-terminus or C-terminus that are not part of the repeat.
[0145] ELP and RLP undergo phase transitions in response to environmental factors. ELP and RLP retain the ability to undergo phase transitions when bound to one or more polypeptides (such as one or more NBPs) or when expressed as fusion proteins with one or more other polypeptides (such as one or more NBPs). Polymers like ELP and RLP have a cloud point temperature (T c The transition temperature (T) is also called the transition temperature. t ) shows. In some embodiments, ELP and RLP are T t A reversible phase transition occurs from the soluble phase to the insoluble phase. The ELP, which transitions from the soluble phase to the insoluble phase due to heating or an increase in salt concentration, has a temperature called the lower critical solution temperature (LCST). t RLP, which transitions from a soluble phase to an insoluble phase due to cooling or a decrease in salt concentration, has a lower critical solution temperature called the lower critical solution temperature (UCST). t In this embodiment, the phase transition is caused by a change in the secondary structure of the ELP and / or RLP. For example, the phase transition of the ELP is caused by a random coil (T tThe transition occurs due to a secondary structure change from (less than) to a type II β-turn. In embodiments, the secondary structure change is characterized by circular dichroism spectropolarimetry, small-angle X-ray scattering, and a method selected from cryo-electron microscopy, ultraviolet-visible spectrophotometry, static light scattering, dynamic light scattering, nuclear magnetic resonance spectroscopy, solid-state nuclear magnetic resonance spectroscopy, infrared spectroscopy, Fourier transform infrared spectroscopy (FTIR), microscopy, and small-angle neutron scattering. In embodiments, the phase transition of ELP is not due to a secondary structure change.
[0146] In some embodiments, the RLP and ELP described herein have a transition temperature of about 0°C to about 100°C. In some embodiments, the RLP and ELP described herein have a transition temperature of about 10°C to about 50°C. In some embodiments, the transition temperature is about 0°C, about 1°C, about 2°C, about 3°C, about 4°C, about 5°C, about 6°C, about 7°C, about 8°C, about 9°C, about 10°C, about 11°C, about 12°C, about 13°C, about 14°C, about 15°C, about 16°C, about 17°C, about 18°C, about 19°C, about 20°C, about 21°C, about 22°C, about 23°C, about 24℃, approximately 25℃, approximately 26℃, approximately 27℃, approximately 28℃, approximately 29℃, approximately 30℃, approximately 31℃, approximately 32℃, approximately 33℃, approximately 34℃, approximately 35℃, approximately 36℃, approximately 37℃, approximately 38℃, approximately 39℃, approximately 40℃, approximately 41℃, approximately 42℃, approximately 43℃, approximately 44℃, approximately 45℃, approximately 46℃, approximately 47℃, approximately 48℃, approximately 49℃, approximately 50℃, approximately 51℃, approximately 52℃, approximately 53℃, approximately 54℃, approximately 55℃, approximately 56℃, approximately 57℃, approximately 58℃, approximately 59℃, approximately 60℃, approximately 61℃, approximately 62℃, approximately 63℃, approximately 64℃, approximately 65℃, approximately 66℃, approximately 67℃, approximately 68℃, approximately 69℃, approximately 70℃, approximately 71℃, approximately 72℃, approximately 73℃, approximately 74℃, approximately 75℃, The transition temperatures are approximately 76°C, 77°C, 78°C, 79°C, 80°C, 81°C, 82°C, 83°C, 84°C, 85°C, 86°C, 87°C, 88°C, 89°C, 90°C, 91°C, 92°C, 93°C, 94°C, 95°C, 96°C, 97°C, 98°C, 99°C, or 100°C. In embodiments, the RLP described herein has a transition temperature of approximately 10°C to approximately 100°C.
[0147] In the embodiments, the RLP and ELP described herein are Tt This is regulated by manipulating the primary structure of RLP and ELP, for example, the amino acid sequence. 。 In the embodiment, the hydrophobicity of ELP or RLP is adjusted. In the embodiment, the hydrophobicity of ELP is modified by changing the properties of guest residue Xaa. In the embodiment, the hydrophobicity of ELP or RLP is increased, and as a result, T t The hydrophobicity of ELP or RLP decreases, and as a result, T t The polarity increases. In the embodiment, the polarity of ELP or RLP is adjusted. In the embodiment, the polarity of ELP is adjusted by changing the properties of the guest residue Xaa. In the embodiment, the polarity of ELP or RLP increases, and as a result, T t The polarity increases. In this embodiment, the polarity of ELP or RLP decreases, and as a result, T t It decreases.
[0148] In this embodiment, the number of ELP pentapeptide repeats (n) is adjusted to T t Modify the following: In this embodiment, a pentapeptide repeat (Val-Pro-Gly-Xaa-Gly) nIn (Sequence ID 217), n is an integer from 1 to 500, including the endpoint. In this embodiment, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108 ,109,110,111,112,113,114,115,116,117,118,119,120,121,122,123,124,125,126,127,128,129,130,131,132,133,134,135,136,137,138,139,140,141,142,143,144,145,146,147,148,149,15 0, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, or 500 (including all values and ranges between them).
[0149] In embodiments, Xaa, also referred to herein as a “guest residue,” is any amino acid that does not exclude the phase behavior of the ELP. In embodiments, Xaa is any amino acid except proline. In embodiments, Xaa is independently selected for each repeat. For example, a given ELP contains the guest residues alanine, glycine, and valine in an 8:7:1 ratio. In some embodiments, Xaa is selected from the group consisting of alanine, arginine, asparagine, aspartic acid, cysteine, glutamic acid, glutamine, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, praline, serine, threonine, tryptophan, tyrosine, and valine. In the embodiment, Xaa is a non-classical amino acid selected from the group consisting of 2,4-diaminobutyric acid, α-aminoisobutyric acid, alloisoleucine, 4-aminobutyric acid, 2-aminobutyric acid (Abu), ε-Ahx, 6-aminohexanoic acid, 2-aminoisobutyric acid (Aib), 3-aminopropionic acid, ornithine, norleucine, norvaline, hydroxyproline, sarcosine, citrulline, homocitrulline, cysteic acid, t-butylglycine, t-butylalanine, phenylglycine, cyclohexylalanine, β-alanine, fluoroamino acids, designer amino acids, such as β-methylamino acids, Cα-methylamino acids, Na-methylamino acids, and amino acid analogs in general. In the embodiment, Xaa is a natural amino acid or a D isomer of a non-classical amino acid.
[0150] In the embodiments, the RLP and ELP described herein are T t This is regulated by introducing one or more environmental factors into a composition having RLP and / or ELP. 。 In this embodiment, the T of ELP and / or RLP t This is adjusted by adjusting the ionic strength of the solvent. In the embodiment, the ionic strength of the solvent is adjusted by adding a salt. In the embodiment, ELP and / or RLP are lower in solvents having anions classified as kosmotropes. tIt has. The cosmotrope anion is highly hydrated and affects the water shield of ELP and / or RLP. In the embodiment, the T of ELP and / or RLP t This can be adjusted by adding an anion that is a chaotrope. At low concentrations, the addition of a chaotrope can adjust the T of ELP and / or RLP. t The concentration increases. At high concentrations, the addition of chaotrope increases the T of ELP and / or RLP. t The T of ELP and / or RLP decreases. In this embodiment, the T of ELP and / or RLP t This can be prepared by introducing one or more reagents that break hydrogen bonds. Non-limiting examples of reagents that break hydrogen bonds include sodium dodecyl sulfate (SDS) and urea. In embodiments, reagents that promote hydrogen bond formation are used to T t This is adjusted. In the embodiment, a reagent that promotes hydrophobic interaction is used to T t This is regulated. Trifluoroethanol promotes both hydrophobic interactions and hydrogen bond formation, and T t This reagent causes a decrease in [value].
[0151] In this embodiment, the concentrations of ELP and / or RLP are adjusted, T t This can be adjusted. In this embodiment, if the ELP and / or RLP concentrations are high, T t The concentration decreases. In this embodiment, when the ELP and / or RLP concentrations are low, T t It will rise.
[0152] In addition, by using the adjustment of pH, light, and ion concentration, T t It can be adjusted.
[0153] In the embodiment, pH adjustment is performed by adjusting the number (e.g., addition or removal) and properties (e.g., positive or negative charge) of charged amino acids (e.g., histidine, lysine, arginine, glutamic acid, aspartic acid, ornithine, or other unnatural charged amino acids). t This will allow for adjustments.
[0154] In embodiments, the ELP and / or RLP described herein are block copolymers. A block copolymer comprises two or more sequence domains or blocks, the two or more blocks having different properties. Non-limiting examples of properties that can be tuned include hydrophilicity, hydrophobicity, polarity, and secondary structure. In embodiments, the block copolymer is an amphiphilic material and comprises, for example, at least one hydrophobic block and at least one hydrophilic block.
[0155] In embodiments, the ELP and / or RLP described herein are formed in various forms. Non-limiting examples of forms include spherical aggregates, micelles, vesicles, fibrils, nanofibrils, nanotubes, and hydrogels. In embodiments, the RLP and / or ELP described herein form various forms after the addition of environmental factors. In embodiments, the RLP and / or ELP described herein change from one form to another after the addition of environmental factors. In embodiments, the RLP and / or ELP described herein change from one form to another after the addition of nucleic acids.
[0156] In the embodiment, the addition of environmental factors causes a phase transition of RLP and / or ELP. In the embodiment, during the phase transition of RLP and / or ELP, RLP and / or ELP are transformed from one form to another.
[0157] In this embodiment, a high-density droplet is formed by a phase transition of RLP and / or ELP.
[0158] In the embodiment, the polypeptide having phase behavior includes an amino acid sequence selected from Table 2. In the embodiment, the polypeptide having phase behavior has at least 90% identity with SEQ ID NO: 56 and one of the polypeptides 262-264. In the embodiment, the polypeptide having phase behavior has at least 90% identity with the polypeptide of SEQ ID NO: 56. In the embodiment, the polypeptide having phase behavior includes SEQ ID NO: 56 and one of the polypeptides 262-264. In the embodiment, the polypeptide having phase behavior includes SEQ ID NO: 56.
[0159] Table 6
[0160] Table 7
[0161] Table 8
[0162] Table 9
[0163] Table 10
[0164] Table 11
[0165] In this embodiment, the fusion protein is 1-500, 1-450, 1-400, 1-350, 1-300, 1-250, 1-200, 1-150, 1-100, 1-95, 1-90, 1-85, 1-80, 1-75, 1-70, 1-65, 1-60, 1-55, 1-50, 1-45, 1-40, 1-35, 1-30 The fusion protein comprises various polypeptides having phase behavior, numbered 1-25, 1-20, 1-15, 1-10, 1-9, 1-8, 1-7, 1-6, 1-5, 1-4, 1-3, 1-2, 2-10, 2-9, 2-8, 2-7, 2-6, 2-5, 2-4, 2-3, 3-10, 4-10, 5-10, 6-10, 7-10, or 8-10. In embodiments, the fusion protein comprises at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten various polypeptides having phase behavior. For example, the fusion protein may comprise a first polypeptide having phase behavior and a second polypeptide having phase behavior. In the embodiment, the fusion protein comprises a third, fourth, fifth, sixth, seventh, eighth, ninth, or tenth polypeptide having phase behavior.
[0166] In an embodiment, a fusion protein containing an amino acid sequence selected from any one of SEQ ID NOs: 1-60 and 217 further contains up to 10, 15, 20, or 25 additional N-terminal and / or C-terminal amino acids. In an embodiment, a fusion protein containing any one of SEQ ID NOs: 1-60 and 217 also contains an additional N-terminal methionine. In an embodiment, a fusion protein containing any one of SEQ ID NOs: 1-60 and 217 also contains an additional C-terminal glycine. In an embodiment, a fusion protein containing any one of SEQ ID NOs: 56 and 262-264 further contains up to 10, 15, 20, or 25 additional N-terminal and / or C-terminal amino acids. In an embodiment, a fusion protein containing any one of SEQ ID NOs: 56 and 262-264 also contains an additional N-terminal methionine. In an embodiment, a fusion protein containing any one of SEQ ID NOs: 56 and 262-264 also contains an additional C-terminal glycine. In the embodiments, the fusion protein containing the amino acid sequence selected from SEQ ID NO: 56 further contains up to 10, up to 15, up to 20, or up to 25 additional N-terminal and / or C-terminal amino acids. In the embodiments, the fusion protein containing the amino acid sequence of SEQ ID NO: 56 also contains an additional N-terminal methionine. In the embodiments, the fusion protein containing the amino acid sequence of SEQ ID NO: 56 also contains an additional C-terminal glycine.
[0167] In the embodiment, the fusion protein contains polypeptide repeat units. In the embodiment, there are 5 to 500 polypeptide repeat units (including all ranges and values in between). In this embodiment, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 14 3, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 2 01, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258,259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289 ,290,291,292,293,294,295,296,297,298,299,300,301,302,303,304,305,306,307,308,309,310,311,312,313,314,315,316,317,318,319,32 0, 321, 322, 323, 324, 325, 326, 327, 328, 329, 330, 331, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 349, 350, 3 51, 352, 353, 354, 355, 356, 357, 358, 359, 360, 361, 362, 363, 364, 365, 366, 367, 368, 369, 370, 371, 372, 373, 374, 375, 376, 377, 378, 379, 380, 381, 3 82, 383, 384, 385, 386, 387, 388, 389, 390, 391, 392, 393, 394, 395, 396, 397, 398, 399, 400, 401, 402, 403, 404, 405, 406, 407, 408, 409, 410, 411, 412, 413, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428, 429, 430, 431, 432, 433, 434, 435, 436, 437, 438, 439, 440, 441, 442, 443 There are 444, 445, 446, 447, 448, 449, 450, 451, 452, 453, 454, 455, 456, 457, 458, 459, 460, 461, 462, 463, 464, 465, 466, 467, 468, 469, 470, 471, 472, 473, 474, 475, 476, 477, 478, 479, 480, 481, 482, 483, 484, 485, 486, 487, 488, 489, 490, 491, 492, 493, 494, 495, 496, 497, 498, 499, or 500 polypeptide repeat units.
[0168] In the embodiment, the fusion protein has the same amino acid composition as ELP and / or RLP, but does not contain repeats. In the embodiment, the fusion protein contains an amino acid sequence that is approximately 80%, approximately 81%, approximately 82%, approximately 83%, approximately 84%, approximately 85%, approximately 86%, approximately 87%, approximately 88%, approximately 89%, approximately 90%, approximately 91%, approximately 92%, approximately 93%, approximately 94%, approximately 95%, approximately 96%, approximately 97%, approximately 98%, approximately 99%, or approximately 100% identical to ELP and / or RLP. In embodiments, the fusion protein contains an amino acid composition that is 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 100% identical to ELP and / or RLP. In embodiments, the fusion protein contains a hydrophobic amino acid composition that is approximately 80%, approximately 81%, approximately 82%, approximately 83%, approximately 84%, approximately 85%, approximately 86%, approximately 87%, approximately 88%, approximately 89%, approximately 90%, approximately 91%, approximately 92%, approximately 93%, approximately 94%, approximately 95%, approximately 96%, approximately 97%, approximately 98%, approximately 99%, or approximately 100% identical to that of ELP and / or RLP.
[0169] In this embodiment, the polypeptide having phase behavior has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one polypeptide of SEQ ID NOs: 1-60, 217, or 262-264. In this embodiment, the polypeptide having phase behavior has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one polypeptide of SEQ ID NO: 56 and 262-264. In this embodiment, the polypeptide having phase behavior has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polypeptide of Sequence ID No. 56.
[0170] In the embodiment, the polypeptide having phase behavior includes a non-repeating unstructured polypeptide. In the embodiment, the non-repeating unstructured polypeptide has an amino acid sequence comprising at least 50 amino acids. In the embodiment, the non-repeating unstructured polypeptide has an amino acid sequence comprising at least 50, at least 60, at least 70, at least 80, at least 90, or at least 100 amino acids. In the embodiment, the polypeptide having phase behavior includes a P and G motif comprising a plurality of P residues and a plurality of G residues. In the embodiment, the P and G motif comprises at least about 10% proline and at least about 20% glycine. In the embodiment, the sequence of the non-repeating unstructured polypeptide is at least about 10% (e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, or 80%) proline and at least about 20% (e.g., at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%) glycine. In the embodiment, the non-repeating unstructured polypeptide has a sequence containing at least about 40% of amino acids selected from the group consisting of valine, alanine, leucine, lysine, threonine, isoleucine, tyrosine, serine, and phenylalanine.
[0171] In the embodiment, the non-repeating unstructured polypeptide comprises a sequence that does not contain three consecutive identical amino acids, any 5 to 10 amino acid subsequences do not appear more than once in the non-repeating unstructured polypeptide, the non-repeating unstructured polypeptide comprises a subsequence that begins with proline and ends with proline, and the subsequence further comprises at least one glycine.
[0172] In embodiments, the ELP and / or RLP described herein are expressed as components of a phase-behaving polypeptide. In embodiments, the phase-behaving polypeptide is expressed in bacterial or mammalian cells. In embodiments, the phase-behaving polypeptide is expressed in Escherichia coli. In embodiments, the phase-behaving polypeptide is expressed in insect cells (e.g., Sf9 cells). In embodiments, the sequence of the non-repeating unstructured polypeptide is at least about 10% (e.g., at least 10%, at least 20%, at least 30%, at least 40%) proline, at least 20% (e.g., at least 20%, at least 30%, at least 40%, or at least 50%) glycine, and at least 40% (e.g., at least 40%, at least 50%, at least 60%, or at least 70%) amino acids selected from the group consisting of valine, alanine, leucine, lysine, threonine, isoleucine, tyrosine, serine, and phenylalanine.
[0173] In embodiments, the ELP and / or RLP described herein are expressed as components of a fusion protein. In embodiments, the fusion protein is expressed in bacterial or mammalian cells. In embodiments, the fusion protein is expressed in Escherichia coli. In embodiments, the fusion protein is expressed in insect cells (e.g., Sf9 cells).
[0174] In the embodiments, the non-repeating unstructured polypeptide does not contain three consecutive identical amino acids. In the embodiments, the non-repeating unstructured polypeptide includes a subsequence (e.g., a fragment of the non-repeating unstructured polypeptide) that appears only once in the non-repeating unstructured polypeptide sequence. In the embodiments, the non-repeating unstructured polypeptide includes a subsequence that begins and ends with proline. In the embodiments, the non-repeating unstructured polypeptide includes a subsequence that contains at least one glycine molecule.
[0175] In the embodiment, the polypeptide having phase behavior includes a signal peptide. In the embodiment, the signal peptide includes an amino acid sequence selected from any one of SEQ ID NOs. 218-220. In the embodiment, the signal peptide is identical to any one of SEQ ID NOs. 218-220 by about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or up to about 100%. In the embodiment, the signal peptide is a polypeptide from Table 3.
[0176] [Table 12]
[0177] In the embodiment, the polypeptide having phase behavior includes the amino acid sequence (GVGVPGLGVPGVGVPGLGVPGVGVP)m (SEQ ID NO: 33), where m is 16. In the embodiment, the polypeptide having phase behavior includes the amino acid sequence of SEQ ID NO: 34. In the embodiment, the fusion protein includes the amino acid sequence of SEQ ID NO: 38. In the embodiment, the fusion protein includes the amino acid sequence of SEQ ID NO: 43. In the embodiment, the fusion protein includes the amino acid sequence of SEQ ID NO: 47. In the embodiment, the fusion protein includes the amino acid sequence of SEQ ID NO: 60.
[0178] Linker between polypeptides with phase behavior and NBPs In the embodiment, the fusion protein includes a linker. In the embodiment, the NBP is bound to the polypeptide having phase behavior via the linker. In the embodiment, any linker that does not interfere with the function of the fusion protein can be used. In the embodiment, the fusion includes one or more NBPs, one or more linkers, and one or more polypeptides having phase behavior, from the C-terminus to the N-terminus. In the embodiment, the fusion includes one or more NBPs, one or more linkers, and one or more polypeptides having phase behavior, from the N-terminus to the C-terminus.
[0179] In the embodiment, the linker links the NBP to the polypeptide having phase behavior. In the embodiment, the linker enables cooperative interaction between the polypeptide having phase behavior and the NBP. In the embodiment, the linker is a peptide. In the embodiment, the linker maintains the phase behavior of the polypeptide having phase behavior. In the embodiment, the linker is the T of the polypeptide having phase behavior t Maintain the structure of the NBP. In the embodiment, the linker preserves the structure of the NBP. In the embodiment, the linker contains 1 to 50 amino acids. In the embodiment, the linker contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acids.
[0180] In this embodiment, the rigidity of the linker is improved by including proline in the amino acid sequence of the linker.
[0181] In some embodiments, the flexibility of the linker is improved by including small polar amino acids, including threonine, serine, and glycine.
[0182] In embodiments, the linker can take on various secondary structures, including, but not limited to, α-helices, β-chains, and random coils. In embodiments, the linker takes on an α-helical structure, (EAAAK) n This sequence contains the amino acid repeats of (SEQ ID NO: 143), where n is the number of repeats, i.e., an integer in the range of 1 to 20 including the endpoint.
[0183] In this embodiment, the linker is (G4S) nIt consists of (SEQ ID NO: 144), where n is an integer from 1 to 30 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30). In embodiments, the polypeptide linker has a repeat of (SGGG)n (SEQ ID NO: 145), where n is an integer from 1 to 50 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20). In embodiments, the polypeptide linker is (GGGS) n It has iterations of (sequence number 146), where n is an integer from 1 to 20 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20).
[0184] In the embodiment, the linker has the amino acid sequence KESGSVSSEQLAQFRSLD (SEQ ID NO: 147). In the embodiment, the linker has the amino acid sequence EGKSSGSGSESKST (SEQ ID NO: 148). In the embodiment, the linker contains only glycine.
[0185] In the embodiment, the peptide linker includes a protease cleavage site. In the embodiment, the protease cleavage site is a furin cleavage site.
[0186] In this embodiment, the polypeptide linker is poly-(Gly) n The linker is a linker in the formula where n is an integer from 1 to 30 (for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30) (SEQ ID NO: 149). In other embodiments, the linker is selected from the group consisting of dipeptides, tripeptides, and quadripeptides. In embodiments, the linker is a dipeptide selected from the group consisting of alanine-serine (AS), leucine-glutamic acid (LE), and serine-arginine (SR).
[0187] In this embodiment, the linker is selected from GKSSGSGSESKS (SEQ ID NO: 150), GTSGSGKSSEGKG (SEQ ID NO: 151), GTSGSGKSSEGSGSTKG (SEQ ID NO: 152), GTSGSGKPGSGEGSTKG (SEQ ID NO: 153), EGKSSGSGSESKEF (SEQ ID NO: 154), SRSSG (SEQ ID NO: 155), and SGSSC (SEQ ID NO: 156).
[0188] In the embodiment, the linker is a self-cleaving peptide. In the embodiment, the self-cleaving peptide is a 2A peptide. A 2A peptide is a class of peptides with an amino acid length of 18 to 22 that induce ribosome skipping during intracellular protein translation. In the embodiment, the 2A peptide is a T2A peptide having the amino acid sequence EGRGSLLTCGDVEENPGP (SEQ ID NO: 157), a P2A peptide having the amino acid sequence ATNFSLLKQAGDVEENPGP (SEQ ID NO: 158), an E2A peptide having the amino acid sequence QCTNYALLKLAGDVESNPGP (SEQ ID NO: 159), or an F2A peptide having the amino acid sequence VKQTLNFDLLKLAGDVESNPGP (SEQ ID NO: 160). In the embodiment, the 2A peptide has at least 80%, at least 85%, at least 90%, at least 95%, or at least 98% identity with any one of SEQ ID NOs. 157 to 160. In this embodiment, the 2A peptide further contains GSG (SEQ ID NOs: 161-164) at its N-terminus.
[0189] In one embodiment, the linker includes one of the amino acid sequences from SEQ ID NOs. 143 to 216. In another embodiment, the linker has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with respect to an amino acid sequence selected from one of SEQ ID NOs. In this embodiment, the linker has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with respect to an amino acid sequence selected from any one of SEQ ID NOs: 143-216 and 261. In one embodiment, the linker has an amino acid sequence that is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the polypeptide of SEQ ID NO: 261.
[0190] In the embodiment, the linker is a polypeptide shown in Table 4.
[0191] [Table 13]
[0192] [Table 14]
[0193] In the embodiment, the linker is a chemical linker. In the embodiment, the chemical linker is selected from the group consisting of carbohydrate linkers, lipid linkers, fatty acid linkers, and polyether linkers.
[0194] In the embodiment, the linker is a direct covalent bond between an amino acid residue of a phase-behaving polypeptide and an NBP. In the embodiment, the fusion protein comprises a phase-behaving polypeptide and an NBP. In the embodiment, the amino acid residue of the phase-behaving polypeptide is covalently bonded to an amino acid in Bs-CspB. In the embodiment, the fusion protein further comprises one or more linkers as described herein. In the embodiment, the fusion protein comprises a phase-behaving polypeptide, a linker, and an NBP, from the N-terminus to the C-terminus. In the embodiment, the fusion protein comprises an NBP, a linker, and a phase-behaving polypeptide, from the N-terminus to the C-terminus. In the embodiment, the fusion protein comprises a phase-behaving polypeptide, a linker, and an NBP, from the C-terminus to the N-terminus. In the embodiment, the fusion protein comprises an NBP, a linker, and a phase-behaving polypeptide, from the C-terminus to the N-terminus. In the embodiment, the fusion protein comprises one or more phase-behaving polypeptides, one or more linkers, and one or more NBPs, from the N-terminus to the C-terminus. In the embodiment, the fusion protein comprises one or more NBPs, one or more linkers, and one or more polypeptides having phase behaviors, arranged from the N-terminus to the C-terminus.
[0195] A fusion protein construct comprising Bs-CspB and a polypeptide exhibiting phase behavior. In some embodiments, the fusion protein described herein comprises Bs-CspB and a polypeptide having phase behavior. In embodiments, the fusion protein comprises, from the N-terminus to the C-terminus, a polypeptide having phase behavior, a linker, and Bs-CspB. In embodiments, the fusion protein comprises, from the N-terminus to the C-terminus, Bs-CspB, a linker, and a polypeptide having phase behavior. In embodiments, the fusion protein comprises, from the C-terminus to the N-terminus, a polypeptide having phase behavior, a linker, and Bs-CspB. In embodiments, the fusion protein comprises, from the C-terminus to the N-terminus, Bs-CspB, a linker, and a polypeptide having phase behavior.
[0196] In one embodiment, Bs-CspB contains an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polypeptide of SEQ ID NO: 90. In this embodiment, the polypeptide having phase behavior has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one polypeptide of SEQ ID NO: 56 and 262-264. In this embodiment, the polypeptide having phase behavior has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polypeptide of Sequence ID No. 56.
[0197] In one embodiment, the fusion protein has at least 90% identity to the polypeptide of SEQ ID NO: 90, and at least 90% identity to one of the polypeptides of SEQ ID NO: 56 and 262-264. In another embodiment, the fusion protein has at least 90% identity to the polypeptide of SEQ ID NO: 90 and at least 90% identity to the polypeptide of SEQ ID NO: 56. In another embodiment, the fusion protein comprises SEQ ID NO: 90 and one of SEQ ID NO: 56 and 262-264. In yet another embodiment, the fusion protein comprises SEQ ID NO: 90 and SEQ ID NO: 56.
[0198] In embodiments, the fusion protein further includes a linker. In embodiments, the linker includes an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 261.
[0199] In an embodiment, the fusion protein contains an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with either one of the polypeptides of SEQ ID NO: 225 or 226. In an embodiment, the fusion protein has at least 90% identity with the polypeptide of SEQ ID NO: 225. In an embodiment, the fusion protein contains SEQ ID NO: 225 or 226. In an embodiment, the fusion protein contains SEQ ID NO: 225. In an embodiment, the fusion protein contains a start codon. In an embodiment, the start codon contains an N-terminal methionine. In the embodiment, the fusion protein contains an N-terminal methionine. In the embodiment, the fusion protein containing an N-terminal methionine has the amino acid sequence of SEQ ID NO: 225. In the embodiment, the fusion protein not containing an N-terminal methionine has the amino acid sequence of SEQ ID NO: 226.
[0200] Table 5 shows an example sequence of the aforementioned fusion protein.
[0201] [Table 15]
[0202] [Table 16]
[0203] How to use fusion proteins This disclosure provides a fusion protein and a method of using the same. In embodiments, the fusion protein comprises a nucleic acid-binding protein (NBP) that binds to a nucleic acid and a polypeptide having phase behavior. In embodiments, the NBP comprises Bs-CspB. In embodiments, the composition containing the nucleic acid is free from one or more contaminants.
[0204] In the embodiment, a method for purifying nucleic acids includes contacting the nucleic acid with a fusion protein, where the nucleic acid binds to the fusion protein to form a complex; the size of the complex is increased by a first environmental factor; the complex is separated from at least one contaminant based on its size; and the nucleic acid is separated from the fusion protein by a second environmental factor.
[0205] In the embodiment, a method for purifying single-stranded nucleic acid is as follows: contacting a fusion protein with a composition containing single-stranded nucleic acid and at least one contaminant, wherein the fusion protein binds to the single-stranded nucleic acid to form a complex; adding a first environmental factor to the composition containing the complex to increase the size of the complex; separating the complex from at least one contaminant; and contacting the complex with a second environmental factor to separate the single-stranded nucleic acid from the fusion protein, thereby forming a product containing single-stranded nucleic acid. In the embodiment, the single-stranded nucleic acid includes ssRNA, ssDNA, or both. In the embodiment, the single-stranded nucleic acid includes ssRNA. In the embodiment, at least one contaminant includes double-stranded RNA.
[0206] In the embodiment, the method includes the step of separating the complex from at least one contaminant based on size.
[0207] In the embodiment, the method for purifying nucleic acids includes the step of contacting the nucleic acid with a fusion protein; where the nucleic acid binds to the fusion protein to form a complex; the size of the complex increases; the complex is separated from at least one contaminant based on its size; the nucleic acid is separated from the fusion protein by an environmental factor, thereby forming a product containing the nucleic acid. In the embodiment, the nucleic acid includes single-stranded nucleic acid. In the embodiment, the nucleic acid includes ssRNA. In the embodiment, the contaminant includes double-stranded nucleic acid. In the embodiment, the contaminant includes dsRNA.
[0208] In the embodiment, a method for removing a contaminant from a composition containing nucleic acid comprises contacting the contaminant with a fusion protein; where the contaminant binds to the fusion protein to form a complex; the size of the complex is increased by a first environmental factor; the complex is separated from the nucleic acid based on its size; and the contaminant is separated from the fusion protein by a second environmental factor, thereby forming a product containing nucleic acid. In the embodiment, the nucleic acid includes single-stranded nucleic acid. In the embodiment, the nucleic acid includes ssRNA. In the embodiment, the contaminant includes double-stranded nucleic acid. In the embodiment, the contaminant includes dsRNA.
[0209] In some embodiments, a method for separating a first nucleic acid from a second nucleic acid includes the steps of: contacting the first nucleic acid with a first fusion protein and contacting the second nucleic acid with a second fusion protein; the first nucleic acid binding to the first fusion protein to form a first complex; the second nucleic acid binding to the second fusion protein to form a second complex; and separating the first nucleic acid from the second nucleic acid by applying an environmental factor. In embodiments, the first nucleic acid includes a single-stranded nucleic acid. In embodiments, the second nucleic acid includes a double-stranded nucleic acid.
[0210] Furthermore, this specification also provides a method for bringing one nucleic acid close to another. In an embodiment, the method for bringing a first nucleic acid close to a second nucleic acid includes the steps of bringing the first nucleic acid into contact with a first fusion protein and the second nucleic acid into contact with a second fusion protein; where the first nucleic acid binds to the first fusion protein to form a first complex; the second nucleic acid binds to the second fusion protein to form a second complex; and then an environmental factor brings the first and second complexes close to each other. In an embodiment, the method described herein brings the first and second nucleic acids close to each other within about 10 μm, about 5 μm, about 1 μm, about 900 nm, about 800 nm, about 700 nm, about 600 nm, about 500 nm, about 400 nm, about 300 nm, about 200 nm, about 100 nm, about 10 nm, about 1 nm, about 0.5 nm, or about 0.1 nm. In the embodiment, the first nucleic acid includes a single-stranded nucleic acid. In the embodiment, the second nucleic acid includes a double-stranded nucleic acid.
[0211] In some embodiments, the method described herein utilizes a fusion protein comprising NBP and a polypeptide having phase behavior. In some embodiments, the method described herein utilizes a fusion protein comprising Bs-CspB and a polypeptide having phase behavior. In some embodiments, the method described herein utilizes two or more different fusion proteins.
[0212] In embodiments, the methods described herein involve the formation of complexes. In embodiments, the methods described herein involve the formation of one or more complexes. In embodiments, the methods described herein involve the formation of one, two, three, four, five, or more complexes. The complexes may be referred to as the "first complex" or the "second complex," and so on.
[0213] In some embodiments, the complex comprises a fusion protein and a nucleic acid. In some embodiments, the complex comprises a fusion protein and a contaminant. In some embodiments, the complex comprises a fusion protein and a second protein, such as an enzyme substrate, a metabolite, or a ligand (e.g., a ligand that binds to a cell receptor).
[0214] In the embodiment, the components of the complex (e.g., the fusion protein and nucleic acid) bind to each other. In the embodiment, this binding is reversible. Reversible binding means that the complex can dissociate, for example, separate into individual components. For example, if a complex is reversibly formed between the fusion protein and the nucleic acid, the fusion protein and nucleic acid can subsequently dissociate. In the embodiment, dissociation is induced by environmental factors. In the embodiment, reversible binding makes it possible to separate the nucleic acid from the fusion protein. In the embodiment, reversible binding makes it possible to separate contaminants from the fusion protein. In the embodiment, reversible binding makes it possible to separate other molecules from the fusion protein. In the embodiment, the nucleic acid includes single-stranded nucleic acid. In the embodiment, the nucleic acid includes ssRNA. In the embodiment, the contaminant includes double-stranded nucleic acid. In the embodiment, the contaminant includes dsRNA.
[0215] In the embodiments, the reversible bond is a non-covalent bond, meaning no covalent bond is formed between the interacting components of the complex (e.g., between the fusion protein and the nucleic acid). In the embodiments, the fusion protein and nucleic acid are formed by non-covalent interactions. Non-limiting examples of non-covalent interactions include dipole forces, van der Waals forces, London dispersion forces, hydrogen bonds, hydrophobic interactions, and electrostatic interactions. In the embodiments, the non-covalent bond is disrupted by the addition of environmental factors. In the embodiments, the nucleic acid includes single-stranded nucleic acid. In the embodiments, the nucleic acid includes ssRNA.
[0216] In the embodiment, the bond between the fusion protein and the target molecule (e.g., nucleic acid) is a covalent bond. In the embodiment, the covalent bond between the fusion protein and the nucleic acid can be cleaved using, for example, a nuclease and / or protease. In the embodiment, the nucleic acid includes single-stranded nucleic acid. In the embodiment, the nucleic acid includes ssRNA.
[0217] In embodiments, the size of the complex described herein increases after the application of environmental factors. In embodiments, the size of the complex formed between the fusion protein and the nucleic acid increases. In embodiments, the size of the initial complex increases as a result of aggregation of multiple complexes. In embodiments, multiple complexes aggregate by self-assembly of the fusion protein. In embodiments, multiple complexes aggregate upon application of environmental factors. In embodiments, the size increase is stabilized by non-covalent interactions between multiple fusion proteins. In embodiments, the size increase is stabilized by non-covalent interactions between polypeptides having phase behavior. In embodiments, non-covalent interactions are dipole forces, van der Waals forces, London dispersion forces, hydrogen bonds, hydrophobic interactions, and / or electrostatic interactions. In embodiments, the nucleic acid includes single-stranded nucleic acid. In embodiments, the nucleic acid includes ssRNA.
[0218] In embodiments, the method of the present disclosure achieves the formation of multiple complexes in a mixture. In embodiments, the size of all complexes increases. In embodiments, the size of some complexes increases while the size of other complexes remains constant. In embodiments, the size of one complex increases while the size of other complexes remains constant.
[0219] In one embodiment, the size of the initial complex is at least about 2 times, at least about 5 times, at least about 10 times, at least about 15 times, at least about 20 times, at least about 25 times, at least about 30 times, at least about 35 times, at least about 40 times, at least about 45 times, at least about 50 times, at least about 55 times, at least about 60 times, at least about 65 times, at least about 70 times, at least about 75 times, at least about 80 times, at least about 85 times, at least about 90 times, at least about 95 times, and at least about 100 times or more. In one embodiment, the size of the initial complex increases by at least about 2 times. In one embodiment, the size of the initial complex increases by at least about 5 times. In one embodiment, the size of the initial complex increases by at least about 10 times. In one embodiment, the size of the initial complex increases by at least about 25 times.
[0220] As used herein, the expression “size increase” may refer to an increase in the diameter of the composite or an increase in the mass of the composite. In embodiments, the size increase is an increase in the molar mass of the composite. In embodiments, the size increase is an increase in the hydrodynamic radius of the composite.
[0221] In the embodiment, the increase in the size of the complex can be visually observed with the naked eye. For example, an increase in the size of the complex may result in changes in the color, transparency, and viscosity of the composition containing the complex, and / or changes in the solubility of the complex (e.g., precipitation from solution), and such changes are observable by humans without the use of any special equipment.
[0222] In embodiments, those skilled in the art can measure the size increase of the complex according to methods known in the art. In embodiments, the size increase of the complex can be measured using techniques selected from the group consisting of X-ray scattering, small-angle X-ray scattering, wide-angle X-ray scattering, dynamic light scattering, analytical ultracentrifugation, size exclusion chromatography, and photon correlation spectroscopy.
[0223] In the embodiment, the enlarged complex is separated from the contaminant. In the embodiment, the enlarged complex containing nucleic acid and fusion protein is separated from the contaminant. In the embodiment, the enlarged complex containing contaminant and fusion protein is separated from the nucleic acid composition. In the embodiment, the enlarged first complex containing the first nucleic acid and the first fusion protein is separated from the second complex containing the second nucleic acid and the second fusion protein. In the embodiment, the enlarged complex containing contaminant and fusion protein is separated from the nucleic acid composition. In the embodiment, the enlarged first complex containing the first nucleic acid and the first fusion protein is separated from the second complex containing the second nucleic acid and the second fusion protein. In the embodiment, the first nucleic acid contains single-stranded nucleic acid. In the embodiment, the second nucleic acid contains single-stranded nucleic acid. In the embodiment, the single-stranded nucleic acid contains ssRNA. In the embodiment, the contaminant contains double-stranded nucleic acid. In the embodiment, the contaminant contains dsRNA.
[0224] In this embodiment, the separation of the complex from the contaminant is visually observable with the naked eye.
[0225] In embodiments, the separation of the complex from contaminants is performed based on size. In embodiments, size-based separation is carried out using techniques selected from the group consisting of tangential flow filtration (TFF), analytical ultracentrifugation, membrane chromatography, high-performance liquid chromatography, size exclusion chromatography, membrane chromatography, total filtration, ultrasonic separation, centrifugation, countercurrent centrifugation, and high-performance protein liquid chromatography. In embodiments, the complex is separated from at least one contaminant based on size using tangential flow filtration. In embodiments, the complex is separated from at least one contaminant based on size using centrifugation. In embodiments, the complex is separated from at least one contaminant by applying relative centrifugal force (RCF) from about 100 to about 16,000 RCF, for example, about 500 to about 16,000 RCF, or about 1,000 RCF to 16,000 RCF. In the embodiment, a relative centrifugal force (RCF) of at least 500, for example, at least about 500 RCF, at least about 600 RCF, at least about 700 RCF, at least about 800 RCF, at least about 900 RCF, at least about 1000 RCF, at least about 2000 RCF, at least about 3000 RCF, at least about 3500 RCF, at least about 4000 RCF, at least about 5000 RCF, at least about 6000 RCF, at least about 7000 RCF, at least about 8000 RCF, and less Apply at least approximately 9,000 RCF, at least approximately 10,000 RCF, at least approximately 11,000 RCF, at least approximately 12,000 RCF, at least approximately 13,000 RCF, at least approximately 14,000 RCF, at least approximately 15,000 RCF, at least approximately 16,000 RCF, at least approximately 17,000 RCF, at least approximately 18,000 RCF, at least approximately 19,000 RCF, or at least approximately 20,000 RCF to separate the fusion protein aggregates from at least one contaminant.
[0226] In the embodiment, size-based separation of the complex from contaminants is carried out using TFF. In the embodiment, TFF can be used to separate the complex from at least one contaminant based on size, and this process is also referred to herein as “diafiltration”. Diafiltration includes both a washing step and an elution step. Washing removes contaminants contained in the composition containing the complex. Elution separates the purified nucleic acid from the fusion protein. In the embodiment, the complex is concentrated using TFF. In the embodiment, TFF can be used to increase the concentration of the complex in the composition, and this process is also referred to herein as “concentration”.
[0227] Tangential flow filtration uses both microfiltration and ultrafiltration membranes to separate molecules. Microfiltration membranes typically have pore sizes of 0.1 μm to 10 μm. Ultrafiltration membranes typically have smaller pore sizes than microfiltration membranes, with pore sizes of 0.001 μm to 0.1 μm. In embodiments, membranes with pore sizes of approximately 0.001 μm to approximately 10 μm are used in the methods of this disclosure. In the embodiment, the film has pore sizes of approximately 0.001 μm, approximately 0.01 μm, approximately 0.05 μm, approximately 0.1 μm, approximately 0.2 μm, approximately 0.3 μm, approximately 0.4 μm, approximately 0.5 μm, approximately 0.6 μm, approximately 0.7 μm, approximately 0.8 μm, approximately 0.9 μm, approximately 1.0 μm, approximately 2 μm, approximately 3 μm, approximately 4 μm, approximately 5 μm, approximately 6 μm, approximately 7 μm, approximately 8 μm, approximately 9 μm, or approximately 10 μm (including all values and ranges in between). In the embodiment, the film has a pore size of approximately 0.1 μm. In the embodiment, the film has a pore size of approximately 0.2 μm.
[0228] In the embodiment, the membrane is formed from hydrophilic poly(vinylidene fluoride) (PVDF), polyethersulfone (PES), cellulose phosphate, diethylaminoethylcellulose, polysulfone, regenerated cellulose, nylon, cellulose nitrate, cellulose acetate, PEGylated PES, modified polyethersulfone, and sulfonated PES, or modified derivatives of these materials.
[0229] In TFF, the membrane is positioned tangentially to the flow of the fluid mixture, so that the fluid mixture flows tangentially along the first side of the membrane. At the same time, the fluid medium is positioned in contact with the second surface of the membrane. Transmembrane pressure is the force that pushes the fluid into the membrane and carries permeable molecules.
[0230] In the embodiment, size-based separation of the complex from the contaminants is performed using a TFF at a transmembrane pressure of approximately 0.1 bar to approximately 3 bar. Assuming the current state, the intermembrane pressure is approximately 0.1 bar, approximately 0.2 bar, approximately 0.3 bar, approximately 0.4 bar, approximately 0.5 bar, approximately 0.6 bar, approximately 0.7 bar, approximately 0.8 bar, approximately 0.9 bar, approximately 1.0 bar, approximately 1.1 bar, approximately 1.2 bar, approximately 1.3 bar, approximately 1.4 bar, approximately 1.5 bar, approximately 1.6 bar, approximately 1.7 bar, approximately 1.8 bar, approximately 1.9 bar, approximately 2.0 bar, approximately 2.1 bar, approximately 2.2 bar, approximately 2.3 bar, approximately 2.4 bar, approximately 2.5 bar, approximately 2.6 bar, approximately 2.7 bar, approximately 2.8 bar, approximately 2.9 bar, or approximately 3.0 bar (including all values and ranges in between). Assuming the current state, the intermembrane pressure is approximately 1.5 bar.
[0231] In embodiments, the cross-flow rate is adjusted to improve the separation of the complex described herein from contaminants. The cross-flow rate is the flow rate of the solution through the membrane via the supply channel. This provides a sweeping force for molecules that may restrict the flow of the filtrate. In embodiments, the cross-flow rate is approximately 500 L / m² / h to approximately 2,000 L / m² / h. In one embodiment, the cross-flow rate is approximately 500 L / m² / h, 600 L / m² / h, 700 L / m² / h, 800 L / m² / h, 900 L / m² / h, 1,000 L / m² / h, 1,100 L / m² / h, 1,200 L / m² / h, 1,300 L / m² / h, 1,400 L / m² / h, 1,500 L / m² / h, 1,600 L / m² / h, 1,700 L / m² / h, 1,800 L / m² / h, 1,900 L / m² / h, or 2,000 L / m² / h (including all values and ranges between them). In another embodiment, the cross-flow rate is approximately 960 L / m² / h. In the embodiment, TFF separation is performed by using a membrane that holds a complex containing the fusion protein and nucleic acid while allowing the contaminant to pass through. In the embodiment, a membrane that holds a complex containing the fusion protein and contaminant while allowing the nucleic acid to pass through is used. In the embodiment, a membrane that holds a complex containing the fusion protein and contaminant while allowing the nucleic acid to pass through is utilized. In the embodiment, a membrane that holds a first complex containing the first fusion protein and the first nucleic acid is utilized while allowing a complex containing a second fusion protein and a second nucleic acid to pass through. In the embodiment, the first and second nucleic acids include single-stranded nucleic acids. In the embodiment, the first and second nucleic acids include ssRNA. In the embodiment, the contaminant includes double-stranded nucleic acids. In the embodiment, the contaminant includes dsRNA.
[0232] In one embodiment, the method described herein enables the purification of nucleic acids in amounts of at least about 0.1 kg, at least about 0.2 kg, at least about 0.3 kg, at least about 0.4 kg, at least about 0.5 kg, at least about 0.6 kg, at least about 0.7 kg, at least about 0.8 kg, at least about 0.9 kg, at least about 1 kg, at least about 2 kg, at least about 3 kg, at least about 4 kg, at least about 5 kg, at least about 6 kg, at least about 7 kg, at least about 8 kg, at least about 9 kg, at least about 10 kg, or more (including all values and ranges in between).
[0233] In embodiments, the method described herein is completed in about 0.5 hours to about 24 hours. In embodiments, the method is completed in about 0.5 hours, about 1 hour, about 2 hours, about 3 hours, about 4 hours, about 5 hours, about 6 hours, about 7 hours, about 8 hours, about 9 hours, about 10 hours, about 11 hours, about 12 hours, about 13 hours, about 14 hours, about 15 hours, about 16 hours, about 17 hours, about 18 hours, about 19 hours, about 20 hours, about 21 hours, about 22 hours, about 23 hours, or about 24 hours. In embodiments, the method described herein is completed in about 0.5 hours to about 8 hours. In embodiments, the method disclosed herein is completed in about 2 hours to about 6 hours.
[0234] In embodiments, the method described herein produces a final product containing a proportion of the target nucleic acid based on the total nucleic acid content. In embodiments, the proportion of the target nucleic acid based on the total nucleic acid content may be calculated as the purification yield. In embodiments, the nucleic acid is a single-stranded nucleic acid.
[0235] In one embodiment, the nucleic acid purification yield is at least 70%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%.
[0236] In embodiments, the nucleic acid is purified to at least 70%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% purity.
[0237] In embodiments, the purification yield of single-stranded nucleic acid has a purity of at least 70%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%.
[0238] In embodiments, the single-stranded nucleic acid is purified to at least 70%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% purity.
[0239] In embodiments, the purification yield of ssRNA is at least 70%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%.
[0240] In embodiments, the ssRNA is purified to at least 70%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% purity.
[0241] In embodiments, the purified nucleic acid retains its biological activity and / or structure. In embodiments, the purified nucleic acid has enhanced biological activity. In embodiments, the purified single-stranded nucleic acid retains its biological activity and / or structure. In embodiments, the purified single-stranded nucleic acid has enhanced biological activity. In embodiments, the ssRNA retains its biological activity and / or structure. In embodiments, the purified ssRNA has enhanced biological activity.
[0242] In embodiments, at least 400 g / m per day 2 (1 m of filter membrane 2 grams of nucleic acid per) is purified. In embodiments, at least 400 g / m per day 2 , at least 500 g / m 2 , at least 600 g / m 2 , at least 700 g / m 2 , at least 800 g / m 2 , at least 900 g / m 2 , or at least 1000 g / m 2 of nucleic acid is purified. In embodiments, the nucleic acid is, for example, mRNA, ssRNA, or a virus. In embodiments, the nucleic acid is ssRNA.
[0243] In one embodiment, at least about 150 g / L (grams of nucleic acid per liter of fusion protein) is purified per day. In another embodiment, at least about 150 g / L, at least about 200 g / L, at least about 250 g / L, at least about 300 g / L, at least about 350 g / L, at least about 400 g / L, at least about 450 g / L, at least about 500 g / L, at least about 550 g / L, at least about 600 g / L, at least about 650 g / L, at least about 700 g / L, at least about 750 g / L, at least about 800 g / L, at least about 850 g / L, at least about 900 g / L, at least about 950 g / L, or at least about 1000 g / L is purified per day. In one embodiment, the nucleic acid is, for example, mRNA, ssRNA, or a virus. In another embodiment, the nucleic acid is ssRNA.
[0244] In the embodiments, the product contains approximately 70% to approximately 100% single-stranded nucleic acids. In the embodiments, the product contains at least 70%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% single nucleic acids. In the embodiments, the single-stranded nucleic acids include ssRNA.
[0245] In the embodiments, the final product contains at least one contaminant in an amount of about 10% or less. In the embodiments, the product contains at least one contaminant in an amount of about 0.1%, about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, or about 10%. In the embodiments, the purification method includes the removal of at least about 3-log of the contaminant in the final product compared to the amount of the contaminant in the original composition. In the embodiments, the purification method includes the removal of at least about 1-log, 2-log, 3-log, 4-log, 5-log, 6-log, 7-log, 8-log, 9-log, or 10-log of the contaminant compared to the amount of the contaminant in the original composition. In the embodiments, the purification method includes the removal of at least about 3-log to 10-log of the contaminant compared to the amount of the contaminant in the original composition. In the embodiments, the contaminant includes dsRNA.
[0246] In the embodiment, the fusion protein is present at a concentration of approximately 1 μM to approximately 200 μM (e.g., 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, or 200 μM (including all values and ranges in between)). In the embodiment, the fusion protein is present at a concentration of approximately 30 μM to approximately 50 μM. In the embodiment, the fusion protein is present at a concentration of approximately 40 μM.
[0247] environmental factors In the embodiments, one or more environmental factors are applied to induce a change in the complex containing the fusion protein and nucleic acid. In the embodiments, one or more environmental factors increase the size of the complex containing the fusion protein and nucleic acid. In the embodiments, one or more environmental factors cause aggregation of polypeptides having phase behavior. In the embodiments, one or more environmental factors cause the separation of the fusion protein from the nucleic acid. In the embodiments, one or more environmental factors cause the separation of the fusion protein from contaminants. In the embodiments, one or more environmental factors allow the nucleic acid to retain its native structure, function, and activity.
[0248] In the embodiment, the environmental factor is a change in temperature. In the embodiment, the temperature is raised by approximately 0.5°C, 1°C, 2°C, 3°C, 4°C, 5°C, 6°C, 7°C, 8°C, 9°C, 10°C, 11°C, 12°C, 13°C, 14°C, 15°C, 16°C, 17°C, 18°C, 19°C, 20°C, 21°C, 22°C, 23°C, 24°C, 25°C, 26°C, 27°C, 28°C, 29°C, 30°C, 31°C, 32°C, 33°C, 34°C, 35°C, 36°C, 37°C, 38°C, 39°C, or 40°C. In the embodiment, the temperature can be lowered by approximately 0.5°C, 1°C, 2°C, 3°C, 4°C, 5°C, 6°C, 7°C, 8°C, 9°C, 10°C, 11°C, 12°C, 13°C, 14°C, 15°C, 16°C, 17°C, 18°C, 19°C, 20°C, 21°C, 22°C, 23°C, 24°C, 25°C, 26°C, 27°C, 28°C, 29°C, 30°C, 31°C, 32°C, 33°C, 34°C, 35°C, 36°C, 37°C, 38°C, 39°C, or 40°C.
[0249] In the embodiment, the environmental factor is a change in pH. In the embodiment, the pH is approximately 0.1 units, approximately 0.2 units, approximately 0.3 units, approximately 0.4 units, approximately 0.5 units, approximately 0.6 units, approximately 0.7 units, approximately 0.8 units, approximately 0.9 units, approximately 1.0 units, approximately 1.1 units, approximately 1.2 units, approximately 1.3 units, approximately 1.4 units, approximately 1.5 units, approximately 1.6 units, approximately 1.7 units, approximately 1.8 units, approximately 1.9 units, approximately 2.0 units, approximately 2.1 units, approximately 2.2 units, approximately 2.3 units, approximately 2.4 units, approximately 2.5 units, approximately 2.6 units, approximately 2.7 units, approximately 2.8 units, approximately 2.9 units, and approximately 3.0 units. The rank can be raised by approximately 3.1 units, 3.2 units, 3.3 units, 3.4 units, 3.5 units, 3.6 units, 3.7 units, 3.8 units, 3.9 units, 4.0 units, 4.1 units, 4.2 units, 4.3 units, 4.4 units, 4.5 units, 4.6 units, 4.7 units, 4.8 units, 4.9 units, 5.0 units, 5.1 units, 5.2 units, 5.3 units, 5.4 units, 5.5 units, 5.6 units, 5.7 units, 5.8 units, 5.9 units, or 6.0 units. In the embodiment, the pH is approximately 0.1 units, approximately 0.2 units, approximately 0.3 units, approximately 0.4 units, approximately 0.5 units, approximately 0.6 units, approximately 0.7 units, approximately 0.8 units, approximately 0.9 units, approximately 1.0 units, approximately 1.1 units, approximately 1.2 units, approximately 1.3 units, approximately 1.4 units, approximately 1.5 units, approximately 1.6 units, approximately 1.7 units, approximately 1.8 units, approximately 1.9 units, approximately 2.0 units, approximately 2.1 units, approximately 2.2 units, approximately 2.3 units, approximately 2.4 units, approximately 2.5 units, approximately 2.6 units, approximately 2.7 units, approximately 2.8 units, approximately 2.9 units, and approximately 3.0 units. The rank can be lowered by approximately 3.1 units, 3.2 units, 3.3 units, 3.4 units, 3.5 units, 3.6 units, 3.7 units, 3.8 units, 3.9 units, 4.0 units, 4.1 units, 4.2 units, 4.3 units, 4.4 units, 4.5 units, 4.6 units, 4.7 units, 4.8 units, 4.9 units, 5.0 units, 5.1 units, 5.2 units, 5.3 units, 5.4 units, 5.5 units, 5.6 units, 5.7 units, 5.8 units, 5.9 units, or 6.0 units.
[0250] In embodiments, the environmental factor is a change in ionic strength. In embodiments, the change in ionic strength is brought about by increasing the concentration of the salt. In embodiments, the change in ionic strength is brought about by decreasing the concentration of the salt. Non-limiting examples of salts include sodium chloride, potassium chloride, ammonium chloride, sodium acetate, sodium citrate, copper sulfate, sodium iodide, ammonium sulfate, and sodium sulfate. In embodiments, dialysis is used to change the concentration of the salt in a composition containing a fusion protein and nucleic acids and / or contaminants. In embodiments, the salt is added at a concentration of about 0.5 M to about 3 M (e.g., 0.5, 1, 1.5, 2, 2.5, or 3 M (including all ranges or values in between)).
[0251] In the embodiments, the environmental factor is the addition of a cofactor. Non-limiting examples of cofactors include calcium, magnesium, cobalt, copper, zinc, iron, manganese, selenium, molybdenum, potassium, coenzyme A (CoA), nucleoside triphosphates, and vitamins. In the embodiments, the cofactor is calcium. In the embodiments, the nucleoside triphosphate is adenosine triphosphate, uridine triphosphate, guanosine triphosphate, cytidine triphosphate, or thymidine triphosphate. In the embodiments, the vitamin is fat-soluble. In the embodiments, the vitamin is water-soluble. Non-exclusive examples of vitamins include vitamin A, vitamin B1 (thiamine), vitamin B2 (riboflavin), vitamin B3 (niacin or niacinamide), vitamin B5 (pantothenic acid), vitamin B6 (pyridoxine, pyridoxal, or pyridoxamine, or pyridoxine hydrochloride), vitamin B7 (biotin), vitamin B9 (folic acid), vitamin B12, vitamin C, vitamin D, vitamin E, vitamin K, K1, and K2, folic acid, and biotin.
[0252] In this embodiment, the environmental factor is a change in the concentration of the fusion protein. In this embodiment, the environmental factor is a change in the concentration of nucleic acid. In this embodiment, the environmental factor is a change in the concentration of pollutants.
[0253] In the embodiment, the environmental factor is a change in pressure of the composition containing the fusion protein and nucleic acid. In the embodiment, the environmental factor is a change in pressure of the composition containing the fusion protein and contaminants. In the embodiment, the change in pressure can be achieved by increasing or decreasing the volume of the composition.
[0254] In the embodiments, the environmental factor is the addition of one or more surfactants. In the embodiments, one or more surfactants are selected from free fatty acid salts, soaps, fatty acid sulfonates, e.g., sodium lauryl sulfate, ethoxylated compounds, e.g., ethoxylated propylene glycol, lecithin, polygluconates, quaternary ammonium salts, lignin sulfonates, 3-((3-colamidopropyl)dimethylammonio)-1-propanesulfonate (CHAPS), sugars including sucrose and glucose, Triton X-100, and NP-40. In the embodiments, the surfactants are anionic, nonionic, or amphoteric.
[0255] In embodiments, the environmental factor is the addition of one or more molecular tungsten agents. Non-limiting examples of molecular tungsten agents include polyethylene glycol, dextran, and Ficol. Non-limiting examples of PEGS include PEG400, PEG1450, PEG3000, PEG8000, and PEG10000.
[0256] In embodiments, the environmental factor is the addition of one or more oxidizing agents. Non-limiting examples of oxidizing agents include hydrogen peroxide, hydrophilic or hydrophobic activated hydrogen peroxide, preformed peracid, persulfate, or hypochlorite.
[0257] In the embodiments, the environmental factor is the addition of one or more reducing agents. In the embodiments, one or more reducing agents are selected from the group consisting of dithiothreitol (DTT), 2-mercaptoethanol (BME), tris(2-carboxyethyl)phosphine (TCEP), hydrazine, boron hydride, amine borane, lower alkyl-substituted amine borane, triethanolamine, and N,N,N',N'-tetramethylethylenediamine (TEMED).
[0258] In embodiments, the environmental factor is the addition of one or more denaturing agents. Non-limiting examples of denaturing agents include urea, guanidine hydrochloride, guanidine, sodium salicylate, dimethyl sulfoxide, and propylene glycol.
[0259] In embodiments, the environmental factor is the addition of one or more enzymes. Non-limiting examples of enzymes include proteases, kinases, phosphatases, synthetases, transferases, nucleases, such as restriction endonucleases, lyases, isomerases, dehydrogenases, decarboxylases, and lipases.
[0260] In some embodiments, the environmental factor is the application of electromagnetic waves. In some embodiments, the environmental factor is the application of light. In some embodiments, the electromagnetic waves have wavelengths ranging from about 0.0001 nm to about 100 m. In some embodiments, the electromagnetic waves are selected from the group consisting of gamma rays, X-rays, ultraviolet rays, visible light, infrared rays, and radio waves. In some embodiments, the electromagnetic waves are gamma rays. In some embodiments, the gamma rays have wavelengths ranging from about 0.0001 nm to about 0.01 nm, for example, 0.0001 nm, 0.0005 nm, 0.001 nm, 0.002 nm, 0.003 nm, 0.004 nm, 0.005 nm, 0.006 nm, 0.007 nm, 0.008 nm, 0.009 nm, and 0.01 nm. In this embodiment, the X-rays have wavelengths ranging from approximately 0.01 nm to 10 nm, for example, approximately 0.01 nm, 0.02 nm, 0.03 nm, 0.04 nm, 0.05 nm, 0.06 nm, 0.07 nm, 0.08 nm, 0.09 nm, 0.10 nm, 0.2 nm, 0.3 nm, 0.4 nm, 0.5 nm, 0.6 nm, 0.7 nm, 0.8 nm, 0.9 nm, 1 nm, 2 nm, 3 nm, 4 nm, 5 nm, 6 nm, 7 nm, 8 nm, 9 nm, or approximately 10 nm. In this embodiment, ultraviolet light has wavelengths ranging from approximately 10 nm to approximately 400 nm, for example, approximately 10 nm, approximately 20 nm, approximately 30 nm, approximately 40 nm, approximately 50 nm, approximately 60 nm, approximately 70 nm, approximately 80 nm, approximately 90 nm, approximately 100 nm, approximately 150 nm, approximately 200 nm, approximately 250 nm, approximately 280 nm, approximately 300 nm, approximately 350 nm, or approximately 400 nm. In this embodiment, visible waves have wavelengths ranging from approximately 400 nm to approximately 800 nm, for example, approximately 400 nm, approximately 450 nm, approximately 500 nm, approximately 550 nm, approximately 600 nm, approximately 650 nm, approximately 700 nm, approximately 750 nm, or approximately 800 nm. In this embodiment, the infrared radiation has wavelengths ranging from approximately 800 nm to approximately 0.1 cm, for example, approximately 800 nm, approximately 1 μm, approximately 2 μm, approximately 3 μm, approximately 4 μm, approximately 5 μm, approximately 6 μm, approximately 7 μm, approximately 8 μm, approximately 9 μm, approximately 10 μm, approximately 20 μm, approximately 30 μm, approximately 40 μm, approximately 50 μm, approximately 60 μm, approximately 70 μm, approximately 80 μm, approximately 90 μm, approximately 100 μm, approximately 200 μm, approximately 300 μm, approximately 400 μm, approximately 500 μm, approximately 600 μm, approximately 700 μm, approximately 800 μm, approximately 900 μm, or approximately 0.1 cm.In an embodiment, the radio waves have a wavelength of from about 0.1 cm to 100 m, for example, about 0.1 cm, about 1 cm, about 10 cm, about 100 cm, about 1,000 cm, about 2,000 cm, about 3,000 cm, about 4,000 cm, about 5,000 cm, about 6,000 cm, about 7,000 cm, about 8,000 cm, about 9,000 cm, or about 100 m.
[0261] In an embodiment, the environmental factor is application of sound waves. In an embodiment, the sound waves have a frequency of from about 1 Hz to 2000 kHz. In an embodiment, the sound waves have a frequency of about 1 Hz, about 5 Hz, about 10 Hz, about 20 Hz, about 30 Hz, about 40 Hz, about 50 Hz, about 60 Hz, about 70 Hz, about 80 Hz, about 90 Hz, about 100 Hz, about 200 Hz, about 300 Hz, about 400 Hz, about 500 Hz, about 600 Hz, about 700 Hz, about 800 Hz, about 900 Hz, about 1 kHz, about 100 kHz, about 200 kHz, about 300 kHz, about 400 kHz, about 500 kHz, about 600 kHz, about 700 kHz, about 800 kHz, about 900 kHz, about 1000 kHz, about 1100 kHz, about 1200 kHz, about 1300 kHz, about 1400 kHz, about 1500 kHz, about 1600 kHz, about 1700 kHz, about 1800 kHz, about 1900 kHz, or about 2000 kHz.
[0262] Numbered Embodiments In addition to the appended claims, the following numbered embodiments also form part of the present disclosure.
[0263] 1. A fusion protein comprising a nucleic acid binding protein (NBP) and a polypeptide having phase behavior.
[0264] 2. NBP contains the following: RNA-specific adenosine deaminase 1 (ADAR1), ADAR1 double-stranded RNA-binding domain 3 (dsRBD3), Bacillus subtilis cold shock protein B (Bs-CspB), cold shock domain Y-box protein (CSD-Ybox), eukaryotic translation initiation factor 4E (eIF4e), Fox-1 protein (FOX1), heteronuclear ribonucleoprotein Q1 (hnRNPQ1), human zinc finger CCCH-type 14 (HsZC3H14), poly(A)-binding protein (PABP), poly(A)-binding protein nucleus 1 (PABPN1), pentatricopeptide repeat protein A (PPRpA), pumilio-like repeat protein A (PUFpA), and Staufen, 12-O-tetradecanoylphorbol-13-acetate-inducible sequence 11. The fusion protein according to claim 1, selected from one of D(TIS11D), Z-type DNA / RNA binding protein 1 (ZBP1), and zinc finger nuclease (ZNF).
[0265] 3. The fusion protein according to claim 1, wherein the NBP comprises an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one polypeptide of SEQ ID NOs.
[0266] 4. The fusion protein according to any one of claims 1 to 3, wherein the NBP is selected from one of the following: ADAR1 double-stranded RNA-binding domain 3 (dsRBD3), heteronuclear ribonucleoprotein Q1 (hnRNPQ1), human (Homo sapiens) zinc finger CCCH-type 14 (HsZC3H14), poly(A)-binding protein nucleus 1 (PABPN1), pentatricopeptide repeat protein A (PPRpA), and pumilio-like repeat protein A (PUFpA).
[0267] 5. The fusion protein according to any one of claims 1 to 4, wherein the NBP comprises an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, or at least 97%, at least 98%, at least 99%, or 100% identity with any one polypeptide of SEQ ID NOs.
[0268] 6. The fusion protein according to any one of claims 1 to 5, wherein the polypeptide having phase behavior has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one polypeptide of SEQ ID NOs: 1 to 60 or 217.
[0269] 7. The fusion protein according to any one of claims 1 to 6, wherein the polypeptide having phase behavior has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with respect to any one polypeptide of SEQ ID NO: 57 or 60.
[0270] 8. A fusion protein according to any one of claims 1 to 7, comprising a linker.
[0271] 9. The fusion protein according to claim 8, wherein the linker comprises an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one polypeptide of SEQ ID NOs.
[0272] 10. The fusion protein according to any one of claims 1 to 9, wherein the NBP binds to RNA, DNA, or both.
[0273] 11. The fusion protein according to any one of claims 1 to 10, wherein the NBP binds to RNA selected from one of the following: double-stranded RNA (dsRNA), single-stranded RNA (ssRNA), mRNA, premRNA, polyadenosine (polyA)RNA, Z-conformation RNA (Z-type RNA), or a combination thereof.
[0274] 12. The fusion protein according to any one of claims 1 to 11, wherein the NBP binds to the 3' end of mRNA, the 3' untranslated region (UTR) of mRNA, the poly-A tail of mRNA, or the AU-rich element of mRNA, or a combination thereof.
[0275] 13. The fusion protein according to any one of claims 1 to 11, wherein the NBP binds to premRNA.
[0276] 14. The fusion protein according to claim 13, wherein the NBP binds to an intron, exon, polyA tail, or a combination thereof of premRNA.
[0277] 15. A fusion protein according to any one of claims 1 to 10, wherein NBP binds to DNA.
[0278] 16. The fusion protein according to claim 15, wherein NBP binds to single-stranded DNA, double-stranded DNA, polyadenosine (polyA)DNA, Z-conformation DNA (Z-type DNA), or a combination thereof.
[0279] 17. A nucleic acid encoding the fusion protein according to any one of claims 1 to 16.
[0280] 18. The nucleic acid according to claim 17, comprising a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one nucleic acid of sequence numbers 119 to 133.
[0281] 19. A vector encoding the fusion protein according to any one of claims 1 to 16.
[0282] 20. A vector comprising the nucleic acid described in any one of claims 17 to 18.
[0283] 21. A vector having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with respect to any one nucleic acid of sequence numbers 104-118, 134, and 135.
[0284] 22. A method for purifying nucleic acids, comprising: (i) contacting a composition comprising nucleic acids and at least one contaminant with a fusion protein according to any one of claims 1 to 16, wherein the fusion protein binds to the nucleic acids to form a complex; (ii) contacting the complex with a first environmental factor to increase the size of the complex; (iii) separating the complex from at least one contaminant; and (iv) separating the nucleic acid from the fusion protein by contacting the complex with a second environmental factor.
[0285] 23. The method according to claim 22, wherein the complex is separated from at least one contaminant based on size.
[0286] 24. The method according to claim 23, wherein size-based separation is carried out using a method selected from one of tangential flow filtration, membrane chromatography, analytical ultracentrifugation, high-performance liquid chromatography, membrane chromatography, total filtration, ultrasonic separation, centrifugation, countercurrent centrifugation, and high-performance protein liquid chromatography.
[0287] 25. The method according to any one of claims 22 to 24, wherein the first environmental factor comprises: (a) a change in one or more of the following: temperature, pH, salt concentration, concentration of the purified matrix, concentration of virus particles, or pressure; (b) the addition of one or more surfactants, cofactors, vitamins, molecular concentrators, reducing agents, oxidizing agents, enzymes, or denaturants; or (c) the application of one or more electromagnetic waves or sound waves.
[0288] 26. The method according to any one of claims 22 to 25, wherein the second environmental factor comprises: (a) a change in one or more of the following: temperature, pH, salt concentration, concentration of the purified matrix, concentration of virus particles, or pressure; (b) the addition of one or more surfactants, cofactors, vitamins, molecular concentrators, reducing agents, oxidizing agents, enzymes, or denaturants; or (c) the application of one or more electromagnetic waves or sound waves.
[0289] 27. The method according to any one of claims 22 to 26, wherein at least one contaminant is selected from: solvents, proteins, peptides, carbohydrates, nucleic acids, viruses, cells (e.g., bacteria, yeast, or mammalian cells), carbohydrates, lipids, or lipopolysaccharides. [Examples]
[0290] Example 1. Design of a fusion protein containing a nucleic acid-binding protein and a polypeptide exhibiting phase behavior. A fusion protein containing an NBP and a polypeptide having phase behavior is generated and characterized. The NBP contains a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one of SEQ ID NOs. The polypeptide exhibiting phase behavior contains a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one polypeptide of SEQ ID NOs: 1-60 or 217. The fusion protein containing the NBP and the polypeptide exhibiting phase behavior is expressed according to a standard protocol. The affinity between the fusion protein and the nucleic acid is evaluated by capture experiments. The transition temperature of the fusion protein is determined using UV-Vis spectrophotometry.
[0291] Example 2. Purification of nucleic acids using the fusion protein from Example 1. The target nucleic acid is purified using the fusion protein from Example 1. The fusion protein is mixed with a sample containing the target nucleic acid and contaminants. The fusion protein binds to the target nucleic acid. The first environmental factor is added to the composition to increase the size of the complex. The complex is separated from the contaminants using TFF. The second environmental factor is added to the solution containing the purified complex to separate the fusion protein from the nucleic acid.
[0292] Example 3. Comparison of ITC generation reagents containing different polypeptides exhibiting phase behavior for nonspecific binding. Objective: The objective of this study is to identify the most effective polypeptides with phase behavior for use in the purification of single-stranded nucleic acids.
[0293] Methods: This study evaluated the nonspecific binding ability of various phase-behaving polypeptides to diverse nucleic acids when used alone. The affinity of phase-behaving polypeptides to nucleic acids was evaluated as described in Example 2. Nucleic acids were synthesized, mixed with various phase-behaving polypeptides, and then purified as described in Example 2 (e.g., using a reverse transfer cycle). The final concentrations of phase-behaving polypeptides, nucleic acids, and salts during the capture reaction were 40 μM, 0.1 mg / ml, and 1.5 M, respectively. Subsequently, 10 μl of the capture reaction product was loaded into test wells. In addition to linearized plasmid DNA, RNA templates of different lengths and strand states were tested for binding to four different phase-behaving polypeptides (20A80 (SEQ ID NO: 262), 50A80 (SEQ ID NO: 263), 100V80 (SEQ ID NO: 56), and 40L80 (SEQ ID NO: 264)). The capture reaction product was loaded directly into a 1% native agarose gel so that good binding of nucleic acid species by phase-behaving polypeptides would be demonstrated by retention in the wells.
[0294] Results: The results are shown in Figure 1. As shown in the figure, 20A80, 50A80, and 100V80 showed little to no nonspecific binding, while 40L80 showed nonspecific binding to a variety of nucleic acid species, indicating a low nucleic acid capture effect. Among the polypeptides exhibiting phase behavior, 100V80 also showed the optimal operating temperature for the reverse transition cycle during production. Therefore, 100V80 was determined to be the most successful candidate polypeptide with phase behavior for developing fusion proteins.
[0295] Example 4. Effective elution of single-stranded RNA in an in vitro transcription reaction using a reagent containing a polypeptide having phase behavior with Bs-CspB. Objective: The objective of this study was to evaluate the purification of single-stranded RNA (ssRNA) during in vitro transcription using a reagent containing a fusion protein (SEQ ID NO: 225) that includes the NBP Bs-CspB (SEQ ID NO: 90) and a polypeptide with a phase behavior called 100V80 (SEQ ID NO: 56).
[0296] Methods: This experiment was performed using ssRNA purification reagents (called "Isotag") that were generated in a shaking flask and purified by reverse transfer cycle (ITC). Next, RNA capture reactions were performed using ssRNA purification reagents obtained from in vitro transcription (IVT) reactions in which various mRNA template sizes, including 1kb, 4kb, and 8kb, were individually transcribed.
[0297] Direct IVT capture was performed using an ITC-generated ssRNA purification reagent. For each IVT reactant, 50 μM of the reagent was added to capture the desired ssRNA present in the reactant. Phase separation was possible by adding NaCl to achieve a final concentration of 1.5 M, which was necessary to separate the captured RNA in the droplets from other IVT contaminants based on size using centrifugation at room temperature (RT). After aspirating the capture supernatant, the RNA was dissociated from the droplets using water while keeping the droplets intact in the presence of heat. The droplets were separated from the final purified RNA by another round of centrifugation. To measure the nucleic acid concentration in the IVT reactants, the RNA present in the reactants was precipitated and purified using lithium acetate precipitation followed by ethanol washing, and then resuspended in water. RNA concentration was calculated by examining the absorbance at 260 nm and 280 nm using a UV spectrophotometer and calculating the concentration using the Beer-Lambert law. Next, the total RNA elution rate was calculated for each mRNA template containing 1kb, 4kb, and 8kb. The elution rate represents the final yield after elution, compared to the total RNA present in the initial IVT reaction product.
[0298] Results: The results are shown in Figure 2. In summary, the ssRNA purification reagent was able to recover more than 80% of the ssRNA during the ITC reaction step. These findings demonstrate that we successfully purified target ssRNA effectively by size-based separation using a fusion protein containing polypeptide 100V80, which exhibits phase behavior with Bs-CspB.
[0299] Example 5. Highly selective capture of purified RNA using the ITC generation reagent of Example 4. Objective: The objective of this study was to evaluate the specificity of the ITC generating reagent of Example 4 in selectively capturing target single-stranded RNA by quantifying the isolation of contaminants exemplified by dsRNA.
[0300] Methods: This study investigated the ability of the ssRNA purification reagent from Example 3 to selectively capture purified RNA. The separation of dsRNA from ssRNA was evaluated at various ssRNA purification reagent:RNA molar ratios.
[0301] To determine the Log10 removal value (LRV) of dsRNA, the dsRNA concentration in solution was evaluated before and after purification using an ITC generating reagent. RNA was synthesized using the HiScribe® T7 Quick High Yield RNA Synthesis Kit (NEB) and a linearized eGFP template with a 120-nucleotide poly-A tail and supporting 5' and 3' UTRs. To ensure sufficient dsRNA synthesis, IVT was performed overnight (O / N) at 20°C. To verify ssRNA specificity while increasing dsRNA impurities, ssRNA was synthesized by IVT using Hi-T7® RNA Polymerase (NEB), and then enriched using a cellulose-based ssRNA purification method. Simultaneously, dsRNA matching the sequence and size of the ssRNA was also synthesized and purified, as described in Biersdorfer et al, Mol. Ther Nucleic Acids, 2019. ssRNA concentrations were evaluated using UV spectrophotometric analysis, and dsRNA concentrations were tested using a multi-species dsRNA ELISA kit (Novus Biologicals). Furthermore, the elution rate of total dsRNA in total RNA containing dsRNA added at various ratios, such as 0%, 0.01%, 0.1%, 1%, and 10%, was evaluated.
[0302] Results: The results are shown in Figures 3A and 3B. In summary, the ssRNA purification reagents were highly selective for ssRNA. As shown in Figure 3A, RNA capture results of purified RNA with added dsRNA demonstrated that the reagents maintained selectivity for ssRNA and were capable of removing dsDNA exceeding 4-log at various ssRNA purification reagent:RNA molar ratios. Furthermore, Figure 3B shows that selectivity was maintained in the capture performed on purified RNA even with the addition of a high concentration of dsRNA, 10% of total RNA. Regardless of the added dsRNA content, the eluted percentage during the purification process was maintained at less than 4%. These findings indicate that reagents containing Bs-CspB and 100V80 are highly selective for target nucleic acids while simultaneously separating most contaminants from the final product.
[0303] Example 6. Comparison of Bs-CspB and 100V80 reagents with other nucleic acid-binding protein reagents regarding binding to ssRNA and other sequences. Objective: The objective of this study was to evaluate the ability of the ITC-generating reagents containing Bs-CspB and 100V80 described in Examples 4 and 5 to bind to a variety of ssRNA templates and other sequences, and to compare this ability with other ITC-generating reagents containing nucleic acid-binding proteins different from 100V80.
[0304] Methods: This study evaluated nucleic acid capture of purified RNA using several different ITC-generating reagents. RNA was synthesized from PCR-generated DNA templates or linearized plasmids using the HiScribe T7 ARCA mRNA kit (NEB), and post-IVT tailing was performed on transcripts supporting a poly(A) tail. 513 bp dsRNA was generated using a luciferase template, as described in Biersdorfer et al, Mol Ther Nucleic Acids, 2019. The RNA was purified by phenol-chloroform extraction followed by ethanol precipitation with ammonium acetate. The final RNA pellet was resuspended in 1 mg / ml ultrapure water. The purity and length of the mRNA transcript were confirmed by UV-Vis spectrophotometric analysis and gel electrophoresis. The final concentrations of biomolecules, RNA, and salts in the capture reaction were 40 μM, 0.1 mg / ml, and 1.5 M, respectively. After centrifugation at 5000 g for 5 minutes, the capture supernatant was aspirated. The success of DNA capture was evaluated by gel electrophoresis of samples separately loaded onto 1% native agarose gels treated with 1×Sybr Safe DNA Gel Stain.
[0305] Candidate ssRNA binding proteins, consisting of various nucleic acid-binding proteins bound to 100V80, were screened for their ability to capture ssRNA and dsRNA targets using independent capture reaction products. The presence of nucleic acid species in the capture supernatant was examined using agarose gel electrophoresis. The presence of bands on the gel indicates nucleic acids that were not captured.
[0306] Results: The results are shown in Figure 4. As illustrated, the Bs-CspB and 100V80 reagents showed no banding of the ssRNA template, indicating that a large amount of dsRNA remained uncaptured. Overall, the Bs-CspB reagent demonstrated excellent capture of diverse ssRNA species without retaining dsRNA within the droplet. In summary, these findings demonstrate the success of using fusion proteins containing Bs-CspB and 100V80 for the purpose of purifying single-stranded nucleic acids.
[0307] Embedding by reference All references, articles, publications, patents, patent gazettes, and patent applications cited herein are incorporated by reference in their entirety for all purposes. However, references to references, articles, publications, patents, patent gazettes, and patent applications cited herein are not, and should not be construed as, an endorsement or in any form of suggestion that they constitute valid prior art or part of common general knowledge applicable in any country in the world. The following patent documents are incorporated herein by reference in their entirety for all purposes: International Publication No. 2021 / 168270 and International Publication No. 2022 / 178537.
Claims
1. A fusion protein containing a polypeptide having phase behavior with Bacillus subtilis cold shock protein B (Bs-CspB).
2. The fusion protein according to claim 1, wherein the Bs-CspB comprises an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polypeptide of SEQ ID NO:
90.
3. The fusion protein according to claim 1 or 2, wherein the polypeptide having phase behavior comprises P and G motifs comprising a plurality of proline residues and a plurality of glycine residues.
4. The fusion protein according to any one of claims 1 to 3, wherein the P and G motifs comprise at least about 10% proline residues and at least about 20% glycine residues.
5. The fusion protein according to any one of claims 1 to 4, wherein the polypeptide having phase behavior comprises a pentapeptide repeat having the sequence (Val-Pro-Gly-Xaa-Gly)n (SEQ ID NO: 217), or a randomized, scrambled analog thereof; wherein Xaa is any amino acid other than proline.
6. The fusion protein according to claim 5, wherein n is an integer from 1 to 360 that includes the endpoint.
7. The polypeptide having phase behavior is: a. (GRGDSPY) n (Sequence ID 1); b. (GRGDSPH) n (Sequence ID 2); c. (GRGDSPV) n (Sequence number: 3); d. (GRGDSPYG) n (Sequence ID 4); e. (RPLGYDS) n (Sequence number: 5); f. (RPAGYDS) n (Sequence ID 6); g. (GRGDSYP) n (Sequence ID 7); h. (GRGDSPYQ) n (Sequence No. 8); i. (GRGNSPYG) n (SEQ ID NO: 9); j. (GVGVP) n (Sequence No. 10); k. (GVGVPGLGVPGVGVPGLGVPGVGVP) m (Sequence No. 11); l. (GVGVPGVGVPGAGVPGVGVPGVGVP) m (Sequence No. 12); m. (GVGVPGWGVPGVGVPGWGVPGVGVP) m (Sequence ID 13); n. (GVGVPGVGVPGVGVPGVGVPGVGVPGVGVPGEGVPFGVPGVGVGVP) m (SEQ ID NO: 14); o. (GVGVPGVGVPGVGVPGVGVPGVGVPGVGVPGKGVPGFGVPGVGVP) m (Sequence ID 15); and p. (GAGVPGVGVPGAGVPGVGVPGAGVP) m (Sequence ID 16) comprises an amino acid sequence selected from or a randomized and scrambled analog thereof. During the ceremony: n is an integer in the range of 20 to 360, including the endpoint; and The fusion protein according to any one of claims 1 to 6, wherein m is an integer in the range of 4 to 25 including the endpoint.
8. The polypeptide having phase behavior is as follows: a. (GVGVP) m (Sequence ID 22); b. (ZZPXXXXGZ) m (Sequence ID 23); c. (ZZPXGZ) m (Sequence ID 24); d. (ZZPXXGZ) m (Sequence ID 25); or e. (ZZPXXXGZ) m (Sequence No. 26) A fusion protein according to any one of claims 1 to 5, comprising an amino acid sequence selected from, wherein m is an integer from 10 to 160 including the endpoint, X is any amino acid except proline or glycine, if present, and Z is any amino acid, if present.
9. The polypeptide having phase behavior is as follows: a. (GVGVPGVGVPGAGVPGVGVPGVGVP) m (Sequence ID 17); or b. (GVGVPGVGVPGLGVPGVGVPGVGVP) m (Sequence No. 18) It includes an amino acid sequence selected from; The fusion protein according to any one of claims 1 to 5, wherein m is an integer from 2 to 32 that includes the endpoint.
10. The polypeptide having phase behavior is: a. (GVGVPGVGVPGAGVPGVGVPGVGVP) m (Sequence ID 19) (wherein m is 8 or 16); b. (GVGVPGAGVP) m (Sequence ID 20) (wherein m is an integer from 5 to 80 including the endpoint); or c. (GXGVP) m (Contains an amino acid sequence selected from Sequence ID No. 21) In the formula, m is an integer between 10 and 160 that includes the endpoint, and The fusion protein according to any one of claims 1 to 5, wherein X in each repeat is independently selected from the group consisting of glycine, alanine, valine, isoleucine, leucine, phenylalanine, tyrosine, tryptophan, lysine, arginine, aspartic acid, glutamic acid, and serine.
11. The fusion protein according to any one of claims 1 to 5, wherein the polypeptide having phase behavior has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one polypeptide of SEQ ID NOs: 1 to 60, 217, or 262 to 264.
12. The fusion protein according to any one of claims 1 to 5, wherein the polypeptide having phase behavior has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one polypeptide of SEQ ID NO: 56 or 262 to 264.
13. The fusion protein according to any one of claims 1 to 5, wherein the polypeptide having phase behavior has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polypeptide of Sequence ID No.
56.
14. The fusion protein according to claim 1 or 2, wherein the fusion protein has at least 90% identity with the polypeptide of SEQ ID NO: 90 and at least 90% identity with any one polypeptide of SEQ ID NO: 56 and 262-264.
15. The fusion protein according to any one of claims 1 to 2 or 14, wherein the fusion protein has at least 90% identity with the polypeptide of SEQ ID NO: 90 and at least 90% identity with the polypeptide of SEQ ID NO:
56.
16. A fusion protein according to any one of claims 1 to 15, comprising a linker.
17. The fusion protein according to claim 16, wherein the linker comprises an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one polypeptide of SEQ ID NOs: 143-216 or 261.
18. The fusion protein according to claim 16 or 17, wherein the linker comprises an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polypeptide of SEQ ID NO:
261.
19. The fusion protein according to claim 1, wherein the fusion protein comprises an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polypeptide of either SEQ ID NO: 225 or 226.
20. The fusion protein according to claim 1, wherein the fusion protein has at least 90% identity with the polypeptide of SEQ ID NO: 225 or 226.
21. The fusion protein according to any one of claims 1 to 20, wherein the Bs-CspB binds to RNA, DNA, or both.
22. The fusion protein according to any one of claims 1 to 21, wherein the Bs-CspB binds to RNA selected from any one of double-stranded RNA (dsRNA), single-stranded RNA (ssRNA), mRNA, premRNA, polyadenosine (polyA)RNA, Z-conformation RNA (Z-type RNA), or a combination thereof.
23. The fusion protein according to any one of claims 1 to 22, wherein the Bs-CspB binds to ssRNA.
24. The fusion protein according to any one of claims 1 to 23, wherein the Bs-CspB binds to the 3' end of mRNA, the 5' end of mRNA, the coding region of mRNA, the non-coding region of mRNA, or a combination thereof.
25. The fusion protein according to any one of claims 1 to 23, wherein the Bs-CspB binds to premRNA.
26. The fusion protein according to claim 25, wherein the Bs-CspB binds to an intron, exon, 5'UTR, 3'UTR, or a combination thereof of the premRNA.
27. The fusion protein according to any one of claims 1 to 21, wherein the Bs-CspB binds to DNA.
28. The fusion protein according to claim 27, wherein the Bs-CspB binds to single-stranded DNA, double-stranded DNA, polyadenosine (polyA) DNA, Z-conformation DNA (Z-type DNA), or a combination thereof.
29. A nucleic acid encoding a fusion protein according to any one of claims 1 to 28.
30. The nucleic acid according to claim 29, comprising a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one nucleic acid of sequence numbers 119 to 133.
31. A vector encoding a fusion protein according to any one of claims 1 to 28.
32. A vector comprising the nucleic acid according to claim 29 or 30.
33. A vector having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with respect to any one nucleic acid of sequence numbers 104-118, 134, and 135.
34. A method for purifying single-stranded nucleic acids, as follows: (i) A step of contacting a composition comprising the single-stranded nucleic acid and at least one contaminant with a fusion protein according to any one of claims 1 to 28, wherein the fusion protein binds to the single-stranded nucleic acid to form a complex; (ii) Adding a first environmental factor to the composition containing the complex, thereby increasing the size of the complex; (iii) the step of separating the complex from at least one contaminant; and (iv) The step of separating the single-stranded nucleic acid from the fusion protein by contacting the complex with a second environmental factor, thereby forming a product containing the single-stranded nucleic acid. A method that includes this.
35. The method according to claim 34, wherein the single-stranded nucleic acid includes ssRNA.
36. The method according to claim 34 or 35, wherein the composite is separated from the at least one contaminant based on size.
37. The method according to claim 36, wherein the size-based separation is carried out using a method selected from one of tangential flow filtration, membrane chromatography, analytical ultracentrifugation, high-performance liquid chromatography, total filtration, ultrasonic separation, centrifugation, countercurrent centrifugation, and high-performance protein liquid chromatography.
38. The method according to claim 37, wherein the separation based on the size is carried out using a method including centrifugation.
39. The first environmental factor is as follows: (a) Changes in one or more of the following: temperature, pH, salt concentration, concentration of purified matrix, concentration of virus particles, or pressure; (b) Addition of one or more surfactants, cofactors, vitamins, molecular concentrators, reducing agents, oxidizing agents, enzymes, or denaturing agents; or (c) Application of electromagnetic waves or sound waves The method according to any one of claims 34 to 38, comprising one or more of the above.
40. The second environmental factor is as follows: (a) Changes in one or more of the following: temperature, pH, salt concentration, concentration of purified matrix, concentration of virus particles, or pressure; (b) Addition of one or more surfactants, cofactors, vitamins, molecular concentrators, reducing agents, oxidizing agents, enzymes, or denaturing agents; or (c) Application of electromagnetic waves or sound waves The method according to any one of claims 34 to 39, comprising one or more of the above.
41. The method according to any one of claims 34 to 40, wherein the at least one contaminant is selected from: solvents, proteins, peptides, carbohydrates, double-stranded nucleic acids, viruses, cells (e.g., bacteria, yeast, or mammalian cells), carbohydrates, lipids, or lipopolysaccharides.
42. The method according to claim 41, wherein the at least one contaminant is a double-stranded nucleic acid.
43. The method according to claim 42, wherein the double-stranded nucleic acid is dsRNA.
44. The method according to claim 39 or 40, wherein the salt is added at a concentration of about 0.5 M to about 3 M.
45. The method according to any one of claims 39 to 40 or 44, wherein the salt is added at a concentration of about 1.2 M to about 1.7 M.
46. The method according to any one of claims 34 to 45, wherein the product comprises about 70% to about 100% of the single-stranded nucleic acid, based on the total nucleic acid content in the product of step (iv).
47. The method according to claim 46, wherein the single-stranded nucleic acid includes ssRNA.
48. The method according to any one of claims 34 to 47, wherein the product contains at least 10% or less of the at least one contaminant based on the total nucleic acid content in the product.
49. The method according to claim 48, wherein the at least one contaminant includes dsRNA.
50. The method according to any one of claims 34 to 49, wherein the purification method includes removing at least about 1-log g of contaminants compared to the amount of contaminants in the composition of step (i).
51. The method according to any one of claims 34 to 50, wherein the purification method includes removing 1 to 10-log g of contaminants compared to the amount of contaminants in the composition of step (i).
52. The method according to any one of claims 34 to 51, wherein in step (i), the fusion protein is present at a concentration of about 1 μM to about 200 μM.
53. The method according to any one of claims 34 to 52, wherein in step (i), the fusion protein is present at a concentration of about 30 μM to about 60 μM.