Synthetic organelle-assisted gene expression method, DNA construct, expression system, and use of synthetic organelle
By constructing synthetic organelles in the host and localizing the expression products to the synthetic organelles, the problem of poor stress on the host and maturation environment of the expression products is solved, and efficient, low stress and correct mature protein expression is achieved.
Patent Information
- Application Number
- PCT/CN2024/085047
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-13
- Filing Date
- 2024-03-29
- Publication Date
- 2025-05-22
AI Technical Summary
There are two main challenges in cell expression: host stress and/or toxicity caused by expression products, and the inability to mature expression products in the environment within the host, resulting in poor expression efficiency and product quality.
By constructing synthetic organelles in the host, the expression product can be localized to the synthetic organelles, thereby isolating the physiological activity of the expression product from the host, reducing host stress, and providing a controlled mature environment for the expression product.
The production of highly efficient expression products is achieved, which reduces the impact of host pressure and toxicity, ensures the correct folding and modification of the products, and improves expression consistency and resource efficiency.
Smart Images

Figure PCTCN2024085047-FTAPPB-I100001 
Figure PCTCN2024085047-FTAPPB-I100002 
Figure PCTCN2024085047-FTAPPB-I100003
Abstract
Description
Synthetic organelle-assisted gene expression method, DNA construct and expression system and their applications Technical Field
[0001] The present invention relates to the field of engineering, and in particular to a synthetic organelle-assisted gene expression method, a DNA construct, an expression system and applications thereof. Background Art
[0002] Gene expression generally refers to the process of synthesizing and expressing a gene of interest in a host. Commonly used hosts include various cell expression systems, such as bacterial expression systems (Escherichia coli, Bacillus subtilis), yeast expression systems, and animal cell expression systems (insect cells, human cells). Alternatively, the host may be a cell-free expression system, such as a cell lysate. This process plays a vital role in fields such as biotechnology, molecular biology, and drug discovery, as it allows for the production of large quantities of target products, particularly recombinant proteins, for research, therapeutic, or industrial applications.
[0003] However, the interaction between the expression process and the host creates a series of complex technical and biological challenges:
[0004] a. Energy and material burden of expression: When attempting to express a foreign gene, the host must allocate energy and resources to meet the expression requirements. This increases the burden on the cell and may limit normal cell growth and health.
[0005] b. Host stress caused by the expression product: This increases the difficulty of gene expression because a balance must be found between normal host activity and protein expression. In particular, when the host is a cell expression system, cell growth may be inhibited or even stopped. According to the mechanism, the sources of host stress include:
[0006] a) Exogenous products occupy space in the host, thereby nonspecifically interfering with normal physiological processes and generating general stress.
[0007] b) Some products are specifically toxic and can cause damage to the host even at extremely low expression levels. For example, the target product is a nuclease, which can cleave and degrade DNA and RNA in the host; or the target product is an exogenous tRNA, which can cause protein translation errors.
[0008] c. The expression product cannot mature under the host environment: The maturation process of the target product requires a specific physical and chemical environment and biochemical reactions, which may not be present in the host, resulting in the target product not being able to mature correctly.
[0009] For example:
[0010] a) When exogenous proteins are overexpressed, misfolding often occurs. These misfolded proteins may not have normal biological activity and tend to aggregate into insoluble inclusion bodies.
[0011] b) When proteins with disulfide bonds are expressed in E. coli, disulfide bonds cannot form due to the reducing environment of the E. coli cytoplasm and the protein cannot fold correctly.
[0012] c) Some target products require post-processing to have correct activity, such as zymogen cleavage, antibody glycosylation, RNA intron excision, and RNA editing.
[0013] These difficulties are common challenges in expression, and researchers have been exploring various methods and strategies to overcome them to achieve efficient and controllable protein expression.
[0014] The present invention aims to address challenges b (host pressure) and c (product maturation), providing an innovative method to improve expression and bring new possibilities to research and application in related fields.
[0015] In this invention, we address two major challenges: host stress caused by the expression product (e.g., expression of toxic proteins) and the inability of the expression product to mature under the host's internal environment (e.g., protein misfolding). These challenges frequently arise in the fields of biotechnology and molecular biology, limiting the efficiency and feasibility of expression. The following are some situations and existing technologies related to these two challenges. These methods are often complex and not always successful, and there is currently no universal solution:
[0016] a. Select an appropriate host system: Based on the characteristics of the target gene, choose the most appropriate host expression system. For example, cell-free systems can be used to express cytotoxic proteins, while human cell lines can be used to express human antibodies. For proteins with disulfide bonds, reductase-deficient mutant strains, such as the E. coli Shuffle, can be used for gene expression. However, many host expression systems are costly and complex to operate. Furthermore, these approaches are only suitable for expressing a single gene or a number of genes with similar properties and cannot be generalized to the expression of more complex DNA constructs.
[0017] b. Expression vector engineering: For example, methods such as selecting an appropriate promoter, adjusting expression strength, and using enhancers can be used to adjust expression levels. However, these methods struggle to simultaneously achieve multiple goals. For example, a weak promoter can prevent excessive expression levels and thus inclusion body formation, but this approach cannot achieve efficient expression of the exogenous protein. On the other hand, a strong promoter can increase the burden on cells, thereby limiting the efficient expression of the target protein.
[0018] c. Optimizing expression conditions: Controlling cell culture conditions, such as temperature, pH, and oxygen supply, as well as using an inducible expression system and adjusting the induction conditions, can affect expression levels. For example, for proteins prone to misfolding and precipitating, lower expression temperatures can be used to promote correct folding. However, in some cases, these optimizations may not effectively improve expression.
[0019] d. Metabolic engineering: By changing the host's metabolic pathways, the yield of the target product can be increased.
[0020] e. Protein Engineering: For toxic proteins or proteins that easily form inclusion bodies, several methods have been proposed, such as altering protein structure, adjusting expression rate, using specific tags, or fragmenting proteins, to mitigate their adverse effects on cells. However, these methods may result in decreased expression efficiency.
[0021] f. Post-processing: The inactive target product produced by the host can be purified and then renatured, cleaved, or modified to restore the correct structure. However, post-processing is a very complex and expensive process and is not suitable for many target products.
[0022] However, there is currently no universal solution that can achieve high expression, low stress / toxicity, correct folding, and post-expression modification.
[0023] Prior art 1 involves optimization of expression vectors and expression conditions.
[0024] The technical solution of the prior art 1 is as follows: In the prior art 1, an engineering strategy of expression vectors and expression conditions is usually adopted to optimize expression.
[0025] Expression vector: Select appropriate promoters, expression vectors, and control elements to control gene expression and optimize expression. This can be achieved by adjusting the copy number in the expression vector, promoter activity, etc.
[0026] Expression conditions: Adjust culture conditions such as temperature, pH, oxygen supply, and medium composition. If the expression vector is inducible, this also includes adjusting the induction conditions. By changing the expression conditions to improve expression efficiency, you can also improve protein folding and stability.
[0027] The existing technology 1 has the following disadvantages: although the existing technology 1 can optimize expression to a certain extent, it lacks clear theoretical guidance, requires a lot of trial and error for different target genes, and is not necessarily effective.
[0028] The second prior art relates to the optimization of the expression host.
[0029] The technical solution of the second existing technology is as follows:
[0030] According to the characteristics of the target gene, select a suitable host for expression, and then match the corresponding post-processing steps.
[0031] For example, human cell lines or insect cells can be used to express animal-derived proteins. Optimizing the host's metabolic network, as well as its transcription and translation systems, can further enhance expression. For example, engineered CHO cell lines can be used for high-expression antibody production. For example, the SHuffle strain, obtained by modifying the reducing environment within E. coli, can be used to produce disulfide-bonded proteins.
[0032] For example, cell-free expression systems are used as an alternative to address the difficulties encountered in cell-based expression. Cell-free expression systems typically include the following key steps:
[0033] ① Cell lysis and preparation of cell extracts: Extract cytoplasm and nucleic acid substances from cells to obtain extracts containing basic biological molecules of the cells.
[0034] ② Adding exogenous DNA or RNA: Adding DNA or RNA encoding the target protein to the extract, usually synthesized by synthetic biology methods.
[0035] ③ Translation and protein synthesis: In an in vitro environment, the target protein is translated and synthesized using cellular machinery in cell extracts, such as ribosomes, tRNA, and amino acids.
[0036] ④ Protein purification and extraction: Purify and extract the target protein from the reaction system, usually using affinity chromatography, gel electrophoresis or other separation techniques.
[0037] ⑤Analysis and confirmation: Analyze and confirm the protein to ensure that its structure and function are as expected.
[0038] The second prior art has the following disadvantages:
[0039] a) High operational complexity: The transformation and transfection steps for different hosts vary significantly, resulting in a high learning threshold. In addition, some hosts are more difficult to genetically manipulate and culture.
[0040] b) High cost: Some host systems are expensive to use. For example, mammalian cell lines require expensive serum for culture, and cell-free systems often require the addition of expensive exogenous factors, such as energy substances, amino acids, ions, and proteolytic factors, to support protein synthesis.
[0041] c) The host may possess some desirable properties but lack others. For example, cell-free systems are not subject to cytotoxicity, but lack the constraints of post-translational modifications. Cell-free expression systems generally cannot provide the same levels of post-translational modifications as cells, such as glycosylation and phosphorylation, which may affect the function and stability of some proteins. In cell-free systems, protein expression levels may be relatively unstable, affected by reaction conditions and component concentrations. For example, shuffle strains can promote disulfide bond formation, but this also reduces their physiological activity and stress tolerance.
[0042] d) A single host is only suitable for a series of genes with similar properties and is difficult to generalize to multi-gene co-expression and the expression of more complex DNA constructs.
[0043] CN112575016A discloses a "construction and application of membraneless organelles in prokaryotes", which successfully constructs membraneless compartments in Escherichia coli using recombinant spider silk protein and arthropod-like elastic protein. Through DNA recombination technology, spider silk protein or arthropod-like elastic protein is fused with different cargo proteins to achieve intracellular co-localization of target functional proteins, thereby constructing membraneless compartments with biological activity.
[0044] WO2023015190A1 discloses a method for controlling cellular processes in mammalian cells by synthesizing organelles. This method utilizes an arginine / glycine-rich (RGG) protein domain capable of liquid-liquid phase separation (LLPS) fused to a high-affinity coiled-coil tag to enrich endogenous proteins also fused to the coiled-coil tag, thereby controlling cell behavior by sequestering or releasing endogenous proteins.
[0045] However, in the above two applications, there is no mention of the possibility of synthetic organelles being able to assist the correct folding of target gene products / target proteins, and existing data do not support such a conclusion. In fact, the descriptions in both applications clearly state that the target protein itself can be correctly expressed and function, of which CN112575016A is described as a protein with biological activity, and WO2023015190A1 is described as an endogenous protein that can regulate cell function. In addition, in CN112575016A, the target protein is connected to spider silk protein and arthropod elastic protein by a fusion method, and the folding and maturation of the target protein may be disturbed by the process of protein aggregation, and this property is difficult to measure and regulate. In WO2023015190A1, the host cell is a mammalian cell, and the protein expressed by mammalian cells is usually always appropriately modified, so the expression product directly has function. Its disadvantage is that the operation technology is difficult, time-consuming and uneconomical.
[0046] Prokaryotic expression systems such as Escherichia coli are currently the most widely used prokaryotic expression systems. Their advantages are simple, rapid, economical cultivation and amplification, and suitability for large-scale production processes; their disadvantages are that the expressed eukaryotic proteins cannot form proper folding or undergo glycosylation modification; the expressed proteins often form insoluble inclusion bodies, and complex renaturation treatment is required to make them active; it is not easy to obtain a large amount of soluble protein in the Escherichia coli expression system.
[0047] In summary, while several approaches have been proposed to optimize expression, this invention aims to provide a comprehensive, more versatile approach that simultaneously addresses these challenges. This innovation will help advance research and applications in related fields, providing new possibilities for a wider range of scientific research and engineering applications.
[0048] Summary of the Invention
[0049] The main purpose of the present invention is to solve two key technical problems in cell expression, namely, eliminating the stress and / or toxicity of gene expression products on the host and providing a suitable and controllable maturation environment for gene expression products, thereby optimizing gene expression.
[0050] The expression product will produce stress and even toxicity on the physiological activities of the host. Therefore, the present invention aims to separate the expression product and the original physiological activity of the host in the host, thereby decoupling the influence of exogenous expression on endogenous physiology. Specifically, the expression product will occupy space in the host, thereby non-specifically interfering with the normal physiological activity process, generating general stress, and may adversely affect the normal physiological activity of the host. Therefore, the present invention aims to provide a method that can significantly increase gene expression while reducing host stress to achieve efficient expression. In particular, some products have specific toxicity and can cause damage to the host even at low expression levels. In particular, when the host is a cell system, expressing these genes will usually have an adverse effect on the cells and even cause cell death. The present invention aims to provide a method that can optimize the expression of toxic products, reduce their adverse effects on the host, and maintain high expression levels.
[0051] The host's internal environment will cause incorrect maturation of the expression product. The maturation process of the target product requires a specific physical and chemical environment and biochemical reactions, which may not be available in the host, which will cause the target product to not mature correctly and produce an inactive expression product. For example, when expressing exogenous proteins, especially when expressed at a high efficiency, the protein often folds into an incorrect three-dimensional structure in the host, thereby losing normal biological activity and forming insoluble aggregates. For example, in Escherichia coli, due to the reducing environment in the cytoplasm, the protein cannot form a stable disulfide bond, so proinsulin produced using Escherichia coli generally exists as misfolded inclusion bodies. The present invention aims to provide a method for providing a maturation environment for the expression product that is decoupled from the host's internal environment, reducing the incorrect maturation of the expression product without destroying the host's original physiology, and thus efficiently obtaining an active target product.
[0052] In summary, the present invention aims to provide a comprehensive and efficient method that simultaneously addresses the key technical issues of host stress and product maturation. By addressing these issues, the present invention aims to provide more effective and feasible solutions for research and applications in bioengineering, drug discovery, protein production, and other related fields, thereby promoting the development of science and technology.
[0053] In order to achieve the above tasks, the present invention adopts the following technical solutions:
[0054] The present invention constructs a synthetic organelle in the host, so that after the target gene is expressed, the expression product can enter the synthetic organelle in a timely manner, thereby achieving:
[0055] 1. Eliminate stress and / or toxicity to the host: Isolating the expression product from the host reduces the probability of contact between the expression product and other molecules in the host. This not only minimizes the impact of the product on the host, but also changes the kinetic balance and feedback inhibition of gene expression by reducing the concentration of the product in the host, thereby promoting the production of the expression product.
[0056] 2. Provide a controllable and suitable maturation environment for expression products: Synthetic organelles provide an orthogonal environment that is different from the host's internal environment and can be programmably regulated. The expression products enriched in the synthetic organelles can complete the correct maturation processes such as folding, modification, and shearing.
[0057] In one aspect, the present invention provides a method for synthetic organelle-assisted gene expression, comprising the following steps:
[0058] a. Constructing synthetic organelles in the host,
[0059] b. co-expressing the synthetic organelle with a gene expression product, wherein the gene expression product can be localized to the synthetic organelle.
[0060] In one embodiment, the host includes a cell-free expression system and a cell host, preferably a prokaryotic cell host, more preferably Escherichia coli.
[0061] In one embodiment, the synthetic organelle refers to an artificially designed compartmentalized structure. Preferably, the synthetic organelle is a membraneless organelle, which is a biomolecular condensate constructed in a host by a polypeptide and / or protein and / or DNA and / or RNA and / or nucleic acid-protein complex. Preferably, the biomacromolecule used to construct the synthetic organelle is a nucleotide sequence repeated 10-1000 times continuously, such as the triplets CAG, CCUG, GGGAA, GGGGCC, etc.; an amino acid sequence repeated 3-500 times continuously, such as the repetition of resilin-like peptides (RLP); an intrinsic non-structural region, such as FUS protein, ParB protein; multimerization of multiple biomacromolecules, such as the repetition of PTB protein and UCUCU sequence, an artificially designed binary protein two-dimensional structure, and an artificially designed binary RNA two-dimensional structure.
[0062] In one embodiment, the gene expression product is RNA, protein, nucleic acid-protein complex, or a small molecule encoded by heterologous DNA, preferably a recombinant protein, for example, superfolded green fluorescent protein, restriction endonuclease DpnI, restriction endonuclease BsaI, human proinsulin or ROR-alpha, preferably, the gene expression product matures in the synthetic organelle.
[0063] In one embodiment, in step b, the gene expression product is localized to the synthetic organelle through non-specific and / or specific binding. Preferably, a specific receptor / ligand pair is used, the synthetic organelle has a receptor, and the gene expression product is connected to a ligand, so that the synthetic organelle can specifically bind to the gene expression product. Preferably, the receptor / ligand pair includes: complementary paired nucleic acid molecules; capsid proteins and translation operon RNA hairpins of RNA phages such as MS2, PP7, and Qbeta; two segments of a split protein, such as split green fluorescent protein, the alpha segment and omega segment of beta-galactosidase, and split T7 RNA polymerase.
[0064] On the other hand, the present invention provides a DNA construct comprising a nucleotide sequence 1 that guides the host to construct the synthetic organelle defined above, and / or a nucleotide sequence 2 that guides and regulates the production of the gene expression product. Preferably, the DNA construct is composed of the nucleotide sequences 1 and 2 and a plasmid vector, such as pACYC184, pACYC177, pET28a(+), pET28b(+), pET-5a(+), pET43.1a, pET-37b(+), pCDFDuet-1, pCOLADuet-1, pRSFDuet-1, pETDuet-1, pUC57, pUC19, pBAD, pBluescript II SK(+), pTrcHis C, pTrcHis A, pTrcHis2C, pBV221, pQE-70, pCold III, pRSET-CFP, pRSET-BFP, pGFPuv, pKD46, pKD4, pTYB1, PinPoint Xa-2, pTWIN1, pRSET C, pSB1C3, pSB3C5, pSB4C5, pSB4K5, preferably, the upstream promoter is a normally expressed promoter or an inducible promoter.
[0065] On the other hand, the present invention provides an expression system comprising the above-mentioned DNA construct. Preferably, the expression system refers to a system that simultaneously constructs the synthetic organelles and generates the gene expression products in prokaryotic cells. More preferably, the expression system is an Escherichia coli expression system.
[0066] In another aspect, the present invention provides the use of synthetic organelles in eliminating the stress and / or toxicity of gene expression on the host and / or decoupling the gene expression products from the host's internal environment, which includes co-expressing the synthetic organelles and the gene expression products in the host, and the gene expression products can be localized to the synthetic organelles.
[0067] In one embodiment, the application includes improving the physiological activity and tolerance of the host, biomass under the same conditions, gene expression rate, and cell growth rate.
[0068] In another embodiment, the above method or application includes promoting the maturation of the gene expression product, preferably the maturation includes one or more of the following: folding of biological macromolecules; post-expression modification, such as disulfide bond formation, glycosylation modification, methylation modification, acetylation modification; cleavage, connection, such as zymogen cleavage activation, RNA cleavage, intein-mediated protein cleavage, self-cyclization; non-covalent assembly, such as dimerization, multimerization. Beneficial effects
[0069] The technical solution of the present invention brings multiple beneficial effects, aims to solve the problems existing in the prior art, and provides a more efficient and feasible method to eliminate the stress and / or toxicity of the expression product on the host and provide a suitable and controllable maturation environment for the expression product, thereby optimizing gene expression. The following are some of the beneficial effects of the technical solution of the present invention:
[0070] ① High yield: The technical solution of the present invention can increase expression without increasing host burden. This means that higher yields of expression products can be achieved, thus meeting the demand for high expression in industrial and research applications.
[0071] ② Reducing the adverse effects of toxic products: The methods of the present invention can optimize the expression of toxic products and mitigate their adverse effects on the host. This can improve host health, reduce mortality, and facilitate the production of more host-toxic expression products, playing an important role in drug development and bioengineering.
[0072] ③ Reduce misfolding: The method of the present invention can prevent the misfolding of exogenous proteins in the cytoplasm and reduce the formation of insoluble substances such as inclusion bodies.
[0073] ④ Providing a correct modification environment: The method of the present invention can provide a suitable environment for post-expression modification of the product, such as disulfide bond formation.
[0074] ⑤ High expression consistency: The technical solution of the present invention can provide more consistent protein expression effects and reduce batch-to-batch variability. This contributes to stable experimental results and repeatable production processes.
[0075] ⑥ Resource efficiency: By reducing the waste of energy and cell resources, the method of the present invention can improve expression efficiency and reduce research and production costs.
[0076] ⑦Wide applicability: Unlike some existing methods that are only applicable to specific types of proteins or cell systems, the technical solution of the present invention has broad applicability and is applicable to various proteins and cell types.
[0077] In summary, the technical solutions of this invention aim to provide a comprehensive and efficient approach to address key issues in cell expression, offering more feasible and efficient solutions for applications in scientific research, biopharmaceuticals, and other related fields. These beneficial effects will promote technological development and provide new opportunities for innovation and progress in the field of biological sciences. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] Figure 1 shows the growth curves of Escherichia coli (E. coli) under three expression conditions. The blue line curve represents the growth curve when GFP is not expressed, the green line represents the growth curve when only tdMCP-GFP is induced to express, and the red line represents the growth curve after co-expression of tdMCP-GFP and rCAG organelles. The vertical line shows the standard deviation std, the horizontal axis represents the growth time, and the vertical axis represents the absorbance value at 600nm.
[0079] Figure 2 shows the gel electrophoresis image when the target protein is DpnI. M represents a protein marker, P represents an insoluble precipitate, S represents a purified soluble protein product, and the red asterisk marks the DpnI protein (30 kDa).
[0080] Figure 3 shows a gel electrophoresis image of the target protein, BsaI. M represents a protein marker, P represents the insoluble precipitate, S represents the purified soluble protein product, a red asterisk marks the BsaI protein (~60 kDa), and a blue asterisk marks the tdMCP-BsaI before cleavage.
[0081] Figure 4 shows the gel electrophoresis image when the target protein is proinsulin. M represents the protein marker, P represents the insoluble precipitate, S represents the purified soluble protein product, and the red asterisk marks the proinsulin protein (12 kDa).
[0082] Figure 5 shows the gel electrophoresis image when the target protein is ROR-alpha. M is a protein marker; S is the purified soluble protein product, and the red asterisk marks the ROR-alpha protein (~60 kDa). DETAILED DESCRIPTION
[0083] The present invention will be described in more detail below.
[0084] 1. Synthetic organelles
[0085] Cells typically possess complex spatial organization, one form of which is called compartmentalization. These compartmentalized structures, often called organelles, are to cells what organs are to individuals.
[0086] Organelles are either individually enclosed in their own lipid bilayers, called membrane-bound organelles, such as mitochondria and chloroplasts; or they are spatially independent functional units without a lipid bilayer around them, called membraneless organelles, which are also compartments observable under an optical microscope. In 2009, Hyman and Brangwynne discovered and revealed that membraneless organelles are formed by phase separation of biological macromolecules within the cell (1.CP Brangwynne et al., Germline P granules are liquid droplets that localize by controlled dissolution / condensation. Science 324, 1729-1732 (2009)).
[0087] Synthetic organelles are an innovative bioengineering concept, also known as "organelle-like structures," "artificial organelles," or "artificial organelles." Synthetic organelles aim to design and construct tiny functional units similar to natural organelles, designed to achieve compartmentalization within cells (or in vitro environments) and to control biosynthesis, metabolic pathways, or other biological processes. The construction of synthetic organelles involves introducing exogenous DNA constructs into host cells. These exogenous DNA, including exogenous coding genes, noncoding DNA, and regulatory elements, is used to guide cell-free expression systems or host cells to synthesize specific peptides, proteins, DNA, RNA, enzymes, or other biomolecules. Some of these molecules contribute to the construction of the synthetic organelle structure, while others carry out the synthetic organelle's intended functions, such as biocatalysis, receptor recognition, and protein sorting. Synthetic organelles can be formed through the self-assembly of heteropolymers of proteins, DNA, and RNA, or by modifying the cell's membrane structure. Since the phase separation mechanism of membraneless organelles was revealed, the construction of synthetic organelles can also be formed by the aggregation of biological macromolecules that can undergo phase separation.
[0088] The present invention provides methods for decoupling gene expression products from the host's internal environment using synthetic organelles. Synthetic organelles can be constructed primarily from polypeptides and / or proteins and / or DNA and / or RNA and / or nucleic acid-protein complexes. For example, after intracellular transcription, the repeating triplet of RNA (CAG)n spontaneously aggregates through liquid-liquid phase separation to form RNA condensates / droplets. Synthetic organelles can also contain other substances depending on the characteristics of the target product. For example, the expression of the P9 and P12 proteins of φ6 bacteriophage can recruit lipid molecules to form a membrane structure.
[0089] The biomacromolecule components for constructing synthetic membraneless organelles are selected from the following groups: nucleotide sequences repeated 10-1000 times continuously, such as triplets CAG, CCUG, GGGAA, GGGGCC, etc.; amino acid sequences repeated 3-500 times continuously, such as resilin-like peptides (RLP) repeats; intrinsic non-structural regions, such as FUS protein, ParB protein; multimerization of multiple biomacromolecules, such as PTB protein and UCUCU sequence repeats, artificially designed binary protein two-dimensional structures, and artificially designed binary RNA two-dimensional structures.
[0090] Of course, the synthetic organelles of the present invention can also be constructed from polypeptides and / or proteins and / or DNA and / or RNA and / or nucleic acid-protein complexes of other sequences, for example:
[0091] RNA:
[0092] o Repeating triplet RNAs, such as (CAG)n, (GGU)n, (AUG)n, (CGG)n, (GUU)n, (CGU)n, (AUU)n, (ACG)n, (AAU)n, (AUC)n, (CCG)n, (AGC)n, (AGU)n, (ACU)n, where n is 10-1000;
[0093] ○ Repeating quadruplex RNA, such as (CCUG)n, where n is 10 to 1000;
[0094] ○ Repeating pentamido RNA, such as (GGGAA)n, where n is 10-1000;
[0095] Repeating hexanucleotides, such as GGGGCC (rGGGGCC), i.e., (GGGGCC)n, (AGGGCC)n, (GGAGCC)n, (GGGGAC)n, (AGGACC)n, where n is 10-1000;
[0096] ○ Monobasic synthetic RNA scaffold, for example (underlined and bold "XX--XX" represent sites where a functional nucleic acid aptamer or ribozyme can be inserted): GGGTAGGCGCCTAGCCTAATGTACATTAAGTTATTTTTCCGGATGAATAGAATATATTCTAATAACGCAGGXX--XXCCTCGAATACGAGCTGGGXX--XXCCCAGGAAGTGTTCGCACTTCTCTCGTATTCGATTGCGACTAGT;
[0097] ○ Binary synthetic RNA scaffolds, for example (underlined and bolded "XX--XX" represent sites where functional aptamers or ribozymes can be inserted): d2-1 GGGTCAGGAATCCTCCTGATAGCTATTTGGACAATTACGTACGTAGTTGATGACAACTACATGAAAATAAGGGXX--XXCCCTCTAGA and d2-2 GGGTAGTTGTTATGGATTCCTGATTTATGGGXX--XXCCACTAGT;
[0098] DNA
[0099] ○ ssDNA sequence identical to the above RNA sequence;
[0100] ○ ssDNA capable of forming nanopolymers, such as the following four ssDNA nanopolymers (the underlined "XX--XX" represents the site where a functional nucleic acid aptamer or ribozyme can be inserted): oligo-1 TGCGCAATCCCXX--XXCTTGAGCACGCCAACATCACCGTATTT, oligo-2 GCGGCTAGCAGAGCATTCGGGAXX--XXGACTATGGCTGTATTT, oligo-3 GGCGTGCTCACXX--XXTCGGATTGCGCATGCTAGCCGCTATTT, oligo-4 CGGTGATGTTCAGCCATAGTAGXX--XXGCCGAATGCTCTATTT, etc.;
[0101] Protein
[0102] Self-assembled protein nanoparticles and nanostructures, such as carboxysome shell proteins; metabolosome shell proteins; bacterial microcompartment (BMC) shell proteins BMC-H, BMC-T, and BMC-P; Haliangium ochraceum bacterial microcompartment capsid protein T1 (HO BMC Shell T1); Thermotoga maritima encapsulation protein EncTm; hybrid bacteriophage Qβ virus-like particles; ferritin; synthetic protein nanocages, such as antibody Fc nanocages; synthetic protein self-assembled fibers, such as CC-Di-B–PduA protein; synthetic protein self-assembled two-dimensional structures; synthetic protein self-assembled three-dimensional structures, such as ATC-HL3 protein;
[0103] Intrinsically disordered proteins (IDPs), such as resilin-like peptides (RLPs); elastin-like peptides (ELPs); RNA helicase LAF-1 RGG domain repeats; RNA helicase DDX4; Caulobacter crescentus nuclease E (RNaseE); IDP homologous protein sequence (GRGDSPYS)n, and mutants (GRGDSPYSGRGDSPYSGRGDSPYSGRGDSPVS)n, (GRGDSPYSGRGDSPVS)n, where n is 3-500; ParB protein; FUS protein; IbpA protein; PopZ protein; PodJ protein; conserved nuclear polar protein fibrillarin FIB-1;
[0104] ○ Protein amyloid deposition, such as AEAEAKAKAAEAEAKAK, LELELKLKLELELKLK, human b-amyloid peptide Ab42 (F19D);
[0105] Other proteins that can aggregate include: Escherichia coli beta-galactosidase; MalE31 maltose binding protein mutant; Clostridium cellulovorans cellulose binding domain (CBD); Clostridium thermocellum endoglucanase D (EGD); Brucella abortus)-derived proteins and fluorescent proteins, PdhS-mCherry, FumC-YFP, DivK-YFP; Escherichia coli (E. coli)-derived proteins and fluorescent proteins, ClpX-sfGFP, ClpP-sfGFP, ClpP-mCherry, ClpX-mCherry, ClpX-mCherry2, ClpX-Venus, LacZ-Dronpa, TetR-Dronpa, P22C2-Dronpa, HKCI-Dronpa; tandem repeat SH3 (polySH3) and tandem repeat PRM (polyPRM); tandem repeat SIM (polySIM) and tandem repeat SUMO (polySUMO); pentameric nuclear polar protein Npm1, etc.
[0106] ●Nucleic acid protein polymer complex:
[0107] Aggregates of long non-coding RNA (lncRNA) and RNA-binding proteins, including nucleic acid scaffolds such as mammalian nuclear paraspeckle assembly transcript 1 isoform 2 in paraspeckles; intergenic spacer lncRNA in amyloid bodies; human satellite III (SatIII) lncRNA in nuclear stress bodies; Drosophila heat shock RNA (Hsr) ω; fission yeast meiRNA; RNA-binding proteins such as PTB protein; FUS protein; and histone BRD4.
[0108] ○ Peptide sequences and polynucleotides, such as the peptide RRASLRRASL and polyuridylic acid (polyU) with a molecular weight of 600–1,000 kDa; synthetic peptides RP3, SR8, and polyU, etc.
[0109] Synthetic organelles can be constructed entirely from biological molecules synthesized by the host under the guidance of exogenous DNA constructs, or they can be constructed by combining some exogenous molecules with endogenous molecules of the host, or they can be modified from the original endogenous organelles or cytoskeleton structures in the host. For example, in a method for achieving codon expansion using synthetic organelles disclosed in EP 19157257.7, the synthetic organelles are formed by exogenously expressed PylRS::FUS::KIF16B1-400 fusion protein, MCP::EWSR1::KIF16B1-400 fusion protein and endogenous microtubule cytoskeleton structure.
[0110] 2. Co-localization of synthetic organelles and target products
[0111] Organelles generally have the properties of selective permeability and biomolecule sorting. Therefore, the biomolecule composition and physicochemical properties within the organelle compartments are significantly different from those of components outside the organelles (such as the cytoplasmic matrix), enabling the organelles to perform special envoy functions.
[0112] In the present invention, the target product can be localized to the artificial organelle through non-specific and / or specific binding. Preferably, a specific receptor / ligand pair is used. Both the receptor and the ligand can be selected as needed. For example, in one embodiment, the recruited ligand is a tandem-dimer MS2 coat protein (tdMCP) and the receptor is an MS2 hairpin RNA aptamer. Other ligand / receptor pairs can also be used as needed, for example:
[0113] ● Nucleic acid molecules with complementary sequences;
[0114] RNA-binding proteins and their corresponding RNAs, such as:
[0115] ○ RNA phage capsid protein and translation operon RNA hairpin, including RNA phages MS2, PP7, Qbeta, etc.
[0116] ○ DNA phage protein N and B box (boxB) RNA hairpin, DNA phages include λ, 21, P22, etc.
[0117] CRISPR Cas proteins and CRISPR RNA hairpins, such as Csy4 (Cas6f) and the PA14 RNA hairpin from Pseudomonas aeruginosa PA14 strain;
[0118] ○ Other binding pairs, such as PUF-8 and 5′-UGUANAUA-3′; FBF-2 and 5′-UGURNNAUA-3′; protein A and protein A binding aptamer; archaeal RNA binding protein L7Ae and RNA C / D box, etc.
[0119] DNA-binding proteins and their corresponding DNA, for example: TetR and the tetO operator; cI repressor protein and the cI operator; Gal4 and its corresponding DNA; Zif268 DNA-binding domain and the GATGCTGCA sequence; Gli-1 DNA-binding domain and the GACCACCCAAGACGA sequence; LexA DNA-binding domain and the CTGTATATATATACAG sequence; Transcription activator-like (TAL) domain and its corresponding DNA;
[0120] ●Homodimers or multimers, such as the reversible green fluorescent protein Dronpa;
[0121] Split proteins, for example: split fluorescent proteins, such as green fluorescent proteins GFP1-10 and GFP11; split β-galactosidase; split T7 polymerase; split esterase; split TEV protease; split dihydrofolate reductase; split β-lactamase; split firefly luciferase; split thymidine kinase; split chorismate mutase; split CRISPR-Cas9; split horseradish peroxidase, etc.
[0122] ●Antigens and specific antibodies, such as: His tag and anti-his antibody; flag tag and anti-flag antibody; HA tag and anti-HA antibody; Myc tag and anti-Myc antibody; GST protein and anti-GST antibody;
[0123] Other protein-protein binding pairs, such as the SYNZIP peptide pair; the SZ17 and SZ18 pair; the Fos and Jun pair; and the Arabidopsis thaliana phytochrome B (PhyB) and PIF3 or PIF6 proteins.
[0124] Of course, receptors and ligands do not necessarily have to be paired between two molecules, but can also form complex multi-component complexes between multiple components, such as the quaternary complex of CRISPR-Cas9 protein, crRNA, tracrRNA, and crRNA targeting DNA; and the ternary complex of FKBP protein, FRB protein, and rapamycin.
[0125] In the synthetic organelle, the receptor can be covalently linked to the component of the synthetic organelle structure, for example, in one embodiment, the receptor MS2 hairpin RNA aptamer and the CAG repeat nucleotides of the synthetic organelle structure constitute a fusion RNA molecule, for example, in another embodiment, the receptor dCsy4 and the protein A of the synthetic organelle structure constitute a fusion protein molecule. Of course, the receptor can also be assembled into the synthetic organelle through non-covalent intermolecular interactions. For example, the CAG repeat nucleotides of the synthetic organelle structure are fused to the MS2 hairpin RNA aptamer, and the MS2 hairpin RNA aptamer non-covalently binds to the fusion protein of tdMCP-His tag, and the His tag non-publicly binds to the fusion protein of anti-His antibody and PhyB, then PhyB acts as a receptor, and the synthetic organelle can specifically bind to the protein connected with the ligand PIF3 or PIF6. The synthetic organelle can have one to multiple receptors of different properties.
[0126] The connection between the ligand and the expression product can be a covalent connection. For example, in one embodiment, the ligand tdMCP and the target product green fluorescent protein constitute a fusion protein. Of course, the ligand can also be connected to the target product through non-covalent intermolecular interactions. For example, through intermolecular interactions, the target product Dronpa145N mutant can form a heterodimer with the ligand tdMCP-Dronpa145K, which can be specifically recruited by synthetic organelles with MS2 hairpin RNA. The target product can also be connected to one or more ligands.
[0127] 3. Gene expression products
[0128] The gene expression product involved in the present invention can be RNA, protein, nucleic acid-protein complex, or a small molecule encoded by heterologous DNA, preferably a recombinant protein, such as superfolded green fluorescent protein, restriction endonuclease DpnI, restriction endonuclease BsaI, human proinsulin, ROR-alpha, red fluorescent protein mKate2, T4 DNA ligase, artificially designed polypeptide; or recombinant RNA, such as Escherichia coli 16S rRNA; or DNA, such as recombinant plasmid, single-stranded DNA; or a complex formed by a recombinant macromolecule and an endogenous macromolecule, such as a 30S ribosomal subunit formed by the assembly of recombinantly expressed 16S rRNA and endogenous ribosomal proteins; or other cellular structures different from the synthetic organelles, such as the cytoskeleton, other organelles.
[0129] 4. DNA constructs
[0130] The present invention introduces an artificially designed DNA construct into the host. The DNA construct is a polynucleotide containing multiple coding genes, non-coding DNA, regulatory elements, etc.
[0131] The DNA construct can guide the host to construct a synthetic organelle. For example, in one embodiment, the synthetic organelle is constructed from an RNA molecule containing CAG triplet nucleotide repeats. The DNA construct comprises a nucleotide sequence that expresses the RNA molecule, including a CAG triplet nucleotide repeat sequence, a transcription terminator, and a promoter that regulates induced expression.
[0132] The DNA construct involved in the present invention can be one or more plasmids, cosmids, artificial chromosomes, or can be integrated into a single or multiple sites in the host genome, as well as permutations and combinations of the above schemes. Preferably, the DNA construct is a plasmid, which is composed of the nucleotide sequence and plasmid vector that guide the host to construct synthetic organelles and produce the target product. The upstream promoter can be a normally expressed promoter or an inducible promoter. The plasmid vector can be a commercial vector, such as pACYC184, pACYC177, pET28a(+), pET28b(+), pET-5a(+), pET43.1a, pET-37b(+), pCDFDuet-1, pCOLADuet-1, pRSFDuet-1, pETDuet-1, pUC57, pUC19, pBAD, pBluescript II SK(+), pTrcHis C, pTrcHis A, pTrcHis2C, pBV221, pQE-70, pCold III, pRSET-CFP, pRSET-BFP, pGFPuv, pKD46, pKD4, pTYB1, PinPoint Xa-2, pTWIN1 or pRSET C, or a modified structure of a commercial vector, or a non-commercial vector, such as pSB1C3, pSB3C5, pSB4C5, pSB4K5, or other artificially designed nucleotide sequences that can stably replicate in host cells.
[0133] 5. Expression System
[0134] The above-mentioned DNA construct is introduced into a host for expression, guiding the host to construct a synthetic organelle, and / or regulating the host to produce the target product. The expression system can be a cell-free expression system and / or a cell expression system. Wherein the cell expression system includes prokaryotic cells and eukaryotic cells, eukaryotic cells such as yeast, fungal cells, animal cells or plant cells, or prokaryotic cells, such as Escherichia coli, Bacillus subtilis, lactic acid bacteria, Bifidobacterium, Streptomyces. The prokaryotic expression system has the advantages of rapid cell propagation, low culture cost and high yield, and is preferred, wherein Escherichia coli is more preferred.
[0135] 6. Methods for decoupling gene expression products from the host's internal environment
[0136] The present invention provides a method for decoupling gene expression products from the host's internal environment using synthetic organelles, which comprises the following steps:
[0137] a. Constructing synthetic organelles in the host,
[0138] b. localizing the gene expression product to the synthetic organelle so that the synthetic organelle and the gene expression product are co-expressed.
[0139] More specifically, step a refers to introducing a DNA construct into the host to guide the host to construct a synthetic organelle; step b refers to localizing the gene expression product to the synthetic organelle so that the synthetic organelle and the gene expression product are co-expressed; preferably, steps a and b are the process of simultaneously constructing a synthetic organelle in a single host cell, biosynthesizing the target product, and specifically enriching the target product by the synthetic organelle.
[0140] 7. Application
[0141] The present invention has two main uses. First, it improves the effects of expression products on the host, including improving host physiological activity and tolerance, biomass, gene expression rate, and cell growth rate under equivalent conditions. For example, in one embodiment, the use of synthetic organelles can eliminate the general host stress caused by sfGFP gene expression. In another embodiment, the use of synthetic organelles can prevent the expression of N-terminally bound toxic proteins.
[0142] On the other hand, it is to avoid the influence of the host's internal environment on the expression product, including promoting the maturation of the gene expression product, preferably, the folding of biological macromolecules; post-expression modification, such as disulfide bond formation, glycosylation modification, methylation modification, acetylation modification; cleavage, connection, such as zymogen cleavage activation, RNA cleavage, intein-mediated protein cleavage, self-cyclization; non-covalent assembly, such as dimerization, polymerization. For example, in one embodiment, the use of synthetic organelles provides an environment for the generation of disulfide bonds to reduce protein misfolding. For example, in another embodiment, the use of synthetic organelles promotes the folding of ROR-alpha to obtain soluble proteins,
[0143] Therefore, the present invention provides a universal solution that can take into account high expression, low stress / toxicity, correct folding, and post-expression modification.
[0144] Example
[0145] Example 1. GFP Example
[0146] System Design
[0147] In this example, superfolder green fluorescent protein (sfGFP) is expressed. It is a widely used fluorescent protein tag and is generally considered non-cytotoxic. The molecule that constitutes the synthetic organelle in this example is an RNA fragment with 46 CAG repeats, which can interact to produce liquid-liquid phase separation. In this example, tandem-dimer MS2 coat protein (tdMCP) serves as a ligand and an MS2 RNA hairpin serves as a receptor to enrich superfolder green fluorescent protein (sfGFP). tdMCP is linked to the N-terminus and C-terminus of sfGFP via protein fusion linkers. The corresponding expression vector in this example is pET28a, and the corresponding selection resistance is kanamycin resistance.
[0148] Experimental Procedure
[0149] pGFP and prCAG-MS2 were co-transformed into Escherichia coli BL21 (DE3) competent cells (Full Gold). Three single clones were picked for each protein and cultured at 37°C and 220 rpm for 16 hours. After the overnight bacteria were diluted 200 times and cultured for 3 hours in 96-well plates and 100 mL shake flasks, the inducer was added to 0.5 mM isopropylthio-β-galactoside (IPTG) (Sangon) and 150 ng / ml anhydrotetracycline hydrochloride (aTc) (Bai Lingwei), and expression was induced and cultured overnight for about 12 hours. The 96-well plate was cultured in a Tecan Spark microplate reader and the growth curve was measured overnight. Expression was induced and cultured overnight for about 12 hours.
[0150] result
[0151] The results, shown in Figure 1, indicate that co-expression of synthetic organelles can significantly enhance host cell growth, demonstrating that the use of synthetic organelles can eliminate the typical host stress caused by gene expression.
[0152] Sequence Details
[0153] The size of sfGFP is 26 kDa, and the protein sequence is shown in SEQ ID NO: 1 below:
[0154] The RNA sequence of rCAG-MS2 is shown in SEQ ID NO: 2 below:
[0155] The protein sequence of tdMCP is shown in SEQ ID NO: 3 below:
[0156] Example 2. DpnI Example
[0157] System Design
[0158] In this example, the restriction endonuclease DpnI, which cleaves DNA chains with 6mA methylated GATC, is expressed and is lethal to the common Escherichia coli expression host. The molecule that constitutes the synthetic organelle in this example is an RNA fragment with 46 CAG repeats, which can undergo liquid-liquid phase separation through interaction. In this example, the tandem-dimer MS2 coat protein (tdMCP) serves as a ligand and the MS2 RNA hairpin serves as a receptor to enrich the restriction endonuclease DpnI, which recognizes and cleaves DNA chains with 6mA methylated GATC. tdMCP is linked to the N-terminus of DpnI via a protein fusion linker, where the fusion hinge utilizes a SNAC-tag that is self-cleavable by divalent nickel ions. The corresponding expression vector in this example is pET28a, and the corresponding selection resistance is kanamycin resistance.
[0159] Experimental Procedure
[0160] pET28a-DpnI and prCAG-MS2 were co-transformed into Escherichia coli BL21 (DE3) competent cells (full gold). Three single clones were picked for each protein and cultured at 37°C and 220 rpm for 16 hours. After the overnight bacteria were diluted 200 times and cultured for 3 hours, the inducer was added to 0.5 mM isopropylthio-β-galactopyranoside (IPTG) (Sangon) and 150 ng / ml anhydrotetracycline hydrochloride (aTc) (B&K), and expression was induced and cultured overnight for about 12 hours. pET28a-DpnI and prCAG-MS2 were co-transformed into Escherichia coli BL21 (DE3) competent cells (full gold). Three single clones were picked for each protein and cultured at 37°C and 220 rpm for 16 hours. After the overnight bacteria were diluted 200 times and cultured for 3 hours, the inducer was added to 0.5mM isopropylthio-β-galactoside (IPTG) (Sangong) and 150ng / ml hydrochloric acid dehydrated tetracycline (aTc) (Bai Lingwei), and the expression was induced and cultured overnight for about 12 hours. The bacterial solution was centrifuged and the supernatant was removed. The lysis solution (one-step bacterial lysis kit, Sangong) was added to the precipitate, pipetted to mix, reacted at room temperature for 30 minutes, and then centrifuged, the supernatant was discarded, and the precipitate was recovered. A cutting eluent containing nickel ions was added to the precipitate obtained in the previous step, and it was placed on a rotary mixer for cutting at room temperature. The sample after the cutting reaction was centrifuged and the supernatant was recovered, which was the solution containing the target product protein. SDS PAGE electrophoresis was used for detection (SurePAGE TM , GenScript), was used to verify the solubility of the target protein and evaluate the expression level of the recovered protein.
[0161] result
[0162] The results, as shown in Figure 2, demonstrate that DpnI is abundantly expressed in the E. coli host as a soluble protein. This indicates that the synthetic organelles fully recruit DpnI expressed in the cytoplasm, enabling the E. coli host to efficiently reproduce and achieve sufficient biomass despite overexpression of DpnI, thereby increasing overall DpnI production. This demonstrates that co-expression of synthetic organelles with N-terminally bound toxic proteins can eliminate the toxicity of toxic proteins to the host.
[0163] Sequence Details
[0164] The protein sequence of DpnI is shown in SEQ ID NO: 4 below:
[0165] The SNAC-tag protein sequence is shown in SEQ ID NO: 5 below:
[0166] Example 3. BsaI Examples
[0167] System Design
[0168] In this example, the restriction endonuclease BsaI, which specifically recognizes GGTCTC and cuts the DNA chain, is expressed and is toxic to common Escherichia coli hosts. The molecule that constitutes the synthetic organelle in this example is an RNA fragment with 46 CAG repeats, which can produce liquid-liquid phase separation through interaction. In this example, the tandem-dimer MS2 coat protein (tdMCP) is used as a ligand and the MS2 RNA hairpin is used as a receptor to enrich the restriction endonuclease BsaI that recognizes and cuts the DNA chain. tdMCP is connected to the N-terminus of BsaI through a protein fusion hinge (fusion linker), in which the fusion hinge uses a SNAC-tag that can be induced to self-cleave by divalent nickel ions. The corresponding expression vector in this example is pET28a, and the corresponding screening resistance is kanamycin resistance.
[0169] Experimental Procedure
[0170] pET28a-tdMCP-BsaI and prCAG-MS2 were co-transformed into competent Escherichia coli BL21(DE3) cells (Full Gold). Three single clones were selected for each protein and cultured at 37°C with shaking at 220 rpm for 16 hours. Overnight cells were diluted 200-fold and cultured for 3 hours. Induction agents were added, including 0.5 mM isopropylthio-β-galactopyranoside (IPTG) (Sangon) and 150 ng / ml anhydrotetracycline hydrochloride (aTc) (J&K), and expression was induced. Culture was continued overnight for approximately 12 hours. The bacterial culture was centrifuged and the supernatant removed. Lysis buffer (One-Step Lysis Kit, Sangon) was added to the pellet, mixed by pipetting, and incubated at room temperature for 30 minutes. The pellet was then centrifuged, the supernatant discarded, and the pellet recovered. A nickel-containing cleavage eluent was added to the pellet obtained in the previous step and cleaved on a rotary mixer at room temperature. The cleavage reaction sample was centrifuged, and the supernatant, containing the target protein, was recovered. SDS PAGE electrophoresis (SurePAGE TM , GenScript), was used to verify the solubility of the target protein and evaluate the expression level of the recovered protein.
[0171] result
[0172] The results, as shown in Figure 3, demonstrate that BsaI achieves substantial soluble expression in the E. coli host. This demonstrates that the synthetic organelles fully recruit BsaI expressed in the cytoplasm, enabling the E. coli host to effectively reproduce and achieve sufficient biomass despite overexpression of BsaI, thereby increasing overall BsaI production. This suggests that co-expression of the synthetic organelles with an N-terminally bound toxic protein can eliminate the toxicity of the toxic protein to the host.
[0173] Sequence Details
[0174] The protein sequence of BsaI is shown in SEQ ID NO: 6:
[0175] Example 4. Proinsulin Example
[0176] Experimental design
[0177] In this example, human proinsulin, which has two disulfide bonds, is expressed in the common Escherichia coli host, resulting in the formation of insoluble inclusion bodies due to the inability to form disulfide bonds. The molecule that constitutes the synthetic organelle in this example is an RNA fragment with 46 CAG repeats. This RNA can interact to produce liquid-liquid phase separation and maintain an oxidative environment by sequestering the enzymes in the E. coli itself that provide reducing potential to the cytoplasm. In this example, proinsulin is enriched using the tandem-dimer MS2 coat protein (tdMCP) as a ligand and the MS2 RNA hairpin as a receptor. tdMCP is linked to the N-terminus of proinsulin via a protein fusion linker, in which the fusion hinge utilizes a SNAC-tag that self-cleaves in response to divalent nickel ions. The corresponding expression vector in this example is pET28a, and the corresponding selection resistance is kanamycin resistance.
[0178] Experimental Procedure
[0179] pET28a-tdMCP-proinsulin and prCAG-MS2 were co-transformed into competent Escherichia coli BL21(DE3) cells (Full Gold). Three single clones were selected for each protein and cultured at 37°C with shaking at 220 rpm for 16 hours. Overnight cells were diluted 200-fold and cultured for 3 hours. Induction agents were added, including 0.5 mM isopropylthio-β-galactopyranoside (IPTG) (Sangon) and 150 ng / mL anhydrotetracycline hydrochloride (aTc) (J&K), and expression was induced. Culture was continued overnight for approximately 12 hours. The bacterial culture was centrifuged and the supernatant removed. Lysis buffer (One-Step Lysis Kit, Sangon) was added to the pellet, mixed by pipetting, and incubated at room temperature for 30 minutes. The pellet was then centrifuged, the supernatant discarded, and the pellet recovered. A nickel-containing cleavage eluent was added to the pellet obtained in the previous step and cleaved on a rotary mixer at room temperature. The cleavage reaction sample was centrifuged, and the supernatant, containing the target protein, was recovered. SDS PAGE electrophoresis (SurePAGE TM , GenScript), was used to verify the solubility of the target protein and evaluate the expression level of the recovered protein.
[0180] result
[0181] The results, as shown in Figure 4, demonstrate that soluble human proinsulin was expressed in the E. coli host. This suggests that the synthetic organelles provide an environment for disulfide bond formation, reduce protein misfolding, promote folding of human proinsulin and disulfide bond formation, and yield soluble human proinsulin.
[0182] Sequence Details
[0183] The protein sequence of human proinsulin is shown in SEQ ID NO: 7:
[0184] Example 5. ROR-alpha Example
[0185] Experimental design
[0186] The protein expressed in this example, ROR-alpha, is a human nuclear receptor. This protein does not have any special modifications, but it is prone to misfolding. When expressed in a common Escherichia coli host, it forms insoluble inclusion body precipitates. The molecule that constitutes the synthetic organelle in this example is an RNA fragment with 46 CAG repeats, which can produce liquid-liquid phase separation through interaction. In this example, tandem-dimer MS2 coat protein (tdMCP) is used as a ligand and the MS2 RNA hairpin is used as a receptor to enrich ROR-alpha. tdMCP is connected to the N-terminus of ROR-alpha through a protein fusion hinge (fusion linker). The corresponding expression vector in this example is pET28a, and the corresponding screening resistance is kanamycin resistance.
[0187] Experimental Procedure
[0188] pET28a-tdMCP-ROR-alpha and prCAG-MS2 were co-transformed into competent Escherichia coli BL21(DE3) cells (Full Gold). Three single clones were selected for each protein and cultured at 37°C with shaking at 220 rpm for 16 hours. Overnight cells were diluted 200-fold and cultured for 3 hours. Induction agents were added, including 0.5 mM isopropylthio-β-galactopyranoside (IPTG) (Sangon) and 150 ng / ml anhydrotetracycline hydrochloride (aTc) (J&K), and expression was induced. Culture was continued overnight for approximately 12 hours. The cell suspension was centrifuged and the supernatant removed. Lysis buffer (One-Step Lysis Kit, Sangon) was added to the pellet, mixed by pipetting, and incubated at room temperature for 30 minutes. The pellet was then centrifuged, the supernatant discarded, and the pellet recovered. A nickel-containing cleavage eluent was added to the pellet obtained in the previous step and cleaved on a rotary mixer at room temperature. The cleavage reaction sample was centrifuged, and the supernatant, containing the target protein, was recovered. SDS PAGE electrophoresis (SurePAGE TM , GenScript), was used to verify the solubility of the target protein and evaluate the expression level of the recovered protein.
[0189] result
[0190] The results are shown in Figure 5, demonstrating that human ROR-alpha was solublely expressed in the E. coli host. This suggests that the synthetic organelles reduced protein misfolding, promoted the folding of ROR-alpha, and produced a soluble protein.
[0191] Sequence Details
[0192] The protein sequence of human ROR-alpha is shown in SEQ ID NO: 8:
Claims
1. A method for gene expression assisted by synthetic organelles, comprising the following steps: a. Construct synthetic organelles in the host, b. co-expressing the synthetic organelle with a gene expression product, wherein the gene expression product can be localized to the synthetic organelle.
2. The method according to claim 1, characterized in that The host includes a cell-free expression system and a cell host, preferably a prokaryotic cell host, more preferably Escherichia coli.
3. The method according to claim 1, characterized in that The synthetic organelle refers to an artificially designed compartmentalized structure. Preferably, the synthetic organelle is a membraneless organelle, which is a biomolecular condensate constructed in a host by a polypeptide and / or protein and / or DNA and / or RNA and / or nucleic acid-protein complex. Preferably, the biomacromolecules used to construct the synthetic organelle are nucleotide sequences that are repeated 10-1000 times continuously, such as triplets CAG, CCUG, GGGAA, GGGGCC, etc.; amino acid sequences that are repeated 3-500 times continuously, such as resilin-like peptides (RLP) repetitions; intrinsic non-structural regions, such as FUS protein, ParB protein; multimerization of multiple biomacromolecules, such as PTB protein and UCUCU sequence repetitions, artificially designed binary protein two-dimensional structures, and artificially designed binary RNA two-dimensional structures.
4. The method according to claim 1, characterized in that The gene expression product is RNA, protein, nucleic acid-protein complex, or a small molecule synthesized by heterologous DNA encoding, preferably a recombinant protein, for example, superfolded green fluorescent protein, restriction endonuclease DpnI, restriction endonuclease BsaI, human proinsulin or ROR-alpha, preferably, the gene expression product matures in the synthetic organelle.
5. The method according to any one of claims 1 to 4, characterized in that In the step b, the gene expression product is localized to the synthetic organelle by non-specific and / or specific binding. Preferably, a specific receptor / ligand pair is used, the synthetic organelle has a receptor, and the gene expression product is connected to a ligand, so that the synthetic organelle can specifically bind to the gene expression product. Preferably, the receptor / ligand pair includes: complementary paired nucleic acid molecules; capsid proteins and translation operon RNA hairpins of RNA phages such as MS2, PP7, and Qbeta; two segments of a split protein, such as split green fluorescent protein, the alpha segment and the omega segment of beta-galactosidase, and split T7 RNA polymerase.
6. A DNA construct comprising a nucleotide sequence 1 that guides a host to construct a synthetic organelle as defined in the method of any one of claims 1 to 5, and / or a nucleotide sequence 2 that guides and regulates the generation of a gene expression product as defined in the method of any one of claims 1 to 5, preferably, the DNA construct is composed of the nucleotide sequences 1, 2 and a plasmid vector, such as pACYC184, pACYC177, pET28a(+), pET28b(+), pET-5a(+), pET43.1a, pET-37b(+), pCDFDuet-1, pCOLADuet-1, pRSFDuet-1, pETDuet-1, pUC57, pUC19, pBAD, pBluescript II SK(+), pTrcHis C, pTrcHis A, pTrcHis2C, pBV221, pQE-70, pCold III, pRSET-CFP, pRSET-BFP, pGFPuv, pKD46, pKD4, pTYB1, PinPoint Xa-2, pTWIN1, pRSET C, pSB1C3, pSB3C5, pSB4C5, pSB4K5, preferably, the upstream promoter is a normally expressed promoter or an inducible promoter.
7. An expression system comprising the DNA construct as claimed in claim 6, preferably, the expression system refers to a system for simultaneously constructing the synthetic organelles and generating the gene expression products in prokaryotic cells, and more preferably, the expression system is an Escherichia coli expression system.
8. The use of synthetic organelles in eliminating the stress and / or toxicity of gene expression on the host and / or decoupling the gene expression product from the host's internal environment, which comprises co-expressing the synthetic organelles and the gene expression product in the host, and the gene expression product can be localized to the synthetic organelles.
9. The use according to claim 8, characterized in that The application includes improving the physiological activity and tolerance of the host, the biomass under the same conditions, the gene expression rate and / or the cell growth rate.
10. The method according to claim 4 or the use according to claim 8, characterized in that: The gene expression product matures in the synthetic organelle, and preferably the maturation includes one or more of the following: folding of biomacromolecules; post-expression modification, such as disulfide bond formation, glycosylation modification, methylation modification, acetylation modification; cleavage and connection, such as zymogen cleavage activation, RNA cleavage, intein-mediated protein cleavage, and self-cyclization; non-covalent assembly, such as dimerization and multimerization.
Citation Information
Patent Citations
Construction and application of membraneless organelles in prokaryotes
CN112575016A
Means and methods for preparing engineered target proteins by genetic code expansion in a target protein selective manner
EP3696189A1
Designer membraneless organelles sequester native factors for control of cell behavior
WO2023015190A1
Synthesis of CO2 immobilized class organelle in bacteria and application thereof
CN110563825A
Synthetic method of biological photosynthetic organelle as well as product and application of biological photosynthetic organelle
CN114634877A