A method for high production of products of biosynthetic genes or gene clusters based on chromatin three-dimensional structure
By integrating target genes or gene clusters into the three-dimensional structure of chromatin, and utilizing chromatin interaction frequency and transcriptional information, the problem of gene expression silencing was solved, and high production of microbial metabolites was achieved, especially a significant increase in the production of Streptomyces secondary metabolite RK-682.
Patent Information
- Application Number
- CN202310149184.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-22
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2043-02-22
AI Technical Summary
Existing technologies struggle to achieve high yields of microbial secondary metabolites by regulating the three-dimensional structure of the genome, especially since target genes or gene clusters are often silent or underexpressed under laboratory culture conditions.
By analyzing global chromatin interaction frequencies and transcriptional information, we can identify chromatin regions with high interaction strength and strong correlation, and integrate target genes or gene clusters into these regions to enhance their expression and yield.
It significantly improves the product yield of target genes or gene clusters, and the method is simple and reproducible, making it suitable for high-yield modification of microbial metabolites.
Smart Images

Figure CN116153406B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of genetic engineering, in particular to the field of high-yield methodology of microbial metabolites, and specifically discloses a method for high-yield of products of biosynthesis genes or gene clusters based on three-dimensional structure of chromatin. BACKGROUND
[0002] In cells, chromatin, as a carrier of genetic information, has a precise and complex three-dimensional spatial structure, and is closely related to various life activities and regulation. Microbial secondary metabolites are a rich source of human drugs, and research shows that many genes or gene clusters responsible for the biosynthesis of secondary metabolites are in a silent or low-expression state under laboratory culture conditions. In order to fully exploit and utilize this valuable resource in nature, the current mainstream methods include optimization of external culture conditions and one-dimensional internal metabolic engineering modification, but there is no method for regulating gene expression by using three-dimensional structure of genome to achieve high-yield of metabolites from a high-dimensional level. The present application utilizes the strong correlation between the higher-order structure of chromatin and the expression of target genes or gene clusters to provide a new method for high-yield of products of target genes or gene clusters. SUMMARY
[0003] The purpose of the present application is to provide a method for high-yield of products of biosynthesis genes or gene clusters based on three-dimensional structure of chromatin.
[0004] The purpose of the present application is achieved by the following technical solutions.
[0005] A method for high-yield of products of biosynthesis genes or gene clusters based on three-dimensional structure of chromatin, which determines the chromatin regions having strong correlation between chromatin interaction frequency and transcription amount by analyzing chromatin global interaction frequency information and global transcription amount information, integrates target genes or gene clusters in the strong correlation chromatin regions having high interaction intensity on the genome, and realizes the yield improvement of products of target genes or gene clusters.
[0006] Further, the method for high-yield of products of biosynthesis genes or gene clusters based on three-dimensional structure of chromatin comprises the following steps.
[0007] (1) Obtain whole genome sequence information through a database or deep sequencing.
[0008] (2) Obtain chromatin interaction frequency information of whole genome through high-throughput chromatin conformation capture technology.
[0009] (3) Obtain whole genome transcription amount information through transcriptome sequencing.
[0010] (4) Calculate the interaction intensity of chromatin regions with minimum resolution and upstream and downstream chromatin regions according to the chromatin interaction frequency information obtained in step (2).
[0011] (5) Based on the information obtained in steps (3) and (4), calculate the Pearson correlation between the interaction strength of several consecutive minimum resolution chromatin regions and the corresponding transcription level.
[0012] (6) Integrate the target gene or gene cluster into the chromatin region with strong correlation between chromatin interaction strength and corresponding transcription level and high interaction strength, so as to improve the yield of the target gene or gene cluster product.
[0013] In some embodiments, the high interaction strength is ranked in the top 30% of interaction strength.
[0014] In some embodiments, the strong correlation is a Pearson correlation ≥ 0.8.
[0015] In some embodiments, the minimum resolution is 5-10 kb, and the upstream and downstream chromatin regions are 15-200 kb upstream and downstream regions.
[0016] In some embodiments, the several consecutive minimum resolution chromatin regions are 7-13 consecutive minimum resolution chromatin regions.
[0017] In some embodiments, the calculation method of interaction strength in step (4) above includes the following steps:
[0018] 1) Calculate the effective fragment length, GC content and mappability of the genome, i.e. F_GC_MAP, using F_GC_MAP.file.sh in HiCnv software;
[0019] 2) Calculate the sum of local chromatin interaction frequency in each minimum resolution chromatin region and the upstream and downstream interval;
[0020] 3) Based on the standardization algorithm of Poisson regression, use the sum of local chromatin interaction frequency obtained in 2) and the F_GC_MAP obtained in 1) to perform HiCNormCis calculation, to obtain the interaction strength of each minimum resolution chromatin region and the upstream and downstream chromatin regions.
[0021] In some embodiments, the calculation method of Pearson correlation in step (5) above includes the following steps: setting several consecutive minimum resolution chromatin regions as a unit, and calculating the Pearson correlation between local chromatin interaction strength and corresponding transcription signal by Z transformation and Gaussian smoothing.
[0022] Further, the method for producing products of biosynthesis genes or gene clusters based on chromatin three-dimensional structure comprises the following steps:
[0023] (1) Obtain the whole genome sequence information of the target strain through a database, or obtain the whole genome sequence information of the target strain through deep sequencing.
[0024] (2) Culture the target strain cells to a specific growth period, prepare a bacterial genome sample, and obtain whole genome interaction frequency information through high-throughput chromatin conformation capture technology.
[0025] (3) Culture the target strain cells to the same growth period, prepare a bacterial mRNA sample, and obtain whole genome transcription data through transcriptome sequencing.
[0026] (4) Calculate the interaction intensity of each minimum resolution chromatin region and the upstream and downstream chromatin regions, and the calculation method is as follows:
[0027] 1) Use F_GC_MAP.file.sh in HiCnv software to calculate the effective fragment length, GC content and mappability of the genome, that is, F_GC_MAP;
[0028] 2) Calculate the sum of local chromatin interaction frequencies in each minimum resolution chromatin region and the upstream and downstream regions;
[0029] 3) Based on the standardization algorithm of Poisson regression, use the sum of local chromatin interaction frequencies obtained in 2) and the F_GC_MAP obtained in 1) to perform HiCNormCis calculation, and obtain the interaction intensity of each minimum resolution chromatin region and the upstream and downstream chromatin regions.
[0030] (5) Set several continuous minimum resolution chromatin regions as a unit, and calculate the Pearson correlation between local chromatin interaction intensity and corresponding transcription signal by Z transformation and Gaussian smoothing.
[0031] (6) Integrate the target gene or gene cluster into the chromatin region with strong correlation between chromatin region interaction intensity and corresponding transcription level and high interaction intensity, so as to realize the improvement of the yield of the product of the target gene or gene cluster.
[0032] The above method can be applied to high-yield modification and silencing activation modification of microbial metabolites.
[0033] The application also provides a high-yield RK-682 Streptomyces, which is obtained by integrating RK-682 biosynthesis genes or gene clusters at any one of the following sites in the chromosome of Streptomyces coelicolor M145 strain: 4248182 nt, 4368050 nt, 5068174 nt, 3477732 nt, 2838948 nt, 4203993 nt and 4423541 nt.
[0034] The present application has the advantages and beneficial effects that the present application can realize the promotion of the target gene or gene cluster product by targeting the integration of the target gene or gene cluster into the strong correlation chromatin region with high frequency interaction on the genome. The method of the present application has the characteristics of simple implementation, good repeatability, remarkable effect, etc., and can be widely applied to improve the yield of target gene or gene cluster product. BRIEF DESCRIPTION OF DRAWINGS
[0035] Figure 1 The results of RK-682 yield changes of RK-682 biosynthetic gene clusters integrated in different interaction intensity chromatin regions in the middle and late exponential growth of Streptomyces coelicolor are shown in the following figure: the horizontal line represents the RK-682 yield of the gene cluster inserted into the attB site as a control. The numerical value above the control line represents the ratio of RK-682 yield of the gene cluster inserted into the chromatin region with different interaction intensity to the gene cluster inserted into the control site attB. DETAILED DESCRIPTION
[0036] The following examples are used to further illustrate the present application, but should not be construed as limiting the present application. If not specifically indicated, the technical means used in the examples are conventional means known to those skilled in the art.
[0037] Example 1
[0038] The present application is described in detail taking the high-yield secondary metabolite RK-682 derived from Streptomyces sp. 88-682 in Streptomyces coelicolor A3(2) M145 strain as an example. The specific steps and results are as follows:
[0039] (1) Obtain the whole genome sequence information of Streptomyces coelicolor M145 through the NCBI database.
[0040] (2) Culture Streptomyces coelicolor M145 in R5-liquid medium to the middle and late exponential growth, prepare the bacterial genome, and obtain the whole genome interaction frequency information of the two periods by high-throughput chromatin conformation capture technology.
[0041] (3) Culture Streptomyces coelicolor M145 in R5-liquid medium to the middle and late exponential growth, prepare mRNA samples, and obtain whole genome transcription data by transcriptome sequencing.
[0042] (4) Calculate the local chromatin interaction intensity of each minimum resolution chromatin region (5kb resolution is used in this embodiment) and the upstream and downstream 15-200kb range, i.e. FIRE value.
[0043] 1) Calculate the effective fragment length, GC content and mappability of the genome, i.e. F_GC_MAP, using the F_GC_MAP.file.sh code in HiCnv software;
[0044] 2) Calculate the sum of local interaction frequency in each minimum resolution chromatin region and the 15-200kb interval upstream and downstream;
[0045] 3) Based on the standardization algorithm of Poisson regression, use the sum of local interaction frequency in each minimum resolution chromatin region and the 15-200kb interval upstream and downstream obtained in 2) and the F_GC_MAP obtained in 1) to perform HiCNormCis calculation to obtain the interaction intensity FIRE value.
[0046] (5) Set each 11 minimum resolution chromatin region as a unit, and after Z transformation and Gaussian smoothing of the FIRE value of the 11 minimum resolution chromatin regions contained in each unit and the corresponding transcription signal, calculate the Pearson correlation between the two, and the unit with Pearson correlation ≥ 0.8 is defined as a strong correlation chromatin region.
[0047] Table 1 Information of strong correlation chromatin regions
[0048]
[0049]
[0050] (6) Integrate the RK-682 biosynthetic gene cluster into each of the 10 chromatin regions with different strong correlations in the middle and late exponential growth of Streptomyces coelicolor M145 (Table 1), respectively, and use the universal heterologous expression chromosomal integration site attB as a control.
[0051] Each strain integrated with the RK-682 biosynthetic gene cluster was subjected to fermentation culture, and the fermentation product was extracted, and high performance liquid chromatography-mass spectrometry was used to detect the RK-682 yield. The results showed that in the middle and late exponential growth, the RK-682 yield integrated in the high FIRE value strong correlation chromatin region was significantly improved by up to 4 times compared with the attB control site ( Figure 1 , the horizontal line represents the yield of the attB site integration control).
[0052] The above examples are the preferred embodiments of the present application, but the embodiments of the present application are not limited by the above examples, and any changes, modifications, substitutions, combinations, simplifications made without departing from the spirit and principles of the present application shall be equivalent replacement methods, and all shall be included in the protection scope of the present application.
Claims
1. A method for producing high-yield biosynthetic genes or gene cluster products based on the three-dimensional structure of chromatin, characterized in that, Includes the following steps: (1) Obtain whole genome sequence information through databases or deep sequencing; (2) Obtain genome-wide chromatin interaction frequency information through high-throughput chromatin conformation capture technology; (3) Obtain whole-genome transcriptional information through transcriptome sequencing; (4) Based on the chromatin interaction frequency information obtained in step (2), calculate the interaction intensity between the minimum resolution chromatin region and the upstream and downstream chromatin regions; The calculation method for interaction strength includes the following steps: 1) Use F_GC_MAP.file.sh in the HiCnv software to calculate the effective fragment length, GC content, and mappability of the genome, i.e., F_GC_MAP; 2) Calculate the sum of the frequencies of local chromatin interactions within each minimum resolution chromatin region and its upstream and downstream intervals; 3) Based on the Poisson regression standardization algorithm, the sum of the local chromatin interaction frequencies obtained in 2) and the F_GC_MAP obtained in 1) are used to calculate HiCNormCis, so as to obtain the interaction intensity between each minimum resolution chromatin region and the upstream and downstream chromatin regions. (5) Based on the information obtained in steps (3) and (4), calculate the Pearson correlation between the interaction strength of several consecutive minimum resolution chromatin regions and the corresponding transcriptional levels. (6) Integrate the target gene or gene cluster into a chromatin region on the genome where the interaction strength is strongly correlated with the corresponding transcription level and the interaction strength is high, thereby increasing the yield of the target gene or gene cluster product.
2. The method for synthesizing high-yield genes or gene cluster products based on the three-dimensional structure of chromatin according to claim 1, characterized in that: The term "high interaction strength" refers to interactions ranked in the top 30%; the term "strong correlation" refers to a Pearson correlation ≥ 0.
8.
3. The method for synthesizing genes or gene cluster products based on the three-dimensional structure of chromatin according to claim 2, characterized in that: The minimum resolution is 5–10 kb, and the upstream and downstream chromatin regions are 15–200 kb regions upstream and downstream; the several consecutive minimum resolution chromatin regions are 7–13 consecutive minimum resolution chromatin regions.
4. The method for synthesizing genes or gene clusters based on the three-dimensional structure of chromatin according to claim 1, characterized in that: The calculation method for Pearson correlation in step (5) includes the following steps: several consecutive minimum resolution chromatin regions are set as a unit, and the Pearson correlation between each unit is calculated by performing Z-transformation and Gaussian smoothing on the local chromatin interaction intensity and the corresponding transcription signal.
5. The method for synthesizing high-yield genes or gene cluster products based on the three-dimensional structure of chromatin according to claim 1, characterized in that: Step (1) is to obtain the whole genome sequence information of the target strain through a database or through deep sequencing. Step (2) is as follows: culture the target strain cells to a specific growth stage, prepare bacterial genome samples, and obtain whole genome interaction frequency information through high-throughput chromatin conformation capture technology; Step (3) involves culturing the target strain cells to the same growth stage, preparing bacterial mRNA samples, and obtaining whole genome transcription data through transcriptome sequencing.
6. The application of the method according to any one of claims 1-5 in the high-yield modification of microbial metabolites.
7. The application of the method according to any one of claims 1-5 in the silencing and activation modification of microbial metabolites.