High-throughput cis-acting element screening system and screening method

By designing a reporter vector system containing a promoter and a reporter gene, and utilizing barcode labeling and high-throughput sequencing, the problem of low efficiency in screening DNA cis-acting elements in existing technologies has been solved, enabling efficient and rapid screening of a variety of regulatory elements applicable to multiple species.

CN121362779AActive Publication Date: 2026-01-20SHANGHAI JIAOTONG UNIV

Patent Information

Application Number
CN202410976595.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-20
Publication Date
2026-01-20
Estimated Expiration
2044-07-20

AI Technical Summary

Technical Problem

In existing technologies, methods for screening DNA cis-acting elements are inefficient, making it difficult to obtain regulatory elements in a high-throughput, efficient, and rapid manner. Furthermore, the screened elements are relatively long, posing biosafety concerns.

Method used

A reporter vector system was designed, comprising a promoter and a reporter gene, into which cis-acting elements to be screened were inserted, and each element was labeled with a barcode sequence. The abundance of the barcodes was analyzed using high-throughput sequencing to calculate the regulatory activity, and a reporter vector library was constructed for screening.

Benefits of technology

This technology enables high-throughput, efficient, and rapid screening of various cis-acting elements with different regulatory levels, applicable to multiple species, meeting the needs of biological research and breeding, and reducing the impact of element length on biosafety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004954743000000051
    Figure BDA0004954743000000051
  • Figure BDA0004954743000000101
    Figure BDA0004954743000000101
  • Figure BDA0004954743000000141
    Figure BDA0004954743000000141
Patent Text Reader

Abstract

The invention relates to a high-throughput cis-acting element screening carrier and a screening method. Specifically, the invention provides a plasmid vector system containing bar codes, each bar code in the system is in one-to-one correspondence with a candidate cis-acting element, and the activation multiple of the candidate cis-acting element can be obtained by measuring the abundance of the bar codes; in order to eliminate the influence of a bar code on the vector on a detection result, an exogenous intron which can be cut off during transcription is inserted into a coding region in the vector and is used for distinguishing vector DNA and RNA obtained by transcription. The screening method of the biological cis-acting element has high efficiency, wide applicability and high throughput, and has outstanding application value in biological research and breeding.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of biotechnology, in particular to a high-throughput cis-acting element screening system and a screening method. BACKGROUND

[0002] For synthetic biology, it is of great application value to obtain DNA cis-acting elements capable of regulating the expression level of genes. At present, there are still relatively few reports on cis-acting elements that can be directly used to regulate the expression level of genes, especially elements inserted in situ into the promoter region to control the expression level of genes.

[0003] In the prior art, some useful candidate cis-acting elements have been obtained by the method of STARR-seq in previous studies, but overall, their ability to regulate the expression level of genes is weak, the range of regulation is narrow, and the length of the elements obtained by screening is relatively long, generally more than 100 bp, which may cause certain biological safety problems.

[0004] In the field of biology, a large number of candidate cis-acting elements can be mined by high-throughput ChIP-Seq (Chromatin immunoprecipitation followed by sequencing), Assay for transposase accessible chromatin with high-throughput sequencing and Hi-C technology, but after obtaining these candidate regions, it is necessary to further evaluate whether the candidate regions have regulatory ability, and overall, the efficiency is low and it is difficult to perform high-throughput.

[0005] Therefore, there is an urgent need in the art to develop a universal high-throughput, efficient and rapid cis-acting element screening method. SUMMARY

[0006] The purpose of the present application is to provide a universal high-throughput, efficient and rapid cis-acting element screening system and a screening method.

[0007] The present application provides a reporter vector for screening cis-acting elements, wherein the reporter vector comprises a promoter and a reporter gene; a cis-acting element to be screened is inserted into the promoter region; and the coding region of the reporter gene comprises (i) a barcode sequence corresponding to the cis-acting element to be screened; and (ii) at least one exogenous intron sequence.

[0008] In another preferred embodiment, the insertion site of the cis-acting element to be screened is located in the region from -1 to -100 bp upstream of the transcription start site of the reporter vector.

[0009] In another preferred embodiment, the barcode sequence consists of 3m bases N, N is selected from A, T, G or C; wherein 2≤m≤5; preferably, m=3.

[0010] In another preferred embodiment, the barcode sequence is a nucleotide sequence of 6-12 bases N, wherein N is selected from A, T, G or C.

[0011] In another preferred embodiment, the barcode sequence is a nucleotide sequence of 9 bases in length NNNNNNNNN, wherein N is selected from A, T, G or C.

[0012] In another preferred embodiment, each of the barcode sequences corresponds to one of the candidate cis-acting element sequences.

[0013] In another preferred embodiment, the barcode sequence is located at the 3' end of the start codon of the reporter gene.

[0014] In another preferred embodiment, the exogenous intron is removed from the transcript produced by the reporter vector.

[0015] In another preferred embodiment, a primer spanning the exogenous intron is used when reverse transcribing the transcript produced by the reporter vector.

[0016] In another preferred embodiment, the reporter vector is a transient expression vector.

[0017] In another preferred embodiment, the reporter vector is a plasmid.

[0018] In another preferred embodiment, the reporter gene is a luciferase gene.

[0019] In another preferred embodiment, the promoter is selected from the group consisting of 35S mini pro, EUI pro, RFT1 pro.

[0020] In another preferred embodiment, the reporter vector further comprises: a UTR sequence located after the promoter.

[0021] In another preferred embodiment, the UTR is derived from the 5' UTR of histone H3.2 of the species corresponding to the cis-acting element.

[0022] In a second aspect of the present application, a library of reporter vectors for high-throughput screening of cis-acting elements is provided, the library of reporter vectors containing X reporter vectors of the first aspect of the present application carrying different cis-acting elements to be screened, wherein 50≤X≤10 7 .

[0023] In another preferred embodiment, 100≤X≤10 6 ; more preferably, 1000≤X≤10 5 .

[0024] In another preferred embodiment, each of the barcode sequences in the library of reporter vectors corresponds to the candidate cis-acting element sequence located in the same vector; one cis-acting element can correspond to multiple barcode sequences.

[0025] In a third aspect of the present application, a method for constructing the library of reporter vectors of the second aspect of the present application is provided, characterized in that it comprises the steps of:

[0026] (i) providing oligonucleotides each comprising a different candidate cis-acting element, a restriction site and a corresponding barcode sequence;

[0027] (ii) inserting the oligonucleotides into a plasmid backbone; the plasmid backbone comprises a reporter gene, and the oligonucleotides are inserted upstream of the reporter gene, forming a reporter vector precursor;

[0028] (iii) performing restriction digestion on the oligonucleotides in the reporter vector precursor, and connecting a promoter to the restriction site, forming the reporter vector.

[0029] In another preferred embodiment, the oligonucleotides have the structure shown in Formula I:

[0030] Z0-STE-Z1-Q-N-L (Formula I)

[0031] In the formula,

[0032] each of “-” is independently a covalent bond or a linker;

[0033] Z0is an amplification primer sequence;

[0034] STE is a candidate cis-acting element sequence;

[0035] Z1is a restriction sequence;

[0036] Q is a start codon;

[0037] N is a barcode sequence;

[0038] L is nothing or a linker.

[0039] In another preferred embodiment, the restriction sequence comprises two Bsal restriction sites.

[0040] In another preferred embodiment, L is a linker, and the linker is a coding sequence of GGGGS short peptide.

[0041] In another preferred embodiment, the candidate cis-acting element sequence is derived from transcriptome data of the species to be screened.

[0042] In another preferred embodiment, the transcriptome data is set as a threshold higher than the median, all candidate genes are set as a standard covering 1 kb upstream of the promoter ± 0.2 kb, and the sequences are truncated at 100 bp ± 10 bp, with an overlap of 30 bp ± 10 bp between the two sequences, to obtain a candidate cis-acting element library.

[0043] In another preferred embodiment, the species to be screened is selected from the group consisting of plants, animals, microorganisms, or a combination thereof.

[0044] In another preferred embodiment, the oligonucleotides are synthesized in high throughput on a chip to form an oligonucleotide library.

[0045] In a fourth aspect of the present application, a method for high-throughput screening of cis-acting elements is provided, comprising the steps of:

[0046] (1) obtaining candidate cis-acting element sequences and reference sequences;

[0047] (2) synthesizing a reporter vector library as described in the second aspect of the present application;

[0048] (3) introducing the reporter vector library into host cells for transcription, extracting total RNA; using an intron-spanning primer for specific RNA reverse transcription to obtain a cDNA library; amplifying sequences containing barcodes from the cDNA library;

[0049] (4) using a primer containing a partial intron sequence to amplify the sequences of the initial input barcodes in the reporter vector library obtained from the host cells;

[0050] (5) using high-throughput sequencing (NGS) to sequence analyze the sequences amplified in steps (3) and (4), respectively, to read the abundance of each barcode; the proportion of the reads of a specific barcode in the cDNA library to the total reads is recorded as R1, and the proportion of the reads of the same barcode in the reporter vector library to the total reads is recorded as R2, and the ratio of R1 to R2 represents the regulatory activity of the cis-acting element corresponding to the barcode.

[0051] In another preferred embodiment, the sequence of the intron-spanning primer is as follows:

[0052] CTCGGCCTTATGCAGTTGCTCTCCAG (SEQ ID NO: 7); and / or

[0053] GCTCGCCTTATGCAGTTGCTCTCCAG (SEQ ID NO: 8).

[0054] In another preferred embodiment, the sequences containing barcodes are amplified from the cDNA library using PCR.

[0055] In another preferred embodiment, the method further comprises:

[0056] (6) Constructing a plurality of reporter vector libraries with different strength promoters in step (2), repeating steps (3)-(5) for each of the plurality of reporter vector libraries independently.

[0057] In another preferred embodiment, the promoter is selected from the group consisting of 35S mini pro, EUI pro, RFT1 pro.

[0058] In another preferred embodiment, the reporter vector in the reporter vector library contains n reference sequences; preferably, 10≤n≤300; more preferably, 100≤n≤200.

[0059] In another preferred embodiment, in step (3), the host cell is selected from the group consisting of a plant cell, an animal cell, or a combination thereof.

[0060] In another preferred embodiment, the host cell is a protoplast.

[0061] In another preferred embodiment, the host cell is derived from rice, soybean, tomato, corn, tobacco, wheat, sorghum, salvia miltiorrhiza, alfalfa, human, mouse, or a combination thereof.

[0062] In another preferred embodiment, in step (4), the reference sequence is assumed to have no enhancing effect on transcription, and the reference sequence is normalized.

[0063] In another preferred embodiment, the regulatory activity of the cis-acting element corresponding to the barcode is calculated using a formula as shown in Formula II:

[0064]

[0065] In the formula,

[0066] Ratio STE-cDNA represents the proportion of the specific barcode read in the cDNA library in the total read;

[0067] Ration STE-plasmid represents the proportion of the same barcode read in the reporter vector library in the total read;

[0068] Ration Ctrl-cDNA represents the proportion of the barcode read corresponding to the reference sequence in the cDNA library in the total read;

[0069] Ration Ctrl-plasmid represents the proportion of the barcode read corresponding to the reference sequence in the reporter vector library in the total read;

[0070] n represents the number of reference sequences;

[0071] Activity STE indicates the fold of the transcription level enhanced by the cis-acting element.

[0072] In a fifth aspect of the present application, a host cell is provided, wherein the host cell comprises the reporter vector of the first aspect of the present application or the reporter vector library of the second aspect of the present application.

[0073] In another preferred embodiment, the host cell comprises a plant cell, an animal cell, or a combination thereof.

[0074] In another preferred embodiment, the host cell is a protoplast.

[0075] In another preferred embodiment, the reporter vector is introduced into the host cell by transient transformation.

[0076] It should be understood that, within the scope of the present application, each of the technical features of the present application described above and each of the technical features specifically described hereinafter (e.g., in the examples) can be combined with each other to form a new or preferred technical solution. Due to the limited space, they will not be listed one by one here. BRIEF DESCRIPTION OF DRAWINGS

[0077] Figure 1 STEM01, a reporter system established in rice using luciferase as a reporter gene, is shown, in which nLUC driven by the UBI promoter serves as an internal reference, and firefly luciferase (fLuc) serves as a dual LUC reporter gene.

[0078] Figure 2 The effects of five cis-acting elements from STARR-seq and four from viruses on three different reporter promoters in the STEM01 reporter system are shown. The data of the analysis results are normalized using nLUC as an internal reference and a vector containing only a promoter as a control. All data were analyzed by one-way ANOVA (N = 3): *** indicates P < 0.001.

[0079] Figure 3 The extremely strong activation effect of vs003, which was screened from viruses, is shown by different truncations to obtain the activation effect of different truncated forms. The data of the analysis results are normalized using nLUC as an internal reference and a vector containing only a promoter as a control. All data were analyzed by one-way ANOVA (N = 3): *** indicates P < 0.001.

[0080] Figure 4A reporter system STEM02 for high-throughput screening of candidate cis-acting elements is shown, in which nLUC driven by UBI promoter as internal control, and firefly luciferase (fLuc) as dual LUC reporter gene. At the 14th amino acid position of fLuc, a 97 bp intron from OsAdhl (LOC_Osl lg10480) is inserted by Gibson assembly, which makes the mRNA different from all DNA sequences of candidate. At the N-terminus of fLuc, a 9 bp barcode DNA sequence encoding a tripeptide is fused to distinguish the transcripts produced by different candidate cis-acting elements.

[0081] Figure 5 A flow chart of genome-wide mining of candidate cis-acting elements is shown. (A) First, all candidate cis-acting elements are screened according to transcriptome data, and the 170-bp fragments are designed according to the rules of candidate cis-acting elements, and the relevant candidate fragments are synthesized by chip synthesis method (Genscript, China) (B) The corresponding cis-acting element library is constructed; (C) plant cells or animal cells are transfected; (D) the strength of different candidate cis-acting elements is further quantified by detecting the content of fLuc transcript containing barcode.

[0082] Figure 6 The results of rice cis-acting elements obtained by genome-wide mining are shown. (A, B) Analysis of biological replicates of two different experiments of the library of weakly transcribed promoter 35Smin and moderately transcribed promoter EUI, 35Smin (A) or EUI (B), randomly selected non-cis-acting element sequences as control group (Null). (C) All cis-acting elements with different activation multiples screened in two candidate libraries in genome-wide screening experiments. Element (-) indicates internal control sequence. All data analysis used Student's t-test (n=2), *** indicates P<0.001. DETAILED DESCRIPTION

[0083] Through extensive and in-depth research, the inventors provide a high-throughput cis-acting element screening vector and screening method. Specifically, the cis-acting element is inserted into the promoter region to regulate the gene expression level. In order to screen cis-acting elements with different regulatory abilities from the genomes of species on a large scale, the inventors provide a plasmid vector system STEM02 containing barcodes, in which each barcode corresponds to a candidate cis-acting element one by one, for distinguishing the transcripts produced by different candidate cis-acting elements, and obtaining the activation fold of the candidate cis-acting element by measuring the abundance of the barcode; in order to exclude the influence of the barcode on the vector on the detection results, an exogenous intron that will be removed during transcription is inserted into the coding region in the vector, for distinguishing the vector DNA and the RNA obtained by transcription. On this basis, the present application is completed.

[0084] Cis-regulatory element (STE)

[0085] As used herein, the terms "cis-regulatory element of the present application", "transcriptional regulatory element of the present application" and "STE element of the present application" are used interchangeably, and all refer to a sequence element that can be inserted in situ into the promoter region to regulate the transcription level. When the STE element of the present application is inserted into the region in front of the transcription start site (TSS), the expression of the target gene operatively linked to the promoter can be significantly increased or up-regulated. The fold of enhancement of the STE element on the transcription level or the gene expression level is also referred to as the activation fold.

[0086] Typically, the lower limit (Lmin) of the length of the STE element of the present application is at least 30 bp, such as ≥ 40 bp, ≥ 50 bp, ≥ 60 bp, ≥ 80 bp, ≥ 100 bp. In addition, the upper limit (Lmax) of the fragment length of the short enhancer of the present application is at most 200 bp, such as ≤ 180 bp, ≤ 160 bp, ≤ 150 bp, ≤ 130 bp, ≤ 120 bp, ≤ 100 bp. Preferably, the length of the STE of the present application can be reasonably determined by any lower limit (Lmin) and upper limit (Lmax) described above, for example, 30-200 bp, preferably 40-150 bp, more preferably 60-130 bp.

[0087] In another preferred embodiment, the STE of the present application comprises a plant-derived STE, a virus-derived STE, or a combination thereof.

[0088] In the present application, the length of the cis-acting element is less than 100 bp, which is an endogenous cis-acting element of a species. The species origin of the cis-acting element is not limited, including but not limited to viruses or rice, etc.

[0089] Preferably, the candidate cis-acting elements are obtained directly by analyzing transcriptome data of different species without focusing on aspects such as histone modification, transcription factor binding sites and chromatin accessibility of DNA.

[0090] In one embodiment, the candidate cis-acting elements are obtained by the following steps: after obtaining the transcriptome data of the species to be screened, all candidate genes are truncated at 100 bp as a segment with 1 kb upstream of the promoter as a standard, with an overlap of 30 bp between the two sequences before and after, to obtain a candidate cis-acting element library.

[0091] The promoters, constructs or expression cassettes containing the STE elements of the present application can be used to express foreign genes in plants, thereby improving the agronomic traits of plants (such as the quality of edible parts, leaf type, leaf color, flower type, flower color, etc.), or enhancing the resistance of plants to adverse environments (such as freezing, high temperature, drought, viruses, bacteria, insect pests or artificial herbicides). The STE elements of the present application can also be used to produce specific protein products (such as antigen proteins and biological agents) in plants, or to study the specific functions of foreign genes in plant parts (such as leaves and flower organs) (such as the effects on mesophyll cell development, or the effects on flower organ formation, etc.).

[0092] STEM01 screening system

[0093] In order to quickly screen STE elements, a double reporter gene system STEM01 is constructed, which contains a first reporter gene and a second reporter gene in one plasmid, driven by a first promoter and a second promoter, respectively. The candidate STE element is inserted into the first promoter, so that the transcription level of the first reporter gene is regulated by the candidate STE, while the transcription level of the second reporter gene is not regulated by the candidate STE. The plasmid carrying the candidate STE is named pSTEM01.

[0094] The ratio of the transcription level T1 of the first reporter gene to the transcription level T2 of the second reporter gene represents the fold of the candidate STE; when T1 / T2>1, the candidate STE has the ability to enhance the transcription level; when T1 / T2<1, the candidate STE has the ability to weaken the transcription level. When T1 / T2>5, the candidate STE has a significant enhancement ability; preferably, T1 / T2>10; more preferably, T1 / T2>20.

[0095] In a preferred embodiment, the first and second reporter genes encode proteins selected from the group consisting of firefly luciferase (fLuc), nanoluciferase (nLuc), renilla luciferase (rLuc), green fluorescent protein (GFP), and the first and second reporter genes encode different proteins. The first promoter is selected from the group consisting of 35S mini pro (SEQ ID NO: 1), EUI pro (SEQ ID NO: 2), RFT1 pro (SEQ ID NO: 3); and the second promoter is a UBI promoter.

[0096] For example, taking nanoluciferase (nLuc) as the second reporter gene, firefly luciferase (fLuc) driven by a promoter containing STE as the first reporter gene, the ratio of fLuc / nLuc represents the regulation strength of STE on the transcription level, and the schematic diagram of the system is as shown in Figure 1

[0097] In order to ensure the stability of the screening element in the species, a STEM01 system containing three promoters with different strengths is constructed for verification, including a general weak promoter 35S mini and two species-specific medium-strength promoters EUI pro and high-strength promoter RFT1 pro.

[0098] STEM02 screening system

[0099] In order to realize high-throughput screening of multiple STEs at the same time, on the basis of the STEM01 system, the application further develops an optimized STE mining method, called STEM02 screening system, and the schematic diagram is as shown in Figure 4

[0100] A reporter vector pSTEM02 for carrying candidate STEs is constructed, which contains a promoter and a reporter gene. The STE to be screened is inserted into the promoter or upstream of the promoter, and a barcode sequence corresponding to the STE is inserted into the coding region of the reporter gene. Each barcode corresponds to a unique STE. Since the candidate STE is placed before the transcription start site, it cannot be transcribed, but the element can change the abundance of mRNA, so the barcode is added to the coding region, and the activation ability of different candidate elements and the abundance of mRNA (i.e. the abundance of the barcode) are one-to-one corresponding, so as to obtain the activation fold of the corresponding STE.

[0101] In order to reduce the problem of incorrect screening results caused by the deletion of the barcode, each candidate STE can correspond to multiple barcode sequences. The average value of the abundance of the multiple barcode sequences corresponding to the same STE is used to calculate the activation fold of the STE.

[0102] ​​In the present application, the amino acid sequence encoded by the barcode sequence is located at the N-terminal of the protein encoded by the reporter gene. The length of the barcode sequence can be 6bp, 9bp, 12bp, or 15bp, etc., and the preferred length of the barcode sequence is 9bp. Preferably, the barcode nucleotide sequence is NNNNNNNNN, wherein N is selected from A, T, G or C.

[0103] In order to exclude the interference of the barcodes possessed by the original pSTEM02 vector library on the process of determining the barcode abundance, at least one exogenous intron element is inserted into the coding region of the reporter gene of the pSTEM02 vector. The intron does not affect the transcription level and is stably removed during the transcription, i.e., the transcript obtained by the expression of the reporter gene does not contain the intron.

[0104] Method for preparing the pSTEM02 vector library

[0105] In the present application, the pSTEM02 vector library is constructed by a two-step method.

[0106] In the first step, oligonucleotides containing different candidate STE sequences, enzyme cutting sites and corresponding barcode sequences, respectively, are provided, and an oligonucleotide library is synthesized on a chip in a high-throughput manner; and the oligonucleotide library is inserted into a plasmid;

[0107] Specifically, the structure of the oligonucleotide is shown in Formula I:

[0108] Z0-STE-Z1-Q-N-L (Formula I)

[0109] In the formula,

[0110] Each of “-” is independently a covalent bond or a linker;

[0111] Z0 is an amplification primer sequence;

[0112] STE is a candidate cis-acting element sequence;

[0113] Z1 is an enzyme cutting sequence; preferably, it contains two BsaI enzyme cutting sites;

[0114] Q is a start codon;

[0115] N is a barcode sequence;

[0116] L is nothing or a linker; preferably, it is the coding sequence of a GGGGS short peptide.

[0117] In a preferred embodiment, the oligonucleotide comprises 18 bp of the amplification primer sequence with a repeat rate lower than 30% in the corresponding species, followed by 100 bp of the candidate cis-acting element, 22 bp of the restriction site containing two Bsal, the translation initiation codon ATG, 9 bp of the barcode and 18 bp of the DNA sequence corresponding to the GGGGS short peptide.

[0118] Illustratively, the nucleotide sequence of Z0 is shown in SEQ ID NO: 4: GACCGCCTCCACGACAAC (SEQ ID NO: 4); the nucleotide sequence of Z1 is shown in SEQ ID NO: 5: AGTAGGAGACCGGTCTCACAGT (SEQ ID NO: 5); and the nucleotide sequence of L is shown in SEQ ID NO: 6: GGTGGAGGCGGTGGTAGT (SEQ ID NO: 6).

[0119] The plurality of oligonucleotides with different STEs constitute an oligonucleotide library. Preferably, the oligonucleotide library is synthesized on a chip in high throughput.

[0120] In the second step, the oligonucleotides in the reporter system precursor are subjected to enzyme digestion, and the promoter is connected to the enzyme digestion site to form the pSTEM02 reporter vector library.

[0121] Screening STEs with the STEM02 screening system

[0122] After transiently transforming cells with the pSTEM02 vector library, the cells are cultured, enriched, and total RNA is harvested. The total RNA is subjected to reverse transcription with primers spanning the intron to prepare cDNA. Sequences containing barcodes are amplified from the cDNA library and the vector library, respectively, and the amplified sequences are subjected to high-throughput sequencing (NGS). The proportion of the reads of a specific barcode in the total reads is calculated, and the proportion of the reads of a specific barcode in the total reads in the cDNA library is denoted as R1, and the proportion of the reads of the same barcode in the total reads in the vector library is denoted as R2. The ratio of R1 to R2 represents the activation fold of the STE corresponding to the barcode. The schematic diagram of the process is shown in FIG. 2. Figure 5 .

[0123] To increase the accuracy of the measurement of the activation fold of STE, the results of the reference sequences in the vector library are normalized. For example, 150 randomly selected reference sequences are assumed to have no enhancement effect on transcription, and the calculated value is 1. Accordingly, the fold change mediated by each STE is calculated using the formula shown in Formula II below

[0124]

[0125] In the formula, Ratio STE-cDNArepresents the proportion of the number of reads of a specific barcode in the cDNA library in the total number of reads; Ration STE-plasmid represents the proportion of the number of reads of the same barcode in the reporter vector library in the total number of reads; Ration Ctrl-cDNA represents the proportion of the number of reads of a specific barcode in the cDNA library in the total number of reads; Ration Ctrl-plasmid represents the proportion of the number of reads of a specific barcode in the cDNA library in the total number of reads; Ration STE represents the fold of the transcription level enhanced by STE, i.e. the activation fold of STE.

[0126] Preferably, a plurality of STEM02 reporter vector libraries are constructed with a plurality of promoters having different strengths. For example, 35S mini pro promoter with weak promoter activity, EUI pro promoter with medium promoter activity.

[0127] The main advantages of the present application include:

[0128] (1) The screening method of the present application can be used for high-throughput, large-scale and efficient screening and identification of cis-acting elements.

[0129] (2) The screening method of the present application can obtain a variety of elements with different transcription regulation levels to meet different needs in biological research and breeding.

[0130] (3) The screening of the present application can be used for screening cis-acting elements from a variety of species, which has wide applicability.

[0131] The present application will be further described in conjunction with specific examples. It should be understood that these examples are only used to illustrate the present application and not used to limit the scope of the present application. The experimental methods in the following examples without specific conditions are generally carried out under conventional conditions, for example, the conditions described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or the conditions recommended by the manufacturer. Unless otherwise specified, percentages and fractions are weight percentages and weight fractions.

[0132] Materials and Methods

[0133] 1. Plant Cultivation

[0134] The japonica rice variety used in this study was Xiushui 134. The rice plants were cultivated in growing rooms, maintained at a 12-hour light cycle of 30°C and a 12-hour dark cycle of 28°C, or planted in paddy fields in Shanghai (N31.13°, E121.28°) from May to October, and in Sanya (N18.25°, E109.50°) from January to April each year.

[0135] 2. Plasmid construction

[0136] The CRISPR / Cas9 plasmid was constructed according to previously reported descriptions (Lu et al., 2017). Optimal spCas9 targets were designed using CRISPR-P and CRISPR-GE (Liu et al., 2017; Xie et al., 2017). The target plasmid pCBSG032 and the oligonucleotide aptamer were ligated via a Golden Gate. A 10 μL assembly reaction containing 50 ng of pCBSG032 vector, 0.05 pmol of oligonucleotide aptamer, T4 ligase, and BsaI was set up and performed under the following thermal cycling conditions: incubation at 37°C for 5 min, followed by incubation at 16°C for 5 min, repeated 30 times. The oligonucleotides were purchased from Sangon Biotech, and the assembled DNA fragment was directly transformed into E. coli to generate the CRISPR / Cas9 plasmid.

[0137] The pSTEM01 plasmid was derived from pCambia1300-LUC. It was pre-cloned in fLuc via Gibson assembly and different promoters were added. Candidate STEs were then inserted into these reporter genes.

[0138] For pSTEM02, a 97 bp intron of OsAdh1 (LOC_Os11g10480) was inserted into the 14th amino acid position of fLuc in the pSTEM01 vector via Gibson assembly. Then, a sequence containing ccdb was inserted to replace the promoter (Lu et al., 2017).

[0139] 3. Protoplast preparation and transformation

[0140] Sterilized rice seeds were germinated on 1 / 2x MS medium and cultured in the dark at 28°C for 8–12 days. Protoplasts were isolated and transformed according to a previously reported protocol (Lin et al., 2018). Approximately 10 μg of plasmid DNA was introduced into approximately 5 × 10⁻⁶ cells per transfection. 5 Transfected protoplasts were incubated in the dark at 25°C for 12–16 hours. After incubation, protoplasts were collected by centrifugation at 100g for 3 minutes for further analysis.

[0141] Example 1: Rapid screening of viral STEs with STEM01 system

[0142] To verify the reliability of the STEM01 system in mining STEs, the sequences screened by STARR-seq were first screened and the fold activation was calculated.

[0143] The pSTEM01 vector was constructed, the schematic diagram of which is shown in Figure 1 Nanoluciferase (nLuc) was used as the second reporter gene, and firefly luciferase (fLuc) driven by the promoter containing the sequence to be screened was used as the first reporter gene to evaluate the strength of STE. Candidate sequences were tested on three promoters, including 35S minipro and rice-derived EUIpro and RFT1pro promoters, to ensure the reliability of the evaluation of STE activation strength.

[0144] The dual luciferase reporter gene detection method is as follows: Nano-Glo dual luciferase reporter gene detection system (Promega, USA) and microplate luminometer were used for luminescence detection, and the method was performed according to the reported method (Shen et al., 2023). After incubation and collection of protoplasts, 100 μL of 1x lysis buffer was added. After 15 minutes of incubation, an equal volume of ONE-Glo reagent was added. After measuring the activity of firefly luciferase, an equal volume of NanoDLR Stop&Glo reagent was added to measure the activity of nLuc.

[0145] The results show that the insertion of these candidate sequences screened by STARR-seq into these three promoters leads to a limited increase in gene expression levels. In addition, the length of these sequences (650-811 bp) is relatively long, which poses a challenge to precise integration.

[0146] To find STEs with higher activity and shorter length, the STEM01 system was used to screen four candidate STE sequences derived from plant virus promoters. Among the four elements tested (STEvs001 to STEvs004), STEvs003 showed the highest up-regulation effect, increasing expression by up to 2283-fold Figure 2 ).

[0147] Further screening of truncations of STEvs003 was performed, and the results are shown in Figure 3 Several short STEs with significant activity were obtained, including 40bp STEvs005 which increased expression by 9.8-fold. Among them, a 123bp truncated fragment (STEvs011) showed comparable ultra-high activity to the original STEvs003, increasing expression by up to about 2000-fold Figure 3 ).

[0148] Example 2: Large scale screening of rice STEs using the STEM02 system

[0149] To enable high-throughput, large-scale and rapid screening of candidate STEs of rice origin, the STEM02 system was used for screening, following the steps below:

[0150] 1. Obtaining of candidate STE sequences

[0151] A total of 284 RNA-seq datasets were retrieved from the IC4R database (http: / / expression.ic4r.org / ) and high expressed genes were identified by analyzing the average and minimum FPKM values in the samples and tissues. After obtaining the promoter regions of these genes from the rice genome (MSU 7.0), they were divided into 10 parts, each containing 100 bp and having a 30 bp overlap between the sequences, starting from the TSS, thus covering a 730 bp promoter sequence. Fifteen genes with lower FPKM values were also randomly selected and truncated to 150 sequences as negative controls. A total of 11,610 candidate STEs were obtained.

[0152] 2. Construction of the STE library

[0153] For each candidate STE sequence, three barcodes were used for labeling, respectively. A total of 34,830 oligonucleotides were generated, which contained two Bsal cleavage sites between the STE and barcode sequences (STE-Bsal-Barcode) for insertion of the promoter when synthesized. All oligonucleotides were synthesized on a chip (Genscript, China) for two-step library construction.

[0154] In the first step, the pSTEM02 precursor amplified by PCR and the oligonucleotide library were used as the vector and the insert, respectively. The vector (5 pg) and the insert (300 ng) were ligated by Gibson assembly. The ligation product was then electroporated into E. coli to obtain about 1 million colonies (~30x coverage). Plasmids were directly extracted from these colonies to obtain the library prepared in the first step.

[0155] In the second step, 10 pg of plasmid from the first step was prepared as a vector by Bsal digestion. The sequence containing the H3.2 (LOC4338913) UTR and the promoter was digested by Bsal and connected to the plasmid vector obtained in the first step by Golden Gate. Similarly, about 1 million colonies (~30x coverage) were generated, and plasmids were directly extracted from these colonies to generate the final pSTEM02 library.

[0156] Two pSTEM02 libraries were constructed using 35S minipro and EUIpro promoters, respectively.

[0157] 3. Transformation of protoplasts and sequencing

[0158] After transformation of rice protoplasts, RNA-seq analysis was performed to evaluate the regulatory activity of all candidate STEs. The detailed method is as follows:

[0159] About 100 pg of plasmid from one STE library (i.e. STE library based on 35S minipro promoter or STE library based on EUIpro promoter) was used to transform 5 x 10 6 protoplasts. The protoplasts were harvested after incubation at 25 °C in the dark for 12-16 hours and immediately resuspended in 1 mL TRIzol. Total RNA was extracted according to the manufacturer's instructions (Vazyme, China). Subsequently, reverse transcription was performed using primers spanning the pSTEM02 intron (protoplast_RT_HR1) to prepare cDNA (TransGen, China). Sequences containing barcodes were amplified from cDNA libraries or corresponding plasmid libraries. Then, next-generation sequencing (NGS, Illumina HiSeq, PE150) was used to sequence these amplicons for subsequent analysis. Two independent biological replicates were performed for each experiment.

[0160] 4. Analysis of the regulatory activity of candidate STEs using barcodes

[0161] NGS reads were classified according to different barcodes corresponding to STEs. The sequencing reads corresponding to the barcodes of each STE were merged and counted. For each sample, more than 99% of the reads were successfully merged. 150 randomly selected sequences were assumed to have no enhancement effect on transcription, and the calculated value was 1. Accordingly, the fold change mediated by each STE was calculated as follows:

[0162]

[0163] where the calculated ratio represents the proportion of the specific barcode reads in the total reads in the sample. By comparing with the control (Ration Ctrl ), the ratio of each STE in the cDNA (Ration STE-cDNA ) or the ratio in the plasmid (Ration STE-plasmid ) was calculated to obtain the fold change (Activity STE ). A bio project of the present application has been created on NCBI, with the bio project ID PRJNA1088974.

[0164] 5. Screening results

[0165] The results are as follows Figure 6 As shown in the figure. The results indicate that the two repeats covering thousands of STEs are highly correlated, demonstrating the reliability of the STEM02 screening system. Compared with the control sequences, a total of 106 STEs in the two pSTEM02 libraries showed significant enhancement activity, ranging from 5.0 to 66.6 times.

[0166] Obtain the corresponding STE sequence from the barcode. Information and sequences of some STEs are shown in Tables 1 and 2.

[0167] Table 1. STEs with activation capabilities screened by the STEM02 system.

[0168]

[0169] Table 2 shows some STE sequences screened by the STEM02 system.

[0170]

[0171]

[0172] The above results demonstrate that the screening method of the present invention has successfully obtained a series of STEs that can achieve different degrees of transcriptional activation in rice.

[0173] All documents mentioned in this invention are incorporated herein by reference as if each document were individually incorporated by reference. Furthermore, it should be understood that after reading the foregoing teachings of this invention, those skilled in the art can make various alterations or modifications to this invention, and these equivalent forms also fall within the scope defined by the appended claims.

Claims

1. A report carrier for screening cis-acting elements, characterized in that, The reporter vector comprises a promoter and a reporter gene; the promoter region contains a cis-acting element to be screened; the coding region of the reporter gene contains: (i) a barcode sequence corresponding to the cis-acting element to be screened; (ii) at least one exogenous intron sequence.

2. The report carrier as described in claim 1, characterized in that, The exogenous introns in the transcripts generated by the reporter vector are excised; when reverse transcribing the transcripts generated by the reporter vector, primers that cross the exogenous introns are used.

3. The report carrier as described in claim 1, characterized in that, The barcode sequence consists of 3m bases N, where N is selected from A, T, G or C; wherein 2≤m≤5.

4. A report carrier library for high-throughput screening of cis-acting elements, characterized in that, The report carrier library contains X report carriers as described in claim 1, each carrying a different cis-acting element to be screened, wherein 50 ≤ X ≤ 10. 7 .

5. The report carrier library as described in claim 4, characterized in that, Each barcode sequence in the report carrier library corresponds to one candidate cis-acting element sequence.

6. A method for constructing the report carrier library as described in claim 4, characterized in that, Including the following steps: (i) Provide oligonucleotides that contain different candidate cis-acting elements, restriction sites and corresponding barcode sequences; (ii) The oligonucleotide is inserted into a plasmid backbone containing a reporter gene, and the oligonucleotide is inserted upstream of the reporter gene to form a reporter vector precursor; (iii) The oligonucleotides in the reporter system precursor are digested with enzymes, and the promoter is linked to the enzyme digestion site to form the reporter vector.

7. The method as described in claim 6, characterized in that, The structure of the oligonucleotide is shown in Formula I: Z0-STE-Z1-QNL(Formula I) In the formula, Each "-" can independently represent a covalent bond or a connector; Z0 is the amplification primer sequence; STE is a candidate cis-acting element sequence; Z1 is the restriction enzyme sequence; Q is the start codon; N is the barcode sequence; L represents no or no connector.

8. A method for high-throughput screening of cis-acting elements, characterized in that, Including the following steps: (1) Obtain the candidate cis-acting element sequence and the reference sequence; (2) Synthesize the report carrier library as described in claim 4; (3) The reporter vector library is introduced into host cells for transcription to extract total RNA; specific RNA reverse transcription is performed using transintron primers to obtain a cDNA library; sequences containing barcodes are amplified from the cDNA library; (4) In the reporter vector library obtained from the host cell, primers containing partial intron sequences are used to amplify the sequence of the barcode initially input into the library; (5) High-throughput sequencing (NGS) was used to perform sequencing analysis on the sequences amplified in steps (3) and (4) respectively, and the abundance of each barcode was read. The proportion of the reading amount of a specific barcode in the cDNA library to the total reading amount was recorded as R1, and the proportion of the reading amount of the same barcode in the reporter vector library to the total reading amount was recorded as R2. The ratio of R1 to R2 indicates the regulatory activity of the cis-acting element corresponding to the barcode.

9. The method as described in claim 8, characterized in that, The regulatory activity of the cis-acting element corresponding to the barcode is calculated using the formula shown in Equation II: In the formula, Ratio STE-cDNA This indicates the proportion of a specific barcode read in the cDNA library out of the total reads; Ration STE-plasmid This indicates the proportion of the number of reads of the same barcode in the report carrier library out of the total number of reads; Ration Ctrl-cDNA This indicates the proportion of barcode reads corresponding to the reference sequence in the cDNA library out of the total reads; Ration Ctrl-plasmid This indicates the proportion of barcode reads corresponding to the reference sequence in the report carrier library out of the total reads; n represents the number of reference sequences; Activity STE This indicates the fold increase in transcriptional levels by cis-acting elements.

10. A host cell, characterized in that, The host cell contains the reporter vector as described in claim 1 or the reporter vector library as described in claim 4.

Citation Information

Patent Citations

  • Method for realizing fixed-point activation of plant genome by directionally inserting enhancer

    CN120005931A

  • A scalable platform for the development of cell-type-specific viruses

    US20220025398A1

Cited By

  • Generated ultrahigh-activity plant core promoter library and application thereof

    CN121801916A