Method for identifying variable transcriptional start site and predictive enhancer by using transcriptional start site sequencing

Through TSS-seq technology and algorithm analysis, the problem of difficulty in identifying enhancer in the existing technology is solved, the precise location of transcription start sites in the genome and the direct detection of enhancer regions is realized, the molecular regulatory mechanism under low temperature stress is revealed, and the understanding of breeding and gene regulation networks is promoted.

CN120279992APending Publication Date: 2025-07-08SICHUAN NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410025891.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-08
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing enhancer identification technology has limitations, and it is impossible to accurately identify new or atypical enhancers, and the experimental cost is high, the operation is complex, and due to cell type and status, it is impossible to directly detect the transcriptional activity of active enhancers.

Method used

TSS-seq technology is used to accurately locate transcription start sites, combine algorithms to analyze bidirectional transcription clusters to predict potential enhancer regions, and use Arabidopsis model to identify variable transcription start sites under low temperature conditions to screen enhancer regions that affect gene expression.

Benefits of technology

The precise localization of transcription start sites in the genome and direct detection of enhancers was realized, revealing the molecular regulatory mechanism of plants under low temperature stress, promoting the development of breeding projects, and providing a comprehensive transcription start site map and new regulatory element prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
Patent Text Reader

Abstract

The invention provides a novel method for predicting a low-temperature induced enhancer by utilizing transcription start site sequencing, and belongs to the fields of bioinformatics and genomics. According to the method, a transcriptional start site is accurately identified by using a high-throughput sequencing technology, then a data bidirectional transcription region is analyzed as a potential enhancer, and finally a bidirectional transcription region near a variable transcriptional start site is predicted as an enhancer. Compared with the traditional enhancer identification technology, the method provided by the invention provides higher precision and resolution, so that researchers can more accurately identify and research the enhancer and the effect of the enhancer in gene regulation and control, and it is helpful to reveal how the plant responds to low-temperature stress at the molecular level. Besides, the method can guide a breeding project so as to cultivate crop varieties with better cold resistance and improve crop characters, and the principle and the method can be popularized to various plants and even non-plant organisms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for identifying transcription start sites based on TSS-seq data and an algorithm for identifying bidirectional transcription regions rich in enhancer characteristics. Background Art

[0002] Enhancers are very important in gene regulation. Although there are many enhancer identification techniques available, each has its own limitations. ChIP-seq involves using antibodies to specifically bind to common transcription factors or epigenetic markers (such as histone modifications) in enhancer regions, followed by sequencing. This method relies on known enhancer-related markers or transcription factors and may not be able to discover new or atypical enhancers. In addition, ChIP-seq experiments are costly, operationally complex, and require a sufficient amount of starting material. DNase-seq and ATAC-seq identify potential regulatory regions by measuring DNA accessibility. Although these methods can identify chromatin open regions, they cannot directly prove that the function of these regions is an enhancer. In addition, these techniques require high-quality DNA samples and may be affected by cell type and state. However, when predicting enhancers using TSS-seq, there are many advantages. The TSS-seq technology can accurately identify transcription start sites in the genome, as well as non-coding regions. Many enhancers can transcribe to produce enhancer RNAs, and TSS-seq can accurately capture these regions. Different from other technologies, TSS-seq can capture actively transcribed enhancers, providing direct functional evidence. TSS-seq data can be used for functional annotation of unannotated regions in the genome and can identify new regulatory elements in non-coding regions. Summary of the Invention

[0003] The object of the present invention is to overcome the shortcomings of the prior art and provide a method for predicting and annotating enhancer regions in the genome.

[0004] The present invention can accurately locate transcription start sites in the genome, including transcription in non-coding regions, can directly detect the transcription of active enhancers, which is information that other methods cannot provide, and can cover the entire genome, improving the comprehensive transcription start site map, can predict many new regulatory elements, and understand the molecular regulation mechanism of plants.

[0005] The present invention can guide breeding projects to cultivate cold-resistant crop varieties, which is of great significance for agricultural production. Its principle and method can be extended to a variety of plants, and even non-plants.

[0006] 1. A method for identifying alternative transcription start sites based on TSS-seq, comprising the following steps: Step 1: Obtain TSS-seq sequencing data. Taking Arabidopsis thaliana as an example, the samples taken are Arabidopsis thaliana leaves at room temperature and Arabidopsis thaliana leaves treated at 4°C for 24 hours during the same growth period, including two groups at room temperature and low temperature, with two biological replicates each.

[0007] Step 2: Preprocess the TSS-seq data. Use cutadapt to remove adapter data.

[0008] Step 3: Detect the quality of TSS-seq data. Use fastqc to detect the data quality.

[0009] Step 4: Map the data. Use hisat2 to perform data mapping to obtain a sam file.

[0010] Step 5: Convert to bam. Use samtools view and q50 to obtain the bam file after q50. Q50 is the quality value, indicating that the data quality is very good.

[0011] Step 6: Sort the bam file. Use samtools sort to obtain the sorted bam file.

[0012] Step 7: Call peaks. Use automatedClustering to obtain the peaks at room temperature and low temperature.

[0013] Step 8: Identify low-temperature variable transcription start sites. Through the peaks at room temperature and low temperature, use bedtools intersect to identify the specific peaks at low temperature. This part of the peaks is defined as the low-temperature specific variable transcription start sites, which is to accurately locate the transcription start sites in the genome using TSS-seq data.

[0014] 2. Predict potential enhancers through TSS-seq, including the following steps Step 1: Convert the bam file obtained in step 6 of claim 1 to the bedgraph format through the algorithm "bamCoverage". Use "--filterRNAstrand forward" to count the positive strand data, use "--filterRNAstrand reverse" to count the negative strand data, and use "-of bedgraph" to set the output file format to the bedgraph format.

[0015] Step 2: Use the bedGraphToBigWig script to convert the bedgraph format file to the bigwig format file.

[0016] Step 3: Use the software R studio (https: / / rstudio.com / ) (v4.3.2) to load the CAGEfightR package. Input the two bigwig files containing positive and negative strand data corresponding to each bam file, and identify the bidirectional transcription start sites in each library through the algorithm "clusterBidirectionally". Set the threshold "window = 51" to limit the scanning range of the program, and only the bidirectional transcription starts within the specified range are identified as a candidate regulatory element. Since these bidirectional transcription starts are obtained through the algorithm "clusterBidirectionally", each pair of bidirectional transcription starts is actually a pair of bidirectional transcription clusters, called BC, which is defined as the potential enhancer region.

[0017] 3. Through the identified alternative transcription sites and BC above, the enhancer region can be predicted. The low-temperature-specific alternative transcription start sites are divided into upstream alternative transcription start sites and downstream alternative transcription start sites, and then the BC situations that appear in the 1k region upstream or downstream of the transcription start site are screened. The appearance of these low-temperature-specific BCs leads to the appearance of low-temperature-specific alternative transcription start sites and affects gene expression. Therefore, this part of BC is predicted as the enhancer region.

[0018] 4. Finally, the tobacco transient transformation technology can be used to verify whether this part of the region has enhancer function. Description of the Drawings

[0019] Figure 1 is a flow chart for predicting enhancers based on TSS-seq data provided by the present invention. Detailed Embodiments

[0020] To make the present invention more obvious and understandable, the following combines the drawings to explain in detail the embodiments of the present invention and describes the steps of analyzing and processing TSS-seq.

[0021] The present invention accurately locates the transcription start sites in the genome and the enhancer RNAs (eRNAs) generated by enhancers. Compared with other methods, TSS-seq can directly identify the actively transcribed enhancers and can reveal subtle transcriptional activities.

[0022] The present invention helps to reveal how plants respond to low-temperature stress at the molecular level, can promote the understanding of the gene regulatory network, and is applicable to a variety of plants.

[0023] 1. A method for identifying alternative transcription start sites based on TSS-seq, comprising the following steps: Step 1: Obtain TSS-seq sequencing data. The samples taken are Arabidopsis thaliana leaves at room temperature and Arabidopsis thaliana leaves treated at 4°C for 24 hours during the same growth period, including two groups at room temperature and low temperature respectively, with two biological replicates. Then, obtain TSS-seq sequencing data through high-throughput sequencing technology.

[0024] Step 2: Preprocess the TSS-seq data. Use cutadapt to remove adapter data.

[0025] Step 3: Detect the quality of TSS-seq data. Use fastqc to detect the data quality. A quality above q20 indicates good data.

[0026] Step 4: Map the data. Use hisat2 to map the data and obtain the sam file.

[0027] Step 5: Convert to bam. Use samtools view -b -q 50 file.sam>out.Q50.bam to obtain the bam file after Q50. Q50 is the quality value, indicating very good data quality.

[0028] Step 6: Sort the bam file. Sort the bam file. Use samtools sort -@ 10 -Toutput_sort -o output_sort.bam out.Q50.bam to obtain the sorted bam file.

[0029] Step 7: Call peaks. Use automatedClustering inputdir outputdir tpm idroutputdir2 project to obtain the peaks at room temperature and low temperature. Among them, inputdir contains the directory path of the input files, with at least 2 bam files; outputdir is the directory path of the output result files, which contains two files, toppeak and bottom peak. The top peak is the cluster extracted at the top of each hierarchical cluster class, and the bottom peak is the cluster extracted at the bottom of each hierarchical cluster class; tpm is the threshold used during clustering, recommended value is 0.1; idr is the threshold for discarding clusters that cannot be reproduced, recommended value is 0.1; outputdir2 is the output file path of the scatter plot, which can analyze the stability between the two replicates; project is the project name, which can be any name you like.

[0030] Step 8, identify low-temperature variable transcription start sites. By using bedtools intersect on the peaks at room temperature and low temperature, specific peaks at low temperature can be identified. These peaks are defined as low-temperature specific variable transcription start sites, and the TSS-seq data is used to precisely locate the transcription start sites in the genome.

[0031] 2. Predict potential enhancers through TSS-seq, including the following steps Step 1, convert the bam file obtained in step 6 of claim 1 into bedgraph format through the algorithm "bamCoverage". Set the offset with "--offset 1", indicating that the first read of the read is 1. Use "--filterRNAstrand forward" to count the positive-strand data. Set the output file format to bedgraph format with "-of bedgraph". Set the size of each bin to 1bp with "-binSize 1". Set the output file name with "-o" to obtain the data on the positive strand. And use "--filterRNAstrand reverse" to count the negative-strand data.

[0032] Step 2, convert the bedgraph format file into a bigwig format file using the bedGraphToBigWig script.

[0033] Step 3, use the R studio (https: / / rstudio.com / ) software (v4.3.2), load the CAGEfightR package. Input the two bigwig files containing positive and negative strand data corresponding to each bam file. Identify the bidirectional transcription start sites in each library through the algorithm "clusterBidirectionally". Set the threshold "window = 51" to limit the scanning range of the program. Only the bidirectional transcription within the specified range is identified as a candidate regulatory element. Since these bidirectional transcription starts are obtained through the algorithm "clusterBidirectionally", each pair of bidirectional transcription starts is actually a pair of bidirectional transcription clusters, called BC, which is defined as the potential enhancer region.

[0034] 3. By using the variable transcription sites identified above and BC, the enhancer region can be predicted. The low-temperature specific variable transcription start sites are divided into upstream variable transcription start sites and downstream variable transcription start sites. The upstream variable transcription start sites are further divided into upstream variable transcription start sites to 1k upstream and upstream variable transcription start sites to the reference transcription start site and then to 1k downstream of the reference transcription start site. The downstream variable transcription start sites can be divided into downstream variable transcription start sites to 1k upstream of the reference transcription start site and downstream variable transcription start sites to 1k downstream. Next, analyze the BC situation in these four regions. It can be analyzed from the following two aspects: the situation where there is no BC during RT in this region and there is BC at 4°C, and the situation where there is BC during RT in this region and there is BC at different positions at 4°C. Finally, screen out BC, and speculate that the appearance of these low-temperature specific BC leads to the appearance of low-temperature specific variable transcription start sites and affects gene expression. Therefore, this part of BC is predicted as the enhancer region.

[0035] 4. Finally, the tobacco transient transformation technology can be used to verify whether this part of the region has enhancer function.

[0036] The method of the present invention can quickly and effectively predict enhancers, is easy to implement, can directly detect the transcriptional activities of active enhancers, and can detect new enhancers.

[0037] By predicting the active enhancers under low-temperature conditions, the present invention helps to reveal how plants respond to low-temperature stress at the molecular level, can provide a deeper understanding of the gene regulatory network, especially the regulatory mechanism during environmental changes. The present invention helps to promote the prediction of enhancers in other plant species.

[0038] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for identifying alternative transcription start sites and predicting enhancers using transcription start site sequencing (TSS-seq), comprising the following steps: (1) Cultivate plants and subject them to low-temperature treatment; (2) Sample, extract RNA, and perform high-throughput sequencing to obtain TSS-seq data; (3) Preprocess the data, and use bioinformatics software to remove adapters, perform quality control, and sequence alignment on the TSS-seq data; (4) Process the sam and bam files. First, convert the sam file to bam, and q50 is required, and then sort the bam file; (5) Peak detection. Perform peak detection on the obtained bam file to obtain transcription start site information; (6) Obtain low-temperature-specific transcription start site data based on the transcription start site data information at low temperature and room temperature.

2. According to the bam analysis obtained in step (4) of claim 1, predict enhancers, comprising the following steps: (1) Convert the bam file to the bedgraph format; (2) Convert the bedgraph file to the bigwig format; (3) Identify significant bidirectional transcription regions through an algorithm and identify this part of the region as potential enhancers.

3. Predict enhancer regions based on the alternative transcription site data and potential enhancer data obtained in claims 1 and 2, comprising the following steps: (1) Divide the alternative transcription start site data into upstream alternative transcription start sites and downstream alternative transcription start sites; (2) Analyze the occurrence of potential enhancers in the 1k region upstream or downstream of the upstream and downstream alternative transcription start sites, and predict the analysis results as enhancers.

4. The plants in claim 1 include but are not limited to Arabidopsis thaliana. Taking Arabidopsis thaliana as an example, the following cultivation can be carried out: (1) Cultivate Arabidopsis thaliana under long-day (16 hours of light / 8 hours of dark) and room temperature (22°C) conditions; (2) Select a group of Arabidopsis thaliana plants for low-temperature treatment, and treat them at 4°C for 24 hours; (3) Sample, extract RNA, and perform high-throughput sequencing to obtain TSS-seq data.

5. Compared with other methods (ChIP-seq or DNase-seq), predicting enhancers from TSS-seq data has its unique advantages, and can accurately identify transcription start sites and nearby active enhancers under low-temperature conditions.

6. By predicting active enhancers under low-temperature conditions, this technology helps to reveal how plants respond to low-temperature stress at the molecular level and can provide a deeper understanding of the gene regulatory network.

7. This technology can play a significant role in deepening the understanding of the plant stress response mechanism and improving the cold resistance breeding of crops. Its principles and methods can be extended to a variety of plants, and even non-plant organisms.