Bioinformatics method for improving number and accuracy of prediction enhancers
By combining the MH-seq and DNase-seq methods and combining specific pattern recognition techniques, the problems of small number of enhancers and low accuracy in the prior art are solved, and more efficient enhancer prediction is achieved.
Patent Information
- Application Number
- CN202410025892.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-08
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, when identifying open chromatin regions in plant genomes, the number of enhancers obtained using a single method is small and the accuracy is low, and it is impossible to fully display the open chromatin regions of the whole genome.
Using two methods combined with MH-seq and DNase-seq, the entire genome open chromatin region was identified and integrated through data processing and bioinformatics tools, and combined with specific patterns to identify potential enhancer regions to improve prediction accuracy.
通过结合两种方法,显著提高了增强子数量和预测准确性,增强子鉴定结果更为全面和可靠。
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bioinformatics and relates to a bioinformatics method for improving the number and accuracy of predicted enhancers. Background Art
[0002] Enhancers refer to special sequences or elements that regulate gene expression in animal and plant cells. These enhancers play a key role in the growth and development of plants, can affect the transcriptional activity of genes, and thus regulate the expression levels of related genes. By bioinformatics methods, analyzing open chromatin regions in the plant genome can identify potential enhancers with regulatory functions. Currently, common methods for identifying open chromatin regions include ATAC-seq, DNase-seq, and MNase-seq. However, there are significant technical differences between methods, and the results obtained using a single method cannot more comprehensively display the open chromatin regions of the entire genome, with fewer predicted enhancers and low accuracy. Summary of the Invention
[0003] Aiming at the deficiencies of the above background art, the present invention provides a bioinformatics method for improving the number and accuracy of predicted enhancers. By combining two methods for identifying open chromatin, a method with a greater number of predicted enhancers and higher accuracy can be obtained.
[0004] A bioinformatics method for improving the number and accuracy of predicted enhancers includes the following steps: S1. Obtain the data of MH-seq and DNase-seq, and perform quality control, adapter sequence excision, and sequence alignment on the data; S2. Use bioinformatics tools to identify open chromatin regions in the whole genome using MH-seq and DNase-seq; S3. Integrate and analyze the data of MH-seq and DNase-seq to obtain the intersection and union of open chromatin regions of the two; S4. Identify possible enhancer regions in the MH-seq and DNase-seq data through specific patterns or features to obtain predicted enhancers; S5. Based on the enhancer subgroups established above, compare the number and accuracy of enhancers predicted using a single method.
[0005] Further, the data processing in S1 is specifically as follows: Use the FastQC software to perform quality control on the original data, and then use the Cutadapt program to excise the adapter sequences.
[0006] Further, the data processing in S1 is also as follows: Use the bwa and bowtie software for sequence alignment, and limit the maximum number of alignments to the genome to 1 and do not allow mismatches.
[0007] Further, the identification of the genome-wide open chromatin regions in S2 is specifically as follows: The Jazz software and the Popera software are respectively used to identify the genome-wide open chromatin regions.
[0008] Further, the identification of the genome-wide open chromatin regions by S2 is also as follows: The bedtools software is used for processing, and the regions with at least 1 bp overlap in the biological replicates are screened to obtain MHS and DHS.
[0009] Further, the identification of the possible enhancer regions in the MH-seq and DNase-seq data by specific patterns or features in S4 is specifically as follows: All DHS and MHS in the intergenic regions are set, and the DHS and MHS 1 kb upstream of the TSS are removed as potential enhancers.
[0010] Further, the identification of the possible enhancer regions in the MH-seq and DNase-seq data by specific patterns or features in S4 is specifically as follows: The DESeq package is used for RNA-seq analysis to obtain the gene expression level situation, and the expression levels of the genes adjacent to the predicted enhancers are jointly analyzed to obtain the enhancers specifically predicted by DHS and MHS and the enhancers with higher prediction accuracy. Brief Description of the Drawings
[0011] Figure 1 is a flowchart of the present invention.
[0012] Figure 2 is the positions of DHS and MHS identified by DNase-seq data and MH-seq on the genome.
[0013] Figure 3 is the number of enhancers specifically predicted by DNase-seq and MH-seq and the number of enhancers jointly predicted. Embodiments
[0014] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be described in detail below in conjunction with specific embodiments and the accompanying drawings. The experimental methods without specific conditions noted in the following examples are usually in accordance with conventional conditions or the conditions recommended by the manufacturer. The test materials used in the following examples are all obtained from regular biochemical reagent stores without special instructions. Unless otherwise stated, percentages and parts are calculated by weight. Unless otherwise defined, all professional and scientific terms used herein have the same meaning as those familiar to those skilled in the art. In addition, any methods and materials similar or equivalent to the described content can be applied to the present invention. The preferred implementation methods and materials described herein are for illustrative purposes only.
[0015] Accordingly, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0016] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0017] In the description of the present invention, it should also be noted that the MH-seq and DNase-seq data obtained in the present invention are all from the inventors. The inventors predict enhancers by obtaining the changes in the expression levels of the neighboring genes of the identified MHSs and DHSs. If the expression of the neighboring gene is up-regulated, it is called a predicted enhancer.
[0018] Example 1: Obtaining DHSs and MHSs Tubers of potatoes stored at normal temperature and low temperature for 15 days were obtained, and their MH-seq and DNase-seq data were obtained. First, FastQC was used to evaluate the quality of the raw data. According to indicators such as read length (the size of a single DNA fragment), the number of reads, the sequencing quality of each base, GC content, and the proportion of duplicate reads, it was judged whether subsequent analysis could be carried out. Then, the adapter sequences were trimmed through the Cutadapt program, reads less than 4 bp were removed, and the base sequences with a sequencing quality greater than 20 were retained. Subsequently, FastQC was used again to evaluate the quality of the data. If the quality was qualified, the next step of analysis was carried out. Then, bwa and bowtie were used for sequence alignment, and the obtained files were converted into the bam file format. The bam files of MH-seq and DNase-seq were respectively input into the Jazz and Popera software for the identification of open chromatin regions in the whole genome, which we call DHSs and MHSs. Then, DHSs and MHSs with at least 1 bp of repetition in the biological replicates were obtained, and these were used for downstream analysis.
[0019] Example 2: Predicting More and More Accurate Enhancers Potato tubers stored at normal temperature and low temperature for 15 days were obtained, and their RNA-seq data were obtained. RNA-seq analysis was performed using the DESeq package to obtain the gene expression levels in potato tubers at low temperature. Then, the "bedtools intersect" command was used to obtain the complement between MHS and DHS at low temperature, obtaining 32,685 MHS specific to MH-seq in potato tubers at low temperature and 17,001 DHS specific to DNase-seq in potato tubers at low temperature. All DHS and MHS located in the intergenic region, excluding DHS and MHS within 1 kb upstream of the TSS, were defined as potential enhancers. A total of 37,915 enhancers were identified using DHS and MHS, including 5,582 enhancers specific to Dnas-seq, 19,506 enhancers specific to MH-seq, and 12,827 common enhancers. DHS identified 18,409 enhancers, including 4,538 specific to low temperature, 6,514 specific to normal temperature, and 7,357 common enhancers. MHS identified 31,681 enhancers, including 4,914 specific to low temperature, 5,849 specific to normal temperature, and 20,918 common enhancers. We believe that the probability of common potential enhancers being true enhancers is higher, improving the accuracy of enhancer identification.
[0020] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A bioinformatics method for increasing the number and accuracy of predicted enhancers, characterized in that It includes the following steps: S1. Obtain the data of MH-seq and DNase-seq, and perform quality control, adapter sequence excision, and sequence alignment on the data; S2. Use bioinformatics tools to identify open chromatin regions in the whole genome using MH-seq and DNase-seq; S3. Integrate and analyze the data of MH-seq and DNase-seq to obtain the intersection and union of open chromatin regions; S4. Identify possible enhancer regions in the MH-seq and DNase-seq data through specific patterns or features to obtain predicted enhancers; S5. Based on the enhancer subgroups established above, compare the number and accuracy of enhancers predicted using a single method.
2. Data processing of the MH-seq and DNase-seq described in claim 1, characterized in that, In the step S1, it specifically includes: Use the FastQC software to perform quality control on the raw data, and then use the Cutadapt program to excise the adapter sequences.
3. Data processing of the MH-seq and DNase-seq described in claim 1, characterized in that, In the step S1, it also includes: Use the bwa and bowtie software for sequence alignment, and limit the maximum number of alignments to the genome to 1 and do not allow mismatches.
4. A method for identifying open chromatin regions across the whole genome as recited in claim 1, characterized in that In the step S2, it specifically includes: Use the Jazz software and Popera software respectively to identify the open chromatin regions in the whole genome of MH-seq and DNase-seq.
5. A method for identifying open chromatin regions across the whole genome as described in claim 1, characterized in that In the step S2, it specifically includes: Then use the bedtools software for processing, screen the regions with at least 1bp overlap in biological replicates to obtain MHS and DHS.
6. A method for identifying possible enhancer regions in MH-seq and DNase-seq data through specific patterns or features as described in claim 1, characterized in that In the step S4, it also includes: Set all DHS and MHS in the intergenic region and remove DHS and MHS 1kb upstream of the TSS as potential enhancers.
7. A method for identifying possible enhancer regions in MH-seq and DNase-seq data through specific patterns or features as described in claim 1, characterized in that, In the step S4, it also includes: Use the DESeq package to perform RNA-seq analysis, obtain the gene expression level, conduct a joint analysis of the expression levels of genes adjacent to the predicted enhancers, and obtain enhancers specifically predicted by DHS and MHS and enhancers with higher prediction accuracy.