Methods, systems, and storage media for analyzing cell-specific transcriptional pause sites

CN118899031BActive Publication Date: 2026-08-18NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411020987.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2026-08-18
Estimated Expiration
2044-07-29

AI Technical Summary

Technical Problem

[0003]因此,造成了关于基因调控方面的知识缺口,不能深入了解各种细胞类型和生理条件下暂停现象的多样性及其功能意义

Benefits of technology

[0051] 1. Abundant data resources and a large amount of information. TransPause integrates sequencing data from 995 newborn RNA samples from three species, annotating over 1.7 million transcriptional pause sites and calculating detailed information such as sequence characteristics and conservation of these sites. It also specifically annotates a large number of known pause sites of significant importance. This provides researchers with an unprecedented large-scale data resource.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118899031B_ABST
    Figure CN118899031B_ABST
Patent Text Reader

Abstract

The application provides an analysis method, system and storage medium of a cell-specific transcription pause site. Through comprehensive database and comprehensive data arrangement and analysis, the diversity of the pause phenomenon and the functional significance thereof under various cell types and physiological conditions can be deeply understood. The application can comprehensively understand the diversity and distribution characteristics of the transcription pause under different cell types and physiological conditions; can deeply analyze the relationship between the transcription pause and the gene expression regulation and the potential functional significance thereof; can explore the role of the transcription pause in cell processes, such as differentiation, development and the like; and can provide new biomarkers and targets for diagnosis and treatment of related diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of transcriptome sequencing technology, specifically relating to all-flash memory system technology, and particularly to a method, system, and storage medium for analyzing cell-specific transcriptional pause sites. Background Technology

[0002] Transcriptional pausing plays a significant regulatory role in many important processes such as cell differentiation, development, and biological rhythms. In recent years, transcriptional pausing has received widespread attention, and numerous reports have documented the existence of pausing sites; however, a comprehensive database to collect and organize this information remains lacking.

[0003] This has created a knowledge gap regarding gene regulation, hindering a deeper understanding of the diversity and functional significance of pause phenomena across various cell types and physiological conditions. These subsequent advancements play a crucial role in understanding transcriptional regulation and its role in cellular processes.

[0004] The foregoing statements are for informational purposes only and are not intended to provide background information in connection with this application. Unless otherwise stated herein, the content described in this section is not prior art to the rest of this application. Summary of the Invention

[0005] The method, system, and storage medium for analyzing cell-specific transcriptional pause sites proposed in this invention, through comprehensive database analysis and data processing, can provide a deeper understanding of the diversity and functional significance of pause phenomena under various cell types and physiological conditions.

[0006] According to a first aspect of the embodiments of this application, a method for analyzing cell-specific transcriptional pausing sites is provided, comprising:

[0007] We collected N nascent RNA sequencing data from different cell lines of three species: human, mouse, and fruit fly; we also collected ChIP-seq peak data of regulatory factors related to transcriptional pausing.

[0008] Based on the transcriptional pausing level reflected in the newborn RNA sequencing data, the pausing index (PI) of transcriptional activity is determined; based on the transcriptional elongation level reflected in N newborn RNA sequencing data, the elongation index (EI) of transcriptional activity is determined; based on the RNA polymerase elongation rate data reflected in N newborn RNA sequencing data, the elongation rate (ES) of transcriptional activity is determined.

[0009] Significant transcriptional pause sites were identified based on newborn RNA sequencing data, and known information about these pause sites was labeled. This known information included: neighboring pause sites near the transcription start site and global pause sites across the genome.

[0010] The sequence and structural features of transcriptional pausing sites were determined; and based on the transcriptional pausing sites, analyses of transcriptional pausing regulatory factors, differential transcriptional activity, and functional genomics were obtained.

[0011] In some embodiments of this application, transcriptional pausing regulatory factor analysis is obtained based on transcriptional pausing sites, including:

[0012] Obtain the distribution of known transcriptional pausing motifs on genes;

[0013] The distribution was compared with the peak data of the ChIP-seq of the regulatory factors;

[0014] Predicting and annotating the regulatory factors that control each pause site provides clues for subsequent mechanistic studies.

[0015] In some embodiments of this application, the regulatory factors controlling each pause site are predicted and annotated, including:

[0016] The regulators of transcriptional arrest include NELF and DSIF. The .bed files of the regulators were obtained from the ENCODE database.

[0017] The intersect command of the bedtools tool was used to determine the intersection of the NELF and DSIF peak locations with transcriptional pause sites;

[0018] Based on the intersection, peak values ​​are determined to overlap with known transcriptional pause sites, and the regulatory factors for each pause site are predicted.

[0019] In some embodiments of this application, differential analysis of transcriptional activity is obtained based on transcriptional pausing sites, including:

[0020] The differences in transcriptional activity of genes among different cell lines were analyzed using statistical methods to identify genes that regulate transcriptional arrest under different cell states. The statistical methods included t-test, Mann-Whitney U test, and Wilcoxon rank-sum test.

[0021] In some embodiments of this application, genes that regulate transcriptional pausing under different cellular states are identified, including:

[0022] Use Python's stats module to determine whether the transcriptional activity data of any two cell lines in multiple cell lines are normally distributed;

[0023] If both groups are normally distributed, use the parametric T-test for difference analysis; if either or both groups deviate from a normal distribution, use the non-parametric Mann-Whitney U test to assess the difference between the two groups.

[0024] Genes with differential transcriptional activity were screened using a p-value less than 0.05 as the significance level.

[0025] In some embodiments of this application, functional genomics analysis is obtained based on transcriptional pause sites, including:

[0026] By identifying genes with similar transcriptional activities, we can predict potential transcriptional or functional associations.

[0027] Gene enrichment analysis was used to identify biological processes or pathways associated with transcriptional arrest.

[0028] Gene set enrichment analysis was used to compare gene sets regulated by transcriptional pausing under different cellular states.

[0029] In some embodiments of this application, finding genes with similar transcriptional activity includes using transcriptional activity data between samples, calculating the Euclidean distance between each gene and other genes, identifying the top 10 nearest neighbor genes of each gene, which may have transcriptional or functional associations, and calculating the Pearson correlation coefficient and Spearman correlation coefficient of each gene with its nearest neighbors and their significance, thereby quantitatively describing their relationship.

[0030] Gene enrichment analysis, including enrichment analysis based on transcriptional activity data and a list of differentially expressed genes, identifies biological processes or pathways that may be associated with transcriptional pausing. The Benjamini-Hochberg method is used to perform multiple hypotheses, and a corrected p-value less than 0.05 is considered significant enrichment.

[0031] According to a second aspect of the embodiments of this application, a system for analyzing cell-specific transcriptional pausing sites is provided, comprising:

[0032] The database was used to collect N nascent RNA sequencing data from different cell lines of three species: human, mouse, and fruit fly; and to collect ChIP-seq peak data of regulatory factors related to transcriptional pausing.

[0033] The transcriptional activity analysis module is used to determine the pause index (PI) of transcriptional activity based on the transcriptional pause level reflected in the nascent RNA sequencing data; to determine the elongation index (EI) of transcriptional activity based on the transcriptional elongation level reflected in N nascent RNA sequencing data; and to determine the elongation rate (ES) of transcriptional activity based on the RNA polymerase elongation rate data reflected in N nascent RNA sequencing data.

[0034] Pause site annotation module: used to identify significant transcription pause sites based on newborn RNA sequencing data and to annotate known information about the transcription pause sites, including: neighboring pause sites near the transcription start site and global pause sites across the genome;

[0035] The differential analysis module is used to determine the sequence and structural features of transcriptional pausing sites; and to obtain analyses of transcriptional pausing regulatory factors, differential transcriptional activity, and functional genomics based on transcriptional pausing sites.

[0036] According to a third aspect of the embodiments of this application, a memory is provided, comprising: a storage unit for storing executable instructions; and a processing unit for being connected to the memory to execute the executable instructions to complete a method for analyzing cell-specific transcriptional pause sites.

[0037] According to a fourth aspect of the embodiments of this application, a computer-readable storage medium is provided having a computer program stored thereon; the computer program is executed by a processor to implement a method for analyzing cell-specific transcriptional pause sites.

[0038] By employing the cell-specific transcriptional pause site analysis method, system, memory, and storage medium of this application, and through comprehensive database analysis and data processing, we can gain a deeper understanding of the diversity and functional significance of pause phenomena under various cell types and physiological conditions.

[0039] Meanwhile, this application also has the following technical effects: it can provide a comprehensive understanding of the diversity and distribution characteristics of transcriptional pausing under different cell types and physiological conditions; it can provide in-depth analysis of the relationship between transcriptional pausing and gene expression regulation and its potential functional significance; it can explore the role of transcriptional pausing in cellular processes, such as differentiation and development; and it can provide new biomarkers and targets for the diagnosis and treatment of related diseases. Attached Figure Description

[0040] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0041] Figure 1 The diagram illustrates the steps of a method for analyzing cell-specific transcriptional pause sites according to an embodiment of this application.

[0042] Figure 2 The diagram illustrates the steps of a method for analyzing cell-specific transcriptional pause sites according to an embodiment of this application.

[0043] Figure 3 The diagram shows a computational gene partitioning diagram used in calculating transcriptional activity data;

[0044] Figure 4 The diagram below illustrates the salience calculation process for pause sites.

[0045] Figures 5-8The diagram below illustrates the results of using the Transpause database.

[0046] Figure 9 The diagram shows a schematic of the structure of a cell-specific transcriptional pause site analysis system according to an embodiment of this application;

[0047] Figure 10 The diagram shows a schematic representation of a memory according to an embodiment of this application. Detailed Implementation

[0048] Regarding this application, it is noted that although research on transcriptional pausing has deepened in recent years and numerous reports document the existence of transcriptional pausing sites, a comprehensive database of transcriptional pausing has not yet been established. This creates a significant knowledge gap regarding gene regulation.

[0049] This application, by establishing a comprehensive database and analysis platform, can fill existing knowledge gaps, enabling researchers to: gain a comprehensive understanding of the diversity and distribution characteristics of transcriptional pausing under different cell types and physiological conditions; conduct in-depth analysis of the relationship between transcriptional pausing and gene expression regulation and its potential functional significance; explore the role of transcriptional pausing in cellular processes, such as differentiation and development; and provide new biomarkers and targets for the diagnosis and treatment of related diseases.

[0050] This application has the following main effects and advantages:

[0051] 1. Abundant data resources and a large amount of information. TransPause integrates sequencing data from 995 newborn RNA samples from three species, annotating over 1.7 million transcriptional pause sites and calculating detailed information such as sequence characteristics and conservation of these sites. It also specifically annotates a large number of known pause sites of significant importance. This provides researchers with an unprecedented large-scale data resource.

[0052] 2. It can comprehensively characterize transcriptional dynamics. This application proposes a unified calculation method across species to measure transcriptional activity indicators such as the pause index, elongation index, and elongation rate of each gene. It can depict the transcriptional pause and elongation dynamics of genes in a multi-dimensional and detailed manner, and comprehensively characterize this key link in transcriptional regulation.

[0053] 3. Powerful and flexible search function. The advanced search function of this application can retrieve pause sites and events of interest based on various conditions such as gene name, location, sequence, and regulatory factors. The search scope is wide and highly flexible, maximizing the extraction of relevant information.

[0054] 4. Cell-specific difference analysis. The statistical analysis methods used in this application can compare the differences in transcriptional arrest among different cell types, analyze the impact of cellular heterogeneity on arrest regulation, and provide new clues for disease diagnosis and mechanism research.

[0055] 5. Multi-level functional analysis tools. The analysis module of this application provides a variety of tools, such as finding related genes, pathway enrichment analysis, and gene set enrichment analysis, which can explore the potential biological functions and mechanisms of transcriptional arrest at different levels.

[0056] 6. User-friendly interface and excellent visualization. The TransPause database in this application features a clear and concise web interface and provides various visualization methods to intuitively display the analysis results, offering a good user experience and maximizing the user's understanding and utilization of this information.

[0057] To make the technical solutions and advantages of the embodiments of this application clearer, the exemplary embodiments of this application will be described in further detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not an exhaustive list of all embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other.

[0058] Example 1

[0059] Figure 1 The diagram illustrates the steps of a method for analyzing cell-specific transcriptional pause sites according to an embodiment of this application.

[0060] like Figure 1 As shown, the method for analyzing cell-specific transcriptional pausing sites in this application includes the following steps:

[0061] S1: Collect N newborn RNA sequencing data from different cell lines of three species: human, mouse, and fruit fly; collect ChIP-seq peak data of regulatory factors related to transcriptional pausing.

[0062] In the specific implementation of this application, 995 nascent RNA sequencing data points from different cell lines of three species—human, mouse, and fruit fly—were collected, including GRO-seq, PRO-seq, NET-seq, and mNET-seq data. These raw sequencing data underwent standard quality control, alignment, and signal value calculation to obtain data in a uniform Bigwig format. Simultaneously, ChIP-seq peak data of regulatory factors related to transcriptional pausing were also collected.

[0063] S2: Determine the pause index (PI) of transcriptional activity based on the level of transcriptional pauses reflected in the sequencing data of newborn RNA; determine the elongation index (EI) of transcriptional activity based on the level of transcriptional elongation reflected in the sequencing data of N newborn RNAs; determine the elongation rate (ES) of transcriptional activity based on the RNA polymerase elongation rate data reflected in the sequencing data of N newborn RNAs.

[0064] In practice, the newly generated RNA sequencing signal consists of a large amount of sequence information in the collected newly generated RNA sequencing files. Through genome alignment, the sequence information is converted into numerical information, i.e., the sequencing signal, for each genomic location. A minimum value of 0 indicates that no gene is being transcribed at the current genomic location, while a value greater than 0 indicates that a gene is being transcribed at the current location.

[0065] Pause Index (PI): The pause index is defined as the ratio of the density of nascent RNA sequencing signal within 200 bases downstream of the gene start site to the signal density within the remaining gene. This index reflects the degree of transcriptional pause, i.e., it measures the extent to which RNA polymerase stops after transcription begins.

[0066] Elongation Index (EI): For genes longer than 2 kb, the elongation index is calculated as the ratio of the signal density of the first 25% to the last 75% of the gene after excluding the first 300 bp. This ratio reflects the degree of transcriptional elongation, i.e., the movement of RNA polymerase within the gene.

[0067] Elongation rate (ES): First, introns overlapping with exons are removed. Then, linear regression is used to calculate the slope of the signal within the intron along the 5'->3' direction. Finally, log10(-1 / slope) is calculated as the polymerase elongation rate within the intron. The median elongation rate of all introns within the gene is selected as the gene's elongation rate. This metric reflects the motility of RNA polymerase.

[0068] These indicators are calculated based on sequencing signals of newly generated RNA in the upstream region of the gene and inside and outside the gene, respectively, and can meticulously characterize the transcriptional dynamics of different genes.

[0069] S3: Identify significant transcriptional pause sites based on newborn RNA sequencing data and label known information about these pause sites, including: neighboring pause sites near transcription start sites and global pause sites across the genome.

[0070] In practice, the location of the highest signal value in the newly generated RNA data calculated in Patent 2, within the region of +20 to +350 bp downstream of the transcription start site, was defined as the proximal pause site. Within a 200 bp sliding window, locations with signal values ​​exceeding the mean by 3 standard deviations were defined as global pause sites. The proximal and global pause sites were compared with manually collected literature-reported pause sites, and these reported sites were labeled in the database.

[0071] S4: Determine the sequence and structural characteristics of transcriptional pausing sites; and based on the transcriptional pausing sites, conduct analyses of transcriptional pausing regulatory factors, differential transcriptional activity, and functional genomics.

[0072] In practice, for each identified pause site, this application calculates its sequence GC content, minimum free energy structure, and evolutionary conservation. These characteristics reflect the stability and accessibility of the pause site, which helps to further explore the impact of the pause site on polymerase motility.

[0073] For the pause sites identified in the preceding steps, a 20-base-pair window centered on the pause site is taken as the pause sequence, and the GC content of this 20 bp sequence is calculated. The GC content is the ratio of G or C nucleotides in the sequence. GC content can affect the stability of RNA secondary structure, thus reflecting the stability and accessibility of the pause site region. Using the ViennaRNA software package, with the 20 bp pause sequence as input, the minimum free energy secondary structure of the sequence is calculated, and the free energy value is given. The minimum free energy structure reflects the structural stability and accessibility of the pause site region, which helps to understand its impact on the RNA polymerase elongation rate.

[0074] Using the PhastCons software, an evolutionary conservation score was calculated for each pause site, ranging from 0 to 1. The higher the score, the more evolutionarily conserved the site is. By analyzing the evolutionary conservation of pause sites, we can gain insights into the functional significance and evolutionary history of these sites.

[0075] Figure 2 The diagram illustrates the steps of a method for analyzing cell-specific transcriptional pause sites according to an embodiment of this application.

[0076] like Figure 2 As shown, S4 includes the analysis of transcriptional pausing regulatory factors based on transcriptional pausing sites, including: S41: obtaining the distribution of known transcriptional pausing motifs on genes; S42: comparing the distribution with the ChIP-seq peak data of regulatory factors; S43: predicting and annotating the regulatory factors that regulate each pausing site, providing clues for subsequent mechanism studies.

[0077] The predicted and annotated regulatory factors controlling each pause site include:

[0078] Regulatory factors of transcriptional arrest include NELF (negative elongation factor) and DSIF (DRB sensitivity-inducing factor). The BED files of these regulatory factors were downloaded from the ENCODE database, and the `intersect` command of the `bedtools` tool was used to determine the intersection between the locations of NELF and DSIF peaks and transcriptional arrest sites. This step identifies which peaks overlap with known transcriptional arrest sites, providing evidence to predict whether these factors regulate arrest.

[0079] In S4, the analysis of differences in transcriptional activity based on transcriptional pausing sites includes: analyzing the differences in transcriptional activity of genes between different cell lines using statistical methods to identify genes that are regulated by transcriptional pausing in different cell states; the statistical methods include t-test, Mann-Whitney U test and Wilcoxon rank-sum test.

[0080] Among them, the genes that were identified as regulating transcriptional pausing under different cellular states include:

[0081] First, for any two cell lines from multiple cell lines, we use the stats module in Python to determine whether the data are normally distributed. If both groups are normally distributed, we will continue to use the parametric T-test for difference analysis.

[0082] However, if any one or both groups deviate from a normal distribution, we will use the nonparametric Mann-Whitney U test to assess the differences between the two groups. Regardless of the test used, a p-value less than 0.05 is used as the significance level to screen genes with differences in transcriptional activity.

[0083] S4 provides functional genomics analysis based on transcriptional pause sites, including:

[0084] By identifying genes with similar transcriptional activities, we can predict potential transcriptional or functional associations; through gene enrichment analysis, we can identify biological processes or pathways associated with transcriptional pausing; and through gene set enrichment analysis, we can compare gene sets regulated by transcriptional pausing under different cellular states.

[0085] The process of finding genes with similar transcriptional activity involves using transcriptional activity data between samples, calculating the Euclidean distance between each gene and other genes, identifying the top 10 nearest neighbor genes for each gene (which may have transcriptional or functional associations), and calculating the Pearson and Spearman correlation coefficients and their significance for each gene with its nearest neighbors to quantitatively describe their relationship.

[0086] The gene enrichment analysis includes transcriptional activity data and a list of differentially expressed genes. Enrichment analysis of the list of differentially expressed genes indicates that the biological processes or pathways enriched may be related to transcriptional pausing. The Benjamini-Hochberg method is used to perform multiple hypotheses, and a corrected p-value of less than 0.05 is considered significant enrichment.

[0087] Finally, this application also includes a website interface. All analysis results are stored in a PostgreSQL database, and researchers can query, filter, and visualize data of interest through a user-friendly web interface.

[0088] To further highlight the beneficial effects of this application, specific implementation methods are described below.

[0089] Figure 3 The diagram shows a computational gene partitioning diagram used in calculating transcriptional activity data. Figure 4 The diagram shows a schematic of the salience calculation process for pause sites.

[0090] like Figure 3 As shown, genes are divided into the proximal region of the transcription start site, the first 25% of the non-proximal region, and the last 75% of the non-proximal region, primarily based on our understanding of the transcription process and its regulatory mechanisms. The principle behind this division is as follows:

[0091] Proximal region of transcription start site: This region is adjacent to the transcription start site and includes hundreds of downstream nucleotides. RNA polymerase II may temporarily stop and accumulate in this region. Observing the signal data in this region can help infer the transcription efficiency.

[0092] The first 25% of the non-proximal region: This region is far from the transcription start site, but is still located in the upstream region of the gene. When RNA polymerase passes through this region, it means that the RNA has broken through the transcription pause and entered the elongation phase.

[0093] The posterior 75% non-proximal region: the gene region from the middle to the rear, is used to determine whether the gene has extended completely.

[0094] pass Figure 3 The formulas for calculating gene partitioning and transcriptional activity were used to calculate cell-line specific transcriptional activity data in 59 human cell lines, 31 mouse cell lines, and 6 Drosophila cell lines, respectively.

[0095] The following is the formula for calculating transcriptional activity:

[0096] Pause Index (PI) = Proximal promoter density / Genome density;

[0097] Elongation Index (EI) = First 25% of genomes / Last 75% of genomes;

[0098] Extension speed (ES) = log10(-1 / tilt value); where the tilt value is based on Figure 4 The slope value is determined by calculating the gene partition.

[0099] like Figure 4 The diagram illustrates the statistical calculation process for the significance of pausing sites, which can be used to calculate and analyze datasets involving N samples, where N is the total number of samples collected. A 4x4 contingency table is constructed. The rows of the contingency table represent the number of samples where a pausing site was detected. The columns represent the number of samples where a pausing site was detected within a 200bp segment. A chi-square test is performed on the contingency table to calculate the p-value. Benjamini-Hochberg (BH) correction is used to control the false discovery rate of multiple tests. The p-value represents the degree of difference between the observed pausing site and the expected theoretical data. If the p-value is less than the significance level (usually set to 0.05), the current pausing site is considered significant, and significance further reflects the confidence of the identified site.

[0100] Table 1 below summarizes the extensive and annotated transcriptional pause site information contained in the TransPause database, covering three species: human, mouse, and fruit fly. The database also calculates and provides transcriptional kinetic indicators such as the pause index, elongation index, and elongation rate to quantify and describe transcriptional regulation processes in different species.

[0101]

[0102]

[0103] Table 1: Correspondence Figure 3 Transcription pause site information in the annotation

[0104] Figures 5-8 The diagram illustrates the process of using the Transpause database. It provides a use case for analyzing the transcriptional pausing characteristics of ACLY, a target gene of the NELF transcriptional pausing factor, using the various functional modules of the Transpause database.

[0105] 1) First, retrieve the list of target genes of NELF in human cell lines through the "target gene search" module, and select the ACLY gene of interest from it.

[0106] 2) Then, detailed transcriptional pause information of ACLY was queried, including gene ID, transcriptional activity index, pause site and pause spectrum, etc. It was found that the proximal pause site of ACLY has a high GC content and special secondary structure.

[0107] like Figure 5 The image shows the query results for transcriptional pausing features of the ACLY gene. ACLY is a known gene that experiences transcriptional pausing. The query yields detailed transcriptional pausing information for ACLY, including gene ID, transcriptional activity index, pausing sites, and pausing patterns. It was found that the proximal pausing sites in ACLY exhibit high GC content and unique secondary structures.

[0108] 3) Next, using the "Cell Line Transcription Activity Changes" tool, the distribution of ACLY pausing index in various cell lines was plotted, and significant differences were found in the ACLY pausing index between the two cell lines A549 and BEA5-2B.

[0109] To understand the distribution of ACLY pause index across various cell lines, the "Cell Line Transcriptional Activity Changes" tool in the Transpause database was used to plot data as follows: Figure 6 The bar chart shown reveals differences in the ACLY pausing index between the two cell lines, A549 and BEA5-2B.

[0110] Furthermore, using Transpause's "Difference Analysis" function, we can calculate whether the difference is significant. The difference calculation results are as follows: Figure 7 As shown, the calculation results indeed show a significant difference.

[0111] 4) Then, by using the "Find Neighboring Genes" function, a batch of genes with similar pause regulation patterns to ACLY were found.

[0112] Finally, these gene lists were imported into the "enrichment analysis" module for GO biological process enrichment analysis. The results showed that these genes were significantly enriched in pathways related to DNA damage, DNA unwinding and other processes, which is consistent with existing research results.

[0113] To identify genes with similar pause regulation patterns to ACLY, the "Find Neighboring Genes" function in the Transpause database was used to search for a batch of genes with similar pause regulation patterns to ACLY. This gene list was then imported into the "Enrichment Analysis" module for GO biological process enrichment analysis. The results are as follows: Figure 8 As shown, these genes are significantly enriched in pathways related to DNA damage, DNA unwinding, and other processes, which is consistent with existing research findings.

[0114] Through the above series of analyses, researchers were able to fully interpret the transcriptional pausing characteristics of the ACLY gene and its differences in different cell lines, and predict the biological processes it may be involved in, demonstrating the powerful data mining and analysis capabilities of the TransPause database.

[0115] In summary, by employing the cell-specific transcriptional pause site analysis method of this application, and through comprehensive database analysis and data processing, we can gain a deeper understanding of the diversity and functional significance of pause phenomena under various cell types and physiological conditions.

[0116] Meanwhile, this application also has the following technical effects: it can provide a comprehensive understanding of the diversity and distribution characteristics of transcriptional pausing under different cell types and physiological conditions; it can provide in-depth analysis of the relationship between transcriptional pausing and gene expression regulation and its potential functional significance; it can explore the role of transcriptional pausing in cellular processes, such as differentiation and development; and it can provide new biomarkers and targets for the diagnosis and treatment of related diseases.

[0117] Example 2

[0118] This embodiment provides an analysis system for cell-specific transcriptional pausing sites. For details not disclosed in this embodiment's analysis system for cell-specific transcriptional pausing sites, please refer to the specific implementation details of the analysis methods for cell-specific transcriptional pausing sites in other embodiments.

[0119] Figure 9 The diagram shows a schematic of the structure of a cell-specific transcriptional pause site analysis system according to an embodiment of this application.

[0120] like Figure 9 As shown, the analysis system for cell-specific transcriptional pausing sites includes:

[0121] Database module 10 is used to collect N nascent RNA sequencing data from different cell lines of three species: human, mouse, and fruit fly; and to collect ChIP-seq peak data of regulatory factors related to transcriptional pausing.

[0122] The transcriptional activity analysis module 20 is used to determine the pause index (PI) of transcriptional activity based on the transcriptional pause level reflected by the nascent RNA sequencing data; to determine the elongation index (EI) of transcriptional activity based on the transcriptional elongation level reflected by N nascent RNA sequencing data; and to determine the elongation rate (ES) of transcriptional activity based on the RNA polymerase elongation rate data reflected by N nascent RNA sequencing data.

[0123] Pause site annotation module 30: used to identify significant transcription pause sites based on newborn RNA sequencing data and to annotate known information about the transcription pause sites, including: neighboring pause sites near the transcription start site and global pause sites across the genome;

[0124] Differential analysis module 40: used to determine the sequence and structural features of transcriptional pausing sites; and to obtain analysis of transcriptional pausing regulatory factors, differential analysis of transcriptional activity, and functional genomics analysis based on transcriptional pausing sites.

[0125] In summary, by using the cell-specific transcriptional pause site analysis system of this application, and through comprehensive database analysis and data processing, we can gain a deeper understanding of the diversity and functional significance of pause phenomena under various cell types and physiological conditions.

[0126] Meanwhile, this application also has the following technical effects: it can provide a comprehensive understanding of the diversity and distribution characteristics of transcriptional pausing under different cell types and physiological conditions; it can provide in-depth analysis of the relationship between transcriptional pausing and gene expression regulation and its potential functional significance; it can explore the role of transcriptional pausing in cellular processes, such as differentiation and development; and it can provide new biomarkers and targets for the diagnosis and treatment of related diseases.

[0127] Example 3

[0128] This embodiment provides a memory. For details not disclosed in the memory of this embodiment, please refer to the specific implementation of the analysis method or system for cell-specific transcriptional pause sites in other embodiments.

[0129] Figure 10 The diagram shows a schematic representation of the structure of a memory 400 or an image recognition device according to an embodiment of this application.

[0130] like Figure 10 As shown, the memory 400 includes: a storage unit 402 for storing executable instructions; and a processing unit 401 for connecting to the storage unit 402 to execute the executable instructions to complete a method for analyzing cell-specific transcriptional pause sites.

[0131] Those skilled in the art will understand that the illustration Figure 10 This is merely an example of memory 400 or an image recognition device and does not constitute a limitation on memory 400 or an image recognition device. It may include more or fewer components than shown, or combine certain components, or different components. For example, memory 400 may also include input / output devices, network access devices, buses, etc.

[0132] The processing unit 401 (Central Processing Unit, CPU) can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or the processing unit 401 can be any conventional processor. The processing unit 401 is the control center of the memory 400, connecting various parts of the memory 400 through various interfaces and lines.

[0133] Storage unit 402 can be used to store computer-readable instructions. Processing unit 401 implements various functions of memory 400 by running or executing computer-readable instructions or modules stored in storage unit 402 and calling data stored in storage unit 402. Storage unit 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of memory 400, etc. In addition, storage unit 402 may include hard disk, memory, plug-in hard disk, smart memory card (SMC), secure digital card (SD) card, flash memory card, at least one disk storage device, flash memory device, read-only memory (ROM), random access memory (RAM), or other non-volatile / volatile storage devices.

[0134] If the module integrated in memory 400 is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by instructing related hardware through computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium, and when executed by a processor, the computer-readable instructions can implement the steps of the various method embodiments described above.

[0135] Example 5

[0136] This embodiment provides a computer-readable storage medium having a computer program stored thereon; the computer program is executed by a processor to implement the cell-specific transcriptional pause site analysis method in other embodiments.

[0137] The memory and storage medium of the embodiments of this application, through comprehensive database and comprehensive data processing and analysis, can provide a deeper understanding of the diversity and functional significance of pause phenomena under various cell types and physiological conditions.

[0138] Meanwhile, this application also has the following technical effects: it can provide a comprehensive understanding of the diversity and distribution characteristics of transcriptional pausing under different cell types and physiological conditions; it can provide in-depth analysis of the relationship between transcriptional pausing and gene expression regulation and its potential functional significance; it can explore the role of transcriptional pausing in cellular processes, such as differentiation and development; and it can provide new biomarkers and targets for the diagnosis and treatment of related diseases.

[0139] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0140] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0141] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0142] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0143] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The singular forms “a,” “the,” and “the” used in this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0144] It should be understood that although the terms first, second, third, etc., may be used in this invention to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first information may also be referred to as second information without departing from the scope of this invention, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0145] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0146] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for analyzing cell-specific transcriptional pausing sites, characterized in that, include: We collected N sequencing data of newly generated RNA from different cell lines of three species: human, mouse, and fruit fly. ChIP-seq peak data of regulatory factors associated with transcriptional arrest were collected; Based on the transcriptional pausing level reflected by the newly generated RNA sequencing data, the pausing index of transcriptional activity is determined; based on the transcriptional elongation level reflected by the N newly generated RNA sequencing data, the elongation index of transcriptional activity is determined. Based on the RNA polymerase elongation rate data reflected by the N newly generated RNA sequencing data, the elongation rate of transcriptional activity is determined. Significant transcriptional pause sites were identified based on the newly generated RNA sequencing data, and known information about these pause sites was labeled. This known information includes: neighboring pause sites near the transcription start site and global pause sites across the genome. The sequence and structural features of the transcriptional pausing sites were determined; and based on the transcriptional pausing sites, analyses of transcriptional pausing regulatory factors, differential transcriptional activity, and functional genomics were obtained.

2. The analysis method according to claim 1, characterized in that, The analysis of transcriptional pausing regulatory factors based on the transcriptional pausing sites includes: Obtain the distribution of known transcriptional pausing motifs on genes; The distribution was compared with the ChIP-seq peak data of the regulatory factors; Predicting and annotating the regulatory factors that control each pause site provides clues for subsequent mechanistic studies.

3. The analysis method according to claim 2, characterized in that, The prediction and annotation modulate the regulatory factors for each pause site, including: The regulatory factors for transcriptional arrest include the arrest factor NELF and the DRB sensitivity inducing factor DSIF. The .bed files of the regulatory factors were obtained from the ENCODE database. The intersect command of the bedtools tool was used to determine the intersection of the NELF and DSIF peak locations with transcriptional pause sites; Based on the intersection, the peak value is determined to overlap with known transcriptional pause sites, and the regulatory factor for each pause site is predicted.

4. The analysis method of claim 1, wherein, Based on the transcriptional pausing sites, differential analysis of transcriptional activity was obtained, including: The differences in transcriptional activity of genes among different cell lines were analyzed using statistical methods to identify genes that regulate transcriptional arrest under different cell states. The statistical methods included t-test, Mann-Whitney U test, and Wilcoxon rank-sum test.

5. The analysis method according to claim 4, characterized in that, The genes identified as being regulated by transcriptional pauses in different cellular states include: Use Python's stats module to determine whether the transcriptional activity data of any two cell lines in multiple cell lines are normally distributed; If both groups are normally distributed, use the parametric T-test for difference analysis; if either or both groups deviate from a normal distribution, use the non-parametric Mann-Whitney U test to assess the difference between the two groups. Genes with differential transcriptional activity were screened using a p-value less than 0.05 as the significance level.

6. The analysis method of claim 1, wherein, Functional genomics analysis was obtained based on the transcriptional pause sites, including: By identifying genes with similar transcriptional activities, we can predict potential transcriptional or functional associations. Gene enrichment analysis was used to identify biological processes or pathways associated with transcriptional arrest. Gene set enrichment analysis was used to compare gene sets regulated by transcriptional pausing under different cellular states.

7. The analysis method according to claim 6, characterized in that, The process of finding genes with similar transcriptional activity includes using transcriptional activity data between samples, calculating the Euclidean distance between each gene and other genes, identifying the top 10 nearest neighbor genes of each gene, which may have transcriptional or functional associations, and calculating the Pearson correlation coefficient and Spearman correlation coefficient of each gene with its nearest neighbors and their significance, thereby quantitatively describing their relationship. The gene enrichment analysis includes enrichment analysis based on transcriptional activity data and a list of differentially expressed genes. The biological processes or pathways enriched by the list of differentially expressed genes may be related to transcriptional pausing. Using the Benjamini-Hochberg method to perform multiple hypotheses, a p-value less than 0.05 after correction is considered significant enrichment.

8. A system for analyzing cell-specific transcriptional pause sites, characterized in that, include: The database module is used to collect N sets of nascent RNA sequencing data from different cell lines of three species: human, mouse, and fruit fly. ChIP-seq peak data of regulatory factors associated with transcriptional arrest were collected; The transcriptional activity analysis module is used to determine the pause index (PI) of transcriptional activity based on the transcriptional pause level reflected by the nascent RNA sequencing data; to determine the elongation index (EI) of transcriptional activity based on the transcriptional elongation level reflected by the N nascent RNA sequencing data; and to determine the elongation rate (ES) of transcriptional activity based on the RNA polymerase elongation rate data reflected by the N nascent RNA sequencing data. Pause site annotation module: used to identify significant transcription pause sites based on the newborn RNA sequencing data, and to annotate the known information of the transcription pause sites, including: neighboring pause sites near the transcription start site and global pause sites across the genome; The differential analysis module is used to determine the sequence and structural features of the transcriptional pausing sites; and to obtain transcriptional pausing regulatory factor analysis, transcriptional activity differential analysis, and functional genomics analysis based on the transcriptional pausing sites.

9. A memory, characterized in that, include: Storage unit, used to store executable instructions; as well as A processing unit is configured to be connected to a memory to execute executable instructions to perform the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, It stores a computer program thereon; the computer program is executed by a processor to implement the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Method for characterizing cell-free nucleic acid fragments

    CN116018646A

  • Cross-species analysis method, equipment and medium for accessibility map of single-cell chromatin

    CN117542414A