A system for predicting transcription factor target genes based on tumor transcriptomic features

By using transcriptomic data at the level of pan-cancer and position weight matrix (PWM) scoring, the problem of complex and costly prediction of transcription factor target genes in the prior art is solved, and efficient and accurate prediction of transcription factor target genes is achieved.

CN114708914BActive Publication Date: 2025-07-29JINSHAN HOSPITAL AFFILIATED TO FUDAN UNIV (EYE DISEASE PREVENTION & TREATMENT CENT OF JINSHAN DISTRICT RES CENT FOR CHEM INJURY EMERGENCY & CRITICAL MEDICINE OF SHANGHAI MUNICIPAL HEALTH COMMISSION)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210213180.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-04
Publication Date
2025-07-29
Estimated Expiration
2042-03-04

AI Technical Summary

Technical Problem

In the prior art, when predicting transcription factor target genes, there are problems such as complex operation, high cost and unstable prediction effects. In particular, the analysis methods based on ChIP-seq experiments and protein expression data have high usage thresholds and sample size limitations.

Method used

Using transcriptomic data at the pan-cancer level and combined with position weight matrix (PWM) scoring, a pan-cancer-wide transcriptomic correlation calculation module, promoter region sequence acquisition module and transcription factor target gene screening module are used to achieve simple and efficient transcription factor target gene prediction.

Benefits of technology

It lowers the experimental workload, technical and funding thresholds, improves the accuracy and efficiency of target gene prediction of transcription factor, and simplifies the operation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114708914B_ABST
    Figure CN114708914B_ABST
Patent Text Reader

Abstract

The present invention relates to a system for predicting transcription factor target genes based on tumor transcriptomic features, which consists of four major modules: a transcriptomic correlation calculation module at the pan-cancer level; a promoter region sequence acquisition module; a position weight matrix scoring module; and a transcription factor target gene screening module. Its advantages are as follows: It innovatively uses transcriptomic data at the pan-cancer level and utilizes the extensive and massive gene expression regulation network perturbations therein to achieve the purpose of simple and efficient prediction of transcription factor target genes in terms of operation; this system can obtain relatively accurate prediction results of transcription factor target genes by introducing position weight matrix (PWM) scoring; all the data and materials used in this system can theoretically be obtained through open public databases, which further reduces the experimental workload, technical and funding thresholds for researchers when using this system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biotechnology, and more particularly, to a system for predicting transcription factor target genes based on tumor transcriptomic features. Background Art

[0002] Transcription factors are DNA-binding proteins that play a key role in many pathophysiological processes, including tumorigenesis. With the development of high-throughput research technologies, an increasing number of differentially expressed transcription factors under different pathophysiological conditions have been identified, but exploring the specific functions of these transcription factors remains a challenging issue.

[0003] It is generally believed that transcription factors can bind to specific regions (i.e., DNA-binding elements) in the promoter region of downstream target genes, regulate the transcription of target genes, thereby causing changes in the expression of these target genes, and further leading to changes in cell biological functions. Therefore, there are mainly two approaches to predicting transcription factor target genes, namely, analysis based on transcription factor genomic binding propensity data and analysis based on protein expression data. While these methods have achieved good prediction results, there are also some difficult-to-solve problems. For the analysis based on genomic binding propensity data, the analysis process relies on data from the operationally complex and expensive ChIP-seq (Chromatin immunoprecipitation followed by sequencing) experiment. On the one hand, this leads to a relatively high threshold for using this analysis method, and on the other hand, the prediction effect of this method is also restricted by the inherent defects of ChIP-seq and the reliability is not high. For the analysis based on protein expression data, they often rely on adding specific perturbations to the gene expression regulatory network of cells and inferring transcription factor target genes from a functional perspective based on the corresponding changes in gene expression profiles. Therefore, the experimental workload is large and the prediction effect is restricted by the sample size.

[0004] Chinese Patent Document: CN201010227078.9, with an application date of July 14, 2010, and a patent name of: A method for predicting transcription factor binding sites. Disclosed is a method for predicting transcription factor binding sites. The basic implementation process of this method is as follows: Step 1, genomic localization of the research object; Step 2, extraction of gene promoter sequences; Step 3, prediction of transcription factor binding sites; Step 4, statistical analysis of prediction results. Through this method, the transcription factor binding sites of target genes can be predicted more accurately, effectively improving the true positive rate of prediction results.

[0005] Chinese Patent Document: CN201910922590.6, application date: September 26, 2019, patent title: A novel tumor-related transcription factor ZSCAN16 and its application in inhibiting tumors, application of ZSCAN16 gene as a therapeutic target for bladder cancer. The advantages of this invention are as follows: This invention reveals that ZSCAN16 may play an important biological role in the development of bladder cancer, and it is also closely related to the development of tumors.

[0006] A method for predicting transcription factor binding sites in the above-mentioned patent document CN201010227078.9 can accurately predict the transcription factor binding sites of the target gene through four steps: genomic localization of the research object; extraction of gene promoter sequences; prediction of transcription factor binding sites; and statistical analysis of prediction results, effectively improving the true positive rate of the prediction results. For a novel tumor-related transcription factor ZSCAN16 and its application in inhibiting tumors in the patent document CN201910922590.6, the results show that ZSCAN16 is a key new oncogene in the occurrence and development of bladder cancer. In vitro experimental results show that the silencing of ZSCAN16 inhibits the proliferation, colony formation, apoptosis, migration and invasion of T24 cells. ZSCAN16 can be used as a research target for bladder cancer regulation, a diagnostic and prognostic evaluation marker for tumors, and a target for developing drugs to inhibit tumors. However, there is currently no relevant report on a system for predicting transcription factor target genes based on tumor transcriptomic characteristics that innovatively uses pan-cancer level transcriptomic data to achieve simple and efficient prediction of transcription factor target genes in terms of operation, and obtains relatively accurate prediction results of transcription factor target genes by introducing position weight matrix (PWM) scoring, further reducing the experimental workload, technical and funding thresholds for researchers when using this system.

[0007] In summary, there is an urgent need for a system for predicting transcription factor target genes based on tumor transcriptomic characteristics that innovatively uses pan-cancer level transcriptomic data to achieve simple and efficient prediction of transcription factor target genes in terms of operation, and obtains relatively accurate prediction results of transcription factor target genes by introducing position weight matrix (PWM) scoring, further reducing the experimental workload, technical and funding thresholds for researchers when using this system. Summary of the Invention

[0008] The object of the present invention is to overcome the deficiencies of the prior art and provide a system for predicting transcription factor target genes based on tumor transcriptomic features. This system innovatively uses transcriptomic data at the pan-cancer level to achieve the purpose of simply and efficiently predicting transcription factor target genes in terms of operation. By introducing the scoring of position weight matrix (PWM), it obtains relatively accurate prediction results of transcription factor target genes, further reducing the experimental workload, technical and funding thresholds for researchers when using this system.

[0009] To achieve the above object, the technical solution adopted by the present invention is:

[0010] A system for predicting transcription factor target genes based on tumor transcriptomic features, which consists of 4 major modules:

[0011] A transcriptomic correlation calculation module for the pan-cancer scope;

[0012] A promoter region sequence acquisition module;

[0013] A position weight matrix scoring module;

[0014] A transcription factor target gene screening module.

[0015] As a preferred technical solution, the transcriptomic correlation calculation module for the pan-cancer scope: is used to calculate and determine the transcriptomic correlation parameters between a given transcription factor gene and all other genes; for a given transcription factor gene x, calculate its Pearson correlation coefficients with the remaining genes (y1, y2,..., yi) in the transcriptomic data of all cancer types respectively; for any pair of x and yi, take the maximum value of its Pearson correlation coefficients in all cancer types as the transcriptomic correlation parameter of this transcription factor x to this gene yi.

[0016] As a preferred technical solution, the promoter region sequence acquisition module: obtains the promoter region sequence of each gene from the human reference genome according to the gene transcription start site position information in the public database.

[0017] As a preferred technical solution, the transcription factor target gene screening module: according to the transcriptomic correlation parameters and the position weight matrix scores of a given transcription factor to all genes, combined with the user's requirements, sets the thresholds of the transcriptomic correlation parameters and the position weight matrix scores, so as to screen out the predicted target genes of this transcription factor.

[0018] The advantages of the present invention are:

[0019] 1. It innovatively uses transcriptomic data at the pan-cancer level and utilizes the extensive and large-scale gene expression regulation network perturbations therein to achieve the purpose of simply and efficiently predicting transcription factor target genes in terms of operation.

[0020] 2. By introducing the position weight matrix (PWM) scoring, this system can obtain relatively accurate prediction results of transcription factor target genes.

[0021] 3. In theory, all the data used in this system can be obtained through open public databases, which further reduces the experimental workload, technical and financial thresholds for researchers when using this system. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Attached Figure 1 is a schematic diagram comparing the prediction effects of this system with other means of predicting transcription factor target genes.

[0023] Attached Figure 2 is a schematic diagram for evaluating the prediction effect of using this system to predict target genes with the transcription factor CREB1 as an example.

[0024] Attached Figure 3 is the program flow chart of the system of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0025] A system for predicting transcription factor target genes based on tumor transcriptomic features of the present invention consists of 4 major modules:

[0026] 1) Transcriptomic correlation calculation module for pan-cancer scope: used to calculate and determine the transcriptomic correlation parameters between a given transcription factor gene and all other genes. For a given transcription factor gene x, calculate its Pearson correlation coefficients with the remaining genes (y1, y2, ……, yi) in the transcriptomic data of all cancer types respectively. For any pair of x and yi, take the maximum value of their Pearson correlation coefficients in all cancer types as the transcriptomic correlation parameter of this transcription factor x to this gene yi.

[0027] 2) Promoter region sequence acquisition module: According to the gene transcription start site position information in the public database, obtain the promoter region sequence of each gene from the human reference genome.

[0028] 3) Position weight matrix (PWM) scoring module: Obtain the position weight matrix of a certain transcription factor from the public database, and then use this matrix to score the promoter regions of all genes.

[0029] 4) Transcription factor target gene screening module: According to the transcriptomic correlation parameters of a given transcription factor to all genes and the scores of the position weight matrix, combined with the needs of the user, set the thresholds of the transcriptomic correlation parameters and the scores of the position weight matrix, so as to screen out the predicted target genes of this transcription factor.

[0030] Specifically, please refer to the appendix Figure 1 , Figure 1 which is a schematic diagram comparing the prediction effects of this system (i.e., TFoTF, Target Finder of Transcription Factor) with other prediction methods for transcription factor target genes (ChIP-seq experiment, PWM scoring). It can be seen that the prediction effect of this system is better.

[0031] Among them, Figure 1 the left area graph in shows the effects of three different prediction methods in searching for target genes of transcription factors (taking STAT1 as an example). The horizontal axis is the number of verified genes included in the top 500 genes predicted by these three methods; the vertical axis is the number of literatures retrieved for each gene. The right scatter plot reflects the differences in the ability of these three methods to predict target genes. The Kolmogorov-Smirnov test was used for comparison between the two groups. *: P<0.05; ****: P<0.0001.

[0032] Please refer to the appendix Figure 2 , Figure 2 which is a schematic diagram for evaluating the prediction effect of using this system to predict target genes with the transcription factor CREB1 as an example. By detecting the change in the expression level of the predicted target genes after knocking down CREB1, the accuracy of the prediction results was verified.

[0033] Among them, Figure 2 the expression changes of the predicted CREB1 target genes (PDS5B, THUMPD1, CNOT6, MAP4K3, SF3B1, CCP110, RBBP6, DDX46, DHX15, YLPM1) after knocking out the transcription factor CREB1; n = 3 independent experiments; ns: no significance; **: P<0.01; ***: P<0.001; ****: P<0.0001.

[0034] Please refer to the appendix Figure 3 , Figure 3 which is the program flow chart of the system of the present invention.

[0035] Among them, Figure 3 Based on the two premise backgrounds of "there are gene expression profile perturbations in tumor tissues" and "transcription factors often have specific binding sites", we designed corresponding calculation modules and achieved the purpose of predicting the target genes of a given transcription factor through mutual cooperation.

[0036] A system for predicting transcription factor target genes based on tumor transcriptomic features of the present invention innovatively uses transcriptomic data at the pan-cancer level and utilizes extensive gene expression regulatory network perturbations therein to achieve the purpose of simple and efficient prediction of transcription factor target genes in operation; this system can obtain relatively accurate prediction results of transcription factor target genes by introducing position weight matrix (PWM) scoring; all data and materials used in this system can theoretically be obtained through open public databases, which further reduces the experimental workload, technical and funding thresholds for researchers when using this system.

[0037] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and supplements can be made, and these improvements and supplements should also be regarded as the protection scope of the present invention.

Claims

1. A system for predicting transcription factor target genes based on tumor transcriptomic features, characterized in that It consists of 4 major modules: 1) Transcriptomic correlation calculation module for pan-cancer scope: used to calculate and determine the transcriptomic correlation parameters between a given transcription factor gene and all other genes; for a given transcription factor gene x, calculate the Pearson correlation coefficients between it and the remaining genes (y1, y2, ……, yi) in the transcriptomic data of all cancer types; for any pair of x and yi, take the maximum value of the Pearson correlation coefficients in all cancer types as the transcriptomic correlation parameter of the transcription factor x for the gene yi; 2) Promoter region sequence acquisition module: obtain the promoter region sequence of each gene from the human reference genome according to the gene transcription start site position information in the public database; 3) Position weight matrix scoring module: obtain the position weight matrix of a certain transcription factor from the public database, and then use this matrix to score the promoter regions of all genes; 4) Transcription factor target gene screening module: according to the transcriptomic correlation parameters and position weight matrix scores of a given transcription factor for all genes, combined with the user's needs, set the threshold for transcriptomic correlation parameters and the threshold for position weight matrix scores, so as to screen out the predicted target genes of this transcription factor.

Citation Information

Patent Citations

  • Method for prediction of transcription factor binding site (TFBS)

    CN102206699A

  • Novel tumor-related transcription factor ZSCAN16 and application thereof in tumor inhibiting

    CN112553330A