Transcription factor analysis method and device, electronic equipment and storage medium

Through an automated transcription factor analysis method, the problems of confusion in the transcription factor database and lack of automated analysis processes in the prior art are solved, and efficient analysis of non-modal plant transcription factor regulation networks are achieved, and data quality and analysis efficiency are improved.

CN120072049AActive Publication Date: 2025-05-30SHENZHEN HUADA GENE INST +1

Patent Information

Application Number
CN202311612037.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-05-30
Estimated Expiration
2043-11-28

AI Technical Summary

Technical Problem

The motif matrix file format and annotation in the existing transcription factor database are confusing, and the lack of automated processes for transcription factor regulation network analysis suitable for non-modal plants, resulting in low analysis efficiency.

Method used

A transcription factor analysis method is proposed, including obtaining the protein sequence and target transcription factor of the species to be analyzed, determining the target motif through homology analysis, co-expression analysis to obtain the co-expression module, screening regulators based on motifs, and obtaining the analysis results by calculating the regulator activity score.

Benefits of technology

Through the automated transcription factor regulation network analysis method, the efficiency of non-modal plant transcription factor regulation network analysis is improved, and data quality and analysis accuracy are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072049A_ABST
    Figure CN120072049A_ABST
Patent Text Reader

Abstract

The invention provides a transcription factor analysis method and device, electronic equipment and a storage medium. The transcription factor analysis method comprises the following steps: acquiring a protein sequence of a to-be-analyzed species and a plurality of target transcription factors of the protein sequence; performing homology analysis on the plurality of target transcription factors and a pre-stored template transcription factor to determine a target motif corresponding to each target transcription factor; performing co-expression analysis on each target transcription factor and the corresponding target motif to obtain a co-expression module of each target transcription factor; screening the co-expression module based on the target motif to obtain a regulator corresponding to each target transcription factor; and calculating the activity score of the regulator to obtain a transcription factor analysis result of the species to be analyzed. The efficiency of transcription factor analysis can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of bioinformatics analysis technology, specifically to the field of transcription factor analysis technology, and particularly to a transcription factor analysis method, device, electronic device, and storage medium. Background Art

[0002] Transcription factors play important roles in biological processes such as the growth and development, environmental response, and growth cycle regulation of species. Analyzing the regulatory network of transcription factors in a species can provide important data support and guidance for gene function research, disease occurrence and development mechanism research, drug development, etc. Currently, the formats and annotations of motif matrix files in existing transcription factor databases are relatively chaotic, and there is no automated process suitable for analyzing the transcription factor regulatory network of non-model plants, resulting in low efficiency in analyzing the transcription factor regulatory network of the species to be analyzed. Summary of the Invention

[0003] In view of the above, it is necessary to propose a transcription factor analysis method, device, electronic device, and storage medium to solve the technical problem of low efficiency in analyzing the transcription factor regulatory network.

[0004] This application provides a transcription factor analysis method, and the method includes: obtaining the protein sequence of the species to be analyzed and multiple target transcription factors of the protein sequence; determining the target motif corresponding to each target transcription factor by performing homology analysis on the multiple target transcription factors and pre-stored template transcription factors; performing co-expression analysis on each target transcription factor and the corresponding target motif to obtain the co-expression module of each target transcription factor; screening the co-expression module based on the target motif to obtain the regulator corresponding to each target transcription factor; and obtaining the transcription factor analysis result of the species to be analyzed by calculating the activity score of the regulator.

[0005] In some embodiments, the method for obtaining the pre-stored template transcription factors includes: obtaining the example transcription factors and the corresponding example motifs of multiple example species from multiple databases; when multiple example transcription factors belong to the same example species and different databases, retaining any one of the multiple example transcription factors and its corresponding example motif; and converting the example transcription factors and example motifs into a unified format to obtain the template transcription factors and the corresponding template motifs.

[0006] In some embodiments, the template transcription factor corresponds to a template motif. Determining the target motif corresponding to each target transcription factor by performing a homology analysis on the multiple target transcription factors and the pre-stored template transcription factors includes: calculating the homology score of each target transcription factor and each template transcription factor; determining the motif of the template transcription factor corresponding to the highest homology score as the target motif.

[0007] In some embodiments, calculating the homology score of each target transcription factor and each template transcription factor includes: dividing the target transcription factor to obtain a plurality of first subsequences corresponding to the target transcription factor; dividing the template transcription factor to obtain a plurality of second subsequences corresponding to the template transcription factor; when any one of the first subsequences is the same as any one of the second subsequences, determining the length of the first subsequence as the candidate homology score; determining the maximum value among the candidate homology scores as the homology score of the target transcription factor and the template transcription factor.

[0008] In some embodiments, the co-expression module includes the target genes of the target transcription factor. Screening the co-expression module based on the target motif to obtain the regulator corresponding to each target transcription factor includes: calculating the enrichment score of the target transcription factor in the corresponding target gene; determining the binding site of the target transcription factor and the corresponding target gene; when the target motif corresponding to the target transcription factor in the co-expression module is upstream of the binding site of the target gene and the enrichment score is greater than a preset threshold, determining the co-expression module as the regulator.

[0009] In some embodiments, calculating the enrichment score of the target transcription factor in the corresponding target gene includes: calculating the first frequency of the target motif of the target transcription factor appearing in the target gene; calculating the second frequency of the target motif of the target transcription factor appearing in the protein sequence; calculating the ratio of the first frequency to the second frequency to obtain the enrichment score of the target transcription factor in the corresponding target gene.

[0010] In some embodiments, obtaining the transcription factor analysis result of the species to be analyzed by calculating the activity score of the regulator includes: determining the target gene in the regulator that contains the binding site of the target transcription factor; calculating the AUC score of the regulator according to the enrichment score to obtain the activity score of the regulator.

[0011] An embodiment of the present application further provides a transcription factor analysis device, which includes: a receiving module for receiving the protein sequence of a species to be analyzed; an analysis module for predicting a plurality of target transcription factors of the protein sequence; the analysis module is further configured to determine the target motif corresponding to each target transcription factor by performing homology analysis on the plurality of target transcription factors and pre-stored template transcription factors; the analysis module is further configured to perform co-expression analysis on each target transcription factor and the corresponding target motif to obtain the co-expression module of each target transcription factor; the analysis module is further configured to screen the co-expression module based on the target motif to obtain a regulator corresponding to each target transcription factor; the analysis module is further configured to obtain the transcription factor analysis result of the species to be analyzed by calculating the activity score of the regulator.

[0012] An embodiment of the present application further provides an electronic device, which includes:

[0013] A memory storing at least one instruction;

[0014] A processor for executing the instructions stored in the memory to implement the transcription factor analysis method described above.

[0015] An embodiment of the present application further provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is executed by a processor in an electronic device to implement the transcription factor analysis method.

[0016] It can be seen from the above technical solutions that in the embodiment of the present application, by performing homology analysis on the target transcription factors of the species to be analyzed and the pre-stored template transcription factors, the motif corresponding to the target transcription factor is determined, providing data support for subsequent analysis of the transcription factor regulatory network of the species to be analyzed. Then, co-expression analysis is performed on the transcription factors and the corresponding motifs to obtain the co-expression module of the transcription factors, and the co-expression module is screened based on the motif to obtain the regulator corresponding to each target transcription factor, which can improve the data quality of the target transcription factor and thus improve the efficiency of transcription factor regulatory network analysis. By calculating the activity score of the regulator, the transcription factor analysis result of the species to be analyzed is obtained, which can realize the automatic analysis of the transcription factor regulatory network of the species to be analyzed, thereby improving the efficiency of transcription factor regulatory network analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is an application scenario diagram of the transcription factor analysis method provided by an embodiment of the present application.

[0018] Figure 2 is a flowchart of the transcription factor analysis method provided by an embodiment of the present application.

[0019] Figure 3 It is a schematic diagram of a transcription factor motif association list provided by an embodiment of the present application.

[0020] Figure 4 It is a heat map of the activity scores of regulators provided by an embodiment of the present application.

[0021] Figure 5 It is a heat map of the activity scores of regulators in different cells provided by an embodiment of the present application.

[0022] Figure 6 It is a flowchart of a method for determining a pre-stored template transcription factor provided by an embodiment of the present application.

[0023] Figure 7 It is a flowchart of a method for determining a target motif corresponding to each target transcription factor provided by an embodiment of the present application.

[0024] Figure 8 It is a flowchart of a method for determining a regulator of a target transcription factor provided by an embodiment of the present application.

[0025] Figure 9 It is a flowchart of a method for calculating an enrichment score of a target transcription factor provided by an embodiment of the present application.

[0026] Figure 10 It is a functional module diagram of a transcription factor analysis device provided by an embodiment of the present application.

[0027] Figure 11 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0028] In order to more clearly understand the purpose, features and advantages of the present application, the present application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other. Many specific details are set forth in the following description in order to fully understand the present application. The described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.

[0029] In addition, the terms "first" and "second" are only used for descriptive purposes, and cannot be construed as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present application, "a plurality" means two or more, unless otherwise specifically defined.

[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0031] An embodiment of this application provides a transcription factor analysis method, which can be applied to one or more electronic devices. An electronic device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.

[0032] An electronic device can be any electronic product that can interact with a customer, such as a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an Internet Protocol Television (IPTV), a smart wearable device, etc.

[0033] The electronic device may also include a network device and / or a client device. Among them, the network device includes, but is not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing.

[0034] The network where the electronic device is located includes, but is not limited to, the Internet, a wide area network, a metropolitan area network, a local area network, a virtual private network (VPN), etc.

[0035] As Figure 1 shown, the transcription factor analysis method provided by this application can be applied to the electronic device 100, and the electronic device 100 is communicatively connected to the sequencer 200. Among them, the sequencer 200 is used to collect sample data (for example, protein sequences) of a user and send the sample data to the electronic device 100.

[0036] In an embodiment of the present application, when a user needs to perform transcription factor regulatory network analysis using the electronic device 100, the electronic device 100 can run a preset analysis instruction, so that the electronic device 100 receives the protein sequence of the species to be analyzed sent by the sequencer 200, predicts multiple target transcription factors of the protein sequence, and performs transcription factor regulatory network analysis on the protein sequence using a preset transcription factor analysis software (for example, SCENIC software).

[0037] Exemplarily, when facing the transcription factor regulatory network analysis of non-model plant single-cell or spatial transcriptome, the electronic device 100 can receive the single-cell protein sequence or spatial transcriptome protein sequence of the non-model plant sent by the sequencer 200, and execute the transcription factor analysis method provided by the present application, so that the electronic device 100 predicts the target transcription factors corresponding to the single-cell protein sequence or spatial transcriptome protein sequence of the non-model plant, thereby realizing one-key transcription factor regulatory network analysis of the single-cell or spatial transcriptome of the non-model plant. In this way, the efficiency of transcription factor regulatory network analysis of non-model plants can be improved.

[0038] As Figure 2 shown, it is a flowchart of the transcription factor analysis method provided by an embodiment of the present application. According to different requirements, the order of steps in this flowchart can be changed, and some steps can be omitted. The transcription factor analysis method provided by the embodiments of the present application includes the following steps.

[0039] S20, receive the protein sequence of the species to be analyzed.

[0040] In an embodiment of the present application, in order to analyze the transcription factors of the species to be analyzed, first obtain the protein sequence of the species to be analyzed, which can provide data support for predicting the transcription factors of the species to be analyzed. Among them, the protein sequence of the species to be analyzed can be the whole genome protein sequence, and the information included in the whole genome protein sequence includes: the sequence and coding sequence of the protein, and the position information of the protein in the genome of the species to be analyzed.

[0041] In an embodiment of the present application, the whole genome protein sequence of the species to be analyzed can be obtained by downloading the genome annotation file of the species to be analyzed from any open-source database (for example, NCBI database). Among them, the genome annotation file includes the sequences and coding sequences of all proteins in the genome of the species to be analyzed, and the position information of these proteins in the genome; it can also collect sample data of the species to be analyzed through a gene sequencer and perform whole genome sequencing (Whole Genome Sequencing, WGS) on the sample data to obtain the whole genome protein sequence of the species to be analyzed. The embodiments of the present application do not limit the method of obtaining the protein sequence of the species to be analyzed.

[0042] S21, Predict multiple target transcription factors of the protein sequence.

[0043] In an embodiment of the present application, in order to perform a regulatory network analysis of transcription factors for a species to be analyzed, it is necessary to first predict the corresponding transcription factors according to the protein sequence of the species to be analyzed. Among them, the transcription factor (TF) corresponding to the species to be analyzed can be a protein that can bind to the DNA of the species to be analyzed and is used to regulate the gene expression of the species to be analyzed. Specifically, the transcription factor binds to the sequence between 500 bases upstream and 100 bases downstream of the transcription start site of the protein, thereby regulating the gene transcription process of the species to be analyzed.

[0044] In an embodiment of the present application, multiple target transcription factors of the species to be analyzed can be obtained by inputting the whole-genome protein sequence into a preset analysis tool. Among them, the preset analysis tool can be the iTAK tool. The iTAK tool predicts transcription factors by identifying protein domains specific to gene families and can identify transcription factors, transcriptional regulators, and protein kinases of the species to be analyzed from protein or nucleotide sequences.

[0045] In an embodiment of the present application, in order to facilitate the invocation of multiple target transcription factors of the species to be analyzed during the subsequent transcription factor regulatory network analysis, the multiple target transcription factors can be stored in a list format.

[0046] S22, Determine the target motif corresponding to each target transcription factor by performing a homology analysis on the multiple target transcription factors and pre-stored template transcription factors.

[0047] In an embodiment of the present application, after obtaining the transcription factors of the species to be analyzed, it is also necessary to determine the motif corresponding to each transcription factor, so as to provide data support for the subsequent transcription factor regulatory network analysis. Among them, the motif of the transcription factor can be a specific region in the transcription factor that binds to DNA, usually composed of 6-12 amino acids, and is used to determine the specific binding of the transcription factor to DNA, thereby regulating gene expression.

[0048] In an embodiment of the present application, a transcription factor motif association list is obtained by performing a homology analysis on the multiple target transcription factors and pre-stored template transcription factors. Among them, the homology analysis of transcription factors is used to determine the similarity between different transcription factors; the pre-stored template transcription factors can be used to store the transcription factors and corresponding motifs of multiple known species. Among them, the pre-stored template transcription factors are composed of transcription factor motif files belonging to multiple different databases. The method for obtaining the pre-stored template transcription factors can refer to the detailed description of the flowchart shown below. Figure 6 For the detailed description.

[0049] In one embodiment of the present application, the BLAST algorithm can be used to perform homology analysis on the target transcription factors of the species to be analyzed and the pre-stored template transcription factors, so as to obtain a transcription factor motif association list. As Figure 3 shown, it is a schematic diagram of the transcription factor motif association list. Among them, each row is used to characterize the corresponding relationship between the target transcription factor and the template transcription factor. The first column is used to record the ID of the motif, the third column is used to record the ID of the template motif corresponding to the template transcription factor, and the last column is used to record the homology score between the target transcription factor and the template transcription factor. The higher the homology score, the higher the similarity between the target transcription factor and the template transcription factor. Exemplarily, as Figure 3 shown in row 3, the ID of the target transcription factor is G1480.1, the ID of the template motif of the template transcription factor corresponding to this target transcription factor is AT5G65410.1, and the homology score between this target transcription factor and the template transcription factor is 89.4. Specifically, the method for obtaining the homology score can refer to the detailed description of the flowchart shown in Figure 8 below.

[0050] In one embodiment of the present application, the motif corresponding to the template transcription factor with the highest homology score can be determined as the target motif.

[0051] S23, perform co-expression analysis on each of the target transcription factors and the corresponding target motifs to obtain the co-expression module of each of the target transcription factors.

[0052] In one embodiment of the present application, co-expression analysis is an analysis method that uses a large amount of gene expression data to construct the correlation between genes. Through the co-expression analysis method, genes that are functionally related can be identified as a module. Through further analysis of the module, advanced analyses such as screening the core genes of the module, associating traits, metabolic pathway modeling, or establishing a gene interaction network can be achieved. In some embodiments, the co-expression analysis methods include WGCNA, GSEA, etc.

[0053] In one embodiment of the present application, the co-expression module includes the target genes of the target transcription factors. The target genes can be genes regulated by transcription factors, and the expression of the target genes can be activated or inhibited by transcription factors. For example, some transcription factors can bind to the promoter region of DNA, thereby promoting the binding and transcription of RNA polymerase, increasing the expression of the target genes; on the contrary, some transcription factors can bind to the enhancer or silencer region, inhibiting the binding and transcription of RNA polymerase, reducing the expression of the target genes. The above corresponding relationship between transcription factors and target genes is an important mechanism for gene expression regulation, enabling cells to respond to different internal and external stimuli by regulating the expression of specific genes.

[0054] In one embodiment of the present application, by inputting the transcription factor and the corresponding motif into a preset analysis tool, a co-expression module of the transcription factor and the target gene can be obtained. The preset analysis tool may be SCENIC (SingleCell Regulatory Network Inference and Clustering) software, which is used to identify the target genes corresponding to the target transcription factor according to the target transcription factor and motif of the species to be analyzed.

[0055] S24, screening the co-expression module based on the target motif to obtain the regulator corresponding to each target transcription factor.

[0056] In one embodiment of the present application, the co-expression module of the transcription factor refers to the set of the transcription factor and a group of genes associated with it, and the expressions of these genes are associated. By analyzing these co-expression modules, it is possible to deeply understand how the transcription factor regulates the expression of the target gene and the role of these regulations in coordinating different biological processes within the cell. A transcription factor may correspond to multiple target genes, and multiple target genes regulated by the same transcription factor are the co-expression module of this transcription factor.

[0057] In one embodiment of the present application, when calculating the co-expression module, the input file is the expression matrix of single-cell or spatial transcriptome and the list of transcription factors, and the output is multiple co-expression modules. The expression conditions of the transcription factor and its target genes are from the single-cell expression matrix. Among them, each row in the expression matrix corresponds to a gene, each column corresponds to a cell, and the value of each element in the expression matrix is used to characterize the expression level of each gene in each cell. For example, when the value of the element in the 3rd row and 3rd column of the expression matrix is 10, it means that the gene in the 3rd column has an expression level of 10 in the cell of the 3rd row.

[0058] In one embodiment of the present application, the co-expression module can be screened according to the target motif corresponding to the target transcription factor, and it is confirmed that the co-expression module obtained after screening is the regulator. Specifically, the method for obtaining the regulator by screening the co-expression module can refer to the detailed description of the flowchart shown below for Figure 8 the flowchart shown.

[0059] S25, by calculating the activity score of the regulator, the transcription factor analysis result of the species to be analyzed is obtained.

[0060] In one embodiment of the present application, the activity score of a regulator can be calculated using the Area Under Curve (AUC) algorithm. Specifically, the target genes containing the binding sites of the target transcription factor in the regulator are determined; the AUC score of the regulator is calculated based on the enrichment score to obtain the activity score of the regulator. Among them, the AUC score can reflect the activity of the regulator, and the higher the AUC score, the stronger the activity of the regulator.

[0061] In one embodiment of the present application, after obtaining the activity score of the regulator, the activity score is further visualized to obtain a heat map of the activity score. Exemplarily, as Figure 4 shown is the heat map of the activity scores of regulators corresponding to different target transcription factors. As Figure 5 shown is the heat map of the activity scores of regulators corresponding to target transcription factors in different cells.

[0062] It can be seen from the above technical solutions that in the embodiment of the present application, by performing homology analysis on the target transcription factor of the species to be analyzed and the pre-stored template transcription factors, the motif corresponding to the target transcription factor is determined, providing data support for subsequent analysis of the transcription factor regulatory network of the species to be analyzed. Then, co-expression analysis is performed on the transcription factor and the corresponding motif to obtain the co-expression module of the transcription factor, and the co-expression module is screened based on the motif to obtain the regulator corresponding to each target transcription factor, which can improve the data quality of the target transcription factor and thus improve the efficiency of transcription factor regulatory network analysis. Furthermore, by calculating the activity score of the regulator, the transcription factor analysis result of the species to be analyzed is obtained, which can realize the automatic analysis of the transcription factor regulatory network of the species to be analyzed, thereby improving the efficiency of transcription factor regulatory network analysis.

[0063] As Figure 6 shown, it is a flowchart of a method for determining a pre-stored template transcription factor provided by an embodiment of the present application. According to different requirements, the order of steps in this flowchart can be changed, and some steps can be omitted. The method for determining a pre-stored template transcription factor provided by an embodiment of the present application includes the following steps.

[0064] S40, obtain the example transcription factors and corresponding example motifs of multiple example species from multiple databases.

[0065] In one embodiment of the present application, in order to provide data support for the analysis of the regulatory network of target transcription factors, it is necessary to first obtain the example transcription factors and example motifs corresponding to multiple example species stored in multiple databases. Exemplarily, the multiple databases include: CisBP database, JASPAR database, PlantTFDB database, and PlantCstrome database. Among them, each database is used to store the example transcription factors and corresponding example motifs of different example species. Among them, the example species can be Arabidopsis thaliana, soybean, rice, wheat, corn, etc.; the example transcription factor refers to the transcription factor corresponding to the protein sequence of the example species; the example motif refers to the motif corresponding to the example transcription factor.

[0066] S41. When multiple of the example transcription factors belong to the same example species and different databases, any one of the example transcription factors and its corresponding example motif are retained.

[0067] In one embodiment of the present application, the example transcription factors and example motifs corresponding to the same example species may be stored in multiple databases. In order to remove redundant data, the example transcription factors and example motifs corresponding to the same example species belonging to different databases can be de-duplicated. Specifically, when multiple of the example transcription factors belong to the same example species and different databases, any one of the example transcription factors and its corresponding example motif are retained, so as to be able to remove redundant data and improve the efficiency of subsequent analysis of the transcription factor regulatory network.

[0068] Exemplarily, when both the CisBP database and the JASPAR database store the example transcription factors and example motifs corresponding to soybeans, any one of the example transcription factors and its corresponding example motif is retained.

[0069] S42. Convert the example transcription factor and the example motif into a unified format to obtain the template transcription factor and the corresponding template motif.

[0070] In one embodiment of the present application, in order to facilitate the storage and invocation of example transcription factors and their corresponding example motifs, all example transcription factors and their corresponding example motifs can be stored in a unified format. Specifically, by inputting the example transcription factor and the example motif into the Cluster-Buster software, the example transcription factor and the example motif can be converted into the Cluster-Buster format. In this way, the formats of the example transcription factor and the example motif can be unified, which is convenient for invocation during the subsequent analysis of the transcription factor regulatory network.

[0071] Such as Figure 7As shown in the figure, it is a flowchart of a method for determining a target motif corresponding to each target transcription factor provided by an embodiment of the present application. According to different requirements, the order of steps in this flowchart can be changed, and some steps can be omitted. The method for determining a target motif corresponding to each target transcription factor provided by the embodiment of the present application includes the following steps.

[0072] S50. Calculate the homology score between each target transcription factor and each template transcription factor.

[0073] In an embodiment of the present application, in order to determine the target motif corresponding to the target transcription factor of the species to be analyzed, first, the homology score between each target transcription factor and each template transcription factor can be calculated, and the target motif corresponding to the target transcription factor can be determined according to the homology score. Specifically, the homology score is obtained by comparing the sequence of the target transcription factor with the sequence of the template transcription factor. The homology score can characterize the similarity between the target transcription factor and the template transcription factor, and can reflect the evolutionary relationship and common function between the target transcription factor and the template transcription factor. By analyzing the homology scores of the target transcription factor and the template transcription factor, their roles and interrelationships in gene expression regulation can be understood, which helps to reveal the complex mechanism of gene expression.

[0074] In an embodiment of the present application, the method for calculating the homology score includes: dividing the target transcription factor to obtain a plurality of first subsequences corresponding to the target transcription factor; dividing the template transcription factor to obtain a plurality of second subsequences corresponding to the template transcription factor; when any one of the first subsequences is the same as any one of the second subsequences, determining the length of the first subsequence as the candidate homology score; and determining the maximum value among the candidate homology scores as the homology score between the target transcription factor and the template transcription factor.

[0075] S51. Determine the template motif of the template transcription factor corresponding to the highest homology score as the target motif.

[0076] In an embodiment of the present application, when the homology score is relatively high, it indicates that the similarity between the sequence of the target transcription factor and the sequence of the template transcription factor is relatively high, then the target transcription factor and the template transcription factor may have similar functions or genes co-regulated. Therefore, for any one target transcription factor, determine the motif of the template transcription factor corresponding to the highest homology score as the target motif.

[0077] As Figure 8 shown in the figure, it is a flowchart of a method for determining the regulator of a target transcription factor provided by an embodiment of the present application. According to different requirements, the order of steps in this flowchart can be changed, and some steps can be omitted. The method for determining the regulator of a target transcription factor provided by the embodiment of the present application includes the following steps.

[0078] S70. Calculate the enrichment score of the target transcription factor in the corresponding target gene.

[0079] In one embodiment of the present application, for the target gene, calculate the number and proportion of genes belonging to a certain transcription factor family, and then use these statistical data to calculate the enrichment degree of the transcription factor family in the given gene list. The enrichment score can reflect the significance of a certain transcription factor family in the given gene list, and further reveal that these genes may be regulated by the transcription factor family. Specifically, the method for calculating the enrichment score can refer to the detailed description of the flowchart shown below for Figure 9 the flowchart shown.

[0080] S71. Determine the binding site of the target transcription factor to the corresponding target gene.

[0081] In one embodiment of the present application, the binding site of the target transcription factor to the corresponding target gene refers to the region where the target transcription factor binds to a specific sequence of the target gene. By recognizing and binding to the specific sequence of the target gene, the target transcription factor can regulate the transcription and expression of the target gene.

[0082] In one embodiment of the present application, the method for determining the binding site of the target transcription factor to the corresponding target gene includes any one of the following: by aligning the sequences of the target transcription factor and the target gene, determining the similarity and matching degree between them, so as to predict the position of the binding site; by analyzing the target motif corresponding to the target transcription factor, predicting the binding site of the target gene and the target transcription factor; using a database search tool (for example, JASPAR tool, TRANSFAC tool, etc.) to search and download the binding site information of the target transcription factor and the target gene that has been predicted.

[0083] S72. When the target motif corresponding to the target transcription factor in the co-expression module is upstream of the binding site of the target gene and the enrichment score is greater than a preset threshold, determine that the co-expression module is a regulon.

[0084] In one embodiment of the present application, when it is determined that the target motif corresponding to the target transcription factor in the co-expression module is upstream of the transcription start site of the target gene and the enrichment score of the target transcription factor in the target gene is greater than a preset threshold, it indicates that the target transcription factor can directly regulate the expression of the corresponding target gene. Therefore, it can be determined that this co-expression module is a regulon.

[0085] In an embodiment of the present application, candidate target genes can be screened by the enrichment ranking of the motif corresponding to the target transcription factor upstream of the target gene. Specifically, candidate target genes can be screened based on the association list of the target transcription factor and its corresponding motif, the list of the target transcription factor and its target genes, and the enrichment ranking position of each motif on each gene in the whole genome. Among them, determining the enrichment ranking position of each motif on each gene in the whole genome includes: using the Cluster-Buster algorithm to calculate the sequence similarity between the motif (i.e., the sequence feature of the binding site of the transcription factor and the target gene) and the upstream sequence (DNA sequence) of the transcription start site of all genes of the species to be analyzed (for example, non-model plants); sorting all genes of the species to be analyzed in descending order according to the sequence similarity to obtain the cisTarget database; among them, the genes with higher sequence similarity are ranked higher; the enrichment ranking of the motif corresponding to the target transcription factor is stored in the form of a matrix in the cisTarget database, each row of the matrix corresponds to a motif, each column of the matrix corresponds to a gene, and each element in the matrix is used to represent the enrichment score of the target transcription factor, that is, the enrichment ranking of the motif of a certain target transcription factor upstream of a certain gene. For example, when the element in the 4th row and 4th column of the cisTarget database is 1, it indicates that the motif in the 4th row has an enrichment ranking of the 1st upstream of the gene in the 4th column.

[0086] In an embodiment of the present application, screening candidate target genes includes: for each motif, determining multiple genes with higher rankings in the genes corresponding to the motif as high-ranking genes (for example, determining the top 50 genes as high-ranking genes); when each candidate target gene of the target transcription factor appears in the high-ranking genes, then retain the high-ranking gene; when any candidate target gene of the target transcription factor does not appear in the high-ranking genes, then do not retain the high-ranking gene; determining the target transcription factor and the high-ranking genes corresponding to the target transcription factor as the regulator corresponding to the target transcription factor.

[0087] As Figure 9 shown, it is a flowchart of a method for calculating the enrichment score of a target transcription factor provided in another embodiment of the present application. According to different requirements, the order of steps in this flowchart can be changed, and some steps can be omitted. The method for calculating the enrichment score of a target transcription factor provided in the embodiments of the present application includes the following steps.

[0088] S80, calculate the first frequency of occurrence of the target motif of the target transcription factor in the target gene.

[0089] In an embodiment of the present application, first, a first binding site of a target transcription factor and a target gene is determined. The first binding site is a DNA fragment in the target gene that binds to the target transcription factor, and the length of the first binding site can be in the range of 5 to 20 base pairs. A DNA fragment having the same length as the first binding site is extracted from the target gene as a candidate binding site. The number of occurrences of the motif of the target transcription factor in the candidate binding sites is counted, and the ratio between the number of occurrences and the total number of candidate binding sites is calculated to obtain the first frequency of the target motif of the target transcription factor in the target gene.

[0090] S81, calculate the second frequency of the target motif of the target transcription factor in the protein sequence.

[0091] In an embodiment of the present application, a second binding site of the target transcription factor and the protein sequence of the species to be analyzed is determined, and a subsequence having the same length as the second binding site is extracted from the protein sequence. The number of occurrences of the motif of the target transcription factor in the subsequence is counted, and the ratio between the number of occurrences and the total number of subsequences is calculated to obtain the second frequency of the motif of the target transcription factor in the protein sequence.

[0092] S82, calculate the ratio of the first frequency to the second frequency to obtain the enrichment score of the target transcription factor in the corresponding target gene.

[0093] In an embodiment of the present application, the higher the enrichment score, the higher the importance of the target transcription factor in the target gene, and it plays an important regulatory role in the transcription and expression of the target gene. Exemplarily, when the first frequency is 0.48 and the second frequency is 0.6, the enrichment score is 0.8.

[0094] Please refer to Figure 10 , Figure 10 is a functional module diagram of a transcription factor analysis device provided in an embodiment of the present application. The transcription factor analysis device 11 includes a receiving module 110 and an analysis module 111. The module / unit referred to in the present application means a series of computer-readable instruction segments that can be executed by a processor 13 and can complete fixed functions, and are stored in a memory 12. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.

[0095] The receiving module 110 is used to receive the protein sequence of the species to be analyzed;

[0096] The analysis module 111 is used to predict multiple target transcription factors of the protein sequence;

[0097] The analysis module 111 is further configured to determine a target motif corresponding to each target transcription factor by performing homology analysis on the multiple target transcription factors and pre-stored template transcription factors.

[0098] The analysis module 111 is further configured to perform co-expression analysis on each target transcription factor and the corresponding target motif to obtain a co-expression module for each target transcription factor.

[0099] The analysis module 111 is further configured to screen the co-expression modules based on the target motifs to obtain a regulator corresponding to each target transcription factor.

[0100] The analysis module 111 is further configured to obtain a transcription factor analysis result of the species to be analyzed by calculating an activity score of the regulator.

[0101] In an embodiment of the present application, the analysis module 111 is further configured to: obtain example transcription factors of multiple example species and corresponding example motifs from multiple databases; when multiple example transcription factors belong to the same example species and different databases, retain any one of the multiple example transcription factors and its corresponding example motif; convert the example transcription factors and example motifs into a unified format to obtain the template transcription factors and corresponding template motifs.

[0102] In an embodiment of the present application, the analysis module 111 is further configured to: calculate a homology score between each target transcription factor and each template transcription factor; determine the motif of the template transcription factor corresponding to the highest homology score as the target motif.

[0103] In an embodiment of the present application, the analysis module 111 is further configured to: divide the target transcription factor to obtain multiple first subsequences corresponding to the target transcription factor; divide the template transcription factor to obtain multiple second subsequences corresponding to the template transcription factor; when any one of the first subsequences is the same as any one of the second subsequences, determine the length of the first subsequence as a candidate homology score; determine the maximum value among the candidate homology scores as the homology score between the target transcription factor and the template transcription factor.

[0104] In an embodiment of the present application, the analysis module 111 is further configured to: calculate an enrichment score of the target transcription factor in the corresponding target gene; determine a binding site between the target transcription factor and the corresponding target gene; when the target motif corresponding to the target transcription factor in the co-expression module is upstream of the binding site of the target gene and the enrichment score is greater than a preset threshold, determine the co-expression module as a regulator.

[0105] In an embodiment of the present application, the analysis module 111 is further configured to: calculate a first frequency of occurrence of the target motif of the target transcription factor in the target gene; calculate a second frequency of occurrence of the target motif of the target transcription factor in the protein sequence; calculate a ratio of the first frequency to the second frequency to obtain an enrichment score of the target transcription factor in the corresponding target gene.

[0106] In an embodiment of the present application, the analysis module 111 is further configured to: determine the target gene in the regulon that contains the binding site of the target transcription factor; calculate the AUC score of the regulon according to the enrichment score to obtain the activity score of the regulon.

[0107] It can be seen from the above technical solutions that in the embodiment of the present application, by performing homology analysis on the target transcription factor of the species to be analyzed and the pre-stored template transcription factor, the motif corresponding to the target transcription factor is determined, providing data support for subsequent analysis of the transcription factor regulatory network of the species to be analyzed. Then, co-expression analysis is performed on the transcription factor and the corresponding motif to obtain the co-expression module of the transcription factor, and the co-expression module is screened based on the motif to obtain the regulon corresponding to each target transcription factor, which can improve the data quality of the target transcription factor, thereby improving the efficiency of transcription factor regulatory network analysis. By calculating the activity score of the regulon, the transcription factor analysis result of the species to be analyzed is obtained, which can realize automatic analysis of the transcription factor regulatory network of the species to be analyzed, thereby improving the efficiency of transcription factor regulatory network analysis.

[0108] Please refer to Figure 11 , which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device 1 includes a memory 12 and a processor 13. The memory 12 is used to store computer-readable instructions, and the processor 13 is used to execute the computer-readable instructions stored in the memory to implement the transcription factor analysis method described in any of the above embodiments.

[0109] In an embodiment of the present application, the electronic device 1 further includes a bus and a computer program stored in the memory 12 and executable on the processor 13, such as a transcription factor analysis program.

[0110] Figure 11 Only the electronic device 1 with the memory 12 and the processor 13 is shown. Those skilled in the art can understand that Figure 11 the shown structure does not constitute a limitation on the electronic device 1, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0111] Combined with Figure 2, the memory 12 in the electronic device 1 stores multiple computer-readable instructions to implement a transcription factor analysis method. The processor 13 can execute the multiple instructions to implement: obtaining a protein sequence of a species to be analyzed and multiple target transcription factors of the protein sequence; determining a target motif corresponding to each target transcription factor by performing a homology analysis on the multiple target transcription factors and pre-stored template transcription factors; performing a co-expression analysis on each transcription factor and the corresponding target motif to obtain a co-expression module of each transcription factor; screening the co-expression modules based on the target motif to obtain a regulator corresponding to each target transcription factor; and obtaining a transcription factor analysis result of the species to be analyzed by calculating an activity score of the regulator.

[0112] Specifically, the specific implementation method of the processor 13 for the above instructions can refer to Figure 2 the description of the relevant steps in the corresponding embodiment, which will not be elaborated here.

[0113] Those skilled in the art can understand that the schematic diagram is only an example of the electronic device 1 and does not constitute a limitation on the electronic device 1. The electronic device 1 can be a bus structure or a star structure. The electronic device 1 can also include more or fewer other hardware or software than shown in the figure, or different component arrangements. For example, the electronic device 1 can also include input / output devices, network access devices, etc.

[0114] It should be noted that the electronic device 1 is only an example. Other existing or future possible electronic products that can be adapted to this application should also be included in the protection scope of this application and are incorporated herein by reference.

[0115] Among them, the memory 12 includes at least one type of readable storage medium. The readable storage medium can be non-volatile or volatile. The readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. The memory 12 can be an internal storage unit of the electronic device 1 in some embodiments, such as the mobile hard disk of the electronic device 1. The memory 12 can also be an external storage device of the electronic device 1 in other embodiments, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 1. The memory 12 can not only be used to store application software installed in the electronic device 1 and various types of data, such as the code of the transcription factor analysis program, etc., but also be used to temporarily store data that has been output or will be output.

[0116] In some embodiments, the processor 13 may be composed of an integrated circuit. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple packaged integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 13 is the control core of the electronic device 1. It uses various interfaces and circuits to connect all components of the entire electronic device 1. By running or executing programs or modules stored in the memory 12 (such as executing a transcription factor analysis program, etc.), and by calling the data stored in the memory 12, it executes various functions of the electronic device 1 and processes data.

[0117] The processor 13 executes the operating system of the electronic device 1 and various installed application programs. The processor 13 executes the application programs to implement the steps in the above-mentioned embodiments of various transcription factor analysis methods. For example Figure 2 the steps shown.

[0118] Exemplarily, the computer program may be divided into one or more modules / units. The one or more modules / units are stored in the memory 12 and executed by the processor 13 to complete the present application. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device 1. For example, the computer program may be divided into a receiving module 110 and an analyzing module 111.

[0119] The above-mentioned integrated units implemented in the form of software function modules may be stored in a computer-readable storage medium. The above-mentioned software function modules are stored in a storage medium, including several instructions for causing a computer device (which may be a personal computer, a computer device, or a network device, etc.) or a processor to execute part of the transcription factor analysis methods described in the various embodiments of the present application.

[0120] If the integrated module / unit of the electronic device 1 is implemented in the form of a software function unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it may also be completed by a computer program instructing relevant hardware devices. The computer program may be stored in a computer-readable storage medium. When the computer program is executed by the processor, it may implement the steps in the above-mentioned various method embodiments.

[0121] Among them, the computer program includes computer program code, which may be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory, and other memories, etc.

[0122] Further, the computer-readable storage medium mainly includes a storage program area and a storage data area. Among them, the storage program area can store an operating system, application programs required for at least one function, etc.; the storage data area can store data created according to the use of the blockchain node, etc.

[0123] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, in Figure 11 only one arrow is used to represent it, but it does not mean that there is only one bus or one type of bus. The bus is set to realize the connection and communication between the memory 12 and at least one processor 13, etc.

[0124] The embodiment of the present application also provides a computer-readable storage medium (not shown in the figure). The computer-readable storage medium stores computer-readable instructions, and the computer-readable instructions are executed by a processor in an electronic device to implement the transcription factor analysis method described in any one of the above embodiments.

[0125] In several embodiments provided by the present application, it should be understood that the disclosed system, device, and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0126] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0127] In addition, in each embodiment of the present application, each functional module can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.

[0128] In addition, it is obvious that the term "including" does not exclude other units or steps, and the singular does not exclude the plural. A plurality of units or devices described in the specification can also be implemented by one unit or device through software or hardware. The terms such as "first" and "second" are used to indicate names, rather than any specific order.

[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit them. Although the present application has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present application can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present application.

Claims

1. A method for analyzing transcription factors, applied to an electronic device, characterized in that, the method includes: obtaining the protein sequence of a species to be analyzed, and multiple target transcription factors of the protein sequence; determining the target motif corresponding to each target transcription factor by performing homology analysis on the multiple target transcription factors and pre-stored template transcription factors; performing co-expression analysis on each target transcription factor and the corresponding target motif to obtain the co-expression module of each target transcription factor; screening the co-expression module based on the target motif to obtain the regulator corresponding to each target transcription factor; obtaining the transcription factor analysis result of the species to be analyzed by calculating the activity score of the regulator.

2. The transcription factor analysis method according to claim 1, characterized in that, the method for obtaining the pre-stored template transcription factors includes: obtaining the example transcription factors and the corresponding example motifs of multiple example species from multiple databases; when multiple of the example transcription factors belong to the same example species and different databases, retaining any one of the example transcription factors and its corresponding example motif among the multiple example transcription factors; converting the example transcription factors and example motifs into a unified format to obtain the template transcription factors and the corresponding template motifs.

3. The transcription factor analysis method according to claim 1, characterized in that, the template transcription factor corresponds to a template motif, and the determining the target motif corresponding to each target transcription factor by performing homology analysis on the multiple target transcription factors and pre-stored template transcription factors includes: calculating the homology score between each target transcription factor and each template transcription factor; determining the motif of the template transcription factor corresponding to the highest homology score as the target motif.

4. The transcription factor analysis method according to claim 3, characterized in that, the calculating the homology score between each target transcription factor and each template transcription factor includes: dividing the target transcription factor to obtain multiple first subsequences corresponding to the target transcription factor; dividing the template transcription factor to obtain multiple second subsequences corresponding to the template transcription factor; when any one of the first subsequences is the same as any one of the second subsequences, determining the length of the first subsequence as the candidate homology score; determining the maximum value among the candidate homology scores as the homology score between the target transcription factor and the template transcription factor.

5. The transcription factor analysis method according to claim 1, characterized in that, the co-expression module includes the target genes of the target transcription factor, and the screening the co-expression module based on the target motif to obtain the regulator corresponding to each target transcription factor includes: calculating the enrichment score of the target transcription factor in the corresponding target gene; determining the binding site between the target transcription factor and the corresponding target gene; when the target motif corresponding to the target transcription factor in the co-expression module is upstream of the binding site of the target gene and the enrichment score is greater than a preset threshold, determining the co-expression module as the regulator.

6. The transcription factor analysis method according to claim 5, characterized in that, Calculating the enrichment score of the target transcription factor in the corresponding target gene includes: Calculating a first frequency of occurrence of the target motif of the target transcription factor in the target gene; Calculating a second frequency of occurrence of the target motif of the target transcription factor in the protein sequence; Calculating a ratio of the first frequency to the second frequency to obtain the enrichment score of the target transcription factor in the corresponding target gene.

7. The transcription factor analysis method according to claim 5, wherein, obtaining the transcription factor analysis result of the species to be analyzed by calculating the activity score of the regulon includes: Determining the target gene in the regulon that contains the binding site of the target transcription factor; Calculating the AUC score of the regulon according to the enrichment score to obtain the activity score of the regulon.

8. A transcription factor analysis device, wherein, the device includes: A receiving module for receiving the protein sequence of the species to be analyzed; An analysis module for predicting a plurality of target transcription factors of the protein sequence; The analysis module is further configured to determine the target motif corresponding to each target transcription factor by performing homology analysis on the plurality of target transcription factors and pre-stored template transcription factors; The analysis module is further configured to perform co-expression analysis on each target transcription factor and the corresponding target motif to obtain the co-expression module of each target transcription factor; The analysis module is further configured to screen the co-expression module based on the target motif to obtain the regulon corresponding to each target transcription factor; The analysis module is further configured to obtain the transcription factor analysis result of the species to be analyzed by calculating the activity score of the regulon.

9. An electronic device, wherein, the electronic device includes: A memory storing computer-readable instructions; and A processor for executing the computer-readable instructions stored in the memory to implement the transcription factor analysis method according to any one of claims 1 to 7.

10. A computer-readable storage medium, wherein, computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by a processor, the transcription factor analysis method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Interfering with hd-zip transcription factor repression of gene expression to produce plants with enhanced traits

    CN105611828A

  • Corn nitrogen response gene regulatory element analysis method

    CN115547411A

  • Transcription factor target gene relation prediction method, system, equipment and medium

    CN116230070A

  • Methods and compositions related to regulation of nucleic acids

    US20160004814A1

  • Methods for predicting transcription factor activity

    US20190385697A1

Cited By

  • Cell specific transcription factor regulatory network analysis method and visualization platform

    CN120748515A