Transcription factor analysis method, apparatus, electronic device, and storage medium

By performing homology and co-expression analyses on the protein sequences of the species to be analyzed, the motifs and regulators of target transcription factors were identified, solving the problems of disordered transcription factor database formats and insufficient automated analysis, and achieving efficient transcription factor regulatory network analysis.

CN120072049BActive Publication Date: 2025-12-09SHENZHEN HUADA GENE INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311612037.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-12-09
Estimated Expiration
2043-11-28

AI Technical Summary

Technical Problem

Existing transcription factor databases suffer from disorganized motif matrix file formats and a lack of automated workflows for analyzing transcription factor regulatory networks in non-model plants, resulting in low analysis efficiency.

Method used

By acquiring the protein sequences of the species to be analyzed, predicting multiple target transcription factors, performing homology analysis with pre-stored template transcription factors, identifying target motifs, conducting co-expression analysis, screening co-expression modules, and calculating the activity scores of regulators, automated transcription factor regulatory network analysis is achieved.

Benefits of technology

It improves the efficiency of transcription factor regulatory network analysis, enhances data quality and the degree of automation in analysis, and significantly improves efficiency, especially in single-cell or spatial transcriptome analysis of non-model plants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072049B_ABST
    Figure CN120072049B_ABST
Patent Text Reader

Abstract

The application provides a transcription factor analysis method and device, electronic equipment and a storage medium. The transcription factor analysis method comprises the following steps: obtaining protein sequences of a species to be analyzed and a plurality of target transcription factors of the protein sequences; determining a target motif corresponding to each target transcription factor by performing homology analysis on the plurality of target transcription factors and pre-stored template transcription factors; performing co-expression analysis on each target transcription factor and the corresponding target motif to obtain a co-expression module of each target transcription factor; screening the co-expression module based on the target motif to obtain a regulatory subunit corresponding to each target transcription factor; and obtaining a transcription factor analysis result of the species to be analyzed by calculating an activity score of the regulatory subunit. The application can improve the efficiency of transcription factor analysis.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of biological information analysis, in particular to the technical field of transcription factor analysis, and especially to a transcription factor analysis method and device, an electronic equipment and a storage medium. BACKGROUND

[0002] Transcription factors play an important role in biological processes such as growth and development of species, environmental response and growth cycle regulation. The regulation network analysis of transcription factors of a species can provide important data support and guidance for gene function research, disease development mechanism research, drug research, etc. At present, the format and annotation of the motif matrix file in the existing transcription factor database are relatively chaotic, and there is no automatic process suitable for non-model plant transcription factor regulation network analysis, resulting in low efficiency of transcription factor regulation network analysis of the species to be analyzed. SUMMARY

[0003] In view of the above, it is necessary to provide a transcription factor analysis method, device, electronic equipment and storage medium to solve the technical problem of low efficiency of transcription factor regulation network analysis.

[0004] The present application provides a transcription factor analysis method, which comprises: obtaining protein sequences of a species to be analyzed and a plurality of target transcription factors of the protein sequences; determining a target motif corresponding to each target transcription factor by performing homology analysis on the plurality of target transcription factors and pre-stored template transcription factors; performing co-expression analysis on each target transcription factor and the corresponding target motif to obtain a co-expression module of each target transcription factor; filtering the co-expression module based on the target motif to obtain a regulatory subunit corresponding to each target transcription factor; and obtaining a transcription factor analysis result of the species to be analyzed by calculating the activity score of the regulatory subunit.

[0005] In some embodiments, the method of obtaining the pre-stored template transcription factor comprises: obtaining example transcription factors and corresponding example motifs of example species from a plurality of databases; when a plurality of the example transcription factors belong to the same example species and belong to different databases, retaining any one of the plurality of example transcription factors and the corresponding example motif thereof; and converting the example transcription factors and example motifs into a unified format to obtain the template transcription factors and corresponding template motifs.

[0006] In some embodiments, the template transcription factor corresponds to a template motif, and the determining a target motif corresponding to each target transcription factor by performing homology analysis on the plurality of target transcription factors and the pre-stored template transcription factors comprises: calculating a homology score of each target transcription factor and each template transcription factor; and determining a motif of a template transcription factor corresponding to a highest homology score as the target motif.

[0007] In some embodiments, the calculating a homology score of each target transcription factor and each template transcription factor comprises: dividing the target transcription factor to obtain a plurality of first sub-sequences corresponding to the target transcription factor; dividing the template transcription factor to obtain a plurality of second sub-sequences corresponding to the template transcription factor; determining a length of a first sub-sequence as a candidate homology score when the first sub-sequence is identical to any second sub-sequence; and determining a maximum value in the candidate homology scores as the homology score of the target transcription factor and the template transcription factor.

[0008] In some embodiments, the co-expression module comprises a target gene of the target transcription factor, and the screening the co-expression module based on the target motif to obtain a regulator corresponding to each target transcription factor comprises: calculating an enrichment score of the target transcription factor in the corresponding target gene; determining a binding site of the target transcription factor and the corresponding target gene; and determining the co-expression module as a regulator when a target motif corresponding to the target transcription factor in the co-expression module is upstream of the binding site of the target gene and the enrichment score is greater than a preset threshold.

[0009] In some embodiments, the calculating the enrichment score of the target transcription factor in the corresponding target gene comprises: calculating a first frequency of a target motif of the target transcription factor in the target gene; calculating a second frequency of the target motif of the target transcription factor in the protein sequence; and calculating a ratio of the first frequency and the second frequency to obtain the enrichment score of the target transcription factor in the corresponding target gene.

[0010] In some embodiments, the obtaining a transcription factor analysis result of the to-be-analyzed species by calculating an activity score of the regulator comprises: determining the target gene containing the binding site of the target transcription factor in the regulator; and calculating an AUC score of the regulator according to the enrichment score to obtain the activity score of the regulator.

[0011] The embodiment of the present application also provides a transcription factor analysis device, which comprises: a receiving module configured to receive a protein sequence of a species to be analyzed; an analysis module configured to predict a plurality of target transcription factors of the protein sequence; the analysis module is further configured to determine a target motif corresponding to each target transcription factor by performing homology analysis on the plurality of target transcription factors and a pre-stored template transcription factor; the analysis module is further configured to perform co-expression analysis on each target transcription factor and the corresponding target motif to obtain a co-expression module of each target transcription factor; the analysis module is further configured to filter the co-expression module based on the target motif to obtain a regulator corresponding to each target transcription factor; and the analysis module is further configured to obtain a transcription factor analysis result of the species to be analyzed by calculating an activity score of the regulator.

[0012] The embodiment of the present application also provides an electronic device, which comprises:

[0013] a memory configured to store at least one instruction;

[0014] a processor configured to execute the instruction stored in the memory to implement the transcription factor analysis method.

[0015] The embodiment of the present application also provides a computer readable storage medium, which stores at least one instruction, and the at least one instruction is executed by a processor in an electronic device to implement the transcription factor analysis method.

[0016] As can be seen from the above technical solutions, the embodiment of the present application determines the motif corresponding to the target transcription factor by performing homology analysis on the target transcription factor of the species to be analyzed and the pre-stored template transcription factor, thereby providing data support for subsequent transcription factor regulatory network analysis of the species to be analyzed. Then, the co-expression module of the transcription factor is obtained by performing co-expression analysis on the transcription factor and the corresponding motif, and the regulator corresponding to each target transcription factor is obtained by filtering the co-expression module based on the motif, so that the data quality of the target transcription factor is improved, and the efficiency of the transcription factor regulatory network analysis is improved. Finally, the transcription factor analysis result of the species to be analyzed is obtained by calculating the activity score of the regulator, so that the transcription factor regulatory network analysis of the species to be analyzed is automatically performed, and the efficiency of the transcription factor regulatory network analysis is improved. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 FIG. 1 is an application scenario diagram of a transcription factor analysis method according to an embodiment of the present application.

[0018] Figure 2 FIG. 2 is a flowchart of a transcription factor analysis method according to an embodiment of the present application.

[0019] Figure 3 is a transcription factor motif association list diagram provided by an embodiment of the present application.

[0020] Figure 4 is a regulatory activity score heat map provided by an embodiment of the present application.

[0021] Figure 5 is a regulatory activity score heat map in different cells provided by an embodiment of the present application.

[0022] Figure 6 is a flowchart of a method for determining a pre-stored template transcription factor provided by an embodiment of the present application.

[0023] Figure 7 is a flowchart of a method for determining a target motif corresponding to each target transcription factor provided by an embodiment of the present application.

[0024] Figure 8 is a flowchart of a method for determining a regulator of a target transcription factor provided by an embodiment of the present application.

[0025] Figure 9 is a flowchart of a method for calculating an enrichment score of a target transcription factor provided by an embodiment of the present application.

[0026] Figure 10 is a functional module diagram of a transcription factor analysis device provided by an embodiment of the present application.

[0027] Figure 11 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0028] In order to more clearly understand the purpose, features and advantages of the present application, the present application will be described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict. In the following description, a large number of specific details are set forth in order to facilitate a full understanding of the present application, and the described embodiments are only some of the embodiments of the present application, not all embodiments.

[0029] In addition, the terms "first", "second" are only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. In the description of the present application, the meaning of "a plurality of" is two or more, unless otherwise specifically limited.

[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0031] This application provides a transcription factor analysis method that can be applied to one or more electronic devices. An electronic device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0032] Electronic devices can be any electronic product that allows human-computer interaction with a customer, such as personal computers, tablets, smartphones, personal digital assistants (PDAs), game consoles, interactive network television (IPTV), smart wearable devices, etc.

[0033] Electronic devices may also include network devices and / or client devices. The network devices include, but are not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.

[0034] The networks in which electronic devices are located include, but are not limited to, the Internet, wide area networks, metropolitan area networks, local area networks, and virtual private networks (VPNs).

[0035] like Figure 1 As shown, the transcription factor analysis method provided in this application can be applied to an electronic device 100, which is communicatively connected to a sequencer 200. The sequencer 200 is used to collect user sample data (e.g., protein sequences) and send the sample data to the electronic device 100.

[0036] In an embodiment of the present application, when a user needs to perform transcription factor regulatory network analysis by using the electronic device 100, the electronic device 100 can be used to run a preset analysis instruction, so that the electronic device 100 receives the protein sequence of the species to be analyzed sent by the sequencer 200, and predicts a plurality of target transcription factors of the protein sequence, and uses a preset transcription factor analysis software (for example, SCENIC software) to perform transcription factor regulatory network analysis on the protein sequence.

[0037] For example, when facing transcription factor regulatory network analysis of non-model plant single cell or spatial transcriptome, the electronic device 100 can receive the single cell protein sequence or spatial transcriptome protein sequence of the non-model plant sent by the sequencer 200, and perform the transcription factor analysis method provided by the present application, so that the electronic device 100 predicts the target transcription factors corresponding to the single cell protein sequence or spatial transcriptome protein sequence of the non-model plant, thereby realizing one-key transcription factor regulatory network analysis of the single cell or spatial transcriptome of the non-model plant. In this way, the efficiency of transcription factor regulatory network analysis of the non-model plant can be improved.

[0038] As shown in Figure 2 FIG. 1 is a flowchart of the transcription factor analysis method provided by an embodiment of the present application. The order of the steps in the flowchart can be changed, and some steps can be omitted according to different requirements. The transcription factor analysis method provided by the embodiment of the present application includes the following steps.

[0039] S20, receiving a protein sequence of a species to be analyzed.

[0040] In an embodiment of the present application, in order to analyze the transcription factors of the species to be analyzed, the protein sequence of the species to be analyzed is first obtained, which can provide data support for predicting the transcription factors of the species to be analyzed. The protein sequence of the species to be analyzed can be a whole genome protein sequence, and the information contained in the whole genome protein sequence includes the sequence and coding sequence of the protein, and the position information of the protein in the genome of the species to be analyzed.

[0041] In an embodiment of the present application, the whole genome protein sequence of the species to be analyzed can be obtained by downloading the genome annotation file of the species to be analyzed from any open source database (for example, NCBI database), wherein the genome annotation file includes the sequence and coding sequence of all proteins in the genome of the species to be analyzed, and the position information of these proteins in the genome; the sample data of the species to be analyzed can also be collected by a gene sequencer, and the sample data is subjected to whole genome sequencing (WGS) to obtain the whole genome protein sequence of the species to be analyzed. The method for obtaining the protein sequence of the species to be analyzed is not limited in the embodiment of the present application.

[0042] S21, predicting a plurality of target transcription factors of the protein sequence.

[0043] In an embodiment of the present application, in order to analyze the transcription factor regulatory network of the species to be analyzed, it is necessary to predict the corresponding transcription factors of the species to be analyzed according to the protein sequence of the species to be analyzed. The transcription factor (Transcription Factor, TF) corresponding to the species to be analyzed can be a protein capable of binding to the DNA of the species to be analyzed, used to regulate the expression of the genes of the species to be analyzed. Specifically, the transcription factor binds to the sequence located between 500 bases upstream and 100 bases downstream of the transcription start site of the protein, thereby regulating the gene transcription process of the species to be analyzed.

[0044] In an embodiment of the present application, the plurality of target transcription factors of the species to be analyzed can be obtained by inputting the whole genome protein sequence into a preset analysis tool. The preset analysis tool can be iTAK tool, which predicts transcription factors by identifying protein domains specific to gene families, and can identify transcription factors, transcription regulators and protein kinases of the species to be analyzed from protein or nucleotide sequences.

[0045] In an embodiment of the present application, in order to facilitate the calling of the plurality of target transcription factors of the species to be analyzed in the subsequent transcription factor regulatory network analysis process, the plurality of target transcription factors can be stored in a list format.

[0046] S22, determining the target motif corresponding to each target transcription factor by performing homology analysis on the plurality of target transcription factors and the pre-stored template transcription factors.

[0047] In an embodiment of the present application, after obtaining the transcription factors of the species to be analyzed, it is also necessary to determine the motif corresponding to each transcription factor, thereby providing data support for subsequent transcription factor regulatory network analysis. The motif of the transcription factor can be a specific region in the transcription factor that binds to DNA, usually composed of 6-12 amino acids, used to determine the specific binding of the transcription factor to DNA, thereby regulating the expression of the gene.

[0048] In an embodiment of the present application, the transcription factor motif association list is obtained by performing homology analysis on the plurality of target transcription factors and the pre-stored template transcription factors. The homology analysis of the transcription factor is used to determine the similarity between different transcription factors; the pre-stored template transcription factor can be a transcription factor and the corresponding motif of a plurality of known species. The pre-stored template transcription factor is composed of transcription factor motif files belonging to a plurality of different databases, and the way to obtain the pre-stored template transcription factor can refer to the detailed description of the flowchart below. Figure 6 ​

[0049] In one embodiment of this application, the BLAST algorithm can be used to perform homology analysis between the target transcription factor of the species to be analyzed and a pre-stored template transcription factor to obtain a transcription factor motif association list. For example... Figure 3 The diagram shows a list of transcription factor motif associations. Each row represents the correspondence between a target transcription factor and a template transcription factor. The first column records the motif ID, the third column records the template motif ID corresponding to the template transcription factor, and the last column records the homology score between the target and template transcription factors. A higher homology score indicates a higher degree of similarity between the target and template transcription factors. For example,... Figure 3 The target transcription factor in the third row has the ID G1480.1, and the template motif ID of the corresponding template transcription factor is AT5G65410.1. The homology score between the target and template transcription factors is 89.4. For details on how to obtain the homology score, please refer to the section below. Figure 8 A detailed description of the flowchart shown.

[0050] In one embodiment of this application, the motif corresponding to the template transcription factor with the highest homology score can be identified as the target motif.

[0051] S23, perform co-expression analysis on each target transcription factor and its corresponding target motif to obtain the co-expression module of each target transcription factor.

[0052] In one embodiment of this application, co-expression analysis is an analytical method for constructing correlations between genes using a large amount of gene expression data. Through co-expression analysis, functionally related genes can be identified as a module. Further analysis of this module enables advanced analyses such as screening core genes, identifying associated traits, modeling metabolic pathways, or establishing gene interaction networks. In some embodiments, co-expression analysis methods include WGCNA, GSEA, etc.

[0053] In one embodiment of this application, the co-expression module includes the target gene of the target transcription factor. The target gene can be a gene regulated by the transcription factor, and its expression can be activated or inhibited by the transcription factor. For example, some transcription factors can bind to the promoter region of DNA, thereby promoting the binding and transcription of RNA polymerase and increasing the expression of the target gene; conversely, some transcription factors can bind to enhancer or silencer regions, inhibiting the binding and transcription of RNA polymerase and reducing the expression of the target gene. The correspondence between the above-mentioned transcription factors and target genes is an important mechanism for gene expression regulation, enabling cells to respond to different internal and external stimuli by regulating the expression of specific genes.

[0054] In an embodiment of the present application, the co-expression modules of transcription factors and target genes can be obtained by inputting the transcription factors and corresponding motifs into a preset analysis tool. The preset analysis tool can be SCENIC (Single Cell Regulatory Network Inference and Clustering) software, which is used to identify target genes corresponding to target transcription factors according to the target transcription factors and motifs of the species to be analyzed.

[0055] S24, filtering the co-expression modules based on the target motifs to obtain a regulatory sub-module corresponding to each target transcription factor.

[0056] In an embodiment of the present application, the co-expression module of a transcription factor refers to a collection of a transcription factor and a group of genes associated with the transcription factor, and the genes are associated in expression. Through analysis of these co-expression modules, it can be understood in depth how the transcription factor regulates the expression of target genes and the role of these regulations in coordinating different biological processes in cells. A transcription factor can correspond to multiple target genes, and multiple target genes regulated by the same transcription factor are the co-expression module of the transcription factor.

[0057] In an embodiment of the present application, when calculating the co-expression module, the input file is an expression matrix of single-cell or spatial transcriptome and a list of transcription factors, and the output is multiple co-expression modules. The expression of transcription factors and their target genes comes from the single-cell expression matrix. In the expression matrix, each row corresponds to a gene, and each column corresponds to a cell. The value of each element in the expression matrix is used to represent the expression amount of each gene in each cell. For example, when the value of the 3rd row and the 3rd column in the expression matrix is 10, it represents that the expression amount of the gene in the 3rd column in the cell in the 3rd row is 10.

[0058] In an embodiment of the present application, the co-expression modules can be filtered according to the target motifs corresponding to the target transcription factors, and the co-expression modules obtained after filtering are confirmed as regulatory sub-modules. Specifically, the way to obtain the regulatory sub-modules from the filtered co-expression modules can refer to the detailed description of the flowchart shown in Figure 8

[0059] S25, obtaining the transcription factor analysis result of the species to be analyzed by calculating the activity score of the regulatory sub-module.

[0060] ​In an embodiment of the present application, the activity score of the regulator can be calculated by using the area under curve (AUC) algorithm. Specifically, the target gene containing the binding site of the target transcription factor in the regulator is determined; the AUC score of the regulator is calculated according to the enrichment score, and the activity score of the regulator is obtained. The AUC score can reflect the activity of the regulator, and the higher the AUC score, the stronger the activity of the regulator.

[0061] In an embodiment of the present application, after obtaining the activity score of the regulator, the activity score is further visualized to obtain a heat map of the activity score. For example, as shown in Figure 4 The heat map of the activity score of the regulator corresponding to different target transcription factors is shown. Figure 5 The heat map of the activity score of the regulator corresponding to the target transcription factor in different cells is shown.

[0062] As can be seen from the above technical solutions, in the embodiments of the present application, the homology analysis is performed on the target transcription factor of the to-be-analyzed species and the pre-stored template transcription factor, the motif corresponding to the target transcription factor is determined, and data support is provided for subsequent transcription factor regulatory network analysis of the to-be-analyzed species. Then, the co-expression analysis is performed on the transcription factor and the corresponding motif, the co-expression module of the transcription factor is obtained, the co-expression module is screened based on the motif, the regulator corresponding to each target transcription factor is obtained, the data quality of the target transcription factor is improved, and thus the efficiency of the transcription factor regulatory network analysis is improved. Then, the activity score of the regulator is calculated to obtain the transcription factor analysis result of the to-be-analyzed species, and the transcription factor regulatory network analysis of the to-be-analyzed species can be automatically performed, and thus the efficiency of the transcription factor regulatory network analysis is improved.

[0063] As shown in Figure 6 The flow chart of the method for determining the pre-stored template transcription factor provided by an embodiment of the present application is shown. According to different requirements, the order of the steps in the flow chart can be changed, and some steps can be omitted. The method for determining the pre-stored template transcription factor provided by the embodiments of the present application includes the following steps.

[0064] S40, obtaining the example transcription factor and the corresponding example motif of the example species from a plurality of databases.

[0065] In an embodiment of the present application, in order to provide data support for the analysis of the regulatory network of the target transcription factor, a plurality of example transcription factors and example motifs corresponding to a plurality of example species stored in a plurality of databases need to be obtained. For example, the plurality of databases include: CisBP database, JASPAR database, PlantTFDB database, and PlantCstrome database. Each of the databases is used to store example transcription factors and corresponding example motifs of different example species. The example species can be Arabidopsis thaliana, soybean, rice, wheat, and corn, etc. The example transcription factor refers to a transcription factor corresponding to the protein sequence of the example species. The example motif refers to a motif corresponding to the example transcription factor.

[0066] S41, when the plurality of example transcription factors belong to the same example species and belong to different databases, retaining any one of the plurality of example transcription factors and the example motif corresponding thereto.

[0067] In an embodiment of the present application, the plurality of databases can store example transcription factors and example motifs corresponding to the same example species. In order to remove redundant data, the example transcription factors and example motifs corresponding to the same example species belonging to different databases can be de-duplicated. Specifically, when the plurality of example transcription factors belong to the same example species and belong to different databases, any one of the plurality of example transcription factors and the example motif corresponding thereto is retained, so as to remove redundant data and improve the efficiency of subsequent transcription factor regulatory network analysis.

[0068] For example, when the CisBP database and the JASPAR database both store example transcription factors and example motifs corresponding to soybean, any one of the example transcription factors and the example motif corresponding thereto is retained.

[0069] S42, converting the example transcription factors and example motifs into a unified format to obtain the template transcription factor and the template motif corresponding thereto.

[0070] In an embodiment of the present application, in order to facilitate the storage and calling of the example transcription factors and the example motifs corresponding thereto, all the example transcription factors and the example motifs corresponding thereto can be stored in a unified format. Specifically, the example transcription factors and the example motifs can be converted into Cluster-Buster format by inputting the example transcription factors and the example motifs into Cluster-Buster software. In this way, the format of the example transcription factors and the example motifs can be unified, which facilitates the calling in the subsequent transcription factor regulatory network analysis process.

[0071] For example, the example transcription factors and the example motifs can be converted into Cluster-Buster format by inputting the example transcription factors and the example motifs into Cluster-Buster software. Figure 7The diagram shown is a flowchart of a method for determining the target motif corresponding to each target transcription factor according to an embodiment of this application. The order of steps in this flowchart can be changed, and some steps can be omitted, depending on different requirements. The method for determining the target motif corresponding to each target transcription factor provided in this embodiment of the application includes the following steps.

[0072] S50, calculate the homology score between each target transcription factor and each template transcription factor.

[0073] In one embodiment of this application, to determine the target motif corresponding to a target transcription factor in a species to be analyzed, a homology score between each target transcription factor and each template transcription factor is first calculated, and the target motif corresponding to the target transcription factor is determined based on the homology score. Specifically, the homology score is obtained by comparing the sequence of the target transcription factor with the sequence of the template transcription factor. The homology score can characterize the similarity between the target transcription factor and the template transcription factor, and can reflect the evolutionary relationship and common function between the target transcription factor and the template transcription factor. By analyzing the homology scores of the target transcription factor and the template transcription factor, their roles and interrelationships in gene expression regulation can be understood, which helps to reveal the complex mechanisms of gene expression.

[0074] In one embodiment of this application, the method for calculating homology scores includes: dividing the target transcription factor to obtain a plurality of first sub-sequences corresponding to the target transcription factor; dividing the template transcription factor to obtain a plurality of second sub-sequences corresponding to the template transcription factor; when any first sub-sequence is identical to any second sub-sequence, determining the length of the first sub-sequence as a candidate homology score; and determining the maximum value among the candidate homology scores as the homology score between the target transcription factor and the template transcription factor.

[0075] S51, the template motif of the template transcription factor corresponding to the highest homology score is identified as the target motif.

[0076] In one embodiment of this application, a high homology score indicates a high similarity between the sequence of the target transcription factor and the sequence of the template transcription factor, suggesting that the target transcription factor and the template transcription factor may have similar functions or jointly regulated genes. Therefore, for any given target transcription factor, the motif of the template transcription factor corresponding to the highest homology score is identified as the target motif.

[0077] like Figure 8 The diagram shown is a flowchart of a method for determining the regulator of a target transcription factor according to an embodiment of this application. The order of steps in this flowchart can be changed, and some steps can be omitted, depending on different needs. The method for determining the regulator of a target transcription factor provided in this embodiment includes the following steps.

[0078] S70, calculating an enrichment score of the target transcription factor in the corresponding target gene.

[0079] In an embodiment of the present application, for a target gene, the number and proportion of genes belonging to a certain transcription factor family are calculated, and then the enrichment degree of the transcription factor family in the given gene list is calculated using these statistical data. The enrichment score can reflect the significance of a certain transcription factor family in the given gene list, and further reveal that these genes can be regulated by the transcription factor family. Specifically, the way to calculate the enrichment score can refer to the detailed description of the flowchart shown in the following. Figure 9

[0080] S71, determining the binding site of the target transcription factor and the corresponding target gene.

[0081] In an embodiment of the present application, the binding site of the target transcription factor and the corresponding target gene refers to the region where the target transcription factor binds to the specific sequence of the target gene. The target transcription factor can regulate the transcription and expression of the target gene by recognizing and binding to the specific sequence of the target gene.

[0082] In an embodiment of the present application, the method for determining the binding site of the target transcription factor and the corresponding target gene includes any one of the following: by aligning the sequences of the target transcription factor and the target gene, the similarity and matching degree between them are determined, so as to predict the position of the binding site; by analyzing the target motif corresponding to the target transcription factor, the binding site of the target gene and the target transcription factor is predicted; using a database search tool (for example, JASPAR tool, TRANSFAC tool, etc.) to search and download the binding site information of the target transcription factor and the target gene which has been predicted.

[0083] S72, when the target motif corresponding to the target transcription factor in the co-expression module is upstream of the binding site of the target gene, and the enrichment score is greater than a preset threshold, determining that the co-expression module is a regulator.

[0084] In an embodiment of the present application, when it is determined that the target motif corresponding to the target transcription factor in the co-expression module is upstream of the transcription start site of the target gene, and the enrichment score of the target transcription factor in the target gene is greater than a preset threshold, it indicates that the target transcription factor can directly regulate the expression of the corresponding target gene, and therefore, the co-expression module can be determined as a regulator.

[0085] ​In one embodiment of this application, candidate target genes can be screened by the enrichment ranking of the motif corresponding to the target transcription factor upstream of the target gene. Specifically, candidate target genes can be screened based on the association list of the target transcription factor and its corresponding motif, the list of the target transcription factor and its target genes, and the enrichment ranking position of each motif on each gene in the whole genome. Determining the enrichment ranking position of each motif on each gene in the whole genome includes: calculating the sequence similarity between the motif (i.e., the sequence characteristics of the binding site between the transcription factor and the target gene) and the sequences (DNA sequences) upstream of the transcription start sites of all genes in the species to be analyzed (e.g., non-model plants) using the Cluster-Buster algorithm; sorting all genes of the species to be analyzed in descending order of sequence similarity to obtain the cisTarget database; wherein, genes with higher sequence similarity are ranked higher; the enrichment ranking of the motif corresponding to the target transcription factor is stored in the cisTarget database in the form of a matrix, where each row of the matrix corresponds to a motif, each column of the matrix corresponds to a gene, and each element in the matrix is ​​used to characterize the enrichment score of the target transcription factor, i.e., the enrichment ranking of a target transcription factor motif upstream of a certain gene. For example, if the element in the 4th row and 4th column of the cisTarget database is 1, it indicates that the motif in the 4th row is ranked 1st in enrichment upstream of the gene in the 4th column.

[0086] In one embodiment of this application, screening candidate target genes includes: for each motif, identifying the top-ranked genes among the genes corresponding to the motif as high-ranking genes (e.g., identifying the top 50 genes as high-ranking genes); if each candidate target gene of the target transcription factor appears among the high-ranking genes, then the high-ranking gene is retained; if any candidate target gene of the target transcription factor does not appear among the high-ranking genes, then the high-ranking gene is not retained; and identifying the target transcription factor and the high-ranking gene corresponding to the target transcription factor as the regulator corresponding to the target transcription factor.

[0087] like Figure 9 The diagram shown is a flowchart of a method for calculating the enrichment score of a target transcription factor according to another embodiment of this application. The order of steps in this flowchart can be changed, and some steps can be omitted, depending on different needs. The method for calculating the enrichment score of a target transcription factor provided in this embodiment of the application includes the following steps.

[0088] S80, calculate the first frequency of the target motif of the target transcription factor appearing in the target gene.

[0089] In an embodiment of the present application, first, a first binding site of a target transcription factor and a target gene is determined, the first binding site being a DNA fragment in the target gene that binds to the target transcription factor, and the length of the first binding site can be in the range of 5-20 base pairs. A DNA fragment with the same length as the first binding site is extracted from the target gene as a candidate binding site. The number of occurrences of a motif of the target transcription factor in the candidate binding site is counted, and the ratio between the number and the total number of candidate binding sites is calculated to obtain a first frequency of the target motif of the target transcription factor in the target gene.

[0090] S81, a second frequency of the target motif of the target transcription factor in the protein sequence is calculated.

[0091] In an embodiment of the present application, a second binding site of a target transcription factor and a protein sequence of a species to be analyzed is determined, and a subsequence with the same length as the second binding site is extracted from the protein sequence. The number of occurrences of a motif of the target transcription factor in the subsequence is counted, and the ratio between the number and the total number of subsequence is calculated to obtain a second frequency of the target motif of the target transcription factor in the protein sequence.

[0092] S82, the ratio between the first frequency and the second frequency is calculated to obtain an enrichment score of the target transcription factor in the corresponding target gene.

[0093] In an embodiment of the present application, the higher the enrichment score, the higher the importance of the target transcription factor in the target gene, and the target transcription factor plays an important role in the transcription and expression of the target gene. For example, when the first frequency is 0.48 and the second frequency is 0.6, the enrichment score is 0.8.

[0094] See Figure 10 , Figure 10 is a functional module diagram of a transcription factor analysis device provided in an embodiment of the present application. The transcription factor analysis device 11 includes a receiving module 110 and an analysis module 111. The module / unit referred to in the present application refers to a series of computer readable instruction segments that can be executed by the processor 13 and can complete a fixed function, which is stored in the memory 12. In the present embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.

[0095] The receiving module 110 is configured to receive a protein sequence of a species to be analyzed.

[0096] The analysis module 111 is configured to predict a plurality of target transcription factors of the protein sequence.

[0097] The analysis module 111 is further configured to determine a target motif corresponding to each target transcription factor by performing homology analysis on the target transcription factor and the pre-stored template transcription factor;

[0098] The analysis module 111 is further configured to perform co-expression analysis on each target transcription factor and the corresponding target motif to obtain a co-expression module of each target transcription factor.

[0099] The analysis module 111 is further configured to filter the co-expression module based on the target motif to obtain a regulatory sub-module corresponding to each target transcription factor.

[0100] The analysis module 111 is further configured to obtain a transcription factor analysis result of the analyzed species by calculating an activity score of the regulatory sub-module.

[0101] In an embodiment of the present application, the analysis module 111 is further configured to obtain example transcription factors and corresponding example motifs of example species from a plurality of databases; retain any one of the example transcription factors and the corresponding example motif when a plurality of the example transcription factors belong to the same example species and belong to different databases; and convert the example transcription factors and the example motifs into a unified format to obtain the template transcription factors and the corresponding template motifs.

[0102] In an embodiment of the present application, the analysis module 111 is further configured to calculate a homology score of each target transcription factor and each template transcription factor; and determine a motif of a template transcription factor corresponding to the highest homology score as the target motif.

[0103] In an embodiment of the present application, the analysis module 111 is further configured to divide the target transcription factors to obtain a plurality of first sub-sequences corresponding to the target transcription factors; divide the template transcription factors to obtain a plurality of second sub-sequences corresponding to the template transcription factors; determine a length of any one of the first sub-sequences as a candidate homology score when any one of the first sub-sequences is identical to any one of the second sub-sequences; and determine a maximum value in the candidate homology score as the homology score of the target transcription factor and the template transcription factor.

[0104] In an embodiment of the present application, the analysis module 111 is further configured to calculate an enrichment score of the target transcription factor in the corresponding target gene; determine a binding site of the target transcription factor and the corresponding target gene; and determine the co-expression module as the regulatory sub-module when the target motif corresponding to the target transcription factor in the co-expression module is upstream of the binding site of the target gene and the enrichment score is greater than a preset threshold.

[0105] In an embodiment of the present application, the analysis module 111 is further configured to: calculate a first frequency of occurrence of a target motif of the target transcription factor in the target gene; calculate a second frequency of occurrence of the target motif of the target transcription factor in the protein sequence; and calculate a ratio of the first frequency to the second frequency to obtain an enrichment score of the target transcription factor in the corresponding target gene.

[0106] In an embodiment of the present application, the analysis module 111 is further configured to: determine the target gene containing the binding site of the target transcription factor in the regulator; and calculate an AUC score of the regulator according to the enrichment score to obtain an activity score of the regulator.

[0107] From the above technical solutions, it can be seen that, in the embodiments of the present application, homology analysis is performed on the target transcription factor of the to-be-analyzed species and the pre-stored template transcription factor to determine the motif corresponding to the target transcription factor, thereby providing data support for subsequent transcription factor regulatory network analysis of the to-be-analyzed species. Then, co-expression analysis is performed on the transcription factor and the corresponding motif to obtain a co-expression module of the transcription factor, and the co-expression module is screened based on the motif to obtain a regulator corresponding to each target transcription factor, which can improve the data quality of the target transcription factor and thus improve the efficiency of transcription factor regulatory network analysis. Finally, the activity score of the regulator is calculated to obtain the transcription factor analysis result of the to-be-analyzed species, which can realize automatic transcription factor regulatory network analysis of the to-be-analyzed species and thus improve the efficiency of transcription factor regulatory network analysis.

[0108] Please refer to Figure 11 FIG. 1 is a structural schematic diagram of an electronic device according to an embodiment of the present application. The electronic device 1 comprises a memory 12 and a processor 13. The memory 12 is configured to store computer readable instructions, and the processor 13 is configured to execute the computer readable instructions stored in the memory to implement the transcription factor analysis method according to any one of the above embodiments.

[0109] In an embodiment of the present application, the electronic device 1 further comprises a bus, a computer program stored in the memory 12 and executable on the processor 13, such as a transcription factor analysis program.

[0110] Figure 11 Only the electronic device 1 with the memory 12 and the processor 13 is shown, and those skilled in the art can understand that, Figure 11 The structure shown does not constitute a limitation on the electronic device 1, and can comprise fewer or more components than shown, or combine certain components, or different component arrangements.

[0111] In combination with Figure 2The memory 12 in the electronic device 1 stores a plurality of computer-readable instructions to implement a transcription factor analysis method. The processor 13 can execute the plurality of instructions to achieve the following: obtaining protein sequences of a species to be analyzed and a plurality of target transcription factors of the protein sequences; determining a target motif corresponding to each target transcription factor by performing homology analysis on the plurality of target transcription factors and pre-stored template transcription factors; performing co-expression analysis on each transcription factor and the corresponding target motif to obtain a co-expression module of each transcription factor; obtaining a regulatory subunit corresponding to each target transcription factor based on the target motif screening of the co-expression module; and obtaining a transcription factor analysis result of the species to be analyzed by calculating an activity score of the regulatory subunit.

[0112] Specifically, the processor 13 can refer to the specific implementation method of the above instructions Figure 2 The descriptions of related steps in corresponding embodiments are not repeated here.

[0113] Those skilled in the art can understand that the schematic diagram is only an example of the electronic device 1 and does not constitute a limitation on the electronic device 1. The electronic device 1 can be a bus type structure or a star type structure. The electronic device 1 can further include more or less other hardware or software or different component arrangements, for example, the electronic device 1 can further include an input / output device, a network access device, etc.

[0114] It should be noted that the electronic device 1 is only an example. Other existing or future electronic products, such as those adaptable to the present application, should also be included in the protection scope of the present application and are hereby incorporated by reference.

[0115] The memory 12 includes at least one type of readable storage medium, which can be non-volatile or volatile. The readable storage medium includes a flash memory, a mobile hard disk, a multimedia card, a card-type memory (e.g., an SD or DX memory, etc.), a magnetic memory, a magnetic disk, an optical disk, etc. The memory 12 can be an internal storage unit of the electronic device 1 in some embodiments, such as a mobile hard disk of the electronic device 1. The memory 12 can also be an external storage device of the electronic device 1 in other embodiments, such as a plug-in mobile hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. The memory 12 can be used to store application software and various data installed in the electronic device 1, such as the code of the transcription factor analysis program, etc., and can also be used to temporarily store data that has been output or will be output.

[0116] The processor 13 can be composed of integrated circuits in some embodiments, for example, can be composed of a single packaged integrated circuit, or can be composed of multiple packaged integrated circuits of the same function or different functions, including one or more central processing units (CPU), microprocessors, digital processing chips, graphics processors, and combinations of various control chips, etc. The processor 13 is the control core of the electronic device 1, which connects all components of the electronic device 1 through various interfaces and lines, executes programs or modules stored in the memory 12 (such as executing transcription factor analysis programs, etc.), and calls data stored in the memory 12, to execute various functions of the electronic device 1 and process data.

[0117] The processor 13 executes the operating system of the electronic device 1 and various installed application programs. The processor 13 executes the application programs to implement the steps in each of the above transcription factor analysis method embodiments, for example Figure 2 The steps shown.

[0118] For example, the computer program can be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to complete the present application. The one or more modules / units can be a series of computer-readable instruction segments that can complete a specific function, which are used to describe the execution process of the computer program in the electronic device 1. For example, the computer program can be divided into a receiving module 110 and an analysis module 111.

[0119] The integrated units implemented in the form of software function modules described above can be stored in a computer-readable storage medium. The software function modules described above are stored in a storage medium, including a number of instructions for causing a computer device (which can be a personal computer, a computer device, or a network device, etc.) or a processor to execute part of the transcription factor analysis method described in each embodiment of the present application.

[0120] The integrated modules / units of the electronic device 1, if implemented in the form of software function units and sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above-mentioned embodiments can also be instructed by a computer program to complete related hardware devices, and the computer program can be stored in a computer-readable storage medium. The computer program, when executed by a processor, can implement the steps of each of the above method embodiments.

[0121] The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory, and other memories, etc.

[0122] Further, the computer readable storage medium can mainly include a storage program area and a storage data area, wherein the storage program area can store an operating system, at least one application required by a function, etc.; and the storage data area can store data created according to the use of the blockchain node, etc.

[0123] The bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For convenience of representation, only one arrow is used in the figure, but it does not mean that there is only one bus or only one type of bus. The bus is arranged to realize the connection and communication between the memory 12, the at least one processor 13, etc. Figure 11

[0124] The embodiment of the application further provides a computer readable storage medium (not shown in the figure), and the computer readable storage medium stores computer readable instructions. The computer readable instructions are executed by a processor in an electronic device to realize the transcription factor analysis method in any of the above embodiments.

[0125] In several embodiments provided in the application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules is only a logical function division. In actual implementation, there can be another division manner.

[0126] The modules described as separated components can or can not be physically separated, and the components displayed as modules can or can not be physical units, i.e. can be located in one place, or can be distributed on multiple network units. According to actual needs, some or all of the modules can be selected to achieve the purpose of the embodiment scheme.

[0127] ​In addition, each function module in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of hardware plus software function modules.

[0128] In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in the description can also be provided by one unit or device, either by software or hardware. The terms first, second, etc. are used to distinguish names and not to indicate any specific order.

[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit the present application. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present application.

Claims

1. A transcription factor analysis method applied to an electronic device, comprising: The method comprises: obtaining protein sequences of a species to be analyzed and a plurality of target transcription factors of the protein sequences; determining a target motif corresponding to each target transcription factor by performing homology analysis on the plurality of target transcription factors and pre-stored template transcription factors; performing co-expression analysis on each target transcription factor and the corresponding target motif to obtain a co-expression module of each target transcription factor; filtering the co-expression module based on the target motif to obtain a regulator corresponding to each target transcription factor; obtaining a transcription factor analysis result of the species to be analyzed by calculating an activity score of the regulator.

2. The transcription factor analysis method of claim 1, wherein, The method for obtaining the pre-stored template transcription factor comprises: obtaining example transcription factors of example species and corresponding example motifs from a plurality of databases; when a plurality of example transcription factors belong to the same example species and belong to different databases, retaining any one of the plurality of example transcription factors and the corresponding example motif; converting the example transcription factors and example motifs into a unified format to obtain the template transcription factors and corresponding template motifs.

3. The transcription factor analysis method of claim 1, wherein, The template transcription factor corresponds to a template motif, and the determination of a target motif corresponding to each target transcription factor by performing homology analysis on the plurality of target transcription factors and pre-stored template transcription factors comprises: calculating a homology score of each target transcription factor and each template transcription factor; determining the motif of the template transcription factor corresponding to the highest homology score as the target motif.

4. The transcription factor analysis method of claim 3, wherein, The calculation of the homology score of each target transcription factor and each template transcription factor comprises: dividing the target transcription factor to obtain a plurality of first sub-sequences corresponding to the target transcription factor; dividing the template transcription factor to obtain a plurality of second sub-sequences corresponding to the template transcription factor; when any one of the first sub-sequences is the same as any one of the second sub-sequences, determining the length of the first sub-sequence as a candidate homology score; determining the maximum value in the candidate homology score as the homology score of the target transcription factor and the template transcription factor.

5. The transcription factor analysis method of claim 1, wherein, The co-expression module comprises target genes of the target transcription factor, and the filtering of the co-expression module based on the target motif to obtain a regulator corresponding to each target transcription factor comprises: calculating an enrichment score of the target transcription factor in the corresponding target gene; determining a binding site of the target transcription factor and the corresponding target gene; when the target motif corresponding to the target transcription factor in the co-expression module is upstream of the binding site of the target gene, and the enrichment score is greater than a preset threshold, determining that the co-expression module is a regulator.

6. The transcription factor analysis method of claim 5, wherein, The calculation of the enrichment score of the target transcription factor in the corresponding target gene comprises: calculating a first frequency of the target motif of the target transcription factor in the target gene; calculating a second frequency of the target motif of the target transcription factor in the protein sequence; calculating a ratio of the first frequency and the second frequency to obtain the enrichment score of the target transcription factor in the corresponding target gene.

7. The transcription factor analysis method of claim 5, wherein, The transcription factor analysis result of the to-be-analyzed species is obtained by calculating the activity score of the regulator. The target gene containing the binding site of the target transcription factor in the regulator is determined. The AUC score of the regulator is calculated according to the enrichment score, and the activity score of the regulator is obtained.

8. A transcription factor analysis device, characterized by, The device comprises: a receiving module configured to receive a protein sequence of a to-be-analyzed species; an analysis module configured to predict a plurality of target transcription factors of the protein sequence; The analysis module is further configured to determine a target motif corresponding to each target transcription factor by performing homology analysis on the plurality of target transcription factors and a pre-stored template transcription factor. The analysis module is further configured to perform co-expression analysis on each target transcription factor and the corresponding target motif to obtain a co-expression module of each target transcription factor. The analysis module is further configured to filter the co-expression module based on the target motif to obtain a regulator corresponding to each target transcription factor. The analysis module is further configured to obtain the transcription factor analysis result of the to-be-analyzed species by calculating the activity score of the regulator.

9. An electronic device, comprising: The electronic device comprises: a memory storing computer readable instructions; and a processor executing the computer readable instructions stored in the memory to implement the transcription factor analysis method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer readable instructions, and the computer readable instructions are executed by the processor to implement the transcription factor analysis method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Interfering with hd-zip transcription factor repression of gene expression to produce plants with enhanced traits

    CN105611828A

  • Corn nitrogen response gene regulatory element analysis method

    CN115547411A