Cell signaling network analysis method based on transcriptional regulation logic decoding
By establishing a transcription factor regulation network and analyzing logical relationships, the problem of low accuracy of the gene regulation network is solved, and higher analysis accuracy and depth are achieved, providing new methods for biomedical research.
Patent Information
- Application Number
- PCT/CN2023/132542
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2023-11-20
- Publication Date
- 2025-05-08
AI Technical Summary
In the prior art, the accuracy of gene regulation networks is low and it is difficult to effectively decode the complex logic of cellular signaling networks.
By collecting dynamic expression data of the target gene, binding data of transcription factors in the genome and genomic data, preprocessing is performed and a transcription factor regulation network is established. Analyze the logical relationships, determine the confidence of each logical relationship, and sort and verify the transcription factor regulation network according to the confidence level.
It improves the accuracy and depth of gene regulation network analysis and modeling, reveals the complex mechanisms and key nodes of transcriptional regulation, enhances the reliability of the model, and provides new ideas and methods for biomedical research, drug research and development and genetic engineering.
Smart Images

Figure CN2023132542_08052025_PF_FP_ABST
Abstract
Description
Cell signaling network analysis method based on transcriptional regulatory logic decoding Technical Field
[0001] The present invention relates to the fields of computer science and biological networks, and in particular to a cell signaling network analysis method based on transcriptional regulation logic decoding. Background Art
[0002] In organisms, genes interact through expression and regulation to fulfill their biological functions and carry out complex life processes. These interactions are a continuous and complex dynamic process that changes over time and in response to environmental changes. Studying the evolution of gene regulatory networks has many important implications. For example, current gene regulatory relationships can be used to predict future ones, and changes in gene regulatory relationships can be used to predict changes in gene function. This can help us understand the pathogenesis of certain diseases, particularly cancer, and provide a basis for disease prediction and treatment.
[0003] Patent document CN109360607A discloses a method for analyzing the network evolution of a dynamic gene regulatory network, comprising: step S1: representing the gene regulatory network in the form of a motif, counting the motif conversion probability between snapshots, and representing the motif conversion probability between two adjacent snapshots with a matrix to obtain a motif conversion probability matrix, wherein the motif is a subgraph consisting of three nodes, and the snapshot is a static structure at a preset moment obtained by sampling the gene regulatory network at a preset time interval. The elements in the motif conversion probability matrix are used to characterize the change of the motif from one moment to the next; step S2: using the motif conversion probability matrix containing the previous T-1 moments Based on the motif transition probability tensor of the matrix, the motif transition probability matrix at time t is predicted to obtain an unsigned network snapshot, where T represents the total number of snapshots and t represents the time; step S3: extract explicit features and implicit features from the edges of the source network and the target network respectively, where the source network is a network with known signs and the target network is a gene regulatory network, and based on the implicit features, the edges of the unsigned gene regulatory network are mapped to the latent space by a preset non-negative matrix three-factor decomposition method, with the coordinates of the edge position in the latent space as features and the signs of the edges as labels, and sample training and prediction are performed through machine learning methods to obtain signed network snapshots at future moments. However, the accuracy of the existing technology for gene regulatory networks is low.
[0004] Summary of the Invention
[0005] To this end, the present invention provides a cell signaling network analysis method based on transcriptional regulatory logic decoding, which can solve the problem of low accuracy of gene regulatory networks.
[0006] To achieve the above objectives, the present invention provides a cell signaling network analysis method based on transcriptional regulatory logic decoding, comprising:
[0007] Collecting first data, second data, and third data, wherein the first data is a dynamic expression data spectrum of a target gene, the second data is binding data of a transcription factor in a genome, and the third data is genome data;
[0008] preprocessing the first data, the second data, and the third data;
[0009] Establish a transcription factor regulatory network based on the pre-processed data;
[0010] Analyzing the logical relationships, determining the confidence of each logical relationship, and sorting the logical relationships existing in the transcription factor regulatory network from large to small according to the confidence;
[0011] Using the logical relationship corresponding to the confidence level to verify whether the combined data of the transcription factor regulatory network and the genomic data are consistent with expectations, a verification conclusion is obtained;
[0012] The transcription factor regulatory network is updated based on the verification conclusion.
[0013] Furthermore, collecting dynamic expression data of target genes includes:
[0014] The transcription kinetic equation of the target gene at time point t is expressed as an ordinary differential equation: dy m (t) / dt=I s (t)-k dm y m (t)
[0015] Among them, I s (t) represents the initial transcription rate of the target gene; k dm represents the mRNA degradation rate, k dm is a constant, y m (t) represents the expression level of the target gene at time point t.
[0016] Furthermore, collecting the binding data of transcription factors in the genome includes:
[0017] When a transcription factor binds to a target gene to form a transcription complex, the transcription initiation rate is related to the probability that the promoter of the target gene is bound by the transcription factor;
[0018] Let Y(t) represent the binding rate of the target gene by the transcription factor at time point t, then the initial transcription rate I s (t) is expressed as:
[0019] Among them, I max k is the maximum transcription rate of the target gene under theoretical conditions, which is affected by the transcription complex or itself; b is the regulatory strength parameter of the transcription factor; T m is the time parameter, which is used to represent the time after the transcription factor binds m Each time unit has a regulatory effect on transcriptional activity;
[0020] When any two transcription factors that are physically close bind to the target gene and form a combination, the logical relationship between the two transcription factors includes AND, OR, and NOT;
[0021] Assume that one of any two transcription factors is A and the other is B;
[0022] The description of the logical function is:
[0023] The description of the logical function is OR is:
[0024] The description of the logical function is not:
[0025] in, represents the regulatory strength of transcription factor A alone; represents the regulatory strength of transcription factor B alone; Indicates the regulatory strength when transcription factors A and B exist simultaneously; Y(tT m ) is the regulatory intensity of transcription factors on genes after considering transcription delay; Y A (tT m ) indicates the regulatory effect of transcription factor A on the target gene; Y B (tT m ) indicates the regulatory effect of transcription factor B on target genes.
[0026] Furthermore, establishing a transcription factor regulatory network based on the pre-processed data includes:
[0027] For the regulation of multiple transcription factors, R i Represents the i-th unique control logic, and uses ω i Represents the probability of the occurrence of this logic, using I i Represents the contribution of this logic to transcriptional activity, and the transcription initiation rate of the target gene is expressed by the equation using a unique regulatory logic:
[0028] Among them, NR represents the number of unique regulatory logics involved in the target gene; for Ii, the following equation is used:
[0029] Among them, Z is a set of composite variables, Z j represents the jth variable, N z Represents the number of all regulatory logics that may be involved in regulating the target gene, a ij is the parameter of each composite variable and corresponds to the jth composite variable and the i-th logic, ε i is the truncation error of the Taylor expansion;
[0030] make and get
[0031] where β j represents a relationship between the transcription kinetic parameters and the probability of occurrence ω i The relevant functions are given only when the number of transcription factors and the Taylor expansion series are determined;
[0032] When two transcription factors (A u ,A v ) participate in regulating a gene, the contribution of this pair of transcription factors to the transcription rate is expressed as:
[0033] By applying Taylor expansion and mathematical transformation, the model equation of gene transcription regulation is obtained:
[0034] in, is time point t l The exponent, β j is the predicted expression value of the target gene in a certain period of time, Z j is the actual expression level in a certain period of time, k dm is the degradation rate of the gene, ε is the model coefficient, and is also a function of the transcription kinetic parameters; m (t l-1 ) is the transcription factor-DNA at time point t l-1 function of the occupancy rate.
[0035] Furthermore, the transcription factors include CLOCK, BMAL1, CRY1, CRY2, PER1, PER2, NPAS2, EP300, CREBBP and POLR2A.
[0036] Furthermore, based on the qTRNs of circadian genes, a gene knockout simulation method was used to evaluate the effects of a certain transcription factor or a combination of transcription factors on regulating circadian gene expression within a dynamic range. In the construction of vKO mutants, first, the LogicTRN method was used to predict the transcription factors responsible for driving gene expression regulation. The logic was set as vsKO mutants with almost single transcription factor knockout, vdKO mutants with almost double transcription factor knockout, and vtKO mutants with almost triple transcription factor knockout.
[0037] A transcription factor was removed from the qTRN to predict the dynamic profile of genes with circadian fluctuations; a transcription factor was removed to make the corresponding transcription factor-gene expression value zero; by comparing the phase shift of the dynamic curve between the original rhythm under WT conditions and the disrupted rhythm under mutant conditions, a score was calculated to represent the degree of effect of different transcription factor knockouts;
[0038] For the construction of vdKO mutants, all possible combinations of two transcription factors were calculated based on the baseline WT logic prediction data, and then, all combinations were subjected to transcriptional simulation using the LogicTRN-based framework;
[0039] For the construction of vtKO mutants, all three possible combinations were knocked out sequentially, and individual transcription factors or combinations of transcription factors were transcriptionally simulated under each condition;
[0040] Calculate the occupancy of all dynamic transcription factors.
[0041] Furthermore, the occupancy of dynamic transcription factors was calculated by the following formula:
[0042] Among them, y TF (t-τ) represents the gene expression of a specific transcription factor TF itself at the time point (t-τ), K Di,Sj Represents the site-specific affinity of transcription factor Di and transcription factor Sj, N B Represents the binding site of transcription factor Di on the target gene, N S It represents the binding site of transcription factor Sj on the target gene. The a value of the transcription factor is estimated by assuming that most genes are within the normal expression range, and a is a constant.
[0043] Furthermore, determining the confidence level of each logical relationship includes:
[0044] Dynamic gene expression data and transcription factor occupancy were used to form a series of binary logistic model equations:
[0045] Use the lasso method to obtain the beta matrix of the regression coefficients:
[0046] Calculate the confidence values of all logical relationships that regulate the target gene:
[0047] Among them, p k (i) It represents the confidence that target gene i participates in the k-th regulatory logic, i represents target gene i, r1 represents the starting number of target genes, and r2 represents the ending number of target genes.
[0048] Furthermore, the logical relationship corresponding to the confidence level is used to verify whether the combined data of the transcription factor regulatory network and the genomic data are consistent with expectations, and the verification conclusions include:
[0049] Identify the dominant logical combinations that include relationships between transcription factors;
[0050] As a result, the logic between multiple transcription factors is identified, and the logic cooperates with the abundance of the transcription factor itself to regulate the expression level of the target gene. The logic of target gene identification is used to explore the verification data of the transcriptional regulation mechanism of the transcription factor on the target gene, and the transcription factor regulatory network is verified with the verification data to obtain the verification conclusion.
[0051] Furthermore, updating the transcription factor regulatory network based on the verification conclusion includes:
[0052] By repeatedly applying the same process to each gene, the entire transcription factor regulatory network is reconstructed, which connects the logical relationships of all transcription factors and their regulated target genes;
[0053] Also included: Visualization of transcription factors and target genes in the transcription factor regulatory network using Cytoscape v3.5.1.
[0054] Compared with the existing technology, the beneficial effects of the present invention are that, by establishing a transcription factor regulatory network with high accuracy, the complex mechanism and key nodes of transcriptional regulation are revealed, and the regulatory logic relationship is analyzed and decoded by a combinational logic analysis method, thereby improving the accuracy and depth of gene regulatory network analysis and modeling; the actual experimental verification method is adopted to improve the reliability of the model; new ideas and methods are provided for in-depth understanding of the regulatory mechanism in organisms, the occurrence and development of diseases, etc.; it can be applied to biomedical research, drug development, genetic engineering and other fields, and has broad application prospects and practical value. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] FIG1 is a schematic diagram of a process for analyzing a cell signaling network based on transcriptional regulatory logic decoding according to an embodiment of the present invention;
[0056] FIG2 is a schematic diagram of a flow chart of a practical application method provided by an embodiment of the present invention;
[0057] FIG3 is a flow chart of another practical application method provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0058] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.
[0059] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0060] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside", and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.
[0061] Furthermore, it should be noted that, in the description of the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0062] Referring to FIG1 , an embodiment of the present invention provides a cell signaling network analysis method based on transcriptional regulatory logic decoding, comprising:
[0063] Step S100: collecting first data, second data, and third data, wherein the first data is a dynamic expression data spectrum of a target gene, the second data is binding data of a transcription factor in a genome, and the third data is genome data;
[0064] Step S200: pre-processing the first data, the second data, and the third data;
[0065] Step S300: establishing a transcription factor regulatory network based on the pre-processed data;
[0066] Step S400: analyzing the logical relationships, determining the confidence level of each logical relationship, and sorting the logical relationships in the transcription factor regulatory network according to the confidence level from large to small;
[0067] Step S500: using the logical relationship corresponding to the confidence level to verify whether the combined data of the transcription factor regulatory network and the genomic data are consistent with expectations, and obtaining a verification conclusion;
[0068] Step S600: updating the transcription factor regulatory network based on the verification conclusion.
[0069] Specifically, the embodiments of the present invention reveal the complex mechanisms and key nodes of transcriptional regulation by establishing a transcription factor regulatory network with high accuracy, and use a combinational logic analysis method to analyze and decode the regulatory logic relationship, thereby improving the accuracy and depth of gene regulatory network analysis and modeling; using actual experimental verification methods to improve the reliability of the model; providing new ideas and methods for in-depth understanding of the regulatory mechanisms in organisms, the occurrence and development of diseases, etc.; can be applied to biomedical research, drug development, genetic engineering and other fields, and has broad application prospects and practical value.
[0070] Specifically, collecting dynamic expression data of target genes includes:
[0071] Let the dynamic expression of the target gene at time t be expressed as an ordinary differential equation: dy m (t) / dt=I s (t)-k dm y m (t)
[0072] Among them, I s (t) is the transcription initiation rate, k dm represents the mRNA degradation rate, y m (t) represents the occupancy of the target gene in time t. The above ordinary differential equation represents the change in the target gene mRNA expression level within a specific time interval, which is equal to the initial transcription amount minus the degradation amount.
[0073] Specifically, the embodiment of the present invention uses ordinary differential equations to represent the dynamic expression of the target gene at time t, thereby achieving an accurate description of the expression of the target gene and improving the efficiency of model construction.
[0074] Specifically, collecting transcription factor binding data on the genome includes:
[0075] When a transcription factor binds to a target gene to form a genome, the transcription initiation rate is controlled by the transcription factor binding data, which is determined by the binding occupancy y at time t. m (t) and the regulatory strength Kb, the transcription initiation rate I s (t) is expressed as the following regulation function:
[0076] When any two transcription factors with logical functions combine with target genes to form a genome, the logical functions of any two transcription factors include AND, OR, and NOT;
[0077] Assume that one of any two transcription factors is A and the other is B;
[0078] The description of the logical function is:
[0079] The description of the logical function is OR is:
[0080] The description of the logical function is not:
[0081] Among them I max There is a physical limit on the transcription initiation rate, determined by the RNA extension rate and the size of the polymerase. represents the regulatory strength of transcription factor A alone; Indicates the regulatory strength of transcription factor B alone; Kb is the regulatory strength of the transcription factor, represents the regulatory strength when transcription factors A and B exist simultaneously, Y(tT m ) is the transcription delay; Y A (tT m ) represents the transcription delay Y of transcription factor A B (tT m ) indicates delayed transcription of transcription factor B.
[0082] Specifically, the embodiments of the present invention implement a cell signaling network analysis method based on transcriptional regulatory logic decoding, which comprehensively analyzes and decodes the complex logical relationships of gene regulatory networks.
[0083] Specifically, establishing a transcription factor regulatory network based on preprocessed data includes:
[0084] When two transcription factors participate in regulating a gene, the model equation of gene transcription regulation is obtained by applying Taylor expansion and mathematical transformation according to the logical functions of the two transcription factors.
[0085] in, is time point t l The exponent, β j is the predicted expression value of the target gene in a certain period of time, Z j is the actual expression level in a certain period of time, k dm is the degradation rate of the gene, ε is the model coefficient, and is also a function of the transcription kinetic parameters; m (t l-1 ) is the transcription factor-DNA at time point t l-1 The function of occupancy, N z is the number of all regulatory logics that may be involved in regulating the target gene.
[0086] Specifically, the embodiments of the present invention reveal the complex mechanisms and key nodes of transcriptional regulation by establishing a transcription factor regulatory network model, which is widely used in the modeling and analysis of gene regulatory networks, and provides a new method for in-depth understanding of the establishment and regulation mechanism of gene regulatory networks. The implementation method is simple and easy to operate, and can be implemented in various forms such as computer software and cloud computing platforms.
[0087] Specifically, the transcription factors include CLOCK, BMAL1, CRY1, CRY2, PER1, PER2, NPAS2, EP300, CREBBP and POLR2A.
[0088] Specifically, the embodiment of the present invention uses 10 core circadian transcription factors to analyze time series gene expression data and ChIP-Seq TF-DNA binding signals, and then we integrate them into the algorithm to calculate the TF-DNA binding occupancy at each time point. On this basis, the dynamic gene expression behavior is predicted, and all predicted gene expression dynamic maps are analyzed together with the original signals from the microarray experiment to represent the predicted dynamics with their original expression profiles. Based on this judgment, a predicted dynamic curve is obtained, in which the dynamic prediction is very consistent with the periodic expression pattern of the original gene expression curve within 24 hours. Since our prediction is based on 10 circadian transcription factors, these well-predicted genes are likely to play a role in circadian rhythms and have circadian rhythms or daily fluctuations.
[0089] Specifically, based on the qTRNs of circadian-fluctuating genes, a gene knockout simulation method was used to evaluate the effects of a certain transcription factor or a combination of transcription factors on regulating circadian-fluctuating gene expression within a dynamic range. In the construction of vKO mutants, first, the LogicTRN method was used to predict the transcription factors responsible for driving gene expression regulation. The logic was set as vsKO mutants with almost single transcription factor knockout, vdKO mutants with almost double transcription factor knockout, and vtKO mutants with almost triple transcription factor knockout.
[0090] A transcription factor was removed from the qTRN to predict the dynamic profile of genes with circadian fluctuations; a transcription factor was removed to make the corresponding transcription factor-gene expression value zero; by comparing the phase shift of the dynamic curve between the original rhythm under WT conditions and the disrupted rhythm under mutant conditions, a score was calculated to represent the degree of effect of different transcription factor knockouts;
[0091] For the construction of vdKO mutants, all possible combinations of two transcription factors were calculated based on the baseline WT logic prediction data, and then, all combinations were subjected to transcriptional simulation using the LogicTRN-based framework;
[0092] For the construction of vtKO mutants, all three possible combinations were knocked out sequentially, and individual transcription factors or combinations of transcription factors were transcriptionally simulated under each condition;
[0093] Calculate the occupancy of all dynamic transcription factors.
[0094] Specifically, the embodiments of the present invention determine the transcription simulation of the transcription factor or transcription factor combination formed by the mutant after the transcription factor is knocked out, calculate the occupancy of all dynamic transcription factors, facilitate the calculation of the realization efficiency of gene expression, and then timely update the regulatory network of the transcription factor, form a new transcription factor regulatory network, and increase the effective control of gene expression.
[0095] Specifically, the occupancy of dynamic transcription factors was calculated by the following formula:
[0096] Where y represents the gene expression of the specific transcription factor itself, K represents the site-specific transcription factor affinity, and N represents the binding sites of the transcription factor on the target gene. The a value of the transcription factor is estimated by assuming that most genes are within the normal expression range.
[0097] Specifically, the embodiment of the present invention calculates the occupancy of dynamic transcription factors by constructing a specific formula, thereby improving the calculation efficiency and facilitating improving the efficiency of constructing the transcription factor network.
[0098] Specifically, determining the confidence level of each logical relationship includes:
[0099] Dynamic gene expression data and transcription factor occupancy were used to form a series of binary logistic model equations:
[0100] Use the lasso method to obtain the beta matrix of the regression coefficients:
[0101] Calculate the confidence values of all logical relationships that regulate the target gene:
[0102] Specifically, β and Beta in the embodiments of the present invention have the same meaning. By calculating the effect of each virtual gene knockout on the expression of cell cycle genes, the embodiments of the present invention can identify the most critical transcription factors for the regulation of the cell circadian clock. These key transcription factors will be used to design regulatory strategies.
[0103] Specifically, the logical relationship corresponding to the confidence level is used to verify whether the combined data of the transcription factor regulatory network and the genomic data are consistent with expectations, and the verification conclusions include:
[0104] Identify the dominant logical combinations that include relationships between transcription factors;
[0105] As a result, the logic between multiple transcription factors is identified, and the logic cooperates with the abundance of the transcription factor itself to regulate the expression level of the target gene. The logic of target gene identification is used to explore the verification data of the transcriptional regulation mechanism of the transcription factor on the target gene, and the transcription factor regulatory network is verified with the verification data to obtain the verification conclusion.
[0106] Specifically, the present invention verifies the regulatory effect of the designed transcription factor combination on the cellular circadian clock. By measuring the expression of cell cycle genes and changes in the cell cycle, it is found that the combination can significantly regulate the cellular circadian rhythmicity.
[0107] Specifically, updating the transcription factor regulatory network based on the verification conclusion includes:
[0108] By repeatedly applying the same process to each gene, the entire transcription factor regulatory network is reconstructed, which connects the logical relationships of all transcription factors and their regulated target genes.
[0109] Specifically, the embodiments of the present invention achieve accurate regulation of the cellular gene expression rhythm by designing a cell circadian clock regulation method based on a transcription factor network. It has the advantages of significant regulation effect and simple operation, and provides new ideas and methods for regulating the cell circadian clock. The accuracy and reliability of the model are verified through simulation experiments and actual experiments.
[0110] Specifically, it also includes: and using Cytoscape v3.5.1 to visualize transcription factors and target genes in the transcription factor regulatory network.
[0111] Specifically, the embodiments of the present invention combine multiple information such as DNA sequences, transcription factor binding sites, epigenetic modifications, etc., and establish a transcription factor regulatory network model to comprehensively analyze and decode the logical relationship of the gene regulatory network, thereby revealing the complex mechanism and key nodes of transcriptional regulation, and visualizing the key data in the transcription factor regulatory network in a more intuitive way.
[0112] In combination with practical applications, as shown in FIG2 and FIG3 , the cell signaling network analysis method based on transcriptional regulation logic decoding provided by an embodiment of the present invention includes the following steps:
[0113] Data collection: Collect various data such as DNA sequences, transcription factor binding sites, epigenetic modifications, etc. to establish a preliminary gene regulatory network.
[0114] Determine the regulatory relationship between transcription factors: Determine the regulatory relationship between transcription factors and target genes by establishing a transcription factor regulatory network model.
[0115] Analyze logical relationships: Analyze logical relationships in gene regulatory networks through methods such as computing logic gates and combinational logic.
[0116] Typically, genes are controlled by multiple TFs in a combinatorial manner. The functions of combinatorial regulation are usually represented by AND, OR, and NOT basic logical functions. AND logic describes the situation where a gene is activated only when two TFs bind to the gene promoter at the same time, OR logic indicates that a gene can be activated by either of the two TFs independently, and NOT logic represents an inhibitory operation. These three logical functions can be described as follows:
[0117] AND (expressed as A&B):
[0118] OR (expressed as A|B):
[0119] NOT (expressed as A>B):
[0120] Among them, Is represents the gene transcription initiation rate, I max Represents the limit of gene transcription initiation rate. TF binding is determined by the binding occupancy Y(t) of transcription factors A and B at time t and the regulatory strength k b composition.
[0121] When a gene is regulated by two TFs, it can form up to six unique regulatory logics according to the definition of transcriptional regulatory logic functions discussed above. Each regulatory logic has an occupancy probability ranging from 0 to 1 in gene regulation.
[0122] Verify the model: Verify the accuracy and reliability of the model through simulation experiments and actual experiments.
[0123] Constructing a transcription factor network: Based on existing transcription factor-target gene relationship data, a transcription network containing the relationships between transcription factors and target genes is constructed. Reliable relationships are then screened through systems biology analysis.
[0124] Perform virtual gene knockout: Based on the relationship between transcription factors and target genes in the transcription network, by deleting the node of a transcription factor or target gene, we can simulate the impact of the node on the entire network.
[0125] Identification of key transcription factors: We can perform virtual gene knockouts and calculate the effect of each virtual gene knockout on cell cycle gene expression to identify the transcription factors that are most critical for cellular circadian clock regulation. These key transcription factors will be used to design regulatory strategies.
[0126] For the effect of virtual gene knockout in step c, we use the following formula to evaluate the impact score s:
[0127] Among them, LAG m and LAG w They are the phase change values of gene transcription before and after knockout relative to the actual data.
[0128] Designing transcription factor combinations: Based on the identification of key transcription factors, we can design a set of transcription factor combinations to regulate the cellular circadian clock. This combination includes three transcription factors: CLOCK, BMAL1, and CRY1. These combinations can be introduced into cells through gene editing or gene transfection.
[0129] Verification of regulatory effects: This study validated the regulatory effects of the designed transcription factor combination on the cellular circadian clock. By measuring cell cycle gene expression and changes in the cell cycle, it was found that the combination can significantly regulate the cellular circadian rhythmicity.
[0130] Through this embodiment, a cell biological clock regulation method based on the transcription factor network is designed, which realizes the accurate regulation of the cell gene expression rhythm, has the advantages of significant regulation effect and simple operation, and provides new ideas and methods for regulating the cell biological clock.
[0131] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.
[0132] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A cell signaling network analysis method based on transcriptional regulation logic decoding, characterized in that: include: Collecting first data, second data and third data, wherein the first data is a dynamic expression data spectrum of a target gene, the second data is binding data of a transcription factor in a genome, and the third data is genome data; preprocessing the first data, the second data, and the third data; Establishing a transcription factor regulatory network based on the pre-processed data; Analyzing the logical relationships, determining the confidence of each logical relationship, and sorting the logical relationships existing in the transcription factor regulatory network from large to small according to the confidence; Using the logical relationship corresponding to the confidence level to verify whether the combined data of the transcription factor regulatory network and the genome data are consistent with expectations, and obtaining a verification conclusion; The transcription factor regulatory network is updated based on the verification conclusion.
2. The cell signaling network analysis method based on transcriptional regulation logic decoding according to claim 1, characterized in that: The dynamic expression data of target genes collected include: The transcription kinetic equation of the target gene at time point t is expressed as an ordinary differential equation: of m (t) / dt=I s (t)-k dm the m (t) Among them, I s (t) represents the initial transcription rate of the target gene; k dm represents the mRNA degradation rate, k dm is a constant, y m (t) represents the expression level of the target gene at time point t.
3. The cell signaling network analysis method based on transcriptional regulation logic decoding according to claim 2, characterized in that: Collecting the binding data of transcription factors in the genome includes: After a transcription factor binds to a target gene to form a transcription complex, the transcription initiation rate is related to the probability that the promoter of the target gene is bound by the transcription factor; Let Y(t) represent the binding rate of the target gene by the transcription factor at time point t, then the initial transcription rate I s (t) is expressed as: Among them, I max k is the maximum transcription rate of the target gene under theoretical conditions, which is affected by the transcription complex or itself; b is the regulatory strength parameter of the transcription factor; T m is the time parameter, which is used to represent the time after the transcription factor binds T m The time unit has a regulatory effect on transcriptional activity; When any two transcription factors that are physically close bind to the target gene and form a combination, the logical relationship between any two transcription factors includes AND, OR, and NOT; Assume that one of any two transcription factors is A and the other is B; The description of the logic function is AND is: The description of the logical function is or is: The description of the logical function is not: in, It indicates the regulatory strength of transcription factor A alone; It indicates the regulatory strength of transcription factor B alone; Indicates the regulation when transcription factors A and B are present at the same time Section strength; Y(tT m ) is the regulatory intensity of transcription factors on genes after considering the transcription delay; Y A (tT m ) indicates the regulatory effect of transcription factor A on the target gene; Y B (tT m ) indicates the regulatory effect of transcription factor B on the target gene.
4. The cell signaling network analysis method based on transcriptional regulation logic decoding according to claim 3, characterized in that: Establishing a transcription factor regulatory network based on preprocessed data includes: For the regulation of multiple transcription factors, R i Represents the i-th unique control logic, and uses ω i Represents the probability of the occurrence of this logic, with I i Represents the contribution of this logic to transcriptional activity, and the transcription initiation rate of the target gene is expressed by the equation using a unique regulatory logic: Among them, N R represents the number of unique regulatory logics involved in the target gene; for I i , expressed using the following equation: Among them, Z is a set of composite variables, Z j represents the jth variable, N z Represents the number of all regulatory logics that may be involved in regulating the target gene, a ij is the parameter of each composite variable and corresponds to the jth composite variable and the i-th logic, ε i is the truncation error of the Taylor expansion; make and get where β j represents a relationship between transcription kinetic parameters and occurrence probability ω i The relevant functions are given only when the number of transcription factors and the Taylor expansion series are determined; When two transcription factors (A u ,A v ) participate in regulating a gene, the contribution of this pair of transcription factors to the transcription rate is expressed as: By applying Taylor expansion and mathematical transformation, the model equation of gene transcription regulation is obtained: in, is the time point t l The exponent, β j is the predicted expression value of the target gene in a certain period of time, Z j is the actual expression level in a certain period of time, k dm is the degradation rate of the gene, ε is the model coefficient, and is also a function of the transcriptional kinetic parameters; y m (t l-1 ) is the transcription factor-DNA at time point t l-1 function of the occupancy rate.
5. The cell signaling network analysis method based on transcriptional regulation logic decoding according to claim 4, characterized in that: The transcription factors include CLOCK, BMAL1, CRY1, CRY2, PER1, PER2, NPAS2, EP300, CREBBP and POLR2A.
6. The cell signaling network analysis method based on transcriptional regulation logic decoding according to claim 5, characterized in that: Based on the qTRNs of circadian genes, the method of simulating gene knockout was used to evaluate the effect of a certain transcription factor or a combination of transcription factors on regulating the expression of circadian genes within the dynamic range. In the construction of vKO mutants, first, the LogicTRN method was used to predict the transcription factors responsible for driving gene expression regulation - logic, set as almost single transcription factor knockout as vsKO mutant, set as almost double transcription factor knockout as vdKO mutant, and set as almost triple transcription factor knockout as vtKO mutant; Extracting a transcription factor from qTRN predicts the dynamics of genes with diurnal fluctuations Outline; knock out a transcription factor so that its corresponding transcription factor-gene expression value is zero; by comparing the phase shift of the dynamic curve between the original rhythm under WT conditions and the disrupted rhythm under mutant conditions, a score is calculated to represent the degree of influence of different transcription factor knockouts; For the construction of vdKO mutants, all possible combinations of two transcription factors were calculated based on the baseline WT logic prediction data, and then, all combinations were subjected to transcriptional simulation using the LogicTRN-based framework; For the construction of vtKO mutants, all three possible combinations were knocked out sequentially, and individual transcription factors or combinations of transcription factors were transcriptionally simulated in each condition; Calculate the occupancy of all dynamic transcription factors.
7. The cell signaling network analysis method based on transcriptional regulation logic decoding according to claim 6, characterized in that: The occupancy of dynamic transcription factors was calculated by the following formula: Among them, y TF (t-τ) represents the gene expression of a specific transcription factor TF itself at time point (t-τ), K Di,Sj Represents site-specific affinity of transcription factor Di and transcription factor Sj, N B represents the binding site of transcription factor Di on the target gene, N S Represents the binding site of transcription factor Sj on the target gene. The a value of the transcription factor is estimated by assuming that most genes are within the normal expression range, and a is a constant.
8. The cell signaling network analysis method based on transcriptional regulation logic decoding according to claim 7, characterized in that: Determining the confidence level of each of the logical relationships includes: Dynamic gene expression data and transcription factor occupancy were used to form a series of binary logistic model equations: Use the lasso method to obtain the beta matrix of regression coefficients: Compute confidence values for all logical relationships that regulate the target gene: Among them, p k (i) It represents the confidence that the target gene i participates in the kth regulatory logic, i represents the target gene i, r1 represents the starting number of the target gene, r2 represents the ending number based on the target, and L represents the length of the expression time.
9. The cell signaling network analysis method based on transcriptional regulation logic decoding according to claim 8, characterized in that: The logical relationship corresponding to the confidence level is used to verify whether the combined data and genomic data of the transcription factor regulatory network are consistent with expectations, and the verification conclusions include: Determine the dominant logical combination including the relationships between transcription factors; As a result, the logic between multiple transcription factors is identified, and the logic cooperates with the abundance of the transcription factor itself to regulate the expression level of the target gene. The logic identified by the target gene is used to explore the verification data of the transcription factor's transcriptional regulation mechanism on the target gene, and the transcription factor regulatory network is verified with the verification data, thereby obtaining a verification conclusion.
10. The cell signaling network analysis method based on transcriptional regulation logic decoding according to claim 9, characterized in that: The updating of the transcription factor regulatory network based on the verification conclusion comprises: By repeatedly applying the same process to each gene, the entire transcription factor regulatory network is reconstructed, which connects the logical relationships of all transcription factors and their regulated target genes; Also included: Visualization of transcription factors and target genes in the transcription factor regulatory network using Cytoscape v3.5.1.
Citation Information
Patent Citations
Microbial growth phenotype predication method based on control-metabolic network integration model
CN105184049A
A differential equation model-based gene regulation and control network construction method
CN109726352A
Gene regulation network database and application thereof in personalized drug screening
CN113130010A
Cancer type specific gene regulatory network construction method
CN115064205A
TF-DNA combination recognition method based on multi-feature fusion
CN115810398A