Gene transcription regulation mechanism prediction method and device, electronic equipment and storage medium
By constructing a matrix of regulatory capacity parameters for transcription factors and target genes, merging multi-omics data, and utilizing sparse optimization and L1 regularization methods, the problem of predicting transcriptional regulation mechanisms in non-model organisms was solved, achieving accurate prediction of transcriptional regulatory networks and improving gene regulation research in non-model organisms.
Patent Information
- Application Number
- CN202510045871.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-01-13
AI Technical Summary
Existing technologies are insufficient to effectively reveal gene transcriptional regulation mechanisms in non-model organisms from high-throughput sequencing data. In particular, the lack of prior knowledge about gene regulation, low sequence similarity in phylogenetic studies, and noise in multi-omics data make it difficult to study transcriptional regulatory networks.
By constructing a matrix of regulatory capacity parameters for transcription factors and target genes, merging the expression matrices of transcription factors and target genes, and utilizing sparse optimization and L1 regularization methods, combined with multi-omics data from model organisms and non-model organisms, the transcriptional regulatory network is inferred, and a method for predicting gene transcriptional regulation mechanisms is constructed.
It enables accurate prediction of transcriptional regulatory networks in different species, improves the understanding of transcriptional regulatory mechanisms in non-model organisms, reduces the false positive rate, and enhances the prediction accuracy of transcriptional regulatory networks.
Smart Images

Figure CN119964641B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of biological technology, and in particular to methods, devices, electronic devices and storage media for predicting gene transcription regulation mechanisms. Background Technology
[0002] Transcriptional regulation refers to the process by which transcription factors (TFs) bind to DNA regulatory elements, activating or inhibiting the transcription of their target genes. Epigenetic factors are also important regulators; for example, histone modifications can control the accessibility of DNA regulatory elements to TFs. The interactions between all TFs, DNA regulatory elements, and their target genes form transcriptional regulatory networks (TRNs). Unveiling the regulatory mechanisms of TRNs is crucial for understanding the molecular mechanisms of biological processes in various species.
[0003] With the development of high-throughput sequencing technology, more and more genomic, transcriptomic, and functional genomic data of species have been sequenced, providing crucial foundational data for further research on transcriptional regulatory mechanisms (TRNs). However, revealing the transcriptional regulatory mechanisms (TRMs) of species from these multi-omics data obtained through sequencing remains a significant challenge. Summary of the Invention
[0004] The main objective of this application is to propose a method, device, electronic device, and storage medium for predicting gene transcription regulation mechanisms, so as to predict gene transcription regulation mechanisms.
[0005] To achieve the above objectives, one aspect of this application proposes a method for predicting gene transcriptional regulation mechanisms, the method comprising the following steps:
[0006] Obtain transcription factors from different species and their corresponding target genes;
[0007] Calculate the regulatory capacity parameter matrix of each transcription factor and target gene pair; wherein each transcription factor and its corresponding target gene constitute a transcription factor and target gene pair.
[0008] The transcription factor expression matrix and target gene expression matrix are determined based on the regulatory ability parameter matrix.
[0009] Construct a linear relationship between the transcription factor expression matrix and the target gene expression matrix;
[0010] Selecting several transcription factors to minimize the difference in the linear relationship, thereby constructing a sparse optimization problem;
[0011] The sparse optimization problem is used to search for sparse solutions, and then the transcriptional regulatory networks of different species are determined based on the sparse solutions; wherein the transcriptional regulatory network includes the transcription factors, DNA regulatory elements, and the target genes corresponding to the transcription factors;
[0012] Predict the gene transcriptional regulation mechanism for the corresponding species based on each of the stated transcriptional regulatory networks.
[0013] In some embodiments, obtaining transcription factors from different species and their corresponding target genes includes the following steps:
[0014] The transcription factors and their binding motifs were obtained from model organisms and non-model organisms in different tissues and at different developmental stages from publicly available databases.
[0015] The active promoters and enhancers of each transcription factor are identified based on transcription-related histone modifications;
[0016] The binding sites of each transcription factor are predicted based on the active promoter and the enhancer; wherein the target gene corresponding to each transcription factor includes the corresponding binding motif and the binding site.
[0017] In some embodiments, calculating the regulatory capacity parameter matrix of each transcription factor-target gene pair includes the following steps:
[0018] The regulatory capacity parameters of each transcription factor and target gene pair are calculated based on the expression of the regulatory capacity parameter.
[0019] The regulation capability parameter matrix is determined based on each of the regulation capability parameters.
[0020] The expression for the regulation capability parameter is:
[0021] ;
[0022] in, Indicates the first The transcription factor and the first The regulatory capacity parameters of the target genes; / Indicates the first The first of the transcription factors The binding site to the first The ratio of the distance to the transcription start site of each target gene to a constant distance; It is the intensity of transcription-related histone modification markers, used to indicate the accessibility of the DNA regulatory element to the binding of the corresponding transcription factor; The number of binding sites on the target gene. is the sequence number of the binding site.
[0023] In some embodiments, determining the transcription factor expression matrix and target gene expression matrix based on the regulatory capacity parameter matrix includes the following steps:
[0024] The transcription factor expression data of all the species in the regulatory capacity parameter matrix are merged into the transcription factor expression matrix;
[0025] The target gene expression data of all the species in the regulatory capacity parameter matrix are merged into the target gene expression matrix.
[0026] In some embodiments, constructing the linear relationship between the transcription factor expression matrix and the target gene expression matrix includes the following steps:
[0027] The linear relationship between the transcription factor expression matrix and the target gene expression matrix is constructed as follows:
[0028] ;
[0029] in, , This represents the transcription factor expression matrix; , This represents the target gene expression matrix; , Represents the error matrix; , Represents the adjustment matrix; m The number of the species, r The number of the transcription factors, n The number of the target genes.
[0030] In some embodiments, selecting several transcription factors to minimize the difference in the linear relationship, thereby constructing a sparse optimization problem, includes the following steps:
[0031] By selecting fewer than a threshold number of the transcription factors to minimize the difference in the linear relationship, the sparse optimization problem is constructed as follows:
[0032] ;
[0033] ;
[0034] in, The Euclidean norm is ; express The number of non-zero items; express The sparsity of is used to represent the adjustment of the first . The first of the target genes The number of the transcription factors.
[0035] In some embodiments, the search for a sparse solution based on the sparse optimization problem includes the following steps:
[0036] The sparse optimization problem is transformed into an L1 regularization problem;
[0037] The problem of converting to L1 regularization is as follows:
[0038] ;
[0039] in, = ;
[0040] The sparse solution is obtained by searching based on the L1 regularization problem.
[0041] The determination of transcriptional regulatory networks for different species based on the sparse solution includes the following steps:
[0042] The transcriptional regulatory networks of different species are determined using an identification model based on the sparse solution;
[0043] The expression for the recognition model is:
[0044] ;
[0045] in, Indicates the first The transcription factor expression matrix of each of the species, Indicates the first The expression matrix of the target genes of each of the species; Indicates the first The regulatory capacity parameter matrix of all transcription factor and target gene pairs in the species; Indicates the first Among the species described, the species with the shortest divergence time in the phylogenetic tree; Indicates the first The transcriptional regulatory networks in the aforementioned species; Representing the The species mentioned above, number 1 The transcription factor mentioned above affects the first The inferred regulatory strength of the target genes, when When it is not 0, the first The aforementioned transcription factors serve as regulatory factors; This is an adjustment parameter used to balance the relative importance of prior knowledge; For the first Regularization parameters for each of the species; For the first The transcriptional regulatory network irrelevance penalty parameter for each of the species; k The number of species, which includes model organisms and non-model organisms.
[0046] To achieve the above objectives, another aspect of this application provides a gene transcription regulation mechanism prediction device, the device comprising:
[0047] The data acquisition unit is used to acquire transcription factors from different species and the target genes corresponding to the transcription factors;
[0048] A matrix calculation unit is used to calculate the regulatory capacity parameter matrix of each transcription factor and target gene pair; wherein each transcription factor and its corresponding target gene constitute a transcription factor and target gene pair.
[0049] A matrix determination unit is used to determine the transcription factor expression matrix and the target gene expression matrix based on the regulatory ability parameter matrix.
[0050] A linear relationship construction unit is used to construct a linear relationship between the transcription factor expression matrix and the target gene expression matrix;
[0051] A problem construction unit is used to select several transcription factors to minimize the difference in the linear relationship, thereby constructing a sparse optimization problem;
[0052] A network determination unit is used to search for sparse solutions based on the sparse optimization problem, and then determine the transcriptional regulatory networks of different species based on the sparse solutions; wherein the transcriptional regulatory network includes the transcription factors, DNA regulatory elements, and the target genes corresponding to the transcription factors;
[0053] The mechanism prediction unit is used to predict the gene transcriptional regulation mechanism of the corresponding species based on each of the transcriptional regulatory networks.
[0054] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method.
[0055] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0056] The embodiments of this application include at least the following beneficial effects:
[0057] This application can obtain transcription factors and their corresponding target genes from different species; calculate the regulatory capacity parameter matrix of each transcription factor and target gene pair; wherein each transcription factor and its corresponding target gene are considered as a transcription factor and target gene pair; determine the transcription factor expression matrix and the target gene expression matrix based on the regulatory capacity parameter matrix; construct a linear relationship between the transcription factor expression matrix and the target gene expression matrix; select several transcription factors to minimize the difference in the linear relationship, thereby constructing a sparse optimization problem; search for sparse solutions based on the sparse optimization problem, and then determine the transcriptional regulatory network of different species based on the sparse solutions; wherein the transcriptional regulatory network includes transcription factors, DNA regulatory elements, and the target genes corresponding to the transcription factors; and predict the gene transcriptional regulation mechanism of the corresponding species based on each transcriptional regulatory network. This application, by using transcription factors and their corresponding target genes from different species, and then constructing a linear relationship between the transcription factor expression matrix and the target gene expression matrix, can determine the transcriptional regulatory network based on this linear relationship, thereby determining the gene transcriptional regulation mechanism, which can be used to reveal the gene transcriptional regulation mechanism of a species. Attached Figure Description
[0058] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0059] Figure 1 A schematic flowchart illustrating the gene transcription regulation mechanism prediction method provided in this application embodiment;
[0060] Figure 2 A flowchart of a bioinformatics method for predicting TRN in non-model organisms using multi-omics data provided in embodiments of this application;
[0061] Figure 3 The performance graph of fused pLASSO provided in the embodiments of this application;
[0062] Figure 4 This is a schematic diagram of the structure of the gene transcription regulation mechanism prediction device provided in the embodiments of this application;
[0063] Figure 5 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0064] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0065] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”
[0066] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.
[0067] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0068] Before providing a detailed description of the embodiments of this application, some related technologies that may be involved in the embodiments of this application will be described first, as follows:
[0069] Transcriptional regulation refers to the process by which transcription factors (TFs) bind to DNA regulatory elements, activating or inhibiting the transcription of their target genes. Epigenetic factors are also important regulators; for example, histone modifications can control the accessibility of DNA regulatory elements to TFs. The interactions between all TFs, DNA regulatory elements, and their target genes form transcriptional regulatory networks (TRNs). Unveiling the regulatory mechanisms of TRNs is crucial for understanding the molecular mechanisms of biological processes in various species.
[0070] With the development of high-throughput sequencing technology, an increasing number of genomes, transcriptomes, and functional genomes of various species have been sequenced, providing crucial foundational data for further research on transcriptional regulatory mechanisms (TRMs). However, revealing the transcriptional regulatory mechanisms (TRMs) of species from these multi-omics data remains a significant challenge, especially for non-model organisms. Non-model organisms possess many characteristics not found in model organisms, lacking sufficient prior knowledge of gene regulation; low sequence similarity between phylogenetically distant species hinders functional annotation of gene regulatory elements; and noise and sequencing bias in multi-omics data can also affect the exploration of TRMs. To understand TRMs, it is necessary to identify TRMs, DNA regulatory elements, and their target genes from these omics data. Currently, bioinformatics tools are available to predict TFs from the genome, but research on DNA regulatory elements and TRNs mainly focuses on model organisms. For non-model organisms, two methods are typically used: (1) comparative analysis, which detects conserved non-coding regions in multiple species, which may be functional elements; and (2) functional analysis, which uses high-throughput sequencing technologies such as ChIP-seq to detect TF binding signals or transcription-related epigenetic markers to identify active DNA regulatory elements in vivo. However, many studies have shown that conserved non-coding regions detected through comparative analysis are not necessarily functional, and functional DNA regulatory elements are not necessarily conserved. Therefore, although researchers have attempted to use the conservation of sequences between genes to identify DNA regulatory elements in phylogenetically similar species, they may only identify a small subset, with a high false positive rate.
[0071] In recent years, the abundant TFs-DNA interactome and epigenomic data obtained from ChIP-seq sequencing have deepened our understanding of model organisms TRNs. However, ChIP-seq results are often complicated by noise generated during immunoprecipitation; and because the detected DNA regulatory elements only have sequence information, the interactions between TFs, DNA regulatory elements, and downstream targets cannot be directly determined, as the DNA regulatory elements and target genes may be thousands of bases apart.
[0072] Therefore, some embodiments of this application provide a bioinformatics scheme (fused pLASSO) to address the aforementioned problems and explore potential TRMs in organisms based on multi-species or multi-source multi-omics data. Multiple types of omics data from different sources can be cross-validated, reducing erroneous signals between them, and the complex interactions of TRNs can be interpreted by studying the cross-information between different types of regulators. Although TRNs in non-model organisms are largely unknown, they can be inferred by considering conserved genes, gene expression, and the interactions of TRNs between well-studied model species and target non-model species. Studies have shown that transcriptional regulatory interactions are more conserved than sequences. For example, in eukaryotes, the binding domains of TFs to DNA and their corresponding TF-binding motifs co-evolve and are highly conserved. Other studies have also shown that the interactions between TFs, DNA regulatory elements, and their target genes (i.e., TRNs) are more conserved than sequences in different mammals.
[0073] Therefore, the inventors of this application believe that conservative TRNs can help transform TRM knowledge from model organisms into TRM knowledge from non-model organisms. Consequently, some embodiments of this application provide a bioinformatics method that combines multi-omics data from model and non-model organisms, utilizing the vast amount of knowledge related to current model organism TRNs, to infer TRNs from non-model organisms.
[0074] This application provides a method, apparatus, electronic device, and storage medium for predicting gene transcription regulation mechanisms. The technical solution includes: obtaining transcription factors and their corresponding target genes from different species; calculating the regulatory capacity parameter matrix of each transcription factor and target gene pair; wherein each transcription factor and its corresponding target gene are considered as a transcription factor and target gene pair; determining the transcription factor expression matrix and target gene expression matrix based on the regulatory capacity parameter matrix; constructing a linear relationship between the transcription factor expression matrix and the target gene expression matrix; selecting several transcription factors to minimize the difference in the linear relationship, thereby constructing a sparse optimization problem; searching for sparse solutions based on the sparse optimization problem, and then determining the transcriptional regulation network of different species based on the sparse solution; wherein the transcriptional regulation network includes transcription factors, DNA regulatory elements, and their corresponding target genes; and predicting the gene transcription regulation mechanism of the corresponding species based on each transcriptional regulation network. This application, by using transcription factors and their corresponding target genes from different species, constructing a linear relationship between the transcription factor expression matrix and the target gene expression matrix, and determining the transcriptional regulation network based on this linear relationship, can determine the gene transcription regulation mechanism, which can be used to reveal the gene transcription regulation mechanism of a species.
[0075] This application provides methods, devices, electronic devices, and storage media for predicting gene transcription regulation mechanisms, relating to the field of biological technology. The methods, devices, electronic devices, and storage media provided in this application can be applied to terminals, servers, or software running on terminals or servers. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or vehicle terminal, but is not limited to these; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application implementing knowledge extraction methods, but is not limited to the above forms.
[0076] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0077] Reference Figure 1 This application provides a method for predicting gene transcription regulation mechanisms, which may include, but is not limited to, S100 to S160, as follows:
[0078] S100: Obtain transcription factors from different species and their corresponding target genes.
[0079] Furthermore, S100 may include S101~S103:
[0080] S101: Obtain the transcription factors and their binding motifs from model organisms and non-model organisms in different tissues and at different developmental stages from publicly available databases;
[0081] S102: Identify the active promoters and enhancers of each of the transcription factors based on transcription-related histone modifications;
[0082] S103: Predict the binding sites of each of the transcription factors based on the active promoter and the enhancer; wherein the target gene corresponding to each of the transcription factors includes the corresponding binding motif and the binding site.
[0083] S110: Calculate the regulatory capacity parameter matrix of each transcription factor and target gene pair; wherein each transcription factor and its corresponding target gene constitute a transcription factor and target gene pair.
[0084] Furthermore, S110 may include S111~S112:
[0085] S111: Calculate the regulatory capacity parameters of each transcription factor and target gene pair based on the expression of the regulatory capacity parameters.
[0086] S112: Determine the control capability parameter matrix based on each of the control capability parameters;
[0087] The expression for the regulation capability parameter is:
[0088] ;
[0089] in, Indicates the first The transcription factor and the first The regulatory capacity parameters of the target genes; / Indicates the first The first of the transcription factors The binding site to the first The ratio of the distance to the transcription start site of each target gene to a constant distance; It is the intensity of transcription-related histone modification markers, used to indicate the accessibility of the DNA regulatory element to the binding of the corresponding transcription factor; The number of binding sites on the target gene. is the sequence number of the binding site.
[0090] S120: Determine the transcription factor expression matrix and target gene expression matrix based on the regulatory ability parameter matrix.
[0091] Furthermore, S120 may include S121~S122:
[0092] S121: Merge the transcription factor expression data of all species in the regulatory capacity parameter matrix into the transcription factor expression matrix;
[0093] S122: Merge the target gene expression data of all species in the regulatory capacity parameter matrix into the target gene expression matrix.
[0094] S130: Construct a linear relationship between the transcription factor expression matrix and the target gene expression matrix.
[0095] Furthermore, S130 may include S131:
[0096] S131: The linear relationship between the transcription factor expression matrix and the target gene expression matrix is constructed as follows:
[0097] ;
[0098] in, , This represents the transcription factor expression matrix; , This represents the target gene expression matrix; , Represents the error matrix; , Represents the adjustment matrix; m The number of the species, r The number of the transcription factors, n The number of the target genes.
[0099] S140: Select several of the transcription factors to minimize the difference in the linear relationship, thereby constructing a sparse optimization problem.
[0100] Furthermore, S140 may include the following step S141:
[0101] S141: Select fewer than a threshold number of the transcription factors to minimize the difference in the linear relationship, thereby constructing the sparse optimization problem as follows:
[0102] ;
[0103] ;
[0104] in, The Euclidean norm is ; express The number of non-zero items; express The sparsity of is used to represent the adjustment of the first . The first of the target genes The number of the transcription factors.
[0105] S150: Search for sparse solutions based on the sparse optimization problem, and then determine the transcriptional regulatory networks of different species based on the sparse solutions; wherein, the transcriptional regulatory network includes the transcription factors, DNA regulatory elements, and the target genes corresponding to the transcription factors.
[0106] Further, the search for a sparse solution based on the sparse optimization problem in S150 includes the following steps S151~S152:
[0107] S151: Transform the sparse optimization problem into an L1 regularization problem;
[0108] The problem of converting to L1 regularization is as follows:
[0109] ;
[0110] in, = ;
[0111] S152: The sparse solution is obtained by searching according to the L1 regularization problem.
[0112] S150, which determines the transcriptional regulatory networks of different species based on the sparse solution, includes the following step S153:
[0113] S153: Using an identification model, determine the transcriptional regulatory networks of different species based on the sparse solutions;
[0114] The expression for the recognition model is:
[0115] ;
[0116] in, Indicates the first The transcription factor expression matrix of each of the species, Indicates the first The expression matrix of the target genes of each of the species; Indicates the first The regulatory capacity parameter matrix of all transcription factor and target gene pairs in the species; Indicates the first Among the species described, the species with the shortest divergence time in the phylogenetic tree; Indicates the first The transcriptional regulatory networks in the aforementioned species; Representing the The species mentioned above, number 1 The transcription factor mentioned above affects the first The inferred regulatory strength of the target genes, when When it is not 0, the first The aforementioned transcription factors serve as regulatory factors; This is an adjustment parameter used to balance the relative importance of prior knowledge; For the first Regularization parameters for each of the species; For the first The transcriptional regulatory network irrelevance penalty parameter for each of the species; k The number of species, which includes model organisms and non-model organisms.
[0117] S160: Predict gene transcriptional regulation mechanisms for the corresponding species based on each of the transcriptional regulatory networks.
[0118] The following section will provide a detailed introduction and explanation of the solutions in the embodiments of this application, using specific application examples.
[0119] This embodiment uses bioinformatics methods for constructing TRNs to comprehensively analyze multi-omics data from model organisms such as humans, mice, rats, fruit flies, worms, and zebrafish, providing most of the model organism omics data for the inference of non-model organism TRNs in this embodiment.
[0120] This embodiment proposes a mathematical optimization model as an identification model. This model combines multiple omics data from multiple species and applies efficient algorithms to solve problems existing in the prior art, such as the alternating direction method of multipliers (ADMM), the proximal gradient algorithm (PGA), and the linearized Bregman method (LBM), among other large-scale optimization algorithms.
[0121] like Figure 2 As shown, this embodiment proposes a bioinformatics method, fused prior least absolute shrinkage and selection operator (fused pLASSO), to predict TRN in non-model organisms lacking prior knowledge of transcriptional regulation. It begins by identifying TFs, TF-binding motifs, and TF-binding sites (TFBSs) in non-model organisms. Then, the genomes, transcriptomes, epigenomes, and predicted TFBSs of both model and non-model organisms are integrated into a fused pLASSO model with the best prior predictive performance for TRN inference.
[0122] Figure 2This is a flowchart of a bioinformatics method for predicting TRNs in non-model organisms using multi-omics data. The process begins with the identification of putative TFs, TF-binding motifs, and TFBSs. Then, the genomes of all model and non-model organisms, the predicted TFBSs, epigenomes (matrix PRS), and transcriptomes (matrixes A and B) are integrated into a fused pLASSO model for TRN inference. Matrix X represents the TRN to be inferred. This method can predict conserved TRNs across different animals as well as species-specific TRNs.
[0123] Specifically, this embodiment may include the following steps:
[0124] 1. Processing and integration of omics data.
[0125] The 'p' in fused pLASSO represents prior knowledge of transcriptional regulation derived from predicted TFBSs from active promoters and enhancers. TFs and TF-binding motifs from different species can be obtained from databases such as AnimalTFDB and CIS-BP. Active promoters and enhancers can be identified through ChIP-seq of transcription-related histone modifications, i.e., epigenomes, and TFBSs of promoters and enhancers can be predicted (Qin et al., 2011). Using the epigenome and predicted TFBSs, TF... i On genes j Potential regulatory strength (PRS):
[0126] ;
[0127] in, / Indicates transcription factor i The TFBS to target gene j The ratio of the distance to the transcription start site (TSS) to a constant distance (10kb), which is artificially set to define a region around the TSS, suggesting that the TFBS plays a regulatory role in gene transcription. The strength of histone modification markers indicates the accessibility of DNA regulatory elements to TF binding. For both model and non-model organisms, this example calculates the PRS for each TF-target gene pair. If species-specific histone modification data is unavailable, then the PRS for all sites is calculated. All are 1. This embodiment uses Indicates species tThe PRS matrix of all TF-target gene pairs in the dataset provides prior information on potential transcriptional regulation.
[0128] Since transcriptomes from different tissues and developmental stages reveal the expression relationships between TFs and target genes, combining TF binding information with transcriptome data can improve TRN inference. Therefore, the transcriptomes of each species will be merged into two expression matrices. Includes species t TFs representation data, matrix This includes expression data of the target genes. These transcriptomic data, along with PRS matrices containing genomic, epigenomic, and TF binding information from model and non-model organisms, will be projected into a fused pLASSO model for TRN inference.
[0129] 2. Infer TRN using fused pLASSO.
[0130] The relationship between TF expression data and target gene expression data can be approximated by a linear relationship as follows:
[0131] ;
[0132] in, This represents the transcription factor expression matrix of this species. Represents the target gene expression matrix. Represents the error matrix. X This represents the regulation matrix, which describes the regulatory relationship between TF expression and target gene expression in a species; m For the number of samples, r For the number of transcription factors, n This represents the number of target genes. In gene transcription regulatory network inference, for each target gene... j This embodiment aims to achieve this by selecting a small number of TFs. and Minimizing the difference between them is a sparse optimization problem, represented as:
[0133] ;
[0134] ;
[0135] in The Euclidean norm is represented as , express The number of non-zero items in China express The high sparsity indicates that the regulatory target genej transcription factors i The quantity.
[0136] Based on the above theoretical description, this embodiment can typically transform the sparse optimization problem into an L1 regularization problem, the L1 regularization model (LASSO), which is expressed as:
[0137] ;
[0138] in = .
[0139] Based on the aforementioned fundamental research theories, this study uses existing multi-omics data from model organisms to infer TRNs in non-model organisms. The algorithm needs to meet the following conditions: (1) integrated PRSs serve as prior information for potential targets of TF regulation; (2) sparse solutions are searched, as each gene is typically regulated by only a few TFs; and (3) conserved TRNs are identified to improve the accuracy of TRNs. To meet these conditions, this embodiment proposes the following fused pLASSO model to achieve [the desired effect]. k Simultaneous inference of TRN for each species:
[0140] ;
[0141] Where the matrix Includes species t TFs representation data, matrix Expression data including target genes Indicates species t PRS matrix of all TF-target gene pairs; Indicates in t Among all species, the species with the shortest divergence time in the phylogenetic tree; Indicates species t The TRN is the variable to be estimated; Representative species t TF i Target genes j The inferred regulatory intensity, when When it is not 0, TF is considered to be i As a regulatory factor; , , To adjust the parameters, k This refers to the number of all species, including both model organisms and non-model organisms.
[0142] The following three penalty terms are important for inferring the three features of TRN:
[0143] (1) The difference between prior information PRSs and inferred TRNs is evaluated, a concept originally proposed in prior LASSO (pLASSO); (2) The sparse-induced L1 norm was first proposed in LASSO and is widely used to approximate sparse solutions of linear systems; (3) This is the first time that the difference in regulatory vectors between each pair of adjacent species in a phylogenetic tree has been proposed in fused pLASSO.
[0144] In summary, fused pLASSO simultaneously restricts the source from PRS The differences in regulatory information with the transcriptome, the predicted number of TFs, and the regulatory differences between neighboring species are relatively small.
[0145] In the fused pLASSO model proposed in this embodiment, with , , To adjust the parameters, a balance is struck between the importance, sparsity, and conservatism of prior information and TRN information. The choice of parameters is crucial to the quality of the inferred TRN. In this embodiment, the following scheme will be used to update the three adjustment parameters:
[0146] (1) The subproblem of the fused LASSO penalty term in mathematical algorithms (PGA and ADMM) can be analytically represented using the thresholding operator. Due to the predefined sparsity level and the difference in regulatory vectors between the two species, the parameters... , It can be adjusted through dynamic schemes.
[0147] in >0 represents a species t The regularization parameter provides a balance between accuracy and sparsity. In this study, this embodiment employs an iterative soft thresholding algorithm to solve the L1 regularized model for gene regulatory network inference based on omics data. In each iteration, the gradient step size of the algorithm is:
[0148] ( );
[0149] Then perform the threshold calculation:
[0150] ;
[0151] inZ and X superscript n For the number of iterations, ν This represents the step size, typically 1 / 2, and the regularization parameters are iteratively updated using the algorithm described above. In order to maintain Sparsity.
[0152] For species t The TRN irrelevance penalty parameter is set as follows:
[0153] ;
[0154] For the first s Large absolute values are used to control the correlation of TRN to not exceed [a certain value]. s .
[0155] (2) In this embodiment, cross-validation technology will be used to select... η .
[0156] η To adjust parameters and balance the relative importance of data and prior knowledge. When η When pLASOO is reduced to LASSO, it is completely independent of prior knowledge; when η At that time, pLASSO relied entirely on prior knowledge to fit the model.
[0157] This embodiment will use cross-validation technology to automatically adjust based on the reliability of prior knowledge. η The value is used to improve the efficiency of the results evaluator. Since the inferred TRNs of non-model organisms are not necessarily conserved across all species, the balance of the four computational terms in the model can detect both species-specific TRNs and TRNs conserved across different species.
[0158] This embodiment will use existing real-world data to evaluate the performance of the proposed method. This embodiment uses previously published gold standard data for mESC TRN and employs the proximal gradient algorithm (PGA) to perform preliminary tests of mESC TRN on human data.
[0159] This application also provides more specific experimental embodiments, as follows:
[0160] This embodiment uses a benchmark dataset to validate that species-conserved TRNs can improve the predictive performance of TRNs. In the experiments, human and mouse data were used to predict TRNs in mouse embryonic stem cells (mESCs), and the predictive performance of fused pLASSO was evaluated using two independent gold standard mESC TRNs (Qin et al.). LASSO integrates only mouse transcriptome data, pLASSO integrates mouse ChIP-seq and transcriptome data, while fused pLASSO integrates both mouse and human ChIP-seq and transcriptome data. The three trained LASSO, pLASSO, and fused pLASSO models were validated on the mouse mESC gold standard TRN dataset. The proximal gradient algorithm (PGA) was used to calculate the predictive performance of the three models, and the area under the operating curve (AUC) was used as the evaluation metric to assess the predictive accuracy of the three models' TRNs. Figure 3 As shown, on two independent gold standard datasets, the AUC of fused pLASSO is around 0.70, which is significantly higher than the other two models. After integrating human data, fused pLASSO significantly improved the accuracy of mouse prediction of TRN.
[0161] Specifically, Figure 3 The performance of fused pLASSO is shown in the graph. The different models were compared by predicting mESC TRNs and using two independent mESC gold standard TRNs: fused pLASSO fused mouse and human ChIP-seq and transcriptome data; pLASSO integrated mouse ChIP-seq and transcriptome data; and LASSO used only mouse transcriptome data. The proximal gradient algorithm (PGA) was used to calculate the performance of the three models, and the area under the curve (AUC) represents the accuracy of TRN prediction. The results show that the evaluation results of the two independent mESC gold standard TRNs are consistent, and fused pLASSO, by integrating human data, improves the accuracy of mouse TRN prediction.
[0162] Reference Figure 4 This application also provides a gene transcription regulation mechanism prediction device, which can realize the above-mentioned gene transcription regulation mechanism prediction method. The device includes:
[0163] The data acquisition unit is used to acquire transcription factors from different species and the target genes corresponding to the transcription factors;
[0164] A matrix calculation unit is used to calculate the regulatory capacity parameter matrix of each transcription factor and target gene pair; wherein each transcription factor and its corresponding target gene constitute a transcription factor and target gene pair.
[0165] A matrix determination unit is used to determine the transcription factor expression matrix and the target gene expression matrix based on the regulatory ability parameter matrix.
[0166] A linear relationship construction unit is used to construct a linear relationship between the transcription factor expression matrix and the target gene expression matrix;
[0167] A problem construction unit is used to select several transcription factors to minimize the difference in the linear relationship, thereby constructing a sparse optimization problem;
[0168] A network determination unit is used to search for sparse solutions based on the sparse optimization problem, and then determine the transcriptional regulatory networks of different species based on the sparse solutions; wherein the transcriptional regulatory network includes the transcription factors, DNA regulatory elements, and the target genes corresponding to the transcription factors;
[0169] The mechanism prediction unit is used to predict the gene transcriptional regulation mechanism of the corresponding species based on each of the transcriptional regulatory networks.
[0170] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0171] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method for predicting gene transcription regulation mechanisms. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0172] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0173] Please see Figure 5 , Figure 5 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:
[0174] The processor 501 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0175] The memory 502 can be implemented in the form of read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 502 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 502 and is called and executed by the processor 501 to execute the gene transcription regulation mechanism prediction method of the embodiments of this application.
[0176] The input / output interface 503 is used to implement information input and output;
[0177] The communication interface 504 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0178] Bus 505 transmits information between various components of the device (e.g., processor 501, memory 502, input / output interface 503, and communication interface 504);
[0179] The processor 501, memory 502, input / output interface 503, and communication interface 504 are connected to each other within the device via bus 505.
[0180] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for predicting gene transcription regulation mechanisms.
[0181] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0182] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0183] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0184] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0185] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0186] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0187] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0188] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0189] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0190] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0191] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0192] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0193] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A method for predicting gene transcriptional regulation mechanisms, characterized in that, The method includes the following steps: Obtain transcription factors from different species and their corresponding target genes; Calculate the regulatory capacity parameter matrix of each transcription factor and target gene pair; wherein each transcription factor and its corresponding target gene constitute a transcription factor and target gene pair. The transcription factor expression matrix and target gene expression matrix are determined based on the regulatory ability parameter matrix. Construct a linear relationship between the transcription factor expression matrix and the target gene expression matrix; Selecting several transcription factors to minimize the difference in the linear relationship, thereby constructing a sparse optimization problem; The sparse optimization problem is used to search for sparse solutions, and then the transcriptional regulatory networks of different species are determined based on the sparse solutions; wherein the transcriptional regulatory network includes the transcription factors, DNA regulatory elements, and the target genes corresponding to the transcription factors; Predict the gene transcriptional regulation mechanism for the corresponding species based on each of the stated transcriptional regulatory networks.
2. The method for predicting gene transcriptional regulation mechanisms according to claim 1, characterized in that, The process of obtaining transcription factors from different species and their corresponding target genes includes the following steps: The transcription factors and their binding motifs were obtained from model organisms and non-model organisms in different tissues and at different developmental stages from publicly available databases. The active promoters and enhancers of each transcription factor are identified based on transcription-related histone modifications; The binding sites of each transcription factor are predicted based on the active promoter and the enhancer; wherein the target gene corresponding to each transcription factor includes the corresponding binding motif and the binding site.
3. The method for predicting gene transcriptional regulation mechanisms according to claim 1, characterized in that, The calculation of the regulatory capacity parameter matrix of each transcription factor and target gene pair includes the following steps: The regulatory capacity parameters of each transcription factor and target gene pair are calculated based on the expression of the regulatory capacity parameter. The regulation capability parameter matrix is determined based on each of the regulation capability parameters. The expression for the regulation capability parameter is: ; in, Indicates the first The transcription factor and the first The regulatory capacity parameters of the target genes; / Indicates the first The first of the transcription factors The binding site to the first The ratio of the distance to the transcription start site of each target gene to a constant distance; It is the intensity of transcription-related histone modification markers, used to indicate the accessibility of the DNA regulatory element to the binding of the corresponding transcription factor; The number of binding sites on the target gene. is the sequence number of the binding site.
4. The method for predicting gene transcriptional regulation mechanisms according to claim 1, characterized in that, Determining the transcription factor expression matrix and target gene expression matrix based on the regulatory capacity parameter matrix includes the following steps: The transcription factor expression data of all the species in the regulatory capacity parameter matrix are merged into the transcription factor expression matrix; The target gene expression data of all the species in the regulatory capacity parameter matrix are merged into the target gene expression matrix.
5. The method for predicting gene transcriptional regulation mechanisms according to claim 1, characterized in that, The construction of the linear relationship between the transcription factor expression matrix and the target gene expression matrix includes the following steps: The linear relationship between the transcription factor expression matrix and the target gene expression matrix is constructed as follows: ; in, , This represents the transcription factor expression matrix; , This represents the target gene expression matrix; , Represents the error matrix; , Represents the adjustment matrix; m The number of the species, r The number of the transcription factors, n The number of the target genes.
6. The method for predicting gene transcriptional regulation mechanisms according to claim 5, characterized in that, The step of selecting several transcription factors to minimize the difference in the linear relationship, thereby constructing a sparse optimization problem, includes the following steps: By selecting fewer than a threshold number of the transcription factors to minimize the difference in the linear relationship, the sparse optimization problem is constructed as follows: ; ; in, Represents the Euclidean norm; express The number of non-zero elements in the middle is used to represent the regulation of the first. The first of the target genes The number of the transcription factors.
7. The method for predicting gene transcriptional regulation mechanisms according to claim 6, characterized in that, The search for a sparse solution based on the sparse optimization problem includes the following steps: The sparse optimization problem is transformed into an L1 regularization problem; The problem of converting to L1 regularization is as follows: ; in, = ; The sparse solution is obtained by searching based on the L1 regularization problem. The determination of transcriptional regulatory networks for different species based on the sparse solution includes the following steps: The transcriptional regulatory networks of different species are determined using an identification model based on the sparse solution; The expression for the recognition model is: ; in, Indicates the first The transcription factor expression matrix of each of the species, Indicates the first The expression matrix of the target genes of each of the species; Indicates the first The regulatory capacity parameter matrix of all transcription factor and target gene pairs in the species; Indicates the first Among the species described, the species with the shortest divergence time in the phylogenetic tree; Indicates the first The transcriptional regulatory networks in the aforementioned species; Representing the The species mentioned above, number 1 The transcription factor mentioned above affects the first The inferred regulatory strength of the target genes, when When it is not 0, the first The aforementioned transcription factors serve as regulatory factors; This is an adjustment parameter used to balance the relative importance of prior knowledge; For the first Regularization parameters for each of the species; For the first The transcriptional regulatory network irrelevance penalty parameter for each of the species; k The number of species, which includes model organisms and non-model organisms.
8. A device for predicting gene transcription regulation mechanisms, characterized in that, The device includes: The data acquisition unit is used to acquire transcription factors from different species and the target genes corresponding to the transcription factors; A matrix calculation unit is used to calculate the regulatory capacity parameter matrix of each transcription factor and target gene pair; wherein each transcription factor and its corresponding target gene constitute a transcription factor and target gene pair. A matrix determination unit is used to determine the transcription factor expression matrix and the target gene expression matrix based on the regulatory ability parameter matrix. A linear relationship construction unit is used to construct a linear relationship between the transcription factor expression matrix and the target gene expression matrix; A problem construction unit is used to select several transcription factors to minimize the difference in the linear relationship, thereby constructing a sparse optimization problem; A network determination unit is used to search for sparse solutions based on the sparse optimization problem, and then determine the transcriptional regulatory networks of different species based on the sparse solutions; wherein the transcriptional regulatory network includes the transcription factors, DNA regulatory elements, and the target genes corresponding to the transcription factors; The mechanism prediction unit is used to predict the gene transcriptional regulation mechanism of the corresponding species based on each of the transcriptional regulatory networks.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Method and system for analyzing cell fate transition key transcription factors
CN112820353A
Inferrence of a gene expression profile via neural network
US20230197194A1