A Method for Constructing a Large-Scale Distributed Transcriptional Regulation Network Based on Transfer Learning
By constructing a large distributed transcriptional regulation network model and utilizing transfer learning and prior knowledge of transcriptional regulation, the problems of computational complexity and data volume limitations in the transcriptional regulation process are solved, enabling efficient prediction of transcriptional regulation relationships and research on target gene regulation.
Patent Information
- Application Number
- CN202411305129.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-19
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-09-19
AI Technical Summary
Existing technologies struggle to accurately simulate the direction of transcription factor regulation of target genes during transcriptional regulation. Furthermore, traditional models are limited by computer memory and data volume, resulting in excessively long training times and an inability to effectively utilize prior knowledge of transcriptional regulation.
We employ a large-scale distributed transcriptional regulation network model construction method based on transfer learning. We construct distributed sub-networks using prior knowledge of transcriptional regulation, pre-train and fine-tune using pan-transcriptome data, extract high-order abstract features, and optimize the model by combining result evaluation metrics.
It enables efficient use of prior knowledge of transcriptional regulation, overcomes limitations in computer memory and data volume, extracts high-order features of transcriptional regulatory relationships, improves the accuracy and efficiency of the model, and supports target gene regulation prediction.
Smart Images

Figure CN119339795B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the interdisciplinary field of biological sciences and artificial intelligence, and in particular to a method for constructing a large-scale distributed transcriptional regulatory network model based on transfer learning. Background Technology
[0002] Transcriptional regulation is one of the core regulatory mechanisms of cellular metabolic activity. In eukaryotic cells, the interactions between numerous transcription factors and target genes form an extremely large and complex transcriptional regulatory network, used to maintain cellular metabolic homeostasis and regulate cellular metabolic differentiation. Given the extreme complexity and variability of the transcriptional regulatory network in eukaryotic cells, there is an urgent need to build fundamental models of the transcriptional regulatory network at both the systemic and global levels to accurately elucidate gene regulation mechanisms and to analyze the molecular mechanisms behind important life activities such as cellular senescence and tumorigenesis at the systemic level. Therefore, the construction of fundamental models of the eukaryotic transcriptional regulatory network is of great research significance for the development of precision medicine and the health industry in my country.
[0003] Currently, methods for constructing transcriptional regulatory networks can be divided into mechanistic modeling and data-driven modeling. Mechanistic modeling involves starting from prior knowledge of a set of core transcriptional regulatory networks and applying mathematical models or signaling pathways to simulate the dynamic expression process of genes. For extensively studied cells and systems, the accumulated knowledge helps in constructing effective transcriptional regulatory network models. However, to date, the transcriptional regulatory relationships in cells are not fully understood, and how they respond to external environmental stimuli remains unknown. Therefore, mechanistic-based transcriptional regulatory network modeling strategies have significant limitations. Data-driven construction methods refer to using machine learning or statistical methods to infer transcriptional regulatory networks on a large scale based on whole-genome or transcriptomics data. This data-driven approach can infer possible gene interactions and predict gene targets affecting key functions or regulatory relationships even without sufficient prior knowledge of regulatory relationships but with abundant transcriptomics data. However, both methods have limitations and cannot accurately simulate the transcriptional regulatory process. Therefore, a fusion of mechanistic and data-driven approaches can fully utilize known prior knowledge of transcriptional regulation and better leverage artificial intelligence methods such as machine learning to model the transcriptional regulatory process in cells.
[0004] A key challenge in transcriptional regulation is identifying the regulatory direction of transcription factors towards their target genes. This determines whether to enhance or suppress the corresponding transcription factor when regulating the expression level of a target gene is necessary. However, due to the extreme complexity and high coupling of microbial cell growth processes, and the variations in regulatory relationships exhibited by the same microorganism under different experimental conditions and at different growth stages, single regulatory relationships are difficult to capture directly. Existing databases of transcriptional regulatory relationships are mostly Boolean, indicating only the existence of a regulatory relationship between the transcription factor and its target gene, while information on the direction of regulation is largely missing. Therefore, learning accurate transcriptional regulatory relationships from batches of data using machine learning methods is of great significance. It allows us to learn the intrinsic connections between data points, thereby helping to understand the regulatory direction between transcription factors and target genes, discovering gene expression regulation mechanisms, and accelerating research on transcriptional regulatory relationships. Therefore, establishing a transcriptional regulatory network model to describe intracellular transcriptional regulatory relationships plays a crucial role in cellular life science research.
[0005] The construction of large-scale models in the current biological field mainly consists of two steps: pre-training and fine-tuning, which is the idea behind transfer learning. Pre-training refers to training the model in advance without a specific task, thus using unsupervised pre-training to learn the potential relationships between genes. Fine-tuning involves initializing the pre-trained parameters and performing specific training based on the data from the downstream task. Therefore, a pan-transcriptome dataset containing multiple strains is used as the pre-training dataset to learn the transcriptional regulatory relationships between genes within a specific cell type or species. As the amount of data fed into the machine learning model increases, the model extracts higher-order, abstract, and more fundamental feature information from this data. In the fine-tuning stage of transfer learning, some parameters of the pre-trained model are fixed, while other parameters are retrained on a more specific dataset. This achieves targeted training based on the abstract features of the pre-trained model and the data of the downstream task, extracting the specific features of the downstream task and obtaining a finely tuned model.
[0006] In transcriptional regulation, a gene typically requires the joint regulation of multiple transcription factors. Furthermore, due to the vast number of genes in a cell—generally ranging from thousands to tens of thousands—the transcriptional regulatory network is an extremely large system. Therefore, when modeling it using artificial neural networks, even if we construct a complete neural network structure by combining all transcription factors and target genes, and discard some connections between layers based on prior knowledge of transcriptional regulation, the number of parameters remains enormous, exceeding the limitations of computer memory. The model computation becomes overly complex, leading to excessively long training times. Moreover, for network models with a large number of parameters, the amount of data required for training is also enormous. Summary of the Invention
[0007] The purpose of this invention is to provide a method for constructing a large model of a distributed transcriptional regulatory network based on transfer learning, so as to solve the problems mentioned in the background art.
[0008] To achieve the aforementioned objectives, this invention provides a method for constructing a large-scale distributed transcriptional regulation network model based on transfer learning. This method uses reliable prior relationships of transcriptional regulation supported by literature as the basis for selecting input and output variables for distributed subnetworks. It employs transfer learning to perform two-stage model training to obtain regulatory relationships during transcription, thereby guiding research on gene expression regulation. Establishing a large-scale distributed transcriptional regulation network model involves extracting high-order abstract features of transcriptional regulatory relationships from data using machine learning methods. The distributed structure avoids the limitations of computer memory and enables the prediction of target gene regulation, providing a reliable tool for studying regulatory relationships. To obtain specific transcriptional regulatory relationships, the method employs the concept of transfer learning, including the following steps:
[0009] Step S1: Obtain transcriptional regulatory relationships through prior knowledge of transcriptional regulation, and obtain the pairing of transcription factors with their corresponding target genes.
[0010] Step S2: Construct a large-scale distributed transcriptional regulatory network model;
[0011] Step S3: Perform model pre-training on pantranscriptome data to obtain high-order features of transcriptional regulatory relationships in various strains.
[0012] Step S4: Using a time-series dataset, fine-tune the model to obtain a specific distributed large model;
[0013] Step S5: Based on the data characteristics, formulate evaluation indicators for the prediction results, divide the prediction results into three levels of accuracy, and observe the impact of the dataset on the model prediction performance and their reliability based on the accuracy of each sub-network.
[0014] Furthermore, the prior knowledge of transcriptional regulation refers to the transcriptional regulatory relationships formed by a pair of transcription factors and their target genes. One transcription factor can regulate multiple target genes, and the same target gene may be regulated by multiple transcription factors. The prior knowledge of transcriptional regulation obtained from the database is selected from studies of transcription factors and target genes that have been reported in the literature and support this knowledge. The obtained regulatory relationships are a pair, that is, they represent the existence of a regulatory relationship between a transcription factor and its target gene, but do not include other information such as the direction of regulation. By pairing all relevant genes in the database, a knowledge graph of transcriptional regulation can be obtained, and a mechanism network of transcriptional regulation can be constructed based on this knowledge graph.
[0015] Furthermore, the large-scale distributed transcriptional regulation network model is composed of distributed sub-networks. These sub-networks are based on transcriptional regulation mechanism networks obtained from a database. The mechanism networks are separated with the target gene as the center, and the transcription factor regulation relationships of the target gene are extracted. The expression level of the transcription factor is used as the input, and the expression level of the target gene is used as the output. The distributed sub-networks are multi-input single-output. The integration of all sub-networks represents the overall transcriptional regulation relationships of the cell.
[0016] Furthermore, when the target gene is regulated by only one transcription factor or itself, the distributed subnetwork is a single-input, single-output structure.
[0017] Furthermore, the pre-training in step S3 is a process of data processing and alignment on a pan-transcriptome dataset based on the determined structure of the distributed network, followed by distributed training. This includes the following steps:
[0018] Step S301 requires processing the pre-training data. The pantranscriptome data is measured in TPM (Transcripts Per Million), a commonly used unit in RNA sequencing data analysis that represents the percentage of a particular transcript per million reads. TPM can eliminate the influence of sequencing depth on gene expression analysis. Since the pantranscriptome dataset collects sequencing data from multiple strains, although the gene similarity among strains of the same species is very high, detection errors and limitations lead to slight differences in the number of genes in the sequencing results for each strain. This means that some genes are not present in all strains. Therefore, data supplementation and alignment are necessary.
[0019] Step S302: Align all TPM data into the logarithmic space to make the data distribution more suitable for machine learning training. The processing formula can be expressed as log2(TPM+1).
[0020] Step S303: Based on the different inputs and outputs of each sub-network, the pre-training data is segmented to adapt to the features of each sub-network.
[0021] In step S304, after determining the basic parameters of the training process, each sub-network model is trained in a distributed manner, with each sub-network operating independently during training. Each sub-network is trained on a pre-processed and segmented subset of the dataset, reducing the size of the training network, improving training efficiency, and overcoming the problem of insufficient data.
[0022] Furthermore, the fine-tuning in step S4 is the process of continuing to train the pre-trained sub-models on the specific dataset of the downstream task. It includes the following steps:
[0023] Step S401: Process the fine-tuning dataset by aligning the pre-training dataset with the fine-tuning dataset in terms of gene count and gene names, and adjusting the structure of each sub-network according to the fine-tuning dataset. Since the fine-tuning dataset is based on time-series detection, and different detection batches use different sampling time intervals, an interpolation method is used to align the time intervals between samples to facilitate model fine-tuning training.
[0024] Step S402: Because the distribution of the fine-tuned dataset is different, it represents the ratio of expression levels before and after gene perturbation, and the unchanged data has been processed to set zero. Therefore, in order to align the fine-tuned data to the logarithmic space, the data processing method used is: log2(2 x +1).
[0025] Step S403: Fix some parameters of each sub-model. Based on transfer learning, freeze some parameters on the pre-training results, while other parameters can continue to be trained based on the pre-training results.
[0026] In step S404, the fine-tuning dataset is also divided to determine the input and output of each sub-network, and then fed into each sub-model for model fine-tuning using a distributed training method.
[0027] Furthermore, the evaluation index for the prediction results is formulated based on different relationships between the output of the fine-tuned model and the label values. Due to the complexity and variability of transcriptional regulation, there are errors between the predicted output values and the actual label values of the fine-tuned distributed model; these errors are acceptable within a certain range. Moreover, obtaining qualitative transcriptional regulatory relationships is one of the main functions of constructing a transcriptional regulatory network model. Therefore, based on the different manifestations of the error between the output of the fine-tuned model and the actual label values, the results are divided into three levels: excellent, acceptable, and unacceptable, representing a small error between the predicted result and the actual label, a large error between the predicted result and the actual label but in the correct direction, and a very large error between the predicted result and the actual label, respectively. Based on the customized evaluation index for prediction results, the performance of each sub-model after fine-tuning can be clearly analyzed, which can be used for further analysis of the sub-models.
[0028] Compared with existing technologies, this system and method have the following advantages:
[0029] (1) By making full use of known prior knowledge of transcriptional regulation, we have achieved mechanism-driven modeling and data fusion, and mined high-order abstract information features in the data;
[0030] (2) The massive transcriptional regulatory network is segmented based on target genes to form a distributed sub-network model framework, overcoming the limitations of computer memory, data size, and computing time.
[0031] (3) Using the idea of transfer learning, abstract features of transcriptional regulation relationships are extracted from different datasets and extracted through fine-tuning process to obtain a specific fine-tuning model for downstream tasks.
[0032] (4) Result evaluation indicators were formulated to evaluate the results of each model based on different error performance. The distributed model structure is conducive to further qualitative research on transcription factors. Attached Figure Description
[0033] Figure 1 This is a flowchart illustrating the method for constructing a large-scale distributed transcriptional regulatory network model based on transfer learning.
[0034] Figure 2 This is an example diagram illustrating prior knowledge of transcriptional regulatory relationships.
[0035] Figure 3 This is a diagram illustrating the overall framework of the distributed transcriptional regulation network model of this invention.
[0036] Figure 4 This is a diagram of the distributed transcriptional regulatory subnetwork structure of the present invention.
[0037] Figure 5 This is a diagram showing the overall results of the pre-training of the distributed transcriptional regulatory subnetwork.
[0038] Figure 6 Example image showing the interpolation results for fine-tuning the dataset.
[0039] Figure 7 This is a diagram showing the overall results of fine-tuning the distributed transcriptional regulatory subnetwork. Detailed Implementation
[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0041] like Figure 1 The diagram shows the flowchart of the method of this invention. An embodiment of this invention models the transcriptional regulation process of *Saccharomyces cerevisiae*. First, all known transcriptional regulatory relationships of *Saccharomyces cerevisiae* in the database are obtained. Then, a distributed network model framework is constructed based on these known regulatory relationships. The pre-training dataset, consisting of sequencing data from 969 different yeast strains, is processed. After pre-training, time-series data interpolated at five-minute intervals is used as the fine-tuning dataset to obtain a large-scale distributed transcriptional regulation network model for a specific task. Finally, evaluation criteria for the model's prediction results are established, and the model is used to predict the direction of transcriptional regulation.
[0042] Step S1: Obtaining the prior relationship of transcriptional regulation.
[0043] Systematic and standard names of all genes in *Saccharomyces cerevisiae* cells were obtained from the *Yeasttract* and *YeastMine* databases, totaling over 6000 genes. If a gene is a transcription factor, the database provides its transcription factor binding domain information. Based on this information, 221 transcription factors were screened from all *Saccharomyces cerevisiae* genes. The *Yeasttract* database supports querying all target genes affected by a given transcription factor, thus providing prior knowledge of all transcriptional regulation in *Saccharomyces cerevisiae* cells. Examples of prior knowledge of transcriptional regulatory relationships can be found in [link to example]. Figure 2 .
[0044] Step S2: Construction of a large-scale model framework for a distributed transcriptional regulatory network.
[0045] Based on the prior relationships of transcriptional regulation obtained through the above process, it is possible to determine which transcription factors regulate target genes, thereby constructing a distributed sub-network model of target gene expression level changes. The large-scale distributed transcriptional regulation network model framework starts from the genotype to determine whether a gene is missing or knocked out, thereby determining whether the gene expression level needs to be fixed at zero. The overall framework of the large-scale distributed transcriptional regulation network model is shown below. Figure 3 The distributed subnetworks targeting specific genes employ a neural network with a dual-hidden-layer structure. Therefore, each subnetwork has a single output, and the number of nodes in each of the two hidden layers is three times the number of input nodes. The subnetwork structure is shown below. Figure 4 The choice of three hidden layers is based on the three regulatory relationships between transcription factors and target genes: enhancement, repression, and no effect. The number of parameters in the constructed subnetworks ranges from several hundred to tens of thousands, due to the varying degrees of transcription factor regulation affecting the genes, resulting in different network complexities. The total number of parameters in all subnetworks can reach millions to tens of millions.
[0046] Step S3. Distributed pre-training of the large-scale transcriptional regulatory network model.
[0047] Since the total number of parameters in all distributed sub-networks can reach millions to tens of millions, training such a massive network as a whole during model training would consume a significant amount of computer memory, potentially leading to memory overflow and training failure. Furthermore, as the model size increases, the training time also increases exponentially, requiring even more time to train the entire network, resulting in a significant decrease in efficiency. Therefore, a distributed training approach is adopted, constructing a unique training dataset for each sub-network based on its structure and input variables, enabling independent training of each sub-network.
[0048] The training dataset was obtained from pan-transcriptome sequencing data, which contains sequencing data from 969 different yeast strains and more than 6,000 genes. This dataset is stored in TPM (transcriptome per million records), and the TPM values vary between different datasets, requiring data preprocessing for model training. First, the gene lists from the 969 strains are not entirely consistent, with some genes missing data in certain strains. Considering that the missing data might be due to low gene expression levels, zero-value imputation was used to complete the gene data for all strains. Subsequently, to align the data to a uniform distribution and ensure that zero values remain zero after scaling, logarithmic scaling (log2(TPM+1)) was used.
[0049] After data processing, the dataset is split into training datasets for each sub-network, and these datasets are then fed into each sub-network for distributed training. Since this is the pre-training phase, relatively conservative training parameter settings are used. Specific training parameter settings are as follows: learning rate lr = 0.005; maximum epochs = 2500; Adam optimizer used, loss function is sum of variances. The overall results of the distributed network pre-training are shown below. Figure 5 .
[0050] Step S4. Fine-tuning training of the large model of the transcriptional regulatory network.
[0051] After completing the model's pre-training, further training is needed based on the pre-trained model parameters to meet the requirements of downstream tasks. The fine-tuning dataset consists of time-series sequencing data from *Saccharomyces cerevisiae*, which, after perturbing the expression levels of different transcription factors, shows changes in the expression levels of other genes at different time points. Similarly, the fine-tuning dataset needs to be preprocessed before retraining.
[0052] First, the fine-tuning dataset is stored as a logarithmically transformed ratio of gene expression levels before and after transcription factor perturbation, and each data set does not contain missing values, so no data imputation is required. Observation revealed some zero values in the fine-tuning dataset, indicating that the gene expression level is too low to be detected at the current time. However, when zero values appear in the training data, the model training process may experience gradient vanishing or gradient exploding; therefore, the dataset is also processed using a logarithmic transformation: log2(2 x +1). Maintain the same approach as the pre-trained dataset to keep the dataset distribution as consistent as possible.
[0053] Secondly, the fine-tuning data is a time series, but the time intervals between each set of detection data are uneven. Therefore, an interpolation method is used to process the data. The radial basis interpolation method of the Gaussian (Exp) function is used to interpolate the dataset at five-minute intervals.
[0054] Finally, distributed fine-tuning training was implemented after splitting the dataset using the same method as in the pre-training stage. The distributions of transcription factors and gene expression levels at the previous and next time steps were used as the network's input and output. Specifically, the hidden layer parameters of the fine-tuned sub-network were frozen and not trained, while the parameters of other layers were trained based on the pre-training parameters. Here, due to the targeted fine-tuning training, more detailed model training parameters were used. The specific training parameter settings are as follows: training and test sets were split in a 9:1 ratio; learning rate lr = 0.001; weight decay = 0.001; number of iterations epoch = 200 × training set size; maximum number of iterations Max_epoch = 20000; Adam optimizer was used; and the loss function was set to mean squared error. The overall training and test errors of the model fine-tuning are shown in [link to documentation]. Figure 6 .
[0055] Step S5. Evaluation indicators for distributed transcriptional regulation prediction results.
[0056] Based on the characteristics of the fine-tuned dataset, the output of the model after fine-tuning is a specific numerical prediction. However, while more accurate numerical predictions are always better, predicting the direction of transcriptional regulatory relationships is also an area of interest for researchers. Therefore, to provide a more detailed evaluation of the fine-tuned model's output, three prediction result levels were established based on different performances of the prediction results relative to the label values.
[0057] When the absolute error between the prediction result and the label value is less than or equal to 0.2, the result is considered to have achieved high accuracy in numerical prediction and can accurately simulate the quantified transcriptional regulatory relationship. The prediction result is then defined as "excellent".
[0058] When the absolute error between the predicted result and the label value is greater than 0.2, the prediction is considered inaccurate in terms of quantitative indicators. However, if both errors are greater than 1.2 or less than 0.8, the predicted trend is considered correct. Therefore, although the prediction accuracy in this case is not high, the prediction result is still considered correct and is defined as "qualified".
[0059] If the relationship between the predicted result and the label value does not meet either of the above two conditions, then the result is considered to lack sufficient accuracy in both numerical and trend prediction, and is therefore deemed incorrect. The predicted result is then defined as "unqualified".
[0060] By analyzing all predictions made by each subnetwork on the fine-tuned dataset, the performance of the network can be evaluated, thereby guiding the regulation of target gene expression levels. The overall evaluation of all prediction results for each subnetwork can be found in [link to evaluation]. Figure 7 .
[0061] In summary, the method of this invention can start from prior knowledge of transcriptional regulation, employ transfer learning to obtain high-order abstract features of transcriptional regulatory relationships, and perform fine-tuning training for downstream tasks, with graded evaluation of prediction results. The distributed network structure can overcome limitations such as computer memory, data size, and computation time. It can predict the direction of transcriptional regulatory relationships, further guide the regulation of target gene expression levels, and accelerate related research on transcriptional regulatory relationships.
[0062] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for constructing a large-scale distributed transcriptional regulatory network model based on transfer learning, characterized in that, Using reliable prior relationships of transcriptional regulation supported by literature as the basis for selecting input and output variables of distributed subnetworks, the model is trained using the concept of transfer learning to obtain the regulatory relationships in the transcription process, including the following steps: Step S1: Obtain transcriptional regulatory relationships through prior knowledge of transcriptional regulation to obtain the pairing of transcription factors with their corresponding target genes; Step S2: Construct a large-scale distributed transcriptional regulatory network model; Step S3: Perform model pre-training on pan-transcriptome data to obtain high-order features of transcriptional regulatory relationships in various strains; Step S4: Using a time-series dataset, fine-tune the model to obtain a specific distributed large model; Step S5: Based on the data characteristics, formulate evaluation indicators for the prediction results, divide the prediction results into three levels of accuracy, and observe the impact of the dataset on the model prediction performance and their reliability based on the accuracy of each sub-network. The large-scale distributed transcriptional regulation network model consists of distributed sub-networks. These sub-networks are based on transcriptional regulation mechanism networks obtained from the database. The mechanism networks are separated around target genes, with the expression levels of transcription factors as inputs and the expression levels of target genes as outputs. Each distributed sub-network has a multi-input, single-output structure. The pre-training in step S3 is a process of data processing and alignment on a pan-transcriptome dataset based on a distributed network, followed by the following steps: Step S301: Supplement and align the data in the pantranscriptome dataset; Step S302: Align all data to logarithmic space to make its data distribution suitable for machine learning training. The processing formula is expressed as follows: TPM stands for number of transcripts per million. Step S303: Based on the different inputs and outputs of each sub-network, the pre-training data is segmented to adapt to the features of each sub-network. Step S304: Train each distributed subnetwork separately. The distributed structure ensures that the subnetworks do not affect each other. The fine-tuning in step S4 is the process of continuing to train each pre-trained sub-model on the specific dataset of the downstream task, including the following steps: Step S401: Process the fine-tuning dataset by aligning the pre-training dataset with the fine-tuning dataset in terms of gene count and gene name, adjusting the structure of each sub-network according to the fine-tuning dataset, and using interpolation to align the time intervals between samples to facilitate model fine-tuning training. Step S402: Align the fine-tuning data to logarithmic space. The data processing formula is: , where x is the time-series sequencing data of changes in the expression levels of other genes detected at different time points after perturbing the expression levels of different transcription factors, and is the ratio of gene expression levels before and after logarithmic transformation of transcription factors; Step S403: Fix some parameters of each sub-model, freeze some parameters on the pre-training results based on transfer learning, and continue to train other parameters based on the pre-training results; Step S404: Segment the fine-tuning dataset, determine the input and output of each sub-network, and feed them into each sub-model for model fine-tuning using a distributed training method.
2. The method for constructing a large-scale distributed transcriptional regulatory network model based on transfer learning according to claim 1, characterized in that, In step S1, all relevant genes in the database are paired to obtain a transcriptional regulation knowledge graph, and a transcriptional regulation mechanism network is constructed based on the transcriptional regulation knowledge graph.
3. The method for constructing a large-scale distributed transcriptional regulatory network model based on transfer learning according to claim 1, characterized in that, When the target gene is regulated by only one transcription factor or itself, the distributed subnetwork is a single-input, single-output structure.
4. The method for constructing a large-scale distributed transcriptional regulatory network model based on transfer learning according to claim 1, characterized in that, The aforementioned evaluation index for prediction results is formulated based on the different relationships between the output results of the fine-tuning model and the actual label values. The results are divided into three levels: excellent, qualified, and unqualified, which respectively represent that the error between the prediction result and the actual label is small, the error between the prediction result and the actual label is large but the direction is correct, and the error between the prediction result and the actual label is very large.
Citation Information
Patent Citations
Transcription factor chip kit and method for high throughput screening of target gene transcription factor
CN108548780A
Dimension reduction modeling method and system for gene regulation and control network
CN117809734A