Large model training method and device for gene regulation network prediction

The gene regulation network is constructed through the large-scale training method, a one-hop sub-map is generated and structural perturbed. Combined with the feature embedding parameters of the graphic and text, the high cost and low throughput problems of gene regulation network prediction are solved, and causal reasoning and high-accurate gene regulation relationship prediction are achieved.

CN120278191AActive Publication Date: 2025-07-08ZHEJIANG LAB
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510775037.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-08
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

The existing technology has problems of high cost, low throughput and coverage limitations in gene regulation network prediction. The traditional experimental methods are expensive and time-consuming. The calculation model relies on statistical correlation rather than causal reasoning, and the false positive rate is high.

Method used

The large-model training method is adopted to construct a gene regulation network, determine a one-hop sub-graph, perform structural perturbation to generate positive and negative samples, use the graph encoding model to obtain the graph embedded features, and combine the text description information to optimize the cross entropy loss and comparison loss, and adjust the model parameters.

Benefits of technology

It realizes low-cost and fast gene regulation network prediction, has causal reasoning ability, and improves the accuracy and robustness of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278191A_ABST
    Figure CN120278191A_ABST
Patent Text Reader

Abstract

The invention discloses a large model training method and device for gene regulatory network prediction, and the method comprises the steps: determining a jump sub-graph with each node as a center in a constructed gene regulatory network, obtaining the graph embedding features of the sub-graph, the positive sample of the sub-graph, and the negative sample of the sub-graph through a preset graph coding model, and carrying out the prediction of the gene regulatory network. And calculating comparison loss based on the graph embedding features. And then, through first prompt information, enabling the large model to learn a corresponding relationship between the graph embedding characteristics of the sub-graphs and the text embedding characteristics of the genes in the sub-graphs, and calculating cross entropy loss. And large model parameters are jointly adjusted through comparison loss and cross entropy loss. In the method, graph structure information represented by graph embedding features obtained by a graph coding model is migrated to a large model, so that the large model performs joint reasoning according to a correspondence relationship between the learned graph structure information of a gene regulation network and text description information of genes, and has the capability of predicting a gene regulation relationship.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and particularly to a large model training method and device for gene regulatory network prediction. Background Art

[0002] A gene regulatory network is a complex system that describes the interaction relationships between genes, and its topological structure directly affects key biological processes such as cell differentiation and disease occurrence. Traditional gene regulatory network (GRN) discovery methods rely on experimental techniques such as gene knockout (Gene Knockout, KO) and chromatin immunoprecipitation (Chromatin Immunoprecipitation, ChIP-seq), and have the following defects: High cost: The cost of a single experiment can reach thousands of dollars, and the cost of large-scale network construction increases exponentially; Low throughput: The experimental period usually takes several weeks to several months, making it difficult to meet the needs of dynamic network analysis; Coverage limitation: It can only capture the regulatory relationships between known transcription factors and target genes, ignoring emerging regulatory elements such as non-coding RNAs. Although existing computational models (such as GENIE3, SCENIC) can partially replace experiments, they are based on statistical correlation rather than causal reasoning, and have a high false positive rate.

[0003] Based on this, this specification provides a large model training method for gene regulatory network prediction. Summary of the Invention

[0004] This specification provides a large model training method, device, storage medium and electronic device for gene regulatory network prediction to at least partially solve the above problems existing in the prior art.

[0005] This specification adopts the following technical solutions: This specification provides a large model training method for gene regulatory network prediction, including: Construct a gene regulatory network according to gene pairs with regulatory relationships, where the nodes in the gene regulatory network represent genes and the edges represent the regulatory relationships between genes; In the gene regulatory network, respectively determine one-hop subgraphs centered on the nodes of each gene; For each subgraph, perform structural perturbation on the subgraph to generate positive and negative samples of the subgraph; Through a preset graph encoding model, respectively obtain the graph embedding features of the subgraph, the positive sample of the subgraph and the negative sample of the subgraph, and obtain a contrast loss according to the similarity of the graph embedding features between the subgraph and the positive sample and the similarity of the graph embedding features between the subgraph and the negative sample; For each sub - graph, through the first hint information, enable the large - model to learn the correspondence between the graph embedding features of the sub - graph and the text embedding features of each gene in the sub - graph, and calculate the cross - entropy loss according to the first result output by the large - model, where the text embedding features are obtained by extracting features from the pre - acquired text description information of genes; Adjust the parameters of the large - model according to the contrast loss and the cross - entropy loss.

[0006] Optionally, for each sub - graph, perform structural perturbation on the sub - graph to generate a positive sample of the sub - graph and a negative sample of the sub - graph, specifically including: For each sub - graph, delete the edges in the sub - graph at a ratio less than the first set value to obtain the positive sample of the sub - graph; Change the edge connection relationship in the sub - graph at a ratio greater than the second set value to obtain the negative sample of the sub - graph, where the first set value is less than the second set value.

[0007] Optionally, the first hint information contains graph feature placeholders, and the first hint information is used to instruct the large - model to sort the shuffled gene description information; Through the first hint information, enable the large - model to learn the correspondence between the graph embedding features of the sub - graph and the text embedding features of each gene in the sub - graph, and calculate the cross - entropy loss, specifically including: Obtain the text embedding features extracted by the large - model for the first hint information, where the text embedding features contain graph feature placeholders and the graph feature placeholders are not encoded; Replace the graph feature placeholders with the graph embedding features to obtain the graph - text joint features; The large - model performs inference based on the graph - text joint features and outputs the sorting prediction results of the gene description information corresponding to each node included in the sub - graph; Determine the cross - entropy loss according to the difference between the sorting prediction result and the correct sorting result.

[0008] Optionally, the first hint information contains the description - based information of each sorted node in the sub - graph and the gene description information of each shuffled node.

[0009] Optionally, after adjusting the parameters of the large - model, the method further includes: Save the parameters of the adjusted large - model; Obtain sample gene pairs, input the sample gene pairs and the second hint information into the large - model, and obtain the second result output by the large - model, where the second hint information is used to instruct the large - model to predict whether there is a regulatory relationship between the input gene pairs; Determine the inference loss according to the difference between the second result and the true regulatory relationship of the sample gene pair; Adjust the parameters of the large model based on the inference loss.

[0010] Optionally, the second result is a discrete confidence level, and the true regulatory relationship is represented by an existence binary label and an existence confidence level label; Determine the inference loss according to the difference between the second result and the true regulatory relationship of the sample gene pair, specifically including: Obtain the predicted probability that the large model obtains to represent the existence of a regulatory relationship between the sample gene pair; Determine the first loss according to the difference between the predicted probability and the existence binary label of the sample gene pair; map the existence confidence level label of the sample gene pair to a continuous value, and determine the second loss according to the difference between the predicted probability and the mapped continuous value; Determine the inference loss based on the first loss and the second loss.

[0011] Optionally, input gene pairs with unknown regulatory relationships into the large model, and the large model outputs a prediction result representing the regulatory relationship between the gene pairs with unknown regulatory relationships.

[0012] This specification provides a large model training device for gene regulatory network prediction, and the device includes: A graph construction module that constructs a gene regulatory network according to gene pairs with regulatory relationships, where the nodes in the gene regulatory network represent genes, and the edges represent the regulatory relationships between genes; A subgraph extraction module that respectively determines one-hop subgraphs centered on the nodes of each gene in the gene regulatory network; A sample generation module that perturbs the structure of each subgraph to generate positive and negative samples of the subgraph; A contrast learning module that respectively obtains the graph embedding features of the subgraph, the positive sample of the subgraph, and the negative sample of the subgraph through a preset graph encoding model, and obtains a contrast loss according to the similarity of the graph embedding features between the subgraph and the positive sample and the similarity of the graph embedding features between the subgraph and the negative sample; A graph-text joint module that, for each subgraph, enables the large model to learn the correspondence between the graph embedding features of the subgraph and the text embedding features of each gene in the subgraph through the first prompt information, and calculates the cross-entropy loss according to the first result output by the large model, where the text embedding features are obtained by extracting features from the pre-obtained text description information of genes; A parameter adjustment module that adjusts the parameters of the large model according to the contrast loss and the cross-entropy loss.

[0013] This specification provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above-described large model training method for gene regulatory network prediction.

[0014] This specification provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the above-described large model training method for gene regulatory network prediction.

[0015] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects: In the large model training method for gene regulatory network prediction provided in this specification, in the constructed gene regulatory network, a one-hop subgraph centered on each node is determined. Through a preset graph encoding model, graph embedding features of the subgraph, the positive sample of the subgraph, and the negative sample of the subgraph are respectively obtained, and a contrastive loss is calculated based on the graph embedding features. Then, through a first prompt message, the large model learns the correspondence between the graph embedding features of the subgraph and the text embedding features of each gene in the subgraph, and calculates a cross-entropy loss. The large model parameters are jointly adjusted through the contrastive loss and the cross-entropy loss. In this method, the graph structure information represented by the graph embedding features obtained by the graph encoding model is migrated to the large model, enabling the large model to perform joint reasoning according to the learned correspondence between the graph structure information of the gene regulatory network and the text description information of the gene, and having the ability to predict gene regulatory relationships. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings described herein are used to provide a further understanding of this specification, and constitute a part of this specification. The illustrative embodiments of this specification and their descriptions are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings: Figure 1 is a schematic flowchart of a large model training method for gene regulatory network prediction in this specification; Figure 2 is a model training architecture diagram provided by an embodiment of this specification; Figure 3 is a data construction process diagram provided by an embodiment of this specification; Figure 4 is a schematic diagram of a large model training device for gene regulatory network prediction provided by this specification; Figure 5 corresponding to Figure 1 is a schematic diagram of an electronic device. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.

[0018] The following will detail the technical solutions provided by each embodiment of this specification in conjunction with the drawings.

[0019] Figure 1 It is a schematic flowchart of a large model training method for gene regulatory network prediction in this specification, specifically including the following steps: S100: Construct a gene regulatory network based on gene pairs with regulatory relationships. The nodes in the gene regulatory network represent genes, and the edges represent the regulatory relationships between genes.

[0020] All steps in the large model training method for gene regulatory network prediction provided by this specification can be implemented by any electronic device with computing functions, such as terminals, servers, etc. For ease of description, only the server is used as the execution subject below to illustrate the large model training method for gene regulatory network prediction provided by this specification.

[0021] In this specification, a gene regulatory network can be constructed based on known gene pairs with regulatory relationships for the training of a large model. The gene regulatory network is a graph data network structure, where nodes represent genes and edges represent regulatory relationships. The gene pairs here can be obtained from public datasets on the network, and the gene regulatory network can be represented as G=(V,E), where V represents the set of gene nodes and E represents the set of edges.

[0022] S102: In the gene regulatory network, respectively determine the one-hop subgraphs centered on each node.

[0023] In a gene regulatory network, if there is an edge connection between two nodes, it means there is a regulatory relationship between the genes represented by these two nodes. For each node in the gene regulatory graph, the one-hop subgraph centered on this node, that is, the subgraph composed of the nodes directly connected to this node, represents the genes that have regulatory relationships with the gene represented by this node.

[0024] Therefore, in order to enable the large model to learn the regulatory relationships between genes, the server extracts the one-hop subgraphs of each node in the gene regulatory network.

[0025] S104: For each subgraph, perform structural perturbation on the subgraph to generate positive and negative samples of the subgraph.

[0026] In this method, the large model is optimized through the contrastive loss. In this step, positive and negative samples required for calculating the contrastive loss are constructed through each sub-graph. The method is to perform structural perturbations on each sub-graph to determine the positive and negative samples of each sub-graph.

[0027] In one embodiment of this specification, the server can perform relatively minor structural perturbations on the sub-graph to generate the positive sample of the sub-graph. The positive sample is highly similar but not exactly the same as the sub-graph itself. Relatively large structural perturbations are performed on the sub-graph to generate the negative sample of the sub-graph, and the negative sample is quite different from the sub-graph itself.

[0028] Specifically, for each sub-graph, the server can delete the edges in the sub-graph at a ratio less than the first set value to obtain the positive sample of the sub-graph. The connection relationship of the edges in the sub-graph is changed at a ratio greater than the second set value to obtain the negative sample of the sub-graph.

[0029] Among them, the first set value is less than the second set value to ensure that the positive sample has a high similarity to the sub-graph itself and the negative sample has a low similarity to the sub-graph itself. Exemplarily, the first set value can be set to a value less than 10%, and the second set value can be set to 50%.

[0030] In this embodiment, when determining the positive and negative samples, structural perturbations are performed, rather than only performing structural perturbations when generating the negative sample. This can reduce the impact of data errors of the gene pairs obtained in step S100 on model inference during the training process, and enhance the robustness and generalization of the large model's gene regulation relationship prediction.

[0031] S106: Through a preset graph encoding model, obtain the graph embedding features of the sub-graph, the positive sample of the sub-graph, and the negative sample of the sub-graph respectively, and obtain the contrastive loss according to the graph embedding feature similarity between the sub-graph and the positive sample and the graph embedding feature similarity between the sub-graph and the negative sample.

[0032] The large model is a model for text encoding and cannot directly process graph data. In this specification, through an additional graph encoding model, each sub-graph, the positive sample of each sub-graph, and the negative sample of each sub-graph are encoded to obtain the graph embedding features corresponding to each sub-graph, the positive sample of each sub-graph, and the negative sample of each sub-graph respectively. This specification does not limit the specific type of the graph encoding model, and it can be selected as needed.

[0033] Then, according to the graph embedding feature similarity between each sub-graph and the positive sample of the sub-graph, and the graph embedding feature similarity between each sub-graph and the negative sample of the sub-graph, calculate the contrastive loss.

[0034] Among them, if the Graph Attention Networks (GAT) is used as the graph encoding model, Represents the graph embedding feature of the subgraph G, Represents the positive sample of this subgraph The graph embedding features of Represents the negative sample of this sub-image The contrast loss is It can be determined by the following formula:

[0035]

[0036]

[0037] In the above formula represents the temperature parameter, which is used to adjust the sharpness of the similarity distribution. The larger the value, the smoother the distribution. The smaller the value, the sharper the distribution and the stronger the contrast. It can usually be set to 0.05~0.2. K represents the number of negative samples.

[0038] By calculating the contrast loss, during the large model inference process, the similarity between the sub-image and the positive sample can be shortened, while the similarity between the sub-image and the negative sample can be increased.

[0039] It should be noted that the graph coding model used in this specification is an existing model that has been trained for book corner coding. The coding results of the graph coding model are used in this specification to train the large model. Therefore, although the contrast loss is calculated based on the coding results of the graph coding model, the contrast loss will be used to adjust the parameters of the large model in order to transfer the subgraph structure knowledge encoded by the graph coding model to the large model.

[0040] S108: For each subgraph, through the first prompt information, the big model learns the correspondence between the graph embedding features of the subgraph and the text embedding features of each gene in the subgraph, and calculates the cross entropy loss according to the first result output by the big model, wherein the text embedding features are obtained by feature extraction of the text description information of the genes obtained in advance.

[0041] The method in this specification can be trained based on any type of large model, which can be a language model with a large parameter scale, such as a generative pre-trained model (GPT), or a lightweight language model with a small parameter scale, such as phi3-mini.

[0042] In this specification, the data for training, in addition to gene pairs with regulatory relationships, also includes the text description information of each gene. This text description information can also be obtained from public datasets and is used to describe the functions of genes.

[0043] The gene regulatory network constructed based on gene pairs with known regulatory relationships is graph data, and the gene description information of each gene is text data. In this step, through the first prompt information, the large model is enabled to learn the correspondence between the graph embedding features of the subgraph and the text embedding features of each gene in the subgraph, and reasoning is performed based on these two types of data related to gene regulatory relationships, namely graph data and text data, to learn the regulatory relationships between gene pairs, so that the large model has the ability to predict the regulatory relationships of gene pairs with unknown gene regulatory relationships.

[0044] Among them, when training the large model to learn the correspondence between the graph embedding features of the subgraph and the text embedding features of each gene in the subgraph, various methods can be adopted. Different training methods result in different contents of the first prompt information and different first results output by the large model.

[0045] In an embodiment of this specification, by enabling the large model to sort the shuffled gene description information, the large model is made to learn the correspondence between the graph embedding features of the subgraph and the text embedding features of each gene in the subgraph. The first prompt information is used to instruct the large model to sort the shuffled gene description information.

[0046] Then in this embodiment, the first prompt information includes a node list and a gene description list. The node list is the nodes of the subgraph in a given order, and the gene description list is the gene description information of each node included in the subgraph in a given order. Among them, the order of each gene description in the gene description list does not exactly correspond to the order of each node in the node list.

[0047] The large model outputs the re-sorted gene description information according to the instruction of the first prompt information, making it correspond to the order of each node in the node list. The sorting of the gene description information output by the large model is used as the first result, and the cross-entropy loss is calculated based on the difference between the first result and the correct sorting result of the gene description information corresponding to the order of each node in the node list determined in advance.

[0048] If the correct sorting result is represented as , and the first result is represented as , then the cross-entropy loss can be represented as . Among them, CE represents the cross-entropy loss calculation function.

[0049] S110: Adjust the parameters of the large model according to the comparison loss and the cross-entropy loss.

[0050] The contrastive loss can enhance the robustness of the large model to the graph embedding features and is used to optimize the ability to understand graph structure information. The cross-entropy loss enables the large model to learn the correspondence between the graph structure of the gene regulatory network and the text description information of genes, enabling the large model to have the ability to jointly reason based on graph data and text data, which is an optimization for language generation tasks.

[0051] In this specification, the total loss is determined by the contrastive loss and the cross-entropy loss, and the large model is jointly optimized, while optimizing the graph structure understanding task and the language generation task, so that the optimized large model achieves a more accurate gene regulatory relationship prediction effect.

[0052] Among them, the total loss can be the weighted sum of the contrastive loss and the cross-entropy loss, and the weight coefficients for weighting are not limited in this specification. Exemplarily, when the weight coefficient of the contrastive loss takes 0.5 and the weight coefficient of the cross-entropy loss takes 1, the total loss can be expressed as: .

[0053] After the total loss converges, the model parameters are saved to obtain the trained large model. At this time, the large model can be used for gene regulatory relationship prediction tasks. Specifically, the nodes corresponding to gene pairs with unknown gene regulatory relationships and the second prompt information can be input into the large model, and the second prompt information is used to instruct the large model to predict the gene regulatory relationship between the input gene pairs and output the prediction result.

[0054] Based on the above Figure 1 shown large model training method for gene regulatory network prediction, in the constructed gene regulatory network, one-hop subgraphs centered on each node are determined, and through a preset graph encoding model, the graph embedding features of the subgraph, the positive samples of the subgraph, and the negative samples of the subgraph are respectively obtained, and the contrastive loss is calculated based on the graph embedding features. Then, through the first prompt information, the large model learns the correspondence between the graph embedding features of the subgraph and the text embedding features of each gene in the subgraph, and calculates the cross-entropy loss. The large model parameters are jointly adjusted through the contrastive loss and the cross-entropy loss. In this method, the graph structure information represented by the graph embedding features obtained by the graph encoding model is migrated to the large model, enabling the large model to jointly reason based on the correspondence between the graph structure information of the learned gene regulatory network and the text description information of genes, and having the ability to predict gene regulatory relationships.

[0055] In the above step S102, when generating positive and negative samples, a dynamic structure perturbation strategy can be adopted to adjust the perturbation rate according to the degree of nodes. High-degree nodes often represent more core regulatory relationships. When generating positive samples, the connection relationships corresponding to lower-degree nodes are preferentially perturbed to ensure a large topological similarity between the positive samples and the original subgraph. When generating negative samples, the connection relationships corresponding to higher-degree nodes are preferentially perturbed to ensure a small topological similarity between the positive samples and the original subgraph, enhancing the training effect of contrastive learning.

[0056] Specifically, when generating positive samples, nodes with degrees lower than a specified value are determined in the subgraph as the first target nodes, and the edges connected to the first target nodes are deleted at a proportion less than the first set value. When generating negative samples, nodes with degrees higher than a second specified value are determined in the subgraph as the second target nodes, and the connection relationships of the edges connected to the second target nodes are changed at a proportion greater than the second set value.

[0057] Alternatively, in another embodiment, when determining whether to perturb an edge, the retention probability corresponding to the edge can also be calculated according to the degrees of the two nodes connected to the edge. The calculation formula for the retention probability P is as follows: For the i-th edge included in the subgraph , according to the node degrees of the two nodes ( and ) connected by the edge, the geometric mean of the degrees of the two nodes is calculated as the importance degree of the edge , and the specific process is as follows: .

[0058] Then, among the importance degrees corresponding to the edges of the subgraph, the maximum importance degree is determined for use in normalizing the retention probability to reduce the influence of the factor graph scale on the calculation of the retention probability.

[0059] Then the retention probability of the i-th edge is calculated by the formula: .

[0060] For each subgraph, the retention probabilities of all the edges included in the subgraph are calculated, and the edges are arranged according to the magnitudes of the retention probabilities.

[0061] When generating positive samples, the first preset proportion of edges is taken from smallest to largest as candidate perturbation edges, and only the candidate perturbation edges are deleted without changing the topological structures of other edges. When generating negative samples, the second preset proportion of edges is taken from largest to smallest as candidate perturbation edges, and connection relationship replacement is performed on the candidate perturbation edges.

[0062] Because the larger the degree, the more complex the connection relationship of the nodes, which represents an important regulatory relationship in the subgraph. The larger the product of the degrees of the two nodes corresponding to an edge, the greater its retention probability. The retention probability indicates the importance of the edge in the topological relationship of the subgraph. The nodes with a larger retention probability have a greater impact on the subgraph topology after being deleted.

[0063] Therefore, when generating positive samples, preferentially select the edges with a smaller retention probability to be deleted, so that the positive samples have a higher topological structure similarity with the original subgraph. When generating negative samples, preferentially change the link relationship of the edges with a larger retention probability, so that the negative samples have a lower topological structure similarity with the original subgraph.

[0064] When determining the cross-entropy loss in the above step S108, the graph embedding features of the subgraph and the text embedding features of the text description information can be fused, so that the large model can perform reasoning based on the fused text-image joint features, better learn the relationship between the graph structure information of the subgraph and the text semantics of the gene description information, and generate more accurate results.

[0065] When the first prompt information is used to instruct the large model to sort the gene description information in a scrambled order, the first prompt information may include a graph feature placeholder. This graph feature placeholder is a special recognition symbol specified in advance during the encoding of the large model. When the large model performs text encoding, it does not encode the graph feature placeholder, so the graph feature placeholder remains in the encoding result.

[0066] In this embodiment, the specific process of determining the cross-entropy loss is as follows.

[0067] First, after inputting the first prompt information into the large model, the large model performs text encoding on the first prompt information to obtain the text embedding features extracted by the large model for the first prompt information. Then, the text embedding features contain the graph feature placeholder that has not been encoded.

[0068] Then, for each subgraph, obtain the graph embedding features extracted by the preset graph encoding model for the subgraph, and replace the graph feature placeholder with the graph embedding features to obtain the text-image joint features.

[0069] If the graph feature placeholder is represented as <graph>, taking as the text embedding feature representing the first prompt message, and taking as the graph embedding feature representing the sub - graph, the text embedding feature before replacement can be expressed as: , the text embedding feature after replacement, (i.e., the text - graph joint feature ) can be expressed as: . Here, the symbol represents feature concatenation, represents the text content before the feature placeholder in the first prompt message, represents the text content after the feature placeholder in the first prompt message.

[0070] After obtaining the text - graph joint feature, input the text - graph joint feature into the large - model again to get the first result of the large - model output. When the large - model selects phi3 - mini, the first result output can be expressed as: .

[0071] Based on this embodiment, there is also a graph feature placeholder in the second prompt message. During the large - model inference, the gene pair with unknown gene regulation relationship input is used as the gene pair to be predicted. The large - model respectively determines the one - hop sub - graphs centered on the two gene nodes in the gene pair to be predicted in the gene regulation network. Then, through the graph encoding model, determine the graph embedding features of the two sub - graphs, and replace the graph feature placeholder in the text embedding feature encoded by the second prompt message with the graph embedding features of the two sub - graphs to obtain the text - graph joint feature. The large - model infers based on the text - graph joint feature and outputs the prediction result.

[0072] The large - model trained in step S110 can already predict gene regulation relationships based on the prompt words representing the edge regulation relationship prediction task.

[0073] However, in order to enhance the inference accuracy of the large - model for the edge prediction task, in an embodiment of this specification, after obtaining the large - model trained in step S110, the large - model can be further optimized for the edge regulation relationship prediction task.

[0074] The sample data for this optimization task are several sample gene pairs with known regulation relationships. For each sample gene pair, input the sample gene pair and the second prompt message into the large - model to obtain the second result output by the large - model. Among them, the second prompt message is used to instruct the large - model to make a prediction on whether there is a regulation relationship for the input gene pair.

[0075] Then, the server determines the inference loss according to the difference between the second result and the true regulation relationship of the sample gene pair, and continues to optimize the large - model based on the inference loss.

[0076] The specific form of the second result obtained by the large model for predicting regulatory relationships can be a binary result representing existence or a discrete confidence level representing the intensity of the existence of regulatory relationships.

[0077] If the second result is a binary label, the corresponding true regulatory relationship is binary data. If the second result is a discrete confidence level, the corresponding true regulatory relationship is also a pre-divided confidence level label.

[0078] Furthermore, in another embodiment, two types of annotations can also be used, that is, the true regulatory relationship includes a binary label representing existence and a discrete confidence level label representing the intensity of the existence of regulatory relationships. Two types of losses are constructed through these two types of annotations to optimize the large model simultaneously.

[0079] First, obtain the predicted probability that the large model shows a regulatory relationship between sample gene pairs. Then, determine the first loss based on the difference between this predicted probability and the binary label of the existence of the sample gene pair. Map the existence confidence level label of the sample gene pair to a continuous value, and determine the second loss based on the difference between this predicted probability and the mapped continuous value.

[0080] Specifically, the first loss is binary cross-entropy loss and can be determined according to the following formula:

[0081] where p represents the predicted probability output by the large model, and y represents the binary label. .

[0082] When calculating the second loss, the discrete confidence level needs to be mapped to a continuous value first. This specification does not limit the specific mapping method. It can be uniformly mapped according to the number of pre-divided confidence levels, or non-uniformly mapped.

[0083] For example, if the confidence level is divided into 6 levels, including levels A - E representing the confidence of the existence of regulatory relationships from weak to strong, and the confidence level N representing that the confidence of the existence of regulatory relationships is zero, that is, there is no regulatory relationship. Then one way of continuous value mapping in this example is as follows: N→0, A→0.2, B→0.4, C→0.6, D→0.8, D→1.0.

[0084] The second loss can be specifically calculated according to the following formula :

[0085] where represents the predicted result of the confidence level output by the model, is the threshold for controlling the error size, usually set to 0.2 - 0.5.

[0086] In the calculation of the second loss, when , that is, when the difference between the predicted result of the confidence level output by the model and the confidence level label is small, the mean squared error (MSE) is adopted. The gradient decreases as the error decreases, and the adjustment of the model parameters is more refined. When the error exceeds the threshold , it indicates that the error is large, and the mean absolute error (MAE) is adopted to prevent the model loss from diverging due to extreme errors, avoid gradient explosion, and enhance the stability of the training gradient.

[0087] Then, after obtaining the first loss and the second loss, based on the first loss and the second loss, the inference loss is determined. Specifically, the inference loss can be determined according to the following formula: .

[0088] After the large model in this embodiment is optimized, the gene pairs to be predicted can be input into the large model, and through the preset second prompt information, the large model is prompted to output the confidence level of the regulatory relationship between the gene pairs .

[0089] If the large model is not optimized by the confidence level prediction task in this embodiment, based on the second prompt information indicating the large model to output the confidence level, the large model can still output the predicted result of the confidence level based on its own semantic understanding ability according to the indication of the second prompt information. However, after being optimized by the confidence level prediction task in this embodiment, the accuracy of the large model for predicting the confidence level of gene regulatory relationships can be significantly increased.

[0090] In the above step S100, when constructing the gene regulatory network based on the gene pairs with known regulatory relationships in the public dataset, it can be specifically constructed in the following manner.

[0091] First, the feature extraction model is used to extract features from each obtained gene to obtain the node embedding features of each gene.

[0092] This specification does not specifically limit the feature extraction model here. It can be a feature extraction model in the general field or a special model in the gene field, such as the single-cell generative pre-trained transformer (scGPT).

[0093] Then, using the node embedding features as the feature identifiers of the nodes, each gene node is generated, and edges are connected between the gene nodes corresponding to the genes with regulatory relationships to obtain a gene regulatory network.

[0094] Figure 2 This is a model training architecture diagram provided by the embodiments of this specification. As Figure 2 shown, the data used for gene regulatory relationship learning are gene pairs with known regulatory relationships, gene description information, and node embedding features of genes. Figure 2 The gene pairs in it are respectively: TF(1)→TF(2), TF(3)→gene(1), TF(4)→miRNA(1), miRNA(2)→TF(1), miRNA(2)→gene(1), where TF represents transcription factor, gene represents gene, miRNA represents microRNA, and → represents regulatory relationship. The numbers in the brackets are used to identify and distinguish different types of TF, gene, and miRNA. Figure 2 The gene regulation information in it is obtained by crawling from publicly available network data, and the content is: USF1 encodes a member of the basic helix-loop-helix leucine zipper family and can function as a cellular transcription factor. The encoded protein can activate transcription through pyrimidine-rich initiator (inr) elements and E-box motifs. This gene has been linked to familial combined hyperlipidemia (FCHL). Alternative splicing of this gene results in multiple transcript variants. A related pseudogene has been defined on chromosome 21. The Chinese interpretation is: The USF1 gene encodes a protein of the basic helix-loop-helix leucine zipper family and can function as a cellular transcription factor. The encoded protein can activate transcription through pyrimidine-rich initiator (Initiator, INR) elements and E-box motifs. This gene is associated with familial combined hyperlipidemia (FCHL). Alternative splicing of this gene can produce multiple transcript variants, and its related pseudogene is located on chromosome 21.

[0095] Construct a gene regulation network based on gene pairs with known regulatory relationships and the node embedding features of genes, encode the one-hop subgraphs of each node in the gene regulation network through a graph encoding model, encode the gene description information through a large model, and through the graph embedding features of the subgraphs , replace the graph feature placeholder that is not encoded in the text embedding feature to obtain the combined text and graph feature, and enable the large model to perform reasoning based on this combined text and graph feature.

[0096] Figure 2 As already shown in , the combined text and graph feature , is composed of the text content before the feature placeholder in the first hint information , the graph embedding feature , and the text content after the feature placeholder in the first hint information

[0097] In one embodiment of this specification, the data input to the large model is organized in a specific data format, and the training process of the above steps S100 to S110 is executed. Figure 3 It is a process diagram of data construction provided by an embodiment of this specification. Figure 5 The process shown specifically includes the following steps: (1.1) Obtain gene pairs with regulatory relationships based on publicly available network data, and use Python code to construct a known gene regulatory network G=(V,E) based on the gene pairs, where V is the set of gene nodes and E is the set of regulatory edges; (1.2) Obtain the feature embedding of the gene in the scGPT model as the node embedding feature of the gene; (1.3) Based on Python crawlers, obtain the text description information of the gene in publicly available network data; (1.4) Use Python to generate JSON data with the following format: { "id": dataid, "graph": { "node_idx": node_idx, "edge_index": edge_index, "node_list": node_list }, "conversations": {"from": "human ", "value": conversation_human}, {"from": "gpt", "value": conversation_gpt} } Among them, "id" is the unique id number corresponding to each piece of data; "graph" represents the one-hop subgraph of the gene related to this piece of data extracted from G for this piece of data, including the central node number "node_idx", the edge index "edge_index", and the node list "node_list" of the nodes other than the central node in the one-hop subgraph. "conversations" is the training data in the form of a conversation, which is used to guide the model to learn. "conversation_human" is the model training prompt, and " <graph>"As an information placeholder for graph embedding features, "conversation_gpt" is the target of model training.

[0098] (1.5) For the regulatory relationship learning task, generate data with a specific structure for this task. The data content is to guide the model to learn the one-hop subgraph gene regulatory network structure centered on a certain gene. Specifically, the data format is the same as the data format described in step 1.4. Among them, "conversation_human" requires the model to sort the scrambled gene description information so that the gene description information corresponds one by one with the order of gene nodes, and the content of "conversation_gpt" is the correct order of gene description information.

[0099] (1.6) For the large model edge prediction inference task, generate data with a specific structure for this task. After the model training is completed, the data format during the inference prediction process is input according to this data structure. The data content is to guide the model to predict the degree of possible regulatory relationship between two genes. Specifically, the data format is the same as the data format described in step 1.4. Among them, "conversation_human" describes two gene nodes, and the content of "conversation_gpt" is the confidence level Confidence of the possible regulatory relationship between the two gene nodes. The levels are divided into N, A - E levels, where the confidence level of having an edge for E is the highest, the confidence level of A is the lowest, and N indicates no regulatory relationship.

[0100] In the above step S108, when enabling the large model to learn the correspondence between the graph embedding features of the subgraph and the text embedding features of each gene in the subgraph, it is not limited to the way of sorting gene description information, and other ways can also be adopted.

[0101] For example, input a one-hop subgraph centered on a certain node and the text description information of several genes. The first prompt message is used to prompt the large model to select the text description information corresponding to the central node of the subgraph from multiple text description information. Or, the gene description information of each gene pair describing the topological regulatory relationship of the subgraph can be randomly masked, and the first prompt message is used to instruct the large model to restore the complete unmasked gene regulatory information according to the graph structure information of the subgraph.

[0102] The above is the large model training X method for gene regulatory network prediction provided in this specification. Based on the same idea, this specification also provides a corresponding large model training device for gene regulatory network prediction, as Figure 4 shown.

[0103] Figure 4 is a schematic diagram of a large model training device for gene regulatory network prediction provided in this specification, specifically including: A graph construction module 200 for constructing a gene regulatory network based on gene pairs with regulatory relationships, where the nodes in the gene regulatory network represent genes and the edges represent the regulatory relationships between genes; A subgraph extraction module 202 for determining one-hop subgraphs centered on the nodes of each gene in the gene regulatory network; A sample generation module 204 for, for each subgraph, performing structural perturbation on the subgraph to generate positive and negative samples of the subgraph; A contrastive learning module 206 for, through a preset graph encoding model, respectively obtaining the graph embedding features of the subgraph, the positive sample of the subgraph, and the negative sample of the subgraph, and obtaining a contrastive loss based on the similarity of the graph embedding features between the subgraph and the positive sample and the similarity of the graph embedding features between the subgraph and the negative sample; A graph-text joint module 208 for, for each subgraph, through a first prompt message, enabling a large model to learn the correspondence between the graph embedding features of the subgraph and the text embedding features of each gene in the subgraph, and calculating a cross-entropy loss based on the first result output by the large model, where the text embedding features are obtained by performing feature extraction on the pre-obtained text description information of genes; A parameter adjustment module 210 for adjusting the parameters of the large model according to the contrastive loss and the cross-entropy loss.

[0104] Optionally, the sample generation module 204 is specifically configured to, for each subgraph, delete the edges in the subgraph at a ratio less than a first set value to obtain the positive sample of the subgraph, and change the edge connection relationship in the subgraph at a ratio greater than a second set value to obtain the negative sample of the subgraph, where the first set value is less than the second set value.

[0105] Optionally, the first prompt message contains a graph feature placeholder, and the first prompt message is used to instruct the large model to sort the scrambled gene description information. The graph-text joint module 208 is specifically configured to obtain the text embedding features extracted by the large model for the first prompt message, where the text embedding features contain the graph feature placeholder, and the graph feature placeholder is not encoded. By replacing the graph feature placeholder with the graph embedding features, graph-text joint features are obtained, and the large model performs inference based on the graph-text joint features and outputs a sorting prediction result of the gene description information corresponding to each node included in the subgraph. According to the difference between the sorting prediction result and the correct sorting result, the cross-entropy loss is determined.

[0106] Optionally, the first prompt message contains the description-based information of each sorted node in the subgraph and the gene description information of each scrambled node.

[0107] Optionally, the device further includes an inference optimization module 212.

[0108] The inference optimization module 212 is configured to save the parameters of the large model after adjustment; Obtain a sample gene pair, input the sample gene pair and the second prompt information into the large model, obtain a second result output by the large model, where the second prompt information is used to instruct the large model to predict whether there is a regulatory relationship between the input gene pairs, and determine an inference loss according to the difference between the second result and the true regulatory relationship of the sample gene pair, and adjust the parameters of the large model based on the inference loss.

[0109] Optionally, the second result is a discrete confidence level, and the true regulatory relationship is represented by an existence binary label and an existence confidence level label. The inference optimization module 212 is specifically configured to obtain the prediction probability that the large model obtains to represent that there is a regulatory relationship between the sample gene pairs, determine a first loss according to the difference between the prediction probability and the existence binary label of the sample gene pair; map the existence confidence level label of the sample gene pair to a continuous value, and determine a second loss according to the difference between the prediction probability and the mapped continuous value, and determine an inference loss based on the first loss and the second loss.

[0110] Optionally, the device further includes a prediction module 214.

[0111] The prediction module 214 is specifically configured to input a gene pair with an unknown regulatory relationship into the large model, and the large model outputs a prediction result representing the regulatory relationship between the gene pairs with the unknown regulatory relationship.

[0112] This specification also provides a computer-readable storage medium, which stores a computer program, and the computer program can be used to execute the above Figure 1 The large model training method for gene regulatory network prediction provided.

[0113] This specification also provides Figure 5 The schematic structural diagram of the electronic device shown. As Figure 5 As described above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 The large model training method for gene regulatory network prediction. Of course, in addition to the software implementation, this specification does not exclude other implementation methods, such as logic devices or the combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or logic devices.

[0114] For an improvement in a technology, it can be clearly distinguished whether it is an improvement in hardware (e.g., improvement in circuit structures such as diodes, transistors, switches, etc.) or an improvement in software (improvement in method processes). However, with the development of technology, many improvements in method processes today can be regarded as direct improvements in hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structures by programming the improved method processes into the hardware circuits. Therefore, it cannot be said that an improvement in a method process cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. Designers can "integrate" a digital system onto a single PLD by programming it themselves, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL), and there is not just one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be clear that as long as the method process is slightly logically programmed using the above-mentioned several hardware description languages and programmed into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method process.

[0115] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.

[0116] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0117] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0118] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0119] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in a block or multiple blocks.

[0120] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in a block or multiple blocks.

[0121] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in a block or multiple blocks.

[0122] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0123] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.

[0124] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0125] It should also be noted that the term "includes", "comprising" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the existence of other identical elements in the process, method, commodity or device including the element.

[0126] It should be understood by those skilled in the art that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0127] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0128] The various embodiments in this specification are described in a progressive manner. For the parts that are the same or similar among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiments.

[0129] The above description is only for the embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various changes and modifications can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this application.< / graph> ​< / graph>

Claims

1. A large model training method for gene regulatory network prediction, characterized in that, Including: Construct a gene regulatory network based on gene pairs with regulatory relationships, where the nodes in the gene regulatory network represent genes, and the edges represent the regulatory relationships between genes; In the gene regulatory network, respectively determine the one-hop subgraphs centered on the nodes of each gene; For each subgraph, perform structural perturbation on the subgraph to generate positive and negative samples of the subgraph; Through a preset graph encoding model, respectively obtain the graph embedding features of the subgraph, the positive sample of the subgraph, and the negative sample of the subgraph, and obtain a contrastive loss according to the similarity of the graph embedding features between the subgraph and the positive sample and the similarity of the graph embedding features between the subgraph and the negative sample; For each subgraph, through the first prompt information, enable the large model to learn the correspondence between the graph embedding features of the subgraph and the text embedding features of each gene in the subgraph, and calculate the cross-entropy loss according to the first result output by the large model, where the text embedding features are obtained by extracting features from the pre-obtained text description information of the genes; According to the contrastive loss and the cross-entropy loss, adjust the parameters of the large model.

2. The method according to claim 1, characterized in that For each subgraph, perform structural perturbation on the subgraph to generate positive and negative samples of the subgraph, specifically including: For each subgraph, delete the edges in the subgraph at a ratio less than the first set value to obtain the positive sample of the subgraph; Change the edge connection relationship in the subgraph at a ratio greater than the second set value to obtain the negative sample of the subgraph, where the first set value is less than the second set value.

3. The method according to claim 1, wherein The first prompt information contains graph feature placeholders, and the first prompt information is used to instruct the large model to sort the shuffled gene description information; Through the first prompt information, enable the large model to learn the correspondence between the graph embedding features of the subgraph and the text embedding features of each gene in the subgraph, and calculate the cross-entropy loss according to the first result output by the large model, specifically including: Obtain the text embedding features extracted by the large model from the first prompt information, where the text embedding features contain graph feature placeholders that are not encoded; Replace the graph feature placeholders with the graph embedding features to obtain a graph-text joint feature; The large model performs inference based on the graph-text joint feature and outputs a sorting prediction result of the gene description information corresponding to each node included in the subgraph; Determine the cross-entropy loss according to the difference between the sorting prediction result and the correct sorting result.

4. The method according to claim 1, characterized in that The first prompt information contains the description-based information of each sorted node in the subgraph and the gene description information of each shuffled node.

5. The method according to claim 1, characterized in that After adjusting the parameters of the large model, the method further includes: Save the parameters of the adjusted large model; Obtain sample gene pairs, input the sample gene pairs and the second prompt information into the large model, and obtain the second result output by the large model, where the second prompt information is used to instruct the large model to predict whether there is a regulatory relationship between the input gene pairs; Determine the inference loss according to the difference between the second result and the true regulatory relationship of the sample gene pairs. Adjust the parameters of the large model based on the inference loss.

6. The method according to claim 5, wherein The second result is a discrete confidence level, and the true regulatory relationship is represented by an existence binary label and an existence confidence level label; Determine the inference loss according to the difference between the second result and the true regulatory relationship of the sample gene pair, specifically including: Obtain the predicted probability that the large model obtains to represent the existence of a regulatory relationship between the sample gene pair; Determine the first loss according to the difference between the predicted probability and the existence binary label of the sample gene pair; map the existence confidence level label of the sample gene pair to a continuous value, and determine the second loss according to the difference between the predicted probability and the mapped continuous value; Determine the inference loss based on the first loss and the second loss.

7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Input gene pairs with unknown regulatory relationships into the large model, and the large model outputs a prediction result representing the regulatory relationship between the gene pairs with unknown regulatory relationships.

8. A large model training device for gene regulatory network prediction, characterized in that, Include: A graph construction module that constructs a gene regulatory network according to gene pairs with regulatory relationships. Nodes in the gene regulatory network represent genes, and edges represent regulatory relationships between genes; A subgraph extraction module that, in the gene regulatory network, respectively determines one-hop subgraphs centered on the nodes of each gene; A sample generation module that, for each subgraph, perturbs the structure of the subgraph to generate positive and negative samples of the subgraph; A contrast learning module that, through a preset graph encoding model, respectively obtains the graph embedding features of the subgraph, the positive sample of the subgraph, and the negative sample of the subgraph, and obtains a contrast loss according to the similarity of the graph embedding features between the subgraph and the positive sample and the similarity of the graph embedding features between the subgraph and the negative sample; A graph-text joint module that, for each subgraph, through the first prompt information, enables the large model to learn the correspondence between the graph embedding features of the subgraph and the text embedding features of each gene in the subgraph, and calculates a cross-entropy loss according to the first result output by the large model, where the text embedding features are obtained by feature extraction of the pre-obtained text description information of the gene; A parameter adjustment module that adjusts the parameters of the large model according to the contrast loss and the cross-entropy loss.

9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 7 above is implemented.

10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the method described in any one of claims 1 to 7 above is implemented.

Citation Information

Patent Citations

  • Model training method and device, node embedding determination method and device, equipment and medium

    CN116702829A

  • Gene regulatory network prediction method based on multi-view attention network

    CN116913390A

  • Node detection model training method and device and computer equipment

    CN117521770A

  • Construction method and application of single cell generation type pre-training basic model

    CN118016163A

  • Method and device for improving prediction accuracy of graph neural network nodes and medium

    CN118036655A