Gene causal relationship prediction method and device, and electronic equipment

CN118230809BActive Publication Date: 2026-09-29SHENZHEN HUADA YONGSHENG INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211604773.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-13
Publication Date
2026-09-29
Estimated Expiration
2042-12-13

AI Technical Summary

Technical Problem

虽然可以根据相关性设置阈值进行过滤,但还是会引入很多的假阳性的边,从而导致基因因果关系预测的准确率与可靠性不高

Benefits of technology

[0039]综上,根据本公开提出的基因因果关系预测方法、装置及电子设备,可首先获取待进行因果关系预测的基因表达数据;利用所述基因表达数据训练基因序列预测模型,基于训练结果确定基因表达数据对应的多个基因预测序列,以及所述多个基因预测序列分别对应的贝叶斯信息准则分数值;在进行基因因果关系预测时,可按照所述贝叶斯信息准则分数值由高到低的顺序,在记录的所述多个基因预测序列中选取预设数量个目标基因预测序列;将所述预设数量个目标基因预测序列分别转换为基因图,对所述预设数量个基因图进行融合处理,得到所述基因表达数据的因果关系预测结果。本公开中的技术方案,可将深度学习算法应用于基因序列预测的模型构建,利用基因序列预测模型在考虑基因之间相关性的同时,选取贝叶斯信息准则分数值较优的多个基因图进行融合处理,可以消除不必要的假阳性的边,最终得到一个较为稳定的因果图结构,进一步增强基因因果关系预测的准确率与可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118230809B_ABST
    Figure CN118230809B_ABST
Patent Text Reader

Abstract

The present disclosure provides a gene causal relationship prediction method and device and electronic equipment, and relates to the technical field of genes, which comprises the following steps: obtaining gene expression data to be subjected to causal relationship prediction; training a gene sequence prediction model by using the gene expression data, determining a plurality of gene prediction sequences corresponding to the gene expression data and Bayesian information criterion score values corresponding to the plurality of gene prediction sequences based on a training result; selecting a preset number of target gene prediction sequences from the plurality of gene prediction sequences in a descending order of the Bayesian information criterion score values; and converting the preset number of target gene prediction sequences into gene graphs respectively, performing fusion processing on the preset number of gene graphs, and obtaining a causal relationship prediction result of the gene expression data. The present disclosure can eliminate unnecessary false positive edges by performing fusion processing on the preset number of gene graphs, and finally obtain a relatively stable directed acyclic graph structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of gene technology, specifically to a method, apparatus, and electronic device for predicting gene causal relationships. Background Technology

[0002] The construction of gene regulatory networks and the discovery of regulatory relationships are hot topics in life science research and are important contributions to advancing biological research. Traditional methods for modeling gene regulation often use statistical algorithms based on co-occurrence networks, without decoding the incidental relationships between TF genes and target genes. Alternatively, algorithms based on gene interactions are linear or tree-based models, which are difficult to apply to nonlinear models and do not fully utilize the computational power provided by deep learning.

[0003] In recent years, deep learning-based gene regulatory network modeling has also attracted widespread attention. Most deep models focus more on the interactions between genes; however, genes with related expression are not necessarily functionally related. Although thresholds can be set based on correlation for filtering, many false positives are still introduced, resulting in low accuracy and reliability of gene causal relationship prediction. Summary of the Invention

[0004] This disclosure provides a method, apparatus, and electronic device for predicting gene causality, which can enhance the accuracy and reliability of gene causality prediction.

[0005] A first aspect of this disclosure provides a method for predicting gene causality, the method comprising:

[0006] Obtain gene expression data for causal relationship prediction;

[0007] A gene sequence prediction model is trained using the gene expression data. Based on the training results, multiple gene prediction sequences corresponding to the gene expression data are determined, as well as Bayesian information criterion scores corresponding to the multiple gene prediction sequences.

[0008] According to the Bayesian information criterion scores from high to low, a preset number of target gene prediction sequences are selected from the plurality of gene prediction sequences.

[0009] The predetermined number of target gene prediction sequences are converted into gene maps, and the predetermined number of gene maps are fused to obtain the causal relationship prediction results of the gene expression data.

[0010] In some embodiments of this disclosure, the gene sequence prediction model includes at least a sampling module, an encoding module, a strategy module, a first evaluation module, and a second evaluation module. Training the gene sequence prediction model using the gene expression data includes:

[0011] The gene expression data is input into the gene sequence prediction model, and the following training process is executed iteratively until the loss function of the gene sequence prediction model is less than a preset threshold, at which point the training of the gene sequence prediction model is considered complete:

[0012] The sampling module is used to sample the gene expression data based on the cell dimension, and the sampled gene expression data is transposed.

[0013] The transposed gene expression data is encoded using the encoding module to obtain gene semantic information.

[0014] The second evaluation module is used to determine the environmental evaluation value of the gene expression data, and the strategy module is dynamically optimized based on the environmental evaluation value.

[0015] The optimized strategy module sorts the gene expression data based on the gene semantic information to obtain multiple gene prediction sequences;

[0016] The Bayesian information criterion scores corresponding to the multiple gene prediction sequences are calculated using the first evaluation module.

[0017] In some embodiments of this disclosure, determining multiple predicted gene sequences corresponding to the gene expression data based on training results, and the Bayesian information criterion scores corresponding to the multiple predicted gene sequences respectively, includes:

[0018] When the training result is training complete, the multiple gene prediction sequences corresponding to the gene expression data output by the strategy module and the Bayesian information criterion scores corresponding to the multiple gene prediction sequences output by the first evaluation module are obtained.

[0019] In some embodiments of this disclosure, converting the preset number of target gene prediction sequences into gene maps includes:

[0020] Each gene in the target gene prediction sequence is abstracted as a node in the gene graph. The edges between nodes in the gene graph are determined according to the order of the genes in the target gene prediction sequence to obtain the gene graph.

[0021] In some embodiments of this disclosure, the step of fusing the preset number of gene maps to obtain the causal relationship prediction results of the gene expression data includes:

[0022] The predetermined number of gene graphs are input into the structural equation model to obtain a set of gene graphs and a gene graph weight matrix;

[0023] Using the gene graph set or the gene graph weight matrix, the preset number of gene graphs are fused to obtain the causal relationship prediction results of the gene expression data.

[0024] In some embodiments of this disclosure, the step of fusing the preset number of gene maps using the gene map set or the gene map weight matrix to obtain the causal relationship prediction result of the gene expression data includes:

[0025] Calculate the generalized selection frequency of the gene map set;

[0026] Based on the generalized selection frequency, the hill-climbing algorithm is used to determine the optimal gene map among the preset number of gene maps, and the optimal gene map is determined as the causal relationship prediction result of the gene expression data.

[0027] In some embodiments of this disclosure, the step of fusing the preset number of gene maps using the gene map set or the gene map weight matrix to obtain the causal relationship prediction result of the gene expression data includes:

[0028] Based on the gene graph weight matrix, the maximum weight value of each edge in the preset number of gene graphs is selected;

[0029] The optimal gene graph is determined based on the maximum weight value of each selected edge, and the optimal gene graph is determined as the causal relationship prediction result of the gene expression data.

[0030] A second aspect of this disclosure provides a gene causality prediction device, characterized in that the device comprises:

[0031] The acquisition module is used to acquire gene expression data for causal relationship prediction.

[0032] The determination module uses the gene expression data to train a gene sequence prediction model, and determines multiple gene prediction sequences corresponding to the gene expression data, as well as Bayesian information criterion scores corresponding to the multiple gene prediction sequences, based on the training results.

[0033] The selection module selects a preset number of target gene prediction sequences from the multiple gene prediction sequences according to the Bayesian information criterion scores from high to low.

[0034] The fusion module converts the preset number of target gene prediction sequences into gene maps, performs fusion processing on the preset number of gene maps, and obtains the causal relationship prediction results of the gene expression data.

[0035] A third aspect of this disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the methods described in the first aspect of this disclosure.

[0036] A fourth aspect of this disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform the methods described in the first aspect of this disclosure.

[0037] A fifth aspect of this disclosure provides a computer program product including a computer program that, when executed by a processor, implements the methods described in the first aspect of this disclosure.

[0038] A sixth aspect of this disclosure provides a chip including one or more interface circuits and one or more processors; the interface circuits are configured to receive signals from the memory of an electronic device and send signals to the processors, the signals including computer instructions stored in the memory, which, when executed by the processors, cause the electronic device to perform the methods described in the first aspect of this disclosure.

[0039] In summary, according to the gene causality prediction method, apparatus, and electronic device proposed in this disclosure, gene expression data to be predicted for causality can be acquired first; a gene sequence prediction model can be trained using the gene expression data; based on the training results, multiple gene prediction sequences corresponding to the gene expression data, and Bayesian information criterion scores corresponding to the multiple gene prediction sequences, can be determined; when predicting gene causality, a preset number of target gene prediction sequences can be selected from the recorded multiple gene prediction sequences in descending order of the Bayesian information criterion scores; the preset number of target gene prediction sequences can be converted into gene graphs respectively; and the preset number of gene graphs can be fused to obtain the causality prediction result of the gene expression data. The technical solution in this disclosure can apply deep learning algorithms to the model construction of gene sequence prediction. By using the gene sequence prediction model to consider the correlation between genes while selecting multiple gene graphs with better Bayesian information criterion scores for fusion, unnecessary false positive edges can be eliminated, ultimately resulting in a more stable causal graph structure, further enhancing the accuracy and reliability of gene causality prediction.

[0040] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0041] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0042] Figure 1 A schematic flowchart illustrating a gene causality prediction method provided in this embodiment of the disclosure;

[0043] Figure 2 A flowchart illustrating another gene causality prediction method provided in this embodiment of the disclosure;

[0044] Figure 3 A schematic diagram illustrating the principle of a gene sequence prediction model provided in this embodiment of the disclosure;

[0045] Figure 4 This is a schematic diagram of the structure of a gene causality prediction device provided in an embodiment of the present disclosure;

[0046] Figure 5 This is a schematic diagram of the structure of a gene causality prediction device provided in an embodiment of this disclosure. Detailed Implementation

[0047] Embodiments of this disclosure are described in detail below, with examples of embodiments shown in the accompanying drawings, wherein the same or similar reference numerals identify the same or similar originals or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this disclosure, and should not be construed as limiting this disclosure.

[0048] In existing technologies, most deep models based on gene regulatory networks in the field of causal discovery focus on the correlation between genes. Using the CORL-2 model (a causal discovery method based on retrospective learning), the model first encodes genes using a Transformer model while simultaneously learning the relationships between them. Then, a Long Short Term Memory (LSTM) network model based on a pointer network mechanism is used for gene search and ranking, i.e., a Markov Decision Process (MDP). The environmental score is obtained by averaging the output of the Transformer model and then predicting it using the corresponding Critic model. The reward score is calculated using the Bayesian information criterion of a Directed Acyclic Graph (DAG). Information Criterion (BIC) yields the final gene causality prediction result. In the above process, it can be seen from the CORL-2 framework that the use of the mean to represent the global semantics and as the start label in the MDP strategy module [START] may not be very appropriate or reflect "fairness". Although reinforcement learning can filter by setting thresholds based on relevance, it will still introduce many false positive edges, resulting in low accuracy and reliability of gene causality prediction.

[0049] To address the aforementioned technical problems, this invention provides a method, apparatus, and electronic device for predicting gene causality, which can eliminate unnecessary false positives and enhance the accuracy and reliability of gene causality prediction.

[0050] Figure 1 This is a flowchart illustrating a gene causality prediction method according to an exemplary embodiment, such as... Figure 1 As shown, it includes the following steps.

[0051] Step 101: Obtain gene expression data for causal relationship prediction.

[0052] Among them, causal relationship prediction is a method of using the causal relationship of the development of things to infer the development trend of things. Generally, based on the historical data obtained in the past, the dependency relationship between the variable of the prediction object and the variable of its related things is found to establish a corresponding mathematical model for causal prediction. Then, the prediction is made by solving the mathematical model. In this embodiment of the disclosure, the aim is to determine the causal relationship between genes contained in the gene expression data. The gene expression data may include cell number, gene type, etc. The gene expression data can be obtained from the dataset, such as the scRNA seq dataset, which includes (1) mouse embryonic stem cells (mESC), (2) mouse dendritic cells (mDC), (3) three mouse hematopoietic stem cell lineages, including erythrocyte lineage (mHSC-E), granulocyte-macrophage lineage (mHSC-GM) and lymphocyte lineage (mHSC-L), (4) human mature hepatocytes (hHep) and (5) human embryonic stem cells (hESC), etc.

[0053] In this embodiment, given gene expression data X∈R N × M Where X represents gene expression data (columns represent genes, rows represent cells or others), N represents the number of cells in the gene expression data, and M represents the gene type in the gene expression data; in the following embodiments of this disclosure, a gene regulation model (i.e., a gene sequence prediction model) f is proposed, and the results generated each time are recorded during the training process, denoted as Y∈Mem(s,S), where S represents the currently predicted sequence, Mem represents the stored dictionary, and s is the key value of the dictionary, that is, the BIC value calculated after each sequence S is converted into a DAG. Then,

[0054] S=f(X)

[0055] Step 102: Train a gene sequence prediction model using gene expression data, and determine multiple gene prediction sequences corresponding to the gene expression data and the Bayesian information criterion scores corresponding to the multiple gene prediction sequences based on the training results.

[0056] Among them, the gene sequence prediction model is a model designed based on deep learning algorithms to transform gene expression data into a directed acyclic graph and its corresponding Bayesian information criterion score. The Bayesian information criterion is to estimate the probability of some unknown states with subjective probability under incomplete information, then correct the probability of occurrence with the Bayesian formula, and finally make the optimal decision using the expected value and the corrected probability.

[0057] For this embodiment, the CORL-2 framework can be referenced in the designed gene sequence prediction model, such as... Figure 3As shown, it may include a sampling module, a transformation module, a strategy module (MDP Decoder), a first evaluation module (Reward), and a second evaluation module (Critic Projection). The first evaluation module can be a strategy evaluation module, and the second evaluation module can be an environmental evaluation module. For example, the first evaluation module can utilize a Markov Decision Process (MDP) to make decisions about gene sequences. When using a gene sequence prediction model to determine multiple predicted gene sequences corresponding to gene expression data, and the Bayesian information criterion scores corresponding to each of the multiple predicted gene sequences, the corresponding input gene expression data X∈R... N × M First, sampling can be used to sample gene expression data. Then, the sampled data can be transposed to obtain the input data X′∈R. M × D Where D is the number of samples or the input dimension of the Transformer module; then, the Transformer module can be used for encoding to learn the interactions between genes. A new label [ENV] is added to the Transformer module, representing the global semantics of all genes at the current input time, or the optimal expected starting label of the MDP Decoder module. At this time, the output is X″∈R N+1 × D The model then proceeds through two branches: one is the Critic Projection module, which predicts the environmental score before sorting; the other is the MDP Decoder module, which generates sorted gene sequences and corresponding probability scores based on the input. The MDP Decoder module is immediately followed by the Reward module, which calculates the predicted score, primarily using the Bayesian Information Criterion (BIC) score to optimize the directed acyclic graph (DAG).

[0058] In this embodiment, the gene sequence prediction model employs a single-step training strategy. After gene expression data is input into the model, the model updates the parameters of the Actor policy function and the Critic value function, recording the results of each step until the loss function decreases to a small value. To continue training, the number of iterations needs to be increased, and the corresponding model parameter file and the Reward model-related record file (Mem) need to be loaded. Once training is complete, multiple predicted gene sequences corresponding to the gene expression data, along with the Bayesian information criterion scores for each predicted gene sequence, can be directly extracted from the record file (Mem).

[0059] Step 103: Select a preset number of target gene prediction sequences from multiple gene prediction sequences according to the Bayesian information criterion scores from high to low.

[0060] The preset quantity is the value of top_k set in advance during prediction; the target gene prediction sequence is the gene prediction sequence ranked among the top K in the Bayesian information criterion score from high to low.

[0061] In this embodiment, the same data is used for both training and prediction, and the specific input for prediction is consistent with the input for the training process. When making a prediction, the value of top_k needs to be set, and then the results recorded in the record file (Mem) are recorded. The results are sorted from high to low according to the BIC value, and the top_k gene prediction sequences are taken as the target gene prediction sequences.

[0062] Step 104: Convert the preset number of target gene prediction sequences into gene maps, and perform fusion processing on the preset number of gene maps to obtain the causal relationship prediction results of gene expression data.

[0063] The gene graph can be a Directed Acyclic Graph (DAG), which consists of nodes and lines (edges) connecting them with unidirectional arrows. "Directed" means that there is a direction, more precisely, the same direction, and "acyclic" means that it cannot form a closed loop.

[0064] In this embodiment, a preset number of target gene prediction sequences are converted into DAGs. After passing through a structural equation modeling (SEM) system, the pruned DAGs and their corresponding weight matrices are returned. By fusion processing, unnecessary false positive edges can be eliminated, and the causal relationship prediction results of gene expression data can be obtained. There are two fusion methods: one is to obtain the optimal DAG result by referring to the DAGBag fusion method in DAGBagST based on the generated DAG set; the other is to pass each DAG through a structural equation modeling (SEM) system, select the edge with the largest weight as the weight value of the current edge, and finally perform sorting and filtering to extract the edges with larger weights.

[0065] In summary, according to the gene causality prediction method proposed in this disclosure, gene expression data to be predicted for causality can be obtained first. The gene expression data is then input into a gene sequence prediction model for training. After training, the model records multiple predicted gene sequences corresponding to the gene expression data, as well as the Bayesian information criterion scores for each sequence. During gene causality prediction, a predetermined number of target gene prediction sequences are selected from the recorded sequences, ordered from highest to lowest Bayesian information criterion score. These predetermined number of target gene prediction sequences are then converted into gene graphs, and these graphs are fused to obtain the causality prediction result for the gene expression data. The technical solution in this disclosure applies deep learning algorithms to gene sequence prediction model construction. By considering the correlation between genes and selecting multiple gene graphs with superior Bayesian information criterion scores for fusion, unnecessary false positive edges can be eliminated, ultimately resulting in a more stable causal graph structure, further enhancing the accuracy and reliability of gene causality prediction.

[0066] based on Figure 1 The embodiments shown are refinements and extensions of the above embodiments. To fully illustrate the specific implementation process of the method in this embodiment, this embodiment provides the following: Figure 2 The specific method is shown. Figure 2 based on Figure 1 The illustrated embodiment further defines step 102. Figure 2 In the illustrated embodiment, step 102 includes steps 202, 203, 204, 205, and 206. For example... Figure 2 As shown, the method includes the following steps:

[0067] Step 201: Obtain gene expression data for causal relationship prediction.

[0068] For the specific implementation process of the embodiments disclosed herein, please refer to the relevant description in step 101 of the embodiment, which will not be repeated here.

[0069] Furthermore, by executing steps 202 to 206 of the subsequent embodiments, a gene sequence prediction model can be trained using gene expression data until the loss function of the gene sequence prediction model is less than a preset threshold, at which point the gene sequence prediction model training is considered complete. After the gene sequence prediction model training is completed, steps 207 to 208 of the embodiments can directly extract multiple gene prediction sequences corresponding to the gene expression data recorded during the gene sequence prediction model training process, as well as the Bayesian information criterion scores corresponding to each of the multiple gene prediction sequences, and execute the subsequent gene causality prediction process.

[0070] Step 202: Input the gene expression data into the gene sequence prediction model, use the sampling module to sample the gene expression data based on the cell dimension, and transpose the sampled gene expression data.

[0071] For embodiments of this disclosure, given input gene expression data X∈R N × M Given dimension D, D rows are randomly sampled from X using a sampling model. The sampled data is then transposed to obtain the input data X. ′ ∈R M×D Where X′ is the transposed gene expression data, D is the number of samples or the input dimension of the Transformer module, N is the number of cells in the gene expression data, M is the gene type in the gene expression data, and X is the cell dimension.

[0072] X ′ =Sample(X)

[0073] X ′ =Transpose(X) ′ )

[0074] Step 203: Use the encoding module to encode the transposed gene expression data to obtain gene semantic information.

[0075] The encoding module can be a Transformer model, a neural network model that learns context and thus meaning by tracking relationships in sequence data (such as words in a sentence). The Transformer model applies a set of evolving mathematical techniques called attention or self-attention.

[0076] For embodiments of this disclosure, the gene sequence prediction model can subsequently be applied to X. ′ Before performing Transformer(T) encoding, to better learn the global semantics of the gene and the optimal starting label of the MDP Decoder module, a variable [ENV] is defined, where [ENV] is related to X. ′ Unrelated custom parameter variables, after being encoded by Transformer, can learn X. ′ Overall semantic information; specifically, first, regarding X ′ The Embedding (E) is performed. The structure of the Embedding consists of the following units: Conv1d+RELU+Conv1d+BatchNorm1d; then it is concatenated with [ENV], and finally the Transformer is used for modeling.

[0077] X′=E(X′)

[0078] X′=Concat([ENV],X′)

[0079] X″=T(X′)

[0080] Where X″∈R N+1 × D X″ is the result of [ENV] concatenated with X′ and then encoded.

[0081] For example, the encoding part of this paper can use a 3-layer, 8-head Transformer with a sampling dimension D = 64 and an embedding dimension of 64.

[0082] Step 204: Use the second evaluation module to determine the environmental evaluation value of the gene expression data, and dynamically optimize the strategy module based on the environmental evaluation value so that the loss function of the gene sequence prediction model is less than a preset threshold.

[0083] The second assessment module can be the Critic Projection module, i.e., the environmental assessment module.

[0084] In this embodiment, the Critic Projection (C) module evaluates the environment before performing MDP Decoder policy prediction and predicts an evaluation value S based on the encoded output. ENV [ENV] represents the overall semantic information of X′. [ENV] can be directly input into the Critic Projection module to predict S. ENV The input only uses the output at the corresponding position of [ENV]. After passing through the MDP Decoder model, the predicted sorted result is evaluated by the Reward module to obtain S. BIC To avoid impacting the training process, the Critic Projection module branch does not update the Transformer model's parameters in reverse. The Critic Projection module consists of the following units: first, a linear mapping; then ReLU; and finally, another mapping to obtain the predicted environment evaluation value S. ENV

[0085] S ENV =C(X″(:,0,:))

[0086] Where S ENV ∈R 1 .

[0087] In this embodiment of the disclosure, the model adopts a deep reinforcement learning Actor-Critic architecture, which mainly consists of two losses: an update policy loss based on the advantage function. Another is the mean squared error loss between the environmental forecast and the actual strategy return. This is used to update the parameters of the Critic Projection module. For example, the batch size during training is 8, meaning it samples 8 times along dimension D, resulting in an input shape of [8, M, D]. The learning rate for the Actor is 1e-5, and the learning rate for the Critic is set to 1e-4. The Adam optimizer is used for parameter updates, with two updates per iteration: first for the Actor branch, then for the Critic branch. After approximately 2000 training iterations, the model's loss converges to a small value. To facilitate further model training, the model parameters, results, and necessary parameter values ​​are stored during training.

[0088]

[0089]

[0090]

[0091] in S represents BIC The momentum average, n represents the input sample batch, s represents the state value obtained after Transformer encoding, a represents the predicted action after passing through the MDP Decoder, i.e., the sorted gene sequence, S ENV This represents the predicted environmental value.

[0092] Step 205: Use the optimized strategy module to sort the gene expression data based on gene semantic information to obtain multiple gene prediction sequences.

[0093] The strategy module can be an MDP Decoder module, which is used to sort gene expression data.

[0094] In this embodiment of the disclosure, the MDP Decoder module mainly consists of two parts: LSTM and Pointer Network. After Transformer encoding, the model obtains X″∈R. N+1×D The output of [ENV] is used as the starting label for the LSTM input. The specific process of the MDP Decoder is as follows:

[0095] 1. The input at each time t is collectively referred to as X. queryAnd solve for the attention weight value a with the output X″ of the Transformer. t

[0096]

[0097] a t =softmax(e t )

[0098] Among them, v and W x W query and b attn These are the learning parameters.

[0099] 2. Then, combine X″ with a t Weighted summation yields the context vector.

[0100]

[0101] Indicates based on X at the current time query The contextual information understood from X″.

[0102] 3. With X query Fusion as the new X query Then solve for the attention with X″. Will Sampling is performed according to a multinomial distribution to obtain the input A at the next time step. t+1 Compared with the currently predicted probability value P t

[0103]

[0104]

[0105]

[0106] Among them, V and W X W q b′ are the learning parameters, and after M iterations, the sorted gene sequence A is obtained. t+1 ∈R M and the corresponding probability score P t ∈R M .

[0107] Step 206: Calculate the Bayesian information criterion scores corresponding to the predicted sequences of multiple genes using the first evaluation module.

[0108] The first evaluation module can be the Reward module.

[0109] In this embodiment, the Reward module is a crucial part of reinforcement learning, primarily responsible for scoring the results of the Policy Module (MDP Decoder). The model uses the Bayesian Information Criterion Score (BIC) as the scoring function for the current DAG. First, the gene sequence A obtained by the MDP Decoder is converted into a DAG(G). Since a recurrent neural network LSTM is used, a node at time t can be considered a child node at time t-1 or an indirect child node at time t-2. Therefore, a node at time t has an edge pointing to a node at time t-1 or time t-2. According to this rule, the sequence can be converted into a DAG. The DAG can be pruned according to a certain threshold, and then the BIC score S is calculated. BlC (G) is calculated, where it can be obtained through S. BIC (G) and Environmental Assessment Value S ENv The difference updates optimize the MDP Decoder model.

[0110]

[0111] |θ j | represents the parameter dimension, where Pa(X) j ) is X j The set of parent nodes, p(X) j |Pa(X j )) is given its parent terms with respect to X j The conditional probability distribution. We assume that the gene expression data X is obtained through a structural equation model (SEM) with additive noise: X j ∶=f j (Pa(X j ))+∈ j ,j=1,..,M, where f j X represents j The functional relationship between it and its parent node, ∈ i Let p(X) represent a jointly independent additive noise variable. j |Pa(X j This can be represented as the simplified information aggregation value (RSS) calculated after SEM.

[0112] RSS j =RSS(X j |Pa(X j )) is represented as X j For Pa(X) j The least squares residual sum of squares.

[0113] Step 207: Select a preset number of target gene prediction sequences from multiple gene prediction sequences according to the Bayesian information criterion scores from high to low.

[0114] In specific application scenarios, after the gene sequence prediction model is trained, multiple predicted gene sequences corresponding to the gene expression data output by the strategy module and the Bayesian information criterion scores corresponding to the multiple predicted gene sequences output by the first evaluation module can be extracted from the model-related record file (Mem). Then, according to the Bayesian information criterion scores from high to low, a predetermined number of predicted gene sequences can be selected from the extracted predicted gene sequences as target predicted gene sequences that can reflect the gene characteristics in the gene expression data.

[0115] Step 208: Convert the preset number of target gene prediction sequences into gene maps, and perform fusion processing on the preset number of gene maps to obtain the causal relationship prediction results of gene expression data.

[0116] In the embodiments of this disclosure, when converting a preset number of target gene prediction sequences into gene graphs, the specific steps may include: abstracting each gene in the target gene prediction sequence into a node in the gene graph, determining the edges between nodes in the gene graph according to the sorting order of the genes in the target gene prediction sequence, and obtaining the gene graph.

[0117] In the embodiments of this disclosure, when fusing a preset number of gene graphs to obtain causal relationship prediction results of gene expression data, the specific steps may include: inputting a preset number of gene graphs into a structural equation model to obtain a set of gene graphs and a gene graph weight matrix; using the set of gene graphs or the gene graph weight matrix to fuse the preset number of gene graphs to obtain causal relationship prediction results of gene expression data.

[0118] In this embodiment of the disclosure, during training, the model stores the BIC value and its corresponding gene sequence for each step. During the prediction phase, the cache is directly read. After sorting the BIC values, the top k optimal sequences are selected. These sequences are then converted into a DAG graph and input into a structural equation model (SEM) to obtain the final DAG result. The DAG set and DAG weight set in the DAG result are used as input to the DAG-Fusion module. The DAG learning process is typically highly variable—even small perturbations in the data can cause significant changes in the learned graph. The DAG-Fusion module is a method for fusing multiple DAGs into a single, superior DAG. We primarily design two methods to achieve this fusion, which are described below:

[0119] 1. DAGBag: DAGBag is similar to bagging fusion in machine learning. First, it solves the generalized selection frequency (GSF) of the input DAG set.e Then, following the steps of the climbing algorithm with aggregated scores, the optimal DAG is calculated. Assume the set G of the DAGs... e ={g b :b=1,2,…,k},gsf e The calculation is as follows

[0120] sf e ={g b :e∈E(g b )} / k

[0121]

[0122] Among them sf e E(g) represents the frequency at which edge e is selected. b ) represents DAG g b Let e* be the edge set, where e* represents the edge in the opposite direction to e, and α is a hyperparameter.

[0123] 2. Take the maximum value: First, merge the weight matrix set and select the largest weight value of the corresponding edge in the k weight matrices as the weight value of the current edge.

[0124] Correspondingly, when using a gene graph set or gene graph weight matrix to fuse a predetermined number of gene graphs to obtain causal relationship prediction results for gene expression data:

[0125] As one possible implementation, the steps of the embodiment may specifically include: calculating the generalized selection frequency of the gene map set; based on the generalized selection frequency, using a hill-climbing algorithm to determine the optimal gene map from a preset number of gene maps, and determining the optimal gene map as the causal relationship prediction result of gene expression data.

[0126] As one possible implementation, the steps of the embodiment may specifically include: filtering the maximum weight value of each edge in a preset number of gene graphs according to the gene graph weight matrix; determining the optimal gene graph based on the maximum weight value of each edge selected, and determining the optimal gene graph as the causal relationship prediction result of gene expression data.

[0127] For the gene causality prediction process of the embodiments of this disclosure, please refer to [link to relevant documentation]. Figure 3The schematic diagram of the gene sequence prediction model shown illustrates the principle. First, gene expression data for causal relationship prediction is acquired. This data is then input into the gene sequence prediction model for training. The training process can be as follows: The Sampling module samples gene expression data at the cell level and transposes the sampled data; the Transformer encoding module encodes the transposed gene expression data to obtain gene semantic information; the Critic Projection module determines the environmental assessment value of the gene expression data; based on this assessment value, the strategy module dynamically optimizes the strategy to ensure the loss function of the gene sequence prediction model is less than a preset threshold; and finally, MDP is used... The Decoder module sorts gene expression data based on gene semantic information to obtain multiple gene prediction sequences. The Reward module calculates the Bayesian information criterion scores for each of the multiple gene prediction sequences. After the gene sequence prediction model is trained, it records the multiple gene prediction sequences corresponding to the gene expression data, as well as the Bayesian information criterion scores for each sequence. When predicting gene causal relationships, the DAG-Fusion module selects a preset number of target gene prediction sequences from the multiple gene prediction sequences, arranged in descending order of Bayesian information criterion scores. These preset number of target gene prediction sequences are then converted into gene graphs. Finally, the preset number of gene graphs are fused to obtain the causal relationship prediction results for the gene expression data.

[0128] In this embodiment, gene expression data for causal relationship prediction is first acquired. The gene expression data is then input into a gene sequence prediction model for training. After training, the model records multiple predicted gene sequences corresponding to the gene expression data, along with their corresponding Bayesian information criterion scores. During gene causal relationship prediction, a predetermined number of target gene prediction sequences are selected from the recorded sequences, ordered from highest to lowest Bayesian information criterion score. These predetermined number of target gene prediction sequences are then converted into gene graphs, and these gene graphs are fused to obtain the causal relationship prediction result for the gene expression data. This approach allows deep learning algorithms to be applied to gene sequence prediction model construction. By considering the correlation between genes and selecting multiple gene graphs with superior Bayesian information criterion scores for fusion, unnecessary false positive edges can be eliminated, ultimately resulting in a more stable causal graph structure, further enhancing the accuracy and reliability of gene causal relationship prediction.

[0129] Based on the above Figures 1-3 The specific implementation of the method shown in this embodiment provides a gene causality prediction device, such as... Figure 4 As shown, the device includes: an acquisition module 41, a determination module 42, a selection module 43, and a fusion module 44.

[0130] The acquisition module 41 is used to acquire gene expression data for causal relationship prediction.

[0131] Module 42 is defined to train a gene sequence prediction model using gene expression data, and to determine multiple gene prediction sequences corresponding to the gene expression data and the Bayesian information criterion scores corresponding to the multiple gene prediction sequences based on the training results.

[0132] Select module 43 and select a preset number of target gene prediction sequences from multiple gene prediction sequences in descending order of Bayesian information criterion scores.

[0133] The fusion module 44 converts a preset number of target gene prediction sequences into gene maps, performs fusion processing on the preset number of gene maps, and obtains the causal relationship prediction results of gene expression data.

[0134] In some embodiments of this disclosure, such as Figure 5 As shown, the device also includes: a training module 45;

[0135] The gene sequence prediction model includes at least a sampling module, an encoding module, a strategy module, a first evaluation module, a second evaluation module, and a training module 45. This module is used to input gene expression data into the gene sequence prediction model and iteratively execute the following training process until the loss function of the gene sequence prediction model is less than a preset threshold, at which point the gene sequence prediction model training is considered complete: The sampling module samples gene expression data based on the cell dimension and transposes the sampled gene expression data; the encoding module encodes the transposed gene expression data to obtain gene semantic information; the second evaluation module determines the environmental evaluation value of the gene expression data and dynamically optimizes the strategy module based on the environmental evaluation value; the optimized strategy module sorts the gene expression data based on the gene semantic information to obtain multiple gene prediction sequences; and the first evaluation module calculates the Bayesian information criterion scores corresponding to each of the multiple gene prediction sequences.

[0136] In some embodiments of this disclosure, the determining module 42 can be used to determine, after the gene sequence prediction model has been trained, multiple gene prediction sequences corresponding to the gene expression data output by the strategy module, and Bayesian information criterion scores corresponding to the multiple gene prediction sequences output by the first evaluation module.

[0137] In some embodiments of this disclosure, the fusion module 44 can be used to abstract each gene in the target gene prediction sequence as a node in a gene graph, determine the edges between nodes in the gene graph according to the sorting order of genes in the target gene prediction sequence, and obtain a gene graph.

[0138] In some embodiments of this disclosure, the fusion module 44 can be used to input a preset number of gene graphs into a structural equation model to obtain a set of gene graphs and a gene graph weight matrix; and use the set of gene graphs or the gene graph weight matrix to perform fusion processing on the preset number of gene graphs to obtain causal relationship prediction results of gene expression data.

[0139] In some embodiments of this disclosure, the fusion module 44 can be used to calculate the generalized selection frequency of the gene map set; based on the generalized selection frequency, the optimal gene map is determined from a preset number of gene maps using a hill-climbing algorithm, and the optimal gene map is determined as the causal relationship prediction result of gene expression data.

[0140] In some embodiments of this disclosure, the fusion module 44 can also be used to filter the maximum weight value of each edge in a preset number of gene graphs according to the gene graph weight matrix; determine the optimal gene graph based on the maximum weight value of each edge selected, and determine the optimal gene graph as the causal relationship prediction result of gene expression data.

[0141] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0142] The technical solution disclosed herein can apply deep learning algorithms to the construction of gene sequence prediction models. By using gene sequence prediction models to select multiple gene graphs with better Bayesian information criterion scores for fusion processing while considering the correlation between genes, unnecessary false positive edges can be eliminated, and a more stable causal graph structure can be obtained, which further enhances the accuracy and reliability of gene causal relationship prediction.

[0143] The methods and apparatus provided in the embodiments of this application have been described above. To implement the functions of the methods provided in the embodiments of this application, the electronic device may include a hardware structure and software modules, and may implement the above functions in the form of a hardware structure, software modules, or a hardware structure plus software modules. One of the above functions may be executed in the form of a hardware structure, software modules, or a hardware structure plus software modules.

[0144] Embodiments of this disclosure also propose a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the gene causality prediction method described in the above embodiments of this disclosure.

[0145] Embodiments of this disclosure also provide a computer program product, including a computer program that is executed by a processor using the gene causality prediction method described in the above embodiments of this disclosure.

[0146] Embodiments of this disclosure also propose a chip including one or more interface circuits and one or more processors; the interface circuits are used to receive signals from the memory of an electronic device and send signals to the processors, the signals including computer instructions stored in the memory, which, when executed by the processor, cause the electronic device to perform the gene causality prediction method described in the above embodiments of this disclosure.

[0147] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0148] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with an embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0149] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of preferred embodiments of this disclosure includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the function involved, as will be understood by those skilled in the art to which embodiments of this disclosure pertain.

[0150] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processing module, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (control method), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic device, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0151] It should be understood that various parts of the embodiments of this disclosure can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0152] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0153] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into a single processing module, or each unit can exist physically separately, or two or more units can be integrated into a single module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The aforementioned storage medium can be a read-only memory, a hard disk, or an optical disk, etc.

[0154] Although embodiments of the present disclosure have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present disclosure.

Claims

1. A method for predicting gene causality, characterized in that, include: Obtain gene expression data for causal relationship prediction; A gene sequence prediction model is trained using the gene expression data. Based on the training results, multiple gene prediction sequences corresponding to the gene expression data are determined, as well as the Bayesian information criterion scores corresponding to the multiple gene prediction sequences. According to the Bayesian information criterion scores from high to low, a preset number of target gene prediction sequences are selected from the plurality of gene prediction sequences. The predetermined number of target gene prediction sequences are converted into gene maps, and the predetermined number of gene maps are fused to obtain the causal relationship prediction results of the gene expression data.

2. The method according to claim 1, characterized in that, The gene sequence prediction model includes at least a sampling module, an encoding module, a strategy module, a first evaluation module, and a second evaluation module. Training the gene sequence prediction model using the gene expression data includes: The gene expression data is input into the gene sequence prediction model, and the following training process is executed iteratively until the loss function of the gene sequence prediction model is less than a preset threshold, at which point the training of the gene sequence prediction model is considered complete: The sampling module is used to sample the gene expression data based on the cell dimension, and the sampled gene expression data is transposed. The transposed gene expression data is encoded using the encoding module to obtain gene semantic information. The second evaluation module is used to determine the environmental evaluation value of the gene expression data, and the strategy module is dynamically optimized based on the environmental evaluation value. The optimized strategy module sorts the gene expression data based on the gene semantic information to obtain multiple gene prediction sequences; The Bayesian information criterion scores corresponding to the multiple gene prediction sequences are calculated using the first evaluation module.

3. The method according to claim 2, characterized in that, The step of determining multiple predicted gene sequences corresponding to the gene expression data based on the training results, and the Bayesian information criterion scores corresponding to the multiple predicted gene sequences respectively, includes: When the training result is training complete, the multiple gene prediction sequences corresponding to the gene expression data output by the strategy module and the Bayesian information criterion scores corresponding to the multiple gene prediction sequences output by the first evaluation module are obtained.

4. The method according to claim 1, characterized in that, The step of converting the preset number of target gene prediction sequences into gene maps includes: Each gene in the target gene prediction sequence is abstracted as a node in the gene graph. The edges between nodes in the gene graph are determined according to the order of the genes in the target gene prediction sequence to obtain the gene graph.

5. The method according to claim 1, characterized in that, The step of fusing the preset number of gene maps to obtain the causal relationship prediction results of the gene expression data includes: The predetermined number of gene graphs are input into the structural equation model to obtain a set of gene graphs and a gene graph weight matrix; Using the gene graph set or the gene graph weight matrix, the preset number of gene graphs are fused to obtain the causal relationship prediction results of the gene expression data.

6. The method according to claim 5, characterized in that, The step of fusing the preset number of gene maps using the gene map set or the gene map weight matrix to obtain the causal relationship prediction result of the gene expression data includes: Calculate the generalized selection frequency of the gene map set; Based on the generalized selection frequency, the hill-climbing algorithm is used to determine the optimal gene map among the preset number of gene maps, and the optimal gene map is determined as the causal relationship prediction result of the gene expression data.

7. The method according to claim 5, characterized in that, The step of fusing the preset number of gene maps using the gene map set or the gene map weight matrix to obtain the causal relationship prediction result of the gene expression data includes: Based on the gene graph weight matrix, the maximum weight value of each edge in the preset number of gene graphs is selected; The optimal gene graph is determined based on the maximum weight value of each selected edge, and the optimal gene graph is determined as the causal relationship prediction result of the gene expression data.

8. A gene causality prediction device, characterized in that, include: The acquisition module is used to acquire gene expression data for causal relationship prediction. The determination module uses the gene expression data to train a gene sequence prediction model, and determines multiple gene prediction sequences corresponding to the gene expression data, as well as Bayesian information criterion scores corresponding to the multiple gene prediction sequences, based on the training results. The selection module selects a preset number of target gene prediction sequences from the multiple gene prediction sequences according to the Bayesian information criterion scores from high to low. The fusion module converts the preset number of target gene prediction sequences into gene maps, performs fusion processing on the preset number of gene maps, and obtains the causal relationship prediction results of the gene expression data.

9. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-7.

11. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-7.

12. A chip, characterized in that, The device includes one or more interface circuits and one or more processors; the interface circuits are configured to receive signals from the memory of the electronic device and send the signals to the processors, the signals including computer instructions stored in the memory, which, when executed by the processors, cause the electronic device to perform the method of any one of claims 1-7.