Synthetic lethal gene pair prediction method and system, electronic equipment and storage medium

By extracting and aggregateing the relevant information of gene pairs from the knowledge graph and generating updated feature vectors, the problem of difficulty in extracting gene pairs in the prior art is solved, and the accuracy of prediction of synthetic lethality probability is improved.

CN120015126APending Publication Date: 2025-05-16PEKING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411569347.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-05
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to effectively extract the interaction characteristics between gene pairs, resulting in low accuracy in predicting the probability of synthesis lethality.

Method used

By extracting joint sub-maps related to the target gene pair from the knowledge graph, performing information aggregation, an updated set of feature vectors is obtained, and the synthetic lethal probability of the target gene pair is calculated based on these feature vectors.

Benefits of technology

The prediction accuracy of synthetic lethality probability is improved, and information in the gene network is systematically integrated, complex intergenic interactions are captured, and more expressive feature representations are generated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015126A_ABST
    Figure CN120015126A_ABST
Patent Text Reader

Abstract

The invention provides a synthetic lethal gene pair prediction method and system, electronic equipment and a storage medium, and belongs to the technical field of bioinformatics, and the method comprises the steps: extracting a first joint subgraph related to a target gene pair from a knowledge graph; aggregating information in the first combined sub-graph to obtain a second combined sub-graph; based on the second joint subgraph, obtaining a node-level feature vector of the target gene; obtaining a final feature vector of the target gene pair based on the node-level feature vector and the second feature vector set; and based on the final feature vector, obtaining the synthesis lethal probability of the target gene pair. According to the method, the gene network information and the multi-omics data in the knowledge graph are integrated, and the multi-level feature representation of the target gene pair is constructed and optimized step by step, so that the prediction accuracy of the synthesis lethal probability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bioinformatics, and in particular to a method, system, electronic device and storage medium for predicting synthetic lethal gene pairs. Background Art

[0002] In the field of genomics, with the rapid development of high-throughput sequencing technology, researchers can identify potential gene interactions from massive amounts of genetic data. Among them, synthetic lethality (SL) has become an important research direction. Synthetic lethality means that when two genes mutate alone, they will not affect cell survival, but when both genes mutate at the same time, they will cause cell death.

[0003] In the current technical scenario, the prediction based on synthetic lethality mainly relies on biological information such as gene expression data and mutation data. However, these data are often high-dimensional and have complex nonlinear relationships. Many existing algorithms have difficulty in effectively extracting the interaction features between gene pairs and ignore the correlation between feature representations, which makes the prediction accuracy of synthetic lethality probability low.

[0004] Therefore, how to improve the prediction accuracy of the probability of synthetic lethality has become a technical problem that needs to be solved urgently. Summary of the invention

[0005] The present invention provides a method, system, electronic device and storage medium for predicting synthetic lethal gene pairs, which are used to solve the defects in the prior art and improve the prediction accuracy of synthetic lethality probability.

[0006] The present invention provides a method for predicting synthetic lethal gene pairs, comprising the following steps: Extracting a first joint subgraph related to the target gene pair from the knowledge graph; the first joint subgraph includes a first feature vector set of all nodes related to the target gene pair; Aggregate the information in the first joint subgraph to obtain a second joint subgraph; the second joint subgraph includes a second feature vector set, and the second feature vector set is the first feature vector set updated during the aggregation process; Based on the second joint subgraph, obtaining a node-level feature vector of the target gene; Based on the node-level feature vector and the second feature vector set, obtaining a final feature vector of the target gene pair; Based on the final feature vector, the synthetic lethality probability of the target gene pair is obtained.

[0007] According to a method for predicting synthetic lethal gene pairs provided by the present invention, aggregating the information in the first joint subgraph to obtain a second joint subgraph specifically includes: Step S101: determining the 0th layer feature vector of each node in the first joint subgraph; Step S102: Calculate the attention weight based on the i-1th layer feature vector of each node and the i-1th layer feature vector of each node's neighboring node, where i∈[1,L], L is a natural number representing the number of layers of the network; Step S103: Based on the attention weight, weight the i-1th layer feature vectors of the neighbor nodes to obtain an aggregated neighbor feature vector; Step S104: combining the i-1th layer feature vector of each node with the aggregated neighbor feature vector to obtain the i-th layer feature vector of each node; Step S105: Repeat steps S102 to S104 until the L-th layer feature vector of each node is obtained, and based on the L-th layer feature vector of each node, a second feature vector set of each node is obtained; Step S106: obtaining the second joint subgraph based on all the second feature vector sets.

[0008] According to a synthetic lethal gene pair prediction method provided by the present invention, the step of determining the 0th layer feature vector of each node in the first joint subgraph specifically includes: Obtaining an initial feature vector for each node in the first joint subgraph using the TransE embedding method; Based on the position information of each node and the target gene pair, a position feature vector of each node is obtained; Each of the position feature vectors is connected to each of the initial feature vectors to generate a 0th layer feature vector for each node.

[0009] According to a synthetic lethal gene pair prediction method provided by the present invention, obtaining a node-level feature vector of a target gene based on the second joint subgraph specifically includes: Determine a second feature vector set for each node in the second joint subgraph, the second feature vector set includes feature vectors of L layers, where L is a natural number used to represent the number of layers of the network; Based on the attention mechanism, the feature vector of the L layer is updated to obtain an updated feature vector of the L layer; Perform vector connection between the updated feature vector of the 0th layer and the updated feature vector of the Lth layer to obtain the node-level feature vector.

[0010] According to a method for predicting synthetic lethal gene pairs provided by the present invention, the second feature vector set includes feature vectors of L layers, wherein L is a natural number used to represent the number of layers of the network; the final feature vector of the target gene pair is obtained based on the node-level feature vector and the second feature vector set, specifically comprising: Aggregating all L-layer feature vectors of each node in the second joint subgraph to generate a third feature vector set; Encoding the omics features of each node in the target gene pair based on a multi-layer perceptron, and generating omics feature vectors of each node in the target gene pair respectively; Combining each of the omics feature vectors by Kronecker product to obtain the omics feature representation of the target gene pair; Connecting the node-level feature vector, the third feature vector set, and the omics feature representation in series to obtain a multi-level feature vector; Based on the multi-level feature vectors, the final feature vector is obtained.

[0011] According to a method for predicting synthetic lethal gene pairs provided by the present invention, obtaining the synthetic lethality probability of the target gene pair based on the final feature vector specifically includes: The final feature vector is input into a synthetic lethal probability prediction model, and the final feature vector is linearly transformed using a weight matrix of a decoder in the synthetic lethal probability prediction model to calculate the synthetic lethal probability; wherein the weight matrix is ​​used to adjust the mapping relationship between the final feature vector and the synthetic lethal probability.

[0012] According to a synthetic lethal gene pair prediction method provided by the present invention, the synthetic lethality probability prediction model is trained by the following method: The total loss function of the synthetic lethality probability prediction model is set, and the total loss function is expressed by the following formula: ; ; ; Among them, K is the total loss; is the cross entropy loss; is the regularization loss; is the regularization coefficient; is the total number of target gene pair samples in the training set; is the true label of the target gene pair sample (u, v), where the target gene pair sample (u, v) is a synthetic lethal gene pair, then ,otherwise, ; The synthetic lethality probability of the target gene for the sample (u, v) predicted by the synthetic lethality probability prediction model; is the set of all weight parameters in the synthetic lethality probability prediction model; The initial model is trained with minimizing the total loss as the training goal, the weight matrix of the decoder is iteratively updated, and the synthetic lethal probability prediction model is obtained through training.

[0013] According to a method for predicting synthetic lethal gene pairs provided by the present invention, the target gene pair includes a first target gene and a second target gene; the step of extracting a first joint subgraph related to the target gene pair from the knowledge graph specifically includes: Based on the knowledge graph, constructing a first gene subgraph of the first target gene and a second gene subgraph of the second target gene respectively; The first gene subgraph and the second gene subgraph are merged to obtain the first joint subgraph.

[0014] The present invention also provides a synthetic lethal gene pair prediction system, comprising the following modules: a joint subgraph extraction module, an information aggregation module, a first processing module, a second processing module and a synthetic lethal probability prediction module; The joint subgraph extraction module is used to extract a first joint subgraph related to the target gene pair from the knowledge graph; the first joint subgraph includes a first feature vector set of all nodes related to the target gene pair; The information aggregation module is used to aggregate the information in the first joint subgraph to obtain a second joint subgraph; the second joint subgraph includes a second feature vector set, and the second feature vector set is the first feature vector set updated during the aggregation process; The first processing module is used to obtain a node-level feature vector of the target gene based on the second joint subgraph; The second processing module is used to obtain a final feature vector of the target gene pair based on the node-level feature vector and the second feature vector set; The synthetic lethality probability prediction module is used to obtain the synthetic lethality probability of the target gene pair based on the final feature vector.

[0015] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for predicting synthetic lethal gene pairs as described above is implemented.

[0016] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the synthetic lethal gene pair prediction methods described above.

[0017] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method for predicting synthetic lethal gene pairs as described in any one of the above is implemented.

[0018] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: By extracting the first joint subgraph related to the target gene pair from the knowledge graph, the first joint subgraph contains all nodes related to the target gene pair and their feature vector sets, so that the gene information directly or indirectly related to the target gene pair in the gene network can be systematically integrated. The complex interaction relationship between genes is effectively captured. Further, by aggregating the information in the first joint subgraph, a second joint subgraph is generated, and an updated second feature vector set is obtained, thereby further refining the association information between genes. Information aggregation updates and integrates the node features related to the target gene pair through multi-level calculations of the graph neural network. This process can dynamically capture the neighborhood information related to the target gene and gradually strengthen important gene nodes. Based on the second joint subgraph, the node-level feature vector of the target gene is obtained, thereby providing accurate local information for synthetic lethality prediction. The node-level feature vector not only reflects the position and direct neighbor information of the target gene in the knowledge graph, but also includes the upstream and downstream gene relationships updated through the aggregation process. This feature extraction ensures that the individual characteristics of each target gene can be fully retained and provides local-level gene function information for subsequent gene pair feature representation. The final feature vector of the target gene pair is further obtained based on the node-level feature vector and the second feature vector set, thereby realizing the multi-level feature fusion of the target gene pair. The final feature vector contains the local information of the gene pair, the global network features, and the functional information extracted through the omics data. This comprehensive feature representation can fully capture the complex interactive relationship between the target gene pairs, providing an informative and expressive feature representation for subsequent prediction tasks, which helps to improve the prediction accuracy of the model. Finally, by obtaining the synthetic lethality probability of the target gene pair based on the final feature vector, the complex feature representation is mapped into an interpretable probability value, thereby improving the prediction accuracy of the synthetic lethality probability. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0020] Figure 1 This is one of the flow charts of the method for predicting synthetic lethal gene pairs provided by the present invention.

[0021] Figure 2 This is the second schematic flow chart of the synthetic lethal gene pair prediction method provided by the present invention.

[0022] Figure 3 It is a schematic diagram comparing the prediction effects of the LSAG provided by the present invention and the existing models.

[0023] Figure 4 It is a structural schematic diagram of the synthetic lethal gene pair prediction model based on graph neural network provided by the present invention.

[0024] Figure 5 It is a schematic diagram for comparing LSAG model variants provided by the present invention.

[0025] Figure 6 It is a schematic diagram of the visualization results of the joint subgraph and attention coefficient provided by the present invention.

[0026] Figure 7 It is a schematic diagram of the enrichment analysis results provided by the present invention.

[0027] Figure 8 It is a schematic diagram of the structure of the synthetic lethal gene pair prediction system provided by the present invention.

[0028] Fig. 9 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0029] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0030] It should be noted that, in the description of the present invention, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "include one..." do not exclude the existence of other identical elements in the process, method, article or device including the elements. The orientation or positional relationship indicated by the terms "upper", "lower", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the system or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances.

[0031] The terms "first", "second", etc. in the present invention are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type, and the number of objects is not limited. For example, the first object can be one or more.

[0032] Combine the following Figure 1-Figure 9 The invention describes a method, system, electronic device and storage medium for predicting synthetic lethal gene pairs.

[0033] Figure 1 FIG. 1 is one of the flow charts of the method for predicting synthetic lethal gene pairs provided by the present invention, such as Figure 1 As shown, including but not limited to the following steps: Step 1: Extract the first joint subgraph related to the target gene pair from the knowledge graph; the first joint subgraph includes the first feature vector set of all nodes related to the target gene pair.

[0034] In the embodiment of step 1, the main purpose is to extract the first joint subgraph related to the target gene pair from the knowledge graph so that the information in these subgraphs can be aggregated and analyzed in subsequent steps, thereby improving the accuracy of synthetic lethal gene pair prediction.

[0035] In a possible implementation, step S1 specifically includes the following steps: Based on the knowledge graph, constructing a first gene subgraph of the first target gene and a second gene subgraph of the second target gene respectively; The first gene subgraph and the second gene subgraph are merged to obtain a first joint subgraph.

[0036] Specifically, first, the neighbor nodes and associated edges of each target gene are extracted from the revised SynLethKG knowledge graph. SynLethKG is a graph containing a variety of biomedical entities (such as genes, proteins, diseases, etc.) and their relationships. Through the rich relationship data it provides, it can comprehensively reflect the potential interactions and functional associations between genes. For each gene in the target gene pair, that is, the first target gene (hereinafter referred to as target gene v) and the second target gene (hereinafter referred to as target gene u), its corresponding gene subgraph needs to be constructed separately. The purpose of constructing these subgraphs is to capture the upstream and downstream gene relationships of the target genes in the knowledge graph, especially those genes that may be related to the synthetic lethal mechanism. In this way, more comprehensive feature data can be provided for subsequent feature aggregation and information update steps.

[0037] In the specific operation, it is first necessary to determine the k-hop neighbor nodes and their associated edges starting from the target gene v to form the subgraph of the target gene v. The k-hop neighbors here represent gene nodes that have a certain path length with the target gene v. These gene nodes may be associated with the target gene through some biological process, signaling pathway or protein interaction. Similarly, the subgraph of gene u is constructed in the same way to extract the k-hop neighbor nodes and their associated edges related to it. This method of expanding through the node neighborhood can further expand its upstream and downstream associations in the biological network while retaining information directly related to the target gene, thereby better capturing the complexity of synthetic lethal relationships. Target gene The subgraph of is represented as: ; in, represents the shortest path distance between two nodes, and Represents the edge in the knowledge graph two nodes.

[0038] Next, the first gene subgraph and the second gene subgraph are merged to obtain the first joint subgraph. Since synthetic lethality is usually not the result of a single gene or direct interaction, but the result of the interaction of multiple genes in a complex biological network. Therefore, analyzing the subgraph of each target gene separately may miss the regulatory pathways or functional modules in which they are jointly involved. By merging the subgraphs of the two target genes into a joint subgraph, the potential functional associations and interactions between them can be captured in the same graph structure, thereby improving the model's perception of the synthetic lethality mechanism.

[0039] Specifically, in the first step, the k-hop neighbor nodes and related edges of the first target gene and the second target gene are extracted from the revised SynLethKG knowledge graph, respectively, to construct their subgraphs. Next, the two subgraphs are merged. The merging process is not a simple superposition of two subgraphs, but to consider their interactive relationships in the knowledge graph, especially those important pathways and genes that may be involved in synthetic lethality. For example, some genes may appear in both the first gene subgraph and the second gene subgraph. These intersections are often the key areas where the two target genes interact in the biological network. By merging these intersection nodes and edges, the potential functional relationship between the two target genes can be better captured. Merge the two nodes and Construct a joint subgraph from the subgraph of, such as: .

[0040] Step 2: Aggregate the information in the first joint subgraph to obtain a second joint subgraph; the second joint subgraph includes a second feature vector set, and the second feature vector set is the first feature vector set updated during the aggregation process.

[0041] Since in the first joint subgraph, the feature vectors of the target genes u and v and their related nodes are only preliminary representations, they mainly reflect their positional relationships and basic properties in the knowledge graph. However, the essence of synthetic lethal relationships usually involves complex interactions between genes and their upstream and downstream regulatory networks. Therefore, relying solely on the initial node features is not enough to fully reveal these potential functional relationships. Through information aggregation, the information of neighboring nodes can be integrated into the feature representation of the target node, so that the feature vector of the target gene can contain a wider range of functional and interactive information related to it. Therefore, in the embodiment of step 2, the purpose is to aggregate the information in the first joint subgraph to generate a more efficient second joint subgraph, which includes an updated second feature vector set. Through this information aggregation process, the complex associations between target gene pairs and their neighborhoods can be better captured, and the prediction accuracy of synthetic lethal gene pairs can be improved.

[0042] In a possible implementation, step 2 specifically includes the following steps: Step S101: Determine the 0th-level feature vector of each node in the first joint subgraph.

[0043] In the embodiment of step S101, the goal is to construct a multi-layer feature vector set for each node in the first joint subgraph, thereby providing a basis for the subsequent attention mechanism and feature aggregation process. By generating multi-level feature vectors, the neighborhood information of different levels of nodes can be gradually captured, thereby improving the ability to characterize the potential synthetic lethal relationship between target gene pairs.

[0044] In a possible implementation, step S101 specifically includes the following steps: Use the TransE embedding method to obtain the initial feature vector of each node in the first joint subgraph; Based on the position information of each node and the target gene pair, the position feature vector of each node is obtained; Connect each position feature vector with each initial feature vector to generate the 0th layer feature vector of each node.

[0045] Specifically, the core reason for constructing a multi-layer feature vector is that the relationship between synthetic lethal gene pairs usually depends not only on direct interactions between genes, but also involves a wide range of upstream and downstream gene networks. The construction of a multi-layer feature vector can capture the neighborhood relationship of the target gene at different levels, thereby better reflecting the complex gene interaction pattern.

[0046] At the initial level, the TransE embedding method is used to generate the initial feature vector of each node. The TransE method embeds the entities and relationships in the knowledge graph into the vector space so that the relationship of each gene can be represented by the vector difference. Specifically, for each node i in the joint subgraph, its initial feature vector is obtained through TransE, and the position vector also needs to be calculated , the formula is as follows: ; in, and Respectively represent the shortest path distance between node i and genes u and v in the target gene pair. In this way, the relative distance between the node and the target gene pair is quantified as a position feature vector. These position feature vectors and the initial feature vector are combined through the concatenation operation to generate the 0th layer feature vector of each node: ; The generated layer 0 feature vector not only contains the basic information of the node in the knowledge graph, but also reflects its relative position relationship with the target gene. This position-aware feature representation can better help the model understand the structural association between the target gene pair and its neighborhood.

[0047] Step S102: Calculate the attention weight based on the i-1th layer feature vector of each node and the i-1th layer feature vector of each node's neighboring node, where i∈[1,L], L is a natural number used to represent the number of layers of the network.

[0048] Specifically, the calculation of attention weights depends on the feature vectors of each node and its neighboring nodes. In the Lth layer, the set of neighboring nodes of node u is recorded as , where r represents the type of relationship in the joint subgraph. For each pair of neighbor nodes u and v, the attention weight of its i-th layer Calculated by the following formula: ; ; Among them, σ represents the sigmoid activation function, and is a trainable attention weight matrix, and is the feature vector of nodes u and v at layer i-1, is the embedding vector of relation r, Represents a feature join operation.

[0049] The core idea of ​​this formula is to calculate the correlation between nodes u and v by performing nonlinear transformation on the embedding of the feature vectors of nodes u and v and the edge (i.e., relationship) between them. The stronger the correlation, the higher the attention weight. This attention mechanism through feature concatenation and embedding operations can capture the complex nonlinear relationship between nodes.

[0050] Next, the attention weights are normalized by the softmax function to ensure that the sum of the weights of all neighboring nodes of each node is 1. In this way, node u can weightedly aggregate the features of all neighboring nodes according to their attention weights.

[0051] Step S103: Based on the attention weight, weight the i-1th layer feature vectors of the neighbor nodes to obtain an aggregated neighbor feature vector.

[0052] Specifically, in the previous step, the attention weight has been calculated by the formula , now we need to perform weighted aggregation on the features of neighboring nodes according to these weights. For node u in the Lth layer, the feature vector of its neighboring nodes in the i-1th layer The attention weight Weighted to get the aggregated neighbor feature vector .

[0053] Step S104: combining the i-1th layer feature vector of each node with the aggregated neighbor feature vectors to obtain the i-th layer feature vector of each node.

[0054] Specifically, next, the feature vector of node u at layer i-1 is and the aggregated neighbor feature vector Combine them to get the updated i-th layer feature vector .

[0055] That is, when L>0, the feature vector will be updated according to the feature vector of the previous layer. Specifically, the feature vector of node u in layer L is It is calculated based on the feature vector of the L-1th layer. The feature vector is updated through the graph neural network and attention mechanism. For the feature representation of the Lth layer, the formula is as follows: Among them, W (L) is the trainable weight matrix of layer L, is the neighbor node feature of node u aggregated at layer L in the joint subgraph. In this way, node u not only retains its own feature vector, but also combines the feature information of its neighbor nodes. This step gradually introduces neighborhood information at different levels into the feature vector by aggregating information layer by layer, thereby enhancing the feature expression ability of the model.

[0056] Finally, after constructing multiple layers of feature vectors, each node in the first joint subgraph has a set of feature vectors containing multiple layers of information. These feature vectors not only contain the initial information and location information of the node, but also undergo multiple information aggregations and gradually incorporate the characteristics of neighboring nodes. This approach greatly improves the model's ability to understand the complex interactions of target gene pairs.

[0057] This feature update method ensures the progressiveness of feature vectors at each layer. By aggregating and updating node features layer by layer, the model can gradually capture the complex relationships between nodes. In the early layers, the model mainly focuses on the direct neighbors of the nodes, and as the number of layers increases, the model can capture deeper relationships between nodes. This method is very suitable for capturing the collaborative effects between multiple genes, thus providing strong support for the prediction of synthetic lethal relationships.

[0058] Step S105: Repeat steps S102 to S104 until the L-th layer feature vector of each node is obtained, and based on the L-th layer feature vector of each node, a second feature vector set of each node is obtained.

[0059] Specifically, this feature update process is repeated at each layer until it reaches the Lth layer. In each layer, the model dynamically adjusts the information transfer weights between nodes through the attention mechanism, so that the model can learn more valuable neighbor information for synthetic lethality prediction. At the same time, the increase in the number of layers can help the model expand its perception range, no longer limited to direct neighbors, but can capture indirect interaction information at a longer distance.

[0060] By repeatedly executing the feature update process, the model can aggregate more neighbor information layer by layer, making the feature representation of each node more and more comprehensive. This layer-by-layer feature update method ensures that the model not only captures local direct relationships, but also captures complex network interactions between genes through multi-level aggregation. Ultimately, at the Lth layer, the feature vectors of all nodes will integrate their own information and the interactive information of their upstream and downstream neighbors to form a more representative feature expression.

[0061] Step S106: Based on the entire set of second eigenvectors, a second joint subgraph is obtained.

[0062] Specifically, during the aggregation process, the feature vector of each node is combined with the information of its neighboring nodes to update its own representation. This layer-by-layer progressive approach can gradually aggregate information from the entire joint subgraph, so that each node not only carries its own attributes, but also contains comprehensive information about the nodes associated with it. Through this integration of information, each node in the joint subgraph can reflect a more global perspective, helping the model to better understand the complex relationships between gene pairs. Ultimately, the construction of the second joint subgraph generates a richer and more biologically meaningful graph structure by fully integrating the information in the updated second feature vector set.

[0063] Step 3: Based on the second joint subgraph, obtain the node-level feature vector of the target gene.

[0064] In the embodiment of step 3, the goal is to extract the node-level feature vectors of the target genes u and v based on the node features in the second joint subgraph. This step is crucial because through the node-level features extracted from the joint subgraph, the model can accurately capture the feature representation of the target gene pair and its related nodes, thereby providing key input for subsequent synthetic lethality prediction.

[0065] In a possible implementation, step 3 specifically includes the following steps: Determine a second feature vector set for each node in the second joint subgraph, the second feature vector set includes feature vectors of L layers, where L is a natural number used to represent the number of layers of the network; Update the feature vector of layer L based on the attention mechanism to obtain the updated feature vector of layer L; The updated feature vector of the 0th layer is connected to the updated feature vector of the Lth layer to obtain the node-level feature vector.

[0066] Specifically, according to the previous steps, each node in the second joint subgraph has undergone multi-layer feature updates to form a multi-level feature vector set. In order to extract the node-level feature vectors of the target genes u and v, it is necessary to extract and combine the feature representations of these nodes from each layer.

[0067] Specifically, the node-level feature vectors of target genes u and v are expressed as: In this case, i∈[1,L], SH u is the node-level feature vector of node u, SH v is the node-level feature vector of node v, is the updated feature vector of the i-th layer of node u, is the updated feature vector of node v at layer i, is the feature vector of the i-th layer of node u, is the feature vector of the i-th layer of node v, is the weight matrix of the i-th layer, d K is a matrix Dimension.

[0068] This multi-layer feature vector extraction method has significant advantages. First, by connecting multiple layers of information, the model can not only capture the direct neighbor information of the target gene, but also its upstream and downstream relationships at a deeper level. This feature extraction method enables the model to gradually learn the complex interactive information of target gene pairs in different layers, which helps to identify possible synthetic lethal mechanisms.

[0069] Step 4: Based on the node-level feature vector and the second feature vector set, the final feature vector of the target gene pair is obtained.

[0070] Specifically, in the prediction of synthetic lethal gene pairs, the target gene pairs not only rely on their own characteristics, but also involve a complex multi-level relationship network. This relationship includes both the omics information such as the expression, mutation, and copy number changes of the gene itself, as well as their structural information in the entire gene network. By aggregating these multidimensional features, a more comprehensive feature vector can be generated to capture the potential regulatory mechanisms and synergies between gene pairs, and ultimately improve the prediction accuracy of the model. Therefore, in the embodiment of step 4, the goal is to generate the final feature vector of the target gene pair u and v based on the features in the node-level feature vector and the second joint subgraph, by combining the node-level features with the omics information, and performing feature aggregation at different levels, the potential synthetic lethality relationship between the target gene pairs can be effectively captured, thereby providing high-quality input for subsequent synthetic lethality probability prediction.

[0071] In one possible implementation, refer to Figure 2 , Figure 2 This is the second flow chart of the synthetic lethal gene pair prediction method provided by the present invention, step 4 specifically includes steps 41-45: Step 41: Aggregate all L-layer feature vectors of each node in the second joint subgraph to generate a third feature vector set.

[0072] Specifically, aggregating all L-layer feature vectors of each node in the second joint subgraph can be expressed by the following formula: , where E is the set of nodes in the second joint subgraph, is the feature vector of the Lth layer of each node. By averaging the features of all nodes, a feature set containing global information can be generated This process integrates the features of different nodes, ensuring that the model does not rely solely on local information, but can understand the potential relationship between gene pairs and their surrounding nodes from a global perspective. Then generate the third feature vector set .

[0073] Step 42: Encode the omics features of each node in the target gene pair based on a multi-layer perceptron, and generate omics feature vectors of each node in the target gene pair.

[0074] Specifically, synthetic lethality prediction not only depends on the gene network structure, but also highly depends on omics data, such as gene expression, mutation, and copy number changes. These omics data directly reflect the biological behavior and functional status of genes, and can provide the model with important information about the functional interaction of genes. However, omics data are usually high-dimensional, and direct use of these data may lead to redundancy and overfitting. Therefore, using a multi-layer perceptron (MLP) to encode these omics features can effectively reduce the data dimension while retaining key feature patterns, so that the model can extract the most predictive information from it.

[0075] In specific operations, the omics data of target genes u and v include gene expression characteristics, Mutation characteristics and copy number variation characteristics Similarly, gene v also has corresponding omics features. In order to encode these high-dimensional omics features, a multi-layer perceptron (MLP) is used to reduce the dimension and fuse these features.

[0076] The encoding process of omics features is as follows: Among them, O u and O v Represent the omics feature vectors of genes u and v respectively, It means connecting different types of omics features together to form a comprehensive input vector. Then, feature encoding is performed through MLP, which consists of multiple fully connected layers and can learn nonlinear relationships in omics features. This process not only reduces the dimensionality of omics features, but also retains key information patterns, making the encoded features more suitable for subsequent prediction tasks.

[0077] The role of MLP is to perform nonlinear transformations on the input high-dimensional omics data through a multi-layer neural network to generate low-dimensional, expressive feature vectors. The parameters of MLP are automatically learned through training data, so it can extract the most valuable features according to the needs of the synthetic lethality prediction task.

[0078] Step 43: Combine each omics feature vector through Kronecker product to obtain the omics feature representation of the target gene pair.

[0079] Specifically, synthetic lethal relationships often involve the synergy of two genes, and these synergies may be reflected in the interaction of omics features (such as gene expression, mutation, copy number changes, etc.). Analyzing the omics features of each gene alone cannot effectively capture the functional associations between them. Therefore, the interactive combination of the omics feature vectors of genes u and v through the Kronecker product can generate more expressive features that reflect the potential interactive relationship between gene pairs. The Kronecker product can expand the two vectors into a higher-dimensional interactive representation, thereby capturing complex gene-to-gene relationships.

[0080] The specific operation is to convert the omics feature vectors O of genes u and v encoded by the multi-layer perceptron (MLP) u and O v Perform Kronecker product operation to generate the omics feature representation of gene pairs. uv .

[0081] The formula is as follows: Among them, O u is the omics feature vector of gene u, O v is the omics feature vector of gene v, represents the Kronecker product operation. The Kronecker product is used to generate a high-dimensional feature vector that combines O u and O v The product between each pair of elements in is obtained by multiplying the two genes together, thereby capturing the complex interaction between the omics information of the two genes.

[0082] This combination can effectively enhance the model's perception of the functional interaction of gene pairs. Through the Kronecker product, the model can capture the coordinated change pattern from the omics features of two genes, which is particularly important when predicting synthetic lethal relationships. For example, if two genes show related expression patterns or mutations under specific conditions, the Kronecker product can reflect this information in the final omics feature representation, improving the model's predictive ability.

[0083] Step 44: Concatenate the node-level feature vector, the third feature vector set, and the omics feature representation to obtain a multi-level feature vector.

[0084] Specifically, the prediction of synthetic lethal gene pairs requires comprehensive consideration of the gene network structure, the characteristics of the nodes themselves, and the interactions between omics data. Therefore, information from a single dimension is not enough to fully reflect the potential functional relationship of gene pairs. By concatenating node-level features, global features of joint subgraphs, and omics features, the model can comprehensively consider information from these different dimensions at multiple levels, thereby enhancing the model's expressive power and understanding of complex biological systems. This feature concatenation process can unify different types of features, ensuring that the model can capture both local and global feature information of gene pairs at the same time.

[0085] These multi-dimensional features are combined and connected in series to obtain a multi-level feature vector, which is expressed as follows: ;in, is a multi-level feature vector.

[0086] Step 45: Based on the multi-level feature vectors, a final feature vector is obtained.

[0087] Specifically, the generated multi-level feature vector already contains all the comprehensive information of the target gene pair. Next, further feature fusion and optimization are needed to extract the final feature representation. In order to make the feature information more prominent, the self-attention mechanism will be used to optimize the multi-level feature vector and extract the final feature representation.

[0088] The optimization process can be achieved through the following formula: ; in, is the final eigenvector, is a multi-level feature vector, , and is a weight matrix. The softmax function is used to calculate the attention weights. By calculating the correlation of feature vectors, it is determined which features have a higher contribution to the prediction of synthetic lethality. is a matrix The dimension of Used to perform normalization operations to ensure that the scale of features is consistent.

[0089] In this process, the self-attention mechanism adjusts the weights to ensure that the model can dynamically assign the importance of each feature according to the different input data, thereby generating a more accurate final feature vector. This step ensures that the model does not rely solely on a single type of feature (such as omics features or node features), but comprehensively considers multiple levels of information to improve the model's prediction accuracy.

[0090] Step 5: Based on the final feature vector, obtain the synthetic lethality probability of the target gene pair.

[0091] In a possible implementation, step 5 specifically includes the following steps: The final feature vector is input into the synthetic lethality probability prediction model, and the final feature vector is linearly transformed using the weight matrix of the decoder in the synthetic lethality probability prediction model to calculate the synthetic lethality; wherein the weight matrix is ​​used to adjust the mapping relationship between the final feature vector and the synthetic lethality probability.

[0092] Specifically, in the previous steps, the final feature vectors of gene pairs u and v have been generated through multi-level feature fusion and self-attention mechanism optimization. , which integrates the network relationship, omics features, and global interaction information of the gene pair. However, the feature vector itself cannot be directly used for synthetic lethality prediction. In order to convert these features into predicted synthetic lethality probabilities, these features need to be further processed by a decoder. The role of the decoder is to map features to probability values ​​through a trainable model to determine whether a gene pair has synthetic lethality.

[0093] The specific calculation process is carried out by the decoder, which converts the final feature vector Input and pass a linear transformation to generate the synthetic lethality probability of gene pairs , the value is between 0 and 1. The formula is as follows: ; in, is the synthetic lethal probability, is the weight matrix of the decoder in the synthetic lethality probability prediction model, is the final feature vector.

[0094] This process maps the high-dimensional final feature vector into a single probability value through linear transformation, indicating the possibility of whether the gene pair u and v is a synthetic lethal gene pair. The probability value can be used to set a threshold according to actual needs to determine the classification results of the gene pair. In this way, the model can convert complex feature vectors into easy-to-understand synthetic lethality probabilities, thereby providing an important reference for practical research and applications.

[0095] In a possible implementation, the synthetic lethality probability prediction model is trained by the following steps: The total loss function of the synthetic lethality probability prediction model is set. The total loss function is expressed by the following formula: ; ; ; Among them, K is the total loss; is the cross entropy loss; is the regularization loss; is the regularization coefficient; is the total number of target gene pair samples in the training set; is the true label of the target gene pair sample (u, v), where the target gene pair sample (u, v) is a synthetic lethal gene pair, then ,otherwise, ; The synthetic lethality probability of the target gene for sample (u, v) predicted by the synthetic lethality probability prediction model; is the set of all weight parameters in the synthetic lethality probability prediction model; The initial model is trained with the goal of minimizing the total loss, and the weight matrix of the decoder is iteratively updated to obtain a synthetic lethal probability prediction model.

[0096] Specifically, the core of training the synthetic lethality prediction model is to minimize the loss function so that the model can accurately predict the synthetic lethality probability of gene pairs. The loss function can measure the difference between the model's prediction results and the true label, so its definition is crucial to the performance of the model. In this step, the total loss function including cross entropy loss and regularization loss is used. Cross entropy loss can effectively measure the error of the binary classification problem (synthetic lethality prediction), while regularization loss can prevent the model from overfitting and ensure the generalization ability of the model on new data.

[0097] The total loss function is expressed by the following formula: ; ; ; Among them, K is the total loss; is the cross entropy loss; is the regularization loss; is the regularization coefficient; is the total number of target gene pair samples in the training set; is the true label of the target gene pair sample (u, v), where the target gene pair sample (u, v) is a synthetic lethal gene pair, then ,otherwise, ; The synthetic lethality probability of the target gene for sample (u, v) predicted by the synthetic lethality probability prediction model; is the set of all weight parameters in the synthetic lethality probability prediction model.

[0098] During the training process of the model, the weight matrix of the decoder is optimized by an optimization algorithm (such as gradient descent or Adam algorithm) with the goal of minimizing the total loss K. Perform iterative updates. At each iteration, based on the gradient of the loss function, the weight matrix Through multiple rounds of training iterations, the weight matrix in the decoder The training process ends when the loss function L reaches the minimum value or performs best on the validation set. It is considered optimal and model training is completed.

[0099] The synthetic lethal gene pair prediction model (LSAG, Layer-wise Self-Attention GraphNeural Network, refers to a model based on graph neural network (GNN) and hierarchical self-attention mechanism, mainly used for the prediction of synthetic lethal gene pairs) of the present invention has a training epoch limit of 30, a batch size of 512, and a gradient-based method such as the Adam optimizer is used to minimize the total loss K. The Adam optimizer has the characteristics of adaptive learning rate, can converge quickly on large-scale data, and is suitable for complex tasks such as synthetic lethality prediction. The learning rate is set to 5×10 -3 , and the model will be further optimized until the training loss coefficient drops to 10 -4 the following.

[0100] The optimization goal of the model during training is to gradually adjust the model weights so that the total loss K gradually decreases. By minimizing the cross entropy loss, the model can improve its prediction accuracy; and by constraining the regularization loss, the model can avoid overfitting on complex data sets.

[0101] To ensure the fairness of the verification results, the verification experiment of the present invention uses the revised knowledge graph and SL data subset as the data set, divides it into training set, test set and validation set in a ratio of 7:2:1, and uses 5-fold cross validation with the following three settings to evaluate various SL prediction models: CV1: The dataset of gene pairs is divided based on genes, allowing both genes of the gene pairs in the test set to appear in the training set; CV2: The gene pair dataset is divided by gene, ensuring that only one gene in the gene pair in the test set appears in the training set; CV3: The gene pair dataset is divided by gene, ensuring that both genes in the gene pair in the test set do not appear in the training set; On this basis, the area under the ROC curve (Area under ROC curve, AUC) and the area under the precision recall curve (Area under precision–recall curve, AUPR) are used as evaluation indicators of model performance.

[0102] Reference Figure 3 , Figure 3 Schematic diagram comparing the prediction effects of the LSAG provided by the present invention and the existing model. Figure 3 As shown, traditional machine learning methods (i.e., XGBoost, KNN, Node2Vec, GAT, GCN, and GraphSAGE) and algorithms with excellent performance published in recent years (i.e., SL2MF, DDGCN, GCATST, KG4SL, and PiLSL) are used as comparative examples to obtain comparative prediction effects. The LSAG of the present invention showed superior performance in all scenarios. In the case of CV1, many models achieved satisfactory results, mainly because both target genes were present in the training set. However, LSAG achieved excellent performance of AUC0.9579 and AUPR0.9613, exceeding the suboptimal model by 0.41% and 0.19%, respectively. Similarly, in the CV2 scenario, LSAG continued to show the highest performance, with an AUC of 0.8162 and an AUPR of 0.8232, exceeding the second-best model by 2.18% and 0.76%, respectively. In the CV3 scenario, LSAG achieved an AUC performance of 0.6962 and an AUPR performance of 0.6943. It is worth noting that the gap between LSAG and the second-best model widens further, reaching 3.03% and 2.34%, respectively.

[0103] The present invention conducts a sensitivity analysis on the number of hidden layers of the LSAG model (i.e., the number of layers L of the feature selection module), which is performed in the CV3 scenario. Figure 4 , Figure 4is a schematic diagram of the structure of the synthetic lethal gene pair prediction model based on graph neural network provided by the present invention, such as Figure 4 As shown, the AUC score increases with The increase first rises and then decreases. On the other hand, the AUPR score shows an upward trend when When , the AUPR score reaches a peak of 0.69434. Because AUPR surpasses AUC in terms of information and persuasiveness when used as a performance evaluation indicator for model prediction, and although The AUC score is higher than The AUC score is The standard deviation of AUC (0.015591) significantly exceeds The AUC standard deviation is (0.00515), so the number of layers of the feature selection module of LSAG in the present invention is Select the value 64.

[0104] The present invention analyzes the role of the dual self-attention module in the LSAG model, removes the modules related to the execution of step S3 and step S4 in the above embodiment in the LSAG model, and obtains two variants, LSAG-WLA and LSAG-WFA. Figure 5 , Figure 5 Schematic diagram of the comparison of LSAG model variants provided by the present invention. Figure 5 As shown in Figure 2, LSAG ranks at least in the top 2 in all three scenarios. Especially in the CV2 scenario, both AUC and AUPR values ​​of LSAG are significantly higher compared with its variants.

[0105] Reference Figure 6 , Figure 6 is a schematic diagram of the visualization results of the joint subgraph and attention coefficient provided by the present invention, Figure 6 The joint subgraph and attention coefficient visualization results of the synthetic lethality gene pair TP53 and USP1 constructed by the LSAG method of the present invention are shown. The attention coefficients of neighboring genes of TP53 and USP1 are calculated, and a weighted average is generated. Then, key genes are identified, and pathway enrichment analysis is performed using the Kyoto Encyclopedia of Genes and Genomes (KEGG) and Gene Ontology (GO) to clarify the synthetic lethality mechanism of these key genes. Figure 7 , Figure 7 Schematic diagram of the enrichment analysis results provided by the present invention. Figure 7As shown, in the joint subgraph, genes showing high attention coefficients showed enrichment characteristics in GOBP and KEGG pathways related to DNA damage repair, including "regulation of response to DNA damage stimuli" and "signal transduction of DNA damage", an observation consistent with the current understanding of the mechanism of synthetic lethality between these two genes. In addition, genes showing significant attention coefficients in the joint subgraph region, such as SFTPA2 (pulmonary surfactant protein), RPA2 (single-stranded DNA binding protein), and GPER1 (damage repair factor of type I interferon signaling during pregnancy), are also related to molecular and cellular growth and repair. In particular, RPA2, as a key protein that binds single-stranded DNA, plays a role in multiple DNA repair pathways. Studies have shown that it can protect inherited DNA damage and promote DNA synthesis after mitosis. Through the analysis of the LSAG model, a deeper understanding of the mechanism of synthetic lethality can be achieved, revealing the importance of gene interactions and functions in synthetic lethality. These findings are of great significance for future studies of the mechanism of synthetic lethality, the identification of novel synthetic lethal gene pairs, and the guidance of related strategies based on synthetic lethality.

[0106] Reference Figure 8 , Figure 8 It is a structural schematic diagram of the synthetic lethal gene pair prediction system provided by the present invention, the system comprising: a joint subgraph extraction module, an information aggregation module, a first processing module, a second processing module and a synthetic lethal probability prediction module; A joint subgraph extraction module is used to extract a first joint subgraph related to the target gene pair from the knowledge graph; the first joint subgraph includes a first feature vector set of all nodes related to the target gene pair; An information aggregation module is used to aggregate the information in the first joint subgraph to obtain a second joint subgraph; the second joint subgraph includes a second feature vector set, and the second feature vector set is the first feature vector set updated during the aggregation process; A first processing module, used for obtaining a node-level feature vector of a target gene based on the second joint subgraph; A second processing module is used to obtain a final feature vector of the target gene pair based on the node-level feature vector and the second feature vector set; The synthetic lethality probability prediction module is used to obtain the synthetic lethality probability of the target gene pair based on the final feature vector.

[0107] In a possible implementation, the joint subgraph extraction module is further used to construct a first gene subgraph of the first target gene and a second gene subgraph of the second target gene respectively based on the knowledge graph; The joint subgraph extraction module is further used to merge the first gene subgraph and the second gene subgraph to obtain a first joint subgraph.

[0108] In a possible implementation, the information aggregation module is further used to determine the 0th layer feature vector of each node in the first joint subgraph; The information aggregation module is also used to calculate the attention weight based on the i-1th layer feature vector of each node and the i-1th layer feature vector of each node's neighboring node, where i∈[1,L], L is a natural number used to represent the number of layers of the network; The information aggregation module is also used to weight the i-1th layer feature vectors of neighbor nodes based on the attention weights to obtain aggregated neighbor feature vectors; The information aggregation module is further used to combine the i-1th layer feature vector of each node with the aggregated neighbor feature vectors to obtain the i-th layer feature vector of each node; The information aggregation module is further used to repeatedly execute the above steps until the L-th layer feature vector of each node is obtained, and based on the L-th layer feature vector of each node, a second feature vector set of each node is obtained; The information aggregation module is further used to obtain a second joint subgraph based on the entire second feature vector set.

[0109] In a possible implementation, the information aggregation module is further used to obtain an initial feature vector of each node in the first joint subgraph using a TransE embedding method; The information aggregation module is also used to obtain the position feature vector of each node based on the position information of each node and the target gene pair; The information aggregation module is also used to connect each position feature vector with each initial feature vector to generate the 0th layer feature vector of each node.

[0110] In a possible implementation, the first processing module is further used to determine a second feature vector set for each node in the second joint subgraph, where the second feature vector set includes feature vectors of L layers, where L is a natural number used to represent the number of layers of the network; The first processing module is further used to update the feature vector of the L layer based on the attention mechanism to obtain an updated feature vector of the L layer; The first processing module is also used to perform vector connection from the updated feature vector of the 0th layer to the updated feature vector of the Lth layer to obtain a node-level feature vector.

[0111] In a possible implementation, the second processing module is further used to aggregate all L-layer feature vectors of each node in the second joint subgraph to generate a third feature vector set; The second processing module is further used to encode the omics features of each node in the target gene pair based on the multi-layer perceptron, and generate the omics feature vectors of each node in the target gene pair respectively; The second processing module is also used to combine each omic feature vector through Kronecker product to obtain the omic feature representation of the target gene pair; The second processing module is also used to connect the node-level feature vector, the third feature vector set and the omics feature representation in series to obtain a multi-level feature vector; The second processing module is also used to obtain a final feature vector based on the multi-level feature vectors.

[0112] In one possible implementation, the synthetic lethality probability prediction module is further used to input the final feature vector into the synthetic lethality probability prediction model, use the weight matrix of the decoder in the synthetic lethality probability prediction model to perform a linear transformation on the final feature vector, and calculate the synthetic lethality; wherein the weight matrix is ​​used to adjust the mapping relationship between the final feature vector and the synthetic lethality probability.

[0113] In a possible implementation, the system further includes: a model building module; a model building module, configured to set a total loss function of the synthetic lethality probability prediction model, wherein the total loss function is represented by the following formula: ; ; ; Among them, K is the total loss; is the cross entropy loss; is the regularization loss; is the regularization coefficient; is the total number of target gene pair samples in the training set; is the true label of the target gene pair sample (u, v), where the target gene pair sample (u, v) is a synthetic lethal gene pair, then ,otherwise, ; The synthetic lethality probability of the target gene for sample (u, v) predicted by the synthetic lethality probability prediction model; is the set of all weight parameters in the synthetic lethality probability prediction model; The model building module is also used to train the initial model with minimizing the total loss as the training objective, iteratively update the weight matrix of the decoder, and train to obtain a synthetic lethality probability prediction model.

[0114] It should be noted that the synthetic lethal gene pair prediction system provided by the present invention can execute the synthetic lethal gene pair prediction method of any of the above embodiments during specific operation, which will not be described in detail in this embodiment.

[0115] Fig. 9 is a schematic diagram of the structure of the electronic device provided by the present invention, such as Fig. 9As shown, the electronic device may include: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus. The processor may call the logic instructions in the memory to execute the synthetic lethal gene pair prediction method, which includes: extracting a first joint subgraph related to the target gene pair from the knowledge graph; the first joint subgraph includes a first feature vector set of all nodes related to the target gene pair; aggregating the information in the first joint subgraph to obtain a second joint subgraph; the second joint subgraph includes a second feature vector set, which is a first feature vector set updated during the aggregation process; based on the second joint subgraph, obtaining the node-level feature vector of the target gene; based on the node-level feature vector and the second feature vector set, obtaining the final feature vector of the target gene pair; based on the final feature vector, obtaining the synthetic lethal probability of the target gene pair.

[0116] In addition, the logic instructions in the above-mentioned memory can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0117] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the synthetic lethal gene pair prediction method provided by the above-mentioned embodiments, and the method includes: extracting a first joint subgraph related to the target gene pair from a knowledge graph; the first joint subgraph includes a first feature vector set of all nodes related to the target gene pair; aggregating the information in the first joint subgraph to obtain a second joint subgraph; the second joint subgraph includes a second feature vector set, and the second feature vector set is a first feature vector set updated during the aggregation process; based on the second joint subgraph, obtaining a node-level feature vector of the target gene; based on the node-level feature vector and the second feature vector set, obtaining a final feature vector of the target gene pair; based on the final feature vector, obtaining the synthetic lethality probability of the target gene pair.

[0118] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the synthetic lethal gene pair prediction method provided in the above-mentioned embodiments, the method comprising: extracting a first joint subgraph related to the target gene pair from a knowledge graph; the first joint subgraph includes a first feature vector set of all nodes related to the target gene pair; aggregating the information in the first joint subgraph to obtain a second joint subgraph; the second joint subgraph includes a second feature vector set, which is a first feature vector set updated during the aggregation process; based on the second joint subgraph, obtaining a node-level feature vector of the target gene; based on the node-level feature vector and the second feature vector set, obtaining a final feature vector of the target gene pair; based on the final feature vector, obtaining the synthetic lethality probability of the target gene pair.

[0119] The system embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, i.e., they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Those of ordinary skill in the art may understand and implement it without creative effort.

[0120] Through the description of the above implementation modes, those skilled in the art can clearly understand that each implementation mode can be implemented by means of software plus a necessary general hardware platform, or of course by hardware. Based on such an understanding, the above technical solution can essentially or in other words be embodied in the form of a software product that contributes to the prior art. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiment.

[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical feature vectors therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for predicting synthetic lethal gene pairs, characterized in that: include: Extracting a first joint subgraph related to the target gene pair from the knowledge graph; the first joint subgraph includes a first feature vector set of all nodes related to the target gene pair; Aggregate the information in the first joint subgraph to obtain a second joint subgraph; the second joint subgraph includes a second feature vector set, and the second feature vector set is the first feature vector set updated during the aggregation process; Based on the second joint subgraph, obtaining a node-level feature vector of the target gene; Based on the node-level feature vector and the second feature vector set, obtaining a final feature vector of the target gene pair; Based on the final feature vector, the synthetic lethality probability of the target gene pair is obtained.

2. The method for predicting synthetic lethal gene pairs according to claim 1, characterized in that: The aggregating the information in the first joint subgraph to obtain the second joint subgraph specifically includes: Step S101: determining the 0th layer feature vector of each node in the first joint subgraph; Step S102: Calculate the attention weight based on the i-1th layer feature vector of each node and the i-1th layer feature vector of each node's neighboring node, where i∈[1,L], L is a natural number representing the number of layers of the network; Step S103: Based on the attention weight, weight the i-1th layer feature vectors of the neighbor nodes to obtain an aggregated neighbor feature vector; Step S104: combining the i-1th layer feature vector of each node with the aggregated neighbor feature vector to obtain the i-th layer feature vector of each node; Step S105: Repeat steps S102 to S104 until the L-th layer feature vector of each node is obtained, and based on the L-th layer feature vector of each node, a second feature vector set of each node is obtained; Step S106: obtaining the second joint subgraph based on all sets of the second feature vectors.

3. The method for predicting synthetic lethal gene pairs according to claim 2, characterized in that: The determining of the 0th layer feature vector of each node in the first joint subgraph specifically includes: Obtaining an initial feature vector for each node in the first joint subgraph using the TransE embedding method; Based on the position information of each node and the target gene pair, a position feature vector of each node is obtained; Each of the position feature vectors is connected to each of the initial feature vectors to generate a 0th layer feature vector for each node.

4. The method for predicting synthetic lethal gene pairs according to claim 1, characterized in that: The step of obtaining a node-level feature vector of the target gene based on the second joint subgraph specifically includes: Determine a second feature vector set for each node in the second joint subgraph, the second feature vector set includes feature vectors of L layers, where L is a natural number used to represent the number of layers of the network; Based on the attention mechanism, the feature vector of the L layer is updated to obtain an updated feature vector of the L layer; Perform vector connection between the updated feature vector of the 0th layer and the updated feature vector of the Lth layer to obtain the node-level feature vector.

5. The method for predicting synthetic lethal gene pairs according to claim 1, characterized in that: The second feature vector set includes feature vectors of L layers, where L is a natural number used to represent the number of layers of the network; the final feature vector of the target gene pair is obtained based on the node-level feature vector and the second feature vector set, specifically including: Aggregating all L-layer feature vectors of each node in the second joint subgraph to generate a third feature vector set; Encoding the omics features of each node in the target gene pair based on a multi-layer perceptron, and generating omics feature vectors of each node in the target gene pair respectively; Combining each of the omics feature vectors by Kronecker product to obtain the omics feature representation of the target gene pair; Connecting the node-level feature vector, the third feature vector set, and the omics feature representation in series to obtain a multi-level feature vector; Based on the multi-level feature vectors, the final feature vector is obtained.

6. The method for predicting synthetic lethal gene pairs according to claim 1, characterized in that: The obtaining the synthetic lethality probability of the target gene pair based on the final feature vector specifically includes: The final feature vector is input into a synthetic lethal probability prediction model, and the final feature vector is linearly transformed using a weight matrix of a decoder in the synthetic lethal probability prediction model to calculate the synthetic lethal probability; wherein the weight matrix is ​​used to adjust the mapping relationship between the final feature vector and the synthetic lethal probability.

7. The method for predicting synthetic lethal gene pairs according to claim 6, characterized in that: The synthetic lethality probability prediction model is trained by the following method: The total loss function of the synthetic lethality probability prediction model is set, and the total loss function is expressed by the following formula: ; ; ; Among them, K is the total loss; is the cross entropy loss; is the regularization loss; is the regularization coefficient; is the total number of target gene pair samples in the training set; is the true label of the target gene pair sample (u, v), where the target gene pair sample (u, v) is a synthetic lethal gene pair, then ,otherwise, ; The synthetic lethality probability of the target gene for the sample (u, v) predicted by the synthetic lethality probability prediction model; is the set of all weight parameters in the synthetic lethality probability prediction model; The initial model is trained with minimizing the total loss as the training goal, the weight matrix of the decoder is iteratively updated, and the synthetic lethality probability prediction model is obtained through training.

8. The method for predicting synthetic lethal gene pairs according to claim 1, characterized in that: The target gene pair includes a first target gene and a second target gene; and extracting a first joint subgraph related to the target gene pair from the knowledge graph specifically includes: Based on the knowledge graph, constructing a first gene subgraph of the first target gene and a second gene subgraph of the second target gene respectively; The first gene subgraph and the second gene subgraph are merged to obtain the first joint subgraph.

9. A synthetic lethal gene pair prediction system, characterized in that: include: A joint subgraph extraction module, an information aggregation module, a first processing module, a second processing module and a synthetic lethality probability prediction module; The joint subgraph extraction module is used to extract a first joint subgraph related to the target gene pair from the knowledge graph; the first joint subgraph includes a first feature vector set of all nodes related to the target gene pair; The information aggregation module is used to aggregate the information in the first joint subgraph to obtain a second joint subgraph; the second joint subgraph includes a second feature vector set, and the second feature vector set is the first feature vector set updated during the aggregation process; The first processing module is used to obtain a node-level feature vector of the target gene based on the second joint subgraph; The second processing module is used to obtain a final feature vector of the target gene pair based on the node-level feature vector and the second feature vector set; The synthetic lethality probability prediction module is used to obtain the synthetic lethality probability of the target gene pair based on the final feature vector.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for predicting synthetic lethal gene pairs according to any one of claims 1 to 8 is implemented.

11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for predicting synthetic lethal gene pairs according to any one of claims 1 to 8 is implemented.

12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for predicting synthetic lethal gene pairs according to any one of claims 1 to 8 is implemented.