Link Prediction Method, Device, and Storage Medium Adapted to Non-Centrally Distributed Data

By introducing mean subtraction preprocessing technology into the link prediction method, the potential representation of nodes is centralized, which solves the calculation efficiency and accuracy problems in non-center distributed data processing in the prior art, and realizes a more stable and efficient link prediction model.

CN119988688BActive Publication Date: 2025-06-24BEI JING NORMAL UNIV HONG KONG BAPTIST UNIV UNITED INT COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510473722.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-06-24
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

Existing link prediction methods perform poorly when processing non-center distributed data, resulting in increased redundant calculations and model complexity, inaccuracy and computational efficiency.

Method used

By obtaining the original input graph and node feature information, the node feature matrix and the real adjacency matrix are generated, and the initial latent representation is encoded to obtain the initial latent representation, and the latent representation is subtracted to the final latent representation. Based on the final potential representation, the potential connection is predicted, the combined loss is calculated and the model is iteratively trained to obtain the link prediction result.

Benefits of technology

Through mean subtraction preprocessing technology, non-informative offsets in the input data are eliminated, the distribution of potential space is simplified, the model is more stable, the redundant calculation is reduced, and the model training efficiency and stability is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988688B_ABST
    Figure CN119988688B_ABST
Patent Text Reader

Abstract

This application relates to the technical field of network link prediction, and particularly to a link prediction method, device, and storage medium adapted to non-centrally distributed data. The link prediction method includes: generating a node feature matrix and a true adjacency matrix based on the original input graph and the feature information of each node in the original input graph; encoding the node feature matrix and the true adjacency matrix to obtain the initial latent representation of each node, subtracting the latent representation mean of the node from the initial latent representation of the node to obtain the final latent representation of each node; predicting the potential connections between entities based on the final latent representation of each node, calculating a combined loss based on the prediction results to iteratively train the model, and obtaining the link prediction result of the original input graph. The method of this application introduces a mean subtraction preprocessing technique, simplifies the distribution of the latent space, makes the model more stable, and reduces redundant calculations, thereby improving the training efficiency and stability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of network link prediction, and particularly to a link prediction method, device and storage medium adapted to non-centrally distributed data. Background Art

[0002] With the rapid development of network science and bioinformatics, the analysis of graph-structured data has become one of the core tasks in many fields. Among these tasks, link prediction is a key issue in graph analysis and is widely applied in fields such as social networks, recommendation systems, biological networks, knowledge graphs, etc. Link prediction aims to predict whether there is a potential connection between node pairs in a graph or predict unobserved edges.

[0003] Traditional link prediction methods mainly rely on similarity metrics of nodes or edges, usually calculating the similarity between nodes and inferring potential connections based on this. However, these methods often ignore the global structure of the graph and the complex high-dimensional feature information between nodes. Especially when dealing with non-centrally distributed data, it often leads to redundant calculations and an increase in model complexity by fitting the mean of the covariance matrix.

[0004] Therefore, there is an urgent need for a link prediction method applicable to non-centrally distributed data to further improve the accuracy and computational efficiency of link prediction through more effective node feature learning, latent representation generation, and decoding processes. Summary of the Invention

[0005] Embodiments of the present application aim to provide a link prediction method, device and storage medium adapted to non-centrally distributed data to solve the problem that the existing link prediction methods perform poorly when dealing with non-centrally distributed data.

[0006] To solve the above technical problems, the embodiments of the present application provide the following technical solutions:

[0007] According to a first aspect of the present application, there is provided a link prediction method adapted to non-centrally distributed data, the method comprising:

[0008] Obtain an original input graph and feature information of each node in the original input graph, the original input graph including a node set and an edge set;

[0009] Generate a node feature matrix and a true adjacency matrix of the original input graph based on the original input graph and the feature information of each node in the original input graph;

[0010] Encode the node feature matrix and the true adjacency matrix to obtain an initial latent representation of each node, and subtract the latent representation mean of the node from the initial latent representation of the node to obtain a final latent representation of each node;

[0011] Predict the potential connections between entities based on the final potential representations of each node to obtain the predicted adjacency matrix of the original input graph;

[0012] Calculate the combined loss based on the true adjacency matrix, the predicted adjacency matrix, and the final potential representations of each node, and iteratively train the link prediction model based on the combined loss to obtain the link prediction result of the original input graph.

[0013] Optionally, encoding the node feature matrix and the true adjacency matrix to obtain the initial potential representation of each node includes:

[0014] Encoding the node feature matrix and the true adjacency matrix based on a first graph convolutional network encoder to obtain the mean of the potential representations of each node;

[0015] Encoding the node feature matrix and the true adjacency matrix based on a second graph convolutional network encoder to obtain the logarithmic variance of the potential representations of each node;

[0016] Generate the initial potential representation of each node from a Gaussian distribution based on the mean of the potential representations and the logarithmic variance of the potential representations of each node.

[0017] Optionally, the predicting the potential connections between entities based on the final potential representations of each node to obtain the predicted adjacency matrix of the original input graph includes:

[0018] Calculate the relationship strength between each node based on the final potential representations of each node;

[0019] Input the relationship strength between each node into a bidirectional activation function, and predict the positive and negative correlation information between nodes through the bidirectional activation function to obtain the predicted adjacency matrix of the original input graph.

[0020] Optionally, the calculation formula for the predicted connection probability between node pairs ( in the predicted adjacency matrix is:

[0021]

[0022] where and are the final potential representations of node and node respectively, denotes and perform a dot product operation, , is the sigmoid function.

[0023] Optionally, calculating the combined loss based on the true adjacency matrix, the predicted adjacency matrix, and the final latent representation of each node includes:

[0024] Calculating a reconstruction loss based on the similarity between the true adjacency matrix and the predicted adjacency matrix;

[0025] Calculating a contrastive loss based on the differences in the final latent representations of different node pairs;

[0026] Calculating the combined loss based on the reconstruction loss and the contrastive loss.

[0027] Optionally, the formula for calculating the reconstruction loss is:

[0028]

[0029] where, represents the set of edges observed in the original input graph, represents the connection relationship between node pairs ([[]] in the true adjacency matrix ), represents the predicted connection probability between node pairs ([[]] in the predicted adjacency matrix );

[0030] The formula for calculating the contrastive loss is:

[0031]

[0032] where, is a weight coefficient, and are the confidence thresholds for positive and negative samples respectively, and are the final latent representations of node and node respectively, represents and performing a dot product operation, , is the sigmoid function.

[0033] Optionally, the original input graph is a cell-cell interaction network graph, and the characteristic information of each cell in the cell-cell interaction network graph is the original expression level of each gene measured in the cell. Generating the node feature matrix of the original input graph based on the original input graph and the characteristic information of each node in the original input graph includes:

[0034] Creating a gene expression matrix based on the original expression levels of each gene measured in each cell in the cell-cell interaction network graph X , where, , is the number of cells, is the number of genes measured in each cell;

[0035] Normalize the gene expression matrix X to obtain the normalized expression levels of each gene measured in each cell;

[0036] Calculate the variance of the original or normalized expression levels of each gene across all cells, and select the top genes with the largest variances as the selection features;

[0037] Based on the normalized expression levels of the cell selection features, construct a reduced-dimensional gene expression matrix as the node feature matrix of the original input graph, where .

[0038] Optionally, generating the true adjacency matrix of the original input graph based on the original input graph and the feature information of each node in the original input graph includes:

[0039] Generate an initial adjacency matrix of the original input graph based on the original input graph;

[0040] Calculate the interaction potential between each pair of cells based on the similarity of the gene expression profiles;

[0041] Use the interaction potential as the weight of the initial adjacency matrix to obtain the true adjacency matrix of the original input graph.

[0042] According to a second aspect of the present application, there is provided an electronic device, including at least one processor and a memory communicatively connected to the at least one processor, the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to execute the method described above.

[0043] According to a third aspect of the present application, there is provided a computer storage medium storing instructions or programs that, when executed by at least one processor, cause the at least one processor to execute the method described above.

[0044] The beneficial effects of the embodiments of this application are as follows: Different from the prior art, in the embodiments of this application, a link prediction method adapted to non-centrally distributed data is provided. First, a node feature matrix and a true adjacency matrix are generated based on the original input graph and the feature information of each node in the original input graph; then, the node feature matrix and the true adjacency matrix are encoded to obtain the initial latent representation of each node, and the initial latent representation of the node is subtracted by the mean of the latent representations of the nodes to obtain the final latent representation of each node; finally, based on the final latent representation of each node, the potential connections between entities are predicted to obtain the predicted adjacency matrix of the original input graph, and a combined loss is calculated based on the true adjacency matrix, the predicted adjacency matrix, and the final latent representation of each node. The link prediction model is iteratively trained based on the combined loss to obtain the link prediction result of the original input graph. The method of this application introduces a mean subtraction preprocessing technique to perform mean centering on the latent representation of the nodes, eliminating the non-informative offset in the input data, simplifying the distribution of the latent space, making the model more stable, and reducing redundant calculations, thereby improving the training efficiency and stability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, unless otherwise stated, and the drawings in the drawings do not constitute a proportional limitation.

[0046] Figure 1 is a schematic flowchart of a link prediction method adapted to non-centrally distributed data provided by an embodiment of this application;

[0047] Figure 2 is a schematic structural diagram of the original input graph provided by an embodiment of this application;

[0048] Figure 3 is an algorithm framework diagram of a link prediction method adapted to non-centrally distributed data provided by an embodiment of this application;

[0049] Figure 4 is a schematic hardware structure diagram of an electronic device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the drawings in the embodiments of this application. Apparently, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts shall fall within the protection scope of this application.

[0051] In addition, the technical features involved in the various embodiments of the present application described below can be combined with each other as long as they do not conflict with each other.

[0052] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0053] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a link prediction method adapted to non-centrally distributed data provided by an embodiment of the present application. The method includes:

[0054] Step S101, obtaining the original input graph and the feature information of each node in the original input graph.

[0055] Specifically, the original input graph is graph-structured data, including a node set and an edge set. Each node in the node set represents an entity, and each edge in the edge set represents the existence of a certain relationship or interaction between the corresponding two entities. The feature information of a node refers to the descriptive data or attributes of each node, used to characterize the characteristics, states, or behaviors of the node.

[0056] As Figure 2 shown, it is an example graph of the original input graph. The original input graph is an undirected graph, including 12 nodes. Node 21 is one of the nodes, and there are a total of 4 edges between Node 21 and other nodes. In other embodiments, the original input graph can also be a directed graph, which is not limited in the present application. The original input graph can be a social network graph representing user relationships in a social network, a paper citation structure graph representing paper citation relationships, a knowledge graph representing knowledge point relationships, a cell-cell interaction network graph representing the interaction between cells, etc. Among them, in the social network graph, each node represents a user, and each edge represents the relationship between users, such as friend relationship, follow relationship, or message exchange. The feature information of a user can include the user's age, activity level, number of posts, number of comments, etc. In the cell-cell interaction network graph, each node represents a cell, and each edge represents the physical or functional interaction between two cells, such as ligand-receptor interaction, cell adhesion interaction, cell signal transduction, etc. The feature information of a cell can include gene expression profiles (transcription levels of each gene in the cell), cell surface molecules, receptor expression, cell type, cell state, metabolic state, etc.

[0057] Step S102, generating a node feature matrix and a true adjacency matrix of the original input graph based on the original input graph and the feature information of each node in the original input graph.

[0058] In some embodiments, before generating the node feature matrix of the input graph based on the feature information of each node in the original input graph, the feature information of each node in the original input graph is preprocessed. The preprocessing includes normalization processing and feature screening. By performing normalization processing on the feature information of each node in the original input graph, the scale difference between features can be eliminated, and the bias of the model towards certain features can be avoided. By performing feature screening on the feature information of each node in the original input graph, the most informative and target variable-representative features can be screened out from the original feature set, which can improve the performance of the model, reduce the computational cost, and enhance the generalization ability of the model.

[0059] Taking the cell-cell interaction network graph as an example below, the generation process of the node feature matrix of the original input graph will be described. Assume that the feature information of each cell is the original expression levels of each gene measured in the cell.

[0060] Step 1, create a gene expression matrix based on the original expression levels of each gene measured in each cell in the cell-cell interaction network graph X 。

[0061] Among them, , is the number of cells, is the number of genes measured in each cell.

[0062] Step 2, perform normalization processing on the gene expression matrix X based on the following formula:

[0063]

[0064] Among them, is the original expression level of gene in cell , is the normalized expression level of gene in cell .

[0065] Step 3, calculate the variance of the original expression level or normalized expression level of each gene in all cells as the variation characteristics of the cell, and screen out the top genes with the largest variation characteristics as the selected features:

[0066]

[0067] Among them, is the original expression level or normalized expression level of gene in cell .

[0068] Step 4, construct a reduced-dimensional gene expression matrix based on the normalized expression levels of the selected features of each cell As the node feature matrix of the original input graph, where .

[0069] It should be noted that the order of normalization processing and feature selection is not limited. Generally, there are two common methods: 1) First, perform feature selection, and then normalize the selected features. The advantage of this method is to reduce the amount of calculation and only normalize important features. The disadvantage is that some feature selection methods may be affected by the feature scale, resulting in inaccurate results. 2) First, normalize all features, and then perform feature selection. The advantage of this method is to ensure that feature selection is not affected by the scale, and the disadvantage is that the amount of calculation is large.

[0070] The true adjacency matrix of the original input graph can be constructed based on the set of points and the set of edges in the original input graph. For example, if the original input graph includes n nodes, then the true adjacency matrix is specifically an n×n dimensional matrix. Exemplarily, let A represent the true adjacency matrix. Suppose node and node are any two nodes in the original input graph respectively. is the data item in the -th row and the -th column of the true adjacency matrix A. Then = 1 indicates that there is an edge from node to node in the original input graph, and = 0 indicates that there is no edge from node to node in the original input graph.

[0071] In one embodiment, the strength of the mutual relationship between nodes can be incorporated into the true adjacency matrix as a weight. Taking the cell-cell interaction network graph as an example, the steps to generate the true adjacency matrix of the original input graph based on the original input graph and the feature information of each node in the original input graph may include: First, generate the initial adjacency matrix of the original input graph based on the set of points and the set of edges of the original input graph. Among them, the value of the matrix element in the -th row and the -th column of the initial adjacency matrix is 0 or 1. Second, calculate the interaction potential between each pair of cells based on the similarity of the expression profiles. Finally, use the interaction potential as the weight of the initial adjacency matrix to obtain the true adjacency matrix of the original input graph.

[0072] In one embodiment, the Pearson correlation coefficient is used to measure the interaction potential between cell pairs, and its calculation formula is:

[0073]

[0074] Where and is a cell point and the cell point in the gene expression level (which can be the original expression level or the normalized expression level), and is a cell point and the cell point is the average expression level of all genes in the cell point.

[0075] The initial adjacency matrix is weighted and calculated using the Pearson correlation coefficient to obtain the true adjacency matrix of the original input graph. Specifically, the value of the matrix element in the row and the column of the true adjacency matrix , where is the value of the matrix element in the row and the column of the initial adjacency matrix.

[0076] Step S103: Encode the node feature matrix and the true adjacency matrix to obtain the initial latent representation of each node, and subtract the mean of the latent representations of the nodes from the initial latent representation of the nodes to obtain the final latent representation of each node.

[0077] Specifically, a link prediction model is pre-constructed, and the link prediction model includes an encoder and a decoder. Among them, the encoder is used to encode and sample the node matrix and the true adjacency matrix to obtain the final latent representation of the nodes; the decoder is used to make predictions based on the final latent representation of the nodes to obtain the predicted adjacency matrix of the original input graph.

[0078] In one embodiment, encoding the node feature matrix and the true adjacency matrix to obtain the initial latent representation of each node specifically includes the following steps:

[0079] First, encode the node feature matrix and the true adjacency matrix based on the graph convolutional network encoder to obtain the mean of the latent representations of the nodes and the log variance of the latent representations.

[0080] Specifically, encode the node feature matrix and the true adjacency matrix based on the first graph convolutional network encoder to obtain the mean of the latent representations of each node; encode the node feature matrix and the true adjacency matrix based on the second graph convolutional network encoder to obtain the log variance of the latent representations of each node. This way can help the model learn the structure of the latent space more flexibly. Each encoder can learn different feature representations, so that the mean and variance can be generated separately through different networks, which can help better represent the distribution of the data.

[0081] Second, generate the initial latent representation of each node from the Gaussian distribution based on the mean of the latent representations of each node and the log variance of the latent representations.

[0082] Specifically, the calculation formula for the initial latent representation of each node is as follows:

[0083]

[0084] Among them, is the initial latent representation of node , is the mean of the latent representations of node , is the log variance of the latent representations of node , is a value sampled from the standard normal distribution.

[0085] In the above embodiments, the initial latent representations of the nodes are generated from the Gaussian distribution through sampling, providing a probabilistic representation for each node, which can make the model more robust and generalization-capable in the prediction task.

[0086] In other embodiments, the initial latent representations of the nodes can also be directly obtained by encoding the node feature matrix and the true adjacency matrix based on a graph convolutional network encoder.

[0087] After obtaining the initial latent of the nodes, subtracting the mean of the latent representations of the nodes from the initial latent representations of the nodes can obtain the final latent representations of the nodes. Taking the cell-cell interaction network graph as an example, the final latent representations of the nodes are the final latent representations of the cells. In the embodiments of the present application, the mean centering process of the latent representations of the nodes is performed through the mean subtraction technique, which can improve the efficiency and accuracy of the model in processing non-centered distribution data.

[0088] Step S104, predicting the latent connections between entities based on the final latent representations of the nodes to obtain the predicted adjacency matrix of the original input graph.

[0089] In one embodiment, the decoder of the link prediction model calculates the relationship strength between each node by computing the inner product of the final latent representations of the nodes, and then inputs the relationship strength between the nodes into a bidirectional activation function to predict the positive and negative correlation information between the nodes through the bidirectional activation function, obtaining the predicted adjacency matrix of the original input graph. Specifically, the calculation formula for the predicted connection probability between node pairs ( in the predicted adjacency matrix is as follows:

[0090]

[0091] Among them, and are the final latent representations of node and node respectively, represent and perform a dot product operation, , is the sigmoid function.

[0092] The traditional GAE model uses the Sigmoid function and can only process the positive correlation information between nodes, making it difficult to capture the negative correlation node relationships. The present invention proposes a bidirectional activation function that can process both positive and negative correlation relationships simultaneously. Through this bidirectional activation function, the model can not only capture the positive connections between nodes but also effectively express the inhibitory or adversarial relationships between nodes, thereby better modeling complex graph structures.

[0093] In step S105, a combined loss is calculated based on the real adjacency matrix, the predicted adjacency matrix, and the final latent representations of each node, and the link prediction model is iteratively trained based on the combined loss to obtain the link prediction result of the original input graph.

[0094] Existing models perform poorly in distinguishing nodes with similar features but different connection patterns, resulting in insufficient prediction accuracy in certain tasks. To solve this problem, the present invention introduces a contrast loss function to enhance the model's ability to distinguish subtle differences between similar nodes, thereby improving the model's prediction performance in complex graph structures.

[0095] In one embodiment, the combined loss includes a reconstruction loss and a contrast loss. Among them, the reconstruction loss is calculated based on the similarity between the real adjacency matrix and the predicted adjacency matrix, and the contrast loss is calculated based on the differences in the final latent representations of different node pairs. Specifically, the calculation formula for the reconstruction loss is:

[0096]

[0097] where, represents the set of edges observed in the original input graph, represents the connection relationship between node pairs ( in the real adjacency matrix, represents the predicted connection probability between node pairs ( in the predicted adjacency matrix.

[0098] The calculation formula for the contrast loss is:

[0099]

[0100] where, is the weight coefficient, and are the confidence thresholds for positive and negative samples respectively, and are nodes Sum node The final latent representation of representation and perform a dot product operation, , is the sigmoid function. Among them, positive samples refer to node pairs with connections in the true adjacency matrix, and negative samples refer to node pairs without connections in the true adjacency matrix.

[0101] The calculation formula for the combined loss is:[[]]

[0102]

[0103] Among them, is the regularization parameter. When iteratively training the link prediction model, the gradient descent method is used to minimize the combined loss L , update the parameters of the link prediction model, so as to obtain a trained link prediction model. Based on this trained link prediction model, the link prediction results of the original input graph can be obtained. Taking the cell-cell interaction network graph as an example, the cell pairs with potential interactions in the cell-cell interaction network graph can be predicted through the trained link prediction model.

[0104] Please refer to Figure 3 , Figure 3 is the algorithm framework diagram of the link prediction method adapted to non-centered distribution data provided by the embodiments of the present application. As Figure 3 shown, the original input graph undergoes mean subtraction preprocessing and introduces random perturbations through Gaussian distribution to obtain the final latent representation of the nodes. The result of multiplying the transposed final latent representation of the nodes by the final latent representation of the nodes is input into the bidirectional activation function to capture the positive and negative correlation relationships between the nodes, and after passing through the sigmoid function, the predicted adjacency matrix is output. Based on the predicted adjacency matrix and the true adjacency matrix, the contrastive learning loss is calculated to enhance the feature discrimination ability of the nodes, and the parameters of the link prediction model are updated through backpropagation, and finally the reconstructed graph of the original input graph is obtained. Compared with the original input graph, there are two more purple edges in the reconstructed graph, and these two purple edges are the potential links predicted based on the original input graph by the link prediction method of the present application.

[0105] The link prediction method for non-centrally distributed data provided by this application first generates a node feature matrix and a true adjacency matrix based on the original input graph and the feature information of each node in the original input graph; then encodes the node feature matrix and the true adjacency matrix to obtain the initial latent representation of each node, and subtracts the mean of the latent representations of the nodes from the initial latent representation of the nodes to obtain the final latent representation of each node; finally, predicts the potential connections between entities based on the final latent representation of each node to obtain the predicted adjacency matrix of the original input graph, and calculates the combined loss based on the true adjacency matrix, the predicted adjacency matrix and the final latent representation of each node, and iteratively trains the link prediction model based on the combined loss to obtain the link prediction result of the original input graph. The method of this application introduces the mean subtraction preprocessing technology to perform mean centering on the latent representations of the nodes, eliminates the non-informative offsets in the input data, simplifies the distribution of the latent space, makes the model more stable, and reduces redundant calculations, thereby improving the training efficiency and stability of the model.

[0106] According to an embodiment of this application, an electronic device is provided. As Figure 4 shown, it is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of this application. The electronic device 100 includes a processor 10, a memory 20, and a communication interface 30. The processor 10, the memory 20, and the communication interface 30 are connected by lines. In Figure 4 the embodiment shown, the processor 10, the memory 20, and the communication interface 30 are communicatively connected to each other through a bus.

[0107] The memory 20 is used to store software programs, computer-executable program instructions, etc. The memory 20 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the electronic device, etc.

[0108] The memory 20 may be a read-only memory (ROM), or may be other types of static storage devices that can store static information and instructions, may also be a random access memory (RAM), or may be other types of dynamic storage devices that can store information and instructions, and may also be an electrically erasable programmable read-only memory (EEPROM). Specifically, it is not limited here.

[0109] Exemplarily, the aforementioned memory 20 may be a double data rate synchronous dynamic random access memory DDR SDRAM (referred to as DDR for short). The memory 20 may exist independently, but is connected to the processor 10. Optionally, the memory 20 may also be integrated with the processor 10 into one body. For example, it is integrated within one or more chips.

[0110] In some embodiments, the memory 20 may optionally include memories remotely arranged relative to the processor 10, and these remote memories may be connected to the electronic device through a network. Examples of the above-mentioned network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.

[0111] The processor 10 connects various parts of the entire electronic device 100 through various interfaces and lines, and by running or executing software programs stored in the memory 20, and calling data stored in the memory 20, performs various functions of the electronic device and processes data, for example, implementing the method described in any embodiment of the present application.

[0112] The processor 10 may be a field programmable gate array (FPGA), a digital signal processor (DSP), a central processing unit (CPU), etc.

[0113] The processor 10 may be a single-core processor or a multi-core processor. For example, the processor 10 may be composed of multiple FPGAs or multiple DSPs. In addition, the processor 10 may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions). The processor 10 may be a separate semiconductor chip, or may be integrated with other circuits into a semiconductor chip. For example, it may form a system on a chip (SoC) with other circuits (such as codec circuits, hardware acceleration circuits, or various bus and interface circuits), or may also be integrated as an embedded processor of an application specific integrated circuit (ASIC) in the ASIC. The ASIC integrated with the processor may be separately packaged or may also be packaged together with other circuits.

[0114] The communication interface 30 may use a transceiver device such as a transceiver to implement communication between the electronic device and other devices or communication networks.

[0115] The embodiments of the present application also provide a computer storage medium. The computer storage medium stores instructions or programs, and these instructions or programs are executed by one or more processors. For example, Figure 4 one of the processors 10 in

[0116] can enable the above one or more processors to execute the link prediction method adapted to non-centrally distributed data in any of the above method embodiments.

[0117] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions. Therefore, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A link prediction method adapted to non-centrally distributed data, characterized in that: The method comprises: Acquire an original input graph and feature information of each node in the original input graph, wherein the original input graph includes a node set and an edge set; Generate a node feature matrix and a true adjacency matrix of the original input graph based on the original input graph and feature information of each node in the original input graph; Encoding the node feature matrix and the true adjacency matrix to obtain an initial potential representation of each node, and subtracting a potential representation mean of the node from the initial potential representation of the node to obtain a final potential representation of each node; Predicting potential connections between entities based on the final potential representation of each node to obtain a predicted adjacency matrix of the original input graph; Calculating a combined loss based on the true adjacency matrix, the predicted adjacency matrix, and a final potential representation of each node, and iteratively training a link prediction model based on the combined loss to obtain a link prediction result of the original input graph; The encoding of the node feature matrix and the true adjacency matrix to obtain the initial potential representation of each node includes: Encoding the node feature matrix and the true adjacency matrix based on a first graph convolutional network encoder to obtain a potential representation mean of each node; Encoding the node feature matrix and the true adjacency matrix based on a second graph convolutional network encoder to obtain the logarithmic variance of the potential representation of each node; Generate an initial latent representation of each node from a Gaussian distribution based on the latent representation mean and latent representation log variance of each node; The calculating of the combined loss based on the true adjacency matrix, the predicted adjacency matrix and the final potential representation of each node comprises: Calculating a reconstruction loss based on a similarity between the true adjacency matrix and the predicted adjacency matrix; Compute contrastive loss based on the difference in final latent representations of different node pairs; Calculating a combined loss based on the reconstruction loss and the contrast loss; The calculation formula of the reconstruction loss is: in, represents the set of edges observed in the original input graph, Represents the node pairs in the real adjacency matrix ( The connection relationship between Represents the node pair in the predicted adjacency matrix ( The predicted connection probability between The calculation formula of the contrast loss is: in, is the weight coefficient, and are the confidence thresholds of positive and negative samples respectively, and The nodes are and nodes The final potential representation of express and Perform a dot product operation, , is the sigmoid function.

2. The method according to claim 1, characterized in that The predicting of the potential connections between the entities based on the final potential representation of each node to obtain the predicted adjacency matrix of the original input graph includes: Calculate the strength of the relationship between each node based on the final potential representation of each node; The relationship strength between each node is input into a bidirectional activation function, and the positive and negative related information between nodes is predicted by the bidirectional activation function to obtain a predicted adjacency matrix of the original input graph.

3. The method according to claim 2, characterized in that The node pairs in the predicted adjacency matrix ( The calculation formula for the predicted connection probability between is: in, and The nodes are and nodes The final potential representation of express and Perform a dot product operation, , is the sigmoid function.

4. The method according to any one of claims 1 to 3, characterized in that: The original input graph is a cell-cell interaction network graph, the characteristic information of each cell in the cell-cell interaction network graph is the original expression amount of each gene measured in the cell, and the node characteristic matrix of the original input graph is generated based on the original input graph and the characteristic information of each node in the original input graph, including: Gene expression matrix is ​​created based on the raw expression of each gene measured in each cell in the cell-cell interaction network diagram X ,in, , is the number of cells, is the number of genes measured in each cell; The gene expression matrix X Perform normalization processing to obtain the normalized expression level of each gene measured in each cell; Calculate the variance of the original expression or normalized expression of each gene in all cells, and select the top gene with the largest variance. genes as selection traits; Based on the normalized expression of each cell selection feature, a reduced-dimensional gene expression matrix is ​​constructed. As the node feature matrix of the original input graph, .

5. The method according to claim 4, characterized in that The generating a true adjacency matrix of the original input graph based on the original input graph and the feature information of each node in the original input graph comprises: Generating an initial adjacency matrix of the original input graph based on the original input graph; The interaction potential between each cell pair was calculated based on the similarity of gene expression profiles; The interaction potential is used as the weight of the initial adjacency matrix to obtain the true adjacency matrix of the original input graph.

6. An electronic device, characterized in that: The method comprises at least one processor and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 5.

7. A computer storage medium, characterized in that: The computer storage medium stores instructions or programs, and when the instructions or programs are executed by at least one processor, the at least one processor is caused to perform the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Alzheimer's disease auxiliary diagnosis system based on fNIRS and graph neural network

    CN111466876A

  • Dynamic link prediction model robustness enhancement method based on reinforcement learning

    CN112580728A