Link prediction method and device suitable for non-central distributed data, and storage medium

By introducing mean subtraction preprocessing technology into the link prediction method, the average centralized processing of the potential representation of the node is solved, and the problem of poor performance in the non-central distributed data processing in the prior art is solved, and the accuracy and computing efficiency of link prediction are improved.

CN119988688AActive Publication Date: 2025-05-13BEI JING NORMAL UNIV HONG KONG BAPTIST UNIV UNITED INT COLLEGE
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510473722.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

Existing link prediction methods perform poorly when processing non-center distributed data, resulting in increased redundant calculations and model complexity, inaccuracy and computational efficiency.

Method used

By obtaining the original input graph and node feature information, the node feature matrix and the real adjacency matrix are generated, the initial latent representation is obtained by encoding, and the final latent representation is obtained through mean subtraction processing. The potential connection prediction is performed based on these representations, and the combination loss is calculated for iterative training.

Benefits of technology

This method eliminates non-informative offsets in the input data through mean-centralized processing, simplifies the distribution of potential space, makes the model more stable, reduces redundant calculations, and improves the training efficiency and stability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988688A_ABST
    Figure CN119988688A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of network link prediction, in particular to a link prediction method and device suitable for non-central distributed data and a storage medium. The link prediction method comprises the following steps: generating a node feature matrix and a real adjacency matrix based on an original input graph and feature information of each node in the original input graph; encoding the node feature matrix and the real adjacent matrix to obtain an initial potential representation of each node, and subtracting the potential representation mean value of the node from the initial potential representation of the node to obtain a final potential representation of each node; and predicting the potential connection between the entities based on the final potential representation of each node, and calculating combination loss based on a prediction result to perform iterative training on the model to obtain a link prediction result of the original input graph. According to the method, the mean value subtraction preprocessing technology is introduced, distribution of potential space is simplified, the model is more stable, redundant calculation is reduced, and therefore the training efficiency and stability of the model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of network link prediction, and particularly to a link prediction method, device, and storage medium adapted to non-centered distributed data. Background Art

[0002] With the rapid development of network science and bioinformatics, the analysis of graph-structured data has become one of the core tasks in many fields. In these tasks, link prediction is a key issue in graph analysis and is widely applied in fields such as social networks, recommendation systems, biological networks, knowledge graphs, etc. Link prediction aims to predict whether there are potential connections between node pairs in a graph or to predict unobserved edges.

[0003] Traditional link prediction methods mainly rely on similarity metrics of nodes or edges, usually calculating the similarity between nodes and inferring potential connections based on this. However, these methods often ignore the global structure of the graph and the complex high-dimensional feature information between nodes. Especially when dealing with non-centered distributed data, it often leads to redundant calculations and an increase in model complexity by fitting the mean of the covariance matrix.

[0004] Therefore, there is an urgent need for a link prediction method adapted to non-centered distributed data to further improve the accuracy and computational efficiency of link prediction through a more effective node feature learning, latent representation generation, and decoding process. Summary of the Invention

[0005] Embodiments of the present application aim to provide a link prediction method, device, and storage medium adapted to non-centered distributed data to solve the problem that the existing link prediction methods perform poorly when dealing with non-centered distributed data.

[0006] To solve the above technical problems, the embodiments of the present application provide the following technical solutions: According to a first aspect of the present application, there is provided a link prediction method adapted to non-centered distributed data, the method comprising: Obtaining an original input graph and feature information of each node in the original input graph, the original input graph including a node set and an edge set; Generating a node feature matrix and a true adjacency matrix of the original input graph based on the original input graph and the feature information of each node in the original input graph; Encoding the node feature matrix and the true adjacency matrix to obtain an initial latent representation of each node, and subtracting the latent representation mean of the node from the initial latent representation of the node to obtain a final latent representation of each node; Predicting potential connections between entities based on the final latent representation of each node to obtain a predicted adjacency matrix of the original input graph; The combined loss is calculated based on the true adjacency matrix, the predicted adjacency matrix and the final potential representation of each node, and the link prediction model is iteratively trained based on the combined loss to obtain the link prediction result of the original input graph.

[0007] Optionally, encoding the node feature matrix and the true adjacency matrix to obtain the initial potential representation of each node includes: Based on the first graph convolutional network encoder, the node feature matrix and the true adjacency matrix are encoded to obtain the potential representation mean of each node; Based on the second graph convolutional network encoder, the node feature matrix and the true adjacency matrix are encoded to obtain the logarithmic variance of the potential representation of each node; Generate the initial potential representation of each node from the Gaussian distribution based on the potential representation mean and potential representation logarithmic variance of each node.

[0008] Optionally, the predicting of the potential connections between the entities based on the final potential representation of each node to obtain the predicted adjacency matrix of the original input graph includes: Calculate the strength of the relationship between nodes based on the final potential representation of each node; The relationship strength between each node is input into the bidirectional activation function, and the positive and negative related information between nodes is predicted by the bidirectional activation function to obtain the predicted adjacency matrix of the original input graph.

[0009] Optionally, the node pairs in the predicted adjacency matrix ( The calculation formula for the predicted connection probability between is:

[0010] Among them, and They are nodes and nodes The final potential representation of , indicates and Perform a dot product operation, , is the sigmoid function.

[0011] Optionally, the calculating the combined loss based on the true adjacency matrix, the predicted adjacency matrix and the final potential representation of each node includes: Calculating reconstruction loss based on the similarity between the true adjacency matrix and the predicted adjacency matrix; Compute contrastive loss based on the difference in final potential representations of different node pairs; A combined loss is calculated based on the reconstruction loss and the contrast loss.

[0012] Optionally, the calculation formula of the reconstruction loss is:

[0013] Among them, represents the set of edges observed in the original input graph, represents the node pairs in the real adjacency matrix ( The connection relationship between indicates the node pair in the predicted adjacency matrix ( The predicted connection probability between The calculation formula of the contrast loss is:

[0014] Among them, is the weight coefficient, and are the confidence thresholds of positive and negative samples, respectively, and They are nodes and nodes The final potential representation of , indicates and Perform a dot product operation, , is the sigmoid function. Optionally, the original input graph is a cell-cell interaction network graph, the characteristic information of each cell in the cell-cell interaction network graph is the original expression amount of each gene measured in the cell, and the node characteristic matrix of the original input graph generated based on the original input graph and the characteristic information of each node in the original input graph includes: Create a gene expression matrix based on the raw expression levels of each gene measured in each cell in the cell-cell interaction network diagram X , among which, , is the number of cells, is the number of genes measured in each cell; For the gene expression matrix X Perform normalization processing to obtain the normalized expression level of each gene measured in each cell; Calculate the variance of the original expression or normalized expression of each gene in all cells, and select the top cells with the largest variance. Genes as selection features; Based on the normalized expression of each cell selection feature, a dimensionality-reduced gene expression matrix is ​​constructed As the node feature matrix of the original input graph, where​ .

[0015] Optionally, generating a true adjacency matrix of the original input graph based on the original input graph and feature information of each node in the original input graph includes: Generating an initial adjacency matrix of the original input graph based on the original input graph; Calculate the interaction potential between each cell pair based on the similarity of gene expression profiles; Use the interaction potential as the weight of the initial adjacency matrix to obtain the true adjacency matrix of the original input graph.

[0016] According to the second aspect of the present application, an electronic device is provided, comprising at least one processor and a memory in communication with the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned method.

[0017] According to the third aspect of the present application, a computer storage medium is provided, wherein the computer storage medium stores instructions or programs, and when the instructions or programs are executed by at least one processor, the at least one processor executes the above-mentioned method.

[0018] The beneficial effect of the embodiment of the present application is as follows: Different from the prior art, the embodiment of the present application provides a link prediction method suitable for non-centrally distributed data, firstly generating a node feature matrix and a true adjacency matrix based on the original input graph and the feature information of each node in the original input graph; then encoding the node feature matrix and the true adjacency matrix to obtain the initial potential representation of each node, and subtracting the mean of the potential representation of the node from the initial potential representation of the node to obtain the final potential representation of each node; finally, predicting the potential connection between each entity based on the final potential representation of each node to obtain the predicted adjacency matrix of the original input graph, and calculating the combined loss based on the true adjacency matrix, the predicted adjacency matrix and the final potential representation of each node, iteratively training the link prediction model based on the combined loss to obtain the link prediction result of the original input graph. The method of the present application introduces the mean subtraction preprocessing technology, performs mean centralization processing on the potential representation of the node, eliminates the non-informative offset in the input data, simplifies the distribution of the potential space, makes the model more stable, and reduces redundant calculations, thereby improving the training efficiency and stability of the model. Brief Description of the Figures

[0019] One or more embodiments are exemplarily described by the pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements. Unless otherwise stated, the drawings in the drawings do not constitute a scale limitation.

[0020] Figure 1 is a flowchart of a link prediction method suitable for non-centrally distributed data provided by an embodiment of the present application; Figure 2 is a schematic diagram of the structure of the original input graph provided in the embodiment of the present application; Figure 3 is an algorithm framework diagram of a link prediction method adapted to non-centrally distributed data provided in an embodiment of the present application; Figure 4 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. Specific implementation method

[0021] To make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of them. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present application.

[0022] In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other.

[0023] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0024] Please refer to Figure 1 , Figure 1 is a flow chart of a link prediction method adapted to non-centrally distributed data provided by an embodiment of the present application, the method comprising: Step S101, obtaining the original input graph and the feature information of each node in the original input graph.

[0025] Specifically, the original input graph is a graph structure data, including a node set and an edge set. Each node in the node set represents an entity, and each edge in the edge set represents a certain relationship or interaction between the corresponding two entities. The feature information of the node refers to the descriptive data or attributes of each node, which is used to characterize the characteristics, state or behavior of the node.

[0026] If Figure 2 As shown, it is an example diagram of the original input graph. The original input graph is an undirected graph, including 12 nodes, node 21 is one of the nodes, and there are 4 edges between node 21 and other nodes. In other embodiments, the original input graph may also be a directed graph, which is not limited in this application. The original input graph may be a social network graph representing user relationships in a social network, a paper citation structure graph representing paper citation relationships, a knowledge graph representing knowledge point relationships, a cell-cell interaction network graph representing interactions between cells, etc. Among them, in the social network graph, each node represents a user, and each edge represents a relationship between users, such as a friend relationship, a follow-up relationship, or a message exchange. The user's characteristic information may include the user's age, activity, number of posts, number of comments, etc. In the cell-cell interaction network graph, each node represents a cell, and each edge represents a physical or functional interaction between two cells, such as a ligand-receptor interaction, a cell adhesion interaction, a cell signaling, etc. The characteristic information of a cell may include a gene expression profile (transcription level of each gene in the cell), cell surface molecules, receptor expression, cell type, cell state, metabolic state, etc.

[0027] Step S102, generating a node feature matrix and a true adjacency matrix of the original input graph based on the original input graph and the feature information of each node in the original input graph.

[0028] In some embodiments, before generating the node feature matrix of the input graph based on the feature information of each node in the original input graph, the feature information of each node in the original input graph is preprocessed. The preprocessing includes normalization and feature screening. By normalizing the feature information of each node in the original input graph, the scale difference between features can be eliminated and the bias of the model on certain features can be avoided. By screening the feature information of each node in the original input graph, the most informative and representative features of the target variable can be screened out from the original feature set, which can improve the performance of the model, reduce the computational cost, and enhance the generalization ability of the model.

[0029] The following takes the cell-cell interaction network diagram as an example to illustrate the generation process of the node feature matrix of the original input diagram. It is assumed that the characteristic information of each cell is the original expression level of each gene measured in the cell.

[0030] Step 1: Create a gene expression matrix based on the raw expression levels of each gene measured in each cell in the cell-cell interaction network diagram X .

[0031] Among them, , is the number of cells, is the number of genes measured in each cell.

[0032] Step 2: Normalize the gene expression matrix X based on the following formula:

[0033] Among them, It's a cell Medium Gene The original expression of , Cells Medium Gene The normalized expression level of .

[0034] Step 3: Calculate the variance of the original expression or normalized expression of each gene in all cells based on the following formula as the cell variation characteristic, and select the top cells with the largest variation characteristic. Genes as selected features:

[0035] Among them, For cells Medium Gene The original expression or normalized expression.

[0036] Step 4: Construct a dimensionality-reduced gene expression matrix based on the normalized expression of each cell selection feature As the node feature matrix of the original input graph, where .

[0037] It should be noted that the order of normalization and feature selection is not limited. There are usually two common methods: 1) Perform feature selection first, and then normalize the selected features. The advantage of this method is that it reduces the amount of calculation and only normalizes important features. The disadvantage is that some feature selection methods may be affected by the feature scale, resulting in inaccurate results. 2) Normalize all features first, and then perform feature selection. The advantage of this method is that it ensures that feature selection is not affected by scale, and the disadvantage is that the amount of calculation is large.

[0038] The real adjacency matrix of the original input graph can be constructed based on the point set and edge set in the original input graph. If the original input graph includes n nodes, the real adjacency matrix is ​​specifically a matrix of n×n dimensions. For example, let A represent the real adjacency matrix, and let the node 、Node are any two nodes in the original input graph, is the th in the real adjacency matrix A Line No. An item in the column, then =1 indicates that the node exists in the original input graph​ Point to Node 's edge, =0 means the node does not exist in the original input graph Point to Node 's edge.

[0039] In one embodiment, the strength of the relationship between nodes can be incorporated into the real adjacency matrix as a weight. Taking the cell-cell interaction network graph as an example, the step of generating the real adjacency matrix of the original input graph based on the original input graph and the feature information of each node in the original input graph may include: first, generating the initial adjacency matrix of the original input graph based on the point set and edge set of the original input graph. The initial adjacency matrix is ​​first Line No. The values ​​of the column matrix elements are 0 or 1. Secondly, the interaction potential between each cell pair is calculated based on the similarity of the expression profiles. Finally, the interaction potential is used as the weight of the initial adjacency matrix to obtain the true adjacency matrix of the original input graph.

[0040] In one embodiment, the Pearson correlation coefficient is used to measure the interaction potential between cell pairs, and its calculation formula is:

[0041] Among them, and For cell points and cell points Medium Gene 's expression level (can be original expression level or normalized expression level), and For cell points and cell points The average expression level of all genes in .

[0042] Use the Pearson correlation coefficient to perform weighted calculation on the initial adjacency matrix to obtain the true adjacency matrix of the original input graph. Specifically, the first Line No. The value of the column matrix element , where is the first in the initial adjacency matrix Line No. Column matrix element values.

[0043] Step S103, encode the node feature matrix and the true adjacency matrix to obtain the initial potential representation of each node, and subtract the mean potential representation of the node from the initial potential representation of the node to obtain the final potential representation of each node.

[0044] ​Specifically, a link prediction model is pre-built, which includes an encoder and a decoder. The encoder is used to encode and sample the node matrix and the true adjacency matrix to obtain the final potential representation of the node; the decoder is used to predict based on the final potential representation of the node to obtain the predicted adjacency matrix of the original input graph.

[0045] In one embodiment, encoding the node feature matrix and the true adjacency matrix to obtain the initial potential representation of each node specifically includes the following steps: First, the node feature matrix and the true adjacency matrix are encoded based on the graph convolutional network encoder to obtain the node's potential representation mean and potential representation logarithmic variance.

[0046] Specifically, the node feature matrix and the true adjacency matrix are encoded based on the first graph convolutional network encoder to obtain the potential representation mean of each node; the node feature matrix and the true adjacency matrix are encoded based on the second graph convolutional network encoder to obtain the potential representation logarithmic variance of each node. This method can help the model learn the structure of the latent space more flexibly. Each encoder can learn different feature representations, thereby generating means and variances through different networks, which can help better represent the distribution of data.

[0047] Secondly, the initial latent representation of each node is generated from the Gaussian distribution based on the latent representation mean and latent representation logarithmic variance of each node.

[0048] Specifically, the calculation formula for the initial potential representation of each node is:

[0049] Among them, Is a node The initial latent representation of , Is a node The latent mean of , Is a node The latent log-variance of , Values ​​sampled from a standard normal distribution.

[0050] In the above embodiments, the initial potential representation of the node is generated from the Gaussian distribution by sampling, and a probabilistic representation is provided for each node, which can make the model more robust and generalizable in the prediction task.

[0051] In other embodiments, the node feature matrix and the true adjacency matrix can also be encoded based on a graph convolutional network encoder to directly obtain the initial potential representation of each node.

[0052] After obtaining the initial potential of the node, the final potential representation of each node is obtained by subtracting the mean of the potential representation of the node from the initial potential representation of the node. Taking the cell-cell interaction network diagram as an example, the final potential representation of each node is the final potential representation of each cell. In the embodiment of the present application, the potential representation of the node is mean-centered by using the mean subtraction technique, which can improve the efficiency and accuracy of the model when processing non-centrally distributed data.

[0053] Step S104, predicting the potential connections between entities based on the final potential representation of each node, and obtaining the predicted adjacency matrix of the original input graph.

[0054] In one embodiment, the decoder of the link prediction model obtains the relationship strength between nodes by calculating the inner product of the final potential representation of each node, and then inputs the relationship strength between nodes into the bidirectional activation function, and predicts the positive and negative correlation information between nodes through the bidirectional activation function to obtain the predicted adjacency matrix of the original input graph. Specifically, the node pairs ( The calculation formula for the predicted connection probability between is:

[0055] Among them, and They are nodes and nodes The final potential representation of , indicates and Perform a dot product operation, , is the sigmoid function.

[0056] The traditional GAE model uses the Sigmoid function, which can only process positive correlation information between nodes and is difficult to capture negative correlation node relationships. This invention proposes a bidirectional activation function that can simultaneously process positive and negative correlation relationships. Through this bidirectional activation function, the model can not only capture the positive connection between nodes, but also effectively express the inhibitory or antagonistic relationship between nodes, thereby better modeling complex graph structures.

[0057] Step S105, calculating the combined loss based on the real adjacency matrix, the predicted adjacency matrix and the final potential representation of each node, iteratively training the link prediction model based on the combined loss, and obtaining the link prediction result of the original input graph.

[0058] Existing models are weak in distinguishing nodes with similar features but different connection patterns, resulting in insufficient prediction accuracy in some tasks. To solve this problem, this paper introduces a contrast loss function to enhance the model's ability to distinguish subtle differences between similar nodes, thereby improving the model's prediction performance in complex graph structures.

[0059] In one embodiment, the combined loss includes reconstruction loss and contrast loss, wherein the reconstruction loss is calculated based on the similarity between the true adjacency matrix and the predicted adjacency matrix, and the contrast loss is calculated based on the final potential representation difference of different node pairs. Specifically, the calculation formula of the reconstruction loss is:

[0060] Among them, represents the set of edges observed in the original input graph, represents the node pairs in the real adjacency matrix ( The connection relationship between indicates the node pair in the predicted adjacency matrix ( The predicted connection probability between

[0061] The calculation formula for contrast loss is:

[0062] Among them, is the weight coefficient, and are the confidence thresholds of positive and negative samples, respectively, and They are nodes and nodes The final potential representation of , indicates and Perform a dot product operation, , is the sigmoid function. Among them, positive samples refer to node pairs that are connected in the true adjacency matrix, and negative samples refer to node pairs that are not connected in the true adjacency matrix.

[0063] The formula for calculating the combined loss is:

[0064] Among them, is the regularization parameter. When iteratively training the link prediction model, the gradient descent method is used to minimize the combined loss L , update the parameters of the link prediction model to obtain a trained link prediction model. Based on the trained link prediction model, the link prediction result of the original input graph can be obtained. Taking the cell-cell interaction network graph as an example, the trained link prediction model can predict the cell pairs with potential interactions in the cell-cell interaction network graph.

[0065] Please refer to Figure 3 , Figure 3This is an algorithm framework diagram of a link prediction method suitable for non-centrally distributed data provided by an embodiment of the present application. Figure 4 As shown, the original input graph is preprocessed by mean subtraction and random perturbations are introduced through Gaussian distribution to obtain the final potential representation of the node. The final potential representation of the node is transposed and multiplied with the final potential representation of the node. The result is input into the bidirectional activation function to capture the positive and negative correlation between the nodes, and the predicted adjacency matrix is ​​output after the sigmoid function. The feature discrimination ability of the nodes is enhanced based on the calculation of the contrast learning loss between the predicted adjacency matrix and the true adjacency matrix, and the parameters of the link prediction model are updated through back propagation, and finally the reconstructed graph of the original input graph is obtained. Compared with the original input graph, the reconstructed graph has two more purple edges. These two purple edges are the potential links predicted by the link prediction method of this application based on the original input graph.

[0066] The link prediction method adapted to non-centrally distributed data provided by the present application first generates a node feature matrix and a true adjacency matrix based on the original input graph and the feature information of each node in the original input graph; then the node feature matrix and the true adjacency matrix are encoded to obtain the initial potential representation of each node, and the initial potential representation of the node is subtracted from the mean of the potential representation of the node to obtain the final potential representation of each node; finally, the potential connection between each entity is predicted based on the final potential representation of each node to obtain the predicted adjacency matrix of the original input graph, and the combined loss is calculated based on the true adjacency matrix, the predicted adjacency matrix and the final potential representation of each node, and the link prediction model is iteratively trained based on the combined loss to obtain the link prediction result of the original input graph. The method of the present application introduces the mean subtraction preprocessing technology to perform mean centralization processing on the potential representation of the node, eliminates the non-informative offset in the input data, simplifies the distribution of the potential space, makes the model more stable, and reduces redundant calculations, thereby improving the training efficiency and stability of the model.

[0067] According to an embodiment of the present application, an electronic device is provided, such as Figure 4 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. The electronic device 100 includes a processor 10, a memory 20 and a communication interface 30. The processor 10, the memory 20 and the communication interface 30 are connected through a line. Figure 4 In the embodiment shown, the processor 10, the memory 20 and the communication interface 30 are connected to each other through a bus.

[0068] The memory 20 is used to store software programs, computer executable program instructions, etc. The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required for at least one function; the data storage area may store data created according to the use of the electronic device, etc. ​​​​

[0069] The memory 20 may be a read-only memory (ROM), or other types of static storage devices that can store static information and instructions, or a random access memory (RAM), or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), which is not specifically limited here.

[0070] Exemplarily, the aforementioned memory 20 may be a double data rate synchronous dynamic random access memory DDRSDRAM (DDR for short). The memory 20 may exist independently but be connected to the processor 10. Optionally, the memory 20 may also be integrated with the processor 10. For example, integrated into one or more chips.

[0071] In some embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the electronic device via a network. Examples of the above network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0072] The processor 10 uses various interfaces and lines to connect various parts of the entire electronic device 100, and executes various functions of the electronic device and processes data by running or executing software programs stored in the memory 20 and calling data stored in the memory 20, such as implementing the method described in any embodiment of the present application.

[0073] The processor 10 can be a field programmable gate array (FPGA), a digital signal processor (DSP), a central processing unit (CPU), etc.

[0074] The processor 10 may be a single-core processor or a multi-core processor. For example, the processor 10 may be composed of multiple FPGAs or multiple DSPs. In addition, the processor 10 may refer to one or more devices, circuits and / or processing cores for processing data (such as computer program instructions). The processor 10 may be a separate semiconductor chip or may be integrated into a semiconductor chip together with other circuits. For example, it may form a system on a chip (SoC) with other circuits (such as codec circuits, hardware acceleration circuits or various buses and interface circuits), or it may be integrated into an application specific integrated circuit (ASIC) as a built-in processor of the ASIC. The ASIC with the integrated processor may be packaged separately or packaged together with other circuits.

[0075] The communication interface 30 can use a transceiver such as a transceiver to achieve communication between the electronic device and other devices or a communication network.

[0076] The present application also provides a computer storage medium, which stores instructions or programs, and the instructions or programs are executed by one or more processors, such as ​ A processor 10 in can enable the one or more processors to execute the link prediction method adapted to non-centrally distributed data in any of the above method embodiments.

[0077] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a general hardware platform, or of course by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the relevant technology, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiment.

[0078] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application. Therefore, any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present application should be included in the protection scope of the present application.

Claims

1. A link prediction method adapted to non-centrally distributed data, characterized in that: The method comprises: Acquire an original input graph and feature information of each node in the original input graph, wherein the original input graph includes a node set and an edge set; Generate a node feature matrix and a true adjacency matrix of the original input graph based on the original input graph and feature information of each node in the original input graph; Encoding the node feature matrix and the true adjacency matrix to obtain an initial potential representation of each node, and subtracting a potential representation mean of the node from the initial potential representation of the node to obtain a final potential representation of each node; Predicting potential connections between entities based on the final potential representation of each node to obtain a predicted adjacency matrix of the original input graph; A combined loss is calculated based on the true adjacency matrix, the predicted adjacency matrix and the final potential representation of each node, and a link prediction model is iteratively trained based on the combined loss to obtain a link prediction result of the original input graph.

2. The method according to claim 1, characterized in that The node feature matrix and the true adjacency matrix are encoded to obtain the initial potential representation of each node, including: Encoding the node feature matrix and the true adjacency matrix based on a first graph convolutional network encoder to obtain a potential representation mean of each node; Encoding the node feature matrix and the true adjacency matrix based on a second graph convolutional network encoder to obtain the logarithmic variance of the potential representation of each node; The initial latent representation of each node is generated from a Gaussian distribution based on the latent representation mean and latent representation log variance of each node.

3. The method according to claim 1, characterized in that The predicting of the potential connections between the entities based on the final potential representation of each node to obtain the predicted adjacency matrix of the original input graph includes: Calculate the strength of the relationship between each node based on the final potential representation of each node; The relationship strength between each node is input into a bidirectional activation function, and the positive and negative related information between nodes is predicted by the bidirectional activation function to obtain a predicted adjacency matrix of the original input graph.

4. The method according to claim 3, characterized in that The node pairs in the predicted adjacency matrix ( The calculation formula for the predicted connection probability between is: in, and The nodes are and nodes The final potential representation of express and Perform a dot product operation, , is the sigmoid function.

5. The method according to claim 1, characterized in that: The calculating of the combined loss based on the true adjacency matrix, the predicted adjacency matrix and the final potential representation of each node comprises: Calculating a reconstruction loss based on a similarity between the true adjacency matrix and the predicted adjacency matrix; Compute contrastive loss based on the difference in final latent representations of different node pairs; A combined loss is calculated based on the reconstruction loss and the contrastive loss.

6. The method according to claim 5, characterized in that The calculation formula of the reconstruction loss is: in, represents the set of edges observed in the original input graph, Represents the node pairs in the real adjacency matrix ( The connection relationship between Represents the node pair in the predicted adjacency matrix ( The predicted connection probability between The calculation formula of the contrast loss is: in, is the weight coefficient, and are the confidence thresholds of positive and negative samples respectively, and The nodes are and nodes The final potential representation of express and Perform a dot product operation, , is the sigmoid function.

7. The method according to any one of claims 1 to 6, characterized in that: The original input graph is a cell-cell interaction network graph, the characteristic information of each cell in the cell-cell interaction network graph is the original expression amount of each gene measured in the cell, and the node characteristic matrix of the original input graph is generated based on the original input graph and the characteristic information of each node in the original input graph, including: Gene expression matrix is ​​created based on the raw expression of each gene measured in each cell in the cell-cell interaction network diagram X ,in, , is the number of cells, is the number of genes measured in each cell; The gene expression matrix X Perform normalization processing to obtain the normalized expression level of each gene measured in each cell; Calculate the variance of the original expression or normalized expression of each gene in all cells, and select the top gene with the largest variance. genes as selection traits; Based on the normalized expression of each cell selection feature, a reduced-dimensional gene expression matrix is ​​constructed. As the node feature matrix of the original input graph, .

8. The method according to claim 7, characterized in that The generating a true adjacency matrix of the original input graph based on the original input graph and the feature information of each node in the original input graph comprises: Generating an initial adjacency matrix of the original input graph based on the original input graph; The interaction potential between each cell pair was calculated based on the similarity of gene expression profiles; The interaction potential is used as the weight of the initial adjacency matrix to obtain the true adjacency matrix of the original input graph.

9. An electronic device, characterized in that: The method comprises at least one processor and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 8.

10. A computer storage medium, characterized in that: The computer storage medium stores instructions or programs, and when the instructions or programs are executed by at least one processor, the at least one processor is caused to perform the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Alzheimer's disease auxiliary diagnosis system based on fNIRS and graph neural network

    CN111466876A

  • Dynamic link prediction model robustness enhancement method based on reinforcement learning

    CN112580728A

  • Molecular generation method and device, equipment and storage medium

    CN116959617A

  • Latent network summarization

    US20200233864A1