Network embedding method and device based on federated learning network, electronic equipment and computer program product
Federated learning with encrypted matrix decomposition addresses privacy and scalability issues in distributed network embedding, ensuring secure and efficient network embedding across multiple data centers.
Patent Information
- Application Number
- CN202510820047.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-18
AI Technical Summary
In a distributed environment, the existing technology cannot effectively solve the problem of poor data security in network embedding, especially in unsupervised network embedding across data centers. Privacy protection is insufficient and cross-data center solutions are lacking, and the computing and storage costs of large-scale networks cannot be dealt with.
The method based on federated learning network is adopted, and the local network matrix is encrypted by using a preset mask matrix to generate a local encryption matrix, and integrated and decomposed on the embedded server side to ensure that the data remains encrypted in the transmission and calculation process, and the original data of the data holder is not leaked.
It realizes the data security of network embedding in a distributed environment, ensures data privacy protection, and reduces the computing and storage costs of large-scale networks, and is suitable for the embedding needs of small and medium-sized and large-scale networks.
Smart Images

Figure CN120321055A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular, to a network embedding method, apparatus, electronic device, and computer program product based on a federated learning network. Background Art
[0002] With the wide application of Federated Learning (FL), it becomes particularly important to implement privacy-preserving Network Embedding (NE) in a distributed environment. Network Embedding is a technology that maps network nodes to a low-dimensional vector space and is widely used in fields such as financial fraud detection and recommendation systems. However, most traditional network embedding methods adopt a centralized approach, which requires data to be centralized on a single server for processing. This is unrealistic in the current distributed big data environment.
[0003] In a distributed scenario, data is usually stored in data centers at multiple geographical locations, and each data center has partial user or subgraph data. In this context, implementing unsupervised network embedding (UNE) across data centers faces many challenges, including privacy protection, multi-party collaborative computing, and large-scale data processing.
[0004] In addition, in a distributed subgraph scenario, if the participating parties only hold some network nodes and edges, traditional centralized methods are not applicable.
[0005] Regarding network embedding, current research mostly focuses on centralized network embedding methods, such as DeepWalk (random walk algorithm), LINE (Large-scale Information Network Embedding), Node2Vec (node vector generation), and NetMF and NetSMF based on Matrix Factorization (MF). These methods perform well in a single data center environment, but have the following drawbacks in a distributed environment:
[0006] 1. Insufficient privacy protection: The centralized method requires all data to be centralized on a single server, which conflicts with the requirements of distributed data privacy protection.
[0007] 2. Lack of cross-data center solutions: Existing methods cannot be directly applied to distributed data scenarios with privacy protection requirements.
[0008] 3. Unable to handle large-scale networks: Traditional methods have too high computational and storage costs when dealing with large-scale networks with millions of nodes.
[0009] Currently, only a few methods have tried to solve the unsupervised network embedding problem in a distributed environment, such as FedWalk (Federated Random Walk). However, these methods are only applicable to the network embedding scenario at the single-node (ego-level), and cannot meet the embedding requirements in the distributed subgraph scenario.
[0010] Regarding the problem of poor data security in network embedding in a distributed environment in the above-mentioned prior art, no effective solution has been proposed yet. Summary of the Invention
[0011] Embodiments of the present invention provide a network embedding method, device, electronic device, and computer program product based on a federated learning network, so as to at least solve the technical problem of poor data security in network embedding in a distributed environment in the prior art.
[0012] According to an aspect of an embodiment of the present invention, a network embedding method based on a federated learning network is provided, including: obtaining local network matrices of multiple data holders in a federated learning network, where the data holders are participants in the federated learning network, the federated learning network is pre-divided into multiple local learning networks, and the local network matrix is a matrix representation of the local learning network; encrypting the local network matrix using a preset mask matrix to obtain a local encrypted matrix; uploading the local encrypted matrices respectively corresponding to the multiple data holders to an embedding server, and receiving a global encrypted matrix integrated by the embedding server based on the multiple local encrypted matrices, where the global encrypted matrix is the result of decomposing the integration result of the multiple local encrypted matrices by the embedding server; decrypting the global encrypted matrix based on the preset mask matrix to obtain a global embedding matrix, where the global embedding matrix is the network embedding result of the federated learning network.
[0013] Optionally, the local network matrix at least includes: a local adjacency matrix and a local degree matrix corresponding to the local learning network. Obtaining local network matrices of multiple data holders in a federated learning network includes: counting the number of nodes of the data holders in the federated learning network; in the case that the number of nodes meets a first order of magnitude, performing graph structure analysis on the local learning network of each data holder to obtain the local adjacency matrix and the local degree matrix corresponding to the local learning network; in the case that the number of nodes meets a second order of magnitude, performing path sampling on the local learning network of each data holder to generate the local adjacency matrix and the local degree matrix corresponding to the local learning network, where the second order of magnitude is greater than the first order of magnitude.
[0014] Optionally, when the number of nodes meets the second order of magnitude, path sampling is performed on the local learning network at each data holder to generate the local adjacency matrix and the local degree matrix corresponding to the local learning network, including: when the number of nodes meets the second order of magnitude, random walks are performed on the local learning network at each data holder to generate a path set, where the local learning network includes a plurality of nodes, the path set includes multiple path information generated by random walks based on a plurality of preset starting points respectively, the preset starting points are the nodes determined in advance as starting points, and each path information is a sequence of a plurality of nodes; according to the path set, the local adjacency matrix and the local degree matrix corresponding to the local learning network are determined, where the elements in the local adjacency matrix are determined according to the adjacent situation of any two nodes in the plurality of paths, and the elements in the local degree matrix are determined according to the number of different nodes adjacent to the node in the plurality of path information.
[0015] Optionally, the local network matrix at least includes: the local adjacency matrix and the local degree matrix corresponding to the local learning network. Encrypting the local adjacency matrix using a preset mask matrix to obtain a local encrypted matrix includes: when the number of nodes in the federated learning network meets the first order of magnitude, each data holder combines the local adjacency matrix and the local degree matrix to obtain a combined polynomial matrix; receiving the preset mask matrix pre-generated by the mask server, where the mask server pre-generates a random orthogonal matrix as the preset mask matrix; combining the random orthogonal matrix with the combined polynomial matrix to obtain the local encrypted matrix of each data holder.
[0016] Optionally, the local network matrix at least includes: the local adjacency matrix and the local degree matrix corresponding to the local learning network. Encrypting the local adjacency matrix using a preset mask matrix to obtain a local encrypted matrix includes: when the number of nodes in the federated learning network meets the second order of magnitude, each data holder uploads the local degree matrix to the embedding server respectively and receives the global degree matrix returned by the embedding server based on the plurality of local degree matrices; determining a random walk polynomial matrix according to the local degree matrix and the Laplacian matrix of the same data holder, and the global degree matrix, where the Laplacian matrix is determined according to the difference between the local adjacency matrix and the local degree matrix of the same data holder, and the random walk polynomial matrix is expressed as: , where i is the data holder, is the random walk polynomial matrix of the data holder, D is the global degree matrix, is the local degree matrix of the data holder is the Laplacian matrix of the data holder; perform federated power iteration on the random walk polynomial matrix according to a preset number of iterations to generate a random walk sub-matrix, where the federated power iteration is used to project the random walk polynomial matrix into a low-dimensional space; encrypt the random walk sub-matrix using the preset mask matrix to obtain a local encryption matrix.
[0017] Optionally, each step of the federated power iteration includes: each data holder encrypts the random walk polynomial matrix using the preset mask matrix to obtain a local pre-encryption matrix; uploads the local pre-encryption matrix to the embedding server, where the embedding server is used to aggregate multiple local pre-encryption matrices to obtain a global pre-encryption matrix, and perform coefficient matrix decomposition on the global pre-encryption matrix to obtain an intermediate result matrix, where the intermediate result matrix is an orthogonal matrix after masking; receives the intermediate result matrix returned by the embedding server, and removes the mask in the intermediate result matrix according to the preset mask matrix to obtain an intermediate encryption matrix, where when the number of iterations reaches the preset number of iterations, the intermediate encryption matrix is the random walk sub-matrix; when the number of iterations does not reach the preset number of iterations, the intermediate encryption matrix is the random walk polynomial matrix for the next iteration.
[0018] Optionally, the method further includes: aggregating the local encryption matrices respectively uploaded by multiple data holders through the embedding server to obtain a global aggregation matrix; performing singular value decomposition on the global aggregation matrix to generate the global encryption matrix.
[0019] According to another aspect of the embodiments of the present invention, there is also provided a network embedding device based on a federated learning network, including: an acquisition module, configured to acquire local network matrices of multiple data holders in a federated learning network, where the data holders are participants in the federated learning network, the federated learning network is pre-divided into multiple local learning networks, and the local network matrix is a matrix representation of the local learning network; an encryption module, configured to encrypt the local network matrix using a preset mask matrix to obtain a local encryption matrix; an upload module, configured to upload the local encryption matrices respectively corresponding to multiple data holders to an embedding server, and receive the global encryption matrix integrated by the embedding server based on multiple local encryption matrices, where the global encryption matrix is the result of decomposition by the embedding server on the integration result of multiple local encryption matrices; a decryption module, configured to decrypt the global encryption matrix based on the preset mask matrix to obtain a global embedding matrix, where the global embedding matrix is the network embedding result of the federated learning network.
[0020] According to another aspect of the embodiments of the present invention, there is also provided an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the above-mentioned network embedding method based on the federated learning network through the computer program.
[0021] According to another aspect of the embodiments of the present invention, there is also provided a computer program product, including computer instructions, and when the computer instructions are executed by a processor, the steps of the above-mentioned network embedding method based on the federated learning network are implemented.
[0022] In the embodiments of the present invention, each data holder in the federated learning network uploads a locally encrypted matrix encrypted by a preset mask matrix to the embedding server, ensuring that all intermediate calculation results during the embedding process are in encrypted form, ensuring that the local data of the data holder will not be leaked, and the embedding server cannot recover the original data of the data holder, achieving the purpose of ensuring data security, thereby realizing the technical effect of data security for network embedding in a distributed environment, and further solving the technical problem of poor data security for network embedding in the prior art in a distributed environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0024] Figure 1 is a flowchart of a network embedding method based on a federated learning network according to an embodiment of the present invention;
[0025] Figure 2 is a schematic diagram of a federated sparse matrix factorization algorithm (FedNetSMF) according to an embodiment of the present invention;
[0026] Figure 3 is a schematic diagram of a privacy-preserving unsupervised federated network embedding method according to an embodiment of the present invention;
[0027] Figure 4 is a schematic diagram of a network embedding device based on a federated learning network according to an embodiment of the present invention;
[0028] Figure 5 is a structural block diagram of a computer terminal according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data used in appropriate cases can be interchanged so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0031] According to an embodiment of the present invention, there is provided an embodiment of a network embedding method based on a federated learning network. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.
[0032] Figure 1 is a flowchart of a network embedding method based on a federated learning network according to an embodiment of the present invention, as Figure 1 shown, the method includes the following steps:
[0033] Step S102, obtain local network matrices of multiple data holders in the federated learning network, where the data holders are participants in the federated learning network, the federated learning network is pre-divided into multiple local learning networks, and the local network matrix is a matrix representation of the local learning network;
[0034] Step S104, encrypt the local network matrix using a preset mask matrix to obtain a local encrypted matrix;
[0035] Step S106, upload the local encrypted matrices corresponding to multiple data holders to the embedding server respectively, and receive the global encrypted matrix integrated by the embedding server based on the multiple local encrypted matrices, where the global encrypted matrix is the result after the embedding server decomposes the integration result of the multiple local encrypted matrices;
[0036] Step S108: Decrypt the global encryption matrix based on a preset mask matrix to obtain a global embedding matrix, where the global embedding matrix is the network embedding result of the federated learning network.
[0037] In the embodiment of the present invention, each data holder in the federated learning network uploads a locally encrypted matrix encrypted by a preset mask matrix to the embedding server, ensuring that all intermediate calculation results during the embedding process are in encrypted form, ensuring that the local data of the data holder will not be leaked and the embedding server cannot recover the original data of the data holder, achieving the purpose of ensuring data security, thereby realizing the technical effect of ensuring data security for network embedding in a distributed environment, and further solving the technical problem of poor data security for network embedding in the prior art in a distributed environment.
[0038] In the above step S102, the federated learning network includes multiple data holders and at least one embedding server. Among them, the federated learning network can be divided into multiple local learning networks, and each data holder is assigned a local learning network. Each data holder can train the assigned local learning network based on local data. The locally trained local learning networks of multiple data holders can be sent to the embedding server for integration to obtain the trained federated learning network, realizing the distributed training of the federated learning network.
[0039] In the above step S102, the federated learning network can perform distributed learning on the graph neural network.
[0040] It should be noted that the learning network can describe the nodes and the connection relationships of the nodes in the learning network through the degree matrix and the adjacency matrix. Among them, the adjacency matrix mainly describes the connection relationships between the nodes, and the degree matrix mainly describes the degrees of the nodes, that is, the number of connections of each node to other nodes.
[0041] Optionally, the Laplacian matrix for describing the overall structure of the learning network can be obtained by combining the degree matrix and the adjacency matrix.
[0042] In the above step S104, the preset mask matrix is a pre-generated orthogonal matrix, and each data holder can use this preset mask matrix as the mask of the local network matrix to encrypt the local network matrix.
[0043] In the above-mentioned step S106, the embedding server can integrate the local network matrices uploaded by multiple data holders respectively to obtain the global network matrix of the federated learning network. Furthermore, since the local encrypted matrices uploaded by the data holders are encrypted using a preset mask matrix, the embedding server can directly integrate the local encrypted matrices without decrypting them, and can integrate the global network matrix that has been encrypted by the preset mask matrix, that is, the global aggregation matrix.
[0044] This application uses a random orthogonal matrix as the mask (preset mask matrix) to ensure that the original data of each participant (i.e., the data holder) will not be leaked during the federated network embedding process.
[0045] Optionally, the data holder generates a local encrypted matrix locally and collaborates with other parties through a preset mask matrix. The whole process is coordinated by a semi-honest central server (such as the embedding server). This mask mechanism not only protects the sensitive information of the participants (i.e., the data holders), but also ensures the accuracy of the calculation results.
[0046] It should be noted that network embedding is a technology that maps nodes or edges in graph-structured data (such as social networks, knowledge graphs, etc.) to a low-dimensional vector space, aiming to preserve the structure and semantic information of the graph for the convenience of processing and analysis in machine learning tasks.
[0047] In the above-mentioned step S106, network embedding needs to be based on the global network of the federated learning network. Therefore, the process of this network embedding needs to be executed by the embedding server. Furthermore, when the embedding server integrates the local network matrices uploaded by multiple data holders respectively into a global network matrix, it can decompose the global network matrix, and the result after decomposition is the global embedding matrix of the federated learning network.
[0048] It should be noted that since the data holders upload local encrypted matrices encrypted by a preset mask matrix, the embedding server cannot decrypt the local encrypted matrices. Therefore, the embedding server can integrate and decompose based on the encrypted local encrypted matrices, and the obtained global encrypted matrix is equivalent to the global embedding matrix encrypted by the preset mask matrix.
[0049] In the above step S108, since the global encryption matrix is determined based on the local encryption matrix, and the local encryption matrix is encrypted by the data holder using a preset mask matrix, the global encryption matrix is equivalent to the globally embedded matrix encrypted by the preset mask matrix. Moreover, the data holder holds the preset mask matrix. Therefore, after receiving the global encryption matrix, the data holder can decrypt the global encryption matrix based on the preset mask matrix to obtain the globally embedded matrix, achieving the acquisition of the network embedding result of the federated learning network and ensuring that the network embedding result cannot be stolen by objects other than the data holder.
[0050] As an optional embodiment, the local network matrix at least includes: the local adjacency matrix and the local degree matrix corresponding to the local learning network. Obtaining the local network matrices of multiple data holders in the federated learning network includes: counting the number of nodes in the federated learning network, where the number of nodes is the number of nodes with data processing capabilities; in the case where the number of nodes meets the first order of magnitude, performing graph structure analysis on the local learning network of each data holder to obtain the local adjacency matrix and the local degree matrix corresponding to the local learning network; in the case where the number of nodes meets the second order of magnitude, performing path sampling on the local learning network of each data holder to generate the local adjacency matrix and the local degree matrix corresponding to the local learning network, where the second order of magnitude is greater than the first order of magnitude.
[0051] In the above embodiments of the present application, the federated learning network can be divided into a small and medium-sized network and a large-scale network according to the number of nodes in the federated learning network. Among them, in a small and medium-sized network, the number of nodes is not large, and the connection relationship and the number of connections between nodes can be directly obtained through graph structure analysis. However, in a large-scale network, the number of nodes is too large, and directly using graph structure analysis, the analysis speed is slow and the efficiency is low. Therefore, different strategies can be used to obtain the local network matrices of each data holder for federated learning networks of different scales, thereby improving the acquisition efficiency of the local network matrices.
[0052] It should be noted that if the number of nodes in the federated learning network is large, the number of nodes in the local learning network segmented from the federated learning network is also large. Therefore, counting the number of nodes in the federated learning network can also be achieved by counting the number of nodes in the local learning network of the data holder.
[0053] Optionally, the first order of magnitude can be an order of magnitude set for small and medium-sized networks. In the case where the number of nodes meets the first order of magnitude, it indicates that the federated learning network is a small and medium-sized network. Furthermore, a strategy pre-configured for small and medium-sized networks can be used to obtain the local learning network of each data holder.
[0054] Optionally, the second order of magnitude can be an order of magnitude set for a large-scale network. When the number of nodes conforms to the second order of magnitude, it indicates that the federated learning network is a large-scale network. Furthermore, the policies pre-configured for the large-scale network can be used to obtain the local learning networks of each data holder locally.
[0055] As an alternative embodiment, when the number of nodes conforms to the second order of magnitude, performing path sampling on the local learning network of each data holder locally to generate the local adjacency matrix and the local degree matrix corresponding to the local learning network includes: when the number of nodes conforms to the second order of magnitude, performing random walks on the local learning network of each data holder locally to generate a set of paths, where the local learning network includes multiple nodes, the set of paths includes multiple path information generated by performing random walks respectively based on multiple preset starting points, the preset starting points are nodes determined in advance as starting points, and each path information is a sequence of multiple nodes; determining the local adjacency matrix and the local degree matrix corresponding to the local learning network according to the set of paths, where the elements in the local adjacency matrix are determined according to the adjacent situation of any two nodes in the multiple paths, and the elements in the local degree matrix are determined according to the number of different nodes adjacent to the node in the multiple path information.
[0056] In the above embodiments of the present application, when the number of nodes conforms to the second order of magnitude, that is, when the federated learning network is a large-scale network, it is possible to perform random walks on the local learning network of the data holder locally based on the path sampling method, generate multiple path information based on the nodes in the local learning network, obtain a set of paths, and then determine the connection relationship between the nodes in the local learning network and the number of connections of each node with other nodes according to the path information in the set of paths. Therefore, according to the multiple path information in the set of paths, the local adjacency matrix and the local degree matrix corresponding to the local learning network can be determined, and thus the local network matrix of each local learning network of the federated learning network belonging to the large-scale network can be obtained by means of path sampling.
[0057] It should be noted that path sampling is a key technology for generating sparse matrices, which effectively reduces the computational complexity of large-scale networks. The path sampling algorithm generates a path sampling matrix (i.e., the matrix representation of multiple path information) by performing random walks on the local network, approximating the global structural relationship of the network. Each data holder independently performs path sampling and randomly samples a certain number of paths for the nodes of the local network (such as the local learning network). By converting the paths into a sparse approximation of the adjacency matrix, the data holder generates a sparse network. The Laplacian matrix calculated based on the sparse network It is further used for the approximate calculation of polynomial matrices. The design of path sampling combines randomness and sparsity, which can effectively reduce memory and computational overhead and is a basic module for realizing large-scale network embedding.
[0058] As an alternative example, for federated learning networks or local learning networks of different scales, different methods can be used for network embedding. For example, when the federated learning network or local learning network is a small or medium-sized network, that is, when the number of nodes in the federated learning network conforms to the first order of magnitude, the federated network matrix factorization algorithm (FedNetMF) can be used; when the federated learning network or local learning network is a large-scale network, that is, when the number of nodes in the federated learning network conforms to the second order of magnitude, the federated sparse matrix factorization algorithm (FedNetSMF) can be used.
[0059] As an alternative embodiment, the local network matrix at least includes: the local adjacency matrix and the local degree matrix corresponding to the local learning network. The local adjacency matrix is encrypted using a preset mask matrix to obtain a local encrypted matrix, including: when the number of nodes in the federated learning network conforms to the first order of magnitude, each data holder combines the local adjacency matrix and the local degree matrix to obtain a combined polynomial matrix; receives the preset mask matrix pre-generated by the mask server, where the mask server pre-generates a random orthogonal matrix as the preset mask matrix; combines the random orthogonal matrix with the combined polynomial matrix to obtain the local encrypted matrix of each data holder.
[0060] In the above embodiments of the present application, when the number of nodes in the federated learning network conforms to the first order of magnitude, that is, when the federated learning network is a small or medium-sized network, the local network matrix can be encrypted locally at the data holder using a preset mask matrix. The encryption process can be to combine the local adjacency matrix and the local degree matrix to obtain a combined polynomial matrix, and then combine the combined polynomial matrix using the preset mask matrix pre-generated by the mask server, so as to obtain the local encrypted matrix of each data holder and realize the encryption of the local network matrix.
[0061] As an alternative example, the federated network matrix factorization algorithm (FedNetMF) is a centralized NetMF algorithm designed specifically for small and medium-sized networks and can achieve lossless embedding results. The federated network matrix factorization algorithm calculates the polynomial matrix of the network (such as the combined polynomial matrix) through a distributed collaboration method, ensuring privacy protection while maintaining embedding accuracy.
[0062] In the above embodiments of the present application, the Federated Network Matrix Factorization algorithm (FedNetMF) implements matrix factorization of the network through a multi-party joint computing method; between data parties, a method combining distributed polynomial matrix calculation and a preset mask matrix is used to ensure data privacy while achieving a lossless embedding result, which is applicable to networks with a medium node scale (for example, several thousand to tens of thousands of nodes).
[0063] Optionally, each data holder calculates the degree matrix of its local network (such as the local network matrix) (i.e., the local degree matrix) and the normalized adjacency matrix (i.e., the local adjacency matrix); subsequently, the data holder encrypts the local matrix through an encrypted mask matrix to generate a polynomial matrix (i.e., the combined polynomial matrix), and sends the encrypted to the embedding server; the embedding server aggregates the encrypted results of all parties, performs singular value decomposition, and calculates the encrypted global embedding matrix (i.e., the global encrypted matrix). The global encrypted matrix is then distributed back to the data holders, and each holder removes the mask to obtain the final embedding result (i.e., the global embedding matrix).
[0064] The above Federated Network Matrix Factorization algorithm of the present application is applicable to networks with a small node scale and can provide embedding quality comparable to that of centralized NetMF under the premise of privacy protection.
[0065] As an alternative embodiment, the local network matrix at least includes: the local adjacency matrix and the local degree matrix corresponding to the local learning network, and encrypts the local adjacency matrix using a preset mask matrix to obtain a local encrypted matrix, including: when the number of nodes in the federated learning network meets the second order of magnitude, each data holder uploads the local degree matrix to the embedding server respectively, and receives the global degree matrix returned by the embedding server based on multiple local degree matrices; according to the local degree matrix and the Laplacian matrix of the same data holder, and the global degree matrix, determine the random walk polynomial matrix, where the Laplacian matrix is determined according to the difference between the local adjacency matrix and the local degree matrix of the same data holder, and the random walk polynomial matrix is expressed as: , where i is the data holder, is the random walk polynomial matrix of the data holder, D is the global degree matrix, is the local degree matrix of the data holder, is the Laplacian matrix of the data holder; perform federated power iteration on the random walk polynomial matrix according to a preset number of iterations to generate a random walk sub-matrix, where the federated power iteration is used to indicate projecting the random walk polynomial matrix into a low-dimensional space; encrypt the random walk sub-matrix using a preset mask matrix to obtain a local encrypted matrix.
[0066] In the above embodiments of the present application, when the number of nodes in the federated learning network conforms to the second order of magnitude, that is, when the federated learning network is a large-scale network, the local adjacency matrix and the local degree matrix in the local network matrix are obtained by random walk based on the path sampling method. Then, the data holder and the embedding server can perform a preset number of iterations of federated power iteration, and in the case of federated power iteration, use a preset mask matrix for encryption to obtain a local encrypted matrix.
[0067] Figure 2 is a schematic diagram of a federated sparse matrix factorization algorithm (FedNetSMF) according to an embodiment of the present invention, as Figure 2 shown, the federated sparse matrix factorization algorithm (FedNetSMF) is an efficient embedding method designed for large-scale networks. It generates a sparse matrix through path sampling technology, effectively reducing the computational and storage complexity.
[0068] Table 1 is a schematic table of a federated sparse matrix factorization algorithm (FedNetSMF) according to an embodiment of the present invention. As shown in Table 1, it explains various descriptions in the federated sparse matrix factorization algorithm (FedNetSMF).
[0069] Table 1
[0070]
[0071] In the above embodiments of the present application, the federated sparse matrix factorization algorithm (FedNetSMF) generates a sparse matrix through the path sampling (PathSampling) method, reducing the computational and storage overhead. At the same time, it ensures that the consistency of the embedding result with the centralized method is within an acceptable range (about 1% accuracy loss), can handle networks with a scale of millions of nodes, and is applicable to scenarios with multiple data participants (i.e., multiple data holders).
[0072] Optionally, each data holder performs path sampling on its local network (i.e., the local learning network) to generate the Laplacian matrix of the sparse network , and calculates the polynomial matrix (such as the random walk polynomial matrix); subsequently, the data holder projects into a low-dimensional space through the federated power iteration method to generate a submatrix ; the data holder sends the encrypted to the embedding server, and the embedding server aggregates all submatrices and performs singular value decomposition to generate the final embedding matrix .
[0073] In the above embodiments of the present application, the federated sparse matrix decomposition algorithm can effectively run when the node scale reaches the million level, and the embedding quality only has a loss of about 1%, which is particularly suitable for large-scale distributed network scenarios with limited resources.
[0074] Optionally, for the random walk polynomial matrix Performing federated power iteration according to a preset number of iterations can obtain the first matrix Q M ; After returning the matrix Q M to the data holder, the data holder can obtain the sub-matrix obtained by projecting the random walk polynomial matrix onto the low-dimensional space ; Then, for the sub-matrix performing federated power iteration according to a preset number of iterations can obtain the second matrix Q B , and further, after returning the second matrix Q B to the data holder, based on the second matrix Q B and the sub-matrix generate a random walk sub-matrix C i , encrypt the random walk sub-matrix C using the preset mask matrix to obtain a local encrypted matrix, and then send the local encrypted matrix to the embedding server for integration and decomposition to obtain a global encrypted matrix.
[0075] Optionally, the preset number of iterations can be 3 times.
[0076] As an alternative embodiment, the steps of each federated power iteration include: each data holder encrypts the random walk polynomial matrix using the preset mask matrix to obtain a local pre-encrypted matrix; uploads the local pre-encrypted matrix to the embedding server, where the embedding server is used to aggregate multiple local pre-encrypted matrices to obtain a global pre-encrypted matrix, and perform coefficient matrix decomposition on the global pre-encrypted matrix to obtain an intermediate result matrix, and the intermediate result matrix is a masked orthogonal matrix; receive the intermediate result matrix returned by the embedding server, and remove the mask in the intermediate result matrix according to the preset mask matrix to obtain an intermediate encrypted matrix, where when the number of iterations reaches the preset number of iterations, the intermediate encrypted matrix is a random walk sub-matrix; when the number of iterations does not reach the preset number of iterations, the intermediate encrypted matrix is the random walk polynomial matrix for the next iteration.
[0077] In the above embodiments of the present application, federated power iteration (Federated Power Iteration, FPI) is a multi-party collaboration algorithm used to calculate the orthogonal decomposition of a global matrix, and is particularly suitable for large-scale sparse matrices. The core of this federated power iteration algorithm is to gradually approximate the first principal eigenvectors of the matrix through multiple rounds of power iteration.
[0078] In the above embodiments of the present application, in sparse matrix factorization, the federated power iteration algorithm is used to implement the orthogonal factorization of the global matrix. Each data holder collaborates through a random mask matrix to complete multiple power iteration steps, avoiding data leakage and improving the factorization efficiency of large-scale sparse matrices.
[0079] Optionally, a random orthogonal mask matrix and (i.e., the preset mask matrix) is generated by the mask server and distributed to all data holders; each data holder calculates the masked version of its local matrix (such as the local network matrix), such as the local pre-encrypted matrix , and sends it to the embedding server; the embedding server aggregates all the local pre-encrypted matrices to generate an intermediate result matrix, and then sends it back to the data holders for the next round of iteration. After multiple rounds of iteration, the embedding server obtains the masked version of the global orthogonal matrix (such as the global encrypted matrix), and the data holders obtain the final orthogonal matrix (i.e., the global embedding matrix) after removing the mask. The design of federated power iteration ensures privacy protection throughout the process while significantly improving the efficiency of sparse matrix factorization.
[0080] As an alternative embodiment, the method further includes: aggregating the local encrypted matrices uploaded by multiple data holders respectively through the embedding server to obtain a global aggregation matrix; performing singular value decomposition on the global aggregation matrix to generate a global encrypted matrix.
[0081] In the above embodiments of the present application, the embedding server can aggregate the local encrypted matrices uploaded by multiple data holders respectively to obtain a global aggregation matrix, and then by performing singular value decomposition on the global aggregation matrix, the network embedding result of the federated learning network can be determined, generating a global encrypted matrix, realizing the generation of an encrypted global embedding matrix.
[0082] As an alternative embodiment, during the entire embedding process, a secure computing protocol that conforms to the MPC privacy definition is adopted. The semi-honest central server is responsible for integrating the encrypted matrices from all parties, and this semi-honest central server cannot recover the original data. Through theoretical proof, the data security of the entire process can be guaranteed.
[0083] It should be noted that MPC stands for Multi-Party Computation, which is multi-party secure computation.
[0084] The present invention also provides a preferred embodiment, which provides an unsupervised federated network embedding algorithm framework based on privacy protection, aiming to solve the problem of secure network embedding in the distributed subgraph scenario. By specifically designing FedNetMF (Federated Network Matrix Factorization) and FedNetSMF (Federated Sparse Matrix Factorization) for the distributed subgraph scenario, it is possible to balance embedding quality, computational efficiency, and privacy protection, thereby enabling efficient and scalable network embedding while protecting data privacy.
[0085] Optionally, this application uses the Federated Network Matrix Factorization (FedNetMF) and the Federated Sparse Matrix Factorization (FedNetSMF) to meet the embedding requirements of medium - small - scale and large - scale networks respectively.
[0086] Optionally, FedNetMF is applicable to medium - small - scale networks and can achieve lossless embedding results with the same accuracy as the centralized method; FedNetSMF is for large - scale networks and uses sparse matrix factorization technology to balance computational efficiency and embedding quality, adapting to the federated scenarios of more participants.
[0087] As an optional example, the unsupervised federated network embedding algorithm framework based on privacy protection is an unsupervised federated network embedding (Federated Network Embedding, FNE) method that supports the distributed subgraph scenario, ensuring the privacy of the participants' data and avoiding the leakage of sensitive information in cross - party communication; while ensuring embedding quality, it should have good computational and communication efficiency, be able to handle large - scale networks; support collaborative embedding calculations of multiple participants, and adapt to the growing network scale.
[0088] Figure 3 It is a schematic diagram of an unsupervised federated network embedding method based on the embodiments of the present invention. As Figure 3 shown, each data holder collaborates to complete the embedding task. At the same time, through the encrypted mask matrix and Federated Power Iteration (FPI), the local adjacency matrix determined based on the local network matrix of each data holder is encrypted, which can protect data privacy and enable the embedding server to only be responsible for integrating the encrypted intermediate results, ensuring end - to - end privacy protection.
[0089] Table 2 is a schematic table of an unsupervised federated network embedding method based on the embodiments of the present invention. As shown in Table 2, it explains the various descriptions of the unsupervised federated network embedding method.
[0090] Table 2
[0091]
[0092] This application uses an encrypted mask matrix mechanism to ensure that data privacy is not leaked during the multi-party collaboration process. Through theoretical analysis and experimental verification, the embedding results are close to those of the centralized method in terms of accuracy.
[0093] This application uses an encrypted mask matrix and federated power iteration technology to ensure that all intermediate calculation results during the embedding process are in encrypted form, so that the local data of the participating parties (i.e., data holders) will not be leaked, and the embedding server cannot recover the original data. This privacy protection mechanism conforms to the privacy definition of multi-party secure computing (MPC) and can meet the strict data privacy requirements in distributed scenarios. Compared with traditional centralized methods, it ensures the data independence and security of all parties while realizing data collaboration.
[0094] The technical solution provided by this application designs a federated sparse matrix factorization algorithm (FedNetSMF) for large-scale network scenarios, which significantly reduces the computational and storage complexity by combining path sampling technology. Sparse matrix factorization can handle networks with node scales reaching millions, and the computational efficiency is improved by several orders of magnitude compared with traditional methods. At the same time, the application of federated power iteration further optimizes the eigenvector calculation process, enabling large-scale network embedding to be efficiently completed with limited computing resources.
[0095] The technical solution provided by this application achieves approximate consistency in accuracy between the embedding results and the centralized method while maintaining privacy protection and computational efficiency. For small and medium-sized networks, the federated network matrix factorization algorithm (FedNetMF) can achieve lossless embedding quality. For large-scale networks, the sparse matrix factorization has only about 1% accuracy loss and shows excellent macro F1 scores in experiments. This makes the embedding results of this invention applicable to various downstream tasks (such as node classification, link prediction, etc.).
[0096] The technical solution provided by this application adopts a modular design of the framework, supports the application of multiple distributed scenarios, can be adapted to multiple participating parties (i.e., data holders), and can be extended to other network embedding methods based on matrix factorization. At the same time, it has a wide range of applications, including multiple fields such as financial fraud detection, social network analysis, and medical and health data analysis. It has good scalability and applicability in both the collaboration of small-scale participating parties and the distributed computing scenarios of large-scale data centers.
[0097] Compared with existing strategies, the technical solution provided by this application finally solves the problem of secure network embedding in distributed subgraph scenarios. Aiming at the problem that existing methods cannot handle the unsupervised network embedding problem in subgraph scenarios, this application fills the technical gap in this field.
[0098] According to an embodiment of the present invention, an embodiment of a network embedding device based on a federated learning network is further provided. It should be noted that the network embedding device based on the federated learning network can be used to execute the network embedding method based on the federated learning network in the embodiment of the present invention, and the network embedding method based on the federated learning network in the embodiment of the present invention can be executed in the network embedding device based on the federated learning network.
[0099] Figure 4 is a schematic diagram of a network embedding device based on a federated learning network according to an embodiment of the present invention. As Figure 4 shown, the device may include: an acquisition module 42, configured to acquire local network matrices of multiple data holders in the federated learning network, where the data holders are participants in the federated learning network, the federated learning network is pre-divided into multiple local learning networks, and the local network matrix is a matrix representation of the local learning network; an encryption module 44, configured to encrypt the local network matrix using a preset mask matrix to obtain a local encrypted matrix; an upload module 46, configured to upload the local encrypted matrices respectively corresponding to the multiple data holders to an embedding server, and receive a global encrypted matrix integrated by the embedding server based on the multiple local encrypted matrices, where the global encrypted matrix is the result of decomposing the integration result of the multiple local encrypted matrices by the embedding server; a decryption module 48, configured to decrypt the global encrypted matrix based on the preset mask matrix to obtain a global embedding matrix, where the global embedding matrix is the network embedding result of the federated learning network.
[0100] It should be noted that the acquisition module 42 in this embodiment can be used to execute step S102 in the embodiment of the present application, the encryption module 44 in this embodiment can be used to execute step S104 in the embodiment of the present application, the upload module 46 in this embodiment can be used to execute step S106 in the embodiment of the present application, and the decryption module 48 in this embodiment can be used to execute step S108 in the embodiment of the present application. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiments.
[0101] In the embodiment of the present invention, each data holder in the federated learning network uploads a local encrypted matrix encrypted by a preset mask matrix to the embedding server, ensuring that all intermediate calculation results during the embedding process are in encrypted form, ensuring that the local data of the data holder will not be leaked, and the embedding server cannot recover the original data of the data holder, achieving the purpose of ensuring data security, thereby realizing the technical effect of ensuring data security for network embedding in a distributed environment, and further solving the technical problem of poor data security for network embedding in the prior art in a distributed environment.
[0102] As an alternative embodiment, the local network matrix at least includes: the local adjacency matrix and the local degree matrix corresponding to the local learning network. The obtaining module includes: a statistical unit for counting the number of nodes in the federated learning network, where the number of nodes is the number of nodes with data processing capabilities; a first generating unit for performing graph structure analysis on the local learning network of each data holder to obtain the local adjacency matrix and the local degree matrix corresponding to the local learning network when the number of nodes meets the first order of magnitude; a second generating unit for performing path sampling on the local learning network of each data holder to generate the local adjacency matrix and the local degree matrix corresponding to the local learning network when the number of nodes meets the second order of magnitude, where the second order of magnitude is greater than the first order of magnitude.
[0103] As an alternative embodiment, when the number of nodes meets the second order of magnitude, the second generating unit includes: a generating subunit for performing random walks on the local learning network of each data holder to generate a path set when the number of nodes meets the second order of magnitude, where the local learning network includes multiple nodes, the path set includes multiple path information respectively generated based on multiple preset starting points, the preset starting point is a node determined in advance as the starting point, and each path information is a sequence of multiple nodes; a determining subunit for determining the local adjacency matrix and the local degree matrix corresponding to the local learning network according to the path set, where the elements in the local adjacency matrix are determined according to the adjacent situation of any two nodes in the multiple paths, and the elements in the local degree matrix are determined according to the number of different adjacent nodes of the node in the multiple path information.
[0104] As an alternative embodiment, the local network matrix at least includes: the local adjacency matrix and the local degree matrix corresponding to the local learning network. The encryption module includes: a combining unit for combining the local adjacency matrix and the local degree matrix by each data holder to obtain a combined polynomial matrix when the number of nodes in the federated learning network meets the first order of magnitude; a receiving unit for receiving a preset mask matrix pre-generated by the mask server, where the mask server pre-generates a random orthogonal matrix as the preset mask matrix; a first encryption unit for combining the random orthogonal matrix with the combined polynomial matrix to obtain the local encryption matrix of each data holder.
[0105] As an alternative embodiment, the local network matrix at least includes: a local adjacency matrix and a local degree matrix corresponding to the local learning network. The encryption module includes: an upload unit, configured to, when the number of nodes in the federated learning network meets the second order of magnitude, each data holder uploads the local degree matrix to the embedding server respectively, and receives the global degree matrix returned by the embedding server based on multiple local degree matrices; a determination unit, configured to determine a random walk polynomial matrix according to the local degree matrix and the Laplacian matrix of the same data holder, and the global degree matrix, where the Laplacian matrix is determined according to the difference between the local adjacency matrix and the local degree matrix of the same data holder, and the random walk polynomial matrix is expressed as: , where i is the data holder, is the random walk polynomial matrix of the data holder, D is the global degree matrix, is the local degree matrix of the data holder, is the Laplacian matrix of the data holder; an iteration unit, configured to perform federated power iteration on the random walk polynomial matrix according to a preset number of iterations to generate a random walk sub-matrix, where the federated power iteration is used to indicate projecting the random walk polynomial matrix into a low-dimensional space; a second encryption unit, configured to encrypt the random walk sub-matrix using a preset mask matrix to obtain a local encrypted matrix.
[0106] As an alternative embodiment, the steps of each federated power iteration include: each data holder encrypts the random walk polynomial matrix using a preset mask matrix to obtain a local pre-encrypted matrix; uploads the local pre-encrypted matrix to the embedding server, where the embedding server is configured to aggregate multiple local pre-encrypted matrices to obtain a global pre-encrypted matrix, and perform coefficient matrix decomposition on the global pre-encrypted matrix to obtain an intermediate result matrix, and the intermediate result matrix is an orthogonal matrix after masking; receives the intermediate result matrix returned by the embedding server, and removes the mask in the intermediate result matrix according to the preset mask matrix to obtain an intermediate encrypted matrix, where when the number of iterations reaches the preset number of iterations, the intermediate encrypted matrix is the random walk sub-matrix; when the number of iterations does not reach the preset number of iterations, the intermediate encrypted matrix is the random walk polynomial matrix for the next iteration.
[0107] As an alternative embodiment, the apparatus further includes: an aggregation sub-module, configured to aggregate local encrypted matrices respectively uploaded by multiple data holders through the embedding server to obtain a global aggregation matrix; a decomposition sub-module, configured to perform singular value decomposition on the global aggregation matrix to generate a global encrypted matrix.
[0108] Embodiments of the present invention may provide an electronic device, which may be a computer terminal, and the computer terminal may be any one of a group of computer terminal devices. Optionally, in this embodiment, the above computer terminal may also be replaced with a terminal device such as a mobile terminal.
[0109] Optionally, in this embodiment, the above computer terminal may be located in at least one of a plurality of network devices in a computer network.
[0110] In this embodiment, the above computer terminal may execute program code for the following steps in a network embedding method based on a federated learning network: obtaining local network matrices of multiple data holders in the federated learning network, where the data holders are participants in the federated learning network, the federated learning network is pre-divided into multiple local learning networks, and the local network matrix is a matrix representation of the local learning network; encrypting the local network matrix using a preset mask matrix to obtain a local encrypted matrix; uploading the local encrypted matrices respectively corresponding to the multiple data holders to an embedding server, and receiving a global encrypted matrix integrated by the embedding server based on the multiple local encrypted matrices, where the global encrypted matrix is the result of decomposing the integration result based on the multiple local encrypted matrices by the embedding server; decrypting the global encrypted matrix based on the preset mask matrix to obtain a global embedding matrix, where the global embedding matrix is the network embedding result of the federated learning network.
[0111] Figure 5 is a structural block diagram of a computer terminal according to an embodiment of the present invention, as Figure 5 shown, the computer terminal 50 may include: one or more (only one is shown in the figure) processors 52 and a memory 54.
[0112] Among them, the memory may be used to store software programs and modules, such as program instructions / modules corresponding to the network embedding method and device based on the federated learning network in the embodiments of the present invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the above-mentioned network embedding method based on the federated learning network. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely disposed relative to the processor, and these remote memories may be connected to the terminal 50 through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0113] The processor can call the information and application programs stored in the memory through the transmission device to execute the following steps: Obtain the local network matrices of multiple data holders in the federated learning network. Here, the data holders are participants in the federated learning network, and the federated learning network is pre-divided into multiple local learning networks. The local network matrix is the matrix representation of the local learning network; Use a preset mask matrix to encrypt the local network matrix to obtain a local encrypted matrix; Upload the local encrypted matrices corresponding to multiple data holders to the embedding server respectively, and receive the global encrypted matrix integrated by the embedding server based on multiple local encrypted matrices. Here, the global encrypted matrix is the result after the embedding server decomposes the integration result based on multiple local encrypted matrices; Decrypt the global encrypted matrix based on the preset mask matrix to obtain the global embedding matrix. Here, the global embedding matrix is the network embedding result of the federated learning network.
[0114] Optionally, the local network matrix at least includes: the local adjacency matrix and the local degree matrix corresponding to the local learning network. The above processor can also execute the program code of the following steps: Count the number of nodes in the federated learning network, where the number of nodes is the number of nodes with data processing capabilities; In the case where the number of nodes conforms to the first order of magnitude, perform graph structure analysis on the local learning network of each data holder locally to obtain the local adjacency matrix and the local degree matrix corresponding to the local learning network; In the case where the number of nodes conforms to the second order of magnitude, perform path sampling on the local learning network of each data holder locally to generate the local adjacency matrix and the local degree matrix corresponding to the local learning network, where the second order of magnitude is greater than the first order of magnitude.
[0115] Optionally, the above processor can also execute the program code of the following steps: In the case where the number of nodes conforms to the second order of magnitude, perform random walks on the local learning network of each data holder locally to generate a path set. Here, the local learning network includes multiple nodes, and the path set includes multiple path information generated by performing random walks based on multiple preset starting points respectively. The preset starting point is a node determined in advance as the starting point, and each path information is a sequence of multiple nodes; According to the path set, determine the local adjacency matrix and the local degree matrix corresponding to the local learning network. Here, the elements in the local adjacency matrix are determined according to the adjacent situation of any two nodes in the multiple paths, and the elements in the local degree matrix are determined according to the number of different nodes adjacent to the node in the multiple path information.
[0116] Optionally, the local network matrix at least includes: the local adjacency matrix and the local degree matrix corresponding to the local learning network. The above-mentioned processor can also execute the program code of the following steps: when the number of nodes in the federated learning network meets the first order of magnitude, each data holder combines the local adjacency matrix and the local degree matrix to obtain a combined polynomial matrix; receives the preset mask matrix pre-generated by the mask server, where the mask server pre-generates a random orthogonal matrix as the preset mask matrix; combines the random orthogonal matrix with the combined polynomial matrix to obtain the local encryption matrix of each data holder.
[0117] Optionally, the local network matrix at least includes: the local adjacency matrix and the local degree matrix corresponding to the local learning network. The above-mentioned processor can also execute the program code of the following steps: when the number of nodes in the federated learning network meets the second order of magnitude, each data holder uploads the local degree matrix to the embedding server respectively and receives the global degree matrix returned by the embedding server based on multiple local degree matrices; determines the random walk polynomial matrix according to the local degree matrix and the Laplacian matrix of the same data holder, and the global degree matrix, where the Laplacian matrix is determined according to the difference between the local adjacency matrix and the local degree matrix of the same data holder. The random walk polynomial matrix is expressed as: , where i is the data holder, is the random walk polynomial matrix of the data holder, D is the global degree matrix, is the local degree matrix of the data holder, is the Laplacian matrix of the data holder; performs federated power iteration on the random walk polynomial matrix according to the preset number of iterations to generate a random walk submatrix, where the federated power iteration is used to indicate projecting the random walk polynomial matrix into a low-dimensional space; encrypts the random walk submatrix using the preset mask matrix to obtain the local encryption matrix.
[0118] Optionally, for each step of the federated power iteration, the above-mentioned processor can also execute the program code of the following steps: each data holder encrypts the random walk polynomial matrix using the preset mask matrix to obtain a local pre-encrypted matrix; uploads the local pre-encrypted matrix to the embedding server, where the embedding server is used to aggregate multiple local pre-encrypted matrices to obtain a global pre-encrypted matrix, and performs coefficient matrix decomposition on the global pre-encrypted matrix to obtain an intermediate result matrix, and the intermediate result matrix is a masked orthogonal matrix; receives the intermediate result matrix returned by the embedding server and removes the mask in the intermediate result matrix according to the preset mask matrix to obtain an intermediate encrypted matrix, where when the number of iterations reaches the preset number of iterations, the intermediate encrypted matrix is the random walk submatrix; when the number of iterations does not reach the preset number of iterations, the intermediate encrypted matrix is the random walk polynomial matrix for the next iteration.
[0119] Optionally, the above-mentioned processor may also execute program code of the following steps: aggregating local encryption matrices respectively uploaded by multiple data holders through an embedding server to obtain a global aggregation matrix; performing singular value decomposition on the global aggregation matrix to generate a global encryption matrix.
[0120] Those of ordinary skill in the art can understand that Figure 5 the structure shown is only illustrative, and the computer terminal can also be a smart phone (such as , tablet computer, palm computer, and terminal devices such as Mobile Internet Devices (MID), PAD, etc. Figure 5 It does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 50 may further include more or fewer components (such as network interfaces, display devices, etc.) than those shown in Figure 5 , or have a different configuration from that shown in Figure 5 .
[0121] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a computer program, and the computer program can be stored in a non-volatile medium. The non-volatile storage medium may include: flash drive, Read-Only Memory (ROM), Random Access Memory (RAM), magnetic disk or optical disc, etc.
[0122] An embodiment of the present invention also provides a non-volatile storage medium. Optionally, in this embodiment, the above non-volatile storage medium can be used to store the program code executed by the network embedding method based on the federated learning network provided in the above embodiment.
[0123] Optionally, in this embodiment, the above non-volatile storage medium may be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.
[0124] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: obtaining local network matrices of multiple data holders in the federated learning network, where the data holders are participants in the federated learning network, the federated learning network is pre-divided into multiple local learning networks, and the local network matrix is a matrix representation of the local learning network; encrypting the local network matrix using a preset mask matrix to obtain a local encrypted matrix; uploading the local encrypted matrices corresponding to the multiple data holders to the embedding server, and receiving the global encrypted matrix integrated by the embedding server based on the multiple local encrypted matrices, where the global encrypted matrix is the result of decomposing the integration result based on the multiple local encrypted matrices by the embedding server; decrypting the global encrypted matrix based on the preset mask matrix to obtain a global embedding matrix, where the global embedding matrix is the network embedding result of the federated learning network.
[0125] Optionally, in this embodiment, the local network matrix at least includes: a local adjacency matrix and a local degree matrix corresponding to the local learning network. The non-volatile storage medium is configured to store program code for performing the following steps: counting the number of nodes in the federated learning network, where the number of nodes is the number of nodes with data processing capabilities; in the case where the number of nodes meets the first order of magnitude, performing graph structure analysis on the local learning network of each data holder to obtain the local adjacency matrix and the local degree matrix corresponding to the local learning network; in the case where the number of nodes meets the second order of magnitude, performing path sampling on the local learning network of each data holder to generate the local adjacency matrix and the local degree matrix corresponding to the local learning network, where the second order of magnitude is greater than the first order of magnitude.
[0126] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: in the case where the number of nodes meets the second order of magnitude, performing random walks on the local learning network of each data holder to generate a path set, where the local learning network includes multiple nodes, the path set includes multiple path information generated by performing random walks based on multiple preset starting points respectively, the preset starting point is a node determined in advance as the starting point, and each path information is a sequence of multiple nodes; determining the local adjacency matrix and the local degree matrix corresponding to the local learning network according to the path set, where the elements in the local adjacency matrix are determined according to the adjacent situation of any two nodes in the multiple paths, and the elements in the local degree matrix are determined according to the number of different nodes adjacent to the node in the multiple path information.
[0127] Optionally, in this embodiment, the local network matrix at least includes: the local adjacency matrix and the local degree matrix corresponding to the local learning network. The non-volatile storage medium is configured to store program code for performing the following steps: when the number of nodes in the federated learning network meets the first order of magnitude, each data holder combines the local adjacency matrix and the local degree matrix to obtain a combined polynomial matrix; receives a preset mask matrix pre-generated by the mask server, where the mask server pre-generates a random orthogonal matrix as the preset mask matrix; combines the random orthogonal matrix with the combined polynomial matrix to obtain the local encryption matrix of each data holder.
[0128] Optionally, in this embodiment, the local network matrix at least includes: the local adjacency matrix and the local degree matrix corresponding to the local learning network. The non-volatile storage medium is configured to store program code for performing the following steps: when the number of nodes in the federated learning network meets the second order of magnitude, each data holder uploads the local degree matrix to the embedding server respectively and receives the global degree matrix returned by the embedding server based on multiple local degree matrices; determines a random walk polynomial matrix according to the local degree matrix and the Laplacian matrix of the same data holder, and the global degree matrix, where the Laplacian matrix is determined according to the difference between the local adjacency matrix and the local degree matrix of the same data holder. The random walk polynomial matrix is expressed as: , where i is the data holder, is the random walk polynomial matrix of the data holder, D is the global degree matrix, is the local degree matrix of the data holder, is the Laplacian matrix of the data holder; performs federated power iteration on the random walk polynomial matrix according to a preset number of iterations to generate a random walk submatrix, where the federated power iteration is used to indicate projecting the random walk polynomial matrix into a low-dimensional space; encrypts the random walk submatrix using a preset mask matrix to obtain a local encryption matrix.
[0129] Optionally, in this embodiment, for each step of the federated power iteration, the non-volatile storage medium is set to store program code for performing the following steps: Each data holder encrypts the random walk polynomial matrix using a preset mask matrix to obtain a local pre-encrypted matrix; uploads the local pre-encrypted matrix to the embedding server, where the embedding server is used to aggregate multiple local pre-encrypted matrices to obtain a global pre-encrypted matrix, and performs coefficient matrix decomposition on the global pre-encrypted matrix to obtain an intermediate result matrix, where the intermediate result matrix is a masked orthogonal matrix; receives the intermediate result matrix returned by the embedding server, and removes the mask in the intermediate result matrix according to the preset mask matrix to obtain an intermediate encrypted matrix, where, when the number of iterations reaches the preset number of iterations, the intermediate encrypted matrix is a random walk sub-matrix; when the number of iterations does not reach the preset number of iterations, the intermediate encrypted matrix is the random walk polynomial matrix for the next iteration.
[0130] Optionally, in this embodiment, the non-volatile storage medium is set to store program code for performing the following steps: Aggregate local encrypted matrices uploaded separately by multiple data holders through the embedding server to obtain a global aggregation matrix; perform singular value decomposition on the global aggregation matrix to generate a global encrypted matrix.
[0131] An embodiment of the present invention also provides a computer program product, including a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, it implements the steps of the network embedding method based on a federated learning network provided in the above embodiment.
[0132] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.
[0133] In the above embodiments of the present invention, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0134] In several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0135] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0136] In addition, in each embodiment of the present invention, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0137] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a non-volatile storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned non-volatile storage medium includes: USB flash drives, read-only memory (ROM), random access memory (RAM), mobile hard disks, magnetic disks, or optical discs, and other various media that can store program codes.
[0138] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A network embedding method based on a federated learning network, characterized in that Including: Obtain the local network matrices of multiple data holders in the federated learning network. Herein, the data holders are participants in the federated learning network, the federated learning network is pre-divided into multiple local learning networks, and the local network matrix is the matrix representation of the local learning network; Use a preset mask matrix to encrypt the local network matrix to obtain a local encrypted matrix; Upload the local encrypted matrices respectively corresponding to multiple data holders to an embedding server, and receive the global encrypted matrix integrated by the embedding server based on multiple local encrypted matrices. Herein, the global encrypted matrix is the result after the embedding server decomposes the integration result based on multiple local encrypted matrices; Decrypt the global encrypted matrix based on the preset mask matrix to obtain a global embedding matrix. Herein, the global embedding matrix is the network embedding result of the federated learning network.
2. The method according to claim 1, characterized in that, The local network matrix at least includes: the local adjacency matrix and the local degree matrix corresponding to the local learning network. Obtaining the local network matrices of multiple data holders in the federated learning network includes: Count the number of nodes in the federated learning network. Herein, the number of nodes is the number of nodes with data processing capabilities; When the number of nodes meets the first order of magnitude, perform graph structure analysis on the local learning network of each data holder locally to obtain the local adjacency matrix and the local degree matrix corresponding to the local learning network; When the number of nodes meets the second order of magnitude, perform path sampling on the local learning network of each data holder locally to generate the local adjacency matrix and the local degree matrix corresponding to the local learning network. Herein, the second order of magnitude is greater than the first order of magnitude.
3. The method according to claim 2, wherein When the number of nodes meets the second order of magnitude, performing path sampling on the local learning network of each data holder locally to generate the local adjacency matrix and the local degree matrix corresponding to the local learning network includes: When the number of nodes meets the second order of magnitude, perform random walks on the local learning network of each data holder locally to generate a path set. Herein, the local learning network includes multiple nodes, the path set includes multiple path information respectively generated by performing random walks based on multiple preset starting points, the preset starting points are the nodes determined in advance as starting points, and each path information is a sequence of multiple nodes; Determine the local adjacency matrix and the local degree matrix corresponding to the local learning network according to the path set. Herein, the elements in the local adjacency matrix are determined according to the adjacent situation of any two nodes in multiple paths, and the elements in the local degree matrix are determined according to the number of different nodes adjacent to the node in multiple path information.
4. The method according to claim 1, wherein The local network matrix at least includes: the local adjacency matrix and the local degree matrix corresponding to the local learning network. Using a preset mask matrix to encrypt the local adjacency matrix to obtain a local encrypted matrix includes: When the number of nodes in the federated learning network meets the first order of magnitude, each data holder combines the local adjacency matrix and the local degree matrix to obtain a combined polynomial matrix; Receive the preset mask matrix pre-generated by the mask server, where the mask server pre-generates a random orthogonal matrix as the preset mask matrix; Combine the random orthogonal matrix with the combined polynomial matrix to obtain the local encryption matrix of each data holder.
5. The method according to claim 1, wherein The local network matrix at least includes: the local adjacency matrix and the local degree matrix corresponding to the local learning network. Encrypting the local adjacency matrix using the preset mask matrix to obtain the local encryption matrix includes: When the number of nodes in the federated learning network meets the second order of magnitude, each data holder uploads the local degree matrix to the embedding server respectively and receives the global degree matrix returned by the embedding server based on the multiple local degree matrices; Determine a random walk polynomial matrix based on the local degree matrix and Laplacian matrix of the same data holder, and the global degree matrix, where the Laplacian matrix is determined based on the difference between the local adjacency matrix and the local degree matrix of the same data holder, and the random walk polynomial matrix is expressed as: , where \(i\) is the data holder, is the random walk polynomial matrix of the data holder, \(D\) is the global degree matrix, is the local degree matrix of the data holder, is the Laplacian matrix of the data holder; Perform federated power iteration on the random walk polynomial matrix according to a preset number of iterations to generate a random walk submatrix, where the federated power iteration is used to project the random walk polynomial matrix into a low-dimensional space; Encrypt the random walk submatrix using the preset mask matrix to obtain the local encryption matrix.
6. The method according to claim 5, wherein Each step of the federated power iteration includes: Each data holder encrypts the random walk polynomial matrix using the preset mask matrix to obtain a local pre-encrypted matrix; Upload the local pre-encrypted matrix to the embedding server, where the embedding server is used to aggregate the multiple local pre-encrypted matrices to obtain a global pre-encrypted matrix, and perform coefficient matrix decomposition on the global pre-encrypted matrix to obtain an intermediate result matrix, and the intermediate result matrix is a masked orthogonal matrix; Receive the intermediate result matrix returned by the embedding server, and remove the mask in the intermediate result matrix according to the preset mask matrix to obtain an intermediate encrypted matrix. When the number of iterations reaches the preset number of iterations, the intermediate encrypted matrix is the random walk submatrix; when the number of iterations does not reach the preset number of iterations, the intermediate encrypted matrix is the random walk polynomial matrix for the next iteration.
7. The method according to claim 1, characterized in that The method further includes: Aggregate the local encryption matrices uploaded by multiple data holders respectively through the embedding server to obtain a global aggregation matrix; Perform singular value decomposition on the global aggregation matrix to generate the global encryption matrix.
8. A network embedding device based on a federated learning network, characterized in that, Includes: An acquisition module, configured to acquire local network matrices of multiple data holders locally in a federated learning network, where the data holder is a participant in the federated learning network, the federated learning network is pre-divided into multiple local learning networks, and the local network matrix is a matrix representation of the local learning network; An encryption module, configured to encrypt the local network matrix using a preset mask matrix to obtain a local encryption matrix; An upload module, configured to upload local encryption matrices respectively corresponding to multiple said data holders to an embedding server, and receive a global encryption matrix obtained by the embedding server based on the integration of multiple said local encryption matrices, wherein the global encryption matrix is the result of the embedding server decomposing the integration result based on multiple said local encryption matrices; A decryption module, configured to decrypt the global encryption matrix based on the preset mask matrix to obtain a global embedding matrix, wherein the global embedding matrix is the network embedding result of the federated learning network.
9. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the network embedding method based on the federated learning network according to any one of claims 1 to 7 through the computer program.
10. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, the steps of the network embedding method based on the federated learning network according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Graph neural network federal recommendation method for privacy protection
CN113420232A
Longitudinal graph federal information recommendation method based on separation learning and related device
CN115952550A
Longitudinal federal learning method and system based on graph neural network
CN118036651A
Federated-learning-based user data classification method and apparatus, and device and medium
WO2021179720A1
Horizontal federated learning modeling optimization method, device, medium and program product
WO2023024368A1