Federated graph data enhancement method and system based on graph limit estimation based on privacy protection

By using graph limit estimation and differential privacy processing methods in federated graph learning, the problem that social network graph data enhancement cannot accurately reflect topological structure and privacy leakage is solved, and high-quality, diversity and privacy-safe graph data enhancement effect is achieved.

CN119443150BActive Publication Date: 2025-05-13CHONGQING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411494663.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-24
Publication Date
2025-05-13
Estimated Expiration
2044-10-24

AI Technical Summary

Technical Problem

The existing social network graph data enhancement methods cannot accurately reflect the topological characteristics of the original social network graph data and pose a risk of privacy leakage.

Method used

The social network graph is obtained through multiple clients, social nodes and edges are extracted to form the original graph dataset, generate graph limit estimates corresponding to local tags, and then shared to the server side after differential privacy processing. The server side calculates the paired cutting distance, generates the nearest neighbor, updates the graph limit estimate through interpolation and returns to the client. The client performs graph sampling based on the updated graph limit estimate to obtain graph enhanced data.

Benefits of technology

It improves the quality and accuracy of the graph data set, enhances the diversity and richness of the data, ensures the privacy and security of social network graphs during the sharing process, and accurately reflects the topological characteristics of the original graph data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119443150B_ABST
    Figure CN119443150B_ABST
Patent Text Reader

Abstract

The present invention relates to graph data processing technology, and discloses a method for enhancing federated graph data based on privacy-preserving graph limit estimation, including: a client obtains a social network graph, and composes an original graph data set to generate a graph limit estimation corresponding to a local label, and performs differential privacy on the graph limit estimation; a server calculates the pairwise cutting distance of each graph limit estimation, and generates multiple nearest neighbors of the graph limit estimation, calculates an adaptive threshold of each nearest neighbor, performs interpolation mixing according to the adaptive threshold, and updates the graph limit estimation; a client receives the updated graph limit estimation to generate a graph data set to be optimized, and performs graph sampling on the graph data set to be optimized according to the proportion of the number of the original graph data sets to obtain graph enhanced data. The present invention also proposes a system for a method for enhancing federated graph data based on privacy-preserving graph limit estimation, which can solve the problem that the existing graph data enhancement cannot accurately reflect the topological structure characteristics of the original graph data, and improve the security of privacy protection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of graph data processing, and in particular to a method and system for enhancing federated graph data based on privacy-preserving graph limit estimation. Background Art

[0002] In the field of online social networking, graph neural networks have shown excellent performance in the field of social network graph data modeling. However, its development still faces many challenges, the most prominent of which is the issue of data privacy. Since social network graph data often contains sensitive information and is difficult to collect centrally, this has brought great difficulties to traditional centralized training methods. Federated Learning (FL), as an emerging machine learning paradigm, allows clients to collaboratively train models under the coordination of a central server without having to share original data. This approach significantly reduces privacy risks and provides data-sensitive industries with the opportunity to collaboratively train shared GNN models, which has led to the concept of Federated Graph Learning (FGL).

[0003] Although FGL provides a promising solution, it still faces key challenges in practical applications, the most prominent of which is the non-independently and identically distributed (Non-IID) characteristic of data between participating nodes. This characteristic mainly stems from the inconsistency of data acquisition equipment specifications and the differences in acquisition objects. These factors lead to large heterogeneity in local data between different clients. In particular, when the client data comes from different fields, it may cause FGL instability and serious degradation of model performance.

[0004] At present, the core of the FGL field is still model training or parameter aggregation based on raw data, which fails to solve the problem of data heterogeneity. The main idea for solving the problem of data heterogeneity is to introduce generative adversarial networks, generate virtual samples and labels and share them between nodes to balance the data distribution differences between nodes; or use shared simple statistical data, but the above two methods will lead to the inability to accurately reflect the topological structure characteristics of the original social network graph data, and will allow attackers to reversely deduce some features of the original social network graph data through these statistical data, leading to user privacy leakage. Summary of the invention

[0005] The present invention provides a method and system for enhancing federated graph data based on privacy-preserving graph limit estimation, which can solve the problem that existing social network graph data enhancement cannot accurately reflect the topological structure characteristics of the original social network graph data and improve the security of privacy protection.

[0006] To achieve the above objectives, the present invention provides a method for enhancing federated graph data based on privacy-preserving graph limit estimation, comprising:

[0007] Multiple clients obtain a social network graph, and extract social nodes and social edges in the social network graph to form an original graph data set;

[0008] According to the local label of each client, the graph limit estimation corresponding to the local label is generated by using the original graph data set, and the graph limit estimation is shared to the server after differential privacy operation;

[0009] The server side calculates the pairwise cutting distances of the graph limit estimate shared by each of the clients, generates a plurality of nearest neighbors of the graph limit estimate according to the pairwise cutting distances, calculates an adaptive threshold of each nearest neighbor using a local distance distribution, updates the graph limit estimate by interpolation mixing according to the adaptive threshold to obtain an updated graph limit estimate, and returns the updated graph limit estimate to the client side;

[0010] The client receives the updated graph limit estimation to generate a graph data set to be optimized, and performs graph sampling on the graph data set to be optimized according to the ratio of the number of original graph data sets of local labels of the client to obtain graph enhancement data.

[0011] Optionally, extracting social nodes and social edges in the social network graph to form an original graph dataset includes:

[0012] The social nodes and the social edges connecting the social nodes are used to form a data graph, and a client is used to perform a labeling operation on the data graph to obtain the original graph data set.

[0013] Optionally, the generating the graph limit estimate corresponding to the local label using the original graph dataset includes:

[0014] Performing graph alignment preprocessing on the original graph dataset of each client to obtain a standard graph dataset;

[0015] Performing averaging processing on the standard graph data set to obtain an average graph data set;

[0016] Performing noise removal processing on the average graph data set by using singular value decomposition and thresholding to obtain a target graph data set;

[0017] The target graph data set is processed with a preset step function to approximate the graph limit, so as to obtain the graph limit estimate.

[0018] Optionally, performing graph alignment preprocessing on the original graph dataset of each client to obtain a standard graph dataset includes:

[0019] Calculating the node degree of the original graph data set of each client, and normalizing the calculated node degree of the original graph data set of each client to obtain the node degree of the same scale;

[0020] The number of nodes in the original graph data set of each client is arranged in descending order using the node degree of the same scale, and the maximum number of nodes in the descending order is extracted, and the graph data set that is not greater than the maximum number of nodes is subjected to node filling processing, and the standard graph data set is obtained after completing the node filling processing.

[0021] Optionally, the server side calculates the pairwise cut distances of the graph limit estimates shared by each of the clients, including:

[0022] Discretizing and sampling the graph limit estimation of each client according to a preset sampling dimension to obtain a plurality of graph limit estimation matrices;

[0023] Using an optimization algorithm to extract the optimal vertex arrangement of each of the graph limit estimation matrices, and optimizing each of the graph limit estimation matrices according to the optimal vertex arrangement to obtain a plurality of optimized graph limit estimation matrices;

[0024] And the cutting norms between the plurality of optimized graph limit estimation matrices are approximately calculated to obtain the pairwise cutting distances of the graph limit estimation.

[0025] Optionally, updating the graph limit estimate by interpolation mixing according to the adaptive threshold to obtain an updated graph limit estimate includes:

[0026] Calculate the local median and median absolute deviation of the graph limit estimates shared by each client, and calculate the local threshold based on the local median and median absolute deviation;

[0027] Determining a neighbor set of graph limit estimates shared by each client according to the local threshold;

[0028] The graph limit estimate shared by each client is interpolated and mixed according to the preset mixing parameters and the neighbor set to obtain an updated graph limit estimate.

[0029] Optionally, sampling the graph data set to be optimized according to the ratio of the number of original graph data sets of local labels of the client to obtain graph enhancement data includes:

[0030] Initialize an empty graph set, generate a uniformly distributed graph sampling matrix according to a preset number of graph samplings, and construct a binary adjacency matrix according to the graph sampling matrix;

[0031] The binary adjacency matrix is ​​optimized according to the symmetry of the matrix to obtain an optimized adjacency matrix, and non-isolated nodes in the graph data set to be optimized are extracted to form a non-isolated node set;

[0032] updating the optimized adjacency matrix using the non-isolated node set, and constructing an edge set according to the optimized adjacency matrix;

[0033] Constructing a graph instance using the edge set and the non-isolated node set, and adding the graph instance to the empty graph set to obtain a generated graph data set;

[0034] Calculate the first node feature and the first degree information of the original graph data set, and calculate the second degree information in the generated graph data set, and match the second degree information with the first degree information to obtain the first node feature corresponding to the first degree information that matches the closest, assign the first node feature to the node feature corresponding to the second degree information, and combine the original graph data set and the generated graph data set to obtain the graph enhancement data.

[0035] In order to solve the above problems, the present invention also provides a system for a federal graph data enhancement method based on privacy-preserving graph limit estimation, including a server side, and one or more clients connected to and communicating with the server side.

[0036] Optionally, the client and the server communicate bidirectionally, the client obtains a social network graph, and extracts social nodes and social edges in the social network graph to form an original graph data set; based on the local label of each of the clients, the original graph data set is used to generate a graph limit estimate corresponding to the local label, and the graph limit estimate is shared to the server after a differential privacy operation; and the client receives the updated graph limit estimate to generate a graph data set to be optimized, and performs graph sampling on the graph data set to be optimized according to the ratio of the number of original graph data sets of the client's local labels to obtain graph enhancement data.

[0037] Optionally, the server calculates the pairwise cutting distances of the graph limit estimate shared by each of the clients, generates multiple nearest neighbors of the graph limit estimate based on the pairwise cutting distances, calculates an adaptive threshold of each nearest neighbor using local distance distribution, updates the graph limit estimate through interpolation mixing based on the adaptive threshold to obtain an updated graph limit estimate, and returns the updated graph limit estimate to the client.

[0038] The present invention performs privacy processing on graph limit estimation through differential privacy operations, which can ensure the privacy security of social network graphs during sharing. In addition, by optimizing the graph data set using graph limit estimation and the nearest neighbor algorithm, the quality of the graph data set can be improved. By calculating the adaptive threshold and performing interpolation mixing, the accuracy of the graph limit estimation can be improved and the balance of multi-category data sets can be maintained. In addition, by sampling the graph data set to be optimized to obtain graph enhancement data, the diversity and richness of the data can be enhanced, and the topological structure characteristics of the original graph data can be accurately reflected. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 A schematic flow chart of a method for enhancing federated graph data based on privacy-preserving graph limit estimation provided by an embodiment of the present invention;

[0040] Figure 2 A flowchart of an optional embodiment of a method for enhancing federated graph data based on privacy-preserving graph limit estimation provided by an embodiment of the present invention;

[0041] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0042] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0043] The embodiment of the present application provides a method for enhancing federal graph data based on privacy-preserving graph limit estimation. The execution subject of the method for enhancing federal graph data based on privacy-preserving graph limit estimation includes but is not limited to at least one of the electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the method for enhancing federal graph data based on privacy-preserving graph limit estimation can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server, or it can be a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and basic cloud computing services such as big data and artificial intelligence platforms.

[0044] Reference Figure 1 As shown, it is a flow chart of a method for enhancing federated graph data based on graph limit estimation with privacy protection provided by an embodiment of the present invention. In this embodiment, the method for enhancing federated graph data based on graph limit estimation with privacy protection includes:

[0045] S1. Multiple clients obtain a social network graph, and extract social nodes and social edges in the social network graph to form an original graph data set.

[0046] In the embodiment of the present invention, the social network graph refers to a connection graph with social edges and social nodes, wherein the social edge refers to the social connection and social behavior between two users. If there is a social connection or social behavior between two users, there is a social edge between the two social nodes representing the two users. If there is no social connection or social behavior between the two users, there is no social edge between the two social nodes representing the two users. For example, if user A likes a passage of user B online, the social edge is all the interactive behaviors of user A. The social node refers to a social entity in the social graph, and the social entity can be an individual, an organization, etc. For example, in the above example, user A and user B in 'user A likes a passage of user B online' are individual entities in the social entity.

[0047] As an embodiment of the present invention, extracting social nodes and social edges in the social network graph to form an original graph data set includes:

[0048] The social nodes and the social edges connecting the social nodes are used to form a data graph, and a client is used to perform a labeling operation on the data graph to obtain the original graph data set.

[0049] Exemplarily, the original graph dataset is represented as:

[0050] Define the number of clients as M, and the original graph dataset of the mth client as G m =(V m ,E m ), G k is a group of social nodes V m and a set of social edges E connecting social nodes m The graphs composed of them have corresponding graph labels C is the number of graph categories in this group. Each node v∈V m With node feature vectors, forming a feature matrix N m is the number of client nodes, F m is the number of feature dimensions.

[0051] In the embodiment of the present invention, the client refers to each device or organization participating in the federated learning process in federated learning, and each of them has data.

[0052] S2. According to the local label of each of the clients, the original graph data set is used to generate a graph limit estimate corresponding to the local label, and the graph limit estimate is shared to the server after performing a differential privacy operation.

[0053] In the embodiment of the present invention, the local label refers to a label for classifying the graph data set in the client.

[0054] As an embodiment of the present invention, the step of generating the graph limit estimate corresponding to the local label by using the original graph data set includes:

[0055] Performing graph alignment preprocessing on the original graph dataset of each client to obtain a standard graph dataset;

[0056] Performing averaging processing on the standard graph data set to obtain an average graph data set;

[0057] Performing noise removal processing on the average graph data set by using singular value decomposition and thresholding to obtain a target graph data set;

[0058] The target graph data set is processed with a preset step function to approximate the graph limit, so as to obtain the graph limit estimate.

[0059] In the embodiment of the present invention, the preset step function refers to a step function predefined when performing graph limit processing, and the preset step function may use a thresholded singular value to reconstruct a transition function.

[0060] Furthermore, the original graph dataset of each client is subjected to graph alignment preprocessing to obtain a standard graph dataset, including:

[0061] Calculating the node degree of the original graph data set of each client, and normalizing the calculated node degree of the original graph data set of each client to obtain the node degree of the same scale;

[0062] The number of nodes in the original graph data set of each client is arranged in descending order using the node degree of the same scale, and the maximum number of nodes in the descending order is extracted, and the graph data set that is not greater than the maximum number of nodes is subjected to node filling processing, and the standard graph data set is obtained after completing the node filling processing.

[0063] In the embodiment of the present invention, the purpose of using singular value decomposition and general singular value thresholding is to decompose the image into main components and secondary components, so as to effectively remove noise.

[0064] Exemplarily, the original graph dataset of each client is subjected to graph alignment preprocessing to obtain a standard graph dataset, and the following implementation steps may be adopted:

[0065] In the original graph dataset of the m-th client, the graph dataset G belonging to class i m,i contains a graph g m,i , where the node set V m (g m,i );

[0066] Step 1: Calculate the node degree d(v) for each graph and perform a normalization operation to obtain the normalized graph node degree to eliminate the influence of graph size differences:

[0067]

[0068]

[0069] Step 2: Use the normalized node degree values to sort the nodes in descending order to obtain the sorted node sequence S m (g m,i ), where the sorting function π can use bubble sort, quick sort, heap sort, etc., and accordingly rearrange the adjacency matrix A(g m,i ):

[0070] S m (g m,i ) = π(V m (g m,i )) = (v 1 , v 2 ,..., v n )

[0071] where n = |V m (g m,i )|, and for any 1 ≤ j < l ≤ n, there is

[0072] A′ m (g m,i )[π(v j ), π(v l )] = A(g m,i )[v, u], v, u ∈ V m (g m,i )

[0073] Step 3: Count the maximum number of nodes N m,i in G max , for graphs with n i < N max , supplement the number of nodes to N max , fill 0 on the right and below the adjacency matrix, and also fill 0 for the degree vector:

[0074]

[0075]

[0076] Exemplarily, the averaging process is performed on the standard image data set to obtain the average image data set, and the following implementation steps may be adopted:

[0077] Compute the average of all standard graph datasets:

[0078]

[0079] Among them, g m,i,j represents the i-th type of graph dataset owned by the k-th client, N represents the number of graphs, Represents the average of all standard graph datasets.

[0080] Exemplarily, the process of removing noise from the average graph dataset by using singular value decomposition and thresholding to obtain a target graph dataset may be implemented by the following steps:

[0081] First, perform singular value decomposition on the average graph dataset:

[0082]

[0083] Among them, U m,i and V m,i is an orthogonal matrix containing left and right singular vectors, S m,i is a vector of singular values;

[0084] Then, the threshold τ is calculated and applied:

[0085]

[0086] S m,i (τ)=max(S m,i -τ,0)

[0087] Among them, λ is a preset parameter, and its default value is the optimal value of 2.02 according to theoretical analysis. is the number of nodes and is used to adjust the threshold to adapt to the size of the graph.

[0088] Exemplarily, the method of performing approximate graph limit processing on the target graph data set using a preset step function to obtain the graph limit estimate may be implemented by the following steps:

[0089] Step 1) Use the thresholded singular values ​​to reconstruct the transition function and calculate the approximate graph limit And normalize it:

[0090]

[0091] Where diag(·) represents the diagonal matrix consisting of the diagonal elements of the matrix. is the transposed matrix of the node features of the i-th type of graph dataset in the m-th client.

[0092] Step 2) retains the main singular values ​​and the corresponding singular vectors. This step actually constructs W P The coefficient matrix w in (x,y) kk′ , which is used to measure the connection probability between subintervals, so it needs to be normalized to limit the estimated value to the range [0,1] to meet the definition of the graph limit:

[0093]

[0094] In the embodiment of the present invention, the differential privacy operation of the graph limit estimation may be performed by the following implementation steps:

[0095] For the calculated Collection, add noise:

[0096]

[0097] in, represents the limit estimate of the graph after adding noise, ∈ represents the noise term, where the distribution of the noise term has a mean of 0 and a covariance matrix of σ 2 Gaussian distribution of I∈~N(0,σ 2 I).

[0098] In the embodiment of the present invention, the server refers to a central server or coordinator in federated learning, which is usually responsible for coordinating and managing various clients in the federated learning process.

[0099] S3. The server side calculates the pairwise cutting distances of the graph limit estimate shared by each of the clients, generates multiple nearest neighbors of the graph limit estimate based on the pairwise cutting distances, calculates an adaptive threshold of each nearest neighbor using the local distance distribution, updates the graph limit estimate by interpolation mixing based on the adaptive threshold to obtain an updated graph limit estimate, and returns the updated graph limit estimate to each client.

[0100] In the embodiment of the present invention, the cut distance is a metric. The cut distance has isomorphism invariance and can effectively capture the essential similarity of graph structures.

[0101] As an embodiment of the present invention, the calculating of the pairwise cutting distances of the graph limit estimates shared by each of the clients includes:

[0102] Discretizing and sampling the graph limit estimation of each client according to a preset sampling dimension to obtain a plurality of graph limit estimation matrices;

[0103] Using an optimization algorithm to extract the optimal vertex arrangement of each of the graph limit estimation matrices, and optimizing each of the graph limit estimation matrices according to the optimal vertex arrangement to obtain a plurality of optimized graph limit estimation matrices;

[0104] And the cutting norms between the plurality of optimized graph limit estimation matrices are approximately calculated to obtain the pairwise cutting distances of the graph limit estimation.

[0105] Exemplarily, the calculation of the pairwise cutting distances of the graph limit estimates shared by each of the clients may be implemented by the following steps:

[0106] For any two graphs, the limiting estimate and The cutting distance is defined as:

[0107]

[0108] in, inf means taking the infimum on all measure-preserving bijections from the unit interval to itself, ∥·∥ is the cut norm, and the cut norm is defined as follows, S and T are measurable subsets of the unit interval [0,1]:

[0109]

[0110] Step 1: Discretization sampling: discretize each encrypted graph limit estimate into an n×n matrix representation, where n is a predefined sampling dimension, and the predefined sampling dimension may be n=1000;

[0111] Step 2: Optimize the vertex arrangement and apply the optimization algorithm to find the best vertex arrangement to minimize the cut norm;

[0112] Step 3: Perform approximate calculations to calculate the cutting norm between the optimized matrices as the approximate cutting distance between the encrypted graph limit estimates, and obtain the cutting distance matrix D i , where D i,j represents the cut distance sequence between the encrypted graph limit estimates provided by the i-th client and the j-th client.

[0113] The method of approximately calculating the cutting distance in the embodiment of the present invention can save time cost.

[0114] In the embodiment of the present invention, the optimization algorithm includes but is not limited to simulated annealing or genetic algorithm.

[0115] In the embodiment of the present invention, the calculation of the cut norm between the optimized matrices is to find a permutation matrix P for two matrices A and B so that ∥A-PBP T ∥Minimum.

[0116] Exemplarily, the generating of the plurality of nearest neighbors of the graph limit estimation according to the pairwise cutting distances may be implemented by the following steps:

[0117] Based on the calculated cut distance, for each graph limit estimate, select the k other graph limit estimates with the smallest distance as its nearest neighbors. The choice of k value can be adjusted according to the specific application scenario and data scale.

[0118] Exemplarily, the method of calculating the adaptive threshold of each nearest neighbor using the local distance distribution may adopt the following implementation steps:

[0119] Calculating a hybrid threshold for each encrypted graph limit estimate based on the k nearest neighbors and the pairwise cut distance matrix;

[0120] For each encrypted graph the limit is estimated calculate Local median and median absolute deviation (MAD):

[0121]

[0122]

[0123] Calculate the local threshold, where α is a predefined anomaly threshold parameter. The choice of α affects the degree of mixing. Larger values ​​of α result in more extensive mixing, while smaller values ​​of α restrict the mixing range:

[0124]

[0125] According to T i To determine the set of neighbors to be mixed with the graph limit estimate

[0126] Further, the updating the graph limit estimate by interpolation mixing according to the adaptive threshold to obtain an updated graph limit estimate includes:

[0127] Calculate the local median and median absolute deviation of the graph limit estimates shared by each client, and calculate the local threshold based on the local median and median absolute deviation;

[0128] Determining a neighbor set of graph limit estimates shared by each client according to the local threshold;

[0129] The graph limit estimate shared by each client is interpolated and mixed according to the preset mixing parameters and the neighbor set to obtain an updated graph limit estimate.

[0130] In the embodiment of the present invention, in order to determine the degree of mixing, a mixing threshold is calculated for each encrypted graph limit estimate based on the k nearest neighbors and the obtained approximate cutting distance matrix.

[0131] In the embodiment of the present invention, interpolating and mixing the graph limit estimate shared by each client according to the preset mixing parameter and the neighbor set includes:

[0132] The following formula is used for interpolation mixing:

[0133]

[0134] Where β∈[0,1] is a predefined mixing parameter used to control the mixing ratio.

[0135] S4. The client receives the updated graph limit estimation to generate a graph data set, obtains the graph data set to be optimized, and samples the graph data set to be optimized according to the ratio of the number of original graph data sets of the client's local labels to obtain graph enhancement data.

[0136] As an embodiment of the present invention, sampling the graph data set to be optimized according to the ratio of the number of original graph data sets of local labels of the client to obtain graph enhancement data includes:

[0137] Initialize an empty graph set, generate a uniformly distributed graph sampling matrix according to a preset number of graph samplings, and construct a binary adjacency matrix according to the graph sampling matrix;

[0138] The binary adjacency matrix is ​​optimized according to the symmetry of the matrix to obtain an optimized adjacency matrix, and non-isolated nodes in the graph data set to be optimized are extracted to form a non-isolated node set;

[0139] updating the optimized adjacency matrix using the non-isolated node set, and constructing an edge set according to the optimized adjacency matrix;

[0140] Constructing a graph instance using the edge set and the non-isolated node set, and adding the graph instance to the empty graph set to obtain a generated graph data set;

[0141] Calculate the first node feature and the first degree information of the original graph data set, and calculate the second degree information in the generated graph data set, and match the second degree information with the first degree information to obtain the first node feature corresponding to the first degree information that matches the closest, assign the first node feature to the node feature corresponding to the second degree information, and combine the original graph data set and the generated graph data set to obtain the graph enhancement data.

[0142] Exemplarily, the sampling of the to-be-optimized graph data set according to the ratio of the number of original graph data sets of the local labels of the client to obtain graph enhancement data may be performed by the following implementation steps:

[0143] Step 1: Discretize the continuous graph limit estimation into a specific graph structure through random sampling and thresholding operations, so that the client can learn the global structural information in the data preprocessing stage;

[0144] Suppose the graph limit estimate obtained by the jth label of the mth client is Initialize an empty graph collection The initial number of nodes is N, and the preset number of samples is n. For each sample i=1,2,...,n, a uniformly distributed random matrix R is generated:

[0145] R u,v ~Uniform(0,1)

[0146] Reconstruct the binary adjacency matrix A, where 1[·] is an indicator function, that is, it only takes 0 or 1:

[0147]

[0148] Step 2: Use the symmetric property of the matrix to ensure the non-directionality of the graph, where triu(·) represents the strictly upper triangular part of the matrix excluding the diagonal, and diag(·) represents the diagonal matrix composed of the diagonal elements of the matrix:

[0149] A=triu(A)+triu(A) T -diag(A)

[0150] Suppose the set of non-isolated nodes is S:

[0151]

[0152] Update the adjacency matrix A, remove the isolated node A = A[S,S], and construct the edge set E based on the final adjacency matrix:

[0153] E={(u,v)|A u,v =1,u <v}

[0154] Construction graph instance G i =(V i ,V i ,j), where the vertex set V i ={1,2,...,|S|}, add it to the graph collection until all preset sampling numbers are completed.

[0155] Step 3: Define the original graph dataset of the mth client as G m , the newly generated graph dataset is G' m First, from G m Extract all node features Ft m and degree information m , where d(v) represents the degree of node v:

[0156] Ft m ={x i |x i ∈X m ,(V m ,X m ,E m )∈G m}

[0157] De m ={d(v i )|v i ∈V m ,(V m ,X m ,E m )∈G m}

[0158] For G' m For each node v∈V' m , calculate its degree d(v), and then in De m Find the node set V with the closest degree value close , randomly select a node u and assign the corresponding feature of u to v:

[0159] x v =Ft m (u)

[0160] The generated feature matrix X' m That is, the feature matrix of the newly generated graph dataset, G m and G' m The combination constitutes the enhanced graph dataset G" m .

[0161] In the embodiment of the present invention, in order to maintain the consistency between the generated graph data set and the original data set, a node feature generation method based on degree matching is adopted.

[0162] In this embodiment, the system of the method for enhancing federated graph data based on privacy-preserving graph limit estimation includes a server and one or more clients connected to and communicating with the server.

[0163] As an embodiment of the present invention, the client and the server communicate bidirectionally, the client obtains a social network graph, and extracts social nodes and social edges in the social network graph to form an original graph data set; based on the local label of each of the clients, the original graph data set is used to generate a graph limit estimate corresponding to the local label, and the graph limit estimate is shared to the server after a differential privacy operation; and the client receives the updated graph limit estimate to generate a graph data set to be optimized, and samples the graph data set to be optimized according to the ratio of the number of original graph data sets of the client's local labels to obtain graph enhancement data.

[0164] As an embodiment of the present invention, the server calculates the pairwise cutting distances of the graph limit estimate shared by each of the clients, generates multiple nearest neighbors of the graph limit estimate based on the pairwise cutting distances, calculates an adaptive threshold of each nearest neighbor using the local distance distribution, updates the graph limit estimate through interpolation mixing based on the adaptive threshold to obtain an updated graph limit estimate, and returns the updated graph limit estimate to the client.

[0165] The embodiment of the present invention performs privacy processing on graph limit estimation through differential privacy operations, which can ensure the privacy security of social network graphs during sharing. In addition, by optimizing the graph data set using graph limit estimation and the nearest neighbor algorithm, the quality of the graph data set can be improved. By calculating the adaptive threshold and performing interpolation mixing, the accuracy of the graph limit estimation can be improved and the balance of multi-category data sets can be maintained. In addition, by sampling the graph data set to be optimized to obtain graph enhancement data, the diversity and richness of the data can be enhanced, and the topological structure characteristics of the original graph data can be accurately reflected.

[0166] It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0167] Therefore, no matter from which point of view, the embodiments should be regarded as illustrative and non-restrictive, and the scope of the present invention is limited by the appended claims rather than the above description, so it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention. Any attached figure mark in the claims should not be regarded as limiting the claims involved.

[0168] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0169] In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in the system claim can also be implemented by one unit or device through software or hardware. The words first, second, etc. are used to indicate names, and do not indicate any particular order.

[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.

Claims

1. A federated graph data enhancement method based on privacy-preserving graph limit estimation, characterized in that: The method comprises: Multiple clients obtain a social network graph, and extract social nodes and social edges in the social network graph to form an original graph data set; According to the local label of each client, the graph limit estimation corresponding to the local label is generated by using the original graph data set, and the graph limit estimation is shared to the server after differential privacy operation, wherein the local label of the client refers to the label for classifying the graph data set in the client; The server side calculates the pairwise cutting distances of the graph limit estimate shared by each of the clients, generates a plurality of nearest neighbors of the graph limit estimate according to the pairwise cutting distances, calculates an adaptive threshold of each nearest neighbor using a local distance distribution, updates the graph limit estimate by interpolation mixing according to the adaptive threshold to obtain an updated graph limit estimate, and returns the updated graph limit estimate to the client side; The client receives the updated graph limit estimation to generate a graph data set to be optimized, and performs graph sampling on the graph data set to be optimized according to the ratio of the number of original graph data sets of local labels of the client to obtain graph enhancement data.

2. The method for enhancing federated graph data based on privacy-preserving graph limit estimation according to claim 1, characterized in that: The step of extracting social nodes and social edges from the social network graph to form an original graph data set includes: The social nodes and the social edges connecting the social nodes are used to form a data graph, and a client is used to perform a labeling operation on the data graph to obtain the original graph data set.

3. The method for enhancing federated graph data based on privacy-preserving graph limit estimation according to claim 1 or 2, characterized in that: The step of generating the graph limit estimate corresponding to the local label by using the original graph data set includes: Performing graph alignment preprocessing on the original graph dataset of each client to obtain a standard graph dataset; Performing averaging processing on the standard graph data set to obtain an average graph data set; Performing noise removal processing on the average graph data set by using singular value decomposition and thresholding to obtain a target graph data set; The target graph data set is processed with a preset step function to approximate the graph limit, so as to obtain the graph limit estimate.

4. The method for enhancing federated graph data based on privacy-preserving graph limit estimation according to claim 3, characterized in that: The process of performing graph alignment preprocessing on the original graph dataset of each client to obtain a standard graph dataset includes: Calculating the node degree of the original graph data set of each client, and normalizing the calculated node degree of the original graph data set of each client to obtain the node degree of the same scale; The number of nodes in the original graph data set of each client is arranged in descending order using the node degree of the same scale, and the maximum number of nodes in the descending order is extracted, and the graph data set that is not greater than the maximum number of nodes is subjected to node filling processing, and the standard graph data set is obtained after completing the node filling processing.

5. The method for enhancing federated graph data based on privacy-preserving graph limit estimation according to claim 1, 2 or 4, characterized in that: The server side calculates the pairwise cut distances of the graph limit estimates shared by each of the clients, including: Discretizing and sampling the graph limit estimation of each client according to a preset sampling dimension to obtain a plurality of graph limit estimation matrices; Using an optimization algorithm to extract the optimal vertex arrangement of each of the graph limit estimation matrices, and optimizing each of the graph limit estimation matrices according to the optimal vertex arrangement to obtain a plurality of optimized graph limit estimation matrices; And the cutting norms between the plurality of optimized graph limit estimation matrices are approximately calculated to obtain the pairwise cutting distances of the graph limit estimation.

6. The method for enhancing federated graph data based on privacy-preserving graph limit estimation according to claim 1, 2 or 4, characterized in that: The updating of the graph limit estimate by interpolation mixing according to the adaptive threshold to obtain an updated graph limit estimate comprises: Calculate the local median and median absolute deviation of the graph limit estimates shared by each client, and calculate the local threshold based on the local median and median absolute deviation; Determining a neighbor set of graph limit estimates shared by each client according to the local threshold; The graph limit estimate shared by each client is interpolated and mixed according to the preset mixing parameters and the neighbor set to obtain an updated graph limit estimate.

7. The method for enhancing federated graph data based on privacy-preserving graph limit estimation according to claim 1, 2 or 4, characterized in that: The step of sampling the image data set to be optimized according to the ratio of the number of original image data sets of the local labels of the client to obtain image enhancement data includes: Initialize an empty graph set, generate a uniformly distributed graph sampling matrix according to a preset number of graph samplings, and construct a binary adjacency matrix according to the graph sampling matrix; The binary adjacency matrix is ​​optimized according to the symmetry of the matrix to obtain an optimized adjacency matrix, and non-isolated nodes in the graph data set to be optimized are extracted to form a non-isolated node set; updating the optimized adjacency matrix using the non-isolated node set, and constructing an edge set according to the optimized adjacency matrix; Constructing a graph instance using the edge set and the non-isolated node set, and adding the graph instance to the empty graph set to obtain a generated graph data set; Calculate the first node feature and the first degree information of the original graph data set, and calculate the second degree information in the generated graph data set, and match the second degree information with the first degree information to obtain the first node feature corresponding to the first degree information that matches the closest, assign the first node feature to the node feature corresponding to the second degree information, and combine the original graph data set and the generated graph data set to obtain the graph enhancement data.

8. A system based on the method for enhancing federated graph data based on privacy-preserving graph limit estimation according to any one of claims 1 to 7, characterized in that: It includes a server and one or more clients connected to and communicating with the server.

9. The system of the method for enhancing federated graph data based on privacy-preserving graph limit estimation as claimed in claim 8, characterized in that: The client and the server communicate bidirectionally, the client obtains a social network graph, and extracts social nodes and social edges in the social network graph to form an original graph data set; based on the local label of each client, the original graph data set is used to generate a graph limit estimate corresponding to the local label, and the graph limit estimate is shared to the server after a differential privacy operation, wherein the local label of the client refers to a label for classifying the graph data set in the client; and the client receives an updated graph limit estimate to generate a graph data set to be optimized, and samples the graph data set to be optimized according to the ratio of the number of original graph data sets of the client's local labels to obtain graph enhancement data.

10. The system of the method for enhancing federated graph data based on privacy-preserving graph limit estimation as claimed in claim 8, characterized in that: The server side calculates the pairwise cutting distances of the graph limit estimate shared by each of the clients, generates multiple nearest neighbors of the graph limit estimate based on the pairwise cutting distances, calculates an adaptive threshold of each nearest neighbor using the local distance distribution, updates the graph limit estimate through interpolation mixing based on the adaptive threshold to obtain an updated graph limit estimate, and returns the updated graph limit estimate to the client.

Citation Information

Patent Citations

  • Network data publishing method based on differential privacy and closeness centrality

    CN115438227A

  • Central disease prediction method and device based on local graph information exchange under privacy protection

    CN116564535A