Privacy Feature Protection Method and System Based on Graph Neural Network

Through the privacy feature protection method based on graph neural network, the topological structure of graph data is disturbed, which solves the problem that it is difficult to protect user privacy features in graph data in the prior art, and achieves efficient privacy protection effect.

CN117056970BActive Publication Date: 2025-05-30ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311024849.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-15
Publication Date
2025-05-30
Estimated Expiration
2043-08-15

AI Technical Summary

Technical Problem

The prior art is difficult to effectively protect user privacy features in graph data, especially when the homogeneity of features and structural homogeneity are considered simultaneously, the computational complexity is high and the performance and applicability are poor.

Method used

The privacy feature protection method based on graph neural network is adopted, and the node feature matrix and the adjacency matrix are constructed, and the topological structure of the graph data is disturbed by the defensive model, reducing the privacy feature classification prediction ability of the opponent model.

Benefits of technology

It realizes effective protection of user privacy characteristics, reduces the risk of privacy leakage, and reduces computing complexity, improving the performance and applicability of the method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117056970B_ABST
    Figure CN117056970B_ABST
Patent Text Reader

Abstract

The present invention discloses a privacy feature protection method and system based on graph neural networks, belonging to the technical field of user privacy security. Given user information and the relationships between users, an original graph data is constructed; the sampling probability of each edge in the original graph data is calculated through a defensive model based on graph neural networks, and the edges are sampled according to the sampling probability to obtain perturbed graph data; a privacy feature classification prediction is performed on the perturbed graph data through an adversary model based on adjacency and structural similarity patterns; the defensive model and the adversary model are adversarially trained to obtain the final sampling probability of each edge, and according to the sampling results, the relationships between users corresponding to the edges not sampled are deleted from the data to be desensitized, so as to achieve privacy feature protection. The present invention not only reduces the prediction ability of the adversary model, but also minimizes the proportion of perturbation as much as possible, that is, the sampled edges are as consistent with the original graph as possible, and effectively protects the privacy features of users in the graph data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of user privacy security, and in particular, to a privacy feature protection method and system based on a graph neural network. Background Art

[0002] In graph data, due to the existence of user information and the relationships between users, such as friendship relationships, following relationships, citation relationships, etc., it is easy to infer users' privacy features, such as gender, age, etc., by using users' dependency relationships. In existing algorithms, the modeling methods adopted by traditional privacy feature protection methods are difficult to prevent adversaries from attacking privacy features by using the topological structure in the graph. In addition, although some protection methods adopt the approach of perturbing the topological structure, they do not consider cutting off the leakage of privacy information from both feature homogeneity and structure homogeneity, and the computational complexity is relatively high, resulting in poor performance and applicability.

[0003] Feature homogeneity and structure homogeneity bring important information to the inference of privacy features on graphs, and existing methods are difficult to well protect this kind of privacy leakage. Therefore, both academia and industry often face security risks when publishing graph data and need to perform data desensitization before publishing. From a microscopic perspective, in a part of graph data, the privacy features of nodes have a distribution pattern of feature homogeneity, that is, two nodes connected by an edge or with a short distance are more likely to have the same privacy feature. For example, people in the same region are more likely to become friends; while in another part of graph data, the privacy features have a distribution pattern of structure homogeneity, that is, nodes with the same sub-structure features are more likely to have the same privacy feature. For example, the friend network of young users is relatively dense, and the friend network of older users is relatively sparse. In common privacy feature inference methods, adversaries often obtain privacy information from the neighbors around a target node or nodes with similar structures to it, so as to more accurately infer the privacy features of the target and nodes.

[0004] Secondly, from a macroscopic perspective, given the privacy features to be protected, existing methods are also difficult to analyze the patterns of privacy leakage from both feature homogeneity and structure homogeneity. On the one hand, although existing methods have proposed methods to measure the strength of the distribution pattern of adjacency, there is still a lack of a quantitative analysis method for the correlation between node privacy features and their structural features, which makes it difficult for existing protection methods to capture the impact of local structural features on privacy leakage. On the other hand, due to different distribution patterns of different nodes in the network, it is also difficult to handle this co-existing distribution mixture pattern well in practice. Therefore, how to cut off privacy leakage by perturbing the topological structure, desensitize the user information to be published and the relationship data between users, and not cause the leakage of users' privacy features while retaining all user information, is still an important research-worthy problem. Summary of the Invention

[0005] To solve the problems of low performance and poor applicability of the existing privacy feature protection methods, the present invention proposes a privacy feature protection method and system based on graph neural networks to effectively protect user privacy features.

[0006] The present invention adopts the following technical solutions:

[0007] In a first aspect, the present invention provides a privacy feature protection method based on graph neural networks, including:

[0008] Given user information and the relationships between users as data to be desensitized;

[0009] Regarding each given user in the data to be desensitized as a node in the network topology structure, using the user information as node features to construct a node feature matrix; constructing an adjacency matrix based on the relationships between users in the data to be desensitized to obtain the original graph data; calculating the sampling probability of each edge in the original graph data through a defense model based on graph neural networks, sampling the edges according to the sampling probability, and updating the adjacency matrix according to the sampled edges to obtain the perturbed graph data;

[0010] Performing privacy feature classification prediction on the perturbed graph data through an adversary model based on adjacency and structural similarity patterns;

[0011] Performing adversarial training on the defense model and the adversary model to obtain the final sampling probability of each edge, sampling the edges according to the final sampling probability, and recording the unsampled edges; deleting the relationships between users corresponding to the unsampled edges from the data to be desensitized to obtain the data after privacy feature protection.

[0012] In a second aspect, the present invention provides a privacy feature protection system based on graph neural networks for implementing the above privacy feature protection method based on graph neural networks.

[0013] Compared with the prior art, the present invention has the beneficial effects that: the present invention establishes graph data based on given user information and the relationships between users, perturbs the topological structure of the original graph data by cutting the edges between nodes in the graph while keeping the node features unchanged, and effectively protects the privacy features of users in the graph data. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 is a schematic diagram of privacy feature protection based on graph neural networks shown according to an exemplary embodiment;

[0015] Figure 2 is a schematic diagram of an adversary model based on feature homogeneity and structural homogeneity shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0016] The present invention will be further described below in conjunction with the accompanying drawings and embodiments. The accompanying drawings are only schematic diagrams of the present invention. Some of the block diagrams shown in the accompanying drawings are functional entities, which do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0017] The goal of the present invention is to construct graph data based on given user information and the relationships between users, and to effectively protect the privacy features (such as gender, age group) of users by perturbing the topological structure of the graph data. Specifically, the present invention designs an end-to-end graph neural network for protecting the privacy features of graph nodes. In the present invention, the method for protecting user privacy features based on graph neural networks is to use adversarial training to degrade the performance of the adversary model based on feature homogeneity and structural homogeneity, while minimizing the perturbation ratio of the depth framework of the graph neural network, which mainly includes two parts: a defensive model based on graph neural networks and an adversary model based on feature homogeneity and structural homogeneity.

[0018] Figure 1 The overall framework of the present invention is shown. The input of the model is the original graph data constructed from the data to be desensitized, which includes the complete topological structure, user features (node features), and the true privacy features of users. The privacy features can be, for example, gender (male, female) or age group (youth, old age), etc.

[0019] In the present invention, the privacy features of nodes in some graph data have a distribution pattern of feature homogeneity, while in another part of the graph data, they have a distribution pattern of structural homogeneity. Therefore, the present invention uses an adversary model based on feature homogeneity and structural homogeneity as the adversary in the worst case of adversarial training. The adversary model uses graph neural networks and feature homogeneity and structural homogeneity coefficients to speculate on the privacy features in the graph data.

[0020] First, each given user is regarded as a node in the network topology. A node feature matrix is obtained from the given user information. This node feature matrix is used as the original input of the nodes in the graph, and an adjacency matrix is constructed based on the relationships between the given users to obtain the original graph data. The edges of the graph data are sampled through a defense model based on graph neural networks to perturb the topological structure of the graph data. Then, a privacy prediction is performed on the perturbed graph data using an adversary model based on feature homogeneity and structural homogeneity. The adversary model aggregates respectively in the perturbed full graph data and in the subgraph constructed by the 2-hop neighbors of each user in the perturbed graph, conducts privacy feature classification prediction, and calculates the feature homogeneity and structural homogeneity coefficients for weighting to obtain the final prediction result. Subsequently, the final sampling probability of each edge is obtained through adversarial training. The ultimate goal of implementing the present invention is that the sampled new graph data will not cause node privacy leakage, thereby achieving the protection of privacy.

[0021] In the defense model based on graph neural networks, first, the representations of each node are aggregated through graph neural networks. Then, the representations of the nodes corresponding to each edge are concatenated to obtain the edge representation. The sampling probability of each edge is calculated through the sigmoid function and the edge is sampled through the Gumbel-Softmax reparameterization method to obtain the perturbed new graph data.

[0022] In a specific implementation of the present invention, in the defense model based on graph neural networks, in order to obtain the edge representation, in this embodiment, GraphSAGE is first used as the convolutional layer. For each node, GraphSAGE takes the average of the representations of its neighbor nodes and concatenates it with the representation of the node itself. After the concatenated representation undergoes a linear transformation and a non-linear activation function, one aggregation is completed. In this embodiment, the graph neural network f in the defense model is designed by stacking two layers of GraphSAGE convolutional layers. d and a two-layer perceptron are used to obtain the representation of each edge, and the sampling probability of each edge is calculated through sigmoid:

[0023] H d = f d (X, A)

[0024]

[0025] where X represents the node feature matrix, A represents the adjacency matrix, f d (·) represents the graph neural network, represents the set of node representations, n represents the number of nodes, represents matrix concatenation, MLP is a multi-layer perceptron, and σ is the sigmoid function. represents the sampling probability of the edge (i, j) in the graph data.

[0026] The sampling process of the opposite side uses Gumbel-Softmax reparameterization, which is expressed as follows:

[0027]

[0028] where A′ ij represents the sampling result of the edge (i, j) between nodes i and j.

[0029] In the adversary model based on feature homogeneity and structural homogeneity, the graph neural network f p (·) is aggregated over the entire graph, aiming to make the privacy classification results between adjacent or nearby nodes tend to be the same, that is, the privacy classification results in the adjacency pattern. The graph neural network f s (·) is aggregated in the subgraph constructed by the second-order neighbors of each user, aiming to make the privacy classification results between nodes with similar substructures tend to be the same, that is, the privacy classification results in the structural similarity pattern. In the subsequent process of continuous learning of the model, f p (·) and f p (·) are adjusted in parallel through gradient backpropagation. The original signals of the nodes pass through f p (·) and f s (·) to obtain the privacy classification results based on the adjacency pattern and the structural similarity pattern respectively. Essentially, the nodes aggregate the information in their neighborhoods in the entire graph to obtain the privacy classification results based on the adjacency pattern aggregate and pool in the node pair second-order subgraph to obtain the privacy classification results based on the structural similarity pattern so that each node can obtain the privacy classification results in two different patterns. Then, the feature homogeneity and structural homogeneity coefficients are used to determine the importance of different classifiers for the nodes.

[0030] In this embodiment, the feature homogeneity and structural homogeneity coefficients at the node level are designed. For each node, the feature homogeneity coefficient calculates the proportion of its neighbors with the same privacy feature as it, and the structural homogeneity coefficient calculates the proportion of nodes with the same degree as it and the same privacy feature as it. To enhance robustness, the adversary model does not know the privacy features of some nodes. Therefore, for unknown nodes, the privacy classification results are used as pseudo-labels and participate in the calculation of the coefficients.

[0031] The above two considerations are respectively from the adjacency pattern and the structural similarity pattern of the nodes, which can make different nodes have different degrees of emphasis on the classifiers of the two patterns. For nodes that conform to the adjacency pattern, since the proportion of nodes with the same privacy characteristics among their neighbors is relatively high, their feature homogeneity coefficient is higher. For nodes that conform to the similarity pattern, the proportion of nodes with the same privacy characteristics among the nodes with the same degree as them is relatively high, so their structural homogeneity coefficient is higher. Finally, the privacy classification results of the two classifiers are weighted to output the final privacy classification result of the unknown node.

[0032] In a specific implementation of the present invention, in the adversary model based on feature homogeneity and structural homogeneity, in order to obtain the privacy classification result under the adjacency pattern, in this embodiment, GraphSAGE is also used as the convolutional layer in the classifier based on the adjacency pattern. By stacking two layers of GraphSAGE convolutional layers, a classifier based on the adjacency pattern is designed, which can aggregate the information of the second-order neighborhood of each node. Finally, the classifier based on the adjacency pattern is designed as the following formula:

[0033]

[0034] where X represents the node feature matrix, A′ represents the updated adjacency matrix, and f p (·) represents the graph neural network based on global graph aggregation, that is, the classifier based on the adjacency pattern, represents the privacy classification result of the node under the adjacency pattern.

[0035] In order to obtain the privacy classification result under the structural similarity pattern, in this embodiment, GIN is used as the convolutional layer in the classifier based on the structural similarity pattern. For each node, GIN essentially aggregates the representations of itself and its neighbors by inputting them into a multi-layer perceptron MLP. In this embodiment, the present invention designs a classifier based on structural similarity by stacking two layers of GIN convolutional layers. Through multi-layer aggregation and pooling operations, feature representations are extracted in the second-order subgraph of each node, effectively capturing the node substructure information. Finally, the classifier based on structural similarity is designed as the following formula:

[0036]

[0037]

[0038] where X i , A i respectively represent the feature matrix and the adjacency matrix of the second-order subgraph of node i, SE i represents the structural encoding in the subgraph, represents matrix concatenation, and f s(·) represents a graph neural network based on subgraph aggregation, i.e., a classifier based on structural similarity. Represents the privacy classification result of node i in the mode of structural similarity.

[0039] In this embodiment, the structure encoding is obtained by performing random walks in the subgraph, and its calculation method is as follows:

[0040] SE i =[Diag(A i D i -1 ),Diag[A i D i -2 ,…,Diag[A i D i -k ,

[0041] Among them, D i represents the diagonal matrix composed of the degrees of each node in the second-order subgraph, Diag(·) represents the vector composed of diagonal elements, and k is a constant.

[0042] In this embodiment, the feature homogeneity coefficient and the structure homogeneity coefficient are calculated by using the privacy classification result of the unknown node as its pseudo-label:

[0043]

[0044]

[0045]

[0046]

[0047] Among them, N(i) represents the first-order neighbors of node i, D(i) represents the degree of node i, represents the privacy classification results of nodes i and j, #{·} represents the cardinality of the set, h i represents the feature homogeneity coefficient of the privacy classification result of node i before normalization, s i represents the structure homogeneity coefficient of the privacy classification result of node i before normalization, represents the feature homogeneity coefficient of the privacy classification result of node i after normalization, represents the structure homogeneity coefficient of the privacy classification result of node i after normalization.

[0048] In particular, at initialization, h i =0.5, s i =0.5.

[0049] Finally, the model weights the privacy classification results using the weight scores to obtain the final privacy classification result of the node, specifically as follows:

[0050]

[0051] Among them, represents the privacy classification result of node i in the mode indicating adjacency. The application of the weight scores effectively measures the degree of compliance of different nodes with different modes, thereby reflecting this difference in the importance of the privacy classification results under different modes. represents the final privacy classification result of node i.

[0052] In adversarial training, the optimization objective of the defense model is to both reduce the prediction ability of the adversary model and minimize the proportion of perturbations, that is, the sampled edges should be as consistent with the original graph as possible. The final sampling probability is obtained through adversarial training, and then a new graph is sampled.

[0053] In adversarial training, the objective of the defense model is to weaken the classification accuracy of the adversary model and prevent the sampling probability from being too small, while the objective of the adversary model is to improve the classification accuracy, which can be specifically expressed as:

[0054]

[0055] Among them, g d , g a represent the patterns of the learnable parameters of the defense model and the adversary model, L A and L util are the optimization objectives, representing the classification accuracy of the adversary model and the magnitude of the perturbation ratio respectively. γ, λ are constants used to control the balance of the two optimization objectives of L A and L util . L A and L util are specifically as follows:

[0056]

[0057]

[0058] Among them, V and E respectively represent the set of nodes and the set of edges of the graph, C represents the number of categories of privacy features, Z i represents the true privacy feature of node i, |·| represents the cardinality of the set, represents the probability that the final privacy classification result of node i is the c-th category of privacy feature, A ij represents the edge (i, j) information of nodes i and j in the original graph data, A ij = 1 indicates that there is an edge between nodes i and j in the original graph data, A ij= 0 indicates that there is no edge between nodes i and j in the original graph data; represents the sampling probability of edge (i, j) in the graph data.

[0059] After the adversarial training ends, reparameterization is performed using the final sampling probability of each edge to obtain the final sampling result of each edge, and the un-sampled edges are recorded. The relationships between users corresponding to the un-sampled edges are deleted from the input original data to obtain the data after privacy feature protection, so that the new graph topology information corresponding to the data after privacy feature protection will not cause the leakage of users' privacy features.

[0060] In summary, the present invention constructs graph data by given user information and the relationships between users, analyzes the relationship between node privacy features in the graph and two distribution patterns, trains end-to-end graph neural network privacy protection, captures adjacency information and structural information in the graph, and can effectively protect the privacy features of users corresponding to nodes by perturbing the topology information in the graph data.

[0061] In this embodiment, a privacy feature protection system based on a graph neural network is also provided. This system is used to implement the above embodiments, and those that have been described will not be repeated. The following terms "module", "unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the system described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible.

[0062] The described system includes:

[0063] A data input module, which is used to give user information and the relationships between users as data to be desensitized;

[0064] A data processing module, which is used to regard each given user in the data to be desensitized as a node in the network topology structure, use the user information as node features to construct a node feature matrix; construct an adjacency matrix according to the relationships between users in the data to be desensitized to obtain the original graph data;

[0065] A defense model module based on a graph neural network, which is used to calculate the sampling probability of each edge in the original graph data, sample the edges according to the sampling probability, update the adjacency matrix according to the sampled edges to obtain the perturbed graph data;

[0066] An adversary model module based on adjacency and structural similarity patterns, which is used to perform privacy feature classification prediction on the perturbed graph data;

[0067] An adversarial training module, which is used to perform adversarial training on the defense model and the adversary model;

[0068] A data desensitization module, which is used to obtain the final sampling probability of each edge according to the adversarial training result, sample the edges according to the final sampling probability, and record the unsampled edges; delete the relationships between users corresponding to the unsampled edges from the data to be desensitized to obtain the data after privacy feature protection.

[0069] In this embodiment, the adversary model module based on the adjacency and structural similarity patterns includes:

[0070] A graph neural network unit based on full-graph aggregation, which is represented as follows:

[0071]

[0072] A graph neural network unit based on subgraph aggregation, which is represented as follows:

[0073]

[0074]

[0075] Among them, f p (·) represents a graph neural network based on full-graph aggregation, and f s (·) represents a graph neural network based on subgraph aggregation, X represents the node feature matrix, A′ represents the updated adjacency matrix, represents the privacy classification result under the adjacency pattern; X i , A i respectively represent the feature matrix and adjacency matrix of the 2-hop subgraph of node i, and X′ i represents the result of concatenating the feature matrix of the 2-hop subgraph of node i with the structure encoding, represents the privacy classification result of node i under the structural similarity pattern, and SE i represents the structure encoding obtained by performing a random walk in the subgraph of node i.

[0076] For the implementation processes of the functions and roles of each module in the above system, please refer to the implementation processes of the corresponding steps in the above method for details, which will not be elaborated here. For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The system embodiments described above are only illustrative. The modules described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0077] Embodiments of the system of the present invention can be applied to any device with data processing capabilities, and such a device with data processing capabilities can be a device or apparatus such as a computer. The system embodiments can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation.

[0078] This embodiment verifies the implementation effect of the present invention through a specific experiment.

[0079] (1) Data description

[0080] The data used in this experiment comes from the publicly available dataset Pokec-n obtained from the Pokec social network. This dataset consists of 66,569 nodes and 729,129 edges. Each node represents a user, and an edge connecting two nodes means that the two users are friends. At the same time, the information of each user includes age, hobbies, work type, education level, etc. In this dataset, age is the privacy feature to be classified and is divided into four categories: young (18 - 24), prime (25 - 34), middle-aged (35 - 49), and elderly (>49), and the corresponding label definitions are 1, 2, 3, 4, and other users are marked as background nodes with 0. The number of users in each category is 17476, 14028, 5635, 726. To test the degree of damage of the defense model to the topological structure, this experiment also tests the prediction performance of the new graph structure after perturbation on the labels of the dataset itself. The label of Pokec-n is the working field of the node and is divided into 2 categories.

[0081] (2) Comparative experiment.

[0082] To comprehensively verify the effectiveness of the model of the present invention, this experiment compares it with several different types of baseline models. The baseline models compared with the model of the present invention include:

[0083] Random sampling (Rand.): Random sampling randomly samples the edges in the graph with a fixed probability to obtain a perturbed graph.

[0084] Degree-based sampling (Deg.): Degree-based sampling calculates the sum of the degrees of the nodes corresponding to each edge and samples the edges in the graph with the sum of the degrees as the weight to obtain a perturbed graph.

[0085] Eigenvalue-based sampling (Eigen.): Eigenvalue-based sampling first calculates the sum of the eigenvalues of the nodes corresponding to each edge and samples the edges in the graph with the sum of the eigenvalues as the weight.

[0086] Sampling by betweenness: First, calculate the sum of the betweenness of the nodes corresponding to each edge, and sample the edges in the graph with the sum of betweenness as the weight.

[0087] The test results are shown in Table 1. Among them, "privacy feature" represents the prediction performance of the adversary model on the privacy features of the new graph after perturbation. The lower it is, the better the defense effect. "Downstream application" represents the prediction performance of the node labels on the new graph after perturbation. The higher it is, the lower the degree of damage to the topological structure, indicating that the defense model is better.

[0088] Test results of the dataset (%)

[0089]

[0090] As can be seen from Table 1, the prediction performance of the adversary model on the privacy features of the new graph perturbed by the method of the present invention is the lowest, and the prediction performance of the labels in the downstream application is the highest through the present invention. This shows that the method of the present invention performs better than other methods participating in the comparison, and the improvement in the inclusion effect is obvious. This result verifies the effectiveness of the present invention in protecting the privacy features of users in the form of graph data and can achieve a good effect.

[0091] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of patent protection. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A privacy feature protection method based on graph neural network, characterized in that, it includes: Given user information and the relationships between users as data to be desensitized; Regarding each given user in the data to be desensitized as a node in the network topology structure, using the user information as node features, and constructing a node feature matrix; Constructing an adjacency matrix according to the relationships between users in the data to be desensitized to obtain the original graph data; calculating the sampling probability of each edge in the original graph data through a defense model based on graph neural network, sampling the edges according to the sampling probability, and updating the adjacency matrix according to the sampled edges to obtain the perturbed graph data; Performing privacy feature classification prediction on the perturbed graph data through an adversary model based on adjacency and structural similarity patterns; the adversary model based on adjacency and structural similarity patterns includes a graph neural network based on full graph aggregation and a graph neural network based on subgraph aggregation, which are expressed as follows: Among them, f p (·) represents a graph neural network based on global graph aggregation, and f s (·) represents a graph neural network based on subgraph aggregation. X represents the node feature matrix, and A′ represents the updated adjacency matrix. represents the privacy classification result in the adjacency pattern; X i , A i respectively represent the feature matrix and the adjacency matrix of the second-order subgraph of node i. X′ i represents the result of concatenating the feature matrix of the second-order subgraph of node i and the structure encoding. represents the privacy classification result of node i in the structural similarity pattern. SE i represents the structure encoding obtained by performing random walks in the subgraph of node i. Performing privacy feature classification prediction on the perturbed graph data includes calculating the feature homogeneity coefficient and the structural homogeneity coefficient corresponding to the perturbed graph data, and using the coefficients to weight the privacy classification results to obtain the final privacy classification results of the nodes. The calculation formula is: Among them, represents the final privacy classification result of node i, represents the privacy classification result of node i under the adjacency pattern, represents the feature homogeneity coefficient of the privacy classification result of node i after normalization, represents the structural homogeneity coefficient of the privacy classification result of node i after normalization; Performing adversarial training on the defense model and the adversary model to obtain the final sampling probability of each edge, sampling the edges according to the final sampling probability, and recording the unsampled edges; deleting the relationships between users corresponding to the unsampled edges from the data to be desensitized to obtain the data after privacy feature protection.

2. The privacy feature protection method based on graph neural network according to claim 1, characterized in that, Calculating the sampling probability of each edge in the original graph data through a defense model based on graph neural network, and sampling the edges according to the sampling probability. The formula is as follows: H d = f d (X, A) Among them, X represents the node feature matrix, A represents the adjacency matrix, and f d (·) represents the graph neural network, and H d represents the set of node representations, represents matrix concatenation, MLP is the multi-layer perceptron, and σ is the sigmoid function, represents the representation of node i, represents the sampling probability of the edge (i, j) in the graph data, and A′ ij represents the sampling result of the edge (i, j) between nodes i and j, and Gumbel-Softmax represents the reparameterization method.

3. The privacy feature protection method based on graph neural network according to claim 1, characterized in that, The calculation formulas for the feature homogeneity coefficient and the structural homogeneity coefficient are: Among them, N(i) represents the first-order neighbors of node i, and D(i) represents the degree of node i. represents the final privacy classification results of nodes i and j, #{·} represents the cardinality of a set, h i represents the feature homogeneity coefficient of the privacy classification result of node i before normalization, s i represents the structural homogeneity coefficient of the privacy classification result of node i before normalization.

4. The privacy feature protection method based on graph neural network according to claim 1, characterized in that, The objective function of the adversarial training is: Among them, g d and g a represent the learnable parameters of the defense model and the learnable parameters of the adversary model, and L D represents the optimization objective of adversarial training. L A is the optimization objective representing the classification accuracy of the adversary model, and L util is the optimization objective representing the size of the perturbation ratio. γ and λ are constants used to control the balance between the two optimization objectives.

5. The privacy feature protection method based on graph neural network according to claim 4, characterized in that, The optimization objective representing the classification accuracy of the adversary model is specifically: where \(V\) represents the set of nodes of the graph data, \(C\) represents the number of categories of privacy features, and \(Z\) i represents the true privacy feature of node \(i\), \(|\cdot|\) represents the cardinality of the set, represents the probability that the final privacy classification result of node \(i\) is the \(c\)-th category of privacy features.

6. The privacy feature protection method based on graph neural network according to claim 4, characterized in that, The optimization objective representing the perturbation ratio size is specifically: Among them, E represents the edge set of the graph data, and A ij represents the edge (i, j) information of nodes i and j in the original graph data, and A ij = 1 indicates that there is an edge between nodes i and j in the original graph data, and A ij = 0 indicates that there is no edge between nodes i and j in the original graph data; represents the sampling probability of the edge (i, j) in the graph data.

7. A privacy feature protection system based on graph neural network, characterized in that, it includes: A data input module, which is used to give user information and the relationships between users as data to be desensitized; A data processing module, which is used to regard each given user in the data to be desensitized as a node in the network topology structure, use the user information as node features, and construct a node feature matrix; Constructing an adjacency matrix according to the relationships between users in the data to be desensitized to obtain the original graph data; The defense model module based on graph neural network, which is used to calculate the sampling probability of each edge in the original graph data, sample the edges according to the sampling probability, update the adjacency matrix according to the sampled edges, and obtain the perturbed graph data; The adversary model module based on adjacency and structural similarity patterns, which is used to perform privacy feature classification prediction on the perturbed graph data; the adversary model based on adjacency and structural similarity patterns includes a graph neural network based on global graph aggregation and a graph neural network based on subgraph aggregation, which are expressed as follows: Among them, f p (·) represents a graph neural network based on global graph aggregation, f s (·) represents a graph neural network based on subgraph aggregation, X represents the node feature matrix, and A′ represents the updated adjacency matrix, represents the privacy classification result in the adjacency pattern; X i and A i respectively represent the feature matrix and the adjacency matrix of the second-order subgraph of node i, and X′ i represents the result after splicing the feature matrix of the second-order subgraph of node i and the structure encoding, represents the privacy classification result of node i in the structural similarity pattern, SE i represents the structure encoding obtained by performing random walks in the subgraph of node i; The privacy feature classification prediction on the perturbed graph data includes calculating the feature homogeneity coefficient and the structural homogeneity coefficient corresponding to the perturbed graph data, and weighting the privacy classification results using the coefficients to obtain the final privacy classification results of the nodes. The calculation formula is: Among them, represents the final privacy classification result of node i, represents the privacy classification result of node i under the adjacency pattern, represents the feature homogeneity coefficient of the privacy classification result of node i after normalization, represents the structural homogeneity coefficient of the privacy classification result of node i after normalization; The adversarial training module, which is used to perform adversarial training on the defense model and the adversary model; The data desensitization module, which is used to obtain the final sampling probability of each edge according to the adversarial training results, sample the edges according to the final sampling probability, and record the unsampled edges; delete the relationships between users corresponding to the unsampled edges from the data to be desensitized to obtain the data after privacy feature protection.

8. The privacy feature protection system based on graph neural network according to claim 7, wherein, the adversary model module based on adjacency and structural similarity patterns includes: a graph neural network unit based on global graph aggregation, which is expressed as follows: a graph neural network unit based on subgraph aggregation, which is expressed as follows: Among them, f p (·) represents a graph neural network based on global graph aggregation, and f s (·) represents a graph neural network based on subgraph aggregation. X represents the node feature matrix, and A′ represents the updated adjacency matrix. represents the privacy classification result in the adjacency pattern; X i , A i respectively represent the feature matrix and the adjacency matrix of the second-order subgraph of node i. X′ i represents the result of concatenating the feature matrix of the second-order subgraph of node i and the structure encoding. represents the privacy classification result of node i in the structural similarity pattern. SE i represents the structure encoding obtained by performing a random walk in the subgraph of node i.

Citation Information

Patent Citations

  • Privacy protection method for social network relationship prediction

    CN112199728A

  • System and method for machine learning architecture with privacy-preserving node embeddings

    US20200356858A1