A method for protecting user privacy based on neural pathways, an electronic device, and a medium

By performing specific transformation and processing on the graph neural network model and hiding the original graph structure information, the problem of privacy leakage risk in graph neural network is solved, and the privacy protection ability and credibility of the model are improved.

CN115659387BActive Publication Date: 2025-06-17ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211184386.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-27
Publication Date
2025-06-17
Estimated Expiration
2042-09-27

AI Technical Summary

Technical Problem

The existing graph neural networks have the risk of privacy leakage. Attackers can infer sensitive information in graph data through extracted low-dimensional features, resulting in user privacy being stolen.

Method used

By reasonably transforming the input graph data, hiding the original graph structure information, improving the robustness of graph neural networks for privacy theft attacks. Specific steps include pre-training the graph neural network model, acquiring key neural pathways, building a mask matrix, selecting backbone graphs and non-backbone graphs, performing node embedding extraction, and weighting combinations to obtain privacy-powered node embedding.

Benefits of technology

It improves the robustness of the graph neural network model for privacy theft attacks, enhances the credibility of the model and the ability to protect private data, and ensures the security of user critical privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115659387B_ABST
    Figure CN115659387B_ABST
Patent Text Reader

Abstract

The present invention discloses a user privacy protection method based on neural pathways, including: (1) using a pre-trained graph convolutional network model to conduct neuron tests on training data to obtain key neural pathways; (2) obtaining important graph structures related to the performance of the main task, thereby extracting the backbone graph; (3) generating non-backbone graphs that are positively correlated with the main task based on the obtained key neural pathways; (4) using the pre-trained graph convolutional network model to extract node embeddings of the backbone graph and the non-backbone graph respectively; (5) performing weighted combination on the obtained node embeddings corresponding to the backbone graph and the non-backbone graph to obtain privacy-robust node embeddings. The method of the present invention can effectively reduce the privacy leakage risk of the graph neural network model, improve the robustness of the graph neural network model against privacy inference attacks, and enhance the privacy data protection ability of the graph neural network model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network security, and particularly relates to a user privacy protection method, an electronic device, and a medium based on a neural pathway. Background Art

[0002] With the development of the Internet, a vast amount of data is generated every moment. Among these data, graph-structured data is common, such as social networks, traffic networks, financial transaction networks, etc. Taking social networks as an example, nodes are usually some social software users or organizations, and the edges can be the friendship relationships existing between user nodes. For such data, how to efficiently analyze and infer a vast amount of graph-structured data reasonably with less manual annotation work and put it into actual production is a key issue. Compared with image data, graph-structured data is often irregular, so convolutional networks and the like commonly used in the image field are difficult to effectively process graph-structured data. Due to its powerful graph-structured relationship extraction ability, graph neural networks have achieved remarkable success in the field of graph representation learning. Graph neural networks can map graph-structured data into a low-dimensional space for downstream application processing (such as node classification, link prediction, community discovery, etc.), and finally be applied to actual systems (such as recommendation systems).

[0003] With the development of research on graph neural networks, the privacy and security issues of graph neural networks have attracted more and more attention from researchers. Research shows that existing graph neural networks have the risk of privacy leakage, that is, attackers can infer sensitive information in graph data through the low-dimensional features extracted by graph neural networks, thereby obtaining private data. In real life, the consequences of such privacy leakage are serious. For example, fraudsters can infer users' key privacy (such as users' income, transaction relationships between users, etc.) through the prediction results of financial models, thereby stealing users' personal information and then carrying out illegal activities such as targeted telecommunications fraud, endangering social security. Therefore, such risks of privacy leakage will pose a threat to people's daily lives. And how to improve the privacy data protection ability of the model has also become a research hotspot.

[0004] In response to the above problems, researchers have proposed different defense strategies. For example, introducing the differential privacy mechanism, adding differential noise to the data to mislead attackers so that they cannot accurately infer private data, but this method may have an adverse impact on the performance of the model; in addition, introducing adversarial noise during the training process and making a trade-off between the main task performance and the risk of privacy leakage, this method will also have a certain impact on the model performance. Therefore, how to effectively defend against privacy stealing attacks on graph neural networks and maintain the model performance as much as possible has important practical significance for improving the privacy security, data protection ability of graph neural networks and their credibility in actual systems. Summary of the Invention

[0005] In view of the risk of privacy leakage in graph neural network models, to improve the robustness of graph neural network models against privacy stealing attacks, enhance the credibility of graph neural network models and the ability to protect privacy data, the present invention provides a user privacy protection method based on neural pathways. This method reasonably transforms the input graph data, and hides the original graph structure information in the node embeddings extracted by the model, thereby improving the robustness of graph neural networks against privacy stealing attacks, and further ensuring the security of users' key privacy on the premise of maintaining the performance of graph neural network models.

[0006] To implement the above invention, the technical solution provided by the present invention is as follows: In the first aspect of the embodiments of the present invention, a user privacy protection method based on neural pathways is provided, and the method specifically includes the following steps:

[0007] Step 1, pre-train a graph neural network model, conduct neuron tests and statistics on the training data, and obtain key neural pathways related to node labels;

[0008] Step 2, fix the parameters of the graph convolutional network model pre-trained in Step 1, add a mask matrix M to the graph adjacency matrix, and use the gradient descent algorithm to train the mask matrix M to obtain the final mask matrix M final , and according to a preset threshold δ, select Q edges from the final mask matrix M final as the backbone graph A b ;

[0009] Step 3, based on the graph convolutional network model pre-trained in Step 1, obtain the hidden representation P of the nodes at the L-th layer L , decode the hidden representation into an approximate adjacency matrix A', construct a loss function based on the key neural pathways obtained in Step 1, calculate the edge gradient matrix based on the approximate adjacency matrix A' and the loss function, and select the top Q edges with the largest gradient values from the edge gradient matrix as the non-backbone graph A n ;

[0010] Step 4, use the graph neural network model pre-trained in Step 1 to extract node embeddings from the backbone graph obtained in Step 2 and the non-backbone graph obtained in Step 3 respectively;

[0011] Step 5, adjust the weights, and perform weighted combination on the node embeddings corresponding to the backbone graph and the non-backbone graph obtained in Step 4 to obtain node embeddings with privacy robustness.

[0012] A second aspect of an embodiment of the present invention provides an electronic device, including a memory and a processor, the memory being coupled to the processor; wherein, the memory is used for storing program data, and the processor is used for executing the program data to implement the above-mentioned user privacy protection method based on neural pathways.

[0013] A third aspect of an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned user privacy protection method based on neural pathways is implemented.

[0014] Compared with the prior art, the beneficial effects of the present invention at least include:

[0015] The user privacy protection method based on neural pathways provided by the present invention, first, conducts neuron testing on training data through a pre-trained graph convolutional network model and makes statistics to obtain key neural pathways related to node labels; secondly, constructs a graph interpretation model, obtains important graph structures related to the main task performance, and extracts the backbone graph; thirdly, constructs a graph generator model based on the key neural pathways to generate non-backbone graphs; then, uses a pre-trained model to extract embeddings from the backbone graph and the non-backbone graph respectively; finally, performs weighted combination on the embeddings corresponding to the obtained backbone graph and non-backbone graph to obtain privacy-robust node embeddings. It improves the robustness of the graph neural network model against privacy theft attacks, improves the credibility of the graph neural network model and the protection ability for privacy data. Description of the Drawings

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0017] Figure 1 It is a schematic diagram of the overall framework of the user privacy protection method based on neural pathways. Detailed Embodiments

[0018] To make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings.

[0019] Refer to Figure 1 , an embodiment of the present invention proposes a user privacy protection method based on neural pathways, and the method includes the following steps:

[0020] Step 1: Pre-train a graph neural network model, conduct neuron testing and statistics on training data, and obtain key neural pathways related to node labels.

[0021] AsFigure 1 As shown, in this example, the graph neural network model takes the graph convolutional network model as an example. First, use the pre-trained graph convolutional model to perform neuron tests on the test nodes to obtain the key neural pathways. Taking the L-layer graph convolutional network model as an example, the key neurons in the last layer related to node t can be defined as:

[0022]

[0023] where is the neuron output of node t in the L-layer graph convolutional network model. denotes taking the th neuron in the k L For the neuron output Z l of the l-th layer (l ≤ L), it can be expressed as:

[0024]

[0025] where A is the adjacency matrix of the input graph, I N is the diagonal matrix, and N represents the number of nodes in the graph. is the degree matrix of A + I N , P l is the hidden representation / node embedding of the l-th layer (P 0 = X, where X is the input node feature), and W l-1 is the model weight matrix of the l - 1-th layer.

[0026] The l-th layer neuron can be defined as:

[0027]

[0028] Obtain the neural pathway P t = [k L , k l-1 ,..., k 1 through formulas (1) and (3). Test the neurons of the training data according to the labels, and count the number of nodes with the same activated neural pathway in the set of nodes with the same label. Select the neural pathway with the largest number of nodes as the key neural pathway related to the label where C is the number of node categories in the graph.

[0029] Step 2: Fix the parameters of the pre-trained graph convolutional network model in Step 1, add the mask matrix M to the graph adjacency matrix, and use the gradient descent algorithm to train the mask matrix M to obtain the final mask matrix M final , and according to the preset threshold δ, select Q edges from the final mask matrix M final as the backbone graph A b .

[0030] As Figure 1 shown, first fix the parameters of the pre-trained graph convolutional network model, and add a trainable mask matrix M to the graph adjacency matrix. Then the training objective of the mask matrix M is:

[0031]

[0032] where MI(·) represents mutual information, and f θ *(·) is the pre-trained graph convolutional network model with fixed parameters θ * , A S and X S are the adjacency matrix and its corresponding feature matrix of the r-order subgraph of node t respectively, Y is the prediction result of the model, ⊙ represents element-wise multiplication, and H(·) is the information entropy (H(Y) is a constant). mse(·) represents the mean square error function, and Zl, is the output of the l-th layer neuron of the graph convolutional network model when the input adjacency matrix is A S , A S ⊙M respectively.

[0033] The purpose of formula (4) is to make the training objective of the mask matrix M have the meaning of two parts, that is, ① maximize the mutual information between the model outputs when the input adjacency matrices are A S , A S ⊙M respectively, and ensure that the adjacency matrix after masking the model input can maintain the main task performance as much as possible; ② minimize the output difference of each layer of neurons, that is, make the input A S ⊙M can activate neural pathways with similar activations as when the input is A S .

[0034] Train the mask matrix M using the gradient descent algorithm and process it as follows:

[0035] M final = sigmoid(M) (5)

[0036] Obtain the final mask matrix M final , and according to the preset threshold δ, select Q edges from M final as the backbone graph A b :

[0037]

[0038] Step 3, based on the pre-trained graph convolutional network model in Step 1, obtain the hidden representation P of the node at the L-th layer L, decode the hidden representation into an approximate adjacency matrix A', construct a loss function based on the key neural pathways obtained in step 1, calculate the edge gradient matrix based on the approximate adjacency matrix A' and the loss function, and select the first Q edges with the largest gradient value from the edge gradient matrix as the non-backbone graph A n .

[0039] like Figure 1 As shown in Figure 2, the original image (A, X) is input into the pre-trained graph convolutional network model to obtain the node L-th layer hidden representation P L . First decode it into an approximate adjacency matrix A':

[0040] A'=sigmoid(P L ·(P L ) T ) (7)

[0041] in(·) T is the transpose operation.

[0042] Then the key neural pathway P c Construct neuron label vector according to the number of layers of the model Among them, K l No. The dimension is 1, and the rest are 0. The loss function is further constructed as follows:

[0043]

[0044] Among them, C is the number of categories of nodes in the graph, Y nc represents the true label of the nth node in the cth category, Y′ nc represents the predicted output value of the graph neural network model for the nth node in the cth category, |F l |,l∈[1,...,L] is the number of neurons in the lth layer of the model.

[0045] Further calculate the edge gradient matrix, the formula is as follows:

[0046]

[0047] Then select the first Q edges with the largest gradient value as the non-backbone graph A that is positively correlated with the main task n , where Q is the number of edges in the backbone graph obtained in step 2).

[0048] 4) The pre-trained model extracts node embeddings from backbone and non-backbone graphs;

[0049] like Figure 1 As shown, the backbone graph A obtained according to step 2) and step 3) b and non-backbone graph A n, use the pre-trained graph convolutional network model to train the backbone graph A b and non-backbone graph A n Extract the embedding P of node t b and P n , the formula is as follows:

[0050]

[0051] in, is the parameter θ * Fixed pre-trained graph convolutional network model, X is the input node feature.

[0052] 5) Weighted combination of node embeddings to obtain privacy-robust node embeddings;

[0053] like Figure 1 As shown, for the node embedding P obtained in step 4) b and P n , and perform weighted combination to obtain privacy-robust node embedding:

[0054] P robust =(1-η)P b +ηP n (11)

[0055] Where η is the weight factor. Adjusting the value of η can control the relationship between backbone graph and non-backbone graph in P robust The proportion in is used to adjust the strength of privacy protection. When η is larger, the privacy protection ability is stronger.

[0056] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0057] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1a device for the functions specified in one or more boxes.

[0058] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one Figure 1 process or more processes and / or boxes Figure 1 a box or more boxes.

[0059] These computer program instructions may also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 process or more processes and / or boxes Figure 1 a box or more boxes.

[0060] The above embodiments are only used to illustrate the design concept and features of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made according to the principles and design concepts disclosed by the present invention are within the protection scope of the present invention.

Claims

1. A method for protecting user privacy based on neural pathways, characterized in that, The method specifically includes the following steps: Step 1, pre-train a graph neural network model, conduct neuron tests and statistics on the training data, and obtain key neural pathways related to node labels; Among them, Step 1 is specifically: The key neurons of the last layer related to node t are defined as: Among them, is the neuron output of node t in the graph convolutional network model at the L-th layer; denotes taking the k-th L neuron in, and for the neuron output Z of the l-th layer l is expressed as: Among them, A is the adjacency matrix of the input graph, and I N is the diagonal matrix, and N represents the number of nodes in the graph; is the degree matrix of A + I N , P l is the hidden representation / node embedding of the l-th layer, and W l-1 is the model weight matrix of the (l - 1)-th layer; The neurons of the l-th layer can be defined as: Obtain the neural pathway P through formula (1) and formula (3). t = [k L , k l-1 ,..., k 1 ; Test the neurons of the training data according to the labels, and count the number of nodes with the same activated neural pathway in the node set with the same label. Select the neural pathway with the largest number of nodes as the key neural pathway related to the label. Where C is the number of categories of nodes in the graph; Step 2: Fix the parameters of the pre-trained graph convolutional network model in Step 1, add a mask matrix M to the graph adjacency matrix, and use the gradient descent algorithm to train the mask matrix M to obtain the final mask matrix M final , and according to a preset threshold δ, select Q edges from the final mask matrix M final as the backbone graph A b ; Step 3: Based on the pre-trained graph convolutional network model in Step 1, obtain the hidden representation P of the nodes at the L-th layer L , decode the hidden representation into an approximate adjacency matrix A', construct a loss function based on the key neural pathways obtained in Step 1, calculate the edge connection gradient matrix based on the approximate adjacency matrix A' and the loss function, and select the top Q edges with the largest gradient values from the edge connection gradient matrix as the non-backbone graph A n ; Step 5, use the graph neural network model pre-trained in Step 1 to extract node embeddings from the backbone graph obtained in Step 2 and the non-backbone graph obtained in Step 3 respectively; Step 6, adjust the weights, and perform weighted combination on the node embeddings corresponding to the backbone graph and the non-backbone graph obtained in Step 4 to obtain node embeddings with privacy robustness.

2. The method for protecting user privacy based on neural pathways according to claim 1, characterized in that, The graph neural network model is a graph convolutional network model.

3. The method for protecting user privacy based on neural pathways according to claim 2, characterized in that, Step 2 is specifically: First, fix the parameters of the graph neural network model pre-trained in Step 1, and add a trainable mask matrix M to the graph adjacency matrix. Then the training objective of the mask matrix M is: where MI(·) represents mutual information, is the parameter θ * a fixed pre-trained graph convolutional network model, A S and X S are respectively the adjacency matrix and its corresponding feature matrix of the r-th order subgraph of node t, Y is the prediction result of the model, ⊙ represents element-wise multiplication, H(·) is the information entropy; mse(·) represents the mean square error function, Z l , is the output of the l-th layer neuron of the graph convolutional network model when the input adjacency matrix is respectively A S , A S ⊙M; Train the mask matrix M using the gradient descent algorithm, and obtain the final mask matrix M after training final , and according to the preset threshold δ, select Q connected edges from M final as the backbone graph A b , the formula is as follows:

4. The method for protecting user privacy based on neural pathways according to claim 2, characterized in that, Step 3 is specifically: Input the original graph (A, X) into the pre-trained graph convolutional network model to obtain the hidden representation P of the nodes at the L-th layer, where A is the adjacency matrix of the input graph and X is the input node feature. First, decode it into an approximate adjacency matrix A': L , A' = sigmoid(P L ·(P L ) T ) (7) where (·) T is the transpose operation; Then, the key neural pathway P obtained in step 1 c Construct a neuron label vector according to the number of layers of the model where K l The dimension is 1 and the rest are 0; further construct the loss function as follows: where C is the number of categories of nodes in the graph, and Y nc represents the true label of the n-th node in the c-th category, and Y' nc represents the predicted output value of the graph neural network model for the n-th node in the c-th category, |F l |, l ∈ [1,..., L] is the number of neurons in the l-th layer of the model; Calculate the edge connection gradient matrix, and the formula is as follows: Then, select the top Q edges with the largest gradient values as the non-backbone graph A n , where Q is the number of edges in the backbone graph obtained in step 2.

5. The method for protecting user privacy based on neural pathways according to claim 2, wherein, Step 4 is specifically: Use the pre-trained graph neural network model in Step 1 to separately perform the extraction of node embeddings on the backbone graph A obtained in Step 2 b and the non-backbone graph A obtained in Step 3 n The formula is as follows: Among them, is the parameter θ * a fixed pre-trained graph convolutional network model, and X is the input node feature.

6. The method for protecting user privacy based on neural pathways according to claim 2, wherein, In step 5, for the node embeddings P b and P n corresponding to the backbone graph and the non-backbone graph obtained in step 4, a weighted combination is performed to obtain a node embedding P robust with privacy robustness. The calculation formula is as follows: P robust = (1 - η)P b + ηP n (11) Where η is a weight factor.

7. The method for protecting user privacy based on neural pathways according to claim 6, wherein, The larger the weight factor η, the stronger the privacy protection ability of the node embeddings with privacy robustness.

8. An electronic device, comprising a memory and a processor, wherein, The memory is coupled to the processor; among them, the memory is used to store program data, and the processor is used to execute the program data to implement the neural-pathway-based user privacy protection method described in any one of the above claims 1-7.

9. A computer-readable storage medium, having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the neural-pathway-based user privacy protection method described in any one of claims 1-7.

Citation Information

Patent Citations

  • Leukocyte five-classification method based on an improved attention convolutional neural network

    CN113887503A

  • Sensitive link privacy protection method based on graph embedding

    CN114662143A