Complex subgraph clustering relation estimation method for graph network analysis under privacy protection
By using two rounds of local perturbation and shuffling models in graph network analysis, combining uniform random sampling and random response mechanism for data perturbation, the problem of poor data utility in the graph network analysis below is solved, and efficient graph network analysis and structural optimization are achieved.
Patent Information
- Application Number
- CN202411966733.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-27
AI Technical Summary
Under privacy protection, it is difficult for the prior art to realize efficient graph network analysis while maintaining the local and overall structural characteristics of the graph network. At the same time, graph data processing based on local differential privacy technology is prone to introduce too much noise when taking into account both local and global features, resulting in poor data utility.
A complex subgraph clustering relationship estimation method for network analysis of the following privacy graph is proposed. Through the combination of two rounds of local perturbation and shuffling models, the global privacy budget is converted into the privacy budget required for local perturbation, and data perturbation is carried out through uniform random sampling and random response mechanisms, and the network structure is finally optimized through noise average and threshold settings.
While ensuring privacy protection, the data utility of graph network analysis is improved, the pyramid count in graph network can be estimated more accurately, the calculation overhead is reduced, and the network structure is optimized to retain global features.
Smart Images

Figure CN120046181A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of graph structure data collection, and more specifically, to a method for estimating complex subgraph clustering relationships in graph network analysis under privacy protection. Background Art
[0002] Graph data essentially uses a graph as the basic data model, which regards entities and relationships as nodes and edges. Graph data has been widely applied in fields such as social network analysis, bioinformatics, network security, and financial investment, and is one of the important tools for data analysis and decision optimization. Graph data has currently been widely used in various fields of life. In addition to representing entity attribute information through nodes, it can also express the link relationships between nodes through edges. Because a large amount of sensitive information is contained in graph data, once leaked, the consequences are extremely serious.
[0003] To protect the privacy of data interaction in graph data calculation, methods such as data encryption, permission control, anonymization processing, and secure transmission are generally used in current work. Among them, the method based on differential privacy is relatively common. It protects privacy by introducing random noise into real data, making it impossible for attackers to directly distinguish the values between any two individuals, thereby protecting privacy. In the traditional differential privacy model, it is assumed that there is a trusted third party, which acts as the collector in data transmission. Due to the restrictive requirements of the trusted third party, to solve this problem, the local differential privacy technology (Local Differential Privacy, abbreviated as: LDP) has been proposed and widely applied. LDP does not require the existence of a trusted third party, and users can process data locally, thereby avoiding the leakage of personal information during the data interaction process. While using LDP to ensure the privacy of the published data, it also requires the practicality and availability of the data, and provides a feasible method for privacy-protected data sharing.
[0004] The graph structure data collection task plays an important role in graph network analysis, and its importance lies in providing insights into the network structure, features, and other aspects. At present, most of such tasks under privacy protection are based on subgraph counting and graph clustering. Simple subgraph counting, such as triangles, k-stars, etc., is commonly used at present, which helps to reveal the basic structure of the network and is easy to implement. However, the above-mentioned subgraph counting cannot reveal important local structures and settlements, so graph clustering algorithms are usually used for settlement detection. However, high computational complexity and excessive information required between users are additional problems that need to be solved urgently in graph clustering. Therefore, the current graph structure data collection task under privacy protection cannot have a certain graph network analysis utility while maintaining good local and overall structural features. In addition, although LDP currently best meets the requirements of real-world applications and has extremely high privacy protection effects, after being perturbed by LDP in graph data, each user cannot balance local and global features, so more noise is introduced resulting in poor data utility. Summary of the Invention
[0005] Object of the Invention: The present invention aims to provide a method for estimating complex subgraph clustering relationships in graph network analysis under privacy protection to solve the problems of privacy protection and data utility in the analysis tasks mentioned in the above background technology.
[0006] To achieve the above object of the invention, the technical solution provided by the present invention is as follows:
[0007] A method for estimating complex subgraph clustering relationships in graph network analysis under privacy protection, comprising the following steps:
[0008] S1, the number of users, the privacy budget required in data perturbation, and the adjacency list of each user;
[0009] Among them, the number of users is n, and the global privacy budgets required for data perturbation are relaxation term δ 1 , δ 2 ∈(0,1), and the adjacency list of each user is a 1 , …, a n ∈{0,1} n ;
[0010] S2, based on the shuffle model, respectively convert the global privacy budget into the privacy budgets required for two rounds of local perturbation
[0011] S3, each user calculates its own node degree d i , and obtains the noisy degree d i ′ through perturbation based on the piecewise mechanism of mean estimation, and at the same time sends the perturbed data to the shuffle model;
[0012] S4. The shuffling model uses uniform random sampling perturbation π on n user nodes to obtain the perturbation sequence [π(1), …, π(n)], and then sends it to each user. At the same time, according to the shuffled sequence order, the shuffled noise degree d π(i) ′ is sent to the collector;
[0013] S5. The collector calculates the noise average degree d π(i) ′ based on the shuffled noise degree d avg ′, and sets a threshold d th based on the noise average degree as the basis for subsequent removal of sparse nodes, thereby reducing the computational overhead and improving the data utility;
[0014] S6. The collector recombines the graph nodes with a uniform random permutation to obtain G′, and at the same time samples disjoint node sets and sends them to all users. The i-th node set after random permutation can be expressed as {v σ(i) , v σ(i+1) , v σ(i+2)}, where σ represents the random permutation;
[0015] S7. For each of the above t disjoint node sets, each user calculates whether it is connected to all nodes in a certain node set and gives the corresponding connection situation. For example, for user v k and the node set {v i , v j , v m}, the connection situation can be estimated as c k . In the present invention, the above connection situation is regarded as a conical data structure;
[0016] S8. Each user perturbs the conical data using the random response mechanism and then rearranges it using the shuffling model to obtain the shuffled and perturbed conical data c π(k) ′, and sends it to the collector. π(k) represents the shuffling operation of user v k ;
[0017] S9. For each point in each disjoint node set, calculate its unilateral relationship and bilateral relationship with the other two points, that is, whether it is separately connected to a certain point and whether it is connected to both points. At the same time, perturb the above relationships using the random response mechanism and send them to the collector. For example, for user v i ∈ {v i , v j , v m}, it calculates the unilateral relationships a ij , a im and the bilateral relationship w i , and obtains the perturbed relationships and
[0018] S10. Each disjoint node set estimates whether three nodes can form a triangle according to the unilateral and bilateral relationships obtained in step S9 above, and obtains a triangle estimate. For example, the node set {v i , v j , v m} estimates the triangle
[0019] S11. According to the shuffled and perturbed conical data obtained in step S8 and the triangle estimate obtained in step S10, estimate the triangular pyramid count of the corresponding node set. For example, the node set {v i , v j , v m} estimates the triangular pyramid count according to c π(k) ′ and estimates the triangular pyramid count
[0020] S12. Judge all node sets according to the threshold d th . When the noise degree of all nodes is higher than the threshold d th , include them in the triangular pyramid estimate of the overall network, and calculate the triangular pyramid count estimate in the graph network
[0021] Further, step S1 is specifically as follows:
[0022] The social network graph is represented as G=(V, E), where V represents the set of all users (nodes), and E represents the relationships (edges) between all users. At the same time, in order to better process the data interaction process, the graph G is usually represented by an adjacency matrix, that is, the adjacency matrix is A=(a i,j ) ∈ {0, 1} n×n . Where a i,j = 1 indicates that there is an edge between nodes i and j, denoted as (v i , v j ) ∈ E; a i,j = 0 indicates that there is no edge between nodes i and j. The adjacency matrix A is composed of the adjacency lists of n users, and the adjacency list is represented as a 1 , …, a n ∈ {0, 1} n , where a i = {a i,1 , a i,2 , …, a i,n}, and in particular a i,i = 0.
[0023] Further, step S2 is specifically as follows:
[0024] In the present invention, a shuffle model is combined with local differential privacy to perform privacy amplification. The specific process is that each user v i∈V uses a local perturbation algorithm encrypts personal data to ensure ε L -LDP and user v i sends the perturbed data to the shuffling model, and then the shuffling model randomly shuffles the perturbed data and sends the result to the data collector. In this model, the common assumption is that the shuffling model and the data collector do not collude with each other. Under the assumption, the shuffling model cannot access the perturbed data, and the data collector cannot link the perturbed data to the user.
[0025] There are n users in the node-degree shuffle and n - 3 users in the cone-shaped shuffle. Therefore, the privacy budgets for the two rounds of shuffling are respectively:[[]]
[0026]
[0027] where ShuffleBudget represents the ε L -LDP and (ε,δ)-DP conversion algorithm, ε << ε L .
[0028] Furthermore, step S3 is specifically as follows:[[]]
[0029] Each user calculates the node degree d according to the adjacency list a i i , the node degree d i The calculation process is as follows:[[]]
[0030]
[0031] Uses the piecewise mechanism to perturb the node degree d i , since it requires the data range of each user to be [-1, 1], the data in different ranges is mapped into the range using a function, and its output can be any value on [-C, C]. Specifically, the value range [-C, C] is divided into non-overlapping segments, and the perturbed output value d i ' of the user will be uniformly and randomly generated within the segments of the [-C, C] range, and its probability density function is as follows:[[]]
[0032]
[0033] where r(d i ) = l(d i ) + C - 1.
[0034] where the privacy budget is represents the privacy budget for one round of shuffling. In addition, similarly, the perturbed node degree d i 'Map it back to the normal real number domain and send it to the shuffling model.
[0035] Furthermore, step S5 is specifically as follows:
[0036] The collector obtains all the noise degrees d after shuffling perturbation π(i) ', calculates the average noise degree d avg ', which is expressed as:
[0037]
[0038] According to the average noise degree d avg ', set the threshold d th , and use the parameter c to adjust the threshold size to adapt to the needs of different graph data sets, which is specifically expressed as:
[0039] d th = cd avg '
[0040] Furthermore, step S7 is specifically as follows:
[0041] For disjoint node sets, data construction is required. In this process, each three-node set is regarded as an external node, and other user nodes are regarded as central nodes. For a given node set {v i , v j , v m}, each user v k calculates the constructible conical data {c k | k ∈ I -(i,j,m)} ∈ {0, 1} by determining whether it is connected to the above node set itself. If all are connected, then assign c k = 1, otherwise c k = 0. In subsequent subgraph counting, only the case where three points are simultaneously connected to the user needs to be considered. Therefore, the multi-dimensional conical data is converted into a one-dimensional binary variable.
[0042] Furthermore, step S8 is specifically as follows:
[0043] Each user v k locally performs random response perturbation on the corresponding conical data, and the perturbation result is and sends it to the shuffling model. The shuffling model uses uniform random sampling perturbation π on the node set I -(i,j,m) and obtains the privacy conical data after shuffling and sends it to the collector.
[0044] For the random response mechanism, its core is the flipping probability. For each result c k , there are two cases. One is the probability p of maintaining the original state, and the other is the probability q of flipping the state. The probabilities are as follows:
[0045]
[0046] For example, when c k = 1 and the flipping probability is p, then Conversely
[0047] Furthermore, step S9 is specifically as follows:
[0048] For each user v in the node set {v i , v j , v m}, the user v x needs to calculate the unilateral relationship and the bilateral relationship with other users respectively. First, the user v x calculates the bilateral relationship w x between itself and the node pair {v i , v j , v m}\{v x}; Second, the user v x calculates the unilateral relationships a x between itself and the remaining two nodes y ∈ {v i , v j , v m}\{v x} respectively. Then, the user v xy perturbs the unilateral and bilateral relationships respectively using the randomized response mechanism to obtain x and and and sends them to the collector;
[0049] Furthermore, step S10 is specifically as follows:
[0050] This step needs to calculate three cases, and the triangle estimations are respectively centered on the nodes v i , v j , v m . Its essence is to estimate the cases containing triangles in the node set {v i , v j , v m} from the perspectives of three users respectively. Since the three cases are similar, only one case is written, and the subsequent estimations follow the same formula. For example, the triangle estimation centered on the node v i is:
[0051]
[0052] Therefore, by aggregating the three estimation cases, the node set {v i , v j , v mThe triangular estimate of} is:
[0053]
[0054] where the privacy budget is ε 2 and is shared with the conical data in the second-round shuffle.
[0055] Furthermore, step S11 is specifically as follows:
[0056] Estimate the triangular pyramid count of the corresponding node set according to the shuffled and perturbed conical data obtained in step S8 and the triangular estimate obtained in step S10. In the node set {v i , v j , v m}}, according to c π(k) ′ and estimate the triangular pyramid count which is expressed as:
[0057]
[0058] where the privacy budget is denotes converting (ε 2 , δ 2 )-DP to
[0059] Furthermore, repeat steps S7 to S11 until all the triangular pyramid count estimates in all disjoint node sets are completed.
[0060] Furthermore, step S12 is specifically as follows:
[0061] Judge all node sets according to the threshold d th and set the range where the noise degree of all nodes in the node set is higher than the threshold d th as:
[0062] D + = {i = 1, 4, …, 3t - 2; d th ≤ min(d σ(i) ’, d σ(i+1) ’, d σ(i+2) ’)
[0063] Calculate the triangular pyramid count estimate in the entire graph network according to the selected node range
[0064]
[0065] Beneficial effects: Compared with the existing technologies, the prominent substantive features and significant progress of the present invention are mainly reflected in the following aspects:
[0066] (1) Compared with traditional graph structure calculations, the present invention proposes a more complex and tighter clustering subgraph structure, namely, a conical structure. By reducing the conical data from multi-dimensional to one-dimensional data, the local characteristics of the data are ensured, and the usability of the calculation structure is ensured. At the same time, in order to estimate the overall count of the triangular pyramid, the external loop structure excluding the core conical data is constructed by combining the single-edge and double-edge relationships between nodes, and the data accuracy of the invention is proved.
[0067] (2) Under the setting of local differential privacy, each user only needs to perform all operations locally, and the relevant data sent is perturbed data, which can prevent attackers from obtaining the relationship between user nodes without a trusted third party. At the same time, in order to improve data utility, a two-round shuffling model is added to amplify privacy under the same privacy budget and prevent the introduction of too much noise.
[0068] (3) Before performing graph network analysis, a threshold is constructed using the node noise degree, and it can be adjusted for different data sets, so as to remove sparse nodes and prevent too many false edges introduced by sparse nodes from affecting the estimation of the data. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 It is the overall flow chart of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0070] To elaborate in detail on the technical solutions disclosed in the present invention, the following further elaboration is made in conjunction with the accompanying drawings and specific embodiments.
[0071] The focus of the present invention is to propose a method for estimating the complex subgraph clustering relationship in graph network analysis under privacy protection. In current common subgraph counting research, such as k-star and triangles, triangles can only represent the connectivity of a small number of nodes, and k-star can only represent the radiation of the central node. However, the application scenarios in life are complex and diverse. In some scenarios with closer relationships, it is necessary to consider the connectivity and radiation relationships of nodes together. Then, how to establish and identify a stronger associated structure in the graph, and at the same time, how to meet better data utility under privacy protection are two core issues for the task of accurately analyzing the overall network framework at present.
[0072] Therefore, in order to enable the analysis task to comprehensively consider the connectivity and radiation relationships, the present invention focuses on the subgraph counting of the conical structure, which is mainly composed of the radiation relationship of the central edge and the circular edge relationship of the remaining nodes. The present invention provides a method for estimating the complex subgraph clustering relationship in graph network analysis under privacy protection. The method described in the present invention is also a method for estimating the triangular cone count under privacy (DTCC, Degree-based Triangular Cone Counting) that combines a two-round shuffle model, and proves its related error. For the two-part structures of connectivity and radioactivity in the conical structure, in order to improve the accuracy of subgraph technology, the present invention proposes to convert the multi-dimensional conical structure data into one-dimensional data, and at the same time construct a peripheral node circular structure with single-edge and double-edge relationships, reducing the error introduced when the relationships between multiple nodes change simultaneously. For the diversity characteristics of social networks, the present invention constructs a noise degree distribution based on mean estimation, and uses the average noise degree to select threshold parameters to optimize the network structure, preventing sparse nodes from introducing incorrect information. For the problem of the trade-off between privacy and practicality, the present invention introduces a two-round perturbed shuffle model, ensuring that its data accuracy is better than that of single-round or non-shuffled data under the same privacy budget.
[0073] Specifically, the present invention is a method for estimating the triangular cone count under privacy (DTCC) based on a two-round shuffle model, including the following steps:
[0074] S1. In the initialization stage, the user needs to determine the required parameters and data structures, including the number of users n, and the privacy budgets required for the two data perturbations are respectively The adjacency list of each user is a 1 ,…,a n ∈{0,1} n ;
[0075] S2. Based on the shuffle model, calculate the privacy budgets required for the two rounds of local data perturbations respectively The shuffle algorithm is defined as follows:
[0076] Let and x is the input data set of each user, and x i is the input data of the i-th user, and Let be the local perturbation algorithm, such that guarantees ε L -LDP. At the same time, let be the shuffle algorithm, that is, given a data set x 1:b , calculate and i ∈ [n]. The specific process is to perform uniform random sampling perturbation π on [n], and the output perturbed data is y π(1),…,y π(n) 。
[0077] After the privacy amplification algorithm of the shuffle model, and has a conversion relationship in terms of privacy budget. If to ensure (ε,δ)-DP, for any δ∈[0,1], there is the following relationship:
[0078]
[0079] After passing through the shuffle model, the output data y π(1) ,…,y π(n) ensures (ε,δ)-DP, and ε < ε L 。
[0080] There are n users in the node-degree shuffle and n - 3 users in the conical shuffle. Therefore, the privacy budgets for the two rounds of shuffles are respectively:
[0081]
[0082] where ShuffleBudget is the inverse function of f, thus calculating the privacy budget
[0083] Steps S1 and S2 are the parameter preparation stage, and the parameters required for the following steps are calculated.
[0084] S3. Each user calculates the node degree d according to the adjacency list a i The calculation process of the node degree d i is as follows: i The process calculation is as follows:
[0085]
[0086] Use the piecewise mechanism to perturb the node degree d i , and use the function to map d i into the range [-1,1]. The perturbed output value d i ' of the user will be uniformly and randomly generated within the segments of the range [-C,C], and its probability density function is as follows:
[0087]
[0088] where, r(d i ) = l(d i ) + C - 1.
[0089] where the privacy budget is represents the privacy budget for one round of shuffle. In addition, similarly, the perturbed node degree di 'Map it back to the normal real number domain and send it to the shuffling model;
[0090] S4. The shuffling model uses uniform random sampling perturbation π on n user nodes to obtain the perturbation sequence [π(1), …, π(n)], and then sends it to each user. At the same time, according to the shuffled sequence order, the shuffled noise degree d π(i) 'is sent to the collector;
[0091] S5. The collector obtains all the shuffled and perturbed noise degrees d π(i) ', calculates the average noise degree d avg ', which is expressed as:
[0092]
[0093] According to the average noise degree d avg ', set the threshold d th , and use the parameter c to adjust the threshold size to adapt to the needs of different graph datasets. Specifically, it is expressed as:
[0094] d th = cd avg '
[0095] Subsequently, use the threshold d th to remove sparse nodes in the network, optimize the network structure, and better retain global features;
[0096] Steps S3 - S5 are the process of generating the node noise degree distribution. Use the segmentation mechanism to perturb the node degree, and at the same time add the shuffling model to amplify privacy. Finally, calculate the average noise degree of the network through mean estimation. Set a threshold according to the average noise degree, so as to provide conditions for subsequent removal of sparse nodes, optimize the network structure, and retain global features.
[0097] S6. In the present invention, the algorithm will solve the counting problem on the entire graph network through multiple iterations of single - round triangular pyramid counting. The core of single - round triangular pyramid counting is that for a selected set of nodes {v i , v j , v m}, determine the number of remaining nodes in the network that form a triangular pyramid with it. Therefore, more directly, the node set statistics can be considered to have a total of choices. However, if this selection method is used, the time complexity required for counting is O(n 4 ), and the time complexity of its selection sampling process is O(n 3 ). In addition, the former will introduce multiple repetitions. For example, if {v 1 , v 2 , v 3 , v 4} form a triangular pyramid. For the node set {v 1 , v 2 , v 3} counting and {v 1 , v 2 , v 4} counting, duplicate cases will be counted.
[0098] To solve the above problems, the present invention uses uncorrelated node sets for counting. The collector randomly permutes the graph nodes uniformly to obtain G', and at the same time samples non - overlapping node sets and sends them to all users. Each node set can be represented as {v σ(i) , v σ(i+1) , v σ(i+2)}, where v σ(i) represents the i - th node after random permutation. After using the above - mentioned selection method, since the complexity of the sample selection process is O(n), the time complexity can be reduced from O(n 4 ) to O(n 2 ). After the data collector completes the extraction of the sampling sequence, it sends the sequence to the users, and the users calculate the number of triangular pyramids according to the node sets.
[0099] S7. For t non - overlapping node sets, data construction needs to be performed, with the nodes in the node set as external nodes and the other user nodes as central nodes. For a given node set {v i , v j , v m}, each user v k calculates the conical data {c k |k ∈ I -(i,j,m)} ∈ {0, 1} that can be constructed by judging whether it is connected to the above - mentioned node set itself. If it is connected to all, then assign c k = 1, otherwise c k = 0.
[0100] S8. Each user v k performs random response perturbation on the corresponding conical data locally, and the perturbation result is and sends it to the shuffle model. The shuffle model uses uniform random sampling perturbation π on the node set I -(i,j,m) , and after shuffling, obtains the private conical data and sends it to the collector.
[0101] S9. For each user v i , v j , v m in the node set {v x , user v x needs to calculate the one - sided relationship and the two - sided relationship with other users respectively. First, user vx Calculate the bilateral relationship w i , v j , v m}\{v x} between itself and the node pair {v x ; second, the user v x respectively calculates the unilateral relationships a i , v j , v m}\{v x} between itself and the remaining two nodes y ∈ {v xy . Then, the user v x uses the randomized response mechanism to perturb the unilateral and bilateral relationships respectively to obtain and and sends them to the collector;
[0102] S10. This step needs to calculate three cases, respectively estimating the triangles centered on the nodes v i , v j , v m . Its essence is to estimate the cases containing triangles in the node set {v i , v j , v m} from the perspectives of three users. Since the three cases are similar, only one case is written, and the subsequent estimations follow the same formula. For example, the triangle estimation centered on the node v i is:
[0103]
[0104] Therefore, by aggregating the three estimation cases, the triangle estimation of the node set {v i , v j , v m} can be obtained as:
[0105]
[0106] where the privacy budget is ε 2 and is shared with the conical data in the second-round shuffle.
[0107] S11. According to the shuffled and perturbed conical data obtained in step S8 and the triangle estimation obtained in step S10, estimate the triangular pyramid count of the corresponding node set. In the node set {v i , v j , v m}, according to c π(k) ′ and estimate the triangular pyramid count which is expressed as:
[0108]
[0109] Among them The privacy budget is It means that using the shuffle model to convert (ε 2 , δ 2 )-DP into
[0110] After the above estimation, it is necessary to prove that is an unbiased estimate of the triangular pyramid count , that is, to prove that The proof process is as follows:
[0111] Proof: First, it is necessary to estimate whether there is a triangle with v i as the core. Then, it is necessary to estimate the bilateral relationship w′ i of v i , and the unilateral relationship between the node pair (v j , v m ). Then, according to v j , v m two different nodes, the following equation exists:
[0112]
[0113] Among them, and are the triangle estimation values of the unilateral evaluation of v j and v m respectively. For the perturbed bilateral relationship w′ x , there is where the flipping probability of w x is q. Therefore, by transforming the above formula, we can get Similarly, for the unilateral relationship a′ xy , since it is perturbed using the random response mechanism, it is easy to get At the same time, since the evaluators of the unilateral relationship are different, it needs to be estimated twice. After merging, the following equation can be obtained (taking v i as the core example):
[0114]
[0115] The triangle estimation centered on v j , v m can also be extended in this way. Therefore, by combining the three cases, the triangle estimation of the node set is obtained Then further estimate the conical structure The perturbed data after the shuffling algorithm only changes the order and has no impact on the subsequent counting work. In the evaluation of conical data, the perturbation mechanism is the same as the above mechanism, except that LDP is used, that is, the flipping probability is ε L . Therefore, for each node v k , the estimated value of the conical structure that can be formed with the node set {v i ,v j ,v m} is expressed as Since the node set is {v i ,v j ,v m}, so k ∈ I -(i,j,m) , an unbiased estimate of the conical structure can be obtained as the equation:
[0116]
[0117] In summary, under the selected node set {v i ,v j ,v m} in graph G, the triangular pyramid count can be expressed as the following equation:
[0118]
[0119] The unbiasedness is proved, and at the same time, the present invention gives the mean square error (= variance) of the above estimation calculation. In the shuffling model, it is necessary to calculate the variance of the triangular pyramid count estimation value to estimate the accuracy of the count:
[0120] For any node set , the estimation provided by OTCC provides the following utility guarantee, as shown in the formula:
[0121]
[0122] where ε L is defined by the shuffling model, and ε L = logn + O(1) can be obtained Therefore, the above formula can be simplified again to obtain the formula:
[0123]
[0124] In each round of triangular pyramid counting process, if the shuffling model is not used, the variance is Through the shuffling model, in this paper, the variance is reduced from to Since most social networks are sparse graphs, this makes d max<<n. At the same time, in this process, the present invention estimates the triangle by constructing a combined relationship of unilateral and bilateral, which can ensure the original model structure to a greater extent and reduce the introduction of noise.
[0125] Repeat steps S7 to S11 until all the triangular pyramid count estimations in all disjoint node sets are completed, that is, complete t rounds of iteration, thus covering all sampling cases.
[0126] S12. Judge all node sets according to the threshold d th and set the range where the noise degree of all nodes in the node set is higher than the threshold d th as:
[0127] D + ={i = 1, 4, …, 3t - 2; d th ≤ min(d σ(i) ’, d σ(i+1) ’, d σ(i+2 )’)
[0128] Calculate the triangular pyramid count estimation in the entire graph network according to the selected node range
[0129]
[0130] Intercept and remove sparse nodes based on the noise degree threshold. Therefore, the optimized algorithm is not an unbiased estimate, so it is necessary to calculate the bias of the algorithm. First, prove that for the entire node set, is an unbiased estimate of the triangular pyramid count f 3* (G), that is, prove
[0131] Proof: First, there should be an equation relationship for the triangular pyramid count in the graph:
[0132]
[0133] where the constant 24 represents the number of times the triangular pyramid is repeatedly counted. Since three nodes are selected to form a node set, the possibility of selecting the same four nodes to form a triangular pyramid is 4 * 3 * 2, that is, it is repeated 24 times.
[0134] Then it can be proved that is an unbiased estimate, expressed as the formula:
[0135]
[0136] Secondly, calculate the bias of the algorithm. For the segmentation mechanism, the bias is as follows:
[0137]
[0138] According to the transformation of the magnitude of c, the deviation of the Piecewise mechanism is divided into four cases. For they are respectively falling in the intervals of [-C, l(d′ i ), [l(d′ i ), r(d′ i ), [r(d′ i ), C), and [C, +∞) for discussion. In the above formula, when d′ avg ∈[C, +∞), the deviation is the largest.
[0139] The proof of the deviation is completed. Next, the error of the DTCC algorithm of the present invention is given as shown in the following formula:
[0140]
[0141] where V + represents all nodes that meet the conditions . In addition,
[0142]
[0143] the maximum term of the variance of the algorithm is By removing the influence of the virtual edges generated by the sparse nodes in the network, the present invention can greatly reduce the variance. At the same time, it is not difficult to find that the variance of the present invention is determined by . When there are more sparse nodes in the network, the variance is smaller, the effect of the optimization scheme is stronger, and the minimum variance is
[0144] According to the combined mechanism of differential privacy, the privacy protection level of DTCC is as follows: DTCC provides (ε 1 +ε 2 , δ 1 +δ 2 )-DP and (2(ε 1 +ε 2 ), 2(δ 1 +δ 2 ))-edge-DP privacy guarantees. In addition, through the transformation of the shuffling algorithm, the present algorithm provides privacy guarantees on local differential privacy.
[0145] The present invention provides a privacy-preserving triangular cone counting estimation method DTCC (Degree-based Triangular Cone Counting) based on a two-round shuffling model, and optimizes the network structure by setting a threshold according to the node noise degree to ensure global features. The present invention ensures that users locally perturb the data, and by reducing the dimensionality of multi-dimensional data, it ensures that local features are not destroyed by introduced noise, and better calculates the estimation query for tightly clustering the graph structure. On the premise of satisfying local differential privacy, the query result of the triangular cone counting for the entire graph still reaches a high accuracy rate.
[0146] Based on the above process, the following are the experimental settings and results of the present invention.
[0147] In the experimental scheme setting, according to the research exploration in this article, there is no relevant method for calculating triangular cones under privacy conditions at present. Therefore, in order to experimentally verify the triangular cone counting under privacy protection, the data utility is mainly proved by the relative error and the mean squared error, and at the same time, the influence of the newly added shuffling model and different perturbation mechanisms on the experimental utility is compared through experiments.
[0148] (1) Evaluation criteria
[0149] In this section, the experiment executes each algorithm 20 times, and calculates the average relative error (Relative Error, abbreviated as: RE) and the mean squared error (Mean Squared Error, abbreviated as: MSE) of these 20 times.
[0150] The mean squared error (Mean Squared Error, abbreviated as: MSE) is calculated by formula (3.32):
[0151]
[0152] When the scale of the social network is large, the subgraph counting will increase with the increase in the number of nodes. At the same time, according to the difference between the estimated value and the true value, the MSE will increase in the form of a square factor, so that the MSE exponent increases exponentially. Therefore, in order to reduce the influence of the order of magnitude, this article calculates its relative error to show the accuracy of the algorithm applicable to different scale data. The MSE converges numerically to the square multiple of the relative error, and the image trends are similar, so the experimental result graph only shows the comparison result of the relative error.
[0153] The relative error (Relative Error, abbreviated as: RE) is calculated by formula (3.33):
[0154]
[0155] Where and f i(G) represents the estimated count and the true count of the i-th result, and the average relative error is taken as the evaluation criterion. In addition, S is a sanity constraint that can mitigate the impact of queries with extremely small selectivity. By convention, S = 0.1%n, which is the same as the settings in previous work.
[0156] (2) Dataset
[0157] To experimentally evaluate the triangular pyramid algorithm under the above local differential privacy, this paper uses the following two datasets:
[0158] Gplus (Google+): This dataset is a dataset released by the SNAP Lab at Stanford University. It contains the "circle" information of users in the Google+ social network. This dataset consists of node features (Profiles), circles, and ego networks. It contains approximately 107,614 nodes and 13,673,453 edges. The statistical properties of this dataset include an average clustering coefficient of 0.4901 and a triangle count as high as 1,073,677,742. Based on this dataset, a graph G = (V, E) is constructed, with an average degree of d avg = 227.4 and a maximum degree of d max = 20,127.
[0159] IMDB (Internet Movie Database): This dataset is sourced from the world's largest movie database website and contains information on a large number of movies, TV shows, documentaries, and other works, as well as user reviews and ratings of these works. The dataset contains 896,308 actors and 428,440 movies, and the relevant information of the dataset is represented by a matrix. The dataset is used to construct a bipartite graph G = (V, E). In this work, actors are regarded as nodes, with a number of 896,308; the edges are used to represent whether two actors have appeared in the same movie, with a number of 57,064,358; the average degree is d avg = 127.3 and a maximum degree of d max = 15,451.
[0160] (3) Comparison schemes
[0161] Based on the research and exploration in this paper, there is no relevant method for calculating the triangular pyramid under privacy conditions at present. And existing related work has proven the effectiveness of the single-round shuffle, so this paper compares the statistical results of the single-round shuffle OneShuffle 3* and the two-round shuffle TwoShuffle 3* .
[0162] OneShuffle: Only uses the shuffle model in the conical shuffle stage and uses the Laplace mechanism to perturb the node degrees.
[0163] Lap_TwoShuffle: Use the shuffle model in both the node degree interaction and conical shuffle phases, and use the Laplace mechanism to perturb the node degree.
[0164] Duchi_TwoShuffle: The precondition is the same as Lap_TwoShuffle, and use the segmentation mechanism to perturb the node degree.
[0165] (4) Experimental configuration
[0166] The experiments in this section were conducted on the Windows 10 operating system, with 16GB of server memory and an Intel(R) Core(TM) i5-10500 CPU @ 3.10GHz configured processor.
[0167] (5) Experimental metrics
[0168] In this paper, δ = 10 -8 , and amplify privacy through the shuffle model while strictly adhering to the upper limit. In addition, divide ε into two parts, where and Since the node degree is only used as a discriminant factor for selecting data and is not included in the exact count calculation, ε 1 is used as the privacy budget for node degree perturbation.
[0169] For the experimental objectives, in the experimental design, to compare the impact of the privacy budget ε on the experimental error, ε is set to [0.2, 5]; in addition, the impact of the threshold parameter c (adjust the threshold for selecting nodes) on the experimental error is also compared, and c is set to [1, 5].
[0170] Experimental results:
[0171] (1) Impact of privacy budget on utility.
[0172] In this experiment, the privacy budget is set to [1, 5], and the case of threshold c = 1 is given simultaneously to explore the impact of the privacy budget on utility. Tables 1 and 2 respectively show the impact of the privacy budget ε on the relative error in the Gplus and IMDB datasets. First, analyze the overall trend of the experimental results. It can be clearly seen that as the privacy budget ε gradually increases, the relative error continuously decreases. The above phenomenon means that when the privacy protection intensity is low, the present invention can obtain better subgraph counting results.
[0173] Table 1. Verification of the impact of the privacy budget ε on the relative error in the Gplus dataset under the threshold parameter c = 1 of the present invention
[0174]
[0175] Table 2. Influence of privacy budget ε on relative error under the threshold parameter c = 1 in the present invention, verifying in the IMDB dataset
[0176]
[0177] Regarding the influence of the newly added shuffling algorithm on data utility, it can be found through experiments comparing algorithms that all perturbation mechanism algorithms under TwoShuffle 3* are almost completely superior to those under OneShuffle 3* in terms of algorithm utility. In addition, the error of the Laplace mechanism perturbation algorithm Lap_TwoShuffle under TwoShuffle is much lower than that of the OneShuffle algorithm (also using the Laplace mechanism). Therefore, the present invention believes that using the shuffling algorithm during the node degree interaction process can better improve the data utility of the algorithm.
[0178] It can be found through experiments comparing algorithms that the piecewise mechanism PM_TwoShuffle is a relatively good perturbation mechanism. Regarding the usability and stability of the present invention, it can be clearly seen that it has significant robustness to changes in the dataset, and the trend of data utility changes little for different datasets. At the same time, when the threshold c = 1, a relative error of the order of magnitude of 10 can be maintained under a relatively low privacy budget, that is, ε≥2. 0 Therefore, it can be proved that the counting method under privacy protection proposed by the present invention provides high counting accuracy while ensuring privacy.
[0179] (2) Influence of the threshold parameter on utility.
[0180] In this experiment, the range of the threshold parameter c is set to [1, 5] to test the influence of deleting different numbers of nodes on the counting results. Tables 3 and 4 respectively show the relative errors of all perturbation algorithms on the Gplus and IMDB datasets. Among them, in order to show the performance under each threshold parameter c, the case of ε = 1 is shown.
[0181] Table 3. Influence of threshold parameter c on relative error under privacy budget ε = 1 in the present invention, verifying in the Gplus dataset
[0182]
[0183] Table 4. Influence of threshold parameter c on relative error under privacy budget ε = 1 in the present invention, verifying in the IMDB dataset
[0184]
[0185] First, analyze the overall trend of the experimental results. It can be clearly seen that as the threshold parameter c continues to increase, the relative errors of all algorithms show a decreasing trend simultaneously. When the threshold parameter c is small, the decreasing trend of the relative error is more obvious; when the threshold parameter c is large, the decreasing trend of the relative error significantly flattens out. This indicates that in a large network structure, there are a large number of sparse nodes, and using the threshold to remove nodes can prevent the introduction of a large number of false edges and ensure the accuracy of counting.
[0186] Regarding the impact of the newly added shuffling algorithm on utility, among all the algorithms compared under the experimental conditions of this experiment, it is easy to find that the two-round shuffling algorithm TwoShuffle 3* is almost completely superior to the one-round shuffling algorithm OneShuffle 3* , and the error of the two-round shuffling Laplace algorithm is much smaller than that of the one-round shuffling. The PM_TwoShuffle algorithm is basically superior to other algorithms when the privacy budget ε = 1. Compared with the Laplace mechanism Laplace_TwoShuffle, the former is superior to the latter in different environments; compared with the Duchi mechanism Duchi_TwoShuffle, PM_TwoShuffle is superior when the threshold parameter is small, while Duchi_TwoShuffle performs better when the threshold parameter is large. However, the experimental results of Duchi_TwoShuffle are even worse than those of OneShuffle in some cases when the threshold parameter is small. 3* . In addition, the PM_TwoShuffle method proposed in the present invention maintains a relative error of the order of 10 0 when the privacy budget ε = 1 and the threshold parameter c ≥ 3. Therefore, according to the above analysis, it can be proved that for the method of the present invention, the threshold parameter c can be used to optimize the algorithm to obtain higher counting accuracy.
[0187] The method for estimating the clustering relationship of complex subgraphs proposed by the present invention mainly focuses on the subgraph counting method for the privacy-preserving conical structure, which has a specific node as the core and the remaining nodes are connected in a cycle. The conical structure has wide applications in real life and can be used in various scenarios such as communication transmission, power systems, and UAV formations. In these application networks, it is often impossible to evaluate important structures through simple subgraph counting. For example, consider a power supply network where the power generation station is the core node, and the substations and users are the peripheral nodes, forming a conical connection relationship. In contrast, for triangles and k-stars in simple subgraphs, a triangle can only describe a closed ternary relationship and cannot reflect the connection between the central node and the peripheral nodes. On the other hand, a k-star can only reflect the central role of the power generation station, but in fact, the connection between the substations and users cannot be ignored. During the process of performing the graph structure data collection task on the above network, information from different users is required. However, this can easily endanger the privacy of data owners and cause serious consequences such as data leakage and privacy infringement. Therefore, the present invention has clear practical application scenarios and strong practical significance.
[0188] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A complex subgraph clustering relationship estimation method for privacy-preserving graph network analysis, characterized by: Includes steps: S1. Initialize the user adjacency table and privacy budget. Suppose the number of users is n. The global privacy budget required for data perturbation is Relaxation term δ1,δ2∈(0,1), the adjacency table of each user is a1,…,a n ∈{0,1} n ; S2. Calculate the privacy budget required for two rounds of local data perturbations based on the shuffle model The local differential privacy is converted to the central differential privacy through two shuffles, so as to amplify the data utility under the same privacy budget. S3. Each user calculates its own node degree d based on its own adjacency table i , and the noise degree d is obtained by perturbing the segmented mechanism based on the mean estimation i ′, and send the perturbed data to the shuffle model at the same time; S4, the shuffle model uses uniform random sampling to perturb π on n user nodes, obtains the perturbation sequence [π(1),…,π(n)] and sends it to each user. At the same time, the shuffled noise degree d is set according to the order of the shuffled sequence. π(i) 'Send to the collector; S5, the collector is based on the noise degree d of the shuffle π(i) Calculate the noise average d avg ′, set the threshold d based on the noise average th As a basis for the subsequent removal of sparse nodes, it reduces computational overhead and improves data utility; S6, the collector uniformly and randomly arranges the graph nodes to obtain G', and samples disjoint sets of nodes and send them to all users. Each node set is represented by {v σ(i) ,v σ(i+1) ,v σ(i+2) }, σ represents random arrangement; S7, each user calculates whether it is connected to all nodes in a certain node set for the above t disjoint node sets, and gives the corresponding connection status, and regards the upper connection status as a cone data structure, and then converts it into one-dimensional binary data for processing; S8. Each user perturbs the cone data using a random response mechanism and rearranges it using a shuffle model to obtain the shuffled perturbed cone data c π(k) ′, send it to the collector; S9. For each point in the disjoint node set, calculate the unilateral relationship and bilateral relationship between itself and the other two points, that is, whether it is connected to a certain point alone or to both points. At the same time, use the random response mechanism to perturb the above relationship and send it to the collector. S10, considering each disjoint node set, estimating whether the three nodes can form a triangle according to the unilateral and bilateral relationships obtained in step S9, and obtaining a triangle estimation; S11, estimating the triangular pyramid count of the corresponding node set according to the shuffled and disturbed cone data obtained in step S8 and the triangle estimate obtained in step S10; S12, all node sets are divided according to the threshold d th Make a judgment, when the noise level of all nodes is higher than the threshold d th Included in the pyramid estimation of the overall network, the pyramid count estimation in the computational graph network 2. The complex subgraph clustering relationship estimation method for privacy-preserving graph network analysis according to claim 1 is characterized in that: In step S2, there are n users in the first round of node degree shuffle and n-3 users in the second round of cone shuffle. Therefore, the privacy budgets of the two rounds of shuffle are: ShuffleBudget represents the inverse function defined by the shuffle model, which is used to convert local differential privacy and central differential privacy to each other. Mathematically, it is expressed as converting (ε,δ)-DP to ε L -LDP.
3. The complex subgraph clustering relationship estimation method for privacy-preserving graph network analysis according to claim 1 is characterized in that: Step S3: For user i, according to the adjacency table a i Calculate the node degree d i , calculated as follows: Using the segmentation mechanism to perturb the node degree d i , the disturbance output value d of user i i 'Will be generated uniformly and randomly in the segment of [-C,C] range, a i,j Indicates whether user i is adjacent to user j. When i=j, a i,j =0; The corresponding probability density function is expressed as: Where x represents the disturbance value, r(d i )=l(d i )+C-1; the privacy budget is represents the privacy budget of the first round of shuffling; in addition, similarly, the node degree d is perturbed i 'Map back into the normal real domain and send it to the shuffle model.
4. The complex subgraph clustering relationship estimation method for privacy-preserving graph network analysis according to claim 1, characterized in that: In step S5, the collector obtains the noise degree d after all shuffle disturbances π(i) ′, calculate the noise average d avg ′, expressed as: According to the noise average d avg 'Set threshold d th ,The parameter c is used to adjust the threshold size to adapt to the needs of different graph data sets, which is specifically expressed as: d th =cd avg ′ The subsequent threshold d th Remove sparse nodes in the network, optimize the network structure, and better preserve global features.
5. The complex subgraph clustering relationship estimation method for privacy-preserving graph network analysis according to claim 1, characterized in that: Step S9: for the node set {v i ,v j ,v m } for each user v x , user v x It is necessary to calculate unilateral and bilateral relationships with other users separately; First, user v x Calculate the node pair {v i ,v j ,v m }\{v x Bilateral relations between x ; Second, user v x Calculate itself and the remaining two nodes y∈{v i ,v j ,v m }\{v x } unilateral relationship a xy ; Next, user v x Using random response mechanism to perturb unilateral and bilateral relations respectively, we can obtain and And send it to the collector.
6. The complex subgraph clustering relationship estimation method for privacy-preserving graph network analysis according to claim 1, characterized in that: In step S10, considering that the nodes v i ,v j ,v m The result of triangle estimation is the same as that of the center, and then it is simplified, considering the node set {v i ,v j ,v m } contains triangles, and the node set {v i ,v j ,v m The triangle estimate of} is: in The privacy budget is ε2, which is shared with the cone data in the second round of shuffling.
7. The complex subgraph clustering relationship estimation method for privacy-preserving graph network analysis according to claim 1, characterized in that: In step S11, in the node set {v i ,v j ,v m }, according to c π(k) 'and Estimated pyramid count It is expressed as: in The privacy budget is It means that the (ε2,δ2)-DP is converted to 8. The complex subgraph clustering relationship estimation method for privacy-preserving graph network analysis according to claim 1, characterized in that: Step S12: All node sets are divided according to the threshold d th Make a judgment and set the noise level of all nodes in the node set to be higher than the threshold d th The range is: D + ={i=1,4,…,3t-2;d th ≤min(d σ(i) ’,d σ(i+1) ’,d σ(i+2) ’) Calculate the pyramid count estimate in the entire graph network based on the selected node range