Trusted efficient graph division method for large-scale privacy graph data in cluster
By identifying and marking privacy nodes and edges in graph data, combining the methods of privacy balanced division and privacy awareness division, the problem of insufficient privacy protection of large-scale graph data in the prior art is solved, and efficient and trustworthy graph data division is achieved.
Patent Information
- Application Number
- CN202510268323.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-20
AI Technical Summary
Existing graph division algorithms lack privacy protection when processing large-scale graph data, especially in sensitive infographic data, which cannot fully guarantee privacy security.
A trusted and efficient graph division method for large-scale privacy graph data in the cluster is proposed. By identifying the privacy of privacy nodes and privacy edges, a balanced privacy division and privacy awareness division method is adopted to ensure privacy protection and security during the graph data division process.
It realizes efficient and trustworthy division of large-scale privacy graph data while ensuring information security, improves computing efficiency and resource utilization, and ensures data security.
Smart Images

Figure CN120180474A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer data processing and information security, and particularly to the partitioning technology of large-scale privacy graph data in a cluster environment.
Background Art
[0002] A graph is an abstract data structure used to represent the association relationships between objects, described using vertices and edges: vertices represent objects, and edges represent the relationships between objects. Graph data is data that can be abstracted into a graph description, and graph computing is the process of expressing and solving problems using a graph as a data model. With the continuous development of intelligent computing, the scale of graph data has been increasing continuously, far exceeding the single-machine storage limit. As a result, distributed graph computing has emerged, which requires partitioning the graph into several subgraphs (a subset of the graph) and distributing them to each computing node for execution.
[0003] In the current situation of the rapid development of graph computing, the computing speed and efficiency of graph computing have been significantly improved. As a prerequisite step of a graph computing system, graph partitioning plays an important role in the graph computing system, and the quality of the partitioning result can directly affect the speed of subsequent computing.
[0004] Graph data has the characteristics of a complex network structure, high heterogeneity and complexity, a large number of nodes, and complex connection relationships. Some graph data (such as biomedical graph data) also involves sensitive information, and privacy protection is an important consideration. Graph partitioning algorithms play an important role in the computing of graph data. Through reasonable partitioning, the computing efficiency can be improved, a large graph can be divided into blocks for parallel computing, the speed can be accelerated, resource utilization can be optimized, communication costs can be reduced, and at the same time, data security can be guaranteed, a privacy protection mechanism and fine-grained access control can be constructed; and distributed computing can be supported, tasks can be evenly distributed, and fault tolerance can be improved.
[0005] Although current graph partitioning algorithms (heuristic, multilevel, spectral graph theory-based, etc.) have improved the performance of graph computing, they lack privacy protection. Research on graph partitioning privacy protection mostly focuses on cryptographic solutions, which will result in a large amount of performance loss in large-scale graph computing scenarios. Therefore, when dealing with sensitive information graph data, existing algorithms cannot fully guarantee privacy security, and new algorithms combining technologies such as differential privacy need to be explored to balance computing performance and privacy protection.
Summary of the Invention
[0006] To address the above problems, the present invention implements a trusted and efficient graph partitioning method for large-scale privacy graph data within a cluster, which can perform efficient and trusted partitioning while ensuring the information security of large-scale privacy graph data within the cluster.
[0007] A trusted and efficient graph partitioning method for large-scale privacy graph data within a cluster, characterized by including the following steps:
[0008] Identify privacy nodes and privacy edges in the identification mark graph data, and mark the privacy degrees of the privacy nodes and privacy edges;
[0009] Based on the privacy degrees, perform privacy balance partitioning among servers and privacy-aware partitioning within a server.
[0010] Preferably, the privacy balance partitioning includes a coarsening stage, a partitioning stage, and a refinement stage; among them,
[0011] In the coarsening stage, confirm candidate nodes for merging based on the label propagation algorithm of privacy degrees, retain the mapping relationships of the original nodes, and accumulate the privacy degrees of the nodes and the connected edges until the node scale is smaller than the set threshold to obtain the compressed graph data;
[0012] Preferably, the coarsening stage includes the following steps:
[0013] Find the neighbor nodes of each node: In the graph data structure, each node has a connection relationship with several other nodes, and these directly connected nodes are the neighbor nodes of the node. Finding neighbor nodes is the starting step of this algorithm. By traversing the graph data structure, the neighbor set of each node can be determined. This step lays the relevant data foundation for calculating the merging cost value in the subsequent steps.
[0014] Obtain the merging cost value: When calculating the merging cost, the sum of the out-degree and the in-degree reflects the connection activity degree of the node in the graph, which can reflect the importance and interaction situation of the node in the graph. Considering the cost factor can reasonably evaluate the impact of the merging on the graph structure and data flow, so as to select the most suitable nodes for the merging operation.
[0015] Select the node pairs with low cost for priority merging according to the merging cost value: After obtaining the merging cost values of all node pairs, by comparing the sizes of these cost values, select the node pairs with low cost for priority merging operations. The low-cost first strategy ensures the correctness and rationality of the coarsening stage. Prioritizing the merging of low-cost nodes protects the original attributes of the data to the greatest extent and facilitates the execution of subsequent refinement operations.
[0016] Preferably, the coarsening stage includes the following steps:
[0017] Use ω(u, v) to represent the edge weight between the nodes to be merged, use deg(u) and deg(v) as the node degrees, and use p u 、p v 、p (u,v) to represent the privacy degrees of nodes u and v and the privacy degree of the edge between u and v respectively, and obtain the merging cost of the candidate nodes
[0018] Select the two candidate nodes with the lowest cost for merging according to the merging cost, retain the original node mapping relationship, accumulate the privacy degrees of the nodes and the connected edges, process the merging work simultaneously using multiple threads, and adopt a local search algorithm to avoid repeatedly merging the same node until the node scale is smaller than the set threshold.
[0019] In the partitioning stage, the graph data is initially partitioned into k partitions, where k represents the number of servers in the cluster; then the partitioning result is adjusted by the FM optimization algorithm based on privacy degree.
[0020] Preferably, the detailed steps of the FM optimization algorithm based on privacy degree are as follows:
[0021] Select the partition edge nodes: In the scenario of partitioning privacy graph data within a cluster, different servers correspond to different partitions, and the partition edge nodes are the data nodes in the edge area of the server. Accurately selecting the edge nodes is the first step for the FM algorithm to perform optimization operations, providing an object basis for calculating the node movement gain subsequently.
[0022] Calculate the gain obtained by moving to other partitions: The movement gain includes load balancing gain, privacy data balance gain, locality gain, and communication overhead gain. The load balancing gain is the difference between the load balancing value after movement and the load balancing value before movement. The privacy data balance gain represents the difference between the privacy data balance index after movement and the privacy data balance index before movement. The locality gain represents the difference between the data locality gain after movement and before movement. The communication overhead gain is the difference between the communication overhead value after movement and the communication overhead value before movement. By calculating the sum of these parts, the movement gain is obtained.
[0023] Select the partition with the largest gain and transfer the node to this partition: After calculating the comprehensive gain of the node moving to each partition, select the partition with the largest gain. By transferring the node to the partition with the largest gain, the optimality of the partitioning operation can be ensured.
[0024] Preferably, the partitioning stage includes the following steps:
[0025] Set the partitioning target
[0026] Select the initial nodes, obtain the node degree deg(u), and calculate the node centrality Select the top k nodes according to the centrality size;
[0027] Taking the k nodes as the center, use the breadth-first search algorithm to construct a subgraph;
[0028] Recursively partition, with θ as the target number of nodes, partition the partitions. If the number of nodes in the partition exceeds 1.5θ, repeat the above process to split the original large partition into new sub-partitions;
[0029] Process the cut edges. If nodes u and v are in different partitions respectively, calculate the node movement gain using the gain function, select the partition with the maximum movement gain and uniformly move the nodes into this partition. If the movement gain is low, give up the movement;
[0030] Output the result, and finally form the preliminary partitioning result of k partitions.
[0031] Preferably, the movement gain includes load balancing gain, private data balance gain, locality gain, and communication overhead gain; the load balancing is defined as the standard deviation of the partition load and the average load The load balancing gain is the difference between the load balancing value after movement and the load balancing value before movement The privacy ratio is defined as the proportion of private data in the partition data volume The private data balance gain represents the difference between the private data balance index after movement and the private data balance index before movement To measure whether the movement is beneficial to the balance of private data, the communication overhead is defined as the weight of the cross-partition edges The communication overhead gain represents the difference between the number and weight of the cross-partition edges after movement and before movement The locality is defined as the number of cross-partition edges, and the locality gain is the difference between the communication overhead value after movement and the communication overhead value before movement, that is Calculate the movement gain G = αg' l (u, i, j) +
[0033] βg' c (u, i, j) + γg' p (u, i, j) + θg' n (u, i, j); where α, β, γ, θ are weight coefficients, and the initial defined values are all 1, which are dynamically adjusted by the user according to the target requirements.
[0034] In the refinement stage, restore the graph data layer by layer to the original scale according to the node mapping relationship.
[0035] Preferably, in the refinement stage, according to the node mapping relationship, for the possible load imbalance problem in the private partition, continue to add semi-private nodes to the private partition, and allocate regular nodes with high priority to the private partition according to the situation to ensure the load balance of the private partition. Optimize the partitioning result with the FM optimization algorithm based on the privacy degree, and restore layer by layer until the graph data is restored to the original scale.
[0036] Preferably, the privacy-aware partitioning includes the following steps:
[0037] Identify data privacy level: It is necessary to read the graph data privacy level to determine the privacy level of each node and edge in the graph structure;
[0038] Update node privacy level using the privacy level propagation algorithm: Nodes updated by the privacy level propagation algorithm can more accurately reflect the actual privacy situation of nodes in the entire data graph, and taking into account the mutual influence between nodes helps to more reasonably allocate and protect data subsequently;
[0039] Use a priority queue to allocate private nodes and private edges to the Trusted Execution Environment (TEE) according to specific rules: Add private nodes to the TEE according to the privacy level of the nodes. If the privacy levels of the nodes are the same, add the private nodes to the TEE in descending order of the node degrees. The TEE is a trusted execution environment that can effectively protect the security of private data. Allocating private nodes and private edges to the TEE can prevent these private data from being leaked or misused;
[0040] Make corresponding adjustments according to the load situation at the TEE end: Semi-private nodes refer to nodes with a privacy level value between 0 and 1 after privacy level update; semi-private nodes have a certain degree of privacy sensitivity, but their sensitivity is less than that of private nodes. When the load at the TEE end is small, add copies of semi-private nodes in the priority queue to the TEE environment until the load is balanced. When the load at the TEE end is large, stop adding nodes and edges;
[0041] Delete private nodes and private edges existing in the normal execution environment: Ensure that private nodes and private edges only exist in the TEE to further ensure the security of private data.
[0042] The detailed steps of the privacy level propagation algorithm involved in this method are as follows:
[0043] Initialize the privacy level: Use the recognition result read by the server as input to initialize the privacy level in the graph and identify private nodes and private edges;
[0044] Propagate the privacy level: For data nodes in the graph, update their privacy level. The specific operation is to accumulate the privacy levels of private nodes and private edges among its neighbor nodes, divide by the degree of the node, and obtain a privacy level value between 0 and 1, and continue to update other vertices in the graph;
[0045] Calculate convergence: When the difference between the current state privacy level and the previous state privacy level of nodes in the graph is less than the specified threshold, it can be considered that the algorithm converges and the algorithm execution is completed;
[0046] The detailed steps of the processing of private edges involved in this method are as follows:
[0047] When processing, the privacy edges and their directly adjacent vertices need to be added to the TEE: For the adjacent vertices of the privacy edges, if they are privacy nodes, they are added to the TEE, and the nodes are deleted in the regular partition. If they are regular nodes, copies of the nodes are added to the TEE, and the privacy edges are deleted in the regular partition.
[0048] Preferably, the privacy-aware partitioning includes the following steps:
[0049] Cache privacy nodes and privacy edges to the TEE;
[0050] Judge the load τ on the TEE side TEE and the load τ on the REE side REE as well as the maximum cache capacity τ on the TEE side max_TEE ;
[0051] If τ TEE < τ REE and τ TEE < τ max_TEE , use the priority queue to cache copies of semi-privacy nodes to the TEE side until the load on the TEE side and the REE side is balanced or the maximum cache capacity of the TEE side is reached;
[0052] If τ TEE > τ REE , or τ TEE = τ max_TEE , stop adding nodes and edges to the TEE side; if there are privacy edges on the REE side, hide or encrypt the privacy edges existing on the REE side;
[0053] Among them, the TEE is a trusted execution environment, the REE is a regular execution environment, the TEE and the REE side rely on shared memory for communication, and privacy data is all calculated on the TEE side; the initial values of the privacy degree are 0, 1, 2, 3, 0 represents regular data, 1 represents secret data, 2 represents confidential data, and 3 represents top-secret data; semi-privacy nodes are the original regular nodes whose privacy degree values are between 0 and 1 after the privacy degree is updated.
[0054] A trusted and efficient graph partitioning system for large-scale privacy graph data in a cluster includes:
[0055] An identification and marking module, used to identify privacy nodes and privacy edges in the graph data and mark the privacy degrees of the privacy nodes and privacy edges;
[0056] A partitioning module, used to perform privacy balance partitioning between servers and privacy-aware partitioning within servers based on the privacy degree.
[0057] A computing device includes a processor and a memory; the memory is used to store computer execution instructions; the processor is used to execute the computer execution instructions so that the computing device executes the method described above.
[0058] Experimental results show that when considering privacy and partitioning privacy data, the algorithm mentioned in this paper can achieve good results in the partitioning of regular data and privacy data, and also shows excellent performance in multiple aspects such as load balancing and locality.
[0059] In terms of load balancing, the number of nodes in each partition is exactly the same, achieving a balanced partition between partitions. This result proves that compared with other traditional graph partitioning algorithms and the multi-level partitioning algorithm METIS, the proposed trusted and efficient graph partitioning method for large-scale privacy graph data within a cluster has obvious advantages.
[0060] In terms of locality, this algorithm uses the strategy of locally merging privacy nodes, which effectively optimizes the locality of privacy nodes, and thus the calculation results of this algorithm in terms of locality are higher than those of other algorithms, showing greater superiority compared with other solutions.
[0061] However, in terms of overhead, due to the introduction of means such as node replicas, the number of nodes increases. Compared with traditional algorithms and the METIS algorithm for multi-level partitioning, this algorithm needs to establish connections between replica nodes and real nodes, which increases the cross-partition communication overhead. Therefore, there is still a certain gap in communication overhead compared with other algorithms, which is also one of the future optimization directions of this algorithm.
[0062] From the above description, it can be concluded that this method realizes the trusted and efficient partitioning of large-scale privacy graph data while ensuring data security.
Description of the Drawings
[0063] Figure 1 It is the architecture diagram of a trusted and efficient graph partitioning method for large-scale privacy graph data within a cluster according to the present invention;
[0064] Figure 2 It is the flow chart of privacy balance partitioning between servers according to the present invention;
[0065] Figure 3 It is the flow chart of privacy-aware partitioning within a server according to the present invention.
Detailed Embodiments
[0066] The present invention discloses a trusted and efficient graph partitioning method for large-scale private graph data within a cluster. This method features load balancing and strong locality, and can be widely applied to various aspects of production and life. Based on domestic servers and their architectures, it runs in a secure environment of domestic servers, effectively protecting the security of domestic servers and their software. The specific implementation of this method will be elaborated in detail below with reference to the accompanying drawings. It should be noted that only a relatively common usage mode of this method is described in this embodiment, and the technology applied by this method should be protected until practitioners in the relevant field can develop more excellent solutions.
[0067] A trusted and efficient graph partitioning method for large-scale private graph data within a cluster, comprising the following steps:
[0068] Identify the private nodes and private edges in the graph data, and mark the privacy degrees of the private nodes and private edges;
[0069] Based on the privacy degrees, perform privacy-balanced partitioning among servers and privacy-aware partitioning within servers.
[0070] In one embodiment, the private nodes and private edges are marked by keyword matching or manual marking.
[0071] In one embodiment, the three steps of coarsening, partitioning, and refinement are sequentially executed to evenly partition the compressed private data to different servers according to the fast partitioning algorithm to achieve load balancing, and then the details of the graph data are restored.
[0072] In one embodiment, the five steps of identifying the data privacy degree, updating the node privacy degree using the privacy degree propagation algorithm, allocating private nodes and private edges to the TEE using the priority queue, adding semi-private node replicas to the TEE according to the load situation at the TEE end, and deleting the private edges existing in the normal execution environment are sequentially executed to further partition the private data inside the server.
[0073] Figure 1 The specific architecture diagram of the present invention is shown and described as follows:
[0074] T1 to T3 are to receive the graph data and identify the privacy situations of the graph data edges and points.
[0075] T4 to T6 are to perform inter-server module partitioning.
[0076] T7 to T14 are to perform intra-server module partitioning.
[0077] T1 is the input of graph data. The graph data input by this method is large-scale private graph data, which is not only huge in scale but also contains rich individual and group information.
[0078] T2 is to mark the privacy degree of the graph data. In this method, manual marking or keyword matching is usually adopted.
[0079] T3 is for the local server to identify the graph data with privacy degree, store it in the local regular environment first, and prepare for further partitioning.
[0080] T4 is the coarsening module, which performs pairwise merging through the label propagation algorithm based on the privacy degree until the data scale is reduced to the set threshold. During the process, the accumulation of the privacy degrees of nodes and connected edges needs to be synchronized.
[0081] T5 is the partitioning module, which uses the fast partitioning algorithm to initially partition the compressed graph data into k partitions, adopts the PBFM optimization algorithm to adjust the partitioning result, and distributes the privacy graph data of the k partitions to k servers.
[0082] T6 is the refinement module, which is used to restore the data to the size before compression.
[0083] T7 is to identify the privacy data, and the server reads down the privacy data for subsequent partitioning operations.
[0084] T8 is to update the privacy degree of the data, and the privacy degree of the privacy data is updated using the privacy degree propagation algorithm.
[0085] T9 is for the privacy data to enter the priority queue according to the privacy degree from high to low.
[0086] T10 is for the semi-private data to queue up according to the load situation. If the TEE load is small, the semi-private data queues up to the TEE load balancing. If the TEE load is large, the semi-private data stops queuing up.
[0087] T11 is for the privacy data to directly enter the TEE executable environment.
[0088] T12 is for the semi-private data to add its copy to the TEE executable environment.
[0089] T13 is the TEE trusted execution environment, which communicates with the regular execution environment only relying on shared memory, can effectively protect the security of privacy data, and all privacy data is stored and calculated at the TEE end.
[0090] T14 is to delete the privacy edges in the regular execution environment to ensure the security of privacy information in the subsequent calculation process.
[0091] Figure 2 It shows the basic process of privacy balance partitioning among servers of the present invention:
[0092] Q1 is the starting step of this method. In this state, the local server has already read the privacy degree of the large-scale privacy graph data.
[0093] Q2 is that the local server determines the merging candidate nodes by using the privacy degree-based label propagation algorithm.
[0094] Q3 is that the local server accumulates the privacy degrees of the nodes and the connected edges while retaining the original node mapping relationship.
[0095] Q4 is that the local server judges the node scale.
[0096] If the node scale is greater than the set threshold, return to step Q2, and re-determine the merging candidate nodes by using the privacy degree-based label propagation algorithm, and continue the node merging operation to further reduce the scale of the graph data.
[0097] If the node scale is less than the set threshold, enter step Q5.
[0098] Q5 uses the fast partitioning algorithm to initially partition the compressed graph data into k partitions. The privacy graph data of the k partitions is allocated to k servers.
[0099] Q6 is that after the partitioning is completed, each server starts to restore the details of the graph data separately.
[0100] Q7 is the end of the whole process, and the privacy balance partitioning between servers is completed.
[0101] Figure 3 Shows the basic process of privacy-aware partitioning within the server of the present invention:
[0102] M1 is the starting step of this method. In this state, the local server starts the privacy-aware partitioning process.
[0103] M2 is that the local server starts to identify the data privacy degree.
[0104] M3 is to update the node privacy degree by using the privacy degree propagation algorithm, and update the privacy degrees of the relevant nodes according to the identified privacy degree data.
[0105] M4 is that the local server uses the priority queue to allocate the privacy nodes and privacy edges to the trusted execution environment TEE.
[0106] M5 is that the local server judges the load situation at the TEE end.
[0107] If the load at the TEE end is small, enter step M6.
[0108] If the load at the TEE end is large or the load balance has been reached after adding semi-private nodes, enter step M7.
[0109] M6 is that the local server uses the priority queue to add copies of semi-private nodes to the trusted execution environment TEE.
[0110] M7 deletes the privacy edge existing in the regular execution environment for the local server.
[0111] M8 indicates the end of the entire process, and the local server completes the privacy-aware partitioning within the server.
[0112] The example of this application is a typical case. That is, in addition to this example, there are many other examples available for use or testing, and this example is only for illustrative description.
[0113] The data sets used in this example are all from the Internet, and testing their performance using the corresponding data sets is not within the scope of the claims of this patent's claim book.
[0114] Except for the items without relevant claims in this patent statement, all other activities related to the use, development, and design of this patent should be carried out only after obtaining the permission of the inventor. The content related to the intellectual property of this invention that is not mentioned in this application should also be protected accordingly.
Claims
1. A reliable and efficient graph partitioning method for large-scale private graph data in a cluster, characterized by: The steps include: Identify privacy nodes and privacy edges in the labeled graph data, and mark the privacy degrees of the privacy nodes and privacy edges; Based on the privacy degree, privacy-balanced partitioning is performed between servers, and privacy-aware partitioning is performed within the server.
2. According to claim 1, a reliable and efficient graph partitioning method for large-scale private graph data in a cluster is characterized in that: The privacy balance partitioning includes the following steps: In the coarsening stage, the privacy-based label propagation algorithm confirms the merged candidate nodes, retains the mapping relationship of the original nodes, and accumulates the privacy of the nodes and connected edges until the node size is less than the set threshold, thus obtaining the compressed graph data; In the partitioning stage, a fast partitioning algorithm is used to initially partition the compressed graph data into k partitions, where k represents the number of servers in the cluster; In the refinement phase, the graph data is restored to its original size layer by layer according to the node mapping relationship.
3. According to claim 1, a reliable and efficient graph partitioning method for large-scale private graph data in a cluster is characterized in that: The privacy-aware partitioning includes the following steps: Identify the privacy of graph data; Update the privacy of nodes based on the privacy propagation algorithm; Use priority queue to allocate privacy nodes and privacy edges to TEE end; Determine TEE end load τ TEE With REE end load τ REE And the maximum cache capacity of the TEE side τ max_TEE ; If τ TEE <τ REE And τ TEE <τ max_TEE , use the priority queue to cache the semi-private node copies to the TEE side until the TEE side and the REE side are load balanced, or the maximum cache capacity of the TEE side is reached; If τ TEE >τ REE , or τ TEE =τ max_TEE , stop adding nodes and edges to the TEE end; if the private edge exists on the REE end, hide or encrypt the private edge on the REE end; Among them, the TEE is a trusted execution environment, REE is a regular execution environment, TEE and REE rely on shared memory to communicate, and privacy data is calculated on the TEE side; the initial value of the privacy degree is 0, 1, 2, 3, 0 represents regular data, 1 represents secret data, 2 represents confidential data, and 3 represents top secret data; the semi-private node is the original regular node whose privacy value is between 0 and 1 after the privacy degree is updated.
4. According to claim 1, a reliable and efficient graph partitioning method for large-scale privacy graph data in a cluster is characterized in that: Use keyword matching to match privacy attributes in the data, or use manual labeling to mark private data.
5. According to claim 2, a reliable and efficient graph partitioning method for large-scale privacy graph data in a cluster is characterized in that: The coarsening stage includes the following steps: Traverse the graph data structure and obtain the neighbor set of each node; Based on the neighbor set, obtaining a node pair merging cost; According to the merging cost, node pairs with low cost are selected to be merged first.
6. According to claim 2, a reliable and efficient graph partitioning method for large-scale private graph data in a cluster is characterized in that: The coarsening stage includes the following steps: ω(u,v) represents the edge weight between the nodes to be merged, deg(u) and deg(v) represent the node degrees, and p represents the edge weight between the nodes to be merged. u 、p v 、p (u,v) Represent the privacy of nodes v,v and the privacy of the edge between u and v, respectively, and obtain the merging cost of candidate nodes According to the merging cost, the two candidate nodes with the lowest cost are selected for merging, the original node mapping relationship is retained, the privacy of nodes and connected edges is accumulated, multi-threading is used to process the merging work simultaneously, and a local search algorithm is used to avoid repeatedly merging the same node until the node size is smaller than the set threshold.
7. According to claim 2, a reliable and efficient graph partitioning method for large-scale privacy graph data in a cluster is characterized in that: The division phase includes the following steps: Select the initial node, obtain the node degree deg(u), and calculate the node centrality Select the top k nodes according to the node centrality; With k nodes as the center, a breadth-first search algorithm is used to construct a subgraph; by The target number of nodes is divided into partitions. If the number of nodes in a partition exceeds 1.5θ, the above process is repeated to split the original large partition into new sub-partitions. If nodes u and v are located in different partitions, the node movement benefit is calculated using the benefit function, and the partition with the maximum movement benefit is selected to move the nodes to this partition. If the movement benefit is low, the movement is abandoned until a preliminary division result of k partitions is formed.
8. According to claim 6, a reliable and efficient graph partitioning method for large-scale private graph data in a cluster is characterized in that: The mobility benefits include load balancing benefits, privacy data balancing benefits, locality benefits, and communication overhead benefits; load balancing is defined as the standard deviation of the partition load and the average load. The load balancing benefit is the difference between the load balancing value after the move and the load balancing value before the move. The privacy ratio is defined as the ratio of private data to the partition data volume The privacy data balance gain represents the difference between the privacy data balance index after the move and the privacy data balance index before the move. To measure whether the movement is conducive to the balance of privacy data, the communication cost is defined as the weight of the cross-partition edge The communication cost benefit represents the difference between the number and weight of cross-partition edges after the move and before the move. Locality is defined as the number of cross-partition edges, and locality benefit is the difference between the communication cost after moving and the communication cost before moving, that is, Calculate the moving benefit G = αg' based on the above weights l (u,i,j)+βg' c (u,i,j)+γg' p (u,i,j)+θg' n (u,i,j); where α, β, γ, and θ are weight coefficients, and their initial defined values are all 1, which are dynamically adjusted by the user according to the target requirements.
9. A trusted and efficient graph partitioning system for large-scale private graph data in a cluster according to any one of claims 1 to 8, characterized in that: include: An identification and marking module, used to identify privacy nodes and privacy edges in graph data, and mark the privacy degrees of the privacy nodes and privacy edges; The partitioning module is used to perform privacy-balanced partitioning between servers and privacy-aware partitioning within servers based on the privacy level.
10. A computing device, characterized in that: The method comprises a processor and a memory; the memory is used to store computer-executable instructions; the processor is used to execute the computer-executable instructions so that the computing device executes the method according to any one of claims 1 to 8.