Privacy protection method and system for community discovery in social network
By using localized differential privacy technology in social networks to noise processing user data and upload it to the blockchain, combined with the PBFT consensus mechanism of multi-layer packets, the problems of privacy leakage and information entropy deviation in the community discovery process in social networks are solved, and consensus efficiency is improved.
Patent Information
- Application Number
- CN202510544444.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-06-20
AI Technical Summary
There are problems such as privacy leakage risks, information entropy deviations, centralized storage risks and inefficiency in large-scale network consensus during the community discovery process in social networks.
Localized differential privacy technology is used to process user data in social networks and upload the noise-added data to the blockchain. The blockchain adopts a PBFT consensus mechanism based on multi-layer packets.
It reduces the risk of privacy leakage, controls information entropy deviation, avoids centralized storage risks, and improves the efficiency of large-scale network consensus.
Smart Images

Figure CN120180504A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of blockchain, and specifically relates to a privacy protection method and system for community discovery in social networks. Background Art
[0002] With the rapid development of network technology and communication technology and the popularization of mobile intelligent devices, people's communication has become more convenient, and various social platforms have emerged. Common social platforms include Weibo, WeChat, Douyin, and Twitter, etc. Through these social platforms, people distributed in different places are connected together, forming a huge human-centered social network. As an extension of the real human society in the virtual network, the social network has the characteristics of strong real-time, huge data, wide coverage, and complex dimensions. Although the academic community has not unified the definition of the attributes of the social network at present, affected by people's living environment, interests and hobbies, etc., the social network follows the community structure characteristics of "people flock together" and reflects the real social relations to a certain extent.
[0003] The social network has accumulated a large amount of user information (such as personal identity information, medical information, and trajectory information, etc.) and user behavior traces. In the process of community discovery, it is necessary to obtain the social information of users. However, among these information, there are private sensitive information. If users upload these data without reservation, or the server does not adopt any privacy protection technology in centralized data analysis, once the privacy data is leaked and obtained and misused by lawbreakers, the user's privacy security and even life will be threatened. This is also an urgent problem to be solved in the process of community discovery. Traditional differential privacy requires global noise addition, which is easy to blur the community structure characteristics and lead to the distortion of node dependency relationships. Moreover, the communication complexity of the traditional Practical Byzantine Fault Tolerance (PBFT) algorithm is O(N 2 ) When the number of nodes exceeds 100, the delay and communication cost increase exponentially. Summary of the Invention
[0004] Object of the Invention: Aiming at the deficiencies of the prior art, the object of the present invention is to provide a privacy protection method and system for community discovery in social networks, which can reduce the risk of privacy leakage and control the information entropy deviation by adopting local differential privacy and hierarchical PBFT algorithm.
[0005] Technical Solution: To achieve the above object of the invention, the present invention adopts the following technical solutions:
[0006] In the first aspect, the present invention provides a privacy protection method for community discovery in social networks, including the following steps:
[0007] Step 1: Use the localized differential privacy technology based on the Laplace mechanism to add noise to the user data in the social network, and then upload the noisy data to the blockchain; the blockchain adopts the PBFT consensus mechanism based on multi-layer grouping;
[0008] Step 2: Take the users on the social platform as nodes of the network model, calculate the information entropy of each node based on the noisy user data, determine the set of central nodes, and use the nodes with information entropy values less than the set threshold as central nodes; calculate the mutual information between the central node and other nodes, and divide the nodes into communities according to the mutual information values.
[0009] Furthermore, in step 1, the differential privacy model adopted selects Laplace distribution to add noise to the data, including the following steps:
[0010] Step 1.1: Obtain the user's public data through the API interface provided by the social platform, and collect the data authorized by the user with the user's consent;
[0011] Step 1.2: Noise the original data by using the noise that complies with the Laplace mechanism to generate a noisy data set to protect user privacy;
[0012] Step 1.3: Upload the noisy data to the blockchain.
[0013] Furthermore, the step 2 includes the following steps:
[0014] Step 2.1: Abstract the social network into an undirected graph;
[0015] Step 2.2: First, calculate the information entropy value of each node. This value is based on the data characteristics of the node and reflects the certainty of the node information. Then set a threshold m and obtain it by statistically analyzing the information entropy values of all nodes. Finally, determine the node with an information entropy value less than the threshold m as the central node.
[0016] The information entropy of each node X is calculated through normal distribution, and the formula is:
[0017]
[0018] Where e is the base of the natural logarithm, m is the number of node attributes; Σ is the covariance matrix of the attributes, with a dimension of m×m; |Σ| is the determinant of the covariance matrix, reflecting the overall uncertainty among multiple attributes;
[0019] Step 2.3: Calculate the mutual information between each central node and the remaining nodes. If the mutual information is greater than a predetermined threshold v, the node and the current central node are classified as the same community.
[0020] The calculation formula of the mutual information between node X and node Y is as follows:
[0021]
[0022] Among them, p(x, y) is the joint probability distribution of X and Y, representing the probability that two nodes simultaneously have attributes x and y, and p(x) and p(y) are marginal probability distributions, representing the probabilities that a node independently has attribute x or y, respectively.
[0023] Further, in step 1, based on the multi-layer grouped PBFT consensus mechanism, the social network user nodes are classified according to node attributes and then consensus is carried out layer by layer.
[0024] Further, the PBFT consensus algorithm based on multi-layer grouping is divided into a preparation stage and a consensus stage. In the preparation stage, regulatory nodes (special nodes responsible for supervising the consensus process, ensuring data validity and security in the hierarchical network architecture) and hierarchical network nodes (ordinary nodes assigned to different levels according to node attributes and participating in local consensus at their respective levels) are classified and selected; Consensus stage: By recursively inserting the submission and response stages of the first N - 1 layer nodes into the PBFT consensus algorithm as sub-layer algorithms, the Nth layer nodes use the DPoS consensus algorithm to select decision-makers, and start the PBFT consensus to achieve global consensus and generate blocks.
[0025] Further, the specific process of the PBFT consensus algorithm based on multi-layer grouping includes:
[0026] Group the nodes according to node attributes, and select a regulatory node in each group to be responsible for coordination and supervision;
[0027] Divide the entire network into N layers, with each layer containing several groups;
[0028] In each of the first N - 1 layers, use the PBFT consensus algorithm to achieve local consensus, and recursively transfer the consensus results of the first N - 1 layers to the next layer;
[0029] The Nth layer uses the DPoS algorithm to elect decision-makers, and the decision-makers start the PBFT consensus, ultimately achieving global consensus and generating blocks.
[0030] In a second aspect, the present invention provides a privacy protection system for community discovery in a social network, including:
[0031] A data processing module, configured to perform noise addition processing on user data in the social network using local differential privacy technology based on the Laplace mechanism, and then upload the noise-added data to the blockchain; the blockchain adopts a multi-layer grouped PBFT consensus mechanism;
[0032] A community discovery module, which takes users on a social platform as nodes of a network model, calculates the information entropy of each node based on the noisy user data, determines a set of central nodes, and takes the nodes with information entropy values less than the set threshold as central nodes; calculates the mutual information between the central nodes and other nodes, and divides the nodes into communities according to the mutual information values.
[0033] In a third aspect, the present invention provides a computer system, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the privacy protection method for community discovery in a social network are implemented.
[0034] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the privacy protection method for community discovery in a social network are implemented.
[0035] In a fifth aspect, the present invention provides a computer program product including a computer program, and when the computer program is executed by a processor, the steps of the privacy protection method for community discovery in a social network are implemented.
[0036] Beneficial effects: Aiming at the problems existing in the existing privacy protection technology in social network community discovery, such as the conflict between privacy and data availability, the risk of centralized storage, the low efficiency of large-scale network consensus, and the insufficient dynamic data processing ability, the present invention combines local differential privacy (LDP) and blockchain technology to perform distributed storage after adding noise to the data at the user end, reducing the risk of privacy leakage and the deviation of information entropy; at the same time, a PBFT consensus mechanism based on multi-layer grouping (F-PBFT) is designed, and through node attribute classification and hierarchical recursive consensus, the communication complexity of traditional PBFT is optimized from O(N 2 ) to O(N). When the number of nodes reaches 1000, the communication cost is reduced and the consensus delay is shortened; combined with dynamic screening of central nodes by information entropy and calculation of local mutual information, the calculation amount is reduced and the accuracy of community division is improved, effectively adapting to the dynamic changes of the social network. Description of the Drawings
[0037] Figure 1 is the flowchart of the method of the embodiment of the present invention.
[0038] Figure 2 is a social network graph.
[0039] Figure 3 is a schematic diagram of a community discovery storage model based on the blockchain idea.
[0040] Figure 4 is a schematic diagram of the PBFT consensus algorithm based on multi-layer grouping (F-PBFT).
[0041] Figure 5 It is the flowchart of F-PBFT.
[0042] Figure 6 It is the comparison graph of the change of communication complexity with the increase of the total number of nodes.
[0043] Figure 7 It is the comparison graph of the change of consensus latency with the increase of the total number of nodes.
[0044] Figure 8 It is the comparison graph of information entropy and mutual information before and after differential privacy.
[0045] Figure 9 It is the comparison graph of information entropy in different value ranges. Specific implementation manners
[0046] The present invention will be described in detail below with reference to the accompanying drawings and specific implementation manners. The following embodiments are used to illustrate the present invention, but are not used to limit the protection scope of the present invention.
[0047] A privacy protection method for community discovery in social networks disclosed in an embodiment of the present invention mainly includes the following steps:
[0048] Step 1: In the preprocessing stage, use the localization differential privacy technology based on the Laplace mechanism to perform noise addition processing on the original user data in the social network, and then store the noisy data in Ethereum with the idea of blockchain to ensure data security with this double guarantee. Among them, the blockchain adopts the PBFT consensus mechanism based on multi-layer grouping.
[0049] Step 2: Take the users on the social platform as the nodes of the network model, calculate the information entropy of each node, determine the set of central nodes, and use the nodes with smaller information entropy values as the central nodes. Calculate the mutual information between the central nodes and other nodes, and divide the nodes into the same community according to the mutual information values.
[0050] The embodiments of the present invention mainly involve the following improvement points: 1) A community discovery algorithm for social networks based on differential privacy technology is proposed. This algorithm adds differential privacy technology to the preprocessing stage of the community discovery algorithm that combines information entropy and mutual information (MI), and uses localization differential privacy (LDP) based on the Laplace mechanism to initialize the data on the social platform, that is, extract the noise that conforms to the Laplace distribution to perturb the original data, so as to realize the protection of the original data; 2) A community discovery data storage mechanism based on the idea of blockchain is proposed, that is, upload the noisy data to the blockchain and borrow the security features of blockchain technology to strengthen data privacy; at the same time, a PBFT consensus mechanism based on multi-layer grouping is proposed to improve the consensus efficiency.
[0051] Specifically, as Figure 1 shown, the above step 1 mainly includes:
[0052] 1) Data collection: Through the API interface provided by the social platform, obtain the public data of users. With the consent of users, collect the data authorized by users for more in-depth community discovery research.
[0053] 2) Data noise addition: Use local differential privacy with the Laplace mechanism to add noise to the data of each user node to generate a noisy dataset to protect user privacy.
[0054] 3) Data uploading to the blockchain: Upload the noisy data to the blockchain, and utilize the characteristics of the blockchain such as decentralization and immutability to ensure the security and integrity of the data.
[0055] The above step 2 mainly includes:
[0056] 1) Network modeling: As Figure 2 shown, abstract the social network as an undirected graph G = 〈V, E〉, where V represents the node set and E represents the edge set;
[0057] 2) Calculate central nodes: First, calculate the information entropy value for each node. This value is based on the data characteristics of the node and reflects the certainty of node information. Then, set a threshold m, which is obtained through statistical analysis of the information entropy values of all nodes. Finally, determine the nodes with information entropy values less than the threshold m as central nodes. These nodes are usually old users in the social network or users with more browsing data, and the information is relatively certain. Through the above steps, users suitable to be central nodes can be effectively screened out, providing a clear operation process and basis for the selection of central nodes.
[0058] The information entropy of each node X is calculated through a normal distribution, and the formula is:
[0059]
[0060] where e is the base of the natural logarithm, m is the number of node attributes; Σ is the covariance matrix of the attributes, with a dimension of m×m; |Σ| is the determinant of the covariance matrix, reflecting the overall uncertainty among multiple attributes.
[0061] 3) Calculate mutual information: Calculate the dependence relationship between each central node and the remaining nodes. The higher the degree of dependence, the higher the possibility of classifying it into the same community as the central node. In this process, borrow the concept of mutual information. If the mutual information is greater than the set threshold v, then classify this node into the same community as the current central node.
[0062] The calculation formula for the mutual information between node X and node Y is as follows:
[0063]
[0064] Among them, p(x, y) is the joint probability distribution of X and Y, representing the probability that two nodes simultaneously have attributes x and y. p(x) and p(y) are marginal probability distributions, representing the probabilities that a node independently has attribute x or y, respectively.
[0065] The PBFT consensus mechanism based on multi-layer grouping includes:
[0066] 1) Storage mechanism: First, in this embodiment, a community discovery data storage mechanism based on the blockchain concept is constructed. As Figure 3 shown, there are three types of nodes involved in this mechanism, namely data nodes, storage nodes, and data analysis nodes. In a social network, when a user registers an account on a social platform, they need to fill in their basic information for real-name authentication, which involves sensitive information such as name, gender, birthday, mobile phone number, and ID card. To ensure the security of this sensitive information, the data nodes are responsible for performing data noise addition processing, the storage nodes are responsible for data aggregation and storage, and the data analysis nodes are responsible for managing and analyzing the uploaded data.
[0067] 2) PBFT consensus mechanism based on multi-layer grouping: The traditional PBFT consensus mechanism solves the problem of low efficiency of the original Byzantine, reducing the algorithm complexity. However, PBFT is only applicable to consortium chains / private chains, and the communication complexity is too high. Therefore, in this embodiment, a PBFT consensus algorithm based on multi-layer grouping is proposed on the basis of the traditional PBFT consensus algorithm to reduce the communication complexity. The traditional PBFT consensus algorithm is divided into three stages: pre-prepare, prepare, and commit stages. The PBFT consensus algorithm based on multi-layer grouping is divided into two stages, and before consensus, the social network user nodes are first classified according to node attributes and then consensus is carried out layer by layer. As Figure 4 shown, in the PBFT consensus mechanism model based on multi-layer grouping, this algorithm is divided into two stages. Preparation stage: Classify according to the attributes of the nodes in the community discovery process, select the supervision nodes to be responsible for supervising the decision-makers, and then layer the various network nodes. Consensus stage: Use a recursive method to insert the submission and response stages of the first N - 1 layer nodes into the PBFT consensus algorithm as sub-layer algorithms; the nodes in the Nth layer use the method in the DPoS consensus algorithm to select the decision-makers, and the decision-makers start the PBFT consensus to achieve the final global consensus and generate blocks. The specific process is as Figure 5As shown, it includes: grouping nodes according to node attributes, selecting regulatory nodes in each group to be responsible for coordination and supervision; dividing the entire network into N layers, with each layer containing several groups; using the PBFT consensus algorithm in each of the first N-1 layers to achieve local consensus, and recursively transmitting the consensus results of the first N-1 layers to the next layer; using the DPoS algorithm in the Nth layer to elect decision-makers, and the decision-makers initiate the PBFT consensus to finally achieve global consensus and generate blocks.
[0068] Exemplarily, Figure 4 Among them, data analysis nodes (S1 - S8): responsible for executing data analysis tasks (such as feature extraction, pattern recognition), usually served by nodes with high computing power. Example: S1 may process the community discovery algorithm, and S2 is responsible for privacy noise addition calculation.
[0069] Data nodes (C1 - C11): store the original social network data (such as user relationships, interaction records), and undertake data preprocessing tasks. Example: C1 stores the user friend list, and C2 caches the dynamically updated interaction data.
[0070] Storage nodes (F1 - F6): provide distributed persistent storage (such as blockchain or IPFS) to ensure data immutability. Example: F1 stores the noise-added encrypted data, and F2 records the consensus log.
[0071] Hierarchical progression: The first line classifies nodes by functional attributes (data, analysis, storage), and the second line starts to layer by network attributes (such as community affiliation, geographical location).
[0072] Consensus initiation: Integrate the classified nodes in the first line into the first layer and execute the local PBFT consensus (such as C1 - C11 data nodes verifying data consistency). Example: In the first layer, C1 - C11 data nodes reach local consensus through PBFT and generate a temporary block summary.
[0073] Recursive consensus: The consensus result of the first layer in the second line is passed as input to the second layer in the third line for further verification and aggregation.
[0074] Hierarchical division of labor: The nodes in the second layer may consist of cross-community regulatory nodes (such as high-reputation nodes among S1 - S8); perform cross-domain verification on the summaries submitted by the first layer to ensure global consistency. Example: The regulatory nodes in the second layer check the block signatures of the first layer and generate a cross-community summary after eliminating malicious data.
[0075] Classification mapping of user nodes:
[0076] After classifying social network user nodes by attributes, they are mapped to three types of nodes in the figure:
[0077] Ordinary users → data nodes (C series): Provide original interaction data;
[0078] Highly active users → data analysis nodes (S series): Execute computing tasks;
[0079] High storage resource users → storage nodes (F series): Provide distributed storage.
[0080] Hierarchical participation of user nodes:
[0081] In the consensus stage, user nodes are classified into different levels according to their categories:
[0082] Level 1: Ordinary users (C series) participate in local consensus;
[0083] Level 2: Highly active users (S series) act as supervision nodes;
[0084] Level N: Decision-makers selected by DPoS (such as high-weight nodes in the F series) complete global confirmation.
[0085] Initial stage (first row): User nodes are classified into three categories: data, analysis, and storage according to their functions, and the role division is clarified.
[0086] Consensus stage (second and third rows):
[0087] Level 1: Data nodes (C series) execute local PBFT consensus to generate temporary blocks;
[0088] Level 2: Analysis nodes (S series) verify and aggregate cross-community data and submit it to a higher level;
[0089] Level N: Storage nodes (decision-makers in the F series) complete global PBFT consensus to generate the final block.
[0090] Recursive logic: Each layer only processes data from adjacent layers, and the communication complexity is reduced from O(N 2 ) to O(N).
[0091] The advantages of this embodiment over the prior art will be described below with reference to several experimental examples.
[0092] Experimental example 1: The communication complexity refers to the number of communications between blockchain nodes. The fewer the number of communications, the lower the communication cost, which means the lower the communication complexity of the system.
[0093] Assume that the total number of nodes in the social network is N. Then the consensus communication cost of traditional PBFT is O(N 2 ), and at the same time, assume that the PBFT consensus algorithm based on multi-layer grouping (F-PBFT) proposed in this embodiment divides the entire network into M layers, and the i-th layer has n i nodes. Then the communication cost (C M)The derivation is as follows:
[0094]
[0095] Based on the above consensus, assuming that each group is assigned 3 nodes, the minimum communication cost can be derived as:
[0096] M max = log3(2N + 1) - 1
[0097]
[0098] where M max is the maximum number of layers, and N is the total number of nodes.
[0099] The above communication cost complexity is linearly increasing with the total number of nodes. Compared with the traditional PBFT consensus algorithm, the PBFT consensus algorithm based on multi-layer grouping (F-PBFT) has lower complexity. As Figure 6 shown, the total number of nodes gradually increases starting from 20, and the communication complexity between nodes is compared.
[0100] It can be seen from the results that in the same network scenario, as the total number of nodes increases, the communication complexities of both the traditional PBFT consensus algorithm and the PBFT consensus algorithm based on multi-layer grouping (F-PBFT) increase. However, as the number of nodes increases, the growth rate of the communication cost of the traditional PBFT consensus algorithm and the CPBFT consensus algorithm is relatively large. The PBFT consensus algorithm based on multi-layer grouping (F-PBFT) proposed in this embodiment increases linearly, and even when the number of nodes increases, the growth rate of its communication cost is relatively small, which is very suitable for scenarios with a large number of nodes in social networks.
[0101] Experimental Example 2: The consensus latency of the consensus algorithm refers to the time from when a request is sent to when the request is confirmed. The less time it takes, the faster the consensus is reached. The consensus latency TD can be expressed as TD = T start - T end , where T start is the request initiation time, and T end is the confirmation time of the request. Assuming that the number of nodes gradually increases starting from 20, the measured consensus latency is obtained, and the results are as Figure 7 shown.
[0102] It can be seen from the results that the consensus delay increases with the increase in the number of nodes in the system. Compared with the traditional PBFT consensus algorithm and the improved PBFT algorithm CPBFT, the PBFT consensus algorithm based on multi-layer grouping (F-PBFT) grows more slowly, which means that in the same number of nodes, the delay of the PBFT consensus algorithm based on multi-layer grouping (F-PBFT) is much smaller than that of the traditional PBFT consensus algorithm and the improved PBFT algorithm CPBFT.
[0103] Experimental Example 3: Select several data in the ranges of [0, 20] and [0, 40] respectively, set the parameter ε = 1, calculate the information entropy and mutual information of the original data and the noisy data and compare them. The results are as Figure 8 shown. In the figure, (a), (b), (c), and (d) select the data in [0, 20], and (e) and (f) select the data in [0, 40]. Figures (a), (c), and (e) are the information entropy values before and after introducing differential privacy, and figures (b), (d), and (f) are the mutual information values before and after introducing differential privacy. Among them, the blue bars represent the information entropy values before introducing differential privacy, and the red bars represent the information entropy values after introducing differential privacy. It can be found from the figure that there is not much difference in the results of the information entropy and mutual information between the original data and the noisy data, which can prove that differential privacy will not cause result deviation to the community division algorithm.
[0104] Experimental Example 4: Taking the normal distribution as the probability density function, as Figure 9 shown, calculate the information entropy values of the data in the numerical ranges of [0, 5], [0, 10], [0, 20], [0, 40], and [40, 100] respectively.
[0105] It can be seen from the above results that: the wider the numerical range of the data, the smaller the information entropy value and the greater the uncertainty. In this embodiment, numerical data is used for experiments, and the information entropy value can be used to judge whether the numerical data has certainty, which can be used as a judgment criterion for the central node, thus proving the feasibility of using the information entropy value to determine the central node.
[0106] In this embodiment, by preferentially establishing the central node, the calculation amount is reduced. Assuming that n nodes are set, the original algorithm needs to calculate n 2 -n mutual informations, while now only the mutual informations between the central node and the remaining nodes need to be calculated. If m central nodes are finally established, only m 2 -m mutual informations need to be calculated (where m << n). By comparison, the calculation amounts of the two algorithms m 2 -m << n 2 -n. The calculation amount of the current algorithm is significantly less than that of the original algorithm, improving the community discovery efficiency.
[0107] Experimental Example 5: In the experiment, 10 nodes were set up. After calculating the information entropy value of each node, the average value of the information entropy was set as the threshold. Thus, the central nodes consisted of nodes smaller than this threshold. Eventually, 3 central nodes were established in the experiment. Next, the mutual information values between the remaining nodes and the central nodes were calculated respectively, and the nodes with high mutual information values were grouped into the same community with the corresponding central node. Table 1 shows the mutual information results with a relatively high degree of dependence between the remaining nodes and the 3 central nodes. It can be seen that the nodes grouped into the same community with central node A are: Node 1, Node 3, Node 4, Node 5, Node 8, Node 10; the nodes grouped into the same community with central node B are: Node 1, Node 2, Node 5; and the nodes in the same community with central node C are: Node 1, Node 2, Node 4, Node 5, Node 10.
[0108] Table 1 Mutual Information Calculation Results of Community Discovery Algorithm
[0109]
[0110] For easy analysis, in this experiment, 10 nodes were first set up, and the entire algorithm process was run. It was found that the community discovery algorithm combining the information entropy value and the mutual information value can effectively divide communities and reduce the computational amount, which is feasible.
[0111] Experimental Example 6: Given n nodes and the data contained in each node, v central nodes (v << n) were determined in the community division. The time complexity of the algorithm in this embodiment is as follows:
[0112] 1) The time complexity required to calculate the information entropy of all nodes is O(n);
[0113] 2) The time complexity required to calculate the mutual information between v central nodes and the remaining n - v nodes is O(v(n - v));
[0114] Based on the above analysis, the time complexity of the algorithm in this embodiment is n + nv - v 2 . Since v << n can be ignored, the time complexity of the entire algorithm is O(n). In addition, as shown in Table 2, the time complexities of the following several types of algorithms were analyzed, and it was found that the algorithm proposed in this embodiment is lower than that of some algorithms.
[0115] Table 2 Time Complexity Comparison
[0116]
[0117] Based on the same inventive concept, a privacy protection system for community discovery in a social network disclosed in an embodiment of the present invention includes: a data processing module, configured to perform noise addition processing on user data in the social network by using local differential privacy technology based on the Laplace mechanism, and then upload the noise-added data to the blockchain; the blockchain adopts a PBFT consensus mechanism based on multi-layer grouping; a community discovery module, configured to use users on the social platform as nodes of a network model, calculate the information entropy of each node based on the noise-added user data, determine a set of central nodes, and use nodes with information entropy values less than a set threshold as central nodes; calculate the mutual information between the central nodes and other nodes, and divide the nodes into communities according to the mutual information values.
[0118] An embodiment of the present invention also discloses a computer system, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of a privacy protection method for community discovery in a social network are implemented.
[0119] An embodiment of the present invention also discloses a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the steps of a privacy protection method for community discovery in a social network are implemented.
[0120] An embodiment of the present invention also discloses a computer program product, including a computer program. When the computer program is executed by a processor, the steps of a privacy protection method for community discovery in a social network are implemented.
[0121] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or a controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, so that when the program codes are executed by the processor or the controller, the steps of the method of the present invention are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed as an independent software package partially on the machine and partially on a remote machine, or executed entirely on a remote machine or server. Where the present invention is not described in detail, it is all well-known techniques to those skilled in the art.
Claims
1. A privacy protection method for community discovery in social networks, characterized in that: The following steps are involved: Step 1: Use the localized differential privacy technology based on the Laplace mechanism to add noise to the user data in the social network, and then upload the noisy data to the blockchain; the blockchain adopts the PBFT consensus mechanism based on multi-layer grouping; Step 2: Use users on the social platform as nodes of the network model, calculate the information entropy of each node based on the noisy user data, determine the central node set, and use the nodes with information entropy values less than the set threshold as the central nodes; Calculate the mutual information between the central node and other nodes, and divide the nodes into communities based on the mutual information value.
2. A privacy protection method for community discovery in a social network according to claim 1, characterized in that: In step 1, the differential privacy model adopted selects Laplace distribution to add noise to the data, including the following steps: Step 1.1: Obtain the user's public data through the API interface provided by the social platform, and collect the data authorized by the user with the user's consent; Step 1.2: Noise the original data by using the noise that complies with the Laplace mechanism to generate a noisy data set to protect user privacy; Step 1.3: Upload the noisy data to the blockchain.
3. The privacy protection method for community discovery in social networks according to claim 1, characterized in that: The step 2 includes the following steps: Step 2.1: Abstract the social network into an undirected graph; Step 2.2: First, calculate the information entropy value of each node. This value is based on the data characteristics of the node and reflects the certainty of the node information. Then set a threshold m and obtain it by statistically analyzing the information entropy values of all nodes. Finally, determine the node with an information entropy value less than the threshold m as the central node. The information entropy of each node X is calculated through normal distribution, and the formula is: Where e is the base of the natural logarithm, m is the number of node attributes; Σ is the covariance matrix of the attributes, with a dimension of m×m; |Σ| is the determinant of the covariance matrix, reflecting the overall uncertainty among multiple attributes; Step 2.3: Calculate the mutual information between each central node and the remaining nodes. If the mutual information is greater than a predetermined threshold v, the node and the current central node are classified as the same community. The calculation formula of the mutual information between node X and node Y is as follows: Where p(x,y) is the joint probability distribution of X and Y, indicating the probability that two nodes have both attributes x and y. p(x) and p(y) are marginal probability distributions, indicating the probability that a node has attribute x or y alone, respectively.
4. The privacy protection method for community discovery in social networks according to claim 1, characterized in that: In step 1, based on the PBFT consensus mechanism of multi-layer grouping, the social network user nodes are classified according to node attributes, and then consensus is performed in layers.
5. The privacy protection method for community discovery in social networks according to claim 4, characterized in that: The PBFT consensus algorithm based on multi-layer grouping is divided into a preparation phase and a consensus phase. In the preparation phase, supervisory nodes are classified and selected and network nodes are layered. Consensus stage: The submission and response stages of the first N-1 layers of nodes are recursively inserted into the PBFT consensus algorithm as a sub-layer algorithm. The Nth layer of nodes uses the DPoS consensus algorithm to select decision makers, start the PBFT consensus to achieve global consensus and generate blocks.
6. The privacy protection method for community discovery in a social network according to claim 5, characterized in that: The specific process of the PBFT consensus algorithm based on multi-layer grouping includes: Nodes are grouped according to their attributes, and supervisory nodes are selected in each group to be responsible for coordination and supervision; Divide the entire network into N layers, each layer contains several groups; Each of the first N-1 layers uses the PBFT consensus algorithm to reach a local consensus, and recursively passes the consensus results of the first N-1 layers to the next layer; The Nth layer uses the DPoS algorithm to elect decision makers, who initiate the PBFT consensus and eventually reach a global consensus and generate blocks.
7. A privacy protection system for community discovery in social networks, characterized in that: include: A data processing module, used to perform noise processing on user data in social networks using a localized differential privacy technology based on the Laplace mechanism, and then upload the noised data to a blockchain; the blockchain adopts a PBFT consensus mechanism based on multi-layer grouping; The community discovery module is used to use users on the social platform as nodes of the network model, calculate the information entropy of each node based on the noisy user data, determine the set of central nodes, and use the nodes with information entropy values less than the set threshold as central nodes; Calculate the mutual information between the central node and other nodes, and divide the nodes into communities based on the mutual information value.
8. A computer system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is executed by a processor, the steps of a privacy protection method for community discovery in a social network are implemented according to any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of a privacy protection method for community discovery in a social network are implemented according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of a privacy protection method for community discovery in a social network are implemented according to any one of claims 1 to 6.