Digraph synthesis method and system based on differential privacy
By community division and sequence processing of directed graphs based on differential privacy, a synthetic social network graph is generated, which solves the shortcomings of directed graph processing in the prior art and improves the usability and privacy protection of graphs.
Patent Information
- Application Number
- CN202510320833.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to effectively process directed graphs and generate synthetic social network graphs, especially in terms of privacy protection and availability.
Using a differential privacy-based method, the nodes in the directed graph are divided community by community-based division, incoming, outgoing and edge sequences are processed, noise is added and constrained inference is performed, target incoming, outgoing and edge sequences are generated, and the target composite graph is finally synthesized.
The availability and accuracy of synthetic social network graphs are improved, the protection of user privacy is enhanced, and the shortcomings of directed graph processing in the prior art are solved.
Smart Images

Figure CN120144833A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of social network privacy protection, and particularly to a method and system for synthesizing a directed graph based on differential privacy. Background Art
[0002] A graph is a very important data structure. Graphs can be used to represent relationships and connections in various practical problems, such as social networks, road networks, organizational structures, etc. The applications of graphs are very extensive. For example, a social network can be represented by a graph, where each person or entity is a node, and the edges represent the relationships between people. A synthetic social network graph refers to a virtual social network graph generated using computer programs and algorithms, which is used to study the social network structure and behavior patterns. However, the synthetic social network graph may indirectly expose user privacy in some cases. For example, an attacker may combine the data in the synthetic social network graph with other public information or external data sources for information recombination and inference, so as to identify the identities and relationships of real users.
[0003] A directed graph is a non-linear data structure used to represent directional relationships. It consists of vertices (nodes) and directed edges (arcs), and each edge has a starting point and an ending point, representing a one-way relationship from one vertex to another. Differential Privacy (DP) is a technical framework for privacy protection. It has a strict mathematical definition and a robust privacy metric, aiming to provide strong privacy protection when publishing or analyzing personal data while maintaining the availability and effectiveness of the data. Differential privacy provides a quantitative privacy protection metric called Privacy Budget (PB).
[0004] The currently commonly used algorithms for synthesizing social network graphs are mainly degree-based synthesis algorithms. However, this method can only ensure that the node degrees in the synthesized social network graph are theoretically consistent with those of the original graph, and the generated synthesized social network graph lacks usability. Other synthesis methods also have their respective defects, such as insufficient privacy protection and low usability due to high sensitivity. At the same time, most of the currently existing algorithms for synthesizing social network graphs are only applicable to undirected graphs. Therefore, how to process a directed graph to obtain a synthesized social network graph is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0005] In view of this, the embodiments of the present invention provide a method and system for synthesizing a directed graph based on differential privacy to solve the problem in the prior art of how to process a directed graph to obtain a synthesized social network graph; that is to say, the embodiments of the present invention can improve the usability and accuracy of the synthesized social network graph.
[0006] According to one aspect of the present invention, a directed graph synthesis method based on differential privacy is provided. The directed graph synthesis method based on differential privacy includes: obtaining a directed graph and a plurality of privacy budget values, and performing community partitioning on the nodes in the directed graph to obtain a community partitioning result; obtaining sub-graph node indication data and sub-graph in-degree indication data corresponding to each community in at least one community, determining an in-degree sequence based on the sub-graph in-degree indication data, and processing the in-degree sequence to obtain a target in-degree sequence; obtaining sub-graph node indication data and sub-graph out-degree indication data corresponding to each community in at least one community, determining an out-degree sequence based on the sub-graph out-degree indication data, and processing the out-degree sequence to obtain a target out-degree sequence; calculating the number of edges between all communities to obtain an edge sequence, and performing noise addition processing on the edge sequence to obtain a target edge sequence; obtaining node indication data of the directed graph, calculating a node similarity set according to the node indication data, and performing noise addition processing on the node similarity set to obtain a target node similarity set; generating an edge set within the community according to the target in-degree sequence, the target out-degree sequence, and the target node similarity set of each community in at least one community, where the edge set within the community includes edges within at least one community; generating an edge set between communities according to the target edge sequence and the target node similarity set of each community in at least one community, where the edge set between communities includes edges between at least one community; and performing merging processing on the edge set within the community and the edge set between communities to obtain a target synthesized graph.
[0007] In one embodiment, the obtaining a directed graph and a plurality of privacy budget values, and performing community partitioning on the nodes in the directed graph to obtain a community partitioning result includes: obtaining the directed graph and a plurality of the privacy budget values; taking each of the nodes as a community; and iterating the partitioning of the communities until an iteration stop condition is reached to obtain the community partitioning result.
[0008] In one embodiment, the obtaining sub-graph node indication data and sub-graph in-degree indication data corresponding to each community in at least one community, determining an in-degree sequence based on the sub-graph in-degree indication data, and processing the in-degree sequence to obtain a target in-degree sequence includes: obtaining the sub-graph node indication data and the sub-graph in-degree indication data corresponding to each community in at least one community, and determining the in-degree sequence based on the sub-graph in-degree indication data; performing noise addition processing on the in-degree sequence to obtain a noise-added in-degree sequence; and performing constraint reasoning on the noise-added in-degree sequence to obtain the target in-degree sequence.
[0009] In one embodiment, the obtaining of the subgraph node indication data and the subgraph out-degree indication data corresponding to each community in at least one of the communities, determining an out-degree sequence based on the subgraph out-degree indication data, and processing the out-degree sequence to obtain a target out-degree sequence includes: obtaining the subgraph node indication data and the subgraph out-degree indication data corresponding to each community in at least one of the communities, and determining the out-degree sequence based on the subgraph out-degree indication data; performing noise addition processing on the out-degree sequence to obtain a noisy out-degree sequence; and performing constrained inference on the noisy out-degree sequence to obtain the target out-degree sequence.
[0010] In one embodiment, the obtaining of the node indication data of the directed graph, calculating a node similarity set according to the node indication data, and performing noise addition processing on the node similarity set to obtain a target node similarity set includes: obtaining the node indication data of the directed graph and calculating the node similarity set according to the node indication data; and performing noise addition processing on the node similarity set to obtain the target node similarity set.
[0011] According to another aspect of the present invention, there is provided a directed graph synthesis system based on differential privacy. The directed graph synthesis system based on differential privacy includes: a community division module, configured to obtain a directed graph and a plurality of privacy budget values, and perform community division on the nodes in the directed graph to obtain a community division result; an in-degree sequence processing module, configured to obtain sub-graph node indication data and sub-graph in-degree indication data corresponding to each community in at least one community, determine an in-degree sequence based on the sub-graph in-degree indication data, and process the in-degree sequence to obtain a target in-degree sequence; an out-degree sequence processing module, configured to obtain sub-graph node indication data and sub-graph out-degree indication data corresponding to each community in at least one community, determine an out-degree sequence based on the sub-graph out-degree indication data, and process the out-degree sequence to obtain a target out-degree sequence; an edge sequence processing module, configured to calculate the number of edges between all communities to obtain an edge sequence, and perform noise addition processing on the edge sequence to obtain a target edge sequence; a node similarity processing module, configured to obtain node indication data of the directed graph, calculate a node similarity set according to the node indication data, and perform noise addition processing on the node similarity set to obtain a target node similarity set; a first edge set generation module, configured to generate an intra-community edge set according to the target in-degree sequence, the target out-degree sequence, and the target node similarity set of each community in at least one community, where the intra-community edge set includes edges within at least one community; a second edge set generation module, configured to generate an inter-community edge set according to the target edge sequence and the target node similarity set of each community in at least one community, where the inter-community edge set includes edges between at least one community; and a synthesis module, configured to perform a merging process on the intra-community edge set and the inter-community edge set to obtain a target synthesized graph.
[0012] In one embodiment, the community division module includes a first acquisition unit, a first division unit, and a second division unit, where: the first acquisition unit is configured to obtain the directed graph and the plurality of privacy budget values; the first division unit is configured to take each of the nodes as a community; and the second division unit is configured to iterate the division of the community until an iteration stop condition is reached to obtain the community division result.
[0013] In one embodiment, the in-degree sequence processing module includes a second acquisition unit, a first noise addition processing unit, and a first constraint reasoning unit, where: the second acquisition unit is configured to acquire the sub-graph node indication data and the sub-graph in-degree indication data corresponding to each community in at least one of the communities, and determine the in-degree sequence based on the sub-graph in-degree indication data; the first noise addition processing unit is configured to perform noise addition processing on the in-degree sequence to obtain a noisy in-degree sequence; the first constraint reasoning unit is configured to perform constraint reasoning on the noisy in-degree sequence to obtain the target in-degree sequence.
[0014] In one embodiment, the out-degree sequence processing module includes a third acquisition unit, a second noise addition processing unit, and a second constraint reasoning unit, where: the third acquisition unit is configured to acquire the sub-graph node indication data and the sub-graph out-degree indication data corresponding to each community in at least one of the communities, and determine the out-degree sequence based on the sub-graph out-degree indication data; the second noise addition processing unit is configured to perform noise addition processing on the out-degree sequence to obtain a noisy out-degree sequence; the second constraint reasoning unit is configured to perform constraint reasoning on the noisy out-degree sequence to obtain the target out-degree sequence.
[0015] In one embodiment, the node similarity processing module includes a fourth acquisition unit and a third noise addition processing unit, where: the fourth acquisition unit is configured to acquire the node indication data of the directed graph and calculate based on the node indication data to obtain the node similarity set; the third noise addition processing unit is configured to perform noise addition processing on the node similarity set to obtain the target node similarity set.
[0016] In summary, in the embodiments of the present invention, the target in-degree sequence is obtained by processing the in-degree sequence, and the target out-degree sequence is obtained by processing the out-degree sequence, and a method for synthesizing directed edges is given, thereby solving the problem of how to process a directed graph to obtain a synthesized social network graph. At the same time, by partitioning the nodes in the directed graph, generating the edge set inside the community according to the target in-degree sequence, the target out-degree sequence, and the target node similarity set of each community in at least one community, and generating the edge set between communities according to the target edge sequence and the target node similarity set of each community in at least one community, the community information of the directed graph is retained, and the usability and accuracy of the synthesized social network graph are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In the following description of the exemplary embodiments in conjunction with the accompanying drawings, more details, features, and advantages of the present invention are disclosed. In the drawings:
[0018] Figure 1Shows a schematic flowchart of a directed graph synthesis method based on differential privacy disclosed in an embodiment of the present application;
[0019] Figure 2 Shows Figure 1 A schematic flowchart of the steps shown in step S110;
[0020] Figure 3 Shows Figure 1 A schematic flowchart of the steps shown in step S120;
[0021] Figure 4 Shows Figure 1 A schematic flowchart of the steps shown in step S130;
[0022] Figure 5 Shows Figure 1 A schematic flowchart of the steps shown in step S150;
[0023] Figure 6 Shows a schematic structural diagram of a directed graph synthesis system based on differential privacy disclosed in an embodiment of the present application. Detailed implementation manners
[0024] Hereinafter, embodiments of the present invention will be described in more detail with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.
[0025] It should be understood that the steps recited in the method embodiments of the present invention can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this regard.
[0026] The term "including" and its variants used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order of functions executed by these devices, modules or units or their interdependent relationships.
[0027] It should be noted that the modifiers "one" and "multiple" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly specified in the context, it should be understood as "one or more".
[0028] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0029] It should be noted that the execution subject of the directed graph synthesis method based on differential privacy provided in the embodiments of the present invention can be one or more electronic devices, and the present invention does not make any limitations in this regard; among them, the electronic device can be a terminal (i.e., a client) or a server. Then, when the execution subject includes multiple electronic devices and at least one terminal and at least one server are included in the multiple electronic devices, the directed graph synthesis method based on differential privacy provided in the embodiments of the present invention can be jointly executed by the terminal and the server. Correspondingly, the terminals mentioned here can include, but are not limited to: smart phones, tablet computers, laptop computers, desktop computers, smart watches, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, and so on. The servers mentioned here can be independent physical servers, or a server cluster or distributed system composed of multiple physical servers, or can also be cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms, and so on.
[0030] Based on the above description, the embodiments of the present invention propose a directed graph synthesis method based on differential privacy. This directed graph synthesis method based on differential privacy can be executed by the above-mentioned electronic devices (terminals or servers); or, this directed graph synthesis method based on differential privacy can be jointly executed by the terminal and the server. For the convenience of description, hereinafter, the case where the electronic device executes this directed graph synthesis method based on differential privacy will be used as an example for illustration.
[0031] Please refer to Figure 1 , which is a schematic flowchart of a directed graph synthesis method based on differential privacy disclosed in the embodiments of the present application. By using the directed graph synthesis method based on differential privacy, the problem of how to process a directed graph to obtain a synthetic social network graph is solved, thereby improving the usability and accuracy of the synthetic social network graph. It should be noted that the directed graph synthesis method based on differential privacy in the embodiments of the present application is not limited to Figure 1The steps and order in the shown flowchart. According to different requirements, the steps in the shown flowchart can be added, removed, or the order can be changed. In the embodiments of the present application, as Figure 1 shown, the process of the directed graph synthesis method based on differential privacy at least includes the following steps.
[0032] S110. Obtain a directed graph and multiple privacy budget values, and perform community partitioning on the nodes in the directed graph to obtain a community partitioning result.
[0033] As Figure 2 shown, in the embodiments of the present invention, Figure 2 the step S110 at least includes the following steps:
[0034] S111. Obtain the directed graph and multiple of the privacy budget values.
[0035] In the embodiments of the present invention, a directed graph and multiple privacy budget values are obtained. Among them, the multiple privacy budget values may include a first privacy budget value ∈ 1 , a second privacy budget value ∈ 2 , a third privacy budget value ∈ 3 , and a fourth privacy budget value ∈ 4 . Specifically, the first privacy budget value ∈ 1 may be the privacy budget used when adding noise to the in-degree sequence, the second privacy budget value ∈ 2 may be the privacy budget used when adding noise to the out-degree sequence, the third privacy budget value ∈ 3 may be the privacy budget used when adding noise to the edge number sequence between communities, and the fourth privacy budget value ∈ 4 may be the privacy budget used when adding noise to the node similarity. The present invention does not make any limitations in this regard.
[0036] S112. Take each of the nodes as a community respectively.
[0037] S113. Iterate the partitioning of the communities until an iteration stop condition is reached to obtain the community partitioning result.
[0038] In the embodiments of the present invention, when iterating the partitioning of the communities, the iteration stop condition may include that the number of iterations reaches a preset number of iterations or the increase amplitude of modularity reaches a preset increase amplitude. Among them, the preset number of iterations and the preset increase amplitude can be set by the user or can also be set by the R & D personnel.
[0039] Modularity, also known as the modularity metric, is a commonly used method to measure the strength of network community structure, first proposed by Mark NewMan. The size of the modularity value mainly depends on the community assignment of nodes in the network, that is, the community division of the network, and can be used to quantitatively measure the quality of network community division. The closer its value is to 1, the stronger the strength of the community structure divided by the network, that is, the better the division quality. Therefore, the optimal network community division can be obtained by maximizing the modularity Q. Taking each node i in the directed graph and each node j followed by node i as an example, the calculation process of the difference in modularity after moving node i into the community corresponding to node j is as follows in the following formula:
[0040]
[0041] D i,in =∑j ∈C A ij Formula (2)
[0042] Among them, ΔQ is the difference in modularity after moving node i into the community corresponding to node j. Taking the community corresponding to node j as community C as an example, D i,in is the number of increased in-degrees when moving node i into community C, A is the adjacency matrix of the graph. When node i follows node j, then A ij =1; when node i does not follow node j, then is the sum of the in-degrees of all nodes in community C, is the sum of the out-degrees of all nodes in community C, and m is the sum of the in-degrees of all nodes in the directed graph.
[0043] Specifically, add node i to the community that can obtain the largest difference in modularity and the difference in modularity is positive. When the iteration stop condition is reached, the community division result is obtained. Taking the community division result as P as an example, P = {C 1 ,C 2 ,…,C n}.
[0044] S120. Obtain the subgraph node indication data and subgraph in-degree indication data corresponding to each community in at least one community, determine the in-degree sequence based on the subgraph in-degree indication data, and process the in-degree sequence to obtain the target in-degree sequence.
[0045] As Figure 3 shown, in the embodiments of the present invention, Figure 3 the step S120 at least includes the following steps:
[0046] S121. Obtain the sub-graph node indication data and the sub-graph in-degree indication data corresponding to each of at least one of the communities, and determine the in-degree sequence based on the sub-graph in-degree indication data.
[0047] In an embodiment of the present invention, obtain, for each community C in at least one community s the corresponding sub-graph node indication data and sub-graph in-degree indication data, and arrange the sub-graph in-degree indication data in non-decreasing order to obtain the in-degree sequence InD s . Wherein, the sub-graph may be a graph composed of nodes in each community in a directed graph and the edges between them, and the sub-graph node indication data includes various information of the sub-graph nodes.
[0048] S122. Perform noise addition processing on the in-degree sequence to obtain a noisy in-degree sequence.
[0049] In an embodiment of the present invention, for each in-degree s in the in-degree sequence InD (which is also the in-degree of node j in community C s ) add independent Laplace noise to obtain the noisy in-degree sequence The distribution of the Laplace noise is generated as described by the following formula:
[0050]
[0051] b 1 = 2 / ∈ 1 Formula (4)
[0052] Wherein, x is the noise value, b 1 is the scale parameter, 2 is the sensitivity, and ∈ 1 is the first privacy budget value. The first privacy budget value may be the privacy budget used when adding noise to the in-degree sequence. Under edge differential privacy, the sensitivity of the in-degree sequence of nodes is always 2.
[0053] S123. Perform constrained inference on the noisy in-degree sequence to obtain the target in-degree sequence.
[0054] In an embodiment of the present invention, taking (which is also the in-degree of node j after adding noise) as an example for illustration, the expression of is as described by the following formula:
[0055]
[0056] Wherein, InD s [j] is the in-degree of node j in community C s , and lap(2 / ∈ 1 ) is the noise addition parameter.
[0057] For the in-degree sequence in, perform constraint reasoning so that the target in-degree sequence is obtained The formula for performing constraint reasoning on the noisy in-degree sequence is as follows:
[0058]
[0059] or
[0060]
[0061] Wherein, is the k-th in-degree obtained after constraint reasoning, InM [i,j] represents to the arithmetic mean of the in-degrees between, InM [i,i] is equal to
[0062] Both formula (6) and formula (7) can be used for the calculation of performing constraint reasoning on the noisy in-degree sequence and the results calculated by both are the same. Taking formula (6) as an example for illustration, its calculation process is as follows:
[0063] ① Traverse i, for each i, calculate the minimum value among each InM [i,j] to InM [i,n] The minimum value is InMin i ;
[0064] ② Find the maximum value among multiple minimum values InMin i which is the calculation result.
[0065] Taking formula (7) as an example for illustration, its calculation process is as follows:
[0066] ① Traverse j, for each j, calculate the maximum value among each InM [1,j] to InM [i,j] The maximum value is InMax j ;
[0067] ② Find the minimum value among multiple maximum values InMax j which is the calculation result.
[0068] S130. Obtain the sub-graph node indication data and sub-graph out-degree indication data corresponding to each community in at least one of the communities, determine the out-degree sequence based on the sub-graph out-degree indication data, and process the out-degree sequence to obtain the target out-degree sequence.
[0069] Such as Figure 4As shown, in the embodiments of the present invention, Figure 4 The step S130 at least includes the following steps:
[0070] S131. Obtain the sub-graph node indication data and the sub-graph out-degree indication data corresponding to each community in at least one of the communities, and determine the out-degree sequence based on the sub-graph out-degree indication data.
[0071] In the embodiments of the present invention, obtain the sub-graph node indication data and the sub-graph out-degree indication data corresponding to each community C s in at least one community, and arrange the sub-graph out-degree indication data in non-decreasing order to obtain the out-degree sequence OutD s .
[0072] S132. Perform noise addition processing on the out-degree sequence to obtain a noisy out-degree sequence.
[0073] In the embodiments of the present invention, for each in-degree s in the in-degree sequence OutD (which is also the out-degree of node j in community C s ) add independent Laplace noise to obtain the noisy out-degree sequence The distribution of the Laplace noise is generated as in the following formula:
[0074]
[0075] b 2 = 2 / ∈ 2 Formula (10)
[0076] where x is the noise value, b 2 is the scale parameter, 2 is the sensitivity, and ∈ 2 is the second privacy budget value. The second privacy budget value can be the privacy budget used when adding noise to the out-degree sequence. Under edge differential privacy, the sensitivity of the out-degree sequence of a node is always 2.
[0077] S133. Perform constrained inference on the noisy out-degree sequence to obtain the target out-degree sequence.
[0078] In the embodiments of the present invention, taking (which is also the out-degree of node j after adding noise) as an example for illustration, the expression of
[0079]
[0080] where OutD S [j] is the out-degree of node j in community C s and lap(2 / ∈2 ) is the noise addition parameter.
[0081] For constrained inference is performed on the out-degree sequence in it so that the target out-degree sequence is obtained The formula for performing constrained inference on the noisy out-degree sequence is as follows:
[0082]
[0083] Or
[0084]
[0085] Wherein, is the k-th out-degree obtained after constrained inference, and OutM [i,j] represents to the arithmetic mean of the out-degrees between them, and the value of OutM [i,i] is equal to
[0086] Both formula (12) and formula (13) can be used for the calculation of performing constrained inference on the noisy out-degree sequence , and the results calculated by both are the same. Taking formula (12) as an example for illustration, its calculation process is as follows:
[0087] ① Traverse i. For each i, calculate the minimum value among each OutM [i,j] to OutM [i,n] , and the minimum value is OutMin i ;
[0088] ② Find the maximum value among multiple minimum values OutMin i as the calculation result.
[0089] Taking formula (13) as an example for illustration, its calculation process is as follows:
[0090] ① Traverse j. For each j, calculate the maximum value among each OutM [1,j] to OutM [i,j] , and the maximum value is OutMax j ;
[0091] ② Find the minimum value among multiple maximum values OutMax j as the calculation result.
[0092] S140. Calculate the edge sequence by calculating the number of edges between all communities, and perform noise addition processing on the edge sequence to obtain the target edge sequence.
[0093] In the embodiment of the present invention, taking community Ci and community C j Take community C as an example for illustration. Community C i and community C j The number of outgoing edges Edge [i][j] (that is, how many edges there are from the nodes in community C i to community C in the directed graph j ), all Edge [i][j] forms an edge sequence Edge. Community C i and community C j The number of outgoing edges Edge [i][[j] The calculation process is as follows:
[0094]
[0095] Add Laplace noise to each item in the edge sequence Edge to obtain the target edge sequence The distribution of Laplace noise is as follows:
[0096]
[0097] b 3 = 2 / ∈ 3 Formula (17)
[0098] where x is the noise value, b 3 is the scale parameter, 2 is the sensitivity, and ∈ 3 is the third privacy budget value. The third privacy budget value can be the privacy budget used when adding noise to the edge number sequence between communities. Under edge differential privacy, the sensitivity of the edge sequence is always 2.
[0099] Take as an example for illustration. The expression of
[0100]
[0101] where Edge [i][j] is how many edges there are from the nodes in community C i to community C j in the directed graph, and lap(2 / ∈ 3 ) is the noise addition parameter.
[0102] S150. Obtain the node indication data of the directed graph, calculate the node similarity set according to the node indication data, and perform noise addition processing on the node similarity set to obtain the target node similarity set.
[0103] As Figure 5 shown, in the embodiment of the present invention, Figure 5The step S150 at least includes the following steps:
[0104] S151. Obtain the node indication data of the directed graph, and perform calculations according to the node indication data to obtain the node similarity set.
[0105] In the embodiment of the present invention, since the similarity between some nodes in the directed graph is almost 0 and cannot be calculated by a formula, the similarity between all nodes in the directed graph is initialized and set to 0. Taking the similarity between node i and node j as an example, the similarity calculation between node i and node j is as follows:
[0106]
[0107]
[0108] Among them, DE(i,j) is the relationship of the direct edge between nodes, is the degree adjustment, IP(i,i) is the relationship of the indirect edge between nodes (that is, the path with a length of 2), CC(i,j) is the relationship that nodes i and j are jointly concerned by other node x, CF(i,j) is the relationship that nodes i and j jointly concern other node x, Ind j is the in-degree of node j (that is, the in-degree of node j in the directed graph), Outd i is the out-degree of node i (that is, the out-degree of node i in the directed graph), OutN(i) is the set formed by the nodes pointed to by node i, InN(j) is the set formed by the nodes pointing to node j, Outd x is the out-degree of node x (that is, the out-degree of node x in the directed graph), Ind i is the in-degree of node i (that is, the in-degree of node i in the directed graph), InN(i) represents the set of nodes that concern node i, OutN(j) represents the set of nodes concerned by node j, and e is the natural constant.
[0109] S152. Perform noise addition processing on the node similarity set to obtain the target node similarity set.
[0110] In the embodiment of the present invention, Laplace noise is added to each node similarity in the node similarity set to obtain the target node similarity set. Taking the similarity between node i and node j as an example, the node similarity S [i][j] , after performing noise addition processing on S [i][j] the target node similarity is obtained The distribution of Laplace noise is as follows:
[0111]
[0112] b 4 = 2 / ∈ 4 Formula (26)
[0113] where x is the noise value, b 4 is the scale parameter, 2 is the sensitivity, ∈ 4 is the fourth privacy budget value, and the fourth privacy budget value can be the privacy budget used when adding noise to the node similarity. Under edge differential privacy, the sensitivity of the similarity between nodes is always 2.
[0114] Take as an example for illustration. The expression of
[0115]
[0116] In this embodiment, the execution order of the above steps can be to first execute steps S120 / S130 / S140 / S150 in parallel and then execute steps S160 / S170, or can be executed in the order of step S120, S130, S140, S150, and then execute steps S160 / S170 or step S120, S140, S130, S150, and then execute steps S160 / S170 or in any order and then execute steps S160 / S170. The present invention does not limit this.
[0117] S160. Generate an edge set inside the community according to the target in-degree sequence, the target out-degree sequence, and the target node similarity set of each community in at least one of the communities. The edge set inside the community includes at least one edge inside the community.
[0118] In the embodiment of the present invention, for each community C in at least one community s the corresponding subgraph G s =(V s , E s ) is initialized, where V s is the node in community C s , and E s is initially an empty set. For each node i in V s , according to the target node similarity set, select the node whose similarity with node i is not 0 and is in community C s , and add this node to the candidate list can(i) of node i. The expression of can(i) is as follows in the following formula:
[0119]
[0120] where both i and j are nodes, is the target node similarity between nodes i and j, C s is a community.
[0121] For each node in the candidate list can(i) of node i, an edge is generated with probability P(e ij |i), and the expression of P(e ij |i) is as follows:
[0122]
[0123] where M is the sum of the in-degrees of all nodes in community C s in, is the target out-degree of node i, is the target in-degree of node j, can(i) is the candidate list of node i, is the target node similarity between nodes i and j, and Attn(i,j) is the distribution coefficient that determines the probability of generating an edge between two nodes based on the similarity between the two nodes, the out-degree of the source node, and the in-degree of the target node.
[0124] S170. Generate an edge set between communities according to the target edge sequence and the target node similarity set of each community in at least one of the communities, and the edge set between communities includes edges between at least one community.
[0125] In the embodiment of the present invention, any two communities are selected. Taking community C i and community C j as an example for illustration, the target edge sequence i and community C j of community C V i is a node in community C i , V j is a node in community C j , and node pairs (u,v) that satisfy the following conditions are found and added to the set pairs. The conditions are as follows:
[0126]
[0127] where, is the target node similarity between node u and node v, that is, the node pair (u,v) satisfies that the target node similarity from node u to node v is not 0, and node u belongs to community C i , and node v belongs to community C j .
[0128] For all node pairs in the set paris, an edge is generated with probability P(e uv |(u,v)∈paris), P(e uvThe expression |(u, v) ∈ paris) is as follows:
[0129]
[0130] Wherein, is the target edge sequence of community C i and community C j Attn(u, v) is the distribution coefficient that determines the probability of generating an edge between two nodes according to the similarity between the two nodes, the out-degree of the source node, and the in-degree of the target node. is the target out-degree of node u, is the target in-degree of node v, is the target node similarity between nodes u and v.
[0131] In this embodiment, the execution order of the above steps may be to first execute steps S160 / S170 in parallel, and then execute step S180, or may be executed in the order of steps S160, S170, S180 or steps S170, S160, S180 or any order. The present invention does not limit this.
[0132] S180. Merge the edge set within the community and the edge set between the communities to obtain a target composite graph.
[0133] In the embodiments of the present invention, the edge set within the community and the edge set between the communities are merged to obtain a target composite graph G res =(V, E res ), where V is the node set of the directed graph, and E res is the merged edge set.
[0134] In the embodiments of the present invention, the data are all from legal sources, and the data may include a directed graph, a privacy budget value, subgraph node indication data, subgraph in-degree indication data, etc.
[0135] In summary, in the directed graph synthesis method based on differential privacy of the present application, the target in-degree sequence is obtained by processing the in-degree sequence, and the target out-degree sequence is obtained by processing the out-degree sequence, and a method for synthesizing directed edges is given, thus solving the problem of how to process a directed graph to obtain a synthesized social network graph. At the same time, by partitioning the nodes in the directed graph, generating the edge set within the community according to the target in-degree sequence, target out-degree sequence, and target node similarity set of each community in at least one community, and generating the edge set between the communities according to the target edge sequence and target node similarity set of each community in at least one community, the community information of the directed graph is retained, and the usability and accuracy of the synthesized social network graph are improved.
[0136] Please refer toFigure 6 , which is a schematic structural diagram of a directed graph synthesis system based on differential privacy disclosed in an embodiment of the present application. In one embodiment, as Figure 6 shown, the present application provides a directed graph synthesis system 100 based on differential privacy. The directed graph synthesis system 100 based on differential privacy may at least include: a community division module 110, an in-degree sequence processing module 120, an out-degree sequence processing module 130, an edge sequence processing module 150, a node similarity processing module 160, a first edge set generation module 170, a second edge set generation module 180, and a synthesis module 190. Among them, there is information interaction between the community division module 110 and the in-degree sequence processing module 120, the out-degree sequence processing module 130, the edge sequence processing module 150, and the node similarity processing module 160. There is information interaction between the first edge set generation module 170 and the in-degree sequence processing module 120, the out-degree sequence processing module 130, the node similarity processing module 160, and the synthesis module 190. There is information interaction between the second edge set generation module 180 and the edge sequence processing module 150, the node similarity processing module 160, and the synthesis module 190.
[0137] The community division module 110 is used to obtain a directed graph and a plurality of privacy budget values, and perform community division on the nodes in the directed graph to obtain a community division result. Among them, the community division module 110 may at least include: a first acquisition unit 111, a first division unit 113, and a second division unit 116.
[0138] The first acquisition unit 111 is used to obtain the directed graph and the plurality of privacy budget values. Specifically, obtain a directed graph and a plurality of privacy budget values. Among them, the plurality of privacy budget values may include a first privacy budget value ∈ 1 , a second privacy budget value ∈ 2 , a third privacy budget value ∈ 3 and a fourth privacy budget value ∈ 4 . Specifically, the first privacy budget value ∈ 1 may be the privacy budget used for adding noise to the in-degree sequence. The second privacy budget value ∈ 2 may be the privacy budget used for adding noise to the out-degree sequence. The third privacy budget value ∈ 3 may be the privacy budget used for adding noise to the edge number sequence between communities. The fourth privacy budget value ∈ 4 may be the privacy budget used for adding noise to the node similarity. The present invention does not limit this.
[0139] The first division unit 113 is used to take each of the nodes as a community.
[0140] The second partitioning unit 116 is used to iterate the partitioning of the community until an iteration stop condition is reached, and the community partitioning result is obtained. Specifically, the partitioning of the community is iterated, and the iteration stop condition may include that the number of iterations reaches a preset number of iterations or the increase amplitude of modularity reaches a preset increase amplitude. Among them, modularity is also called the modularity metric value, which is a commonly used method to measure the strength of the network community structure at present and was first proposed by Mark Newman. The size of the modularity value mainly depends on the community assignment C of the nodes in the network, that is, the community partitioning situation of the network, and can be used to quantitatively measure the quality of the network community partitioning. The closer its value is to 1, the stronger the strength of the community structure divided by the network, that is, the better the partitioning quality. Therefore, the optimal network community partitioning can be obtained by maximizing the modularity Q. Taking each node i in the directed graph and each node j followed by node i as an example, the calculation process of the difference in modularity after moving node i into the community corresponding to node j is as follows in the following formula:
[0141]
[0142] D i,in =∑j ∈C A ij Formula (2)
[0143] Among them, ΔQ is the difference in modularity after moving node i into the community corresponding to node j. Taking the community corresponding to node j as community C as an example, D i,in is the number of increased in-degrees when moving node i into community C, A is the adjacency matrix of the graph. When node i follows node j, then A ij =1; when node i does not follow node j, then A ij =0, is the sum of the in-degrees of all nodes in community C, is the sum of the out-degrees of all nodes in community C, and m is the sum of the in-degrees of all nodes in the directed graph.
[0144] Specifically, node i is added to the community that can obtain the largest difference in modularity and the difference in modularity is a positive number. When the iteration stop condition is reached, the community partitioning result is obtained. Taking the community partitioning result as P as an example, P = {C 1 , C 2 , …, C n}, where C i is community i.
[0145] The in-degree sequence processing module 120 is used to obtain the subgraph node indication data and subgraph in-degree indication data corresponding to each community in at least one community, determine the in-degree sequence based on the subgraph in-degree indication data, and process the in-degree sequence to obtain the target in-degree sequence. Among them, the in-degree sequence processing module 120 can at least include: a second acquisition unit 121, a first noise addition processing unit 123, and a first constraint reasoning unit 125.
[0146] The second acquisition unit 121 is used to obtain the subgraph node indication data and subgraph in-degree indication data corresponding to each community in at least one community, and determine the in-degree sequence based on the subgraph in-degree indication data. Specifically, it obtains the subgraph node indication data and subgraph in-degree indication data corresponding to each community C s in at least one community, and arranges the subgraph in-degree indication data in non-decreasing order to obtain the in-degree sequence InD s . Among them, the subgraph can be a graph composed of the nodes in each community in the directed graph and the edges between them, and the subgraph node indication data includes various information of the subgraph nodes.
[0147] The first noise addition processing unit 123 is used to perform noise addition processing on the in-degree sequence to obtain a noisy in-degree sequence. Specifically, for each in-degree s in the in-degree sequence InD (that is, the in-degree of node j in community C s ), independent Laplace noise is added to obtain the noisy in-degree sequence The distribution of the Laplace noise is generated as follows in the formula:
[0148]
[0149] b 1 =2 / ∈ 1 Formula (4)
[0150] Among them, x is the noise value, b 1 is the scale parameter, 2 is the sensitivity, ∈ 1 is the first privacy budget value. The first privacy budget value can be the privacy budget used when adding noise to the in-degree sequence. Under edge differential privacy, the sensitivity of the in-degree sequence of nodes is always 2.
[0151] The first constraint reasoning unit 125 is used to perform constraint reasoning on the noisy in-degree sequence to obtain the target in-degree sequence. Specifically, taking (that is, the in-degree of node j after adding noise) as an example for illustration, The expression of is as follows in the formula:
[0152]
[0153] Among them, InD s [j] is community C s The in-degree of node j, lap(2 / ∈ 1 ) is the noise adding parameter.
[0154] right Constraint reasoning is performed on the in-degree sequence in so that Get the target in-degree sequence The formula for constrained reasoning on the noisy in-degree sequence is as follows:
[0155]
[0156] or
[0157]
[0158] in, is the kth in-degree obtained after constraint reasoning, InM [i,j] Represents arrive The arithmetic mean of the in-degree, InM [i,i] The value is equal to
[0159] Formula (6) and formula (7) can be used to calculate the noise-added degree sequence The constraint reasoning calculation is performed, and the results obtained by the two calculations are the same. Taking formula (6) as an example, the calculation process is as follows:
[0160] ①Traverse i, for each i, calculate each InM [i,j] To InM [i,n] The minimum value in the i ;
[0161] ②Find multiple minimum values InMin i The maximum value among is the calculation result.
[0162] Taking formula (7) as an example, the calculation process is as follows:
[0163] ① Traverse j, for each j, calculate each InM [1,j] To InM [i,j] The maximum value in is InMax j ;
[0164] ②Find multiple maximum values InMax j The minimum value among is the calculation result.
[0165] The out-degree sequence processing module 130 is used to obtain the subgraph node indication data and the subgraph out-degree indication data corresponding to each community in at least one community, determine the out-degree sequence based on the subgraph out-degree indication data, and process the out-degree sequence to obtain the target out-degree sequence. Among them, the out-degree sequence processing module 130 may at least include: a third acquisition unit 131, a second noise addition processing unit 133, and a second constraint reasoning unit 135.
[0166] The third acquisition unit 131 is used to obtain the subgraph node indication data and the subgraph out-degree indication data corresponding to each community in at least one community, and determine the out-degree sequence based on the subgraph out-degree indication data. Specifically, it obtains the subgraph node indication data and the subgraph out-degree indication data corresponding to each community C s in at least one community, and arranges the subgraph out-degree indication data in non-decreasing order to obtain the out-degree sequence OutD s .
[0167] The second noise addition processing unit 133 is used to perform noise addition processing on the out-degree sequence to obtain a noisy out-degree sequence. Specifically, for each in-degree s in the in-degree sequence OutD (that is, the out-degree of node j in community C s ), independent Laplace noise is added to obtain the noisy out-degree sequence The distribution of the Laplace noise is generated as shown in the following formula:
[0168]
[0169] b 2 = 2 / ∈ 2 Formula (10)
[0170] where x is the noise value, b 2 is the scale parameter, 2 is the sensitivity, and ∈ 2 is the second privacy budget value. The second privacy budget value can be the privacy budget used when adding noise to the out-degree sequence. Under edge differential privacy, the sensitivity of the out-degree sequence of a node is always 2.
[0171] The second constraint reasoning unit 135 is used to perform constraint reasoning on the noisy out-degree sequence to obtain the target out-degree sequence. Specifically, taking (that is, the out-degree of node j after adding noise) as an example for illustration, the expression of
[0172]
[0173] is as shown in the following formula: S [j] is for community C sThe out-degree of the middle node j, lap(2 / ∈ 2 ) is the noise addition parameter.
[0174] For the out-degree sequence in, perform constraint reasoning so that the target out-degree sequence is obtained The formula for performing constraint reasoning on the noisy out-degree sequence is as follows:
[0175]
[0176] Or
[0177]
[0178] Among them, is the k-th out-degree obtained after constraint reasoning, OutM [i,j] represents to the arithmetic mean of the out-degrees between, OutM [i,i] is equal to
[0179] Both formula (12) and formula (13) can be used for the calculation of performing constraint reasoning on the noisy out-degree sequence , and the results calculated by both are the same. Taking formula (12) as an example for illustration, its calculation process is as follows:
[0180] ① Traverse i, for each i, calculate the minimum value in each OutM [i,j] to OutM [i,n] , and the minimum value is OutMin i ;
[0181] ② Find the maximum value among multiple minimum values OutMin i as the calculation result.
[0182] Taking formula (13) as an example for illustration, its calculation process is as follows:
[0183] ① Traverse j, for each j, calculate the maximum value in each OutM [1,j] to OutM [i,j] , and the maximum value is OutMax j ;
[0184] ② Find the minimum value among multiple maximum values OutMax j as the calculation result.
[0185] The edge sequence processing module 150 is used to calculate the edge sequence for the number of edges between all communities and perform noise addition processing on the edge sequence to obtain the target edge sequence. Specifically, taking community C iand community C j as an example for illustration, community C i and community C j The number of out-edges Edge [i][j] (that is, how many edges there are from the nodes in community C i to community C j ) in the directed graph. All Edge [i][j] form an edge sequence Edge. The calculation process of the number of out-edges Edgr i of community C j and community C [i][j] is as follows in the following formula:
[0186]
[0187] Add Laplace noise to each item in the edge sequence Edge to obtain the target edge sequence The distribution of the Laplace noise is as follows in the following formula:
[0188]
[0189] b 3 = 2 / ∈ 3 Formula (17)
[0190] where x is the noise value, b 3 is the scale parameter, 2 is the sensitivity, and ∈ 3 is the third privacy budget value. The third privacy budget value can be the privacy budget used when adding noise to the edge number sequence between communities. In edge differential privacy, the sensitivity of the edge sequence is always 2.
[0191] Take as an example for illustration, The expression of
[0192]
[0193] where Edge [i][j] is how many edges there are from the nodes in community C i to community C j , and lap(2 / ∈ 3 ) is the noise addition parameter.
[0194] The node similarity processing module 160 is used to obtain the node indication data of the directed graph, calculate a node similarity set according to the node indication data, and perform noise addition processing on the node similarity set to obtain a target node similarity set. Among them, the node similarity processing module 160 can at least include: a fourth acquisition unit 161 and a third noise addition processing unit 163.
[0195] The fourth acquisition unit 161 is configured to acquire the node indication data of the directed graph, and perform calculations according to the node indication data to obtain a node similarity set. Specifically, since the similarity between some nodes in the directed graph is almost 0 and cannot be calculated using a formula, the similarities between all nodes in the directed graph are initialized and all set to 0. Taking the similarity between node i and node j as an example, the similarity calculation between node i and node j is as follows:
[0196]
[0197] where DE(i,j) is the relationship of the direct edge between nodes, is the degree adjustment, IP(i,j) is the relationship of the indirect edge (i.e., the path of length 2) between nodes, CC(i,j) is the relationship that nodes i and j are jointly concerned by other node x, CF(i,j) is the relationship that nodes i and j jointly concern other node x, Ind j is the in-degree of node j (i.e., the in-degree of node j in the directed graph), Outd i is the out-degree of node i (i.e., the out-degree of node i in the directed graph), OutN(i) is the set formed by the nodes pointed to by node i, InN(j) is the set formed by the nodes pointing to node j, Outd x is the out-degree of node x (i.e., the out-degree of node x in the directed graph), Ind i is the in-degree of node i (i.e., the in-degree of node i in the directed graph), InN(i) represents the set of nodes that concern node i, OutN(j) represents the set of nodes concerned by node j, and e is the natural constant.
[0198] The third noise addition processing unit 163 is configured to perform noise addition processing on the node similarity set to obtain a target node similarity set. Specifically, Laplace noise is added to each node similarity in the node similarity set to obtain the target node similarity set. Taking the similarity between node i and node j as an example, the node similarity S [ i ][ j ] , after performing noise addition processing on S [ i ] [ j ] the target node similarity is obtained. The distribution of the Laplace noise is as follows: The distribution of the Laplace noise is as follows:
[0199]
[0200] b 4 = 2 / ∈ 4Formula (26)
[0201] where x is the noise value, b 4 is the scale parameter, 2 is the sensitivity, ∈ 4 is the fourth privacy budget value. The fourth privacy budget value can be the privacy budget used when adding noise to the node similarity. Under edge differential privacy, the sensitivity of the similarity between nodes is always 2.
[0202] Take as an example for illustration. The expression of
[0203]
[0204] The first edge set generation module 170 is used to generate the edge set inside the community according to the target in-degree sequence, the target out-degree sequence, and the target node similarity set of each community in at least one community. The edge set inside the community includes edges inside at least one community. Specifically, for each community C in at least one community s The corresponding subgraph G s = (V s , E s ) is initialized, where V s is the node in community C s , and E s is initially an empty set. For each node i in V s , according to the target node similarity set, select the node whose similarity with node i is not 0 and is in community C s , and add this node to the candidate list can(i) of node i. The expression of can(i) is as follows:
[0205]
[0206] where both i and j are nodes, is the target node similarity between nodes i and j, and C s is the community.
[0207] For each node in the candidate list can(i) of node i, generate an edge with probability P(e ij | i). The expression of P(e ij | i) is as follows:
[0208]
[0209] where M is the sum of the in-degrees of all nodes in community C s , is the target out-degree of node i, is the target in-degree of node j, and can(i) is the candidate list of node i. is the target node similarity between nodes i and j, and Attn(i, j) is the allocation coefficient that determines the probability of generating an edge between two nodes according to the similarity between the two nodes, the out-degree of the source node, and the in-degree of the target node.
[0210] The second edge set generation module 180 is used to generate an edge set between communities according to the target edge sequence of each community in at least one community and the set of target node similarities. The edge set between communities includes edges between at least one community. Specifically, any two communities are selected. Taking community C i and community C j as an example for illustration, the target edge sequence i of community C j is V i is a node in community C i , V j is a node in community C j . Node pairs (u, v) that satisfy the following conditions are found and added to the set pairs. The conditions are as follows in the following formula:
[0211]
[0212] where, is the target node similarity between node u and node v, that is, the node pair (u, v) satisfies that the target node similarity from node u to node v is not 0, and node u belongs to community C i , and node v belongs to community C j .
[0213] For all node pairs in the set paris, an edge is generated with probability P(e uv |(u, v) ∈ paris). The expression of P(e uv |(u, v) ∈ paris) is as follows in the following formula:
[0214]
[0215] where, is the target edge sequence of community C i and community C j , Attn(u, v) is the allocation coefficient that determines the probability of generating an edge between two nodes according to the similarity between the two nodes, the out-degree of the source node, and the in-degree of the target node, is the target out-degree of node u, is the target in-degree of node v, is the set of target node similarities.
[0216] The synthesis module 190 is configured to perform a merging process on the edge set within the community and the edge set between the communities to obtain a target synthesized graph. Specifically, the edge set within the community and the edge set between the communities are merged to obtain a target synthesized graph G res =(V, E res ), where V is the node set of the directed graph and E res is the merged edge set.
[0217] In summary, in the directed graph synthesis system based on differential privacy of the present application, the in-degree sequence processing module 120 processes the in-degree sequence to obtain a target in-degree sequence, and the out-degree sequence processing module 130 processes the out-degree sequence to obtain a target out-degree sequence, providing a method for synthesizing directed edges, thereby solving the problem of how to process a directed graph to obtain a synthesized social network graph. At the same time, the community division module 110 divides the nodes in the directed graph into communities, the first edge set generation module 170 generates an edge set within the community according to the target in-degree sequence, the target out-degree sequence, and the target node similarity set of each community in at least one community, and the second edge set generation module 180 generates an edge set between the communities according to the target edge sequence and the target node similarity set of each community in at least one community, retaining the community information of the directed graph and improving the usability and accuracy of the synthesized social network graph.
[0218] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", "one implementation manner", "one preferred implementation manner" or "some examples", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0219] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.
Claims
1. A directed graph synthesis method based on differential privacy, characterized in that: The directed graph synthesis method based on differential privacy includes: Obtaining a directed graph and a plurality of privacy budget values, and performing community division on the nodes in the directed graph to obtain a community division result; Obtaining subgraph node indication data and subgraph in-degree indication data corresponding to each community in at least one community, determining an in-degree sequence based on the subgraph in-degree indication data, and processing the in-degree sequence to obtain a target in-degree sequence; Acquire the subgraph node indication data and subgraph out-degree indication data corresponding to each community in at least one of the communities, determine an out-degree sequence based on the subgraph out-degree indication data, and process the out-degree sequence to obtain a target out-degree sequence; The number of edges between all communities is calculated to obtain an edge sequence, and the edge sequence is subjected to noise processing to obtain a target edge sequence; Obtaining node indication data of the directed graph, calculating according to the node indication data to obtain a node similarity set, and performing noise processing on the node similarity set to obtain a target node similarity set; Generate an edge set within the community according to the target in-degree sequence, the target out-degree sequence and the target node similarity set of each community in at least one of the communities, wherein the edge set within the community includes at least one edge within the community; Generate an edge set between communities according to the target edge sequence of each community in at least one of the communities and the target node similarity set, wherein the edge set between communities includes at least one edge between communities; The edge set within the community and the edge set between the communities are merged to obtain a target composite graph.
2. The directed graph synthesis method based on differential privacy according to claim 1, characterized in that: The obtaining of a directed graph and a plurality of privacy budget values, and performing community division on nodes in the directed graph to obtain a community division result, includes: Obtaining the directed graph and a plurality of the privacy budget values; Each of the nodes is regarded as a community; The community division is iterated until an iteration stop condition is reached to obtain the community division result.
3. The directed graph synthesis method based on differential privacy according to claim 2, characterized in that: The obtaining of subgraph node indication data and subgraph in-degree indication data corresponding to each community in at least one community, determining an in-degree sequence based on the subgraph in-degree indication data, and processing the in-degree sequence to obtain a target in-degree sequence includes: Acquire the subgraph node indication data and the subgraph in-degree indication data corresponding to each community in at least one of the communities, and determine the in-degree sequence based on the subgraph in-degree indication data; Performing noise processing on the in-degree sequence to obtain a noisy in-degree sequence; Constraint reasoning is performed on the noisy in-degree sequence to obtain the target in-degree sequence.
4. The directed graph synthesis method based on differential privacy according to claim 3, characterized in that: The obtaining of the subgraph node indication data and the subgraph out-degree indication data corresponding to each community in at least one of the communities, determining an out-degree sequence based on the subgraph out-degree indication data, and processing the out-degree sequence to obtain a target out-degree sequence includes: Acquire the subgraph node indication data and the subgraph out-degree indication data corresponding to each community in at least one of the communities, and determine the out-degree sequence based on the subgraph out-degree indication data; Performing noise processing on the out-degree sequence to obtain a noisy out-degree sequence; Constraint reasoning is performed on the noisy out-degree sequence to obtain the target out-degree sequence.
5. The directed graph synthesis method based on differential privacy according to claim 4, characterized in that: The acquiring node indication data of the directed graph, calculating according to the node indication data to obtain a node similarity set, and performing noise processing on the node similarity set to obtain a target node similarity set includes: Obtaining node indication data of the directed graph, and performing calculations based on the node indication data to obtain the node similarity set; The node similarity set is subjected to noise addition processing to obtain the target node similarity set.
6. A directed graph synthesis system based on differential privacy, characterized in that: The directed graph synthesis system based on differential privacy includes: A community division module, used to obtain a directed graph and multiple privacy budget values, and perform community division on the nodes in the directed graph to obtain a community division result; an in-degree sequence processing module, configured to obtain subgraph node indication data and subgraph in-degree indication data corresponding to each community in at least one community, determine an in-degree sequence based on the subgraph in-degree indication data, and process the in-degree sequence to obtain a target in-degree sequence; an out-degree sequence processing module, configured to obtain subgraph node indication data and subgraph out-degree indication data corresponding to each community in at least one community, determine an out-degree sequence based on the subgraph out-degree indication data, and process the out-degree sequence to obtain a target out-degree sequence; An edge sequence processing module is used to calculate the number of edges between all communities to obtain an edge sequence, and to perform noise processing on the edge sequence to obtain a target edge sequence; A node similarity processing module is used to obtain node indication data of the directed graph, calculate a node similarity set according to the node indication data, and perform noise processing on the node similarity set to obtain a target node similarity set; A first edge collection generation module, configured to generate an edge set within a community according to the target in-degree sequence, the target out-degree sequence, and the target node similarity set of each community in at least one community, wherein the edge set within the community includes at least one edge within the community; A second edge set generation module is used to generate an edge set between communities according to the target edge sequence of each community in at least one community and the target node similarity set, wherein the edge set between communities includes at least one edge between communities; The synthesis module is used to merge the edge set within the community and the edge set between the communities to obtain a target synthetic graph.
7. The directed graph synthesis system based on differential privacy according to claim 6, characterized in that: The community division module includes a first acquisition unit, a first division unit and a second division unit, wherein: The first acquisition unit is used to acquire the directed graph and the plurality of privacy budget values; The first division unit is used to respectively regard each of the nodes as a community; The second division unit is used to iterate the division of the community until an iteration stop condition is reached to obtain the community division result.
8. The directed graph synthesis system based on differential privacy according to claim 7, characterized in that: The in-degree sequence processing module includes a second acquisition unit, a first noise processing unit and a first constraint reasoning unit, wherein: The second acquisition unit is used to acquire the subgraph node indication data and the subgraph in-degree indication data corresponding to each community in at least one of the communities, and determine the in-degree sequence based on the subgraph in-degree indication data; The first noise processing unit is used to perform noise processing on the in-degree sequence to obtain a noisy in-degree sequence; The first constraint reasoning unit is used to perform constraint reasoning on the noisy in-degree sequence to obtain the target in-degree sequence.
9. The directed graph synthesis system based on differential privacy according to claim 7, characterized in that: The out-degree sequence processing module includes a third acquisition unit, a second noise processing unit and a second constraint reasoning unit, wherein: The third acquisition unit is used to acquire the subgraph node indication data and the subgraph out-degree indication data corresponding to each community in at least one of the communities, and determine the out-degree sequence based on the subgraph out-degree indication data; The second noise processing unit is used to perform noise processing on the out-degree sequence to obtain a noisy out-degree sequence; The second constraint reasoning unit is used to perform constraint reasoning on the noisy out-degree sequence to obtain the target out-degree sequence.
10. The directed graph synthesis system based on differential privacy according to claim 7, characterized in that: The node similarity processing module includes a fourth acquisition unit and a third noise processing unit, wherein: The fourth acquisition unit is used to acquire node indication data of the directed graph, and perform calculation according to the node indication data to obtain the node similarity set; The third noise processing unit is used to perform noise processing on the node similarity set to obtain the target node similarity set.