Community detection microscale confrontation method and system based on node relative importance disturbance

By considering the relative importance of nodes in community detection, screening and adding false links, the problems of user node status and information loss in the confrontation process in the prior art are solved, and effective privacy protection and community detection confrontation effects are achieved.

CN120162679AActive Publication Date: 2025-06-17Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510318512.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-17
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

The existing community detection adversarial methods fail to fully consider the importance of nodes, resulting in the reduction of user node status during adversarial process, and the cost of deleting user connections and missing information.

Method used

A micro-scale adversarial method for community detection based on the perturbation of relative importance of nodes is adopted. By obtaining the community topology of the target node in the social network, adding fake user nodes and establishing a database of relative importance candidate links, eliminating inferior genes, retaining preparatory genes, and using genetic algorithms to screen the optimal false links and adding them to the community topology structure.

Benefits of technology

It realizes reducing the relative importance of target nodes at a minimum cost, changing the results of community division, protecting user privacy information, and performing better than traditional methods on multiple real network data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162679A_ABST
    Figure CN120162679A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of network data processing, in particular to a community detection micro-scale confrontation method and a community detection micro-scale confrontation system based on node relative importance disturbance. Establishing a false node relative importance candidate link library according to the link relationship between the false user nodes and the original nodes in the community; dividing nodes in the community into poor-quality nodes and preparatory nodes, and obtaining poor-quality genes and preparatory genes; inferior genes are removed through pruning operation, preparatory genes are reserved, node relative importance serves as fitness, and excellent genes are screened through a roulette selection method; a false user node is taken as a core, an excellent gene chromosome is obtained by utilizing cross selection, and an optimal relative importance false link generated by the chromosome is added to a community topological structure to which a target node belongs. The method can achieve a confrontation effect on overlapping community detection and non-overlapping community detection, and can better achieve the purpose of hiding target nodes or communities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network data processing, and particularly to a community detection micro-scale adversarial method and system based on node relative importance perturbation. Background Art

[0002] As community detection algorithms are gradually and maturely applied to various fields, people's information is increasingly exposed to the public. Some institutions or individuals may use community detection algorithms to steal others' privacy information. For example, a core figure in an enterprise with occupational secrets that are not convenient to disclose is accurately located by the community detection method. To prevent such problems from occurring, adversarial community detection methods have gradually become the focus of people's protection of privacy information.

[0003] Adversarial community detection is to change the community structure through a certain means or method to achieve the effect of confusing the community detection result, so as to protect the community from being correctly detected and solve the problem of privacy leakage. Currently, the popular adversarial community detection methods mainly focus on community hiding or individual hiding. Most studies mainly disrupt the community structure by adjusting the connections between nodes, deleting the original links and re-adding new links, so as to hide individuals or even communities. However, this method ignores the actual problems in real life. First, most methods only target the relationship between specific nodes and their surrounding nodes during the adversarial process, without fully utilizing the characteristics associated with the nodes and their communities, resulting in the effectiveness of the adversarial effect to be verified; second, it is difficult to delete the connections between actual users, and a large cost needs to be paid to achieve this. Methods that do not consider the cost issue but only consider the concealment or adversarial effect cannot comprehensively show the advantages and disadvantages of the strategy method; third, the purpose of the confrontation is to protect privacy, but deleting the connections between users will cause the loss of originally important information, violating the original intention of protecting user privacy. Summary of the Invention

[0004] Therefore, the present invention provides a community detection micro-scale adversarial method and system based on node relative importance perturbation to solve the problems existing in the existing community detection confrontation, such as the reduction of the status of user nodes during the adversarial process due to the lack of consideration of node importance, and the high cost / information loss caused by deleting user connections.

[0005] According to the design scheme provided by the present invention, on the one hand, a community detection micro-scale adversarial method based on node relative importance perturbation is provided, including:

[0006] Obtain the topological structure of the community to which the target node belongs in the social network, add fake user nodes to the topological structure of the community to which the target node belongs, and establish a candidate link library for the relative importance of fake nodes based on the link relationship between the fake user nodes and the original nodes in the community. The candidate link library for relative importance is used to store the links that are expected to exist between the fake user nodes and all nodes in the community according to the relative importance of the nodes. The relative importance of the nodes represents the relative importance degree of the nodes in the community based on the magnitude of the node importance;

[0007] Divide the nodes in the community into inferior nodes and reserve nodes according to the magnitude of the importance of all nodes in the community, and divide the links in the candidate link library for relative importance into inferior genes and reserve genes. The inferior nodes are used to represent the nodes whose magnitude of node importance is greater than the threshold, the reserve nodes are used to represent the nodes whose node importance is less than the threshold, the inferior genes are used to represent the links that are expected to exist between the fake user nodes and the inferior nodes, and the reserve genes are used to represent the links that are expected to exist between the fake user nodes and the reserve nodes;

[0008] Eliminate the inferior genes in the candidate link library for relative importance through pruning operations, retain the reserve genes, and form an importance gene library. Take the relative importance of the nodes as the fitness, and screen out excellent genes from the importance gene library based on the fitness of the user nodes and through the roulette wheel selection method;

[0009] Taking the fake user node as the core, use the crossover selection in the genetic algorithm to obtain the chromosomes formed by the excellent genes, and add the optimal relative importance fake links generated by the chromosomes to the topological structure of the community to which the target node belongs.

[0010] As the community detection micro-scale adversarial method based on the perturbation of node relative importance of the present invention, further, obtaining the topological structure of the community to which the target node belongs in the social network includes:

[0011] Use an undirected graph to represent the topological structure of the social network;

[0012] Through community detection of the social network, divide the undirected graph of the social network into several community subgraphs, and each community subgraph is the corresponding community population;

[0013] Select the target node according to the user's needs, and obtain the community population to which the target node belongs.

[0014] As the community detection micro-scale adversarial method based on the perturbation of node relative importance of the present invention, further, dividing the nodes in the community into inferior nodes and reserve nodes according to the magnitude of the importance of all nodes in the community includes:

[0015] Set the magnitude of the importance of the target node in the community as the threshold, and the magnitude of the node importance is calculated according to the degree of the node in the topological structure;

[0016] Nodes in the community that are not lower than the threshold are classified as inferior nodes, and nodes lower than the threshold are classified as preliminary nodes.

[0017] As the micro-scale adversarial method for community detection based on node relative importance perturbation of the present invention, further, excellent genes are screened from the importance gene pool based on the user node fitness and by the roulette selection method, including:

[0018] Taking the normalized node relative importance function as the node fitness function, the normalized node relative importance function is constructed by using the node relative importance, the maximum value of the node relative importance in the community, and the minimum value of the relative importance. The node relative importance is represented by the ratio of the node importance to the number of nodes in the community that are not less than the importance of this node;

[0019] Calculating the fitness of each node by using the node fitness function, and obtaining the total fitness of all nodes in the community by summing according to the fitness of each node;

[0020] Taking the ratio of the fitness of the user node to the total fitness as the probability of each user node being selected, and obtaining the cumulative probability before each user node is selected according to the probability of the user node being selected;

[0021] Generating a random number, and selecting user nodes based on the cumulative probability and the random number, and screening out the predicted links between the selected user nodes and the false nodes from the importance gene pool as excellent genes.

[0022] As the micro-scale adversarial method for community detection based on node relative importance perturbation of the present invention, further, using the cross selection in the genetic algorithm to obtain the chromosomes formed by excellent genes, including:

[0023] Taking the false node as the core, connecting the excellent genes to each other to form a new gene pool;

[0024] Using single-point crossover and multi-point crossover to form chromosomes from the genes in the new gene pool. The single-point crossover is the crossover composed of two genes, and the multi-point crossover is the mutual crossover composed of more than two genes.

[0025] As the micro-scale adversarial method for community detection based on node relative importance perturbation of the present invention, further, adding the optimal relative importance false links generated by the chromosomes to the community topology structure of the target node, including:

[0026] Inserting the chromosome into the community topology structure, and calculating the target node hiding effect index according to the relative importance size of the target node before and after the confrontation, the number of the same nodes before and after the confrontation, and the number of all nodes before and after the confrontation included in the community to which the target node belongs;

[0027] Calculate the overall adversarial effect based on the target node hiding effect index and the difference in community division before and after the confrontation;

[0028] Select the optimal chromosome according to the overall adversarial effect, take the false link corresponding to the optimal chromosome as the optimal relatively important false link, and add the optimal relatively important false link to the topological structure of the community to which the target node belongs.

[0029] As the community detection micro-scale adversarial method based on node relative importance perturbation of the present invention, further, the calculation process of the target node hiding effect index is expressed as: Where α and β are weight coefficients, and α + β = 1, C(n) represents the change in the relative importance of the target node n before and after the confrontation, Q(m) represents the number of the same nodes before and after the confrontation, and Q(n) represents the number of all nodes before and after the confrontation included in the community to which the target node belongs.

[0030] On the other hand, the present invention also provides a community detection micro-scale adversarial system based on node relative importance perturbation, including: a community detection module, a node division module, a link screening module, and a link addition module, where

[0031] The community detection module is used to obtain the topological structure of the community to which the target node belongs in the social network, add false user nodes to the topological structure of the community to which the target node belongs, and establish a false node relatively important candidate link library according to the link relationship between the false user nodes and the original nodes in the community. The relatively important candidate link library is used to store the links that are expected to exist between the false user nodes and all nodes in the community according to the node relative importance. The node relative importance represents the relative importance degree of the node in the community based on the size of the node importance;

[0032] The node division module is used to divide the nodes in the community into inferior nodes and preparatory nodes according to the size of the importance of all nodes in the community, and divide the links in the relatively important candidate link library into inferior genes and preparatory genes. The inferior nodes are used to represent the nodes whose node importance size is greater than the threshold, the preparatory nodes are used to represent the nodes whose node importance is less than the threshold, the inferior genes are used to represent the links that are expected to exist between the false user nodes and the inferior nodes, and the preparatory genes are used to represent the links that are expected to exist between the false user nodes and the preparatory nodes;

[0033] The link screening module is used to eliminate the inferior genes in the relatively important candidate link library through pruning operations, retain the preparatory genes, and form an important gene library. Take the node relative importance as the fitness, and select excellent genes from the important gene library based on the user node fitness by the roulette wheel selection method;

[0034] A link addition module, which uses the crossover selection in the genetic algorithm with a fake user node as the core to obtain chromosomes formed by excellent genes, and adds the optimal relative importance fake links generated by the chromosomes to the community topology structure where the target node belongs.

[0035] Advantages of the present invention:

[0036] Based on the genetic algorithm, the present invention changes the community division result by reducing the relative importance of the target node at the minimum cost through the method of only adding a small number of fake nodes and links within the community without deleting the original structure, establishes a candidate link library for the relative importance of nodes, avoids confusion with real links, uses the pruning principle to eliminate the inferior genes composed of nodes with high importance and fake nodes within the community, then uses the roulette wheel method to double-screen the preparatory genes, comprehensively crosses the preparatory genes, finds the optimal fake links, improves the accuracy of gene selection, maximally reduces the relative importance of the target user, and changes the community division result of the target node, so as to achieve the purpose of protecting the privacy information of the target user. Further, evaluation indicators are used to conduct experimental tests on various community detection algorithms on multiple real network data sets, and compared with four existing adversarial community detection methods. The experimental test results prove that the GIM scheme in this case is significantly better than the traditional baseline method, and can achieve adversarial effects for both overlapping community detection and non-overlapping community detection, and has good application prospects in the field of community detection. Description of the Drawings

[0037] Figure 1 Schematic diagram of the micro-scale adversarial process of community detection based on node relative importance perturbation in the embodiment;

[0038] Figure 2 Schematic diagram of the relative importance and adversarial effect of nodes within the community in the embodiment;

[0039] Figure 3 Schematic diagram of the GIM algorithm process of the present invention in the embodiment;

[0040] Figure 4 Schematic diagram of the fake node candidate link library and pruning principle in the embodiment;

[0041] Figure 5 Schematic diagram of comprehensive crossover and iterative selection in the embodiment;

[0042] Figure 6 Schematic diagram of comparison of hiding effect parameters in the embodiment;

[0043] Figure 7 Schematic diagram of the performance of NMI values of various methods on various different data sets in the embodiment. Detailed Embodiment

[0044] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and technical solutions.

[0045] The importance of users is one of the important bases for community detection. Different users show different degrees of importance in the community. However, social networks often aim to highlight users and cannot reduce the popularity of users by deleting the connections between users. In the embodiments of the present invention, refer to Figure 1 as shown, a micro-scale adversarial method for community detection based on node relative importance perturbation is provided, including:

[0046] S101. Obtain the community topology to which the target node in the social network belongs, add a false user node to the community topology structure to which the target node belongs, and establish a false node relative importance candidate link library according to the link relationship between the false user node and the original nodes in the community. The relative importance candidate link library is used to store the links that are expected to exist between the false user node and all nodes in the community according to the node relative importance. The node relative importance represents the relative importance degree of a node in the community based on the node importance size.

[0047] The importance of a node in the community not only affects the community division result but also affects the overall structure, function, and dynamic characteristics of the community. Therefore, in-depth analysis of node importance is a key factor in improving the division quality and accuracy, which can be defined and measured from multiple angles. Common measurement methods include:

[0048] 1. Centrality

[0049] (1) Degree centrality: The number of connections of a node. The more connections, the higher its importance in the community.

[0050] (2) Betweenness centrality: The frequency of a node appearing in the shortest paths between other nodes. The more frequent, the more important the node.

[0051] (3) Closeness centrality: The average distance between a node and other nodes. The shorter the distance, the higher the importance.

[0052] 2. Community influence

[0053] (1) Leader node: A node that can influence other nodes or play a guiding role within the community.

[0054] (2) Bridge node: A node that connects different communities and can play a key role in information flow.

[0055] 3. Aggregation degree

[0056] The aggregation degree of a node within the community, that is, the connection strength with its neighbors, reflects the closeness of the node within the community.

[0057] In addition to the above three measurement methods, there are also similarity, propagation ability, stability, and so on.

[0058] Without affecting the importance of users, the solution in this case uses the relative importance of nodes in the community to evaluate the relative importance degree of nodes in the community.

[0059] Among them, the calculation process of the relative importance of nodes can be summarized as: within this community, the ratio of the importance of a node to the number of nodes whose importance is not less than that of this node. It is expressed by I(n):

[0060]

[0061] Among them, n represents the user node, i(n) represents the importance of the node, and N(n) represents the number of nodes in the community where the importance is not less than i(n) of this node.

[0062] Such as Figure 2 As shown in the blue community on the left, assuming that the degree is used as the importance, for node 1, i(1)=4, N(1)=1, I(1)=4; for node 2, i(2)=3, N(2)=2, I(2)=1.5.

[0063] The relative importance of a node can effectively reflect the relative importance degree of this node in the community, and at the same time, it will also affect the community division result. Nodes with high relative importance usually play an important role in the community detection process and may also become the core or bridge nodes of the community. Community detection algorithms often give priority to these nodes to ensure that the divided community structure is reasonable and representative. Therefore, by analyzing the relative importance of nodes, the functions and dynamics within the community can be better understood, which helps to optimize the community management strategy and resource allocation. The relative importance size of nodes plays an important role in community division, and the reduction of the relative importance of nodes will not cause the original important status of users in the social network to decline.

[0064] There are two methods to reduce the relative importance of nodes: one is to delete some links of the target node to reduce the connection between the target node and other nodes, and this method does not conform to the original intention of protection; the other is to increase the importance of nodes with lower importance, thereby reducing the relative importance of the target node, which conforms to the original intention of the solution in this embodiment.

[0065] In obtaining the community topology of the target node in the social network, an undirected graph can be used to represent the social network topology structure; through community detection of the social network, the undirected graph of the social network is divided into several community subgraphs, and each community subgraph is a corresponding community population; according to user needs, the target node is selected, and the community population to which the target node belongs is obtained.

[0066] As Figure 3 shown, the entire social network structure G is regarded as an overall community. First, use the community detection algorithm to partition the network into communities, and select the target node according to user needs. Add fake user nodes within the community to which the target node belongs. Use the possible links between the fake nodes and the original nodes within the community to establish a candidate link library of relative importance. Use the pruning principle to eliminate the inferior genes established by the nodes with higher importance and the fake nodes, and then use the roulette method to perform a second screening on the remaining candidate genes. Then use the method of comprehensive crossover to form chromosomes, and finally use the method of double-index limitation to generate the optimal relative importance fake links and insert them into the original structure. Reduce the relative importance of the target node at the lowest cost, affect the result of community partitioning, and thus achieve the hiding of the target node.

[0067] Use the undirected graph G = (N, E) to represent the complete network set in the real world. Use the community detection algorithm Aa to partition the original graph G = (N, E) into K communities C i =(N i ,E i ), i = 1, 2,..., K, where C i ={C1, C2,..., C a} represents the set of partitioned communities, and each community is regarded as a population. N i ={n1, n2,..., n α} represents the set of nodes in the community, and E i ={e1, e2,..., e β} represents the set of links between nodes. The meanings of each symbol are shown in Table 1.

[0068] Table 1 Symbol Table

[0069]

[0070]

[0071] Adopt the method of adding only fake nodes N f and links E f within the community to counter the structure. Add fake nodes within the community to which the target node belongs. This community is not limited to overlapping and non-overlapping, but only limited to the community to which the target node belongs. After adding, update the community to which the target node belongs, which can be expressed as:

[0072] C t =C t ∪{N f}(2)

[0073] While adding the fake node N f , this node may be linked to all nodes N jare connected, and this possible link is set as gene G f . Define the set formed by this gene as the relatively important candidate link library RIL.

[0074] RIL = {(N f , N j ) | N j ∈ C t} (3)

[0075] This library is a set composed of a possible link. This definition method can more comprehensively cover all node users in the community to avoid losing important user link information, and at the same time avoid confusion between real links, which affects the selection of the optimal individual. As Figure 4 shown, assume that node 1 is the target node, n is a false node, and n may be connected to nodes 2, 3, 4, 5, 6, 7, 8, and 9. This possible link is defined as the relatively important candidate link library of node 1.

[0076] S102. Divide the nodes in the community into inferior nodes and reserve nodes according to the importance of all nodes in the community, and divide the links in the relatively important candidate link library into inferior genes and reserve genes. The inferior nodes are used to represent nodes with importance greater than the threshold, and the reserve nodes are used to represent nodes with importance less than the threshold. The inferior genes are used to represent the expected links between false user nodes and inferior nodes, and the reserve genes are used to represent the expected links between false user nodes and reserve nodes.

[0077] Specifically, the importance I(n) of all nodes in the community can be calculated, and the importance T of the target node is set as the threshold. Divide the inferior nodes and reserve nodes according to the threshold. Define the nodes not lower than the threshold as inferior nodes, and the nodes lower than the threshold as reserve nodes.

[0078]

[0079] Call the link between the false node and the inferior node the inferior gene I bad , and the link with the reserve node is called the reserve gene I pre .

[0080]

[0081] S103. Eliminate the inferior genes in the relatively important candidate link library through pruning operations, retain the reserve genes, and form an importance gene library. Use the node relative importance as the fitness, and select excellent genes from the importance gene library based on the user node fitness through the roulette wheel selection method.

[0082] Use the pruning principle to cut off the inferior genes and leave the candidate genes. Therefore, the important gene pool RIL after pruning can be updated to

[0083] RIL' = I pre (6)

[0084] As Figure 4 shown, the importance within the communities of nodes 3, 5, 8, and 9 is greater than that of the target node No. 1. These nodes are called inferior nodes, and the possible links between the false nodes and these nodes are called inferior genes. Use the pruning principle to eliminate these inferior genes and retain the candidate genes.

[0085] Using the pruning principle to initially screen the candidate genes can reduce the workload of the genetic algorithm in selecting the optimal nodes and links, and improve the selection efficiency and accuracy.

[0086] Among them, excellent genes are screened from the important gene pool based on the fitness of user nodes and through the roulette wheel selection method, which can be designed to include:

[0087] Take the normalized node relative importance function as the node fitness function. The normalized node relative importance function is constructed using the node relative importance, the maximum value of the node relative importance in the community, and the minimum value of the relative importance. The node relative importance is represented by the ratio of the node importance to the number of nodes in the community that are not less than the importance of this node;

[0088] Calculate the fitness of each node using the node fitness function, and obtain the total fitness of all nodes in the community by summing according to the fitness of each node;

[0089] Take the ratio of the fitness of the user node to the total fitness as the probability of each user node being selected, and obtain the cumulative probability of each user node before being selected according to the probability of the user node being selected;

[0090] Generate a random number, and select user nodes based on the cumulative probability and the random number. Screen out the expected links between the selected user nodes and the false nodes from the important gene pool as excellent genes.

[0091] Take the normalized relative importance function as the fitness function. Use I norm (n) to represent.

[0092]

[0093] Among them, I(n) represents the relative importance magnitude of the target node, I max (n) represents the node value with the maximum relative importance in this community, and I min (n) represents the node value with the minimum relative importance in this community.

[0094] The process of secondary screening of genes is also a process of selecting the best. In the implementation of this case, the roulette wheel selection method is used to further select genes. Each individual is assigned a "sector" whose area is proportional to its fitness value. Then, by spinning the roulette wheel and randomly selecting a point, it is determined which sector is selected. The steps of the roulette wheel selection method can be summarized as follows:

[0095] (1) First, calculate the fitness I of each user node norm (i). Suppose there are n user nodes, and their fitness values are I norm (1), I norm (2),..., I norm (n).

[0096] (2) Secondly, calculate the total fitness F of all users in the community.

[0097]

[0098] (3) Calculate the selection probability P of each user node i which is the ratio of the fitness of the user node to the total fitness.

[0099]

[0100] (4) Calculate the cumulative probability Q i .

[0101]

[0102] Among them, Q1 = P1, Q2 = P1 + P2, and so on.

[0103] (5) Generate a random number r in [0, 1], and then find the first i that satisfies Q i > r, and this individual i is selected.

[0104] The sectors with larger areas correspond to individuals with higher fitness, so the probability of being selected is higher.

[0105] Through two rounds of screening, the probability of selecting excellent genes is increased. This two-round screening method is more efficient than ordinary screening methods, can effectively reduce the complexity of the entire process, and improve the confrontation efficiency of GIM.

[0106] S104: Using the false user node as the core, utilize the crossover selection in the genetic algorithm to obtain the chromosomes formed by excellent genes, and add the optimal relative importance false links generated by the chromosomes to the topological structure of the community to which the target node belongs.

[0107] According to the number of nodes and links to be added, in the embodiments of this case, a method combining single-point and multi-point crossovers is comprehensively used to form chromosomes from the genes in the new gene pool, calculate the change in the degree of the original target node caused by the chromosomes formed by the crossover, and retain the most qualified offspring chromosomes. During this crossover process, using the false node N f as the core, the selected preparatory genes are connected to each other respectively. Among them, the crossover formed by two genes is called single-point crossover, and the way of crossover of more than two genes is called multi-point crossover. It can be expressed as:

[0108] E = {(G qualityi , N f ), (N f , G qualityj )} (11)

[0109] When i = j = 1, it means that this crossover is a single-point crossover with only two genes using the false node N f as the core. When it means that this crossover is a multi-point crossover with at least three or more genes using the false node N f as the core. As shown in Figure 5 .

[0110] This method has stronger applicability and can select a more appropriate crossover method according to the number of added nodes and links. This method can apply to more scenarios and situations compared with ordinary single crossover methods, and is also conducive to generating the optimal chromosomes.

[0111] Among them, adding the optimal relatively important false links generated by the chromosomes to the topological structure of the community to which the target node belongs may include:

[0112] Insert the chromosome into the community topological structure, and calculate the target node hiding effect index according to the relative importance of the target node before and after confrontation, the number of the same nodes before and after confrontation, and the number of all nodes before and after confrontation included in the community to which the target node belongs;

[0113] Calculate the overall confrontation effect according to the target node hiding effect index and the difference in community division before and after confrontation;

[0114] Select the optimal chromosome according to the overall confrontation effect, take the false link corresponding to the optimal chromosome as the optimal relatively important false link, and add the optimal relatively important false link to the topological structure of the community to which the target node belongs.

[0115] Insert the selected chromosome into the original structure, perform community detection again, and calculate the hiding effect R(n). Set the number of iterations t and the threshold r of the hiding effect. Among them, the qualified chromosome is the chromosome that meets the hiding effect R(n) not less than the threshold r, and the termination condition is to reach the iteration number constant t. This iterative selection can further narrow the range of high-quality chromosomes.

[0116] Use C(n) to represent the change in relative importance, where I'(n) represents the relative importance of the target node after the confrontation.

[0117]

[0118] Use R(n) to represent the hiding effect of the target node. By comparing the nodes in the community to which the target node belongs after the confrontation with the nodes in the original community, the sum of the proportion of the number of identical nodes and the change in relative importance represents the hiding effect.

[0119]

[0120] Among them, α + β = 1, Q(m) represents the number of identical nodes m after comparison, and Q(n) represents the number of all nodes in the target node community before and after the confrontation.

[0121] In the embodiments of this case, the hiding effect index D(n) is used to reflect the change in the number of original nodes and the change in relative importance in the community where the target node is located in the graph structure after the confrontation. The more the number of nodes in the community that are not in the original community, and the more the relative importance decreases, the better the hiding effect. NMI is used to constrain the number of added nodes, minimize the change in the graph structure as much as possible, and achieve the secrecy of the confrontation.

[0122] D(n) = γR(n) + δNMI (14)

[0123] Among them, γ + δ = 1, and the values of γ and δ can be adjusted according to the proportion of the hiding effect and the strategy advantages and disadvantages of NMI in practice. NMI represents the normalized mutual information, which is used to evaluate the difference between the two community partitions before and after the confrontation and also represents the secrecy of the confrontation. The value range of D(n) is (0, 1). The closer it is to 0, the worse the overall confrontation effect; on the contrary, the closer it is to 1, the better the overall confrontation effect.

[0124] Insert the high-quality genes obtained from the above process into the original structure to generate the adversarial network G'=(N', E'). Use the same detection algorithm Aa to partition the graph again to obtain a new structure graph, and obtain the new community C'=(c′1, c′2,..., c′ a ). The community containing the target node N t is represented as Cx'=(n j , N t , e j ), where n i ≠ n j , e i ≠ e j, the relative importance of the target node is I'(x), and I'(x) < I(x). After the confrontation, the number of community nodes and edges containing the target node changes, and the relative importance of the target node decreases.

[0125] B(n) can be used to represent the cost paid by the strategy, and A(n) can be used to represent the superiority and inferiority of the strategy. This kind of index makes up for the shortcomings of the one-sidedness of traditional evaluation indexes, incorporates three elements: hidden effect, structural influence size and confrontation cost, and can more comprehensively evaluate the superiority and inferiority of the strategy. At the same time, this index formula can be applied to various different scenarios for different confrontation measurement strategies, and has universality.

[0126] The cost of adding and deleting edges can be expressed as:

[0127]

[0128] Among them, N(E + ) represents the number of added edges, N(E - ) represents the number of deleted edges, and N is the total number of edges in the entire network structure. k is a constant. Since it is difficult to delete edges in reality and the greater the cost paid, therefore, the k value can be used to adjust the difficulty parameter of deleting edges.

[0129] The total evaluation index can be defined as:

[0130] A (n) = D (n) - B (n) (16)

[0131] B (n) is a pure burden behavior parameter, which is used to measure the cost used in the confrontation. This formula can evaluate the superiority and inferiority of the method.

[0132] Furthermore, based on the above method, an embodiment of the present invention further provides a community detection micro-scale confrontation system based on node relative importance perturbation, including: a community detection module, a node division module, a link screening module and a link addition module, wherein,

[0133] The community detection module is used to obtain the community topology of the target node in the social network, add false user nodes to the community topology structure of the target node, and establish a false node relative importance candidate link library according to the link relationship between the false user nodes and the original nodes in the community. The relative importance candidate link library is used to store the links that are expected to exist between the false user nodes and all nodes in the community according to the node relative importance. The node relative importance represents the relative importance degree of the node in the community based on the size of the node importance;

[0134] The node division module is used to divide the nodes in the community into inferior nodes and candidate nodes according to the importance of all nodes in the community, and divide the links in the relatively important candidate link library into inferior genes and candidate genes. The inferior nodes are used to represent the nodes with node importance greater than the threshold, the candidate nodes are used to represent the nodes with node importance less than the threshold, the inferior genes are used to represent the links expected to exist between false user nodes and inferior nodes, and the candidate genes are used to represent the links expected to exist between false user nodes and candidate nodes;

[0135] The link screening module is used to eliminate the inferior genes in the relatively important candidate link library through pruning operations, retain the candidate genes, and form an importance gene library. Taking the node relative importance as the fitness, excellent genes are screened from the importance gene library based on the user node fitness through the roulette wheel selection method;

[0136] The link addition module is used to take the false user node as the core, use the crossover selection in the genetic algorithm to obtain the chromosomes formed by excellent genes, and add the optimal relatively important false links generated by the chromosomes to the community topology structure of the target node.

[0137] To verify the effectiveness of the solution in this case, the following further explains with experimental data:

[0138] The proposed algorithm GIM in this case is experimented on 4 real-world networks using 3 common community detection methods to detect the hiding effect of the target nodes.

[0139] Among them, the Louvain algorithm is a community detection algorithm based on modularity, aiming to identify the modular structure in the network and divide the community according to the identification of the modular structure. The LPA algorithm is a community detection algorithm based on label propagation. By assigning labels to nodes and making these labels spread in the network, nodes with the same label are divided into a community after reaching stability. The CPM algorithm is an overlapping community detection algorithm that detects communities by searching for adjacent cliques. The basic idea is to discover the community structure by finding the maximum complete subgraph (clique) in the network.

[0140] The dataset and the overlapping community division results are shown in Table 2.

[0141] Table 2 Real social networks and the number of community divisions

[0142]

[0143] The proposed algorithm GIM in this case is compared with the following four baseline methods.

[0144] Safeness - Based Attacks(SBA): By identifying the member links of a certain number of communities and perturbing the community structure by deleting the edges within the communities and adding new links between the communities, the effect of hiding the community where the target node is located is achieved.

[0145] Q - Attack: This method uses modularity Q to design a fitness function and uses a genetic algorithm to delete old links and reconnect a small number of new links to attack the community detection method, achieving community deception and thus achieving the effect of countering community detection.

[0146] DICE: This method uses the idea of heuristic rewiring. By deleting the specified internal links and connecting adjacent nodes in the community to distant nodes, the target is made invisible, achieving the hiding of the target node in the community.

[0147] ProHiCo: This method introduces the idea of likelihood minimization, randomly and fairly allocates perturbation resources, and then selects appropriate edges for perturbation through the method of likelihood minimization.

[0148] The hiding effect index R(n) includes the change in relative importance C(n) and the change in the community where the target node is located. The influence of different parameter values of α and β on the hiding effect can be compared by adding different edges. As Figure 6 shown. It can be seen that when adding different numbers of edges, the influence of parameter changes on the hiding effect shows a normal distribution. From this, it can be inferred that when adding different numbers of edges, when α = β = 0.5, the hiding effect reaches the optimal effect. Therefore, in the experiment, the parameters α and β are set to 0.5.

[0149] The purpose of the experiment is to achieve the best comprehensive effect with a small cost, adding fewer nodes and edges and obtaining a better hiding effect. Therefore, it is necessary to find the value of the added edges. After partitioning different real - world networks using different community detection algorithms, different numbers of optimal nodes and edges are added through a genetic algorithm to calculate the number of communities to which the target node belongs and the impact on the structure. For the NMI value ranges from 0 to 1, the closer it is to 1, the more similar the two community structures are. From Figure 7From the bar chart, it can be seen that the NMI values of these real network datasets vary greatly. Among them, the Q-attack method is the most obvious, indicating that this method has the greatest impact on the structure. Since GIM needs to consider both the adversarial effect and the impact on the structure, and more importantly, the cost of the adversarial attack, its NMI value is not fully prominent in the comparison of this group of data, but it still shows relatively excellent performance. To balance the hiding effect and the impact on the structure, both γ and δ are set to 0.5. This setting can ensure that the hiding effect is achieved while taking into account the impact on the structure. If γ > δ, it focuses on the hiding effect and the proportion of the impact on the structure is relatively small, ignoring the concealment of the overall adversarial attack in the real world and resulting in a large deviation in the overall community structure. If γ < δ, it means that it focuses on the impact on the structure. Although the change in the structure is small, the hiding effect of the target user nodes cannot meet the requirements. Therefore, setting γ = δ = 0.5 can achieve a better hiding effect for the target user as much as possible while having a relatively small impact on the overall community structure, with the best comprehensive effect.

[0150] The optimal number of edges used for adversarial attacks on different datasets and algorithms by selecting different nodes as target user nodes is shown in Table 3.

[0151] Table 3 The optimal number of edges used for adversarial attacks on different algorithms and datasets (1) The optimal number of edges to be added in the adversarial attacks of different community detection algorithms with the 13th node as the target user

[0152]

[0153] (2) The optimal number of edges to be added in the adversarial attacks of different community detection algorithms with the 1st node as the target user

[0154]

[0155]

[0156] (3) The optimal number of edges to be added in the adversarial attacks of different community detection algorithms with the 2nd node as the target user

[0157]

[0158] During the experiment, when conducting adversarial attacks on different target nodes, it can be found that the results shown by different datasets vary greatly. For communities with a smaller difference in degrees between nodes and a denser connection, fewer edges are used for adversarial attacks. While for communities with a larger difference in node degrees, more links are used for adversarial attacks and the cost is greater.

[0159] The value of NMI ranges from 0 to 1. The closer it is to 1, the more similar the two community structures are. From Figure 7From the bar chart, it can be seen that the NMI values of these real network datasets vary greatly. Among them, the Q-attack method is the most obvious, indicating that this method has the greatest impact on the structure.

[0160] Since GIM needs to consider both the adversarial effect and the impact on the structure, and more importantly, the cost of the adversary, the NMI value cannot be fully prominent in the comparison of this group of data, but it can still have a relatively excellent performance.

[0161] According to the experimental results, the overall evaluation A(n) is calculated to evaluate the effect of the proposed solution in this case against the adversarial method, and it is compared with four overlapping community detection methods. As shown in Table 4.

[0162] Table 4 Overall evaluation indicators of different methods

[0163]

[0164] The value range of A(n) is (0, 1). The closer A(n) is to 0, the worse the overall effect of the strategy on the target node, community structure change, and cost paid. On the contrary, the closer it is to 1, the better. It is not difficult to see from the data results in the table that in the process of adversarial against different community detection methods in real network data, the experimental results of the superiority and inferiority indicators of the proposed solution GIM in this case are 17% higher than the optimal SBA indicator, which is significantly better than other adversarial methods. It is found in the adversarial process that there are large errors in the other several methods when adversarial against the CPM overlapping community detection algorithm. It can be judged that these methods are more targeted at the adversarial of non-overlapping community detection algorithms. Thus, it can be concluded that the GIM method is also applicable in the adversarial of overlapping community detection algorithms. It can be seen that this method is applicable to both the adversarial of overlapping community detection and non-overlapping community detection.

[0165] The above experimental results show that the proposed solution in this case takes the target node as an opportunity, and only uses the method of adding false nodes and edges to reconstruct the data structure for overlapping communities, which is closer to reality and has a smaller cost. It can achieve the adversarial effect for both overlapping community detection and non-overlapping community detection, and thus can achieve the purpose of hiding the target node or community.

[0166] Unless otherwise specifically stated, the relative steps, numerical expressions, and values of the components and steps described in these embodiments do not limit the scope of the present invention.

[0167] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0168] The units and method steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation is not considered to exceed the scope of the present invention.

[0169] Those of ordinary skill in the art can understand that all or part of the steps in the above methods can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disc, etc. Optionally, all or part of the steps of the above embodiments can also be implemented using one or more integrated circuits. Correspondingly, each module / unit in the above embodiments can be implemented in the form of hardware or in the form of a software functional module. The present invention is not limited to any specific form of the combination of hardware and software.

[0170] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than to limit it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or make equivalent replacements for some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A micro-scale adversarial method for community detection based on node relative importance perturbation, characterized in that: Include: Obtaining the community topology to which the target node belongs in the social network, adding a fake user node to the community topology structure to which the target node belongs, and establishing a fake node relative importance candidate link library based on the link relationship between the fake user node and the original node in the community, wherein the relative importance candidate link library is used to store the links expected to exist between the fake user node and all nodes in the community based on the relative importance of the node, wherein the relative importance of the node represents the relative importance of the node in the community based on the size of the node importance; According to the importance of all nodes in the community, the nodes in the community are divided into inferior nodes and reserve nodes, and the links in the relative importance candidate link library are divided into inferior genes and reserve genes, wherein the inferior nodes are used to represent nodes whose node importance is greater than a threshold, and the reserve nodes are used to represent nodes whose node importance is less than a threshold, the inferior genes are used to represent the links that are expected to exist between the false user nodes and the inferior nodes, and the reserve genes are used to represent the links that are expected to exist between the false user nodes and the reserve nodes; Through pruning operations, inferior genes in the relative importance candidate link library are removed, and the preliminary genes are retained to form an important gene library. The relative importance of the node is used as the fitness. Based on the user node fitness and the roulette wheel selection method, excellent genes are screened from the important gene library; Taking the false user node as the core, the crossover selection in the genetic algorithm is used to obtain the chromosome formed by excellent genes, and the optimal relative importance false links generated by the chromosome are added to the community topology structure to which the target node belongs.

2. According to claim 1, the community detection micro-scale adversarial method based on node relative importance perturbation is characterized in that: Get the community topology to which the target node belongs in the social network, including: Use undirected graphs to represent social network topology; By performing community detection on the social network, the undirected graph of the social network is divided into several community subgraphs, each of which is a corresponding community population; Select the target node according to user needs and obtain the community population to which the target node belongs.

3. According to claim 1, the community detection micro-scale adversarial method based on node relative importance perturbation is characterized in that: According to the importance of all nodes in the community, the nodes in the community are divided into inferior nodes and reserve nodes, including: The importance of the target node in the community is set as a threshold, wherein the importance of the node is calculated based on the degree of the node in the topological structure; Nodes in the community that are not lower than the threshold are classified as poor-quality nodes, and nodes that are lower than the threshold are classified as reserve nodes.

4. The community detection micro-scale adversarial method based on node relative importance perturbation according to claim 1 is characterized in that: Based on the user node fitness, excellent genes are screened from the important gene pool through the roulette selection method, including: A normalized node relative importance function is used as a node fitness function, wherein the normalized node relative importance function is constructed using node relative importance, a maximum relative importance value of nodes in a community, and a minimum relative importance value, and the node relative importance is represented by a ratio of the node importance to the number of nodes in the community whose importance is not less than that of the node; The node fitness function is used to calculate the fitness of each node, and the total fitness of all nodes in the community is obtained by summing up the fitness of each node; The ratio of the user node's fitness to the total fitness is used as the probability of each user node being selected, and the cumulative probability of each user node before being selected is obtained based on the probability of the user node being selected; Generate random numbers, select user nodes based on cumulative probability and random numbers, and screen out the expected links between the selected user nodes and false nodes as excellent genes from the importance gene library.

5. According to claim 1, the community detection micro-scale adversarial method based on node relative importance perturbation is characterized in that: Using crossover selection in genetic algorithms to obtain chromosomes formed by excellent genes, including: With false nodes as the core, excellent genes are connected to each other to form a new gene pool; The genes in the new gene pool are formed into chromosomes by using single-point crossover and multi-point crossover. The single-point crossover is a crossover between two genes, and the multi-point crossover is a crossover between more than two genes.

6. The community detection micro-scale adversarial method based on node relative importance perturbation according to claim 1 or 5, characterized in that: Add the optimal relative importance pseudo link generated by the chromosome to the community topology structure to which the target node belongs, including: Insert the chromosome into the community topology, and calculate the target node hiding effect index based on the relative importance of the target node before and after the confrontation, the number of identical nodes before and after the confrontation, and the number of all nodes before and after the confrontation contained in the community to which the target node belongs; The overall effect of the confrontation is calculated based on the target node hidden effect index and the difference in community division before and after the confrontation; The optimal chromosome is selected according to the overall confrontation effect, the false link corresponding to the optimal chromosome is used as the false link with the optimal relative importance, and the false link with the optimal relative importance is added to the community topology structure to which the target node belongs.

7. The community detection micro-scale adversarial method based on node relative importance perturbation according to claim 6 is characterized in that: The calculation process of the target node hidden effect index is expressed as: Among them, α and β are weight coefficients, and α+β=1, C(n) represents the change in the relative importance of the target node n before and after the confrontation, Q(m) represents the number of identical nodes m before and after the confrontation, and Q(n) represents the number of all nodes before and after the confrontation contained in the community to which the target node belongs.

8. A community detection micro-scale adversarial system based on node relative importance perturbation, characterized in that: It includes: community detection module, node division module, link screening module and link adding module, among which, A community detection module, used to obtain the community topology to which the target node belongs in the social network, add a fake user node to the community topology structure to which the target node belongs, and establish a fake node relative importance candidate link library based on the link relationship between the fake user node and the original node in the community, wherein the relative importance candidate link library is used to store the links expected to exist between the fake user node and all nodes in the community based on the relative importance of the node, wherein the relative importance of the node represents the relative importance of the node in the community based on the size of the node importance; A node division module is used to divide the nodes in the community into inferior nodes and reserve nodes according to the importance of all nodes in the community, and divide the links in the relative importance candidate link library into inferior genes and reserve genes, wherein the inferior nodes are used to represent nodes whose node importance is greater than a threshold, and the reserve nodes are used to represent nodes whose node importance is less than a threshold, the inferior genes are used to represent the links that are expected to exist between the false user nodes and the inferior nodes, and the reserve genes are used to represent the links that are expected to exist between the false user nodes and the reserve nodes; The link screening module is used to remove inferior genes from the candidate link library of relative importance through pruning operations, retain the preliminary genes, and form an important gene library. The relative importance of the node is used as the fitness, and the excellent genes are screened from the important gene library based on the user node fitness and through the roulette wheel selection method; The link adding module is used to use the false user node as the core, to obtain the chromosome formed by excellent genes by using the crossover selection in the genetic algorithm, and to add the false link with the best relative importance generated by the chromosome to the community topology structure to which the target node belongs.

9. An electronic device, characterized in that: include: at least one processor, and a memory coupled to the at least one processor; The memory stores a computer program, and the computer program can be executed by the at least one processor to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 7 can be implemented.

Citation Information

Patent Citations

  • Method for discovering community structure oriented to directed-weighting network

    CN104391889A

  • Important node identification method based on improved genetic algorithm

    CN115831386A

  • Single-target hard tag community detection countermeasure attack method based on graph neural network

    CN116668060A

  • User community hiding method based on attribute weakening in overlapping communities

    CN117494201A

  • Method for detecting community structure of complicated network

    US20200210864A1