Multi-objective large-scale community detection method based on adaptive selection of surrogate model
By adopting a multi-objective community detection method based on adaptive selection of agent models in large-scale complex networks, the problems of community detection time and resource waste in the existing technology are solved, and efficient community division and detection efficiency are achieved.
Patent Information
- Application Number
- CN202411624897.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-14
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-11-14
AI Technical Summary
When prior art conducts community detection of large-scale complex networks, algorithms are time-consuming and resource waste, and detection efficiency is low.
A multi-objective large-scale community detection method based on adaptive selection of the proxy model is adopted. By obtaining the core node set of the network to be detected, building a similarity matrix, initializing the population, and optimizing using multi-objective genetic algorithm and proxy model, the community division is finally realized.
This method can quickly locate community center nodes, save resources, compress the decision space of the proxy model, reduce the consumption of computer resources and time, improve community detection efficiency, and ensure the quality of community division.
Smart Images

Figure CN119151703B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large-scale community detection, and in particular to a multi-target large-scale community detection method based on adaptive selection of an agent model. Background Art
[0002] In real life, complex systems are modeled as complex networks for analysis by defining each entity in the system as a node and the interaction relationships between entities as edges. The community structure of a complex network is defined as a topological structure that is tight inside and loose outside, that is, a set of nodes, where the nodes within the set interact closely and the nodes between different sets interact loosely. Common complex networks include power networks, aviation networks, computer networks, transportation networks and social networks.
[0003] Community detection is a technique used to reveal network aggregation behavior. It is used to parse modular community structures from complex networks in order to understand the organizational principles and topological structures of complex systems in the real world corresponding to the complex networks. Specifically, community detection divides complex networks into communities. After the division, the nodes in the same community are closely connected, and the denser the better. The connections between nodes in different communities are sparse, and the sparser the better. Due to the increasing amount of data related to the real world, community detection of large-scale complex networks has attracted considerable attention in various fields.
[0004] Community detection can not only divide complex networks into communities, but also solve other problems that require considering the relationship between nodes. For example, the friend recommendation problem in online social networks. Complex online social networks are usually sparse structures, representing a small part of the user's potential online social circle; therefore, the friend recommendation problem can be formalized as a probabilistic model, and the probability of future friendship relationships can be evaluated by detecting the dense social communities of users, so as to find the probability of forming friendships between two non-adjacent users. For example, the Sybil defense problem, Sybil is a fake account with many malicious activities, such as ticket swiping, and controlling certain accounts by providing misleading information through fake identities; users use Sybil to create multiple identities to establish friendships with legitimate accounts, and then attack social networks, and community detection can detect social networks and identify such fake users based on the type of nodes connected to each node. For example, community detection can also be applied to Web community detection, applying the concept of community to the World Wide Web to enhance Web search results and further improve the accuracy of web page retrieval.
[0005] At present, scientists have developed many community detection algorithms based on different principles, which can be divided into three categories. The first category uses traditional graph partitioning and clustering methods, including graph partitioning, hierarchical clustering, partition clustering, and spectral clustering. The second category is the splitting method, which removes the edges linking nodes in different communities. The most popular algorithm was proposed by Girvan and Newman. The third category is optimization algorithms, such as mathematical programming. Among these algorithms, most non-heuristic algorithms are deterministic, and have certain limitations in detection quality and computational efficiency. They are still challenging when performing community detection in complex networks.
[0006] Evolutionary algorithms (EA) belong to meta-heuristic algorithms. Currently, several effective evolutionary algorithms (EA) have been developed for community detection in complex networks. These algorithms can be divided into two types according to the number of optimization objectives: single-objective optimization and multi-objective optimization. In single-objective optimization, modularity is usually used as the optimization objective. Multi-objective optimization usually optimizes multiple conflicting objectives at the same time, so it is widely used in community detection in complex networks in theoretical research. However, in actual network applications, there is little research on the use of multi-objective optimization methods to effectively detect complex large-scale community structures. Therefore, using large-scale community detection algorithms based on multi-objective optimization to improve the application performance of networks will be a very interesting issue for the scientific community.
[0007] Although the evolutionary algorithm EA has shown good performance in community detection, the mechanism of evolutionary algorithm based on population iteration optimization requires more objective function evaluation times, which inevitably leads to time-consuming problems and cannot be popularized to large-scale complex networks. Moreover, as the network size increases, the search space of evolutionary algorithm EA grows exponentially. Therefore, the existing large-scale community detection algorithm based on evolutionary algorithm EA is only suitable for network community detection with relatively few nodes. For community detection in large-scale complex networks, evolutionary algorithms have the problems of time-consuming detection, waste of resources, and low detection efficiency. Summary of the invention
[0008] To this end, the technical problem to be solved by the present invention is to overcome the problems of the prior art in which the algorithm is time-consuming, resources are wasted, and the detection efficiency is low when performing community detection on large-scale complex networks.
[0009] In order to solve the above technical problems, the present invention provides a multi-target large-scale community detection method based on agent model adaptive selection, comprising:
[0010] Obtain the network to be detected and construct an adjacency matrix of the network to be detected;
[0011] Based on the adjacency matrix of the network to be detected, the degree of each node in the network to be detected and the initial average degree of the network to be detected are obtained, core nodes are selected, and a core node set is constructed;
[0012] Based on the adjacency matrix of the network to be detected, a similarity matrix of nodes in the network to be detected is constructed;
[0013] Initialize the parent population based on the core node set and the similarity matrix, and construct an initial parent population containing multiple individuals; each individual corresponds to a candidate solution, which represents a partitioning method of the network to be tested;
[0014] The multi-objective genetic algorithm NSGA-II is used to select, crossover and mutate the initial parent population to obtain the updated parent population until the number of individuals in the updated parent population reaches the preset candidate solution set space size, and the target parent population is obtained. The kernel k-means objective function value KKM and the relative change objective function value RC of the partitioning method represented by each individual in the target parent population are calculated, and the KKM sample space and RC sample space are constructed.
[0015] Based on the KKM sample space and the RC sample space, multiple proxy models in the candidate proxy model pool are trained separately, and the Kendall coefficient of each proxy model is calculated and validated by cross validation. and Spearman coefficient , obtain the performance evaluation index value for evaluating the prediction accuracy of each proxy model, and select the proxy model corresponding to the maximum performance evaluation index value as the target proxy model;
[0016] Using the target proxy model, based on the multi-objective genetic algorithm NSGA-II, the kernel k-means objective function value KKM and the relative change objective function value RC are minimized as the optimization objectives to obtain the optimal target proxy model; the optimal target proxy model is used to perform community detection on the network to be detected to obtain the non-dominated solution set; the individuals contained in the non-dominated optimal layer with the smallest ordinal value in the non-dominated solution set are combined into the optimal solution set output;
[0017] Based on the number of communities in the partitioning method represented by each individual in the optimal solution set, the total number of edges in each community, and the sum of the node degrees, the modularity index value corresponding to each individual is calculated;
[0018] The division method corresponding to the individual with the largest modularity index value is obtained, and the community division is performed on the network to be detected to complete multi-target large-scale community detection.
[0019] Preferably, after obtaining the network to be detected, the network to be detected is simplified, and nodes with a degree of 1 are deleted to obtain a simplified network to be detected.
[0020] Preferably, the method of obtaining the degree of each node in the network to be detected and the initial average degree of the network to be detected based on the adjacency matrix of the network to be detected, selecting core nodes, and constructing a core node set includes:
[0021] Calculate the average degree of all nodes in the network to be detected as the initial average degree;
[0022] Sort the degree of each node in the network to be detected, select the node with the largest degree among the nodes with a degree greater than the initial average degree as the core node, and add it to the core node set;
[0023] The core node and its connected neighbor nodes are deleted from the network to be detected, and an updated network to be detected is obtained;
[0024] Repeatedly select the node with the largest degree from the nodes with a degree greater than the initial average degree in the updated network to be detected, as the new core node, add it to the core node set, and delete the core node and its connected neighbor nodes from the updated network to be detected, until there are no nodes with a degree greater than the initial average degree in the updated network to be detected, and complete the construction of the core node set.
[0025] Preferably, the similarity matrix of the nodes in the network to be detected is constructed based on the adjacency matrix of the network to be detected, which is expressed as:
[0026] ;
[0027] in, represents the diffusion kernel similarity between the i-th node and the j-th node in the network to be detected, and n represents the total number of nodes in the network to be detected; is the preset weight, Represents the Laplace matrix of the network to be tested, and the expression is , represents the degree matrix of the network to be detected, Represents the adjacency matrix of the network to be detected.
[0028] Preferably, the initializing the parent population based on the core node set and the similarity matrix to construct an initial parent population including a plurality of individuals includes:
[0029] Based on the network to be detected, randomly generate Individuals represented by trajectory coding; Decode the trajectory coding using the decoding function to obtain the label coding corresponding to each individual; Select a core node in each community of each individual as the community center to complete the initialization of half of the individuals in the parent population; Indicates the preset population size;
[0030] Select a value that is less than the number of core nodes in the core node set. The random number rank is used as the number of communities; a core node is selected for each community, and based on the similarity matrix, all nodes in the network to be detected are divided into the community corresponding to the core node with the highest similarity to its diffusion kernel, and an individual is obtained; the random number rank is updated until an individual is obtained. individuals, complete the initialization of the other half of the individuals in the parent population, and obtain the initial parent population.
[0031] Preferably, the multi-objective genetic algorithm NSGA-II is used to select, crossover and mutate the initial parent population to obtain an updated parent population, until the number of individuals in the updated parent population reaches a preset candidate solution set space size, and a target parent population is obtained, and the kernel k-means objective function value KKM and the relative change objective function value RC of the partitioning method represented by each individual in the target parent population are calculated, and the KKM sample space and the RC sample space are constructed, including:
[0032] Select, crossover and mutate the current parent population to obtain the current child population;
[0033] Merge the current child population with the current parent population, use the kernel k-means objective function value KKM and the relative change objective function value RC to evaluate the quality of each candidate solution in the merged population, use the non-dominated sorting and crowding distance calculation operations of the multi-objective genetic algorithm NSGA-II to perform survival selection on the merged population, and obtain the updated parent population;
[0034] The updated parent population is selected, crossed and mutated until the number of individuals in the sample space reaches the preset candidate solution space size to obtain the target parent population;
[0035] Taking each individual in the target parent population as a sample and the kernel k-means objective function value KKM corresponding to each individual as a label, a KKM sample space is constructed; taking each individual in the target parent population as a sample and the relative change objective function value RC corresponding to each individual as a label, a RC sample space is constructed.
[0036] Preferably, calculating the kernel k-means objective function value KKM and the relative change objective function value RC corresponding to the division mode represented by each individual includes:
[0037] Construct the kernel k-means objective function value KKM corresponding to the division method represented by the individual, expressed as: ;
[0038] Construct the relative change objective function value RC corresponding to the division method represented by the individual, expressed as: ;
[0039] in, Indicates the total number of nodes in the network to be detected, represents the number of communities; Indicates the i-th community in the currently calculated partitioning method The number of edges is expressed as: ; Indicates the i-th community in the currently calculated partitioning method The number of nodes in ; Indicates the i-th community in the currently calculated partitioning method Other communities connected The number of, the expression is: ; It represents the connection relationship between the i-th node and the j-th node in the network to be detected. The expression is: .
[0040] Preferably, performing selection, crossover and mutation on the parent population to obtain the offspring population includes:
[0041] Selecting father and mother individuals from the initial parent population based on a preset selection algorithm;
[0042] Get a crossover bitmask with a length equal to the number of nodes of individuals in the initial parent population;
[0043] Obtain the genotype corresponding to each node in the parent individual and the mother individual, including the initial genotype and the community center genotype;
[0044] If the mask value in the intersection bit mask is 1, determine the attributes of the nodes at the positions corresponding to the mask value of the parent and mother individuals:
[0045] If both are core nodes, the attributes of the nodes at the positions corresponding to the mask values in the parent and mother individuals are exchanged. At this time:
[0046] If a node changes from a community center node to a non-community center node, the genotypes of the node and the nodes in the community to which the node belongs are modified to the initial genotype;
[0047] If a node changes from a non-community center node to a community center node, its genotype is modified to the community center genotype, and the genotypes of the remaining nodes in its community are modified to the index of the node;
[0048] If not all of them are core nodes, determine whether the core node corresponding to the genotype of the node is the community center:
[0049] If it is a community center, the genotype of the node is set to the index of the community center;
[0050] If it is not a community center, the genotype of the node is set to the community center genotype;
[0051] The remaining nodes with the initial genotype are divided into the communities most similar to themselves according to the similarity matrix to obtain the parent population after crossover;
[0052] Get a mutation point bit mask with a length equal to the number of nodes of individuals in the initial parent population;
[0053] For some individuals in the parent population after crossover, determine whether the node corresponding to the position where the mutation value is 1 in the mutation point mask is a core node:
[0054] If it is a core node, determine whether the core node is a community center node:
[0055] If it is a community center node, the genotype of this node and all nodes in its community is set as the initial genotype;
[0056] If it is not a community center node, the genotype of the node is set to the community center genotype, and the genotype of the neighbor node of the node is set to the index of the node;
[0057] If it is not a core node, the index of the community center node is randomly selected as the mutated genotype.
[0058] Preferably, the candidate proxy model pool includes multiple proxy models, including: regression tree , radial basis function , Support Vector Regression , K-nearest neighbor algorithm With Graph Neural Networks .
[0059] Preferably, the calculation is based on the Kendall coefficient of each proxy model. and Spearman coefficient , obtain the performance evaluation index values for evaluating the prediction accuracy of each surrogate model, including:
[0060] Construct the Kendall coefficient corresponding to each proxy model , expressed as:
[0061] ;
[0062] Construct the Spearman coefficient corresponding to each surrogate model , expressed as: ;
[0063] Kendall coefficients based on each proxy model and Spearman coefficient , calculate the performance evaluation index value corresponding to each proxy model , expressed as: ;
[0064] in, is the model prediction value of the proxy model for multiple networks, is the true label value corresponding to multiple networks; for and The number of consistent pairs after sorting; for and The number of divergent pairs after sorting; Respectively represent data and data The number of tied rankings in ; for and Sequential interpolation, is the total number of nodes in the network to be detected; Represents the weight coefficient, the value range is .
[0065] Preferably, the modularity index value corresponding to each individual is calculated based on the number of communities in the partitioning method represented by each individual in the optimal solution set, the total number of edges in each community, and the sum of the node degrees, and is expressed as:
[0066] ;
[0067] in, represents the modularity index value, Indicates the connection The total number of edges between nodes in a community, , represents the total number of communities in the partitioning method represented by the individual; Indicates The sum of the degrees of the nodes in a community; Represents the total number of edges in the network to be detected.
[0068] Preferably, if there is a real label in the network to be detected, the normalized mutual information NMI corresponding to each individual in the optimal solution set is calculated based on the division method corresponding to the real label and the confusion matrix of each individual in the optimal solution set, which is expressed as:
[0069] ;
[0070] Obtain the division method corresponding to the individual with the largest normalized mutual information NMI, perform community division on the network to be tested, and obtain the community division result;
[0071] in, represents the number of communities in the partition method A represented by the solution in the optimal solution set, Indicates the number of communities in the partition method B corresponding to the true label; represents the confusion matrix; Indicates the The community and the The number of common nodes between communities, Represents the confusion matrix Middle The sum of the row elements, Represents the confusion matrix Middle The sum of the column elements, Indicates the total number of nodes in the network to be detected.
[0072] The above technical solution of the present invention has the following beneficial effects compared with the prior art:
[0073] The multi-objective large-scale community detection method based on adaptive selection of proxy models disclosed in the present invention selects core nodes from a network to be detected so as to quickly locate nodes that may become community centers based on the core nodes, thereby saving resources consumed by community division of large-scale complex networks and compressing the decision space of training network proxy models; the present invention constructs a candidate proxy model pool containing multiple proxy models, takes individual kernel k-means objective function values KKM and relative change objective function values RC as optimization targets, and uses NSGA-II multi-objective genetic algorithm to update KKM sample space and RC sample space to calculate Kendall coefficients and Spearman coefficients of different proxy models in the candidate proxy model pool, and selects the optimal target proxy model so as to use the optimal target proxy model to perform community division on the network to be detected; the present invention uses proxy models to approximate evaluation of KKM and RC objective function values to replace real KKM and RC calculations, thereby avoiding a large number of real function evaluations caused by evolutionary algorithms based on population optimization, reducing the number of calculations of the objective function of each proxy model, realizing community detection of large-scale complex networks, reducing the consumption of computer resources and time, improving community detection efficiency, and ensuring the quality of community division. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below according to specific embodiments of the present invention in conjunction with the accompanying drawings, wherein
[0075] Figure 1It is a flowchart of the steps of the multi-target large-scale community detection method based on adaptive selection of the proxy model provided by the present invention;
[0076] Figure 2 It is a simplified schematic diagram of the network provided by the present invention;
[0077] Figure 3 It is a community detection flow chart of a simplified network to be detected provided by the present invention;
[0078] Figure 4 It is a schematic diagram of obtaining the first core node provided by the present invention;
[0079] Figure 5 is a schematic diagram of obtaining a second core node provided by the present invention;
[0080] Figure 6 It is a schematic diagram of obtaining the third core node provided by the present invention;
[0081] Figure 7 It is a schematic diagram of core node hybrid coding provided by the present invention;
[0082] Figure 8 is a flow chart of the population initialization steps provided by the present invention;
[0083] Fig. 9 It is a schematic diagram of individual generation provided by the present invention;
[0084] Fig.10 It is a schematic diagram of converting an individual into a sample provided by the present invention;
[0085] Fig.11 It is a schematic diagram of population crossover provided by the present invention;
[0086] Fig.12 It is a schematic diagram of population variation provided by the present invention;
[0087] Fig.13 It is a schematic diagram of a similarity matrix of a network provided by the present invention;
[0088] Fig.14 It is a schematic diagram of the crossover process of two parent core nodes provided by the present invention;
[0089] Fig.15 It is the crossover process of two parent common nodes provided by the present invention;
[0090] Fig.16 is a schematic diagram of the mutation provided by the present invention;
[0091] Fig.17 It is a flow chart of the adaptive selection agent model provided by the present invention;
[0092] Fig.18 It is a schematic diagram comparing the predicted value and the original value of the proxy model provided by the present invention. DETAILED DESCRIPTION
[0093] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it, but the embodiments are not intended to limit the present invention.
[0094] Reference Figure 1 As shown, the flowchart of the multi-target large-scale community detection method based on agent model adaptive selection of the present invention specifically includes the following steps:
[0095] S101: Obtain a network to be detected and construct an adjacency matrix of the network to be detected;
[0096] S102: based on the adjacency matrix of the network to be detected, obtaining the degree of each node in the network to be detected and the initial average degree of the network to be detected, selecting core nodes, and constructing a core node set;
[0097] S103: constructing a similarity matrix of nodes in the network to be detected based on the adjacency matrix of the network to be detected;
[0098] S104: Initializing the parent population based on the core node set and the similarity matrix, and constructing an initial parent population including multiple individuals; each individual corresponds to a candidate solution, representing a partitioning method of the network to be detected;
[0099] S105: using the multi-objective genetic algorithm NSGA-II, selecting, crossing over and mutating the initial parent population to obtain an updated parent population, until the number of individuals in the updated parent population reaches the preset candidate solution set space size, obtaining the target parent population, calculating the kernel k-means objective function value KKM and the relative change objective function value RC of the partitioning method represented by each individual in the target parent population, and constructing the KKM sample space and the RC sample space;
[0100] S106: Based on the KKM sample space and the RC sample space, train the multiple proxy models in the candidate proxy model pool separately, and use cross-validation to calculate and calculate the Kendall coefficient based on each proxy model. and Spearman coefficient , obtain the performance evaluation index value for evaluating the prediction accuracy of each proxy model, and select the proxy model corresponding to the maximum performance evaluation index value as the target proxy model;
[0101] The candidate proxy model pool includes multiple proxy models, including: regression tree , radial basis function , Support Vector Regression , K-nearest neighbor algorithm With Graph Neural Networks .
[0102] S107: Using the target agent model, based on the multi-objective genetic algorithm NSGA-II, taking the kernel k-means objective function value KKM and the relative change objective function value RC as the minimum optimization goal, using the agent model to predict the values of KKM and RC, optimizing the community division, and obtaining the non-dominated solution set; the individuals included in the non-dominated optimal layer with the smallest ordinal value in the non-dominated solution set are combined into the optimal solution set output;
[0103] S108: Calculate the modularity index value corresponding to each individual based on the number of communities in the partitioning method represented by each individual in the optimal solution set, the total number of edges in each community, and the sum of the node degrees;
[0104] S109: Obtain the division method corresponding to the individual with the largest modularity index value, divide the network to be detected into communities, and complete multi-target large-scale community detection.
[0105] Among them, the kernel k-means objective function value KKM (kernel k-means Cost) is a common indicator for measuring clustering effects. In community detection, it is an indicator for measuring the closeness of connections between nodes within the community, usually expressed as the ratio of the number of connections between nodes within the community to all possible connections within the community. The relative change objective function value RC is an indicator for measuring the degree of population change, usually used in optimization algorithms such as genetic algorithms; in community detection, it is an indicator for measuring the closeness of connections between nodes outside the community, usually expressed as the ratio of the number of connections between nodes outside the community to all possible connections. KKM and RC play a mutually restrictive role in community discovery; on the one hand, a high KKM value means that the community is closely connected, but it may lead to isolation between communities; on the other hand, a high RC value means that the communities are closely connected, but it may weaken the cohesion within the community. Therefore, the optimization of the KKM indicator and the RC indicator has a certain conflicting relationship. In practical applications, multi-objective optimization is required to find a balance between KKM and RC.
[0106] Specifically, in step S102, based on the adjacency matrix of the network to be detected, the degree of each node in the network to be detected and the initial average degree of the network to be detected are obtained, core nodes are selected, and a core node set is constructed, including:
[0107] S102-1: Calculate the average degree of all nodes in the network to be detected as the initial average degree;
[0108] S102-2: sorting the degree of each node in the network to be detected, selecting the node with the largest degree among the nodes with a degree greater than the initial average degree as the core node, and adding it to the core node set;
[0109] S102-3: Delete the core node and its connected neighbor nodes from the network to be detected, and obtain an updated network to be detected;
[0110] S102-4: Repeatedly select the node with the largest degree from the nodes with a degree greater than the initial average degree in the updated network to be detected, as a new core node, add it to the core node set, and delete the core node and its connected neighbor nodes from the updated network to be detected, until there are no nodes with a degree greater than the initial average degree in the updated network to be detected, and the construction of the core node set is completed.
[0111] Specifically, the similarity matrix of nodes in the network to be detected is expressed as: ;in, represents the diffusion kernel similarity between the i-th node and the j-th node in the network to be detected, and n represents the total number of nodes in the network to be detected; is the preset weight, Represents the Laplace matrix of the network to be tested, and the expression is , represents the degree matrix of the network to be detected, Represents the adjacency matrix of the network to be detected.
[0112] Specifically, in step S104, the parent population is initialized, including:
[0113] S104-1: Based on the network to be detected, use the preset encoding format to randomly generate Individuals represented by trajectory coding; Decode the trajectory coding using the decoding function to obtain the label coding corresponding to each individual; Select a core node in each community of each individual as the community center to complete the initialization of half of the individuals in the parent population; Indicates the preset population size;
[0114] S104-2: Select a number that is less than the number of core nodes in the core node set The random number rank is used as the number of communities; a core node is selected for each community, and based on the similarity matrix, all nodes in the network to be detected are divided into the community corresponding to the core node with the highest similarity to its diffusion kernel, and an individual is obtained; the random number rank is updated until an individual is obtained. individuals, complete the initialization of the other half of the individuals in the parent population, and obtain the initial parent population.
[0115] Specifically, in step S105, constructing the KKM sample space and the RC sample space includes:
[0116] S105-1: Select, crossover and mutate the current parent population to obtain the current child population;
[0117] S105-2: merge the current child population with the current parent population, use the kernel k-means objective function value KKM and the relative change objective function value RC to evaluate the quality of each candidate solution in the merged population, use the non-dominated sorting and crowding distance calculation operations of the multi-objective genetic algorithm NSGA-II to perform survival selection on the merged population, and obtain the updated parent population;
[0118] S105-3: Select, crossover and mutate the updated parent population until the number of individuals in the sample space reaches the preset candidate solution set space size, and obtain the target parent population;
[0119] S105-4: Taking each individual in the target parent population as a sample and the kernel k-means objective function value KKM corresponding to each individual as a label, a KKM sample space is constructed; taking each individual in the target parent population as a sample and the relative change objective function value RC corresponding to each individual as a label, an RC sample space is constructed.
[0120] Among them, the kernel k-means objective function value KKM corresponding to the division method represented by the constructed individual is expressed as: ;
[0121] Construct the relative change objective function value RC corresponding to the division method represented by the individual, expressed as: ;
[0122] in, Indicates the total number of nodes in the network to be detected, represents the number of communities; Indicates the i-th community in the currently calculated partitioning method The number of edges is expressed as: ; Indicates the i-th community in the currently calculated partitioning method The number of nodes in ; Indicates the i-th community in the currently calculated partitioning method Other communities connected The number of, the expression is: ; It represents the connection relationship between the i-th node and the j-th node in the network to be detected. The expression is: .
[0123] Among them, the parent population is selected, crossed and mutated to obtain the offspring population, including:
[0124] Selecting father and mother individuals from the initial parent population based on a preset selection algorithm;
[0125] Get a crossover bitmask with a length equal to the number of nodes of individuals in the initial parent population;
[0126] Obtain the genotype corresponding to each node in the parent individual and the mother individual, including the initial genotype and the community center genotype;
[0127] If the mask value in the intersection bit mask is 1, determine the attributes of the nodes at the positions corresponding to the mask value of the parent and mother individuals:
[0128] If both are core nodes, the attributes of the nodes at the positions corresponding to the mask values in the parent and mother individuals are exchanged. At this time:
[0129] If a node changes from a community center node to a non-community center node, the genotypes of the node and the nodes in the community to which the node belongs are modified to the initial genotype;
[0130] If a node changes from a non-community center node to a community center node, its genotype is modified to the community center genotype, and the genotypes of the remaining nodes in its community are modified to the index of the node;
[0131] If not all of them are core nodes, determine whether the core node corresponding to the genotype of the node is the community center:
[0132] If it is a community center, the genotype of the node is set to the index of the community center;
[0133] If it is not a community center, the genotype of the node is set to the community center genotype;
[0134] The remaining nodes with the initial genotype are divided into the communities most similar to themselves according to the similarity matrix to obtain the parent population after crossover;
[0135] Get a mutation point bit mask with a length equal to the number of nodes of individuals in the initial parent population;
[0136] For some individuals in the parent population after crossover, determine whether the node corresponding to the position where the mutation value is 1 in the mutation point mask is a core node:
[0137] If it is a core node, determine whether the core node is a community center node:
[0138] If it is a community center node, the genotype of this node and all nodes in its community is set as the initial genotype;
[0139] If it is not a community center node, the genotype of the node is set to the community center genotype, and the genotype of the neighbor node of the node is set to the index of the node;
[0140] If it is not a core node, the index of the community center node is randomly selected as the mutated genotype.
[0141] In the embodiment of the present invention, obtaining a performance evaluation index value for evaluating the prediction accuracy of each proxy model includes:
[0142] Construct the Kendall coefficient corresponding to each proxy model , expressed as:
[0143] ;
[0144] Construct the Spearman coefficient corresponding to each surrogate model , expressed as: ;
[0145] Kendall coefficients based on each proxy model and Spearman coefficient , calculate the performance evaluation index value corresponding to each proxy model , expressed as: ;
[0146] in, is the model prediction value of the proxy model for multiple networks, is the true label value corresponding to multiple networks; for and The number of consistent pairs after sorting; for and The number of divergent pairs after sorting; Respectively represent data and data The number of tied rankings in ; for and Sequential interpolation, is the total number of nodes in the network to be detected; Represents the weight coefficient, the value range is .
[0147] Specifically, in step S108, the modularity index value is expressed as: ;in, represents the modularity index value, Indicates the connection The total number of edges between nodes in a community, , represents the total number of communities in the partitioning method represented by the individual; Indicates The sum of the degrees of the nodes in a community; Represents the total number of edges in the network to be detected.
[0148] In this embodiment of the present invention, if there is a real label in the network to be detected, the normalized mutual information NMI corresponding to each individual in the optimal solution set is calculated based on the division method corresponding to the real label and the confusion matrix of each individual in the optimal solution set, which is expressed as:
[0149] ;
[0150] Obtain the division method corresponding to the individual with the largest normalized mutual information NMI, perform community division on the network to be tested, and obtain the community division result; among them, represents the number of communities in the partition method A represented by the solution in the optimal solution set, Indicates the number of communities in the partition method B corresponding to the true label; represents the confusion matrix; Indicates the The community and the The number of common nodes between communities, Represents the confusion matrix Middle The sum of the row elements, Represents the confusion matrix Middle The sum of the column elements, Indicates the total number of nodes in the network to be detected.
[0151] The multi-objective large-scale community detection method based on adaptive selection of proxy models disclosed in the present invention selects core nodes from a network to be detected so as to quickly locate nodes that may become community centers based on the core nodes, thereby saving resources consumed by community division of large-scale complex networks and compressing the decision space of training network proxy models; the present invention constructs a candidate proxy model pool containing multiple proxy models, takes individual kernel k-means objective function values KKM and relative change objective function values RC as optimization targets, and uses NSGA-II multi-objective genetic algorithm to update KKM sample space and RC sample space to calculate Kendall coefficients and Spearman coefficients of different proxy models in the candidate proxy model pool, and selects the optimal target proxy model so as to use the optimal target proxy model to perform community division on the network to be detected; the present invention uses proxy models to approximate evaluation of KKM and RC objective function values to replace real KKM and RC calculations, thereby avoiding a large number of real function evaluations caused by evolutionary algorithms based on population optimization, reducing the number of calculations of the objective function of each proxy model, realizing community detection of large-scale complex networks, reducing the consumption of computer resources and time, improving community detection efficiency, and ensuring the quality of community division.
[0152] Specifically, since the community to which a node with a degree of 1 belongs must be the community to which its only neighboring node belongs, after obtaining the network to be detected, this embodiment simplifies the network to be detected, deletes the nodes with a degree of 1, and obtains a simplified network to be detected. Figure 2 The following is a simplified diagram of the network.
[0153] Based on the above embodiment, in an embodiment of the present invention, a multi-target large-scale community detection method based on adaptive selection of a proxy model provided in this embodiment is used to perform community detection on a simplified network to be detected. Figure 3 As shown in the figure, it is a simplified flowchart of community detection of the network to be detected. The specific steps include:
[0154] S201: Simplify the network to be detected, obtain a simplified network to be detected, construct an adjacency matrix of the simplified network to be detected, calculate the network similarity matrix SM, and then obtain a core node set, including:
[0155] S201-1: Input the adjacency table EDGE of the network to be detected with n nodes, where each row represents an edge and the two columns represent the two endpoints of the edge. Obtain the adjacency matrix A of the network according to EDGE; wherein, It represents the connection relationship between the i-th node and the j-th node in the network to be detected. The expression is: ;
[0156] Among all nodes, the community where the node with degree 1 finally belongs must be the community where its only neighbor belongs, so these nodes are not optimized subsequently, and the size can be obtained as The new adjacency matrix An is de, which is the number of nodes with degree 1.
[0157] S201-2: Calculate the similarity matrix SM: Similarity matrix , each element A similarity measure based on the diffusion kernel similarity is calculated as: , ; Generally, the value is 1, and L is the Laplace matrix of the network. and The similarity between The larger the value of , the more similar the two nodes are.
[0158] S201-3: Obtain the core node set: sort the average degree of the simplified network to be detected and the degree of each node therein, select the node with the largest degree as the core node, and add it to the core node set; delete the core node and its connected neighbors from the simplified network to be detected, and obtain the updated simplified network to be detected; repeatedly select the node with the largest degree from the updated simplified network to be detected as the new core node, add it to the core node set, and delete the core node and its connected neighbors from the updated simplified network to be detected, until there are no nodes in the updated simplified network to be detected, and complete the construction of the core node set. After that, use a binary vector to represent the core node set of the network:
[0159] ;
[0160] in, Representation Node Whether it is a core node, specifically, It means It is the core node. It means It is a normal node.
[0161] According to the vector, the nodes can be divided into two sets, one is the core node set CN, and the other is the ordinary node set NN, which can be expressed as:
[0162] ;
[0163] ;
[0164] in represents the i-th node in CN, The corresponding element in is 1, is the jth node in NN, middle The corresponding element is 0, and r is the number of core nodes, s is the number of ordinary nodes, and , n represents the total number of nodes in the simplified network to be detected.
[0165] Reference Figure 4 As shown, it is a schematic diagram for obtaining the first core node; refer to Figure 5 As shown, it is a schematic diagram for obtaining the second core node; refer to Figure 6 As shown, it is a schematic diagram for obtaining the third core node; refer to Figure 4 , Figure 5 and Figure 6 As shown, and so on, a core node set is constructed.
[0166] S202: Establishing a proxy model adaptive selection model for the objective functions KKM and RC. The initial model includes: sample space size , model pool , the number of cross validation ;
[0167] S202-1: Set the sample space size based on the number of core nodes in the core node set , expressed as:
[0168] ;
[0169] in, is the number of core nodes in the network, is the population size during the optimization process.
[0170] S202-2: Candidate proxy model pool, including: regression tree , radial basis function , Support Vector Regression , K-nearest neighbor algorithm With Graph Neural Networks ;
[0171] ① Based on the residual sum of squares formula , build a regression tree model, expressed as: ; Represents the sample prediction result of the regression tree model, and the expression is ; represents the jth non-overlapping region in the feature space, represents the prediction result of the objective function, Represents the total number of samples.
[0172] ② Based on the cubic kernel function, a radial basis function model is constructed, which can be expressed as: ; r is the Euclidean distance between the input data points;
[0173] ③Build a support vector regression model, including:
[0174] The loss function is expressed as: ;
[0175] The optimization function is expressed as: ;
[0176] Indicates the actual value, represents the predicted value of the support vector regression model, represents the tolerable bandwidth, represents the regression model weight, and is the slack variable, represents the regularization parameter;
[0177] ④Build a K-nearest neighbor algorithm model, including:
[0178] The similarity between nodes is calculated based on the Euclidean distance, expressed as: ;
[0179] For each node, get the closest nodes, this The output values of the nodes are averaged as the prediction value corresponding to each node, expressed as: ;
[0180] in, and are two nodes in the input space, is the dimension of the input space, It is The output of the neighboring nodes.
[0181] S203: Initialize a population with N individuals;
[0182] S203-1: Hybrid coding: In hybrid coding, a positive number represents the index of a node, indicating that the node at that location and the node of the genotype are in the same community, "-1" indicates that the node is the core node of the community center, and "-2" indicates that the node has not yet been assigned to a community. Figure 7 As shown in the figure, it is a schematic diagram of core node hybrid coding. Nodes 5 and 9 are the centers of two communities. These nodes are divided into two communities. and .
[0183] S203-2: Initialize the population: Design the number of individuals is 100; reference Figure 8 As shown in the figure, it is a flow chart of population initialization steps. The specific steps are as follows:
[0184] Based on the simplified network to be tested, using the preset encoding mode, it is generated Individuals represented by trajectory coding; Decode the trajectory coding using the decoding function to obtain the label coding corresponding to each individual; Select a core node in each community of each individual as the community center to complete the initialization of half of the individuals in the parent population;
[0185] For the rest individuals, select less than the number of core nodes in the core node set The random number rank is used as the number of communities, and the setting The genotype of the core node is "-1"; select a core node for each community, and based on the similarity evidence, simplify all the nodes in the network to be detected and divide them into the community corresponding to the core node with the highest similarity to its own diffusion kernel, and obtain an individual; update the random number rank until the individual is obtained. individuals, complete the initialization of the other half of the individuals in the parent population, and obtain the initialized parent population.
[0186] Reference Fig. 9 As shown, it is a schematic diagram of individual generation; all genotypes of Ind1 are initialized to "-2", and then two core nodes are randomly selected As the core of the community, each node The SM value determines their genotype.
[0187] S203-3: Calculate the two objective functions KKM and RC for each individual, perform fast non-dominated sorting on the individuals and record the number of evaluations of the objective function .
[0188] The calculation formulas for KKM and RC are as follows:
[0189] ;
[0190] in, is the number of communities, Representing the community The number of internal connections in Representing the community The number of connections with other communities, is the adjacency matrix, Representation Node The number of edges between .
[0191] S203-4: Constructing the initial KKM sample space of the proxy model and the initial RC sample space of the proxy model, including:
[0192] Get the initial parent population, compress it, keep only the core nodes, and get the compressed parent population;
[0193] The kernel k-means objective function value KKM of the compressed parent population and each individual in it is constructed as the initial KKM sample space of the surrogate model;
[0194] The relative change objective function value RC of the compressed parent population and each individual in it is constructed as the initial RC sample space of the surrogate model.
[0195] Specifically, refer to Fig.10 As shown, it is a schematic diagram of the transformation of individuals into samples; for the initial population, each individual Corresponding to KKM and RC, The genotype of the core node is used as the feature vector, KKM and RC are used as the feature values, and the samples of the adaptive selection model are added with KKM and RC proxy models. and .
[0196] S204: Through multi-objective genetic algorithm Optimize the partitioning of large-scale community networks;
[0197] S204-1: Preset the number of evaluations of the optimization objective function, use the variable Ass as a counter, and record the number of times the current objective function is evaluated ;
[0198] S204-2: Perform selection, crossover, and mutation operations on the parent population to generate offspring, update the kernel k-means objective function value KKM and the relative change objective function value RC of the offspring individuals, and record the number of evaluations as , and compress the offspring individuals according to different objective functions and add them to the sample and And update the minimum value of the objective function in the sample;
[0199] Specifically, the parent population is subjected to selection, crossover, and mutation operations to generate offspring, including:
[0200] ① Crossover: Randomly select two individuals from the parent population as the parents of the crossover; generate a mask for the crossover point, 1 means that the crossover is performed at this position, and 0 means no crossover. If the corresponding mask value is 1, determine the current node attribute: if the current node is in the core node set, and the core node status at the same position of the crossover parents is different, then the position will undergo gene crossover; if a core node of a community center becomes a non-community center, the genotype becomes "-2", and the genotype of the nodes in the same community is also set to "-2"; if a node that is not a community center becomes a community center, the genotype is set to "-1", and the genotype of the nodes in the parent generation that are in the same community as the node and are not community centers is set to the index of the crossover position. If the current node is not in the core node set, if the core node corresponding to the genotype is the community center, the genotype of the ordinary node is set to the index of the core node at the same position; if the core node corresponding to the genotype is not the community center, the genotype of the crossover position is set to the genotype of the core node. After the crossover is completed, the remaining nodes with a genotype of "-2" are assigned to the community that is most similar to the community center according to the similarity matrix. Reference Fig.11 The figure shows a schematic diagram of population crossover.
[0201] ② Mutation: For all individuals in the parent population, a mask corresponding to the mutation position is generated. If the position is a core node and the mutation occurs at the community center, the node genotype of the same community is set to "-2". If it is a core node that is not a community center, the genotype is set to "-1". At the same time, according to SM, the genotype of the neighbor of the core node at the mutation position is set to the index of the core node. If the position is not a core node, the NN node is mutated, that is, the community center is randomly selected as the genotype. Fig.12 The following is a schematic diagram of population variation.
[0202] Reference Fig.13 As shown, it is a schematic diagram of the similarity matrix of a network; refer to Fig.14 As shown, it is a schematic diagram of the crossover process of two parent core nodes; refer to Fig.15 The figure shows the crossover process of two parent common nodes. Fig.16 Shown is a schematic diagram of the mutation.
[0203] S204-3: Merge the parent population with the child population, and perform fast non-dominated sorting and crowding calculation on the merged individuals;
[0204] S204-4: Use environmental selection to screen the merged parent population and offspring population to obtain a new parent population;
[0205] S204-5: construct the current parent population and the objective function KKM of each individual therein as a proxy model to update the KKM sample space;
[0206] S204-6: Construct the current updated parent population and the objective function RC of each individual therein as a proxy model to update the RC sample space.
[0207] S204-7: Determine whether the set maximum number of evaluations has been reached. If so, stop iterating and execute step S204-8. Otherwise, continue to determine whether the set sample space size has been reached. If so, execute step S204-5. Otherwise, , return to step S204-2.
[0208] S204-8: Take the set of non-dominated solutions with the smallest order value in the current population as the solution set output for community partitioning.
[0209] S205: Training the adaptive surrogate model selection model for KKM and RC, refer to Fig.17 As shown in the figure, it is a flow chart of the adaptive selection agent model, which specifically includes:
[0210] S205-1: Set the number of cross-validation times , create two initial arrays of evaluation criteria and , is the total number of models.
[0211] S205-2: Randomly divide the sample space into group, extract from The remaining group is used as a test set, and the number of training times is recorded. ;
[0212] S205-3: Training The proxy model represented by is successfully trained and Tau and Rho are calculated and recorded in and Otherwise, and ;
[0213] S205-4: Determine whether all models have been trained. If so, execute step S205-5. Otherwise, execute step S205-3. ;
[0214] S205-5: Determine whether the maximum number of cross-validation times has been reached. If so, execute step S205-6. Otherwise, execute step S205-2. ;
[0215] S205-6: Calculate TR. The TR calculation formula can be expanded to:
[0216] ;in, . for The matrix of .
[0217] S205-7: Calculate the corresponding The average value of:
[0218] ;in is the number of cross validations.
[0219] S205-8: Get the target proxy model:
[0220] ;
[0221] in, is the target proxy model, for The label for the maximum value.
[0222] Specifically, the Kendall coefficient Tau and Spearman coefficient Rho corresponding to each proxy model are constructed, including:
[0223] Kendall coefficient ;
[0224] Spearman coefficient ;
[0225] Among them, and Calculate Tau and Rho from two data. is the surrogate model value predicted by the surrogate model, is the sample objective function value. for and The number of consistent pairs after sorting, for and The number of divergent pairs after sorting. Respectively represent data The number of tied rankings and data The number of tied rankings. for and Sequential interpolation, is the number of data. Fig.18 As shown in Figure 1, it is a schematic diagram comparing the predicted value of the proxy model with the original value.
[0226] The two standards are integrated to select the target proxy model. The calculation formula is as follows:
[0227] ;in, is the coefficient, and its value range is .
[0228] S206: Use the agent model to perform genetic algorithm to optimize community division. After obtaining the agent model, most of the calculation of the objective function is replaced by the agent model. The specific steps are as follows:
[0229] S206-1: The evaluation number Ass is inherited from step S204;
[0230] S206-2: Perform selection, crossover, and mutation operations on the parent population to generate offspring, and use the proxy model to calculate the objective function value corresponding to the offspring individuals;
[0231] S206-3: Check the individuals whose objective function calculated by the agent model is less than the recorded minimum value, and record the set of these individuals ;
[0232] S206-4: Gather The individuals in are recalculated with the original objective function, the number of objective function evaluations ass is recorded, and these individuals are added to the sample space;
[0233] S206-5: Update the proxy model. Bring the new sample space into step S205 to obtain a new target proxy model.
[0234] S206-6: The parent population is merged with the child population, and the merged individuals are quickly sorted and the crowding degree is calculated;
[0235] S206-7: Use environmental selection to screen the merged parent population and offspring population to obtain a new parent population;
[0236] S206-8: Determine whether the set maximum number of evaluations has been reached. If so, stop iterating and execute step S206-9. Otherwise, , return to step S206-2;
[0237] S206-9: Take the set of non-dominated solutions with the smallest order value in the current population as the solution set output for community partitioning.
[0238] S207: According to the solution set obtained in step S206 , and get the final community division. The specific steps are as follows:
[0239] S207-1: Create a complete network solution ,Will The labels of the nodes are filled in according to the positions of the nodes in the original network. middle;
[0240] S207-2: Extraction Individuals in , nodes with degree 1 in the original network , looking for his neighbor in Tags , do the following: ;
[0241] S207-3: If the network has a true label, calculate NMI to obtain the final optimal solution, otherwise execute step S207-4:
[0242] ;
[0243] in and Is the division method and The number of corresponding communities, is the confusion matrix, For partition Central Community With partition Central Community The same number of nodes, yes No. The sum of the row elements, yes No. The sum of the column elements, is the number of nodes in the network, The larger the value, the more community division and The more similar. The largest one is the final community division result;
[0244] S207-4: Calculate Q to obtain the final optimal solution:
[0245] ;
[0246] in, is the number of communities, is the total number of edges in the network, To connect the community The total number of edges in the node, It's a community The sum of the node degrees. The larger the Q, the higher the quality of the detected community. The community with the largest Q value is the final community division result.
[0247] The multi-objective large-scale community detection method based on adaptive selection of proxy models disclosed in the present invention selects core nodes from a network to be detected so as to quickly locate nodes that may become community centers based on the core nodes, thereby saving resources consumed by community division of large-scale complex networks and compressing the decision space of training network proxy models; the present invention constructs a candidate proxy model pool containing multiple proxy models, takes individual kernel k-means objective function values KKM and relative change objective function values RC as optimization targets, and uses NSGA-II multi-objective genetic algorithm to update KKM sample space and RC sample space to calculate Kendall coefficients and Spearman coefficients of different proxy models in the candidate proxy model pool, and selects the optimal target proxy model so as to use the optimal target proxy model to perform community division on the network to be detected; the present invention uses proxy models to approximate evaluation of KKM and RC objective function values to replace real KKM and RC calculations, thereby avoiding a large number of real function evaluations caused by evolutionary algorithms based on population optimization, reducing the number of calculations of the objective function of each proxy model, realizing community detection of large-scale complex networks, reducing the consumption of computer resources and time, improving community detection efficiency, and ensuring the quality of community division.
[0248] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0249] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0250] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0251] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0252] Obviously, the above embodiments are merely examples for the purpose of clear explanation and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived therefrom are still within the scope of protection of the present invention.
Claims
1. A multi-target large-scale community detection method based on agent model adaptive selection, characterized in that: include: Obtain the network to be detected and construct an adjacency matrix of the network to be detected; Based on the adjacency matrix of the network to be detected, the degree of each node in the network to be detected and the initial average degree of the network to be detected are obtained, core nodes are selected, and a core node set is constructed; Based on the adjacency matrix of the network to be detected, a similarity matrix of nodes in the network to be detected is constructed; Initialize the parent population based on the core node set and the similarity matrix, and construct an initial parent population containing multiple individuals; each individual corresponds to a candidate solution, which represents a partitioning method of the network to be tested; The multi-objective genetic algorithm NSGA-II is used to select, crossover and mutate the initial parent population to obtain the updated parent population until the number of individuals in the updated parent population reaches the preset candidate solution set space size, and the target parent population is obtained. The kernel k-means objective function value KKM and the relative change objective function value RC of the partitioning method represented by each individual in the target parent population are calculated, and the KKM sample space and RC sample space are constructed, including: Select, crossover and mutate the current parent population to obtain the current child population; Merge the current child population with the current parent population, use the kernel k-means objective function value KKM and the relative change objective function value RC to evaluate the quality of each candidate solution in the merged population, use the non-dominated sorting and crowding distance calculation operations of the multi-objective genetic algorithm NSGA-II to perform survival selection on the merged population, and obtain the updated parent population; The updated parent population is selected, crossed and mutated until the number of individuals in the sample space reaches the preset candidate solution space size to obtain the target parent population; Taking each individual in the target parent population as a sample and the kernel k-means objective function value KKM corresponding to each individual as a label, a KKM sample space is constructed; taking each individual in the target parent population as a sample and the relative change objective function value RC corresponding to each individual as a label, a RC sample space is constructed; Based on the KKM sample space and the RC sample space, multiple proxy models in the candidate proxy model pool are trained separately, and the Kendall coefficient of each proxy model is calculated and validated by cross validation. and Spearman coefficient , obtain the performance evaluation index value for evaluating the prediction accuracy of each proxy model, and select the proxy model corresponding to the maximum performance evaluation index value as the target proxy model; Among them, the Kendall coefficient corresponding to each proxy model is constructed , expressed as: ; Construct the Spearman coefficient corresponding to each proxy model , expressed as: ; Kendall coefficients based on each proxy model and Spearman coefficient , calculate the performance evaluation index value corresponding to each proxy model , expressed as: ; is the model prediction value of the proxy model for multiple networks, is the true label value corresponding to multiple networks; for and The number of consistent pairs after sorting; for and and The number of divergent pairs after sorting; Respectively represent data and data The number of tied rankings in ; for and Sequential interpolation, is the total number of nodes in the network to be detected; Represents the weight coefficient, the value range is ; Using the target proxy model, based on the multi-objective genetic algorithm NSGA-II, the kernel k-means objective function value KKM and the relative change objective function value RC are minimized as the optimization objectives to obtain the optimal target proxy model; the optimal target proxy model is used to perform community detection on the network to be detected to obtain the non-dominated solution set; the individuals contained in the non-dominated optimal layer with the smallest ordinal value in the non-dominated solution set are combined into the optimal solution set output; Based on the number of communities in the partitioning method represented by each individual in the optimal solution set, the total number of edges in each community, and the sum of the node degrees, the modularity index value corresponding to each individual is calculated, expressed as: ;in, represents the modularity index value, Indicates the connection The total number of edges between nodes in a community, , represents the total number of communities in the partitioning method represented by the individual; Indicates The sum of the degrees of the nodes in a community; Represents the total number of edges in the network to be detected; The division method corresponding to the individual with the largest modularity index value is obtained, and the community division is performed on the network to be detected to complete multi-target large-scale community detection.
2. The multi-target large-scale community detection method based on agent model adaptive selection according to claim 1, characterized in that: After obtaining the network to be detected, the network to be detected is simplified, and nodes with a degree of 1 are deleted to obtain a simplified network to be detected.
3. The multi-target large-scale community detection method based on agent model adaptive selection according to claim 1, characterized in that: The method of obtaining the degree of each node in the network to be detected and the initial average degree of the network to be detected based on the adjacency matrix of the network to be detected, selecting core nodes, and constructing a core node set includes: Calculate the average degree of all nodes in the network to be detected as the initial average degree; Sort the degree of each node in the network to be detected, select the node with the largest degree among the nodes with a degree greater than the initial average degree as the core node, and add it to the core node set; The core node and its connected neighbor nodes are deleted from the network to be detected, and an updated network to be detected is obtained; Repeatedly select the node with the largest degree from the nodes with a degree greater than the initial average degree in the updated network to be detected, as the new core node, add it to the core node set, and delete the core node and its connected neighbor nodes from the updated network to be detected, until there are no nodes with a degree greater than the initial average degree in the updated network to be detected, and complete the construction of the core node set.
4. The multi-target large-scale community detection method based on agent model adaptive selection according to claim 1, characterized in that: Based on the adjacency matrix of the network to be detected, a similarity matrix of nodes in the network to be detected is constructed, which is expressed as: ; in, represents the diffusion kernel similarity between the i-th node and the j-th node in the network to be detected, and n represents the total number of nodes in the network to be detected; is the preset weight, Represents the Laplace matrix of the network to be tested, and the expression is , represents the degree matrix of the network to be detected, Represents the adjacency matrix of the network to be detected.
5. The multi-target large-scale community detection method based on agent model adaptive selection according to claim 4 is characterized in that: The method of initializing the parent population based on the core node set and the similarity matrix to construct an initial parent population including multiple individuals includes: Based on the network to be tested, randomly generate Individuals represented by trajectory coding; Decode the trajectory coding using the decoding function to obtain the label coding corresponding to each individual; Select a core node in each community of each individual as the community center to complete the initialization of half of the individuals in the parent population; Indicates the preset population size; Select a value that is less than the number of core nodes in the core node set. The random number rank is used as the number of communities; a core node is selected for each community, and based on the similarity matrix, all nodes in the network to be detected are divided into the community corresponding to the core node with the highest similarity to its diffusion kernel, and an individual is obtained; the random number rank is updated until the individual is obtained. individuals, complete the initialization of the other half of the individuals in the parent population, and obtain the initial parent population.
6. The multi-target large-scale community detection method based on agent model adaptive selection according to claim 1, characterized in that: Calculate the kernel k-means objective function value KKM and the relative change objective function value RC corresponding to the division method represented by each individual, including: Construct the kernel k-means objective function value KKM corresponding to the division method represented by the individual, expressed as: ; Construct the relative change objective function value RC corresponding to the division method represented by the individual, expressed as: ; in, Indicates the total number of nodes in the network to be detected, represents the number of communities; Indicates the i-th community in the currently calculated partitioning method The number of edges is expressed as: ; Indicates the i-th community in the currently calculated partitioning method The number of nodes in ; Indicates the i-th community in the currently calculated partitioning method Other communities connected The number of, the expression is: ; It represents the connection relationship between the i-th node and the j-th node in the network to be detected. The expression is: .
7. The multi-target large-scale community detection method based on agent model adaptive selection according to claim 1, characterized in that: Perform selection, crossover and mutation on the parent population to obtain the offspring population, including: Selecting father and mother individuals from the initial parent population based on a preset selection algorithm; Get a crossover bitmask with a length equal to the number of nodes of individuals in the initial parent population; Obtain the genotype corresponding to each node in the parent individual and the mother individual, including the initial genotype and the community center genotype; If the mask value in the intersection bit mask is 1, determine the attributes of the nodes at the positions corresponding to the mask value of the parent and mother individuals: If both are core nodes, the attributes of the nodes at the positions corresponding to the mask values in the parent and mother individuals are exchanged. At this time: If a node changes from a community center node to a non-community center node, the genotypes of the node and the nodes in the community to which the node belongs are modified to the initial genotype; If a node changes from a non-community center node to a community center node, its genotype is modified to the community center genotype, and the genotypes of the remaining nodes in its community are modified to the index of the node; If not all of them are core nodes, determine whether the core node corresponding to the genotype of the node is the community center: If it is a community center, the genotype of the node is set to the index of the community center; If it is not a community center, the genotype of the node is set to the community center genotype; The remaining nodes with the initial genotype are divided into the communities most similar to themselves according to the similarity matrix to obtain the parent population after crossover; Get a mutation point bit mask with a length equal to the number of nodes of individuals in the initial parent population; For some individuals in the parent population after crossover, determine whether the node corresponding to the position where the mutation value is 1 in the mutation point mask is a core node: If it is a core node, determine whether the core node is a community center node: If it is a community center node, the genotype of this node and all nodes in its community is set as the initial genotype; If it is not a community center node, the genotype of the node is set to the community center genotype, and the genotype of the neighbor node of the node is set to the index of the node; If it is not a core node, the index of the community center node is randomly selected as the mutated genotype.
8. The multi-target large-scale community detection method based on agent model adaptive selection according to claim 1, characterized in that: The candidate proxy model pool includes multiple proxy models, including: regression tree , radial basis function , Support Vector Regression , K-nearest neighbor algorithm With Graph Neural Networks .
9. The multi-target large-scale community detection method based on agent model adaptive selection according to claim 1, characterized in that: Also includes: If the network to be detected has a true label, then based on the division method corresponding to its true label and the confusion matrix of each individual in the optimal solution set, the standardized mutual information NMI corresponding to each individual in the optimal solution set is calculated, which is expressed as: ; Obtain the division method corresponding to the individual with the largest normalized mutual information NMI, perform community division on the network to be tested, and obtain the community division result; in, represents the number of communities in the partition method A represented by the solution in the optimal solution set, Indicates the number of communities in the partition method B corresponding to the true label; represents the confusion matrix; Indicates the The community and the The number of common nodes between communities, Represents the confusion matrix Middle The sum of the row elements, Represents the confusion matrix Middle The sum of the column elements, Indicates the total number of nodes in the network to be detected.
Citation Information
Patent Citations
Heterogeneous social network community detection method based on genetic algorithm
CN103605793A
Multi-target community detection method based on k node updating and a similarity matrix
CN106934722A