Multi-agent reinforcement learning assisted multi-target community detection method

Through the multi-objective community detection method assisted by multi-agent reinforcement learning, combined with reinforcement learning and multi-objective evolution algorithm, the problems of insufficient utilization of historical information and insufficient random search orientation in community detection are solved, and more efficient and accurate community detection is achieved.

CN120373342APending Publication Date: 2025-07-25GUANGZHOU INSTITUTE OF TECHNOLOY XIDIAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510226232.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing community detection algorithms are using insufficient historical information and insufficient random search orientation, resulting in low search efficiency and poor accuracy.

Method used

The multi-objective community detection method assisted by multi-agent reinforcement learning is adopted. By building a dynamic strategy optimization framework, combining reinforcement learning and multi-objective evolution algorithm, it uses sliding windows and external shared archives to store historical information and reward mechanisms to guide the global evolution direction of community detection.

Benefits of technology

It improves the convergence speed and division robustness of Pareto's frontier, and improves the search efficiency and accuracy of community detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373342A_ABST
    Figure CN120373342A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-agent reinforcement learning assisted multi-target community detection method, which relates to the technical field of reinforcement learning, and comprises the following steps: coding community network nodes; initializing population parameters; constructing and initializing a sliding window, an external shared file and a threshold value; a two-dimensional table Q-table is constructed and initialized, and a multi-layer learning strategy is adopted to initialize a T-table; selecting a neighbor node for each node of each individual according to a sliding window and a threshold triggering selection condition, and completing updating of the individuals; storing the non-dominated solution to an external shared file A; updating the two two-dimensional tables; and calculating MIRgen and adding the MIRgen to the sliding window. According to the method, reinforcement learning and a multi-objective evolutionary algorithm are fused, the defects of insufficient utilization of historical information and low efficiency of random search in a traditional algorithm are overcome, community detection global evolution is accurately regulated and controlled through a dynamic strategy optimization framework, the Pareto frontier convergence speed and robustness are improved, and the problems of low search efficiency and poor precision are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of reinforcement learning, and particularly to a multi-objective community detection method assisted by multi-agent reinforcement learning. Background Art

[0002] In recent years, with the increasing complexity of the topological structures of various network systems in the real world (such as social networks, biomolecular networks, and communication networks), how to effectively analyze their internal organizational characteristics has become a key technical challenge. Community Detection is a core means of analyzing the organizational structure of complex networks. By identifying clusters of nodes in the network with high cohesion (i.e., dense internal node connections) and low coupling (i.e., sparse cross-community connections), it can decompose the complex global network into multiple functionally or structurally independent sub-modules, thereby realizing the simplified expression of network topology and the mining of deep laws. In practical applications, community detection plays an important role in multiple fields such as social software, personalized recommendation, disease control, and illegal drug control.

[0003] Community detection essentially belongs to an NP-hard optimization problem, and the scale of its solution space expands exponentially with the growth of the number of network nodes. Currently, the solution framework based on evolutionary algorithms has been widely used in this field. As a heuristic search algorithm, evolutionary algorithms optimize the solution strategy by simulating natural selection and genetic mechanisms. The typical process of using evolutionary algorithms to solve community detection problems usually includes defining a suitable fitness function to evaluate the quality of community detection, and designing effective genetic operations (such as selection, crossover, and mutation) to explore the solution space. The fitness function usually measures the effectiveness of community partitioning based on the modularity of the network or other structural features.

[0004] In recent years, multi-objective evolutionary algorithms have also been applied to solve community detection problems, by simultaneously optimizing multiple objectives, such as the balance between modularity and community size, to find the optimal network partitioning structure. Although these algorithms show good performance in solving community detection, there are still some limitations in previous studies that have not been fully resolved:

[0005] (1) First, most existing evolutionary algorithms only collect and utilize a small amount of historical data. These algorithms usually only record the individuals with good performance and their evolutionary directions, while ignoring other important historical information, such as the historical community selection of each node. The lack of information makes it difficult for the algorithm to obtain knowledge about the correct evolutionary direction;

[0006] (2) Second, most existing algorithms usually adopt randomly selected favorable evolutionary directions to improve the community structure, while ignoring the mutual influence between individual differences and community partitioning, resulting in low search efficiency of the algorithm. Summary of the Invention

[0007] Aiming at the problems existing in the prior art, the present invention provides a multi-objective community detection method assisted by multi-agent reinforcement learning. Through the collaborative mechanism of reinforcement learning and multi-objective evolutionary algorithm, aiming at the bottleneck of one-sided historical information mining and insufficient random search orientation in the traditional evolutionary algorithm, a multi-agent collaborative decision-making framework based on dynamic policy optimization is constructed to achieve precise control of the global evolutionary direction of community detection, effectively improve the convergence speed of the Pareto front and the partitioning robustness, and solve the technical problems of slow search efficiency and low search accuracy.

[0008] The technical solution of the present invention is realized as follows:

[0009] A multi-objective community detection method assisted by multi-agent reinforcement learning, applied to a community network, the community network including a plurality of nodes; comprising the following steps:

[0010] S1. Encode the nodes; initialize the population parameters, and randomly initialize and generate a plurality of individuals for the population; the information of one individual includes the association of each node with other nodes; the node connection relationship in the individual has nothing to do with the original community network, but is a predicted solution for the future; initialize the individual by randomly generating connection relationships for each node; construct a sliding window W, the sliding window W including a plurality of values MIR i , where i is an index and i>0; initialize MIR1 = 1; construct an external shared archive A and initialize it as an empty set; initialize a threshold t;

[0011] S2. Construct a two-dimensional table Q-table and T-table; initialize the values of the Q-table to 0; initialize the T-table using a multi-layer learning strategy.

[0012] The Q-table is the core data structure in reinforcement learning (especially the Q-learning algorithm), used to store the expected cumulative reward values of state-action pairs. Each cell records the long-term value (Q value) of taking action a in a specific state s, guiding the agent to select the optimal action.

[0013] The data structures of the Q-table and the T-table are the same;

[0014] In the early stage of evolution, since the Q-table has not been established yet, the agent mainly relies on the guidance of the multi-layer learning strategy and focuses on learning the local high-order information of the network. Therefore, the T-table is used to guide the evolution direction of the agent, and it extracts and stores the high-order topological information in the community network. As the agent interacts with the environment, it will gradually turn to using the accumulated historical information and reward mechanism to explore the solution space. The Q-table records the historical information in the evolution process and uses it as experience to guide the evolution direction and complete the reinforcement learning.

[0015] In a community network (such as a social network, a biological network, or a communication network), a node is the basic unit that constitutes the network topology. Physically, a node refers to an independent terminal device in the network. Abstractly, a node can be abstracted as any object in the network, such as a user in a social network, etc.

[0016] S3. For each of the individuals, denoted as the parent individual, trigger the selection condition according to the sliding window W and the threshold t. The conditions include a first condition and a second condition. If the first condition is triggered, then based on the T-table, use the roulette wheel strategy to select neighbor nodes for each node of the parent individual. If the second condition is triggered, then based on the Q-table, use the ε-greedy strategy to select neighbor nodes for each node of the parent individual.

[0017] Construct a new individual using the selected neighbor nodes, denoted as the offspring individual. The offspring individual replaces the parent individual to complete the update of the population. Calculate the fitness value of the offspring individual.

[0018] The roulette wheel strategy is to allocate selection probabilities according to the fitness values of individuals. The higher the fitness, the greater the probability of being selected. The core steps include calculating individual probabilities, cumulative probabilities, and simulating the "roulette wheel" selection process through random numbers.

[0019] The ε-greedy strategy (ε-Greedy Strategy) is a classic method in reinforcement learning to solve the exploration-exploitation dilemma. Its core idea is to balance the exploitation of the current optimal action and the exploration of unknown actions by dynamically adjusting the weights of exploration and exploitation.

[0020] S4. Use the fast non-dominated sorting method to select non-dominated solutions from the population and store the non-dominated solutions in the external shared archive A.

[0021] In the present invention, the fast non-dominated solution sorting method adopts a classic method in the field of evolution. The specific steps are as follows:

[0022] S4-1. Assign two key quantities to each individual p in the population: the number N of solutions that dominate p p , and the solution set S dominated by p p ; Set up the storage set F;

[0023] S4-2. Set i = 1, and classify the individuals with N p = 0 into F i ;

[0024] S4-3. For the individuals in F i , traverse S p of each solution p, and subtract 1 from the N p of each solution therein;;

[0025] S4-4. Increment i by 1; Classify the solutions with N p = 0 into F i ;

[0026] S4-5. Repeat S4-3 and S4-4 until all individuals in the population are classified into F;

[0027] S4-6. Take out the individuals in F1 and put them into the external shared archive A, thus completing the update of the external shared archive A.

[0028] S5. Update the Q-table and T-table;

[0029] S6. Calculate MIR gen , and add it to the sliding window W; where gen represents the current iteration round;

[0030] S7. Repeat S3 to S6 until the preset iteration end condition is reached.

[0031] The present invention adopts a multi-agent reinforcement learning method, and assigns an agent (i.e., T-table and Q-table) to each node. The multi-agents select appropriate coding labels for the current node by using historical evolutionary information and environmental rewards, so as to provide a promising evolutionary direction for individuals and populations. The action reward of each agent is calculated by the consensus feedback mechanism proposed by the present invention, and this mechanism feeds back on the behavior decision of the agent by using the elite solution information generated by the evolutionary population.

[0032] The final external shared archive A is the detection result of the algorithm, including the Pareto front composed of several non-dominated solutions. The Pareto front is the set of all Pareto optimal solutions in the objective function space of a multi-objective optimization problem. Its core feature is that for any solution on the Pareto front, it is impossible to make all objective functions improve simultaneously by adjusting parameters (that is, there is no other solution that is not inferior to it in all objectives and is strictly better in at least one objective).

[0033] Preferably, the threshold value t can be customized according to requirements. The iteration end condition is the customizable number of iterations. When the iteration round reaches the preset number of iterations, the iteration ends.

[0034] As a further optimization of the above solution, in S1, the coding representation of the node is v, v ∈ {1, 2, 3,..., n}; an individual is represented as X = {x1, x2,..., x i ,…,x n}; where, x i = j means that node v i and node v j are connected, and node v j and node v i are neighbor nodes to each other;

[0035] The initialization of the individual is to randomly assign the neighbor nodes to any one of the nodes; several nodes with a connection relationship form a community, and nodes that are not connected belong to different communities.

[0036] Several nodes with direct or indirect connection relationships form a connected component (connected graph), and multiple nodes belonging to the same connected component belong to the same community; a community network can be divided into one or more communities.

[0037] As a further optimization of the above solution, the multi-layer learning strategy is to calculate the node local similarity, node degree influence, and node central density according to the community network; here it refers to calculating based on the original community network data, rather than the predicted node data (i.e., the individuals / solutions of the population). Specifically, the original community network is represented by an adjacency matrix M.

[0038] The calculation process of the node local similarity S is as follows:

[0039]

[0040] Among them, τ(i) represents the set of node v i , including the neighbor nodes of node v i and node v i ; M represents the adjacency matrix of the community network, M ij = 1 means that node v i and node v j are connected; M ij = 0 means that node v i and node v j are not connected; ∪ and ∩ respectively represent taking the union and intersection of two sets; |·| represents the number of nodes in a set; S ij represents node v i and node v jLocal similarity between; node local similarity represents the distance between neighbor nodes; generally, a node is more likely to be assigned to the same community as similar neighbor nodes.

[0041] The calculation process of the node degree influence D is as follows:

[0042]

[0043] where d i represents the degree of node v i ; D i represents the node degree influence of node v i ; the degree of a node is the number of nodes connected to the specified node; the node degree influence can reveal the importance of neighbor nodes; specifically here, it refers to the degree in the original community network.

[0044] Select an individual from the population, and calculate the node central density C according to the selected individual. The calculation process is as follows:

[0045]

[0046] where V i represents the set of nodes in the community to which node v i belongs in the selected individual; represents the degree of node v i in community V i ; C i represents the node central density of node v i . Similarly, represents the degree of node v j in community V j . The node central density can measure the centrality of neighbor nodes in the community.

[0047] The local density refers to the number of nodes within a specific range in the same community. Generally speaking, nodes with higher local density are closer to the community center. Using local density can dynamically represent the centrality of a node in the community. It should be noted that if the sum of the degrees of a node and its neighbor nodes in the same partition is directly used as the local density, it may lead to an incorrect evaluation of the local density because some nodes rely heavily on their neighbor nodes and thus obtain a large density. To solve this problem, this solution uses the above formula to calculate the node central density.

[0048] The central density comprehensively considers the degree of a node and the contribution weight of its neighbor nodes, thereby enhancing the distinguishability of central nodes.

[0049] Preferably, an individual can be randomly selected from the population, or an individual with the largest modularity MOD can be selected.

[0050] As a further optimization of the above solution, in S2, the initialization process of the T-table is expressed as:

[0051]

[0052] where NS, ND, and NC respectively represent the normalized node local similarity, node degree influence, and node central density; α, β, and γ are all weight values; T ij represents the value at the i-th row and j-th column in the T-table;

[0053] In S5, the T-table is updated using the multi-layer learning strategy.

[0054] The node local similarity, node degree influence, and node central density form three layers of data; the T-table is constructed by normalizing the information of each layer of data and then integrating the three layers of information through weighted summation. Among them, the node local similarity is normalized for each row, and after normalizing the node degree influence and node central density, they are combined with each row of the node local similarity respectively.

[0055] According to the roulette wheel strategy, the larger the value of T ij in the process of evolution, the greater the probability that node v i selects node v j as its neighbor node.

[0056] By combining the three layers of information, each node can utilize the local high-order information in the network, including the local static information and dynamic information of the network.

[0057] After each iterative update of the population, the central density of the nodes will change. Therefore, each time the T-table is updated, only the node central density needs to be recalculated, while the values of the node local similarity and node degree influence remain unchanged.

[0058] As a further optimization of the above solution, in S3, for node v i , a T ij corresponds to the probability of selecting a v j ; when the first condition is triggered, according to the probability, a neighbor node v i is selected for node v j , until the offspring individual is generated.

[0059] Each row of the two-dimensional table (T-table or Q-table) represents a node, and each column represents the adjacent nodes that the current node may connect to. After selecting an adjacent node for each row according to the T-table or Q-table and its related strategy, a new individual can be constructed. The greater the probability, the greater the probability of being selected as an adjacent node.

[0060] As a further optimization of the above solution, in S3, when there is any MIR in the sliding window W i greater than the threshold t, the first condition is triggered; when all MIRs in the sliding window W i are less than the threshold t, the second condition is triggered.

[0061] As a further optimization of the above solution, for each individual in the external shared file A, the modularity MOD is calculated respectively, that is:

[0062]

[0063] where m is the sum of the edges of any two nodes in the community network; σ(v i , v j ) represents the community relationship between node v i and node v j , σ(v i , v j ) = 1 means that node v i and node v j belong to the same community, σ(v i , v j ) = 0 means that node v i and node v j do not belong to the same community;

[0064] In S6, the calculation of MIR gen is as follows:

[0065]

[0066] where Arc gen represents the average modularity of the external shared file A in the gen-th iteration; |A| represents the number of individuals in the external shared file A;

[0067] The sliding window W is a first-in-first-out queue with a fixed length of w.

[0068] In the early stage of evolution, since the Q-table has not been established yet, the agent mainly relies on the guidance of the multi-layer learning strategy (T-table) and focuses on learning the local high-order information of the network. With the interaction with the environment, the agent will gradually turn to using the accumulated historical information and reward mechanism to explore the solution space. In this process, how to balance the learning of the local structure of the network and the utilization of evolutionary knowledge is crucial. The present invention proposes an adaptive control strategy for controlling the transition between the two stages.

[0069] The modularity improvement amount is a method for directly measuring the evolutionary state, but the variation range of the modularity improvement is significantly different in different problems and stages. Generally speaking, the modularity improvement is much greater in the early stage of evolution than in the later stage. To address this issue, an adaptive control strategy based on the modularity improvement rate is proposed, which uses a sliding window W to store the value of the modularity improvement rate MIR gen 。

[0070] As a further optimization of the above solution, in S5, the update of the Q-table is expressed as:

[0071]

[0072] where M represents the adjacency matrix of the community network, M ij = 1 indicates that node v i and node v j are connected;

[0073] M ij = 0 indicates that node v i and node v j are not connected; d i represents the degree of node v i ;

[0074] R(a, b) represents the reward value when node v a selects node v b as a neighbor node;

[0075] V a is the set of nodes in the community network that belong to the same community as node v a ; |V a | represents the number of nodes in set V a ;

[0076] m is the sum of the edges of any two of the nodes in the community network; L is a preset learning rate; |A| represents the number of individuals in the external shared archive A.

[0077] Considering that the community change of a certain node will not only affect the original community where the node is located, but also affect the community that the node is about to join. Therefore, the present invention proposes a multi-agent reinforcement learning reward strategy, which uses group consensus information to evaluate the behavior of each agent. During the population evolution process, an external shared archive A is used to store the non-dominated solutions generated during the evolution process, and the archive is updated in each generation. The elite solutions stored in the external shared archive A have good performance on both dual objectives. Therefore, the actions selected by the solutions in the archive should receive higher rewards. To quantify the contribution of the actions in the external shared archive A, the present invention rewards each action according to its contribution to the improvement of the overall community modularity. The agent a obtains a reward R(a, b) for taking the action b, which means that when the node v a selects the node v b as the neighbor node, the reward value. The parameter L is the learning rate, which represents the influence degree of the new evolutionary experience on updating the existing knowledge.

[0078] The present invention proposes an environmental reward strategy based on a consensus feedback mechanism for updating the reinforcement learning strategy. This strategy evaluates the behavior of the agent through group consensus information, not only effectively utilizes the community division information of the nodes, but also can provide more accurate feedback for the actions of the reinforcement learning, thereby optimizing the learning process of the behavior strategy.

[0079] As a further optimization of the above solution, when the second condition is triggered, the process of generating the offspring individual is as follows:

[0080]

[0081] where ε is a preset adaptive parameter; rand is a random value; Q(v i , A c ) represents the Q value obtained when the node v i executes the action A c , and forms a neighbor node with the node v j ; Argmax[Q(v i , A c )] represents selecting the action A c to maximize the Q value; v r represents randomly selecting a neighbor node for the node v i .

[0082] When the second condition is triggered, reinforcement learning is used to select neighbor nodes for each node, and then offspring individuals are generated. ε is used to balance the utilization of known information and the exploration of unknown space by the reinforcement learning. A c represents all the actions that can be taken in the current state. In the present invention, the Q value is the reward value in the Q-table, that is, Q(v i , A c) = R(i, j).

[0083] As a further optimization of the above solution, in S3, the offspring individuals are divided into q communities, namely {V1, V2, …, V k , …, V q}, where V k represents the set of nodes in the k-th community of the offspring individuals;

[0084] The fitness value includes f1 and f2, and the calculation process is as follows:

[0085]

[0086] Among them, M represents the adjacency matrix of the community network, M ij = 1 indicates that node v i and node v j are connected; M ij = 0 indicates that node v i and node v j are not connected; V represents the set of all nodes in the community network; |·| represents the size of the set; L(V k , V k ) represents the internal connection strength of community V k ; represents the external connection strength of community V k .

[0087] The present invention uses two indicators, Kernel K-Means (KKM) and Ratio Cut (RC), to evaluate the fitness of individuals in the solution space. Among them, KKM is obtained by calculating the sum of the internal connection densities of communities. The smaller the KKM value, the higher the internal connection density of the communities of this individual; RC represents the sum of the connection densities between communities. The smaller the RC value, the lower the connection density between the communities of this individual.

[0088] The smaller the values of f1 and f2, the better the quality of the individuals (solutions) obtained after update.

[0089] Among two individuals, if the two fitness values of the first individual are both smaller than those of the second individual, the second individual is considered to be dominated by the first individual; if the two fitness values of the first individual are both greater than those of the second individual, the first individual is considered to be dominated by the second individual; otherwise, they are non-dominated solutions to each other.

[0090] Compared with the prior art, the present invention has the following beneficial effects:

[0091] (1) The present invention proposes a novel framework combining reinforcement learning and evolutionary algorithms to solve the community detection problem. This framework guides the optimization direction of the evolutionary algorithm through a reinforcement learning mechanism and uses the historical information generated during the population evolution process to update the reinforcement learning strategy in real-time. Compared with traditional evolutionary algorithms, this framework can make full use of the historical evolution information and environmental rewards in reinforcement learning to explore and accumulate effective evolutionary behavior knowledge.

[0092] (2) The present invention proposes an environmental reward strategy based on a consensus feedback mechanism to update the reinforcement learning strategy. This strategy evaluates the behavior of agents through group consensus information, not only effectively utilizes the community division information of nodes, but also can provide more accurate feedback for the actions of reinforcement learning, thereby optimizing the learning process of the behavior strategy.

[0093] (3) The present invention proposes a multi-layer learning strategy to learn the local high-order information in the network. By dynamically mining the local similarity of nodes, the influence of node degrees, and the central density of nodes in the network, it helps multi-agents quickly learn the local structure of the network, thereby providing guidance for the early evolution of the population and effectively improving the evolutionary efficiency of the algorithm in the initial stage.

[0094] (4) The present invention proposes a stage adaptive adjustment strategy to dynamically balance the learning of the local structure of the network by agents and the exploration of historical evolutionary experience based on the modular improvement rate. This strategy can reduce redundant learning processes and improve the overall performance and efficiency of the algorithm. Brief Description of the Drawings

[0095] Figure 1 is a schematic flow chart of a multi-objective community detection method assisted by multi-agent reinforcement learning provided by an embodiment of the present invention;

[0096] Figure 2 is a schematic diagram of the relationship between population, individual, node, and community provided by an embodiment of the present invention;

[0097] Figure 3 is a data processing demonstration diagram of the multi-layer learning strategy provided by an embodiment of the present invention;

[0098] Figure 4 is a schematic diagram of the process of constructing offspring individuals from parent individuals provided by an embodiment of the present invention;

[0099] Figure 5 is a schematic diagram of the data structure of the sliding window provided by an embodiment of the present invention. Detailed Embodiment

[0100] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts belong to the scope of protection of the present invention.

[0101] As Figure 1 shown, this embodiment provides a multi-objective community detection method assisted by multi-agent reinforcement learning, which is applied to a community network. The community network includes multiple nodes; the method includes the following steps:

[0102] S1. Encode the nodes; in this embodiment, in S1, the encoding of the nodes is represented as v, where v ∈ {1, 2, 3,..., n}.

[0103] Initialize the population parameters and randomly initialize multiple individuals for the population; the information of one individual includes the association between each node and other nodes; specifically, the initialization of an individual is to randomly assign neighbor nodes to any one node; an individual is represented as X = {x1, x2,..., x i ,..., x n}; where x i = j means that node v i and node v j are connected, and node v j and node v i are neighbor nodes to each other. As Figure 2 shown, an individual includes 7 nodes, which are respectively represented as node 1, node 2, node 3, node 4, node 5, node 6, and node 7; the adjacent nodes selected by each individual are node 2, node 3, node 1, node 3, node 6, node 7, and node 5; several nodes with a connection relationship form a community, and nodes that are not connected to each other belong to different communities. For example, nodes 1, 2, 3, and 4 form a community, and nodes 5, 6, and 7 form another community.

[0104] Several nodes with direct or indirect connection relationships form a connected component (connected graph), and multiple nodes belonging to the same connected component belong to the same community; a community network can be divided into one or more communities.

[0105] The node connection relationship in the individual has nothing to do with the original community network, but is the solution to the future prediction; initializing the individual is to randomly generate a connection relationship for each node.

[0106] Construct a sliding window W. The sliding window W is a first-in-first-out queue with a fixed length of w. The sliding window W includes several values MIR i, where i is an index and i > 0; initialize MIR1 = 1; construct an external shared archive A and initialize it as an empty set; initialize a threshold t;

[0107] S2. Construct two-dimensional tables Q-table and T-table; initialize the values of Q-table to 0; initialize T-table using a multi-layer learning strategy;

[0108] In this embodiment, the multi-layer learning strategy is to calculate the local similarity of nodes, the influence of node degrees, and the central density of nodes based on the community network; here it refers to calculating based on the original community network data, rather than the predicted node data (i.e., the individuals / solutions of the population). Specifically, the original community network is represented by an adjacency matrix M.

[0109] The calculation process of the local similarity S of nodes is as follows:

[0110]

[0111] where τ(i) represents the set of nodes v i including the node v i and the neighbor nodes of node v i ; M represents the adjacency matrix of the community network, M ij = 1 indicates that node v i and node v j are connected; M ij = 0 indicates that node v i and node v j are not connected; ∪ and ∩ respectively represent taking the union and intersection of two sets; |·| represents the number of nodes in a set; S ij represents the local similarity between node v i and node v j ; the local similarity of nodes represents the distance between neighbor nodes; generally, a node is more likely to be assigned to the same community as similar neighbor nodes.

[0112] The calculation process of the node degree influence D is as follows:

[0113]

[0114] where d i represents the degree of node v i ; D i represents the node degree influence of node v i ; the degree of a node is the number of nodes connected to the specified node; the node degree influence can reveal the importance of neighbor nodes; here it specifically refers to the degree in the original community network;

[0115] Select an individual with the largest modularity MOD from the population, and calculate the node central density C according to the selected individual. The calculation process is as follows:

[0116]

[0117] Among them, V i represents the set of nodes in the community to which node v i belongs in the selected individual; represents the degree of node v i in community V i ; C i represents the node central density of node v i . Similarly, represents the degree of node v j in community V j . The node central density can measure the centrality of neighbor nodes in the community.

[0118] The local density refers to the number of nodes within a specific range in the same community. Generally speaking, nodes with higher local density are closer to the community center. Using the local density can dynamically represent the centrality of nodes in the community. It should be noted that if the sum of the degrees of nodes and their neighbor nodes in the same partition is directly used as the local density, some nodes may obtain a large density because they rely heavily on their neighbor nodes, which will lead to an incorrect evaluation of the local density. To solve this problem, this solution uses the above formula to calculate the node central density.

[0119] The central density comprehensively considers the degree of the node and the contribution weight of its neighbor nodes, thereby improving the distinguishability of central nodes.

[0120] In this embodiment, the initialization process of the T-table is expressed as:

[0121]

[0122] Among them, NS, ND, and NC respectively represent the normalized node local similarity, node degree influence, and node central density; α, β, and γ are all weight values; T ij represents the value in the i-th row and j-th column of the T-table.

[0123] The node local similarity, node degree influence, and node central density constitute three layers of data; the construction method of the T-table is to normalize the information of each layer of data, and then integrate the three layers of information through weighted summation. Among them, the node local similarity is normalized for each row, and after normalizing the node degree influence and node central density, they are combined with each row of the node local similarity respectively.

[0124] Such asFigure 3 As shown, the node local similarity values of node 6 for nodes 1, 7, 8, and 9 are 0.29, 0.6, 0.6, and 0.8 respectively, and the node local similarity values for nodes 2, 3, 4, and 5 are all 0. The sum of the node local similarity values is 0.29 + 0.6 + 0.6 + 0.8 = 2.29. Then, after normalization, the node local similarities of node 6 for the remaining nodes are 0.29 / 2.29, 0, 0, 0, 0, 0, 0.6 / 2.29, 0.6 / 2.29 (i.e., NS 61 to NS 69 ), 208 / 2.29; similarly, after normalization, the degree influence of node 6 on other nodes is 3 / 10, 0, 0, 0, 0, 0, 2 / 10, 2 / 10, 3 / 10 (i.e., ND 61 to ND 69 ), and the central densities are 3.43 / 15.53, 0, 0, 0, 0, 0, 3.5 / 15.53, 3.5 / 15.53, 51. / 15.53 (i.e., NC 61 to NC 69 ). Based on this, the probability value T 61 to T 69 can be obtained by integration.

[0125] According to the roulette wheel strategy, the larger the value of T ij , the greater the probability that node v i selects node v j as a neighbor node during the evolution process.

[0126] By combining the three-layer information, each node can utilize the local high-order information in the network, including the local static information and dynamic information of the network.

[0127] The Q-table is the core data structure in reinforcement learning (especially the Q-learning algorithm), which is used to store the expected cumulative reward values of state-action pairs (State-Action Pair). Each cell records the long-term value (Q-value) of taking action a in a specific state s, guiding the agent to select the optimal action.

[0128] The data structures of the Q-table and the T-table are the same; the T-table stores the local high-order information between nodes; the Q-table stores the historical experience information of individuals during the evolution process. Each row of the two-dimensional table represents an agent for a node, which is used to guide the node to select neighbor nodes and construct new individuals.

[0129] In the early stage of evolution, since the Q-table has not been established yet, the agent mainly relies on the guidance of the multi-layer learning strategy and focuses on learning the local high-order information of the network. Therefore, the T-table is used as the agent to guide the evolution direction, and it extracts and saves the high-order topological information in the community network. With the interaction with the environment, the agent will gradually turn to using the accumulated historical information and reward mechanism to explore the solution space. The Q-table records the historical information in the evolution process and uses it as experience to guide the evolution direction and complete the reinforcement learning.

[0130] In a community network (such as a social network, a biological network, or a communication network), a node is the basic unit that constitutes the network topology. Physically, a node refers to an independent terminal device in the network. Abstractly, a node can be abstracted as any object in the network, such as a user in a social network, etc.

[0131] S3. For each individual, denoted as the parent individual, trigger the selection condition according to the sliding window W and the threshold t. The conditions include the first condition and the second condition. Specifically, when there is any MIR i in the sliding window W that is greater than the threshold t, the first condition is triggered; when all MIRs i in the sliding window W are less than the threshold t, the second condition is triggered.

[0132] If the first condition is triggered, then based on the T-table, adopt the roulette wheel strategy to select neighbor nodes for each node of the parent individual. Specifically, for the node v i , a T ij corresponds to selecting a probability for a v j . When the first condition is triggered, select the neighbor node v i for the node v j according to the probability until the offspring individual is generated.

[0133] If the second condition is triggered, then based on the Q-table, adopt the ε-greedy strategy to select neighbor nodes for each node of the parent individual. Specifically, when the second condition is triggered, the selection process is as follows:

[0134]

[0135] where ε is a preset adaptive parameter; rand is a random value; Q(v i , A c ) represents the Q value obtained when the node v i executes the action A c and forms a neighbor node with the node v j ; Argmax[Q(v i , A c )] represents selecting the action A c, maximize the Q value; v r Denoted as node v i Randomly select a neighbor node.

[0136] When the second condition is triggered, use reinforcement learning to select neighbor nodes for each node, and then generate offspring individuals. ε is used to balance the utilization of known information and the exploration of unknown space by reinforcement learning. A c Denotes all actions that can be taken in the current state. In the present invention, the Q value is the reward value in the Q-table, that is, Q(v i , A c ) = R(i, j).

[0137] Construct a new individual using the selected neighbor nodes, denoted as the offspring individual; the offspring individual replaces the parent individual to complete the update of the population. As Figure 4 shown, among the information originally stored by the individual, for nodes v1 to v6 (i.e., nodes 1 to 6), the connected neighbor nodes are 5, 6, 2, 1, 6, 5, 6 respectively. After being guided and selected by the Q-table or T-table, among the generated new individuals, the neighbor nodes selected by nodes 1 to 6 are 2, 3, 1, 3, 6, 7, 5 respectively.

[0138] Each row of the two-dimensional table (T-table or Q-table) represents a node, and each column represents the adjacent nodes that the current node may connect to. After selecting an adjacent node for each row according to the T-table or Q-table and its related strategies, a new individual can be constructed to complete the evolution of the individual, that is, the evolution of the population. The greater the probability, the greater the probability of being selected as an adjacent node.

[0139] The roulette wheel strategy assigns selection probabilities according to the fitness values of individuals. The higher the fitness, the greater the probability of being selected. The core steps include calculating individual probabilities, cumulative probabilities, and simulating the "roulette wheel" selection process through random numbers.

[0140] The ε-greedy strategy is a classic method in reinforcement learning to solve the exploration-exploitation dilemma. Its core idea is to balance the utilization of the current optimal action and the exploration of unknown actions by dynamically adjusting the weights of exploration and exploitation.

[0141] Calculate the fitness value of the offspring individual.

[0142] Specifically, first divide the offspring individual into q communities, that is, {V1, V2, …, V k , …, V q}, V k represents the set of nodes in the kth community of the offspring individual;

[0143] The fitness values include f1 and f2, and the calculation process is as follows:

[0144]

[0145] Among them, M represents the adjacency matrix of the community network, M ij = 1 indicates that node v i and node v j are connected; M ij = 0 indicates that node v i and node v j are not connected; V represents the set of all nodes in the community network; |·| represents the size of the set; L(V k , V k ) represents the internal connection strength of community V k ; represents the external connection strength of community V k .

[0146] The present invention uses two indicators, Kernel K-Means (KKM) and Ratio Cut (RC), to evaluate the fitness of individuals in the solution space. Among them, KKM is obtained by calculating the sum of the internal connection densities of communities. The smaller the KKM value, the higher the internal connection density of the community of this individual; RC represents the sum of the connection densities between communities. The smaller the RC value, the lower the connection density between the communities of this individual.

[0147] The smaller the values of f1 and f2, the better the quality of the individual (solution) obtained after updating.

[0148] Among two individuals, if the two fitness values of the first individual are both smaller than those of the second individual, it is considered that the second individual is dominated by the first individual; if the two fitness values of the first individual are both greater than those of the second individual, it is considered that the first individual is dominated by the second individual; otherwise, the two are non-dominated solutions to each other.

[0149] S4. Select non-dominated solutions from the population using the fast non-dominated sorting method, and store the non-dominated solutions in the external shared archive A;

[0150] In the present invention, the fast non-dominated solution sorting method uses a classic method in the evolutionary field, and the specific steps are as follows:

[0151] S4-1. Assign two key quantities to each individual p in the population: the number N p of solutions that dominate p, and the solution set S p dominated by p; set the storage set F;

[0152] S4-2. Set i = 1, and set N pIndividuals with a value of 0 are classified into F i ;

[0153] S4-3. For the individuals in F i traverse the S of each solution p p and subtract 1 from the N of each solution therein;; p

[0154] S4-4. Increment i by 1; Classify solutions with N p = 0 into F i ;

[0155] S4-5. Repeat S4-3 and S4-4 until all individuals in the population are classified into F;

[0156] S4-6. Take out the individuals in F1 and put them into the external shared archive A, thus completing the update of the external shared archive A.

[0157] S5. Update the Q-table and T-table.

[0158] Among them, the T-table is updated using the multi-layer learning strategy in S2. That is, re-select an individual with the largest modularity MOD from the population, calculate the node central density C, and then complete the calculation of each dimension data of the T-table.

[0159] After each iterative update of the population, the node central density will change. Therefore, each time the T-table is updated, only the node central density needs to be recalculated, while the values of the node local similarity and the node degree influence remain unchanged.

[0160] In this embodiment, the update of the Q-table is expressed as:

[0161]

[0162] d i represents the degree of node v i ;

[0163] R(a,b) represents the reward value when node v a selects node v b as its neighbor node;

[0164] V a is the set of nodes in the community network that belong to the same community as node v a ; |V a | represents the number of nodes in set V a ;

[0165] m is the sum of the edges of any two nodes in the community network; L is the preset learning rate; |A| represents the number of individuals in the external shared archive A. ​

[0166] Considering that the community change of a certain node will not only affect the original community where the node is located, but also affect the community that the node is about to join, the present invention proposes a multi-agent reinforcement learning reward strategy, which uses group consensus information to evaluate the behavior of each agent. During the population evolution process, an external shared archive A is used to store the non-dominated solutions generated during the evolution process, and the archive is updated in each generation. The elite solutions stored in the external shared archive A perform well on both dual objectives. Therefore, the actions selected by the solutions in the archive should receive higher rewards. To quantify the contribution of the actions in the external shared archive A, the present invention rewards each action according to its contribution to the improvement of the overall community modularity. The agent a receives a reward R(a, b) for taking the action b, which represents the node v a selects the node v b as the reward value when it is a neighbor node. The parameter L is the learning rate, which represents the influence degree of new evolutionary experience on updating the existing knowledge.

[0167] The present invention proposes an environmental reward strategy based on a consensus feedback mechanism for updating the reinforcement learning strategy. This strategy evaluates the behavior of agents through group consensus information, not only effectively utilizes the community partition information of nodes, but also can provide more accurate feedback for the actions of reinforcement learning, thereby optimizing the learning process of the behavior strategy.

[0168] S6. Calculate MIR gen , and add it to the sliding window W, as Figure 5 shown; where gen represents the current iteration round. Specifically, for each individual in the external shared archive A, the modularity MOD is calculated respectively, that is:

[0169]

[0170] where m is the sum of the edges of any two nodes in the community network; σ(v i , v j ) represents the community relationship between node v i and node v j in the individual, σ(v i , v j ) = 1 means that node v i and node v j belong to the same community, σ(v i , v j ) = 0 means that node v i and node v j do not belong to the same community;

[0171] The calculation of MIR gen is as follows:

[0172]

[0173] Among them, Arc gen represents the average modularity of the external shared archive A at the gen-th iteration; |A| represents the number of individuals in the external shared archive A.

[0174] In the early stage of evolution, since the Q-table has not been established, the agent mainly relies on the guidance of the multi-layer learning strategy (T-table) and focuses on learning the local high-order information of the network. With the interaction with the environment, the agent will gradually turn to using the accumulated historical information and reward mechanism to explore the solution space. In this process, how to balance the learning of the local structure of the network and the utilization of evolutionary knowledge is crucial. The present invention proposes an adaptive control strategy for controlling the transition between the two stages.

[0175] The modularity improvement amount is a direct method to measure the evolutionary state, but the variation range of the modularity improvement is significantly different in different problems and stages. Generally speaking, the modularity improvement is much larger in the early stage of evolution than in the later stage. To address this problem, an adaptive control strategy based on the modularity improvement rate is proposed, which uses a sliding window W to store the value MIR of the modularity improvement rate gen .

[0176] S7. Repeat S3 to S6 until the preset iteration end condition is reached. In this embodiment, the iteration end condition is the custom number of iterations. When the iteration round reaches the preset number of iterations, the iteration ends.

[0177] The present invention adopts a multi-agent reinforcement learning method, and assigns an agent (i.e., T-table and Q-table) to each node. The multi-agent selects a suitable coding label for the current node by using the historical evolutionary information and environmental rewards, so as to provide a promising evolutionary direction for individuals and populations. The action reward of each agent is calculated by the consensus feedback mechanism proposed by the present invention, which feeds back the behavior decision of the agent by using the elite solution information generated by the evolutionary population.

[0178] The final external shared archive A is the detection result of the algorithm, including the Pareto front composed of several non-dominated solutions. The Pareto front is the set of all Pareto optimal solutions in the objective function space of a multi-objective optimization problem. Its core feature is that for any solution on the Pareto front, it is impossible to improve all objective functions simultaneously by adjusting parameters (that is, there is no other solution that is not inferior to it in all objectives and is strictly better in at least one objective).

[0179] Based on the disclosure and teachings of the above specification, those skilled in the art to which the present invention pertains can also make changes and modifications to the above embodiments. Therefore, the present invention is not limited to the specific embodiments disclosed and described above, and some modifications and changes to the present invention should also fall within the protection scope of the claims of the present invention. In addition, although some specific terms are used in this specification, these terms are only for convenience of description and do not constitute any limitation to the present invention.

Claims

1. A multi-objective community detection method assisted by multi-agent reinforcement learning, which is applied to a community network, and the community network includes a plurality of nodes; characterized in that, It includes the following steps: S1. Encode the nodes; initialize the population parameters and randomly initialize multiple individuals for the population; the information of one individual includes the association between each node and other nodes. Construct a sliding window W, where the sliding window W includes a number of values MIR i , where i is an index and i > 0; initialize MIR1 = 1; construct an external shared file A and initialize it as an empty set; initialize a threshold t; S2. Construct two-dimensional tables Q-table and T-table; initialize the values of Q-table to 0; initialize T-table using a multi-layer learning strategy. S3. For each individual, denoted as the parent individual, trigger the selection conditions according to the sliding window W and the threshold t; the conditions include the first condition and the second condition; if the first condition is triggered, based on the T-table, use the roulette wheel strategy to select neighbor nodes for each node of the parent individual; if the second condition is triggered, based on the Q-table, use the ε-greedy strategy to select neighbor nodes for each node of the parent individual. Construct a new individual using the selected neighbor nodes, denoted as the offspring individual. Replace the parent individual with the offspring individual to complete the update of the population. Calculate the fitness value of the offspring individual. S4. Use the fast non-dominated sorting method to select non-dominated solutions from the population and store the non-dominated solutions in the external shared archive A. S5. Update Q-table and T-table. S6. Calculate MIR gen and add it to the sliding window W, where gen represents the current iteration round; S7. Repeat S3 to S6 until the preset iteration end condition is reached.

2. The multi-objective community detection method assisted by multi-agent reinforcement learning according to claim 1, wherein, In S1, the coding representation of the node is v, where v ∈ {1, 2, 3, …, n}; an individual is represented as X = {x1, x2, …, x i , …, x n}; where x i = j means that node v i is connected to node v j , and node v j and node v i are neighbor nodes to each other; The initialization of the individual is to randomly assign neighbor nodes to any one of the nodes; several nodes with connection relationships form a community, and nodes that are not connected to each other belong to different communities.

3. A multi-objective community detection method assisted by multi-agent reinforcement learning according to claim 2, characterized in that The multi-layer learning strategy is to calculate the node local similarity, node degree influence, and node central density according to the community network. The calculation process of the node local similarity S is as follows: Among them, τ(i) represents the set of nodes v i , including the node v i and the neighbor nodes of the node v i ; M represents the adjacency matrix of the community network, M ij = 1 indicates that the node v i and the node v j are connected; M ij = 0 indicates that the node v i and the node v j are not connected; ∪ and ∩ respectively represent taking the union and intersection of two sets; |·| represents the number of nodes in a said set; S ij represents the local similarity between the node v i and the node v j . The calculation process of the node degree influence D is as follows: Among them, d i represents the degree of node v i ; D i represents the node degree influence of node v i ; Select an individual from the population and calculate the node central density C according to the selected individual. The calculation process is as follows: Among them, V i represents the set of nodes in the community to which node v i belongs among the selected individuals; represents node v i in community V i degree within; C i represents the node central density of node v i ​ 4. A multi-objective community detection method assisted by multi-agent reinforcement learning according to claim 3, characterized in that In S2, the initialization process of T-table is expressed as: Among them, NS, ND, and NC respectively represent the normalized local similarity of nodes, the influence of node degrees, and the node central density; α, β, and γ are all weight values; T ij represents the value in the i-th row and j-th column in the T-table; In S5, update T-table using the multi-layer learning strategy.

5. A multi-objective community detection method assisted by multi-agent reinforcement learning according to claim 3, characterized in that, In S3, for node v i , a T ij correspondingly selects a probability of a v j ; when the first condition is triggered, according to the probability, a neighbor node v i is selected for node v j until the offspring individual is generated.

6. A multi-objective community detection method assisted by multi-agent reinforcement learning according to claim 3, characterized in that In S3, when there is any MIR in the sliding window W i greater than the threshold t, the first condition is triggered; When all MIRs in the sliding window W i are less than the threshold t, the second condition is triggered.

7. A multi-objective community detection method assisted by multi-agent reinforcement learning according to claim 6, characterized in that For each individual in the external shared archive A, calculate the modularity MOD respectively, that is: where m is the sum of the edges of any two nodes in the community network; σ(v i ,v j ) represents the community relationship between nodes v i and v j in the individual. σ(v i ,v j ) = 1 indicates that nodes v i and v j belong to the same community, and σ(v i ,v j ) = 0 indicates that nodes v i and v j do not belong to the same community; In S6, the calculation of MIR gen is as follows: Among them, Arc gen represents the average modularity of the external shared archive A at the gen-th iteration; |A| represents the number of individuals in the external shared archive A; The sliding window W is a first-in-first-out queue with a fixed length of w.

8. A multi-objective community detection method assisted by multi-agent reinforcement learning according to claim 2, characterized in that, In S5, the update of Q-table is expressed as: Among them, M represents the adjacency matrix of the community network, M ij = 1 indicates that node v i and node v j are connected; M ij = 0 indicates that node v i and node v j are not connected; d i represents the degree of node v i . R(a,b) represents the reward value when node v a selects node v b as a neighbor node; V a For the community network, the set of nodes that belong to the same community as node v a ; |V a | represents the number of nodes in the set V a ; m is the sum of the edges of any two nodes in the community network; L is the preset learning rate; |A| represents the number of individuals in the external shared archive A.

9. A multi-objective community detection method assisted by multi-agent reinforcement learning according to claim 1, characterized in that When the second condition is triggered, the process of generating the offspring individual is as follows: where ε is a preset adaptive parameter; rand is a random value; Q(v i ,A c ) represents the Q value obtained when node v i executes action A c and forms a neighbor node with node v j ; Argmax[Q(v i ,A c )] represents selecting action A c to maximize the Q value; v r represents randomly selecting a neighbor node for node v i .

10. A multi-objective community detection method assisted by multi-agent reinforcement learning according to claim 2, characterized in that, In S3, the offspring individuals are divided into q communities, namely {V1, V2, …, V k , …, V q}, where V k represents the set of nodes within the k-th community of the offspring individuals; The fitness value includes f1 and f2. The calculation process is as follows: Among them, M represents the adjacency matrix of the community network, M ij = 1 indicates that node v i and node v j are connected; M ij = 0 indicates that node v i and node v j are not connected; V represents the set of all nodes in the community network; |·| represents the size of the set; L(V k , V k ) represents the internal connection strength of community V k ; represents the external connection strength of community V k .