Community detection method and device, equipment and storage medium
By obtaining optimization parameters and queue processing based on node importance in the Leiden algorithm, combined with local movement and refinement processing of multiple community division indicators, the complexity and efficiency of parameter adjustment during community detection of Leiden algorithm is solved, and the comprehensive effect of community detection is improved.
Patent Information
- Application Number
- CN202510042855.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-27
AI Technical Summary
The existing Leiden algorithm has problems such as high parameter adjustment complexity and low efficiency of fast moving nodes during community detection, resulting in poor overall results.
A community detection method is proposed, by obtaining the optimization parameters in the target Leiden algorithm, adding nodes to the preset queue according to the comprehensive importance of nodes, and locally moving nodes based on multiple community division indicators, partition refinement and network aggregation are performed until it meets the iteration conditions.
The efficiency of local node movement and the efficiency of Leiden algorithm parameter generation are improved, thereby improving the comprehensive effect of community detection.
Smart Images

Figure CN120045956A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a community detection method, device, equipment and storage medium. Background Art
[0002] In complex networks, community detection plays an indispensable role. It is a technique used to reveal the aggregation behavior of networks, which can mine hidden correlation relationships and structures. Community detection is actually a method of network clustering, and one of its key roles is to extract useful information from the network.
[0003] Currently, when performing community detection, there are various common community detection methods, such as hierarchical clustering algorithms, spectral clustering algorithms, density-based clustering algorithms, Leiden algorithms, etc. Among them, the Leiden algorithm is an algorithm for detecting communities in large networks. It greedily optimizes modularity, recursively merges communities into a single node, and repeats this process in the compressed graph. The Leiden algorithm takes the modularity of the network as its core driving force, follows a greedy strategy, and continuously iteratively moves nodes to different communities to improve modularity, thus has become a leader in the field of community detection research.
[0004] Although the Leiden algorithm already has good community detection effects, there are still some limitations, including the high complexity of parameter adjustment, the low efficiency of nodes forming queues in a random order when quickly moving nodes, etc. Therefore, the existing Leiden algorithm has problems with poor comprehensive effects when performing community detection. Summary of the Invention
[0005] This application aims to at least solve the technical problems existing in the prior art. To this end, a first aspect of this application proposes a community detection method, which includes:
[0006] Obtain a target Leiden algorithm; wherein, the target Leiden algorithm includes a variety of optimization parameters, and each optimization parameter is generated based on a preset parameter prediction model;
[0007] When performing community detection based on the target Leiden algorithm, add each node to a preset queue according to the comprehensive importance corresponding to each node, traverse all nodes through the preset queue, and perform local movement of nodes based on a variety of preset community division indicators to obtain a first community division result; wherein, the variety of preset community division indicators include the modularity, compactness, intra-community homogeneity and the reciprocal of community separation of the community;
[0008] Perform partition refinement processing on the first community division result to obtain a second community division result;
[0009] Based on the first community division result and the second community division result, perform network aggregation processing to obtain the current iteration result, and repeat the process of generating the current iteration result until the preset iteration condition is met. Take the current iteration result generated last time as the target community detection result.
[0010] In a possible implementation manner, adding each node to a preset queue according to the comprehensive importance corresponding to each node includes:
[0011] For each node, obtain the degree centrality, betweenness centrality, and eigenvector centrality corresponding to the node;
[0012] Based on the degree centrality, betweenness centrality, and eigenvector centrality, calculate the comprehensive importance corresponding to the node;
[0013] Based on the comprehensive importance corresponding to each node, add each node to a preset queue.
[0014] In a possible implementation manner, calculating the comprehensive importance corresponding to a node based on the degree centrality, betweenness centrality, and eigenvector centrality includes:
[0015] Obtain the first preset weight corresponding to the degree centrality, the second preset weight corresponding to the betweenness centrality, and the third preset weight corresponding to the eigenvector centrality;
[0016] Based on the first preset weight, the second preset weight, and the third preset weight, perform weighted summation processing on the degree centrality, betweenness centrality, and eigenvector centrality to obtain the comprehensive importance corresponding to the node.
[0017] In a possible implementation manner, performing node local movement based on multiple preset community division indicators to obtain the first community division result includes:
[0018] Obtain the preset index weights corresponding to the respective preset community division indicators;
[0019] Based on the respective preset index weights and the respective preset community division indicators, calculate the target comprehensive index;
[0020] When performing node local movement, perform community division on each node based on the target comprehensive index and the preset community division threshold to obtain the first community division result.
[0021] In a possible implementation manner, the optimization parameters include the optimized community selection degree, the optimized resolution, the optimized internal connection tightness, the optimized community quality threshold, the optimized community quality approximation threshold, and the optimized balance coefficient.
[0022] In a possible implementation manner, the construction of the preset parameter prediction model includes:
[0023] Obtain training sample data; wherein, the training sample data includes various feature data of the historical network structure, Leiden algorithm parameter data corresponding to the feature data, and community detection results.
[0024] Train the initial parameter prediction model based on the training sample data to generate a preset parameter prediction model.
[0025] In a possible implementation manner, training the initial parameter prediction model based on the training sample data to generate a preset parameter prediction model includes:
[0026] Train the initial parameter prediction model based on the training sample data to obtain an intermediate parameter prediction model.
[0027] Evaluate the intermediate parameter prediction model using a preset evaluation index, adjust the model parameters of the intermediate parameter prediction model according to the evaluation results, and generate a preset parameter prediction model based on the adjusted model parameters.
[0028] A second aspect of the present application proposes a community detection device, and the device includes:
[0029] An acquisition module, configured to acquire a target Leiden algorithm; wherein, the target Leiden algorithm includes various optimization parameters, and each optimization parameter is generated based on a preset parameter prediction model.
[0030] A first partitioning module, configured to, when performing community detection based on the target Leiden algorithm, add each node to a preset queue according to the comprehensive importance corresponding to each node, traverse all nodes through the preset queue, and perform local movement of the nodes based on multiple preset community partitioning indexes to obtain a first community partitioning result; wherein, the multiple preset community partitioning indexes include the modularity, compactness, intra-community homogeneity, and the reciprocal of the community separation degree of the community.
[0031] A second partitioning module, configured to perform partition refinement processing on the first community partitioning result to obtain a second community partitioning result.
[0032] A processing module, configured to perform network aggregation processing based on the first community partitioning result and the second community partitioning result to obtain a current iteration result, and repeat the process of generating the current iteration result until a preset iteration condition is met, and use the current iteration result generated last as the target community detection result.
[0033] In a possible implementation manner, the above first partitioning module is specifically configured to:
[0034] Adding each node to a preset queue according to the comprehensive importance corresponding to each node includes:
[0035] For each node, obtain the degree centrality, betweenness centrality, and eigenvector centrality corresponding to the node;
[0036] Based on the degree centrality, betweenness centrality, and eigenvector centrality, calculate the comprehensive importance corresponding to the node;
[0037] Based on the comprehensive importance corresponding to each node, add each node to a preset queue.
[0038] In a possible implementation manner, the above-mentioned first partitioning module is further configured to:
[0039] Obtain a first preset weight corresponding to the degree centrality, a second preset weight corresponding to the betweenness centrality, and a third preset weight corresponding to the eigenvector centrality;
[0040] Perform a weighted summation process on the degree centrality, betweenness centrality, and eigenvector centrality based on the first preset weight, the second preset weight, and the third preset weight to obtain the comprehensive importance corresponding to the node.
[0041] In a possible implementation manner, the above-mentioned first partitioning module is further configured to:
[0042] Obtain preset index weights corresponding to each preset community partitioning index;
[0043] Based on each preset index weight and each preset community partitioning index, calculate a target comprehensive index;
[0044] When performing node local movement, perform community partitioning on each node based on the target comprehensive index and a preset community partitioning threshold to obtain a first community partitioning result.
[0045] In a possible implementation manner, the optimization parameters include the optimized community selection degree, the optimized resolution, the optimized internal connection tightness, the optimized community quality threshold, the optimized community quality approximation threshold, and the optimized balance coefficient.
[0046] In a possible implementation manner, the above-mentioned community detection device is further configured to:
[0047] Obtain training sample data; wherein, the training sample data includes various feature data of the historical network structure, Leiden algorithm parameter data corresponding to the feature data, and community detection results;
[0048] Train an initial parameter prediction model based on the training sample data to generate a preset parameter prediction model.
[0049] In a possible implementation manner, the above-mentioned community detection device is further configured to:
[0050] Train an initial parameter prediction model based on training sample data to obtain an intermediate parameter prediction model;
[0051] Evaluate the intermediate parameter prediction model using a preset evaluation metric, adjust the model parameters of the intermediate parameter prediction model according to the evaluation results, and generate a preset parameter prediction model based on the adjusted model parameters.
[0052] A third aspect of this application proposes an electronic device, which includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the community detection method as described in the first aspect.
[0053] A fourth aspect of this application proposes a computer-readable storage medium, in which at least one instruction, at least one program, a code set, or an instruction set is stored, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the community detection method as described in the first aspect.
[0054] The embodiments of this application have the following beneficial effects:
[0055] The community detection method provided by the embodiments of this application includes: obtaining a target Leiden algorithm. When performing community detection based on the target Leiden algorithm, each node is added to a preset queue according to the comprehensive importance corresponding to each node, all nodes are traversed through the preset queue, and node local movement is performed based on multiple preset community division metrics to obtain a first community division result. The first community division result is subjected to partition refinement processing to obtain a second community division result. Network aggregation processing is performed based on the first community division result and the second community division result to obtain the current iteration result, and the process of generating the current iteration result is repeated until a preset iteration condition is met. The last generated current iteration result is used as the target community detection result. In this solution, each node is added to a preset queue according to the comprehensive importance corresponding to each node, and multiple preset community division metrics are considered during node local movement instead of only considering the modularity metric, which improves the efficiency of node local movement. In addition, by automatically generating optimization parameters based on a preset parameter prediction model, the efficiency of generating Leiden algorithm parameters is improved, thereby improving the efficiency of performing community detection and further improving the comprehensive effect during community detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is a block diagram of a computer device provided by an embodiment of this application;
[0057] Figure 2The flowchart of steps of a community detection method provided by an embodiment of this application;
[0058] Figure 3 The flowchart of steps of a preset parameter prediction model construction method provided by an embodiment of this application;
[0059] Figure 4 The flowchart of steps of a preset parameter prediction model generation method provided by an embodiment of this application;
[0060] Figure 5 The flowchart of steps of adding each node to a preset queue provided by an embodiment of this application;
[0061] Figure 6 The flowchart of steps of calculating comprehensive importance provided by an embodiment of this application;
[0062] Figure 7 The flowchart of steps of obtaining a first community division result provided by an embodiment of this application;
[0063] Figure 8 The structural block diagram of a community detection device provided by an embodiment of this application. Detailed implementation manners
[0064] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0065] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of this disclosure, unless otherwise stated, the meaning of "a plurality" is two or more. Additionally, the use of "based on" or "according to" is meant to be open and inclusive, since a process, step, calculation, or other action "based on" or "according to" one or more of the stated conditions or values may in practice be based on additional conditions or values beyond those stated.
[0066] The community detection method provided by this application can be applied to a computer device (electronic device). The computer device can be a server or a terminal. Among them, the server can be a single server or a server cluster composed of multiple servers. The embodiments of this application do not make specific limitations in this regard. The terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices.
[0067] Taking the computer device as a server as an example, Figure 1 A block diagram of a server is shown, as Figure 1 shown, the server may include a processor and a memory connected by a system bus. Among them, the processor of the server is used to provide computing and control capabilities. The memory of the server includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. When the computer program is executed by the processor, it is used to implement a community detection method.
[0068] Those skilled in the art can understand that, Figure 1 the structure shown in
[0069] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the server to which the solution of this application is applied. Optionally, the server may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0070] Figure 2 is a step flow chart of a community detection method provided by an embodiment of this application. As Figure 2 shown, the method includes the following steps:
[0071] Step 202, obtain the target Leiden algorithm.
[0072] Among them, the node local movement strategy in the traditional Leiden algorithm uses a queue to implement the access to the changed nodes. This strategy effectively reduces unnecessary node access and improves the efficiency of the algorithm to a certain extent, but there is still room for further optimization. For example, the computational efficiency problem, that is, when dealing with a very large network, significant computational resources are still required; the complexity problem of parameter settings, that is, it requires the support of experienced engineers and continuous debugging; the problem of local optimum rather than global optimum, that is, Leiden ensures the connectivity of the community through local optimization, and may find a local optimum but not a global optimum, etc.
[0073] Therefore, by improving the traditional Leiden algorithm, it is necessary to first obtain the target Leiden algorithm, which includes various optimization parameters. Each optimization parameter is generated based on a preset parameter prediction model. Optionally, the optimization parameters may include the optimized community selection degree, the optimized resolution, the optimized internal connection tightness, the optimized community quality threshold, the optimized community quality approximation threshold, and the optimized balance coefficient.
[0074] Among them, by adjusting the community selection degree parameter θ, it can control the degree of randomness. The selection of the θ value has an important impact on the exploration ability of the algorithm and the quality of the final community division. Different θ values can be tested on different datasets through methods such as cross-validation to find the optimal parameter settings. The Leiden algorithm introduces the parameter θ to control the degree of randomly selecting communities, which allows users to trade off between the randomness and determinacy of the algorithm according to their needs.
[0075] The resolution parameter p has an important impact on the coarseness and fineness of the community. By adaptively adjusting this parameter, community structures of different granularities can be found at different stages. For example, a larger resolution value is used in the initial division to discover large community structures, and the resolution value is gradually decreased in the refinement stage to identify more detailed community patterns.
[0076] In the Leiden algorithm, the internal connection tightness parameter σ controls the tightness of the internal connections within the community. By adaptively adjusting this parameter, while ensuring the internal connectivity of the community, the problem of excessive communities caused by over-segmentation can be avoided. During the iterative process of the algorithm, the value of this parameter can be dynamically adjusted according to the stability of the community and the change of modularity.
[0077] At each step of the algorithm, the quality of the current community division should be evaluated. If the community has reached a certain quality standard, that is, it is necessary to preset the community quality threshold parameter q, and further refinement will not significantly improve the modularity, then the algorithm can be terminated in advance, thus saving computing resources.
[0078] The community quality approximation threshold can be pre-customized with a default value and can be denoted as s.
[0079] The balance coefficient r is the ratio of the comprehensive quality to the efficiency. The relationship between the balance coefficient r and other thresholds is r = f(θ, p, σ, q, s, a), where a is the balance correction parameter. The role of a is to balance various hyperparameters, which is generally composed of a relatively complex polynomial, and its value is obtained through calculation.
[0080] In some optional embodiments, such as Figure 3 shown Figure 3The flowchart of steps for constructing a preset parameter prediction model provided by an embodiment of this application includes:
[0081] Step 302, obtain training sample data.
[0082] Step 304, train an initial parameter prediction model based on the training sample data to generate a preset parameter prediction model.
[0083] Among them, the training sample data includes various feature data of historical network structures, Leiden algorithm parameter data corresponding to the feature data, and community detection results. These data can come from multiple different historical network structures to ensure the generalization ability of the model.
[0084] Then, feature engineering can be performed based on the feature data of the historical network structure to extract key features that affect the community detection performance. The key features can include global attributes and local attributes of the network. Optionally, the global attributes can include the average path length, diameter, etc., and the local attributes can include the local clustering coefficient of nodes, the distribution of neighbor nodes, etc.
[0085] It is also necessary to select a suitable machine learning model for parameter prediction. The machine learning model can include decision trees, random forests, gradient boosting machines, support vector machines, or neural networks, etc. The selection of the model depends on the characteristics of the data and the complexity of the problem.
[0086] In some optional embodiments, as Figure 4 shown, Figure 4 The flowchart of steps for generating a preset parameter prediction model provided by an embodiment of this application includes:
[0087] Step 402, train an initial parameter prediction model based on the training sample data to obtain an intermediate parameter prediction model.
[0088] Step 404, evaluate the intermediate parameter prediction model using a preset evaluation index, adjust the model parameters of the intermediate parameter prediction model according to the evaluation results, and generate a preset parameter prediction model based on the adjusted model parameters.
[0089] Among them, when training an initial parameter prediction model based on the training sample data to obtain an intermediate parameter prediction model, the hyperparameters of the model, such as the learning rate, regularization parameter, etc., need to be adjusted to obtain the best prediction performance. The prediction performance of the model can also be evaluated by methods such as cross-validation.
[0090] Next, a preset evaluation index can be used to evaluate the intermediate parameter prediction model. According to the evaluation results, the model parameters of the intermediate parameter prediction model are adjusted, and a preset parameter prediction model is generated based on the adjusted model parameters. Among them, the preset evaluation index can include accuracy, F1 value, etc., and can also include indexes such as retrieval efficiency and retrieval effect. Each index can be determined according to the scenario in different business scenarios. Compared with the traditional process of determining thresholds and evaluation indexes based on experience and other algorithms, the algorithm uses a machine learning-based method to learn and determine thresholds and evaluation indexes, with high accuracy and intelligence.
[0091] Step 204, when performing community detection based on the target Leiden algorithm, each node is added to a preset queue according to the comprehensive importance corresponding to each node. All nodes are traversed through the preset queue, and local movement of nodes is performed based on a variety of preset community division indexes to obtain a first community division result.
[0092] Among them, when performing community detection based on the target Leiden algorithm, each node can be added to a preset queue according to the comprehensive importance corresponding to each node first. In some optional embodiments, as Figure 5 shown Figure 5 FIG. is a flowchart of steps for adding each node to a preset queue provided by an embodiment of the present application, including:
[0093] Step 502, for each node, obtain the degree centrality, betweenness centrality, and eigenvector centrality corresponding to the node.
[0094] Step 504, based on the degree centrality, betweenness centrality, and eigenvector centrality, calculate the comprehensive importance corresponding to the node.
[0095] Step 506, based on the comprehensive importance corresponding to each node, add each node to a preset queue.
[0096] Among them, the degree centrality indicates that the greater the degree of a node, the more important this node is, and the degree centrality of each node is denoted as c1. The betweenness centrality is an index used to describe the number of shortest paths passing through a certain node to characterize the importance of the node, and the betweenness centrality of each node is denoted as c2. The eigenvector centrality considers the importance of its neighbor nodes, that is, the importance of a node depends not only on the number of its neighbor nodes, that is, the degree centrality of this node, but also on the importance of its neighbor nodes, and the eigenvector importance of each node is denoted as c3.
[0097] Thus, the comprehensive importance corresponding to the node can be calculated based on the degree centrality, betweenness centrality, and eigenvector centrality. In some optional embodiments, as Figure 6 shown Figure 6A flowchart of steps for calculating the comprehensive importance provided by an embodiment of this application includes:
[0098] Step 602: Obtain the first preset weight corresponding to degree centrality, the second preset weight corresponding to betweenness centrality, and the third preset weight corresponding to eigenvector centrality.
[0099] Step 604: Perform a weighted summation process on degree centrality, betweenness centrality, and eigenvector centrality based on the first preset weight, the second preset weight, and the third preset weight to obtain the comprehensive importance corresponding to the node.
[0100] Among them, the first preset weight corresponding to degree centrality can be denoted as wn1, the second preset weight corresponding to betweenness centrality can be denoted as wn2, and the third preset weight corresponding to eigenvector centrality can be denoted as wn3. wn1, wn2, and wn3 are hyperparameters of the algorithm and can be dynamically set according to different application scenarios and network environments.
[0101] Thus, the node importance c0 can be calculated by c0 = softmax(c1) * wn1 + softmax(c2) * wn2 + softmax(c3) * wn3, where softmax() represents the normalization function.
[0102] Finally, based on the comprehensive importance corresponding to each node, each node can be added to the preset queue. Optionally, the nodes with greater comprehensive importance can be added to the preset queue first.
[0103] Optionally, in some cases, such as when the network is huge, the time limit requirement is high but not particularly accurate, etc., an approximation algorithm can be used to quickly estimate the community structure, which can significantly improve the calculation speed. The approximation algorithm can include first setting an approximation threshold, that is, setting an acceptable approximation threshold, and then setting an approximation strategy, that is, according to the approximation threshold, selecting different approximation algorithms or mechanisms. For example, when calculating the community modularity, select a function with lower complexity, reduce the accuracy of the participating technical data, etc.
[0104] Then, all nodes can be traversed through the preset queue, and node local movement can be performed based on multiple preset community division metrics to obtain the first community division result. In some optional embodiments, as Figure 7 shown Figure 7 A flowchart of steps for obtaining the first community division result provided by an embodiment of this application includes:
[0105] Step 702: Obtain the preset metric weights corresponding to each preset community division metric.
[0106] Step 704: Calculate the target comprehensive metric based on each preset metric weight and each preset community division metric.
[0107] Step 706: When performing local node movement, partition each node based on the target comprehensive metric and a preset community partition threshold to obtain a first community partition result.
[0108] Among them, the preset community partition metric can be denoted as target_com. Multiple preset community partition metrics can include the modularity, compactness, intra-community homogeneity, and the reciprocal of the community separation degree of the community. These metrics can be obtained using existing metric definitions and calculation methods, which will not be elaborated here.
[0109] And obtain the preset metric weights corresponding to each preset community partition metric, which can be denoted as target_w. The preset metric weights are hyperparameters of the algorithm and can be dynamically set according to different application scenarios and network environments.
[0110] Thus, based on each preset metric weight and each preset community partition metric, calculate the target comprehensive metric. The target comprehensive metric can be denoted as target_score, and its calculation method is target_score = ∑target_com(i) * target_w(i), where target_com(i) represents the i-th preset community partition metric and target_w(i) represents the i-th preset metric weight.
[0111] Next, when performing local node movement, partition each node based on the target comprehensive metric and a preset community partition threshold to obtain a first community partition result.
[0112] Step 206: Perform partition refinement processing on the first community partition result to obtain a second community partition result.
[0113] Among them, after obtaining the first community partition result, partition refinement processing can be performed on the first community partition result to ensure that all communities have good connectivity, thereby obtaining a second community partition result. The specific process of performing refinement processing can refer to the existing process, which will not be elaborated here.
[0114] Step 208: Perform network aggregation processing based on the first community partition result and the second community partition result to obtain the current iteration result, and repeat the process of generating the current iteration result above until the preset iteration condition is met. Take the current iteration result generated last time as the target community detection result.
[0115] Among them, when performing network aggregation processing, the current iteration result is created based on the second community partition result, but the community partition in the current iteration result is obtained based on the first community partition result. After completing the current iteration, another iteration can be performed, that is, repeat the process of generating the current iteration result above until the preset iteration condition is met, and use the current iteration result generated last as the target community detection result.
[0116] This application provides a community detection method, which includes: obtaining a target Leiden algorithm. When performing community detection based on the target Leiden algorithm, each node is added to a preset queue according to the comprehensive importance corresponding to each node, all nodes are traversed through the preset queue, and node local movement is performed based on multiple preset community partition metrics to obtain a first community partition result. The first community partition result is subjected to partition refinement processing to obtain a second community partition result. Network aggregation processing is performed based on the first community partition result and the second community partition result to obtain the current iteration result, and the process of generating the current iteration result above is repeated until the preset iteration condition is met, and the current iteration result generated last is used as the target community detection result. In this solution, each node is added to the preset queue according to the comprehensive importance corresponding to each node, and multiple preset community partition metrics are considered during node local movement instead of only considering the modularity metric, which improves the efficiency of node local movement. In addition, by automatically generating optimization parameters based on a preset parameter prediction model, the efficiency of generating Leiden algorithm parameters is improved, thereby improving the efficiency of community detection and further improving the comprehensive effect during community detection.
[0117] Figure 8 It is a structural block diagram of a community detection device provided by an embodiment of this application.
[0118] As Figure 8 shown, the community detection device 800 includes:
[0119] An acquisition module 802, configured to acquire a target Leiden algorithm; among them, the target Leiden algorithm includes multiple optimization parameters, and each optimization parameter is generated based on a preset parameter prediction model.
[0120] A first partitioning module 804, configured to, when performing community detection based on the target Leiden algorithm, add each node to a preset queue according to the comprehensive importance corresponding to each node, traverse all nodes through the preset queue, and perform node local movement based on multiple preset community partition metrics to obtain a first community partition result; among them, the multiple preset community partition metrics include the modularity, compactness, internal homogeneity of the community, and the reciprocal of the community separation degree of the community.
[0121] The second partitioning module 806 is configured to perform partition refinement processing on the first community partitioning result to obtain a second community partitioning result.
[0122] The processing module 808 is configured to perform network aggregation processing based on the first community partitioning result and the second community partitioning result to obtain a current iteration result, and repeat the process of generating the current iteration result until a preset iteration condition is met, and use the last generated current iteration result as the target community detection result.
[0123] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein. Each module in the above community detection device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form so that the processor can call and execute the operations of the above modules.
[0124] In an embodiment of the present application, a computer device is provided. The computer device includes a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:
[0125] Obtain a target Leiden algorithm; wherein, the target Leiden algorithm includes a variety of optimization parameters, and each optimization parameter is generated based on a preset parameter prediction model;
[0126] When performing community detection based on the target Leiden algorithm, add each node to a preset queue according to the comprehensive importance corresponding to each node, traverse all nodes through the preset queue, and perform local movement of nodes based on a variety of preset community partitioning metrics to obtain a first community partitioning result; wherein, the variety of preset community partitioning metrics include the modularity, compactness, intra-community homogeneity, and the reciprocal of the community separation degree of the community;
[0127] Perform partition refinement processing on the first community partitioning result to obtain a second community partitioning result;
[0128] Perform network aggregation processing based on the first community partitioning result and the second community partitioning result to obtain a current iteration result, and repeat the process of generating the current iteration result until a preset iteration condition is met, and use the last generated current iteration result as the target community detection result.
[0129] In an embodiment of the present application, when the processor executes the computer program, the following steps are further implemented:
[0130] For each node, obtain the degree centrality, betweenness centrality, and eigenvector centrality corresponding to the node;
[0131] Calculate the comprehensive importance corresponding to the node based on degree centrality, betweenness centrality, and eigenvector centrality;
[0132] Add each node to a preset queue based on the comprehensive importance corresponding to each node.
[0133] In an embodiment of the present application, when the processor executes the computer program, the following steps are further implemented:
[0134] Obtain the first preset weight corresponding to degree centrality, the second preset weight corresponding to betweenness centrality, and the third preset weight corresponding to eigenvector centrality;
[0135] Perform a weighted summation process on degree centrality, betweenness centrality, and eigenvector centrality based on the first preset weight, the second preset weight, and the third preset weight to obtain the comprehensive importance corresponding to the node.
[0136] In an embodiment of the present application, when the processor executes the computer program, the following steps are further implemented:
[0137] Obtain the preset index weights corresponding to the respective preset community division indicators;
[0138] Calculate the target comprehensive index based on the respective preset index weights and the respective preset community division indicators;
[0139] When performing node local movement, perform community division on each node based on the target comprehensive index and a preset community division threshold to obtain a first community division result.
[0140] In an embodiment of the present application, the optimization parameters include the optimized community selection degree, the optimized resolution, the optimized internal connection tightness, the optimized community quality threshold, the optimized community quality approximation threshold, and the optimized balance coefficient.
[0141] In an embodiment of the present application, when the processor executes the computer program, the following steps are further implemented:
[0142] Obtain training sample data; wherein, the training sample data includes various feature data of the historical network structure, Leiden algorithm parameter data corresponding to the feature data, and community detection results;
[0143] Train the initial parameter prediction model based on the training sample data to generate a preset parameter prediction model.
[0144] In an embodiment of the present application, when the processor executes the computer program, the following steps are further implemented:
[0145] Train the initial parameter prediction model based on the training sample data to obtain an intermediate parameter prediction model;
[0146] Evaluate the intermediate parameter prediction model using preset evaluation metrics, adjust the model parameters of the intermediate parameter prediction model according to the evaluation results, and generate a preset parameter prediction model based on the adjusted model parameters.
[0147] The computer device provided by the embodiments of the present application has a similar implementation principle and technical effect to the above method embodiments, which will not be elaborated here.
[0148] In an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0149] Obtain the target Leiden algorithm; wherein, the target Leiden algorithm includes a variety of optimization parameters, and each optimization parameter is generated based on a preset parameter prediction model;
[0150] When performing community detection based on the target Leiden algorithm, add each node to a preset queue according to the comprehensive importance corresponding to each node, traverse all nodes through the preset queue, and perform local movement of nodes based on a variety of preset community division metrics to obtain a first community division result; wherein, the variety of preset community division metrics includes the modularity, compactness, intra-community homogeneity, and the reciprocal of the community separation degree of the community;
[0151] Perform partition refinement processing on the first community division result to obtain a second community division result;
[0152] Perform network aggregation processing based on the first community division result and the second community division result to obtain the current iteration result, and repeat the process of generating the current iteration result until the preset iteration condition is satisfied, and use the last generated current iteration result as the target community detection result.
[0153] In an embodiment of the present application, when the computer program is executed by a processor, the following steps are also implemented:
[0154] For each node, obtain the degree centrality, betweenness centrality, and eigenvector centrality corresponding to the node;
[0155] Calculate the comprehensive importance corresponding to the node based on the degree centrality, betweenness centrality, and eigenvector centrality;
[0156] Add each node to the preset queue based on the comprehensive importance corresponding to each node.
[0157] In an embodiment of the present application, when the computer program is executed by a processor, the following steps are also implemented:
[0158] Obtain the first preset weight corresponding to degree centrality, the second preset weight corresponding to betweenness centrality, and the third preset weight corresponding to eigenvector centrality;
[0159] Perform weighted summation processing on degree centrality, betweenness centrality, and eigenvector centrality based on the first preset weight, the second preset weight, and the third preset weight to obtain the comprehensive importance corresponding to the node.
[0160] In an embodiment of the present application, when the computer program is executed by a processor, the following steps are further implemented:
[0161] Obtain the preset index weights corresponding to the respective preset community division indicators;
[0162] Calculate the target comprehensive index based on the respective preset index weights and the respective preset community division indicators;
[0163] When performing node local movement, perform community division on each node based on the target comprehensive index and the preset community division threshold to obtain the first community division result.
[0164] In an embodiment of the present application, the optimization parameters include the optimized degree of community selection, the optimized resolution, the optimized internal connection tightness, the optimized community quality threshold, the optimized community quality approximation threshold, and the optimized balance coefficient.
[0165] In an embodiment of the present application, when the computer program is executed by a processor, the following steps are further implemented:
[0166] Obtain training sample data; wherein, the training sample data includes various feature data of the historical network structure, Leiden algorithm parameter data corresponding to the feature data, and community detection results;
[0167] Train the initial parameter prediction model based on the training sample data to generate a preset parameter prediction model.
[0168] In an embodiment of the present application, when the computer program is executed by a processor, the following steps are further implemented:
[0169] Train the initial parameter prediction model based on the training sample data to obtain an intermediate parameter prediction model;
[0170] Evaluate the intermediate parameter prediction model using a preset evaluation index, adjust the model parameters of the intermediate parameter prediction model according to the evaluation results, and generate a preset parameter prediction model based on the adjusted model parameters.
[0171] The computer-readable storage medium provided in this embodiment has the same implementation principle and technical effects as the above method embodiment, and will not be elaborated here.
[0172] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. This computer program can be stored in a non-volatile computer-readable storage medium. When this computer program is executed, it can include the processes of the embodiments of the above various methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0173] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0174] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A community detection method, characterized in that: The method comprises: Obtaining a target Leiden algorithm; wherein the target Leiden algorithm includes a plurality of optimization parameters, each of which is generated based on a preset parameter prediction model; When performing community detection based on the target Leiden algorithm, each node is added to a preset queue according to the comprehensive importance corresponding to each node, all nodes are traversed through the preset queue, and nodes are locally moved based on multiple preset community division indicators to obtain a first community division result; wherein the multiple preset community division indicators include community modularity, compactness, internal homogeneity of the community, and the inverse of community separation; Performing partition refinement processing on the first community division result to obtain a second community division result; Based on the first community division result and the second community division result, network aggregation processing is performed to obtain the current iteration result, and the above process of generating the current iteration result is repeated until the preset iteration condition is met, and the current iteration result generated for the last time is used as the target community detection result.
2. The method according to claim 1, characterized in that The adding each node to a preset queue according to the comprehensive importance corresponding to each node includes: For each node, obtain the degree centrality, betweenness centrality and eigenvector centrality corresponding to the node; Based on the degree centrality, the betweenness centrality and the eigenvector centrality, calculating the comprehensive importance corresponding to the node; Based on the comprehensive importance corresponding to each node, each node is added to the preset queue.
3. The method according to claim 2, characterized in that The calculating the comprehensive importance corresponding to the node based on the degree centrality, the betweenness centrality and the eigenvector centrality includes: Obtaining a first preset weight corresponding to the degree centrality, a second preset weight corresponding to the betweenness centrality, and a third preset weight corresponding to the eigenvector centrality; Based on the first preset weight, the second preset weight, and the third preset weight, a weighted summation process is performed on the degree centrality, the betweenness centrality, and the eigenvector centrality to obtain a comprehensive importance corresponding to the node.
4. The method according to any one of claims 1 to 3, characterized in that: The node is locally moved based on a plurality of preset community division indicators to obtain a first community division result, including: Obtaining preset indicator weights corresponding to each of the preset community division indicators; Calculate the target comprehensive index based on the preset index weights and the preset community division indexes; When performing local node movement, each node is divided into communities based on the target comprehensive index and a preset community division threshold to obtain the first community division result.
5. The method according to any one of claims 1 to 3, characterized in that: The optimization parameters include the optimized community selection degree, the optimized resolution, the optimized internal connection density, the optimized community quality threshold, the optimized community quality approximate threshold, and the optimized balance coefficient.
6. The method according to any one of claims 1 to 3, characterized in that: The construction of the preset parameter prediction model includes: Acquire training sample data; wherein the training sample data includes multiple feature data of the historical network structure, Leiden algorithm parameter data corresponding to the feature data, and community detection results; The initial parameter prediction model is trained based on the training sample data to generate the preset parameter prediction model.
7. The method according to claim 6, characterized in that The training of the initial parameter prediction model based on the training sample data to generate the preset parameter prediction model includes: Training the initial parameter prediction model based on the training sample data to obtain an intermediate parameter prediction model; The intermediate parameter prediction model is evaluated using a preset evaluation index, the model parameters of the intermediate parameter prediction model are adjusted according to the evaluation result, and the preset parameter prediction model is generated based on the adjusted model parameters.
8. A community detection device, characterized in that: The device comprises: An acquisition module, used for acquiring a target Leiden algorithm; wherein the target Leiden algorithm includes a plurality of optimization parameters, each of which is generated based on a preset parameter prediction model; A first partitioning module is used to add each node to a preset queue according to the comprehensive importance corresponding to each node when performing community detection based on the target Leiden algorithm, traverse all nodes through the preset queue, and perform local node movement based on multiple preset community partitioning indicators to obtain a first community partitioning result; wherein the multiple preset community partitioning indicators include community modularity, compactness, internal homogeneity of the community, and the inverse of community separation; A second division module is used to perform a partition refinement process on the first community division result to obtain a second community division result; A processing module is used to perform network aggregation processing based on the first community division result and the second community division result to obtain a current iteration result, and repeat the above process of generating the current iteration result until a preset iteration condition is met, and the current iteration result generated for the last time is used as the target community detection result.
9. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the community detection method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the community detection method according to any one of claims 1 to 7.