Multi-target influence maximization method based on cost constraint in hypergraph
By introducing independent cascading IC models and multi-objective optimization functions into the hypergraph model, combined with the NSGA-II framework, the balance problem of diffusion and cost in the maximization of influence in the hypergraph is solved, and a better Pareto solution set is generated, suitable for rapid decision-making of social networks and the Internet of Things.
Patent Information
- Application Number
- CN202510678050.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-05-26
AI Technical Summary
The existing hypergraph model has failed to effectively balance the information diffusion effect and seed selection costs in the research on maximizing influence, making it difficult to obtain the best solution in resource-constrained scenarios.
Using an independent cascading IC model as the basis, combining multi-objective optimization functions and NSGA-II framework, a diversified Pareto cutting-edge solution set is generated through a three-mode initialization strategy, a fast non-dominant sorting algorithm, a two-point crossover operator and an adaptive mutation operator, to reduce Monte Carlo simulation dependence and improve evaluation efficiency.
It realizes the simultaneous optimization of affecting diffusion and seed costs in the hypergraph, and generates a more uniform Pareto solution set, which significantly improves the economicality of diffusion effect and seed selection, and is suitable for the rapid execution of social platforms and the Internet of Things.
Smart Images

Figure CN120524820A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of social network analysis, and in particular to a multi-objective influence maximization method based on cost constraints in a hypergraph. Background Art
[0002] Complex networks, a fundamental tool for modeling complex systems, have garnered increasing research and attention in recent years. Social networks, a typical example of complex networks, have attracted widespread attention due to their wide application in information dissemination, product marketing, and other fields. The influence maximization (IM) problem is typically defined as selecting a set of seed nodes within a network to maximize the spread of information. Traditional IM research primarily uses binary graphs to represent networks, which only capture point-to-point relationships. In contrast, in scenarios such as collaborative networks, transportation systems, and social groups, information often propagates in higher-order "one-to-many" or "many-to-many" fashions. Hypergraphs allow a single hyperedge to connect any number of nodes simultaneously, enabling more accurate modeling of higher-order interactions. This makes them a powerful tool for studying cascading propagation and identifying key nodes. In practical applications such as product recommendations and news delivery, it is crucial to maximize the diffusion effect while limiting the "seed selection cost" (e.g., the operational investment required for the number of groups a node participates in). Existing hypergraph IM research often focuses on a single objective (diffusion) without systematically considering cost factors. In resource-constrained settings, balancing these two objectives becomes a key challenge. To simulate the cascade process in high-order systems, studies usually use three types of models: SI (Susceptible-Infected model), LT (Linear Threshold model), and IC (Independent Cascade model). Among them, the IC model considers the bidirectional propagation probability of "node→hyperedge" and "hyperedge→node" at the same time, and has been proven to more accurately characterize the fault / information diffusion mechanism in the hypergraph, and therefore serves as the modeling basis for the present invention. The complex dual-objective conflict makes the problem difficult to solve using analytical methods. The evolutionary multi-objective optimization (EMO) algorithm has been successfully applied in the field of IM of graph models (such as NSGA-II) due to its advantages such as no need for differentiability assumptions and strong search capabilities, providing a feasible path for solving dual-objective problems in hypergraph scenarios.
[0003] In existing research, the technical solutions most similar to the present invention can be roughly divided into the following categories: First, Hao et al. proposed a fairness-aware influence maximization method based on an evolutionary multi-objective algorithm. By introducing a new definition of "fairness," they incorporated prior knowledge into specific regions of the Pareto front (PF), thereby improving the quality of the solution in this region. Fabían et al. reformulated the IM problem as an extreme adversarial problem of "maximizing the seed set and minimizing the diffusion capacity" and solved it using a swarm intelligence metaheuristic method. Perrault et al. introduced a budget framework constrained by total advertising cost for online IM and designed a new algorithm based on the IC model and edge-level semi-bandit feedback. Although theoretical analysis and experimental verification have been provided, many of the new algorithms' assumptions are difficult to verify in practical applications. The second category is single-objective heuristic algorithms based on hypergraphs, including the High Hyper Degree Adaptation algorithm (HHDA), eigenvector centrality, the hypergraph collective influence algorithm (HCI), and k-core decomposition. Xie et al. proposed the High-Degree Adaptation (HHDA) method, Kovalenko et al. extended the eigenvector centrality of graphs to hypergraphs, and Lee et al. analyzed the percolation characteristics of hypergraphs through (k, q)-core decomposition and self-consistent equations.
[0004] Objective influence maximization methods for ordinary graphs are difficult to directly extend to hypergraph scenarios because they only support binary edge relationships. They also suffer from drawbacks such as uneven distribution of solutions on the PF and slow convergence. Single-objective heuristic algorithms for hypergraphs, while leveraging the high-order structure of the hypergraph to identify key nodes, focus solely on diffusion capacity, disregarding seed selection costs, and exhibit poor adaptability to weak nodes or highly heterogeneous structures. While these methods have made some progress in hypergraph IM research, they generally suffer from limitations such as neglect of weak nodes, poor adaptability, and crude influence ranking due to their strongly coupled structures and complex cascading dynamics. Furthermore, these studies focus on a single objective—influence diffusion—while ignoring the cost of seed selection, rendering them ineffective in multi-objective scenarios. Therefore, this paper introduces two conflicting objectives, "diffusion" and "selection cost," into the hypergraph model and investigates them within the NSGA-II framework. Summary of the Invention
[0005] In view of the above shortcomings of the prior art, the present invention provides a multi-objective influence maximization method based on cost constraints in a hypergraph.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: A cost-constrained multi-objective influence maximization method in a hypergraph includes the following: 1. Select the independent cascade IC model as the basic model; Second, use a multi-objective optimization function to evaluate the dual objectives; 3. Three-mode initialization strategy: Generate diverse initial individuals in different areas of the Pareto frontier PF through the HCI-based initialization module, the unit collective influence-based initialization module, and the random initialization module: 4. Operator Selection 4.1. Use the fast non-dominated sorting algorithm to divide the current population into several non-dominated levels according to the dominance relationship. Then, screen individuals based on the non-dominated level and crowding distance to construct the next generation population. 4.2. Generate offspring using two-point crossover operator; 4.3. When a mutation occurs, random mutation, HCI mutation or UCI mutation is performed randomly; 4.4. Introduce a mechanism for adaptively adjusting crossover and mutation rates to adjust crossover and mutation rates based on the current frontier diversity; 5. Output the Pareto frontier solution set.
[0007] Furthermore, the independent cascade IC model is selected as the basic model. Assume that the hypergraph H(V, E) has N nodes and M hyperedges; where V represents the node set and E represents the hyperedge set; the association between nodes and hyperedges can be described by the association matrix H: when node i is associated with a hyperedge When there is a correlation, ;otherwise, ; The number of hyperedges associated with node i is denoted as ki, which is called the hyperdegree of the node; The number of associated nodes is recorded as , which is the cardinality of the hyperedge.
[0008] Furthermore, the independent cascade IC model is used to characterize the cascading failure process in high-order systems: when node i fails, it will fail with probability The hyperedge that causes its association Failure; on the contrary, when the super edge When it fails, it will return to the state with probability This causes the associated node i to fail.
[0009] Furthermore, the multi-objective optimization function is defined as follows:
[0010] Where σ0(S) represents the number of seed nodes, σ1(S) represents the expected number of activations of one-hop neighbor nodes excluding the seed set S, and σ2(S) represents the expected number of activations of two-hop neighbor nodes.
[0011] Among them, |S| represents the number of seed nodes, Indicates that node j is a one-hop neighbor of seed set S, N S(1) represents the set of one-hop neighbor nodes of the seed set S; Indicates the Super edge, Represents a hyperedge Contains node j, N S (2) represents the set of two-hop neighbor nodes of the seed set S, indicates that node k is a two-hop neighbor, but neither a one-hop neighbor nor a seed node, Represents the hyperedge e μ Contains two-hop neighbor node k, Indicates that any seed node to the hyperedge The probability of propagation, Represents a hyperedge The number of seed nodes that intersect with the seed set S; Indicates that on the hyperedge There are When there are seed nodes, the probability of the hyperedge successfully propagating to the one-hop neighbor node j; q μk Represents a hyperedge The number of active nodes that intersect with the one-hop neighbor set, p eμ Represents a hyperedge The probability of being activated, p eμ q μk Indicates that on the hyperedge There are When there are already activated nodes, the probability of the hyperedge successfully propagating to the two-hop neighbor node k;
[0012] in, Indicates that the one-hop neighbor node j also belongs to the hyperedge e μ , p j represents the probability of node j being activated, s jeμ Indicates that node j is activated via hyperedge e μ The success probability of propagating to the two-hop neighbor;
[0013] in, Represents a hyperedge A one-hop neighbor hyperedge belonging to the seed set S and containing node j, Represents a hyperedge The probability of being activated, Indicates the activation of the hyperedge The success probability of propagating to one-hop neighbor node j;
[0014] in, Represents the hyperedge to seed node i The probability of propagating and activating this hyperedge, Indicates that node i belongs to the seed set S.
[0015] Furthermore, the HCI-based initialization module is used to calculate the first-level collective influence HCI value of all nodes in the hypergraph:
[0016] in, - represents the collective influence of node i under the first-level communication (ICM), k i represents the hyperdegree of node i (the number of hyperedges connected to i), Represents a hyperedge The cardinality (with The number of connected nodes), ∂ i represents the set of hyperedges associated with node i, r i Represents a hyperedge to node i The probability of propagating and activating the hyperedge; Subsequently, the HCI value of each node is multiplied by a random number between [0, 1], and then sorted in descending order, and the first n nodes are taken to form the initial individual.
[0017] Furthermore, the unit collective influence-based initialization module is used to define the unit cost collective influence as:
[0018] Then multiply the UCI value of each node by a random number in the interval [0, 1] and sort them in descending order, and select the first n nodes to generate individuals; The random initialization module is used to randomly select n different nodes from the entire hypergraph to form individuals.
[0019] Furthermore, a two-point crossover operator is used to generate offspring. This operator randomly selects two crossover points on two parent individuals and exchanges these two gene fragments to generate two offspring with better quality. If duplicate nodes appear in the offspring after crossover, they are randomly replaced with nodes that have not been used by the individual.
[0020] Furthermore, random mutation: each gene in each individual mutates with a probability of 1 / len and can be replaced by any node that has not appeared before; if a duplicate node appears after mutation, it is randomly replaced by a node that has not yet been included in the individual; HCI mutation: Calculate the first-level HCI value of the node corresponding to each gene in the individual and multiply this value by a random number in the interval [0, 1]. After all genes have obtained HCI values, select the gene position with the smallest HCI as the mutation point. Randomly select n nodes from the remaining nodes, calculate their first-level HCI values, and replace the gene at the mutation point with the node with the largest HCI. UCI mutation: Calculate the UCI value of each gene in the individual and multiply the value by a random number in the interval [0, 1]; select the gene position with the smallest UCI as the mutation point; randomly select n nodes from the remaining nodes, calculate their UCI values, and replace the gene at the mutation point with the node with the largest UCI.
[0021] Furthermore, given the Pareto frontier PF, its diversity is defined as:
[0022] Among them, x i represents the seed set corresponding to the i-th solution in the Pareto frontier PF, x j represents the seed set corresponding to the j-th solution; Initially, let the initial crossover rate C init =0.8, initial mutation rate M init =0.2; when population diversity decreases, the mutation rate M is dynamically increased p To enhance exploration while reducing the crossover rate C p ; Set the minimum crossover rate C min =0.6, maximum mutation rate M max =0.4, the crossover rate and mutation rate of each generation are calculated as follows: .
[0023] Furthermore, it also includes S6, result evaluation: Monte Carlo simulation performs accurate influence evaluation on the Pareto front solution set obtained by the above method: for each candidate seed set S, 10,000 independent IC cascade propagation simulations are performed, and the average number of activated nodes in each simulation is calculated as the expected influence σ(S) of the solution, and this is used as the final influence score to sort the solution set within the Pareto front and select the optimal solution.
[0024] Compared with the existing technology, the present invention has the following beneficial effects: 1. Dual-objective modeling: For the first time, the present invention simultaneously incorporates the two conflicting objectives of "influence diffusion" and "seed cost" into the hypergraph IM, forming a multi-objective optimization task, which fundamentally makes up for the shortcomings of single-objective research.
[0025] 2. Use low-complexity evaluation function: The present invention proposes an approximate calculation of the impact based on the two-hop local structure, which significantly reduces the dependence on Monte Carlo simulation and improves the evaluation efficiency.
[0026] 3. This paper adopts an improved NSGA-II framework for hypergraphs (EMOIM-HG): first, three-mode (HCI / UCI / random) initialization, covering high-impact-high-cost, unit-impact-priority, and low-cost areas, respectively, to ensure the initial diversity of the population; second, a crossover-mutation operator based on diversity adaptive adjustment, combined with dedicated UCI / HCI mutation, avoids falling into local optimality and accelerates convergence.
[0027] 4. The present invention is user-friendly and easy to distribute or use for edge computing. Evaluation does not require multiple rounds of full-graph simulation. It can be quickly executed locally on social platforms or the Internet of Things, reducing communication and computing power overhead.
[0028] 5. The present invention has been verified through experiments that, on synthetic and real hypergraph datasets, the Pareto solution sets generated by the method of the present invention are more uniform and have greater diffusion when the cost is the same, which is significantly better than existing single-objective centrality algorithms and general multi-objective frameworks. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.
[0030] Figure 1 This is a schematic diagram of a multi-objective influence maximization method based on cost constraints in a hypergraph of the present invention.
[0031] Figure 2 Results of cost-influence experiments on six simulated hypergraphs.
[0032] Figure 3 Results of cost-influence experiments on four real hypergraphs.
[0033] Figure 4 is the IGD (non-intergenerational distance) of each generation of the three algorithms.
[0034] Figure 5 Schematic diagram of the three-mode initialization strategy.
[0035] Figure 6 Schematic diagram of the mutation operator.
[0036] Figure 7 Schematic diagram of the two-point crossover operator. DETAILED DESCRIPTION
[0037] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0038] Example 1: In some embodiments, please refer to the accompanying drawings of the specification. Figure 1 , a cost-constrained multi-objective influence maximization method in hypergraphs, named EMOIM-HG, includes the following contents: First, in a hypergraph, the influence maximization (IM) problem aims to select a fixed number n of seed nodes to maximize the influence spread range under a given propagation model. This problem is often considered a single-objective optimization problem and has been proven to be NP-hard. Minimizing the cost of seed node selection is also crucial in many practical applications. For example, in the news dissemination scenario, maximizing influence spread while reducing cost is a typical multi-objective optimization problem, whose goal is to maximize the spread impact while minimizing the seed selection cost. This single-objective optimization problem can be formalized as:
[0039] Where S is the set of seed nodes, σ(S) represents the expected impact propagation scale under the selected seed node set S, and c i represents the selection cost of node i, and C(S) represents the total selection cost of the selected seed set S. In the present invention, the selection cost of a seed node can be quantified by its superdegree (the number of hyperedges to which the node belongs). The larger the superdegree, the more groups the node participates in, and the higher its selection cost.
[0040] 2. Model Selection: Due to the strong correlation and coupling between nodes and hyperedges in hypergraphs during information propagation, this paper introduces the Independent Cascade (IC) model to describe the information propagation rules in hypergraphs. This helps EMOIM-HG accurately measure the collective influence of multiple nodes.
[0041] Suppose a hypergraph H(V, E) has N nodes and M hyperedges. Where V represents the node set and E represents the hyperedge set. The association between nodes and hyperedges can be described by the association matrix H: when node i is associated with a hyperedge When there is a correlation, ;otherwise, The number of hyperedges associated with node i is denoted as ki, which is called the hyperdegree of the node; The number of associated nodes is recorded as , which is the cardinality of the hyperedge. In a hypergraph, a node can be regarded as an “agent”, and a hyperedge is equivalent to a functional module composed of multiple agents. The IC model is used to characterize the cascading failure process in a high-order system: when a node (agent) i fails, it will fail with probability Hyperedges (functional modules) that lead to their association Failure; on the contrary, when the super edge When it fails, it will return to the state with probability This causes the associated node i to fail.
[0042] 3. Design a low-complexity impact evaluation function to evaluate dual objectives: A low-complexity, high-precision multi-objective optimization function can improve the efficiency of the algorithm and guide the population to converge towards a better solution. In the independent cascade (IC) model, in order to estimate the expected influence propagation, a large number of Monte Carlo simulations are usually required to take into account the randomness of the propagation process, resulting in high computational complexity. To address this problem, the present invention proposes a new multi-objective optimization function that uses local structural information to evaluate the collective influence of a seed set. Given the significance of this function in practical applications, its value is not normalized in this paper. The multi-objective optimization function consists of the following two objectives: the first objective uses the local structural information of the two-hop neighborhood in the hypergraph to approximate the influence of the seed set; the second objective evaluates the seed selection cost based on the node's hyperdegree ki.
[0043] The multi-objective optimization function is defined as follows:
[0044] Where σ0(S) represents the number of seed nodes, σ1(S) represents the expected number of activations of one-hop neighbor nodes excluding the seed set S, and σ2(S) represents the expected number of activations of two-hop neighbor nodes.
[0045] Among them, |S| represents the number of seed nodes, Indicates that node j is a one-hop neighbor of seed set S (excluding the seed node itself), N S (1) Represents the set of one-hop neighbor nodes of the seed set S (excluding the seed set itself). Indicates the Super edge, Represents a hyperedge Contains node j, N S (2) represents the set of two-hop neighbor nodes of the seed set S (excluding one-hop neighbors and the seed set), indicates that node k is a two-hop neighbor, but neither a one-hop neighbor nor a seed node, Represents the hyperedge e μ Contains two-hop neighbor node k, Indicates that any seed node to the hyperedge The probability of propagation, Represents a hyperedge The number of seed nodes that intersect with the seed set S (used when affecting one-hop neighbor j); Indicates that on the hyperedge There are When there are seed nodes, the probability of the hyperedge successfully propagating (activating) to the one-hop neighbor node j. μk Represents a hyperedge The number of activated nodes that intersect with the one-hop neighbor set (used when affecting the two-hop neighbor k), p eμ Represents a hyperedge The probability of being activated, p eμ q μk Indicates that on the hyperedge There are The probability that the hyperedge will successfully propagate to the two-hop neighbor node k when there are already activated nodes.
[0046]
[0047] in, Indicates that the one-hop neighbor node j also belongs to the hyperedge e μ , p j represents the probability of node j being activated, s jeμ Indicates that node j is activated via hyperedge e μ The success probability of propagating to the two-hop neighbor;
[0048] in, Represents a hyperedge A one-hop neighbor hyperedge belonging to the seed set S and containing node j, Represents a hyperedge The probability of being activated, Indicates the activation of the hyperedge The success probability of propagating to one-hop neighbor node j;
[0049] in, Represents the hyperedge to seed node i The probability of propagating and activating this hyperedge, Indicates that node i belongs to the seed set S.
[0050] The present invention adopts a low-complexity impact assessment function and introduces a two-hop local structure approximation. Under the IC model, only the 0 / 1 / 2-hop neighborhood probability of the seed node is used in a closed form (without global Monte Carlo) to calculate σ(S), and the complexity is reduced to O(|S|d 2 ).
[0051] 4. Three-mode initialization strategy ( Figure 5 In EMOIM-HG, a candidate solution represents a seed set, represented by a vector of length nnn. Each gene corresponds to the index of a node in the hypergraph, and nodes are guaranteed to be unique within the same vector. A high-quality initial population generally accelerates convergence and improves algorithm performance. To ensure a uniform distribution of solutions on the Pareto front (PF), EMOIM-HG generates diverse initial individuals in different regions of the PF using three modules: 4.1. HCI-based initialization module, used to calculate the first-level collective influence (HCI) value of all nodes in the hypergraph:
[0052] in, - represents the collective influence of node i under the first-level communication (ICM), k i represents the hyperdegree of node i (the number of hyperedges connected to i), Represents a hyperedge The cardinality (with The number of connected nodes), ∂ i represents the set of hyperedges associated with node i, r i Represents a hyperedge to node i The probability of propagating and activating the hyperedge; Subsequently, the HCI value of each node is multiplied by a random number between [0, 1], and then sorted in descending order, and the first n (n is the number of seeds to be selected) nodes are taken to form the initial individual.
[0053] This module aims to generate solutions with high influence spread but relatively high seed selection cost.
[0054] 4.2. Initialization module based on unit collective influence (UCI), used to define unit cost collective influence as:
[0055] Then, the UCI value of each node is multiplied by a random number in the interval [0, 1] and sorted in descending order, and the first n nodes are selected to generate individuals.
[0056] The solutions generated by this module have a moderate seed selection cost while maintaining a high influence spread, thus achieving a balance between influence and cost.
[0057] 4.3. Random initialization module is used to randomly select n different nodes from the entire hypergraph to form individuals.
[0058] This module can generate solutions with low influence diffusion and seed cost to enhance population diversity.
[0059] This method uses a three-module initialization strategy: the HCI module prioritizes the generation of high-impact, high-cost individuals, the UCI module generates high-unit benefit (diffusion / cost) individuals, and the random module covers low-cost areas. This can improve convergence speed and solution uniformity.
[0060] 5. Operator Selection 5.1. Use the Fast Non-dominated Sorting algorithm to partition the current population into several non-dominated fronts based on dominance relationships. Then, select individuals based on their non-dominated fronts and crowding distance to construct the next generation of the population, stopping after a given number of iterations. Non-dominated sorting and crowding distance calculations are performed after evaluating the dual objectives.
[0061] The fast non-dominated rank sorting method is as follows: for each individual, calculate the number of individuals that dominate it and the set of individuals it dominates. Individuals that are not dominated by any individual are assigned to the first non-dominated rank. All individuals in the first rank are removed from the population, and the above process is repeated for the remaining individuals, resulting in the second and third rank, until all individuals have been assigned a rank.
[0062] The crowding distance measures the density of individual solutions in the target space. A large Euclidean distance between an individual and its adjacent solutions (sorted by target values) indicates a sparse solution region; a small distance indicates a concentrated solution region. During the selection phase, populations are prioritized from low to high non-dominated levels. If incorporating all individuals from a particular level would exceed the population size, individuals within that level are sorted in descending order of crowding distance, and those with the largest crowding distances are selected for the next generation, ensuring a diverse and even distribution of solutions.
[0063] 5.2, such as Figure 7 As shown in Figure 2, the two-point crossover operator is used to generate offspring. This operator randomly selects two crossover points on two parent individuals and swaps these two gene segments to produce two offspring of higher quality. If duplicate nodes appear in the offspring after the crossover, they are randomly replaced with nodes that have not been used by the individual.
[0064] 5.3, such as Figure 6 As shown, a mutation operator is selected when a mutation occurs. To prevent the algorithm from falling into a local optimum, the present invention designs a new mutation operator consisting of three modules: random mutation, HCI mutation, and UCI mutation. When a mutation occurs, one of the modules is randomly selected to execute: Random mutation: Each gene in each individual mutates with a probability of 1 / len(individual) and can be replaced by any node that has not appeared before; if a duplicate node appears after the mutation, it is randomly replaced by a node that has not yet been included in the individual.
[0065] HCI mutation: Calculate the first-level HCI value of the node corresponding to each gene in the individual and multiply the value by a random number in the interval [0, 1]. After all genes have obtained the HCI values, select the gene position with the smallest HCI as the mutation point to increase the probability of generating a low-influence node. Randomly select n nodes from the remaining nodes, calculate their first-level HCI values, and replace the gene at the mutation point with the node with the largest HCI.
[0066] UCI mutation: Calculate the UCI value of each gene in the individual and multiply it by a random number in the interval [0, 1]. Select the gene position with the smallest UCI as the mutation point to increase the probability of generating a node with low unit influence. Randomly select n nodes from the remaining nodes, calculate their UCI values, and replace the gene at the mutation point with the node with the largest UCI.
[0067] The above mutation operator can guide the population to evolve in a direction that achieves a better balance between influence diffusion and seed selection cost while ensuring search diversity.
[0068] 5.4. Adaptive Adjustment of Crossover Rate and Mutation Rate Mechanism: To further prevent falling into local optimality, EMOIM-HG introduces an adaptive probability adjustment strategy based on the diversity of individual genotypes on the Pareto front (PF). Given a Pareto front PF, its diversity is defined as:
[0069] Among them, x i represents the seed set corresponding to the i-th solution in the Pareto frontier PF, x j represents the seed set corresponding to the j-th solution; Initially, let the initial crossover rate C init =0.8, initial mutation rate M init =0.2. When population diversity decreases, the mutation rate M is dynamically increased. p To enhance exploration while reducing the crossover rate C p (and satisfy M p +C p =1). In order to take into account the convergence, the minimum crossover rate C is set min =0.6, maximum mutation rate M max =0.4, the crossover rate and mutation rate of each generation are calculated as follows:
[0070] The present invention adopts an adaptive evolution operator and uses three operators, random / UCI / HCI, to dynamically select according to individual scores to improve the probability of jumping out of the local optimum. The crossover rate C is adjusted based on the current frontier diversity D. p , mutation rate M p , avoid premature puberty or shock.
[0071] 6. Results Evaluation In the EMOIM-HG framework, Monte Carlo simulation is used to accurately assess the influence of the Pareto front solution set obtained by the above method: for each candidate seed set S, 10,000 independent IC cascade propagation simulations are performed, and the average number of activated nodes in each simulation is calculated as the expected influence σ(S) of the solution. This is used as the final influence score to sort the solution set within the Pareto front and select the optimal solution.
[0072] The present invention performs Monte Carlo simulation, calculates the impact range and finds the Pareto front, and takes the Pareto front as the final result.
[0073] Figure 2 As can be seen, on six synthetic and four real hypergraphs, EMOIM-HG outputs over 600 non-dominated solutions, covering high, medium, and low cost zones, while traditional centrality or single-objective algorithms only produce single-point solutions. This demonstrates that the proposed method balances high influence diffusion with low seed cost, providing decision makers with a complete set of Pareto compromise solutions.
[0074] Figure 3 It can be seen that the diffusion-cost advantage of the present invention exceeds 25% in both restaurant-reviews (commercial) and geometry-questions (scientific research) datasets, and it has high generalization and scalability across scenarios.
[0075] Figure 4 It can be seen that the proposed method accelerates convergence and makes the frontier distribution more uniform. The IGD (inverse generation distance) index decreases by 40–60% on average in ER, SF, and UF networks, and approaches the final frontier within 10 generations, while the randomly initialized NSGA-II requires >30 generations.
[0076] The present invention has the following advantages: Dual-objective modeling: This invention is the first to incorporate the two conflicting objectives of "influence diffusion" and "seed cost" into the hypergraph IM at the same time, forming a multi-objective optimization task, fundamentally making up for the shortcomings of single-objective research.
[0077] Use low-complexity evaluation functions: This invention proposes an approximate calculation of the impact based on the two-hop local structure, which significantly reduces the dependence on Monte Carlo simulation and improves the evaluation efficiency.
[0078] This paper adopts an improved NSGA-II framework for hypergraphs (EMOIM-HG): first, three-mode (HCI / UCI / random) initialization, covering high-impact-high-cost, unit-impact-priority, and low-cost areas, respectively, to ensure the initial diversity of the population; second, a crossover-mutation operator based on diversity adaptive adjustment, combined with dedicated UCI / HCI mutation, avoids falling into local optimality and accelerates convergence.
[0079] The present invention is deployment-friendly and easy to distribute or edge compute. Evaluation does not require multiple rounds of full-graph simulation and can be quickly executed locally on social platforms or the Internet of Things, reducing communication and computing power overhead.
[0080] The present invention has been experimentally verified to have a more uniform Pareto solution set generated by the method on synthetic and real hypergraph datasets, with a larger diffusion at the same cost, which is significantly better than existing single-objective centrality algorithms and general multi-objective frameworks.
[0081] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A multi-objective influence maximization method based on cost constraints in a hypergraph, characterized by: Includes the following: S1. Select the independent cascade IC model as the basic model; S2, using a multi-objective optimization function to evaluate dual objectives; S3, three-mode initialization strategy: Generate diverse initial individuals in different areas of the Pareto frontier PF through the HCI-based initialization module, the unit collective influence-based initialization module, and the random initialization module: S4. Operator selection: S4.
1. Use the fast non-dominated sorting algorithm to divide the current population into several non-dominated levels according to the dominance relationship. Then, select individuals based on the non-dominated level and crowding distance to construct the next generation population. S4.2, using the two-point crossover operator to generate offspring; S4.3, when a mutation occurs, random mutation, HCI mutation, or UCI mutation is performed randomly; S4.
4. Introduce a mechanism for adaptively adjusting crossover and mutation rates to adjust crossover and mutation rates based on the current frontier diversity; S5. Output the Pareto frontier solution set.
2. The multi-objective influence maximization method based on cost constraints in a hypergraph according to claim 1 is characterized in that: The independent cascade IC model is selected as the basic model. Assume that the hypergraph H(V, E) has N nodes and M hyperedges. V represents the node set and E represents the hyperedge set. The association between nodes and hyperedges can be described by the association matrix H: when node i is connected to hyperedge When there is a correlation, ;otherwise, ; The number of hyperedges associated with node i is denoted as ki, which is called the hyperdegree of the node; The number of associated nodes is recorded as , which is the cardinality of the hyperedge.
3. The multi-objective influence maximization method based on cost constraints in a hypergraph according to claim 2 is characterized in that: The independent cascade IC model is used to characterize the cascading failure process in high-order systems: when node i fails, it will fail with probability The hyperedge that causes its association Failure; on the contrary, when the super edge When it fails, it will return to the state with probability This causes the associated node i to fail.
4. The multi-objective influence maximization method based on cost constraints in a hypergraph according to claim 3 is characterized in that: The multi-objective optimization function is defined as follows: ; Where σ0(S) represents the number of seed nodes, σ1(S) represents the expected number of activations of one-hop neighbor nodes excluding the seed set S, and σ2(S) represents the expected number of activations of two-hop neighbor nodes. ; Among them, |S| represents the number of seed nodes, Indicates that node j is a one-hop neighbor of seed set S, N S (1) represents the set of one-hop neighbor nodes of the seed set S; Indicates the Super edge, Represents a hyperedge Contains node j, N S (2) represents the set of two-hop neighbor nodes of the seed set S, indicates that node k is a two-hop neighbor, but neither a one-hop neighbor nor a seed node, Represents the hyperedge e μ Contains two-hop neighbor node k, Indicates that any seed node to the hyperedge The probability of propagation, Represents a hyperedge The number of seed nodes that intersect with the seed set S; Indicates that on the hyperedge There are When there are seed nodes, the probability of the hyperedge successfully propagating to the one-hop neighbor node j; q μk Represents a hyperedge The number of active nodes that intersect with the one-hop neighbor set, p eμ Represents a hyperedge The probability of being activated, p eμ q μk Indicates that on the hyperedge There are When there are already activated nodes, the probability of the hyperedge successfully propagating to the two-hop neighbor node k; ; in, Indicates that the one-hop neighbor node j also belongs to the hyperedge e μ , p j represents the probability of node j being activated, s jeμ Indicates that node j is activated via hyperedge e μ The success probability of propagating to the two-hop neighbor; ; in, Represents a hyperedge A one-hop neighbor hyperedge belonging to the seed set S and containing node j, Represents a hyperedge The probability of being activated, Indicates the activation of the hyperedge The success probability of propagating to one-hop neighbor node j; ; in, Represents the hyperedge to seed node i The probability of propagating and activating this hyperedge, Indicates that node i belongs to the seed set S.
5. The multi-objective influence maximization method based on cost constraints in a hypergraph according to claim 4 is characterized in that: The HCI-based initialization module is used to calculate the first-level collective influence HCI value of all nodes in the hypergraph: ; in, - represents the collective influence of node i under the first-level propagation ICM, k i represents the excess degree of node i, Represents a hyperedge The cardinality of ∂ i represents the set of hyperedges associated with node i, r i Represents a hyperedge to node i The probability of propagating and activating the hyperedge; Subsequently, the HCI value of each node is multiplied by a random number between [0, 1], and then sorted in descending order, and the first n nodes are taken to form the initial individual.
6. The multi-objective influence maximization method based on cost constraints in a hypergraph according to claim 5, characterized in that: The unit collective influence-based initialization module is used to define the unit cost collective influence as: ; Then multiply the UCI value of each node by a random number in the interval [0, 1] and sort them in descending order, and select the first n nodes to generate individuals; The random initialization module is used to randomly select n different nodes from the entire hypergraph to form individuals.
7. The multi-objective influence maximization method based on cost constraints in a hypergraph according to claim 6, characterized in that: The two-point crossover operator is used to generate offspring. The operator randomly selects two crossover points on two parent individuals and exchanges the two gene segments to generate two offspring with better quality. If duplicate nodes appear in the offspring after crossover, they are randomly replaced with nodes that have not been used by the individual.
8. The multi-objective influence maximization method based on cost constraints in a hypergraph according to claim 7, characterized in that: Random mutation: Each gene in each individual mutates with a probability of 1 / len and can be replaced by any node that has not appeared before; if a duplicate node appears after the mutation, it is randomly replaced by a node that has not yet been included in the individual; HCI mutation: Calculate the first-level HCI value of the node corresponding to each gene in the individual and multiply this value by a random number in the interval [0, 1]. After all genes have obtained HCI values, select the gene position with the smallest HCI as the mutation point. Randomly select n nodes from the remaining nodes, calculate their first-level HCI values, and replace the gene at the mutation point with the node with the largest HCI. UCI mutation: Calculate the UCI value of each gene in the individual and multiply the value by a random number in the interval [0, 1]; select the gene position with the smallest UCI as the mutation point; randomly select n nodes from the remaining nodes, calculate their UCI values, and replace the gene at the mutation point with the node with the largest UCI.
9. The multi-objective influence maximization method based on cost constraints in a hypergraph according to claim 8, characterized in that: Given a Pareto frontier PF, its diversity is defined as: ; Among them, x i represents the seed set corresponding to the i-th solution in the Pareto frontier PF, x j represents the seed set corresponding to the j-th solution; Initially, let the initial crossover rate C init =0.8, initial mutation rate M init =0.2; when population diversity decreases, the mutation rate M is dynamically increased p To enhance exploration while reducing the crossover rate C p ; Set the minimum crossover rate C min =0.6, maximum mutation rate M max =0.4, the crossover rate and mutation rate of each generation are calculated as follows: 。 10. The multi-objective influence maximization method based on cost constraints in a hypergraph according to claim 9, characterized in that: It also includes S6, result evaluation: Monte Carlo simulation performs accurate influence evaluation on the Pareto front solution set obtained by the above method: for each candidate seed set S, 10,000 independent IC cascade propagation simulations are performed, and the average number of activated nodes in each simulation is calculated as the expected influence σ(S) of the solution, which is then used as the final influence score to sort the solution set within the Pareto front and select the optimal solution.
Citation Information
Patent Citations
Dynamic consistency maintenance method for copies in mobile edge computing
CN112612422A
Hypergraph influence propagation method and influence maximization method based on conformity psychology
CN113269652A
Node influence maximization method based on hypergraph
CN114691938A
Method for solving seed node with maximum influence based on mixture of hypergraph and common graph
CN116894740A
Marketing network optimal influence customer group identification method based on hypergraph threshold model
CN117575678A