Fitness terrain-driven hyper-heuristic influence maximization method and system
Patent Information
- Application Number
- CN202611180876.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-05
- Publication Date
- 2026-09-25
AI Technical Summary
传统贪心类方法的核心过程依赖大量蒙特卡洛模拟评估节点边际增益,计算开销随网络规模呈非线性剧增,面临严重的可扩展性瓶颈;而启发式方法仅依赖局部网络拓扑,未能充分量化种子节点间的传播重叠效应,导致求解精度受限
(1)搜索策略的多样性与调度灵活性增强。通过将多类扰动器与选择器进行正交组合,构建了丰富的原子操作库,实现了高层调度与低层执行的解耦。这种解耦机制使算法能够根据演化阶段的需求,灵活地从操作库中选取并组合原子操作,克服了传统元启发式方法中搜索模式固化的局限,在搜索效率与解质量之间实现了更优的平衡。
Smart Images

Figure CN122820198A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of social network analysis and intelligent optimization technology, specifically relating to a fitness terrain-driven hyperheuristic influence maximization method and system. Background Technology
[0002] The Influence Maximization (IM) problem aims to select the optimal seed set in complex social networks to maximize the reach of information under a specific propagation model. Existing solutions mainly encompass approximate acceleration methods, heuristic methods, and metaheuristic methods. Among these, metaheuristic optimization methods have become an important paradigm for solving this problem due to their ability to achieve a certain balance between computational efficiency and solution quality. However, facing the high-dimensional, discrete, multi-modal, and significantly overlapping IM search space, existing methods still generally suffer from the following drawbacks: (1) There is an inherent contradiction between solution accuracy and computational efficiency. The core process of traditional greedy methods relies on a large number of Monte Carlo simulations to evaluate the marginal gain of nodes. The computational cost increases nonlinearly with the network size, which faces a serious scalability bottleneck. On the other hand, heuristic methods rely only on the local network topology and fail to fully quantify the propagation overlap effect between seed nodes, resulting in limited solution accuracy. Neither of them can balance evaluation accuracy and solution efficiency with limited resources.
[0003] (2) Existing metaheuristic solutions suffer from a lack of understanding of the discrete solution space, which can easily lead to stagnation in local optima and degradation of global exploration capabilities. This approach treats the discrete combinatorial solution space as a "black box" lacking prior structural information, and cannot actively perceive the macroscopic skewness characteristics and microscopic search behavior of the population fitness distribution. This lack of quantitative information on evolutionary states results in a lack of clear directional guidance in the search process, and computational resources are easily wasted in low-potential regions. As evolution progresses, the population is prone to over-aggregating in local high-fitness regions, losing the ability to continue exploring better regions, ultimately causing the obtained seed set to deviate from the global optimum.
[0004] (3) Multi-strategy metaheuristic scheduling mechanisms lack dynamic adaptability, which can easily lead to a mismatch between search behavior and evolutionary requirements. Multi-strategy metaheuristic solutions improve solution quality by introducing multiple search operators, but their strategy switching logic is usually based on empirical settings, fixed periods, or random probabilities, failing to systematically learn the cooperative patterns between different strategies from historical effective search trajectories. When facing complex network structures, such rigid scheduling methods are difficult to achieve effective cooperation between strategies, seriously affecting the global optimization robustness and convergence quality of the algorithm. Summary of the Invention
[0005] The purpose of this invention is to provide a fitness-terrain-driven hyperheuristic influence maximization method and system. It constructs a structured heuristic space through the orthogonal combination of perturbers and selectors, and relies on the synergistic mechanism of fitness terrain perception and multidimensional distribution estimation learning to quantify the macroscopic distribution skewness and microscopic solution structure characteristics in the population evolution process. In turn, it dynamically guides high-level policies to adaptively generate atomic operation sequences to achieve multi-policy collaborative optimization.
[0006] To achieve the above objectives, this invention provides a fitness-terrain-driven hyperheuristic influence maximization method, comprising the following steps: S1. Collect user relationship data in the target social network, model it as a social network graph, and determine the input parameters; S2. Initialize the low-level population; S3. Construct an atomic operation library composed of orthogonal combinations of perturber set and selector set to form a structured heuristic space for high-level policy learning and scheduling, and set a global node structure scoring mechanism. S4. Characterize each chromosome in the population, and sequentially execute the atomic operations in the atomic operation sequence to update the current solution and determine the optimal solution for the trajectory; S5. The current population evolutionary state is quantified through the fitness terrain perception mechanism and transformed into a feedback signal that can be used for high-level policy learning. S6. Perform high-level policy learning and sequence adaptive generation based on multidimensional distribution estimation; S7. Employs a master-slave MPI parallel architecture to perform distributed processing of low-level chromosome evolution tasks.
[0007] Preferably, in step S1, the user relationship data in the input target social network is modeled as a social network graph. ,in Represents the set of user nodes in the network. This represents the set of interaction relationships or social connections between users; the input parameters include the size constraint of the target seed node set. Population size Length of operation sequence for high-level strategies Elite retention rate Maximum iteration algebra .
[0008] Preferably, step S2 specifically includes: The lower-level population is initialized randomly, and for each individual in the population... From the user node set Random selection without replacement 1. Construct initial candidate solutions from 10 distinct nodes. This ensures that each initial candidate solution satisfies Furthermore, there are no duplicate nodes, and the initial low-level evolutionary population is composed of all initial candidate solutions. .
[0009] Preferably, the global node structure scoring mechanism in step S3 is used to guide the perturber to generate high-quality candidate solutions, and the system pre-calculates the comprehensive influence score of all nodes in the network. This score is based on the bridging potential index of the fusion nodes. and neighbor's degree and Composed of: ; in, For nodes The neighborhood group, For nodes The neighborhood group, The weighting coefficients are used to adjust the relative contributions of the bridging potential index and neighbor degree to the overall influence score, and are set with equal weights. Represents a node The bridging potential index is used to measure the bridging potential of nodes. The number of different nodes that can be reached by a two-hop path outside a closed neighborhood. Represents a node The degree of the neighbor and, For nodes The bridging potential index For nodes The sum of neighbor degrees is used. An equal-weighting method is employed to ensure the bridging potential index contributes equally to the overall influence score as the sum of neighbor degrees. All scores are calculated and stored once during algorithm initialization for sharing by the perturber.
[0010] Preferably, the set of disturbances and the set of selectors in step S3 are as follows: Perturber set: Let the current solution be The temporary candidate solution generated after perturbation is denoted as This includes the following six types of disturbances: RandomReplace: Randomly selects the number of replacements. }, and then from Random selection Replace different nodes Randomly selected Positions to generate ; NeighborExpansion: will The nodes in the data are scored based on their overall influence. Sort in ascending order and select the one with the lowest score. The set consists of nodes that need to be replaced. For each node in this set... From its neighbor set Select one that is not in Nodes in The probability of selection and The overall influence score is directly proportional to the selected score. replace To generate ; EliteGuidedReplace: Statistics on the previous generation of elites Historical frequency of all nodes .make This refers to nodes in the current solution that do not appear in the elite set. For Each node in the process, from and The middle is proportional to The probability of selecting a replacement node; DEProbabilistic: Randomly select three distinct individuals from the current population. ,by As the fundamental solution, calculate the difference set. Determine the number of replacements. ,in This is the scaling factor. Execute on the basis The mutation solution is obtained through this replacement. Finally, the crossover probability is used to determine the mutated solution. Cross the current solution with the mutated solution and call the repair operator to eliminate duplicate nodes, so that the solution size remains the same. .
[0011] SampledTabu: From the candidate node set The sampling size is proportional to the overall score. ( Candidate set (for network average degree) The system constructs single-node replacement neighborhoods based on nodes in the current solution and nodes in the candidate set. During the neighborhood search, if a neighborhood solution satisfies the desire criterion, it is selected as a candidate solution and the neighborhood search ends immediately; otherwise, the solution with the highest fitness from the non-taboo neighborhood solutions with better fitness than the current solution is selected as a candidate solution. After obtaining a valid move, the corresponding move is added to the taboo list of SampledTabu, and the taboo term is [duration missing]. exist Random value selection; DiversityGuidedSearch: Position-by-position optimization. First, let... For each node, calculate a hybrid score that balances coverage diversity and individual quality: ,in For diversity weights. Then for... Each position According to The probability is proportional to Nodes are selected for incremental replacement, and external metrics of affected nodes are dynamically adjusted. Selector set: Let the current solution be... Candidate solutions fitness and Three selectors are used: Greedy selector: only when When new interpretations are accepted.
[0012] Tabu-aware selector: If the transition from the current solution to a candidate solution is not locked by the tabu list and satisfies... If the desire criterion is triggered, then a new solution will be forced to be accepted.
[0013] Metropolis Selector: By Probability Accepting inferior solutions, the temperature decreases linearly with iteration.
[0014] Atomic operations are ordered pairs of perturbers and selectors. ,in, For disturbance, For selectors; perform Cartesian products of all 6 perturbators and 3 selectors to generate atomic operations to form an atomic operation library. Each atomic operation establishes a context dynamic counter to record its execution frequency and the overlap and difference between the generated new solution and the current global optimal solution in real time.
[0015] Preferably, step S4 specifically includes: To achieve decoupling and unified representation of high-level strategies and low-level problem spaces, the system will use each chromosome... Represented as a binary tuple ;in For length is The atomic operation sequences are sampled from the atomic operation library during the initialization phase, with each sequence position sampled from the library in a uniform distribution. For the reason The low-level solution consists of nodes, given a chromosome. The evolution process of its lower-level solutions is as follows: First, let the current solution be... Current fitness ; For atomic operation sequences Each atomic operation in Execute in sequence: Call the perturbator Generate candidate solutions ; Call selector According to the current solution Candidate solutions and current fitness and Decision on a new interpretation ; Update the current solution and fitness to , ; Update the optimal trajectory solution during the execution of the current atomic operation sequence; if the fitness of the current solution... Higher than the optimal fitness of the trajectory Then let , ; Record the improvement flags, structural overlap, and differences of this operation (for subsequent feedback); after the atomic operation sequence is completed, return the lowest-level solution with the highest fitness in the sequence's execution trajectory, denoted as the optimal trajectory solution. and fitness ; This mechanism transforms high-level strategies into an ordered sequence of operations on low-level solutions, thus decoupling the high-level control strategy from the low-level problem space. For each chromosome in the population, its optimal trajectory solution is obtained after evolution according to the above process. After all chromosomes have evolved, compare all The fitness is used to determine the optimal solution for the current generation, and the optimal solution for the current generation is superior to the current global optimal solution. Updated regularly .
[0016] Preferably, the fitness terrain perception mechanism in step S5 includes macro-fitness distribution perception and micro-search behavior measurement, specifically: Macro level: Let the fitness value of each chromosome in the current population be... ; Calculate the Medcouple robust skewness index, which characterizes the asymmetry of population distribution. and unbiased sample skewness index ;in, The skew direction used to characterize the fitness distribution, with a value range of [value range missing]. ; This indicates that the fitness distribution of the population is symmetrical. This indicates a right-skewed distribution (most individuals are poor, while a few elites lead the way). This indicates a left skew (most individuals cluster in high-fitness areas, which carries the risk of getting trapped in local optima). It is used to characterize the intensity of the distribution deviating from the symmetry state. The larger the absolute value, the more extreme the population aggregation or dispersion state. At the microscopic level: Record the new solutions generated by each atomic operation. Compared with the current global optimal solution The structural relationship between them, for a given atomic operation, is defined by the degree of structural overlap as: ; Structural dissimilarity is defined as: ; At the end of the current generation cycle, based on the execution record of atomic operations on the elite chromosome, the statistics of each atomic operation are compiled. Average structural overlap and average difference ; Then, the direction factor is constructed based on the skewness information. and intensity factor ,in, yes The sliding window smoothing value is used to suppress the influence of single-generation random fluctuations on the feedback signal; the direction factor is used to adjust the relative contribution of structural overlap and structural difference in the operating weights; and the intensity factor is used to adjust the learning rate of the subsequent probabilistic model.
[0017] Finally, based on the direction factor, the overall guidance weight of each atomic operation under the current terrain is calculated: These weights will be used for subsequent updates to the probabilistic model. When the distribution is right-skewed, the model prefers operations with high overlap; when the distribution is left-skewed, the model prefers operations with high dissimilarity. The intensity factor is used to adjust the learning rate of the subsequent probabilistic model, allowing the model update magnitude to adaptively adjust according to the degree to which the fitness distribution deviates from the symmetric state.
[0018] Preferably, step S6 specifically includes: High-level policy learning maintains a three-dimensional probability matrix This describes the transition probability between adjacent positions in a sequence of atomic operations; it is defined as follows: , indicating in sequence The current operation being performed is At that time, the next position Call operation The conditional probability, and , During initialization, all elements of the matrix are set to be uniformly distributed, where, This represents the number of atomic operations in the atomic operation library.
[0019] After each generation of population evolution ends, the top-fittest individuals are selected from the current population. A proportion of chromosomes as an elite set ,extract The atomic operation sequence of all individuals, and the statistics of each position. From operation Transferred to frequency Based on this, the transfer frequency matrix is calculated. ; Subsequently, the first The overall guiding weight of each atomic operation The transition frequencies are weighted to obtain a weighted frequency matrix, which is then normalized to obtain the target transition distribution. To achieve a balance between fast model convergence and preserving diversity, the Kullback-Leibler divergence between the target distribution and the current model distribution is further calculated, and the base learning rate is adaptively determined based on this divergence under different positions and operating conditions. When the target distribution differs significantly from the current model, the base learning rate is increased; conversely, the base learning rate is decreased to avoid excessive model fluctuations. Subsequently, an intensity factor is used... according to The base learning rate is adjusted to obtain the final learning rate used for updating the probabilistic model. .
[0020] Finally, the probability model is updated using exponential smoothing: ; in, Indicates the position in the sequence before the update. The current atomic operation is At that time, the next position selects an atomic operation. The conditional probability; This represents the normalized target transition probability; This represents the base learning rate determined by the Kullback-Leibler divergence. This represents the final learning rate used for updating the probabilistic model after adjustment by the intensity factor.
[0021] After the update, each row is renormalized to ensure that the sum of conditional probabilities is 1. When generating the next generation of chromosomes, for the retained elite individuals, their atomic operation sequences are directly inherited; for non-elite individuals, their atomic operation sequences are generated through probability sampling: the first atomic operation is weighted based on the positive fitness gain accumulated from each atomic operation, and the remaining positions are recursively sampled based on the three-dimensional probability matrix: the current position of the known sequence is given. and current operation Next operation With probability Select.
[0022] Preferably, step S7 specifically includes: Considering the ever-growing scale of social networks, population evolution and fitness evaluation in a single-machine environment face significant computational pressure. Therefore, a master-slave MPI parallel architecture is adopted to distribute the low-level chromosome evolution task. In this architecture, the master node is responsible for population task partitioning, dynamic scheduling, and global state maintenance. Specifically, the master node divides the chromosomes in the current population into several sub-tasks based on the number of worker nodes and distributes them to each worker node. Each worker node independently executes the low-level evolution process of the received chromosome, namely, the execution of the atomic operation sequence of complete step S4, determination of the optimal trajectory solution, and return of fitness. After all worker nodes have completed their computations, the master node aggregates the optimal trajectory solutions and their fitness returned by each sub-task, updates the current generation's optimal solution uniformly, and ensures that the current generation's optimal solution is superior to the current global optimal solution. Updated regularly This triggers the fitness terrain perception mechanism in step S5 and the high-level policy learning mechanism in step S6. This parallel design can significantly reduce the overall evaluation time in population evolution and improve the engineering scalability and practical deployment efficiency of the algorithm in large-scale social network scenarios without affecting the solution quality.
[0023] This invention also provides a fitness-terrain-driven hyperheuristic influence maximization system, comprising: A low-level parallel evaluation module is used for the construction of atomic operation libraries and the evolution of low-level solutions; The fitness terrain perception module is used to perceive the population fitness distribution status and solution structure changes; The multidimensional distribution estimation learning module is used to learn the cooperative relationships between atomic operations and adaptively generate new operation sequences to guide the next generation of evolution.
[0024] Therefore, the present invention employs the above-mentioned adaptive terrain-driven hyperheuristic influence maximization method and system, which has the following beneficial effects: (1) Enhanced diversity of search strategies and scheduling flexibility. By orthogonally combining various types of perturbers and selectors, a rich library of atomic operations is constructed, achieving decoupling between high-level scheduling and low-level execution. This decoupling mechanism enables the algorithm to flexibly select and combine atomic operations from the operation library according to the needs of the evolution stage, overcoming the limitation of fixed search patterns in traditional metaheuristic methods, and achieving a better balance between search efficiency and solution quality.
[0025] (2) Enhance the directional guidance capability of the search process. By calculating the skewness characteristics of the population fitness distribution online, as well as the structural overlap and difference between the new solution and the current best solution, dynamic directional and intensity factors are constructed, thereby transforming the abstract evolutionary state into a quantifiable feedback signal. This feedback signal can directly guide the high-level strategy to adjust the exploration and development behavior in a targeted manner, effectively reducing the proportion of ineffective searches and suppressing premature convergence.
[0026] (3) Adaptive scheduling of multi-strategy cooperative search is realized. A three-dimensional transition probability model is constructed through a multidimensional distribution estimation learning mechanism to learn the operator cooperative pattern from the historical effective search trajectory. The learning rate is adaptively adjusted by combining KL divergence and intensity factor to generate new operation scheduling sequence, thus overcoming the problem of rigid strategy switching.
[0027] (4) It combines high solution quality with engineering scalability. The master-slave MPI parallel architecture is used to perform low-level chromosome evolution and fitness evaluation, which effectively breaks through the computing power bottleneck in the single-machine environment. While ensuring the solution quality, it greatly improves the algorithm's running efficiency and practical engineering deployment capability.
[0028] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0029] Figure 1 This is a basic topological structure diagram of the social network graph in an embodiment of the present invention; Figure 2 This is an overall framework diagram of FLHAM according to an embodiment of the present invention; Figure 3 The figure shows the comparison of the final influence propagation range of the embodiments of the present invention under different seed set sizes, where (a) is Blog, (b) is CA-GrQc, (c) is Lastfm, (d) is CA-HepTh, (e) is PGP, (f) is SisterCities, (g) is NetHEHT, (h) is Deezer, and (i) is Gemsec-RO. Figure 4 The present invention provides a heat map of the disturbance-selector combination in an embodiment of the invention, wherein (a) is CA-GrQc k=30, (b) is CA-GrQc k=60, (c) is CA-GrQc k=100, (d) is PGP k=30, (e) is PGP k=60, (f) is PGP k=100, (g) is Deezer k=30, (h) is Deezer k=60, and (i) is Deezer k=100. Figure 5 The graph shows a comparison of the algorithm running time in an embodiment of the present invention, where (a) is k=50 and (b) is k=100. Figure 6 This is a schematic diagram of the atomic operation sequence flow in an embodiment of the present invention. (a) shows the transition probability distribution of atomic operations between adjacent sequence positions 1-2 in the high-level three-dimensional probabilistic model at the initial generation; (b) shows the transition probability distribution of atomic operations between adjacent sequence positions 2-3 in the high-level three-dimensional probabilistic model at the initial generation; (c) shows the transition probability distribution of atomic operations between adjacent sequence positions 3-4 in the high-level three-dimensional probabilistic model at the initial generation; (d) shows the transition probability distribution of atomic operations between adjacent sequence positions 1-2 in the high-level three-dimensional probabilistic model at the 100th generation; and (e) shows the transition probability distribution of atomic operations between adjacent sequence positions 1-2 in the high-level three-dimensional probabilistic model at the 100th generation. (f) is the transition probability distribution of atomic operations between adjacent sequence positions 2-3 in the probabilistic model after 100 generations; (g) is the transition probability distribution of atomic operations between adjacent sequence positions 1-2 in the high-level three-dimensional probabilistic model after 200 generations; (h) is the transition probability distribution of atomic operations between adjacent sequence positions 2-3 in the high-level three-dimensional probabilistic model after 200 generations; and (i) is the transition probability distribution of atomic operations between adjacent sequence positions 3-4 in the high-level three-dimensional probabilistic model after 200 generations. Detailed Implementation
[0030] The following detailed description of embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0031] In practical engineering implementation, the fitness-terrain-driven hyperheuristic influence maximization method described in steps S1 to S7 can be encapsulated as an influence maximization problem-solving component and deployed in a distributed computing cluster containing processors and memory. This component provides a standardized calling interface to receive... Figure 1 The social network topology data and seed set size constraints shown are used to return the optimal set of seed nodes obtained in the current search process after hyperheuristic adaptive solution, which can be called by business decision-making systems such as advertising targeting and public opinion intervention.
[0032] The component is internally divided according to a hyperheuristic framework. Figure 2The three functional modules shown are low-level parallel evaluation, fitness terrain perception, and multidimensional distribution estimation learning. These modules interact through predefined data structures and can be independently configured and replaced. The low-level parallel evaluation module is responsible for building and managing the atomic operation library and, based on the MPI parallel architecture, distributes individual operation sequences from the population to multiple worker nodes for parallel execution. Each node independently completes perturber and selector operations, and the master node summarizes the results and updates the current optimal solution. The fitness terrain perception module analyzes the skewness characteristics of the population fitness distribution in real time, as well as the structural overlap and difference between the new solution and the current optimal solution. Based on this, it dynamically generates direction factors and intensity factors. The direction factor is used to calculate the comprehensive guidance weight of each atomic operation, and the intensity factor is used to adjust the learning rate of the probabilistic model. The multidimensional distribution estimation learning module maintains a three-dimensional matrix describing the transition probabilities of adjacent positions in an atomic operation sequence. It updates this matrix based on the transition statistics of elite individuals, the comprehensive guidance weight, and the learning rate adjusted by the intensity factor, thereby adaptively sampling to generate a new generation of operation sequences. This allows it to learn from historical effective search trajectories and strengthen effective operator cooperative patterns.
[0033] At the engineering deployment level, the three modules mentioned above are encapsulated using a decoupled architecture, allowing for the replacement or extension of specific perturbers and selectors according to specific business scenarios. The solution engine supports MPI parallel acceleration, enabling it to adapt to the computational needs of large-scale social network environments.
[0034] To verify the effectiveness of the adaptive terrain-driven hyperheuristic influence maximization method and system of this invention, this embodiment uses nine real social networks for verification, including Blog, CA-GrQc, Lastfm, CA-HepTh, PGP, Sister Cities, NetHEHT, Deezer, and Gemsec-RO. The network node size is approximately 4000 to 42000, and the edge size is approximately 6800 to 126000. The experiment is based on an independent cascading propagation model, with the propagation probability set to... =0.05, seed set size The value range is from 10 to 100. Meanwhile, six representative algorithms—CELF, LIDDE, DHHO, PHEE, DSMO, and ASVEA—were selected as comparison algorithms. Figure 3 The results show the comparison of the final influence propagation range between the FLHAM algorithm of this invention and the comparison algorithm under different seed set sizes.
[0035] Experimental results show that FLHAM exhibits stable and excellent influence propagation performance on all nine real-world social networks. Compared with various comparative metaheuristic methods, this invention demonstrates strong consistency across different network structures. For example, compared to LIDDE, this invention achieves influence improvements of 23.67% and 22.09% on PGP and CA-GrQc networks, respectively; on Sister Cities, CA-HepTh, and NetHEHT networks, the improvements are 13.15%, 12.86%, and 11.21%, respectively; on Blog and Deezer networks, the improvements are 5.65% and 5.04%, respectively; even on the Gemsec-RO network, where the improvement is relatively small, a performance gain of 3.04% is still achieved. Compared to DHHO, the improvement ranges from 2.25% to 19.64%, with particularly significant advantages on PGP and CA-GrQc networks. Compared to PHEE and DSMO, the improvements are concentrated between 0.78% and 10.54% and 0.92% and 7.26%, respectively. Compared to ASVEA, although the improvement is relatively modest, it still represents an improvement of over 3% on networks such as CA-GrQc and PGP.
[0036] Compared with CELF, the present invention has achieved similar or even slightly better experimental results on some networks, and can keep the gap within a small range on other networks, indicating that the present invention can maintain a high level of propagation performance while significantly improving computational efficiency.
[0037] The aforementioned performance advantages can be attributed to three core mechanisms: the hyperheuristic framework effectively overcomes the local optima limitation caused by operator solidification in traditional metaheuristics by dynamically scheduling perturbers and selectors through an atomic operation library; the fitness terrain-aware mechanism utilizes the robustness of population distribution and the structural overlap and difference of solutions to construct direction factors and intensity factors, which respectively adjust the weighted direction of atomic operations and the update intensity of the probabilistic model, effectively suppressing premature convergence; and the multidimensional distribution estimation learning mechanism extracts the atomic operation transition probabilities from elite sequences, determines the basic learning rate through KL divergence, and adjusts the final learning rate used for probabilistic model updates in conjunction with the intensity factor, learning and reinforcing efficient operator cooperative patterns from historical successful trajectories. Figure 3 The results demonstrate that FLHAM possesses excellent solution quality and generalization ability across different network topologies, proving the effectiveness and robustness of this invention in maximizing influence in complex social networks.
[0038] Figure 4This is a heatmap showing the average frequency of atomic operations in the optimal chromosome under different network and seed set sizes. The vertical axis represents six perturbators, and the horizontal axis represents three selectors; the darker the color, the higher the frequency of that combination in the optimal sequence. This figure illustrates the preference of the multidimensional distribution estimation learning mechanism for different operator combinations, and the reinforcing effect of the elite preservation strategy on efficient combinations, verifying that the algorithm can learn and reinforce efficient operator cooperative patterns from the search history.
[0039] Figure 5 The graph showing the comparison of algorithm runtime between the FLHAM algorithm and the comparison algorithm is presented on nine real social networks and two seed set sizes, with the vertical axis representing a logarithmic scale. The results demonstrate that FLHAM achieves several orders of magnitude speedup compared to CELF while maintaining comparable runtime efficiency to the optimal metaheuristic method, validating the high computational efficiency and engineering scalability of this invention in large-scale social network scenarios.
[0040] Figure 6 Using the Deezer network as an example, this paper demonstrates the atomic operation transition probabilities between adjacent sequence positions in a high-level three-dimensional probabilistic model under different evolutionary generations. Nodes in the diagram represent atomic operations, and the width of directed edges is proportional to the transition probability. Initially, all transition probabilities are uniformly distributed, reflecting the unbiased initialization of the high-level policy. As evolution progresses, several high-probability transition paths gradually emerge, indicating that the multidimensional distribution estimation learning mechanism can effectively learn and reinforce high-quality operation connection patterns from elite sequences.
[0041] Therefore, the present invention adopts the above-mentioned fitness terrain-driven hyperheuristic influence maximization method and system, which effectively overcomes the problems of premature convergence and scheduling rigidity while significantly reducing computational overhead, highlighting its excellent performance and broad application potential in solving the influence maximization problem of large-scale complex networks.
[0042] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A fitness-terrain-driven hyperheuristic influence maximization method, characterized in that, Includes the following steps: S1. Collect user relationship data in the target social network, model it as a social network graph, and determine the input parameters; S2. Initialize the low-level population; S3. Construct an atomic operation library composed of orthogonal combinations of perturber set and selector set to form a structured heuristic space for high-level policy learning and scheduling, and set a global node structure scoring mechanism. S4. Characterize each chromosome in the population, and sequentially execute the atomic operations in the atomic operation sequence to update the current solution and determine the optimal solution for the trajectory; S5. The current population evolutionary state is quantified through the fitness terrain perception mechanism and transformed into a feedback signal that can be used for high-level policy learning. S6. Perform high-level policy learning and sequence adaptive generation based on multidimensional distribution estimation; S7. Employs a master-slave MPI parallel architecture to perform distributed processing of low-level chromosome evolution tasks.
2. The fitness-terrain-driven hyperheuristic influence maximization method according to claim 1, characterized in that: In step S1, the user relationship data in the input target social network is modeled as a social network graph. ,in Represents the set of user nodes in the network. This represents a set of edges representing the interaction relationships or social connections between users. Input parameters include the size constraint of the target seed node set. Population size Length of operation sequence for high-level strategies Elite retention rate Maximum iteration algebra .
3. The fitness-terrain-driven hyperheuristic influence maximization method according to claim 2, characterized in that, Step S2 is as follows: The lower-level population is initialized randomly, and for each individual in the population... From the user node set Random selection without replacement 1. Construct initial candidate solutions from 10 distinct nodes. This ensures that each initial candidate solution satisfies Furthermore, there are no duplicate nodes, and the initial low-level evolutionary population is composed of all initial candidate solutions. .
4. The fitness-terrain-driven hyperheuristic influence maximization method according to claim 1, characterized in that, The global node structure scoring mechanism in step S3 is used to guide the perturber to generate high-quality candidate solutions. The system pre-calculates the comprehensive influence score of all nodes in the network. Its bridging potential index through fusion nodes and neighbor's degree and Composed of: ; in, For nodes The neighborhood group, For nodes The neighborhood group, These are the weighting coefficients. For nodes The bridging potential index For nodes The neighbor's degree and, For nodes The bridging potential index For nodes The degree of the neighbor.
5. The fitness-terrain-driven hyperheuristic influence maximization method according to claim 1, characterized in that, The perturber set and selector set in step S3 are specifically as follows: Perturber set: Let the current solution be The temporary candidate solution generated after perturbation is denoted as It includes six types of perturbators: RandomReplace, NeighborExpansion, EliteGuidedReplace, DEProbabilistic, SampledTabu, and DiversityGuidedSearch. Selector set: Let the current solution be... Candidate solutions Current fitness , ,in, This represents the fitness function used to evaluate the quality of low-level solutions; it includes greedy selectors, tabu-aware selectors, and Metropolis selectors. Atomic operations are ordered pairs of perturbers and selectors. ,in, For disturbance, For selectors; perform Cartesian products of all 6 perturbators and 3 selectors to generate atomic operations to form an atomic operation library. Each atomic operation establishes a context dynamic counter to record its execution frequency and the overlap and difference between the generated new solution and the current global optimal solution in real time.
6. The adaptive terrain-driven hyperheuristic influence maximization method according to claim 5, characterized in that, Step S4 is as follows: The system will divide each chromosome Represented as a binary tuple ;in For length is atomic operation sequence, For the reason The low-level solution consists of nodes, given a chromosome. The evolution process of its lower-level solutions is as follows: First, let the current solution be... Current fitness ; For atomic operation sequences Each atomic operation in Execute in sequence: Call the perturbator Generate candidate solutions ; Call selector According to the current solution Candidate solutions and current fitness and Decision on a new interpretation ; Update the current solution and fitness to , ; Update the optimal trajectory solution during the execution of the current atomic operation sequence; if the fitness of the current solution... Higher than the optimal fitness of the trajectory Then let , ; Record the improvement indicators, structural overlap, and differences of this operation; After the atomic operation sequence is completed, the lowest-level solution with the highest fitness in the execution trajectory of that sequence is returned, which is denoted as the optimal solution of the trajectory. and fitness ; For each chromosome in the population, its optimal trajectory is obtained by evolving according to the above process. After all chromosomes have evolved, compare all The fitness is used to determine the optimal solution for the current generation, and the optimal solution for the current generation is superior to the current global optimal solution. Updated regularly .
7. The fitness-terrain-driven hyperheuristic influence maximization method according to claim 6, characterized in that, The fitness terrain perception mechanism in step S5 includes macro-level fitness distribution perception and micro-level search behavior measurement, specifically: Macro level: Let the fitness value of each chromosome in the current population be... ; Calculate the Medcouple robust skewness index, which characterizes the asymmetry of population distribution. and unbiased sample skewness index ;in, The skew direction used to characterize the fitness distribution, with a value range of [value range missing]. ; This indicates that the fitness distribution of the population is symmetrical. This indicates a right-skewed distribution. Indicates leftward deviation; Used to characterize the intensity of a distribution deviating from a symmetrical state; At the microscopic level: Record the new solutions generated by each atomic operation. Compared with the current global optimal solution The structural relationship between them, for a given atomic operation, is defined by the degree of structural overlap as: ; Structural dissimilarity is defined as: ; At the end of the current generation cycle, based on the execution record of atomic operations on the elite chromosome, the statistics of each atomic operation are compiled. Average structural overlap and average difference ; Then, the direction factor is constructed based on the skewness information. and intensity factor ,in, yes The sliding window smoothing value is used to suppress the influence of single-generation random fluctuations on the feedback signal; the direction factor is used to adjust the relative contribution of structural overlap and structural difference in the operational weights; and the intensity factor is used to adjust the learning rate of the subsequent probabilistic model. Finally, based on the direction factor, the overall guidance weight of each atomic operation under the current terrain is calculated: 。 8. The fitness-terrain-driven hyperheuristic influence maximization method according to claim 7, characterized in that, Step S6 is as follows: High-level policy learning maintains a three-dimensional probability matrix This describes the transition probability between adjacent positions in a sequence of atomic operations; it is defined as follows: , indicating in sequence The current operation being performed is At that time, the next position Call operation The conditional probability, and , During initialization, all elements of the matrix are set to be uniformly distributed, where, This represents the number of atomic operations in the atomic operation library. After each generation of population evolution ends, the top-fittest individuals are selected from the current population. A proportion of chromosomes as an elite set ,extract The atomic operation sequence of all individuals, and the statistics of each position. From operation Transferred to frequency Based on this, the transfer frequency matrix is calculated. ; Subsequently, the first The overall guiding weight of each atomic operation The transition frequencies are weighted to obtain a weighted frequency matrix, which is then normalized to obtain the target transition distribution. The Kullback-Leibler divergence between the target distribution and the current model distribution is further calculated, and the base learning rate is determined based on this divergence under different positions and operating conditions. ; and then use the intensity factor according to The base learning rate is adjusted to obtain the final learning rate used for updating the probabilistic model. Finally, the probability model is updated using exponential smoothing: ; in, Indicates the position in the sequence before the update. The current atomic operation is At that time, the next position selects an atomic operation. The conditional probability; This represents the normalized target transition probability; After the update, each row is renormalized to ensure that the sum of conditional probabilities is 1. When generating the next generation of chromosomes, for the retained elite individuals, their atomic operation sequences are directly inherited; for non-elite individuals, their atomic operation sequences are generated through probability sampling: the first atomic operation is weighted based on the positive fitness gain accumulated from each atomic operation, and the remaining positions are recursively sampled based on the three-dimensional probability matrix: the current position of the known sequence is given. and current operation Next operation With probability Select.
9. The fitness-terrain-driven hyperheuristic influence maximization method according to claim 1, characterized in that, Step S7 is as follows: In the master-slave MPI parallel architecture, the master node divides the chromosomes in the current population into several sub-tasks according to the number of worker nodes and distributes them to each worker node; each worker node independently executes the low-level evolution process of the received chromosomes. After all worker nodes have completed their calculations, the master node aggregates the optimal trajectory solutions and their fitness values returned by each subtask, updates the current generation's optimal solution uniformly, and ensures that the current generation's optimal solution is superior to the current global optimal solution. Updated regularly This then triggers the adaptive terrain perception mechanism in step S5 and the high-level policy learning mechanism in step S6.
10. A fitness-terrain-driven hyperheuristic influence maximization system, applied to the fitness-terrain-driven hyperheuristic influence maximization method according to any one of claims 1-9, characterized in that, include: A low-level parallel evaluation module is used for the construction of atomic operation libraries and the evolution of low-level solutions; The fitness terrain perception module is used to perceive the population fitness distribution status and solution structure changes; The multidimensional distribution estimation learning module is used to learn the cooperative relationships between atomic operations and adaptively generate new operation sequences to guide the next generation of evolution.