Influence maximization method based on direct connection and autoregression influence estimation
By using direct and autoregressive influence estimation methods, combined with graph neural networks and gradient descent optimization, an end-to-end propagation agent model is constructed, which solves the problems of low efficiency, weak generalization ability and poor scalability of influence maximization methods in existing technologies, and realizes efficient and flexible seed node selection and information diffusion.
Patent Information
- Application Number
- CN202510704999.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-05-29
AI Technical Summary
Existing influence maximization methods have low efficiency, weak generalization ability, insufficient global optimality and poor scalability in large-scale networks, making them difficult to adapt to diverse practical application needs.
A method based on direct and autoregressive influence estimation is adopted to model the node state evolution through graph neural network, build an end-to-end propagation agent model, use gradient descent to optimize the initial state of the node, and combine the direct estimator to optimize the seed node set in continuous space.
It improves the computational efficiency and diffusion effect, enhances the flexibility and adaptability of the model, enables it to adapt to complex and changing practical application needs, and significantly improves the global quality and diffusion capacity of seed node selection.
Smart Images

Figure CN120632808A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information processing, and in particular to an influence maximization method based on direct and autoregressive influence estimation. Background Art
[0002] With the rapid development of social media and content-sharing platforms, the means and scale of information dissemination have undergone profound changes. The influence maximization problem (IM) is a key topic in the study of information diffusion on social networks. Its goal is to select a small number of seed nodes in the network and maximize the overall influence through a diffusion process.
[0003] Existing solutions to the influence maximization (IM) problem can be divided into three categories according to their solution ideas: greedy algorithms, heuristic algorithms, and learning / reinforcement learning algorithms.
[0004] Greedy algorithms were originally based on Monte Carlo simulations to estimate node influence. Based on the submodularity of the propagation function, they select the node with the greatest marginal benefit in each round. This approach has theoretical guarantees and can achieve a near-optimal solution (1-1 / e). However, Monte Carlo simulations are computationally expensive, making them particularly inefficient on large-scale networks, a bottleneck for practical applications.
[0005] To this end, a greedy algorithm based on reverse influence sampling (RIS) has been proposed. This method significantly improves solution efficiency by randomly sampling sets of influenced nodes (RR sets) in the network and then searching for the nodes most likely to influence these sets. RIS-based methods are efficient and scalable, and theoretically provide probabilistic approximation guarantees. However, their performance is highly dependent on specific propagation models (such as independent cascade models and linear threshold models), limiting their adaptability in practical applications with complex or unknown propagation mechanisms.
[0006] Heuristic algorithms primarily fall into two categories: centrality-based and community-structure-based. Centrality-based heuristic algorithms measure the "importance" of nodes in the network structure by designing various centrality metrics (such as degree centrality and betweenness centrality) and directly select the top k nodes with the highest centrality as the seed set. However, such methods often overlook the issue of overlapping influence between nodes, resulting in significant redundancy among seed nodes and poor diffusion effectiveness. Improved voting mechanisms (such as discounting and sparsification) introduce strategies that account for overlap in node influence areas. While these improve the rationality of seed distribution to some extent, they remain locally greedy and lack global optimality. Community-structure-based heuristic algorithms first divide the network into communities and then select important nodes within or across communities as seeds. While this approach can ensure seed dispersion to a certain extent, it often relies on locally greedy node selection based on heuristic experience, lacks accurate modeling of the actual diffusion capabilities of nodes, and thus struggles to maximize influence globally, ultimately resulting in mediocre performance in actual diffusion.
[0007] Learning-based methods (including reinforcement learning and supervised learning) have also been widely used in recent years to solve the influence maximization problem. Reinforcement learning-based methods are often combined with graph neural networks (GNNs) to learn seed node selection strategies through interaction with the environment. However, reinforcement learning methods suffer from difficulties in training convergence, sensitivity to reward design, and insufficient generalization in complex real-world communication environments. Supervised learning methods based on node ranking attempt to select seeds by directly predicting node importance, but they also struggle to effectively address the overlap of influence between nodes, resulting in limited overall diffusion effectiveness. Furthermore, the DeepIM model, based on latent space optimization, introduces a generative model to reduce the dimensionality of the node space and then optimizes in the low-dimensional latent space to solve the influence maximization problem. While this improves efficiency to some extent, due to the weak generalization ability of deep learning models on out-of-distribution data, the decoder of the generative model tends to map latent variables into invalid or low-quality seed sets, affecting the ultimate diffusion effect.
[0008] On the other hand, most existing methods are designed and optimized around the standard influence maximization (IM) setting, and lack good adaptability to diverse practical application requirements (such as target influence maximization, expanded node budget constraints, target node isolation, etc.), resulting in insufficient flexibility and scalability, making it difficult to effectively cope with complex and changeable real-world tasks. Summary of the Invention
[0009] In order to overcome the defects in the above-mentioned prior art, the present invention provides an influence maximization method based on direct and autoregressive influence estimation. In view of the problems of low efficiency, weak generalization ability, insufficient global optimality and poor scalability in the prior art, an influence maximization method is proposed that can take into account diffusion effect, computational efficiency and application flexibility.
[0010] To achieve the above object, the present invention adopts the following technical solutions, including:
[0011] Influence maximization methods based on direct and autoregressive influence estimation, including:
[0012] Given a propagation network G, the propagation dynamics is modeled based on the propagation trajectory data of the nodes in the network, and the propagation model M(x, G; θ) is obtained to model the node state evolution process, where x represents the node state and θ is the model parameter.
[0013] Use the propagation model M(x,G;θ) to transform the initial state x of the node into 0 Mapped to the final state y of the spread, that is, the final infection probability of the node, and the initial state x of the node 0 Mapped to a continuous state vector z, construct an end-to-end propagation agent model y = M s (z,G;θ);
[0014] Fixed model parameters θ, with the goal of maximizing influence, the initial state x of the node is quantified by gradient descent method. 0 The corresponding continuous state vector z is optimized, and the seed node set is obtained according to the optimization result.
[0015] Preferably, the node propagation trajectory data consists of infected nodes and their corresponding infection time; the node state trajectory x at different time steps is divided using a time window 0 ,…x t ,…,x T-1 ; Node state trajectory at time step t in, represents the state of node i at time step t;
[0016] According to the evolution of node states, we use graph neural networks to model the given propagation network G and construct a propagation model M(x, G; θ) through autoregression. The model parameter θ is calculated as follows:
[0017]
[0018] Where E[·] represents the expectation; p θ (x t+1 ∣x t ,…,x 0 ,G) represents the propagation model in xt ,…,x 0 Under the conditions of and G, generate x t+1 probability.
[0019] Preferably, a propagation model is constructed based on the propagation mechanism, and the node state trajectory is updated in the following manner:
[0020]
[0021] Among them, f(·) and g(·) are functions of the propagation model, N(·) represents the node’s neighbor set, and Δt is the time step size.
[0022] Preferably, a propagation model is constructed based on a neural network, and a gated recurrent unit is used in combination with a graph neural network to model the evolution of the node state, as shown below:
[0023]
[0024]
[0025] Where N(·) represents the node’s neighbor set; is the hidden state of the gated recurrent unit; represents the hidden state of node i at time step t; represents edge-level MLPs, f out (·) denotes the MLPs of the output layer;
[0026] At time step t, the graph neural network updates its own hidden representation by aggregating the hidden representations of neighboring nodes and its own state. The output layer predicts the state of the node in the next time step based on the hidden representation of the node.
[0027] Preferably, the propagation model M(x,G;θ) is used to transform the initial state x of the node into 0 Mapping to the propagation final state y = x T-1 =M(x 0 ,G;θ), the initial state of the node x 0 From discrete values, namely binary vectors, to continuous values, namely continuous state vector z, the continuous state vector z of a node represents the probability of each node being selected, and an end-to-end propagation agent model y = M is constructed. s (z, G; θ); fix the model parameter θ, set the propagation final state y with the goal of maximizing influence, and optimize the initial state x of the node by gradient descent method 0 The corresponding continuous state vector z is mapped back to a discrete binary vector through the straight-through estimator STE, thereby determining the initial state of the node and obtaining the seed node set.
[0028] Preferably, node budget is added as a constraint.
[0029] The present invention also provides a readable storage medium having a computer program stored thereon, which implements the influence maximization method based on direct and autoregressive influence estimation when the computer program is executed.
[0030] The present invention also provides a computer program product comprising a computer program / instruction, which implements the influence maximization method based on direct and autoregressive influence estimation when the computer program / instruction is executed by a processor.
[0031] The advantages of the present invention are:
[0032] (1) This invention is a further development of the existing technology. It solves the influence maximization (IM) problem by combining straight-through estimation with an autoregressive information diffusion model based on a graph neural network. By designing an end-to-end model and optimization objectives, this invention can extend the typical influence maximization (IM) problem to other scenarios, including expanding the definition of budget constraints, target influence maximization, and target isolation.
[0033] (2) The present invention proposes a method for constructing an end-to-end propagation agent model based on graph neural networks. By learning historical propagation trajectories, it autoregressively simulates the evolution of node states in continuous time steps, and then accurately approximates the information propagation process in complex networks, providing a derivable and efficient propagation agent model for subsequent seed set inference.
[0034] (3) The present invention proposes a continuous relaxation and direct estimation mechanism for the initial state of the node, which relaxes the initial state of the node from a discrete binary vector to a continuous state vector to support gradient-based optimization, and introduces a direct estimator to achieve mapping to discrete seed selection while maintaining the trainability of the propagation model.
[0035] (4) The present invention proposes a continuous inference strategy for jointly optimizing seed sets. Based on the frozen propagation agent model parameters and combined with a differentiable optimization objective function, the seed probability vector is iteratively updated in the continuous space to achieve joint optimization of the seed set in high-dimensional space, thereby improving the global quality and diffusion capability of seed selection.
[0036] (5) The present invention proposes a seed inference mechanism that can adapt to a variety of task settings. The inference mechanism has flexible target adaptation capabilities and can dynamically adjust the optimization strategy according to different actual application scenarios (such as changes in node budget constraints, adjustments to the influence target area, modifications to constraints, etc.), thereby achieving support for a variety of diffusion optimization tasks under a unified framework.
[0037] (6) Compared with the greedy algorithm based on Monte Carlo simulation, the heuristic method with improved voting mechanism, and the reinforcement learning method, the present invention optimizes the seed node selection process, improves the convergence speed and stability of the model training phase, and realizes a more efficient decision-making process for the target scenario in the inference phase, thereby having a faster solution speed and better resource utilization efficiency in large-scale networks, thereby improving computing efficiency.
[0038] (7) The design of the present invention introduces an adaptation mechanism for different propagation models and diverse practical scenarios (such as budget constraints, specific target node propagation requirements, etc.), which can maintain stable performance under different propagation environments and application conditions, has good scalability and practicality, and enhances generalization and adaptability.
[0039] (8) While maintaining theoretical interpretability, the present invention also takes into account the feasibility of engineering implementation and can effectively solve core technical problems in real-world applications such as information diffusion, advertising placement, and public opinion intervention.
[0040] (9) The present invention not only overcomes the shortcomings of existing methods in terms of efficiency, diffusion effect and scalability, but also can adapt to more complex and changeable practical application needs, and has significant application value and promotion prospects.
[0041] (10) We conducted extensive experiments on real-world datasets and compared them with various advanced methods. The experimental results show that our model achieves excellent performance on almost all datasets, significantly surpassing existing advanced solutions, thus verifying the advantages and effectiveness of our model and improving the impact maximization effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 Schematic diagram of the influence maximization method based on direct and autoregressive influence estimation of the present invention.
[0043] Figure 2 2 is a comparison chart of experimental results of an embodiment of the present invention. DETAILED DESCRIPTION
[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0045] Example 1
[0046] The influence maximization method based on straight-through and autoregressive influence estimation of the present invention adopts a model-based optimization framework STARIM (Straight-Through and Autoregressive Influence Estimation for Influence Maximization) to solve the problem of influence maximization. Under the premise of assuming that information propagation is probabilistic, STARIM uses graph neural networks to model the continuous evolution of node states in information propagation in an autoregressive manner, and combines straight-through estimation to build an end-to-end propagation agent model, thereby efficiently inferring the seed node set in the continuous optimization space. The overall framework is as follows Figure 1 As shown, STARIM consists of two parts:
[0047] The propagation dynamics modeling stage, such as Figure 1 As shown in part (a) of the paper. First, the propagation network G and multiple propagation trajectory data are input. Each propagation trajectory consists of a series of infected nodes and their corresponding infection time. Then, the propagation trajectory data is divided into node state trajectories x at different time steps using a time window. 0 ,…,x T-1 The propagation model uses graph neural networks (GNN) to fit these propagation trajectories in an autoregressive manner, that is, predicting the infection probability of the node at each time step based on the historical state, thereby approximating the actual propagation trajectory on a given network.
[0048] Seed node inference stage, such as Figure 1 As shown in part (b) of the figure, the initial state of the node x 0 Relax from the binary vector to a continuous state vector z to support gradient-based optimization, and then use a straight-through estimator to map the continuous state vector z back to the discrete state x 0 The GNN-based propagation model can use autoregressive method to predict the initial state of the node x 0 The corresponding final infection probability y is combined with the direct estimation to construct the end-to-end propagation agent model y = M s (z, G; θ). Finally, freeze the propagation agent model M s The model parameters are updated according to the custom optimization objectives, and the continuous state vector z of the node is updated. The seed node set is selected through multiple rounds of iterative optimization to maximize the diffusion influence of the seed node set.
[0049] 1. The specific method of propagation dynamics modeling is as follows:
[0050] Given a propagation network G, the dynamics of information propagation is modeled based on the propagation trajectory data of the nodes in the network. Each trajectory consists of a series of infected nodes and their corresponding infection time. The propagation trajectory data of the node is divided into node state trajectories x at different time steps using a time window. 0 ,…x t ,…,x T-1 , where T is the total number of time steps and the node state trajectory at time step t is use Represents the state of node i at time step t, the state trajectory of node i
[0051] According to the evolution process of node status, by using graph neural network to model under the condition of given propagation network G, p θ (x t+1 ∣x t ,…,x 1 ,G), and finally construct the propagation model M(x,G;θ) by autoregression. The model parameter θ is calculated by maximizing the following probability:
[0052]
[0053] Where E[·] represents the expectation; p θ (x t+1 ∣x t ,…,x 0 ,G) represents the propagation model in x t ,…,x 0 Under the conditions of and G, generate x t+1 probability.
[0054] In this embodiment, the propagation model can be constructed in two ways:
[0055] Implementation method 1: Construct a propagation model based on the propagation mechanism. Both information propagation and message passing of graph neural networks rely on the topological structure of the graph and realize the update and diffusion of information through the interaction between neighbors. The two are highly similar in mechanism. The mechanism propagation model updates the state of the node by performing matrix multiplication, which can be regarded as a form of message passing. Under the premise that the propagation mechanism is known and it is probabilistic propagation, the message passing framework of the graph neural network is used to represent the quenched mean field equation corresponding to the propagation model, and the model describes the evolution of the node state in the complex network. Under this setting, represents the infection probability of node i at time step t. In fact, for the node i at time step t in the propagation network G, The synchronous state update caused by one propagation of the mechanism propagation model is equivalent to each node in the graph aggregating the states of its neighbors and performing one operation according to the mechanism model. The state update of the entire network can be uniformly expressed as:
[0056]
[0057] Among them, f(·) and g(·) are functions of the propagation model, N(·) represents the neighbor set of the node, Δt is the time step size, and Δt = 1 is used as the default value in the future. The message passing of the graph neural network is used to simulate the propagation dynamics of the propagation model in discrete time steps.
[0058] The following describes the above process using a typical propagation model in a social network scenario as an example. A typical individual-based SIR model can use a graph neural network to construct the following propagation dynamics in discrete time steps:
[0059]
[0060] in, is the adjacency matrix of the network, is a vector representing three different states of the SIR model of all nodes in the network, is the representation of the three states at time t, ⊙ represents the element product, the infection rate of the edge is represented as β∈[0,1], and the recovery rate of the node is represented as γ∈[0,1].
[0061] Similarly, the typical independent cascade model (IC) can be viewed as a form where γ is 1 and the infection probability between different pairs of nodes is randomly different. Graph neural networks can be used to construct the following propagation dynamics in discrete time steps:
[0062]
[0063] in, is the infection probability between node pairs, that is, B i,j is the infection probability of node i to node j.
[0064] Implementation method 2: Constructing a propagation model based on a neural network (data-driven model). For real-world scenarios where the diffusion model is complex and the model is unknown, the propagation network G and the propagation trajectory data are used to train the neural network model to model the implicit dynamics of the actual propagation data. The task of the neural network model is to predict the future node state dynamics evolution p in the propagation scenario. θ (x t+1 ∣x t ,…,x 1,G). Since the propagation of real-world scenarios is usually non-Markov, that is, the current state not only depends on the previous state, but may also be affected by all previous states, the GRU (Gated Recurrent Unit) structure is used in combination with the graph neural network to model the evolution of the node state x 0 ,…,x T-1 , as shown below:
[0065]
[0066] Among them, N(·) represents the neighbor set of the node, is the hidden state of GRU (Gated Recurrent Unit); represents the hidden state of node i at time step t; represents edge-level MLPs, f out (·) denotes the MLPs of the output layer.
[0067] At time step t, the graph neural network updates its own hidden representation by aggregating the hidden representations of neighboring nodes and its own state. The output layer predicts the state of the node in the next time step based on the hidden representation of the node.
[0068] Assuming that the infection probability of the model follows a Gaussian distribution, the mean square error between the predicted value and the true value is used as the loss function to model the infection probability of the modeled node at each time step, so that the model models the evolution of the node infection probability, which is similar to mean field dynamics.
[0069] 2. The specific method of seed node inference is as follows:
[0070] The propagation model M(x,G;θ) can be used to model the node state evolution process, that is, to model p θ (x t+1 ∣x t ,…,x 1 ,G), so the initial state x of the node can be autoregressively 0 Mapped to the propagation final state y, so given the initial state x 0 An end-to-end propagation model can be constructed:
[0071] y=x T-1 =M(x 0 ,G;θ)
[0072] Where y is the node infection probability at time step T-1.
[0073] In the scenario of the influence maximization problem on a graph with N nodes (communication network G), each node can be in one of two discrete states: selected or unselected. 0Relaxation from discrete values (binary vectors of 0 or 1) to continuous values, namely the continuous state vector z, where the continuous state vector z of a node represents the probability of each node being selected and can be any real number between 0 and 1. This relaxation method supports gradient-based optimization. To bridge the continuous and discrete domains, a straight-through estimator (STE) is used to map the continuous probabilities back to discrete binary values (0 or 1) during evaluation.
[0074] In the forward propagation, the input z represents the continuous state vector of the node, and a threshold operation is used to map the continuous value to a discrete binary value. Specifically, when z>σ, the output is 1 (indicating that the node is selected), otherwise the output is 0 (indicating that the node is not selected). This operation can be formalized as:
[0075] x 0 =I(z>σ)
[0076] Among them, I is the indicator function, and the output is converted to floating point type to be compatible with subsequent calculations. In this way, through the straight-through estimator STE and the diffusion model M(x 0 ,G;θ), an end-to-end propagation agent model can be constructed to map the node's continuous state vector z to the final infection probability through the end-to-end propagation agent model, that is:
[0077] y=x T-1 =M s (z,G;θ)
[0078] In the backpropagation phase, since the threshold operation is inherently non-differentiable (i.e., its derivative is undefined at the threshold point and zero elsewhere), directly calculating the gradient will hinder optimization. To this end, STE ignores the discontinuity of this operation, assumes that the discretization in the forward propagation has no effect on the gradient, and directly passes the output gradient to the input. When the model calculates the loss function as Loss according to the optimization objective, the gradient is calculated as follows:
[0079]
[0080] This pass-through strategy allows gradients to flow through discrete layers, enabling the continuous node state vector z to be efficiently updated via gradient descent.
[0081] Furthermore, we design a scalable optimization objective. We use the end-to-end proxy model M trained with historical propagation data. s, which can estimate the propagation influence of node combinations in the target scenario and infer the optimal node set using gradient descent according to the optimization objective. By expanding the form of the optimization objective, the STEIM of the present invention can solve various IM variants in different scenarios. The expanded optimization objective is to find an optimal combination of seed nodes, under the expanded node constraint conditions to maximize the estimated value of the expanded propagation influence objective of this model , that is:
[0082]
[0083]
[0084] x 0 = I(z > σ)
[0085] where V is the node set, and V i is the i-th node. is the generalized propagation influence objective, for example, by setting the target set of expected influence, maximizing the influence of the seed set on the target population. Various algorithms have been proposed for the target diffusion problem in the prior art, such as the reverse local path algorithm, which defines the propagation ability of nodes to the target nodes. is the generalized budget constraint applicable to a single node. For example, the budget can be set as different centrality metrics of the nodes, and K is the actual budget. For the IM problem of setting the node degree as the budget
[64] , can be set as ‖x·A‖1, is the adjacency matrix of the network G, and ‖x·A‖1 < K represents that the L1-norm of the total seed node degree is restricted by the budget K.
[0086] Using the penalty function method, the constraint condition is added as a penalty term to the loss function, and the loss function used in the final optimization process is:
[0087]
[0088] where, is the negative value of the objective function, making the estimated value of the generalized propagation influence objective larger; μ > 0 is the penalty coefficient, used to control the importance of the constraint; ensures that the constraint is satisfied as much as possible, and the loss increases when deviating from the constraint.
[0089] During the optimization process, freeze the parameters θ of the forward diffusion model M s , and iteratively update the continuous state vector z of the seed nodes by the gradient descent method. This reverse optimization utilizes M sThe gradient of (z,G;θ) with respect to z enables the model to explore the continuous space of node selection probabilities and converge to an optimal configuration that maximizes influence.
[0090] In summary, the present invention proposes:
[0091] 1. A seed node inference method based on an end-to-end propagation agent model, which efficiently infers the seed node set under multiple practical constraints through continuous optimization.
[0092] 2. A method for relaxing the initial state of a node from a discrete space to a continuous probability space and combining it with a straight-through estimator for optimization, which is used to perform differentiable optimization of a set of seed nodes in a continuous space.
[0093] 3. A task-adaptable seed node inference mechanism that supports various application scenarios such as node budget and influence target changes, improving the flexibility and scalability of the inference stage.
[0094] In this example, 10 real-world network datasets were collected, including Soc-Dolphins, Celegans, Fb-Pages-Food, Cora, Ego-Facebook, Ca-GrQc, Wiki-Vote, Deezer-Europe, Cit-HepPh, and Soc-Douban. In a typical IM scenario, the proposed method, STARIM, was compared with several existing methods, including centrality methods such as H-index and NumCycle; improved heuristic methods such as VoteRank and VoteRank++; methods based on reverse reachable sets such as OPIM and SubSim; methods for solving the general CO problem based on reinforcement learning such as S2V-DQN; methods for node importance ranking based on supervised learning such as RCNN and MRCNN; methods for solving IM problems based on reinforcement learning such as ToupleGDD; and methods based on latent space optimization such as DeepIM.
[0095] STARIM was experimented under three different settings: STAR-M is a propagation agent model based on direct estimation and mechanism propagation, STAR-N is a propagation agent model based on direct estimation and using trajectory data to train a neural network, and GSO-M is a propagation agent model based on Gumbel-Softmax and mechanism propagation.
[0096] Table 1 below shows the IM performance of different methods on multiple datasets under the IC propagation model. The values in the table represent the percentage of nodes ultimately infected by the seed sets found by each method at different seed ratios. The STARIM framework models information propagation based on the mechanism data generated by the IC propagation model and combines it with direct estimation to construct a robust end-to-end proxy model. By directly optimizing the selection of node sets through gradient descent, more stable and efficient experimental results are achieved.
[0097] Table 1 Performance on IC propagation model
[0098]
[0099] Table 2 below shows the IM performance of different methods on multiple datasets under the SIR propagation model. Overall, the STARIM framework still achieves a significant advantage in this scenario.
[0100] Table 2 Performance on the SIR propagation model
[0101]
[0102] In the IM scenario with enhanced budget constraints, by replacing Customizing node centrality as the node selection cost allows the selection of seed sets to maximize influence under the expanded node budget. STAR-M was compared with the typical solution BCT. As shown in Table 3, STAR-M was compared with the typical method BCT under the IC propagation model using node centrality as the budget constraint. Considering that the influence of a node in real scenarios is usually positively correlated with the number of its fans, the node out-degree d out As the selection cost of a single node, the total budget constraint is set to in Where ρ is the average out-degree of nodes in the network, and ρ is set to 0.02, 0.04, 0.06, and 0.08. The budget-constrained IM capabilities of different methods were quantified by comparing the proportion of nodes ultimately infected by initially infected nodes. STAR-M exhibits significant advantages over BCT, indicating that the STARIM framework can be extended to the budget-constrained influence maximization scenario.
[0103] Table 3 IM with node centrality constraints
[0104]
[0105] In the scenario of maximizing target influence, by Setting it as the sum of the infection probabilities of the target node set can make the optimization direction of the seed set move towards maximizing the influence on the target node set. STAR-M is compared with two target diffusion problem solutions, RLP and Katz index. Figure 2 As shown, under the SIR propagation model, the ratio ρ of the initial infected nodes is set to 0.01 (as shown in Figure 2 (a)-(e) in Figure 2) and 0.05 (as shown in Figure 2) Figure 2 As shown in (f)-(j) of the figure, the target node set ratio is set to 0.1-0.5. The influence of target nodes by different methods is quantified by comparing the proportion of target nodes ultimately infected by the initially infected node. Katzindex, because it considers the more global nature of the node's relationship to the target node set, achieves better experimental results than RLP. STAR-M shows significant advantages over the other two methods, demonstrating that the STARIM framework can be extended to scenarios where target influence is maximized.
[0106] Example 2
[0107] In addition to the above method, an embodiment of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the decision-making behavior decision-making method according to various embodiments of the present application described in the above embodiment 1 of this specification.
[0108] The computer program product may be written in any combination of one or more programming languages to implement the program code for performing the operations of the embodiments of the present application, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0109] Example 3
[0110] An embodiment of the present application may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, causes the processor to execute the steps of the decision-making behavior decision-making method according to various embodiments of the present application described in the above-mentioned embodiment 1.
[0111] The computer-readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0112] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. An influence maximization method based on direct and autoregressive influence estimation, characterized in that: include: Given a propagation network G, the propagation dynamics is modeled based on the propagation trajectory data of the nodes in the network, and the propagation model M(x,G; θ), used to model the node state evolution process, x represents the node state, and θ is the model parameter; Use the propagation model M(x,G;θ) to transform the initial state x of the node into 0 Mapped to the final state y of the spread, that is, the final infection probability of the node, and the initial state x of the node 0 Mapped to a continuous state vector z, construct an end-to-end propagation agent model y = M s (z,G;θ); Fixed model parameters θ, with the goal of maximizing influence, the initial state x of the node is quantified by gradient descent method. 0 The corresponding continuous state vector z is optimized, and the seed node set is obtained according to the optimization result.
2. The influence maximization method based on direct and autoregressive influence estimation according to claim 1, characterized in that: The node propagation trajectory data consists of infected nodes and their corresponding infection time; the node state trajectory x at different time steps is divided using the time window 0 ,…x t ,…,x T-1 ; Node state trajectory at time step t in, represents the state of node i at time step t; According to the evolution of node states, we use graph neural networks to model the given propagation network G and construct a propagation model M(x, G; θ) through autoregression. The model parameter θ is calculated as follows: Where E[·] represents the expectation; p θ (x t+1 ∣x t ,…,x 0 ,G) represents the propagation model in x t ,…,x 0 Under the conditions of and G, generate x t+1 probability.
3. The influence maximization method based on direct and autoregressive influence estimation according to claim 1 or 2, characterized in that: The propagation model is constructed based on the propagation mechanism, and the node status trajectory is updated as follows: Among them, f(·) and g(·) are functions of the propagation model, N(·) represents the neighbor set of the node, and Δt is the time step size.
4. The influence maximization method based on direct and autoregressive influence estimation according to claim 1 or 2, characterized in that: The propagation model is built based on a neural network, and the evolution of node states is modeled using a gated recurrent unit combined with a graph neural network, as shown below: Where N(·) represents the node’s neighbor set; is the hidden state of the gated recurrent unit; represents the hidden state of node i at time step t; represents edge-level MLPs, f out (·) denotes the MLPs of the output layer; At time step t, the graph neural network updates its own hidden representation by aggregating the hidden representations of neighboring nodes and its own state. The output layer predicts the state of the node in the next time step based on the hidden representation of the node.
5. The influence maximization method based on direct and autoregressive influence estimation according to claim 1, characterized in that: Using the propagation model M(x,G;θ) and through autoregression, the initial state x of the node is transformed into 0 Mapping to the propagation final state y = x T-1 =M(x 0 ,G;θ), the initial state of the node x 0 From discrete values, namely binary vectors, to continuous values, namely continuous state vector z, the continuous state vector z of a node represents the probability of each node being selected, and an end-to-end propagation agent model y = M is constructed. s (z, G; θ); fix the model parameter θ, set the propagation final state y with the goal of maximizing influence, and optimize the initial state x of the node by gradient descent method 0 The corresponding continuous state vector z is mapped back to a discrete binary vector through the straight-through estimator STE, thereby determining the initial state of the node and obtaining the seed node set.
6. The influence maximization method based on direct and autoregressive influence estimation according to claim 1 or 5, characterized in that: Add node budget as a constraint.
7. A readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed, the influence maximization method based on direct and autoregressive influence estimation according to any one of claims 1 to 6 is implemented.
8. A computer program product, characterized in that It includes a computer program / instruction, which, when executed by a processor, implements the influence maximization method based on direct and autoregressive influence estimation according to any one of claims 1 to 6.
Citation Information
Patent Citations
Influence maximization method having unwanted users under independent cascade model (IC)
CN108596777A
Influence maximization method and system for sequential network
CN113378470A
Influence maximization seed node set selection method and device
CN114065914A
Node identification method and device, electronic equipment and computer readable storage medium
CN115134247A
Method and system for generating target text sequence for natural language text sequence
CN117556787A