Node solution method and device for the problem of maximizing social network influence
By calculating the spread influence and selection cost of nodes in social networks, node grouping and optimization operations are carried out, the problem of seed node aggregation is solved, and the quality and influence of node sets are maximized.
Patent Information
- Application Number
- CN202211330371.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-27
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-10-27
AI Technical Summary
The prior art tends to aggregation when selecting seed nodes in social networks, resulting in low quality of seed node sets and failure to effectively consider the selection cost differences between different nodes.
By obtaining social network data sets, calculate the propagation influence and selection cost of nodes, sort the node importance, group and initialize the population, generate new individuals using cross-operations and mutated operations, calculate the fitness, and select the individual with the highest fitness as the optimal node set to avoid the aggregation of seed nodes.
This improves the probability of selecting high-influence nodes within the cost range, avoids seed node aggregation, and improves the quality of node sets.
Smart Images

Figure CN115689794B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of social networks, and in particular to a node solution method and device for the problem of maximizing social network influence. Background Art
[0002] With the development of the Internet, online social platforms have developed rapidly, providing people with extremely convenient and fast communication channels. At the same time, relying on a large user base and convenient communication means, such platforms have gradually developed into the main promotion platforms for new products and new concepts. The existence of an efficient dissemination method in these social networks is largely due to the "word-of-mouth effect" phenomenon in social networks. This special information dissemination phenomenon has brought great application potential to the field of viral marketing. Therefore, the problem of maximizing influence (IM) based on the business background of viral marketing has become a key algorithm problem in the study of information dissemination.
[0003] The goal of the influence maximization problem is to find the most influential k users (seed nodes) in the social network according to the input total budget, and the total selection cost of these k users does not exceed the total budget. When the information spreads in the network through these k users, the expected number of affected users in the network can be maximized. The Independent Cascade (IC) model is a classic information diffusion model. When the information dissemination process is modeled using the independent cascade model, the influence maximization problem is proven to be an NP-hard problem. When solving this problem, evaluating the dissemination influence of the seed node set is one of the most critical steps. Traditional methods tend to use a greedy strategy to select the seed node set and use the Monte Carlo simulation method to calculate the dissemination influence of the seed node set in the network during the process of selecting the seed node set. This simulation method requires a large number of simulations of the information dissemination process, so it is very time-consuming. At the same time, due to the lack of optimization of the seed node set in the greedy strategy, this leads to the aggregation of the seed nodes obtained by the algorithm based on the greedy strategy, resulting in a low quality of the seed node set. On the other hand, when dealing with the problem of node selection cost, traditional methods usually set the node selection cost to the same value, that is, they do not consider the situation that there may be different selection costs between different nodes. However, in reality, when a company or group wants to find users in the social network to promote products, the incentives required for different users are generally different. Therefore, it is more realistic to have a relevant selection cost for each node in the social network. Therefore, not considering the situation that there may be different selection costs between different nodes when solving the influence maximization problem will lead to a very limited application scenario of the algorithm.
[0004] The prior art discloses a method and system for maximizing the influence of groups in a social network, including: Step 1: In the social network, through the random walk method, map nodes to the representation space and retain the influence propagation attributes of the nodes; Step 2: Define and calculate the propagation affinity between nodes, and successively merge adjacent node pairs with the highest propagation affinity until a coarsened network is obtained that meets the set compression ratio, where each node corresponds to a group in the original network; Step 3: According to the attributes of the influence of nodes propagating within a group and across groups, construct an influence propagation function for the candidate seed set, and select the maximum influence user set containing a preset number of nodes according to the greedy algorithm. Since the prior art uses the greedy algorithm to select the seed node set, the nodes are prone to aggregation, resulting in a low quality of the seed node set. Summary of the Invention
[0005] The object of the present invention is to provide a node solution method and device for the problem of maximizing the influence of a social network, so as to solve the problem that the division of the existing seed node set in the above background technology is prone to aggregation, resulting in a low quality of the seed node set.
[0006] To achieve the above object, the present invention provides a node solution method for the problem of maximizing the influence of a social network, including:
[0007] S1. Obtain a social network data set, use the users in the social network data set as nodes in the social network, calculate the importance of the nodes from the propagation influence and selection cost of the nodes, sort the nodes in descending order of importance to obtain a node sequence nodelist, and group each node starting from the first node in the node sequence nodelist to obtain multiple node groups;
[0008] S2. Set a total budget and a first candidate node set, the first candidate node set is initially an empty node set, starting from the first node in each node group, successively select loadsize nodes in the node group, and the total cost of the loadsize nodes does not exceed the total budget, and then randomly select one node from these loadsize nodes to join the first candidate node set. After all the nodes obtained by completing the above operations for all the node groups obtained in step S1 are added to the first candidate node set;
[0009] S3. Population initialization: Randomly select a node from the first candidate node set and add it to individual i. At the same time, form the first shielding node set with the nodes of individual i, reset the first candidate node set to an empty node set, and starting from the first node of each node group, sequentially select loadsize nodes that do not exist in the first shielding node set, and the total cost of the loadsize nodes does not exceed the total budget. Then randomly select a node from these loadsize nodes and add it to the first candidate node set. After all the nodes obtained by performing the above operations on all the node groups obtained in step S1 are added to the first candidate node set, repeat all the above operations in this step until the total cost of all the nodes in individual i does not exceed the total budget and no new nodes can be added, and obtain multiple individuals in the same process as obtaining individual i. All the obtained individuals form a population;
[0010] S4. Each individual in the population generates new individuals through crossover operations or mutation operations, and a new population is composed of multiple new individuals;
[0011] S5. Calculate the fitness of each new individual, and select the individuals with high fitness from the population and the new population through the binary tournament selection method to form the next-generation population, and select the individual with the highest fitness in the next-generation population as the current optimal individual;
[0012] S6. Repeat S4 and S5 until the iteration ends. Select the individual with the highest fitness from all the current optimal individuals as the historical optimal individual, and the node set represented by this historical optimal individual is the user combination that maximizes the social network influence and meets the cost budget.
[0013] Preferably, in step S1, the propagation influence is the expected propagation value of the node set within two hops, and the propagation influence calculation formula is as follows:
[0014]
[0015] Where NB one (S) and NB two (S) respectively represent the sets of one-hop and two-hop neighbor nodes of the seed node set S; δ(S, v) represents the expected propagation value of the seed node set S to node v, N in (v) represents the set of nodes with edges pointing to node v, and p(u, v) represents the propagation probability of node u to node v. When calculating, substitute the node set S into the formula, and at this time the node set only contains a single node v.
[0016] Preferably, in step S1, the importance is the ratio of the node propagation influence to the node selection cost, and the formula is as follows:
[0017]
[0018] Among them, f(v) represents the propagation influence of node v, which is calculated using the propagation influence calculation formula; cost(v) represents the selection cost of node v, that is, the selection cost of node v is determined by the input cost function cost.
[0019] Preferably, in step S4, a new individual is generated through the crossover operator, and the process is as follows:
[0020] S4-1. First, randomly select an individual from the population as the target individual P, select half of the nodes from the target individual P to form a new individual Pnew, randomly select two individuals Pr1 and Pr2 from the population, and form a second candidate node set with the individuals Pr1, Pr2, and the target individual P;
[0021] S4-2. Delete the nodes that are repeated with the new individual Pnew from the second candidate node set. B represents the total budget, cost(Pnew) represents the total cost of the nodes in the new individual Pnew, and the remaining budget RB of the new individual Pnew = B - cost(Pnew). Load cs nodes from the second candidate node set according to the remaining budget RB of the new individual Pnew, and the cost of each node does not exceed the remaining budget RB;
[0022] S4-3. Select the node with the maximum distance from the new individual Pnew from the loaded cs nodes and add it to the new individual Pnew while updating the remaining budget RB;
[0023] S4-4. Repeat S4-2 and S4-3 until the remaining budget RB is exhausted, then the crossover operator ends, and finally the new individual Pnew is obtained.
[0024] Preferably, in step S4, a new individual is generated through the mutation operator, and the process is as follows:
[0025] S4-1. First, randomly select an individual from the population as the target individual P, and randomly select n nodes from the target individual P to form a new individual Pnew. Among them, the randomly selected number of nodes n adopts an adaptive value-taking method, that is, the value of n will increase with the increase of the running generation number, and the formula is as follows:
[0026] n = g / maxG · |P|
[0027] Among them, g represents the current generation number, maxG represents the maximum generation number, and |P| represents the number of nodes of the target individual P;
[0028] S4-2. Let B denote the total budget, and cost(Pnew) denote the total cost of the nodes in the new individual Pnew. Therefore, the remaining budget RB of the new individual Pnew is RB = B - cost(Pnew). Select multiple nodes from all individuals in the population whose costs do not exceed the remaining budget RB to form the third candidate node set and the fourth candidate node set;
[0029] S4-3. Randomly select two nodes from the combination of the third candidate node set and the fourth candidate node set, and select the node with the maximum distance from the new individual Pnew from these two nodes to join the new individual Pnew and update the remaining budget RB simultaneously;
[0030] S4-4. Repeat S4-2 and S4-3 until the remaining budget RB is exhausted, then the mutation operator ends, and finally the new individual Pnew is obtained.
[0031] Preferably, the third candidate node set can be generated in the following way: Obtain m node groups through the same process as in step S1, set the total budget and the first shielding node set, where the first shielding node set is initially an empty node set. Add the nodes in the target individual P that have not formed the new individual Pnew to the first shielding node set. Then, starting from the first node of each node group, sequentially select loadsize nodes in this node group that do not exist in the first shielding node set, and the total cost of the loadsize nodes does not exceed the total budget. Then randomly select one node from these loadsize nodes. Perform the above operations for each of the m node groups to obtain m nodes, and the cost of each of the m nodes does not exceed the remaining budget RB of the new individual Pnew. The m nodes form the third candidate node set.
[0032] Preferably, the fourth candidate node set can be generated in the following way: Form the second shielding node set from all individuals in the population. Obtain m node groups through the same process as in step S1, set the total budget. Starting from the first node of each node group, sequentially select loadsize nodes in this node group that do not exist in the second shielding node set, and the total cost of the loadsize nodes does not exceed the total budget. Then randomly select one node from these loadsize nodes. The m nodes obtained after performing the above operations for each of the m node groups form the fourth candidate node set.
[0033] Preferably, when selecting the node with the maximum distance from the new individual Pnew among the cs nodes loaded in step S4-3, or when randomly selecting two nodes from the third candidate node set in step S4-3 and selecting the node with the maximum distance from the new individual Pnew from them to join the new individual Pnew and update the remaining budget RB simultaneously, the formula for calculating the node distance is as follows:
[0034]
[0035] In the formula, NB two (v) represents the set of neighbor nodes of node v within two-hop range, where the one-hop neighbors of a node refer to the nodes directly connected to the node, and the two-hop neighbors refer to the neighbors of the node's neighbors.
[0036] Preferably, since the propagation influence of nodes is mainly concentrated within two-hop range, in step S5, the expected propagation value within two-hop range of the seed node set used by each newly generated individual is used as the fitness, and the formula is as follows:
[0037]
[0038] Where NB one (S) and NB two (S) respectively represent the sets of one-hop and two-hop neighbor nodes of the seed node set S, and δ(S, v) represents the expected propagation value of the seed node set S to node v. N in (v) represents the set of nodes with edges pointing to node v, and p(u, v) represents the propagation probability of node u to node v. When calculating, the node set S is substituted into the formula.
[0039] This application also proposes a computer device, which includes a memory and a processor. The memory is used to store a computer program. When the computer program is executed by the processor, the processor executes the node grouping method for the problem of maximizing social network influence according to any one of claims 1-9.
[0040] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0041] The present invention selects multiple most influential users from a social network dataset as nodes, calculates the importance of the nodes based on the propagation influence and selection cost of the nodes, sorts the nodes in descending order of importance and divides them into node groups, sets a total budget, starts from the first node of each node group, sequentially selects multiple nodes of the node group, and the total cost of the nodes does not exceed the total budget, then randomly selects nodes from them to join the first candidate node set, randomly selects a node from the first candidate node set to join an individual, and at the same time forms the first blocked node set with the nodes of the individual, reset the first candidate node set to an empty node set, and start from the first node of each node group, sequentially select multiple nodes of the node group that do not exist in the first blocked node set, and the total cost of the nodes does not exceed the total budget, then randomly select a node from them to join the first candidate node set. After all node groups complete the above operations, all the obtained nodes are added to the first candidate node set. Repeat the above operations until the total cost of all nodes in the individual does not exceed the total budget and no new nodes can be added, to obtain multiple individuals by the same process for an individual, and all the obtained individuals form a population. Each individual in the population generates new individuals through crossover operations or mutation operations, and multiple new individuals form a new population. Calculate the fitness of each new individual, select individuals with high fitness from the population and the new population through binary tournament selection to form the next-generation population, and select the individual with the highest fitness in the next-generation population as the current optimal individual until the iteration ends. Select the individual with the highest fitness from all the current optimal individuals as the historical optimal individual. The node set represented by this historical optimal individual is the user combination that maximizes the social network influence within the cost budget. In each node group, the nodes are arranged in descending order of importance. Since the importance of the nodes is determined by the influence and cost of the nodes, the selection probability of high-influence nodes within the cost range is increased. By using the crossover operator or mutation operator, the situation where seed nodes will aggregate is avoided, and the fitness of the individual can continuously improve during the optimization process, and finally the obtained node set can have higher quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is a flowchart of a node solution method for the social network influence maximization problem proposed in an embodiment of the present invention;
[0043] Figure 2 is a flowchart of the main method of an embodiment of the present invention;
[0044] Figure 3 is a flowchart of step S1 of an embodiment of the present invention;
[0045] Figure 4 is a flowchart of step S2 of an embodiment of the present invention;
[0046] Figure 5 is a flowchart of step S3 of an embodiment of the present invention;
[0047] Figure 6 is a flowchart of the crossover operation in step S4 of an embodiment of the present invention;
[0048] Figure 7 is a flowchart of the mutation operation in step S4 of an embodiment of the present invention. Detailed implementation manners
[0049] The following further describes in detail the specific implementation manners of the present invention in conjunction with the drawings and embodiments. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0050] In the description of the present invention, it should be noted that the terms "center", "longitudinal", "transverse", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0051] In the description of the present invention, it should be noted that unless otherwise clearly specified and defined, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0052] In addition, in the description of the present invention, unless otherwise stated, the meaning of "a plurality of" is two or more.
[0053] Embodiment 1
[0054] As Figure 1-7 shown, the node solution method for the problem of maximizing social network influence in a preferred embodiment of the present invention includes:
[0055] S1. Obtain a social network dataset, use the users in the social network dataset as nodes in the social network, calculate the importance of the nodes based on the propagation influence and selection cost of the nodes, sort the nodes in descending order of importance to obtain a node sequence nodelist, and group each node starting from the first node in the node sequence nodelist to obtain a plurality of node groups;
[0056] In step S1, the propagation influence is the expected propagation value of the node set within two hops, and the calculation formula for the propagation influence is as follows:
[0057]
[0058] where NB one (S) and NB two (S) respectively represent the sets of one-hop and two-hop neighbor nodes of the seed node set S; δ(S, v) represents the expected propagation value of the seed node set S for node v, N in (v) represents the set of nodes with edges pointing to node v, and p(u, v) represents the propagation probability of node u for node v. When calculating, the node set S is substituted into the formula, and at this time the node set only contains a single node v;
[0059] The importance is the ratio of the node propagation influence to the node selection cost, and the formula is as follows:
[0060]
[0061] where f(v) represents the propagation influence of node v, which is calculated using the propagation influence calculation formula; cost(v) represents the selection cost of node v, that is, the selection cost of node v is determined by the input cost function cost. It can be seen from the formula that the larger this ratio is, the higher the importance of the node;
[0062] In this embodiment, through the method of node grouping, nodes with different ICR values are evenly distributed in m node groups. At the same time, the nodes in each node group are arranged in descending order according to the ICR value. In the node grouping operation, the node set needs to be divided into m node groups in total. These m node groups are represented by the node group list Groups = [g0, g1,..., gi,...., gm-1], and each node group stores nodes in an array structure. Before grouping the node set, the node set is sorted in descending order according to the node ICR value to obtain an ordered list nodelist, where nodelist[i] represents the (i + 1)-th node stored in nodelist. The steps for grouping the nodes are as follows: Starting from the first node in the list nodelist, each node is grouped in turn, and the node nodelist[i] will be assigned to the node group gg(i), where g(i) = i % m. Therefore, after grouping the nodes through the above operations, the nodes in each node group will be arranged in descending order according to the ICR value.
[0063] S2. Set the total budget and the first candidate node set. The first candidate node set is initially an empty node set. Starting from the first node of each node group, select loadsize nodes of this node group in sequence, and the total cost of the loadsize nodes does not exceed the total budget. Then randomly select one node from these loadsize nodes and add it to the first candidate node set. After all the node groups obtained in step S1 complete the above operations, all the obtained nodes are added to the first candidate node set;
[0064] After m node groups complete the above operations, a first candidate node set containing m candidate nodes is obtained. It can be seen from this that when the number of node groups m remains unchanged, the larger the load number of nodes loadsize, the larger the range of node search each time. At the same time, since the nodes in the node group are arranged in descending order of ICR value, the higher the ICR value, the higher the probability of being selected.
[0065] S3. Population initialization: Randomly select one node from the first candidate node set and add it to individual i. At the same time, form the first shield node set with the nodes of individual i, reset the first candidate node set to an empty node set, and starting from the first node of each node group, select loadsize nodes that do not exist in the first shield node set in sequence, and the total cost of the loadsize nodes does not exceed the total budget. Then randomly select one node from these loadsize nodes and add it to the first candidate node set. After all the node groups obtained in step S1 complete the above operations, all the obtained nodes are added to the first candidate node set. Repeat all the above operations in this step until the total cost of all nodes in individual i does not exceed the total budget and no new nodes can be added, to obtain multiple individuals in the same process as individual i. All the obtained individuals form a population;
[0066] Each individual in the population represents a node set and the individuals use an integer encoding method, that is, each individual stores the node numbers in the form of a set. Given the total budget B, the purpose of the population initialization operation is to generate a population Pop containing NP individuals, and the total cost of the seed set represented by each individual in the population does not exceed the total budget B. In the population initialization operation, Groups = [g0, g1,..., gi,...., gm - 1] represents the m node groups obtained through the node grouping operation, the initial remaining budget RB = B, and individual i (0 ≤ i < NP).
[0067] S4. Each individual in the population generates new individuals through crossover operations, and a new population is composed of multiple new individuals;
[0068] Generate new individuals through the crossover operator. The process is as follows:
[0069] S4-1. First, randomly select an individual from the population as the target individual P. Select half of the nodes from the target individual P to form a new individual Pnew. Randomly select two individuals Pr1 and Pr2 from the population, and form a second candidate node set with the individuals Pr1, Pr2, and the target individual P.
[0070] S4-2. Delete the nodes that are repeated with the new individual Pnew from the second candidate node set. Let B represent the total budget, and cost(Pnew) represent the total cost of the nodes in the new individual Pnew. The remaining budget RB of the new individual Pnew is RB = B - cost(Pnew). Load cs nodes from the second candidate node set according to the remaining budget RB of the new individual Pnew, and the cost of each node does not exceed the remaining budget RB.
[0071] S4-3. Select the node with the maximum distance from the new individual Pnew from the loaded cs nodes and add it to the new individual Pnew while updating the remaining budget RB.
[0072] S4-4. Repeat S4-2 and S4-3 until the remaining budget RB is exhausted, then the crossover operator ends, and finally obtain the new individual Pnew.
[0073] Since the crossover operator generates the new individual Pnew through local search and uses the distance measurement method in the process of selecting nodes for the new individual Pnew, it avoids the situation where the nodes in the new individual Pnew gather.
[0074] In step S4-2, before the new individual Pnew selects nodes from the second candidate node set, the second candidate node set has already deleted the nodes that are repeated with the new individual Pnew. Therefore, the new individual Pnew will not waste the remaining budget due to selecting duplicate nodes during the process of selecting nodes.
[0075] Among the cs nodes loaded in step S4-3, when selecting the node with the maximum distance from the new individual Pnew, or in step S4-3, randomly select two nodes from the third candidate node set and select the node with the maximum distance from the new individual Pnew and add it to the new individual Pnew while updating the remaining budget RB, the formula for calculating the node distance is as follows:
[0076]
[0077] In the formula, NB two (v) represents the set of neighbor nodes of node v within two-hop range. Among them, the one-hop neighbor of a node refers to the node directly connected to the node, and the two-hop neighbor refers to the neighbor of the node's neighbor. From the definition of this formula, it can be seen that when the number of common neighbor nodes between node v and the seed set S within two-hop range is less, the distance distance(v, S) between node v and the node set S is greater.
[0078] In the process of selecting nodes from the new individual Pnew, it can be seen that the mutation operation generates a candidate node set in two ways, realizes searching for candidate nodes in a larger range, generates a new individual Pnew through global search, and adopts an adaptive method in the process of selecting nodes from the new individual Pnew, so that the new individual Pnew generated by the mutation operator can jump out of the local optimum.
[0079] S5. Calculate the fitness of each new individual, select the individuals with high fitness from the population and the new population through the binary tournament selection method to form the next generation population, and select the individual with the highest fitness in the next generation population as the current optimal individual;
[0080] Since the propagation influence of nodes is mainly concentrated within two-hop range, in step S5, the expected propagation value of the seed node set within two-hop range of each newly generated individual is used as the fitness, and the formula is as follows:
[0081]
[0082] Among them, NB one (S) and NB two (S) respectively represent the sets of one-hop and two-hop neighbor nodes of the seed node set S, δ(S, v) represents the expected propagation value of the seed node set S to node v, N in (v) represents the set of nodes with edges pointing to node v, p(u, v) represents the propagation probability of node u to node v, and the node set S is substituted into the formula when calculating.
[0083] S6. Repeat S4 and S5 until the iteration ends, select the individual with the highest fitness from all current optimal individuals as the historical optimal individual, and the node set represented by this historical optimal individual is the user combination that maximizes the influence of the social network meeting the cost budget.
[0084] Embodiment 2
[0085] The difference between this embodiment and Embodiment 1 is that in step S4, each individual in the population generates a new individual through the mutation operation, and a new population is composed of multiple new individuals;
[0086] Generate a new individual through the mutation operator, and the process is as follows:
[0087] S4-1. First, randomly select an individual from the population as the target individual P, and randomly select n nodes from the target individual P to form a new individual Pnew. Among them, the randomly selected number of nodes n adopts an adaptive value-taking method, that is, the value of n will increase with the increase of the running generation number, and the formula is as follows:
[0088] n = g / maxG · |P|
[0089] Among them, g represents the current generation number, maxG represents the maximum generation number, and |P| represents the number of nodes of the target individual P;
[0090] S4-2: B represents the total budget, and cost(Pnew) represents the total cost of the nodes in the new individual Pnew. Therefore, the remaining budget RB of the new individual Pnew is RB = B - cost(Pnew). Select multiple nodes with costs not exceeding the remaining budget RB from all individuals in the population to form the third candidate node set and the fourth candidate node set;
[0091] S4-3: Randomly select two nodes from the combination of the third candidate node set and the fourth candidate node set, and select the node with the largest distance from the new individual Pnew from these two nodes to join the new individual Pnew and update the remaining budget RB at the same time;
[0092] S4-4: Repeat S4-2 and S4-3 until the remaining budget RB is exhausted, then the mutation operator ends, and finally the new individual Pnew is obtained;
[0093] In step S4-2, the third candidate node set can be generated in the following way: Obtain m node groups through the same process as in step S1, set the total budget and the first shielding node set. The first shielding node set is initially an empty node set. Add the nodes in the target individual P that have not formed the new individual Pnew to the first shielding node set. Then, starting from the first node of each node group, sequentially select loadsize nodes in this node group that do not exist in the first shielding node set, and the total cost of the loadsize nodes does not exceed the total budget. Then randomly select one node from these loadsize nodes. Perform the above operations on m node groups to obtain m nodes, and the cost of each node in the m nodes does not exceed the remaining budget RB of the new individual Pnew. The m nodes form the third candidate node set;
[0094] In step S4-2, the fourth candidate node set can be generated in the following way: The second shielding node set is composed of all individuals in the population. Obtain m node groups through the same process as in step S1, set the total budget. Starting from the first node of each node group, sequentially select loadsize nodes in this node group that do not exist in the second shielding node set, and the total cost of the loadsize nodes does not exceed the total budget. Then randomly select one node from these loadsize nodes. The m nodes obtained after m node groups complete the above operations form the fourth candidate node set;
[0095] The other processes of this embodiment are the same as those of Embodiment 1 and will not be elaborated here.
[0096] Embodiment 3
[0097] In this embodiment, a computer device is proposed. The computer device includes a memory and a processor. The memory is used to store a computer program. When the computer program is executed by the processor, the processor executes the method for grouping problem nodes aiming at maximizing the influence of the social network according to any one of claims 1-9.
[0098] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and substitutions can be made, and these improvements and substitutions should also be regarded as the protection scope of the present invention.
Claims
1. A node solution method for the problem of maximizing social network influence, characterized in that, Including: S1. Obtain a social network dataset. Use the users in the social network dataset as nodes in the social network. Calculate the importance of each node based on the propagation influence and selection cost of the node. Sort the nodes in descending order of importance to obtain a node sequence nodelist, and start from the first node in the node sequence nodelist to group each node to obtain multiple node groups; In step S1, the propagation influence is the expected propagation value of the node set within two hops. The formula for calculating the propagation influence is as follows: Among them, NB one (S) and NB two (S) respectively represent the sets of one-hop and two-hop neighbor nodes of the seed node set S; δ(S, v) represents the expected propagation value of the seed node set S to the node v, represents the set of nodes with edges pointing to the node v, and p(u, v) represents the propagation probability of the node u to the node v. When calculating, the node set S is substituted into the formula, and at this time the node set only contains a single node v; S2. Set the total budget and the first candidate node set. The first candidate node set is initially an empty node set. Starting from the first node of each node group, select loadsize nodes in sequence, and the total cost of the loadsize nodes does not exceed the total budget. Then randomly select one node from these loadsize nodes and add it to the first candidate node set. After all the node groups obtained in step S1 complete the above operations, all the obtained nodes are added to the first candidate node set; S3. Population initialization: Randomly select a node from the first candidate node set and add it to individual i. At the same time, form the first shielding node set with the nodes of individual i. Reset the first candidate node set to an empty node set. Starting from the first node of each node group, select loadsize nodes that do not exist in the first shielding node set in sequence, and the total cost of the loadsize nodes does not exceed the total budget. Then randomly select one node from these loadsize nodes and add it to the first candidate node set. After all the node groups obtained in step S1 complete the above operations, all the obtained nodes are added to the first candidate node set. Repeat all the above operations in this step until the total cost of all the nodes in individual i does not exceed the total budget and no new nodes can be added, to obtain multiple individuals in the same process as individual i. All the obtained individuals form a population; S4. Each individual in the population generates new individuals through crossover operations or mutation operations, and multiple new individuals form a new population; S5. Calculate the fitness of each new individual. Select the individuals with high fitness from the population and the new population through the binary tournament selection method to form the next-generation population, and select the individual with the highest fitness in the next-generation population as the current optimal individual; S6. Repeat S4 and S5 until the iteration ends. Select the individual with the highest fitness from all the current optimal individuals as the historical optimal individual. The node set represented by this historical optimal individual is the user combination that maximizes the social network influence meeting the cost budget.
2. The node solution method for the social network influence maximization problem according to claim 1, characterized in that, In step S1, the importance is the ratio of the node propagation influence to the node selection cost. The formula is as follows: Where f(v) represents the propagation influence of node v, which is calculated using the propagation influence calculation formula; cost(v) represents the selection cost of node v, that is, the selection cost of node v is determined by the input cost function cost.
3. The node solution method for the social network influence maximization problem according to claim 1, wherein In step S4, new individuals are generated through the crossover operator. The process is as follows: S4-1. First, randomly select an individual from the population as the target individual P. Select half of the nodes from the target individual P to form a new individual Pnew. Randomly select two individuals Pr1 and Pr2 from the population, and combine the individuals Pr1, Pr2, and the target individual P to form a second candidate node set; S4-2. Delete the nodes that are repeated with the new individual Pnew from the second candidate node set. B represents the total budget, and cost(Pnew) represents the total cost of the nodes in the new individual Pnew. The remaining budget RB of the new individual Pnew is RB = B - cost(Pnew). Load cs nodes from the second candidate node set according to the remaining budget RB of the new individual Pnew, and the cost of each node does not exceed the remaining budget RB; S4-3. Select the node with the maximum distance from the new individual Pnew from the loaded cs nodes and add it to the new individual Pnew while updating the remaining budget RB; S4-4. Repeat S4-2 and S4-3 until the remaining budget RB is exhausted, then the crossover operator ends, and finally a new individual Pnew is obtained.
4. The node solution method for the social network influence maximization problem according to claim 1, wherein In step S4, a new individual is generated through the mutation operator, and the process is as follows: S4-1. First, randomly select an individual from the population as the target individual P. Randomly select n nodes from the target individual P to form a new individual Pnew. Among them, the randomly selected number of nodes n adopts an adaptive value-taking method, that is, the value of n will increase as the number of running generations increases. The formula is as follows: n = g / maxG · |P| where g represents the current generation, maxG represents the maximum number of generations, and |P| represents the number of nodes of the target individual P; S4-2. B represents the total budget, and cost(Pnew) represents the total cost of the nodes in the new individual Pnew. Therefore, the remaining budget RB of the new individual Pnew is RB = B - cost(Pnew). Select multiple nodes with costs not exceeding the remaining budget RB from all individuals in the population to form a third candidate node set and a fourth candidate node set; S4-3. Randomly select two nodes from the combination of the third candidate node set and the fourth candidate node set, and select the node with the maximum distance from the new individual Pnew from these two nodes and add it to the new individual Pnew while updating the remaining budget RB; S4-4. Repeat 4-2 and S4-3 until the remaining budget RB is exhausted, then the mutation operator ends, and finally a new individual Pnew is obtained.
5. The node solution method for the problem of maximizing social network influence according to claim 4, characterized in that, The third candidate node set can be generated in the following way: Obtain m node groups through the same process as in step S1, set the total budget and the first shielding node set, where the first shielding node set is initially an empty node set. Add the nodes in the target individual P that have not formed the new individual Pnew to the first shielding node set. Then, starting from the first node of each node group, sequentially select loadsize nodes in the node group that do not exist in the first shielding node set, and the total cost of the loadsize nodes does not exceed the total budget. Then, randomly select one node from these loadsize nodes. Perform the above operations on all m node groups to obtain m nodes, and the cost of each node in the m nodes does not exceed the remaining budget RB of the new individual Pnew. The m nodes form the third candidate node set.
6. The node solution method for the social network influence maximization problem according to claim 4, characterized in that, The fourth candidate node set can be generated in the following way: Form a second shielding node set from all individuals in the population. Obtain m node groups through the same process as in step S1, set the total budget. Starting from the first node of each node group, sequentially select loadsize nodes in the node group that do not exist in the second shielding node set, and the total cost of the loadsize nodes does not exceed the total budget. Then, randomly select one node from these loadsize nodes. The m nodes obtained after performing the above operations on all m node groups form the fourth candidate node set.
7. The node solving method for the social network influence maximization problem according to claim 3 or 4, characterized in that, Among the cs nodes loaded in step S4-3, select the node with the maximum distance from the new individual Pnew, or randomly select two nodes from the third candidate node set in step S4-3 and select the node with the maximum distance from the new individual Pnew to add to the new individual Pnew while updating the remaining budget RB. The formula for calculating the node distance is as follows: where NB two (v) represents the set of neighbor nodes of node v within two-hop range, where the one-hop neighbors of a node refer to the nodes directly connected to the node, and the two-hop neighbors refer to the neighbors of the node's neighbors.
8. The node solution method for the social network influence maximization problem according to claim 1, characterized in that Since the propagation influence of nodes is mainly concentrated within two hops, in step S5, the expected propagation value within two hops of the seed node set used by each newly generated individual is used as the fitness, and the formula is as follows: Among them, NB one (S) and NB two (S) respectively represent the sets of one-hop and two-hop neighbor nodes of the seed node set S. δ(S, v ) represents the expected propagation value of the seed node set S to the node v, represents the set of nodes with edges pointing to the node v. p(u, v) represents the propagation probability of the node u to the node v. When calculating, the node set S is substituted into the formula.
9. A computer device, characterized in that: The computer device includes a memory and a processor. The memory is used to store a computer program. When the computer program is executed by the processor, the processor executes the node grouping method for the problem of maximizing social network influence according to any one of claims 1-8.
Citation Information
Patent Citations
Social network influence maximizing method based on culture gene algorithm
CN104361462A