Influence maximization method based on deep learning and community division

By applying deep learning and community division models in social networks, the problems of high time complexity of seed node selection and insufficient generalization ability in the existing technology are solved, and an impact maximization algorithm with more efficient and strong generalization ability is achieved.

CN119941427APending Publication Date: 2025-05-06QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510081908.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The time complexity of the prior art when looking for seed nodes in social networks is too high, and can only be solved in attribute networks or non-attribute networks, making it difficult to generalize on complex data.

Method used

A community division model based on deep learning is adopted to divide nodes in social networks through graph neural networks, and a node attribute matrix is ​​generated in combination with Node2Vec method to reduce time complexity and improve generalization capabilities.

Benefits of technology

The efficiency of the influence maximization algorithm is improved, and better generalization capabilities can be achieved on complex data. The quality of selected seed nodes has been improved, and the scalability and efficiency have been significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FDA0005248908880000011
    Figure FDA0005248908880000011
  • Figure FDA0005248908880000033
    Figure FDA0005248908880000033
  • Figure FDA0005248908880000043
    Figure FDA0005248908880000043
Patent Text Reader

Abstract

The invention discloses an influence maximization method based on deep learning and community division, relates to the technical field of social networks, and combines a community division model based on deep learning with a traditional community division influence maximization algorithm to form a new influence maximization algorithm. In the community division process, according to the method, attributes are generated for nodes through node embedding, and an automatic encoder, modularity and a self-training module are fused. The node embedding not only enables the model to be suitable for a common network, but also can be used for solving an influence maximization problem in an attribute network, so that the universality is improved. And during node selection, a sampling method is adopted, so that the accuracy is ensured, and the solving efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of social network, and in particular to a method for maximizing influence based on deep learning and community division. Background Art

[0002] In real life, groups in different fields such as business, politics and society use social networks as the main medium to spread information. These groups adopt specific strategies to expand their influence in order to maximize their interests. Influence maximization means selecting appropriate nodes (i.e. seed nodes) in social networks to spread information, so that the scope and depth of information dissemination are maximized, thereby achieving the best dissemination effect.

[0003] However, in the field of influence maximization, due to the large number of nodes in social networks, the time complexity of existing technologies in finding seed nodes is too high, resulting in a long search process and loss of timeliness. At the same time, existing influence maximization algorithms can only be solved in attribute networks or non-attribute networks.

[0004] Since social networks have become the most important platform for people to connect and communicate with each other, and also the core position for information dissemination, maximizing influence plays an important role in improving the effectiveness of marketing activities, promoting social change, or accelerating the dissemination of health information in the field of public health. However, how to improve the quality of seed node sets while further improving efficiency and generalization ability is a problem that deserves further research. Summary of the invention

[0005] In order to overcome the deficiencies of the above technologies, the present invention provides a method for reducing the time complexity of the influence maximization algorithm and achieving generalization on complex data.

[0006] The technical solution adopted by the present invention to overcome the technical problems is:

[0007] A method for maximizing influence based on deep learning and community division, including:

[0008] a) Obtain n users and m relationships between them, generate graph data G, G = {V, E}, V is a node set, V = {v1, v2, ..., v i ,...,v n},v i is the i-th user, i∈{1,...,n}, E is the edge set, E={e1,e2,...,e i ,...,e m},e i is the i-th edge, i∈{1,...,m}, e i represents the uth user vu With the vth user v v There is a relationship between them, u∈{1,...,n}, v∈{1,...,n}, u≠v, P is the edge attribute set, P={p1,p2,...,p i ,...,p m}, p i For the uth user v u With the vth user v v The probability of successfully sharing goods between them, d′ is the i-th edge e i The number of edges that share the same endpoint;

[0009] b) Generate a node attribute matrix x based on the edge attribute set P;

[0010] c) constructing an adjacency matrix adjl according to the edge set E, and constructing a proximity matrix M according to the adjacency matrix adjl;

[0011] d) Construct a community division model consisting of an encoder and a decoder;

[0012] e) Input the node attribute matrix x, the adjacency matrix adjl, and the proximity matrix M into the encoder of the community partition model, and output the encoding result z;

[0013] f) Input the encoding result z into the decoder of the community division model, and output the node label sequence arr;

[0014] g) Obtain the important community set C based on the node label sequence arr s , C s ={C s1 ,C s2 ,...,C si ,...,C sf}, C si is the i-th important community, i∈{1,...,f}, f is the number of important communities;

[0015] h) Calculate the number of important communities C in the i-th si The number of users k selected i ;

[0016] i) According to the number of users k i Generate a reverse reachable set, and get the i-th important community C according to the reverse reachable set si The seed set S i , all f important communities constitute the final seed set S.

[0017] Furthermore, in step a), n users and m relationships between the n users are obtained from the email-Eu-core network dataset.

[0018] Furthermore, in step b), the edge attribute set P is calculated using the Node2vec algorithm to obtain a node attribute matrix x with n rows and a columns.

[0019] Further, step c) comprises the following steps:

[0020] c-1) Generate an n*n matrix A, where each element value in the matrix A is 0;

[0021] c-2) If the uth user v u With the vth user v v There is a relationship between i , then modify the element value of the u-th row and v-th column in matrix A to 1, and obtain matrix A′;

[0022] c-3) Modify all element values ​​in the diagonal of the matrix A′ from the upper left element to the lower right element to 1, and obtain the matrix A″;

[0023] c-4) Normalize the matrix A″ using the L1 norm to obtain the adjacency matrix adjl;

[0024] c-5) Through the formula M = adjl 2 +adjl calculates the proximity matrix M.

[0025] Further, step e) comprises the following steps:

[0026] e-1) The encoder is composed of the first image attention network and the second image attention network. Both the first image attention network and the second image attention network are composed of an attention score unit, a LeakyReLU activation function, and a softmax function;

[0027] e-2) Input the node attribute matrix x, the adjacency matrix adjl, and the proximity matrix M into the first graph attention network of the encoder, and initialize a weight matrix W1 with a rows and b columns according to the node attribute matrix x through the Xavier initialization method. Multiply the node attribute matrix x and the weight matrix W1 to obtain a linear transformation matrix h0 with n rows and b columns;

[0028] e-3) Initialize the matrix attn_self with b rows and 1 column and the matrix attn_neighs with b rows and 1 column respectively through the Xavier initialization method, input the linear transformation matrix h0, the matrix attn_self and the matrix attn_neighs into the attention score unit of the first graph attention network, multiply the linear transformation matrix h0 with the matrix attn_self to obtain the attention score s1 of the self with n rows and 1 column, and multiply the linear transformation matrix h0 with the matrix attn_neighs to obtain the attention score l1 of the neighborhood with n rows and 1 column;

[0029] e-4) Transpose the attention score l1 of the neighborhood to obtain the attention score of the neighborhood with 1 row and n columns after transposition Copy the 1 column data of the attention score s1 of n rows and 1 column to n-1 columns to get the attention score s1′ of n rows and n columns, and then transpose the attention score of the neighborhood of 1 row and n columns. Copy n-1 rows of 1 row of data to get the attention score of the neighborhood of n rows and n columns The attention score s1′ of the self and the attention score of the neighboring domain Perform the addition operation to obtain the initial attention score matrix attn_dense with n rows and n columns 01 , the initial attention score matrix attn_dense 01 Multiply it with the proximity matrix M to get the attention matrix attn_dense 11 , the attention matrix attn_dense 11 Input into the LeakyReLU activation function of the attention network of the first figure, and output the attention matrix attn_dense1;

[0030] e-5) Create an n-row and n-column attention score feature matrix attention 01 , if the value of the i-th row and j-th column of the adjacency matrix adjl is greater than 0, then the attention score feature matrix attention 01 The value of the i-th row and j-th column of is equal to the value of the i-th row and j-th column of the attention matrix attn_dense1. If the value of the i-th row and j-th column of the adjacency matrix adjl is less than or equal to 0, then the attention score feature matrix attention 01 The value of the i-th row and j-th column of the attention score feature matrix is ​​set to 0, 01 Input into the softmax function of the attention network of the first figure, and output the normalized attention score feature matrix attention1 with n rows and n columns;

[0031] e-6) Multiply the normalized attention score feature matrix attention1 with the linear transformation matrix h0 to obtain the n-row and b-column node feature representation h_prime;

[0032] e-7) Input the node feature representation h_prime, the adjacency matrix adjl, and the proximity matrix M into the second graph attention network of the encoder, initialize the node feature representation h_prime using the Xavier initialization method to obtain a b-row and c-column weight matrix W2, and multiply the node feature representation h_prime with the weight matrix W2 to obtain an n-row and c-column linear transformation matrix h1;

[0033] e-8) Initialize the matrix attn_self2 with c rows and 1 column and the matrix attn_neighs2 with c rows and 1 column respectively by the Xavier initialization method, input the linear transformation matrix h1, the matrix attn_self2 and the matrix attn_neighs2 into the attention score unit of the second graph attention network, multiply the linear transformation matrix h1 with the matrix attn_self to obtain the attention score s2 of the self with n rows and 1 column, and multiply the linear transformation matrix h1 with the matrix attn_neighs2 to obtain the attention score l2 of the neighborhood with n rows and 1 column;

[0034] e-9) Transpose the attention score l2 of the neighborhood to obtain the attention score of the neighborhood with 1 row and n columns after transposition Copy the 1 column data of the attention score s2 of n rows and 1 column to n-1 columns to get the attention score s2′ of n rows and n columns, and then transpose the attention score of the neighborhood of 1 row and n columns Copy n-1 rows of 1 row of data to get the attention score of the neighborhood of n rows and n columns The attention score s2′ of the self and the attention score of the neighboring domain Perform the addition operation to obtain the initial attention score matrix attn_dense with n rows and n columns 02 , the initial attention score matrix attn_dense 02 Multiply it with the proximity matrix M to get the attention matrix attn_dense 12 , the attention matrix attn_dense 12 Input into the LeakyReLU activation function of the attention network of the second figure, and output the attention matrix attn_dense2;

[0035] e-10) Create an n-row and n-column attention score feature matrix attention 02, if the value of the i-th row and j-th column of the adjacency matrix adjl is greater than 0, then the attention score feature matrix attention 02 The value of the i-th row and j-th column of is equal to the value of the i-th row and j-th column of the attention matrix attn_dense2. If the value of the i-th row and j-th column of the adjacency matrix adjl is less than or equal to 0, then the attention score feature matrix attention 02 The value of the i-th row and j-th column of the attention score feature matrix is ​​set to 0, 02 Input into the softmax function of the attention network of the second figure, and output the normalized attention score feature matrix attention2 with n rows and n columns;

[0036] e-11) Multiply the normalized attention score feature matrix attention2 with the linear transformation matrix h0 to obtain the encoding result z.

[0037] Further, step f) comprises the following steps:

[0038] f-1) Input the encoding result z into the decoder of the community division model and transpose the encoding result z to obtain the transposed matrix z T ;

[0039] f-2) Compare the encoding result z with the transposed matrix z T Perform multiplication operation to obtain matrix A_pred0;

[0040] f-3) Through the formula The prediction matrix A_pred1 is calculated and input into the Sigmoid function to obtain the prediction matrix A_pred.

[0041] f-4) Using the encoding result z, the cluster center kmeans is calculated by the k-means algorithm;

[0042] f-5) Use the encoding result z and the cluster center kmeans to obtain the similarity matrix q of n rows and Y columns through t distribution, q iy For the i-th user v i The probability of being assigned to community y, i∈{1,...,n}, y∈{1,...,Y}, Y is the number of communities, and the i-th row [q i1 ,q i2 ,...,q iy ,...,q iY ] is the one with the largest probability value as the i-th user v i Community tags c i, all n users’ community labels constitute a node label sequence arr, arr = {c1, c2, ..., c i ,...,c n}.

[0043] Furthermore, the method further comprises the following steps:

[0044] f-6) Through the formula Calculate the element value of the i-th row and y-th column in the target distribution matrix p;

[0045] f-7) Use the target distribution matrix p and the similarity matrix q to obtain the divergence loss L through the divergence loss function c ;

[0046] f-8) Calculate the loss L between the adjacency matrix adjl and the prediction matrix A_pred through the cross entropy loss function R ;

[0047] f-9) Calculate the modularity Q of the adjacency matrix adjl and the prediction matrix A_pred by the modularity calculation formula;

[0048] f-10) By the formula L=L R -βQ+γL c The total loss function L is calculated, where β and γ are both hyperparameters;

[0049] f-11) Use the total loss function L to train the community partition model through the Adam optimizer to obtain the optimized community partition model.

[0050] Further, step g) comprises the following steps:

[0051] g-1) According to the i-th user v i Community tags c i The i-th user v i Join the community corresponding to the community label to obtain the community set C, C = {C1, C2, ..., C i ,...,C Y}, C i is the i-th community, i∈{1,...,Y}, each community contains a number of users;

[0052] g-2) Determine the i-th community C i The number of users n i Is it greater than or equal to If yes, then the i-th community C i For important communities C si , all f important communities constitute the important community set C s , C s ={Cs1 ,C s2 ,...,C si ,...,C sf}, where k is the number of pre-selected seeds.

[0053] Further, in step h), by formula Calculate the number of important communities C si The number of users k selected i , where n j is the jth important community C sj The number of users in , j∈{1,...,f}.

[0054] Further, step i) comprises the following steps:

[0055] i-1) Set the error tolerance ε, 0<ε<1, and set the parameter l, l≥1;

[0056] i-2) Through the formula Calculate the quantity θ i , from the jth important community C sj Select θ i Each user generates a reverse reachable set RR;

[0057] i-3) is the jth important community C sj Initialize a seed set S i , S i =φ, φ is an empty set;

[0058] i-4) All n users {v1,v2,...,v i ,...,v n Substitute into formula F R (S i ∪v i )-F R (S i ) and add the user corresponding to the maximum difference into the seed set S i In the formula, F R (·) is the number of RRs covering the reverse reachable set;

[0059] i-5) Judge F R (S i ) is greater than or equal to If yes, then execute step i-6), if not, repeat step i-4) until F R (S i ) is greater than or equal to

[0060] i-6) Through the formula Calculate the quantity θ and use the quantity θ to replace the quantity θ i , update the reverse reachable set RR;

[0061] i-7) Repeat step i-4) until the seed set S i Contains k i Users;

[0062] i-8) The seed sets of all f important communities constitute the final seed set S, The beneficial effect of the present invention is that a community partitioning model based on deep learning is combined with a traditional community partitioning influence maximization algorithm to form a new influence maximization algorithm. In the process of community partitioning, the method uses node embedding to generate attributes for nodes, and integrates autoencoders, modularity and self-training modules. Node embedding not only makes the model applicable to ordinary networks, but can also be used in attribute networks to solve the influence maximization problem, thereby improving versatility. When selecting nodes, a sampling method is used to improve the solution efficiency while ensuring accuracy. DETAILED DESCRIPTION

[0063] The present invention is further described below.

[0064] A method for maximizing influence based on deep learning and community division, including:

[0065] a) Obtain n users and m relationships between them, generate graph data G, G = {V, E}, V is a node set, V = {v1, v2, ..., v i ,...,v n},v i is the i-th user, i∈{1,...,n}, E is the edge set, E={e1,e2,...,e i ,...,e m},e i is the i-th edge, i∈{1,...,m}, e i represents the uth user v u With the vth user v v There is a relationship between them, u∈{1,...,n}, v∈{1,...,n}, u≠v, P is the edge attribute set, P={p1,p2,...,p i ,...,p m}, p i For the uth user v u With the vth user v v The probability of successfully sharing goods between them, d′ is the number of edges e i The number of edges that share the same endpoint.

[0066] b) Generate the node attribute matrix x based on the edge attribute set P.

[0067] c) Construct an adjacency matrix adjl based on the edge set E, and construct a proximity matrix M based on the adjacency matrix adjl.

[0068] d) Construct a community partition model consisting of an encoder and a decoder.

[0069] e) Input the node attribute matrix x, the adjacency matrix adjl, and the proximity matrix M into the encoder of the community partition model, and output the encoded result z.

[0070] f) Input the encoding result z into the decoder of the community partitioning model and output the node label sequence arr.

[0071] g) Obtain the important community set C based on the node label sequence arr s , C s ={C s1 ,C s2 ,...,C si ,...,C sf}, C si is the i-th important community, i∈{1,...,f}, and f is the number of important communities.

[0072] h) Calculate the number of important communities C in the i-th si The number of users k selected i .

[0073] i) According to the number of users k i Generate a reverse reachable set, and get the i-th important community C according to the reverse reachable set si The seed set S i , all f important communities constitute the final seed set S.

[0074] The present invention maps nodes to representation space through the Node2Vec method in a social network and retains the structural properties of the node neighborhood. According to the structure of the graph, the nodes in the graph are divided into communities using a graph neural network (GNN), and each node is divided into a specified community. According to the divided communities, important communities are selected, and quotas are allocated according to a given number of seeds. Sampling is performed in each important community, and the reverse reachable set of the extracted nodes is obtained, and a group of nodes with the highest reverse reachable set coverage score in the community are used as seed nodes of the community. All seed nodes together constitute the user set with the greatest influence.

[0075] In order to verify the reliability of the patented method, the expansion degree that can be achieved by the finally selected seed nodes is compared with the existing influence maximization algorithm based on community division, as shown in Table 1.

[0076] Table 1 Comparison with CIM algorithm in terms of scalability

[0077] Seed quantity\method Method of the present invention CIM 5 98.3 93.6 10 146.7 142.0 15 188.4 185.7 20 210.1 204.4 25 245.9 244.6

[0078] At the same time, in order to compare the efficiency of the algorithm, it is also compared with the CIM algorithm, as shown in Table II.

[0079] Table 2 Comparison with CIM algorithm in terms of time

[0080] Seed quantity\method Method of the present invention C1M 5 23.49s 1301.45s 10 23.65s 1292.51s 15 24.13s 1286.88s 20 25.01s 1283.15s 25 25.83s 1279.72s

[0081] It can be seen from Table 1 and Table 2 that the traditional influence maximization algorithm based on community division is compared in our experimental method. The greater the expansion, the better the result, and the smaller the time, the higher the efficiency. It can be seen from the data table that the method of the present invention has certain improvements in different aspects when selecting different quantities, which is better than the previous method. In terms of expansion, the expansion of the method of the present invention is improved by about 2% compared with the CIM algorithm, and the time has been significantly improved. It shows that this method has certain advantages, the selected seeds are relatively ideal, and especially the efficiency has been significantly improved.

[0082] In one embodiment of the present invention, in step a), n users and m relationships between the n users are obtained from the email-Eu-core network dataset.

[0083] In one embodiment of the present invention, in step b), the edge attribute set P is calculated using the Node2vec algorithm to obtain a node attribute matrix x with n rows and a columns.

[0084] In one embodiment of the present invention, step c) comprises the following steps:

[0085] c-1) Generate an n*n matrix A, in which each element value is 0.

[0086] c-2) If the uth user v u With the vth user v v There is a relationship between i , then modify the element value of the u-th row and v-th column in matrix A to 1, and obtain matrix A′.

[0087] c-3) Add a self-loop to the matrix A′. Specifically, change the values ​​of all elements in the diagonal from the upper left element to the lower right element of the matrix A′ to 1 to obtain the matrix A″.

[0088] c-4) Normalize the matrix A″ using the L1 norm to obtain the adjacency matrix adjl.

[0089] c-5) Through the formula M = adjl 2 +adjl calculates the proximity matrix M.

[0090] In one embodiment of the present invention, step e) comprises the following steps:

[0091] e-1) The encoder consists of a first-image attention network and a second-image attention network. Both the first-image attention network and the second-image attention network consist of an attention score unit, a LeakyReLU activation function, and a softmax function.

[0092] e-2) Input the node attribute matrix x, adjacency matrix adjl, and proximity matrix M into the first graph attention network of the encoder, and initialize a weight matrix W1 with a rows and b columns according to the node attribute matrix x through the Xavier initialization method. Multiply the node attribute matrix x and the weight matrix W1 to obtain a linear transformation matrix h0 with n rows and b columns.

[0093] e-3) Initialize the matrix attn_self with b rows and 1 column and the matrix attn_neighs with b rows and 1 column respectively through the Xavier initialization method, input the linear transformation matrix h0, the matrix attn_self and the matrix attn_neighs into the attention score unit of the first graph attention network, multiply the linear transformation matrix h0 with the matrix attn_self to obtain the attention score s1 of itself with n rows and 1 column, and multiply the linear transformation matrix h0 with the matrix attn_neighs to obtain the attention score l1 of the neighborhood with n rows and 1 column.

[0094] e-4) Transpose the attention score l1 of the neighborhood to obtain the attention score of the neighborhood with 1 row and n columns after transposition Copy the 1 column of data of the attention score s1 of n rows and 1 column to n-1 columns to get n rows

[0095] The attention score s1′ of n columns itself is converted into the attention score of the neighborhood of 1 row and n columns after transposition. Copy n-1 rows of 1 row of data to get the attention score of the neighborhood of n rows and n columns The attention score s1′ of the self and the attention score of the neighboring domain Perform the addition operation to obtain the initial attention score matrix attn_dense with n rows and n columns 01 , the initial attention score matrix attn_dense 01 Multiply it with the proximity matrix M to get the attention matrix attn_dense 11 , the attention matrix attn_dense 11Input into the LeakyReLU activation function of the attention network of the first figure, and output the attention matrix attn_dense1.

[0096] e-5) Create an n-row and n-column attention score feature matrix attention 01 , if the value of the i-th row and j-th column of the adjacency matrix adjl is greater than 0, then the attention score feature matrix attention 01 The value of the i-th row and j-th column of is equal to the value of the i-th row and j-th column of the attention matrix attn_dense1. If the value of the i-th row and j-th column of the adjacency matrix adjl is less than or equal to 0, then the attention score feature matrix attention 01 The value of the i-th row and j-th column of the attention score feature matrix is ​​set to 0, 01 Input into the softmax function of the attention network of the first figure, and output the normalized attention score feature matrix attention1 with n rows and n columns.

[0097] e-6) Multiply the normalized attention score feature matrix attention1 with the linear transformation matrix h0 to obtain the n-row and b-column node feature representation h_prime.

[0098] e-7) Input the node feature representation h_prime, the adjacency matrix adjl, and the proximity matrix M into the second graph attention network of the encoder, and initialize a b-row and c-column weight matrix W2 according to the node feature representation h_prime through the Xavier initialization method. Multiply the node feature representation h_prime with the weight matrix W2 to obtain an n-row and c-column linear transformation matrix h1.

[0099] e-8) Initialize the matrix attn_self2 with c rows and 1 column and the matrix attn_neighs2 with c rows and 1 column respectively through the Xavier initialization method, input the linear transformation matrix h1, the matrix attn_self2 and the matrix attn_neighs2 into the attention score unit of the second graph attention network, multiply the linear transformation matrix h1 with the matrix attn_self to obtain the attention score s2 of itself with n rows and 1 column, and multiply the linear transformation matrix h1 with the matrix attn_neighs2 to obtain the attention score l2 of the neighborhood with n rows and 1 column.

[0100] e-9) Transpose the attention score l2 of the neighborhood to obtain the attention score of the neighborhood with 1 row and n columns after transposition Copy the 1 column data of the attention score s2 of n rows and 1 column to n-1 columns to get the attention score s2′ of n rows and n columns, and then transpose the attention score of the neighborhood of 1 row and n columns Copy n-1 rows of 1 row of data to get the attention score of the neighborhood of n rows and n columns The attention score s2′ of the self and the attention score of the neighboring domain Perform the addition operation to obtain the initial attention score matrix attn_dense with n rows and n columns 02 , the initial attention score matrix attn_dense 02 Multiply it with the proximity matrix M to get the attention matrix attn_dense 12 , the attention matrix attn_dense 12 Input into the LeakyReLU activation function of the attention network of the second figure, and output the attention matrix attn_dense2.

[0101] e-10) Create an n-row and n-column attention score feature matrix attention 02 , if the value of the i-th row and j-th column of the adjacency matrix adjl is greater than 0, then the attention score feature matrix attention 02 The value of the i-th row and j-th column of is equal to the value of the i-th row and j-th column of the attention matrix attn_dense2. If the value of the i-th row and j-th column of the adjacency matrix adjl is less than or equal to 0, then the attention score feature matrix attention 02 The value of the i-th row and j-th column of the attention score feature matrix is ​​set to 0, 02 Input into the softmax function of the attention network of the second graph, and output the normalized attention score feature matrix attention2 with n rows and n columns.

[0102] e-11) Multiply the normalized attention score feature matrix attention2 with the linear transformation matrix h0 to obtain the encoding result z.

[0103] In one embodiment of the present invention, step f) comprises the following steps:

[0104] f-1) Input the encoding result z into the decoder of the community division model and transpose the encoding result z to obtain the transposed matrix z T .

[0105] f-2) Compare the encoding result z with the transposed matrix z T Perform multiplication operation to obtain matrix A_pred0.

[0106] f-3) Through the formula The prediction matrix A_pred1 is calculated and input into the Sigmoid function to obtain the prediction matrix A_pred.

[0107] f-4) Use the encoding result z to calculate the cluster center kmeans through the k-means algorithm.

[0108] f-5) Use the encoding result z and the cluster center kmeans to obtain the similarity matrix q of n rows and Y columns through t distribution, q iy For the i-th user v i The probability of being assigned to community y, i∈{1,...,n}, y∈{1,...,Y}, Y is the number of communities, and the i-th row [q i1 ,q i2 ,...,q iy ,...,q iY ] is the one with the largest probability value as the i-th user v i Community tags c i , the community labels of all n users constitute the node label sequence arr, arr = {c1, c2, ..., c i ,...,c n}.

[0109] In one embodiment of the present invention, the following steps are also included:

[0110] f-6) Through the formula Calculate the element value of the i-th row and y-th column in the target distribution matrix p.

[0111] f-7) Use the target distribution matrix p and the similarity matrix q to obtain the divergence loss L through the divergence loss function c f-8) Calculate the loss L between the adjacency matrix adjl and the prediction matrix A_pred through the cross entropy loss function R .

[0112] f-9) The modularity calculation formula is used to calculate the modularity Q of the adjacency matrix adjl and the prediction matrix A_pred.

[0113] f-10) By the formula L=L R -βQ+γL c The total loss function L is calculated, where β and γ are hyperparameters.

[0114] f-11) Use the total loss function L to train the community partition model through the Adam optimizer to obtain the optimized community partition model.

[0115] In one embodiment of the present invention, step g) comprises the following steps:

[0116] g-1) According to the i-th user v i Community tags c i The i-th user v i Join the community corresponding to the community label to obtain the community set C, C = {C1, C2, ..., C i ,...,C Y}, C i is the i-th community, i∈{1,...,Y}, and each community contains several users.

[0117] g-2) Determine the i-th community C i The number of users n i Is it greater than or equal to If yes, then the i-th community C i For important communities C si , all f important communities constitute the important community set C s , C s ={C s1 ,C s2 ,…,C si ,…,C sf}, where k is the number of pre-selected seeds.

[0118] In one embodiment of the present invention, in step h), the formula Calculate the number of important communities C si The number of users k selected i , where n j is the jth important community C sj The number of users in , j∈{1,…,f}.

[0119] In one embodiment of the present invention, step i) comprises the following steps:

[0120] i-1) Set the error tolerance ε, 0<ε<1, and set the parameter l, l≥1.

[0121] i-2) Through the formula Calculate the quantity θ i , from the jth important community C sj Select θ i Each user generates a reverse reachable set RR.

[0122] i-3) is the jth important community C sj Initialize a seed set S i , S i =φ, φ is the empty set.

[0123] i-4) All n users {v1,v2,...,v i ,…,v n Substitute into formula F R (S i ∪v i )-F R (S i ) and add the user corresponding to the maximum difference into the seed set S i In the formula, F R (·) is the number of covered reverse reachable sets RR.

[0124] i-5) Judge F R (S i ) is greater than or equal to If yes, then execute step i-6), if not, repeat step i-4) until F R (S i ) is greater than or equal to

[0125] i-6) Through the formula Calculate the quantity θ and use the quantity θ to replace the quantity θ i , update the reverse reachable set RR;

[0126] i-7) Repeat step i-4) until the seed set S i Contains k i Users;

[0127] i-8) The seed sets of all f important communities constitute the final seed set S, Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for maximizing influence based on deep learning and community division, characterized in that: include: a) Obtain n users and m relationships between them, generate graph data G, G = {V, E}, V is a node set, V = {v1, v2, ..., v i ,...,v n },v i is the i-th user, i∈{1,...,n}, E is the edge set, E={e1,e2,...,e i ,...,e m },e i is the i-th edge, i∈{1,...,m}, e i represents the uth user v u With the vth user v v There is a relationship between them, u∈{1,…,n}, v∈{1,…,n}, u≠v, P is the edge attribute set, P={p1,p2,…,p i ,…,p m }, p i For the uth user v u With the vth user v v The probability of successfully sharing goods between them, d′ is the number of edges e i The number of edges that share the same endpoint; b) Generate a node attribute matrix x based on the edge attribute set P; c) constructing an adjacency matrix adjl according to the edge set E, and constructing a proximity matrix M according to the adjacency matrix adjl; d) Construct a community division model consisting of an encoder and a decoder; e) Input the node attribute matrix x, the adjacency matrix adjl, and the proximity matrix M into the encoder of the community partition model, and output the encoding result z; f) Input the encoding result z into the decoder of the community division model, and output the node label sequence arr; g) Obtain the important community set C based on the node label sequence arr s , C s ={C s1 ,C s2 ,...,C si ,...,C sf }, C si is the i-th important community, i∈{1,…,f}, f is the number of important communities; h) Calculate the number of important communities C in the i-th si The number of users k selected i ; i) According to the number of users k i Generate a reverse reachable set, and get the i-th important community C according to the reverse reachable set si The seed set S i , all f important communities constitute the final seed set S.

2. The method for maximizing influence based on deep learning and community division according to claim 1, characterized in that: In step a), n users and m relationships between n users are obtained from the email-Eu-core network dataset.

3. The method for maximizing influence based on deep learning and community division according to claim 1, characterized in that: In step b), the edge attribute set P is calculated using the Node2vec algorithm to obtain a node attribute matrix x with n rows and a columns.

4. The method for maximizing influence based on deep learning and community division according to claim 1, characterized in that: Step c) comprises the following steps: c-1) Generate an n*n matrix A, where each element value in the matrix A is 0; c-2) If the uth user v u With the vth user v v There is a relationship between i , then modify the element value of the u-th row and v-th column in matrix A to 1, and obtain matrix A′; c-3) Modify all element values ​​in the diagonal of the matrix A′ from the upper left element to the lower right element to 1, and obtain the matrix A″; c-4) Normalize the matrix A″ using the L1 norm to obtain the adjacency matrix adjl; c-5) Through the formula M = adjl 2 +adjl calculates the proximity matrix M.

5. The method for maximizing influence based on deep learning and community division according to claim 3 is characterized in that: Step e) comprises the following steps: e-1) The encoder is composed of the first image attention network and the second image attention network. Both the first image attention network and the second image attention network are composed of an attention score unit, a LeakyReLU activation function, and a softmax function; e-2) Input the node attribute matrix x, the adjacency matrix adjl, and the proximity matrix M into the first graph attention network of the encoder, and initialize a weight matrix W1 with a rows and b columns according to the node attribute matrix x through the Xavier initialization method. Multiply the node attribute matrix x and the weight matrix W1 to obtain a linear transformation matrix h0 with n rows and b columns; e-3) Initialize the matrix attn_self with b rows and 1 column and the matrix attn_neighs with b rows and 1 column respectively through the Xavier initialization method, input the linear transformation matrix h0, the matrix attn_self and the matrix attn_neighs into the attention score unit of the first graph attention network, multiply the linear transformation matrix h0 with the matrix attn_self to obtain the attention score s1 of the self with n rows and 1 column, and multiply the linear transformation matrix h0 with the matrix attn_neighs to obtain the attention score l1 of the neighborhood with n rows and 1 column; e-4) Transpose the attention score l1 of the neighborhood to obtain the attention score of the neighborhood with 1 row and n columns after transposition Copy the 1 column data of the attention score s1 of n rows and 1 column to n-1 columns to get the attention score s1′ of n rows and n columns, and then transpose the attention score of the neighborhood of 1 row and n columns. Copy n-1 rows of 1 row of data to get the attention score of the neighborhood of n rows and n columns The attention score s1′ of the self and the attention score of the neighboring domain Perform the addition operation to obtain the initial attention score matrix attn_dense with n rows and n columns 01 , the initial attention score matrix attn_dense 01 Multiply it with the proximity matrix M to get the attention matrix attn_dense 11 , the attention matrix attn_dense 11 Input into the LeakyReLU activation function of the attention network of the first figure, and output the attention matrix attn_dense1; e-5) Create an n-row and n-column attention score feature matrix attention 01 , if the value of the i-th row and j-th column of the adjacency matrix adjl is greater than 0, then the attention score feature matrix attention 01 The value of the i-th row and j-th column of is equal to the value of the i-th row and j-th column of the attention matrix attn_dense1. If the value of the i-th row and j-th column of the adjacency matrix adjl is less than or equal to 0, then the attention score feature matrix attention 01 The value of the i-th row and j-th column of the attention score feature matrix is ​​set to 0, 01 Input into the softmax function of the attention network of the first figure, and output the normalized attention score feature matrix attention1 with n rows and n columns; e-6) Multiply the normalized attention score feature matrix attention1 with the linear transformation matrix h0 to obtain the n-row and b-column node feature representation h_prime; e-7) Input the node feature representation h_prime, the adjacency matrix adjl, and the proximity matrix M into the second graph attention network of the encoder, initialize the node feature representation h_prime using the Xavier initialization method to obtain a b-row and c-column weight matrix W2, and multiply the node feature representation h_prime with the weight matrix W2 to obtain an n-row and c-column linear transformation matrix h1; e-8) Initialize the matrix attn_self2 with c rows and 1 column and the matrix attn_neighs2 with c rows and 1 column respectively by the Xavier initialization method, input the linear transformation matrix h1, the matrix attn_self2 and the matrix attn_neighs2 into the attention score unit of the second graph attention network, multiply the linear transformation matrix h1 with the matrix attn_self to obtain the attention score s2 of the self with n rows and 1 column, and multiply the linear transformation matrix h1 with the matrix attn_neighs2 to obtain the attention score l2 of the neighborhood with n rows and 1 column; e-9) Transpose the attention score l2 of the neighborhood to obtain the attention score of the neighborhood with 1 row and n columns after transposition Copy the 1 column data of the attention score s2 of n rows and 1 column to n-1 columns to get the attention score s2′ of n rows and n columns, and then transpose the attention score of the neighborhood of 1 row and n columns. Copy n-1 rows of 1 row of data to get the attention score of the neighborhood of n rows and n columns The attention score s2′ of the self and the attention score of the neighboring domain Perform the addition operation to obtain the initial attention score matrix attn_dense with n rows and n columns 02 , the initial attention score matrix attn_dense 02 Multiply it with the proximity matrix M to get the attention matrix attn_dense 12 , the attention matrix attn_dense 12 Input into the LeakyReLU activation function of the attention network of the second figure, and output the attention matrix attn_dense2; e-10) Create an n-row and n-column attention score feature matrix attention 02 , if the value of the i-th row and j-th column of the adjacency matrix adjl is greater than 0, then the attention score feature matrix attention 02 The value of the i-th row and j-th column of is equal to the value of the i-th row and j-th column of the attention matrix attn_dense2. If the value of the i-th row and j-th column of the adjacency matrix adjl is less than or equal to 0, then the attention score feature matrix attention 02 The value of the i-th row and j-th column of the attention score feature matrix is ​​set to 0, 02 Input into the softmax function of the attention network of the second figure, and output the normalized attention score feature matrix attention2 with n rows and n columns; e-11) Multiply the normalized attention score feature matrix attention2 with the linear transformation matrix h0 to obtain the encoding result z.

6. The method for maximizing influence based on deep learning and community division according to claim 1, characterized in that: Step f) comprises the following steps: f-1) Input the encoding result z into the decoder of the community division model and transpose the encoding result z to obtain the transposed matrix z T ; f-2) Compare the encoding result z with the transposed matrix z T Perform multiplication operation to obtain matrix A_pred0; f-3) Through the formula The prediction matrix A_pred1 is calculated and input into the Sigmoid function to obtain the prediction matrix A_pred. f-4) Using the encoding result z, the cluster center kmeans is calculated by the k-means algorithm; f-5) Use the encoding result z and the cluster center kmeans to obtain the similarity matrix q of n rows and Y columns through t distribution, q iy For the i-th user v i The probability of being assigned to community y, i∈{1,…,n}, y∈{1,…,Y}, Y is the number of communities, and the i-th row [q i1 ,q i2 ,...,q iy ,...,q iY ] is the one with the largest probability value as the i-th user v i Community tags c i , all n users’ community labels constitute a node label sequence arr, arr = {c1, c2, …, c i ,…,c n }.

7. The method for maximizing influence based on deep learning and community division according to claim 6, characterized in that: The following steps are also included: f-6) Through the formula Calculate the element value of the i-th row and y-th column in the target distribution matrix p; f-7) Use the target distribution matrix p and the similarity matrix q to obtain the divergence loss L through the divergence loss function c ; f-8) Calculate the loss L between the adjacency matrix adjl and the prediction matrix A_pred through the cross entropy loss function R ; f-9) Calculate the modularity Q of the adjacency matrix adjl and the prediction matrix A_pred by the modularity calculation formula; f-10) By the formula L=L R -βQ+γL c The total loss function L is calculated, where β and γ are both hyperparameters; f-11) Use the total loss function L to train the community partition model through the Adam optimizer to obtain the optimized community partition model.

8. The method for maximizing influence based on deep learning and community division according to claim 6, characterized in that: Step g) comprises the following steps: g-1) According to the i-th user v i Community tags c i The i-th user v i Join the community corresponding to the community label to obtain the community set C, C = {C1, C2, ..., C i ,…,C Y }, C i is the i-th community, i∈{1,...,Y}, each community contains a number of users; g-2) Determine the i-th community C i The number of users n i Is it greater than or equal to If yes, then the i-th community C i For important communities C si , all f important communities constitute the important community set C s , C s ={C s1 ,C s2 ,...,C si ,...,C sf }, where k is the number of pre-selected seeds.

9. The method for maximizing influence based on deep learning and community division according to claim 8, characterized in that: In step h), the formula Calculate the number of important communities C si The number of users k selected i , where n j is the jth important community C sj The number of users in , j∈{1,...,f}.

10. The method for maximizing influence based on deep learning and community division according to claim 8, characterized in that: Step i) comprises the following steps: i-1) Set the error tolerance ε, 0<ε<1, and set the parameter l, l≥1; i-2) Through the formula Calculate the quantity θ i , from the jth important community C sj Select θ i Each user generates a reverse reachable set RR; i-3) is the jth important community C sj Initialize a seed set S i , S i =φ, φ is an empty set; i-4) All n users {v1,v2,...,v i ,...,v n Substitute into formula F R (S i ∪v i )-F R (S i ) and add the user corresponding to the maximum difference to the seed set Si, where F R (·) is the number of RRs covering the reverse reachable set; i-5) Judge F R (S i ) is greater than or equal to If yes, then execute step i-6), if not, repeat step i-4) until F R (S i ) is greater than or equal to i-6) Through the formula Calculate the quantity θ and use it to replace the quantity θ i , update the reverse reachable set RR; i-7) Repeat step i-4) until the seed set S i Contains k i Users; i-8) The seed sets of all f important communities constitute the final seed set S,

Citation Information

Cited By

  • Large-scale symbol graph personalized influence community search method based on lightweight index and multi-level pruning

    CN121073690A