A project recommendation method for graph neural networks based on directed and undirected structural information
By adopting a graph neural network based on directed and undirected structural information in the recommendation system, the structural information in the session sequence diagram is extracted and accurate project implicit vectors are generated, and the problem of ignoring repeated clicks on project and structural information in the prior art is solved, and more accurate project recommendation is achieved.
Patent Information
- Application Number
- CN202111447363.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-01
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-12-01
AI Technical Summary
The prior art ignores items that appear repeatedly in click sequences and does not make good use of structural information in the session sequence graph when generating vector representations of items.
A graph neural network based on directed and undirected structural information is adopted to extract undirected and directed structural information in the session sequence diagram through graph convolutional networks and gated graph neural networks, generate accurate project implicit vectors, and recommend them through the target attention network.
Effectively utilize the structural information in the session sequence diagram, especially the importance of repeatedly clicking on items, improving the accuracy and effectiveness of project recommendations.
Smart Images

Figure CN114117229B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of recommendation methods, and particularly to an item recommendation method based on a graph neural network with directed and undirected structural information. Background Art
[0002] The recommendation system is one of the most important downstream applications in the fields of data mining and machine learning. It can help platform users alleviate the problem of information overload and select valuable information in many web applications such as e-commerce platforms, music websites, and so on. In most recommendation systems, the user's behavior sequence is arranged in time and exhibits the characteristics of anonymity and large data volume. In order to predict the user's behavior information at the next moment, session sequence-based recommendation learns the user's preferences by mining the sequential order feature information in the user's historical behavior.
[0003] A session sequence refers to a sequence of items generated by user clicks within a period of time; and session sequence-based recommendation can capture the importance of the dependencies within the sequence for sequence prediction. In other words, users usually have a common purpose in a certain session sequence, such as buying summer clothing; while the behavioral characteristics between different sequences may have little relevance, for example, the user's purpose in other sessions is to buy mobile phone accessories, etc. Session sequence-based recommendation is to predict the user's next click, that is, the sequence label v in session s. n+1 . Using the session sequence-based recommendation model, for each session s, the probabilities of all possible items can be obtained. Wherein Probability vector includes all possible situations of the next clicked item that appears after the current session, and the value of each element represents the recommendation score of the corresponding item. The top K items with the highest rankings in are the candidate items to be recommended.
[0004] Due to its high practical value, session sequence-based recommendation has received great attention in recent years, and many research results with good effects have emerged. Early methods were mainly based on Markov chains and recurrent neural networks. With the rise of graph neural networks in recent times and their excellent performance in many downstream tasks, some research work has applied GNNs to session sequence-based recommendation. Although these GNN-based methods have good performance, these methods also have some problems.
[0005] (1) Ignore the items that repeatedly appear in the click sequence. In fact, the importance of the items that appear multiple times is different from that of other items, and these items can reflect user preference information to a certain extent.
[0006] (2) When generating the vector representation of items, the structural information in the session sequence diagram is not well utilized. In fact, only considering the directionality between items is not sufficient. Introducing the undirected relationship between items can better learn the user's behavior information.
[0007] For example, the paper "Improving music recommendation in session-based collaborative filtering by using temporal context" by FONSECA et al. proposed using a clustering method to convert sparse session vectors into dense vectors; the paper "Session-based collaborative filtering for predicting the next song" by PARK.S.E et al. proposed a method to convert session sequences into vectors and then calculate the cosine similarity between session vectors. It can be seen that the neighborhood-based method is simple but effective; at the same time, it is also affected by data sparsity. More importantly, the above methods do not consider the complex transformation relationship between items within the session vector. Summary of the Invention
[0008] The technical problem to be solved by the present invention is to provide an item recommendation method based on a graph neural network with directed and undirected structural information to solve the problems existing in the prior art, namely, ignoring the items that repeatedly appear in the click sequence and not making good use of the structural information in the session sequence diagram when generating the vector representation of items.
[0009] To solve the above technical problems, the present invention provides an item recommendation method based on a graph neural network with directed and undirected structural information. The overall framework is as Figure 1 shown and includes the following steps:
[0010] S1. Let V represent the set of all items that appear in all session sequences, that is Then an anonymous session sequence of length n can be represented by and the items in session s are arranged in chronological order. Each v i ∈V represents the item clicked by the user in session s. The recommendation is to predict the user's next click, that is, to predict the sequence label v n+1 in session s;
[0011] S2. Receive the historical session sequence and convert the historical session sequence into a directed session sequence graph G s =(V s , ε s , A s ), where V sRepresentative point set, ε s Representative edge set, A s Representative set of adjacency matrices. Define A s as the concatenation of three adjacency matrices and , where represents the weighted adjacency matrix of an undirected graph, while and represent the weighted in-degree adjacency matrix and out-degree adjacency matrix respectively;
[0012] S3. Map the node v i ∈V to a d-dimensional vector representation in a random embedding vector space Extract the first intermediate hidden vector of items in the session sequence graph using a graph convolutional network, extract the second intermediate hidden vector of item transitions in the session sequence graph using a gated graph neural network, and obtain the item hidden vector through a first linear transformation;
[0013] S4. Input the item hidden vector into the target attention network to obtain the session hidden vector corresponding to the session sequence s;
[0014] S5. Obtain the global information and local information of the session sequence s, and construct the session vector representation through a second linear transformation;
[0015] S6. Predict the probabilities of all possible items being clicked for the session sequence s through the softmax function, and thus recommend the item with a high probability.
[0016] Furthermore, extract the first intermediate hidden vector of items in the session sequence graph using a graph convolutional network, that is, extract the attention-based undirected structural information through a graph convolutional network. The steps are as follows:
[0017] S31. Generate the feature matrix X of the session sequence graph; the d-dimensional feature vectors i corresponding to each node v in the session sequence graph are stacked to form the feature matrix X = [x 1 , …, x n T ;
[0018] S32. For the k-th layer of the graph convolutional layer, use the matrix H (k-1) to represent the input vectors of all nodes, and use H (k) to represent the output vectors of the nodes. Among them, the initial d-dimensional node vector is the feature initially input to the first layer of the graph convolutional layer network. The formula is:
[0019] H (0) = X, (1)
[0020] Before the input of each graph convolutional layer, for each node v i the features are averaged with the feature vectors of its local neighbors, and the calculation formula is:
[0021]
[0022] where, a ij is the edge weight between node v i and v j , d i = ∑ j a ij ;
[0023] S33. The output of the graph convolutional network is the first intermediate hidden vector.
[0024] Furthermore, for formula (2), it is simplified using simple matrix operations on the entire graph. Let S represent the result of symmetric normalization, as follows:
[0025]
[0026]
[0027]
[0028] where, for formula (3), and is 's degree matrix, and ⊙ is the dot product operator. is the weighted adjacency matrix of the undirected graph. I is the identity matrix.
[0029] Furthermore, in order to increase the weight of items with edges and reduce the interference of noise from other items, for the items connected by edges in the propagation matrix the weight is increased through the left part of formula (4), that is, the attention to self-information in the matrix is increased, as Figure 2 shown.
[0030] Furthermore, for formula (4), α and β are hyperparameters respectively to control the proportion of propagation matrix information and identity matrix information, so as to control the absorption proportion of node information with attention during the propagation process. As Figure 2 shown, in the adjacency matrix and the propagation matrix the item v 2 with repeated clicks and the item v 3 transformed during the repeated clicks will have higher attention information, that is, weight.
[0031] The above steps perform local smoothing on the implicit vector representation of nodes along the upper edge of the graph. After using the graph convolutional network as a feature preprocessing method to propagate the features, the nodes can absorb the attention information of adjacent nodes, and finally, the locally connected nodes can have similar prediction performances;
[0032] Furthermore, a gated graph neural network is used to extract the second intermediate implicit vector of item transitions in the session sequence graph, that is, to extract the attention-directed structural information through the gated graph neural network. The steps are as follows:
[0033] For the nodes in the session sequence graph, the node vector update steps are as follows:
[0034]
[0035]
[0036]
[0037]
[0038]
[0039] Among them, for formula (6), and control the magnitudes of the weights and bias terms, is used to represent the result of the interaction between a node and its adjacent nodes, and are the reset gate and the update gate respectively. The weight matrices W z , U z , W r , U r and W o , U o represent the learnable network parameters in the reset gate, the update gate, and the output gate respectively. represents the second intermediate implicit vector of node v i , is the sequence of node vectors in the session, and is the first intermediate implicit vector output by the graph convolutional network. σ(·) is the sigmoid function, and ⊙ is the dot product operator. The adjacency matrix represents the communication of nodes in the graph, represents the two-column matrix block of node v i in .
[0040] Among them, the matrix is defined as the in-degree matrix and the out-degree matrix The splicing of them respectively represents the weighted connection of the input edge and the output edge in the session sequence diagram. For example, given the session sequence s = [v 1 , v 2 , v 3 , v 2 , v 4 , the corresponding session sequence diagram G s and the adjacency matrix are as Figure 3 shown;
[0041] Furthermore, generate the item implicit vector, and the steps are as follows:
[0042] After the information processing of the graph convolutional network GCN and the gated graph neural network GGNN, the second intermediate implicit vector is obtained. In order to balance the ratio of the undirected structure information with attention to the directed structure information, the following formula is used for control:
[0043]
[0044] where γ is a hyperparameter, and thus the final accurate item implicit vector H is obtained;
[0045] After obtaining the implicit vector representation of each item, further construct the target vector, so as to analyze the relevance of historical behaviors considering the target item. The target item is all items to be predicted.
[0046] Use the local target attention model to calculate the attention score β i of all items v t in the session s for each target item v i,t , where and are respectively the second intermediate implicit vector representations of the item v i and v t .
[0047]
[0048] In formula (12), all items in the session sequence are respectively matched with the target item, and a weighted matrix is used for pairwise non-linear transformation; then the softmax function is used to normalize the obtained self-attention scores, so as to obtain the final attention score β i,t ;
[0049] Finally, for each session sequence s, the user's interest in the target item v t can be expressed as That is, the vector based on target attention
[0050]
[0051] It represents the degree of interest generated by the user among different target items;
[0052] Furthermore, generate a session sequence vector, and the steps are as follows:
[0053] Express the short-term interest of the user as a local vector Use the last item in the session sequence to represent this local vector, as shown in formula (14):
[0054]
[0055] Define the long-term preference of the user as a global vector which aggregates all item vectors that appear in session s. At the same time, use the attention mechanism to introduce the item of the last interaction and the dependencies between the items [v 1 , v 2 , …, v n that appear in the entire session.
[0056]
[0057]
[0058] where q, and W 1 , are the corresponding weight parameters, and α i represents the dependency between the last item and the items that appear in the entire session sequence.
[0059] Finally, for the local vector, global vector, and target-attention-based vector obtained above, concatenate the three and use linear transformation to obtain the session vector corresponding to session sequence s.
[0060]
[0061] where the weight parameter projects the result of concatenating the three vectors into the vector space ;
[0062] Furthermore, the steps for step S6 to generate recommendations are as follows:
[0063] Multiply the second hidden vector of all items v i ∈V with its corresponding session vector s h ,
[0064]
[0065] Then, the output vector of the obtained model is processed through the softmax function.
[0066]
[0067] where represents the predicted recommendation scores for all target items, and represents the probability that the target item will be clicked at the next moment in the session sequence s. The top K items with the highest rankings in
[0068] are the items to be recommended. For each session sequence graph, the loss function is defined as the cross-entropy between the predicted value and the actual value.
[0069]
[0070] where y i represents the one-hot encoded vector of the actual clicked item at the next moment in the session sequence.
[0071] Finally, the time-based backpropagation algorithm (BPTT algorithm) is used to iteratively train steps S2 - S6 to generate the parameters involved, such as W, α, and β, etc. The training method is a prior art and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0073] Figure 1 It is a schematic diagram of the overall framework of the model provided by the present invention;
[0074] Figure 2 It is a schematic diagram of the session sequence graph of the example provided by the present invention and its corresponding weighted undirected adjacency matrix and weighted propagation matrix schematic diagram;
[0075] Figure 3 It is a schematic diagram of the session sequence graph of the example provided by the present invention and its corresponding adjacency matrix schematic diagram;
[0076] Figure 4The performance of different components of the present invention in terms of the P@20 metric among the structural information components in the ablation experiment;
[0077] Figure 5 The performance of different components of the present invention in terms of the MRR@20 metric among the structural information components in the ablation experiment;
[0078] Figure 6 The performance of different components of the present invention in terms of the P@20 metric among the attention information components in the ablation experiment;
[0079] Figure 7 The performance of different components of the present invention in terms of the MRR@20 metric among the attention information components in the ablation experiment;
[0080] Figure 8 The flowchart of the item recommendation method according to an embodiment of the present invention. Detailed implementation manners
[0081] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0082] In view of the neglect of the items that repeatedly appear in the click sequence by the existing methods and the failure to make good use of the structural information in the session sequence graph when generating the vector representation of the items, the present invention proposes an item recommendation method based on a graph neural network with undirected structural information and directed structural information. This method gives a prediction of the item that the user will click at the next moment according to the user's current session sequence data, and in this process, it completely does not rely on the user's long-term preference information.
[0083] For session sequence-based recommendation, first construct a directed session sequence graph from the information in the historical session sequence; use the graph convolutional network GCN and the gated graph neural network GGNN to extract the undirected structural information and directed structural information of the item transitions in the session sequence graph respectively, then generate accurate item latent vectors, and then input the obtained item latent vectors into the attention network, considering both the global information and local information of the session, so as to construct a more reliable session representation and infer the next click item based on this.
[0084] As Figure 1 and Figure 8 shown, the item recommendation method based on a graph neural network with undirected structural information and directed structural information includes:
[0085] S1. Let \(V\) represent the set of items that have appeared in all session sequences, that is Then an anonymous session sequence of length \(n\) can be represented by and the items in session \(s\) are arranged in chronological order. Each \(v\) i \(\in V\) represents an item that the user clicks on in session \(s\). The recommendation is to predict the user's next click, that is, to predict the sequence label \(v\) in session \(s\) n+1 .
[0086] S2. Convert the historical session sequence into a directed session sequence graph \(G\) s \(=(V\) s , \(\varepsilon\) s , \(A\) s ), where \(V\) s represents the set of nodes, \(\varepsilon\) s represents the set of edges, and \(A\) s represents the set of adjacency matrices. In the session sequence graph \(G\) s each node represents an item \(v\) i \(\in V\), and each edge \((v\) i-1 , \(v\) i ) \(\in \varepsilon\) s represents that the user has clicked on item \(v\) i-1 and item \(v\) i in sequence, and \(A\) s is defined as the concatenation of three adjacency matrices and . represents the weighted adjacency matrix of an undirected graph, while and represent the weighted in-degree adjacency matrix and out-degree adjacency matrix respectively.
[0087] S3. Process the session sequence graph \(G\) s to obtain the item implicit vectors of all nodes in the session sequence graph;
[0088] S4. Input the item implicit vectors into the target attention network to obtain the attention-based vector of the session sequence;
[0089] S5. Obtain the global information and local information of the session sequence \(s\), and construct the session vector representation through a second linear transformation;
[0090] S6. Predict the probabilities of all target items being clicked for the session sequence \(s\) through the softmax function, and thus recommend the items with high probabilities.
[0091] In step S3, for each node \(v\) in the session sequence graph \(G\) s iThe ∈s mapping is projected into a random embedding vector space to obtain a d-dimensional vector representation Then, through a graph convolutional network, a gated graph neural network, and linear processing, the project implicit vector H is obtained.
[0092] Specifically, it includes:
[0093] (1) According to the conversation sequence graph G s Generate the initial project implicit vector x corresponding to each project v i ∈V, where V represents the set of all projects that have appeared in the conversation sequences; the d-dimensional feature vectors corresponding to each node v in the conversation sequence graph i are stacked to form the feature matrix of the conversation sequence graph X = [x i , …, x 1 n . T i The specific steps for generating x
[0094] include:
[0095] 1) The weighted undirected adjacency matrix s in the conversation sequence graph G ij is a sparse and symmetric adjacency matrix, where a i represents the edge weight between nodes v j and v ij ; if there is no connection between nodes, it is represented as a 1 = 0. n
[0096] 2) Define the degree matrix D as a diagonal matrix D = diag(d i , …, d j ), and the values on the diagonal are equal to the sum of the row elements of the adjacency matrix d ij = ∑ i a 1 .
[0097] 3) Through each node v in the graph has a corresponding d-dimensional feature vector n Therefore, the feature matrix of the conversation sequence is the stacking of the feature vectors corresponding to each node in the graph, that is, X = [x T , …, x i i .
[0098] (2) Input x (k-1) into a graph convolutional network with undirected structural information to obtain the first intermediate implicit vectors of all nodes in the graph (the first intermediate implicit vectors carry undirected structural information).
[0099] (3) Obtain the second intermediate hidden vector (with directed structure information) of all nodes in the graph through a gated graph neural network with directed structure information.
[0100] (4) Input the second intermediate hidden vector into the first linear transformation to obtain the precise item hidden vector.
[0101] Graph neural networks have a natural adaptability to session sequence-based recommendations because they can automatically extract the features of session sequence graphs while taking into account rich node connection relationships.
[0102] Similar to convolutional neural networks (CNNs) and multi-layer perceptrons (MLPs), graph convolutional networks (GCNs) learn the features of each node v i in a multi-layer structure and obtain a new feature representation, and then input it into the corresponding linear classifier.
[0103] Among them, in step (2), the steps of generating the first intermediate hidden vector include:
[0104] S31. For the k-th layer of graph convolutional layer, use the matrix H (k-1) to represent the input vectors h i of all nodes, and use H (k) to represent the output vectors of the nodes. The initial d-dimensional node vectors are the initially input features and are input into the first layer of GCN:
[0105] H (0) = X, (1)
[0106] The GCN with K layers is equivalent to applying a K-layer MLPs model to the feature vectors x i of all nodes in the graph, except that the hidden vector representations of each node are averaged with its neighbor nodes at the beginning of each layer. In each graph convolutional layer, the vector representation of the node has three update stages: feature propagation, linear transformation, and pointwise non-linear activation. Only the feature propagation stage is used in the present invention for learning the item hidden vector.
[0107] S32. At the beginning of each layer, the features of each node s h are averaged with the feature vectors of its local (random) neighbors:
[0108]
[0109] The input of the graph convolutional network is the first intermediate hidden vector.
[0110] Preferably, use the simple matrix operation of the whole graph to simplify formula (2), specifically as follows:
[0111] Let \(S\) represent the adjacency matrix with self-loops after "symmetric normalization":
[0112]
[0113]
[0114]
[0115] where and is the degree matrix, \(\odot\) is the dot product operator, and \(I\) is the identity matrix. Since the adjacency matrix considers the weights of duplicate items, the matrix has additional attention to the items with duplicate clicks. At the same time, because \(S\) has self-loops, the process of symmetric normalization results in less weight for items connected by multiple edges compared to items connected by a single edge or no edge. \(\alpha\) and \(\beta\) are hyperparameters.
[0116] To increase the weights of items with edges and reduce the interference of noise from other items, in the propagation matrix the weights of items connected by edges are increased through the left part of formula (4), that is, the attention to self-information in the matrix is increased.
[0117] The hyperparameters \(\alpha\) and \(\beta\) are used to control the ratio of the information in the propagation matrix to the information in the identity matrix, so as to control the absorption ratio of the node information with attention during the propagation process. According to Figure 2 the specific examples given in it can be seen that for the adjacency matrix and the propagation matrix 2 items with repeated clicks \(\alpha\) 3 and items that are transformed during repeated clicks will have higher attention information, that is, weights.
[0118] Therefore, the equivalent update form of formula (2) can be transformed into a simple sparse matrix multiplication for all nodes.
[0119] The above steps perform local smoothing on the implicit vector representation of nodes along the edges of the graph. After propagating the features using the graph convolutional network as a feature preprocessing method, the nodes can absorb the attention information of adjacent nodes, and finally, the locally connected nodes can have similar prediction performances.
[0120] In step (3), the method of Li et al. is used to construct a gated graph neural network GGNN.
[0121] For the node \(v\) s in the session sequence graph \(G\) i , the update steps of the node vector are as follows:
[0122]
[0123]
[0124]
[0125]
[0126]
[0127] In formula (6), the adjacency matrix represents the communication of nodes in the graph, represents node v i in the two-column matrix block, is the sequence of node vectors in the session, and that is, the final output of the graph convolutional network is used as the initial input of the gated graph neural network, and controls the magnitudes of the weights and bias terms, which is used to represent the result of the interaction between a node and its adjacent nodes.
[0128] For formula (7), that is, the reset gate is obtained through the σ(·) sigmoid function respectively and the update gate weight matrices W z , U z , W r , U r and W o , U o respectively represent the learnable network parameters in the reset gate, update gate and output gate, and ⊙ is the dot product operator. Finally represents the hidden vector of node v i generated by the GGNN gated graph neural network.
[0129] The matrix is defined as the concatenation of the in-degree matrix and the out-degree matrix , which respectively represent the weighted connections of the input edges and output edges in the session sequence graph. For example, given the session sequence s = [v 1 , α 2 , v 3 , v 2 , v 4 , the corresponding session sequence graph G s and the adjacency matrix are as Figure 3 shown. It can be seen that the weighting in the directed adjacency matrix is set according to the degree of closeness between nodes. For example, v 2From α 3 to α 4 each has an edge, but the reason why their weights are different is that there are more edges connecting α 2 and α 3 which means a higher similarity between them. To achieve a better prediction effect, from the perspective of α 2 more information of α 3 should be absorbed, so the model should pay more attention to v 3 rather than v 4 .
[0130] Therefore, for each session sequence graph G s , the GGNN model propagates the node information with attention between adjacent nodes, while the reset gate and the update gate respectively determine the next information to be discarded or retained.
[0131] In step (4), after the information processing of the graph convolutional network GCN and the gated graph neural network GGNN, we respectively obtain The former performs undirected structure information processing with attention on the initial embedding vector, and the latter more precisely extracts the directed structure information with attention in the graph structure based on the former.
[0132] To balance the ratio of the undirected structure information with attention and the directed structure information with attention, the following formula is used for control:
[0133]
[0134] where γ is a hyperparameter, and thus the final accurate item latent vector H is obtained.
[0135] In step S4, the specific steps include:
[0136] S41. Use the local target attention model (this model is a prior art and will not be introduced in detail) to calculate the attention score β i for each target item v t ∈V in session s, where i,t and and are respectively the latent vector representations of item v i and v t .
[0137]
[0138] In the above formula, the items in the session are respectively matched with the candidate targets, and a weighted matrix Perform paired non - linear transformations. Then, use the softmax function to normalize the obtained self - attention scores and obtain the final attention scores.
[0139] S42. For each session sequence s, the user's interest in the target item v t can be expressed as
[0140]
[0141] Finally, obtain the target - attention - based vector which represents the degree of interest generated by the user among different target items.
[0142] In step S5, use the item vectors involved in session s to further explore the user's short - term and long - term preferences, thereby obtaining the local vector, global vector in the session, and combining the target - attention - based vector calculated in the previous section to generate the final session vector.
[0143] S51. Obtain the local vector. In a session sequence s, the user's final behavior is often determined by the last interacted item in the current sequence. Therefore, represent the user's short - term interest as the local vector and this local vector is the vector representation of the last item in the session sequence.
[0144]
[0145] S52. Obtain the global vector. Define the user's long - term preference as the global vector which aggregates all the item vectors that appear in session s. At the same time, use the attention mechanism to introduce the last interacted item and the dependencies between the item [v 1 , v 2 , …, v n that appears in the whole session.
[0146]
[0147]
[0148] where q, and W 1 , are the corresponding weight parameters.
[0149] S53. Concatenate the obtained local variable, global variable, and target - attention - based vector, and use linear transformation to obtain the session vector corresponding to the session sequence s.
[0150]
[0151] where the weight parameter Project the result of concatenating the three vectors into the vector space Note that different session vectors will be correspondingly generated for different target items (items in the session sequence).
[0152] In step S6, after obtaining the session vector s corresponding to the session sequence s h then, for all items v i ∈V, calculate the score that is, multiply the candidate item vector by the session vector s h and then
[0153]
[0154] pass the obtained output vector of the model through the softmax function
[0155]
[0156] where represents the predicted recommendation scores for all target items, and represents the probability that the target item will be clicked at the next moment in the session sequence s The top K items with the highest rankings in
[0157] For each session sequence graph G s define the loss function as the cross-entropy between the predicted value and the actual value
[0158]
[0159] where y i represents the one-hot embedding vector of the true clicked item at the next moment of the session sequence
[0160] In use, the dataset can be used to iteratively train steps S2 - S6, such as using the Backpropagation Through Time (BPTT) algorithm for training, so as to obtain the parameters in the above steps, such as W, W1, W2, W3, etc. These parameters can be initially set randomly and then learned during training
[0161] During training, each sequence is used as a training sample, so the total error is the sum of the errors at each time step (recommended). Note that in the recommendation scenario based on session sequences, most sessions are relatively short sequences. To prevent overfitting, a relatively small number of training epochs is adopted.
[0162] Experimental analysis:
[0163] 1. Experimental datasets
[0164] In practical applications, the publicly available datasets Yoochoose and the publicly available dataset Diginetica released by CIKM Cup 2016 are used to evaluate the method. The Yoochoose dataset contains the user click stream on an e - shopping platform within 6 months, while the Diginetica dataset only contains the data of successful transactions, that is, the user purchase stream.
[0165] Meanwhile, the input sequence data is further segmented to generate corresponding sequences and labels. For the input session sequence s = [v 1 , v 2 , …, v n , as a data augmentation strategy, a series of sequences and labels ([v 1 , v 2 ), ([v 1 , v 2 , v 3 ), …, ([v 1 , v 2 , …, v n-1 , v n ) are generated. Among them, [v 1 , v 2 , …, v n-1 is the generated sequence, and v n represents the item clicked at the next moment, that is, the label of the sequence.
[0166] The specific situation of the finally used dataset is shown in Table 1.
[0167] Table 1 Statistics of experimental datasets
[0168]
[0169] 2. Evaluation criteria
[0170] After determining the dataset, two commonly used metrics in session - sequence - based recommendation are adopted as the evaluation indicators of the algorithm.
[0171] (1) P@20 (Precision) is a widely used metric for prediction accuracy. It represents the proportion of correctly recommended items among the top 20 recommended results of the algorithm.
[0172] (2) MRR@20 (Mean Reciprocal Rank) is the mean of the reciprocal ranks of the correctly recommended items in the algorithm's recommendation results. When the true result exceeds 20 in the algorithm's recommendation ranking, the corresponding reciprocal rank is 0. The MMR metric is a method that takes into account the recommendation order. A larger MRR value indicates that the true result is at the top of the ranking list in the recommendation list, which proves the effectiveness of the recommendation system.
[0173] 3. Experimental Setup
[0174] In both datasets, the dimension of the latent vector is set to d = 100. All hyperparameter settings are initialized using a Gaussian distribution function with a mean of 0 and a standard deviation of 0.1. At the same time, the mini-batch Adam optimizer is also used to optimize these parameters, and the initial learning rate η is set to 0.001 and decays by 0.1 every three training epochs. In addition, the batch size is set to 100, and the L2 regularization parameter is set to 10 -5 。
[0175] 4. Analysis of Experimental Results
[0176] The performance of several methods on the two metrics of P@20 and MRR@20 is shown in Table 2, where the best results are shown in bold. The method proposed in the present invention can flexibly construct the connection between items on the session sequence graph, and extract the directed and undirected structural information with attention, so that the subsequent learning of target attention can be more accurate, and the final recommendation is given by integrating the global interest and local interest of the user in the session. According to the experimental data in Table 2, it is obvious that this method has obtained the best performance results on the two metrics for the three datasets, which proves the effectiveness of this method.
[0177] Traditional recommendation methods such as POP and S-POP perform unsatisfactorily on session sequence-based problems because they ignore the user's preference in the current session and only consider the top K most popular items. BPR-MF shows that using the semantic information in the session is meaningful, while the better-performing FPMC shows that using the first-order Markov chain to model the session sequence is a relatively effective method. Similarly, as a traditional recommendation method, Item-KNN is better than the former two. It is worth noting that Item-KNN only relies on calculating the similarity between items, which shows that the co-occurrence of items is also an important piece of information. However, Item-KNN does not take into account the temporal information in the session, so it cannot capture the information about the transition between items.
[0178] Different from traditional methods, methods based on deep learning generally have better performance on all dataset metrics. GRU4Rec is a method based on recurrent neural networks, and its performance can be better than most traditional methods and comparable to some traditional methods. This shows that recurrent neural networks have a certain ability to model sequential data. However, GRU4Rec mainly focuses on modeling session sequences and cannot capture user preferences in sessions. Subsequently, methods such as NARM and STAMP have significantly improved GRU4Rec. NARM explicitly captures the main preferences of users in sessions, while STAMP uses the attention mechanism to consider users' short-term interests, which are the reasons why they are superior to GRU4Rec. RepeatNet is also an algorithm based on recurrent neural networks. It achieves good prediction results by considering users' repeated click behaviors, which also shows the importance of modeling users' behavioral habits. However, the improvement of RepeatNet compared to NARM and STAMP is limited, possibly because modeling users' repeated click habits only through item features is insufficient, and the RNN-based structure cannot capture some common dependencies within sessions.
[0179] Methods based on graph neural networks construct each session sequence into a subgraph and encode all items in the session through graph neural networks. SR-GNN and TAGNN both have better results than all RNN-based models. SR-GNN uses gated graph neural networks to learn the dependencies between items in session sequences, while TAGNN further uses the attention mechanism to explore the dependencies between items in the session and the target item. However, these methods all learn completely according to the directed relationship of the session sequence graph and do not comprehensively consider the undirected relationship in the session sequence graph. Because the relationship between items in the session sequence is sometimes not a one-way relationship but a two-way relationship, using undirected structural information can often capture a more comprehensive relationship between items. Moreover, they all ignore the repeated click features that appear in the session sequence. Intuitively, the importance of repeated items in a sequence should be greater. In addition, in the actual recommendation scenario, the degree of association between items is variable, and these methods use an averaging method for the dependencies between items in the session and cannot reflect the different degrees of dependence of one item on other items through weights or attention methods.
[0180] The method proposed in this paper performs better than the methods introduced in the previous text. Specifically, on three datasets, there are relative improvements of 3.55%, 1.38%, and 1.18% for P@20 compared to the best-performing related methods, and relative improvements of 1.92%, 4.34%, and 1.98% for MRR@20 compared to the best-performing related methods. This method can well extract the structural information in the conversation sequence graph. First, it uses a graph convolutional network and a gated graph neural network to extract the undirected and directed structural information in the graph respectively, and then linearly combines the two to achieve an accurate representation of the vector. And it takes into account the repeated click items in the conversation sequence, increasing the weight of the repeated information through an attention network; at the same time, it increases the self-information ratio of the nodes in the conversation sequence graph by adding self-loops and matrix operations, making the nodes less vulnerable to the noise interference of other nodes. Then, according to the different dependencies between items, it uses the attention network to assign different weights, so that the network can generate an accurate vector representation.
[0181] Table 2 Comparison of experimental results
[0182]
[0183] Ablation experiment:
[0184] This method can flexibly capture the structural information in the conversation sequence graph and the relationships between items. To verify the actual roles of the components in the model, several model variants were set up for ablation experiments. In the experiment, SR-GNN was selected as the baseline method for comparison, and the data in the experiment was shown in the form of the relative improvement percentage compared to SR-GNN.
[0185] First, perform the combined analysis of directed and undirected structural information: (a)-GCN, which only extracts the undirected structural information in the conversation sequence graph. (b)-GNN, which only extracts the directed structural information in the conversation sequence graph. (c)-GCN+GNN, which inputs a random initial vector into two neural networks at the same time, and then linearly combines the model output results. (d)-GCN+GNN(GCN), which inputs a random initial vector into GCN, then uses the output vector of the GCN model as the input of the GNN model, and finally linearly combines the output results of the two models. The experimental comparison results are as Figure 4 、 Figure 5 shown.
[0186] Among them, AVG represents the average performance of the four combination conditions on the three datasets. From Figure 4 、 Figure 5It can be seen that the method GCN+GNN (GCN) that combines directed structure information and undirected structure information has achieved the best results in two metrics, P@20 and MRR@20, on three datasets, which proves the importance of considering both directed structure information and undirected structure information. Figure 5 The average data AVG also shows that considering only undirected structure information performs better than considering only directed structure information. However, in terms of the performance of a single dataset, in the MRR@20 metric of the Yoochoose 1 / 4 dataset and the Diginetica dataset, considering directed structure information performs slightly better than considering undirected structure information. This reflects to a certain extent that the preferences of users and the connections between items, as well as the conversion directions between items, have different degrees of importance in different scenarios in session sequence-based recommendation. However, on average, undirected structure information is more important. This shows that in the scenario of session sequence-based recommendation, although the conversion direction between items by users is worthy of consideration, the connections between the items browsed by users still need to be considered as the key point in order to better learn the preferences of users. And a better approach than the previous two is to consider both directed structure information and undirected structure information. Figure 4 、 Figure 5 In general, the method of combining GCN and GNN performs better than using GCN or GNN alone in the comprehensive performance and average performance AVG on three datasets. In the method of GCN+GNN, the input data of the two network models are random embedding vectors, while in the method of GCN+GNN (GCN), the input of the GNN model is the vector obtained by extracting undirected structure information by the GCN model. This shows that compared with directly using random vectors, extracting undirected structure information first and then extracting directed structure information can obtain more accurately represented embedding vectors.
[0187] Then, a combined analysis of repeated click attention information and item dependencies is carried out: (a) GCN+GNN, without considering repeated click attention information and different dependencies between items. (b) AttGCN+GNN, only considering the attention information of repeated click items in GCN. (c) GCN+AttGNN, only considering different degrees of dependencies between items in GNN. (d) AttGCN+AttGNN, fusing the attention information of repeated clicks in GCN and the attention-based item dependencies in GNN at the same time. The experimental results are as Figure 6 、 Figure 7 shown.
[0188] From Figure 6 、 Figure 7It can be seen from the data that AttGCN+AttGNN, which comprehensively considers the repeated click attention information and the dependencies between items, has achieved the best experimental results in two metrics on the three datasets, indicating that the repeated click behavior and the dependencies between items are of certain importance in session sequence-based recommendation. At the same time, according to Figure 6 the experimental performance of the P@20 metric in, the attention that separately considers the repeated click attention and the relationships between items generally performs better than GCN+GNN that does not consider any attention information. The AttGCN+AttGNN that comprehensively considers the two has the best effect, indicating that the attention information can indeed make important information be retained as much as possible and obtain more accurately represented embedding vectors. However, according to Figure 7 the experimental performance of the MRR@20 metric in, although the AttGCN+AttGNN that comprehensively considers the two types of attention can still obtain the best experimental performance, the performance of separately considering one of the two types of attention is slightly worse than that of GCN+GNN that does not consider any attention information. That is to say, although using attention information alone can improve the precision of the recommendation results, it does not perform well in predicting the recommendation ranking. The possible reason is that when the vectors are input into the GCN and GNN models, if one model considers attention while the other does not, the representation patterns of the adjacency matrices in the two models are not unified, resulting in the vectors not being able to utilize the attention to retain the structural information simultaneously when inputting and outputting between the two models. Instead, the structural information is disturbed by the inconsistent attention patterns, so the embedding vectors with accurate representations are not generated in the end, and the accurate prediction scores of each item cannot be calculated in the prediction stage.
[0189] In the session sequence-based recommendation scenario, the user's repeated click behavior and the graph structure information are both worthy of key consideration because they can well predict the user's behavior without knowing the user's historical preferences. This paper proposes a graph neural network sequence recommendation method based on the fusion of directed and undirected information with co-attention. The present invention not only extracts the directed and undirected structural information in the session sequence graph using the GCN and GNN models and performs a linear combination, but also introduces an attention mechanism when generating the item implicit vectors to effectively extract the user's repeated clicks and the complex transition information between items, making the generated session vectors more accurate in prediction during the recommendation process. On the actual datasets in three real-world scenarios, the present invention verifies that the proposed algorithm is superior to other state-of-the-art methods, and validates the effectiveness of the attention mechanism and the complex structural information through detailed ablation experiments.
[0190] Finally, it should be noted that the above description is the preferred embodiment of the present invention. It should be pointed out that although the preferred embodiments of the present invention have been described, for those skilled in the art of this technology, once they know the basic creative concept of the present invention, without departing from the principle described in the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.
Claims
1. A project recommendation method based on a graph neural network with directed and undirected structural information, characterized in that, the method includes: S1. Let V represent the set of items that have appeared in all session sequences, i.e., Then an anonymous session sequence of length n is represented by and the items in session s are arranged in chronological order. Each v i ∈V represents an item that the user clicked on in session s; S2. Receive the historical session sequence and convert the historical session sequence into a directed session sequence graph G s =(V s , ε s , A s ), where V s represents the set of vertices, ε s represents the set of edges, and A s represents the set of adjacency matrices. Define A s as the concatenation of three adjacency matrices and . Among them, represents the weighted adjacency matrix of an undirected graph, while and represent the weighted in-degree adjacency matrix and out-degree adjacency matrix respectively; S3. Map the node v i ∈ V to the random embedding vector space to obtain a d-dimensional vector representation Use a graph convolutional network to extract the first intermediate hidden vector of the items in the session sequence graph, use a gated graph neural network to extract the second intermediate hidden vector of the item transitions in the session sequence graph, and weighted sum the first intermediate hidden vector and the second intermediate hidden vector to obtain the item hidden vector; wherein, the output of the graph convolutional network is the first intermediate hidden vector; the input of the gated graph neural network is the first intermediate hidden vector, and the output is the second intermediate hidden vector; S4. Input the project implicit vector into the target attention network to obtain the target attention-based vector of the session sequence s; S5. Obtain the global information and local information of the session sequence s, and splice them with the target attention-based vector, and then construct the session vector representation through a second linear transformation; S6. Predict the probabilities of all target projects being clicked for the session sequence s through the softmax function, and thus recommend the projects with high probabilities.
2. The project recommendation method according to claim 1, characterized in that, the steps of using a graph convolutional network to extract the first intermediate implicit vector of the projects in the session sequence graph are as follows: S31. Generate the feature matrix X of the session sequence diagram; for each node v in the session sequence diagram i The corresponding d-dimensional feature vector The stacking of forms the feature matrix of the session sequence diagram S32. For the graph convolutional layer of the k-th layer, use the matrix H (k-1) to represent the input vectors of all nodes, and use H (k) to represent the output vectors of the nodes. Among them, the initial d-dimensional node vectors are the features initially input to the first layer of the graph convolutional layer network. The formula is as follows: H (0) = X, (1) Before the input to each graph convolutional layer, the features of each node v i are averaged with the feature vectors of its local neighbors, and the calculation formula is: Among them, a ij is the edge weight between nodes v i and v j , d i = ∑ j a ij ; S33. The output of the graph convolutional network is the first intermediate implicit vector.
3. The project recommendation method according to claim 2, characterized in that, in the step of using a graph convolutional network to extract the first intermediate implicit vector of the projects in the session sequence graph, formula (2) is replaced with the following formula: where S represents the adjacency matrix with self-loops after "symmetric normalization", and is the degree matrix, ⊙ is the dot product operator, is the propagation matrix of the graph convolutional network, α and β are hyperparameters, I is the identity matrix, is the first intermediate hidden vector of the output.
4. The project recommendation method according to claim 3, characterized in that, For the projects connected by edges, by increasing the weight of the left half of formula (4), the attention of the self-information in the matrix is increased, and the hyperparameters α and β are used to control the proportion of the propagated matrix information and the identity matrix information, so as to control the absorption proportion of the node information with attention in the propagation process.
5. The project recommendation method according to claim 1, characterized in that, the steps of using a gated graph neural network to extract the second intermediate implicit vector of the project transformation in the session sequence graph are: Among them, and are used to control the magnitudes of the weights and bias terms, which is used to represent the result of the interaction between a node and its adjacent nodes, and are the reset gate and update gate respectively, and the weight matrices W z , U z , W r , U r and W o , U o represent the learnable network parameters in the reset gate, update gate and output gate respectively, represents the second intermediate hidden vector of node v i , is the sequence of node vectors in the session, and is the first intermediate hidden vector output by the graph convolutional network, σ(·) is the sigmoid function, and ⊙ is the dot product operator; the adjacency matrix represents the communication of nodes in the graph, represents the two-column matrix block of node v i in ; the matrix is defined as the concatenation of the in-degree matrix and the out-degree matrix .
6. The project recommendation method according to claim 1, characterized in that, in step S4, the steps of obtaining the target attention-based vector of the session sequence s are as follows: S41. Calculate all items v in the conversation sequence s using the local target attention model i For each target item v t ∈V, the attention distribution e i,t , and then obtain the attention score β through the softmax(·) function i,t . The formula is as follows: Among them, and are respectively the item implicit vector representations of all items v i and the target item v t , is the weight matrix, generated through training, and exp(·) is the function for calculating the nth power of the constant e; S42. For each session sequence s, the user's interest in the target item v t is expressed as The calculation formula is:
7. The project recommendation method according to claim 6, characterized in that, step S5 includes: S51. Obtain a local vector The local vector is the vector representation of the last item in the session sequence s , that is, its hidden vector, and the formula is: S52. Obtain the global vector The global vector aggregates all item vectors that appear in the session sequence s, and also uses the attention mechanism to introduce the dependencies between the items [v 1 , v 2 , …, v n that appear in the entire session sequence s. The formula is as follows: wherein is the last item in the conversation, and is the corresponding weight parameter, ɑ i represents the dependency between the last item and the items that appear in the entire conversation sequence; S53. Splice the local vector, the global vector and the target attention-based vector, and use a second linear transformation to obtain the session vector corresponding to the session sequence s. The formula is: Among them, the weight parameter Project the result of concatenating the three vectors into the vector space ; The W1, W2, and W3 are weight parameter matrices, which are generated through training.
8. The project recommendation method according to claim 7, characterized in that, step S6 includes: S61. After obtaining the session vector s corresponding to each session sequence s h then, for all items v i ∈ V, calculate the score that is, multiply the second hidden vector of all items v i ∈ V by the transpose of its corresponding session vector s h The formula is as follows: S62. Output the vector through the second softmax function The formula is as follows: where represents the predicted recommendation scores of all candidate target items, and represents the probability that the target item will be clicked at the next moment in the session sequence s, the top K items with the highest rankings in s are the items to be recommended; among them, for the session sequence graph G s , the loss function is defined as the cross-entropy between the predicted value and the actual value.
9. The project recommendation method according to any one of claims 1-8, characterized in that, Repeat steps S2-S6, and use the time-based backpropagation algorithm to train the parameters.
Citation Information
Patent Citations
A session sequence recommendation method and system based on a graph convolutional neural network
CN109816101A
Session-based project recommendation method, device and apparatus and storage medium
CN110119467A