A Contrastive Learning Session Recommendation Method Incorporating Hypergraphs
By using hypergraph structure and contrast learning methods in the session recommendation system, the problems of project advanced relation modeling and data sparseness are solved, and the accuracy and performance of the recommendation system are improved.
Patent Information
- Application Number
- CN202211692089.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-28
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-12-28
AI Technical Summary
Existing session recommendation methods fail to adequately model the high-order relationships between projects and the highly sparse session data, resulting in low recommendation system performance.
The hypergraph structure is used to model the high-order relationship between projects, and the embeddings of global hypergraphs and local conversation graphs are compared through comparative learning methods to enhance the complementarity of conversation information.
It improves the accuracy of the recommendation system, makes up for the shortcomings of the sparsity of session data, and enhances the interactive information modeling ability between projects.
Smart Images

Figure CN116186390B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a contrastive learning session recommendation method integrating hypergraphs, and specifically to a method for enhancing the interaction information between hypergraph information and items in a session-based recommendation task, belonging to the technical field of data mining and applications. Background Art
[0002] In the information age, with the development of Internet technology, the information received by users every day has grown explosively. The massive information often leads to a high degree of information redundancy. The recommendation system can combine user interests, filter and select information that meets user needs, and recommend it to users. Therefore, the recommendation system has been widely applied in scenarios such as e-commerce and video websites.
[0003] Traditional recommendation systems mainly utilize known user information. However, in many real-world scenarios, due to a series of reasons such as privacy protection, user information is unknown. Therefore, session-based recommendation systems, as a branch of information recommendation, have gradually emerged. The session-based recommendation system uses the behavior sequence of anonymous users over a period of time to predict the next item, and can learn user preferences based on the items interacted by users in the session, providing accurate personalized recommendations for users.
[0004] Existing session recommendation methods mainly include four categories: session recommendation methods based on Markov chains, session recommendation methods based on factorization, session recommendation methods based on recurrent neural networks, and session recommendation methods based on graph neural networks. Among them, the first two model-based methods recommend the next item based on the previous item and can capture first-order dependencies. The method based on recurrent neural networks usually models session data as an ordered sequence and models the order of items in the session. The method based on graph neural networks constructs session data as a directed graph and constructs pairwise relationships between items.
[0005] However, in real-world scenarios, the occurrence of an item is often the result of the joint action of a series of previous items, and the relationship between items is not a simple binary relationship but a more complex many-to-many relationship. Summary of the Invention
[0006] The object of the present invention is to creatively propose a contrastive learning session recommendation method integrating hypergraphs for the technical problems that the existing technology fails to fully model the high-order relationships between items and the session data is highly sparse, resulting in low performance of the session recommendation system.
[0007] Compared with the ordinary graph structure, the hypergraph structure can better model this kind of high-order relationship. The hyperedges of a hypergraph can connect multiple vertices, and the correlation between items encoded by the hyperedges is no longer binary, but ternary, quaternary or higher-order. Therefore, the present invention uses the hypergraph structure to model the conversation, which can capture the complex high-order relationship between items and more conform to the characteristics of conversation data.
[0008] Meanwhile, in a conversation, the interaction information between items is very important. The present invention constructs a local conversation graph for each conversation and obtains the characteristics of the conversation by using the interaction information between items in the conversation. To solve the problem that conversation data is highly sparse, the present invention compares the conversation embeddings obtained by two different graph construction methods through contrastive learning, so that the two conversation embeddings can provide new information for each other, thereby improving the accuracy of recommendation.
[0009] First, the relevant concepts involved in the present invention are explained:
[0010] 1. Item set and conversation set
[0011] Define V as the set of all items, that is, the item set, V = {v1, v2, v3,..., v n}, where n is the number of items. Each conversation is represented as a set s, L is the length of the conversation.
[0012] 2. Hypergraph
[0013] A hypergraph is an extension of the ordinary graph structure. In the ordinary graph structure, an edge can only connect two vertices. While the edges (i.e., hyperedges) of a hypergraph can connect any number of vertices.
[0014] The present invention is implemented by the following technical solutions.
[0015] A contrastive learning conversation recommendation method integrating hypergraphs, comprising the following steps:
[0016] Step 1: Construct a global hypergraph and a local conversation graph according to the conversation data.
[0017] Specifically, Step 1 includes the following steps:
[0018] Step 1.1: Construct a global hypergraph G g , G g = (V g , E g ); V g represents the set of vertices in the hypergraph; E g represents the set of hyperedges in the hypergraph, which is composed of all historical conversation sequences.
[0019] According to the definition of a hypergraph, a hyperedge can connect multiple vertices. Therefore, each session is constructed as a hyperedge, that is and each item
[0020] Step 1.2: Construct a local session graph G according to the session data s , G s =(V s , E s ); V s represents the vertex set composed of all items in the session; E s represents the set of edges in the session graph.
[0021] The local session graph is constructed according to the current session. Given a session where the set of all items in the session constitutes V s , a directed edge e in the local session graph is formed according to the interaction order between items i , that is represents the item of the i-th interaction in session s.
[0022] Step 2: Input the obtained global hypergraph data into the hypergraph attention network for learning to obtain the item embedding representation under the global hypergraph.
[0023] Specifically, Step 2 includes the following steps:
[0024] Step 2.1: At the beginning layer of the network, embed each item from a high-dimensional space into a low-dimensional continuous vector space to obtain the initial embedding representation of each item where represents the initial embedding of the n-th item.
[0025] Step 2.2: Input the initial embedding representation of each node into the first layer of the hypergraph attention network, give different degrees of attention to each node according to the different node information, and obtain the feature embedding of the hyperedge by aggregating the node information on the hyperedge.
[0026] Specifically as follows:
[0027]
[0028]
[0029]
[0030]
[0031] Among them, represents the feature embedding of hyperedge e j in the l-th layer of the hypergraph attention network, Denote the information passed by node \(i\) through hyperedge \(e\) in the \(l\)-th layer hypergraph attention network; j N j represents all nodes connected by hyperedge \(e\), j \(\mu\) (l) denotes the trainable node-level context vector in the \(l\)-th layer hypergraph attention network. \(W_1\) (l) and are trainable transformation matrices; \(\alpha\) i,j represents the attention score of node \(i\) on hyperedge \(e\), j ; represents the feature embedding of node \(i\) in the \((l - 1)\)-th layer hypergraph attention network; \(S(·,·)\) represents calculating the similarity between the node embedding and the context vector; \(D\) is the dimension size, \(a\) and \(b\) are parameters without practical significance, and \(T\) represents the transpose operation.
[0032] Step 2.3: According to the obtained hyperedge feature embeddings, use the attention mechanism to aggregate the information of all hyperedges containing a certain node to obtain the information of the node.
[0033] Specifically as follows:
[0034]
[0035]
[0036]
[0037] Among them, represents the feature embedding of node \(i\) in the \(l\)-th layer hypergraph attention network, represents the information passed from hyperedge \(e\) to node \(i\) in the \(l\)-th layer hypergraph attention network; \(Y\) j represents the set of all hyperedges connected to node \(i\); i ; and \(W_3\) (l) are trainable transformation matrices. \(\beta\) i,j represents the attention score of hyperedge \(e\) on node \(i\). j
[0038] After stacking multiple hypergraph attention layers to obtain high-order information, the embedding representation of each node, that is, each item, is output at the last layer, and finally the item embedding \(h\) under the global hypergraph is obtained. g .
[0039] Step 3: Input the obtained local session graph data into the gated graph neural network for learning to obtain the item embedding representation under the local session graph.
[0040] Specifically, Step 3 includes the following steps:
[0041] Step 3.1: At the beginning layer of the network, each item is embedded from a high-dimensional space into a low-dimensional continuous vector space to obtain the initial embedding representation of each item.
[0042] Step 3.2: Use a gated graph neural network to learn each constructed local session graph.
[0043] Specifically as follows:
[0044]
[0045]
[0046]
[0047]
[0048]
[0049] Among them, A s is the adjacency matrix, and H is the weight matrix; is the update gate in the gated graph neural network, and r i (l) is the reset gate in the gated graph neural network; is an intermediate quantity in the calculation process, represents the feature embedding of node i in the l-th layer of the hypergraph attention network, represents the feature embedding of node i in the (l - 1)-th layer of the hypergraph attention network; ⊙ represents element-wise product, tanh represents the hyperbolic tangent function; σ represents the sigmoid function; W4, W5, W6, U1, U2, U3, b1 are all trainable parameters.
[0050] Take the final output result as the feature embedding of the items in the session, and process all local session graphs in this way to obtain the item embedding representation h s .
[0051] Step 4: Aggregate the item embeddings under the obtained global hypergraph and the item embedding information under the local session graph using the attention mechanism respectively to obtain the session embedding representation under the global hypergraph and the session embedding representation under the local session graph.
[0052] Specifically as follows:
[0053] x i = tanh(W7(h i || p L-i+1 )) + b2) (8)
[0054]
[0055] χ i = q T (W8x i + W9s + b3) (10)
[0056]
[0057] where x i represents the feature embedding representation after fusing the position information of the i-th item in the conversation, p L-i+1 represents the position embedding representation of the i-th item in the conversation; h i represents the feature embedding representation of the i-th item in the conversation; χ i represents the attention score of item i for the entire conversation; s is an intermediate quantity in the calculation process; L represents the length of the conversation, i.e., the number of items in the conversation; S represents the embedding representation of the conversation; W7, W8, W9, q, b2, b3 are all trainable parameters. T represents the transpose operation.
[0058] Perform the operation of the attention mechanism on the item embedding h g under the global hypergraph and the item embedding h s under the local conversation graph respectively, and finally obtain the conversation embedding representation S g under the global hypergraph and the conversation embedding representation S s .
[0059] Step 5: Use contrastive learning to maximize the mutual information between the two conversation embeddings.
[0060] Compare the two conversation embeddings. If the conversation embedding under the global hypergraph and the conversation embedding under the local conversation graph represent the same conversation, then this pair of conversation embeddings is recorded as a positive sample, otherwise it is recorded as a negative sample. Maximize the mutual information between the two conversation embeddings to obtain the contrastive loss under contrastive learning.
[0061] Specifically as follows:
[0062] L s = -logσ(ρ h · ρ l ) - logσ(1 - ρ h · ρ l ) (12)
[0063] where L s represents the contrastive learning loss function, ρ h represents the positive sample, ρ l represents the negative sample, and σ represents the sigmod function.
[0064] Step 6: Calculate the recommendation probability of the candidate item and give the loss function.
[0065] Specifically as follows:
[0066] S = S g + S s (13)
[0067]
[0068]
[0069] L = L c + λL s (16)
[0070] Among them, S is the embedded representation of the session, S g is the session embedded representation under the global hypergraph, S s is the session embedded representation under the local session graph; z i represents the one-hot encoding of the i-th candidate item, represents the predicted probability of the i-th candidate item; L c represents the cross-entropy loss function, L s represents the contrastive learning loss function, L represents the final loss function, and λ is a learnable parameter.
[0071] From step 1 to step 6, the recommended probability of candidate items for a given session sequence is obtained. According to the recommended probability, the contrastive learning session recommendation of the fusion hypergraph is realized.
[0072] Beneficial effects
[0073] The method of the present invention has the following advantages compared with the prior art:
[0074] This method uses a hypergraph structure to model the global structure, can better obtain the high-order information between items, and at the same time uses a local session graph to model the interaction information between items to provide in-session information. For the two different session information, this method uses contrastive learning to compare the two session information, so that the two session information complement each other, effectively making up for the shortage of sparse session data and improving the recommendation performance. Description of the drawings
[0075] Figure 1 is the flowchart of the method of the present invention. Detailed implementation manners
[0076] The present invention will be further described in detail below in conjunction with the specification drawings and embodiments.
[0077] As Figure 1 shown, a contrastive learning session recommendation method for a fusion hypergraph includes the following steps:
[0078] Step A: Construct a global hypergraph and a session graph based on session data;
[0079] Specifically in this embodiment, it is the same as step 1 of the invention content;
[0080] Step B: Input the global hypergraph data into the hypergraph attention network to obtain item embeddings;
[0081] Specifically in this embodiment, it is the same as step 2 of the invention content;
[0082] Step C: Input the session graph data into the gated graph neural network to obtain item embeddings;
[0083] Specifically in this embodiment, it is the same as step 3 of the invention content;
[0084] Step D: Use the attention mechanism to obtain session embeddings under the global hypergraph and the session graph;
[0085] Specifically in this embodiment, it is the same as step 4 of the invention content;
[0086] Step E: Use the contrastive learning method to enhance the two session embeddings;
[0087] Specifically in this embodiment, it is the same as step 5 of the invention content;
[0088] Step F: Calculate the candidate item recommendation probability;
[0089] Specifically in this embodiment, it is the same as step 6 of the invention content.
[0090] Embodiment
[0091] Taking the session sequences "Session 1: [Item 1, Item 3, Item 2, Item 5, Item 7]; Session 2: [Item 2, Item 6, Item 8, Item 10]; Session 3: [Item 4, Item 7, Item 9, Item 11]" as an example, this embodiment will detail the specific operation steps of the contrastive learning session recommendation method with hypergraph fusion of the present invention through specific examples;
[0092] A contrastive learning session recommendation method with hypergraph fusion, as Figure 1 shown, includes the following steps:
[0093] Step A: Construct a global hypergraph and a session graph based on session data;
[0094] Specifically in this embodiment, a global hypergraph and local session graphs are constructed according to the order of items in the session. The global hypergraph is constructed by forming each session into a hyperedge. Therefore, the constructed global hypergraph has 3 hyperedges, and each hyperedge connects all items in sessions 1, 2, and 3 respectively. The local session graph constructs an item interaction graph for each session. Taking session 1 as an example, the nodes in the graph are item 1, item 3, item 2, item 5, and item 7;
[0095] Step B: Input the global hypergraph data into the hypergraph attention network to obtain item embeddings;
[0096] Specifically in this embodiment, first, all items in sessions 1, 2, and 3 are embedded into a low-dimensional space to obtain 100-dimensional low-dimensional embedding representations of each item. The embedding representations are input into the constructed hypergraph attention network, and through learning, 100-dimensional embedding representations of all items in sessions 1, 2, and 3 are obtained;
[0097] Step C: Input the session graph data into the gated graph neural network to obtain item embeddings;
[0098] Specifically in this embodiment, the local session graphs of sessions 1, 2, and 3 constructed in step A are respectively input into the gated graph neural network, and through learning, 100-dimensional embedding representations of each item are obtained;
[0099] Step D: Use the attention mechanism to obtain session embeddings under the global hypergraph and session graph;
[0100] Specifically in this embodiment, the item embeddings under the global hypergraph obtained in step B and the item embeddings under the local session graph obtained in step C are respectively processed. The position information is fused in the item embeddings, and the average value of all items in the session is used as the initial embedding of the session. Taking session 1 as an example, the average value of the embeddings of item 1, item 3, item 2, item 5, and item 7 is used as the global embedding representation of session 1, and the last item in the session, that is, item 7, is used as the intent embedding representation of the session. The attention mechanism is used to obtain the final representation of the session. Finally, session embeddings under the global hypergraph and session graph are obtained;
[0101] Step E: Use the contrastive learning method to enhance the two types of session embeddings;
[0102] Specifically in this embodiment, the embeddings of sessions 1, 2, and 3 under the global hypergraph and the embeddings of sessions 1, 2, and 3 under the session graph are obtained from step D. If the session embeddings under the global hypergraph and the session embeddings under the local session graph represent the same session, then this pair of session embeddings is recorded as a positive sample. That is, the embedding of session 1 under the global hypergraph and the embedding of session 1 under the session graph are a pair of positive samples, the embedding of session 2 under the global hypergraph and the embedding of session 2 under the session graph are a pair of positive samples, and the embedding of session 3 under the global hypergraph and the embedding of session 3 under the session graph are a pair of positive samples. The rest are negative samples. Maximize the mutual information of the two session embeddings to obtain the loss function under contrastive learning and add it to the final loss function for training;
[0103] Step F: Calculate the candidate item recommendation probability;
[0104] Specifically in this embodiment, the recommendation probability is calculated by taking the inner product of the feature representations of sessions 1, 2, and 3 and the candidate items, and the item most likely to be recommended for each session is obtained.
Claims
1. A contrastive learning conversation recommendation method integrating hypergraphs, characterized in that, Including the following steps: Define \(V\) as the set of all items, i.e., the item set, \(V = \{v_1, v_2, v_3, \ldots, v\) n \(\}\), where \(n\) is the number of items; each session is represented as a set \(s\), \(L\) is the length of the session; Step 1: Construct a global hypergraph and a local session graph according to the session data; Step 2: Input the obtained global hypergraph data into a hypergraph attention network for learning to obtain the item embedding representation under the global hypergraph; Step 2.1: At the beginning layer of the network, each item is embedded from a high-dimensional space into a low-dimensional continuous vector space to obtain the initial embedding representation of each item where represents the initial embedding of the nth item; Step 2.2: Input the initial embedding representation of each node into the first-layer hypergraph attention network, give different degrees of attention to each node according to different node information, and obtain the feature embedding of the hyperedge by aggregating the node information on the hyperedge; Step 2.3: According to the obtained hyperedge feature embedding, use the attention mechanism to aggregate the information of all hyperedges containing a certain node to obtain the node information; Obtain high-order information by stacking multiple hypergraph attention layers, output the embedding representation of each node, that is, each item, at the last layer, and finally obtain the item embedding under the global hypergraph; Step 3: Input the obtained local session graph data into a gated graph neural network for learning to obtain the item embedding representation under the local session graph; Take the final output result as the feature embedding of the item in the session, and process all local session graphs in this way to obtain the item embedding representation under the local session graph; Step 4: Aggregate the item embedding under the global hypergraph and the item embedding information under the local session graph using the attention mechanism respectively to obtain the session embedding representation under the global hypergraph and the session embedding representation under the local session graph; Perform operations of the attention mechanism on the item embedding under the global hypergraph and the item embedding under the local session graph respectively, and finally obtain the session embedding representation under the global hypergraph and the session embedding representation under the local session graph; Step 5: Use contrastive learning to maximize the mutual information of the two session embeddings; Compare the two session embeddings. If the session embedding under the global hypergraph and the session embedding under the local session graph represent the same session, then record this pair of session embeddings as positive samples, otherwise record them as negative samples; maximize the mutual information of the two session embeddings to obtain the contrastive loss under contrastive learning; Step 6: Calculate the recommendation probability of candidate items and give the loss function; From Step 1 to Step 6, obtain the recommendation probability of candidate items for the given session sequence; according to the recommendation probability, implement the contrastive learning session recommendation of the fusion hypergraph.
2. The contrastive learning session recommendation method integrating a hypergraph according to claim 1, wherein Step 1 includes the following steps: Step 1.1: Construct a global hypergraph G based on the session data g , G g =(V g , E g ); V g represents the set of vertices in the hypergraph; E g represents the set of hyperedges in the hypergraph, which is composed of all historical session sequences; each session is constructed into a hyperedge, that is and each item Step 1.2: Construct a local session graph G based on the session data s , G s = (V s , E s ); V s represents the vertex set composed of all items in the session; E s represents the set of edges in the session graph; The local conversation graph is constructed based on the current conversation, given a conversation where the set of all items in the conversation constitutes V s , and the directed edges e in the local conversation graph are formed according to the interaction order between the items i , that is represents the item of the i-th interaction in the conversation s 3. The contrastive learning session recommendation method integrating a hypergraph according to claim 1, characterized in that The implementation method of Step 2.2 is as follows: Among them, represents the feature embedding of the hyperedge e j in the l-th layer hypergraph attention network, represents the information passed by the node i through the hyperedge e j in the l-th layer hypergraph attention network; N j represents all the nodes connected by the hyperedge e j ; μ (l) represents the trainable node-level context vector in the l-th layer hypergraph attention network; W1 (l) and are trainable transformation matrices; α i,j represents the attention score of the node i on the hyperedge e j ; represents the feature embedding of the node i in the (l - 1)-th layer hypergraph attention network; S(.,.) represents calculating the similarity between the node embedding and the context vector; D is the dimension size, a and b are parameters without practical significance, and T represents transpose; The implementation method of Step 2.3 is as follows: Among them, h i (l) represents the feature embedding of node i in the l-th layer of the hypergraph attention network, represents the information passed from the hyperedge e j to node i in the l-th layer of the hypergraph attention network; Y i represents the set of all hyperedges connected to node i; and are trainable transformation matrices; β i,j represents the attention score of the hyperedge e j on node i.
4. The contrastive learning session recommendation method integrating a hypergraph according to claim 1, wherein, Step 3 includes the following steps: Step 3.1: At the beginning layer of the network, embed each item from a high-dimensional space into a low-dimensional continuous vector space to obtain the initial embedding representation of each item Step 3.2: Use a gated graph neural network to learn each constructed local session graph, specifically as follows: Among them, A s is the adjacency matrix, and H is the weight matrix; is the update gate in the gated graph neural network, and r i (l) is the reset gate in the gated graph neural network; is an intermediate quantity in the calculation process, represents the feature embedding of node i in the l-th layer of the hypergraph attention network, represents the feature embedding of node i in the (l - 1)-th layer of the hypergraph attention network; ⊙ represents the element-wise product, tanh represents the hyperbolic tangent function; σ represents the sigmoid function; W4, W5, W6, U1, U2, U3, b1 are all trainable parameters.
5. The contrastive learning conversation recommendation method integrating a hypergraph according to claim 1, characterized in that Step 4 includes: x i = tanh(W7(h i || p L-i+1 ) + b2) (13) χ i = q T (W8x i + W9s + b3) (15) Among them, x i represents the feature embedding representation after fusing the position information of the i-th item in the conversation, p L-i+1 represents the position embedding representation of the i-th item in the conversation; h i represents the feature embedding representation of the i-th item in the conversation; χ i represents the attention score of item i for the entire conversation; s is an intermediate quantity in the calculation process; L represents the length of the conversation, that is, the number of items in the conversation; S represents the embedding representation of the conversation; W7, W8, W9, q, b2, b3 are all trainable parameters; T represents the transpose.
6. The contrastive learning session recommendation method integrating a hypergraph according to claim 1, wherein, Step 5 includes: L s = -logσ(ρ h ·ρ l ) - logσ(1 - ρ h ·ρ l ) (17) Among them, L s represents the contrastive learning loss function, ρ h represents the positive sample, ρ l represents the negative sample, and σ represents the sigmod function.
7. The contrastive learning conversation recommendation method integrating a hypergraph according to claim 1, characterized in that, Step 6 includes: S = S g + S s (18) L = L c + λL s (21) Among them, S represents the embedded representation of the session, S g is the session embedded representation under the global hypergraph, S s is the session embedded representation under the local session graph; z i represents the one-hot encoding of the i-th candidate item, represents the predicted probability of the i-th candidate item; L c represents the cross-entropy loss function, L s represents the contrastive learning loss function, L represents the final loss function, and λ is a learnable parameter.
Citation Information
Patent Citations
Session recommendation method based on three-channel graph neural network
CN114547276A
Friend and interest point recommendation method and terminal
CN115146180A