Session recommendation method and device based on graph representation learning, medium and product
Through the method of graph representation learning, a multi-view diagram representation learning session recommendation architecture is constructed, which solves the problems of insufficient recommendation and sparse data of the existing session recommendation method in the absence of user historical data, and achieves a more accurate and dynamic recommendation effect.
Patent Information
- Application Number
- CN202510550994.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-29
AI Technical Summary
Existing session recommendation methods are difficult to achieve dynamic and accurate recommendations when processing lack of user historical data, and face the problems of sparse data and noisy data.
Using a graph representation learning method, a multi-view diagram representation learning session recommendation architecture is constructed, including a dual-branch co-occurrence-semantic feature extraction fusion module, a cross-session graph construction module, a collaborative information extraction module and a multi-intention prediction fusion module, to capture complex relationships between projects and collaborative information between sessions, and to reduce the impact of noise data.
It improves the rationality and accuracy of item recommendations, effectively alleviates the problem of data sparsity, and enhances the performance and robustness of the session recommendation system.
Smart Images

Figure CN120070014A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of conversational recommendation, and in particular to a conversational recommendation method, device, medium and product based on graph representation learning. Background Art
[0002] With the prosperity and popularization of the Internet, a large number of enterprises tend to provide their products or services in the e-commerce mode, such as Amazon, YouTube, etc. How to quickly screen out the content that users are interested in from these massive amounts of information has become an urgent problem to be solved. As an efficient information filtering technology, a recommendation system can mine the interest points of users based on the historical interaction records of users (such as clicks, views, reads, adding to the shopping cart and purchases, etc.), and recommend the content that users are interested in in an automated manner. Specifically, in the past few decades, the traditional recommendation system methods can be divided into content-based models, collaborative filtering-based models, and hybrid models (i.e., the combination of content-based and collaborative filtering models). However, most of the above methods are dedicated to using the complete historical information of users to capture the long-term and statistical preferences of users for recommendation. However, these methods may not be applicable in actual application scenarios: (1) The preferences of users will change and evolve over time, and these methods are difficult to capture the current behavior of users to achieve real-time recommendation. (2) Most importantly, when the historical information of users cannot be obtained, the traditional recommendation system cannot make effective recommendations. Therefore, modeling the short-term preferences and intentions of users based on the current interaction behavior is crucial for achieving dynamic and accurate recommendations.
[0003] To make up for the deficiencies of traditional recommendation methods, session-based recommendation has received increasing attention in recent years. By analyzing the real-time interaction data of users in a short period of time, it can provide relatively accurate personalized recommendations in the absence of user historical data. Existing session-based recommendation methods can be roughly divided into two categories. One is the sequence-based method, which mainly mines users' short-term behaviors and intentions by modeling the sequential information in the session. Such methods usually rely on recurrent neural networks (such as RNN, LSTM, GRU) and attention mechanisms. Although the sequence-based method shows good performance in session-based recommendation, it mainly focuses on sequential information learning and ignores the complex relationships between items. Since the graph structure can flexibly model the non-linear relationships in the session and is introduced into session-based recommendation, the graph-based method can capture the complex relationships between items by constructing the sequence data into a graph, making up for the limitations of the sequence model. At the same time, the graph neural network (GNN) can capture context information faster through the information propagation and aggregation process, showing significant advantages. Due to the special application scenario of session-based recommendation, there is a serious data sparsity problem, and the noisy data in the session has an undeniable impact on the recommendation task. Most existing methods infer users' intentions based on the data of a single session. However, it is difficult to capture the real user intentions when the session is short. Although some methods try to use global session information to alleviate sparsity, the consideration of the connection patterns between sessions is relatively simple, and the potential collaborative information cannot be captured. Secondly, most existing methods do not effectively pay attention to the impact of noisy data in the session, so they cannot effectively use node representations to infer users' intentions, resulting in poor recommendation effects. Summary of the Invention
[0004] The purpose of this application is to provide a session-based recommendation method, device, medium and product based on graph representation learning, which can improve the rationality and accuracy of item recommendation.
[0005] To achieve the above purpose, this application provides the following solutions:
[0006] In the first aspect, this application provides a session-based recommendation method based on graph representation learning, including:
[0007] Build an initial architecture for session-based recommendation with multi-perspective graph representation learning; the initial architecture for session-based recommendation with multi-perspective graph representation learning includes: a dual-branch co-occurrence-semantic feature extraction and fusion module, a cross-session graph construction module, a collaborative information extraction module, and a multi-intention prediction and fusion module; the output end of the dual-branch co-occurrence-semantic feature extraction and fusion module and the output end of the cross-session graph construction module are both connected to the input end of the collaborative information extraction module; the output end of the collaborative information extraction module is connected to the input end of the multi-intention prediction and fusion module;
[0008] Train the initial architecture of multi-view graph representation learning session recommendation using the session dataset and the total loss function to obtain the multi-view graph representation learning session recommendation architecture;
[0009] Obtain the session generated by the interaction between the target user and the recommendation system; the session is an item sequence formed when the target user continuously interacts with multiple items;
[0010] Input the session into the multi-view graph representation learning session recommendation architecture to obtain the recommended item sequence of the target user.
[0011] In a second aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the above-mentioned session recommendation method based on graph representation learning.
[0012] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the above-mentioned session recommendation method based on graph representation learning.
[0013] In a fourth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the above-mentioned session recommendation method based on graph representation learning.
[0014] According to the specific embodiments provided by the present application, the present application discloses the following technical effects:
[0015] The present application provides a session recommendation method, device, medium and product based on graph representation learning. In order to use graph neural networks to capture the complex relationships between items, the present application constructs graphs from two perspectives. For the input session sequence, a directed edge is connected between adjacent items to describe the co-occurrence relationship between items (such as a user buying mobile phone accessories after buying a mobile phone). When reaching the last item, the construction of the co-occurrence relationship item graph is completed. In order to capture the relationships between items from the semantic perspective (such as price, brand, category), the model first initializes the features of the items in the session, and then dynamically constructs a semantic item graph according to the initial features (semantics) of the items. The weight of the edge in the graph is the semantic relevance between items. Finally, the co-occurrence item graph and the semantic item graph are output; for the two item graphs constructed based on the session sequence, on the one hand, the model uses the co-occurrence information extraction network to capture the co-occurrence features of items from the co-occurrence item graph; at the same time, the semantic information extraction network is used to model the semantic associations of items on the semantic item graph, and then the output item features are further integrated to obtain a preliminary representation of the session. Specifically, the model uses a mean aggregation function to obtain order-invariant session features, and introduces order relationships through a recurrent neural network to obtain order-sensitive session features; all sessions are input into the cross-session construction module, which constructs a cross-session graph according to the number of shared items between sessions (the weight of the edge is the number of item overlaps between sessions), and at the same time inputs the order-invariant session features into the co-occurrence collaborative information extraction network to capture the co-occurrence collaborative information between sessions. From the perspective of semantic collaborative information, the model inputs the order-invariant session features and the cross-session graph into the semantic collaborative information extraction network to capture the semantic collaborative information between sessions. Finally, the order-sensitive session features obtained by the recurrent neural network are input into the ordered collaborative information extraction network, which samples multiple collaborative sessions according to the session feature similarity and captures the ordered collaborative information between sessions; for the learned item co-occurrence features and item semantic features, the model further generates the local interest and overall interest of the session, combines various collaborative information, and dynamically integrates multiple prediction results through a multi-intent fusion module to obtain the prediction probability distribution of the session. Finally, considering the hard negative mining auxiliary task to further enhance the performance of the session recommendation system. Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the following-described drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 Schematic diagram of explicit collaborative relationships in an embodiment of the present application;
[0018] Figure 2 Schematic diagram of potential collaborative relationships in an embodiment of the present application;
[0019] Figure 3 Flowchart of a session recommendation method based on graph representation learning in an embodiment of the present application;
[0020] Figure 4 Schematic diagram of multi - perspective cross - session collaborative information in an embodiment of the present application;
[0021] Figure 5 Overall architecture diagram of the session recommendation method of multi - perspective graph representation learning in an embodiment of the present application;
[0022] Figure 6 Schematic diagram for describing the session recommendation problem in an embodiment of the present application;
[0023] Figure 7 Item co - occurrence graph of a session in an embodiment of the present application and adjacency matrix schematic diagram;
[0024] Figure 8 Flowchart for extracting item co - occurrence information in an embodiment of the present application;
[0025] Figure 9 Flowchart for extracting item semantic information in an embodiment of the present application;
[0026] Figure 10 Flowchart for extracting cross - session co - occurrence collaborative information in an embodiment of the present application;
[0027] Figure 11 Flowchart for extracting cross - session semantic collaborative information in an embodiment of the present application;
[0028] Figure 12 Flowchart for extracting cross - session ordered collaborative information in an embodiment of the present application;
[0029] Figure 13 Flowchart for multi - intent prediction and fusion in an embodiment of the present application. Detailed implementation manners
[0030] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0031] To make the above - mentioned objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the drawings and specific implementation manners.
[0032] Most existing session recommendation methods only focus on the behaviors within the current session, ignoring the potential collaborative relationships that may exist between sessions, and at the same time facing the dilemma of data sparsity. To make up for this deficiency, this application aims to mine the collaborative information between sessions from multiple perspectives, provide multiple potential contexts, and effectively alleviate the data sparsity problem. Specifically, it is divided into three perspectives:
[0033] Based on co-occurrence relationships: This perspective captures cross-session collaborative information based on the co-occurrence relationships of user behaviors. If there are the same user behaviors between sessions (such as Figure 1 the same product is clicked in different sessions), it can be considered that there is a collaborative relationship between them. By connecting these sessions through the common behaviors, session information can be propagated between sessions. Through the information propagation mechanism, sessions with the same behavior can share each other's context information, thereby alleviating the data sparsity problem.
[0034] Based on semantic relationships: In addition to introducing collaborative information through the same behaviors between sessions, collaborative information can also be considered through the similarity of behaviors. Specifically, although there are no identical click behaviors in two sessions, different users may be interested in similar categories of products, such as the same brand, similar price ranges, or the same type of products (such as Figure 2 , there is a similar association with electronic products in two sessions). From this perspective, by extracting behavior attributes (such as product category, brand, price, etc.) and measuring the similarity of sessions in the feature space, collaborative information that has no overlap in behavior but has potential semantic connections can be captured.
[0035] Based on the orderliness of behaviors: In some scenarios, user behaviors usually show a certain temporal order, reflecting the dynamic pattern of user interests (such as Figure 1 , the user may pay attention to shoes after buying pants or a top). By considering the order relationship of behaviors in the session and transmitting this order information between different sessions, the ordered dependence relationship between sessions can be captured. Especially in scenarios where user behaviors have obvious temporal dependence, it can more accurately reflect the behavior changes and interest trends.
[0036] In summary, by considering this session collaborative information, the data sparsity problem of session recommendation can be effectively alleviated, and the performance of the recommendation system can be improved.
[0037] When the recommendation system infers the conversation intention, it is necessary to comprehensively consider the various intentions that may exist in the conversation. Such as multiple collaborative information, as well as the long-term and short-term intentions of the current conversation. It is worth noting that the conversation representations from different perspectives represent different intentions. Therefore, this application introduces an adaptive multi-intention fusion module to effectively fuse the conversation representations from different perspectives to ensure the reasonable reasoning of user intentions. In short, each conversation representation needs to be normalized first, which helps to unify the feature representations from different perspectives and avoid information distortion caused by inconsistent feature scales. On this basis, the module predicts the conversation representation of each perspective separately, and then weightedly fuses the prediction results. Through this adaptive fusion module, the model can integrate multiple conversation representations, thereby improving the accuracy of the conversation recommendation system. When dealing with long-tail users (demand tends to be niche products or unpopular content) and conversations with complex content, the recommendation accuracy of the conversation recommendation system is often low. In order to meet this challenge, this application introduces an auxiliary task that enables the model to pay more attention to such conversation samples. This mechanism increases the optimization efforts for sessions with large prediction errors, prompting the model to better learn the characteristic expressions of these sessions, thereby improving the modeling capabilities for complex and rare samples, avoiding overfitting of the model to common samples, and further enhancing the overall performance and robustness of the recommendation system. The Chinese meanings of the symbols involved in this application are shown in Table 1.
[0038] Table 1 Main symbol table
[0039]
[0040] In an exemplary embodiment, Figure 3 As shown, a conversation recommendation method based on graph representation learning is provided, including:
[0041] Step 301: Build an initial architecture for multi-view graph representation learning session recommendation.
[0042] Step 302: Using the conversation data set and the total loss function, the multi-view graph representation learning conversation recommendation initial architecture is trained to obtain the multi-view graph representation learning conversation recommendation architecture.
[0043] Step 303: Acquire a session generated by the target user interacting with the recommendation system; a session is a sequence of items formed when the target user continuously interacts with multiple items.
[0044] Step 304: Input the conversation into the multi-view graph representation learning conversation recommendation framework to obtain a sequence of recommended items for the target user.
[0045] The initial architecture for multi-view graph representation learning session recommendation includes: a dual-branch co-occurrence-semantic feature extraction and fusion module, a cross-session graph construction module, a collaborative information extraction module, and a multi-intent prediction and fusion module. The output ends of both the dual-branch co-occurrence-semantic feature extraction and fusion module and the cross-session graph construction module are connected to the input end of the collaborative information extraction module. The output end of the collaborative information extraction module is connected to the input end of the multi-intent prediction and fusion module.
[0046] The dual-branch co-occurrence-semantic feature extraction and fusion module includes: a co-occurrence branch, a semantic branch, and a preliminary representation integration module.
[0047] The output ends of both the co-occurrence branch and the semantic branch are connected to the input end of the preliminary representation integration module. The output end of the preliminary representation integration module is connected to the input end of the collaborative information extraction module.
[0048] The co-occurrence branch includes a co-occurrence item graph construction unit and a co-occurrence information extraction unit connected in sequence. The co-occurrence information extraction unit uses a gated graph neural network. The output end of the co-occurrence information extraction unit is connected to the input end of the preliminary representation integration module.
[0049] The semantic branch includes a semantic item graph construction unit and a semantic information extraction unit connected in sequence. The semantic information extraction unit uses a multi-level graph convolutional network. The output end of the semantic information extraction unit is connected to the input end of the preliminary representation integration module.
[0050] The preliminary representation integration module includes: a sequential-invariant session feature extraction module and a sequential-sensitive session feature extraction module connected in sequence; the output ends of the sequential-sensitive session feature extraction module and the sequential-invariant session feature extraction module are connected to the input end of the collaborative information extraction module; the sequential-invariant session feature extraction module uses a mean aggregation function; the sequential-sensitive session feature extraction module uses a recurrent neural network.
[0051] The co-occurrence item graph construction unit is used to generate a co-occurrence item graph based on the input session.
[0052] The co-occurrence information extraction unit is used to extract co-occurrence features based on the co-occurrence item graph.
[0053] The semantic item graph construction unit is used to generate a semantic item graph based on the input session.
[0054] The semantic information extraction unit is used to extract semantic features based on the semantic item graph.
[0055] The sequential-invariant session feature extraction module is used to generate a first sequential-invariant session feature and a second sequential-invariant session feature based on the co-occurrence features and semantic features.
[0056] The sequential-sensitive session feature extraction module is used to generate sequential-sensitive session features based on the co-occurrence features and semantic features.
[0057] The co-occurrence item graph is a directed item graph, where represents the set of items in the co-occurrence item graph ; represents a session , represents the first item in session S; represents the session the second item in; represents the session the th item in; represents the session the Lth item in; represents the ith item in session S; represents the co-occurrence item graph the set of all edges in; represents the session the th item in; item and item in the co-occurrence item graph the corresponding edge weight of is: ; The adjacency matrix of the co-occurrence item graph includes an out-degree neighbor information matrix and an in-degree neighbor information matrix ;
[0058] The element in the out-degree neighbor information matrix is:
[0059] ;
[0060] The element in the in-degree neighbor information matrix is:
[0061] ;
[0062] The semantic item graph is an undirected item graph, where represents all items in the session, represents the edge between item pairs ;
[0063] The adjacency matrix of the semantic item graph the element represents item and item the weight of the corresponding edge in the semantic item graph ;
[0064] ;
[0065] wherein is the initial weight of the corresponding edge between item and item in the semantic item graph ;
[0066] ;
[0067] wherein is the initial vector obtained by encoding item ; is the transpose of the initial feature vector of item ; is the initial feature vector of item ;
[0068] The cross - session graph construction module is used to generate a cross - session graph based on a session sequence ;
[0069] wherein is the point corresponding to each session in the session sequence represents the i - th session in the session set constituted by the session sequence is the edge set in the cross - session graph represents the edge constituted by the session and the session with interaction items; the element in the adjacency matrix of the cross - session graph represents the weight of the edge ; ; represents the item set of the session ; represents the item set of the session ;
[0070] The collaborative information extraction module includes co - occurrence collaborative information extraction branches, semantic collaborative information extraction branches, and ordered collaborative information extraction branches in parallel
[0071] The co - occurrence collaborative information extraction branch is used to extract co - occurrence collaborative information based on co - occurrence features, first - order invariant session features, and the adjacency matrix of the cross - session graph
[0072] The semantic collaborative information extraction branch is used to extract semantic collaborative information based on semantic features, second-order invariant session features, and the adjacency matrix of the cross-session graph ;
[0073] The ordered collaborative information extraction branch is used to extract ordered collaborative information based on co-occurrence features and order-sensitive session features.
[0074] The multi-intent prediction fusion module is used to determine the local interest feature, the first overall interest feature, and the second overall interest feature of the session, and decode the local interest feature, the first overall interest feature, the second overall interest feature, the co-occurrence collaborative information, the semantic collaborative information, and the ordered collaborative information respectively, and multiple probability predictions can be obtained. The weighted max pooling and average pooling are used to fuse the multiple probability predictions to obtain the prediction result.
[0075] The total loss function is:
[0076] ;
[0077] ;
[0078] ;
[0079] ;
[0080] where is the total loss function, is the cross-entropy loss function is the dynamic loss value function; is the parameter for L2 regularization of all trainable parameters of the model; is the prediction result of the i-th item in the candidate item set; is the true result of the i-th item in the candidate item set; N is the total number of item categories is the bias value control factor; is the final prediction score of the item corresponding to the label.
[0081] Most of the existing session recommendation methods only focus on the user behavior within the session, while ignoring the collaborative information existing between sessions. Due to the serious problem of session data sparsity, especially when the session is short, it is difficult to use the session data for effective user intention reasoning.
[0082] Table 2 Statistical Information Table of Public Datasets
[0083]
[0084] Table 2 shows the basic statistics of the session recommendation popular datasets in five real scenarios. It can be seen that although the average session length of Last.fm is relatively long (9.16), the number of short sessions is still 1,136,909, indicating that the sparsity problem still exists; while short sessions dominate in the remaining four datasets. These data show that short sessions dominate in the session recommendation scenario, and the sparsity of the data poses a great challenge to capturing user preferences and modeling complex relationships.
[0085] This application considers using global information to enhance the performance of session recommendation and introduces collaborative information from multiple cross-session perspectives to alleviate the data sparsity problem. For example Figure 4 , based on the co-occurrence relationship, sessions with the same items (trousers, shoes) are screened to introduce potential contexts for purchasing tops for the current session; or based on the semantic (category) representation of items, sessions with the intention of purchasing clothing are filtered to provide collaborative information for purchasing tops for the current session.
[0086] Data noise: Session recommendation is mainly applied in the instant recommendation scenario. The content that users are interested in in the future may be highly relevant to their recent behaviors. Therefore, existing methods tend to regard the last item as the session intention. In this case, the model is easily affected by noisy data, resulting in incorrect inference of user intentions. For example Figure 4 , if the model mainly focuses on the recent behaviors of the session, then the racket will be predicted as the session intention. However, in fact, the similarity between the racket and the labeled item top is very low. In fact, the racket has nothing to do with the overall session intention (clothing items). To more accurately understand and predict the needs of users, this application considers both the overall and local representations of the session, enabling the model to capture both instant intentions and overall intentions and reducing the impact of noisy items. For example Figure 4 , specifically, by modeling the overall intention of the session, a more general session intention expression (clothing) can be obtained. At the same time, relevant collaborative sessions can also be screened to provide contexts related to the overall intention (clothing). Finally, in the intention inference stage, by comprehensively considering various session representations, the impact of noisy data on intention inference can be reduced.
[0087] The overall architecture of the model is as shown in Figure 5 , and the framework mainly consists of the following four parts:
[0088] The description of the session recommendation problem is as shown in Figure 6 , the item set is defined as ; where N represents the total number of items. The session set is represented as where M represents the total number of sessions in the dataset. Given a session S represented as , represents the interaction terms, and L is the length of the session. Session-based recommendation aims to predict the user's behavior at the next time step (i.e., ). Specifically, given a session, the goal of session recommendation is to recommend the top μ items in the item set that are most likely to be clicked or purchased by the user at the next timestamp, where .
[0089] Step 1: Local item graph construction.
[0090] (1) Co-occurrence item graph construction.
[0091] Session recommendation needs to obtain the user's intention based on the item information in the session. To capture the complex relationships between items and infer the user's intention, this application constructs an item graph based on co-occurrence relationships for each session and uses a graph neural network to learn item representations. Specifically, according to the order of item appearance in the session sequence, a co-occurrence item graph is constructed, which means that if two items appear successively in the session, these two items will be connected in the item graph. The direction of the connection represents their order relationship, and the strength of the connection can reflect the weight of their co-occurrence. For example, for the session , define the directed item graph , where represents the set of items in the directed item graph , represents the set of all edges in the graph. The adjacency matrix in the directed item graph is represented by , is used to generate two adjacency relationship matrices, namely the in-degree neighbor matrix and the out-degree neighbor matrix . represents the value of the i-th row and j-th column, that is, the weight of the corresponding edge and the item in the directed item graph corresponding to the edge . is expressed as follows:
[0092] (1)
[0093] To focus on the directionality of the edges in the directed item graph , then decompose the adjacency matrix into two parts: the in-degree neighbor matrix and the out-degree neighbor matrix , and perform normalization operations on these two matrices respectively: Since the graph is a directed graph where the neighbors of a node can be divided into in-degree neighbors (nodes pointing to this node) and out-degree neighbors (nodes pointed to by this node), then introduce the in-degree neighbor matrix Out-degree neighbor matrix Consider these two adjacency relationships separately. Both are generated by the adjacency matrix of the graph . The edge weight corresponding to nodes i and j in the out-degree neighbor matrix is represented as follows:
[0094] (2)
[0095] Similarly represents the edge weight between nodes i and j in the in-degree neighbor matrix , which is specifically represented as .
[0096] Given a session , adding a directed edge to each pair of adjacent items can obtain the edge set , item set . A co-occurrence item graph can be constructed based on the edge set and item set. Then, a directed graph is generated based on the item set and edge set. The nodes in the graph represent items, and the edges represent the relationships between adjacent items in the session. The directed item graph of the session and its adjacency matrix are as shown in Figure 7 .
[0097] (2)Semantic item graph construction.
[0098] To capture the relationships between items from the semantic perspective of items, the model dynamically constructs a semantic item graph , represents all items in the session, represents the edge between item pair . Specifically, the weight of item pair is defined as follows, where , representing the initial feature vector of item pair :
[0099]
[0100] Then, the function is used to normalize the weights to obtain the adjacency matrix of the semantic item graph . Finally, the weight of edge in graph is represented as follows:
[0101] (4)
[0102] Step 2: Item representation learning.
[0103] For the input session sequence, each item is first mapped into a high-dimensional vector, and the item features are normalized to mitigate the impact of item popularity bias (the tendency to recommend popular, high-frequency, and trendy items while ignoring unpopular items). For a session , after encoding, the item feature matrix is obtained, where , L represents the maximum session length, and d is the feature dimension size. Finally, the initial feature of the session is , which is expressed as follows:
[0104] (5)
[0105] Among them, represents calculating the vector norm. Then, a graph neural network (GNN) is used to learn the feature representation of items on two types of item graphs. The process of extracting item co-occurrence information is as Figure 8 shown.
[0106] (1) Co-occurrence information extraction.
[0107] Specifically, the model uses a gated graph neural network (co-occurrence information extraction network) to process the in-degree neighbor matrix and the out-degree neighbor matrix obtained in step 1 to facilitate the message passing process between items. The context information received by item on the directed item graph is defined as , where is a learnable matrix, and the symbol [;] represents the concatenation operation, is the initial feature representation of the session . (the i-th row of the out-degree neighbor matrix and the in-degree neighbor matrix respectively) represents the neighbor information of the out-edge and in-edge of item in the directed item graph . The feature representation of item is updated using the context information as follows:
[0108] (6)
[0109] (7)
[0110] (8)
[0111] (9)
[0112] Among them, are learnable parameters, and ⊙ represents element-wise multiplication. is a function, representing the candidate features of node in the τ-th layer of the gated graph neural network, which are used to update the feature representation of node in the current network layer. while represents updating the features of node using context information. The control of the network is regulated by the update gate and the reset gate , which are crucial throughout the network and control the processing and filtering of information. Subsequently, let , and finally the output representation of the co-occurrence information extraction network is , which represents the co-occurrence features of the project.
[0113] The process of project semantic information extraction is as shown in Figure 9 .
[0114] (2) Semantic information extraction.
[0115] From a semantic perspective, the model uses a multi-level graph convolutional network (semantic information extraction network) to process the adjacency matrix of the semantic item graph in step 1 to capture the semantic relationships between items. Specifically, the propagation process of this graph network is defined as follows:
[0116] (10)
[0117] (11)
[0118] (12)
[0119] where is the initial feature of node , represents the semantic information that node obtains from adjacent items, are learnable parameters used to generate the contribution weights of the item features output by each layer of the network and the initial item features, is a function, represents the i-th row of the adjacency matrix of the -th layer. is composed of Dynamically generated. For the detailed process, please refer to the construction process of the semantic project graph. Finally, the output of the semantic information extraction network is represented as , which represents the semantic features of the project.
[0120] Step 3: Cross-session collaborative information extraction.
[0121] For the co-occurrence features obtained in Step 2 and semantic features , first, feature fusion is performed to obtain the preliminary representation of the session. On the one hand, the order-invariant intention of the session can be extracted based on the mean aggregation function. Considering the ordered relationship of the projects, the model further adopts a gated recurrent neural network ) to extract the order-sensitive session intention, and both are applied to the following collaborative information extraction network. The cross-session co-occurrence collaborative information extraction process is as Figure 10 shown.
[0122] (1) Co-occurrence collaborative information extraction.
[0123] To obtain the explicit collaborative information between sessions (sessions sharing the same project can provide session information for each other), the model constructs a cross-session graph: each session is regarded as a supernode, and the edge weight between any two nodes in the graph is quantified by the overlap degree of the corresponding session item types. Specifically, the global co-occurrence session graph is defined as , where , the adjacency matrix of the graph is defined as , M represents the total number of sessions. The edge weight corresponding to session and in the graph is (for the i-th row and j-th column of ) and is calculated as follows:
[0124] (13)
[0125] where represents the item set of session . This collaborative information extraction network uses the order-invariant features of the session as the representation of the session node. For session S, the initial representation is defined as:
[0126] (14)
[0127] where, L represents the session length, is the item co-occurrence feature. The identity matrix is used to enhance to obtain . And the diagonal matrix is obtained, where each diagonal element Represents a session And the weight sum of adjacent sessions. Using Represents the initial feature matrix of the session, and , the convolution operation is defined as follows:
[0128] (15)
[0129] Generally speaking, each session can obtain collaborative information from sessions sharing items. The data sparsity problem can be alleviated by capturing cross-session information. At the same time, in order to avoid the over-smoothing phenomenon in the multi-layer convolutional network, the residual connection method is further considered, so the finally captured co-occurrence collaborative information Is represented as follows:
[0130] (16)
[0131] (2) Semantic collaborative information extraction.
[0132] Since there may be no direct shared items between sessions, considering that there may be similar behaviors or preferences between sessions, in order to capture this potential collaborative information, this application introduces a semantic collaborative information network, aiming to capture cross-session potential collaborative information to further improve the session representation quality. Similarly, the sequential invariant features of the session will be used as the initial representation of the session and input into the semantic collaborative information network, Represents the item semantic information:
[0133] (17)
[0134] Specifically, this network uses a session-aware attention mechanism to capture the correlation between sessions. Each session is linearly combined with the attention scores of other sessions, and the session The cross-session information that can be obtained is as follows:
[0135] (18)
[0136] Among them, Represents the neighbor session of session , Then represents session And the importance weight of the neighbor session, and the weight is specifically represented as follows:
[0137] (19)
[0138] Among them Is a trainable parameter, Represents session On the global session graph The weight corresponding to the edge, represents the element-wise product. Then, the function is used to normalize all the neighbor weights connected to the session :
[0139] (20)
[0140] To capture deeper semantic collaborative information, a gated mechanism is further used to enhance the session-aware attention network based on the above:
[0141] (21)
[0142] (22)
[0143] (23)
[0144] where is the session feature with the initial order unchanged, is the sigmoid activation function, represents the contribution weight of the output of the current layer network to the initial session feature , and the final output is expressed as , representing the captured cross-session semantic collaborative information. The cross-session semantic collaborative information extraction process is as shown in Figure 11 , where the superscript τ represents the network output quantity.
[0145] (3) Ordered collaborative information extraction.
[0146] Considering that user behaviors in real scenarios show a certain order (for example, after purchasing a mobile phone, users may purchase mobile phone accessories such as chargers and phone cases), this application further considers passing the order-sensitive representation of sessions in a global perspective to capture the ordered dependencies between sessions. At this time, the order-sensitive features of sessions are used as the input of the network. Specifically, the initial session representation obtained by using the item co-occurrence information is:
[0147] (24)
[0148] The item features are processed by a gated recurrent neural network (GRU), and then its average representation is obtained as the network input. Then, by calculating the similarity scores between session representations, several sessions with higher similarity are selected for collaboration. The calculation formula for the session similarity score is as follows:
[0149] (25)
[0150] where , M represents the total number of sessions, They are the features of the current session S. The sessions are sorted according to the Sim value. For session S, F samples with higher Sim values are selected as collaborative sessions, and cross-session ordered collaborative information is obtained as follows:
[0151] (26)
[0152] (27)
[0153] Among them, is the temperature parameter for adjusting the contribution intensity, represents the contribution weight of the collaborative session , represents the captured cross-session ordered collaborative information. The cross-session ordered collaborative information extraction process is as shown in Figure 12 .
[0154] Step 4: Multi-perspective session representation and fusion.
[0155] The model has obtained cross-session co-occurrence collaborative information , semantic collaborative information and ordered collaborative information . However, the intention of the current session has not been inferred. Therefore, it is necessary to consider the two types of item features obtained in the second step to infer the user interest.
[0156] Since the session recommendation system is mostly applied in the instant recommendation scenario, the user's future intention has a strong correlation with the recent behavior. Therefore, the model considers using the item representation at the last time step in the item co-occurrence feature and item semantic feature as the local interest of the session. MLP(*) represents the multi-layer perceptron, which is specifically represented as follows:
[0157] (28)
[0158] Among them represents the local interest, and [;] represents feature concatenation. Select the k most recent user behaviors in session S as the overall intention of the session , which can capture a more immediate overall intention and avoid overconsidering early interfering behaviors. To alleviate the possible noisy data in the recent behaviors, the model further considers the overall interest of the session. From the item semantic perspective (item semantic feature):
[0159] (29)
[0160] Among them Represents the overall intention of the conversation. By considering multiple recent actions simultaneously to obtain the overall interest representation of the conversation, it can alleviate the impact of noisy data and introduce richer user intentions.
[0161] Similarly, from the perspective of item co-occurrence (item co-occurrence representation), a soft attention mechanism is used to obtain the overall interest. Specifically, a learnable position matrix is used Then the reversed position features are combined with the learned item co-occurrence features The feature of the i-th item is represented as:
[0162] (30)
[0163] where are learnable parameters, and tanh(*) is the activation function. The overall user intention is represented as:
[0164] (31)
[0165] (32)
[0166] Here and are attention parameters for learning the weights of each item in the conversation is the order-invariant representation of session S after introducing position information:
[0167] (33)
[0168] In summary, the overall interest captured by the attention mechanism is . So far, the model has obtained local, overall interests and cross-session collaborative information.
[0169] However, these session representations are in different feature spaces, and directly fusing them during the prediction stage will instead lead to confusion in user intentions. Therefore, this application introduces an adaptive multi-intention fusion module to effectively fuse these session representations from different perspectives and ensure reasonable inference of user intentions. Specifically, for the candidate item set feature matrix represents the number of items. First, perform L2 normalization on it to obtain , and at the same time perform the same normalization operation on the obtained multiple session representations, which can effectively alleviate the popularity bias. Then, for each representation, perform prediction respectively to obtain the corresponding prediction distribution (using the session representation to predict the probability score of each item in the candidate item set when being recommended):
[0170] (34)
[0171] Here represents the predicted distribution decoded from the session local interest. Similarly, for the two session global interests , the collaborative information from the three perspectives is decoded separately, and multiple probability predictions can be obtained. To effectively aggregate the multiple prediction results, the module uses weighted max pooling and average pooling to fuse them. Among them, max pooling retains the individual intentions of multiple session representations (highlighting the most significant features from multiple session representations and retaining the uniqueness of different perspectives), while average pooling captures the common intentions (paying attention to the common points between different perspectives and reflecting the general trend in the session representations). The final prediction result is expressed as follows:
[0172] (35)
[0173] Among them, is the activation function is the learnable fusion parameter, and represent max pooling and average pooling respectively. The final prediction result represents the probability that each item in the candidate item set will be selected by the user at the next time step of the current session. The larger the probability value of an item, the more it conforms to the user's current interest tendency. By sorting the probability scores of the candidate items and selecting multiple items with the highest predicted probabilities to generate the recommendation list, session recommendation is achieved. The multi-intention prediction and fusion process is as Figure 13 shown. The final model converts the true label label of the session into a one-hot vector , and uses the cross-entropy loss function as an index to optimize the model parameters:
[0174] (36)
[0175] Considering the differences between session samples, the prediction difficulty of different samples may vary greatly. For session samples that are difficult to predict, the prediction scores of the model may deviate significantly from the true labels. To enable the model to handle these difficult-to-classify session samples more effectively, we introduce an additional dynamic loss value in the loss function as an auxiliary task of the model (hard sample mining) to increase the model's attention to difficult-to-classify samples, which is defined as follows:
[0176] (37)
[0177] (38)
[0178] Among them, Represents the final predicted score for the item corresponding to the label, represents the deviation between the predicted value and the true value of the sample, that is, the prediction difficulty of the sample in the current round. The model controls the magnitude of this deviation value through . The model unifies the main recommendation task and the auxiliary task to enhance the performance of session recommendation. The final loss can be expressed as:
[0179] (39)
[0180] where, represents the parameter for L2 regularization of all trainable parameters of the model. Among them, directly reflects the gap between the predicted value of the model and the true label. The goal of is to minimize this gap as much as possible, so that the recommended results of the model are closer to the actual needs of users. The optimization process of this loss is the core of model training. By dynamically assigning higher loss values to samples with greater prediction difficulty (i.e., samples with a large deviation between the predicted result and the true result), the model is forced to pay more attention to these samples, thereby improving the overall model performance.
[0181] Table 3 Comparison effect table with existing baseline models
[0182]
[0183] Table 3 shows the recommendation effects of the model on four public datasets, including: Tmall and Dig-inetica datasets (shopping and transaction data on e-commerce platforms), RetailRocket (click and purchase dataset of users on retail platforms), and Last.fm, a classic benchmark dataset for music recommendation. The evaluation metrics used are HR@20 and MRR@20. HR@20 (hit rate) measures whether the labeled item appears in the top 20 items of the recommendation list, regardless of its specific ranking. MRR@20 (mean reciprocal rank) calculates the reciprocal average of the rankings of the labeled item in the top 20 items of the recommendation list, which is used to evaluate the recommendation accuracy and ranking quality of the model. The higher the HR metric, the higher the probability that the recommended list contains content of interest to users. The higher the MRR metric, the higher the ranking of the content of most interest to users in the recommendation list. It can be seen from the last column of the table that the experimental results of this application have a significant improvement compared with the baseline model. Further indicating the effectiveness of the model proposed in this application in the session recommendation task.
[0184] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements a session recommendation method based on graph representation learning.
[0185] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, which, when executed by a processor, implements the steps in the above method embodiments.
[0186] In an exemplary embodiment, a computer program product is provided, including a computer program, which, when executed by a processor, implements the steps in the above method embodiments.
[0187] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0188] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0189] The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.
[0190] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0191] Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A conversation recommendation method based on graph representation learning, characterized in that: include: Build the initial architecture for multi-view graph representation learning and conversational recommendation; The multi-view graph representation learning conversation recommendation initial architecture includes: a dual-branch co-occurrence-semantic feature extraction fusion module, a cross-session graph construction module, a collaborative information extraction module and a multi-intention prediction fusion module; the output end of the dual-branch co-occurrence-semantic feature extraction fusion module and the output end of the cross-session graph construction module are both connected to the input end of the collaborative information extraction module; the output end of the collaborative information extraction module is connected to the input end of the multi-intention prediction fusion module; Using the conversation dataset and the total loss function, the multi-view graph representation learning conversation recommendation initial architecture is trained to obtain the multi-view graph representation learning conversation recommendation architecture; Acquire a session generated by the target user interacting with the recommendation system; the session is a sequence of items formed when the target user continuously interacts with multiple items; The session is input into the multi-view graph representation learning session recommendation architecture to obtain a sequence of recommended items for the target user.
2. The conversation recommendation method based on graph representation learning according to claim 1, characterized in that: The dual-branch co-occurrence-semantic feature extraction fusion module includes: a co-occurrence branch, a semantic branch and a preliminary representation integration module; The output end of the co-occurrence branch and the output end of the semantic branch are both connected to the input end of the preliminary representation integration module; the output end of the preliminary representation integration module is connected to the input end of the collaborative information extraction module; The co-occurrence branch includes a co-occurrence item graph construction unit and a co-occurrence information extraction unit connected in sequence; the co-occurrence information extraction unit adopts a gated graph neural network; the output end of the co-occurrence information extraction unit is connected to the input end of the preliminary representation integration module; The semantic branch includes a semantic item graph construction unit and a semantic information extraction unit connected in sequence; the semantic information extraction unit adopts a multi-level graph convolutional network; the output end of the semantic information extraction unit is connected to the input end of the preliminary representation integration module; The preliminary representation integration module includes: an order-invariant session feature extraction module and an order-sensitive session feature extraction module connected in sequence; the output ends of the order-sensitive session feature extraction module and the order-invariant session feature extraction module are connected to the input end of the collaborative information extraction module; the order-invariant session feature extraction module adopts a mean aggregation function; the order-sensitive session feature extraction module adopts a recurrent neural network; The co-occurrence item graph construction unit is used to generate a co-occurrence item graph based on the input conversation; The co-occurrence information extraction unit is used to extract co-occurrence features based on the co-occurrence item graph; The semantic item graph construction unit is used to generate a semantic item graph based on the input conversation; The semantic information extraction unit is used to extract semantic features based on the semantic item graph; The order-invariant conversation feature extraction module is used to generate a first order-invariant conversation feature and a second order-invariant conversation feature based on the co-occurrence feature and the semantic feature; The order-sensitive conversation feature extraction module is used to generate order-sensitive conversation features based on co-occurrence features.
3. The conversation recommendation method based on graph representation learning according to claim 2, characterized in that: The co-occurrence item graph is a directed project graph, where Representing a co-occurrence item graph The collection of items in ; Represents a session, , represents the first item in session S; represents session The second item in; Represents a session The Items; represents a session The Lth item in ; represents the i-th item in session S; Representing a co-occurrence item graph The set of all edges in ; represents a session The Projects and projects in the co-occurrence project graph The weight of the corresponding edge in is: ; The adjacency matrix of the co-occurrence item graph Including out-degree neighbor information matrix and the in-degree neighbor information matrix; The out-degree neighbor information matrix The elements in are: ; The in-degree neighbor information matrix The elements in are: ; The semantic item graph is an undirected item graph, wherein, represents all items in the session, Indicates the project The edge between The semantic project map The adjacency matrix of The meta in represents items and items In the semantic project diagram The corresponding side The weight of ; in, For Project and Projects In the semantic project diagram The corresponding side The initial weight of ; Among them, for the project The initial vector obtained after encoding; For Project The transpose of the initial eigenvector of ; For Project The initial eigenvector of .
4. The conversation recommendation method based on graph representation learning according to claim 3 is characterized in that: The cross-session graph building module is used to generate a cross-session graph based on the session sequence ; in, is the point corresponding to each session in the session sequence; Represents a session set consisting of a sequence of sessions The i-th and session in the graph; is the edge set in the cross-session graph; represents the session with interaction terms and Session The edges formed; cross-session graph The adjacency matrix of Elements in Represents edge The weight of ; Represents a session A collection of items; Represents a session A collection of items.
5. The conversation recommendation method based on graph representation learning according to claim 4, characterized in that: The collaborative information extraction module includes a parallel co-occurrence collaborative information extraction branch, a semantic collaborative information extraction branch and an ordered collaborative information extraction branch; The co-occurrence collaborative information extraction branch is used to extract information based on co-occurrence features, first order invariant session features and cross-session graphs. Adjacency matrix of , extracting co-occurrence information; The semantic collaborative information extraction branch is used to extract information based on semantic features, second order invariant session features and cross-session graphs. Adjacency matrix of , extracting semantic collaboration information; The ordered collaborative information extraction branch is used to extract ordered collaborative information based on co-occurrence features and order-sensitive session features.
6. The conversation recommendation method based on graph representation learning according to claim 5, characterized in that: The multi-intention prediction fusion module is used to determine the local interest features, the first overall interest features, the second overall interest features of the conversation, and decode the local interest features, the first overall interest features, the second overall interest features, the co-occurrence collaborative information, the semantic collaborative information and the ordered collaborative information respectively to obtain multiple probability predictions. Weighted maximum pooling and average pooling are used to fuse the multiple probability predictions to obtain the prediction result.
7. The conversation recommendation method based on graph representation learning according to claim 6, characterized in that: The total loss function is: ; ; ; ; in, is the total loss function, is the cross entropy loss function; is the dynamic loss value function; For all trainable parameters of the model Parameters for L2 regularization; is the prediction result of the i-th item in the candidate item set; is the true result of the i-th project in the candidate project set; N is the total number of project types; is the deviation value control factor; is the final prediction score of the item corresponding to the label.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the conversation recommendation method based on graph representation learning described in any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the conversation recommendation method based on graph representation learning described in any one of claims 1 to 7 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the conversation recommendation method based on graph representation learning described in any one of claims 1 to 7 is implemented.