A Session Recommendation Method, Device, Medium and Product Based on Graph Representation Learning

Through the multi-view diagram, learning the session recommendation architecture is represented, combining the perspectives of co-occurrence, semantics and behavioral order, the data sparsity and noise problems in session recommendation are solved, achieving a more accurate and stable recommendation effect.

CN120070014BActive Publication Date: 2025-07-08ZHEJIANG NORMAL UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510550994.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-07-08
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

Existing session recommendation methods have data sparsity problems when dealing with user short-term preferences and intentions, and fail to effectively utilize the impact of synergistic information and noise data between sessions, resulting in poor recommendation results.

Method used

A multi-perspective diagram is constructed to represent learning session recommendation architecture, mine the coordinated information between sessions through the perspectives of co-occurrence relationship, semantic relationship and behavioral order, and introduce an adaptive multi-intention fusion module to comprehensively consider local and overall intentions, alleviate data sparsity and reduce noise impact.

Benefits of technology

Improves the accuracy and robustness of session recommendations, especially when dealing with long-tail users and complex sessions, and improves the performance of the recommendation system and resistance to noise data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070014B_ABST
    Figure CN120070014B_ABST
Patent Text Reader

Abstract

The present application discloses a session recommendation method, device, medium and product based on graph representation learning, which relates to the technical field of session recommendation. The method includes: training an initial architecture for session recommendation of multi-perspective graph representation learning by using a session data set and a total loss function to obtain an architecture for session recommendation of multi-perspective graph representation learning; obtaining a session generated by a target user interacting with a recommendation system; and inputting the session into the architecture for session recommendation of multi-perspective graph representation learning to obtain a recommended item sequence of the target user. By constructing an architecture for session recommendation of multi-perspective graph representation learning, the present application can improve the rationality and accuracy of item recommendation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of session recommendation, and particularly relates to a session recommendation method, device, medium and product based on graph representation learning. Background Art

[0002] With the prosperity and popularization of the Internet, a large number of enterprises tend to provide their products or services in the e-commerce mode, such as Amazon, YouTube, etc. How to quickly screen out the content that users are interested in from these massive amounts of information has become an urgent problem to be solved. As an efficient information filtering technology, the recommendation system can mine the interest points of users based on the historical interaction records of users (such as clicks, views, readings, adding to the shopping cart and purchases, etc.), and recommend the content that users are interested in in an automated manner. Specifically, in the past few decades, the traditional recommendation system methods can be divided into content-based models, collaborative filtering-based models, and hybrid models (that is, the combination of content and collaborative filtering models). However, most of the above methods are dedicated to using the complete historical information of users to capture the long-term and statistical preferences of users for recommendation. However, these methods may not be applicable in actual application scenarios: (1) The preferences of users will change and evolve over time, and it is difficult for these methods to capture the current behavior of users and achieve real-time recommendation. (2) Most importantly, when the historical information of users cannot be obtained, the traditional recommendation system cannot make effective recommendations. Therefore, modeling the short-term preferences and intentions of users based on the current interaction behavior is crucial for achieving dynamic and accurate recommendations.

[0003] To make up for the deficiencies of traditional recommendation methods, session-based recommendation has received increasing attention in recent years. By analyzing the real-time interaction data of users in a short period of time, it can provide relatively accurate personalized recommendations in the absence of user historical data. Existing session-based recommendation methods can be roughly divided into two categories. One is the sequence-based method, which mainly mines users' short-term behaviors and intentions by modeling the sequential information in the session. Such methods usually rely on recurrent neural networks (such as RNN, LSTM, GRU) and attention mechanisms. Although the sequence-based method shows good performance in session-based recommendation, it mainly focuses on sequential information learning and ignores the complex relationships between items. Since the graph structure can flexibly model the non-linear relationships in the session and is introduced into session-based recommendation, the graph-based method can capture the complex relationships between items by constructing the sequence data into a graph, making up for the limitations of the sequence model. At the same time, the graph neural network (GNN) can capture context information faster through the information propagation and aggregation process, showing significant advantages. Due to the special application scenario of session-based recommendation, there is a serious data sparsity problem, and at the same time, the noisy data in the session has an inestimable impact on the recommendation task. Most existing methods infer users' intentions based on the data of a single session. However, it is difficult to capture the real user intentions when the session is short. Although some methods try to use global session information to alleviate sparsity, the consideration of the connection patterns between sessions is relatively simple, and the potential collaborative information cannot be captured. Secondly, most existing methods do not effectively pay attention to the impact of noisy data in the session, so they cannot effectively use node representations for user intention inference, resulting in poor recommendation effects. Summary of the Invention

[0004] The purpose of this application is to provide a session-based recommendation method, device, medium and product based on graph representation learning, which can improve the rationality and accuracy of item recommendation.

[0005] To achieve the above purpose, this application provides the following solutions:

[0006] In the first aspect, this application provides a session-based recommendation method based on graph representation learning, including:

[0007] Construct a multi-perspective graph representation learning session recommendation initial architecture; the multi-perspective graph representation learning session recommendation initial architecture includes: a dual-branch co-occurrence-semantic feature extraction and fusion module, a cross-session graph construction module, a collaborative information extraction module, and a multi-intention prediction and fusion module; the output end of the dual-branch co-occurrence-semantic feature extraction and fusion module and the output end of the cross-session graph construction module are both connected to the input end of the collaborative information extraction module; the output end of the collaborative information extraction module is connected to the input end of the multi-intention prediction and fusion module;

[0008] Train the initial architecture of the multi-view graph representation learning session recommendation using the session dataset and the total loss function to obtain the multi-view graph representation learning session recommendation architecture;

[0009] Obtain the session generated by the interaction between the target user and the recommendation system; the session is an item sequence formed when the target user continuously interacts with multiple items;

[0010] Input the session into the multi-view graph representation learning session recommendation architecture to obtain the recommended item sequence of the target user.

[0011] In a second aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the above-mentioned session recommendation method based on graph representation learning.

[0012] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the above-mentioned session recommendation method based on graph representation learning.

[0013] In a fourth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the above-mentioned session recommendation method based on graph representation learning.

[0014] According to the specific embodiments provided by the present application, the following technical effects are disclosed:

[0015] The present application provides a session recommendation method, device, medium and product based on graph representation learning. In order to use graph neural networks to capture the complex relationships between items, the present application constructs graphs from two perspectives. For the input session sequence, a directed edge is connected between adjacent items to describe the co-occurrence relationship of items (such as a user purchasing mobile phone accessories after purchasing a mobile phone). When the last item is reached, the construction of the co-occurrence relationship item graph is completed. In order to capture the relationships between items from the semantic perspective (such as price, brand, category), the model first initializes the features of the items in the session, and then dynamically constructs a semantic item graph according to the initial features (semantics) of the items. The weight of the edge in the graph is the semantic relevance between items. Finally, the co-occurrence item graph and the semantic item graph are output; for the two item graphs constructed based on the session sequence, on the one hand, the model uses the co-occurrence information extraction network to capture the co-occurrence features of items from the co-occurrence item graph; at the same time, the semantic information extraction network is used to model the semantic associations of items on the semantic item graph, and then the output item features are further integrated to obtain a preliminary representation of the session. Specifically, the model uses a mean aggregation function to obtain order-invariant session features, and introduces order relationships through a recurrent neural network to obtain order-sensitive session features; all sessions are input into the cross-session construction module, which constructs a cross-session graph according to the number of shared items between sessions (the weight of the edge is the number of item overlaps between sessions), and at the same time inputs the order-invariant session features into the co-occurrence collaborative information extraction network to capture the co-occurrence collaborative information between sessions. From the perspective of semantic collaborative information, the model inputs the order-invariant session features and the cross-session graph into the semantic collaborative information extraction network to capture the semantic collaborative information of the session. Finally, the order-sensitive session features obtained by the recurrent neural network are input into the ordered collaborative information extraction network, which samples multiple collaborative sessions according to the session feature similarity and captures the ordered collaborative information between sessions; for the learned item co-occurrence features and item semantic features, the model further generates the local interest and overall interest of the session, combines various collaborative information, and dynamically integrates multiple prediction results through a multi-intent fusion module to obtain the prediction probability distribution of the session. Finally, considering the hard sample mining auxiliary task to further enhance the performance of the session recommendation system. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0017] Figure 1 It is a schematic diagram of explicit collaborative relationships in an embodiment of the present application;

[0018] Figure 2 It is a schematic diagram of potential collaborative relationships in an embodiment of the present application;

[0019] Figure 3 Flowchart of a session recommendation method based on graph representation learning in an embodiment of the present application;

[0020] Figure 4 Schematic diagram of multi - perspective cross - session collaborative information in an embodiment of the present application;

[0021] Figure 5 Overall architecture diagram of the session recommendation method of multi - perspective graph representation learning in an embodiment of the present application;

[0022] Figure 6 Schematic diagram for describing the session recommendation problem in an embodiment of the present application;

[0023] Figure 7 Item co - occurrence graph of a session in an embodiment of the present application and adjacency matrix schematic diagram;

[0024] Figure 8 Flowchart of item co - occurrence information extraction in an embodiment of the present application;

[0025] Figure 9 Flowchart of item semantic information extraction in an embodiment of the present application;

[0026] Figure 10 Flowchart of cross - session co - occurrence collaborative information extraction in an embodiment of the present application;

[0027] Figure 11 Flowchart of cross - session semantic collaborative information extraction in an embodiment of the present application;

[0028] Figure 12 Flowchart of cross - session ordered collaborative information extraction in an embodiment of the present application;

[0029] Figure 13 Flowchart of multi - intent prediction and fusion in an embodiment of the present application. Detailed implementation manners

[0030] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0031] To make the above - mentioned objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0032] Most existing session recommendation methods only focus on the behaviors within the current session, ignoring the potential collaborative relationships that may exist between sessions, and at the same time facing the dilemma of data sparsity. To make up for this deficiency, this application aims to mine the collaborative information between sessions from multiple perspectives, provide multiple potential contexts, and effectively alleviate the data sparsity problem. Specifically, it is divided into three perspectives:

[0033] Based on co-occurrence relationships: This perspective captures cross-session collaborative information based on the co-occurrence relationships of user behaviors. If there are the same user behaviors between sessions (such as Figure 1 the same product is clicked in different sessions), it can be considered that there is a collaborative relationship between them. By connecting these sessions through the common behaviors, session information can be propagated between sessions. Through the information propagation mechanism, sessions with the same behavior can share each other's context information, thereby alleviating the data sparsity problem.

[0034] Based on semantic relationships: In addition to introducing collaborative information through the same behaviors between sessions, collaborative information can also be considered through the similarity of behaviors. Specifically, although there are no identical click behaviors in two sessions, different users may be interested in similar categories of products, such as the same brand, similar price ranges, or the same type of products (such as Figure 2 , there is a similar association with electronic products in two sessions). From this perspective, by extracting behavior attributes (such as product category, brand, price, etc.) and measuring the similarity of sessions in the feature space, collaborative information that has no overlap in behavior but has potential semantic connections can be captured.

[0035] Based on the orderliness of behaviors: In some scenarios, user behaviors usually show a certain temporal order, reflecting the dynamic pattern of user interests (such as Figure 1 , the user may pay attention to shoes after buying pants or a top). By considering the sequential relationship of behaviors in a session and transmitting this sequential information between different sessions, the sequential dependence relationship between sessions can be captured. Especially in scenarios where user behaviors have obvious temporal dependencies, it can more accurately reflect behavior changes and interest trends.

[0036] In summary, by considering this session collaborative information, the data sparsity problem of session recommendation can be effectively alleviated, and the performance of the recommendation system can be improved.

[0037] When the recommendation system infers the conversation intention, it is necessary to comprehensively consider the various intentions that may exist in the conversation. Such as multiple collaborative information, as well as the long-term and short-term intentions of the current conversation. It is worth noting that the conversation representations from different perspectives represent different intentions. Therefore, this application introduces an adaptive multi-intention fusion module to effectively fuse the conversation representations from different perspectives to ensure the reasonable reasoning of user intentions. In short, each conversation representation needs to be normalized first, which helps to unify the feature representations from different perspectives and avoid information distortion caused by inconsistent feature scales. On this basis, the module predicts the conversation representation of each perspective separately, and then weightedly fuses the prediction results. Through this adaptive fusion module, the model can integrate multiple conversation representations, thereby improving the accuracy of the conversation recommendation system. When dealing with long-tail users (demand tends to be niche products or unpopular content) and conversations with complex content, the recommendation accuracy of the conversation recommendation system is often low. In order to meet this challenge, this application introduces an auxiliary task that enables the model to pay more attention to such conversation samples. This mechanism increases the optimization efforts for sessions with large prediction errors, prompting the model to better learn the characteristic expressions of these sessions, thereby improving the modeling capabilities for complex and rare samples, avoiding overfitting of the model to common samples, and further enhancing the overall performance and robustness of the recommendation system. The Chinese meanings of the symbols involved in this application are shown in Table 1.

[0038] Table 1 Main symbol table

[0039]

[0040] In an exemplary embodiment, Figure 3 As shown, a conversation recommendation method based on graph representation learning is provided, including:

[0041] Step 301: Build an initial architecture for multi-view graph representation learning session recommendation.

[0042] Step 302: Using the conversation data set and the total loss function, the multi-view graph representation learning conversation recommendation initial architecture is trained to obtain the multi-view graph representation learning conversation recommendation architecture.

[0043] Step 303: Acquire a session generated by the target user interacting with the recommendation system; a session is a sequence of items formed when the target user continuously interacts with multiple items.

[0044] Step 304: Input the conversation into the multi-view graph representation learning conversation recommendation framework to obtain a sequence of recommended items for the target user.

[0045] The initial architecture of multi-perspective graph representation learning for session recommendation includes: a dual-branch co-occurrence-semantic feature extraction and fusion module, a cross-session graph construction module, a collaborative information extraction module, and a multi-intent prediction and fusion module. The output ends of both the dual-branch co-occurrence-semantic feature extraction and fusion module and the cross-session graph construction module are connected to the input end of the collaborative information extraction module. The output end of the collaborative information extraction module is connected to the input end of the multi-intent prediction and fusion module.

[0046] The dual-branch co-occurrence-semantic feature extraction and fusion module includes: a co-occurrence branch, a semantic branch, and a preliminary representation integration module.

[0047] The output ends of both the co-occurrence branch and the semantic branch are connected to the input end of the preliminary representation integration module. The output end of the preliminary representation integration module is connected to the input end of the collaborative information extraction module.

[0048] The co-occurrence branch includes a co-occurrence item graph construction unit and a co-occurrence information extraction unit connected in sequence. The co-occurrence information extraction unit uses a gated graph neural network. The output end of the co-occurrence information extraction unit is connected to the input end of the preliminary representation integration module.

[0049] The semantic branch includes a semantic item graph construction unit and a semantic information extraction unit connected in sequence. The semantic information extraction unit uses a multi-level graph convolutional network. The output end of the semantic information extraction unit is connected to the input end of the preliminary representation integration module.

[0050] The preliminary representation integration module includes: a sequential-unchanged session feature extraction module and a sequential-sensitive session feature extraction module connected in sequence; the output ends of the sequential-sensitive session feature extraction module and the sequential-unchanged session feature extraction module are connected to the input end of the collaborative information extraction module; the sequential-unchanged session feature extraction module uses a mean aggregation function; the sequential-sensitive session feature extraction module uses a recurrent neural network.

[0051] The co-occurrence item graph construction unit is used to generate a co-occurrence item graph based on the input session.

[0052] The co-occurrence information extraction unit is used to extract co-occurrence features based on the co-occurrence item graph.

[0053] The semantic item graph construction unit is used to generate a semantic item graph based on the input session.

[0054] The semantic information extraction unit is used to extract semantic features based on the semantic item graph.

[0055] The sequential-unchanged session feature extraction module is used to generate a first sequential-unchanged session feature and a second sequential-unchanged session feature based on the co-occurrence features and semantic features.

[0056] The sequential-sensitive session feature extraction module is used to generate a sequential-sensitive session feature based on the co-occurrence features and semantic features.

[0057] The co-occurrence item graph is a directed item graph, where represents the set of items in the co-occurrence item graph ; represents a session , represents the first item in session S; represents the session the second item in; represents the session the th item in; represents the session the Lth item in; represents the ith item in session S; represents the co-occurrence item graph the set of all edges in; represents the session the th item in; item and item in the co-occurrence item graph the corresponding edge weight is: ; The adjacency matrix of the co-occurrence item graph includes an out-degree neighbor information matrix and an in-degree neighbor information matrix ;

[0058] The element in the out-degree neighbor information matrix is:

[0059] ;

[0060] The element in the in-degree neighbor information matrix is:

[0061] ;

[0062] The semantic item graph is an undirected item graph, where represents all items in the session, represents the edge between item pairs ;

[0063] The adjacency matrix of the semantic item graph the element represents item and item The weight of the corresponding edge in the semantic item graph ;

[0064] ;

[0065] Among them, is the initial weight of the corresponding edge between item and item in the semantic item graph ;

[0066] ;

[0067] Among them, is the initial vector obtained by encoding item ; is the transpose of the initial feature vector of item ; is the initial feature vector of item ;

[0068] The cross - session graph construction module is used to generate a cross - session graph based on the session sequence ;

[0069] Among them, is the point corresponding to each session in the session sequence; represents the i - th session in the session set constituted by the session sequence; is the edge set in the cross - session graph; represents the edge formed by session and session where there is an interaction item; the adjacency matrix of the cross - session graph The element in it represents the weight of edge ; ; represents the item set of session ; represents the item set of session ;

[0070] The collaborative information extraction module includes a co - occurrence collaborative information extraction branch, a semantic collaborative information extraction branch, and an ordered collaborative information extraction branch connected in parallel;

[0071] The co - occurrence collaborative information extraction branch is used to extract co - occurrence collaborative information based on co - occurrence features, first - order invariant session features, and the adjacency matrix of the cross - session graph;

[0072] The semantic collaborative information extraction branch is used to extract semantic collaborative information based on semantic features, second-order invariant session features, and the adjacency matrix of the cross-session graph ;

[0073] The ordered collaborative information extraction branch is used to extract ordered collaborative information based on co-occurrence features and order-sensitive session features.

[0074] The multi-intent prediction fusion module is used to determine the local interest feature, the first overall interest feature, and the second overall interest feature of the session, and decode the local interest feature, the first overall interest feature, the second overall interest feature, co-occurrence collaborative information, semantic collaborative information, and ordered collaborative information respectively, and multiple probability predictions can be obtained. The weighted max pooling and average pooling are used to fuse the multiple probability predictions to obtain the prediction result.

[0075] The total loss function is:

[0076] ;

[0077] ;

[0078] ;

[0079] ;

[0080] where is the total loss function, is the cross-entropy loss function is the dynamic loss value function; is the parameter for L2 regularization of all trainable parameters of the model; is the prediction result of the i-th item in the candidate item set; is the true result of the i-th item in the candidate item set; N is the total number of item categories is the bias value control factor; is the final prediction score of the item corresponding to the label.

[0081] Most of the existing session recommendation methods only focus on the user behavior within the session, while ignoring the collaborative information existing between sessions. Due to the serious problem of session data sparsity, especially when the session is short, it is difficult to use session data for effective user intention inference.

[0082] Table 2 Statistical Information Table of Public Datasets

[0083]

[0084] Table 2 shows the basic statistics of the session recommendation popularity datasets in five real scenarios. It can be seen that although the average session length of Last.fm is relatively long (9.16), the number of short sessions is still 1,136,909, indicating that the sparsity problem still exists; while short sessions dominate in the remaining four datasets. These data show that short sessions dominate in the session recommendation scenario, and the sparsity of the data poses a great challenge to capturing user preferences and modeling complex relationships.

[0085] This application considers using global information to enhance the performance of session recommendation, and introduces collaborative information through multiple cross-session perspectives to alleviate the data sparsity problem. For example, Figure 4 , according to the co-occurrence relationship, sessions with the same items (trousers, shoes) are screened to introduce potential contexts for purchasing tops for the current session; or according to the semantic (category) representation of items, sessions with the intention of purchasing clothing are filtered out to provide collaborative information for purchasing tops for the current session.

[0086] Data noise: Session recommendation is mainly applied in the instant recommendation scenario. The content that users are interested in in the future may be highly relevant to their recent behaviors. Therefore, existing methods tend to take the last item as the session intention. In this case, the model is easily affected by noisy data, resulting in incorrect inference of user intentions. For example, Figure 4 , if the model mainly focuses on the recent behaviors of the session, then the racket will be predicted as the session intention. However, in fact, the similarity between the racket and the labeled item top is very low. In fact, the racket has nothing to do with the overall session intention (clothing items). In order to more accurately understand and predict users' needs, this application considers both the overall and local representations of the session, enabling the model to capture both instant intentions and overall intentions, and weakening the impact of noisy items. For example, Figure 4 , specifically, by modeling the overall intention of the session, a more general session intention expression (clothing) can be obtained. At the same time, relevant collaborative sessions can also be screened to provide contexts related to the overall intention (clothing). Finally, at the intention inference stage, by comprehensively considering various session representations, the impact of noisy data on intention inference can be weakened.

[0087] The overall architecture of the model is as shown in Figure 5 , and the framework mainly includes the following four parts:

[0088] The description of the session recommendation problem is as shown in Figure 6 , the item set is defined as ; where N represents the total number of items. The session set is represented as where M represents the total number of sessions in the dataset. A given session S is represented as , represents the Interaction terms, where \(L\) is the length of the session. Session-based recommendation aims to predict the user's behavior at the next time step (i.e., ). Specifically, given a session, the goal of session recommendation is to recommend the top \(\mu\) items in the item set that are most likely to be clicked or purchased by the user at the next timestamp, where .

[0089] Step 1: Local item graph construction.

[0090] (1) Co-occurrence item graph construction.

[0091] Session recommendation needs to obtain the user's intention based on the item information in the session. To capture the complex relationships between items and infer the user's intention, this application constructs an item graph based on co-occurrence relationships for each session and uses a graph neural network to learn item representations. Specifically, according to the order of item appearance in the session sequence, a co-occurrence item graph is constructed, which means that if two items appear successively in the session, these two items will be connected in the item graph. The direction of the connection represents their order relationship, and the strength of the connection can reflect the weight of their co-occurrence. For example, for the session , a directed item graph is defined, where represents the set of items in the directed item graph , and represents the set of all edges in the graph. The adjacency matrix in the directed item graph is represented by , which is used to generate two types of adjacency relationship matrices, namely the in-degree neighbor matrix and the out-degree neighbor matrix . represents the value of the \(i\)-th row and \(j\)-th column of , that is, the weight of the corresponding edge between item and item in the directed item graph . is represented as follows:

[0092] (1)

[0093] To focus on the directionality of the edges in the directed item graph , the adjacency matrix is then decomposed into two parts: the in-degree neighbor matrix and the out-degree neighbor matrix , and normalization operations are performed on these two matrices respectively: Since the graph is a directed graph where the neighbors of a node can be divided into in-degree neighbors (nodes pointing to this node) and out-degree neighbors (nodes pointed to by this node), then the in-degree neighbor matrix is introduced Out-degree neighbor matrix Consider these two adjacency relations separately. Both are generated by the adjacency matrix of the graph . The out-degree neighbor matrix . The edge weight corresponding to node i and node j in the out-degree neighbor matrix is represented as follows:

[0094] (2)

[0095] Similarly represents the in-degree neighbor matrix and the edge weight between node i and node j in it, which is specifically represented as .

[0096] Given a session , adding a directed edge to each pair of adjacent items can obtain the edge set , the item set . A co-occurrence item graph can be constructed based on the edge set and the item set. Then, a directed graph is generated based on the item set and the edge set. The nodes in the graph represent items, and the edges represent the relationships between adjacent items in the session. The directed item graph of the session and its adjacency matrix are as shown in Figure 7 .

[0097] (2)Semantic item graph construction.

[0098] To capture the relationships between items from the semantic perspective of items, the model dynamically constructs a semantic item graph , represents all items in the session, represents the edge between the item pair . Specifically, the weight of the item pair is defined as follows, where , representing the initial feature vector of the item pair :

[0099]

[0100] Then use the function to normalize the weight to obtain the adjacency matrix of the semantic item graph . Finally, the weight of the edge in the graph is represented as follows:

[0101] (4)

[0102] Step 2: Item representation learning. ​

[0103] For the input session sequence, each item is first mapped into a high-dimensional vector, and the item features are normalized to mitigate the impact of item popularity bias (tendency to recommend popular, high-frequency, and trendy items while ignoring unpopular items). For a session , after encoding, an item feature matrix is obtained, where , L represents the maximum session length, and d is the size of the feature dimension. Finally, the initial feature of the session is , which is expressed as follows:

[0104] (5)

[0105] where represents calculating the vector norm. Then, a graph neural network (GNN) is used to learn the feature representation of items on two types of item graphs. The process of extracting item co-occurrence information is as Figure 8 shown.

[0106] (1)Co-occurrence information extraction.

[0107] Specifically, the model uses a gated graph neural network (co-occurrence information extraction network) to process the in-degree neighbor matrix and the out-degree neighbor matrix obtained in step 1 to facilitate the message passing process between items. The context information received by item on the directed item graph is defined as , where is a learnable matrix, and the symbol [;] represents the concatenation operation, is the initial feature representation of the session . (Respectively, the i-th row of the out-degree neighbor matrix and the in-degree neighbor matrix )represents the neighbor information of the out-edge and in-edge of item in the directed item graph . The feature representation of item is updated using the context information as follows:

[0108] (6)

[0109] (7)

[0110] (8)

[0111] (9)

[0112] where are learnable parameters, ⊙ represents element-wise multiplication, is a function, representing the candidate features of node in the τ-th layer of the gated graph neural network, used to update the feature representation of node in the current network layer while represents using context information to update the features of node . The control of the network is regulated by the update gate and the reset gate , which are crucial throughout the network and control the processing and filtering of information. Subsequently, let , and finally the output representation of the co-occurrence information extraction network is , which represents the co-occurrence features of the project.

[0113] The process of project semantic information extraction is as shown in Figure 9 .

[0114] (2) Semantic information extraction.

[0115] From a semantic perspective, the model uses a multi-level graph convolutional network (semantic information extraction network) to process the adjacency matrix of the semantic item graph in step 1 to capture the semantic relationships between items. Specifically, the propagation process of this graph network is defined as follows:

[0116] (10)

[0117] (11)

[0118] (12)

[0119] where is the initial feature of node , represents the semantic information obtained by node from adjacent items, are learnable parameters used to generate the contribution weights of the item features output by each layer of the network and the initial item features, is a function, represents the i-th row of the adjacency matrix of the -th layer. is composed of Dynamically generated. For the detailed process, please refer to the construction process of the semantic project graph. Finally, the output of the semantic information extraction network is represented as , which represents the semantic features of the project.

[0120] Step 3: Cross-session collaborative information extraction.

[0121] For the co-occurrence features obtained in Step 2 and semantic features , first, feature fusion is performed to obtain a preliminary representation of the session. On the one hand, based on the mean aggregation function, the order-invariant intent of the session can be extracted. Considering the ordered relationship of the projects, the model further adopts a gated recurrent neural network ) to extract the order-sensitive session intent, and apply these two to the following collaborative information extraction network. The cross-session co-occurrence collaborative information extraction process is as shown in Figure 10 .

[0122] (1) Co-occurrence collaborative information extraction.

[0123] To obtain the explicit collaborative information between sessions (sessions sharing the same project can provide session information for each other), the model constructs a cross-session graph: each session is regarded as a supernode, and the edge weight between any two nodes in the graph is quantified by the overlap degree of the corresponding session item types. Specifically, the global co-occurrence session graph is defined as , where , the adjacency matrix of graph is defined as , M represents the total number of sessions. The edge weight corresponding to session and in graph is (for the i-th row and j-th column of ) and is calculated as follows:

[0124] (13)

[0125] where represents the set of items in session . This collaborative information extraction network uses the order-invariant features of the session as the representation of the session node. For session S, the initial representation is defined as:

[0126] (14)

[0127] where, L represents the session length, is the item co-occurrence feature. The identity matrix is used to enhance to obtain . And the diagonal matrix is obtained, where each diagonal element Represents a session And the weight sum of adjacent sessions. Using To represent the initial feature matrix of the session, and , the convolution operation is defined as follows:

[0128] (15)

[0129] Generally speaking, each session can obtain collaborative information from sessions sharing items. The data sparsity problem can be alleviated by capturing cross-session information. At the same time, in order to avoid the over-smoothing phenomenon in the multi-layer convolutional network, the residual connection method is further considered, so the finally captured co-occurrence collaborative information Is represented as follows:

[0130] (16)

[0131] (2) Semantic collaborative information extraction.

[0132] Since there may be no direct shared items between sessions, considering that there may be similar behaviors or preferences between sessions, in order to capture this potential collaborative information, this application introduces a semantic collaborative information network, aiming to capture cross-session potential collaborative information and further improve the session representation quality. Similarly, the sequential invariant features of the session will be used as the initial representation of the session and input into the semantic collaborative information network, Representing the item semantic information:

[0133] (17)

[0134] Specifically, this network uses a session-aware attention mechanism to capture the correlation between sessions. Each session is linearly combined with the attention scores of other sessions, and the session The cross-session information that can be obtained is as follows:

[0135] (18)

[0136] Among them, Represents the neighbor session of session , Then represents the importance weight of session And its neighbor session, and the weight is specifically represented as follows:

[0137] (19)

[0138] Among them Is a trainable parameter, Represents session On the global session graph The weight of the corresponding edge, represents the element-wise product. Then, the function is used to normalize all the neighbor weights connected to the session:

[0139] (20)

[0140] To capture deeper semantic collaborative information, a gating mechanism is further used to enhance the session-aware attention network on the basis of the above:

[0141] (21)

[0142] (22)

[0143] (23)

[0144] where is the session feature with the initial order unchanged, is the sigmoid activation function, represents the contribution weight of the output of the current layer network to the initial session feature , and the final output is expressed as , representing the captured cross-session semantic collaborative information. The cross-session semantic collaborative information extraction process is as shown in Figure 11 , where the superscript τ represents the network output quantity.

[0145] (3) Ordered collaborative information extraction.

[0146] Considering that user behavior in the real scenario shows a certain order (for example, after purchasing a mobile phone, the user may purchase mobile phone accessories such as chargers and phone cases), this application further considers passing the order-sensitive representation of the session in the global perspective to capture the ordered dependencies between sessions. At this time, the order-sensitive feature of the session is used as the input of the network. Specifically, the initial session representation obtained by using the item co-occurrence information is:

[0147] (24)

[0148] The item features are processed by a gated recurrent neural network (GRU), and then its average representation is obtained as the network input. Then, by calculating the similarity score between session representations, several sessions with higher similarity are selected for collaboration. The calculation formula of the session similarity score is as follows:

[0149] (25)

[0150] where , M represents the total number of sessions, are the features of the current session S. The sessions are sorted according to the Sim value, and for session S, F samples with higher Sim values are selected as collaborative sessions, and cross-session ordered collaborative information is obtained as follows:

[0151] (26)

[0152] (27)

[0153] Among them, is the temperature parameter for adjusting the contribution intensity, represents the contribution weight of the collaborative session , represents the captured cross-session ordered collaborative information. The cross-session ordered collaborative information extraction process is as Figure 12 shown.

[0154] Step 4: Multi-perspective session representation and fusion.

[0155] The model has obtained cross-session co-occurrence collaborative information , semantic collaborative information and ordered collaborative information . However, the intention of the current session has not been inferred, so it is necessary to consider the two types of item features obtained in step two to infer the user interest.

[0156] Since most session recommendation systems are applied in the instant recommendation scenario, the user's future intention has a strong correlation with the recent behavior. Therefore, the model considers using the item representation at the last time step in the item co-occurrence feature and the item semantic feature as the local interest of the session. MLP(*) represents a multi-layer perceptron, and the specific representation is as follows:

[0157] (28)

[0158] Among them represents the local interest, and [;] represents feature concatenation. Select the k most recent user behaviors in session S as the overall intention of the session, which can capture a more immediate overall intention and avoid overemphasizing early interfering behaviors. To alleviate the possible noisy data in the recent behaviors, the model further considers the overall interest of the session from the item semantic perspective (item semantic feature):

[0159] (29)

[0160] Among them Represents the overall intention of the conversation. By considering multiple recent actions simultaneously to obtain the overall interest representation of the conversation, it can alleviate the influence of noisy data and introduce richer user intentions.

[0161] Similarly, from the perspective of item co-occurrence (item co-occurrence representation), a soft attention mechanism is used to obtain the overall interest. Specifically, a learnable position matrix is used Then, the reversed position features are combined with the learned item co-occurrence features Combined, the feature of the i-th item is represented as:

[0162] (30)

[0163] Where are learnable parameters, and tanh(*) is the activation function. The overall user intention is represented as:

[0164] (31)

[0165] (32)

[0166] Here and are attention parameters used to learn the weights of each item in the conversation . is the order-invariant representation of session S after introducing position information:

[0167] (33)

[0168] In summary, the overall interest captured by the attention mechanism is . At this point, the model has obtained local, overall interests, and cross-session collaborative information.

[0169] These session representations are in different feature spaces. Directly fusing them during the prediction stage will instead lead to confusion in user intentions. Therefore, this application introduces an adaptive multi-intention fusion module to effectively fuse these session representations from different perspectives and ensure reasonable inference of user intentions. Specifically, for the candidate item set feature matrix represents the number of items. First, perform L2 normalization on it to obtain . At the same time, perform the same normalization operation on the obtained multiple session representations, which can effectively alleviate the popularity bias. Then, for each representation, perform prediction respectively to obtain the corresponding prediction distribution (using the session representation to predict the probability score of each item in the candidate item set when being recommended):

[0170] (34)

[0171] Here represents the predicted distribution decoded from the session local interest. Similarly, for the two session global interests , the collaborative information from the three perspectives is decoded separately to obtain multiple probability predictions. To effectively aggregate the multiple prediction results, the module uses weighted max pooling and average pooling to fuse them. Among them, max pooling retains the individual intentions of multiple session representations (highlighting the most significant features from multiple session representations and retaining the uniqueness of different perspectives), while average pooling captures the common intentions (paying attention to the common points between different perspectives and reflecting the general trends in session representations). The final prediction result is expressed as follows:

[0172] (35)

[0173] Among them, is the activation function is a learnable fusion parameter, and represent max and average pooling respectively. The final prediction result represents the probability that each item in the candidate item set will be selected by the user at the next time step of the current session. The larger the probability value of an item, the more it conforms to the user's current interest tendency. By sorting the probability scores of the candidate items and selecting multiple items with the highest predicted probabilities to generate a recommendation list, session recommendation is achieved. The multi-intention prediction and fusion process is as Figure 13 shown. The final model converts the true label label of the session into a one-hot vector , and uses the cross-entropy loss function as an index to optimize the model parameters:

[0174] (36)

[0175] Considering the differences between session samples, the prediction difficulty of different samples may vary greatly. For session samples that are difficult to predict, the prediction scores of the model may deviate significantly from the true labels. To enable the model to handle these difficult-to-classify session samples more effectively, we introduce an additional dynamic loss value in the loss function as an auxiliary task for the model (hard sample mining) to increase the model's attention to difficult-to-classify samples, defined as follows:

[0176] (37)

[0177] (38)

[0178] Among them, Represents the final predicted score for the item corresponding to the label, Represents the deviation between the predicted value and the true value of the sample, that is, the prediction difficulty of this sample in the current round. The model controls the magnitude of this deviation value through . The model unifies the main recommendation task and the auxiliary task to enhance the performance of session recommendation. The final loss can be expressed as:

[0179] (39)

[0180] Among them, Represents the parameter for L2 regularization of all trainable parameters of the model. Among them directly reflects the gap between the predicted value of the model and the true label, and the goal of is to minimize this gap as much as possible, so that the recommended results of the model are closer to the actual needs of users. The optimization process of this loss is the core of model training. By dynamically assigning higher loss values to samples with greater prediction difficulty (i.e., samples with a large deviation between the predicted result and the true result), the model is forced to pay more attention to these samples, thereby improving the overall model performance.

[0181] Table 3 Comparison effect table with existing baseline models

[0182]

[0183] Table 3 shows the recommendation effects of the model on four public datasets, including: Tmall and Dig-inetica datasets (shopping and transaction data on e-commerce platforms), RetailRocket (click and purchase dataset of users on retail platforms), and Last.fm, a classic benchmark dataset for music recommendation. The evaluation metrics used are HR@20 and MRR@20. HR@20 (hit rate) measures whether the labeled item appears in the top 20 items of the recommendation list, without paying attention to its specific ranking. MRR@20 (mean reciprocal rank) calculates the reciprocal average of the rankings of the labeled item in the top 20 items of the recommendation list, and is used to evaluate the recommendation accuracy and ranking quality of the model. The higher the HR metric, the higher the probability that the recommended list will contain content of interest to users. The higher the MRR metric, the higher the ranking of the content of most interest in the recommendation list. It can be seen from the last column of the table that the experimental results of this application have a significant improvement compared to the baseline model. Further indicating the effectiveness of the model proposed in this application in the session recommendation task.

[0184] In an exemplary embodiment, a computer device is provided, which may be a server or a terminal. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a session recommendation method based on graph representation learning.

[0185] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0186] In an exemplary embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0187] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0188] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0189] The databases involved in the various embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the various embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.

[0190] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0191] In this article, specific examples are used to elaborate on the principles and implementation manners of this application. The descriptions of the above embodiments are only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A conversation recommendation method based on graph representation learning, characterized in that Including: Construct an initial architecture for session recommendation with multi-view graph representation learning; The initial architecture for session recommendation with multi-view graph representation learning includes: a dual-branch co-occurrence-semantic feature extraction and fusion module, a cross-session graph construction module, a collaborative information extraction module, and a multi-intention prediction and fusion module; the output end of the dual-branch co-occurrence-semantic feature extraction and fusion module and the output end of the cross-session graph construction module are both connected to the input end of the collaborative information extraction module; the output end of the collaborative information extraction module is connected to the input end of the multi-intention prediction and fusion module; The dual-branch co-occurrence-semantic feature extraction and fusion module includes: a co-occurrence branch, a semantic branch, and a preliminary representation integration module; The output end of the co-occurrence branch and the output end of the semantic branch are both connected to the input end of the preliminary representation integration module; the output end of the preliminary representation integration module is connected to the input end of the collaborative information extraction module; The co-occurrence branch includes a co-occurrence item graph construction unit and a co-occurrence information extraction unit connected in sequence; the co-occurrence information extraction unit uses a gated graph neural network; the output end of the co-occurrence information extraction unit is connected to the input end of the preliminary representation integration module; The semantic branch includes a semantic item graph construction unit and a semantic information extraction unit connected in sequence; the semantic information extraction unit uses a multi-level graph convolutional network; the output end of the semantic information extraction unit is connected to the input end of the preliminary representation integration module; The preliminary representation integration module includes: a sequential-invariant session feature extraction module and a sequential-sensitive session feature extraction module connected in sequence; the output ends of the sequential-sensitive session feature extraction module and the sequential-invariant session feature extraction module are connected to the input end of the collaborative information extraction module; the sequential-invariant session feature extraction module uses a mean aggregation function; the sequential-sensitive session feature extraction module uses a recurrent neural network; Use the session dataset and the total loss function to train the initial architecture for session recommendation with multi-view graph representation learning to obtain an architecture for session recommendation with multi-view graph representation learning; Obtain the session generated by the interaction between the target user and the recommendation system; the session is an item sequence formed when the target user continuously interacts with multiple items; Input the session into the architecture for session recommendation with multi-view graph representation learning to obtain the recommended item sequence of the target user.

2. The session recommendation method based on graph representation learning according to claim 1, wherein The co-occurrence item graph construction unit is used to generate a co-occurrence item graph based on the input session; The co-occurrence information extraction unit is used to extract co-occurrence features based on the co-occurrence item graph; The semantic item graph construction unit is used to generate a semantic item graph based on the input session; The semantic information extraction unit is used to extract semantic features based on the semantic item graph; The sequential-invariant session feature extraction module is used to generate a first sequential-invariant session feature and a second sequential-invariant session feature based on the co-occurrence features and semantic features; The sequential-sensitive session feature extraction module is used to generate a sequential-sensitive session feature based on the co-occurrence features.

3. The session recommendation method based on graph representation learning according to claim 2, wherein The co-occurrence item graph is a directed item graph, where represents the set of items in the co-occurrence item graph ; represents a session , represents the first item in session S; represents a session the second item in; represents a session in the th item; represents a session the Lth item in; represents the ith item in session S; represents the co-occurrence item graph the set of all edges in; represents a session in the th item; item and item in the co-occurrence item graph the corresponding edge weight is: ; The adjacency matrix of the co-occurrence item graph includes an out-degree neighbor information matrix and an in-degree neighbor information matrix ; The out-degree neighbor information matrix The elements in are as follows: ; The in-degree neighbor information matrix The elements in are as follows: ; The semantic item graph is an undirected item graph, where represents all items in the conversation, represents item pairs and the edges between them; The semantic item graph adjacency matrix of the element in represents the item and the item in the semantic item graph the corresponding edge in weight; ; Among them, is for project and project in the semantic project graph the initial weight of the corresponding edge ; ; Among them, is the initial vector obtained after encoding the project ; is the transpose of the initial feature vector of the project ; is the initial feature vector of the project .

4. The session recommendation method based on graph representation learning according to claim 3, characterized in that The cross-session graph construction module is used to generate a cross-session graph based on the session sequences ; Among them, is the point corresponding to each session in the session sequence; represents the session set composed of the session sequence the i-th session in; is the edge set in the cross-session graph; represents the session with interaction items and session constitute the edge; cross-session graph adjacency matrix of the element in represents the edge weight; ; represents the item set of session ; represents the item set of session ; 5. The session recommendation method based on graph representation learning according to claim 4, wherein The co-occurrence information extraction module includes a co-occurrence co-information extraction branch, a semantic co-information extraction branch, and an ordered co-information extraction branch connected in parallel; The co-occurrence collaborative information extraction branch is used to extract co-occurrence collaborative information based on co-occurrence features, first-order invariant session features, and the adjacency matrix of the cross-session graph ; The semantic collaboration information extraction branch is used to extract semantic collaboration information based on semantic features, second-order invariant conversation features, and the adjacency matrix of the cross-conversation graph ; The ordered co-information extraction branch is used to extract ordered co-information based on co-occurrence features and order-sensitive session features.

6. The session recommendation method based on graph representation learning according to claim 5, wherein The multi-intention prediction fusion module is used to determine the local interest feature, the first overall interest feature, and the second overall interest feature of the session, decode the local interest feature, the first overall interest feature, the second overall interest feature, the co-occurrence co-information, the semantic co-information, and the ordered co-information respectively, and multiple probability predictions can be obtained. The multiple probability predictions are fused using weighted max pooling and average pooling to obtain the prediction result.

7. The session recommendation method based on graph representation learning according to claim 6, wherein The total loss function is: ; ; ; ; Among them, is the total loss function, is the cross-entropy loss function; is the dynamic loss value function; is for all trainable parameters of the model parameters for L2 regularization; is the prediction result of the i-th item in the candidate item set; is the true result of the i-th item in the candidate item set; N is the total number of item categories; is the bias value control factor; is the final prediction score of the item corresponding to the label.

8. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the session recommendation method based on graph representation learning according to any one of claims 1-7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the session recommendation method based on graph representation learning according to any one of claims 1-7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the session recommendation method based on graph representation learning according to any one of claims 1-7.

Citation Information

Patent Citations

  • Neighbor enhancement-based graph neural network session recommendation method, system and equipment

    CN115130001A

  • Motor imagery electroencephalogram signal classification method based on graph neural network

    CN117216631A