Cross-session Recommendation Method and System for a Multi-interest Integrated Multimodal Knowledge Graph
By constructing multiple historical interest graphs and current interest graphs, using the multimodal knowledge graph attention mechanism and similar conversation retrieval networks, cross-session recommendations are generated, which solves the problem of ignoring the continuity and relevance of users in the prior art, and improves the accuracy and personalization of recommendations.
Patent Information
- Application Number
- CN202410897065.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-05
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-07-05
AI Technical Summary
Existing session recommendation methods ignore the continuity and correlation between users among multiple sessions, resulting in recommendations based on a single interest that cannot fully meet the diverse needs of users, and recurrent neural networks ignore the randomness and complexity of user behavior, resulting in poor recommendation results.
By obtaining multiple historical session sequences and current session sequences of the current user, multiple historical interest graphs and current interest graphs are constructed, and aggregation embedded representations are generated using the multimodal knowledge graph attention mechanism, and similar historical sessions are retrieved through the similar session search network to generate cross-session recommendations.
Effectively capture multiple user interests and obtain the context information of project entities through multimodal knowledge graphs, improving the effect of cross-session recommendations and enhancing the accuracy and personalization of recommendations.
Smart Images

Figure CN118861238B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and particularly to a cross-session recommendation method and system for a multi-interest fusion multi-modal knowledge graph. Background Art
[0002] In the era of information explosion, session recommendation methods play an important role in alleviating the problem of information overload. Session recommendation methods analyze the historical interaction data of users and recommend personalized products and services to users. Session recommendation methods are currently widely used in fields such as social media, movies and music, e-commerce, news, marketing, and advertising. For example, in the field of short videos, session recommendation methods can judge the preferences of users by analyzing the list of videos clicked by users in history, predict the types of videos that users will click in the future, and thus recommend videos preferred by users in a targeted manner to improve the user experience.
[0003] Existing recommendations focus on the interests of users in the current session and ignore the continuity and relevance of users among multiple sessions. However, users' interests are often diverse, and recommendations based on a single interest often cannot fully meet the needs of users. Currently, the next item based on session recommendation mainly uses a recurrent neural network to search for dependencies between items as key features. However, the recurrent neural network ignores the randomness and complexity of user behavior, regards the item relationship as a rigid dependency, and the effect of session recommendation is poor.
[0004] Therefore, there is a need to provide a cross-session recommendation method and system for a multi-interest fusion multi-modal knowledge graph to improve the effect of cross-session recommendation. Summary of the Invention
[0005] The present invention provides a cross-session recommendation method for a multi-interest fusion multi-modal knowledge graph, including: obtaining a plurality of historical session sequences and a current session sequence of a current user; constructing a plurality of historical interest graphs of the current user based on the plurality of historical session sequences of the current user; constructing a plurality of current interest graphs of the current user based on the current session sequence; generating an aggregated embedding representation based on the plurality of historical interest graphs and the plurality of current interest graphs of the current user through a multi-modal knowledge graph attention mechanism; retrieving similar historical sessions of other users corresponding to the current session of the current user through a similar session retrieval network; and generating a cross-session recommendation based on the aggregated embedding representation and the similar historical sessions of other users.
[0006] Further, based on multiple historical session sequences of the current user, multiple historical interest graphs of the current user are constructed, including: converting all items included in the multiple historical session sequences into multi-dimensional embedding representations, generating a historical embedding matrix corresponding to the items, and determining a correlation score between the items and interests based on the historical embedding matrix; constructing multiple historical interest graphs of the current user based on the correlation scores of the correlation scores between the items and interests, where one historical interest graph corresponds to one interest.
[0007] Further, based on the current session sequence, multiple current interest graphs of the current user are constructed, including: converting all items in the current session sequence into multi-dimensional embedding representations, generating a current embedding matrix corresponding to the items; for each interest, calculating a relationship score between two items of the interest based on the current embedding matrix, generating an attention matrix corresponding to the interest as an adjacency matrix based on the relationship score, and generating a current interest graph corresponding to the interest based on the adjacency matrix.
[0008] Further, through a multi-modal knowledge graph attention mechanism, an aggregated embedding representation is generated based on multiple historical interest graphs and multiple current interest graphs of the current user, including: generating an interest historical representation based on multiple historical interest graphs of the current user through the multi-modal knowledge graph attention mechanism; generating an interest current representation based on multiple current interest graphs of the current user through the multi-modal knowledge graph attention mechanism; generating the aggregated embedding representation based on the interest historical representation and the interest current representation.
[0009] Further, through the multi-modal knowledge graph attention mechanism, an interest historical representation is generated based on multiple historical interest graphs of the current user, including: for each historical interest graph, aggregating information of K-hop neighbors through a multi-interest gated graph neural network based on an initial node embedding through K update steps to obtain a node embedding vector of an item of an interest corresponding to the historical interest graph; for each interest, aggregating node embedding vectors of all items in the multiple historical session sequences according to an interest recognition vector corresponding to the interest to obtain a historical sequence embedding of the interest, and generating an interest historical representation based on the historical sequence embedding of the interest.
[0010] Further, through the multi-modal knowledge graph attention mechanism, an interest current representation is generated based on multiple current interest graphs of the current user, including: for each current interest graph, generating an interest node embedding through a multi-interest gated graph neural network based on the adjacency matrix of the current interest graph; generating the interest current representation based on the interest node embedding, where the interest current representation consists of a session embedding based on the interest, a global preference session embedding, and a local preference session embedding.
[0011] Further, based on the interest history representation and the interest current representation, generate the aggregated embedding representation according to the following formula: C m = W 4 (c m + h m ), where C m is the context representation of the m-th interest, and W 4 is the projection matrix based on the sum of the interest current representation C m of the m-th interest and the interest history representation h m of the m-th interest.
[0012] Further, through a similar session retrieval network, retrieve the similar historical sessions of other users corresponding to the current session of the current user, including: obtaining the user preference representation of the current user; for each candidate similar session, determining the session similarity corresponding to the candidate similar session based on the user preference representation of the current user; and retrieving the similar historical sessions of other users corresponding to the current session of the current user based on the session similarity corresponding to the candidate similar session.
[0013] Further, generate a cross-session recommendation based on the aggregated embedding representation and the similar historical sessions of other users, including: generating a cross-session recommendation according to the session similarity, user preference, and the aggregated embedding representation of the similar historical sessions.
[0014] The present invention provides a cross-session recommendation system for a multi-interest fusion multi-modal knowledge graph, including: a sequence acquisition module for acquiring multiple historical session sequences and a current session sequence of the current user; a graph generation module for constructing multiple historical interest graphs of the current user based on the multiple historical session sequences of the current user, and further for constructing multiple current interest graphs of the current user based on the current session sequence; an embedding representation module for generating an aggregated embedding representation based on the multiple historical interest graphs and multiple current interest graphs of the current user through a multi-modal knowledge graph attention mechanism; a session retrieval module for retrieving the similar historical sessions of other users corresponding to the current session of the current user through a similar session retrieval network; and a session recommendation module for generating a cross-session recommendation based on the aggregated embedding representation and the similar historical sessions of other users.
[0015] Compared with the prior art, the cross-session recommendation method and system for a multi-interest fusion multi-modal knowledge graph provided by the present invention at least have the following beneficial effects:
[0016] Capture various interests of users from their current and historical behavior sequences, design a cross-session recommendation network to find historical sessions similar to the current session, and supplement the limited context information. Secondly, combine the overall interests of users with the current and historical behavior sequences to construct multiple interest graphs of the current and historical sequences, and generate accurate context representations for each interest. Finally, introduce a multi-modal knowledge graph in the process of interest modeling to better obtain the context information of project entities, and recommend the obtained aggregated embedding representations, improving the effect of cross-session recommendation. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] This specification will be further described in the form of exemplary embodiments, which will be described in detail through the drawings. These embodiments are not restrictive. In these embodiments, the same numbers represent the same structures, where:
[0018] Figure 1 is a schematic flow chart of a cross-session recommendation method for a multi-interest fusion multi-modal knowledge graph according to some embodiments of this specification;
[0019] Figure 2 is a schematic flow chart of a cross-session recommendation method for a multi-interest fusion multi-modal knowledge graph according to some embodiments of this specification;
[0020] Figure 3 is a schematic structural diagram of a multi-modal knowledge graph encoder according to some embodiments of this specification;
[0021] Figure 4 is a schematic module diagram of a cross-session recommendation system for a multi-interest fusion multi-modal knowledge graph according to some embodiments of this specification;
[0022] Figure 5a is a schematic diagram of the experimental results of the recall rate on dataset A according to some embodiments of this specification;
[0023] Figure 5b is a schematic diagram of the experimental results of the MMR on dataset A according to some embodiments of this specification;
[0024] Figure 5c is a schematic diagram of the experimental results of the recall rate on dataset B according to some embodiments of this specification;
[0025] Figure 5d is a schematic diagram of the experimental results of the MMR on dataset B according to some embodiments of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] To more clearly illustrate the technical solutions of the embodiments of this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some examples or embodiments of this specification. For those of ordinary skill in the art, without creative efforts, this specification can also be applied to other similar scenarios based on these drawings. Unless obvious from the language context or otherwise stated, the same reference numerals in the drawings represent the same structure or operation.
[0027] Figure 1 It is a schematic flowchart of a cross-session recommendation method for a multi-interest fusion multi-modal knowledge graph according to some embodiments of this specification. As Figure 1 shown, the cross-session recommendation method for a multi-interest fusion multi-modal knowledge graph may include the following steps.
[0028] Step 110, obtain multiple historical session sequences and the current session sequence of the current user.
[0029] Let the user U = {u 1 , …, u N} represent the set of all N users in the dataset, and the point set V = {v 1 , …, v M} represent the set of all M items in the dataset. Each user u ∈ U has a session sequence |S u | represents the number of sequences of the user. Each sequence contains interaction items in order and is represented as The current sequence contains t - 1 items in order and is represented as The purpose is to predict the next item of the current sequence All sequences that occurred before constitute the historical sequence set, denoted as For the current user u c 's current session 's target item In the current session All items that occurred before form the in-session context The historical session sets of other users are represented as Specifically, train a probability classifier to predict the conditional probability p(v|e l,m , S}) of each candidate item v ∈ V, where is the set containing all historical sessions and the current session during training.
[0030] Step 120, based on multiple historical session sequences of the current user, construct multiple historical interest graphs of the current user.
[0031] Figure 2 It is a schematic flowchart of a cross - session recommendation method for a multi - interest fusion multi - modal knowledge graph shown in some embodiments of this specification, as Figure 2 shown, specifically including:
[0032] Convert all items included in multiple historical session sequences into multi - dimensional embedding representations, generate a historical embedding matrix corresponding to the items, and based on the historical embedding matrix, determine the correlation scores between the items and interests;
[0033] Based on the correlation scores of the correlation scores between the items and interests, construct multiple historical interest graphs of the current user, where one historical interest graph corresponds to one interest.
[0034] For the historical set corresponding to multiple historical session sequences of the current user Construct a large historical cross - session graph G h , which contains items from all historical sequences of the current user, represents the unique set of items represented by nodes, and the connection between two consecutive items in the sequence is represented by a directed edge. By inputting the adjacency matrix and the output adjacency matrix to form a connection matrix A h , which is used to model the complex information propagation between nodes, and the connection matrix A h is shared by multiple historical interest graphs.
[0035] Convert all items included in multiple historical session sequences into D - dimensional embedding representations v ∈ W e , where W e ∈ R D×|V| is the embedding matrix of the items and is used as an input to calculate the D - dimensional embedding representation and the correlation between the items of k interests and a specific interest. The above process is expressed as follows:
[0036]
[0037] where, a i,m is the correlation between the item v i and the m - th interest, W I [:,m] is the m - th column of W I , and k is the total number of interests.
[0038] To identify the interests that drive the user to browse items, use the calculated correlation scores as inputs to perform the following operations:
[0039]
[0040] where, y i,mis a scalar close to 0 or 1, depending on whether the item searched by the user is the m-th interest. π j is noise with a mean of 0 and a standard deviation of 1 obtained from a distribution, y i,k represents a one-hot vector, indicating that each item is associated with a specific interest. k is the total number of interests. In this model, y i,k is set to 0.01.
[0041] Multiple historical interest graphs, denoted as In the m-th interest graph, each node represents a unique item v in all historical sequences i , and items that do not belong to the m interests have no impact on the m-th historical interest graph and will not appear in the m-th historical interest graph.
[0042] The interest-based items related to the m-th interest are generated as follows:
[0043] v i,m = y i,m * v i
[0044] where y i,m is a scalar close to 0 or 1, depending on whether the item searched by the user is the m-th interest. If v i,m is approximately 0, then v i,m is an approximately zero vector in the m-th historical interest graph, and the item v m will not appear in the m-th historical interest graph and will not be propagated on this interest graph. Conversely, the item v i aggregates the information of other items that appear in the m-th historical interest graph, preventing insufficient information aggregation between items associated with different interests. The m-th historical interest graph is generated through the following initial node embeddings: In , items that do not belong to the m-th interest are represented by zero vectors.
[0045] Step 130, based on the current session sequence, construct multiple current interest graphs of the current user.
[0046] As Figure 2 shown, specifically including:
[0047] Convert all items in the current session sequence into a multi-dimensional embedding representation to generate the current embedding matrix corresponding to the items;
[0048] For each interest, based on the current embedding matrix, calculate the relationship score between two items of the interest. Based on the relationship score, generate the attention matrix corresponding to the interest as the adjacency matrix, and based on the adjacency matrix, generate the current interest graph corresponding to the interest.
[0049] Model the complex item conversion patterns in the current session sequence and convert the current session sequence into a directed session graph G c , for G c , denote the nodes of the items in the current session sequence s c in G c The directed edges in G are represented by the transitions between an item and the next item, and the connection matrix A is constructed by combining the input and output and c .
[0050] First, use W e as the embedding matrix for the item conversions in the current session sequence. Subsequently, calculate the relationship scores between two items of k interests. The formula is as follows
[0051]
[0052] where denotes the relationship score between node i and node j for the m-th interest are the query part and the key part of the projection matrix for the m-th interest respectively. Perform masked attention using the incoming and outgoing adjacency matrices, i.e., and are normalized as follows
[0053]
[0054] where * represents the incoming or outgoing relationship denotes the normalized relationship score of node v i for the m-th interest, generate the α values for all nodes for each interest, and form the attention matrix [A' c,1 ,…,A' c,k based on k interests, which is used as the adjacency matrix of the corresponding current interest graph
[0055] The construction of multiple current interest graphs is as follows In the m-th current interest graph, each node represents an item v i in the current session sequence. When using the interest-based attention matrix as the adjacency matrix, each current interest graph has different attention weights between the same item node pairs. If the node pairs do not have a strong relationship with a specific interest, they will not propagate a large amount of information to each other in the interest graph, and vice versa. When the above conditions are met, the nodes of the same interest will aggregate information with each other, and the nodes of different interests will not propagate unnecessary information
[0056] Step 140: Generate an aggregated embedding representation based on multiple historical interest graphs and multiple current interest graphs of the current user through a multi-modal knowledge graph attention mechanism.
[0057] Specifically, it includes:
[0058] Generate a historical interest representation based on multiple historical interest graphs of the current user through a multi-modal knowledge graph attention mechanism;
[0059] Generate a current interest representation based on multiple current interest graphs of the current user through a multi-modal knowledge graph attention mechanism;
[0060] Generate an aggregated embedding representation based on the historical interest representation and the current interest representation.
[0061] In some embodiments, generating a historical interest representation based on multiple historical interest graphs of the current user through a multi-modal knowledge graph attention mechanism includes:
[0062] For each historical interest graph, use a multi-interest gated graph neural network to aggregate the information of K-hop neighbors based on the initial node embedding through K update steps to obtain the node embedding vector of the item of interest corresponding to the historical interest graph;
[0063] For each interest, aggregate the node embedding vectors of all items in multiple historical session sequences according to the interest recognition vector corresponding to the interest to obtain the historical sequence embedding of the interest, and generate a historical interest representation based on the historical sequence embedding of the interest.
[0064] The multi-interest gated graph neural network (GGNN) constructs an adjacency matrix A h to appropriately aggregate the information of adjacent nodes. For the m-th interest, GGNN uses as the initial node embedding, and aggregates the information of K-hop neighbors through K update steps to obtain This is the interest-based node embedding vector of the item of the m-th interest. The processing of the historical interest graph by GGNN is as follows:
[0065]
[0066]
[0067]
[0068]
[0069]
[0070]
[0070] Among them, is the project The relevance m between the project and the m-th interest, where K represents the information aggregation of K-hop neighbors in K update steps. H m ∈R D×2D and b m ∈R D are the weight and bias respectively, and z i and r i are the reset gate and update gate respectively. After K aggregations, the historical interest graph is obtained, and the interest-based node embeddings
[0071] The multi-modal encoder is an encoder that fuses multi-modal data, mainly used for fusing image, video, and text information. Figure 3 is the structural schematic diagram of the multi-modal knowledge graph encoder shown in some embodiments of this specification. As Figure 3 shown, first, the triple information (h, r, t) in the knowledge graph is generated into dense vectors. Second, the classical feature extraction network ResNet50 is used to extract image semantics. Then, the Google open-source model is used to extract text semantic features. Finally, all modalities are unified into a unified dimension and trained on the multi-modal model.
[0072] The multi-modal knowledge graph attention layer includes a propagation layer and an aggregation layer. The output information calculation formula of the propagation layer is as follows,
[0073]
[0074] where N h represents a triple space, e agg is the output information of the propagation layer, e(h, r, t) represents the embedding of the triple π(h, r, t), and π(h, r, t) represents the corresponding attention score. However, the calculation formula of the embedding representation e(h, r, t) is as follows,
[0075] e(h, r, t) = W 1tr (e h ||e r ||e t )
[0076] where W 1tr is the training matrix, and e h , e t , e r represent entity and relation embeddings respectively. And the calculation formula of π(h, r, t) is as follows,
[0077] π(h, r, t) = Leaky ReLU(W 2tr e(h, r, t))
[0078] Among them, W 2tr is similar to W 1tr and both represent trainable matrices. Then, all triple coefficients connected to entity h are normalized through the normalization function softmax. The formula is as follows,
[0079]
[0080] However, the aggregation layer linearly transforms e h and e agg through a concatenation method. This layer is mainly responsible for the concatenation operation, and the concatenation operation can be expressed as,
[0081]
[0082] where || is the concatenation operation, and W 3tr is similar to the previous W 2tr and W 1tr and is a training matrix.
[0083] Aggregate the embeddings of all items in the sequence according to the corresponding interest recognition vector y i to obtain the historical sequence embedding of the m-th interest in s h .
[0084]
[0085] where s l is the l-th historical session sequence in the set of historical session sequences s h , e l,m represents the sequence embedding of s l for the m-th interest, y i,m is the interest recognition y i of the m-th interest. After obtaining the embeddings of the historical sequences of each interest, these embeddings are used to generate the historical embeddings of each interest.
[0086] Based on the historical sequence embedding e m ={e 1,m ,…,e n,m} to ensure the consistency between historical interests and the current interest, the self-attention method is implemented with the representation c m of the current interest as the query,
[0087]
[0088] where d represents the dimension of the space R, h m ∈R d is the interest-based historical representation of the m-th interest, and W Q , W K , W V∈R d×d They are the projection matrices for query, key, and value respectively.
[0089] In some embodiments, through a multi-modal knowledge graph attention mechanism, based on multiple current interest graphs of the current user, an interest current representation is generated, including:
[0090] For each current interest graph, through a multi-interest gated graph neural network, based on the adjacency matrix of the current interest graph, an interest node embedding is generated;
[0091] Based on the interest node embedding, an interest current representation is generated, where the interest current representation consists of an interest-based session embedding, a global preference session embedding, and a local preference session embedding.
[0092] Specifically, the modeling formula for the complex item conversion pattern of a specific interest is as follows:
[0093]
[0094] Where Contains the interest-based node embeddings of the items in the current sequence, used to generate the interest-based current representation, A' c,m Is the attention matrix for the m-th interest.
[0095] Based on the interest node embedding To generate an interest-based current representation, which consists of three types of session embeddings: interest-based session embedding, global preference session embedding, and local preference session embedding. Interest-based session embedding: For the items in the current sequence, the relationship scores between different items and interests are different. Considering the relationship scores between items and interests, a weight vector is generated, and weighted to obtain the interest-based session embedding for the m-th interest, which is expressed as follows:
[0096]
[0097]
[0098] Where Represents the attention score of a certain node i at the m-th interest point, Is the set of interest-based node embeddings for the m-th interest in the current sequence, a m ∈R is the weight vector for the m-th interest, And Are the projection matrices for the m-th interest, and the superscript T represents the transpose of a matrix or vector, Represents the interest-based session embedding for the m-th interest, and n is the number of all embedding nodes.
[0099] Global preference session embedding: Global preference session embedding is based on the last clicked session embedding, and the weighted sum of all node embeddings is obtained as the global preference session embedding. All attention scores are calculated as follows,
[0100]
[0101]
[0102] where b represents the bias, represents the global session embedding of the m-th interest, and the parameter q ∈ R d , W 1 , W 2 ∈ R d×d is the weight matrix of the attention scores.
[0103] Local preference session embedding is used to represent the local preference of a sequence, as follows:
[0104]
[0105] where, represents the local preference session embedding of the m-th interest.
[0106] Finally, the session embedding of the interest, the global preference session embedding, and the local session embedding are added to obtain the sum vector, and W 3 ∈ R d×2d is used for transformation to obtain the final interest-based current representation of the m-th interest.
[0107]
[0108] where c m ∈ R d is the interest-based current representation of the m-th interest, and W 3 ∈ R d×d is the projection matrix of the sum vector. In some embodiments, the aggregated embedding representation is generated based on the interest history representation and the interest current representation according to the following formula:
[0109] C m = W 4 (c m + h m )
[0110] where C m is the context representation of the m-th interest, and W 4 is the projection matrix based on the sum of the interest current representation C m of the m-th interest and the interest history representation h m of the m-th interest.
[0111] After obtaining the interest-based context representation C = {C 1 , …, C k}, they are integrated to generate the target-aware context representation S t . The interest recognition network is used to calculate the relationship weights between the target item and the interests, and the relationship weights are used to obtain the weighted sum of C to generate S t .
[0112]
[0113]
[0114]
[0115] Among them, a t,m represents the correlation between item v t and the m-th interest, and ω t,m is the relevant weight of the target item v t of the m-th interest.
[0116] Step 150, through the similar session retrieval network, retrieve the similar historical sessions of other users corresponding to the current session of the current user.
[0117] Specifically, it includes:
[0118] Obtain the user preference representation of the current user;
[0119] For each candidate similar session, based on the user preference representation of the current user, determine the session similarity corresponding to the candidate similar session;
[0120] Based on the session similarity corresponding to the candidate similar session, retrieve the similar historical sessions of other users corresponding to the current session of the current user.
[0121] To alleviate the information limitation of the current session, by retrieving the historical sessions of the current user and other users from other sessions, that is, according to the user preference representation h learned from the interest-based local preference sessions c , and for the candidate sessions e or in the above, the interest-based historical sequence preference set h l,m of each user has been obtained i ∈ e l,m , and its similarity λ c with h i,c is calculated by the formula
[0122]
[0123] where \(i\in[1,t]\) is the position of the item in \(e\), by taking the maximum similarity between items in \(e\) l,m and to obtain the similarity between \(e\) l,m and , as the minimum distance between \(h\) c and the candidate similar session \(e\) l,m :
[0124]
[0125] Once the candidate session \(e\) l,m is similar to \(h\) c , the preferences of the user in \(e\) l,m are used to supplement the current session, and an attention-based session encoder is used to represent the user preferences in \(e\) l,m . Embed \(u\) e into the \(d\)-dimensional user preference embedding \(\theta\) e . Then, the preference calculation formula for the item \(v\) l,m in \(e\) i is,
[0126]
[0127] where, is the normalization factor. And the calculation formula for the user preference \(e\) l,m is,
[0128]
[0129] where, \(w\) e represents the value of the user preference \(e\) l,m .
[0130] Step 160, generate cross-session recommendations based on the aggregated embedding representation and the similar historical sessions of other users.
[0131] Specifically, it includes:
[0132] Generate cross-session recommendations according to the session similarity, user preferences, and aggregated embedding representation of the similar historical sessions.
[0133] As Figure 2 shown, after the session similarity and user preferences of each session in and are ready, respectively take the similarity between and as the weight to aggregate the user preferences of all sessions in the two candidate sets,
[0134]
[0135]
[0136] Finally, apply the softmax function to calculate the probability that the target item becomes the next item.
[0137]
[0138] Select the items with the top-K conditional probabilities to form the recommendation list.
[0139] Next, the performance of the cross-session recommendation method for the multi-interest fusion multi-modal knowledge graph verified by experimental data is presented.
[0140] Dataset A comes from an e-commerce website. This dataset covers user behaviors, including browsing, purchasing, and commenting.
[0141] Dataset B is collected from an online forum. This dataset includes the list of sub-forums commented by users and the timestamps of user comments, etc.
[0142] First, sequences with fewer than 5 items are filtered out from Datasets A and B to obtain a processed dataset of appropriate size. Dataset B includes session IDs and user IDs, which are used to generate the historical and current behavior sequences of users. Dataset A does not contain sufficient information to identify each session, and the sequences in this dataset are manually divided into parts corresponding to 1 day. For Datasets A and B, at least four items are used as the input of the current sequence, and the last item is the target item. Items that appear fewer than 100 and 5 times in Datasets A and B are respectively removed from the sequences. For each current sequence, the number of corresponding historical sequences generated for Datasets A and B is at least three. Subsequently, all sequences are sorted according to the timestamps, and all datasets are respectively divided into an 80% training set and a 20% test set. The summary statistics of the datasets are shown in Table 1. Among them, user represents the number of users, item represents the number of items, train session represents the number of training sessions, test session represents the number of test sessions, avg.length represents the average length of sessions, interactions represents the number of relationships, interactions persession represents the average number of relationships per session, and interactions per user represents the average number of relationships per user.
[0143] Table 1
[0144] dataset Dataset A Dataset B user 81643 18173 item 44749 9102 trainsession 443108 66283 testsession 139052 27480 avg. length 14.05 7.63 interactions 1257639 188050 interactionspersession 6.4 3.6 interactionsperuser 165.7 159.2
[0145] To evaluate the performance of the proposed method, it is compared with the following baseline methods: GRU4REC, STAMP, and SASRec methods based on RNN; SR-GNN, GC-SAN, and TA-GNN methods based on GNN; and SURGE and TGSRec, which are multi-interest graph neural network modeling methods that utilize high-order correlations to derive user interests. Among them, SURGE integrates the user's current core interests in the behavior sequence. STAMP captures the general interests and current interests in the session. SASRec is a self-attention-based model that captures long-term semantics and focuses on important behaviors. GC-SAN is a method that combines GGNN with the self-attention mechanism. TA-GNN is a target-aware attention modeling method with GGNN. TGSRec considers the temporal dynamics within the sequential pattern. GRU4REC captures the general interests in the session through the RNN method. SR-GNN constructs a session graph to simulate complex item transitions.
[0146] Select K items to generate the top-K recommendation list, and let K be 20. Rec@K (recall): represents the proportion of true target items that appear in the top-K recommendations; MRR@K (mean reciprocal rank): the ranking in the top-K recommendation list.
[0147]
[0148]
[0149] Among them, N is the total number of sessions, and n hit is the number of true target items in the top-K recommendations, and Rank(v) is the ranking of item v in the top-K recommendations.
[0150] Set the dimension of the embedding vector to 100. By adjusting the test data, set the number of interests in dataset A to 7 and the number of interests in dataset B to 9. All parameters are initialized using a Gaussian distribution with a mean of 0 and a standard deviation of 0.1 to ensure that all calculations are within the same range. Select the Adam optimizer for training. The batch sizes for datasets A and B are 30, the initial learning rate is 0.001, and 20% dropout is used to avoid overfitting. The comparison of the performance of the proposed method and the baselines is shown in Table 2.
[0151] Table 2
[0152] Serial number algorithm Dataset A Dataset A Dataset B Dataset B Rec@20 MRR@20 Rec@20 MRR@20 1 GRU4REC 31.09 11.37 44.76 22.28 2 STAMP 35.42 12.38 61.55 36.57 3 SASRec 19.89 6.24 61.91 36.94 4 SR-GNN 35.17 15.66 62.71 38.49 5 GC-SAN 33.59 14.29 57.39 35.73 6 TA-GNN 38.61 16.83 64.14 41.02 7 TGSRec 32.28 13.62 52.73 35.06 8 SURGE 34.72 14.19 58.45 36.29 9 This method 40.83 19.02 65.87 42.85
[0153] As shown in Table 2, this method is superior to the baseline method. Since RNN models the sequential dependencies in time series data, RNN models such as GRU4REC and STAMP are widely used to identify features in sequences. GRU4REC uses GRU to simulate the general interests in a session, and STAMP uses short-term memory to model the last clicked item as the current interest. Experiments show the importance of modeling short-term behavior in next item prediction. SASRec does not use RNN, but uses self-attention mechanism to capture long-term semantics and identify previous actions at each time step. The above methods integrate the information of the entire session and only consider the relationships between items when modeling the transition patterns between consecutive items.
[0154] For traditional sequential models of RNN and attention mechanism, it is difficult to achieve better performance compared with time-aware models and multi-interest models, and it is not suitable for modeling complex and diverse user interests. As graph convolution methods, TGSRec and SURGE models select multi-level correlations between historical items to refine user preferences, but neither of them realizes more accurate multi-interest extraction and obtains better recommendations by aggregating multi-level user preferences. SURGE forms dense clusters in the interest graph to distinguish users' core interests, and performs graph convolution propagation of cluster awareness and query graph to fuse users' current core interests and behavior sequences. TGSRec constructs a global graph and uses the interactions at different time points as edges. This design choice requires more computational cost for graph retrieval and convolution. The above two methods do not achieve good results compared with the method proposed in this specification.
[0155] SR-GNN converts sequences into graph structures to capture complex item dependencies. In GC-SAN, GGNN is combined with the self-attention mechanism to retain relevant information and ignore noise; TA-GNN uses target-aware attention to construct a session embedding and is superior to other GNN-based methods on the adopted datasets, which shows the importance of considering the target item in recommendations. The above GNN-based methods only aggregate various information sequentially by obtaining the representation of a single session, and these representations cannot be combined with the context related to different interests in the sequence, lacking consideration of multi-session data. The method proposed in this specification involves extracting multi-interest context and obtaining the correlations between different items and interests. In addition, due to the cold start problem with insufficient historical items and sessions, and connecting consecutive sessions into an item sequence will destroy the intrinsic structure of the session, this specification designs a retrieval module for similar sessions to find the session users who are truly similar to the current session of the current user from the historical sessions of the current user and other users, so as to effectively supplement the limited context information in the current session. Therefore, the method proposed in this specification is superior to the baseline methods in the research.
[0156] The experiment was conducted with the interest numbers set to {1, 3, 5, 7, 9}. The results obtained by the cross-session recommendation method for the multi-interest fusion multi-modal knowledge graph on two datasets with different numbers of interests are as Figures 5a to 5d shown.
[0157] The results show that the method proposed in this specification captures multi-interest context from sequential data, examines different numbers of interests through experiments, and the multi-interest structure is superior to the single-interest structure. At the same time, it shows that the multi-interest structure can capture comprehensive features from the user's behavior sequence and avoid information loss.
[0158] For dataset A, when the number of interests is set to 7, the performance of the method is the best. When the number of interests increases to 9, the performance of the method slightly decreases, and more interests do not necessarily bring superior performance. For dataset B, the model with a single-interest structure is slightly better than all models with a multi-interest structure in terms of Recall@50. Since the number of target items included is the smallest, it is easier to select the correct target item into the recommendation list than to correctly rank the target items. The multi-interest structure proposed in this specification is superior to the single-interest structure, extracts multiple contexts from the sequence and represents them with multiple vectors to prevent information loss.
[0159] To verify the effectiveness of the cross-session recommendation method for multi-interest fusion multi-modal knowledge graph (hereinafter referred to as "CRFMKG") proposed in this specification, three simplified versions were designed, namely CRFMKG-c, CRFMKG-o, and CRFMKG-h. They represent removing steps such as the current session, the global module, and the historical session respectively.
[0160] The experimental results of the three simplified versions and the complete version are shown in Table 3. It can be seen that CRFMKG-c performs poorly because it only uses limited context information in the current session for recommendation; the performance of CRFMKG-h is better than that of CRFMKG-c because it obtains sessions similar to the current session from the historical sessions of the current user to supplement the limited context information in the current session; the result of CRFMKG-o is slightly better than that of CRFMKG-h, indicating that the historical sessions of other users similar to the current user can play an important role. The CRFMKG model performs the best because it finds the sessions similar to the current session from the historical sessions of the current user and other users.
[0161] Table 3
[0162] algorithm Rec@5 Recl@20 MRR@5 MRR@20 CRFMKG-c 36.17 37.79 15.69 16.95 CRFMKG-h 36.94 38.53 16.17 17.63 CRFMKG-o 37.62 39.21 16.82 18.58 CRFMKG 38.12 40.83 17.96 19.02
[0163] This specification presents a cross-session recommendation method for multi-interest fusion multi-modal knowledge graphs. First, in view of the idea of few-shot learning, a novel cross-session recommendation network is designed. By using global and local modules, it captures the user's preferences from the current session and optimizes them using the learned prior knowledge to accurately recommend the next item in the current session. Second, an interest graph is constructed based on the basic graph neural network to capture the complex item transition patterns of various interests. Then, a multi-modal knowledge graph is introduced, and through the use of the multi-modal graph attention mechanism for information propagation, a multi-interest representation aggregated based on the user's multiple interests is generated. Finally, extensive experiments are conducted on two real-world datasets, and the experiments show that the proposed cross-session recommendation method for multi-interest fusion multi-modal knowledge graphs outperforms the baseline algorithms.
[0164] Figure 4 It is a schematic diagram of the modules of a cross-session recommendation system for multi-interest fusion multi-modal knowledge graphs according to some embodiments of this specification. As Figure 4 shown, the cross-session recommendation system for multi-interest fusion multi-modal knowledge graphs may include a sequence acquisition module, a graph generation module, an embedding representation module, a session retrieval module, and a session recommendation module.
[0165] The sequence acquisition module can be used to obtain multiple historical session sequences and the current session sequence of the current user.
[0166] The graph generation module can be used to construct multiple historical interest graphs of the current user based on multiple historical session sequences of the current user, and can also be used to construct multiple current interest graphs of the current user based on the current session sequence;
[0167] The embedding representation module can be used to generate an aggregated embedding representation based on multiple historical interest graphs and multiple current interest graphs of the current user through the multi-modal knowledge graph attention mechanism;
[0168] The session retrieval module can be used to retrieve similar historical sessions of other users corresponding to the current session of the current user through a similar session retrieval network;
[0169] The session recommendation module can be used to generate cross-session recommendations based on the aggregated embedding representation and the similar historical sessions of other users.
[0170] The cross-session recommendation system for multi-interest fusion multi-modal knowledge graphs can be used to execute the cross-session recommendation method for multi-interest fusion multi-modal knowledge graphs. For more descriptions of the cross-session recommendation system for multi-interest fusion multi-modal knowledge graphs, reference can be made to the relevant descriptions of the cross-session recommendation method for multi-interest fusion multi-modal knowledge graphs, which will not be elaborated here.
[0171] Finally, it should be understood that the embodiments described in this specification are only used to illustrate the principles of the embodiments of this specification. Other variations may also fall within the scope of this specification. Therefore, by way of example and not limitation, alternative configurations of the embodiments of this specification may be regarded as consistent with the teachings of this specification. Accordingly, the embodiments of this specification are not limited to the embodiments explicitly presented and described in this specification.
Claims
1. A cross-session recommendation method based on multi-interest fusion and multi-modal knowledge graph, characterized in that: include: Get multiple historical session sequences and current session sequence of the current user; Based on the multiple historical conversation sequences of the current user, construct multiple historical interest graphs of the current user; Based on the current session sequence, construct multiple current interest graphs of the current user; Generate a clustered embedding representation based on the multiple historical interest graphs and the multiple current interest graphs of the current user through a multimodal knowledge graph attention mechanism; Retrieving similar historical sessions of other users corresponding to the current session of the current user through a similar session retrieval network; generating cross-session recommendations based on the clustered embedding representation and similar historical sessions of the other users; Among them, generating a clustered embedding representation includes: Through the multimodal knowledge graph attention mechanism, the interest history representation is generated based on multiple historical interest graphs of the current user; For each current interest graph, generate interest node embedding based on the adjacency matrix of the current interest graph through a multi-interest gated graph neural network, and generate interest current representation based on the interest node embedding, wherein the interest current representation consists of interest-based session embedding, global preference session embedding and local preference session embedding; Generate a clustered embedding representation based on the historical interest representation and the current interest representation; For the items in the current session sequence, the relationship scores between different items and interests are different. Considering the relationship scores between items and interests, a weight vector is generated, and the weighting is determined to obtain the interest-based session embedding of the mth interest: represents the weight of the i-th node in the m-th interest, is the set of interest-based node embeddings of the mth interest in the current session sequence, is the weight vector of the mth interest, and is the projection matrix of the mth interest, the superscript T indicates the transpose of the matrix or vector, represents the interest-based session embedding of the mth interest, where n is the number of entire embedding nodes; The global preference session embedding is based on the last clicked session embedding. The weighted sum of all node embeddings can be obtained as the global preference session embedding. All its attention scores are calculated as follows: represents the attention score of the i-th node at the m-point of interest, b represents the bias, Represents the global session embedding of the mth interest, with parameters , is the weight matrix of attention scores; ; in, The local preference session embedding representing the mth interest; The interest session embedding, global preference session embedding, and local preference session embedding are added together to get the sum vector, and the Transform to obtain the final interest-based current representation of the mth interest: is the current interest-based representation of the mth interest, is the projection matrix of the summed vector.
2. The cross-session recommendation method of multi-interest fusion multi-modal knowledge graph according to claim 1 is characterized in that: Based on the multiple historical conversation sequences of the current user, constructing multiple historical interest graphs of the current user, including: Convert all items included in the multiple historical conversation sequences into multi-dimensional embedding representations, generate historical embedding matrices corresponding to the items, and determine the relevance scores between the items and the interests based on the historical embedding matrix; Based on the correlation scores between the items and interests, a plurality of historical interest graphs of the current user are constructed, wherein one historical interest graph corresponds to one interest.
3. The cross-session recommendation method of multi-interest fusion multi-modal knowledge graph according to claim 1 is characterized in that: Based on the current session sequence, constructing multiple current interest graphs of the current user, including: Convert all items in the current conversation sequence into multi-dimensional embedding representations to generate current embedding matrices corresponding to the items; For each interest, based on the current embedding matrix, the relationship score between the two items of the interest is calculated, based on the relationship score, the attention matrix corresponding to the interest is generated as an adjacency matrix, and based on the adjacency matrix, the current interest graph corresponding to the interest is generated.
4. The cross-session recommendation method of multi-interest fusion multi-modal knowledge graph according to claim 1 is characterized in that: By using the multimodal knowledge graph attention mechanism, an interest history representation is generated based on the multiple historical interest graphs of the current user, including: For each of the historical interest graphs, a node embedding vector of the items of interest corresponding to the historical interest graph is obtained by aggregating information of F-hop neighbors through F update steps based on initial node embedding through a multi-interest gated graph neural network; For each interest, node embedding vectors of all items in the multiple historical conversation sequences are aggregated according to the interest identification vector corresponding to the interest to obtain the historical sequence embedding of the interest, and based on the historical sequence embedding of the interest, a historical representation of the interest is generated.
5. The cross-session recommendation method of multi-interest fusion multi-modal knowledge graph according to claim 1 is characterized in that: The aggregated embedding representation is generated based on the interest history representation and the interest current representation according to the following formula: in, is the context representation of the mth interest, is the current representation of the mth interest based on interest and the interest history representation of the mth interest The projection matrix of the sum.
6. The cross-session recommendation method of multi-interest fusion multi-modal knowledge graph according to any one of claims 1 to 5, characterized in that: Retrieving similar historical sessions of other users corresponding to the current session of the current user through a similar session retrieval network, including: Get the user preference representation of the current user; For each candidate similar session, determining a session similarity corresponding to the candidate similar session based on the user preference representation of the current user; Based on the session similarities corresponding to the candidate similar sessions, similar historical sessions of other users corresponding to the current session of the current user are retrieved.
7. The cross-session recommendation method of multi-interest fusion multi-modal knowledge graph according to claim 6 is characterized in that: Generating a cross-session recommendation based on the clustered embedding representation and similar historical sessions of the other users, including: Cross-session recommendations are generated according to session similarities and user preferences of similar historical sessions and the aggregated embedding representation.
8. A cross-session recommendation system based on multi-interest fusion and multi-modal knowledge graph, characterized by: The cross-session recommendation method using the multi-interest fusion multimodal knowledge graph described in any one of claims 1 to 7 comprises: A sequence acquisition module is used to acquire multiple historical session sequences and current session sequence of the current user; A graph generation module, configured to construct multiple historical interest graphs of the current user based on multiple historical conversation sequences of the current user, and further configured to construct multiple current interest graphs of the current user based on the current conversation sequence; An embedding representation module, configured to generate a clustered embedding representation based on a plurality of historical interest graphs and a plurality of current interest graphs of the current user through a multimodal knowledge graph attention mechanism; A session retrieval module, used to retrieve similar historical sessions of other users corresponding to the current session of the current user through a similar session retrieval network; The session recommendation module is used to generate cross-session recommendations based on the clustered embedding representation and similar historical sessions of the other users.
Citation Information
Patent Citations
Session recommendation method and system based on session data
CN114282077A
Session recommendation method, system and device and storage medium
CN114691981A
Graph neural network session recommendation method based on graph embedding and attention mechanism
CN116485501A