Session recommendation method for heterogeneous multi-view enhanced subsequence unit learning

Through the heterogeneous multi-view enhancer sequence unit learning method, the information deviation problem caused by the construction of graph structure in a single perspective in the prior art is solved, and the global context information of multiple sessions is better captured, and the performance of the recommendation system is improved.

CN120045780AActive Publication Date: 2025-05-27NANTONG UNIV

Patent Information

Application Number
CN202510063432.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-27
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

The prior art adopts a single perspective method when building graph structures, resulting in information bias, unable to integrate user intentions from a higher dimension, and it is difficult to capture global context information in multiple sessions.

Method used

Using heterogeneous multi-view enhancer sub-sequence unit learning method, multiple consecutive projects are treated as one sub-sequence unit, and learn at the sub-sequence level, dynamically adjust the number of items in the sub-sequence, combine single-view and heterogeneous multi-view structures, and use a mixed read function and a multi-head attention mechanism to generate session-level embedding vectors.

Benefits of technology

It better captures the context information in the global session, provides more accurate user intention understanding and personalized recommendations, and improves the performance of the recommendation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045780A_ABST
    Figure CN120045780A_ABST
Patent Text Reader

Abstract

The invention discloses a session recommendation method for heterogeneous multi-view enhanced subsequence unit learning, particularly relates to the technical field of session-based recommendation, and solves the technical problems that in the prior art, interaction between single items lacks rich global context information, and it is difficult to understand the intention of a user from a higher-dimensional angle. According to the technical scheme, a plurality of continuous items are regarded as a sub-sequence unit, and learning is carried out on the sub-sequence level; according to the method, intentions of users are explored in a specific range through subsequences instead of only paying attention to direct relationships among items; the number of items in the subsequences can be dynamically adjusted, so that the influence of the subsequences with different lengths on the recommendation performance is explored; according to the invention, user intentions can be learned by applying the sub-sequence units in a plurality of sessions, and context information in a global session is better captured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of session-based recommendation, and particularly to a session recommendation method for heterogeneous multi-view enhanced subsequence unit learning. Background Art

[0002] With the rapid development of the Internet and the explosive growth of digital information, a vast amount of data has flooded into people's lives. To help users discover and obtain the content they want from this huge amount of information, recommendation systems have emerged. These systems build user models by analyzing users' historical behaviors and provide personalized recommendation services. Session-based recommendation (SBR) is an emerging recommendation paradigm, different from sequential recommendation. Its goal is to capture users' short-term dynamic preferences by analyzing the dependencies between items in a session and accurately predict the next item that the user is most likely to interact with.

[0003] Methods based on graph neural networks (GNNs) model session sequences as graphs to show the interaction relationships between items and learn user behaviors through multiple layers of GNNs. Graph neural networks can be applied to local sessions to consider the relationships between items within a session; they can also be applied to multiple global sessions to study the relationships between items in multiple sessions, thereby expanding the scope of item influence. However, most existing work usually adopts a single-perspective approach when constructing the graph structure, that is, regarding each item as an independent individual. These methods may introduce information bias because they only rely on the relationship between two items and cannot integrate users' intentions from a higher-dimensional perspective. Information from sessions of different lengths can reveal different user intentions. If the user's interests can be globally grasped from a higher dimension, it will help to provide better recommendations. Although some methods also consider learning from sessions of different granularities, they only operate in local sessions and cannot obtain rich information from the global context.

[0004] Therefore, how to apply subsequence units to learn user intentions in multiple sessions and better capture the context information in global sessions is the technical problem to be solved by the present invention. Summary of the Invention

[0005] Therefore, the present invention solves the technical problem in the prior art that the interaction between individual items lacks rich global context information and it is difficult to understand users' intentions from a higher-dimensional perspective. A session recommendation method for heterogeneous multi-view enhanced subsequence unit learning provided by the present invention regards multiple consecutive items as a subsequence unit and conducts learning at the subsequence level; the number of items in the subsequence can be dynamically adjusted to explore the influence of subsequences of different lengths on the recommendation performance; subsequence units can be applied to multiple sessions to learn user intentions and better capture the context information in global sessions.

[0006] The present invention provides a session recommendation method for heterogeneous multi-view enhancer subsequence unit learning, which consists of four main stages. The first stage is the view construction stage, where the session sequence is constructed into a single view and heterogeneous multi-views. The second stage is the embedding learning stage, which uses a hybrid reading function to learn the embeddings between different scale views. Next, the fusion layer aggregates the information from the two views to generate a session-level embedding vector. Finally, in the result prediction stage, all items are scored through a sample adaptive loss function, and the item with the highest score is recommended.

[0007] Further, first, the global session is modeled as a single view and a heterogeneous multi-view structure. The single view is constructed based on a single session s and is defined as G s =(X s , E s ), where X s is the set of items interacted with by the session sequence s, is the set of directed edges; each edge records the interaction between items, indicating that an anonymous user clicks on items x i and x j in sequence; each item has a self-loop, which is represented by a self-connecting edge r self ; in addition, to represent the influence from other items, two relationships are defined: forward information propagation r in and reverse information propagation r out , r in represents the transmission of information from x i to x j , while r out represents the information transmission from x j to x j under the same conditions. The connection strength in the single view is stored through the adjacency matrix A s , which records the interaction relationships between |X s | non-repeating items, and the interaction strength between two items is defined as:

[0008]

[0009] k 1 represents the result after row normalization of the adjacency matrix A s .

[0010] The heterogeneous multi-view structure is used to represent the relationships between items at multiple granularities. Each node in the heterogeneous multi-view consists of multiple consecutive items, which come from the subsequences in the same batch of sessions; the heterogeneous multi-view where is the set of subsequences ; any one subsequence where l * represents the number of consecutive items in the current subsequence; when l * = 1, the multi-view degrades to a single view; is the set of edges, representing the number of connections of the subsequence; represents a directed edge from the subsequence to , and the lengths of both subsequences are l * ; the adjacency matrix stores the interaction frequency between the subsequences and :

[0011]

[0012] k 2 is the result after row normalization in the adjacency matrix .

[0013] Furthermore, after the session sequence is constructed into a single view and a heterogeneous multi-view, the single-view attention network (SV-GAT) and the heterogeneous multi-view neural network (HMV-GNN) are used to obtain the embeddings of items and subsequences. SV-GAT is based on the transition relationship between items in the single view and aims to learn the correlation between individual items. To achieve this goal, any item x i in the single view is mapped to the embedding space to obtain the corresponding embedding vector. The graph attention network is used to capture the degree of association e ij between an item and other items:

[0014] e ij = LeakyRelu(W ij (x i ⊙x j ))

[0015] The LeakyRelu activation function is used to calculate the influence degree between items, representing the element-wise product of the vector sum, and W is a learnable parameter.

[0016] Normalization is performed through the Softmax function:

[0017]

[0018] N xi is the set of 1-hop neighbors of item x i , and α ij represents the transformation relationship between an item and its neighbors; the neighbor information is weighted and summed to obtain the final feature representation of the item:

[0019]

[0020] The set of all item-level embedding vectors in a single view is denoted as X SV-GAT ; The item embedding vectors are updated through a multi-layer attention network, and the formula is:

[0021]

[0022] l 1 represents the number of attention layers of SV-GAT, and A s represents the normalized value storing the interaction relationship between items.

[0023] Furthermore, different from a single view, each node in the heterogeneous multi-view consists of multiple consecutive items; for a subsequence of length l * To learn the complete intention of this subsequence, a hybrid reading function R H is designed to generate the embedding vector m i of the heterogeneous subsequence, and the formula is as follows:

[0024]

[0025] Specifically, both set-based reading functions (such as Mean and Max) and sequence-based reading functions (such as the gated recurrent unit GRU) are considered. The former learns the intention with invariant order, and the latter is better at capturing order-sensitive intentions. For a subsequence of length l * of the subsequence The final intention of this subsequence is calculated by the following formula:

[0026]

[0027] and respectively retain the intention with invariant order and order-sensitive intention of the subsequence; the set of all subsequences is denoted as M.

[0028] Furthermore, HMV-GNN captures high-dimensional user intention information based on the transition relationship in the heterogeneous multi-view. To achieve this effect, a multi-head attention mechanism is adopted, which not only considers the influence of subsequences but also enhances the generalization ability of the present invention. For the attention of a single head, its calculation formula is as follows:

[0029]

[0030] head i = Attention(MW i Q ,MW i K ,MW i V );

[0031] M represents the input of the subsequence, W i Q , W i K and W i V are the learnable parameters of the i-th head. The operation of the multi-head attention mechanism is to concatenate the results of n single-head attention networks and normalize them through the softmax function; to form the latent vector representation of the session sequence:

[0032] M HMV-GNN = Multi-Attention(M; W att ) = [head 1 , …, head n W att

[0033] W att represents the learnable parameters of the multi-head attention network;

[0034] After passing the information to the l 2 layer, the final embedded vector representation is:

[0035]

[0036] M and M HMV-GNN respectively represent the initial vector and the learned vector of all subsequences in the heterogeneous multi-view.

[0037] Furthermore, in order to avoid overfitting of the training model, the Dropout operation is introduced when learning two-level information:

[0038]

[0039] The item-level and subsequence-level embedded representations corresponding to the single-view attention network and the heterogeneous multi-view neural network are extracted from and respectively; for each session, these two types of embedded representations are denoted as X SV-GAT and M HMV-GAT respectively; in order to better utilize the information in the subsequence, a separate session embedded representation is generated at the subsequence level; given a session s and its corresponding subsequence unit where l * represents the number of consecutive items in the subsequence, select the last subsequence as the local representation; in order to capture rich comprehensive information, calculate the influence weights of the local representation and the context , and use the soft attention mechanism to calculate the attention score Thus, a global representation of the subsequence is generated

[0040]

[0041] and is a learnable weight matrix, is a bias vector, and an embedded representation of heterogeneous multi-view content is generated through information splicing The formula is as follows:

[0042]

[0043] || represents the splicing operation, is a learnable parameter; the subsequence-level embedding contains information on the global view and local view of the subsequence;

[0044] The last item clicked by the user often highly coincides with their current preference and significantly helps with the recommendation task. Therefore, the last clicked item x l ∈X SV-GAT is used to generate the final session-level embedded representation s f :

[0045] μ is an adjustable parameter used to control the weight distribution between session information and the information of the last clicked item.

[0046] Furthermore, in order to calculate the probabilities of all candidate items, the session embedding s f is interacted with the initial item vector x i and the next item most likely to be interacted with in the current session is selected through the softmax function

[0047] In addition, since the prediction difficulty of different sessions varies due to samples, a sample-adaptive loss function is proposed. All items are scored through the sample-adaptive loss function, and this loss function introduces a modulation factor SA i , which is added to the cross-entropy loss function; weights are assigned according to the prediction deviation of the samples:

[0048]

[0049] L SA =-∑(2 - 2·SA i ) τ log(SA i )

[0050] y iIt represents the true label for the next item click, and τ is the temperature coefficient; the SA loss function assigns weights according to the prediction deviation of the samples; when τ decreases, the modulation factor reduces the loss contribution of simple samples; conversely, when τ increases, more attention is paid to high-confidence samples and their contributions.

[0051] In the above technical solution, the technical effects and advantages provided by the present invention are as follows:

[0052] 1. The present invention regards multiple consecutive items as a subsequence unit and conducts learning at the subsequence level. It explores the user's intention within a specific range through the subsequence, rather than simply focusing on the direct relationship between items. In addition, the number of items in the subsequence can be dynamically adjusted to explore the impact of subsequences of different lengths on the recommendation performance.

[0053] 2. The present invention uses a hybrid reading function to fuse the intentions of each item in the corresponding subsequence unit, considering both order-insensitive and order-sensitive intentions. In addition, the sample adaptive loss function differentiates the weights of items in the session, enhancing the fitting ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings.

[0055] Figure 1 It is the overall framework diagram of a session recommendation method for heterogeneous multi-view enhanced subsequence unit learning according to the present invention;

[0056] Figure 2 It is the flowchart of a session recommendation method for heterogeneous multi-view enhanced subsequence unit learning according to the present invention;

[0057] Figure 3 It is the construction example diagram of single view and heterogeneous multi-view according to the present invention;

[0058] Figure 4 It is the recommendation performance diagram of the recommendation method according to the present invention for subsequences of different lengths. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0059] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the following will further introduce the present invention in detail in conjunction with the drawings.

[0060] Embodiment 1:

[0061] Please refer to Figure 1, as shown in the overall framework diagram of a session recommendation method for heterogeneous multi-view enhanced subsequence unit learning. The present invention provides a technical solution, which consists of four main stages: The first stage is the view construction stage, where multiple session sequences are constructed into single views and heterogeneous multi-views. The second stage is the embedding learning stage, where a hybrid reading function is used to learn the embeddings between views of different scales. Next, the fusion layer aggregates the information from two graph neural networks to generate a session-level embedding vector. Finally, in the result prediction stage, all items are scored through sample adaptive loss, and the item with the highest score is recommended.

[0062] A flowchart of a session recommendation method for heterogeneous multi-view enhanced subsequence unit learning is as Figure 2 shown.

[0063] First, the global session is modeled as a single view and a heterogeneous multi-view structure. Please refer to Figure 3 the construction example diagrams of the single view and heterogeneous multi-view of the present invention. The single view is constructed based on a single session s, defined as G s =(X s , E s ), where X s is the set of items interacted with by the session sequence s, is the set of directed edges; each edge records the interaction between items, indicating that an anonymous user clicks on items x i and x j in sequence; in order to retain the information of the item itself, each item has a self-loop, represented by the self-connecting edge r self ; in addition, in order to represent the influence from other items, two relationships are defined: the forward information propagation r in and the reverse information propagation r out , r in represents the transfer of information from x i to x j , while r out represents the information transfer from x j to x j under the same conditions. The connection strength in the single view is stored through the adjacency matrix A s , which records the interaction relationships between |X s | non-repeating items, and the interaction strength between two items is defined as:

[0064]

[0065] k 1 represents the result after row normalization of the adjacency matrix A s .

[0066] The heterogeneous multi-view structure is used to represent the relationships between items at multiple granularities. Different from the single view, each node in the heterogeneous multi-view consists of multiple consecutive items, which are subsequences from the same batch session; the heterogeneous multi-view where is the set of subsequences ; any one subsequence where l * represents the number of consecutive items in the current subsequence; when l * = 1, the multi-view degenerates into a single view; is the set of edges, representing the connection number of subsequences; represents a directed edge from subsequence to , and the lengths of both subsequences are l * ; the adjacency matrix stores the interaction frequency between subsequences and :

[0067]

[0068] k 2 is the result after row normalization in the adjacency matrix .

[0069] After the session sequence is constructed into a single view and a heterogeneous multi-view, the single-view attention network (SV-GAT) and the heterogeneous multi-view neural network (HMV-GNN) are used to obtain the embeddings of items and subsequences. SV-GAT is based on the transfer relationships between items in the single view and aims to learn the correlations between individual items. To achieve this goal, any item x i in the single view is mapped to the embedding space to obtain the corresponding embedding vector. The graph attention network is used to capture the association degree e ij between an item and other items:

[0070] e ij = LeakyRelu(W ij (x i ⊙x j ))

[0071] Calculating the influence degree between items using the LeakyRelu activation function, representing the element-wise product of vector sums, and is a learnable parameter.

[0072] Normalized through the Softmax function:

[0073]

[0074] Nxi is the set of 1 - order neighbors of item x i , and α ij represents the conversion relationship between an item and its neighbors; weighted summation is performed on the neighbor information to obtain the final feature representation of the item:

[0075]

[0076] The set of all item - level embedding vectors in a single view is denoted as X SV-GAT ; the item embedding vectors are updated through a multi - layer attention network, and the formula is:

[0077]

[0078] l 1 represents the number of attention layers of SV - GAT, and A s represents the normalized value storing the interaction relationship between items.

[0079] Different from a single view, each node in a heterogeneous multi - view consists of multiple consecutive items; for a subsequence of length l * , in order to learn the complete intention of the subsequence, a hybrid reading function R H is designed to generate the embedding vector m i of the heterogeneous subsequence, and the formula is as follows:

[0080]

[0081] Specifically, both set - based reading functions (such as Mean and Max) and sequence - based reading functions (such as the gated recurrent unit GRU) are considered. The former learns the intention with unchanged order, and the latter is better at capturing order - sensitive intentions. For a subsequence of length l * the final intention of the subsequence is calculated by the following formula:

[0082]

[0083] and respectively retain the order - invariant intention and order - sensitive intention of the subsequence; the set of all subsequences is denoted as M.

[0084] HMV - GNN captures high - dimensional user intention information based on the transfer relationship in the heterogeneous multi - view. To achieve this effect, a multi - head attention mechanism is adopted, which not only considers the influence of subsequences but also enhances the generalization ability of the present invention. For the attention of a single head, its calculation formula is as follows:

[0085]

[0086] head i = Attention(MW i Q , MW i K , MW i V );

[0087] M represents the input of the subsequence, W i Q , W i K and W i V are the learnable parameters of the i-th head. The operation of the multi-head attention mechanism is to concatenate the results of n single-head attention networks and normalize them through the softmax function; to form the latent vector representation of the session sequence:

[0088] M HMV-GNN = Multi - Attention(M; W att ) = [head 1 , …, head n W att

[0089] W att represents the learnable parameters of the multi-head attention network;

[0090] After passing the information to the l 2 -th layer, the final embedded vector representation is:

[0091]

[0092] M and M HMV-GNN represent the initial vector and the learned vector of all subsequences in the heterogeneous multi-view respectively.

[0093] To avoid overfitting of the training model, Dropout operation is introduced when learning two-level information:

[0094]

[0095] The single-view attention network and the heterogeneous multi-view neural network are learned based on the same batch of sessions. The corresponding item-level and subsequence-level embedded representations are extracted from and respectively. For each session, these two types of embedded representations are denoted as X SV-GAT and M HMV-GAT respectively; To better utilize the information in the subsequence, a separate session embedded representation is generated at the subsequence level; Given a session s and its corresponding subsequence unit where l *Indicates the number of consecutive items in the subsequence, and selects the last subsequence As the local representation; to capture rich comprehensive information, calculate the local representation and the context of the influence weights, and use the soft attention mechanism to calculate the attention scores to generate the global representation of the subsequence

[0096]

[0097]

[0098] and are learnable weight matrices, is the bias vector, and the embedded representation of heterogeneous multi-view content is generated through information concatenation The formula is as follows:

[0099]

[0100] || represents the concatenation operation, is a learnable parameter; the subsequence-level embedding contains the information of the global view and the local view of the subsequence;

[0101] The last item clicked by the user is often highly consistent with their current preference and is significantly helpful for the recommendation task. Therefore, through the item x l ∈X SV-GAT clicked by the user for the last time in the single-view graph, generate the final session-level embedded representation s f :

[0102]

[0103] μ is an adjustable parameter used to control the weight allocation between session information and the information of the last clicked item.

[0104] To calculate the probabilities of all candidate items, the session embedding s f is interacted with the initial item vector x i and the next item most likely to interact with the current session is selected through the softmax function

[0105] In addition, since the prediction difficulty of different sessions varies from sample to sample, a sample-adaptive loss function is proposed. All items are scored through the sample-adaptive loss function, and this loss function introduces the modulation factor SA i , and is added to the cross-entropy loss function; weights are assigned according to the prediction deviation of the samples:

[0106]

[0107] L SA =-∑(2 - 2·SA i ) τ log(SA i )

[0108] y i represents the true label of the next clicked item, τ is the temperature coefficient; the SA loss function assigns weights according to the prediction deviation of the samples; when τ decreases, the modulation factor reduces the loss contribution of simple samples; on the contrary, when τ increases, more attention is paid to high-confidence samples and their contributions.

[0109] Example 2:

[0110] In this example, three basic datasets are selected for experiments: Tmall, Gowalla, and Diginetica. These datasets are used to evaluate the recommendation performance of the present invention in different fields and scenarios. The Tmall dataset was initially released in the IJCAI-15 competition and records the product interaction information of users on the Tmall e-commerce platform. Gowalla is a real-world dataset that includes the check-in behaviors of users in different geographical locations, reflecting the spatial distribution of users' activities and interests. The Diginetica dataset is from the CIKM Cup 2016 and provides the transaction records and purchase behaviors of users on the online shopping platform.

[0111] To evaluate the performance of the present invention, a comparative experiment is conducted between the present invention and the baseline model, and P@20 and MRR@20 are used as evaluation metrics. P@20 is used to measure the proportion of products actually interacted by users among the top 20 recommended products. MRR@20 focuses on the ranking of the first relevant product in the top 20 list and calculates the average of the reciprocal of its ranking. The larger the values of P@20 and MRR@20, the higher the recommendation quality.

[0112] In this example, a single view based on a single item and a heterogeneous multi-view based on multiple consecutive items are constructed. In the heterogeneous multi-view, the number of items for each node can be dynamically adjusted to explore the impact of different sub-session sequence lengths on the recommendation performance. In all three datasets, the present invention performs excellently, and the specific experimental results are shown in Table 1 below. Table 1 verifies that the overall performance of the present invention is better than that of the baseline model, indicating that the single-view and heterogeneous multi-view modeling methods proposed by the present invention can better capture user behaviors and thus generate more accurate predictions that conform to user intentions.

[0113] Table 1: Comparison of the overall performance of the recommendation method and the baselines model

[0114]

[0115] Example 3:

[0116] To understand the impact of subsequence length on recommendation performance in the heterogeneous multi-view structure, this example conducts experiments on the subsequence length, with the value range from 1 to 7. Figure 4 The recommendation effects of different subsequence lengths on three datasets are shown. In the Tmall dataset, the recommendation performance fluctuates with the change of subsequence length, which indicates that the subsequence length is closely related to the actual intention of users. When users perform click behaviors, the dynamic change of interests will affect their click results. It is observed that when the subsequence is 6, the P@20 and MRR@20 of the Tmall dataset reach the peak. For the Gowalla dataset, the change of P@20 is not significant, but when the subsequence length is 5, the recommendation performance of MRR@20 is the worst. In the Diginetica dataset, when the subsequence length is 3, both P@20 and MRR@20 achieve the best results.

[0117] Only some exemplary embodiments of the present invention are described by way of illustration above. Undoubtedly, for those of ordinary skill in the art, the described embodiments can be modified in various different ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the protection scope of the claims of the present invention.

Claims

1. A conversation recommendation method based on heterogeneous multi-view enhanced subsequence unit learning, characterized in that: The following steps are included: S1: View construction phase, constructing the session sequence into single view and heterogeneous multi-view; S2: Embedding learning stage, using a mixed reading function to learn embeddings between views of different scales; S3: Use the fusion layer to aggregate the information from the two views and generate a session-level embedding vector; S4: Result prediction stage, all items are scored using a sample adaptive loss function, and the items with the highest scores are recommended.

2. The conversation recommendation method based on heterogeneous multi-view enhanced subsequence unit learning according to claim 1, characterized in that: The single view in step S1 is constructed based on a single session s, defined as G s =(X s ,E s ), where X s is the set of items interacted by the conversation sequence s, is a set of directed edges; each edge The interaction between items is recorded, indicating that the anonymous user clicked on item x in sequence i and x j ; Each item has a self-loop, connected by a self-edge r self Represents; defines forward information propagation r in and reverse information propagation out , r in Indicates that information is from x i to x j The transmission of r out It means that under the same conditions, from x j to x j The connection strength in a single view is obtained by the adjacency matrix A s Storage, the matrix records |X s | The interaction relationship between non-repeated items, and the interaction strength between two items Defined as: k1 represents the adjacency matrix A s The result after row normalization.

3. The conversation recommendation method based on heterogeneous multi-view enhanced subsequence unit learning according to claim 2, characterized in that: The heterogeneous multi-view structure in step S1 is used to represent the relationship between items at multiple granularities. Each node in the heterogeneous multi-view consists of multiple consecutive items, which are from subsequences in the same batch session. in is a subsequence A set of; any subsequence Among them l * Indicates the number of consecutive items in the current subsequence; when l * =1, multi-view degenerates to single view; is the set of edges, Indicates the number of connections of the subsequence; Represents a subsequence arrive The length of both subsequences is l * ; Adjacency matrix Stored subsequence and Frequency of interaction between: k2 is in the adjacency matrix The result after row normalization.

4. The conversation recommendation method based on heterogeneous multi-view enhanced subsequence unit learning according to claim 3 is characterized in that: The embedding learning phase after the session sequence is constructed into single view and heterogeneous multi-view in step S2 includes a single view attention network and a heterogeneous multi-view neural network to obtain embeddings of items and subsequences.

5. The conversation recommendation method based on heterogeneous multi-view enhanced subsequence unit learning according to claim 4, characterized in that: The single-view attention network is based on the transfer relationship between items in a single view, aiming to learn the correlation between individual items; any item x in a single view i is mapped to the embedding space to obtain the corresponding embedding vector; a graph attention network is used to capture the correlation between items and other items. ij : and ij =LeakyRelu(W ij (x i ⊙x j )); The LeakyRelu activation function is used to calculate the influence between items, x i ⊙x j Represents vector x i and x j The element-wise product of ij is a learnable parameter; Normalized by Softmax function: Is item x i The first-order neighbor set of α ij Represents the transformation relationship between an item and its neighbors; weighted summation of neighbor information is performed to obtain the final feature representation of the item: The set of all item-level embedding vectors in a single view is denoted as X SV-GAT ; The item embedding vector is updated through a multi-layer attention network, and the formula is: l1 represents the number of attention layers of SV-GAT, A s A normalized value representing the interaction relationship between stored items.

6. The conversation recommendation method based on heterogeneous multi-view enhanced subsequence unit learning according to claim 5, characterized in that: Each node in the heterogeneous multi-view neural network consists of multiple consecutive items; for a length of l * In order to learn the complete intention of the subsequence, a hybrid reading function R is designed H To generate the embedding vector m of the heterogeneous subsequence i , the formula is as follows: For a length of l * subsequence of The final intent of this subsequence is calculated by the following formula: and The order-invariant intent and order-sensitive intent of subsequences are preserved respectively; the set of all subsequences is denoted as M.

7. The conversation recommendation method based on heterogeneous multi-view enhanced subsequence unit learning according to claim 6, characterized in that: The heterogeneous multi-view neural network captures high-dimensional user intention information based on the transfer relationship in the heterogeneous multi-views and adopts a multi-head attention mechanism. The calculation formula for the attention of a single head is as follows: and is the learnable parameter of the i-th head. The operation of the multi-head attention mechanism is to concatenate the results of n single-head attention networks and normalize them through the softmax function; Form a latent vector representation of the conversation sequence: M HMV-GNN =Multi-Attention(M;W att )=[head1,…,head n ]W att W att represents the learnable parameters of the multi-head attention network; After passing the information to the l2 layer, the final embedding vector is represented as: M and M HMV-GNN They represent the initial vectors and learned vectors of all subsequences in heterogeneous multi-views respectively.

8. The conversation recommendation method based on heterogeneous multi-view enhanced subsequence unit learning according to claim 7, characterized in that: In step S3, before information aggregation, a Dropout operation is introduced when learning the single-view and multi-view two-level information: The item-level and subsequence-level embedding representations of the single-view attention network and the heterogeneous multi-view neural network are respectively and For each session, these two types of embedding representations are recorded as X SV-GAT and M HMV-GAT ; Generate separate session embedding representations for the subsequence level; Given a session s and its corresponding subsequence unit Among them l * Indicates the number of consecutive items in a subsequence, select the last subsequence As a local representation; Compute a local representation and context The influence weight of , and the soft attention mechanism is used to calculate the attention score Thus generating a global representation of the subsequence and is the learnable weight matrix, is a bias vector that generates an embedded representation of heterogeneous multi-view content through information concatenation The formula is as follows: || represents the concatenation operation. is a learnable parameter; subsequence level embedding Contains information about the global view and local view of the subsequence; By the last item clicked by the user in the single view graph x l ∈X SV-GAT Generate the final session-level embedding representation μ is an adjustable parameter used to control the weight distribution between session information and last clicked item information.

9. The conversation recommendation method based on heterogeneous multi-view enhanced subsequence unit learning according to claim 8, characterized in that: In step S4, the session embedding sf and the initial item vector x are combined. i Interact and select the next item that is most likely to interact with the current session through the softmax function All items are scored using a sample adaptive loss function, which introduces a modulation factor SA i , and added to the cross entropy loss function; weights are assigned according to the prediction bias of the sample: L SA =-∑(2-2·SA i ) τ log(IN i ) y i represents the true label of the next clicked item, and τ is the temperature coefficient; the SA loss function assigns weights according to the prediction bias of the sample; when τ decreases, the modulation factor reduces the loss contribution of simple samples; on the contrary, when τ increases, more attention is paid to high confidence samples and their contributions.

Citation Information

Patent Citations

  • User preference modeling method based on lightweight graph convolution attention network

    CN115438258A

  • Graph neural network session recommendation method based on dual-channel information fusion

    CN115470406A

  • Session recommendation enhancement method based on double composition

    CN117807281A

  • Session-based recommendation method and device

    US20220374962A1

Cited By

  • Inplanatable node classification prediction method based on adversarial causal graph learning

    CN120524163A

  • Explainable node classification prediction method based on adversarial causal diagram learning

    CN120524163B