Sequence Recommendation Method and System for Multi-Behavior and Multi-Comparison Views
By constructing a multi-behavior multi-view training set in the recommendation system, using deep learning networks to combine graphs and sequence information for data augmentation and comparison learning, the problems of user multi-behavior patterns and sequence sparseness are solved, and the effectiveness and performance of the recommendation system are improved.
Patent Information
- Application Number
- CN202310794792.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-30
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2043-06-30
AI Technical Summary
When existing recommendation systems deal with user multi-behavior patterns and sequence sparseness, it is difficult for them to make full use of graph structure information and sequence information, resulting in poor recommendation results.
Using a sequence recommendation method of multi-behavior and multi-contrast views, a multi-behavior and multi-view training set is constructed, and a deep learning network model is used to complement each other with graph information and sequence information, and data enhancement and comparison learning are carried out to improve the user's and item representation ability.
It improves the satisfaction of user recommendation results, combines graphs and sequence information, enhances the characterization ability, solves the problem of sequence sparseness, and improves the overall performance of the recommendation system.
Smart Images

Figure CN116776002B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of recommendation systems, and particularly relates to a sequential recommendation method and system with multi-behavior and multi-comparison views. Background Art
[0002] A recommendation system searches for information that meets its own needs from a vast amount of news, audio and video, commodities and other information to achieve the effect of personalized recommendation. Different from the search system, the basic idea of the search system is similar to the classification directory. The search engine more often bases on the information that the user wants to search for, and then classifies, integrates and sorts the information by the system itself, and finally presents the information most relevant to the user's search to the user. If effective keywords cannot be provided, the search engine will not be able to serve the user well. While the recommendation system provides useful information to the user according to the user's interaction history under the condition of insufficient information. Generally speaking, the recommendation system extracts the information that the user is interested in from the big data according to the user's historical interaction behavior without the user's active behavior. In some cases, it can even discover the user's potential interests, bringing greater benefits to the platform and a great sense of experience to the user at the same time.
[0003] Modern recommendation systems more often use deep learning models to model user representations. Especially in the models based on the attention mechanism are popular in various fields, because they can notice the most important relationships in the whole sequence and are thus applied to sequential recommendation. Some researchers combine the attention mechanism with CNN for recommendation, combining a powerful tool for extracting sequences with a powerful tool for extracting features. There are even models based on Transformer proposed, but these methods are all for sequences. Due to the sparsity of the sequence itself, such models often cannot make full use of the information of the sequence itself. To address the sparsity of the sequence, existing models have proposed self-supervised and contrastive learning strategies, such as adding multiple self-supervised tasks, using pre-training and fine-tuning strategies for model training, improving the representation to enhance the data to obtain contrastive loss, proposing an explanation-guided contrastive learning method, proposing contrastive learning in terms of time concept, etc. But these models only utilize the information at the sequence level to obtain the user's personal representation, while ignoring the more comprehensive representation on the graph structure. Therefore, considering how to fuse sequence information and graph information is the key problem. At the same time, with the continuous development of artificial intelligence and deep learning, the industrial community's requirements for recommendation systems are increasing day by day. Simply modeling the user behavior sequence pattern before can no longer meet the requirements. In most platforms, the user's behavior types usually have multiple forms. For the already sparse sequence, adding auxiliary behaviors will surely enhance the representation of the sequence. Therefore, the existing sequential recommendation problem needs to comprehensively consider the multi-behavior patterns of users and then recommend to users. Summary of the Invention
[0004] The purpose of the present invention is to provide a sequence recommendation method and system with multiple behaviors and multiple comparison views, which is beneficial to improving the satisfaction of user recommendation results.
[0005] To achieve the above object, the technical solution adopted by the present invention is: a sequence recommendation method with multiple behaviors and multiple comparison views, including the following steps:
[0006] Step A: Collect multi-behavior data generated by users during interaction with items, and construct a multi-behavior multi-view training set;
[0007] Step B: Use the training set to train a deep learning network model G for sequence recommendation. The deep learning network model G utilizes the globality of graph information and the individuality of sequence information to complement and enhance each other. At the same time, data augmentation is performed on the graph data for contrastive learning to learn more robust representations, and data augmentation is performed on the sequence data for contrastive learning to solve the sequence sparsity problem and further improve the representation ability of the graph and the sequence itself;
[0008] Step C: Input the user behavior data into the deep learning network model G in sequence, and output the corresponding recommendation results for the current user.
[0009] Furthermore, step B specifically includes the following steps:
[0010] Step B1: Process the training set and divide it according to different behaviors to obtain a general behavior sequence S c and the sequence data S b of specific behaviors uv,b and graph data G
[0011] Step B2: Convert the general behavior sequence data S c obtained in step B1 into general behavior graph data G uv,c , perform data augmentation operations to obtain an extended matrix M uu and M vv , and then construct corresponding behavior-independent enhanced graphs G uu and G vv ;
[0012] Step B3: Input the general behavior sequence data S c and the specific behavior sequence data S b obtained in step B1 into a sequence network for training, and at the same time perform contrastive learning operations to obtain a sequence contrast loss L SeqCL ;
[0013] Step B4: The specific behavior graph data G uv,b and the behavior-independent enhanced graph data G uu and G vv, use LightGCN for message propagation to obtain the Embeddings of users and items under each graph; then perform contrastive learning operations to obtain the graph contrast loss L GraphCL ;
[0014] Step B5: Perform attention fusion operations on the user and item Embeddings of each behavior under each graph obtained in Step B4, that is, the Embedding of each behavior after fusion in the graph view; perform concatenation fusion operations on the user and item Embeddings of each behavior in the sequence obtained in Step B3, that is, the Embedding of each behavior after fusion in the sequence view;
[0015] Step B6: Perform contrastive learning operations on the Embeddings of each behavior after fusion in the sequence view and the graph view obtained in Step B5 to obtain the cross-view contrast loss L CrossCL ;
[0016] Step B7: Train after fusing the sequence view and the graph view. When the loss value generated by the deep learning network model is less than the set threshold or reaches the maximum number of iterations, terminate the training of the deep learning model M.
[0017] Furthermore, Step B1 specifically includes the following steps:
[0018] Step B11: Generate corresponding data sets for the general behavior sequences of each user in the training set. For the nth data in the user's general behavior data, construct the first n - 1 data to predict the nth data of the user sequence. Each user generates MAX_LEN_u data, where MAX_LEN_u represents the maximum sequence length of user u;
[0019] Step B12: Clean the data obtained in Step B11. For the data generated by each user, record the last target behavior time as t, and then delete the items that appeared in the target behavior before the t behavior from the auxiliary behavior;
[0020] Step B13: Perform partitioning operations on the data obtained in Step B12 according to specific behaviors to obtain the sequence data S of each behavior b ;
[0021] Step B14: For the sequence data S obtained in Step B13 b , divide it into user nodes and item nodes according to nodes, and connect edges on the items with which the user has interacted to construct a graph matrix differentiated by specific behaviors represents the interaction matrix between user i and item j under specific behavior b, thereby constructing graph data G differentiated by specific behaviors uv,b .
[0022] Further, step B2 specifically includes the following steps:
[0023] Step B21: Convert the data obtained in step B11 into a matrix according to step B14 indicating the number of interactions between user i and item j in the user-item view without behavior differentiation, thereby constructing the graph data G of general behavior uv,c ;
[0024] Step B22: Perform matrix multiplication according to the matrix obtained in step B21 through M uu =(M uv )(M uv ) T to obtain the co-occurrence matrix of users; similarly, through M vv =(M uv ) T (M uv ) to obtain the co-occurrence matrix of items, thereby constructing the graph data G uu , G vv .
[0025] Further, step B3 specifically includes the following steps:
[0026] Step B31: Let the historical behavior sequence of a certain user u be where B represents B types of behaviors {b 1 , …, b B}; for each behavior, there is an item sequence representing the behavior sequence under the specific behavior b of user u, the target behavior sequence is represented as the general behavior sequence is represented as Input the above item sequence into the sequence network;
[0027] Step B32: Let the item Embedding and the position Embedding P∈R T×d , where T is the maximum sequence length, is the size of the item set, so given the i-th item v i , its representation is After being transformed by the Embedding layer, the sequence input in step B31 becomes the Embedding matrix of the sequence
[0028] Step B33: The item matrix obtained in step B32 is input into the Transformer layer, and through multi-head attention:
[0029]
[0030] Where X is the input Embedding matrix, W is the learnable matrix, d is the Embedding dimension, and h is the number of attention heads. In addition, Concatenate the obtained multi-head attention (MSA):
[0031] MSA(X) = Concat(head i , head 2 ,..., head h )W o
[0032] Where Concat is the concatenation operation, and W o is the learnable matrix; introduce non-linearity and perform feature transformation between MSA layers, and use the feed-forward (PFF) layer:
[0033] PFF(X) = FC(σ(FC(X))), FC(x) = XW + b
[0034] Where FC is the fully connected layer, and σ is the sigmoid activation function, Use Dropout technology, residual connection, and layer normalization to obtain the final output Embedding:
[0035] H (l) = LayerNorm(X (l-1) + MSA(X (l-1) ))
[0036] X (l) = LayerNorm(H (l) + PFF(H (l) ))
[0037] Where H (l) is the intermediate representation of the l-th layer, and X (l) is the final hidden vector of the l-th layer; combine the above attention layer operations and call it Trm(), and the sequence passes through L layers of Trm() to obtain the final output:
[0038]
[0039] Step B34: Mark the multi-layer Trm() operations passed in Step B33 as SeqEnc(), and calculate its Embedding for each behavior sequence:
[0040]
[0041] Where u s,b represents the Embedding of the user under behavior b; similarly, use SeqEnc() to calculate the general behavior sequence:
[0042]
[0043] where u s,c represents the Embedding of the user under the general behavior;
[0044] Step B35: Compare the user Embedding u s,t under the target behavior obtained in Step B34 with the user Embedding u s,c under the general behavior:
[0045]
[0046]
[0047] f(x, y, z) = log(σ(x T y - x T z))
[0048] where MLP maps and to the fully connected layers in the same space, represents the sequential Embedding of the i-th user under the target behavior, represents the sequential Embedding of the i-th user under the general behavior, represents the sequential Embedding of the j-th user under the general behavior, σ is the sigmoid activation function, and N is the number of users.
[0049] Furthermore, Step B4 specifically includes the following steps:
[0050] Step B41: Use the data G uv,b , G uu , G vv obtained in Steps B14 and B22 for message propagation:
[0051]
[0052] where A is the adjacency matrix, D ii = ∑ j=0 A ij is the diagonal matrix, X is the node Embedding, and the user-item view node Embedding after propagation is denoted as u uv,b , v uv,b , the user-user view node Embedding is denoted as u uu , and the item-item view node Embedding is denoted as v vv, particularly, the user-item view node Embedding under the target behavior is denoted as u uv,t , v uv,t ;
[0053] Step B42: Use the node Embeddings of the three views obtained in Step B41 to perform user comparison:
[0054]
[0055]
[0056] Among them, represents the user Embedding of the i-th user under the target behavior of the user-item view, represents the user Embedding of the i-th user under the user-user view, represents the user Embedding of the j-th user under the user-user view, and N is the number of users;
[0057] Step B43: Use the node Embeddings of the three views obtained in Step B41 to perform item comparison:
[0058]
[0059]
[0060] Among them, represents the item Embedding of the i-th item under the target behavior of the user-item view, represents the item Embedding of the i-th item under the item-item view, represents the item Embedding of the j-th item under the item-item view;
[0061] Step B44: Add the losses in Steps B42 and B43 to obtain the final graph comparison loss:
[0062] L GraphCL = L GraphUCL + L GraphICL .
[0063] Furthermore, Step B5 specifically includes the following steps:
[0064] Step B51: Fuse the behavior Embeddings obtained in Step B34:
[0065]
[0066] Among them, u sIt is the fused Embedding of the sequence view, and "||" is the concatenation operation;
[0067] Step B52: Perform a concatenation operation on the Embeddings of each row obtained in Step B41:
[0068]
[0069] where u g-t represents the user Embedding under other auxiliary behaviors except the target behavior, and u 0 represents the original Embedding of the user-item view, and ";" represents Embedding stacking;
[0070] Step B53: Perform an attention operation on the result of Step B52 using the user Embedding under the target behavior as in Step B33:
[0071] u uv,att = ATT(u uv,t , u g-t , u g-t )
[0072] where ATT represents the attention operation, and u uv,att is the output Embedding of the target behavior's attention to the auxiliary behavior;
[0073] Step B54: Concatenate the original vector of the user-item view and the attention output vector as the final fused vector:
[0074] u g = MLP g (u 0 | u uv,att )
[0075] where u g is the fused Embedding of the graph view.
[0076] Furthermore, in Step B6, a comparison operation is performed on the different view Embeddings obtained in Step B51 and Step B54:
[0077]
[0078]
[0079] where, represents the sequence fused Embedding of the i-th user, represents the graph fused Embedding of the i-th user, represents the graph fused Embedding of the j-th user.
[0080] Further, step B7 specifically includes the following steps:
[0081] Step B71: Perform a splicing and fusion operation on two views:
[0082] u = MLP U (u s |u g ), v = MLP V (v s |v g )
[0083] where u s , u g , v s , v g respectively represent the user Embedding of the sequence and the graph, and the item Embedding of the sequence and the graph;
[0084] Step B72: Calculate the loss value using the BPR loss:
[0085]
[0086] where u represents the user vector, v i represents the positive sample of the user, and v j represents the negative sample of the user;
[0087] Step B73: Update the learning rate through the gradient optimization algorithm Adam, and use backpropagation to iteratively update the model parameters to train the model by minimizing the loss function; the total loss of the model is the weighted sum of the above losses:
[0088] L = λ o L o + λ 1 L SeaCL + λ 2 L GraphCL + λ 3 L crossCL
[0089] where λ o , λ 1 , λ 2 , λ 3 are hyperparameters, L o is the main loss, L SeqCL is the sequence contrast loss, L GraphCL is the graph contrast loss, L CrossCL is the cross-view contrast loss.
[0090] The present invention also provides a multi-behavior multi-contrast view sequence recommendation system adopting the above method, including:
[0091] A training set construction module, which is used to collect multi-behavior data generated by users during the interaction with items and construct a multi-behavior and multi-view training set;
[0092] A model training module, which is used to train a deep learning network model G for sequential recommendation and output the trained deep learning network model G to the model prediction module; and
[0093] A model prediction module, which is used to input user behavior data into the deep learning network model G and output the corresponding recommendation results for the current user.
[0094] Compared with the prior art, the present invention has the following beneficial effects: for the problem of sparse user behavior sequences, it is proposed to use the globality of graph information and the individuality of sequence information to complement and enhance each other. In addition, the present invention also performs data augmentation operations on graph data for contrastive learning to learn more robust representations. At the same time, data augmentation operations are also performed on sequence data for contrastive learning to solve the problem of sequence sparsity from another aspect and further improve the representation capabilities of the graph and the sequence itself. The present invention also starts from reality and uses auxiliary behavior sequences to enhance sequence representations. Finally, the present invention also performs contrastive learning on the vector representations obtained from the graph view and the sequence view to further enhance the representation results of users and items. Description of the Drawings
[0095] Figure 1 is a flowchart of the method implementation of the embodiment of the present invention;
[0096] Figure 2 is an architecture diagram of the deep learning network model in the embodiment of the present invention;
[0097] Figure 3 is a schematic diagram of the system structure of the embodiment of the present invention. Detailed Embodiments
[0098] The present invention will be further described below in conjunction with the drawings and embodiments.
[0099] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further descriptions of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0100] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0101] As Figure 1 shown, this embodiment provides a sequence recommendation method for multi-behavior and multi-comparison views, including the following steps:
[0102] Step A: Collect multi-behavior data generated by users during interaction with items, and construct a multi-behavior and multi-view training set.
[0103] Step B: Use the training set to train the deep learning network model G for sequence recommendation. The deep learning network model G utilizes the globality of graph information and the individuality of sequence information to complement and enhance each other. At the same time, data augmentation is performed on the graph data for contrastive learning to learn more robust representations, and data augmentation is performed on the sequence data for contrastive learning to solve the sequence sparsity problem and further improve the representation ability of the graph and the sequence itself. The architecture of the deep learning network model G in this embodiment is as Figure 2 shown.
[0104] Step C: Input the user behavior data into the deep learning network model G in sequence, and output the corresponding recommendation results for the current user.
[0105] In this embodiment, the specific steps of step B include the following steps:
[0106] Step B1: Process the training set and divide it according to different behaviors to obtain the general behavior sequence S c and the sequence data S b of specific behaviors uv,b and the graph data G
[0107] In this embodiment, the specific steps of step B1 include the following steps:
[0108] Step B11: Generate corresponding data sets for the general behavior sequences of each user in the training set. For the nth data in the user's general behavior data, construct a user sequence that uses the first n - 1 data to predict the nth data. Each user generates MAX_LEN_u pieces of data, where MAX_LEN_u represents the maximum sequence length of user u.
[0109] Step B12: Perform a cleaning operation on the data obtained in step B11. For the data generated by each user, record the last target behavior time as t, and then delete the items that appeared in the target behavior before the t behavior from the auxiliary behaviors.
[0110] Step B13: Perform an operation of dividing the data obtained in step B12 according to specific behaviors to obtain the sequence data S b under each behavior.
[0111] Step B14: The sequence data S obtained in step B13b , divided into user nodes and item nodes according to nodes, and edges are connected on the items with which the user has interacted to construct a graph matrix distinguished by specific behaviors represents the interaction matrix between user i and item j under specific behavior b, thereby constructing the graph data G distinguished by specific behaviors uv,b .
[0112] Step B2: Convert the general behavior sequence data S obtained in Step B1 c into general behavior graph data G uv,c , perform data augmentation operations to obtain the extended matrix M uu and M vv , and then construct the corresponding enhanced graph G without behavior distinction uu and G vv .
[0113] In this embodiment, the specific steps of Step B2 include the following steps:
[0114] Step B21: Convert the data obtained in Step B11 into a matrix according to Step B14 represents the number of interactions between user i and item j in the user-item view without behavior distinction, thereby constructing the graph data G of general behavior uv,c ;
[0115] Step B22: Perform matrix multiplication operations according to the matrix obtained in Step B21 through M uu =(M uv )(M uv ) T to obtain the co-occurrence matrix of users; similarly, through M vv =(M uv )(M T (M uv ) to obtain the co-occurrence matrix of items, thereby constructing the graph data G without behavior distinction uu , G vv .
[0116] Step B3: Input the general behavior sequence data S obtained in Step B1 c and the specific behavior sequence data S b into the sequence network for training, and at the same time perform contrastive learning operations to obtain the sequence contrast loss L SeqCL .
[0117] In this embodiment, the specific steps of Step B3 include the following steps:
[0118] Step B31: Assume that the historical behavior sequence of a certain user u is where B represents B types of behaviors {b 1 , …, bB}; For each behavior, there is an item sequence denotes the behavior sequence under the specific behavior b of user u, and the target behavior sequence is denoted as The general behavior sequence is denoted as Input the above item sequence into the sequence network.
[0119] Step B32: Let the item Embedding and the position Embedding P ∈ R T×d , where T is the maximum sequence length, is the size of the item set. Thus, given the i-th item v i , its representation is After being transformed by the Embedding layer, the sequence input in Step B31 becomes the Embedding matrix of the sequence
[0120] Step B33: The item matrix obtained in Step B32 is input into the Transformer layer and undergoes multi-head attention:
[0121]
[0122] where X is the input Embedding matrix, W is the learnable matrix, d is the Embedding dimension, h is the number of attention heads. In addition, Concatenate the obtained multi-head attention (MSA):
[0123] MSA(X) = Concat(head i , head 2 ,..., head h )W o
[0124] where Concat is the concatenation operation, W o is the learnable matrix; Introduce non-linearity and perform feature transformation between MSA layers, using the feed-forward (PFF) layer:
[0125] PFF(X) = FC(σ(FC(X))), FC(x) = XW + b
[0126] where FC is the fully connected layer, σ is the sigmoid activation function, Use Dropout technology, residual connection, and layer normalization to obtain the final output Embedding:
[0127] H (l) = LayerNorm(X (l-1) + MSA(X (l-1)))
[0128] X (l) = LayerNorm(H (l) + PFF(H (l) ))
[0129] where H (l) is the intermediate representation of the l-th layer, and X (l) is the final hidden vector of the l-th layer; combining the above attention layer operations is called Trm(), and the sequence passes through L layers of Trm() to obtain the final output:
[0130]
[0131] Step B34: Mark the multi-layer Trm() operations passed in Step B33 as SeqEnc(), and calculate the Embedding for each behavior sequence:
[0132]
[0133] where u s,b represents the Embedding of the user under behavior b; similarly, use SeqEnc() to calculate the general behavior sequence:
[0134]
[0135] where u s,c represents the Embedding of the user under the general behavior.
[0136] Step B35: Compare the user Embedding u s,t under the target behavior obtained in Step B34 with the user Embedding u s,c under the general behavior:
[0137]
[0138]
[0139] f(x, y, z) = log(σ(x T y - x T z))
[0140] where MLP maps and to fully connected layers in the same space, represents the sequence Embedding of the i-th user under the target behavior, represents the sequence Embedding of the i-th user under the general behavior, It represents the sequence Embedding of the j-th user under the general behavior, σ is the sigmoid activation function, and N is the number of users.
[0141] Step B4: Use the graph data G of the specific behavior obtained in steps B2 and B1 uv,b and the enhanced graph data G without behavior distinction uu and G vv , perform message propagation using LightGCN to obtain the Embedding of users and items under each graph; then perform contrastive learning operations to obtain the graph contrast loss L GraphCL .
[0142] In this embodiment, the specific steps of step B4 include the following steps:
[0143] Step B41: Use the data G obtained in steps B14 and B22 uv,b , G uu , G vv to perform message propagation:
[0144]
[0145] where A is the adjacency matrix, D ii =∑ j=0 A ij is the diagonal matrix, X is the node Embedding, and the user-item view node Embedding after propagation is denoted as u uv,b , v uv,b , the user-user view node Embedding is denoted as u uu , and the item-item view node Embedding is denoted as v vv , in particular, the user-item view node Embedding under the target behavior is denoted as u uv,t , v uv,t .
[0146] Step B42: Use the node Embedding of the three views obtained in step B41 to perform user comparison:
[0147]
[0148]
[0149] where, represents the user Embedding of the i-th user under the target behavior in the user-item view, represents the user Embedding of the i-th user in the user-user view, denotes the user Embedding of the j-th user in the user-user view, and N is the number of users.
[0150] Step B43: Use the node Embeddings of the three views obtained in Step B41 to perform item comparison:
[0151]
[0152]
[0153] Among them, denotes the item Embedding of the i-th item under the target behavior in the user-item view, denotes the item Embedding of the i-th item in the item-item view, denotes the item Embedding of the j-th item in the item-item view.
[0154] Step B44: Add the losses obtained in Steps B42 and B43 to obtain the final graph comparison loss:
[0155] L GraphCL = L GraphUCL + L GraphICL .
[0156] Step B5: Perform attention fusion operation on the user and item Embeddings of each behavior under each graph obtained in Step B4, that is, the Embedding after fusion of each behavior in the graph view; perform splicing fusion operation on the user and item Embeddings of each behavior under the sequence obtained in Step B3, that is, the Embedding after fusion of each behavior in the sequence view.
[0157] In this embodiment, Step B5 specifically includes the following steps:
[0158] Step B51: Fuse the Embeddings of each behavior obtained in Step B34:
[0159]
[0160] Among them, u s is the fused Embedding in the sequence view, and "||" is the splicing operation.
[0161] Step B52: Perform splicing operation on the Embeddings of each behavior obtained in Step B41:
[0162]
[0163] Among them, u g-tDenote the user Embedding under other auxiliary behaviors except the target behavior as u 0 Denote the original Embedding of the user-item view. ";" indicates Embedding stacking.
[0164] Step B53: Perform the attention operation on the result of step B52 using the user Embedding under the target behavior as in step B33:
[0165] u uv,att = ATT(u uv,t , u g-t , u g-t )
[0166] where ATT represents the attention operation, and u uv,att is the output Embedding of the attention of the target behavior to the auxiliary behavior.
[0167] Step B54: Concatenate the original vector of the user-item view and the attention output vector as the final fusion vector:
[0168] u g = MLP g (u 0 | u uv,att )
[0169] where u g is the fusion Embedding of the graph view.
[0170] Step B6: Perform contrastive learning operation on the Embeddings of each behavior of the sequence view and the graph view obtained in step B5 to obtain the cross-view contrastive loss L CrossCL .
[0171] In the said step B6, perform contrastive operation on the Embeddings of different views obtained in step B51 and step B54:
[0172]
[0173]
[0174] where represents the sequence fusion Embedding of the i-th user, represents the graph fusion Embedding of the i-th user, represents the graph fusion Embedding of the j-th user.
[0175] Step B7: After fusing the sequence view and the graph view, perform training. When the loss value generated by the deep learning network model is less than the set threshold or reaches the maximum number of iterations, terminate the training of the deep learning model M.
[0176] In this embodiment, step B7 specifically includes the following steps:
[0177] Step B71: Perform a splicing and fusion operation on the two views:
[0178] u = MLP U (u s |u g ),v = MLp V (v s |v g )
[0179] where u s ,u g ,v s ,v g represent the user Embedding of the sequence and the graph, and the item Embedding of the sequence and the graph respectively;
[0180] Step B72: Calculate the loss value using the BPR loss:
[0181]
[0182] where u represents the user vector, v i represents the positive sample of the user, and v j represents the negative sample of the user;
[0183] Step B73: Update the learning rate through the gradient optimization algorithm Adam, and use backpropagation to iteratively update the model parameters to train the model by minimizing the loss function; the total loss of the model is the weighted sum of the above losses:
[0184] L = λ o L o + λ 1 L SeaCL + λ 2 L GraphCL + λ 3 L crossCL
[0185] where λ o ,λ 1 ,λ 2 ,λ 3 are hyperparameters, L o is the main loss, L SeqCL is the sequence contrast loss, L GraphCL is the graph contrast loss, L CrossCL is the cross-view contrast loss.
[0186] As shown Figure 3 in the figure, this embodiment also provides a sequence recommendation system with multi-behavior and multi-comparison views adopting the above method, including: a training set construction module, a model training module, and a model prediction module.
[0187] The training set construction module is used to collect multi-behavior data generated by users during their interaction with items and construct a multi-behavior and multi-view training set.
[0188] The model training module is used to train a deep learning network model G for sequence recommendation and output the trained deep learning network model G to the model prediction module.
[0189] The model prediction module is used to input user behavior data into the deep learning network model G and output the corresponding recommendation results for the current user.
[0190] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0191] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0192] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including instruction means, and the instruction means implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0193] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to generate a computer-implemented process, thereby providing instructions for implementing the steps of the process Figure 1 one process or a plurality of processes and / or boxes Figure 1 steps of the functions specified in one box or a plurality of boxes.
[0194] As mentioned above, it is only the preferred embodiment of the present invention, and it is not intended to limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications to equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A sequential recommendation method for multi-behavior and multi-comparison views, characterized in that, it includes the following steps: Step A: Collect multi-behavior data generated by users during their interaction with items, and construct a multi-behavior and multi-view training set; Step B: Use the training set to train a deep learning network model G for sequential recommendation. The deep learning network model G utilizes the globality of graph information and the individuality of sequential information to complement and enhance each other. At the same time, data augmentation is performed on the graph data for contrastive learning to learn more robust representations, and data augmentation is performed on the sequential data for contrastive learning to solve the problem of sequential sparsity and further improve the representation ability of the graph and the sequence itself; Step C: Input the user behavior data into the deep learning network model G in sequence, and output the corresponding recommendation results for the current user; The specific steps of step B include the following steps: Step B1: After processing the training set, divide it according to different behaviors to obtain the general behavior sequence S c and the sequence data S of specific behaviors b and the graph data G uv,b ; Step B2: Convert the general behavior sequence data S obtained in Step B1 c into general behavior graph data G uv,c , perform data augmentation operations to obtain an extended matrix M uu and M vv , and then construct corresponding augmented graphs G without behavior distinction uu and G vv ; Step B3: Input the general behavior sequence data S c obtained in Step B1 b and the specific behavior sequence data S SeqCL into the sequence network for training, and at the same time perform contrastive learning operations to obtain the sequence contrast loss L Step B4: Use the graph data G of specific behaviors obtained in Steps B2 and B1 uv,b and the enhanced graph data G without behavior distinction uu and G vv , perform message propagation using LightGCN to obtain the Embeddings of users and items under each graph; then perform contrastive learning operations to obtain the graph contrastive loss L GraphCL ; Step B5: Perform attention fusion operations on the user and item Embeddings of each behavior under each graph obtained in step B4, that is, the Embedding after the fusion of each behavior in the graph view; perform splicing fusion operations on the user and item Embeddings of each behavior under the sequence obtained in step B3, that is, the Embedding after the fusion of each behavior in the sequence view; Step B6: Perform contrastive learning on the Embeddings obtained by fusing each row of the sequence view and the graph view obtained in Step B5 to obtain a cross-view contrastive loss L CrossCL ; Step B7: Train after fusing the sequence view and the graph view. When the loss value generated by the deep learning network model is less than the set threshold or reaches the maximum number of iterations, terminate the training of the deep learning model M; The specific steps of step B1 include the following steps: Step B11: Generate corresponding data sets for the general behavior sequences of each user in the training set. For the nth piece of data in the user's general behavior data, construct the first n - 1 pieces of data to predict the nth piece of data in the user sequence. Each user generates MAX_LEN_u pieces of data, where MAX_LEN_u represents the maximum sequence length of user u; Step B12: Perform cleaning operations on the data obtained in step B11. For the data generated by each user, record the last target behavior time as t, and then delete the items that appeared in the target behavior before the t behavior from the auxiliary behaviors; Step B13: Divide the data obtained in Step B12 according to specific behaviors to obtain the sequence data S under each behavior b ; Step B14: For the sequence data S obtained in step B13 b , divide it into user nodes and item nodes according to nodes, and connect edges on the items with which the user has interacted to construct a graph matrix for distinguishing specific behaviors represents the interaction matrix between user i and item j under specific behavior b, so as to construct graph data G for distinguishing specific behaviors uv,b ; The specific steps of step B2 include the following steps: Step B21: Convert the data obtained in Step B11 into a matrix according to Step B14 represents the number of interactions between user i and item j in the user-item view without behavior differentiation, thereby constructing the graph data G of general behavior uv,c ; Step B22: According to the matrix obtained in Step B21 perform matrix multiplication operations, through M uu =(M uv )(M uv ) T obtain the co-occurrence matrix of user and user; similarly, through M vv =(M uv ) T (M uv ) obtain the co-occurrence matrix of item and item, thereby constructing the graph data G uu , G vv ; The specific steps of step B3 include the following steps: Step B31: Let the historical behavior sequence of a certain user u be where B represents B types of behaviors {b 1 , …, b B}; for each behavior, there is an item sequence representing the behavior sequence under the specific behavior b of user u, and the target behavior sequence is represented as The general behavior sequence is represented as Input the above item sequences into the sequence network; Step B32: Set item Embedding and position Embedding P ∈ R T×d , where T is the maximum sequence length, is the size of the item set, so for the i-th item v i , its representation is After being transformed by the Embedding layer, the sequence input in Step B31 becomes the Embedding matrix of the sequence Step B33: The item matrix obtained in step B32 is input into the Transformer layer, and through multi-head attention: where X is the input Embedding matrix, W is the learnable matrix, d is the Embedding dimension, and h is the number of attention heads. In addition, concatenate the obtained multi-head attention MSA: MSA(X) = Concat(head i , head 2 , …, head h )W o where Concat is the concatenation operation, and W o is a learnable matrix; introducing non-linearity and performing feature transformation between MSA layers, using a feed-forward PFF layer: PFF(X) = FC(σ(FC(X))), FC(X) = XW + b where FC is the fully connected layer, σ is the sigmoid activation function, W ∈ R d×d , Using Dropout technology, residual connections, and layer normalization, the final output Embedding is obtained: H (l) = LayerNorm(X (l-1) + MSA(X (l-1 )) X (l) = LayerNorm(H (l) + PFF(H (l )) Among them, H (l) as the intermediate representation of the l-th layer, X (l) is the final hidden vector of the l-th layer; combining the above attention layer operations is called Trm(), and the sequence passes through L layers of Trm() to obtain the final output: Step B34: Mark the multi-layer Trm() operations passed in step B33 as SeqEnc(), and calculate its Embedding for each behavior sequence: Among them, u s,b represents the Embedding of the user under behavior b; similarly, SeqEnc() is used to calculate the general behavior sequence: where u s,c represents the Embedding of the user under general behaviors; Step B35: Compare the user Embedding u under the target behavior obtained in Step B34 s,t with the user Embedding u under the general behavior s,c for feature comparison: f(x,y,z) = log(σ(x T y - x T z)) where the MLP maps and to a fully connected layer in the same space, represents the sequence Embedding of the i-th user under the target behavior, represents the sequence Embedding of the i-th user under the general behavior, represents the sequence Embedding of the j-th user under the general behavior, σ is the sigmoid activation function, and N is the number of users; The specific steps of step B4 include the following steps: Step B41: Using the data G obtained in Steps B14 and B22 uv,b , G uu , G vv Perform message propagation: Among them, A is the adjacency matrix, and D ii = ∑ j=0 A ij is the diagonal matrix, X is the node Embedding, and the user-item view node Embedding after propagation is denoted as u uv,b , v uv,b , the user-user view node Embedding is denoted as u uu , and the item-item view node Embedding is denoted as v vv , the user-item view node Embedding under a specific behavior is denoted as u uv,t , v uv,t ; Step B42: Perform user comparison using the node Embeddings of the three views obtained in step B41; Among them, represents the user Embedding of the i-th user under the target behavior in the user-item view, represents the user Embedding of the i-th user in the user-user view, represents the user Embedding of the j-th user in the user-user view, and N is the number of users; Step B43: Perform item comparison using the node Embeddings of the three views obtained in step B41; Among them, represents the item Embedding of the i-th item under the target behavior in the user-item view, represents the item Embedding of the i-th item in the item-item view, represents the item Embedding of the j-th item in the item-item view; Step B44: Add the losses in steps B42 and B43 to obtain the final graph comparison loss; L GraphCL = L GraphUCL + L GraphICL ; The specific steps of step B5 include the following steps: Step B51: Fuse the Embeddings of each row obtained in Step B34: where u s is the fused Embedding of the sequence view, and "||" is the concatenation operation; Step B52: Perform a concatenation operation on the Embeddings of each row obtained in Step B41: where u g-t represents the user Embedding under other auxiliary behaviors except the target behavior, and u 0 represents the original Embedding of the user-item view, and ";" represents Embedding stacking; Step B53: Perform an attention operation on the result of Step B52 using the user Embedding under the target behavior as in Step B33: u uv,att = ATT(u uv,t , u g-t , u g-t ) Among them, ATT represents the attention operation, and u uv,att is the output Embedding of the target behavior's attention to the auxiliary behavior; Step B54: Concatenate the original vector of the user-item view and the attention output vector as the final fused vector: u g = MLP g (u 0 |u uv,att ) Among them, u g is the fused Embedding of the view view; In Step B6, perform a comparison operation on the different view Embeddings obtained in Step B51 and Step B54: Among them, represents the sequence fusion Embedding of the i-th user, represents the graph fusion Embedding of the i-th user, represents the graph fusion Embedding of the j-th user; Step B7 specifically includes the following steps: Step B71: Perform a concatenation and fusion operation on the two views: u = MLP U (u s |u g ), v = MLP V (v s |v g ) Among them, u s , u g , v s , v g respectively represent the user Embedding of the sequence and the graph, and the item Embedding of the sequence and the graph; Step B72: Calculate the loss value using the BPR loss: Among them, u represents the user vector, and v i represents the positive sample of the user, and v j represents the negative sample of the user; Step B73: Update the learning rate through the gradient optimization algorithm Adam, and use backpropagation to iteratively update the model parameters to train the model by minimizing the loss function; the total loss of the model is the weighted sum of the above losses: L = λ o L o + λ 1 L SeqCL + λ 2 L GraphCL + λ 3 L CrossCL Among them, λ o , λ 1 , λ 2 , λ 3 are hyperparameters, L o is the main loss, L SeqCL is the sequence contrast loss, L GraphCL is the graph contrast loss, L CrossCL is the cross-view contrast loss.
2. A multi-behavior multi-contrast view sequence recommendation system using the method as described in Claim 1, characterized in that it includes: A training set construction module for collecting multi-behavior data generated by users during interactions with items and constructing a multi-behavior multi-view training set; A model training module for training a deep learning network model G for sequence recommendation and outputting the trained deep learning network model G to the model prediction module; and A model prediction module for inputting user behavior data into the deep learning network model G and outputting the corresponding recommendation results for the current user.
Citation Information
Patent Citations
Serialization recommendation method based on multi-task learning
CN114168845A
Customized automatic recommendation method and system according to travel itinerary
KR1020220138774A