Sequence recommendation model based on bidirectional simplified gated loop network and linear attention
By introducing a bidirectional simplified gating recurrent network, additive attention module and hybrid expert module into the sequence recommendation model, the problem of difficulty in capturing the long-term preferences of user behavior and periodic dynamic preferences is solved, and efficient sequence modeling and recommendation performance improvements are achieved.
Patent Information
- Application Number
- CN202510241697.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-20
AI Technical Summary
Existing sequence recommendation models are difficult to effectively capture long-term preferences and periodic dynamic preferences in user behavior, and are highly computationally complex and difficult to deal with long sequences.
A sequence recommendation model based on bidirectional simplified gating recurrent network and linear attention is proposed. Short-term dynamic preference is captured through bidirectional SGRN module, additive attention module reduces computational complexity and models long-term preferences, and hybrid expert module deals with the diversity of user behavior.
Effectively capture the front-and-back dependencies of user behavior, reduce the computational complexity, improve the modeling ability of users' long-term and short-term preferences, and improve the generalization ability of the model and the capture ability of diversified features.
Smart Images

Figure CN120179896A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sequential recommendation, in particular to the technical field of a sequential recommendation model based on a bidirectional simplified gated recurrent network and linear attention. Background Art
[0002] In the past decade, sequential recommendation models (SRs) have shown great potential and value in multiple fields such as content streaming platforms and e-commerce; SRs build models based on the behavior sequences of users (such as the order of purchased goods) to capture the dynamic changes in user preferences and then predict the next behavior of users; the core of sequential recommendation is user behavior modeling technology, which has undergone a series of changes over time; early matrix factorization methods, such as BPR-MF, mainly made recommendations based on the binary interactions between users and items, but they could not capture the temporality in the interactions; sequential recommendation methods based on RNNs and CNNs, such as GRU4Rec, NARM, and Caser, introduced gating mechanisms or convolutional operations to enhance the sequential modeling effect. However, such methods are prone to the problems of gradient vanishing and gradient explosion, have limited ability to model long-distance dependencies, and are insufficient in capturing complex behaviors and diverse preferences; in recent years, Transformer-based sequential modeling methods, such as SASRec and BERT4Rec, have shown good long-distance dependency modeling ability and parallel computing ability; however, the self-attention mechanism adopted by such methods has a quadratic computational complexity and is too computationally expensive when dealing with long sequences; sequential modeling methods based on linear attention mechanisms, such as LinRec and LightSAN, effectively reduce the computational cost; but since the linear attention mechanism does not depend on the time order, the ability to capture users' periodic dynamic preferences is slightly insufficient; Mamba2Rec introduces a Mamba block based on a variant of the selective state space model and uses a structured state tensor to solve the long-distance dependency problem, and is superior to the existing attention mechanisms in terms of performance; MaTrRec processes the dynamic preferences in the sequence by fusing Transformer and Mamba, enhancing the modeling ability of the model on both long and short sequences while ensuring the model efficiency; but this method does not consider the periodic patterns in user behaviors and is slightly insufficient in capturing users' dynamic interests.
[0003] With the development of sequence modeling methods, bidirectional sequence modeling methods have been proposed; Bi-LSTM combines two long short-term memory networks, one processes data from the forward direction of the sequence, and the other processes data from the reverse direction of the sequence, improving the accuracy of sequence modeling; EchoMamba4Rec enhances the model's overall understanding of user behavior by simultaneously processing the past (previous items) and future (subsequent items) of the sequence; SIGMA constructs a bidirectional architecture by introducing partially flipped Mamba blocks to enhance the model's modeling ability; the application of the gating mechanism in sequence recommendation models can effectively improve the ability to capture users' dynamic preferences, such as GRU4Rec and NARM ] The attention to historical behavior is dynamically adjusted through a gating unit, which can better capture users' short-term preferences; however, this type of method is difficult to reflect users' long-term global preferences; LRURec introduces a linear recurrent unit (LRU) to improve the modeling efficiency of the model, but due to the introduction of sequence padding, some long-term dependence information may be weakened, thus affecting the modeling performance; RecBLR introduces a behavior-dependent linear recurrent unit (BD-LRU), which can capture long-distance dependence relationships to a certain extent, but due to the lack of an explicit noise filtering mechanism, noise information may be amplified during the loop process, thus affecting the accurate capture of real user preferences. Summary of the Invention
[0004] The purpose of the present invention is to solve the problems in the prior art and propose a sequence recommendation model based on a bidirectional simplified gated recurrent network and linear attention, which can solve the above problems.
[0005] To achieve the above purpose, the present invention proposes a sequence recommendation model based on a bidirectional simplified gated recurrent network and linear attention, including an embedding layer, a bidirectional simplified gated recurrent network module, an additive attention module, a mixture of experts module, and a prediction layer;
[0006] The bidirectional simplified gated recurrent network module simultaneously captures the forward and backward dependence relationships in the user behavior sequence and is used for modeling users' short-term dynamic preferences;
[0007] The additive attention module reduces the computational complexity of attention weights from quadratic to linear and is used for modeling users' long-term preferences;
[0008] The mixture of experts module adaptively selects appropriate experts to participate in the calculation according to the input features and is used for dealing with the diversity of user behavior and complex behavior modeling in specific scenarios.
[0009] Preferably, the working steps of the bidirectional simplified gated recurrent network module are as follows:
[0010] Adopt a simplified gated recurrent network as the core unit to capture the dynamic changes of user preferences:
[0011] For a given input sequence x t , the simplified gated recurrent network first decomposes it into real and imaginary parts,
[0012] The decomposition formula is as follows:
[0013] Re(c t ) = SiLU(x t W r + b r ) (2)
[0014] Im(c t ) = SiLU(x t W i + b i ) (3)
[0015] where x t represents the input vector at time step t, c t represents the complex-domain representation of the current input, Re(c t ) and Im(c t ) represent the real and imaginary parts of c t respectively, and the complex domain extends the hidden state using the real and imaginary parts; W r and W i represent the weight matrices corresponding to the real and imaginary parts respectively, b r and b i represent the bias vectors corresponding to the real and imaginary parts respectively, and SiLU(x) = x + sigmoid(x) is a non-linear activation function used to enhance the model's expressive power;
[0016] Next, the hidden state is updated through the forget gate, and the formula is as follows:
[0017] μ t = Sigmoid(x t W μ + b μ ) (4)
[0018] h t = μ t exp(iθ)·h t-1 + (1 - μ t )·c t (5)
[0019] where μ t represents the forget gate, h t-1 represents the hidden state of the previous time step, and the rotation angle exp(iθ)
[0020] introduces a phase change;
[0021] Finally, the corresponding output vector is obtained through the gating mechanism and projection transformation, and the formula is as follows:
[0022] g t = τ(W g x t + b g )(6)
[0023] o′ t = LayerNorm(g t ⊙ [Re(h t ), Im(h t )] (7)
[0024] o t = o′ t W o + b o (8)
[0025] Among them, g t represents the gating vector, which is used to control the real and imaginary parts of the output. τ(·) represents the Sigmoid activation function. W o and b o represent the projection matrix and bias vector respectively;
[0026] The bidirectional SGRN module is introduced to model the short-term dynamic preferences of users from both positive and negative directions. The sequence modeling processes in the forward and reverse orders are as follows:
[0027]
[0028]
[0029] Among them, H0 and represent the embedding matrices corresponding to the user interaction sequence S u and the reverse interaction sequence respectively. S0 and represent the modeling results of the forward and reverse sequences respectively. Here, we design a gating convolutional layer to fuse the modeling results of the forward and reverse sequences;
[0030] First, the local features of the sequence are extracted through one-dimensional convolution, and the formula is as follows:
[0031] G0 = Conv1d(H0W c + b c )(11)
[0032] Among them, W δ and b c represent the projection matrix and bias vector of the convolution operation respectively;
[0033] Then, feature selection is performed through a dual-channel architecture. The left channel acts as a gating mechanism, and the execution process is as follows:
[0034] δ0(G0) = G0W δ + b δ (12)
[0035] L(G0) = Sigmoid(δ0(G0)) (13)
[0036] Among them, W δ and b δ respectively represent the projection matrix and the bias vector, and the Sigmoid activation function acts as a forgetting gate; the right channel directly models the input features using a non-linear activation function to increase the expressive power of the features. The formula is as follows:
[0037] R(G0) = SiLU(G0) (14)
[0038] Next, the modeling results of the left and right channels are fused. The formula is as follows:
[0039]
[0040] Among them, ⊙ represents the dot product operation between vectors, that is, the corresponding elements of two vectors are multiplied;
[0041] Finally, the fusion of the forward and backward sequence modeling results is achieved through GCL. The formula is as follows:
[0042]
[0043] Among them, and respectively represent the outputs obtained by the forward and backward sequences through GCL, and serve as the weight vectors of the forward and backward sequence modeling results in the fusion process.
[0044] Preferably, the working steps of the additive attention module are as follows:
[0045] First, the input vector is linearly transformed to generate low-rank representations of the query matrix and the key matrix. The formula is as follows:
[0046]
[0047]
[0048] Among them, H0 represents the output of the embedding layer, and respectively represent the low-rank transformation matrices of the query and the key; to balance the computational complexity and the information extraction effect, we introduce a learnable weight vector W aCalculate the attention distribution and normalize the result to obtain the attention weights. The formula is as follows:
[0049]
[0050] Next, based on the attention weights, perform a weighted sum on each element of the query matrix to obtain the global query vector q. The formula is as follows:
[0051]
[0052] Finally, directly multiply the global query vector by the key matrix and then perform a linear combination with the query vector to obtain the sequence modeling result. The formula is as follows: The formula is as follows:
[0053] H Att = Proj(q·K)+Q (21)
[0054] Among them, Proj(·) represents a linear mapping layer;
[0055] Here, we perform adaptive fusion on the outputs of the two modules. The formula is as follows:
[0056] Z = Sigmoid(H Att W Att +H BiSGRN W BiSGRN ) (22)
[0057] Y = Z⊙H Att +(1-Z)⊙H BiSGRN (23)
[0058] Among them, H BiSGRN and H Att respectively represent the outputs of the bidirectional SGRN module and the additive attention module, W Att and W BiSGRN represent learnable weight matrices, and the fusion weight Z is obtained through the activation function Sigmoid;
[0059] Furthermore, we perform residual processing on the fused features to enhance the feature expression ability of the model:
[0060] H F = Y⊙GeLU(H0W1+b1) (24)
[0061] Among them, W1 and b1 are respectively the weight matrix and bias vector of the linear transformation. The activation function GeLU(x) = x + φ(x) provides a smooth non-linear transformation, avoiding the risk of gradient vanishing, and φ(x) represents the cumulative distribution function of the normal distribution.
[0062] Preferably, the working steps of the mixture-of-experts module are as follows:
[0063] The mixture-of-experts module contains a group of experts. Each expert is independently responsible for different parts of the input data and uses a gating mechanism to selectively activate different experts to handle different input features; each expert consists of a feedforward neural network. Different experts share the same structure but have their own independent parameters, as follows:
[0064] h i =GeLU(W i H F +b i ), i=1,2,...,E (25)
[0065] Among them, represents the input feature, W i and b i respectively represent the weight and bias of the i-th expert network, E represents the number of experts, and the expert network learns the representation of the input feature through a fully connected layer;
[0066] We adopt a routing mechanism called Top2Gating to dynamically activate the expert module. The specific process is as follows: First, calculate the basic score g i of each expert according to the input feature, and add random noise to it to improve the diversity of expert selection. The formula is as follows:
[0067]
[0068] Among them, ∈ i represents the noise sampled from the normal distribution ;
[0069] Then, sort the scores of all experts in descending order, select the two experts with the highest scores and keep their scores, and set the scores of the remaining experts to negative infinity; Next, normalize all the scores g i through the Softmax function to obtain the weight of each expert. The formula is as follows:
[0070]
[0071] Finally, select the two activated experts for feature modeling to obtain the final output. The formula is as follows:
[0072]
[0073] Advantages of the present invention:
[0074] The present invention designs a bidirectional gated recurrent unit module: simplifies the recurrent network structure, only retaining a single-layer recurrent unit; models in both forward and backward directions to effectively capture the forward and backward dependencies of sequential data; the gated unit can selectively control the information flow, retaining key features and removing noise; the present invention designs an additive attention module: realizes key-value interaction through linear operations, effectively reducing the computational complexity of the attention network while maintaining the global feature selection ability for effectively modeling the long-term preferences of users; the present invention introduces a mixture of experts mechanism: the mixture of experts mechanism adaptively selects a combination of multiple experts according to different input features, thus effectively dealing with the diversity of user behaviors and complex behavior modeling in specific scenarios; the present invention conducts a large number of experiments on a public dataset to verify the effectiveness and generalization ability of the model in this paper, and compares it with the current state-of-the-art sequential modeling methods. The results show that the model in this paper performs excellently in various indicators.
[0075] The features and advantages of the present invention will be described in detail through embodiments in conjunction with the accompanying drawings. Brief Description of the Drawings
[0076] Figure 1 is the architecture diagram of the bidirectional simplified gated recurrent network-attention fusion model;
[0077] Figure 2 is the model structure diagram of the simplified graph recurrent network;
[0078] Figure 3 is the architecture of the gated convolutional layer model;
[0079] Figure 4 is the diagram showing the influence of the Dropout rate on the model performance;
[0080] Figure 5 is the diagram showing the influence of the maximum sequence length on the model performance;
[0081] Figure 6 is the diagram showing the influence of the number of experts on the model performance. Detailed Description of the Preferred Embodiments
[0082] Problem Description:
[0083] Let U = {u1, u2,... u |U|} represent the set of users, V = {v1, v2,... v |V|} represent the set of items, and S u = {v1, v2,..., v n} represent the interaction sequence of user u ∈ U arranged in chronological order, where n is the length of the sequence; given the historical interaction sequence S u of the user, the task of sequential recommendation is to predict the next item that user u will interact with, denoted as v n+1。
[0084] Overall framework:
[0085] The sequence recommendation model proposed in this paper is called the Bidirectional Simplified Gated Recurrent Network-Attention Fusion Model, abbreviated as BiSGR-Att; the overall framework of the model is as Figure 1 shown, mainly including three components: the bidirectional gated SGRN module, the additive attention module, and the mixture of experts module; the bidirectional SGRN module captures the forward and backward dependencies of sequence data from both positive and negative directions, realizing accurate modeling of users' short-term dynamic preferences; the additive attention module focuses on identifying important interactions in the sequence, helping the model better understand context information, and using the linear attention calculation method to effectively model users' long-term preferences; the MoE module adaptively selects different expert networks according to different feature inputs, further optimizes the results of preference modeling, and improves the expressive ability of the model.
[0086] Bidirectional gated SGRN layer
[0087] Simplified Gated Recurrent Network (SGRN)
[0088] Traditional recurrent neural networks, such as RNN, GRU, and LSTM, etc., are difficult to capture the periodic fluctuations and diverse features in signals, and these features are very important for modeling dynamic user preferences in sequence recommendation; in the proposed model, we adopt the Simplified Gated Recurrent Network (SGRN) as the core unit to capture the dynamic changes of user preferences; SGRN is a simplification of the Hierarchical Gated Recurrent Network (HGRN)
[13] , abandoning its inherent multi-layer recurrent structure framework and only using a single-layer gated recurrent structure for modeling, while reducing the computational complexity, maintaining its ability to model long-term and short-term dependencies; the model structure of SGRN is as Figure 1 (left) shown, where SRU represents the gated recurrent unit, and its structure is as Figure 1 (right) shown.
[0089] For a given input sequence x t , SGRN first decomposes it into the real part and the imaginary part. The complex-form input is more sensitive to capturing features such as frequency changes and amplitude fluctuations in the signal and is more effective in dealing with periodic changes in user preferences; the decomposition formula is as follows:
[0090] Re(c t ) = SiLU(x t W r + b r ) (2)
[0091] Im(ct ) = SiLU(x t W i + b i ) (3)
[0092] where x t represents the input vector at time step t, c t represents the complex domain representation of the current input, Re(c t ) and Im(c t ) represent the real and imaginary parts of c t respectively, and the complex domain extends the hidden state by using the real and imaginary parts; W r and W i represent the weight matrices corresponding to the real and imaginary parts respectively, b r and b i represent the bias vectors corresponding to the real and imaginary parts respectively, and SiLU(x) = x + sigmoid(x) is a non - linear activation function used to enhance the model's expressive power.
[0093] Next, the hidden state is updated through the forget gate, and the formula is as follows:
[0094] μ t = Sigmoid(x t W μ + b μ ) (4)
[0095] h t = μ t exp(iθ)·h t-1 + (1 - μ t )·c t (5)
[0096] where μ t represents the forget gate, h t-1 represents the hidden state of the previous time step, and the rotation angle exp(iθ) introduces a phase change, which is crucial for capturing periodic preferences.
[0097] Finally, the corresponding output vector is obtained through the gating mechanism and projection transformation, and the formula is as follows:
[0098] g t = τ(W g x t + b g ) (6)
[0099] o′ t = LayerNorm(g t ⊙ [Re(h t ), Im(h t )]) (7)
[0100] o t = o' t W o + b o (8)
[0101] Among them, g t represents the gating vector, which is used to control the real and imaginary parts of the output. τ(·) represents the Sigmoid activation function. W o and b o represent the projection matrix and the bias vector respectively.
[0102] Bidirectional Simplified Graph Recurrent Network (Bidirectional SGRN) structure
[0103] To capture the bidirectional dependencies of user behavior patterns, we introduce a Bidirectional SGRN module to model the short-term dynamic preferences of users in both forward and backward directions. Forward modeling is used to capture the influence of users' past behaviors and preferences on the current preference, while backward modeling provides context information for the current moment by modeling the subsequent behaviors in the sequence. The forward and backward sequence modeling processes are as follows:
[0104]
[0105]
[0106] Among them, H0 and represent the embedding matrices corresponding to the user interaction sequence S u and the backward interaction sequence respectively. S0 and represent the modeling results of the forward and backward sequences respectively.
[0107] Here, we design a Gated Convolutional Layer (GCL) to fuse the results of the forward and backward sequence modeling. The model architecture of the GCL is as Figure 3 shown.
[0108] First, the local features of the sequence are extracted through one-dimensional convolution (Conv1d). The formula is as follows:
[0109] G0 = Conv1d(H0W c + b c ) (11)
[0110] Among them, W δ and b c represent the projection matrix and the bias vector of the convolution operation respectively.
[0111] Then, feature selection is performed through a two-channel architecture. The left channel plays a role of a gating mechanism. The execution process is as follows:
[0112] δ0(G0) = G0W δ +b δ (12)
[0113] L(G0) = Sigmoid(δ0(G0)) (13)
[0114] where, W δ and b δ represent the projection matrix and the bias vector respectively, and the Sigmoid activation function acts as a forgetting gate; the right channel directly models the input features using a non-linear activation function to increase the feature expression ability, and the formula is as follows:
[0115] R(G0) = SiLU(G0) (14)
[0116] Next, the modeling results of the left and right channels are fused, and the formula is as follows:
[0117]
[0118] where ⊙ represents the dot product operation between vectors, that is, the corresponding elements of two vectors are multiplied; the essence of the above operation is to use the left channel as a gate to achieve an adaptive selection of the right channel feature encoding, retain useful information, and remove noise.
[0119] Finally, the forward and backward sequence modeling results are fused through GCL, and the formula is as follows:
[0120]
[0121] where, and represent the outputs obtained by the forward and backward sequences through GCL respectively, and serve as the weight vectors of the forward and backward sequence modeling results in the fusion process.
[0122] Additive Attention Module
[0123] Bidirectional SGRN can capture the forward and backward dependencies of the sequence and realize the modeling of the user's dynamic behavior characteristics. However, relying solely on the bidirectional recurrent network is not enough to capture the global information in the sequence; the self-attention mechanism can effectively extract the global context in the sequence and realize the modeling of the user's long-term preferences; however, the traditional self-attention mechanism has a high computational complexity, and the computational cost and memory consumption will increase significantly when processing long sequences; in order to reduce the computational complexity, we designed a linear additive attention mechanism that can effectively model the user's long-term preferences.
[0124] First, the input vector is linearly transformed to generate low-rank representations of the query matrix and the key matrix, and the formula is as follows:
[0125]
[0126]
[0127] Among them, H0 represents the output of the embedding layer, and respectively represent the low-rank transformation matrices of the query and the key.
[0128] Then, in order to balance the computational complexity and the information extraction effect, we introduce a learnable weight vector W a to calculate the attention distribution and perform a normalization operation on the result to obtain the attention weights. The formula is as follows:
[0129]
[0130] Next, based on the attention weights, each element of the query matrix is weighted and summed to obtain the global query vector q. The formula is as follows:
[0131]
[0132] Finally, directly multiply the global query vector by the key matrix and then linearly combine it with the query vector to obtain the sequence modeling result. The formula is as follows:
[0133] H Att = Proj(q·K)+Q (21)
[0134] Among them, Proj(·) represents a linear mapping layer; the above process realizes the key-value interaction through linear operations, reduces the computational complexity of the attention model from quadratic to linear level, and at the same time retains the ability of the model to capture context information.
[0135] Combining the bidirectional SGRN gating module and the additive attention module can effectively model the long-term and short-term preferences of users; here, we perform adaptive fusion on the outputs of the two modules. The formula is as follows:
[0136] Z = Sigmoid(H Att W Att + H BiSGRN W BiSGRN ) (22)
[0137] Y = Z⊙H Att +(1-Z)⊙H BiSGRN (23)
[0138] Among them, H BiSGRN and H Att respectively represent the outputs of the bidirectional SGRN module and the additive attention module, WAtt and W BiSGRN represents a learnable weight matrix, and the fused weight Z is obtained through the activation function Sigmoid.
[0139] Furthermore, we perform residual processing on the fused features to enhance the feature expression ability of the model, as follows:
[0140] H F = Y ⊙ GeLU(H0W1 + b1) (24)
[0141] where W1 and b1 are the weight matrix and bias vector of the linear transformation respectively, and the activation function GeLU(x) = x + φ(x) provides a smooth non-linear transformation, avoiding the risk of gradient vanishing, and φ(x) represents the cumulative distribution function of the normal distribution.
[0142] Mixture of Experts (MoE) module
[0143] To enhance the generalization ability of the model and the ability to capture diverse features, we designed a Mixture of Experts (MoE) module; MoE contains a group of experts, each expert is independently responsible for different parts of the input data, and the gating mechanism is used to selectively activate different experts to handle different input features; each expert consists of a feed-forward neural network, different experts share the same structure but have their own independent parameters, as follows:
[0144] h i = GeLU(W i H F + b i ), i = 1, 2,..., E (25)
[0145] where represents the input feature, W i and b i represent the weight and bias of the i-th expert network respectively, E represents the number of experts, and the expert network learns the representation of the input feature through a fully connected layer.
[0146] We adopt a routing mechanism called Top2Gating
[32] to dynamically activate the expert module, and the specific process is as follows: First, the basic score g of each expert is calculated according to the input feature i , and random noise is added to it to enhance the diversity of expert selection. The formula is as follows:
[0147]
[0148] where ∈ i represents the noise sampled from the normal distribution .
[0149] Then, sort the scores of all experts in descending order, select the top 2 experts with the highest scores and keep their scores, and set the scores of the remaining experts to negative infinity; next, normalize all scores g i Normalize through the Softmax function to obtain the weights of each expert. The formula is as follows:
[0150]
[0151] Finally, select the two activated experts for feature modeling to obtain the final output. The formula is as follows:
[0152]
[0153] Prediction layer
[0154] In the prediction layer, we use the vector multiplication method to calculate the predicted scores of candidate items. The formula is as follows:
[0155]
[0156] Among them, E represents the embedding of the item, represents the probability distribution of different items in the item set V as the next visited item.
[0157] During the model training process, we use the cross-entropy loss function to minimize the difference between the predicted item and the true item. The formula is as follows:
[0158]
[0159] Experiment
[0160] To verify the effectiveness of the proposed method BiSGR-Att in this paper, we conducted a large number of experiments under the Recbole
[23] framework, aiming to answer the following research questions.
[0161] RQ1: How does the performance of BiSGR-Att compare with the current state-of-the-art sequential recommendation baseline models?
[0162] RQ2: Can the core components in BiSGR-Att play a positive role in the sequential recommendation process?
[0163] RQ3: What is the impact of hyperparameter settings on the model performance? Especially in the MoE module, how many experts should be selected to achieve the best performance?
[0164] Experiment preparation
[0165] Dataset
[0166] To verify the effectiveness of the BiSGR-Att model, we conducted experiments on four publicly available datasets with different categories and different sparsity levels:
[0167] Amazon Beauty (Beauty): A dataset of product reviews and ratings from the "Beauty" category on Amazon. Amazon Video-Games (Video-Games): A dataset of product reviews and ratings from the "Video Games" category on Amazon.
[0168] MovieLens 100K (ML-100K): A movie rating dataset from the MovieLens website, containing 100,000 rating records of 1,682 movies by 943 users.
[0169] Ali_Display_Ad_Click (AliEC): A dataset for predicting the click-through rate of Taobao display ads from Alibaba.
[0170] We preprocessed the original datasets, filtered out users and items with fewer than 5 interactions, and sorted the interaction records of users according to timestamps to generate interaction sequences. The detailed statistics of the preprocessed datasets are shown in Table 1:
[0171] Table 1: Statistics of the datasets
[0172]
[0173] Evaluation metrics
[0174] We used Hit Rate (HR), Normalized Discounted Cumulative Gain (NDCG), and Mean Reciprocal Rank (MRR) as evaluation metrics, and tested the top 10 items in the recommendation list, namely HR@10, NDCG@10, and MRR@10; to prove the effectiveness of the proposed BiSGR-Att model in this paper, we compared BiSGR-Att with a series of baseline models, including Caser based on CNN, models based on RNN (NARM and GRU4Rec), models based on Transformer (BERT4Rec and SASRec), a model based on linear attention LinRec, models based on linear recurrent units RecBLR and GLINT-RU; the selected baseline models adopted different sequence modeling methods and are currently the sequence recommendation models with the best performance.
[0175] Experimental settings
[0176] We conducted experiments on the NVIDIA RTX 3090 with 24GB of video memory and completed model encoding using the PyTorch environment; for the Beauty and Video Games datasets, the hidden layer size was set to 128, and for other datasets, it was set to 64; for the ML-100K dataset, the maximum sequence length was set to 150, the AliEC dataset was set to 40, and other datasets were set to 100; the Adam optimizer was used for model training, and the learning rate was set between 5×10 -4 and 1×10 -3 . The batch sizes for both training and testing were set to 2048, and other implementation details followed the settings of the original paper.
[0177] Overall performance (RQ1)
[0178] In this experimental session, we comprehensively evaluated the recommendation performance of different sequence modeling methods. The experimental results are shown in Table 2. On the Beauty and Games datasets, the Caser, GRU4Rec, and NARM models performed very poorly, indicating that traditional deep learning methods (such as CNN and DNN) have relatively poor performance in dealing with complex dependencies. Transformer-based models (such as SASRec and BERT4Rec) benefit from the powerful sequence modeling ability of the self-attention mechanism and perform well in this regard, but they still lack in short sequence modeling (corresponding to the AliEC dataset). The recommendation performance of models based on linear attention and linear recurrent units (LinRec, RecBLR, and GLINT-RU) has been further improved, fully demonstrating the positive role of linear attention mechanisms and linear recurrent units in sequence modeling. The proposed model BiSGR-Att in this paper combines linear attention mechanisms and linear recurrent units and further adds a mixture of experts model on this basis. From the experimental results, BiSGR-Att performs excellently on both simple datasets (such as ML-100K) and complex and sparse e-commerce datasets (such as Beauty and Games). Compared with the comparison models, the proposed model BiSGR-Att has significant performance improvements in all three metrics of HR@10, NDCG@10, and MRR@10. The reason is that when dealing with short sequences, the gated bidirectional SGRN is used to model both forward and backward dependencies simultaneously, capturing more comprehensive context information in a balanced way and realizing the modeling of users' short-term dynamic preferences. The simple and efficient additive attention mechanism is used to process complex long sequences, extracting important information to realize the modeling of users' long-term preferences. Finally, the MoE component is used for encoding optimization, further enhancing the generalization ability of the model and the ability to capture diverse features.
[0179] Table 2: Performance comparison of different methods
[0180]
[0181] Ablation Experiment (RQ2)
[0182] To understand the contributions of each core component in the model of this paper, we conducted ablation experiments; specifically, we implemented the following model variants:
[0183] w / o ATT: Remove the additive attention module in the BiSGR-Att model, and keep the remaining components unchanged.
[0184] w / o SGRN: Remove the bidirectional SGRN module in the BiSGR-Att model, and keep the remaining components unchanged. w / o GCL: Remove the GCL module in the BiSGR-Att model, and keep the remaining components unchanged.
[0185] w / o MoE: Remove the MoE module in the BiSGR-Att model, and keep the remaining components unchanged.
[0186] Table 3: Results of Ablation Experiment
[0187]
[0188] The performance comparison of the model of this paper and its four variants on different datasets is shown in Table 3; it can be seen from the table that after removing the bidirectional SGRN from the model, the overall performance drops most significantly, especially on the ML-100K dataset, where HR@10 drops from 0.1453 to 0.1209, indicating that the bidirectional SGRN plays the most crucial role in the model of this paper because this module plays an important role in both modeling users' long-term and short-term preferences and feature expressions, ensuring that the model can capture the complex dynamics of user behavior and the trend of preference changes; after removing the additive attention module, the model performance drops on all datasets, indicating that the additive attention mechanism plays an indispensable role in modeling user behavior features; by introducing the additive attention mechanism, the model can effectively learn the long-range dependencies in the user sequence and improve the understanding of user preferences and complex behavior patterns; the GCL module and the MoE module also have a positive effect on the model performance; the GCL module effectively filters out noise signals by dynamically adjusting the weights of channels; in the absence of the GCL module, the model will not capture important behaviors accurately and its adaptability to short-term preferences will also decrease; the MoE module plays a role similar to that of the feed-forward network in the whole model, and it improves the model's ability to handle users' diverse preferences and complex behavior features by introducing multiple expert networks; deleting the MoE module will affect the model's ability to learn complex feature representations and cannot achieve fine-grained modeling of user behavior, thus affecting the overall performance of the model.
[0189] Parameter Sensitivity Analysis (RQ3)
[0190] In this section, we analyze the impact of hyperparameter settings in the model on the recommendation performance through experiments. First, we analyze the Dropout parameter. Dropout is a strategy that randomly discards some neurons during training to prevent the model from overfitting. When the Dropout rate is low, the model may be prone to overfitting the details in the training data, resulting in poor performance on new data in the test set. When the Dropout rate is high, the model may not be able to effectively learn from the sparse training data, thus losing important feature information and affecting the recommendation effect. The impact of the Dropout parameter setting on the recommendation performance is as Figure 4 shown, which presents the performance of the proposed model on different datasets when the Dropout rate ranges from 0.1 to 0.9. For the ML-100K dataset, the model performance reaches its peak when the Dropout rate is 0.4, indicating that in the case of relatively dense user-item interaction data, moderate Dropout can effectively prevent overfitting while retaining sufficient information to capture users' behavior patterns. On the Beauty and Games datasets, the model performs best when the Dropout rate is 0.7, suggesting that moderately increasing Dropout helps improve the generalization ability of the model in a sparse interaction data environment. For the AliEC dataset, due to its more complex and diverse features, the model performs well when the Dropout rates are 0.4 and 0.7, showing two performance peaks. Generally speaking, the setting of the Dropout parameter needs to be carefully adjusted according to the characteristics of the dataset to achieve a balance between preventing overfitting and maintaining the model's learning ability.
[0191] Next, we analyze the impact of the maximum sequence length (N) on the model performance on three sparse datasets, namely Beauty, Games, and AliEC. The experimental results are as Figure 5As shown; it can be seen from the figure that when N is set to be small, the model cannot fully capture the long-distance dependence information in the sequence, resulting in unsatisfactory performance of the model; while when N is set too large, the added irrelevant noise information increases, affecting the model's ability to extract effective information and leading to performance degradation; in the Beauty and Games datasets, the model achieves the best performance when N = 100, at which time it can effectively capture the key information in user behavior while avoiding introducing too much noise; when N exceeds 100, the performance of the model begins to decline rapidly, which is due to the introduction of too much irrelevant information interfering with the learning of effective information; in addition, as N increases, the demand for computing resources also increases, so in practical applications, the model performance and computational overhead should be balanced, and the maximum sequence length should be reasonably selected; for the AliEC dataset, the optimal value of N is 40, which indicates that a smaller sequence length performs better in sparse and fast-changing behavior scenarios, because a shorter sequence length can help the model better focus on the key features in user behavior and reduce the impact of irrelevant noise; generally speaking, the differences in the optimal sequence lengths of different datasets reflect their respective data characteristics. In practical applications, selecting an appropriate maximum sequence length can effectively improve the performance of the model.
[0192] Finally, we analyzed the impact of the number of experts in the MoE module on the performance of sequence recommendation; by setting different numbers of experts, experimental evaluations were conducted on four datasets, and the results are as Figure 6 shown; it can be seen from the figure that on all datasets, the best results are achieved when the number of experts is 4 - 5; in the AliEC dataset, when the number of experts increases from 3 to 5, the performance gradually increases; on other datasets, the performance is optimal when the number of experts increases to 4; this indicates that increasing the number of experts can improve the model's ability to capture user behavior characteristics; when the number of experts continues to increase to 7 or 8, the performance of the model levels off or even shows a slight decline; this indicates that increasing the number of experts is not always beneficial and may also introduce additional noise, resulting in an increased risk of overfitting; generally speaking, selecting an appropriate number of experts is crucial for model performance. An appropriate number of experts can not only effectively capture the complex patterns in user behavior but also avoid the overfitting problems and computational resource load caused by too many experts, thus achieving the best balance in the recommendation task.
[0193] The above embodiments are illustrative of the present invention and not restrictive thereof. Any scheme obtained by simply transforming the present invention falls within the protection scope of the present invention.
Claims
1. A sequence recommendation model based on bidirectional simplified gated recurrent network and linear attention, characterized by: It includes an embedding layer, a bidirectional simplified gated recurrent network module, an additive attention module, a hybrid expert module, and a prediction layer; The bidirectional simplified gated recurrent network module simultaneously captures the forward and backward dependencies in the user behavior sequence, and is used to model the user's short-term dynamic preferences; The additive attention module reduces the computational complexity of attention weights from quadratic to linear, and is used to model users’ long-term preferences. The hybrid expert module adaptively selects appropriate experts to participate in the calculation according to the input features, so as to cope with the diversity of user behaviors and complex behavior modeling in specific scenarios.
2. The sequence recommendation model based on bidirectional simplified gated recurrent network and linear attention as claimed in claim 1, characterized in that: The working steps of the embedding layer are as follows: Assume the item embedding matrix is Where |V| is the size of the item set, D is the dimension of the embedding vector; for the user's historical interaction sequence S u ={v1,v2,...,v n }, each item v is embedded in it through the embedding layer i Mapped to the corresponding vector representation e i =E(v i ), and get the sequence embedding matrix In order to enhance the robustness of embedding and prevent overfitting, we performed a Dropout operation on top of the embedding layer, as follows:
3. The sequence recommendation model based on bidirectional simplified gated recurrent network and linear attention as claimed in claim 1, characterized in that: The working steps of the bidirectional simplified gated recurrent network module are as follows: A simplified gated recurrent network is used as the core unit to capture the dynamic changes of user preferences: For a given input sequence x t , To simplify the gated recurrent network, first decompose it into real and imaginary parts. The decomposition formula is as follows: Re(c t )=SiLU(x t W r +b r ) (2) Im(c t )=SiLU(x t IN i +b i ) (3) Among them, x t represents the input vector at time step t, c t Represents the complex domain representation of the current input, Re(c t ) and Im(c t ) represent c t The real and imaginary parts of the complex domain are used to expand the hidden state; W r and W i Represents the weight matrices corresponding to the real and imaginary parts, b r and b i Represent the bias vectors corresponding to the real part and the imaginary part respectively. SiLU(x)=x+sigmoid(x) is a nonlinear activation function used to enhance the expressiveness of the model. Next, the hidden state is updated through the forget gate, and the formula is as follows: μ t =Sigmoid(x t W μ +b μ ) (4) h t =μ t exp(iθ)·h t-1 +(1-m t )·c t (5) Among them, μ t represents the forget gate, h t-1 Represents the hidden state of the previous time step, and the rotation angle exp(iθ) introduces a phase change; Finally, the corresponding output vector is obtained through the gating mechanism and projection transformation. The formula is as follows: g t =τ(W g x t +b g ) (6) o' t =LayerNorm(g t ⊙[Re(h t ),Im(h t )]) (7) o t =o′ t W o +b o (8) Among them, g t represents the gate vector, which is used to control the real and imaginary parts of the output, τ(·) represents the Sigmoid activation function, and W o and b o Represent the projection matrix and bias vector respectively; The bidirectional SGRN module is introduced to model users’ short-term dynamic preferences from both positive and negative directions. The sequence modeling process of forward and reverse order is as follows: Among them, H0 and Represent the user interaction sequence S u and reverse interaction sequences The corresponding embedding matrices, S0 and Represent the modeling results of the forward and reverse sequences, respectively; Here, we designed a gated convolutional layer to fuse the results of forward and reverse sequence modeling; First, the local features of the sequence are extracted through one-dimensional convolution. The formula is as follows: G0=Conv1d(H0W c +b c ) (11) Among them, W δ and b c Represent the projection matrix and bias vector of the convolution operation respectively; Then, feature selection is performed through a dual-channel architecture, with the left channel acting as a gating mechanism. The execution process is as follows: δ0(G0)=G0W δ +b δ (12) L(G0)=Sigmoid(δ0(G0)) (13) Among them, W δ and b δ Represent the projection matrix and the bias vector respectively. The Sigmoid activation function acts as a forget gate. The right channel directly uses the nonlinear activation function to model the input features to increase the expressiveness of the features. The formula is as follows: R(G0)=SiLU(G0) (14) Next, the modeling results of the left and right channels are fused, and the formula is as follows: Where ⊙ represents the dot product operation between vectors, that is, the corresponding elements of two vectors are multiplied; Finally, GCL is used to achieve the fusion of the forward and reverse sequence modeling results. The formula is as follows: in, and They represent the outputs of the forward and reverse sequences obtained by GCL respectively, and serve as the weight vectors of the forward and reverse sequence modeling results in the fusion process.
4. The sequence recommendation model based on bidirectional simplified gated recurrent network and linear attention as claimed in claim 1, characterized in that: The working steps of the additive attention module are as follows: First, the input vector is linearly transformed to generate the representation of the query matrix and the key matrix. The formula is as follows: Among them, H0 represents the output of the embedding layer, and Represent the transformation matrices of query and key respectively; Then, in order to balance the computational complexity and information extraction effect, we introduce a learnable weight vector W a To calculate the attention distribution, and normalize the result to get the attention weight, the formula is as follows: Next, we weight each element of the query matrix based on the attention weight to obtain the global query vector q, as follows: Finally, the sequence modeling result is obtained by directly multiplying the global query vector with the key matrix and then linearly merging it with the query vector. The formula is as follows: H Att =Proj(q·K)+Q (21) Among them, Proj(·) represents a linear mapping layer; Here, we adaptively fuse the outputs of the two modules, the formula is as follows: Z=Sigmoid(H Att W Att +H BiSGRN W BiSGRN ) (22) Y=Z⊙H Att +(1-Z)⊙H BiSGRN (23) Among them, H BiSGRN and H Att denote the outputs of the bidirectional SGRN module and the additive attention module, respectively, and W Att and W BiSGRN Represents a learnable weight matrix, and the fusion weight Z is obtained through the activation function Sigmoid. Furthermore, we perform residual processing on the fused features to enhance the feature expression ability of the model: H F =Y⊙GeLU(H0W1+b1) (24) Among them, W1 and b1 are the weight matrix and bias vector of the linear transformation respectively, the activation function GeLU(x)=x+φ(x) provides a smooth nonlinear transformation and avoids the risk of gradient disappearance, and φ(x) represents the cumulative distribution function of the normal distribution.
5. The sequence recommendation model based on bidirectional simplified gated recurrent network and linear attention as claimed in claim 1, characterized in that: The working steps of the hybrid expert module are as follows: The hybrid expert module contains a group of experts, each of which is responsible for different parts of the input data. The gating mechanism is used to selectively activate different experts to deal with different input features. Each expert consists of a feedforward neural network. Different experts share the same structure but have their own independent parameters, as shown below: h i =GeLU(W i A F +b i ),i=1,2,...,E (25) in, represents the input features, W i and b i They represent the weight and bias of the i-th expert network respectively, E represents the number of experts, and the expert network learns the representation of input features through the fully connected layer; We use a routing mechanism called Top2Gating to dynamically activate the expert module. The specific process is as follows: First, the basic score g of each expert is calculated based on the input features. i On this basis, random noise is added to improve the diversity of expert selection. The formula is as follows: Among them, ∈ i Represents a normal distribution The noise obtained by sampling; Then, sort the scores of all experts in descending order, select the two experts with the highest scores and keep their scores, and set the scores of the remaining experts to negative infinity; next, set all the scores g i The weight of each expert is obtained by normalizing through the Softmax function. The formula is as follows: Finally, two activated experts are selected for feature modeling to obtain the final output. The formula is as follows:
6. The sequence recommendation model based on bidirectional simplified gated recurrent network and linear attention as claimed in claim 1, characterized in that: The working steps of the prediction layer are as follows: We use the vector product method to calculate the predicted score of the candidate items. The formula is as follows: Where E represents the embedding of the item, Represents the probability distribution of different items in the item set V as the next access item; During model training, we used the cross entropy loss function to minimize the difference between the predicted items and the true items. The formula is as follows:
Citation Information
Cited By
Hybrid expert model sparse reasoning method and system for generative recommendation
CN121835927A
Wafer process fault early warning method and system based on mixed sequence decomposition and Mama architecture
CN122112929A
Wafer processing fault early warning method and system based on hybrid sequence decomposition and mamba architecture
CN122112929B
Sequence recommendation method for decoupling modeling of long and short term preferences of user based on frequency domain analysis
CN122196275A