A Deep Neural Network Recommendation Model and Method Based on the Multi-Head Self-Attention Mechanism
By introducing a multi-head self-attention mechanism into the deep neural network recommendation model, combining the sequential information and contextual relationships at the session level, the limitations of user interest feature representation and interest change capture are solved, and more efficient user interest analysis and recommendation accuracy are achieved.
Patent Information
- Application Number
- CN202210529804.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-05-16
AI Technical Summary
The existing deep neural network recommendation model has limitations in user interest feature representation and interest change capture, and cannot effectively extract user interest features in multiple dimensions, ignoring the relationship between users and item advertisements and the evolution of user interests.
A deep neural network recommendation model based on multi-head self-attention mechanism is designed, sequential position information is introduced through session division layer, multi-head self-attention network is used to mine multi-dimensional implicit relationships between behaviors in the session interest interaction layer, and to capture the evolutionary law of user interest in the session interest activation layer combined with contextual relationships.
This model can analyze deep implicit features in multiple dimensions, improve the accuracy of the recommendation system, and more accurately characterize the diversity of user interest characteristics, and capture the relationship between users and item advertisements.
Smart Images

Figure CN115203529B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of recommendation, and particularly relates to a deep neural network recommendation model and method based on a multi-head self-attention mechanism. Background Art
[0002] Recommendation algorithms infer what users may like based on differences in their historical behaviors and personal preferences. Traditional recommendation algorithms are represented by collaborative filtering models. The two most common recommendation models are item-based collaborative filtering algorithms and user-based collaborative filtering algorithms. The former realizes recommendation by calculating the similarity of implicit vectors between items and sorting to take the TopN. The latter realizes recommendation by extracting the implicit features of users from the historical evaluations of users and sorting the feature similarities to take the TopN. The above two methods are simple and easy to implement by extracting static features. One only needs to focus on the historical behaviors of users to extract implicit features and sort them according to similarity. However, the high-dimensional cross-combination between implicit features and the influence of the user behavior time series, an important dimension, are ignored.
[0003] Deep learning models have gradually been applied in the field of recommendation. Deep neural network recommendation models can fit high-order non-linear relationships through the high-dimensional combination of input features. However, directly using DNN (Shan Y, Hoens T R, Jiao J, et al. Deep crossing: web-scale modeling without manually crafted combinatorial features [C]. In Proceedings of the 2 2nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), 2016: 255-262.) to extract user interests is doomed to be unable to model the changes in the time series, because the behavior of a user sequence at a certain moment is doomed to be related not only to the current behavior but also to the previous behaviors. In addition, the existing deep neural network recommendation models also face the following problems: First, the dimensionality of the interest features representing users is limited. The same dimensionality of user features extracted from different user behavior sequences makes it impossible to represent the diversity of user interest features in a personalized way. Second, the relationship between users and item advertisements is ignored. The same feature vector is used to represent the interests of users for different recommended candidate advertisements, which obviously limits the expressive power of the model. Third, the changes in user interests are ignored. The features extracted from the user behavior sequence should have a process of interest evolution to present a continuous change trend.
[0004] Therefore, how to design a deep neural network recommendation model so that the model can extract users' interest features in multiple dimensions and thus complete more accurate recommendation results has become the research focus. Summary of the Invention
[0005] Aiming at the problem in the background technology that most existing session-based recommendation algorithms do not fully consider the implicit features of other dimensions in multiple different spaces and the problem of limited dimensionality for personalized user interest analysis, the purpose of the present invention is to provide a deep neural network recommendation model based on the multi-head self-attention mechanism (Multi-head Self-attention Deep Neural Network for Session-based Recommendation Model, MSDN) and method. This model designs a session partitioning layer to introduce the sequential position information between different behaviors within and between sessions. In the session interest interaction layer, a multi-head self-attention network is used to mine the multi-dimensional implicit relationships between various behaviors of the user in the session. Then, in the session interest activation layer, the evolution law of the user's interest between sessions is captured by combining the context relationship for block activation, enabling the model to analyze deep implicit features in multiple dimensions, thereby improving the accuracy of the recommendation system.
[0006] To achieve the above object, the technical solution of the present invention is as follows:
[0007] A deep neural network recommendation model based on the multi-head self-attention mechanism, including a behavior data preprocessing module, a session partitioning module, a session interest interaction module, a session interest activation module, and a fully connected module;
[0008] The behavior data preprocessing module is used to obtain the user's historical behavior data and preprocess the historical data to obtain the user behavior sequence and the training data set;
[0009] The session partitioning module is used to perform fine-grained partitioning on the user behavior sequence obtained after preprocessing at the session level, and divide the user behavior sequence sorted in chronological order into different sessions according to the time interval threshold. Each session contains the user behavior sequence with different time intervals;
[0010] The session interest interaction module includes a multi-head self-attention sublayer and a multi-head self-attention network formed by residual connections. By extracting the multi-dimensional correlations between the current behavior and other behaviors within the session from multiple angles, and then stacking multiple layers through the multi-head self-attention sublayer to construct a multi-head self-attention network for extracting the correlation relationships between behaviors within the same session;
[0011] The session interest activation module adopts a bidirectional long short-term memory structure, which is used to capture the interest drift evolution of the user between sessions and comprehensively extract the interest evolution features from the context;
[0012] The fully connected module is used to make predictions based on the user behavior sequence obtained through preprocessing and the interest evolution features output by the session interest activation module, and then perform TopN sorting on the prediction results, and recommend the product with the highest score to the user.
[0013] Furthermore, the user historical behavior data is specifically the behavior sequence information accessed by the user extracted from the buried point log data such as browsers and APPs, including user registration information, product information, user evaluations, and access timestamps.
[0014] Furthermore, the specific process of preprocessing is as follows: First, extract the information of corresponding columns such as product numbers, user names, and comment numbers from the user historical behavior dataset as needed, and associate the rating data and product data according to the product numbers; then generate the initial user behavior sequence in the order of the corresponding timestamps, and limit the length of each user interaction sequence to obtain the user behavior sequence; finally, construct the positive and negative samples required for model training. Among them, the positive sample is the sample data in which the user participates in the comment in the comment dataset, that is, the real user behavior sequence. On the contrary, the negative sample is the sample data that is not actually interacted by the user but is simulated, that is, the user behavior sequence constructed artificially.
[0015] Furthermore, the number of positive and negative samples is kept at 1:1 to avoid the situation that the loss of the entire model biases to one side due to the imbalance of the input sample ratio.
[0016] The present invention also provides a recommendation method for a deep neural network recommendation model based on a multi-head self-attention mechanism, including the following steps:
[0017] Step 1. Obtain user behavior data, then preprocess the user behavior data to obtain a user behavior sequence, and construct an input sample set for the recommendation model;
[0018] Step 2. Perform fine-grained partitioning on the user behavior sequence obtained in Step 1 at the session level, partition the user behavior sequence into different sessions according to the time interval threshold, and then perform embedding mapping processing on the user behavior sequence in each session to obtain a low-dimensional dense feature vector Q embedding , and then add the position bias between sessions to obtain the user behavior embedding vector Q embedding_pos ;
[0019] Step 3. Calculate the single-head attention distribution vector based on the user behavior embedding vector Q obtained in Step 2 embedding_pos , then splice the single-head attention distribution vectors to obtain a multi-head self-attention distribution vector MultiHead(Q, K, V) with the same dimension as the user behavior embedding vector, and then perform hierarchical normalization and stacking processing to obtain the user multi-dimensional interest vector I in the same session k;
[0020] Step 4. Calculate the user interest evolution feature H between sessions t , and at the same time, for the user interest evolution feature H between sessions t and the user multi-dimensional interest vector I in the session obtained in Step 3 k perform partial activation to obtain the global expression U of the user interest evolution feature between sessions H and the global expression U of the multi-dimensional interest vector in the session I ;
[0021] Step 5. Use the concatenated user behavior feature vector as the input of multiple fully connected layers, make a preliminary prediction based on the target loss function, then add an auxiliary loss function to correct the prediction results at each time step, and finally sort the prediction results by TopN to recommend the product with the highest score to the user.
[0022] Furthermore, the specific process of extracting the correlation relationship between behaviors in the same session in Step 3 is as follows:
[0023] Step 3.1. Divide the user behavior embedding vector Q embedding_pos obtained in Step 2 evenly into multiple single-head structures, and calculate the correlation weight vectors in each subspace using scaled dot product;
[0024] Step 3.2. After concatenating each correlation weight vector, convert it into a multi-dimensional correlation weight vector MultiHead(Q, K, V) with the same dimension as the user behavior embedding vector;
[0025] Step 3.3. Use layer normalization to normalize the data in the same layer of the multi-dimensional correlation weight vector MultiHead(Q, K, V) to obtain the multi-head attention sublayer;
[0026] Step 3.4. Through residual connection of the multi-head self-attention sublayer, stack to form a network, and extract deep cross implicit features, that is, the user multi-dimensional interest vector I in the same session k .
[0027] Furthermore, the specific process in Step 4 is as follows:
[0028] Step 4.1. Adopt the bidirectional long short-term memory method to extract the user's interest drift evolution at the session level. Input two long short-term memory neural networks in the forward and reverse orders for feature extraction to obtain the forward candidate hidden layer state and the reverse candidate hidden layer state Finally, concatenate the hidden layer states in both directions to form the user interest evolution feature H between sessions that combines context information t ;
[0029] Step 4.2. Reassign the weight scores according to the matching degree of the target project, and activate the user interests in parts, specifically: The first part is to locally activate the multi-dimensional user interest vector I of the current session layer extracted from the interest interaction layer K to obtain the global expression U of the multi-dimensional interest vector I ; The second part is to locally activate the user interest evolution relationship H between sessions mined by two-way modeling t to obtain the global expression U of the user interest evolution characteristics between sessions H .
[0030] Furthermore, in step 5, the target loss function L target is specifically
[0031]
[0032] where x is a set concatenated by a series of interest vectors, x = [Q embedding , U I , U H , including Q embedding the user basic behavior embedding vector, U I is the global expression of the multi-dimensional interest vector, U H is the global expression of the interest evolution feature; D is a training set of size N, y is the true interest degree value of the user, y = {0, 1}, and p(x) is the final output of the model, indicating the predicted interest degree of the user in the commodity;
[0033] L aux is the auxiliary loss function, specifically
[0034]
[0035] where e b i [t + 1] is the positive sample at the next moment, is the negative sample at the next moment.
[0036] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are:
[0037] Aiming at the problem that most existing session-based recommendation algorithms do not fully consider the implicit features of other dimensions in multiple different spaces, resulting in limited dimensions for personalized user interest analysis, the present invention proposes a deep neural network recommendation model (MSDN) with a multi-head self-attention mechanism. A session division layer is designed to introduce the sequential position information between different behaviors within and between sessions. In the session interest interaction layer, a multi-head self-attention network is used to mine the multi-dimensional implicit relationships between the behaviors of users in a session. Then, in the session interest activation layer, the evolution law of user interests between sessions is captured by combining the context relationship for partial activation. The partial activation can generate diverse user interest feature vectors according to different combinations of features. The construction form of the entire model enables the model to finally have a high recommendation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a schematic structural diagram of the deep neural network recommendation model of the present invention.
[0039] Figure 2 It is a schematic flow diagram of the recommendation method based on the deep neural network recommendation model of the present invention.
[0040] Figure 3 It is a schematic structural diagram of the session interest interaction layer in the deep neural network recommendation model of the present invention.
[0041] Figure 4 It is a schematic structural diagram of the auxiliary loss function in the deep neural network recommendation model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the embodiments and the drawings.
[0043] A deep neural network recommendation model based on a multi-head self-attention mechanism, the structural diagram of which is as Figure 1 shown, including a behavior data preprocessing module, a session division module, a session interest interaction module, a session interest activation module, and a fully connected module;
[0044] The behavior data preprocessing module is used to obtain the historical behavior data of users and preprocess the historical data to obtain the user behavior sequence and the training data set;
[0045] The session division module is used to perform fine-grained division on the user behavior sequence obtained after preprocessing at the session level, and divide the user behavior sequence sorted by time order into different sessions according to the time interval threshold. Each session contains the user behavior sequence with different time intervals;
[0046] The session interest interaction module includes a multi-head self-attention sublayer and a multi-head self-attention network formed by residual connections. By extracting the multi-dimensional correlations between the current behavior and other behaviors in the session from multiple perspectives, and then stacking the multi-head self-attention sublayers to construct a multi-head self-attention network for extracting the correlation relationships between behaviors within the same session;
[0047] The session interest activation module adopts a bidirectional long short-term memory structure to capture the interest drift evolution of the user between sessions and extract interest evolution features by integrating the context;
[0048] The fully connected module is used to make predictions based on the user behavior sequence obtained by preprocessing and the interest evolution features output by the session interest activation module, and then sort the prediction results by TopN to recommend the product with the highest score to the user.
[0049] Embodiment 1
[0050] A recommendation method for a deep neural network recommendation model based on the multi-head self-attention mechanism. The schematic flow chart of this method is as Figure 2 shown, including the following steps:
[0051] Step 1. Obtain user behavior data, and then preprocess the user behavior data to obtain a user behavior sequence and construct an input sample set for the recommendation model. The specific process is as follows:
[0052] Step 1.1. Obtain user behavior data:
[0053] Extract the behavior sequence information accessed by the user from the buried point log data such as browsers and APPs, including user registration information (such as user ID, gender, age, permanent address, etc.), product information (such as product ID, category ID, store ID, etc.), user evaluations (such as ratings, evaluation messages, etc.) and access timestamps (timestamp);
[0054] Since the Amazon Electronics dataset is used, the relevant data is a compressed file in json format.
[0055] Step 1.2. Preprocess the user behavior data obtained in Step 1.1:
[0056] First, extract the information of the corresponding columns such as product numbers, user names, and comment numbers from the compressed dataset file in json format, associate the rating data and product data according to the product number, then generate an initial user behavior sequence in the order of the corresponding access timestamps, and then limit the length of each user behavior sequence to obtain the user behavior S; finally, construct the input sample set required for training the recommendation model;
[0057] The input sample set includes positive and negative samples. Among them, the positive samples are the sample data of the actual comments made by users in the comment dataset, that is, the real user behavior sequences, and the negative samples are the sample data artificially simulated without actual comments by users;
[0058] At the same time, it is necessary to perform data balancing operations on the positive and negative samples. The ratio of the number of positive and negative samples is 1:1 to avoid the situation that the loss of the entire model biases to one side due to the imbalance of the input sample ratio.
[0059] Step 2. Perform fine-grained partitioning on the user behavior sequence obtained in Step 1 at the session level. Divide the user behavior sequence into different sessions according to the time interval threshold, and then perform embedding mapping processing on the user behavior sequence in each session to obtain the user basic behavior embedding vector Q embedding , and then add the position bias value between sessions to obtain the user behavior embedding vector Q embedding_pos , and the specific process is as follows:
[0060] Step 2.1. Perform fine-grained partitioning on the user behavior sequence at the session level.
[0061] Divide the user behavior sequence into different sessions according to the time interval threshold. The entire user behavior sequence S is divided into multiple sessions, as specifically shown in formula (1).
[0062]
[0063] Among them, K is the number of sessions into which the user behavior sequence S is divided, T represents the number of behaviors in the session, b i represents the i-th click behavior of the user in the session, d model is the dimension of the behavior embedding vector, and the dimension of the session set Q is
[0064] Step 2.2. Map the user behavior sequence with high-dimensional sparse one-hot encoding into a low-dimensional dense feature vector to obtain the user basic behavior embedding vector Q embedding ;
[0065] Step 2.3. By adding a bias value to the user basic behavior embedding vector Q obtained in Step 2.2, that is, adding the position information in different sessions, to obtain the user behavior embedding vector Q embedding , and the specific addition form is shown in formula (2) and formula (3); embedding_pos Specifically, the addition form is shown in formula (2) and formula (3);
[0066]
[0067] Q embedding_pos = Q embedding + BE (k,t,c)(3)
[0068] Among them, BE (k,t,c) represents the position information corresponding to the c-th dimension of the t-th behavior embedding vector within the k-th session, represents the position information at the session level, represents the position information of each behavior in the session, represents the user's basic behavior embedding vector Q embedding at the dimension level;
[0069] Step 3. Based on the session division result in Step 2, extract the correlation relationship between behaviors in the same session. The specific process is as follows:
[0070] Step 3.1. Evenly distribute the user behavior embedding vector Q with additional position encoding obtained in Step 2 embedding_pos to h single-head structures to obtain respective correlation weight vectors Q k , which can be specifically expressed as Q k = [Q k1 ; …; Q ki ; …; Q kh . The dimension of the single-head attention distribution vector is shown by formula (4).
[0071]
[0072] Then, in each single-head structure, use scaled dot product to calculate the correlation weight vectors {head 1 , head 2 , …, head h} in each subspace. Among them, the output result head i of a certain single-head structure is shown by formula (5). Calculate the similarity matrix of the key vector W K and the query vector W Q , and then multiply it corresponding to the value vector W V to obtain the single-head attention distribution vector.
[0073]
[0074] Among them, d model represents the dimension of Q embedding_pos , T represents the transpose, W i Q , W i V are weight coefficients that need to be trained and learned, and
[0075] Step 3.2. Concatenate the single-head attention distribution vectors obtained in Step 3.1 to obtain the multi-head self-attention distribution vector MultiHead(Q, K, V) with the same dimension as the input feature vector. The multi-head self-attention distribution vector is specifically shown by the calculation formula (6).
[0076] MultiHead(Q,K,V)=Concat(head 1 ,head 2 ,...,head h )W o (6)
[0077] where W o is the coefficient matrix;
[0078] Step 3.3. Use hierarchical normalization to process each layer of data of the multi-head self-attention distribution vector MultiHead(Q, K, V), and normalize the data of the same layer for each batch to reduce overfitting and accelerate the model training to convergence, generating the multi-head attention sublayer S′. Specifically,
[0079]
[0080]
[0081] where μ l ,σ l are the mean and variance of the input sample x i l respectively, x i l is the data of the same layer in the multi-head self-attention distribution vector MultiHead(Q, K, V), z is the simplified expression form of the multi-head self-attention distribution vector MultiHead(Q, K, V), H is the number of elements in x i l , and α and β are the scaling factor and bias term added in the regularization calculation respectively.
[0082] Step 3.4. Through the residual connection of the multi-head attention sublayer S′ obtained in Step 3.3, stack them to form a network for extracting deep cross implicit features and the association relationship between behaviors in the same session, that is, the user multi-dimensional interest vector I k , and the specific calculation formula is as follows. The structural schematic diagram of the session interest interaction layer is as Figure 3 shown.
[0083] I k Q =FFN(S′+Dropout(S′)) (9)
[0084] I k= Avgpooling(I k Q ) (10)
[0085] Wherein, I k Q represents the behavior relationship vector extracted by the user for a certain behavior within the k-th session divided. FFN is a feed-forward neural network, Dropout is to accelerate network training, and Avgpooling is average pooling;
[0086] Step 4. Extract the relationship between sessions and mine the evolution law of user interests. The specific process is as follows:
[0087] Step 4.1. Use the bidirectional long short-term memory method (Bi-LSTM) to extract the interest drift evolution of the user at the session level. Input two long short-term memory neural (LSTM) networks in the forward and reverse orders for feature extraction to obtain the forward candidate hidden layer state and the reverse candidate hidden layer state Finally, concatenate the hidden layer states in both directions to form the inter-session user interest evolution feature H t , and the related calculations are shown in Formulas (13) to (18):
[0088] i t = σ(W xi I t + W hi h t-1 + W ci c t-1 + b i ) (13)
[0089] f t = σ(W xf I t + W hf h t-1 + W cf c t-1 + b f ) (14)
[0090] o t = σ(W xo I t + W ho h t-1 + W co c t-1 + b o ) (15)
[0091] c t = f t c t-1 + i t tanh(Wxc I t + W hc h t-1 + b c ) (16)
[0092] h t = o t tanh(c t ) (17)
[0093]
[0094] Among them, i t is the input gate, o t is the output gate, f t is the forget gate, c t is the cell state, b c is the bias term; the input sequence is the multi-dimensional user interest vector {I 1 , I 2 ,.., I k} obtained in step 3. The dimension size of the weight matrix W xx related to the gate structure calculation is indicated by subscripts, and x is the dimension size of the multi-dimensional user interest vector I k .
[0095] Step 4.2. Reassign the weight scores according to the target item matching degree, and activate the user interests in parts, specifically:
[0096] The first part is to locally activate the current session layer multi-dimensional user interest vector I k extracted by the interest interaction layer to obtain the global expression U I of the multi-dimensional interest vector. The specific calculation is shown in formulas (19) and (20),
[0097]
[0098] U I = ∑ K a k I I k (20)
[0099] Among them, a k I is the normalized weight for scaling the calculation of softmax, and W I is a parameter to be learned during network training;
[0100] The second part is to locally activate the inter-session interest evolution relationship H t mined by bidirectional modeling to obtain the global expression U H, The specific calculation is shown in Formulas (21) and (22).
[0101]
[0102] U H = ∑ K a k H H k (2)
[0103] where a k H is the normalization weight for scaling and calculating softmax, and W H is the parameter to be learned during network training;
[0104] Step 5. Perform prediction and recommendation. The specific process is as follows:
[0105] Step 5.1. Use the concatenated user behavior feature vector as the input of multiple fully connected layers, and perform preliminary prediction based on the target loss function. The target loss function is shown in Formula (23).
[0106]
[0107] where x is a set concatenated by a series of interest vectors, including Q embedding user basic behavior embedding vector, U I session layer user multi-dimensional interest vector, U H interest evolution relationship between sessions. Specifically, x = [Q embedding , U I , U H , D is the training set of size N, y = {0, 1} indicates whether the user is interested in the product, and p(x) is the final output of the model, indicating the degree of interest of the predicted user in the product;
[0108] Step 5.2. Add an auxiliary loss function to correct the prediction results at each time step. The specific structure of the auxiliary loss function is as Figure 4 shown;
[0109] The auxiliary loss function takes the true behavior e b i [t + 1] of the next moment as the positive sample, and the artificially constructed behavior of the next moment as the negative sample. That is, the positive and negative sampling results of the next moment are used as the correction of the current moment loss function L at the same time. The relevant calculations of the auxiliary loss function are shown in Formulas (24) and (25).
[0110]
[0111] L = Ltarget +α×L aux (25)
[0112] α is the proportion for correcting the auxiliary loss function;
[0113] Step 5.3. Finally, perform TopN sorting on the corrected prediction results, and recommend the product with the highest score to the user.
[0114] The present invention is based on two widely used Yoochoose and Diginetica datasets in the session recommendation model, and respectively compares and verifies the effectiveness of different recommendation models in various situations on two different datasets. The relevant experimental results are shown in Table 1.
[0115] Table 1
[0116]
[0117] It can be seen from the table that the MSDN recommendation model designed by the present invention has a high accuracy rate in both datasets.
[0118] As described above, it is only the specific implementation manner of the present invention. Any feature disclosed in this specification, unless specifically described, can be replaced by other equivalent or similar-purpose alternative features; all the features disclosed, or all the steps in any method or process, except for mutually exclusive features and / or steps, can be combined in any way.
Claims
1. A recommendation method for a deep neural network recommendation model based on the multi-head self-attention mechanism, characterized in that, it includes the following steps: Step 1. Obtain user behavior data, then preprocess the user behavior data to obtain a user behavior sequence, and construct an input sample set for the recommendation model; Step 2. Perform fine-grained partitioning on the user behavior sequence obtained in Step 1 at the session level. Divide the user behavior sequence into different sessions according to the time interval threshold, and then perform embedding mapping processing on the user behavior sequence in each session to obtain low-dimensional dense feature vectors , and then add the position bias between sessions to obtain the user behavior embedding vector ; Step 3. Based on the user behavior embedding vectors obtained in Step 2 calculate to obtain a single-head attention distribution vector, and then splice the single-head attention distribution vectors to obtain a multi-head self-attention distribution vector with the same dimension as the user behavior embedding vector , and then perform hierarchical normalization and stacking processing to obtain the user's multi-dimensional interest vectors in the same conversation ; The specific process is: Step 3.
1. Divide the user behavior embedding vectors obtained in Step 2 evenly among multiple single-head structures, and calculate the correlation weight vectors in each subspace using scaled dot product; Step 3.
2. Concatenate each relevance weight vector and convert it into a multi-dimensional relevance weight vector with the same dimension as the user behavior embedding vector ; Step 3.
3. Use hierarchical normalization to normalize the data of the same layer in the multi-dimensional correlation weight vector to obtain a multi-head attention sub-layer; Step 3.
4. Stack the residual connection multi-head self-attention sub-layers to form a network, extract deep cross implicit features, that is, the multi-dimensional interest vectors of users in the same conversation ; Step 4. Calculate the user interest evolution features between sessions Meanwhile, for the user interest evolution features between sessions and the multi-dimensional user interest vectors in the sessions obtained in Step 3 partial activation is performed to obtain the global expression of the user interest evolution features between sessions and the global expression of the multi-dimensional interest vectors in the sessions ; The specific process is: Step 4.
1. Use the bidirectional long short-term memory method to extract the user's interest drift evolution at the session level. Input two long short-term memory neural networks in the forward and reverse orders for feature extraction to obtain the forward candidate hidden layer states and the reverse candidate hidden layer states . Finally, concatenate the hidden layer states in both directions to form the session-level user interest evolution features that combine context information H t ; Step 4.
2. Re - assign the weight scores according to the matching degree of the target project, and activate the user interests in parts. Specifically: The first part is to perform local activation on the multi - dimensional user interest vectors of the current session layer extracted from the interest interaction layer I K to obtain the global expression of the multi - dimensional interest vectors ; The second part is to perform local activation on the user interest evolution relationship between sessions mined by two - way modeling H t to obtain the global expression of the user interest evolution characteristics between sessions ; Step 5. Use the concatenated user behavior feature vectors as the input of multiple fully connected layers, make a preliminary prediction based on the target loss function, then add an auxiliary loss function to correct the prediction results at each time step, and finally perform a TopN sorting on the prediction results to recommend the product with the highest score to the user.
2. The recommendation method according to claim 1, characterized in that, In step 5, the target loss function Specifically, Among them, x is a set spliced by a series of interest vectors, , including user basic behavior embedding vectors, is the global expression of multi-dimensional interest vectors, is the global expression of interest evolution features; D is a training set of size N, y is the true interest degree value of the user, y = {0, 1}, and p(x) is the final output of the model, indicating the predicted interest degree of the user in the commodity; The auxiliary loss function is specifically Among them, is the positive sample at the next moment, is the negative sample at the next moment.
3. A deep neural network recommendation system adopted by the recommendation method according to any one of claims 1 or 2, characterized in that, it includes a behavior data preprocessing module, a session partitioning module, a session interest interaction module, a session interest activation module and a fully connected module; The behavior data preprocessing module is used to obtain the user's historical behavior data and preprocess the historical data to obtain a user behavior sequence and a training data set; The session partitioning module is used to perform fine-grained partitioning of the user behavior sequence obtained after preprocessing at the session level, and divide the user behavior sequence sorted by time order into different sessions according to a time interval threshold. Each session contains user behavior sequences with different time intervals; The session interest interaction module includes a multi-head self-attention sublayer and a multi-head self-attention network formed by residual connections. By extracting the multi-dimensional correlations between the current behavior and other behaviors in the session from multiple angles, and then stacking the multi-head self-attention sublayers, a multi-head self-attention network is constructed to extract the correlation relationships between behaviors within the same session; The session interest activation module adopts a bidirectional long short-term memory structure to capture the interest drift evolution between sessions of the user, and comprehensively extracts interest evolution features from the context; The fully connected module is used to make a prediction based on the user behavior sequence obtained by preprocessing and the interest evolution features output by the session interest activation module, and then perform a TopN sorting on the prediction results to recommend the product with the highest score to the user.
4. The deep neural network recommendation system according to claim 3, characterized in that, the user's historical behavior data is specifically the behavior sequence information accessed by the user extracted from the browser and APP buried point log data of the user, including user registration information, product information, user evaluations and access timestamps.
5. The deep neural network recommendation system according to claim 3, characterized in that, The specific process of preprocessing is as follows: First, the information of the corresponding columns of product numbers, user names, and comment numbers is retrieved from the user historical behavior dataset as needed, and the rating data and product data are associated according to the product numbers; then, an initial user behavior sequence is generated in the order of the corresponding timestamps, and the length of each user interaction sequence is limited to obtain the user behavior sequence; finally, positive and negative samples required for system training are constructed. Among them, the positive samples are the sample data of users participating in comments in the comment dataset, that is, the real user behavior sequence. On the contrary, the negative samples are the sample data that are simulated without actual interaction by the user, that is, the user behavior sequence constructed artificially.
6. The deep neural network recommendation system according to claim 5, characterized in that the number of positive and negative samples is maintained at 1:1 to avoid the situation that the loss of the entire system is biased towards one side due to the imbalance of the input sample ratio.
Citation Information
Patent Citations
Self-attention mechanism-based POI recommendation method integrated with multiple views
CN113886451A
METHOD AND SYSTEM FOR CREATING A PERSONALIZED USER INTEREST PARAMETER FOR IDENTIFYING A PERSONALIZED TARGET CONTENT ELEMENT
RU2017126525A