Method for User Preference Modeling Based on Lightweight Graph Convolutional Attention Network

Through the lightweight graph convolution attention network LightGCAN, combined with lightweight GCN and time-aware GAT, the problems of low user preference modeling efficiency and insufficient dynamic change capture in the prior art are solved, and efficient user preference modeling and recommendation performance improvements are achieved.

CN115438258BActive Publication Date: 2025-07-18HUZHOU UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211044168.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-30
Publication Date
2025-07-18
Estimated Expiration
2042-08-30

AI Technical Summary

Technical Problem

The existing graph neural networks are inefficient in user preference modeling and cannot effectively capture dynamic changes in user preferences. The traditional GCN and GAT models are complex in computing and ignore time factors, resulting in the user preference model being static and unable to reflect dynamic changes.

Method used

Lightweight graph convolution attention network LightGCAN is adopted, combining lightweight GCN and time-aware GAT, and input dual-channel deep neural networks for feature interaction learning and matching score prediction to capture the user's static and dynamic preferences.

Benefits of technology

It realizes efficiently capturing users' static and dynamic preferences in an end-to-end manner, significantly better than existing recommendation methods and improves the performance of recommendation systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115438258B_ABST
    Figure CN115438258B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for user preference modeling based on a lightweight graph convolutional attention network, comprising the following steps: S1. Use a lightweight GCN with only neighborhood aggregation to model static user preferences; S2. Use a time-aware GAT based on the most recent interaction items to model dynamic user preferences; S3. Combine the static user preferences and the dynamic user preferences and input them into a dual-channel deep neural network model for feature interaction learning and matching score prediction. The invention can capture the static and dynamic preferences of users in an end-to-end manner, and can effectively capture the static and dynamic user preferences by using different GNN methods. This method is significantly superior to the current state-of-the-art recommendation methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of user preference modeling and personalized recommendation, and particularly to a method for user preference modeling based on a lightweight graph convolutional attention network.

Background Art

[0002] With the rapid development of computing resources and the availability of large amounts of training data, researchers have started to apply deep learning (DL) techniques to machine learning tasks such as speech recognition, machine translation, and recommendation. Currently, DL models have achieved great success in Euclidean data. However, data generated from non-Euclidean domains (such as graph structures) are now widely used and require additional analysis. For example, in chemistry, molecules are represented as graph-based structures, and their biological activities need to be determined for drug discovery. In e-commerce, the interactions between users and products are modeled as graph structures for accurate product recommendation. A graph is a data format with "nodes" and "edges", where "nodes" represent individuals in the network and "edges" represent the connections between individuals. With the great success of DL, researchers have attempted to design the architecture of GNNs based on the ideas of deep autoencoders, recurrent networks, and convolutional networks. Early studies learned the representation of target nodes by iteratively propagating neighborhood information until a stable state was reached. This learning process is also called graph embedding, which aims to transform graph nodes into low-dimensional vectors by preserving the network topology and node features so that subsequent graph processing tasks (classification, clustering, recommendation, etc.) can be achieved by simple statistical and machine learning methods (dot product, cosine similarity, etc.).

[0003] Graph Neural Network (GNN) is a promising graph data representation learning technology. Graph Convolution Network (GCN) and Graph Attention Network (GAT) are two main representative technologies in GNN, which can learn the embedding representation of the target node by aggregating the embedding representations of adjacent nodes. Thanks to the advantages of GNN in graph representation learning, many GNN-based recommendation technologies have emerged to address different challenges on various graphs. The first method is GCN, which learns user and item embeddings by aggregating information from neighbors in the graph through convolutional and pooling operations. Due to its strong feature extraction and learning capabilities, GCN has been widely applied in remote sensing and achieved great success. Another GNN-based recommendation technology is GAT, which introduces the attention mechanism into GNN to learn the different influence degrees of neighbor nodes on the target node in the interaction graph. Due to its good discrimination ability, GAT has been widely used to build high-performance recommendation models. Although GNN-based recommendation methods have achieved great success, the complex structures of existing GNN methods make them less efficient in complex tasks with multiple graph nodes. The computational processes of GCN and GAT are not very friendly to tasks with complex graph structures. In addition, traditional GCN and GAT ignore the time factor in representation learning, and the obtained user preference model is static and cannot reflect the dynamic changes of user preferences.

Summary of the Invention

[0004] The objective of the present invention is to solve the problems in the prior art and propose a method for user preference modeling based on a lightweight graph convolutional attention network, which can effectively capture static and dynamic user preferences by using different GNN methods.

[0005] To achieve the above objective, the present invention proposes a method for user preference modeling based on a lightweight graph convolutional attention network, including the following steps:

[0006] S1. Use a lightweight GCN with only neighborhood aggregation to model static user preferences;

[0007] S2. Use a time-aware GAT based on the most recent interaction items to model dynamic user preferences;

[0008] S3. Combine the static user preferences and the dynamic user preferences, and input them into a two-channel deep neural network model for feature interaction learning and matching score prediction.

[0009] Preferably, the implementation of this method is based on the Time-Aware Lightweight Graph Convolutional Attention Network (LightGCAN), which includes: an input layer, an embedding layer, a representation layer, an interaction layer, and an output layer; the input layer includes two matrices: the user-item interaction matrix and the interaction time matrix where m and n represent the number of users and items respectively, R is an implicit feedback matrix, if there is an interaction between user u and item i, then r ui = 1, otherwise r ui = 0, T records the interaction time between the user and the item through timestamps, and its dimension is the same as that of R. The input layer provides the initial feature representations of the user and the item and x u 、x i are both multi-hot vectors, corresponding to the u-th row and the i-th column of R respectively; the embedding layer is a fully connected layer, which is used to convert the sparse user and item representations into dense latent embedding representations, and then used as the input of the user preference modeling representation layer; the representation layer includes two GNN models: LightGCN and TGAT, which are used for static and dynamic user preference modeling respectively, combine the obtained static and dynamic user preferences and send them to the interaction layer for high-order feature interaction learning; the interaction layer includes two DNN models: DMF and MLP, which are used to learn different feature interactions according to different deep learning strategies, and finally concatenate the obtained feature interaction vectors and send them to the output layer for predicting the user-item matching score.

[0010] Preferably, the attention elements of the Time-Aware Lightweight Graph Convolutional Attention Network (LightGCAN) include the current user, the target item, the last k interaction items, and the interaction time.

[0011] Preferably, the hyperparameters of the Time-Aware Lightweight Graph Convolutional Attention Network (LightGCAN) include the number of latent factors, the number of last interaction items, and the number of hidden layers in the DL model.

[0012] Preferably, the number of hidden layers of both the DMF and MLP models is set to 3 layers.

[0013] Preferably, in step S1, the lightweight GCN model LightGCN is used for static user preference modeling, and LightGCN only retains the neighbor aggregation operation in GCN, without the two operations of feature transformation and non-linear activation.

[0014] Preferably, in step S1, the modeling process of static user preference includes the following steps:

[0015] S1.1 Lightweight Graph Convolution: The user embedding of the (k+1)-th layer is defined as an aggregation operation based on weighted summation:

[0016]

[0017]

[0018] where, and represent the embedding representations of user u and item i at the k-th layer respectively, N u and N i represent the neighbor nodes of user u and item i respectively; is used as a normalization term; the hyperparameters to be learned are the user and item embeddings of the first layer, and the user and item embeddings of the higher layers are automatically learned layer by layer through the above iterative process;

[0019] The user and item embeddings of the first layer are represented as:

[0020]

[0021]

[0022] where, W u and W v represent the weight matrices for converting the initial feature vectors of users and items into latent embedding representations respectively;

[0023] S1.2 Layer Aggregation: After K layers of lightweight graph convolution operations, K different user / item embeddings are generated, and the embeddings of each layer represent different latent semantic information; the embeddings of each layer are combined to generate the embedding of the target user / item:

[0024]

[0025]

[0026] where, and represent static user preferences and item features respectively; α k ≥0 represents the importance weight of the k-th layer embedding.

[0027] Preferably, in the step S1.2, α k is set to 1 / (K + 1).

[0028] Preferably, in step S2, the modeling process of dynamic user preferences includes the following steps:

[0029] S2.1 Combine the embedding representations of the current user, target item, recent interaction item, and interaction time, and input them into the attention network; the attention network is responsible for learning the importance weights of the recent interaction item for user dynamic preference modeling:

[0030]

[0031]

[0032] Among them, W k and b k represent the weight matrix and bias vector of the k-th layer of the attention network respectively; b k represents the number of layers of the attention network; σ(·) represents the activation function; e u , e v , and represent the embeddings of the current user, target item, recent interaction item, and interaction time respectively, and x j is the combined vector obtained by concatenating the above four embeddings and is used as the input of the attention network;

[0033] S2.2 Divide the time interval between the interaction time of the historical interaction item and the current time, and then obtain the embedding representation of the interaction time through a linear transformation:

[0034] ts j = min((T - t j ) / 60, δ)

[0035]

[0036] Among them, t j represents the interaction time between the current user and item j, T is the current time at the time of prediction, ts j represents the time interval between the interaction time and the prediction time, and the min function is used to set the threshold of the time interval to δ, and W t represents the transformation matrix of the time embedding;

[0037] S2.3 Normalize the attention coefficients through the softmax function:

[0038]

[0039] Among them, RK u represents the recent k interaction items of user u;

[0040] S2.4 The user's dynamic preference vector is modeled as the weighted sum of the embedding representations of the current user's recent k interaction items:

[0041]

[0042] Among them and respectively represent the embeddings of user u and historical interaction item j.

[0043] Preferably, step S3 specifically includes the following steps:

[0044] S3.1 The static and dynamic user preference representations are combined in the form of vector concatenation to obtain the user preference representation; the item embedding vectors obtained from the embedding layer and LightGCN are also combined to obtain the item feature representation:

[0045]

[0046]

[0047] Among them, and respectively represent the static user preference and the item feature, represents the embedding of user u, e u and e i respectively represent the user preference vector and the item feature vector; represents the embedding generated based on item i; the generated user and item representations will be used as the inputs of the high-order feature interaction learning models DMF and MLP;

[0048] S3.2 Feature interaction learning based on DMF: The DMF model has a multi-layer two-channel structure based on user components and item components. In each component, the output of the current layer is used as the input of the next layer; in each layer, the input vector is projected into a hidden vector through linear transformation and non-linear activation operations:

[0049]

[0050]

[0051] Among them, and respectively represent the hidden representations of user u and item i in the k-th layer; here and respectively represent the weight matrix and bias vector of the k-th layer of the user component; and respectively represent the weight matrix and bias vector of the k-th layer of the item component;

[0052] Through the iterative learning of the multi-layer DMF model, the user and item representations are mapped into a low-dimensional latent embedding space:

[0053]

[0054]

[0055] Among them, L1 represents the number of layers of the DMF model, p u and q i respectively represent the latent representations of user u and item i learned;

[0056] The user-item feature interaction is defined as the product of the user and item latent representation vectors:

[0057]

[0058] Among them represents the high-order feature interaction vector learned by the DMF model;

[0059] S3.3 Feature Interaction Learning Based on MLP: MLP is a typical deep learning model. First, the feature vectors of the user and the item are combined, and then multiple hidden layers are passed through to learn high-order user-item feature interactions:

[0060] z0 = [e u ||e i

[0061]

[0062]

[0063] Among them, and α k respectively represent the weight matrix, bias vector, and activation function of the k-th layer. H1 represents the weight matrix of the output layer, and L2 represents the number of layers of the model, represents the high-order feature interaction vector learned by the MLP model;

[0064] S3.4 Matching Score Prediction: First, let DMF and MLP run separately, then combine the output vectors of the two models through vector concatenation, and finally input the combined embedding vector into the output layer of LightGCAN for matching score prediction:

[0065]

[0066] where H2 represents the weight matrix of the output layer;

[0067] S3.5 Model Training: LightGCAN is a CF model based on implicit feedback information. It uses binary cross-entropy loss as the objective function to minimize the difference between the predicted matching score and the implicit feedback information:

[0068]

[0069] Among them,​ is the predicted matching score between user u and item i, r ui is the implicit feedback information observed in the interaction matrix R, R + and R - are the positive sample set and negative sample set respectively, and Θ is the hyperparameter of the model.

[0070] The present invention proposes a time-aware graph convolutional attention network, which effectively captures static and dynamic user preferences by using different GNN methods. Specifically, static user preferences are captured by a lightweight GNN with only node aggregation, and dynamic user preferences are captured based on the time-aware GAT of the most recent interaction items. These two types of user preferences are combined and input into a two-channel DNN model composed of DMF and MLP for feature interaction learning and matching score prediction.

[0071] Advantages of the present invention:

[0072] 1. An efficient user preference learning model is designed, and this framework can capture the static and dynamic preferences of users in an end-to-end manner.

[0073] 2. A time-aware attention network model is proposed, which estimates the contribution weight of each historical interaction item to the dynamic user preference modeling based on the current user, target item, most recent historical interaction items and their interaction times.

[0074] 3. Experiments are conducted on four datasets to evaluate the performance of the method of the present invention in collaborative filtering (CF) recommendation. The experimental results show that the method of the present invention is significantly better than the current state-of-the-art recommendation methods.

[0075] The features and advantages of the present invention will be described in detail through embodiments in conjunction with the accompanying drawings.

Description of the Drawings

[0076] Figure 1 is the overall architecture diagram of the time-aware lightweight graph convolutional attention network LightGCAN;

[0077] Figure 2 is the comparison diagram of HR@10 with different numbers of latent factors;

[0078] Figure 3 is the comparison diagram of NDCG@10 with different numbers of latent factors;

[0079] Figure 4 is the comparison diagram of HR@10 with different historical interaction items;

[0080] Figure 5 is the comparison diagram of NDCG@10 with different historical interaction items;

[0081] Figure 6 is a comparison graph of HR@10 with different numbers of hidden layers;

[0082] Figure 7 is a comparison graph of NDCG@10 with different numbers of hidden layers.

Specific implementation manner

[0083] 1 Basic knowledge

[0084] Some symbols used in this article will be introduced below. Bold italic uppercase letters (such as X) and bold italic lowercase letters (such as x) are used to represent matrices and vectors respectively, and x ij represents the entry in row i and column j of matrix X. The symbols ⊙ and || are used to represent element-wise multiplication and vector concatenation operations respectively. Table 1 summarizes some symbols used in the rest of this article and their descriptions:

[0085] Table 1 Symbol descriptions used in the article

[0086]

[0087] 1.1 Graph Convolutional Network (GCN)

[0088] GCN is a neural network model for graph data structures. Taking graph data as input, each layer learns the graph node representation in a low-dimensional embedding space and uses the output of the previous layer as the input of the next layer. For simplicity, only the implementation details of one layer of GCN will be introduced next.

[0089] The node embedding in the GCN layer includes two main operations: node aggregation and feature transformation. The node embedding process can be abstracted as:

[0090]

[0091] Among them, and represent the embedding representations of the target node i and its adjacent node j respectively; represents the potential embedding d << f of the target node; and represent the node aggregation function and the feature transformation function respectively.

[0092] Node aggregation improves the representation of the target node by collecting information from its adjacent nodes. The basic principle behind node aggregation is that the attributes of the target node can usually be reflected to some extent by the attributes of its neighbors. In recent years. The research on GCN has mainly focused on constructing different node aggregation functions to capture information from the neighborhood. For example: The average pooling function is used to filter out the common attributes of neighbor nodes, and the max pooling function is used to extract representative features from neighbors.

[0093] By mapping the target nodes from the input representation space to the latent embedding space, the feature transformation makes the representation of the target nodes more comprehensive. In traditional GCNs, the feature transformation is usually defined as a process with matrix mapping and a non-linear activation function, abstracted as follows:

[0094]

[0095] where and represent the mapping matrix and the bias vector respectively, and σ is the activation function.

[0096] 1.2 Graph Attention Network (GAT)

[0097] In the embedding learning of target nodes, GAT differentiates the different roles of its neighbor nodes based on the attention mechanism. Similar to GCN, GAT also consists of multiple layers, learning the hidden representations of nodes layer by layer. For simplicity, only one layer of GAT will be introduced in detail next.

[0098] Taking a set of node features as the input of GAT, each layer of GAT generates a new set of node representations with different dimensions as its output:

[0099] x i = σ(∑ j∈neighbor(i) α ij Wx j ) (3)

[0100] where represents the shared weight matrix for feature transformation; f and f' represent the dimensions of the input and output node feature vectors respectively; α ij represents the importance weight of neighbor node j for the representation learning of target node i.

[0101] Generally, the shared self-attention mechanism is used to calculate the influence weights (attention coefficients) of the neighborhood:

[0102] e ij = α(Wx i ,Wx j ) (4)

[0103] To make it more convenient to compare the attention coefficients among different neighbors of the target node, the softmax function is usually used to normalize the attention coefficients:

[0104]

[0105] 2 Lightweight Graph Convolutional Attention Network (LightGCAN)

[0106] 2.1 Overall Framework

[0107] The overall framework of LightGCAN is as Figure 1 shown, consisting of five layers: the input layer, the embedding layer, the representation layer, the interaction layer, and the output layer. The input layer consists of two matrices, namely the user-item interaction matrix and the interaction time matrix where m and n represent the number of users and items respectively. R is an implicit feedback matrix. If there is an interaction between user u and item i, then r ui = 1, otherwise r ui = 0. T records the interaction time between users and items through timestamps, and its dimension is the same as that of R. The input layer provides the initial feature representations of users and items and which are multi-hot vectors corresponding to the u-th row and the i-th column of R respectively. The embedding layer is a fully connected layer used to convert the sparse user and item representations into dense latent embedding representations, which are then used as the input to the representation layer for user preference modeling. The representation layer contains two GNN models, namely LightGCN and TGAT, which are used for static and dynamic user preference modeling respectively. The obtained static and dynamic user preferences are combined and fed into the interaction layer for high-order feature interaction learning. The interaction layer contains two DNN models, namely DMF and MLP, which are used to learn different feature interactions according to different deep learning strategies. Finally, the obtained feature interaction vectors are concatenated and sent to the output layer for predicting the user-item matching scores.

[0108] 2.2 Static User Preference Modeling

[0109] User preferences can generally be divided into two categories: static preferences and dynamic preferences. Static preferences refer to the relatively fixed interests and hobbies formed by users in the long term. Dynamic preferences refer to the short-term preferences of users at the current moment. In CF recommendation, the most common method to capture user static preferences is to use historical interaction information. In the present invention, implicit feedback information is used as the data source for static user preference modeling.

[0110] In a CF recommendation system, the relationship between users and items is usually described as a graph structure. Given the powerful representation learning ability of GCN in graph-structured data, it has been widely used for user and item representation learning based on interaction graphs. In the present invention, a lightweight GCN model, LightGCN, is used for static user preference modeling. LightGCN only retains the neighbor aggregation operation in GCN and discards the two operations of complex feature transformation and non-linear activation that are meaningless for the recommendation task.

[0111] 2.2.1 Lightweight Graph Convolution

[0112] LightGCN is a simplified GCN and a multi - layer model. Formally, the user (item) embedding at the k + 1 layer is defined based on a weighted - sum aggregation operation:

[0113]

[0114]

[0115] where, and respectively represent the embedding representations of user u and item i at the k - th layer, N u and N i respectively represent the neighbor nodes of user u and item i. is used as a normalization term to avoid the explosion phenomenon in feature aggregation. The only hyper - parameter to be learned here is the user and item embeddings in the first layer, because the user and item embeddings in the higher layers can be automatically learned layer - by - layer through the above iterative process.

[0116] Specifically, the user and item embeddings in the first layer are represented as:

[0117]

[0118]

[0119] where W u and W v respectively represent the weight matrices that transform the initial feature vectors of users and items into latent embedding representations.

[0120] 2.2.2 Layer Aggregation

[0121] After K - layer lightweight graph convolutional operations, K different user (item) embeddings are generated. The embeddings of each layer represent different latent semantic information. The intuitive idea is to combine the embeddings of each layer to generate the embedding of the target user (item):

[0122]

[0123]

[0124] where, and respectively represent the static user preference and item features; α k ≥0 represents the importance weight of the k - th layer embedding. In our experiments, α k is uniformly set to 1 / (K + 1), which achieves good performance. This setting strategy can avoid complicating LightGCN while maintaining its simplicity and efficiency.

[0125] 2.3 Dynamic User Preference Modeling

[0126] Static user preference modeling is based on the entire interaction history of the target user and completely ignores the drift of user preferences over time. When a user faces different items at different times, his interests and preferences are different, and this kind of preference belongs to short-term preference. To capture short-term user preferences, one should rely on the items that the user has interacted with recently, which can reflect the user's current interests and preferences better than the items that the user interacted with a long time ago. In addition, the interaction time between the user and the item should be considered in dynamic user preference modeling, which can reveal the drift of user preferences over time.

[0127] In this paper, we propose a time-aware GAT model (TGAT) to capture dynamic user preferences. Different from existing attention networks that model dynamic user preferences based on the entire interaction history or only use the current session, TGAT uses the user's most recent k interaction items to model the user's dynamic preferences. The items interacted with recently can reflect the user's current interests and preferences better than the items interacted with a long time ago. In addition, using fewer items can reduce the computational complexity of user preference modeling. To make the lengths of the historical interaction vectors of all users equal, when the number of the user's historical interaction items is less than k, a meaningless constant (such as -1) is usually used to pad its historical interaction vector.

[0128] First, the embedding representations of the current user, the target item, the most recent interaction items, and the interaction time are combined and input into the attention network. The attention network is responsible for learning the importance weights of the most recent interaction items for modeling the user's dynamic preferences:

[0129]

[0130]

[0131] where, W k and b k represent the weight matrix and the bias vector of the k-th layer of the attention network respectively; b k represents the number of layers of the attention network; σ(·) represents the activation function; e u , e v , and represent the embeddings of the current user, the target item, the most recent interaction items, and the interaction time respectively, and x j is the combined vector obtained by concatenating the above four embeddings and is used as the input of the attention network.

[0132] The time interval between the interaction time of the historical interaction items and the current time is divided in minutes to discretize the continuous time, and then the embedding representation of the interaction time is obtained through a linear transformation:

[0133] ts j = min((T - t j ) / 60, δ) (14)

[0134]

[0135] where t j represents the interaction time between the current user and item j (in seconds), T is the current time at the time of prediction, and ts j represents the time interval (in minutes) between the interaction time and the prediction time. The min function is used to set the threshold of the time interval to δ, and W t represents the transformation matrix of the time embedding.

[0136] Then, to facilitate comparison between different historical interaction items of the user, the attention coefficients are normalized through the softmax function:

[0137]

[0138] where RK u represents the last k interaction items of user u.

[0139] Finally, the user's dynamic preference vector is modeled as the weighted sum of the embedding representations of the last k interaction items of the current user:

[0140]

[0141] where and represent the embeddings of user u and historical interaction item j, respectively.

[0142] 2.4 Feature Interaction Learning and Score Prediction

[0143] The static and dynamic user preference representations are combined in the form of vector concatenation to obtain the user preference representation. In addition, the item embedding vectors obtained from the embedding layer and LightGCN are also combined to obtain the item feature representation:

[0144]

[0145]

[0146] where, and represent the static user preference and item feature, respectively, represents the embedding of user u, e u and e i represent the user preference vector and item feature vector, respectively; Denote the embedding based on item i; the generated user and item representations will be used as the inputs of the high-order feature interaction learning models DMF and MLP.

[0147] 2.4.1 Feature Interaction Learning Based on DMF

[0148] The DMF model is a multi-layer and two-channel structure based on user components and item components. In each component, the output of the current layer is used as the input of the next layer. In each layer, the input vector is projected into a hidden vector through linear transformation and non-linear activation operations:

[0149]

[0150]

[0151] where and respectively denote the hidden representations of user u and item i in the k-th layer; here and respectively denote the weight matrix and bias vector of the k-th layer of the user component; and denote the weight matrix and bias vector of the k-th layer of the item component respectively.

[0152] Through the iterative learning of the multi-layer DMF model, the user and item representations are mapped into a low-dimensional latent embedding space:

[0153]

[0154]

[0155] where L1 represents the number of layers of the DMF model, p u and q i respectively denote the learned latent representations of user u and item i.

[0156] The user-item feature interaction is defined as the product of the user and item latent representation vectors:

[0157]

[0158] where denotes the high-order feature interaction vector learned by the DMF model.

[0159] 2.4.2 Feature Interaction Learning Based on MLP

[0160] DMF uses two independent channels to learn the latent representations of users and items respectively. Finally, it is intuitive to concatenate the latent representations of users and items. However, simply using vector concatenation cannot well describe the interaction between user and item latent factors. Here, another deep learning model MLP will be utilized. First, the feature vectors of users and items are combined, and then multiple hidden layers are added on it to learn high-order user-item feature interactions:

[0161] Z0 = [e u ||e i (25)

[0162]

[0163]

[0164] Among them, and α k represent the weight matrix, bias vector, and activation function of the k-th layer respectively. H1 represents the weight matrix of the output layer, L2 represents the number of layers of the model, represents the high-order feature interaction vector learned by the MLP model.

[0165] 2.4.3 Matching Score Prediction

[0166] So far, two types of high-order feature interaction vectors have been obtained. The DMF model uses a two-channel structure to model the latent representations of users and items, and then calculates the interaction vector. The MLP model first integrates the user and item representations, and then uses a typical DNN model to learn the interaction vector. To retain the advantages of both models, DMF and MLP can be fused into an integrated model. To provide great flexibility to the integrated model, first let DMF and MLP run separately, then combine the outputs of the two models through vector concatenation, and finally input the combined embedding vector into the output layer of LightGCAN for matching score prediction:

[0167]

[0168] Among them, H2 represents the weight matrix of the output layer.

[0169] 2.4.4 Model Training

[0170] The point-wise and pair-wise objective functions are usually used for model training of recommendation systems. For simplicity, the point-wise method is used in the present invention. Since LightGCAN is a CF model based on implicit feedback information, the binary cross-entropy loss is used here as the objective function to minimize the difference between the predicted matching score and the implicit feedback information:

[0171]

[0172] Among them, is the predicted matching score between user u and item i, and r ui is the implicit feedback information observed in the interaction matrix R, and R + and R - are the positive sample set and negative sample set respectively, and Θ is the hyperparameter of the model.

[0173] 3 Experiments and Analysis

[0174] To answer the following research questions, a series of experiments will be conducted below, and the experimental results will be analyzed in detail:

[0175] RQ1. In the Top-k recommendation task, is the proposed recommendation model LightGCAN better than the existing CF recommendation models?

[0176] RQ2. Do different components in LightGCAN play a role in the recommendation task?

[0177] RQ3. Do different settings of the hyperparameters in LightGCAN affect the recommendation performance?

[0178] 3.1 Experimental Settings

[0179] 3.1.1 Datasets

[0180] To evaluate the recommendation performance of the model LightGCAN, experiments are conducted using four real-world datasets in different domains and of different scales, namely MovieLens 100K (ml-100k), Movielen 1M (ml-1m), Amazonmusic (Amusic), and Amazon toys (Atoy). Table 2 gives the detailed statistical information of the experimental datasets:

[0181] Table 2 Statistical Information of Datasets

[0182]

[0183] 3.1.2 Evaluation Strategies and Performance Metrics

[0184] In this experiment, the leave-one-out strategy was used for performance evaluation. The most recent interaction item of each user was used as the test sample, and the remaining interaction items were used for training. Since it is very time-consuming to sort all items in each test, 100 non-interaction items were randomly selected for each user, and then they were sorted together with the test items according to the predicted matching scores. In the experiment, two popular test metrics, Hit Rate (HR) and Normalized Discounted Cumulative Gain (NDCG), were used to evaluate the ranking performance. The length of the ranking list was set to 10 to evaluate the Top-10 recommendation performance.

[0185] 3.1.3 Comparative Methods

[0186] The following CF methods were selected for comparison with the method of the present invention:

[0187] · ItemPop is a statistics-based method that ranks items according to the popularity (access frequency) of the items. It is usually used as a baseline method for CF recommendations.

[0188] · eALS (Xiangnan He, Hanwang Zhang, Min-Yen Kan, and Tat-Seng Chua, 2016. Fast matrix factorization for online recommendation with implicit feedback. In SIGIR. 549–558.) is a CF method based on Matrix Factorization (MF) with fast parameter learning techniques. It takes the observed and unobserved interaction items as positive and negative samples respectively.

[0189] · DMF (Hongjian Xue, Xinyu Dai, Jianbing Zhang, Shuijian Huang, and Jiajun Chen, 2017. Deep matrix factorization models for recommender systems. In IJCAI. 3203–3209.) is a CF method for representation learning based on DL. It uses normalized binary cross-entropy as the loss function for model training.

[0190] · NeuMF (Xiangnan He, Lizi Liao, Hanwang Zhang, Liqiang Nie, Xia Hu, and Tat-Seng Chua, 2017. Neural collaborative filtering. In WWW. 173–182.) is a CF method that predicts matching scores based on DL. It learns feature interactions using different learning strategies based on two major models, GMF and MLP. The outputs of the two models are combined for matching score prediction.

[0191] · DeepCF (Zhihong Deng, Ling Huang, Changdong Wang, Jianhuang Lai, and Philip S. Yu, 2019. DeepCF: a unified framework of representation learning and matching function learning in recommender system. In AAAI. 61–68.) is the current state-of-the-art DL-based CF method. It uses the DMF model to learn user and item embedding representations and the MLP model to learn user-item high-order feature interactions.

[0192] 3.2 Recommendation Performance Comparison (RQ1)

[0193] Table 3 shows the Top-10 recommendation performance of different recommendation methods. Since there is no time information in the Amusic and Atoy datasets, the tests on these two datasets do not include the time embedding component. The following results can be observed from the table:

[0194] · All methods perform best on the ml-100k dataset, followed by the ml-1m dataset, then the Amusic dataset, and finally the Atoy dataset. This result indicates that the sparsity of the training data has a great impact on the recommendation performance of CF methods because the data sparsity of the above datasets increases in turn.

[0195] · The MF-based method eALS outperforms the statistic-based method ItemPop. This phenomenon shows that although the principle of the latent factor model is simple and it only has linear modeling ability, it plays a certain role in feature interaction learning and CF recommendation.

[0196] · The DL-based MF methods (DMF and NeuMF) outperform the traditional MF method eALS, which indicates that deep learning has strong advantages in both representation learning and matching function learning, which is very beneficial for CF.

[0197] · MF-based ensemble methods (NeuMF and DeepCF) outperform MF-based standalone methods (DMF) on all datasets. This is because each component in the ensemble model makes predictions according to different strategies, and when combined, they can maintain the advantages of each component, making the model more robust.

[0198] · Among DL-based methods, DeepCF outperforms NeuMF. Analysis reveals that DeepCF uses two different deep learning strategies for feature interaction learning, while NeuMF uses a linear model and a DL model. This further shows that deep learning models are significantly superior to linear models in feature interaction learning.

[0199] · The method of the present invention, LightGCAN, performs best on all datasets. Compared with the comparative methods, the performance improvement is significant and stable. On the ml-100k dataset, the HR and NDCG of LightGCAN are 6.3% and 9.7% higher than those of the current state-of-the-art DeepCF method respectively; on the ml-1m dataset, the HR and NDCG are 4.7% and 16.1% higher respectively.

[0200] Table 3 Performance comparison of different recommendation methods

[0201]

[0202] 3.3 Ablation experiment analysis (RQ2)

[0203] The method LightGCAN proposed in the present invention consists of two key parts: user static preference modeling and user dynamic preference modeling. To verify the roles of these two parts in the CF recommendation task, ablation experiments were conducted. Recommendations were made by only using static modeling, only using dynamic modeling, and integrating static modeling and dynamic modeling. The experimental results are shown in Table 4. It can be seen from the table that the recommendation performance obtained based on dynamic modeling is better than that based on static modeling, which indicates that dynamic preference modeling considering time factors can better capture users' current preferences. However, if time information is not available, dynamic modeling will not work, and only static modeling can be used to obtain long-term user preferences. The data source of static modeling is easily available. Therefore, static modeling is a general modeling method in CF recommendation, which is why static and dynamic modeling are combined in the model of the present invention. The experimental results show that the combination of the two modeling methods can achieve the best recommendation performance.

[0204] Table 4 Performance comparison of different user modeling methods

[0205]

[0206] In LightGCAN, a novel attention network for dynamic user preference modeling is designed, which comprehensively considers multiple attention elements, including the current user, the target item, the last k interaction items, and the interaction time. To prove the effectiveness of different attention elements in dynamic user preference modeling, experiments are conducted on different combinations of these elements. Since traditional methods only use historical interaction items for user preference modeling, the other three elements should be combined with historical interaction items to view the results. The experimental results are shown in Table 5. It can be seen from the table that all attention elements have a positive impact on dynamic user preference modeling, and the best performance can be obtained by combining these factors. Among them, the interaction time plays the most crucial role, but existing methods often ignore this point, which is the novelty of the attention mechanism proposed in this invention.

[0207] Table 5 Performance Comparison of Different Attention Mechanisms

[0208]

[0209] 3.4 Parameter Sensitivity Analysis (RQ3)

[0210] Next, the impact of different settings of the model hyperparameters on the recommendation performance will be studied. The hyperparameters of the model in this invention include the number of latent factors, the number of last interaction items, and the number of hidden layers in the DL model. Figure 2 and Figure 3 show the HR and NDCG performance with different numbers of latent factors in user (item) representation learning. It can be seen from the figure that as the number of latent factors increases, the recommendation performance becomes better and better. This result indicates that by considering more influencing factors in representation learning, the model of this invention can obtain better performance. However, considering more factors will inevitably lead to higher computational complexity. Therefore, the common practice is not to increase the number of latent factors when the performance improvement is not significant. In this invention, the number of latent factors in the user and item embedding representations is set to 64.

[0211] Figure 4 and Figure 5 show the HR and NDCG with different numbers of historical interaction items during the process of dynamic user preference modeling on the ml-100k dataset. It can be seen from the figure that the most suitable number of historical interaction items on the ml-100k dataset is 20. Considering too many or too few interaction items will lead to performance degradation. Because when too few historical interaction items are used, there is not enough information to learn user preferences, while using too many historical interaction items will inevitably introduce noise. After all, items interacted with a long time ago cannot reflect the user's current interest preferences. On different datasets, the optimal number of historical interaction items is different and needs to be obtained through experiments.

[0212] Generally, the number of hidden layers of a DL model has an important impact on model prediction. To show the impact of the depth of the DL model on representation learning and feature interaction learning, the performance of LightGCAN with different hidden layers was evaluated. Figure 6 and Figure 7 respectively show the HR and NDCG performance of DMF and MLP with different numbers of hidden layers. It can be seen from the figure that at first, as the number of hidden layers increases, the performance gets better and better; when the number of hidden layers reaches three, the performance begins to decline. This result indicates that when there are three hidden layers, the learning ability of the model has reached saturation, and adding more hidden layers will lead to overfitting. Therefore, the number of hidden layers of the DMF and MLP models in the present invention is set to 3 layers.

[0213] 4 Conclusions

[0214] As is well known, CF recommendation includes two stages: representation learning and matching function learning. Representation learning has experienced a process from matrix factorization to deep learning, and matching function learning has experienced a process from dot product, factorization to in-depth learning. For matching function learning, the currently best-performing strategy is the double-path DL-based strategy, so this structure is continued to be used in the present invention. Modeling static and dynamic user preferences is the innovation of the present invention. In the present invention, the possibility of combining long-term and short-term user preferences for user representation learning and CF recommendation is explored. A new user representation learning method is proposed, which includes two parts: static and dynamic user preference modeling. In the static user preference modeling stage, a lightweight GCN is used to extract long-term user preferences; in the dynamic user preference modeling stage, a time-aware GAT is used to model short-term user preferences. The long-term and short-term user preferences are combined and sent to a two-channel DL model for feature interaction learning and matching score prediction. The experimental results on four datasets show that the method of the present invention is significantly better than the existing CF recommendation methods.

[0215] The above embodiments are illustrative of the present invention and not restrictive thereof. Any simple transformation of the present invention falls within the protection scope of the present invention.

Claims

1. A method for user preference modeling based on a lightweight graph convolutional attention network, characterized in that: It includes the following steps: S1. Use a lightweight GCN with only neighborhood aggregation to model static user preferences; S2. Use a time-aware GAT based on the most recent interaction items to model dynamic user preferences; The process of modeling dynamic user preferences includes the following steps: S2.1 Combine the embedding representations of the current user, target item, most recent interaction items, and interaction time, and input them into the attention network; the attention network is responsible for learning the importance weights of the most recent interaction items for modeling the dynamic preferences of the user: Among them, W k and b k represent the weight matrix and the bias vector of the k-th layer of the attention network respectively; b k represents the number of layers of the attention network; σ(·) represents the activation function; and represent the embeddings of the current user, the target item, the most recent interaction item, and the interaction time respectively, and x j is a combined vector obtained by concatenating the above four embeddings and is used as the input of the attention network; S2.2 Divide the time interval between the interaction time of the historical interaction items and the current time, and then obtain the embedding representation of the interaction time through a linear transformation: ts j = min((T - t j ) / 60, δ) Among them, t j represents the interaction time between the current user and project j, T is the current time at the time of prediction, and ts j represents the time interval between the interaction time and the prediction time. The min function is used to set the threshold of the time interval to δ, and W t represents the transformation matrix of the time embedding; S2.3 Normalize the attention coefficients through the softmax function: Among them, RK u represents the last k interaction items of user u; S2.4 The dynamic user preference vector is modeled as the weighted sum of the embedding representations of the user's k most recent interaction items: where and represent the embeddings of user u and historical interaction item j, respectively S3. Combine the static user preferences and dynamic user preferences, input them into a two-channel deep neural network model, and perform feature interaction learning and matching score prediction; the implementation of this method is based on the time-aware lightweight graph convolutional attention network LightGCAN, including: an input layer, an embedding layer, a representation layer, an interaction layer, and an output layer; The input layer includes two matrices: the user-item interaction matrix and the interaction time matrix where m and n represent the number of users and items respectively, R is an implicit feedback matrix, and if there is an interaction between user u and item i, then r ui = 1, otherwise r ui = 0. T records the interaction time between the user and the item through timestamps, and its dimension is the same as that of R. The input layer provides the initial feature representations of the user and the item and x u and x i are both multi-hot vectors, corresponding to the u-th row and the i-th column of R respectively; The embedding layer is a fully connected layer, which is used to convert the sparse user and item representations into dense latent embedding representations, and then used as the input of the user preference modeling representation layer; The representation layer includes two GNN models: LightGCN and TGAT, which are used for static and dynamic user preference modeling respectively. Combine the obtained static and dynamic user preferences and send them to the interaction layer for high-order feature interaction learning; The interaction layer includes two DNN models: DMF and MLP, which are used to learn different feature interactions according to different deep learning strategies. Finally, concatenate the obtained feature interaction vectors and send them to the output layer for predicting the user-item matching score; Step S3 specifically includes the following steps: S3.1 Combine the static and dynamic user preference representations in the form of vector concatenation to obtain the user preference representation; also combine the item embedding vectors obtained from the embedding layer and LightGCN to obtain the item feature representation: Among them, and respectively represent static user preferences and item features, represents the embedding of user u, e u and e i respectively represent the user preference vector and the item feature vector; represents the embedding based on item i; the generated user and item representations will be used as inputs for the high-order feature interaction learning models DMF and MLP; S3.2 Feature interaction learning based on DMF: The DMF model has a multi-layer two-channel structure based on user components and item components. In each component, the output of the current layer is used as the input of the next layer; in each layer, project the input vector into a hidden vector through a linear transformation and a non-linear activation operation: Among them, and respectively represent the hidden representations of user u and item i in the k-th layer; here and respectively represent the weight matrix and bias vector of the k-th layer of the user component; and represent the weight matrix and bias vector of the k-th layer of the item component respectively; Through the iterative learning of the multi-layer DMF model, the user and item representations are mapped to a low-dimensional latent embedding space: where, L1 represents the number of layers of the DMF model, p u and q i respectively represent the learned latent representations of user u and item i; The user-item feature interaction is defined as the product of the user and item latent representation vectors: Among them represents the high-order feature interaction vector learned by the DMF model; S3.3 Feature interaction learning based on MLP: MLP is a typical deep learning model. First, combine the feature vectors of the user and item, and then learn high-order user-item feature interactions through multiple hidden layers on it: z0 = [e u ||e i ​ Among them, and α k respectively represent the weight matrix, bias vector, and activation function of the k-th layer. H1 represents the weight matrix of the output layer, and L2 represents the number of layers of the model. represents the high-order feature interaction vector learned by the MLP model; S3.4 Matching Score Prediction: First, let DMF and MLP run separately, then combine the output vectors of the two models through vector concatenation, and finally input the combined embedded vector into the output layer of LightGCAN for matching score prediction: where H2 represents the weight matrix of the output layer; S3.5 Model Training: LightGCAN is a CF model based on implicit feedback information, using binary cross-entropy loss as the objective function to minimize the difference between the predicted matching score and the implicit feedback information: Among them, is the predicted matching score between user u and item i, r ui is the implicit feedback information observed in the interaction matrix R, R + and R - are the positive sample set and the negative sample set respectively, and Θ is the hyperparameter of the model.

2. The method for user preference modeling based on a lightweight graph convolutional attention network according to claim 1, characterized in that: The attention elements of the time-aware lightweight graph convolutional attention network LightGCAN include the current user, the target item, the last k interaction items, and the interaction time.

3. The method for user preference modeling based on a lightweight graph convolutional attention network according to claim 1, characterized in that: The hyperparameters of the time-aware lightweight graph convolutional attention network LightGCAN include the number of latent factors, the number of last interaction items, and the number of hidden layers in the DL model.

4. The method for user preference modeling based on a lightweight graph convolutional attention network according to claim 1, characterized in that: The number of hidden layers of both the DMF and MLP models is set to 3.

5. The method for user preference modeling based on a lightweight graph convolutional attention network according to claim 1, characterized in that: In step S1, a lightweight GCN model LightGCN is used for static user preference modeling. LightGCN only retains the neighbor aggregation operation in GCN, without the two operations of feature transformation and non-linear activation.

6. The method for user preference modeling based on a lightweight graph convolutional attention network according to claim 5, wherein: In step S1, the modeling process of static user preference includes the following steps: S1.1 Lightweight Graph Convolution: The user embedding of the (k + 1)-th layer is defined as an aggregation operation based on weighted summation: Among them, and respectively represent the embedding representations of user u and item i at the k-th layer. N u and N i respectively represent the neighbor nodes of user u and item i; is used as a normalization term; the hyperparameters to be learned are the user and item embeddings of the first layer, and the user and item embeddings of the higher layers are automatically learned layer by layer through the above iterative process; The user and item embeddings of the first layer are represented as: Among them, W u and W v respectively represent the weight matrices for converting the initial feature vectors of users and items into latent embedding representations; S1.2 Layer Aggregation: After K layers of lightweight graph convolution operations, K different user / item embeddings are generated, and the embeddings of each layer represent different latent semantic information; combine the embeddings of each layer to generate the embedding of the target user / item: Among them, and represent static user preferences and item characteristics respectively; α k ≥ 0 represents the importance weight of the k-th layer embedding.

7. The method for user preference modeling based on a lightweight graph convolutional attention network according to claim 6, wherein: In the step S1.2, set α k to 1 / (K + 1).

Citation Information

Patent Citations

  • Collaborative filtering recommendation method based on enhanced graph learning

    CN112905894A

  • Mobile application recommendation method based on lightweight graph convolutional network

    CN113688974A