Neural Network Recommendation Method and System Based on Side Information and Attention Mechanism

The neural network-based recommendation system effectively addresses the limitations of existing methods by using edge information and attention mechanisms to extract and fuse user and item features, enhancing prediction accuracy and personalization.

CN115408605BActive Publication Date: 2025-07-15Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210945721.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-08
Publication Date
2025-07-15
Estimated Expiration
2042-08-08

AI Technical Summary

Technical Problem

When the existing recommendation method based on heterogeneous information networks makes full use of edge information for recommendation, it faces the problems of incomplete feature extraction and inaccurate user preference characterization, resulting in limited recommendation performance.

Method used

Using a neural network recommendation method based on edge information and attention mechanism, by constructing a heterogeneous information network, the hidden features, attribute features and relational features of users and projects are extracted, and these features are fused with attention network for scoring prediction.

Benefits of technology

Efficiently extracting and fusing user and project features in heterogeneous information networks improves recommendation performance, especially on sparse datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115408605B_ABST
    Figure CN115408605B_ABST
Patent Text Reader

Abstract

The present invention discloses a neural network recommendation method and system based on side information and attention mechanism. The method includes: constructing a heterogeneous information network from multi-source data; the data includes user-item rating data and side information; extracting the latent features, attribute features and relationship features of users and items in the heterogeneous information network; wherein the latent features are dense feature representations obtained by embedding the numbers of users and items, the attribute features are learned from the attribute matrices of users and items through a multi-layer perceptron network, and the relationship features are learned from the heterogeneous information network through meta-path-based network embedding; predicting the ratings of users for items based on the extracted latent features, attribute features and relationship features of users and items, and making recommendations based on the predicted ratings. The present invention can effectively extract and fuse the features of users and items in the heterogeneous information network to improve the recommendation performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of personalized recommendation, and particularly to a neural network recommendation method and system based on side information and attention mechanism. Background Art

[0002] With the explosive growth of Internet information, it has become increasingly difficult for users to obtain personalized information. As an indispensable tool for information filtering, the recommendation system can quickly and accurately find out the content that users may be interested in from a large amount of data. It is no exaggeration to say that today's recommendation systems are ubiquitous on the Internet and play a crucial role in many user-oriented online services, such as music (Last.fm and Spotify), video (Netflix and YouTube), news (Bing and Toutiao), and e-commerce (Amazon and Taobao), etc. Due to its outstanding performance in improving the user experience of network applications, increasing network service traffic and revenue, the recommendation system has attracted extensive attention and interest from the academic and industrial circles.

[0003] In the past few decades, researchers have proposed a large number of recommendation algorithms. Among these existing recommendation technologies, collaborative filtering is the most prominent method and has been widely tried in the industrial community. Specifically, collaborative filtering builds a user preference model by using the user's past behaviors, that is, the interaction between the user and the item, for personalized recommendation. The user-item interaction is usually represented as a matrix, where one row corresponds to a user, one column corresponds to an item (such as music, video, news, and products), and the value in the matrix represents the user's preference for the item. Among them, the user preference can be implicit feedback (click, browse, and purchase), or explicit feedback (rating). According to the type of user-item interaction, the recommendation task can be divided into two categories, namely Top-N recommendation (implicit feedback) and rating prediction (explicit feedback). The task of Top-N recommendation is to generate a ranked list of items that the user is interested in, while the rating prediction task aims to predict the user's rating for the item. The present invention mainly studies the collaborative filtering algorithm for rating prediction, which is an important task in many real-world recommendation scenarios such as movie recommendation and product recommendation. After obtaining the predicted rating of the user for the item, personalized recommendation can be achieved according to the rating level.

[0004] Traditional rating prediction algorithms only use users' historical ratings of items to model user preferences for personalized recommendation. Among them, matrix factorization achieved great success in the Netflix Prize competition. Matrix factorization first decomposes the user-item rating matrix into two low-dimensional user latent matrices and item latent matrices, and then uses the inner product between the user latent vector and the item latent vector to predict unknown ratings. However, due to data sparsity and cold start problems, the recommendation performance of matrix factorization algorithms is limited. To alleviate these problems, social recommendation introduces social relationships between users when building a user preference model. However, simply using social relationships cannot properly characterize the similarity between users. With the rapid development of network technology, in addition to social relationships, more and more side information can be obtained through network services, such as user attributes, item attributes, text descriptions, etc. To use this side information for recommendation, it is usually modeled as a heterogeneous information network. As a powerful data modeling tool, the Heterogeneous Information Network (HIN) can effectively model the heterogeneity and complexity of side information. Such recommendation methods are called heterogeneous information network-based recommendation.

[0005] Recently, heterogeneous information network-based recommendation systems have attracted much attention due to their good recommendation performance. Since a meta-path describes how to connect two entity types through a specific semantic path, it is usually used to describe the semantic relevance between users and items in heterogeneous information network-based recommendation. Some early methods use the semantic similarity between users and items based on meta-paths for recommendation. However, these methods may suffer from the problems of sparse path connections and noise. Another class of methods uses meta-path-based network embedding to extract useful information from the heterogeneous information network for recommendation. Since this class of methods samples the structural and semantic patterns in the heterogeneous information network through multiple meta-paths, the recommendation performance is greatly affected by the meta-path selection.

[0006] Although heterogeneous information network-based recommendation algorithms have achieved some improvement in recommendation performance to a certain extent, there are still two challenges in making full use of side information for recommendation. First, how to comprehensively extract useful features of users and items from side information. Existing methods only extract meta-path-based relational features from the heterogeneous information network, which may lead to irreversible loss of useful features for recommendation. Second, how to accurately characterize users' preferences for various features of items, and the attractiveness of items to various feature users. Existing methods assume that users have the same preference for the same features of different items, which is usually unreasonable and may lead to incorrect recommendations. For example, users may have different expectations for two similar products with different prices: for the more expensive product, users require all its features to be excellent; however, for the cheaper product, users only focus on its basic functions. Summary of the Invention

[0007] In view of the problem that the recommendation performance of existing recommendation methods based on heterogeneous information networks still needs to be improved, the present invention proposes a neural network recommendation method and system based on side information and attention mechanism, which can effectively extract and fuse the features of users and items in the heterogeneous information network to improve the recommendation performance.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] On the one hand, the present invention proposes a neural network recommendation method based on side information and attention mechanism, including:

[0010] Step 1, constructing a heterogeneous information network from multi-source data; the data includes user-item rating data and side information;

[0011] Step 2, extracting the latent features, attribute features and relationship features of users and items in the heterogeneous information network; wherein the latent features are dense feature representations obtained by embedding the numbers of users and items, the attribute features are learned from the attribute matrices of users and items through a multi-layer perceptron network, and the relationship features are learned from the heterogeneous information network through meta-path-based network embedding;

[0012] Step 3, predicting the ratings of users for items according to the extracted latent features, attribute features and relationship features of users and items, and making recommendations based on the predicted ratings.

[0013] Further, in the step 2, the latent features of users and items in the heterogeneous information network are extracted in the following manner: converting the corresponding numbers of users and items into one-hot encoded binary vectors; mapping the obtained binary vectors into dense user / item embeddings through an embedding layer; the obtained user / item embeddings are the latent features of users / items.

[0014] Further, in the step 2, using meta-path-based network embedding to extract the relationship features of users and items from the heterogeneous information network, including: first, for each meta-path, converting the heterogeneous information network into a corresponding homogeneous information network by using a connectivity matrix; then, using a network embedding algorithm to extract the feature representations of users or items in the homogeneous information network.

[0015] Further, in the step 2, the attribute features of users and items in the heterogeneous information network are extracted in the following manner: converting the attribute values of users and items into one-hot or multi-hot encoded binary vectors; mapping the obtained binary vectors into dense feature representations through a two-layer fully connected network with a ReLU activation function, that is, obtaining a user attribute feature matrix and an item attribute feature matrix.

[0016] Further, step 3 includes: fusing the latent features, attribute features, and relationship features of each user and item through an attention network into a user feature vector and an item feature vector, and then predicting the user's rating for the item using the user feature vector and the item feature vector, and recommending the top n items with the highest predicted ratings to the user.

[0017] On the other hand, the present invention proposes a neural network recommendation system based on side information and attention mechanism, including:

[0018] A heterogeneous information network construction module for constructing multi-source data into a heterogeneous information network; the data includes user-item rating data and side information;

[0019] A feature extraction module for extracting the latent features, attribute features, and relationship features of users and items in the heterogeneous information network; where the latent features are dense feature representations obtained by embedding the numbers of users and items, the attribute features are learned from the attribute matrices of users and items through a multi-layer perceptron network, and the relationship features are learned from the heterogeneous information network through meta-path-based network embedding;

[0020] A predicted rating and recommendation module for predicting the user's rating for the item based on the extracted latent features, attribute features, and relationship features of the user and the item, and making recommendations based on the predicted ratings.

[0021] Further, in the feature extraction module, the latent features of users and items in the heterogeneous information network are extracted in the following manner: converting the corresponding numbers of users and items into one-hot encoded binary vectors; mapping the obtained binary vectors into dense user / item embeddings through an embedding layer; the obtained user / item embeddings are the latent features of the user / item.

[0022] Further, in the feature extraction module, the relationship features of users and items are extracted from the heterogeneous information network using meta-path-based network embedding, including: first, for each meta-path, converting the heterogeneous information network into a corresponding homogeneous information network using a connectivity matrix; then, using a network embedding algorithm to extract the feature representations of users or items in the homogeneous information network.

[0023] Further, in the feature extraction module, the attribute features of users and items in the heterogeneous information network are extracted in the following manner: converting the attribute values of users and items into one-hot or multi-hot encoded binary vectors; mapping the obtained binary vectors into dense feature representations through a two-layer fully connected network with a ReLU activation function, that is, obtaining the user attribute feature matrix and the item attribute feature matrix.

[0024] Furthermore, the prediction scoring and recommendation module is specifically configured to: fuse the latent features, attribute features, and relationship features of each user and item through an attention network into a user feature vector and an item feature vector, and then use the user feature vector and the item feature vector to predict the user's score for the item, and recommend the top n items with the highest predicted scores to the user.

[0025] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0026] The present invention comprehensively describes users and items from three perspectives of latent features, attributes, and relationships. In addition, the present invention proposes a feature fusion method based on the attention mechanism, which can effectively characterize the preferences of users for different items and the attractiveness of items to different users. The present invention can effectively extract and fuse the features of users and items in the heterogeneous information network to improve the recommendation performance.

[0027] The present invention is evaluated and tested on three real datasets of Yelp, Douban Books, and Douban Movies, and the experimental results show that the present invention is superior to the existing rating prediction algorithms. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is an overall framework diagram of a neural network recommendation method based on side information and attention mechanism according to an embodiment of the present invention;

[0029] Figure 2 is a schematic diagram of the network structure of a neural network recommendation method based on side information and attention mechanism according to an embodiment of the present invention;

[0030] Figure 3 is a schematic diagram of the attention network structure of a neural network recommendation method based on side information and attention mechanism according to an embodiment of the present invention;

[0031] Figure 4 is the influence of the feature representation dimension on the recommendation performance;

[0032] Figure 5 is the influence of the prediction vector dimension on the recommendation performance;

[0033] Figure 6 is the influence of the learning rate and regularization coefficient on the recommendation performance;

[0034] Figure 7 is the influence of the regularization coefficient and the number of iterations on the recommendation performance;

[0035] Figure 8 is a schematic diagram of the architecture of a neural network recommendation system based on side information and attention mechanism according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] For ease of understanding, the following explanations are provided for some terms that appear in the specific embodiments of the present invention:

[0037] Table 1 Symbol Explanation

[0038]

[0039]

[0040] Definition 1 (Heterogeneous Information Network) An information network is a directed graph and can be defined as where is a set of entities, ε is a set of connections, is a predefined set of entity types, is a predefined set of connection types, is an entity type mapping function, is a connection type mapping function. Each entity belongs to a specific entity type and at the same time each connection e ∈ ε belongs to a specific connection type If the number of entity types or the number of connection types i.e., then the information network is called a heterogeneous information network (Heterogeneous Information Network, HIN); otherwise, i.e., and the information network is a homogeneous information network, denoted as

[0041] Definition 2 (Meta-Path) Given a heterogeneous information network A meta-path ρ is defined as a path template on the set of entity types and the set of connection types represented in the form of (abbreviated as T1T2…T l+1 ), where l is the length of the meta-path, is a specific entity type, is a specific connection type, i = 1, 2, …, l + 1. The meta-path ρ defines a new composite relationship l+1 between the entity types T1 and T where represents a composite operator for two connection relationships.

[0042] Definition 3 (Network Embedding) Network embedding is also known as network representation learning. Given an information network Network embedding aims to learn a mapping function This function projects each node into a d-dimensional vector space while maximizing the similarity between each node and its neighbors in the original network. The low-dimensional vector x is called the embedding or representation vector of node v, which can effectively capture the structural and semantic information of the network.

[0043] Definition 4 (Scaled Dot-Product Attention) Scaled dot-product attention first calculates the dot product between a query and all keys, and then applies the softmax function to obtain the attention weights for the corresponding values, which can be expressed as

[0044]

[0045] where Q, K, and V represent the matrices of query, key, and value respectively, and d is the dimension of the feature vectors. Each row of these three matrices is a query / key / value, and there is a one-to-one correspondence between keys and values. The attention output is the weighted sum of all values, where the weight assigned to each value is related to the similarity between the corresponding key and query, and a scaling factor is used to avoid too small a gradient of the softmax function.

[0046] Problem Definition (Rating Prediction Based on Side Information) The recommendation problem based on explicit feedback can be represented as a rating prediction task for completing the user-item rating matrix. The rating data and side information can be modeled as a heterogeneous information network Let and represent the sets of users, items, and known ratings respectively, where the triple <u, v, r u,v > means that user rates item as r u,v . Given the set of known ratings and the heterogeneous information network G constructed from side information, the objective of the present invention is to achieve accurate prediction of the unknown ratings in the rating matrix.

[0047] The following further explains and illustrates the present invention with reference to the accompanying drawings and specific embodiments:

[0048] The overall framework of the Neural Attention Recommendation (NARec) method based on side information and attention mechanism proposed by the present invention is as Figure 1 shown. This method consists of three main parts, namely data preprocessing ( Figure 1 (a)), feature extraction ( Figure 1 (b)), and rating prediction ( Figure 1(c)). Among them, the network structures of the last two parts are as Figure 2 shown.

[0049] Data preprocessing aims to construct multi-source data into a heterogeneous information network, as Figure 1 (a) shown. The dataset used in the experiments of the present invention not only contains user-item rating data, but also contains rich side information. Since the heterogeneous information network has various types of entities and connected structures, it can effectively model the heterogeneity and complexity of multi-source data.

[0050] Feature extraction aims to extract the features of users and items in the heterogeneous information network, as Figure 1 (b) and Figure 2 (a) shown. To describe users and items more comprehensively, the present invention characterizes them from three aspects: latent features, attribute features, and relational features. Specifically, the latent features are the dense feature representations obtained by embedding the numbers of users and items; the attribute features are learned from the attribute matrices of users and items through a multi-layer perceptron network; the relational features are learned from the heterogeneous information network through meta-path-based network embedding. The different relational features and attribute features of users / items constitute the feature matrix of users / items.

[0051] Rating prediction aims to predict the ratings of users for items based on the latent features and feature matrices of users and items, as Figure 1 (c) and Figure 2 (b) shown. Specifically, first, the latent features, attribute features, and relational features of each user and item are fused into a user feature vector and an item feature vector through an attention network, as Figure 3 shown; then, these two feature vectors are used to predict the ratings of users for items. Next, the present invention will introduce the feature extraction and rating prediction parts of NARec in detail.

[0052] 1. Feature Extraction of Users and Items

[0053] The heterogeneous information network composed of side information and rating data contains all the information of users and items. How to comprehensively and accurately extract the features of users and items from it is the key to rating prediction. Due to the complex non-Euclidean graph structure of the heterogeneous information network, a typical method is to use meta-path-based heterogeneous information network embedding to extract the features of users and items. A meta-path downsamples the heterogeneous information network, retaining only the features consistent with the structure and semantics of the meta-path while discarding all other features. At the same time, a meta-path can also be regarded as the relationship between users and items in the heterogeneous information network. Considering only the relational features of users and items may cause the irreversible loss of useful features for recommendation. Therefore, the present invention characterizes users and items from three aspects, namely latent features, attribute features, and relational features, as Figure 2(as shown in (a))

[0054] (1.1) Embedding layer

[0055] Suppose there are M users N items For each user and each item Their corresponding numbers are converted into one - hot encoded binary vectors, that is and Since and They are usually of high dimension and very sparse, so they are mapped into dense feature representations through the embedding layer:

[0056]

[0057] Wherein, is the embedding representation of user u, is the embedding representation of item v, d is the dimension of node embedding, d << M and d << N; and represent the embedding matrices of users and items respectively. They are the parameters of the embedding layer and can be learned through the model. The obtained user (item) embeddings can be regarded as the latent features of users (items).

[0058] Table 2 User relationships in the Douban movie dataset

[0059]

[0060] Table 3 Item relationships in the Douban movie dataset

[0061] Project Relationship Description MUM Movies Watched by the Same User MDM Movies by the Same Director MAM Movies Starring the Same Actor MTM Movies of the Same Genre MUGUM Movies Watched by the Same Group of Users

[0062] (1.2) Relationship feature extraction

[0063] Relationship features characterize the similarity between users and the similarity between items. For a specific user, items interacted by other similar users can be recommended, or items similar to those interacted by the user in the past can be recommended. Therefore, relationship features can be used to improve the recommendation performance. Given a heterogeneous information network Meta - paths can be used to represent the relationships between users and items. Taking the Douban movie dataset as an example, the user relationships and item relationships are shown in Table 2 and Table 3 respectively. Use to represent the set of meta - paths in the heterogeneous information network G, where and They are the user meta-path set and the item path set respectively. The present invention extracts relationship features from a heterogeneous information network by using meta-path-based network embedding, which mainly includes two steps: First, for each meta-path, the heterogeneous information network is converted into a corresponding homogeneous information network by using a connectivity matrix; then, a network embedding algorithm is used to extract the feature representation of users or items in the homogeneous information network, and the specific process is as follows.

[0064] Given a meta-path The binary connectivity matrix of this meta-path can be defined as

[0065]

[0066] where, is the adjacency matrix between entity types T i and T j and Bin is a binary function. If x>0, Bin(x)=1, otherwise Bin(x)=0. To construct a homogeneous information network based on , the first and last entity types of the meta-path must be the same, either both users or both items. Denote the homogeneous information network under the meta-path ρ as The binary connectivity matrix is the adjacency matrix of G ρ .

[0067] Several existing network embedding methods can be directly used for node representation learning of homogeneous information networks, such as DeepWalk, LINE, and node2vec. These methods can be formalized as a maximum likelihood optimization problem, that is, maximizing the co-occurrence probability of each node and its adjacent nodes. Under the node independence assumption, the objective function of network embedding can be expressed as

[0068]

[0069] where, is the context of node v in network Gρ generated by a specific neighborhood sampling strategy S. f is a mapping function that maps each node into a d-dimensional vector space, which can be equivalent to a matrix.

[0070] By using the stochastic gradient descent method to optimize the objective function (4), the node representation of the homogeneous information network Gρ can be obtained. If ρ is a user relationship, the user relationship feature matrix can be learned, and its i-th row represents the relationship feature vector of the i-th user u i . If ρ is an item relationship, the item relationship feature matrix can be learned, and its i-th row represents the relationship feature vector of the i-th item v iThe relational feature vectors. For different relationships in the heterogeneous information network G, different user relational feature matrices can be obtained and different item relational feature matrices The user and item relationships for each dataset are shown in Table 4.

[0071] (1.3) Attribute feature extraction

[0072] User attributes are portraits of users, such as age, occupation, gender, etc.; item attributes describe the content characteristics of items, such as category, color, usage, etc. Users with different attributes may have different preferences, and different users may focus on different attributes of items. Therefore, attribute features can be used to improve the performance of recommendations.

[0073] The attribute set is denoted as where and are the attribute sets of users and items respectively. The attribute values of users and items are first converted into binary vectors encoded in one-hot or multi-hot. For any user attribute a one-hot / multi-hot user attribute matrix can be obtained whose i-th row represents the attribute vector of the i-th user u i under the user attribute α k . For any item attribute a one-hot / multi-hot encoded item attribute matrix can be obtained whose j-th row represents the attribute vector of the j-th item v j under the item attribute β t .

[0074] Through a two-layer fully connected network with a ReLU (Rectified Linear Unit) activation function, these sparse attribute vectors and are mapped to dense feature representations, namely

[0075]

[0076]

[0077] where, is the weight matrix of the user attribute α k and the item attribute β t , is the bias vector, is the attribute feature vector of the user u i under the user attribute α k , is for the item vj In the project attribute β t The attribute feature vector below. That is, the sparse attribute vector and are mapped to the dense feature vectors and i = 1, 2, …, M, j = 1, 2, …, N,

[0078] Through formula (5), the user attribute feature matrix under the user attribute α k can be obtained. Each row of it is the attribute feature vector of the corresponding user; through formula (6), the project attribute feature matrix under the project attribute β t can be obtained. Each row of it u is the attribute feature vector of the corresponding project. For different attributes in the attribute set

[0079]

[0080] 2. Rating Prediction Based on Attention Mechanism

[0080] The present invention extracts useful information of users and projects in the heterogeneous information network from three aspects: latent features, attribute features, and relationship features. Specifically, through the embedding layer, the latent feature matrices P and Q of users and projects can be obtained; through heterogeneous information network embedding, the relationship feature matrices of users and projects can be obtained and Through the multi-layer perceptron network, the attribute feature matrices of users and projects can be obtained and This section will study how to use these extracted features for rating prediction. Due to the powerful non-linear modeling ability of neural networks, this section proposes a rating prediction model based on the attention mechanism, as shown in Figure 2 (b).

[0081] (2.1) Feature Fusion

[0082] For the sake of convenience of expression, the relationship feature matrix and the attribute feature matrix of users and projects are uniformly represented as the user feature matrix and the project feature matrix Among them, is the number of user features, including user relationship features and user attribute features; is the number of project features, including project relationship features and project attribute features. For any user the latent feature vector p can be obtainedu , and L u different feature vectors For any project the latent feature vector q can be obtained v , and L v different feature vectors The various feature vectors of user u or project v can be represented as a feature matrix:

[0083]

[0084] wherein, is the user feature matrix, is the project feature matrix.

[0085] Accurately and adaptively characterizing the personalized features of users and projects is crucial for score prediction. On the one hand, users with different features may have different preferences for the same project, and projects with different features may also have different attractions to the same user; on the other hand, a user may have different degrees of emphasis on the same feature of different projects, and a project may also have different attractions to different users with the same feature. Since the interaction between users and projects may be very complex, the present invention introduces an attention mechanism to model them. The attention mechanism has been successfully applied in fields such as computer vision and natural language processing, and its basic idea is that the output depends on the most relevant part in the input. Therefore, the attention mechanism can be used to learn the adaptive personalized feature representations of users and projects. Based on scaled dot-product attention, the present invention designs a user feature attention network and a project feature attention network to fuse various features of users and projects, such as Figure 3 shown.

[0086] Essentially, the attention operation can be regarded as a function that maps a query and a set of key-value pairs to an attention output. For the user feature attention network, the query is the project latent feature vector q v , and the key and value are the linear mappings of the user feature matrix X u , and the output can be expressed as

[0087] x u = Attention(q v , X u W uk , X u W uv ) (8)

[0088] wherein, is the user feature mapping matrix, which makes the attention network more flexible. The attention output Namely, it is the personalized user feature vector, which characterizes the preference of user u for item v by weighting the user preferences under different user features. Similarly, for the item feature attention network, the query is the user latent feature vector p u , the key and value are the linear mappings of the item feature matrix Y v , and the output can be expressed as

[0089] y v =Attention(p u ,Y v W ik ,Y v W iv ) (9)

[0090] Among them, is the item feature mapping matrix. The attention output is the personalized item feature vector, which characterizes the attraction of item v to user u by weighting the item attractions under different item features.

[0091] To comprehensively describe users and items, the latent feature vector and the personalized feature vector are concatenated to represent the features of users and items:

[0092]

[0093] Among them, is the concatenated feature vector of user u, is the concatenated feature vector of item v. Although the feature attention network can fuse multiple features of users and items through adaptive weights, it is still a linear transformation. To endow the network with non-linearity and introduce the interaction between personalized features and latent features, a two-layer fully connected network with ReLU activation function is used for the concatenated feature vector:

[0094]

[0095] Among them, is the weight matrix, is the bias vector, is the fused feature vector of user u, is the fused feature vector of item v.

[0096] (2.2) Rating prediction

[0097] After obtaining the fused feature vectors, the Hadamard product operation is used to introduce the feature interaction between users and items:

[0098]

[0099] Among them, the operator ⊙ represents the Hadamard product, z′u,v is the product feature vector. To introduce the interaction between different feature dimensions and reduce the feature dimension, a two-layer fully connected network with ReLU activation function is applied to the product feature vector:

[0100] z u,v = ReLU(z′ u,v W p1 + b p1 )W p2 + b p2 (13)

[0101] where, and are weight matrices, and are bias vectors. Since the output vector determines the performance of the score prediction, it is called the prediction vector, where D is the dimension of the prediction vector. Finally, the prediction vector is imported into the output layer to obtain the predicted score:

[0102]

[0103] where, is the weight vector, represents the predicted score of user u for item v. To reduce the prediction error, the output is restricted to the range of [1, 5].

[0104] 3. Model Training

[0105] Specifically, in the present invention, a score prediction model is constructed, simply referred to as NARec. The inputs of the model include the numbers of users and items (latent feature matrix), one-hot / multi-hot attribute matrix, and user-item relationship feature matrix. Among them, the relationship features are pre-extracted from the heterogeneous information network through network embedding based on meta-paths. The output of the model is the predicted score. By ranking the candidate items according to the predicted score, the recommendation list of users can be obtained. Formally, the model can be represented as a complex function F:

[0106]

[0107] where, is the one-hot / multi-hot attribute matrix of user u, and its i-th row represents the attribute vector of user u under the i-th user attribute; is the one-hot / multi-hot attribute matrix of item v, and its i-th row represents the attribute vector of item v under the i-th item attribute; is the relationship feature matrix of user u, and its i-th row represents the feature vector of user u under the i-th user relationship; is the relational feature matrix of project v, and its i-th row represents the feature vector of project v under the i-th project relationship; Θ represents all trainable model parameters, including the weight matrices and bias vectors of the embedding layer, fully connected network, attention network, and output layer.

[0108] To learn the model parameters Θ, through a regression framework, to minimize the error between the predicted score and the actual score, the objective function for model optimization is

[0109]

[0110] where r u,v is the true score of user u for project v, Ω(·) represents the regularizer, λ is the regularization coefficient, and l(·) represents the loss function. According to different optimization objectives, mean squared error or mean absolute error can be selected as the loss function. denotes the score training set, which is obtained by dividing the known score set into a training set and a test set.

[0111] The objective function (16) can be optimized by the Adam optimizer, which is a variant of the stochastic gradient descent method using adaptive moment estimation. Specifically, the present invention uses mini-batch Adam with a batch size of 256 to accelerate the training process, adopts L2 regularization to prevent the model from overfitting, and updates the model parameters using backpropagation. The detailed training process of the model is shown in Algorithm 1. Since mainstream deep learning libraries (such as TensorFlow, Keras, PyTorch) can automatically implement the optimization process through automatic differentiation, the calculation of the partial derivatives of the model parameters is omitted here. After learning all the parameters of the score prediction model, the model prediction can be performed through Equation (15).

[0112]

[0113]

[0114] To verify the effectiveness of the present invention, the following experiments are carried out:

[0115] a. Dataset

[0116] The proposed recommendation model NARec is evaluated on publicly available datasets with different sparsities in three different domains (commerce, books, and movies), including the Yelp dataset, the Douban Books dataset, and the Douban Movies dataset. These three datasets are based on explicit feedback in the form of ratings, which are integers from 1 to 5. All three datasets contain rich side information and are widely used for rating prediction tasks. Specifically, for the Yelp dataset, an item refers to a merchant; for the Douban Books dataset, an item refers to a book; and for the Douban Movies dataset, an item refers to a movie.

[0117] The Yelp dataset comes from the Yelp Dataset Challenge, a competition that allows students to explore and conduct research using data provided by Yelp. Yelp is a well-known merchant review website in the United States where users can rate the products or services of merchants. The Douban Books and Movies datasets were crawled from Douban by Ishikawa et al. [Adaptive Deep Modeling of Users and Items Using Side Information for Recommendation]. Douban is a Chinese community website that allows users to create, record, and review content related to movies, books, music, etc.

[0118] All three datasets contain user attributes, item (merchant / book / movie) attributes, and user-item ratings. Their detailed statistics are shown in Table 4, where #A and #(A - B) represent the number of entities of type A and the number of relationships of type A - B in the dataset, respectively, denotes the average degree of entities of type A in the relationship type A - B, calculated as shown in Equation (18). In addition, the ratings of these three datasets have different densities, i.e., Yelp (0.09%) < Douban Books (0.27%) < Douban Movies (0.63%), calculated as shown in Equation (17). Theoretically, the denser the dataset, the more useful information it contains, and thus the higher the accuracy of rating prediction.

[0119]

[0120]

[0121] Table 4 Statistics of the Datasets

[0122]

[0123] b. Comparative Algorithms

[0124] To verify the effectiveness of the method of the present invention, the present invention compares NARec with three types of scoring prediction methods: (1) Matrix factorization-based methods, namely PMF (Probabilistic Matrix Factorization). (2) Deep learning-based methods, including NeuMF (Neural Matrix Factorization) and CFNet (Collaborative Filtering Network). Since these two methods focus on implicit feedback, they use sigmoid as the activation function of the output layer and binary cross-entropy as the loss function. For scoring prediction, the present invention redefines the output unit of these two methods as a linear layer and resets its loss function to mean squared error loss. (3) Heterogeneous information network-based methods, including HERec (HIN Embedding based Recommendation) and HopRec (Outer Product Enhanced HIN embedding for Recommendation). The detailed descriptions of these 5 comparison algorithms are as follows:

[0125] PMF: PMF is a traditional matrix factorization-based recommendation model that models the user-item rating matrix as the product of two low-rank user and item matrices. In this embodiment, the model is implemented based on the Python toolkit Surprise 1 to implement this model.

[0126] NeuMF: NeuMF learns the user-item interaction function by integrating matrix factorization and multi-layer perceptron under the neural network collaborative filtering framework. In the experiment of the present invention, the code provided by the author 2 is used to implement this model.

[0127] CFNet: CFNet is a collaborative filtering model based on multi-layer perceptron that combines representation learning and matching function learning for recommendation. In the experiment of the present invention, the code provided by the author 3 is used to implement this model.

[0128] HERec: HERec is a heterogeneous information network-based recommendation method that uses heterogeneous information network embedding and extended matrix factorization to learn and fuse user / item embeddings for scoring prediction. In the experiment of the present invention, the code provided by the author 4 is used to implement this model.

[0129] HopRec: HopRec is an improved version of HERec that models the pairwise relationship between user embeddings and item embeddings using the outer product. In the experiment of the present invention, the code provided by the author is used to implement this model.

[0130] 1 http: / / surpriselib.com

[0131] 2 https: / / github.com / hexiangnan / neural_collaborative_filtering

[0132] 3 https: / / github.com / familyld / DeepCF

[0133] 4 https: / / github.com / librahu / HERec

[0134] c. Evaluation Metrics

[0135] To evaluate the performance of NARec, this embodiment uses the Root Mean Square Error (RMSE) and Mean Absolute Error (MAE), which are widely used in rating prediction tasks, as evaluation metrics and are defined as follows:

[0136]

[0137]

[0138] where r u,v is the true rating of user u for item v, is the predicted rating of the test model; represents the test set of ratings and is obtained by dividing the known rating set into a training set and a test set. Obviously, these two metrics characterize the rating prediction error of the test model, so the smaller the value, the better the recommendation performance.

[0139] d. Experimental Environment and Parameter Settings

[0140] The experiments of the present invention were conducted on a workstation with an Intel Xeon Gold 6230R 2.10GHz CPU, 128GB of memory, and an NVIDIA GeForce RTX 2080Ti GPU. The proposed model NARec was implemented on the PyCharm platform based on Python 3.7 and Keras using the TensorFlow backend.

[0141] In terms of relationship feature extraction, the model NARec uses the meta-paths in the last column of Table 4, where the capital letters are the abbreviations of the corresponding entity types. In terms of attribute feature extraction, the Yelp dataset adopts three attributes: Compliment, City, and Category; the Douban Book dataset adopts four attributes: Group, Author, Publisher, and Year; the Douban Movie dataset adopts four attributes: Group, Director, Actor, and Type. During model training, the mean squared error is selected as the loss function to obtain the best RMSE, and the mean absolute error is selected as the loss function to obtain the best MAE. Meanwhile, mini-batch Adam with a batch size of 256 is used to accelerate the training process. To obtain the best prediction performance, the grid search method is used to find the best learning rate η from {0.0005, 0.001, 0.002}, the best feature representation dimension d from {4, 8, 16, 32, 64, 128, 256}, and the best prediction vector dimension D from {4, 8,..., d / 2}. In particular, for the Yelp dataset and the Douban dataset, the best regularization coefficient λ is found in the intervals [0.0002, 0.002] and [0.00001, 0.0001], respectively. For each set of hyperparameters, the model NARec is iterated 30 times, and the predicted score with the minimum RMSE or MAE is taken as the final output. For fairness, the hyperparameters of the comparison algorithms are also optimized to obtain the best performance on the three datasets.

[0142] e. Analysis of experimental results

[0143] In this subsection, extensive experiments are conducted on three public datasets (Yelp, Douban Book, and Douban Movie), and the proposed model NARec is compared with five benchmark models (PMF, NeuMF, CFNet, HERec, and HopRec) in terms of performance. For each dataset, the known rating set is randomly divided into a training set and a test set To consider the impact of different rating densities on the recommendation performance, the present invention sets three splitting ratios for these three datasets, namely 80%, 50%, and 20%. The splitting ratio here refers to the training ratio. For example, a splitting ratio of 80% means randomly selecting 80% of the known ratings as the training set to predict the remaining 20% of the ratings. The comparison of the rating prediction performance of the proposed model NARec and the benchmark models on the three datasets is shown in Table 5. For each training ratio of each dataset, the optimal result is represented in bold, and the sub-optimal result is underlined. For easy comparison, the performance improvement of the model NARec over each benchmark model in terms of MAE and RMSE is given in the table. Analyzing the experimental results, the following conclusions can be drawn:

[0144] (1) Among these recommendation models, NeuMF and HopRec achieved the second-best performance under different training data settings respectively, while the proposed model NARec achieved the best performance at all training ratios on all datasets. This is because NeuMF is a state-of-the-art neural network-based collaborative filtering method that can effectively model the complex non-linear interactions between users and items for recommendation; while HopRec is a state-of-the-art heterogeneous information network-based recommendation method that can effectively extract useful information from side information for recommendation. Since NARec combines the advantages of deep learning and heterogeneous information networks in recommendation, its rating prediction performance has been significantly improved compared to the baseline models. In particular, compared with PMF, NARec has greatly improved the rating prediction performance. These experimental results demonstrate the effectiveness of the proposed model NARec.

[0145] (2) As the training ratio of each dataset decreases, the rating training data becomes sparser and sparser, resulting in a decline in recommendation performance. However, the lower the training ratio, the higher the performance improvement of the proposed model NARec. For the Yelp dataset with a training ratio of 80%, in terms of RMSE (MAE), the performance improvement of NARec compared to PMF is 15.60% (20.91%); however, when the training ratio is 20%, the performance improvement reaches 22.62% (26.97%). For the Douban Book dataset with a training ratio of 80%, in terms of RMSE (MAE), the performance improvement of NARec compared to PMF is 8.28% (14.10%); however, when the training ratio is 20%, the performance improvement reaches 19.92% (23.30%). For the Douban Movie dataset with a training ratio of 80%, in terms of RMSE (MAE), the performance improvement of NARec compared to PMF is 9.41% (13.72%); however, when the training ratio is 20%, the performance improvement reaches 12.81% (16.60%). These experimental results show the effectiveness of NARec for sparse datasets.

[0146] (3) The proposed model NARec outperforms all comparison models on different real-world datasets. Since the Yelp dataset is too sparse, the training ratio of 20% was not experimented in the original texts of HERec and HopRec. The present invention conducted experiments on all recommendation models with a training ratio of 20% on the Yelp dataset, where NARec achieved the best RMSE (MAE) of 1.0670 (0.6849). In contrast, the Douban Movie dataset is the densest, and when the training ratio is 80%, NARec also achieved the best RMSE (MAE) of 0.6894 (0.5042). These experimental results demonstrate the effectiveness of NARec for datasets with different densities.

[0147] (4) The proposed model NARec achieved a higher performance improvement in terms of MAE than in terms of RMSE. For example, on the Douban Book dataset with a training ratio of 50%, the performance improvement of NARec compared to PMF in terms of RMSE was 11.40%, but the performance improvement in terms of MAE reached 16.91%. This is because when choosing the absolute error as the loss function instead of the squared error, the model can obtain a better MAE.

[0148] Table 5 Comparison of rating prediction performance of different recommendation algorithms

[0149]

[0150] f. Hyperparameter analysis

[0151] In this subsection, the impact of hyperparameters on the model's recommendation performance is analyzed, including the feature representation dimension d, the prediction vector dimension D, the learning rate η, and the regularization coefficient λ. For simplicity but without loss of generality, the training ratio of 0.8 is taken as an example to analyze the hyperparameters.

[0152] f.1 Feature representation dimension

[0153] The model NARec proposed in this embodiment characterizes users and items from three aspects, namely latent features, attribute features, and relationship features. To analyze the impact of the feature representation dimension d on the recommendation performance, in this embodiment, d is respectively set to {4, 8, 16, 32, 64, 128, 256}, and the prediction vector dimension D is set to d / 2. At the same time, HERec and HopRec are selected as comparison models because these two models also use relationship features for recommendation. For each feature representation dimension d, the learning rate and the regularization coefficient are adjusted so that NARec achieves the best performance on the three datasets. The experimental results are as Figure 4As shown. It can be seen that the performance of the model NARec exceeds that of the comparison models in terms of different feature representation dimensions on different datasets. For the Yelp dataset, the prediction accuracy of NARec first increases and then decreases with the increase of the feature representation dimension d, and the best performance is obtained when d is 32. For the Douban Book and Douban Movie datasets, the prediction accuracy of NARec gradually increases with the increase of the feature representation dimension d, but the model complexity also increases accordingly. As a compromise between performance and efficiency, the feature representation dimension d of the Douban dataset is set to 128 in this embodiment.

[0154] f.2 Prediction vector dimension

[0155] In the score prediction part of the model NARec, the prediction vector z u,v determines the score prediction ability. To analyze the influence of the prediction vector dimension D on the recommendation performance, the feature representation dimension d is fixed at 128 in this embodiment, and then D is set to {4, 8, 16, 32, 64} respectively. For each prediction vector dimension D, the learning rate and the regularization coefficient are adjusted to make NARec obtain the best performance on the three datasets. The experimental results are as Figure 5 shown. It can be seen that the performance of the model NARec remains stable for different prediction vector dimensions, that is, the model NARec is robust to the prediction vector dimension. At the same time, NARec can obtain good prediction performance on different prediction vector dimensions of the three datasets. The prediction vector dimension is set to d / 2 in this embodiment.

[0156] f.3 Learning rate and regularization coefficient

[0157] To optimize the model NARec, the Adam optimizer is used to minimize the objective function (16) in this embodiment, where the learning rate η determines the step size of each iteration. When setting the learning rate, a trade-off needs to be made between the convergence speed and overshoot. A higher learning rate will cause the model to skip the optimal value, while a lower learning rate will result in a longer convergence time or convergence to a local optimal value. To prevent the model from overfitting, L2 regularization is performed on the embedding layer and the fully connected layer of the model NARec. By imposing a penalty on the complexity of the model, the regularization term improves the generalization ability of the model, where the regularization coefficient controls the importance of the regularization term in the objective function.

[0158] Due to the mutual coupling between the learning rate and the regularization coefficient, it is necessary to analyze their impacts on the recommendation performance simultaneously. Among them, the dimension d of the feature representation is set to 128, the dimension D of the prediction vector is set to 64, the regularization coefficients λ of the Yelp dataset are respectively set to {0.0005, 0.001, 0.002, 0.005, 0.01}, the regularization coefficients λ of the Douban dataset are respectively set to {0.00001, 0.00002, 0.00005, 0.0001, 0.0002}, and the learning rates η are respectively set to {0.0005, 0.001, 0.002, 0.005}. The experimental results are as Figure 6 shown, where the black dots represent the best performance at different learning rates. It can be seen that as the learning rate increases, the best regularization coefficient gradually decreases, that is, the optimal values of η and λ are negatively correlated. Taking the Yelp dataset as an example, the best regularization coefficient when the learning rate η = 0.001 is λ = 0.005, the best regularization coefficient when the learning rate η = 0.002 is λ = 0.001, and the best regularization coefficient when the learning rate η = 0.005 is λ = 0.0005. Considering the performance of different learning rates comprehensively, the Yelp dataset achieves the best performance when η = 0.001 and λ = 0.005, and the Douban dataset achieves the best performance when η = 0.001 and λ = 0.0001. Therefore, in this embodiment, the learning rate η of the three datasets is all set to 0.001.

[0159] To analyze the impact of the regularization coefficient at different iteration rounds on the recommendation performance, this embodiment tests the performance of the model NARec when the regularization coefficient λ takes the optimal value, a larger value, and a smaller value respectively. For the Yelp dataset, the values of λ are {0.002, 0.005, 0.01}; for the Douban dataset, the values of λ are {0.00005, 0.0001, 0.0002}. Since the fluctuation of RMSE is relatively large at different iteration rounds, the optimal value of RMSE for every two iterations is taken as the prediction error for these two iterations. The prediction errors of NARec at different regularization coefficients and different iteration rounds on the training set and test set of the three datasets are as Figure 7As shown, it can be seen that on the training sets of the three datasets, as the number of iterations increases, the RMSE of model NARec gradually decreases, and the larger the regularization coefficient, the slower the RMSE decreases; on the test sets of the three datasets, the most effective update of RMSE occurs in the first 10 iterations. When the regularization coefficient takes a smaller value (λ = 0.002 for the Yelp dataset and λ = 0.00005 for the Douban dataset), as the number of iterations increases, the RMSE on the test set gradually increases instead, indicating that the model is overfitting. When the regularization coefficient takes a larger value (λ = 0.01 for the Yelp dataset and λ = 0.0002 for the Douban dataset), as the number of iterations increases, the RMSE on the test set will eventually tend to be stable, avoiding model overfitting, but the recommendation performance of the model is poor. When the regularization coefficient takes the optimal value, the model achieves the optimal performance and avoids model overfitting except for the Douban Book dataset. By comparing the optimal regularization coefficients of the three datasets, it can be seen that due to the lack of training data, sparse datasets require larger regularization coefficients to prevent model overfitting.

[0160] As Figure 8 shown, on the other hand, the present invention proposes a neural network recommendation system based on side information and attention mechanism, including:

[0161] A heterogeneous information network construction module for constructing multi-source data into a heterogeneous information network; the data includes user-item rating data and side information;

[0162] A feature extraction module for extracting the latent features, attribute features, and relationship features of users and items in the heterogeneous information network; among them, the latent features are dense feature representations obtained by embedding the numbers of users and items, the attribute features are learned from the attribute matrices of users and items through a multi-layer perceptron network, and the relationship features are learned from the heterogeneous information network through meta-path-based network embedding;

[0163] A predicted rating and recommendation module for predicting the rating of a user for an item based on the extracted latent features, attribute features, and relationship features of the user and the item, and making recommendations based on the predicted rating.

[0164] Furthermore, in the feature extraction module, the latent features of users and items in the heterogeneous information network are extracted in the following manner: converting the corresponding numbers of users and items into one-hot encoded binary vectors; mapping the obtained binary vectors into dense user / item embeddings through an embedding layer; the obtained user / item embeddings are the latent features of the user / item.

[0165] Further, in the feature extraction module, relationship features of users and items are extracted from the heterogeneous information network by using network embedding based on meta-paths, including: First, for each meta-path, the heterogeneous information network is converted into a corresponding homogeneous information network by using the connectivity matrix; then, the feature representation of users or items in the homogeneous information network is extracted by using the network embedding algorithm.

[0166] Further, in the feature extraction module, the attribute features of users and items in the heterogeneous information network are extracted in the following manner: the attribute values of users and items are converted into binary vectors of one-hot or multi-hot encoding; through a two-layer fully connected network with ReLU activation function, the obtained binary vectors are mapped into dense feature representations, that is, the user attribute feature matrix and the item attribute feature matrix are obtained.

[0167] Further, the prediction score and recommendation module is specifically used for: fusing the latent features, attribute features and relationship features of each user and item through an attention network into a user feature vector and an item feature vector, and then predicting the score of the user for the item by using the user feature vector and the item feature vector, and recommending the top n items with the highest predicted scores to the user.

[0168] The present invention proposes a neural network recommendation method and system based on side information and attention mechanism to improve the recommendation performance of collaborative filtering based on side information. The rating data and side information are first modeled as a heterogeneous information network. In order to comprehensively extract useful recommendation features from the heterogeneous information network, the present invention characterizes users and items from three aspects: latent features, attribute features and relationship features. Specifically, these three types of features are respectively learned from the numbers of users and items, the attribute matrix and the connectivity matrix based on meta-paths. Generally, a user may attach different degrees of importance to the same feature of different items, and at the same time, an item may also have different attractions to different users with the same feature. In order to adaptively characterize the personalized features of users and items, the present invention uses an attention mechanism-based rating prediction model to fuse different features of users and items for recommendation. The present invention has conducted a large number of experiments on three real-world data sets, and the experimental results verify the effectiveness and superiority of the present invention in the rating prediction task.

[0169] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A neural network recommendation method based on side information and attention mechanism, characterized in that Including: Step 1: Construct a heterogeneous information network from multi-source data; the data includes user-item rating data and edge information; Step 2: Extract the latent features, attribute features, and relationship features of users and items in the heterogeneous information network; wherein the latent features are dense feature representations obtained by embedding the numbers of users and items, the attribute features are learned from the attribute matrices of users and items through a multi-layer perceptron network, and the relationship features are learned from the heterogeneous information network through network embedding based on meta-paths; Step 3: Predict the ratings of users for items based on the extracted latent features, attribute features, and relationship features of users and items, and make recommendations based on the predicted ratings; In the said Step 2, the latent features of users and items in the heterogeneous information network are extracted in the following manner: convert the corresponding numbers of users and items into one-hot encoded binary vectors; map the obtained binary vectors into dense user / item embeddings through an embedding layer; the obtained user / item embeddings are the latent features of users / items; In the said Step 2, network embedding based on meta-paths is used to extract the relationship features of users and items from the heterogeneous information network, including: First, for each meta-path, convert the heterogeneous information network into a corresponding homogeneous information network using a connectivity matrix; Then, use a network embedding algorithm to extract the feature representations of users or items in the homogeneous information network; In the said Step 2, the attribute features of users and items in the heterogeneous information network are extracted in the following manner: convert the attribute values of users and items into one-hot or multi-hot encoded binary vectors; map the obtained binary vectors into dense feature representations through a two-layer fully connected network with a ReLU activation function, that is, obtain the user attribute feature matrix and the item attribute feature matrix.

2. The neural network recommendation method based on side information and attention mechanism according to claim 1, characterized in that Step 3 includes: fusing the latent features, attribute features, and relationship features of each user and item through an attention network into a user feature vector and an item feature vector, and then using the user feature vector and the item feature vector to predict the user's rating for the item, and recommending the top n items with the highest predicted ratings to the user.

3. A neural network recommendation system based on side information and attention mechanism, characterized in that, Including: A heterogeneous information network construction module for constructing a heterogeneous information network from multi-source data; the data includes user-item rating data and edge information; A feature extraction module for extracting the latent features, attribute features, and relationship features of users and items in the heterogeneous information network; wherein the latent features are dense feature representations obtained by embedding the numbers of users and items, the attribute features are learned from the attribute matrices of users and items through a multi-layer perceptron network, and the relationship features are learned from the heterogeneous information network through network embedding based on meta-paths; A predicted rating and recommendation module for predicting the ratings of users for items based on the extracted latent features, attribute features, and relationship features of users and items, and making recommendations based on the predicted ratings; In the said feature extraction module, the latent features of users and items in the heterogeneous information network are extracted in the following manner: convert the corresponding numbers of users and items into one-hot encoded binary vectors; map the obtained binary vectors into dense user / item embeddings through an embedding layer; the obtained user / item embeddings are the latent features of users / items; In the feature extraction module, relationship features of users and items are extracted from the heterogeneous information network by using meta-path-based network embedding, including: First, for each meta-path, the heterogeneous information network is converted into a corresponding homogeneous information network by using a connectivity matrix; Then, the feature representation of users or items in the homogeneous information network is extracted by using a network embedding algorithm. In the feature extraction module, the attribute features of users and items in the heterogeneous information network are extracted in the following way: the attribute values of users and items are converted into binary vectors of one-hot or multi-hot encoding; through a two-layer fully connected network with a ReLU activation function, the obtained binary vectors are mapped into dense feature representations, that is, the user attribute feature matrix and the item attribute feature matrix are obtained.

4. The neural network recommendation system based on side information and attention mechanism according to claim 3, characterized in that, The prediction score and recommendation module is specifically used to: fuse the latent features, attribute features, and relationship features of each user and item through an attention network into a user feature vector and an item feature vector, and then use the user feature vector and the item feature vector to predict the score of the user for the item, and recommend the top n items with the highest predicted scores to the user.