Project prediction method based on project attribute perception
By splitting the user interaction sequence into attribute interaction sequences and predicting the next possible attribute of the user on each attribute, the problem of ignoring project semantics and text information in the prior art is solved, and a more efficient sequence recommendation effect and attention to different attributes is achieved.
Patent Information
- Application Number
- CN202510351585.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-10
AI Technical Summary
The existing sequence recommendation algorithm ignores the semantic and textual information carried by the project itself, resulting in insufficient recommendation capabilities for cold-start projects and users, and it is difficult to use multi-domain data to improve the target area.
A project prediction method based on project attribute awareness is proposed, which splits the user's interaction sequence into each attribute interaction sequence, predicts the user's next possible interaction attribute on each attribute, and weighted and aggregates these prediction results to generate a prediction for the next project.
By refining the prediction granularity to project attributes, the effect of sequence recommendations is improved, users can better pay attention to their preferences for different attributes, and effectively use multi-field data to improve recommendation performance.
Smart Images

Figure CN120123596A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of personalized recommendation systems, and particularly relates to an item prediction method based on item attribute perception. Background Art
[0002] With the development of the Internet and information technology, the amount of information in online services such as social media, news media, and e-commerce platforms has increased exponentially. Individual users can hardly process all the information they may face, thus facing the problem of information overload. The existence of recommendation systems is to help users filter out low-quality or irrelevant information to personal interests in a short time.
[0003] The sequential recommendation task focuses on modeling the user profile through the user's interaction sequence, and improving the recommendation effect by capturing the user's time-varying interests. Its problem definition can be simply described as: predicting the item that the user may interact with at the next moment based on the user's historical behavior. Therefore, how to efficiently mine the user's preferences has become a difficult point in the research of sequential recommendation algorithms. In the era of deep learning, many recommendation algorithms adopt the model framework of deep learning and are committed to improving the processing effect of sequential information. However, most existing algorithms still rely on item IDs to distinguish different items and learn embedding vectors based on the interaction history between users and items. Although this method can play a fundamental role, it ignores the semantic and text information carried by the items themselves, such as product attributes, descriptions, comments, and multi-modal information. This not only limits the recommendation ability of the algorithm model for cold-start items and users, but also leads to insufficient attention to long-tail items.
[0004] Although there are some sequential recommendation algorithms that consider the attribute information of items at present, most of them simply splice the item texts and input them into the model for processing. For example, in the existing technology, the ZESRec model uses the general content information of items, such as natural language descriptions, to generate item embeddings. This kind of embedding can be used as a general item ID in different fields. However, the ZESRec model still has two main problems to be solved: one is that the text semantic space is not directly applicable to the recommendation task; the other is that it is difficult to use multi-domain data to improve the target domain, where the seesaw phenomenon (referring to the conflict or oscillation between the knowledge learned from multiple specific domain patterns) often appears.
[0005] In addition, a sequence recommendation model UniSRec has also been proposed in the prior art. In the model, a lightweight item encoding architecture is designed based on parameter whitening and mixture of experts, and two contrastive pre-training tasks are introduced to learn general item and sequence representations. However, this cannot learn the importance of different attributes for items and ignores the preferences of different users for different attributes of items. For example, high-spending users may be more inclined to the brand value and popularity trends of goods when choosing goods, while ordinary users may be more concerned about whether the price of goods is appropriate. Summary of the Invention
[0006] In view of the problems mentioned in the background art, the present invention proposes an item prediction method based on item attribute perception, which refines the prediction granularity to the attributes of items. The user's item interaction sequence is split into each attribute interaction sequence, and the next possible interacting attribute of the user is predicted on each attribute. Finally, these prediction results are weighted and aggregated into a prediction of the next item, effectively improving the effect of sequence recommendation.
[0007] Technical Solution: To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0008] An item prediction method based on item attribute perception, comprising the following steps:
[0009] S1: Personalized attribute preference representation: Expand the user interaction sequence into a user interaction attribute matrix;
[0010] S2: Input embedding: Generate attribute embeddings and item embeddings, and fuse learnable weights and position embeddings;
[0011] S3: Attribute sequence preference learning: Based on the Transformer architecture, learn the user preference of a single attribute sequence through the self-attention mechanism and the feed-forward neural network;
[0012] S4: Preference learning and prediction between attributes: Weightedly aggregate the prediction results of each attribute, calculate the item score, and generate a recommendation list.
[0013] Preferably, in S1, first truncate or pad the user interaction sequence to a fixed length n, and then expand the user's interaction sequence into a user interaction attribute matrix composed of the attribute information of each item:
[0014]
[0015] Wherein, represents the user interaction attribute matrix, represents the nth element of the jth attribute in the interaction sequence of user u, Each row in represents an interaction attribute sequence, and each column represents an interacting item.
[0016] Preferably, in S2, the embedding of each item is obtained by concatenating the respective attributes in the item. Denote the item embedding as , the attribute embedding as , and the embedding of the -th item is represented as:
[0017] ,
[0018] where represents the embedding of the k-th item; represents the embedding of the j-th attribute in the k-th item, and its specific representation is:
[0019] ,
[0020] where represents the element in the embedding vector obtained after performing the embedding operation on the j-th attribute element; T represents the transpose operation.
[0021] Preferably, in S2, learnable weights and positional embeddings are introduced for each attribute sequence, specifically:
[0022] ,
[0023] ,
[0024] where represents the weight parameter on different attributes; represents the parameter category; represents adding the learnable weights and positional embeddings to each element in the attribute sequence; represents the positional embedding at the n-th position; T represents the transpose operation; represents the embedding of the j-th attribute sequence in the n-th item.
[0025] Preferably, in S3, for the learning of a single attribute sequence, the calculation formula of scaled dot-product attention is:
[0026] ,
[0027] Using the attribute embedding and three learnable parameter matrices, input them into the self-attention mechanism through linear transformation to calculate the attention scores between elements, specifically:
[0028] ,
[0029] where Q, K, and V represent query, key, and value respectively; represents the scaling operation; T represents the transpose operation; Denotes the vector output after the self-attention mechanism processes; Denotes adding learnable weights and position embeddings to each element in the attribute sequence; Denotes the query matrix; Denotes the key matrix; Denotes the value matrix.
[0030] Preferably, in S3, in order to learn the information between different embedding dimensions and enhance the effect of the self-attention mechanism, a feed-forward neural network composed of two fully connected layers is added, and the specific calculation formula is:
[0031] ,
[0032] where, Denotes the feed-forward neural network; Denotes the input tensor obtained by further processing the output of the self-attention mechanism and input into the feed-forward neural network; and Denotes the learnable weight matrix; and Denotes the learnable parameters.
[0033] Preferably, in S3, layer normalization, Dropout operation, and residual connection are performed on the output of each self-attention mechanism module, specifically:
[0034] ,
[0035] where, Denotes the output of the previous module during the stacking operation; Denotes the input when stacking the self-attention mechanism module and the feed-forward neural network, Denotes regularization; Denotes the self-attention mechanism; Denotes the layer normalization operation;
[0036] The calculation process of layer normalization is:
[0037] ,
[0038] where, Denotes performing layer normalization on the input tensor x; Denotes the scaling factor; Denotes the bias term; Denotes the Hadamard product, and respectively denote the mean and variance of the input tensor ; Denotes a constant.
[0039] Preferably, layer normalization, Dropout operation, and residual connection are performed on each feed-forward neural network module, specifically:
[0040] ,
[0041] Among them, represents the input when the feed-forward neural network; represents the output of the feed-forward neural network; represents layer normalization on the input tensor ; represents regularization; represents the feed-forward neural network.
[0042] Preferably, in S4, the preferences of the user among different attributes are learned, specifically:
[0043] Suppose the input sequence is , and the input subsequence is . Before predicting the next item, the output sequence on each attribute is . By weighted concatenating the outputs at the -th position on different attributes, the prediction of the item can be obtained, specifically:
[0044] ,
[0045] Among them, represents the prediction of the item embedding at the t-th position; represents the weight parameters on different attributes; represents the prediction vector of attribute j at the t position;
[0046] Taking the item embedding at the -th position as the ground truth, then the sum of the errors between the output sequence of the model and the ground truth is expressed as:
[0047] ,
[0048] Among them, represents the prediction of the item embedding at the t-th position; n represents the length of the interaction sequence.
[0049] Preferably, in S4, the score of each item is calculated using the dot product of vectors, specifically:
[0050] ,
[0051] Among them, represents at the position the The score of an item; Denote the predicted item embedding at the t-th position; Denote the embedding of the k-th item.
[0052] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0053] (1) In the present invention, the prediction granularity is refined to the attributes of items. The user's item interaction sequence is split into each attribute interaction sequence. The attribute that the user may interact with next is predicted on each attribute, and finally these prediction results are weighted and aggregated into the prediction of the next item, effectively improving the effect of sequential recommendation.
[0054] (2) In the present invention, the user's interaction sequence is split into multiple attribute subsequences, and the user's preferences are learned separately on each subsequence, and then weighted and aggregated according to the weight parameters between the attributes to obtain the final recommendation result. This method can not only prove that the information contained in the item attributes can optimize the effect of the sequential recommendation task, but also pay attention to the importance that the user attaches to different attributes.
[0055] (3) Although some sequential recommendation algorithms in the prior art consider the attribute information of items, most of them simply splice the item texts and input them into the model for processing. This cannot learn the importance of different attributes for items, and at the same time ignores the preferences of different users for different attributes of items. The present invention constructs a new sequential recommendation algorithm, an attribute-aware sequential recommendation algorithm based on the self-attention mechanism. This algorithm can recommend the next item for the user based on multiple attributes of the item. The whole recommendation process is as follows: Independently predict the characteristics of the item that the user may interact with next on each attribute, and at the same time assign a learnable weight parameter to each attribute. Finally, according to the learned weight parameter and the prediction results on each attribute, weighted aggregation is performed to predict the next item. Description of the Drawings
[0056] Figure 1 is the flowchart of the item prediction method based on item attribute awareness of the present invention;
[0057] Figure 2 is the architecture diagram of the item prediction method based on item attribute awareness of the present invention. Detailed Embodiments
[0058] The following further clarifies the present invention in conjunction with specific embodiments. The embodiments are implemented on the premise of the technical solution of the present invention. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.
[0059] The project prediction method based on project attribute perception provided in this embodiment is applied to improving the sequence recommendation effect, and mainly includes the following steps:
[0060] S1: Personalized attribute preference representation: Adjust the length of the interaction sequence and expand the user interaction sequence into a user interaction attribute matrix;
[0061] First, the interaction sequence needs to be preprocessed for subsequent batch processing. In this step, all sequences are adjusted to a fixed length , if the length of the sequence exceeds , the model only considers the nearest interaction items. If the sequence length is less than , padding items will be added to the left of the sequence until the sequence length is equal to . Similarly, a length limit will also be set for the text length within the attribute to prevent excessive irrelevant information from interfering with the learning of user preferences.
[0062] Since this method needs to make predictions at the attribute level, the user's interaction sequence can be expanded into a user interaction attribute matrix composed of the attribute information of each item:
[0063]
[0064] Among them, represents the user interaction attribute matrix, is the number of attributes considered in each item; n represents that there are n elements in the interaction sequence; represents the nth element of the jth attribute in the interaction sequence of user u. Obviously, each row in is the interaction attribute sequence, and each column represents an interaction item.
[0065] S2: Input embedding: Generate attribute embeddings and item embeddings, and fuse learnable weights and position embeddings;
[0066] The algorithm proposed in this embodiment is based on attribute-level prediction. Therefore, the embedding of each item is obtained by concatenating the various attributes in the item. Here, the item embedding is denoted as , and the attribute embedding is denoted as , where represents the size of the item set; represents the dimension of the embedding vector; represents that there are j attributes.
[0067] So the embedding of the th item can be expressed as:
[0068] ,
[0069] Among them, represents the embedding of the k-th item; represents the embedding of the j-th attribute in the k-th item, and its specific representation is:
[0070] ,
[0071] Among them, represents the element in the embedding vector obtained after the embedding operation on the j-th attribute element, which is a numerical value; T represents the transpose operation.
[0072] Since the algorithm in this embodiment is trained independently on each attribute, the dimensionality of the feature vectors between different attributes can be different. However, due to simplifying the complexity of the operation, the dimensionality of the feature vectors on each attribute is uniformly designed as . In a user interaction sequence , the embedding representation of the -th attribute sequence can be written as:
[0073] ,
[0074] Among them, represents the embedding operation; represents the embedding of the j-th attribute sequence in the n-th item; represents the n-th element of the j-th attribute in the interaction sequence of user u.
[0075] Based on the subsequent learning of the user's preferences between different attributes, a learnable weight parameter is multiplied by each attribute sequence, and all weight coefficients can be written as:
[0076] ,
[0077] Among them, represents the weight parameter on different attributes; represents the parameter category.
[0078] The embedding representation of the corresponding attribute sequence is modified to: ,
[0079] In order to enable the model to perceive the position of each element in the sequence, a position embedding is also added to each element in the attribute sequence:
[0080] ,
[0081] Among them, represents adding a learnable weight and position embedding to each element in the attribute sequence; represents the position embedding at the n-th position; T represents the transpose operation; Represents the weight parameters on different attributes; represents the embedding of the j-th attribute sequence in the n-th item.
[0082] at last, It can be input into the model for attribute-level prediction.
[0083] S3: Attribute sequence preference learning: Based on the Transformer architecture, the user preference of a single attribute sequence is modeled through the self-attention mechanism and feedforward neural network, and the model stability is improved by combining layer normalization, residual connection and Dropout;
[0084] In the learning of a single attribute sequence, the present invention is based on the Transformer architecture and the self-attention mechanism. The calculation formula of the scaled dot product attention is defined as:
[0085] ,
[0086] Among them, Q, K, and V represent "query", "key", and "value" respectively; This is to prevent the inner product of the vector from causing the result to be too large. T represents the transpose operation.
[0087] The difference between the self-attention mechanism and the attention mechanism is that the "query", "key" and "value" in the model come from the same sequence object. Use the attribute embedding obtained in the embedding layer , through linear transformation with three learnable parameter matrices and input into the self-attention mechanism, the attention scores between elements are calculated, specifically:
[0088] ,
[0089] in, Represents the vector output after self-attention mechanism processing; , , are all learnable parameter matrices in the self-attention mechanism, specifically: represents the query matrix; represents the bond matrix; Represents a matrix of values.
[0090] Through the self-attention mechanism, it is possible to take into account all previously interacted elements when predicting the next element that the user may interact with. In order to learn the information between different embedding dimensions and enhance the effect of the self-attention mechanism, a feedforward neural network consisting of two fully connected layers is added, using the activation function ReLU. The specific calculation formula is:
[0091] ,
[0092] Among them, represents a feedforward neural network; represents the input tensor that is further processed from the output of the self-attention mechanism and then input into the feedforward neural network; and represent learnable weight matrices; and represent learnable parameters.
[0093] Next, how to perform some necessary operations on the input and output of the self-attention mechanism and the feedforward neural network will be introduced.
[0094] As the model scale gets larger and the number of network layers gets deeper, the training of the model often becomes unstable, such as the phenomenon of gradient disappearance, and there is also a possibility of overfitting. To prevent these problems, corresponding operations need to be performed on each core module.
[0095] The self-attention mechanism module is the core of this method. Here, layer normalization, Dropout operation, and residual connection are performed on the output of each module. Specifically:
[0096] ,
[0097] Among them, represents the output of the previous module during the stacking operation; represents the input when stacking the self-attention mechanism module and the feedforward neural network, is a regularization technique; represents the self-attention mechanism, an abbreviation of Self-Attention; represents the layer normalization operation.
[0098] Since most researchers have limited computing power resources, layer normalization is used in the model to normalize the input of each layer, which can improve the generalization ability of the model, make the training of the model more stable, and also be able to use a smaller batch size, reducing the video memory required for training the model. The formula for layer normalization is:
[0099] ,
[0100] Among them, represents performing layer normalization on the input tensor x; represents the scaling factor; represents the bias term; represents the Hadamard product, and respectively represent the mean and variance of the input tensor . is an extremely small positive constant used to prevent the denominator from being zero and ensure numerical stability.
[0101] To address the risk of overfitting caused by increasing the number of network layers, the Dropout regularization technique is used. During training, several neurons are randomly discarded, preventing the network from relying on any single neuron and thus preventing overfitting.
[0102] Residual connections alleviate the problems of vanishing gradients and exploding gradients that may occur during the training of deep neural networks by directly adding the input information to the output information, and also accelerate the convergence speed during network training.
[0103] Similarly, layer normalization, Dropout operations, and residual connections are performed on each feedforward neural network module, specifically:
[0104] ,
[0105] where, represents the input when it comes to the feedforward neural network, represents the output of the feedforward neural network module; represents layer normalization on the input tensor ; represents regularization; represents the feedforward neural network.
[0106] S4: Preference learning and prediction among attributes: Aggregate the prediction results of each attribute sequence with weights, calculate the item scores, and generate a recommendation list.
[0107] After learning the user's interaction history on each attribute sequence, the user's preferences among different attributes are now learned.
[0108] In this method, the input sequence is , so as to predict the interaction item at the th position. Then, in each attribute sequence, the input subsequence is . Before predicting the next item, the output sequence on each attribute is . Then, the weighted concatenation of the outputs at the th position on different attributes can obtain the prediction of the item:
[0109] ,
[0110] where, represents the prediction of the item embedding at the t-th position; represents the weight parameters on different attributes; represents the prediction vector for attribute j at position t.
[0111] Embed the item at the th position as the ground truth, then the sum of the errors between the output sequence of the model and the ground truth can be expressed as:
[0112] ,
[0113] where represents the predicted embedding of the item at the t-th position; n represents the length of the interaction sequence mentioned above. represents the embedding of the k-th item.
[0114] During the training process, the sum of errors is minimized by adjusting the weight parameters to achieve the goal of learning the user's preferences for different attributes. In this method, is used to predict the next possible interacting item. First, the scores of each item are calculated using the dot product of vectors:
[0115] ,
[0116] where represents the score of the th item at the position ; represents the predicted embedding of the item at the t-th position; represents the embedding of the k-th item.
[0117] The higher the score of an item, the more likely the algorithm thinks it is to interact with the user. After calculating the scores of all candidate items, recommendations are made in descending order of scores.
[0118] In this embodiment, two datasets in different domains on the Amazon platform are selected for experimental evaluation, namely Office Products (Office) and CDs and Vinyl (CDs). The evaluation metrics used are Recall@K and NDCG@K, where K = 10, 20.
[0119] Table 1 Comparison of model evaluation results for datasets in different domains
[0120]
[0121] Compared with 7 baseline models on 2 datasets, the experimental results show that the newly proposed method generally has improvements in evaluation metrics compared with existing models. However, the improvement effects on different datasets are not the same. Next, the existing experimental results will be analyzed in detail. The bold data in the table are the model data with the best performance in the corresponding metrics. For the models with sub-optimal performance, the table also marks them with underlines at the corresponding positions.
[0122] Among all the baseline models, SRGNN has the best performance and is mostly better than other baseline models on Office and CDs, which also demonstrates the effectiveness of the attention network in the sequential recommendation task. Among the remaining baseline models, FDSA and UniSRec are modeled based on item attribute information and perform better than most models on the Office and CDs datasets.
[0123] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A project prediction method based on project attribute perception, characterized in that: The following steps are involved: S1: Personalized attribute preference representation: Expand the user interaction sequence into a user interaction attribute matrix; S2: Input embedding: Generate attribute embedding and item embedding, and fuse learnable weights and position embedding; S3: Attribute sequence preference learning: Based on the Transformer architecture, it learns user preferences for a single attribute sequence through self-attention mechanism and feedforward neural network; S4: Preference learning and prediction between attributes: Weighted aggregation of attribute prediction results, calculation of item scores and generation of recommendation lists.
2. The project prediction method based on project attribute perception according to claim 1 is characterized in that: In S1, the user interaction sequence is first truncated or padded to a fixed length n, and then the user interaction sequence is expanded into a user interaction attribute matrix consisting of the attribute information of each item: , in, represents the user interaction attribute matrix, represents the nth element of the jth attribute in the interaction sequence of user u, Each row in represents an interaction attribute sequence, and each column represents an interaction item.
3. The project prediction method based on project attribute perception according to claim 1 is characterized in that: In S2, the embedding of each item is obtained by concatenating the attributes in the item, and the item embedding is recorded as , attribute embedding is recorded as , No. The embedding representation of each item is: , in, represents the embedding of the kth item; represents the embedding of the jth attribute in the kth item, which is specifically expressed as: , in, represents the element in the embedding vector obtained after the j-th attribute element is embedded; T represents the transposition operation.
4. The project prediction method based on project attribute perception according to claim 3 is characterized in that: In S2, learnable weights and position embeddings are introduced for each attribute sequence, specifically: , , in, Represents the weight parameters on different attributes; Indicates the parameter category; Represents adding learnable weights and position embedding to each element in the attribute sequence; represents the position embedding at the nth position; T represents the transposition operation; represents the embedding of the j-th attribute sequence in the n-th item.
5. The project prediction method based on project attribute perception according to claim 1 is characterized in that: In S3, the calculation formula of scaled dot product attention on the learning of a single attribute sequence is: , Using attribute embedding The three learnable parameter matrices are linearly transformed and input into the self-attention mechanism to calculate the attention scores between elements, specifically: , Among them, Q, K, and V represent query, key, and value respectively; represents a scaling operation; T represents a transposition operation; Represents the vector output after self-attention mechanism processing; Represents adding learnable weights and position embedding to each element in the attribute sequence; represents the query matrix; represents the bond matrix; Represents a matrix of values.
6. The project prediction method based on project attribute perception according to claim 5 is characterized in that: In S3, in order to learn the information between different embedding dimensions and enhance the effect of the self-attention mechanism, a feedforward neural network consisting of two fully connected layers is added. The specific calculation formula is: , in, represents a feed-forward neural network; Represents the input tensor of the feedforward neural network after further processing of the output of the self-attention mechanism; and represents the learnable weight matrix; and represents learnable parameters.
7. The project prediction method based on project attribute perception according to claim 6 is characterized in that: In S3, the output of each self-attention mechanism module is normalized, dropout operated, and residual connected, specifically: , in, Represents the output of the self-attention mechanism module during the stacking operation; represents the input when stacking the self-attention mechanism module and the feedforward neural network, represents regularization; Represents the self-attention mechanism; Representation layer normalization operation; The calculation process of layer normalization is: , in, Indicates layer normalization of the input tensor x; represents the scaling factor; represents the bias term; represents the Hadamard product, and Represents the input tensor respectively The mean and variance of Represents a constant.
8. The project prediction method based on project attribute perception according to claim 1 is characterized in that: Each feedforward neural network module is subjected to layer normalization, Dropout operation and residual connection, specifically: , in, represents the input of the feedforward neural network, Represents the output of a feedforward neural network; Represents the input tensor Perform layer normalization; represents regularization; represents a feed-forward neural network.
9. The project prediction method based on project attribute perception according to claim 1, characterized in that: In S4, the user's preferences between different attributes are learned, specifically: Suppose the input sequence is , the input subsequence is , before predicting the next item, the output sequence on each attribute is , the different attributes are The weighted concatenation of the outputs of the positions can get the prediction of the item, specifically: , in, represents the embedding prediction of the item at the t-th position; Represents the weight parameters on different attributes; Represents the prediction vector for attribute j at position t; The first Item embedding at position As the true value, then the sum of the errors between the model's output sequence and the true value It is expressed as: , in, represents the embedding prediction of the item at the t-th position; n represents the length of the interaction sequence.
10. The project prediction method based on project attribute perception according to claim 9, characterized in that: In S4, the score of each item is calculated using vector dot product, specifically: , in, Indicates at location Previous The score of each item; represents the embedding prediction of the item at the t-th position; represents the embedding of the k-th item.