Multi-modal sequence recommendation method fusing user price preference and interest preference
By introducing product price information into the multimodal sequence recommendation model, capturing user's price preferences and weighting them with interest preferences, the problem of existing models neglecting price factors is solved, and a more accurate and persuasive recommendation effect is achieved.
Patent Information
- Application Number
- CN202510447217.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-06-24
AI Technical Summary
The existing multimodal sequence recommendation model mainly focuses on factors that affect user interest preferences, ignores the important impact of product prices on users' purchasing intentions, resulting in the recommendation effect being not convincing enough.
By introducing product price information, the user's price preferences are weighted and fused, the user's price preferences are captured using a unidirectional attention sequence encoder, and combined with a pre-trained multimodal item encoder to capture the user's interest preferences, and finally the preference score is calculated by matrix multiplication.
It achieves more comprehensive user preference capture and improves the accuracy and effectiveness of sequence recommendations. Especially when considering user price preferences, the recommendation results are more in line with the user's purchasing behavior.
Smart Images

Figure CN120198202A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of e-commerce recommendation systems, and in particular to a multi-modal sequential recommendation method that integrates user price preferences and interest preferences. Background Art
[0002] There are two main research directions in current sequential recommendation systems. The first direction is to enrich the representation of items by introducing item attribute information based on explicit item IDs, thereby improving recommendation performance. For example, the NOVA model in AAAI2021 effectively utilizes auxiliary information under the BERT framework by using a non-invasive self-attention mechanism to model user preferences; the CARCA model in RecSys2022 combines context features and commodity attributes using multi-head self-attention; DIF-SR in SIGIR2022 decouples the attention calculation of various auxiliary information and item representations and proposes an auxiliary attribute predictor to further activate the beneficial interaction between auxiliary information and item representation learning. These models only regard the type or attribute information of commodities as auxiliary information, while ignoring the important factor affecting users' purchase willingness - commodity price. At the same time, models based on explicit item IDs all face the problems of cold start and cross-domain.
[0003] The second direction is to utilize modal information such as commodity pictures and texts to improve sequential recommendation performance and solve the problems of cross-domain and cold start at the same time. For example, UniSRec in KDD2022 uses relevant description texts of items to learn transferable representations across different recommendation scenarios; TransRec in Arxiv2022 encodes visual and text features of items using a modal encoder and realizes true general recommendation through a feedback with MoM mechanism; MMSRec in Arxiv2023 designs a two-tower retrieval architecture for sequential recommendation. In this architecture, the predicted embedding of the user encoder is used to retrieve the embedding generated by the item encoder, thereby alleviating the inconsistency problem between the high-level feature vectors and low-level feature embeddings of item modalities. However, all of the above models only focus on users' interest preferences affected by commodity visual or text features and ignore users' price preferences affected by commodity prices, so they lack a certain degree of persuasiveness.
[0004] In recent years, to address these issues, some sequential recommendation models have started incorporating the price information of products into the recommendation process. CoHHN in SIGIR2022 proposed a heterogeneous hypergraph network to model various heterogeneous information involved in session recommendation, including the price information of products, and proposed a co-guided learning mechanism to fuse user price preferences and interest preferences, guiding each other's representation learning process. The ANT model in Arxiv2023 found that the existing pre-training and fine-tuning sequential recommendation paradigms have serious negative transfer problems. Therefore, it integrates multimodal item information, including item text, images, and prices, to effectively learn more transferable knowledge from auxiliary tasks and uses a re-learning-based adaptation strategy to better capture task-specific knowledge in the target task.
[0005] In summary, existing multimodal sequential recommendations mainly rely on modal information of products such as pictures and texts for recommendation, mainly focusing on factors that affect user interest preferences. However, different from other attributes of products, such as color and style, which affect users' degree of liking for products, that is, interest preferences, product price often directly determines whether a user will purchase a product. Many economic studies have shown that product price plays a crucial role in users' product selection. For example, a person whose consumption product price has been fluctuating between $1 and $10 is also very likely not to buy a product priced at $200. We are more willing to recommend items close to their historical consumption price for them. Therefore, in addition to capturing users' interest preferences, users' price preferences should also be considered in multimodal sequential recommendations. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a multimodal sequential recommendation method that fuses user price preferences and interest preferences in view of the above-mentioned deficiencies of the prior art. By introducing the price information of products into multimodal sequential recommendations to capture users' price preferences, and weighted fusing interest preferences and price preferences, more effective sequential recommendations are achieved.
[0007] To solve the above technical problems, the technical solution adopted by the present invention is:
[0008] A multi-modal sequential recommendation method that fuses users' price preferences and interest preferences. When capturing users' price preferences, the price and category sequences of goods are input into a sequential encoder with unidirectional attention to obtain the user price representation, and then matrix multiplication is performed with the price representation of candidate items to complete the similarity calculation and obtain the price preference score. When capturing users' interest preferences, the pre-trained multi-modal item encoder Clip is used to process the visual and text information of goods to obtain the item representation, and the generated item representation is input into a Transformer-based sequential encoder according to the corresponding relationship of the interaction sequence to obtain the user interest representation. By inputting the candidate item representation into the sequential encoder and then performing matrix multiplication with the user interest representation to complete the similarity calculation, the user's interest preference score for the candidate item is obtained. In the preference fusion stage, the interest preference score and the price preference score are weighted and fused in a suitable proportion to complete more comprehensive preference capture and achieve more effective sequential recommendation.
[0009] Further, the specific implementation steps of the method include:
[0010] Step 1: Data preprocessing. Mark a single item as v, and use to represent the text information, visual information, price information, and category information of the i-th item respectively; the data preprocessing process includes filtering users and items with less than 5 ratings, determining the information content of the text information, visual information, price information, and category information of the items, encoding the text information and visual information to obtain the text representation and visual representation, and data splitting to obtain the test set, validation set, and training set;
[0011] Step 2: Model construction and loss calculation, the specific steps include:
[0012] Step 2.1: Respectively input the text feature vector and the visual feature vector of the item into the linear mapping module Linear to extract the item representation that is more suitable for recommendation semantics from the general semantic information, and then parallelly splice the visual representation and text representation of the item into a whole;
[0013] Step 2.2: According to the historical interaction behavior sequence corresponding to user u, splice the corresponding item modality representations into the item representation sequence S u ={e1, e2, e3…, e n}, e i represents the modality information representation of item i; then input the representation sequence S uThe modal information representation of each item i in is masked with a probability of p = 0.2, that is, replaced with the special token [mask]; for the masked item, all its associated modal representation embeddings are replaced with the token [mask]; where the item representation e at the last position n must be masked; the sequence of user historical interaction item representations after masking is
[0014] Step 2.3: The sequence of user historical interaction item representations obtained in Step 2.2 is summed with the position embedding and the modal embedding Modal u and then input into the sequence encoder Transformer. The specific formula is as follows:
[0015]
[0016] where, represents the context encoding result of the masked sequence;
[0017] The maximum pooling layer is used to extract significant feature information from all modalities to obtain the unified representation of item i The formula is as follows:
[0018]
[0019] where, is the text channel feature, is the visual channel feature, H i is the dual-channel feature. MaxPooling represents the maximum pooling calculation, that is, the maximum value is taken from each dimension of the dual-channel feature and merged into a single-channel feature as the unified representation of item i;
[0020] Step 2.4: The modal information representations e of all candidate items i are concatenated into a representation sequence and added with the position embedding and the modal embedding Modal u and input into the sequence encoder in Step 2.3. Different from Step 2.3, the representation sequence of candidate items is not masked; then the maximum pooling layer is also used to obtain the unified representation E of all candidate items I , where n represents the number of candidate items;
[0021] Step 2.5: The unified representation of item i in Step 2.3 is compared with the unified representation E of all candidate items in Step 2.4 IPerform matrix multiplication to obtain the prediction of the interest preference score for the item at the masked position. The specific formula is as follows:
[0022]
[0023] where τ is the temperature coefficient, Matmul represents matrix multiplication, represents the prediction of the interest preference score for the item at the masked position;
[0024] Step 2.6: Use the cross-entropy loss function to calculate the loss L for predicting the item at the masked position based on the user's interest preference I , and the specific formula is as follows:
[0025]
[0026] where y i represents the true item label at the masked position and is also the supervision signal for recommendation;
[0027] Step 2.7: According to the historical interaction behavior sequence corresponding to user u, convert the corresponding item price information and category information into the vector representation pirce of the price i and the vector representation category of the category i , respectively, and then add the price representation sequence, the category representation sequence, and the position embedding and input them into another one-way attention sequence encoder Transformer. The specific formula is as follows:
[0028]
[0029] h u = Last(H u );
[0030] where Price u = {pirce1, pirce2, pirce3…, pirce n}, represents the price representation sequence of the interaction items of user u, Category u = {category1, category2, category3…, category n}, represents the category representation sequence of the interaction items of user u, H u represents the result encoded by the sequence encoder after adding the three sequences, and Last represents taking the vector at the last position of H u as the price preference representation h u of the user;
[0031] Step 2.8: Take the price preference representation h of the useru Multiply with the price vector matrix and category vector matrix of all candidate items respectively to obtain the preference scores for the price and category of the next item. The specific formula is as follows:
[0032]
[0033] where E P and E C represent the candidate price matrix and category matrix respectively, represent the preference scores for the price and category of the next item respectively;
[0034] Step 2.9: Use the cross-entropy loss function to calculate the loss L P for price prediction and the loss L C for category prediction of the next item based on the user's historical interaction price and category information. The specific formula is as follows:
[0035]
[0036] where represent the true price level label and category label of the next item in the user interaction sequence respectively, which are the supervision signals for prediction;
[0037] Step 2.10: After obtaining the loss L I from Step 2.6 and the loss L P from Step 2.9, as well as the loss L C combine the three parts into a multi-task learning framework. Finally, the objective loss function of the model is expressed as follows:
[0038]
[0039] where θ represents all trainable parameters, λ P and λ C are regularization factors that balance the importance of the price prediction task and the category prediction task;
[0040] Step 3: Model evaluation, the specific steps include:
[0041] Step 3.1: When performing inference, the item representation sequence S u of user u is no longer randomly masked, but only the item id at the last position is masked. At this time, the item representation sequence is For perform the same operations as in Step 2.3 and input it into the sequence encoder to obtain the user interest preference representation at the last position and the unified representation E IPerform matrix multiplication to obtain the interest preference scores of user u for all candidate items. The specific formula is as follows:
[0042]
[0043] Among them, represents the context encoding result of the mask sequence, is the interest preference score of user u for all candidate items;
[0044] Step 3.2: When performing inference in the price module, the price preference representation h of the user u is respectively matrix-multiplied with the price matrix E of all candidate items P and the category matrix E C to obtain the price preference scores of user u for all candidate items and the category preference scores Different from the training stage, there is no need to add the temperature coefficient τ. The specific formula is as follows:
[0045]
[0046] Step 3.3: According to the three preference score vectors obtained in Step 3.1 and Step 3.2, perform preference fusion. The specific formula is as follows:
[0047]
[0048] Among them, β P and β C are influence factors that weigh the importance of the user's price preference and category preference. Total_Preference u is the comprehensive preference score of user u for all candidate items;
[0049] Step 3.4: Sort the comprehensive preference scores of all candidate items obtained in the previous step in descending order to generate a sorted list, and then take the top K items as the final recommendation results. The specific formula is as follows:
[0050] Ranked_List u = argsort(Total_Preference u );
[0051] Rec_List u = {j | j ∈ Ranked_List u [1:K]};
[0052] Among them, argsort is a sorting function that returns the sorted index values of the elements in the array, corresponding to the labels of the candidate items; Ranked_List uIt is a list of tags after all candidate items are sorted in descending order according to the preference scores, and the item corresponding to the first tag in the list has the highest score; Rec_List u It is a recommended list for the user obtained by intercepting the tags of the first K items;
[0053] Step 3.5: According to the recommended list given by the model and the true tag of the user's next interactive item, calculate the model evaluation metrics, including Recall@5, Recall@10, Recall@20, Recall@50, NDCG@5, NDCG@10, NDCG@20, NDCG@50, MRR@5, MRR@10, MRR@20, MRR@50;
[0054] Step 4: Implementation of model training and evaluation, and the specific process includes:
[0055] Step 4.1: Execute Step 2 on the training dataset, perform forward propagation, and calculate the loss, that is, the loss in Step 2.10;
[0056] Step 4.2: Backward propagation and gradient update; perform backward propagation according to the loss obtained from forward propagation, execute gradient update, and update the model training parameters;
[0057] Step 4.3: Execute Step 3 model evaluation on the validation set to obtain the model accuracy performance, that is, the evaluation metrics in Step 3.6, and then compare it with the current optimal Recall@10. If it exceeds the optimal metric, update the optimal Recall@10 to the current value and save the current model training parameters; otherwise, use the early-stop strategy, that is, determine whether the accuracy has not improved for 10 consecutive training rounds. If so, stop training and enter Step 4.4 to perform testing. If not, return to Step 4.1 for the next round of training;
[0058] Step 4.4: Perform testing, load the optimal model parameters, perform forward calculation on the test set samples, generate prediction results, calculate the accuracy metrics, and obtain the final model evaluation metrics as the performance of this training model to end the process.
[0059] Furthermore, Step 1 specifically includes the following steps:
[0060] Step 1.1: The training dataset for sequential recommendation includes several sequences. Each sequence includes users, items, and the ratings of users for items. Delete the users and items with less than 5 ratings, and then group them according to the interactive items of users and sort them in ascending order of time;
[0061] Step 1.2: For the items obtained after the previous filtering, use a crawler tool to crawl the pictures of the items as visual information according to the url provided by the dataset, then concatenate the title, category and description of the product as text information, and finally record the price information and category information of the product;
[0062] Step 1.3: Use the pre-trained multi-modal item encoder Clip to encode the text information and visual information of the item respectively, and obtain the text representation and visual representation of the item respectively, as follows:
[0063]
[0064] where, and represent the feature vectors of the text information and visual information of item i after encoding respectively, represent the text data and picture data of item i respectively;
[0065] Step 1.4: Adopt the leave-one-out data splitting strategy to divide the data, that is, use the last sample of each interaction sequence as the test set, the second-to-last sample as the validation set, and the remaining all interaction sequences as the training set; for the historical interaction sequences of each user, according to the definition in the baseline model, define its maximum length as 20, and the interaction sequences with less than 20 items are filled with <pad>Pad the interaction sequences, and for sequences with more than 20 items, truncate the most recent 20 interactions;
[0066] Step 1.5: During model training, each sequence and its corresponding modality information are used as input data, and the next interaction item in each interaction sequence is used as the supervision signal for the sequence recommendation model.
[0067] Furthermore, in Step 2.1, the visual representation and text representation of the item are concatenated in parallel to obtain the modality information representation of the item, as shown in the following formula:
[0068]
[0069] where Linear represents the linear function, Concat represents the concatenation operation, and represent the text and visual representations mapped to a suitable recommendation semantics, and e i represents the modality information representation of item i.
[0070] Furthermore, in Step 2.2, a specific masking mechanism is adopted, that is, the last position in the input item sequence is always replaced with the [mask] token during training; for the [mask] tokens in other positions in the sequence, three replacement strategies are adopted: First, with an 80% probability, the [mask] token is replaced with the original item; second, with a 10% probability, it is replaced with a randomly selected item; third, with a 10% probability, the original item remains unchanged.
[0071] Furthermore, the price information of the item in Step 2.7 is not the original price data, but the price level after conversion; since the price distribution of items within a certain category conforms to a logical distribution, price discretization is achieved by making the probabilities of each interval equal when dividing the intervals; specifically, for an item i with an absolute price of p i whose price range for its category is [min, max], when the price needs to be divided into ρ price levels, then its price level information is calculated as follows:
[0072]
[0073] where round represents the rounding integer calculation, and Φ(x) represents the cumulative distribution function of the logical distribution, and its formula is as follows:
[0074]
[0075] where μ is the expected value and δ is the standard deviation.
[0076] Furthermore, the specific calculation formula for the model evaluation index in Step 3.6 is as follows:
[0077] For Recall@10, that is, recall rate @10, it measures the ability of the model to cover the true item labels among the top 10 in the recommendation list. The formula is as follows:
[0078]
[0079] When calculating, for each user, a list of the top 10 recommended items is given; then it is statistically determined whether the true item label, that is, the next item the user actually interacts with, is included in the predicted list. If it is included, it is recorded as 1 hit, otherwise as 0; the sum of the hit times for all users is divided by the total number of users to obtain the value of Recall@10;
[0080] For NDCG@10, that is, Normalized Discounted Cumulative Gain @10, it evaluates the position-weighted quality of the target item among the top 10 in the model's recommendation list, emphasizing that the relevance contribution of items ranked higher is greater. The calculation formula is as follows:
[0081]
[0082] Among them, DCG@10 is the Discounted Cumulative Gain, and IDCG@10 is the Ideal DCG, which is the DCG value after sorting the true item labels in the optimal order; the specific calculation formula of DCG@10 is as follows:
[0083]
[0084] Among them, rel i represents the relevance of the item in the i-th position. If it is a true item label, it is 1, otherwise it is 0;
[0085] For MRR@10, that is, Mean Reciprocal Rank @10, it measures the average ranking quality of the true item label in the Top10 recommendation list predicted by the model. The higher the ranking, the higher the score. The calculation formula is as follows:
[0086]
[0087] Among them, rank i is the position of the true item label of user i in the recommendation list. If it does not enter the top 10, it will not be included in the calculation; N is the total number of users.
[0088] Furthermore, in step 4.2, based on the loss L, the parameter gradients are calculated layer by layer through the backpropagation algorithm and the model parameters are updated; the core formula is:
[0089]
[0090] Among them, represents the weight gradient, represents the bias gradient, W old represents the old weight matrix, W new represents the updated weight matrix, b old represents the old bias vector, b new represents the updated bias vector, η represents the learning rate;
[0091] When performing gradient updates, the AdamW optimizer is used, which combines momentum and adaptive learning rate adjustment to dynamically adjust the learning rate of each parameter through first-order moment and second-order moment estimation; and as training progresses, dynamic learning rate scheduling ExponentialLR is used, and the learning rate decays to the original 0.9 after each training; gradient clipping is used to limit the maximum norm of the gradient.
[0092] The beneficial effect of adopting the above technical solution is that the multimodal sequence recommendation method provided by the present invention integrates user price preference and interest preference. While considering the user interest preference affected by the visual and textual information of the product, it focuses on the key factor affecting the user's purchase of the product - price, introduces the user price preference affected by the price, and makes recommendations based on the fused preference. The performance of the present invention on Amazon's four data sets is better than the baseline model that does not introduce price preference. At the same time, thanks to the above technical solution, the present invention can achieve model convergence more quickly in a short time, ensuring that the model is optimized in the right direction. The overall improved accuracy and higher training efficiency prove the rationality and advancement of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0093] Figure 1 An overall flow chart of the multimodal sequence recommendation method provided by an embodiment of the present invention;
[0094] Figure 2 A data preprocessing flow chart provided for an embodiment of the present invention;
[0095] Figure 3 A flow chart of model training provided by an embodiment of the present invention;
[0096] Figure 4 A schematic diagram of multi-task learning provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0097] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0098] When analyzing the historical price data of user interactions in this embodiment, it is found that price factors are also important factors affecting users' purchase intentions. By introducing commodity prices and corresponding category information into the existing multi-modal sequence recommendation to achieve more effective recommendations. Specifically, the method of this embodiment divides users' preferences into two categories, namely interest preferences, which are mainly affected by the visual and text modality information of commodities; and price preferences, which are affected by commodity prices. First, use a multi-modal item encoder to encode the visual information and text information of the commodity and input it into a sequence encoder to perform the Mask Item Prediction task to mine users' interest preferences. Then, input the price information and category information in the historical sequence into an encoder and then into a sequence encoder with unidirectional attention, and then perform the price and category predictions of the next item respectively to mine users' price preferences. Finally, organically integrate the mined interest preferences and price preferences in the inference stage to achieve more effective recommendations. As Figure 1 shown, the method of this embodiment is described as follows.
[0099] Step 1: Data preprocessing; Mark a single item as v, and use to represent the text information, visual information, price information, and category information of the i-th item respectively; As Figure 2 shown, it specifically includes the following steps:
[0100] Step 1.1: The training dataset for sequence recommendation includes several sequences. Each sequence includes users, items, and users' ratings of items. Delete users and items with less than 5 ratings, and then group the users according to the interactive items and sort them in ascending order of time.
[0101] Step 1.2: According to the items obtained after the filtering in the previous step, use a crawler tool to crawl the pictures of the items as visual information according to the url provided by the dataset, then concatenate the title, category, and description of the commodity as text information, and finally record the price information and category information of the commodity.
[0102] Step 1.3: Use the pre-trained multi-modal item encoder Clip to encode the text information and visual information of the item respectively, and obtain the text representation and visual representation of the item respectively, as follows:
[0103]
[0104] Among them, and respectively represent the feature vectors after encoding the text information and visual information of item i, respectively represent the text data and picture data of item i.
[0105] Step 1.4: Divide the data using the leave-one-out data splitting strategy, that is, use the last sample of each interaction sequence as the test set, the penultimate sample as the validation set, and the remaining all interaction sequences as the training set; for the historical interaction sequences of each user, according to the definition in the baseline model (MMSRec), define its maximum length as 20, and for interaction sequences with less than 20 items, use <pad>To complete the interaction sequence of more than 20 items, the 20 most recent interactions are captured.
[0106] Step 1.5: When training the model, each sequence and the corresponding modal information are used as input data, and the next interactive item in each interactive sequence is used as the supervision signal of the sequence recommendation model.
[0107] Step 2: Model construction and loss calculation; Figure 3 As shown, the specific steps are as follows:
[0108] Step 2.1: Separately transform the text feature vectors of the items and the visual feature vector The input is sent to the linear mapping module Linear, which extracts the item representation that is more suitable for the recommendation semantics from the general semantic information, and then splices the visual representation and textual representation of the item into a whole in parallel, as shown in the following formula:
[0109]
[0110] Among them, Linear represents a linear function, Concat represents a concatenation operation, and The representation is mapped into textual and visual representations suitable for recommendation semantics, e i Represents the modal information representation of item i.
[0111] Step 2.2: According to the historical interaction behavior sequence corresponding to user u, the corresponding item modal representations are concatenated into the item representation sequence S of user u u ={e1,e2,e3…,e n }; Then the representation sequence S u The modal information representation of each item i in is replaced with a special tag [mask] with a probability of p = 0.2. The task of the model is to predict the original item based on the contextual content information of the sequence in which the original item is located. Since the model uses multiple modal representation embeddings as input to represent each item, for the masked item, all its associated modal representation embeddings are replaced with the tag [mask]; the item at the last position is represented by e n must be masked; for the [mask] tags at other positions in the sequence, this embodiment adopts three replacement strategies: the first, 80% probability of replacing the [mask] tag with the original item; the second, 10% probability of replacing it with a randomly selected item; the third, 10% probability of keeping the original item unchanged. The user history interaction item representation sequence after masking is The possible expressions are as follows:
[0112]
[0113] Step 2.3: The user historical interaction item representation sequence obtained in Step 2.2 is summed with the position embedding and the modality embedding Modal u , and then input into the sequence encoder, which is essentially a Transformer. The specific formula is as follows:
[0114]
[0115] where represents the context encoding result of the masked sequence.
[0116] Since the modality representation of the commodity is concatenated, the output result of the final Transformer is also two-channel. The maximum pooling layer is used to extract significant feature information from all modalities, so as to obtain the unified representation of item i The formula is as follows:
[0117]
[0118] where is the text channel feature, is the visual channel feature, and H i is the two-channel feature. MaxPooling represents the maximum pooling calculation, that is, the maximum value of each dimension in the two-channel feature is taken and merged into a single-channel feature as the unified representation of item i.
[0119] Step 2.4: The modality information representations e of all candidate items i are concatenated into a representation sequence, added with the position embedding and the modality embedding Modal u , and then input into the sequence encoder in Step 2.3. Different from Step 2.3, the representation sequence of candidate items does not perform the masking operation. Then, the maximum pooling layer is also used to obtain the unified representation E of all candidate items I , where n represents the number of candidate items.
[0120] Step 2.5: The unified representation of item i in Step 2.3 is multiplied by the unified representations of all candidate items in Step 2.4 to obtain the prediction of the interest preference score of the model for the item at the masked position. The specific formula is as follows:
[0121]
[0122] where τ is the temperature coefficient, and Matmul represents matrix multiplication; Indicates the prediction of the interest preference score for the item at the masked position.
[0123] Step 2.6: Calculate the loss L of predicting the masked item based on the user's interest preference using the cross-entropy loss function I , and the specific formula is as follows:
[0124]
[0125] where y i represents the true item label of the masked item and is also the supervision signal for recommendation.
[0126] Step 2.7: According to the historical interaction behavior sequence corresponding to user u, find the original price and category id of the item. Before using the price, the original price needs to be converted into a price level. Since the price distribution of items within a certain category conforms to a logical distribution, in this embodiment, the discretization of the price is achieved by making the probabilities of each interval equal when dividing the intervals. Specifically, for item i with an absolute price of p i , the price range of its category is [min, max]. When the price needs to be divided into ρ price levels, then its price level information is calculated by the following formula:
[0127]
[0128] where round represents the rounding integer calculation, and Φ(x) represents the cumulative distribution function of the logical distribution, and its formula is as follows:
[0129]
[0130] where μ is the expected value and δ is the standard deviation.
[0131] Then, the transformed price level and category id are respectively passed through the embedding layer to obtain the vector representation price i of the price and the vector representation category i of the category. Then, the price representation sequence, the category representation sequence, and the position embedding are added and input into another one-way attention sequence encoder, which is essentially a Transformer, and the specific formula is as follows:
[0132]
[0133] h u = Last(H u );
[0134] where Price u = {price1, price2, price3…, pirce n}, representing the sequence of interaction item price representations of user u, Category u = {category1, category2, category3…, category n}, representing the sequence of interaction item category representations of user u, H u represents the encoded result after adding the three sequences using an encoder. Last represents taking the representation at the last position as the price preference representation h of the user u .
[0135] Step 2.8: Multiply the price preference representation h of the user u by the price vectors and category vectors of all candidate items respectively to obtain the preference scores for the price and category of the next item. The specific formula is as follows:
[0136]
[0137] where E P , E C represent the candidate price matrix and category matrix respectively. represent the preference scores for the price and category of the next item respectively.
[0138] Step 2.9: Use the cross-entropy loss function to calculate the loss L P for predicting the price of the next item and the loss L C for predicting the category based on the user's historical interaction price and category information. The specific formula is as follows:
[0139]
[0140] where represent the true price level label and category label of the next item in the user interaction sequence respectively, which are the supervision signals for prediction.
[0141] Step 2.10: After obtaining the loss L I from Step 2.6 and the loss L P and the loss L C from Step 2.9, combine the three parts into a multi-task learning framework as shown in Figure 4 . Finally, the objective loss function of the model is expressed as follows:
[0142]
[0143] where θ represents all trainable parameters, λ P and λ C is a regularization factor that weighs the importance of the price prediction task and the category prediction task.
[0144] Step 3: Evaluate the model. The specific steps are as follows:
[0145] Step 3.1: When performing inference, the item representation sequence S of user u u No longer performs random masking, but only masks the item id at the last position. At this time, the item representation sequence is For Input it into the sequence encoder using the same operation as in Step 2.3 to obtain the user interest preference representation at the last position Perform matrix multiplication with the unified representation E of all candidate items I to obtain the interest preference scores of user u for all candidate items. The specific formula is as follows:
[0146]
[0147] where represents the context encoding result of the masked sequence. is the interest preference score of user u for all candidate items.
[0148] Step 3.2: When performing inference in the price module, the price preference representation h of the user u Perform matrix multiplication with the price matrix E of all candidate items respectively P and the category matrix E C to obtain the price preference scores of user u for all candidate items and the category preference scores Different from the training stage, there is no need to add the temperature coefficient τ. The specific formula is as follows:
[0149]
[0150] Step 3.3: According to the three preference score vectors obtained in Step 3.1 and Step 3.2, perform preference fusion. The specific formula is as follows:
[0151]
[0152] where β P and β C are influence factors that weigh the importance of the user's price preference and category preference.
[0153] Step 3.4: Sort the comprehensive preference scores of all candidate items obtained in the previous step in descending order to generate a sorted list, and then take the top K items (e.g., K = 10) as the final recommendation results. The specific formula is as follows:
[0154] Ranked_List u = argsort(Total_Preference u );
[0155] Rec_List u = {j | j ∈ Ranked_List u [1:K]};
[0156] Among them, argsort is a sorting function whose core function is to return the sorted index values of the elements in the array, corresponding to the labels of the candidate items. Ranked_List u is the list of labels of all candidate items sorted in descending order according to the preference scores. The item corresponding to the first label in the list has the highest score. Rec_List u is the recommendation list for the user obtained by intercepting the labels of the first K items.
[0157] Step 3.5: Calculate the model evaluation metrics based on the recommendation list given by the model and the true label of the user's next interaction item. Calculate the recall rate (Recall), normalized discounted cumulative gain (NDCG), and mean reciprocal rank (MRR) respectively.
[0158] For Recall@10 (recall rate @10), it measures the ability of the model to cover the true item labels among the top 10 items in the recommendation list. The formula is as follows:
[0159]
[0160] When calculating, for each user, a list of the top 10 recommended items is given; then it is statistically determined whether the true item label (i.e., the next item the user actually interacts with) is included in the predicted list; if it is included, it is recorded as 1 hit, otherwise it is recorded as 0; the sum of the hit times for all users is divided by the total number of users to obtain the value of Recall@10. And so on, calculate Recall@5, Recall@20, Recall@50.
[0161] For NDCG@10 (normalized discounted cumulative gain @10), it evaluates the position-weighted quality of the target items among the top 10 items in the model's recommendation list, emphasizing that the correlation contribution of the items ranked higher is greater. The calculation formula is as follows:
[0162]
[0163] Among them, DCG@10 is the discounted cumulative gain, and IDCG@10 is the ideal DCG, which is the DCG value after sorting the true item labels in the optimal order. The specific formula for DCG@10 is as follows:
[0164]
[0165] Among them, rel i represents the relevance of the i-th item. If it is the true item label, it is 1; otherwise, it is 0. And so on, calculate NDCG@5, NDCG@20, and NDCG@50.
[0166] For MRR@10 (Mean Reciprocal Rank@10), it measures the average ranking quality of the true item label in the Top10 recommended list predicted by the model. The higher the ranking, the higher the score. The calculation formula is as follows:
[0167]
[0168] Among them, rank i is the position of the true item label of user i in the recommended list (if it does not enter the top 10, it is not involved in the calculation), and N is the total number of users. And so on, calculate MRR@5, MRR@20, and MRR@50.
[0169] Step 4: Implementation of model training and evaluation. The specific process includes:
[0170] Step 4.1: Execute Step 2 on the training dataset, perform forward propagation, and calculate the loss loss, that is, the loss in Step 2.10.
[0171] Step 4.2: Backward propagation and gradient update. When performing backward propagation, calculate the parameter gradients layer by layer through the backward propagation algorithm and update the model parameters. The core formula is:
[0172]
[0173] Among them, represents the weight gradient, represents the bias gradient, W old represents the old weight matrix, W new represents the updated weight matrix, b old represents the old bias vector, b new represents the updated bias vector, and η represents the learning rate.
[0174] When performing gradient updates, this embodiment uses the AdamW optimizer, which combines Momentum and adaptive learning rate adjustment (such as RMSProp), and dynamically adjusts the learning rate of each parameter through first-order moment (mean) and second-order moment (variance) estimation. And as the training progresses, this embodiment uses the dynamic learning rate scheduler ExponentialLR, and the learning rate decays to 0.9 of the original after each training. A larger learning rate is used at the beginning to accelerate convergence, and the learning rate is reduced later to finely adjust the parameters and avoid oscillations. This embodiment also uses Gradient Clipping to limit the maximum norm (max_grad_norm) of the gradient, prevent the gradient explosion problem, ensure the stability of the optimization direction, and avoid the out-of-control parameter update caused by extreme gradient values.
[0175] Step 4.3: Perform the model evaluation in step 3 on the validation set to obtain the model accuracy performance, that is, the evaluation metrics in step 3.6, and compare it with the optimal Recall@10 obtained currently. If it exceeds the optimal metric, update the optimal Recall@10 and save the current model training parameters. Otherwise, use the early-stop strategy, that is, judge whether the accuracy has not improved for 10 consecutive training epochs. If so, stop the training and enter step 4.4 for testing. If not, return to step 4.1 to continue training.
[0176] Step 4.4: Perform testing, load the optimal model parameters, perform forward calculation on the test set samples, generate prediction results, calculate the accuracy metrics, and obtain the final model evaluation metrics as the performance of this training model to end the process.
[0177] This embodiment compares with other methods on four publicly available real-scenario datasets, namely: Amazon Beauty, Amazon Sports, Amazon Home, and Amazon Games. The dataset statistics are shown in Table 1.
[0178] Table 1 Dataset Statistics
[0179] Beauty Sports Home Games Number of users 16772 11010 35201 18875 Number of items 11386 13765 23788 10333 Number of interactions 123517 62248 242193 157600 Average number of interactions 7.36 5.65 6.88 8.35
[0180] The method of this embodiment is compared with the current SOTA model MMSRec for multi-modal sequential recommendation that does not introduce user price preference.
[0181] MMSRec: This method proposes a two-tower retrieval architecture for sequential recommendation to alleviate the inconsistency problem between the high-level feature vectors and low-level feature embeddings of item modalities.
[0182] To evaluate the performance of sequential recommendation, three commonly used metrics are employed: Recall (denoted as R), NDCG (Normalized Discounted Cumulative Gain), and MRR (Mean Reciprocal Rank).
[0183] Table 2 and Table 3 present all the metric results of this embodiment for sequential recommendation on four datasets. It can be seen that the method proposed in this embodiment has significantly improved in almost all evaluation metrics on all datasets.
[0184] Experimental Results I of Table 2 for Sequential Recommendation
[0185]
[0186]
[0187] Experimental Results II of Table 3 for Sequential Recommendation
[0188]
[0189] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope defined by the claims of the present invention.< / pad> < / pad>
Claims
1. A multimodal sequence recommendation method integrating user price preference and interest preference, characterized by: When capturing user price preference, the method inputs the price and category sequence of the goods into a unidirectional attention sequence encoder to obtain the user price representation, and then performs matrix multiplication with the price representation of the candidate item to complete the similarity calculation to obtain the price preference score; when capturing user interest preference, the pre-trained multimodal item encoder Clip is used to process the visual and textual information of the goods to obtain the item representation, and the generated item representation is input into a Transformer-based sequence encoder according to the corresponding relationship of the interaction sequence to obtain the user interest representation, and the candidate item representation is input into the sequence encoder and then matrix multiplied with the user interest representation to complete the similarity calculation, thereby obtaining the user's interest preference score for the candidate item; in the preference fusion stage, the interest preference score and the price preference score are weighted fused with an appropriate ratio to complete more comprehensive preference capture and achieve more effective sequence recommendation.
2. According to claim 1, a multimodal sequence recommendation method integrating user price preference and interest preference is characterized by: The specific implementation steps of the method include: Step 1: Data preprocessing: label each item as v and use Represent the text information, visual information, price information and category information of the ith item; the data preprocessing process includes filtering users and items with less than 5 ratings, determining the information content of the text information, visual information, price information and category information of the items, encoding the text information and visual information to obtain text representation and visual representation, and data segmentation to obtain a test set, a validation set and a training set; Step 2: Model construction and loss calculation. The specific steps include: Step 2.1: Separately transform the text feature vectors of the items and the visual feature vector The input is sent to the linear mapping module Linear, which extracts the item representation that is more suitable for recommendation semantics from the general semantic information, and then splices the visual representation and textual representation of the item into a whole in parallel; Step 2.2: According to the historical interaction behavior sequence corresponding to user u, the corresponding item modal representations are concatenated into the item representation sequence S of user u u ={e1,e2,e3…,e n },e i Represents the modal information representation of item i; then the representation sequence S u The modal information representation of each item i in is masked with a probability of p = 0.2, that is, replaced with a special tag [mask]; for the masked items, all their associated modal representation embeddings are replaced with the tag [mask]; the item at the last position is represented by e n must be masked; the user's historical interaction items after masking are represented as Step 2.3: Represent the user's historical interaction items obtained in step 2.2 as a sequence With positional embedding Modal u Sum and then input into the sequence encoder Transformer. The specific formula is as follows: in, represents the context encoding result of the mask sequence; Use the maximum pooling layer to extract salient feature information from all modalities to obtain a unified representation of item i The formula is as follows: in, is the text channel feature, is the visual channel feature, H i is a dual-channel feature, and MaxPooling represents the maximum pooling calculation, that is, taking the maximum value from each dimension of the dual-channel feature and merging it into a single-channel feature as a unified representation of item i; Step 2.4: Represent the modal information of all candidate items i The concatenated representation sequence plus position embedding and Modal u Input to the sequence encoder in step 2.
3. The difference from step 2.3 is that the representation sequence of candidate items is not masked. Then the maximum pooling layer is used to obtain the unified representation E of all candidate items. I , Where n represents the number of candidate items; Step 2.5: Unify the representation of item i in step 2.3 The unified representation E of all candidate items in step 2.4 I Perform matrix multiplication to obtain the interest preference score prediction for the items at the mask position. The specific formula is as follows: Among them, τ is the temperature coefficient, Matmul represents matrix multiplication, Represents the prediction of interest preference score for items at mask positions; Step 2.6: Use the cross entropy loss function to calculate the loss L for predicting items at masked locations based on user interest preferences I , the specific formula is as follows: Among them, y i The true object label representing the masked location is also the supervisory signal for recommendation; Step 2.7: According to the historical interaction behavior sequence corresponding to user u, the corresponding item price information and category information are converted into price vector representations through the embedding layer. i and the vector representation of the category i , and then embed the price representation sequence and category representation sequence as well as the position The added input is fed into another unidirectional attention sequence encoder Transformer. The specific formula is as follows: h u =Last(H u ); Among them, Price u ={price1,pirce2,pirce3…,pirce n }, represents the price sequence of interactive items of user u, Category u ={category1,category2,category3…,category n }, represents the interactive item category representation sequence of user u, H u It represents the result of encoding the three sequences using the sequence encoder after adding them together. Last represents the result of taking H u The vector of the last position is used as the user's price preference representation h u ; Step 2.8: Represent the user’s price preference as h u Perform matrix multiplication with the price vector matrix and category vector matrix of all candidate items respectively to obtain the preference score for the price and category of the next item. The specific formula is as follows: Among them, E P 、E C denote the candidate price matrix and category matrix respectively, Respectively represent the preference scores for the price and category of the next item; Step 2.9: Use the cross entropy loss function to calculate the loss L for predicting the next item price based on the user's historical interaction price and category information P and the loss L for category prediction C , the specific formula is as follows: in, They represent the true price level label and category label of the next item in the user interaction sequence, which are the supervisory signals for prediction; Step 2.10: Add the loss L obtained from step 2.6 I and the loss L obtained in step 2.9 P And the loss L C Afterwards, the three parts are combined into a multi-task learning framework. Finally, the objective loss function of the model is expressed as follows: Among them, θ represents all trainable parameters, λ P and λ C is a regularization factor that weighs the importance of the price prediction task and the category prediction task; Step 3: Model evaluation, the specific steps include: Step 3.1: When performing inference, the item representation sequence S of user u u Instead of random masking, only the product ID at the last position is masked. The product representation sequence is right The same operation as in step 2.3 is used to input into the sequence encoder to obtain the user interest preference representation at the last position and the unified representation E of all candidate items I Perform matrix multiplication to obtain the interest preference scores of user u for all candidate items. The specific formula is as follows: in, represents the context encoding result of the mask sequence, Score the interest preference of user u for all candidate items; Step 3.2: When the price module performs reasoning, the user's price preference representation h u Respectively with the price matrix E of all candidate items P With the category matrix E C Perform matrix multiplication to obtain the price preference scores of user u for all candidate items and category preference scores Unlike the training phase, there is no need to add the temperature coefficient τ. The specific formula is as follows: Step 3.3: According to the three preference score vectors obtained in step 3.1 and step 3.2, preference fusion is performed. The specific formula is as follows: Among them, β P and β C is a factor that weighs the importance of user price preference and category preference, Total_Preference u is the comprehensive preference score of user u for all candidate items; Step 3.4: Arrange the comprehensive preference scores of all candidate items obtained in the previous step in descending order to generate a sorted list, and then take the first K items as the final recommendation results. The specific formula is as follows: Ranked_List u =argsort(Total_Preference u ); Rec_List u ={j∣j∈Ranked_List u [1:K]}; Among them, argsort is a sorting function that returns the index value of the sorted elements in the array, corresponding to the label of the candidate item; Ranked_List u Rec_List is a list of labels of all candidate items sorted in descending order according to preference scores. The item corresponding to the first label in the list has the highest score. u It is a recommendation list for users obtained by intercepting the first K item labels; Step 3.5: Based on the recommendation list given by the model and the true label of the user's next interactive item, calculate the model evaluation indicators, including Recall@5, Recall@10, Recall@20, Recall@50, NDCG@5, NDCG@10, NDCG@20, NDCG@50, MRR@5, MRR@10, MRR@20, and MRR@50; Step 4: Implementation of model training and evaluation. The specific process includes: Step 4.1: Perform step 2 on the training data set, perform forward propagation, and calculate the loss, which is the loss in step 2.10; Step 4.2: Back propagation and gradient update: Back propagate according to the loss obtained by forward propagation, perform gradient update, and update model training parameters; Step 4.3: Perform the model evaluation in step 3 on the validation set to obtain the model accuracy performance, that is, the evaluation index in step 3.6, and then compare it with the current optimal Recall@10. If it exceeds the optimal index, update the optimal Recall@10 to the current value and save the current model training parameters; otherwise, use the early-stop strategy, that is, determine whether the accuracy has not improved for 10 consecutive training rounds. If so, stop training and go to step 4.4 to perform testing. If not, return to step 4.1 for the next round of training; Step 4.4: Execute the test, load the optimal model parameters, perform forward calculations on the test set samples, generate prediction results, calculate accuracy indicators, and obtain the final model evaluation indicators as the performance of this training model, ending the process.
3. According to claim 2, a multimodal sequence recommendation method integrating user price preference and interest preference is characterized by: The step 1 specifically comprises the following steps: Step 1.1: The training data set for sequence recommendation includes several sequences, each of which includes users, items, and user ratings of items. Users and items with less than 5 ratings are deleted, and then the interactive items are grouped by user and sorted in ascending time order. Step 1.2: Based on the items obtained after filtering in the previous step, use the crawler tool to crawl the pictures of the items as visual information according to the URL provided by the dataset, then splice the title, category and description of the product as text information, and finally record the price information and category information of the product; Step 1.3: Use the pre-trained multimodal object encoder Clip to encode the text information and visual information of the object respectively, and obtain the text representation and visual representation of the object respectively, as shown in the following two formulas: in, and Respectively represent the feature vectors of the text information and visual information of item i after encoding, Represent the text data and image data of item i respectively; Step 1.4: Use the leave-one-out data segmentation strategy to divide the data, that is, the last sample of each interaction sequence is used as the test set, the second to last sample is used as the validation set, and the remaining interaction sequences are used as the training set; for each user's historical interaction sequence, according to the definition in the baseline model, its maximum length is defined as 20, and the interaction sequence with less than 20 items is used as <pad> To complete the interaction sequence of more than 20 items, the 20 most recent interactions are intercepted;< / pad> Step 1.5: When training the model, each sequence and the corresponding modal information are used as input data, and the next interactive item in each interactive sequence is used as the supervision signal of the sequence recommendation model.
4. According to claim 2, a multimodal sequence recommendation method integrating user price preference and interest preference is characterized by: In step 2.1, the visual representation and textual representation of the object are concatenated in parallel to obtain the modal information representation of the object, as shown in the following formula: Among them, Linear represents a linear function, Concat represents a concatenation operation, and The representation is mapped into textual and visual representations suitable for recommendation semantics, e i Represents the modal information representation of item i.
5. The multimodal sequence recommendation method integrating user price preference and interest preference according to claim 2, characterized in that: In step 2.2, a specific mask mechanism is used, that is, the last position in the input item sequence is always replaced by the [mask] tag during training; for the [mask] tags at other positions in the sequence, three replacement strategies are adopted: the first, 80% probability of replacing the [mask] tag with the original item; the second, 10% probability of replacing it with a randomly selected item; the third, 10% probability of keeping the original item unchanged.
6. The multimodal sequence recommendation method integrating user price preference and interest preference according to claim 2, characterized in that: The price information of the items in step 2.7 is not the original price data, but the converted price level. Since the price distribution of commodities in a certain category conforms to the logical distribution, the price is discretized by making the probability of each interval equal when dividing the interval. Specifically, for the absolute price p i The price range of item i in its category is [min, max]. When the price needs to be divided into ρ price levels, then its price level information The calculation formula is as follows: Among them, round represents the rounding calculation, and Φ(x) represents the cumulative distribution function of the logical distribution, and its formula is as follows: Among them, μ is the expected value and δ is the standard deviation.
7. The multimodal sequence recommendation method integrating user price preference and interest preference according to claim 2, characterized in that: The specific calculation formula of the model evaluation index in step 3.6 is as follows: Recall@10 is a measure of the model’s ability to cover the labels of real items in the top 10 recommendation lists. The formula is as follows: During the calculation, for each user, a list of the top 10 recommended items is given; then the predicted list is counted to see whether it contains the real item label, that is, the next item that the user actually interacts with; if it does, it is recorded as 1 hit, otherwise it is recorded as 0; the number of hits for all users is summed up and divided by the total number of users to get the Recall@10 value; NDCG@10, or Normalized Discounted Cumulative Gain@10, is the position-weighted quality of the target item in the top 10 recommendation list of the evaluation model, emphasizing that the top rankings have a greater relevance contribution. The calculation formula is as follows: Among them, DCG@10 is the discounted cumulative gain, IDCG@10 is the ideal DCG, which is the DCG value after the real object labels are optimally sorted; the specific calculation formula of DCG@10 is as follows: Among them, rel i Indicates the relevance of the i-th item, if it is a real item label, it is 1, otherwise it is 0; MRR@10, or average reciprocal ranking@10, measures the average ranking quality of real item labels in the Top10 recommendation list predicted by the model. The higher the ranking, the higher the score. The calculation formula is as follows: Among them, rank i is the position of the real item label of user i in the recommendation list. If it is not in the top 10, it is not included in the calculation; N is the total number of users.
8. The multimodal sequence recommendation method integrating user price preference and interest preference according to claim 2, characterized in that: In step 4.2, based on the loss L, the parameter gradient is calculated layer by layer through the back propagation algorithm. and And update the model parameters; the core formula is: in, represents the weight gradient, represents the bias gradient, W old represents the old weight matrix, W new represents the updated weight matrix, b old represents the old bias vector, b new represents the updated bias vector, η represents the learning rate; When performing gradient updates, the AdamW optimizer is used, which combines momentum and adaptive learning rate adjustment to dynamically adjust the learning rate of each parameter through first-order moment and second-order moment estimation; and as training progresses, dynamic learning rate scheduling ExponentialLR is used, and the learning rate decays to the original 0.9 after each training; gradient clipping is used to limit the maximum norm of the gradient.
Citation Information
Cited By
Multi-modal sequence recommendation method based on large language model and related device
CN121233844A
Collaborative filtering recommendation method and system based on state space model
CN121597919A