Multi-modal interpretable recommendation method based on attention mechanism
By introducing attention mechanism and multimodal feature extraction into the recommendation system, the problems of difficult-to-understand multimodal data fusion and recommendation reasons in the prior art are solved, and the recommendation effect of high accuracy and interpretability is achieved.
Patent Information
- Application Number
- CN202510172903.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-17
AI Technical Summary
Existing recommendation systems are difficult to effectively integrate multimodal data such as text and images, and are difficult to provide easy-to-understand recommendation reasons, and varying degrees of impact on user decisions are underutilized.
The multimodal interpretable recommendation method based on attention mechanism is adopted to extract the text and image features of the product through text feature extractors and image feature extractors, and the attention mechanism is used to identify specific visual and text elements that affect users' preferences, integrate and give heterogeneous information different degrees of importance, and thus predict users' purchasing behavior and comment text.
It achieves high accuracy and interpretability recommendations, helping merchants choose appropriate product information display media, improve the efficiency of e-commerce platform recommendations, and reduce the cost of recommendations.
Smart Images

Figure CN120070002A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of recommendation systems, and particularly relates to a multi-modal interpretable recommendation method based on an attention mechanism. Background Art
[0002] Recommendation systems are the core technologies of online user-oriented platforms such as e-commerce, social media, and content sharing websites. Recommendations in such platforms usually need to meet the personalized needs of users, thereby increasing the likelihood of user purchases and satisfaction, and creating more revenue for the platforms.
[0003] Currently, recommendation systems utilize multi-source heterogeneous data, deep learning, and attention mechanisms to improve the accuracy and interpretability of recommendations. Multi-source heterogeneous data includes the text descriptions and images of products, and user reviews. Among them, the text descriptions of products can provide detailed parameter information and product details; the images of products can provide immediate and intuitive visual information to assist users in analyzing the applicability of products; user reviews can provide users' evaluations of products and purchase bases. Although existing studies have comprehensively used text and images to enhance recommendation performance and used attention mechanisms to identify important text or image information that affects user decisions, they often ignore the fusion of multi-modal data and its different degrees of influence on user decisions, and it is difficult to provide easily understandable recommendation reasons. In view of the above background and technologies, there is an urgent need for a recommendation method that can effectively extract and fuse multi-modal features related to products, and at the same time make full use of the synergy of user preferences and decisions, and design a recommendation method with high accuracy and interpretability. Summary of the Invention
[0004] The present invention is proposed to solve the deficiencies of the above-mentioned existing technologies, and provides a multi-modal interpretable recommendation method based on an attention mechanism, aiming to comprehensively utilize multi-modal data of text and images, introduce an attention mechanism to identify specific visual and text elements that affect user preferences, integrate and assign different degrees of importance to heterogeneous information, and then predict users' purchase behaviors and review texts, and give efficient and interpretable recommendations.
[0005] To achieve the above invention purpose, the present invention adopts the following technical solutions:
[0006] A multi-modal interpretable recommendation method based on an attention mechanism of the present invention is characterized by including the following steps:
[0007] Step 1, obtain the basic information in the product recommendation scenario, including: the text description of the product, image information, and the interaction behavior between the user and the product; where the user set is denoted as , let any user in be denoted as ; the product set is denoted as , let Any production record among them is denoted as ; Let represent the set of user-product interaction behaviors, where represents the user 's interaction behavior of purchasing product i; Let represent the set of product text descriptions of all products in the set, where represents the text description of product i, and , represents the c-th word in the text description of product i, represents the total number of words in the text description of product i. Let represent the j-th sub-word of the c-th word in the text description of product i, ; n represents the total number of sub-words of the c-th word in the text description of product i; Let represent the image of product i, where represents the k-th region of the image of product i, and h represents the total number of image regions; Let represent the representation of user u; Let represent the representation of product i;
[0008] Step 2: Based on the basic information in the product recommendation scenario, use the text feature extractor to obtain the semantic feature of the c-th word in the text description of product i, and use the image feature extractor to obtain the visual feature of the k-th region of product i;
[0009] Step 3: Use the attention mechanism module to fuse the representation of user u with the semantic feature to obtain the text representation of user after fusing attention to product i; Use the attention mechanism module to fuse the representation of user u with the visual feature to obtain the image representation of user after fusing attention to product i;
[0010] Step 4: Based on the text representation , image representation , interaction representation of user u and product i, obtain the feature incorporating user attention at the , using the improved GRU network to generate user The predicted probability of the th word in the predicted review of product i is used as the basis for recommending product i to user u; where represents a linear transformation operation;
[0011] Step 5, based on the text representation and the image representation , predict the preference degree of user u for product i , which is used to generate a list of products with high preference for user u.
[0012] The characteristics of the multi-modal interpretable recommendation method based on the attention mechanism described in the present invention also lie in that the step 2 includes:
[0013] Step 2.1, construct a text feature extractor based on the pre-trained BERT model;
[0014] Step 2.1.1, the text feature extractor uses formula (1) to calculate the embedding vector of the j-th sub-word of the c-th word in the text description of product i ;
[0015] (1)
[0016] In formula (1), represents the embedding function based on Wordpiece; represents the embedding function based on position;
[0017] Step 2.1.2, the text feature extractor uses the Transformer converter layer to process and obtains the semantic features of ;
[0018] Step 2.1.3, the text feature extractor uses formula (2) to obtain the semantic features of the c-th word in the text description of product i
[0019] (2)
[0020] Step 2.2, construct an image feature extractor based on the pre-trained VGGNet-19 network, and the image feature extractor uses formula (3) to extract the visual features of the k-th region of product i
[0021] (3).
[0022] Further, the said step 3 includes:
[0023] Step 3.1, calculate the attention degree of the c-th word in the text description of product i by user using formula (4), and normalize it using the softmax function to obtain the normalized attention degree of the c-th word ; ;
[0024] (4)
[0025] In formula (4), and represent two weight parameters of the first text fully connected layer; represents the ReLU function; represents the weight parameter of the second text fully connected layer; represents the bias of the first text fully connected layer; represents the bias of the second text fully connected layer; T represents transpose;
[0026] Step 3.2, calculate the text representation of product i after fusing the attention of user using formula (5):
[0027] (5)
[0028] Step 3.3, calculate the attention degree of user u to the k-th region in the image of product i using formula (6), and normalize it using the softmax function to obtain the normalized attention degree of the k-th region ;
[0029] (6)
[0030] In formula (6), and represent two weight parameters of the first image fully connected layer; represents the weight parameter of the second image fully connected layer; represents the bias of the first image fully connected layer; represents the bias of the second image fully connected layer;
[0031] Step 3.2.2: Calculate the user image representation after fusing attention for product i :
[0032] (7).
[0033] Furthermore, the said step 4 includes:
[0034] Step 4.1: Define and initialize the current time step of the improved GRU network ; Define the hidden state at the time step as and randomly initialize ; Define the predicted review text word at the time step as and initialize to an empty character; Use the text feature extractor to process according to steps 2.1 - 2.3 to obtain the semantic feature of ;
[0035] Step 4.2: The improved GRU network uses the stick-breaking method to assign different weights to the text representation , the image representation , and the interaction representation to obtain the feature incorporating user attention at the time step, and capture the feature and the semantic feature into the hidden state to update , obtaining the hidden state at the time step as ;
[0036] Step 4.2.1: Use equations (8) - (9) to calculate the weight of the d-th representation for the feature incorporating user attention at the time step, where the weight of the text representation is , the weight of the image representation is , the weight of the interaction representation is , and :
[0037] (8)
[0038] (9)
[0039] In equations (8) and (9), represents the weight parameter of the d-th representation at the th time step; ; represents the sigmoid function; represents the weight matrix of the hidden state at the th time step; represents the th weight parameter of the th representation at the th time step;
[0040] Step 4.2.2. Calculate the feature incorporating user attention at the th time step using equation (10):
[0041] (10)
[0042] In equation (10), RELU represents the non-linear activation function; represents the linear transformation operation;
[0043] Step 4.2.3. The improved GRU network obtains the hidden state at the th time step using equations (11) - (14): :
[0044] (11)
[0045] (12)
[0046] (13)
[0047] (14)
[0048] In equations (11) - (14), represents the update gate at the th time step; represents the reset gate at the th time step; represents the candidate set of the hidden state at the th time step; tanh represents the hyperbolic tangent function; represents the Hadamard product; represents the semantic feature of the predicted review word at the th time step;
[0049] Step 4.3. Obtain the th time step user Comment on product i in the th word predicted probability :
[0050] (15)
[0051] In formula (15), represents the weight matrix of the hidden state ; represents the word sequence predicted at the 1st to the time step; softmax represents the normalized exponential function.
[0052] Furthermore, step 5 includes:
[0053] Step 5.1, obtaining the final representation of product i using formula (16) :
[0054] (16)
[0055] In formula (16), represents the weight matrix for dimensional transformation; represents the parameter matrix for dimensional transformation;
[0056] Step 5.2, predicting the preference value of user for product i :
[0057] (17)
[0058] In formula (17), PredictionLayer represents a multi-layer perceptron;
[0059] Step 5.3, calculating the loss function using formula (18) , and performing backpropagation on the attention mechanism module and the improved GRU network according to the loss function to update the parameters, obtaining the trained optimal model for realizing the prediction of user preferences for products and review texts:
[0060] (18)
[0061] In formula (18), represents the set of products purchased by user u; represents the set of products not purchased by user u; is a hyperparameter; is the regularization coefficient; denotes the regularization parameter; L denotes the total number of layers of the improved GRU network.
[0062] An electronic device according to the present invention includes a memory and a processor, characterized in that the memory is used to store a program for supporting the processor to execute the multi-modal interpretable recommendation method, and the processor is configured to execute the program stored in the memory.
[0063] A computer-readable storage medium according to the present invention, characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, it executes the steps of the multi-modal interpretable recommendation method.
[0064] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0065] 1. The present invention introduces multi-modal data including text and images, and uses a feature extractor and an attention mechanism to identify specific visual and text elements that affect user preferences, which helps merchants select appropriate display media for specific product information and facilitates users to obtain product information.
[0066] 2. The present invention develops a new gated recurrent unit (GRU) network, introduces the stick-breaking method for efficiently allocating weights, integrates and assigns different degrees of importance to heterogeneous information, effectively predicts the user's comment content, and provides suggestions for merchants to improve marketing strategies and update products.
[0067] 3. The present invention incorporates image representation and text representation, accurately predicts the user's preference degree for products, improves the efficiency of recommendation on the e-commerce platform, and reduces the cost of recommendation. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 is a flowchart of the multi-modal interpretable recommendation method based on the attention mechanism of the present invention;
[0069] Figure 2 is a flowchart of the text feature extraction of the present invention;
[0070] Figure 3 is a flowchart of the visual feature extraction of the present invention.
[0071] Figure 4 is the time step flowchart of the improved GRU network of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0072] In this example, a multi-modal interpretable recommendation method based on the attention mechanism is an efficient and interpretable recommendation method designed for multi-modal data including the text description and image of a product and user reviews. The method introduces a feature extractor and an attention mechanism to identify specific visual and text elements that affect user preferences, and then predicts the degree of user preference. An improved GRU network is used to integrate text representations, image representations, and interaction representations to predict user review text. Specifically, as Figure 1 shown, the method is carried out according to the following steps:
[0073] Step 1: Obtain the basic information in the product recommendation scenario, including: the text description of the product, image information, and the interaction behavior between the user and the product; among them, the user set is denoted as , let any user in be denoted as ; the product set is denoted as , let any product in be denoted as ; let represent the set of interaction behaviors between the user and the product, where represents the interaction behavior of user purchasing product i; let represent the set of text descriptions of all products in the product set , where represents the text description of product i, and , represents the c-th word in the text description of product i, represents the total number of words in the text description of product i, let represent the j-th sub-word of the c-th word in the text description of product i, ; n represents the total number of sub-words of the c-th word in the text description of product i ; let represent the image of product i, where represents the k-th region of the image of product i, h represents the total number of image regions; let represent the representation of user u; let represent the representation of product i;
[0074] Step 2: Based on the basic information in the product recommendation scenario, use the text feature extractor to obtain the semantic feature of the c-th word in the text description of product i, and use the image feature extractor to obtain the visual feature of the k-th region of product i;
[0075] Step 2.1: Construct a text feature extractor based on the pre-trained BERT model; BERT is a deep bidirectional encoder that uses self-supervised learning to represent text as a series of vectors, as Figure 2 shown;
[0076] Step 2.1.1: The text feature extractor uses Equation (1) to calculate the embedding vector of the j-th sub-word of the c-th word in the text description of product i in; ; ;
[0077] (1)
[0078] In Equation (1), represents the Wordpiece-based embedding function; represents the position-based embedding function; the embedding vector is represented by a 768-dimensional vector.
[0079] Step 2.1.2: The text feature extractor uses the Transformer encoder layer to process and obtains the semantic feature of; ;
[0080] Step 2.1.3: The text feature extractor uses Equation (2) to obtain the semantic feature of the c-th word in the text description of product i :
[0081] (2)
[0082] Step 2.2: Construct an image feature extractor based on the pre-trained VGGNet-19 network. The image feature extractor uses Equation (3) to extract the visual feature of the k-th region of product i . VGGNet-19 is a convolutional neural network widely used in the field of computer vision and can obtain complex image features, as Figure 3 shown;
[0083] (3).
[0084] Step 3: Use the attention mechanism module to fuse the representation of user u with the semantic feature to obtain the text representation of user after fusing attention for product i; use the attention mechanism module to represent the representation Fused with visual features to obtain the user the image representation of product i after fusing attention ;
[0085] Step 3.1: Calculate the attention degree of the c-th word in the text description of product i by the user using Equation (4), and normalize it using the softmax function to obtain the normalized attention degree of the c-th word ; ;
[0086] (4)
[0087] In Equation (4), and represent two weight parameters of the first text fully-connected layer, used to map the user representation and the semantic feature to the same-dimensional space; represents the ReLU function; represents the weight parameter of the second text fully-connected layer; represents the bias of the first text fully-connected layer; represents the bias of the second text fully-connected layer; T represents transpose; Capturing the user's fine-grained text preferences at the regional level helps to clarify the impact of specific texts in the product text description on the user's decision-making process;
[0088] Step 3.2: Calculate the text representation of product i after fusing attention by the user using Equation (5):
[0089] (5)
[0090] In Equation (5), represents the attention degree of the c-th word in the text description of product i by the user after being normalized by the softmax function ;
[0091] Step 3.3: Calculate the attention degree of the k-th region in the image of product i by user u using Equation (6), and normalize it using the softmax function to obtain the normalized attention degree of the k-th region ;
[0092] (6)
[0093] In formula (6), and represent two weight parameters of the first image fully-connected layer, which are used to map the user representation and visual features into the same-dimensional space; represents the weight parameter of the second image fully-connected layer; represents the bias of the first image fully-connected layer; represents the bias of the second image fully-connected layer; Capturing the user's fine-grained visual preferences at the regional level helps to clarify the impact of specific regions in product pictures on the user's decision-making process;
[0094] Step 3.2.2, calculate the user's image representation after fusing attention to product i :
[0095] (7)
[0096] In formula (7), represents the degree of attention of user u to the k-th region in the image of product i after being normalized by the softmax function;
[0097] Step 4, based on the text representation , image representation , the interaction representation of user u and product i , obtain the feature incorporating user attention at the t-th time step , and use the improved GRU network to generate the prediction probability of the n-th word in the predicted comment of user u on product i , as the basis for recommending product i to user u; where represents a linear transformation operation.
[0098] Step 4.1, define and initialize the current time step of the improved GRU network ; define the hidden state at the t-th time step as , and randomly initialize ; define the word of the predicted comment text at the n-th time step as , and initialize to an empty character; use the text feature extractor to process according to steps 2.1 - step 2.3 to obtain , and get Semantic features ;
[0099] Step 4.2: The improved GRU network uses the stick-breaking method to assign different weights to text representations , image representations , and interaction representations to obtain the features incorporating user attention at the t-th time step , and capture the features and semantic features into the hidden state, thereby updating to obtain the hidden state at the t-th time step as ;
[0100] Step 4.2.1: Calculate the weight of the d-th representation for the features incorporating user attention at the t-th time step using Equations (8) - (9), where the weight of the text representation is , the weight of the image representation is , the weight of the interaction representation is , and . The weight calculation uses the stick-breaking method, which decomposes the generation process of the review words into the weighting of these three types of representations. This process starts with a "stick" of unit length, which represents the probability distributed among different representations. At each step, a part of the "stick" is assigned to a specific representation, and the remaining part goes to the next step. Mathematically, this process can be expressed as:
[0101] (8)
[0102] (9)
[0103] In Equations (8) and (9), represents the weight parameter of the d-th representation at the t-th time step, ; represents the sigmoid function; represents the weight matrix of the hidden state at the t-th time step; represents the weight parameter of the d-th representation at the t-th time step, .
[0104] Step 4.2.2, calculate the feature integrating user attention at the -th
[0105] time step using Equation (10):
[0106] In Equation (10), RELU represents a non-linear activation function; represents a linear transformation operation;
[0107] Step 4.2.3, the improved GRU network obtains the hidden state at the -th time step using Equations (11)-(14). The improved GRU network not only depends on the historical hidden state and semantic features , but also introduces the feature , as shown in Figure 4 . This design enables the gating mechanism of GRU to dynamically adjust the contribution of multi-modal information to the generation of the current word, enhancing the model's ability to model complex user behaviors. Mathematically, this process can be expressed as:
[0108] (11)
[0109] (12)
[0110] (13)
[0111] (14)
[0112] In Equations (11)-(14), represents the update gate at the -th time step, which helps capture long-term dependencies in the sequence; represents the reset gate at the -th time step, which helps capture short-term dependencies in the sequence; represents the Hadamard product; represents the -th semantic feature of the comment word predicted at the
[0113] Step 4.3, obtain the predicted probability of the -th word in user 's comment on product i at the -th time step using Equation (15). :
[0114] (15)
[0115] In formula (15), represents the weight matrix of the hidden state ; represents the word sequence predicted from the 1st to the time step; softmax represents the softmax function.
[0116] Step 5. Based on the text representation and the image representation , predict the preference degree of user u for product i , and use it to generate a list of products with high preference for user u.
[0117] Step 5.1. Obtain the final representation of product i using formula (16) :
[0118] (16)
[0119] In formula (16), represents the weight matrix for dimensional transformation; represents the parameter matrix for dimensional transformation; represents the image representation of product i after the user fuses attention; represents the text representation of product i after the user fuses attention;
[0120] Step 5.2. Predict the preference value of user for product i using formula (17) :
[0121] (17)
[0122] In formula (17), PredictionLayer represents a multi-layer perceptron;
[0123] Step 5.3. Calculate the loss function using formula (18) , and perform backpropagation on the attention mechanism module and the improved GRU network according to the loss function to update the parameters , , , , , , , , , , , , an optimal model after training is obtained for predicting the user's preference for products and review texts:
[0124] (18)
[0125] In formula (18), represents the set of products purchased by user u; represents the set of products not purchased by user u; is a hyperparameter used to balance the modeling of user purchase feedback and user review information; is the regularization coefficient, a hyperparameter used to control the complexity of the model; represents the regularization parameter; L represents the total number of layers of the improved GRU network.
[0126] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.
[0127] In this embodiment, a computer-readable storage medium stores a computer program, and when the computer program is run by a processor, it executes the steps of the above method.
[0128] In summary, the present invention comprehensively utilizes multi-modal data of text and pictures, introduces an attention mechanism to identify specific visual and text elements that affect user preferences, integrates and assigns different degrees of importance to heterogeneous information, and then predicts the user's purchase behavior and review texts, giving efficient and interpretable recommendations.
Claims
1. A multimodal explainable recommendation method based on attention mechanism, characterized in that: The steps include: Step 1: Obtain basic information in the product recommendation scenario, including: product text description, image information, and user interaction with the product; where the user set is recorded as ,make Any user is recorded as ; The product set is denoted as ,make Any one of the products is recorded as ;make Represents the set of interactive behaviors between users and products, where Indicates user The interactive behavior of purchasing product i; Represents a product set The set of text descriptions of all products in , where represents the text description of product i, and , represents the cth word in the text description of product i, represents the total number of words in the text description of product i, let Represents the text description of product i The jth subword of the cth word in ; n represents the text description of product i The cth word in The total number of subwords of represents an image of product i, where represents the kth region of the image of product i, h represents the total number of image regions; let represents the representation of user u; let represents the representation of product i; Step 2: Based on the basic information in the product recommendation scenario, use the text feature extractor to obtain the text description of product i The cth word in Semantic features of , use the image feature extractor to obtain the kth region of product i Visual features ; Step 3: Use the attention mechanism module to represent user u and semantic features Integration, acquisition of users Text representation after integrating attention on product i ; Use the attention mechanism module to represent user u With visual features Fusion, acquisition of users Image representation after integrating attention on product i ; Step 4: Text-based representation , Image Representation , the interaction representation between user u and product i , get the Time step into the user attention feature , using the improved GRU network to generate user The predicted reviews for product i Words The predicted probability , as the basis for recommending product i to user u; among them, Represents a linear conversion operation; Step 5: Text-based representation and image representation , predict user u’s preference for product i , used to generate a list of highly preferred products for user u.
2. The multimodal explainable recommendation method based on the attention mechanism according to claim 1, characterized in that: The step 2 comprises: Step 2.1: Build a text feature extractor based on the pre-trained BERT model; Step 2.1.1: The text feature extractor calculates the text description of product i using formula (1): The jth subword of the cth word in The embedding vector ; (1) In formula (1), represents the embedding function based on Wordpiece; represents the position-based embedding function; Step 2.1.2: The text feature extractor uses the Transformer converter layer to Process and obtain Semantic features of ; Step 2.1.3: The text feature extractor uses formula (2) to obtain the text description of product i The semantic features of the cth word in : (2) Step 2.2: construct an image feature extractor based on the pre-trained VGGNet-19 network. The image feature extractor extracts the kth region of product i using formula (3): Visual features ; (3)。 3. The multimodal explainable recommendation method based on the attention mechanism according to claim 2, characterized in that: The step 3 comprises: Step 3.1: Calculate the user Text description of product i The cth word in Degree of attention , and use the softmax function to Normalize and get the cth word Normalized attention level ; (4) In formula (4), and Represents the two weight parameters of the first text fully connected layer; Represents the ReLU function; Represents the weight parameter of the second text fully connected layer; Represents the bias of the first text fully connected layer; represents the bias of the second text fully connected layer; T represents transposition; Step 3.2: Calculate the user using formula (5) Text representation after integrating attention on product i : (5) Step 3.3: Use formula (6) to calculate the degree of attention paid by user u to the kth region in the image of product i , and use the softmax function to Normalize and get the normalized attention level of the kth region ; (6) In formula (6), and Represents the two weight parameters of the first image fully connected layer; Represents the weight parameters of the second image fully connected layer; Represents the bias of the first image fully connected layer; Represents the bias of the second image fully connected layer; Step 3.2.2: Calculate the user Image representation after integrating attention on product i : (7)。 4. The multimodal explainable recommendation method based on the attention mechanism according to claim 2, characterized in that: The step 4 comprises: Step 4.1: Define and initialize the current time step of the improved GRU network ; Definition The hidden state at the time step is , and Perform random initialization; define the The comment text words predicted at time step are , and Initialize to an empty character; use the text feature extractor according to steps 2.1-2.3 Process and obtain Semantic features of ; Step 4.2: The improved GRU network uses the stick-breaking method to represent text , Image Representation , interactive representation Assign different weights to obtain Time step into the user attention feature , and capture features and semantic features to the hidden state, thereby updating , get the The hidden state at the time step is ; Step 4.2.
1. Use equations (8) and (9) to calculate the dth representation of the Time step into the user attention feature Weight , where the text representation The weight of , Image Representation The weight of , interactive representation Weight ,and : (8) (9) In formulas (8) and (9), Indicates The weight parameter of the dth representation at time step, ; Represents the sigmoid function; Indicates Hidden state at time step The weight matrix of Indicates The time step The weight parameter of the representation, ; Step 4.2.2: Use formula (10) to calculate Time step into the user attention feature : (10) In formula (10), RELU represents the nonlinear activation function; Represents a linear conversion operation; Step 4.2.3: The improved GRU network uses equations (11) to (14) to obtain the first Hidden state at time step : (11) (12) (13) (14) In formula (11) to formula (14), Indicates Update gate for the time step; Indicates reset gate for the time step; Indicates The candidate set of hidden states at the time step; tanh represents the hyperbolic tangent function; represents the Hadamard product; Indicates Review words predicted at time step The semantic features of Step 4.3: Use formula (15) to get Time step user Reviews on product i Middle Words The predicted probability : (15) In formula (15), Indicates hidden state The weight matrix of Indicates the 1st to The sequence of words predicted at time step; softmax represents the normalized exponential function.
5. The multimodal explainable recommendation method based on attention mechanism according to claim 3, characterized in that: The step 5 comprises: Step 5.1: Use formula (16) to obtain the final representation of product i : (16) In formula (16), represents the weight matrix used for dimensional transformation; represents the parameter matrix used for dimension transformation; Step 5.2: Use formula (17) to predict user Preference value for product i : (17) In formula (17), PredictionLayer represents a multi-layer perceptron; Step 5.3: Calculate the loss function using formula (18) , according to the loss function Back-propagation is performed on the attention mechanism module and the improved GRU network to update the parameters and obtain the optimal model after training, which is used to predict user preferences for products and comment texts: (18) In formula (18), represents the set of products purchased by user u; represents the set of products that user u has not purchased; is a hyperparameter; is the regularization coefficient; represents the regularization parameter; L represents the total number of layers of the improved GRU network.
6. An electronic device, comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the multimodal explainable recommendation method described in any one of claims 1 to 5, and the processor is configured to execute the program stored in the memory.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the multimodal explainable recommendation method according to any one of claims 1 to 5 are performed.
Citation Information
Patent Citations
Interpretable recommendation method fusing implicit item preference and explicit feature preference of user
CN113420221A
Multi-view multi-mode commodity recommendation method based on attention neural network
CN114693397A
Commodity recommendation method, system and equipment based on comment information and score matrix
CN114926239A
Interpretable recommendation method based on graph neural network inference
WO2022222037A1
Cited By
Interpretable recommendation method based on user preference similarity and concept predetermination relation
CN122470760A