Interpretable recommendation method and device based on multi-modal feature enhancement

By introducing distillation learning and feature purification technologies into the multimodal recommendation model, the multimodal feature extraction and data sparsity problems are solved, high-precision recommendation and high-quality interpretation are achieved, and user experience is improved.

CN120067440APending Publication Date: 2025-05-30ZHONGBEI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510131571.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing multimodal recommendation models are difficult to effectively extract information when processing multimodal features. Especially when facing data sparsity problems, deep neural network models cannot be well trained, resulting in limited recommendation performance.

Method used

Using distillation learning and feature purification ideas, through feature distillation training of teacher model and student model, combined with CLIP image encoder and Transformer text encoder, multimodal features are extracted and purified, and high-precision recommendation lists are generated and high-quality recommendation explanations are provided.

Benefits of technology

Through feature distillation and purification technology, the quality of user and product characterization is improved, the quality of recommendation performance and interpretation is improved, and the acceptance and satisfaction of users are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067440A_ABST
    Figure CN120067440A_ABST
Patent Text Reader

Abstract

The invention relates to an interpretable recommendation method based on multi-modal feature enhancement. The method comprises the following steps: extracting original image features, text features and behavior features in a system; a teacher model with rich knowledge is utilized to guide a student model to train text features with recommendation semantics; integrating local features and global features of the text by adopting an attention mechanism, and purifying and denoising original image features and text features by adopting a feature purifier; fusing the purified image features, text features and behavior features; and predicting a recommendation score by using a decoder and generating a textual personalized recommendation reason. According to the method, multi-modal features are fully utilized, distillation learning and feature purification thoughts are introduced, recommendation accuracy can be improved, high-quality recommendation explanation can be provided, and acceptability and satisfaction of users are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to an interpretable recommendation method and device based on multi-modal feature enhancement. Background Art

[0002] In the era of information explosion, recommendation systems are an effective way to solve information overload and have achieved success in many fields. The key lies in deeply mining users' behavior habits and interest preferences to achieve personalized and efficient information push services. Traditional recommendation algorithms usually focus on improving the accuracy of recommendation results, but their ability to explain the reasons for recommendations is insufficient. With the continuous progress of technology, the public's requirements for personalized services are getting higher and higher. Currently, users' demands are not only for accurate recommendations, but they also hope to understand the decision-making basis behind them, and interpretable recommendation algorithms have emerged as the times require. Interpretable recommendation algorithms can generate various forms of explanations, among which explanations in natural language form are the most intuitive and easy to understand, so they have gradually become the mainstream. Most early natural language explanation methods were based on predefined general text templates, filling key prediction features into the templates to generate explanations. However, this method often leads to relatively single generated explanation content, with all users using the same template, lacking personalization and diversity. To achieve more personalized text explanations, researchers began to use advanced natural language generation techniques, especially conducting in-depth exploration in multi-modal data modeling. In recommendation systems, multi-modal data includes product images, user ratings, and review texts, etc., which can show product features and user preferences from different perspectives. Although recurrent neural networks and their improved models have achieved certain results in text generation, they perform poorly in dealing with long-distance dependency relationships and have low computational efficiency. In contrast, models based on Transformer perform better in the field of natural language generation, but their application in recommendation systems still needs to be further developed.

[0003] Extracting effective features from multi-modal data is the basis of multi-modal recommendation models. The following uses the visual feature extraction process to illustrate the traditional process of feature extraction in multi-modal recommendation methods. First, a designed feature extractor is used to extract visual features. This feature extractor is generally a shallow neural network model or a deep neural network model. Subsequently, the visual features extracted from the feature extractor are used as the input of the recommendation model to learn the features of users and items. Due to the complexity and high dimensionality of multi-modal features, it is difficult for a feature extractor using a shallow neural network to extract effective information from each modality. Therefore, most recent multi-modal recommendation models adopt a deep neural network as the feature extractor, taking advantage of its powerful representation learning ability. However, the problem is that deep neural network models require sufficient training data to obtain good performance, but recommendation models often face serious data sparsity problems. Therefore, the feature extractor with a deep neural network model cannot be well trained, which will have a negative impact on the final recommendation performance. Summary of the Invention

[0004] The purpose of the present invention is to provide an interpretable recommendation method and device based on multi-modal feature enhancement. By introducing the idea of distillation learning and feature purification, making full use of multi-modal information, and purifying and denoising the extracted multi-modal features, better user and product representations can be obtained, generating a high-precision recommendation list while providing high-quality recommendation explanations.

[0005] To achieve the above object, the technical solution adopted by the present invention is:

[0006] An interpretable recommendation method based on multi-modal feature enhancement, comprising the following steps:

[0007] Obtain the user review text and the corresponding item image, obtain the user-item pair review text data, item historical review text data, and user historical review text data, and utilize the interaction information of the user and item to obtain the behavioral characteristics of the user and item;

[0008] Input the item historical review text data into the first teacher model to obtain the first item global text feature, and input the item historical review text data into the first student model to obtain the second item global text feature and item local text features;

[0009] Based on the first item global text feature and the second item global text feature, perform feature distillation training on the first student model and the first teacher model to obtain the trained first student model;

[0010] Input the user historical review text data into the second teacher model to obtain the first user global text feature, and input the user historical review text data into the second student model to obtain the second user global text feature and user local text features;

[0011] Based on the first user's global text features and the second user's global text features, perform feature distillation training on the second student model and the second teacher model to obtain the trained second student model;

[0012] Input the item image into the pre-trained CLIP image encoder to obtain the initial item image features; utilize the message propagation mechanism on the user-item interaction graph to obtain the initial user image features; respectively purify and denoise the initial user image features and the initial item image features by using the user's behavioral features and the item's behavioral features to obtain the final user and item image features;

[0013] Input the item's local text features into the first attention mechanism layer, use the second item's global text features to constrain the item's local text features, and use the item's behavioral features to purify and denoise the item's text features to obtain the final item text features; input the user's local text features into the second attention mechanism layer, use the second user's global text features to constrain the user's local text features, and use the user's behavioral features to purify and denoise the user's text features to obtain the final user text features;

[0014] Combine the item's behavioral features, the final item image features, and the final item text features to obtain the item's fusion features; combine the user's behavioral features, the final user image features, and the final user text features to obtain the user's fusion features;

[0015] Jointly model the user's fusion features, the user's behavioral features, the item's fusion features, the item's behavioral features, and the user-item pair review text data by inputting them into the decoder to obtain the output list of the decoder; input the output vectors in the decoder output list into the multi-layer perceptron to obtain the predicted score of the user for the item, and input them into the linear layer to generate the recommendation explanation.

[0016] Preferably, the first teacher model and the second teacher model are both CLIP text encoders; the first student model and the second student model are both Transformer encoders.

[0017] Preferably, the feature distillation loss function is:

[0018]

[0019] where t idist is the first item's global text feature, t icls is the second item's global text feature, t udist is the first user's global text feature, t ucls is the second user's global text feature.

[0020] Preferably, the user initial image features are obtained by using a message propagation mechanism on the user-item interaction graph as follows:

[0021]

[0022] where N u represents the set of interaction items of a specific user, is the initial image feature of the item, represents the user's initial image feature.

[0023] Preferably, the user initial image features and the item initial image features are purified and denoised respectively by using the user's behavior features and the item's behavior features to obtain the final user and item image features, which are expressed as:

[0024]

[0025] where is the initial image feature of the item, is the user's initial image feature, i id is the behavior feature of the item, u id is the behavior feature of the user, W 1 represents the weight vector, b 1 represents the bias coefficient, ⊙ represents element-wise multiplication, σ(·) is the ReLU function, i img represents the final image feature of the item, u img represents the final image feature of the user.

[0026] Preferably, the item text features are purified and denoised by using the item's behavior features to obtain the final item text features; the user text features are purified and denoised by using the user's behavior features to obtain the final user text features, which are expressed as:

[0027]

[0028] where W 2 represents the weight vector, b 2 represents the bias coefficient, σ(·) is the ReLU function, is the intermediate feature obtained by passing the item local text features through the attention mechanism layer, is the intermediate feature obtained by passing the user local text features through the attention mechanism layer, T i represents the final text feature of the item, T u represents the final text feature of the user.

[0029] Preferably, the behavioral features of the item, the final item image features, and the final item text features are combined to obtain the fused features of the item; the behavioral features of the user, the final user image features, and the final user text features are combined to obtain the fused features of the user, which is expressed as:

[0030] m i = Concat(i id , T i , i img )

[0031] m u = Concat(u id , T u , u img )

[0032] Among them, m i represents the fused features of the item, and m u represents the fused features of the user.

[0033] Preferably, the fused features of the user, the behavioral features of the user, the fused features of the item, the behavioral features of the item, and the user-item pair review text data are input into the decoder for joint modeling to obtain the output list of the decoder; the output vector in the decoder output list is input into the multi-layer perceptron to obtain the predicted score of the user for the item, and input into the linear layer to generate the recommendation explanation. The specific steps include:

[0034] First, the review text data of the user-item pair passes through the embedding layer to obtain the initial representation of the user-item pair review text s represents the number of review words. Then, the behavioral features i id of the item, the behavioral features u id of the user, the fused features m i of the item, the fused features m u of the user, and the initial representation of the user-item pair review text are used as the initial sequence and input into the decoder for joint modeling to obtain the output of the decoder, where L represents passing through L layers of Transformer. Then, the first output vector is input into the multi-layer perceptron network to predict the score of the user for the product. The calculation process is as follows:

[0035]

[0036] Among them, is the weight vector, is the bias coefficient, It represents the predicted score of the user for the item, and then the fifth to the last output vectors are input into the linear layer to calculate the generation probability of each word, and then the generated explanatory text is obtained.

[0037] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it executes the steps in the above-mentioned interpretable recommendation method based on multi-modal feature enhancement.

[0038] The present invention also provides an interpretable recommendation device based on multi-modal feature enhancement, including:

[0039] A memory for storing software application programs;

[0040] A processor for executing the software application program, and each program of the software application program correspondingly executes the steps in the above-mentioned interpretable recommendation method based on multi-modal feature enhancement.

[0041] The present invention introduces the idea of feature distillation, combines the teacher model, i.e., the large-scale pre-trained model, and the student model, i.e., the small-scale un-pre-trained model. The pre-trained model contains richer knowledge, and the extracted features can guide the student model more effectively and accurately, enabling it to obtain better training and thus improving the final recommendation performance. Moreover, as a pre-trained model, CLIP is only used for feature extraction and does not participate in the training iteration of parameters, which can save time in the overall training process.

[0042] The present invention introduces the ideas of distillation learning and feature purification, makes full use of multi-modal information, and purifies and denoises the extracted multi-modal features, so as to obtain better user and product representations, generate a high-precision recommendation list while providing high-quality recommendation explanations, and improve the acceptance and satisfaction of users. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is the architecture diagram of the algorithm proposed by the present invention;

[0044] Figure 2 is an example diagram of user-item interaction;

[0045] Figure 3 is the schematic diagram of the network structure of the item attention mechanism layer. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The following further describes the present invention in conjunction with the accompanying drawings and specific embodiments.

[0047] In the following description, the terms "first / second" only distinguish similar objects and do not represent a specific order for the objects. Understandably, "first / second" can be interchanged in a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0048] Before further elaborating on the embodiments of the present invention, the nouns and terms involved in the embodiments of the present invention are described. The nouns and terms involved in the embodiments of the present invention are applicable to the following explanations.

[0049] 1) User-item review text data: The review text of an item by a user.

[0050] 2) Item historical review text data: The review texts of an item by all users.

[0051] 3) User historical review text data: The review texts of all items by a user.

[0052] 4) Feature distillation: Feature distillation is a deep learning model compression and knowledge distillation technology that aims to improve the performance and generalization ability of a student model (usually a smaller and more computationally efficient model) by extracting feature information from a teacher model (usually a large and complex model) and transferring it to the student model.

[0053] 5) CLIP model: The CLIP model is a multi-modal pre-training model based on contrastive learning. The CLIP model mainly consists of two core components: an image encoder and a text encoder. The CLIP model is trained with a large amount of image and text data, enabling the model to understand the semantic relationship between image content and related text.

[0054] As Figure 1 shown, an interpretable recommendation method based on multi-modal feature enhancement of the present invention includes the following steps:

[0055] S1. Obtain the user review text and the corresponding item image, obtain the user-item review text data, the item historical review text data, and the user historical review text data, and use the interaction information of the user and the item to obtain the behavioral characteristics of the user and the item.

[0056] Specifically, a real dataset containing rich review texts is obtained through the Internet, and an image of each item is obtained using a crawler algorithm. After digitally encoding and preprocessing the review texts in the dataset, the preprocessed user-item pair review text data, user historical review text data, and item historical review text data are obtained. The data encoding and preprocessing specifically refer to: digitally encoding the words in the review text according to the word numbers in the vocabulary V composed of all review words. At the same time, the reviews made by each user and each item are preprocessed to obtain the user-item pair review text data, the historical review text data of this user and this item;

[0057] Specifically, the ID information and ratings of users and items respectively pass through the embedding layer to obtain the behavioral characteristics of users and the behavioral characteristics of items.

[0058] S2. Input the item historical review text data into the first teacher model to obtain the first item global text feature, and input the item historical review text data into the first student model to obtain the second item global text feature and the item local text feature;

[0059] Based on the first item global text feature and the second item global text feature, perform feature distillation training on the first student model and the first teacher model to obtain the trained first student model;

[0060] Input the user historical review text data into the second teacher model to obtain the first user global text feature, and input the user historical review text data into the second student model to obtain the second user global text feature and the user local text feature;

[0061] Based on the first user global text feature and the second user global text feature, perform feature distillation training on the second student model and the second teacher model to obtain the trained second student model;

[0062] Specifically, the first teacher model and the second teacher model are pre-trained CLIP text encoders, and the first student model and the second student model are un-pre-trained Transformer encoders,

[0063] More specifically, the preprocessed item historical review text data is respectively input into the first teacher model, that is, the pre-trained CLIP text encoder, and the first student model, that is, the un-pre-trained Transformer encoder, to obtain the first item global text feature, the second item global text feature, and the item local text feature. In this embodiment, a 2-layer Transformer encoder is used, and each layer includes two sub-layers: a multi-head self-attention layer and a feed-forward neural network layer. The calculation process is as follows:

[0064]

[0065] Among them, represents the preprocessed item historical review text data, t idist ∈R 1×d represents the first item global text feature, l ∈ [1, L], h ∈ [1, H] represents the h-th attention head of the corresponding layer, represents the output of the l-th layer of the multi-head self-attention, is the output of the (l - 1)-th layer encoder, d z = d / H represents the dimension of each attention head, d represents the feature embedding dimension, FFN represents the feed-forward neural network. Therefore, the text feature of the final item in the first student model can be expressed as t icls ∈R 1×d represents the second item global text feature, t ik ∈R 1×d represents the vector representation of the k-th word where k ∈ [1, n]. Similarly, the first user global text feature t uidst ∈R 1×d and the text feature of the user in the second student model can be obtained

[0066] More specifically, by performing feature distillation on the global text features in the teacher model and the global text features in the student model, the student model is promoted to learn the knowledge of the large-scale pre-trained teacher model, the difference between the output text features of the two is reduced, and the encoding ability of the student model is improved. The following distillation loss function is designed as the optimization objective:

[0067]

[0068] where t idist is the first item global text feature, t icls is the second item global text feature, t udist is the first user global text feature, t ucls is the second user global text feature;

[0069] S3. Input the item image into the pre-trained CLIP image encoder to obtain the initial item image feature; use the message propagation mechanism on the user-item interaction graph to obtain the initial user image feature; adopt the behavioral features of the user and the behavioral features of the item to purify and denoise the initial user image feature and the initial item image feature respectively to obtain the final user and item image features;

[0070] Specifically, input the item image into the pre-trained CLIP image encoder to obtain the initial item image feature, and the calculation process is as follows:

[0071]

[0072] Among them, p represents an image of an item, represents the initial image feature of the item. Then, combined with the user-item interaction graph, the initial image feature of the user is further obtained, where the user-item interaction graph is as Figure 2 shown, and the calculation process is as follows:

[0073]

[0074] Among them, N u represents the set of interaction items of a specific user, where the set of interaction items of a specific user refers to the IDs of all items that have interacted with this user, represents the initial image feature of the user. Finally, with the help of latent factors, the image modal features are purified to obtain the final image feature of the item and the final image feature of the user, and the calculation process is as follows:

[0075]

[0076] Among them, is the initial image feature of the item, is the initial image feature of the user, i id is the behavior feature of the item, u id is the behavior feature of the user, W 1 represents the weight vector, b 1 represents the bias coefficient, ⊙ represents element-wise multiplication, σ(·) is the ReLU function, i img represents the final image feature of the item, u img represents the final image feature of the user.

[0077] S4. Input the item local text feature into the first attention mechanism layer, use the second item global text feature to constrain the item local text feature, and use the behavior feature of the item to purify and denoise the text feature of the item to obtain the final item text feature; input the user local text feature into the second attention mechanism layer, use the second user global text feature to constrain the user local text feature, and use the behavior feature of the user to purify and denoise the text feature of the user to obtain the final user text feature;

[0078] Specifically, after the first student model and the second student model are trained, the local item text features obtained by using the trained first student model are input into the first attention mechanism layer, and the second item global text features obtained by using the trained first student model are used to constrain the local item text features. Similarly, the local user text features obtained by using the trained second student model are input into the second attention mechanism layer, and the second user global text features obtained by using the trained second student model are used to constrain the local user text features.

[0079] Figure 3 As shown in the network structure diagram of the item attention mechanism layer, the calculation process is as follows:

[0080]

[0081] where T represents the transpose operation, and α ik represents the attention coefficient of each word, represents the intermediate feature obtained by the item passing through the attention mechanism layer. Similarly, the intermediate feature of the user is obtained.

[0082] Then, the item's behavioral features are used to purify and denoise the item's text features to obtain the final item text features, and the user's behavioral features are used to purify and denoise the user's text features to obtain the final user text features. The calculation process is as follows:

[0083]

[0084] where W 2 represents the weight vector, b 2 represents the bias coefficient, σ(·) is the ReLU function, T i , T u ∈R 1×d represent the final item text features and the final user text features respectively.

[0085] S5. Combine the item's behavioral features, the final item image features, and the final item text features to obtain the item's fusion features; combine the user's behavioral features, the final user image features, and the final user text features to obtain the user's fusion features;

[0086] Specifically, the calculation process is as follows:

[0087] m i = Concat(i id , T i , i img )

[0088] m u = Concat(u id,T u ,u img )

[0089] Among them, m i ,m u ∈R 1×d respectively represent the fusion feature of the item and the fusion feature of the user.

[0090] S6. Jointly model the fusion feature of the user, the behavioral feature of the user, the fusion feature of the item, the behavioral feature of the item, and the user-item pair review text data into the decoder to obtain the output list of the decoder; input the output vector in the decoder output list into the multi-layer perceptron to obtain the predicted score of the user for the item, and input it into the linear layer to generate a recommendation explanation.

[0091] Specifically, first, the review text data of the user-item pair passes through the embedding layer to obtain the initial representation of the review text of the user-item pair s represents the number of review words. Then, the behavioral feature i of the item id , the behavioral feature u of the user id , the fusion feature m of the item i , the fusion feature m of the user u and the initial representation of the review text of the user-item pair are used as the initial sequence to be jointly modeled in the decoder to obtain the output of the decoder where L represents passing through L layers of Transformer. In this embodiment, L is 2. Then, the first output vector is input into the multi-layer perceptron network to predict the score of the user for the commodity. The calculation process is as follows:

[0092]

[0093] Among them, is the weight vector, is the bias coefficient, represents the predicted score of the user for the item. Then, the fifth to the last output vectors are input into the linear layer to calculate the generation probability of each word, and then the generated explanation text is obtained.

[0094] In this embodiment, a start token bos is set, and the next possible word is predicted according to the probability distribution generated by it. To achieve this, a greedy decoding method is adopted, and the word with the highest probability is directly selected each time of generation. Subsequently, this predicted word is connected to the end of the word sequence to form a new sequence and input to the model. This process can be repeated continuously until the model obtains a special end token eos, or when the generated explanation statement reaches the pre-set length of 15, and finally a recommendation explanation sentence containing 15 words is generated.

[0095] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-described interpretable recommendation method based on multi-modal feature enhancement are executed.

[0096] The present invention also provides an interpretable recommendation device based on multi-modal feature enhancement, including:

[0097] A memory for storing software application programs;

[0098] A processor for executing the software application programs, and each program of the software application programs correspondingly executes the steps in the above-described interpretable recommendation method based on multi-modal feature enhancement.

Claims

1. An explainable recommendation method based on multimodal feature enhancement, characterized in that: The following steps are involved: Obtain user comment text and corresponding item images, obtain user-item pair comment text data, item historical comment text data and user historical comment text data, and use user-item interaction information to obtain user and item behavior characteristics; Inputting the historical review text data of the item into the first teacher model to obtain the global text features of the first item, and inputting the historical review text data of the item into the first student model to obtain the global text features of the second item and the local text features of the item; Based on the global text features of the first item and the global text features of the second item, the first student model and the first teacher model are subjected to feature distillation training to obtain a trained first student model; Inputting the user's historical comment text data into the second teacher model to obtain the first user's global text features, and inputting the user's historical comment text data into the second student model to obtain the second user's global text features and the user's local text features; Based on the global text features of the first user and the global text features of the second user, the second student model and the second teacher model are subjected to feature distillation training to obtain a trained second student model; The object image is input into the pre-trained CLIP image encoder to obtain the object initial image features; the message propagation mechanism is used on the user-item interaction graph to obtain the user initial image features; the user's behavior features and the item's behavior features are used to purify and denoise the user's initial image features and the item's initial image features respectively to obtain the final user and item image features; The local text features of the item are input into the first attention mechanism layer, the local text features of the item are constrained by the second global text features of the item, and the text features of the item are purified and denoised by the behavioral features of the item to obtain the final item text features; the local text features of the user are input into the second attention mechanism layer, the local text features of the user are constrained by the second global text features of the user, and the text features of the user are purified and denoised by the behavioral features of the user to obtain the final user text features; The behavior features of the item, the final item image features and the final item text features are combined to obtain the fusion features of the item; the behavior features of the user, the final user image features and the final user text features are combined to obtain the fusion features of the user; The user's fusion features, user's behavior features, item's fusion features, item's behavior features and user-item review text data are input into the decoder for joint modeling to obtain the decoder's output list; the output vector in the decoder's output list is input into the multi-layer perceptron to obtain the user's predicted score for the item, and then input into the linear layer to generate a recommendation explanation.

2. The explainable recommendation method based on multimodal feature enhancement according to claim 1, characterized in that: The first teacher model and the second teacher model are both CLIP text encoders; the first student model and the second student model are both Transformer encoders.

3. The explainable recommendation method based on multimodal feature enhancement according to claim 1, characterized in that: The feature distillation loss function is: where t idist is the global text feature of the first item, t icls is the global text feature of the second item, t udist is the first user's global text feature, t ucls is the second user's global text feature.

4. The explainable recommendation method based on multimodal feature enhancement according to claim 1, characterized in that: The message propagation mechanism is used on the user-item interaction graph to obtain the user's initial image features, as shown below: Among them, N u represents a collection of interactive items for a specific user, is the initial image feature of the object, Represents the initial image features of the user.

5. The explainable recommendation method based on multimodal feature enhancement according to claim 1, characterized in that: The user's behavior characteristics and the item's behavior characteristics are used to purify and denoise the user's initial image characteristics and the item's initial image characteristics respectively, to obtain the final user and item image characteristics, which are expressed as: in, is the initial image feature of the object, is the user's initial image feature, i id is the behavioral characteristic of the item, u id is the user's behavior feature, W1 represents the weight vector, b1 represents the bias coefficient, ⊙ represents the element-level multiplication, σ(·) is the ReLU function, i img Represents the final image features of the object, u img Represents the final image features of the user.

6. The explainable recommendation method based on multimodal feature enhancement according to claim 1, characterized in that: The behavior characteristics of the item are used to purify and de-noise the text characteristics of the item to obtain the final item text characteristics; the behavior characteristics of the user are used to purify and de-noise the text characteristics of the user to obtain the final user text characteristics, which can be expressed as: Where W2 represents the weight vector, b2 represents the bias coefficient, σ(·) is the ReLU function, It is the intermediate feature obtained by the local text feature of the item through the attention mechanism layer. is the intermediate feature obtained by the user's local text feature through the attention mechanism layer, T i Represents the final text feature of the item, T u The final text features representing the user.

7. The explainable recommendation method based on multimodal feature enhancement according to claim 1, characterized in that: The behavior feature of the item, the final item image feature and the final item text feature are combined to obtain the fusion feature of the item; the behavior feature of the user, the final user image feature and the final user text feature are combined to obtain the fusion feature of the user; expressed as: m i =Concat(i id ,T i ,i img ) m u =Concat(u id ,T u ,u img ) Among them, m i Indicates the fusion characteristics of the item, m u Represents the fusion features of the user.

8. The explainable recommendation method based on multimodal feature enhancement according to claim 1, characterized in that: The user's fusion features, the user's behavior features, the item's fusion features, the item's behavior features, and the user-item pair comment text data are input into the decoder for joint modeling to obtain the decoder's output list; the output vector in the decoder's output list is input into the multilayer perceptron to obtain the user's predicted score for the item, and then input into the linear layer to generate a recommendation explanation. The specific steps include: First, the comment text data of the user-item pair is passed through the embedding layer to obtain the initial representation of the comment text of the user-item pair. s represents the number of comment words, and then the behavior feature i of the item id , user behavior characteristics u id 、The fusion feature m of the item i , user's fusion feature m u and the initial representation of the review text for user-item pairs As the initial sequence input to the decoder for joint modeling, the output of the decoder is obtained Where L means passing through L layers of Transformer, and then inputting the output vector into the multi-layer perceptron network to predict the user's rating of the product. The calculation process is as follows: in, is the weight vector, is the bias coefficient, Represents the user's predicted rating of the item, and then the output vector is input into the linear layer to calculate the generation probability of each word, thereby obtaining the generated explanation text.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the explainable recommendation method based on multimodal feature enhancement described in any one of claims 1 to 8 are performed.

10. An explainable recommendation device based on multimodal feature enhancement, characterized in that: include: Memory for storing software applications; A processor is used to execute the software application, and each program of the software application correspondingly executes the steps of the explainable recommendation method based on multimodal feature enhancement as described in any one of claims 1 to 8.