Scoring and comment collaborative enhancement interpretable recommendation method and system

By integrating BERTopic and graph convolutional neural networks into the recommendation system, combining sentiment analysis and dual-perspective comparative learning, correcting ratings and screening user-preferred topics, the problem of insufficient modeling of user multi-dimensional preferences in existing recommendation methods is solved, and more accurate and explainable recommendation results are achieved.

CN120670655APending Publication Date: 2025-09-19SHANXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510684372.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing recommendation methods that integrate ratings and reviews fail to effectively model users' multi-dimensional and fine-grained preferences, resulting in limited recommendation effectiveness and a lack of interpretability of recommendation results.

Method used

BERTopic-based topic analysis extracts topics from user reviews. This is then combined with sentiment analysis to collaboratively process bimodal information, correct ratings, and filter trending topics. A graph convolutional neural network learns node representations from the perspective of user-item interaction, and user representations are optimized through dual-perspective comparative learning. Finally, an acceptance metric is defined based on user attention and item recognition, re-ranking the recommendation list and providing multi-dimensional graded feedback.

Benefits of technology

The accuracy and explainability of recommendation results have been improved, which can more accurately match user needs and enhance the core competitiveness and commercial value of the platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670655A_ABST
    Figure CN120670655A_ABST
Patent Text Reader

Abstract

The invention provides a score and comment collaborative enhancement interpretable recommendation method and system, and belongs to the technical field of deep learning. According to the technical scheme, on the basis of subject extraction and sentiment analysis on comment information, score correction and subject screening are achieved by comprehensively utilizing the synergistic effect between scores and comments, and a user tendency subject set is established. And then, in combination with the user attention degree and the project recognition degree under each theme, establishing bipartite graphs under three perspectives of user-theme, project-theme and user-project, and generating node representation rich in information through multi-perspective fusion. After an initial recommendation list is obtained, an acceptability index is defined based on the attention of a user to a tendency theme and the acceptance of a group to a project, reordering of the project is achieved, meanwhile, a multi-dimensional grading feedback mechanism containing the user acceptability, the project acceptance and a recommendation grade is established, and a high-quality and explainable recommendation result is provided for the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of deep learning technology, and specifically relates to an explainable recommendation method and system with collaborative enhancement of ratings and comments. Background Art

[0002] Recommendation systems that integrate ratings and reviews aim to provide users with personalized recommendations by analyzing their ratings and reviews of items. Currently, relevant research focuses on two aspects: representation learning and modal fusion. Representation learning focuses on mining user preferences and item features from ratings and reviews. Modal fusion integrates the embedding vectors of different modalities obtained through representation learning using methods such as splicing or weighting. Considering the complex relationships between different modalities, contrastive learning is introduced to model the similarities and differences between representations in different modalities. This not only enriches the semantics of the representations, but also enhances the fusion effect of multimodal information and further optimizes the recommendation performance.

[0003] However, current research on recommendation methods that integrate ratings and reviews typically only considers the modifier effect of reviews on ratings, ignoring the guiding and supervisory role of ratings on reviews. This insufficient processing of review information makes it difficult to model users' multidimensional and fine-grained preferences, limiting recommendation effectiveness. Furthermore, existing methods often present recommendation results as lists of items based on predicted scores, lacking feedback and providing limited user reference value. Summary of the Invention

[0004] The purpose of the present invention is to solve the above problems and provide an explainable recommendation method and system with synergistic enhancement of ratings and comments. Make full use of the synergy between ratings and comments, and while correcting the original user ratings, filter user-inclined topics to comprehensively improve the quality of information. On this basis, define user attention and project recognition under each topic, construct corresponding bipartite graphs from the three perspectives of user-project, user-topic, and project-topic, learn more informative node representations, and design dual-perspective comparative learning to optimize user representations. Finally, define the acceptance index in combination with user attention and project recognition, and rearrange the initial recommendation list accordingly. At the same time, establish a multi-dimensional hierarchical feedback mechanism including user acceptance, project recognition and recommendation level to present more explainable recommendation results to users.

[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0006] An explainable recommendation method based on collaborative enhancement of ratings and reviews includes the following steps:

[0007] Step 1: Extract topics from user comments based on BERTopic technology and establish associations between users and topics, and between items and topics;

[0008] Step 2: Combine the sentiment analysis model to collaboratively process the bimodal information, implement score correction and topic screening, and obtain the user's preferred topics;

[0009] Step 3: Construct a weighted bipartite graph based on the user's revised rating of the item, and learn the node representation from the perspective of user-item interaction through a graph convolutional neural network;

[0010] Step 4: Based on the definition of user attention and project recognition, construct user-topic and project-topic bipartite graphs and learn node representations from two perspectives;

[0011] Step 5: Use adaptive fusion and dual-view contrastive learning to model the comprehensive representation of items and users respectively;

[0012] Step 6: Perform an inner product operation on the comprehensive representation of users and items to obtain the prediction score, and use a multi-task learning method to jointly optimize the recommendation loss and contrastive learning loss;

[0013] Step 7: Define acceptance indicators based on users' attention to trending topics and the group's recognition of the project, re-rank the projects, and form multi-dimensional graded feedback information.

[0014] Furthermore, the implementation method of step 1 is as follows:

[0015] First, user comments are divided into sentences, using S = {s1,s2,…,s h} represents the set of all sentences. All sentences in S are encoded using the BERT pre-training model to obtain the sentence representation set e(S) = {e(s1), e(s2),…, e(s h )};

[0016] Then, the sentences are clustered according to their semantic similarity. To avoid the problem of dimensional disaster caused by representation in high-dimensional space during clustering, the UMAP dimensionality reduction technique is used to reduce the sentence representation to d dimensions before clustering. The sentence representation set after dimensionality reduction is recorded as {u(s1),u(s2),…,u(s h )},in Represents the d-dimensional representation of the j-th sentence. The density-based HDBSCAN algorithm is used to cluster the sentence representations after dimensionality reduction. The specific formula is as follows:

[0017] HDBSCAN({u(s1),u(s2),…,u(s h )})→{C1,C2,…,C k} (1)

[0018] C g →T g,g=1,2,…,k (2)

[0019] Among them, k represents the number of generated clusters, and each cluster Corresponding to a topic T g , Represents the subject T g The d-dimensional representation of the next j-th sentence, |C g | indicates the subject T g The number of sentence tokens below;

[0020] Finally, for C g The sentence representations in are average pooled to obtain the topic T g Characterization of e(T g ), the specific formula is as follows:

[0021]

[0022] The representation of the subject e(T)={e(T1),e(T2),…,e(T k )} form a matrix

[0023]

[0024] According to the correspondence between users, sentences and topics, derive the topic set involved by user u Similarly, according to the correspondence between items, sentences and topics, the topic set involved in item i is derived

[0025] Furthermore, the implementation method of step 2 is as follows:

[0026] On the one hand, the comments are used to modify the ratings. First, the comment d of user u on item i is divided into ui Convert it into a format suitable for BERT processing. The specific formula is as follows:

[0027] e(d ui )=Tokenizer(d ui ) (5)

[0028] Next, BERT is used to calculate the sentiment score w of user u for item i ui , the specific formula is as follows:

[0029] w ui =BertModel(e(d ui )) (6)

[0030] to w ui Normalize it to be consistent with the range of scores in the dataset. The specific formula is as follows:

[0031]

[0032] Taking the Amazon dataset as an example, since the range of user ratings is {1, 2, 3, 4, 5}, the comment-level sentiment score of user u on item i is

[0033] Score the comment-level sentiment Compared with the original score r ui Weighted fusion to obtain the revised score The specific formula is as follows:

[0034]

[0035] Among them, α is the weight parameter that balances the user's original rating and the comment-level sentiment score;

[0036] The modified user-item rating matrix is ​​denoted as

[0037] On the other hand, the revised scores are used to filter topics. First, the relationship between users, items and topics is used to construct the user-topic-item ternary graph G UTI ={(u,T g ,i)|u∈U,T g ∈T u ∩T i ,i∈I}, where (u,T g ,i) indicates that user u and project i have a common topic T g ,user u has a good understanding of item i in topic T g All the sentences below constitute a comment After inputting it into the BERT model and normalizing it, we get the topic-level sentiment score of user u on item i.

[0038] Then, the Pearson correlation coefficient was used to calculate T g The topic-level sentiment score of user u and revised ratings The correlation between The specific formula is as follows:

[0039]

[0040] in, Represents the subject T g The set of items that user u has interacted with, Represents the subject T g The number of items that user u has interacted with, and Represents the topic Tg The average sentiment score of user u and the average modified score of all items he interacted with;

[0041] Finally, the topics involved by users are sorted by relevance Sort from high to low and select the top N topics to form the tendency topic set of user u

[0042] Furthermore, the implementation method of step 3 is as follows:

[0043] Based on the revised scoring matrix Constructing a weighted bipartite graph G from the perspective of user-item interaction ui (V ui ,E ui ), where V ui =U∪V represents the node set, E ui Represents the edge set, and the modified rating is used as the edge feature; the initial representation matrix of users and items is recorded as and Where d is the embedding dimension; the user representation matrix P (0) and the item representation matrix Q (0) Stack and get the initial representation matrix of user-item union Input it into GCN to update the user and item representations. The specific formula is as follows:

[0044]

[0045] Among them, E (l) represents the node representation matrix of the lth layer, W1 and W2 are learnable parameter matrices, I represents the identity matrix, The Laplace matrix representing the user-item bipartite graph is as follows:

[0046]

[0047] in, is the modified rating matrix, D is the degree matrix of the node;

[0048] Perform average pooling on the representation matrix of each layer to obtain the updated representation matrix

[0049]

[0050] Among them, L is the number of convolution layers;

[0051] This results in a user representation from the perspective of user-item interaction. and project representation

[0052] Furthermore, the implementation method of step 4 is as follows:

[0053] The more sentences involving a certain topic in a user's historical comments, the higher the user's attention to the topic. The user's attention to the topic reflects the user's own preferences to a certain extent. g ∈T′ u Attention The definition is as follows:

[0054]

[0055] Among them, S′ u Represents the set of sentences involving the tendency topic of user u;

[0056] Based on the user-topic attention matrix Constructing a bipartite graph G from the perspective of user-topic association ut (V ut ,E ut ), where V ut =U∪T represents the node set, E ut Represents the edge set, takes the user's attention to the topic as the edge feature; represents the user representation matrix P (0) and the topic representation matrix E T Stack to get the initial representation matrix of user topics It is also input into GCN. The specific formula is as follows:

[0057]

[0058] in, represents the node representation matrix of the lth layer, W3 and W4 are learnable parameter matrices, The Laplace matrix representing the user-topic bipartite graph is as follows:

[0059]

[0060] Perform average pooling on the representation matrix of each layer to obtain the updated representation matrix The specific formula is as follows:

[0061]

[0062] This results in the user representation from the perspective of user-topic association.

[0063] For a project, the degree of its recognition can be measured by analyzing the sentiment scores of the user groups that have interacted with it under various topics. Taking project i as an example, its sentiment scores under topic T gThe following sentence set is Calculated using the BERT model Each sentence in The sentiment score is obtained after normalization in will be collected The average sentiment score of the sentences in T is defined as the average sentiment score of item i in topic T g Recognition under The specific formula is as follows:

[0064]

[0065] Based on the project-theme recognition matrix Constructing a bipartite graph G from the perspective of project-topic association it (V it ,E it ), where V it =I∪T represents the node set, E it Represents the edge set, and takes the recognition degree of the project under the topic as the feature of the edge; similar to obtaining the user representation from the user-topic association perspective, the project representation matrix Q (0) and the topic representation matrix E T Stacking is performed to obtain the initial representation matrix of the project theme By inputting it into GCN, we can obtain the project representation from the perspective of project-topic association. Laplace matrix corresponding to the project-topic bipartite graph The specific formula is as follows:

[0066]

[0067] Furthermore, the implementation method of step 5 is as follows:

[0068] The project representation from the perspective of user-item interaction and the perspective of item-topic association are adaptively integrated to obtain the comprehensive project representation e i , the specific formula is as follows:

[0069]

[0070] Among them, β is a trainable parameter;

[0071] To further enrich the semantic information in user representation, a dual-perspective contrastive learning task is designed. This task optimizes user representation by maximizing the consistency of the representation of the same user under the user-item perspective and the user-topic perspective, as well as the differences in the representation of different users under the two perspectives. The specific formula is as follows:

[0072]

[0073] Among them, sim(·) represents the cosine similarity function, τ represents the temperature coefficient, represents the user representation from the perspective of user-item interaction, represents the user representation from the perspective of user-topic association, Representation of other users from the perspective of user-topic association;

[0074] The optimized user representation from the user-item perspective is used as the final comprehensive representation of the user, denoted as e u .

[0075] Furthermore, the implementation method of step 6 is as follows:

[0076] By comprehensively characterizing the user u Comprehensive representation of the project i Perform inner product operation to obtain the predicted score of user u for item i. The specific formula is as follows:

[0077]

[0078] The multi-task learning method is used to jointly optimize the recommendation loss and contrastive learning loss. The recommendation loss function adopts the Bayesian personalized ranking method. This method defaults to predicting that the project with which the user has interacted in the past has a higher score than the project with which the user has not interacted. The specific formula is as follows:

[0079]

[0080] Among them, ln(·) and sigmoid(·) represent the logarithmic function and activation function respectively, O represents the training data set, i + Indicates the items that user u has interacted with, i - Indicates items that user u has not interacted with;

[0081] The total loss function Loss is composed of the recommended loss function L bpr Compared with the dual-view learning loss function L cl Mixed with:

[0082]

[0083] Among them, γ represents the hyperparameter that controls the weight of contrastive learning loss, λ is the regularization coefficient, represents L2 regularization.

[0084] Furthermore, the implementation method of step 7 is as follows:

[0085] After obtaining the initial recommendation list, we first define the acceptance index based on the user's attention to the trending topics and the group's recognition of the project;

[0086] For user u, the set of tendency topics is Use formula (14) to calculate the attention of user u to each topic in T′ Get the user's attention distribution A(u):

[0087]

[0088] For candidate item i, use formula (19) to calculate the item i’s u Recognition of each theme Get the project recognition distribution P(i):

[0089]

[0090] User u's acceptance of item i, fav ui The definition is as follows:

[0091]

[0092] Then, according to the acceptance index fav ui Reorder the initial recommendation list of user u to obtain the final item recommendation list, which is and Know fav ui ∈(0,5), according to fav ui The value of further determines the recommendation grade of item i for user u. The specific formula is as follows:

[0093]

[0094] The multi-dimensional hierarchical feedback mechanism composed of acceptance indicators, project recognition and recommendation level can effectively enhance the interpretability of recommendation results and provide a reference basis for user decision-making.

[0095] An explainable recommendation system with collaborative enhancement of ratings and reviews, including the following modules:

[0096] The topic analysis module based on BERTopic: divide user comments into sentences, forming a sentence set S = {s1, s2, ..., s h}, use the BERT pre-training model to encode all sentences in S, and get the sentence representation set (S) = {e(s1), e(s2),…, e(s h )}, use UMAP dimensionality reduction technology to reduce the sentence representation to d dimensions, and the sentence representation set after dimensionality reduction is recorded as {u(s1),u(s2),…,u(s h )}, HDBSCAN technology is used to cluster sentence representations to generate multiple clusters C1, C2, ..., C k , each cluster Corresponding to a topic T g , the subject set is T={T1,T2,…,T k} means that the sentence representations in the topic are average pooled to obtain the topic representation e(T) = {e(T1), e(T2),…, e(T k )}, the subject representation matrix is ​​denoted as According to the correspondence between users, sentences and topics, derive the topic set involved by user u Similarly, according to the correspondence between items, sentences and topics, the topic set involved in item i is derived

[0097] Bimodal information collaborative processing module: BERT is used to calculate the sentiment score w of user u on item i ui , and standardize it to obtain Score the comment-level sentiment Compared with the original score r ui Weighted fusion to obtain the revised score This constitutes the revised user-item rating matrix User u's response to item i in topic T g All the sentences below constitute a comment After inputting it into the BERT model and normalizing it, we get the topic-level sentiment score of user u for item i. Pearson correlation coefficient was used to calculate T g The topic-level sentiment score of user u and revised ratings The correlation between Sort the topics users are involved in by relevance Sort from high to low and select the top N topics to form the tendency topic set of user u

[0098] Representation learning module from the perspective of user-item interaction: based on the revised rating matrix Construct a weighted bipartite graph from the perspective of user-item interaction, and the initial representation matrices of users and items are denoted as and The user representation matrix P (0) and the item representation matrix Q (0) Stacking into the initial representation matrix of user-item union And compare it with the revised scoring matrix Input into GCN, and obtain the updated representation matrix by averaging the representations obtained in each layer Then we can get the user representation from the perspective of user-item interaction and project representation

[0099] Representation learning module from the perspective of user-topic association: define the relationship between user u and topic T g ∈T′ u Attention Based on the user-topic attention matrix Construct a bipartite graph from the perspective of user-topic association and represent the user matrix P (0) and the topic representation matrix E T Stack to get the initial representation matrix of user topics And compare it with the user-topic attention matrix Input into GCN, and obtain the updated representation matrix by averaging the representations obtained in each layer Then we can get the user representation from the perspective of user-topic association

[0100] Representation learning module from the perspective of project-topic association: define project i in topic T g Recognition under Based on the project-theme recognition matrix Constructing a bipartite graph G from the perspective of project-topic association it (V it ,E it ), where V it =I∪T represents the node set, E it Represents the edge set, takes the recognition degree of the project under the topic as the feature of the edge, similar to obtaining the user representation from the user-topic association perspective, and transforms the project representation matrix Q (0) and the topic representation matrix E T Stacking is performed to obtain the initial representation matrix of the project theme By inputting it into GCN, we can obtain the project representation from the perspective of project-topic association.

[0101] Dual-perspective representation fusion module: Adaptively fuses the item representations from the user-item interaction perspective and the item-topic association perspective to obtain the item comprehensive representation e i , design dual-perspective contrastive learning to maximize the consistency of the representation of the same user in the user-item perspective and the user-topic perspective, as well as the difference between the representations of different users in the two perspectives, and obtain the user representation in the user-item perspective as the final comprehensive representation of the user. u ;

[0102] Model prediction and optimization module: Through comprehensive characterization of users u Comprehensive representation of the projecti Perform inner product operation to obtain the predicted score of user u for item i After obtaining the prediction score, a multi-task learning method is used to jointly optimize the recommendation loss and contrastive learning loss to generate an initial item recommendation list;

[0103] Re-ranking and feedback module: calculate user u’s response to T′ u The attention of each topic Get the user's attention distribution A(u), item i in T′ u Recognition of each theme Get the project recognition distribution P(i), and get user u’s acceptance of project i fav based on A(u) and P(i) ui , according to the acceptance index fav ui Reorder the initial recommendation list of user u, and further determine the recommendation grade of item i for user u.

[0104] Compared with the prior art, the present invention has the following beneficial effects:

[0105] 1. This invention effectively utilizes the collaborative enhancement of bimodal information to correct, filter, and improve the quality of raw data, and establishes a multi-dimensional hierarchical feedback mechanism that includes user acceptance, project recognition, and recommendation level, providing users with richer and more valuable recommendation results;

[0106] 2. The present invention can be widely used in various Internet service scenarios such as e-commerce, online education, online medical care, and financial management. It can enhance the accuracy and explainability of recommendation results through the collaborative enhancement of multimodal information, accurately match user needs, and enhance the core competitiveness and commercial value of the platform while improving the platform's retention rate, conversion rate, and activity indicators. BRIEF DESCRIPTION OF THE DRAWINGS

[0107] Figure 1 Schematic diagram of the process of the present invention;

[0108] Figure 2 This is a schematic diagram of the framework of the method recommended by the present invention;

[0109] Figure 3 Schematic diagram of the structure of the system recommended by the present invention. DETAILED DESCRIPTION

[0110] The present invention will be further described below with reference to the accompanying drawings and examples.

[0111] like Figure 1-3 As shown in FIG, an explainable recommendation method for collaborative enhancement of ratings and reviews includes the following steps:

[0112] Step 1: Extract topics from user comments based on BERTopic technology and establish associations between users and topics, and between items and topics. The implementation method is as follows:

[0113] The present invention first divides user comments into sentences, using S = {s1, s2, ..., s h} represents the set of all sentences. All sentences in S are encoded using the BERT pre-training model to obtain the sentence representation set e(S) = {e(s1), e(s2),…, e(s h )};

[0114] Then, the sentences are clustered according to their semantic similarity. To avoid the problem of dimensional disaster caused by representation in high-dimensional space during clustering, the UMAP dimensionality reduction technique is used to reduce the sentence representation to d dimensions before clustering. The sentence representation set after dimensionality reduction is recorded as {u(s1),u(s2),…,u(s h )},in Represents the d-dimensional representation of the j-th sentence. The density-based HDBSCAN algorithm is used to cluster the sentence representations after dimensionality reduction. The specific formula is as follows:

[0115] HDBSCAN({u(s1),u(s2),…,u(s h )})→{G1,C2,…,C k} (1)

[0116] C g →T g ,g=1,2,…,k (2)

[0117] Among them, k represents the number of generated clusters, and each cluster Corresponding to a topic T g , Represents the subject T g The d-dimensional representation of the next j-th sentence, |C g | indicates the subject T g The number of sentence tokens below;

[0118] Finally, for C g The sentence representations in are average pooled to obtain the topic T g Characterization of e(T g ), the specific formula is as follows:

[0119]

[0120] The representation of the subject e(T)={e(T1),e(T2),…,e(T k )} form a matrix

[0121]

[0122] According to the correspondence between users, sentences and topics, derive the topic set involved by user u Similarly, according to the correspondence between items, sentences and topics, the topic set involved in item i is derived

[0123] Step 2: Combine sentiment analysis with bimodal information for collaborative processing. The implementation method is as follows:

[0124] On the one hand, the comments are used to modify the ratings. First, the comment d of user u on item i is divided into ui Convert it into a format suitable for BERT processing. The specific formula is as follows:

[0125] e(d ui )=Tokenizer(d ui ) (5)

[0126] Next, BERT is used to calculate the sentiment score w of user u for item i ui , the specific formula is as follows:

[0127] w ui =BertModel(e(d ui )) (6)

[0128] to w ui Normalize it to be consistent with the range of scores in the dataset. The specific formula is as follows:

[0129]

[0130] Taking the Amazon dataset as an example, since the range of user ratings is {1, 2, 3, 4, 5}, the comment-level sentiment score of user u on item i is

[0131] Score the comment-level sentiment Compared with the original score r ui Weighted fusion to obtain the revised score The specific formula is as follows:

[0132]

[0133] Among them, α is the weight parameter that balances the user's original rating and the comment-level sentiment score;

[0134] The modified user-item rating matrix is ​​denoted as

[0135] On the other hand, the revised scores are used to filter topics. First, the relationship between users, items and topics is used to construct the user-topic-item ternary graph G UTI ={(u,T g ,i)|u∈U,T g ∈T u ∩T i ,i∈I}, where (u,T g ,i) indicates that user u and project i have a common topic T g ,user u has a good understanding of item i in topic T g All the sentences below constitute a comment After inputting it into the BERT model and normalizing it, we get the topic-level sentiment score of user u on item i.

[0136] Then, the Pearson correlation coefficient was used to calculate T g The topic-level sentiment score of user u and revised ratings The correlation between The specific formula is as follows:

[0137]

[0138] in, Represents the subject T g The set of items that user u has interacted with, Represents the subject T g The number of items that user u has interacted with, and Represents the topic T g The average sentiment score of user u and the average modified score of all items he interacted with;

[0139] Finally, the topics involved by users are sorted by relevance Sort from high to low and select the top N topics to form the tendency topic set of user u

[0140] Step 3: Construct a weighted bipartite graph based on the user's revised rating of the item, and use a graph convolutional neural network to learn node representations from the perspective of user-item interaction. The implementation method is as follows:

[0141] Based on the revised scoring matrix Constructing a weighted bipartite graph G from the perspective of user-item interaction ui (V ui ,E ui ), where V ui =U∪V represents the node set, E uiRepresents the edge set, and the modified rating is used as the edge feature; the initial representation matrix of users and items is recorded as and Where d is the embedding dimension; the user representation matrix P (0) and the item representation matrix Q (0) Stack and get the initial representation matrix of user-item union Input it into GCN to update the user and item representations. The specific formula is as follows:

[0142]

[0143] Among them, E (l) represents the node representation matrix of the lth layer, W1 and W2 are learnable parameter matrices, I represents the identity matrix, The Laplace matrix representing the user-item bipartite graph is as follows:

[0144]

[0145] in, is the modified rating matrix, D is the degree matrix of the node;

[0146] Perform average pooling on the representation matrix of each layer to obtain the updated representation matrix

[0147]

[0148] Among them, L is the number of convolution layers;

[0149] This results in a user representation from the perspective of user-item interaction. and project representation

[0150] Step 4: Based on the definition of user attention and project recognition, construct user-topic and project-topic bipartite graphs and learn node representations from two perspectives. The implementation method is as follows:

[0151] The more sentences involving a certain topic in a user's historical comments, the higher the user's attention to the topic. The user's attention to the topic reflects the user's own preferences to a certain extent. g ∈T′ u Attention The definition is as follows:

[0152]

[0153] Among them, S′ u Represents the set of sentences involving the tendency topic of user u;

[0154] Based on the user-topic attention matrix Constructing a bipartite graph G from the perspective of user-topic association ut (V ut ,E ut ), where V ut =U∪T represents the node set, E ut Represents the edge set, takes the user's attention to the topic as the edge feature; represents the user representation matrix P (0) and the topic representation matrix E T Stack to get the initial representation matrix of user topics It is also input into GCN. The specific formula is as follows:

[0155]

[0156] in, represents the node representation matrix of the lth layer, W3 and W4 are learnable parameter matrices, The Laplace matrix representing the user-topic bipartite graph is as follows:

[0157]

[0158] Perform average pooling on the representation matrix of each layer to obtain the updated representation matrix The specific formula is as follows:

[0159]

[0160] This results in the user representation from the perspective of user-topic association.

[0161] For a project, the degree of its recognition can be measured by analyzing the sentiment scores of the user groups that have interacted with it under various topics. Taking project i as an example, its sentiment scores under topic T g The following sentence set is Calculated using the BERT model Each sentence in The sentiment score is obtained after normalization in will be collected The average sentiment score of the sentences in T is defined as the average sentiment score of item i in topic T g Recognition under The specific formula is as follows:

[0162]

[0163] Based on the project-theme recognition matrix Constructing a bipartite graph G from the perspective of project-topic associationit (V it ,E it ), where V it =I∪T represents the node set, E it Represents the edge set, and takes the recognition degree of the project under the topic as the feature of the edge; similar to obtaining the user representation from the user-topic association perspective, the project representation matrix Q (0) and the topic representation matrix E T Stacking is performed to obtain the initial representation matrix of the project theme By inputting it into GCN, we can obtain the project representation from the perspective of project-topic association. Laplace matrix corresponding to the project-topic bipartite graph The specific formula is as follows:

[0164]

[0165] Step 5: Use adaptive fusion and dual-view contrastive learning to model the comprehensive representation of the project and user, respectively. The implementation method is as follows:

[0166] The project representation from the perspective of user-item interaction and the perspective of item-topic association are adaptively integrated to obtain the comprehensive project representation e i , the specific formula is as follows:

[0167]

[0168] Among them, β is a trainable parameter;

[0169] To further enrich the semantic information in user representation, a dual-perspective contrastive learning task is designed. This task optimizes user representation by maximizing the consistency of the representation of the same user under the user-item perspective and the user-topic perspective, as well as the differences in the representation of different users under the two perspectives. The specific formula is as follows:

[0170]

[0171] Among them, sim(·) represents the cosine similarity function, τ represents the temperature coefficient, represents the user representation from the perspective of user-item interaction, represents the user representation from the perspective of user-topic association, Representation of other users from the perspective of user-topic association;

[0172] The optimized user representation from the user-item perspective is used as the final comprehensive representation of the user, denoted as e u .

[0173] Step 6: Perform an inner product operation on the comprehensive representation of users and items to obtain the prediction score. Use a multi-task learning method to jointly optimize the recommendation loss and contrastive learning loss. The implementation method is as follows:

[0174] By comprehensively characterizing the user u Comprehensive representation of the project i Perform inner product operation to obtain the predicted score of user u for item i. The specific formula is as follows:

[0175]

[0176] The multi-task learning method is used to jointly optimize the recommendation loss and contrastive learning loss. The recommendation loss function adopts the Bayesian personalized ranking method. This method defaults to predicting that the project with which the user has interacted in the past has a higher score than the project with which the user has not interacted. The specific formula is as follows:

[0177]

[0178] Among them, ln(·) and sigmoid(·) represent the logarithmic function and activation function respectively, O represents the training data set, i + Indicates the items that user u has interacted with, i - Indicates items that user u has not interacted with;

[0179] The total loss function Loss is composed of the recommended loss function L bpr Compared with the dual-view learning loss function L cl Mixed with:

[0180]

[0181] Among them, γ represents the hyperparameter that controls the weight of contrastive learning loss, λ is the regularization coefficient, represents L2 regularization.

[0182] Step 7: Define acceptance indicators based on users' attention to trending topics and the group's recognition of the project, re-rank the projects, and form multi-dimensional graded feedback information. The implementation method is as follows:

[0183] After obtaining the initial recommendation list, we first define the acceptance index based on the user's attention to the trending topics and the group's recognition of the project;

[0184] For user u, the set of tendency topics is Use formula (14) to calculate user u to T′ u The attention of each topic Get the user's attention distribution A(u):

[0185]

[0186] For candidate item i, use formula (19) to calculate the item i’s u Recognition of each theme Get the project recognition distribution P(i):

[0187]

[0188] User u's acceptance of item i, fav ui The definition is as follows:

[0189]

[0190] Then, according to the acceptance index fav ui Reorder the initial recommendation list of user u to obtain the final item recommendation list, which is and Know fav ui ∈(0,5), according to fav ui The value of further determines the recommendation grade of item i for user u. The specific formula is as follows:

[0191]

[0192] The multi-dimensional hierarchical feedback mechanism composed of acceptance indicators, project recognition and recommendation level can effectively enhance the interpretability of recommendation results and provide a reference basis for user decision-making.

[0193] An explainable recommendation system with collaborative enhancement of ratings and reviews, including the following modules:

[0194] Divide user comments into sentences to form a sentence set S = {s1, s2, ..., s h}, use the BERT pre-training model to encode all sentences in S, and get the sentence representation set (S) = {e(s1), e(s2),…, e(s h )}, use UMAP dimensionality reduction technology to reduce the sentence representation to d dimensions, and the sentence representation set after dimensionality reduction is recorded as {u(s1),u(s2),…,u(s h )}, HDBSCAN technology is used to cluster sentence representations to generate multiple clusters C1, C2, ..., C k , each cluster Corresponding to a topic T g , the subject set is T={T1,T2,…,T k} means that the sentence representations in the topic are average pooled to obtain the topic representation e(T) = {e(T1), e(T2),…, e(T k )}, the subject representation matrix is ​​denoted as According to the correspondence between users, sentences and topics, derive the topic set involved by user u Similarly, according to the correspondence between items, sentences and topics, the topic set involved in item i is derived

[0195] Bimodal information collaborative processing module: BERT is used to calculate the sentiment score w of user u on item i ui , and standardize it to obtain Score the comment-level sentiment Compared with the original score r ui Weighted fusion to obtain the revised score This constitutes the revised user-item rating matrix User u's response to item i in topic T g All the sentences below constitute a comment After inputting it into the BERT model and normalizing it, we get the topic-level sentiment score of user u on item i. Pearson correlation coefficient was used to calculate T g The topic-level sentiment score of user u and revised ratings The correlation between Sort the topics users are involved in by relevance Sort from high to low and select the top N topics to form the tendency topic set of user u

[0196] Representation learning module from the perspective of user-item interaction: based on the revised rating matrix Construct a weighted bipartite graph from the perspective of user-item interaction, and the initial representation matrices of users and items are denoted as and The user representation matrix P (0) and the item representation matrix Q (0) Stacking into the initial representation matrix of user-item union And compare it with the revised scoring matrix Input into GCN, and obtain the updated representation matrix by averaging the representations obtained in each layer Then we can get the user representation from the perspective of user-item interaction and project representation

[0197] Representation learning module from the perspective of user-topic association: define the relationship between user u and topic T g ∈T′ u Attention Based on the user-topic attention matrix Construct a bipartite graph from the perspective of user-topic association and represent the user matrix P (0) and the topic representation matrix E T Stack to get the initial representation matrix of user topics And compare it with the user-topic attention matrix Input into GCN, and obtain the updated representation matrix by averaging the representations obtained in each layer Then we can get the user representation from the perspective of user-topic association

[0198] Representation learning module from the perspective of project-topic association: define project i in topic T g Recognition under Based on the project-theme recognition matrix Constructing a bipartite graph G from the perspective of project-topic association it (V it ,E it ), where V it =I∪T represents the node set, E it Represents the edge set, takes the recognition degree of the project under the topic as the feature of the edge, similar to obtaining the user representation from the user-topic association perspective, and transforms the project representation matrix Q (0) and the topic representation matrix E T Stacking is performed to obtain the initial representation matrix of the project theme By inputting it into GCN, we can obtain the project representation from the perspective of project-topic association.

[0199] Dual-perspective representation fusion module: Adaptively fuses the item representations from the user-item interaction perspective and the item-topic association perspective to obtain the item comprehensive representation e i , design dual-perspective contrastive learning to maximize the consistency of the representation of the same user in the user-item perspective and the user-topic perspective, as well as the difference between the representations of different users in the two perspectives, and obtain the user representation in the user-item perspective as the final comprehensive representation of the user. u ;

[0200] Model prediction and optimization module: Through comprehensive characterization of users u Comprehensive representation of the project i Perform inner product operation to obtain the predicted score of user u for item i After obtaining the prediction score, a multi-task learning method is used to jointly optimize the recommendation loss and contrastive learning loss to generate an initial item recommendation list;

[0201] Re-ranking and feedback module: Calculate the attention of user u to each topic in T′ Get the user's attention distribution A(u), item i in T′ u Recognition of each theme Get the project recognition distribution P(i), and get user u’s acceptance of project i fav based on A(u) and P(i) ui , according to the acceptance index fav ui Reorder the initial recommendation list of user u, and further determine the recommendation grade of item i for user u.

Claims

1. An explainable recommendation method with collaborative enhancement of ratings and reviews, characterized by: The following steps are involved: Step 1: Extract topics from user comments based on BERTopic technology and establish associations between users and topics, and between items and topics; Step 2: Combine the sentiment analysis model to collaboratively process the bimodal information, implement score correction and topic screening, and obtain the user's preferred topics; Step 3: Construct a weighted bipartite graph based on the user's revised rating of the item, and learn the node representation from the perspective of user-item interaction through a graph convolutional neural network; Step 4: Based on the definition of user attention and project recognition, construct user-topic and project-topic bipartite graphs and learn node representations from two perspectives; Step 5: Use adaptive fusion and dual-view contrastive learning to model the comprehensive representation of items and users respectively; Step 6: Perform an inner product operation on the comprehensive representation of users and items to obtain the prediction score, and use a multi-task learning method to jointly optimize the recommendation loss and contrastive learning loss; Step 7: Define acceptance indicators based on users' attention to trending topics and the group's recognition of the project, re-rank the projects, and form multi-dimensional graded feedback information.

2. The explainable recommendation method for collaborative enhancement of ratings and reviews according to claim 1, characterized in that: The implementation method of step 1 is as follows: First, user comments are divided into sentences, using S = {s1,s2,…,s h } represents the set of all sentences. All sentences in S are encoded using the BERT pre-training model to obtain the sentence representation set e(S) = {e(s1), e(s2),…, e(s h )}; Then, the sentences are clustered according to their semantic similarity. To avoid the problem of dimensional disaster caused by representation in high-dimensional space during clustering, the UMAP dimensionality reduction technique is used to reduce the sentence representation to d dimensions before clustering. The sentence representation set after dimensionality reduction is recorded as {u(s1),u(s2),…,u(s h )},in Represents the d-dimensional representation of the j-th sentence. The density-based HDBSCAN algorithm is used to cluster the sentence representations after dimensionality reduction. The specific formula is as follows: HDBSCAN({u(s1),u(s2),…,u(s h )})→{C1,C2,…,C k } (1) C g →T g ,g=1,2,…,k (2) Among them, k represents the number of generated clusters, and each cluster Corresponding to a topic T g , Represents the subject T g The d-dimensional representation of the next j-th sentence, |C g | indicates the subject T g The number of sentence tokens below; Finally, for C g The sentence representations in are average pooled to obtain the topic T g Characterization of e(T g ), the specific formula is as follows: The representation of the subject e(T)={e(T1),e(T2),…,e(T k )} form a matrix According to the correspondence between users, sentences and topics, derive the topic set involved by user u Similarly, according to the correspondence between items, sentences and topics, the topic set involved in item i is derived 3. The explainable recommendation method for collaborative enhancement of ratings and reviews according to claim 1, characterized in that: The implementation method of step 2 is as follows: On the one hand, the comments are used to modify the ratings. First, the comment d of user u on item i is divided into ui Convert it into a format suitable for BERT processing. The specific formula is as follows: e(d ui )=Tokenizer(d ui ) (5) Next, BERT is used to calculate the sentiment score w of user u for item i ui , the specific formula is as follows: w ui =BertModel(e(d ui )) (6) to w ui Normalize it to be consistent with the range of scores in the dataset. The specific formula is as follows: Taking the Amazon dataset as an example, since the range of user ratings is {1, 2, 3, 4, 5}, the comment-level sentiment score of user u on item i is Score the comment-level sentiment Compared with the original score r ui Weighted fusion to obtain the revised score The specific formula is as follows: Among them, α is the weight parameter that balances the user's original rating and the comment-level sentiment score; The modified user-item rating matrix is ​​denoted as On the other hand, the revised scores are used to filter topics. First, the relationship between users, items and topics is used to construct the user-topic-item ternary graph G UTI ={(u,T g ,i)|u∈U,T g ∈T u ∩T i ,i∈I}, where (u,T g ,i) indicates that user u and project i have a common topic T g ,user u has a good understanding of item i in topic T g All the sentences below constitute a comment After inputting it into the BERT model and normalizing it, we get the topic-level sentiment score of user u on item i. Then, the Pearson correlation coefficient was used to calculate T g The topic-level sentiment score of user u and revised ratings The correlation between The specific formula is as follows: in, Represents the subject T g The set of items that user u has interacted with, Represents the subject T g The number of items that user u has interacted with, and Represents the topic T g The average sentiment score of user u and the average modified score of all items he interacted with; Finally, the topics involved by users are sorted by relevance Sort from high to low and select the top N topics to form the tendency topic set of user u 4. The explainable recommendation method for collaborative enhancement of ratings and reviews according to claim 1, characterized in that: The implementation method of step 3 is as follows: Based on the revised scoring matrix Constructing a weighted bipartite graph G from the perspective of user-item interaction ui (V ui ,E ui ), where V ui =U∪V represents the node set, E ui Represents the edge set, and the modified rating is used as the edge feature; the initial representation matrix of users and items is recorded as and Where d is the embedding dimension; the user representation matrix P (0) and the item representation matrix Q (0) Stack and get the initial representation matrix of user-item union Input it into GCN to update the user and item representations. The specific formula is as follows: Among them, E (l) represents the node representation matrix of the lth layer, W1 and W2 are learnable parameter matrices, I represents the identity matrix, The Laplace matrix representing the user-item bipartite graph is as follows: in, is the modified rating matrix, D is the degree matrix of the node; Perform average pooling on the representation matrix of each layer to obtain the updated representation matrix Among them, L is the number of convolution layers; This results in a user representation from the perspective of user-item interaction. and project representation 5. The explainable recommendation method with collaborative enhancement of ratings and reviews according to claim 1, characterized in that: The implementation method in step 4 is as follows: The more sentences involving a certain topic in a user's historical comments, the higher the user's attention to the topic. The user's attention to the topic reflects the user's own preferences to a certain extent. g ∈T′ u Attention The definition is as follows: Among them, S′ u Represents the set of sentences involving the tendency topic of user u; Based on the user-topic attention matrix Constructing a bipartite graph G from the perspective of user-topic association ut (E ut ,E ut ), where V ut =U∪T represents the node set, E ut Represents the edge set, takes the user's attention to the topic as the edge feature; represents the user representation matrix P (0) and the topic representation matrix E T Stack to get the initial representation matrix of user topics It is also input into GCN. The specific formula is as follows: in, represents the node representation matrix of the lth layer, W3 and W4 are learnable parameter matrices, The Laplace matrix representing the user-topic bipartite graph is as follows: Perform average pooling on the representation matrix of each layer to obtain the updated representation matrix The specific formula is as follows: This results in the user representation from the perspective of user-topic association. For a project, the degree of its recognition can be measured by analyzing the sentiment scores of the user groups that have interacted with it under various topics. Taking project i as an example, its sentiment scores under topic T g The following sentence set is Calculated using the BERT model Each sentence in The sentiment score is obtained after normalization in will be collected The average sentiment score of the sentences in T is defined as the average sentiment score of item i in topic T g Recognition under The specific formula is as follows: Based on the project-theme recognition matrix Constructing a bipartite graph G from the perspective of project-topic association it (V it ,E it ), where V it =I∪T represents the node set, E it Represents the edge set, and takes the recognition degree of the project under the topic as the feature of the edge; similar to obtaining the user representation from the user-topic association perspective, the project representation matrix Q (0) and the topic representation matrix E T Stacking is performed to obtain the initial representation matrix of the project theme By inputting it into GCN, we can obtain the project representation from the perspective of project-topic association. Laplace matrix corresponding to the project-topic bipartite graph The specific formula is as follows:

6. The explainable recommendation method for collaborative enhancement of ratings and reviews according to claim 1, characterized in that: The implementation method in step 5 is as follows: The project representation from the perspective of user-item interaction and the perspective of item-topic association are adaptively integrated to obtain the comprehensive project representation e i , the specific formula is as follows: Among them, β is a trainable parameter; To further enrich the semantic information in user representation, a dual-perspective contrastive learning task is designed. This task optimizes user representation by maximizing the consistency of the representation of the same user under the user-item perspective and the user-topic perspective, as well as the differences in the representation of different users under the two perspectives. The specific formula is as follows: Among them, sim(·) represents the cosine similarity function, τ represents the temperature coefficient, represents the user representation from the perspective of user-item interaction, represents the user representation from the perspective of user-topic association, Representation of other users from the perspective of user-topic association; The optimized user representation from the user-item perspective is used as the final comprehensive representation of the user, denoted as e u .

7. The explainable recommendation method for collaborative enhancement of ratings and reviews according to claim 1, characterized in that: The implementation method in step 6 is as follows: By comprehensively characterizing the user u Comprehensive representation of the project i Perform inner product operation to obtain the predicted score of user u for item i. The specific formula is as follows: The multi-task learning method is used to jointly optimize the recommendation loss and contrastive learning loss. The recommendation loss function adopts the Bayesian personalized ranking method. This method defaults to predicting that the project with which the user has interacted in the past has a higher score than the project with which the user has not interacted. The specific formula is as follows: Among them, ln(·) and sigmoid(·) represent the logarithmic function and activation function respectively, O represents the training data set, i + Indicates the items that user u has interacted with, i - Indicates items that user u has not interacted with; The total loss function Loss is composed of the recommended loss function L bpr Compared with the dual-view learning loss function L cl Mixed with: Among them, γ represents the hyperparameter that controls the weight of contrastive learning loss, λ is the regularization coefficient, represents L2 regularization.

8. The explainable recommendation method for collaborative enhancement of ratings and reviews according to claim 1, characterized in that: The implementation method in step 7 is as follows: After obtaining the initial recommendation list, we first define the acceptance index based on the user's attention to the trending topics and the group's recognition of the project; For user u, the set of tendency topics is Use formula (14) to calculate user u to T′ u The attention of each topic Get the user's attention distribution A(u): For candidate item i, use formula (19) to calculate the item i’s u Recognition of each theme Get the project recognition distribution P(i): User u's acceptance of item i, fav ui The definition is as follows: Then, according to the acceptance index fav ui Reorder the initial recommendation list of user u to obtain the final item recommendation list, which is and Know fav ui ∈(0,5), according to fav ui The value of further determines the recommendation grade of item i for user u. The specific formula is as follows: The multi-dimensional hierarchical feedback mechanism composed of acceptance indicators, project recognition and recommendation level can effectively enhance the interpretability of recommendation results and provide a reference basis for user decision-making.

9. An explainable recommendation system with collaborative enhancement of ratings and reviews, characterized by: Includes the following modules: Divide user comments into sentences to form a sentence set S = {s1, s2, ..., s h }, use the BERT pre-training model to encode all sentences in S, and get the sentence representation set (S) = {e(s1), e(s2),…, e(s h )}, use UMAP dimensionality reduction technology to reduce the sentence representation to d dimensions, and the sentence representation set after dimensionality reduction is recorded as {u(s1),u(s2),…,u(s h )}, HDBSCAN technology is used to cluster sentence representations to generate multiple clusters C1, C2, ..., C k , each cluster Corresponding to a topic T g , the subject set is T={T1,T2,…,T k } means that the sentence representations in the topic are average pooled to obtain the topic representation e(T) = {e(T1), e(T2),…, e(T k )}, the subject representation matrix is ​​denoted as According to the correspondence between users, sentences and topics, derive the topic set involved by user u Similarly, according to the correspondence between items, sentences and topics, the topic set involved in item i is derived Bimodal information collaborative processing module: BERT is used to calculate the sentiment score w of user u on item i ui , and standardize it to obtain Score the comment-level sentiment Compared with the original score r ui Weighted fusion to obtain the revised score This constitutes the revised user-item rating matrix User u's response to item i in topic T g All the sentences below constitute a comment After inputting it into the BERT model and normalizing it, we get the topic-level sentiment score of user u on item i. Pearson correlation coefficient was used to calculate T g The topic-level sentiment score of user u and revised ratings The correlation between Sort the topics users are involved in by relevance Sort from high to low and select the top N topics to form the tendency topic set of user u Representation learning module from the perspective of user-item interaction: based on the revised rating matrix Construct a weighted bipartite graph from the perspective of user-item interaction, and the initial representation matrices of users and items are denoted as and The user representation matrix P (0) and the item representation matrix Q (0) Stacking into the initial representation matrix of user-item union And compare it with the revised scoring matrix Input into GCN, and obtain the updated representation matrix by averaging the representations obtained in each layer Then we can get the user representation from the perspective of user-item interaction and project representation Representation learning module from the perspective of user-topic association: define the relationship between user u and topic T g ∈T′ u Attention Based on the user-topic attention matrix Construct a bipartite graph from the perspective of user-topic association and represent the user matrix P (0) and the topic representation matrix E T Stack to get the initial representation matrix of user topics And compare it with the user-topic attention matrix Input into GCN, and obtain the updated representation matrix by averaging the representations obtained in each layer Then we can get the user representation from the perspective of user-topic association Representation learning module from the perspective of project-topic association: define project i in topic T g Recognition under Based on the project-theme recognition matrix Constructing a bipartite graph G from the perspective of project-topic association it (V it ,E it ), where V it =I∪T represents the node set, E it Represents the edge set, takes the recognition degree of the project under the topic as the feature of the edge, similar to obtaining the user representation from the user-topic association perspective, and transforms the project representation matrix Q (0) and the topic representation matrix E T Stacking is performed to obtain the initial representation matrix of the project theme By inputting it into GCN, we can obtain the project representation from the perspective of project-topic association. Dual-perspective representation fusion module: Adaptively fuses the item representations from the user-item interaction perspective and the item-topic association perspective to obtain the item comprehensive representation e i , design dual-perspective contrastive learning to maximize the consistency of the representation of the same user in the user-item perspective and the user-topic perspective, as well as the difference between the representations of different users in the two perspectives, and obtain the user representation in the user-item perspective as the final comprehensive representation of the user. u ; Model prediction and optimization module: Through comprehensive characterization of users u Comprehensive representation of the project i Perform inner product operation to obtain the predicted score of user u for item i After obtaining the prediction score, a multi-task learning method is used to jointly optimize the recommendation loss and contrastive learning loss to generate an initial item recommendation list; Re-ranking and feedback module: calculate user u’s response to T′ u The attention of each topic Get the user's attention distribution A(u), item i in T′ u Recognition of each theme Get the project recognition distribution P(i), and get user u’s acceptance of project i fav based on A(u) and P(i) ui , according to the acceptance index fav ui Reorder the initial recommendation list of user u, and further determine the recommendation grade of item i for user u.