Multi-type context-aware dialogue recommendation method based on hybrid expert model

By constructing a multi-type context-aware dialogue recommendation method of hybrid expert model, the heterogeneity and semantic gap between structured and unstructured data are solved, and efficient fusion of multi-type context information is achieved, and recommendation performance and diversity are improved.

CN120277182APending Publication Date: 2025-07-08UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510312106.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

现有对话推荐方法难以有效整合和融合结构化数据与非结构化数据的异质性与语义鸿沟问题,以及对上下文信息利用过于单一的问题,导致推荐性能不足。

Method used

A multi-type context-aware dialogue recommendation method based on a hybrid expert model is constructed. By constructing a dialogue expert model, a knowledge graph expert model and a comment expert model, it is specially modeled for different types of context information, and a coordination system is introduced for fusion, dynamically allocating the weights of different context information to generate the final recommendation result.

Benefits of technology

It improves the performance of dialogue recommendation methods, enhances the adaptability and robustness to diversified data, improves the relevance and diversity of recommendation results, and has good scalability and interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277182A_ABST
    Figure CN120277182A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-type context-aware dialogue recommendation method based on a hybrid expert model. The aim of the invention is fulfilled by a hybrid expert framework comprising a plurality of expert module coordination systems. The knowledge graph expert model is responsible for extracting and coding structured entity relationship information, the dialogue expert model is used for deeply mining key contents and user intentions in dialogues, and the comment experts are used for analyzing user emotion and preference information in commodity comments. In this way, the system can make full use of semantic information of the structured data and the unstructured data, and the problem of semantic gaps caused by data heterogeneity in the prior art is solved. On the basis, a coordination system is introduced to serve as a core module for coordination and integration. The coordination system can dynamically allocate the output of the expert module according to the weights and contributions of different context information to generate a final recommendation result, so that efficient fusion of multiple types of context information is realized, the problems of semantic alignment and inconsistency between different types of data are solved, and the performance of the dialogue recommendation method is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of dialogue recommendation. More specifically, it relates to a multi-type context-aware dialogue recommendation method based on a mixture of experts model. Background Art

[0002] Dialogue recommendation methods have become a prominent research topic in recent years. Different from traditional recommendation methods that mainly rely on historical data and user behavior, dialogue recommendation methods utilize the powerful function of natural language dialogue to provide personalized and context-aware recommendations. Through dialogue interaction, users can experience more attractive and effective recommendations.

[0003] With the increasing research interest in dialogue recommendation methods, many methods have been proposed to promote the development of this field. Since the dialogues in dialogue recommendation methods are usually short and contain limited context information, in many existing studies, integrating external data sources to enrich context information has gradually become the norm. Among these external data sources, the most commonly used are knowledge graphs and item reviews. However, how to combine different types of context information remains a challenge.

[0004] One of the main challenges is the heterogeneity of different external data. For example, knowledge graphs are structured data, while dialogue content and item reviews are unstructured data. There is a semantic gap between multi-type context information and heterogeneous data. In addition, different external data usually belong to different semantic spaces, so it is difficult to align them. A possible solution is to use contrastive learning to fuse different external data. However, contrastive learning requires the same entries in all external data to calculate the contrastive loss, which is not always achievable. Summary of the Invention

[0005] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a multi-type context-aware dialogue recommendation method based on a mixture of experts model. By constructing and mixing three expert models, it effectively fuses multi-type context information to solve the heterogeneity and semantic gap problems caused by the difficulty of efficiently integrating and fusing structured data and unstructured data in existing dialogue recommendation methods, as well as the problem of overly single use of context information in existing dialogue recommendation methods, and addresses the particularity and challenges of dialogue recommendation data, thereby improving the performance of dialogue recommendation methods.

[0006] To achieve the above object of the invention, the multi-type context-aware dialogue recommendation method based on a mixture of experts model of the present invention is characterized by including the following steps:

[0007] (1) Construct three expert models

[0008] For the dialogue data, knowledge graph data, and project review data in the dataset, corresponding expert models are constructed to reason about the three types of data, specifically including:

[0009] 1.1) Construct a dialogue expert model using dialogue data

[0010] 1.1.1) For the j-th dialogue in the dialogue data, extract the project and entities related to the project in sequence and use them as elements constituting a sequence to form a user sequence. Based on the method of the cloze task, randomly mask some content in the user sequence and input it into the Transformer network to predict the masked project, thereby modeling the user sequence: The k-th element of the randomly masked user sequence is embedded as the position embedding s k and the sequence embedding p k to form a combination h k = s k + p k , and all element embeddings h k form a dialogue representation matrix H 1 . Through the multi-head self-attention mechanism and the feed-forward neural network activated by the Gaussian error linear unit, perform N-layer updates. The dialogue representation matrix H n+1 after the n-th layer update is:

[0011] H n+1 = MultiHead(PFFN(LayerNorm(H n + Dropout(sublayer(H n )))))

[0012] where n represents the number of layers, n = 1, 2,..., N, H n is the dialogue representation matrix input to the n-th layer, PFFN represents the feed-forward neural network, MultiHead represents the multi-head self-attention mechanism, Dropout represents random inactivation, sublayer represents the feed-forward neural network or the multi-head self-attention mechanism, and LayerNorm represents layer normalization;

[0013] The element embedding corresponding to the i-th project in the dialogue representation matrix H N+1 after the last layer update of the Transformer network is used as its hidden embedding, denoted as

[0014] 1.1.2) In the prediction stage, calculate the probability of the i-th project through a linear layer and activation by the Gaussian error linear unit

[0015]

[0016] Among them, W is a trainable transformation matrix, b is a trainable bias vector, and i is the index of the item;

[0017] 1.1.3), Use the cross-entropy loss function L rec , to train the dialogue expert model, where the cross-entropy loss function L rec is:

[0018]

[0019] Among them, j represents the dialogue index, and y ij is the true label of the i-th item in the j-th dialogue, and I is the set of all items;

[0020] 1.1.4), After training, input the dialogue corresponding to the i-th item into the trained dialogue expert model to obtain the probability and the hidden embedding as the input of the subsequent coordination system;

[0021] 1.2), Use the knowledge graph data to construct a graph expert model

[0022] 1.2.1), Map the items in all dialogues in the dialogue data and the entities related to the items to the knowledge graph G. Among them, the knowledge graph G contains the entity set E and the relationship set composed of the items and the entities related to the items

[0023] Adopt a relational graph convolutional network to encode the entities e, e ∈ E in the knowledge graph G to obtain the entity representation Then update, the entity representation of entity e after the l-th layer update is:

[0024]

[0025] Among them, l represents the number of layers, l = 1, 2,..., L, is the entity representation input to the l-th layer of entity e, is the entity representation input to the l-th layer of the neighbor entity e' of entity e, is the neighbor set of entity e under the relationship r, and W l is the learnable parameter matrix used to transform the entity representation at the l-th layer, is the learnable parameter matrix used to transform the entity representation under the relationship r at the l-th layer, and Z e,t is the normalization factor of entity e under the relationship r, σ(·) is the activation function, and the entity representation of the first layer is obtained by inputting entity e into the pre-trained network;

[0026] 1.2.2) Entity representation of entity e after the update of the L-th layer Denoted as entity representation n e After obtaining the entity representations of all entities e, e ∈ E, extract the items and the entities related to the items from the user conversation to form the user entity set E u And use the self-attention mechanism to calculate the user interest representation n u :

[0027]

[0028] Among them, α e Represents the important weight of entity e, e ∈ E u In the user interest, W a Is the attention weight matrix, e′ is an entity in the user entity set E u n e′ Is the entity representation of entity e′, exp is the exponential function;

[0029] 1.2.3) Finally, based on the user interest representation n u Calculate the probability of the i-th item

[0030]

[0031] Among them, n i Is the entity representation of the i-th item;

[0032] Take the entity representation n i As the hidden embedding of the i-th item Together with the probability As the input of the subsequent coordination system;

[0033] 1.3) Use the item review data to construct a review expert model

[0034] 1.3.1) For the i-th item, corresponding to m reviews, use the standard Transformer model to encode each review to obtain the corresponding sentence representation, and place the sentence representations of all reviews in order to form the review representation matrix D i ;

[0035] 1.3.2) Calculate the overall review representation of the i-th item

[0036]

[0037] Among them, Transformer represents the Transformer network, and SelfAttention represents the self-attention mechanism;

[0038] 1.3.3), Calculate the user comment interest representation

[0039]

[0040] Among them: I u is the set of items that the user has interacted with, and β i is the attention weight;

[0041] 1.3.4), Based on the user comment interest representation Calculate the probability of the i-th item

[0042]

[0043] 1.3.5), Use the cross-entropy loss function to optimize the comment expert model, and then calculate the probability P R (i) of the i-th item and the overall comment representation of the i-th item And use the overall comment representation as the hidden embedding of the i-th item Together with the probability P R (i) as the input to the subsequent coordination system;

[0044] (2), Construct a mixture-of-experts recommendation module

[0045] 2.1), The mixture-of-experts recommendation module fuses the results of three expert models by introducing a coordination system:

[0046] First, generate the following representation through a concatenation operation:

[0047]

[0048] Then, generate a normalized importance score for each expert model

[0049]

[0050] Finally, calculate the recommendation probability of the i-th item:

[0051]

[0052] 2.2), Use the cross-entropy loss function to fine-tune the mixture recommendation module;

[0053] (3), Item recommendation

[0054] Send the user conversation into the trained dialogue expert model in step 1.1) to obtain the probability and the hidden embedding Send the user conversation into the trained dialogue expert model in step 1.2), and calculate the hidden embedding according to steps 1.2.2) and 1.2.3). And probability Calculate the probability P of the user according to the comment expert model in step 1.3). R (i) And hidden embedding Then send it into the hybrid expert recommendation module fine-tuned in step (2) to calculate the recommendation probability of the i-th item. If the recommendation probability of the i-th item is obtained, recommend this item to the user.

[0055] (4) Response generation

[0056] In the response generation module, first construct the initial representation B 0 , as the input of the 0th layer:

[0057] B 0 = X + W rec

[0058] Among them, X is the dialogue history matrix extracted from the Transformer encoder, providing dialogue context information, and W rec is the recommended item bias;

[0059] In the k-th, k = 1, 2,..., K layer decoder of the Transformer decoder, fuse the entity representations of the experts and the output representations of the Transformer encoder in the following way:

[0060]

[0061] Among them, K is the number of layers of the Transformer decoder, MultiHead[·, ·, ·] represents the multi-head attention mechanism, FFN(·) represents the feed-forward neural network, F C , F G , F R represent the matrices composed of all entity representations from the dialogue expert model, the graph expert model, and the comment expert model respectively;

[0062] The response generation module generates a response based on the output representation of the K-th layer decoder of the Transformer decoder Generate a response.

[0063] The object of the present invention is achieved in this way.

[0064] The present invention innovatively designs a multi-type context-aware dialogue recommendation method based on a mixture of experts model, which includes a mixture of experts framework with multiple expert modules to coordinate the system to achieve the object of the present invention. Each expert model specializes in modeling different types of context information. Among them, the knowledge graph expert model is responsible for extracting and encoding structured entity relationship information, the dialogue expert model deeply mines the key content and user intentions in the dialogue, and the review expert analyzes the user sentiment and preference information in the product reviews. In this way, the system can make full use of the semantic information of structured and unstructured data, and solves the semantic gap problem caused by data heterogeneity in the prior art. On this basis, the present invention introduces a coordination system as the core module for coordination and integration. The coordination system can dynamically allocate the outputs of expert modules according to the weights and contributions of different context information to generate the final recommendation result, so as to achieve the efficient fusion of multi-type context information and solve the problems of semantic alignment and inconsistency between different types of data. In addition, the present invention relaxes the strict requirements of contrastive learning for data consistency by combining multi-type data modeling, and enhances the adaptability and robustness of the model to diverse data. The present invention further improves the comprehensive utilization of context information through deep modeling and semantic fusion, significantly enhancing the relevance and diversity of the recommendation results, and overcoming the problem of single utilization of context information in the existing dialogue recommendation methods. At the same time, by means of modular modeling, the present invention has good scalability, can flexibly add new expert models to integrate more context information sources, and the independence of each module makes the source of the recommendation result more interpretable, providing a clear direction for system optimization. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 is a flowchart of a specific embodiment of the multi-type context-aware dialogue recommendation method based on the mixture of experts model of the present invention;

[0066] Figure 2 is a schematic diagram of the architecture of a specific embodiment of the multi-type context-aware dialogue recommendation method based on the mixture of experts model of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0068] The following describes the specific embodiments of the present invention with reference to the accompanying drawings, so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed description of known functions and designs may obscure the main content of the present invention, these descriptions will be omitted here.

[0069] Figure 1 is a flowchart of a specific embodiment of the multi-type context-aware dialogue recommendation method based on the mixture of experts model of the present invention.

[0070] In this embodiment, as Figure 1 shown, the multi-type context-aware dialogue recommendation method based on the mixture of experts model of the present invention includes the following steps:

[0071] Step S1: Construct three expert models

[0072] First, for the input data set, the present invention extracts three different types of context information therefrom, namely dialogue data, knowledge graph data, and review data. For the dialogue data, knowledge graph data, and project review data in the data set, corresponding expert models are respectively constructed to reason about the three types of data, specifically including:

[0073] Step S1.1: Construct a dialogue expert model using dialogue data

[0074] Step S1.1.1: Obtain a dialogue representation matrix and the hidden embedding of the i-th project

[0075] For the j-th dialogue in the dialogue data, the project and the entities related to the project are sequentially extracted and used as elements constituting a sequence to form a user sequence. Based on the method of the Cloze Task, by randomly masking part of the content in the user sequence and inputting it into the Transformer network to predict the masked project, the user sequence is thus modeled: the k-th element embedding of the randomly masked user sequence is the combination h k of the position embedding s k and the sequence embedding p k = s k + p k , and all the element embeddings h k form a dialogue representation matrix H 1 , which is updated through the multi-head self-attention mechanism (MultiHead) and the feed-forward neural network (PFFN) activated by the Gaussian error linear unit (GELU) for N layers. The dialogue representation matrix H n+1 after the n-th layer update is:

[0076] H n+1 = MultiHead(PFFN(LayerNorm(H n + Dropout(sublayer(H n )))))

[0077] where n represents the number of layers, n = 1, 2,..., N, and H nThe dialogue representation matrix input for the nth layer, PFFN represents the feed-forward neural network, MultiHead represents the multi-head self-attention mechanism, Dropout represents random inactivation, sublayer represents the feed-forward neural network or the multi-head self-attention mechanism, and LayerNorm represents layer normalization.

[0078] To solve the problem of difficult training of deep networks, the module introduces a residual connection, layer normalization, and random inactivation (Dropout) mechanism in each layer, that is, H n = H n + Dropout(sublayer(H n ))

[0079] The element embedding corresponding to the i-th item in the updated dialogue representation matrix H of the last layer of the Transformer network N+1 is used as its hidden embedding, denoted as

[0080] Step S1.1.2: Calculate the probability of the i-th item

[0081] In the prediction stage, calculate the probability of the i-th item through a linear layer and a Gaussian error linear unit activation

[0082] where W is a trainable transformation matrix, b is a trainable bias vector, and i is the index of the item.

[0083] Step S1.1.3: Train the dialogue expert model

[0084] Use the cross-entropy loss function L rec , and train the dialogue expert model, where the cross-entropy loss function L rec is:

[0085]

[0086] where j represents the dialogue index, y ij is the true label of the i-th item in the j-th dialogue, and I is the set of all items.

[0087] Step S1.1.4: Calculate the probability and the hidden embedding as the input to the subsequent coordination system

[0088] After training, input the dialogue corresponding to the i-th item into the trained dialogue expert model to obtain the probability and the hidden embedding as the input to the subsequent coordination system.

[0089] Step S1.2: Construct a graph expert model using knowledge graph data

[0090] The graph expert model models user preferences using an external knowledge graph, specifically:

[0091] Step S1.2.1: Calculate entity representations

[0092] Map all items and entities related to the items in the dialogue data as nodes to the knowledge graph G, where the knowledge graph G contains an entity set E composed of items and entities related to the items and a relationship set

[0093] In this embodiment, the knowledge graph G is constructed based on the DBpedia dataset. The relationships in the knowledge graph G are represented by triples <e1, r, e2>, where e1, e2 ∈ E are entities, is the relationship.

[0094] Use a relational graph convolutional network to encode the entities e, e ∈ E in the knowledge graph G to obtain entity representations Then update. The entity representation of entity e after the l-th layer update is:

[0095]

[0096] where l represents the layer number, l = 1, 2,..., L, is the entity representation of entity e input at the l-th layer, is the entity representation of the neighbor entity e' of entity e input at the l-th layer, is the neighbor set of entity e under the relationship r, W l is the learnable parameter matrix for transforming the entity representation at the l-th layer, is the learnable parameter matrix for transforming the entity representation at the l-th layer under the relationship r, Z e,r is the normalization factor of entity e under the relationship r, σ(·) is the activation function, and the entity representation of the first layer is obtained by inputting entity e into the pre-trained network.

[0097] Step S1.2.2: Calculate the user interest representation n u

[0098] The entity representation of entity e after the L-th layer update is denoted as the entity representation n e , and after obtaining the entity representations of all entities e, e ∈ E, extract the items and entities related to the items from the user dialogue to form the user entity set Eu and calculate the user interest representation n using the self-attention mechanism u :

[0099]

[0100] where α e represents the important weight of entity e, e ∈ E u in the user interest, W a is the attention weight matrix, e' is an entity in the user entity set E u , n e′ is the entity representation of entity e', and exp is the exponential function

[0101] Step S1.2.3: Calculate the probability of the i-th item and use it as the input to the subsequent coordination system together with the hidden embedding

[0102] Finally, calculate the probability of the i-th item based on the user interest representation n u

[0103]

[0104] where n i is the entity representation of the i-th item

[0105] Use the entity representation n i as the hidden embedding of the i-th item together with the probability as the input to the subsequent coordination system

[0106] Step S1.3: Construct a comment expert model using item review data

[0107] The comment expert model improves the performance of the dialogue recommendation method by encoding the review data related to the item. The review data consists of sentences written by users about the item. First, use the Transformer network to encode the review sentences to obtain sentence representations; then, aggregate the representations of all sentences through a sentence-level self-attention layer to generate an overall representation based on the reviews, specifically:

[0108] Step S1.3.1: Construct a comment representation matrix D i :

[0109] For the i-th item, corresponding to m reviews, use the standard Transformer model to encode each review to obtain the corresponding sentence representation, and place the sentence representations of all reviews in order to form the comment representation matrix D i .​​

[0110] Step S1.3.2: Calculate the overall comment representation of the \(i\)-th item

[0111]

[0112] Among them, Transformer represents the Transformer network, and SelfAttention represents the self-attention mechanism.

[0113] Overall comment representation Extract the deep semantic features of the comment through Transformer, and perform weighted aggregation through the self-attention mechanism to capture the key information in the comment.

[0114] Step S1.3.3: Calculate the user comment interest representation

[0115] User comment interest representation It is obtained through the aggregation calculation of the comment representations of the user's historical interaction items. The specific calculation is as follows:

[0116]

[0117] Among them: \(I\) u is the set of items that the user has interacted with, and \(\beta\) i is the attention weight, which measures the importance of the comment of item \(i\) in the user interest modeling. The contribution degree of the comments of different items to the user interest is learned through the self-attention mechanism to obtain the user comment interest representation

[0118] Step S1.3.4: Based on the user comment interest representation Calculate the probability of the \(i\)-th item

[0119] Step S1.3.5: Optimize the comment expert model and calculate the probability \(P\) R (i) And the hidden embedding As the input of the subsequent coordination system

[0120] Use the cross-entropy loss function to optimize the comment expert model, and then calculate the probability \(P\) of the \(i\)-th item R (i) And the overall comment representation of the \(i\)-th item And use the overall comment representation As the hidden embedding of the \(i\)-th item Together with the probability \(P\) R (i) As the input of the subsequent coordination system.

[0121] Step S2: Construct a Hybrid Expert Recommendation Module

[0122] Step S2.1: The hybrid expert recommendation module fuses the results of three expert models by introducing a coordination system to generate more accurate and relevant recommendations. The coordination system coordinates the embeddings and prediction results of the three expert models (i.e., the dialogue expert model, the graph expert model, and the review expert model) and processes them.

[0123] First, collect the hidden embeddings of the i-th item from all expert models and the prediction results, i.e., probabilities and generate the following representation through a concatenation operation:

[0124]

[0125]

[0126] Then, generate a normalized importance score for each expert model

[0127]

[0128] Finally, calculate the recommendation probability of the i-th item:

[0129]

[0130] Step S2.2: Fine-tune the hybrid recommendation module using the cross-entropy loss function

[0131] To optimize the entire recommendation module, the hybrid recommendation module is fine-tuned using the cross-entropy loss function to optimize the representation and improve performance.

[0132] Step S3: Item Recommendation

[0133] Send the user dialogue into the dialogue expert model trained in Step S1.1 to obtain the probability and the hidden embedding Send the user dialogue into the dialogue expert model trained in Step S1.2 and calculate the hidden embedding according to Steps S1.2.2 and S1.2.3 and the probability Calculate the probability P R (i) and the hidden embedding of this user according to the review expert model in Step S1.3 Then send it into the hybrid expert recommendation module fine-tuned in Step S2 to calculate the recommendation probability of the i-th item. If the recommendation probability of the i-th item is [probability value], recommend this item to the user.

[0134] Step S4: Response Generation

[0135] The response generation module improves the effect of response generation through the recommended item bias provided by the recommendation method, making the generated responses more consistent and diverse. Based on the pre-trained representation, in the standard Transformer decoder architecture, the present invention integrates multiple cross-attention layers to effectively fuse the entity representations from the dialogue expert model, the graph expert model, and the review expert model.

[0136] In the response generation module, first construct the initial representation B 0 , as the input of the 0th layer:

[0137] B 0 = X + W rec

[0138] where X is the dialogue history matrix extracted from the Transformer encoder, providing dialogue context information, and W rec is the recommended item bias, enhancing the bias of the model when recommending relevant items.

[0139] In the kth, k = 1, 2,..., K layer decoder of the Transformer decoder, fuse the entity representation of the expert and the output representation of the Transformer encoder in the following way:

[0140]

[0141]

[0142] where K is the number of layers of the Transformer decoder, MultiHead[·, ·, ·] represents the multi-head attention mechanism, FFN(·) represents the feed-forward neural network, and F C , F G , F R represent the matrices respectively composed of all entity representations from the dialogue expert model, the graph expert model, and the review expert model;

[0143] The above cross representation is processed by a feed-forward neural network (FFN) and an activation function to generate the output of the decoder, which is used to generate the user response. The response generation module generates a response based on the output representation of the Kth layer decoder of the Transformer decoder to generate a response. The generated natural language response can combine the context information of the dialogue recommendation system and provide high-quality recommendation suggestions related to the target item, thereby improving the user experience.

[0144] Experimental verification

[0145] To verify the effectiveness of the present invention, the effects of the present invention and nine existing conversational recommendation algorithms were first compared. The experiment was mainly verified from two aspects: recommendation effect and response generation, which specifically included the following parts: experimental dataset, comparison method, experimental results and analysis.

[0146] Experimental Dataset

[0147] The experiment was verified based on two publicly available conversational recommendation datasets: the ReDial dataset and the INSPIRED dataset. The ReDial dataset contains 10,006 conversations and 182,150 conversation utterances, covering recommendation requests for user target items and context information related to the recommended items. The INSPIRED dataset is relatively small, containing 1,001 conversations and 35,811 conversation utterances, and also provides more detailed user intention annotations. To enhance the diversity of the experiment, we also extracted relevant movie reviews from IMDb as an additional source of text context data. These datasets were all divided into training set, validation set and test set according to the ratio of 8:1:1.

[0148] Comparison Method

[0149] To comprehensively evaluate the performance of the present invention, nine classic and state-of-the-art methods were selected as comparison baselines in the experiment. These comparison methods cover the mainstream technical directions in the current conversational recommendation system, so as to fully reflect the performance advantages of the method of the present invention.

[0150] Experimental Metrics

[0151] The following metrics were used in the experiment to evaluate the recommendation performance and generation effect: Recommendation performance metrics: Recall@k (k = 1, 10, 50), which is used to measure the accuracy of the target item in the recommendation list. Generation effect metrics: Distinct-n (n = 2, 3, 4), which is used to evaluate the diversity and innovation of the generated response. Manual evaluation metrics: The generated response was scored from three dimensions: fluency, informativeness and relevance, and the score range was from 1 to 5.

[0152] Experimental Results and Analysis

[0153] Model Recall@1 Recall@10 Recall@50 Popularity 0.012 0.061 0.179 TextCNN 0.013 0.068 0.191 BERT 0.014 0.117 0.191 ReDial 0.023 0.129 0.287 KBRD 0.031 0.150 0.336 KGSF 0.039 0.183 0.378 RevCore 0.046 0.220 0.396 VRICR 0.054 0.244 0.406 <![CDATA[C 2 -CRS]]> 0.053 0.233 0.407 The present invention 0.057* 0.250* 0.473*

[0154] Table 1

[0155]

[0156]

[0157] Table 2

[0158] Tables 1 and 2 respectively show the recommendation performance of the present invention and nine baseline methods on the ReDial dataset and the INSPIRED dataset. As can be seen from Tables 1 and 2, the present invention is significantly superior to the baseline methods in the three metrics of Recall@1, Recall@10, and Recall@50. On the ReDial dataset, the present invention improves by 16.2% in the Recall@50 metric compared to the optimal baseline C2-CRS; on the INSPIRED dataset, the improvement in this metric reaches 24.6%. These results indicate that the multi-type context-aware framework proposed by the present invention can effectively integrate multi-modal information, thereby significantly improving the recommendation performance.

[0159] Model Distinct-2 Distinct-3 Distinct-4 Transformer 0.148 0.151 0.137 ReDial 0.225 0.236 0.228 KBRD 0.263 0.368 0.423 KGSF 0.330 0.417 0.521 RevCore 0.424 0.558 0.612 VRICR 0.382 0.453 0.496 <![CDATA[C 2 -CRS]]> 0.631 0.932 0.909 The present invention 0.680* 0.976* 0.981*

[0160] Table 3

[0161] Model Distinct-2 Distinct-3 Distinct-4 Transformer 1.020 2.248 3.582 ReDial 1.347 1.521 3.445 KBRD 1.369 2.259 3.592 KGSF 1.608 2.719 4.929 RevCore 2.419 3.820 4.648 VRICR 1.937 3.248 4.965 <![CDATA[C 2 -CRS]]> 2.456 4.432 5.092 The present invention 2.584* 4.579* 5.251*

[0162] Table 4

[0163] Tables 3 and 4 show the performance of each method in the metrics of Distinct-2, Distinct-3, and Distinct-4. The experimental results show that the dialogue responses generated by the present invention are significantly superior to all baseline methods in terms of diversity, especially with the largest improvement in the Distinct-4 metric, indicating that the method of the present invention can generate more diverse dialogue response content.

[0164] Model Fluency Informativeness Transformer 0.82 0.91 ReDial 1.25 1.09 KBRD 1.31 1.22 KGSF 1.53 1.32 RevCore 1.55 1.38 VRICR 1.52 1.34 <![CDATA[C 2 -CRS]]> 1.58 1.51 The present invention 1.66* 1.59*

[0165] Table 5

[0166] In addition, the results of the human evaluation on the ReDial dataset are shown in Table 5. The average scores of the present invention in the two dimensions of fluency and informativeness are higher than those of the baseline methods, further verifying the quality of the responses generated by it.

[0167] Model Recall@1 Recall@10 Recall@50 Dialogue only 0.022 0.113 0.207 Knowledge graph only 0.051 0.213 0.373 Comment only 0.021 0.093 0.369 Remove dialogue 0.053 0.216 0.471 Remove knowledge graph 0.025 0.121 0.382 Remove comment 0.052 0.228 0.472 The present invention 0.057* 0.250* 0.473*

[0168] Table 6

[0169] As can be seen from Table 6, through ablation experiments, we verified the independent contributions of each module in the MCCRS framework to the recommendation performance. The results show that removing any one module will lead to a performance decline, indicating that each module plays an important role in the recommendation task. In particular, the structured information of the knowledge graph expert module plays a key role in modeling user preferences and ranking target items, significantly improving the recommendation performance. At the same time, the dialogue expert module can capture users' short-term interests, while the review expert module enhances the diversity and relevance of recommendations by supplementing sentiment and evaluation information. The experimental results further prove that the three expert modules have strong complementarity in the framework, and their collaborative work effectively improves the overall performance of the system.

[0170] Although the above description of the illustrative specific embodiments of the present invention is provided for the convenience of those skilled in the art to understand the present invention, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

Claims

1. A multi-type context-aware dialogue recommendation method based on a mixture of experts model, characterized by comprising the following steps: (1) Construct three expert models For the dialogue data, knowledge graph data, and project review data in the dataset, construct corresponding expert models to reason about the three types of data, specifically including: 1.1) Use the dialogue data to construct a dialogue expert model 1.1.1) For the j-th conversation in the conversation data, the items and the entities related to the items are sequentially extracted and used as the elements constituting a sequence to form a user sequence. Based on the method of the cloze task, part of the content in the user sequence is randomly masked and input into the Transformer network to predict the masked items, thereby modeling the user sequence: the k-th element of the randomly masked user sequence is embedded as the positional embedding s k and the sequence embedding p k to form the combination h k = s k + p k , and all the element embeddings h k form the conversation representation matrix H 1 , which is updated for N layers through the multi-head self-attention mechanism and the feed-forward neural network based on the Gaussian error linear unit activation. The conversation representation matrix H n+1 after the n-th layer update is as follows: H n+1 = MultiHead(PFFN(LayerNorm(H n + Dropout(sublayer(H n )))) Among them, n represents the number of layers, n = 1, 2, …, N, H n is the dialogue representation matrix input for the n-th layer, PFFN represents a feed-forward neural network, MultiHead represents a multi-head self-attention mechanism, Dropout represents random inactivation, sublayer represents a feed-forward neural network or a multi-head self-attention mechanism, and LayerNorm represents layer normalization; The dialogue representation matrix H after the update of the last layer of the Transformer network N+1 The element embedding corresponding to the i-th item in is used as its hidden embedding, denoted as 1.1.2), in the prediction stage, calculate the probability of the i-th item through linear layer and Gaussian error linear unit activation Among them, W is a trainable transformation matrix, b is a trainable bias vector, and i is the index of the item; 1.1.3), use the cross-entropy loss function L rec , to train the dialogue expert model, where Cross-entropy loss function L rec is as follows: where j represents the dialogue index, and y ij is the true label of the i-th item in the j-th dialogue, and I is the set of all items; 1.1.4), after training, input the conversation corresponding to the i-th item into the trained conversation expert model to obtain the probability and the hidden embedding as the input for the subsequent coordination system; 1.2), constructing a graph expert model using knowledge graph data 1.2.1), map the items in all conversations and the entities related to the items in the conversation data to the knowledge graph G, where, The knowledge graph G includes an entity set E and a relationship set R composed of items and entities related to the items; Encode the entities \(e, e\in E\) in the knowledge graph \(G\) using a relational graph convolutional network to obtain entity representations Then perform an update. The entity representation of entity \(e\) after the update in the \(l\)-th layer is as follows: where \(l\) represents the number of layers, \(l = 1, 2, \ldots, L\), is the entity representation of entity \(e\) input at the \(l\)-th layer, is the entity representation of the neighbor entity \(e'\) of entity \(e\) input at the \(l\)-th layer, is the set of neighbors of entity \(e\) under relation \(r\), \(W\) l is the learnable parameter matrix for transforming the entity representation at the \(l\)-th layer, is the learnable parameter matrix for transforming the entity representation at the \(l\)-th layer under relation \(r\), \(Z\) e,r is the normalization factor of entity \(e\) under relation \(r\), \(\sigma(\cdot)\) is the activation function, and the entity representation of the first layer is obtained by inputting entity \(e\) into the pre-trained network; 1.2.2) Entity representation of entity e after the update of the L-th layer Denoted as entity representation n e , after obtaining the entity representations of all entities e, e ∈ E, extract items and entities related to the items from the user conversation to form the user entity set E u , and use the self-attention mechanism to calculate the user interest representation n u : Among them, α e represents the entity e, e ∈ E u the important weight in the user interest, W a is the attention weight matrix, e′ is the set of user entities E u one entity in e′ is the entity representation of the entity e′, exp is the exponential function; 1.2.3), Finally, based on the user interest representation n u Calculate the probability of the i-th item where n i is the representation of the i-th project entity; Represent the entity as n i as the hidden embedding of the i-th item along with the probability as the input to the subsequent coordination system; 1.3), constructing a review expert model using item review data 1.3.1), For the i-th project, corresponding to m comments, use the standard Transformer model to encode each comment to obtain the corresponding sentence representation. Place the sentence representations of all comments by example to form the comment representation matrix D i ; 1.3.2), Calculate the overall comment representation of the i-th project where, Transformer represents the Transformer network, and SelfAttention represents the self-attention mechanism; 1.3.3), Calculate the user comment interest representation Where: I u is the set of items the user has interacted with, and β i is the attention weight; 1.3.4), User Comment Interest Representation Calculate the probability of the i-th item 1.3.5), optimize the comment expert model using the cross - entropy loss function, and then calculate the probability P R (i) of the i - th item and the overall comment representation of the i - th item And use the overall comment representation as the hidden embedding of the i - th item along with the probability P R (i) as the input to the subsequent coordination system; (2), constructing a hybrid expert recommendation module 2.1), the hybrid expert recommendation module fuses the results of the three expert models by introducing a coordination system: First, generate the following representations through a concatenation operation: Then, generate a normalized importance score for each expert model Finally, calculate the recommendation probability of the i-th item: 2.2), using the cross-entropy loss function to fine-tune the hybrid recommendation module; (3), item recommendation Send the user conversation into the trained dialogue expert model in step 1.1) to obtain probabilities and hidden embeddings Send the user conversation into the trained dialogue expert model in step 1.2), and calculate the hidden embeddings according to steps 1.2.2) and 1.2.3) and probabilities Calculate the probability P of this user according to the comment expert model in step 1.3) R (i) and hidden embeddings Then send it into the fine-tuned mixture-of-experts recommendation module in step (2) to calculate the recommendation probability of the i-th item. If the recommendation probability of the i-th item is [probability value], recommend this item to the user; (4), response generation In the response generation module, first construct the initial representation B 0 , as the input of the 0th layer: B 0 = X + W rec Among them, X is the dialogue history matrix extracted from the Transformer encoder, providing dialogue context information, and W rec is the recommendation item bias; In the k-th, k = 1, 2,..., K layer decoder of the Transformer decoder, fuse the entity representation of the expert and the output representation of the Transformer encoder in the following way: where K is the number of layers of the Transformer decoder, MultiHead[·,·,·] represents the multi-head attention mechanism, FFN(·) represents the feed-forward neural network, and F C ,F G ,F R represent matrices respectively composed of all entity representations from the dialogue expert model, the graph expert model, and the review expert model; The response generation module generates a response based on the output representation of the K-th layer of the Transformer decoder to generate a response.

Citation Information

Cited By

  • Anti-data migration recommendation method based on hybrid expert and dynamic graph neural network

    CN120910364A

  • Intelligent customer service automatic answering method and system

    CN121009169A

  • Multimodal knowledge fusion personalized recommendation method and device based on context and memory information, medium, program product and terminal

    CN121901429A