A knowledge tree enhanced conversation recommendation method and a model training method
Patent Information
- Application Number
- CN202610962333.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]本申请提出了一种知识树增强的会话推荐方案,旨在解决基于预训练语言模型与外部知识图谱的现有技术中存在的知识噪声过滤不足、图结构信息利用低效、用户深层偏好建模困难、以及模型训练成本与多任务性能难以兼顾这四项核心技术缺陷
[0016]综上所述,本申请各实施例提供的知识树增强的会话推荐模型的训练方法和知识树增强的会话推荐方法,通过依据对话表示对知识图谱进行分层语义筛选以构建对话知识树,从源头实现知识的动态、精细化过滤,从而攻克了现有技术知识噪声过滤不足的难题;通过将对话知识树序列化为自然语言文本,并借助对比学习损失对齐其表示与实体表示,将图结构转化为预训练语言模型可直接推理的符号序列,弥合了表征鸿沟,从而解决了图结构信息利用低效的问题;通过基于对话与实体表示进行交互计算与上下文聚合以生成用户协同偏好表示,并计算协同偏好损失,该特征显式地建模了对话上下文中多个提及实体之间的复杂关联,并通过辅助的监督信号确保所提取的表示能够有效反映用户的协同兴趣,从而克服了现有技术用户深层偏好建模困难的缺陷;通过采用参数冻结的生成式预训练语言模型作为核心引擎,并构建融合协同偏好、对比学习、推荐与生成损失的总损失,以联合优化提示、偏好提取及知识编码参数,在保留大模型通用能力的同时高效协调多任务,从而打破了现有技术模型训练成本与多任务性能难以兼顾的困境。此外,上述方案通过区分用户反馈的方向性信息并对偏好变化进行有效追踪,进一步提升了在冷启动或反馈不确定场景下对用户真实意图的建模精度。上述技术特征,动态知识过滤、结构序列化与对齐、监督式偏好提取、以及冻结模型下的多任务提示学习,并非简单叠加,而是相互耦合、深度协同,共同实现了会话推荐系统在推荐精准度、回复相关性、知识利用率及用户意图理解深度上的协同提升。
Smart Images

Figure CN122840170A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of conversation recommendation systems, specifically relating to a knowledge tree-enhanced conversation recommendation method and model training method. Background Technology
[0002] Conversational recommendation systems aim to understand user intent through multi-turn dialogues and simultaneously provide personalized item recommendations and fluent natural language responses. Currently, solutions based on pre-trained language models and external knowledge graphs have become the mainstream technical paradigm in this field. This approach typically utilizes graph neural networks to encode entities and relationships in the knowledge graph and injects knowledge into the pre-trained language model through mechanisms such as attention to jointly optimize the recommendation and response generation tasks.
[0003] However, this paradigm suffers from a series of specific technical shortcomings in practice: First, the knowledge retrieval and fusion process introduces a large amount of semantic noise. Existing methods typically retrieve knowledge graph subgraphs based on a fixed number of hops, resulting in a large number of entities and relationships unrelated to the current dialogue context being input into the model. This noisy knowledge interferes with the model's focus on key evidence. Second, there is a representational gap between the structured information of the knowledge graph and the sequence processing capabilities of pre-trained language models. Existing methods encode the knowledge graph as dense vectors, losing explicit relationships and hierarchical structures between entities. The serialized and symbolic logical descriptions that pre-trained language models excel at handling are not effectively utilized, thus limiting the model's ability to perform deep reasoning. Third, there is insufficient mining of deep and collaborative user preference signals. Existing methods typically simply splice together multi-turn dialogue histories, failing to explicitly model the collaborative interests implied between multiple entities mentioned by the user in different turns. This results in a superficial understanding of the user's dynamic intent, especially in cold-start or ambiguous user preference scenarios. The system lacks a basis for judging the direction and confidence of user preferences and cannot effectively track cross-turns when preferences change, further exacerbating the bias in user intent modeling. Finally, model training faces a trade-off between efficiency and multi-task performance. Fine-tuning all parameters of the entire large pre-trained language model is computationally expensive and may impair its general language capabilities; while simple prompting learning methods struggle to reconcile the conflicting optimization objectives of recommendation and generation tasks during the shared representation learning process.
[0004] Therefore, existing technical solutions have long faced problems such as insufficient knowledge noise filtering, inefficient utilization of graph structure information, difficulty in modeling deep user preferences, and the inability to balance model training costs with multi-task performance. The technical solution of this application aims to specifically address the above-mentioned deficiencies. Summary of the Invention
[0005] This application proposes a knowledge tree-enhanced conversation recommendation scheme, aiming to address four core technical shortcomings in existing technologies based on pre-trained language models and external knowledge graphs: insufficient knowledge noise filtering, inefficient utilization of graph structure information, difficulty in modeling deep user preferences, and difficulty in balancing model training costs and multi-task performance.
[0006] A first aspect of this application provides a training method for a knowledge tree-enhanced conversation recommendation model, the conversation recommendation model including a generative pre-trained language model, comprising: The multi-turn dialogue history and external knowledge graph are encoded to obtain dialogue representation and global entity representation. The mentioned entities are identified from the dialogue history, and the representation of the mentioned entities is obtained from the global entity representation. Feedback state identification is performed on each round of user utterance in the multi-turn dialogue history to determine the feedback type of the user utterance in that round and calculate the state confidence. Based on the feedback type and the state confidence, interactive calculation and context aggregation are performed on the dialogue representation and the representation of the mentioned entity to generate a user collaborative preference representation and calculate the collaborative preference loss. The knowledge graph is subjected to hierarchical semantic filtering to construct a dialogue knowledge tree. The dialogue knowledge tree is serialized into text, the text is encoded to obtain its representation, and a contrastive learning loss is calculated based on the representation of the mentioned entity and the representation of the text. Performing a recommendation task includes: concatenating a learnable soft cue, the text, the multi-turn dialogue history, and the user collaborative preference representation into a combined cue, inputting the generative pre-trained language model with frozen parameters, and calculating the recommendation loss and response generation loss, wherein the learnable soft cue is used to guide the generative pre-trained language model to adapt to the conversation recommendation task; The total loss is constructed based on the collaborative preference loss, the contrastive learning loss, the recommendation loss, and the response generation loss, and the parameters of the conversation recommendation model are updated accordingly.
[0007] In some embodiments of this application, the step of performing interactive calculations and context aggregation on the dialogue representation and the representation of the mentioned entity based on the feedback type and the feedback state confidence to generate a user collaborative preference representation includes: User collaborative preference representation is generated based on the following formula: , in, This is a vector concatenation operation, where Asum(·) is the self-attention aggregation function. It is the first The dialogue indicated that, The entities mentioned in this round indicate that It is the confidence level of the feedback state; The feedback direction weights are determined based on the feedback type, wherein positive preferences correspond to positive weights, negative preferences correspond to suppression weights, weak preferences and ambiguous preferences correspond to lower weights, and preference shifts correspond to increasing the weight of the current round and decreasing the weight of previous rounds.
[0008] In some embodiments of this application, the feedback types include at least positive preference, negative preference, weak preference, fuzzy preference, and preference shift; No. Confidence of user feedback status Calculate according to the following formula: , in, For the Sigmoid function, For the first The semantic vector of user discourse, For historical dialogue semantic vectors, and These are learnable parameters.
[0009] In some embodiments of this application, it also includes: Conflict detection is performed on preferences in different rounds based on the feedback type and the feedback state confidence level, wherein historical round preferences are calculated according to the following formula. With current preferences Level of conflict: , in, and These are historical round preferences and current preferences, respectively. The preference polarity is determined by the corresponding feedback type. For indicator functions; When the degree of conflict exceeds a threshold, the weights of the feedback directions from previous rounds are reduced: , in, The feedback direction weights for historical rounds, For the interval between rounds, The current preference confidence level, This is the attenuation parameter.
[0010] In some embodiments of this application, constructing the dialogue knowledge tree includes: Based on the feedback type, the mentioned entities are categorized into a set of entities with positive preferences. Negative preference entity set and uncertain entity set , In the knowledge graph, candidate nodes are selected, and for each candidate node, a screening score is calculated according to the following formula. : , in, For dialogue pooling vectors, Representing neighboring entities, , This is a node noise penalty term. , , λ4 are weighting coefficients; Based on the screening score, neighboring nodes are selected from the candidate nodes to construct a dialogue knowledge tree.
[0011] In some embodiments of this application, serializing the dialogue knowledge tree into text includes: The tree structure is transformed into text by performing a depth-first traversal of the dialogue knowledge tree and inserting markers to distinguish entity names from relation names during the traversal.
[0012] In some embodiments of this application, the contrastive learning loss is calculated using the following formula: , in, Indicates the contrast learning loss. This represents the mentioned entity. The representation of the serialized text. For temperature coefficient, This represents the representation of other samples in the same training batch.
[0013] In some embodiments of this application, the total construction loss includes: The collaborative preference loss The contrastive learning loss The recommended loss and the loss generated by the response We perform a weighted summation to obtain the total loss L: , in, , , These are the weighting coefficients for the corresponding loss terms.
[0014] In some embodiments of this application, a two-stage training strategy is employed to update the parameters of the conversation recommendation model other than the generative pre-trained language model. The two-stage training strategy includes: In the first stage, the parameters of the generative pre-trained language model are kept frozen, and the parameters of the conversation recommendation model other than the generative pre-trained language model are updated based on the collaborative preference loss, the contrastive learning loss and the recommendation loss. In the second phase, all parameters of the session recommendation model are jointly fine-tuned.
[0015] A second aspect of this application provides a knowledge tree-enhanced conversation recommendation method, including: Obtain the multi-turn dialogue history between the current user and the recommendation system, as well as external knowledge graphs related to the recommended items; The dialogue history and the knowledge graph are input into a conversation recommendation model, wherein the conversation recommendation model is a model obtained by the training method according to the first aspect of the embodiments of this application; Obtain the recommended item ranking results and natural language responses for the current dialogue output by the conversation recommendation model.
[0016] In summary, the training method and the method for knowledge tree-enhanced conversation recommendation provided in the embodiments of this application overcome the problem of insufficient knowledge noise filtering in existing technologies by constructing a conversation knowledge tree through hierarchical semantic filtering of the knowledge graph based on the dialogue representation, thereby achieving dynamic and refined filtering of knowledge from the source; by serializing the conversation knowledge tree into natural language text and aligning its representation with entity representation using contrastive learning loss, the graph structure is transformed into a symbol sequence that can be directly reasoned about by the pre-trained language model, bridging the representation gap and solving the problem of inefficient use of graph structure information; and by performing interactive computation and context aggregation based on dialogue and entity representation... This approach generates user collaborative preference representations and calculates collaborative preference losses. This explicitly models the complex relationships between multiple mentioned entities in the dialogue context and uses auxiliary supervision signals to ensure that the extracted representations effectively reflect the user's collaborative interests, thus overcoming the difficulty of deep user preference modeling in existing technologies. By employing a parameter-frozen generative pre-trained language model as the core engine and constructing a total loss that integrates collaborative preferences, contrastive learning, recommendation, and generation losses, it jointly optimizes the parameters for prompts, preference extraction, and knowledge encoding. This efficiently coordinates multiple tasks while retaining the general capabilities of a large model, breaking the dilemma of balancing training cost and multi-task performance in existing technologies. Furthermore, by distinguishing the directional information of user feedback and effectively tracking preference changes, the above scheme further improves the modeling accuracy of the user's true intent in cold-start or uncertain feedback scenarios. These technical features—dynamic knowledge filtering, structured serialization and alignment, supervised preference extraction, and multi-task prompt learning under a frozen model—are not simply superimposed but mutually coupled and deeply collaborative, jointly achieving a synergistic improvement in recommendation accuracy, response relevance, knowledge utilization, and depth of user intent understanding in the conversational recommendation system. Attached Figure Description
[0017] The features and advantages of this application will become clearer with reference to the accompanying drawings, which are illustrative and should not be construed as limiting the application in any way. In the drawings: Figure 1 This is a schematic diagram of the overall architecture of the knowledge tree-enhanced conversation recommendation model in this application; Figure 2 This is a flowchart illustrating a training method for a knowledge tree-enhanced conversation recommendation model according to some embodiments of this application; Figure 3 This is a flowchart illustrating a knowledge tree-enhanced conversation recommendation method according to some embodiments of this application; Figure 4 This is an ablation experiment result of a recommendation task on the INSPIRED dataset, based on an embodiment of this application. Detailed Implementation
[0018] In the following detailed description, numerous specific details of this application are illustrated by example to provide a thorough understanding of the relevant disclosure. However, it will be apparent to those skilled in the art that this application can be practiced without these details. It should be understood that the terms “system,” “apparatus,” “unit,” and / or “module” used in this application are one way of distinguishing different parts, elements, sections, or components at different levels in a sequential arrangement. However, these terms may be replaced with other expressions if other expressions can achieve the same purpose.
[0019] It should be understood that when a device, unit, or module is referred to as being "on," "connected to," or "coupled to" another device, unit, or module, it may be directly connected to or coupled to or communicate with other devices, units, or modules, or there may be intermediate devices, units, or modules present, unless the context explicitly indicates otherwise. For example, the term "and / or" as used herein includes any one and all combinations of one or more of the relevant listed items.
[0020] The terminology used in this application is for the purpose of describing specific embodiments only and is not intended to limit the scope of this application. As shown in the specification and claims of this application, unless the context clearly indicates otherwise, words such as "a," "an," "an," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate that explicitly identified features, integrals, steps, operations, elements, and / or components are included, and such expressions do not constitute an exclusive list, and other features, integrals, steps, operations, elements, and / or components may also be included.
[0021] Referring to the following description and accompanying drawings, these and other features and characteristics, operating methods, functions of related structural elements, combinations of parts, and economics of manufacture of this application can be better understood, wherein the description and drawings form part of the specification. However, it is clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. It is understood that the drawings are not drawn to scale.
[0022] Various structural diagrams are used in this application to illustrate various variations of the embodiments according to this application. It should be understood that the preceding or following structures are not intended to limit this application. The scope of protection of this application is determined by the claims.
[0023] In conversational recommendation systems, achieving personalized recommendations and coherent response generation based on natural language interaction is crucial to ensuring a high level of user satisfaction. Currently, typical conversational recommendation solutions employ a paradigm combining pre-trained language models with external knowledge graphs. The knowledge graph is encoded using graph neural networks, and an attention mechanism is used to inject retrieved knowledge into the pre-trained language model to jointly optimize the recommendation and response generation tasks.
[0024] A knowledge graph is a directed graph structure composed of entities and relations, where each entity is connected to other entities through various relation types, providing rich associative knowledge for recommendations. A dialogue knowledge tree is a tree-like knowledge structure formed by using entities mentioned in the dialogue as root nodes and hierarchically filtering neighboring entities in the knowledge graph based on the current dialogue semantics. Compared to fully connected subgraphs, it has more precise knowledge focusing capabilities and lower noise interference. Learnable soft cues consist of a set of trainable continuous vectors, unlike traditional non-trainable hard cues. They can be optimized through backpropagation during training, thus more flexibly guiding pre-trained language models to adapt to specific tasks.
[0025] As described in the background section, the above-mentioned solutions have the following drawbacks: a large number of noisy entities unrelated to the current dialogue semantics are introduced during the knowledge retrieval and fusion process, interfering with the model's focus on key evidence; there is a representation gap between the structured information of the knowledge graph and the sequence processing capability of the pre-trained language model, which limits the deep reasoning capability; there is insufficient mining of deep collaborative preference signals in multi-turn user dialogues, especially in cold start or ambiguous user feedback scenarios, making it difficult to effectively utilize the directional information of feedback and track changes in preferences across turns; and there are also problems such as high cost of full parameter fine-tuning and conflict with multi-task optimization objectives.
[0026] To systematically address the aforementioned shortcomings, this application proposes a knowledge tree-enhanced conversation recommendation scheme. This scheme aims to construct a conversational knowledge tree by combining user feedback signals to perform differentiated filtering of the knowledge graph, and then serialize the tree structure into text that can be understood by a pre-trained language model, thereby achieving a synergistic improvement in recommendation accuracy and response quality.
[0027] like Figure 1 (The overall architecture diagram of the knowledge tree-enhanced conversation recommendation model of this application is shown.) The core technical idea of this solution is: First, by introducing the directionality (positive / negative) and confidence recognition mechanism of user feedback, the model can accurately capture the user's true intention in multi-turn dialogues and explicitly distinguish between positive and negative preferences. This fundamentally overcomes the key technical defect of traditional methods that rely solely on semantic similarity, which leads to the incorrect introduction of negative entities into the conversation knowledge tree.
[0028] Based on this, the initial knowledge graph is dynamically filtered and paths are constructed according to the identified positive and negative preferences. This ensures that the generated dialogue knowledge tree contains both positive reasoning paths that need to be "closed" and negative interference paths that need to be "suppressed," thereby effectively reducing noise interference at the source of knowledge and improving the model's ability to represent user preferences with high fidelity.
[0029] Furthermore, this application proposes a modular processing paradigm of "serialization-alignment": by efficiently serializing the structured dialogue knowledge tree into a text sequence compatible with the pre-trained language model, and introducing a contrastive learning objective for representation alignment, the structural gap between graph data structures and language models in the representation space is bridged, enabling the pre-trained language model with frozen parameters to directly perform explicit reasoning on structured knowledge rich in user preferences.
[0030] Ultimately, within the parameter-frozen multi-task prompting learning framework, the model is able to collaboratively complete the two core tasks of conversation recommendation and response generation, ensuring both the accuracy of the recommendation results and maintaining the natural fluency of the responses. The aforementioned steps are not simply sequential sequences, but rather form a complete technical loop of "perception—differentiation—filtering—alignment—generation," with close input-output dependencies between each step: directional signals from user feedback drive the differentiated construction of the dialogue knowledge tree; sequentialization and contrastive learning transform graph-structured knowledge into representations directly usable by the language model; and finally, multi-task joint optimization is achieved within the frozen model's cue-based learning framework. The absence or weakening of any step will lead to a significant decline in overall performance.
[0031] The following section will provide a detailed description of the specific implementation methods for each step, in conjunction with the accompanying drawings.
[0032] Figure 2 This is a flowchart illustrating a training method for a knowledge tree-enhanced conversation recommendation model according to some embodiments of this application. Figure 2 As shown, the training method for the knowledge tree-enhanced conversation recommendation model specifically includes: S210, the multi-turn dialogue history and external knowledge graph are encoded respectively to obtain dialogue representation and global entity representation, the mentioned entities are identified from the dialogue history, and the representation of the mentioned entities is obtained from the global entity representation.
[0033] S210 is used to encode multi-turn dialogue history and external knowledge graphs to obtain the dialogue semantic representation and entity semantic representation required for subsequent steps. It is the data foundation for building dialogue knowledge trees and extracting user collaborative preferences.
[0034] In practice, firstly, regarding the history of multi-round dialogues... ( For the first Each round of dialogue is concatenated with special tags and then input into a pre-trained language model such as RoBERTa for encoding, resulting in a token-by-token dialogue semantic representation matrix. ( The total number of conversation tokens. (For the hidden layer dimension), and then average pooling or taking the [CLS] flag bit of the output vector as the overall semantic vector of the dialogue. .
[0035] For external knowledge graphs ( For a collection of entities, For the set of triple edges, For a set of relations), use Relational Graph Convolutional Network (RGCN) aggregates neighbor information according to relation type to obtain a global entity representation matrix. And locate the entities mentioned in the conversation through entity links. ,from Retrieve the corresponding entity representation by index This provides initial dialogue semantics and entity semantics for subsequent collaborative preference modeling and knowledge tree construction.
[0036] S220, perform feedback state identification on each round of user utterance in the multi-turn dialogue history to determine the feedback type of the user utterance in that round and calculate the state confidence. Based on the feedback type and the state confidence, perform interactive calculation and context aggregation on the dialogue representation and the representation of the mentioned entity to generate a user collaborative preference representation and calculate the collaborative preference loss.
[0037] S220 is used to extract user collaborative preference representations from multi-turn dialogue history, providing user interest representations for subsequent recommendation tasks.
[0038] In conversational recommendation, user interests often change dynamically as the conversation progresses, and user expressions may be uncertain; for example, users may give vague or contradictory feedback. Therefore, it is necessary to first identify the feedback state of each round of user utterances to determine the feedback type of that round and calculate the confidence level of the feedback state, thereby providing a basis for subsequent user preference modeling.
[0039] In practice, the feedback state of each round of user speech in the multi-turn dialogue history is first identified. In some embodiments of this application, the feedback types include at least five categories: positive preference, negative preference, weak preference, fuzzy preference, and preference shift.
[0040] Let the first User speech is represented as Its semantic vector is The semantic vector of historical dialogue is Feedback type is Feedback status confidence for: , in, For the Sigmoid function, and For learnable parameters, This indicates element-by-element interaction.
[0041] In obtaining the feedback type in each round and feedback status confidence Then, the feedback direction weight ρ(z) is determined based on the feedback type. t Among them, positive preferences correspond to positive weights, negative preferences correspond to suppressive weights, weak preferences and ambiguous preferences correspond to lower weights, and preference shifts correspond to increasing the weight of the current round and decreasing the weight of previous rounds.
[0042] Based on this, user collaboration preferences are represented as: , in, This is a vector concatenation operation, where Asum(·) is the self-attention aggregation function. It is the first The dialogue indicated that, This refers to the entity mentioned in this round.
[0043] Will With candidate item set Each item entity representation Calculate the inner product matching score And according to the actual recommended items Constructing cross-entropy collaborative preference loss This serves as an auxiliary monitoring signal. This task forces... Even in the absence of prior generative language models, it can still give higher rankings to the correct items, thus explicitly encoding collaborative preferences in multi-turn dialogues.
[0044] To address the issue of inconsistent preferences in multi-turn dialogues, after obtaining the user's collaborative preference representation, the preferred embodiment of this application further includes a cross-turn preference conflict detection and resolution step.
[0045] Specifically, if the current round of preferences conflicts with historical preferences at the entity, attribute, or relationship level, the historical preferences are attenuated or transferred based on the time interval, current feedback confidence, and preference polarity. Among these, single-round preferences... By the Round-robin dialogue t The entities mentioned in this round indicate t and feedback status confidence The calculated result can be expressed as: , in, For the first Feedback type, For the feedback direction weight, This represents the confidence weight of the preferences in this round. The confidence level for the feedback status.
[0046] Historical preferences and current preferences Received from different rounds Its preference polarity According to the corresponding feedback type Defined preferences, for example, positive preferences are of positive polarity, negative preferences are of negative polarity, and fuzzy preferences are of uncertain polarity. Let's assume historical preferences... With current preferences The degree of conflict is: , in, and These are historical round preferences and current preferences, respectively. The preference polarity is determined by the corresponding feedback type. For indicator functions; When the level of conflict exceeds a threshold, the weights of the feedback directions from previous rounds are reduced. , Then, user collaboration preferences are recalculated based on the feedback direction weights of the historical rounds after decay.
[0047] in, The feedback direction weights for historical rounds, For the interval between rounds, The current preference confidence level, This is the attenuation parameter.
[0048] After the conflict is resolved in this step, the knowledge tree constructed subsequently will be based on the updated set of positive and negative preferences, thus preventing expired preferences from continuing to affect the recommendation results.
[0049] S230, perform hierarchical semantic filtering on the knowledge graph to construct a dialogue knowledge tree, serialize the dialogue knowledge tree into text, encode the text to obtain its representation, and calculate the contrastive learning loss based on the representation of the mentioned entities and the representation of the text.
[0050] S230 is used to bridge the representation gap between graph data structures and language models by constructing a dialogue-specific knowledge tree and converting it into a text form that a pre-trained language model can understand in a serialized manner, providing structured external knowledge input for the subsequent prompting learning stage.
[0051] First, construct the dialogue knowledge tree: For each mentioned entity As the root node, in the knowledge graph Perform hierarchical neighbor retrieval with filtering in the middle: In the Layer, for the current layer node All the neighbors who jumped Take its representation after RGCN encoding. With dialogue pooling vector Calculate the cosine similarity score as a semantic relevance score between the neighbor and the current dialogue: , Select by score in descending order Each neighbor is used as a next-level expansion node, thus prioritizing the allocation of the retrieval budget to entities most relevant to the dialogue semantics and filtering out irrelevant common-sense noise. .
[0052] Recursively expand to the maximum depth Obtain the dialogue-specific knowledge tree .
[0053] The dialogue knowledge tree is then serialized into text: Explicitly insert virtual relationship nodes (annotating the relationship type of the edges) between parent and child entities. Then... Perform a depth-first traversal and use entity tags. Relationship markers Distinguish between entity names and relation names, and linearize the tree structure into a plain text sequence that can be directly read by a pre-trained language model: .
[0054] Subsequently, pre-trained language models such as RoBERTa were used again to... The encoding process involves first aggregating the node representations within each subtree using two levels of self-attention, and then aggregating global information across multiple subtrees to obtain a tree-level representation. and overall knowledge tree representation This serves as structured external knowledge to prompt the learning phase.
[0055] Then, the comparative learning loss is calculated: To alleviate physical level ( ) and knowledge tree level ( This indicates inconsistencies in information fusion caused by drift in the semantic space, where the same dialogue sample corresponds to... Positive sample pairs are constructed, and negative sample pairs are constructed from the corresponding representations of different dialogue samples within a batch. A contrastive learning loss in the form of InfoNCE is used. This ensures that the model maintains semantic consistency between explicit graph structure compression and serialized re-representation: , in, Indicates the contrast learning loss. This represents the mentioned entity. The representation of the serialized text. For temperature coefficient, This represents the representation of other samples in the same training batch.
[0056] When user historical profiles are insufficient or dialogue feedback is unstable (including negative, weak, or ambiguous feedback), relying solely on semantic similarity for Top-N neighbor filtering may still diffuse negatively favored entities and their neighborhoods into the knowledge tree, introducing erroneous paths. Therefore, the preferred embodiment of this application replaces simple semantic similarity filtering with a filtering score constrained by positive and negative preferences during the knowledge tree construction process. This ensures that the knowledge tree simultaneously includes both "knowledge paths that should be approached" and "knowledge paths that should be avoided," reducing noise caused by highly connected knowledge nodes.
[0057] Specifically, the feedback type obtained based on S220 and feedback status confidence The dialogue entities are divided into: a set of entities with positive preferences. Negative preference entity set and uncertain entity set .
[0058] For candidate neighbor nodes The screening score has been modified as follows: , in, For dialogue pooling vectors, Representing neighboring entities, , , , λ4 is the weighting coefficient. This is a node noise penalty term, which can be determined by node degree, relation generalization degree, historical round decay factor, or entity ambiguity.
[0059] System basis Select each layer of neighbors and prune nodes that are below a preset threshold.
[0060] During dialogue knowledge tree serialization, in addition to retaining entity tags # and relationship tags $, preference direction tags are further added, for example: <pos>#Movie $starring #Actor represents a positive preference path <neg>#Movie $genre #Horror indicates a negative inhibition path <uncertain>#Movie $similarTo #Movie indicates an uncertain path that needs to be confirmed. Serialized text with preference direction markers is input into pre-trained language models such as RoBERTa for encoding, resulting in tree-level representations and overall knowledge tree representations, which are used in subsequent contrastive learning, alignment, and cue learning stages.
[0061] S240, Perform the recommendation task, including: concatenating a learnable soft cue, the text, the multi-turn dialogue history, and the user collaborative preference representation into a combined cue, inputting the generative pre-trained language model with frozen parameters, and calculating the recommendation loss and response generation loss, wherein the learnable soft cue is used to guide the generative pre-trained language model to adapt to the conversation recommendation task.
[0062] S240 is used to concatenate the dialogue history obtained in S210, the user collaborative preference representation obtained in S220, the serialized knowledge tree obtained in S230, and the learnable soft prompts into a unified prompt input according to a predefined template, and input it into the generative pre-trained language model with frozen parameters to jointly complete the candidate item ranking and natural language response generation, and calculate the recommendation loss and response generation loss.
[0063] Specifically, it will be able to learn soft tips Serialized knowledge tree Dialogue with History With collaborative preference vector The input prompts are assembled according to a predefined template. ,in Used to supplement task-level priors Provide knowledge of dialogue-related graph structures, Provide multi-round collaborative preferences: , Will A generative pre-trained language model (such as DialoGPT) with input frozen parameters completes two sub-tasks through a shared language model head: calculating the ranking distribution on candidate items to obtain the recommendation loss. The generation loss is obtained by autoregressive generation of responses on the vocabulary. .
[0064] S250, construct the total loss based on the collaborative preference loss, the contrastive learning loss, the recommendation loss, and the response generation loss, and update the parameters of the conversation recommendation model.
[0065] S250 is used to construct the total loss and update the model parameters to complete model training.
[0066] Specifically, the total loss is the weighted sum of all items: , in, , , These are the weighting coefficients for the corresponding loss terms.
[0067] This application employs a two-stage training strategy to stabilize optimization and avoid interference between tasks: First, freeze the generative model and train only the recommendation side. Then, all parameters of the conversation recommendation model are fine-tuned.
[0068] The above training data was used to perform a two-stage joint training of the conversation recommendation model based on knowledge tree-enhanced prompt learning proposed in this invention. After the recommendation side was trained, the generation side was fine-tuned. The contribution of each module to the recommendation performance during the training process is as follows: Figure 4 As shown.
[0069] Figure 3 This is a flowchart illustrating a knowledge tree-enhanced conversation recommendation method according to some embodiments of this application. Figure 3 As shown, the knowledge tree-enhanced conversation recommendation method specifically includes: S310: Obtain the multi-turn dialogue history between the current user and the recommendation system, as well as the external knowledge graph related to the recommended items.
[0070] S320, the dialogue history and the knowledge graph are input into a conversation recommendation model, wherein the conversation recommendation model is a model obtained by training as described in S210-S250.
[0071] S330, Obtain the recommended item ranking results and natural language responses for the current dialogue output by the conversation recommendation model.
[0072] Specifically, the system acquires the multi-turn dialogue history between the current user and the recommendation system, as well as an external knowledge graph. This dialogue history and external knowledge graph are then input into a conversation recommendation model trained using the methods described in S210-S250. The model outputs a ranking of recommended items and a natural language response to the current dialogue. The conversation recommendation model encodes the input data, identifies feedback states, constructs and serializes a knowledge tree, concatenates prompts, and generates inferences, ultimately outputting the recommendation results and responses.
[0073] The technical solution of this application will be further described below with reference to specific embodiments.
[0074] The embodiment selected the INSPIRED and ReDial public conversation recommendation datasets in the film domain as experimental data, and DBpedia was used as the external knowledge graph. The overall framework of the model is as follows: Figure 1 As shown.
[0075] A two-stage training strategy is employed to train the knowledge tree-enhanced conversation recommendation model proposed in this application: the first stage freezes the generative pre-trained language model and trains only the recommendation-related modules; the second stage jointly fine-tunes all parameters of the conversation recommendation model. The contribution of each module to the recommendation performance during training is as follows: Figure 4 (The ablation experiment results of one embodiment of this application on the INSPIRED dataset for the recommendation task are shown.)
[0076] To verify the actual contribution of each module proposed in this application to recommendation performance, ablation experiments were conducted on the model on the INSPIRED dataset. In the experiments, models S210-S250 and several variant models were constructed, including: removing the user preference module, removing the knowledge tree module, removing the semantic alignment module, using triples to replace the knowledge tree, and using an unfiltered knowledge tree, and their performance was tested respectively. Experimental results are as follows... Figure 4 As shown in Table 1.
[0077] Table 1: Comparison of ablation experimental results based on the INSPIRED dataset
[0078] From the data in Table 1 and Figure 4 It can be seen that: The complete model of this application achieved optimal values in all metrics, with recall@50 reaching 0.438 and ndcg@50 reaching 0.222, significantly outperforming other variant models.
[0079] Key module validity: 1. Removing the knowledge tree module had the greatest impact on performance, causing its recall@10 to drop significantly to 0.223 (a decrease of approximately 19.7%) and ndcg@10 to drop to 0.150. This indicates that aggregating high-order neighbor information through the knowledge tree can effectively capture users' deep interests.
[0080] 2. Removing the semantic alignment module caused ndcg@50 to drop to 0.208, demonstrating the importance of aligning the dialogue context with the semantics of the knowledge graph.
[0081] 3. The performance of "using unfiltered knowledge trees" is slightly inferior (e.g., mrr@50 is only 0.162), which verifies that the knowledge tree neighbor filtering mechanism in this application can effectively remove graph noise and improve recommendation accuracy.
[0082] In summary, the training method and the method for knowledge tree-enhanced conversation recommendation provided in the embodiments of this application overcome the problem of insufficient knowledge noise filtering in existing technologies by constructing a conversation knowledge tree through hierarchical semantic filtering of the knowledge graph based on the dialogue representation, thereby achieving dynamic and refined filtering of knowledge from the source; by serializing the conversation knowledge tree into natural language text and aligning its representation with entity representation using contrastive learning loss, the graph structure is transformed into a symbol sequence that can be directly reasoned about by the pre-trained language model, bridging the representation gap and solving the problem of inefficient use of graph structure information; and by performing interactive computation and context aggregation based on dialogue and entity representation... This approach generates user collaborative preference representations and calculates collaborative preference losses. This explicitly models the complex relationships between multiple mentioned entities in the dialogue context and uses auxiliary supervision signals to ensure that the extracted representations effectively reflect the user's collaborative interests, thus overcoming the difficulty of deep user preference modeling in existing technologies. By employing a parameter-frozen generative pre-trained language model as the core engine and constructing a total loss that integrates collaborative preferences, contrastive learning, recommendation, and generation losses, it jointly optimizes the parameters for prompts, preference extraction, and knowledge encoding. This efficiently coordinates multiple tasks while retaining the general capabilities of a large model, breaking the dilemma of balancing training cost and multi-task performance in existing technologies. Furthermore, by distinguishing the directional information of user feedback and effectively tracking preference changes, the above scheme further improves the modeling accuracy of the user's true intent in cold-start or uncertain feedback scenarios. These technical features—dynamic knowledge filtering, structured serialization and alignment, supervised preference extraction, and multi-task prompt learning under a frozen model—are not simply superimposed but mutually coupled and deeply collaborative, jointly achieving a synergistic improvement in recommendation accuracy, response relevance, knowledge utilization, and depth of user intent understanding in the conversational recommendation system.
[0083] Although the foregoing embodiments are described within the general context of computer systems running in conjunction with operating systems and applications, those skilled in the art will recognize that other implementations can also be performed by incorporating other types of program modules. Generally, program modules include routines, programs, components, data structures, and other types of structures that perform specific tasks or implement specific abstract data types. Those skilled in the art will understand that the subject matter described herein can be practiced using other computer system configurations, including handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframes, etc., and can also be used in distributed computing environments where tasks are performed by remote processing devices connected via communication networks. In a distributed computing environment, program modules may reside on both local and remote memory storage devices.
[0084] Those skilled in the art will recognize that the units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0085] It should be understood that the specific embodiments described above are merely illustrative or explanatory of the principles of this application and do not constitute a limitation thereof. Therefore, any modifications, equivalent substitutions, improvements, etc., made without departing from the spirit and scope of this application should be included within the protection scope of this application. Furthermore, the appended claims are intended to cover all variations and modifications falling within the scope and boundaries of the appended claims, or equivalent forms of such scope and boundaries.< / uncertain> < / neg> < / pos>
Claims
1. A training method for a knowledge tree-enhanced conversation recommendation model, wherein the conversation recommendation model includes a generative pre-trained language model, characterized in that, include: The multi-turn dialogue history and external knowledge graph are encoded to obtain dialogue representation and global entity representation. The mentioned entities are identified from the dialogue history, and the representation of the mentioned entities is obtained from the global entity representation. Feedback state identification is performed on each round of user utterance in the multi-turn dialogue history to determine the feedback type of the user utterance in that round and calculate the feedback state confidence. Based on the feedback type and the feedback state confidence, interactive calculation and context aggregation are performed on the dialogue representation and the representation of the mentioned entity to generate a user collaborative preference representation and calculate the collaborative preference loss. The knowledge graph is subjected to hierarchical semantic filtering to construct a dialogue knowledge tree. The dialogue knowledge tree is serialized into text, the text is encoded to obtain its representation, and a contrastive learning loss is calculated based on the representation of the mentioned entity and the representation of the text. Performing a recommendation task includes: concatenating a learnable soft cue, the text, the multi-turn dialogue history, and the user collaborative preference representation into a combined cue, inputting the generative pre-trained language model with frozen parameters, and calculating the recommendation loss and response generation loss, wherein the learnable soft cue is used to guide the generative pre-trained language model to adapt to the conversation recommendation task; The total loss is constructed based on the collaborative preference loss, the contrastive learning loss, the recommendation loss, and the response generation loss, and the parameters of the conversation recommendation model are updated accordingly.
2. The method according to claim 1, characterized in that, The step of performing interactive calculations and context aggregation on the dialogue representation and the representation of the mentioned entities based on the feedback type and the feedback state confidence level to generate a user collaborative preference representation includes: User collaborative preference representation is generated based on the following formula: , in, This is a vector concatenation operation, where Asum(·) is the self-attention aggregation function. It is the first The dialogue indicated that, The entities mentioned in this round indicate that It is the confidence level of the feedback state; The feedback direction weights are determined based on the feedback type, wherein positive preferences correspond to positive weights, negative preferences correspond to suppression weights, weak preferences and ambiguous preferences correspond to lower weights, and preference shifts correspond to increasing the weight of the current round and decreasing the weight of previous rounds.
3. The method according to claim 2, characterized in that: The feedback types include at least positive preference, negative preference, weak preference, fuzzy preference, and preference shift; No. Confidence of user feedback status Calculate according to the following formula: , in, For the Sigmoid function, For the first The semantic vector of user discourse, For historical dialogue semantic vectors, and These are learnable parameters.
4. The method according to claim 3, characterized in that, Also includes: Conflict detection is performed on preferences in different rounds based on the feedback type and the feedback state confidence level, wherein historical round preferences are calculated according to the following formula. With current preferences Level of conflict: , in, and These are historical round preferences and current preferences, respectively. The preference polarity is determined by the corresponding feedback type. For indicator functions; When the degree of conflict exceeds a threshold, the weights of the feedback directions from previous rounds are reduced: , in, The feedback direction weights for historical rounds, For the interval between rounds, The current preference confidence level, This is the attenuation parameter.
5. The method according to claim 1, characterized in that, The construction of the dialogue knowledge tree includes: Based on the feedback type, the mentioned entities are categorized into a set of entities with positive preferences. Negative preference entity set and uncertain entity set , In the knowledge graph, candidate nodes are selected, and for each candidate node, a screening score is calculated according to the following formula. : , in, For dialogue pooling vectors, Representing neighboring entities, , This is a node noise penalty term. , , λ4 are weighting coefficients; Based on the screening score, neighboring nodes are selected from the candidate nodes to construct a dialogue knowledge tree.
6. The method according to claim 1, characterized in that, The step of serializing the dialogue knowledge tree into text includes: The tree structure is transformed into text by performing a depth-first traversal of the dialogue knowledge tree and inserting markers to distinguish entity names from relation names during the traversal.
7. The method according to claim 1, characterized in that, The contrastive learning loss is calculated using the following formula: , in, Indicates the contrast learning loss. This represents the mentioned entity. The representation of the serialized text. For temperature coefficient, This represents the representation of other samples in the same training batch.
8. The method according to claim 1, characterized in that, The total construction loss includes: The collaborative preference loss The contrastive learning loss The recommended loss and the loss generated by the response We perform a weighted summation to obtain the total loss L: , in, , , These are the weighting coefficients for the corresponding loss terms.
9. The method according to claim 1, characterized in that, The parameters of the conversation recommendation model, excluding the generative pre-trained language model, are updated using a two-stage training strategy. This two-stage training strategy includes: In the first stage, the parameters of the generative pre-trained language model are kept frozen, and the parameters of the conversation recommendation model other than the generative pre-trained language model are updated based on the collaborative preference loss, the contrastive learning loss and the recommendation loss. In the second phase, all parameters of the session recommendation model are jointly fine-tuned.
10. A knowledge tree-enhanced conversation recommendation method, characterized in that, include: Obtain the multi-turn dialogue history between the current user and the recommendation system, as well as external knowledge graphs related to the recommended items; The dialogue history and the knowledge graph are input into a conversation recommendation model, wherein the conversation recommendation model is a model obtained by the training method described in any one of claims 1-9; Obtain the recommended item ranking results and natural language responses for the current dialogue output by the conversation recommendation model.