Dialogue recommendation method and device, equipment and storage medium

By analyzing user emotions through a dialogue recommendation model and combining them with external knowledge sources to generate emotional responses, this approach solves the problem of neglecting user emotions in dialogue recommendation technology, achieving more natural dialogue transitions and higher recommendation accuracy.

CN120821792APending Publication Date: 2025-10-21TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410437114.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-11
Publication Date
2025-10-21

Smart Images

  • Figure CN120821792A_ABST
    Figure CN120821792A_ABST
Patent Text Reader

Abstract

The invention discloses a dialogue recommendation method and device, equipment and a storage medium, and relates to the technical field of computers. The method comprises the steps of obtaining at least one historical dialogue statement; generating a first answer statement according to the at least one historical dialogue statement through a dialogue recommendation model, the first answer statement including a recommendation object for the at least one historical dialogue statement; a second answer statement is generated through a dialogue recommendation model according to the first answer statement and entity vocabularies which are obtained from an external knowledge source and related to the recommendation object, and the second answer statement comprises the first answer statement and an emotional answer statement for the recommendation object. According to the method, the recommendation object conforming to expectation and preference of the user is generated, the accuracy of the recommendation object is improved, emotional resonance between the recommendation object and the user is considered, the content richness of the second answer statement is improved, and the emotional sharing ability of the dialogue recommendation model for the user demand is enhanced, so that the persuasion of the recommended dialogue is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a conversation recommendation method, apparatus, device, and storage medium. Background Art

[0002] Through real-time, dynamic and multiple online discourse interactions, conversational recommendation technology can quickly reveal users' current preferences and effectively explain their click and purchase behaviors.

[0003] Related technologies use generation-driven conversational recommendation technology. This technology focuses on generating fluent responses and accurate recommendations. This technology makes suggestions through natural language text, where the flexibility of the text has a significant impact on the discourse process.

[0004] However, the above method tends to ignore the user's emotions in the conversation, resulting in a clunky connection between the generated conversation and the previous text. Summary of the Invention

[0005] The present invention provides a method, apparatus, device, and storage medium for recommending conversations. The technical solutions provided by the present invention are as follows:

[0006] According to one aspect of an embodiment of the present application, a conversation recommendation method is provided, the method comprising:

[0007] Get at least one historical dialogue sentence;

[0008] generating, by a dialogue recommendation model, a first answer sentence based on the at least one historical dialogue sentence, wherein the first answer sentence includes a recommendation object for the at least one historical dialogue sentence;

[0009] The dialogue recommendation model generates a second answer statement based on the first answer statement and entity vocabulary related to the recommendation object obtained from an external knowledge source. The second answer statement includes the first answer statement and an emotional answer statement for the recommendation object.

[0010] According to one aspect of an embodiment of the present application, a method for training a conversation recommendation model is provided, the method comprising:

[0011] Obtaining a basic dataset for training the dialogue recommendation model, the basic dataset including at least one data group, each of the data groups including at least one sample dialogue sentence, a first label sentence corresponding to the at least one sample dialogue sentence, and a second label sentence corresponding to the at least one sample dialogue sentence, wherein the first label sentence includes a label object for the at least one sample dialogue sentence, and the second label sentence includes the first label sentence and an emotional response sentence for the label object;

[0012] generating, based on the at least one data group, a first training sample corresponding to a reply task, wherein the reply task uses an initial labeled sentence as label data and determines sample data based on the at least one sample dialogue sentence, wherein the initial labeled sentence includes a blank object for the at least one sample dialogue sentence;

[0013] generating, based on the at least one data group, a second training sample corresponding to an object recommendation task, wherein the object recommendation task uses the labeled object as label data and determines sample data based on the at least one sample dialogue sentence and the first label sentence;

[0014] generating, based on the at least one data group, a third training sample corresponding to a sentiment alignment task, wherein the sentiment alignment task uses the second labeled sentence as labeled data and determines sample data based on the at least one sample dialogue sentence and the first labeled sentence;

[0015] The first training sample, the second training sample, and the third training sample are used to train the dialogue recommendation model to obtain a trained dialogue recommendation model.

[0016] According to one aspect of an embodiment of the present application, a conversation recommendation device is provided, the device comprising:

[0017] A sentence acquisition module, used to acquire at least one historical conversation sentence;

[0018] A first generating module, configured to generate a first answer sentence based on the at least one historical dialogue sentence using a dialogue recommendation model, wherein the first answer sentence includes a recommendation object for the at least one historical dialogue sentence;

[0019] The second generation module is used to generate a second answer statement based on the first answer statement and entity vocabulary related to the recommendation object obtained from an external knowledge source through the dialogue recommendation model, wherein the second answer statement includes the first answer statement and an emotional answer statement for the recommendation object.

[0020] According to one aspect of an embodiment of the present application, a training device for a dialogue recommendation model is provided, the device comprising:

[0021] a data acquisition module, configured to acquire a basic data set for training the dialogue recommendation model, wherein the basic data set includes at least one data group, each of which includes at least one sample dialogue sentence, a first label sentence corresponding to the at least one sample dialogue sentence, and a second label sentence corresponding to the at least one sample dialogue sentence, wherein the first label sentence includes a label object for the at least one sample dialogue sentence, and the second label sentence includes the first label sentence and an emotional response sentence for the label object;

[0022] a first generating module, configured to generate, based on the at least one data group, a first training sample corresponding to a reply task, wherein the reply task uses an initial labeled sentence as labeled data and determines sample data based on the at least one sample dialogue sentence, wherein the initial labeled sentence includes a blank object for the at least one sample dialogue sentence;

[0023] a second generating module, configured to generate, based on the at least one data group, a second training sample corresponding to an object recommendation task, wherein the object recommendation task uses the labeled object as label data and determines sample data based on the at least one sample dialogue sentence and the first label sentence;

[0024] a third generating module, configured to generate, based on the at least one data group, a third training sample corresponding to a sentiment alignment task, wherein the sentiment alignment task uses the second labeled sentence as labeled data and determines sample data based on the at least one sample dialogue sentence and the first labeled sentence;

[0025] A training module is used to train the dialogue recommendation model using the first training sample, the second training sample, and the third training sample to obtain a trained dialogue recommendation model.

[0026] According to one aspect of an embodiment of the present application, a computer device is provided, comprising a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the above-mentioned conversation recommendation method, or the training method of the conversation recommendation model.

[0027] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is loaded and executed by a processor to implement the above-mentioned dialogue recommendation method, or the training method of the dialogue recommendation model.

[0028] According to one aspect of an embodiment of the present application, a computer program product is provided, which includes a computer program, and the computer program is loaded and executed by a processor to implement the above-mentioned dialogue recommendation method, or the training method of the dialogue recommendation model.

[0029] The technical solutions provided in the embodiments of the present application can bring the following beneficial effects:

[0030] A conversational recommendation model is employed to generate a first response based on at least one historical conversational statement. The model also generates a second response based on the first response and entity vocabulary related to the recommended object obtained from an external knowledge source. Compared to related conversational recommendation techniques that tend to ignore user emotions during conversations, the conversational recommendation model provided in this application analyzes user emotions in historical conversational statements to recommend objects that meet user expectations and preferences, thereby improving the accuracy of recommended objects. Furthermore, based on entity vocabulary related to the recommended object obtained from an external knowledge source, a second response containing an emotional response is generated. This not only considers emotional resonance with the user, making the model-generated recommended conversation more naturally connected to the preceding context, but also enhances the content richness of the second response, strengthening the conversational recommendation model's ability to empathize with user needs, thereby enhancing the persuasiveness of the recommendation conversation. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 This is a schematic diagram of an implementation environment for a solution provided by an embodiment of the present application;

[0032] Figure 2 This is a schematic diagram of dialogue recommendation for film and television works provided by an embodiment of the present application;

[0033] Figure 3 This is a comparison and recommendation diagram of the related technology provided by an embodiment of the present application and the solution of the present application;

[0034] Figure 4 This is a flowchart of a conversation recommendation method provided by an embodiment of the present application;

[0035] Figure 5 This is a schematic diagram of a conversation recommendation process provided by an embodiment of the present application;

[0036] Figure 6 This is a flowchart of a method for training a conversation recommendation model provided by one embodiment of the present application;

[0037] Figure 7 This is a block diagram of a conversation recommendation device provided by one embodiment of the present application;

[0038] Figure 8 This is a block diagram of a conversation recommendation device provided by one embodiment of the present application;

[0039] Figure 9 This is a structural block diagram of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0040] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0041] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0042] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0043] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning by demonstration. Pretrained models are the latest development in deep learning, integrating these techniques.

[0044] A Large Language Model (LLM) is an artificial intelligence algorithm based on deep learning technology, whose goal is to enable computers to understand and generate natural language. It analyzes large amounts of language data, such as text, speech, or images, to learn the structure and patterns of language and use this knowledge to perform various natural language processing tasks, such as machine translation, speech recognition, text classification, and question-answering systems. Large language models typically use the Transformer architecture used in deep learning to model text sequences to understand context and semantics. Their training process typically involves vast amounts of data and computing resources, such as large-scale corpora and high-performance computing platforms. During training, large language models gradually learn the characteristics and patterns of language, developing the ability to understand and express it.

[0045] The Transformer architecture is a deep learning model that uses a self-attention mechanism, which assigns different weights to different parts of the input data based on their importance. This architecture is primarily used in natural language processing and computer vision (CV). It typically includes components such as self-attention, multi-head attention, positional encoding, residual connections and normalization (Add & Norm), a feed-forward network, and a position-wise feed-forward network, which together form the encoder and decoder.

[0046] A pre-training model (PTM), also known as a cornerstone model or large model, refers to a deep neural network (DNN) with large parameters. It is trained on massive amounts of unlabeled data. Leveraging the function approximation capabilities of large-parameter DNNs, the PTM extracts common features from the data. Through techniques such as fine tuning, parameter-efficient fine-tuning (PEFT), and prompt-tuning, the PTM is then adapted for downstream tasks. Therefore, pre-trained models can achieve ideal results in few-shot or zero-shot scenarios. Based on the data modality processed, PTMs can be categorized into language models (ELMO, BERT, GPT), vision models (swin-transformer, ViT, V-MOE), speech models (VALL-E), and multimodal models (ViBERT, CLIP, Flamingo, Gato). Multimodal models are those that represent features from two or more data modalities. Pre-trained models are important tools for outputting artificial intelligence generated content (AIGC) and can also serve as a universal interface for connecting multiple specific task models.

[0047] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence generated content (Artificial Intelligence Generated Content, AIGC), conversational interaction, smart medical care, intelligent customer service, game AI, virtual reality (VR), augmented reality (AR), etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0048] The technical solution of this application mainly involves machine learning technology in artificial intelligence technology, mainly involving the training and use process of dialogue recommendation models.

[0049] Please refer to Figure 1 , which shows a schematic diagram of a solution implementation environment provided by an embodiment of the present application. The solution implementation environment can be implemented as a dialogue recommendation system. The solution implementation environment may include: a model training device 10 and a model use device 20.

[0050] The model training device 10 can be an electronic device such as a mobile phone, tablet computer, laptop computer, desktop computer, smart TV, multimedia player, vehicle-mounted terminal, server, intelligent robot, or other electronic device with strong computing power. The model training device 10 is used to train the dialogue recommendation model.

[0051] In the present embodiment, the conversation recommendation model is a machine learning model trained using a conversation recommendation model training method, and is used to generate emotional responses to historical conversation sentences. The model training device 10 can use machine learning to train the conversation recommendation model, enabling it to generate emotional responses to historical conversation sentences. The specific model training method can be found in the following embodiments.

[0052] In one embodiment of the present application, the input data of the conversation recommendation model includes historical conversation sentences and a first feature set corresponding to a first vocabulary set in the historical conversation sentences. The output data is an initial response sentence. The initial response sentence is a response sentence to the historical conversation sentence. The initial response sentence contains a blank object, that is, a masked object.

[0053] In one embodiment of the present application, the input data of the dialogue recommendation model includes historical dialogue sentences, emotional features corresponding to the historical dialogue sentences, and initial response sentences, and the output data is the recommended objects. The recommended objects are used to fill in the blank objects in the initial response sentence.

[0054] In one embodiment of the present application, the input data of the conversational recommendation model includes triples obtained from an external knowledge source, entity vocabulary obtained from an external knowledge source, recommended objects, and an initial response sentence, and the output data is a second response sentence. The second response sentence contains the initial response sentence, the recommended object, and an emotional response sentence for the recommended object.

[0055] The trained plot extraction model can be deployed in model-using device 20 for use. Model-using device 20 can be an electronic device such as a mobile phone, tablet computer, laptop computer, desktop computer, smart TV, multimedia player, in-vehicle terminal, server, intelligent robot, or other electronic device with strong computing capabilities. When it is necessary to generate emotional responses to historical conversation sentences, model-using device 20 can achieve this function using the trained conversation recommendation model.

[0056] The model training device 10 and the model using device 20 can be two independent devices or the same device. If the model training device 10 and the model using device 20 are the same device, the model training device 10 can be deployed in the model using device 20.

[0057] In the embodiment of the present application, the execution subject of each step may be a computer device, which may be as follows: Figure 1 The model training device 10 may also be a model using device 20. The server may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0058] Advances in conversational recommendation technology have greatly enhanced human-computer interaction capabilities and sparked a new trend in integrating natural language processing with recommendation technology, spurring the development of conversational recommendation technology. Through dynamic, real-time conversations, conversational recommendation models capture users' real-time preferences and understand the underlying reasons behind their consumption behaviors. They can then predict user needs by analyzing the intent revealed in their conversations. Even if the initial recommendations don't fully match user expectations, they can be modified based on user feedback.

[0059] In some embodiments, the dialogue recommendation model provided in the embodiments of the present application can be applied to dialogue recommendations for film and television works, such as Figure 2 As shown in the figure, the conversational recommendation model (recommender) is on the left side of the conversational sentence, and the user is on the right. Based on the user's recommendation request, the user's desired movie is a romantic movie with a comedic tone. In response, the conversational recommendation model recommends a movie to the user and generates an emotionally charged response, explaining the reasoning behind the recommendation. This enhances the conversational recommendation model's empathy for the user's needs, thereby increasing the persuasiveness of its recommendations.

[0060] In some embodiments, the conversational recommendation model provided by the embodiments of this application can be applied to conversational product recommendations. In each round of conversation, the user expresses their desired product to the conversational recommendation model. The conversational recommendation model analyzes the user's needs based on historical conversational statements, recommends products to the user, and provides reasons for the recommendation, i.e., emotional responses.

[0061] The related art of conversational recommendation technology fails to fully guide users’ actual needs and often ignores the importance of user emotions, which leads to the related technology simply repeating the behaviors recorded in the dataset rather than deeply exploring the user’s true preferences. For recommendation tasks, related technologies often assume that all entities mentioned in the conversation sentences meet the user’s expectations. However, in Figure 3In the example, although the recommender (dialogue recommendation model) mentioned two entities in the dialogue sentence "Have you seen XXXX's classic movie "XXXX"?", the user's dialogue sentence showed no interest in these entities. In terms of generating responses, the standard responses in traditional dialogue recommendation datasets are often concise and lack description, resulting in a clumsy connection with the previous text and ignoring the emotional resonance with the user, which in turn reduces user satisfaction. Figure 3 As shown on the left side of the "System Interaction", the system response, which is trained with the standard response of the dataset, only contains object names such as movie names. It has poor interpretability and user experience, and is inconsistent with the actual needs of users. Even if it performs well in some evaluation indicators, it is difficult to apply to actual business scenarios. The conversation recommendation model provided by this application, such as Figure 3 As shown on the right side of "System Interaction", it covers two core components: "emotion-aware object recommendation" and "emotion-aligned response generation". In the emotion-aware object recommendation part, user emotions are combined with entities mentioned in dialogue sentences to learn user preferences. In the emotion-aligned response generation part, the dialogue recommendation model can generate responses consistent with human emotions and integrate entity vocabulary related to the recommended objects from external knowledge sources as generation prompts to generate emotional responses. For example Figure 3 As shown in the figure, the answer sentence of the dialogue recommendation model contains the recommended objects obtained by perceiving the emotions in the historical dialogue sentences, and contains the emotional answer sentence aligned with the emotions, "You should watch the romantic love movie "XXXX". I declare that this movie is my favorite! In my opinion, this is the best work of XXXX. Everything about this movie is beautiful and perfect!".

[0062] In some embodiments, the conversation recommendation model provided in the embodiments of the present application may also be referred to as an ECR (Empathetic Conversational Recommender) model.

[0063] Please refer to Figure 4 , which shows a flowchart of a conversation recommendation method provided by an embodiment of the present application. The execution subject of each step of the method can be a computer device. The method can include at least one of the following steps 410 to 430.

[0064] Step 410: Obtain at least one historical conversation sentence.

[0065] Historical conversation sentences refer to conversation sentences previously exchanged between the user and the conversation recommendation model, with the last conversation sentence in the historical conversation sentence being one uttered by the user. The conversation recommendation method provided in this application is used to generate a response sentence to at least one historical conversation sentence based on at least one historical conversation sentence. This can also be understood as generating a response sentence to the last conversation sentence uttered by the user based on at least one historical conversation sentence.

[0066] For example, Figure 3 As shown, at least one historical dialogue sentence includes "What type of movies do you like?" issued by the dialogue recommendation model, "I like romantic movies." issued by the user, "Have you seen XXXX's classic movie "XXXX"? " issued again by the dialogue recommendation model, and "I have seen it, but I don't like it very much." The dialogue recommendation method provided in this application is used to generate the final reply sentence of the dialogue recommendation model based on the historical dialogue sentences issued by the user and the dialogue recommendation model, "You should watch the romantic love movie "XXXX". I declare this movie my favorite! In my opinion, this is XXXX's best work. Everything in this movie is beautiful and perfect!"

[0067] Step 420 : Generate a first answer statement based on at least one historical dialogue statement through a dialogue recommendation model, wherein the first answer statement includes a recommendation object for the at least one historical dialogue statement.

[0068] Recommended items are items that meet the user's expectations and preferences, as determined by the conversational recommendation model based on user sentiment analysis of at least one historical conversational statement. The category of the recommended item is determined by the category of the item targeted by the historical conversational statement between the user and the conversational recommendation model, and is not limited in this application. For example, if the historical conversational statement is about movie recommendations, the recommended item category is movies; if the historical conversational statement is about drug recommendations, the recommended item category is drugs.

[0069] The first reply is a preliminary reply to at least one historical conversation statement. It includes a recommended object for the at least one historical conversation statement and an initial reply to the recommended object. The portion of the initial reply that refers to the recommended object is masked, and the semantics of the initial reply are identical to those of the recommended object.

[0070] For example, Figure 3In the example dialogue for movie recommendations, the user expresses their expectations and preferences for the recommended movie. Specifically, the user requests a romantic movie from the dialogue recommendation model, and the user dislikes the movie initially recommended by the dialogue recommendation model. In response, the dialogue recommendation model recommends a romantic movie to the user based on the previous dialogue. For example, the first response is "You should watch the romantic movie 'XXXX'." The recommended movie is 'XXXX,' and the initial response is "You should watch the romantic movie XXXX [mask]."

[0071] In some embodiments, the generation of the first answer statement can refer to UniCRS (Towards Unified Conversational Recommender Systems via Knowledge-Enhanced Prompt Learning, a unified conversational recommendation system achieved through knowledge-enhanced prompt learning). UniCRS mainly includes three sub-processes: a semantic fusion process, a response generation process, and an object recommendation process. The semantic fusion process is used to fuse the semantic space of historical conversation sentences and external knowledge sources to obtain the semantic features of the fused historical conversation sentences. The response generation process is used to generate an initial answer statement based on the semantic features of the fused historical conversation sentences and the historical conversation sentences. The object recommendation process is used to generate a recommended object based on the semantic features of the fused historical conversation sentences, the historical conversation sentences, and the initial answer statement.

[0072] In some embodiments, the above step 420 includes at least one sub-step of steps 421 to 423 (not shown in the figure).

[0073] Step 421 : Generate an initial answer sentence based on at least one historical dialogue sentence through a dialogue recommendation model, where the initial answer sentence contains a blank object for at least one historical dialogue sentence.

[0074] The initial response statement is a response statement to at least one historical conversation statement. It is generated by analyzing the semantics of at least one historical conversation statement and meets the user's recommendation requirements. In other words, the initial response statement contains the semantics corresponding to the user's recommendation requirements.

[0075] Blank objects are used to indicate recommendation objects that the conversational recommendation model has not yet determined. Therefore, the initial response sentence contains a blank object for at least one historical conversation sentence, but no recommended object. Blank objects are objects that the conversational recommendation model has yet to determine and that are tailored to the user's recommendation needs. In the initial conversation sentence, blank objects refer to the masked regions of the sentence.

[0076] For example, Figure 3As shown in the figure, the initial response is "You should watch the romantic movie XXXX[mask]." This response contains the semantics corresponding to the user's recommendation request, meaning that the user wants the conversational recommendation model to recommend a romantic movie. The initial response also contains a blank object for at least one historical conversational statement. This blank object represents the romantic movie category to be determined by the conversational recommendation model.

[0077] In some embodiments, the dialogue recommendation model is used to analyze the semantic features of at least one historical dialogue sentence based on at least one historical dialogue sentence to generate an initial answer sentence.

[0078] In some embodiments, the process of generating the initial answer sentence includes a semantic fusion process and a response generation process in UniCRS.

[0079] Step 422: Generate a recommendation object based on at least one historical dialogue sentence and an initial answer sentence using a dialogue recommendation model.

[0080] For example, Figure 2 As shown, the recommendation object is the romantic love movie "XXXX" (movie name).

[0081] In some embodiments, the conversation recommendation model is used to analyze user emotions in at least one historical conversation sentence based on at least one historical conversation sentence and an initial answer sentence to generate a recommendation object.

[0082] In some embodiments, the process of generating the initial answer sentence includes an object recommendation process in UniCRS.

[0083] Step 423: Fill the blank object in the initial answer statement with the recommended object to obtain the first answer statement.

[0084] The recommended object obtained in step 422 is filled into the blank object in the initial answer sentence obtained in step 421, that is, the recommended object is filled into the masked sentence area in the initial answer sentence to obtain the first answer sentence.

[0085] For example, Figure 3 As shown in the figure, the initial answer sentence is "You should watch the romantic love movie XXXX[mask]". If the recommended object is movie A, movie A is filled into the masked sentence area in the initial answer sentence, and the first answer sentence is obtained, "You should watch the romantic love movie - movie A."

[0086] The conversation recommendation model first generates an initial conversation statement that meets the user's recommendation requirements based on the semantic features of at least one historical conversation statement. It then further analyzes the user's emotions within at least one historical conversation statement to generate recommendations that meet the user's expectations and preferences. By analyzing the semantic features of historical conversation statements first, then the emotional features of historical conversation statements, and then progressively determining the initial response statement, this model avoids in-depth analysis of historical conversation statements, which could affect the model's efficiency and reduces the difficulty of sentence analysis.

[0087] In step 430, a second answer statement is generated through the dialogue recommendation model based on the first answer statement and entity vocabulary related to the recommendation object obtained from an external knowledge source. The second answer statement includes the first answer statement and an emotional answer statement for the recommendation object.

[0088] External knowledge sources include knowledge graphs and review databases. The knowledge graph is a knowledge database built based on objective facts, and the data in the knowledge graph will not change or transfer due to human will. For example, the data about a movie in the knowledge graph includes the director, screenwriter, actors, etc. of the movie. The data related to the movie is factual data that already exists and will not change, that is, the director of the movie will not change due to human will. The review database is a database built based on the subjective will of each user. The review database contains the review statements made by each user on different objects. The data in the review database can change due to changes in the user's will. For example, a user originally commented on movie A and expressed his liking for movie A. After a period of time, the user changed his feedback on movie A and could modify the comment statement on movie A to express his dislike for movie A.

[0089] Entity vocabulary related to the recommended object obtained from external knowledge sources, including entity vocabulary related to the recommended object obtained from the knowledge graph and entity vocabulary related to the recommended object obtained from the review database.

[0090] A vocabulary is at least one character in a sentence used to describe semantic information; in other words, a vocabulary is a meaningful string of characters in a sentence. Vocabulary can be used to describe entity information, character emotions, character actions, and so on. Entity vocabulary is the vocabulary used in a sentence to describe entity information. Entities refer to objectively existing people or objects. Objects can be real objects or virtual objects with symbolic meaning. Entity vocabulary primarily includes names of people, places, organizations, proper nouns, and so on. The types of entity vocabulary contained in different types of sentences often vary. For example, in news sentences, entity vocabulary primarily includes names of people, places, and organizations. In product-related sentences, entity vocabulary primarily includes brand names, product names, and item attribute names. In movie-related sentences, entity vocabulary primarily includes movie titles, director names, and actor names.

[0091] The emotional response sentence for the recommended object is aligned with human emotions and is used to express the reason for recommending the recommended object. The entity vocabulary in the emotional response sentence for the recommended object is extracted from the entity vocabulary related to the recommended object obtained from an external knowledge source, and is used to increase the richness of the answer content of the emotional response sentence to enhance the persuasiveness of the emotional response sentence. The emotional vocabulary contained in the emotional response sentence for the recommended object is generated based on the emotional features contained in the entity vocabulary related to the recommended object obtained from an external knowledge source, and is aligned with human emotions. Generally speaking, emotional vocabulary is a descriptive vocabulary used to express positive emotions towards the recommended object. For example, it can express the dialogue recommendation model's liking for the recommended object, express a positive evaluation of the recommended object, or express the target population of the recommended object, etc.

[0092] For example, Figure 3 As shown, the emotional response sentence for the recommended object is "I declare that this movie is my favorite! In my opinion, this is the best work of XXXX. Everything about this movie is beautiful and perfect!" This emotional response sentence contains entity words such as "movie", "XXXX", and "work", as well as emotional words such as "favorite", "best", "beautiful", "very", and "perfect", showing the style and emotional expression of the recommended movie.

[0093] In some embodiments, the dialogue recommendation model is used to generate a response consistent with human emotions based on the first answer statement and entity vocabulary related to the recommendation object obtained from an external knowledge source, thereby obtaining a second answer statement including an emotional answer statement.

[0094] In some embodiments, the conversational recommendation model is configured to generate a response consistent with human emotions based on the first answer statement and entity vocabulary related to the recommended object obtained from an external knowledge source, thereby obtaining an emotional answer statement for the recommended object. The first answer statement and the emotional answer statement are concatenated to obtain a second answer statement.

[0095] The technical solution provided by the embodiments of the present application utilizes a conversational recommendation model to generate a first response statement based on at least one historical conversational statement. The conversational recommendation model also utilizes the first response statement and entity vocabulary related to the recommended object obtained from an external knowledge source to generate a second response statement. Compared to related conversational recommendation techniques that tend to ignore user emotions during conversations, the conversational recommendation model provided by the present application analyzes user emotions in historical conversational statements to obtain recommended objects that meet user expectations and preferences, thereby improving the accuracy of recommended objects. Furthermore, based on entity vocabulary related to the recommended object obtained from an external knowledge source, a second response statement containing an emotional response statement is generated. This not only considers emotional resonance with the user, making the model-generated recommended conversation more naturally connected to the preceding context, but also enhances the content richness of the second response statement, strengthening the conversational recommendation model's ability to empathize with user needs, thereby enhancing the persuasiveness of the recommendation conversation.

[0096] In some embodiments, the above step 421 includes at least one sub-step of steps 4211 to 4213 (not shown in the figure).

[0097] Step 4211: Obtain a first vocabulary set based on at least one historical conversation sentence. The first vocabulary included in the first vocabulary set is a vocabulary in the at least one historical conversation sentence.

[0098] Retrieve vocabulary from at least one historical conversation sentence to obtain a first vocabulary set, where the first vocabulary contained in the first vocabulary set refers to the vocabulary in the at least one historical conversation sentence. The first vocabulary is a string of at least one character describing semantic information in the historical conversation sentence. Typically, the first vocabulary is a meaningful string in the historical conversation sentence. The first vocabulary can be a vocabulary describing entity information, a vocabulary describing a character's emotions, a vocabulary describing a character's actions, a vocabulary describing any object, and so on. This application does not limit the type of vocabulary.

[0099] In some embodiments, at least one historical conversation sentence is marked as D, the first vocabulary set is marked as W, w1, w2, Indicates the first word, n w Indicates the number of first words in the first vocabulary.

[0100] Step 4212: Obtain a first feature set corresponding to the first vocabulary set based on the first vocabulary set. The first feature included in the first feature set is a feature representation after encoding the first vocabulary set.

[0101] Optionally, the first feature may be a feature representation in the form of a vector, a feature representation in the form of a matrix, or a feature representation in the form of a binary value, which is not limited in this application.

[0102] In some embodiments, based on the first vocabulary set, a second vocabulary set corresponding to the first vocabulary set is obtained, and the second vocabulary contained in the second vocabulary set is a vocabulary in at least one first vocabulary used to describe entity information; the first vocabulary set and the second vocabulary set are fused using a first linear transformation to obtain a first feature set.

[0103] Entity words in the first vocabulary set are obtained to obtain a second vocabulary set, where the second words included in the second vocabulary set are words used to describe entity information in at least one of the first vocabulary sets. In other words, the second words are entity words in at least one historical conversation sentence.

[0104] In some embodiments, the second vocabulary set is labeled E, e1, e2, Indicates the second word, n e represents the number of the second vocabulary in the second vocabulary. The number of the second vocabulary is less than or equal to the number of the first vocabulary, that is, n e Less than or equal to n w .

[0105] The first linear transformation can be any linear transformation method and is not limited in this application. The first linear transformation is used to integrate each second word in the second vocabulary set into each first word in the first vocabulary set to obtain a first feature corresponding to each first word. The first feature is a feature representation of the first word after the first word is integrated with the encoded second words.

[0106] In some embodiments, the first feature set is labeled Indicates the first feature.

[0107] By adopting the first linear transformation to fuse the first vocabulary set and the second vocabulary set, each first feature in the obtained first feature set is a feature representation of the first vocabulary fused with the encoding of each second vocabulary, thereby enhancing the feature representation of the first feature relative to the first vocabulary, which is beneficial for the subsequent model to analyze the user emotions in the historical dialogue sentences according to the first feature, and improves the accuracy of the generated recommendation objects.

[0108] Step 4213: Generate an initial answer statement based on at least one historical dialogue statement and the first feature set through the dialogue recommendation model.

[0109] In some embodiments, the input data of the dialogue recommendation model includes at least one historical dialogue sentence and a first feature set, and the output data is an initial answer sentence.

[0110] In some embodiments, the process of generating an initial answer statement can be called a reply task of the dialogue recommendation model, and the task prompt of the reply task is Among them, S gen represents the random vector corresponding to the reply task, Denotes the first feature set, and D denotes at least one historical dialogue sentence. The response output of the reply task can be expressed as Right now Indicates the initial answer statement.

[0111] By obtaining the first feature set corresponding to the first vocabulary set, the dialogue recommendation model can obtain semantic features in at least one historical dialogue sentence based on the feature representation in the first feature set, thereby generating an initial dialogue sentence that meets the user's recommendation requirements. Compared with the method of directly analyzing the semantic features in the historical dialogue sentences, the difficulty of using the model is reduced and the generation efficiency of the model is improved.

[0112] In some embodiments, the above step 422 includes at least one sub-step in steps 4221 to 4222 (not shown in the figure).

[0113] Step 4221: Determine, based on the second vocabulary set, an emotional feature corresponding to at least one historical dialogue sentence.

[0114] The emotion feature corresponding to the at least one historical conversation sentence is used to characterize the user emotion associated with the at least one historical conversation sentence. The user emotion associated with the at least one historical conversation sentence may include the user emotion contained in the at least one historical conversation sentence, and may also include user emotions in other conversation sentences associated with the at least one historical conversation sentence.

[0115] Optionally, the first feature may be a feature representation in the form of a vector, a feature representation in the form of a matrix, or a feature representation in the form of a binary value, which is not limited in this application.

[0116] Step 4222: Generate a recommendation object through the dialogue recommendation model based on at least one historical dialogue sentence, emotional features, and initial answer sentence.

[0117] In some embodiments, the input data of the dialogue recommendation model includes at least one historical dialogue sentence, emotional features, and initial answer sentences, and the output data is the recommendation object.

[0118] In some embodiments, the process of generating recommended objects can be referred to as an object recommendation task of a dialogue recommendation model, and the task prompt of the object recommendation task is Among them, E′ represents the emotional feature corresponding to at least one historical dialogue sentence, S rec represents a random vector corresponding to the object recommendation task, D represents at least one historical dialogue sentence, Represents the initial answer sentence. The response output of the object recommendation task can be expressed as i r , i.e. i r Indicates the recommended object.

[0119] By determining the emotional features corresponding to at least one historical conversation sentence based on the second vocabulary set, the conversation recommendation model can analyze the user emotions in at least one historical conversation sentence based on the emotional features, thereby generating recommendation objects that meet the user's expectations and preferences. Compared with the method of directly analyzing the user emotions in the historical conversation sentences, the difficulty of using the model is reduced and the generation efficiency of the model is improved.

[0120] Next, two methods for determining the emotional features corresponding to at least one historical dialogue sentence are introduced.

[0121] Solution 1: Based on the second vocabulary set, determine the local emotional features corresponding to at least one historical dialogue sentence, where the local emotional features are used to represent the emotional features of at least one historical dialogue sentence; and determine the local emotional features as the emotional features corresponding to at least one historical dialogue sentence.

[0122] In some embodiments, the emotion feature corresponding to at least one historical conversation sentence includes a local emotion feature corresponding to at least one historical conversation sentence, that is, the emotion feature corresponding to at least one historical conversation sentence is used to characterize the user emotion contained in the at least one historical conversation sentence.

[0123] In some embodiments, the determination of the local emotion feature includes at least one sub-step in steps A1 to A4 (not shown in the figure).

[0124] Step A1: Use an emotion recognition model to perform emotion recognition on at least one historical dialogue sentence to obtain an emotion label set corresponding to the at least one historical dialogue sentence, and an emotion probability set corresponding to the emotion label set. The emotion label set contains at least one emotion label, and the emotion probability set contains the probability corresponding to the at least one emotion label.

[0125] In some embodiments, the emotion recognition model can be any publicly available large language model, such as a natural language model based on a transformer structure obtained by training with a large amount of data. The large amount of data can reach a sample level of more than 100 million, and this application does not limit this.

[0126] The emotion recognition model is used to identify user emotions in dialogue sentences and obtain an emotion tag set corresponding to each dialogue sentence. The emotion tag set contains at least one emotion tag.

[0127] In some embodiments, multiple emotion labels are pre-set, including labels representing the user's positive emotions, labels representing the user's negative emotions, and labels representing the user's neutral emotions. One or more labels representing each emotion can be set. Normally, one label representing the user's neutral emotions is set, and multiple labels representing the user's positive emotions are set. Based on the demand for emotion recognition in the actual application of the model, one or more labels representing the user's negative emotions are set. For example, nine labels can be pre-set, including labels such as "like", "curiosity", "pleasure", "gratitude", "nostalgia", "identification" and "surprise" for representing the user's positive emotions, labels such as "dissatisfaction" for representing the user's negative emotions, and labels such as "neutral" for representing the user's neutral emotions.

[0128] The emotion recognition model is used to perform emotion recognition on historical dialogue sentences, and the judgment probability of historical dialogue sentences relative to each emotion label is obtained.

[0129] In some embodiments, the emotion tag set refers to a set of multiple preset emotion tags, and the emotion probability set is a probability set corresponding to the emotion tag set. Each emotion probability is used to represent the judgment probability that the emotion recognition model recognizes the historical dialogue sentence as an emotion tag.

[0130] In some embodiments, emotion tags having an emotion probability greater than or equal to a second threshold are obtained, at least one emotion tag having an emotion probability greater than or equal to the second threshold is constructed into an emotion tag set, and emotion probabilities greater than or equal to the second threshold are constructed into an emotion probability set. The second threshold is any value between 0 and 1. For example, the second threshold may be 0.1, and emotion tags having an emotion probability greater than or equal to 0.1 are constructed into an emotion tag set, i.e., emotion probabilities greater than or equal to 0.1 are constructed into an emotion probability set.

[0131] In some embodiments, each historical conversation sentence in at least one historical conversation sentence D may be marked as Then the emotion label set can be expressed as Among them, f j represents the i-th emotion label, represents the number of emotion tags in the emotion tag set. The emotion probability set can be expressed as Among them, p j represents the probability corresponding to the i-th emotion label, represents the number of emotion probabilities in the emotion probability set, and The sizes of are the same, and i is a positive integer.

[0132] Step A2: obtaining the emotion tag set and the emotion probability set corresponding to the local entity vocabulary according to the emotion tag set and the emotion probability set, where the local entity vocabulary is a vocabulary in the second vocabulary set.

[0133] The local entity vocabulary can be understood as the second vocabulary, that is, the entity vocabulary in the historical dialogue sentence.

[0134] The sentiment tag set reflects the user's emotions towards the entity words in the historical conversation sentences, as well as the emotions and feelings towards the entity words in the previous round of historical conversation sentences. Therefore, each local entity word in the historical conversation sentences is associated with the user's emotions in the historical conversation sentences.

[0135] In some embodiments, the emotion tag set and emotion probability set corresponding to a historical conversation sentence can be assigned to the local entity vocabulary in the historical conversation sentence. For example, conversation sentence a corresponds to emotion tag set a and emotion probability set a, conversation sentence b corresponds to emotion tag set b and emotion probability set b, conversation sentence a contains local entity vocabulary 1 and local entity vocabulary 2, and conversation sentence b contains local entity vocabulary 3. Then, local entity vocabulary 1 and local entity vocabulary 2 correspond to emotion tag set a and emotion probability set a, and local entity vocabulary 3 corresponds to emotion tag set b and emotion probability set b.

[0136] Therefore, the historical dialogue sentence The jth local entity vocabulary in The emotion label set can be expressed as The emotion probability set can be expressed as j is a positive integer.

[0137] Step A3: obtaining the emotion features of the local entity vocabulary according to the emotion tag set corresponding to the local entity vocabulary and the emotion probability set corresponding to the local entity vocabulary.

[0138] The sentiment features of local entity vocabulary are used to represent the user sentiment contained in the local entity vocabulary.

[0139] In some embodiments, the initial emotion features of the local entity vocabulary are obtained based on the emotion tag set corresponding to the local entity vocabulary and the emotion probability set corresponding to the local entity vocabulary.

[0140] Adopt feature conversion model based on local entity vocabulary The emotion tag set is obtained to obtain the local entity vocabulary Feature representation of the emotion label set Optionally, the feature representation of the emotion tag set can be a feature representation in the form of a vector, a feature representation in the form of a matrix, or a feature representation in the form of a binary value. This application does not limit the feature conversion form of the feature conversion model.

[0141] Based on local entity vocabulary Feature representation of the emotion label set And the emotion probability set corresponding to the local entity vocabulary Get local entity vocabulary The initial emotional characteristics can be expressed as:

[0142]

[0143] in, Represents the jth local entity word in the historical dialogue sentence The initial emotional characteristics, p i represents the probability corresponding to the i-th emotion label, v(f i ) represents the feature representation of the i-th emotion label, Representing local entity vocabulary The number of emotion labels in the corresponding emotion label set.

[0144] In some embodiments, the first vocabulary set and the second vocabulary set are fused using a second linear transformation to obtain a second feature set, and the second features included in the second feature set are feature representations after the second vocabulary is encoded.

[0145] The second linear transformation is used to integrate each first word in the first vocabulary set into each second word in the second vocabulary set to obtain a second feature corresponding to each second word. The second feature is a feature representation of the second word after integrating the encoding of each first word.

[0146] The second linear transformation can be the same linear transformation as the first linear transformation, and the parameters of the second linear transformation are different from the parameters of the first linear transformation. Alternatively, the second linear transformation can also be a linear transformation different from the first linear transformation, which is not limited in this application.

[0147] In some embodiments, the second feature set is labeled Indicates the second feature.

[0148] In some embodiments, the initial emotion feature of the local entity vocabulary and the second feature corresponding to the local entity vocabulary are fused to obtain the emotion feature of the local entity vocabulary, and the feature dimension of the emotion feature of the local entity vocabulary is the same as the feature dimension of the second feature.

[0149] The second feature corresponding to the local entity vocabulary is the second feature of the second vocabulary corresponding to the local entity vocabulary.

[0150] Local entity vocabulary Initial emotional characteristics and local entity vocabulary The corresponding second feature Fusion to obtain local entity vocabulary The emotional characteristics can be expressed as:

[0151]

[0152] in, Represents the jth local entity word in the historical dialogue sentence emotional characteristics, Representing local entity vocabulary The second characteristic, Representing local entity vocabulary The initial emotional characteristics of Representing local entity vocabulary The second feature and local entity vocabulary The merging operation of the initial emotional features. t and b represent learnable parameters used to map the merged feature dimension back to the second feature The original feature dimension of .

[0153] Therefore, local entity vocabulary Emotional characteristics The characteristic dimension and second feature The feature dimensions are the same.

[0154] By constructing the emotional features of local entity vocabulary through the above method, the emotional features of local entity vocabulary fully integrate the user emotions in historical conversation sentences, and can more accurately represent the user emotions contained in local entity vocabulary. Therefore, when the model generates recommendation objects based on local emotional features, it can fully understand the user emotions in historical conversation sentences and improve the accuracy of recommendation objects.

[0155] Step A4: obtaining local emotion features based on the second vocabulary set, where the local emotion features include emotion features of at least one second vocabulary.

[0156] The emotion features of the second words in the second vocabulary set are combined, for example, the emotion features of at least one second word can be concatenated to obtain a local emotion feature. The local emotion feature is used to represent the emotion feature contained in at least one historical dialogue sentence.

[0157] In some embodiments, the local emotion feature can be expressed as Optionally, the local emotion feature may be a feature representation in the form of a vector, a feature representation in the form of a matrix, or a feature representation in the form of a binary value, which is not limited in this application.

[0158] Through steps A1 to A4, based on the emotional features of the second vocabulary in at least one dialogue sentence, a local emotional feature corresponding to at least one dialogue sentence is obtained, so that the local emotional feature includes the user's emotions towards the entity vocabulary in the historical dialogue sentence, as well as the emotions towards the entity vocabulary in the previous round of historical dialogue sentences. In this way, the dialogue recommendation model can analyze the user's emotions in the historical dialogue sentences based on the local emotional features, obtain recommended objects that meet the user's expectations and preferences, and improve the accuracy of the recommended objects.

[0159] Based on the above embodiment, the emotion feature E′ corresponding to at least one historical dialogue sentence includes Then the task prompt of the object recommendation task is

[0160] Solution 2: Determine local emotional features corresponding to at least one historical conversation sentence based on a second vocabulary set. Determine global emotional features corresponding to at least one historical conversation sentence based on the second vocabulary set and entity vocabulary related to the second vocabulary obtained from an external knowledge source. The global emotional features are used to represent emotional features associated with the at least one historical conversation sentence. The local emotional features and the global emotional features are determined as the emotional features corresponding to the at least one historical conversation sentence.

[0161] In some embodiments, the emotion feature corresponding to the at least one historical conversation sentence includes a local emotion feature corresponding to the at least one historical conversation sentence and a global emotion feature corresponding to the at least one historical conversation sentence. That is, the emotion feature corresponding to the at least one historical conversation sentence is used to characterize the user emotion contained in the at least one historical conversation sentence, as well as the user emotion in other conversation sentences associated with the at least one historical conversation sentence.

[0162] The global emotion feature is used to represent the emotion feature associated with at least one historical dialogue sentence. It can also be understood that the global emotion feature is used to represent the emotion feature in other dialogue sentences associated with at least one historical dialogue sentence in an external knowledge source.

[0163] In some embodiments, the entity vocabulary related to the second vocabulary obtained from the external knowledge source refers to entity vocabulary in other conversation sentences related to the second vocabulary obtained from the comment database.

[0164] By adding global emotional features to the emotional features corresponding to at least one historical dialogue sentence, and taking into account the entity vocabulary related to the second vocabulary in the external knowledge source, the emotional features corresponding to at least one historical dialogue sentence are made more comprehensive, the influence of uncertain emotions in local emotional features is reduced, user preferences are better explored, and the accuracy of recommended objects is improved.

[0165] In some embodiments, the determination of the global emotion feature includes at least one sub-step in steps B1 to B4 (not shown in the figure).

[0166] Step B1: Acquire at least one global entity vocabulary related to a local entity vocabulary from an external knowledge source. The global entity vocabulary is an entity vocabulary that appears in the same dialogue sentence as the local entity vocabulary.

[0167] In some embodiments, a global entity vocabulary is an entity vocabulary that appears in the same conversation sentence as a local entity vocabulary.

[0168] In some embodiments, the global entity vocabulary is an entity vocabulary whose highest probability sentiment label is the same as the highest probability sentiment label of the local entity vocabulary.

[0169] In some embodiments, the global entity vocabulary is an entity vocabulary whose highest probability sentiment tags intersect with the highest probability sentiment tag of the local entity vocabulary.

[0170] In some embodiments, the global entity vocabulary is an entity vocabulary whose highest probability sentiment tag intersects with several highest probability sentiment tags of the local entity vocabulary.

[0171] In some embodiments, the external knowledge source is related to the local entity vocabulary The relevant global entity vocabulary can be expressed as k is a positive integer. k represents the global entity vocabulary, Representation and Local Entity Vocabulary The number of related global entity words.

[0172] Step B2: Calculate the co-occurrence probability corresponding to at least one global entity vocabulary. The co-occurrence probability refers to the probability that a global entity vocabulary and a local entity vocabulary appear in the same dialogue sentence.

[0173] According to the number of review sentences in the review database and the global entity vocabulary e k and local entity vocabulary The number of times it appears in the same conversation sentence is calculated for the local entity vocabulary Global entity vocabulary k The corresponding co-occurrence probability P(e k|e j ).

[0174] Step B3: obtaining the global sentiment feature of the local entity vocabulary according to the co-occurrence probabilities corresponding to the local entity vocabulary and at least one global entity vocabulary.

[0175] Based on local entity vocabulary and for local entity vocabulary At least one global entity vocabulary e k Corresponding co-occurrence probability, get the local entity vocabulary The global sentiment characteristics can be expressed as:

[0176]

[0177] in, Representing local entity vocabulary The global emotional characteristics, v(e k ) is the global entity vocabulary e k The characteristic representation of P(e k |e j ) represents the global entity vocabulary e k and local entity vocabulary The probability of appearing in the same dialogue sentence. Global entity vocabulary e k The feature representation v(e k ) can be obtained using the feature transformation model.

[0178] Step B4: obtaining a global emotion feature based on the second vocabulary set, where the global emotion feature includes a global emotion feature of at least one second vocabulary.

[0179] The global emotion features of the second words in the second vocabulary set are combined, for example, the global emotion features of at least one second word can be concatenated to obtain a global emotion feature. The global emotion feature is used to represent the emotion feature associated with at least one historical conversation sentence.

[0180] In some embodiments, the global sentiment feature can be expressed as Optionally, the global emotion feature may be a feature representation in the form of a vector, a feature representation in the form of a matrix, or a feature representation in the form of a binary value, which is not limited in this application.

[0181] Through steps B1 to B4, the global emotion feature corresponding to at least one dialogue sentence is obtained based on the global emotion feature of the second word in at least one dialogue sentence, so that the global emotion feature includes the user emotion of at least one global entity word related to the local entity word. Therefore, the dialogue recommendation model analyzes the user emotion corresponding to the historical dialogue sentences based on the local emotion feature and the global emotion feature, and can generate recommendation objects based on the user expectations and preferences and the associated emotions in the big data, so that the generated recommendation objects can simultaneously meet the user needs and the big data associated needs.

[0182] Based on the above embodiment, the emotion feature E′ corresponding to at least one historical dialogue sentence includes and E′ g , then the task prompt of the object recommendation task is

[0183] In some embodiments, the above step 430 includes at least one sub-step of steps 431 to 433 (not shown in the figure).

[0184] Step 431: Acquire at least one triple related to the recommended object from an external knowledge source. The triple includes a head entity, a tail entity, and a relationship type connecting the head entity and the tail entity. The head entity is used to indicate the recommended object.

[0185] In some embodiments, the recommended object i r At least one triplet can be represented as Among them, i r Indicates the head entity, e s Represents the tail entity, l represents the type of relationship connecting the head entity and the tail entity, and s is a positive integer.

[0186] A triple represents a head entity, a tail entity, and the relationship type connecting the head and tail entities. The head entity is the recommended object. The relationship type connecting the recommended object and the tail entity can contain one or more relationship types, and each relationship type can be connected to one or more tail entities. If a relationship type connects multiple tail entities, the relationship type also corresponds to the triples corresponding to the multiple tail entities.

[0187] For example, if the head entity is the title of a movie, l can be the movie director, and the tail entity is the director's name. This relationship type is connected to a tail entity, and the corresponding triple is represented as <movie name, director, director name>. If l is a movie actor, the tail entity is the name of the actor appearing in the movie. This relationship type can be connected to multiple tail entities, that is, movie actor can correspond to multiple actor names, and l can correspond to multiple triples corresponding to actor names, for example, <movie name, actor, actor A's name>, <movie name, actor, actor B's name>, <movie name, actor, actor C's name>, and so on.

[0188] In some embodiments, obtaining at least one triple related to the recommended object from an external knowledge source refers to obtaining at least one triple related to the recommended object from a knowledge graph.

[0189] Step 432: Obtain at least one comment sentence for the recommended object from an external knowledge source, select entity words whose occurrence times are greater than a first threshold from the at least one comment sentence, and obtain a set of associated words.

[0190] In some embodiments, obtaining at least one comment statement for the recommended object from an external knowledge source refers to obtaining at least one comment statement for the recommended object from a comment database.

[0191] The at least one comment statement regarding the recommended object refers to a comment statement published by each user regarding the recommended object. This application does not limit the value of the first threshold. For example, the first threshold can be 2, meaning that entity words that appear more than twice in at least one comment statement are selected as associated words to obtain an associated word set. For example, if entity word 1 appears in both comment statement A and comment statement B, entity word 1 is added to the associated word set.

[0192] In some embodiments, the associated vocabulary set can be represented as t is a positive integer.

[0193] Step 433: Generate a second answer statement through the dialogue recommendation model based on at least one triple, at least one associated word in the associated word set, the recommended object, and the first answer statement.

[0194] In some embodiments, the input data of the dialogue recommendation model includes at least one triple, at least one associated word in the associated word set, a recommendation object and an initial answer sentence, and the output data is a second answer sentence.

[0195] In some embodiments, the process of generating recommendation objects can be regarded as the sentiment alignment task of the dialogue recommendation model, and the task prompt of the sentiment alignment task is in, A sequence of words representing at least one triple transformation, A word sequence representing at least one associated word transformation, A word sequence representing the conversion of the recommended object. represents the initial answer sentence. The response output of the sentiment alignment task can be expressed as Right now Indicates the second answer statement.

[0196] By obtaining at least one triple related to the recommended object from an external knowledge source, and obtaining a set of associated vocabulary for the recommended object, a response consistent with human emotions can be generated based on the triple and the set of associated vocabulary, thereby enhancing the emotional resonance with the user and thus enhancing the persuasiveness of the recommendation dialogue.

[0197] In some embodiments, a second answer statement including an emotional answer statement is generated through a dialogue recommendation model based on at least one triple, at least one associated word in an associated word set, a recommendation object, and a first answer statement.

[0198] In some embodiments, the input data of the dialogue recommendation model includes at least one triple, at least one related word in the related word set, a recommended object, and an initial answer sentence, and the output data is a second answer sentence. The second answer sentence includes an emotional answer sentence.

[0199] In some embodiments, through a dialogue recommendation model, an emotional answer statement is generated based on at least one triple, at least one related word in the related word set, the recommended object and the first answer statement; the first answer statement and the emotional answer statement are spliced ​​to obtain a second answer statement.

[0200] In some embodiments, the input data of the conversation recommendation model includes at least one triple, at least one related word from a related vocabulary set, a recommended object, and an initial answer sentence, and the output data is an emotional answer sentence. The first answer sentence and the emotional answer sentence are concatenated to generate a second answer sentence.

[0201] According to the different functions of the dialogue recommendation model, different schemes for generating the second answer statement are adopted, so that the second answer statement containing emotional answer statements can be obtained in the end, which improves the content richness of the second answer statement and enhances the persuasiveness of the recommendation dialogue.

[0202] The diagram of the dialogue recommendation process provided by this application can be referred to Figure 5As shown, the dialogue recommendation process includes three subtasks: the reply task, the object recommendation task, and the sentiment alignment task. The task prompt for the reply task includes a first feature set, a random vector corresponding to the reply task, and at least one historical dialogue sentence. The response output of the reply task is the initial answer sentence. The response output of the reply task is used as a task prompt for the subsequent object recommendation task and sentiment alignment task. The task prompt for the object recommendation task includes local sentiment features, global sentiment features, a random vector corresponding to the object recommendation task, at least one historical dialogue sentence, and the response output of the reply task (initial answer sentence). The response output of the object recommendation task is the recommended object. The task prompt for the sentiment alignment task includes at least one triple-converted word sequence, at least one associated word sequence, the response output of the reply task (initial answer sentence), and the response output of the object recommendation task (recommended object). The response output of the sentiment alignment task is the second answer sentence.

[0203] Figure 5 The dialogue recommendation process shown is for Figure 5 In the example of the conversation recommendation based on the historical conversation data shown, the generated initial answer is "Me too. You should watch the comedy [mask]." The generated recommendation object is "XXXX" (movie name), and the generated second answer is "I've seen XXXX several times. It never gets old! I'm a big fan of XXXX."

[0204] Please refer to Figure 6 , which shows a flowchart of a method for training a conversation recommendation model provided by one embodiment of the present application. The execution entity of each step of the method can be a computer device. The method can include at least one of the following steps 610 to 650.

[0205] Step 610: Obtain a basic data set for training a dialogue recommendation model, where the basic data set includes at least one data group, each data group includes at least one sample dialogue sentence, a first label sentence corresponding to at least one sample dialogue sentence, and a second label sentence corresponding to at least one sample dialogue sentence, wherein the first label sentence includes a label object for at least one sample dialogue sentence, and the second label sentence includes the first label sentence and an emotional response sentence for the label object.

[0206] Step 620: Generate a first training sample corresponding to a reply task based on at least one data group. The reply task uses an initial label sentence as label data, and determines sample data based on at least one sample dialogue sentence. The initial label sentence contains a blank object for at least one sample dialogue sentence.

[0207] In some embodiments, a first vocabulary set is obtained based on at least one sample conversation sentence, wherein a first vocabulary contained in the first vocabulary set is a vocabulary in the at least one sample conversation sentence. A first feature set corresponding to the first vocabulary set is obtained based on the first vocabulary set, wherein a first feature contained in the first feature set is a feature representation after encoding the corresponding first vocabulary. A first training sample corresponding to a reply task is generated based on the at least one sample conversation sentence, the first feature set, a random vector, and an initial label sentence, wherein the reply task uses the initial label sentence as label data and uses the at least one sample conversation sentence, the first feature set, and the random vector as sample data.

[0208] Step 630: Generate a second training sample corresponding to the object recommendation task based on at least one data group, where the object recommendation task uses the label object as label data, and determines the sample data based on at least one sample dialogue sentence and the first label sentence.

[0209] In some embodiments, a second vocabulary set is obtained based on at least one sample conversation sentence, where second vocabulary included in the second vocabulary set is vocabulary used to describe entity information in the at least one sample conversation sentence. Based on the second vocabulary set, an emotional feature corresponding to the at least one sample conversation sentence is determined. A second training sample corresponding to an object recommendation task is generated based on the at least one sample conversation sentence, the emotional feature, the random vector, the initial label sentence, and the label object, wherein the object recommendation task uses the second label sentence as label data and uses the at least one sample conversation sentence, the emotional feature, the random vector, and the initial label sentence as sample data.

[0210] Step 640: Generate a third training sample corresponding to the emotion alignment task based on at least one data group. The emotion alignment task uses the second label sentence as label data, and determines sample data based on at least one sample dialogue sentence and the first label sentence.

[0211] In some embodiments, at least one triple related to the label object is obtained from an external knowledge source, the triple including a head entity, a tail entity, and a relationship type connecting the head entity and the tail entity, the head entity is used to indicate the label object, the tail entity is obtained from a comment statement for the label object, and the relationship type is obtained from the external knowledge source.

[0212] In some embodiments, the tag object i r′ At least one triplet can be represented as Among them, i r′ Indicates the head entity, e s′ Represents the tail entity, l represents the type of relationship connecting the head entity and the tail entity, and s′ is a positive integer.

[0213] In some embodiments, based on at least one comment statement for a label object in a comment database, at least one entity word is obtained from the comment statement, the entity word is used as the tail entity, and the type of relationship between the head entity and the tail entity is obtained from an external knowledge source to obtain at least one triple.

[0214] In some embodiments, based on at least one comment statement about a tagged object in a comment database, at least one entity word is retrieved from the comment statement, the entity word is used as the tail entity, and the type of relationship between the head entity and the tail entity is obtained from an external knowledge source to obtain at least one triple. Furthermore, at least one triple related to the tagged object is obtained from a knowledge graph, thereby obtaining the at least one triple mentioned above.

[0215] In some embodiments, based on at least one comment statement for a tag object in a comment database, at least one entity word is obtained from the comment statement, the entity word is used as a relationship type, and at least one triple corresponding to the relationship type is obtained from an external knowledge source.

[0216] In some embodiments, based on at least one comment statement regarding a tagged object in a comment database, at least one entity word is retrieved from the comment statement, the entity word is used as a relationship type, and at least one triple corresponding to the relationship type is retrieved from an external knowledge source. Furthermore, at least one triple related to the tagged object is retrieved from a knowledge graph, thereby obtaining the at least one triple mentioned above.

[0217] In some embodiments, a second vocabulary set is obtained based on at least one sample dialogue sentence, and the second vocabulary included in the second vocabulary set is a vocabulary used to describe entity information in the at least one sample dialogue sentence.

[0218] In some embodiments, a third training sample corresponding to the sentiment alignment task is generated based on at least one triple, a second vocabulary set, a first label sentence, and a second label sentence, wherein the sentiment alignment task uses the second label sentence as label data, and uses at least one triple, a second vocabulary set, an initial label sentence, and a label object as sample data.

[0219] The task prompts for the emotion alignment task are in, A sequence of words representing at least one triple transformation, a word sequence representing at least one second lexical transformation, The word sequence representing the conversion of the label object, Represents an initial labeled statement.

[0220] The data acquisition process of steps 620 to 640 may refer to the above embodiment and will not be described in detail here.

[0221] Step 650 : Use the first training sample, the second training sample, and the third training sample to train the dialogue recommendation model to obtain a trained dialogue recommendation model.

[0222] In some embodiments, the above step 650 includes at least one sub-step of steps 651 to 653 (not shown in the figure).

[0223] Step 651: Adjust the parameters of the dialogue recommendation model according to the first loss function value corresponding to the reply task to obtain the dialogue recommendation model trained for the reply task.

[0224] In some embodiments, the first loss function is a cross-entropy loss function.

[0225] In some embodiments, a first loss function value is obtained based on the difference between the vocabulary at each position of the first predicted sentence obtained from the reply task and the initial label sentence.

[0226] The first loss function value can be expressed as:

[0227]

[0228] in, Indicates the task prompt for a given reply task In the case of the first predicted sentence, the qth word w q , generated as the probability of the corresponding vocabulary in the initial label sentence, |w q | represents the number of words in the second predicted sentence, and q is a positive integer.

[0229] For example, if the corresponding word in the initial label sentence is "comedy", then represents the qth word w in the first predicted sentence q , generating the probability of “comedy”.

[0230] Step 652 : Adjust the parameters of the dialogue recommendation model trained for the reply task according to the second loss function value corresponding to the object recommendation task to obtain the dialogue recommendation model trained for the object recommendation task.

[0231] In some embodiments, the second loss function is a cross-entropy loss function.

[0232] In some embodiments, a second loss function value is obtained based on the difference between the predicted object and the labeled object obtained from the object recommendation task.

[0233] In some embodiments, a second loss function value is calculated based on the predicted probability of at least one predicted object obtained from the object recommendation task.

[0234] The second loss function value can be expressed as:

[0235]

[0236] Among them, Pr(i u |C′ rec ) indicates the task prompt for a given reply task In the case of u The probability of , N represents the number of predicted objects, and u is a positive integer.

[0237] In some embodiments, based on the feedback label of at least one prediction object, a weight value corresponding to at least one prediction object is obtained, where the feedback label is a pre-labeled label used to indicate each user's preference for the prediction object; based on the prediction probability of at least one prediction object and the weight value corresponding to at least one prediction object, a second loss function value is calculated.

[0238] In some embodiments, the preference tags for the predicted object in the big data can be used as feedback tags for the predicted object. For example, preference tags in the big data that are greater than a first ratio can be used as feedback tags for the predicted object. This application does not limit the first ratio; for example, the first ratio can be 50%.

[0239] In some embodiments, the feedback labels include "like", "dislike" and "no opinion expressed". The weight value corresponding to "like" can be set to 1, the weight value corresponding to "dislike" can be set to 0, and the weight value corresponding to "no opinion expressed" can be set to 0.5.

[0240] In some embodiments, the weight value corresponding to "like" is set to 1, and the weight values ​​corresponding to "dislike" and "no position expressed" are both set to 0.5.

[0241] This application does not limit the mapping relationship between feedback labels and corresponding weight values, which can be set according to actual conversation recommendation needs.

[0242] The second loss function value can be expressed as:

[0243]

[0244] in, Represents the predicted object i u The corresponding weight value, Pr(i u |C′ rec ) indicates the task prompt for a given reply task In the case of u The probability of N is the number of predicted objects.

[0245] By weighting the predicted objects, the feedback labels of the big data for the predicted objects are mapped to weight values, so that the trained dialogue recommendation model can filter out recommended objects with negative feedback labels and recommend objects with positive feedback labels to users, so that the recommended objects can meet the user's preferences.

[0246] Step 653 : Adjust the parameters of the conversation recommendation model trained for the object recommendation task according to the third loss function value corresponding to the emotion alignment task to obtain a trained conversation recommendation model.

[0247] In some embodiments, the third loss function is a cross-entropy loss function.

[0248] In some embodiments, a third loss function value is obtained based on the difference between the vocabulary at each position of the second predicted sentence and the second label sentence obtained in the sentiment alignment task.

[0249] The third loss function value can be expressed as:

[0250]

[0251] in, Indicates the task prompt for a given reply task In the case of the second predicted sentence, the hth word w h , generated as the probability of the corresponding word in the second label sentence, |w h | represents the number of words in the second predicted sentence, and h is a positive integer.

[0252] The first loss function value, second loss function value, and third loss function value obtained by the above method are used to adjust the parameters of the dialogue recommendation model in turn, which is more in line with the usage process of the dialogue recommendation model. The goal is to minimize the loss function value, so that the trained dialogue recommendation model can better meet user needs and maximize the accuracy of the recommended dialogues of the dialogue recommendation model.

[0253] The conversation recommendation method and conversation recommendation model training method provided in the embodiments of the present application are corresponding model training processes and usage processes. For details not described in detail on one side, please refer to the introduction and description on the other side.

[0254] Below, the performance of the dialogue recommendation model provided in the embodiment of the present application and the models in the related technology are tested and compared.

[0255] 1. Evaluation Indicators

[0256] In this application, the ECR evaluation metrics include both subjective and objective evaluation metrics. In particular, considering the user's emotional experience in the actual conversation recommendation environment, the following evaluation metrics are designed:

[0257] (1) Objective indicators: For the evaluation of recommendation tasks, the recall rate Recall@m (R@m, m takes values ​​of 1, 10, 50) can be used. In order to measure the accuracy of the model in predicting user needs and filter out the negative impact of incorrect recommendations in the data set, this application also introduces the Recall_True@m (RT@m, m takes values ​​of 1, 10, 50) indicator. This indicator further refines Recall@m and only considers products that receive user "like" feedback as the correct answer. In addition, this application also adds the AUC (Area Under Curve, the area under the curve and the coordinate axis) indicator. The AUC indicator is used to indicate the proportion of correctly predicted positive emotions being greater than negative emotions, in order to emphasize the ranking accuracy of products that receive positive and negative user feedback.

[0258] (2) Subjective Indicators: When evaluating the generation effect, since the ECR model has deviated from the standard answers provided by the dataset, we use five subjective indicators to measure model performance: emotional intensity, emotional persuasiveness, logical persuasiveness, information richness, and human-likeness. (a) Emotional intensity is responsible for evaluating the strength of the emotions conveyed by the model to the user. (b) Emotional persuasiveness is used to measure the model's ability to influence the user's emotions and persuade the user. (c) Logical persuasiveness evaluates the efficiency of the model's use of logical reasoning to influence user decision-making. (d) Information richness considers the amount of information provided by the model. (e) Human-likeness evaluates how close the technical response is to natural human communication.

[0259] These metrics are scored on a scale of [0, 9]. We use a large language model-based scoring technique, which scores samples based on carefully designed prompts. We randomly select 200 samples for evaluation. Any large language model can be used for subjective scoring. However, given the instability of the large language model's output, we also hired two human reviewers to provide scoring. Each reviewer is responsible for evaluating 100 samples, and their results are compared with those of the large language model-based scoring technique.

[0260] 2. Comparison Method

[0261] In the evaluation of the recommendation task, in addition to the ECR basic model UniCRS, this application also selected several other dialogue recommendation models in related technologies, including:

[0262] (1) The KBRD (Towards Knowledge-Based Recommender Dialog System) model introduces knowledge graphs for the first time and proposes an end-to-end dialogue recommendation framework, which uses Transformer to generate responses with enhanced recommendation content.

[0263] (2) KGSF (Improving Conversational Recommender Systems via KnowledgeGraph based Semantic Fusion) model, which enhances the ability to mine user preferences by integrating vocabulary and entity-oriented knowledge graphs, and adopts a knowledge-enhanced Transformer framework for response generation.

[0264] (3) RevCore (Review-augmented Conversational Recommendation) model, which introduces review data, optimizes recommendation and response generation through additional product review information, and places special emphasis on the sentiment-based review screening process.

[0265] (4) UCCR (User-Centric Conversational Recommendation) model, which conducts comprehensive user modeling around the user's real needs, focusing on utilizing current and historical discourse information and similar user data.

[0266] In the response generation task, the model presented in this application was primarily compared against large language models that do not rely on benchmark datasets. These comparison models included the base model UniCRS, Large Language Model 1, and two callable large language models: Large Language Model 3 and Large Language Model 4. To ensure controllable output from the comparison models, the temperature parameter was set to 0 when calling the API (Application Programming Interface). These large language models provide users with conversational recommendations through carefully designed prompts. They are aware of utterance history to ensure fair evaluation and are required to recommend products predicted by the ECR recommendation module.

[0267] 3. Dataset

[0268] This application uses the benchmark REDIAL dataset, which is widely used in the field of dialogue recommendation technology, to test the model's effectiveness. REDIAL is an English dataset that was manually constructed through crowdsourcing on a network platform. The dataset contains conversations between users and recommenders about movie recommendations. It contains 10,006 utterances, a total of 182,150 utterance sentences, involving 51,699 movies and 64,362 entities. The user feedback on the products in the dataset is divided into three types: "like", "dislike" and "no statement". This application uses the knowledge graph to extract the entities mentioned in each dialogue sentence for utilization.

[0269] IV. Performance Analysis of Recommended Tasks

[0270] The performance evaluation results of the model proposed in this application on the object recommendation task are shown in Table 1 below:

[0271] Table 1 Performance evaluation results of object recommendation tasks

[0272]

[0273]

[0274] The bold text represents the optimal result. * indicates significant improvement (t-test, p<0.05).

[0275] By analyzing the performance of the baseline models, we can find that by incorporating more sufficient external knowledge sources into the conversational recommendation process, KGSF and RevCore achieve more accurate recommendation results compared to KBRD, which proves the value of incorporating richer knowledge into conversational recommendation technology. By analyzing the user's multiple interaction histories, UCCR is able to extract comprehensive user-centric data and demonstrates excellent performance in the RT@50 and R@50 indicators. At the same time, UniCRS achieves the best performance among all the compared baseline models by integrating the knowledge recorded in the pre-trained language model parameters into the conversational recommendation task. However, these models still have shortcomings in distinguishing user feedback on recommended products, which is reflected in their AUC values ​​being close to 0.5.

[0276] Among all the baseline models compared, ECR demonstrated significant advantages. Specifically, ECR achieved improvements of 3.9% and 3.0% in RT@10 and RT@50 over the top-performing model, UniCRS, respectively. Furthermore, ECR significantly outperformed UniCRS in ranking recommended items, achieving a 6.9% improvement in the AUC. This result demonstrates that mining user sentiment can enhance a model's ability to accurately capture user preferences.

[0277] 5. Generation Task Performance Analysis

[0278] To evaluate ECR's ability to express emotions, a large language model and human reviewers conducted subjective evaluations on the response generation task. The specific results are shown in Table 2 below:

[0279] Table 2 Performance evaluation results of response generation task

[0280]

[0281]

[0282] The bold text represents the optimal result. * indicates significant improvement (t-test, p<0.05).

[0283] The results show that the scores of the large language model and human reviewers are generally similar. Among all baseline comparison models, the zero-shot large language model significantly outperforms UniCRS, which is supervised learning using the full amount of data from the REDIAL dataset. This suggests that the quality of the standard responses provided in the dataset is unsatisfactory. Furthermore, Large Language Model 1, having undergone instruction fine-tuning on high-quality chat data, is more suitable for the dialogue recommendation task and thus performs the best.

[0284] For the model ECR proposed in this application, it can be seen from the table that ECR[Large Language Model 1] outperforms all the comparison models in most evaluation indicators. At the same time, the performance of ECR[Large Language Model 1] is on par with that of Large Language Model 4, although the number of parameters is very different. In the scoring results given by the large language model, ECR has significantly improved in terms of emotional expression. ECR[Large Language Model 1] and ECR[Large Language Model 2] have improved by 59.6% and 10.1% respectively compared with Large Language Model 1. This shows that after fine-tuning with supervision using emotional comment data, ECR has gained stronger emotional expression. Similarly, ECR[Large Language Model 1] and ECR[Large Language Model 2] have 10.7% and 2.9% optimizations in terms of emotional persuasiveness compared with Large Language Model 1 and Large Language Model 4 respectively. These findings all prove that ECR greatly optimizes the user experience in the emotional dimension. At the same time, ECR improves response generation by incorporating relevant knowledge into generation prompts. This is verified by the information richness metric, where ECR[Large Language Model 1] and ECR[Large Language Model 2] achieve improvements of 9.1% and 1.7%, respectively, compared to the comparison models, Large Language Model 1 and Large Language Model 4. Finally, in terms of human-likeness, ECR[Large Language Model 1] achieves a 4.59% improvement over Large Language Model 1. Taken together, these results demonstrate that the proposed ECR framework can bring technology closer to human-likeness, thereby enhancing user emotional experience and improving satisfaction.

[0285] Given the wide range of ECR ​​scores and the variability in absolute scores across metrics, there's no objective, absolute evaluation standard. Therefore, we calculated the Cohen's kappa coefficient by ranking the models. The inter-rater agreement score was 0.73, indicating high consistency. However, the average agreement between the human raters and the large language model was 0.49, indicating only moderate consistency. Analysis revealed significant differences between human and large language model scores on emotional persuasion, human-likeness, and logical persuasion. Emotional persuasion had the lowest agreement, at only 0.43. Compared to the large language model-based evaluation, the human rater results showed that ECR [Large Language Model 2] outperformed Large Language Model 1 by 29.9% in emotional persuasion and significantly surpassed Large Language Model 4 in human-likeness by 16.2%. ECR [Large Language Model 1] also outperformed Large Language Model 1 in logical persuasion. This is because large language models tend to generate repetitive introductory content, causing reviewers to feel bored after evaluating a large number of examples. This is also a problem that real users may encounter in conversational recommendation tasks. Therefore, the evaluation results of human reviewers better reflect user satisfaction in real conversational recommendation tasks. At the same time, human reviewers tend to prefer responses that include personalized experiences and opinions, believing these to be more human. However, large language models are more likely to generate objective content, such as plot summaries. These results show that large language models also differ from human reviewers when performing subjective tasks, further confirming the importance of the proposed empathy-centered conversational recommendation framework in meeting user needs in real scenarios.

[0286] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0287] Please refer to Figure 7 , which shows a block diagram of a conversation recommendation device provided by an embodiment of the present application. The device has the function of implementing the above-mentioned conversation recommendation method, and the function can be implemented by hardware or by hardware executing corresponding software. The device can be the computer device described above, or it can be set in a computer device. Figure 7 As shown, the apparatus 700 may include: a statement acquisition module 710 , a first generation module 720 and a second generation module 730 .

[0288] The sentence acquisition module 710 is used to acquire at least one historical conversation sentence.

[0289] The first generating module 720 is configured to generate a first answer statement based on the at least one historical dialogue statement through a dialogue recommendation model, wherein the first answer statement includes a recommendation object for the at least one historical dialogue statement.

[0290] The second generation module 730 is used to generate a second answer statement based on the first answer statement and entity vocabulary related to the recommendation object obtained from an external knowledge source through the dialogue recommendation model, wherein the second answer statement includes the first answer statement and an emotional answer statement for the recommendation object.

[0291] In some embodiments, the second generating module 730 is configured to:

[0292] Acquire at least one triple related to the recommended object from the external knowledge source, wherein the triple includes a head entity, a tail entity, and a relationship type connecting the head entity and the tail entity, wherein the head entity is used to indicate the recommended object;

[0293] Acquiring at least one comment sentence for the recommended object from the external knowledge source, selecting entity words with a number of occurrences greater than a first threshold from the at least one comment sentence, and obtaining a set of associated words;

[0294] The second answer sentence is generated by the dialogue recommendation model according to the at least one triple, at least one associated word in the associated word set, the recommendation object and the first answer sentence.

[0295] In some embodiments, the second generating module 730 is configured to:

[0296] generating the emotional answer statement using the dialogue recommendation model based on the at least one triple, at least one associated word in the associated word set, the recommended object, and the first answer statement; and concatenating the first answer statement and the emotional answer statement to obtain the second answer statement;

[0297] or,

[0298] The second answer statement including the emotional answer statement is generated through the dialogue recommendation model according to the at least one triple, at least one associated word in the associated word set, the recommendation object and the first answer statement.

[0299] In some embodiments, the first generating module 720 includes:

[0300] An initial generation unit is configured to generate an initial answer sentence based on the at least one historical dialogue sentence using the dialogue recommendation model, wherein the initial answer sentence includes a blank object for the at least one historical dialogue sentence.

[0301] An object generation unit is used to generate the recommended object according to the at least one historical dialogue sentence and the initial answer sentence through the dialogue recommendation model.

[0302] The first generating unit is configured to fill the blank object in the initial answer sentence with the recommended object to obtain the first answer sentence.

[0303] In some embodiments, the initial generation unit is configured to:

[0304] obtaining a first vocabulary set according to the at least one historical conversation sentence, wherein a first vocabulary contained in the first vocabulary set is a vocabulary in the at least one historical conversation sentence;

[0305] Obtaining, based on the first vocabulary set, a first feature set corresponding to the first vocabulary set, wherein a first feature included in the first feature set is a feature representation of the first vocabulary after encoding;

[0306] The initial answer sentence is generated by the dialogue recommendation model according to the at least one historical dialogue sentence and the first feature set.

[0307] In some embodiments, the initial generation unit is configured to:

[0308] Obtaining, based on the first vocabulary set, a second vocabulary set corresponding to the first vocabulary set, wherein second words included in the second vocabulary set are at least one word in the first vocabulary set used to describe entity information;

[0309] The first vocabulary set and the second vocabulary set are fused using a first linear transformation to obtain the first feature set.

[0310] In some embodiments, the object generation unit is configured to:

[0311] determining, based on the second vocabulary set, an emotional feature corresponding to the at least one historical conversation sentence;

[0312] The recommended object is generated through the dialogue recommendation model according to the at least one historical dialogue sentence, the emotional feature and the initial answer sentence.

[0313] In some embodiments, the object generation unit is configured to:

[0314] determining, based on the second vocabulary set, a local emotion feature corresponding to the at least one historical conversation sentence, the local emotion feature being used to represent an emotion feature of the at least one historical conversation sentence; and determining the local emotion feature as the emotion feature corresponding to the at least one historical conversation sentence;

[0315] or,

[0316] Based on the second vocabulary set, local emotional features corresponding to the at least one historical dialogue sentence are determined; based on the second vocabulary set and entity vocabulary related to the second vocabulary obtained from the external knowledge source, global emotional features corresponding to the at least one historical dialogue sentence are determined, wherein the global emotional features are used to represent emotional features associated with the at least one historical dialogue sentence; the local emotional features and the global emotional features are determined as emotional features corresponding to the at least one historical dialogue sentence.

[0317] In some embodiments, the object generation unit is configured to:

[0318] Performing emotion recognition on the at least one historical conversation sentence using an emotion recognition model to obtain an emotion label set corresponding to each of the at least one historical conversation sentence, and an emotion probability set corresponding to the emotion label set, wherein the emotion label set includes at least one emotion label, and the emotion probability set includes probabilities corresponding to each of the at least one emotion label;

[0319] Obtaining, based on the emotion tag set and the emotion probability set, an emotion tag set corresponding to a local entity vocabulary and an emotion probability set corresponding to the local entity vocabulary, wherein the local entity vocabulary is a vocabulary in the second vocabulary set;

[0320] Obtaining the emotional features of the local entity vocabulary according to the emotional tag set corresponding to the local entity vocabulary and the emotional probability set corresponding to the local entity vocabulary;

[0321] The local emotion feature is obtained according to the second vocabulary set, and the local emotion feature includes an emotion feature of at least one of the second vocabulary.

[0322] In some embodiments, the object generation unit is configured to:

[0323] Obtaining initial emotion features of the local entity vocabulary according to the emotion tag set corresponding to the local entity vocabulary and the emotion probability set corresponding to the local entity vocabulary;

[0324] fusing the first vocabulary set and the second vocabulary set using a second linear transformation to obtain a second feature set, wherein second features included in the second feature set are feature representations of the second vocabulary after encoding;

[0325] The initial emotion feature of the local entity vocabulary and the second feature corresponding to the local entity vocabulary are fused to obtain the emotion feature of the local entity vocabulary, and the feature dimension of the emotion feature of the local entity vocabulary is the same as the feature dimension of the second feature.

[0326] In some embodiments, the object generation unit is configured to:

[0327] Acquire at least one global entity vocabulary related to the local entity vocabulary from the external knowledge source, wherein the global entity vocabulary is an entity vocabulary that appears in the same dialogue sentence as the local entity vocabulary;

[0328] Calculating a co-occurrence probability corresponding to each of the at least one global entity vocabulary, wherein the co-occurrence probability refers to a probability that the global entity vocabulary and the local entity vocabulary appear in the same dialogue sentence;

[0329] Obtaining a global sentiment feature of the local entity vocabulary according to the co-occurrence probabilities corresponding to the local entity vocabulary and the at least one global entity vocabulary;

[0330] The global emotion feature is obtained according to the second vocabulary set, and the global emotion feature includes a global emotion feature of at least one of the second vocabulary.

[0331] Please refer to Figure 8 , which shows a block diagram of a training device for a conversation recommendation model provided by an embodiment of the present application. The device has the function of implementing the training method of the conversation recommendation model described above, and the function can be implemented by hardware or by hardware executing corresponding software. The device can be the computer device described above, or it can be set in a computer device. Figure 8 As shown, the apparatus 800 may include: a data acquisition module 810 , a first generation module 820 , a second generation module 830 , a third generation module 840 and a training module 850 .

[0332] The data acquisition module 810 is used to obtain a basic data set for training the dialogue recommendation model, where the basic data set includes at least one data group, each of which includes at least one sample dialogue sentence, a first label sentence corresponding to the at least one sample dialogue sentence, and a second label sentence corresponding to the at least one sample dialogue sentence, wherein the first label sentence includes a label object for the at least one sample dialogue sentence, and the second label sentence includes the first label sentence and an emotional response sentence for the label object.

[0333] The first generation module 820 is used to generate a first training sample corresponding to a reply task based on the at least one data group, wherein the reply task uses an initial label sentence as label data and determines sample data based on the at least one sample dialogue sentence, wherein the initial label sentence contains a blank object for the at least one sample dialogue sentence.

[0334] The second generation module 830 is used to generate a second training sample corresponding to the object recommendation task based on the at least one data group, where the object recommendation task uses the label object as label data and determines the sample data based on the at least one sample dialogue sentence and the first label sentence.

[0335] The third generation module 840 is used to generate a third training sample corresponding to the emotion alignment task based on the at least one data group, where the emotion alignment task uses the second label sentence as label data and determines sample data based on the at least one sample dialogue sentence and the first label sentence.

[0336] The training module 850 is configured to train the dialogue recommendation model using the first training sample, the second training sample, and the third training sample to obtain a trained dialogue recommendation model.

[0337] In some embodiments, the third generating module 840 is configured to:

[0338] Acquire at least one triple related to the label object from an external knowledge source, wherein the triple includes a head entity, a tail entity, and a relationship type connecting the head entity and the tail entity, wherein the head entity is used to indicate the label object, the tail entity is acquired from a comment sentence for the label object, and the relationship type is acquired from the external knowledge source;

[0339] obtaining a second vocabulary set according to the at least one sample dialogue sentence, wherein second vocabulary included in the second vocabulary set is a vocabulary used to describe entity information in the at least one sample dialogue sentence;

[0340] Based on the at least one triple, the second vocabulary set, the first label sentence and the second label sentence, a third training sample corresponding to the sentiment alignment task is generated, wherein the sentiment alignment task uses the second label sentence as label data and uses the at least one triple, the second vocabulary set, the initial label sentence and the label object as sample data.

[0341] In some embodiments, the first generating module 820 is configured to:

[0342] obtaining a first vocabulary set according to the at least one sample dialogue sentence, wherein a first vocabulary contained in the first vocabulary set is a vocabulary in the at least one sample dialogue sentence;

[0343] Obtaining, based on the first vocabulary set, a first feature set corresponding to the first vocabulary set, wherein a first feature included in the first feature set is a feature representation of the first vocabulary after encoding;

[0344] A first training sample corresponding to the reply task is generated based on the at least one sample dialogue sentence, the first feature set, the random vector, and the initial label sentence, wherein the reply task uses the initial label sentence as label data and uses the at least one sample dialogue sentence, the first feature set, and the random vector as sample data.

[0345] In some embodiments, the second generating module 830 is configured to:

[0346] obtaining a second vocabulary set according to the at least one sample dialogue sentence, wherein second vocabulary included in the second vocabulary set is a vocabulary used to describe entity information in the at least one sample dialogue sentence;

[0347] determining, based on the second vocabulary set, an emotional feature corresponding to the at least one sample dialogue sentence;

[0348] A second training sample corresponding to the object recommendation task is generated based on the at least one sample dialogue sentence, the emotional feature, the random vector, the initial label sentence, and the label object, wherein the object recommendation task uses the second label sentence as label data and uses the at least one sample dialogue sentence, the emotional feature, the random vector, and the initial label sentence as sample data.

[0349] In some embodiments, the training module 850 is configured to:

[0350] Adjusting the parameters of the conversation recommendation model according to the first loss function value corresponding to the reply task to obtain the conversation recommendation model trained for the reply task;

[0351] Adjusting parameters of the conversation recommendation model trained for the reply task according to a second loss function value corresponding to the object recommendation task to obtain a conversation recommendation model trained for the object recommendation task;

[0352] The parameters of the conversation recommendation model trained for the object recommendation task are adjusted according to the third loss function value corresponding to the emotion alignment task to obtain the trained conversation recommendation model.

[0353] In some embodiments, the training module 850 is configured to:

[0354] Obtaining a first loss function value based on differences between the first predicted sentence obtained from the reply task and the vocabulary at each position of the initial label sentence;

[0355] Obtaining a second loss function value according to a difference between the predicted object obtained from the object recommendation task and the labeled object;

[0356] The third loss function value is obtained according to the difference between the vocabulary at each position of the second predicted sentence obtained by the sentiment alignment task and the second label sentence.

[0357] In some embodiments, the training module 850 is configured to:

[0358] Calculating the second loss function value according to the predicted probability of at least one predicted object obtained from the object recommendation task;

[0359] or,

[0360] Based on the feedback label of the at least one prediction object, a weight value corresponding to each of the at least one prediction object is obtained, where the feedback label is a pre-labeled label used to indicate each user's preference for the prediction object; and based on the predicted probability of the at least one prediction object and the weight value corresponding to the at least one prediction object, the second loss function value is calculated.

[0361] It should be noted that the apparatus provided in the above embodiments, when implementing its functions, is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the content structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0362] Please refer to Figure 9 , which shows a block diagram of a computer device 900 provided in one embodiment of the present application. Computer device 900 can be any electronic device capable of data computing, processing, and storage. Computer device 900 can be used to implement the conversation recommendation method or conversation recommendation model training method provided in the above embodiments.

[0363] Typically, the computer device 900 includes a processor 901 and a memory 902 .

[0364] The processor 901 may include one or more processing cores, such as a 4-core processor, an 8-core processor, and the like. The processor 901 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), or PLA (Programmable Logic Array). The processor 901 may also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 901 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 901 may also include an AI processor for processing computing operations related to machine learning.

[0365] Memory 902 may include one or more computer-readable storage media, which may be non-transitory. Memory 902 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory storage devices. In some embodiments, the non-transitory computer-readable storage media in memory 902 is used to store a computer program, which is configured to be executed by one or more processors to implement the above-mentioned conversation recommendation method or conversation recommendation model training method.

[0366] Those skilled in the art will understand that Figure 9 The structure shown in the figure does not constitute a limitation on the computer device 900, and the computer device 900 may include more or fewer components than shown in the figure, or combine some components, or adopt a different component arrangement.

[0367] In an illustrative embodiment, a computer-readable storage medium is also provided, storing a computer program that, when executed by a processor of a computer device, implements the aforementioned conversation recommendation method or conversation recommendation model training method. Optionally, the computer-readable storage medium may be a ROM (Read-Only Memory), RAM (Random Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk, or optical data storage device.

[0368] In an exemplary embodiment, a computer program product is also provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the aforementioned conversation recommendation method or conversation recommendation model training method.

[0369] It should be noted that this application can display a prompt interface, pop-up window or output voice prompt information before collecting the user's relevant data and during the process of collecting the user's relevant data. The prompt interface, pop-up window or voice prompt information is used to remind the user that its relevant data is currently being collected, so that this application only starts to execute the relevant steps of obtaining the user's relevant data after obtaining the user's confirmation operation on the prompt interface or pop-up window. Otherwise (that is, when the user's confirmation operation on the prompt interface or pop-up window is not obtained), the relevant steps of obtaining the user's relevant data are terminated, that is, the user's relevant data is not obtained. In other words, all user data (including conversation data) collected by this application are processed strictly in accordance with the requirements of relevant national laws and regulations. The informed consent or separate consent of the personal information subject is obtained only when the user agrees and authorizes it to be collected, and the subsequent data use and processing behavior is carried out within the scope of authorization of laws and regulations and the personal information subject. The collection, use and processing of relevant user data must comply with the relevant laws, regulations and standards of relevant countries and regions.

[0370] It should be understood that the "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship. In addition, the step numbers described in this article only illustrate a possible execution sequence between the steps. In some other embodiments, the above steps may not be executed in the order of the numbers, such as two steps with different numbers are executed at the same time, or two steps with different numbers are executed in the opposite order to the diagram. The embodiments of the present application do not limit this.

[0371] The above description is merely an exemplary embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A conversation recommendation method, characterized in that: The method comprises: Get at least one historical dialogue sentence; generating, by a dialogue recommendation model, a first answer sentence based on the at least one historical dialogue sentence, wherein the first answer sentence includes a recommendation object for the at least one historical dialogue sentence; The dialogue recommendation model generates a second answer statement based on the first answer statement and entity vocabulary related to the recommendation object obtained from an external knowledge source. The second answer statement includes the first answer statement and an emotional answer statement for the recommendation object.

2. The method according to claim 1, characterized in that The step of generating a second answer sentence based on the first answer sentence and entity vocabulary related to the recommendation object obtained from an external knowledge source by the dialogue recommendation model includes: Acquire at least one triple related to the recommended object from the external knowledge source, wherein the triple includes a head entity, a tail entity, and a relationship type connecting the head entity and the tail entity, wherein the head entity is used to indicate the recommended object; Acquiring at least one comment sentence for the recommended object from the external knowledge source, selecting entity words with a number of occurrences greater than a first threshold from the at least one comment sentence, and obtaining a set of associated words; The second answer sentence is generated by the dialogue recommendation model according to the at least one triple, at least one associated word in the associated word set, the recommendation object and the first answer sentence.

3. The method according to claim 2, characterized in that The step of generating the second answer statement by the dialogue recommendation model according to the at least one triple, at least one associated word in the associated word set, the recommended object, and the first answer statement includes: generating the emotional answer statement using the dialogue recommendation model based on the at least one triple, at least one associated word in the associated word set, the recommended object, and the first answer statement; and concatenating the first answer statement and the emotional answer statement to obtain the second answer statement; or, The second answer statement including the emotional answer statement is generated through the dialogue recommendation model according to the at least one triple, at least one associated word in the associated word set, the recommendation object and the first answer statement.

4. The method according to claim 1, wherein The step of generating a first answer statement according to the at least one historical dialogue statement using the dialogue recommendation model includes: generating an initial answer sentence based on the at least one historical dialogue sentence using the dialogue recommendation model, wherein the initial answer sentence includes a blank object for the at least one historical dialogue sentence; generating the recommended object according to the at least one historical dialogue sentence and the initial answer sentence by the dialogue recommendation model; The blank object in the initial answer sentence is filled with the recommended object to obtain the first answer sentence.

5. The method according to claim 4, characterized in that Generating an initial answer statement according to the at least one historical dialogue statement by the dialogue recommendation model includes: obtaining a first vocabulary set according to the at least one historical conversation sentence, wherein a first vocabulary contained in the first vocabulary set is a vocabulary in the at least one historical conversation sentence; Obtaining, based on the first vocabulary set, a first feature set corresponding to the first vocabulary set, wherein a first feature included in the first feature set is a feature representation of the first vocabulary after encoding; The initial answer sentence is generated by the dialogue recommendation model according to the at least one historical dialogue sentence and the first feature set.

6. The method according to claim 5, characterized in that The obtaining, based on the first vocabulary set, a first feature set corresponding to the first vocabulary set includes: Obtaining, based on the first vocabulary set, a second vocabulary set corresponding to the first vocabulary set, wherein second words included in the second vocabulary set are at least one word in the first vocabulary set used to describe entity information; The first vocabulary set and the second vocabulary set are fused using a first linear transformation to obtain the first feature set.

7. The method according to claim 6, characterized in that Generating the recommendation object according to the at least one historical dialogue sentence and the initial answer sentence by the dialogue recommendation model includes: determining, based on the second vocabulary set, an emotional feature corresponding to the at least one historical conversation sentence; The recommended object is generated through the dialogue recommendation model according to the at least one historical dialogue sentence, the emotional feature and the initial answer sentence.

8. The method according to claim 7, characterized in that The determining, based on the second vocabulary set, an emotional feature corresponding to the at least one historical dialogue sentence includes: determining, based on the second vocabulary set, a local emotion feature corresponding to the at least one historical conversation sentence, the local emotion feature being used to represent an emotion feature of the at least one historical conversation sentence; and determining the local emotion feature as the emotion feature corresponding to the at least one historical conversation sentence; or, Based on the second vocabulary set, local emotional features corresponding to the at least one historical dialogue sentence are determined; based on the second vocabulary set and entity vocabulary related to the second vocabulary obtained from the external knowledge source, global emotional features corresponding to the at least one historical dialogue sentence are determined, wherein the global emotional features are used to represent emotional features associated with the at least one historical dialogue sentence; the local emotional features and the global emotional features are determined as emotional features corresponding to the at least one historical dialogue sentence.

9. The method according to claim 8, characterized in that Determining, based on the second vocabulary set, a local emotion feature corresponding to the at least one historical dialogue sentence includes: Performing emotion recognition on the at least one historical conversation sentence using an emotion recognition model to obtain an emotion label set corresponding to each of the at least one historical conversation sentence, and an emotion probability set corresponding to the emotion label set, wherein the emotion label set includes at least one emotion label, and the emotion probability set includes probabilities corresponding to each of the at least one emotion label; Obtaining, based on the emotion tag set and the emotion probability set, an emotion tag set corresponding to a local entity vocabulary and an emotion probability set corresponding to the local entity vocabulary, wherein the local entity vocabulary is a vocabulary in the second vocabulary set; Obtaining the emotional features of the local entity vocabulary according to the emotional tag set corresponding to the local entity vocabulary and the emotional probability set corresponding to the local entity vocabulary; The local emotion feature is obtained according to the second vocabulary set, and the local emotion feature includes an emotion feature of at least one of the second vocabulary.

10. The method according to claim 9, characterized in that The obtaining of the emotion features of the local entity vocabulary according to the emotion tag set corresponding to the local entity vocabulary and the emotion probability set corresponding to the local entity vocabulary includes: Obtaining initial emotion features of the local entity vocabulary according to the emotion tag set corresponding to the local entity vocabulary and the emotion probability set corresponding to the local entity vocabulary; fusing the first vocabulary set and the second vocabulary set using a second linear transformation to obtain a second feature set, wherein second features included in the second feature set are feature representations of the second vocabulary after encoding; The initial emotion feature of the local entity vocabulary and the second feature corresponding to the local entity vocabulary are fused to obtain the emotion feature of the local entity vocabulary, and the feature dimension of the emotion feature of the local entity vocabulary is the same as the feature dimension of the second feature.

11. The method according to claim 9, characterized in that The determining, based on the second vocabulary set and entity vocabulary related to the second vocabulary obtained from the external knowledge source, a global sentiment feature corresponding to the at least one historical dialogue sentence includes: Acquire at least one global entity vocabulary related to the local entity vocabulary from the external knowledge source, wherein the global entity vocabulary is an entity vocabulary that appears in the same dialogue sentence as the local entity vocabulary; Calculating a co-occurrence probability corresponding to each of the at least one global entity vocabulary, wherein the co-occurrence probability refers to a probability that the global entity vocabulary and the local entity vocabulary appear in the same dialogue sentence; Obtaining a global sentiment feature of the local entity vocabulary according to the co-occurrence probabilities corresponding to the local entity vocabulary and the at least one global entity vocabulary; The global emotion feature is obtained according to the second vocabulary set, and the global emotion feature includes a global emotion feature of at least one of the second vocabulary.

12. A method for training a dialogue recommendation model, characterized in that: The method comprises: Obtaining a basic dataset for training the dialogue recommendation model, the basic dataset including at least one data group, each of the data groups including at least one sample dialogue sentence, a first label sentence corresponding to the at least one sample dialogue sentence, and a second label sentence corresponding to the at least one sample dialogue sentence, wherein the first label sentence includes a label object for the at least one sample dialogue sentence, and the second label sentence includes the first label sentence and an emotional response sentence for the label object; generating, based on the at least one data group, a first training sample corresponding to a reply task, wherein the reply task uses an initial labeled sentence as label data and determines sample data based on the at least one sample dialogue sentence, wherein the initial labeled sentence includes a blank object for the at least one sample dialogue sentence; generating, based on the at least one data group, a second training sample corresponding to an object recommendation task, wherein the object recommendation task uses the labeled object as label data and determines sample data based on the at least one sample dialogue sentence and the first label sentence; generating, based on the at least one data group, a third training sample corresponding to a sentiment alignment task, wherein the sentiment alignment task uses the second labeled sentence as labeled data and determines sample data based on the at least one sample dialogue sentence and the first labeled sentence; The first training sample, the second training sample, and the third training sample are used to train the dialogue recommendation model to obtain a trained dialogue recommendation model.

13. The method according to claim 12, characterized in that The step of training the conversation recommendation model using the first training sample, the second training sample, and the third training sample to obtain a trained conversation recommendation model includes: Adjusting the parameters of the conversation recommendation model according to the first loss function value corresponding to the reply task to obtain the conversation recommendation model trained for the reply task; Adjusting parameters of the conversation recommendation model trained for the reply task according to a second loss function value corresponding to the object recommendation task to obtain a conversation recommendation model trained for the object recommendation task; The parameters of the conversation recommendation model trained for the object recommendation task are adjusted according to the third loss function value corresponding to the emotion alignment task to obtain the trained conversation recommendation model.

14. The method according to claim 13, characterized in that The method further comprises: Obtaining a first loss function value based on differences between the first predicted sentence obtained from the reply task and the vocabulary at each position of the initial label sentence; Obtaining a second loss function value according to a difference between the predicted object obtained from the object recommendation task and the labeled object; The third loss function value is obtained according to the difference between the vocabulary at each position of the second predicted sentence obtained by the sentiment alignment task and the second label sentence.

15. The method according to claim 14, characterized in that Obtaining the second loss function value by the difference between the predicted object obtained according to the object recommendation task and the labeled object includes: Calculating the second loss function value according to the predicted probability of at least one predicted object obtained from the object recommendation task; or, Based on the feedback label of the at least one prediction object, a weight value corresponding to each of the at least one prediction object is obtained, where the feedback label is a pre-labeled label used to indicate each user's preference for the prediction object; and based on the predicted probability of the at least one prediction object and the weight value corresponding to the at least one prediction object, the second loss function value is calculated.

16. A conversation recommendation device, characterized in that: The device comprises: A sentence acquisition module, used to acquire at least one historical conversation sentence; A first generating module, configured to generate a first answer sentence based on the at least one historical dialogue sentence using a dialogue recommendation model, wherein the first answer sentence includes a recommendation object for the at least one historical dialogue sentence; The second generation module is used to generate a second answer statement based on the first answer statement and entity vocabulary related to the recommendation object obtained from an external knowledge source through the dialogue recommendation model, wherein the second answer statement includes the first answer statement and an emotional answer statement for the recommendation object.

17. A training device for a dialogue recommendation model, characterized in that: The device comprises: a data acquisition module, configured to acquire a basic data set for training the dialogue recommendation model, wherein the basic data set includes at least one data group, each of which includes at least one sample dialogue sentence, a first label sentence corresponding to the at least one sample dialogue sentence, and a second label sentence corresponding to the at least one sample dialogue sentence, wherein the first label sentence includes a label object for the at least one sample dialogue sentence, and the second label sentence includes the first label sentence and an emotional response sentence for the label object; a first generating module, configured to generate, based on the at least one data group, a first training sample corresponding to a reply task, wherein the reply task uses an initial labeled sentence as labeled data and determines sample data based on the at least one sample dialogue sentence, wherein the initial labeled sentence includes a blank object for the at least one sample dialogue sentence; a second generating module, configured to generate, based on the at least one data group, a second training sample corresponding to an object recommendation task, wherein the object recommendation task uses the labeled object as label data and determines sample data based on the at least one sample dialogue sentence and the first label sentence; a third generating module, configured to generate, based on the at least one data group, a third training sample corresponding to a sentiment alignment task, wherein the sentiment alignment task uses the second labeled sentence as labeled data and determines sample data based on the at least one sample dialogue sentence and the first labeled sentence; A training module is used to train the dialogue recommendation model using the first training sample, the second training sample, and the third training sample to obtain a trained dialogue recommendation model.

18. A computer device, characterized in that: The computer device includes a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the dialogue recommendation method according to any one of claims 1 to 11, or to implement the training method of the dialogue recommendation model according to any one of claims 12 to 15.

19. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which is loaded and executed by a processor to implement the conversation recommendation method according to any one of claims 1 to 11, or the training method of the conversation recommendation model according to any one of claims 12 to 15.

20. A computer program product, characterized in that The computer program product includes a computer program, which is loaded and executed by a processor to implement the dialogue recommendation method according to any one of claims 1 to 11, or the training method of the dialogue recommendation model according to any one of claims 12 to 15.