Multi-modal common-condition reply generation method based on tool learning

By constructing the TOOL-STCIKERCONV multimodal empathy response generation framework, the decoupling empathy response task is two parts: multimodal chat and empathy package content description, which solves the problem of chatbots lacking the ability to actively reply empathy packages, and achieves the end-to-end generation of multimodal empathy responses and improves emotional expression.

CN119938994AActive Publication Date: 2025-05-06NORTHEASTERN UNIV CHINA
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510027568.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-06
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

The current chatbots lack the ability to actively reply to emoticons or pictures, which leads to a greater room for improvement in emotional expression and empathy expression.

Method used

Using the multimodal empathy replies generation method based on tool learning, the TOOL-STCIKERCONV multimodal empathy replies generation framework is constructed, and the data transformation module, the tool call module and the evaluation training module are used to decouple the empathy replies task for multimodal chat subtasks and emoticon package content description generation subtasks to realize end-to-end training of the multimodal base model.

Benefits of technology

It realizes end-to-end generation of multimodal empathy replies, simulates the frequency of sentiment of emoticon packets in human chats, and improves the emotional expression and empathy expression capabilities of chat robots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938994A_ABST
    Figure CN119938994A_ABST
Patent Text Reader

Abstract

The invention provides a tool learning-based multi-modal common-situation reply generation method, and relates to the technical field of common-situation dialogue generation. According to the method, a tool learning-based multi-modal common-situation reply generation framework TOOL-STCIKERCONV is constructed, and a data transformation module, a tool calling module and an evaluation training module are included; according to the method, a constructed dialogue data set is transformed, a method of inserting a special token is used, and a fragment containing the special token is added in a content description field Prompt of an expression package, so that a multi-modal base model has the capability of thinking generation of the expression package, thinking in an implicit vector space is realized, learning of the multi-modal base model actively calls an expression package generation tool, and the expression package generation efficiency is improved. The problem that the expression package sending frequency is abnormal due to the fact that a large model in the multi-mode common-situation dialogue generation field lacks autonomous thinking ability is solved; in addition, according to the method, a plug-and-play emoji package generation tool is utilized, iterative updating of an emoji package generation model is achieved, and different styles of text graph models are replaced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of empathy dialogue generation, and in particular to a multimodal empathy response generation method based on tool learning. Background Art

[0002] In today's era, chatting has become one of the important ways of communication. During the chat, the two parties express their emotions by sending text messages to each other. But in addition to text messages, pictures and emoticons can also accurately express personal emotions and actively promote emotional exchanges between the two parties. With the development of artificial intelligence technology, the two parties in the chat are no longer limited to people. How to achieve human-computer interaction has become a hot issue in the field of current dialogue systems. At this stage, chatbots have achieved initial results and are widely used in customer service, entertainment, medical care, voice assistants and other fields. However, at this stage, chatbots are mainly based on text replies and do not have the ability to actively reply to anthropomorphic information such as emoticons or pictures. Therefore, there is still a lot of room for improvement in the emotional expression and empathy expression of chatbots during the conversation process. Large language models such as Chatgpt and Qwen-vl are powerful artificial intelligence technologies that have outstanding performance in paper writing, dialogue question and answer, and intelligent medical care, but research in the field of multimodal empathy dialogue remains to be explored.

[0003] Qian et al. conducted a comprehensive empirical study on the performance of LLMs represented by ChatGPT in empathetic response generation in the paper "Harnessing the Power of Large Language Models for Empathetic Response Generation: Empirical Investigations and Improvements". Based on LLMs, they proposed three improvement methods (semantically similar context learning, two-stage interactive generation, and integration with knowledge base) and verified the effectiveness of the three improvement methods through experiments. Judging from the objective evaluation indicators in the paper, the three improvement strategies can significantly improve the model's empathetic response ability and achieve better response effects.

[0004] In the paper "A Neural Network Approach to Context-Sensitive Generation of Conversational Responses", Sordoni et al. proposed the need to use a multi-layer feedforward neural network to replace the original RNN structure in the encoder part. This will introduce contextual information into the model, thereby alleviating the problem that the model cannot capture long-distance dependencies, so that the final model reply content can reflect more chat conversation history, thereby improving the user's chat experience.

[0005] Miao et al. proposed a seq2seq-based empathy dialogue generation model in the paper "Emotional Dialogue Generation with Emotion Embedding", and embedded emotional information into the model decoder to improve the empathy expression ability during the dialogue process. However, since the model filled in the seq2seq framework at that time was RNN, it was unable to maximize the effect of the generative model.

[0006] Liu et al. proposed an empathetic response model based on the attention mechanism in the paper "Empathetic Dialogue Generation with Pre-trained RoBERTa-GPT2 and External Knowledge". They chose to use the autoencoder model RoBERTa as the encoder, hoping to capture more contextual information through the advantages of the autoencoder pre-training model. Correspondingly, the autoregressive model GPT2 was selected as the decoder, hoping to generate higher quality content through the advantages of the autoregressive pre-training model. However, it does not extract or process any user emotional information, resulting in the overall model only responding to small talk based on the context, ignoring the user's emotional information.

[0007] Zhang et al. proposed an agent Agent4SC with multimodal empathetic response capabilities in the paper "STICKERCONV: Generating Multimodal Empathetic Responses from Scratch". It mainly consists of six parts: tool module, role module, planning module, memory module, behavior module and agent management module. This agent successfully simulates the behavior of humans using emoticons during chat. After sorting, filtering and screening the interaction data generated by this agent, a multimodal empathetic dialogue dataset STICKERCONV was obtained, which includes 12.9K dialogues, 5.8K non-repeated emoticons and 2K different dialogue scenes. Considering the high cost of Agent4SC in deployment, the agent was distilled and a lightweight version of the multimodal empathetic response generation framework PEGS was designed. PEGS uses the STICKERCONV dataset for fine-tuning training to give it multimodal empathetic response capabilities. The base model ensures its normal text response capabilities, and uses general retrieval, generation and retrieval-enhanced generation methods to obtain image response capabilities. However, in actual applications, PEGS has problems such as abnormal frequency of emoticon replies and huge cost consumption caused by the need to re-train with the large model when replacing the emoticon generation model. Summary of the invention

[0008] The technical problem to be solved by the present invention is to provide a multimodal empathy response generation method based on tool learning in view of the deficiencies of the above-mentioned prior art. Through tool learning, a large language model is allowed to learn the way humans use emoticons, and multimodal empathy responses are implemented end-to-end, simulating the frequency of humans sending emoticons during chats.

[0009] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0010] The present invention provides a multimodal empathy reply generation method based on tool learning, constructs a multimodal empathy reply generation framework TOOL-STCIKERCONV based on tool learning, including a data transformation module, a tool calling module and an evaluation training module, wherein the data transformation module is used to transform the constructed dialogue data set, and call the evaluation training module to screen and filter the transformed data set to ensure data quality; the tool calling module decouples the empathy reply task into a multimodal chatting subtask based on a multimodal base model and a multimodal emoticon package content description generation subtask based on tool calling; the evaluation training module uses the transformed data set to train the multimodal base model, and evaluates the effect of the training result, objectively evaluating the quality and usability of the reply content;

[0011] The following steps are involved:

[0012] Step 1: Construct a conversation dataset and call the data transformation module to transform the constructed conversation dataset, dividing the conversation dataset into a training set, a validation set, and a test set in proportion;

[0013] Step 1.1: Obtain the conversation content entered by the user, the description of the emoticon package used by the chatbot to reply, the storage path of the emoticon package, and the text information replied by the chatbot to construct a conversation dataset;

[0014] The conversation dataset includes: "user", "function_call", "observation" and "assistant". The "user" field corresponds to the conversation content entered by the user, including text and emoticon information; the "function_call" field corresponds to the emoticon description used by the chatbot for reply; the "observation" field corresponds to the storage path of the emoticon used by the chatbot for reply; and the "assistant" field corresponds to the text information replied by the chatbot.

[0015] Step 1.2: Call the data transformation module to transform the constructed dialogue dataset, reorganize the content and splice the instructions of different fields in the dialogue dataset, and obtain the transformed dialogue dataset;

[0016] The specific method for data transformation of the constructed dialogue dataset is:

[0017] Perform tokenization on the conversation data in the constructed dataset, decompose the natural language text into the smallest vocabulary unit token, define the token composed of vocabulary as text token, define the token composed of special symbols <·>,< / ·> The token composed of the text is a special token, which is used to express a specific meaning. The token is converted into the corresponding index token_id under the function of vocabulary mapping, which is used to obtain the word vector corresponding to the token in the pre-training stage.

[0018] For the user input content contained in the "user" field, use irrelevant tokens to adaptively fill the beginning and end of the user input text, so that the content length contained in the "user" field is consistent after the token is filled. The irrelevant token itself has no meaning and is only used to ensure that the conversation dataset meets the fixed parameter specifications;

[0019] For the "function_call" field, use the method of inserting special tokens to add the model thinking field though and the emoticon content description field prompt to the content corresponding to the "function_call" field. The content corresponding to the though field includes: fixed starting token " <though>", fixed prefix, corresponding emotion and fixed end token"< / though> ", "prompt" field corresponds to the detailed description of the emoticon package, which serves as a guide for generating emoticon package content for the target emotion; <thought> and< / thought> Used to prompt whether to reply to the emoticon package and the content of the emoticon package to be replied at this stage; <thought> and< / thought> The middle part is the process of determining whether an emoticon needs to be output. For the current conversation scenario, the judgment result is that an emoticon picture expressing the corresponding emotion needs to be replied, and a complete description of the content of the emoticon picture is made. The complete description of the emoticon picture content is then converted into token_id, and irrelevant tokens are used for adaptive padding to make the content length corresponding to the function_call field consistent;

[0020] For the "observation" emoji path field, use irrelevant tokens to adaptively fill it, and keep the emoji path length corresponding to the emoji path field consistent;

[0021] For the "assistant" plain text reply field, convert the text reply content into the corresponding token_id, and process the input and output formats so that the content length corresponding to the "assistant" field is consistent;

[0022] Step 1.3: Call the evaluation training module to screen and filter the transformed data set, and divide the transformed dialogue data set into training set, validation set, and test set in proportion;

[0023] Step 2: Call the tool calling module to decouple the empathy reply task into a multimodal small talk subtask based on the multimodal base model and a multimodal emoticon package content description generation subtask based on tool calling, and generate corresponding text reply content and emoticon package reply content respectively;

[0024] Step 2.1: Call the tool calling module to decouple the empathy reply task into a multimodal small talk subtask based on the multimodal base model and a multimodal emoticon package content description generation subtask based on tool calling, and select the multimodal base model to generate text reply content;

[0025] In the multimodal small talk subtask based on the multimodal base model, the multimodal base model generates text reply content according to the multimodal dialogue history;

[0026] In the tool-based multimodal emoji content description generation subtask, the multimodal base model determines whether it is necessary to call the emoji generated by the tool based on the current conversation scene and conversation history, and then calls the emoji generation tool GenerateSticker to generate the emoji reply content corresponding to the text reply content;

[0027] The multimodal base model generates text reply content according to the multimodal conversation history and determines whether to reply with an emoticon. If the multimodal base model determines that the current scene needs to reply with an emoticon package, the multimodal emoticon package content description generation subtask based on tool call is entered, and the emoticon package generation tool is called to directly generate an emoticon package reply content corresponding to the text reply content. If the multimodal base model determines that the current scene does not need to reply with an emoticon package, the multimodal emoticon package content description generation subtask based on tool call is skipped, and only a text reply is generated;

[0028] Step 2.2 The multimodal base model calls the emoticon generation tool GenerateSticker, regards the conversation context as a parameter input to the emoticon generation tool, generates the emoticon description content description field prompt, and then generates the emoticon reply content corresponding to the text reply content;

[0029] The emoji generation tool GenerateSticker uses a plug-and-play emoji generation model. It selects emoji generation models of different styles according to the different emoji style requirements of users.

[0030] Step 3: Call the evaluation training module and use the modified dialogue dataset to train and evaluate the multimodal base model to obtain a trained multimodal base model;

[0031] Step 3.1: The current user input content X, the conversation history H, and the emoticon content description field Q are transferred to the multimodal base model LLMs to obtain the text reply content Y and the emoticon reply content Z, as shown in the following formula:

[0032] Y,Z=LLMs(X,Q,H)

[0033] The loss function L of the multimodal base model LLMs is:

[0034]

[0035]

[0036] L=L y +L z

[0037] Among them, h is the generated vector feature of the input content after passing through the multimodal base model, θ is a fixed parameter in the large model that does not participate in the update, and E LoRA is a parameter that can be updated during training, y i ,z i The i-th token of the content is generated for the reply content and the emoji content description field Prompt, k is the number of tokens, L y ,L z They are the text reply part loss function in the training process and the emoji content description field Prompt generation part loss function respectively, and L is the final model loss function;

[0038] Step 3.2: Select the pure text evaluation index, emoticon sending frequency evaluation index, and picture-text reply content consistency evaluation index to evaluate the quality and usability of the reply content of the trained multimodal base model, and select the multimodal base model with the best evaluation index as the multimodal base model of the TOOL-STCIKERCONV multimodal empathy reply framework to generate the corresponding text replies and emoticon replies;

[0039] Step 4: Based on the user input, use the trained TOOL-STCIKERCONV multimodal empathy response framework to generate corresponding text responses and emoticon responses.

[0040] The beneficial effects of adopting the above technical solution are as follows: the present invention provides a multimodal empathy response generation method based on tool learning, which transforms the dialogue data set and uses the special token insertion method to add a fragment containing a special token in the Prompt field of the emoticon content description field, so that the multimodal base model has the ability to think about emoticon generation, realizes thinking in the latent vector space, and allows the multimodal base model to learn to actively call the emoticon generation tool, thereby solving the problem of abnormal emoticon sending frequency caused by the lack of autonomous thinking ability of large models in the field of multimodal empathy dialogue generation; in addition, the present invention uses a plug-and-play emoticon generation tool to realize the iterative update of the emoticon generation model and the replacement of different styles of literary and graphic models. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 A schematic diagram of a multimodal empathy dialogue including an emoticon package provided in an embodiment of the present invention;

[0042] Figure 2 A diagram of the TOOL-STCIKERCONV multimodal empathic response framework provided by an embodiment of the present invention;

[0043] Figure 3 A schematic diagram of tool calling in the generation of multimodal empathic responses provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0044] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0045] Empathic response dialogue was first proposed in FaceBook's empathic dialogue benchmark in 2019 and is one of the core tasks in the emotional field. In the traditional empathic response field, researchers have focused their efforts on the pure text field. Psychological studies have shown that if other modal contents, such as images and emoticons, are introduced during empathic dialogue, better empathy effects will be achieved. Therefore, this embodiment simulates the normal chat state of humans and implements a multi-modal empathic response mode of "text + emoticons" to enhance the user's interactive experience. Multi-modal empathic dialogues such as Figure 1 shown.

[0046] There are two main challenges in the field of multimodal empathy dialogue generation: the first is the lack of multimodal empathy dialogue datasets, that is, most of the current empathy dialogue tasks are single-mode pure text, and there is no large-scale multimodal dialogue dataset. The second is that most of the current dialogue models do not have the ability to actively send emoticons, and lack relevant thinking on how to send emoticons, resulting in problems such as boring dialogue content and abnormal frequency of emoticon sending.

[0047] In view of the above problems, this embodiment provides a multimodal empathy response method based on tool learning, and constructs a tool learning-based TOOL-STCIKERCONV multimodal empathy response generation framework, including a data transformation module, an evaluation training module, and a tool calling module. Figure 2 As shown, the data transformation module is used to transform the constructed dialogue dataset, and call the evaluation training module to screen and filter the transformed dataset to ensure data quality; the tool calling module decouples the empathy reply task into a multimodal chat subtask based on the multimodal base model and a multimodal emoticon package content description generation subtask based on tool calling; the evaluation training module uses the transformed dataset to train the multimodal base model, and evaluates the training results, objectively evaluating the quality and usability of the reply content;

[0048] The multimodal empathic response generation method based on tool learning includes the following steps:

[0049] Step 1: Construct a conversation dataset and call the data transformation module to transform the constructed conversation dataset, dividing the conversation dataset into a training set, a validation set, and a test set in proportion;

[0050] Step 1.1: Obtain the conversation content entered by the user, the description of the emoticon package used by the chatbot to reply, the storage path of the emoticon package, and the text information replied by the chatbot to construct a conversation dataset;

[0051] The conversation dataset includes the "user", "function_call", "observation" and "assistant" fields. The "user" field corresponds to the conversation content entered by the user, including text and emoticon information. The "function_call" field corresponds to the emoticon description used by the chatbot to reply. The "observation" field corresponds to the storage path of the emoticon used by the chatbot to reply. The "assistant" field corresponds to the text information replied by the chatbot.

[0052] Step 1.2: Call the data transformation module to transform the constructed dialogue dataset, reorganize the content and splice the instructions of different fields in the dialogue dataset, and obtain the transformed dialogue dataset;

[0053] The specific method for data transformation of the constructed dialogue dataset is:

[0054] Perform tokenization on the conversation data in the constructed dataset, decompose the natural language text into the smallest vocabulary unit token, define the token composed of vocabulary as text token, define the token composed of special symbols <·>,< / ·> The token composed of the text is a special token, which is used to express a specific meaning. The token is converted into the corresponding index token_id under the function of vocabulary mapping, which is used to obtain the word vector corresponding to the token in the pre-training stage.

[0055] For the user input content contained in the "user" field, use irrelevant tokens to adaptively fill the beginning and end of the user input text, so that the content length contained in the "user" field is consistent after the token is filled. The irrelevant token itself has no meaning and is only used to ensure that the conversation dataset meets the fixed parameter specifications;

[0056] For the "function_call" field, use the method of inserting special tokens to add the model thinking field though and the emoticon content description field prompt to the content corresponding to the "function_call" field. The content corresponding to the though field includes: fixed starting token " <though> ", fixed prefix, corresponding emotion and fixed end token"< / though> ", "prompt" field corresponds to the detailed description of the emoticon package, which serves as a guide for generating emoticon package content for the target emotion; <thought> and< / thought>Used to prompt whether to reply to the emoticon package and the content of the emoticon package to be replied at this stage; <thought> and< / thought> The middle part is the process of determining whether an emoticon needs to be output. For the current conversation scenario, the judgment result is that an emoticon picture expressing the corresponding emotion needs to be replied, and a complete description of the content of the emoticon picture is made. The complete description of the emoticon picture content is then converted into token_id, and irrelevant tokens are used for adaptive padding to make the content length corresponding to the function_call field consistent;

[0057] For the "observation" emoji path field, a similar method is used to process the "user" field. After using irrelevant tokens for adaptive filling, the emoji path length corresponding to the emoji path field is kept consistent.

[0058] For the "assistant" plain text reply field, use a similar processing method to the "function_call" field to convert the text reply content into the corresponding token_id, and process the input and output formats so that the content length corresponding to the "assistant" field is consistent;

[0059] Step 1.3: Call the evaluation training module to screen and filter the transformed data set, and divide the transformed dialogue data set into training set, validation set, and test set in proportion;

[0060] In this embodiment, the preprocessed conversation data set is divided into a training set, a validation set, and a test set in a ratio of 8:1:1;

[0061] Step 2: Call the tool calling module to decouple the empathy reply task into a multimodal small talk subtask based on the multimodal base model and a multimodal emoticon package content description generation subtask based on tool calling, and generate corresponding text reply content and emoticon package reply content respectively;

[0062] Step 2.1: Call the tool calling module to decouple the empathy reply task into a multimodal small talk subtask based on the multimodal base model and a multimodal emoticon package content description generation subtask based on tool calling, and select the multimodal base model to generate text reply content;

[0063] In the multimodal small talk subtask based on the multimodal base model, the multimodal base model generates text reply content according to the multimodal dialogue history;

[0064] In the tool-based multimodal emoji content description generation subtask, the multimodal base model determines whether it is necessary to call the emoji generated by the tool based on the current conversation scene and conversation history, and then calls the emoji generation tool GenerateSticker to generate the emoji reply content corresponding to the text reply content;

[0065] The multimodal base model generates text reply content according to the multimodal conversation history and determines whether to reply with emoticons. If the multimodal base model determines that the current scenario requires an emoticon reply, it enters the multimodal emoticon content description generation subtask based on tool call, and calls the emoticon generation tool to directly generate emoticon reply content corresponding to the text reply content. If the multimodal base model determines that the current scenario does not require an emoticon reply, it skips the multimodal emoticon content description generation subtask based on tool call and only generates a text reply. The tool call diagram is shown in the figure. Figure 3 As shown;

[0066] Step 2.2 The multimodal base model calls the emoticon generation tool GenerateSticker, regards the conversation context as a parameter input to the emoticon generation tool, generates the emoticon description content description field prompt, and then generates the emoticon reply content corresponding to the text reply content;

[0067] The emoji generation tool GenerateSticker uses a plug-and-play emoji generation model. It selects emoji generation models of different styles according to different emoji style requirements of users.

[0068] In this embodiment, the 1.6 version of Stable Diffusion is used as the emoticon package generation tool, which can also be replaced by a higher version of Stable Diffusion, or other styles of text graph models such as Midjourney.

[0069] Step 3: Call the evaluation training module and use the modified dialogue dataset to train and evaluate the multimodal base model to obtain a trained multimodal base model;

[0070] Step 3.1: The current user input content X, the conversation history H, and the emoticon content description field Q are transferred to the multimodal base model LLMs to obtain the text reply content Y and the emoticon reply content Z, as shown in the following formula:

[0071] Y,Z=LLMs(X,Q,H)

[0072] The loss function L of the multimodal base model LLMs is:

[0073]

[0074]

[0075] L=L y +L z

[0076] Among them, h is the generated vector feature of the input content after passing through the multimodal base model, θ is a fixed parameter in the large model that does not participate in the update, and E LoRA is a parameter that can be updated during training, y i ,z i The i-th token of the content is generated for the reply content and the emoji content description field Prompt, k is the number of tokens, L y ,L z They are the text reply part loss function in the training process and the emoji content description field Prompt generation part loss function respectively, and L is the final model loss function;

[0077] In this embodiment, the training strategy is to use the LLaMA-factory framework with LoRA parameters to efficiently fine-tune the model, ensure that the original parameters of the multimodal base model are fixed, and only train the newly added parameters. Qwen-VL is used as a pre-trained model, where the hidden layer dimension is 3584, and the input and output vocabularies are both 15K. The initial learning rate is 1e-5, the optimizer is Adam, the batch size is set to 4, the annealing strategy step number is 5K, and all experiments are performed on 8 NVIDIA A6000 GPUs, which takes about 72 hours to reach convergence.

[0078] Step 3.2: Select the pure text evaluation index, emoticon sending frequency evaluation index, and picture-text reply content consistency evaluation index to evaluate the quality and usability of the reply content of the trained multimodal base model, and select the multimodal base model with the best evaluation index as the multimodal base model of the TOOL-STCIKERCONV multimodal empathy reply framework to generate the corresponding text replies and emoticon replies;

[0079] In this embodiment, Rouge-L, BERT-score, BLEU-1 / 2 / 3, METEOR, CIDEr, Dist-1 / 2 / 3 / 4 are selected as the evaluation indicators of the plain text class, Accuracy, F1, Recall, and Precision are selected as the evaluation indicators of the frequency of emoticon package sending, Dist-1 and Dist-2 are selected as the consistency evaluation indicators of the prompt field of the emoticon package content description, Precision, recall, F1, and clip score are selected as multimodal consistency evaluation indicators, and empathy-txt, empathy-mm, and consistency are selected to evaluate the quality and usability of the base model reply content;

[0080] Step 4: Based on the user input, use the trained TOOL-STCIKERCONV multimodal empathy response framework to generate corresponding text responses and emoticon responses.

[0081] This example is based on the original open source data set of STICKERCONV, focusing on the chat content, emoticon description and emotional expression field content. Based on the constructed dialogue data set, the model thinking field though and the emoticon content description field prompt are added to the content corresponding to the "function_call" field. The content corresponding to the though field includes: a fixed starting token " <though> ", fixed prefix "I need to generate a sticker to express", corresponding emotion and fixed end token"< / though> ", "The content corresponding to the "prompt" field is a detailed description of the emoticon package, which serves as a guide for generating emoticon package content for the target emotion. The specific examples of the dialogue dataset are shown in Table 1;

[0082] Table 1 Example of conversation data

[0083]

[0084] In this embodiment, three large language models, Vicuna, ChatGLM3 and Qwen-VL, are selected as base models, and Rouge-L, BERT-score, BLEU-1 / 2 / 3, METEOR, CIDEr, Dist-1 / 2 / 3 / 4 are selected as evaluation indicators for the plain text class; Accuracy, F1, Recall, and Precision are used as evaluation indicators for the frequency of emoticon sending; dist-1 and dist-2 are used as prompt consistency evaluation indicators; precision, recall, f1, and clip score are used as multimodal consistency evaluation indicators; empathy-txt, empathy-mm, and consistency are used as large model evaluation indicators. The evaluation results are shown in Table 2-6. Qwen-VL achieved good performance in most experiments.

[0085] Table 2 Experimental results on text

[0086]

[0087] As shown in Table 2, the TOOL-Qwen-VL model achieves the best results. In terms of specific indicators, the three base models have a small gap in the Dist series evaluation indicators, so the Dist series indicators do not have a strong reference significance. The TOOL-Qwen-VL model has a great advantage in the three indicators of BLEU-1 / 2 / 3. Compared with the TOOL-Vincuna model, the effect improvement has increased from 11.63% to 29.41%, proving that the TOOL-Qwen-VL model has a great advantage in the standardization of generated content. The TOOL-Qwen-VL model has improved by 14.29% in the Rouge-L indicator compared with TOOL-Vicuna. Its indicator meaning is similar to the BLEU value, which also proves the standardization of the TOOL-Qwen-VL model in generating content. In terms of the METEOR indicator, the TOOL-Qwen-VL model has improved by 22.22% compared with TOOL-Vicuna, proving the standardization of the TOOL-Qwen-VL model in vocabulary-level content. Finally, in terms of the CIDEr indicator, the TOOL-Qwen-VL model is 50% higher than the TOOL-Vicuna, indicating that the TOOL-Qwen-VL model has better intent recognition and context understanding capabilities in the multimodal field, and the generated image description content is also closest to the reference content.

[0088] The base model of this embodiment is evaluated in terms of the frequency of sending emoticons, and the results are shown in Table 3:

[0089] Table 3 Experimental results of emoticon sending frequency

[0090]

[0091] As shown in Table 3, the TOOL-Qwen-VL model achieves the best results. In terms of specific indicators, the TOOL-Qwen-VL model has achieved great advantages in Accuracy, F1, Recall and Precision, indicating that the TOOL-Qwen-VL model is most suitable for the sending of emoticons in actual application scenarios. At the same time, the core indicator pred_freq has the smallest error with the standard frequency of 0.42 (a survey in the United States showed that the frequency of college students replying to emoticons when chatting is 0.42), which also proves that the TOOL-Qwen-VL model is closest to the reference content in the distribution of emoticon sending frequency, and can best simulate the frequency of humans sending emoticons during normal chats.

[0092] In this embodiment, the base model is subjected to consistency evaluation of the prompt field of the emoticon package content description and multimodal consistency evaluation. The consistency evaluation experimental results of the prompt field of the emoticon package content description are shown in Table 4, and the multimodal consistency evaluation results are shown in Table 5:

[0093] Table 4 Prompt consistency test results

[0094]

[0095] As can be seen from Table 4, the three models achieved relatively close results in terms of the consistency of the Prompt field in the emoji content description, but the TOOL-Qwen-VL model has some advantages and is closest to the ideal Dist-1 and Dist-2 values.

[0096] Table 5 Multimodal consistency experimental results

[0097]

[0098] As shown in Table 5, the TOOL-Qwen-VL model improves by 11.03%, 25.07% and 8.36% in Precision, Recall and F1 respectively compared with TOOL-ChatGLM3, proving that the TOOL-Qwen-VL model has a higher text-image relevance in its reply content. Finally, in terms of clip-score, the TOOL-Qwen-VL model improves by 2.09% compared with TOOL-Vicuna, proving that the emojis replied by the TOOL-Qwen-VL model have a higher text-image consistency, indicating that the model's ability to understand user emotions and empathize with responses has been improved.

[0099] In this embodiment, a large model scoring evaluation is performed on the base model, and the evaluation results are shown in Table 6.

[0100] Table 6 Large model scoring

[0101]

[0102] As shown in Table 6, the TOOL-Qwen-VL model has improved 2.04% in text empathy ability and 1.32% in multimodal empathy ability compared to TOOL-Vicuna, which proves that under the comprehensive judgment of the large model, the empathy ability of the TOOL-Qwen-VL model has improved compared to TOOL-Vicuna. In addition, in terms of consistency, the TOOL-Qwen-VL model has improved by 1.25% compared to TOOL-Vicuna, which proves that the TOOL-Qwen-VL model's reply content maintains good contextual consistency and has a low probability of generating jumping topics.

[0103] In order to verify the effectiveness of the large model thinking ability proposed in this paper, this embodiment conducted an ablation experiment and selected multimodal consistency for effect comparison. The results are shown in Table 7.

[0104] Table 7 Ablation experiment

[0105]

[0106]

[0107] As shown in Table 7, the TOOL-Qwen-VL model with tool thinking ability has improved by 0.92%, 25.07%, 8.36% and 3.69% in Precision, Recall, F1 and clip-score respectively compared with the control model without tool thinking ability. After adding tool thinking ability, the recall rate of multimodal consistency has increased most significantly, proving that the corresponding multimodal generation ability has been greatly improved, and the image and text generation ability is closer to the actual chat state. In addition, clip-score has also been significantly improved, proving that the introduction of tool thinking ability can improve the consistency of generated images and texts, that is, it can better understand user emotions and generate corresponding emotional responses, and improve the empathy effect during the entire chat process.

[0108] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.

Claims

1. A multimodal empathic response generation method based on tool learning, characterized by: A tool-learning-based multimodal empathy response generation framework TOOL-STCIKERCONV is constructed, which includes a data transformation module, a tool calling module, and an evaluation and training module. The data transformation module is used to transform the constructed dialogue dataset, and call the evaluation and training module to screen and filter the transformed dataset to ensure data quality; the tool calling module decouples the empathy response task into a multimodal small talk subtask based on a multimodal base model and a multimodal emoticon package content description generation subtask based on tool calling; the evaluation and training module uses the transformed dataset to train the multimodal base model, and evaluates the training results to objectively evaluate the quality and usability of the response content; The following steps are involved: Step 1: Construct a conversation dataset and call the data transformation module to transform the constructed conversation dataset, dividing the conversation dataset into a training set, a validation set, and a test set in proportion; Step 2: Call the tool calling module to decouple the empathy reply task into a multimodal small talk subtask based on the multimodal base model and a multimodal emoticon package content description generation subtask based on tool calling, and generate corresponding text reply content and emoticon package reply content respectively; Step 3: Call the evaluation training module and use the modified dialogue dataset to train and evaluate the multimodal base model to obtain a trained multimodal base model; Step 4: Based on the user input, use the trained TOOL-STCIKERCONV multimodal empathy response framework to generate corresponding text responses and emoticon responses.

2. The method for generating multimodal empathic responses based on tool learning according to claim 1, characterized in that: The step 1 comprises: Step 1.1: Obtain the conversation content entered by the user, the description of the emoticon package used by the chatbot to reply, the storage path of the emoticon package, and the text information replied by the chatbot to construct a conversation dataset; Step 1.2: Call the data transformation module to transform the constructed dialogue dataset, reorganize the content and splice the instructions of different fields in the dialogue dataset, and obtain the transformed dialogue dataset; Step 1.3: Call the evaluation training module to screen and filter the transformed data set, and divide the transformed dialogue data set into training set, validation set, and test set in proportion.

3. The method for generating multimodal empathic responses based on tool learning according to claim 2, characterized in that: The conversation data set described in step 1.1 includes: "user" field, "function_call" field, "observation" field and "assistant" field, wherein the "user" field corresponds to the conversation content input by the user, including text and emoticon information, the "function_call" field corresponds to the emoticon description used by the chatbot for reply, the "observation" field corresponds to the storage path of the emoticon used by the chatbot for reply, and the "assistant" field corresponds to the text information replied by the chatbot.

4. The method for generating multimodal empathic responses based on tool learning according to claim 3, characterized in that: The specific method for data transformation of the constructed dialogue dataset described in step 1.2 is: Perform tokenization on the conversation data contained in the constructed dataset, decompose the natural language text into the smallest vocabulary unit token, define the token composed of vocabulary as text token, define the token composed of special symbols <·>,< / ·> The token composed of the text is a special token, which is used to express a specific meaning. The token is converted into the corresponding index token_id under the function of vocabulary mapping, which is used to obtain the word vector corresponding to the token in the pre-training stage. For the user input content contained in the "user" field, use irrelevant tokens to adaptively fill the beginning and end of the user input text, so that the content length contained in the "user" field is consistent after the token is filled. The irrelevant token itself has no meaning and is only used to ensure that the conversation dataset meets the fixed parameter specifications; For the "function_call" field, use the method of inserting special tokens to add the model thinking field though and the emoticon content description field prompt to the content corresponding to the "function_call" field. The content corresponding to the though field includes: fixed starting token" <though> ", fixed prefix, corresponding emotion and fixed end token"< / though> ", the content corresponding to the "prompt" field is a detailed description of the emoticon package, which serves as a guide for generating emoticon package content for the target emotion; <thought> and< / thought> Used to prompt whether to reply to the emoticon package and the content of the emoticon package to be replied at this stage; <thought> and< / thought> The middle part is the process of determining whether an emoticon needs to be output. For the current conversation scenario, the judgment result is that an emoticon picture expressing the corresponding emotion needs to be replied, and a complete description of the content of the emoticon picture is made. The complete description of the emoticon picture content is then converted into token_id, and irrelevant tokens are used for adaptive padding to make the content length corresponding to the function_call field consistent; For the "observation" emoji path field, use irrelevant tokens to adaptively fill it, and keep the emoji path length corresponding to the emoji path field consistent; For the "assistant" plain text reply field, convert the text reply content into the corresponding token_id, and handle the input and output formats so that the content length corresponding to the "assistant" field is consistent.

5. The method for generating multimodal empathic responses based on tool learning according to claim 4, characterized in that: The step 2 comprises: Step 2.1: Call the tool calling module to decouple the empathy reply task into a multimodal small talk subtask based on the multimodal base model and a multimodal emoticon package content description generation subtask based on tool calling, and select the multimodal base model to generate text reply content; Step 2.2 The multimodal base model calls the emoticon generation tool GenerateSticker, regards the conversation context as a parameter input to the emoticon generation tool, generates the emoticon description content description field prompt, and then generates the emoticon reply content corresponding to the text reply content.

6. The method for generating multimodal empathic responses based on tool learning according to claim 5, characterized in that: The specific method of step 2.1 is: In the multimodal small talk subtask based on the multimodal base model, the multimodal base model generates text reply content according to the multimodal dialogue history; In the tool-based multimodal emoji content description generation subtask, the multimodal base model determines whether it is necessary to call the emoji generated by the tool based on the current conversation scene and conversation history, and then calls the emoji generation tool GenerateSticker to generate the emoji reply content corresponding to the text reply content; The multimodal base model generates text reply content according to the multimodal conversation history and determines whether to reply with emoticons. If the multimodal base model determines that the current scene requires an emoticon reply, it enters the multimodal emoticon content description generation subtask based on tool call, and calls the emoticon generation tool to directly generate emoticon reply content corresponding to the text reply content. If the multimodal base model determines that the current scene does not require an emoticon reply, it skips the multimodal emoticon content description generation subtask based on tool call and only generates a text reply.

7. The method for generating multimodal empathic responses based on tool learning according to claim 6, characterized in that: The emoticon generation tool GenerateSticker described in step 2.2 uses a plug-and-play emoticon generation model and selects emoticon generation models of different styles according to different emoticon style requirements of users.

8. The method for generating multimodal empathic responses based on tool learning according to claim 7, characterized in that: The step 3 comprises: Step 3.1: The current user input content X, the conversation history H, and the emoticon content description field Q are transferred to the multimodal base model LLMs to obtain the text reply content Y and the emoticon reply content Z, as shown in the following formula: Y,Z=LLMs(X,Q,H) The loss function L of the multimodal base model LLMs is: L=L y +L z Among them, h is the generated vector feature of the input content after passing through the multimodal base model, θ is a fixed parameter in the large model that does not participate in the update, and E LoRA is a parameter that can be updated during training, y i ,z i The i-th token of the content is generated for the reply content and the emoji content description field Prompt, k is the number of tokens, L y ,L z They are the text reply part loss function in the training process and the emoji content description field Prompt generation part loss function respectively, and L is the final model loss function; Step 3.2: Select pure text evaluation indicators, emoticon sending frequency evaluation indicators, and picture and text reply content consistency evaluation indicators to evaluate the quality and usability of the reply content of the trained multimodal base model, and select the multimodal base model with the best evaluation indicators as the multimodal base model of the TOOL-STCIKERCONV multimodal empathy reply framework, which is used to generate corresponding text replies and emoticon replies.

Citation Information

Patent Citations

  • UniLM model and Copy mechanism-based Chinese emotion-sharing statement training method and system

    CN116150334A

  • Emotion recognition method and system based on voice text cross-modal fusion

    CN117765981A

  • Emotion support dialogue generation method based on multi-clue prompt learning

    CN118410131A

  • Multi-mode psychological counselor skill recommendation method, device and equipment

    CN119202281A

  • Method and system for generating conversation summary

    US11709989B1