Skill Recommendation System for Open-Domain Human-Machine Conversation

The skill recommendation system addresses incoherent responses and disjointed transitions in open-domain conversational agents by using a pre-trained Electra model and prompt learning to ensure coherent and contextually relevant skill transitions.

CN116340488BActive Publication Date: 2025-07-15HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310298629.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-24
Publication Date
2025-07-15
Estimated Expiration
2043-03-24

AI Technical Summary

Technical Problem

The existing open domain human-computer dialogue system may cause incoherent dialogue context and lack specific skill guidance statements when there are errors or ambiguity in user input, affecting the user experience.

Method used

A skill recommendation system for open domain human-computer dialogue is adopted, including a skill recognition module, a chat reply module and a skill recommendation module. The pre-trained Electra model and Bert model are used for semantic representation and skill requirement recognition, and a candidate chat reply is generated by combining the generation and search models, and a smooth skill recommendation reply is generated through prompt learning.

Benefits of technology

Effectively identify the skill needs in user input, avoid incoherence of dialogue, improve user experience, and solve the problem of lack of guidance statements in the existing technology through smooth skill recommendation reply, improving the coherence and richness of dialogue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116340488B_ABST
    Figure CN116340488B_ABST
Patent Text Reader

Abstract

A skill recommendation system for open-domain human-computer dialogue, which belongs to the field of computer artificial intelligence technology. The present invention solves the problems existing in the existing open-domain human-computer dialogue, that is, when there are errors or ambiguous information in the user input, the robot may make a response that is incoherent with the dialogue context, and there is no specific skill guidance statement. The present invention uses a skill recognition module based on weakly supervised learning to recognize the skill requirements in the user input text. The chit-chat reply module generates candidate replies respectively by using generative and retrieval models according to the user input. In the ranking stage, a text relevance scorer based on Bert is used to rank and score the candidate replies, and the reply with the highest score is selected as the optimal chit-chat reply. The skill recommendation module actively recommends appropriate skills according to the optimal chit-chat reply and generates a fluent reply containing the recommended skills. The method of the present invention can be applied to skill recommendation in open-domain human-computer dialogue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer artificial intelligence, and specifically relates to a skill recommendation system for open-domain human-computer dialogue. Background Art

[0002] Human-computer dialogue belongs to the research category of natural language processing in artificial intelligence and is a hot topic in recent years for the research and product implementation of artificial intelligence (Wang Haochang, Li Bin. Research progress of chatbot systems [J]. Computer Applications and Software, 2018, 35(12): 1-6.). A relatively typical application is a chatbot, which can be divided into a domain-limited intelligent dialogue assistant and an open-domain casual dialogue robot according to different scenarios.

[0003] Domain-limited intelligent dialogue assistants such as Microsoft Xiaoice, Google Assistant, Apple Siri, and Amazon Alexa usually help users complete some specific tasks, such as querying itineraries, playing music, and setting reminders. Open-domain chatbots such as Microsoft Xiaobing (Li Zhou, Jianfeng Gao, Di Li, and Heung-Yeung Shum. 2020. The Design and Implementation of XiaoIce, an Empathetic Social Chatbot. Computational Linguistics, 46(1): 53–93.), Harbin Institute of Technology Benben (Wei-Nan Zhang, Ting Liu, Bing Qin, Yu Zhang, Wanxiang Che, Yanyan Zhao, and Xiao Ding. 2017. Benben: A Chinese Intelligent Conversational Robot. In Proceedings of ACL 2017, System Demonstrations, pages 13–18, Vancouver, Canada. Association for Computational Linguistics.), Baidu PLATO-XL (Siqi Bao, Huang He, Fan Wang, Hua Wu, Haifeng Wang, Wenquan Wu, Zhihua Wu, Zhen Guo, Hua Lu, Xinxian Huang, Xin Tian, Xinchao Xu, Yingzhan Lin and Zheng-Yu Niu. 2022. PLATO-XL: Exploring the Large-scale Pre-training of Dialogue Generation. Findings of AACL-IJCNLP 2022, online.). Such chatbots can provide emotional companionship for users. With the development of dialogue systems, open-domain human-computer dialogue has become a research hotspot and has gradually been applied in various fields.Although Baidu's PLATO-XL currently has the best performance in various metrics of open-domain dialogue, it does not possess multiple skills (Siqi Bao, Huang He, Fan Wang, Hua Wu, Haifeng Wang, Wenquan Wu, Zhihua Wu, Zhen Guo, Hua Lu, Xinxian Huang, Xin Tian, Xinchao Xu, Yingzhan Lin and Zheng-Yu Niu. 2022. PLATO-XL: Exploring the Large-scale Pre-training of Dialogue Generation. Findings of AACL-IJCNLP 2022, online.). For Microsoft XiaoIce (in the industry) and HIT Benben (in the academia) which possess multiple skills, their chit-chat response generation technology mainly involves the following two aspects:.

[0004] Skill switching mechanism. The key problem is how to switch to different skills. In this regard, Microsoft XiaoIce analyzes the input information according to the dialogue strategy to judge whether to trigger chit-chat or domain-specific chat. Then it calls the corresponding skill module and switches different skill modules (Li Zhou, Jianfeng Gao, Di Li, and Heung-Yeung Shum. 2020. The Design and Implementation of XiaoIce, an Empathetic Social Chatbot. Computational Linguistics, 46(1): 53–93.). HIT Benben uses a series of predefined rules to trigger multiple modules and finally evaluates the response quality of multiple modules to decide the final response (Li Huan, Xu Hui. Topic Recommendation for Chatbots Based on Collaborative Filtering [J]. Artificial Intelligence and Robotics Research, 2020, 9(2): 154-162.). Both switch to different skills according to the user input of "intelligent". If there are error messages or ambiguous information in the user input, the robot may reply with responses that are incoherent in the dialogue context, affecting the user experience.

[0005] Skill guiding statements. The key issue is how to smoothly guide the conversation content to the corresponding skills. In this regard, after determining the skills to be triggered, both XiaoIce of Microsoft and Benben of Harbin Institute of Technology call the corresponding skill modules to obtain candidate responses, and generate the final response after sorting and other processing. Neither of them has specific skill guiding statements, but directly jumps to the starting response of the corresponding module. A characteristic of casual conversations is that users pour out their hearts and the robot provides company. Switching different skills can enhance the richness of the conversation, but the lack of guiding statements will lead to a large change in information when the topic jumps from casual chat to specific skills. This may make users feel very abrupt and reduce their desire to confide in the dialogue robot.

[0006] In summary, the following main problems exist in the existing open-domain human-machine conversations:

[0007] I. If there are error messages or ambiguous messages in the user input, the robot may make responses that are not coherent with the conversation context;

[0008] II. There are no specific skill guiding statements, but directly jump to the starting response of the corresponding module, which may make users feel very abrupt. Summary of the Invention

[0009] The purpose of the present invention is to solve the problems existing in the existing open-domain human-machine conversations, that is, when there are errors or ambiguous messages in the user input, the robot may make responses that are not coherent with the conversation context, and there are no specific skill guiding statements, and to propose a skill recommendation system for open-domain human-machine conversations.

[0010] The technical solution adopted by the present invention to solve the above technical problems is:

[0011] A skill recommendation system for open-domain human-machine conversations, the skill recommendation system includes a skill recognition module, a casual chat response module, and a skill recommendation module, wherein:

[0012] The skill recognition module is used to extract the semantic representation of the user input text, and then obtain the skill requirement recognition result of the user input text according to the extracted semantic representation;

[0013] The casual chat response module generates candidate casual chat responses according to the user input text, and then selects the optimal casual chat response from the candidate casual chat responses;

[0014] The skill recommendation module generates a skill recommendation response according to the skill requirement recognition result and the optimal casual chat response;

[0015] The working principle of the skill recommendation module is:

[0016] If the recognition result of the skill recognition module indicates that the current user input text contains a skill requirement, a skill recommendation response containing the optimal casual chat response is generated using the prompting learning method;

[0017] If the recognition result of the skill recognition module indicates that there is no skill requirement in the current user input text, the optimal casual chat response is directly used as the response to the current user input text.

[0018] Furthermore, the working process of the skill recognition module is as follows:

[0019] Using the pre-trained Electra model as the encoder, with the user input text as the input of the encoder, the output of the encoder is the semantic representation of the user input text;

[0020] Using a fully connected layer to map the semantic representation to the probability distribution of skill requirements, and in the probability distribution result, the category with the highest probability is used as the skill requirement recognition result of the user input text.

[0021] Furthermore, the Electra model is pre-trained based on weak supervision learning, and the construction method of the training data is as follows:

[0022] Step 1: Obtain an unlabeled text corpus;

[0023] Step 2: For each skill, generate a keyword dictionary for the skill; then use the generated keyword dictionary to extract text from the text corpus obtained in Step 1, respectively obtaining the corpus candidate sets corresponding to each skill;

[0024] Use the RoBERTa model to generate the semantic representations of the text in the corpus candidate sets corresponding to each skill, and then cluster all the semantic representations to delete abnormal text;

[0025] Step 3: Delete the keywords in the remaining text, and then assign a skill requirement as a label to each text respectively. Use the text after deleting the keywords and the corresponding label as part of the training data;

[0026] Step 4: Extract skill-irrelevant corpus from the text corpus obtained in Step 1, and the extracted skill-irrelevant corpus has the same scale as the corpus extracted in Step 2. Assign the label of no skill requirement to the extracted skill-irrelevant corpus, and use the skill-irrelevant corpus and the corresponding label together as another part of the training data;

[0027] Step 5: Use the training data obtained in Step 3 and Step 4 to pre-train the Electra model.

[0028] Furthermore, when clustering all the semantic representations, the K-means algorithm is adopted.

[0029] Furthermore, the recognition result of the skill requirements of the user input text is: the specific skill requirements to which the user input text belongs or there are no skill requirements in the user input text.

[0030] Furthermore, the chat response module includes a recall sub-module and a ranking sub-module, where:

[0031] The recall sub-module generates candidate chat responses according to the user input text by using a generation model and a retrieval model. Among them, the generation model uses a Chinese pre-trained model. The specific process of generating candidate chat responses by using the retrieval model is as follows:

[0032] Step ①: For the retrieval model in the form of user input - user input, if the corpus of the retrieval model stores dialogue data, then select the chat response corresponding to the user input with the most similar semantics to the current user input text from the corpus of the retrieval model;

[0033] For the retrieval model in the form of user input - response, the corpus of the retrieval model only contains responses. Directly match the text with the most similar semantics in the corpus according to the current user input text as the chat response;

[0034] Step ②: Perform word segmentation and vectorization processing on the current user input text in sequence, and then input the processing results into the encoder Bert for encoding. Take the output at the [CLS] position as the semantic vector of the current user input text;

[0035] Calculate the distance between the semantic vector of the current user input text and the semantic vector of each chat response obtained in Step ① respectively, and take the calculated distance as the correlation score between the current user input text and the corresponding chat response;

[0036] Take the chat response with the maximum correlation score and the chat response generated by the generation model as candidate chat responses;

[0037] The ranking sub-module is used to select the optimal chat response from the candidate chat responses. The specific selection process is as follows:

[0038] Concatenate the current user input text with each candidate chat response respectively to obtain the concatenation result corresponding to each candidate chat response; then input each concatenation result into the encoder Bert respectively to obtain the semantic vector corresponding to each concatenation result;

[0039] Then pass the semantic vector corresponding to each concatenation result through a fully connected layer to obtain the correlation score of each concatenation result. Take the obtained correlation score as the correlation score of the candidate chat response corresponding to the concatenation result, and take the candidate chat response with the highest correlation score as the optimal chat response.

[0040] Further, the semantic vectors of the chat responses are obtained as follows:

[0041] After tokenizing and vectorizing each response in the corpus, the processed responses are sequentially input into the encoder Bert to obtain the semantic vectors of each response in the corpus.

[0042] For the chat responses selected in step ①, faiss is used for semantic vector indexing, that is, the semantic vectors corresponding to the chat responses selected in step ① are obtained.

[0043] Furthermore, the Chinese pre-trained model adopted by the generation model is plato-mini.

[0044] The beneficial effects of the present invention are as follows:

[0045] The present invention uses a skill recognition module based on weak supervision learning to recognize the skill requirements in the user input text. When there are errors or ambiguous information in the user input, the present invention can still accurately recognize the skill requirements in the user input text, avoiding incoherent responses in the dialogue context. The chat response module generates candidate responses using generative and retrieval models respectively according to the user input. In the sorting stage, a text relevance scorer based on Bert is used to score and sort the candidate responses, and the response with the highest score is selected as the optimal chat response. The skill recommendation module actively recommends appropriate skills according to the optimal chat response and generates a fluent response containing the recommended skills, solving the problem of no specific skill guidance statements in the existing open-domain human-computer dialogue and improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 This is a Chinese example of the open-domain human-computer dialogue for actively recommending skills according to the present invention;

[0047] Figure 2 This is a Chinese example of the association between the keyword "weather query" and a large number of different user descriptions according to the present invention;

[0048] Figure 3 This is a working schematic diagram of the skill recommendation system for open-domain human-computer dialogue according to the present invention;

[0049] Figure 4 This is a working schematic diagram of the construction of weak supervision training data according to the present invention;

[0050] Figure 5 This is a Chinese example of the skill-specific prompt template according to the present invention;

[0051] Figure 6 This is a working schematic diagram of the similarity calculation of the recall sub-module in the chat response module according to the present invention;

[0052] Figure 7 This is a schematic diagram of the similarity calculation of the sorting sub-module in the chat response module of the present invention. Detailed implementation manners

[0053] Detailed implementation manner 1: In combination with Figure 3 This implementation manner is described. A skill recommendation system for open-domain human-computer dialogue according to this implementation manner, the skill recommendation system includes a skill recognition module, a chat response module, and a skill recommendation module, wherein:

[0054] The skill recognition module is used to extract the semantic representation of the user input text, and then obtain the skill requirement recognition result of the user input text according to the extracted semantic representation;

[0055] The chat response module generates candidate chat responses according to the user input text, and then selects the optimal chat response from the candidate chat responses;

[0056] The skill recommendation module generates a skill recommendation response according to the skill requirement recognition result and the optimal chat response;

[0057] The working principle of the skill recommendation module is:

[0058] If the recognition result of the skill recognition module is that the current user input text contains a skill requirement, the method of prompt learning is used to generate a skill recommendation response including the optimal chat response;

[0059] On the premise of recognizing the skill requirement, the simplest way for skill recommendation is to use a fixed template, but such a response is relatively rigid and cannot be naturally and smoothly integrated into the chat response. The present invention adopts the method of prompt learning, configures a dedicated prompt template for each skill, and guides the model to generate a skill recommendation response that takes into account the chat response.

[0060] Such as Figure 5 shown is an example of the prompt template, which includes dialogue response examples and current session information. The response examples are a small amount of dialogue data for each skill, used to guide the model to generate responses similar to its style. The current session information includes three parts: user input, chat response, and skill requirement. The user input and chat response provide the context of the conversation, and the skill requirement guides the model to reply to relevant text. The dialogue response examples and the current session information have the same structure, which is convenient for the model to identify the rules of the examples and generate responses consistent with the examples.

[0061] If the recognition result of the skill recognition module is that the current user input text has no skill requirement, the optimal chat response is directly used as the response to the current user input text.

[0062] In existing methods, user keywords are mainly used to trigger corresponding skills, mainly involving keyword matching technology, which may not be applicable to open-domain conversations. For example, for the "weather query" skill, the model needs to know which information mentioned by the user is for "weather query" and what the relationship between them is. In open-domain human-computer conversations, the descriptions of users are often diverse. Taking Figure 2 as an example, the number of keywords related to "weather query" is very large. In real-life applications, there may be dozens of modules, and maintaining a large number of keywords is time-consuming and laborious. Regarding the cold start problem in the recommendation system (referring to how to provide accurate recommendations for new users or new products without historical interaction information), in the open-domain conversation scenario, after adding a new skill module to the system, without relevant information on user usage, how to make the model recommend this module is a technical difficulty. Considering the characteristics of open-domain conversations, the present invention proposes a skill recommendation method based on weak supervision learning, modeling the association relationship between the conversation context and candidate skills. After the user input text passes through the skill recognition module, the skill requirement recognition result can be output.

[0063] In existing methods, fixed templates are mainly used to contain skills. This method mainly has two problems: the recommendation language based on fixed templates may make users lose interest; moreover, taking the "weather query" module as an example, if the user input is "going out to play" and the skill recommended by the system is "weather query", there may not be a strong association between the two. Directly replying with "weather query" will seem abrupt and not smooth. Considering the coherence of the conversation context, the present invention proposes a recommendation discourse generation method based on a prompt template, ensuring the relevance between the recommendation language and the recommended skill as well as the diversity of the recommendation language. At the same time, a casual reply that is relatively relevant to both the context and the recommended skill is used to connect the two parts, improving the overall coherence. In Figure 1 the example, "pay attention to the weather conditions" is related to both "going out to play" and the "weather query" module, improving the fluency of the reply.

[0064] In summary, the present invention is mainly aimed at open-domain chatbots with multiple skills (including Microsoft Xiaoice and HIT Benben). The skill recommendation system of the present invention can understand the implicit user needs in chat conversations, evaluate whether other skills of the system can meet this need, and use smooth and appropriate natural language to recommend the user to use the corresponding skill under the condition of judging that it can be met. At the same time, the skill recommendation system of the present invention adopts a pipeline method. The pipeline solution has the advantages of strong interpretability and small data requirements, and can effectively reduce the high cost of data annotation in the process of system implementation.

[0065] Specific Embodiment 2: This embodiment further limits Specific Embodiment 1. The working process of the skill recognition module is as follows:

[0066] Use the pre-trained Electra model as the encoder, take the user input text as the input of the encoder, and the output of the encoder is the semantic representation of the user input text;

[0067] The Electra model draws on the idea of a generative adversarial network. A generator samples and replaces the input text, and a discriminator determines which positions of the tokens have been replaced. Finally, the encoder part of the discriminator is retained for other classification tasks;

[0068] Use a fully connected layer to map the semantic representation to the probability distribution of skill requirements. In the probability distribution result, the category with the highest probability is used as the skill requirement recognition result of the user input text.

[0069] The skill recognition module needs to identify possible skill requirements from the user input text. For example, "Planning to go out on the weekend" implies a need for a weather query skill. However, if keyword detection is used for skill recognition, the skill recommendation process will be rigid and unable to automatically identify implicit skill requirements in the user's conversation. The skill recognition module of the present invention exactly solves this problem.

[0070] Specific Embodiment 3: Combine Figure 4 Describe this embodiment. This embodiment is a further limitation of Specific Embodiment 2. The Electra model is pre-trained based on weak supervision learning, and the construction method of the training data is as follows:

[0071] Step 1: Obtain unlabeled text corpora;

[0072] The text corpus data can come from the Internet, social media, or existing unlabeled conversation datasets. The construction process can be roughly divided into the construction of training corpora for each skill and the construction of corpora without skill requirements;

[0073] Step 2: For each skill, generate a keyword dictionary for the skill; then use the generated keyword dictionary to extract text from the text corpus obtained in Step 1 to obtain the corpus candidate sets corresponding to each skill respectively;

[0074] Taking the "weather query" skill as an example, the keywords related to weather query are "weather query", "check weather", "how's the weather", "weather forecast", etc. Use these keywords that can indicate the skill requirement of weather query to grab relevant text from the corpus. These texts contain rich semantic knowledge of this skill requirement and constitute the corpus candidate set related to the skill.

[0075] Use the RoBERTa model to generate the semantic representations of the texts in the corpus candidate sets corresponding to each skill, and then cluster all the semantic representations to delete abnormal texts;

[0076] Since abnormal samples are likely to be mixed in the candidate corpus obtained by keyword extraction, the present invention clusters all semantic representations through a clustering algorithm, sets the number of categories to the number of skills, and removes the samples far from the cluster center in each skill corpus according to the clustering result, that is, removes the abnormal samples;

[0077] Step 3: Delete the keywords in the remaining text, and then assign skill requirements to each text as labels, and use the text after deleting the keywords and the corresponding labels as part of the training data;

[0078] Step 4: Extract the skill-independent corpus from the text corpus obtained in Step 1. The source of the skill-independent corpus is the complement of the original corpus and the skill-related corpus, and the scale of the extracted skill-independent corpus is the same as that of the corpus extracted in Step 2. Then assign the label of no skill requirement to the extracted skill-independent corpus, and use the skill-independent corpus and the corresponding labels together as another part of the training data;

[0079] Considering that there may be no skill requirement in the conversation, a "no skill requirement" category is added outside all skill requirement categories. This part of the corpus is used to train the model to recognize the context of no skill requirement and prevent the model from over-recommending.

[0080] Step 5: Use the training data obtained in Step 3 and Step 4 to pre-train the Electra model.

[0081] Based on the weakly-supervised learning method instead of the pre-training fine-tuning method, it solves the problem that in the fine-tuning process, it is necessary to label the conversation data with skill requirements, and such data is very scarce.

[0082] Specific Embodiment 4: This embodiment further limits Specific Embodiment 3. The clustering of all semantic representations adopts the K-means algorithm.

[0083] Specific Embodiment 5: This embodiment further limits Specific Embodiment 4. The recognition result of the skill requirement of the user input text is: the specific skill requirement to which the user input text belongs or no skill requirement in the user input text.

[0084] Specific Embodiment 6: This embodiment further limits Specific Embodiment 5. The chatting reply module includes a recall sub-module and a ranking sub-module, where:

[0085] The recall sub-module generates candidate chatting replies according to the user input text by using a generation model and a retrieval model. Among them, the generation model adopts a Chinese pre-trained model. The specific process of generating candidate chatting replies by using the retrieval model is:

[0086] Step ①. For a retrieval model in the form of user input (query) - user input (query), if the corpus of the retrieval model stores dialogue data, select the chat response (response) corresponding to the user input that is semantically most similar to the current user input text from the corpus of the retrieval model;

[0087] For a retrieval model in the form of user input (query) - response (response), the corpus of the retrieval model only contains responses. Directly match the text with the most similar semantics in the corpus according to the current user input text as the chat response;

[0088] Step ②. As Figure 6 shown, perform word segmentation and vectorization processing on the current user input text in sequence, then input the processing results into the encoder Bert for encoding, and use the output at the [CLS] position as the semantic vector of the current user input text;

[0089] Calculate the distances between the semantic vector of the current user input text and the semantic vectors of each chat response obtained in Step ① respectively, and use the calculated distances as the correlation scores between the current user input text and the corresponding chat responses;

[0090] Take the chat response with the highest correlation score and the chat response generated by the generation model as candidate chat responses;

[0091] As Figure 7 shown, the sorting sub-module is used to select the optimal chat response from the candidate chat responses. The specific selection process is as follows:

[0092] Concatenate the current user input text with each candidate chat response respectively to obtain the concatenation result corresponding to each candidate chat response; then input each concatenation result into the encoder Bert respectively to obtain the semantic vector corresponding to each concatenation result;

[0093] The encoder of the sorting sub-module processes the query and the response simultaneously, concatenates the query and the response, so that the information between the two can be fully interacted, and can better capture semantic relevance; in the recall stage, if the query and the response are concatenated, the response cannot be pre-encoded and indexed. In order to improve efficiency, the information interaction between the query and the response must be sacrificed. In the sorting stage, the number of candidate responses has been greatly reduced, so a more accurate correlation scoring model can be used.

[0094] Then, the semantic vector corresponding to each splicing result is passed through a fully connected layer to obtain a relevance score for each splicing result. The obtained relevance score is used as the relevance score of the candidate chat reply corresponding to the splicing result, and the candidate chat reply with the highest relevance score is used as the optimal chat reply.

[0095] The present invention uses a multi-round dialogue data set of Weibo when constructing a training set for the encoder Bert, takes the context of each round of multi-round dialogue as a positive example for training, and constructs appropriate negative examples in the training samples. For the construction of negative examples, the present invention adopts a sampling method within a batch, and all samples except positive examples in the same batch are used as negative examples. The random sampling of the batch ensures the diversity and richness of the negative examples, and better utilizes the data set.

[0096] The sorting submodule can be trained through a binary classification task of related / unrelated, using the probability of correlation between query and response as the correlation score, or it can be trained through a regression task using a conversation dataset annotated with correlation scores. Since there is such conversation data annotated with correlation, in order to make full use of the annotated information in the data, the present invention uses a regression task to train the sorting model.

[0097] Specific implementation method 7: This implementation method is a further limitation of specific implementation method 6. The semantic vector of the chat reply is obtained in the following manner:

[0098] After word segmentation and vectorization processing for each reply in the corpus, each reply is input into the encoder Bert in sequence to obtain the semantic vector of each reply in the corpus.

[0099] For the chat reply selected in step ①, faiss is used to perform semantic vector indexing, that is, the semantic vector corresponding to the chat reply selected in step ① is obtained.

[0100] The calculation of semantic vectors requires a lot of overhead, and there are a large number of candidate response texts in the corpus of the model. It is unreasonable to encode the response during recall. After the encoder training is completed, the present invention completes the generation of semantic vectors for the candidate texts in the corpus in advance, omitting the semantic vector calculation process of the response during recall, thereby improving the efficiency of the recall process.

[0101] Specific implementation eight: This implementation is a further limitation of specific implementation seven, and the Chinese pre-training model used in the generation model is plato-mini.

[0102] At the system level, compared with the existing dialogue system, the present invention can enable the chat dialogue robot to actively recommend skills, which can improve the skill utilization rate of the task-based dialogue system and enhance the user experience.

[0103] At the technical level, compared with traditional supervised skill recommendation models, the weakly supervised learning method adopted by the present invention can better handle the cold start scenario with data scarcity and support the access of new skills. At the same time, the recommendation discourse generation model based on the prompt template improves the richness and fluency of the dialogue and enhances the user experience.

[0104] The present invention can be directly applied to the open-domain chatbot system and is a core module of a chatbot.

[0105] First, the central control module hands over the input and control right to the skill recommendation module. This module recommends appropriate skills according to the input and generates skill recommendation discourse to return to the central control module to complete a skill recommendation task.

[0106] In terms of the deployment method, the technology of the present invention can be independently used as a computing node and deployed on a cloud computing platform. The communication with other modules can be carried out by binding IP addresses and port numbers.

[0107] In the specific implementation of the technology of the present invention, because deep learning related technologies are used, corresponding deep learning frameworks need to be used. The relevant experiments of the technology of the present invention are implemented based on the open-source framework Pytorch. If necessary, other frameworks can be used, such as the equally open-source tensorflow, or PadlePadle used within the enterprise, etc.

[0108] The above examples of the present invention are only used to illustrate in detail the calculation model and calculation process of the present invention, rather than to limit the implementation manner of the present invention. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is impossible to list all the implementation manners here. Any obvious changes or variations derived from the technical solution of the present invention still fall within the protection scope of the present invention.

Claims

1. A skill recommendation system for open-domain human-machine dialogue, characterized in that, The skill recommendation system includes a skill recognition module, a chat response module, and a skill recommendation module, where: The skill recognition module is used to extract the semantic representation of the user input text, and then obtain the skill requirement recognition result of the user input text according to the extracted semantic representation; The working process of the skill recognition module is as follows: Using the pre-trained Electra model as the encoder, taking the user input text as the input of the encoder, and the output of the encoder is the semantic representation of the user input text; Using a fully connected layer to map the semantic representation to the probability distribution of skill requirements. In the probability distribution result, the category with the highest probability is used as the skill requirement recognition result of the user input text; The chat response module generates candidate chat responses according to the user input text, and then selects the optimal chat response from the candidate chat responses; The chat response module includes a recall sub-module and a ranking sub-module, where: The recall sub-module generates candidate chat responses according to the user input text using a generation model and a retrieval model. Among them, the generation model uses a Chinese pre-trained model. The specific process of generating candidate chat responses using the retrieval model is as follows: Step ①: For the retrieval model in the form of user input - user input, the corpus of the retrieval model stores conversation data. Then, select the chat response corresponding to the user input that is most similar in semantics to the current user input text from the corpus of the retrieval model; For the retrieval model in the form of user input - response, the corpus of the retrieval model only contains responses. Directly match the text with the most similar semantics in the corpus according to the current user input text as the chat response; Step ②: Perform word segmentation and vectorization processing on the current user input text in sequence, and then input the processing result into the encoder Bert for encoding. The output at the [CLS] position is used as the semantic vector of the current user input text; Calculate the distance between the semantic vector of the current user input text and the semantic vector of each chat response obtained in Step ①, and use the calculated distance as the correlation score between the current user input text and the corresponding chat response; The chat response with the highest correlation score and the chat response generated by the generation model are used as candidate chat responses; The ranking sub-module is used to select the optimal chat response from the candidate chat responses. The specific selection process is as follows: Concatenate the current user input text with each candidate chat response respectively to obtain the concatenation result corresponding to each candidate chat response; then input each concatenation result into the encoder Bert respectively to obtain the semantic vector corresponding to each concatenation result; Then pass the semantic vector corresponding to each concatenation result through a fully connected layer to obtain the correlation score of each concatenation result. Use the obtained correlation score as the correlation score of the candidate chat response corresponding to the concatenation result, and use the candidate chat response with the highest correlation score as the optimal chat response; The skill recommendation module generates a skill recommendation response according to the skill requirement recognition result and the optimal chat response; The working principle of the skill recommendation module is: If the recognition result of the skill recognition module is that the current user input text contains a skill requirement, a skill recommendation reply containing the optimal casual chat reply is generated using the prompt learning method; If the recognition result of the skill recognition module is that there is no skill requirement in the current user input text, the optimal casual chat reply is directly used as the reply to the current user input text.

2. The skill recommendation system for open-domain human-computer dialogue according to claim 1, wherein The Electra model is pre-trained based on weak supervision learning, and the training data is constructed as follows: Step 1: Obtain an unlabeled text corpus; Step 2: For each skill, generate a keyword dictionary for the skill; then use the generated keyword dictionary to extract texts from the text corpus obtained in Step 1 to obtain a corpus candidate set corresponding to each skill respectively; Use the RoBERTa model to generate semantic representations of the texts in the corpus candidate set corresponding to each skill, and then cluster all the semantic representations to delete abnormal texts; Step 3: Delete the keywords in the remaining texts, and then assign a skill requirement as a label to each text respectively. Use the texts after deleting the keywords and the corresponding labels as part of the training data; Step 4: Extract skill-independent corpus from the text corpus obtained in Step 1, and the extracted skill-independent corpus has the same scale as the corpus extracted in Step 2. Assign a label of no skill requirement to the extracted skill-independent corpus, and use the skill-independent corpus and the corresponding labels together as another part of the training data; Step 5: Use the training data obtained in Steps 3 and 4 to pre-train the Electra model.

3. The skill recommendation system for open-domain human-machine dialogue according to claim 2, wherein, The K-means algorithm is used for clustering all the semantic representations.

4. The skill recommendation system for open-domain human-computer dialogue according to claim 3, wherein The recognition result of the skill requirement of the user input text is: the specific skill requirement to which the user input text belongs or there is no skill requirement in the user input text.

5. The skill recommendation system for open-domain human-machine dialogue according to claim 4, characterized in that The method for obtaining the semantic vector of the casual chat reply is as follows: After performing word segmentation and vectorization processing on each reply in the corpus respectively, input the processed replies into the encoder Bert in sequence to obtain the semantic vector of each reply in the corpus; For the casual chat reply selected in Step ①, use faiss for semantic vector indexing, that is, obtain the semantic vector corresponding to the casual chat reply selected in Step ①.

6. The skill recommendation system for open-domain human-computer dialogue according to claim 5, characterized in that The Chinese pre-trained model used by the generation model is plato-mini.

Citation Information

Patent Citations

  • Multi-round dialogue processing method and device, electronic equipment and storage medium

    CN114722171A

  • Test method and device of dialogue management system, server and storage medium

    CN115562980A