Zero-Fine-Tuning Anthropomorphic Conversation Generation Method and Device Based on Pre-Trained Language Model
Through the zero-fine-tuning anthropomorphic conversation generation method based on the pre-trained language model, the relevant corpus and conversation history are obtained using keywords for concept extraction, which solves the problems of dialogue system deployment complexity and computing resource requirements in the existing technology, and achieves efficient and fast high-quality dialogue generation.
Patent Information
- Application Number
- CN202210315679.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-03-28
AI Technical Summary
The existing technology is difficult to quickly deploy high-quality anthropomorphic dialogue systems, mainly due to the difficulty of integrating high-quality data, the need for a large number of specific domain dialogue corpus, and high computing resource requirements.
The zero-fine-tuning anthropomorphic conversation generation method based on the pre-trained language model is adopted. The concept set is expanded by obtaining keyword-related corpus in the description field, the resources are aggregated to provide relevant knowledge resources, and concept extraction and knowledge resource supplementation are carried out based on the user's conversation history. The guide language is constructed as the input of a large-scale pre-trained language model to generate dialogue replies.
It realizes unsupervised automatic knowledge supplementation and efficient high-quality dialogue generation, reduces the complexity and computing resource requirements of model deployment, and enables rapid deployment of anthropomorphic dialogue systems.
Smart Images

Figure CN114780694B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of dialogue systems, and in particular to a zero-fine-tuning anthropomorphic conversation generation method, device, computer device and storage medium based on a pre-trained language model. Background Technique
[0002] The rise of ultra-large-scale pre-trained language models is considered a paradigm revolution in the field of AI. A large number of application systems for the next generation of artificial intelligence will rely on large models to be built, so as to utilize their super strong modeling capabilities and the massive data knowledge contained in advance. Among these practical applications, building a highly realistic dialogue system that can communicate with humans is an important task that has received particular attention. Especially by introducing some external knowledge related to the current dialogue, so as to complete a dialogue generation with higher information content and more concretization, which has been urgently needed to be put into industrial use. Through investigation, it can be known that this task that has received much attention in recent years is defined as "knowledge-concretized dialogue generation", that is, given the historical content of the dialogue, relevant knowledge resources are searched and selected from an external database for supplementation, and finally a high-quality reply that conforms to the context is generated, which requires full exploration and use of large language models, and high-quality relevant knowledge resources need to be collected.
[0003] However, although there are many existing research results related to this task, such as the large model PLATO-XL dedicated to simulation dialogue, it is still very difficult for developers to actually deploy ultra-large-scale models and build dialogue systems. First of all, the integration of high-quality data for building a highly realistic dialogue system is very difficult. On the one hand, it is not easy to obtain high-quality knowledge itself. And if according to the strategy of existing methods, if the large model is to be fine-tuned, a large amount of specific domain dialogue corpus needs to be prepared additionally. The collection and effective maintenance of these data have increased the complexity of deploying large models to complete realistic dialogue; secondly, existing methods often need to make a trade-off between efficiency and performance. For example, the method Inverse Prompt with excellent results has a very high time cost due to the large model queries that require multiple reverse searches, and the computational resource requirements for fine-tuning ultra-large-scale models are also relatively high. These are the main obstacles for developers to use large models. Summary of the Invention
[0004] The present invention provides a zero-fine-tuning anthropomorphic conversation generation method, device, computer device and storage medium based on a pre-trained language model, aiming to support unsupervised automatic knowledge supplementation and efficient high-quality dialogue generation functions, and facilitate developers to quickly deploy their own anthropomorphic dialogue systems.
[0005] To this end, the first object of the present invention is to propose a zero-fine-tuning anthropomorphic conversation generation method based on a pre-trained language model, including:
[0006] Based on the keywords in the given description domain, obtain the relevant corpus of the keywords, perform concept set expansion, and aggregate resources to provide relevant knowledge resources;
[0007] Based on the user session history, select the conversation turns related to the keywords, perform concept extraction on the conversations in the conversation turns, find relevant resources from the knowledge resource library based on the concepts obtained from the extraction results, splice and integrate the conversation turn text and the resources to construct a guiding statement, which is used as the input of the large-scale pre-trained language model, and the output result is the conversation reply for the corresponding conversation turn.
[0008] Among them, based on the keywords in the given description domain, obtaining the relevant corpus of the keywords, performing concept set expansion, and aggregating resources to provide relevant knowledge resources includes:
[0009] Taking the keywords as seed concepts, obtain the concept descriptions and knowledge resources related to the seed concepts from the external knowledge graph, and at the same time obtain the text content corresponding to the concept descriptions and knowledge resources related to the seed concepts, to obtain the resource data corresponding to the keywords;
[0010] Standardize the format of the collected resource data; among them, the standardized formats include the Q&A pair form and the text description form;
[0011] Perform data expansion on the resource data with standardized format as the expansion candidate set of the resource data;
[0012] Cluster the expansion candidate set, calculate the similarity with the seed concepts in units of the clustering results, and complete the concept ranking.
[0013] Among them, the Q&A pair form is obtained from the collected triple content, and the text description form is obtained from the entity descriptions in the knowledge graph.
[0014] Among them, based on the user session history, selecting the conversation turns related to the keywords, performing concept extraction on the conversations in the conversation turns, finding relevant resources from the knowledge resource library based on the concepts obtained from the extraction results, splicing and integrating the conversation turn text and the resources to construct a guiding statement, which is used as the input of the large-scale pre-trained language model, and the output result is the conversation reply for the corresponding conversation turn, including steps:
[0015] Select k relevant conversation turns from the entire session history as the input basis according to the current user input turn as the query;
[0016] Perform concept extraction on the k relevant conversation turns as the input basis, and find relevant conversation-type resources from the knowledge resource library based on the concepts obtained from the extraction results, and supplement them in the corresponding conversation turns;
[0017] Taking the multi-turn dialogue history as a whole, find the most relevant concepts from the knowledge base and insert their descriptive resources before the chat as the background knowledge of the dialogue;
[0018] Using a large-scale pre-trained language model, take the obtained overall dialogue as the input and use the Beam Search method to generate the response for the current turn of the dialogue.
[0019] Among them, when clustering the extended candidate set, the K-means clustering algorithm is adopted, and the similarity calculation formula is as follows:
[0020]
[0021] where cosine represents the cosine similarity, s k represents a specific seed concept, represents the clustering category.
[0022] Among them, selecting k relevant dialogue turns from the entire conversation history includes the steps of:
[0023] For the sentences in the entire conversation history, use SentenceBert as the encoder to map them into 768-dimensional space vectors;
[0024] Calculate the similarity based on the following formula;
[0025] α t-i *cosine([U i ; S i , U t )
[0026] where cosine represents the cosine similarity, α = 0.7, representing the distance attenuation coefficient, U i and S i are the statements in the dialogue, corresponding to the dialogue statements generated by the user and the dialogue system in the i-th turn, and U i represents the input statement of the user in this turn.
[0027] Among them, based on the relevant dialogue turns selected from the historical conversation, use a named entity recognition tool to extract the corresponding concepts, search for the corresponding concepts from the relevant knowledge resources, obtain a triple dialogue, and insert it into the conversation content.
[0028] The second object of the present invention is to propose a zero-fine-tuning anthropomorphic conversation generation device based on a pre-trained language model, including:
[0029] An offline knowledge acquisition module, used to obtain the relevant corpus of the given keywords describing the domain, perform concept set expansion, and aggregate resources to provide relevant knowledge resources;
[0030] An online dialogue generation module, which is used to select the conversation turns related to the keyword from the user conversation history, extract concepts from the conversations in the conversation turns, find relevant resources from the knowledge resource library based on the concepts obtained from the extraction results, splice and integrate the conversation turn text and the resources to construct a guiding statement, use it as the input of a large-scale pre-trained language model, and the output result is the dialogue response for the corresponding conversation turn.
[0031] The third object of the present invention is to propose a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method of the foregoing technical solution is implemented.
[0032] The fourth object of the present invention is to propose a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method of the foregoing technical solution is implemented.
[0033] Different from the prior art, the zero-fine-tuning anthropomorphic conversation generation method based on a pre-trained language model provided by the present invention obtains relevant corpora of the keyword based on the given keyword describing the domain, and performs concept set expansion, aggregates resources to provide relevant knowledge resources; based on the user conversation history, selects the conversation turns related to the keyword therefrom, extracts concepts from the conversations in the conversation turns, finds relevant resources from the knowledge resource library based on the concepts obtained from the extraction results, splices and integrates the conversation turn text and the resources to construct a guiding statement, uses it as the input of a large-scale pre-trained language model, and the output result is the dialogue response for the corresponding conversation turn. Through the present invention, it is possible to automatically construct a large model guiding statement template suitable for multi-round realistic conversations, and generate a dialogue result based on the large model. Description of the Drawings
[0034] The aspects and advantages of the present invention and / or additional aspects will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, wherein:
[0035] Figure 1 is a schematic flowchart of a zero-fine-tuning anthropomorphic conversation generation method based on a pre-trained language model provided by the present invention.
[0036] Figure 2 is a schematic diagram of historical conversation input in a zero-fine-tuning anthropomorphic conversation generation method based on a pre-trained language model provided by the present invention.
[0037] Figure 3 is a schematic diagram of the modified historical conversation input in a zero-fine-tuning anthropomorphic conversation generation method based on a pre-trained language model provided by the present invention.
[0038] Figure 4It is a schematic structural diagram of a zero-fine-tuning anthropomorphic conversation generation device based on a pre-trained language model provided by the present invention.
[0039] Figure 5 It is a schematic structural diagram of a non-transitory computer-readable storage medium provided by the present invention. Detailed implementation manners
[0040] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, in which the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as a limitation to the present invention.
[0041] Figure 1 It is a zero-fine-tuning anthropomorphic conversation generation method provided by an embodiment of the present invention. The method includes the following steps:
[0042] Step 101, based on the given keywords describing the field, obtain the relevant corpus of the keywords, and perform concept set expansion to aggregate resources to provide relevant knowledge resources.
[0043] The present invention aims to construct a multi-turn chat conversation for a robot, and defines the problem to be solved as:
[0044] The input is a t-turn conversation history where U t and S t are both sentences, corresponding to the conversation content of the user and the system in the i-th turn respectively, where U t is also the user's question in this turn, and the system needs to output a machine-generated reply S for this turn based on external knowledge resources t , where represents a series of external resources related to the conversation.
[0045] Step 101 specifically includes:
[0046] Taking the keywords as seed concepts, obtaining the concept descriptions and knowledge resources related to the seed concepts from an external knowledge graph, and at the same time obtaining the text content corresponding to the concept descriptions and knowledge resources related to the seed concepts, to obtain the resource data corresponding to the keywords;
[0047] Normalize the format of the collected resource data; among them, the normalized format includes the Q&A pair form and the text description form; the Q&A pair form is obtained from the collected triple content, and the text description form is obtained from the entity description in the knowledge graph.
[0048] Perform data expansion on the format-standardized resource data to obtain an expansion candidate set of the resource data;
[0049] Cluster the expansion candidate set, calculate the similarity with the seed concept for each clustering result unit, and complete the concept ranking.
[0050] Step 102: Based on the user session history, select the conversation turns related to the keyword, extract concepts from the conversations in the conversation turns, find relevant resources from the knowledge resource library based on the concepts obtained from the extraction results, splice and integrate the conversation turn text and the resources to construct a guiding statement, which is used as the input of the large-scale pre-trained language model, and the output result is the conversation response for the corresponding conversation turn
[0051] Specifically include:
[0052] Select k relevant conversation turns from the entire session history according to the current user input turn as the input basis;
[0053] Extract concepts from the k relevant conversation turns used as the input basis, and find relevant dialogue resources from the knowledge resource library based on the concepts obtained from the extraction results, and supplement them in the corresponding conversation turns;
[0054] Take the multi-turn conversation history as a whole, find the most relevant concept from the knowledge base, and insert its descriptive resources before the chat as the dialogue background knowledge;
[0055] Use the large-scale pre-trained language model, take the obtained overall conversation as the input, and use the Beam Search method to generate the conversation response for the current turn.
[0056] When clustering the expansion candidate set, use the K-means clustering algorithm, and the similarity calculation formula is as follows:
[0057]
[0058] Where cosine represents the cosine similarity, s k represents a specific seed concept, represents the clustering category.
[0059] Selecting k relevant conversation turns from the entire session history includes the steps of:
[0060] For the sentences in the entire session history, use SentenceBert as the encoder to map them into 768-dimensional space vectors;
[0061] Calculate the similarity based on the following formula;
[0062] α t-i*cosine([U i ; S i , U t )
[0063] where cosine represents cosine similarity, α = 0.7, representing the distance attenuation coefficient, U i and S i are the statements in the conversation, corresponding to the conversation statements generated by the user and the dialogue system in the i-th round, and U i represents the input statement of the user in this round.
[0064] Based on the relevant conversation turns selected from the historical conversation, use the named entity recognition tool to extract the corresponding concepts, search for the corresponding concepts from the relevant knowledge resources, obtain the triple conversation, and insert it into the conversation content.
[0065] Next, the implementation method of completing zero-shot fine-tuning and realistic dialogue generation based on the pre-trained language model will be introduced in detail. In this embodiment, taking the open dialogue scenario of the sports theme in the Chinese context as an example, k is set to 2, and the Chinese GLM model is used as the pre-trained model, which is a generative pre-trained language model with 10 billion parameters.
[0066] Given {skiing, figure skating, short track speed skating} as seeds, the content of the seeds is obtained from two directions: 1) Link these entries to the large-scale encyclopedic knowledge graph Xlore2, and for example, for "skiing", obtain its corresponding triple and text content; 2) Search these entries using the Bing search engine to obtain the corresponding text content.
[0067] Standardize the format of the resource data, and the specific method is as follows:
[0068] Question-answer pair format: 1) Rewrite the collected triple content into the form of question-answer pairs using a rule-based method. For example, <skiing, country of origin, Country A> is transcribed as "Q: What is the country of origin of skiing? A: The country of origin of skiing is Country A." and collect it accordingly; 2) Use the open-source question generation tool based on T5 to complete the conversion of the declarative sentences containing the seed concepts. For example, the entity description text can generate the question "What is skiing?"
[0069] Description text format: Save the entity description and the text paragraphs containing the seed concepts in the search engine.
[0070] These formats of resources are indexed and constructed through their corresponding seed concepts, so that they can be conveniently queried using Elastic Search.
[0071] For the text corresponding to these seed concepts, use NER tools and encyclopedia entries to discover the knowledge concepts they contain. For example, in the skiing interface, concepts such as "skis", "slalom", "ski poles", "alpine skiing", "metal materials", etc. can be found. These concepts are temporarily stored as candidates, but not all of them should be retained.
[0072] For these obtained candidate concepts, cluster them. In this implementation, these candidate concepts are aggregated into 15 categories using K-means, and each category is represented by . The confidence of the concepts in these categories can be obtained by the following formula, where cosine represents the cosine similarity, and s k represents a specific seed concept. Finally, retain the candidate concepts in the cluster with the highest score, thereby expanding the seed concept set. Concepts such as "skis", "slalom", "ski poles", "alpine skiing", etc. in this example are retained after being clustered into one category, while concepts such as "metal materials", "plastics", etc. are excluded.
[0073] Select historical conversations, as Figure 2 shown, including 6 rounds of conversation history. The seventh-round question "Then how do skiers train every day?" is the current user query, and the system's goal is to give a suitable response to this question.
[0074] All sentences use SentenceBert as the encoder and are mapped into a 768-dimensional vector space. Then, complete the similarity calculation based on the following formula α t-i *cosine([U i ; S i , U t ), where α = 0.7. According to the current question sentence, this patent selects the 2 most relevant historical conversations as the conversation history to retain, namely Q4, S4 and Q5, S5.
[0075] Based on the selected 2 groups of conversations, extract the corresponding concepts respectively. Using entities, Q4, S4 and Q5, S5 respectively obtain the concept "skiing". In this way, from the offline resource library, use Elastic Search that has been indexed before to search for "skiing", and obtain its relevant triple conversations, and insert them before the conversation content, that is, Q: What is skiing? A: Skiing is a sport in which athletes attach skis to the soles of their boots...
[0076] Taking the multi-turn conversation history as a whole, search for the most relevant concepts from the knowledge base, and insert their descriptive resources before the chat as the background knowledge of the conversation. Since the Q&A history mentions not only "skiing" and "freestyle skiing" but also concepts such as "athlete", through similarity calculation, the most relevant concept is obtained as "Winter Games", and its description is placed at the forefront of the conversation as the background. Thus, the input to the large model is as follows Figure 3 .
[0077] Using the large-scale pre-trained language model GLM, take the obtained overall conversation as the input, and use methods such as Beam Search to generate the response for the current turn of the conversation. Thus, the corresponding response result is obtained. In this example, the obtained response result is "They all exercise according to different skiing environments every day. When skiing, they also need to continuously adjust their speed and sliding direction, which requires athletes to have a solid foundation and good coordination ability. Coupled with continuous learning and training, they can acquire good skills."
[0078] Collect in the open-domain conversation scenario and two specific domains of "travel" and "sports". 75,000 conversation contents are collected for open-domain Q&A, and 6,000 conversation contents are collected for domain-specific scenarios, and manual evaluation is completed from five dimensions: coherence, coordination, informativeness, hallucination, and attractiveness.
[0079] The experimental results show that without model training or fine-tuning, this method can achieve comparable results to existing specifically trained models in dimensions such as coherence, coordination, and attractiveness, and has a significant advantage in informativeness, exceeding the current algorithm by about 30%.
[0080] In addition, based on the actual machine test effect of deploying and online interface calling of existing methods, the response speed of this method is close to that of the large model's single query, exceeding the response speed of other existing controllable generation algorithms by about 50%.
[0081] In addition, as Figure 4 shown, the present invention provides a zero-fine-tuning anthropomorphic conversation generation device based on a pre-trained language model, including:
[0082] An offline knowledge acquisition module 310, configured to obtain the relevant corpus of the keyword based on the given keyword describing the domain, perform concept set expansion, and aggregate resources to provide relevant knowledge resources;
[0083] The online conversation generation module 320 is used to select the conversation turns related to the keyword from the user conversation history, extract concepts from the conversations in the conversation turns, find relevant resources from the knowledge resource library based on the concepts obtained from the extraction results, splice and integrate the conversation turn text and the resources to construct a guiding statement, which is used as the input of the large-scale pre-trained language model, and the output result is the conversation reply for the corresponding conversation turn.
[0084] To implement the embodiment, the present invention also proposes another computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it realizes the zero-shot anthropomorphic conversation generation as in the embodiment of the present invention.
[0085] As Figure 5 shown, the non-temporary computer-readable storage medium includes a memory 810 with instructions, an interface 830, and the instructions can be executed by a processor 820 according to zero-shot anthropomorphic conversation generation to complete the method. Optionally, the storage medium can be a non-temporary computer-readable storage medium. For example, the non-temporary computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0086] To implement the embodiment, the present invention also proposes a non-temporary computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the zero-shot anthropomorphic conversation generation as in the embodiment of the present invention.
[0087] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0088] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of the features. In the description of the present invention, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically and clearly defined.
[0089] Any process or method description represented in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logic function or process. The scope of the preferred embodiments of the present invention includes additional implementations where functions may be executed not in the order shown or discussed, including in a substantially simultaneous manner according to the functions involved or in a reverse order, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0090] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered a sequenced list of executable instructions for implementing a logical function and can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. As used in this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection having one or more wires (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.
[0091] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.
[0092] Those of ordinary skill in the art can understand that all or part of the steps carried out in implementing the method of the above-described embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0093] In addition, in each of the embodiments of the present invention, the functional units can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0094] The storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present invention have been shown and described above, it can be understood that the embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the embodiments within the scope of the present invention.
Claims
1. A zero-fine-tuning anthropomorphic conversation generation method based on a pre-trained language model, characterized in that, it includes: Based on the given keywords describing the domain, obtain the relevant corpus of the keywords, and perform concept set expansion, aggregating resources to provide relevant knowledge resources, including: using the keywords as seed concepts, obtaining concept descriptions and knowledge resources related to the seed concepts from an external knowledge graph, and at the same time obtaining the text content corresponding to the concept descriptions and knowledge resources related to the seed concepts, to obtain the resource data corresponding to the keywords; standardize the format of the collected resource data; wherein, the standardized format includes question-and-answer pair form and text description form; expand the format-standardized resource data as the extended candidate set of the resource data; cluster the extended candidate set, and calculate the similarity with the seed concepts in units of the clustering results to complete concept ranking. Based on the user conversation history, select the conversation turns related to the keywords, extract concepts from the conversations in the conversation turns, find relevant resources from the knowledge resource library based on the concepts obtained from the extraction results, and splice and integrate the conversation turn text and the resources to construct a guiding statement as the input of the large-scale pre-trained language model, and the output result is the conversation reply for the corresponding conversation turn, including: select k relevant conversation turns from the entire conversation history as the input basis according to the current user input turn as the query; extract concepts from the k relevant conversation turns as the input basis, and find relevant dialogue resources from the knowledge resource library based on the concepts obtained from the extraction results, and supplement them in the corresponding conversation turns; regard the multi-turn conversation history as a whole, find the most relevant concepts from the knowledge base, and insert their descriptive resources before the chat as the dialogue background knowledge; use the large-scale pre-trained language model, use the obtained overall conversation as the input, and use the Beam Search method to generate the conversation reply for the current turn.
2. The zero-fine-tuning anthropomorphic conversation generation method based on a pre-trained language model according to claim 1, characterized in that, the question-and-answer pair form is obtained from the collected triple content, and the text description form is obtained from the entity description in the knowledge graph.
3. The zero-fine-tuning anthropomorphic conversation generation method based on a pre-trained language model according to claim 1, characterized in that, when clustering the extended candidate set, the K-means clustering algorithm is used, and the similarity calculation formula is as follows: where cosine represents cosine similarity, and s k represents a specific seed concept, and c j represents the clustering category.
4. The zero-fine-tuning anthropomorphic conversation generation method based on a pre-trained language model according to claim 2, characterized in that, selecting k relevant conversation turns from the entire conversation history includes the steps of: For the sentences in the entire conversation history, use SentenceBert as the encoder and map them into a 768-dimensional space vector; Calculate the similarity based on the following formula; α t-i *cosine([U i ; S i , U t ) Among them, cosine represents the cosine similarity, α = 0.7, representing the distance attenuation coefficient, U i and S i are the statements in the dialogue, corresponding to the dialogue statements generated by the user and the dialogue system in the i-th round. U t represents the input statement of the user in this round.
5. The zero-fine-tuning anthropomorphic conversation generation method based on a pre-trained language model according to claim 4, characterized in that, Based on relevant dialogue turns selected from historical conversations, use a named entity recognition tool to extract corresponding concepts therein, search for the corresponding concepts from relevant knowledge resources, obtain triple conversations, and insert them into the conversation content.
6. A zero-fine-tuning anthropomorphic conversation generation device based on a pre-trained language model, characterized in that, comprising: An offline knowledge acquisition module, which is used to obtain relevant corpora of the keywords based on the given keywords describing the domain, perform concept set expansion, and aggregate resources to provide relevant knowledge resources, including: using the keywords as seed concepts, obtaining concept descriptions and knowledge resources related to the seed concepts from an external knowledge graph, and at the same time obtaining the text content corresponding to the concept descriptions and knowledge resources related to the seed concepts to obtain the resource data corresponding to the keywords; standardizing the format of the collected resource data; wherein, the standardized format includes the form of question-and-answer pairs and the form of text descriptions; expanding the resource data with standardized format as the expansion candidate set of the resource data; clustering the expansion candidate set, calculating the similarity with the seed concepts in units of the clustering results, and completing concept ranking; An online dialogue generation module, which is used to select the conversation turns related to the keywords from the user conversation history, extract concepts from the conversations in the conversation turns, find relevant resources from the knowledge resource library based on the concepts obtained from the extraction results, splice and integrate the conversation turn text and the resources to construct a guiding language, and use it as the input of a large-scale pre-trained language model, and the output result is the dialogue reply for the corresponding conversation turn, including: selecting k relevant conversation turns from the entire conversation history as the input basis according to the current user input turn as a query; extracting concepts from the k relevant conversation turns as the input basis, and finding relevant dialogue-type resources from the knowledge resource library based on the concepts obtained from the extraction results, and supplementing them in the corresponding conversation turns; regarding the multi-turn conversation history as a whole, finding the most relevant concepts from the knowledge base and inserting their descriptive resources before the chat as the dialogue background knowledge; using a large-scale pre-trained language model, using the obtained overall conversation as the input, and using the Beam Search method to generate the dialogue reply for the current turn.
7. A computer device, characterized in that, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the method described in any one of claims 1-5 is implemented.
8. A non-transitory computer-readable storage medium, on which a computer program is stored, characterized in that, when the computer program is executed by a processor, the method described in any one of claims 1-5 is implemented.
Citation Information
Patent Citations
Man-machine conversation method based on artificial intelligence, and model training method and device
CN111309883A