Multi-round dialogue reply method and device, equipment and storage medium
By segmenting historical dialogue records and extracting key information, combining natural language understanding and reordering models, the problems of context restriction and intention conversion in multiple rounds of dialogue are solved, and more accurate and smooth dialogue replies are achieved.
Patent Information
- Application Number
- CN202510298753.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-05-30
AI Technical Summary
The existing multi-round dialogue technology relies on LLM and faces issues of contextual restrictions and user intention conversion, resulting in outdated or inaccurate responses.
The historical dialogue records are segmented using preset chapter analysis tools and text segmentation algorithms, and the target key information is extracted and stored in the vector database. Convert user questions through natural language understanding tools, recall relevant text slice key information, and use the reordering model to generate the final answer.
It improves the accuracy and fluency of multiple rounds of conversations, enhances LLM's understanding, optimizes the quality of reply, and creates a more efficient conversation experience.
Smart Images

Figure CN120067272A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing, and particularly to a multi-turn dialogue response method, device, equipment and storage medium. Background Art
[0002] Existing multi-turn dialogue technologies mainly rely on LLM (Large Language Model), which is a major advancement in the field of artificial intelligence. By leveraging deep learning techniques and learning from a large amount of text data, LLM has achieved efficient parsing and generation capabilities for natural language. This enables machines to not only understand the context information in the dialogue but also generate smooth and natural responses based on the dialogue history, thus approaching a more human-like communication style. The application of LLM has significantly improved the quality of human-machine interaction and demonstrated broad application prospects in many fields such as customer service, intelligent assistants, and online education.
[0003] Despite the remarkable achievements of LLM in multi-turn dialogue systems, it still faces some challenges. Firstly, LLM has limitations in context, and it is impossible to use all historical dialogue content as a Prompt for the large model. Secondly, the intention of the user's multi-turn dialogue may change, and using all historical dialogue content as a Prompt for LLM will result in outdated or inaccurate answers. If LLM cannot accurately capture the user's current intention, it will directly affect the user experience.
[0004] To address these issues, RAG (Retrieval Augmented Generation) technology has emerged. As an advanced artificial intelligence solution that combines retrieval and generation capabilities, RAG technology is driving the transformation in the field of multi-turn dialogue. By introducing keyword extraction technology and RAG technology, not only the efficiency of text processing has been improved, but also the understanding ability of LLM has been significantly enhanced, enabling it to more accurately capture the actual needs of users, effectively optimizing the response quality of the system, and paving the way for creating a more smooth and efficient dialogue experience.
[0005] In summary, how to improve multi-turn dialogue ability is an urgent problem to be solved currently. Summary of the Invention
[0006] In view of this, the purpose of the present invention is to provide a multi-turn dialogue response method, device, equipment and storage medium, which can improve multi-turn dialogue ability. The specific solutions are as follows:
[0007] In the first aspect, the present application provides a multi-turn dialogue response method, including:
[0008] Use a preset discourse analysis tool and a preset text segmentation algorithm to segment the historical conversation records to obtain corresponding text slices, extract the target key information in the text slices, and store the target key information in the text slices in a preset vector database;
[0009] Convert the obtained target question into a preset vector form to obtain a target vector question, use a preset natural language understanding tool to perform a preset language understanding operation on the target vector question to obtain a corresponding target question intention, and recall the target key information in the corresponding text slices from the preset vector database according to the target question intention;
[0010] Use a preset re-ranking model to perform a preset re-ranking operation on the target key information in the recalled text slices to obtain sorted key information, and generate a corresponding target answer based on the sorted key information according to a preset prompt word to complete the multi-turn conversation reply.
[0011] Optionally, the using a preset discourse analysis tool and a preset text segmentation algorithm to segment the historical conversation records to obtain corresponding text slices includes:
[0012] Use a preset discourse analysis tool to identify and extract the main logical relationships in the historical conversation records;
[0013] Based on the main logical relationships, merge the conversation records that meet the preset close association conditions;
[0014] Use a preset text segmentation algorithm to segment the historical conversation records according to a preset segmentation standard and the merging result to obtain corresponding text slices; wherein, the preset segmentation standard is a segmentation standard constructed based on the length of the segmented paragraphs and / or the overlapping degree between the segmented paragraphs.
[0015] Optionally, the extracting the target key information in the text slices includes:
[0016] Establish a target two-level index; wherein the target two-level index includes a target first-level index and / or a target second-level index;
[0017] Use the target first-level index to determine the key information from the text slices of the historical conversation records, and perform Embedding processing on the key information to obtain the target key information;
[0018] Use the target second-level index to obtain the original conversation record corresponding to the target key information, and store the original conversation record in the preset prompt word so as to generate a corresponding target answer according to the preset prompt word.
[0019] Optionally, the extracting the target key information in the text slices includes:
[0020] Use the preset constituency parsing and preset named entity recognition techniques to extract the target key information in the text slice; wherein the target key information includes any one or more of noun phrases, verb phrases, and specific entity names.
[0021] Optionally, the using the preset natural language understanding tool to perform a preset language understanding operation on the target vector question to obtain the corresponding target question intention includes:
[0022] Use the preset natural language understanding tool to perform a preset language understanding operation on the target vector question to determine the knowledge domain, intention category, and entity information corresponding to the target vector question;
[0023] Based on the knowledge domain and the intention category, fill the entity information into a preset task template to obtain the corresponding target question intention.
[0024] Optionally, the using the preset re-ranking model to perform a preset re-ranking operation on the target key information in the recalled text slice to obtain the re-ranked key information includes:
[0025] Use the preset re-ranking model to evaluate the target key information in the recalled text slice according to the preset evaluation criteria;
[0026] Based on the obtained evaluation results, perform a preset re-ranking operation on the target key information in the recalled text slice in a preset order to obtain the re-ranked key information.
[0027] Optionally, after generating the corresponding target answer based on the re-ranked key information according to the preset prompt words, it further includes:
[0028] Judge whether the target answer meets the preset rationality verification condition;
[0029] If the target answer does not meet the preset rationality verification condition, then re-execute the step of generating the corresponding target answer based on the re-ranked key information according to the preset prompt words;
[0030] If the target answer meets the preset rationality verification condition, then directly output the target answer, and store the target question and the target answer in the preset vector database.
[0031] In a second aspect, the present application provides a multi-turn dialogue reply device, including:
[0032] An information storage module, configured to use a preset discourse analysis tool and a preset text segmentation algorithm to segment the historical conversation record to obtain corresponding text slices, extract the target key information in the text slices, and store the target key information in the text slices into a preset vector database;
[0033] An information recall module, configured to convert the obtained target question into a preset vector form to obtain a target vector question, use a preset natural language understanding tool to perform a preset language understanding operation on the target vector question to obtain a corresponding target question intention, and recall the target key information in the corresponding text slices from the preset vector database according to the target question intention;
[0034] An answer generation module, configured to use a preset re-ranking model to perform a preset re-ranking operation on the target key information in the recalled text slices to obtain the re-ranked key information, and generate a corresponding target answer based on the re-ranked key information according to a preset prompt word to complete the multi-round conversation reply.
[0035] In a third aspect, the present application provides an electronic device, including:
[0036] A memory, configured to store a computer program;
[0037] A processor, configured to execute the computer program to implement the multi-round conversation reply method as described above.
[0038] In a fourth aspect, the present application provides a computer-readable storage medium, configured to store a computer program; wherein, when the computer program is executed by a processor, the multi-round conversation reply method as described above is implemented.
[0039] In summary, the present application first uses a preset discourse analysis tool and a preset text segmentation algorithm to segment the historical conversation record to obtain corresponding text slices, extracts the target key information in the text slices, and stores the target key information in the text slices in a preset vector database; converts the obtained target question into a preset vector form to obtain a target vector question, uses a preset natural language understanding tool to perform a preset language understanding operation on the target vector question to obtain a corresponding target question intention, and recalls the target key information in the corresponding text slices from the preset vector database according to the target question intention; uses a preset re-ranking model to perform a preset re-ranking operation on the recalled target key information in the text slices to obtain the re-ranked key information, and generates a corresponding target answer based on the re-ranked key information according to a preset prompt word to complete the multi-turn conversation reply. As can be seen from the above, the present application first uses a preset discourse analysis tool and a preset text segmentation algorithm to segment the historical conversation record to obtain text slices, extracts the target key information therein and stores it in a preset vector database. Then, the target question proposed by the user is converted into a preset vector form to obtain a target vector question, and a preset natural language understanding tool is used to perform a preset natural language understanding operation on the target vector question to obtain a corresponding target question intention. The target key information in the text slices related to the target question is recalled from the preset vector database according to the target question intention. Subsequently, a preset re-ranking model is used to perform a re-ranking operation on the recalled target key information, and based on the recalled re-ranked key information, a corresponding target answer is generated according to a preset prompt word to complete the multi-turn conversation reply. In this way, by introducing keyword extraction technology and RAG technology, not only the efficiency of text processing is improved, but also the understanding ability of the LLM is significantly enhanced, enabling it to more accurately capture the actual needs of users, thereby effectively optimizing the reply quality of the system, and also paving the way for creating a more fluent and efficient conversation experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings according to the provided drawings without creative efforts.
[0041] Figure 1 It is a flowchart of a multi-turn conversation reply method disclosed in the present application;
[0042] Figure 2 It is a flowchart of the construction of a specific historical conversation pair record database disclosed in the present application;
[0043] Figure 3A specific flowchart for obtaining key information after sorting disclosed in the present application;
[0044] Figure 4 A specific flowchart for a multi-round dialogue reply method disclosed in the present application;
[0045] Figure 5 A schematic structural diagram of a multi-round dialogue reply device disclosed in the present application;
[0046] Figure 6 A structural diagram of an electronic device disclosed in the present application. Detailed implementation manners
[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0048] Currently, existing multi-round dialogue technologies mainly rely on LLM, which is a major advancement in the field of artificial intelligence. LLM utilizes deep learning technologies and, through learning a large amount of text data, achieves efficient parsing and generation capabilities for natural language. This enables machines not only to understand the context information in the dialogue but also to generate smooth and natural responses based on the dialogue history, thus being closer to the human communication method. The application of LLM has significantly improved the quality of human-machine interaction and demonstrated broad application prospects in many fields such as customer service, intelligent assistants, and online education. Despite the remarkable achievements of LLM in multi-round dialogue systems, it still faces some challenges. First, LLM has limitations on context and it is impossible to submit all historical dialogue contents as Prompts to the large model. Second, the multi-round dialogue intentions of users may change, and submitting all historical dialogue contents as Prompts to LLM will result in outdated or inaccurate answers. If LLM cannot accurately capture the current intentions of users, it will directly affect the user experience. To solve the above technical problems, the present application discloses a multi-round dialogue reply method, device, equipment, and storage medium, which can enhance the multi-round dialogue ability.
[0049] See Figure 1 As shown, the embodiments of the present invention disclose a multi-round dialogue reply method, including:
[0050] Step S11: Use a preset discourse analysis tool and a preset text segmentation algorithm to segment the historical dialogue record to obtain corresponding text slices, extract the target key information in the text slices, and store the target key information in the text slices in a preset vector database.
[0051] In this embodiment, different vector databases are first constructed for different problems so as to store historical conversation records into different preset vector databases. Meanwhile, it is necessary to use a preset discourse analysis tool to identify and extract the main logical relationships in the historical conversation records; merge the conversation records that meet the preset close association conditions based on the main logical relationships; use a preset text segmentation algorithm to segment the historical conversation records according to the preset segmentation criteria and the merging result to obtain corresponding text slices; wherein the preset segmentation criteria are segmentation criteria constructed based on the length of the paragraphs after segmentation and / or the overlapping degree between the paragraphs after segmentation. Specifically, as Figure 2 shown, a discourse analysis tool in natural language processing technology is used to identify and extract the main logical relationships in the historical conversation records, such as subordinate structures. On this basis, paragraphs with close relevance are merged into the same unit to maintain the coherence and integrity of the content. Next, with the help of an effective text segmentation algorithm, the historical conversation records are reasonably divided according to preset criteria, such as paragraph length and the overlapping degree of adjacent paragraphs, to ensure that each segment can independently carry its core meaning and is also convenient for subsequent processing. For different dataset characteristics, it is crucial to select an appropriate chunk size. Too small a chunk may lead to omission of important information, while too large a chunk may introduce unnecessary noise and affect the performance of the model. Therefore, it is necessary to finely optimize and adjust the historical conversation dataset to find the best chunking strategy to ensure that each text segment can discuss around the same theme.
[0052] Furthermore, use preset constituency syntactic analysis and preset named entity recognition technologies to extract the target key information in the text slices; wherein the target key information includes any one or more of noun phrases, verb phrases, and specific entity names. Specifically, by combining constituency syntactic analysis and named entity recognition technologies in the field of natural language processing, the text content can be further refined to extract the target key information therein, including but not limited to noun phrases, verb phrases, and specific entity names, such as currency units, personal names, or company names.
[0053] In addition, to ensure the accuracy of the index, a target two-level index needs to be established; the target two-level index includes a target first-level index and / or a target second-level index; the target first-level index is used to determine key information from the text slices of the historical conversation records, and the key information is processed by Embedding to obtain target key information; the target second-level index is used to obtain the original conversation records corresponding to the target key information, and the original conversation records are stored in the preset prompt word, so as to generate a corresponding target answer according to the preset prompt word. Specifically, a two-level index system is established, where the first-level index records the key information, and the second-level index retains the complete original text, and there is a one-to-one correspondence between the two. In practical applications, only the key information in the first-level index is processed by Embedding and used as the basis for similarity calculation, so as to efficiently screen out the historical conversation records most relevant to the current query. Subsequently, the complete texts corresponding to these records will be integrated into the preset prompt word (Prompt) to further assist in relevance evaluation.
[0054] Step S12: Convert the obtained target question into a preset vector form to obtain a target vector question, use a preset natural language understanding tool to perform a preset language understanding operation on the target vector question to obtain a corresponding target question intention, and recall the target key information in the corresponding text slice from the preset vector database according to the target question intention.
[0055] In this embodiment, first, the target question obtained from the user side is converted into a preset vector form, which usually uses the same means as the vectorization of the historical conversation text to ensure effective comparison with the existing data in the same vector space. Then, NLU (Natural Language Understanding) is used to perform a preset language understanding operation on the target vector question, so that the corresponding target question intention can be obtained. Finally, the target key information in the corresponding text slice is recalled from the preset vector database according to the target question intention.
[0056] In this embodiment, a preset natural language understanding tool can be used to perform a preset language understanding operation on the target vector question to determine the knowledge field, intention category, and entity information corresponding to the target vector question; based on the knowledge field and the intention category, the entity information is filled into a preset task template to obtain a corresponding target question intention.
[0057] In a specific embodiment, according to the user's natural language input, the field or topic that the user is concerned about is determined. For example, if the target question is about a travel-related question, it can be recognized that this is a consultation in the travel field, and then the vector database in the travel field is retrieved. Field recognition helps the dialogue system focus on a specific knowledge field and provide more professional and accurate answers.
[0058] In another specific embodiment, by performing NLU analysis on the user's input, the user's intention, that is, the specific goal that the user hopes to achieve through the dialogue, can be recognized. For example, when the target question is "Dad wants to book a flight to Beijing tomorrow", this sentence will be classified as the intention of "flight booking", and then the vector database related to flight booking is retrieved.
[0059] In a third specific embodiment, NLU is used to extract specific entity information from the user's input, such as person names, locations, times, items, etc. For example, in the sentence "Dad wants to book a flight to Beijing tomorrow", "Dad" and "tomorrow" are respectively recognized as entities of person and time.
[0060] Furthermore, based on the results of the first three steps, the NLU module can fill the recognized entity information into the corresponding slots in the predefined task template. For example, in the flight booking scenario, "Dad" can be filled into the "person" slot, and "tomorrow" can be filled into the "departure time" slot. In this way, it provides a structured input for subsequent dialogue management and action execution, and can obtain the corresponding target question intention based on this information. Next, the target question intention is transformed into a vector form. Then, the vector most similar to the target question intention is searched for in the corresponding vector database, so as to recall a series of relevant text slices.
[0061] Step S13: Use a preset re-ranking model to perform a preset re-ranking operation on the target key information in the recalled text slices to obtain the re-ranked key information, and generate a corresponding target answer based on the re-ranked key information according to a preset prompt word to complete the multi-round dialogue reply.
[0062] In this embodiment, as Figure 3As shown, by using the Rerank model to perform a preset reordering operation on the target key information in the recalled text information, the sorted key information can be obtained. At the same time, use the preset reordering model to evaluate the target key information in the recalled text slices according to the preset evaluation criteria; based on the obtained evaluation results, perform a preset reordering operation on the target key information in the recalled text slices in a preset order to obtain the sorted key information. Specifically, the Rerank model will re-evaluate the quality of each recalled result according to multiple criteria, such as content relevance, authority of information sources, diversity and novelty of results, etc., and adjust its order accordingly. Finally, ensure that the information that best meets the requirements is ranked first to improve the user experience.
[0063] Next, generate a target answer based on the sorted key information according to the Prompt, and determine whether the target answer meets the preset rationality verification conditions; if the target answer does not meet the preset rationality verification conditions, then re-execute the step of generating the corresponding target answer based on the sorted key information according to the preset prompt words; if the target answer meets the preset rationality verification conditions, then directly output the target answer, and store the target question and the target answer in the preset vector database. Specifically, submit the Prompt to the LLM to summarize and generate the target answer. To ensure the efficiency and accuracy of task execution, it is necessary to judge the rationality of the answer and perform loop optimization. First, perform a rationality verification on the target answer generated by the LLM, which is achieved by setting a series of rules, constraints, and secondary reasoning mechanisms. If the target answer does not meet the preset rationality verification conditions, the Prompt will be reorganized or the query strategy will be investigated according to the feedback information, and the target answer will be requested from the large model again until a satisfactory result is obtained. To continuously improve the performance of the system and the richness of the knowledge base, if the target answer meets the preset rationality verification conditions, the target question and the target answer will be added to the vector database. This process not only accumulates more historical conversation records but also provides richer background support for future conversations.
[0064] As can be seen from the above, in the embodiment of the present application, first, a preset discourse analysis tool and a preset text segmentation algorithm are used to segment the historical conversation record to obtain text slices, and the target key information therein is extracted and stored in a preset vector database. Then, the target question proposed by the user is converted into a preset vector form to obtain a target vector question, and a preset natural language understanding tool is used to perform a preset natural language understanding operation on the target vector question to obtain the corresponding target question intention. The target key information in the text slices related to the target question is recalled from the preset vector database according to the target question intention. Subsequently, a preset re-ranking model is used to perform a re-ranking operation on the recalled target key information, and based on the recalled and sorted key information, a corresponding target answer is generated according to the preset prompt words to complete the multi-turn conversation reply. In this way, in the embodiment of the present application, by introducing keyword extraction technology and RAG technology, not only the efficiency of text processing is improved, but also the understanding ability of the LLM is significantly enhanced, enabling it to more accurately capture the actual needs of users, thereby effectively optimizing the reply quality of the system, and also paving the way for creating a more fluent and efficient conversation experience.
[0065] Based on the previous embodiment, it can be known that the present application discloses a multi-turn conversation reply method, which can improve the multi-turn conversation ability. Next, a detailed description will be given of the multi-turn conversation reply optimization method as Figure 4 shown.
[0066] The present application first constructs different preset vector databases, uses a preset discourse analysis tool to identify and extract the main logical relationships in the historical conversation record, such as subordinate structures, and then segments the user's historical conversation record according to the main logical relationships in the historical conversation record to obtain text slices. At the same time, it is also necessary to reasonably divide the text according to preset criteria, such as paragraph length, overlap degree of adjacent paragraphs, etc., to ensure that each segment can independently carry its core meaning. Then, the corresponding target key information is determined in the text slices, and the target key information is stored in the preset vector database.
[0067] Secondly, the target question obtained from the user side is converted into a preset vector form, and then NLU is used to perform a preset language understanding operation on the target vector question, so that the corresponding target question intention can be obtained. Finally, the vector most similar to the target question intention is searched in the corresponding vector database, so as to recall the target key information of a series of related text slices.
[0068] Finally, use the Rerank model to perform a preset reordering operation on the target key information in the recalled text information, and the sorted key information can be obtained. At the same time, use the Rerank model to reorder the target key information in the recalled text slices, select the top k results as the final output to obtain the sorted key information, generate the target answer based on the sorted key information according to the Prompt, and verify the rationality of the target answer, and add the target question and the target answer to the corresponding preset vector database.
[0069] See Figure 5 As shown, an embodiment of the present invention discloses a multi-round dialogue reply device, which may include:
[0070] An information storage module 11, configured to use a preset discourse analysis tool and a preset text segmentation algorithm to segment the historical dialogue record to obtain corresponding text slices, extract the target key information in the text slices, and store the target key information in the text slices in a preset vector database;
[0071] An information recall module 12, configured to convert the obtained target question into a preset vector form to obtain a target vector question, perform a preset language understanding operation on the target vector question by using a preset natural language understanding tool to obtain a corresponding target question intention, and recall the target key information in the corresponding text slices from the preset vector database according to the target question intention;
[0072] An answer generation module 13, configured to use a preset reordering model to perform a preset reordering operation on the target key information in the recalled text slices to obtain sorted key information, and generate a corresponding target answer based on the sorted key information according to a preset prompt word to complete a multi-round dialogue reply.
[0073] As can be seen from the above, in the embodiment of the present application, first, a preset discourse analysis tool and a preset text segmentation algorithm are used to segment the historical conversation record to obtain text slices, and the target key information therein is extracted and stored in a preset vector database. Then, the target question proposed by the user is converted into a preset vector form to obtain a target vector question, and a preset natural language understanding tool is used to perform a preset natural language understanding operation on the target vector question to obtain the corresponding target question intention. The target key information in the text slices related to the target question is recalled from the preset vector database according to the target question intention. Subsequently, a preset re-ranking model is used to perform a re-ranking operation on the recalled target key information, and based on the recalled and sorted key information, a corresponding target answer is generated according to a preset prompt word to complete the multi-round conversation reply. In this way, in the embodiment of the present application, by introducing keyword extraction technology and RAG technology, not only the efficiency of text processing is improved, but also the understanding ability of the LLM is significantly enhanced, enabling it to more accurately capture the actual needs of users, thereby effectively optimizing the reply quality of the system, and also paving the way for creating a more fluent and efficient conversation experience.
[0074] In some specific embodiments, the information storage module 11 may specifically include:
[0075] A logical relationship extraction unit, configured to identify and extract the main logical relationships in the historical conversation record by using a preset discourse analysis tool;
[0076] A record merging unit, configured to merge the conversation records that meet the preset tight association condition based on the main logical relationship;
[0077] A text slice obtaining unit, configured to segment the historical conversation record according to a preset segmentation standard and the merging result by using a preset text segmentation algorithm to obtain corresponding text slices; wherein, the preset segmentation standard is a segmentation standard constructed based on the length of the segmented paragraphs and / or the overlapping degree between the segmented paragraphs.
[0078] In some specific embodiments, the information storage module 11 may specifically include:
[0079] An index establishment unit, configured to establish a target two-level index; wherein the target two-level index includes a target first-level index and / or a target second-level index;
[0080] An information obtaining unit, configured to determine key information from the text slices of the historical conversation record by using the target first-level index, and perform Embedding processing on the key information to obtain target key information;
[0081] A record storage unit for obtaining the original conversation record corresponding to the target key information by using the target secondary index and storing the original conversation record into the preset prompt word, so as to generate a corresponding target answer according to the preset prompt word.
[0082] In some specific embodiments, the information storage module 11 may specifically include:
[0083] A target key information extraction unit for extracting the target key information in the text slice by using a preset constituency syntactic analysis and a preset named entity recognition technology; wherein the target key information includes any one or more of a noun phrase, a verb phrase, and a specific entity name.
[0084] In some specific embodiments, the information recall module 12 may specifically include:
[0085] A preset language understanding operation execution unit for performing a preset language understanding operation on the target vector question by using a preset natural language understanding tool to determine the knowledge domain, intention category, and entity information corresponding to the target vector question;
[0086] A target question intention acquisition unit for filling the entity information into a preset task template based on the knowledge domain and the intention category to obtain a corresponding target question intention.
[0087] In some specific embodiments, the answer generation module 13 may specifically include:
[0088] An information evaluation unit for evaluating the target key information in the recalled text slice according to a preset evaluation criterion by using a preset re-ranking model;
[0089] A sorted key information acquisition unit for performing a preset re-ranking operation on the target key information in the recalled text slice in a preset order based on the obtained evaluation result to obtain sorted key information.
[0090] In some specific embodiments, the multi-turn dialogue reply device may further include:
[0091] A target answer judgment module for judging whether the target answer meets a preset rationality verification condition;
[0092] A step re-execution module for re-executing the step of generating a corresponding target answer according to a preset prompt word based on the sorted key information if the target answer does not meet the preset rationality verification condition;
[0093] An answer storage module, configured to directly output the target answer if the target answer meets a preset rationality verification condition, and store the target question and the target answer in the preset vector database.
[0094] Furthermore, an embodiment of the present application also discloses an electronic device. Figure 6 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure should not be considered as any limitation on the scope of use of the present application.
[0095] Figure 6 It is a schematic structural diagram of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the multi-round dialogue reply method disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0096] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.
[0097] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc., and the resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be temporary storage or permanent storage.
[0098] Among them, the operating system 221 is used to manage and control each hardware device on the electronic device 20 and the computer program 222, and it may be Windows Server, Netware, Unix, Linux, etc. The computer program 222 may further include a computer program capable of completing other specific tasks in addition to the computer program capable of implementing the multi-round dialogue reply method executed by the electronic device 20 disclosed in any of the foregoing embodiments.
[0099] Further, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the multi-round dialogue reply method disclosed above. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.
[0100] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and reference can be made to the description in the method part for related parts.
[0101] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0102] The steps of the method or algorithm described in combination with the embodiments disclosed herein can be directly implemented by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0103] Finally, it should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the presence of additional identical elements in the process, method, article or device including the element.
[0104] The above has introduced the technical solution provided by this application in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A multi-round dialogue reply method, characterized in that: include: Using a preset chapter analysis tool and a preset text segmentation algorithm to segment the historical conversation records to obtain corresponding text slices, extracting target key information from the text slices, and storing the target key information in the text slices in a preset vector database; The acquired target question is converted into a preset vector form to obtain a target vector question, a preset natural language understanding tool is used to perform a preset language understanding operation on the target vector question to obtain a corresponding target question intention, and according to the target question intention, the corresponding target key information in the text slice is recalled from the preset vector database; A preset reordering model is used to perform a preset reordering operation on the target key information in the recalled text slice to obtain the sorted key information, and based on the sorted key information, a corresponding target answer is generated according to the preset prompt words to complete multiple rounds of dialogue responses.
2. The multi-round dialogue reply method according to claim 1, characterized in that: The method of using a preset chapter analysis tool and a preset text segmentation algorithm to segment the historical conversation records to obtain corresponding text slices includes: Use the preset text analysis tool to identify and extract the main logical relationships in the historical dialogue records; Based on the main logical relationship, the conversation records that meet the preset close association condition are merged; The historical conversation records are segmented using a preset text segmentation algorithm according to a preset segmentation standard and a merging result to obtain corresponding text slices; wherein the preset segmentation standard is a segmentation standard constructed based on the length of the segmented paragraphs and / or the degree of overlap between the segmented paragraphs.
3. The multi-round dialogue reply method according to claim 1, characterized in that: The extracting target key information from the text slice includes: Establish a target two-level index; wherein the target two-level index includes a target primary index and / or a target secondary index; Determine key information from text slices of historical conversation records using the target primary index, and perform embedding processing on the key information to obtain target key information; The target secondary index is used to obtain the original conversation record corresponding to the target key information, and the original conversation record is stored in the preset prompt word, so as to generate the corresponding target answer according to the preset prompt word.
4. The multi-round dialogue reply method according to claim 1, characterized in that: The extracting target key information from the text slice includes: The target key information in the text slice is extracted by using preset component syntactic analysis and preset named entity recognition technology; wherein the target key information includes any one or more of noun phrases, verb phrases and specific entity names.
5. The multi-round dialogue reply method according to claim 1, characterized in that: The using a preset natural language understanding tool to perform a preset language understanding operation on the target vector question to obtain a corresponding target question intention includes: Using a preset natural language understanding tool to perform a preset language understanding operation on the target vector question to determine the knowledge domain, intent category, and entity information corresponding to the target vector question; Based on the knowledge domain and the intent category, the entity information is filled into a preset task template to obtain the corresponding target question intent.
6. The multi-round dialogue reply method according to claim 1, characterized in that: The using of a preset reordering model to perform a preset reordering operation on the target key information in the recalled text slice to obtain the sorted key information includes: Using a preset re-ranking model to evaluate the target key information in the recalled text slice according to a preset evaluation criterion; Based on the obtained evaluation results, a preset reordering operation is performed on the target key information in the recalled text slices in a preset order to obtain the sorted key information.
7. The multi-round dialogue reply method according to any one of claims 1 to 6, characterized in that: After generating the corresponding target answer according to the preset prompt words based on the sorted key information, the method further includes: Determine whether the target answer meets the preset rationality verification conditions; If the target answer does not meet the preset rationality verification condition, re-execute the step of generating the corresponding target answer based on the sorted key information according to the preset prompt word; If the target answer meets the preset rationality verification condition, the target answer is directly output, and the target question and the target answer are stored in the preset vector database.
8. A multi-round dialogue reply device, characterized in that: include: An information storage module, used to segment the historical conversation records using a preset chapter analysis tool and a preset text segmentation algorithm to obtain corresponding text slices, extract target key information from the text slices, and store the target key information in the text slices into a preset vector database; An information recall module is used to convert the acquired target question into a preset vector form to obtain a target vector question, use a preset natural language understanding tool to perform a preset language understanding operation on the target vector question to obtain a corresponding target question intent, and recall the corresponding target key information in the text slice from the preset vector database according to the target question intent; The answer generation module is used to use a preset reordering model to perform a preset reordering operation on the target key information in the recalled text slice to obtain the sorted key information, and generate a corresponding target answer based on the sorted key information according to the preset prompt words to complete multiple rounds of dialogue responses.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the multi-round dialogue reply method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: Used to store a computer program; wherein, when the computer program is executed by a processor, the multi-round dialogue reply method as described in any one of claims 1 to 7 is implemented.