Multi-round dialogue processing method and device and storage medium
The method improves dialogue systems by using a vehicle-tailored language model to ensure relevant responses through keyword matching and historical context, addressing answer mismatches and enhancing user experience.
Patent Information
- Application Number
- CN202410051391.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-12
- Publication Date
- 2025-07-15
AI Technical Summary
In the existing multi-round dialogue system, the question-and-answer library matching mechanism and task-based dialogue implemented by the card slot are prone to answer questions that are not asked or cannot be answered, resulting in poor user experience.
The preset thesaurus matches the target keywords to determine the answer sentences. When there is no match, the answers are generated through historical dialogue and the language model of the in-vehicle text language style to improve the relevance and flexibility of the dialogue scene.
Improve the flexibility and user experience of multiple rounds of conversations, avoiding the limited number of question-and-answer libraries or the answers to unquestioned questions after task-based conversations are separated from the task.
Smart Images

Figure CN120316205A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a multi-turn dialogue processing method, apparatus, and storage medium. Background Art
[0002] With the continuous development and application of artificial intelligence technology, more and more robots and intelligent devices have entered our lives. One important application field is the multi-turn dialogue system. The multi-turn dialogue system is an artificial-intelligence-based natural interaction system that can simulate the natural language interaction process and achieve intelligent dialogue between humans and machines.
[0003] Currently, in the multi-turn dialogue system, a simple Q&A library matching mechanism is usually used to implement multi-turn dialogue, or task-based dialogue based on slots. Among them, for the multi-turn dialogue implemented by the Q&A library matching mechanism, due to the limited number of answer sentences in the Q&A library, there may be situations where the answer is off-topic or there is no response during the dialogue process. In the task-based dialogue based on slots, the interaction between the user and the device is usually based on a specific task. After the task is removed, there may also be problems such as the answer being off-topic or there being no response. Summary of the Invention
[0004] To solve the above technical problems, this application provides a multi-turn dialogue processing method, apparatus, and storage medium, which can improve the flexibility of multi-turn dialogue.
[0005] In a first aspect, this application provides a multi-turn dialogue processing method, including: obtaining the user's current dialogue sentence; when a target keyword is matched in a preset word library according to the current dialogue sentence, using the target keyword and a preset correspondence to determine the final answer sentence; the preset word library includes at least one keyword in the in-vehicle text language style, and the preset correspondence is the correspondence between the keyword and the answer sentence; when a target keyword is not matched in the preset word library according to the current dialogue sentence, obtaining the historical dialogue sentence; extracting summary information from the historical dialogue sentence, and generating the final answer sentence by a language model using the summary information; the language model is a language model pre-trained according to the text in the in-vehicle text language style.
[0006] In a second aspect, the present application provides a multi-turn dialogue processing device, including: an acquisition module configured to acquire the current dialogue statement of the user; a determination module configured to, when a target keyword is matched in a preset word library according to the current dialogue statement, determine a final answer statement using the target keyword and a preset correspondence relationship; the preset word library includes at least one keyword in the in-vehicle text language style, and the preset correspondence relationship is the correspondence relationship between the keyword and the answer statement; the acquisition module is further configured to acquire historical dialogue statements when no target keyword is matched in the preset word library according to the current dialogue statement; a processing module configured to extract summary information from the historical dialogue statements, and a language model generates a final answer statement using the summary information; the language model is a language model pre-trained according to texts in the in-vehicle text language style.
[0007] In a third aspect, the present application provides an electronic device, including: a processor, a memory, and a computer program stored on the memory and executable on the processor, where the computer program, when executed by the processor, implements the multi-turn dialogue processing method as in the first aspect.
[0008] In a fourth aspect, the present application provides a computer-readable storage medium, including: a computer program stored on the computer-readable storage medium, where the computer program, when executed by the processor, implements the multi-turn dialogue processing method as in the first aspect.
[0009] In a fifth aspect, the present application provides a computer program product, including: when the computer program product runs on a computer, enabling the computer to implement the multi-turn dialogue processing method as in the first aspect.
[0010] The technical solution provided by this application has the following advantages compared with the prior art: First, obtain the user's current conversation statement. Then, when a target keyword is matched in the preset word library according to the current conversation statement, use the target keyword and the preset correspondence to determine the final answer statement, where the preset word library includes at least one keyword in the in-vehicle text language style, and the preset correspondence is the correspondence between the keyword and the answer statement. Then, when no target keyword is matched in the preset word library according to the current conversation statement, obtain the historical conversation statement and determine the final answer statement according to the historical conversation and the language model. The language model is a language model pre-trained according to the text in the in-vehicle text language style. In this way, since the keywords are keywords in the in-vehicle text language style, that is, the matched keywords and answer statements are all related to the vehicle, it is possible to determine whether the conversation scenario is related to the vehicle according to the keywords in the current conversation statement, and when the conversation scenario is related to the vehicle, the final answer statement can be determined according to the historical statement and the language model. It avoids the multi-round conversation implemented through the Q&A library matching mechanism. Due to the limited number of answer statements in the Q&A library, or in the task-based conversation implemented based on the card slot, the interaction between the user and the device is based on a specific task. After the task is separated, it may lead to irrelevant answers or inability to answer during the conversation process, improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0012] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0013] Figure 1 It is a schematic diagram of the application scenario of the multi-round conversation processing method provided by the embodiment of this application;
[0014] Figure 2 It is one of the flow diagrams of the multi-round conversation processing method provided by the embodiment of this application;
[0015] Figure 3 It is one of the architecture diagrams of the multi-round conversation processing method provided by the embodiment of this application;
[0016] Figure 4 It is the second architecture diagram of the multi-round conversation processing method provided by the embodiment of this application;
[0017] Figure 5The second flowchart of the multi-turn dialogue processing method provided by the embodiment of the present application;
[0018] Figure 6 The third flowchart of the multi-turn dialogue processing method provided by the embodiment of the present application;
[0019] Figure 7 The structural schematic diagram of a multi-turn dialogue processing device provided by the embodiment of the present application;
[0020] Figure 8 The structural schematic diagram of an electronic device provided by the embodiment of the present application. Detailed implementation manners
[0021] In order to more clearly understand the above objects, features, and advantages of the present application, the solutions of the present application will be further described below. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other.
[0022] Many specific details are set forth in the following description in order to fully understand the present application, but the present application may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present application, rather than all the embodiments.
[0023] Figure 1 The scenario architecture schematic diagram of a multi-turn dialogue processing method provided by the embodiment of the present application. As Figure 1 shown, the scenario architecture provided by the embodiment of the present application includes: a server 100 and an electronic device 200.
[0024] The electronic device 200 provided by the embodiment of the present application may have various implementation forms. For example, it may be an in-vehicle device, a mobile phone, a personal computer (PC), a smart TV, a laser projection device, a TV, a monitor, a wearable device, an electronic table, etc.
[0025] In some embodiments, when the electronic device 200 receives an instruction to process a multi-turn dialogue, it may perform data communication with the server 100. The electronic device 200 is allowed to communicate with the server 100 through a local area network or a wireless local area network.
[0026] The server 100 may be a server that provides various services. For example, it is a server that provides support for the current dialogue statement obtained by the electronic device 200. The server may perform processing such as keyword matching on the received current dialogue statement and feedback the processing result to the electronic device 200. The server 100 may be a server cluster or multiple server clusters, and may include one type or multiple types of servers.
[0027] It should be noted that the multi-turn dialogue processing method provided by the embodiments of the present application can be executed by the electronic device 200, or can be jointly executed by the server 100 and the electronic device 200. The present application does not make any limitations on this.
[0028] The multi-turn dialogue processing device provided by the embodiments of the present application can be hardware or software. When the multi-turn dialogue processing device is hardware, it can be various electronic devices, including but not limited to intelligent vehicles, in-vehicle devices, smart phones, TVs, tablet computers, smart watches, computers, AI devices, robots, and so on. When the multi-turn dialogue processing device is software, it can be installed in the above-listed electronic devices. It can be implemented as multiple software or software modules, or can be implemented as a single software or software module, and no specific limitations are made here.
[0029] Figure 2 It is a schematic flowchart of the multi-turn dialogue processing method provided by the embodiments of the present application. As Figure 4 shown, the multi-turn dialogue processing method can include the following steps.
[0030] S11. Obtain the user's current dialogue statement.
[0031] Among them, the format of the current dialogue statement is in text format.
[0032] In some embodiments, the way to obtain the current dialogue statement can be to obtain the voice information input by the user, and then perform text conversion on the voice information to obtain the current dialogue statement in text format.
[0033] In some other embodiments, the way to obtain the current dialogue statement can also be to directly obtain the current dialogue statement in text format, that is, the multi-turn dialogue processing device directly receives the current dialogue statement in text format without the need for a voice-to-text conversion process.
[0034] Of course, the multi-turn dialogue processing device may also receive dialogue data in other formats, and all of them can be converted into the current dialogue statement in text format in the multi-turn dialogue processing device. The present application does not make any limitations on this.
[0035] Exemplarily, taking the obtained dialogue data as the voice information input by the user as an example, Figure 3 It is an architecture of a multi-turn dialogue processing method provided by the embodiments of the present application. Refer to Figure 3, including an audio input / output module for collecting the input voice information and outputting the processing result of the voice information; an automatic speech recognition (ASR) module is deployed with a speech recognition service for recognizing the audio stream as text; a natural language understanding (NLU) module is deployed with a semantic understanding service for semantic parsing of the text; a dialog management (DM) module is deployed with a business instruction management service for providing business instructions; a natural language generation (NLG) module is deployed with a language generation service for converting the instructions indicating the execution of the multi-turn dialog processing device into text language; a speech synthesis module is deployed with a speech synthesis service for processing the text language corresponding to the instructions and outputting it to the user through the audio input / output module.
[0036] For another example, taking the current dialog statement in the text format input by the user obtained as an example, Figure 4 This is the architecture of another multi-turn dialog processing method provided by the embodiments of the present application. Refer to Figure 4 , including a text input / output module for receiving the input text information and outputting the processing result of the text information; a semantic understanding module is deployed with a semantic understanding service for semantic parsing of the text; a dialog management module is deployed with a business instruction management service for providing business instructions; a language generation module is deployed with a language generation service for converting the instructions indicating the execution of the multi-turn dialog processing device into text language and outputting it to the user through the text input / output module.
[0037] Of course, Figure 2 or Figure 3 In the architecture shown, there may be multiple entity service devices deployed with different business services, or one or more entity service devices may integrate one or more functional services. The present application does not make any limitations in this regard.
[0038] S12. Match the target keyword in the preset word library according to the current dialog statement, and execute step S13 when the target keyword is matched; execute step S14 when the target keyword is not matched.
[0039] Among them, the preset word library includes at least one keyword in the in-vehicle text language style. The keyword in the in-vehicle text language style refers to a keyword related to the vehicle, such as vehicle model, vehicle name, device name that makes up the vehicle, etc.
[0040] S13. Use the target keyword to determine the final answer statement.
[0041] In some embodiments, such as Figure 5As shown, the method of using the target keyword to determine the final answer statement can be to use the target keyword and the preset correspondence to determine the final answer statement. More specifically, it can include the following steps:
[0042] S131. Retrieve at least one target answer statement from the preset correspondence according to the target keyword, and determine the score of each target answer statement.
[0043] Among them, the preset correspondence is the correspondence between the keyword and the answer statement.
[0044] In some embodiments, before retrieving at least one target answer statement from the preset correspondence according to the target keyword, the multi-round dialogue processing method further includes: obtaining at least one keyword and the answer statement corresponding to each keyword; establishing and storing the preset correspondence between the keyword and the answer statement. Among them, the way to obtain the keyword and the answer statement can be the manually marked keyword and / or answer statement, or the keyword extracted from the text in the in-vehicle text language style using the keyword extraction algorithm, and / or the answer statement extracted from the text in the in-vehicle text language style, or directly taking the text in the in-vehicle text language style as a whole as the answer statement. This application does not make any limitations in this regard. The text in the in-vehicle text language style is text related to the vehicle, for example, the car user manual, the car operation manual, etc.
[0045] After that, retrieve at least one target answer statement from the preset correspondence according to the target keyword, and determine the score of each target answer statement.
[0046] In some embodiments, the way to determine the score of each target answer statement can be: performing the following processing operations on each target answer statement to obtain the score of each target answer statement.
[0047] The processing operations include: scoring the first answer statement using the text recall method to obtain the first score; scoring the first answer statement using the vector recall method to obtain the second score; and performing a weighted sum of the first score and the second score to obtain the score of the first answer statement. Among them, the first answer statement is any target answer statement.
[0048] First, the way to score the first answer statement using the text recall method to obtain the first score can be to score the first answer statement using the BM25 formula to obtain the first score.
[0049] Among them, the BM25 formula is: Score(Q, d) is used to represent the first score, Q is used to represent the current dialogue statement, q i is used to represent the target keyword, d is used to represent the first answer statement, W iUsed to represent the weight of the target keyword, where k1 and b are adjustment factors, usually set according to experience. For example, k1 = 2 and b = 0.75; f i Used to represent the frequency of the target keyword appearing in the first answer sentence, dl is used to represent the length of the first answer sentence, avgdl is used to represent the average length of all target answer sentences corresponding to the target keyword, i is the i-th target keyword in the current dialogue sentence for the target keyword, and n is the upper limit of the number of target keywords.
[0050] After that, the method of using vector recall to score the first answer sentence to obtain the second score can be to first vectorize the target keyword and the first answer sentence to obtain a keyword vector and an answer sentence vector. Then, use a similarity calculation algorithm to calculate the similarity between the keyword vector and the answer sentence vector to obtain the second score. The higher the similarity between the keyword vector and the answer sentence vector, the higher the second score obtained.
[0051] Among them, the method of vectorizing the target keyword and the first answer sentence can be to use word embedding and sentence embedding for vectorization, or use a pre-trained neural network model for vectorization. This application does not make a limitation. The similarity calculation formula can be cosine similarity, sine similarity, etc., and this application also does not make a limitation.
[0052] In some embodiments, before weighted summing the first score and the second score, the multi-round dialogue processing method further includes: normalizing the first score and the second score to limit the first score and the second score between 0 and 1 to avoid too large numerical values affecting the calculation result and saving computing resources.
[0053] Finally, the method of weighted summing the first score and the second score to obtain the score of the first answer sentence can be: first obtain the first weight and the second weight, and then perform weighted summing on the first score and the second score according to the following formula.
[0054] S = v1×A + v2×B, where S is used to represent the score of the first answer sentence, v1 is used to represent the first weight, v2 is used to represent the second weight, A is used to represent the first score, and B is used to represent the second score; the first weight is the weight of the first score, and the second weight is the weight of the second score; the values of the first weight and the second weight are both preset. For example, they can be default values, or values set by relevant personnel according to the actual situation. For another example, the first weight is 0.3 and the second weight is 0.7.
[0055] In the above solution, text recall and vector recall can be combined, that is, the question-and-answer matching logic and semantic analysis logic are combined to score the answer sentences, solving the problem that the results of the simple Q&A library matching mechanism are not ideal, and further improving the user experience.
[0056] S132. Determine a target score from the scores of at least one target answer sentence according to a preset condition.
[0057] The preset condition includes: the target score is the largest score among at least one target answer sentence, and the target score is greater than the score threshold. Among them, the score threshold is preset. For example, it can be a default value, or a value set by relevant personnel according to the actual situation. For another example, the score threshold is 0.75.
[0058] In some embodiments, at least one target answer sentence is sorted in descending order of scores, and it is determined whether the score of the target answer sentence ranked first is greater than the score threshold. When the score of the target answer sentence ranked first is greater than the score threshold, the score of the target answer sentence ranked first is determined as the target score.
[0059] In some embodiments, when the score of the target answer sentence ranked first is less than or equal to the score threshold, that is, the scores of the target answer sentences do not meet the preset conditions, a takeover request is sent. Among them, the takeover request is used to request manual acceptance of the current conversation.
[0060] S133. Determine the target answer sentence corresponding to the target score as the final answer sentence.
[0061] In some embodiments, when there are multiple target answer sentences corresponding to the target score, any one of the target answer sentences corresponding to the target score is taken as the final answer sentence.
[0062] S14. Obtain historical conversation sentences.
[0063] In some embodiments, the way to obtain historical conversation sentences can be to obtain historical conversation sentences by calling historical conversation records, or to obtain historical conversation sentences in the way of step S11. This application does not make any limitations on this.
[0064] S15. Determine the final answer sentence according to the historical conversation and the language model.
[0065] Among them, the language model is a language model pre-trained according to texts in the in-vehicle text language style.
[0066] In some embodiments, before step S15, the multi-turn dialogue processing method further includes: obtaining text in the in-vehicle text language style, and fine-tuning the natural language processing model according to the text in the in-vehicle text language style to obtain a language model. Among them, the manner of fine-tuning the natural language processing model according to the text in the in-vehicle text language style to obtain a language model may be to first extract a training set from the text in the in-vehicle text language style, and then use the training set to perform adaptive training on the natural language processing model to obtain a language model. The natural language processing model may be the current open-source large model, for example, models such as ChatGLM and ChatGPT.
[0067] In some embodiments, as Figure 6 shown, the manner of determining the final answer statement according to the historical dialogue and the language model may be to extract summary information from the historical dialogue statements, and the language model uses the summary information to generate the final answer statement. More specifically, it may include the following steps:
[0068] S151. Obtain the current dialogue count.
[0069] Among them, when receiving a user's dialogue statement once, the dialogue count is incremented by one, or when replying with an answer statement once, the dialogue count is incremented by one.
[0070] For example, the process of this conversation is as follows, where A is used to represent the user's dialogue statement and B is used to represent the answer statement of the multi-turn dialogue processing device.
[0071] [A: Hello. B: Hello, I'm X. / / At this time, the dialogue count is 1 time.
[0072] A: What is XX? B: XX is XXX. / / At this time, the dialogue count is 1 + 1 = 2 times.
[0073] A: Then how is XXXX done? B: The steps of XXXX are 1234. / / At this time, the dialogue count is 2 + 1 = 3 times.
[0074] A: Explain step 2. B:... ] / / At this time, the dialogue count is 3 + 1 = 4 times.
[0075] In this way, the current dialogue count is 4 times.
[0076] S152. Determine whether the current dialogue count is greater than the dialogue threshold. If the current dialogue count is greater than the dialogue threshold, execute step S143; if the current dialogue count is less than or equal to the dialogue threshold, execute step S146.
[0077] Among them, the dialogue threshold is preset. For example, it can be the default value, or a value set by relevant personnel according to the actual situation. For another example, the dialogue threshold is 3.
[0078] S153. Determine that the historical dialogue statement includes at least one round of historical dialogue.
[0079] In some embodiments, the specific number of rounds of the historical dialogue statement is preset. For example, it can be the default number of rounds, or the number of rounds set by relevant personnel according to the actual situation. For another example, the specific number of rounds of the historical dialogue statement is 3 rounds.
[0080] S154. Extract the summary information from at least one round of historical dialogue.
[0081] In some embodiments, the method for extracting the summary information from at least one round of historical dialogue can be to extract the summary information from at least one round of historical dialogue based on a pre-trained summary extraction model, or to extract the summary information from at least one round of historical dialogue based on a summary extraction algorithm. This application does not make any limitations in this regard.
[0082] S155. Input the summary information into a language model to obtain the final answer statement.
[0083] In some embodiments, before inputting the summary information into the language model, the multi-round dialogue processing method further includes: generating a guiding statement according to at least one round of historical dialogue, and constructing the input information of the language model according to the summary information, the guiding statement, and the current dialogue statement.
[0084] In some embodiments, the method for generating a guiding statement according to at least one round of historical dialogue can be to input at least one round of historical dialogue into a pre-trained generation model to obtain the guiding statement. Among them, the generation model is used to generate the guiding statement according to the following generation conditions. The generation conditions are: determining the problem points mentioned by the user in the dialogue; if there is no answer statement for the problem points in at least one round of historical dialogue, the information is retained; if the question information in at least one round of historical dialogue is a greeting, the information is ignored.
[0085] In some embodiments, the structure of the input information of the language model is: custom guiding statement + summary information + current dialogue statement.
[0086] After that, input the input information into the language model to obtain the final answer statement.
[0087] In the above solution, a guiding statement can be generated according to at least one round of historical dialogue, and then the input information of the language model can be constructed according to the summary information, the guiding statement, and the current dialogue statement. Finally, the input information is input into the language model to obtain the final answer statement, avoiding the limitation problem of only using the current dialogue statement for retrieval and matching, and being able to obtain a more accurate answer statement by combining the context of the conversation, improving the user experience.
[0088] S156. Concatenate the historical dialogue statement and the current dialogue statement to obtain a concatenated statement.
[0089] Exemplarily, the process of this conversation is as follows, where A is used to represent the user's conversation statement, and B is used to represent the answer statement of the multi-turn conversation processing device.
[0090] [A: What is XX? B: XX is XXX.
[0091] A: Then how is XXXX done? B: The steps of XXXX are 1234.
[0092] A: Explain step 2. B: ……]
[0093] Then the historical conversation statements include: What is XX, XX is XXX, Then how is XXXX done, The steps of XXXX are 1234; The current conversation statement is: Explain step 2. In this way, the concatenated statement is: What is XX, XX is XXX, Then how is XXXX done, The steps of XXXX are 1234; Explain step 2.
[0094] S157. Input the concatenated statement into the language model to obtain the final answer statement.
[0095] In the above solution, first, obtain the user's current conversation statement. Then, when a target keyword is matched in the preset word library according to the current conversation statement, use the target keyword to determine the final answer statement, where the preset word library includes at least one keyword in the in-vehicle text language style. Then, when no target keyword is matched in the preset word library according to the current conversation statement, obtain the historical conversation statements, and determine the final answer statement according to the historical conversation and the language model. The language model is a language model pre-trained according to the text in the in-vehicle text language style. In this way, since the keywords are keywords in the in-vehicle text language style, that is, the matched keywords and answer statements are all related to the vehicle, it is possible to determine whether the conversation scenario is related to the vehicle according to the keywords in the current conversation statement, and when the conversation scenario is related to the vehicle, it is possible to determine the final answer statement according to the historical statements and the language model. It avoids the multi-turn conversation implemented through the Q&A library matching mechanism. Due to the limited number of answer statements in the Q&A library, or in the task-based conversation implemented based on the card slot, the interaction between the user and the device is based on a specific task. After leaving the task, it may lead to the problem of answering irrelevantly or being unable to respond during the conversation process, improving the user experience.
[0096] In the embodiments of the present application, the multi-turn dialogue processing device can be divided into functional modules according to the above method examples. For example, each functional module can be corresponding to each function, or two or more functions can be integrated into one processing unit. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there can be other division methods in actual implementation.
[0097] As Figure 7 shown, an embodiment of the present application provides a schematic structural diagram of a multi-turn dialogue processing device 200, and the multi-turn dialogue processing device 200 includes an acquisition module 201, a determination module 202, and a processing module 203.
[0098] The acquisition module 201 is configured to acquire the current dialogue statement of the user; the determination module 202 is configured to, when a target keyword is matched in the preset word library according to the current dialogue statement, use the target keyword and the preset correspondence to determine the final answer statement; the preset word library includes at least one keyword in the in-vehicle text language style, and the preset correspondence is the correspondence between the keyword and the answer statement; the acquisition module 201 is further configured to acquire the historical dialogue statement when no target keyword is matched in the preset word library according to the current dialogue statement; the processing module 203 is configured to extract summary information from the historical dialogue statement, and the language model uses the summary information to generate the final answer statement; the language model is a language model pre-trained according to the text in the in-vehicle text language style.
[0099] In some embodiments, the determination module 202 is specifically configured to: retrieve at least one target answer statement in the preset correspondence according to the target keyword, and determine the score of each target answer statement; determine the target score from the scores of at least one target answer statement according to the preset condition; the preset condition includes: the target score is the largest score among at least one target answer statement, and the target score is greater than the score threshold; determine the target answer statement corresponding to the target score as the final answer statement.
[0100] In some embodiments, the determination module 202 is specifically configured to: perform the following processing operations on each target answer statement to obtain the score of each target answer statement; the processing operation includes: scoring the first answer statement using the text recall method to obtain the first score; scoring the first answer statement using the vector recall method to obtain the second score; performing a weighted sum on the first score and the second score to obtain the score of the first answer statement; the first answer statement is any one of the target answer statements.
[0101] In some embodiments, the determining module 202 is specifically configured to: when the scores of the target answer statements do not meet the preset conditions, send a takeover request; the takeover request is used to request manual acceptance of the current conversation.
[0102] In some embodiments, the processing module 203 is specifically configured to: obtain the current conversation count; wherein, when a user's conversation statement is received, the conversation count is incremented by one; when the current conversation count is greater than the conversation threshold, determine that the historical conversation statements include at least one round of historical conversation; extract the summary information from at least one round of historical conversation; input the summary information into the language model to obtain the final answer statement.
[0103] In some embodiments, the processing module 203 is specifically configured to: when the current conversation count is less than or equal to the conversation threshold, splice the historical conversation statements and the current conversation statement to obtain a spliced statement; input the spliced statement into the language model to obtain the final answer statement.
[0104] In some embodiments, the obtaining module 201 is further configured to obtain text in the in-vehicle text language style; the processing module 203 is further configured to fine-tune the natural language processing model according to the text in the in-vehicle text language style to obtain the language model.
[0105] The multi-round conversation processing device provided in this embodiment can execute the multi-round conversation processing method provided in the above method embodiment. The implementation principle and technical effect are similar to those of the above method, and will not be elaborated here.
[0106] Figure 8 It is a schematic structural diagram of an electronic device provided in an embodiment of the present application.
[0107] As Figure 8 shown, an embodiment of the present application provides an electronic device, which includes: a processor 1201, a memory 1202, and a computer program stored on the memory 1202 and executable on the processor 1201. When the computer program is executed by the processor 1201, it implements each process of the multi-round conversation processing method in the above method embodiment. And it can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0108] An embodiment of the present application provides a computer-readable storage medium, which is characterized in that a computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements each process of the multi-round conversation processing method in the above method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0109] Among them, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disc, or the like.
[0110] An embodiment of the present application provides a computer program product. The computer program product stores a computer program. When the computer program is executed by a processor, it implements each process of the multi-round dialogue processing method in the above method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0111] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.
[0112] In the present application, the processor can be a central processing unit (CPU), or can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0113] In the present application, the memory may include non-permanent memory in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0114] In this application, computer-readable media include both permanent and non-permanent, removable and non-removable storage media. The storage media can implement information storage by any method or technology, and the information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media do not include transitory media such as modulated data and carrier waves.
[0115] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.
[0116] The above description is only the specific implementation manners of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to these embodiments described herein, but rather will conform to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multi-turn dialogue processing method, characterized in that, including: obtaining the current conversation statement of the user; when a target keyword is matched in a preset keyword library according to the current conversation statement, determining a final answer statement by using the target keyword and a preset corresponding relationship; at least one keyword in the preset keyword library has a vehicle-mounted text language style, and the preset corresponding relationship is the corresponding relationship between the keyword and the answer statement; when a target keyword is not matched in the preset keyword library according to the current conversation statement, obtaining historical conversation statements; extracting summary information from the historical conversation statements, and generating a final answer statement by using the summary information by a language model; the language model is a language model pre-trained according to texts with a vehicle-mounted text language style.
2. The multi-round conversation processing method according to claim 1, wherein, The determining the final answer statement by using the target keyword and the preset corresponding relationship includes: retrieving at least one target answer statement in the preset corresponding relationship according to the target keyword, and determining the score of each target answer statement; determining a target score from the scores of the at least one target answer statement according to a preset condition; the preset condition includes: the target score is the largest score among the at least one target answer statement, and the target score is greater than a score threshold; determining the target answer statement corresponding to the target score as the final answer statement.
3. The multi-round dialogue processing method according to claim 2, wherein The determining the score of each target answer statement includes: performing the following processing operations on each target answer statement to obtain the score of each target answer statement; the processing operations include: scoring a first answer statement by using a text recall method to obtain a first score; scoring the first answer statement by using a vector recall method to obtain a second score; weighted summing the first score and the second score to obtain the score of the first answer statement; the first answer statement is any target answer statement.
4. The multi-round dialogue processing method according to claim 2, wherein After determining the scores of each target answer statement, the method further includes: when the scores of the target answer statements do not satisfy the preset condition, sending a takeover request; the takeover request is used to request manual acceptance of the current conversation.
5. The multi-round dialogue processing method according to claim 1, characterized in that The extracting summary information from the historical conversation statements and generating a final answer statement by using the summary information by a language model includes: obtaining the current conversation count; wherein, when a conversation statement of the user is received, the conversation count is incremented by one; when the current conversation count is greater than a conversation threshold, determining that the historical conversation statements include at least one round of historical conversation; extracting summary information from the at least one round of historical conversation; inputting the summary information into the language model to obtain a final answer statement.
6. The multi-round dialogue processing method according to claim 5, wherein After obtaining the current conversation count, the method further includes: when the current conversation count is less than or equal to the conversation threshold, splicing the historical conversation statements and the current conversation statement to obtain a spliced statement; inputting the spliced statement into the language model to obtain a final answer statement.
7. The multi-round dialogue processing method according to claim 5 or 6, characterized in that, Before determining the final answer statement according to the historical conversation and the language model, the method further includes: obtaining texts with a vehicle-mounted text language style; fine-tuning a natural language processing model according to the texts with a vehicle-mounted text language style to obtain the language model.
8. A multi-turn dialogue processing device, characterized in that, including: an obtaining module, configured to obtain the current conversation statement of the user; A determination module, configured to, when a target keyword is matched in a preset keyword library according to the current dialogue statement, determine a final answer statement by using the target keyword and a preset corresponding relationship; at least one keyword in the preset keyword library is in the language style of in-vehicle texts, and the preset corresponding relationship is the corresponding relationship between a keyword and an answer statement; The acquisition module is further configured to acquire historical dialogue statements when no target keyword is matched in the preset keyword library according to the current dialogue statement; A processing module, configured to extract summary information from the historical dialogue statements, and a language model generates a final answer statement by using the summary information; The language model is a language model pre-trained according to texts in the language style of in-vehicle texts.
9. An electronic device, characterized in that, It includes: A processor, a memory, and a computer program stored on the memory and executable on the processor, where when the computer program is executed by the processor, the multi-round dialogue processing method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that, It includes: A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the multi-round dialogue processing method according to any one of claims 1 to 7 is implemented.