Conversation interaction method and device

By extracting user intention and context information and generating and replacing reply texts, the poor dialogue quality problem caused by dialogue interaction dependence on machine translation in the prior art is solved, and the dialogue interaction quality and user satisfaction are improved.

CN119988549APending Publication Date: 2025-05-13北京云迹科技股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510062588.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, dialogue interaction depends on machine translation, resulting in poor dialogue quality.

Method used

By obtaining user input and dialogue configuration information, user intention and context information are extracted, and the user language entered by the user is judged. Reply text is generated based on user intent, context information and dialogue configuration information, using a preset language. Then, the reply text is replaced in sentence form according to the user's language to obtain the replacement result, and finally the reply output corresponding to the user input is determined based on the replacement result.

Benefits of technology

It improves the quality of conversation interaction and user satisfaction, avoiding the problems of missing information and poor dialogue experience caused by machine translation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988549A_ABST
    Figure CN119988549A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and provides a dialogue interaction method and device. The method comprises the following steps: acquiring user input and dialogue configuration information; extracting a user intention and context information input by the user, and judging a user language input by the user; a reply text is generated according to the user intention, the context information and the dialogue configuration information, and the language of the reply text is a preset language; performing sentence pattern replacement on the reply text according to a user language to obtain a replacement result; and determining reply output corresponding to the user input based on the replacement result. By means of the technical means, the problem that in the prior art, dialogue interaction depends on machine translation, and consequently dialogue quality is poor is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a conversation interaction method and device. Background Art

[0002] In order to improve the efficiency of interaction with users, technicians have been developing high-level intelligent customer service. At present, the method of intelligent customer service to interact with users relies too much on machine translation, which will lead to missing information and poor conversation experience. Summary of the invention

[0003] In view of this, embodiments of the present application provide a conversation interaction method, device, electronic device, and computer-readable storage medium to solve the problem in the prior art that conversation interaction relies on machine translation, resulting in poor conversation quality.

[0004] According to a first aspect of an embodiment of the present application, a method for dialogue interaction is provided, including: obtaining user input and dialogue configuration information; extracting user intent and context information of the user input, and determining the user language of the user input; generating a reply text according to the user intent, context information and dialogue configuration information, wherein the language of the reply text is a preset language; performing sentence replacement on the reply text according to the user language to obtain a replacement result; and determining a reply output corresponding to the user input based on the replacement result.

[0005] According to a second aspect of an embodiment of the present application, a conversation interaction device is provided, including: an acquisition module configured to acquire user input and conversation configuration information; an extraction module configured to extract user intent and context information of the user input, and to determine the user language of the user input; a generation module configured to generate a reply text according to the user intent, context information and conversation configuration information, wherein the language of the reply text is a preset language; a replacement module configured to perform sentence replacement on the reply text according to the user language to obtain a replacement result; and a determination module configured to determine the reply output corresponding to the user input based on the replacement result.

[0006] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.

[0007] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, which stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0008] Compared with the prior art, the embodiments of the present application have the following beneficial effects: obtaining user input and dialogue configuration information; extracting user intent and context information of the user input, and determining the user language of the user input; generating a reply text according to the user intent, context information and dialogue configuration information, wherein the language of the reply text is a preset language; performing sentence replacement on the reply text according to the user language to obtain a replacement result; and determining the reply output corresponding to the user input based on the replacement result. The above technical means can solve the problem in the prior art that dialogue interaction relies on machine translation, resulting in poor dialogue quality, thereby improving the dialogue interaction quality and user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0010] Figure 1 It is a flowchart of a conversation interaction method provided in an embodiment of the present application;

[0011] Figure 2 is a flowchart of another dialog interaction method provided in an embodiment of the present application;

[0012] Figure 3 is a structural diagram of a conversation interaction device provided in an embodiment of the present application;

[0013] Figure 4 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0014] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.

[0015] A method and device for dialog interaction according to an embodiment of the present application will be described in detail below with reference to the accompanying drawings.

[0016] Figure 1 It is a flowchart of a dialogue interaction method provided in an embodiment of the present application. Figure 1 The dialog interaction method can be executed by a computer or a server, or software on a computer or a server. Figure 1 As shown, the dialogue interaction method includes:

[0017] S101, obtaining user input and dialogue configuration information;

[0018] S102, extracting user intention and context information input by the user, and determining the user language input by the user;

[0019] S103, generating a reply text according to the user intention, context information and dialogue configuration information, wherein the language of the reply text is a preset language;

[0020] S104, performing sentence replacement on the reply text according to the user language to obtain a replacement result;

[0021] S105: Determine a reply output corresponding to the user input based on the replacement result.

[0022] The embodiment of the present application is used to control the interaction between intelligent customer service and users. User input can be voice or text. The dialogue configuration information is the information or answer template related to the user input set by the owner of the intelligent customer service according to his own information, such as information about store products or answers to questions that users may ask. The user intent and context information of the user input are extracted by the large language model to determine the user language of the user input. The user intent can be the question the user wants to ask, the context information is the context features or key information extracted by the large language model, and the user language is the type of language input by the user. According to the user intent, context information and dialogue configuration information, the reply text is generated by the large language model. For example, the user input is English and the reply text is Chinese. The sentence structure of the reply text in Chinese is replaced according to English to obtain the replacement result. Finally, the reply output corresponding to the user input is determined based on the replacement result.

[0023] According to the technical solution provided in the embodiment of the present application, user input and dialogue configuration information are obtained; user intent and context information of the user input are extracted, and the user language of the user input is determined; a reply text is generated according to the user intent, context information and dialogue configuration information, wherein the language of the reply text is a preset language; the sentence structure of the reply text is replaced according to the user language to obtain a replacement result; and the reply output corresponding to the user input is determined based on the replacement result. The above technical means can solve the problem in the prior art that dialogue interaction relies on machine translation, resulting in poor dialogue quality, thereby improving the dialogue interaction quality and user satisfaction.

[0024] Furthermore, user intent and context information of the user input are extracted, and the user language of the user input is determined, including: determining whether the user input is voice or text; if the user input is voice, determining the user language of the user input, performing voice recognition on the user input to obtain target content, and extracting user intent and context information of the target content; if the user input is text, taking the user input as the target content, determining the user language of the target content, and extracting user intent and context information of the target content.

[0025] If the user input is voice, the user input is recognized to convert the user input into the target content. If the user input is text, the user input is directly used as the target content. The user intent and context information of the target content are extracted using a large language model.

[0026] Furthermore, sentence replacement is performed on the reply text according to the user language to obtain a replacement result, including: replacing the speech of the reply text according to the user language; and replacing the entity of the reply text according to the user language to obtain a replacement result.

[0027] The embodiment of the present application adopts a unified dialogue management process to realize reply generation and necessary sentence replacement and machine translation, and supports arbitrary language switching during the dialogue process, so that the development and maintenance of intelligent customer service is more efficient and the dialogue experience is better. Among them, sentence pattern refers to the words and entities that represent a certain specific meaning, and different languages ​​have different expression patterns. Sentence replacement refers to the replacement of words and entity expressions in sentence patterns of different languages ​​to achieve a more accurate and natural dialogue effect. Machine translation is only used for bottom-up processing.

[0028] Furthermore, determining the reply output corresponding to the user input based on the replacement result includes: judging whether there is a segment that does not belong to the user language in the replacement result; if so, translating the segment that does not belong to the user language in the replacement result to obtain a translation result, and determining the reply output corresponding to the user input according to the translation result; if not, determining the reply output corresponding to the user input based on the replacement result.

[0029] The segment can be a sentence, a phrase or a word. The segments that do not belong to the user's language in the replacement result are partially translated, and the untranslated segments plus the translated segments are taken as the translation result.

[0030] Furthermore, determining the reply output corresponding to the user input based on the replacement result includes: determining whether the user input is voice or text; if the user input is voice, performing voice synthesis on the replacement result, and outputting the synthesis result as the reply corresponding to the user input; if the user input is text, outputting the replacement result as the reply corresponding to the user input.

[0031] If the user input is voice, the reply output is also voice, so the replacement result is voice-synthesized. If the user input is text, the reply output is also text, so the replacement result is directly output as the reply.

[0032] In one embodiment, user input in the current round of dialogue is obtained; user intent and context information of the user input in the current round of dialogue are extracted; a reply text corresponding to the current round of dialogue is generated according to dialogue configuration information and user intent and context information corresponding to the current round of dialogue, wherein the language of the reply text in all rounds of dialogue is a preset language; sentence replacement is performed on the reply text corresponding to the current round of dialogue according to the user language to obtain a replacement result corresponding to the current round of dialogue, wherein the user language is determined based on the user input in the first round of dialogue; a reply output corresponding to the current round of dialogue is determined based on the replacement result corresponding to the current round of dialogue; the current round of dialogue is ended and the next round of dialogue is carried out.

[0033] Obtain user input in the next round of dialogue; extract user intent and context information of the user input in the next round of dialogue; generate a reply text corresponding to the next round of dialogue according to the dialogue configuration information and the user intent and context information corresponding to the next round of dialogue, wherein the language of the reply text in all rounds of dialogue is a preset language; perform sentence replacement on the reply text corresponding to the next round of dialogue according to the user language to obtain a replacement result corresponding to the next round of dialogue, wherein the user language is determined based on the user input in the first round of dialogue; determine a reply output corresponding to the next round of dialogue based on the replacement result corresponding to the next round of dialogue; and end the next round of dialogue.

[0034] In one embodiment, the following loop is executed: obtaining the ith user input, where i is a positive integer used to mark the number of user inputs; extracting the user intent and context information of the ith user input; generating the ith reply text according to the dialogue configuration information and the user intent and context information of the ith user input, wherein the language of all reply texts is a preset language; performing sentence replacement on the ith reply text according to the user language to obtain the ith replacement result, wherein the user language is determined based on the first user input; determining the ith reply output based on the ith replacement result, determining to end processing of the ith user input, and updating i with the value of i plus 1; and terminating the loop until no user input is obtained.

[0035] Determine the i-th reply output based on the i-th replacement result, and end the processing of the i-th user input. Check whether there is the i+1-th user input. If so, update i with the value of i+1, that is, process the i+1-th user input. If not, it means that no user input can be obtained, and end the loop.

[0036] Figure 2FIG. 1 is a flow chart of another method for dialog interaction provided in an embodiment of the present application. Figure 2 As shown, the method includes:

[0037] S201, obtaining user input in the current round of dialogue;

[0038] S202, determining whether the user input is voice or text;

[0039] S203, if the user input is voice, determine the user language of the user input, and perform voice recognition on the user input to obtain the target content;

[0040] S204, if the user input is text, taking the user input as the target content and determining the user language of the target content;

[0041] S205, extracting user intent and context information of target content;

[0042] S206, generating a reply text according to the dialogue configuration information, the user's intention and the context information;

[0043] S207, replacing the words in the reply text according to the user language, and replacing the entities in the reply text according to the user language to obtain a replacement result;

[0044] S208, determining whether there is a segment that does not belong to the user's language in the replacement result;

[0045] S209, translating the segments in the replacement result that do not belong to the user language to obtain a translation result;

[0046] S210, if the user input is voice, performing voice synthesis on the translation result, and outputting the synthesized result as a reply;

[0047] S211, if the user input is text, the translation result is directly output as a reply;

[0048] S212, proceed to the next round of dialogue.

[0049] According to the technical solution provided in the embodiment of the present application, user input and dialogue configuration information are obtained; user intent and context information of the user input are extracted, and the user language of the user input is determined; a reply text is generated according to the user intent, context information and dialogue configuration information, wherein the language of the reply text is a preset language; the sentence structure of the reply text is replaced according to the user language to obtain a replacement result; and the reply output corresponding to the user input is determined based on the replacement result. The above technical means can solve the problem in the prior art that dialogue interaction relies on machine translation, resulting in poor dialogue quality, thereby improving the dialogue interaction quality and user satisfaction.

[0050] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present application, which will not be described one by one here.

[0051] The following is an embodiment of the device of the present application, which can be used to execute the embodiment of the method of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the method of the present application.

[0052] Figure 3 Schematic diagram of a conversation interaction device provided in an embodiment of the present application. Figure 3 As shown, the dialogue interaction device includes:

[0053] The acquisition module 301 is configured to acquire user input and dialog configuration information;

[0054] An extraction module 302 is configured to extract user intent and context information input by the user, and to determine the user language input by the user;

[0055] A generation module 303 is configured to generate a reply text according to the user intention, the context information and the dialogue configuration information, wherein the language of the reply text is a preset language;

[0056] The replacement module 304 is configured to replace the reply text with a sentence according to the user language to obtain a replacement result;

[0057] The determination module 305 is configured to determine a reply output corresponding to the user input based on the replacement result.

[0058] The embodiment of the present application is used to control the interaction between intelligent customer service and users. User input can be voice or text. The dialogue configuration information is the information or answer template related to the user input set by the owner of the intelligent customer service according to his own information, such as information about store products or answers to questions that users may ask. The user intent and context information of the user input are extracted by the large language model to determine the user language of the user input. The user intent can be the question the user wants to ask, the context information is the context features or key information extracted by the large language model, and the user language is the type of language input by the user. According to the user intent, context information and dialogue configuration information, the reply text is generated by the large language model. For example, the user input is English and the reply text is Chinese. The sentence structure of the reply text in Chinese is replaced according to English to obtain the replacement result. Finally, the reply output corresponding to the user input is determined based on the replacement result.

[0059] According to the technical solution provided in the embodiment of the present application, user input and dialogue configuration information are obtained; user intent and context information of the user input are extracted, and the user language of the user input is determined; a reply text is generated according to the user intent, context information and dialogue configuration information, wherein the language of the reply text is a preset language; the sentence structure of the reply text is replaced according to the user language to obtain a replacement result; and the reply output corresponding to the user input is determined based on the replacement result. The above technical means can solve the problem in the prior art that dialogue interaction relies on machine translation, resulting in poor dialogue quality, thereby improving the dialogue interaction quality and user satisfaction.

[0060] In some embodiments, the extraction module 302 is also configured to determine whether the user input is voice or text; if the user input is voice, the user language of the user input is determined, voice recognition is performed on the user input to obtain the target content, and the user intent and context information of the target content are extracted; if the user input is text, the user input is used as the target content, the user language of the target content is determined, and the user intent and context information of the target content are extracted.

[0061] If the user input is voice, the user input is recognized to convert the user input into the target content. If the user input is text, the user input is directly used as the target content. The user intent and context information of the target content are extracted using a large language model.

[0062] In some embodiments, the replacement module 304 is further configured to perform word replacement on the reply text according to the user language; and perform entity replacement on the reply text according to the user language to obtain a replacement result.

[0063] The embodiment of the present application adopts a unified dialogue management process to realize reply generation and necessary sentence replacement and machine translation, and supports arbitrary language switching during the dialogue process, so that the development and maintenance of intelligent customer service is more efficient and the dialogue experience is better. Among them, sentence pattern refers to the words and entities that represent a certain specific meaning, and different languages ​​have different expression patterns. Sentence replacement refers to the replacement of words and entity expressions in sentence patterns of different languages ​​to achieve a more accurate and natural dialogue effect. Machine translation is only used for bottom-up processing.

[0064] In some embodiments, the determination module 305 is also configured to determine whether there are segments that do not belong to the user language in the replacement result; if so, the segments that do not belong to the user language in the replacement result are translated to obtain a translation result, and the reply output corresponding to the user input is determined based on the translation result; if not, the reply output corresponding to the user input is determined based on the replacement result.

[0065] The segment can be a sentence, a phrase or a word. The segments that do not belong to the user's language in the replacement result are partially translated, and the untranslated segments plus the translated segments are taken as the translation result.

[0066] In some embodiments, the determination module 305 is also configured to determine whether the user input is voice or text; if the user input is voice, the replacement result is voice synthesized and the synthesis result is output as a reply corresponding to the user input; if the user input is text, the replacement result is output as a reply corresponding to the user input.

[0067] If the user input is voice, the reply output is also voice, so the replacement result is voice-synthesized. If the user input is text, the reply output is also text, so the replacement result is directly output as the reply.

[0068] In some embodiments, the determination module 305 is further configured to obtain user input in the current round of dialogue; extract user intent and context information of the user input in the current round of dialogue; generate a reply text corresponding to the current round of dialogue according to the dialogue configuration information and the user intent and context information corresponding to the current round of dialogue, wherein the language of the reply text in all rounds of dialogue is a preset language; perform sentence replacement on the reply text corresponding to the current round of dialogue according to the user language to obtain a replacement result corresponding to the current round of dialogue, wherein the user language is determined based on the user input in the first round of dialogue; determine the reply output corresponding to the current round of dialogue based on the replacement result corresponding to the current round of dialogue; end the current round of dialogue and proceed to the next round of dialogue.

[0069] In some embodiments, the determination module 305 is further configured to obtain user input in the next round of dialogue; extract user intent and context information of the user input in the next round of dialogue; generate a reply text corresponding to the next round of dialogue according to the dialogue configuration information and the user intent and context information corresponding to the next round of dialogue, wherein the language of the reply text in all rounds of dialogue is a preset language; perform sentence replacement on the reply text corresponding to the next round of dialogue according to the user language to obtain a replacement result corresponding to the next round of dialogue, wherein the user language is determined based on the user input in the first round of dialogue; determine the reply output corresponding to the next round of dialogue based on the replacement result corresponding to the next round of dialogue; and end the next round of dialogue.

[0070] In some embodiments, the determination module 305 is also configured to execute the following loop: obtain the ith user input, where i is a positive integer used to mark the number of user inputs; extract the user intent and context information of the ith user input; generate the ith reply text according to the dialogue configuration information and the user intent and context information of the ith user input, wherein the language of all reply texts is a preset language; perform sentence replacement on the ith reply text according to the user language to obtain the ith replacement result, wherein the user language is determined based on the first user input; determine the ith reply output based on the ith replacement result, determine to end the processing of the ith user input, and update i with the value of i plus 1; and end the loop until no user input is obtained.

[0071] Determine the i-th reply output based on the i-th replacement result, and end the processing of the i-th user input. Check whether there is the i+1-th user input. If so, update i with the value of i+1, that is, process the i+1-th user input. If not, it means that no user input can be obtained, and end the loop.

[0072] In some embodiments, the determination module 305 is also configured to obtain user input in the current round of dialogue; determine whether the user input is voice or text; if the user input is voice, determine the user language of the user input, and perform voice recognition on the user input to obtain the target content; if the user input is text, use the user input as the target content, determine the user language of the target content; extract the user intention and context information of the target content; generate a reply text according to the dialogue configuration information, user intention and context information; replace the reply text with words according to the user language, and replace the reply text with entities according to the user language to obtain a replacement result; determine whether there are fragments that do not belong to the user language in the replacement result; translate the fragments that do not belong to the user language in the replacement result to obtain a translation result; if the user input is voice, perform voice synthesis on the translation result, and output the synthesis result as a reply; if the user input is text, output the translation result directly as a reply; and proceed to the next round of dialogue.

[0073] According to the technical solution provided in the embodiment of the present application, user input and dialogue configuration information are obtained; user intent and context information of the user input are extracted, and the user language of the user input is determined; a reply text is generated according to the user intent, context information and dialogue configuration information, wherein the language of the reply text is a preset language; the sentence structure of the reply text is replaced according to the user language to obtain a replacement result; and the reply output corresponding to the user input is determined based on the replacement result. The above technical means can solve the problem in the prior art that dialogue interaction relies on machine translation, resulting in poor dialogue quality, thereby improving the dialogue interaction quality and user satisfaction.

[0074] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0075] Figure 4 Schematic diagram of an electronic device 4 provided in an embodiment of the present application. Figure 4 As shown, the electronic device 4 of this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. When the processor 401 executes the computer program 403, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor 401 executes the computer program 403, the functions of the modules / units in the above-mentioned device embodiments are implemented.

[0076] The electronic device 4 may be a desktop computer, a notebook, a PDA, a cloud server, or other electronic device. The electronic device 4 may include, but is not limited to, a processor 401 and a memory 402. Those skilled in the art will appreciate that Figure 4 The electronic device 4 is merely an example and does not limit the electronic device 4 , and may include more or less components than those shown in the figure, or different components.

[0077] The processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0078] The memory 402 may be an internal storage unit of the electronic device 4, for example, a hard disk or memory of the electronic device 4. The memory 402 may also be an external storage device of the electronic device 4, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 4. The memory 402 may also include both an internal storage unit and an external storage device of the electronic device 4. The memory 402 is used to store computer programs and other programs and data required by the electronic device.

[0079] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units.

[0080] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. The computer program may include computer program code, and the computer program code may be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electric carrier signals and telecommunication signals.

[0081] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A conversation interaction method, characterized in that: include: Get user input and dialog configuration information; Extracting user intention and context information of the user input, and determining the user language of the user input; Generate a reply text according to the user intention, the context information and the dialog configuration information, wherein the language of the reply text is a preset language; Perform sentence replacement on the reply text according to the user language to obtain a replacement result; A reply output corresponding to the user input is determined based on the replacement result.

2. The method according to claim 1, characterized in that Extracting user intention and context information of the user input, and determining the user language of the user input, including: Determining whether the user input is voice or text; If the user input is voice, determining the user language of the user input, performing voice recognition on the user input to obtain target content, and extracting user intent and context information of the target content; If the user input is text, the user input is used as the target content, the user language of the target content is determined, and the user intention and context information of the target content are extracted.

3. The method according to claim 1, characterized in that: The reply text is subjected to sentence replacement according to the user language to obtain a replacement result, including: Replacing the reply text with words according to the user language; And performing entity replacement on the reply text according to the user language to obtain the replacement result.

4. The method according to claim 1, characterized in that Determining a reply output corresponding to the user input based on the replacement result includes: Determining whether there is a segment that does not belong to the user language in the replacement result; If so, translating the segment in the replacement result that does not belong to the user language to obtain a translation result, and determining a reply output corresponding to the user input according to the translation result; If not, determining a reply output corresponding to the user input based on the replacement result.

5. The method according to claim 1, characterized in that Determining a reply output corresponding to the user input based on the replacement result includes: Determining whether the user input is voice or text; If the user input is speech, performing speech synthesis on the replacement result, and outputting the synthesis result as a reply corresponding to the user input; If the user input is text, the replacement result is output as a reply corresponding to the user input.

6. The method according to claim 1, characterized in that The method further comprises: Get the user input in the current round of dialogue; Extract the user intent and context information input by the user in the current round of dialogue; Generate a reply text corresponding to the current round of dialogue according to the dialogue configuration information and the user intention and context information corresponding to the current round of dialogue, wherein the language of the reply text in all rounds of dialogue is a preset language; Performing sentence replacement on the reply text corresponding to the current round of dialogue according to the user language to obtain a replacement result corresponding to the current round of dialogue, wherein the user language is determined based on the user input in the first round of dialogue; Determine the reply output corresponding to the current round of dialogue based on the replacement result corresponding to the current round of dialogue; End the current round of dialogue and proceed to the next round of dialogue.

7. The method according to claim 1, characterized in that The method further comprises: Execute the following loop: Get the i-th user input, where i is a positive integer used to mark the number of user inputs; Extract the user intent and context information of the i-th user input; Generate an i-th reply text according to the dialogue configuration information and the user intention and context information of the i-th user input, wherein the language of all reply texts is a preset language; Perform sentence replacement on the i-th reply text according to the user language to obtain the i-th replacement result, wherein the user language is determined based on the first user input; Determine the i-th reply output based on the i-th replacement result, determine the end of processing the i-th user input, and update i with the value of i plus 1; The loop ends when no user input is obtained.

8. A conversational interaction device, characterized in that: include: An acquisition module, configured to acquire user input and dialog configuration information; An extraction module, configured to extract the user intention and context information of the user input, and determine the user language of the user input; A generation module, configured to generate a reply text according to the user intention, the context information and the dialogue configuration information, wherein the language of the reply text is a preset language; A replacement module is configured to perform sentence replacement on the reply text according to the user language to obtain a replacement result; The determination module is configured to determine a reply output corresponding to the user input based on the replacement result.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.