Information interaction method and device, storage medium and program product
By combining AI voice assistants with customer service programs and utilizing voice and head movement recognition technologies, the system assists customer service personnel with disabilities in completing interactive operations, solving the problem of low interaction efficiency caused by physical disabilities and improving user experience.
Patent Information
- Application Number
- CN202410997326.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2044-07-23
AI Technical Summary
Customer service staff with disabilities are unable to use peripherals such as keyboards and mice effectively due to physical limitations, resulting in low interaction efficiency and low user satisfaction.
An AI voice assistant is integrated with the customer service program. Through voice recognition and head movement recognition modules, customer service personnel can input their voice and head movements to assist in completing interactive operations.
It improves the interaction efficiency of customer service personnel and solves the problem of low interaction efficiency caused by the inconvenience of manual input, and is especially applicable to interaction scenarios involving people with disabilities.
Smart Images

Figure CN119025203B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to an information interaction method and device, a storage medium and a program product. BACKGROUND
[0002] At present, some enterprises have launched a new customer service solution, that is, a cloud customer service based on cloud computing technology, to help solve the employment problem of the disabled population. The cloud customer service can work anywhere with network, and can create an accessible work platform for the disabled, so that they can provide professional customer service for enterprise users at home or even in bed through computers or smart phones.
[0003] The customer service personnel enter the customer service interface through the terminal device such as computer or smart phone, understand the user's problem information in the customer service interface, and interact with the terminal device through peripherals such as keyboard and mouse to answer various questions for the user, such as helping the user to query order information or logistics information. However, the disabled customer service personnel have various physical obstacles, such as may not be able to flexibly use peripherals such as keyboard and mouse, so there is a problem of inconvenient operation, resulting in low customer service efficiency and low user satisfaction. SUMMARY
[0004] The aspects of the present application provide an information interaction method, device, storage medium and program product, to provide an efficient and fast interaction mode, solve the problem of low interaction efficiency in various interaction scenarios, especially the interaction scenarios participated by the disabled.
[0005] The information interaction method provided by the embodiment of the present application is applied to an AI voice assistant on a terminal device, the AI voice assistant is associated with a customer service program on the terminal device, and the method comprises the following steps: displaying a customer service interface provided by the customer service program, the customer service interface comprising a question and answer area, the question and answer area being used to display at least one target problem information of a target user for consulting a customer service personnel; receiving at least one voice interaction instruction initiated by the customer service personnel for any target problem information based on a voice recognition module in the AI voice assistant; for any voice interaction instruction, executing the voice interaction instruction and outputting the execution result information of the voice interaction instruction in a display and / or voice manner; for a target voice interaction instruction in the at least one voice interaction instruction, selecting at least part of the execution result information of the target voice interaction instruction as target reply information of the any target problem information and displaying the target reply information in the question and answer area.
[0006] The embodiment of the application further provides an information interaction method, which is applied to an AI voice assistant on a terminal device, the AI voice assistant is associated with a target program on the terminal device, and the method comprises the following steps: displaying an interaction interface provided by the target program, the interaction interface comprises an interaction area, and the interaction area is used for displaying at least one interaction information initiated by a first user to a second user; receiving at least one voice interaction instruction initiated by the second user for any interaction information based on a voice recognition module in the AI voice assistant; executing any voice interaction instruction and outputting execution result information of the any voice interaction instruction in a display and / or voice manner for the any voice interaction instruction; and selecting at least part of information in the execution result information of a target voice interaction instruction as target reply information of the any interaction information and displaying the target reply information in the interaction area for the target voice interaction instruction in the at least one voice interaction instruction.
[0007] The embodiment of the application further provides an electronic device, which comprises a memory and a processor; the memory is used for storing a computer program; the processor is coupled with the memory and is used for executing the computer program in the memory so as to implement steps in the information interaction method.
[0008] The embodiment of the application further provides a computer readable storage medium storing a computer program, when the computer program is executed by a processor, the processor can implement steps in any one of the information interaction methods.
[0009] The embodiment of the application further provides a computer program product, which comprises a computer program / instruction, when the computer program / instruction is executed by a processor, the processor can implement steps in the information interaction method.
[0010] In the embodiment, an AI voice assistant is provided for a customer service personnel, the AI voice assistant is associated with a customer service program, is responsible for displaying a customer service interface provided by the customer service program on one hand, and provides a voice function for the customer service personnel on the other hand, can respond to and execute voice interaction instructions of the customer service personnel and then give execution result information, and can select content capable of being used as reply information from corresponding execution result information and display the content in a question and answer area for voice instructions that need to be replied to a user, can assist the customer service personnel to complete interaction with the user to a certain extent, and in this process, the customer service personnel can complete interaction through simple voice input with the assistance of the AI voice assistant, can solve the problem that a mouse, a keyboard and other peripherals cannot be used flexibly due to inconvenience of manual input, and then solve the problem of low interaction efficiency caused thereby. The scheme has strong universality, is applicable to various interaction scenes, and is especially applicable to an interaction scene in which a disabled person participates.
[0011] The embodiment of the present application also provides an information interaction method, which provides an AI voice assistant for a user, the AI voice assistant is associated with a target program used for interaction between the AI voice assistant and the user, is responsible for displaying an interaction interface provided by the target program on one hand; and provides a voice function for the user on the other hand, can respond to and execute a voice interaction instruction of the user and then give an execution result information, and can select content capable of being used as reply information from the corresponding execution result information and display the content in the interaction interface for a voice instruction of another user that needs to be replied, can assist the user in interaction to a certain extent, and in this process, the user can complete the interaction through simple voice input under the assistance of the AI voice assistant, can solve the problem that the user cannot flexibly use a mouse, a keyboard and other peripherals to interact due to inconvenience of manual input, and then solve the problem of low interaction efficiency caused thereby. BRIEF DESCRIPTION OF DRAWINGS
[0012] The accompanying drawings, which are included to provide a further understanding of the present application, form a part of the present application and illustrate the illustrative embodiments of the present application and together with the description serve to explain the present application. In the drawings:
[0013] Figure 1a A flowchart of an information interaction method provided by an illustrative embodiment of the present application;
[0014] Figure 1b A schematic diagram of a customer service interface provided by an illustrative embodiment of the present application;
[0015] Figure 2 A schematic diagram of an interaction interface displayed in a floating window form provided by an illustrative embodiment of the present application;
[0016] Figure 3 An interface state schematic diagram of an interaction interface in a hidden state provided by an illustrative embodiment of the present application;
[0017] Figure 4 An interface state schematic diagram of an AI voice assistant provided by an illustrative embodiment of the present application;
[0018] Figure 5 An interface state schematic diagram of a refund card provided by an illustrative embodiment of the present application;
[0019] Figure 6 An interface state schematic diagram of a dialogue selection and sending provided by an illustrative embodiment of the present application;
[0020] Figure 7 An interface state schematic diagram of voice input in a question and answer area provided by an illustrative embodiment of the present application;
[0021] Figure 8Fig. 1 is a schematic diagram of an interface state for displaying at least one user session information according to an example embodiment of the present application;
[0022] Figure 9 Fig. 2 is a schematic diagram of interface states before and after session switching according to an example embodiment of the present application;
[0023] Figure 10 Fig. 3 is a flowchart of another information interaction method according to an example embodiment of the present application;
[0024] Figure 11 Fig. 4 is a schematic diagram of an information interaction device according to an example embodiment of the present application;
[0025] Figure 12 Fig. 5 is a schematic diagram of an electronic device according to an example embodiment of the present application. DETAILED DESCRIPTION
[0026] In order to make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described below in detail with reference to the embodiments of the present application and corresponding drawings. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0027] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of the country and region, and provide corresponding operation portal for user to choose authorization or refusal. In addition, the various models (including but not limited to language models or large models) involved in the present application are in line with relevant legal and standard regulations.
[0028] In some existing customer service scenarios, there may be situations where it is not flexible to use or temporarily inconvenient to use peripherals such as keyboards and mice for interaction. For example, when a disabled customer service personnel interacts with a terminal device to answer various questions for users, due to various physical disabilities, such as the inability to flexibly use peripherals such as keyboards and mice, there is a problem of inconvenience in operation, resulting in low customer service efficiency and low user satisfaction. For another example, when a non-disabled customer service personnel is holding something or occupied, it may be temporarily inconvenient to use peripherals such as keyboards and mice, which will also result in low customer service efficiency and low user satisfaction.
[0029] To solve the above technical problems, an AI (Artificial Intelligence) voice assistant solution is provided for a customer service personnel. An AI voice assistant and a customer service program are installed on a terminal device. The AI voice assistant can be associated with the customer service program. The customer service personnel can make various interactive controls on a customer service interface provided by the customer service program with the assistance of the AI voice assistant through simple voice input. Thus, the customer service personnel can still provide customer service for users more efficiently in the case of being unable to flexibly use or inconvenient to use peripherals such as a keyboard and a mouse, and user experience is improved.
[0030] The technical solutions provided by the embodiments of the present application are described in detail below with reference to the drawings.
[0031] Figure 1a A flowchart of an information interaction method provided by an example embodiment of the present application is shown. The method can be executed by an AI voice assistant on a terminal device. The terminal device can be a smart phone, a tablet computer, a computer, a smart watch or the like. The terminal device can be configured with the AI voice assistant. The AI voice assistant is associated with a customer service program on the terminal device.
[0032] In the embodiments of the present application, the product implementation forms of the AI voice assistant and the customer service program are not limited. The AI voice assistant and the customer service program can be implemented as independent program products. For example, the customer service program is a website or an APP (Application), and the AI voice assistant is implemented as an APP. During use of the customer service program, the AI voice assistant can be started to take over part of the functions of the customer service program to assist the customer service personnel to complete information interaction, for example, to display a customer service interface and perform related control operations required for assisting the customer service personnel to complete information interaction on the basis of the customer service interface. Alternatively, the AI voice assistant and the customer service program can be implemented as one program product. For example, the AI voice assistant is embedded or integrated in the customer service program. The AI voice assistant can be implemented as a functional module of the customer service program to assist the customer service personnel to complete interaction. Alternatively, the AI voice assistant and the customer service program can be implemented as independent program products coupled with each other. For example, the AI voice assistant and the customer service program are independent program products, but an access portal to the customer service program is embedded in the AI voice assistant. The access portal can be an access link provided by the customer service program to the outside, so as to enter the customer service program through the AI voice assistant, display a customer service interface provided by the customer service program, and perform related control operations required for assisting the customer service personnel to complete information interaction on the basis of the customer service interface.
[0033] Regardless of the product implementation forms mentioned above, the process of the AI voice assistant assisting the customer service personnel to complete information interaction is the same or similar. Specifically, as shown in Figure 1a the method can include the following steps:
[0034] Step 11, display the customer service interface provided by the customer service program, and the customer service interface includes a question and answer area for displaying at least one target question information of the target user for consulting the customer service personnel.
[0035] Step 12, based on the voice recognition module in the AI voice assistant, receive at least one voice interaction instruction initiated by the customer service personnel for any target question information.
[0036] Step 13, for any voice interaction instruction, execute the voice interaction instruction and output the execution result information of the voice interaction instruction in a display and / or voice manner.
[0037] Step 14, for a target voice interaction instruction in the at least one voice interaction instruction, select at least part of the execution result information of the target voice interaction instruction as target reply information of any target question information and display in the question and answer area.
[0038] In this embodiment, the AI voice assistant can be used to assist the customer service personnel to make interactive control on the customer service interface. The AI voice assistant can include a non-manual recognition module. The non-manual recognition module refers to a module that does not rely on manual input for information recognition, and can rely on other information input modes other than manual input for information recognition. The non-manual recognition module at least includes a voice recognition module, which refers to a module that recognizes information based on voice input. Further optionally, the non-manual recognition module can also include a head movement recognition module, which refers to a module that recognizes information based on head movements or movements of other parts on the head (such as eye movements, mouth movements, etc.). In this embodiment, the voice recognition module is specifically used to recognize the voice interaction instruction issued by the customer service personnel during the running of the AI voice assistant. The head movement recognition module is specifically used to recognize the head movement or movement of other parts on the head issued by the customer service personnel during the running of the AI voice assistant, and further recognize the instruction intended to be issued by the customer service personnel. In this embodiment, the AI voice assistant can release the customer service personnel from manual input operation through these non-manual recognition modules, so that the customer service personnel can quickly complete information input on the customer service interface in a non-manual input mode such as voice, head movement, etc., and improve the interaction efficiency.
[0039] In this embodiment, the AI voice assistant can display the customer service interface provided by the customer service program, as shown in Figure 1b The customer service interface at least includes a question and answer area, and further can include a function area and a conversation area. Herein, Figure 1bThe customer service interface style shown is only an example and is not limited thereto. The question and answer area is used to display at least one target question information of a target user for consulting a customer service personnel. In this embodiment, the user currently responsible for by the customer service personnel is referred to as the target user, and the question information consulted by the target user is referred to as the target question information, and the target question information is used to describe the question consulted by the target user to the customer service personnel. It should be noted that according to different customer service scenarios, the target user and the target question information consulted by the target user will be different. For example, if it is an e-commerce scenario, the target user can be a consumer of an e-commerce platform, and the target question information can be question information related to goods, logistics, orders, refunds, returns, etc. For example, the target question information can be "When can this product be shipped?" or "Why not return it?" and so on. If it is a tourism scenario, the target user can be a tourist or a traveler, and the target question information can be question information related to travel consultation, hotel reservation, scenic spot inquiry, etc. If it is a financial industry, the target user can be a financial or savings user, and the target question information can be question information related to account inquiry, transaction consultation, loan application, etc.
[0040] Among them, the target user can consult multiple target question information to the customer service personnel, and the multiple target question information can be independent of each other or can have an association relationship, for example, the next target question information is a subsequent question of the previous target question information. For multiple target question information, the target user can consult the customer service personnel at the same time or in the same batch, or can consult the customer service personnel in batches or at different times, which is not limited.
[0041] In this embodiment, the target user consults the target question information to the customer service personnel through the customer service interface displayed on the terminal device of the target user, the terminal device obtains the target question information and sends the target question information to the server, the server sends the target question information to the customer service program on the terminal device of the customer service personnel, and the customer service program displays the target question information in the question and answer area of the customer service interface to present to the customer service personnel. Correspondingly, the customer service personnel can view at least one target question information through the question and answer area of the customer service interface, and can process any target question information. In this embodiment, the processing of the target question information by the customer service personnel mainly refers to the process of obtaining reply information related to the target question information and sending the reply information to the question and answer area for the target user. Among them, according to different target question information, the customer service personnel obtains the reply information related to the target question information in different ways, but the entire reply process needs the customer service personnel to interact with the terminal device.
[0042] In this embodiment, in order to simplify the interactive operation of the customer service personnel, an AI voice assistant is provided, on the one hand, so that the customer service personnel can input information to the terminal device in a simple voice manner, and on the other hand, the AI voice assistant can process the reply information instead of the customer service personnel, and the AI voice assistant assists the customer service personnel from the information input link to the problem processing link, thereby simplifying the interactive operation required by the customer service personnel to reply to the target problem information. Based on this, when replying to any target problem information, the customer service personnel can initiate a voice interaction instruction to the AI voice assistant for any target problem information. The number of voice interaction instructions initiated by the customer service personnel for any target problem information can be one or more. In the case of initiating multiple voice interaction instructions for any target problem information, the multiple voice interaction instructions can be initiated in sequence, and further optionally, the multiple voice interaction instructions are all for the purpose of serving the any target problem information. The AI voice assistant can receive at least one voice interaction instruction initiated by the customer service personnel for any target problem information based on a voice recognition module, which can call the audio module (such as a microphone) of the terminal device to receive the voice interaction instruction initiated by the customer service personnel; then, for any received voice interaction instruction, any voice interaction instruction is executed, and the execution result information of any voice interaction instruction is output. The execution result information is used to describe the result obtained after executing the voice interaction instruction of the customer service personnel. In the case of multiple voice interaction instructions, the AI voice assistant can receive and execute them one by one according to the order of the multiple voice interaction instructions.
[0043] In this embodiment, in addition to supporting voice functions, the AI voice assistant can also integrate a large language model for information processing. The large language model (LLM) is a deep learning model trained based on massive text data. In this embodiment, the large language model is referred to as a target language model. The AI voice assistant can use the target language model to process information sent by the customer service personnel or the user, for example, when executing part of the voice interaction instruction, which is not limited. The target language model can be a language model developed according to a specific customer service scenario, or a language model obtained by fine-tuning an existing large language model according to a specific customer service scenario, or an existing large language model can be directly used, which is not limited. The existing large language models that can be used in this embodiment include but are not limited to: LaMDA (Language Models for Dialog Applications) model, LLaMa (Large Language Model Family), Qwen-Max model, etc. The detailed process of the AI voice assistant executing the voice interaction instruction can be referred to the description of the subsequent embodiments, which is not described here.
[0044] The output execution result information is used to assist the customer service personnel to determine the target reply information of the target question information. The customer service personnel can generate the reply information of the target question information according to the execution result information of any voice interaction instruction, or the AI voice assistant can directly obtain the reply information of the target question information from the execution result information instead of the customer service personnel, for example, the AI voice assistant can take the whole execution result information or part of the execution result information as the reply information of the target question information.
[0045] The output execution result information can be in a display mode, that is, the execution result information is displayed, for example, on the current customer service interface or the interactive interface provided by the AI voice assistant, which depends on the type of the voice interaction instruction. The output execution result information can also be in a voice mode, that is, the execution result information is output in a voice broadcast mode. The output execution result information can also be in a combination of the display mode and the voice mode, that is, the execution result information is displayed and output in a voice broadcast mode. Alternatively, a corresponding relationship between the type of the voice interaction instruction and the output mode of the execution result information can be established in advance, and based on this, the specific mode of the output execution result information can be determined according to the type of the voice interaction instruction. Alternatively, the display mode, the voice mode, or the combination of the display mode and the voice mode can be used by default.
[0046] In the embodiment, the at least one voice interaction instruction initiated by the customer service personnel for any target question information can include: a voice interaction instruction that needs to be replied to the target user and a voice interaction instruction that does not need to be replied to the target user. The voice interaction instruction that needs to be replied to the target user is referred to as a target voice interaction instruction; otherwise, the voice interaction instruction that does not need to be replied to the target user is referred to as a non-target voice interaction instruction. For example, the target voice interaction instruction that needs to be replied to the target user can be a refund instruction or a script instruction, etc. The refund instruction is used to instruct the AI voice assistant to obtain refund operation information and send the refund operation information to the target user to guide the user to initiate a refund operation. The script instruction is used to instruct the AI voice assistant to generate relevant script information and select script information matching the target question information issued by the target user from the script information as reply information and send the script information to the target user. The non-target voice interaction instruction that does not need to be replied to the target user can be a background data calling instruction, a page calling instruction, a search engine calling instruction or a customer service program login instruction, etc. The background data calling instruction can be used to call and display the data stored in the background for the customer service personnel to make corresponding decisions based on the background data. The page calling instruction can be used to call certain subpages in the customer service program to obtain relevant information from the subpages, such as a product detail page, etc. The search engine calling instruction is used to call a search engine to assist the customer service personnel to make corresponding decisions. Unlike the target voice interaction instruction, these voice interaction instructions do not need to be replied to the target user.
[0047] In the embodiment, for any target question information, at least part of the execution result information of the target voice interaction instruction in the at least one voice interaction instruction initiated by the customer service personnel for the any target question information can be selected as the target reply information of the any target question information and displayed in the question and answer area. The entire execution result information of the target voice interaction instruction can be selected as the target reply information of the any target question information and displayed in the question and answer area, or part of the execution result information of the target voice interaction instruction can be selected as the target reply information of the any target question information and displayed in the question and answer area, which can be determined according to the type of the target voice interaction instruction.
[0048] For example, in a case where the target voice interaction instruction is a script instruction, the execution result information corresponding to the target voice interaction instruction can include multiple script information, script information 1 is "I am sorry for the inconvenience, please forgive me", script information 2 is "I will help you query the logistics, please wait", and script information 3 is "We will refund you as soon as possible, please wait patiently". The customer service personnel can select a certain script information, for example, script information 1, as the target reply information of the target question information and display it in the Q&A area. For example, in a case where the target voice interaction instruction is a refund instruction, the execution result information corresponding to the target voice interaction instruction can be a refund card, and the AI voice assistant can display the refund card as the target reply information of the target question information in the Q&A area.
[0049] In this embodiment, an AI voice assistant associated with a customer service program is provided for a customer service personnel, which is responsible for displaying a customer service interface provided by the customer service program on one hand, and can respond to and execute a voice interaction instruction of the customer service personnel to give execution result information, and for a voice instruction that needs to be replied to the user, can select content that can be used as reply information from the execution result information and display it in the Q&A area of the customer service interface, assisting the customer service personnel to complete the interaction with the user. In this process, the customer service personnel can complete the interaction through simple voice input with the assistance of the AI voice assistant, which can solve the problem that the interaction cannot be flexibly performed by using peripherals due to the inconvenience of manual input, and further solve the problem of low interaction efficiency caused thereby. This scheme is applicable to various interaction scenarios, especially interaction scenarios involving disabled personnel.
[0050] In the embodiments of the present application, the AI voice assistant can serve as an assistant for the customer service personnel, execute various voice interaction instructions issued by the customer service personnel, and provide corresponding services for the customer service personnel. In the embodiments of the present application, the voice interaction instructions issued by the customer service personnel can be divided into different types of instructions according to the types of services provided by the AI voice assistant for the customer service personnel, such as a first type of instruction, a second type of instruction, and a third type of instruction.
[0051] The first type of instruction can make the AI voice assistant execute a corresponding transaction in the data source associated with the customer service interface based on the target voice model, such as a refund information pushing transaction and a historical price query transaction. After executing the corresponding transaction, the transaction execution result can be displayed in the function area of the customer service interface, such as a refund card corresponding to the transaction execution result of the refund information pushing transaction, and a price card corresponding to the transaction execution result of the historical price query transaction. It is explained that the AI voice assistant can obtain the relevant information of the target user currently displayed in the customer service interface, and has access to and operation on the data source associated with the customer service interface, and on this basis, the first type of instruction issued by the customer service personnel can be executed to execute a corresponding transaction in the data source. In this embodiment, the relevant information of the target user includes but is not limited to: identification information, account information, order information, etc. of the target user; the data source associated with the customer service interface includes but is not limited to: order information system, logistics information system, commodity information system, and customer service knowledge base, etc.; the customer service knowledge base stores various question information and corresponding solutions, such as order exception and its solution (such as canceling the order and initiating the refund process), logistics exception and its solution (such as canceling the order and initiating the refund process), etc.
[0052] The first type of instruction can be understood as a standard question type instruction, that is, a standard question type instruction, such as an order status query instruction, a logistics information query instruction, a historical price query instruction, an evaluation information query instruction, a promotion activity query instruction, a refund information query instruction, and a pre-sale size recommendation instruction, etc. The target voice model in the AI voice assistant can be associated with a data source in the customer service interface, such as historical refund data in the data source which can be associated with the refund control of the customer service interface, and historical price data in the data source which can be associated with the historical price viewing control of the customer service interface.
[0053] The second type of instruction is a voice instruction for the AI voice assistant to generate corresponding speech information based on the target voice model and display the speech information on the interactive interface provided by the AI voice assistant. In other words, the AI voice assistant can respond to the second type of instruction to generate corresponding speech information, so as to be displayed on the interactive interface provided by the AI voice assistant for the customer service personnel to select the speech and reply to the target user. For example, the second type of instruction can be an appeasement speech generation instruction, and the AI voice assistant can respond to the second type of instruction to generate the following speech information using the target voice model: "Please select one of the following replies: 1, We are deeply sorry for the misunderstanding; 2, I am sorry for the inconvenience, please forgive me". The second type of instruction can also be an end speech generation instruction, and the AI voice assistant can respond to the second type of instruction to generate the following speech information using the target voice model: "Please select one of the following replies: 1, Welcome to come again next time; 2, Thank you for your support, I hope you have a pleasant shopping experience".
[0054] The third type of instruction is a standard instruction directly executed by the AI voice assistant, such as an AI voice assistant interface hiding instruction, a session switching instruction, an AI voice assistant closing instruction, an upstroke instruction, a downstroke instruction, a favorites opening instruction, a hosting instruction, a click instruction, a work instruction, or an off-duty instruction, and the like. The work instruction can be understood as a customer service program login instruction, and the off-duty instruction can be understood as a customer service program exit instruction or a terminal shutdown instruction.
[0055] Further optionally, the voice interaction instruction of the embodiment can further include a fourth type of instruction and a fifth type of instruction. The fourth type of instruction is an instruction sent by the AI voice assistant to the terminal device, such as an assistant update instruction or an assistant restart instruction, and the like. The fifth type of instruction is an instruction sent by the AI voice assistant to the customer service program, such as a customer service program restart instruction or a customer service program update instruction.
[0056] In some optional embodiments, the AI voice assistant can remain in a dormant state before voice interaction with the customer service personnel. Before receiving at least one voice interaction instruction initiated by the customer service personnel for any target problem information, the AI voice assistant can also be awakened based on any of the following embodiments:
[0057] Embodiment one, based on a voice recognition module, in response to a wake-up instruction issued by the customer service personnel through voice, the AI voice assistant is awakened. Specifically, the customer service personnel can issue a wake-up instruction through voice, the voice recognition module can perform voice recognition on the wake-up instruction to obtain text information corresponding to the wake-up instruction, and compare the text information corresponding to the wake-up instruction with a preset wake-up text. If the matching degree between the text information corresponding to the wake-up instruction and the preset wake-up text is greater than a preset matching degree threshold, the AI voice assistant can be awakened. The preset wake-up text refers to the text information required to wake up the AI voice assistant, which can be, for example, "Hello, Special Assistant", but is not limited thereto.
[0058] In the embodiment, the head movement recognition module can capture the face image or video of the customer service personnel by calling the camera of the terminal device, and detect the head key points (such as eyes, nose tip, and mouth corner) of the face image or video, and then estimate the head posture (such as the pitch angle or the yaw angle) according to the positions of the detected head key points. The head movement recognition module can determine the line-of-sight gaze position of the customer service personnel in the current head posture according to the head posture of the customer service personnel and the mapping relationship between the head posture and the line-of-sight gaze position, and if the line-of-sight gaze position is on the icon of the AI voice assistant, the icon of the AI voice assistant can be simulated to be clicked by the customer service personnel to wake up the AI voice assistant. Alternatively, the head movement recognition module can also recognize the type of head movement according to the estimated head posture, for example, nodding or shaking, and if the type of head movement is the type of movement required to wake up the AI voice assistant (for example, nodding movement), the AI voice assistant can be woken up.
[0059] After the AI voice assistant is woken up in the above manner, the interactive interface provided by the AI voice assistant can be displayed, and the text information corresponding to the wake-up instruction, such as "Hello, Special Assistant", can be displayed on the interactive interface. The interactive interface can be displayed in any position above the customer service interface in the form of a floating window or a drawer, and the embodiment does not limit the display form and display position of the interactive interface. For example, Figure 2 Figure 2 The case of displaying the interactive interface in the form of a floating window is exemplarily shown. Further, the AI voice assistant can actively ask the customer service personnel what help is needed after being woken up, and the inquiry information can be output in the form of text display and voice broadcast at the same time, or only in one of the forms, which is not limited. In the interactive interface shown in Figure 2 , the words "Hello, what help do you need" are displayed, which is the inquiry information initiated by the AI voice assistant to the customer service personnel after being woken up.
[0060] Further, during the running of the AI voice assistant, the customer service personnel can hide the interactive interface at any time according to the use demand by using a hiding instruction. The state switching process of the interactive interface between the hidden state and the displayed state will be described below in combination with Figure 3 and Figure 4 . For example, Figure 3 Before receiving the first type of instruction, if the interactive interface is in the displayed state, the interactive interface can be hidden in response to the hiding instruction issued by the customer service personnel through voice or head movement, and the icon of the AI voice assistant can be displayed on the customer service interface. The display position of the icon of the AI voice assistant on the customer service interface is not limited in the embodiment, and the icon can be displayed at any position on the customer service interface, such as the upper right corner or the lower left corner. Alternatively, as shown in Figure 4 As shown, before receiving other types of instructions, if the interactive interface is in a hidden state, the interactive interface is displayed in response to the wake-up instruction issued by the customer service personnel through voice or head movement, and the text information corresponding to the wake-up instruction is displayed on the interactive interface. In this way, the customer service personnel only needs to make a simple voice input or move the head to complete the switching between the hidden state and the display state of the interactive interface.
[0061] In this way, the customer service personnel can be freed from manual input operations, and the customer service personnel can quickly wake up the AI voice assistant in a voice manner or a head movement manner, thereby improving the interaction efficiency.
[0062] In some optional embodiments, the foregoing step 13 "for any voice interaction instruction, executing any voice interaction instruction, and outputting the execution result information of any voice interaction instruction in a display and / or voice manner", can be implemented based on the following steps:
[0063] Step 131, identifying the instruction type of any voice interaction instruction according to the semantic information of any voice interaction instruction and / or the type marking word contained in any voice interaction instruction. The semantic information of the voice interaction instruction is used to describe the meaning, context and semantic relationship of the voice interaction instruction. The type marking word of the text instruction refers to the word in the plurality of words obtained by performing word segmentation processing on the text instruction, which is used to mark the type of the voice interaction instruction. The type marking word refers to the marking word appearing in the voice interaction instruction for distinguishing the types of the voice interaction instruction. In succession to the first type of instruction, the second type of instruction and the third type of instruction divided above, the type marking words corresponding to the three types of instructions can be "question mark", "dialogue" and "operation" in turn, or can be "A", "B", "C" and any other marking information capable of distinguishing the three types of instructions. Taking "question mark", "dialogue" and "operation" as an example, the first type of instruction issued by the customer service personnel includes "question mark" and specific instruction content; the second type of instruction issued by the customer service personnel includes "dialogue" and specific instruction content; and the third type of instruction issued by the customer service personnel includes "operation" and specific instruction content. A specific example, if the customer service personnel wants to query logistics information, the customer service personnel can issue "question mark: query logistics".
[0064] Among them, the instruction type of any voice interaction instruction can be identified according to the semantic information of any voice interaction instruction, the instruction type of any voice interaction instruction can also be identified according to the type marking word contained in any voice interaction instruction, and the instruction type of any voice interaction instruction can also be identified according to the semantic information of any voice interaction instruction and the type marking word contained in any voice interaction instruction. The present embodiment does not make any limitation.
[0065] Different instruction types have different execution manners. For example, some instruction types are executed in a manner dependent on the target language model, such as generating dialogue information by using the target language model; some instruction types are executed in a manner independent of the target language model, such as calling historical price information in the data source corresponding to the customer service interface and displaying the information. For another example, some instruction types are executed in a manner independent of the data source corresponding to the customer service interface, such as generating emotion pacification information for the customer service personnel by using the target language model; some instruction types are executed in a manner dependent on the data source corresponding to the customer service interface, such as calling historical refund information in the data source corresponding to the customer service interface and displaying the information.
[0066] Different instruction types have different result display manners. For example, if the instruction is of the first type, the result display manner is to be displayed in the function area of the customer service interface, and if the instruction is of the second type or the third type, the result display manner is to be displayed on the interactive interface of the AI voice assistant.
[0067] In step 131, when identifying the instruction type of any voice interaction instruction according to the semantic information of the voice interaction instruction, the voice recognition module can be used to perform text information conversion processing on any voice interaction instruction, that is, to convert the voice interaction instruction in audio form into text form to obtain a text instruction corresponding to any voice interaction instruction. Based on the semantic information of the text instruction, the instruction type of any voice interaction instruction is identified. The semantic information of the text instruction can be extracted and the instruction type can be identified by the following steps: using a tokenizer to split the text instruction into multiple text units (such as words), and annotating the grammatical role (noun, verb, adjective, etc.) of each word; constructing a syntax tree or dependency relationship to understand the relationship between sentence components, such as subject-predicate-object and modification relationship; performing entity recognition (Named Entity Recognition, NER) to extract specific information such as names, places, organizations, times, and proper nouns; extracting entity relationship, such as what someone did and when and where, to deeply understand the context; understanding the sentiment polarity (positive, negative, neutral) of the text; semantic role labeling to understand the meaning of words in the sentence context, such as the metaphorical meaning of verbs; semantic parsing: deeply understanding the context by using the target language model and associating the semantic relationship between words. After the semantic information of the text instruction is extracted by the above steps, the instruction type of any voice interaction instruction can be identified based on the semantic information of the text instruction.
[0068] In step 131, when identifying the instruction type of any voice interaction instruction according to the type mark word contained in any voice interaction instruction, the text information conversion processing can be performed on any voice interaction instruction based on the voice recognition module to obtain the text instruction corresponding to any voice interaction instruction; and the instruction type of any voice interaction instruction is identified based on the type mark word contained in the text instruction. Specifically, the text instruction can be split into multiple words by using a word segmenter, and the instruction type of the voice interaction instruction is determined according to the type mark word in the multiple words, for example, the type mark word "refund" is contained in the multiple words split from the text instruction, then the instruction type of the voice interaction instruction is determined as the first type of instruction according to the type mark word. Through the above-mentioned manner, the instruction type of any voice interaction instruction can be accurately identified.
[0069] In step 132, according to the instruction type of any voice interaction instruction, the execution of any voice interaction instruction is performed in the execution mode adapted to the instruction type of any voice interaction instruction to obtain the execution result information of any voice interaction instruction. Specifically, it can include the following multiple cases:
[0070] Case 1, in the case where any voice interaction instruction is the first type of instruction, the text instruction is input into the target language model in the AI voice assistant, so that the target voice model executes the transaction corresponding to any voice interaction instruction in the data source associated with the customer service interface and outputs the transaction execution result as the execution result information of any voice interaction instruction. For example, in the case where the voice interaction instruction is a logistics information query instruction, the target voice model can search for logistics information from the data source associated with the customer service interface according to the text instruction corresponding to the voice interaction instruction, and output the searched logistics information as the execution result information of the voice interaction instruction. For another example, in the case where the voice interaction instruction is a refund information query instruction, the target voice model can search for refund information from the data source associated with the customer service interface according to the text instruction corresponding to the voice interaction instruction, and output the searched refund information as the execution result information of the voice interaction instruction.
[0071] Case 2, in the case where any voice interaction instruction is the second type of instruction, the text instruction is input into the target language model in the AI voice assistant for dialogue requirement identification and dialogue information generation processing, and at least one piece of dialogue information corresponding to any voice interaction instruction is output as the execution result information of any voice interaction instruction. Wherein, the target voice model can identify the dialogue requirement existing in the text instruction according to the semantic information of the text instruction, and determine the dialogue information corresponding to the dialogue requirement from the preset dialogue library as the execution result information according to the dialogue requirement.
[0072] In case 3, if any of the voice interaction instructions is a third type of instruction, the text instruction is matched in the standard type instruction library supported by the AI voice assistant. Specifically, the text instruction can be similarity matched with the standard type instructions stored in the standard type instruction library. If there is a standard type instruction in the standard type instruction library that has a similarity greater than a preset similarity threshold with the text instruction, the standard type instruction can be taken as a target standard type instruction. If there is no standard type instruction in the standard type instruction library that has a similarity greater than the preset similarity threshold with the text instruction, it indicates that there is no target standard type instruction in the standard type instruction library.
[0073] If the target standard type instruction is matched, the target standard type instruction can be executed to obtain the execution result information of any of the voice interaction instructions. If the target standard type instruction is not matched, the text instruction can be input into a target language model in the AI voice assistant for instruction conversion processing. If the text instruction is converted into a standard type instruction, the converted standard type instruction is output. Specifically, the target language model can determine a standard type instruction that is most semantically close to the semantic information of the text instruction from the standard type instructions stored in the standard type instruction library.
[0074] If the text instruction cannot be converted into a standard type instruction, i.e., there is no standard type instruction in the standard type instruction library that is semantically close to the semantic information of the text instruction, in this case, the prompt information can be output to the customer service personnel in the following ways: way one, voice output of the prompt information to remind the customer service personnel to re-input any of the voice interaction instructions; way two, display of the prompt information on the interactive interface to remind the customer service personnel to re-input any of the voice interaction instructions; way three, voice output of the prompt information and display of the prompt information on the interactive interface to remind the customer service personnel to re-input any of the voice interaction instructions.
[0075] Step 133, according to the instruction type of any of the voice interaction instructions, the execution result information of any of the voice interaction instructions is displayed in the function area of the customer service interface or on the interactive interface provided by the AI voice assistant. Specifically, according to the correspondence between the instruction type and the result display mode, the result display mode corresponding to the instruction type can be determined, so that the execution result information can be displayed in the function area of the customer service interface or on the interactive interface provided by the AI voice assistant. In some optional embodiments, in the case where any of the voice interaction instructions is a first type of instruction, the execution result information of any of the voice interaction instructions can be displayed in the function area of the customer service interface. The function area is different from the question and answer area. In the case where any of the voice interaction instructions is another type of instruction, the text information corresponding to any of the voice interaction instructions and the execution result information thereof can be displayed on the interactive interface, and the interactive interface is located above the customer service interface.
[0076] In this way, the voice interaction instruction can be executed in different execution manners and result display manners according to different instruction types, and the execution result information can be displayed in a personalized manner.
[0077] In some optional embodiments, the target voice interaction instruction can include at least one of the following voice instructions: a voice instruction in the at least one voice interaction instruction that belongs to the first type of instruction and whose execution result information contains the sending control, and a voice instruction in the at least one voice interaction instruction that belongs to the second type of instruction.
[0078] The AI voice assistant can identify the target voice interaction instruction in the following manner: for any voice interaction instruction in the at least one voice interaction instruction, if the voice interaction instruction is of the first type and its execution result information contains the sending control, the voice interaction instruction is taken as the target voice interaction instruction; or if the voice interaction instruction is of the second type, the voice interaction instruction is taken as the target voice interaction instruction.
[0079] Based on this, for the target voice interaction instruction in the at least one voice interaction instruction, at least part of the execution result information of the target voice interaction instruction is selected as the target reply information of any target question information and displayed in the Q&A area, which can include the following cases:
[0080] Case 4: In the case where the target voice interaction instruction belongs to the first type of instruction and its execution result information contains the sending control, the execution result information of the target voice interaction instruction is displayed in the function area of the customer service interface, and the customer service personnel can issue a sending instruction in a voice manner, or trigger the sending control in a head movement manner, or trigger the sending control in a manual clicking manner. In response to the sending instruction issued by the customer service personnel in a voice or sending control manner, part of the execution result information of the target voice interaction instruction associated with the sending control is directly sent from the function area to the Q&A area as the target reply information.
[0081] Figure 5 Exemplarily, the case where the target voice interaction instruction belongs to the first type of instruction, i.e., the refund information query instruction, is shown. Figure 5 As shown, in response to the sending instruction issued by the customer service personnel in a voice or sending control 1 manner, the refund card in the execution result information of the target voice interaction instruction associated with the sending control 1 is directly sent from the function area to the Q&A area as the target reply information.
[0082] Case 5, in the case of the target voice interaction instruction belonging to the second type of instruction, at least one piece of dialogue information corresponding to the target voice interaction instruction is displayed on the interaction interface provided by the AI voice assistant. Optionally, at least one piece of dialogue information can be directly displayed on the interaction interface, at least one piece of dialogue information can be dynamically displayed on the interaction interface, or the voice input operation of the customer service personnel can be simulated, and the microphone control in the interaction interface is pressed. This embodiment is not limited.
[0083] In response to the selection instruction issued by the customer service personnel in the form of voice, the target dialogue information is selected from the dialogue information, and the target dialogue information is displayed in the input box corresponding to the question and answer area. In response to the sending instruction issued by the customer service personnel, the target dialogue information in the input box is sent to the question and answer area. Among them, the customer service personnel can issue the sending instruction in the form of voice, or can trigger the sending key in the input box by head movement to issue the sending instruction, or can issue the sending instruction by manually clicking the sending key, and this embodiment is not limited. As shown in Figure 6 , the customer service personnel can select the dialogue information M2 "very sorry" from the dialogue information M1 to dialogue information M3 as the target dialogue information, and issue the sending instruction to send the target dialogue information in the input box to the question and answer area.
[0084] Case 6, for at least one non-target voice interaction instruction belonging to the first type of instruction in the voice interaction instruction, the execution result information of the non-target voice interaction instruction can be displayed in the function area of the customer service interface, as Figure 7 shown. Among them, the non-target voice interaction instruction refers to the voice instruction that does not need to reply to the target user. For example, since the historical price can be referred to by the customer service personnel, it does not need to be sent to the user, so the non-target voice interaction instruction of the first type of instruction can be a historical price query instruction, in addition to this, the non-target voice interaction instruction of the first type of instruction can also be a search engine result acquisition instruction, a customer service personnel information acquisition instruction, etc.
[0085] In response to the voice input operation initiated by the customer service personnel in the input box of the question and answer area according to the execution result information of the non-target voice interaction instruction, the text information corresponding to the voice information input by the customer service personnel is displayed in the input box. In other words, the customer service personnel can refer to the execution result information (such as search engine results or historical prices, etc.) of the non-target voice interaction instruction, so as to make decisions with the aid of the execution result information and input corresponding voice information.
[0086] In response to the sending instruction issued by the customer service personnel, the text information in the input box is sent to the question and answer area as the target reply information of any target question information. Among them, the customer service personnel can initiate a voice input operation by triggering the voice input control in the question and answer area in the form of voice or head movement, and this embodiment is not limited.
[0087] Through the above case 4-case 6, the customer service personnel can select at least part of the execution result information of the target voice interaction instruction as the target reply information of any target question information in an efficient and fast interactive manner, and display the target reply information in the Q&A area.
[0088] In some optional embodiments, executing the target standard class instruction to obtain the execution result information of any voice interaction instruction can include the following cases:
[0089] Case 7, in the case where the target standard class instruction is a work instruction, at least one user session information that needs to be taken care of by the customer service personnel is displayed in the conversation area of the customer service interface, such as Figure 8 As shown. Any user session information can include the nickname of any user and the latest message sent by the user.
[0090] Case 8, in the case where the target standard class instruction is a session switching instruction, the switching of user session information in the Q&A area of the customer service interface is performed, and the user session information at least includes the target question information of the user to which the switching is currently performed. As Figure 9 As shown, the dashed box represents the currently selected user session information, the switching of user session information in the Q&A area of the customer service interface is performed, that is, user session information 1 is switched to user session information 2, so that the target question information of user 1 displayed in the Q&A area is switched to the target question information of user 2.
[0091] Based on the above case 7 and case 8, the customer service personnel can control the customer service interface more quickly and efficiently.
[0092] It is explained here that considering that in the service process of the customer service personnel, the target user may encounter a very dissatisfied situation, the target user may verbally abuse the customer service personnel, which may cause emotional fluctuations of the customer service personnel. Therefore, the AI voice assistant provided in the embodiments of the present application can not only serve as an assistant of the customer service personnel, but also as a good friend of the customer service personnel to appease the emotions of the customer service personnel. The following will be further explained:
[0093] When any target question contains a specific emotional word, an emotional reassurance message can be generated for customer service personnel based on a target language model. Here, the specific emotional word refers to sensitive words that affect the customer service personnel's emotions. The target language model can extract semantic information from the target question containing the specific emotional word and select a target emotional reassurance statement that matches the semantic information from a set of preset emotional reassurance statements. This emotional reassurance message can then be displayed on the interactive interface provided by the AI voice assistant to soothe the customer service personnel's emotions. For example, when any target question contains a specific emotional word, the interactive interface provided by the AI voice assistant can display the emotional reassurance message, "Dear customer, the customer is currently quite agitated. Please calm down. I will always be there for you. Keep going!" to soothe the customer service personnel's emotions.
[0094] It should be noted that, considering customer service personnel may suddenly become unwell or need to temporarily leave their posts for other reasons during service, the AI voice assistant provided in this application embodiment can not only serve as an assistant and friend to customer service personnel, but also as an emergency care provider, autonomously acting as a customer service representative. This will be further explained below:
[0095] When customer service staff need to leave their posts due to illness or other reasons, they can issue an emergency takeover command via voice to allow the AI voice assistant to continue completing the corresponding work tasks. Correspondingly, the AI voice assistant can respond to the emergency takeover command and switch to emergency work mode. In emergency work mode, it can identify unanswered target questions in the question-and-answer area, generate target responses based on the target language model, and display the target responses in the question-and-answer area. Specifically, the target language model can extract semantic information from the unanswered target questions and select a target response that matches the semantic information from a set of preset response statements. The target language model can be pre-trained on a large number of training samples to learn the contextual relationships of question information and related response rules, and then generate target responses for unanswered target questions based on these learned contextual relationships and response rules. Furthermore, the target language model can classify the unanswered target questions, thereby generating target responses based on the type of target question and the corresponding response method. For example, when dealing with a specific type of target question information, a three-part response format can be used: "summary answer - detailed answer - reference link"; or, when dealing with a specific type of target question information, a reference link or summary answer can be provided directly.
[0096] The main functions of the AI voice assistant provided by the embodiments of the present application are introduced and described above. The complete working process of the AI voice assistant is described in detail below by taking the example of embedding the AI voice assistant in a customer service platform (a customer service program implemented in the form of a website):
[0097] The terminal device displays a home page of the customer service platform and a login control. The customer service personnel can use a head movement method to trigger the login control to log in. After logging in, the customer service interface can be displayed. The customer service personnel can issue a "hello, special assistant" wake-up instruction in a voice manner to wake up the AI voice assistant to display an interaction interface. The customer service personnel can issue a "I want to work" work instruction in a voice manner to display the customer service interface, display at least one conversation information that needs to be handled by the customer service personnel in a conversation area of the customer service interface, and display the conversation information that needs to be handled at present in a question and answer area. The conversation information includes at least one target question information consulted by a target user. On this basis, the customer service personnel can issue various voice interaction instructions to the AI voice assistant. The AI voice assistant responds to various voice interaction instructions issued by the customer service personnel and controls the content change on the customer service interface based on the execution result information of the voice interaction instructions.
[0098] For example, the customer service personnel can issue a "next conversation" conversation switching instruction in a voice manner. The AI voice assistant can switch the current conversation to the next conversation in response to the instruction to display at least one target question information consulted by a target user involved in the next conversation. Of course, the customer service personnel can issue a "previous conversation" conversation switching instruction in a voice manner. The AI voice assistant can switch the current conversation to the previous conversation in response to the instruction to display at least one target question information consulted by a target user involved in the previous conversation.
[0099] For example, the customer service personnel can issue a "calm down the user" soothing speech generation instruction in a voice manner. The AI voice assistant can generate the following text information "Please select one of the following replies: 1. We are deeply sorry for the misunderstanding this time. 2. I am sorry for the inconvenience, please forgive me. 3. We will try our best to correct the mistake and apologize again" in response to the instruction, and prompt the customer service personnel in a voice manner "Please select one of the following replies". The customer service personnel can select among the soothing speeches in a voice manner. The AI voice assistant can respond to the selection operation of the customer service personnel to display the selected soothing speech in the question and answer area of the customer service interface, that is, send the selected soothing speech to the target user.
[0100] For example, the customer service personnel can issue a "hide" hiding instruction in a voice manner. The AI voice assistant can hide the interaction interface in response to the instruction and display the icon of the AI voice assistant on the customer service interface.
[0101] For example, the customer service personnel can issue a "check logistics" logistics information query instruction in a voice manner, and the AI voice assistant can respond to the instruction and display the queried logistics information in the function area for the customer service personnel to refer to. In response to a voice input operation initiated by the customer service personnel in the input box of the question and answer area according to the queried logistics information, the text information corresponding to the voice information input by the customer service personnel is displayed in the input box: This side verifies that your express logistics is abnormal, and will help you refund immediately, please wait. In response to the sending instruction issued by the customer service personnel, the text information in the input box is sent to the question and answer area as target reply information.
[0102] For example, the customer service personnel can issue a "help refund" refund instruction in a voice manner, and the AI voice assistant can respond to the instruction and display a refund card and a corresponding sending control in the function area of the customer service interface. The customer service personnel can issue a sending instruction in a voice manner or by triggering the sending control through head movement, and the AI voice assistant can respond to the sending instruction of the customer service personnel and directly send the refund card from the function area to the question and answer area.
[0103] Optionally, the AI voice assistant can provide an AI expansion service, which can expand the text information corresponding to the voice information input by the customer service personnel. For example, the text information corresponding to the voice information input by the user is: Don't be anxious, and the target text information obtained after expansion can be: Dear, don't be anxious, I will help you check the previous communication record first.
[0104] Optionally, the AI voice assistant can detect specific emotional words in the target problem information consulted by the target user. In the case that any target problem information contains specific emotional words, the AI voice assistant can generate emotional soothing information for the customer service personnel based on a target language model, such as "Your heart is full of endless energy. Face difficulties bravely, and you will find that your strength far exceeds your imagination", and display the emotional soothing information on the interactive interface provided by the AI voice assistant to soothe the emotions of the customer service personnel.
[0105] Optionally, when the customer service personnel suddenly feels unwell and needs to leave the post, an emergency hosting instruction can be issued in a head movement manner or a voice manner to continue the corresponding work task by using the AI voice assistant.
[0106] In this way, the AI voice assistant can more efficiently assist the customer service personnel to complete the interaction with the target user. In this process, the customer service personnel can complete the interaction through simple voice input with the assistance of the AI voice assistant, which can solve the problem of being unable to flexibly use peripherals to interact due to the inconvenience of manual input, and thus solve the problem of low interaction efficiency. The scheme is applicable to various interaction scenarios, especially interaction scenarios involving people with disabilities.
[0107] In addition to the information interaction method in the foregoing embodiments, the application further provides an information interaction method applied to an AI voice assistant on a terminal device, the AI voice assistant being associated with a target program on the terminal device. As shown in Figure 10 The method can include the steps as shown in Figure 10
[0108] Step 1001, display an interaction interface provided by the target program, the interaction interface including an interaction area, the interaction area being used to display at least one interaction information initiated by a first user to a second user.
[0109] Step 1002, based on a voice recognition module in the AI voice assistant, receive at least one voice interaction instruction initiated by the second user to any interaction information.
[0110] Step 1003, for any voice interaction instruction, execute the voice interaction instruction and output the execution result information of the voice interaction instruction in a display and / or voice manner.
[0111] Step 1004, for a target voice interaction instruction in the at least one voice interaction instruction, select at least part of the execution result information of the target voice interaction instruction as target reply information of any interaction information and display the target reply information in the interaction area.
[0112] Optionally, the target program can be an instant messaging type program, the interaction interface can be a chat window, the interaction area can be an information display area in the chat window, and the chat window can further include an information input area.
[0113] Optionally, the AI voice assistant further includes a head movement recognition module; before receiving the at least one voice interaction instruction initiated by the second user to any target question information, the method further includes: based on the voice recognition module or the head movement recognition module, in response to a wake-up instruction issued by the second user in a voice or head movement manner, wake up the AI voice assistant; and display an interaction interface provided by the AI voice assistant and display text information corresponding to the wake-up instruction on the interaction interface.
[0114] Optionally, for any voice interaction instruction, the any voice interaction instruction is executed, and the execution result information of the any voice interaction instruction is output in a display and / or voice manner, including: according to the semantic information of the any voice interaction instruction and / or the type marking word contained in the any voice interaction instruction, identifying the instruction type of the any voice interaction instruction, the execution mode and the result display mode of different instruction types are different; according to the instruction type of the any voice interaction instruction, executing the any voice interaction instruction according to the execution mode adapted to the instruction type of the any voice interaction instruction to obtain the execution result information of the any voice interaction instruction; and according to the instruction type of the any voice interaction instruction, displaying the execution result information of the any voice interaction instruction in the function area of the interactive interface or on the interactive interface provided by the AI voice assistant.
[0115] Optionally, according to the semantic information of the any voice interaction instruction and / or the type marking word contained in the any voice interaction instruction, the instruction type of the any voice interaction instruction is identified, including: based on the voice recognition module, the text information conversion processing is performed on the any voice interaction instruction to obtain the text instruction corresponding to the any voice interaction instruction; based on the semantic information of the text instruction and / or the type marking word contained in the text instruction, the instruction type of the any voice interaction instruction is identified.
[0116] Optionally, according to the instruction type of the any voice interaction instruction, the execution mode adapted to the instruction type of the any voice interaction instruction is executed to obtain the execution result information of the any voice interaction instruction, including: in the case that the any voice interaction instruction is a first type instruction, the text instruction is input into a target language model in the AI voice assistant, so that the target voice model executes a transaction corresponding to the any voice interaction instruction in a data source associated with the interactive interface and outputs a transaction execution result as the execution result information of the any voice interaction instruction; wherein the first type instruction is a voice instruction based on which the AI voice assistant executes a corresponding transaction in a data source associated with the interactive interface and displays a transaction execution result in a function area of the interactive interface; in the case that the any voice interaction instruction is a second type instruction, the text instruction is input into the target language model in the AI voice assistant for dialogue requirement identification and dialogue information generation processing, and at least one piece of dialogue information corresponding to the any voice interaction instruction is output as the execution result information of the any voice interaction instruction; wherein the second type instruction is a voice instruction based on which the AI voice assistant generates corresponding dialogue information and displays the dialogue information on the interactive interface provided by the AI voice assistant; in the case that the any voice interaction instruction is a third type instruction, the text instruction is matched in a standard instruction library supported by the AI voice assistant, if a target standard instruction is matched, the target standard instruction is executed to obtain the execution result information of the any voice interaction instruction; wherein the third type instruction is a standard instruction directly executed by the AI voice assistant.
[0117] Optionally, in the case that any voice interaction instruction is a third type of instruction, the method further comprises: if the target standard type instruction is not matched, inputting the text instruction into an AI voice assistant target language model for instruction conversion processing, if the text instruction is converted into a standard type instruction, outputting the converted standard type instruction; if the text instruction cannot be converted into a standard type instruction, outputting a prompt information by voice and / or displaying the prompt information on the interaction interface to remind the second user to re-input any voice interaction instruction.
[0118] Optionally, according to the instruction type of any voice interaction instruction, the execution result information of any voice interaction instruction is displayed in the function area of the interaction interface or on the interaction interface provided by the AI voice assistant, comprising: in the case that any voice interaction instruction is a first type of instruction, the execution result information of any voice interaction instruction is displayed in the function area of the interaction interface, and the function area is different from the question and answer area; in the case that any voice interaction instruction is other type of instruction, the text information corresponding to any voice interaction instruction and the execution result information thereof are displayed on the interaction interface, and the interaction interface is located above the interaction interface.
[0119] Optionally, it further comprises: before receiving the first type of instruction, if the interaction interface is in a display state, responding to a hiding instruction issued by the second user by voice or head movement, hiding the interaction interface, and displaying an icon of the AI voice assistant on the interaction interface; or before receiving other type of instruction, if the interaction interface is in a hidden state, responding to a wake-up instruction issued by the second user by voice or head movement, displaying the interaction interface, and displaying the text information corresponding to the wake-up instruction on the interaction interface.
[0120] Optionally, the target voice interaction instruction includes at least one voice interaction instruction belonging to the first type of instruction and having execution result information containing the sending control, and / or at least one voice interaction instruction belonging to the second type of instruction; for the target voice interaction instruction in the at least one voice interaction instruction, at least part of the execution result information of the target voice interaction instruction is selected as the target reply information of any target question information and displayed in the Q&A area, including: in the case that the target voice interaction instruction belongs to the first type of instruction and its execution result information contains the sending control, the execution result information of the target voice interaction instruction is displayed in the function area of the interaction interface, and in response to a sending instruction issued by the second user in the form of voice or the sending control, part of the execution result information of the target voice interaction instruction associated with the sending control is directly sent from the function area to the Q&A area as the target reply information; in the case that the target voice interaction instruction belongs to the second type of instruction, at least one piece of dialogue information corresponding to the target voice interaction instruction is displayed on the interaction interface provided by the AI voice assistant, and in response to a selection instruction issued by the second user in the form of voice, the target dialogue information is selected therefrom and displayed in the input box corresponding to the Q&A area; and in response to a sending instruction issued by the second user, the target dialogue information in the input box is sent to the Q&A area.
[0121] Optionally, it further includes: for a non-target voice interaction instruction belonging to the first type of instruction in the at least one voice interaction instruction, the execution result information of the non-target voice interaction instruction is displayed in the function area of the interaction interface, in response to a voice input operation initiated by the second user in the input box of the Q&A area according to the execution result information of the non-target voice interaction instruction, text information corresponding to the voice information input by the second user is displayed in the input box, and in response to a sending instruction issued by the second user, the text information in the input box is sent to the Q&A area as the target reply information of any target question information.
[0122] Optionally, in the case that the target standard type of instruction is a work instruction, at least one user conversation information to which the second user is responsible is displayed in the conversation area of the interaction interface; in the case that the target standard type of instruction is a conversation switching instruction, the Q&A area of the interaction interface is switched to perform user conversation information, and the user conversation information at least includes target question information of a user to which the switching is performed.
[0123] Optionally, the method further comprises at least one of the following operations: in the case that any of the target question information contains a specific emotional word, generating emotional pacification information for the second user based on the target language model, and displaying the emotional pacification information on the interactive interface provided by the AI voice assistant; in response to an emergency hosting instruction issued by the second user in a voice manner, switching to an emergency working mode; in the emergency working mode, identifying target question information in the question and answer area that has not been replied, generating target reply information of the target question information that has not been replied based on the target language model, and displaying the target reply information in the question and answer area.
[0124] Optionally, the AI voice assistant is embedded in the target program, or the AI voice assistant and the target program are independent of each other, and the AI voice assistant integrates an access portal of the target program, and the AI voice assistant displays the interactive interface provided by the target program through the access portal.
[0125] The detailed implementation and beneficial effects of each step in the method of the embodiment have been described in detail in the foregoing embodiments, and will not be described in detail here.
[0126] In the embodiment, an AI voice assistant associated with a customer service program is provided for a customer service personnel. On the one hand, the AI voice assistant is responsible for displaying a customer service interface provided by the customer service program; on the other hand, the AI voice assistant can respond to and execute a voice interaction instruction of the customer service personnel to give an execution result information, and for a voice instruction of a user that needs to be replied, the AI voice assistant can select content that can be used as reply information from the execution result information and display the content in a question and answer area of the customer service interface, thereby assisting the customer service personnel to complete interaction with the user. In this process, the customer service personnel can complete the interaction through simple voice input with the assistance of the AI voice assistant, which can solve the problem that the interaction cannot be flexibly performed using peripherals due to inconvenience of manual input, and further solve the problem of low interaction efficiency caused thereby. The scheme is suitable for various interaction scenarios, especially interaction scenarios involving disabled persons.
[0127] It should be noted that the execution subject of each step of the method provided in the above embodiments can be the same device, or the method can also be executed by different devices as the execution subject. For example, the execution subject of steps 11 to 14 can be device A; for another example, the execution subject of steps 11 and 12 can be device A, and the execution subject of steps 13 and 14 can be device B; and so on.
[0128] In addition, in some of the processes described in the above embodiments and accompanying drawings, a plurality of operations are included in a specific order, but it should be clear that these operations can be executed in the order in which they appear in this document or in parallel, and the serial numbers of the operations such as 11, 12, etc. are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and the operations can be executed in sequence or in parallel. It should be noted that the "first", "second", etc. described herein are used to distinguish different messages, devices, modules, etc. and do not represent the order of precedence. Also, "first" and "second" are not of different types.
[0129] Figure 11 A structural schematic diagram of an information interaction device is provided for another exemplary embodiment of the present application. As shown in the figure, the device includes a display module 1101 for displaying a customer service interface provided by the customer service program, the customer service interface including a question and answer area for displaying at least one target question information for a target user to consult a customer service personnel; a receiving module 1102 for receiving at least one voice interaction instruction initiated by the customer service personnel for any target question information based on a voice recognition module in the AI voice assistant; an execution module 1103 for executing any voice interaction instruction and outputting the execution result information of the any voice interaction instruction in a display and / or voice manner; a selection module 1104 for selecting at least part of the execution result information of a target voice interaction instruction in the at least one voice interaction instruction as target reply information of the any target question information and displaying it in the question and answer area. Figure 11
[0130] Optionally, the AI voice assistant further includes a head movement recognition module; the receiving module 1102, before receiving at least one voice interaction instruction initiated by the customer service personnel for any target question information, is further configured to: based on the voice recognition module or the head movement recognition module, in response to a wake-up instruction issued by the customer service personnel in a voice or head movement manner, wake up the AI voice assistant; and display an interaction interface provided by the AI voice assistant and display text information corresponding to the wake-up instruction on the interaction interface.
[0131] Optionally, the execution module 1103 is specifically configured to, for any voice interaction instruction, execute the any voice interaction instruction, and output the execution result information of the any voice interaction instruction in a display and / or voice manner, by: identifying the instruction type of the any voice interaction instruction according to the semantic information of the any voice interaction instruction and / or the type label word contained in the any voice interaction instruction, the execution manners of different instruction types being different from each other and the result display manners of different instruction types being different from each other; executing the any voice interaction instruction according to the execution manner adapted to the instruction type of the any voice interaction instruction to obtain the execution result information of the any voice interaction instruction according to the instruction type of the any voice interaction instruction; and displaying the execution result information of the any voice interaction instruction in the function area of the customer service interface or on the interactive interface provided by the AI voice assistant according to the instruction type of the any voice interaction instruction.
[0132] Optionally, the execution module 1103 is specifically configured to, for any voice interaction instruction, execute the any voice interaction instruction, and output the execution result information of the any voice interaction instruction in a display and / or voice manner, by: identifying the instruction type of the any voice interaction instruction according to the semantic information of the any voice interaction instruction and / or the type label word contained in the any voice interaction instruction, the execution manners of different instruction types being different from each other and the result display manners of different instruction types being different from each other; executing the any voice interaction instruction according to the execution manner adapted to the instruction type of the any voice interaction instruction to obtain the execution result information of the any voice interaction instruction according to the instruction type of the any voice interaction instruction; and displaying the execution result information of the any voice interaction instruction in the function area of the customer service interface or on the interactive interface provided by the AI voice assistant according to the instruction type of the any voice interaction instruction.
[0133] Optionally, the execution module 1103 is configured to execute the any voice interaction instruction according to the instruction type of the any voice interaction instruction, and obtain execution result information of the any voice interaction instruction in a manner adapted to the instruction type of the any voice interaction instruction, specifically configured to: in a case where the any voice interaction instruction is a first type of instruction, input the text instruction into a target language model in the AI voice assistant, so that the target language model executes a transaction corresponding to the any voice interaction instruction in a data source associated with the customer service interface and outputs a transaction execution result as the execution result information of the any voice interaction instruction; wherein the first type of instruction is a voice instruction in which the AI voice assistant executes a corresponding transaction in a data source associated with the customer service interface based on the target language model and displays a transaction execution result in a functional area of the customer service interface; in a case where the any voice interaction instruction is a second type of instruction, input the text instruction into the target language model in the AI voice assistant for dialogue requirement identification and dialogue information generation processing, and output at least one piece of dialogue information corresponding to the any voice interaction instruction as the execution result information of the any voice interaction instruction; wherein the second type of instruction is a voice instruction in which the AI voice assistant generates corresponding dialogue information based on the target language model and displays the dialogue information on an interactive interface provided by the AI voice assistant; in a case where the any voice interaction instruction is a third type of instruction, match the text instruction in a standard instruction library supported by the AI voice assistant, if a target standard instruction is matched, execute the target standard instruction to obtain the execution result information of the any voice interaction instruction; wherein the third type of instruction is a standard instruction directly executed by the AI voice assistant.
[0134] Optionally, in a case where the any voice interaction instruction is a third type of instruction, the execution module 1103 is further configured to: if the target standard instruction is not matched, input the text instruction into the target language model in the AI voice assistant for instruction conversion processing, if the text instruction is converted into a standard instruction, output the converted standard instruction; if the text instruction cannot be converted into a standard instruction, output prompt information by voice and / or display prompt information on the interactive interface, to remind the customer service personnel to re-input the any voice interaction instruction.
[0135] Optionally, the execution module 1103 is configured to display the execution result information of the any voice interaction instruction in a function area of the customer service interface or on an interaction interface provided by the AI voice assistant according to the instruction type of the any voice interaction instruction, and specifically configured to: in a case where the any voice interaction instruction is a first type of instruction, display the execution result information of the any voice interaction instruction in the function area of the customer service interface, the function area being different from the question and answer area; and in a case where the any voice interaction instruction is another type of instruction, display text information corresponding to the any voice interaction instruction and the execution result information thereof on the interaction interface, the interaction interface being located above the customer service interface.
[0136] Optionally, the execution module 1103 is further configured to: before receiving the first type of instruction, if the interaction interface is in a display state, hide the interaction interface in response to a hiding instruction issued by the customer service personnel through voice or head movement, and display an icon of the AI voice assistant on the customer service interface; or before receiving the other type of instruction, if the interaction interface is in a hidden state, display the interaction interface in response to a wake-up instruction issued by the customer service personnel through voice or head movement, and display text information corresponding to the wake-up instruction on the interaction interface.
[0137] Optionally, the target voice interaction instruction includes a voice instruction that is a first type of instruction in the at least one voice interaction instruction and whose execution result information contains a sending control, and / or a voice instruction that is a second type of instruction in the at least one voice interaction instruction; and the selection module 1104 is configured to, when selecting at least part of information in the execution result information of the target voice interaction instruction as target reply information of the target question information and displaying the target reply information in the question and answer area, specifically configured to: in a case where the target voice interaction instruction is a first type of instruction and its execution result information contains a sending control, display the execution result information of the target voice interaction instruction in a function area of the customer service interface, and in response to a sending instruction issued by the customer service personnel through voice or the sending control, directly send part of information associated with the sending control in the execution result information of the target voice interaction instruction to the question and answer area as the target reply information; in a case where the target voice interaction instruction is a second type of instruction, display at least one piece of text information corresponding to the target voice interaction instruction on an interaction interface provided by the AI voice assistant, in response to a selection instruction issued by the customer service personnel in a voice manner, select target text information therefrom, and display the target text information in an input box corresponding to the question and answer area; and in response to a sending instruction issued by the customer service personnel, send the target text information in the input box to the question and answer area.
[0138] Optionally, the selection module 1104 is further configured to: for a non-target voice interaction instruction in the at least one voice interaction instruction belonging to the first type of instruction, display execution result information of the non-target voice interaction instruction in a function area of the customer service interface; in response to a voice input operation initiated by the customer service personnel in an input box of the question and answer area according to the execution result information of the non-target voice interaction instruction, display text information corresponding to voice information input by the customer service personnel in the input box; and in response to a sending instruction issued by the customer service personnel, send the text information in the input box to the question and answer area as target reply information of the any target question information.
[0139] Optionally, when the execution module 1103 executes the target standard type instruction to obtain the execution result information of the any voice interaction instruction, the execution module 1103 is specifically configured to: in a case where the target standard type instruction is a work instruction, display at least one user session information that needs to be taken care of by the customer service personnel in a conversation area of the customer service interface; and in a case where the target standard type instruction is a session switching instruction, switch user session information in a question and answer area of the customer service interface, the user session information at least including target question information of a user currently switched to.
[0140] Optionally, the selection module 1104 is further configured to: in a case where the any target question information contains a specific emotional word, generate emotional soothing information for the customer service personnel based on a target language model, and display the emotional soothing information on an interaction interface provided by the AI voice assistant; in response to an emergency hosting instruction issued by the customer service personnel in a voice manner, switch to an emergency working mode; in the emergency working mode, identify target question information that has not been replied to in the question and answer area, generate target reply information of the target question information that has not been replied to based on a target language model, and display the target reply information in the question and answer area.
[0141] Optionally, the AI voice assistant is embedded in the customer service program to be implemented, or the AI voice assistant and the customer service program are independent of each other, and the AI voice assistant integrates an access portal of the customer service program, and the AI voice assistant displays a customer service interface provided by the customer service program through the access portal.
[0142] In this embodiment, an AI voice assistant associated with a customer service program is provided for a customer service personnel, which is responsible for displaying a customer service interface provided by the customer service program on one hand, and can respond to and execute voice interaction instructions of the customer service personnel to give execution result information, and for voice instructions that need to be replied to the user, can select content that can be used as reply information from the execution result information to display in the question and answer area of the customer service interface, assisting the customer service personnel to complete the interaction with the user, in this process, the customer service personnel can complete the interaction through simple voice input with the assistance of the AI voice assistant, which can solve the problem that it is not flexible to use peripherals for interaction due to the inconvenience of manual input, and further solve the problem of low interaction efficiency caused thereby. The scheme is suitable for various interaction scenarios, especially the interaction scenarios involving disabled personnel.
[0143] The internal functions and structures of the information interaction device are described above, as shown in Figure 12 In practice, the information interaction device can be implemented as an electronic device, including a memory 1201, a processor 1202, and a communication component 1203.
[0144] The memory 1201 is used to store computer programs and can be configured to store other various data to support operations on the computing platform. Examples of these data include instructions for any application or method operating on the computing platform, contact data, phonebook data, messages, pictures, videos, etc.
[0145] The memory 1201 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0146] The processor 1202 is coupled to the memory 1201 and is used to execute computer programs in the memory 1201, for displaying a customer service interface provided by the customer service program, the customer service interface including a question and answer area for displaying at least one target question information of a target user for consulting a customer service personnel; based on a voice recognition module in the AI voice assistant, receiving at least one voice interaction instruction initiated by the customer service personnel for any target question information; for any voice interaction instruction, executing the voice interaction instruction and outputting the execution result information of the voice interaction instruction in a display and / or voice manner; for a target voice interaction instruction in the at least one voice interaction instruction, selecting at least part of the execution result information of the target voice interaction instruction as target reply information of the any target question information and displaying in the question and answer area.
[0147] Optionally, the AI voice assistant further comprises a head movement recognition module; before receiving the at least one voice interaction instruction initiated by the customer service personnel for any target question information, the receiving module 1102 is further used for: based on the voice recognition module or the head movement recognition module, in response to the wake-up instruction issued by the customer service personnel through voice or head movement, waking up the AI voice assistant; and displaying the interaction interface provided by the AI voice assistant, and displaying the text information corresponding to the wake-up instruction on the interaction interface.
[0148] Optionally, when the processor 1202 executes any voice interaction instruction and outputs the execution result information of the any voice interaction instruction in a display and / or voice manner, it is specifically used for: according to the semantic information of the any voice interaction instruction and / or the type marking word contained in the any voice interaction instruction, identifying the instruction type of the any voice interaction instruction, the execution mode and the result display mode of different instruction types are different; according to the instruction type of the any voice interaction instruction, executing the any voice interaction instruction according to the execution mode adapted to the instruction type of the any voice interaction instruction to obtain the execution result information of the any voice interaction instruction; and according to the instruction type of the any voice interaction instruction, displaying the execution result information of the any voice interaction instruction in the function area of the customer service interface or on the interaction interface provided by the AI voice assistant.
[0149] Optionally, when the processor 1202 identifies the instruction type of the any voice interaction instruction according to the semantic information of the any voice interaction instruction and / or the type marking word contained in the any voice interaction instruction, it is specifically used for: based on the text information conversion processing of the any voice interaction instruction by the voice recognition module, to obtain the text instruction corresponding to the any voice interaction instruction; based on the semantic information of the text instruction and / or the type marking word contained in the text instruction, identifying the instruction type of the any voice interaction instruction.
[0150] Optionally, the processor 1202 is configured to execute the any voice interaction instruction according to the instruction type of the any voice interaction instruction in an execution mode adapted to the instruction type of the any voice interaction instruction to obtain execution result information of the any voice interaction instruction, specifically configured to: in a case where the any voice interaction instruction is a first type of instruction, input the text instruction into a target language model in the AI voice assistant, so that the target language model executes a transaction corresponding to the any voice interaction instruction in a data source associated with the customer service interface and outputs a transaction execution result as the execution result information of the any voice interaction instruction; wherein the first type of instruction is a voice instruction in which the AI voice assistant executes a corresponding transaction in a data source associated with the customer service interface based on the target language model and displays a transaction execution result in a function area of the customer service interface; in a case where the any voice interaction instruction is a second type of instruction, input the text instruction into the target language model in the AI voice assistant for dialogue requirement identification and dialogue information generation processing and output at least one piece of dialogue information corresponding to the any voice interaction instruction as the execution result information of the any voice interaction instruction; wherein the second type of instruction is a voice instruction in which the AI voice assistant generates corresponding dialogue information based on the target language model and displays the dialogue information on an interaction interface provided by the AI voice assistant; in a case where the any voice interaction instruction is a third type of instruction, match the text instruction in a standard instruction library supported by the AI voice assistant, if a target standard instruction is matched, execute the target standard instruction to obtain the execution result information of the any voice interaction instruction; wherein the third type of instruction is a standard instruction directly executed by the AI voice assistant.
[0151] Optionally, in a case where the any voice interaction instruction is a third type of instruction, the processor 1202 is further configured to: if the target standard instruction is not matched, input the text instruction into the target language model in the AI voice assistant for instruction conversion processing, if the text instruction is converted into a standard instruction, output the converted standard instruction; if the text instruction cannot be converted into a standard instruction, output prompt information by voice and / or display prompt information on the interaction interface to remind the customer service personnel to re-input the any voice interaction instruction.
[0152] Optionally, the processor 1202 is configured to display the execution result information of the any voice interaction instruction in a function area of the customer service interface or on an interaction interface provided by the AI voice assistant according to the instruction type of the any voice interaction instruction, specifically: in the case that the any voice interaction instruction is a first type of instruction, displaying the execution result information of the any voice interaction instruction in the function area of the customer service interface, the function area being different from the question and answer area; in the case that the any voice interaction instruction is another type of instruction, displaying text information corresponding to the any voice interaction instruction and the execution result information thereof on the interaction interface, the interaction interface being located above the customer service interface.
[0153] Optionally, the processor 1202 is further configured to: before receiving the first type of instruction, if the interaction interface is in a display state, in response to a hiding instruction issued by the customer service personnel through voice or head movement, hiding the interaction interface and displaying an icon of the AI voice assistant on the customer service interface; or before receiving the other type of instruction, if the interaction interface is in a hidden state, in response to a wake-up instruction issued by the customer service personnel through voice or head movement, displaying the interaction interface and displaying text information corresponding to the wake-up instruction on the interaction interface.
[0154] Optionally, the target voice interaction instruction includes a voice instruction belonging to the first type of instruction and having execution result information containing a sending control among the at least one voice interaction instruction, and / or a voice instruction belonging to the second type of instruction; and the processor 1202 is configured to, for the target voice interaction instruction among the at least one voice interaction instruction, select at least part of information in the execution result information of the target voice interaction instruction as target reply information of the target question information and display the target reply information in the question and answer area, specifically: in the case that the target voice interaction instruction belongs to the first type of instruction and its execution result information contains the sending control, the function area of the customer service interface displays the execution result information of the target voice interaction instruction, in response to a sending instruction issued by the customer service personnel through voice or the sending control, sending part of information associated with the sending control in the execution result information of the target voice interaction instruction as the target reply information from the function area to the question and answer area; in the case that the target voice interaction instruction belongs to the second type of instruction, the interaction interface provided by the AI voice assistant displays at least one piece of script information corresponding to the target voice interaction instruction, in response to a selection instruction issued by the customer service personnel in the form of voice, selecting target script information therefrom and displaying the target script information in an input box corresponding to the question and answer area; and in response to a sending instruction issued by the customer service personnel, sending the target script information in the input box to the question and answer area.
[0155] Optionally, the processor 1202 is further configured to: for a non-target voice interaction instruction belonging to the first type of instruction in the at least one voice interaction instruction, display execution result information of the non-target voice interaction instruction in a function area of the customer service interface, in response to a voice input operation initiated by the customer service personnel in an input box of the question and answer area according to the execution result information of the non-target voice interaction instruction, display text information corresponding to voice information input by the customer service personnel in the input box as target reply information of any target question information, and in response to a sending instruction issued by the customer service personnel, send the text information in the input box to the question and answer area as the target reply information of the any target question information.
[0156] Optionally, when the processor 1202 executes the target standard type instruction to obtain the execution result information of the any voice interaction instruction, the processor 1202 is specifically configured to: in a case where the target standard type instruction is a work instruction, display at least one user session information that needs to be taken care of by the customer service personnel in a conversation area of the customer service interface; and in a case where the target standard type instruction is a session switching instruction, switch user session information in a question and answer area of the customer service interface, the user session information at least including target question information of a user currently switched to.
[0157] Optionally, the processor 1202 is further configured to: in a case where the any target question information contains a specific emotional word, generate emotional soothing information for the customer service personnel based on a target language model, and display the emotional soothing information on an interaction interface provided by the AI voice assistant; in response to an emergency hosting instruction issued by the customer service personnel in a voice manner, switch to an emergency working mode; and in the emergency working mode, identify target question information that has not been replied to in the question and answer area, generate target reply information of the target question information that has not been replied to based on a target language model, and display the target reply information in the question and answer area.
[0158] Optionally, the AI voice assistant is embedded in the customer service program to be implemented, or the AI voice assistant and the customer service program are independent of each other, and the AI voice assistant integrates an access portal of the customer service program, and the AI voice assistant displays a customer service interface provided by the customer service program through the access portal.
[0159] Further, as shown in Figure 12 The electronic device further includes a display 1204, a power component 1205, an audio component 1206, and other components. Figure 12 Some components are only schematically shown in the electronic device, and it does not mean that the electronic device only includes Figure 12 the components shown. In addition, Figure 12The components in the dashed box are optional components, not mandatory components, and can be determined according to the product form of the working node. The working node in the embodiment can be implemented as a terminal device such as a desktop computer, a notebook computer, a smart phone or an IOT device, or a server device such as a general server, a cloud server or a server array. If the working node in the embodiment is implemented as a terminal device such as a desktop computer, a notebook computer or a smart phone, it can include Figure 12 components in the dashed box; if the working node in the embodiment is implemented as a server device such as a general server, a cloud server or a server array, it can not include Figure 12 components in the dashed box.
[0160] Correspondingly, the embodiment of the application further provides a computer readable storage medium storing a computer program, which can implement each step that can be executed by the electronic device when the computer program is executed.
[0161] In the embodiment, an AI voice assistant associated with a customer service program is provided for a customer service personnel, which is responsible for displaying a customer service interface provided by the customer service program on one hand, and can respond to and execute a voice interaction instruction of the customer service personnel to give an execution result information, and can select content that can be used as reply information from the execution result information and display the content in a question and answer area of the customer service interface for a voice instruction of the user that needs to be replied, so as to assist the customer service personnel to complete the interaction with the user. In this process, the customer service personnel can complete the interaction through simple voice input with the assistance of the AI voice assistant, which can solve the problem that the interaction cannot be flexibly performed by using peripherals due to inconvenience of manual input, and further solve the problem of low interaction efficiency caused thereby. The scheme is applicable to various interaction scenarios, especially interaction scenarios involving disabled personnel.
[0162] The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0163] The communication component is configured to facilitate wired or wireless communication between the device on which the communication component is installed and other devices. The device on which the communication component is installed can access a wireless network based on a communication standard, such as WiFi, a 2G, 3G, 4G / LTE, 5G, or the like cellular communication network, or a combination thereof. In an example embodiment, the communication component receives a broadcast signal or broadcast related information from an external broadcast management system via a broadcast channel. In an example embodiment, the communication component further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Blue Tooth (BT) technology, and other technologies.
[0164] The display includes a screen, which can include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide, and a gesture on the touch panel. The touch sensor can not only sense a boundary of a touch or a slide action, but also detect a duration and a pressure associated with a touch or a slide operation.
[0165] The power supply component provides power to various components of the device on which the power supply component is installed. The power supply component can include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the device on which the power supply component is installed.
[0166] The audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) that is configured to receive an external audio signal when the device on which the audio component is installed is in a particular mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in memory or transmitted via the communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0167] Those skilled in the art will appreciate that embodiments of the application can be supplied as a method, a system, or a computer program product. Accordingly, the application can be embodied in the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the application can be in the form of a computer program product on one or more computer readable storage media (including, but not limited to, disk memory, Compact Disc Read-Only Memory (CD-ROM), optical memory, and the like) embodying computer readable program code.
[0168] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks or in the flowchart one or more blocks and / or in the block or blocks of the block diagram.
[0169] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart block or blocks or in the flowchart one or more blocks and / or in the block or blocks of the block diagram.
[0170] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks or in the flowchart one or more blocks and / or in the block or blocks of the block diagram.
[0171] In one typical arrangement, the computing device includes one or more processors (Central Processing Units, CPUs), input / output interfaces, network interfaces, and memory.
[0172] The memory can include non-persistent memory and / or volatile memory, such as a random access memory (RAM) including a cache area for the temporary storage of data. The memory can also include non-volatile memory, such as read only memory (ROM), electrically programmable read only memory (EPROM), or electrically erasable programmable memory (EEPROM), for storing structural information and instructions delivered from a program file that can include software that can be executed by the memory. The memory can also include a compact disc read only memory (CDROM) or other optical memory, for storing structural information and instructions delivered from a program file that can include software that can be executed by the memory.
[0173] Computer-readable media includes permanent and non-permanent, movable and non-movable media, which can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital video disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.
[0174] It should also be noted that the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, so that processes, methods, articles or devices including a series of elements not only include those elements, but also include other elements not explicitly listed, or inherent to such processes, methods, articles or devices. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or device including the element.
[0175] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application shall be included in the scope of claims of the present application.
Claims
1. An information interaction method, characterized in that, The application is applied to an AI voice assistant on a terminal device, the AI voice assistant is associated with a customer service program on the terminal device, and the method comprises the following steps: Displaying a customer service interface provided by the customer service program, wherein the customer service interface comprises a question and answer area for displaying at least one target question information for a target user to consult a customer service personnel; Based on a voice recognition module in the AI voice assistant, receiving at least one voice interaction instruction initiated by the customer service personnel for any target question information; For any voice interaction instruction, executing the voice interaction instruction and outputting the execution result information of the voice interaction instruction in a display and / or voice manner; For a target voice interaction instruction in the at least one voice interaction instruction, selecting at least part of the execution result information of the target voice interaction instruction as target reply information of the any target question information and displaying the target reply information in the question and answer area.
2. The method of claim 1, wherein, The AI voice assistant further comprises a head movement recognition module; before receiving at least one voice interaction instruction initiated by the customer service personnel for any target question information, the method further comprises the following steps: Based on the voice recognition module or the head movement recognition module, responding to a wake-up instruction issued by the customer service personnel in a voice or head movement manner to wake up the AI voice assistant; and Displaying an interaction interface provided by the AI voice assistant and displaying text information corresponding to the wake-up instruction on the interaction interface.
3. The method of claim 1, wherein, For any voice interaction instruction, executing the voice interaction instruction and outputting the execution result information of the voice interaction instruction in a display and / or voice manner, comprising the following steps: According to the semantic information of the any voice interaction instruction and / or the type marking word contained in the any voice interaction instruction, identifying the instruction type of the any voice interaction instruction, the execution mode and the result display mode of different instruction types are different; According to the instruction type of the any voice interaction instruction, executing the any voice interaction instruction in an execution mode adapted to the instruction type of the any voice interaction instruction to obtain the execution result information of the any voice interaction instruction; and According to the instruction type of the any voice interaction instruction, displaying the execution result information of the any voice interaction instruction in a function area of the customer service interface or on an interaction interface provided by the AI voice assistant.
4. The method of claim 3, wherein, According to the semantic information of the any voice interaction instruction and / or the type marking word contained in the any voice interaction instruction, identifying the instruction type of the any voice interaction instruction, comprising the following steps: Based on the voice recognition module, performing text information conversion processing on the any voice interaction instruction to obtain a text instruction corresponding to the any voice interaction instruction; Based on the semantic information of the text instruction and / or the type marking word contained in the text instruction, identifying the instruction type of the any voice interaction instruction.
5. The method of claim 4, wherein, According to the instruction type of the any voice interaction instruction, executing the any voice interaction instruction in an execution mode adapted to the instruction type of the any voice interaction instruction to obtain the execution result information of the any voice interaction instruction, comprising the following steps: In the case that the any voice interaction instruction is a first type of instruction, input the text instruction into a target language model in the AI voice assistant, so that the target language model executes a transaction corresponding to the any voice interaction instruction in a data source associated with the customer service interface and outputs a transaction execution result as execution result information of the any voice interaction instruction; wherein the first type of instruction is a voice instruction that the AI voice assistant executes a corresponding transaction in the data source associated with the customer service interface based on the target language model and displays a transaction execution result in a function area of the customer service interface; In the case that the any voice interaction instruction is a second type of instruction, input the text instruction into a target language model in the AI voice assistant for dialogue requirement identification and dialogue information generation processing and output at least one piece of dialogue information corresponding to the any voice interaction instruction as execution result information of the any voice interaction instruction; wherein the second type of instruction is a voice instruction that the AI voice assistant generates a corresponding dialogue information based on the target language model and displays the dialogue information on an interactive interface provided by the AI voice assistant; In the case that the any voice interaction instruction is a third type of instruction, match the text instruction in a standard instruction library supported by the AI voice assistant, if a target standard instruction is matched, execute the target standard instruction to obtain execution result information of the any voice interaction instruction; wherein the third type of instruction is a standard instruction directly executed by the AI voice assistant.
6. The method of claim 5, wherein, In the case that the any voice interaction instruction is a third type of instruction, the method further comprises: If the target standard instruction is not matched, input the text instruction into a target language model in the AI voice assistant for instruction conversion processing, if the text instruction is converted into a standard instruction, output the converted standard instruction; if the text instruction cannot be converted into a standard instruction, output a prompt information by voice and / or display a prompt information on the interactive interface, so as to remind the customer service personnel to re-input the any voice interaction instruction.
7. The method of claim 5, wherein, According to the instruction type of the any voice interaction instruction, display execution result information of the any voice interaction instruction in a function area of the customer service interface or on an interactive interface provided by the AI voice assistant, comprising: In the case that the any voice interaction instruction is a first type of instruction, display execution result information of the any voice interaction instruction in a function area of the customer service interface, and the function area is different from the question and answer area; In the case that the any voice interaction instruction is a second type of instruction, display text information corresponding to the any voice interaction instruction and its execution result information on the interactive interface, and the interactive interface is located above the customer service interface.
8. The method of claim 7, wherein, Further comprising: Before receiving the first type of instruction, if the interactive interface is in a display state, hide the interactive interface in response to a hiding instruction issued by the customer service personnel by voice or head movement, and display an icon of the AI voice assistant on the customer service interface; Or Before receiving other types of instructions, if the interactive interface is in a hidden state, a wake-up instruction issued by the customer service personnel through voice or head movement is responded to, the interactive interface is displayed, and text information corresponding to the wake-up instruction is displayed on the interactive interface.
9. The method of claim 5, wherein, The target voice interaction instruction includes a voice instruction in the at least one voice interaction instruction that belongs to the first type of instruction and has execution result information containing a sending control, and / or a voice instruction in the at least one voice interaction instruction that belongs to the second type of instruction. For the target voice interaction instruction in the at least one voice interaction instruction, at least part of the information in the execution result information of the target voice interaction instruction is selected as target reply information of the any target question information and displayed in the question and answer area, including: In the case where the target voice interaction instruction belongs to the first type of instruction and its execution result information contains a sending control, the execution result information of the target voice interaction instruction is displayed in the function area of the customer service interface, and in response to a sending instruction issued by the customer service personnel in the form of voice or the sending control, part of the information in the execution result information of the target voice interaction instruction associated with the sending control is directly sent from the function area to the question and answer area as the target reply information. In the case where the target voice interaction instruction belongs to the second type of instruction, at least one piece of script information corresponding to the target voice interaction instruction is displayed on the interactive interface provided by the AI voice assistant, target script information is selected in response to a selection instruction issued by the customer service personnel in the form of voice, and the target script information is displayed in the input box corresponding to the question and answer area; and in response to a sending instruction issued by the customer service personnel, the target script information in the input box is sent to the question and answer area.
10. The method of claim 9, wherein, Also includes: For the non-target voice interaction instruction that belongs to the first type of instruction in the at least one voice interaction instruction, the execution result information of the non-target voice interaction instruction is displayed in the function area of the customer service interface, in response to a voice input operation initiated by the customer service personnel in the input box of the question and answer area according to the execution result information of the non-target voice interaction instruction, text information corresponding to the voice information input by the customer service personnel is displayed in the input box, and in response to a sending instruction issued by the customer service personnel, the text information in the input box is sent to the question and answer area as target reply information of the any target question information. The non-target voice interaction instruction is a voice interaction instruction that does not need to reply to the target user, and the target voice interaction instruction is a voice interaction instruction that needs to reply to the target user.
11. The method according to any one of claims 1 to 10, characterized in that, Also includes at least one of the following operations: In the case where the any target question information contains a specific emotional word, emotion soothing information for the customer service personnel is generated based on a target language model, and the emotion soothing information is displayed on the interactive interface provided by the AI voice assistant; Switch to an emergency working mode in response to an emergency hosting instruction issued by the customer service personnel in a voice mode; in the emergency working mode, identify target question information in the question and answer area that has not been replied, generate target reply information of the target question information that has not been replied based on a target language model, and display the target reply information in the question and answer area.
12. An information interaction method, characterized in that, An AI voice assistant applied to a terminal device, the AI voice assistant being associated with a target program on the terminal device, the method comprising: displaying an interactive interface provided by the target program, the interactive interface including an interactive area for displaying at least one interactive information initiated by a first user to a second user; receiving at least one voice interactive instruction initiated by the second user for any interactive information based on a voice recognition module in the AI voice assistant; for any voice interactive instruction, executing the any voice interactive instruction and outputting the execution result information of the any voice interactive instruction in a display and / or voice manner; for a target voice interactive instruction in the at least one voice interactive instruction, selecting at least part of the execution result information of the target voice interactive instruction as target reply information of the any interactive information and displaying the target reply information in the interactive area.
13. The method of claim 12, wherein, The target program is an instant messaging program, the interactive interface is a chat window, the interactive area is an information display area in the chat window, and the chat window further includes an information input area.
14. An electronic device, comprising: comprising: a memory and a processor; the memory is configured to store a computer program; the processor is coupled to the memory and is configured to execute the computer program in the memory to implement the steps in the method of any one of claims 1-11 and claims 12-13.
15. A computer readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the processor can implement the steps in the method of any one of claims 1-11 and claims 12-13.
16. A computer program product, characterised in that, comprising computer programs / instructions, when the computer programs / instructions are executed by the processor, the processor can implement the steps in the method of any one of claims 1-11 and claims 12-13.
Citation Information
Patent Citations
Method for providing question consultation service and electronic equipment
CN113724036A
Information interaction method, device and system
CN114500419A