Voice interaction method and device

By configuring multiple agents and corresponding large models and skill knowledge bases in the voice customer service system, and combining the historical data of the shared knowledge base, the method of automatically selecting agents to reply based on user questions is realized, solving the interaction delay problem caused by the complex model in the existing technology, and improving the efficiency and accuracy of voice interaction.

CN120164459APending Publication Date: 2025-06-17BAIRONG ZHIXIN (BEIJING) TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510394885.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-06-17

Smart Images

  • Figure CN120164459A_ABST
    Figure CN120164459A_ABST
Patent Text Reader

Abstract

The invention provides a voice interaction method and device. The method comprises the following steps: receiving voice information input by a user; if it is determined that the intelligent agents need to be switched based on the voice information, determining a target intelligent agent from other intelligent agents except the current intelligent agent in the multiple intelligent agents; processing content corresponding to the voice information by adopting a large model corresponding to the target agent, a skill knowledge base and a shared knowledge base to obtain a reply; and outputting voice corresponding to the reply to the user. As the skill knowledge base queried when the user questions is processed is split into a plurality of sub-knowledge bases with different professional knowledge from a complete knowledge base, the data query amount is all reduced, and the voice interaction efficiency is improved. Moreover, the switched intelligent agent can check the voice information sent by the user to the intelligent agent before switching through the shared knowledge base, user historical voice information transmission between the intelligent agents is not needed, the speed of information sharing between the intelligent agents is increased, and the efficiency and accuracy of information interaction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a voice interaction method, a voice interaction device, a computer device, a computer-readable storage medium, and a computer program product. Background Art

[0002] With the rapid development of artificial intelligence, it has become a trend to introduce artificial intelligence in voice customer service to improve the service efficiency and service quality of voice customer service.

[0003] Currently, when introducing artificial intelligence in voice customer service, a pre-trained model is mainly configured in the voice customer service system. When a user interacts with the voice customer service system, the user inputs voice information to the voice customer service system, and the voice customer service system calls the configured model to process the voice information and generate a reply voice, so as to output the reply voice to the user.

[0004] However, if the voice customer service system wants to improve the accuracy of the reply to the user, it needs to introduce more complex structures or algorithms in this configured model, but this will lead to the model being too complex, the calculation time for the user's voice information being too long, and the delay in outputting the reply voice to the user being relatively obvious, thereby reducing the interaction efficiency of the voice customer service system. Summary of the Invention

[0005] The purpose of the embodiments of this application is to provide a voice interaction method, a voice interaction device, a computer device, a computer-readable storage medium, and a computer program product to improve the interaction efficiency between the voice customer service system and the user.

[0006] To solve the above technical problems, the embodiments of this application provide the following technical solutions:

[0007] In the first aspect of this application, a voice interaction method is provided. The method is applied to a voice customer service system. The voice customer service system includes multiple agents and a shared knowledge base. Each agent corresponds to at least one large model and at least one skill knowledge base. The shared knowledge base stores historical data of the interaction between the user and the voice customer service system. The method includes: receiving the voice information input by the user; if it is determined based on the voice information that an agent needs to be switched, determining a target agent from other agents except the current agent among the multiple agents; processing the content corresponding to the voice information by using the large model and the skill knowledge base corresponding to the target agent and the shared knowledge base to obtain a reply; and outputting the voice corresponding to the reply to the user.

[0008] Compared with the prior art, the voice interaction method provided in the first aspect of the present application configures multiple agents, and configures a corresponding large model and a skill knowledge base for each agent, and then automatically selects the corresponding agent to reply according to different questions of the user. Since the skill knowledge base queried when processing the user's question is split from a complete knowledge base into multiple sub-knowledge bases with different professional knowledge, the data query volume is reduced, so the query of the current knowledge can be completed faster, the reply can be generated faster, and the voice interaction efficiency can be improved. Moreover, the switched agent can view the voice information sent by the user to the agent before switching through the shared knowledge base, without the need for the agents to transmit the user's historical voice information to each other, which improves the speed of information sharing between the agents, enables the switched agent to generate a reply faster and more accurately based on the historical data, and improves the efficiency and accuracy of information interaction.

[0009] In other embodiments provided by the present application, before determining that an agent needs to be switched if the voice information indicates that an agent needs to be switched, or if there is no reply corresponding to the voice information in the skill knowledge base corresponding to the current agent, the method further includes: if the voice information indicates that an agent needs to be switched, or if there is no reply corresponding to the voice information in the skill knowledge base corresponding to the current agent, then determine that an agent needs to be switched; if the voice information indicates that no agent needs to be switched currently, or if there is a reply corresponding to the voice information in the skill knowledge base corresponding to the current agent, then determine that no agent needs to be switched currently.

[0010] According to the content of the appropriate voice information itself and the content in the skill knowledge base, it is possible to determine whether an agent needs to be switched faster and more accurately, improving the efficiency and accuracy of agent switching.

[0011] In other embodiments provided by the present application, before the voice information indicates that an agent needs to be switched, or if there is no reply corresponding to the voice information in the skill knowledge base corresponding to the current agent, the method further includes: if the voice information matches a preset keyword, and the preset keyword is used to indicate another agent, then determine that the voice information indicates that an agent needs to be switched; and / or, if the semantics recognized from the voice information indicates another agent, then determine that the voice information indicates that an agent needs to be switched; and / or, if the intent recognized from the voice information and historical data using the large model indicates another agent, then determine that the voice information indicates that an agent needs to be switched.

[0012] By dynamically using keyword matching, semantic recognition, and intent recognition, the switching opportunity of the agent can be determined more precisely, realizing the precise and efficient switching of the agent.

[0013] In other embodiments provided by the present application, if the voice information matches a preset keyword, and the preset keyword is used to indicate another intelligent agent, determining that the voice information indicates a need to switch the intelligent agent includes: if the voice information matches the preset keyword and the emotional intensity of the voice information is higher than the emotional intensity threshold, determining that the voice information indicates a need to switch the intelligent agent.

[0014] With this emotional intensity threshold in the voice information, it is possible to more accurately determine whether the user needs to switch the intelligent agent, improving the accuracy of intelligent agent switching.

[0015] In other embodiments provided by the present application, each intelligent agent corresponds to a priority weight; determining a target intelligent agent from other intelligent agents except the current intelligent agent among multiple intelligent agents includes: determining multiple candidate intelligent agents from other intelligent agents based on the voice information; selecting the candidate intelligent agent with the highest priority weight from the multiple candidate intelligent agents and determining it as the target intelligent agent.

[0016] With the priority weights of the multiple candidate intelligent agents, it is possible to select the optimal target intelligent agent among the intelligent agents that meet the conditions, thereby improving the accuracy of intelligent agent switching.

[0017] In other embodiments provided by the present application, the skill knowledge base includes a rule knowledge base and a content knowledge base. The rule knowledge base stores various keywords and the content corresponding to various keywords, and the content knowledge base stores various knowledge for reference when the large model generates a reply; processing the content corresponding to the voice information by using the large model corresponding to the target intelligent agent, the skill knowledge base, and the shared knowledge base to obtain a reply includes: matching the content corresponding to the voice information with the keywords in the rule knowledge base and determining the content corresponding to the successfully matched keyword as the reply; and / or, identifying the semantics in the content corresponding to the voice information, matching the semantics with the keywords in the rule knowledge base, and determining the content corresponding to the successfully matched keyword as the reply; and / or, using the large model to identify the intention of the historical data and the content corresponding to the voice information, matching the intention with the keywords in the rule knowledge base, and determining the content corresponding to the successfully matched keyword as the reply; and / or, inputting the content corresponding to the voice information and the historical data into the large model corresponding to the target intelligent agent, and generating a reply based on the content knowledge base by the large model corresponding to the target intelligent agent.

[0018] By matching the content, semantics, and intention of the voice information with the expressions in the rule knowledge base, and generating a reply to the voice information based on the large model and the content knowledge base, it is possible to generate more comprehensive and accurate reply content, improving the accuracy of voice interaction.

[0019] In other embodiments provided by the present application, after receiving the voice information input by the user, the method further includes: storing the voice information in a shared knowledge base, and generating a user profile based on the voice information and historical data, and storing the profile in the shared knowledge base for the agent to call.

[0020] Storing the user's previous voice information in the shared knowledge base and generating a user profile based on the previous voice information can use the profile during the next interaction with the user, enabling comprehensive and rapid acquisition of user information and improving the accuracy and efficiency of interaction with the user.

[0021] In other embodiments provided by the present application, the voice customer service system further includes a language adaptation model; before processing the voice information using the large model, skill knowledge base, and shared knowledge base corresponding to the target agent, the method further includes: identifying the language type of the voice information through the language adaptation model, and determining a target large model from at least one large model corresponding to the target agent according to the language type; converting the voice information into text information through the language adaptation model, and converting the text information into a semantic vector; processing the voice information using the large model, skill knowledge base, and shared knowledge base corresponding to the target agent, including: processing the semantic vector using the target large model, skill knowledge base, and shared knowledge base corresponding to the target agent.

[0022] Identifying the language type of the voice information through the language adaptation model can select a large model that is more proficient in processing voice information, improving the accuracy of voice information processing. Moreover, converting the voice information into text information first and then into a semantic vector can eliminate interference factors in the voice information, improving the precision of voice information processing and thus enhancing the accuracy of interaction with the user.

[0023] In other embodiments provided by the present application, outputting the voice corresponding to the reply to the user includes: converting the reply into voice according to the language habits corresponding to the language type through the language adaptation model, and outputting the voice to the user.

[0024] Converting the obtained reply into a voice that conforms to the user's language habits before outputting it to the user can enable the user to obtain a localized reply voice, enhancing the user's experience.

[0025] In a second aspect of the present application, a voice interaction device is provided. The device is applied to a voice customer service system, which includes multiple agents and a shared knowledge base. Each agent corresponds to at least one large model and at least one skill knowledge base. The shared knowledge base stores historical data of the interaction between the user and the voice customer service system. The device includes: a receiving module, configured to receive the voice information input by the user; a determining module, configured to determine a target agent from other agents except the current agent among the multiple agents if it is determined based on the voice information that an agent switch is required; a processing module, configured to process the content corresponding to the voice information by using the large model and the skill knowledge base corresponding to the target agent and the shared knowledge base to obtain a reply; and an output module, configured to output the voice corresponding to the reply to the user.

[0026] In a third aspect of the present application, a computer device is provided, which includes a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the method in the first aspect.

[0027] In a fourth aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method in the first aspect are implemented.

[0028] In a fifth aspect of the present application, a computer program product is provided, which includes a computer program. When the computer program is executed by a processor, the steps of the method in the first aspect are implemented.

[0029] The voice customer service system provided in the second aspect of the present application, the computer device provided in the third aspect, the computer-readable storage medium provided in the fourth aspect, and the computer program product provided in the fifth aspect have the same or similar beneficial effects as the voice interaction method provided in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] By referring to the accompanying drawings and reading the following detailed description, the above and other objects, features, and advantages of the exemplary embodiments of the present application will become readily understood. In the drawings, several embodiments of the present application are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, where:

[0031] Figure 1 Schematic diagram of the application scenario of the voice interaction method in the embodiment of the present application Figure 1 ;

[0032] Figure 2 Schematic diagram of the process of the voice interaction method in the embodiment of the present application Figure 1 ;

[0033] Figure 3 Schematic diagram of the application scenario of the voice interaction method in the embodiment of the present application Figure 2 ;

[0034] Figure 4 Schematic flow of the voice interaction method in the embodiments of the present application Figure 2 ;

[0035] Figure 5 Schematic structure of the voice interaction device in the embodiments of the present application Figure 1 ;

[0036] Figure 6 Schematic structure of the voice interaction device in the embodiments of the present application Figure 2 ;

[0037] Figure 7 Schematic diagram of the structure of the computer device in the embodiments of the present application. Detailed implementation manners

[0038] The exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be fully conveyed to those skilled in the art.

[0039] It should be noted that, unless otherwise specified, the technical terms or scientific terms used in the present application should have the ordinary meanings understood by those skilled in the art to which the present application belongs.

[0040] Currently, in a voice customer service system, a model with complete functions is often configured. The structure of this model is relatively complex, and when replying based on the user's question, the processing process is relatively long, thereby reducing the efficiency of the interaction between the voice customer service system and the user.

[0041] In view of this, the embodiments of the present application provide a voice interaction method, a voice interaction device, a computer device, a computer-readable storage medium, and a computer program product, which split the complete functions in the voice customer service system of the related technology, and each functional unit corresponds to an intelligent agent. The intelligent agents used to implement each function process the user's question by combining a large model and a corresponding knowledge base. When facing the user's question, the intelligent agent communicating with the user can query in the corresponding skill knowledge base. Compared with a knowledge base that completely covers all knowledge, the data query volume can be reduced, and the reply efficiency can be improved. Moreover, when replying based on the user's current question, the historical data of the user in the shared knowledge base can also be combined for the reply, avoiding the output of replies that are not strongly targeted to the user's needs. By using the historical data in the shared knowledge base, it is also possible to avoid the need to transfer the user's historical data between the intelligent agents before and after the switch when the intelligent agent switches, reducing data transmission, improving the reply accuracy while also improving the reply efficiency, and improving the accuracy and efficiency of the interaction between the voice customer service system and the user.

[0042] It should be noted here that all components, data, and related processing methods involved in this application are authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0043] First, the application scenarios of the voice interaction method provided by some embodiments of this application will be described.

[0044] Figure 1 Schematic diagram of the application scenario of the voice interaction method in the embodiments of this application Figure 1 , see Figure 1 As shown, this scenario may include: a user and a voice customer service system (which can be simply referred to as the system).

[0045] The voice customer service system includes multiple agents and a shared knowledge base. Each agent corresponds to at least one large model and at least one skill knowledge base. The shared knowledge base stores the historical data of the interaction between the user and the voice customer service system.

[0046] Here, an agent can be understood as a customer service robot. Different agents can be customer service robots that are good at handling problems in different fields. In an agent, a customer service process is configured. The customer service process can be a process that guides the agent to interact with the user according to a predefined process. For example: when the agent is a debt collector, the customer service process is "send a debt collection reminder voice to the user - make corresponding replies according to the user's response - when the user indicates that they will repay on time, send an ending statement to the user - end the call". Different agents can be configured with different customer service processes or the same general customer service process, which can be specifically configured according to actual needs and is not limited here.

[0047] Here, the large model can be any pre-trained model with natural language understanding ability. The type and quantity of at least one large model are not limited here.

[0048] Here, the skill knowledge base can be a knowledge base that stores content, Q&A pairs, and knowledge corresponding to various keywords. The fields or details of the conversation scripts, Q&A pairs, and knowledge stored in different skill knowledge bases are different. For example: the debt collection knowledge base stores conversation scripts, Q&A pairs, and knowledge related to repayment, overdue, etc. Another example: the financial management knowledge base stores conversation scripts, Q&A pairs, and knowledge related to deposits, funds, etc. The specific content stored in different skill knowledge bases is not limited here.

[0049] In practical applications, an agent can correspond to one large model or multiple large models. When an agent corresponds to multiple large models, the multiple corresponding large models can be multiple large models with the same data processing capabilities. The agent can select one large model from the multiple large models for use. Alternatively, the multiple corresponding large models can be multiple large models with different data processing capabilities. The agent can comprehensively utilize the processing capabilities of each large model to give a response. Also, an agent can correspond to one skill knowledge base or multiple skill knowledge bases. In some scenarios, when an agent corresponds to one skill knowledge base, the agent is a customer service robot capable of handling domain problems or requirements corresponding to the skill knowledge base. When an agent corresponds to several skill knowledge bases, the agent is a composite customer service robot capable of handling domain problems or requirements corresponding to these skill knowledge bases. For the specific types and quantities of the large models and skill knowledge bases corresponding to each agent, they can be configured according to actual needs and are not limited here.

[0050] The shared knowledge base here can be a knowledge base that stores all the voice information sent by the user to the voice customer service system and the data obtained after processing all the user's voice information.

[0051] When the user establishes a connection with the voice customer service system, the voice customer service system can first allocate an agent to provide services to the user. The user can ask questions to the system. The current agent in the system processes the user's questions by calling the corresponding large model and skill knowledge base, generates a response, and feeds the response back to the user. At the same time, the system stores the user's questions in the shared knowledge base.

[0052] If the user asks a new question again, the current agent can continue to process the user's new question by calling the corresponding large model and skill knowledge base. If the system finds at this time that the current agent cannot handle the user's new question, it can switch to the next agent, that is, the target agent, based on the user's question, and then use the large model and skill knowledge base corresponding to the target agent and the user's historical questions in the shared knowledge base to process the user's new question, generate a response, and feed the response back to the user.

[0053] In practical applications, the voice customer service system can be any system that can interact with the user through voice and provide corresponding services. The services here can refer to customer service services in various industries or scenarios, such as: financial consultation, medical consultation, e-commerce scenarios, and so on. The specific type of the voice customer service system is not limited here.

[0054] For the large model, skill knowledge base, and shared knowledge base corresponding to the agent in the voice customer service system, they can be integrated into the voice customer service system and stored locally, or stored in the cloud. The specific storage location of the large model, skill knowledge base, and shared knowledge base corresponding to the agent in the voice customer service system is not limited here.

[0055] Next, a detailed description will be given to the voice interaction method provided by some embodiments of the present application.

[0056] Figure 2 Schematic diagram of the process of the voice interaction method in the embodiments of the present application Figure 1 , this method can be applied to a voice customer service system, see Figure 2 As shown, this method may include S21 - S24.

[0057] S21: Receive the voice information input by the user.

[0058] The voice customer service system sends a call request to the user. After the user passes the request and the voice customer service system sends a piece of voice to the user, if the user has a question based on the voice, the user can send voice information to the voice customer service system. For example: The user has handled a service reminder. The system will call the user actively at the corresponding time point and give a service reminder to the user. If the user has a question about the service reminder or other content at this time, the user can ask the system. At this time, the user's question to the system is the voice information input by the user to the system. Or, the user has some questions, the user sends a call request to the corresponding voice customer service system. After the voice customer service system passes the request, the voice customer service system establishes a connection with the user. At this time, the user can also send voice information to the voice customer service system.

[0059] The voice information here may refer to the information in the form of voice containing consultation content. For example: What is the current fixed deposit interest rate. The specific content of the voice information needs to be determined according to the actual needs of the user and the user's usual expression method, and is not limited here.

[0060] After the user accesses the voice customer service system, the voice customer service system first assigns a default agent to the user to process the user's voice information. In the user's view, the default agent provided by the system serves the user. The default agent can be an agent of the general desk type or other types of agents that match the user.

[0061] If the default agent can process the user's voice information, the large model and skill knowledge base corresponding to the default agent will be used to process the voice information. If the default agent cannot process the user's voice information, the system needs to switch to another agent to process the user's voice information. In other words, before using the large model and skill knowledge base corresponding to a certain agent to process the voice information, the system needs to first determine whether the certain agent can process the voice information.

[0062] In this embodiment, the agent currently providing services to the user is the current agent. The voice information sent by the user to the system is processed by the current agent.

[0063] S22: If it is determined based on the voice information that the agent needs to be switched, a target agent is determined from among the multiple agents except the current agent.

[0064] After the system receives the user's voice information, it first temporarily stores the voice information and then determines whether the current intelligent agent is capable of processing the voice information.

[0065] Since the system processes voice information through an agent, it is possible to determine whether the current agent is capable of processing voice information through voice information. If the current agent is capable of processing voice information, it is determined that there is no need to switch agents, and the current agent continues to be used to process voice information. If the current agent is not capable of processing voice information, it is determined that an agent needs to be switched, and the switched agent is used to process voice information.

[0066] In judging whether the current intelligent agent is capable of processing voice information, since the intelligent agent mainly generates responses through its skill knowledge base, it is possible to judge whether the voice information has corresponding knowledge in the knowledge base corresponding to the current intelligent agent. If it has corresponding knowledge, it is determined that the current intelligent agent is capable of processing the voice information. If it does not have corresponding knowledge, it is determined that the current intelligent agent is unable to process the voice information.

[0067] In addition to the above judgment methods, it is also possible to determine whether the current agent can process voice information based on the user's willingness. Since the user's willingness is expressed through the voice information input to the system, if the user is satisfied with the current agent, it will be expressed in the voice information. Therefore, the user's willingness can be determined by the content of the voice information, and thus the user's willingness can be used to determine whether the current agent can continue to process the voice information.

[0068] When the system determines that it is necessary to switch the intelligent agent, it can determine the target intelligent agent that can process the voice information from the multiple intelligent agents except the current intelligent agent.

[0069] When specifically determining, the target agent can be matched from the tags corresponding to multiple other agents based on the content in the voice information, or an agent with more comprehensive or more refined knowledge reserves can be selected from multiple other agents as the target agent.

[0070] S23: Use the large model, skill knowledge base, and shared knowledge base corresponding to the target agent to process the content corresponding to the voice information to obtain a reply.

[0071] The content corresponding to the voice information can be the voice information itself or the text after the voice information is processed into text.

[0072] After the system switches the current agent to the target agent, the temporarily stored voice information can be matched with the skill knowledge base corresponding to the target agent to generate a reply. Alternatively, the voice information can also be processed by the large model. During the process of processing the voice information by the large model, when the large model generates a reply corresponding to the voice information, it needs to refer to the knowledge in the skill knowledge base.

[0073] The scheme for matching with the skill knowledge base to generate a reply can be: the corresponding content of the voice information can be matched with the content in the skill knowledge base, and the successfully matched content can be used as the reply; or the semantics of the voice information can be recognized first, and then the recognized semantics can be matched with the content in the skill knowledge base, and the successfully matched content can be used as the reply; or the voice information can be combined with the user's historical data in the shared knowledge base first to obtain the user's intention, and then it can be matched with the skill knowledge base, and the successfully matched content in the skill knowledge base can be used as the reply.

[0074] When generating a reply through the large model, the voice information can be input into the large model, and the output of the large model is the reply. Or the voice information and the user's historical data in the shared knowledge base can be input into the large model, and the output of the large model is the reply. Or the user's intention can be determined first based on the voice information and the user's historical data in the shared knowledge base, and then the determined intention can be input into the large model, and the output of the large model is the reply.

[0075] During the process of the large model generating an output based on the input, the output can be generated using the knowledge reserves inherent in the large model, or the output can be generated based on the large model's understanding ability of natural language and the content in the skill knowledge base.

[0076] In the process of generating a response corresponding to the voice information, by invoking the shared knowledge base, the voice information can be understood more precisely, avoiding repeated responses to questions that the user is not concerned about, or avoiding the user from repeatedly restating the question, and more precisely understanding the user's intention. Also, in the use of the skill knowledge base and the large model, by using feature matching, semantic recognition, intention recognition, knowledge base extraction, and large model generation alone or in combination according to different situations, a more accurate response can be generated, improving the accuracy of voice interaction.

[0077] It should be noted here that in order to improve the processing efficiency of the large model, when using the historical data of the user in the shared knowledge base, relevant data can be selected from all the historical data of the user for use. The relevant data can refer to several pieces of the latest stored data in all the historical data of the user, or data with a relevance higher than the preset relevance to the voice information.

[0078] S24: Output the voice corresponding to the response to the user.

[0079] After generating a response using the skill knowledge base corresponding to the target agent and the large model, if the response is an audio, the target agent can directly output the response generated by the large model to the user. If the response is a text, the target agent needs to first convert the text to voice and then output the converted voice to the user.

[0080] Since the input from the user to the voice customer service system is voice, and the voice customer service system also outputs voice to the user, it can match the user's input and interact with the user in the same type, improving the user experience.

[0081] As can be seen from the above content, the voice interaction method provided by the embodiments of the present application configures multiple agents, and configures a corresponding large model and skill knowledge base for each agent, and then automatically selects the corresponding agent to generate a response according to different questions of the user. Since the skill knowledge base queried when processing the user's questions is split from a complete knowledge base into multiple sub-knowledge bases with different professional knowledge, the data query volume is reduced, so the query of the current knowledge can be completed faster, and the response can be generated faster, improving the voice interaction efficiency. And the switched agent can view the voice information sent by the user to the agent before switching through the shared knowledge base, without the need for the agents to transmit the user's historical voice information to each other, improving the speed of information sharing between the agents, enabling the switched agent to generate a response faster and more accurately based on the historical data, and improving the efficiency and accuracy of information interaction.

[0082] In some embodiments, as Figure 2 a refinement and extension of the method shown, the embodiments of the present application also provide a voice interaction method.

[0083] Figure 3Schematic diagram of the application scenario of the voice interaction method in the embodiments of this application Figure 2 , see Figure 3 As shown, this scenario may include: a user and a voice customer service system.

[0084] In the voice customer service system, in addition to including multiple agents and a shared knowledge base, it also includes a language adaptation model. Moreover, the skill knowledge base corresponding to the agent can be refined into a rule knowledge base and a content knowledge base. The rule knowledge base stores various keywords and the content corresponding to various keywords. The content knowledge base stores various knowledge for reference when the large model generates a reply.

[0085] The language adaptation model here may refer to a model that can perform language type recognition and text processing on voice information. The specific type of the language adaptation model is not limited here.

[0086] For the language adaptation model, it can be integrated into each agent. In different agents, the integrated language adaptation models can be the same or different. When the system receives voice information, it uses the language adaptation model corresponding to the current agent to process the voice information, obtains the text content and language type corresponding to the voice information, and then generates a reply of the same language type based on the skill knowledge base corresponding to the current agent and the large model based on the processed voice information. In some embodiments, the language type can be classified by country, such as Chinese, English, etc.; the language type can also be classified by region, such as Mandarin, Minnan dialect, Sichuan dialect, etc.

[0087] For the language adaptation model, it can also be independently set in the system separately from the agent. When the system receives voice information, it first uses the separately set language adaptation model to process the voice information, obtains the text content and language type corresponding to the voice information, and then generates a reply of the same language type based on the skill knowledge base corresponding to the current agent and the large model based on the processed voice information.

[0088] In the process of generating a reply based on the skill knowledge base and the large model: when using the skill knowledge base alone to generate a reply, it is to match the processed voice information with the content in the skill knowledge base, and then output a reply based on the successfully matched content. Therefore, the content in the skill knowledge base at this time needs to be various keywords and the content or answers corresponding to various keywords, that is, the rule knowledge base; when using the skill knowledge base and the large model comprehensively to generate a reply, although the large model has the ability of natural language understanding and processing, and can generate a reply corresponding to the voice information alone, the system generally provides more professional services. In order to improve the accuracy of the reply, the large model can mainly refer to the content in the skill knowledge base during the process of generating a reply. Therefore, the content in the skill knowledge base at this time needs to be various professional or specific field-related knowledge, that is, the content knowledge base.

[0089] After generating a response based on the current agent, generally speaking, the generated response is text. To adapt to the voice information input by the user, the generated text-based response can be sent to a language adaptation model, which converts the text into audio, and then outputs the response in audio form to the user.

[0090] For example, the user inputs voice to the voice customer service system. The system calls the Automatic Speech Recognition (ASR) model to convert the voice into text and obtain the language identification. Then it calls the skill knowledge base corresponding to the current agent and the large model to process the text based on the language identification to get a response. Furthermore, it calls the Text-to-Speech (TTS) component to process the response into a response voice based on the language identification, and then outputs the response voice to the user. The ASR model and the TTS component here are the specific forms of the language adaptation model.

[0091] Figure 4 Flow schematic of the voice interaction method in the embodiments of this application Figure 2 , see Figure 4 As shown, the method may include S41 - S48.

[0092] S41: Receive the voice information input by the user.

[0093] The specific implementation manner of step S41 here is the same as that of step S21 in the foregoing embodiments. For relevant descriptions, refer to the foregoing embodiments and will not be elaborated here.

[0094] After the voice customer service system receives the voice information input by the user, it first determines whether the current agent can process the voice information. When the user first accesses the system, the current agent is the initial default agent. After the user accesses the system for a period of time, the current agent may still be the initial default agent or may be the agent after switching. The specific current agent changes dynamically according to the user's voice information.

[0095] If it is determined that the current agent can process the voice information, then use the current agent to process the voice information. If it is determined that the current agent cannot process the voice information, then switch the agent and use the switched agent to process the voice information.

[0096] S42: If the voice information indicates that the agent needs to be switched, or there is no response corresponding to the voice information in the skill knowledge base corresponding to the current agent, then determine that the agent needs to be switched.

[0097] If there is no reply corresponding to the voice information in the skill knowledge base corresponding to the current agent, it indicates that the current agent cannot reply to the user's question, and it is determined that the agent needs to be replaced. For example, the voice information is "Introduce the specific allocation of the 500,000 financial management", and the current agent is the general desk customer service, and the corresponding skill knowledge base is the general desk knowledge base. The general desk knowledge base mainly stores knowledge such as reception, transfer, reservation, etc. Obviously, it cannot give an accurate reply to the specific allocation of the 500,000 financial management. At this time, it is determined that the general desk agent needs to be switched to other agents.

[0098] If there is content in the voice information to switch the agent, it indicates that the user feels that the current agent cannot meet his needs, and it is also determined that the agent needs to be switched. For example, the voice information is "Transfer me to the financial management customer service", and the current agent is the general desk customer service. At this time, the user has clearly expressed in the voice information that the current agent, the general desk customer service, needs to be switched to another agent, the financial management customer service.

[0099] For the content in the voice information that can represent switching the agent, in addition to the user clearly expressing in the voice information, it can also be obtained through semantic recognition and intention recognition of the voice information.

[0100] Specifically, the above step S42 may include:

[0101] S42a: If the voice information matches the preset keyword, and the preset keyword is used to indicate other agents, it is determined that the voice information indicates that the agent needs to be switched.

[0102] For each agent in the voice customer service system, there is a corresponding preset keyword. For example: For example, the agent of the financial management customer service corresponds to the preset keywords "deposit, fund, etc.", and the agent of the general desk customer service corresponds to the preset keywords "chat in person, offline, reservation, etc.".

[0103] If the voice information input by the user matches a certain preset keyword, and the agent corresponding to the preset keyword is not the current agent, then the agent corresponding to the preset keyword is the target agent that needs to be switched.

[0104] In some embodiments, sometimes the current agent can also answer the user's questions, but the answering effect is not very ideal, and the user can barely accept the reply of the current agent, but inadvertently expresses the need to switch the agent. In order to improve the accuracy of agent switching, it is possible to combine the emotion threshold in the voice information to judge whether to switch the agent.

[0105] Specifically, the above step S42a may include: If the voice information matches the preset keyword, and the emotion intensity of the voice information is higher than the emotion intensity threshold, it is determined that the voice information indicates that the agent needs to be switched.

[0106] That is to say, when the voice information matches the preset keyword, continue to obtain the emotional intensity from the voice information. When the obtained emotional intensity is higher than the emotional intensity threshold, it is determined that the intelligent agent needs to be switched.

[0107] The emotional intensity here can refer to the intensity of emotions in the voice information. Emotions can include anger, anxiety, etc. For example, the user says "Change me to a professional financial advisor" in a very angry tone. At this time, there is not only an indication to switch the intelligent agent in the voice information, but also a relatively high emotional intensity. The specific value of the emotional intensity threshold can be determined according to the actual situation and is not limited here.

[0108] The emotional intensity here can also refer to the number of occurrences of specified words in the voice information. The specified words can include words such as affirmation, for example: "Yes", "Right", etc. And the number of occurrences can be once, twice, multiple times, etc. For example, the current intelligent agent, the general desk customer service, asks the user whether to switch to the target intelligent agent, the financial advisor. The voice information replied by the user to the general desk customer service is "Yes Yes". In the voice information, the specified word "Yes" appears, and the specified word appears twice. At this time, it can be determined that the emotional intensity in the voice information is higher than the preset intensity threshold, and then it is determined that the intelligent agent needs to be switched.

[0109] S42b: If the semantic indication recognized from the voice information indicates other intelligent agents, it is determined that the voice information indicates that the intelligent agent needs to be switched.

[0110] When performing semantic recognition, any one of the following methods or a combination of any methods can be specifically used:

[0111] Rule-based methods, such as: rule engines, regular expressions, etc.;

[0112] Statistics-based methods, such as: N-gram models, Hidden Markov Model (HMM), etc.;

[0113] Machine learning-based methods, such as: Support Vector Machine (SVM), decision trees and random forests, naive Bayes, etc.;

[0114] Deep learning-based methods, such as: Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM), Gated Recurrent Unit (GRU), Convolutional Neural Network (CNN), Transformer model, etc.;

[0115] Knowledge graph-based methods, such as: knowledge graph, ontology reasoning, etc.;

[0116] Pre-trained language model-based methods, specifically, they can be Natural Language Processing (NLP) methods, such as: BERT, GPT, RoBERTa, XLNet, etc.;

[0117] Semantic agent annotation-based methods, such as: Semantic Role Labeling (SRL), etc.;

[0118] Word vector-based methods, such as: Word2Vec, GloVe, FastText, etc.;

[0119] Attention mechanism-based methods, such as: self-attention mechanism, multi-head attention mechanism, etc.

[0120] After identifying the semantics corresponding to the speech information, match the identified semantics with the labels of other agents except the current agent among multiple agents. If the match is successful, the agent corresponding to the matched label is the target agent. If the match fails, it means that there is no more suitable agent at present, and continue to interact with the user through the current agent.

[0121] Since there may be some differences in the words used between the identified semantics and the preset labels in the agent, in order to ensure a successful match and achieve a smooth switch of the agent, the condition for a successful match can be set to a similarity greater than the preset similarity. The preset similarity here can be any value less than 100%, such as: 99%, 90%, etc.

[0122] S42c: If the intention recognized from the speech information and historical data using a large model indicates other agents, it is determined that the speech information indicates a need to switch agents.

[0123] To improve the accuracy of intent recognition, in the intent recognition of voice information, the context information of the voice information can be referred to. However, if the amount of data in the context information of the voice information is large, it will increase the data processing burden on the large model. To balance the accuracy and efficiency of user intent recognition, more relevant information to the voice information can be selected from the context information of the voice information. That is, data with a relevance higher than a preset relevance to the voice information is selected from historical data. Here, the relevance can refer to the number of recent data entries, that is, the first few voice messages of the voice information are selected. Here, the relevance can also refer to the similarity of information, that is, voice messages with similar content sent by the user before the voice information are selected (not necessarily the first few voice messages of the voice information).

[0124] The voice information, all or part of the data in the historical data, and the requirement for intent acquisition are input into the large model, and the large model outputs the user's intent.

[0125] Here, the large model can be a model with natural language understanding ability. The specific type of the large model is not limited here.

[0126] After the user's intent is recognized, the recognized intent is matched with the labels of other agents except the current agent among multiple agents. If the match is successful, the agent corresponding to the matched label is the target agent. If the match fails, it means that there is no more suitable agent currently, and continue to interact with the user through the current agent.

[0127] Since there may be some differences in the words used between the recognized intent and the preset labels in the agent, to ensure a successful match and achieve a smooth switch of the agent, the condition for a successful match can be set to a similarity greater than a preset similarity. Here, the preset similarity can be any value less than 100%, for example: 99%, 90%, etc.

[0128] After determining that the agent needs to be switched, it is necessary to determine the target agent that can process the voice information from other agents except the current agent among multiple agents in the voice customer service system.

[0129] S43: Based on the voice information, multiple candidate agents are determined from other agents, and then the candidate agent with the highest priority weight is selected from the multiple candidate agents to be determined as the target agent.

[0130] When determining multiple candidate agents, the preset keywords, semantics, or intents extracted from the voice information before can be matched with the labels of each other agent, and other agents corresponding to the labels with a match degree greater than the preset match degree are used as candidate agents.

[0131] To ensure that multiple candidate agents can be determined, the preset matching degree can be set lower, for example: 50%, 60%, etc. The specific value of the preset matching degree can be determined according to the actual situation and is not limited here.

[0132] After determining multiple candidate agents, since each agent is correspondingly configured with a priority weight, and the priority weight is set manually according to the importance of each agent in the actual business, therefore, a target agent that is more suitable for processing voice information can be selected from multiple candidate agents based on the priority weight.

[0133] For example, assume the voice message is "Introduce the financial management configuration above 500,000". Two candidate agents are determined. One candidate agent is a financial advisor, and the corresponding priority weight is 2. The other candidate agent is a large-amount financial advisor, and the corresponding priority weight is 3. Generally speaking, users who need to switch agents often have relatively high requirements. Therefore, selecting the large-amount financial advisor corresponding to the priority weight of 3 to answer the user's question about the financial management configuration above 500,000 can provide a more accurate response to the user.

[0134] S44: If the voice message indicates that there is no need to switch agents currently, or the user does not input a voice message within the preset time, it is determined that there is no need to switch agents currently.

[0135] If there is content in the voice message indicating that there is no need to switch agents, it means that the user feels that the current agent can meet their needs and is willing to continue interacting with the current agent. Then it is determined that there is no need to switch agents. For example: The voice message is "I still want to ask you another question". No matter what kind of customer service the current agent is, it means that the user is willing to continue with it. At this time, it is determined that there is no need to switch agents.

[0136] If there is a reply corresponding to the voice message in the skill knowledge base corresponding to the current agent, it means that the current agent can answer the user's question. Then it is determined that there is no need to replace the agent. For example: The voice message is "Introduce the specific allocation of 500,000 in financial management", and the current agent is a financial management customer service, and the corresponding skill knowledge base is a financial management knowledge base. The financial management knowledge base stores various professional financial management knowledge and can accurately answer the specific allocation of 500,000 in financial management. At this time, it is determined that there is no need to switch the main agent to other agents.

[0137] Regardless of whether it is necessary to switch agents currently, after the current agent obtains the voice message, at the same time, the voice message can also be stored in the shared knowledge base. And based on the voice message and historical data of this user in the shared knowledge base, a portrait of this user is generated for the agent to call when processing the voice messages sent by this user subsequently. The portrait has a small amount of data and can also comprehensively represent the user, which can improve the accuracy and efficiency of the reply.

[0138] S45: Store the voice information in the shared knowledge base, and generate a user profile based on the voice information and historical data, and store the profile in the shared knowledge base for the agent to call.

[0139] The profile here can refer to various features that can help quickly understand the user, such as: areas of interest, personality, location, etc.

[0140] After the current voice information is stored in the shared knowledge base, the current agent can extract the feature information corresponding to each feature requirement from the voice information in the shared knowledge base according to the feature requirements of the profile. If the feature information corresponding to a certain feature requirement is not extracted, it can be ignored, and then the extracted feature information is updated in the existing profile of the user. In this way, the update efficiency of the user profile can be improved.

[0141] In addition, the profile of the user can also be regenerated based on all the voice information of the user in the shared knowledge base using a large model. In this way, the relevance between the various voice information of the user can be comprehensively considered, and the accuracy of the user profile generation can be improved.

[0142] In the shared knowledge base, not only the voice information of the user is stored, but also the profile of the user is stored. When the agent processes the current voice information, in addition to the historical data related to the current voice information in the shared knowledge base, the profile of the user in the shared knowledge base can also be used.

[0143] After the voice information is stored in the shared knowledge base and the system switches to the target agent, the target agent can obtain the voice information, historical data, user profile, etc. from the shared knowledge base, so as to process the voice information to generate a reply. At this time, the voice customer service system has switched to the target agent to interact with the user.

[0144] S46: Identify the language type of the voice information through a language adaptation model, determine the target large model from at least one large model corresponding to the target agent according to the language type, and convert the voice information into text information through the language adaptation model, and convert the text information into a semantic vector.

[0145] The voice adaptation model here can be any model that can identify the language type of the voice information and convert the voice information into text information and then into a semantic vector. The specific type of the language adaptation model is not limited here.

[0146] The language types here include regional dialects and languages. Regional dialects such as Sichuan dialect, Cantonese, etc. Languages such as English, German, etc.

[0147] Since an agent may be pre-configured with multiple large models, and the language types that each large model is good at processing are different, it is possible to select a large model that is good at processing the language type of the voice information according to the language type of the voice information, which can improve the accuracy of voice information processing. For example: The target agent is pre-configured with large model a and large model b, where large model a is good at processing English and large model b is good at processing various dialects. If the language type corresponding to the voice information is Sichuan dialect, then large model b is selected to process the voice information.

[0148] When determining the target large model, the voice information can also be converted into text information through a language adaptation model, and then the text information is converted into a semantic vector. In this way, some environmental interference factors in the voice information can be excluded, and the accuracy of the large model in processing the voice information can be improved.

[0149] The semantic vector here can refer to the representation of mapping high-dimensional discrete data such as words, sentences or documents into a low-dimensional continuous vector space. Through the semantic vector, the large model can more accurately grasp the user's intentions, needs, etc., so as to generate more accurate responses.

[0150] After determining the target large model that can more accurately process the voice information and the semantic vector that can accurately represent the essential content in the voice information, the semantic vector is input into the target large model. The target large model can accurately generate response content by referring to the skill knowledge base and the shared knowledge base corresponding to the target agent.

[0151] S47: Process the semantic vector using the target large model, the skill knowledge base corresponding to the target agent, and the shared knowledge base.

[0152] Since the target agent corresponds to a large model and a skill knowledge base, and the skill knowledge base and the large model can either process the voice information separately or in combination. When processing the voice information separately based on the skill knowledge base, it is actually to match the voice information with the content in the knowledge base. Therefore, the skill knowledge base includes a rule knowledge base. Various keywords and the corresponding content are stored in the rule knowledge base. When processing the voice information in combination based on the large model and the skill knowledge base, the knowledge base is actually the content that the large model needs to refer to when generating a response. Therefore, the skill knowledge base includes a content knowledge base. Various knowledge for the large model to refer to when generating a response are stored in the content knowledge base. It can be seen that there are multiple ways for the target agent to process the voice information and generate a response.

[0153] Specifically, the above step S47 may include:

[0154] Step S47a: Match the content corresponding to the voice information with the keywords in the rule knowledge base, and determine the content corresponding to the successfully matched keywords as the reply.

[0155] In the rule knowledge base, various keywords and the content corresponding to various keywords are pre-configured. That is to say, it is pre-set for the intelligent agent what kind of reply to give to what kind of questions from the user.

[0156] Specifically, the text information output by the language adaptation model can be matched with the keywords in the rule knowledge base. If the match fails, it means that there is no corresponding statement for the user's question, and other methods such as semantic recognition and intention recognition can be used to generate a reply. If the match is successful, it means that there is a corresponding statement for the user's question, and the content corresponding to the successfully matched keywords can be used as the reply.

[0157] Step S47b: Identify the semantics in the content corresponding to the voice information, match the semantics with the keywords in the rule knowledge base, and determine the content corresponding to the successfully matched keywords as the reply.

[0158] Specifically, the semantics can be obtained from the text information or semantic vector output by the language adaptation model, and then the semantics are matched with the keywords in the rule knowledge base. If the match fails, it means that there is no corresponding statement for the user's question, and other methods such as keyword matching and intention recognition can be used to generate a reply. If the match is successful, it means that there is a corresponding statement for the user's question, and the content corresponding to the successfully matched keywords can be used as the reply.

[0159] Step S47c: Use the large model to identify the intention of the historical data and the content corresponding to the voice information, match the intention with the keywords in the rule knowledge base, and determine the content corresponding to the successfully matched keywords as the reply.

[0160] Specifically, the text information or semantic vector, historical data, and the language for the large model to identify the intention can be input into the large model. Through the large model's ability to understand natural language and combining the context information of the voice information, the user's intention can be accurately obtained, and then the intention is matched with the keywords in the rule knowledge base. If the match fails, it means that there is no corresponding statement for the user's question, and other methods such as keyword matching and semantic recognition can be used to generate a reply. If the match is successful, it means that there is a corresponding statement for the user's question, and the content corresponding to the successfully matched keywords can be used as the reply.

[0161] Step S47d: Input the content corresponding to the voice information and the historical data into the large model corresponding to the target intelligent agent, and generate a reply based on the content knowledge base through the large model corresponding to the target intelligent agent.

[0162] Specifically, text information or semantic vectors, historical data, and prompts for the large model to refer to the content knowledge base to generate responses can be input into the large model. Based on the prompts, when generating responses based on the text information or semantic vectors and historical data, the large model will preferentially use the knowledge in the content knowledge base as the response, so that the generated response is more accurate and professional.

[0163] In the specific process of the large model generating responses, the large model also has its own knowledge reserve. Generally speaking, the large model uses its understanding ability of natural language combined with its own knowledge reserve to generate responses. In this embodiment, based on the prompts for generating responses from the reference content knowledge base, after understanding the input text information or semantic vectors and historical data, the large model first refers to the knowledge in the content knowledge base to generate a response, and then refers to its own knowledge reserve to enrich the response content. Or, after understanding the input text information or semantic vectors and historical data, the large model first generates an initial response based on its own knowledge reserve, and then, based on the prompts for generating responses from the reference content knowledge base, replaces the content in the generated initial response that is similar to the content in the content knowledge base with the corresponding knowledge in the content knowledge base to obtain the final response. The similar content here can refer to a sentence or a paragraph with the same one or more keywords, or a sentence or a paragraph with a similarity higher than a preset threshold.

[0164] It should be noted here that for the respective response generation methods corresponding to the above steps S47a, S47b, S47c, and S47d, they can be used selectively, or several methods can be used simultaneously. When multiple response generation methods are adopted, the responses generated by each response method can be combined to generate the final response, or the response generated by one response method can be selected as the final response according to a certain standard. Here, the standard can be based on the highest priority corresponding to the response method, the largest number of identical or similar responses, and so on.

[0165] Generally speaking, the output response is standard text content. To make the content output by the system to the user more in line with the user's habits, the generated response can be processed again through a language adaptation model to obtain a voice output that is more in line with the user's habits.

[0166] S48: Convert the response into speech according to the language habits corresponding to the language type through the language adaptation model, and output the speech to the user.

[0167] The language habits here can refer to the language type itself, or the localized idioms, cultural taboos, etc. corresponding to the language type.

[0168] In the process of converting the generated reply into speech according to the determined language habits, the reply can be first converted according to idioms, cultural taboos, etc. corresponding to the language type, and then converted into reply speech according to the language type, and finally the reply speech is output to the user.

[0169] For example, assume that the voice message is "Evaluate xx stock" in Cantonese, and the generated reply is "The stock price is below HK$1". In Cantonese idioms, "The stock price is below HK$1" is called "penny stock". Therefore, the language adaptation model can first convert the text "The stock price is below HK$1" into the text "penny stock", and then convert the text "penny stock" into the Cantonese "penny stock", and output the Cantonese "penny stock" to the user.

[0171] In this way, the user can obtain a reply speech that conforms to their language habits.

[0172] Finally, a complete embodiment is used to clearly and completely illustrate the voice interaction method provided by the embodiments of the present application again.

[0173] First, relevant personnel need to configure the agent in the voice customer service system.

[0174] 1. Create a customer service robot on the interface of the agent creation platform.

[0175] 2. Configure the customer service robot. Specifically, it includes:

[0176] (1) Configure basic parameters. The basic parameters can include the large model used, the role of the robot, responsibilities, whether there is a long-term memory function, etc.

[0177] (2) Configure the knowledge base.

[0178] It can include whether to enable the knowledge base. After enabling, knowledge retrieval will be performed for each round of conversation, and the content recalled by the knowledge retrieval will be added to the prompt. If it is closed, this process will not be performed.

[0179] It can also include whether to use the hit Q&A content for reply. In some scenarios, if it is necessary to reply strictly according to the Q&A content in the knowledge base, this function can be enabled. After enabling, when the relevance of the knowledge retrieval is greater than [0.65 - 1], and the number of knowledge that meets this threshold condition is in the range of [1 - 10], this logic will take effect. The relevance and the number of knowledge can be adjusted according to the actual effect.

[0180] (3) Configure the dialogue flow. It can include dialogue nodes, reply words for each node, etc. Specifically, when the user's intention in the conversation is higher than a certain threshold, the dialogue flow can be jumped to. Of course, the dialogue flow can also be directly opened in a new session.

[0181] (4) Configure the rules.

[0182] It can include synonym replacement and is suitable for mapping configuration between popular expressions and professional terms. If the "original word" configured is included in the user's question, the platform will automatically change the "original word" in the user's question to the "replacement word" configured (this function takes effect after being enabled).

[0183] It can also include keyword hit replies and is applicable to scenarios where direct replies are made according to specified phrases for keywords or sensitive words. If the "keyword" configured is included in the user's question, the platform will directly use the content of the "specified reply" configured, and the hit method can be selected (this function takes effect after being enabled).

[0184] After configuring each agent in the system, the configured agents in the system can be used to provide services to users.

[0185] The following provides an actual call scenario for illustration.

[0186] 1. The user makes a call to the voice customer service system.

[0187] 2. The customer service robot 1 in the system (which is the pre-configured agent 1) answers the call, and the user states the requirements to the customer service robot 1.

[0188] 3. If the system believes that the customer service robot 1 can answer the user's requirements, it will answer based on the user's requirements through the customer service robot 1. If the system believes that the customer service robot 1 cannot answer the user's requirements, it will prompt the user to transfer to the new customer service robot 2 (which is the pre-configured agent 2).

[0189] When the system determines whether the customer service robot 1 can answer the user's requirements, it can adopt any of the following methods:

[0190] (1) Based on process configuration. That is, the user's requirements are matched with the configuration content of the customer service robot 1. If the user's requirements match the configured content corresponding to the customer service robot 1 successfully, it is determined that the customer service robot 1 can answer the user's requirements, and then the successfully matched content in the configuration content is used as the reply. If the user's requirements do not match the configured content corresponding to the customer service robot 1, it is determined that the customer service robot 1 cannot answer the user's requirements, and then it switches to the customer service robot 2.

[0191] (2) Identify the semantics of the current round of conversation based on the NLP model. That is, input the user's current needs into the NLP model so that the NLP model can identify the user's semantics, and then match the user's semantics with the configuration content of the customer service robot 1. If the identified semantics match the corresponding configuration content of the customer service robot 1 successfully, it is determined that the customer service robot 1 can answer the user's needs, and then the successfully matched content in the configuration content is used as the reply. If the identified semantics do not match the corresponding configuration content of the customer service robot 1, it is determined that the customer service robot 1 cannot answer the user's needs, and then switch to the customer service robot 2.

[0192] (3) Identify the context intention based on the large model. That is, input the user's current needs and multiple previous needs into the large model so that the large model can identify the user's intention, and then match the user's intention with the configuration content of the customer service robot 1. If the identified intention matches the corresponding configuration content of the customer service robot 1 successfully, it is determined that the customer service robot 1 can answer the user's needs, and then the successfully matched content in the configuration content is used as the reply. If the identified intention does not match the corresponding configuration content of the customer service robot 1, it is determined that the customer service robot 1 cannot answer the user's needs, and then switch to the customer service robot 2.

[0193] 4. The user does not hang up the phone, and the customer service robot 2 answers based on the user's needs.

[0194] When the customer service robot 2 answers the user, it can obtain the user's needs and their context information from the shared knowledge base and generate a reply. The specific ways to generate a reply can be any of the following:

[0195] (1) Match the user's needs with the configuration content, and then use the successfully matched content in the configuration content as the reply, specifically including using the Q&A reply hit in the configuration content, using the conversation flow words and expressions hit in the configuration content, jumping to the next link according to the rules hit in the configuration content, etc.

[0196] (2) Identify the semantics based on the user's needs using the NLP model, and then match the identified semantics with the configuration content, so as to use the successfully matched content in the configuration content as the reply, specifically including using the Q&A reply hit in the configuration content, using the conversation flow words and expressions hit in the configuration content, jumping to the next link according to the rules hit in the configuration content, etc.

[0197] (3) Identify the user's intention using a large model based on the user's needs and the context information of the needs, and then match the identified intention with the configured content, so as to use the successfully matched content in the configured content as the reply. Specifically, it includes using the Q&A replies hit in the configured content, using the conversation flow words and phrases hit in the configured content, and jumping to the next link according to the rules hit in the configured content, etc.

[0198] (4) Input the user's needs, the context information of the needs, and the prompt information for generating a reply by referring to a specified knowledge base into the large model. When generating a reply based on the needs and the context information of the needs, the large model refers to the knowledge in the specified knowledge base and generates the reply content, so as to use the reply content output by the large model as the reply.

[0199] 5. The user continues to state their needs. If the customer service robot 2 cannot answer, the system transfers the user to a new customer service robot. If the customer service robot 2 can answer or meet the needs, the system continues to use the customer service robot 2 to answer or call the corresponding component based on the needs.

[0200] The corresponding component here may refer to the functional module corresponding to the need. For example: When the user wants to purchase a certain financial product in an application (Application, APP), the user expresses the purchase need of the financial product to the voice customer service system in the application, and the customer service robot in the voice customer service system can call out the purchase interface of the financial product in the application and display it to the user in the application, and the user can conveniently purchase the financial product in the application. Another example: When the user wants to purchase a certain financial product, the user calls the voice customer service system and tells the voice customer service system the need to purchase the financial product, and the customer service robot in the voice customer service system broadcasts the purchase guidance to the user by voice.

[0201] 6. The user's needs are met, and the user hangs up the call with the voice customer service system.

[0202] So far, the description of the voice interaction method provided by the embodiments of the present application has been completed.

[0203] Based on the same inventive concept, the embodiments of the present application also provide a voice interaction device.

[0204] The voice interaction device is applied to a voice customer service system. The voice customer service system includes multiple agents and a shared knowledge base. Each agent corresponds to at least one large model and at least one skill knowledge base. The shared knowledge base stores the historical data of the interaction between the user and the voice customer service system.

[0205] Figure 5 For the structural schematic of the voice interaction device in the embodiments of the present application Figure 1 , see Figure 5As shown, the device may include:

[0206] A receiving module 51, configured to receive voice information input by a user.

[0207] A determining module 52, configured to determine a target agent from other agents except the current agent among multiple agents if it is determined based on the voice information that an agent needs to be switched.

[0208] A processing module 53, configured to process the content corresponding to the voice information by using the large model, skill knowledge base, and shared knowledge base corresponding to the target agent to obtain a reply.

[0209] An output module 54, configured to output the voice corresponding to the reply to the user.

[0210] In some embodiments, as Figure 5 a refinement and extension of the shown device, an embodiment of the present application further provides a voice interaction device.

[0211] The voice interaction device is applied to a voice customer service system. The voice customer service system includes multiple agents, a shared knowledge base, and a language adaptation model.

[0212] Each agent corresponds to at least one large model and at least one skill knowledge base. Each agent has a priority weight. The skill knowledge base includes a rule knowledge base and a content knowledge base. Various keywords and the content corresponding to various keywords are stored in the rule knowledge base. Various knowledge for reference when the large model generates a reply is stored in the content knowledge base.

[0213] Historical data of the interaction between the user and the voice customer service system is stored in the shared knowledge base.

[0214] Figure 6 For the structural schematic of the voice interaction device in the embodiment of the present application Figure 2 , see Figure 6 As shown, the device may include:

[0215] A receiving module 61, configured to receive voice information input by a user.

[0216] A storage module 62, configured to store the voice information in the shared knowledge base, and generate a user portrait based on the voice information and historical data, and store the portrait in the shared knowledge base for the agent to call.

[0217] A switching module 63, configured to determine that an agent needs to be switched if the voice information indicates that an agent needs to be switched, or if there is no reply corresponding to the voice information in the skill knowledge base corresponding to the current agent; determine that there is no need to switch the current agent if the voice information indicates that there is no need to switch the current agent, or if there is a reply corresponding to the voice information in the skill knowledge base corresponding to the current agent.

[0218] The switching module 63 is specifically configured to determine that the voice message indicates a need to switch intelligent agents if the voice message matches a preset keyword, and the preset keyword is used to indicate other intelligent agents; and / or, if the semantics recognized from the voice message indicates other intelligent agents, then determine that the voice message indicates a need to switch intelligent agents; and / or, if the intent recognized from the voice message and historical data using a large model indicates other intelligent agents, then determine that the voice message indicates a need to switch intelligent agents.

[0219] The switching module 63 is specifically configured to determine that the voice message indicates a need to switch intelligent agents if the voice message matches a preset keyword and the emotional intensity of the voice message is higher than the emotional intensity threshold.

[0220] The determination module 64 is configured to, if it is determined based on the voice message that a switch of intelligent agents is required, determine multiple candidate intelligent agents from other intelligent agents based on the voice message; and select the candidate intelligent agent with the highest priority weight from the multiple candidate intelligent agents and determine it as the target intelligent agent.

[0221] The adaptation module 65 is configured to identify the language type of the voice message through a language adaptation model, and determine a target large model from at least one large model corresponding to the target intelligent agent according to the language type; convert the voice message into text information through the language adaptation model, and convert the text information into a semantic vector.

[0222] The processing module 66 is configured to process the semantic vector using the target large model, the skill knowledge base corresponding to the target intelligent agent, and the shared knowledge base to obtain a reply.

[0223] The processing module 66 is specifically configured to match the content corresponding to the voice message with the keywords in the rule knowledge base, and determine the content corresponding to the successfully matched keywords as the reply; and / or, identify the semantics in the content corresponding to the voice message, match the semantics with the keywords in the rule knowledge base, and determine the content corresponding to the successfully matched keywords as the reply; and / or, use a large model to identify the intent of the content corresponding to the historical data and the voice message, match the intent with the keywords in the rule knowledge base, and determine the content corresponding to the successfully matched keywords as the reply; and / or, input the content corresponding to the voice message and the historical data into the large model corresponding to the target intelligent agent, and generate a reply based on the content knowledge base through the large model corresponding to the target intelligent agent.

[0224] The output module 67 is configured to convert the reply into voice according to the language habit corresponding to the language type through the language adaptation model, and output the voice to the user.

[0225] It should be noted here that the description of the above device embodiments is similar to that of the above method embodiments and has similar beneficial effects to those of the method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0226] Based on the same inventive concept, an embodiment of the present application also provides a computer device.

[0227] Figure 7 For the structural schematic diagram of the computer device in the embodiments of the present application, see Figure 7 As shown, the computer device may include: a memory 71, a processor 72, and a computer program stored on the memory 71. The processor 72 executes the computer program to implement the method in the foregoing embodiments.

[0228] It should be noted here that the description of the above computer device embodiments is similar to that of the above method embodiments and has similar beneficial effects to those of the method embodiments. For the technical details not disclosed in the computer device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0229] Based on the same inventive concept, an embodiment of the present application also provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the method in the foregoing embodiments is implemented.

[0230] It should be noted here that the description of the above computer-readable storage medium embodiments is similar to that of the above method embodiments and has similar beneficial effects to those of the method embodiments. For the technical details not disclosed in the computer-readable storage medium embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0231] Based on the same inventive concept, an embodiment of the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the method in the foregoing embodiments is implemented.

[0232] It should be noted here that the description of the above computer program product embodiments is similar to that of the above method embodiments and has similar beneficial effects to those of the method embodiments. For the technical details not disclosed in the computer program product embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0233] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A voice interaction method, characterized in that: The method is applied to a voice customer service system, the voice customer service system includes multiple agents and a shared knowledge base, each agent corresponds to at least one large model and at least one skill knowledge base, and the shared knowledge base stores historical data of the user's interaction with the voice customer service system. The method includes: Receive voice information input by the user; If it is determined based on the voice information that an agent needs to be switched, a target agent is determined from other agents among the multiple agents except the current agent; Using the large model and skill knowledge base corresponding to the target intelligent agent and the shared knowledge base to process the content corresponding to the voice information and obtain a reply; Outputting a voice message corresponding to the reply to the user.

2. The method according to claim 1, characterized in that Before determining that the agent needs to be switched based on the voice information, the method further includes: If the voice information indicates that the agent needs to be switched, or there is no answer corresponding to the voice information in the skill knowledge base corresponding to the current agent, it is determined that the agent needs to be switched; If the voice information indicates that there is no need to switch the agent at present, or there is a response corresponding to the voice information in the skill knowledge base corresponding to the current agent, it is determined that there is no need to switch the agent at present.

3. The method according to claim 2, characterized in that If the voice information indicates that the agent needs to be switched, or before there is no answer corresponding to the voice information in the skill knowledge base corresponding to the current agent, the method further includes: If the voice information matches a preset keyword, and the preset keyword is used to indicate the other agent, then determining that the voice information indicates that an agent needs to be switched; and / or, If the semantics identified from the voice information indicates the other agent, determining that the voice information indicates a need to switch agents; and / or, If the intention identified using the large model from the voice information and the historical data indicates the other agent, it is determined that the voice information indicates the need to switch agents.

4. The method according to claim 3, characterized in that If the voice information matches a preset keyword, and the preset keyword is used to indicate the other agent, determining that the voice information indicates that an agent needs to be switched includes: If the voice information matches a preset keyword and the emotion intensity of the voice information is higher than the emotion intensity threshold, it is determined that the voice information indicates that the intelligent agent needs to be switched.

5. The method according to claim 1, characterized in that Each agent has a corresponding priority weight; the step of determining a target agent from other agents among the multiple agents except the current agent includes: Determine a plurality of candidate intelligent agents from the other intelligent agents based on the voice information; The candidate intelligent agent with the highest priority weight is selected from the multiple candidate intelligent agents and determined as the target intelligent agent.

6. The method according to claim 1, characterized in that The skill knowledge base includes a rule knowledge base and a content knowledge base. The rule knowledge base stores various keywords and the content corresponding to the keywords. The content knowledge base stores various knowledge referenced when providing a large model to generate a response. The adopting of the large model and skill knowledge base corresponding to the target agent and the shared knowledge base to process the content corresponding to the voice information and obtain a reply includes: Matching the content corresponding to the voice information with the keywords in the rule knowledge base, and determining the content corresponding to the successfully matched keywords as the reply; and / or, Identifying the semantics in the content corresponding to the voice information, matching the semantics with keywords in the rule knowledge base, and determining the content corresponding to the successfully matched keywords as the reply; and / or, Using a large model to identify the intent of the content corresponding to the historical data and the voice information, and matching the intent with keywords in the rule knowledge base, and determining the content corresponding to the successfully matched keywords as the reply; and / or, The content corresponding to the voice information and the historical data are input into the big model corresponding to the target intelligent agent, and the reply is generated based on the content knowledge base through the big model corresponding to the target intelligent agent.

7. The method according to any one of claims 1 to 6, characterized in that After receiving the voice information input by the user, the method further includes: The voice information is stored in the shared knowledge base, and a portrait of the user is generated based on the voice information and the historical data, and the portrait is stored in the shared knowledge base for easy invocation by the intelligent agent.

8. The method according to any one of claims 1 to 6, characterized in that The voice customer service system also includes a language adaptation model; before using the large model and skill knowledge base corresponding to the target agent and the shared knowledge base to process the voice information, the method also includes: Identifying the language type of the speech information through the language adaptation model, and determining a target large model from at least one large model corresponding to the target agent according to the language type; Converting the speech information into text information through the language adaptation model, and converting the text information into a semantic vector; The method of using the large model and skill knowledge base corresponding to the target agent and the shared knowledge base to process the voice information includes: The target macro model, the skill knowledge base corresponding to the target agent, and the shared knowledge base are used to process the semantic vector.

9. The method according to claim 8, characterized in that The outputting the speech corresponding to the reply to the user includes: The response is converted into speech according to the language habits corresponding to the language type through the language adaptation model, and the speech is output to the user.

10. A voice interaction device, characterized in that: The device is applied to a voice customer service system, the voice customer service system includes multiple agents and a shared knowledge base, each agent corresponds to at least one large model and at least one skill knowledge base, the shared knowledge base stores historical data of the user's interaction with the voice customer service system, and the device includes: A receiving module, used for receiving voice information input by a user; A determination module, configured to determine a target agent from other agents in the plurality of agents except the current agent if it is determined based on the voice information that an agent needs to be switched; A processing module, used to process the content corresponding to the voice information by using the large model and skill knowledge base corresponding to the target intelligent agent and the shared knowledge base to obtain a reply; An output module is used to output the voice corresponding to the reply to the user.

11. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 9.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Cited By

  • Intelligent agent abnormal input detection method and device based on large model and electronic equipment

    CN120706568A

  • Data processing method and system based on AI question and answer robot

    CN121092677A

  • Intelligent agent control method and system based on smart screen scene recognition, smart television and storage medium

    CN121126033A

  • Dialogue generation method and device, electronic equipment, storage medium and product

    CN121388110A