Voice interaction method and device, computer program product and electronic equipment
Through the purpose recognition result-driven private application calls and internal application calls, combined with the agent to generate reply content, the problem that voice interaction devices cannot meet special service scenarios is solved, and a more comprehensive and personalized voice interaction service is achieved.
Patent Information
- Application Number
- CN202510480008.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-18
AI Technical Summary
Existing voice interaction devices cannot provide voice interaction services that meet special service scenarios, resulting in incomplete response content and poor accuracy.
The privatized application call or internal application call is determined through the intent identification results, and the reply content is generated by the agent, providing adaptation and access to the privatized application, realizing personalized interaction in special service scenarios.
It improves the comprehensiveness, accuracy and flexibility of voice interaction, expands the scope of application, enhances the compatibility between privatized applications and internal applications, and reduces system maintenance and operation costs.
Smart Images

Figure CN120340484A_ABST
Abstract
Description
Background Art
[0002] In the related art, a voice interaction device can only generate general response content depending on a preset general list, and cannot provide voice interaction services for special service scenarios (such as diet, medical services, etc.). As a result, the response content generated by the voice interaction device is incomplete, the voice interaction has limitations, and the accuracy is poor. Summary of the Invention
[0003] The purpose of the present disclosure is to provide a voice interaction method, a voice interaction device, a computer program product, and an electronic device, so as to at least to some extent overcome the problems of incomplete generated response content and limitations of voice interaction caused by the limitations and defects of the related art.
[0004] According to one aspect of the present disclosure, a voice interaction method is provided, including: responding to an interaction request of a target user acting on a voice interaction device, and obtaining query information of the target user; performing intent recognition on the query information to determine an intent recognition result, and according to the intent recognition result, determining to perform a privatized application call or an internal application call on the query information, and determining a target application according to the intent recognition result; determining a call path corresponding to the target application, and generating content based on the call path for the intent recognition result to obtain a response content to the query information; returning the response content to the voice interaction device, so that the voice interaction device outputs the response content.
[0005] In an exemplary embodiment of the present disclosure, the performing a privatized application call or an internal application call on the query information according to the intent recognition result includes: if the intent in the intent recognition result requires a privatized application call, performing a privatized application call on the query information; if the intent in the intent recognition result requires an internal application call, performing an internal application call on the query information.
[0006] In an exemplary embodiment of the present disclosure, the determining a call path corresponding to the target application, and generating content based on the call path for the intent recognition result to obtain a response content to the query information includes: if the target application is a privatized application, forwarding the intent recognition result to a unified interface of an application server, and controlling the unified interface to send the intent recognition result to a distribution decision center module; based on the distribution decision center module, distributing and routing the intent recognition result to an agent of the target application, so that the agent generates a response content to the query information.
[0007] In an exemplary embodiment of the present disclosure, generating the response content for the query information includes: obtaining the input parameters required to invoke the agent; and controlling the agent to generate the response content corresponding to the query information based on the input parameters, the query information, the context content, and the attribute information of the target application.
[0008] In an exemplary embodiment of the present disclosure, determining the call path corresponding to the target application and generating the content for the intent recognition result based on the call path to obtain the response content for the query information includes: if the target application is an internal application, forwarding the intent recognition result to the unified interface of the application server, and controlling the unified interface to send the target application to the access layer; the access layer sends the intent recognition result to the central control module, and determines the response content for the query information through the central control module and the vertical service module.
[0009] In an exemplary embodiment of the present disclosure, the response content includes the broadcast content and the display content; returning the response content to the voice interaction device includes: performing content recognition on the response content, and recognizing the display content in the response content as structured information; sending the structured information to the voice interaction device, keeping the broadcast content other than the structured information unchanged, and transmitting the broadcast content to the voice interaction device.
[0010] In an exemplary embodiment of the present disclosure, returning the response content to the voice interaction device includes: if the target application is a privatized application, according to the call policy configured in the access layer, when the response content returned by the privatized application is the first identifier, returning the response content of the privatized application to the voice interaction device; when the response content returned by the privatized application is the second identifier, returning the response content of the internal application to the voice interaction device.
[0011] According to one aspect of the present disclosure, there is provided a voice interaction device, including: a query information acquisition module, configured to respond to an interaction request of a target user acting on a voice interaction device, and acquire the query information of the target user; a target application recognition module, configured to perform intent recognition on the query information to determine an intent recognition result, determine to perform a privatized application call or an internal application call on the query information according to the intent recognition result, and determine a target application according to the intent recognition result; a response content generation module, configured to determine the call path corresponding to the target application, and generate the content for the intent recognition result based on the call path to obtain the response content for the query information; and a response content return module, configured to return the response content to the voice interaction device so that the voice interaction device outputs the response content.
[0012] According to one aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the voice interaction method described in any one of the above.
[0013] According to one aspect of the present disclosure, there is provided an electronic device including: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the voice interaction method described in any one of the above by executing the executable instructions.
[0014] In the technical solution provided in the embodiments of the present disclosure, on the one hand, since it is determined to perform a privatized application call or an internal application call on the query information according to the intent recognition result, and the target application is determined according to the intent recognition result; then the content generation is performed on the intent recognition result according to the call path corresponding to the target application to obtain the reply content of the query information. Among them, two different types of calls, namely privatized application call and internal application call, are provided, which can realize the adaptation and access of privatized applications, realize the convenient expansion of applications, can provide reply content matching the service scenario for special service scenarios, increase the application scope of voice interaction, improve the flexibility and comprehensiveness of the generated reply content, and also improve the compatibility between privatized applications and internal applications. On the other hand, it is possible to generate content based on the intent recognition result according to the call path, so as to obtain the target application of the query information, and then it is possible to call the appropriate module based on the call path to generate the reply content of the target application, improve the comprehensiveness, coherence and accuracy of the reply content, be able to provide more complex and comprehensive voice interaction, improve the accuracy and experience of voice interaction, realize personalized interaction and customized interaction in more special service scenarios, and increase the application scope.
[0015] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0017] Figure 1 Schematically showing a flowchart of a voice interaction method in an embodiment of the present disclosure.
[0018] Figure 2 Schematically showing a schematic diagram of a system architecture for performing voice interaction in an embodiment of the present disclosure.
[0019] Figure 3 A flowchart showing the process of generating a response content when the target application is a privatized application in an embodiment of the present disclosure is schematically shown.
[0020] Figure 4 A flowchart showing the process of reserving a shuttle bus in an embodiment of the present disclosure is schematically shown.
[0021] Figure 5 A block diagram showing a voice interaction device in an embodiment of the present disclosure is schematically shown.
[0022] Figure 6 A block diagram showing an electronic device in an embodiment of the present disclosure is schematically shown. Detailed implementation manners
[0023] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be used. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring the various aspects of the present disclosure.
[0024] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0025] Voice interaction devices generally include two forms, forming two major categories of solutions: FAQ (Frequently Asked Questions) - style conversations and task - style conversations.
[0026] The FAQ-style dialogue is an interactive model based on pre-set question-and-answer pairs. In this solution, the system processes the user's questions through a pre-prepared list of common questions and corresponding answers. After the user asks a question, the system attempts to find the most matching question-and-answer pair in the list and returns the relevant answer. Its main drawbacks are as follows: Dependence on pre-set questions and answers: The system can only handle pre-defined questions. If the user's question is not in the FAQ list, the system may not be able to give an appropriate answer. Lack of flexibility: For variants of a question or complex multi-level questions, the FAQ-style solution may not be able to accurately match and provide effective answers. Difficulty in expansion: Adding new question-and-answer pairs requires manual maintenance. As the number of question-and-answer pairs increases, management becomes complex. Insufficient context understanding: The FAQ-style solution usually does not understand the dialogue context and cannot provide appropriate answers based on previous interactions.
[0027] The generative text summary of the task-style dialogue can generate new text that does not exist in the original document. However, there are the following problems: Due to the pre-defined set of processes and intents, the task-style dialogue system performs poorly in handling complex and non-predefined questions. Once the user's needs exceed the preset scope, the system may not be able to respond effectively, and the flexibility is limited. The task-style dialogue system sometimes cannot accurately understand the context of the dialogue. If the amount of information in the dialogue process is too large or too complex, the system may not be able to effectively track the context and intent. The system's dialogue is usually relatively mechanical and restricted, lacking the ability of natural human interaction. The user experience may not be as natural and smooth as interpersonal dialogue. To maintain the effectiveness of the system, frequent maintenance and updates are required, including adding new intent recognition, changing the dialogue process, etc., which increases the development and operation costs. In the related art, for special service scenarios, the voice interaction device cannot provide voice interaction that meets the special service scenarios, which has certain limitations.
[0028] To solve the above technical problems, in the embodiments of the present disclosure, a voice interaction method is provided, which can be used in application scenarios for interacting with a voice interaction device.
[0029] Next, refer to Figure 1 for a specific description of the voice interaction method in the embodiments of the present disclosure.
[0030] In step S110, in response to an interaction request of a target user acting on the voice interaction device, query information of the target user is obtained.
[0031] In the embodiments of the present disclosure, the target user can be any type of user. For example, it can be a user in a certain retirement community, or a user in a certain work park, etc. Here, a user in a retirement community is taken as an example for illustration.
[0032] The voice interaction device can be any device capable of voice interaction, such as a smart speaker, a smart TV, a smart watch, a smart phone, a robot, etc. The interaction request can be a voice request directly sent by the target user to the voice interaction device, or a voice request input by the target user on an application installed on the mobile terminal for controlling the voice interaction device, or a text request input, which is not specifically limited herein.
[0033] After the voice interaction device receives the interaction request, it can further obtain the query information input by the target user. The query information can be voice query information directly sent by the target user to the voice interaction device, or voice query information input by the target user on an application installed on the mobile terminal for controlling the voice interaction device or text query information input. Exemplarily, the voice recognition ability of the cloud can be used to obtain the query information of the target user. The voice recognition ability can be implemented through ASR (Automatic Speech Recognition). The application can be, for example, a separate application, or a small program embedded in an instant messaging application, or a public account or other forms, which is not specifically limited herein.
[0034] In step S120, perform intent recognition on the query information to determine the intent recognition result, determine to perform a private application call or an internal application call on the query information according to the intent recognition result, and determine the target application according to the intent recognition result.
[0035] In the embodiments of the present disclosure, different types of application calls can be configured for the voice interaction device. The application calls of different paths can include private application calls or internal application calls. Among them, the private application call refers to calling a private service additionally configured on the voice interaction device. The internal application call refers to calling a general service that the voice interaction device itself has.
[0036] First, intent recognition can be performed on the query information to obtain the intent recognition result. Further, the call type of the query information can be determined according to the type of the intent recognition result. Among them, the query information can be converted into text information, natural language processing can be performed on the text information to generate the initial intent result of the target user; the context intent of the target user can be determined by combining the multi-round context information of the query information, and the initial intent can be corrected according to the context intent to obtain the intent recognition result. If there is no context information, the initial intent is directly used as the intent recognition result.
[0037] The intention recognition result may include an intention and key information. The intention can represent the desired goal and can be used to determine the application to be finally invoked. The key information is used to represent the specific content to be achieved in the target application. Based on this, the type of the intention recognition result can be determined according to the type of application to be invoked according to the intention in the intention recognition result. For example, if the intention in the intention recognition result needs to invoke a privatized application, the query information is invoked for the privatized application. If the intention in the intention recognition result needs to invoke an internal application, the query information is invoked for the internal application.
[0038] In some embodiments, the application server may pre-configure multiple privatized applications in advance. The privatized applications can be different types of customized services. The types of the privatized applications can be determined according to the user characteristics and scenario characteristics in the scenario where the voice interaction device is used. For example, when the scenario where the voice interaction device is used is a retirement community, the privatized applications may include shuttle buses, catering, medical treatment, etc. Based on this, the voice interaction device can implement any one or more of asking questions based on the information within the retirement community, asking questions based on the knowledge base, multi-turn voice interaction based on privatized applications, interacting with the displayed content, asking questions based on external Internet information, asking questions based on poster picture knowledge, and asking questions based on video recognition. For example, the multi-turn voice interaction based on privatized applications can be shuttle bus reservation, medication management, restaurant reservation, etc. Interacting with the displayed content can be purchasing the content displayed on the screen of the intelligent interaction device or other devices connected to the intelligent interaction device.
[0039] In addition, the voice interaction device also includes internal applications it has. The internal applications can include built-in conversation capabilities such as weather query, music playback, alarm setting, smart home control, and FAQ answering.
[0040] For each configured privatized application and each internal application, an intention mapping table can be established to clarify the intentions and key information supported by each application. The key information can be represented by keywords. After determining the intention recognition result of the query information, the target application to be invoked for the query information can be determined according to the intention in the intention recognition result. Exemplarily, the intention in the intention recognition result of the target user is matched with the intention mapping tables of the privatized applications and the internal applications. Calculate the similarity between the intention in the intention recognition result and the intention of each privatized application. When the similarity is greater than the similarity threshold, it can be considered that the intention recognition result is for invoking a privatized application; when the similarity between the intention in the intention recognition result and the intention of each internal application is greater than the similarity threshold, it can be considered that the intention recognition result is for invoking an internal application.
[0041] Exemplarily, a call policy can be configured in the access layer of the cloud. The call policy can be to first execute the call of the privatized application, and then execute the call of the internal application when the call of the privatized application cannot be executed. That is, the priority of calling the privatized application is higher than that of calling the internal application. Based on this, when the reply content returned by the privatized application is the first identifier, the reply content of the privatized application is executed; when the reply content returned by the privatized application is the second identifier, that is, after the privatized application clearly returns an identifier indicating that it cannot be processed, the reply content of the internal application is executed. Among them, the first identifier can be data.confidence = 100, and the second identifier can be data.confidence = 0.
[0042] Furthermore, when the similarity between the intent in the intent recognition result and the intent of a certain privatized application is greater than the similarity threshold, the privatized application with a similarity greater than the similarity threshold to the intent in the intent recognition result can be determined as the target application. When the similarity between the intent in the intent recognition result and the intent of a certain internal application is greater than the similarity threshold, the internal application with a similarity greater than the similarity threshold to the intent in the intent recognition result can be determined as the target application. If neither the privatized application nor the internal application is hit, the target user can be reminded to re-enter the query information.
[0043] Next, continue to refer to Figure 1 As shown in, in step S130, determine the call path corresponding to the target application, and generate content for the intent recognition result based on the call path to obtain the reply content of the query information.
[0044] In the embodiments of the present disclosure, after determining the intent recognition result and the target application, the call path of the target application can be determined. It should be noted that for privatized applications and internal applications, their corresponding call paths are completely different.
[0045] Figure 2 Schematically shows the system architecture diagram on which the voice interaction method depends. Refer to Figure 2 As shown in, the system structure can include a voice interaction device end, a cloud, and an application server end. Among them, the interaction device end can be a voice interaction device or an application program for controlling the voice interaction device. The cloud can be the cloud corresponding to the voice interaction device, which can include an identification module, an access layer, a vertical service module, and a central control module of the voice interaction device. The application server end can be an additional configured server end, which can include a unified interface, an intent classification module, and a distribution decision center module. In addition, it can also include an application service module. The application service module can include multiple agents, such as a shuttle bus agent, a catering agent, and so on.
[0046] In some embodiments, the recognition module is used for ASR recognition. The access layer makes two calls asynchronously. The two calls respectively return their respective intent recognition results and the content of the next reply. The content of the next reply can also be the content of the next reply, that is, the reply content output by the voice interaction device for the query information. A call policy can be configured in the access layer. The call policy can be, for example: give priority to the result returned by calling the privatized application. Only when the privatized application clearly returns an unprocessable identifier (such as data.confidence = 0), then execute according to the reply content of calling the internal application. By flexibly configuring the execution priority and fallback policy in the access layer policy, seamless cooperation between the privatized application and the internal application is ensured, and the network transmission performance is optimized.
[0047] Unified interface: Provides an interface for the access layer and stipulates the return format. The bottom layer of the unified interface is the intent classification module and the distribution decision center module. The main purpose is to avoid the separate interaction between the intent classification module, the distribution decision center module and the access layer, saving network transmission time and optimizing performance.
[0048] The intent classification module is used to determine the intent recognition result of the query information and judge whether the intent recognition result of the target user hits the privatized application supported by the voice interaction device. The intent classification module has the ability of multi-round semantic recognition and can eliminate the interference caused by factors such as semantic drift and ASR typos. The intent classification module can be responsible for a large language model that has been prompt-engineered or instruction-tuned. The parameter scale of the intent classification model can be comprehensively considered according to the computing power and performance requirements. In addition, retrieval and recall and other strategies can be combined to further improve the accuracy. Through the intent classification module, the accuracy and robustness are improved.
[0049] The distribution decision center module can be used to distribute and route the intent recognition results, and isolate the session states and maintenance policies of different users and the same user on different dates. Exemplarily, it can support maintaining the session state and historical maintenance policies according to the device identifier and application identifier, ensuring the isolation of session histories between different devices and applications. Specifically, the refresh frequency of the database can be configured in the distribution decision center module to achieve the isolation of session histories, and the refresh frequency can vary according to different users. In addition, the distribution decision center module can also process the reply content returned by the privatized application.
[0050] Based on the above architecture, the distribution decision center module can distribute the intent recognition result to the agent corresponding to the target application in the application service module, such as distributing it to the shuttle bus agent or the catering agent, etc., so that the agent corresponding to the target application can generate a reply content corresponding to the query information based on the execution of the target application. By distributing the intent recognition result to the agent corresponding to the target application through the distribution decision center module, it can ensure the accurate invocation and execution of the target application, and can provide isolation of sessions and data between different devices and different applications, avoiding information interleaving and data pollution.
[0051] Based on the above system architecture, the call path corresponding to the target application can be determined. The call path refers to the order chain in which modules and components are called during the control process of the voice interaction device. In some embodiments, if the intent recognition result is a private application call and the target application is a private application, referring to Figure 3 as shown in, its call path can be: sending the intent recognition result to the unified interface, sending it to the distribution decision center module through the unified interface, and distributing the intent recognition result to the agent corresponding to the target application through the distribution decision center module, so that the agent corresponding to the target application can generate the reply content of the target application.
[0052] In other embodiments, if the intent recognition result is an internal application call and the target application is an internal application, its call path can be: sending the intent recognition result to the unified interface, sending it to the access layer through the unified interface, distributing the intent recognition result to the central control module of the voice interaction device through the access layer, and then sending it to the vertical service module of the voice interaction device based on the central control module to call the vertical service to generate the reply content of the target application.
[0053] For the target application being a private application, the reply content of the target application can be a private reply content. For the target application being an internal application, the reply content of the target application can be an internal reply content.
[0054] The reply content of the target application can include a broadcast content and a display content. The broadcast content refers to the content played by voice through the voice interaction device, and the display content refers to the picture displayed through the display screen of the voice interaction device, such as a video or an image, etc.
[0055] In some embodiments, the input parameters required for the agent to call the target application can be obtained; through the input parameters, the query information, the context content, and the attribute information of the target application, the agent is controlled to generate a reply content corresponding to the query information.
[0056] Exemplarily, the input parameters can be user names, dates, and so on. The input parameters can be parsed to extract the parsing results. For example, keywords or specific instructions can be extracted from the input parameters. Entities related to the intent in the intent recognition result can be extracted. The entity can be the object targeted by the intent. The context content can include the historical record before the query information, the context conversation, user preferences, the session state, etc. Among them, the user preferences can be determined according to the user information of the target user and the historical behavior of the target user. The user information can be the basic situation of the user. The historical behavior can be the behavior that matches the current intent recognition result. The attribute information of the target application can include the application content of the target application, or the application features and configuration. The intelligent agent generates content that meets the user preferences, is executed by the target application, and conforms to the query information for each target user according to the query information, the context content, and the attribute information of the target application. For example, the generated broadcast content and display content can be generated. The broadcast content can be the voice reply information for the query information, and the display content can be the picture for the query information. Further, the intelligent agent can further adjust the generated reply content according to the feedback information of the target user.
[0057] In some other embodiments, the vertical service can be called to generate the reply content of the target application according to the intent recognition result and the context content of the target user.
[0058] When the target application is an internal application or a privatized application, the returned reply content can include both broadcast content and display content. For a privatized application, the distribution decision center module can identify the reply content returned by the privatized application and identify the display content in the reply content as structured information. The structured information can be JSON information. The distribution decision center module can match the broadcast content and the structured information to the specified fields to realize the association between the content and the display through field mapping.
[0059] In step S140, the reply content is returned to the voice interaction device so that the voice interaction device outputs the reply content.
[0060] In the embodiments of the present disclosure, for different target applications, the return path for transmitting the reply content returned by the target application to the voice interaction device is also different, and the return path can correspond to the call path. For example, as shown in Figure 3 When the target application is a privatized application, the return path can be: the reply content is returned to the distribution decision center module, the distribution decision center module returns the reply content to the unified interface, and the unified interface sends it to the access layer, and then the reply content is transmitted to the voice interaction device through the access layer for output.
[0061] When the target application is an internal application, the return path can be as follows: The reply content is returned to the central control module of the voice interaction device through the vertical service module, further sent to the access layer, and then transmitted to the voice interaction device through the access layer for output.
[0062] Specifically, according to the call policy configured in the access layer, the reply content returned by the privatized application and the internal application can be processed in the access layer to determine the reply content to be returned to the voice interaction device.
[0063] Exemplarily, when transmitting the reply content to the voice interaction device through the access layer to enable the voice interaction device to output, according to the call policy included in the access layer, when the reply content returned by the privatized application is the first identifier, the reply content of the privatized application is returned to the voice interaction device; when the reply content returned by the privatized application is the second identifier, the reply content of the internal application is returned to the voice interaction device. It should be noted that the first identifier can be data.confidence = 100, and the second identifier can be data.confidence = 0. That is, when the privatized application clearly returns an unprocessable identifier, the reply content of the internal application can continue to be output.
[0064] It should be noted that when transmitting the reply content returned by the privatized application to the voice interaction device, the broadcast content other than the display content in the reply content returned by the privatized application can be kept unchanged, and the structured information and the broadcast content are sent to the voice interaction device. When sending the structured information and the broadcast content to the voice interaction device, the data block size transmitted to the voice interaction device each time can be determined, and the structured information and the broadcast content are streamed based on the determined data block size. Among them, the data block size CHUNK SIZE can be dynamically adjusted according to the size of the structured information and the broadcast content and the network status, and controlling the data block size can optimize the data transmission efficiency and reduce latency.
[0065] In the embodiments of the present disclosure, by providing two call paths for privatized application calls and internal application calls, and performing priority control through the call policy of the access layer. In the configuration policy, the privatized application is preferentially called to process the requests of the privatized application. When the privatized application call path is clearly unable to process, the internal application call path is then called to process the requests of the internal application. This not only ensures the priority processing of the privatized application but also ensures that the reply content of the internal application is not missed when the privatized application cannot be processed, thus realizing the compatibility of the two application systems.
[0066] After the voice interaction device receives the returned reply content, it can wake up the voice interaction device to output the returned reply content. For example, it can output the reply content returned by the privatized application or the reply content returned by the internal application. For example, it wakes up the application installed on the user terminal to perform customized operations such as jumping, displaying, broadcasting, and reminding for the voice interaction. For example, when the target application is shuttle reservation, it can play the broadcast content of shuttle reservation and display the reservation interface.
[0067] In the embodiments of the present disclosure, by providing two call paths for privatized application invocation and internal application invocation, the voice interaction device can be effectively connected to and adapted to various privatized applications, avoiding the problem in the related art that it is impossible to provide matching voice interaction services according to special service scenarios, improving the comprehensiveness and accuracy of the reply content of the voice interaction, and increasing the application scope of the voice interaction. Through the access layer, the call policy configuration for privatized application invocation and internal application invocation is carried out, realizing priority control, ensuring the priority processing of privatized applications, and also ensuring that when the privatized applications cannot be processed, the execution reply content of the built-in applications will not be missed, thus realizing the compatibility of different types of applications. By providing a standard unified interface and a prescribed return format, the separate interaction between the intent classification module and the distribution decision center module and the access layer is avoided, saving network transmission time and improving the overall performance. The intent classification module improves the accuracy of intent classification through multi-round semantic recognition ability, avoiding the influence of ASR typos and semantic drift, thus reducing the error handling and response time. The distribution decision center module distributes and routes the intent recognition result to the corresponding agent when the privatized application is hit, and also ensures the data isolation between the device ID and the application ID to provide a smooth session history and status management function. In addition, it can also post-process the reply content returned by each application to accurately display the content shown on the screen and the broadcast content. The agent can generate accurate reply content for the target user by combining dimensions such as user preferences, context information, intent recognition results, and attribute information of the target application. Since different applications can be implemented through the agent and can be extended according to actual needs, the process of frequent maintenance and update is avoided, which is convenient for management and reduces the system development and operation costs.
[0068] Based on this, the above voice interaction method can implement general schedule management and reservation management capabilities, support vital sign data collection / modification, etc.; for dish consultation, new prices, allergen ingredients, etc. are added, and it can automatically recommend dining restaurants according to the dishes without the user having to search by themselves. In addition, it can build a user social circle, for example, recommend alumni associations, etc.; realize the recommendation of value-added service products, etc., with strong scalability and improved user experience.
[0069] Next, take the privatized application for reserving a shuttle bus as an example for illustration. First, the target user wakes up the voice interaction device through the application program and identification module installed on the user terminal, and expresses the need to go out and reserve a shuttle bus through voice.
[0070] The query information corresponding to the demand flows to the joining layer, is transmitted to the unified interface through the access layer, and can correctly identify the multi-round intention recognition result through the intention classification module. And the intention recognition result is sent to the distribution decision center module through the unified interface. After the distribution decision center module sends the intention recognition result to the shuttle agent, the shuttle agent collects the input parameters required to call the shuttle agent, and then calls the corresponding interface. The shuttle agent processes the query information, and replies to the target user's query information through the large model dialogue ability and relevant tool abilities to generate the reply content.
[0071] The distribution decision center module recognizes the display content in the reply content returned by the shuttle agent as a JSON structure, keeps the broadcast content in the reply content unchanged, and matches the broadcast content and the JSON structure displayed on the control screen to the specified fields, so as to be transmitted to the voice interaction device through the access layer for output.
[0072] Figure 4 The interaction flow chart of reserving a shuttle bus is schematically shown in Figure 4 As shown in, when receiving the query information, the first reply content played can be whether to take the shuttle bus. If the target user further replies yes, they can ask about the time of taking the shuttle bus, and can query the shuttle bus information through the unified interface. The broadcast content can be whether the route is appropriate and select the next stop. At the same time, the information about leaving the park can be updated. If not, they can ask about the time of leaving the park. In the case of taking the shuttle bus, if the selected is a loop shuttle bus, the generated broadcast content can be whether to take the same shuttle bus back. If yes, the information about returning to the park is updated. If not, ask about the time of returning to the park. In the case of not taking the shuttle bus, they can ask about the time of returning to the park and update the information about returning to the park. Further, they can ask about the reason for going out and the number of accompanying people, and update the information about going out. Further play the overall confirmation content and update the confirmation information.
[0073] For example, based on the above interaction flow chart, during the process of reserving a shuttle bus, the interaction process between the target user and the voice interaction device can be as follows:
[0074] User: Xiao A, I want to go out.
[0075] Xiao A: Here, please wait a moment. May I ask when you plan to go out? Specific to the date and time.
[0076] User: Go out at 12 o'clock tomorrow noon.
[0077] Xiao A: May I ask if you need to take the shuttle bus when you leave the park?
[0078] User: Yes, I do.
[0079] Xiaoming: Okay. You need to reserve the shuttle bus out of the park at 12:00 noon tomorrow. Please wait a moment while I check the available shuttle bus information for you.
[0080] There is a pharmacy pick-up shuttle bus tomorrow. It will depart from the North Hall of the community at 12:20 noon tomorrow, pass by Hospital B and Hospital C, and there are currently 24 seats available. May I ask if this shuttle bus meets your needs?
[0081] User: Yes, it does.
[0082] Xiaoming: Okay. You have selected the pharmacy pick-up shuttle bus at 12:20 noon tomorrow. Next, please select the stop where you want to get off: Hospital B or Hospital C?
[0083] User: Hospital B.
[0084] Xiaoming: Okay. You have selected to get off at Hospital B. May I ask when you expect to return? Please specify the date and time.
[0085] User: The day after tomorrow at 5:00 pm.
[0086] Xiaoming: You plan to return at 5:00 pm on Thursday. Do you need to take the shuttle bus when you return?
[0087] User: No, I don't.
[0088] Xiaoming: Okay. Do you have an accompanying person when you go out? If so, please let me know the number of accompanying people.
[0089] User: 1.
[0090] Xiaoming: Okay. You have one accompanying person. What is the reason for your going out this time? For example, going home, traveling within the country, seeking medical treatment, shopping, etc.
[0091] User: Going home.
[0092] Xiaoming: Okay. The reason for your going out that you have selected is "going home". Next, please confirm the reservation information:
[0093] You plan to go out at 12:00 noon tomorrow, that is, Wednesday, take the outbound shuttle bus at 12:20 noon, and get off at Hospital B. The return time is 5:00 pm the day after tomorrow, and you don't need to take the shuttle bus. You have one accompanying person, and the reason for going out is "going home".
[0094] Is the above information correct?
[0095] User: Correct.
[0096] Xiaoming: Okay, your reservation information has been confirmed and recorded. Have a great trip! If you have any questions, feel free to come to me anytime.
[0097] In the embodiments of the present disclosure, by providing an agent, accurate response content can be generated for each target application, improving the coherence and accuracy of the response content, enabling the adaptation and access of privatized applications, realizing the convenient expansion of applications, improving flexibility, and also improving the compatibility between privatized applications and internal applications.
[0098] In the embodiments of the present disclosure, a voice interaction device is provided. Referring to Figure 5 as shown in, the voice interaction device 500 mainly includes the following modules:
[0099] A query information acquisition module 501, configured to respond to an interaction request of a target user acting on a voice interaction device, and acquire query information of the target user;
[0100] A target application recognition module 502, configured to perform intent recognition on the query information to determine an intent recognition result, determine to perform a privatized application call or an internal application call on the query information according to the intent recognition result, and determine a target application according to the intent recognition result;
[0101] A response content generation module 503, configured to determine a call path corresponding to the target application, and generate content based on the call path for the intent recognition result to obtain a response content to the query information;
[0102] A response content return module 504, configured to return the response content to the voice interaction device, so that the voice interaction device outputs the response content.
[0103] In an exemplary embodiment of the present disclosure, the performing a privatized application call or an internal application call on the query information according to the intent recognition result includes: if the intent in the intent recognition result requires a privatized application call, performing a privatized application call on the query information; if the intent in the intent recognition result requires an internal application call, performing an internal application call on the query information.
[0104] In an exemplary embodiment of the present disclosure, determining the call path corresponding to the target application and generating content based on the call path for the intent recognition result to obtain the reply content of the query information includes: if the target application is a privatized application, forwarding the intent recognition result to the unified interface of the application server, and controlling the unified interface to send the intent recognition result to the distribution decision center module; based on the distribution decision center module, distributing and routing the intent recognition result to the agent of the target application so that the agent generates the reply content of the query information.
[0105] In an exemplary embodiment of the present disclosure, generating the reply content of the query information includes: obtaining the input parameters required to call the agent; controlling the agent to generate the reply content corresponding to the query information through the input parameters, the intent recognition result, the context content, and the attribute information of the target application.
[0106] In an exemplary embodiment of the present disclosure, determining the call path corresponding to the target application and generating content based on the call path for the intent recognition result to obtain the reply content of the query information includes: if the target application is an internal application, forwarding the intent recognition result to the unified interface of the application server, and controlling the unified interface to send the target application to the access layer; the access layer sends the intent recognition result to the central control module, and determines the reply content of the query information through the central control module and the vertical service module.
[0107] In an exemplary embodiment of the present disclosure, the reply content includes a broadcast content and a display content; returning the reply content to the voice interaction device includes: performing content recognition on the reply content, and recognizing the display content in the reply content as structured information; sending the structured information to the voice interaction device, and keeping the broadcast content other than the structured information unchanged, and transmitting the broadcast content to the voice interaction device.
[0108] In an exemplary embodiment of the present disclosure, returning the reply content to the voice interaction device includes: if the target application is a privatized application, according to the call policy configured in the access layer, when the reply content returned by the privatized application is the first identifier, returning the reply content of the privatized application to the voice interaction device; when the reply content returned by the privatized application is the second identifier, returning the reply content of the internal application to the voice interaction device.
[0109] It should be noted that the specific details of each module in the above voice interaction device have been described in detail in the corresponding method, so they will not be repeated here.
[0110] It should be noted that although several modules or units of the device for content execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-mentioned modules or units can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0111] In addition, although the steps of the method in the present disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be executed in that specific order, or that all the shown steps must be executed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0112] In an exemplary embodiment of the present disclosure, there is also provided an electronic device capable of implementing the above method.
[0113] Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, method, or program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.
[0114] The following refers to Figure 6 to describe the electronic device 600 according to this embodiment of the present disclosure. Figure 6 The shown electronic device 600 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0115] As Figure 6 shown, the electronic device 600 is presented in the form of a general-purpose computing device. The components of the electronic device 600 may include, but are not limited to: at least one of the above-mentioned processing units 610, at least one of the above-mentioned storage units 620, a bus 630 connecting different system components (including the storage unit 620 and the processing unit 610), and a display unit 640.
[0116] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 610, so that the processing unit 610 executes the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification. For example, the processing unit 610 can execute the steps as Figure 1 shown in
[0117] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 6201 and / or a cache storage unit 6202, and may further include a read-only storage unit (ROM) 6203.
[0118] The storage unit 620 may also include a program / utilities 6204 having a set (at least one) of program modules 6205. Such program modules 6205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.
[0119] The bus 630 may represent one or more of several types of bus structures, including a storage unit bus or storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of a variety of bus structures.
[0120] The electronic device 600 may also communicate with one or more external devices 700 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 600, and / or may communicate with any device that enables the electronic device 600 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be through an input / output (I / O) interface 650. And, the electronic device 600 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 660. As shown in the figure, the network adapter 660 communicates with other modules of the electronic device 600 through the bus 630. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0121] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or an electronic device, etc.) to execute the method according to the embodiments of the present disclosure.
[0122] It should be noted that in some embodiments of the present disclosure, a computer program product is also provided. The computer program product includes a computer program, and when the computer program is executed by a processor, the above-mentioned method is implemented.
[0123] In one implementation, the computer program product may be a tangible product containing a computer program, such as a computer-readable storage medium storing the computer program. The readable storage medium may be a storage medium based on signals such as electricity, magnetism, light, electromagnetic, infrared, etc., including but not limited to: random access memory (RAM), read-only memory (ROM), magnetic tape, floppy disk, flash memory (Flash), hard disk drive (HDD), solid state drive (SSD), and so on. Exemplarily, the computer program product may be implemented as a non-volatile storage medium storing the computer program, such as read-only memory, NAND flash memory, etc.
[0124] In one implementation, the computer program product may be an intangible product containing a computer program. Exemplarily, the computer program product may be implemented as a virtual digital product, such as an executable file storing the computer program, a digital file such as an installation package.
[0125] The code of the computer program can be written in one or more programming languages. Programming languages such as C language, Java, C++, etc. The program code can be executed entirely on the user's computing device, or partially on the user's computing device, or executed as an independent software package, or partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, such as a local area network (LAN), a wide area network (WAN), etc., or can be connected to an external computing device (for example, through an Internet connection provided by an operator).
[0126] The computer program can be carried or transmitted by signals such as electricity, magnetism, light, electromagnetic, infrared, etc. The electronic device can convert the signal carrying the computer program into a digital signal, and then run the computer program. When the computer program runs on the electronic device, its code is used to cause the electronic device to execute (more specifically, to cause the processor of the electronic device to execute) the method steps of various exemplary embodiments of the present disclosure.
[0127] From the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0128] After considering the specification and practicing the content disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed herein. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0129] It should be understood that the present disclosure is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A voice interaction method, characterized in that, Including: In response to an interaction request from a target user acting on a voice interaction device, obtain the query information of the target user; Perform intent recognition on the query information to determine the intent recognition result, determine to perform a privatized application call or an internal application call on the query information according to the intent recognition result, and determine the target application according to the intent recognition result; Determine the call path corresponding to the target application, and generate content based on the call path for the intent recognition result to obtain the reply content of the query information; Return the reply content to the voice interaction device so that the voice interaction device outputs the reply content.
2. The voice interaction method according to claim 1, wherein The performing a privatized application call or an internal application call on the query information according to the intent recognition result includes: If the intent in the intent recognition result requires a privatized application call, perform a privatized application call on the query information; If the intent in the intent recognition result requires an internal application call, perform an internal application call on the query information.
3. The voice interaction method according to claim 1, wherein The determining the call path corresponding to the target application and generating content based on the call path for the intent recognition result to obtain the reply content of the query information includes: If the target application is a privatized application, forward the intent recognition result to the unified interface of the application server, and control the unified interface to send the intent recognition result to the distribution decision center module; Based on the distribution decision center module, distribute and route the intent recognition result to the agent of the target application so that the agent generates the reply content of the query information.
4. The voice interaction method according to claim 3, wherein The generating the reply content of the query information includes: Obtain the input parameters required to call the agent; Control the agent to generate a reply content corresponding to the query information through the input parameters, the query information, the context content, and the attribute information of the target application.
5. The voice interaction method according to claim 1, wherein The determining the call path corresponding to the target application and generating content based on the call path for the intent recognition result to obtain the reply content of the query information includes: If the target application is an internal application, forward the intent recognition result to the unified interface of the application server, and control the unified interface to send the target application to the access layer; The access layer sends the intent recognition result to the central control module, and determines the reply content of the query information through the central control module and the vertical service module.
6. The voice interaction method according to claim 1, wherein The reply content includes a broadcast content and a display content; the returning the reply content to the voice interaction device includes: Perform content recognition on the reply content, and recognize the display content in the reply content as structured information; Send the structured information to the voice interaction device, keep the broadcast content other than the structured information unchanged, and transmit the broadcast content to the voice interaction device.
7. The voice interaction method according to claim 1, wherein The returning the reply content to the voice interaction device includes: If the target application is a privatized application, according to the call policy configured in the access layer, when the reply content returned by the privatized application is the first identifier, return the reply content of the privatized application to the voice interaction device; When the reply content returned by the privatized application is the second identifier, return the reply content of the internal application.
8. A voice interaction device, characterized in that, including: A query information acquisition module, configured to respond to an interaction request of a target user on a voice interaction device and acquire query information of the target user; A target application identification module, configured to perform intent recognition on the query information to determine an intent recognition result, determine to call a privatized application or an internal application for the query information according to the intent recognition result, and determine a target application according to the intent recognition result; A reply content generation module, configured to determine a call path corresponding to the target application, and generate content based on the intent recognition result based on the call path to obtain a reply content to the query information; A reply content return module, configured to return the reply content to the voice interaction device so that the voice interaction device outputs the reply content.
9. A computer program product, characterized in that, Including a computer program, characterized in that when the computer program is executed by a processor, it implements the voice interaction method according to any one of claims 1-7.
10. An electronic device, characterized in that, including: A processor; and A memory, configured to store executable instructions of the processor; wherein, the processor is configured to execute the voice interaction method according to any one of claims 1-7 by executing the executable instructions.
Citation Information
Patent Citations
Human-computer interaction method and system, electronic equipment and storage medium
CN117891922A
Display device and task processing method based on large language model
CN118586501A
Interactive method and device of robot, and device
US20200005772A1
Methods and systems for sharing private data
US20240346175A1
Cited By
Voice interaction-oriented multi-agent task cooperative processing system and processing method
CN120932652A