Voice interaction method and apparatus
By identifying and matching user intentions with services in the on-board voice system, the problem of multi-wheel dialogue command processing in the whole vehicle and full-time and multi-person interactive environment is solved, and the efficient distribution and execution of voice intentions is achieved, and interactive flexibility is improved.
Patent Information
- Application Number
- PCT/CN2024/140120
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-12-17
- Publication Date
- 2025-06-26
AI Technical Summary
In the vehicle-wide voice environment with full-time and multi-person interaction, the existing technology is difficult to effectively process multi-wheel dialogue instructions, resulting in the next multi-wheel dialogue voice instructions after processing one multi-wheel dialogue voice instructions, and it is easy to confuse different business executors and cannot correctly distribute voice instructions in multiple multi-wheel dialogues.
By identifying user intent, matching instructions with user intent, determining the target service, and executing instructions using the target service. Upon receiving the second instruction, the adapted target service is determined by matching the second instruction with the received user intention, thereby automatically executing the second instruction.
It realizes voice intention distribution in multiple rounds of dialogue scenarios, improves the flexibility of human-computer interaction, and ensures the correct distribution and execution of multiple voice commands.
Smart Images

Figure CN2024140120_26062025_PF_FP_ABST
Abstract
Description
Voice interaction method and device
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on December 19, 2023, with application number 202311756133.7 and application name “A Voice Interaction Method and Device”. The entire contents of the above Chinese patent application are incorporated by reference into this application. Technical Field
[0003] The present application relates to the field of data processing, and in particular to a voice interaction method and device. Background Art
[0004] Current in-vehicle voice interaction has evolved to enable full-car, all-time, multi-person interaction. This means that commands can be issued via voice at any time and from any location (in multiple audio zones) within the vehicle, and multiple people can participate in the same conversation. In this scenario of full-car, all-time, multi-person interaction, voice command data containing intent needs to be delivered to the appropriate service executor to complete the command execution. For example, if the driver says "Navigate home," the passenger says "I want to listen to crosstalk," and the person in the back seat says "I want to watch a movie," these commands need to be identified and then executed by the corresponding service executor. However, in related art voice interaction methods, when processing commands requiring multiple rounds of dialogue, they can only process the next multi-round voice command after completing one multi-round dialogue. For example, related art methods require completing the driver's "Navigate home" command before continuing to process the passenger's "I want to listen to crosstalk" command. If multiple multi-round voice commands are processed simultaneously, it is easy to confuse different service executors in the multi-round dialogue, and it is impossible to distribute voice commands containing different intents received in multiple multi-round dialogues to the appropriate service executor for execution. Summary of the Invention
[0005] The present application provides a voice interaction method, device, electronic device, storage medium, and program product.
[0006] According to a first aspect of the present application, a voice interaction method is provided, the method comprising:
[0007] Identify at least one first instruction to obtain at least one first user intention; each first user intention corresponds to a target service;
[0008] When the second instruction is received, the target service corresponding to the second instruction is determined by matching the second instruction with the at least one first user intention, and the second instruction is executed using the target service corresponding to the second instruction.
[0009] In some embodiments, each of the first user intents corresponds to a registration scenario, and the registration scenario includes a target service; and determining the target service corresponding to the second instruction by matching the second instruction with the at least one first user intent includes: identifying a second user intent corresponding to the second instruction;
[0010] Matching the second user intent with at least one registration scenario to determine a target registration scenario corresponding to the second user intent;
[0011] The target service corresponding to the target registration scenario is determined as the target service corresponding to the second instruction.
[0012] In some embodiments, the identifying at least one first instruction to obtain at least one first user intention includes: in response to obtaining multiple user intentions through identifying the first instruction, outputting first prompt information;
[0013] Acquire selection information input by the user based on the first prompt information, determine the first user intention corresponding to the first instruction from the multiple user intentions according to the selection information, and thereby obtain the at least one first user intention based on the at least one first instruction.
[0014] In some embodiments, before determining the target service corresponding to the second instruction by matching the second instruction with the at least one first user intent, the method further includes:
[0015] Matching each of the first user intents with the registered intents of multiple services to obtain a matching result, and determining a target service corresponding to each of the first user intents based on the matching result, wherein the target service is used to execute at least one instruction corresponding to the first user intent;
[0016] When it is determined that the preset conditions are met by executing the target service, scene registration is performed according to the first user intention to obtain a registration scene corresponding to the first user intention, thereby obtaining at least one registration scene based on the at least one first user intention; the at least one registration scene is used to match the second instruction.
[0017] In some embodiments, the registration scenario includes: intent type, scenario identifier, and the correspondence between the scenario identifier and the target service; the preset condition includes the intention feature data that the user intention is of a preset type, and the preset type includes at least one of dialogue interaction and acquisition of data information.
[0018] In some embodiments, matching each of the first user intents with registration intents of multiple services to obtain a matching result, and determining a target service corresponding to each of the first user intents based on the matching result, includes:
[0019] matching the registration intents of the plurality of services with each of the first user intents;
[0020] In response to two or more registration intentions of the plurality of services matching each of the first user intentions, outputting second prompt information, wherein the second prompt information is used to guide the user to select at least one from the two or more services;
[0021] Acquire reply information input by the user based on the second prompt information, determine a corresponding service according to the reply intent of the reply information, and determine the service as a target service corresponding to each of the first user intentions.
[0022] In some embodiments, before matching each of the first user intents with registration intents of multiple services to obtain a matching result, the method further includes:
[0023] Obtain the intent of the services based on their functions and business scenarios;
[0024] At least one intent is registered for each of the multiple services to obtain the registered intents of the multiple services.
[0025] In some embodiments, the registration scenario further includes scenario parameter information, and performing scenario registration according to the first user intent to obtain the registration scenario corresponding to the first user intent includes:
[0026] Obtaining an intent type; the intent type is determined based on the instruction intent expected to be received in the registration scenario;
[0027] generating a scenario identifier corresponding to the first user intention;
[0028] Establishing a correspondence between the scenario identifier and the target service;
[0029] The target service determines the scene parameter information according to the first user intention, and the scene parameter information includes: one or more of display device information, an identifier that needs to receive generative artificial intelligence data, a participant whitelist, and registration time.
[0030] In some embodiments, the second user intent includes: intention feature data of a preset type of conversational interaction; matching the second user intent with at least one registration scenario to determine a target registration scenario corresponding to the second user intent includes:
[0031] Matching the intent feature data of the preset type of conversation interaction with the intent type corresponding to the at least one registered scenario to obtain an intent matching result;
[0032] According to the intention matching result, the target registration scene is determined from the at least one registration scene.
[0033] In some embodiments, determining the target registration scenario from the at least one registration scenario based on the intent matching result includes:
[0034] If the preset type of intention feature data for conversational interaction matches the intention types corresponding to two or more registered scenarios in the at least one registered scenario, obtaining new parameter information corresponding to the preset type of intention feature data for conversational interaction;
[0035] Matching the new parameter information with the scene parameter information of the two or more registered scenes to determine a parameter information matching result;
[0036] According to the parameter information matching result, the target registration scene is determined from the two or more registration scenes.
[0037] In some embodiments, the second user intent includes: a preset type of intent feature data for obtaining data information; the data information includes generative artificial intelligence data, the generative artificial intelligence data being obtained by processing the first user intent and a scene identifier corresponding to the first user intent through an artificial intelligence macro model, wherein the generative artificial intelligence data is bound to the scene identifier corresponding to the first user intent; and matching the second user intent with at least one registration scene to determine a target registration scene corresponding to the second user intent includes:
[0038] Matching the scene identifier corresponding to the second user intention with the scene identifier of the at least one registered scene to obtain an identifier matching result;
[0039] The target registration scenario is determined according to the identification matching result.
[0040] In some embodiments, the method further comprises:
[0041] Determining a display device corresponding to the registered scene according to the scene parameter information of the registered scene;
[0042] Under a preset logout situation, the registration scene is logged out, wherein the preset logout situation includes that the intention feature data of a preset type matching the registration scene is not received for a preset time, and the user turns off at least one of the display devices corresponding to the registration scene.
[0043] According to a second aspect of the present application, a voice interaction device is provided, which includes: an identification module, configured to identify at least one first instruction and obtain at least one first user intention; each first user intention corresponds to a target service; a matching module, configured to, upon receiving a second instruction, determine the target service corresponding to the second instruction by matching the second instruction with the at least one first user intention; and a distribution module, configured to execute the second instruction using the target service corresponding to the second instruction.
[0044] According to the third aspect of the present application, an electronic device is provided, comprising: at least one processor; and a memory connected to the at least one processor for voice interaction; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of the aforementioned first aspect.
[0045] According to a fourth aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the method of the aforementioned first aspect.
[0046] According to a fifth aspect of the present application, a computer program product is provided, comprising a computer program, which implements the method of the first aspect when executed by a processor.
[0047] The voice interaction method provided by the embodiment of the present application includes identifying at least one first instruction to obtain at least one first user intent; each first user intent corresponds to a target service; when a second instruction is received, the target service corresponding to the second instruction is determined by matching the second instruction with at least one first user intent, and the second instruction is executed using the target service corresponding to the second instruction. When the method of the present application receives the second instruction, it can match the second instruction with at least one first user intent corresponding to at least one received first instruction. Since each first user intent corresponds to a target service, it is possible to determine the target service adapted to the second instruction, and execute the second instruction using the target service adapted to the second instruction. In this way, in the scenario of processing multiple multi-round conversations, based on at least one first instruction that has been received, if a second instruction is continued to be received, the second instruction can be automatically matched to the appropriate target service for execution by matching the second instruction with at least one first user intent of at least one first instruction, thereby realizing the distribution of voice intent under full-vehicle and full-time interaction and improving the flexibility of human-computer interaction.
[0048] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0050] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0051] FIG1 is a flow chart of a voice interaction method according to an embodiment of the present application;
[0052] FIG2 is a flow chart of a voice interaction method provided in an embodiment of the present application;
[0053] FIG3 is an example diagram of a voice interaction system provided in an embodiment of the present application;
[0054] FIG4 is a flow chart of a voice interaction method provided in an embodiment of the present application;
[0055] FIG5 is a flow chart of a voice interaction method provided in an embodiment of the present application;
[0056] FIG6 is a schematic diagram of the structure of a voice interaction device provided in an embodiment of the present application;
[0057] FIG7 is a schematic block diagram of an example electronic device 600 provided in an embodiment of the present application. DETAILED DESCRIPTION
[0058] The following description of exemplary embodiments of the present application is made in conjunction with the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0059] Current in-vehicle voice has developed to the point where it can interact with multiple people at all times throughout the vehicle. This means that commands can be issued via voice at any time and in any location (multiple audio zones) within the vehicle, and multiple people can participate in the same conversation. Voice command data containing intent needs to be sent to the appropriate execution side to complete the execution of the command. As large language models demonstrate significant advantages in areas such as intelligent question-answering, their introduction into in-vehicle voice will be an inevitable trend.
[0060] This application proposes a voice interaction system and method, which introduces generative artificial intelligence into in-vehicle voice, realizes the distribution of user intentions in full-vehicle and full-time interaction, and is compatible with the distribution of generative content of large language models.
[0061] The following describes in detail, with reference to the accompanying drawings, a method, device, electronic device, storage medium, and program product for implementing the voice interaction proposed in this application.
[0062] Figure 1 illustrates a voice interaction method provided in an embodiment of the present application. This method is applied to a voice interaction system, and its execution subject may be a processor. The processor may be, for example, a processor for mobile devices such as mobile phones, computers, and tablets, a processor for smart wearable devices such as smart glasses and smart watches, or a processor for vehicles and other transportation vehicles, such as the main control chip in a vehicle's onboard system.
[0063] As shown in Figure 1, the method includes the following steps:
[0064] Step 101: Identify at least one first instruction to obtain at least one first user intention; each first user intention corresponds to a target service.
[0065] In an embodiment of the present application, at least one first instruction may include at least one voice instruction and / or text instruction input by the same user, or at least one first instruction may include at least one voice instruction and / or text instruction input by different users simultaneously or at different times, which is not specifically limited here.
[0066] The semantic intent of each first instruction in at least one first instruction is identified to obtain the semantic intent corresponding to each first instruction as the first user intent corresponding to each first instruction, thereby obtaining at least one first user intent.
[0067] In some embodiments, the voice instruction input by the user can be used as the first instruction, and the voice instruction can be recognized by the artificial intelligence large model to obtain the first user intention.
[0068] In some embodiments of the present application, the artificial intelligence big model includes a large language model (LLM), ChatGPT (Chat Generative Pre-trained Transformer), a multimodal big model and a multimodal cognitive big model, etc.
[0069] In the embodiments of the present application, the artificial intelligence large model used for speech recognition can handle a variety of natural language tasks, such as text classification, question and answer, dialogue, etc., to meet the needs of voice interaction in multiple scenarios, such as intelligent customer service, smart home and autonomous driving.
[0070] In some embodiments, in addition to obtaining the first user intention through the artificial intelligence big model, for scenarios that require receiving multiple rounds of voice commands or conversations, generative artificial intelligence data can also be obtained through the artificial intelligence big model to achieve intelligent question and answer. Generative artificial intelligence data includes pictures, text, audio, etc.
[0071] For example, in the full-time voice interaction in the vehicle, for example, the co-pilot says "I want to listen to crosstalk" and the back row says "What is the difference between L8 and L9", after the artificial intelligence large model obtains and processes the voice command, it can obtain the corresponding first user intention, such as the intention to play crosstalk, the dialogue intention, etc. For the co-pilot's intention to play crosstalk, a prompt message "Which one do you want to listen to" is generated to guide the user to select a specific playback target. For the back row's "What is the difference between L8 and L9", corresponding dialogue information is generated, such as "The difference between L8 and L9 is..." to realize intelligent question and answer.
[0072] In an embodiment of the present application, each first user intention corresponds to a target service, and the target service is used to execute at least one instruction corresponding to the first user intention.
[0073] In some embodiments, the first user intention may be matched with registration intentions of multiple services to obtain a matching result, and the target service corresponding to the first user intention may be determined based on the matching result.
[0074] In some embodiments, after the voice command is understood by the artificial intelligence model, it is necessary to find the corresponding service to execute it. This is achieved by matching the first user intention with the registered intentions of multiple services, and assigning the first user intention to the target service that is adapted to it.
[0075] In some embodiments, in an in-vehicle voice interaction scenario, for example, the driver says "navigate home", the co-driver says "I want to listen to crosstalk", and the back row says "what is the difference between L8 and L9", etc., after these voice commands are understood by the artificial intelligence big model, they need to find the corresponding service to execute. For the driver, a navigation application is needed to execute the intention of navigating home, and for the co-driver's command, a media application is needed to execute the intention of playing crosstalk. For the back row, a business processing unit is needed to receive the dialogue text or voice response generated by the artificial intelligence big model and display it to the back row users.
[0076] Step 102 : upon receiving the second instruction, determining a target service corresponding to the second instruction by matching the second instruction with at least one first user intention, and executing the second instruction using the target service corresponding to the second instruction.
[0077] In some embodiments, upon receiving a second instruction, the second instruction is matched with at least one first user intent. The first instruction and the second instruction corresponding to the first user intent matched by the second instruction belong to multiple rounds of conversations under the same intent and need to be executed using the same target service. Therefore, the target service corresponding to the first user intent matched by the second instruction can be determined as the target service corresponding to the second instruction, and the second instruction can be executed using the target service corresponding to the second instruction.
[0078] In some embodiments, the second user intent corresponding to the second instruction can be identified, and the target first user intent corresponding to the second user intent can be determined by matching the second user intent with at least one first user intent, and the target service corresponding to the target first user intent can be determined as the target service corresponding to the second instruction.
[0079] In some embodiments, the second instruction may include a voice instruction and / or text instruction input by the user, or may include a data instruction containing preset data information. The second instruction may be subjected to semantic intent recognition based on the artificial intelligence big model to obtain the second user intention; or, the second user intention corresponding to the second instruction may be identified based on the data information contained in the second instruction. Exemplarily, the second instruction may include the reply content generated by the artificial intelligence big model based on a first user intention to obtain the data information, such as the user's inquiry intention. In the case where the second instruction contains specific data information, such as generative artificial intelligence data, it can be considered that the intention of the second instruction is to respond to a first user intention for obtaining the data information, and thus, the second user intention corresponding to the second instruction can be determined as the preset type of intention feature data for obtaining data information.
[0080] In some embodiments, a registration scenario can be pre-generated for each first user intent, and the registration scenario includes a target service corresponding to the first user intent. In this case, the second user intent corresponding to the second instruction can be identified; the second user intent can be matched with at least one registration scenario to determine the target registration scenario corresponding to the second user intent; and the target service corresponding to the target registration scenario can be determined as the target service corresponding to the second instruction.
[0081] In some instances, when executing the target service corresponding to the first user intent, if it is determined that a preset condition is satisfied, registration can be performed based on the first user intent to obtain a registration scenario corresponding to the first user intent. Similarly, processing is performed on at least one first user intent to obtain at least one registration scenario.
[0082] Exemplarily, when preset conditions are met, a scene identifier corresponding to each first user intention can be generated and the scene can be registered to obtain a registered scene corresponding to each first user intention, and the registered scene includes: the intention type corresponding to the scene, the scene identifier, and the correspondence between the scene identifier and the target service; the preset conditions include the intention feature data of the first user intention being a preset type, and the preset type includes at least one of dialogue interaction and acquisition of data information.
[0083] In some embodiments, the conversational interaction type includes the need to continuously receive user voice commands. For example, when a user says "I want to watch a movie", multiple movies are searched, and the user is prompted to which one he wants to watch, it is necessary to receive the user's "first one" and similar commands such as the name of the movie.
[0084] In some embodiments, obtaining data information types includes obtaining relevant data information based on current intent feature data, such as obtaining the answer "The difference between L8 and L9 is..." based on the question "What is the difference between L8 and L9", and obtaining a picture of a Corgi based on the intent "What does a Corgi look like".
[0085] In some embodiments, after the first user intention is distributed to the target service, in the process of executing the first user intention through the target service, it can be determined whether the first user intention is a preset type of intention feature data by judging whether the execution of the first user intention requires obtaining further intentions of the user, and whether the execution of the user intention requires receiving generative artificial intelligence data, etc.
[0086] In some embodiments, the intent type corresponding to the scenario is determined based on the instruction intent expected to be received in the scenario.
[0087] For example, the user says "I want to listen to crosstalk", and the media service performs a search for crosstalk and finds multiple results, which the user needs to select. The media service hopes to receive instructions such as "first" or "next page" from the user. At this time, the intent type corresponding to the registration scenario is selection intent.
[0088] In some embodiments, the execution of the first user intention requires obtaining further intentions of the user. For example, when the execution of the first user intention requires the user to make further selections, a scene is registered for the first user intention, a scene identifier is generated, and the intention type corresponding to the scene is generated.
[0089] For example, the driver says "navigate home", and the driver's intention to navigate home is assigned to the navigation service, which displays the navigation route selection interface. The navigation needs to obtain the user's further intention, such as which route the user wants to choose. At this time, the scene is registered for the driver's intention to navigate home, and a scene identifier is generated. The intent type corresponding to the scene and the scene identifier are bound to the navigation service.
[0090] In some embodiments, when the execution of the first user intention requires receiving generative artificial intelligence data, for example, the execution of the first user intention requires obtaining dialogue information generated according to the user's voice instructions, registering a scene for the first user intention, generating a scene identifier, the intent type corresponding to the scene is a dialogue intention, an identifier that requires receiving generative artificial intelligence data, etc.
[0091] For example, the person in the back row said "What is the difference between L8 and L9?", and the conversation intention of the back row is assigned to the conversation service. At this time, it is necessary to receive the content generated based on the question in the back row, such as "The difference between L8 and L9 is...". At this time, the scene is registered for the intent, and the scene identifier is bound to the conversation service.
[0092] In this embodiment, in response to the received second instruction, intent recognition is performed on the second instruction to obtain a second user intent corresponding to the second instruction. Here, the second user intent may include a preset type of intent feature data. Based on the second user intent, a match is performed with the registered scenario. If a match is found with a corresponding target registered scenario, the second instruction is distributed to the target service corresponding to the target registered scenario for execution based on the scenario identifier of the target registered scenario.
[0093] In some embodiments, for multi-round conversation scenarios, at least one registered scenario is registered based on at least one first user intent, such as intent feature data of a preset type of conversational interaction and / or intent feature data of a preset type of acquisition of data information. Upon receiving a second instruction, if the second user intent of the second instruction includes new intent feature data of a preset type of conversational interaction and / or new intent feature data of a preset type of acquisition of data information, the second user intent is matched to a scenario, such as receiving a new user intent and / or acquiring generative artificial intelligence data in multiple rounds of voice instructions.
[0094] In some embodiments, when a new user intention is received, such as the second user intention corresponding to the second instruction, it is matched with at least one registration scenario, and the target registration scenario that matches the new user intention in at least one registration scenario is determined, so that it can be assigned to the target service corresponding to the target registration scenario, thereby realizing the distribution of user intentions throughout the vehicle and at all times.
[0095] For example, when the main driver's intention is "navigate home", the navigation displays multiple routes and asks the user which route to choose. The co-driver says "I want to watch a movie". At this time, multiple movies are displayed on the media and the user is asked which one to watch. Since the execution of both requires obtaining further user intentions, scenes are registered for them respectively, and scene 1 and scene 2 are obtained. If a new intention "I want to watch movie A" is received at this time, the user intention is matched with scene 1 and scene 2. The matching result is that the user intention has the highest match with scene 2, and the media service corresponding to scene 2 is assigned to it.
[0096] In some embodiments, when the execution of the first user intention requires the reception of generative artificial intelligence data, a scene is registered for the first user intention, a scene identifier is generated, the first user intention and the corresponding scene identifier are sent to the artificial intelligence big model for processing at the same time, and generative artificial intelligence data is obtained. The generative artificial intelligence data is bound to the scene identifier corresponding to the first user intention. In response to receiving the generative artificial intelligence data, the corresponding registered scene is matched according to the bound scene identifier, and the generative artificial intelligence data is allocated to the target service corresponding to the registered scene.
[0097] For example, to execute the rear intention "What is the difference between L8 and L9", a dialogue question and answer is required. The dialogue service registers the scenario for the user's intention, generates a scenario identifier, and sends "What is the difference between L8 and L9" and the corresponding scenario identifier to ChatGPT for processing to generate the content "What is the difference between L8 and L9 is...". The generated content is bound to the scenario identifier corresponding to "What is the difference between L8 and L9". Therefore, when the generated content is received, it can be matched to the corresponding registered scenario according to its bound scenario identifier and assigned to the dialogue service corresponding to the registered scenario.
[0098] In summary, according to the embodiments of the present application, when a second instruction is received, the second instruction can be matched with at least one first user intent corresponding to at least one first instruction that has been received. Since each first user intent corresponds to a target service, the target service adapted to the second instruction can be determined, and the second instruction can be executed using the target service adapted to the second instruction. In this way, in the scenario of processing multiple multi-round conversations, based on at least one first instruction that has been received, if a second instruction is continued to be received, the second instruction can be automatically matched to a suitable target service for execution by matching the second instruction with at least one first user intent of at least one first instruction, thereby realizing the distribution of voice intent under full-vehicle and full-time interaction and improving the flexibility of human-computer interaction. The method of the present application can be implemented by introducing a large artificial intelligence model into voice interaction, distributing user intent and generative artificial intelligence data to the appropriate executor based on scene registration, while satisfying the distribution of voice intent under full-vehicle and full-time interaction, and solving the problem of distributing generative content of the large model.
[0099] Based on the embodiment shown in Figure 1, Figure 2 is a flow chart of a voice interaction method provided in an embodiment of the present application. This method can be applied to the voice interaction system shown in Figure 3 as an example, for example, it can be run in the vehicle's on-board system.
[0100] The method includes the following steps 201-207.
[0101] Step 201: obtain the voice command input by the user, recognize the voice command through the artificial intelligence large model, and obtain the first user intention.
[0102] Here, the voice instruction input by the user is taken as the first instruction, and the voice instruction is recognized by the artificial intelligence large model to obtain the first user intention corresponding to the first instruction.
[0103] In some embodiments, for each of the at least one first instruction, in response to identifying the first instruction and obtaining multiple user intents, first prompt information can be output; selection information input by the user based on the first prompt information is obtained, and the first user intent corresponding to the first instruction is determined from the multiple user intents based on the selection information. Thus, at least one first user intent can be obtained based on the at least one first instruction.
[0104] In some embodiments, in the voice interaction system shown in FIG3 , the AI capability layer recognizes and processes the voice commands issued by the user to obtain the user's intent, wherein the AI capability layer provides AI capabilities including natural language understanding (NLU), natural language generation (NLG), and generative artificial intelligence.
[0105] In some embodiments, the first instruction includes a single-sentence voice instruction. In response to recognizing multiple user intents from the single-sentence voice instruction, first prompt information is output, the first prompt information being used to guide the user to select at least one user intent from the multiple user intents; selection information input by the user is obtained, and the single user intent of the voice instruction is determined from the multiple user intents based on the selection information as the first user intent corresponding to the first instruction.
[0106] In some embodiments, as shown in FIG3 , a multiple user intentions corresponding to a voice instruction are semantically arbitrated by the dialogue management (hereinafter referred to as DM) of the voice interaction system to determine a single user intention of the voice instruction. After determination, the intention is passed to the artificial intelligence assistant management (hereinafter referred to as CM), which assigns it to the AI assistant of the target service.
[0107] For example, when a user says "I want to see the sky", there may be multiple intentions, such as one is the car control intention to open the sunroof, and the other is the movie search intention (i.e. the movie "The Sky"). At this time, the DM can ask the user whether he needs to open the sunroof or watch a movie. If it is determined that the first user intention is to open the car window.
[0108] Step 202: Obtain the intent corresponding to the services based on the functions and business scenarios of the multiple services.
[0109] Step 203: Register at least one intent for each of the multiple services to obtain the registered intents for the multiple services.
[0110] In some embodiments, the voice interaction system shown in FIG3 includes an AI assistant layer, which includes multiple artificial intelligence assistants (hereinafter referred to as AI assistants), such as the media AI assistant, the vehicle control AI assistant, and the navigation AI assistant in FIG3 . Each AI assistant corresponds to a service. For example, for the navigation service, a navigation AI assistant manages one or more navigation applications to execute user navigation-related intents such as the "navigate home" intent. For the media service, a media AI assistant manages one or more media applications to execute user media-related intents such as the "play movie" intent. In addition, the AI assistant layer also includes a vehicle control AI assistant and a general conversational AI assistant.
[0111] In some embodiments, when an application is started, the corresponding intent is registered for the AI assistant responsible for the application. For example, when a media application is started, the media AI assistant is registered with intents related to media search and playback control, such as "play music", "I want to watch a movie", "I want to listen to crosstalk", and "fast forward 30 seconds". For the car control AI assistant, the intents such as "open the trunk", "turn on the reading light", and "turn on seat massage" are registered for it. Other AI assistants are similar and will not be repeated here.
[0112] Step 204 : Match the first user intent with the registration intents of multiple services to obtain a matching result, and determine a target service based on the matching result, where the target service is used to execute at least one instruction corresponding to the first user intent.
[0113] In some embodiments, for each first user intent in at least one first user intent, each first user intent can be matched with the registration intentions of multiple services to obtain a matching result, and the target service corresponding to each first user intent can be determined based on the matching result.
[0114] In some embodiments, each first user intention is matched with the registration intentions of multiple services to obtain a matching result, and the target service corresponding to each first user intention is determined based on the matching result, including: matching the registration intentions of multiple services with each first user intention; in response to the presence of two or more registration intentions of multiple services that match the first user intention, outputting a second prompt message, the second prompt message being used to guide the user to select at least one from the two or more services; obtaining reply information input by the user based on the second prompt information, determining the corresponding service based on the reply intention of the reply information, and determining the service as the target service corresponding to the first user intention.
[0115] That is to say, when matching a first user intention with the registration intentions of multiple services, if there are two or more services among the multiple services whose registration intentions match the first user intention, a second prompt message is output. According to the reply information input by the user based on the second prompt message, the service corresponding to the first user intention can be determined as the target service corresponding to the first user intention.
[0116] It should be noted that multiple services can correspond to one registration intent, one user intent corresponds to one target service, and one target service can correspond to one or more user intents.
[0117] In some embodiments, such as applied to the voice interaction system shown in FIG3 , when multiple AI assistants register the same intent and a user triggers the intent, the CM arbitrates the intent and asks the user which AI assistant to use to execute it.
[0118] For example, if a user says "I want to watch a movie", and if Video Application 1 and Video Application 2 each register the video search intent as an AI assistant, then CM can ask the user which application to use for the search. It should be noted that in other implementations, there is also a situation where a media AI assistant manages all media applications, and this application is not limited to this.
[0119] Step 205, by executing the target service, it is determined that the preset conditions are met and the scene identifier corresponding to the first user's intention is generated and the scene is registered to obtain a registered scene.
[0120] Among them, the registration scenario includes: the intention type corresponding to the scenario, the scenario identifier, and the correspondence between the scenario identifier and the target service; the preset conditions include the intention feature data of the first user intention being a preset type, and the preset type includes at least one of dialogue interaction and acquisition of data information.
[0121] In some embodiments of the present application, after generative artificial intelligence is introduced into the full-vehicle, full-time voice interaction, the user voice commands are distributed to the target service for execution, including three situations: one is that the command is directly executed and the execution result is returned; the second is that the command requires further instructions from the user, including scenarios that need to support multiple rounds of dialogue; the third is that it is necessary to receive artificial intelligence-generated content, including scenarios that need to support intelligent question and answer. In the latter two cases, the intent registration scenario corresponds to the user voice command.
[0122] In some embodiments of the present application, when it is necessary to receive further user intentions and / or to receive artificial intelligence-generated content during the execution of the first user intention, a scene identifier corresponding to the first user intention is generated and the scene registration is performed.
[0123] In some embodiments, the registration scenario also includes scenario parameter information. The above-mentioned scenario registration based on the first user intent and obtaining the registration scenario corresponding to the first user intent include: obtaining the intent type under preset conditions and generating a scenario identifier corresponding to the first user intent; establishing a correspondence between the scenario identifier and the target service; and determining the scenario parameter information based on the first user intent through the target service. The scenario parameter information includes: display device information corresponding to the scenario, an identifier for receiving generative artificial intelligence data, a participant whitelist, and one or more of the registration time. The intent type is determined based on the instruction intent expected to be received in the registration scenario.
[0124] In some embodiments, the registration scenario is often accompanied by a UI display, that is, associated with a screen inside the car.
[0125] In some embodiments, the scene ID generation rule is taken as an example: "sceneid"-ScreenId-timestamp-6-digit random number; each field is an identification header, the screen corresponding to the scene, a timestamp, and a random number. In addition, other parameters can also be defined to achieve specific purposes. For example, the following parameters can be defined: the intent whitelist includes the intention of the AI assistant to receive instructions at this time, the screen information includes the screen corresponding to the scene, the participant whitelist includes whether the scene is a multi-person scene, that is, whether to receive instructions from other people, the screen information includes the screen corresponding to the scene, and parameters indicating that the scene needs to receive generated artificial data, etc.
[0126] In some embodiments, applied to the voice interaction system shown in Figure 3, when the AI assistant needs to dynamically receive intent or receive generative artificial intelligence data, perform scene registration, and generate a scene ID, the relationship between the scene ID and the AI assistant will be recorded in the CM, and the scene information will be passed to the DM.
[0127] For example, the user says "I want to listen to crosstalk", and the media AI assistant performs a search for crosstalk and finds multiple results. The user needs to choose. Then the AI assistant displays the searched media resources and announces "Multiple results found, which one do you want to listen to?" At this time, the AI assistant hopes to receive instructions such as "first" and "next page" from the user. Then the AI assistant will register the scene, generate a scene ID such as B, and the intent whitelist includes the intents that the AI assistant wants to receive instructions such as "first" and "next page". The relationship between scene B and the media AI assistant will be recorded in CM, and the scene information will be passed to DM.
[0128] For example, a user says "compare L8 and L9" and is assigned to a general conversational AI assistant. At this time, the user hopes that the AI assistant will receive the intelligent answer generated for the question. The general conversational AI assistant will register the scene and generate a scene ID such as C, which identifies the parameter information of the scene that needs to receive the generated artificial data. The relationship between scene C and the general conversational AI assistant will be recorded in CM, and the scene information will be passed to DM.
[0129] Step 206, in response to receiving the new preset type of intention feature data, matching is performed based on the new preset type of intention feature data and the registration scene. If a match is found with the corresponding target registration scene, the new preset type of intention feature data is distributed to the target service corresponding to the target registration scene for execution according to the scene identifier of the target registration scene.
[0130] In some embodiments, in response to receiving new intention feature data of a preset type of dialogue interaction and / or in response to receiving new intention feature data of a preset type of obtaining data information, scene matching is performed, for example, new user intentions are received in multiple rounds of voice commands and / or generative artificial intelligence data are obtained.
[0131] In some embodiments, as shown in the scenario in Figure 3, when the AI assistant registers the scene, the CM is responsible for generating a scene identifier, and the relationship between the scene identifier and the AI assistant is recorded in the CM, and the scene information is passed to the DM; when there is a new user intention or AIGC related to the scene, the DM is responsible for associating the new user intention or generative artificial intelligence data with the scene ID and passing it to the CM; the CM routes the new user intention or generative artificial intelligence data to the AI assistant that registered the scene according to the scene ID.
[0132] In some embodiments, the second user intention includes: intention feature data of a preset type of conversational interaction; matching the second user intention with at least one registration scenario to determine the target registration scenario corresponding to the second user intention, including: matching the intention feature data of a preset type of conversational interaction with the intention type corresponding to at least one registration scenario to obtain an intention matching result; determining the target registration scenario from at least one registration scenario based on the intention matching result.
[0133] In some embodiments, as shown in FIG4 , in response to receiving new intention feature data of a preset type of conversation interaction, the process of matching the new intention feature data of the preset type with the registered scenario includes steps 301 - 303 .
[0134] Step 301: obtain a new voice command input by the user, recognize the new voice command through the artificial intelligence model, and obtain new intention feature data of the preset type of dialogue interaction.
[0135] Here, the new voice instruction input by the user is equivalent to the second instruction. The new preset intention feature data of the conversational interaction type is equivalent to the second user intention.
[0136] Step 302: Match the new preset intent feature data of the conversation interaction type with the intent types corresponding to the multiple registered scenarios to obtain an intent matching result.
[0137] In some embodiments, the main driver says "navigate home", the navigation AI assistant displays the navigation route selection interface, prompts "Which route do you want to choose", the registered scene ID is A, and the intent list is the selection intent. The co-driver says "I want to watch a movie", the media AI assistant pops up the power resource card, prompts "Multiple results found, which one do you want to watch" and the registered scene ID is B, the intent list is the selection intent. If the main driver says "the first one", the voice command is processed to obtain a new intent as the selection intent. After intent matching, it can be seen that the intent matches the intent list of registered scene A and registered scene B.
[0138] Step 303: Determine a target registration scenario from multiple registration scenarios based on the intent matching result.
[0139] In some embodiments, determining a target registration scene from multiple registration scenes based on the intent matching results includes: if the new preset type of intent feature data for conversational interaction matches the intent types corresponding to two or more registration scenes in the multiple registration scenes, obtaining parameter information of the new preset type of intent feature data for conversational interaction; matching the parameter information of the new preset type of intent feature data for conversational interaction with the scene parameter information of the two or more registration scenes to obtain a parameter information matching result; and determining the target registration scene from the two or more registration scenes based on the parameter information matching result.
[0140] In some embodiments, the parameter information of the new preset type of intention feature data for dialogue interaction includes the voice zone or user corresponding to the source of the voice command, and the voice zones include the main driver, co-driver, back row, etc.
[0141] In some embodiments, when the intent whitelists corresponding to multiple scenarios include the same intent and the user triggers the intent, scenario arbitration is performed on the multiple scenarios through DM, one of the scenarios is selected, the intent is associated with the scenario identifier of the scenario and then passed to CM, which then sends it to the AI assistant that registered the scenario.
[0142] In some embodiments, the DM performs scene arbitration based on scene parameters, including performing scene arbitration based on one or more of screen information, participant whitelist, and registration time.
[0143] In some embodiments, it can be determined whether the registration scenario is a multi-person scenario based on the participant whitelist. If it is a multi-person scenario, instructions from other people can be received, and whether the registration scenario matches is determined based on the participant whitelist and the voice zone or user corresponding to the voice instruction source of the new preset type of intention feature data for dialogue interaction.
[0144] For example, the driver says "Navigate home", the navigation AI assistant displays the navigation route selection interface, prompting "Which route do you want to choose", the registered scene ID is A, and the intent list is the selection intent. The co-driver says "I want to watch a movie", the media AI assistant pops up the power resource card, tts prompts "Multiple results found, which one do you want to watch" and the registered scene ID is B, and the intent list is also the selection intent. At this time, the back row says "first" according to the intent matching results, and both registered scenes A and B are matched. At this time, if the registered scene A is not a multi-person scene, and the registered scene B is a multi-person scene, the selection intent of "first" will be sent to the media AI assistant corresponding to scene B, and the operation of selecting to play the first movie will be executed.
[0145] In some embodiments, the screen information includes the screen associated with the scene and the sound zone where the screen is located. The sound zones include the main driver's seat, the co-driver's seat, the back row, etc. It is determined whether the sound zone corresponding to the voice instruction source of the new preset type of intention feature data for dialogue interaction is the same as the sound zone where the screen associated with the registration scene is located. If they are the same sound zone, the registration scene is determined to be the target registration scene.
[0146] For example, the driver says "Navigate home", the navigation AI assistant displays the navigation route selection interface, prompting "Which route do you want to choose", the registered scene ID is A, and the intent list is the selection intent. The person in the back row says "I want to watch a movie", and the media AI assistant pops up the power resource card. The voice prompt system (Text-to-Speech, tts) prompts "Multiple results found, which one do you want to watch" and the registered scene ID is B. The intent list is also the selection intent. At this time, the driver says "first" according to the intent matching result, and both registered scenes A and B are matched. At this time, if the screen associated with registered scene A is on the driver's side and the screen associated with registered scene B is in the back row, the selection intent of "first" is sent to the navigation AI assistant corresponding to scene A, and the operation of selecting the first route is executed to start navigation and complete the command execution.
[0147] In some embodiments, if the new user intent matches the intent type corresponding to two or more registration scenarios among multiple registration scenarios, the registration scenario with the earlier registration time is determined as the target registration scenario based on the registration time of the registration scenario.
[0148] In some embodiments, the first user intention includes: a preset type of intention feature data for obtaining data information; the data information includes generative artificial intelligence data, and the second user intention includes: a preset type of intention feature data for obtaining data information; the first user intention and the second user intention can be processed through the process shown in Figure 5 to achieve the distribution of generative artificial intelligence data, including steps 401-403.
[0149] Step 401, in response to the first user intention being the intention feature data of a preset type and the preset type being obtaining data information, the first user intention and the scene identifier corresponding to the first user intention are processed by an artificial intelligence big model to obtain generative artificial intelligence data, wherein the generative artificial intelligence data is bound to the scene identifier corresponding to the first user intention.
[0150] For example, based on Figure 3, in response to the first user intent being a preset type of intent feature data, and the preset type being obtaining data information, the DM generates a request to obtain data information based on the scenario ID and the first user intent transmitted by the CM, and transmits it to the artificial intelligence model. The artificial intelligence model processes the first user intent and the scenario identifier corresponding to the first user intent to obtain generative artificial intelligence data, and then sends the generative artificial intelligence data bound to the scenario identifier to the DM. Alternatively, the artificial intelligence model can process the first user intent to obtain generative artificial intelligence data, and send the generated artificial intelligence data as the reply content or response information to the request to obtain data information to the DM, which then binds the generative artificial intelligence data to the scenario identifier.
[0151] Step 402: Match the scene identifier corresponding to the second user intention with the scene identifier of at least one registered scene to obtain an identifier matching result.
[0152] The artificial intelligence model processes the first user intent and the scene identifier corresponding to the first user intent to obtain generative artificial intelligence data, and then returns the generative artificial intelligence data bound to the scene identifier corresponding to the first user intent as a second instruction. Upon receiving the second instruction including the generative artificial intelligence data and the scene identifier, it is determined that the second user intent includes intent feature data of a preset type of obtaining data information. The scene identifier corresponding to the second user intent is matched with the scene identifier of at least one registered scene to obtain an identifier matching result.
[0153] Step 403: Determine a target registration scene from at least one registration scene according to the identification matching result.
[0154] Based on the above process, it can be understood that the scene identifier corresponding to the second user intention and the scene identifier corresponding to the first user intention are the same scene identifier. Therefore, according to the identifier matching result, the target registration scene can be determined from at least one registration scene.
[0155] In some embodiments, in the voice interaction system shown in Figure 3, the back row says "Compare L8 and L9", which requires receiving generative artificial intelligence data. The general conversation AI assistant registers the scene ID C for the intent, and sends the intent and the registered scene ID to GPT for processing to generate comparison information of L8 and L9. The comparison information will be identified in DM as AI data with scene ID C. If the current registered scenes include scene A corresponding to the navigation AI assistant, scene B corresponding to the media AI assistant, and scene C corresponding to the general conversation AI assistant, then according to the identified scene ID C, the AI data will be distributed to the general conversation AI assistant by CM.
[0156] Step 207: Under the preset deregistration situation, deregister the registration scene.
[0157] Among them, the preset logout situations include at least one of not receiving new preset type of intention feature data matching the registration scene for a preset period of time and the user turning off the display device corresponding to the registration scene.
[0158] In some embodiments, the display device corresponding to the registered scene is determined based on the scene parameter information of the registered scene.
[0159] In some embodiments, when the AI assistant does not receive the AI layer processing results, i.e., new user intentions and / or generative artificial intelligence data, within a preset time, and when the user turns off the display device corresponding to the registered scene, the scene can be deregistered; for example, the user says "I want to watch a movie", the screen displays the UI, showing the searched movie resources, and asking the user which one to watch. At this time, the user does not issue subsequent instructions for a long time, the UI interface disappears, or the user directly closes the UI interface, and the AI assistant deregisters the scene.
[0160] To sum up, according to the embodiments of the present application, the distribution of AI capabilities is carried out based on scenarios, which realizes the full-vehicle, full-time and multi-person interaction of in-vehicle voice while supporting the generative content distribution of large models.
[0161] Corresponding to the above-mentioned voice interaction method, the present application also proposes a voice interaction device. FIG6 is a structural diagram of a voice interaction device 500 provided in an embodiment of the present application. As shown in FIG4 , it includes:
[0162] The identification module 510 is configured to identify at least one first instruction and obtain at least one first user intention; each first user intention corresponds to a target service; the matching module 520 is configured to determine the target service corresponding to the second instruction by matching the second instruction with the at least one first user intention when receiving the second instruction; the distribution module 530 is configured to execute the second instruction using the target service corresponding to the second instruction.
[0163] In some embodiments, the matching module 520 is further configured to: identify a second user intention corresponding to the second instruction; match the second user intention with at least one registration scenario to determine a target registration scenario corresponding to the second user intention; and determine the target service corresponding to the target registration scenario as the target service corresponding to the second instruction.
[0164] In some embodiments, the recognition module 510 is further configured to output a first prompt message in response to identifying multiple user intentions obtained by identifying the first instruction; obtain selection information input by the user based on the first prompt message, and determine the first user intention corresponding to the first instruction from the multiple user intentions according to the selection information, thereby obtaining the at least one first user intention based on the at least one first instruction.
[0165] In some embodiments, the device also includes a registration module, and the registration module is configured to match each of the first user intentions with the registration intentions of multiple services to obtain a matching result before determining the target service corresponding to the second instruction by matching the second instruction with the at least one first user intention, and determine the target service corresponding to each of the first user intentions based on the matching result, and the target service is used to execute at least one instruction corresponding to the first user intention; when it is determined that the preset conditions are met by executing the target service, the scene is registered according to the first user intention to obtain the registration scene corresponding to the first user intention, thereby obtaining at least one registration scene based on the at least one first user intention; the at least one registration scene is used to match the second instruction.
[0166] In some embodiments, the registration scenario includes: intent type, scenario identifier, and the correspondence between the scenario identifier and the target service; the preset condition includes the intention feature data that the user intention is of a preset type, and the preset type includes at least one of dialogue interaction and acquisition of data information.
[0167] In some embodiments, the matching module 520 is further configured to match the registration intentions of the multiple services with each of the first user intentions; in response to the presence of two or more registration intentions of the multiple services that match each of the first user intentions, output a second prompt message, wherein the second prompt message is used to guide the user to select at least one from two or more services; obtain the reply information input by the user based on the second prompt information, determine the corresponding service according to the reply intention of the reply information, and determine the service as the target service corresponding to each of the first user intentions.
[0168] In some embodiments, the registration module is further configured to obtain the intents corresponding to the services based on the functions and business scenarios of multiple services before matching the results; register at least one intent for each of the multiple services to obtain the registration intents of the multiple services.
[0169] In some embodiments, the registration module is further configured to obtain an intent type; the intent type is determined based on the instruction intent expected to be received in the registration scenario; a scene identifier corresponding to the first user intent is generated; a correspondence between the scene identifier and the target service is established; and scene parameter information is determined based on the first user intent through the target service, the scene parameter information including: display device information, an identifier that needs to receive generative artificial intelligence data, a participant whitelist, and one or more of the registration time.
[0170] In some embodiments, the second user intention includes: intention feature data of a preset type of conversational interaction; the matching module 520 is further configured to match the intention feature data of a preset type of conversational interaction with the intention type corresponding to the at least one registration scenario to obtain an intention matching result; and determine the target registration scenario from the at least one registration scenario based on the intention matching result.
[0171] In some embodiments, the matching module 520 is further configured to obtain new parameter information corresponding to the intention feature data of the preset type of conversational interaction if the intention feature data of the preset type matches the intention type corresponding to two or more registration scenarios in the at least one registration scenario; match the new parameter information with the scene parameter information of the two or more registration scenarios to determine the parameter information matching result; and determine the target registration scenario from the two or more registration scenarios based on the parameter information matching result.
[0172] In some embodiments, the second user intention includes: a preset type of intention feature data for obtaining data information; the data information includes generative artificial intelligence data, and the generative artificial intelligence data is obtained by processing the first user intention and the scene identification corresponding to the first user intention through an artificial intelligence big model, wherein the generative artificial intelligence data is bound to the scene identification corresponding to the first user intention; the matching module 520 is further configured to match the scene identification corresponding to the second user intention with the scene identification of the at least one registered scene to obtain an identification matching result; and determine the target registration scene based on the identification matching result.
[0173] In some embodiments, the device also includes a logout module, which is configured to determine the display device corresponding to the registration scene based on the scene parameter information of the registration scene; and log out of the registration scene under a preset logout situation, wherein the preset logout situation includes not receiving intention feature data of a preset type matching the registration scene for a preset time, and the user turning off at least one of the display devices corresponding to the registration scene.
[0174] In summary, according to the embodiments of the present application, the device distributes AI capabilities based on scenarios through the recognition module, matching module, registration module and distribution module, realizing full-vehicle, full-time and multi-person interaction of in-vehicle voice while supporting generative content distribution of large models.
[0175] It should be noted that since the device embodiment of the present application corresponds to the above-mentioned method embodiment, the above-mentioned explanation of the method embodiment is also applicable to the device of this embodiment, and the principle is the same. For details not disclosed in the device embodiment, refer to the above-mentioned method embodiment, and no further details will be given in this application.
[0176] According to an embodiment of the present application, the present application also provides an electronic device, a readable storage medium and a computer program product.
[0177] FIG7 shows a schematic block diagram of an example electronic device 600 that can be used to implement an embodiment of the present application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.
[0178] As shown in FIG7 , the device 600 includes a computing unit 601 that can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 602 or a computer program loaded from a storage unit 608 into a RAM (Random Access Memory) 603. Various programs and data required for the operation of the device 600 can also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An I / O (Input / Output) interface 605 is also connected to the bus 604.
[0179] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as a keyboard, mouse, etc.; output unit 607, such as various types of displays, speakers, etc.; storage unit 608, such as a magnetic disk, optical disk, etc.; and voice interaction unit 609, such as a network card, modem, wireless voice interaction transceiver, etc. Voice interaction unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0180] The computing unit 601 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units that run machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as the voice interaction method. For example, in some embodiments, the voice interaction method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 600 via the ROM 602 and / or the voice interaction unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the method described above can be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to execute the aforementioned voice interaction method in any other appropriate manner (for example, by means of firmware).
[0181] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System on Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0182] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow charts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0183] In the context of the present application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0184] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0185] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by voice interaction of digital data in any form or medium (e.g., a voice interaction network). Examples of voice interaction networks include: LANs (Local Area Networks), WANs (Wide Area Networks), the Internet, and blockchain networks.
[0186] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a voice-interactive network. This client-server relationship is established by computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within a cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.
[0187] It's important to note that artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). This encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.
[0188] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this application can be achieved. This is not a limitation herein.
[0189] The above specific embodiments do not constitute a limitation on the scope of protection of this application. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the scope of protection of this application. Industrial Applicability
[0190] When receiving a second instruction, the present application can match the received second instruction to the target service corresponding to the previously received first instruction based on the matching of the instruction intent for processing. In this way, in the scenario where at least one first instruction is received and then the second instruction is received, the second instruction can be automatically matched to the appropriate target service for execution, thereby realizing the distribution of voice intentions in full-vehicle and full-time interaction, and improving the flexibility of human-computer interaction.
Claims
1. A voice interaction method, the method comprising: Identify at least one first instruction to obtain at least one first user intention; Each of the first user intentions corresponds to a target service; When the second instruction is received, the target service corresponding to the second instruction is determined by matching the second instruction with the at least one first user intention, and the second instruction is executed using the target service corresponding to the second instruction.
2. The method according to claim 1, wherein: Each of the first user intentions corresponds to a registration scenario, and the registration scenario includes a target service; The determining a target service corresponding to the second instruction by matching the second instruction with the at least one first user intention includes: identifying a second user intention corresponding to the second instruction; Matching the second user intention with at least one registration scenario to determine a target registration scenario corresponding to the second user intention; The target service corresponding to the target registration scenario is determined as the target service corresponding to the second instruction.
3. The method according to claim 1 or 2, wherein: The identifying at least one first instruction to obtain at least one first user intention includes: In response to identifying the first instruction and obtaining multiple user intentions, outputting first prompt information; Acquire selection information input by the user based on the first prompt information, determine the first user intent corresponding to the first instruction from the multiple user intents according to the selection information, and thereby obtain the at least one first user intent based on the at least one first instruction.
4. The method according to any one of claims 1 to 3, wherein: Before determining the target service corresponding to the second instruction by matching the second instruction with the at least one first user intention, the method further includes: Matching each of the first user intents with the registered intents of multiple services to obtain a matching result, and determining a target service corresponding to each of the first user intents according to the matching result, wherein the target service is used to execute at least one instruction corresponding to the first user intent; When it is determined that the preset conditions are met by executing the target service, scene registration is performed according to the first user intention to obtain a registration scene corresponding to the first user intention, thereby obtaining at least one registration scene based on the at least one first user intention; the at least one registration scene is used to match the second instruction.
5. The method according to claim 4, wherein: The registration scenario includes: intent type, scenario identifier, and the correspondence between the scenario identifier and the target service; the preset condition includes the intention feature data that the user intention is of a preset type, and the preset type includes at least one of dialogue interaction and acquisition of data information.
6. The method according to claim 4 or 5, wherein: The step of matching each of the first user intentions with the registration intentions of multiple services to obtain a matching result, and determining a target service corresponding to each of the first user intentions according to the matching result, includes: matching the registration intents of the plurality of services with each of the first user intents; In response to the presence of two or more registration intentions of the multiple services matching each of the first user intentions, outputting second prompt information, wherein the second prompt information is used to guide the user to select at least one from the two or more services; Acquire reply information input by the user based on the second prompt information, determine a corresponding service according to the reply intent of the reply information, and determine the service as a target service corresponding to the first user intent.
7. The method according to any one of claims 4 to 6, wherein: Before matching each of the first user intentions with registration intentions of multiple services to obtain a matching result, the method further includes: Obtain the intent of the services based on the functions and business scenarios of multiple services; At least one intent is registered for each of the multiple services to obtain the registered intents of the multiple services.
8. The method according to any one of claims 4 to 7, wherein: The registration scene also includes scene parameter information, and the scene registration is performed according to the first user intention to obtain the registration scene corresponding to the first user intention, including: Obtaining an intent type; the intent type is determined according to the instruction intent expected to be received in the registration scenario; generating a scenario identifier corresponding to the first user intention; Establishing a correspondence between the scenario identifier and the target service; The target service determines scene parameter information according to the first user intention, and the scene parameter information includes: one or more of display device information, an identifier that needs to receive generative artificial intelligence data, a participant whitelist, and a registration time.
9. The method according to any one of claims 2 to 8, wherein: The second user intention includes: intention feature data of a preset type of dialogue interaction; matching the second user intention with at least one registration scenario to determine a target registration scenario corresponding to the second user intention includes: Matching the intention feature data of the preset type of dialogue interaction with the intention type corresponding to the at least one registered scenario to obtain an intention matching result; According to the intention matching result, the target registration scene is determined from the at least one registration scene.
10. The method according to claim 9, wherein: The determining the target registration scene from the at least one registration scene according to the intention matching result includes: If the preset type of the intention feature data of the conversation interaction matches the intention types corresponding to two or more registered scenarios in the at least one registered scenario, obtaining new parameter information corresponding to the preset type of the intention feature data of the conversation interaction; Matching the new parameter information with the scene parameter information of the two or more registered scenes to determine a parameter information matching result; According to the parameter information matching result, the target registration scene is determined from the two or more registration scenes.
11. The method according to any one of claims 2 to 8, wherein: The second user intention includes: a preset type of intention feature data for obtaining data information; the data information includes generative artificial intelligence data, the generative artificial intelligence data is obtained by processing the first user intention and the scene identifier corresponding to the first user intention through an artificial intelligence big model, wherein the generative artificial intelligence data is bound to the scene identifier corresponding to the first user intention; the matching of the second user intention with at least one registration scene to determine the target registration scene corresponding to the second user intention includes: Matching the scene identifier corresponding to the second user intention with the scene identifier of the at least one registered scene to obtain an identifier matching result; The target registration scenario is determined according to the identification matching result.
12. The method according to claim 8, wherein: The method further comprises: Determining a display device corresponding to the registered scene according to the scene parameter information of the registered scene; Under a preset logout situation, the registration scene is logged out, wherein the preset logout situation includes that intention feature data of a preset type matching the registration scene is not received for a preset period of time, and the user turns off at least one of the display devices corresponding to the registration scene.
13. A voice interaction device, comprising: an identification module, configured to identify at least one first instruction and obtain at least one first user intention; Each of the first user intentions corresponds to a target service; a matching module configured to, upon receiving the second instruction, determine a target service corresponding to the second instruction by matching the second instruction with the at least one first user intention; The distribution module is configured to execute the second instruction using the target service corresponding to the second instruction.
14. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 12.
15. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-12.
16. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 12.
17. A vehicle, comprising: The voice interaction device as claimed in claim 13, or the electronic device as claimed in claim 14.
Citation Information
Patent Citations
Interaction method and device, medium and operating system
CN110874202A
Voice interaction method, device and system
CN111312235A
Voice interaction method and server
CN115116450A
Voice interaction method, device and equipment of vehicle and storage medium
CN115641845A
Vehicle intelligent scenarized interaction system and method, electronic equipment and vehicle
CN115756166A
Cited By
Intelligent assistant-based teaching assistance method, device, and computer program product
CN122509639A