Voice interaction method and device

Through the artificial intelligence model, it recognizes user intentions and matches them, generates scene identifiers for scene registration, solving the problem of difficulty in distributing voice commands in full-time and multi-person interactive scenarios, and achieving efficient command distribution and execution.

CN120183392APending Publication Date: 2025-06-20BEIJING CHJ AUTOMOTIVE TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202311756133.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the existing technology, in the full-vehicle and full-time and multi-person interaction scenarios, it is difficult to process multiple voice commands containing different consents at the same time, resulting in confusion among the business executors and unable to effectively distribute the instructions.

Method used

The user's intent is identified through the artificial intelligence model and match it with the registration intentions of multiple services, generate scene identification for scene registration, establish target services and scene relationships, and then distribute intent feature data to the appropriate target services for execution.

Benefits of technology

It realizes efficient distribution and execution of multiple voice commands containing different ideas in full-vehicle, full-time, multi-person interaction scenarios, avoids confusion among business executors and improves the system's processing capabilities and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183392A_ABST
    Figure CN120183392A_ABST
Patent Text Reader

Abstract

The invention provides a voice interaction method, system and device, and the method comprises the steps: obtaining a voice instruction inputted by a user, carrying out the recognition of the voice instruction through an artificial intelligence large model, and obtaining a user intention; the user intention is matched with the registration intentions of the multiple services to obtain a matching result, a target service is determined according to the matching result, and the target service is used for executing at least one instruction corresponding to the user intention; the method comprises the steps of generating a scene identifier corresponding to a user intention under a preset condition and performing scene registration to obtain a registered scene, the registered scene comprising an intention type corresponding to the scene, the scene identifier and a corresponding relationship between the scene identifier and a target service. The user intention and the generative content are distributed to a proper executor based on scene registration, and the problem of distribution of the generative content of a large model is solved while the requirement for voice intention distribution under full-vehicle full-time interaction is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing, and in particular, to a voice interaction method and apparatus. Background Art

[0002] The current in-vehicle voice has developed to the ability of full-vehicle full-time and multi-person interaction, that is, voice commands can be issued at any time and any position (multiple sound zones) in the vehicle, and multiple people can also participate in the same conversation. In the case of full-vehicle full-time and multi-person interaction, the voice command data containing intentions needs to be sent to the appropriate execution side to complete the execution of the command. For example, the driver says "Navigate to Tiananmen", the co-driver says "I want to listen to cross talk", and the rear seat passenger says "I want to watch a movie". After these commands are recognized, the corresponding service executor needs to be found to execute them. However, in the related art, after receiving multiple commands including different intentions, the voice interaction method can only process a single voice command containing an intention at the same time. For example, in the related art, after processing the driver's command "Navigate to Tiananmen", it can continue to process the co-driver's command "I want to listen to cross talk". If multiple voice commands containing different intentions are processed simultaneously, it will confuse the service executors and prevent the multiple voice commands containing different intentions from being distributed to the appropriate service executors for execution. Summary of the Invention

[0003] The present disclosure provides a voice interaction method, apparatus, electronic device, storage medium, and program product. According to a first aspect of the present disclosure, a voice interaction method is provided. The method includes: obtaining a voice command input by a user, and recognizing the voice command through an artificial intelligence large model to obtain a user intention; matching the user intention with the registered intentions of multiple services to obtain a matching result, and determining a target service according to the matching result, where the target service is used to execute at least one command corresponding to the user intention; generating a scene identifier corresponding to the user intention and performing scene registration under a preset condition to obtain a registered scene, where the registered scene includes: the intention type corresponding to the scene, the scene identifier, and the corresponding relationship between the scene identifier and the target service; the preset condition includes that the user intention is intention feature data of a preset type, and the preset type includes at least one of dialogue interaction and obtaining data information; in response to receiving new intention feature data of the preset type, matching the new intention feature data of the preset type with the registered scene, and if a corresponding target registered scene is matched, distributing the new intention feature data of the preset type to the target service corresponding to the target registered scene for execution.

[0004] In some embodiments, after obtaining a voice command input by a user and identifying the voice command through an artificial intelligence large model to obtain the user intention, the method further includes: in response to identifying multiple user intentions for a single-sentence voice command, outputting a first prompt message for guiding the user to select at least one user intention from the multiple user intentions; obtaining selection information input by the user, and determining a single user intention of the voice command from the multiple user intentions according to the selection information.

[0005] In some embodiments, before matching the user intention with the registered intentions of multiple services to obtain a matching result, the method further includes: obtaining the intention corresponding to a service according to the functions and business scenarios of the multiple services; registering at least one intention for each of the multiple services respectively to obtain the registered intentions of the multiple services.

[0006] In some embodiments, matching the user intention with the registered intentions of multiple services to obtain a matching result and determining a target service according to the matching result includes: matching the registered intentions of multiple execution services with the user intention; in response to there being two or more registered intentions of services among the multiple services that match the user intention, outputting a second prompt message for guiding the user to select at least one from the two or more services; obtaining reply information input by the user, determining the corresponding service according to the reply intention of the reply information, and determining the service as the target service.

[0007] In some embodiments, the registration scenario further includes scenario parameter information. Generating a scenario identifier corresponding to the user intention and performing scenario registration under a preset condition to obtain a registered scenario includes: under the preset condition, obtaining the intention type and generating a scenario identifier corresponding to the user intention; establishing a correspondence between the scenario identifier and the target service; determining scenario parameter information according to the user intention through the target service, where the scenario parameter information includes one or more of the display device information corresponding to the scenario, the identifier for receiving generative artificial intelligence data, the participant whitelist, and the registration time.

[0008] In some embodiments, in response to receiving new intention feature data of a preset type for dialogue interaction, matching the new preset type of intention feature data with the registered scenarios includes: obtaining a new voice command input by the user and identifying the new voice command through an artificial intelligence large model to obtain new intention feature data of a preset type for dialogue interaction; matching the new preset type of intention feature data with the intention types corresponding to multiple registered scenarios to obtain an intention matching result; and determining a target registered scenario from the multiple registered scenarios according to the intention matching result.

[0009] In some embodiments, determining a target registration scenario from multiple registration scenarios according to the intention matching result includes: if the intention feature data of the new preset type for dialogue interaction matches the intention types corresponding to two or more registration scenarios among the multiple registration scenarios, obtaining the parameter information of the new user intention; matching the parameter information of the intention feature data of the new preset type for dialogue interaction with the scenario parameter information of two or more registration scenarios to obtain a parameter information matching result; and determining a target scenario from two or more registration scenarios according to the parameter information matching result.

[0010] In some embodiments, the data information includes generative artificial intelligence data. In response to receiving the intention feature data of the new preset type, matching the intention feature data of the new preset type with the registration scenario includes: in response to the user intention being the intention feature data of the preset type and the preset type being to obtain data information, processing the user intention and the scenario identifier corresponding to the user intention through an artificial intelligence large model to obtain generative artificial intelligence data, where the generative artificial intelligence data is bound to the scenario identifier corresponding to the user intention; matching the scenario identifier corresponding to the user intention with the scenario identifiers of multiple registration scenarios to obtain an identifier matching result; and determining a target registration scenario from multiple registration scenarios according to the identifier matching result.

[0011] In some embodiments, the method further includes: determining the display device corresponding to the registration scenario according to the scenario parameter information of the registration scenario; and canceling the registration scenario in a preset cancellation situation, where the preset cancellation situation includes at least one of not receiving the intention feature data of the new preset type that matches the registration scenario for a preset duration and the user turning off the display device corresponding to the registration scenario.

[0012] According to a second aspect of the present disclosure, a voice interaction device is provided. The device includes: an identification unit configured to obtain a voice command input by a user, identify the voice command through an artificial intelligence large model to obtain a user intention; a matching unit configured to match the user intention with the registered intentions of multiple services to obtain a matching result, and determine a target service according to the matching result, the target service being used to execute at least one command corresponding to the user intention; a registration unit configured to generate a scene identifier corresponding to the user intention and perform scene registration under a preset condition to obtain a registered scene, the registered scene including: the intention type corresponding to the scene, the scene identifier, and the corresponding relationship between the scene identifier and the target service; the preset condition includes that the user intention is intention feature data of a preset type, and the preset type includes at least one of dialogue interaction and obtaining data information; a distribution unit configured to, in response to receiving new intention feature data of the preset type, match the new intention feature data of the preset type with the registered scenes. If a corresponding target registered scene is matched, the new intention feature data of the preset type is distributed to the target service corresponding to the target registered scene for execution.

[0013] According to a third aspect of the present disclosure, an electronic device is provided, including: at least one processor; and a memory in voice interaction connection with the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method of the foregoing first aspect.

[0014] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method of the foregoing first aspect.

[0015] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program which, when executed by a processor, implements the method of the foregoing first aspect.

[0016] The voice interaction method provided by the embodiments of the present disclosure includes obtaining a voice command input by a user, recognizing the voice command through an artificial intelligence large model to obtain the user intention; matching the user intention with the registered intentions of multiple services to obtain a matching result, and determining a target service according to the matching result, where the target service is used to execute at least one command corresponding to the user intention; generating a scene identifier corresponding to the user intention and performing scene registration under a preset condition to obtain a registered scene, where the registered scene includes: the intention type corresponding to the scene, the scene identifier, and the corresponding relationship between the scene identifier and the target service; the preset condition includes that the user intention is intention feature data of a preset type, and the preset type includes at least one of dialogue interaction and obtaining data information; in response to receiving new intention feature data of the preset type, matching the new intention feature data of the preset type with the registered scene, and if a corresponding target registered scene is matched, distributing the new intention feature data of the preset type to the target service corresponding to the target registered scene for execution according to the scene identifier of the target registered scene. The method of the present disclosure realizes introducing the artificial intelligence large model into the vehicle-mounted voice interaction throughout the vehicle and at all times, and distributing the user intention and the generative content to the appropriate executor based on scene registration, while satisfying the voice intention distribution during the vehicle-mounted interaction throughout the vehicle and at all times, and solving the problem of distributing the generative content of the large model.

[0017] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0019] Figure 1 It is a schematic flowchart of a voice interaction method provided by an embodiment of the present disclosure;

[0020] Figure 2 It is a schematic flowchart of a voice interaction method provided by an embodiment of the present disclosure;

[0021] Figure 3 It is an example diagram of a voice interaction system provided by an embodiment of the present disclosure;

[0022] Figure 4 It is a schematic flowchart of a voice interaction method provided by an embodiment of the present disclosure;

[0023] Figure 5 It is a schematic flowchart of a voice interaction method provided by an embodiment of the present disclosure;

[0024] Figure 6 It is a schematic structural diagram of a voice interaction device provided by an embodiment of the present disclosure;

[0025] Figure 7 Schematic block diagram of the exemplary electronic device 600 provided by the embodiments of the present disclosure. Detailed implementation manners

[0026] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0027] The current in-vehicle voice has developed to the ability of full-vehicle, full-time, and multi-person interaction, that is, voice commands can be issued at any time and any position (multiple sound zones) in the vehicle, and multiple people can also participate in the same conversation; the voice command data containing intentions needs to be sent to the appropriate execution side to complete the execution of the command; with the obvious advantages of large language models in fields represented by intelligent question answering, it is an inevitable trend to introduce it into in-vehicle voice.

[0028] The present disclosure proposes a voice interaction system and method, which introduces generative artificial intelligence into in-vehicle voice, realizes the distribution of user intentions under full-vehicle, full-time interaction, and at the same time, is compatible with the distribution of generative content of large language models.

[0029] The following describes in detail a voice interaction method, device, electronic device, storage medium, and program product proposed by the present disclosure with reference to the accompanying drawings.

[0030] Figure 1 A voice interaction method provided by the embodiments of the present disclosure, which is applied to a voice interaction system, and its execution subject can be a processor. Among them, the processor can be, for example, the processor corresponding to mobile devices such as mobile phones, computers, and tablets, the processor corresponding to intelligent wearable devices such as smart glasses and smart watches, and can also be the processor of vehicles and other means of transportation, for example, the main control chip carried by the vehicle's in-vehicle system.

[0031] As Figure 1 shown, the method includes the following steps:

[0032] Step 101, obtain the voice command input by the user, and recognize the voice command through an artificial intelligence large model to obtain the user intention.

[0033] In some embodiments of the present disclosure, the large artificial intelligence model includes a large language model (LLM), ChatGPT (Chat Generative Pre-trained Transformer), a multimodal large model, a multimodal cognitive large model, etc.

[0034] In an embodiment of the present disclosure, the large artificial intelligence model for speech recognition can process a variety of natural language tasks, such as text classification, question answering, dialogue, etc., to meet the needs of speech interaction in multiple scenarios, such as intelligent customer service, smart home, and autonomous driving.

[0035] In some embodiments, in addition to obtaining the user's intention through the large artificial intelligence model, for scenarios that require receiving multi-round voice commands or dialogues, the large artificial intelligence model can also obtain generative artificial intelligence data to achieve intelligent question answering. The generative artificial intelligence data includes pictures, texts, audio, etc.

[0036] For example, in the full-time voice interaction of the whole vehicle, such as the co-pilot saying "I want to listen to cross-talk" and the rear row saying "What's the difference between L8 and L9", after the large artificial intelligence model obtains and processes the voice commands, it can obtain corresponding user intentions such as the intention to play cross-talk and the intention to have a dialogue. For the co-pilot's intention to play cross-talk, a prompt message "Which one do you want to listen to" is generated to guide the user to select a specific playback target. For the "What's the difference between L8 and L9" said by the rear row, corresponding dialogue information such as "The difference between L8 and L9 is..." is generated to achieve intelligent question answering.

[0037] Step 102: Match the user intention with the registered intentions of multiple services to obtain a matching result, and determine the target service according to the matching result. The target service is used to execute at least one instruction corresponding to the user intention.

[0038] In some embodiments, after the voice command is understood by the large artificial intelligence model, it is necessary to find the corresponding service to execute. This is achieved by matching the user intention with the registered intentions of multiple services, and the user intention is assigned to the target service.

[0039] In some embodiments, in the in-vehicle voice interaction scenario, such as the driver saying "Navigate to Tiananmen", the co-pilot saying "I want to listen to cross-talk", and the rear row saying "What's the difference between L8 and L9", etc., after these voice commands are understood by the large artificial intelligence model, it is necessary to find the corresponding service to execute. For the driver, a navigation application is required to execute the intention of navigating to Tiananmen. For the co-pilot's command, a media application is required to execute the intention of playing cross-talk. For the rear row, a business processing unit is required to receive the dialogue text or voice reply generated by the large artificial intelligence model and display it to the rear row users.

[0040] Step 103, generate a scene identifier corresponding to the user intention under preset conditions and perform scene registration to obtain a registered scene. The registered scene includes: the intention type corresponding to the scene, the scene identifier, and the corresponding relationship between the scene identifier and the target service; the preset conditions include that the user intention is intention feature data of a preset type, and the preset type includes at least one of dialogue interaction and obtaining data information.

[0041] In some embodiments, the dialogue interaction type includes continuously receiving the user's voice instructions. For example, when the user says "I want to watch a movie" and multiple movies are searched, and the user is prompted to indicate which one they want to watch, at this time, instructions similar to "the first one" or "the movie 'Three Teams'" need to be received from the user.

[0042] In some embodiments, the type of obtaining data information includes obtaining relevant data information according to the current intention feature data. For example, according to the question "What is the difference between L8 and L9", the answer "The difference between L8 and L9 is..." is obtained, and according to the intention "What does a corgi look like", a picture of a corgi is obtained, etc.

[0043] In some embodiments, after the user intention is distributed to the target service, during the execution process of the user intention, it can be determined whether the user intention is intention feature data of a preset type by judging whether the execution of the user intention requires obtaining further intentions from the user and whether the execution of the user intention requires receiving generative artificial intelligence data, etc.

[0044] In some embodiments, the intention type corresponding to the scene is determined according to the instruction intention expected to be received in this scene.

[0045] For example, when the user says "I want to listen to cross-talk", the media service performs a search for cross-talk and obtains multiple results, and the user needs to make a selection. The media service hopes to receive instructions such as "the first one" or "next page" from the user. At this time, the intention type corresponding to the registered scene is the selection intention.

[0046] In some embodiments, when the execution of the user intention requires obtaining further intentions from the user, for example, when the execution of the user intention requires the user to make a further selection, a scene is registered for this user intention, and a scene identifier and the intention type corresponding to the scene are generated.

[0047] For example, the driver says "Navigate to Tiananmen". The intention of the driver to navigate to Tiananmen is distributed to the navigation service, and a navigation route selection interface is displayed. The navigation needs to obtain further intentions from the user, such as which route the user wants to choose. At this time, a scene is registered for the driver's intention to navigate to Tiananmen, a scene identifier and the intention type corresponding to the scene are generated, and the scene identifier is bound to the navigation service.

[0048] In some embodiments, when the execution of a user intention requires receiving generative AI data, for example, when the execution of a user intention requires obtaining conversation information generated based on a user voice command, a scene is registered for this user intention, generating a scene identifier, the intention type corresponding to the scene is a conversation intention, an identifier for receiving generative AI data, etc.

[0049] For example, for the statement "the differences between L8 and L9" said by the rear row, the conversation intention of the rear row is assigned to the conversation service. At this time, content generated based on the question of the rear row needs to be received, such as generating "The differences between L8 and L9 are...". At this time, a scene is registered for this intention, and the scene identifier is bound to the conversation service.

[0050] Step 104, in response to receiving new intention feature data of a preset type, match the new intention feature data of the preset type with the registered scenes. If a corresponding target registered scene is matched, distribute the new intention feature data of the preset type to the target service corresponding to the target registered scene for execution according to the scene identifier of the target registered scene.

[0051] In some embodiments, in response to receiving new intention feature data of a preset type for conversation interaction and / or in response to receiving new intention feature data of a preset type for obtaining data information, scene matching is performed, such as receiving a new user intention and / or obtaining generative AI data in multi-round voice commands, etc.

[0052] In some embodiments, when a new user intention is received, it is matched with the registered scenes. If the new user intention is related to the registered scenes, it can be assigned to the target service corresponding to the registered scenes, thereby realizing the distribution of user intentions for the entire vehicle at all times.

[0053] For example, for the execution of the main driver's intention "navigate to Tiananmen", multiple routes are displayed on the navigation and the user is asked to select one. When the co-driver says "I want to watch a movie", multiple movies are displayed on the media and the user is asked which one to watch. Since the execution of both requires obtaining further user intentions, scenes are registered for them respectively, obtaining Scene 1 and Scene 2. If a new intention "I want to watch the movie 'Invincible'" is received at this time, this user intention is matched with Scene 1 and Scene 2, and the matching result shows that the matching degree of this user intention with Scene 2 is the highest, then it is assigned to the media service corresponding to this Scene 2.

[0054] In some embodiments, when the execution of a user intent requires receiving generative AI data, a scenario is registered for the user intent, a scenario identifier is generated, and the user intent and the corresponding scenario identifier are sent to a large AI model for processing to obtain generative AI data. The generative AI data is bound to the scenario identifier corresponding to the user intent. Thus, in response to receiving the generative AI data, the corresponding registered scenario is matched according to the bound scenario identifier, and the generative AI data is allocated to the target service corresponding to the registered scenario.

[0055] For example, for the execution of the rear-row intent "the difference between L8 and L9", a dialogue and question-and-answer session is required. The dialogue service registers a scenario for this user intent, generates a scenario identifier, and sends "the difference between L8 and L9" and the corresponding scenario identifier to ChatGPT for processing, generating the content "The difference between L8 and L9 is...". This generated content is bound to the scenario identifier corresponding to "the difference between L8 and L9". Thus, when this generated content is received, the corresponding registered scenario can be matched according to its bound scenario identifier, and it is allocated to the dialogue service corresponding to the registered scenario.

[0056] In summary, according to the embodiments of the present disclosure, it includes obtaining a voice command input by a user, identifying the voice command through a large AI model to obtain a user intent; matching the user intent with the registered intents of multiple services to obtain a matching result, and determining a target service according to the matching result, where the target service is used to execute at least one command corresponding to the user intent; generating a scenario identifier corresponding to the user intent and performing scenario registration under a preset condition to obtain a registered scenario, where the registered scenario includes: the intent type corresponding to the scenario, the scenario identifier, and the corresponding relationship between the scenario identifier and the target service; the preset condition includes that the user intent is intent feature data of a preset type, and the preset type includes at least one of dialogue interaction and obtaining data information; in response to receiving a new user intent and / or generative AI data, based on the new user intent and / or generative AI data being matched with the registered scenario, if a corresponding target registered scenario is matched, the new user intent and / or generative AI data are distributed to the target service corresponding to the target registered scenario for execution according to the scenario identifier of the target registered scenario. The method of the present disclosure realizes the introduction of a large AI model in voice interaction, and distributes the user intent and generative AI data to the appropriate execution party for execution based on scenario registration, solving the problem of distributing the generative content of the large model while satisfying the voice intent distribution during full-vehicle and full-time interaction.

[0057] Based on Figure 1 the embodiments shown, Figure 2 is a flowchart of a voice interaction method provided by an embodiment of the present disclosure, and this method can be applied to Figure 3The voice interaction system shown, as an example, can run in the in-vehicle system of a vehicle, for instance.

[0058] The method includes the following steps 201 - 2010.

[0059] Step 201: Obtain the voice command input by the user, and identify the voice command through an artificial intelligence large model to obtain the user intention.

[0060] In some embodiments, such as Figure 3 in the voice interaction system shown, the voice command issued by the user is identified and processed through the AI capability layer to obtain the user intention. Obtain the user intention, wherein the AI capability layer provides AI capabilities including Natural Language Understanding (NLU), Natural Language Generation (NLG), and generative artificial intelligence.

[0061] In some embodiments, after obtaining the voice command input by the user and identifying the voice command through an artificial intelligence large model to obtain the user intention, the method further includes: in response to identifying multiple user intentions for a single-sentence voice command, output a first prompt message, where the first prompt message is used to guide the user to select at least one user intention from the multiple user intentions; obtain the selection information input by the user, and determine the single user intention of the voice command from the multiple user intentions according to the selection information.

[0062] In some embodiments, such as Figure 3 shown, semantic arbitration processing is performed on the multiple user intentions corresponding to the voice command through the dialogue management (hereinafter referred to as DM) of the voice interaction system to determine the single user intention of the voice command. After determination, the intention is passed to the artificial intelligence assistant management (hereinafter referred to as CM), and is allocated by CM to the Copilot of the target service.

[0063] For example, when the user says "I want to look at the sky", there may be multiple intentions. For example, one is the vehicle control intention to open the sunroof, and the other is the movie search intention (i.e., the movie "Sky"). At this time, DM can ask the user whether they want to open the sunroof or watch a movie. If it is determined that the user intention is to open the window.

[0064] Step 202: Obtain the intention corresponding to the service according to the functions and business scenarios of multiple services.

[0065] Step 203: Register at least one intention for each of the multiple services to obtain the registered intentions of the multiple services.

[0066] In some embodiments, such as Figure 3The voice interaction system shown includes multiple artificial intelligence assistants (hereinafter referred to as Copilot), and each Copilot corresponds to a service. For example, for the navigation service, there is a navigation Copilot that manages one or more navigation applications to execute user navigation-related intents such as the intent of "navigate to Tiananmen". For the media service, there is a media Copilot that manages one or more media applications to execute user media-related intents such as the intent of "play a movie". In addition, it also includes a vehicle control Copilot, a general dialogue Copilot, etc.

[0067] In some embodiments, when an application starts up, corresponding intents are registered for the Copilot responsible for the application. For example, when a media application starts up, intents related to media search and playback control such as "play music", "I want to watch a movie", "I want to listen to cross-talk", "fast forward 30 seconds", etc. are registered for the media Copilot. For the vehicle control Copilot, intents such as "open the trunk", "turn on the reading light", "turn on the seat massage", etc. are registered for it. Other Copilots are similar and will not be elaborated here.

[0068] Step 204, match the user intent with the registered intents of multiple services to obtain a matching result, and determine a target service according to the matching result. The target service is used to execute at least one instruction corresponding to the user intent.

[0069] In some embodiments, matching the user intent with the registered intents of multiple services to obtain a matching result, and determining a target service according to the matching result includes: matching the registered intents of multiple execution services with the user intent; in response to the situation that the registered intents of two or more services among multiple services match the user intent, output a second prompt message, where the second prompt message is used to guide the user to select at least one from two or more services; obtain the reply information input by the user, determine the corresponding service according to the reply intent of the reply information, and determine the service as the target service.

[0070] In some embodiments, such as applied to Figure 3 the voice interaction system shown, when multiple Copilots register the same intent and the user triggers the intent, the human CM arbitrates the intent and asks the user which Copilot to use for execution.

[0071] For example, when the user says "I want to watch a movie", if iQIYI and Tencent Video respectively register the intent of video search as a Copilot, then the CM can ask the user which application to use for the search. It should be noted that in other implementation manners, there is also a situation where one media Copilot manages all media applications, and the present disclosure does not limit this.

[0072] Step 205: Generate a scene identifier corresponding to the user intention under preset conditions and perform scene registration to obtain a registered scene. The registered scene includes: the intention type corresponding to the scene, the scene identifier, and the correspondence between the scene identifier and the target service; the preset conditions include that the user intention is intention feature data of a preset type, and the preset type includes at least one of dialogue interaction and obtaining data information.

[0073] In some embodiments of the present disclosure, after introducing generative artificial intelligence into the full-vehicle and full-time voice interaction, distributing the user voice command to the target service for execution includes three cases. One is that the command is directly executed and the execution result is returned. The second is that the command requires further instructions from the user, including scenarios that require multi-round dialogue support. The third is that it is necessary to receive generative content generated by artificial intelligence, including scenarios that require intelligent question-and-answer support. In the latter two cases, register a scene for the intention corresponding to the user voice command.

[0074] In some embodiments of the present disclosure, when receiving further intentions from the user and / or generative content generated by artificial intelligence during the execution of the user intention, generate a scene identifier corresponding to the user intention and perform scene registration.

[0075] In some embodiments, the registered scene further includes scene parameter information. Generating a scene identifier corresponding to the user intention under preset conditions and performing scene registration to obtain a registered scene includes: under preset conditions, obtaining the intention type, generating a scene identifier corresponding to the user intention; establishing the correspondence between the scene identifier and the target service; determining the scene parameter information according to the user intention through the target service. The scene parameter information includes one or more of the display device information corresponding to the scene, the identifier for receiving generative artificial intelligence data, the participant whitelist, and the registration time.

[0076] In some embodiments, the registered scene is often accompanied by UI display, that is, associated with a screen in the vehicle.

[0077] In some embodiments, the rule for generating the scene ID is, for example: "sceneid"-ScreenId-timestamp-6-digit random number; each field is the identifier header, the screen corresponding to the scene, the timestamp, and the random number. In addition, other parameters can also be defined to achieve specific purposes. For example, the following parameters can be defined: the intention whitelist includes the intentions that Copilot hopes to receive commands at this time, the screen information includes the screen corresponding to this scene, the participant whitelist includes whether this scene is a multi-person scene, that is, whether to receive commands from others, the screen information includes the screen corresponding to this scene, the parameter indicating that this scene needs to receive generative artificial data, etc.

[0078] In some embodiments, it is applied to Figure 3The voice interaction system shown, when Copilot needs to dynamically receive intents or receive generative AI data, perform scenario registration, generate a scenario ID, the relationship between the scenario ID and Copilot will be recorded in the CM, and the scenario information will be passed to the DM.

[0079] For example, the user says "I want to listen to cross-talk". The media Copilot executes a search for cross-talk and finds multiple results. The user needs to make a selection. Then the Copilot displays the searched media resources and announces "Multiple results found. Which one do you want to listen to?". At this time, the Copilot hopes to receive instructions from the user such as "the first one", "next page", etc. Then the Copilot will perform scenario registration, generate a scenario ID such as B, and the intent whitelist includes the intents that the Copilot hopes to receive instructions at this time, such as "the first one", "next page", etc. The relationship between scenario B and the media Copilot will be recorded in the CM, and the scenario information will be passed to the DM.

[0080] For example, the user says "Compare L8 and L9", which is assigned to the general dialogue Copilot. At this time, it is hoped that the Copilot will receive an intelligent answer generated for the question. The general dialogue Copilot will perform scenario registration, generate a scenario ID such as C, indicating that this scenario needs to receive the parameter information of generative artificial data. The relationship between scenario C and the general dialogue Copilot will be recorded in the CM, and the scenario information will be passed to the DM.

[0081] Step 206, in response to receiving new preset-type intent feature data, match the new preset-type intent feature data with the registered scenarios. If a corresponding target registered scenario is matched, distribute the new preset-type intent feature data to the target service corresponding to the target registered scenario for execution according to the scenario identifier of the target registered scenario.

[0082] In some embodiments, in response to receiving new preset-type intent feature data for dialogue interaction and / or in response to receiving new preset-type intent feature data for obtaining data information, perform scenario matching, such as receiving a new user intent and / or obtaining generative AI data in a multi-round voice instruction, etc.

[0083] In some embodiments, such as Figure 3In the scenario shown, when Copilot registers a scenario, the CM is responsible for generating a scenario identifier. The relationship between the scenario identifier and Copilot is recorded in the CM, and the scenario information is passed to the DM. When there is a new user intent or AIGC is related to the scenario, the DM is responsible for associating the new user intent or generative AI data with the scenario ID and then passing it to the CM. The CM then routes the new user intent or generative AI data to the Copilot that registered the scenario based on the scenario ID.

[0084] In some embodiments, as Figure 4 shown, in response to receiving new intent feature data of a preset type of dialogue interaction, matching the new preset type of intent feature data with registered scenarios includes steps 301-303.

[0085] Step 301, obtain a new voice command input by the user, and identify the new voice command through an artificial intelligence large model to obtain new intent feature data of a preset type of dialogue interaction.

[0086] Step 302, match the new intent feature data of the preset type of dialogue interaction with the intent types corresponding to multiple registered scenarios to obtain an intent matching result.

[0087] In some embodiments, the driver says "Navigate to Tiananmen", the navigation Copilot displays a navigation route selection interface and prompts "Which route do you want to choose". The registered scenario ID is A, and the intent list is the selection intent. The co-driver says "I want to watch a movie", and the media Copilot pops up a power resource card and prompts "Multiple results found. Which one do you want to watch" and registers the scenario ID as B, and the intent list is the selection intent. If the driver says "The first one", after processing this voice command, a new intent is obtained as the selection intent. After intent matching, it can be known that this intent matches the intent lists of registered scenario A and registered scenario B.

[0088] Step 303, determine a target registered scenario from multiple registered scenarios according to the intent matching result.

[0089] In some embodiments, determining a target registered scenario from multiple registered scenarios according to the intent matching result includes: if the new intent feature data of the preset type of dialogue interaction matches the intent types corresponding to two or more registered scenarios among multiple registered scenarios, obtain the parameter information of the new intent feature data of the preset type of dialogue interaction; match the parameter information of the new intent feature data of the preset type of dialogue interaction with the scenario parameter information of two or more registered scenarios to obtain a parameter information matching result; determine the target scenario from two or more registered scenarios according to the parameter information matching result.

[0090] In some embodiments, the parameter information of the intention feature data with the new preset type of dialogue interaction includes the sound zone or user corresponding to the source of the voice command. The sound zones include the driver's seat, the co-driver's seat, the rear row, etc.

[0091] In some embodiments, when the intention whitelist corresponding to multiple scenarios includes the same intention and the user triggers the intention, the DM performs scenario arbitration on multiple scenarios, selects one of the scenarios, associates the intention with the scenario identifier of the scenario, and then passes it to the CM, and then the CM distributes it to the Copilot registered for the scenario.

[0092] In some embodiments, the DM performs scenario arbitration according to scenario parameters, including performing scenario arbitration based on one or more of the screen information, the participant whitelist, and the registration time.

[0093] In some embodiments, according to the participant whitelist, it can be determined whether the registered scenario is a multi-person scenario. If it is a multi-person scenario, instructions from others can be received, and it is judged whether the registered scenario matches according to the participant whitelist and the sound zone or user corresponding to the voice command source of the intention feature data with the new preset type of dialogue interaction.

[0094] For example, the driver says "Navigate to Tiananmen", the navigation Copilot displays the navigation route selection interface and prompts "Which route do you want to choose". The registered scenario ID is A, and the intention list is the selection intention. The co-driver says "I want to watch a movie", the media Copilot pops up the power resource card, and the tts prompts "Multiple results found. Which one do you want to watch" and registers the scenario ID as B, and the intention list is also the selection intention. At this time, the person in the rear row says "The first one". According to the intention matching result, both the registered scenarios A and B match. At this time, if the registered scenario A is not a multi-person scenario and the registered scenario B is a multi-person scenario, the selection intention of "The first one" is sent to the media Copilot corresponding to the scenario B to perform the operation of selecting and playing the first movie.

[0095] In some embodiments, the screen information includes the screen associated with the scenario, the sound zone where the screen is located. The sound zones include the driver's seat, the co-driver's seat, the rear row, etc. It is judged whether the sound zone corresponding to the voice command source of the intention feature data with the new preset type of dialogue interaction is the same as the sound zone where the screen associated with the registered scenario is located. If it is the same sound zone, it is determined that the registered scenario is the target scenario.

[0096] For example, the driver says "Navigate to Tiananmen", and the navigation Copilot displays the navigation route selection interface, prompting "Which route do you want to select". The registered scene ID is A, and the intent list is the selection intent. The rear passenger says "I want to watch a movie", and the media Copilot pops up the power resource card, with the tts prompting "Multiple results found. Which one do you want to watch" and registering the scene ID as B, and the intent list is also the selection intent. At this time, the driver says "The first one". According to the intent matching result, both registered scenes A and B are matched. At this time, if the screen associated with registered scene A is in the driver's seat and the screen associated with registered scene B is in the rear row, then the selection intent of "The first one" is sent to the navigation Copilot corresponding to scene A, and the operation of selecting the first route is executed to start navigation, completing the instruction execution.

[0097] In some embodiments, if the new user intent matches the intent types corresponding to two or more registered scenes among multiple registered scenes, according to the registration time of the registered scenes, the registered scene with the earlier registration time is determined as the target scene.

[0098] In some embodiments, as Figure 5 shown, the data information includes generative artificial intelligence data. In response to receiving the intent feature data of a new preset type, the matching of the intent feature data of the new preset type with the registered scenes includes steps 401 - 403.

[0099] Step 401, in response to the user intent being the intent feature data of a preset type and the preset type being to obtain data information, the user intent and the scene identifier corresponding to the user intent are processed by an artificial intelligence large model to obtain generative artificial intelligence data, where the generative artificial intelligence data is bound to the scene identifier corresponding to the user intent.

[0100] Step 402, the scene identifier corresponding to the user intent is matched with the scene identifiers of multiple registered scenes to obtain the identifier matching result.

[0101] Step 403, according to the identifier matching result, the target registered scene is determined from multiple registered scenes.

[0102] In some embodiments, as Figure 3In the voice interaction system shown, when the person in the back row says "Compare L8 and L9", it is necessary to receive generative AI data. The general dialogue Copilot registers the scenario ID as C for this intention, and sends this intention and the registered scenario ID to GTP for processing to generate comparison information of L8 and L9. The comparison information will be marked as AI data with scenario ID C in the DM. If the current registered scenarios include scenario A corresponding to the navigation Copilot, scenario B corresponding to the media Copilot, and scenario C corresponding to the general dialogue Copilot, then according to the marked scenario ID C, the AI data will be distributed by the CM to the general dialogue Copilot.

[0103] Step 207, in a preset cancellation situation, cancel the registered scenario, where the preset cancellation situation includes not receiving new intention feature data of a preset type that matches the registered scenario for a preset period of time, and the user turning off at least one of the display devices corresponding to the registered scenario.

[0104] In some embodiments, determine the display device corresponding to the registered scenario according to the scenario parameter information of the registered scenario.

[0105] In some embodiments, when the Copilot does not receive the AI layer processing result (including new user intentions and / or generative AI data) within a preset time, and when the user turns off the display device corresponding to the registered scenario, the scenario can be cancelled; for example, when the user says "I want to watch a movie", the screen displays the UI, shows the searched movie resources, and asks the user which one to watch. At this time, if the user does not issue a subsequent instruction for a long time, the UI interface disappears, or the user directly closes the UI interface. At this time, the Copilot cancels the scenario.

[0106] In summary, according to the embodiments of the present disclosure, the distribution of AI capabilities based on scenarios realizes the full-time and multi-person interaction of in-vehicle voice for the whole vehicle, and at the same time supports the distribution of generative content of large models.

[0107] Corresponding to the above voice interaction method, the present disclosure also proposes a voice interaction device. Figure 6 It is a schematic structural diagram of a voice interaction device 500 provided by an embodiment of the present disclosure. As Figure 4 shown, it includes:

[0108] An identification unit 510, configured to obtain a voice command input by a user, identify the voice command through an artificial intelligence large model, and obtain a user intention; a matching unit 520, configured to match the user intention with the registered intentions of multiple services to obtain a matching result, and determine a target service according to the matching result, where the target service is used to execute at least one command corresponding to the user intention; a registration unit 530, configured to generate a scenario identifier corresponding to the user intention and perform scenario registration under a preset condition to obtain a registered scenario, where the registered scenario includes: an intention type corresponding to the scenario, a scenario identifier, and a correspondence between the scenario identifier and the target service; the preset condition includes that the user intention is intention feature data of a preset type, and the preset type includes at least one of dialogue interaction and obtaining data information; a distribution unit 540, configured to, in response to receiving new intention feature data of the preset type, match the new intention feature data of the preset type with the registered scenario, and if a corresponding target registered scenario is matched, distribute the new intention feature data of the preset type to the target service corresponding to the target registered scenario for execution.

[0109] In some embodiments, the device further includes a prompt unit. After obtaining the voice command input by the user and identifying the voice command through the artificial intelligence large model to obtain the user intention, the arbitration module is configured to: in response to identifying multiple user intentions for a single-sentence voice command, output a first prompt message, where the first prompt message is used to guide the user to select at least one user intention from the multiple user intentions; obtain the selection information input by the user, and determine a single user intention of the voice command from the multiple user intentions according to the selection information.

[0110] In some embodiments, before matching the user intention with the registered intentions of multiple services to obtain a matching result, the registration module 530 is further configured to: obtain the intentions corresponding to the services according to the functions and business scenarios of the multiple services; register at least one intention for each of the multiple services to obtain the registered intentions of the multiple execution services.

[0111] In some embodiments, the matching unit 520 is specifically configured to: match the registered intentions of the multiple execution services with the user intention; in response to the registered intentions of two or more services among the multiple services matching the user intention, output a second prompt message, where the second prompt message is used to guide the user to select at least one from the two or more services; obtain the reply information input by the user, determine the corresponding service according to the reply intention of the reply information, and determine the service as the target service.

[0112] In some embodiments, the registration scenario further includes scenario parameter information. The registration unit 530 is specifically configured to: under preset conditions, obtain the intent type, generate a scenario identifier corresponding to the user intent; establish a correspondence between the scenario identifier and the target service; and determine the scenario parameter information according to the user intent through the target service. The scenario parameter information includes one or more of the following: display device information corresponding to the scenario, an identifier for receiving generative AI data, a whitelist of participants, and the registration time.

[0113] In some embodiments, the distribution unit 540 is specifically configured to: obtain a new voice command input by the user, identify the new voice command through an AI large model to obtain intent feature data of a new preset type for dialogue interaction; match the intent feature data of the new preset type for dialogue interaction with the intent types corresponding to multiple registered scenarios to obtain an intent matching result; and determine a target registered scenario from the multiple registered scenarios according to the intent matching result.

[0114] In some embodiments, the distribution unit 540 is further configured to: if the intent feature data of the new preset type for dialogue interaction matches the intent types corresponding to two or more registered scenarios among the multiple registered scenarios, obtain the parameter information of the intent feature data of the new preset type for dialogue interaction; match the parameter information of the intent feature data of the new preset type for dialogue interaction with the scenario parameter information of the two or more registered scenarios to obtain a parameter information matching result; and determine a target scenario from the two or more registered scenarios according to the parameter information matching result.

[0115] In some embodiments, the distribution unit 540 is further configured to: in response to the user intent being intent feature data of a preset type and the preset type being to obtain data information, process the user intent and the scenario identifier corresponding to the user intent through an AI large model to obtain generative AI data, where the generative AI data is bound to the scenario identifier corresponding to the user intent; match the scenario identifier corresponding to the user intent with the scenario identifiers of the multiple registered scenarios to obtain an identifier matching result; and determine a target registered scenario from the multiple registered scenarios according to the identifier matching result.

[0116] In some embodiments, the device further includes a cancellation unit. The cancellation unit is configured to: determine the display device corresponding to the registered scenario according to the scenario parameter information of the registered scenario; and cancel the registered scenario under preset cancellation circumstances, where the preset cancellation circumstances include at least one of not receiving new preset type intent feature data matching the registered scenario for a preset duration and the user turning off the display device corresponding to the registered scenario.

[0117] In summary, according to the embodiments of the present disclosure, the device distributes AI capabilities based on scenarios through the recognition unit, matching unit, registration unit, and distribution unit, achieving full-time and multi-person interaction of in-vehicle voice throughout the vehicle, while supporting the distribution of generative content of large models.

[0118] It should be noted that since the device embodiments of the present disclosure correspond to the above method embodiments, the foregoing explanations of the method embodiments also apply to the devices of this embodiment. The principles are the same. For details not disclosed in the device embodiments, reference may be made to the above method embodiments, and no further elaboration will be provided in the present disclosure.

[0119] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0120] Figure 7 FIG. shows a schematic block diagram of an exemplary electronic device 600 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0121] As Figure 7 shown, the device 600 includes a computing unit 601, which can execute various appropriate actions and processes according to the computer program stored in the ROM (Read-Only Memory) 602 or the computer program loaded from the storage unit 608 into the RAM (Random Access Memory) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The I / O (Input / Output) interface 605 is also connected to the bus 604.

[0122] Multiple components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disk, etc.; and a voice interaction unit 609, such as a network card, a modem, a wireless voice interaction transceiver, etc. The voice interaction unit 609 allows the device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0123] The computing unit 601 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as a voice interaction method. For example, in some embodiments, the voice interaction method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 600 via the ROM 602 and / or the voice interaction unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to execute the aforementioned voice interaction method in any other appropriate manner (eg, by means of firmware).

[0124] The various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SoCs (System On Chip), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0125] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on a remote machine or server.

[0126] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only Memory), or a flash memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0127] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0128] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data voice interaction (e.g., a voice interaction network). Examples of voice interaction networks include: a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, and a blockchain network.

[0129] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a voice interaction network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS" for short). The server may also be a server of a distributed system or a server combined with blockchain.

[0130] Among them, it should be noted that artificial intelligence is a discipline that studies how to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and there are both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0131] It should be understood that various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recorded in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0132] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A voice interaction method, characterized in that, The method includes: Obtain a voice command input by a user, and identify the voice command through an artificial intelligence large model to obtain the user intention; Match the user intention with the registered intentions of multiple services to obtain a matching result, and determine a target service according to the matching result. The target service is used to execute at least one instruction corresponding to the user intention; Generate a scene identifier corresponding to the user intention and perform scene registration under preset conditions to obtain a registered scene. The registered scene includes: the intention type corresponding to the scene, the scene identifier, and the correspondence between the scene identifier and the target service; the preset conditions include that the user intention is intention feature data of a preset type, and the preset type includes at least one of dialogue interaction and obtaining data information; In response to receiving new intention feature data of the preset type, match the new intention feature data of the preset type with the registered scene. If a corresponding target registered scene is matched, distribute the new intention feature data of the preset type to the target service corresponding to the target registered scene for execution according to the scene identifier of the target registered scene.

2. The method according to claim 1, characterized in that, After obtaining the voice command input by the user and identifying the voice command through the artificial intelligence large model to obtain the user intention, the method further includes: In response to identifying multiple user intentions for the single-sentence voice command, output a first prompt message for guiding the user to select at least one of the multiple user intentions; Obtain the selection information input by the user, and determine the single user intention of the voice command from the multiple user intentions according to the selection information.

3. The method according to claim 1, characterized in that, Before matching the user intention with the registered intentions of multiple services to obtain a matching result, the method further includes: Obtain the intentions corresponding to the services according to the functions and business scenarios of the multiple services; Register at least one intention for each of the multiple services to obtain the registered intentions of the multiple services.

4. The method according to claim 3, characterized in that, The matching of the user intention with the registered intentions of multiple services to obtain a matching result and determining the target service according to the matching result includes: Match the registered intentions of the multiple execution services with the user intention; In response to the situation that the registered intentions of two or more services among the multiple services match the user intention, output a second prompt message for guiding the user to select at least one from the two or more services; Obtain the reply information input by the user, determine the corresponding service according to the reply intention of the reply information, and determine the service as the target service.

5. The method according to claim 1, characterized in that, The registered scene further includes scene parameter information. The generating a scene identifier corresponding to the user intention and performing scene registration under preset conditions to obtain a registered scene includes: Under preset conditions, obtain the intention type and generate a scene identifier corresponding to the user intention; Establish the correspondence between the scene identifier and the target service; The target service determines scenario parameter information according to the user intention, and the scenario parameter information includes one or more of the following: display device information corresponding to the scenario, an identifier for receiving generative artificial intelligence data, a participant whitelist, and a registration time.

6. The method according to claim 1, characterized in that, The matching of the new intention feature data of the preset type of dialogue interaction based on the received new intention feature data with the registered scenarios includes: Obtain a new voice command input by the user, and identify the new voice command through the artificial intelligence large model to obtain new intention feature data of the preset type of dialogue interaction; Match the new intention feature data of the preset type of dialogue interaction with the intention types corresponding to multiple registered scenarios to obtain an intention matching result; Determine a target registered scenario from the multiple registered scenarios according to the intention matching result.

7. The method according to claim 6, characterized in that, The determining of the target registered scenario from the multiple registered scenarios according to the intention matching result includes: If the new intention feature data of the preset type of dialogue interaction matches the intention types corresponding to two or more registered scenarios among the multiple registered scenarios, obtain the parameter information of the new user intention; Match the parameter information of the new intention feature data of the preset type of dialogue interaction with the scenario parameter information of the two or more registered scenarios to obtain a parameter information matching result; Determine a target scenario from the two or more registered scenarios according to the parameter information matching result.

8. The method according to claim 1, characterized in that, The data information includes generative artificial intelligence data, and the matching of the new intention feature data of the preset type based on the received new intention feature data with the registered scenarios includes: In response to the user intention being intention feature data of a preset type and the preset type being to obtain data information, process the user intention and the scenario identifier corresponding to the user intention through the artificial intelligence large model to obtain generative artificial intelligence data, where the generative artificial intelligence data is bound to the scenario identifier corresponding to the user intention; Match the scenario identifier corresponding to the user intention with the scenario identifiers of multiple registered scenarios to obtain an identifier matching result; Determine a target registered scenario from the multiple registered scenarios according to the identifier matching result.

9. The method according to claim 5, characterized in that,The method further includes: Determine the display device corresponding to the registered scenario according to the scenario parameter information of the registered scenario; Under a preset cancellation condition, cancel the registered scenario, where the preset cancellation condition includes not receiving new intention feature data of the preset type that matches the registered scenario for a preset period of time, and the user turning off at least one of the display devices corresponding to the registered scenario.

10. A voice interaction device, characterized in that, The device includes: An identification unit, configured to obtain a voice command input by the user, and identify the voice command through an artificial intelligence large model to obtain a user intention; A matching unit, configured to match the user intention with the registered intentions of multiple services to obtain a matching result, and determine a target service according to the matching result, where the target service is used to execute at least one instruction corresponding to the user intention; A registration unit, configured to generate a scenario identifier corresponding to the user intention and perform scenario registration under a preset condition to obtain a registered scenario, where the registered scenario includes: the intention type corresponding to the scenario, the scenario identifier, and the correspondence between the scenario identifier and the target service; the preset condition includes that the user intention is intention feature data of a preset type, and the preset type includes at least one of dialogue interaction and obtaining data information; A distribution unit, configured to, in response to receiving new intention feature data of the preset type, match the new intention feature data of the preset type with the registered scenario. If a corresponding target registered scenario is matched, distribute the intention feature data of the preset type to the target service corresponding to the target registered scenario for execution according to the scenario identifier of the target registered scenario.

11. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-9.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-9.

13. A computer program product, comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Voice processing method, distributed system, voice interaction device and method

    CN112652301A

  • Voice recognition method and device, electronic equipment, medium and program product

    CN113051895A

  • Interaction control method and device, intelligent voice equipment and storage medium

    CN114356275A

  • Voice instruction processing method and device, electronic equipment, vehicle and storage medium

    CN115424610A

  • Vehicle-mounted voice instruction recommendation method and device and model training method

    CN115547302A