Voice control method and device, vehicle, electronic equipment and storage medium
By responding to voice request commands and registering processing results in the vehicle voice system, the problem of command identification restrictions in multiple rounds of conversations is solved, and multiple rounds of conversations are supported through multi-terminal interaction to meet the functional needs of different users and improve the user experience.
Patent Information
- Application Number
- CN202311747920.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-18
- Publication Date
- 2025-06-20
AI Technical Summary
After entering multiple rounds of dialogue, the existing vehicle voice system refuses to recognize other commands other than selection, page turn, and confirmation, and the user experience is poor. At the same time, in the multi-terminal interaction scenario, a new voice interaction interface cannot be opened, limiting the user's functional needs to control the vehicle through voice.
By responding to the voice request command, the corresponding vertical field is found according to the correspondence between the request command and the vertical field, the list page corresponding to the voice request command is registered with the dialogue management module, and the registration processing results are output at the display terminal. At the same time, the target operation instructions are received and executed, and the display terminal identifier is carried to support parallel multi-round dialogue in multi-terminal interactive scenarios.
It realizes the multi-round or single-round dialogue request instructions of different users in multiple rounds of dialogue, and supports multiple single-round or multi-round dialogue through multi-terminal interaction, meeting the different functional needs of different users and improving user experience.
Smart Images

Figure CN120183391A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of vehicles, and particularly to a voice control method and device, a vehicle, an electronic device, and a storage medium. Background Art
[0002] In-vehicle voice is an in-vehicle intelligent interaction system that can realize human-machine dialogue and vehicle control through technologies such as speech recognition, natural language understanding, and speech synthesis. Compared with traditional in-vehicle voice systems, modern in-vehicle voice systems can achieve multi-turn conversations, that is, after each round of conversation, even if the voice interface exits, within a certain period of time, users can send control instructions to the vehicle computer again through the return results of voice control.
[0003] The multi-turn voice in related technologies is a closed-domain multi-turn, that is, after entering the multi-turn voice, the vehicle receives instructions such as selection, paging, and confirmation, and refuses to recognize other instructions, resulting in a poor user experience. And when the in-vehicle voice system enters a multi-terminal interaction scenario, passengers in the vehicle can participate in the current multi-turn voice conversation to control the vehicle, but cannot start a new multi-turn voice conversation to control other functions of the vehicle, thus failing to meet the different functional requirements of different users through multi-turn voice conversations. Summary of the Invention
[0004] The present disclosure provides a voice control method and device, an electronic device, and a storage medium. Its main purpose is to solve the problems that after entering the multi-turn voice, the vehicle receives instructions such as selection, paging, and confirmation, refuses to recognize other instructions, resulting in a poor user experience, and when the in-vehicle voice system is in a multi-terminal interaction scenario, other display terminals in the vehicle cannot start a new voice interaction interface, and thus cannot control the vehicle by voice.
[0005] According to a first aspect of the present disclosure, a voice control method is provided, which includes:
[0006] In response to a first voice request instruction, find a first vertical domain corresponding to the first voice request instruction according to the corresponding relationship between the request instruction and the vertical domain;
[0007] Register a first list page corresponding to the first voice request instruction to a dialogue management module based on the first vertical domain, and output a first registration processing result of the first list page on a first display terminal and other display terminals, where the dialogue management module is used to store the corresponding relationship between the request instruction and the vertical domain and the list page, and the first display terminal is the display terminal that triggers the first voice request instruction;
[0008] In response to a second voice request instruction, find a second vertical domain corresponding to the second voice request instruction according to the corresponding relationship between the request instruction and the vertical domain;
[0009] Register the second list page corresponding to the second voice request instruction to the dialogue management module based on the second vertical domain, and output the second registration processing result of the second list page on the second display terminal, where the second display terminal is the display terminal that triggers the second voice request instruction;
[0010] Receive and execute a target operation instruction for the first registration processing result or the second registration processing result, where the target operation instruction carries a first display terminal identifier or a second display terminal identifier for the first registration processing result or the second registration processing result.
[0011] Optionally, before receiving and executing the target operation instruction for the first registration processing result or the second registration processing result, the method includes:
[0012] In response to a third voice request instruction, execute the third voice request instruction, where the third voice request instruction is a voice request instruction that does not need to output the corresponding registration processing result and is used to directly control the vehicle state.
[0013] Optionally, the step of finding the first vertical domain corresponding to the first voice request instruction according to the correspondence between the request instruction and the vertical domain includes:
[0014] Identify the first voice request instruction as the corresponding target semantic intent information;
[0015] Find the first vertical domain corresponding to the target semantic intent information according to the correspondence between the semantic intent information and the vertical domain;
[0016] Send the first voice request instruction to the first vertical domain.
[0017] Optionally, the step of registering the first list page corresponding to the first voice request instruction to the dialogue management module based on the first vertical domain includes:
[0018] Find the first list page corresponding to the first voice request instruction based on the first vertical domain;
[0019] Register the first list page to the dialogue management module.
[0020] Optionally, the step of receiving and executing the target operation instruction for the first registration processing result or the second registration processing result includes:
[0021] Receive a target operation instruction, where the target operation instruction includes at least one of a voice instruction, a touch instruction, or a gaze instruction;
[0022] Determine the first display terminal identifier or the second display terminal identifier carried by the target operation instruction;
[0023] If it is determined that the display terminal identifier is the first display terminal identifier, then determine the first registration processing result as the target registration processing result;
[0024] If it is determined that the display terminal identifier is the second display terminal identifier, then determine the second registration processing result as the target registration processing result;
[0025] Execute the target operation instruction on the target registration processing result.
[0026] Optionally, the executing the target operation instruction on the target registration processing result includes:
[0027] When it is determined that the second registration processing result is the target registration processing result, after executing the target operation instruction on the target registration processing result, the second display terminal displays the first registration processing result of the first list page.
[0028] Optionally, the determining to cancel or register a new list page in response to the first list page according to the change of the first list page includes:
[0029] Judge the target category of the voice request instruction. The first registration processing result has the control authority for the first category of voice request instructions, and the second registration processing result has the control authority for the second category of voice request instructions. The first category includes the second category;
[0030] If it is determined that the target category of the voice request instruction does not belong to the second category, then control the first registration processing result to execute the voice request instruction.
[0031] According to a second aspect of the present disclosure, there is provided a voice control device, including:
[0032] A first search unit, configured to, in response to a first voice request instruction, search for a first vertical domain corresponding to the first voice request instruction according to the correspondence between the request instruction and the vertical domain;
[0033] A first registration unit, configured to register a first list page corresponding to the first voice request instruction into a dialogue management module based on the first vertical domain;
[0034] A first output unit, configured to output a first registration processing result of the first list page on a first display terminal and other display terminals. The dialogue management module is configured to store the correspondence between the request instruction and the vertical domain and the list page. The first display terminal is the display terminal that triggers the first voice request instruction;
[0035] A second search unit, configured to, in response to a second voice request instruction, search for a second vertical domain corresponding to the second voice request instruction according to the correspondence between the request instruction and the vertical domain;
[0036] A second registration unit, configured to register a second list page corresponding to the second voice request instruction into a dialogue management module based on the second vertical domain;
[0037] A second output unit, configured to output a second registration processing result of the second list page on a second display terminal, where the second display terminal is the display terminal that triggers the second voice request instruction;
[0038] A first execution unit, configured to receive and execute a target operation instruction for the first registration processing result or the second registration processing result, where the target operation instruction carries a first display terminal identifier or a second display terminal identifier for the first registration processing result or the second registration processing result.
[0039] In an embodiment of the present application, the device further includes a second execution unit, and the second execution unit is configured to:
[0040] Before receiving and executing the target operation instruction for the first registration processing result or the second registration processing result, in response to a third voice request instruction, execute the third voice request instruction, where the third voice request instruction is a voice request instruction that does not need to output a corresponding registration processing result and is used to directly control the vehicle state.
[0041] In an embodiment of the present application, the first search unit includes:
[0042] An identification module, configured to identify the first voice request instruction as corresponding target semantic intention information;
[0043] A search module, configured to search for a first vertical domain corresponding to the target semantic intention information according to the correspondence between the semantic intention information and the vertical domain;
[0044] A sending module, configured to send the first voice request instruction to the first vertical domain.
[0045] Optionally, the first registration unit includes:
[0046] An acquisition module, configured to acquire a first list page corresponding to the first voice request instruction based on the first vertical domain;
[0047] A registration module, configured to register the first list page into the dialogue management module.
[0048] In an embodiment of the present application, the first execution unit includes:
[0049] A receiving module, configured to receive a target operation instruction, where the target operation instruction includes at least one of a voice instruction, a touch instruction, or a gaze instruction;
[0050] A first determination module, configured to determine a first display terminal identifier or a second display terminal identifier carried by the target operation instruction;
[0051] A second determination module, configured to, if it is determined that the display terminal identifier is the first display terminal identifier, determine the first registration processing result as the target registration processing result;
[0052] A third determination module, configured to, if it is determined that the display terminal identifier is the second display terminal identifier, determine the second registration processing result as the target registration processing result;
[0053] An execution module, configured to execute the target operation instruction on the target registration processing result.
[0054] In an embodiment of the present application, the device further includes a judgment unit, and the judgment unit includes:
[0055] A judgment module, configured to judge the target category of the voice request instruction before looking up the first vertical domain corresponding to the first voice request instruction according to the correspondence between the request instruction and the vertical domain, where the first registration processing result has the control authority for the first category of voice request instructions, and the second registration processing result has the control authority for the second category of voice request instructions, and the first category includes the second category;
[0056] A control module, configured to, if it is determined that the target category of the voice request instruction does not belong to the second category, control the first registration processing result to execute the voice request instruction.
[0057] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0058] At least one processor; and
[0059] A memory communicatively connected to the at least one processor; wherein,
[0060] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in the foregoing first aspect.
[0061] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause the computer to execute the method described in the foregoing first aspect.
[0062] According to a fifth aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the method described in the foregoing first aspect.
[0063] The voice control method, device, vehicle, electronic device and storage medium provided by the present disclosure, in response to a first voice request instruction, find a first vertical domain corresponding to the first voice request instruction according to the correspondence between the request instruction and the vertical domain; register a first list page corresponding to the first voice request instruction to a dialogue management module based on the first vertical domain, and output a first registration processing result of the first list page on a first display terminal and other display terminals, where the dialogue management module is used to store the correspondence between the request instruction and the vertical domain and the list page, and the first display terminal is the display terminal that triggers the first voice request instruction; in response to a second voice request instruction, find a second vertical domain corresponding to the second voice request instruction according to the correspondence between the request instruction and the vertical domain; register a second list page corresponding to the second voice request instruction to the dialogue management module based on the second vertical domain, and output a second registration processing result of the second list page on a second display terminal, where the second display terminal is the display terminal that triggers the second voice request instruction; receive and execute a target operation instruction for the first registration processing result or the second registration processing result, where the target operation instruction carries a first display terminal identifier or a second display terminal identifier for the first registration processing result or the second registration processing result. Compared with the related art, the voice control method provided by the present disclosure can realize that in entering a multi-round dialogue, it can continue to respond to multi-round dialogue request instructions or single-round dialogue request instructions of different users, and support multiple single-round dialogues or multiple multi-round dialogues in parallel in a multi-terminal interaction scenario by the method of carrying a display terminal identifier in the target operation instruction to meet different needs of different users.
[0064] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0066] Figure 1 is a schematic flowchart of a voice control method provided by an embodiment of the present disclosure;
[0067] Figure 2 is a schematic diagram of the position of a display terminal provided by an embodiment of the present application;
[0068] Figure 3Schematic diagram of a voice control system architecture provided by an embodiment of the present application;
[0069] Figure 4 Schematic diagram of the structure of a voice control device provided by an embodiment of the present disclosure;
[0070] Figure 5 Schematic diagram of the structure of another voice control device provided by an embodiment of the present disclosure;
[0071] Figure 6 Schematic block diagram of an exemplary electronic device 300 provided by an embodiment of the present disclosure. Detailed implementation manners
[0072] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.
[0073] The following describes a voice control method, device, electronic device, and storage medium according to embodiments of the present disclosure with reference to the accompanying drawings.
[0074] Figure 1 Schematic flow diagram of a voice control method provided by an embodiment of the present disclosure. This method is applied to a vehicle, and the control method can be executed by a control device or equipment for information prompting. The device or equipment can be configured in a server, a processor, or a main control chip. Exemplarily, it can be deployed on the in-vehicle unit side, domain controller side, etc. of the vehicle, such as Figure 1 As shown, the method includes the following steps:
[0075] Step 101: In response to a first voice request instruction, search for a first vertical domain corresponding to the first voice request instruction according to the correspondence between the request instruction and the vertical domain.
[0076] In an embodiment of the present application, the first voice request instruction is a voice signal carrying control information received by a terminal. The vertical domain is a vertical field, which is a term in the Internet industry and refers to providing specific services for a defined group, including industries such as entertainment, medical care, environmental protection, education, and sports. In the present invention, it refers to different application modules, such as vertical fields like maps and weather.
[0077] Step 102: Based on the first vertical domain, register the first list page corresponding to the first voice request instruction to the dialogue management module, and output the first registration processing result of the first list page on the first display terminal and other display terminals. The dialogue management module is used to store the corresponding relationship between the request instruction and the vertical domain and the list page. The first display terminal is the display terminal that triggers the first voice request instruction.
[0078] In an embodiment of the present application, for better understanding of the display terminal, taking a six-seater car as an example, please refer to FIG. 2. Figure 2 It is a schematic diagram of the position of the display terminal provided by the embodiment of the present application. As Figure 2 shown, the first display terminal includes a main driver position display screen and devices associated with the main driver position voice control system. The second display terminal includes a co-driver position display screen and devices associated with the co-driver position voice control system, a rear row position display screen and devices related to the rear row position voice control system.
[0079] For example, the first voice request instruction may be "Navigate to Location A". The request instruction is the instruction information obtained after the first display terminal recognizes the received first voice request instruction. The first display terminal outputs a corresponding first registration processing result according to the first voice request instruction. The first registration processing result may be multiple different methods that can implement the request instruction obtained according to the request instruction. It can be provided to the user for selection in the form of a registration list page, but is not limited thereto, and the first registration processing result is displayed on the first display terminal and other display terminals, so that in a multi-terminal interaction scenario, users in other positions can also view and select through other display terminals.
[0080] Step 103: In response to the second voice request instruction, find the second vertical domain corresponding to the second voice request instruction according to the corresponding relationship between the request instruction and the vertical domain.
[0081] In an embodiment provided by the present application, after receiving the first voice request instruction and outputting the corresponding first registration processing result, if no operation instruction for the first registration processing result is received, the second vertical domain can also be found in response to the second voice request instruction.
[0082] Step 104: Based on the second vertical domain, register the second list page corresponding to the second voice request instruction to the dialogue management module, and output the second registration processing result of the second list page on the second display terminal. The second display terminal is the display terminal that triggers the second voice request instruction.
[0083] In an embodiment of the present application, after registering the second list page corresponding to the second voice request instruction to the dialogue management module based on the second vertical domain, the second registration processing result is output on the second display terminal that triggers the second voice request instruction. Before outputting the second registration processing result, the first registration processing result is displayed on the second display terminal.
[0084] Step 105, receive and execute a target operation instruction for the first registration processing result or the second registration processing result, where the target operation instruction carries the first display terminal identifier or the second display terminal identifier for the first registration processing result or the second registration processing result.
[0085] In an embodiment of the present application, the target operation instruction is used to control the selection of the registration processing result corresponding to the first registration processing result or the second registration processing result. Taking the example of step 101, when the registration processing result is presented in the form of a registration list page, assume that there are 6 registration processing results in the registration list page, and the target operation instruction is an operation instruction triggered by selecting any one of the 6 registration processing results.
[0086] In practical applications, the target operation instruction includes multiple triggering forms. For example, the target operation instruction can be issued through a voice instruction or by touching the interaction interface of the terminal. The specific way of issuing the target operation instruction is not limited in the embodiments of the present application. However, regardless of the form of the target operation instruction, the target operation instruction carries the corresponding display terminal identifier, and the display terminal identifier includes the first display terminal identifier and the second display terminal identifier, which is used to determine the source of the target operation instruction.
[0087] In some embodiments of the present application, the receiving and executing the target operation instruction for the first registration processing result or the second registration processing result may adopt, but is not limited to, the following method: receiving the target operation instruction, where the target operation instruction includes at least one of a voice instruction, a touch instruction, or a gaze instruction; determining the first display terminal identifier or the second display terminal identifier carried by the target operation instruction; if it is determined that the display terminal identifier is the first display terminal identifier, then determining the first registration processing result as the target registration processing result; if it is determined that the display terminal identifier is the second display terminal identifier, then determining the second registration processing result as the target registration processing result; and executing the target operation instruction on the target registration processing result.
[0088] The voice control method provided by the present disclosure, in response to a first voice request instruction, searches for a first vertical domain corresponding to the first voice request instruction according to the corresponding relationship between the request instruction and the vertical domain; registers a first list page corresponding to the first voice request instruction to a dialogue management module based on the first vertical domain, and outputs a first registration processing result of the first list page to a first display terminal and other display terminals, where the dialogue management module is used to store the corresponding relationship between the request instruction and the vertical domain and the list page, and the first display terminal is the display terminal that triggers the first voice request instruction; in response to a second voice request instruction, searches for a second vertical domain corresponding to the second voice request instruction according to the corresponding relationship between the request instruction and the vertical domain; registers a second list page corresponding to the second voice request instruction to the dialogue management module based on the second vertical domain, and outputs a second registration processing result of the second list page to a second display terminal, where the second display terminal is the display terminal that triggers the second voice request instruction; receives and executes a target operation instruction for the first registration processing result or the second registration processing result, where the target operation instruction carries a first display terminal identifier or a second display terminal identifier for the first registration processing result or the second registration processing result. Compared with the related art, the voice control method provided by the present disclosure can realize directly controlling the vehicle state by inserting a voice request instruction that does not need to output the corresponding registration processing result in a multi-round dialogue. And by the method of carrying the display terminal identifier in the target operation instruction, it is possible to support multiple single-round dialogues or multiple single-person multi-round dialogues in parallel in a multi-terminal interaction scenario to meet different needs of different users.
[0089] In some embodiments, before receiving and executing the target operation instruction for the first registration processing result or the second registration processing result, if a third voice request instruction is received, the third voice request instruction is directly executed, where the third voice request instruction does not need to output the corresponding registration processing result and is used to directly control the vehicle state. For better understanding of the third voice request instruction, illustration by way of example may be used but is not limited thereto. For example, if the first voice request instruction is "Navigate to Location A", the first display terminal and other display terminals output the first registration processing result and issue a voice prompt "Please select the nth one", if instead of inputting the target operation instruction, a third voice request instruction "Turn on seat ventilation" is input, then the request instruction "Turn on seat ventilation" is directly controlled to be executed without being displayed on any display terminal.
[0090] In some embodiments, since the control authorities of the first registration processing result and the second registration processing result are different, before looking up the first vertical domain corresponding to the first voice request instruction according to the correspondence between the request instruction and the vertical domain, it is also necessary to determine the target category of the voice request instruction. The first registration processing result has the control authority for the voice request instructions of the first category, and the second registration processing result has the control authority for the voice request instructions of the second category. The first category includes the second category. If it is determined that the target category of the voice request instruction does not belong to the second category, the first registration processing result is controlled to execute the voice request instruction. For example, taking map navigation as an example, in some embodiments, map navigation is a request instruction of the first category, while music playback is a request instruction of the second category. Therefore, map navigation can be controlled and executed based on the first registration processing result, but cannot be controlled and executed based on the second registration processing result. And music playback is a request instruction of the second category, and the first category includes the second category. Therefore, music playback can be controlled and executed based on either the first registration processing result or the second registration processing result.
[0091] In some embodiments, after obtaining the first voice request instruction, if the target operation instruction is not received within a preset time period, a prompt message is output to prompt the user to make a selection.
[0092] In some embodiments, to better understand the output of the registration processing result corresponding to the first voice request instruction in response to the first voice request instruction, taking the first voice request instruction "Navigate to Location A" as an example, please refer to Figure 3 , Figure 3 which is a schematic diagram of a system architecture for voice control provided by an embodiment of the present application. As Figure 3 shown, after the terminal receives the first voice request instruction "Navigate to Location A", based on Figure 3 Automatic Speech Recognition (ASR) in the voice interaction module shown, the first voice request instruction is recognized to obtain the corresponding text information. Based on offline and online Natural Language Understanding (NLU), the text information is parsed to obtain the semantic request corresponding to the first voice request instruction. Subsequently, the semantic request is sent to the Dialogue Manage (DM) in the voice skill service. The semantic request is arbitrated by the scenario arbitration engine in the dialogue management, and the actuator in the dialogue management calls the Agent Manager to route to the corresponding service for execution. For example, instructions containing travel navigation semantics such as "Navigate to Location A" are arbitrated to the map service for execution. The map service includesFigure 1 Hedi Figure 2 For the user to select. After obtaining the corresponding registration processing result, send the registration processing result to the scenario state machine in the dialogue management, and convert the text into speech through a text-to-speech (TTS) system to issue a prompt message, such as "Multiple results found", "Please select the nth one", etc.
[0093] To better understand the voice control process in a multi-terminal interaction scenario, taking the first voice request instruction "Navigate to Location A" as an example, when the first display terminal receives the first voice request instruction "Navigate to Location A", display the registration processing result on the first display terminal of the first display terminal and other display terminals located in the co-pilot position and the rear row position, and the first display terminal emits a prompt voice "Multiple results found, please select". Before entering the first control instruction, the second display terminal located in the co-pilot receives the second voice request instruction "Turn on the seat ventilation", and directly controls the execution of the request instruction corresponding to "Turn on the seat ventilation". Then, the second display terminal located in the co-pilot receives another first voice request instruction "I want to listen to Jacky Cheung's songs", thus opening a new voice interaction interface on the second display terminal, display the corresponding registration processing result on the second display terminal in the co-pilot position, and the second display terminal in the co-pilot position emits the corresponding voice prompt, while the first display terminal and other display terminals in the rear row position still display the registration processing result of "Navigate to Location A" unchanged. At this time, since two different voice interaction interfaces are started, control the two different registration processing results respectively according to the display terminal identifier carried by the first control instruction. For example, the first control instruction carrying the first display terminal identifier controls the registration processing result corresponding to the first display terminal, and the first control instruction carrying the co-pilot second display terminal identifier controls the registration processing result corresponding to the second display terminal in the co-pilot position. Therefore, when the in-vehicle voice system is in a multi-terminal interaction scenario, a new in-vehicle voice interaction interface can be opened on the target device, and the vehicle can be controlled by voice.
[0094] In summary, the embodiments of the present disclosure can achieve the following effects:
[0095] When implementing multi-terminal interaction scenarios by means of carrying a display terminal identifier, identify the target registration processing result corresponding to the voice request instruction, can open a new voice interaction interface on the second display terminal in a multi-terminal interaction scenario, and further realize controlling the vehicle by voice to meet the different needs of different users.
[0096] Corresponding to the above voice control method, the present invention also provides a voice control device. Since the device embodiment of the present invention corresponds to the above method embodiment, for the details not disclosed in the device embodiment, reference may be made to the above method embodiment, and no further elaboration will be made in the present invention.
[0097] Figure 4 As shown in the structure diagram of a voice control device provided by an embodiment of the present disclosure, Figure 5 it includes: a first search unit 21, a first registration unit 22, a first output unit 23, a second search unit 24, a second registration unit 25, a second output unit 26, and a first execution unit 27.
[0098] The first search unit 21 is configured to, in response to a first voice request instruction, search for a first vertical domain corresponding to the first voice request instruction according to the correspondence between the request instruction and the vertical domain;
[0099] The first registration unit 22 is configured to register a first list page corresponding to the first voice request instruction into the dialogue management module based on the first vertical domain;
[0100] The first output unit 23 is configured to output a first registration processing result of the first list page on a first display terminal, where the dialogue management module is configured to store the correspondence between the request instruction and the vertical domain and the list page, and the first display terminal is the display terminal that triggers the first voice request instruction;
[0101] The second search unit 24 is configured to, in response to a second voice request instruction, search for a second vertical domain corresponding to the second voice request instruction according to the correspondence between the request instruction and the vertical domain;
[0102] The second registration unit 25 is configured to register a second list page corresponding to the second voice request instruction into the dialogue management module based on the second vertical domain;
[0103] The second output unit 26 is configured to output a second registration processing result of the second list page on a second display terminal, where the second display terminal is the display terminal that triggers the second voice request instruction;
[0104] The first execution unit 27 is configured to receive and execute a target operation instruction for the first registration processing result or the second registration processing result, where the target operation instruction carries a first display terminal identifier or a second display terminal identifier for the first registration processing result or the second registration processing result.
[0105] The voice control device provided by the present disclosure, in response to a first voice request instruction, searches for a first vertical domain corresponding to the first voice request instruction according to the correspondence between the request instruction and the vertical domain; registers a first list page corresponding to the first voice request instruction into the dialogue management module based on the first vertical domain, and outputs a first registration processing result of the first list page on a first display terminal, where the dialogue management module is used to store the correspondence between the request instruction and the vertical domain and the list page, and the first display terminal is the display terminal that triggers the first voice request instruction; in response to a second voice request instruction, searches for a second vertical domain corresponding to the second voice request instruction according to the correspondence between the request instruction and the vertical domain; registers a second list page corresponding to the second voice request instruction into the dialogue management module based on the second vertical domain, and outputs a second registration processing result of the second list page on a second display terminal, where the second display terminal is the display terminal that triggers the second voice request instruction; receives and executes a target operation instruction for the first registration processing result or the second registration processing result, and the target operation instruction carries a first display terminal identifier or a second display terminal identifier for the first registration processing result or the second registration processing result. Compared with the related art, the voice control method provided by the present disclosure can implement directly controlling the vehicle state by inserting a voice request instruction that does not need to output the corresponding registration processing result in a multi-round dialogue. And by the method of carrying the display terminal identifier in the target operation instruction, it is possible to support multiple single-round dialogues or multiple single-person multi-round dialogues in parallel in a multi-terminal interaction scenario to meet the different needs of different users.
[0106] In an embodiment of the present application, the device further includes a second execution unit 28, and the second execution unit 28 is configured to:
[0107] Before receiving and executing the target operation instruction for the first registration processing result or the second registration processing result, in response to a third voice request instruction, execute the third voice request instruction, where the third voice request instruction is a voice request instruction that does not need to output the corresponding registration processing result and is used to directly control the vehicle state.
[0108] In an embodiment of the present application, the first search unit 21 includes:
[0109] An identification module 211, configured to identify the first voice request instruction as corresponding target semantic intention information;
[0110] A search module 212, configured to search for a first vertical domain corresponding to the target semantic intention information according to the correspondence between the semantic intention information and the vertical domain;
[0111] A sending module 213, configured to send the first voice request instruction to the first vertical domain.
[0112] Optionally, the first registration unit 22 includes:
[0113] An obtaining module 221, configured to obtain a first list page corresponding to the first voice request instruction based on the first vertical domain;
[0114] A registration module 222, configured to register the first list page to the dialogue management module.
[0115] In an embodiment of the present application, the first execution unit 27 includes:
[0116] A receiving module 271, configured to receive a target operation instruction, where the target operation instruction includes at least one of a voice instruction, a touch instruction, or a gaze instruction;
[0117] A first determination module 272, configured to determine a first display terminal identifier or a second display terminal identifier carried by the target operation instruction;
[0118] A second determination module 273, configured to, if it is determined that the display terminal identifier is the first display terminal identifier, determine the first registration processing result as the target registration processing result;
[0119] A third determination module 274, configured to, if it is determined that the display terminal identifier is the second display terminal identifier, determine the second registration processing result as the target registration processing result;
[0120] An execution module 275, configured to execute the target operation instruction on the target registration processing result.
[0121] In an embodiment of the present application, the device further includes a judgment unit 28, and the judgment unit 28 includes:
[0122] A judgment module 281, configured to judge the target category of the voice request instruction before looking up the first vertical domain corresponding to the first voice request instruction according to the correspondence between the request instruction and the vertical domain, where the first registration processing result has the control authority for the first category of voice request instructions, and the second registration processing result has the control authority for the second category of voice request instructions, and the first category includes the second category;
[0123] A control module 282, configured to, if it is determined that the target category of the voice request instruction does not belong to the second category, control the first registration processing result to execute the voice request instruction.
[0124] It should be noted that the foregoing explanations of the method embodiments also apply to the device in this embodiment, and the principles are the same, and are not limited in this embodiment.
[0125] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.
[0126] Figure 6 FIG. shows a schematic block diagram of an exemplary electronic device 300 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0127] As Figure 6 shown, the device 300 includes a computing unit 301 that can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 302 or a computer program loaded from a storage unit 308 into a RAM (Random Access Memory) 303. In the RAM 303, various programs and data required for the operation of the device 300 can also be stored. The computing unit 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An I / O (Input / Output) interface 305 is also connected to the bus 304.
[0128] A plurality of components in the device 300 are connected to the I / O interface 305, including: an input unit 306, such as a keyboard, a mouse, etc.; an output unit 307, such as various types of displays, speakers, etc.; a storage unit 308, such as a magnetic disk, an optical disk, etc.; and a communication unit 309, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 309 allows the device 300 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0129] The computing unit 301 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 301 executes the various methods and processes described above, such as the voice control method. For example, in some embodiments, the voice control method can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 308. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 300 via the ROM 302 and / or the communication unit 309. When the computer program is loaded into the RAM 303 and executed by the computing unit 301, one or more steps of the methods described above can be executed. Alternatively, in other embodiments, the computing unit 301 can be configured to execute the foregoing voice control method by any other suitable means (e.g., by means of firmware).
[0130] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SoCs (System On Chip), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0131] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program codes may be executed entirely on the machine, partially on the machine, executed partially on the machine as a stand-alone software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0132] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only-Memory), or a flash memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0133] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (Cathode-Ray Tube) or an LCD (Liquid Crystal Display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user may be in any form (including acoustic input, voice input, or tactile input).
[0134] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0135] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with blockchain.
[0136] It should be noted that artificial intelligence is a discipline that studies to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and has both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0137] It should be understood that various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is made herein.
[0138] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A voice control method, characterized in that, Including: In response to a first voice request instruction, look up a first vertical domain corresponding to the first voice request instruction according to the correspondence between the request instruction and the vertical domain; Register a first list page corresponding to the first voice request instruction into a dialogue management module based on the first vertical domain, and output a first registration processing result of the first list page on a first display terminal and other display terminals, where the dialogue management module is used to store the correspondence between the request instruction and the vertical domain and the list page, and the first display terminal is the display terminal that triggers the first voice request instruction; In response to a second voice request instruction, look up a second vertical domain corresponding to the second voice request instruction according to the correspondence between the request instruction and the vertical domain; Register a second list page corresponding to the second voice request instruction into the dialogue management module based on the second vertical domain, and output a second registration processing result of the second list page on a second display terminal, where the second display terminal is the display terminal that triggers the second voice request instruction; Receive and execute a target operation instruction for the first registration processing result or the second registration processing result, where the target operation instruction carries a first display terminal identifier or a second display terminal identifier for the first registration processing result or the second registration processing result.
2. The method according to claim 1, characterized in that, Before receiving and executing the target operation instruction for the first registration processing result or the second registration processing result, the method includes: In response to a third voice request instruction, execute the third voice request instruction, where the third voice request instruction is a voice request instruction that does not need to output a corresponding registration processing result and is used to directly control the vehicle state.
3. The method according to claim 1, characterized in that, The step of looking up a first vertical domain corresponding to the first voice request instruction according to the correspondence between the request instruction and the vertical domain includes: Identify the first voice request instruction as corresponding target semantic intention information; Look up a first vertical domain corresponding to the target semantic intention information according to the correspondence between the semantic intention information and the vertical domain; Send the first voice request instruction to the first vertical domain.
4. The method according to claim 1, characterized in that, The step of in response to a first voice request instruction includes: Obtain a first list page corresponding to the first voice request instruction based on the first vertical domain; Register the first list page into the dialogue management module.
5. The method according to claim 1, characterized in that, The step of receiving and executing the target operation instruction for the first registration processing result or the second registration processing result includes: Receive a target operation instruction, where the target operation instruction includes at least one of a voice instruction, a touch instruction, or a gaze instruction; Determine the first display terminal identifier or the second display terminal identifier carried by the target operation instruction; If it is determined that the display terminal identifier is the first display terminal identifier, determine the first registration processing result as the target registration processing result; If it is determined that the display terminal identifier is the second display terminal identifier, determine the second registration processing result as the target registration processing result; Execute the target operation instruction on the target registration processing result.
6. The method according to claim 5, characterized in that, The step of executing the target operation instruction on the target registration processing result includes: In the case where it is determined that the second registration processing result is the target registration processing result, after executing the target operation instruction on the target registration processing result, the second display terminal displays the first registration processing result of the first list page.
7. The method according to any one of claims 1 - 6, characterized in that, Before looking up the first vertical domain corresponding to the first voice request instruction according to the correspondence between the request instruction and the vertical domain, the method further includes: Judging the target category of the voice request instruction, the first registration processing result has the control authority for the first category of voice request instructions, and the second registration processing result has the control authority for the second category of voice request instructions, where the first category includes the second category; If it is determined that the target category of the voice request instruction does not belong to the second category, control the first registration processing result to execute the voice request instruction.
8. A voice control device, characterized in that, Includes: A first lookup unit, configured to, in response to a first voice request instruction, look up the first vertical domain corresponding to the first voice request instruction according to the correspondence between the request instruction and the vertical domain; A first registration unit, configured to register the first list page corresponding to the first voice request instruction into the dialogue management module based on the first vertical domain; A first output unit, configured to output the first registration processing result of the first list page on the first display terminal and other display terminals, where the dialogue management module is used to store the correspondence between the request instruction and the vertical domain and the list page, and the first display terminal is the display terminal that triggers the first voice request instruction; A second lookup unit, configured to, in response to a second voice request instruction, look up the second vertical domain corresponding to the second voice request instruction according to the correspondence between the request instruction and the vertical domain; A second registration unit, configured to register the second list page corresponding to the second voice request instruction into the dialogue management module based on the second vertical domain; A second output unit, configured to output the second registration processing result of the second list page on the second display terminal, where the second display terminal is the display terminal that triggers the second voice request instruction; A first execution unit, configured to receive and execute a target operation instruction on the first registration processing result or the second registration processing result, where the target operation instruction carries the first display terminal identifier or the second display terminal identifier of the first registration processing result or the second registration processing result.
9. An electronic device, characterized in that, Includes: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-7.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.