Voice interaction method and device
Patent Information
- Application Number
- CN202380092875.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-09-05
AI Technical Summary
When using terminal devices, users need to remember the jump path of the display interface with deep nesting levels, resulting in poor user experience and low efficiency.
A voice interaction method and device are provided. By obtaining the user's voice commands, the display device is controlled to switch from the first interface to the target interface above the second level, and to realize the direct jump from the current interface to the deeper nested interface.
It improves the user's voice interaction experience and efficiency, reduces the user's learning costs and interface switching delay, and provides a more humane interface interaction process.
Smart Images

Figure CN120604209A_ABST
Abstract
Description
Voice interaction method and device Technical Field
[0001] The present application relates to the field of terminal technology, and more specifically, to a voice interaction method and device. Background Art
[0002] Currently, when using a terminal device, if a user wants to view some deeply nested display interfaces, the user needs to actively memorize the jump path to the deeply nested display interface in advance. This will result in low efficiency when the user views the target interface, resulting in a poor user experience.
[0003] Summary of the Invention
[0004] The present application provides a voice interaction method and device, which helps to improve the user's human-computer interaction experience and human-computer interaction efficiency.
[0005] In a first aspect, a voice interaction method is provided, which includes: controlling a display device to display a first interface, which is a first-level interface; obtaining a first voice command from a user; and according to the first voice command, controlling the display device to display a second interface corresponding to the first voice command, which is an interface above the second level.
[0006] Based on the above technical solution, when a user issues a voice command, the terminal device can control the display device to switch from displaying the first-level interface to displaying interfaces at the second level or above. In this way, switching from the current interface to an interface with a deeper nesting level can be achieved, which helps to improve the user's voice interaction experience.
[0007] In the embodiments of the present application, the first and second levels may be relative concepts. For example, the interface displayed by the terminal device when receiving a voice command may be an interface of the first level, and the target interface indicated by the voice command may be an interface of the second level or an interface above the second level. The second level interface can also be understood as an interface displayed after one jump from the current interface, and the interface above the second level may be an interface displayed after at least two jumps from the current interface.
[0008] In some possible implementations, the content of the first voice instruction includes content indicating an interface at the second level or above.
[0009] In combination with the first aspect, in some implementations of the first aspect, the second interface is an interface of a settings application, a video application, a music application, or a car owner's guide application.
[0010] Based on the above technical solution, the terminal device can switch from the current interface to a deeper nested interface in applications such as the settings application, video application, music application, or car owner guide application, which helps to improve the user's voice interaction experience.
[0011] In some possible implementations, the second interface may be an interface of a pre-installed application (such as a settings application, a car owner's guide application, a video application, a music application, a navigation map application, etc.) or a third-party application (such as an application downloaded from an application store, etc.).
[0012] In combination with the first aspect, in certain implementations of the first aspect, controlling the display device to display a second interface corresponding to the first voice instruction according to the first voice instruction includes: controlling the display device to display at least one third interface according to the first voice instruction, and then displaying the second interface.
[0013] Based on the above technical solution, the terminal device can control the display device to switch from the first interface to at least one third interface first, and then switch from the at least one third interface to the second interface. In this way, the user can be shown the interface switching process and the specific location of the second interface (or the target interface), reducing the user's learning cost for using the application. At the same time, the interface switching process is more human-like, which helps to improve the user's voice interaction experience.
[0014] In combination with the first aspect, in some implementations of the first aspect, the method further includes: controlling the display device to display the jump progress on the first interface and / or the second interface.
[0015] Based on the above technical solution, the jump progress can be displayed on each interface during the jump from the current interface to the target interface. This makes it convenient for users to view the jump progress from the current interface to the target interface, which helps to improve the user's voice interaction experience.
[0016] In some possible implementations, controlling the display device to display the jump progress on the first interface and the second interface includes: displaying the total number of interface switches and the number of the current interface in the switching process on the first interface and the second interface.
[0017] In combination with the first aspect, in certain implementations of the first aspect, controlling the display device to display the second interface corresponding to the first voice instruction according to the first voice instruction includes: controlling the display device to display the second interface according to the first voice instruction and the first interface jump information, the first interface jump information including the jump path information of the multi-level interface, and the multi-level interface including the first interface and the second interface.
[0018] Currently, when a user asks how to perform an operation, the terminal device might pull up a tutorial interface and announce, "Let's go check out the tutorial." This doesn't directly solve the user's problem. When the user wants to perform an operation, they still have to understand the tutorial document, making them search for the corresponding solution in a complex tutorial, which increases the user's learning cost.
[0019] Based on the above technical solution, automatic switching from the first interface to the second interface can be achieved through interface jump information and user voice commands. This avoids tedious user operation processes and helps to improve the intelligence level of the terminal device. At the same time, by controlling the display device to switch from the current interface to the display interface the user expects to see, rather than forcing the user to search for solutions in a study manual, it helps reduce the user's learning cost and thus helps to improve the user experience.
[0020] In addition, by using the interface jump information of the entire application, all interfaces of the entire application can be made accessible, and then the target interface can be accurately jumped to.
[0021] In some possible implementations, the jump path information of the multi-level interface includes the jump order of the multi-level interface and the identification information of each interface in the multi-level interface. According to the first voice command and the first interface jump information, the display device is controlled to display the second interface, including: according to the jump order from the first interface to the second interface and the identification information of each interface in the process of jumping from the first interface to the second interface, the display device is controlled to switch from displaying the first interface to displaying the second interface.
[0022] In combination with the first aspect, in certain implementations of the first aspect, controlling the display device to display the second interface based on the first voice command and the first interface jump information includes: determining the second interface based on the first voice command; determining the first position of the second interface in the jump path information of the multi-level interface; determining the jump path information from the first interface to the second interface based on the second position of the first interface in the jump path information of the multi-level interface, the first position and the jump path information of the multi-level interface; controlling the display device to switch from displaying the first interface to displaying the second interface based on the jump path information from the first interface to the second interface.
[0023] Based on the above technical solution, the positions of the initial interface and the target interface in the jump path information of the multi-level interface can be determined respectively, and then the jump path from the initial interface to the target interface can be determined according to their respective positions and the jump path information of the multi-level interface.
[0024] In combination with the first aspect, in certain implementations of the first aspect, the multi-level interface also includes at least one third interface, which controls the display device to display the second interface according to the first voice command and the first interface jump information, including: controlling the display device to switch from displaying the first interface to displaying the at least one third interface according to the first voice command and the first interface jump information; controlling the display device to switch from displaying the at least one third interface to displaying the second interface.
[0025] Based on the above technical solution, when the second interface indicated by the voice command is an interface that cannot be directly jumped to from the currently displayed first interface, it is possible to switch from the first interface to at least one third interface based on the interface jump information, and then switch from at least one third interface to the second interface. In this way, the terminal device can jump between two directly related display interfaces, or it can jump from the current interface to a display interface with a deeper nesting level, which helps to improve the user experience.
[0026] In combination with the first aspect, in certain implementations of the first aspect, the jump path information of the multi-level interface consists of interaction paths of multiple interface elements and operation instructions for each of the multiple interface elements.
[0027] Exemplarily, taking the example of a jump process from a first interface to a second interface including a third interface, the multiple interface elements include a first interface element and a second interface element, the first interface element is located on the first interface, and the second interface element is located on the third interface, wherein, according to the first voice instruction and the first interface jump information, controlling the display device to display the second interface includes: performing an operation on the first interface element on the first interface to control the display device to switch from displaying the first interface to displaying the third interface; controlling the display device to switch from displaying the third interface to displaying the second interface includes: performing an operation on the second interface element on the third interface to control the display device to switch from displaying the third interface to displaying the second interface.
[0028] Based on the above technical solution, the terminal device can simulate the user's operation process based on the information of the interaction path of multiple interface elements and the operation of each of the multiple interface elements. This can make the jump process of the display interface more human-like and help further improve the user experience.
[0029] In combination with the first aspect, the above-mentioned interface jump information may be preset.
[0030] In combination with the first aspect, in certain implementations of the first aspect, controlling the display device to display a second interface corresponding to the first voice command based on the first voice command includes: inputting the information of the first interface and the content of the first voice command into a model to obtain a first interface element and an operation on the first interface element, the multiple interface elements including the first interface element; performing the operation on the first interface element to control the display device to switch from displaying the first interface to displaying the second interface.
[0031] Based on the above technical solution, the terminal device obtains interface jump information (or the interaction path of interface elements) in real time through model reasoning, so that the display device gradually switches to the target interface, providing generalization of voice interaction applications, thereby helping to improve the user's voice interaction experience.
[0032] In combination with the first aspect, in certain implementations of the first aspect, the second interface is a display interface of a first application, and the method further includes: obtaining jump information of the first interface corresponding to the first application.
[0033] Based on the above technical solution, the terminal device can first obtain the interface jump information corresponding to the first application. In this way, after receiving the voice command, the terminal device can directly execute the interface switching process based on the obtained interface jump information and the voice command. This can reduce the response delay between the terminal device receiving the voice command and displaying the second interface, avoiding the user waiting too long for the interface to change, and helping to improve the user experience.
[0034] In combination with the first aspect, in some implementations of the first aspect, the first interface jump information is included in a system update data packet received from the cloud server.
[0035] Based on the above technical solution, the terminal device can receive a system update data packet sent by the cloud server, which carries the second interface jump information for the first application. In this way, the interface jump information of the first application can be updated in a timely manner during the system update, so that the terminal device can control the display device to accurately display the target interface.
[0036] In combination with the first aspect, in certain implementations of the first aspect, the method also includes: sending a first request message to a cloud server, where the first request message is used to request interface jump information corresponding to the first version of the second application; and receiving the first interface jump information sent by the cloud server.
[0037] Based on the above technical solution, after receiving the request information from the terminal device, the corresponding interface jump information can be sent to the terminal device according to the association relationship. In this way, the terminal device can receive the interface jump information corresponding to the current version of the second application. The terminal device can perform the interface switching operation according to the voice command and the interface jump information, thereby controlling the display device to display the display interface that the user expects to see, which helps to improve the user's voice interaction experience. For example, the association relationship between the application version and the interface jump information can be saved in the cloud server.
[0038] In some possible implementations, the first request information includes the first version information.
[0039] In some possible implementations, the cloud server may store the correspondence between different versions of the second application and the interface jump information. When the first request information is obtained, the interface jump information corresponding to the first version can be searched based on the correspondence. In this way, the terminal device can obtain the first interface jump information corresponding to the current version of the second application from the cloud server by sending a request information, thereby helping the user accurately find the display interface the user wants to view based on the first interface jump information.
[0040] In some possible implementations, the first request information may include semantic information corresponding to the first voice command, and the first interface jump information may include jump path information from the first interface to the second interface corresponding to the semantic information. After receiving the first request information, the cloud server may determine the jump path information from the first interface to the second interface corresponding to the semantic information based on the semantic information and the interface jump information corresponding to the first version, and transmit the information to the terminal device.
[0041] Based on the above technical solution, the amount of data sent from the cloud server to the terminal device can be reduced. At the same time, the terminal device does not need to analyze the received interface jump information, which helps to reduce the delay from the terminal device receiving the voice command to controlling the display device to display the second interface, thereby helping to improve the user experience.
[0042] In combination with the first aspect, in certain implementations of the first aspect, when the third version of the third application does not match the second interface jump information, a second request message is sent to the cloud server, and the second request message is used to request the interface jump information corresponding to the third version; and the first interface jump information sent by the cloud server is received.
[0043] Based on the above technical solution, when the version of the third application in the terminal device does not match the interface jump information of the third application, the interface jump information corresponding to the current version of the third application can be requested from the cloud server. In this way, when the application in the terminal device is updated and the updated version does not match the previous interface jump information, the interface jump information corresponding to the latest version can also be updated accordingly. This helps to avoid the user from being unable to find the user's desired display interface through the interface jump information corresponding to the old version after issuing a voice command, which helps to improve the user experience.
[0044] In some possible implementations, when the third version of the third application does not match the second interface jump information, a second request message is sent to the cloud server, including: when the third version does not match the second interface jump information and the first voice instruction is obtained, a second request message is sent to the cloud server.
[0045] In some possible implementations, the second request information includes the third version information.
[0046] In combination with the first aspect, in some implementations of the first aspect, the method further includes: updating the second interface jump information according to the first interface jump information.
[0047] Based on the above technical solution, after receiving the interface jump information for the third version from the cloud server, the interface jump information of the third version can be used to update or overwrite the interface jump information corresponding to the previous version, so that the terminal device always stores the interface jump information corresponding to the current version of the application.
[0048] In combination with the first aspect, in certain implementations of the first aspect, before controlling the display device to display the second interface based on the first voice instruction and the first interface jump information, the method also includes: determining that the first interface jump information includes at least part of the text content corresponding to the first voice instruction.
[0049] Based on the above technical solution, the terminal device can first determine whether at least part of the text content corresponding to the voice command is included in the first interface jump information. If included, the interface switching can be executed according to the first interface jump information; otherwise, the terminal device can use other methods to perform the operation corresponding to the voice command.
[0050] In combination with the first aspect, in certain implementations of the first aspect, the first interface and the second interface are display interfaces of different applications, wherein, according to the first voice instruction, controlling the display device to display the second interface corresponding to the first voice instruction includes: according to the first voice instruction, controlling the first display area in the display device to display the first interface and controlling the second display area in the display device to display the second interface.
[0051] Based on the above technical solution, when the first and second interfaces are display interfaces of different applications, it is possible to switch from displaying the first interface on the display screen to displaying the first and second interfaces on a split screen. For example, the first interface is displayed in the first display area of the display screen, and the second interface is displayed in the second display area of the display screen. This prevents the display interface of the new application from overwriting the display interface of the previous application, thereby preventing the user from being able to view the display interface of the previous application, thereby helping to improve the user experience.
[0052] For example, if the first interface is the display interface of an in-vehicle navigation application and the second interface is the display interface of a settings application, the user's need to view the display interface of the settings application can be further met by maintaining the display interface of the in-vehicle navigation application in the first display area and switching from the display interface of the in-vehicle navigation application to the display interface of the settings application in the second display area without affecting the user's viewing of navigation information.
[0053] In some possible implementations, according to the first voice command, the display device is controlled to display a second interface corresponding to the first voice command, including: according to the first voice command, determining that the user wants to view the second interface; when it is determined that the first interface and the second interface are interfaces of different applications, controlling the first display area in the display device to display the first interface and controlling the second display area in the display device to display the second interface.
[0054] In the second aspect, a voice interaction device is provided, which includes: a control unit for controlling a display device to display a first interface, which is a first-level interface; an acquisition unit for acquiring a first voice instruction of a user, wherein the content of the first voice instruction includes content indicating an interface above the second level; the control unit is also used to control the display device to display a second interface corresponding to the first voice instruction according to the first voice instruction, wherein the second interface is an interface above the second level.
[0055] In combination with the second aspect, in some implementations of the second aspect, the second interface is an interface of a settings application, a video application, a music application, or a car owner's guide application.
[0056] In combination with the second aspect, in certain implementations of the second aspect, the control unit is used to: control the display device to display at least one third interface and then display the second interface according to the first voice instruction.
[0057] In combination with the second aspect, in some implementations of the second aspect, the control unit is further used to control the display device to display the jump progress on the first interface and / or the second interface.
[0058] In combination with the second aspect, in certain implementations of the second aspect, the control unit is used to: control the display device to display the second interface according to the first voice command and the first interface jump information, the first interface jump information includes the jump path information of the multi-level interface, and the multi-level interface includes the first interface and the second interface.
[0059] In combination with the second aspect, in certain implementations of the second aspect, the jump path information of the multi-level interface consists of interaction paths of multiple interface elements and operation instructions for each of the multiple interface elements.
[0060] In combination with the second aspect, in certain implementations of the second aspect, the acquisition unit is used to: input the information of the first interface and the content of the first voice instruction into the model to obtain a first interface element and an operation on the first interface element, and the multiple interface elements include the first interface element.
[0061] In combination with the second aspect, in some implementations of the second aspect, the second interface is a display interface of the first application, and the acquisition unit is further used to obtain the first interface jump information corresponding to the first application.
[0062] In combination with the second aspect, in some implementations of the second aspect, the device further includes: a receiving unit, configured to receive a system update data packet sent by a cloud server, wherein the system update data packet includes updated second interface jump information of the first application.
[0063] In combination with the second aspect, in certain implementations of the second aspect, the device also includes: a sending unit, used to send a first request message to the cloud server according to the first voice instruction, and the first request message is used to request the interface jump information corresponding to the first version of the second application; a receiving unit, used to receive the first interface jump information sent by the cloud server.
[0064] In conjunction with the second aspect, in certain implementations of the second aspect, the apparatus further includes: a sending unit configured to send a second request message to a cloud server when the third version of the third application does not match the second interface jump information, the second request message being configured to request the interface jump information corresponding to the third version; and a receiving unit configured to receive the first interface jump information sent by the cloud server. Exemplarily, the terminal device stores an association between the version of the third application and the interface jump information.
[0065] In combination with the second aspect, in certain implementations of the second aspect, the first interface and the second interface are display interfaces of different applications, wherein the control unit is used to: control the first display area in the display device to display the first interface and control the second display area in the display device to display the second interface according to the first voice command.
[0066] In a third aspect, a voice interaction device is provided, which includes a processing unit and a storage unit, wherein the storage unit is used to store instructions, and the processing unit executes the instructions stored in the storage unit so that the device executes any possible method in the first aspect.
[0067] In a fourth aspect, a system is provided, comprising a display device and a computing platform, wherein the computing platform comprises any possible device in the second aspect or the third aspect.
[0068] In a fifth aspect, a terminal device is provided, which includes any possible device in the second aspect, or includes the device described in the third aspect, or includes the system described in the fourth aspect.
[0069] In some possible implementations, the terminal device may be a vehicle.
[0070] In a sixth aspect, a computer program product is provided, comprising: a computer program code, which, when executed by one or more processors, enables a terminal device to execute any possible method in the first aspect.
[0071] It should be noted that the above-mentioned computer program code can be stored in whole or in part on the first storage medium, wherein the first storage medium can be packaged together with the processor or separately packaged with the processor, and the embodiments of the present application do not specifically limit this.
[0072] In a seventh aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores program code, and when the computer program code is executed by one or more processors, the terminal device executes any possible method in the above-mentioned first aspect.
[0073] In an eighth aspect, an embodiment of the present application provides a chip system, which includes a processor for calling a computer program or computer instructions stored in a memory so that the processor executes any possible method in the above-mentioned first aspect.
[0074] In combination with the eighth aspect, in a possible implementation, the processor is coupled to the memory through an interface.
[0075] In combination with the eighth aspect, in one possible implementation, the chip system also includes a memory, in which a computer program or computer instructions are stored. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] FIG1 is a functional block diagram of a terminal device provided in an embodiment of the present application.
[0077] FIG2 is a set of graphical user interfaces GUI provided by an embodiment of the present application.
[0078] FIG3 is another set of GUIs provided in an embodiment of the present application.
[0079] FIG4 is another set of GUIs provided in an embodiment of the present application.
[0080] FIG5 is a schematic diagram of the system architecture provided in an embodiment of the present application.
[0081] FIG6 is a schematic diagram of a layout tree stored in a large prediction model LLM provided in an embodiment of the present application.
[0082] FIG7 is another schematic diagram of the layout tree stored in the LLM provided in an embodiment of the present application.
[0083] FIG8 is a schematic flowchart of the voice interaction method provided in an embodiment of the present application.
[0084] FIG9 is a schematic block diagram of a voice interaction device according to an embodiment of the present application.
[0085] FIG10 is a schematic block diagram of a chip provided in an embodiment of the present application. DETAILED DESCRIPTION
[0086] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is only a way to describe the association relationship of associated objects, indicating that there can be three kinds of relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. "At least one" means one or more. For example, "at least one of A and B" is similar to "A and / or B", describing the association relationship of associated objects, indicating that there can be three kinds of relationships, for example, at least one of A and B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0087] In the embodiments of the present application, prefixes such as "first" and "second" are used only to distinguish different description objects and have no limiting effect on the position, order, priority, quantity or content of the described objects. The use of prefixes such as ordinal numbers to distinguish description objects in the embodiments of the present application does not constitute a restriction on the described objects. For the statement of the described objects, please refer to the description in the context of the claims or embodiments, and the use of such prefixes should not constitute an unnecessary restriction. In addition, in the description of this embodiment, unless otherwise specified, the meaning of "plurality" is two or more.
[0088] The following describes a terminal device, a graphical user interface (GUI) for such a terminal device, and embodiments for using such a terminal device. In some embodiments, the terminal device can be a mobile phone, a tablet computer, a wearable electronic device with wireless communication capabilities (such as a smart watch), a vehicle, an on-board domain controller (for example, a computing platform, a cockpit domain controller, a vehicle domain controller, etc.), an on-board terminal, etc. Exemplary embodiments of the terminal device include but are not limited to a terminal device equipped with Devices running HarmonyOS or other operating systems.
[0089] 1 is a functional block diagram of a terminal device 100 provided in an embodiment of the present application. The terminal device 100 may include a display device 110 and a computing platform 120.
[0090] Some or all functions of the terminal device 100 can be controlled by the computing platform 120. The computing platform 120 may include one or more processors, such as processors 121 to 12n (n is a positive integer). A processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and execution capabilities, such as a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a microprocessor), or a digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationship of a hardware circuit. The logical relationship of the hardware circuit is fixed or reconfigurable. For example, the processor is a hardware circuit implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as a field programmable gate array (FPGA). In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the hardware circuit configuration can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, the processor may also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a neural network processing unit (NPU), a tensor processing unit (TPU), a deep learning processing unit (DPU), etc. In addition, the computing platform 120 may also include a memory for storing instructions, and some or all of the processors 121 to 12n may call the instructions in the memory to implement corresponding functions.
[0091] Taking the terminal device 100 as an example, which is a vehicle, the display device 110 in the vehicle cabin is mainly divided into two categories. The first category is the vehicle-mounted display screen; the second category is a projection display screen, such as a head-up display (HUD). The vehicle-mounted display screen is a physical display screen and an important part of the vehicle infotainment system. There can be multiple display screens in the cabin, such as a digital instrument display screen, a central control screen, a display screen in front of the passenger on the co-pilot seat (also called the front passenger), a display screen in front of the left rear passenger, and a display screen in front of the right rear passenger. Even the windows can be used as display screens for display. Head-up display, also known as head-up display system. It is mainly used to display driving information such as speed and navigation on a display device in front of the driver (such as a windshield). This reduces the driver's line of sight diversion time, avoids pupil changes caused by the driver's line of sight diversion, and improves driving safety and comfort. HUDs include, for example, combiner-HUD (C-HUD), windshield-HUD (W-HUD), and augmented reality HUD (AR-HUD). It should be understood that other types of HUDs may emerge as technology evolves, and this application does not limit this.
[0092] The above display device 110 is described by taking a vehicle display screen and a projection display screen as examples, but the embodiments of the present application are not limited thereto. For example, the display device 110 can also be a light display screen or a projection screen.
[0093] The following GUI is described using the example of the terminal device 100 being a vehicle.
[0094] Exemplarily, FIG2 shows a set of GUIs provided in an embodiment of the present application.
[0095] The GUI shown in (a) of Figure 2 is the vehicle settings interface. The left side of the settings interface includes multiple tabs, such as the vehicle status tab, display tab, connection tab, sound tab, and smart assistant tab. The right side of the settings interface includes a description of the functions under each tab. For example, as shown in (a) of Figure 2, when initially entering the settings interface, the vehicle can display the display interface of the vehicle status tab on the display screen.
[0096] The vehicle setting interface shown in (a) of FIG. 2 above can also be understood as the display interface of the vehicle status tab.
[0097] When detecting that the user issues a voice command "Xiao A Xiao A, how do I set a custom wake-up word", the vehicle can convert the voice command into text and identify the user's intention based on the text content, thereby determining that the user's intention is to set the wake-up word for the voice assistant. The vehicle can control the display screen to switch from the display interface displaying the vehicle status tab to the display interface displaying the smart assistant tab as shown in (b) in Figure 2, from the display interface displaying the smart assistant tab to the display interface of the my smart assistant tab as shown in (c) in Figure 2, and from the display interface displaying the my smart assistant tab to the display interface displaying the custom wake-up word input as shown in (d) in Figure 2.
[0098] In one embodiment, the vehicle can control the display screen to switch from (a) to (d) as shown in Figure 2 based on locally stored interface jump information and the user's voice command.
[0099] Exemplarily, the interface jump information may include jump path information of a multi-level interface.
[0100] In one embodiment, the jump path information of the multi-level interface may include jump path information of the multi-level interface under the same application.
[0101] For example, taking the settings application as an example, the jump path information of the multi-level interface under the settings application includes the jump path information of interface 1-interface 2-interface 3-interface 4, where interface 1 can be the display interface of the vehicle status tab, interface 2 can be the display interface of the smart assistant tab, interface 3 can be the display interface of the my smart assistant tab, and interface 4 can be the display interface for entering a custom wake-up word.
[0102] For another example, the jump path information of the multi-level interface under the setting application also includes the interactive path information of interface 1-interface 5-interface 6, where interface 1 can be the display interface of the vehicle status tab, interface 5 can be the display interface of the maintenance and inspection sub-tab, and interface 6 can be the display interface of the wiper detection and maintenance mode.
[0103] In one embodiment, the jump path information of the multi-level interface may include jump path information of the multi-level interfaces under different application programs.
[0104] For example, taking the multiple applications that may include a music application and a settings application as an example, the jump path information of the multi-level interface under the different applications includes interface 7-desktop-interface 1-interface 2-interface 3-interface 4, where interface 7 is any display interface under the music application.
[0105] For example, Table 1 shows an interface jump information provided by an embodiment of the present application.
[0106] Table 1
[0107] The first-level interface above may be the interface displayed when the vehicle receives the user's voice command, the second-level interface may be the interface displayed after a single jump from the first-level interface, and the interfaces above the second level may be the interfaces displayed after at least two jumps from the first-level interface. For example, the third-level interface may be the interface displayed after two jumps from the first-level interface.
[0108] Exemplarily, the jump path information of the multi-level interface can be represented by an interactive path of multiple interface elements. The interface elements can include interface elements that can respond to operations, such as tabs, buttons, progress bars, or sliders on the interface.
[0109] Taking the Smart Assistant tab as an example, the Smart Assistant tab's display interface can include multiple sub-tabs, including the Smart Listening tab, My Smart Assistant tab, Voice Wake-up tab, Voice on Screen tab, Incoming Call Voice Announcement tab, and Voice Skills tab. The next level of the display interface associated with the My Smart Assistant tab can include controls for customizing wake-up words, customizing response words, and multiple audio mode selection controls. The custom wake-up word control is associated with the display interface for entering a custom wake-up word.
[0110] For example, the interaction path of the multiple interface elements can be the smart assistant tab - my smart assistant tab - custom wake-up word control. The interaction path of the multiple interface elements can be used to represent the jump path information of the above interface 1 - interface 2 - interface 3 - interface 4.
[0111] For example, on the interface 1 shown in (a) of FIG. 2 , the vehicle can simulate a user clicking on the smart assistant tab 201 , thereby displaying the interface 2 shown in (b) of FIG. 2 through the display screen.
[0112] For example, on the interface 2 shown in (b) of FIG. 2 , the vehicle can simulate the user clicking the My Smart Assistant tab 202 on the interface 2 , thereby displaying the interface 3 shown in (c) of FIG. 2 through the display screen.
[0113] For example, on the interface 3 shown in (c) of FIG. 2 , the vehicle can simulate the user clicking the custom wake-up word control 203 on the interface 3 , thereby displaying the interface 4 shown in (d) of FIG. 2 through the display screen.
[0114] In one embodiment, after the vehicle switches from interface 1 to interface 4 through the display screen, a prompt message "You can enter a custom wake-up word here" can also be displayed on interface 4 and voice broadcast can be performed.
[0115] In one embodiment, if the voice command issued by the user is "Xiao A Xiao A, how to set a custom wake-up word for a male voice", the vehicle can simulate the user first clicking on the male voice in the official voice on interface 3, and then simulate the user clicking on the custom wake-up word control 203 on interface 3, thereby displaying interface 4 through the display screen.
[0116] In one embodiment, during the process of switching from interface 1 to interface 4, each interface may also display the progress of switching from the initial interface to the target interface. This helps to improve the interactivity with the user, thereby helping to improve the user experience.
[0117] In one embodiment, displaying the progress from the initial interface to the target interface includes: displaying the number of display interfaces that need to be switched during the switching process from the initial interface to the target interface and the number of the display interface in the switching process that the current display interface is.
[0118] For example, on the interface 1 shown in (a) in FIG2 , a prompt message “Switching to the display interface for inputting a custom wake-up word requires three display interface switches, and the first display interface is currently being displayed” may be displayed.
[0119] For another example, on the interface 2 shown in (b) in FIG2 , a prompt message “Switching to the display interface for inputting a custom wake-up word requires three display interface switches, and the second display interface is currently being displayed” may be displayed.
[0120] For another example, on interface 3 as shown in (c) in FIG2 , a prompt message “Switching to the display interface for inputting a custom wake-up word requires three display interface switches, and the third display interface is currently being displayed” may be displayed.
[0121] For another example, on the interface 4 shown in (d) in FIG2 , a prompt message “Switching to the display interface for inputting a custom wake-up word requires three display interface switches, and the fourth display interface is currently being displayed” may be displayed.
[0122] In one embodiment, during the jump process of the above multi-level interfaces, the switching time between adjacent interfaces may be a first preset time, for example, 500ms.
[0123] In one embodiment, during the jump process of the above multi-level interfaces, the total duration of switching from the current interface to the target interface is the second preset duration.
[0124] For example, the second preset duration is 3 seconds. Taking the above-mentioned switching from interface 1 to interface 4 as an example, after determining that the number of jumps is 3, the jump duration between each two adjacent interfaces is 1 second. For another example, if it is determined based on the voice command and interface jump information that 5 interface jumps need to be performed, then the jump duration between adjacent interfaces can be 600ms. This helps to avoid the total jump duration being too long and also helps to avoid the user's waiting time being too long.
[0125] In one embodiment, after receiving the user's voice command, the vehicle can control the display screen to switch interfaces according to the interface jump information and the voice command; or, the display screen can be controlled to switch interfaces by combining the interface jump information with other interface switching methods. For example, taking the deeplink technology as an example, the deeplink technology can achieve the purpose of switching from one interface to another through a uniform resource identifier (URI). The vehicle can switch from the current interface to the interface closest to the target interface in combination with the deeplink technology; and then switch from the interface closest to the target interface to the target interface according to the interface jump information.
[0126] For example, the interface jump information includes the jump path information from Interface 1 to Interface 8 to Interface 9 to Interface 10. Deeplink technology can be used to switch from Interface 1 to Interface 9. When the vehicle receives a voice command to open Interface 10 while the display screen is displaying Interface 1, Deeplink technology can be used to first control the display screen to switch from Interface 1 to Interface 9, and then control the display screen to switch from Interface 9 to Interface 10 based on the interface jump information. This ensures that the number of jumps from the current interface to the target interface is minimized, making it easier for users to view the target interface in a timely manner and helping to improve the user experience.
[0127] In the embodiment of the present application, through the user's voice command and the above-mentioned interface jump information, the vehicle can control the display screen to jump from the initial settings interface to the display interface of the smart assistant tab, then from the display interface of the smart assistant tab to the display interface of the my smart assistant tab, and finally from the display interface of the my smart assistant tab to the display interface for entering a custom wake-up word. Thus, by issuing a single voice command, it is possible to jump from the current interface to an interface with a deeper nesting level, eliminating the user's tedious operation and learning process, and helping to improve the user's voice interaction experience.
[0128] At the same time, accurately finding the target interface based on interface jump information can not only expand the range of interfaces that can be reached by voice interaction, but also avoid users from searching step by step, improving the convenience of interaction. In addition, it can also guide users to understand the location of the target interface, especially for users who are unfamiliar with the system or application, thus helping to improve the user experience.
[0129] Exemplarily, FIG3 shows a set of GUIs provided in an embodiment of the present application.
[0130] As shown in (a) in FIG3 , the GUI is the vehicle's desktop, wherein the desktop includes multiple cards, such as a card corresponding to settings, a card corresponding to music applications, a card corresponding to the vehicle's remaining power and cruising range, and a card 301 corresponding to the owner's guide.
[0131] When detecting that the user issues a voice command "Xiao A Xiao A, I want to check the emergency mode of the remote control key", the vehicle can convert the voice command into text content and identify the user's intention based on the text content, so as to determine that the user's intention is to view the description content corresponding to the emergency mode of the remote control key in the user manual. The vehicle can switch from displaying the desktop through the display screen to displaying the initial interface of the owner's guide application as shown in (b) of Figure 3, from displaying the initial interface of the owner's guide application to the display interface of the user manual tab 302 as shown in (c) of Figure 3, from displaying the user manual tab 302 to displaying the remote control key tab 304 as shown in (d) of Figure 3, and from displaying the remote control key tab 304 to displaying the description content of the emergency mode of the remote control key as shown in (e) of Figure 3.
[0132] In one embodiment, the vehicle can control the display screen to switch from (a) to (e) in FIG. 3 according to the interface jump information and the user's voice command.
[0133] For example, Table 2 shows another interface jump information provided in an embodiment of the present application.
[0134] Table 2
[0135] The above description is based on an example in which the first-level display interface is the desktop shown in (a) of Figure 3 . The content of the voice command issued by the user includes content for instructing a sixth-level display interface (e.g., the display interface of the remote control key's emergency mode, not shown in Table 2). If the terminal device displays the display interface of the vehicle control tab when receiving the user's voice command, then the display interface of the vehicle control tab may be the first-level display interface, and the content of the voice command issued by the user includes content for instructing a fifth-level display interface (e.g., the display interface of the remote control key's emergency mode, not shown in Table 2).
[0136] The division of interfaces into different levels in Table 2 above is merely illustrative and is not intended to limit the present invention. For example, the interface jump information may not include the display interface for the key tab or the display interface for the door tab. In this case, the display interface for the remote key tab can be considered a fourth-level interface, and the interface describing the remote key's emergency mode can be considered a fifth-level interface.
[0137] The Owner's Guide app includes multiple tabs, such as a Highlights tab, a Voice Skills tab, and a User Manual tab 302. The Highlights tab includes a Getting Started tab, a Smart Voice tab, a Smart Connectivity tab, an Assisted Driving tab, and a Featured Apps tab; the User Manual tab includes a Vehicle Overview tab, a Driving Safety tab, a Vehicle Control tab 303, a Driving tab, an Assisted Driving tab, a Travel & Entertainment tab, and a Smart Car tab. The Vehicle Control tab 303 includes a Keys tab and a Doors tab. The Keys tab includes a Remote Key tab 304, a Phone Key tab, a Watch Key tab, and a Card Key tab. The Remote Key tab includes descriptions of button functions, the remote key's emergency mode, and instructions for replacing the remote key's battery.
[0138] For example, the interface jump information may include information about the interaction path of multiple interface elements, where the interaction path includes the owner guide application icon 301, the user manual tab 302, the vehicle control tab 303, the remote control key tab 304, and a preset point (or progress bar) on the display interface. The interaction path of multiple interface elements can be used to represent the jump path information from the desktop to the initial interface of the owner guide application, the display interface of the user manual tab, the display interface of the remote control key tab, and the display interface of the remote control key emergency mode description.
[0139] On the desktop as shown in (a) in Figure 3, the vehicle can simulate the user clicking on the icon 301 of the owner's guide application, thereby displaying the initial interface of the owner's guide application as shown in (b) in Figure 3 or the display interface of the highlight function tab in the owner's guide application through the display screen.
[0140] On the interface shown in (b) of FIG. 3 , the vehicle may simulate a user clicking on the user manual tab 302 , thereby displaying the display interface of the user manual tab 302 shown in (c) of FIG. 3 through the display screen.
[0141] On the interface shown in FIG. 3 ( c ), the vehicle can simulate a user clicking on the remote key tab 304 in the vehicle control tab 303 , thereby displaying the display interface of the remote key tab 304 shown in FIG. 3 ( d ) on the display screen.
[0142] The above switching process from (c) in Figure 3 to (d) in Figure 3 is merely illustrative. The vehicle control tab 303 may include a key tab and a door tab, etc., and the key tab may include a remote control key tab, a mobile phone car key tab, a watch car key tab, and a card key tab, etc. For example, the vehicle may first simulate a user clicking on the vehicle control tab 303, thereby displaying the expanded display interface of the vehicle control tab 303 through the display screen, and the expanded display interface of the vehicle control tab 303 includes a key tab and a door tab (at this time, the key tab 304 is in an unfolded state). The vehicle may then simulate a user clicking on the key tab 304, thereby displaying the display interface of the remote control key tab 304 as shown in (d) in Figure 3 through the display screen.
[0143] On the interface shown in (d) in Figure 3, the vehicle can simulate the user's sliding operation at a certain point on the display interface or the sliding operation on the progress bar, thereby displaying the description content of the remote control key emergency mode as shown in (e) in Figure 3 through the display screen.
[0144] In this embodiment of the present application, voice commands issued by the user can help the user accurately locate deeply nested display interfaces, such as the display interface for the description of a specific function in the owner's manual. This eliminates the need for the user to search for the corresponding function description in a large and complex owner's manual, thus improving the user experience.
[0145] Exemplarily, FIG4 shows a set of GUIs provided in an embodiment of the present application.
[0146] As shown in (a) of FIG4 , the GUI is a display interface of a video application, wherein the display interface includes information of “Movie A” to “Movie H”.
[0147] The GUI shown in (a) of FIG. 4 above may also be the initial display interface of the video application.
[0148] When the user issues a voice command "Xiao A Xiao A, I want to watch the movie starring actor 1", the vehicle can convert the voice command into text and identify the user's intention based on the text content, thereby determining that the user wants to watch the movie starring actor 1. Based on the user's intention, the vehicle can control the display screen to switch from the initial display interface of the video application to the video playback interface of "Movie D" as shown in (d) in Figure 4.
[0149] In one embodiment, the vehicle may send a request to the cloud server for information about the video application's interface jump. After receiving the interface jump information from the cloud server, the vehicle may switch from (a) to (b) in FIG4 based on the user's voice command and the interface jump information.
[0150] For example, the cloud server can analyze the display interfaces of different levels in the video application to obtain the jump path information of the multi-level interface of the video application. For example, the cloud server can analyze the poster information corresponding to each movie (for example, the movie poster information includes the list of starring actors, plot summary, and movie genre) to determine the interface jump information of the video application. The above interface jump information may include the correspondence between the movie name and the poster information.
[0151] The above description uses the example of a cloud server analyzing the display interfaces at different levels in a video application to obtain interface jump information, but the embodiments of the present application are not limited to this. The vehicle can analyze the display interfaces at different levels in the video application to obtain jump path information for the multi-level interfaces in the video application. For example, the vehicle can dynamically capture the content displayed on the screen by taking a screenshot, thereby obtaining the poster information corresponding to each movie and determining the correspondence between the movie title and the poster information.
[0152] For example, Table 3 shows another interface jump information provided in an embodiment of the present application.
[0153] Table 3
[0154] The vehicle can determine that the user wants to watch "Movie D" based on the above interface jump information and the user's voice command, and can control the display screen to display the video playback interface of "Movie D" as shown in (b) of Figure 4.
[0155] Figure 5 shows a schematic diagram of the system architecture provided by an embodiment of the present application. As shown in Figure 5, the system architecture includes a large language model (LLM), a robotic process automation (RPA) policy generator, and an RPA execution engine.
[0156] Exemplarily, the layout information of each display interface is used to form interface jump information through UiAutomator testing, optical character recognition (OCR), etc. When the vehicle receives a voice command from the user, the voice command and interface jump information can be input into the LLM. The LLM can output the jump path information of the multi-level interface from the current interface to the target interface to the RPA policy generator based on the input voice command and interface jump information. The representation format of the jump path information can be a json string. The RPA policy generator can generate a combined instruction based on the jump path information output by the LLM and send the combined instruction to the RPA execution engine. The RPA execution engine can execute the combined instruction to enable the vehicle to simulate user operations, thereby controlling the display screen to switch from displaying the current interface to displaying the target interface.
[0157] The above UiAutomator can be used to obtain multiple interface elements on an interface and the operations of each interface element in multiple interface elements. The interaction paths between different interfaces and the interaction paths of interface elements between different interfaces can be obtained using other testing tools or scripts. For example, a testing tool or script can be used to obtain the current interface and one or more next interfaces reached after operating at least one interface element of the current interface, thereby establishing the jump path information from the current interface to one or more next interfaces.
[0158] The above description uses the example of testing the interaction paths of interface elements on different interfaces using a test tool, but the embodiments of this application are not limited to this. The interaction paths of interface elements on different interfaces can also be obtained through model reasoning. The following description will focus on the model training stage and the model reasoning stage.
[0159] Model training phase
[0160] For example, the model can be trained using the page layout information of different interfaces in the video application 1. During the model training process, the model can be trained using labeled data. The labeled data can include an interface in the video application 1 that has been labeled (i.e., an identifier is labeled for each interface element on the interface. The identifier indicates that the interface element can respond to an operation on the interface. The identifier can be embodied as a number, a string, etc., and the identifier corresponding to each interface element is different), text content (text content corresponding to a voice command), a target interface element, and an operation on the target interface element as a true value. For example, the labeled data can include the homepage of the video application 1, the text content "Open my recently watched videos", and the corresponding true value is the "My" tab on the homepage of the video application 1 and the click operation on the "My" tab. For another example, the labeled data can also include the display interface of the "My" tab, the text content "Open my recently watched videos", and the corresponding true value is the viewing history control on the display interface of the "My" tab and the click operation on the viewing history control. For another example, the labeled data may also include a display interface for the history record corresponding to a viewing history control, with the text "Open my recently watched videos" and the corresponding true value being the first result in the history record and a click operation on that first result. The model can be trained using this labeled data. After the model is trained, it can be deployed on the terminal device.
[0161] Model inference stage
[0162] For example, after receiving the user's voice command, the terminal device can convert the voice command into text content and annotate the current interface, for example, assigning an identifier to each interface element on the current interface (for example, interface a). After the annotation is completed, the terminal device can input the annotated current interface and text content into the model, so that the model can infer the target interface element and the operation on the target interface element. The terminal device can perform the operation on the target interface element on interface a, thereby controlling the display screen to jump to the next interface of the current interface (for example, interface b).
[0163] The terminal device can annotate interface b. After the annotation is completed, the terminal device can input the annotated interface b and text content into the model, so that the model can infer another target interface element and the operation of the other target element. The terminal device can perform the operation on the other target interface element on interface b, thereby controlling the display screen to jump to the next level interface of interface b (for example, interface c). Similarly, the terminal device can control the display screen to jump to the display interface that the user expects to see.
[0164] For example, while displaying the homepage of Video Application 2, a terminal device receives a user's voice command, "Continue watching the video I watched previously." The terminal device can annotate the homepage of Video Application 2. After annotating, the terminal device can input the annotated homepage of Video Application 2 and the text content corresponding to the voice command into the model, so that the model can infer the "My" tab on the homepage of Video Application 2 and the click operation on the "My" tab. The terminal device can click the "My" tab to control the display screen to switch from displaying the homepage of Video Application 2 to displaying the "My" tab. The terminal device can annotate the display screen of the "My" tab. After annotating, the terminal device can input the annotated display screen of the "My" tab and the text content into the model, so that the model can infer the viewing history control on the display screen of the "My" tab and the click operation on the viewing history control. The terminal device can click the viewing history control to control the display screen to switch from displaying the "My" tab to displaying the viewing history. The terminal device can annotate the viewing history display screen. After the annotation is complete, the terminal device can input the viewing history display interface and the text content into the model, so that the model can infer the first video content on the viewing history display interface and the click operation on the first video content. The terminal device can execute the click operation on the first video content, thereby controlling the display screen to switch from the viewing history display interface to the playback interface of the first video.
[0165] For example, a terminal device receives a user's voice command "Send file B to user A" while viewing the homepage of a social application. The terminal device can annotate the homepage of the social application. After annotating, the terminal device can input the annotated homepage of the social application and the text content corresponding to the voice command into a model. The model can then infer the chat bar with user A on the homepage of the social application and a click operation on the chat bar. The terminal device can click the chat bar to control the display screen to switch from displaying the homepage of the social application to displaying the chat interface with user A. The terminal device can annotate the chat interface with user A. After annotating, the terminal device can input the annotated chat interface with user A and the text content into a model. The model can then infer the file transfer control on the chat interface with user A and a click operation on the file transfer control. The terminal device can click the file transfer control to control the display screen to switch from displaying the chat interface with user A to displaying the file selection interface. The terminal device can annotate the file selection interface. After the annotation is completed, the terminal device can input the annotated file selection interface and the text content into the model, so that the model can infer the location of file B and the click operation on the location of file B. The terminal device can perform the click operation on the location of file B, thereby controlling the display screen to switch from displaying the file selection interface to displaying the chat interface with user A. At this time, the chat interface with user A shows that file B has been successfully transferred to user A.
[0166] The following describes the process of generating interface jump information through UiAutomator, OCR, etc. Taking the application scenario shown in Figure 2 as an example, UiAutomator can test each interface element on the setting interface shown in (a)-(d) in Figure 2.
[0167] For example, UiAutomator can test and determine that clicking the Smart Assistant tab displays the Smart Listening tab, My Smart Assistant tab, Voice Wake-up tab, Voice On-Screen tab, Incoming Call Announcement tab, and Voice Skills tab. Clicking the My Smart Assistant tab displays a custom wake-up word control and a custom response control. UiAutomator can record the coordinates of each tab and control. For example, the pixel coordinates for the Smart Assistant tab are [200, 2130], the pixel coordinates for the Smart Listening tab are [280, 430], the pixel coordinates for the My Smart Assistant tab are [800, 430], the pixel coordinates for the custom wake-up word control are [274, 462], and the pixel coordinates for the custom response control are [700, 462]. UiAutomator can send information such as the interaction paths between various interface elements and the operations performed on each interface element to the LLM. This allows the LLM to pre-configure interface jump information within the vehicle. In this way, after the user issues a voice command, the LLM can input the voice command into the LLM. Based on the semantic information corresponding to the voice command and the pre-built interface jump information, the LLM can output a JSON string used to generate the command to the RPA policy generator. For example, the JSON string can include the interaction path between each interface element (for example, root-interface element a-interface element b-...) and the corresponding operation of each interface element (for example, click, slide, zoom in or out, etc.).
[0168] The interaction paths between these interface elements and the operations performed on each element can be used to represent interface transition information. Use tools such as UiAutomator to construct interface transition information offline and send it to the LLM as a prompt. Upon receiving a voice command, the LLM generates a JSON string based on the voice command and interface transition information and sends it to the RPA policy generator.
[0169] For example, FIG6 shows a schematic diagram of the interface jump information stored in the LLM provided in an embodiment of the present application.
[0170] After receiving the JSON string sent by the LLM, the RPA strategy generator can generate a combination instruction (65, 22, 13), which indicates that instructions 65, 22, and 13 are executed in sequence. Table 4 shows the information of each instruction.
[0171] Table 4
[0172] The PRA executor can switch from the interface (a) to (d) in FIG2 by executing the combined instruction (65, 22, 13).
[0173] Taking the application scenario shown in FIG3 as an example, UiAutomator can test each interface element on the setting interface shown in (b)-(e) in FIG3.
[0174] For example, through testing, UiAutomator can determine that after clicking the User Manual tab, the Vehicle Overview tab, Driving Safety tab, Vehicle Control tab, Driving Vehicle tab, Assisted Driving tab, etc. can be displayed. After clicking the Vehicle Control tab, the Key tab and Door tab can be displayed. After clicking the Key tab, the Remote Control Key tab, Mobile Phone Car Key tab, Watch Car Key tab and Card Key tab can be displayed. After clicking the Remote Control Key tab, the description content corresponding to the button function can be displayed. After clicking the Remote Control Key tab, swiping up a first distance at a preset point on the screen (or at the progress bar) can display the description content corresponding to the emergency mode of the remote control key. After clicking the Remote Control Key tab, swiping up a second distance at a preset point on the screen (or at the progress bar) can display the description content corresponding to replacing the remote control key battery. UiAutomator can record the coordinates of each tab or preset point. For example, the pixel coordinates corresponding to the User Manual tab are pixel 1, the pixel coordinates corresponding to the Vehicle Overview tab are pixel 2, the pixel coordinates corresponding to the Driving Safety tab are pixel 3, the pixel coordinates corresponding to the Vehicle Control tab are pixel 4, the pixel coordinates corresponding to the Key tab are pixel 5, the pixel coordinates corresponding to the Door tab are pixel 6, the pixel coordinates corresponding to the Remote Key tab are pixel 7, the pixel coordinates corresponding to the Phone Key tab are pixel 8, the pixel coordinates corresponding to the Watch Key tab are pixel 9, the pixel coordinates corresponding to the Card Key tab are pixel 10, and the pixel coordinates of the preset point are pixel 11. UiAutomator can send information such as the interaction paths between various interface elements and the operations performed on each interface element to the LLM. This allows the LLM to pre-configure interface transition information within the vehicle. This way, when a user issues a voice command and inputs it into the LLM, the LLM can output a JSON string to the RPA policy generator for generating instructions based on the semantic information corresponding to the voice command and the pre-configured interface transition information.
[0175] For example, FIG7 shows a schematic diagram of the interface jump information stored in the LLM provided in an embodiment of the present application.
[0176] After the LLM generates a JSON string for the RPA policy generator, the RPA policy generator can generate a combined instruction (53, 48, 12, 79, 93), which indicates that instructions 53, 48, 12, 79, and 93 should be executed in sequence. Table 5 shows the information of each instruction.
[0177] Table 5
[0178] The PRA executor can realize the display change from (b) to (e) in Figure 3 by executing the combined instructions (53, 48, 12, 79, 93).
[0179] In one embodiment, when obtaining a user's voice command, the vehicle can first determine whether the semantic information corresponding to the voice command is included in the pre-determined interface jump information. For example, when the LLM obtains the voice command "turn up the volume", it can determine that the semantic information corresponding to the voice command is not included in the interface jump information determined in advance by the LLM. For another example, when the LLM obtains the voice command "how to set a custom wake-up word", it can determine that the semantic information corresponding to the voice command is included in the interface jump information shown in Figure 6, so that the corresponding json string can be output to the RPA policy generator, so that the RPA policy generator can generate the corresponding instructions.
[0180] In one embodiment, to ensure the stability of LLM output results, a scoring mechanism can be implemented. After multiple executions, the majority rule prevails. Multiple executions can refer to the LLM output or inference process of node link information. For example, repeating the process five times, according to the scoring mechanism, ensures the stability of the model output and improves the confidence of the LLM output results.
[0181] In one embodiment, the LLM can also constrain the format of the JSON string it outputs. This JSON string can include information about the interaction path and the command ID. For example, using the application scenario shown in Figure 3, the JSON string output by the LLM can include information about the interaction path: User Manual tab - Vehicle Control tab - Key tab - Remote Key tab - Preset Point, as well as the command ID information shown in Table 5. This allows the RPA policy generator to generate the corresponding combined command upon receiving this JSON string.
[0182] The above description uses the example of interface jump information including the interaction path between multiple interface elements and the operation of each of the multiple interface elements. The embodiments of the present application are not limited to this. For example, the interface jump information can also include the jump order of multi-level interfaces and the identification information of each interface in the multi-level interface.
[0183] Taking the application scenario shown in Figure 2 as an example, the interface jump information may include an identifier 1 corresponding to the desktop, an identifier 2 corresponding to the display interface of the vehicle status tab, an identifier 3 corresponding to the display interface of the smart assistant, an identifier 4 corresponding to the display interface of my smart assistant, and an identifier 5 corresponding to the display interface for inputting a custom wake-up word, as well as the jump order of these interfaces. For example, if the current display interface is the desktop and the target display interface is the display interface for inputting a custom wake-up word, then the vehicle can switch the display interface in the order of the desktop, the display interface of the vehicle status tab, the display interface of the smart assistant, the display interface of my smart assistant, and the display interface for inputting a custom wake-up word. Exemplarily, the computing platform 120 can determine that the identifier of the target interface is identifier 5 and the identifier of the current display interface is identifier 1 based on the voice command. The computing platform 120 can send the identifiers of the multi-level interfaces from the starting interface to the target interface and the jump order from the starting interface to the target interface to the setting application based on the interface jump information, so that the setting application can execute the switch from the desktop to the display interface for inputting a custom wake-up word.
[0184] When the above interface jump information includes the jump order of multi-level interfaces and the identification information of each interface, the interface switching process may not depend on interface elements, but on interface switching based on different operating system ecosystems (for example, interface switching based on Android ecosystem, switching based on Harmony ecosystem, or interface switching based on iOS ecosystem).
[0185] For example, as shown in (a) of Figure 2, when the vehicle control display screen displays the initial display interface of the settings application, it is detected that the user issues a voice command "Xiao A Xiao A, how to set a custom wake-up word with a male voice". At this time, the vehicle can determine that the user wants to enter a custom wake-up word based on the voice command issued by the user. Based on the jump order of the above-mentioned multi-level interfaces and the identification information of each interface in the multi-level interface, the vehicle first switches to the display interface displaying the smart assistant according to identification 3, then switches to the display interface displaying my smart assistant according to identification 4, and finally switches to the display interface displaying the input custom wake-up word according to identification 5.
[0186] For example, the display interface of the Vehicle Status tab shown in FIG2(a) and the display interface of the Smart Assistant tab shown in FIG2(b) can be display interfaces at the same level. For example, the display interface of the Smart Assistant tab shown in FIG2(b) and the display interface of the My Smart Assistant tab shown in FIG2(c) can be display interfaces at different levels.
[0187] The above interface jump information may also include the order of other display interfaces of the same level. Exemplarily, the display interface of the display tab and the display interface of the smart assistant tab shown in (b) of Figure 2 may also be display interfaces of the same level. When the vehicle controls the display screen to display the display interface of the display tab and detects that the user issues a voice command "Xiao A Xiao A, how to set a custom wake-up word", it can determine that the user wants to view the display interface for inputting a custom wake-up word based on the voice command issued by the user. According to the jump order of the above-mentioned multi-level interfaces and the identification information of each interface in the multi-level interface, the vehicle can switch from the display interface displaying the display tab to the display interface displaying the smart assistant according to identification 3, then switch to the display interface displaying my smart assistant according to identification 4, and finally switch to the display interface displaying the input custom wake-up word according to identification 5.
[0188] FIG8 shows a schematic flow chart of a voice interaction method 800 provided in an embodiment of the present application. The method 800 can be executed by the terminal device 100, or the method 800 can be executed by the computing platform 120, or the method 800 can be executed by a system consisting of the computing platform 120 and the display device 110, or the method 800 can be executed by a system-on-a-chip (SoC) in the computing platform 120, or the method 800 can be executed by a processor, chip, or circuit in the computing platform 120. The method 800 includes:
[0189] S810, controlling the display device to display a first interface.
[0190] Optionally, the first interface is a first-level interface. In the embodiment of the present application, the interface displayed by the display device when receiving the user's voice command can be called a first-level interface.
[0191] Exemplarily, as shown in (a) in FIG. 2 , the first interface may be an initial display interface of a settings application, or the first interface may be a display interface of a vehicle status tab.
[0192] Exemplarily, as shown in (a) in FIG3 , the first interface may be a desktop.
[0193] Exemplarily, as shown in (a) of FIG4 , the first interface may be an initial display interface of a video application.
[0194] S820: Acquire a first voice command from the user.
[0195] Optionally, the content of the first voice instruction includes content indicating an interface of the second level or above.
[0196] For example, as shown in (a) of FIG2 , the vehicle may receive a voice command from the user, “Xiao A Xiao A, how do I set a custom wake-up word?” The first voice command includes the text content “custom wake-up word,” which indicates an interface for inputting a custom wake-up word. The interface for inputting a custom wake-up word, as shown in (d) of FIG2 , may be a fourth-level interface. Three interface jumps may be performed from the first-level interface to the fourth-level interface.
[0197] For example, as shown in (a) of FIG3 , a vehicle may receive a user's voice command "Xiao A Xiao A, I want to check the remote control key's emergency mode." This first voice command includes the text "Remote control key's emergency mode," which indicates the remote control key's emergency mode interface. The interface for inputting a custom wake-up word, as shown in (e) of FIG3 , may be a fifth-level interface. Four interface jumps are required to go from the first-level interface to the fifth-level interface.
[0198] For example, as shown in FIG4(a), a vehicle may receive a user's voice command "Xiao A Xiao A, I want to watch the movie starring actor 1." This first voice command includes the text "The movie starring actor 1," which indicates the playback interface of "Movie D." The playback interface of "Movie D," as shown in FIG4(b), may be a second-level interface. A single interface jump may be required to transition from the first-level interface to the second-level interface.
[0199] S830, according to the first voice command, controlling the display device to display a second interface corresponding to the first voice command, the second interface being an interface of the second level, or the second interface being an interface above the second level.
[0200] Optionally, the first interface and the second interface are display interfaces of different applications. According to the first voice command, the display device is controlled to display the second interface corresponding to the first voice command, including: controlling the first display area in the display screen to display the first interface and controlling the second display area in the display screen to display the second interface.
[0201] For example, when the voice command is received, the vehicle's display screen is displaying the interface of an in-vehicle map application. After receiving the voice command, the vehicle can control the first display area of the display screen to display the interface of the in-vehicle map application and control the second display area of the display screen to display the interface for inputting a custom wake-up word, or control the second display area to switch from displaying the interface of the in-vehicle map application to displaying the interface for inputting a custom wake-up word.
[0202] Optionally, the method 800 further includes: determining display positions of the first display area and the second display area on the display device according to the area where the user who issued the first voice command is located.
[0203] For example, when the area where the user who issued the first voice command is located is the main driving area, the second display area can be controlled to be displayed in a position close to the main driving area and the first display area can be controlled to be displayed in a position away from the main driving area. In this way, it is convenient for the user to view the interface he wants to view through the display area close to his area.
[0204] Optionally, controlling the display device to display a second interface corresponding to the first voice instruction according to the first voice instruction includes: controlling the display device to display at least one third interface according to the first voice instruction, and then displaying the second interface.
[0205] For example, taking the application scenario shown in Figure 2 as an example, the first interface can be the display interface of the vehicle control tab, the second interface can be the display interface for inputting a custom wake-up word, and the at least one third interface can be two third interfaces, namely the display interface of the smart assistant tab and the display interface of the my smart assistant tab.
[0206] As shown in (a) in Figure 2, when the vehicle receives the voice command "Xiao A Xiao A, how to set a custom wake-up word", it can analyze the voice command to obtain the user's intention to enter a custom wake-up word. In this way, the vehicle can determine that the target interface (second interface) that the user wants to see is the display interface for entering a custom wake-up word. According to the voice command, the vehicle can control the display screen to switch from displaying the vehicle status tab as shown in (a) in Figure 2 to displaying the display interface for entering a custom wake-up word as shown in (d) in Figure 2, wherein the interfaces shown in (b) and (c) in Figure 2 can be the above-mentioned third interface.
[0207] Optionally, the method 800 further includes: controlling the display device to display the jump progress on the first interface and / or the second interface.
[0208] Optionally, according to the first voice command, the display device is controlled to display the second interface corresponding to the first voice command, including: according to the first voice command and the first interface jump information, the display device is controlled to display the second interface, the first interface jump information includes the jump path information of the multi-level interface, and the multi-level interface includes the first interface and the second interface.
[0209] For example, taking the application scenario shown in Figure 2 as an example, the jump path information of the multi-level interface may include the interactive path information between the desktop, the display interface of the vehicle status tab, the display interface of the smart assistant tab, the display interface of the my smart assistant tab, and the display interface for entering a custom wake-up word. In this way, after the vehicle obtains the voice command, it can first confirm the interface displayed on the current display screen (for example, the display interface of the vehicle status tab) and the target interface (for example, the display interface for entering a custom wake-up word), and then control the display screen to switch from the display interface displaying the vehicle status tab to the display interface displaying the custom wake-up word input through the jump path information of the above multi-level interface.
[0210] Optionally, the jump path information of the multi-level interface may also include the interaction path information between the display interface of other applications and the desktop. For example, when the voice command is obtained, the display screen of the vehicle is displaying the display interface of the in-vehicle map application. After the vehicle obtains the voice command, it can first control the display screen to switch from displaying the display interface of the in-vehicle map application to displaying the desktop based on the jump path information between the display interface of other applications and the desktop; and then control the display screen to switch from displaying the desktop to displaying the display interface for inputting custom wake-up words based on the jump path information between the desktop, the display interface of the vehicle status tab, the display interface of the smart assistant tab, the display interface of the my smart assistant tab, and the display interface for inputting custom wake-up words.
[0211] For example, using the application scenario shown in FIG3 as an example, the first interface may be a desktop, the second interface may be an interface displaying a description of the remote control key emergency mode, and the at least one third interface may be three third interfaces, namely, an initial interface of the owner's guide application, an interface displaying a user manual tab, and an interface displaying a remote control key tab.
[0212] As shown in (a) in Figure 3, when the vehicle receives the voice command "Xiao A Xiao A, I want to check the emergency mode of the remote control key", it can analyze the voice command to obtain the user's intention to view the description content of the remote control key emergency mode. In this way, the vehicle can determine that the target interface (second interface) that the user wants to see is the display interface of the description content of the remote control key emergency mode. The first interface jump information may include the jump path information of the multi-level interface in the process of switching from the desktop to the display interface of the description content of the remote control key emergency mode. The vehicle can control the display screen to switch from displaying the desktop as shown in (a) in Figure 3 to displaying the display interface of the description content of the remote control key emergency mode as shown in (e) in Figure 3 based on the voice command and the first interface jump information.
[0213] Exemplarily, taking the application scenario shown in FIG4 as an example, the first interface may be the initial display interface of the video application, and the second interface may be the display interface of “Movie D”.
[0214] As shown in (a) in Figure 4, when the vehicle receives the voice command "Xiao A Xiao A, I want to watch the movie starring actor 1", it can analyze the voice command to obtain the user's intention of watching the movie starring actor 1. In this way, the vehicle can determine that the target interface (second interface) that the user wants to see is the display interface of the movie starring actor 1. The first interface jump information may include the jump path information of the multi-level interface in the process of switching from the initial display interface of the video application to the display interface of "Movie D" starring actor 1. The vehicle can control the display screen to switch from displaying the initial display interface of the video application as shown in (a) in Figure 4 to displaying the display interface of "Movie D" as shown in (b) in Figure 4 based on the voice command and the first interface jump information.
[0215] Optionally, the jump path information of the multi-level interface may consist of interaction paths of multiple interface elements and operation instructions for each of the multiple interface elements.
[0216] Exemplarily, taking the example of a third interface included in the jump process from the first interface to the second interface, the multiple interface elements include a first interface element and a second interface element, the first interface element is located on the first interface, and the second interface element is located on the third interface, wherein, according to the first voice instruction, controlling the display device to display at least one third interface includes: according to the first voice instruction and the first interface jump information, performing an operation on the first interface element on the first interface to control the display device to switch from displaying the first interface to displaying the third interface; and then displaying the second interface includes: performing an operation on the second interface element on the third interface to control the display device to switch from displaying the third interface to displaying the second interface.
[0217] For example, the multi-level interface jump path information includes jump path information between the desktop, the vehicle control tab display interface, the smart assistant tab display interface, the My Smart Assistant tab display interface, and the display interface for entering a custom wake-up word. This jump path information can be indicated by the interaction path of the Settings app icon - Smart Assistant tab - My Smart Assistant tab - Custom wake-up word control.
[0218] The interaction paths of the above multiple interface elements may include the location information (e.g., pixel coordinates) of each of the multiple interface elements on the interface where they are located. For example, the icon of the Settings application is located on the desktop and has pixel coordinates of [300, 1760], the Smart Assistant tab is located on the display interface of the Vehicle Control tab and has pixel coordinates of [200, 2130], the My Smart Assistant tab is located on the display interface of the Smart Assistant tab and has pixel coordinates of [800, 430], and the custom wake-up word control is located on the display interface of the My Smart Assistant tab and has pixel coordinates of [274, 462].
[0219] Based on the interaction paths of the above-mentioned multiple interface elements and the operations of each of the multiple interface elements, the vehicle can simulate the user's operations to achieve the jump of the display interface. For example, when the vehicle receives the voice command "Xiao A Xiao A, how to set a custom wake-up word", it can analyze the voice command to obtain the user's intention to enter a custom wake-up word, or the user wants to view the display interface for entering a custom wake-up word. The vehicle can determine that the interface element associated with the target interface (the display interface of the custom wake-up word) is the custom wake-up word control based on the user's intention. When the vehicle receives the voice command, the display screen displays the display interface of the vehicle status tab and the display interface of the vehicle status tab includes the smart assistant tab. Then, the vehicle can perform operations on the smart assistant tab, my smart assistant tab and the custom wake-up word control in sequence according to the interaction path information of the above-mentioned setting application icon-smart assistant tab-my smart assistant tab-custom wake-up word control and the operation of each interface element, and finally control the display screen to display the display interface for entering a custom wake-up word.
[0220] Optionally, the display device is controlled to display the second interface according to the first voice command and the first interface jump information, including: determining the second interface according to the first voice command; determining the first position of the second interface in the jump path information of the multi-level interface; determining the jump path from the first interface to the second interface according to the second position of the first interface in the jump path information of the multi-level interface, the first position and the jump path information of the multi-level interface; and controlling the display device to switch from displaying the first interface to displaying the second interface according to the jump path from the first interface to the second interface.
[0221] Exemplary, take the example where the first interface jump information includes the jump order of the multi-level interface and the identification information of each interface in the multi-level interface. After obtaining the voice command, the vehicle can determine the target interface according to the voice command. The vehicle can determine the position of the target interface in the multi-level interface and the position of the current interface in the multi-level interface. The jump path from the current interface to the target interface can be determined by the position of the target interface, the position of the current interface and the jump order. The vehicle can control the display screen to switch from displaying the current interface to displaying the target interface based on the jump path and the identification information of each interface on the jump path.
[0222] For example, the first interface jump information is composed of an interaction path of multiple interface elements and an operation instruction for each of the multiple interface elements. After receiving the voice command, the vehicle can determine the target interface according to the voice command, and determine the position of the interface element 1 associated with the target interface in the interaction path according to the target interface. The vehicle can determine that the interaction path includes an interaction path from interface element 2 to interface element 1, and the interface element 2 is the interface element on the current interface. In this way, the vehicle can operate the multiple interface elements on the interaction path from interface element 2 to interface element 1 in sequence, thereby controlling the display screen to switch from displaying the current interface to displaying the target interface. The first position of the second interface in the jump path information of the multi-level interface can be understood as the position of interface element 1, and the second position of the first interface in the jump path information of the multi-level interface can be understood as the position of interface element 2. The jump path can be the interaction path from interface element 2 to interface element 1.
[0223] The above is explained using the example of an interactive path including two or more interface elements, and the embodiments of the present application are not limited to this. For example, the interactive path may also include one interface element. For example, when the vehicle controls the display screen to display the interface shown in (d) in Figure 3, it detects that the user issues a voice command "Xiao A Xiao A, I want to check the emergency mode of the remote control key". The vehicle can determine that the target interface is the interface shown in (e) in Figure 3 based on the voice command, and the target interface includes a progress bar. In the interface shown in (d) in Figure 3, the progress bar is located at the initial position, and in the interface shown in (e) in Figure 3, the progress bar is located at a position 10 cm downward from the initial position. The vehicle can perform a sliding operation on the progress bar to control the display screen to jump from the current interface to the target interface. The first position of the second interface in the jump path information of the multi-level interface can be understood as the position of the progress bar after it slides down 10 cm from the initial position, and the second position of the first interface in the jump path information of the multi-level interface can be understood as the initial position of the progress bar. The jump path can be switched from the initial position of the progress bar to the position after the progress bar slides down 10 cm.
[0224] Optionally, when the jump path information of the multi-level interface includes multiple jump paths from the first interface to the second interface, the shortest jump path may be selected from them.
[0225] Exemplarily, the multi-level interface jump path information includes a jump path of interface a - interface b - interface c - interface d and a jump path of interface a - interface e - interface f - interface c - interface d. When the initial interface is determined to be interface a and the target interface is interface d, the shortest jump path (the jump path of interface a - interface b - interface c - interface d) can be selected from these two jump paths as the jump path.
[0226] Optionally, the second interface is a display interface for the first application. Before obtaining the user's first voice instruction, the method 800 further includes: obtaining the first interface jump information corresponding to the first application.
[0227] Optionally, the method 800 further includes: receiving a system update data packet sent by a cloud server, where the system update data packet includes updated second interface jump information of the first application.
[0228] For example, when the vehicle is performing a system upgrade, it can obtain a system update data packet from the cloud server. The system update data packet includes the updated interface jump information of the first application. After receiving the updated interface jump information, the vehicle can use the updated interface jump information to overwrite the previous interface jump information.
[0229] The vehicle system update can be performed by the vehicle requesting a system update data packet from the cloud server after detecting a user triggering a system update. Alternatively, the vehicle detects that the user has set automatic system updates, and the cloud server sends the system update data packet to the vehicle during the system update.
[0230] Optionally, the method 800 also includes: sending a first request message to the cloud server, the first request message being used to request interface jump information corresponding to the first version of the second application, the cloud server storing an association relationship between the first version and the first interface jump information; and receiving the first interface jump information sent by the cloud server.
[0231] For example, the second application may be Video Application 1. When the vehicle receives the voice command "Xiao A Xiao A, I want to watch the movie starring Actor 1," it can analyze the user's intent and determine that the user wishes to view the movie starring Actor 1 on the initial interface of Video Application 1 currently displayed on the screen. The version of Video Application 1 is Version 1. The vehicle can send a request to the cloud server, requesting the interface jump information corresponding to Version 1 of Video Application 1. Based on the stored correspondence between the application version information and the interface jump information, the cloud server can send Interface Jump Information 1 corresponding to Version 1 to the vehicle.
[0232] For example, Table 6 shows the correspondence between the version information of the application stored in the cloud server provided in an embodiment of the present application and the interface jump information.
[0233] Table 6
[0234] In one embodiment, the cloud server may analyze the hierarchical relationship of the interface of each application program in advance.
[0235] Thus, the corresponding relationship in the above table 6 is obtained. After receiving the request information sent by the vehicle, the interface jump information 1 corresponding to the version 1 is sent to the vehicle.
[0236] In one embodiment, when the cloud server determines that a version of an application has been updated, it can re-determine the interface jump information corresponding to the updated version based on the hierarchical relationship of different interfaces in the updated application. The cloud server can also use the updated interface jump information to overwrite the previous interface jump information.
[0237] In one embodiment, the request information sent by the vehicle to the cloud server may include version 1 information. Upon receiving the version 1 information, the cloud server may determine that the vehicle wishes to obtain the interface jump information corresponding to version 1 of the video application. The cloud server may analyze the interface jump information of version 1 of the video application 1 to obtain the interface jump information 1 and send the interface jump information 1 to the vehicle.
[0238] In one embodiment, the cloud server may store the correspondence between different versions of the second application and the interface jump information.
[0239] When the first request information is obtained, the interface jump information corresponding to the first version can be searched based on the corresponding relationship. In this way, when the first version is not the latest version of the second application, the terminal device can obtain the first interface jump information corresponding to the first version from the cloud server by sending a request information, thereby helping the user accurately find the display interface that the user wants to view based on the first interface jump information.
[0240] For example, taking the second application being video application 1 as an example, Table 7 shows another correspondence between the version information of the application stored in the cloud server and the interface jump information provided in an embodiment of the present application.
[0241] Table 7
[0242] For example, version 5 is the latest version of video application 1, and the version of video application 1 in the terminal device may be version 1 (a non-latest version). The request information sent by the vehicle to the cloud server may include information about version 1. Upon receiving the information about version 1, the cloud server may send interface jump information 1 corresponding to version 1 to the vehicle according to the corresponding relationship shown in Table 7.
[0243] Optionally, the terminal device stores an association between the second version of the third application and the second interface jump information, and the method 800 also includes: when the third version of the third application does not match the second interface jump information, sending a second request message to the cloud server, the second request message being used to request the interface jump information corresponding to the third version; and receiving the first interface jump information sent by the cloud server.
[0244] For example, taking the third application being video application 1 as an example, the second version can be version 1 of video application 1 (which can be understood as the initial version of video application 1 installed on the vehicle is version 1), and the second interface jump information can be interface jump information 1 corresponding to version 1. When the current version 5 of video application 1 (version 5 can be a version of video application 1 updated through an app store or other means) does not match interface jump information 1, the vehicle can send a request message to the cloud server, requesting the interface jump information corresponding to version 5. Upon receiving the request message, the cloud server can analyze the hierarchical relationship of the interface corresponding to version 5 to obtain the interface jump information 5. The cloud server can send the interface jump information 5 corresponding to version 5 to the vehicle. Alternatively, the cloud server has already determined the interface jump information 5 corresponding to version 5 in advance, and upon receiving the request message, can directly send the interface jump information 5 to the vehicle.
[0245] The above mismatch between the third version and the second interface jump information can be understood as the jump path information when jumping from interface a to interface b in the third version of the third application is the first jump path information, and the jump path information when jumping from interface a to interface b in the second version is the second jump path information, and the first jump path information is different from the second jump path information.
[0246] The above vehicle can send the request information to the cloud server when the version 5 of the video application 1 does not match the interface jump information 1; or, it can also send the request information to the cloud server when the version 5 of the video application 1 does not match the interface jump information 1 and receives the user's voice command "Xiao A Xiao A, I want to watch the movie starring actor 1".
[0247] The interface jump information of the video application 1 sent by the above cloud server to the vehicle can be the interface jump information of different levels of interfaces in the entire video application 1, and can also be the interface jump information for the user's voice command. For example, the vehicle can obtain the user's intention after analyzing the user's voice command, and the vehicle can send a request message to the cloud server, which includes the information of the display interface (the homepage of the video application 1) on the current display screen and the user's intention. When the cloud server obtains the request information, it can send the jump path information of the homepage of the video application 1, the details page display interface of "Movie D", and the playback interface of "Movie D" to the vehicle based on the information of the display interface on the current display screen and the user's intention. For example, the jump path information includes the interactive path of the "Movie D" icon on the homepage of the video application 1 - the playback control on the details page display interface of "Movie D", as well as the operation of the "Movie D" icon and playback control. After receiving the interaction path, the vehicle can, according to the interaction path, first execute a click operation on the "Movie D" icon on the homepage of the video application 1, and switch the control display screen from displaying the homepage of the video application 1 to displaying the details page display interface of "Movie D"; then execute a click operation on the play control on the details page display interface of "Movie D", and switch the control display screen from displaying the details page display interface of "Movie D" to displaying the play interface of "Movie D".
[0248] Optionally, the method 800 further includes: updating the second interface jump information according to the first interface jump information.
[0249] For example, when the vehicle receives the interface jump information 5 corresponding to the version 5, the vehicle can use the interface jump information 5 to overwrite the previous interface jump information 1.
[0250] Optionally, the method 800 further includes: determining that the first interface jump information includes at least part of the text content corresponding to the first voice instruction.
[0251] For example, when the vehicle receives the voice command "Xiao A Xiao A, how do I set a custom wake-up word?", it can analyze the voice command to determine that the user intends to enter a custom wake-up word, or that the user wishes to view the display interface for entering a custom wake-up word. The vehicle can first determine that the first interface jump information includes information for indicating the display interface for entering a custom wake-up word, and then control the display screen to switch from the currently displayed display interface to the display interface for entering a custom wake-up word based on the first interface jump information.
[0252] For example, when a vehicle receives the voice command "Xiao A Xiao A, please open the sunroof," it can analyze the voice command to determine that the user's intention is to open the sunroof. Because the first interface jump information does not include information corresponding to this intention, the vehicle may not perform an operation based on the first interface jump information. Instead, after determining that the voice command is a preset voice control command, it may perform the sunroof opening operation.
[0253] Figure 9 shows a schematic block diagram of a voice interaction device 900 provided in an embodiment of the present application. The device 900 includes: a control unit 910, configured to control a display device to display a first interface, where the first interface is a first-level interface; an acquisition unit 920, configured to acquire a first voice command from a user, where the content of the first voice command includes content indicating an interface at or above the second level; and the control unit 910 is further configured to control the display device to display a second interface corresponding to the first voice command, where the second interface is an interface at or above the second level, based on the first voice command.
[0254] Optionally, the second interface is an interface of a settings application, a video application, a music application, or a car owner's guide application.
[0255] Optionally, the control unit 910 is used to: control the display device to display at least one third interface and then display the second interface according to the first voice instruction.
[0256] Optionally, the control unit 910 is further configured to control the display device to display the jump progress on the first interface and / or the second interface.
[0257] Optionally, the control unit 910 is used to: control the display device to display the second interface according to the first voice command and the first interface jump information, the first interface jump information includes the jump path information of the multi-level interface, and the multi-level interface includes the first interface and the second interface.
[0258] Optionally, the jump path information of the multi-level interface consists of interaction paths of multiple interface elements and operation instructions for each of the multiple interface elements.
[0259] Optionally, the acquisition unit 920 is used to: input the information of the first interface and the content of the first voice instruction into the model to obtain a first interface element and an operation on the first interface element, and the above-mentioned multiple interface elements include the first interface element.
[0260] Optionally, the second interface is a display interface of a first application, and the acquisition unit 920 is further configured to acquire jump information of the first interface for the first application.
[0261] Optionally, the device 900 further includes: a receiving unit, configured to receive a system update data packet sent by a cloud server, where the system update data packet includes interface jump information corresponding to the updated first application.
[0262] Optionally, the device 900 also includes: a sending unit, used to send a first request message to the cloud server, the first request message is used to request interface jump information corresponding to the first version of the second application, and the cloud server stores the association relationship between the first version and the first interface jump information; a receiving unit, used to receive the first interface jump information sent by the cloud server.
[0263] Optionally, the terminal device stores an association between the second version of the third application and the second interface jump information, and the device 900 also includes: a sending unit, used to send a second request message to the cloud server when the third version of the third application does not match the second interface jump information, and the second request message is used to request the interface jump information corresponding to the third version; a receiving unit, used to receive the first interface jump information sent by the cloud server, and the first interface jump information is the interface jump information corresponding to the third version of the third application.
[0264] Optionally, the first interface and the second interface are display interfaces of different applications, wherein the control unit 910 is used to: control the first display area in the display device to display the first interface and control the second display area in the display device to display the second interface according to the first voice command.
[0265] For example, the control unit 910 can be the computing platform in Figure 1 or a processing circuit, processor, or controller in the computing platform. Taking the control unit 910 as the processor 121 in the computing platform as an example, the processor 121 can control the display device to display the first interface. The processor 121 can also control the display device to switch from displaying the first interface to displaying the second interface based on a user's voice command.
[0266] For another example, the acquisition unit 920 may be the computing platform or a processing circuit, processor, or controller in the computing platform in Figure 1. For example, if the acquisition unit 920 is the processor 122 in the computing platform, the processor 122 may acquire the user's voice command.
[0267] The functions implemented by the above control unit 910 and the functions implemented by the acquisition unit 920 can be implemented by different processors, or can also be implemented by the same processor, which is not limited in the embodiment of the present application.
[0268] It should be understood that the division of the various units in the above device is merely a division of logical functions. In actual implementation, they may be fully or partially integrated into a single physical entity, or they may be physically separated. Furthermore, the units in the device may be implemented in the form of a processor calling software; for example, the device includes a processor connected to a memory storing instructions, and the processor calls the instructions stored in the memory to implement any of the above methods or the functions of the various units in the device, where the processor is, for example, a general-purpose processor such as a CPU or a microprocessor, and the memory is a memory within the device or a memory external to the device. Alternatively, the units in the device may be implemented in the form of hardware circuits, and the functions of some or all of the units may be implemented through the design of the hardware circuits. The hardware circuits may be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units may be implemented through the design of the logical relationships between the components within the circuits. In another implementation, the hardware circuit may be implemented using a PLD, such as an FPGA, which may include a large number of logic gate circuits, and the connections between the logic gate circuits may be configured using a configuration file to implement the functions of some or all of the above units. All units of the above apparatus may be implemented entirely in the form of software called by a processor, or entirely in the form of hardware circuits, or partially in the form of software called by a processor and the rest in the form of hardware circuits.
[0269] Each unit in the above device can be one or more processors (or processing circuits) configured to implement the above method, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.
[0270] In addition, the various units in the above apparatus may be fully or partially integrated together, or may be implemented independently. In one implementation, these units are integrated together and implemented in the form of a system-on-chip (SoC). The SoC may include at least one processor for implementing any of the above methods or implementing the functions of the various units of the apparatus. The at least one processor may be of different types, for example, including a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.
[0271] An embodiment of the present application also provides a device, which includes a processing unit and a storage unit, wherein the storage unit is used to store instructions, and the processing unit executes the instructions stored in the storage unit so that the device executes the method or steps performed by the above embodiment.
[0272] Optionally, if the apparatus is located in a terminal device, the processing unit may be the processor 121 - 12n shown in FIG. 1 .
[0273] An embodiment of the present application further provides a system, which includes a display device and a computing platform, and the computing platform includes the above-mentioned device 900.
[0274] An embodiment of the present application further provides a terminal device, which may include the above-mentioned apparatus 900, or the terminal device may include the above-mentioned system.
[0275] Exemplarily, the terminal device may be a vehicle.
[0276] An embodiment of the present application further provides a computer program product, which includes: computer program code, which, when executed by one or more processors, enables a terminal device to execute the above method.
[0277] An embodiment of the present application further provides a computer-readable storage medium, which stores program code. When the computer program code is executed by one or more processors, the terminal device executes the above method.
[0278] Figure 10 shows a schematic block diagram of a chip 1000 provided in an embodiment of the present application. The chip 1000 includes a processor 1010, a data interface 1020, and a memory 1030. The processor 1010 reads instructions stored in the memory 1030 through the data interface 1020 to enable the terminal device to execute the above method.
[0279] During implementation, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor or by instructions in the form of software. The method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or a power-on erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in conjunction with its hardware. To avoid repetition, it will not be described in detail here.
[0280] It should be understood that in the embodiment of the present application, the memory may include a read-only memory and a random access memory, and provide instructions and data to the processor.
[0281] It should also be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0282] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0283] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0284] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0285] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0286] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0287] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0288] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be covered and fall within the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A voice interaction method, It is characterized in that include: Controlling the display device to display a first interface, where the first interface is a first-level interface; Acquire a first voice command of a user, wherein the content of the first voice command includes content indicating an interface above the second level; According to the first voice instruction, the display device is controlled to display a second interface corresponding to the first voice instruction, where the second interface is an interface above the second level.
2. The method according to claim 1, It is characterized in that The second interface is an interface of a car owner guide application, a setting application, a video application or a music application.
3. The method according to claim 1 or 2, It is characterized in that The controlling the display device to display a second interface corresponding to the first voice command according to the first voice command includes: According to the first voice instruction, the display device is controlled to display at least one third interface, and then display the second interface.
4. The method according to any one of claims 1 to 3, It is characterized in that The method further comprises: Control the display device to display the jump progress on the first interface and / or the second interface.
5. The method according to claim 1 or 2, It is characterized in that The controlling the display device to display a second interface corresponding to the first voice command according to the first voice command includes: According to the first voice command and the first interface jump information, the display device is controlled to display the second interface, the first interface jump information includes the jump path information of the multi-level interface, and the multi-level interface includes the first interface and the second interface.
6. The method according to claim 5, It is characterized in that The jump path information of the multi-level interface consists of interaction paths of multiple interface elements and operation instructions for each of the multiple interface elements.
7. The method according to claim 6, It is characterized in that Also includes: The information of the first interface and the content of the first voice instruction are input into a model to obtain a first interface element and a first operation on the first interface element, wherein the plurality of interface elements include the first interface element.
8. The method according to claim 5 or 6, It is characterized in that The second interface is an interface of the first application program, and the method further includes: Obtain the first interface jump information corresponding to the first application.
9. The method according to claim 8, It is characterized in that The first interface jump information is included in a system update data packet received from a cloud server.
10. The method according to claim 8, It is characterized in that The obtaining the first interface jump information corresponding to the first application comprises: Sending first request information to a cloud server, where the first request information is used to request interface jump information corresponding to a first version of a first application; The first interface jump information sent by the cloud server is received, where the first interface jump information is interface jump information corresponding to the first version of the first application.
11. The method according to claim 8, It is characterized in that The obtaining the first interface jump information corresponding to the first application comprises: When the first version of the first application does not match the second interface jump information, sending second request information to the cloud server, where the second request information is used to request the interface jump information corresponding to the first version; The first interface jump information sent by the cloud server is received, where the first interface jump information is interface jump information corresponding to the first version of the first application.
12. The method according to any one of claims 1 to 11, It is characterized in that The first interface and the second interface are interfaces of different application programs. Wherein, controlling the display device to display a second interface corresponding to the first voice command according to the first voice command further includes: According to the first voice command, a first display area in the display device is controlled to display the first interface and a second display area in the display device is controlled to display the second interface.
13. A voice interaction device, It is characterized in that include: A control unit, used to control the display device to display a first interface, where the first interface is a first-level interface; An acquisition unit, configured to acquire a first voice instruction of a user, wherein the content of the first voice instruction includes content indicating an interface above the second level; The control unit is further used to control the display device to display a second interface corresponding to the first voice instruction according to the first voice instruction, where the second interface is an interface above the second level.
14. The device according to claim 13, It is characterized in that The second interface is an interface of a car owner guide application, a setting application, a video application or a music application.
15. The device according to claim 13 or 14, It is characterized in that The control unit is used to control the display device to display at least one third interface and then display the second interface according to the first voice instruction.
16. The device according to any one of claims 13 to 15, It is characterized in that The control unit is further used to control the display device to display the jump progress on the first interface and the second interface.
17. The device according to claim 13 or 14, It is characterized in that The control unit is used for: According to the first voice command and the first interface jump information, the display device is controlled to display the second interface, the first interface jump information includes the jump path information of the multi-level interface, and the multi-level interface includes the first interface and the second interface.
18. The device according to claim 17, It is characterized in that The jump path information of the multi-level interface consists of interaction paths of multiple interface elements and operation instructions for each of the multiple interface elements.
19. The device according to claim 18, It is characterized in that The acquisition unit is used to: The information of the first interface and the content of the first voice instruction are input into a model to obtain a first interface element and an operation on the first interface element, wherein the plurality of interface elements include the first interface element.
20. The device according to claim 17 or 18, It is characterized in that The second interface is the interface of the first application program, The acquisition unit is further used to acquire the first interface jump information corresponding to the first application.
21. The device according to claim 20, It is characterized in that The first interface jump information is included in a system update data packet received from a cloud server.
22. The device according to claim 20, It is characterized in that The device also includes: A sending unit, sending first request information to a cloud server, wherein the first request information is used to request interface jump information corresponding to a first version of a first application; The receiving unit is used to receive the first interface jump information sent by the cloud server, where the first interface jump information is the interface jump information corresponding to the first version of the first application.
23. The device according to claim 20, It is characterized in that The device also includes: A sending unit, configured to send second request information to a cloud server when the first version of the first application does not match the second interface jump information, wherein the second request information is used to request the interface jump information corresponding to the first version; The receiving unit is used to receive the first interface jump information sent by the cloud server, where the first interface jump information is the interface jump information corresponding to the first version of the first application.
24. The device according to any one of claims 13 to 23, It is characterized in that The first interface and the second interface are interfaces of different application programs. Wherein, the control unit is also used for: According to the first voice command, a first display area in the display device is controlled to display the first interface and a second display area in the display device is controlled to display the second interface.
25. A voice interaction device, It is characterized in that include: Memory for storing computer programs; A processor, configured to execute the computer program stored in the memory, so that the apparatus performs the method according to any one of claims 1 to 12.
26. A system, It is characterized in that It comprises a display device and a computing platform, wherein the computing platform comprises the device as claimed in any one of claims 13 to 25.
27. A terminal device, It is characterized in that Comprising an apparatus as claimed in any one of claims 13 to 25, or comprising a system as claimed in claim 26.
28. The terminal device according to claim 27, It is characterized in that The terminal device is a vehicle.
29. A computer-readable storage medium, It is characterized in that A computer program is stored thereon, and when the computer program is executed by a computer, the method according to any one of claims 1 to 12 is implemented.
30. A chip, It is characterized in that The chip includes a processor and a data interface, and the processor reads instructions stored in a memory through the data interface to execute the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Voice control method and device, electronic equipment and storage medium
CN111968639A
Page tag processing method and device, computer equipment and medium
CN112559101A
Voice control method and device, electronic equipment and storage medium
CN114121013A
Display apparatus
US20210409832A1