Interface display method, apparatus and system during execution of automatic task
By displaying the identifiers of meta-tasks and execution steps on the display interface, the problem of poor user experience during the execution of automated tasks is solved, task transparency and user interaction are achieved, and the efficiency and security of task execution are improved.
Patent Information
- Application Number
- PCT/CN2025/087776
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-18
- Filing Date
- 2025-04-08
- Publication Date
- 2025-10-23
AI Technical Summary
During the execution of automated tasks, existing technologies make it difficult for users to transparently display and interact with task progress in real time, resulting in a poor user experience and the risk of privacy leakage.
By displaying the identifiers of meta-tasks and execution steps on the display interface, users are allowed to interact with the interface in real time, adjust or modify task content, and improve user experience.
It realizes transparent display and user interaction of automated tasks, meets users' needs for understanding task progress, and improves the efficiency and safety of task execution.
Smart Images

Figure CN2025087776_23102025_PF_FP_ABST
Abstract
Description
Interface display method, device and system in automatic task execution process TECHNICAL FIELD
[0001] The present application relates to the technical field of computers, and particularly relates to an interface display method, device and system in an automatic task execution process. BACKGROUND
[0002] An AI large model has strong natural language processing capability, can learn the rules and knowledge of language from a large amount of text data, and can accurately understand the intention instruction input by a user and give a reply corresponding to the intention instruction. An intelligent device is usually configured with a digital assistant. A user can realize a series of automatic tasks by calling the digital assistant.
[0003] When the AI large model is combined with the digital assistant, the capability of the AI model is given to the digital assistant, so that the digital assistant has strong natural language processing capability, thereby making the digital assistant more usable, capable of automatically understanding voice input, automatically determining a task solution after intention understanding, and then executing an automatic task.
[0004] At present, in the process of executing an automatic task, the feedback mode of the digital assistant mainly includes two modes. One is to directly feed back a final result page. This mode directly saves the interaction layer between the digital assistant and the user, and the user is difficult to intervene in the process of the automatic task, and many times cannot meet the actual needs of the user. The other mode is to simulate the operation process of the user, open an APP, display all action processes by simulating the screen clicking of the user, and the user can clearly view the execution progress of the automatic interface. However, this mode exposes the privacy of the user, such as viewing historical travel records.
[0005] Therefore, how to design a safe and reliable display interface that can meet the needs of the user to understand and interact with the execution progress of the automatic task is a problem to be solved at present. SUMMARY
[0006] The application provides a method for displaying an interface during execution of an automated task by a digital assistant. The digital assistant receives a user-specified automated task, and executes the automated task, wherein the automated task comprises at least one subtask determined based on the automated task. The at least one subtask is different for different automated tasks. The at least one identifier is displayed on the display interface, and the at least one identifier is mainly used to represent the at least one subtask and / or an execution step in the at least one subtask. The method enables the user to see the execution process of the entire automated task on the display interface, and enables the user to interact with the display interface in real time to modify or adjust the subtask or the execution step in the automated task. The method improves the user experience, and meets the requirement of real-time interaction of the user with the display interface, thereby making the executed automated task more efficient and meeting the requirement of the user.
[0007] In a first aspect, the application provides a method for displaying an interface during execution of an automated task, applied to an electronic device comprising a digital assistant and a display interface, and comprising the following steps: the digital assistant receives a user instruction and generates an automated task; the digital assistant executes at least one subtask determined based on the automated task; and at least one identifier is displayed on the display interface, and the identifier is used to represent the subtask and / or an execution step in the subtask.
[0008] In the present solution, the digital assistant receives a user-specified automated task, determines at least one subtask based on the automated task, executes the at least one subtask, and displays the at least one subtask and an execution step in the at least one subtask on the display interface. The method for transparently displaying the execution process of the automated task on the display interface of the user can facilitate the user to interact with the display interface in real time to adjust or modify the content in the subtask or the execution step, improve the user experience, and meet the requirement of real-time interaction of the user.
[0009] In a possible implementation manner of the first aspect, each subtask comprises at least one key factor, and the identifier of the subtask is determined based on the at least one key factor, wherein the key factor is at least one of an application icon of a called application program and a virtual control in an application interface.
[0010] In the present solution, the key factor is an application icon of a called application program and / or a virtual control in an application interface in the automated task. For example, the task of finding a charging pile generally comprises calling a map APP, and therefore the identifier of the corresponding map software, such as Baidu Map, Gaode Map, etc., can be the key factor. The subtask can be determined based on the at least one key factor.
[0011] In a possible implementation manner of the first aspect, the key factor further includes a generalization icon, and the generalization icon is a predefined icon.
[0012] In the solution, the key factor can also be a self-defined icon. For example, for some documents, an icon of a document type can be used, and for some pictures, an icon of a picture type can be used. Both of them are self-defined. In a possible implementation manner, the function of the corresponding execution step can be directly associated with the icon.
[0013] In a possible implementation manner of the first aspect, in the process of executing the at least one meta task, the user interacts with the display interface by at least one of voice and touch screen to modify or replace the at least one meta task.
[0014] In a possible implementation manner of the first aspect, the method further includes that the display interface further includes a pop-up window or a card, and the pop-up window or the card is used to show the execution process of the corresponding meta task.
[0015] In the solution, the meta task and some execution steps in the automation task are displayed on the interface. Some meta tasks are displayed on the interface in the form of a card or a pop-up window. The entire execution process can be shown, and the modification and adjustment of the information in the meta task or the interface display content or the execution step can be supported for real-time interaction with the user, so as to better meet the needs of the user and improve the user experience.
[0016] In a possible implementation manner of the first aspect, the at least one meta task is determined based on at least one of the number of called APPs and the number of page jumps.
[0017] In the solution, how to determine the meta task is mainly described. Usually, the number of meta tasks is determined based on the number of called APPs and the number of page jumps. One APP can be used as one meta task, and one execution page can also be used as a separate meta task. In this way, the user can completely see the execution process of the entire automation task, which is more transparent and meets the needs of real-time interaction and modification of the user. In some scenarios, the interaction of the user can guide the digital assistant to better and more efficiently execute other automation tasks.
[0018] In a possible implementation manner of the first aspect, the key factor is determined based on the execution steps in the meta task. Each meta task includes at least one execution step, and the at least one execution step collectively completes all functions of the meta task.
[0019] In a possible implementation manner of the first aspect, the user interacts with the window or the card by at least one of voice and touch screen to specify or modify content in the window or the card.
[0020] In this solution, the user can access and modify the corresponding execution process in real time by implementing the interactive manner, which not only effectively improves the execution efficiency of the digital assistant, but also realizes the transparentization of the execution process and receives real-time supervision and guidance of the user, so as to obtain more satisfactory results.
[0021] In a possible implementation manner of the first aspect, the method further includes: detecting a preset operation on the display area, and displaying a dialogue interface of the digital assistant, the dialogue interface including task description information of the automated task.
[0022] In a second aspect, applied to an electronic device with a display interface, the method includes: receiving a user instruction by the electronic device, and generating an automated task, the automated task including at least one execution action; executing the at least one execution action by the electronic device; and displaying at least one identifier on the display interface, the identifier being used to represent an executed or executing execution action.
[0023] In a possible implementation manner of the second aspect, the execution action at least includes calling an APP, page jumping, clicking a control, inputting voice, and typing text.
[0024] In a possible implementation manner of the second aspect, in a process of executing the at least one execution action by the electronic device, the user interacts with the display interface by at least one of voice and touch screen to modify or replace the execution action.
[0025] In this solution, the automated task execution process is transparently displayed on the display interface of the user, which can facilitate the user to interact with the display interface in real time to adjust or modify content in the meta task or the execution step, improve the user experience, and meet the real-time interaction requirement of the user.
[0026] In a possible implementation manner of the second aspect, the method further includes: the display interface further includes a window or a card, and the window or the card is used to display the execution process of the automated task.
[0027] In this solution, some execution steps in the automated task are displayed on the interface, and the execution steps are displayed on the interface in the form of a card or a window, the entire execution process can be displayed, and the modification and adjustment of information in the meta task or the interface display content or the execution step can be supported to meet the requirement of the user and improve the user experience.
[0028] In a third aspect, an electronic device with a display interface is provided, comprising: a digital assistant configured to receive a user instruction and generate an automated task, and further configured to execute at least one subtask determined based on the automated task; and the display interface configured to display at least one identifier representing the subtask and / or an execution step in the subtask.
[0029] In a possible implementation of the third aspect, each subtask comprises at least one key factor based on which the identifier of the subtask is determined, and the key factor is at least one of an application icon of an invoked application and a virtual control in an application interface.
[0030] In a possible implementation of the third aspect, the key factor further comprises a generalized icon, and the generalized icon is a predefined icon.
[0031] In a possible implementation of the third aspect, during execution of the at least one subtask, the user interacts with the display interface by at least one of voice and touch screen to modify or replace the at least one subtask.
[0032] In a possible implementation of the third aspect, the method further comprises that the display interface further comprises a pop-up window or a card configured to show an execution process of a corresponding subtask.
[0033] In a possible implementation of the third aspect, the at least one subtask is determined based on at least one of a number of invoked applications and a number of page jumps.
[0034] In a possible implementation of the third aspect, the key factor is determined based on an execution step in the subtask, and each subtask comprises at least one execution step which collectively completes all functions of the subtask.
[0035] In a possible implementation of the second aspect, the user interacts with the pop-up window or the card by at least one of voice and touch screen to specify or modify content in the pop-up window or the card.
[0036] In a possible implementation of the third aspect, further comprising: detecting a preset operation on the identifier display area, and displaying a dialogue interface of the digital assistant, and the dialogue interface comprises task description information of the automated task.
[0037] In a fourth aspect, the method is applied to an electronic device with a display interface, and includes: receiving a user instruction by the electronic device, and generating an automated task, the automated task including at least one execution action; executing the at least one execution action by the electronic device; and displaying at least one identifier on the display interface, the identifier indicating the executed or executing execution action.
[0038] In a possible implementation of the fourth aspect, the execution action includes at least invoking an APP, page jumping, clicking a control, inputting a voice, and typing a text.
[0039] In a possible implementation of the fourth aspect, during the execution of the at least one execution action by the electronic device, the user interacts with the display interface by at least one of a voice and a touch screen, and the execution action is modified or replaced.
[0040] In a possible implementation of the fourth aspect, the method further includes: the display interface further includes a pop-up window or a card, and the pop-up window or the card is used to show an execution process of the automated task.
[0041] In a fifth aspect, an electronic device with a display interface is provided, and the electronic device includes a processor coupled with a memory, and the processor is configured to store instructions, and when the instructions are executed by the processor, the electronic device performs the method in any one of the first aspect or the second aspect.
[0042] In a sixth aspect, a computer readable storage medium is provided, and includes computer instructions, and when the computer instructions are executed on a computer system, the computer system implements the method in any one of the first aspect or the second aspect.
[0043] In a seventh aspect, a computer program product is provided, and includes instructions, and when the instructions are executed, the computer implements the method in any one of the first aspect or the second aspect.
[0044] The advantages of the second aspect to the seventh aspect are as described in the first aspect, and will not be described here. BRIEF DESCRIPTION OF DRAWINGS
[0045] FIG. 1 is a schematic diagram of a center screen interface of a car machine provided in the embodiment;
[0046] FIG. 2 is a schematic diagram of a division of an automated task based on a key factor provided in the embodiment;
[0047] Fig. 3(a), Fig. 3(b), Fig. 3(c), Fig. 3(d) and Fig. 3(e) are schematic diagrams of a process of receiving a voice instruction by a digital assistant and then performing an automated task according to the present embodiment;
[0048] Fig. 4 is a schematic diagram of an automated task execution process of searching for a charging pile according to the present embodiment;
[0049] Fig. 5(a) and Fig. 5(b) are schematic diagrams of an automated task of online ordering movie tickets according to the present embodiment;
[0050] Fig. 6(a) and Fig. 6(b) are schematic diagrams of displaying a key factor by a floating card according to the present embodiment;
[0051] Fig. 7(a) and Fig. 7(b) are schematic diagrams of an interface of user interaction with a key factor according to the present embodiment. DETAILED DESCRIPTION
[0052] In order to make the objects, technical solutions and advantages of the present application clearer, the embodiments of the present application are described below with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Those skilled in the art can know that the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems as new application scenarios appear.
[0053] The terms "first", "second", and the like in the specification of the present application, claims, and above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the descriptions used in this way can be interchanged under appropriate circumstances, so that the embodiments can be implemented in an order other than that illustrated or described in the present application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or modules does not necessarily limit to those steps or modules clearly listed, but can include other steps or modules not clearly listed or inherent to these processes, methods, products or devices. The naming or numbering of the steps appearing in the present application does not mean that the steps in the method flow must be executed in the order / time sequence indicated by the naming or numbering. The flow steps that have been named or numbered can change the execution order according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved.
[0054] The division of units appearing in the present application is a logical division. In actual application, another division mode can be used, for example, a plurality of units can be combined or integrated in another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be through an interface, and the indirect coupling or communication connection between the units can be electrical or other similar forms, which are not limited in the present application. In addition, the units or sub-units described as separate components can or can not be physically separated, can or can not be physical units, or can be distributed in a plurality of circuit units. According to actual needs, part or all of the units can be selected to achieve the purpose of the present application.
[0055] For ease of understanding, some technical terms related to the embodiments of the present application are introduced below.
[0056] (1) Intelligent cockpit
[0057] The intelligent cockpit integrates advanced information technology and automation systems in the car interior, upgrades and optimizes the in-car environment, and interconnects with the outside environment to provide a higher level of driving experience and riding comfort. This concept covers multiple aspects, including entertainment, information acquisition, vehicle control for drivers and passengers, and interaction with the outside world.
[0058] (2) Digital assistant
[0059] The digital assistant is a high-level computer program that usually relies on the Internet to run and can simulate human conversation with users. It can use advanced artificial intelligence (AI), natural language processing, natural language understanding, and machine learning techniques for self-learning to provide personalized conversation experience for users.
[0060] (3) Meta task
[0061] In the present embodiment, in one implementation, the following steps will be performed: finding a target application -> starting the target application -> the task of jumping to the target page and obtaining information on the target page is a meta task; in another implementation, the meta task can be a key stage task in the process of achieving a certain target task, and all meta tasks are linked together to complete the target task. Therefore, the number of meta tasks is determined based on the automation task. For different automation tasks, since the complexity of the process execution is different, the number of divided meta tasks is also different.
[0062] (4) Key factor
[0063] In the embodiment, the key factor is that in the execution step of each meta task, it has a decisive influence on the subsequent step, and directly affects the application called by the subsequent step or the interface of the application. The key factor is usually an application icon or a virtual control of the application interface of the application program. Generally, at least one key factor can be included in a meta task, and each key factor corresponds to a key execution step.
[0064] (5) Generalization icon
[0065] In the embodiment, the generalization icon refers to an icon that can be customized in advance by the user, such as a text icon, a general icon for representing a picture, or an icon for representing a document, etc. The most important feature of the generalization icon is that it can associate the corresponding meta task function according to the icon.
[0066] The present application provides an interface display method in an automatic task execution process. The method meets the user interaction needs in the task execution process by displaying the automatic task process on the interface. The user can intervene or instruct the execution of the task through interaction with the digital assistant. The process does not expose the user's privacy and can effectively improve the user experience.
[0067] Please refer to FIG. 1, which is an architecture schematic diagram of an application scenario provided by an embodiment of the present application. As shown in FIG. 1, the application scenario includes an electronic device 001 and a user 002. The scenario is mainly used to show the interaction between the electronic device 001 and the user 002. In the embodiment, the electronic device 001 is a smart cockpit or a smart car: it is built-in with a car operating system, the operating system can be built-in with a digital assistant, and the digital assistant can be connected to the cloud through the communication device of the car to call the large model capability of the cloud; or the digital assistant itself has certain large model capability. The user interacts with the digital assistant through voice and other ways, so that the digital assistant creates an automatic task to complete the user's expressed intention and meet the user's needs. The user 002 is usually the car owner, who interacts with the digital assistant to realize a series of automatic operations during driving.
[0068] In a possible implementation manner of the embodiment, the smart cockpit is built-in with a display interface, and the user 002 interacts with the digital assistant through the display interface.
[0069] In another embodiment, the electronic device 001 can also be a terminal device with a display interface, such as a mobile phone, a tablet, a computer, a watch, a robot, etc., which carries a digital assistant. The user 001 can interact with the digital assistant through the display interface. The digital assistant has the ability of an AI large model, which can be implemented by at least one of calling a cloud large model and a built-in large model. The electronic device can call the digital assistant to provide services in any interface. In order to facilitate viewing and setting, the execution process of the automated task is also displayed on the display interface, so as to facilitate the user 001 to view in real time, human intervention and setting.
[0070] In the present embodiment, the electronic device receives input information and creates an automated task based on the understanding of the input information. Specifically, the electronic device can call the digital assistant to provide services in any display interface. Since the digital assistant is usually resident in the background and in standby state, it can detect and receive the input of the user in real time. Some digital assistants can also be activated by voice wake-up and other operations. The activated digital assistant can also detect and receive the input of the user in order to better provide services for the user.
[0071] In one possible implementation of the present embodiment, the digital assistant has the ability to receive user input. The user can provide input information by at least one of voice input and text input. For different input messages received, the digital assistant can perform semantic understanding and intent recognition on the input information based on the large model capability. Finally, according to the semantic and intent information, a series of automated tasks are generated. The automated task usually includes a plurality of steps, which are usually completed by a plurality of computer instructions. The digital assistant can generate computer instructions and execute each step layer by layer until all steps are finally executed.
[0072] Please refer to FIG. 1, which is a schematic diagram of a center screen interface of a car machine provided in the embodiment. In the application scenario of the car machine, the digital assistant can be displayed on the front-end interface or not, but the digital assistant is always resident in the background and can be awakened by voice or real-time monitoring of the input of text or voice on the front end. In the application scenario of the car machine, the way of voice input can be more, especially during the driving process of the user, the way of using voice input to interact is more secure and more convenient. Of course, for the user waiting for a red light, the way of text input or touch screen interaction is also a possible way. As shown in FIG. 1, the digital assistant can call the microphone to receive the voice input of the user in real time. The user inputs the voice "help me continue to play the last anime", and the digital assistant receives the voice. Based on the intention understanding of the input voice information, an automatic task of opening the historical play anime is created. The execution process of the automatic task can be divided into the following steps:
[0073] 1) Find the target application: the digital assistant will first execute the meta task of finding the target software. In the process of finding the target software, the user's favorite software will be selected according to the user's preference. For example, the digital assistant will consider selecting the target software that the user uses most frequently, and secondly, it will also consider the last play software used by the user.
[0074] 2) Start the target application: the digital assistant executes the action of opening the historical play record page based on the target software found;
[0075] 3) Execute instructions in the target application: find the historical record of the last played anime in the historical play record page, and execute the action of playing the first video of the historical play list.
[0076] The above is an example of the simplest task scenario, which can be described as three actions / steps in a meta task. The next action depends on the execution result of the previous action / step, and the digital assistant can complete the automatic task through the automatic task flow.
[0077] However, the user's intention can be a simple intention or a complex intention. The simple intention of "continue to play" described above only needs to call an application to achieve. However, in the implementation of complex intention, the above simple logic is difficult to achieve the user's intention, and often needs the cooperation of multiple applications to achieve the user's intention, that is, the automatic task can be logically divided into multiple meta tasks, each meta task corresponds to the call of an application program.
[0078] It should be noted that the meta task is a logical concept defined for ease of description in the embodiment, and the electronic device does not actually have the concept or definition of the meta task. When the electronic device performs the automated task, it only needs to execute according to the task steps created based on the intention recognition result. Each task step can correspond to a meta task, and a meta task can also be executed across task steps.
[0079] The scenario in which one automated task is divided into multiple meta tasks, for example, in the scenario of booking a ticket, the ticket booking application may not have information such as flight punctuality rate statistics. When booking a ticket, if the flight punctuality rate is used as one of the reference factors, the flight punctuality rate information needs to be obtained through an application that has flight punctuality rate statistics. Therefore, the ticket booking application and the flight punctuality rate query application need to be used to complete the ticket booking task.
[0080] Taking the scenario of finding a charging pile in a vehicle as an example, after the electronic device receives the input instruction of "finding a nearby charging pile for charging, supporting fast charging, and trying to be as cheap as possible", an automated task is created. The automated task can generally include the following steps:
[0081] 1) Selecting a charging pile to find a charging pile using the nearby function on the map;
[0082] 2) Obtaining fast charging and price information of all charging piles nearby;
[0083] 3) Filtering charging piles with fast charging function and sorting them according to price to obtain multiple candidate charging piles;
[0084] 4) Marking the candidate charging piles on the map and determining a target charging pile;
[0085] 5) Setting the target charging pile as the navigation end point and starting the map navigation function.
[0086] In the above steps, steps 1, 4, and 5 need to call the map application to complete, but the map application usually only has the function of positioning the location of the nearby charging pile, and does not have the information of fast charging, price, etc. of the charging pile. Therefore, step 2 needs to input the address information of each charging pile on the map to the local life application to query whether each charging pile has the fast charging function and the price information.
[0087] That is, the above-mentioned automated task includes two sub-tasks, one sub-task is to find the target charging pile on the map and navigate, and the other sub-task is to obtain the fast charging, price and other information of the charging pile. The first sub-task corresponds to steps 1, 4 and 5, and the second sub-task corresponds to step 2. After the first sub-task executes step 1, it is suspended first, and then the second sub-task is started to grab the fast charging, price and other information. Then, after the AI or digital assistant completes the judgment of step 3, the first sub-task continues.
[0088] The embodiment aims to enhance the transparency of the AI or digital assistant in the task execution process, so that the user can understand the key nodes of the actions being executed by the AI or digital assistant, which is usually performed in a black box manner in the prior art. In order to facilitate the user to understand the execution process of the automated task, in the embodiment, the key factors of the execution actions in the automated task process are displayed to the user in a visual manner, so that the user can generally understand the execution process of the automated task of the AI or digital assistant and the actions being executed.
[0089] As for the execution action, it can be understood as a more detailed specific execution operation than the task steps of the above-mentioned decomposed automated task, for example, starting an application is an execution action, and entering the historical playback record page from the home page of the video playback application is also an action. It can be understood that one execution action corresponds to one simulated click event triggered by the AI or digital assistant.
[0090] Based on this understanding, the key factor of the execution action can also be described as the key factor of the task step or the key factor of the entire automated task, and from the perspective of the sub-task, it can also be described as the key factor of the sub-task. The difference lies in that one execution action corresponds to one key factor, while the implementation of a task step may require one or more actions, so it can correspond to one or more key factors, and the entire automated task and the sub-task usually have multiple key factors.
[0091] When displaying the key factors, if all the key factors of the execution actions are displayed, there will be too many key factors, which is not conducive to the user to quickly capture information, therefore, the key factors of part of the execution actions can be displayed. For example, but not limited to, one key factor is displayed for each task step. Specifically, the AI or digital assistant can judge the importance of the execution action, and only display the key factors of the execution actions with high importance.
[0092] In the following, the key factor will be described as the key factor of the task step or the key factor of the sub-task. Each task step can correspond to one or more key factors, and each sub-task can correspond to one or more key factors, which is determined by the execution actions included in the task step and the sub-task, and the importance of the execution action.
[0093] The AI or digital assistant can be simply understood as simulating a person operating an application installed on an electronic device when performing an automated task. The AI or digital assistant enters a target page by opening the application, interface jumping, and other simulated operations, and finally obtains target information. In this process, interface jumping and information acquisition are achieved by simulating clicking of the control in the application interface. The simulated clicking is an execution action, and in the operating system of the electronic device, the simulated clicking is usually achieved by simulating the clicking event of the control. Therefore, it can also be understood that a simulated clicking event trigger corresponds to an execution action. The object of the simulated clicking is usually the key factor, or the simulated clicking object corresponding to the simulated clicking event is the key factor. In this embodiment, the process of simulating clicking and jumping of the interface is still black-boxed and not shown to the user, but only the key factor is shown to the user.
[0094] In this embodiment, when the key factor is shown to the user, the associated identifier of the key factor is displayed. The associated identifier is used to represent the corresponding key factor, and the associated identifier is displayed on the screen. Therefore, after determining the key factor, the associated identifier of the key factor needs to be displayed on the screen based on the key factor.
[0095] Specifically, after the key factor is determined, the associated identifier corresponding to the key factor can be determined based on the key factor, and when the step of performing an automated task is executed, the associated identifier of the key factor of the process interface corresponding to the step is displayed. The key factor is usually a control, and therefore the associated identifier can be a thumbnail of the icon corresponding to the control. For example, if the execution action is to click an application icon to start the application, the key factor corresponding to the execution action is the application icon of the application; if the execution action is to click the historical playback record control, the key factor corresponding to the execution action is the historical playback record control
[0096] For controls, in general, they can be divided into the following categories according to types. The first category is an icon control, which is mainly displayed in the interface by an icon, such as common search, return, and other controls. The second category is a text control, which indicates the function of the control by text, such as common confirm, cancel, and other controls. The third category is an icon + text control, which displays text below the icon, and common application icons. The fourth category is a file zoom control, which is common in chat windows for sending files and historical video thumbnail windows in historical playback records.
[0097] When the control corresponding to the key factor includes a graphic file, the graphic file of the control can be extracted from the process interface as the association identifier. For example, for the first, third and fourth categories described above, there is a graphic file, and the graphic file can be extracted as the association identifier. When there is a graphic file and text, the graphic file and text can also be extracted together, and the layout of the graphic and text in the display style of the key factor in the process interface is used as the association identifier. For example, the application icon is followed by the application name, and the graphic of the application icon displaying the application name at the bottom can be used as the association identifier.
[0098] For text controls, the text can also be extracted as an association identifier, but text is not as helpful as graphics for users to quickly understand. Therefore, for text controls, a preset generic icon can be used to represent, or a generic icon can be used in conjunction with the graphic style of the extracted text from the control to represent the association identifier. Specifically, when selecting a generic icon, the generic icon can be selected according to the function tag corresponding to the control, for example, for a settings class, a generic icon for a settings class is used, and for a video class control, a generic icon for a video class is used.
[0099] The association identifier of the key factor is displayed on the interface, and the user can interact with the association identifier to intervene in the execution of the automated task. For example, when the association identifier of the key factor is an application icon, the user can click the application icon, and the AI or digital assistant can display application icons of the same application on the interface according to the operation, and change another application to play the historical video by selecting a different application icon.
[0100] In order to reflect the continuity of the automated task, the displayed association identifier will always be displayed before the completion of an automated task, and the association identifier of the key factor will be displayed in sequence according to the step execution order of the automated task, so that the display of the association identifier can help the user understand the real-time task progress of the digital assistant, and support the user to make settings and interventions during the task execution process according to actual needs.
[0101] Please refer to FIG. 2 and FIG. 3, which show all the process interfaces of performing the aforementioned “continue playing” automation task. FIG. 2 is an embodiment in which the automation task is receiving the voice instruction “help me continue playing the last episode of the TV series” in the car machine automatic driving interface. FIG. 3 shows the sequential jump changes of all interfaces. In FIG. 3(a), the original interface of the car machine task receiving the user input voice instruction is shown, that is, the automatic driving interface, in which the road being driven, the surrounding vehicles, and the driving model of the car on the road are displayed, as well as some building facilities and real-time traffic conditions around the driving section, and there is a suspended window of real-time map navigation in the lower right corner. In FIG. 3(b), the interface of selecting and clicking the video application B icon in the main interface is shown, which corresponds to the task step of finding the target application in the aforementioned steps; in this interface, application icons and widgets are mainly displayed, including the widgets of application A, application B, application C, application D, application E, video software A, video software B, video software C, application F, application G, and application H. In FIG. 3(c), the home page interface after entering the video application B is shown, in which the “movie name” is displayed on the home page to represent the recommended movies. The position of the recommended movies displayed on the home page also includes a sliding control, which supports the user to view the recommended movies and TV programs by sliding. Some recommended movies and TV programs are displayed in the middle of the video software B. In addition, some other movies and TV programs are displayed below. At the top of the interface, the “history” and “search” controls are displayed for entering the history page and the resource search page, respectively. Below these two controls, there is a column of different types of selection page controls, including selection, recommendation, home page, 4K channel, variety show, documentary, TV series, and animation. Each control represents a type of data information, for example, clicking the variety show control from the current home page interface will switch the current interface to a variety show related display interface, which supports the user to select the desired and favorite variety show by sliding up and down. FIG. 3(c) corresponds to the steps of starting the target application and opening the history playing record page, which shows the home page after starting the video software B application and the operation of clicking the history control. In FIG. 3(d), the interface of the history playing record in the video application B is shown, and the “my collection” and “purchase record” controls are displayed side by side with the playing record, which supports the user to select different controls based on different needs. Of course, in the current playing record interface, three types of information are displayed: movies, TV series, and animation. When no selection is made on these three types, all types of playing records will be displayed on the interface, as shown in the figure, including animation 1, animation 2, animation 3, movie 1, movie 2, movie 3, variety show 1, variety show 2, and variety show 3. After entering the page, clicking the last played TV series can enter the playing interface, that is, the interface in FIG. 3(e).
[0102] FIG. 3 shows all execution process interfaces of the automation task, but in the actual screen display, the actual screen display changes only from FIG. 3(a) to FIG. 3(e), and the interfaces from FIG. 3(b) to FIG. 3(d) are black box execution processes.
[0103] FIG. 4 shows the presentation interface of the association identifier of the key factor of the automation task of “continuing playing” in FIG. 2 and FIG. 3. The association identifier of the application icon of the video software B corresponding to FIG. 3(b) is presented in sequence, which is the graphic file contained in the application icon of the video software B; and the association identifier of the historical playing record control corresponding to FIG. 3(c) is presented, which is the graphic file contained in the historical playing record control; and the association identifier of the video thumbnail window of the recently played video corresponding to FIG. 3(d) is presented, which is the thumbnail of the video thumbnail window.
[0104] FIG. 4 is an example of the application of the embodiment to the intelligent cockpit scene. During driving, the user inputs a real-time voice instruction, the digital assistant receives the voice instruction for intent understanding and analysis, and then divides the automation task to obtain multiple task steps, each task step corresponding to only one execution action, and thus each task step also corresponds to a key factor. After determining the key factor, a corresponding association icon of each key factor will be displayed on the current interface (FIG. 4 is an autonomous driving interface), thus meeting the user's demand for perception and interaction of the entire automation task execution process. As can be seen from FIG. 4, in addition to the original interface information such as the real-time traffic situation around the vehicle and the surrounding building facilities during the current vehicle driving process, the association icons corresponding to the key factors are also displayed, which indicate the execution process of the automation task corresponding to the key factors.
[0105] In FIG. 4, the association identifier of the corresponding key factor is displayed, and the association identifier of each key factor also represents an execution action of an automation task. In addition, the arrangement of the association identifiers of the key factors is also arranged according to the execution order of the corresponding execution actions. In FIG. 4, three execution actions are required to complete the task, and thus the association icons of three key factors are displayed.
[0106] In some embodiments, in addition to displaying the association identifier of the key factor on the screen, an interface card of the process interface corresponding to the automation task step being executed can also be displayed. The interface card can be a thumbnail view of the process interface corresponding to the execution action, or a part of the interface elements in the process interface. Common scenarios for displaying interface cards include scenarios where the process interface is a train ticket / movie ticket / flight seat selection interface, and scenarios where the process interface is a navigation destination selection interface.
[0107] In some automatic task execution processes, although many schemes can be automatically determined by user preferences in many scenarios, there are some that are difficult to determine by preferences, at which time user manual intervention is needed, and therefore an interface card can be provided, and the user can determine by interacting with the interface card or by observing the interface card and using voice interaction to input, and after obtaining the input result, the automatic task continues to execute.
[0108] Referring to FIG. 5, FIG. 5 shows the interface of the display interface card in two classic scenarios of navigation destination selection and movie theater seat selection. In FIG. 5, (a) is a screen interface when the automatic task execution reaches the destination selection scenario, and (b) is a screen interface when the automatic task execution reaches the movie theater seat selection scenario.
[0109] In FIG. 5(a), the current execution action is to select a navigation destination based on the destination parsed from the user input. When the user has not been to the destination or has been there a few times, it is difficult to determine the preference according to the historical record. Moreover, if the destination is a large area, the distance between different navigation destinations can be very far, and if the destination is not selected correctly, it will bring a bad user experience. Therefore, user intervention is needed to help complete the execution action. Therefore, an interface card is provided, which displays the navigation destination selection page of the map application or part of the navigation destination selection page in the interface card, and can notify the user to intervene by voice prompt “Please select the destination you want to go to”. When the user input “for example, the destination marked as 1” is received, the automatic task continues to execute. After receiving the user input, the interface card can disappear. That is, when the execution action corresponding to the interface card is completed, the interface card is no longer displayed on the screen. In FIG. 5(a), the interface card is displayed above the associated identifier. In other examples, it can also be displayed below the associated identifier or in other positions on the screen.
[0110] Similarly, in FIG. 5(b), when the process interface of the automatic task is the movie theater seat selection interface, the corresponding execution action is to select a seat. The interactive element (such as a seat card) corresponding to the seat selection in the ticket purchase software interface is extracted and displayed on the screen in the form of a card, that is, an interface card is obtained and displayed on the current interface.
[0111] The real-time method described above is to display the seat selection indication as a key factor on the interface, extract the interactive element (such as a seat card) corresponding to the seat selection in the ticket purchase software interface, and display it on the display interface in the form of a card to obtain an interface card in the movie theater seat selection scenario.
[0112] The association identifier is displayed in a suspended manner on the upper layer of the current interface. In an implementation, referring to FIGS. 4 and 5, a suspended card or a suspended tray is displayed on the current interface, and the association identifier is displayed on the suspended card or the suspended tray, so as to avoid the content of the current interface affecting the user's viewing of the association identifier.
[0113] In some implementations, the conversation interface of the AI / digital assistant is displayed by interacting with the area on the screen where the key factor is displayed, and the conversation interface displays the task step information executed by the AI / digital assistant. For example, the association identifier can be displayed on the suspended card, and an upswipe gesture is performed near the display area of the suspended card, and then the conversation interface of the AI / digital assistant is displayed. Of course, other gesture operations or voice inputs and the like can also be used to display the conversation interface of the AI / digital assistant on the screen.
[0114] Referring to FIG. 6, FIG. 6 is a schematic diagram of a key factor displayed on a suspended card in the embodiment. The embodiment is mainly used to show that the key factor can be displayed by using the suspended card, and the suspended card is expanded into a conversation card by swiping at any position of the suspended card, the complete conversation flow is displayed in the conversation card, the task step description information corresponding to each execution step of the automated task of the digital assistant is displayed, and the corresponding key factor is displayed before the task step description information in the conversation interface, each key factor corresponds to an execution step, and multiple key factors are used to represent the task flow that has been executed or is being executed.
[0115] As shown in (a) of FIG. 6, the execution flow of the entire cinema seat selection is shown, which includes (from left to right in sequence): the digital assistant identifier (the first identifier on the left), the association identifier (the map identifier, the cinema identifier, and the seat to be selected identifier), wherein the digital assistant receives the voice input of the user, and based on semantic understanding and analysis of the voice of the user, the following flow is obtained (here, the flow is a simple example, and the specific flow is obtained according to the analysis of the digital assistant): first, locate the current position, call the map application to search for nearby cinemas, and find a cinema that meets the user's demand; then, enter the movie ticket purchase application to select a corresponding movie and a playing time period, and determine the movie session to be ordered; and then, enter the ticket purchase interface to select a seat. In (a) of FIG. 6, the three task steps each correspond to an association identifier of a key factor. From the perspective of the meta task, the first association identifier is the association identifier of the key factor of the meta task of calling the map application, and the second and third association identifiers are the association identifiers of the key factors of the meta task of calling the ticket purchase application. After performing an upswipe gesture on the suspended card in (a) of FIG. 6, the conversation interface in (b) of FIG. 6 is obtained.
[0116] (b) of FIG. 6 shows a dialogue process of a movie ticket booking in progress. In the interface, first, the user inputs the voice "Help me book a movie ticket for Avatar, IMAX hall", then the corresponding digital assistant displays the voice "Finding movie theaters near destination A 999 for you", the map icon displays the voice "Filtering IMAX theaters", and the candidate seat icon displays the voice "Found the nearest IMAX theater, and according to viewing habits, booked a seat for the 20:00 session". During the execution of the meta task, the digital assistant can play the voice. The task step information displayed on the interface can be the text of the voice broadcast. Through the dialogue interface, the task step information can explain the execution steps corresponding to each key factor.
[0117] In an implementation form of the present embodiment, the association identifier of the key factor supports interaction. Since the association identifier and the key factor have a mapping relationship, the key factor of the task execution step can be directly changed through interaction with the association identifier, and the key factor is replaced by an element of the same type in the process interface. First, the display interface should allow receiving the user's correction input. Based on the user's correction input, the target step is determined, and the updated key factor of the target step is determined based on the correction input. The key factor corresponding to the identifier of the target step is modified. The target step can be a step that has been executed or is being executed. In this way, the user can intervene in the automated task step that the digital assistant is executing or has executed through correction input, so that the automated task can be modified in real time according to the user's intention. These modifications include but are not limited to: changing the key factor of the step, adding one or more steps. For example, modifying the file to be sent, using another video software to continue playing the task, etc. In the present embodiment, the correction input refers to at least one of detecting a preset input acting on the key factor, detecting a voice input, or detecting a touch screen input.
[0118] In a possible implementation form of the present embodiment, the user can realize correction by directly interacting with the key factor. The user can click the key factor on the screen to adjust the execution step or execution action corresponding to the key factor. Of course, the preset input can be long press, swipe, and other common gestures in addition to click. After performing the preset input on the key factor, since the preset input is directly acting on the key factor of the step, the execution step or execution action corresponding to the key factor is the target step. After detecting the preset input, the electronic device displays other key factors of the same type in the process interface / application on the interface, so as to facilitate the user to replace other key factors to execute the task of the present step.
[0119] Referring to FIG. 7, FIG. 7 is a schematic diagram of an interface for user interaction with a key factor according to an embodiment. Taking the example of continuing to watch an animation in (a) of FIG. 7, the automation task automatically selects video software B to perform the task of continuing to watch the animation according to the user's preference when the automation task is executed. At this time, if the user clicks the first key factor, the process interface corresponding to the key factor is the desktop home screen, and thus the application icons of the same type are found in the desktop home screen (if there are multiple screens, the application icons in the multiple screens can be searched, that is, the data of the entire desktop application). There are three application icons with the label of "video software": video software A, video software B, and video software C. The icon of video software B is already displayed on the interface, and thus the electronic device displays the icons of video software A and video software C on the desktop to allow the user to select and replace. For example, after clicking, the key factors of the same type are displayed in a list. In addition, the application icons of the same type can be sorted, for example, according to the preference of the use time length.
[0120] In (b) of FIG. 7, an automation task of sending the files in group A to Zhang San is shown. By default, the latest file in the chat record is sent. In the step of searching for the file, the latest file "1215 meeting minutes.doc" is the key factor. When the user wants to replace it, the key factor is clicked, and then the same type of file in the chat record of group A is searched, that is, the file in the group chat record is searched, and the file is expanded and displayed on the screen in a list.
[0121] In one scheme, if no key factor of the same type is found, the process interface is displayed on the screen in the form of a floating window. For example, in the seat selection link when booking a movie theater, in the step of continuing to watch the last watched video, the history record / my collection is viewed, and the corresponding steps are executed. The corresponding key factor is unique in the process interface. At this time, there is no element of the same type in the process interface or in the application. The process interface is displayed on the screen in the form of a floating window. The process interface can be completely displayed or only partially displayed.
[0122] In one possible implementation, the user can also correct the step being executed or the step that has been executed by voice input during the execution of the automation task. After receiving the voice input during the execution of the automation task, semantic understanding is first performed, then the target step is determined based on the semantic understanding, and then the target step is re-executed according to the latest semantic understanding. Correspondingly, the subsequent steps are also re-executed. The corresponding key factor also needs to be re-determined, and the associated identifier of the new key factor is displayed.
[0123] Specifically, when the user makes changes through voice input, the user usually directly indicates the key factor to be replaced. For example, in the aforementioned example of continuing to play the last viewed video, the user can directly give voice input of "do not open video software B, open video software A" or "watch video software A", and the key factor is directly replaced from the icon of video software B to the icon of video software A based on semantic understanding. If the key factor to be replaced is not identified in semantic understanding, the processing scheme of the screen interaction input can be referred to, and other key factors of the same type in the process interface / application are displayed on the interface, or a floating window of the process interface is displayed. The change of the key factor usually causes the change of the subsequent step, and therefore, if the key factor of the completed step is changed, the automated task usually needs to be executed again from the step, and the key factor of the subsequent step is also replaced accordingly (for example, the history record icon of video software B is replaced by the history record icon of video software A). For the automated task being executed (not yet completed), if the correction input is detected, the automated task is interrupted.
[0124] In a possible implementation, the corresponding step has been completed before the user completes the modification, and therefore, for this short-range conflict, there are generally two scenarios, one of which is that the running speed of the automated task is very fast, causing the entire automated task to be completed before the user completes the modification, and the other of which is that the key factor modified by the user is the key factor of the last step of the last meta-task, and the automated task is also very likely to be completed before the user completes the modification. Therefore, there are two solutions for this short-range conflict:
[0125] ① Before jumping from the current interface to the result interface after the completion of the execution of the automated task, the jump can be interrupted, and the user can be asked whether to jump. For example, in the example of continuing to play the video of the first anime type in the history play record of video software A, the user is asked "whether to jump to video software A to play the video of the first anime type in the history play record of video software A". If the user wants to replace the video software, the change can be made in the manner of the above embodiment. If the video is directly played, "yes" can be directly replied.
[0126] ② The jump to the result interface is performed after the completion of the execution of the automated task, and after the jump is completed, the key factor is displayed for a preset time (for example, the key factor is still displayed for 10 seconds after the interface displays the result interface), so that the user can modify the completed automated task, and the modification can be made in the manner of the above embodiment.
[0127] The above embodiment is a flowchart of the division of the automated task by the digital assistant and the execution of the automated task in different scenarios. However, this embodiment covers the implementation of all automated tasks, which will not be listed one by one here.
[0128] The functions or features described in the above embodiments can be implemented by software, hardware, or a combination of software and hardware or firmware. When implemented by software, the functions or features can be implemented in the form of a computer program product.
[0129] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the computer program instructions produce, in whole or in part, the processes or functions according to the embodiments of the present application. The computer can be a general-purpose computer, a special-purpose computer, a computer language, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transferred from one website, computer, training device, or data center to another website, computer, training device, or data center through a wired (for example, coaxial cable, optical fiber, digital subscriber line) or wireless (for example, infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be stored by the computer or a data storage device such as a training device, a data center, etc. integrated with one or more available media sets. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
Claims
1. An interface display method in an automated task execution process, characterized by, Applied to an electronic device containing a digital assistant and a display interface, comprising: The electronic device receives a user instruction and generates an automated task; The electronic device executes at least one meta task, which is determined based on the automated task; At least one identifier is displayed on the display interface, which is used to represent the meta task and / or the execution step of the meta task.
2. The method of claim 1, wherein, Each meta task includes at least one key factor, and the identifier of the meta task is determined based on the at least one key factor, and the key factor is at least one of an application icon for calling an application program and a virtual control in an application interface.
3. The method of claim 2, wherein, The key factor further includes a generalized icon, and the generalized icon is a predefined icon.
4. The method of claim 1, wherein, During execution of the at least one meta task, the user interacts with the display interface through at least one of voice and touch screen to modify or replace the at least one meta task.
5. The method of claim 1, wherein, The method further comprises that the display interface further comprises a pop-up window or a card, and the pop-up window or the card is used to show the execution process of the corresponding meta task.
6. The method of claim 1, wherein, The at least one meta task is determined based on at least one of the number of called APPs and the number of page jumps.
7. The method according to any one of claims 2-6, characterized in that, The key factor is determined based on the execution step in the meta task, and each meta task includes at least one execution step, and the at least one execution step collectively completes all functions of the meta task.
8. The method of claim 5, wherein, The user interacts with the pop-up window or the card through at least one of voice and touch screen to specify or modify the content in the pop-up window or the card.
9. The method of claim 1, wherein, Further comprising: A preset operation for an identifier display area is detected, and a dialogue interface of a digital assistant is displayed, and the dialogue interface includes task description information of an automated task.
10. An interface display method in an automated task execution process, characterized by, Applied to an electronic device containing a display interface, comprising: The electronic device receives a user instruction and generates an automated task, and the automated task includes at least one execution action; The electronic device executes at least one execution action; At least one identifier is displayed on the display interface, which is used to represent the execution action that has been executed or is being executed.
11. The method of claim 10, wherein, The execution action at least includes calling an APP, page jump, clicking a control, inputting voice, and typing text.
12. The method of claim 10, wherein, During execution of the at least one execution action by the electronic device, the user interacts with the display interface through at least one of voice and touch screen to modify or replace the execution action.
13. The method of claim 10, wherein, The method further comprises that the display interface further comprises a pop-up window or a card, and the pop-up window or the card is used to show the execution process of the automated task.
14. An electronic device with a display interface, comprising: Comprising: A digital assistant is used to receive a user instruction and generate an automated task, and is further used to execute at least one meta task, which is determined based on the automated task; A display interface is used to display at least one identifier, which is used to represent the meta task and / or an execution step in the meta task.
15. The electronic device of claim 14, wherein, Each meta task includes at least one key factor, and the identifier of the meta task is determined based on the at least one key factor, and the key factor is at least one of an application icon for calling an application program and a virtual control in an application interface.
16. The electronic device of claim 15, wherein, The key factor further includes a generalized icon, and the generalized icon is a predefined icon.
17. The electronic device of claim 14, wherein, During execution of the at least one subtask, the user interacts with the display interface by at least one of voice and touch screen to modify or replace the at least one subtask.
18. The electronic device of claim 14, wherein, The method further includes that the display interface further includes a pop-up window or a card, and the pop-up window or the card is used to show the execution process of the corresponding subtask.
19. The electronic device of claim 14, wherein, The at least one subtask is determined based on at least one of the number of called APPs and the number of page jumps.
20. The electronic device of any of claims 14-19, wherein, The key factor is determined based on the execution steps in the subtask, and each subtask includes at least one execution step, and the at least one execution step collectively completes all functions of the subtask.
21. The electronic device of claim 18, wherein, The user interacts with the pop-up window or the card by at least one of voice and touch screen to specify or modify the content in the pop-up window or the card.
22. The electronic device of claim 14, wherein, Further include: A preset operation for an identification display area is detected, and a dialog interface of a digital assistant is displayed, and the dialog interface includes task description information of an automated task.
23. An electronic device with a display interface, comprising: Include: The electronic device receives a user instruction and generates an automated task, and the automated task includes at least one execution action; The electronic device executes the at least one execution action; At least one identification is displayed on the display interface, and the identification is used to represent an executed or executing execution action.
24. The electronic device of claim 23, wherein, The execution action at least includes calling an APP, page jump, clicking a control, inputting voice, and typing text.
25. The electronic device of claim 23, wherein, During execution of the at least one execution action, the user interacts with the display interface by at least one of voice and touch screen to modify or replace the execution action.
26. The electronic device of claim 23, wherein, Further include: The display interface further includes a pop-up window or a card, and the pop-up window or the card is used to show the execution process of the automated task.
27. An electronic device with a display interface, the electronic device comprising: Include a processor coupled with a memory, and the processor is used to store instructions, and when the instructions are executed by the processor, the electronic device executes the method in any one of claims 1-9 or 10-13.
28. A computer-readable storage medium, characterized in that, Include computer instructions, and when the computer instructions run on a computer system, the computer system implements the method in any one of claims 1-9 or 10-13.
29. A computer program product, comprising instructions therein, characterised in that, The instructions are executed to make the computer implement the method in any one of claims 1-9 or 10-13.
Citation Information
Patent Citations
Page intelligent response interaction system and method
CN109522083A
Conversation method and model based on interactive interface
CN114895999A
Interactive itinerary planning method and device, electronic equipment and storage medium
CN117349521A
GUI-oriented task automatic execution plug-in generation method and service acquisition method
CN117608552A
Human-in-the-loop voice automation system
WO2023234931A1