Task interaction method and device, equipment, storage medium and program product

By introducing a collaborative interaction mechanism into the intelligent system, users can view and adjust in real time during task execution, solving the problems of opacity and poor controllability in the existing technology, and improving the accuracy and user experience of task execution.

CN120409683APending Publication Date: 2025-08-01BEIJING ZITIAO NETWORK TECH CO LTD

Patent Information

Application Number
CN202510497452.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

When existing intelligent systems perform complex tasks, their interactions are opaque and poor controllability, making it difficult for users to understand and intervene in the task execution process in real time, resulting in inefficiency in tasks.

Method used

By introducing a collaborative interaction mechanism during task execution, users can view and adjust task execution information in the interface, and provide interface elements to support users' real-time intervention and control of task processes.

Benefits of technology

It enhances the transparency and controllability of task execution, improves the accuracy and user experience of task execution, and ensures that the task process meets user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120409683A_ABST
    Figure CN120409683A_ABST
Patent Text Reader

Abstract

According to the embodiment of the invention, a task interaction method and device, equipment, a storage medium and a program product are provided. The method comprises the steps that in response to a request for a task of an intelligent agent, first execution information used for executing a first action of the task is presented in an interface interacting with the intelligent agent, and the request comprises task information indicating the requirement of the task; in response to the execution of the first action, receiving an operation on the execution of the first action, the operation indicating an adjustment or confirmation of the execution of the first action, in which the task is suspended; and in response to receiving an operation on execution of the first action, restoring execution of the task by the agent based on the operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to methods, devices, electronic devices, computer-readable storage media, and computer program products for task interaction. Background Art

[0002] With the development of information technology, more and more intelligent systems can complete various complex tasks based on user input. Such systems usually interact with users in the form of conversations, generating response content according to users' natural language requests, thus forming question-and-answer conversations. However, in practical applications, the user's goal is often not just to obtain a single reply, but rather to hope that the system can complete a complete plan and execution around a certain task. Therefore, how to complete the execution of a specific task based on a task-goal-driven model and improve the structural degree and user controllability of the execution process has become an issue worthy of attention. Summary of the Invention

[0003] In a first aspect of the present disclosure, there is provided a method for task interaction. The method includes: in response to a request for a task of an agent, presenting first execution information for a first action for executing the task in an interface for interacting with the agent, where the request includes task information indicating the requirements of the task; in response to the execution of the first action, receiving an operation on the execution of the first action, the operation indicating an adjustment or confirmation of the execution of the first action, where the task is paused; and in response to receiving the operation on the execution of the first action, resuming the execution of the task by the agent based on the operation.

[0004] In a second aspect of the present disclosure, there is provided a device for task interaction. The device includes: an execution information presentation module configured to, in response to a request for a task of an agent, present first execution information for a first action for executing the task in an interface for interacting with the agent, where the request includes task information indicating the requirements of the task; an operation receiving module configured to, in response to the execution of the first action, receive an operation on the execution of the first action, the operation indicating an adjustment or confirmation of the execution of the first action, where the task is paused; and an execution module configured to, in response to receiving the operation on the execution of the first action, resume the execution of the task by the agent based on the operation.

[0005] In a third aspect of the present disclosure, there is provided an electronic device. The device includes at least one processor; and at least one memory, where the at least one memory is coupled to the at least one processor and stores instructions for execution by the at least one processor. The instructions, when executed by the at least one processing unit, cause the electronic device to execute the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. Computer instructions are stored on the medium, and when the computer instructions are executed by a processor, the method of the first aspect is implemented.

[0007] In a fifth aspect of the present disclosure, a computer program product is provided. The product includes a computer program, and when the computer program is executed by a processor, the method according to the first aspect of the present disclosure is implemented.

[0008] It should be understood that the content described in this part is not intended to define the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In combination with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0010] Figure 1 A schematic diagram showing an example environment in which the embodiments of the present disclosure can be implemented;

[0011] Figures 2A to 2G A schematic diagram showing an interface for interacting with an agent according to some embodiments of the present disclosure;

[0012] Figure 3 A flowchart showing a method for task interaction according to some embodiments of the present disclosure;

[0013] Figure 4 A schematic structural block diagram showing a device for task interaction according to some embodiments of the present disclosure; and

[0014] Figure 5 A block diagram showing an electronic device in which one or more embodiments of the present disclosure can be implemented. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0016] In the description of the embodiments of the present disclosure, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter.

[0017] In this document, unless otherwise specified, performing a step "in response to A" does not mean that the step is immediately performed after "A", but may include one or more intermediate steps.

[0018] It can be understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of data) should comply with the requirements of the corresponding laws, regulations and related provisions.

[0019] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, scope of use, usage scenarios, etc. of the information involved in the present disclosure should be informed to the relevant users in an appropriate manner and the authorization of the relevant users should be obtained in accordance with the relevant laws and regulations. Among them, the relevant users may include any type of right subject, such as individuals, enterprises, and groups.

[0020] For example, when receiving an active request from a user, a prompt message is sent to the relevant user to clearly prompt the relevant user that the operation requested by them will require obtaining and using the information of the relevant user, so that the relevant user can autonomously choose whether to provide information to software or hardware such as an electronic device, application program, server or storage medium that performs the operations of the technical solutions of the present disclosure according to the prompt message.

[0021] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the relevant user in response to receiving an active request from the relevant user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide information to the electronic device.

[0022] It can be understood that the above process of notifying and obtaining user authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0023] As used herein, the term "model" can learn the corresponding association between inputs and outputs from training data, so that after training, for a given input, the corresponding output can be generated. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this article, "model" can also be referred to as "machine learning model", "learning model", "machine learning network", or "learning network", and these terms are used interchangeably herein.

[0024] A "neural network" is a machine learning network based on deep learning. A neural network can process inputs and provide corresponding outputs, and it generally includes an input layer and an output layer and one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications usually include many hidden layers, thus increasing the depth of the network. The layers of the neural network are connected in sequence, so that the output of the previous layer is provided as the input of the next layer, where the input layer receives the input of the neural network, and the output of the output layer is the final output of the neural network. Each layer of the neural network includes one or more nodes (also called processing nodes or neurons), and each node processes the input from the previous layer.

[0025] Generally, machine learning can roughly include three stages, namely the training stage, the testing stage, and the application stage (also called the inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values are continuously iteratively updated until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association from input to output (also called the mapping from input to output) from the training data. The parameter values of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, so as to determine the performance of the model. In the application stage, the model can be used to process the actual input based on the parameter values obtained from training and determine the corresponding output.

[0026] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. In this example environment 100, an application 125 is installed in the terminal device 110. The user 140 can interact with the application 125 via the terminal device 110 and / or the attached devices of the terminal device 110.

[0027] In some embodiments, the application 125 can be downloaded and installed in the terminal device 110. In some embodiments, the application 125 can also be accessed in other ways, such as by web access, etc. In Figure 1In the environment 100, in response to the application 125 being launched, the terminal device 110 can present the interface 150 of the application 125.

[0028] In some embodiments, the terminal device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the terminal device 110 can also support any type of user interface (such as a "wearable" circuit, etc.).

[0029] In some embodiments, the terminal device 110 communicates with the server 130 to implement the service provision of the application 125. The terminal device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, television receivers, radio broadcast receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the terminal device 110 can also support any type of user interface (such as a "wearable" circuit, etc.). The application 125 can be various types of computing systems / servers capable of providing computing capabilities, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, and the like.

[0030] In the embodiments of the present disclosure, the application 125 can provide an interaction function with the agent. The application 125 can be an application dedicated to providing agent services, or an application integrated with an agent (that is, it can provide other functions or services in addition to the agent). Although Figure 1 a single application is shown, in fact, multiple applications can be installed on the terminal device 110.

[0031] In the present disclosure, the agents 160-1, 160-2,..., 160-N (collectively or individually referred to as the agent 160) can be deployed locally on the terminal device 110 or remotely deployed. In the case of remote deployment, the terminal device 110 can directly call the agent, or can call the agent via the server 130.

[0032] In an embodiment of the present disclosure, the agent 160 may have intelligent dialogue and task processing capabilities. The terminal device 110 provides an interface 150 that can present the interaction with the agent 160. In the interface 150, the user 140 can input a task request for the agent 160 by inputting natural language (text input or voice input), and can also upload an input online or offline file dialogue to instruct the agent 160 to assist in completing various tasks.

[0033] In an embodiment of the present disclosure, during the interaction with the user 140, the agent 160 can respond to the request of the user 140 and process the task indicated by the user. In some embodiments, during the task processing, the agent 160 can call one or more tools 165-1, 165-2,... 165-M (collectively or individually referred to as tools 165) according to the task needs to assist in the execution of the task and the provision of the task result. These tools 165 can be any type of tool, such as a weather query tool, a flight query tool, an information search tool, an online or offline database, an image processing tool, a chart generation tool, a web page production tool, and so on.

[0034] In some embodiments, the environment 100 may further include a management node for a plurality of agents 160-1, 160-2,... 160-N, and the management node can interact with the plurality of agents 160-1, 160-2,... 160-N. In some examples, the management node can determine the task requirements corresponding to the task request in response to the task request of the user 140. Thereafter, the management node can assign the task request to the agent 160 that matches the task requirements based on the task requirements to request the agent 160 to execute the task. In other examples, the management node can also determine the execution plan of the task based on the task requirements. The execution plan of the task can indicate one or more subtasks required to complete the task. The management node can assign the one or more subtasks to one or more agents 160, and the one or more agents 160 can execute their respective corresponding subtasks. Regarding this management node, in some examples, the management node can be implemented by one of the plurality of agents 160-1, 160-2,... 160-N (in this case, this agent 160 is also referred to as the management agent 160). In other examples, the management node can be implemented by a machine learning model, for example, can be implemented by a language model (LM) or a large language model (LLM).

[0035] In some embodiments, the agent 160 can be built based on one or more machine learning models. In some embodiments, the machine learning model based on the agent 160 can include at least a language model (LM). These machine learning models include content generation models that can generate corresponding outputs based on model inputs. In some embodiments, a machine learning model based on a language model can receive model inputs in textual modalities (e.g., natural language and / or machine language) and / or model inputs in non-textual modalities (e.g., images, voice, video, etc.), and can obtain corresponding model outputs based on the model inputs and prompt words, thereby completing the execution of the task. The prompt words here are used to guide the machine learning model to generate user queries that can solve the model inputs. In an application scenario for supporting user dialogue, the input of the user 140 can be provided to the machine learning model 160 as at least a part of the model input (the other part may include prompt words). The user input is regarded as a question or query request. Based on the model output, a corresponding response can be provided to the user 140.

[0036] It should be understood that the structure and function of the various elements in the environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present disclosure.

[0037] As briefly mentioned above, with the development of information technology, natural language-driven intelligent systems have been widely used in scenarios such as content generation, information query, and logical reasoning. Conventional intelligent systems often employ a "conversational" approach to human-computer interaction. Users input requests in natural language, and the system instantly generates responses, completing several rounds of dialogue. This approach is suitable for conversational interaction tasks such as question-answering, translation, summarization, and information query.

[0038] However, in real-world scenarios, users often don't simply want to obtain a single piece of information; rather, they want the system to complete multi-step operations around a more complex goal, such as itinerary planning, project report generation, research and analysis, and decision support. To meet these requirements, the system must not only understand the user's needs but also possess comprehensive capabilities such as task organization, external call processing, and result processing. Therefore, building a "task-centric" intelligent interactive system that supports multi-step execution and user collaborative control has become a key direction in the development of intelligent systems.

[0039] Furthermore, while some intelligent systems have attempted to expand the functionality of language models through mechanisms such as plugin systems or tool invocations to achieve more complex goals, their interaction model still primarily relies on "automatic response." Specifically, after a user enters a task objective via natural language, the intelligent model is triggered to process it, automatically executing a series of operations and providing the user with a single feedback loop of the final results.

[0040] This type of approach has a certain degree of automation, but its core feature remains closed-loop execution, that is, from receiving instructions to outputting results. The entire process is autonomously completed by the intelligent system, making it difficult for users to understand the task execution logic in real time and impossible to intervene or adjust in the intermediate steps. This approach may have problems of opaque interaction and poor controllability when dealing with complex, dynamic, or highly accurate tasks. Moreover, during the interaction process, users often wait for a period of time to view the output result, but then find that the current result is not satisfactory, so they need to initiate a new round of dialogue again. Repeating this way will seriously affect the task efficiency.

[0041] In view of this, according to an embodiment of the present disclosure, an improved solution for task interaction is provided. According to this solution, in response to a request for a task of an agent, first execution information for executing a first action for the task is presented in an interface for interacting with the agent, where the request includes task information indicating the requirements of the task. Then, in response to the execution of the first action, an operation on the execution of the first action can be received, and the operation indicates an adjustment or confirmation of the execution of the first action, where the task is paused. Further, in response to receiving the operation on the execution of the first action, the execution of the task by the agent is resumed based on this operation.

[0042] Thus, by introducing a collaborative interaction mechanism during the task execution process, users can view the execution information in real time during multiple actions of the task execution and confirm or adjust through user interaction, thereby achieving effective control and flexible intervention in the task execution process. In this way, the problems of opacity and uncontrollability brought by the "black box" task execution process are avoided, making the task execution process more accurate and more in line with user needs.

[0043] In the present disclosure, "trigger" refers to one or more interaction operations of a user on a terminal device. Further, these interaction operations can be triggered within the same user interface / popup window or within different user interfaces / popup windows. The present disclosure is not limited in this regard.

[0044] For ease of understanding, the following will refer to Figure 2A and Figure 2G to describe examples of interfaces for interacting with an agent in some embodiments of the present disclosure. These example interfaces can be presented by an application 125 in the terminal device 110. It should be understood that the interfaces shown in the drawings are only examples, and various interface designs can actually exist. Each graphical element in the interface can have different arrangements and different visual representations, one or more of which can be omitted or replaced, and there can also be one or more other elements. The embodiments of the present disclosure are not limited in this regard.

[0045] In Figure 2AThe interface 200A shown can present an interaction interface 200A between the user 140 and the agent 160. In this interaction interface 200A, a request for a new task initiated by the user can be received. As Figure 2A shown, the user can trigger a new task by triggering the "New Task" control 211. In response to the triggering of the "New Task" control 211, a task-specific interaction interface is presented in the interface 200A, such as Figure 2B the interaction interface 200B shown, to initiate a new task in the interaction interface 200B. In some embodiments, an input area 215 for task input is also provided in the interface 200A. In the input area 215, the user can input a request for the task to be triggered and initiate the task request through the "Send" control 216. The input area 215 can support text input, such as entering text in a text input box, and can also support voice input, such as by triggering the voice control 217. In addition, the input area 215 can also provide an upload control 216 to support uploading an attachment to indicate the user's task request. In some embodiments, the interaction interface 200A also provides a task list 212 initiated by the agent 160. The task list 212 indicates each initiated task and the execution status of the task (e.g., interrupted status, in-progress status, completed status, etc.).

[0046] In some embodiments, multiple operation modes of the user 140 and the agent 160 can be provided, and flexible switching can be performed between the multiple operation modes, such as via Figure 2A the mode switching control 218 shown. When a certain mode is triggered, a corresponding interaction area is presented to facilitate the interaction between the user 140 and the agent 160. The interaction methods between the user 140 and the agent 160 are different in different operation modes, so that the interaction requirements in different application scenarios can be flexibly adapted.

[0047] After triggering a request for a task of the agent, the terminal device 110 presents Figure 2BThe interactive interface 200B shown. In the interactive interface 200B, requests for tasks initiated by the user can be presented, such as message 221, and task execution information during the execution of tasks by the agent can also be presented. The task execution information can be presented in the interactive interface 200B in the form of one or more messages. During the execution of the task, the agent 160 can analyze the task request initiated by the user to determine the task requirements, thereby generating an execution plan for the task. The execution plan of the task can indicate one or more subtasks required to complete the task. To enable the user to more clearly understand the execution process of the task by the agent, a message 221 indicating the generation of the execution plan can be presented in the interactive interface 200B. After the execution plan is determined, messages 223 indicating each subtask in the execution plan can be presented in the interactive interface 200B. Next, the agent 160 will automatically or in response to the user's confirmation continue to execute each subtask, and the execution process and execution results of each subtask can be presented in the interactive interface 200B. After the execution of the entire task is completed, the execution result of the task or an access entry to the execution result (if the execution result needs to be presented in another interface) can be presented in the interactive interface 200B. In this way, during the execution of the entire task, the user can intuitively understand how the agent specifically decomposes the user's task requirements and the specific execution steps of each subtask, so as to facilitate the user to confirm or adjust the execution process of the task or subtask, and thus obtain the desired execution result.

[0048] In some embodiments, the process of determining one or more subtasks required to complete the task in the execution plan of the task based on the task request of the user 140 can be implemented by a machine learning model deployed at the server 130 or called by the server 130. Further, after determining the specific task execution plan, the machine learning model can assign each subtask to one or more suitable agents 160 for specific task execution.

[0049] Alternatively or additionally, the user's task request can also be sent to the management agent 160. The process of determining one or more subtasks in the execution plan of the task can be implemented by the management agent 160. The management agent 160 can be used to parse the user's task request based on the task request of the user 140, thereby determining one or more subtasks required to complete the task in the execution plan of the task. Further, the management agent 160 can assign the subtasks to one or more suitable agents 160 for specific task execution according to different task execution requirements and capabilities.

[0050] In some embodiments, during the process of task execution, the agent 160 can be in a paused state. Options such as modification and resume can be provided during the paused state so that the user can check whether the execution of the current action meets expectations and whether the execution of the action needs to be modified. Then, in the case of modifying the task, the modified task can be passed to the management agent 160. The management agent 160 can allocate the modification of the task to other agents. During the task execution, the user 140 can add a new task. For example, add a sub-task XX. In this scenario, the terminal device 110 can call the management agent 160 to add the sub-task XX. After receiving the sub-task XX, the management agent 160 can modify the entire task. Subsequently, the agent 160 can re-allocate.

[0051] In some embodiments, the management agent 160 can determine a prompt word corresponding to additional task information indicating a modification or addition to a sub-task. This prompt word can be concatenated as a part of the system prompt (System Prompt, SP) to the original system prompt, so as to provide context guidance for subsequent task execution. Alternatively or additionally, the prompt word corresponding to the additional task information can be directly concatenated into the task prompt word of the agent corresponding to the sub-task, so that the corresponding agent can take into account the new context when executing the task.

[0052] Some example embodiments of the present disclosure will continue to be described below with reference to the accompanying drawings. The embodiments related to the present disclosure can be implemented on the terminal device 110. It should be noted that the operations performed by the terminal device 110 can specifically be performed by related applications (such as application 120) installed on the terminal device 110. Some operations described with reference to the terminal device 110 may require the assistance of the server 130 to complete.

[0053] In some embodiments, in response to a request for an agent's task, the terminal device 110 presents first execution information for a first action for performing the task in an interface for interacting with the agent. The request for the agent's task includes task information indicating the requirements of the task. For example, the request includes information related to user requirements such as task objectives, constraints, output expectations, etc. The task information serves as the key semantic basis for task driving, guiding the subsequent task decomposition and execution decision-making of the agent.

[0054] As an example, refer to Figure 2A, in the input area 215, the user can input a request for a task to be triggered and initiate the request for the task by sending the control 216, thus serving as the starting point for triggering the task execution process of the intelligent agent 160. The request may include a natural language request for the task objective entered by the user in the input area 215 (such as "Help me organize a research report on XX"). Alternatively or additionally, the request may also include multimodal forms such as voice, files, etc. For example, the user can provide the task request by triggering the voice control 217 or the upload control 216.

[0055] As another example, referring to Figure 2B , after triggering the request for the task of the intelligent agent 160, the terminal device 110 presents Figure 2B the interactive interface 200B shown in the figure. When the terminal device 110 receives the request for the task of the intelligent agent 160 initiated by the user, it can perform semantic parsing on the request to identify the task requirements and constraint information contained therein. Based on the parsing result, the terminal device 110 can determine and execute the first action corresponding to the task, and present the first execution information associated with the first action (such as the running status or intermediate result of the first action) in the interface 200B for the user to visually view and understand. Exemplarily, the process described above in which the intelligent agent 160 analyzes the task request initiated by the user to determine the task requirements and thus generates an execution plan for the task can be regarded as an example process of performing the first action.

[0056] In some embodiments, the execution of a task includes multiple actions. A task is the overall goal or requirement initiated by the user, such as "Help me make a travel plan". An action refers to a phased operation performed to complete the user's task, which can be an operation unit that the intelligent agent 160 performs during the task execution process and has a clear purpose and observable execution result. Each action can constitute a stage in the task process and has a phased execution result.

[0057] In some embodiments, the multiple actions included in the execution of a task may include an action indicating the generation of an execution plan for the task, and the execution plan may indicate at least one subtask of the task. For example, the process described above in which the intelligent agent 160 analyzes the task request initiated by the user to generate an execution plan for the task can be regarded as one of the multiple actions.

[0058] In some embodiments, the multiple actions included in the execution of a task may also include at least one subtask indicated by the execution plan. The execution of each subtask can be represented by an action, and the intelligent agent 160 can gradually execute these subtasks based on the determined execution plan, thereby gradually completing the execution of the entire task.

[0059] In some embodiments, in order to achieve human-machine collaboration and dynamic control during task execution and enhance the user's controllability over the task execution process of the agent, when the execution of the first action of the agent 160 is completed or it is detected during the execution of the first action that the action process requires user participation, the intermediate interaction stage can be entered. In this stage, the terminal device 110 can cache the execution result of the first action and pause the automatic advancement of subsequent actions in the task process. Further, if the task process requires user participation, the terminal device 110 can receive an operation for the execution of the first action, and this operation can indicate an adjustment to the execution of the first action or a confirmation of the execution of the first action. The operation can be implemented in various forms. For example, a first interface element and a second interface element are presented in the interface to guide the user to participate in task control.

[0060] In some embodiments, the first interface element can indicate an adjustment to the execution of the current action (i.e., the first action), such as modifying, replacing, or regenerating the execution content. The second interface element can indicate a confirmation of the execution of the first action, that is, it indicates recognition of the execution result of the agent 160 in the interaction stage of the current action and can be used as a basis for continuing to advance the task execution process.

[0061] In some embodiments, after the user's operation is completed, the terminal device 110 can resume the agent's execution process of the task based on the user's operation. For example, the terminal device 110 can resume the execution of the task by the agent 160 based on the user's trigger of the first interface element or the second interface element. If it is detected that the first interface element is triggered, the adjustment interaction process is entered. In this process, the terminal device 110 can guide the user to supplement, modify, or re-plan the currently executed action. After the adjustment is completed, the terminal device 110 can resume the execution process of the subsequent task. This processing mode ensures that the user can intervene and reconstruct key task nodes to achieve flexible control and precise guidance. If it is detected that the second interface element is triggered, the terminal device 110 can determine that the user has recognized the execution result or intermediate state of the current action. At this time, the terminal device 110 can instruct the agent 160 to continue to execute the next action of the task, that is, there is no need to interrupt or modify the current task process, and only the task execution flow needs to be advanced.

[0062] As an example, Figure 2C An example of the interaction interface 200C after the execution of the first action is shown, where the first action can indicate the generation of an execution plan for the task. Refer to Figure 2C, after the user triggers a task request for the agent 160, the terminal device 110 can, based on the reasoning ability of the agent 160, perform semantic understanding and requirement recognition on the natural language task. Further, the terminal device 110 can generate a structured execution plan based on the user's requirements. The process of generating this execution plan can be understood as the execution process of the first action. During the execution process of the first action, the agent 160 can analyze the task request initiated by the user to determine the task requirements, thereby generating an execution plan for the task. The execution plan of the task can indicate one or more subtasks required to complete the task.

[0063] Continue to refer to Figure 2C , the terminal device 110 can present a status information area 222 for executing the first action in the interaction interface 200C to prompt the specific action being currently executed and the execution status. Each action can include several steps, thus constituting the basic operation process of the action. For example, for the action of "information collection", it can include multiple steps such as "invoking the search engine interface", "retrieving with different search terms", and "sorting out materials". The execution plan of the task determined by executing the first action, as well as the specific steps included in each subtask of the execution plan, can be displayed through the message 223.

[0064] Continue to refer to Figure 2C , the terminal device 110 can also present a first interface element 231 (such as a "Modify Task" button) in the interaction interface 200C. When the first action is to disassemble the user's requirements to generate a specific task plan, the execution result of the first action can be a task execution plan including at least one subtask. To support the user's real-time intervention in the task execution plan, the first interface element can be used to trigger an operation to adjust one or more subtasks generated by executing the first action.

[0065] In some embodiments, in response to the triggering of the first interface element 231, the terminal device 110 can present an input control in the interface. The input control can be used to guide the user to input additional task information, such as supplementing constraint conditions, explaining the expected structure, adding missing subtasks, etc. Exemplarily, the input control can include a text input box, a task parameter editing area, a task priority setting control, etc.

[0066] Further, in response to receiving additional task information via the input control, the terminal device 110 may control the agent 160 to re - execute the first action. Then, the terminal device 110 may present the updated execution information of the first action in the interface, where the updated first execution information is generated by the agent 160 through re - executing the first action based on the additional task information. In other words, the agent 160 can, on the basis of the cached execution information, combine the additional task information provided by the user, re - analyze and disassemble the task, so as to generate an adjusted task execution plan. The updated task execution plan will be presented in the interactive interface for the user to further confirm or adjust again.

[0067] Figure 2D An example of the interactive interface 200D after triggering the first interface element is shown, as Figure 2D shown, after the user clicks the modify task button based on the need to adjust the subtasks, the terminal device 110 may present the input control 241 (such as an input box) in the interactive interface 200D. The user can, through the input control 241, input supplementary instructions or adjustment suggestions for the current task execution plan, for example, "Please add another task A", etc.

[0068] Continue to refer to Figure 2D , the terminal device 110 may also present a cancel control 242 and a confirmation control 243 in the interactive interface 200D. The cancel control 242 is used to revoke the modification process carried by the current input control 241 and return to the presentation interface of the original task plan when the user abandons the modification operation. The confirmation control 243 is used to confirm the submission of the modified content after the user completes the input.

[0069] If the user clicks the confirmation control 243, the terminal device 110 will utilize the agent 160 to re - execute the task disassembling logic based on the additional input information of the user, that is, re - execute the first action. The agent 160 accordingly generates an updated task plan. The terminal device 110 will then present the updated first execution information in the interactive interface 200D, showing new subtask lists, task structures or dependency relationships, etc. in a structured form for the user to continue reviewing and confirming.

[0070] In some embodiments, when receiving additional task information provided by the user, the management agent 160 may, based on this, update or adjust the existing task execution plan. On the basis of the updated task flow, the management agent 160 further decomposes the task into several subtasks and assigns them to appropriate agents for execution respectively to achieve the efficient collaborative processing of the overall task.

[0071] Return to refer to Figure 2C, the terminal device 110 can also present a second interface element 232 (such as a "start task" button) in the interaction interface 200C. If it is detected in the interaction interface 200C that the user triggers the second interface element 232, the terminal device 110 can instruct the agent 160 to continue to perform the second action of the task. Further, the terminal device 110 can present the second execution information of the second action in the interface 200C.

[0072] The second action can be one of at least one subtask determined in the first action (task plan generation). Further, during the execution of the second action, the terminal device can present the second execution information in the user interface, and the second execution information is used to display the content or intermediate result generated by the agent 160 when performing the current subtask, so as to realize the visual presentation and transparent interaction of the task execution process. In this way, the processing process of each subtask can be observed and intervened by the user, so as to enhance the accuracy, controllability and interaction experience of the task process.

[0073] As another example, Figure 2E shows an example of the interaction interface 200E after performing the first action, where the first action can instruct the execution of the first subtask in the task plan. Refer to Figure 2E , the agent 160 executes each subtask in the task plan according to the preset task order. For example, the subtasks can include operation contents such as information collection, content writing, and data analysis.

[0074] During the execution of the subtask, the agent 160 can combine its reasoning ability and tool ability to process the subtask step by step. Each subtask can include several steps, thus constituting the basic operation process of the subtask. For example, the first subtask can include steps such as accessing a search engine, retrieving data sources related to keywords, reading web page content, and organizing and caching the obtained materials.

[0075] Continue to refer to Figure 2E , the terminal device 110 can present an execution information area 251 for executing the first subtask in the interaction interface 200E to prompt the specific execution steps of the current subtask. In this way, the visibility of the user to the task execution process can be improved. Alternatively or additionally, after the subtask is completed, the terminal device 110 can display the execution result of the first subtask in a structured form in the preview area 255 on the right side of the interface. In this area, the user can view the execution of the current subtask (such as multiple web page links, text summaries, picture materials, etc. collected), so as to evaluate the execution effect of the subtask.

[0076] The terminal device 110 can also present a first interface element 252 (such as a "Modify" button) and a second interface element 253 (such as a "Continue" button) in the interaction interface 200E. The first interface element 252 can be used to guide the user to adjust the execution of the current subtask, such as modifying the subtask target, resetting parameters, supplementing operation conditions, etc. The second interface element is used to confirm the validity of the execution result of the current subtask and instruct the agent 160 to continue to promote the execution of the subsequent subtasks.

[0077] Figure 2F An example of the interaction interface 200F after triggering the first interface element is shown, as Figure 2F shown. After the user clicks the modify task button, the terminal device 110 can present an input control 261 (such as an input box) in the interaction interface 200F. The user can input opinions on adjusting the execution method and execution result of the current subtask through this input control, such as "Please change to display in a table", "Add data source X", etc.

[0078] Continue to refer to Figure 2F and the terminal device 110 can also present a cancel control 262 and a confirmation control 263 in the interaction interface 200D. The cancel control 262 is used to revoke the modification process carried by the current input control 261 and return to the original presentation interface when the user abandons the modification operation. The confirmation control 263 is used to confirm the submission of the modified content after the user completes the input.

[0079] Furthermore, if the user triggers the confirmation control 263, the agent 160 can re-execute the first action based on the cached execution result and the additional task information provided by the user, that is, re-execute the first subtask. For example, the result of the re-execution (i.e., the updated first execution information) may include an updated link list, more demand-compliant material content, accurately matched analysis materials, etc.

[0080] Return to refer to Figure 2E and the terminal device 110 can also present a second interface element 253 (such as a "Continue" button) in the interaction interface 200E. If it is detected in the interaction interface 200E that the user triggers the second interface element 253, it indicates that the user has confirmed and recognized the execution result of the first subtask and no further adjustment is required. The agent 160 can continue to execute the second action of the task, that is, the next subtask arranged after the first subtask in the task plan. For example, if the first subtask is information collection, the second subtask may be preliminary content analysis or extraction of key information. Furthermore, the terminal device 110 can present the second execution information of the second action in the interface 200E.

[0081] In the embodiments of the present disclosure, by introducing an intermediate interaction stage during the task execution process and setting interface elements for adjusting the execution content and for confirming the execution result, the user can, after generating a task execution plan or after any subtask is completed, achieve real-time intervention and dynamic control of the task process based on the execution information presented on the interface.

[0082] Return reference Figure 2A To improve the controllability of the behavior mode of the intelligent agent by the user, the terminal device 110 may receive a selection of a mode applicable to the intelligent agent, and this mode may indicate an execution strategy or an interaction strategy adopted by the intelligent agent during the task execution process. For example, the terminal device 110 may present, in the initial stage after creating a new task, setting options related to the task execution mode in the interface 200A, such as the setting control 218.

[0083] In some embodiments, in response to the triggering of the setting control 218, the terminal device 110 may present a list of modes applicable to the intelligent agent, and the mode list at least includes a first mode and a second mode. The first mode indicates that the task is executed based on the triggering of a first interface element or a second interface element. In response to receiving the selection of the first mode of the intelligent agent, the intelligent agent executes the task based on the first mode. When the first mode of the intelligent agent is selected, in response to the completion of the execution of the first action, the first interface element and the second interface element may be presented in the interface (such as the interaction interface 200C or the interaction interface 200E) of the terminal device 110. The second mode indicates sequential automatic execution of multiple actions included in the execution of the task.

[0084] As an example, reference Figure 2G After the user triggers the setting control 218 in the interface, the terminal device 110 may present a list of modes 271 applicable to the intelligent agent in the interface 200G, which is used to indicate the optional execution modes for the current task execution. The mode list 271 at least includes two modes: the first mode (also called the collaborative mode) and the second mode (also called the automatic mode). The user can, when creating a task, select, according to actual needs, the way in which the task is to be executed. In the first mode, during the task execution process of the intelligent agent 160, for the execution result of each action, it is necessary to go through the operation of the user before the subsequent task execution can be advanced. The task advancement process in this mode has high controllability and interactivity, and is suitable for scenarios where the user needs to participate in key decisions or has high requirements for execution details. In the second mode, the intelligent agent 160 can automatically execute multiple actions included in the task in a sequential manner according to the established task process without waiting for user confirmation or intervention. This mode is suitable for scenarios where the user already has sufficient trust in the execution process or expects to complete the task quickly.

[0085] In some embodiments, the user may not explicitly select any mode at the initial stage of task creation. When the user does not select a specific mode, the agent 160 can automatically determine the execution mode of the task based on the default configuration or the user's needs. For example, by default, the agent 160 can execute the task based on the automatic mode.

[0086] In summary, according to the embodiments of the present disclosure, through the execution mechanism of agent task-driven and user interaction intervention, the efficient cooperation between the user and the agent in the whole process of the task is realized, thus solving the problems of closed task execution process and low user participation. In this way, the user can dynamically intervene in the processes of task plan generation and subtask execution, thereby enhancing the usability of the agent in complex task scenarios and improving the overall user experience.

[0087] Figure 3 The flowchart of a method 300 for task interaction according to some embodiments of the present disclosure is shown. The method 300 can be implemented at the terminal device 110. The following refers to Figure 1 Describe the method 300.

[0088] In block 310, the terminal device 110 presents first execution information for a first action for performing a task in an interface for interacting with the agent in response to a request for a task of the agent. The request includes task information indicating the requirements of the task.

[0089] In block 320, the terminal device 110 receives an operation on the execution of the first action in response to the execution of the first action, the operation indicating an adjustment or confirmation of the execution of the first action, where the task is paused.

[0090] In block 330, the terminal device 110 resumes the execution of the task by the agent based on the operation in response to receiving the operation on the execution of the first action.

[0091] In some embodiments, resuming the execution of the task includes: in response to receiving an operation indicating an adjustment of the execution of the first action, receiving additional task information for adjusting the execution of the first action; in response to receiving the additional task information, re-executing the first action by the agent; and presenting updated execution information of the first action in the interface, where the updated first execution information is generated by the agent re-executing the first action based on the additional task information.

[0092] In some embodiments, resuming the execution of the task further includes: in response to receiving an operation indicating confirmation of the execution of the first action, continuing the execution of a second action of the task by the agent; and presenting second execution information of the second action in the interface.

[0093] In some embodiments, method 300 further includes: receiving a selection of a mode applicable to the agent, the mode including at least a first mode and a second mode, wherein the first mode indicates performing a task based on a received operation; and in response to receiving a selection of the first mode of the agent, performing, by the agent, the task based on the first mode.

[0094] In some embodiments, method 300 further includes, when the first mode of the agent is selected, presenting, in an interface, a first interface element and a second interface element in response to completion of execution of a first action, the first interface element being for receiving an operation indicating adjustment of the execution of the first action, and the second interface element being for receiving an operation indicating confirmation of the execution of the first action.

[0095] In some embodiments, the second mode indicates performing a plurality of actions included in the sequential automatic execution of a task.

[0096] In some embodiments, the execution of a task includes a plurality of actions, which respectively indicate: generating an execution plan for the task, the execution plan indicating at least one subtask of the task, and the at least one subtask.

[0097] In some embodiments, the first action indicates generating an execution plan for the task, and wherein, in response to execution of the first action, receiving an operation for the execution of the first action includes: presenting the execution plan in an interface in response to completion of generating the execution plan for the task; and receiving an operation for the execution of the first action, the operation indicating adjustment or confirmation of the execution plan.

[0098] In some embodiments, the first action indicates a first subtask among at least one subtask, and wherein, in response to execution of the first action, receiving an operation of the user for the execution of the first action includes: presenting an execution result of the first subtask in an interface in response to completion of execution of the first subtask; and receiving an operation for the execution of the first action, the operation indicating adjustment or confirmation of the execution result of the first subtask.

[0099] Embodiments of the present disclosure also provide corresponding apparatuses for implementing the above methods or processes. Figure 4 A schematic structural block diagram of an apparatus 400 for task interaction according to some embodiments of the present disclosure is shown. The apparatus 400 may be implemented in or included in a terminal device 110, for example. Each module / component in the apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.

[0100] As Figure 4As shown, the device 400 includes an execution information presentation module 410 configured to present, in an interface for interacting with an agent, first execution information for a first action to perform a task in response to a request for a task of the agent, where the request includes task information indicating the requirements of the task; an operation receiving module 420 configured to receive an operation for the execution of the first action in response to the execution of the first action, the operation indicating an adjustment or confirmation of the execution of the first action, where the task is paused; and an execution module 430 configured to resume the agent's execution of the task based on the operation in response to receiving the operation for the execution of the first action.

[0101] In some embodiments, the execution module 430 is further configured to receive additional task information for adjusting the execution of the first action in response to receiving an operation indicating an adjustment of the execution of the first action; re - execute the first action by the agent in response to receiving the additional task information; and present updated execution information of the first action in the interface, where the updated first execution information is generated by the agent re - executing the first action based on the additional task information.

[0102] In some embodiments, the execution module 430 is also further configured to continue the agent's execution of a second action of the task in response to receiving an operation indicating confirmation of the execution of the first action; and present second execution information of the second action in the interface.

[0103] In some embodiments, the device 400 further includes a mode selection module configured to receive a selection of a mode applicable to the agent, the mode including at least a first mode and a second mode, where the first mode indicates performing a task based on the received operation; and in response to receiving a selection of the first mode of the agent, perform the task by the agent based on the first mode.

[0104] In some embodiments, the mode selection module is further configured to, in the case where the first mode of the agent is selected, present a first interface element and a second interface element in the interface in response to the completion of the execution of the first action, the first interface element for receiving an operation indicating an adjustment of the execution of the first action, and the second interface element for receiving an operation indicating confirmation of the execution of the first action.

[0105] In some embodiments, the second mode indicates sequentially and automatically performing multiple actions included in the execution of a task.

[0106] In some embodiments, the execution of a task includes multiple actions, which respectively indicate: generating an execution plan for the task, the execution plan indicating at least one subtask of the task, and at least one subtask.

[0107] In some embodiments, the first action indicates the generation of an execution plan for a task, and the operation receiving module 420 is further configured to present the execution plan in the interface in response to the completion of the generation of the execution plan for the task; and receive an operation for the execution of the first action, where the operation indicates the adjustment or confirmation of the execution plan.

[0108] In some embodiments, the first action indicates the first subtask among at least one subtask, and the operation receiving module 420 is further configured to present the execution result of the first subtask in the interface in response to the completion of the execution of the first subtask; and receive an operation for the execution of the first action, where the operation indicates the adjustment or confirmation of the execution result of the first subtask.

[0109] The units and / or modules included in the apparatus 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to the machine-executable instructions, some or all of the units and / or modules in the apparatus 400 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.

[0110] It should be understood that one or more steps in the above methods can be executed by a suitable electronic device or a combination of electronic devices. Such an electronic device or a combination of electronic devices can include, for example, Figure 1 the terminal device 110 in

[0111] Figure 5 A block diagram of an electronic device 500 in which one or more embodiments of the present disclosure can be implemented is shown. It should be understood that Figure 5 the shown electronic device 500 is merely exemplary and should not constitute any limitation to the functions and scopes of the embodiments described herein. Figure 5 The shown electronic device 500 can include or be implemented as Figure 1 the terminal device 110 of Figure 4 the apparatus 400 of <\

[0112] As Figure 5As shown, the electronic device 500 is in the form of a general-purpose electronic device. The components of the electronic device 500 may include, but are not limited to, one or more processors or processing units 510, a memory 520, a storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 550. The processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to the programs stored in the memory 520. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the electronic device 500.

[0113] The electronic device 500 generally includes multiple computer storage media. Such media can be any accessible media that can be obtained by the electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 can be a volatile memory (such as registers, caches, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 can be a removable or non-removable medium and can include machine-readable media, such as flash drives, magnetic disks, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 500.

[0114] The electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 5 a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 520 may include a computer program product 525 having one or more program modules that are configured to perform various methods or actions of the various embodiments of the present disclosure.

[0115] The communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 500 may be implemented in a single computing cluster or multiple computer machines that are capable of communicating via a communication connection. Thus, the electronic device 500 may operate in a networked environment using a logical connection with one or more other servers, network personal computers (PCs), or another network node.

[0116] The input device 550 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 550 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 500 can also communicate with one or more external devices (not shown) via the communication unit 540 as needed. The external devices such as a storage device, a display device, etc., communicate with one or more devices that enable a user to interact with the electronic device 500, or communicate with any device (e.g., a network card, a modem, etc.) that enables the electronic device 500 to communicate with one or more other electronic devices. Such communication can be performed via an input / output (I / O) interface (not shown).

[0117] According to an exemplary implementation of the present disclosure, there is provided a computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, there is also provided a computer program product, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0118] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0119] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is produced that implements the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which causes a computer, a programmable data processing device, and / or other devices to operate in a specific manner. Thus, the computer-readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0120] The computer-readable program instructions can be loaded onto a computer, other programmable data processing device, or other device, such that a series of operation steps are performed on the computer, other programmable data processing device, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing device, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0121] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.

[0122] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art in the field without departing from the scope and spirit of the described implementations. The choice of the terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of the technology in the market, or to enable other ordinary skilled persons in the art to understand the various implementations disclosed herein.

Claims

1. A method for task interaction, comprising: In response to a request for a task of an agent, presenting first execution information for a first action for performing the task in an interface for interacting with the agent, wherein the request includes task information indicating the requirements of the task; In response to the execution of the first action, receiving an operation on the execution of the first action, the operation indicating an adjustment or confirmation of the execution of the first action, wherein the task is paused; And In response to receiving the operation on the execution of the first action, resuming the agent's execution of the task based on the operation.

2. The method according to claim 1, wherein resuming the execution of the task includes: In response to receiving an operation indicating an adjustment of the execution of the first action, receiving additional task information for adjusting the execution of the first action; In response to receiving the additional task information, having the agent re-execute the first action; And Presenting updated execution information of the first action in the interface, wherein the updated first execution information is generated by the agent re-executing the first action based on the additional task information.

3. The method according to claim 1, wherein resuming the execution of the task further includes: In response to receiving an operation indicating confirmation of the execution of the first action, having the agent continue to execute a second action of the task; And Presenting second execution information of the second action in the interface.

4. The method according to claim 1, further comprising: Receiving a selection of a mode applicable to the agent, the mode including at least a first mode and a second mode, wherein the first mode indicates executing the task based on the received operation; And In response to receiving the selection of the first mode of the agent, having the agent execute the task based on the first mode.

5. The method according to claim 4, wherein presenting a first interface element and a second interface element in the interface includes: In the case where the first mode of the agent is selected, in response to the execution of the first action, presenting a first interface element and a second interface element in the interface, the first interface element being for receiving an operation indicating an adjustment of the execution of the first action, and the second interface element being for receiving an operation indicating confirmation of the execution of the first action.

6. The method according to claim 4, wherein the second mode indicates sequentially and automatically executing multiple actions included in the execution of the task.

7. The method according to claim 1, wherein the execution of the task includes multiple actions, the multiple actions respectively indicating: Generating an execution plan for the task, the execution plan indicating at least one subtask of the task, and The at least one subtask.

8. The method according to claim 7, wherein the first action indicates generating an execution plan for the task, and wherein in response to the execution of the first action, receiving an operation on the execution of the first action includes: In response to the completion of the generation of the execution plan for the task, present the execution plan in the interface; and Receive an operation for the execution of the first action, where the operation indicates an adjustment or confirmation of the execution plan.

9. The method according to claim 7, wherein the first action indicates a first subtask among the at least one subtask, and wherein in response to the execution of the first action, receiving an operation for the execution of the first action by the user includes: In response to the completion of the execution of the first subtask, present the execution result of the first subtask in the interface; and Receive an operation for the execution of the first action, where the operation indicates an adjustment or confirmation of the execution result of the first subtask.

10. A device for task interaction, comprising: An execution information presentation module, configured to, in response to a request for a task of an agent, present first execution information for a first action for executing the task in an interface for interacting with the agent, where the request includes task information indicating the requirements of the task; An operation receiving module, configured to, in response to the execution of the first action, receive an operation for the execution of the first action, where the operation indicates an adjustment or confirmation of the execution of the first action, and where the task is paused; and An execution module, configured to, in response to receiving an operation for the execution of the first action, resume the execution of the task by the agent based on the operation.

11. An electronic device, comprising: At least one processor; and At least one memory, the at least one memory being coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, causing the electronic device to execute the method according to any one of claims 1 to 9.

12. A computer-readable storage medium, having stored thereon computer-executable instructions that can be executed by a processor to implement the method according to any one of claims 1 to 9.

13. A computer program product, comprising computer-executable instructions, where the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Multi-agent thinking chain negotiation enhancement generation method

    CN118821834A

  • Unmanned equipment autonomous task planning and execution method based on large model agent framework

    CN119090307A

Cited By

  • Task interaction method and device based on intelligent agent, equipment and storage medium

    CN121094129A