Task processing method and device, equipment and storage medium
By deploying an agent and configuration interface focused on specific tasks for the intelligent system, the problem of opaque interaction and poor controllability of intelligent systems when performing complex tasks is solved, and the quality and controllability of task execution are improved.
Patent Information
- Application Number
- CN202510497280.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-08-01
AI Technical Summary
When existing intelligent systems perform complex tasks, their interactions are opaque and poor controllability, making it difficult for users to understand and adjust the task execution logic in real time.
By deploying agents focused on specific task types, and configuring corresponding task configuration interfaces and prompt word templates, receiving task input, generating task execution information, and providing task execution plans and results.
It improves the execution quality and process controllability of specific tasks, and enhances users' understanding and adjustment of task execution process.
Smart Images

Figure CN120409681A_ABST
Abstract
Description
Technical Field
[0001] Example embodiments of the present disclosure generally relate to the field of computers, and particularly to methods, devices, equipment, computer-readable storage media, and computer-executable instruction products for task processing. Background Art
[0002] With the development of information technology, more and more intelligent systems can complete various complex tasks based on user input. Such systems usually interact with users in the form of conversations, generating response content according to users' natural language requests, thus forming question-and-answer conversations. However, in practical applications, the user's goal is often not just to obtain a single reply, but rather to hope that the system can complete a complete plan and execution around a certain task. Therefore, how to complete the execution of a specific task based on a task goal-driven model and improve the structural degree and user controllability of the execution process has become an issue worthy of attention. Summary of the Invention
[0003] In a first aspect of the present disclosure, there is provided a method for task processing. The method includes: in response to a selection of a first agent related to a first task type, presenting a first interface corresponding to the first agent, the first interface being configured to receive task inputs required for the first task type; receiving, via the first interface, a first task request for the first agent, the first task request including at least one information item required for the first task type; and in response to the first task request, presenting a second interface, the second interface presenting the first task request and task execution information for the first task request, where the task execution information is generated by the first agent based on at least one information item and a prompt word template corresponding to the first task type.
[0004] In a second aspect of the present disclosure, there is provided a device for task processing. The device includes: a first presentation module configured to present, in response to a selection of a first agent related to a first task type, a first interface corresponding to the first agent, the first interface being configured to receive task inputs required for the first task type; a receiving module configured to receive, via the first interface, a first task request for the first agent, the first task request including at least one information item required for the first task type; and a second presentation module configured to present, in response to the first task request, a second interface, the second interface presenting the first task request and task execution information for the first task request, where the task execution information is generated by the first agent based on at least one information item and a prompt word template corresponding to the first task type.
[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory, the at least one memory being coupled to the at least one processor and storing instructions for execution by the at least one processor. The instructions, when executed by the at least one processor, cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. Computer-executable instructions are stored on the computer-readable storage medium and can be executed by a processor to implement the method of the first aspect.
[0007] In a fifth aspect of the present disclosure, a computer program product is provided, including computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to the first aspect of the present disclosure.
[0008] It should be understood that the content described in this part is not intended to define the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In combination with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0010] Figure 1 A schematic diagram showing an example environment in which embodiments according to the present disclosure can be implemented is shown;
[0011] Figures 2A to 2I Schematic diagrams showing interfaces for interacting with an agent according to some embodiments of the present disclosure are respectively shown;
[0012] Figure 3 A schematic diagram showing an example architecture for task processing according to some embodiments of the present disclosure is shown;
[0013] Figure 4 A flowchart showing a process for task processing according to some embodiments of the present disclosure is shown;
[0014] Figure 5 A schematic structural block diagram showing an example device for task processing according to some embodiments of the present disclosure is shown; and
[0015] Figure 6 A block diagram of an electronic device capable of implementing multiple embodiments of the present disclosure is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0017] In the description of the embodiments of the present disclosure, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions below.
[0018] In this article, unless otherwise specified, performing a step "in response to A" does not mean that the step is immediately performed after "A", but may include one or more intermediate steps.
[0019] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.
[0020] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner according to the relevant laws and regulations.
[0021] For example, when a user's active request is received, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require the acquisition and use of the user's personal information, so that the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server or a storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.
[0022] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0023] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other ways that meet the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0024] As used herein, the term "model" can learn the corresponding association relationship between inputs and outputs from training data, so that after training is completed, for a given input, a corresponding output can be generated. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this article, "model" can also be referred to as "machine learning model", "learning model", "machine learning network", or "learning network", and these terms are used interchangeably herein.
[0025] A "neural network" is a machine learning network based on deep learning. A neural network can process inputs and provide corresponding outputs, and it generally includes an input layer and an output layer, as well as one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications generally include many hidden layers, thereby increasing the depth of the network. The layers of the neural network are connected in sequence, so that the output of the previous layer is provided as the input of the next layer, where the input layer receives the input of the neural network, and the output of the output layer is used as the final output of the neural network. Each layer of the neural network includes one or more nodes (also referred to as processing nodes or neurons), and each node processes the input from the previous layer.
[0026] Generally, machine learning can be roughly divided into three stages, namely, the training stage, the testing stage, and the application stage (also referred to as the inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values are continuously iteratively updated until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association from input to output (also referred to as the mapping from input to output) from the training data. The parameter values of the trained model are determined. In the testing stage, test inputs are applied to the trained model to test whether the model can provide correct outputs, so as to determine the performance of the model. In the application stage, the model can be used to process actual inputs based on the parameter values obtained through training and determine the corresponding outputs.
[0027] Figure 1 FIG. shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In this example environment 100, an application 125 is installed in the terminal device 110. The user 140 can interact with the application 125 via the terminal device 110 and / or an attached device of the terminal device 110.
[0028] In some embodiments, the application 125 can be downloaded and installed in the terminal device 110. In some embodiments, the application 125 can also be accessed in other ways, such as through web access, etc. InFigure 1 In an environment 100, in response to an application 125 being launched, a terminal device 110 can present an interface 150 of the application 125.
[0029] In some embodiments, the terminal device 110 communicates with a server 130 to implement service provision for the application 125. The terminal device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, desktop computer, laptop computer, notebook computer, netbook computer, tablet computer, media computer, multimedia tablet, personal communication system (PCS) device, personal navigation device, personal digital assistant (PDA), audio / video player, digital camera / video camera, television receiver, radio broadcast receiver, e-book device, gaming device, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the terminal device 110 can also support any type of user interface (such as a "wearable" circuit, etc.). The application 125 can be various types of computing systems / servers capable of providing computing capabilities, including but not limited to mainframes, edge computing nodes, computing devices in a cloud environment, and the like.
[0030] In embodiments of the present disclosure, the application 125 can provide an interaction function with an agent. The application 125 can be an application dedicated to providing agent services, or an application integrated with an agent (that is, it can provide other functions or services in addition to the agent). Although Figure 1 a single application is shown, in fact, multiple applications can be installed on the terminal device 110.
[0031] In the present disclosure, agents 160-1, 160-2,... 160-N (collectively or individually referred to as agent 160) can be deployed locally on the terminal device 110 or remotely deployed. In the case of remote deployment, the terminal device 110 can directly call the agent, or can call the agent via the server 130.
[0032] In embodiments of the present disclosure, the agent 160 can have intelligent dialogue and task processing capabilities. The terminal device 110 provides an interface 150 that can present the interaction with the agent 160. In the interface 150, the user 140 can input a task request for the agent 160 through natural language input (text input or voice input), and can also upload input online or offline file conversations to instruct the agent 160 to assist in completing various tasks.
[0033] In an embodiment of the present disclosure, during the interaction between the agent 160 and the user 140, the agent 160 can process the task indicated by the user in response to the user's request. In some embodiments, during the task processing, the agent 160 can call one or more tools 165-1, 165-2, …… 165-M (collectively or individually referred to as tools 165) according to the task requirements to assist in the execution of the task and the provision of the task result. These tools 165 can be any type of tool, such as a weather query tool, a flight query tool, an information search tool, an online or offline database, an image processing tool, a chart generation tool, a web page production tool, and so on.
[0034] In some embodiments, the environment 100 may further include a management node for a plurality of agents 160-1, 160-2, …… 160-N, and the management node can interact with the plurality of agents 160-1, 160-2, …… 160-N. In some examples, the management node can determine the task requirements corresponding to the task request in response to the task request of the user 140. After that, the management node can allocate the task request to the agent 160 that matches the task requirements based on the task requirements to request the agent 160 to execute the task. In other examples, the management node can also determine the execution plan of the task based on the task requirements. The execution plan of the task can indicate one or more subtasks required to complete the task. The management node can allocate the one or more subtasks to one or more agents 160, and the one or more agents 160 can execute their respective corresponding subtasks respectively. Regarding this management node, in some examples, the management node can be implemented by one of the plurality of agents 160-1, 160-2, …… 160-N (in this case, this agent 160 is also referred to as the management agent 160). In other examples, the management node can be implemented by a machine learning model, for example, it can be implemented by a language model (LM) or a large language model (LLM).
[0035] In some embodiments, the agent 160 can be built based on one or more machine learning models. In some embodiments, the machine learning model based on the agent 160 can include at least a language model (LM). These machine learning models include content generation models that can generate corresponding outputs based on model inputs. In some embodiments, a machine learning model based on a language model can receive model inputs in textual modalities (e.g., natural language and / or machine language) and / or model inputs in non-textual modalities (e.g., images, voice, video, etc.), and can obtain corresponding model outputs based on the model inputs and prompt words, thereby completing the execution of the task. The prompt words here are used to guide the machine learning model to generate user queries that can solve the model inputs. In an application scenario for supporting user dialogue, the input of the user 140 can be provided to the machine learning model 160 as at least a part of the model input (the other part may include prompt words). The user input is regarded as a question or query request. Based on the model output, a corresponding response can be provided to the user 140.
[0036] It should be understood that the structure and function of the various elements in the environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present disclosure.
[0037] As mentioned above, with the development of information technology, more and more intelligent systems are able to complete various tasks based on user input. Such systems typically interact with users in a dialogue format, generating responses based on users' natural language requests, thus forming a question-and-answer dialogue. Specifically, traditional intelligent systems typically interact with users in a dialogue format. Users input requests in natural language, and the system promptly generates responses, completing tasks through several rounds of dialogue. These traditional intelligent systems can provide relatively satisfactory task quality for some conventional information processing tasks (such as question-and-answering, translation, summarization, information query, etc.), but the task quality for some specific fields or some more complex tasks (such as planning, generating project reports, research and analysis, decision support, etc.) still needs to be improved.
[0038] Furthermore, while some intelligent systems have attempted to expand the functionality of language models through mechanisms such as plugin systems or tool invocations to achieve more complex goals, their interaction model still primarily relies on "automatic response." Specifically, after a user enters a task objective via natural language, the intelligent model is triggered to process it, automatically executing a series of operations and providing the user with a single feedback loop of the final results.
[0039] While this approach offers some automation capabilities, its core characteristic remains closed-loop execution, from instruction reception to output. The entire process is autonomously completed by the intelligent system, making it difficult for users to understand the task execution logic in real time, nor to intervene or adjust intermediate steps. This approach can suffer from opaque interactions and poor controllability when handling complex, dynamic, or accuracy-critical tasks. Furthermore, during the interaction process, users often wait for a while before viewing the output, only to find it unsatisfactory, necessitating a new round of dialogue. This cycle can severely impact task efficiency.
[0040] In view of this, an embodiment of the present disclosure proposes an improved scheme for task processing. In the improved scheme, if a selection of a first agent related to a first task type is received, a first interface corresponding to the first agent is presented. The first interface is configured to receive and match the task input required for the first task type. A first task request for the first agent is received via the first interface, and the first task request includes at least one information item required for the first task type. In response to the first task request, a second interface is presented. The second interface presents the first task request and task execution information for the first task request, and the task execution information is generated by the first agent based on at least one information item and a prompt word template corresponding to the first task type.
[0041] In the disclosed embodiments, not only are agents dedicated to specific types of tasks deployed, but corresponding task configuration interfaces and prompt word templates are also deployed for these tasks. The task configuration interface accurately captures the task inputs required for these tasks. Subsequently, agents dedicated to these tasks are used to execute these tasks based on matching task inputs and prompt word templates. This improves the quality of task execution and the controllability of the task execution process for the specific tasks the agents are focused on.
[0042] In this disclosure, "trigger" refers to one or more interactive operations performed by a user on a terminal device. Furthermore, these interactive operations can be triggered within the same user interface / pop-up window or within different user interfaces / pop-up windows. This disclosure is not limited in this respect.
[0043] For ease of understanding, the following will refer to Figures 2A to 2I 1 and 2 to describe examples of interfaces for interacting with an agent in some embodiments of the present disclosure. These example interfaces can be presented by the application 125 in the terminal device 110. It should be understood that the interfaces shown in the drawings are merely examples, and various interface designs can actually exist. The various graphical elements in the interface can have different arrangements and different visual representations, one or more elements can be omitted or replaced, and one or more other elements can also be present. The embodiments of the present disclosure are not limited in this respect.
[0044] In Figure 2A the interface 200A shown can present an interaction interface 200A between the user 140 and the agent 160-1. In this interaction interface 200A, a request for a new task initiated by the user can be received. As Figure 2A shown, the user can trigger a new task by triggering the "New Task" control 211. In response to the triggering of the "New Task" control 211, a task-specific interaction interface is presented in the interface 200A, such as Figure 2B the interaction interface 200B shown, to initiate a new task in the interaction interface 200B. In some embodiments, an input area 215 for task input is also provided in the interface 200A. In the input area 215, the user can input a request for the task to be triggered and initiate the task request through the "Send" control 216. The input area 215 can support text input, such as entering text in a text input box, and can also support voice input, such as by triggering a voice control 217. In addition, the input area 215 can also provide an upload control 216 to support uploading an attachment to indicate the user's task request. In some embodiments, the interaction interface 200A also provides a task list 212 initiated by the agent 160-1. The task list 212 indicates each initiated task and the execution status of the task (e.g., interrupted status, in-progress status, completed status, etc.).
[0045] In some embodiments, multiple operation modes of the user 140 and the agent 160-1 can be provided, and flexible switching can be performed between the multiple operation modes, such as via Figure 2A the mode switching control 218 shown. When a certain mode is triggered, a corresponding interaction area is presented to facilitate the interaction between the user 140 and the agent 160-1. The interaction methods between the user 140 and the agent 160-1 are different under different operation modes, so that the interaction requirements in different application scenarios can be flexibly adapted. Different operation modes can include an automatic mode and a collaborative mode. In the automatic mode, the agent 160-1 can continuously execute the task automatically until the task is completed. In the coordination mode, after the agent 160-1 completes an action of task execution, it can pause the task execution and provide modification and confirmation options so that the user can check whether the execution of the current action meets the expectations and whether the execution of the action needs to be modified. If no modification is required, after receiving the user's confirmation, the agent 160-1 will continue to execute the next action of the task until the task is completed.
[0046] After triggering a request for a task of the agent, the terminal device 110 presents Figure 2BThe interactive interface 200B shown. In the interactive interface 200B, requests for tasks initiated by the user can be presented, such as message 221, and task execution information when the agent executes the task can also be presented. The task execution information can be presented in the interactive interface 200B in the form of one or more messages. During the execution of the task, the agent 160-1 can analyze the task request initiated by the user to determine the task requirements, thereby generating an execution plan for the task. The execution plan of the task can indicate one or more subtasks required to complete the task. To enable the user to more clearly understand the execution process of the task by the agent, a message 221 indicating the generation of the execution plan can be presented in the interactive interface 200B. After the execution plan is determined, messages 223 indicating each subtask in the execution plan can be presented in the interactive interface 200B. Next, the agent 160-1 will automatically or in response to the user's confirmation continue to execute each subtask, and the execution process and execution results of each subtask can be presented in the interactive interface 200B. After the execution of the entire task is completed, the execution result of the task or an access entry to the execution result (if the execution result needs to be presented on another interface) can be presented in the interactive interface 200B. In this way, during the execution of the entire task, the user can intuitively understand how the agent specifically decomposes the user's task requirements and the specific execution steps of each subtask, so as to facilitate the user to confirm or adjust the execution process of the task or subtask, and thus obtain the desired execution result.
[0047] In some embodiments, the process of determining one or more subtasks required to complete the task in the execution plan of the task based on the task request of the user 140 can be implemented by a machine learning model deployed at the server 130 or called by the server 130. Further, after determining the specific task execution plan, the machine learning model can assign each subtask to one or more suitable agents 160 for specific task execution.
[0048] Alternatively or additionally, the task request of the user can also be sent to the management agent 160. The process of determining one or more subtasks in the execution plan of the task can be implemented by the management agent 160. The management agent 160 can be used to parse the task request of the user based on the task request of the user 140, thereby determining one or more subtasks required to complete the task in the execution plan of the task. Further, the management agent 160 can assign the subtasks to one or more suitable agents 160 for specific task execution according to different task execution requirements and capabilities.
[0049] In some embodiments, during the task execution with different agents 160, pausing of the task can be supported at various execution stages of the task. This allows the user to determine whether to continue task execution or input additional task information while the task is paused. The pausing of the task can be manually triggered by the user. For example, after the user views the task execution information presented by the agent and determines that the task needs to be paused to input modification suggestions. In this case, a pause control can be provided in the interface for the user to initiate the pause. In some embodiments, the pausing of the task can also be automatic. For example, the agent automatically pauses after determining that certain pause conditions are met. Examples of automatic pausing can include the execution time of a subtask disassembled from the task exceeding a threshold (which may be due to task reasons or network link reasons), and the execution of a subtask requiring additional input from the user (e.g., authorization for certain information sources or the need for the user to provide auxiliary information), and so on.
[0050] In some embodiments, during the task execution by the agent 160, it can be in a paused state. Options such as modification and resume can be provided during the paused state so that the user can check whether the execution of the current action meets expectations and whether the execution of the action needs to be modified. Then, in the case of modifying the task, the modified task can be passed to the management agent 160. The management agent 160 can allocate the modification of the task to other agents. During the task execution, the user 140 can add a new task. For example, add a XX subtask. In this scenario, the terminal device 110 can call the management agent 160 to add the XX subtask. After the management agent 160 receives the XX subtask, it can modify the entire task. Subsequently, the agent 160 can re - allocate the task to each agent. In some embodiments, the terminal device 110 can call the agent that can solve the XX subtask to process it.
[0051] In some embodiments, the management agent 160 can determine a prompt word corresponding to the additional task information indicating the modification or addition of a subtask. This prompt word can be spliced into the original prompt word as a component of the system prompt (SP), so as to provide context guidance for subsequent task execution. Alternatively or additionally, the prompt word corresponding to the additional task information is directly spliced into the task prompt word of the agent corresponding to the subtask, so that the corresponding agent can take into account the new context content when executing the task.
[0052] Some example embodiments of the present disclosure will be further described below with reference to the accompanying drawings. The embodiments related to the present disclosure can be implemented in the terminal device 110. It should be noted that the operations performed by the terminal device 110 can specifically be performed by relevant applications (such as application 120) installed on the terminal device 110. Some operations described with reference to the terminal device 110 may require the assistance of the server 130 to complete.
[0053] In some embodiments, if a selection of the agent 160-2 (sometimes also referred to as the first agent in this article) related to the first task type is detected, the terminal device 110 presents a first interface corresponding to the first agent. The first task type can be any suitable task type, such as a management task, an analysis task, a creation task, etc. The agent 160-2 here is constructed for tasks of the first task type and can be understood as an agent focusing on tasks of the first task type. In some cases, the agent 160-2 can also be referred to as, for example, an "expert agent" focusing on a specific task type. The first interface can be understood as an interaction interface deployed for the first task type, and the first interface can be configured to receive task inputs required for the first task type.
[0054] In an actual application scenario, the application 125 can provide multiple entrances for selecting the agent 162-. As an example, Figure 2C shows an interaction interface 200C between the user 140 and the agent 160-1. As Figure 2C shown, in this interaction interface 200C, controls 230-1, 230-2, and 230-3 are presented. The controls 230-1, 230-2, and 230-3 correspond to the agents 160-2, 160-3, and 160-4 respectively. The agent 160-2 can be configured to focus on tasks in the "AAA" field and can be called the "expert agent A". The agent 160-3 can be configured to focus on tasks in the "BBB" field and can be called the "expert agent B". The agent 160-4 can be configured to focus on tasks in the "CCC" field and can be called the "expert agent C". The user 140 can select the agent 160-2 by triggering the control 230-1. The terminal device 110 can determine that a selection of the agent 160-2 is detected in response to the triggering of the control 230-1 and present the first interface.
[0055] In some embodiments, the terminal device 110 receives a task request for the agent 160-2 (sometimes also referred to herein as the first task request) via a first interface. The task request includes at least one information item required for the first task type. Specifically, the user 140 may input at least one information item required for the first task type via the first interface. The terminal device 110 receives the at least one information item as all or part of the content of the task request. It can be understood that the specific content (i.e., information item) included in the task request is related to the task type targeted by the agent 160-2. The following provides an exemplary description of the content included in the task request with some examples, but it should not be understood that the task request is limited to including the content shown below.
[0056] In some embodiments, the first interface may include at least one input control corresponding to the at least one information item. The user 140 may input the at least one information item via the at least one input control respectively. The input control may be configured in any suitable type or form, such as an input box, an option box, an upload control (which can upload, for example, text, image or video), etc. As an example, the agent 160-2 may be configured to perform, for example, a content analysis task. Figure 2D An interaction interface 200D (i.e., the first interface) between the user 140 and the agent 160-2 is shown. As Figure 2D shown, for the "content analysis task", the interaction interface 300D may present input controls 241, input controls 242, input controls 243, and input controls 245, etc. The user 140 may input the at least one information item via the input controls 241 to input controls 245. It can be understood that the above first interface is only exemplary, and the first interface for different task types may be different. In an actual application scenario, any suitable input control may be selected for the first interface according to the task type. The embodiments of the present disclosure do not limit this.
[0057] In some embodiments, the task request for the agent 160-2 may include a first information item, and the first information item may indicate the task requirements of the first task corresponding to the task request. The terminal device 110 may determine the task requirements of the first task by accepting the first information item.
[0058] In some examples, the first interface may present a first input control for inputting keywords. The user 140 may input at least one keyword via the first input control, and the terminal device 110 may receive the at least one keyword input by the user 140 via the first input control. As an example, in combination with Figure 2D and Figure 2EAs shown, the interaction interface 200D can present an input control 241, and the user 140 can use the terminal device 110 or an attached device of the terminal device 110 to input one or more keywords to the input control 241. The terminal device 110 can receive the keywords input by the user via the input control 241, and present the keywords input by the user 140 through the input control 241, such as keyword 1, keyword 2, as Figure 2E shown.
[0059] Alternatively or additionally, the first interface can also present a second input control. The terminal device 110 can receive description information about the task requirements via the second input control. As an example, in combination with Figure 2D and Figure 2E shown, the user 140 can use the terminal device 110 or an attached device of the terminal device 110 to input a text description 251 of the task requirements to the input control 242 (such as a text box). The terminal device 110 can present the text description 251 through the input control 242. It should be noted that in actual applications, it is not limited to describing the task requirements by text, and the task requirements can also be indicated by, for example, images, audio, or video. In this case, the second input control can also be configured as an upload control for images, audio, or video, a trigger control for an image acquisition device, or a trigger control for an audio acquisition device. That is, the terminal device 110 can also respond to the triggering of the second input control to activate the image acquisition device to acquire images or videos, or activate the audio acquisition device to acquire audio, as the description information of the task requirements.
[0060] Alternatively or additionally, the first interface can also present a third input control. The terminal device 110 can receive at least one associated word via the third input control, and the at least one associated word can be generated based on at least one of the keywords or the description information. The associated word here can be a word formed by combining multiple keywords, a word formed by combining a keyword and a word in the description information, a word formed by combining a keyword and other words (that is, other words outside the keyword and the description information), a word formed by combining words in the description information and other words, a word with a semantic proximity to the keyword or a word in the description information, or it can also be a word formed by combining multiple words with a semantic proximity to the keyword or a word in the description information.
[0061] In some examples, if at least one keyword and / or description information is received, and a request for generating associated keywords is received, the terminal device 110 may generate multiple candidate associated keywords. The terminal device 110 may present the multiple candidate associated keywords in the third input control. The user 140 may select all or part of the associated keywords (i.e., at least one associated keyword) from the multiple candidate associated keywords, and the terminal device 110 receives at least one associated keyword selected from the multiple candidate associated keywords. In this way, the user 140 can flexibly select associated keywords according to actual needs to supplement or expand the description of the task requirements, which is beneficial to improving the task execution quality.
[0062] As an example, continuing to combine Figure 2E and Figure 2F shown, the interaction interface 200E may present an input control 243. Before the terminal device receives a request for generating associated keywords, the input control 243 may present a "Generate Associated Keywords" control 244. If the user 140 needs to generate associated keywords, the control 244 may be triggered (such as clicking, double-clicking, long-pressing, etc.). The terminal device 110 may determine that a request for generating associated keywords is received in response to the triggering of the control 244. The terminal device 110 may generate multiple candidate associated keywords by using the intelligent agent 160-2 based on the keyword presented in the input control 241 and the description information presented in the input control 242. After that, the terminal device 110 may control the input control 243 to switch from the display state shown in the interaction interface 200E to the display state shown in the interaction interface 200F. Specifically, the control 244 in the input control 243 may be removed, and associated keywords 261, associated keywords 262, etc. may be presented in the input control 243. Of course, the terminal device 110 may generate keywords based on, for example, an algorithm or a search engine. The embodiments of the present disclosure do not limit this.
[0063] In some examples, the terminal device 110 may determine one or more information dimensions indicating the type or characteristics of the associated keywords. The terminal device 110 may generate multiple candidate associated keywords based on at least one of the keyword or the description information and the determined one or more information dimensions. Each candidate associated keyword conforms to the corresponding information dimension. As an example, as Figure 2F shown, the terminal device 110 may generate associated keywords 261 and associated keywords 262 that conform to the first information dimension, associated keywords 263 and associated keywords 264 that conform to the second information dimension, and associated keywords 265 and associated keywords 266 that conform to the third information dimension. After that, the terminal device 110 may present multiple information dimensions (i.e., the first information dimension, the second information dimension, and the third information dimension) in the input control 243, and present the associated keywords corresponding to each information dimension.
[0064] In some examples, the first interface may also present a fourth input control. The terminal device 110 may receive the execution mode of the first task via the fourth input control. The execution mode may indicate at least one of the execution time or execution frequency of the first task. In practical applications, multiple execution modes may be preset, and the user 140 may input or select an execution mode via the fourth input control. The terminal device 110 may determine the selected or input execution mode as the execution mode of the first task.
[0065] As an example, as shown in Figure 2E and Figure 2F , the terminal device 110 may preset a first execution mode and a second execution mode for the first task. The first execution mode may indicate that the first task can be executed repeatedly at regular intervals (for example, repeated daily). The second execution mode may indicate that the first task is executed once currently. The interaction interface 200E may present an input control 245, and the input control 245 may include an option 246 corresponding to the first execution mode and an option 247 corresponding to the second execution mode. If the user 140 selects the option 246 through the terminal device 110 or an attached device of the terminal device 110, the terminal device 110 may highlight (such as bold, highlight, etc.) the option 246, as shown in the interaction interface 200F, to indicate the selection of the first execution mode. Thus, the user can define the execution mode of the task independently, which can improve the flexibility and autonomy of task execution.
[0066] In some examples, the first interface may also present a fifth input control. The terminal device 110 may receive the task output requirement for the first task via the fifth input control. The task output requirement may indicate the output content of the first task, the format of the output content, etc. As an example, multiple candidate output contents and multiple candidate output formats may be determined in advance. The terminal device 110 may present multiple options corresponding to the multiple candidate output contents respectively, and multiple options corresponding to the multiple candidate output formats on the first interface. The user 140 may select one or more output contents via the first interface and may select the output type. The terminal device 110 may determine the task output requirement according to the selection of the output content and the output type. In this way, the task output can better meet the actual needs of the user 140.
[0067] In some embodiments, the task request for agent 160-2 may further include a second information item, which may indicate at least one data source for the first task. The at least one data source is used to provide the data required for the first task. Specifically, agent 160-2 may obtain the data required for the first task from the at least one data source according to the second information item. Thus, user 140 can independently select the data source according to actual needs, which can improve the controllability of the task and make the execution result of the task meet the actual needs of user 140. The data source here may include, but is not limited to, applications, websites, databases, cloud platforms, servers, or terminal devices, etc. Of course, the above data sources are only exemplary, and any appropriate source can be selected according to actual needs.
[0068] In some examples, if a confirmation of the first information item is received, terminal device 110 may present a third interface. The third interface may present one or more candidate data sources. User 140 may select at least one data source from the one or more candidate data sources, and terminal device 110 receives the second information item indicating the selected at least one data source. Further, terminal device 110 may determine one or more candidate data sources based on at least one of the first task type or the first information item. For example, if the first task type indicates an analysis task of real-time hot news. Terminal device 110 may select one or more applications or websites that provide real-time news as candidate data sources. It should be noted here that terminal device 110 needs to obtain the authorization of user 140 and the data provider to obtain data from the data source. For example, in some cases, logging in to the registered account of user 140 on the application or website may be a way to obtain the authorization of user 140 and the data provider. Another example is that accessing the database through the data interface opened by the service provider of the database can also be regarded as a way to obtain the authorization of the service provider of the database.
[0069] As an example, in combination with Figure 2F and Figure 2GAs shown, the interactive interface 200F presents a control 249. If the user confirms that the content of the first information item presented in the interactive interface 200F is correct, the control 249 can be triggered by the terminal device 110 or an attached device of the terminal device 110. In response to the triggering of the control 249, the terminal device 110 determines that it has received the confirmation of the first information item, and switches from presenting the interactive interface 200F to presenting the interactive interface 200G. The interactive interface 200G presents multiple candidate data sources such as candidate data sources 271 to 276. The user 140 can select one or more data sources from the multiple candidate data sources. For example, if the user 140 selects candidate data sources 271, 272, 273, and 275, the terminal device 110 can highlight (e.g., fill the corresponding option boxes) the candidate data sources 271, 272, 273, and 275 in the interactive interface 200G.
[0070] In some embodiments, the task request for the agent 160-2 may further include a third information item, and the third information item may indicate the data required for the first task. Specifically, the user 140 can independently provide the data required for the first task. Suppose the first task type indicates a data analysis task. The user 140 can independently provide the data to be analyzed, such as text, images, videos, audio, and so on. In this way, the flexibility, diversity, and controllability of the agent 160-2 in performing tasks can be further improved. As an example, an upload control for data can be presented in the interactive interface between the user 140 and the agent 160-2. The terminal device 110 can receive the data required for the first task through the upload control.
[0071] In some embodiments, if an update request for the first task request is received, the terminal device 110 presents at least a fourth interface. The fourth interface can present at least one information item received via the first interface, and at least one information item in the fourth interface is editable. The user 140 edits at least some of the at least one information item via the fourth interface, such as updating keywords, associated words, execution modes, and so on. After determining that the update of the first task request is completed, the user 140 can trigger the update confirmation of the at least one information item, such as selecting the "Next" or "Confirm" button. If an update confirmation of the at least one information item is received, the terminal device 110 receives the updated first task request. The updated first task request includes at least the updated at least one information item, such as updated keywords, updated description information, or updated execution modes, and so on.
[0072] As an example, in combination with Figure 2G 、 Figure 2H and Figure 2IAs shown, if user 140 triggers control 277 in interactive interface 200G, terminal device 110 determines that it has received a first task request, and terminal device 110 may present interactive interface 200H. Interactive interface 200H may be a trial run interface for the first task, and the execution mode and trial run information of the first task, such as the trial run process and trial run results of the first task, etc., may be presented through the trial run interface. User 140 may determine whether each information item in the first task request meets the requirements based on the trial run results. If user 140 determines that the first task request needs to be modified, control 281 in interactive interface 200H may be triggered. Terminal device 110 may, in response to the triggering of control 281, present interactive interface 200I. Task setting interface 290 (i.e., the fourth interface) is presented in interactive interface 200I. Input controls 291, 292, 293, and 294 are presented in task setting interface 290, and keywords, description information, associated words, and execution modes input via interactive interface 200F are respectively presented in input controls 291, 292, 293, and 294. User 140 may update the above-mentioned information items via input controls 291, 292, 293, and 294. If user 140 determines that the update of the first task request is completed, control 295 may be triggered. Terminal device 110 receives the updated first task request in response to the triggering of control 295.
[0073] As another example, terminal device 110 may also, in response to the triggering of control 295, present a fifth interface or a sixth interface. The fifth interface may present a second information item indicating at least one data source, and the sixth interface may present the data required for the first task (such as uploaded data). User 140 may update the data source via the fifth interface, and may also update the data required for the first task via the sixth interface, such as updating uploaded files, images, videos, audios, etc. If user 140 determines that the update of the data source or the data required for the first task is completed, the confirmation option of the corresponding interface may be triggered. Terminal device 110 may receive the updated first task request in response to the triggering of the confirmation option.
[0074] In some embodiments of the present disclosure, in response to the first task request, terminal device 110 presents a second interface. The second interface presents the first task request and task execution information for the first task request, and the task execution information is generated by the first agent based on at least one information item and a prompt word template corresponding to the first task type.
[0075] In some embodiments, as introduced in the foregoing analysis, the task execution information may indicate at least one of the following: a task plan generated for the first task request, the execution process of the task plan, or the execution result of the task plan. As an example, in combination with Figure 2G and Figure 2BAs shown, if user 140 triggers control 277 in interactive interface 200G, terminal device 110 determines that a first task request is received and presents interactive interface 200B (i.e., the second interface). In interactive interface 200B, an execution plan for the first task can be presented. During the execution of the first task, interactive interface 200B can present the execution process of the task plan. After the entire task is completed, interactive interface 200B can present the execution result of the task plan. Since interactive interface 200B has been introduced in detail above, only a brief description of the information presented by interactive interface 200B is given here. For specific details, refer to the detailed introduction of interactive interface 200B above.
[0076] In some embodiments, terminal device 110 generates a task prompt word corresponding to the first task type based on at least one information item and a prompt word template. After that, terminal device 110 can generate task execution information based on the task prompt word by using agent 160-2, such as formulating a task plan, executing a task plan, generating an execution result, and so on. In some examples, the prompt word template can be configured to have at least one position corresponding to the at least one information item. Terminal device 110 can generate a task prompt word by adding the at least one information item to the corresponding positions in the prompt word template respectively. After that, terminal device 110 can provide the task prompt word to agent 160-2 to request agent 160-2 to formulate a task plan, execute a task plan, or generate an execution result, and so on.
[0077] In some embodiments, in order to improve the task execution quality of the task type that agent 160-2 focuses on (i.e., the first task type), one or more tools that can provide relatively better task execution quality (such as higher accuracy, higher operating efficiency, or lower system resource consumption, etc.) can be pre-selected for the first task type, and a first tool set is constructed based on the one or more tools. The tools here can include, but are not limited to, application programs, plugins, machine learning models, and so on. During the task execution process, terminal device 110 can use agent 160-2 to call one or more tools from the first tool set related to the first task type based on the task prompt word to generate task execution information. Thus, the task execution instruction of the task type that agent 160-2 focuses on can be significantly improved.
[0078] As an example, Figure 3 shows a schematic diagram of an example architecture 300 for task processing according to some embodiments of the present disclosure. As Figure 3 shown, for the first task type (such as a data analysis task), a tool set 350 can be pre-constructed. Tool set 350 includes multiple tools such as tool 352-1, tool 352-2, tool 352-3, …, tool 350-K.
[0079] At block 310, the terminal device 110 may perform task configuration to obtain a first task request. The process of receiving the first task request has been described in detail in the foregoing analysis and will not be elaborated here.
[0080] At block 320, the terminal device 110 may obtain the data required for the first task. For example, the terminal device 110 may obtain the data required for the first task from data sources A, B, C, and D shown in the interaction interface 200G respectively based on at least one of keywords, associated words, and descriptive information, so as to form data sets 322, 324, 326, and 328.
[0081] At block 330, the terminal device 110 may perform preprocessing on the obtained data to obtain a task data set 340 corresponding to the first task. As an example, the terminal device 110 may perform operations such as data cleaning, data classification, and adding labels on the data in the data sets 322, 324, 326, and 328, and construct the task data set 340 based on the processed data.
[0082] After that, the agent 160-2 may call one or more tools 352 in the tool set 350 corresponding to the first task type to perform the first task, such as formulating a task plan, executing a task plan, or generating a task result 360. Thus, a relatively high task execution quality can be obtained for the task type that the agent 160-2 focuses on. Of course, the agent 160-2 may also use the tools 352 in the tool set 350 to perform operations such as data collection and preprocessing.
[0083] In the embodiments of the present disclosure, not only agents focusing on specific types of tasks are deployed, but also corresponding task configuration interfaces and prompt word templates are deployed for such tasks. Through the task configuration interface, the task inputs required for such tasks can be accurately obtained. After that, using the agents focusing on such tasks, such tasks are performed based on the mutually matching task inputs and prompt word templates. Thus, the task execution quality and the controllability of the task execution process of the specific tasks that the agents focus on can be improved.
[0084] Example processes, apparatuses, and devices
[0085] Figure 4 A flowchart of a process 400 for task processing according to some embodiments of the present disclosure is shown. The process 400 may be implemented or included in the terminal device 110. The following will refer to Figure 1 describe the process 400.
[0086] At block 410, in response to the selection of a first agent associated with a first task type, the terminal device 110 presents a first interface corresponding to the first agent, and the first interface is configured to receive task inputs required for the first task type.
[0087] At block 420, the terminal device 110 receives, via the first interface, a first task request for the first agent, and the first task request includes at least one information item required for the first task type.
[0088] At block 430, in response to the first task request, the terminal device 110 presents a second interface, and the second interface presents the first task request and task execution information for the first task request, where the task execution information is generated by the first agent based on at least one information item and a prompt word template corresponding to the first task type.
[0089] In some embodiments, the task execution information indicates at least one of the following: a task plan generated for the first task request, a process of executing the task plan, or a result of executing the task plan.
[0090] In some embodiments, the first interface includes at least one input control corresponding to at least one information item, and receiving the first task request includes: receiving at least one information item input via the at least one input control.
[0091] In some embodiments, receiving the first task request includes at least one of the following: receiving a first information item indicating a task requirement of a first task corresponding to the first task request, receiving a second information item indicating at least one data source of the first task, where the at least one data source is used to provide data required for the first task, or receiving a third information item indicating data required for the first task.
[0092] In some embodiments, the first information item includes at least one of the following: at least one keyword related to the task requirement, descriptive information about the task requirement, at least one associated word, where the at least one associated word is generated based on at least one of the at least one keyword or the descriptive information, an execution mode of the first task, or a task output requirement of the first task.
[0093] In some embodiments, receiving at least one associated word includes: in response to a generation request for the associated word, presenting a plurality of candidate associated words within a third input control, where the plurality of candidate associated words are generated based on at least one of the at least one keyword or the descriptive information; and receiving at least one associated word selected from the plurality of candidate associated words.
[0094] In some embodiments, receiving the second information item includes: presenting a third interface in response to confirmation of the first information item, the third interface presenting one or more candidate data sources; and receiving the second information item indicating the at least one selected data source in response to selection of at least one of the one or more candidate data sources.
[0095] In some embodiments, process 400 further includes: presenting at least a fourth interface in response to an update request for the received first task request, the fourth interface presenting at least one information item received via the first interface and at least one information item in the fourth interface being editable; and receiving the updated first task request in response to confirmation of the update of the at least one information item, the updated first task request including at least the updated at least one information item.
[0096] In some embodiments, the task execution information is generated by: generating a task prompt word corresponding to the first task type based on at least one information item and a prompt word template; and generating the task execution information using a first agent based on the task prompt word.
[0097] In some embodiments, generating the task execution information using a first agent based on the task prompt word includes: invoking one or more tools from a first tool set related to the first task type using the first agent based on the task prompt word to generate the task execution information.
[0098] In some embodiments, generating the task execution information using a first agent based on the task prompt word includes: determining, using the first agent based on the task prompt word, a task plan corresponding to the first task request, the task plan indicating at least a plurality of subtasks; and determining an execution process and an execution result of the task plan based on execution of the plurality of subtasks.
[0099] Embodiments of the present disclosure also provide corresponding apparatuses for implementing the above methods or processes. Figure 5 A schematic structural block diagram of an example apparatus 500 for task processing according to certain embodiments of the present disclosure is shown. Apparatus 500 may be implemented as or included in a terminal device 110. Each module / component in apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.
[0100] As Figure 5As shown, the apparatus 500 includes: a first presentation module 510 configured to present, in response to a selection of a first agent associated with a first task type, a first interface corresponding to the first agent, the first interface being configured to receive task inputs required for the first task type; a receiving module 520 configured to receive, via the first interface, a first task request for the first agent, the first task request including at least one information item required for the first task type; and a second presentation module 530 configured to present, in response to the first task request, a second interface that presents the first task request and task execution information for the first task request, where the task execution information is generated by the first agent based on at least one information item and a prompt word template corresponding to the first task type.
[0101] In some embodiments, the task execution information indicates at least one of the following: a task plan generated for the first task request, a process of executing the task plan, or a result of executing the task plan.
[0102] In some embodiments, the first interface includes at least one input control corresponding to at least one information item, and the receiving module 520 is further configured to: receive at least one information item input via the at least one input control.
[0103] In some embodiments, the receiving module 520 is further configured to perform at least one of the following: receive a first information item indicating a task requirement of a first task corresponding to the first task request, receive a second information item indicating at least one data source of the first task, the at least one data source being used to provide data required for the first task, or receive a third information item indicating data required for the first task.
[0104] In some embodiments, the first information item includes at least one of the following: at least one keyword related to the task requirement, descriptive information about the task requirement, at least one correlation word, the at least one correlation word being generated based on at least one of the at least one keyword or the descriptive information, an execution mode of the first task, or a task output requirement of the first task.
[0105] In some embodiments, the receiving module 520 is further configured to: present, in response to a generation request for a correlation word, a plurality of candidate correlation words in a third input control, the plurality of candidate correlation words being generated based on at least one of the at least one keyword or the descriptive information; and receive at least one correlation word selected from the plurality of candidate correlation words.
[0106] In some embodiments, the receiving module 520 is further configured to: present, in response to confirmation of the first information item, a third interface that presents one or more candidate data sources; and receive, in response to a selection of at least one data source from the one or more candidate data sources, a second information item indicating the selected at least one data source.
[0107] In some embodiments, the apparatus 500 further includes: an update module configured to, in response to an update request for the received first task request, present at least a fourth interface, the fourth interface presenting at least one information item received via the first interface, and at least one information item in the fourth interface being editable; and in response to an update confirmation for at least one information item, receive the updated first task request, the updated first task request including at least the updated at least one information item.
[0108] In some embodiments, the apparatus 500 further includes: an execution module configured to generate task execution information by: generating a task prompt word corresponding to a first task type based on at least one information item and a prompt word template; and generating task execution information by using a first agent based on the task prompt word.
[0109] In some embodiments, the execution module is further configured to: call one or more tools from a first tool set related to the first task type by using a first agent based on the task prompt word to generate task execution information.
[0110] In some embodiments, the execution module is further configured to: determine a task plan corresponding to the first task request by using a first agent based on the task prompt word, the task plan at least indicating a plurality of subtasks; and determine an execution process and an execution result of the task plan based on the execution of the plurality of subtasks.
[0111] The units and / or modules included in the apparatus 500 may be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules may be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to the machine-executable instructions, some or all of the units and / or modules in the apparatus 500 may be at least partially implemented by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0112] Figure 6 A block diagram of an electronic device 600 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that Figure 6 The electronic device 600 shown is merely exemplary and should not constitute any limitation to the functions and scopes of the embodiments described herein. Figure 6 The electronic device 600 shown may include or be implemented as Figure 1the terminal device 110, or Figure 5 the device 500.
[0113] As Figure 6 shown, the electronic device 600 is in the form of a general-purpose electronic device. The components of the electronic device 600 may include, but are not limited to, one or more processors 610, a memory 620, a storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processor 610 may be a physical or virtual processor and is capable of performing various processes according to computer-executable instructions stored in the memory 620. In a multi-processor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing ability of the electronic device 600.
[0114] The electronic device 600 generally includes multiple computer storage media. Such media can be any accessible media that can be obtained by the electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 620 may be a volatile memory (such as registers, caches, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 630 may be a removable or non-removable medium and may include machine-readable media, such as a flash drive, a magnetic disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 600.
[0115] The electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 6 it, a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 620 may include a computer program product 625 having one or more computer-executable instruction modules configured to perform various methods or actions of the various embodiments of the present disclosure.
[0116] The communication unit 640 enables communication with other electronic devices through a communication medium. Additionally, the functions of the components of the electronic device 600 may be implemented by a single computing cluster or multiple computer machines that are capable of communicating through a communication connection. Thus, the electronic device 600 may operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or another network node.
[0117] The input device 650 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 660 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 600 can also communicate with one or more external devices (not shown) as needed through the communication unit 640. The external devices such as a storage device, a display device, etc., communicate with one or more devices that enable a user to interact with the electronic device 600, or communicate with any device that enables the electronic device 600 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).
[0118] According to an exemplary implementation of the present disclosure, there is provided a computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, there is also provided a computer program product, the computer program product being tangibly stored on a non-transitory computer-readable storage medium and including computer-executable instructions, and the computer-executable instructions being executed by a processor to implement the method described above.
[0119] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-executable instructions.
[0120] These computer-executable instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is produced that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-executable instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, the programmable data processing device, and / or other devices to work in a specific manner. Thus, the computer-readable storage medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0121] Computer-executable instructions can be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to generate a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0122] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-executable instruction products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, an executable instruction, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0123] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of the technology in the market, or to enable other ordinary skilled persons in the art to understand the various implementation manners disclosed herein.
Claims
1. A method for task processing, comprising: In response to the selection of a first agent associated with a first task type, presenting a first interface corresponding to the first agent, the first interface being configured to receive task inputs required for the first task type; Receiving, via the first interface, a first task request for the first agent, the first task request including at least one information item required for the first task type; And In response to the first task request, presenting a second interface, the second interface presenting the first task request and task execution information for the first task request, wherein the task execution information is generated by the first agent based on the at least one information item and a prompt word template corresponding to the first task type.
2. The method according to claim 1, wherein the task execution information indicates at least one of the following: A task plan generated for the first task request, The execution process of the task plan, or The execution result of the task plan.
3. The method according to claim 1, wherein the first interface includes at least one input control corresponding to the at least one information item, and wherein receiving the first task request includes: Receiving the at least one information item input via the at least one input control.
4. The method according to claim 1, wherein receiving the first task request includes at least one of the following: Receiving a first information item indicating the task requirements of a first task corresponding to the first task request, Receiving a second information item indicating at least one data source of the first task, the at least one data source being used to provide data required for the first task, or Receiving a third information item indicating the data required for the first task.
5. The method according to claim 4, wherein the first information item includes at least one of the following: At least one keyword related to the task requirements, Descriptive information about the task requirements, At least one correlation word, the at least one correlation word being generated based on at least one of the at least one keyword or the descriptive information, The execution mode of the first task, or The task output requirements of the first task.
6. The method according to claim 5, wherein receiving the at least one correlation word includes: In response to a request for generating correlation words, presenting a plurality of candidate correlation words, the plurality of candidate correlation words being generated based on at least one of the at least one keyword or the descriptive information; And Receiving the at least one correlation word selected from the plurality of candidate correlation words.
7. The method according to claim 4, wherein receiving the second information item includes: In response to the confirmation of the first information item, presenting a third interface, the third interface presenting one or more candidate data sources; And In response to the selection of at least one data source from the one or more candidate data sources, receiving the second information item indicating the selected at least one data source.
8. The method according to claim 1, further comprising: In response to an update request for the received first task request, at least present a fourth interface that presents the at least one information item received via the first interface, and the at least one information item is editable in the fourth interface; and In response to an update confirmation for the at least one information item, receive an updated first task request that at least includes the updated at least one information item.
9. The method according to claim 1, wherein the task execution information is generated by: generating a task prompt word corresponding to the first task type based on at least the at least one information item and the prompt word template; and generating the task execution information by using the first agent based on the task prompt word.
10. The method according to claim 9, wherein generating the task execution information by using the first agent based on the task prompt word includes: Based on the task prompt word, using the first agent to call one or more tools from a first tool set related to the first task type to generate the task execution information.
11. The method according to claim 9, wherein generating the task execution information by using the first agent based on the task prompt word includes: Based on the task prompt word, using the first agent to determine a task plan corresponding to the first task request, the task plan at least indicating a plurality of subtasks; and Based on the execution of the plurality of subtasks, determining the execution process and execution result of the task plan.
12. An apparatus for task processing, comprising: A first presentation module configured to present a first interface corresponding to the first agent in response to a selection of the first agent related to the first task type, the first interface being configured to receive a task input required for the first task type; A receiving module configured to receive a first task request for the first agent via the first interface, the first task request at least including at least one information item required by the first task type; and A second presentation module configured to present a second interface in response to the first task request, the second interface presenting the first task request and task execution information for the first task request, wherein the task execution information is generated by the first agent based on the at least one information item and a prompt word template corresponding to the first task type.
13. An electronic device, comprising: At least one processor; and At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions when executed by the at least one processor cause the electronic device to execute the method according to any one of claims 1 to 11.
14. A computer-readable storage medium having computer-executable instructions stored thereon, the computer-executable instructions being executable by a processor to implement the method according to any one of claims 1 to 11.
15. A computer program product comprising computer-executable instructions which, when executed by a processor, implement the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Intelligent task discovery
CN107491469A
Intelligent agent-based information processing method and device, electronic equipment and storage medium
CN118885592A
Task processing method and device, equipment and storage medium
CN119003023A
Intelligent agent configuration method and device, electronic equipment, storage medium and computer program product
CN119025184A