Intelligent agent generation method, data processing method, device, equipment and medium

By identifying interactive pages and operational actions, analyses that perform target tasks are generated, which solves the problem that traditional agent development relies on large-scale data annotations, and achieves efficient and low-cost agent development.

CN120215922AInactive Publication Date: 2025-06-27BEIJING ZHONGBING ZHIHANG SOFTWARE TECH CO LTD

Patent Information

Application Number
CN202510559267.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The development of traditional agents relies on the preparation and annotation of large-scale training data sets, resulting in high costs and inefficient development.

Method used

By obtaining the interactive page displayed during the execution of the target task and its operational actions, the interactive page is identified to obtain the attribute information of the page element, determine the target page element for which the operational action is targeted, and generate an agent that performs the target task based on this information.

Benefits of technology

No manual encoding and labeling of data is required, which reduces the dependence of the agent generation process on data, reduces development costs, improves development efficiency, and enables rapid construction and deployment of agents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215922A_ABST
    Figure CN120215922A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to an agent generation method, a data processing method, a data processing device, equipment and a medium. The intelligent agent generation method comprises the steps of obtaining an interaction page displayed in an execution process of a target task and action information of an operation action aiming at the interaction page; identifying the interactive page to obtain attribute information of page elements included in the interactive page; based on the attribute information and the action information of the page elements, determining a target page element which the operation action aims at in the page elements included in the interactive page; and generating an intelligent agent for executing the target task based on the target page element and the action information. According to the method, the intelligent agent can be automatically generated in combination with identification of the interaction page and the action information of the operation action, manual coding and data labeling are not needed, dependence on data in the intelligent agent generation process is reduced, the workload of manual coding and debugging can be greatly reduced, and the development cost of the intelligent agent is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to a method and apparatus for generating an agent, a method and apparatus for data processing, an electronic device, and a computer-readable storage medium. Background Art

[0002] With the rapid development of artificial intelligence and computer software technologies, an agent, as a software entity capable of autonomously completing specific tasks, has been widely used in fields such as automation and process optimization. However, the development of traditional agents relies on the preparation and annotation of large-scale training data sets, which is not only costly but also significantly reduces the development efficiency. Summary of the Invention

[0003] At least one embodiment of the present disclosure provides a method for generating an agent, including: obtaining an interaction page displayed during the execution of a target task and action information of operation actions for the interaction page; identifying the interaction page to obtain attribute information of page elements included in the interaction page; determining, based on the attribute information of the page elements and the action information, a target page element among the page elements included in the interaction page that the operation action is directed to; and generating an agent for executing the target task based on the target page element and the action information.

[0004] In at least one embodiment of the present disclosure, the method for generating an agent further includes: obtaining annotation information, where the annotation information includes variable description information for describing variables involved in the target task; wherein, generating an agent for executing the target task based on the target page element and the action information includes: generating an agent for executing the target task based on the target page element, the action information, and the annotation information.

[0005] In at least one embodiment of the present disclosure, the method for generating an agent further includes: obtaining agent description information of the agent for executing the target task; wherein, generating an agent for executing the target task based on the target page element and the action information includes: generating an agent for executing the target task based on the agent description information, the target page element, and the action information.

[0006] In at least one embodiment of the present disclosure, generating an agent for executing the target task includes: generating an agent for executing the target task that conforms to the model context protocol.

[0007] In at least one embodiment of the present disclosure, determining a target page element among the page elements included in the interaction page that the operation action is directed to includes: determining a page element among the page elements included in the interaction page whose attribute information matches the action information to obtain the target page element.

[0008] In at least one embodiment of the present disclosure, the interactive page presented during the execution of the target task includes a plurality of sub - pages presented in sequence; the action information includes the sub - action information of a plurality of sub - actions included in the operation action, and each sub - action corresponds to one of the plurality of sub - pages; determining the page elements in the interactive page whose attribute information matches the action information to obtain the target page elements includes: determining the page elements in each sub - page whose attribute information matches the sub - action information of the corresponding sub - page of each sub - page, and obtaining a plurality of target page elements corresponding to the plurality of sub - actions respectively.

[0009] In at least one embodiment of the present disclosure, the sub - action information includes the execution time information of the sub - action; wherein, based on the target page elements and the action information, generating an agent for executing the target task includes: determining the operation order of the plurality of target page elements based on the execution time information of the plurality of sub - actions; and generating an agent for executing the target task based on the operation order, the plurality of target page elements, and the sub - action information of the plurality of sub - actions.

[0010] In at least one embodiment of the present disclosure, at least one of the plurality of sub - pages is a target sub - page, and the target sub - page includes at least two target page elements corresponding to at least two sub - actions respectively.

[0011] In at least one embodiment of the present disclosure, obtaining the action information of the operation action for the interactive page includes: obtaining the input event log for the interactive page; and determining the operation action for the interactive page and the action information of the operation action based on the input event log, wherein the attribute information of the page element includes the first position information of the location where the page element is located; the action information includes the second position information of the position corresponding to the operation action; wherein, determining the page elements in the interactive page whose attribute information matches the action information to obtain the target page elements includes: determining the page elements in the interactive page whose first position information matches the second position information, and obtaining the target page elements.

[0012] In at least one embodiment of the present disclosure, obtaining the action information of the operation action for the interactive page includes: in response to obtaining voice information describing the operation action during the process of obtaining the interactive page, recognizing the voice information to obtain the action information of the operation action, wherein the attribute information of the page element includes the first element name of the page element; the action information includes the second element name of the target page element targeted by the operation action; wherein, determining the page elements in the interactive page whose attribute information matches the action information to obtain the target page elements includes: determining the page elements in the interactive page whose first element name matches the second element name, and obtaining the target page elements.

[0013] In at least one embodiment of the present disclosure, obtaining an interaction page displayed during the execution of a target task includes: in response to a start operation for the target task, monitoring and obtaining the displayed page; in response to a pause operation for the target task, stopping the monitoring and obtaining of the displayed page; and in response to a completion operation for the target task, using the obtained page as the interaction page displayed during the execution of the target task.

[0014] At least one embodiment of the present disclosure provides a data processing method, including: in response to receiving prompt information describing a task to be executed, processing the prompt information using a large language model to determine a target agent in an agent library corresponding to the task to be executed; and invoking the target agent to execute the task to be executed, where the agent library includes at least one specified agent, and the specified agent is generated based on the agent generation method provided in at least one embodiment of the present disclosure.

[0015] In at least one embodiment of the present disclosure, an agent in the agent library has corresponding agent description information; processing the prompt information using a large language model to determine a target agent in the agent library corresponding to the task to be executed includes: inputting the agent description information corresponding to each agent in the agent library and the prompt information into the large language model, and determining the target agent and the invocation information of the target agent based on the output information of the large language model.

[0016] In at least one embodiment of the present disclosure, the task to be executed involves at least one variable; the invocation information includes at least one variable value corresponding to the at least one variable.

[0017] In at least one embodiment of the present disclosure, the target agent includes at least two agents; determining the target agent and the invocation information of the target agent based on the output information of the large language model includes: determining at least two subtasks included in the task to be executed, at least two agents corresponding to the at least two subtasks respectively, and the invocation information of each agent in the at least two agents based on the output information of the large language model.

[0018] At least one embodiment of the present disclosure provides an agent generation device, including: an information acquisition module configured to acquire an interaction page displayed during the execution of a target task and action information of an operation action for the interaction page; a page recognition module configured to recognize the interaction page to obtain attribute information of page elements included in the interaction page; a target determination module configured to determine a target page element in the page elements included in the interaction page that the operation action is directed to based on the attribute information of the page elements and the action information; and a generation module configured to generate an agent for executing the target task based on the target page element and the action information.

[0019] At least one embodiment of the present disclosure provides a data processing device, including: an agent determination module configured to process the prompt information using a large language model in response to receiving the prompt information describing the task to be executed, and determine a target agent corresponding to the task to be executed in the agent library; and an agent invocation module configured to invoke the target agent to execute the task to be executed, wherein the agent library includes at least one designated agent, and the designated agent is generated based on the agent generation method provided by at least one embodiment of the present disclosure.

[0020] At least one embodiment of the present disclosure provides an electronic device, including: a processing device; and a storage device including one or more computer program instructions; wherein, when the one or more computer program instructions are run by the processing device, they execute the agent generation method or data processing method provided by at least one embodiment of the present disclosure.

[0021] At least one embodiment of the present disclosure provides a computer-readable storage medium that non-temporarily stores computer-readable instructions, wherein when the computer-readable instructions are executed by a processor, they implement the agent generation method or data processing method provided by at least one embodiment of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.

[0023] Figure 1 Schematically shows an application scenario diagram of the agent generation method and data processing method of at least one embodiment of the present disclosure;

[0024] Figure 2 Schematically shows a flowchart of the agent generation method of at least one embodiment of the present disclosure;

[0025] Figure 3 Schematically shows a schematic diagram of the principle of obtaining an interaction page of at least one embodiment of the present disclosure;

[0026] Figure 4 Schematically shows a schematic diagram of the principle of generating an agent of at least one embodiment of the present disclosure;

[0027] Figure 5 Schematically shows a schematic diagram of the principle of generating an agent of at least another embodiment of the present disclosure;

[0028] Figure 6 Schematically shows a schematic diagram of the principle of generating an agent of at least another embodiment of the present disclosure;

[0029] Figure 7 Schematic flow diagram showing the data processing method according to at least one embodiment of the present disclosure;

[0030] Figure 8 Schematic diagram showing the principle of performing a task to be executed based on prompt information according to at least one embodiment of the present disclosure;

[0031] Figure 9 Schematic communication diagram showing the execution of a task to be executed based on prompt information according to at least one embodiment of the present disclosure;

[0032] Figure 10 Schematic block diagram showing the generating device of an agent according to at least one embodiment of the present disclosure;

[0033] Figure 11 Schematic block diagram showing the data processing device according to at least one embodiment of the present disclosure;

[0034] Figure 12 Schematic block diagram showing the electronic device according to at least one embodiment of the present disclosure; and

[0035] Figure 13 Schematic diagram showing the computer-readable storage medium according to at least one embodiment of the present disclosure. Detailed implementation manners

[0036] To make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are some but not all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0037] Unless otherwise defined, technical terms or scientific terms used in this disclosure shall have the ordinary meanings as understood by those of ordinary skill in the art to which this disclosure pertains. The "first", "second" and similar terms used in this disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Similarly, words such as "a", "an" or "the" do not denote a limitation of quantity, but mean that there is at least one. Words such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Upper", "lower", "left", "right", etc. are only used to indicate relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0038] An agent, as an important part of the field of AI, is an entity that can perceive the environment, make autonomous decisions and execute tasks, and is widely used in fields such as automated task execution and human-computer interaction. With the progress of deep learning and natural language processing technologies, the application ability of agents in complex tasks has been significantly improved, but they also face problems such as high development costs and insufficient adaptability.

[0039] For example, it is possible to rely on manual coding or semi-automated data annotation processes to generate agents. However, this generation method requires the use of a large amount of labeled data to train and optimize the agent model, and the workload of manual coding and data annotation is large, which greatly limits the development efficiency of agents.

[0040] To at least partially solve the above technical problems, at least one embodiment of this disclosure provides a method for generating an agent, the method including: obtaining an interaction page displayed during the execution of a target task and action information of operation actions for the interaction page; identifying the interaction page to obtain attribute information of page elements included in the interaction page; based on the attribute information and action information of the page elements, determining target page elements in the page elements included in the interaction page that the operation actions are directed to; and generating an agent for executing the target task based on the target page elements and the action information. This method for generating an agent can automatically generate an agent by combining the identification of the interaction page and the action information of the operation actions, without manual coding and without data annotation, reducing the dependence on data in the agent generation process, and can greatly reduce the workload of manual coding and debugging, reducing the development cost of the agent, and can quickly build and deploy the agent, significantly improving the development efficiency.

[0041] Based on the generation method of the agent, at least another embodiment of the present disclosure also provides a data processing method, which includes: in response to receiving prompt information describing a task to be executed, using a large language model to process the prompt information to determine a target agent corresponding to the task to be executed in an agent library; and invoking the target agent to execute the task to be executed, where the agent library includes at least one specified agent, and the specified agent is generated based on the generation method of the agent provided in at least one embodiment of the present disclosure. This data processing method can automatically execute a target task by combining an agent and a large model, enabling the automatically generated agent to cooperate with the large model through a question-and-answer mode, reducing the need for manual intervention, and improving the automation level and intelligence degree of executing the target task.

[0042] Correspondingly, at least one embodiment of the present disclosure also provides an agent generation device, a data processing device, a processor, and a computer-readable storage medium to reduce the development cost of the agent and achieve automatic execution of tasks.

[0043] Figure 1 Schematically shows an application scenario diagram of the agent generation method and the data processing method of at least one embodiment of the present disclosure.

[0044] As Figure 1 shown, the application scenario 100 may include, for example, a user 101 and a terminal device 102. The terminal device 102 may be an electronic device such as a smart phone, a tablet computer, a portable laptop computer, or a desktop computer.

[0045] For example, the terminal device 102 may be installed with various client applications, such as instant messaging applications, music playback applications, video playback applications, or shopping applications. The terminal device 102 may provide an interaction page, which may be a page of a client application or an interaction page provided by the operating system of the terminal device 102. For example, the terminal device 102 may, in response to a user operation, record the operation actions of the user on one or at least two interaction pages during the process of using one or at least two client applications to complete a target task, identify the page elements in the interaction page, determine the target page elements targeted by the operation actions based on the operation actions and the page elements in the interaction page, thereby record the operation record of the user on the terminal device when completing the target task, and generate an agent based on the operation record. In this way, when the target task needs to be executed again, for example, the target task can be automatically executed by invoking the agent without user intervention.

[0046] For example, the target task may include any of the following tasks: a task of querying information, a task of sending information, a task of outputting information, etc. The information may include at least one of the following information: text information, picture information, video information, voice information, etc. Embodiments of the present disclosure are not limited thereto.

[0047] For example, based on a browser application or an agent generation application installed in the terminal device 102, an operation record of the user on the terminal device when completing the target task can be implemented, and the function of the agent can be generated based on the operation record.

[0048] In one embodiment, the application scenario 100 may further include a server 103. The server 103 may be, for example, a background management server that provides support for the operation of a client application running in the terminal device 102. For example, the server 103 may be a background management server that provides support for the operation of an agent generation application and / or an artificial intelligence application.

[0049] For example, the generated agent can be stored in the agent library 104. The terminal device 102 can, for example, provide an artificial intelligence interaction page for the user. The user 101 or any other user can put forward an assistance requirement through the artificial intelligence interaction page, expecting the terminal device 102 to provide assistance to the user. For example, task information for executing the target task can be input through the artificial intelligence interaction page. The terminal device 102 can, for example, send the task information to the background management server 103, and the background management server 103 uses a large model to execute the target task. For example, the target task can be directly executed by the large model, or the large model can determine the agent to be called for executing the target task and call the agent in the agent library 104 to complete the target task.

[0050] For example, the terminal device providing the artificial intelligence interaction page can also be any other terminal device different from the terminal device 102. When the computing power permits, the terminal device 102 can also use the large model to execute the target task. Embodiments of the present disclosure are not limited thereto.

[0051] The following will be combined with Figures 2 to 6 A method for generating an agent provided in at least one embodiment of the present disclosure will be described in detail.

[0052] Figure 2 A schematic flow chart of a method for generating an agent according to at least one embodiment of the present disclosure is schematically shown.

[0053] As Figure 2 shown, the method 200 for generating an agent in this embodiment may include step S210 to step S210 to step S240.

[0054] Step S210, obtain the interactive page displayed during the execution of the target task and the action information of the operation actions on the interactive page.

[0055] Step S220, identify the interactive page to obtain the attribute information of the page elements included in the interactive page.

[0056] Step S230, based on the attribute information of the page elements and the action information, determine the target page element among the page elements included in the interactive page that the operation action targets.

[0057] Step S240, generate an agent for executing the target task based on the target page element and the action information.

[0058] In at least one embodiment of the present disclosure, the target task may, for example, include any task performed by a user when using a terminal device, such as a task of controlling the terminal device to play song A, a task of controlling the terminal device to send a message to friend B in an instant messaging application, etc. For example, during the execution of the target task, the operation actions may, for example, at least include the task of opening the target client application.

[0059] For example, the obtained interactive page is the interactive page displayed by the terminal device during the execution of the target task. This interactive page may be a page that has been operated on during the execution of the target task. For example, if the target task is to open the target client application, the interactive page may include the page where the shortcut icon of the target client application is located, and the interactive page may not include the opening page of the target client application.

[0060] For example, in response to the start of the execution of the target task, monitor the interactive page displayed by the terminal device, and in response to the completion of the target task, terminate the monitoring of the interactive page. Use the monitored page as the interactive page displayed during the execution of the obtained target task. Alternatively, the page uploaded by the user may be used as the interactive page displayed during the execution of the obtained target task. The present disclosure embodiments do not limit the acquisition method of the interactive page.

[0061] In at least one embodiment of the present disclosure, the action information of the operation actions on the interactive page may be obtained by acquiring the input events of the screen of the terminal device. This input event may be obtained, for example, in response to an operation on an input device. The input device may, for example, include a keyboard, a mouse, a touch screen, a stylus, etc., and the present disclosure embodiments do not limit this. For example, the input event may be monitored through an event listener, obtained from an operation log, or obtained by any existing method, and the present disclosure embodiments do not limit this.

[0062] For example, the action information of an operation action may include the position information of the operation action in the display screen of the terminal device, the type of the operation action, the duration of the operation action, the execution time information of the operation action, etc. When the operation action is an action of inputting information through a keyboard, the action information of the operation action may further include, for example, the key code value of the key or the name of the key, etc.

[0063] In at least one embodiment of the present disclosure, a positioning technology based on image recognition may be used to analyze an interaction page, identify the area where a page element is located therefrom, and use the position information of the area where the identified page element is located as the attribute information of the page element. For example, methods such as optical character recognition, template area, or deep learning may be used to identify the area where the page element is located.

[0064] In at least one embodiment of the present disclosure, a deep learning algorithm may be used to recognize an interaction page to recognize various page elements included in the interaction page and generate positioning information. The category and positioning information of the page element are used as the attribute information of the page element. For example, the deep learning algorithm may include a target detection algorithm or an image classification algorithm based on deep learning, etc.

[0065] For example, a vision large model or a multimodal large model may be used to process the interaction page to obtain the attribute information of the page element.

[0066] For example, the attribute information of the page element may include the position of the page element in the interaction page, the type of the page element, the name of the page element, the trigger type of the page element, etc. For example, the page element refers to a user interface (UI) element, and the page element may include, for example, graphical interface elements such as buttons, text boxes, menus, check boxes, drop-down lists, sliders, etc.

[0067] In at least one embodiment of the present disclosure, the page element at the position of the operation action may be determined according to the position information of the operation action in the display screen of the terminal device in the action information and the position of the page element in the interaction page in the attribute information of the page element, and used as the target page element targeted by the operation action. Alternatively, the page element whose trigger type matches the action type may be determined according to the action type of the operation action in the action information and the trigger type of the page element, and used as the target page element targeted by the operation action. It can be understood that the above principle of determining the target page element is only an example for the convenience of understanding the present disclosure, and the embodiments of the present disclosure do not limit this.

[0068] For example, by determining the target page element for an operation action, the operation action can be associated with the target page element. Based on this association relationship, an agent for executing a target task can be generated in this embodiment. For example, an agent can be generated based on the target page element, action information, and service interface. The service interface can be, for example, an interface that provides a virtual operation execution service. For example, when invoking the agent, the interface of the virtual operation execution service can be invoked according to the service interface, and the action information and the information indicating the target page element (such as the name of the target page element) are used as inputs to the interface, so that the terminal device executes a virtual operation by invoking the interface of the virtual operation execution service to simulate a user operation, thereby executing the target task.

[0069] In the technical solution of at least one embodiment of the present disclosure, by identifying the attribute information of page elements from the interaction page displayed during the execution of the target task and combining the action information of the operation actions on the interaction page, an agent is generated without manual coding and without data annotation, reducing the dependence on data in the agent generation process, greatly reducing the workload of manual coding and debugging, lowering the development cost of the agent, enabling rapid construction and deployment of the agent, and significantly improving the development efficiency.

[0070] Figure 3 Schematically shows a schematic diagram of the principle of obtaining an interaction page in at least one embodiment of the present disclosure.

[0071] In an embodiment of the present disclosure, the terminal device may be installed with a client application for generating an agent, for example. The method for generating an agent provided in at least one embodiment of the present disclosure can be implemented based on the client application.

[0072] For example, the client application for generating an agent can provide an interaction page 310. A control 311 is displayed on the interaction page 310, and the control name of the control 311 can be, for example, "+ Create Agent". In response to an operation on the control 311, an identifier 312 of the agent to be created is displayed below the control 311. The initial name of the agent to be created can be a default name, for example, "XXXX". At the same time, a settings sub-page can be displayed in any area on the interaction page 310. The settings sub-page can have, for example, an input box 313 for annotating the name of the agent. In response to entering information in the input box 313, the name of the agent to be created can be modified from the initial name to the information entered in the input box.

[0073] For example, a recording button 314 can be displayed on the sub-page. The client application for generating the agent can use the trigger operation on the recording button 314 as the start operation for the target task, and start monitoring and obtaining the page displayed on the terminal device. After the recording button 314 is triggered, for example, the trigger operation on the recording button 314 again can be used as the completion operation for the target task, and the obtained page can be used as the interaction page displayed during the execution of the target task.

[0074] In one embodiment, in response to the trigger operation on the recording button 314, the interaction page 310 can be updated to the interaction page 320. In this interaction page 320, the form of the recording button is updated, so that the recording button can be equivalent to the termination button 324. For example, the trigger operation on the termination button 324 can be used as the completion operation for the target task.

[0075] In one embodiment, a pause button 325 can also be displayed on the interaction page 320, for example. During the process of monitoring and obtaining the displayed page, the trigger operation on the pause button 325 can be used as the pause operation for the target task. The client application for generating the agent can respond to this pause operation and stop monitoring and obtaining the displayed page. In this way, when the user needs to interrupt the execution of the target task, the pause button 325 can be triggered to avoid the interaction page outside the execution process of the target task being included in the obtained interaction page, thereby ensuring the accuracy of the finally generated agent.

[0076] It can be understood that Figure 3 The layout of the interaction page shown and the quantity and type of each control and button in the interaction page are only examples for the convenience of understanding the present disclosure, and the embodiments of the present disclosure do not limit this.

[0077] In at least one embodiment of the present disclosure, when determining the target page element, the attribute information of the page element included in the interaction page can be matched with the action information of the operation action, and the page element whose attribute information matches the action information is determined as the target page element.

[0078] For example, the attribute information and action information of each page element can be converted into feature vectors, and then the similarity between the feature vector obtained by converting the attribute information of each page element and the feature vector obtained by converting the action information is determined, and the page element corresponding to the highest similarity is used as the target page element.

[0079] For example, the attribute information and action information of each page element can also be input into a neural network model (such as a two-tower model, etc.) in the form of an information pair. The information output by the neural network model is used as the similarity for this information pair. The page element corresponding to the attribute information in the information pair with the highest similarity is used as the target page element.

[0080] Figure 4 Schematically shows a schematic diagram of the principle of generating an intelligent agent according to at least one embodiment of the present disclosure.

[0081] In at least one embodiment of the present disclosure, as Figure 4 shown, when obtaining the action information of the operation action, the input event log 401 recording the input events for the interaction page during the execution of the target task can be obtained first. For example, for the Android system, the log generated during the process of converting the raw input events obtained from various input devices by the Input system into KeyEvent objects and MotionEvent objects can be used as the input event log. For example, the log obtained by monitoring the screen input events can also be used as the input event log.

[0082] For example, after obtaining the input event log 401, the operation action for the interaction page and the action information 402 of the operation action can be determined based on the input event log. For example, the action information can be obtained by extracting keywords from the input event log or converting the input event log into structured information. For example, if the extracted keywords or the converted structured information are in the form of key-value pairs, the information of the value in the key-value pair can be used as the action information. For example, a pre-trained text processing model can also be used to process the input event log, and the text processing model outputs the indication information of the operation action and the action information. It can be understood that the above method for determining the action information 402 of the operation action based on the input event log is only an example for the convenience of understanding the present disclosure, and the embodiments of the present disclosure do not limit this.

[0083] For example, in one embodiment, a multimodal large model 410 can be used to identify and process the obtained interaction page 403, so as to obtain the attribute information 404 of the page elements.

[0084] For example, in one embodiment, the attribute information of the page elements may include the first position information of the location where the page elements are located, and the action information may include the second position information of the position corresponding to the operation action. For example, both the first position information and the second position information represent position coordinates in the interaction page. In this embodiment, when matching the action information 402 of the operation action with the attribute information 404 of the page elements, the first position information and the second position information can be matched. The page element whose first position information and second position information in the page elements included in the interaction page match is used as the target page element 405. The matching of the position information can be understood, for example, as the positions indicated by the two position information overlapping.

[0085] After obtaining the target page element 405, an agent 406 for executing the target task can be generated based on the target page element 405 and the action information 402 of the operation action.

[0086] Figure 5 Schematically shows a schematic diagram of the principle of generating an agent according to at least another embodiment of the present disclosure.

[0087] In at least one embodiment of the present disclosure, the client application for generating an agent can, for example, be provided with a function of inputting voice information. For example, during the process in which a user executes a touch operation on an interactive page displayed by a terminal device to execute a target task, the user can synchronously input voice information to describe the touch operation being executed in real time. For example, as Figure 5 shown, when determining the action information of the operation action, in response to obtaining voice information 501 describing the operation action during the process of obtaining the interactive page, the voice information 501 can be recognized to obtain the action information 502 of the operation action.

[0088] For example, the voice information 501 can first be converted into text information, and then semantic recognition and extraction of key information are performed on the text information, and the extracted key information is used as the action information 502 of the operation action. For example, sequence annotation can also be performed on the text information, and the annotated position information, the name of the page element, etc. are used as the action information 502 of the operation action.

[0089] For example, in one embodiment, a multimodal large model 510 can be used to perform recognition processing on the obtained interactive page 503 to obtain the attribute information 504 of the page element.

[0090] For example, in one embodiment, the attribute information of the page element can include the element name of the page element. The action information of the operation action can include the element name of the target page element targeted by the operation action. When matching the action information 502 of the operation action with the attribute information 504 of the page element in this embodiment, the element name of the page element can be matched with the element name included in the action information. The page element whose element name in the page element matches the element name included in the action information is used as the target page element 505. The matching of the element names can, for example, be understood as the two element names being exactly the same, and according to actual requirements, it can also be understood as the similarity between the two element names being greater than a predetermined similarity threshold.

[0091] After obtaining the target page element 505, an agent 506 for executing the target task can be generated based on the target page element 505 and the action information 502 of the operation action.

[0092] Through the technical solution of the embodiments of the present disclosure, the action information of the operation action can be determined in combination with the voice information provided by the user, so as to determine the target page element, without listening to the input event. In this way, it is beneficial to meet the requirement of generating an intelligent agent based on the pre-intercepted interactive page uploaded.

[0093] Figure 6 Schematically shows a schematic diagram of the principle of generating an intelligent agent according to at least another embodiment of the present disclosure.

[0094] In at least one embodiment of the present disclosure, the interactive page displayed during the execution of the target task may include a plurality of sub-pages displayed in sequence, that is, the obtained interactive page includes a plurality of sub-pages. Correspondingly, the operation action may include the operation actions for each of the plurality of sub-pages. For example, the action information of the obtained operation action may include the sub-action information of the plurality of sub-actions included in the operation action. There are, for example, a plurality of sub-action information, and the plurality of sub-action information corresponds to the plurality of sub-actions one by one. It can be understood that each of the plurality of sub-actions corresponds to a sub-page. For example, each of the sub-actions is an operation action generated by an input operation for a sub-page.

[0095] For example, the sub-action information of each sub-action may include the execution time information of the sub-action. The correspondence between the sub-action and the interactive page can be determined according to the display time of the interactive page and the execution time of the sub-action.

[0096] In at least one embodiment of the present disclosure, when matching the attribute information of the page element with the action information, it is possible to determine the page element among the page elements included in each sub-page whose attribute information matches the sub-action information of the sub-action corresponding to the sub-page, so as to obtain a plurality of sub-page elements corresponding to the plurality of sub-actions respectively.

[0097] For example, as Figure 6 shown, when matching the attribute information of the page element with the action information, for each sub-action 601, the sub-action information 602 of the sub-action 601 can be matched with the attribute information 604 of each page element in the sub-page 603 corresponding to the sub-action 601 to obtain the target page element 605 corresponding to the sub-action 601. Through this embodiment, a plurality of target page elements corresponding to the plurality of sub-actions respectively can be obtained.

[0098] For example, after obtaining a plurality of target page elements, the plurality of target page elements may be sorted according to the acquisition order of the interaction pages to which the plurality of target page elements belong, so as to obtain the operation order of the plurality of target page elements. Alternatively, in the case where the sub-action information of each sub-action includes the execution time information of each sub-action, the operation order of the plurality of target page elements may also be determined based on the execution time information of the plurality of sub-actions 606. For example, the execution order of the plurality of sub-actions may be determined according to the execution time information, and the execution order may be used as the operation order 606 of the plurality of target page elements respectively corresponding to the plurality of sub-actions.

[0099] After obtaining the operation order 606, for example, an agent 607 for executing the target task may be generated based on the operation order 606, the plurality of target page elements 605, and the action information of the plurality of sub-actions.

[0100] For example, in this embodiment, the sub-action information of each sub-action and the target page element corresponding to each sub-action may be combined into a <sub-action information, element> pair, and a plurality of <sub-action information, element> pairs are obtained in total. Subsequently, the plurality of <action, element> pairs are sorted based on the sequential operation 606 to obtain a <sub-action information, element> sequence. Finally, an agent may be generated based on the <sub-action information, element> sequence and the service interface. For example, when calling the agent, the interface of the service for executing virtual operations may be called according to the service interface, and the <sub-action information, element> sequence, etc. may be used as the input of the interface, so that the terminal device executes a plurality of virtual operations respectively corresponding to the plurality of sub-actions according to the operation order by calling the interface of the service for executing virtual operations, so as to simulate user operations and thus execute the target task.

[0101] In at least one embodiment of the present disclosure, the plurality of sub-pages obtained may include a target sub-page, for example. The target sub-page includes at least two target page elements respectively corresponding to at least two sub-actions. That is, the plurality of sub-pages obtained may include a sub-page in which the operations performed by the user are at least two. Alternatively, the obtained interaction page may also be a single page, and this page includes at least two target page elements respectively corresponding to at least two sub-actions. In this embodiment, for example, the operation order of the target page elements may be determined based on the execution time information of the sub-actions.

[0102] In at least one embodiment of the present disclosure, the client application for generating the agent may, for example, provide an annotation information input box in an interaction page Figure 3 similar to be used to input annotation information for the agent. For example, the input annotation information may include variable description information for describing the variables involved in the target task (i.e., variable description information of the variables involved when the agent executes the target task). According to actual requirements, the annotation information may also include the name of the agent, etc.

[0103] For example, in the case where annotation information is obtained in response to an input operation, an agent for performing a target task can be generated based on a target page element, action information, and the annotation information. For example, taking the target task of opening an instant messaging software to send a message to a friend as an example, the variables involved in the target task may include the friend's name, the content of the message to be sent, etc. When calling the generated agent to perform the target task, for example, the values of the friend's name and the content of the message to be sent can be obtained first, and the obtained values and these variables are formed into key-value pairs. Subsequently, these key-value pairs, the target page element, and the action information are all used as inputs to the service interface, so that the terminal device performs a virtual operation by calling the interface for executing the virtual operation service to simulate the user operation, thereby performing the target task.

[0104] For example, the variables described by the annotation information can be any variables that do not affect the position of the target page element. By obtaining the annotation information and generating an agent based on the annotation information in the embodiments of the present disclosure, the generalization ability of the generated agent can be improved. By obtaining the annotation information through the interactive page, it is convenient for the user to set parameters for the agent through intuitive operations.

[0105] In at least one embodiment of the present disclosure, the client application for generating an agent may provide a description information input box in an interactive page Figure 3 similar to be used to input agent description information for the agent. The agent description information can be used to describe at least one of the following information: the role type of the agent, the target task that the agent can perform, the image of the agent, etc. The agent description information can be any information that is conducive to distinguishing this agent from other agents, and the embodiments of the present disclosure do not limit this.

[0106] For example, in the case where the description information of the agent for performing the target task is obtained in response to an input operation, an agent for performing the target task can be generated based on the agent description information, the target page element, and the action information. For example, the agent description information can be used as a basis for determining whether to call the agent.

[0107] By obtaining the description information and generating an agent based on the agent description information in the embodiments of the present disclosure, the agent can be accurately called during subsequent use, which is conducive to improving the accuracy of performing the target task.

[0108] In at least one embodiment of the present disclosure, when generating an agent, the agent can be generated according to a standard protocol such as the Model Context Protocol (MCP), that is, an agent that meets the standard protocol such as the context protocol and executes the target task is generated. In this way, it is beneficial to realize the flexible application of the agent at the operating system level, so that the agent can be connected in series with the large language model or be called by the large language model.

[0109] Based on the method for generating an agent provided by at least one embodiment of the present disclosure, at least one embodiment of the present disclosure further provides a data processing method. The following will be combined with Figure 7 to describe this data processing method in detail.

[0110] Figure 7 The flowchart of the data processing method according to at least one embodiment of the present disclosure is schematically shown.

[0111] As Figure 7 shown, the data processing method 700 according to the embodiment of the present disclosure may include step S710 to step S720.

[0112] Step S710, in response to receiving a prompt message describing the task to be executed, use a large language model to process the prompt message and determine the target agent corresponding to the task to be executed in the agent library.

[0113] Step S720, call the target agent to execute the task to be executed.

[0114] In at least one embodiment of the present disclosure, the agent library may include multiple agents, and these agents may include, for example, the specified agents generated based on the method for generating an agent described above.

[0115] For example, the prompt message describing the task to be executed may include the description information of the task to be executed. For example, the prompt message may be "Please help me send the following information to XXX friends in XX: XXXXX". This prompt message can be used as a prompt word (prompt) and input into the large language model, and the large language model performs semantic understanding on this prompt message and determines whether to call an agent according to the result of the semantic understanding.

[0116] For example, the large language model may be a fine-tuned model. Through fine-tuning, the model has learned, for example, the tasks that each agent in the agent library can execute. After performing semantic understanding on the prompt message, it is determined whether the task to be executed belongs to the learned tasks. If so, the target agent in the agent library that can execute this task to be executed is determined.

[0117] For example, when determining the target agent or after determining the target agent, the large language model can also extract the variable values required when calling the agent from the prompt information. The variable values and the identification information of the agent are used as call information to call the target agent to execute the task to be executed. By extracting the variable values required when calling the agent, it is beneficial to the smooth execution of the task to be executed.

[0118] In at least one embodiment of the present disclosure, for example, the agent description information corresponding to each agent in the agent library and the prompt information can be used as prompt words and input into the large model. In this way, the large model can determine whether to call an agent based on the input agent description information corresponding to each agent, and when it is determined that an agent needs to be called, determine the target agent corresponding to the task to be executed.

[0119] For example, by processing the prompt information, the output information of the large language model can indicate the target agent to be called and the call information of the target agent. For example, the output information of the large language can include the identification information of the target agent, and the information that needs to be input or provided when calling the target agent. The information that needs to be input or provided can include, for example, the variable values of the variables involved in the task to be executed. The information that needs to be input or provided can be extracted by the large language model from the prompt information, for example.

[0120] Figure 8 Schematically shows the principle diagram of executing the task to be executed based on the prompt information in at least one embodiment of the present disclosure.

[0121] In at least one embodiment of the present disclosure, the task to be executed may be relatively complex and may be realized by calling at least two agents. For example, the large language model can split the task to be executed into at least two subtasks by performing semantic analysis on the prompt information and combining the description information of each agent in the intelligent question bank.

[0122] For example, as Figure 8 shown, after inputting the prompt information 801 and the description information 802 of the agent into the large language model 810, the output information 803 output by the large language model 810 can indicate at least two subtasks obtained by splitting, the agent corresponding to each subtask among the at least two subtasks, and the call information of the agent.

[0123] For example, the call information indicated by the output information 803 of the large language model 810 may only include, for example, the call information of the agent corresponding to the subtask that needs to be executed first among at least two subtasks that need to be executed sequentially. Or, after inputting the prompt information 801 and the description information 802 of the agent into the large language model 810, the output information 803 of the large language model may only include, for example, the identification information 804 of the agent corresponding to a subtask and the call information 805 of the agent. The identification information 804 of the agent may be, for example, the number of the agent in the agent library, etc., and the embodiments of the present disclosure do not limit this. When inputting the description information 802 of the agent, the description information of different agents may carry, for example, the identification information of the agent.

[0124] After obtaining the output information 803, for example, the agent 806 may be called first according to the identification information 804 and the call information 805 of the agent corresponding to the subtask that needs to be executed first, so as to execute the subtask that needs to be executed first. For example, the agent 806 may be selected from the agent library according to the identification information 804 of the agent first, and then the call information may be passed to the agent 806 so that the agent 806 executes the subtask.

[0125] Subsequently, the feedback information obtained by executing the subtask that needs to be executed first (such as the execution result 807 feedback by the agent 806), the prompt information 801, and the description information 802 of the agent may all be used as prompt words and input into the large language model 810. In this round of processing, the output of the large language model 810 may indicate the identification information of the agent corresponding to the next subtask to be executed and call the agent with the call information 805. And so on, until the large language model 810 determines that the execution of the task to be executed has been completed according to the execution result. For example, after inputting the execution result, the prompt information, and the description information of the agent that executes the last subtask into the large language model, the output information of the large language model is information indicating end, termination, or completion, etc.

[0126] In at least one embodiment of the present disclosure, when generating an agent, for example, the MCP protocol or other standard protocols may be used to encapsulate the agent, that is, an agent that conforms to the MCP protocol or other standard protocols is generated.

[0127] Figure 9 Schematically shows the communication principle diagram of executing the task to be executed based on the prompt information in at least one embodiment of the present disclosure.

[0128] As Figure 9 shown, taking the agent conforming to the MCP protocol as an example, in the process of processing the prompt information describing the task to be executed, the large language model communicates and collaborates with the agent to be called through the MCP protocol.

[0129] For example, as Figure 9 shown, the host is an application that runs a large language model. The server provides the function of features or data access function. The server can, for example, provide an MCP interface, which can, for example, include the interface for executing virtual operation services described above. The client acts as an intermediary between the host and the server, responsible for forwarding the requests of the host to the server and returning the responses of the server to the host.

[0130] For example, the MCP interface provided by the server can access the agents that conform to the MCP protocol generated by the system for generating agents in the embodiments of the present disclosure, or can also access the processing tools generated by other systems. The embodiments of the present disclosure do not make any limitations in this regard.

[0131] For example, the host can use the large language model to process the prompt information, determine the agent to be called, and based on the agent to be called, determine the server that has accessed the agent to be called. Subsequently, the call information can be sent to the determined server, and the server calls the accessed agent and executes the task to be executed. The server can also, for example, feedback the execution result of the task to be executed to the host.

[0132] In one embodiment, the prompt information and the description information of the agent can be sent to the host together, and the host processes the input information. The large language model analyzes the prompt information and decides whether to call an agent. If no call is needed, the large language model directly generates a natural language response. If a call is needed, the large language model can, for example, output a call request in a structured format. If the output information of the large language model includes a call request, the client will execute the call request to call the agent accessed by the server. After the server feeds back the call result, the call result, together with the prompt information and the description information of the agent, will be re-input into the large language model to determine whether the task to be executed is completed and whether other agents need to be called again.

[0133] In one embodiment, the server can also directly access the prompt words or data sources to call the accessed agents or tools based on the prompt words, or can also query data from the data sources.

[0134] Based on the above content, it can be known that at least one embodiment of the present disclosure provides a method for generating agents without relying on manually preparing a large-scale training data set, significantly reducing the costs of data collection and annotation and improving the development efficiency. Specifically, at least one embodiment of the present disclosure has the following beneficial technical effects:

[0135] Reduce data dependence. Instead of generating and annotating a large number of training data sets, the function information is directly extracted from the UI interface through a multi-modal model, and agents are generated by combining manual annotation and operation records, reducing the dependence on data and lowering the development cost.

[0136] Improve development efficiency. The function of automatically generating agents greatly reduces the workload of manual coding and debugging, enabling developers to quickly build and deploy agents and significantly enhancing development efficiency.

[0137] Enhance personalization and adaptability. Through manual annotation and operation records, agents can precisely adapt to the personalized needs and operation habits of different users, improving the user experience.

[0138] Support intelligent task execution. Agents work in collaboration with large language models through MCP or other standard protocols, can autonomously understand and execute complex tasks, and significantly enhance the automation level and intelligence of the system.

[0139] Reduce maintenance costs. Automatically generated agents can collaborate with large models through a question-and-answer mode, reducing the need for manual intervention and lowering the maintenance costs of the system.

[0140] Improve the scalability of the system. Automatically generated agents follow MCP and other standard protocols, can be seamlessly integrated with multiple systems and large language models, and enhance the scalability and flexibility of the system.

[0141] Based on the method for generating an agent provided in at least one embodiment of the present disclosure, at least one embodiment of the present disclosure further provides a device for generating an agent. The following will be combined with Figure 10 to describe this device in detail.

[0142] Figure 10 Schematically shows a schematic block diagram of the device for generating an agent according to at least one embodiment of the present disclosure.

[0143] As Figure 10 shown, the device 1000 for generating an agent in this embodiment may include an information acquisition module 1010, a page recognition module 1020, a target determination module 1030, and a generation module 1040.

[0144] The information acquisition module 1010 is configured to acquire the interaction page displayed during the execution of the target task and the action information of the operation actions for the interaction page. In at least one embodiment of the present disclosure, the information acquisition module 1010 may be configured to execute step S210 described above, which will not be elaborated here.

[0145] The page recognition module 1020 is configured to recognize the interaction page and obtain the attribute information of the page elements included in the interaction page. In at least one embodiment of the present disclosure, the page recognition module 1020 may be configured to execute step S220 described above, which will not be elaborated here.

[0146] The target determination module 1030 is configured to determine, based on the attribute information and action information of the page elements, the target page elements for the operation actions among the page elements included in the interactive page. In at least one embodiment of the present disclosure, the target determination module 1030 may be configured to perform the above-described step S230, which will not be elaborated herein.

[0147] The generation module 1040 is configured to generate an agent for executing the target task based on the target page elements and the action information. In at least one embodiment of the present disclosure, the generation module 1040 may be configured to perform the above-described step S240, which will not be elaborated herein.

[0148] In at least one embodiment of the present disclosure, the above-described agent generation device 1000 may further include a first information acquisition module, which is configured to acquire annotation information, and the annotation information includes variable description information for describing the variables involved in the target task. For example, the above-described generation module 1040 may be configured to: generate an agent for executing the target task based on the target page elements, the action information, and the annotation information.

[0149] In at least one embodiment of the present disclosure, the above-described agent generation device 1000 may further include a second information acquisition module, which is configured to acquire agent description information of the agent for executing the target task. For example, the above-described generation module 1040 may be configured to: generate an agent for executing the target task based on the agent description information, the target page elements, and the action information.

[0150] In at least one embodiment of the present disclosure, the above-described generation module 1040 may specifically be configured to: generate an agent for executing the target task that conforms to the model context protocol.

[0151] In at least one embodiment of the present disclosure, the above-described target determination module 1030 may specifically be configured to: determine the page elements among the page elements included in the interactive page whose attribute information matches the action information, and obtain the target page elements.

[0152] In at least one embodiment of the present disclosure, the interactive page displayed during the execution of the target task includes a plurality of sub-pages displayed in sequence; the action information includes sub-action information of a plurality of sub-actions included in the operation action, and each sub-action corresponds to one of the plurality of sub-pages. The above-described target determination module 1030 may specifically be configured to: determine the page elements among the page elements included in each sub-page whose attribute information matches the sub-action information of the sub-action corresponding to each sub-page, and obtain a plurality of target page elements corresponding to the plurality of sub-actions respectively.

[0153] In at least one embodiment of the present disclosure, the sub-action information includes the execution time information of the sub-actions. The above-mentioned generation module 1040 may include, for example, an order determination sub-module and a generation sub-module. The order determination sub-module is configured to determine the operation order of multiple target page elements based on the execution time information of the multiple sub-actions. The generation sub-module is configured to generate an agent for executing the target task based on the operation order, the multiple target page elements, and the sub-action information of the multiple sub-actions.

[0154] In at least one embodiment of the present disclosure, at least one sub-page among the multiple sub-pages is a target sub-page, and the target sub-page includes at least two target page elements respectively corresponding to at least two sub-actions.

[0155] In at least one embodiment of the present disclosure, the above-mentioned information acquisition module 1010 may include a log acquisition sub-module and an information determination sub-module. The log acquisition sub-module is configured to acquire the input event log for the interactive page; the information determination sub-module is configured to determine the operation action for the interactive page and the action information of the operation action based on the input event log. For example, the attribute information of the page element includes the first position information of the location where the page element is located; the action information includes the second position information of the location corresponding to the operation action. For example, the above-mentioned target determination module 1030 may be specifically configured to: determine the page elements in the interactive page whose first position information matches the second position information, and obtain the target page elements.

[0156] In at least one embodiment of the present disclosure, the above-mentioned information acquisition module 1010 may be specifically configured to: in response to obtaining voice information describing the operation action during the process of obtaining the interactive page, recognize the voice information to obtain the action information of the operation action. For example, the attribute information of the page element includes the first element name of the page element; the action information includes the second element name of the target page element targeted by the operation action. For example, the above-mentioned target determination module 1030 may be specifically configured to: determine the page elements in the interactive page whose first element name matches the second element name, and obtain the target page elements.

[0157] In at least one embodiment of the present disclosure, the above-mentioned information acquisition module 1010 may include a page acquisition sub-module and a page determination sub-module. The page acquisition sub-module is configured to: in response to the start operation for the target task, monitor and acquire the displayed page; and in response to the pause operation for the target task, stop monitoring and acquiring the displayed page. The page determination sub-module is configured to, in response to the completion operation for the target task, use the acquired page as the interactive page displayed during the execution of the target task.

[0158] Based on the data processing method provided by at least one embodiment of the present disclosure, at least one embodiment of the present disclosure further provides a data processing apparatus, which will be described in detail below in conjunction with Figure 11 This will be described in detail below.

[0159] Figure 11 The schematic block diagram of the data processing apparatus according to at least one embodiment of the present disclosure is schematically shown.

[0160] As Figure 11 shown, the data processing apparatus 1100 of this embodiment may include an agent determination module 1110 and an agent invocation module 1120.

[0161] The agent determination module 1110 is configured to, in response to receiving a prompt message describing a task to be executed, process the prompt message using a large language model to determine a target agent corresponding to the task to be executed in the agent library. For example, the agent library includes at least one specified agent, and the specified agent is generated based on the agent generation method provided by at least one embodiment of the present disclosure. In at least one embodiment of the present disclosure, the agent determination module 1110 may be configured to execute step S710 described above, which will not be elaborated here.

[0162] The agent invocation module 1120 is configured to invoke the target agent to execute the task to be executed. In at least one embodiment of the present disclosure, the agent invocation module 1120 may be configured to execute step S720 described above, which will not be elaborated here.

[0163] In at least one embodiment of the present disclosure, the agents in the agent library have corresponding agent description information. Specifically, the agent determination module 1110 is configured to: input the agent description information corresponding to each agent in the agent library and the prompt message into the large language model, and determine the target agent and the invocation information of the target agent based on the output information of the large language model.

[0164] In at least one embodiment of the present disclosure, the task to be executed involves at least one variable; the invocation information includes at least one variable value corresponding to the at least one variable.

[0165] In at least one embodiment of the present disclosure, the target agent includes at least two agents. Specifically, the agent determination module 1110 is configured to: determine at least two subtasks included in the task to be executed, at least two agents corresponding to the at least two subtasks respectively, and the invocation information of each of the at least two agents based on the output information of the large language model.

[0166] For example, each unit included in the agent generation device or data processing device can be implemented by a hardware (such as a circuit) module, a software module, etc., which will not be elaborated here. For example, these units can be implemented by a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a field programmable gate array (FPGA), or other forms of processing units with data processing capabilities and / or instruction execution capabilities, along with corresponding computer instructions.

[0167] It should be noted that in the embodiments of the present disclosure, the agent generation device or data processing device may include more or fewer circuits or units, and the connection relationships between the respective circuits or units are not limited and can be determined according to actual needs. The specific composition manner of each circuit is not limited and can be composed of analog devices according to circuit principles, or composed of digital chips, or in other applicable manners.

[0168] At least one embodiment of the present disclosure further provides an electronic device, including: a processing device; and a storage device including one or more computer program instructions; for example, when the one or more computer program instructions are run by the processing device, they execute the agent generation method or data processing method provided in any embodiment of the present disclosure.

[0169] Figure 12 It is a schematic structural diagram of an electronic device provided in at least one embodiment of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 12 The illustrated electronic device is merely an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present disclosure.

[0170] For example, in some examples, the electronic device includes the agent generation device provided in any embodiment of the present disclosure (for example, Figure 12 the processing device 1201 and the input device 1206 shown) to generate an agent. For example, the input device 1206 acquires the interaction page displayed during the execution of the target task and the action information of the operation actions for the interaction page. The processing device 1201 identifies the interaction page to obtain the attribute information of the page elements included in the interaction page; based on the attribute information of the page elements and the action information, determines the target page element in the page elements included in the interaction page that the operation action is directed to; and based on the target page element and the action information, generates an agent for executing the target task.

[0171] For example, in some examples, the electronic device includes a data processing device provided in any embodiment of the present disclosure (e.g., Figure 12 the processing device 1201 and the communication device 1209 shown in FIG. 1) to perform a task to be executed. For example, in response to receiving a prompt message describing the task to be executed, the processing device 1201 processes the prompt message using a large language model to determine a target agent in the agent library corresponding to the task to be executed. The processing device 1201 can also call the target agent via the communication device 1209 to execute the task to be executed.

[0172] For example, as Figure 12 shown, in some examples, the electronic device 1200 includes a processing device (such as a central processing unit, a graphics processing unit, etc.) 1201, which can perform various appropriate actions and processes according to a program stored in the read-only memory (ROM) 1202 or a program loaded from the storage device 1208 into the random access memory (RAM) 1203. In the RAM 1203, various programs and data required for the operation of the computer system are also stored. The processing device 1201, the ROM 1202, and the RAM 1203 are connected to each other via a bus 1204. The input / output (I / O) interface 1205 is also connected to the bus 1204.

[0173] For example, the following components can be connected to the I / O interface 1205: an input device 1206 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1207 including, such as a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1208 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1209 including a network interface card such as a LAN card, a modem, etc. The communication device 1209 can allow the electronic device 1200 to communicate with other devices wirelessly or wiredly to exchange data and perform communication processing via a network such as the Internet. The drive 1210 is also connected to the I / O interface 1205 as needed. A removable medium 1211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1210 as needed so that a computer program read from it can be installed into the storage device 1208 as needed. Although Figure 12 the electronic device 1200 including various devices is shown, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices can be implemented or included.

[0174] For example, the electronic device 1200 may further include a peripheral interface (not shown in the figure), etc. The peripheral interface may be various types of interfaces, such as a USB interface, a Lightning interface, etc. The communication device 1209 may communicate with a network and other devices through wireless communication. The network may be, for example, the Internet, an intranet, and / or a wireless network such as a cellular phone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN). The wireless communication may use any one of a variety of communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wi-Fi (e.g., based on IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n standards), Voice over Internet Protocol (VoIP), WiMAX, protocols for email, instant messaging, and / or Short Message Service (SMS), or any other suitable communication protocol.

[0175] For example, the electronic device may be any device such as a mobile phone, a tablet computer, a laptop computer, an e-book, a game console, a television, a digital photo frame, a navigator, etc., or may be any combination of an electronic device and hardware. The embodiments of the present disclosure are not limited thereto.

[0176] For example, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium. The computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device 1209, or installed from the storage device 1208, or installed from the ROM 1202. When the computer program is executed by the processing device 1201, the method for generating the above-mentioned agent or the data processing method defined in the method of the embodiment of the present disclosure is executed.

[0177] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiments of the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the embodiments of the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0178] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with digital data communication in any form or medium (for example, a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (for example, the Internet), and end-to-end networks (for example, ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0179] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; it can also exist separately without being assembled into the electronic device.

[0180] The above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: obtain an interaction page displayed during the execution of a target task and action information of operation actions on the interaction page; identify the interaction page to obtain attribute information of page elements included in the interaction page; determine a target page element in the page elements included in the interaction page that the operation action targets based on the attribute information of the page elements and the action information; and generate an agent for executing the target task based on the target page element and the action information.

[0181] Alternatively, the above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: in response to receiving prompt information describing a task to be executed, process the prompt information using a large language model to determine a target agent corresponding to the task to be executed in an agent library; and call the target agent to execute the task to be executed, where the agent library includes at least one specified agent, and the specified agent is generated based on the agent generation method provided in any embodiment of the present disclosure.

[0182] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, execute as a stand-alone software package, execute partially on the user's computer and partially on a remote computer, or execute entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by connecting through an Internet service provider using the Internet).

[0183] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.

[0184] In various embodiments of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0185] At least one embodiment of the present disclosure further provides a storage medium. Figure 13 It is a schematic diagram of a storage medium provided by at least one embodiment of the present disclosure. For example, as Figure 13 shown, the storage medium 1300 non-temporarily stores computer-readable instructions 1301, and when the non-temporary computer-readable instructions are executed by a computer (including a processor), the generation method or data processing method of the agent provided by any embodiment of the present disclosure can be executed.

[0186] For example, the storage medium may be any combination of one or more computer-readable storage media. For example, one computer-readable storage medium contains computer-readable program code for the generation method of the agent, and another computer-readable storage medium contains computer-readable program code for the data processing method. For example, when the program code is read by a computer, the computer can execute the program code stored in the computer storage medium and execute, for example, the generation method or data processing method of the agent provided by any embodiment of the present disclosure.

[0187] For example, the storage medium may include a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), portable compact disk read-only memory (CD-ROM), flash memory, or any combination of the foregoing storage media, and may also be other applicable storage media.

[0188] In addition to the above exemplary description, the following points need to be noted:

[0189] (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures can refer to the general design.

[0190] (2)Where there is no conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other to obtain new embodiments.

[0191] As described above, the foregoing is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be subject to the protection scope of the claims.

Claims

1. A method for generating an intelligent agent, characterized in that: The method comprises: Acquire the interactive page displayed during the execution of the target task and the action information of the operation action on the interactive page; Identify the interactive page, and obtain attribute information of page elements included in the interactive page; Determining, based on the attribute information of the page element and the action information, a target page element targeted by the operation action among the page elements included in the interactive page; and Based on the target page elements and the action information, an agent that performs the target task is generated.

2. The method for generating an intelligent agent according to claim 1, characterized in that: The method further comprises: Acquire annotation information, wherein the annotation information includes variable description information describing the variables involved in the target task; Wherein, generating an agent for executing the target task based on the target page element and the action information includes: Based on the target page element, the action information and the annotation information, an agent that performs the target task is generated.

3. The method for generating an intelligent agent according to claim 1, characterized in that: The method further comprises: Obtaining agent description information of an agent that performs the target task; Wherein, generating an agent for executing the target task based on the target page element and the action information includes: Based on the agent description information, the target page elements and the action information, an agent that performs the target task is generated.

4. The method for generating an intelligent agent according to claim 1, characterized in that: Generating an agent to perform the target task, including: Generate an agent that performs the target task in accordance with the model context protocol.

5. The method for generating an intelligent agent according to claim 1, characterized in that: Determining a target page element targeted by the operation action among the page elements included in the interactive page includes: The page elements included in the interactive page are determined, and the page elements whose attribute information matches the action information are obtained, to obtain the target page elements.

6. The method for generating an intelligent agent according to claim 5, characterized in that: The interactive page displayed during the execution of the target task includes a plurality of sub-pages displayed in sequence; the action information includes sub-action information of a plurality of sub-actions included in the operation action, and each of the sub-actions corresponds to a sub-page among the plurality of sub-pages; The step of determining the page element whose attribute information matches the action information among the page elements included in the interactive page, and obtaining the target page element, comprises: The page elements included in each sub-page whose attribute information matches the sub-action information of the sub-action corresponding to each sub-page are determined, and a plurality of target page elements respectively corresponding to the plurality of sub-actions are obtained.

7. The method for generating an intelligent agent according to claim 6, characterized in that: The sub-action information includes execution time information of the sub-action; Wherein, generating an agent for executing the target task based on the target page element and the action information includes: Determining the operation sequence of the plurality of target page elements based on the execution time information of the plurality of sub-actions; and Based on the operation sequence, the plurality of target page elements and the sub-action information of the plurality of sub-actions, an intelligent agent for executing the target task is generated.

8. The method for generating an intelligent agent according to claim 6 or 7, characterized in that: At least one sub-page among the multiple sub-pages is a target sub-page, and the target sub-page includes at least two target page elements corresponding to at least two sub-actions respectively.

9. The method for generating an intelligent agent according to claim 5, characterized in that: Acquiring action information of an operation action on the interactive page, including: Obtaining an input event log for the interactive page; and Determine an operation action for the interactive page and action information of the operation action based on the input event log, The attribute information of the page element includes first position information of the location of the page element; the action information includes second position information of the location corresponding to the operation action; The step of determining the page element whose attribute information matches the action information among the page elements included in the interactive page to obtain the target page element comprises: Determine the page element whose first position information matches the second position information among the page elements included in the interactive page, and obtain the target page element.

10. The method for generating an intelligent agent according to claim 5, characterized in that: Acquiring action information of an operation action on the interactive page, including: In response to acquiring voice information describing the operation action in the process of acquiring the interactive page, recognizing the voice information and obtaining action information of the operation action, The attribute information of the page element includes the first element name of the page element; the action information includes the second element name of the target page element targeted by the operation action; The step of determining the page element whose attribute information matches the action information among the page elements included in the interactive page to obtain the target page element comprises: A page element whose first element name matches the second element name among page elements included in the interactive page is determined to obtain the target page element.

11. The method for generating an intelligent agent according to claim 1, characterized in that: Get the interactive page displayed during the execution of the target task, including: In response to a start operation for the target task, monitoring and acquiring a displayed page; In response to a pause operation on the target task, stopping monitoring and acquiring the displayed page; In response to a completion operation for the target task, the acquired page is used as an interactive page displayed during the execution of the target task.

12. A data processing method, characterized in that: The method comprises: In response to receiving prompt information describing a task to be performed, using a large language model to process the prompt information and determine a target agent in an agent library corresponding to the task to be performed; and Invoke the target agent to execute the task to be executed, The agent library includes at least one designated agent, and the designated agent is generated based on the agent generation method described in any one of claims 1 to 11.

13. The data processing method according to claim 12, characterized in that: The agents in the agent library have corresponding agent description information; The prompt information is processed using a large language model to determine a target agent in an agent library corresponding to the task to be performed, including: The agent description information corresponding to each agent in the agent library and the prompt information are input into the large language model, and the target agent and the calling information of the target agent are determined based on the output information of the large language model.

14. The data processing method according to claim 13, characterized in that: The task to be executed involves at least one variable; the calling information includes at least one variable value corresponding to the at least one variable.

15. The data processing method according to claim 13, characterized in that: The target agent includes at least two agents; Determining the target agent and the calling information of the target agent based on the output information of the large language model includes: Based on the output information of the large language model, at least two subtasks included in the task to be executed, the at least two agents corresponding to the at least two subtasks respectively, and the calling information of each of the at least two agents are determined.

16. A device for generating an intelligent agent, characterized in that: The device comprises: An information acquisition module is configured to acquire the interactive page displayed during the execution of the target task and action information of the operation action on the interactive page; A page identification module is configured to identify the interactive page and obtain attribute information of page elements included in the interactive page; a target determination module configured to determine, based on the attribute information of the page element and the action information, a target page element targeted by the operation action among the page elements included in the interactive page; and A generation module is configured to generate an intelligent agent that performs the target task based on the target page elements and the action information.

17. A data processing device, characterized in that: The device comprises: an agent determination module, configured to, in response to receiving prompt information describing a task to be performed, process the prompt information using a large language model to determine a target agent in an agent library corresponding to the task to be performed; and An agent calling module is configured to call the target agent to execute the task to be executed, The agent library includes at least one designated agent, and the designated agent is generated based on the agent generation method described in any one of claims 1 to 11.

18. An electronic device, characterized in that: The device comprises: processing equipment; and a storage device including one or more computer program instructions; The one or more computer program instructions are executed by the processing device to perform the method according to any one of claims 1 to 15.

19. A computer-readable storage medium, characterized in that: Non-temporarily storing computer-readable instructions, wherein when the computer-readable instructions are executed by a processor, the method of any one of claims 1 to 15 is implemented.

Citation Information

Patent Citations

  • Intelligent agent generation method and device, electronic equipment and storage medium

    CN117764107A

  • Interaction method and device based on large model, electronic equipment and storage medium

    CN118606590A

  • Large model driven Web task automatic execution method and system

    CN119248379A

  • Business processing method and device, equipment and storage medium

    CN119536854A

  • Task processing method and device, equipment, storage medium and product

    CN119850146A

Cited By

  • Generation method and system of task execution engine, electronic equipment and program product

    CN120687264A

  • Intelligent agent collaborative response and task execution method and system oriented to operation scene and storage medium

    CN122364581A