Interaction method, intelligent agent training method and device and electronic equipment
By receiving user instructions and combining user interface information to generate target actions, the problem of inefficient user interface interaction is solved, intelligent and automated interaction response is achieved, the consistency and satisfaction of the user experience is improved, and different interface designs are adapted to different interface designs.
Patent Information
- Application Number
- CN202510551689.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-08
AI Technical Summary
In the prior art, user interface interaction is inefficient, users are inefficient when reviewing user manuals or online help documents, and traditional agent training schemes cannot adapt to dynamic UI environments and new tasks or new interfaces, resulting in interaction failure.
By receiving user instructions, combining the page description information of the user interface to generate target actions, and executing actions on the user interface, multi-modal instruction input and flexible instruction acquisition methods are used to train the agent to generate actions that match the user's intentions, and realize intelligent and automated interactive responses.
It improves the speed and accuracy of user interface interaction, enhances the consistency and satisfaction of user experience, and can flexibly deal with user interfaces of different structures and designs, providing a smooth and natural operating experience.
Smart Images

Figure CN120447798A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to an interaction method, an intelligent agent training method, a device, and an electronic device. Background Art
[0002] User Interface (UI) has become the mainstream way for people to interact with various software, systems and devices.
[0003] However, with the continuous expansion of software functions and the increasing complexity of interface design, different software have different UI styles. Users often encounter various difficulties when interacting with the UI. When users encounter problems when using the UI, traditional solutions are often inefficient. For example, consulting the user manual or online help documents requires users to spend a lot of time searching for relevant information, and the document content may not be intuitive enough and difficult to understand, which not only affects the user experience but also reduces the interaction efficiency. Summary of the Invention
[0004] This application aims to solve one of the technical problems in the related art to a certain extent.
[0005] To this end, the present application proposes an interaction method, an intelligent agent training method, a device and an electronic device to realize the intelligence and automation of user interface interaction, significantly improving the speed and accuracy of responding to user commands, and at the same time being able to flexibly respond to user interfaces of different structures and designs (including complex multi-element interfaces or simple single-function interfaces), thereby enhancing the consistency and satisfaction of the user experience and making operations in various application scenarios smoother and more natural.
[0006] In one aspect, an embodiment of the present application provides an interaction method, including:
[0007] Receive user instructions;
[0008] generating a first target action associated with the first user interface according to the first page description information of the first user interface and the user instruction;
[0009] The first target action is executed on the first user interface to implement an interactive response to the user instruction.
[0010] Another embodiment of the present application provides a method for training an intelligent agent, including:
[0011] Obtaining training samples; wherein the training samples include sample instructions and sample page description information of a sample user interface;
[0012] Using an initial intelligent agent to generate a predicted action associated with the sample instruction based on the training sample;
[0013] The initial intelligent agent is trained according to the sample actions marked by the sample instructions and the predicted actions to obtain a trained intelligent agent.
[0014] Another embodiment of the present application provides an interactive device for implementing the interactive method described in the first embodiment, including:
[0015] A receiving module, used for receiving user instructions;
[0016] a generating module, configured to generate a first target action associated with the first user interface according to the first page description information of the first user interface and the user instruction;
[0017] An execution module is used to execute the first target action on the first user interface to achieve an interactive response to the user instruction.
[0018] Another embodiment of the present application provides a training device for an intelligent agent, which is used to implement the training method for an intelligent agent as described in another embodiment, including:
[0019] An acquisition module is used to acquire training samples; wherein the training samples include sample instructions and sample page description information of a sample user interface;
[0020] A generation module, configured to generate a predicted action associated with the sample instruction based on the training sample using an initial intelligent agent;
[0021] A training module is used to train the initial intelligent agent according to the sample actions marked by the sample instructions and the predicted actions to obtain a trained intelligent agent.
[0022] In another aspect, an embodiment of the present application provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; and the processor executes the computer-executable instructions stored in the memory to implement the method described in the above embodiment.
[0023] In another aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the method described in the above embodiment.
[0024] In another aspect, an embodiment of the present application provides a computer program product, which stores a computer program. When the program is executed by a processor, the method described in the above embodiment is implemented.
[0025] The interaction method, intelligent agent training method, device and electronic device proposed in this application receive user instructions and, based on the page description information of the first user interface and the user instructions, intelligently generate a first target action that precisely matches the first user interface, and then execute the first target action on the first user interface, thereby achieving efficient and accurate interactive response to user instructions. This mechanism not only promotes the development of user interface interaction towards intelligence and automation, significantly improves the speed and accuracy of responding to user instructions, but also can flexibly respond to user interfaces of different structures and designs (including complex multi-element interfaces or simple single-function interfaces), enhances the consistency and satisfaction of the user experience, and makes operations in various application scenarios smoother and more natural.
[0026] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0028] Figure 1 A flowchart of an interaction method provided in an embodiment of the present application;
[0029] Figure 2 A flowchart of another interactive method provided in an embodiment of the present application;
[0030] Figure 3 A flowchart of another interactive method provided in an embodiment of the present application;
[0031] Figure 4 A flowchart of a method for training an intelligent agent provided in an embodiment of the present application;
[0032] Figure 5 A flowchart of a method for training an intelligent agent provided in an embodiment of the present application;
[0033] Figure 6 A schematic diagram of the principles of an agent training method provided in an embodiment of the present application;
[0034] Figure 7 A schematic diagram of the structure of an interactive device provided in an embodiment of the present application;
[0035] Figure 8 A schematic diagram of the structure of an intelligent agent training device provided in an embodiment of the present application;
[0036] Figure 9 A block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0037] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0038] It should be noted that the acquisition, storage, use, and processing of data in the technical solution of this application comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0039] It should also be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.
[0040] Accessibility service is one of the important system services of mobile smart terminals. For example, when driving or taking care of children, users may need additional interactive feedback (such as voice, etc.) as a supplement. In related technologies, by training the intelligent agent for UI interaction, it is executed according to the following logic: 1) Page description: This step is used to enable the intelligent agent to understand the interface information of the current UI, describe the key UI elements, interactive components and page status on the page; 2) Action thinking (Action-Think): This step simulates the thinking process of humans before performing an operation, that is, analyzing the current interface status and inferring possible interaction strategies; 3) Action (Next Action): This step generates specific operations based on the results of the "Action Thinking" stage, such as clicking a button, entering text, etc.; 4) Action Result (Action Result) This step is used to record the page feedback after performing the operation to help the intelligent agent learn the cause and effect relationship of the operation.
[0041] While this approach can help agents reason about tasks according to a certain logic, its current agent training scheme is primarily based on a static Chain of Action Thought (CoAT) construct and trained using supervised learning. However, this approach suffers from significant generalization issues, specifically in the following aspects: 1) Predefined CoATs cannot adapt to dynamic UI environments: Existing approaches rely on a predefined fixed chain of thought (e.g., page description → action thought → specific action → action result). However, this static approach struggles to adapt to new environments when the UI structure or layout changes. For example, the location and style of the shopping cart button may vary across different e-commerce websites, and static CoATs may not correctly identify and adjust the action strategy, leading to execution failure. 2) Static CoATs cannot generalize to new tasks or new interfaces: Because existing approaches primarily use fixed task datasets for training, the model cannot infer appropriate interaction methods when encountering new UI tasks or interface layouts. For example, during training, an agent may learn to "click on the shopping cart to view items," but if it switches to a new application (app) or a different version of the interface, the style or location of the shopping cart icon may change, and the agent may not be able to adapt, resulting in task failure.
[0042] In response to the above problems, the present application proposes an interaction method, an intelligent agent training method, a device and an electronic device.
[0043] The following describes the interaction method, agent training method, device and electronic device of the embodiments of the present application with reference to the accompanying drawings.
[0044] Figure 1 A flowchart of an interaction method provided in an embodiment of the present application.
[0045] It should be noted that the interaction method of the embodiments of the present application can be applied to an interaction device. In some possible embodiments, the interaction device can be configured in an electronic device or chip so that the electronic device or chip can perform interaction functions. In addition, in some possible embodiments, the interaction device can also be software in an electronic device, etc.
[0046] In any embodiment of the present application, the chip can be integrated into an electronic device. Among them, the chip includes a central processing unit (CPU), an image signal processing (ISP), an application-specific integrated circuit (ASIC), a microprocessor (DSP), a field programmable gate array (FPGA), a system on a chip (SOC), a reduced instruction set computer (RISC), etc., which are not listed one by one here.
[0047] Among them, electronic devices include but are not limited to: terminals, personal computers, etc. A terminal is an entity on the user side for receiving or transmitting signals, such as a mobile phone. A terminal can also be called a terminal device (terminal), user equipment (UE), mobile station (MS), mobile terminal (MT), etc. A terminal can be a mobile phone with communication functions, a wearable device, a tablet computer (Pad), a computer with wireless transceiver functions, a virtual reality (VR) terminal, an augmented reality (AR) terminal, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical surgery, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, a wireless terminal in smart home, etc. The embodiments of this application do not limit the specific technology and specific device form adopted by the terminal.
[0048] like Figure 1 As shown, the interaction method may include the following steps:
[0049] Step 101: Receive user instructions.
[0050] In order to improve the accuracy of interaction with the user interface, in an embodiment of the present application, instructions issued by the user for interacting with the user interface are captured and identified to understand the user's needs.
[0051] In order to improve the convenience and efficiency of user interaction with the user interface, as a possible implementation method, multimodal command input methods and flexible command acquisition methods are used to meet the usage habits and scenario requirements of different users.
[0052] As an example, in response to a text input operation on a first interactive control in a first user interface, a user instruction is obtained from the first interactive control.
[0053] That is to say, in scenarios where users are accustomed to manual input or need to express their needs precisely, when the user inputs instructions through the first interactive control (such as an input box), the user's text content can be directly captured and processed as a user instruction.
[0054] As another example, in response to a voice input operation on a second interactive control in the first user interface, an input user instruction is obtained.
[0055] That is to say, in scenarios where manual operation is inconvenient, such as on mobile devices or while driving, users can issue commands in the form of voice through a second interactive control (such as a voice button), and then use voice recognition technology to convert the voice content into text commands.
[0056] As another example, in response to a selection operation, a user instruction is determined from at least one candidate instruction displayed in the first user interface.
[0057] That is to say, in a scenario where the user is not clear about the instruction description or wants to get help quickly, at least one candidate instruction (such as a list of common instructions) is displayed in the first user interface, and the user can determine the instruction through a simple click or selection operation without manual input or voice expression.
[0058] Step 102: Generate a first target action associated with the first user interface according to first page description information of the first user interface and a user instruction.
[0059] In order to accurately obtain the specific operation that needs to be performed in response to the user instruction, in the embodiment of the present application, after receiving the user instruction, the user intention can be analyzed and a specific target action (i.e., the first target action) can be generated by combining the page description information of the first user interface (i.e., the layout, functional modules, status, etc. of the current interface). The first user interface can be the UI page (such as the homepage or homepage) of any application (Application, abbreviated as APP).
[0060] Step 103: Execute a first target action on the first user interface to implement an interactive response to the user instruction.
[0061] As a possible implementation method, the intelligent agent is called to perform corresponding operations on the first user interface according to the generated first target action, thereby completing the response to the user instruction.
[0062] For example, the first target action is "jump to the settings page", and the agent performs a page jump on the first user interface and jumps to the settings page.
[0063] In summary, by receiving user instructions and intelligently generating a first target action that precisely matches the first user interface based on the page description information and user instructions of the first user interface, and then executing the first target action on the first user interface, efficient and accurate interactive response to user instructions is achieved. This mechanism not only promotes the development of user interface interaction towards intelligence and automation, significantly improves the speed and accuracy of responding to user instructions, but also can flexibly respond to user interfaces of different structures and designs (including complex multi-element interfaces or simple single-function interfaces), enhances the consistency and satisfaction of the user experience, and makes operations in various application scenarios smoother and more natural.
[0064] In order to clearly illustrate how the target action is executed on the first user interface in the above embodiment to achieve an interactive response to the user instruction, the present application proposes another interaction method.
[0065] Figure 2 A flowchart of another interaction method provided in an embodiment of the present application.
[0066] like Figure 2 As shown, the interaction method may include the following steps:
[0067] Step 201: Receive user instructions.
[0068] Step 202: Generate a first target action associated with the first user interface according to first page description information of the first user interface and a user instruction.
[0069] Step 203: Execute at least one round of page interaction process according to the first user interface and the first target action to achieve interactive response to the user instruction.
[0070] In the embodiment of the present application, the process of the first user interface and the first target action implementing an interactive response to the user instruction may involve one or more rounds of page interaction processes.
[0071] In an embodiment of the present application, the first round of interaction process in at least one round of page interaction process includes: executing a first target action on a first user interface to obtain and display a second user interface to which the first round of page interaction process in at least one round of page interaction process jumps; when it is determined according to the second user interface that the user instruction response is completed, the page interaction process is ended; when it is determined that the user instruction response is not completed, the next round of interaction process is continued.
[0072] That is to say, when the first target action (such as page jump) is executed on the first user interface, an interface jump occurs, jumping to the second user interface, and displaying the second user interface. Then, based on the status of the second user interface or the page content, it is determined whether the user instruction has been fully responded to. For example, the user instruction is "open the settings page". When the settings page is successfully loaded, the system can consider that the instruction has been completed and end the page interaction process.
[0073] In an embodiment of the present application, the non-first round interaction process in the multi-round page interaction process includes: when it is determined that the user instruction has not been responded to according to the second user interface to which the i-1th round of page interaction process jumps, the second target action of the i-th round of page interaction process is determined according to the second page description information and user instruction of the second user interface to which the i-1th round of page interaction process jumps; wherein i is a positive integer greater than 1; on the second user interface to which the i-1th round of page interaction process jumps, the second target action of the i-th round of page interaction process is executed to obtain the second user interface to which the i-th round of page interaction process jumps, and whether the user instruction has been responded to according to the second user interface to which the i-th round of page interaction process jumps, until it is determined that the user instruction has been responded to, the page interaction process is ended.
[0074] That is, if the second user interface, to which the page interaction process jumps in the (i-1) round, determines that the user instruction has not yet been fully responded to, the description of the second user interface and the original user instruction can be combined to determine the specific action to be performed in the i-th round of page interaction, namely the second target action. Here, i represents a positive integer greater than 1. Then, this second target action is performed on the second user interface to which the page interaction process jumps in the (i-1) round, resulting in a new interface after the i-th round of interaction. The system then checks this new interface again to determine whether the user instruction has been fully responded to. This process repeats until it is determined that the user instruction has been fully responded to, at which point the entire page interaction process ends.
[0075] For example, if the user's instruction is to "purchase a product," multiple rounds of interaction are involved, such as completing the following interactions in sequence: clicking "Add to Cart" on the product details page, jumping to the shopping cart page, clicking "Checkout," filling in payment information, and confirming the order. The first round of interaction includes clicking "Add to Cart" on the product details page, and non-first rounds of interaction include jumping to the shopping cart page, clicking "Checkout," filling in payment information, and confirming the order.
[0076] It should be noted that the execution process of steps 201 to 202 can be implemented in any of the embodiments of the present application. The embodiments of the present application do not limit this and will not be described in detail.
[0077] In summary, by executing at least one round of page interaction process according to the first user interface and the first target action, a comprehensive and dynamic response to user instructions is achieved. This mechanism can not only intelligently go through a series of orderly interface changes or operation steps according to user instructions until an interactive response to user instructions is achieved, but also this multi-round interaction method allows the system to handle more complex user interfaces, thereby improving the flexibility and adaptability of the system.
[0078] In order to clearly illustrate how the first target action associated with the user instruction is generated according to the first page description information of the first user interface and the user instruction in the above embodiment, the present disclosure proposes another interaction method.
[0079] Figure 3 A flowchart of another interaction method provided in an embodiment of the present application.
[0080] like Figure 3 As shown, the interaction method may include the following steps:
[0081] Step 301: Receive user instructions.
[0082] Step 302: Based on the first page description information, call the intelligent agent to interactively reason on the user instruction to obtain first interactive information.
[0083] The first interaction information is used to indicate a page interaction element in the first page that is used to respond to a user instruction.
[0084] In an embodiment of the present application, the first page description information is obtained by performing a page description on the first user interface. In order to improve the effectiveness and accuracy of the first page description information, as a possible implementation method, detailed description information of the first user interface is extracted and generated based on the target image and page description instructions of the first user interface.
[0085] As an example, a target image and a page description instruction showing a first user interface are obtained; wherein the page description instruction is used to prompt an intelligent agent to perform a page description task; the intelligent agent is used to perform a page description of the first user interface based on the target image and the page description instruction to obtain first page description information of the first user interface.
[0086] That is, a target image showing the first user interface and related page description instructions are obtained. The target image provides the visual content of the user interface, while the page description instructions are used to guide the agent to complete a specific page description task. The target image and page description instructions are then input into the agent, which then performs a comprehensive analysis of the first user interface based on the target image and page description instructions, extracting key information and generating structured first page description information. The first page description information may include static content such as the interface layout and control locations, as well as the interface's functional logic and interaction rules.
[0087] In an embodiment of the present application, the first page description information provides the detailed structure and functional logic of the current user interface (including interface elements, functional modules and their attributes, etc.), while the user instructions express the user's operational intentions or needs. Based on these inputs, the intelligent agent, combined with the context and rules in the page description information, deeply analyzes and infers the user instructions. For example, by analyzing the relationship between the user instructions and the page description information, the intelligent agent identifies the specific page interaction elements (such as buttons, input boxes or menu items, etc.) in the first user page that can respond to user instructions. These interaction elements are encapsulated as the first interaction information (also known as the action thinking chain). For example, the first interaction information is "The current page is on the xxx page, and there is an xxx control at the xxx position on the page. You need to click the plus button at the xxx position to add it to the shopping cart."
[0088] In order to enable the user to more clearly understand the currently available interactive elements and their functions, in an embodiment of the present application, the first interaction information can be presented to the user.
[0089] As a possible implementation manner, the first interaction information is visually displayed in a scenario where it is convenient for the user to view the screen.
[0090] As another possible implementation, the first interaction information is voice broadcasted while driving, exercising, or in other scenarios where it is inconvenient to view the screen.
[0091] It should be noted that, in actual application, one of the two possible implementation methods mentioned above can be executed, or both possible implementation methods can be executed at the same time, and this application does not make specific limitations.
[0092] Step 303: Call the intelligent agent to generate first action information based on the first interaction information and multiple candidate interaction actions.
[0093] The first action information at least includes a first target action.
[0094] As one possible implementation, the first interaction information is received and combined with multiple candidate interaction actions (i.e., possible response options), and an agent is called to analyze and make a decision. The agent comprehensively considers the compatibility of the first interaction information with the candidate actions, selects the optimal action plan, and generates first action information. The first action information includes at least a first target action, i.e., an operation instruction that best meets the user's needs (e.g., "click a button" or "jump to a page").
[0095] In order to allow the user to understand the interactive response more intuitively, in an embodiment of the present application, the first action information can be presented to the user.
[0096] As a possible implementation manner, the first action information is visually displayed in a scenario where it is convenient for the user to view the screen.
[0097] As another possible implementation, the first action information is voice broadcasted while driving, exercising, or in other scenarios where it is inconvenient to view the screen.
[0098] It should be noted that, in actual application, one of the two possible implementation methods mentioned above can be executed, or both possible implementation methods can be executed at the same time, and this application does not make specific limitations.
[0099] Step 304: Execute a first target action on the first user interface to implement an interactive response to the user instruction.
[0100] It should be noted that the execution process of step 301 and step 304 can be implemented in any of the embodiments of the present application. The embodiments of the present application do not limit this and will not be described in detail.
[0101] In summary, based on the first page description information, the intelligent agent is called to interactively reason about the user instructions and generate the first interaction information, which ensures that the system can accurately locate the page interaction elements corresponding to the user instructions and avoid ambiguous or erroneous operations; then, based on the first interaction information and multiple candidate interaction actions, the intelligent agent is called to generate the first action information, which realizes the selection of the optimal action plan according to different scenarios and enhances the flexibility and adaptability of the system; finally, the first target action is executed on the first user interface, which not only ensures the immediacy of the interactive response, but also optimizes the user's operating experience.
[0102] The above are various embodiments corresponding to the application method of the intelligent agent. This application also proposes a training method for the intelligent agent. Figure 4 A flowchart of a method for training an intelligent agent provided in an embodiment of the present application.
[0103] It should be noted that the training method of the intelligent agent can be executed alone, or it can be executed in combination with any embodiment of the present application or a possible implementation method in the embodiment, or it can be executed in combination with any technical solution in the relevant technology, and the embodiments of the present application do not limit this.
[0104] like Figure 4 As shown, the training method of the agent may include the following steps:
[0105] Step 401: Obtain training samples.
[0106] The training samples include sample instructions and sample page description information of a sample user interface.
[0107] In an embodiment of the present application, the sample instruction may represent an operation request issued by a user, and the sample page description information of the sample user interface is used to describe the structure and function of the sample user interface.
[0108] It should be noted that the training samples include at least sample instructions and sample page descriptions of the sample user interface. To ensure the diversity and practicality of the training data, the training samples may also include sample interaction data associated with the sample user interface. The sample interaction data includes sample images of the sample user interface and behavioral data generated during user interaction with the sample user interface, such as operation records such as clicking a button, filling out a form, and selecting a menu item. The sample interaction data can reflect how users interact with the interface elements of the sample user interface in actual use.
[0109] It should be noted that, in order to ensure the diversity of the intelligent agent output, as a possible implementation method, instructions under at least one dimension associated with the sample user interface can be obtained as sample instructions.
[0110] In an embodiment of the present application, a first instruction under a first dimension associated with a sample user interface is obtained; wherein the first dimension includes at least one of the following: component anchoring, component context, and page description; a second instruction under a second dimension associated with the sample user interface is obtained; wherein the second dimension includes at least one of the following: component function description and component nesting relationship; a third instruction under a third dimension associated with the sample user interface is obtained; wherein the third dimension includes at least one of the following: page structure and page jump; a sample instruction is determined based on at least one of the first instruction, the second instruction, and the third instruction.
[0111] That is to say, by obtaining instructions associated with the sample user interface from multiple dimensions (first dimension, second dimension and third dimension) respectively, and determining sample instructions based on at least one of the instructions of multiple dimensions, the diversity and completeness of the training samples are ensured; wherein, the first dimension focuses on the basic interface elements and their contextual information in the sample user interface, including but not limited to the following: component anchoring, component context and page description, that is, the first instructions under the first dimension include general sample instructions; the second dimension focuses on the functions of the interface elements and their mutual relationships, including but not limited to the following: component function description and component nesting relationship, that is, the second instructions under the second dimension include instructions focusing on component function description and nesting and hierarchical relationships between components; the third dimension focuses on the overall structure of the page and the jump relationship between pages, including but not limited to the following: page structure and page jump, that is, the third instructions under the third dimension include instructions focusing on partial description of the page structure framework and control interaction triggered page jump prediction.
[0112] Step 402: Use the initial intelligent agent to generate predicted actions associated with sample instructions based on training samples.
[0113] As a possible implementation method, the training samples are input into the intelligent agent to obtain the predicted action output by the intelligent agent.
[0114] Step 403: Train the initial intelligent agent according to the sample actions and predicted actions marked by the sample instructions to obtain a trained intelligent agent.
[0115] In order to avoid overfitting or underfitting problems caused by training, as a possible implementation method, the intelligent agent is trained in a stage-by-stage manner.
[0116] As an example, the initial agent is trained based on the difference between the correct sample actions marked by the sample instructions and the predicted actions generated by the initial agent to obtain an intermediate agent with improved performance; in order to further improve the generalization ability of the agent, the intermediate agent is further trained to finally obtain a trained agent.
[0117] In summary, by obtaining training samples containing sample instructions and sample page description information, the initial intelligent agent generates predicted actions for interacting with the sample user interface based on the training samples. The diverse training samples enable the intelligent agent to learn different sample user interface layouts and interaction modes, and acquire the ability to understand user intentions and parse interface structures. Furthermore, the predicted actions generated by the intelligent agent are compared and analyzed with the sample actions marked by the sample instructions to optimize the algorithm model of the intelligent agent and gradually reduce the prediction error. As a result, the trained intelligent agent can respond to user instructions quickly and accurately, provide a smooth and natural interactive experience, and improve user satisfaction and usage efficiency.
[0118] In order to clearly illustrate how the initial intelligent agent is trained according to the sample actions and predicted actions marked by the sample instructions in the above embodiment to obtain a trained intelligent agent, the present disclosure proposes another intelligent agent training method.
[0119] Figure 5 A flowchart of a method for training an intelligent agent provided in an embodiment of the present application.
[0120] like Figure 5 As shown, the training method of the agent may include the following steps:
[0121] Step 501: Obtain training samples.
[0122] The training samples include sample instructions, sample interaction data associated with the sample user interface, and sample page description information of the sample user interface.
[0123] Step 502: Use the initial intelligent agent to generate predicted actions associated with sample instructions based on training samples.
[0124] Step 503: Train the initial intelligent agent according to the sample actions and predicted actions marked by the sample instructions to obtain an intermediate intelligent agent.
[0125] As a possible implementation, a loss function is generated based on the difference between the sample actions annotated by the sample instructions and the predicted actions. The initial agent is trained based on this loss function to obtain an intermediate agent. For example, the agent parameters can be optimized by maximizing the probability that the agent predicts the standard action, thereby reducing the difference between the predicted action and the sample action.
[0126] It should be noted that in order to ensure the diversity of the agent output, an early checkpoint (e.g., the inflection point of the loss function curve) is selected as the intermediate agent.
[0127] Step 504: Based on the sample interaction data, perform at least one round of training on the intermediate agent to obtain a trained agent.
[0128] It should be understood that since the sample interaction data reflects the actual interaction method between the user and the interface elements, it can supplement the semantic information in the training samples and enable the intelligent agent to better understand the relationship between instructions and interface elements. Therefore, in an embodiment of the present application, the intermediate intelligent agent is trained for at least one round based on the sample interaction data to obtain a trained intelligent agent, which can enable the intelligent agent to accurately understand user instructions and accurately respond to user instructions.
[0129] As an example, a sample image of a sample user interface in the sample interaction data is obtained; based on the sample image and the sample instructions, at least one round of training is performed on the intermediate intelligent agent to obtain a trained intelligent agent.
[0130] In order to further improve the understanding and reasoning abilities of the intelligent agent, in an embodiment of the present application, sample images of sample user interfaces are extracted from sample interaction data. The sample images can supplement the deficiencies in the training data and provide more comprehensive contextual information, thereby improving the understanding and reasoning abilities of the intelligent agent.
[0131] In an embodiment of the present application, the intermediate intelligent agent is an intelligent agent that has undergone preliminary training and has acquired certain understanding and reasoning abilities, but there may still be errors or limitations. Furthermore, the intermediate intelligent agent combines sample instructions and sample images for at least one round of training, so that the intelligent agent can quickly deduce the correct interactive actions based on the sample instructions, that is, learn how to map the sample instructions to specific interface interactions. For example, the intelligent agent can recognize the "Add to Cart" button through the image and associate it with the sample instruction "How to purchase a product?"; at the same time, ensure that the intelligent agent can adapt to changes in different interface layouts and visual styles.
[0132] In an embodiment of the present application, for the i-th round of training in at least one round of training, the i-th round of intelligent agent to be trained is used to perform at least one page description of the sample user interface based on the sample image to obtain at least one second page description information of the sample user interface; the i-th round of intelligent agent to be trained is used to interactively reason on the sample instructions based on any second page description information to obtain second interactive information of the sample instructions under any second page description information; the i-th round of intelligent agent to be trained is used to generate second action information of the sample instructions under any second page description information based on the second interactive information and multiple candidate interactive actions under any second page description information; based on the sample image, each second page description information and the second interactive information and second action information under each second page description information, the i-th round of intelligent agent to be trained is trained to obtain an intelligent agent after the i-th round of training; wherein, in response to i being equal to 1, the i-th round of intelligent agent to be trained is the intermediate intelligent agent, and in response to i not being equal to 1, the i-th round of intelligent agent to be trained is the intelligent agent after the i-1-th round of training.
[0133] That is, for the first round of training, a sample image and a page description instruction are input into the intermediate intelligent agent, so that the intermediate intelligent agent is used to perform page description of the sample user interface to obtain a second page description information of the sample user interface. This process is performed multiple times to obtain multiple second page description information of the sample user interface; then, for any second page description information, the intermediate intelligent agent is called, and interactive reasoning is performed on the sample instruction based on the any second page description information and the sample instruction to obtain the second interactive information of the sample instruction under the any second page description information; then, the intermediate intelligent agent outputs the second action information of the sample instruction under the any second page description information based on the second interactive information of the sample instruction under the any second page description information and multiple candidate interactive actions; then, based on the sample image, each second page description information, and the second interactive information and second action information under each second page description information, the intermediate intelligent agent is trained to obtain the intelligent agent after the first round of training;
[0134] For non-first round training, a sample image and a page description instruction are input into the intelligent agent after the i-1th round of training, so that the sample user interface is described by the intelligent agent after the i-1th round of training, and a second page description information of the sample user interface output by the intelligent agent after the i-1th round of training is obtained. This process is performed multiple times to obtain multiple second page description information of the sample user interface; wherein i is a positive integer greater than 1, and then, for any second page description information, the intelligent agent after the i-1th round of training, based on the any second page description information, performs interactive reasoning on the sample instruction to obtain the second interactive information of the sample instruction under the any second page description information; and then, the intelligent agent after the i-1th round of training outputs the second action information of the sample instruction under any second page description information based on the second interactive information of the sample instruction under the any second page description information and multiple candidate interactive actions. Then, based on the sample image, each second page description information and the second interactive information and second action information under each second page description information, the intelligent agent after the i-1th round of training is trained to obtain the intelligent agent after the i-th round of training.
[0135] In order to further improve the understanding and reasoning abilities of the agent, as a possible implementation method, as an example, based on the sample images, each second page description information, and the second interaction information and second action information under each second page description information, the agent to be trained in the i-th round is trained to obtain the agent after the i-th round of training. The steps are as follows:
[0136] 1. Create a dependency tree based on the sample image, each second page description information, and the dependency relationship between the second interaction information and the second action information under each second page description information;
[0137] In order to improve the accuracy of action generation, in an embodiment of the present application, the sample image serves as the root node of the dependency tree, the child nodes under the root node are multiple second page description information of the sample image, the child nodes of the multiple second page description information are the corresponding second interaction information, and the child nodes of the second reasoning information are the corresponding second action information. Through the dependency tree, the intelligent agent can more accurately identify the logical relationship between different interactive elements, thereby generating actions that are more in line with the user's intentions.
[0138] 2. Determine the evaluation value of each leaf node based on the degree of match between each leaf node in the dependency tree and the sample actions corresponding to each leaf node;
[0139] In the embodiment of the present application, for any leaf node, the evaluation value v(s t ) is determined using the following formula:
[0140]
[0141] Among them, d(s t ,a * ) [x,y] represents the distance between the leaf node and the corresponding sample action, F1(s t ,a * ) text In text-related tasks, the text matching score between the leaf node and the corresponding sample action, dir(s t ~a * ) seroll Indicates the matching degree between the leaf node and the corresponding sample action in the scrolling direction, type(s t ~a * ) indicates the degree of matching between the leaf node and the corresponding sample action type;
[0142] It should be noted that the distance between the leaf node and the corresponding sample action is less than d min (e.g., 0.05), the evaluation value of the leaf node is assigned to 1; when the text matching score between the leaf node and the corresponding sample action exceeds 0.5, the evaluation value of the leaf node is assigned to 1; when the degree of matching between the leaf node and the corresponding sample action in the scrolling direction is greater than the set threshold, the evaluation value of the leaf node is assigned to 1; when the degree of matching between the leaf node and the corresponding sample action in the scrolling direction is less than the set threshold, the evaluation value of the leaf node is assigned to 0.
[0143] 3. Based on the evaluation value of each leaf node, the intelligent agent to be trained in the i-th round is trained to obtain the intelligent agent after the i-th round of training.
[0144] In order to improve the preference learning ability of the intelligent agent, in an embodiment of the present application, based on the evaluation value of each leaf node, the evaluation value of other nodes in the dependency tree except the leaf nodes is determined, and based on the evaluation value of each node in the dependency tree, the preference training of the intelligent agent is performed.
[0145] As an example, based on the evaluation value of each leaf node, the evaluation value of other nodes in the dependency tree except the leaf nodes is determined; for any parent node in the dependency tree, according to the evaluation value of the child node of any parent node, a candidate data pair is determined from the child nodes of any parent node; wherein the evaluation values of the two child nodes in the candidate data pair meet the set preference requirements; from the corresponding candidate data pairs of each parent node, a target data pair is selected; the target data pair is used to train the intelligent agent to be trained in the i-th round to obtain the intelligent agent after the i-th round of training.
[0146] That is, the evaluation values of the nodes other than the leaf nodes in the dependency tree are calculated based on the evaluation value of each leaf node through a recursive calculation formula, wherein the recursive calculation formula is as follows:
[0147]
[0148] Among them, t represents the depth of the node in the dependency tree, s t-1 It is t The parent node of , K represents the number of child nodes, and V represents the evaluation value of the node.
[0149] Furthermore, for any parent node in the dependency tree, the child node with the highest evaluation value and the child node with the lowest evaluation value are selected according to the evaluation value of the child nodes of any parent node to form a candidate data pair (also called preference data pair (Direct Preference Optimization, DPO)). Furthermore, in order to ensure the effectiveness and diversity of preference training, the candidate data pairs with small differences in quality and content are filtered out from the candidate data pairs to retain the target data pairs with sufficient differences in quality and content. Finally, the target data pairs are used to train the intelligent agent to be trained in the i-th round to obtain the intelligent agent after the i-th round of training.
[0150] In summary, based on sample images and sample instructions, at least one round of training is performed on the intermediate intelligent agent, which enables the intelligent agent to learn how to match interface information with user needs in the context of multiple rounds of asking actual user instructions, thereby improving the accuracy of its interactive reasoning and interactive response, thereby improving user experience and operational efficiency.
[0151] In any embodiment of the present application, Figure 6 As shown, the training method of the intelligent agent in the embodiment of the present application can also be implemented based on the following steps:
[0152] Phase 1: Instruction Evolution & Supervised Fine-tuning
[0153] Specifically, obtaining sample instructions associated with the sample user interface;
[0154] Among them, the sample instructions mainly include instructions in the following three dimensions:
[0155] (1) a first instruction in a first dimension associated with the sample user interface; wherein the first dimension includes at least one of the following: a component anchor, a component context, and a page description, i.e., a general question-answer pair associated with the sample user interface;
[0156] (2) a second instruction under a second dimension associated with the sample user interface; wherein the second dimension includes at least one of the following: a component function description and a component nesting relationship;
[0157] (3) a third instruction in a third dimension associated with the sample user interface; wherein the third dimension includes at least one of the following: page structure and page jump;
[0158] To train an agent capable of basic tasks and expand the diversity of its output, sample interaction data (also known as trajectory data) is mixed with sample instructions and sample page description information as training samples. The initial agent is trained based on these training samples, that is, supervised fine-tuning is performed on the training samples. The specific formula is as follows:
[0159] L SFT (θ)=-E (o,q)~D [logπ θ (a|o,q)];
[0160] Among them, L SFT (θ) represents the loss function of supervised fine-tuning, o represents the sample image, q represents the sample instruction, and a represents the labeled sample action; π θ (a|o,q) represents the probability of the agent generating a sample action a given the inputs o and q;
[0161] In addition, it should be noted that in order to ensure the diversity of the agent output, an early checkpoint (e.g., the inflection point of the loss function curve) is selected as the intermediate agent;
[0162] Phase 2: Action-level CoaT sampling
[0163] Sample each action in the sample interaction data and combine it with CoaT for modeling;
[0164] Specifically, the steps include:
[0165] (1) For each sample image in the sample interaction data, call the agent obtained in the first stage, perform page description on the sample user interface according to the sample image and the sample page description instruction, so as to obtain second page description information of the sample user interface, and execute this process multiple times to obtain multiple page description information of the sample user interface;
[0166] (2) For each second page description information of the sample user interface, call the agent obtained in the first stage, and obtain a second interaction information (also called action thinking chain) according to the sample instruction and the second page description information;
[0167] (3) For each second interaction information, call the agent obtained in the first stage, and determine the second target action based on the second interaction information and multiple candidate interaction actions;
[0168] (4) creating a dependency tree (also called a sampling tree) based on the dependency relationships among the plurality of second page description information, the second interaction information, and the second target action corresponding to each sample image;
[0169] (5) For each leaf node under the dependency tree, the matching degree between the target interaction action and the real action corresponding to each leaf node is calculated to calculate the value of each leaf node;
[0170] (6) For the value of each leaf node, calculate the value of each node in the dependency tree based on the recursive formula;
[0171] Stage 3: Preference Learning
[0172] Based on the value of each node, a DPO pair is constructed and the selected DPO is used to train the agent obtained in the first stage;
[0173] The second and third phases are repeated until convergence.
[0174] Furthermore, it should be noted that trained AI agents can be applied to a variety of practical application scenarios, including but not limited to the following areas: 1) Automated Testing of Intelligent Graphical User Interfaces (GUIs): Automated testing of the functional integrity of GUI applications improves test coverage and reduces manual intervention. 2) GUI Task Planning and Intelligent Navigation: AI agents can autonomously perform complex tasks, such as cross-page operations and task flow automation. 3) Enhanced Interactive AI Agents: Trained AI agents are built with the ability to reason about context and interact autonomously, improving the intelligence of GUI task execution.
[0175] In order to achieve the above Figures 1 to 3 In an embodiment, the present application proposes an interactive device.
[0176] Figure 7 A schematic diagram of the structure of an interactive device provided in an embodiment of the present application.
[0177] like Figure 7 As shown, the interaction device 700 includes: a receiving module 710 , a generating module 720 and an executing module 730 .
[0178] Among them, the receiving module 710 is used to receive user instructions; the generating module 720 is used to generate a first target action associated with the first user interface based on the first page description information of the first user interface and the user instructions; the executing module 730 is used to execute the first target action on the first user interface to achieve an interactive response to the user instructions.
[0179] As a possible implementation manner, the execution module 730 is configured to execute at least one round of page interaction process according to the first user interface and the first target action, so as to implement an interactive response to the user instruction.
[0180] As a possible implementation method, the execution module 730 is used to execute the first target action on the first user interface to obtain and display the second user interface to which the first round of page interaction process in at least one round of page interaction process jumps; when it is determined according to the second user interface that the user instruction response is completed, the page interaction process is ended.
[0181] As a possible implementation method, the execution module 730 is also used to determine the second target action of the i-th round of page interaction process based on the second page description information and user instructions of the second user interface to which the i-1th round of page interaction process jumps, when it is determined that the user instruction has not been fully responded to according to the second user interface to which the i-1th round of page interaction process jumps; wherein, i is a positive integer greater than 1; execute the second target action of the i-th round of page interaction process on the second user interface to which the i-1th round of page interaction process jumps to, so as to obtain the second user interface to which the i-th round of page interaction process jumps, and determine whether the user instruction has been fully responded to according to the second user interface to which the i-1th round of page interaction process jumps, until it is determined that the user instruction has been fully responded to, and end the page interaction process.
[0182] As a possible implementation method, generation module 720 is used to call the intelligent agent to perform interactive reasoning on the user instruction based on the first page description information to obtain first interaction information; wherein the first interaction information is used to indicate the page interaction element in the first page for responding to the user instruction; based on the first interaction information and multiple candidate interaction actions, call the intelligent agent to generate first action information; wherein the first action information includes at least a first target action.
[0183] As a possible implementation manner, the interaction device 700 further includes: a first processing module.
[0184] Among them, the first processing module is used to visually display and / or voice broadcast the first interaction information; and visually display and / or voice broadcast the first action information.
[0185] As a possible implementation manner, the interaction device 700 further includes: a second processing module.
[0186] Among them, the second processing module is used to obtain the target image and page description instructions showing the first user interface; wherein the page description instructions are used to prompt the intelligent agent to perform the page description task; the intelligent agent is used to perform page description of the first user interface according to the target image and page description instructions to obtain the first page description information of the first user interface.
[0187] As a possible implementation manner, the interactive device 700 further includes: an acquisition module.
[0188] Among them, the acquisition module is used to obtain user instructions from the target control in response to a text input operation on the first interactive control in the first user interface; obtain input user instructions in response to a voice input operation on the second interactive control in the first user interface; and determine the user instruction from at least one candidate instruction displayed in the first user interface in response to a selection operation.
[0189] The interactive device of the embodiment of the present disclosure receives user instructions and, based on the page description information of the first user interface and the user instructions, intelligently generates a first target action that precisely matches the first user interface, and then executes the first target action on the first user interface, thereby achieving efficient and accurate interactive response to user instructions. This mechanism not only promotes the development of user interface interaction towards intelligence and automation, significantly improves the speed and accuracy of responding to user instructions, but also can flexibly respond to user interfaces of different structures and designs (including complex multi-element interfaces or simple single-function interfaces), enhances the consistency and satisfaction of user experience, and makes operations in various application scenarios smoother and more natural.
[0190] In order to achieve the above Figures 4 to 6 In an embodiment, the present application proposes an interactive device.
[0191] Figure 8 A schematic structural diagram of an intelligent agent training device provided in an embodiment of the present application.
[0192] like Figure 8 As shown, the training device 800 of the intelligent agent includes: an acquisition module 810, a generation module 820 and a training module 830.
[0193] Among them, the acquisition module 810 is used to obtain training samples; wherein the training samples include sample instructions and sample page description information of the sample user interface; the generation module 820 is used to use the initial intelligent agent to generate predicted actions associated with the sample instructions based on the training samples; the training module 830 is used to train the initial intelligent agent based on the sample actions and predicted actions marked by the sample instructions to obtain a trained intelligent agent.
[0194] As a possible implementation method, the training sample also includes: sample interaction data associated with the sample user interface, a training module 830, which is used to train the initial intelligent agent according to the sample actions marked by the sample instructions and the predicted actions to obtain an intermediate intelligent agent; based on the sample interaction data, at least one round of training is performed on the intermediate intelligent agent to obtain a trained intelligent agent.
[0195] As a possible implementation method, the training module 830 is used to obtain sample images of sample user interfaces in the sample interaction data; based on the sample images and sample instructions, perform at least one round of training on the intermediate agent to obtain a trained agent.
[0196] As a possible implementation method, the training module 830 is used to use the intelligent agent to be trained in the i-th round to perform at least one page description of the sample user interface based on the sample image to obtain at least one second page description information of the sample user interface; use the intelligent agent to be trained in the i-th round to perform interactive reasoning on the sample instructions based on any second page description information to obtain the second interactive information of the sample instructions under any second page description information; use the intelligent agent to be trained in the i-th round to generate the second action information of the sample instructions under any second page description information based on the second interactive information and multiple candidate interactive actions under any second page description information; train the intelligent agent to be trained in the i-th round based on the sample image, each second page description information and the second interactive information and second action information under each second page description information to obtain the intelligent agent after the i-th round of training; wherein, in response to i being equal to 1, the intelligent agent to be trained in the i-th round is the intermediate intelligent agent, and in response to i not being equal to 1, the intelligent agent to be trained in the i-th round is the intelligent agent after the i-1-th round of training.
[0197] As a possible implementation method, the training module 830 is used to create a dependency tree based on the sample pictures, each second page description information, and the dependency between the second interaction information and the second action information under each second page description information; determine the evaluation value of each leaf node based on the degree of matching between each leaf node in the dependency tree and the sample action corresponding to each leaf node; and train the intelligent agent to be trained in the i-th round based on the evaluation value of each leaf node to obtain the intelligent agent after the i-th round of training.
[0198] As a possible implementation method, the training module 830 is used to determine the evaluation values of other nodes in the dependency tree except for each leaf node based on the evaluation value of each leaf node; for any parent node in the dependency tree, according to the evaluation value of the child node of any parent node, determine the candidate data pair from the child nodes of any parent node; wherein the evaluation values of the two child nodes in the candidate data pair meet the set preference requirements; select the target data pair from the corresponding candidate data pairs of each parent node; use the target data pair to train the intelligent agent to be trained in the i-th round to obtain the intelligent agent after the i-th round of training.
[0199] As a possible implementation method, the acquisition module 810 is also used to obtain at least one of the following: obtaining a first instruction under a first dimension associated with the sample user interface; wherein the first dimension includes at least one of the following: component anchoring, component context, and page description; obtaining a second instruction under a second dimension associated with the sample user interface; wherein the second dimension includes at least one of the following: component function description and component nesting relationship; obtaining a third instruction under a third dimension associated with the sample user interface; wherein the third dimension includes at least one of the following: page structure and page jump.
[0200] The training device for the intelligent agent of the embodiment of the present disclosure obtains training samples containing sample instructions and sample page description information. The initial intelligent agent generates predicted actions for interacting with the sample user interface based on the training samples. The diversified training samples enable the intelligent agent to learn different sample user interface layouts and interaction modes, and acquire the ability to understand user intentions and parse interface structures. Furthermore, the predicted actions generated by the intelligent agent are compared and analyzed with the standard sample actions marked with the sample instructions, so as to optimize the algorithm model of the intelligent agent and gradually reduce the prediction error. Thus, the trained intelligent agent can execute user instructions quickly and accurately, provide a smooth and natural interactive experience, reduce waiting time and error rate in user operations, and improve user satisfaction and usage efficiency.
[0201] In order to implement the above embodiment, the present application also proposes an electronic device, comprising a processor, and a memory connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the above embodiment. Figures 1 to 3 The interactive method described in the embodiment, or, implementing the above Figures 4 to 6 The training method of the intelligent agent described in the embodiment.
[0202] In order to implement the above embodiment, the present application also proposes a computer-readable storage medium, in which computer-executable instructions are stored, and the computer-executable instructions are used to implement the above embodiment when executed by the processor. Figures 1 to 3 The interactive method described in the embodiment, or, implementing the above Figures 4 to 6 The training method of the intelligent agent described in the embodiment.
[0203] In order to implement the above embodiments, the present application also proposes a computer program product having a computer program stored thereon, which implements the above embodiments when the computer program is executed by a processor. Figures 1 to 3 The interactive method described in the embodiment, or, implementing the above Figures 4 to 6 The training method of the intelligent agent described in the embodiment.
[0204] Figure 9This is a block diagram of an electronic device provided in an embodiment of the present application. For example, the electronic device 900 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0205] Reference Figure 9 , the electronic device 900 may include one or more of the following components: a processing component 902 , a memory 904 , a power component 906 , a multimedia component 908 , an audio component 910 , an input / output (I / O) interface 912 , a sensor component 914 , and a communication component 916 .
[0206] The processing component 902 generally controls the overall operation of the electronic device 900, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 902 may include one or more processors 920 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 902 may include one or more modules to facilitate interaction between the processing component 902 and other components. For example, the processing component 902 may include a multimedia module to facilitate interaction between the multimedia component 908 and the processing component 902.
[0207] The memory 904 is configured to store various types of data to support operations on the electronic device 900. Examples of such data include instructions for any application or method operating on the electronic device 900, contact data, phone book data, messages, pictures, videos, etc. The memory 904 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0208] The power component 906 provides power to the various components of the electronic device 900. The power component 906 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 900.
[0209] The multimedia component 908 includes a screen that provides an output interface between the electronic device 900 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 908 includes a front camera and / or a rear camera. When the electronic device 900 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0210] The audio component 910 is configured to output and / or input audio signals. For example, the audio component 910 includes a microphone (MIC), and when the electronic device 900 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 904 or transmitted via the communication component 916. In some embodiments, the audio component 910 also includes a speaker for outputting audio signals.
[0211] I / O interface 912 provides an interface between processing component 902 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.
[0212] The sensor assembly 914 includes one or more sensors for providing various aspects of status assessment for the electronic device 900. For example, the sensor assembly 914 can detect the open / closed state of the electronic device 900, the relative positioning of components, such as the display and keypad of the electronic device 900. The sensor assembly 914 can also detect changes in the position of the electronic device 900 or a component of the electronic device 900, the presence or absence of user contact with the electronic device 900, the orientation or acceleration / deceleration of the electronic device 900, and temperature changes of the electronic device 900. The sensor assembly 914 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 914 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 914 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0213] The communication component 916 is configured to facilitate wired or wireless communication between the electronic device 900 and other devices. The electronic device 900 can access a wireless network based on a communication standard, such as WiFi, 4G or 5G, or a combination thereof. In an exemplary embodiment, the communication component 916 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 916 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0214] In an exemplary embodiment, the electronic device 900 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described methods.
[0215] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 904 including instructions, and the instructions can be executed by the processor 920 of the electronic device 900 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0216] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0217] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of such features. Throughout the description of this application, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0218] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.
[0219] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.
[0220] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0221] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0222] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0223] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. An interactive method, characterized in that: include: Receive user instructions; generating a first target action associated with the first user interface according to the first page description information of the first user interface and the user instruction; The first target action is executed on the first user interface to implement an interactive response to the user instruction.
2. The method according to claim 1, characterized in that The executing the target action on the first user interface to achieve an interactive response to the user instruction includes: At least one round of page interaction process is performed according to the first user interface and the first target action to achieve an interactive response to the user instruction.
3. The method according to claim 2, characterized in that The performing at least one round of page interaction process according to the first user interface and the first target action to implement an interactive response to the user instruction includes: Executing the first target action on the first user interface to obtain and display a second user interface to which the first round of page interaction process in the at least one round of page interaction process jumps; When it is determined according to the second user interface that the response to the user instruction is complete, the page interaction process is ended.
4. The method according to claim 3, characterized in that The performing of at least one round of page interaction process according to the first user interface and the first target action to complete the interactive response to the user instruction further includes: If it is determined that the user instruction has not been fully responded to according to the second user interface to which the page interaction process jumps in the (i-1) round, determining a second target action of the i-th round of page interaction process based on the second page description information of the second user interface to which the page interaction process jumps in the (i-1) round and the user instruction; wherein i is a positive integer greater than 1; On the second user interface to which the i-1th round of page interaction process jumps, the second target action of the i-th round of page interaction process is executed to obtain the second user interface to which the i-th round of page interaction process jumps, and based on the second user interface to which the i-th round of page interaction process jumps, it is determined whether the user instruction has been responded to, and the page interaction process is terminated until it is determined that the user instruction has been responded to.
5. The method according to claim 1, wherein The generating, according to the first page description information of the first user interface and the user instruction, a first target action associated with the user instruction includes: Invoking an agent to interactively reason on the user instruction based on the first page description information to obtain first interaction information; wherein the first interaction information is used to indicate a page interaction element in the first page used to respond to the user instruction; Based on the first interaction information and multiple candidate interaction actions, the intelligent agent is called to generate the first action information; wherein the first action information at least includes the first target action.
6. The method according to claim 5, characterized in that The method further comprises at least one of the following: Visually display and / or voice broadcast the first interactive information; The first action information is visually displayed and / or voice broadcasted.
7. The method according to claim 1, characterized in that The first page description information of the first user interface is obtained by using the following steps: Obtaining a target image showing the first user interface and a page description instruction; wherein the page description instruction is used to prompt the agent to perform a page description task; The intelligent agent is used to perform page description on the first user interface according to the target image and the page description instruction to obtain first page description information of the first user interface.
8. The method according to claim 1, characterized in that The user instruction is obtained by using any of the following methods: In response to a text input operation on a first interactive control in the first user interface, acquiring the user instruction from the target control; In response to a voice input operation on a second interactive control in the first user interface, acquiring the input user instruction; In response to a selection operation, the user instruction is determined from at least one candidate instruction displayed in the first user interface.
9. A method for training an intelligent agent, characterized in that: include: Obtaining training samples; wherein the training samples include sample instructions and sample page description information of a sample user interface; Using an initial intelligent agent to generate a predicted action associated with the sample instruction based on the training sample; The initial intelligent agent is trained according to the sample actions marked by the sample instructions and the predicted actions to obtain a trained intelligent agent.
10. The method according to claim 9, characterized in that The training sample also includes: sample interaction data associated with the sample user interface, the sample actions annotated by the sample instructions and the predicted actions, training the initial intelligent agent to obtain a trained intelligent agent, including: Training the initial intelligent agent according to the sample actions marked by the sample instructions and the predicted actions to obtain an intermediate intelligent agent; Based on the sample interaction data, at least one round of training is performed on the intermediate agent to obtain a trained agent.
11. The method according to claim 10, characterized in that The step of performing at least one round of training on the intermediate agent based on the sample interaction data to obtain a trained agent includes: Obtaining a sample image of the sample user interface in the sample interaction data; Based on the sample images and the sample instructions, at least one round of training is performed on the intermediate agent to obtain a trained agent.
12. The method according to claim 11, characterized in that The i-th round of training in the at least one round of training includes: Using the intelligent agent to be trained in the i-th round, perform at least one page description of the sample user interface based on the sample image to obtain at least one second page description information of the sample user interface; Using the intelligent agent to be trained in the i-th round, interactive reasoning is performed on the sample instruction according to any second page description information to obtain second interactive information of the sample instruction under the any second page description information; Using the intelligent agent to be trained in the i-th round, generating second action information of the sample instruction under any second page description information according to the second interaction information under any second page description information and the multiple candidate interaction actions; Based on the sample images, each second page description information, and the second interaction information and second action information under each second page description information, the agent to be trained in the i-th round is trained to obtain an agent after the i-th round of training; In which, in response to i being equal to 1, the intelligent agent to be trained in the i-th round is the intermediate intelligent agent, and in response to i not being equal to 1, the intelligent agent to be trained in the i-th round is the intelligent agent after the i-1-th round of training.
13. The method according to claim 12, characterized in that The training of the intelligent agent to be trained in the i-th round based on the sample images, each second page description information, and the second interaction information and second action information under each second page description information to obtain an intelligent agent after the i-th round of training includes: Creating a dependency tree based on the dependency relationships among the sample images, each piece of second page description information, and the second interaction information and second action information under each piece of second page description information; Determining the evaluation value of each leaf node based on the degree of matching between each leaf node in the dependency tree and the sample action corresponding to each leaf node; Based on the evaluation value of each leaf node, the intelligent agent to be trained in the i-th round is trained to obtain an intelligent agent after the i-th round of training.
14. The method according to claim 13, wherein: The step of training the intelligent agent to be trained in the i-th round based on the evaluation value of each leaf node to obtain an intelligent agent after the i-th round of training includes: Determining the evaluation values of other nodes in the dependency tree except for each leaf node based on the evaluation value of each leaf node; For any parent node in the dependency tree, determining a candidate data pair from the child nodes of any parent node according to the evaluation values of the child nodes of the parent node; wherein the evaluation values of the two child nodes in the candidate data pair meet the set preference requirements; Selecting a target data pair from the candidate data pairs corresponding to each of the parent nodes; The target data pair is used to train the intelligent agent to be trained in the i-th round to obtain an intelligent agent after the i-th round of training.
15. The method according to claim 9, characterized in that The sample instruction is obtained by using at least one of the following: Obtaining a first instruction under a first dimension associated with the sample user interface; wherein the first dimension includes at least one of the following: component anchoring, component context, and page description; Obtaining a second instruction under a second dimension associated with the sample user interface; wherein the second dimension includes at least one of the following: component function description and component nesting relationship; Obtain a third instruction under a third dimension associated with the sample user interface; wherein the third dimension includes at least one of the following: page structure and page jump.
16. An interactive device, characterized in that: For implementing the interaction method according to any one of claims 1 to 8, the device comprises: A receiving module, used for receiving user instructions; a generating module, configured to generate a first target action associated with the first user interface according to the first page description information of the first user interface and the user instruction; An execution module is used to execute the first target action on the first user interface to achieve an interactive response to the user instruction.
17. A training device for an intelligent agent, characterized in that: For implementing the training method of an intelligent agent according to any one of claims 9 to 15, the device comprises: An acquisition module is used to acquire training samples; wherein the training samples include sample instructions and page description information of the sample user interface; A generation module, configured to use an initial agent to generate a predicted action associated with the sample instruction according to the page description information and the sample instruction; A training module is used to train the initial intelligent agent according to the sample actions marked by the sample instructions and the predicted actions to obtain a trained intelligent agent.
18. An electronic device, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 8, or implements the method according to any one of claims 9 to 15.
19. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method according to any one of claims 1 to 8; or, to implement the method according to any one of claims 9 to 15.
20. A computer program product, characterized in that, The invention comprises a computer program, which, when executed by a processor, implements the method described in any one of claims 1 to 8; or implements the method described in any one of claims 9 to 15.
Citation Information
Cited By
Interface operation instruction generation method, electronic equipment, storage medium and program product
CN120704792A