Action sequence execution method, electronic device, storage medium, and program product
By generating a global action sequence and monitoring anomalies in real time, and dynamically adjusting the execution path, the problem of low efficiency, insufficient reliability and flexibility in task execution in existing technologies is solved, and efficient, stable and flexible execution of automated processes is achieved.
Patent Information
- Application Number
- CN202511149469.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-08-18
AI Technical Summary
Existing technologies suffer from inefficiency, insufficient reliability and flexibility in multi-step operations, multi-action sequence combinations, and complex process control. They are also unable to cope with dynamically changing interfaces, which limits the generalization ability and practical value of automated intelligent agents.
By generating a global action sequence based on semantic features, anomalies are monitored and identified in real time, the execution path is dynamically adjusted, and a target action sequence is generated to update the global action sequence, thereby enabling the processing of the user interface.
It improves the success rate of automated processes, reduces process interruptions caused by unexpected situations, enhances the robustness and flexibility of task execution, and adapts to various complex and dynamic environments.
Smart Images

Figure CN120653167B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of human-computer interaction, and particularly relates to a motion sequence execution method, an electronic device, a storage medium and a program product. BACKGROUND
[0002] With the development of intelligent automation and human-computer interaction technology, more and more agents need to complete complex tasks through a graphical user interface (Web GUI), such as automatic testing, automatic form filling, web operation process automation, and the like. Existing visual and language-based agents still have significant bottlenecks in multi-step operations, multi-motion sequence combination, and complex process control.
[0003] The related art adopts single-step operations, which are too coarse or too fine in granularity, resulting in low efficiency in complex tasks and difficulty in forming high-level abstract motions. Moreover, the existing solutions lack effective modeling and reasoning of long motion sequences or multi-step combinations, causing insufficient reliability and flexibility of task execution, difficulty in coping with large-scale and dynamically changing interfaces, and limitation of multi-motion continuous coordination, which limits the generalization ability and practical value of automated agents. SUMMARY
[0004] Embodiments of the present application provide a motion sequence execution method, an electronic device, a storage medium and a program product to alleviate or solve the technical problem that the success rate of automatic execution process is not ideal in the related art.
[0005] In a first aspect, embodiments of the present application provide a motion sequence execution method, comprising:
[0006] Based on semantic features representing user instructions, a global motion sequence is generated, and the global motion sequence includes a plurality of node motions with an execution order;
[0007] After each node motion is executed, an exception is identified for each node motion to obtain an identification result corresponding to each node motion;
[0008] In a case where the identification result indicates that a target node motion has an exception, a target motion sequence is generated with the target node motion as a starting point; and the target motion sequence is used to update the global motion sequence;
[0009] According to the updated global motion sequence, a processing is performed on interface elements in an operation interface until a user instruction execution target is completed.
[0010] In a second aspect, embodiments of the present application provide an electronic device, comprising a memory, a processor and a computer program stored in the memory, and the processor implements any method of embodiments of the present application when executing the computer program.
[0011] In a third aspect, an embodiment of the present application provides a computer readable storage medium, and the computer readable storage medium stores a computer program. The computer program is executed by a processor to implement the method of any of the embodiments of the present application.
[0012] In a fourth aspect, an embodiment of the present application provides a computer program product, comprising a computer program. The computer program is executed by a processor to implement the method of any of the embodiments of the present application.
[0013] Based on the action sequence execution method of the first aspect, the present application has at least the following advantages:
[0014] Through the semantic-driven intelligent planning and dynamic adjustment mechanism, a closed-loop system with self-correction ability is constructed. After each operation, the system monitors the execution state in real time through multi-modal perception (visual recognition, interface structure analysis and semantic understanding), can accurately identify various abnormal situations such as element occlusion and state timeout, and when detecting abnormal situations, the system does not simply terminate the process, but intelligently generates a targeted correction scheme. For example, when it is found that the interface element is occluded, the operation action will be automatically adjusted to eliminate the occlusion effect. This dynamic adjustment capability enables the system to flexibly cope with various interface changes and unexpected situations, significantly improving the success rate of automated processes. Compared with traditional automation technology, the embodiments of the present application greatly reduce the process interruption caused by unexpected situations, and through intelligent local optimization strategy, the abnormal processing is more efficient and smooth. This dynamic adjustment capability enables the system to adapt to various complex and dynamic environments, and improves the robustness of task execution.
[0015] The above description is only a summary of the technical solutions of the present application. In order to more clearly understand the technical means of the present application, the embodiments can be implemented in accordance with the content of the description, and in order to make the above and other purposes, characteristics and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application are described below. BRIEF DESCRIPTION OF DRAWINGS
[0016] In the drawings, the same reference numbers in the several drawings represent the same or similar elements or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments according to the present application, and should not be regarded as limiting the scope of the present application.
[0017] Figure 1 A flowchart of an action sequence execution method according to an embodiment of the present application is shown;
[0018] Figure 2 A schematic diagram of an action sequence execution method according to an embodiment of the present application is shown;
[0019] Figure 3A block diagram of an action sequence execution apparatus of an embodiment of the present application is shown.
[0020] Figure 4 A block diagram of an electronic device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION
[0021] In the following, only certain exemplary embodiments are simply described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the concept or scope of the present application. Therefore, the drawings and the description are considered to be exemplary in nature, rather than limiting.
[0022] To facilitate understanding of the technical solutions of the embodiments of the present application, the related technologies of the embodiments of the present application are described below. The following related technologies can be combined with the technical solutions of the embodiments of the present application in any manner as optional solutions, and all of them belong to the protection scope of the embodiments of the present application.
[0023] In the following, the following terms will be used:
[0024] A graphical user interface (GUI) is an interface design that allows users to interact with computer systems or software applications through graphical elements such as windows, icons, menus, buttons, text boxes, etc. without relying on complex command line inputs. Users can interact with the system or application through mouse clicks, keyboard inputs, touch operations, etc.
[0025] Interface elements are components of a graphical user interface that enable interaction and information display. Through proper layout and design, users can intuitively interact with software or devices.
[0026] An agent is a software program that simulates human interaction with a software system to automatically complete tasks in the field of robotic process automation (RPA).
[0027] A vision-language-action (VLA) model is a multi-modal artificial intelligence model that integrates visual perception, natural language understanding, and action execution.
[0028] It should be noted that the above application scenarios or application examples provided in the embodiments of the present application are for the convenience of understanding, and the application of the technical solutions in the embodiments of the present application is not specifically limited. In addition, the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.
[0029] The prior art lacks effective modeling and reasoning of long action sequences or multi-step combinations, is easily disturbed by different interface layouts, styles, dynamic changes and other factors, has weak generalization ability, and has limited adaptability.
[0030] The technical solutions of the present application and how the technical solutions of the present application solve the foregoing technical problems will be described in detail below with specific embodiments. Several specific embodiments listed can be combined with each other, and the same or similar concepts or processes can not be described again in some embodiments. The embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0031] Figure 1 A flowchart of an action sequence execution method according to an embodiment of the present application is shown in FIG. 1, which can include steps S101-S104. Figure 1
[0032] Step S101: Based on semantic features representing user instructions, a global action sequence is generated, and the global action sequence includes a plurality of node actions with execution order;
[0033] Step S102: After each node action is executed, an exception identification is performed on each node action to obtain an identification result corresponding to each node action;
[0034] Step S103: In the case where the identification result indicates that the target node action has an exception, a target action sequence is generated with the target node action as a starting point; the target action sequence is used to update the global action sequence;
[0035] Step S104: According to the updated global action sequence, a processing is performed on the interface elements in the operation interface until the execution target of the user instruction is completed.
[0036] Exemplarily, the execution subject is in the form of an agent, and the agent can be a multi-action combination VLA (Vision-Language-Action) model or a multi-modal Transformer model for a Web GUI. The agent is used for page operation automation, GUI agent, and the like, adopts an end-to-end visual, language, and action joint modeling manner, and realizes multi-action serialization control of a complex Web page. Optional application scenarios can be, for example, a software robot (RPA robot), a browser plug-in, an application program extension, an intelligent interaction system, an Internet of Things (IoT) device controller, a collaborative robot, and the like.
[0037] In the case of a software robot (RPA robot), a browser plug-in, and an application program extension, an interaction agent for an end user understands user requirements through natural language processing and operates an interface to complete a service. The exceptions of the node actions can include various types, such as a sudden pop-up window (such as a system upgrade notification) in the operation interface, a verification code recognition error, an account lock prompt, an unresponsive state of an operating system or an interactive interface element, a dynamic loading content that does not appear on time, a session timeout that causes an operation permission to be lost, an operation frequency limit that is triggered, an application crash caused by insufficient memory, a response timeout caused by network delay, and the like.
[0038] The intelligent interaction system can be, for example, an intelligent cockpit or an intelligent home control center. The intelligent interaction system is integrated into an interaction hub (such as a vehicle-mounted control screen or a smart home panel) of a hardware device. In the case of an intelligent interaction system, an Internet of Things (IoT) device controller, and a collaborative robot, the exceptions of the action nodes can include various types, such as a device sudden offline prompt, a communication interruption alarm, a device function abnormality prompt, a clock desynchronization between devices, and manual interruption, and the like.
[0039] In the embodiments provided in the present application, the semantic features of the user instruction are analyzed, and based on the semantic understanding of the instruction intention, a global execution sequence containing multiple ordered node actions is automatically generated, and the node actions form a complete operation chain according to the business logic. After completing each node action, real-time monitoring and abnormality identification of the execution result are performed, including but not limited to interface element state verification, operation feedback analysis and expected effect comparison. Abnormality identification is performed after the execution of the node action, which can timely find the problems that may occur in the execution process, enhances the stability and reliability of the system, and avoids the failure or error result of the whole operation due to abnormal conditions. When an abnormality of a target node action is detected, the flow is not simply terminated, but the abnormality reason is intelligently analyzed, and a new action sequence that is re-planned based on the current context environment and takes the abnormal node as a starting point, that is, a target action sequence, is generated. The target action sequence will be used as a local correction scheme to dynamically update the original global execution plan. The introduced dynamic monitoring and re-planning mechanism enables the system to have self-adaptive ability, can flexibly cope with various unexpected abnormal conditions, and greatly improves the robustness of the flow. The subsequent operation is continued according to the updated action sequence, and the target task set by the user instruction is completely implemented.
[0040] According to the embodiments provided in the present application, in step S101: based on the semantic features representing the user instruction, the global action sequence is generated, which can include the following specific steps:
[0041] An initial screenshot of the operation interface is obtained;
[0042] Based on the initial screenshot, the semantic features, and the historical action sequence, the action execution strategy is determined, and the historical action sequence is obtained based on the historical node actions of the completed historical instructions;
[0043] According to the action execution strategy, a reference action is selected in a predetermined action space to generate a global action sequence.
[0044] In the embodiments provided in the application, the initial screenshot of the operation interface is acquired as a visual reference, the intention recognition and parameter extraction of the semantic features of the user instruction are combined, and the prior knowledge of the historical action sequence is referred, the historical action sequence provides a successful action mode in the historical execution record, and the optimal action execution strategy is determined through triple information fusion. The accumulation of the historical action sequence forms a reusable knowledge base, facilitating the continuous learning and optimization of the action strategy of the intelligent agent, and with the increase of the number of uses, the success rate and efficiency of execution will be continuously improved. By analyzing the distribution of interface elements in the initial screenshot, the core demands in the semantic features are understood, and effective experience in the historical action sequence is referred, multi-dimensional information is provided as the basis for decision-making for the decision-making of the intelligent agent, and then the execution strategy is obtained. Based on the execution strategy, the reference action most suitable for the current context is intelligently selected from the predefined action space (such as basic operations such as clicking, inputting, and scrolling), and finally a structured global action sequence is generated, which not only contains necessary operation nodes, but also clearly defines the time sequence execution order between nodes.
[0045] According to the embodiments provided in the application, the method further comprises:
[0046] For a predetermined type of operation, a plurality of reference actions are extracted;
[0047] According to the action types corresponding to the plurality of reference actions, the positioning of each interface element, and / or the reference fill-in content, an action space is generated.
[0048] In the embodiments provided in the application, for a specific type of operation (such as form filling, data query, navigation switching, etc.), the intelligent agent pre-collects and organizes a series of representative reference actions. The above-mentioned reference actions can come from historical execution records, artificially annotated cases or domain expert experience. Based on the extracted reference actions, the intelligent agent is structured and organized, each reference action includes multiple dimensions of information, including the basic type of operation (such as clicking, inputting, selecting, scrolling, etc.), the position information of the interface element recording the action effect in the operation interface, and for the action requiring input data (such as form filling), a typical input example or template is provided. By combining the information of the above-mentioned dimensions, an action space containing possible actions and their parameters is generated, providing a candidate set for the generation of the subsequent global action sequence. The structured action space can integrate the experience knowledge of the domain experts, provide a clear action candidate set, reduce the exploration space, and accelerate the model convergence speed and optimization efficiency.
[0049] Exemplarily, for the agent model, its input includes the GUI screenshot of the current page (i.e., the operation interface) as a visual feature input, a historical action sequence, and a current task text description as a language feature input. The model body of the agent can be a multi-modal Transformer or a large visual-language-action model, which has cross-modal feature alignment and action reasoning capability, and can understand interface element visual features, text semantics, and action intentions.
[0050] The agent output is a selected reference action in the action space, and a normalized action combination sequence is obtained to generate a global action sequence. For example, clicking, double-clicking, dragging, inputting, scrolling, shortcut key operation, etc., are expressed in the form of "action type + coordinate / content", and have a time sequence structure.
[0051] Exemplarily, based on Web page operation, there can be multiple common actions, and the reference actions abstracted in the above action space can include: click (start_box=' (x1, y1) ') representing performing a left mouse button single click at the specified coordinate (x1, y1), start_box representing the coordinate of the click position in the form of a string '(x1, y1)', which is used to trigger buttons, links, drop-down menus, and other interactive elements.
[0052] doubleclick (start_box=' (x1, y1) ') represents performing a left mouse button double click at the specified coordinate (x1, y1), start_box represents the coordinate of the double-click position in the form of a string '(x1, y1)', which is used to open files, edit text, or trigger elements that require double-click activation.
[0053] right_single (start_box=' (x1, y1) ') represents performing a right mouse button single click at the specified coordinate (x1, y1), start_box represents the coordinate of the right-click position in the form of a string '(x1, y1)', which is used to open a context menu or trigger a right menu option.
[0054] drag (start_box=' (x1, y1) ', end_box=' (x2, y2) ') represents dragging from the starting coordinate (x1, y1) to the target coordinate (x2, y2), start_box represents the starting coordinate of the drag in the form of a string '(x1, y1)', end_box represents the ending coordinate of the drag in the form of a string '(x2, y2)', which is used for drag sorting, adjusting element size, or performing operations that require dragging (such as slider control).
[0055] hotkey (key='') can execute keyboard shortcuts, key represents hotkey, format is string, quickly execute system or application functions, replace mouse operation, for example, the shortcut key for refreshing the page is F5, and the shortcut key for copying and pasting is Ctrl + C / V.
[0056] type (content='', start_box='(x1,y1)') indicates inputting text content at the specified coordinates (x1,y1), content represents the text content to be input, format is string, start_box represents the coordinates of the input position, format is string '(x1,y1)', used to fill in information in input box, text area, etc.
[0057] scroll down () indicates scrolling down the page, scroll up () scrolls up the page, scroll left () scrolls left the page, scroll right () scrolls right the page, used for browsing long pages.
[0058] scrollmenu (start_box='(x1,y1,x2,y2)') indicates scrolling in the specified area (x1,y1,x2,y2) (such as drop-down menu, list), start_box represents the coordinates of the top left and bottom right of the scrolling area, format is string '(x1,y1,x2,y2)', browse content in limited area (such as selector, scroll list).
[0059] finish () is used to mark the completion of the task and terminate the current operation sequence.
[0060] wait () is used to pause execution and wait for specific conditions, such as interface element loading, system response waiting time.
[0061] call_user () is used to pause and request user intervention, such as entering verification code, selecting preference content by user, performing face recognition to obtain operation permission.
[0062] stop (reason='') indicates terminating the task due to specified reasons, such as error occurrence, which can record operation log by recording text in reason parameter, add current execution state (such as executed actions, current interface screenshot, variable value), facilitate subsequent problem troubleshooting.
[0063] The above reference actions can be combined into an ordered sequence to achieve complex high-order task flow, in order to facilitate understanding, taking searching for an item as a task for example:
[0064] click('(100,200)') / / Click the search box;
[0065] type('item', '(100,200)') / / Enter the search term;
[0066] click('(300,200)') / / Click the search button;
[0067] scrolldown() / / Scrolls the page down;
[0068] wait() / / Wait for search results to load;
[0069] click('(200,400)') / / Click the first search result;
[0070] finish() / / Complete the task.
[0071] In the above embodiments, the positioning of interface elements is represented in coordinate form. In order to improve the recognition ability of interface elements, the positioning of interface elements can also be represented in the form of region boxes.
[0072] According to the embodiments provided in this application, the method may further include the following specific steps:
[0073] Based on the visual features of the user interface and the semantic features of user commands, fused features are obtained.
[0074] Based on the fusion features, a region bounding box is generated for each interface element as its location. The region bounding box is used to represent the interactive area of the corresponding interface element.
[0075] In the embodiments provided in this application, visual features of the user interface are extracted using computer vision technology, including but not limited to low-level features such as color distribution, texture patterns, and shape contours. Simultaneously, natural language processing results from the semantic feature parsing of user commands are combined to fuse these two types of heterogeneous features, resulting in fused features. Based on these fused features, the agent intelligently identifies interactive elements in the interface, generating interactive boundary labels for each element, i.e., obtaining a bounding box for each element. These bounding boxes not only define the physical location of the element in the screen coordinate system but also define the interactive range. Clicking any coordinate within the bounding box successfully enables interaction with that element. This approach demonstrates good adaptability to the labeling requirements of small icons in the user interface, improving the robustness of the model.
[0076] By analyzing the visual features of the operation interface (such as element shape, color, texture, layout relationship) and the semantic features of the user instruction (such as keywords, intent categories), the two pieces of information are fused to generate more comprehensive fusion features. The semantic intent and the interface visual element can be associated, for example, the semantic of "clicking the search button" is linked to the visual area of the search icon or the "search" text label in the interface. In the face of interface style mutations or element changes, the recognition accuracy can still be maintained.
[0077] Exemplarily, the visual features can be extracted into high-dimensional vectors by a convolutional neural network, the semantic features are encoded into semantic space representations via a language model, and cross-modal alignment is realized through a feature cross-attention mechanism to finally generate the fusion feature representation.
[0078] According to some embodiments provided in the present application, the above-mentioned generation of the region frame corresponding to each interface element as positioning according to the fusion feature can include the following steps:
[0079] Based on the fusion feature, the initial corner point coordinates of each interface element are predicted.
[0080] The initial corner point coordinates of each interface element are normalized to obtain the target corner point coordinates of each interface element, and the target corner point coordinates are at least two, which are used to anchor the range of the region frame.
[0081] In the embodiments provided in the present application, the output layer of the agent model includes multiple regression heads, and each 2 regression heads are specially responsible for predicting the initial coordinates of one type of corner point. Since one coordinate includes the horizontal coordinate x and the vertical coordinate y, for a rectangular region frame, there are 2 types of corner points (upper left and lower right), and therefore 2 regression heads are used to predict the horizontal coordinate and the vertical coordinate of the upper left corner point. Similarly, 2 regression heads are also used to predict the lower right corner point, allowing the model to learn the features of each corner point respectively, thereby improving the prediction accuracy. Each regression head predicts the initial coordinates of the corresponding corner point based on the input fusion feature. The above-mentioned initial corner point coordinates can be regarded as the original prediction value without normalization processing. The normalization processing makes the output of the model independent of the specific size of the input image, thereby improving the generalization ability of the model. According to the multiple target corner point coordinates after normalization, the region frame of each interface element is constructed, and the target corner point coordinates define the shape and position of the region frame, thereby framing the interactive region of the current interface element.
[0082] Exemplarily, a fully connected layer is arranged in the agent model, which is used to further map the initial corner point coordinates of each corner point to the final target corner point coordinates. In the mapping process, the fully connected layer normalizes the initial corner point coordinates to a predetermined size range.
[0083] According to some embodiments provided in the present application, the global action sequence is composed of process actions, and the plurality of node actions belong to the process actions included in the global action sequence. Among all the process actions, there are some actions that implement the key role of user instructions, namely node actions. For the plurality of node actions included in the global action sequence in step S101, the selection method can be realized by the following steps, and the method further comprises:
[0084] The node recognition model is used to determine the process actions for changing the execution state among the plurality of process actions as the plurality of node actions, and the node recognition model is obtained by training based on a plurality of historical action sequences with historical nodes marked.
[0085] In the embodiments provided in the present application, the node recognition model obtained by pre-training based on a plurality of historical action sequences with historical nodes marked is used to analyze the plurality of process actions contained in the global action sequence, and the process actions capable of changing the execution state are identified therefrom and determined as the plurality of node actions. These node actions are the core steps of the global action sequence for implementing user instructions, and the key nodes are screened from the process actions by model training, ensuring the effectiveness and pertinence of subsequent action anomaly recognition. By eliminating non-key process actions, the global action sequence is simplified, the number of actions that need to be monitored and analyzed is reduced, the complexity of subsequent anomaly recognition is reduced, and the abnormality calculation overhead in the action process is reduced.
[0086] Exemplarily, the above node recognition model can be continuously iteratively updated according to historical data, the recognition rules of the node actions are continuously optimized according to a predetermined period, the global action sequence is dynamically adjusted according to changes in business scenarios, and the above changes in business scenarios can represent changes in business rules, including interface revision and process adjustment, and the efficient execution capability of user instructions is maintained.
[0087] According to some embodiments provided in the present application, based on the heterogeneity of node actions, different node actions can be classified, the plurality of node actions correspond to different abnormal patterns, and the subsequent abnormal recognition processing of the node action can be performed according to a certain abnormal pattern. In step S102, the abnormality of each node action is recognized to obtain the recognition result corresponding to each node action, which can include the following specific steps:
[0088] Obtaining a real-time screenshot of the operation interface under the current execution state;
[0089] Based on the real-time screenshot, the abnormality of each node action is recognized according to the recognition strategy to determine the recognition result corresponding to each node action. The recognition strategy is determined based on the predetermined abnormal pattern corresponding to the plurality of node actions, and the predetermined abnormal pattern is associated with the operation type of the corresponding node action.
[0090] In some embodiments provided in the present application, a real-time screenshot of the operation interface in the current execution state is acquired, and the interface state change can be captured instantly based on the real-time screenshot, and the abnormality can be quickly found, which is more intuitive and efficient than the traditional code or log-based detection method. Abnormality recognition is performed on the node action based on the recognition strategy determined according to the predetermined abnormality mode associated with the operation type of the node action, so as to determine the corresponding recognition result. Different node actions adopt customized recognition strategies, which reduces unnecessary detection steps and reduces computational overhead. In the interactive process of the multi-action combination, the agent automatically customizes the abnormality mode according to the characteristics of each node action, which can cope with various abnormal situations in complex business scenarios and ensure the stability of the automated task.
[0091] Specifically, the agent can adapt to real-time detection of interface changes and dynamically adjust the operation path, and can handle scenarios such as element position offset and business rule mutation.
[0092] Exemplarily, the types of the node actions can include various types such as click action, input action, drag action, scroll action, judgment action, flow control action, element interaction operation, hotkey operation, and the like, which will be described one by one below.
[0093] The abnormality mode of the click action includes the following: interface element not found, click non-response, and expected effect after clicking not appearing. The element not found can be caused by the update of the operation interface, so that the interface element does not exist in the DOM (Document Object Model) of the interface, or the interface element can be covered by a pop-up window or a floating layer. The click non-response can be caused by the disabled state or the non-clickable attribute, so that the interface element cannot be interacted with, or the loading of the interface element is not completed. The expected effect after clicking not appearing can be caused by the problem that the page is not jumped due to the non-reaction of the click link, the modal box is not displayed due to the non-triggering of the modal box by the click button, or the state is not updated due to the non-switching of the selected state of the check box.
[0094] The abnormality mode of the input action includes the following: input box not available and input content not meeting the expectation. The input box not available can be caused by the attribute setting problem, so that the input box is prohibited from being used, or is in the read-only state without input permission.
[0095] The abnormal patterns of the drag type action include abnormal start or end element, interruption during dragging, and unexpected result. The abnormal start or end element can be caused by the fact that some interface elements do not support the drag operation, for example, the start / end element does not exist or is invisible. The interruption during dragging can be caused by the position offset due to page scrolling, which makes the dragging distance exceed the threshold. The unexpected result can include that the dragged element is not moved to the target or that the action effect is not triggered after the dragging.
[0096] The abnormal patterns of the scroll type action include no effect, wrong position, and timeout. The no effect can be caused by the failure of the driver to interact with the page, which makes the scroll operation not triggered. The wrong position can be caused by the fact that the target element is not in the viewport due to the fact that the target element is not scrolled to or the fact that the distance of the scroll exceeds the expectation. The timeout can be caused by the slow page loading, which makes the scroll wait timeout.
[0097] The abnormal patterns of the judgment action include timeout and condition judgment failure. The timeout can be caused by the fact that the interface element does not appear or refresh within the specified time. The condition judgment failure can include the fact that the expected interface element does not exist, for example, the drop-down menu should automatically retract but still remains in the drop-down state, which makes the interface elements below it not be recognized. Or the fact that the state of the interface element does not meet the expectation, for example, the interface element is in an uninteractive state.
[0098] The abnormal patterns of the flow control action include interruption, for example, the interface element is not responded due to manual intervention or network interruption.
[0099] The abnormal patterns of the element interaction operation include file upload failure, drop-down failure, and pop-up window processing failure. The file upload failure can be caused by the fact that the file path does not exist or the fact that the upload times out. The drop-down failure can be caused by the fact that the drop-down menu is not correctly expanded, the fact that the drop-down menu is empty, or the fact that the drop-down menu fails to perform the predetermined operation (such as sorting or filtering the predetermined content). The pop-up window processing failure can include the fact that the pop-up window cannot be closed or the fact that the content prompt information of the pop-up window is incorrect.
[0100] The abnormal patterns of the hot key operation include the fact that the page intercepts the default keyboard event, for example, the copy and paste of the mailbox or password are disabled. Or the fact that the character input is abnormal, for example, the input switching is abnormal.
[0101] In some embodiments provided in the application, in step S103, a target action sequence starting from the target node action is generated, including:
[0102] determining the coping action corresponding to the predetermined abnormal pattern of the target node action;
[0103] According to the coping action, a target action sequence is generated.
[0104] In the embodiments provided in the present application, when it is identified that the target node action is abnormal, the agent can match the corresponding coping action from the preset coping strategy library according to the predetermined abnormal mode corresponding to the target node action. The above-mentioned coping action is a repair or retry logic designed for a specific abnormal mode. The agent takes the target node action as a starting point, organizes the matched coping action into an ordered execution sequence to form a target action sequence. The target action sequence is used to replace or supplement the problem part in the original global action sequence, so as to ensure that the process can continue to execute after the exception is processed.
[0105] Through the above processing of the generation of the target action sequence, the system can dynamically adjust the execution path according to the real-time abnormal situation, and the flexibility and adaptability of the automatic process are enhanced. The mapping relationship between the abnormal mode and the coping action is centrally managed, and when the business scenario changes, only the coping strategy library needs to be updated, without the need to modify the entire automatic process, thereby improving the code maintainability.
[0106] Exemplarily, by recording the frequency of occurrence of the exception and the success rate of the coping action, the exception handling strategy can be continuously optimized, and the demand for manual intervention can be gradually reduced.
[0107] Exemplarily, the above-mentioned setting of the application action can be based on the following modes: preferentially trying to retry, waiting, and other lightweight repairs, adding a confirmation mechanism for irreversible operations (such as submitting a form), saving the context data (screenshot, log, state variable) at the time of exception, so as to facilitate the backtracking and optimization of the agent.
[0108] The application action corresponding to the above-mentioned example of the abnormal mode is described.
[0109] For the interface element not found, the coping action can include the following modes: retrying to locate the element (increasing the waiting time); checking whether there is a pop-up window / floating layer on the interface and closing it; triggering page refresh and repositioning; calling the OCR (Optical Character Recognition) technology to identify the element position in the screenshot to replace the DOM positioning.
[0110] For the click non-response, the coping action can include the following modes: checking whether the element is in a disabled state, obtaining administrator permission, and trying to clear the disabled attribute; waiting for the element to be loaded.
[0111] For the expected effect not appearing after clicking, the coping action can include the following modes: checking whether the page URL changes, retrying to click or switching the link if it does not jump; waiting for a pop-up window / modal box to appear (setting a timeout threshold); verifying whether the interface element state is updated; and executing the clicking process again after falling back to the page.
[0112] For input box unavailable, the response action can include the following ways: get permission by login, or remove the disabled or read-only attribute of the input box with administrator permission; check if the pre-operation (such as checking the agreement) needs to be completed first; if the input box is hidden, try to scroll to the visible area and then operate.
[0113] For input content not meeting expectations, the response action can include the following ways: verify if the input format matches the regular expression; re-enter after clearing the original content in the input box; check if there is automatic filling interference and clear the cache.
[0114] For abnormal start or end element of drag, the response action can include the following ways: handle element offset caused by page scrolling, scroll to the target position first and then drag, or dynamically adjust the drag path to compensate for the scrolling offset.
[0115] For drag interruption, the response action can include the following ways: increase the timeout time of the drag operation; segment the drag (such as first dragging to the intermediate point and then dragging to the end); detect if there are obstacles (such as pop-up windows) in the drag path that cause early closure; use explicit waiting to ensure that the element is interactive; record the current drag progress and continue from the breakpoint after interruption.
[0116] For drag result not meeting expectations, the response action can include the following ways: verify if the target position coordinates are correct, such as comparing the screenshot pixel coordinates; check if additional events need to be triggered after dragging; rollback the page state and re-execute the drag.
[0117] For scroll with no effect or incorrect position, the response action can include the following ways: first click on the blank area of the page to get focus and then scroll; replace one-time scrolling with multiple small-scale scrolling; calculate the offset of the target element relative to the viewport; disable the infinite scrolling function; record the page height before scrolling and rollback in case of abnormality; combine OCR to identify if the page content after scrolling matches the expected result.
[0118] For scroll timeout, the response action can include the following ways: increase the page load timeout threshold; check the network status and switch to stable network.
[0119] For waiting timeout, the response action can include the following ways: extend the waiting time or switch the network; use event listening to check the interface element state; trigger page refresh and then wait again.
[0120] For condition judgment failure, the response action can include the following ways: reacquire the interface element state; compare the interface screenshot to confirm if the element really exists; check if the business logic requires preconditions (such as login status).
[0121] For process interruption exceptions, the coping actions can include the following ways: automatically retry the interrupted process, set the maximum number of retries; send a notification to remind manual intervention, save the current process state, continue running from the breakpoint next time, record the interface screenshot and log at the time of interruption, and facilitate manual troubleshooting.
[0122] For file upload failures, the coping actions can include the following ways: verify the validity of the file path; check the file format / size limit, compress or convert the format; use chunked upload instead of one-time upload.
[0123] For pull-down failures, the coping actions can include the following ways: repeat the attempt to interact with the interface element according to the predetermined number of retries; check if the options are dynamically loaded and wait for data to return.
[0124] For pop-up processing failures, the coping actions can include the following ways: reopen the page after closing the browser tab; click on the blank page to trigger the pop-up to automatically close.
[0125] For page keyboard event interception, the coping action can be to use a virtual keyboard plugin to input special characters.
[0126] For character input exceptions, the coping actions can include the following ways: restart the input method; clear the default value in the input box and then input; verify if there is an input length limit and truncate the content.
[0127] According to the above embodiments and optional embodiments, the application further provides an optional implementation, Figure 2 The action sequence execution method of the embodiment of the application is shown in a schematic diagram, as Figure 2 As shown, taking the user instruction "click on the yellow braised chicken and rice below 30, and want speed and high score" as an example, the intelligent agent executes the processing flow in the takeout application, extracts semantic constraints through the VLA model, determines that the dish type is yellow braised chicken and rice, the price needs to be less than or equal to 30 yuan, and the sorting priority is speed greater than score.
[0128] According to the historical action sequence, the intelligent agent identifies 5 key node actions, including: node 1: sorting, node 2: selecting a store, node 3: filtering takeout, node 4: ordering, and node 5: paying.
[0129] For node 1 sorting, the execution action is to sort by speed and score, there are nodes that interact with pop-up type and pull-down type menus, and the abnormal mode is that the pull-down menu is not automatically retracted, blocking part of the store display. The coping action is to analyze the visual blocking of the menu retraction, or DOM node coverage detection, to determine whether a certain page element is visually blocked or layout overlaid by other elements. If an exception occurs, the action of clicking on the blank interface to retract the menu can be added.
[0130] For node 2, the action is to enter the target store page of Huangmianjimesifan, and the abnormal mode is that the store is closed or the store information suddenly changes, such as the delivery price suddenly changes to 40 yuan. The response action can be to determine the store operating status based on the operating status icon identification combined with the OCR text prompt reading, and to determine the store information change through the DOM value extraction combined with the logic judgment of the predetermined constraint condition.
[0131] For node 3, the action is to select the filter label below 30 yuan, and the abnormal mode includes filter empty set, filter condition drift, and actual filter exceeding the acceptable price range. The response action can be to count the list items or perform visual detection for the filter result empty set, and to check the numerical range of the price through OCR recognition for the condition drift anomaly.
[0132] For node 4, the action is to submit an order after adding it to the shopping cart, and the abnormal mode includes sudden out-of-stock of goods / other goods in the shopping cart. The application action can be to determine the inventory anomaly through the state of the selectable button combined with visual detection, trigger action backtracking, and reselect the store; for the problem of historical goods in the shopping cart, directly get the cart list whether it is an empty set.
[0133] For node 5, the action is to select the default payment method or preferred payment method, and the abnormal mode can be default payment balance insufficient / payment interface loading timeout. The response can be to identify the payment obstacle problem through error prompt text recognition or payment button state detection; for the timeout problem, determine through the predetermined timer combined with the loading animation recognition method.
[0134] In the above real-time exception cases, the initial planned process is no longer applicable, and the process needs to be re-planned to replace the subsequent processing method, that is, for different node exceptions, adaptively add steps to solve the problem. Only the affected node and subsequent steps are modified, such as node 1 exception, insert the drop-down menu recall action and continue the original process.
[0135] All the above steps are executed by selecting actions in the predetermined standard action space, including steps in the historical action sequence, which are also standardized actions in the form of: action type + coordinates / content. And the above action execution can increase the reliability of the operation process by using the predicted box of GUI elements.
[0136] Through the above processing, the multi-step task processing efficiency in the Web automation scene is greatly improved, the generality and flexibility of the agent are significantly improved, various page changes and dynamic content are adapted, the robustness and generalization ability are strong, and the agent is easy to migrate to different Web systems. By using the multi-modal deep model, the complex GUI information can be understood, and reliable human-machine interface process automation can be realized. The operation output is standardized, which is convenient for extension and connection with other automation platforms, and supports subsequent secondary development and custom operation and maintenance.
[0137] Figure 3 A block diagram of an action sequence execution device of an embodiment of the present application is shown as Figure 3 Corresponding to the application scenarios and methods of the method provided by the embodiments of the present application, the embodiments of the present application further provide an action sequence execution device, which comprises:
[0138] The generation module 301 is configured to generate a global action sequence based on semantic features representing user instructions, the global action sequence comprising a plurality of node actions with execution sequences.
[0139] The anomaly identification module 302 is configured to identify each node action after each node action is executed, to obtain an identification result corresponding to each node action.
[0140] The re-planning module 303 is configured to generate a target action sequence with a target node action as a starting point in a case where the identification result indicates that the target node action is abnormal; and the target action sequence is used to update the global action sequence.
[0141] The update module 304 is configured to execute processing on interface elements in an operation interface according to the updated global action sequence until the execution target of the user instructions is completed.
[0142] The functions of each module in each device of the embodiments of the present application can be referred to the corresponding description in the above method, and have the corresponding beneficial effects, which will not be repeated here.
[0143] Figure 4 A block diagram of an electronic device for implementing the embodiments of the present application is shown as Figure 4 The electronic device comprises a memory 401 and a processor 402, and the memory 401 stores a computer program capable of running on the processor 402. The processor 402 implements the method in the above embodiments when executing the computer program. The number of the memory 401 and the processor 402 can be one or more. In specific implementation, the electronic device can further comprise a communication interface 403 for communicating with external devices and transmitting data.
[0144] In a specific implementation, if the memory 401, the processor 402 and the communication interface 403 are independently implemented, the memory 401, the processor 402 and the communication interface 403 can be connected with each other through a bus and complete communication between each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For convenience of representation, Figure 4 Only one thick line is used to represent the bus in the figure, but it does not mean that there is only one bus or only one type of bus.
[0145] Optionally, in a specific implementation, if the memory 401, the processor 402 and the communication interface 403 are integrated on a chip, the memory 401, the processor 402 and the communication interface 403 can complete communication between each other through an internal interface.
[0146] The embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the method provided in the embodiment of the present application.
[0147] The embodiment of the present application provides a computer program product, which includes a computer program, and the program is executed by a processor to implement the method provided in the embodiment of the present application.
[0148] The embodiment of the present application further provides a chip, which includes a processor, is used for calling and running instructions stored in a memory from the memory, and makes a communication device installed with the chip execute the method provided in the embodiment of the present application.
[0149] The embodiment of the present application further provides a chip, which includes an input interface, an output interface, a processor and a memory, the input interface, the output interface, the processor and the memory are connected through an internal connection path, and the processor is used for executing code in the memory, and when the code is executed, the processor is used for executing the method provided in the embodiment of the present application.
[0150] It is to be understood that the above-described processor can be a Central Processing Unit (CPU), but can also be other general purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field Programmable Gate Arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic components, discrete hardware components, or the like. The general purpose processor can be a microprocessor or any conventional processor, or the like. It is to be appreciated that the processor can be an Advanced RISC Machines (ARM) architecture processor.
[0151] Further, the memory can optionally include a read-only memory and a random access memory. The memory can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memory. The non-volatile memory can include a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory, among others. The volatile memory can include a Random Access Memory (RAM), which is used as an external cache. By way of example, and not limitation, many forms of RAM are available. For example, a Static Random Access Memory (SRAM), a Dynamic Random Access Memory (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Sync Link DRAM (SLDRAM), and a Direct Rambus RAM (DR RAM), among others.
[0152] In the above-described embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded on a computer, all or part of the processes or functions according to the present disclosure are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium.
[0153] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. Also, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, a person skilled in the art can combine and combine the different embodiments or examples described in the present specification and the features of the different embodiments or examples without contradiction.
[0154] In addition, the terms "first", "second", etc. are used only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the technical features indicated. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present disclosure, the meaning of "a plurality of" is two or more, unless otherwise explicitly specified.
[0155] Any process or method described in the flowchart or otherwise described herein can be understood as a representation of code including one or more executable instructions for performing a specific logical function or process. Also, the scope of the preferred embodiments of the present disclosure includes additional implementations, in which the functions can be performed in an order different from that shown or discussed, including functions performed in a substantially simultaneous manner, or in reverse order.
[0156] The logic and / or steps described in the flowchart or otherwise described herein, for example, can be considered a list of executable instructions for implementing logic functions, and can be specifically embodied in any computer-readable medium for use by an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other system that can take instructions from an instruction execution system, apparatus, or device, or in conjunction with such instruction execution system, apparatus, or device.
[0157] It should be understood that each part of the present application can be realized by hardware, software, firmware or a combination thereof. In the above embodiments, a plurality of steps or methods can be realized by software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the above-mentioned embodiment methods can be completed by a program instructing the relevant hardware, which can be stored in a computer readable storage medium and includes one or a combination of the steps of the embodiment methods when executed.
[0158] In addition, each functional unit in each embodiment of the present application can be integrated in one processing module, or each unit can be physically present separately, or two or more units can be integrated in one module. The above-mentioned integrated module can be realized in the form of hardware or in the form of a software functional module. The above-mentioned integrated module, if realized in the form of a software functional module and sold or used as an independent product, can also be stored in a computer readable storage medium. The storage medium can be a read-only memory, a magnetic disk or an optical disk, etc.
[0159] The above is only an exemplary embodiment of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of various changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for executing an action sequence, characterized in that, include: Based on the semantic features representing user instructions, a global action sequence is generated, which includes multiple node actions that have an execution order; After each node action is executed, anomaly identification is performed on each node action to obtain the identification result corresponding to each node action; When the identification result indicates that the target node action is abnormal, a target action sequence is generated starting from the target node action. The target action sequence is used to update the global action sequence. According to the predetermined abnormality pattern corresponding to the target node action, the corresponding response action is matched from the preset response strategy library. The matched response actions are organized into an ordered execution sequence to form the target action sequence. The response action is generated by the repair or retry logic designed for the predetermined abnormality pattern. The target action sequence is used to replace or complete the abnormal part in the global action sequence. Based on the heterogeneity of node actions, different node actions are classified. The multiple node actions correspond to different abnormality patterns. The predetermined abnormality pattern is associated with the operation type of the corresponding node action. For click-related actions, abnormal patterns include: interface element not found, no response to click, and no expected effect after clicking; for input-related actions, abnormal patterns include: input box unavailable, input content not meeting expectations; for drag-related actions, abnormal patterns include: drag start or end element abnormal, drag process interrupted, and drag result not meeting expectations. The exception patterns for scrolling actions include: no scrolling effect, incorrect scrolling position, and scrolling timeout; the exception patterns for judgment actions include: waiting timeout and condition judgment failure; the exception patterns for flow control actions include: flow interruption exception; the exception patterns for element interaction operations include: file upload failure, drop-down failure, and pop-up handling failure; by recording the frequency of exception occurrences and the success rate of response actions, the exception handling strategy can be optimized. The node actions and the response actions are optional actions in a standard action space; the method further includes: obtaining fusion features based on the visual features of the operation interface and the semantic features of the user instructions; generating a region box corresponding to each interface element as a location based on the fusion features, the region box being used to represent the interactive area of the corresponding interface element; the region box marking the physical position of the interface element in the screen coordinate system and selecting the interactive range, clicking any coordinate in the interactive range to perform interactive processing on the interface element; According to the updated global action sequence, the interface elements in the operation interface are processed until the execution target of the user instruction is completed; The method further includes: For each type of operation, extract multiple reference actions; The action space is generated according to the action type corresponding to the multiple reference actions, the position of each interface element, and / or the reference input content.
2. The method according to claim 1, characterized in that, The generation of a global action sequence based on semantic features representing user instructions includes: Obtain an initial screenshot of the operation interface; Based on the initial screenshot, the semantic features, and the historical action sequence, an action execution strategy is determined, wherein the historical action sequence is obtained based on the historical node actions that complete historical instructions; According to the action execution strategy, a reference action is selected in the predetermined action space to generate the global action sequence.
3. The method according to claim 1, characterized in that, The multiple node actions belong to the process actions included in the global action sequence, and the method further includes: A node recognition model is used to identify process actions that change the execution state among multiple process actions, which are referred to as the multiple node actions. The node recognition model is trained based on multiple historical action sequences labeled with historical nodes.
4. The method according to any one of claims 1 to 3, characterized in that, The multiple node actions each correspond to different anomaly patterns. Anomaly identification is performed on each node action to obtain the identification result corresponding to each node action, including: Get a real-time screenshot of the operation interface in the current execution state; Based on the real-time screenshot, anomaly identification is performed on each node action according to the identification strategy to determine the identification result corresponding to each node action; the identification strategy is determined based on the predetermined anomaly patterns corresponding to the multiple node actions respectively, and the predetermined anomaly patterns are associated with the operation type of the corresponding node action.
5. An electronic device comprising a memory, a processor, and a computer program stored in the memory, wherein the processor, when executing the computer program, implements the method of any one of claims 1 to 4.
6. A computer-readable storage medium storing a computer program that, when executed by a processor, implements the method of any one of claims 1 to 4.
7. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Voice interaction method and device, and storage medium
CN119580707A
Webpage action execution method and device, electronic equipment and storage medium
CN119719553A
Automatic test script dynamic generation method and system based on multi-modal AI identification
CN120011247A
APP automatic testing method based on multi-modal perception and Agent system
CN120086109A