Web page action execution method, device, electronic device and storage medium

By inputting the HTML text data of the web page and the page screenshots into the visual language model, determining the functional node tree and semantic matching, the problem of handling complex hierarchical operations and page state tracking in the prior art is solved, and the accuracy of operation element recognition and the efficiency of automation tasks are improved.

CN119719553BActive Publication Date: 2025-06-13SHANDONG INSPUR SCI RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510238447.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-13
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

The prior art has limited capabilities in handling dynamic modeling of web page structure, and it is difficult to effectively handle complex hierarchical operations and page state tracking. Due to excessive reliance on HTML text parsing, visual features such as icon position and button style are ignored, resulting in the accuracy of operation element recognition.

Method used

By inputting the HTML text data of the target website's web page and page screenshots into the visual language model, the results of the structured analysis of the page function are obtained, the function node tree is determined, and the task target vector corresponding to the task text is semantically matched with the function description vector to obtain the candidate operation path. The web page action is executed based on the candidate operation path and the execution result is status verified. If the verification fails, the functional node tree is backtracked to obtain a more accurate operation path.

Benefits of technology

It improves the accuracy of element recognition in the structured analysis results of page functions, can effectively handle complex hierarchical operations, and improves the accuracy and reliability of page state tracking corresponding to the operation path, ensuring that the user's natural language instructions can be accurately converted into structured operation intentions, thereby improving the efficiency of automated tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119719553B_ABST
    Figure CN119719553B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and provides a method, device, electronic device and storage medium for executing web page actions. The method includes: determining a function node tree based on the structured analysis result of page functions, semantically matching the task target vector corresponding to the task text with the function description vector to obtain a candidate operation path; based on the candidate operation path and the matching degree sorting corresponding to the candidate operation path, executing the web page action. When the status verification result corresponding to the web page action fails to pass, performing path backtracking on the function node tree to obtain an operation path, and executing the web page action based on the operation path. When the execution result corresponding to the web page action passes the verification, displaying the execution result of the web page action. Determining the function node tree based on the structured analysis result of page functions can effectively handle complex hierarchical operations. The operation path is obtained by semantic matching of the task target vector and the function description vector, improving the accuracy of page status tracking corresponding to the operation path.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method, device, electronic device and storage medium for executing web page actions. Background Art

[0002] Web automation technology has been continuously developing in recent years, but still faces some technical limitations. Traditional web automation tools (such as Selenium, Puppeteer) mainly define operation steps and element positioning rules by manually writing scripts. This method has significant limitations when facing web page dynamic changes (such as element ID changes and asynchronous loading content), and scripts need to be frequently updated to adapt to the functional changes of the web page.

[0003] Rule-based intelligent operating systems encounter challenges when processing complex user instructions. Such systems usually need to pre-define detailed task decomposition logics and operation rules, and it is difficult to accurately execute tasks that require multiple conditional judgments (such as "book a hotel with the cheapest price and a rating of 4 stars or above"). At the same time, these systems have limited understanding of web page functions and there is a certain degree of uncertainty in cross-page operation path planning.

[0004] Existing large model application solutions attempt to use large models to parse user instructions and generate operation codes, but still face technical problems in practice. These solutions have limited capabilities in dynamically modeling web page structures, and it is difficult to effectively handle complex hierarchical operations and page state tracking. Due to excessive reliance on HTML text parsing, these solutions often ignore visual features such as icon positions and button styles, resulting in affected accuracy of operation element recognition. Summary of the Invention

[0005] The present invention provides a method, device, electronic device and storage medium for executing web page actions to solve the defects in the prior art that have limited capabilities in dynamically modeling web page structures, are difficult to effectively handle complex hierarchical operations and page state tracking, and furthermore, due to excessive reliance on HTML text parsing, these solutions often ignore visual features such as icon positions and button styles, resulting in affected accuracy of operation element recognition.

[0006] The present invention provides a method for executing web page actions, including the following steps:

[0007] Input the HTML text data of the web page of the target website and the page screenshot of the web page into a vision-language large model to obtain a page function structured analysis result output by the vision-language large model;

[0008] Based on the page function structured analysis result, a function node tree is determined, and a task target vector corresponding to the task text is semantically matched with a function description vector to obtain a candidate operation path; the function description vector is obtained by converting the function description text in the function node tree;

[0009] Based on the candidate operation paths and the matching degree rankings corresponding to the candidate operation paths, a web page action is executed, and a status verification is performed on the execution result corresponding to the web page action to obtain a status verification result. When the status verification result is a verification failure, the function node tree is backtraced to obtain a first operation path. Based on the first operation path, a web page action is executed. When the execution result corresponding to the web page action is a verification success, the execution result of the web page action is displayed.

[0010] According to a web page action execution method provided by the present invention, the state verification of the execution result corresponding to the web page action to obtain the state verification result includes:

[0011] The task text, the current page screenshot corresponding to the execution result and the function description text are input into the multimodal large model to obtain the state verification result output by the multimodal large model.

[0012] According to a webpage action execution method provided by the present invention, when the state verification result is verification failure, the function node tree is backtracked to obtain a first operation path, including:

[0013] When the status verification result is verification failure, determining a node function description set of the function node tree and a tree structure level depth;

[0014] Starting from the leaf node corresponding to the hierarchical depth of the tree structure, the functional description of the parent node is revised upward layer by layer to obtain the first operation path corresponding to the revised target functional node tree.

[0015] According to a web page action execution method provided by the present invention, starting from the leaf node corresponding to the hierarchical depth of the tree structure, the function description of the parent node is revised upward layer by layer to obtain the first operation path corresponding to the revised target function node tree, including:

[0016] Perform conflict detection on the function descriptions of each sub-node in the function node tree, and input the conflicting first sub-node function description, the second sub-node function description, the HTML text data, and the page screenshot into the visual language model to obtain the target analysis result output by the visual language model;

[0017] Modify the functional node tree based on the target analysis result, and convert the specific operation function description in the functional description of the child nodes in the functional node tree into a task target description to obtain a first functional node tree;

[0018] Delete the duplicate child nodes and / or invalid child nodes in the first functional node tree to obtain a second functional node tree;

[0019] Based on the user task execution frequency, perform priority sorting on the node function descriptions of the second functional node tree to obtain a third functional node tree;

[0020] Input the functional descriptions of all child nodes under the same parent node in the third functional node tree into a multi-modal large model for semantic aggregation to obtain a high-level semantic description output by the multi-modal large model, and use the high-level semantic description as the functional description of the parent node to obtain a first operation path corresponding to the target functional node tree.

[0021] According to a web page action execution method provided by the present invention, before the step of performing status verification on the execution result corresponding to the web page action to obtain a status verification result, it further includes:

[0022] Obtain the page status of the web page and the waiting operation duration;

[0023] In the case where the page status is the unchanged status or the waiting operation duration is greater than the first threshold, perform path backtracking on the functional node tree to obtain a second operation path.

[0024] According to a web page action execution method provided by the present invention, the step of determining the page status includes:

[0025] Obtain the current page and the previous page of the current page;

[0026] Determine a first hash value of the current page and a second hash value of the previous page;

[0027] Determine the hash comparison result of the first hash value and the second hash value, and in the case where the hash comparison result is less than the second threshold, determine the page status as the unchanged status, and in the case where the hash comparison result is greater than or equal to the second threshold, determine the page status as the changed status.

[0028] According to a web page action execution method provided by the present invention, the method further includes:

[0029] In the case where the status verification result is verification failed, obtain the executed historical operation path;

[0030] Determine adjacent nodes adjacent to each node in the historical operation path from the function node tree, and regenerate a third operation path based on the adjacent nodes.

[0031] According to a web page action execution method provided by the present invention, in the case where the status verification result is verification failed, performing path backtracking on the function node tree to obtain a first operation path, including:

[0032] In the case where the status verification result is verification failed, perform path backtracking on the function node tree based on the node confidence of the function node tree to obtain the first operation path.

[0033] The present invention also provides a web page action execution device, including the following modules:

[0034] An input unit, configured to input HTML text data of a web page of a target website and a page screenshot of the web page into a vision-language large model, and obtain a page function structured analysis result output by the vision-language large model;

[0035] A semantic matching unit, configured to determine a function node tree based on the page function structured analysis result, and perform semantic matching between a task target vector corresponding to a task text and a function description vector to obtain a candidate operation path; the function description vector is obtained by converting function description text in the function node tree;

[0036] An execution unit, configured to execute a web page action based on the candidate operation path and the matching degree sorting corresponding to the candidate operation path, perform status verification on an execution result corresponding to the web page action to obtain a status verification result, in the case where the status verification result is verification failed, perform path backtracking on the function node tree to obtain a first operation path, execute the web page action based on the first operation path, and display the execution result of the web page action in the case where the execution result corresponding to the web page action is verification passed.

[0037] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, it implements the web page action execution method as described in any one of the above.

[0038] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the web page action execution method as described in any one of the above.

[0039] The present invention also provides a computer program product, including a computer program, where when the computer program is executed by a processor, it implements the web page action execution method as described in any one of the above.

[0040] The web page action execution method, device, electronic device and storage medium provided by the present invention, on the one hand, determine the page function structured analysis result based on the HTML text data of the web page and the page screenshot of the web page, improving the accuracy of element recognition in the page function structured analysis result; on the other hand, determine the function node tree based on the page function structured analysis result, which can effectively process complex hierarchical operations, and then semantically match the task target vector corresponding to the task text with the function description vector to determine the operation path, further improving the accuracy and reliability of page state tracking corresponding to the operation path, and converting the user's natural language instruction into a structured operation intention, facilitating subsequent semantic matching and automated task generation, and improving efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0042] Figure 1 is one of the flow diagrams of the web page action execution method provided by the present invention.

[0043] Figure 2 is the flow diagram of the construction of the function node tree provided by the present invention.

[0044] Figure 3 is the second flow diagram of the web page action execution method provided by the present invention.

[0045] Figure 4 is the structural diagram of the web page action execution device provided by the present invention.

[0046] Figure 5 is the structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0048] The terms "first", "second", etc. in the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same category.

[0049] Figure 1 is one of the flow diagrams of the web page action execution method provided by the present invention, as Figure 1 shown, the method includes step 110, step 120, and step 130.

[0050] Step 110: Input the HTML text data of the web page of the target website and the page screenshot of the web page into the vision-language large model to obtain the page function structured analysis result output by the vision-language large model.

[0051] Specifically, the HTML (HyperText Markup Language) text data of the web page of the target website and the page screenshot of the web page can be obtained. For example, the target website can be loaded through a headless browser, and the HTML text data of the web page can be synchronously obtained by using the browser developer tool interface, and the page screenshot of the web page can be captured.

[0052] The HTML text data is the source code of the web page, which contains all the elements (such as titles, paragraphs, buttons, forms, etc.) of the web page and their attributes. The page screenshot is the visual representation of the web page in the current state, capturing the layout, style, and content of the web page in the form of an image. It can help analyze the visual effects of the web page and verify the visibility and status of the elements.

[0053] The page screenshot can be obtained by calling a browser driver (such as WebDriver) to perform operations such as clicking, inputting, and scrolling, and capturing in real time.

[0054] Then, the HTML text data of the web page of the target website and the page screenshot of the web page can be input into the vision-language large model (Vision-Language Models, VLMs) to obtain the page function structured analysis result output by the vision-language large model.

[0055] Here, the page function structured analysis result refers to the analysis of the HTML text data and visual representation (page screenshot) of the web page, and the extractable operation elements (such as buttons, input boxes, links, etc.) in the page and their function descriptions.

[0056] Step 120: Based on the page function structured analysis result, determine a function node tree, and semantically match the task target vector corresponding to the task text with the function description vector to obtain candidate operation paths; the function description vector is obtained by converting the function description text in the function node tree.

[0057] Specifically, after obtaining the page function structured analysis result, the actionable elements and their function descriptions in the page function structured analysis result can be organized hierarchically to form a function node tree.

[0058] Then, semantically match the task target vector corresponding to the task text with the function description vector to obtain candidate operation paths. Among them, the task text can be "book a hotel with the cheapest price and a rating of 4 stars or above", or "query the weather of XX today", etc. Here, calculating the semantic match between the task target vector and the function description vector can be performed using methods such as cosine similarity and Pearson correlation coefficient, and the embodiments of the present invention do not make specific limitations in this regard.

[0059] The function description vector is obtained by converting the function description text in the function node tree. For example, the function description text can be input into a BERT (Bidirectional Encoder Representations from Transformers) model to obtain the function description text output by the BERT model.

[0060] Among them, the task target vector can be obtained by parsing the task text input by the user, that is, a natural language task instruction, using a large model.

[0061] It can be understood that based on the page function structured analysis result, determining a function node tree, and semantically matching the task target vector corresponding to the task text with the function description vector to obtain candidate operation paths can effectively integrate the visual information of the web page screenshot and the structural information of the HTML text data, thereby converting the user's vague text instruction into an executable operation sequence.

[0062] Step 130: Based on the candidate operation paths and the matching degree sorting corresponding to the candidate operation paths, execute web page actions, and perform status verification on the execution results corresponding to the web page actions to obtain a status verification result. In the case where the status verification result is verification failed, perform path backtracking on the function node tree to obtain a first operation path, execute the web page action based on the first operation path, and in the case where the execution result corresponding to the web page action is verification passed, display the execution result of the web page action.

[0063] Specifically, based on the candidate operation paths and the matching degree sorting corresponding to the candidate operation paths, web page actions are performed, that is, the candidate paths are executed according to the matching degree sorting to simulate user operations.

[0064] Then, the execution result corresponding to the web page action is subjected to status verification to obtain a status verification result. Here, a multi-modal large model can be combined to perform status verification on the execution result corresponding to the web page action. Among them, the status verification includes the page jump completion degree, such as whether the target page title appears, the visibility of key elements, such as the order submission button being highlighted, and the asynchronous loading status, such as the data table being rendered completely. The embodiments of the present invention do not make specific limitations on this.

[0065] In the case where the status verification result is that the verification fails, path backtracking is performed on the function node tree to obtain the first operation path, and web page actions are performed based on the first operation path. In the case where the execution result corresponding to the web page action is verified to pass, the execution result of the web page action is displayed. For example, when the path execution fails, specific failure reasons are generated, and failure cases are recorded for subsequent system optimization.

[0066] In addition, the function node tree can be stored in the function node database, and the function node tree in the process of performing web page tree analysis and backtracking correction can also be stored in the function node database.

[0067] The method provided by the embodiment of the present invention inputs the HTML text data of the web page of the target website and the page screenshot of the web page into the vision-language large model to obtain the page function structured analysis result output by the vision-language large model. Then, based on the page function structured analysis result, a function node tree is determined. The task target vector corresponding to the task text is semantically matched with the function description vector to obtain a candidate operation path. The function description vector is obtained by converting the function description text in the function node tree. Finally, based on the candidate operation path and the matching degree sorting corresponding to the candidate operation path, a web page action is executed, and the execution result corresponding to the web page action is verified for its status to obtain a status verification result. In the case where the status verification result is that the verification fails, a path backtracking is performed on the function node tree to obtain a first operation path, and the web page action is executed based on the first operation path. In the case where the execution result corresponding to the web page action is that the verification passes, the execution result of the web page action is displayed. On the one hand, based on the HTML text data of the web page and the page screenshot of the web page, the page function structured analysis result is determined, which improves the accuracy of element recognition in the page function structured analysis result. On the other hand, based on the page function structured analysis result, a function node tree is determined, which can effectively handle complex hierarchical operations. Then, the task target vector corresponding to the task text is semantically matched with the function description vector to determine the operation path, further improving the accuracy and reliability of page status tracking corresponding to the operation path, and converting the user's natural language instruction into a structured operation intention, which is convenient for subsequent semantic matching and automated task generation, improving efficiency.

[0068] Based on the above embodiment, the step of verifying the status of the execution result corresponding to the web page action in step 130 to obtain a status verification result includes:

[0069] Step 131, input the task text, the current page screenshot corresponding to the execution result, and the function description text into a multimodal large model to obtain the status verification result output by the multimodal large model.

[0070] Specifically, the task text, the current page screenshot corresponding to the execution result, and the function description text can be input into a multimodal large model to obtain the status verification result output by the multimodal large model. Among them, the multimodal large model can be CLIP (Contrastive Language–Image Pretraining), etc., and the embodiments of the present invention do not make specific limitations on this.

[0071] It can be understood that the task text provides the background and requirements, the function description text clarifies the goals, and the page screenshots provide intuitive visual information. The multimodal large model can integrate this information to better understand complex scenarios. Further, the multimodal model can match and compare the semantic information in the task text and function description text with the visual information in the current page screenshot, such as verifying whether the buttons are displayed as required and whether the text is correctly displayed, etc., so as to more comprehensively evaluate the status.

[0072] It should be noted that the traditional manual verification method requires a lot of time and effort, while the multimodal large model can quickly process text and image information and automatically verify whether the function meets the expectations, greatly saving time and labor costs.

[0073] The method provided by the embodiments of the present invention adopts an interactive dynamic execution mechanism, decomposes the complex web page automation execution task into multiple operation steps, obtains page screenshots and HTML text data at each key node, verifies the corresponding operation status through a multimodal large model, and obtains a status verification result, thereby significantly improving the intelligence level of the web page automation system, providing an innovative solution for the automatic execution of complex web page tasks, and enhancing the human-computer interaction efficiency.

[0074] Based on the above embodiments, in the case where the status verification result is verification failed in step 130, performing path backtracking on the function node tree to obtain a first operation path, including:

[0075] Step 210, in the case where the status verification result is verification failed, determining the node function description set of the function node tree and the hierarchical depth of the tree structure;

[0076] Step 220, starting from the leaf node corresponding to the hierarchical depth of the tree structure, gradually correcting the function descriptions of the parent nodes upward to obtain the first operation path corresponding to the corrected target function node tree.

[0077] Specifically, in the case where the status verification result is verification failed, determining the node function description set {Dleaf} of the function node tree and the hierarchical depth N of the tree structure (a preset value, such as N = 5). Here, recursively traverse: repeat the analysis process of the function node tree for the new page generated by the clickable element until the preset exploration depth is reached.

[0078] Then, starting from the leaf node corresponding to the hierarchical depth N of the tree structure, gradually correct the function descriptions of the parent nodes upward to obtain the first operation path corresponding to the corrected target function node tree.

[0079] For example, the function description of the child node:

[0080] Enter the email address;

[0081] Enter the password;

[0082] Click the "Login" button;

[0083] Function description of the aggregated parent node: Complete user login authentication.

[0084] Based on the above embodiments, the first operation path corresponding to the corrected target function node tree obtained by correcting the function descriptions of the parent nodes layer by layer starting from the leaf nodes corresponding to the hierarchical depth of the tree structure in step 220 includes:

[0085] Step 221, perform conflict detection on the function descriptions of each child node in the function node tree, and input the conflicting first child node function description, second child node function description, the HTML text data, and the page screenshot into the vision-language large model to obtain the target analysis result output by the vision-language large model;

[0086] Step 222, correct the function node tree based on the target analysis result, and convert the specific operation function description in the function description of the child nodes in the function node tree into a task objective description to obtain the first function node tree;

[0087] Step 223, delete the duplicate child nodes and / or invalid child nodes in the first function node tree to obtain the second function node tree;

[0088] Step 224, perform priority sorting on the node function descriptions of the second function node tree based on the user task execution frequency to obtain the third function node tree;

[0089] Step 225, input the function descriptions of all child nodes under the same parent node in the third function node tree into the multi-modal large model for semantic aggregation to obtain the high-level semantic description output by the multi-modal large model, and use the high-level semantic description as the function description of the parent node to obtain the first operation path corresponding to the target function node tree.

[0090] Specifically, perform conflict detection on the function descriptions of each child node in the function node tree. For example, if there are contradictions in the child node function descriptions (such as "Submit Order" and "Cancel Order"), input the conflicting first child node function description, second child node function description, HTML text data, and page screenshot into the vision-language large model to obtain the target analysis result output by the vision-language large model. The vision-language large model will analyze the page screenshot and HTML text data of the web page to judge the actually executable operations. According to the output result (target analysis result) of the vision-language large model, preferentially retain the high-frequency operations or the functions related to the main process (such as preferentially retaining "Submit Order") to resolve the conflicts.

[0091] Then, based on the target analysis result, the function node tree is corrected, and the specific operation function descriptions in the function descriptions of the child nodes in the function node tree are converted into task target descriptions to obtain the first function node tree. For example, generalizing the operation of clicking the "Search" button to "Execute flight search" makes the function description more general.

[0092] Further, the duplicate child nodes and / or invalid child nodes in the first function node tree are deleted to obtain the second function node tree. For example, the duplicate child nodes in the first function node tree can be deleted to obtain the second function node tree, or the invalid child nodes in the first function node tree can be deleted to obtain the second function node tree, or the duplicate child nodes and invalid child nodes in the first function node tree can be deleted to obtain the second function node tree. The embodiments of the present invention do not make specific limitations in this regard. For example, duplicate pop - ups generated due to page refreshing can be deleted, and this process helps to reduce redundant information and optimize the function description.

[0093] Further, based on the user task execution frequency, the function descriptions of the nodes in the second function node tree can be sorted by priority to obtain the third function node tree. For example, record the user's operation path and task execution frequency in the system. Sort the function descriptions of the nodes according to the task execution frequency, and assign weights to each node based on the priority sorting. The operation of high - frequency tasks or key task nodes with a higher priority in the sorting will be given a higher weight. In addition, the weights of the nodes can be dynamically adjusted according to the user's real - time behavior or historical data to ensure that the function node tree can reflect the user's real needs.

[0094] For example, the output format of the third function node tree can be:

[0095] {"node_id": "P1",

[0096] "children": ["C1", "C2", "C3"],

[0097] "function_desc": "Complete the user registration process",

[0098] "confidence_score": 0.93};

[0099] node_id: The unique identifier of the node.

[0100] children: The list of identifiers of the child nodes.

[0101] function_desc: The corrected function description.

[0102] confidence_score: The confidence level of the model for the function description.

[0103] Finally, input the function descriptions of all child nodes under the same parent node in the third function node tree into the multimodal large model for semantic aggregation to obtain the high-level semantic description output by the multimodal large model, and use the high-level semantic description as the function description of the parent node to obtain the first operation path corresponding to the target function node tree.

[0104] It can be understood that semantic constraints are performed during the operation path planning stage to ensure the accuracy of the operation path. Moreover, the method provided by the embodiments of the present invention realizes the dynamic correction of function nodes through bottom-up semantic aggregation and multimodal conflict resolution.

[0105] The method provided by the embodiments of the present invention: 1) Through conflict detection, contradictions or inconsistencies between function descriptions can be discovered, and the vision-language large model is used for analysis and correction, which ensures the accuracy of the function description and avoids user misunderstandings or system errors caused by description conflicts. 2) Convert the specific operation function description into a task objective description, making the function node tree closer to the user's real needs, helping the user quickly understand the functions and operation paths of the system, and improving user satisfaction. 3) Sort the function descriptions according to the user task execution frequency to ensure that high-frequency tasks are easier to access and execute, thereby improving the response speed and efficiency of the system and reducing the user waiting time. 4) By deleting duplicate and invalid child nodes, the redundancy of the function node tree is reduced, making the system structure clearer and facilitating subsequent maintenance and expansion. 5) Through semantic aggregation of the child node function descriptions by the multimodal large model to generate a high-level parent node description, it can more accurately reflect the overall objective and logical structure of the function, and avoid affecting the overall understanding due to the fragmentation of local descriptions.

[0106] Based on the above embodiments, before step 130 of performing status verification on the execution result corresponding to the web page action to obtain a status verification result, it further includes:

[0107] Step 310, obtain the page status of the web page and the waiting operation duration;

[0108] Step 320, in the case where the page status is the unchanged status or the waiting operation duration is greater than the first threshold, perform path backtracking on the function node tree to obtain a second operation path.

[0109] Specifically, obtain the page status of the web page and the waiting operation duration. Among them, the page status includes the unchanged status and the changed status, and the page status can be determined based on the hash value between the current page and the previous page of the current page.

[0110] When the page status is the unchanged status or the waiting operation duration is greater than the first threshold, perform path backtracking on the function node tree to obtain a second operation path. The first threshold can be 5 seconds, 6 seconds, etc., and the embodiments of the present invention do not make specific limitations on this.

[0111] Based on the above embodiments, the determining step of the page status includes:

[0112] Step 410, obtain the current page and the previous page of the current page;

[0113] Step 420, determine the first hash value of the current page and the second hash value of the previous page;

[0114] Step 430, determine the hash comparison result of the first hash value and the second hash value, and when the hash comparison result is less than the second threshold, determine the page status as the unchanged status, and when the hash comparison result is greater than or equal to the second threshold, determine the page status as the changed status.

[0115] Specifically, to obtain the current page and the previous page of the current page, for example, a browser driver (such as WebDriver) can be called to perform operations such as clicking, inputting, and scrolling, and the current page can be intercepted in real time.

[0116] Then, determine the first hash value of the current page and the second hash value of the previous page. The hash comparison result of the first hash value and the second hash value can be determined by the Hamming distance. When the hash comparison result is less than the second threshold, determine the page status as the unchanged status, and when the hash comparison result is greater than or equal to the second threshold, determine the page status as the changed status.

[0117] In the method provided by the embodiments of the present invention, the generation of hash values and the calculation of Hamming distance generally have relatively low computational complexity. Compared with directly comparing the complete content of the page, this method can significantly reduce the consumption of computing resources.

[0118] Based on the above embodiments, the method further includes:

[0119] Step 510, when the status verification result is verification failed, obtain the executed historical operation path;

[0120] Step 520, determine the adjacent nodes adjacent to each node in the historical operation path from the function node tree, and regenerate a third operation path based on the adjacent nodes.

[0121] Specifically, when the status verification result is verification failed, obtain the executed historical operation path.

[0122] Then, adjacent nodes adjacent to each node in the historical operation path can be determined from the functional node tree, and based on the adjacent nodes, a third operation path is regenerated.

[0123] For example, determine the position of the failed node. For example, assume the current operation path is [P1 -> C1 -> C1.1], and the operation of C1.1 fails. Check the parent node C1 of this node and its sibling nodes (i.e., other children of C1).

[0124] For example, if C1 has children C1.1 and C1.2, and C1.1 fails, then try C1.2. If there are no unattempted nodes at the current level, backtrack to the previous level and repeat the above process.

[0125] The method provided by the embodiments of the present invention, in the case where the status verification result fails the verification, obtains the executed historical operation path, then determines the adjacent nodes adjacent to each node in the historical operation path from the functional node tree, and based on the adjacent nodes, regenerates the third operation path. By automatically adjusting the operation path, the need for manual intervention by the user is reduced, the user experience is improved, and moreover, by regenerating the path, the system can find a better operation path and improve the efficiency of task execution.

[0126] Based on the above embodiments, in the case where the status verification result fails the verification, performing path backtracking on the functional node tree in step 130 to obtain the first operation path includes:

[0127] In the case where the status verification result fails the verification, based on the node confidence of the functional node tree, perform path backtracking on the functional node tree to obtain the first operation path.

[0128] Specifically, in the case where the status verification result fails the verification, based on the node confidence of the functional node tree, perform path backtracking on the functional node tree to obtain the first operation path. For example, preferentially select functional nodes with a node confidence (confidence_score) higher than a threshold (such as 0.85) to generate the operation path. In addition, for low-confidence nodes, strengthen the result verification after performing the operation (such as repeating the screenshot analysis 3 times).

[0129] Based on any of the above embodiments, a web page action execution method can be implemented based on a web page action execution system, and the system includes:

[0130] 1. User interface module: Receive the user task text (such as "query the weather of XX today"), and return the operation result or the reason for the error (such as "the weather query function is temporarily unavailable").

[0131] 2. Independent Exploration Module: Perform web page tree analysis and backtracking correction to generate a functional node database.

[0132] 3. Task Processing Module: Parse the semantics of the task based on the large model, match functional nodes, and generate an operation path.

[0133] 4. Operation Execution Module: Call a browser driver (such as WebDriver) to perform operations such as clicking, inputting, and scrolling, and capture the page status in real time.

[0134] 5. Verification and Backtracking Module: Judge whether the operation result meets the expectation through the multimodal large model, and trigger path backtracking when it fails.

[0135] 6. Database: Store the functional node tree, operation history records, and user task context.

[0136] Based on any of the above embodiments, Figure 2 is a schematic flowchart of the process for constructing a functional node tree provided by the present invention. As Figure 2 shown, first, obtain the HTML text data of the web page of the target website and the page screenshot of the web page, and then input the HTML text data of the web page of the target website and the page screenshot into the vision-language large model for multimodal semantic analysis to obtain the page function structured analysis result output by the vision-language large model, and generate a functional description of the interactive element based on the page function structured analysis result, so as to construct a functional node tree based on the functional description of the interactive element.

[0137] Further, recursively traverse the child nodes of the functional node tree, further correct the semantic description of the parent node, and then backtrack from bottom to top to store the functional node tree in the graph database.

[0138] Based on any of the above embodiments, Figure 3 is the second schematic flowchart of the web page action execution method provided by the present invention. As Figure 3 shown, first, receive a natural language instruction, that is, the user task input, and convert the user task input into a target vector through the large model. Then obtain the functional node tree from the database, and then perform semantic matching between the target vector corresponding to the user task input and the functional description vector to obtain a candidate operation path.

[0139] Based on the candidate operation path and the matching degree sorting (priority) corresponding to the candidate operation path, perform web page actions, and based on the new page state, perform multi-modal state verification on the execution results corresponding to the web page actions to obtain a state verification result. In the case where the state verification result fails the verification, perform path backtracking on the function node tree to obtain the first operation path, that is, perform exception handling and try alternative paths. That is to say, use the first operation path to replace the failed operation path and perform web page actions based on the first operation path. In the case where the execution result corresponding to the web page action passes the verification, display the execution result of the web page action.

[0140] The method provided by the embodiments of the present invention can improve the intelligent level of human-computer interaction by improving web page automation technology, and provide more effective technical support for the automatic execution of complex web page tasks.

[0141] In summary, (1) a comprehensive web page autonomous exploration mechanism is designed, which can dynamically construct a website function node tree.

[0142] (2) A semantic alignment method based on a multi-modal large model is proposed to improve the accuracy of operation element recognition.

[0143] (3) The large model is innovatively applied to complex task decomposition and path planning to achieve intelligent web page operations.

[0144] The method provided by the embodiments of the present invention can significantly improve the intelligent level of the web page automation system and provide a new solution for the automatic execution of complex web page tasks.

[0145] The web page action execution device provided by the present invention will be described below. The web page action execution device described below can be mutually corresponded and referred to the web page action execution method described above.

[0146] Based on any of the above embodiments, the present invention provides a web page action execution device, Figure 4 which is a schematic structural diagram of the web page action execution device provided by the present invention, as Figure 4 shown. The device includes:

[0147] An input unit 01, configured to input the HTML text data of the web page of the target website and the page screenshot of the web page into the vision-language large model, and obtain the page function structured analysis result output by the vision-language large model;

[0148] A semantic matching unit 02, configured to determine a function node tree based on the page function structured analysis result, and perform semantic matching on the task target vector corresponding to the task text and the function description vector to obtain a candidate operation path; the function description vector is obtained by converting based on the function description text in the function node tree;

[0149] An execution unit 03 is configured to perform a web page action based on the candidate operation path and the matching degree sorting corresponding to the candidate operation path, verify the status of the execution result corresponding to the web page action to obtain a status verification result, and in the case where the status verification result fails the verification, perform path backtracking on the function node tree to obtain a first operation path, perform a web page action based on the first operation path, and in the case where the execution result corresponding to the web page action passes the verification, display the execution result of the web page action.

[0150] On the one hand, the device provided by the embodiment of the present invention determines the page function structured analysis result based on the HTML text data of the web page and the page screenshot of the web page, improving the accuracy of element recognition in the page function structured analysis result; on the other hand, based on the page function structured analysis result, a function node tree is determined, which can effectively handle complex hierarchical operations, and then the task target vector corresponding to the task text is semantically matched with the function description vector to determine the operation path, further improving the accuracy and reliability of page state tracking corresponding to the operation path, and converting the user's natural language instruction into a structured operation intention, facilitating subsequent semantic matching and automated task generation, and improving efficiency.

[0151] Based on any of the above embodiments, the execution unit 03 is specifically configured to:

[0152] Input the task text, the current page screenshot corresponding to the execution result, and the function description text into a multi-modal large model to obtain the status verification result output by the multi-modal large model.

[0153] Based on any of the above embodiments, the execution unit 03 specifically includes:

[0154] A determination unit configured to determine the node function description set of the function node tree and the tree structure level depth in the case where the status verification result fails the verification;

[0155] A correction unit configured to start from the leaf node corresponding to the tree structure level depth and correct the function descriptions of the parent nodes layer by layer upward to obtain the first operation path corresponding to the corrected target function node tree.

[0156] Based on any of the above embodiments, the correction unit is specifically configured to:

[0157] Detect conflicts in the function descriptions of each child node in the function node tree, and input the conflicting first child node function description, second child node function description, the HTML text data, and the page screenshot into the vision language large model to obtain the target analysis result output by the vision language large model;

[0158] Modify the functional node tree based on the target analysis result, and convert the specific operation function description in the sub-node function description of the functional node tree into a task target description to obtain a first functional node tree;

[0159] Delete the duplicate sub-nodes and / or invalid sub-nodes in the first functional node tree to obtain a second functional node tree;

[0160] Based on the user task execution frequency, perform priority sorting on the node function descriptions of the second functional node tree to obtain a third functional node tree;

[0161] Input the function descriptions of all sub-nodes under the same parent node in the third functional node tree into the multimodal large model for semantic aggregation to obtain the high-level semantic description output by the multimodal large model, and use the high-level semantic description as the function description of the parent node to obtain the first operation path corresponding to the target functional node tree.

[0162] Based on any of the above embodiments, it further includes a path backtracking unit, and the path backtracking unit is specifically used for:

[0163] Obtain the page state of the web page and the waiting operation duration;

[0164] In the case where the page state is an unchanged state or the waiting operation duration is greater than the first threshold, perform path backtracking on the functional node tree to obtain a second operation path.

[0165] Based on any of the above embodiments, it further includes a page state determination unit, and the page state determination unit is specifically used for:

[0166] Obtain the current page and the previous page of the current page;

[0167] Determine the first hash value of the current page and the second hash value of the previous page;

[0168] Determine the hash comparison result of the first hash value and the second hash value, and in the case where the hash comparison result is less than the second threshold, determine the page state as the unchanged state, and in the case where the hash comparison result is greater than or equal to the second threshold, determine the page state as the changed state.

[0169] Based on any of the above embodiments, it further includes an operation path generation unit, and the operation path generation unit is specifically used for:

[0170] In the case where the status verification result is verification failed, obtain the executed historical operation path;

[0171] Determine adjacent nodes adjacent to each node in the historical operation path from the function node tree, and regenerate a third operation path based on the adjacent nodes.

[0172] Based on any of the above embodiments, the execution unit 03 is specifically configured to:

[0173] In the case where the status verification result fails the verification, perform path backtracking on the function node tree based on the node confidence of the function node tree to obtain the first operation path.

[0174] Figure 5 is a schematic structural diagram of an electronic device provided by the present invention. As Figure 5 shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call the logical instructions in the memory 530 to execute a web page action execution method, which includes: inputting the HTML text data of the web page of the target website and the page screenshot of the web page into a vision-language large model to obtain a page function structured analysis result output by the vision-language large model; determining a function node tree based on the page function structured analysis result, and performing semantic matching between the task target vector corresponding to the task text and the function description vector to obtain a candidate operation path; the function description vector is obtained by converting the function description text in the function node tree; performing a web page action based on the candidate operation path and the matching degree sorting corresponding to the candidate operation path, and verifying the execution result corresponding to the web page action to obtain a status verification result. In the case where the status verification result fails the verification, perform path backtracking on the function node tree to obtain the first operation path, perform a web page action based on the first operation path, and display the execution result of the web page action in the case where the execution result corresponding to the web page action passes the verification.

[0175] In addition, when the logical instructions in the above-mentioned memory 530 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0176] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the web page action execution method provided by the above-mentioned various methods. The method includes: inputting the HTML text data of the web page of the target website and the page screenshot of the web page into a vision-language large model to obtain the page function structured analysis result output by the vision-language large model; based on the page function structured analysis result, determining a function node tree, and semantically matching the task target vector corresponding to the task text with the function description vector to obtain a candidate operation path; the function description vector is obtained by converting the function description text in the function node tree; based on the candidate operation path and the matching degree sorting corresponding to the candidate operation path, executing a web page action, and verifying the status of the execution result corresponding to the web page action to obtain a status verification result. In the case where the status verification result is that the verification fails, performing a path backtracking on the function node tree to obtain a first operation path, and executing the web page action based on the first operation path. In the case where the execution result corresponding to the web page action is that the verification passes, displaying the execution result of the web page action.

[0177] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the web page action execution method provided by the above-mentioned various methods. The method includes: inputting the HTML text data of the web page of the target website and the page screenshot of the web page into a vision-language large model to obtain a page function structured analysis result output by the vision-language large model; determining a function node tree based on the page function structured analysis result, and performing semantic matching between the task target vector corresponding to the task text and the function description vector to obtain a candidate operation path; the function description vector is obtained by converting the function description text in the function node tree; based on the candidate operation path and the matching degree sorting corresponding to the candidate operation path, execute a web page action, and perform status verification on the execution result corresponding to the web page action to obtain a status verification result. In the case where the status verification result is verification failed, perform path backtracking on the function node tree to obtain a first operation path, and execute the web page action based on the first operation path. In the case where the execution result corresponding to the web page action is verification passed, display the execution result of the web page action.

[0178] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0179] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0180] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A web page action execution method, characterized in that: include: Inputting HTML text data of the webpage of the target website and the page screenshot of the webpage into the visual language model to obtain the page function structured analysis result output by the visual language model; The page function structured analysis result refers to the analysis of the HTML text data of the web page and the page screenshot, and the extracted operable elements and their functional descriptions in the page; the operable elements include buttons, input boxes and links; Based on the page function structured analysis result, a function node tree is determined, and a task target vector corresponding to the task text is semantically matched with a function description vector to obtain a candidate operation path; the function description vector is obtained by converting the function description text in the function node tree; Based on the candidate operation paths and the matching degree rankings corresponding to the candidate operation paths, a webpage action is executed, and a status verification is performed on the execution result corresponding to the webpage action to obtain a status verification result. If the status verification result is that the verification fails, a path backtracking is performed on the function node tree to obtain a first operation path. Based on the first operation path, a webpage action is executed. If the execution result corresponding to the webpage action is that the verification passes, the execution result of the webpage action is displayed. When the state verification result is verification failure, performing path backtracking on the function node tree to obtain a first operation path includes: When the status verification result is verification failure, determining a node function description set of the function node tree and a tree structure level depth; Starting from the leaf node corresponding to the hierarchical depth of the tree structure, the functional description of the parent node is revised upward layer by layer to obtain the first operation path corresponding to the revised target functional node tree.

2. The web page action execution method according to claim 1, characterized in that: The performing status verification on the execution result corresponding to the webpage action to obtain the status verification result includes: The task text, the current page screenshot corresponding to the execution result and the function description text are input into the multimodal large model to obtain the state verification result output by the multimodal large model.

3. The web page action execution method according to claim 1, characterized in that: The step of taking the leaf node corresponding to the hierarchical depth of the tree structure as the starting point and revising the function description of the parent node layer by layer to obtain the first operation path corresponding to the revised target function node tree includes: Perform conflict detection on the function descriptions of each sub-node in the function node tree, and input the conflicting first sub-node function description, the second sub-node function description, the HTML text data, and the page screenshot into the visual language model to obtain the target analysis result output by the visual language model; The function node tree is modified based on the target analysis result, and the specific operation function description in the sub-node function description in the function node tree is converted into a task target description to obtain a first function node tree; Deleting duplicate child nodes and / or invalid child nodes in the first function node tree to obtain a second function node tree; Based on the user task execution frequency, the node function descriptions of the second function node tree are prioritized to obtain a third function node tree; The function descriptions of all child nodes under the same parent node in the third function node tree are input into the multimodal large model for semantic aggregation to obtain a high-level semantic description output by the multimodal large model, and the high-level semantic description is used as the function description of the parent node to obtain the first operation path corresponding to the target function node tree.

4. The web page action execution method according to any one of claims 1 to 3, characterized in that: The state verification of the execution result corresponding to the webpage action to obtain the state verification result also includes: Obtaining the page status of the web page and the waiting operation time; When the page state is unchanged or the waiting operation time is longer than a first threshold, the function node tree is backtracked to obtain a second operation path.

5. The web page action execution method according to claim 4, characterized in that: The step of determining the page status includes: Get the current page and the previous page of the current page; Determine a first hash value of the current page and a second hash value of the previous page; Determine a hash comparison result of the first hash value and the second hash value, and when the hash comparison result is less than a second threshold, determine that the page state is the unchanged state; when the hash comparison result is greater than or equal to the second threshold, determine that the page state is the changed state.

6. The web page action execution method according to any one of claims 1 to 3, characterized in that: The method further comprises: When the status verification result is verification failure, obtaining the executed historical operation path; Adjacent nodes adjacent to each node in the historical operation path are determined from the function node tree, and a third operation path is regenerated based on the adjacent nodes.

7. The web page action execution method according to any one of claims 1 to 3, characterized in that: When the state verification result is verification failure, performing path backtracking on the function node tree to obtain a first operation path includes: When the status verification result is verification failure, the function node tree is backtracked based on the node confidence of the function node tree to obtain the first operation path.

8. A web page action execution device, characterized in that: include: An input unit, used to input HTML text data of a web page of a target website and a page screenshot of the web page into the visual language model, and obtain a page function structured analysis result output by the visual language model; The page function structured analysis result refers to the analysis of the HTML text data of the web page and the page screenshot, and the extracted operable elements and their functional descriptions in the page; the operable elements include buttons, input boxes and links; A semantic matching unit, for determining a function node tree based on the page function structured analysis result, and semantically matching a task target vector corresponding to the task text with a function description vector to obtain a candidate operation path; the function description vector is obtained by converting the function description text in the function node tree; an execution unit, configured to execute a webpage action based on the candidate operation paths and the matching degree rankings corresponding to the candidate operation paths, and to perform status verification on the execution result corresponding to the webpage action to obtain a status verification result, and to perform path backtracking on the function node tree to obtain a first operation path when the status verification result is a verification failure, and to execute a webpage action based on the first operation path, and to display the execution result of the webpage action when the execution result corresponding to the webpage action is a verification success; When the state verification result is verification failure, performing path backtracking on the function node tree to obtain a first operation path includes: When the status verification result is verification failure, determining a node function description set of the function node tree and a tree structure level depth; Starting from the leaf node corresponding to the hierarchical depth of the tree structure, the functional description of the parent node is revised upward layer by layer to obtain the first operation path corresponding to the revised target functional node tree.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the web page action execution method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the web page action execution method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Pre-training method and device for multi-task model of webpage and electronic equipment

    CN116049597A

  • Large model driven Web task automatic execution method and system

    CN119248379A