Method for realizing element positioning in RPA system

By building a view relationship map and combining local retrieval and commercial multimodal model, the problems of inaccurate element positioning and poor user experience in traditional RPA systems are solved, and the effect of fast and accurate element positioning and cost reduction is achieved.

CN119962674APending Publication Date: 2025-05-09HANGZHOU BRANCH INTELLIGENT TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510020347.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In traditional RPA systems, element positioning relies on manual clicks and JS mapping, resulting in poor user experience and element xpath/selector is prone to failure; while the end-to-end multimodal model positioning method is efficient, but it has high requirements for large model capabilities, and has high user-side latency and cost.

Method used

By obtaining user input and web pages, building a view relationship tree and converting it into a view relationship map, combining local search and commercial multimodal big models, identifying similar elements and obtaining element coordinates, reducing dependence on big models and improving positioning accuracy.

Benefits of technology

It realizes fast and accurate element positioning in the RPA system, reduces misjudgment situations, improves user experience, and reduces the frequency and cost of calling commercial models through local retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962674A_ABST
    Figure CN119962674A_ABST
Patent Text Reader

Abstract

The invention discloses a method for realizing element positioning in an RPA system, which comprises the following steps of: 1, obtaining an original request input by a user and a current webpage of the user, and obtaining a view relation graph by utilizing the webpage; 2, recognizing similar elements according to the view relation graph; splitting the original request into a condition judgment instruction or a loop execution instruction; step 3, acquiring element coordinates for the loop execution instruction based on the view relation graph; step 4, disassembling the condition judgment instruction into an element query; then recalling visible elements to be recognized in the similar elements; and step 5, for the visible elements to be identified, acquiring element coordinates by using the view relation graph or the multi-modal large model. According to the invention, element positioning can be rapidly realized in the RPA working process, and the misjudgment condition is reduced at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robotic process automation technology, and specifically to a method for implementing element positioning in an RPA system. Background Art

[0002] RPA stands for Robotic Process Automation. Its core capability is to automatically execute the user's business process through robots and improve human work efficiency. In traditional RPA systems, locating / capturing elements is a prerequisite for completing various instructions. This step usually requires manual clicking of elements on the page and using JS to map the coordinates of the web page elements back to the tree nodes of the elements in the accessibility tree. Finally, the xpath / selector and other paths of the elements are obtained as the id that uniquely identifies the element and is used when executing instructions later. The disadvantage of traditional solutions is that users need to capture elements by manually clicking on the screen coordinates, and the xpath / selector and other information of the elements are easily invalid when the page is revised or the accessibility tree changes drastically. With the rapid improvement of AI capabilities, there are prototype demos for office automation (such as computer use released by Anthropic) that can directly respond to the user's intentions input in natural language, such as "If there is a login button, fill in the login information and click the login button", and input the screenshots (images) and the accessibility tree of the web page into GPT-4o / Claude3.5 and other multimodal large models to locate elements through end-to-end solutions. The main disadvantages of the end-to-end method of using multimodal models to locate elements are: 1. High requirements for the capabilities of large models: the model is required to have strong business understanding capabilities, the ability to directly convert text descriptions of screen elements into coordinates, and the ability to find accurate element text blocks for large-scale text input (10w+tokens) (same as the "needle in a haystack" experiment). If the model makes mistakes in any link, the final result will be unavailable) (the current optimal positioning accuracy of the end-to-end large model tested is about 80%). 2. High latency and cost on the user side: As a commercial service, the cost of large models has dropped exponentially, but in the element positioning scenario, the computing power on the terminal side is not reasonably and fully utilized. Moreover, for systems such as RPA that require multiple rounds of interaction with users, the communication + inference latency of commercial models in seconds (2-10 seconds) per step is not the best choice to provide customers with a smooth product experience. Summary of the invention

[0003] The object of the present invention is to provide a method for realizing element positioning in an RPA system. The present invention can enable the RPA working process to realize element positioning quickly while reducing the possibility of misjudgment.

[0004] The technical solution of the present invention is a method for realizing element positioning in an RPA system, comprising the following steps:

[0005] Step 1: Obtain the original request input by the user and the current web page of the user, use the visible elements in the web page and the element boxes corresponding to the visible elements to build a view relationship tree, and then convert the view relationship tree into a view relationship graph;

[0006] Step 2: Perform environmental perception based on the view relationship graph to identify similar elements in the web page; at the same time, split the original request input by the user into conditional judgment instructions or loop execution instructions;

[0007] Step 3: For the loop execution instruction, the multimodal large model is called to identify the overall area of ​​similar elements, and then the element coordinates of all similar elements are obtained based on the overall area and the view relationship map;

[0008] Step 4: For the conditional judgment instruction, the conditional statement of the conditional judgment instruction is decomposed into a natural language description of a single element, which is named element query; then the local search service is used to recall the visible elements to be identified that are semantically identical, similar and / or related to the element query among similar elements;

[0009] Step 5: For visible elements to be identified that have the same or similar semantics as the element query, the element coordinates are directly obtained based on the view relationship graph; for visible elements to be identified that are related to the element query, the visible elements to be identified are input into the multimodal large model to determine the visible element id, and then the element coordinates are determined by the visible element id.

[0010] In the above-mentioned method for implementing element positioning in the RPA system, in step 1, the process of constructing a view relationship tree using visible elements and element frames in a web page is: sorting the element frames of the visible elements from large to small according to the area, and taking the visible element corresponding to the element frame with the largest area as the root node of the view relationship tree, and taking the corresponding visible elements of the remaining element frames as descendant nodes of the view relationship tree according to the inclusion relationship of the area size; wherein the descendant nodes are judged starting from the root node of the view relationship tree, if the element frame of the currently to-be-inserted visible element C is completely contained by the element frame of the root node visible element, then the currently visible element C is a descendant node of the root node visible element, Next, recursively observe the relationship between all child node visible elements of the root node visible element and visible element C. If the element box of a child node visible element A of the root node visible element still completely contains the element box of the visible element C to be inserted, the visible element C to be inserted is regarded as the descendant node of the child node visible element A. Repeat this process until a child node visible element B of a certain layer is encountered. At this time, visible element C is not the descendant node of any other child node visible element of child node visible element B. Then visible element C is the direct child node of child node visible element B and is in a parallel relationship with other child node visible elements of child node visible element B.

[0011] In the aforementioned method for implementing element positioning in the RPA system, in step 1, the process of converting the view relationship tree into a view relationship graph is:

[0012] Step 1.1, based on the multimodal large model, take the coordinates of the element box and the web page screenshot as input, and output the style category of the visible element;

[0013] Step 1.2, based on a predefined semantic model, taking the text of the visible element as input, outputting the semantic category of the visible element;

[0014] Step 1.3, storing the style category and semantic category of the visible element as node attributes in the node of the view relationship tree;

[0015] Step 1.4: Based on spatial location, group visible elements with the same function, close location and / or similar pattern into the same group of elements;

[0016] Step 1.5: Based on the semantics and spatial location of visible elements, consider the elements that are functionally and conceptually generalized to other elements as conceptually related elements;

[0017] Step 1.6: construct bidirectional edges between nodes of the same group of elements, and construct unidirectional edges or semantically relative bidirectional edges between nodes of conceptually related elements, so as to expand the view relationship tree structure into a view relationship graph structure.

[0018] In the aforementioned method for implementing element positioning in the RPA system, in step 2, the process of performing environmental perception based on the view relationship map to identify similar elements in the web page is to use the RPA software to capture the target large frame of the target element in the web page, and regard the element frame with the same or similar size as the target large frame in the view relationship map as a similar frame; then, within the scope of the web page, traverse all similar frames, identify the target small frame inside the similar frame with a fixed relative position to the similar frame, and regard the visible element corresponding to the target small frame as a high-confidence element; finally, traverse all similar frames, and use the high-confidence elements to regard the visible elements in the similar frames that meet the position relationship as similar elements.

[0019] The method for implementing element positioning in the aforementioned RPA system, the similar frame judgment process is as follows: when using the RPA software to capture the target large frame of the target element in the web page, when the length or width of the element frame in the view relationship map is equal to the length or width of the target large frame captured by the RPA software, the distance between the four sides of the target small frame corresponding to all visible elements in all candidate element frames and the target large frame is calculated, and at this time, it is determined whether the candidate element frame and the target small frame of the target element captured by the user in the target large frame meet a specific combination of the following conditions:

[0020] i. The distances of n ≥ 3 of the 4 edges are equal;

[0021] ii. The distances of n ≥ 2 of the 4 edges are equal, and the length and width of the target box of the visible element are exactly equal;

[0022] If so, the target frame and the element frame are considered similar frames.

[0023] In the aforementioned method for implementing element positioning in the RPA system, the high-confidence element judgment process is to calculate the distance between the four edges of the target small box and the similar box of all visible elements in all similar boxes, when one of the following conditions is met:

[0024] a1. At least three margins are equal, and at least one of the width or height is equal;

[0025] a2. At least 2 margins are equal, width and height are equal, and the text of the visible elements is the same;

[0026] The visible elements inside the similarity box are considered high-confidence elements, and the others that do not meet the conditions are considered non-high-confidence elements. Finally, the same group_id is assigned to high-confidence elements in different similarity boxes.

[0027] The method for implementing element positioning in the aforementioned RPA system, using high-confidence elements to regard visible elements in similar boxes that satisfy positional relationships as similar elements, includes the following conditions:

[0028] b1. All visible elements in similar boxes with the same group_id are directly regarded as similar elements;

[0029] b2. For target elements whose target small boxes do not overlap with the target small boxes of other high-confidence elements in the target large box to which they belong, obtain the relative position y1 and distance d1 of the target element and the high-confidence element; then search for the visible element N and the corresponding high-confidence element in other similar boxes, obtain the relative position y1 and distance d2 of the visible element N and the high-confidence element, and if y1=y2 and d1=d2, then the target element and element N are similar elements;

[0030] b3. If there are target elements contained in the target small box and the target small boxes of other high-confidence elements in the target large box to which it belongs, calculate the distance between the four edges of the target element and the high-confidence element; then find the visible element N and the corresponding high-confidence element in other similar boxes, and obtain the distance between the four edges of the visible element N and the high-confidence element; if there are at least three edges with equal margins, and at least one of the width or height is equal, then the target element and element N are similar elements;

[0031] b4. For a target element whose length and width are exactly the same as those of the visible element N in the target small box and the similar box, and whose texts are the same or similar to those of the visible element N, the target element and element N are similar elements.

[0032] In the aforementioned method for implementing element positioning in the RPA system, in step 5, for visible elements associated with the element query, first delete the visible elements that do not conform to the category related to the element query, then delete the visible elements outside the current target large box, and then delete the elements whose similarity with the element query is lower than the threshold, so as to reduce the number of visible elements input into the multimodal large model.

[0033] The method for implementing element positioning in the aforementioned RPA system, in the process of inputting the visible elements to be identified into the multimodal large model to determine the visible element id, if the number of remaining visible elements to be identified is greater than 10, the nodes corresponding to the visible elements to be identified in the view relationship graph are input into the multimodal large model; if the number of remaining visible elements to be identified is less than or equal to 10, the visible elements to be identified are marked with a red frame and id in the screenshot of the current web page, and are input into the multimodal large model in the form of a screenshot.

[0034] Compared with the prior art, the present invention designs a solution that combines local retrieval with calling a commercial multimodal large model, combining the powerful business understanding ability of the large model with the local efficient retrieval ability, so that users can obtain the product experience of "generating short processes through natural language descriptions". The present invention completes "requests where the user's intention directly includes element text" locally through a solution that combines local retrieval with calling a commercial multimodal large model. This solution can bring the following advantages: 1. Improve the accuracy of "requests where the user's intention directly includes element text" and significantly reduce misjudgments. 2. Reduce the calling frequency and cost of model services (local indexing also requires calling the semantic vector service interface, but a page only needs to be called once, which is advantageous compared to the end-to-end calling of a multimodal model that requires multiple calls). 3. Delete the text elements that are determined to be processable locally before calling the large model to reduce the input complexity of the model service and reduce model consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a schematic diagram of the application of the element positioning of the present invention in the RPA system;

[0036] Figure 2 Represents the element boxes corresponding to different visible elements in a web page

[0037] Figure 3 A schematic diagram of the present invention for sorting visible elements from large to small in terms of area;

[0038] Figure 4 A schematic diagram of the present invention using a visible element as a descendant node of a root node visible element;

[0039] Figure 5 A schematic diagram of a child node of a visible element as a child node of the present invention;

[0040] Figure 6 It is a data structure diagram of the view relationship graph;

[0041] Figure 7 This is a schematic diagram of the target frame;

[0042] Figure 8 It is a schematic diagram of condition b for similar element determination;

[0043] Fig. 9 It is a schematic diagram of condition c for similar element determination;

[0044] Fig.10 It is a schematic diagram of the condition d for similar element determination;

[0045] Fig.11 It is a schematic diagram of the overall area of ​​similar elements;

[0046] Fig.12 It is a schematic diagram of splitting the original request input by the user into conditional judgment instructions or loop execution instructions;

[0047] Fig.13 It is a flow chart for determining the element coordinates of the visible elements to be identified that are associated with the element query. DETAILED DESCRIPTION

[0048] The present invention is further described below in conjunction with the accompanying drawings and embodiments, but they are not intended to limit the present invention.

[0049] Embodiment: A method for implementing element positioning in an RPA system is applied to a user to generate a corresponding short process through natural language description, and the process is implemented in the RPA system. In the process of responding to user requests and generating processes, the RPA system needs to go through multiple steps. For different product forms, not all steps are required, but the element positioning stated in the present invention is a mandatory module in the overall process generation process, such as Figure 1 As shown, the specific implementation steps are as follows:

[0050] Step 1: Obtain the original request input by the user and the current web page of the user, use the visible elements in the web page and the element boxes corresponding to the visible elements to build a view relationship tree, and then convert the view relationship tree into a view relationship graph;

[0051] In this step, the user input intent is limited to requests that can be responded to when the current web page display content remains unchanged, that is, requests that do not require cross-application or page jumps. For example: "Please find keywords with a CPM less than 5 yuan and that match the account user profile, and promote them within the current plan."

[0052] Web pages include computer web pages, software clients or mobile clients; visible elements refer to controls or content on web pages that can be perceived by users, such as buttons, icons, input boxes, and text. Due to the presence of multiple layers of panels or areas on the page, elements edited in advance in the accessibility tree (dom code) are not always visible. There are special functions in js to distinguish whether the elements in the current web page or client interface are visible. After obtaining the visible element, a unique-id will be assigned to the visible element for subsequent access to the element. In addition, all visible elements can call js functions to calculate the coordinate position and size (x, y, width width, height height). These four values ​​can be used to draw a rectangular box, called an element box. The element box can just cover the visible elements in the screenshot of the target page, such as Figure 1 As shown, Figure 2Indicates the element boxes corresponding to different visible elements in a web page. The coordinates of the element boxes can be obtained through Selenium and Microsoft Inspect tools. Selenium and Microsoft Inspect tools can also be used to obtain screenshots after the front-end code of the target page is rendered.

[0053] In this step, the basic data of the visible element is also obtained, that is, the text of the visible element. The text is generally used to describe the function of the visible element. When obtaining the text (innerText) attribute of the visible element in the accessibility tree, the attributes that are invisible to the user are filtered out, such as class, id, imgurl (image hyperlink, etc.).

[0054] When the element frame is obtained, such as Figure 3 As shown, the element frames of the visible elements are sorted from large to small by area, and the visible element corresponding to the element frame with the largest area is used as the root node of the view relationship tree. The remaining element frames use the corresponding visible elements as descendant nodes of the view relationship tree according to the inclusion relationship of the area size, thereby forming a complete view relationship tree. Among them, the specific process of constructing the view relationship tree is to start from the root node of the view relationship tree. If the element frame of the currently inserted visible element C is completely contained by the element frame of the root node visible element, then the current visible element C is a descendant node of the root node visible element. Next, recursively observe the relationship between all the child node visible elements of the root node visible element and the visible element C. If the element frame of a child node visible element A of the root node visible element still completely contains the element frame of the currently inserted visible element C (corresponding Figure 4 ), then the visible element C to be inserted is the descendant node of the child node visible element A, and the process is repeated until a child node visible element B of a certain layer is encountered. At this time, the visible element C is not the descendant node of any other child node visible element of the child node visible element B. Then the visible element C is the direct child node of the child node visible element B, and is in a parallel relationship with other child node visible elements of the child node visible element B (corresponding to Figure 5 ).

[0055] After the view relationship tree is constructed, the view relationship graph is constructed using the view relationship tree. The specific process is as follows:

[0056] Step 1.1, based on the multimodal large model, take the element frame coordinates and interface screenshots as input, and output the category of visible elements;

[0057] In this step, the multimodal large model can use local CV models predefined based on RPA business, such as classification (ViT) and target detection (Yolo). In these local CV models, the element box coordinates and interface screenshots obtained in step 1 are used as input, and the style category of the element is output, such as "input box", "button", "icon", "date panel", etc.;

[0058] Step 1.2, based on a predefined semantic model, taking the text and style of the visible element as input, outputting the semantic category of the visible element;

[0059] In this step, the semantic model predefined based on the RPA business, such as regular expressions and large language models (LLM), can be used to input element text and style, and output the semantic category of the element, such as title, navigation bar, number, time, and location;

[0060] Step 1.3, storing the style category and semantic category of the element as node attributes in the node of the view relationship tree;

[0061] Step 1.4: Based on spatial location, group visible elements with the same function, close location and / or similar pattern into the same group of elements;

[0062] In this step, the similarity pattern includes the element box layout and the semantics of the visible elements;

[0063] Whether the element frame layout is similar is determined based on the coordinates and size of the element frame; for example, if the x-coordinate and width of the element frame are completely equal, it means that the corresponding elements are similar, which can be vertical list content, including Weibo, Baidu search results, etc. in an example;

[0064] Whether the semantics of the visible elements are similar is determined by the text attributes of the visible elements. For example, if the text of a group of elements are all place names (such as Hangzhou, Shanghai, Ningbo), then this group of text elements can be regarded as a "place name" element group, and this element group can be used as an input variable for a screening action in subsequent applications.

[0065] Step 1.5: Based on the semantics and spatial positions of visible elements, elements that are functionally and conceptually generalized to other elements are considered as conceptually related elements;

[0066] For example, if the element "Remember password" is visible on the target page, there is usually a "check box" on the left side of it, so "Remember password" and "check box" are conceptually related elements;

[0067] Step 1.6: Build bidirectional edges between nodes of the same group of elements, and build unidirectional edges or semantically related bidirectional edges between nodes of conceptually related elements to expand the view relationship tree structure into a view relationship graph structure. The final data structure is as follows: Figure 6 shown.

[0068] For example, take "number of visitors" as a visible element on a page. When it is a node in the view relationship graph, "payment amount", "number of paid sub-orders", ..., "number of views", etc. are elements in the same group as "number of visitors". At this time, there is a bidirectional edge named "same group element" from the "number of visitors" node to the "number of views" node. "0" is a conceptual association element of "number of visitors", where the "number of visitors" node is the key and the "0" node is the value. At this time, the "number of visitors" node has an edge pointing to the "0" node, named "key", and the "0" node has an edge pointing to the "number of visitors" node, named "value".

[0069] Step 2: Perform environmental perception based on the view relationship graph to identify similar elements in the web page; at the same time, split the original request input by the user into conditional judgment instructions or loop execution instructions;

[0070] The environment in this step refers to the content displayed on the current user terminal, including the accessibility tree, style information, JS action definition, etc. required to render these contents; the environment perception in this step includes but is not limited to target detection (used to obtain the category of elements and the semantic information of some images), similar element recognition (obtaining similar elements on the page through the view relationship map to identify lists or multiple-choice button groups in specific areas, etc.) and associated element (key-value) recognition (identifying the association relationship between elements, used to describe elements such as icons, single / check boxes, input boxes, etc. that have no available text in the elements themselves).

[0071] Among them, the process of performing environmental perception based on the view relationship graph to identify similar elements in a web page is to use RPA software to capture the target large frame of the target element in the web page, and regard the element frame in the view relationship graph that is the same or similar in size to the target large frame as a similar frame; then, within the scope of the web page, traverse all similar frames, identify the target small frame inside the similar frame that is fixed in position relative to the similar frame, and regard the visible element corresponding to the target small frame as a high-confidence element; finally, traverse all similar frames, and use high-confidence elements to regard the visible elements in the similar frames that meet the position relationship as similar elements.

[0072] Furthermore, the target large frame contains the content required for one cycle of RPA software batch operation. Figure 7As shown, for example, when using RPA software to obtain multiple products displayed on an e-commerce page, each product will correspond to product pictures, product introductions, product prices, sales volume and other contents. The target large frame captured by the RPA software will include product pictures, product introductions, product prices and sales volume and other contents. These product pictures, product introductions, product prices and sales volume are visible elements, and the element frame corresponding to the visible elements is the target small frame, which is contained in the target large frame.

[0073] Preferably, the process of determining the element frame with a similar size to the target large frame in the view relationship graph is as follows:

[0074] When the length or width of an element box in the view relationship graph is equal to the length or width of the target large box captured by the RPA software, the distance between the four sides (up, down, left, and right) of the target small boxes corresponding to all visible elements in all candidate element boxes and the target large box is calculated. At this time, it is determined whether the candidate element box and the target small box of the target element captured by the user in the target large box meet a specific combination of the following conditions:

[0075] i. The distances of n ≥ 3 of the 4 edges are equal;

[0076] ii. The distances of n ≥ 2 of the 4 edges are equal, and the length and width of the target box of the visible element are exactly equal;

[0077] If so, the target frame and the element frame are considered similar frames.

[0078] Preferably, the judgment process of the high confidence element is as follows:

[0079] Step 3.1. Calculate the distances between all visible elements inside all similar boxes and the four edges between the similar boxes;

[0080] Step 3.2. When one of the following conditions is met:

[0081] a. At least 3 margins are equal, and at least one of the width or height is equal;

[0082] b. At least 2 margins are equal, width and height are equal, and the text of the visible elements is the same;

[0083] The visible elements inside the similarity box are considered high-confidence elements, and the others that do not meet the conditions are considered non-high-confidence elements;

[0084] Step 3.3. Assign the same group_id to high confidence elements in different similarity boxes.

[0085] Preferably, using high-confidence elements to regard visible elements in similar frames that satisfy the positional relationship as similar elements includes the following conditions:

[0086] a. All visible elements in similar boxes with the same group_id are directly regarded as similar elements;

[0087] b. For target elements whose target small boxes do not overlap with the target small boxes of other high-confidence elements in the target large box to which they belong, obtain the relative position y1 and distance d1 of the target element and the high-confidence element; then find the visible element N and the corresponding high-confidence element in other similar boxes, obtain the relative position y1 and distance d2 of the visible element N and the high-confidence element, and y1 = y2, d1 = d2, then the target element and element N are similar elements, such as Figure 8 As shown;

[0088] c. For the target small box and other high-confidence elements in the target large box to which it belongs, there are target elements that are mutually contained, such as Fig. 9 As shown, the distances between the four edges of the target element and the high-confidence element are calculated; then the visible element N and the corresponding high-confidence element are searched in other similar boxes. If the visible element N and the high-confidence element have at least three edges with equal margins (such as d1=d4, d2=d5, d3=d6), and at least one of the width or height is equal (x1=x2), then the target element and element N are similar elements;

[0089] d. For target elements whose length and width are exactly the same as those of visible element N in the target small box and similar box (x1=x2, y1=y2), and whose texts are the same or similar to those of visible element N (the text contents are both place names, such as Hangzhou and Shanghai), then the target element and element N are similar elements, such as Fig.10 shown.

[0090] Therefore, when the RPA software captures any product introduction, it can obtain other product introductions as its similar elements. When the RPA software captures any product picture, it can obtain other product pictures as its similar elements. When the RPA software captures any product price, it can obtain other product prices as its similar elements. When the RPA software captures any product sales, it can obtain other product sales as other elements.

[0091] Furthermore, in this step, splitting the original request input by the user into conditional judgment instructions or loop execution instructions means that when the original user's requirements are complex, the original requirements need to be split, and the split requirements can be directly converted into conditional judgment instructions or loop execution instructions, such as Fig.12 shown.

[0092] Among them, loop execution instructions: obtain multiple UI elements (similar elements / groups) through user intent and logical conditions, and loop the same operation on this group of elements. For example, if the user's intent is to grab all JD.com self-operated products with a price higher than 20 yuan and add them to the shopping list, the original request can be split into loop execution instructions.

[0093] Conditional judgment instructions: Make Boolean judgments on the visibility and type of UI elements, or make logical expression judgments on the text and value of UI elements. For example, if the user's intention is: if there is a login button, fill in the login information and click the login button, then the original request can be split into conditional judgment instructions.

[0094] The user's original request splitting step can be completed by calling a commercial multimodal large model (the multimodal large model can adopt a local cv model predefined based on RPA business such as classification (ViT) and target detection (yolo).

[0095] Step 3: For the loop execution instruction, the multimodal large model is called to identify the overall area of ​​similar elements, and then the element coordinates of all similar elements are obtained based on the overall area and the view relationship map;

[0096] In this step, since the loop execution instructions are generally repetitive, the overall area of ​​similar elements can be identified by calling the multimodal large model, such as Fig.11 As shown in the thick black box in , after the recognition of the entire area is completed, the existing view relationship map can be called to complete the batch processing and acquisition of all similar elements. At the same time, since the view relationship map will contain the coordinates of the visible elements, the corresponding visible elements can be directly located when similar elements are obtained.

[0097] Step 4: For the conditional judgment instruction, the conditional statement of the conditional judgment instruction is decomposed into a natural language description of a single element, which is named element query; then the local search service is used to recall the visible elements to be identified that are semantically identical, similar and / or related to the element query among similar elements;

[0098] In this step, the element query is a natural sentence including the corresponding text of a visible element, such as "click the button to use it immediately", where "use it immediately" is the corresponding text of the visible element, which can be seen on the corresponding web page. Therefore, the local search service can recall the visible elements to be identified that are semantically identical, similar and / or related to the text "use it immediately" among similar elements;

[0099] Step 5: For visible elements to be identified that have the same or similar semantics as the element query, the element coordinates are directly obtained based on the view relationship graph; for visible elements to be identified that are related to the element query, the visible elements to be identified are input into the multimodal large model to determine the visible element id, and then the element coordinates are determined by the visible element id.

[0100] Taking the "click the button to use immediately" as an example, for visible elements with the same or similar semantics as the "use immediately" text (for example, the text is "use immediately"), the existing view relationship map can be called to complete the acquisition of the visible elements. At the same time, since the view relationship map will contain the coordinates of the visible elements, the positioning of the corresponding visible elements can be directly achieved.

[0101] For the element query which is a natural language description of "enter the mobile phone number of the person in charge", since the visible element associated with it is an "input box", it has no relevant text description, because this type of element query can only recall the visible elements to be identified that are associated with it (i.e., different input boxes), among which the types of input boxes include "name", "gender", "age", "address", "mobile phone number", "QQ number", etc. Due to the large number of input boxes, before calling the multimodal large model, first delete the visible elements that do not conform to the category of the element query (if the query is an input box, delete the visible elements other than the input box, such as the "search box"), and then delete the visible elements other than the current target large box (if the target large box is an "information filling" large box, the input boxes in other parts of the web page can be deleted), and elements whose similarity with the element query is lower than the threshold can also be deleted (the similarity determination can refer to the similar element determination process in the previous article, and the threshold setting can be set according to actual needs) to reduce the number of visible elements input to the multimodal large model. After the deletion, the number of remaining visible elements to be identified is determined. If the number of remaining visible elements to be identified is greater than 10, the nodes corresponding to the visible elements to be identified in the view relationship graph are input into the multimodal large model; if the number of remaining visible elements to be identified is less than or equal to 10, the visible elements to be identified are marked with red boxes and ids in the screenshot of the current web page (id refers to the serial number of the corresponding red box, such as 1, 2, 3, 4...), and are input into the multimodal large model in the form of screenshots; the multimodal large model determines the id of each visible element (such as a name input box, a gender input box, etc.). After knowing the id of the visible element, the coordinates of the visible element can be determined in the corresponding code and page by calling JavaScript, thereby realizing the positioning of the element. The process is as follows: Fig.13 shown.

[0102] In this step, node input is used, which is to pass the page context into the multimodal large model in the form of text (DOM / accessibility tree). The advantage is that the multimodal large model can understand more text at one time. When the node to be identified is "node input", the number of nodes that support identification is relatively large; but its disadvantage is that the number of input tokens is relatively large (high cost) for one call, and the style information is difficult to fully represent. Screenshot input is used to add a mask or a red frame to the identified element on the picture. The advantage is that the screenshot method is more intuitive, and the number of input tokens is only related to the resolution of the screenshot; and the disadvantage is that the multimodal large model can pay attention to a limited number of elements in the form of a picture (the current model GPT-4O / O1, the number of elements that can be identified with high quality in one call is limited, and the experiment is <10). Therefore, this application selects different input methods according to the capabilities of the multimodal large model, thereby realizing the efficient use of the multimodal large model.

[0103] In summary, the present invention designs a solution that combines local retrieval with calling a commercial multimodal large model, combining the powerful business understanding ability of the large model with the local efficient retrieval ability, so that the RPA work process can quickly locate elements while reducing misjudgments.

Claims

1. A method for implementing element positioning in an RPA system, characterized in that: The steps include: Step 1: Obtain the original request input by the user and the current web page of the user, use the visible elements in the web page and the element boxes corresponding to the visible elements to build a view relationship tree, and then convert the view relationship tree into a view relationship graph; Step 2: Perform environmental perception based on the view relationship graph to identify similar elements in the web page; At the same time, the original request input by the user is split into conditional judgment instructions or loop execution instructions; Step 3: For the loop execution instruction, the multimodal large model is called to identify the overall area of ​​similar elements, and then the element coordinates of all similar elements are obtained based on the overall area and the view relationship map; Step 4: For the conditional judgment instruction, the conditional statement of the conditional judgment instruction is decomposed into a natural language description of a single element, which is named element query; then the local search service is used to recall the visible elements to be identified that are semantically identical, similar and / or related to the element query among similar elements; Step 5: For visible elements to be identified that have the same or similar semantics as the element query, the element coordinates are directly obtained based on the view relationship graph; for visible elements to be identified that are related to the element query, the visible elements to be identified are input into the multimodal large model to determine the visible element id, and then the element coordinates are determined by the visible element id.

2. The method for implementing element positioning in the RPA system according to claim 1, characterized in that: In step 1, the process of constructing a view relationship tree using visible elements and element frames in the web page is as follows: sort the element frames of the visible elements from large to small according to the area, and use the visible element corresponding to the element frame with the largest area as the root node of the view relationship tree. The remaining element frames use the corresponding visible elements as descendant nodes of the view relationship tree according to the inclusion relationship of the area size; wherein the descendant nodes are determined starting from the root node of the view relationship tree. If the element frame of the currently inserted visible element C is completely contained by the element frame of the root node visible element, then the currently visible element C is a descendant node of the root node visible element. Next, recursively observe the root node. The relationship between all child node visible elements of the point visible element and the visible element C, if the element box of a child node visible element A of the root node visible element still completely contains the element box of the visible element C to be inserted, then the visible element C to be inserted is regarded as the descendant node of the child node visible element A, and the process is repeated until a child node visible element B of a certain layer is encountered. At this time, the visible element C is not the descendant node of any other child node visible element of the child node visible element B, then the visible element C is regarded as the direct child node of the child node visible element B, and is in a parallel relationship with other child node visible elements of the child node visible element B.

3. The method for implementing element positioning in the RPA system according to claim 1, characterized in that: In step 1, the process of converting the view relationship tree into a view relationship graph is: Step 1.1, based on the multimodal large model, take the coordinates of the element box and the web page screenshot as input, and output the style category of the visible element; Step 1.2, based on a predefined semantic model, taking the text of the visible element as input, outputting the semantic category of the visible element; Step 1.3, storing the style category and semantic category of the visible element as node attributes in the node of the view relationship tree; Step 1.4: Based on spatial location, group visible elements with the same function, close location and / or similar pattern into the same group of elements; Step 1.5: Based on the semantics and spatial location of visible elements, consider the elements that are functionally and conceptually generalized to other elements as conceptually related elements; Step 1.6: construct bidirectional edges between nodes of the same group of elements, and construct unidirectional edges or semantically relative bidirectional edges between nodes of conceptually related elements, so as to expand the view relationship tree structure into a view relationship graph structure.

4. The method for implementing element positioning in an RPA system according to claim 1, characterized in that: In step 2, the process of performing environmental perception based on the view relationship graph to identify similar elements in the web page is to use the RPA software to capture the target large frame of the target element in the web page, and regard the element frame in the view relationship graph that is the same or similar in size to the target large frame as a similar frame; Then, within the scope of the webpage, all similar frames are traversed, and the target small frame with a fixed relative position to the similar frame is identified, and the visible element corresponding to the target small frame is regarded as a high-confidence element; Finally, all similar boxes are traversed, and the visible elements in the similar boxes that satisfy the positional relationship are regarded as similar elements using high-confidence elements.

5. The method for implementing element positioning in the RPA system according to claim 4, characterized in that: The similarity frame judgment process is as follows: when the RPA software is used to capture the target large frame of the target element in the web page, when the length or width of the element frame in the view relationship map is equal to the length or width of the target large frame captured by the RPA software, the distance between the four sides of the target small frame corresponding to all visible elements in all candidate element frames and the target large frame is calculated. At this time, it is determined whether the candidate element frame and the target small frame of the target element captured by the user in the target large frame meet a specific combination of the following conditions: i. The distances of n ≥ 3 of the 4 edges are equal; ii. The distances of n ≥ 2 of the 4 edges are equal, and the length and width of the target box of the visible element are exactly equal; If so, the target frame and the element frame are considered similar frames.

6. The method for implementing element positioning in the RPA system according to claim 4, characterized in that: The high confidence element judgment process is to calculate the distance between the four edges of the target small box and the similar box of all visible elements in all similar boxes, when one of the following conditions is met: a1. At least three margins are equal, and at least one of the width or height is equal; a2. At least 2 margins are equal, width and height are equal, and the text of the visible elements is the same; The visible elements inside the similarity box are considered high-confidence elements, and the others that do not meet the conditions are considered non-high-confidence elements. Finally, the same group_id is assigned to high-confidence elements in different similarity boxes.

7. The method for implementing element positioning in the RPA system according to claim 6, characterized in that: Using high-confidence elements to regard visible elements in similar boxes that satisfy positional relationships as similar elements includes the following conditions: b1. All visible elements in similar boxes with the same group_id are directly regarded as similar elements; b2. For target elements whose target small boxes do not overlap with the target small boxes of other high-confidence elements in the target large box to which they belong, obtain the relative position y1 and distance d1 of the target element and the high-confidence element; then search for the visible element N and the corresponding high-confidence element in other similar boxes, obtain the relative position y1 and distance d2 of the visible element N and the high-confidence element, and if y1=y2 and d1=d2, then the target element and element N are similar elements; b3. If there are target elements contained in the target small box and the target small boxes of other high-confidence elements in the target large box to which it belongs, calculate the distance between the four edges of the target element and the high-confidence element; then find the visible element N and the corresponding high-confidence element in other similar boxes, and obtain the distance between the four edges of the visible element N and the high-confidence element; if there are at least three edges with equal margins, and at least one of the width or height is equal, then the target element and element N are similar elements; b4. For a target element whose length and width are exactly the same as those of the visible element N in the target small box and the similar box, and whose texts are the same or similar to those of the visible element N, the target element and element N are similar elements.

8. The method for implementing element positioning in an RPA system according to claim 4, characterized in that: In step 5, for the visible elements associated with the element query, first delete the visible elements that do not conform to the category of the element query, then delete the visible elements outside the current target large box, and then delete the elements whose similarity with the element query is lower than the threshold, so as to reduce the number of visible elements input into the multimodal large model.

9. The method for implementing element positioning in the RPA system according to claim 7, characterized in that: In the process of inputting the visible elements to be identified into the multimodal large model to determine the visible element id, if the number of remaining visible elements to be identified is greater than 10, the nodes corresponding to the visible elements to be identified in the view relationship graph are input into the multimodal large model; if the number of remaining visible elements to be identified is less than or equal to 10, the visible elements to be identified are marked with a red box and id in the screenshot of the current web page and input into the multimodal large model in the form of a screenshot.

Citation Information

Cited By

  • Visual semantic fusion page element positioning method based on large model

    CN120257213A

  • Icon and text positioning method based on multi-modal large model

    CN121561451A

  • Target element positioning method and device based on interface structure perception

    CN122363805A

  • Target element positioning method and device based on interface structure perception

    CN122363805B