Data structure construction method for understanding graphical interactive interface in RPA scene

By building a view relationship tree and graph, combining multimodal and semantic models, the shortcomings of graphical interface understanding in the existing technology are solved, cross-platform interface data structure is realized, and the development efficiency of RPA applications and the accuracy of AI models are improved.

CN119938040APending Publication Date: 2025-05-06HANGZHOU BRANCH INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510001703.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

There are two main problems in the existing technology when understanding graphical interfaces: one is that the method based on the accessibility dom tree fails with the website or software revision, and cannot perfectly reproduce the user's intentions; the other is that the method based on screenshot occupies a lot of resources, and the AI ​​model has low accuracy when element positioning and layout alignment.

Method used

By obtaining screenshots and barrier-free trees of the target page, building a view relationship tree and view relationship map, combining multimodal models and predefined semantic models, identifying and classifying interface elements, forming a cross-platform data structure.

Benefits of technology

This method retains interface information consistent with user perception, reduces the difficulty of training AI models, improves the accuracy of element positioning and layout alignment, and enhances the development efficiency and capabilities of RPA applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938040A_ABST
    Figure CN119938040A_ABST
Patent Text Reader

Abstract

The invention discloses a data structure construction method for understanding a graphical interactive interface under an RPA scene. The method comprises the following steps: step 1, obtaining a target page screenshot and a barrier-free tree of a target page back-end code; step 2, obtaining visible elements, element boxes and basic data according to the barrier-free tree and the target page screenshot; 3, sorting the element boxes of the visible elements from large to small according to the areas, taking the visible element corresponding to the element box with the largest area as a root node of the view relation tree, and taking the visible elements corresponding to the other element boxes as descendant nodes of the view relation tree according to the inclusion relation of the areas; and step 4, constructing a view relation graph by using the view relation tree and the basic data, thereby taking the obtained view relation graph as a data structure for understanding the graphical interaction interface. The data structure constructed by the method can be convenient for the RPA robot to understand the graphical interaction interface, and is suitable for AI model training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robotic process automation technology, and specifically to a method for constructing a data structure for understanding a graphical interactive interface in an RPA scenario. Background Art

[0002] RPA stands for Robotic Process Automation. Its core capability is to automatically execute the user's business process through robots and improve human work efficiency. Most RPA applications require robots to complete business needs through interaction with graphical interfaces. Therefore, a better understanding of graphical interfaces is one of the prerequisites for improving RPA application development efficiency and expanding the boundaries of RPA capabilities. There are currently two known solutions for understanding graphical interfaces, and their disadvantages are: 1. The source code of the graphical interface (such as web pages, software clients, and mobile apps) is used as input. The "barrier-free DOM tree" is analyzed to achieve functions such as functional classification of interface elements and grouping of similar elements. The disadvantages of this type of solution are: a. The "barrier-free tree" structure often changes with the revision of websites and software, resulting in the failure of rules or algorithm systems that only use the "barrier-free tree" structure as input. For example, web crawlers and the positioning of the same and similar elements of software. b. The "barrier-free tree" structure is a data structure for website and software developers, which is different from the information that users see on the screen after the browser rendering is completed, resulting in RPA applications developed for the "barrier-free tree" structure cannot perfectly reproduce the user's intentions. For example, batch data capture and operation. 2. Using screenshots of graphical interfaces as input, AI models are used to identify the functions of interface elements or group them. The disadvantages of this method are: a. Compared with the effective information of the "accessibility tree", screenshots contain a lot of redundant information and are difficult to be losslessly compressed from the perspective of RPA task requirements. Therefore, more resources are required when transmitting or storing system inputs, which increases system overhead. b. When AI models use images as input, they are more prone to hallucinations, and the most advanced models in the industry (such as GPT-4o) are aligned with human perception during training, and do not regard pixel-level positioning as a high-priority task, resulting in low accuracy in answering questions such as precise positioning of elements and alignment of interface layouts. Summary of the invention

[0003] The purpose of the present invention is to provide a method for constructing a data structure for understanding a graphical interactive interface in an RPA scenario. The data structure constructed by the present invention can facilitate the RPA robot to understand the graphical interactive interface and is suitable for AI model training.

[0004] The technical solution of the present invention is a method for constructing a data structure for understanding a graphical interactive interface in an RPA scenario, comprising the following steps:

[0005] Step 1: Get the screenshot of the target page and the accessibility tree of the target page backend code;

[0006] Step 2: Obtain the visible elements and basic data of the visible elements in the backend code of the target page according to the accessibility tree; obtain the element frames and basic data of the element frames in the target page corresponding to the visible elements according to the screenshot of the target page;

[0007] Step 3: Sort the element boxes of the visible elements from large to small according to their area, and use the visible element corresponding to the element box with the largest area as the root node of the view relationship tree. The corresponding visible elements of the remaining element boxes are used as descendant nodes of the view relationship tree according to the inclusion relationship of the area size, thereby forming a complete view relationship tree;

[0008] Step 4: Use the view relationship tree, the basic data of visible elements, and the basic data of element frames to construct a view relationship graph, and use the obtained view relationship graph as a data structure for understanding the graphical interactive interface.

[0009] In the above RPA scenario, the data structure construction method for understanding the graphical interactive interface is as follows: In step 1, Selenium and Microsoft Inspect tools are used to obtain screenshots of the target page's front-end code rendering and element boxes of visible elements.

[0010] In the aforementioned RPA scenario, the data structure construction method of the graphical interactive interface is understood. In step 3, the remaining element boxes use the corresponding visible elements as descendant nodes of the view relationship tree according to the inclusion relationship of the area size:

[0011] Starting from the root node of the view relationship tree, if the element box of the currently to-be-inserted visible element C is completely contained by the element box of the root node visible element, then the currently visible element C is a descendant node of the root node visible element. Next, recursively observe the relationship between all the child node visible elements of the root node visible element and the visible element C. If the element box of a child node visible element A of the root node visible element still completely contains the element box of the currently to-be-inserted visible element C, then the to-be-inserted visible element C is the descendant node of the child node visible element A. Repeat this process until a child node visible element B of a certain layer is encountered. At this time, the visible element C is not a descendant node of any other child node visible element of the child node visible element B. Then, the visible element C is the direct child node of the child node visible element B, and is in a parallel relationship with other child node visible elements of the child node visible element B.

[0012] Under the aforementioned RPA scenario, the data structure construction method of the graphical interactive interface is understood. When the element box of the visible element judges the descendant nodes according to the inclusion relationship of the area size: if the element box of the visible element C to be inserted completely overlaps with the element box of the child node, or the intersection and union ratio of the two areas is greater than 0.95, the view relationship tree is constructed normally at this time, but the visible element is assigned a value and regarded as a hidden node.

[0013] In the aforementioned RPA scenario, the data structure construction method of the graphical interactive interface is understood, wherein the basic data of the visible element is the text of the visible element; and the basic data of the element frame is the coordinates of the element frame.

[0014] In the aforementioned RPA scenario, we understand the data structure construction method of the graphical interactive interface. In step 4, the view relationship tree, the basic data of the visible elements, and the basic data of the element box are used to construct the view relationship graph. The specific process is as follows:

[0015] Step 4.1, based on the multimodal model, taking the element frame coordinates and the interface screenshot as input, output the style category of the visible element;

[0016] Step 4.2, based on a predefined semantic model, taking the text of the visible element as input, outputting the semantic category of the visible element;

[0017] Step 4.3, storing the style category and semantic category of the visible element as node attributes in the node of the view relationship tree;

[0018] Step 4.4, based on spatial position, group visible elements with the same function, close position and / or similar pattern into the same group of elements;

[0019] Step 4.5: Based on the semantics and spatial positions of visible elements, elements that are functionally and conceptually generalized to other elements are considered as conceptually related elements;

[0020] Step 4.6: construct bidirectional edges between nodes of the same group of elements, and construct unidirectional edges or semantically relative bidirectional edges between nodes of conceptually related elements, so as to expand the view relationship tree structure into a view relationship graph structure.

[0021] The data structure construction method for understanding the graphical interactive interface in the aforementioned RPA scenario, wherein the similarity pattern includes the element box layout and the semantics of the visible elements;

[0022] Whether the element frame layouts are similar is determined based on the coordinates and sizes of the element frames;

[0023] Whether the semantics of the visible elements are similar is determined by the texts of the visible elements.

[0024] Compared with the prior art, the present invention has the following beneficial effects:

[0025] 1. Compared with the "barrier-free DOM tree" structure, the data structure designed by this solution retains the interface information (such as interface element layout, style, and text semantics) that is consistent with user perception, and contains the necessary information for a complete understanding of user needs.

[0026] 2. The data structure designed by the present invention has cross-platform characteristics, thereby reducing the difficulty of model training or prompting when inputting the AI ​​model.

[0027] 3. Compared with screenshots, the data structure designed by the present invention only retains the information required to execute RPA automation tasks, and ids the layout, style and other information, so that the AI ​​model can easily complete operations such as element positioning and element size and element position judgment, thereby improving the accuracy of the AI ​​model's answer results. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 A schematic diagram of the present invention for sorting visible elements from large to small in terms of area;

[0029] Figure 2 This is a schematic diagram of the present invention using a visible element as a child node of a node;

[0030] Figure 3 This is a schematic diagram of the present invention using a visible element as a direct child node of a node;

[0031] Figure 4 This is a schematic diagram of the situation when the visible elements do not overlap with the tree nodes;

[0032] Figure 5 This is a schematic diagram of the situation when the visible element coincides with the tree node;

[0033] Figure 6 The local fragment graph after the view relationship tree is completed;

[0034] Figure 7 A schematic diagram of the data structure of the view relationship graph. DETAILED DESCRIPTION

[0035] The present invention is further described below in conjunction with the accompanying drawings and embodiments, but they are not intended to limit the present invention.

[0036] Embodiment: A method for constructing a data structure for understanding a graphical interactive interface in an RPA scenario includes the following steps:

[0037] Step 1: Get the screenshot of the target page and the accessibility tree of the target page backend code;

[0038] In this step, the target page includes a web page, a software client or a mobile client;

[0039] Step 2: Obtain the visible elements and basic data of the visible elements in the backend code of the target page according to the accessibility tree; obtain the element frames and basic data of the element frames in the target page corresponding to the visible elements according to the screenshot of the target page;

[0040] In this step, visible elements refer to controls or contents on the page that can be perceived by users, such as buttons, icons, input boxes, and text. Due to the presence of multiple layers of panels or areas on the page, the elements edited in advance in the accessibility tree (dom code) are not always visible. There are special functions in js to distinguish whether the elements in the current web page and client interface are visible. All visible elements can call js functions to calculate the coordinate position and size (x, y, width width, height height). These four values ​​can be used to draw a rectangular box, called an element box, which can just cover the visible elements in the screenshot of the target page.

[0041] In this step, the basic data of the visible element is also obtained, that is, the text of the visible element. The text is generally used to describe the function of the visible element. When obtaining the text (innerText) attribute of the visible element in the accessibility tree, the attributes that are invisible to the user are filtered out, such as class, id, imgurl (image hyperlink, etc.); at the same time, the coordinates corresponding to the element frame are also obtained, that is, the (x, y) mentioned in the previous text. The coordinates of the element frame can be obtained through Selenium and Microsoft Inspect tools. Selenium and Microsoft Inspect tools can also be used to obtain screenshots after the front-end code of the target page is rendered.

[0042] Furthermore, after obtaining the visible element, a unique-id is assigned to the visible element to facilitate subsequent access to the element.

[0043] Step 3: Sort the element frames of the visible elements from large to small by area, and use the visible element corresponding to the element frame with the largest area as the root node of the view relationship tree. The remaining element frames use the corresponding visible elements as descendant nodes of the view relationship tree according to the inclusion relationship of the area size, thereby forming a complete view relationship tree;

[0044] In this step, the specific process of constructing the view relationship tree is to judge from the root node of the view relationship tree. If the element box of the visible element C to be inserted is completely contained by the element box of the root node visible element, then the current visible element C is a descendant node of the root node visible element (corresponding to Figure 2), then recursively observe the relationship between all the child node visible elements of the root node visible element and the visible element C. If the element box of a child node visible element A of the root node visible element still completely contains the element box of the visible element C to be inserted, then the visible element C to be inserted is regarded as the descendant node of the child node visible element A. Repeat this process until a child node visible element B of a certain layer is encountered. At this time, the visible element C is not the descendant node of any other child node visible element of the child node visible element B. Then the visible element C is regarded as the direct child node of the child node visible element B, and is in a parallel relationship with other child node visible elements of the child node visible element B (corresponding to Figure 3 ).

[0045] During the construction process, there may be a situation where the element box of the visible element C to be inserted completely overlaps with the element box of the child node, or the intersection and union ratio of the two areas is greater than 0.95. In this case, the visible element is assigned a value and regarded as a hidden node. It is still constructed normally when building the tree, but it will be treated differently in subsequent use. For example, the situation when the visible element unique-element-2 does not overlap with the tree node unique-element1 is as follows Figure 4 As shown in the figure, unique-element-2 is the child node of unique-element-1, and the node type attribute of string type is assigned the value of normal. It can be seen that the situation when element unique-element-2 overlaps with tree node unique-element1 is as follows Figure 5 As shown, unique-element-2 is the child node of unique-element-1. Since the two nodes overlap, the node type attribute of unique-element-2 is assigned a value of hidden.

[0046] The formatted fragment after the view relationship tree is constructed is as follows Figure 6 As shown in the figure, in the view relationship tree, by adding data structures such as dictionaries and hash tables, the search for elements can be completed in O(n*log(n)) (n is the number of visible elements) in a corresponding time.

[0047] Step 4: Use the view relationship tree, the basic data of visible elements, and the basic data of element frames to construct a view relationship graph, and use the obtained view relationship graph as a data structure for understanding the graphical interactive interface.

[0048] In this step, the process of constructing a view relationship graph based on the view relationship tree is as follows:

[0049] Step 4.1, based on the multimodal model, take the element frame coordinates and the interface screenshot as input, and output the category of the visible element;

[0050] In this step, the multimodal model can use local CV models predefined based on RPA business, such as classification (ViT) and object detection (Yolo). In these local CV models, the element box coordinates and interface screenshots obtained in step 1 are used as input, and the style category of the element is output, such as "input box", "button", "icon", "date panel", etc.

[0051] Step 4.2, taking the text and style of the visible element as input based on a predefined semantic model, outputting the semantic category of the visible element;

[0052] In this step, the semantic model predefined based on the RPA business, such as regular expressions and large language models (LLM), can be used to input element text and style, and output the semantic category of the element, such as title, navigation bar, number, time, and location;

[0053] Step 4.3, storing the style category and semantic category of the element as node attributes in the node of the view relationship tree;

[0054] Step 4.4, based on spatial position, group visible elements with the same function, close position and / or similar pattern into the same group of elements;

[0055] In this step, the similar patterns include element box layout and semantics of visible elements;

[0056] Whether the element frame layout is similar is determined based on the coordinates of the element frame and the size of the element frame; for example, if the x-coordinate and width of the element frame are completely equal, it means that the corresponding elements are similar elements, which can be vertical list content, including Weibo, Baidu search results, etc. in an example;

[0057] Whether the semantics of the visible elements are similar is determined by the text attributes of the visible elements. For example, if the text of a group of elements are all place names (such as Hangzhou, Shanghai, Ningbo), then this group of text elements can be regarded as a "place name" element group, and this element group can be used as an input variable for a screening action in subsequent applications.

[0058] Step 4.5: Based on the semantics and spatial positions of visible elements, elements that are functionally and conceptually generalized to other elements are considered as conceptually related elements;

[0059] For example, if the element "Remember password" is visible on the target page, there is usually a "check box" on the left side of it, so "Remember password" and "check box" are conceptually related elements;

[0060] Step 4.6: construct bidirectional edges between nodes of the same group of elements, and construct unidirectional edges or semantically related bidirectional edges between nodes of conceptually related elements, so as to expand the view relationship tree structure into a view relationship graph structure. The final data structure is as follows: Figure 7 shown.

[0061] For example, take "number of visitors" as a visible element on a page. When it is a node in the view relationship graph, "payment amount", "number of paid sub-orders", ..., "number of views", etc. are elements in the same group as "number of visitors". At this time, there is a bidirectional edge named "same group element" from the "number of visitors" node to the "number of views" node. "0" is a conceptual association element of "number of visitors", where the "number of visitors" node is the key and the "0" node is the value. At this time, the "number of visitors" node has an edge pointing to the "0" node, named "key", and the "0" node has an edge pointing to the "number of visitors" node, named "value".

[0062] Furthermore, based on the data structure in the above embodiment, it can be applied to the analysis of specific areas of web pages, software, and app interfaces. Including but not limited to:

[0063] a. Automatic generation of element anchors (Explanation: Element anchor is an RPA term. The element anchor itself is also an element. It assists in locating the target element by calculating its relationship with the target element in terms of interface layout, semantics, etc.)

[0064] b. Element retrieval (Explanation: When the target element fails, find the target element through images and elements in other areas of the interface)

[0065] c. Batch processing (Explanation: "Batch data capture" and "loop similar elements" are functions that are frequently called by RPA applications. The core of these functions is to find "similar elements". The graph can realize the function of "finding similar elements")

[0066] 2. The data structure of the present invention can also be applied to the understanding and analysis of the entire interface. Including but not limited to:

[0067] a. Interface understanding (Explanation: By extracting the information required for RPA automation tasks, the complete context of the interactive interface can be saved without affecting user understanding. At the same time, the interface screenshot is converted into a tree / graph-based substructure, so that specific elements can be accurately located by ID and interface coordinates)

[0068] b. Process annotation (Explanation: Usually, users do not write enough comments when writing RPA applications, which makes it difficult to debug and maintain the application. Since the graph has saved the complete context of the interactive interface, the operation object of the unannotated process can be determined, thereby adding annotations to the process)

[0069] In summary, the data structure designed by the present invention, compared with the "barrier-free DOM tree" structure, retains interface information that is consistent with user perception (such as: interface element layout, style, text semantics), and contains the necessary information to fully understand user needs. The data structure designed by the present invention has a cross-platform feature, which can reduce the difficulty of model training or prompting when inputting the AI ​​model. Compared with screenshots, the data structure designed by the present invention only retains the information required to perform RPA automation tasks, and ids the layout, style and other information, so that the AI ​​model can easily complete operations such as element positioning and element size and element position judgment, thereby improving the accuracy of the AI ​​model's answer results.

Claims

1. A method for constructing a data structure for understanding a graphical interactive interface in an RPA scenario, characterized in that: The following steps are involved: Step 1: Get the screenshot of the target page and the accessibility tree of the target page backend code; Step 2: Obtain the visible elements and basic data of the visible elements in the backend code of the target page according to the accessibility tree; obtain the element frames and basic data of the element frames in the target page corresponding to the visible elements according to the screenshot of the target page; Step 3: Sort the element frames of the visible elements from large to small by area, and use the visible element corresponding to the element frame with the largest area as the root node of the view relationship tree. The remaining element frames use the corresponding visible elements as descendant nodes of the view relationship tree according to the inclusion relationship of the area size, thereby forming a complete view relationship tree; Step 4: Use the view relationship tree, the basic data of visible elements, and the basic data of element frames to construct a view relationship graph, and use the obtained view relationship graph as a data structure for understanding the graphical interactive interface.

2. According to claim 1, the method for constructing a data structure for understanding a graphical interactive interface in an RPA scenario is characterized in that: In step 1, Selenium and Microsoft Inspect tools are used to obtain screenshots of the target page's front-end code rendering and element boxes of visible elements.

3. According to claim 1, the method for constructing a data structure for understanding a graphical interactive interface in an RPA scenario is characterized in that: In step 3, the remaining element frames use the corresponding visible elements as descendant nodes of the view relationship tree according to the inclusion relationship of the area size. Specifically: Starting from the root node of the view relationship tree, if the element box of the currently to-be-inserted visible element C is completely contained by the element box of the root node visible element, then the currently visible element C is a descendant node of the root node visible element. Next, recursively observe the relationship between all the child node visible elements of the root node visible element and the visible element C. If the element box of a child node visible element A of the root node visible element still completely contains the element box of the currently to-be-inserted visible element C, then the to-be-inserted visible element C is the descendant node of the child node visible element A. Repeat this process until a child node visible element B of a certain layer is encountered. At this time, the visible element C is not a descendant node of any other child node visible element of the child node visible element B. Then, the visible element C is the direct child node of the child node visible element B, and is in a parallel relationship with other child node visible elements of the child node visible element B.

4. The method for constructing a data structure for understanding a graphical interactive interface in an RPA scenario according to claim 1, characterized in that: When the element box of the visible element judges the descendant nodes according to the inclusion relationship of the area size: if the element box of the visible element C to be inserted completely overlaps with the element box of the child node, or the intersection and union ratio of the two areas is greater than 0.95, the view relationship tree is constructed normally, but the visible element is assigned a value and regarded as a hidden node.

5. According to claim 1, the method for constructing a data structure for understanding a graphical interactive interface in an RPA scenario is characterized in that: The basic data of the visible element is the text of the visible element; the basic data of the element frame is the coordinates of the element frame.

6. The method for constructing a data structure for understanding a graphical interactive interface in an RPA scenario according to claim 4, characterized in that: In step 4, the view relationship graph is constructed using the view relationship tree, the basic data of visible elements, and the basic data of element frames. The specific process is as follows: Step 4.1, based on the multimodal model, taking the element frame coordinates and the interface screenshot as input, output the style category of the visible element; Step 4.2, based on a predefined semantic model, taking the text of the visible element as input, outputting the semantic category of the visible element; Step 4.3, storing the style category and semantic category of the visible element as node attributes in the node of the view relationship tree; Step 4.4, based on spatial position, group visible elements with the same function, close position and / or similar pattern into the same group of elements; Step 4.5: Based on the semantics and spatial positions of visible elements, elements that are functionally and conceptually generalized to other elements are considered as conceptually related elements; Step 4.6: construct bidirectional edges between nodes of the same group of elements, and construct unidirectional edges or semantically relative bidirectional edges between nodes of conceptually related elements, so as to expand the view relationship tree structure into a view relationship graph structure.

7. The method for constructing a data structure for understanding a graphical interactive interface in an RPA scenario according to claim 6, characterized in that: The pattern similarity includes similarity of element box layout and semantics of visible elements; Whether the element frame layouts are similar is determined based on the coordinates and sizes of the element frames; Whether the semantics of the visible elements are similar is determined by the texts of the visible elements.