An automatic method for finding RPA element anchor points based on multimodal model
Through the multimodal model-based method, web page screenshots and element attributes are used to identify anchor elements, the shortcomings of traditional methods in identifying anchor points in dynamic environments are solved, and more accurate and reliable automatic anchor points are achieved to adapt to more complex web page changes.
Patent Information
- Application Number
- CN202510185798.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-20
AI Technical Summary
The traditional rules-based automatic anchor point search method cannot fully obtain the global structure and context information of the web page, making it difficult to meet the needs of dynamic adjustment in frequently updated application environments.
The multimodal model is used to dynamically identify the anchor elements of the web page through web page screenshots, element coordinates and element categories, and the multimodal model is used to perform vector transformation, vector alignment, element distinction, element attention and element judgment, and the target anchor elements are detected and determined.
It realizes more accurate search for anchor points in all scenarios, adapts to more complex web page changes, improves the reliability and scope of application of recognition, and ensures the stable operation of the RPA system in a dynamic environment.
Smart Images

Figure CN119669600B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robotic process automation, and in particular to an automatic search method for RPA element anchors based on a multimodal model. Background Art
[0002] In today's wave of digital transformation, Robotic Process Automation (RPA) has become an important tool for many companies to optimize business processes and improve efficiency. RPA simulates human operations and automatically performs various repetitive tasks to reduce human intervention. However, when designing and deploying RPA systems, accurately identifying and locating elements of the user interface becomes a key step. The accuracy and stability of elements directly affect the execution efficiency and reliability of RPA. The traditional rule-based automatic anchor point search method only inputs fixed information such as element coordinates and attributes, and cannot fully obtain the global structure and context information of the web page. Therefore, the rule-based method cannot generalize the rules to all web pages, and can only cover known scenarios that can be exhaustively enumerated by the rules. In addition, the traditional multimodal large model only has text and image input, and the attributes such as coordinates, which are very important in web page elements, are not taken seriously by the model when input in natural language. If the system interface is adjusted or updated, the RPA script may not run normally, which will affect business continuity. Especially in the application environment with frequent updates, the traditional method is difficult to meet the needs of dynamic adjustment. Summary of the invention
[0003] The purpose of the present invention is to provide an automatic RPA element anchor search method based on a multimodal model. The present invention dynamically identifies the anchor elements of a web page through web page screenshots, element coordinates and element categories, and has the advantages of wide application range and reliable recognition.
[0004] The technical solution of the present invention is a method for automatically finding RPA element anchor points based on a multimodal model, which is performed according to the following steps:
[0005] Step S1: obtaining the coordinates, text information and area screenshots of the target element and candidate anchor elements in the web page;
[0006] Step S2: according to the region screenshot, the element categories of the target element and the candidate anchor point elements are obtained through the object detection model;
[0007] Step S3: input the coordinates, text information, area screenshots and element categories of the target element and the candidate anchor element into the multimodal model, and use the multimodal model to perform vector conversion, vector alignment, element distinction, element attention and element judgment to detect and determine the target anchor element;
[0008] Step S4: Record the relative relationship between the target element and the anchor element for use in locating the target element during the RPA runtime.
[0009] In the above-mentioned RPA element anchor automatic search method based on the multimodal model, the coordinates and text information of the target element and the candidate anchor element are obtained from the DOM tree of the web page.
[0010] In the aforementioned RPA element anchor point automatic search method based on the multimodal model, the detection process of the target detection model in step S2 is performed according to the following steps:
[0011] Step S2.1: converting the webpage screenshot into a feature map through a feature extraction network;
[0012] Step S2.2: Generate a detection box in the feature map based on the coordinates of the target element and the candidate anchor element through the region proposal network;
[0013] Step S2.3: perform bounding box regression on the detection box to adjust the boundary position;
[0014] Step S2.4: Determine whether there is a target in the detection frame through the classification layer, and if there is a target, distinguish the object category in the detection frame;
[0015] Step S2.5: Remove duplicate and overlapping detection frames from the detection frames.
[0016] In the aforementioned RPA element anchor automatic search method based on the multimodal model, the detection box of step S2.5 is processed by the non-maximum suppression method.
[0017] In the aforementioned RPA element anchor point automatic search method based on a multimodal model, the multimodal model detection process of step S3 is performed according to the following steps:
[0018] Step S3.1: combining the text information of the target element and the candidate anchor element with their corresponding categories to form a text input, taking a screenshot of the target element and the candidate anchor element area as an image input, and extracting coordinate features from the coordinates of the target element and the candidate anchor element as a coordinate input;
[0019] Step S3.2: Use the embedding conversion model to convert the text input and image input into text vectors and image vectors, and then align the text vectors, image vectors and coordinate feature vectors through a fully connected layer;
[0020] Step S3.3: The target element and the candidate anchor point elements are distinguished by rotating the position encoding, and then the text vector, image vector and coordinate feature vector of the target element are made to pay attention to the text vector, image vector and coordinate feature vector of the candidate anchor point elements through the self-attention module, and the weight of each element is obtained by similarity and the vector is weighted summed;
[0021] Step S3.4: Output the candidate anchor element with the highest sum value through the fully connected layer as the anchor element.
[0022] In the aforementioned RPA element anchor point automatic search method based on a multimodal model, in step S3.1, coordinate features including distance, angle and margin are calculated by using the coordinate difference between each candidate anchor point element and the target element.
[0023] In the aforementioned RPA element anchor automatic search method based on a multimodal model, the embedded conversion model in step S3.2 includes a BERT model and a ViT model. The text input is converted by the BERT model, and the image input is converted by the ViT model.
[0024] In the aforementioned RPA element anchor automatic search method based on the multimodal model, the fully connected layer in step S3.4 generates a one-dimensional vector with an information vector representing the target element and a key vector representing the candidate anchor element from the candidate anchor point, and obtains the scores of the information vector and the key vector through the Softmax function. If the information vector has the highest score, it is judged that it has no anchor element, otherwise the key vector with the highest score is used as the found anchor element.
[0025] Compared with the prior art, the present invention uses web page screenshots as input to perform an overall analysis of the web page structure, cooperates with the target detection model to further obtain element categories, and inputs the regional screenshots, element coordinates and element categories as features into the multimodal model for anchor point detection, which is used to help more accurately locate and distinguish elements; the multimodal large model of the present invention can use the entire web page screenshot as input, and analyze the web page structure as a whole, so as to more accurately find anchor points in all scenarios, and the additional input feature attributes can effectively adapt to web page changes including scaling and updating, making the model more sensitive to elements, thereby achieving better results in finding anchor points. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 Schematic diagram of the process of step S1 of the present invention;
[0027] Figure 2 Schematic diagram of the process of step S2 of the present invention;
[0028] Figure 3 Schematic diagram of the process of step S3.1 and step S3.2 of the present invention;
[0029] Figure 4 It is a schematic diagram of the process of step S3.3 and step S3.4 of the present invention. DETAILED DESCRIPTION
[0030] The present invention is further described below in conjunction with the accompanying drawings and embodiments, but they are not intended to limit the present invention.
[0031] Embodiment: A method for automatically finding RPA element anchor points based on a multimodal model, as shown in the attached Figure 1 As shown, follow the steps below:
[0032] Step S1: As shown in the attached Figure 1 As shown, the coordinates, text information and area screenshots of the target element and the candidate anchor element are obtained from the DOM tree of the web page; the DOM tree (Document Object Model Tree) of the web page is a tree model used to represent the structure of an HTML or XML document. It organizes the web page content (such as tags, text, attributes) in a tree structure, which facilitates the browser to parse and render the web page, and allows programming languages such as JavaScript to dynamically operate the web page content. The DOM tree example is shown below:
[0033] Bash
[0034] html
[0035] │
[0036] ├── head
[0037] │ └── title
[0038] │ └── Example page (text node)
[0039] │
[0040] └── body
[0041] └── div id="container"
[0042] ├── h1
[0043] │ └── Page title (text node)
[0044] ├── p class="intro"
[0045] │ └── Introduction paragraph. (text node)
[0046] └── ul
[0047] ├── li
[0048] │ └── text1 (text node)
[0049] ├── li
[0050] │ └── text2 (text node)
[0051] └── li
[0052] └── text3 (text node);
[0053] Get several leaf nodes that are close to the target element from the DOM tree as candidate anchor elements, and get the coordinates and text information of the target element and the candidate anchor elements in the DOM. The coordinates of the elements can be obtained with the help of the getBoundingClientRect() method in the DOM interface. At the same time, get a screenshot of the area on the web page based on the coordinates of the target element and the candidate anchor elements.
[0054] Step S2: As shown in the attached Figure 2 As shown in the figure, the target element and the element category of the candidate anchor element are obtained through the target detection model; Object Detection is a computer vision task that aims to identify and locate specific objects in images or videos. Unlike image classification, target detection not only needs to identify the object category in the image, but also needs to accurately find the location of these objects on the image; the target detection model includes a feature extraction network, a region proposal network, a bounding box regression module, a classification layer, and a non-maximum suppression module. The detection process is carried out in the following steps:
[0055] Step S2.1: Convert the webpage screenshot into a feature map through a feature extraction network; the feature extraction network uses a convolutional neural network (CNN) model, such as ResNet, VGG, and MobileNet, which converts image data into a feature map and carries spatial and semantic information;
[0056] Step S2.2: Generate a detection box in the feature map based on the coordinates of the target element and the candidate anchor element through the region proposal network; the region proposal network (RPN) is a component for detecting specific objects, which can generate potential candidate boxes (i.e., detection boxes) containing the target or background. The RPN will then screen and classify each candidate box. The region proposal methods include sliding window-based frameworks, anchor-based methods (such as YOLO and SSD), and anchor-free methods (such as FCOS);
[0057] Step S2.3: performing bounding box regression on the detection box to adjust the boundary position; the bounding box regression module calculates the precise boundary position of the object through regression and fine-tunes the position and size of the detection box;
[0058] Step S2.4: Determine whether there is a target in the detection frame through the classification layer, and if there is a target, distinguish the object category in the detection frame;
[0059] Step S2.5: Remove duplicate and overlapping detection frames in the detection frame by non-maximum suppression. For example, when the IoU (intersection over union) of two detection frames exceeds the threshold of 0.5, retain the detection frame with higher confidence. The detection frame that is most likely to be a real object is retained, and the overlapping low-score detection frames are eliminated to improve the accuracy of the detection results.
[0060] Before detection, the open source YOLO model is used as the target detection model for training on the web page. Each element on the web page is divided into text, icons, buttons, input boxes, drop-down boxes, and check boxes. These elements are manually labeled with categories in the dataset, and these categories are used for model training. In addition, the coordinates of the target element and the candidate anchor element obtained in advance can be input into the target detection model, so that it can more accurately return the type of element in the target area.
[0061] Step S3: input the coordinates, text information, area screenshots and element categories of the target element and the candidate anchor element into the multimodal model, and use the multimodal model to perform vector conversion, vector alignment, element distinction, element attention and element judgment on the input information to detect and determine the target anchor element;
[0062] The detection process of the multimodal model is carried out in the following steps:
[0063] Step S3.1: As attached Figure 3 As shown, the text information of the target element and the candidate anchor element is combined with their corresponding categories to form a text input, the target element and the candidate anchor element area screenshots are taken as image input, and coordinate features are extracted from the coordinates of the target element and the candidate anchor element as coordinate input, and the coordinate features include distance, angle and margin, which are calculated according to the coordinate difference between each candidate anchor element and the target element;
[0064] Step S3.2: Use the BERT model and the ViT model to embed the text input and image input respectively, convert them into text vectors and image vectors, and then align the text vectors, image vectors and coordinate feature vectors through the fully connected layer; The BERT (Bidirectional Encoder Representations from Transformers) model is based on the Transformer architecture and is mainly composed of multi-layer bidirectional Transformer encoders; Transformer contains components such as multi-head attention mechanism, position encoding, and feedforward neural network. It first performs word segmentation on the input text, converts each word into a word vector, and adds position vectors and sentence vectors to represent the position of the word in the text and the sentence information to which it belongs. These vectors are input into the multi-layer Transformer encoder. Each layer of Transformer uses the multi-head attention mechanism to allow the model to pay attention to different parts of the text in parallel, capture the contextual dependency between words, and after multi-layer processing, finally output the context representation vector of each word. These vectors integrate the semantic information of the entire text and can be used for a variety of natural language processing tasks; ViT (Vision The ViT model divides the image into multiple fixed-size image blocks, linearly embeds these image blocks, adds position encoding, and then inputs them into the Transformer encoder. The Transformer encoder also contains a multi-head attention mechanism, a feedforward neural network, etc. Unlike traditional convolutional neural networks, the ViT model directly applies the Transformer on the image block sequence and does not rely on convolution operations to extract features; the ViT model divides the input image into a series of image blocks, regards each image block as a "word", and converts it into a vector through linear projection. In order to retain the position information of the image block, position encoding is added. These image block vectors with position information are input into the Transformer encoder. The multi-head attention mechanism establishes connections between image blocks and captures long-distance dependencies in the image. After multi-layer Transformer encoding, the image feature vector is output; the fully connected layer (Fully Connected Layer, FC Layer is a common layer and a core component in traditional neural networks. Its characteristic is that each neuron is connected to all neurons in the previous layer to form a dense connection structure. It receives the output vector of the previous layer, sums the signal of each input neuron with the corresponding weight, adds the bias term, and finally performs a nonlinear transformation through the activation function to get the output of the current layer.
[0065] Step S3.3: As shown in the attached Figure 4As shown in the figure, the element coordinates are converted into angles and radii in the polar coordinate system by rotating the position encoding, and the position feature vector is generated by combining the sine function encoding to distinguish the spatial relationship between the target element and the candidate anchor point element. Then, the text vector, image vector and coordinate feature vector of the target element and the text vector, image vector and coordinate feature vector between the candidate anchor point elements are mutually concerned by the self-attention module, and the weight of each element is obtained by similarity and the vector is weighted and summed. The self-attention module is a mechanism for dynamically paying attention to different parts of the input sequence, which is used to capture the association between elements in sequence data (such as text vectors, image vectors and coordinate feature vectors), and assign different weights to each element in the sequence so that the model can automatically pay attention to important information and ignore irrelevant information.
[0066] Step S3.4: The output of the self-attention module is used to generate a one-dimensional vector containing an information vector representing the target element and a key vector representing the candidate anchor element through a fully connected layer. The scores of the information vector and the key vector are obtained through a Softmax function. If the information vector has the highest score, it is judged that it has no anchor element. Otherwise, the key vector with the highest score is used as the found anchor element.
[0067] Step S4: If there is an anchor point, record the relative relationship between the target element and the anchor element for target element positioning during the run of Robotic Process Automation (RPA); by recording the relatively fixed anchor element information and the relatively fixed relative position of the anchor point and the target element, the less stable target element can be positioned more stably; the relative position of the anchor point and the target element includes the angle and distance from the anchor point to the target element, where the distance is based on the size of the element itself and can adapt to the scaling changes of the page; when locating the target element, first find the anchor element in a traditional way, and then calculate the cosine similarity of the angle and the distance similarity, so that the most similar element that meets the threshold can be regarded as the target element.
[0068] In summary, the present invention uses a multimodal large model, which can take a screenshot of the entire web page as input and analyze the structure of the web page as a whole, so that it can find anchor points more accurately in all scenarios, and can have more complete input when searching, and can adapt to more complex web page changes. The traditional rule-based automatic anchor point search method only inputs fixed information such as element coordinates and attributes, and cannot see the full picture of the web page. Therefore, the rule-based method cannot generalize the rules to all web pages, and can only cover known scenarios that can be exhaustively enumerated by the rules; the multimodal large model architecture of the present invention adds the attributes of element coordinates and element categories, making the model more sensitive to the above-mentioned element attributes, thereby playing a better effect in the task of finding anchor points. The traditional multimodal large model only has text and image inputs, and the attributes such as coordinates, which are very important in web page elements, are input in a natural language manner and cannot be taken seriously by the model.
Claims
1. A method for automatically finding RPA element anchor points based on a multimodal model, characterized by: Follow these steps: Step S1: obtaining the coordinates, text information and area screenshots of the target element and candidate anchor elements in the web page; Step S2: according to the region screenshot, the element categories of the target element and the candidate anchor point elements are obtained through the object detection model; Step S3: input the coordinates, text information, area screenshots and element categories of the target element and the candidate anchor element into the multimodal model, and use the multimodal model to perform vector conversion, vector alignment, element distinction, element attention and element judgment to detect and determine the target anchor element; Step S4: Record the relative relationship between the target element and the anchor element for use in locating the target element during the RPA runtime; The detection process of the target detection model in step S2 is performed according to the following steps: Step S2.1: converting the webpage screenshot into a feature map through a feature extraction network; Step S2.2: Generate a detection box in the feature map based on the coordinates of the target element and the candidate anchor element through the region proposal network; Step S2.3: perform bounding box regression on the detection box to adjust the boundary position; Step S2.4: Determine whether there is a target in the detection frame through the classification layer, and if there is a target, distinguish the object category in the detection frame; Step S2.5: removing duplicate and overlapping detection frames in the detection frames; The detection process of the multimodal model in step S3 is performed according to the following steps: Step S3.1: combining the text information of the target element and the candidate anchor element with their corresponding categories to form a text input, taking a screenshot of the target element and the candidate anchor element area as an image input, and extracting coordinate features from the coordinates of the target element and the candidate anchor element as a coordinate input; Step S3.2: Use the embedding conversion model to convert the text input and image input into text vectors and image vectors, and then align the text vectors, image vectors and coordinate feature vectors through a fully connected layer; Step S3.3: The target element and the candidate anchor point elements are distinguished by rotating the position encoding, and then the text vector, image vector and coordinate feature vector of the target element are made to pay attention to the text vector, image vector and coordinate feature vector of the candidate anchor point elements through the self-attention module, and the weight of each element is obtained by similarity and the vector is weighted summed; Step S3.4: Output the candidate anchor element with the highest sum value through the fully connected layer as the anchor element.
2. The method for automatically finding RPA element anchor points based on a multimodal model according to claim 1 is characterized in that: The coordinates and text information of the target element and the candidate anchor element are obtained from the DOM tree of the web page.
3. The method for automatically finding RPA element anchor points based on a multimodal model according to claim 1 is characterized in that: The detection box of step S2.5 is processed by the non-maximum suppression method.
4. The method for automatically finding RPA element anchor points based on a multimodal model according to claim 1 is characterized in that: In step S3.1, coordinate features including distance, angle and margin are calculated by using the coordinate difference between each candidate anchor point element and the target element.
5. The method for automatically finding RPA element anchor points based on a multimodal model according to claim 1 is characterized in that: The embedded conversion models in step S3.2 include a BERT model and a ViT model. The text input is converted by the BERT model, and the image input is converted by the ViT model.
6. The method for automatically finding RPA element anchor points based on a multimodal model according to claim 1 is characterized in that: In step S3.4, the fully connected layer generates a one-dimensional vector with an information vector representing the target element and a key vector representing the candidate anchor element from the candidate anchor point, and obtains the scores of the information vector and the key vector through the Softmax function. If the information vector has the highest score, it is judged that it has no anchor element, otherwise the key vector with the highest score is used as the found anchor element.
Citation Information
Patent Citations
Webpage login method and system based on fixed candidate box target detection algorithm
CN115905767A
Construction method and system of universal RPA check box operation component
CN116168405A