An Anchor-based Automatic Repair Method for RPA Element Anchors
By combining vector and detached vector repair methods, object detection and image feature matching are used to solve the problem of unstable element positioning in the RPA system, and efficient repair of input boxes, drop-down boxes, check boxes, icons and text elements is achieved, improving the stability and accuracy of the automation process.
Patent Information
- Application Number
- CN202510518198.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-24
AI Technical Summary
In existing RPA systems, element repair technology is difficult to effectively locate target elements when the target system interface changes dynamically, resulting in interruption or failure of automation processes, especially for input boxes, drop-down boxes, check boxes, icons and text elements. The repair coverage of input boxes, drop-down boxes, check boxes, icons and text elements is not wide or the error rate is high.
The automatic repair method of RPA element anchor points based on anchor points is adopted. By combining vector repair and disengagement vector repair, candidate elements are obtained using object detection and image feature matching, scaling ratio and region overlap are calculated during vector repair, and rules or models are repaired according to trusted anchor elements to ensure the accuracy of element positioning.
It realizes extensive repair of multiple elements, reduces the repair error rate, improves the stability and scope of application of the process, especially the repair effect of input boxes, drop-down boxes, check boxes, icons and text elements.
Smart Images

Figure CN120045377B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information detection, and particularly relates to an automatic repair method for RPA element anchors based on anchors. Background Art
[0002] At present, Robotic Process Automation (RPA) has become an important tool for many enterprises to optimize business processes and improve efficiency; in RPA, the stability and reliability of automated processes highly depend on the ability of robots to accurately locate and interact with target interface elements. However, due to the dynamic changes in the target system interface (such as interface layout adjustment, element attribute update, or version upgrade), the robot may not be able to locate the originally set elements, resulting in the interruption or even failure of the automated process. This problem is particularly common in the production environment, posing challenges to the process efficiency and business continuity of enterprises. To solve this problem, element repair technology has gradually become a key technology in the RPA system; the element repair technology can automatically adjust the robot positioning logic in an intelligent way when the target element changes, so as to re-match and locate the correct target element, thus ensuring the continuous operation of the process. Currently, the main methods for automatic element repair are as follows: 1. Intelligent repair through element paths: Calculate a relatively stable element path through rules, and then when operating on a similar page, the corresponding element can also be found through the corresponding rules to achieve element repair; however, since many pages often modify the page path structure and naming to prevent crawling, there will often be a situation where the path changes while the page visual remains unchanged. Therefore, the element repair relying only on element paths has a narrow coverage. 2. Re-positioning through element images: Based on the screenshot of the target element, perform cv template matching on the current page to achieve element repair; however, for many elements (such as input boxes), when the internal content changes, it cannot be retrieved through template matching. In addition, for many similar elements, such as check boxes, template matching will return multiple elements and cannot achieve element repair either. 3. Re-positioning through anchors and vectors: Starting from the anchor point of the current page through a normalized vector, re-position the target element, but it can only cover the situation where the relative position between the target element and the anchor point has not changed, and the applicable range is limited. Summary of the Invention
[0003] The purpose of the present invention is to provide an automatic repair method for RPA element anchors based on anchors. The present invention can repair various elements by combining vector repair and off-vector repair, and has the advantage of wide coverage; at the same time, the present invention also has the characteristic of low repair error rate.
[0004] The technical solution of the present invention is: an anchor-based RPA element anchor point automatic repair method, which obtains candidate elements through target detection and image feature matching, first performs combined vector repair on the candidate elements, and performs detached vector repair if the combined vector repair fails;
[0005] The process of combined vector repair is first to obtain the coordinates of the original target element, the original anchor element and the current anchor element, then calculate the scaling ratio according to the coordinates of the original anchor element and the current anchor element, and perform scaling calculation on the vector from the current anchor element to the original target element according to the scaling ratio to restore the area of the original target element; calculate the similarity between the area of the original target element and the area of the candidate element, and take the candidate element with the highest similarity value that reaches the set threshold as the repaired target element;
[0006] The process of the off-vector repair is to perform rule repair or model repair on the candidate elements based on the trusted anchor element; the rule repair is to calculate the score according to the distance and angle between the candidate element and the trusted anchor element, and regard the candidate element that meets the specified score threshold as the repaired target element; the model repair searches for the pending anchor element for each candidate element through the multimodal model. If the pending anchor element is the same as the trusted anchor element, the candidate element corresponding to the pending anchor element is the repaired target element.
[0007] In the above-mentioned anchor-based RPA element anchor automatic repair method, the candidate elements include input boxes, drop-down boxes, check boxes, icons and text elements; the rule repair is applied to input boxes, drop-down boxes and check boxes; the model repair is applied to icons and text elements.
[0008] In the aforementioned anchor-based RPA element anchor automatic repair method, for input boxes and drop-down boxes, the rule repair is to rank the candidate elements according to the relative distance and absolute distance from the trusted anchor element to obtain a comprehensive distance score, and obtain the angle score by calculating the angle between the candidate element and the trusted anchor element; the candidate elements whose comprehensive distance score and angle score both meet the specified score threshold are regarded as the repaired target elements; for check boxes, the rule repair obtains the absolute distance score of the candidate elements according to the absolute distance from the trusted anchor element, and the candidate elements whose absolute distance score meets the specified score threshold are regarded as the repaired target elements.
[0009] In the aforementioned anchor-based RPA element anchor point automatic repair method, the comprehensive distance score is calculated as shown in the following formula:
[0010] Comprehensive distance score = relative distance ranking score × absolute distance score;
[0011] Among them, the relative distance ranking score = 1 - (the relative distance ranking of the candidate element and the trusted anchor element / 10);
[0012] The absolute distance score = min(1, the absolute distance between the original target element and the trusted anchor element / the absolute distance between the candidate element and the trusted anchor element);
[0013] The angle score is calculated as shown in the following formula:
[0014] Angle score = 1 - ((the vector angle between the candidate element and the trusted anchor element - the minimum angle directly below and directly to the right of the candidate element) / 90).
[0015] In the aforementioned anchor - based RPA element anchor automatic repair method, the image feature matching is performed according to the following steps:
[0016] The image feature matching is performed according to the following steps:
[0017] Step A1: Obtain the current interface image and the screenshot image of the original target element, and perform pre - processing of grayscale conversion and denoising on the two images;
[0018] Step A2: Detect the points that meet high recognition and high stability in the two images through the Scale - Invariant Feature Transform (SIFT) method to obtain feature points;
[0019] Step A3: Extract the feature descriptors of the feature points through the Scale - Invariant Feature Transform descriptor. The feature descriptor is used to describe the regional features around the feature points;
[0020] Step A4: Compare the feature descriptors of different feature points in the two images through the K - Nearest Neighbor (KNN) algorithm to obtain multiple similar feature points for each feature point in the interface image;
[0021] Step A5: Calculate the distance ratio between the feature descriptor of each feature point and the feature descriptor of the corresponding similar feature point, compare the distance ratio with the set threshold, and the similar feature points exceeding the threshold are incorrect matches, and the rest are correct matches and are used for the next step;
[0022] Step A6: Apply the geometric relationship between the correctly matched feature points and the similar feature points to the interface image by solving the transformation matrix to obtain multiple similar regions, and use the similar regions as candidate elements.
[0023] In the aforementioned anchor - based RPA element anchor automatic repair method, the scaling ratio is the ratio of the short - side length of the current anchor element to the short - side length of the original anchor element; the scaling calculation is as shown in the following formula:
[0024] The coordinates of the current target element = the coordinates of the current anchor element + the scaling ratio × (the coordinates of the original target element - the coordinates of the original anchor element).
[0025] In the aforementioned method for automatically repairing the anchor points of RPA elements based on anchor points, the calculation of the similarity is as follows: first, calculate the regional overlap degree and the edge alignment degree between the region of the original target element and the region of the candidate element, and then obtain the similarity through the product of the regional overlap degree and the edge alignment degree;
[0026] The degree of regional overlap is obtained by calculating the intersection over union (IoU) of the regions of the candidate element and the original target element; the IoU is the ratio of the area of the intersection of the regions of the candidate element and the original target element to the area of the union of the regions of the candidate element and the original target element;
[0027] The calculation of the edge alignment degree is shown by the following formula:
[0028] Edge alignment degree = 1 - min (the minimum margin on both sides of the left side of the candidate element / the length of the region of the original target element, the minimum margin on both sides of the right side of the candidate element / the length of the region of the original target element, the minimum margin on both sides of the upper side of the candidate element / the width of the region of the original target element, the minimum margin on both sides of the lower side of the candidate element / the width of the region of the original target element).
[0029] In the aforementioned method for automatically repairing the anchor points of RPA elements based on anchor points, the process of the multi-modal model to find the pending anchor point elements is carried out according to the following steps:
[0030] Step S1: Obtain the coordinates, text information, and regional screenshots of the candidate elements and the corresponding candidate anchor point elements;
[0031] Step S2: According to the regional screenshots, obtain the element categories of the candidate elements and the candidate anchor point elements through the object detection model;
[0032] Step S3: Input the coordinates, text information, regional screenshots, and element categories of the candidate elements and the candidate anchor point elements into the multi-modal model, and use the multi-modal model to detect and determine the pending anchor point elements.
[0033] In the aforementioned method for automatically repairing the anchor points of RPA elements based on anchor points, the detection process of the multi-modal model is carried out according to the following steps:
[0034] Step S3.1: Form a text input from the text information of the candidate element, the text information of the candidate anchor point element, the category of the candidate element, and the category of the candidate anchor point element; use the regional screenshots of the candidate element and the candidate anchor point element as the image input; extract the coordinate features from the coordinates of the candidate element and the candidate anchor point element as the coordinate input;
[0035] Step S3.2: Use the embedding transformation model to transform the text input and the image input into text vectors and image vectors, and then align the text vectors, image vectors, and coordinate feature vectors through the fully connected layer;
[0036] Step S3.3: Distinguish candidate elements and candidate anchor elements through rotational position encoding, and then, through the self-attention module, enable the text vectors, image vectors, and coordinate feature vectors of the candidate elements to mutually attend to the text vectors, image vectors, and coordinate feature vectors of the candidate anchor elements, obtain the weights of each element through similarity, and perform weighted summation of the vectors.
[0037] Step S3.4: Output the candidate anchor element with the highest summation value through the fully connected layer as the to-be-determined anchor element.
[0038] Compared with the prior art, the present invention first obtains candidate elements through object detection and image feature matching, and first performs combined vector repair on the candidate elements. If the combined vector repair fails, vector detachment repair is performed; the combined vector repair calculates the region of the original target element through scaling calculation, and further obtains the similarity through the coincidence degree of the combined region and the edge alignment degree, and compares it with the threshold to obtain the repaired target element, realizing the repair of elements with scaling changes; the process of vector detachment repair is to perform regular repair or model repair on the candidate elements based on trustworthy anchor elements. Regular repair formulates rules using the position rules of elements, calculates their distance and angle scores according to the rules, and realizes element repair when the scores reach the set threshold; model repair searches for to-be-determined anchor elements through a multimodal model, and the candidate element corresponding to the to-be-determined anchor element found and the trustworthy anchor element is the repaired target element. Vector detachment repair realizes the repair of target elements with position and content changes, and has a wide repair coverage; at the same time, the present invention obtains candidate elements through object detection and image feature matching, narrows the range of candidate elements, and thus reduces the error rate of element repair. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is the flowchart of the present invention;
[0040] Figure 2 is the positional relationship between the original anchor element and the original target element during the development of the present invention;
[0041] Figure 3 is the repair schematic diagram of candidate elements in the combined vector repair of the present invention;
[0042] Figure 4 is the repair schematic diagram of regular repair of the present invention;
[0043] Figure 5 is the repair schematic diagram of model repair of the present invention;
[0044] Figure 6 is the input schematic diagram of the multimodal model of the present invention;
[0045] Figure 7It is a schematic diagram of the detection process of step S3.-1 to step S3.2 of the present invention;
[0046] Figure 8 It is a schematic diagram of the detection process of step S3.3-step S3.4 of the present invention. DETAILED DESCRIPTION
[0047] The present invention is further described below in conjunction with the accompanying drawings and embodiments, but they are not intended to limit the present invention.
[0048] Embodiment: An RPA element anchor point automatic repair method based on anchor points, as shown in the attached Figure 1 As shown, candidate elements are obtained through target detection and image feature matching, and the candidate elements are first repaired by combining vectors. If the combined vector repair fails, the trusted anchor element is obtained. If a trusted anchor element exists, the detached vector repair is performed. If the group exists, the repair fails.
[0049] Object detection is a computer vision task that aims to identify and locate specific objects in images or videos. Unlike image classification, object detection not only needs to identify the categories of objects in the image, but also needs to accurately find the locations of these objects on the image. The types of target elements should be recorded in advance during the application development period. When repairing, all elements of this type on the page can be obtained as candidate elements through object detection.
[0050] The image feature matching is performed according to the following steps:
[0051] Step A1: Obtain the current interface image and the original target element screenshot image, and perform grayscale and denoising preprocessing on the two images. Grayscale is to convert a color image into a grayscale image to reduce computational complexity, and denoising is to use filtering technology (such as Gaussian filtering) to remove noise in the image to avoid affecting the effect of feature extraction;
[0052] Step A2: Feature points are points with high recognition and stability in the image, usually at the corners and edges of the image. The points with high recognition and high stability in the two images are detected by the scale-invariant feature transformation method to obtain feature points with scale and rotation invariance.
[0053] Step A3: Extract the feature descriptors of the feature points to describe the features of the surrounding areas through Scale-Invariant Feature Transform (SIFT) descriptors, which are rotation and scale invariant based on local image gradients. The SIFT algorithm first constructs the scale space of the image through Gaussian blur to detect potential key points. To make the features rotation invariant, the SIFT algorithm assigns a main direction to each key point by calculating the gradient direction histogram of the area around the key point to determine the main direction. After determining the position and direction of the key point, the algorithm generates a feature descriptor to describe the local image area around the key point. After generating the feature descriptors of all key points, the feature points in different images can be matched by calculating the distances between the feature descriptors.
[0054] Step A4: Compare the feature descriptors of different feature points in the two images through the K-Nearest Neighbor (KNN) algorithm to obtain multiple similar feature points for each feature point in the interface image. That is, in the feature space, if most of the k nearest (i.e., the closest in the feature space) samples near a sample belong to a certain category, then this sample also belongs to this category.
[0055] Step A5: Calculate the distance ratio (such as the distance ratio between the nearest neighbor and the second nearest neighbor) between the feature descriptor of each feature point and the feature descriptor of the corresponding similar feature point, and compare the distance ratio with a set threshold. The similar feature points that exceed the threshold are mis-matches, and the rest are correct matches and are used for the next step. For example, there are and two images, and their SIFT feature points and feature descriptors are extracted respectively. The set of feature descriptors of image is: , image 's set of feature descriptors is: , where the feature descriptor is a 128-dimensional vector; for the feature descriptor in image , find its nearest neighbor and second nearest neighbor in image , and calculate the ratio of the nearest neighbor distance to the second nearest neighbor distance through . If the ratio is less than the threshold (0.7 - 0.8), it is a correct match.
[0056] Step A6: Apply the geometric relationship between the correctly matched feature points and similar feature points to the interface image by solving the transformation matrix (such as affine transformation, perspective transformation) to obtain multiple similar regions, and use the similar regions as candidate elements.
[0057] The process of solving the transformation matrix is as follows:
[0058] Step A6.1: Prepare matching point pairs. For example, there are N pairs of correctly matched feature points and similar feature points. The point set in image is , and the point set in image is ;
[0059] Step A6.2: The perspective transformation matrix satisfies:
[0060] ;
[0061] After expansion, we get:
[0062] ;
[0063] ;
[0064] Further linearize the equation to get:
[0065] ;
[0066] ;
[0067] After arrangement, we get the linear equation:
[0068] ;
[0069] For N pairs of correctly matched feature points and similar feature points, construct a 2N×8 matrix A and a 2N×1 vector , and the two satisfy:
[0070] ;
[0071] where , (fixed as 1 to eliminate scale uncertainty);
[0072] Step A6.3: Solve the transformation matrix. Use the least squares method to solve the linear equations:
[0073] ;
[0074] Step A6.4: Convert the obtained into a perspective transformation matrix :
[0075] .
[0076] As shown in Appendix Figure 2 and Appendix Figure 3As shown, the process of binding vector repair is to obtain the coordinates of the original target element, the original anchor element (obtained during the development period), and the current anchor element, and then calculate the scaling ratio based on the coordinates of the original anchor element and the current anchor element. The scaling ratio is the ratio of the length of the short side of the current anchor element to the length of the short side of the original anchor element. The vector from the current anchor element to the original target element is scaled according to the scaling ratio to restore the original target element area. The scaling calculation is shown in the following formula:
[0077] Coordinates of the current target element = Coordinates of the current anchor element + Scaling ratio × (Coordinates of the original target element - Coordinates of the original anchor element);
[0078] The similarity between the candidate element and the original target element is obtained by multiplying the overlapping degree of the areas of the original target element and the candidate element and the edge alignment degree. The candidate element with the highest similarity that reaches the set threshold is taken as the repaired target element. Among them, the overlapping degree of the areas is obtained by calculating the intersection-over-union ratio of the areas of the candidate element and the original target element. The intersection-over-union ratio is the ratio of the intersection area of the areas of the candidate element and the original target element to the union area of the areas of the candidate element and the original target element. The calculation of the edge alignment degree is shown in the following formula:
[0079] Edge alignment degree = 1 - min (Minimum margin of the left two sides of the candidate element / Length of the area of the original target element, Minimum margin of the right two sides of the candidate element / Length of the area of the original target element, Minimum margin of the upper two sides of the candidate element / Width of the area of the original target element, Minimum margin of the lower two sides of the candidate element / Width of the area of the original target element).
[0080] The process of unbinding vector repair is to perform regular repair or model repair on the candidate element based on the trusted anchor element. The trusted anchor element is manually captured by the user or automatically found by the model;
[0081] As shown in the appendix Figure 4 For input boxes and dropdown boxes, the regular repair ranks the candidate elements according to the relative distance and absolute distance between the candidate element and the trusted anchor element to obtain a comprehensive distance score. The angle score is obtained by calculating the angle between the candidate element and the trusted anchor element. Since for input boxes and dropdown boxes, the target element is usually directly below, directly to the right, or to the lower right of the anchor element, candidate elements that meet the specified score threshold are regarded as the repaired target elements. Among them, the calculation of the comprehensive distance score is shown in the following formula:
[0082] Comprehensive distance score = Relative distance ranking score × Absolute distance score;
[0083] Among them, Relative distance ranking score = 1 - (Relative distance ranking of the candidate element and the trusted anchor element (ranked from 0 in ascending order / 10));
[0084] Absolute distance score = min(1, absolute distance between the original target element and the trusted anchor element / absolute distance between the candidate element and the trusted anchor element);
[0085] The angle score is calculated as shown in the following formula:
[0086] Angle score = 1 - ((vector angle between the candidate element and the trusted anchor element - minimum angle directly below and directly to the right of the candidate element) / 90). When the angle exceeds this region (i.e., 90°), the score is zero;
[0087] For checkboxes, the target element is usually the first element directly to the left of the anchor element. Therefore, only the absolute distance score of the first element on the left needs to be calculated, and the candidate element that meets the specified score threshold is considered the repaired target element.
[0088] As shown in the appendix Figure 5 As shown, the model repair uses a multimodal model to find a pending anchor element for each candidate element. If the pending anchor element is the same as the trusted anchor element, the candidate element corresponding to the pending anchor element is the repaired target element;
[0089] As shown in the appendix Figure 6 As shown, the process of the multimodal model finding the pending anchor element is carried out according to the following steps:
[0090] Step S1: Obtain several leaf nodes close to the target element from the DOM tree as candidate anchor elements. At the same time, obtain the coordinates and text information of the candidate element and the candidate anchor element in the DOM. The coordinates of the element can be obtained by using the getBoundingClientRect() method in the DOM interface. At the same time, obtain the regional screenshot of them on the web page according to the coordinates of the candidate element and the candidate anchor element;
[0091] Step S2: Convert the web page screenshot into a feature map through a feature extraction network; the feature extraction network uses a convolutional neural network (CNN) model, such as ResNet, VGG, and MobileNet. This network converts the image data into a feature map, which carries spatial information and semantic information; generate detection boxes in the feature map according to the coordinates of the target element and the candidate anchor element through a region proposal network; the region proposal network (RPN) is a component for detecting specific objects and can generate potential candidate boxes (i.e., detection boxes). The candidate boxes contain the target or background. Subsequently, the RPN will screen and classify each candidate box. Region proposal methods include sliding window-based frameworks, Anchor-based methods (such as YOLO and SSD), and Anchor-free methods (such as FCOS); perform bounding box regression on the detection boxes to adjust the boundary positions; the bounding box regression module calculates the exact boundary position of the object through regression and fine-tunes the position and size of the detection box; determine whether there is a target inside the detection box through a classification layer. If there is a target, distinguish the object category inside the detection box; remove duplicate and overlapping detection boxes in the detection boxes through non-maximum suppression. For example, when the IoU (Intersection over Union) of two detection boxes exceeds the threshold of 0.5, retain the detection box with a higher confidence; retain the detection box that is most likely to be a real object and eliminate overlapping low-score detection boxes to improve the accuracy of the detection results;
[0092] Step S3: Input the coordinates, text information, regional screenshots, and element categories of the candidate elements and candidate anchor elements into a multi-modal model, and use the multi-modal model for vector transformation, vector alignment, element discrimination, element attention, and element judgment to detect and determine the target anchor element.
[0093] As shown in the appendix Figure 7 and the appendix Figure 8 shown, the detection process of the multi-modal model is carried out according to the following steps:
[0094] Step S3.1: Combine the text information of the candidate elements and candidate anchor elements with their corresponding categories to form a text input, use the regional screenshots of the candidate elements and candidate anchor elements as an image input, and extract coordinate features from the coordinates of the candidate elements and candidate anchor elements as a coordinate input. The coordinate features include distance, angle, and margin, which are calculated according to the coordinate differences between each candidate anchor element and the candidate element;
[0095] Step S3.2: Use an embedding transformation model (BERT model and ViT model) to convert the text input and image input into text vectors and image vectors, and then align the text vectors, image vectors, and coordinate feature vectors through a fully connected layer;
[0096] Step S3.3: Convert the element coordinates into angles and radii in the polar coordinate system through rotational position encoding, and generate a position feature vector by combining sine function encoding to distinguish the spatial relationship between the candidate element and the candidate anchor element. Subsequently, through the self-attention module, the text vector, image vector, and coordinate feature vector of the candidate element and the text vector, image vector, and coordinate feature vector of the candidate anchor element pay attention to each other. Obtain the weights of each element through similarity and perform weighted summation of the vectors; The self-attention module is a mechanism for dynamically paying attention to different parts of the input sequence, used to capture the associations between elements in sequence data (such as text vectors, image vectors, and coordinate feature vectors), assign different weights to each element in the sequence, so that the model can automatically focus on important information and ignore irrelevant information;
[0097] Step S3.4: The output of the self-attention module passes through a fully connected layer to generate a one-dimensional vector of the candidate anchor with an information vector representing the target element and a key vector representing the candidate anchor element. After passing through the Softmax function, the scores of the information vector and the key vector are obtained. If the score of the information vector is the highest, it is judged that there is no pending anchor element, otherwise the key vector with the highest score is used as the found pending anchor element.
[0098] In summary, the present invention first obtains candidate elements through object detection and image feature matching. Traditional element repair methods do not filter the candidate range of the target element. However, the initial element type and screenshot of the target element are attributes that well depict the portrait of the target element. Use these attributes to obtain candidate elements to effectively narrow the element range. And after filtering by type and image features, the probability of repairing to the wrong element can be reduced; The present invention adds a scaling calculation on the basis of the original vector positioning, obtains the similarity by combining the degree of regional overlap and the degree of edge alignment, compares it with the threshold to obtain the repaired target element, and increases the applicable range; The present invention includes rule repair applied to input boxes, dropdown boxes, and checkboxes, and model repair applied to icons and text elements. Rule repair formulates rules using the position rules of elements, calculates their distance and angle scores according to the rules, and realizes element repair when the score reaches the set threshold; Model repair searches for pending anchor elements through a multi-modal model, and the candidate elements corresponding to the found pending anchor elements and the trusted anchor elements are the repaired target elements. It realizes the repair of target elements whose positions and contents have changed without vector repair, and has a wide repair coverage.
Claims
1. An anchor-based automatic repair method for RPA element anchors, characterized in that: Candidate elements are obtained through object detection and image feature matching. First, the combination vector is repaired for the candidate elements. If the combination vector repair fails, the detachment vector is repaired instead. The process of repairing the combination vector first involves obtaining the coordinates of the original target element, the original anchor element, and the current anchor element. Then, the scaling ratio is calculated based on the coordinates of the original anchor element and the current anchor element. The vector from the current anchor element to the original target element is scaled according to the scaling ratio to restore the area of the original target element. The similarity between the area of the original target element and the area of the candidate element is calculated, and the candidate element with the highest similarity reaching the set threshold is taken as the repaired target element. The process of repairing the detachment vector is to perform rule-based repair or model-based repair on the candidate element based on the trusted anchor element. The rule-based repair calculates scores based on the distance and angle between the candidate element and the trusted anchor element, and the candidate element that meets the specified score threshold is regarded as the repaired target element. The model-based repair uses a multi-modal model to find a pending anchor element for each candidate element. If the pending anchor element is the same as the trusted anchor element, the candidate element corresponding to the pending anchor element is the repaired target element. The candidate elements include input boxes, dropdown boxes, checkboxes, icons, and text elements. The rule-based repair is applied to input boxes, dropdown boxes, and checkboxes. The model-based repair is applied to icons and text elements. The process of the multi-modal model finding the pending anchor element is carried out according to the following steps: Step S1: Obtain the coordinates, text information, and area screenshots of the candidate element and the corresponding candidate anchor element. Step S2: Based on the area screenshot, obtain the element categories of the candidate element and the candidate anchor element through an object detection model. Step S3: Input the coordinates, text information, area screenshot, and element category of the candidate element and the candidate anchor element into the multi-modal model, and use the multi-modal model to detect and determine the pending anchor element. The detection process of the multi-modal model is carried out according to the following steps: Step S3.1: Form a text input by combining the text information of the candidate element, the text information of the candidate anchor element, the category of the candidate element, and the category of the candidate anchor element. Use the area screenshots of the candidate element and the candidate anchor element as the image input. Extract coordinate features from the coordinates of the candidate element and the candidate anchor element as the coordinate input. Step S3.2: Use an embedding transformation model to transform the text input and the image input into text vectors and image vectors, and then align the text vectors, image vectors, and coordinate feature vectors through a fully connected layer. Step S3.3: Distinguish the candidate element and the candidate anchor element through rotational position encoding. Then, through the self-attention module, the text vectors, image vectors, and coordinate feature vectors of the candidate element are made to mutually attend to the text vectors, image vectors, and coordinate feature vectors between the candidate element and the candidate anchor factor. The weights of each element are obtained through similarity, and the vectors are weighted and summed. Step S3.4: Output the candidate anchor element with the highest sum value as the pending anchor element through a fully connected layer.
2. The method for automatically repairing RPA element anchors based on anchors according to claim 1, characterized in that: For the input box and the drop-down box, the rule repair ranks the candidate elements according to the relative distance and the absolute distance from the trusted anchor element to obtain a comprehensive distance score, and obtains an angle score by calculating the angle between the candidate element and the trusted anchor element; the candidate elements that satisfy the specified score threshold for both the comprehensive distance score and the angle score are regarded as the repaired target elements. For the checkbox, the rule repair obtains the absolute distance score of the candidate element according to the absolute distance from the trusted anchor element, and the candidate elements that satisfy the specified score threshold for the absolute distance score are regarded as the repaired target elements.
3. The method for automatically repairing RPA element anchors based on anchors according to claim 1, characterized in that: The comprehensive distance score is calculated as shown in the following formula: Comprehensive distance score = relative distance ranking score × absolute distance score; Among them, the relative distance ranking score = 1 - (relative distance ranking of the candidate element and the trusted anchor element / 10); Absolute distance score = min(1, absolute distance between the original target element and the trusted anchor element / absolute distance between the candidate element and the trusted anchor element); The angle score is calculated as shown in the following formula: Angle score = 1 - ((vector angle between the candidate element and the trusted anchor element - minimum angle of the due south and due east of the candidate element) / 90).
4. The method for automatically repairing RPA element anchors based on anchors according to claim 1, characterized in that: The image feature matching is carried out according to the following steps: Step A1: Obtain the current interface image and the screenshot image of the original target element, and perform preprocessing of grayscale conversion and denoising on the two images; Step A2: Detect the points that meet the high recognition and high stability in the two images by the scale-invariant feature transform method to obtain feature points; Step A3: Extract the feature descriptors of the feature points through the scale-invariant feature transform descriptor, and the feature descriptor is used to describe the regional features around the feature points; Step A4: Compare the feature descriptors of different feature points in the two images by the K-nearest neighbor algorithm to obtain multiple similar feature points for each feature point in the interface image; Step A5: Calculate the distance ratio between the feature descriptor of each feature point and the feature descriptor of the corresponding similar feature point, compare the distance ratio with the set threshold, and the similar feature points that exceed the threshold are incorrect matches, and the rest are correct matches and are used for the next step; Step A6: Apply the geometric relationship between the correctly matched feature points and the similar feature points to the interface image by solving the transformation matrix to obtain multiple similar regions, and regard the similar regions as candidate elements.
5. The method for automatically repairing RPA element anchors based on anchors according to claim 1, characterized in that: The scaling ratio is the ratio of the short side length of the current anchor element to the short side length of the original anchor element; the scaling is calculated as shown in the following formula: Current target element coordinates = current anchor element coordinates + scaling ratio × (original target element coordinates - original anchor element coordinates).
6. The method for automatically repairing RPA element anchors based on anchors according to claim 1, characterized in that: The calculation of the similarity is to first calculate the region overlap degree and the edge alignment degree with the region of the original target element and the region of the candidate element, and then obtain the similarity through the product of the region overlap degree and the edge alignment degree; The region overlap degree is obtained by calculating the intersection-over-union ratio of the regions of the candidate element and the original target element; the intersection-over-union ratio is the ratio of the area of the intersection of the regions of the candidate element and the original target element to the area of the union of the regions of the candidate element and the original target element; The calculation of the edge alignment degree is shown in the following formula: Edge alignment degree = 1 - min(minimum margin on both sides of the left side of the candidate element / length of the original target element area, minimum margin on both sides of the right side of the candidate element / length of the original target element area, minimum margin on both sides of the upper side of the candidate element / width of the original target element area, minimum margin on both sides of the lower side of the candidate element / width of the original target element area).
Citation Information
Patent Citations
RPA process element path intelligent repairing method and system
CN116630990A
RPA element anchor point automatic searching method based on multi-modal model
CN119669600A