Target detection anchor frame alignment optimization and recall rate improvement method based on RPA
The interface element borders are obtained through RPA, the object detection model is trained, and the confidence and coordinate overlap are optimized, which solves the deviation and missed call problems in object detection, and improves the recall and accuracy of the detection system.
Patent Information
- Application Number
- CN202510460467.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-29
AI Technical Summary
In the existing object detection technology, there is a deviation between the prediction box and the real box, and a high confidence threshold leads to missed calls, affecting the overall performance and recall of the detection system.
The element borders in the interface screenshot are obtained through RPA, and the object detection model is trained after data annotation is performed. Confidence threshold adjustment and coordinate overlap optimization are used to recall and correct detection borders, improving recall and improving detection accuracy.
The anchor box alignment of target detection is optimized, the recall rate and detection accuracy are improved, and the automation operation accuracy of the RPA system is ensured.
Smart Images

Figure CN120388361A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information detection, and particularly to a method for optimizing the alignment degree of target detection anchor boxes and improving the recall rate based on RPA. Background Art
[0002] In the field of Robotic Process Automation (RPA), the integration of target detection technology has significantly improved the intelligence level and efficiency of automated tasks, especially in dealing with Graphical User Interfaces (GUIs). The accurate recognition of interface elements (buttons, texts, icons, input boxes, dropdown boxes, etc.) by target detection technology enables the RPA system to understand the distribution and functions of interface elements. The RPA system performs automated operations (clicking, inputting, text extraction, etc.) with different functions on different regions of the interface according to the element borders and categories provided by target detection. Target detection technology is a core component in modern automation and computer vision systems and is widely used in various scenarios from simple image classification to complex scene parsing. This technology realizes automated processing and analysis by identifying and locating specific objects in images. However, despite the significant progress of target detection technology, there are still some key technical challenges in practical applications: 1: There is a deviation between the predicted boxes and the ground truth boxes in target detection. Although target detection algorithms optimize the deviation between prediction boxes and ground truth boxes by introducing Distributed Focal Loss (DFL), this deviation cannot be completely eliminated, thus still limiting the overall performance of the detection system; 2: Target detection has missed detections due to the setting of the confidence threshold. In existing target detection systems, in order to reduce the false detection rate, a relatively high confidence threshold is often set. Although this method can reduce false detections, it also leads to the missed detection of real targets. In addition, an overly high confidence threshold also significantly reduces the recall rate of the system, affecting the comprehensiveness of detection performance. Summary of the Invention
[0003] The purpose of the present invention is to provide a method for optimizing the alignment degree of target detection anchor boxes and improving the recall rate based on RPA. The present invention uses element borders to recall and correct target detection boxes, effectively optimizing the accuracy of detection.
[0004] The technical solution of the present invention: The method for optimizing the alignment degree of target detection anchor boxes and improving the recall rate based on RPA is carried out according to the following steps:
[0005] Step S1: Use RPA software to obtain a screenshot of the application interface, obtain the element borders in the interface screenshot through element capture, then annotate the border data of the element borders, and use the border data and the interface screenshot to train a target detection model;
[0006] Step S2: Use the trained object detection model to perform object detection on the upper border of the interface screenshot to obtain the detection border and the corresponding category and confidence level; compare the confidence level of the detection border with the set confidence level threshold, and retain the detection borders with a confidence level greater than the confidence level threshold.
[0007] Step S3: Locate the coordinates of the detection borders not retained in Step S2 and determine whether there are element borders within the detection borders. If there are no element borders, the confidence level threshold remains unchanged. If there are element borders, reduce the confidence level threshold within the located area and recall the detection borders in this located area.
[0008] Step S4: Obtain the coordinates of the detection borders retained in Step S2, the coordinates of the detection borders recalled in Step S3, and the coordinates of the element borders in Step S1. Calculate the coordinate coincidence degree between the retained and recalled detection borders and the element borders, and compare the coordinate coincidence degree with the set coordinate coincidence degree threshold. If it is less than the coincidence degree threshold, retain the original detection border. If it is greater than the coincidence degree threshold, use the element border to replace the corresponding detection border to complete the correction of the detection border.
[0009] In the above method for optimizing the object detection anchor box alignment degree and improving the recall rate based on RPA, the element capture process in Step S1 is carried out according to the following steps:
[0010] Step S1.1.1: Use the RPA software to continuously monitor the mouse position and call the automated element interface. The automated element interface obtains the root of the UI elements of the application window, traverses the UI element tree, and the smallest element at the mouse position is the target element.
[0011] Step S1.1.2: Obtain the upper left corner coordinates, width, and height in the target element information as the border of the element selected by the mouse, that is, the element border.
[0012] Step S1.1.3: Move the mouse position evenly and repeat Step S1.1.1 and Step S1.1.2. After traversing the application interface, obtain all element borders.
[0013] In the above method for optimizing the object detection anchor box alignment degree and improving the recall rate based on RPA, in Step S1.1.3, the mouse is moved every 5 pixels to traverse the application interface.
[0014] In the above method for optimizing the object detection anchor box alignment degree and improving the recall rate based on RPA, the data annotation process in Step S1 is carried out according to the following steps:
[0015] Step S1.2.1: Import the interface screenshot and the element border data into the annotation tool.
[0016] Step S2.2.2: Annotate the category according to the style of the element boxed in the interface screenshot based on the element border.
[0017] Step S2.2.3: Export the element border data marked with categories into the input format required by the object detection model for training.
[0018] In the aforementioned method for optimizing the alignment degree of object detection anchor boxes and improving the recall rate based on RPA, the annotation tool is Label Studio.
[0019] In the aforementioned method for optimizing the alignment degree of object detection anchor boxes and improving the recall rate based on RPA, the object detection model in step S2 is the YOLOv8 model; the YOLOv8 model includes an input layer, a feature extraction module, a feature fusion module, and an output module; the input layer receives the interface screenshot and processes it to the required size; the feature extraction module extracts the features in the interface screenshot; the feature fusion module fuses the features of different scales; the output module predicts the detection border and the corresponding category and confidence level for the fused features.
[0020] In the aforementioned method for optimizing the alignment degree of object detection anchor boxes and improving the recall rate based on RPA, step S4 calculates the coordinate coincidence degree through the intersection over union. The specific process is to represent the upper left coordinate and the lower right coordinate of the element border as Represent the upper left coordinate and the lower right coordinate of the detection border as Calculate the coincidence degree through the following formula:
[0021]
[0022] Area of Union = Area ele + Area pre - Area of Intersect;
[0023]
[0024] In the formula, is the abscissa of the upper left corner of the element border, is the ordinate of the upper left corner of the element border, is the abscissa of the lower right corner of the element border, is the ordinate of the lower right corner of the element border; is the abscissa of the upper left corner of the detection border, is the ordinate of the upper left corner of the detection border, is the abscissa of the lower right corner of the detection border, is the ordinate of the lower right corner of the detection border; Area ofIntersect is the area of the intersection region between the element border and the detection border; Area ele is the area of the element border; Area preTo detect the area of the bounding box; Area of Union is the area of the union of the element bounding box and the detected bounding box; IOU is the intersection over union.
[0025] In the above-mentioned method for optimizing the alignment degree and improving the recall rate of target detection anchor boxes based on RPA, in step S4, the way to replace the corresponding detected bounding box with the element bounding box is to assign the coordinates of the element bounding box to the detected bounding box.
[0026] Compared with the prior art, the present invention first obtains the element bounding boxes in the interface screenshot through RPA, then performs data annotation and imports them into the target detection model along with the interface screenshot for training. The trained target detection model is used for target detection of the upper bounding boxes on the interface screenshot. It detects whether there are element bounding boxes within the range of the detected bounding boxes with low confidence, and recalls the low-confidence detected bounding boxes with element bounding boxes by reducing the confidence. Utilizing the known information of the existence of element bounding boxes, it improves the recall rate of target detection; it detects the detected bounding boxes through the coordinate coincidence degree, and corrects the offset detected bounding boxes by replacing them with element bounding boxes with similar coincidence degrees, avoiding the situation where the RPA instruction fails to act on the correct area due to the offset of the detected bounding box or only selecting part of the area of the element. Brief Description of the Drawings
[0027] Figure 1 is the flow chart of the present invention;
[0028] Figure 2 is the cross-sectional screenshot of the example of the present invention;
[0029] Figure 3 is the visualization diagram of the element bounding boxes of the example of the present invention;
[0030] Figure 4 is the visualization diagram of the original output of the target detection model of the example of the present invention;
[0031] Figure 5 is the visualization diagram of the target detection of the example of the present invention;
[0032] Figure 6 is the visualization diagram after recall and correction of the example of the present invention. Detailed Embodiments
[0033] The following further illustrates the present invention in conjunction with the drawings and embodiments, but it shall not be used as the basis for limiting the present invention.
[0034] Embodiment: A method for optimizing the alignment degree and improving the recall rate of target detection anchor boxes based on RPA, as shown in the appendix Figure 1 is carried out according to the following steps:
[0035] Step S1: As shown in the appendix Figure 2As shown, use RPA software to obtain screenshots of the application interface, capture the element borders in the screenshots through element capture, then annotate the border data of the element borders, and use the border data and the screenshots to train the YOLOv8 object detection model;
[0036] The element capture process is carried out according to the following steps:
[0037] Step S1.1.1: Enter the RPA element capture state. RPA continuously monitors the mouse position and calls the UI automatic automation element interface. The automation element interface obtains the root of the UI elements of the application window, traverses the UI element tree, and the smallest element at the mouse position is the target element;
[0038] Step S1.1.2: Obtain the upper left coordinate, width, and height in the target element information as the border of the element selected by the mouse, that is, the element border;
[0039] Step S1.1.3: Move the mouse position every 5 pixels and repeat Step S1.1.1 and Step S1.1.2. After traversing the application interface, all element borders are obtained. After the instance is captured by the element, as shown in the appendix Figure 3 as shown.
[0040] The data annotation process is carried out according to the following steps:
[0041] Step S1.2.1: Import the interface screenshot and the element border data into the Label Studio annotation tool and ensure that they correspond one by one;
[0042] Step S1.2.2: Manually annotate the category according to the style of the element framed in the interface screenshot by the element border. The categories include Text, Icon, Image, Inputbox, and Button.
[0043] Step S1.2.3: Export the element border data marked with categories into the YOLO format required by the object detection model for training.
[0044] The target detection model includes an input layer, a feature extraction module Backbone, a feature fusion module Neck, and an output module Head; the input layer receives an image of size 640×640 as input; the structure of the feature extraction module Backbone is Conv(k3,s2)+Conv(k3,s2)+C2f(n = 3)+Conv(k3,s2)+C2f(n = 6)+Conv(k3,s2)+C2f(n = 6)+Conv(k3,s2)+C2f(n = 3), and multiple Conv and C2f layers cooperate to extract image features; the feature fusion module Neck adopts the PANet structure to fuse feature maps of different scales; the output module Head adopts a dual-head structure, which is responsible for outputting the bounding box position and the bounding box category respectively; the V100 graphics card is selected as the training environment, and the screenshot of the element border interface with labeled categories is used as the training data, and the training parameters are: image size: 640×640, batch size: 16, number of training epochs: 50.
[0045] Step S2: Use the trained target detection model to perform target detection on the bounding boxes in the interface screenshot, and obtain the detected bounding boxes and the corresponding categories and confidences; compare the confidence of the detected bounding boxes with the set confidence threshold, and retain the detected bounding boxes with a confidence greater than the confidence threshold. After object detection, the instance is as shown in the appendix Figure 4 shown, and after filtering by the confidence threshold as shown in the appendix Figure 5 shown, the retained detected bounding boxes do not maintain an alignment relationship with the element bounding boxes, and some of the detected bounding boxes corresponding to the element bounding boxes are not recalled.
[0046] The process of performing target detection on the bounding boxes by the target detection model is carried out according to the following steps:
[0047] Step S2.1.1: Preprocessing, the input interface screenshot is scaled to the size required by the target detection model.
[0048] Step S2.1.2: Feature extraction, the interface screenshot extracts important visual features layer by layer through the feature extraction module.
[0049] Step S2.1.3: Feature fusion, fuse feature maps of different scales through the feature fusion module to enhance the model's detection ability for objects of different sizes.
[0050] Step S2.1.4: Raw output, the output module generates predictions for the feature maps of each scale at each position, including the detected bounding box coordinates and the corresponding categories and confidences.
[0051] Step S2.1.5: Post-processing, filter the bounding boxes through the confidence threshold;
[0052] Step S2.1.6: Output the detected bounding boxes with high confidence and the corresponding categories and confidences.
[0053] Step S3: Locate the coordinates of the detected bounding boxes not retained in Step S2 and determine whether there are element bounding boxes within the detected bounding boxes. If there are no element bounding boxes, the confidence threshold remains unchanged at 0.25. If there are element bounding boxes, the confidence threshold within the located area is reduced to 0.001, and the detected bounding boxes within the located area are recalled.
[0054] Step S4: Obtain the coordinates of the detected bounding boxes retained in Step S2, the coordinates of the detected bounding boxes recalled in Step S3, and the coordinates of the element bounding boxes in Step S1. Calculate the coordinate coincidence degree between the retained and recalled detected bounding boxes and the element bounding boxes. Compare the coordinate coincidence degree with the set coordinate coincidence threshold. If it is less than the coincidence threshold, retain the original detected bounding box. If it is greater than the coincidence threshold, use the element bounding box to replace the corresponding detected bounding box to complete the correction of the detected bounding box;
[0055] The coincidence degree is calculated through the intersection over union. The specific process is to represent the upper left coordinate and the lower right coordinate of the element bounding box as Represent the upper left coordinate and the lower right coordinate of the detected bounding box as The coincidence degree is calculated by the following formula:
[0056]
[0057]
[0058]
[0059]
[0060] Area of Union = Area ele + Area pre - Area of Intersect;
[0061]
[0062] In the formula, is the abscissa of the upper left corner of the element bounding box, is the ordinate of the upper left corner of the element bounding box, is the abscissa of the lower right corner of the element bounding box, is the ordinate of the lower right corner of the element bounding box; is the abscissa of the upper left corner of the detected bounding box, is the ordinate of the upper left corner of the detected bounding box, is the abscissa of the lower right corner of the detected bounding box, is the ordinate of the lower right corner of the detected bounding box; Area of Intersect is the area of the intersection region between the element bounding box and the detected bounding box; Areaele is the area of the element border; Area pre is the area of the detection border; Area of Union is the area of the union region of the element border and the detection border; IOU is the intersection over union ratio.
[0063] If the IOU between the element border and the detection border > 0.9, then assign the coordinates of the element border to the prediction box. After the instance is rectified and recalled, as shown in the appendix Figure 6 The prediction box fits closely with the element.
[0064] In summary, the present invention first obtains the element border in the interface screenshot through RPA, then performs data annotation and imports it into the object detection model along with the interface screenshot for training. The trained object detection model is used for object detection of the border on the interface screenshot. It detects whether there is an element border within the range of the detection border with low confidence, and recalls the low-confidence detection border with an element border by reducing the confidence. By using the known information of the existence of the element border, the recall rate of object detection is improved; the detection border is detected through the coordinate coincidence degree, and the offset detection border is rectified by replacing it with an element border with a close coincidence degree, avoiding the situation where the RPA instruction fails to act on the correct area due to the offset of the detection border or only selecting a partial area of the element.
Claims
1. Method for optimizing alignment degree of target detection anchor boxes and improving recall rate based on RPA, characterized in that: The steps are as follows: Step S1: Use RPA software to obtain a screenshot of the application interface, capture the element borders in the screenshot through element capture, then label the border data of the element borders, and train a target detection model using the border data and the screenshot. Step S2: Use the trained target detection model to perform target detection on the borders in the screenshot to obtain the detected borders and the corresponding categories and confidences; compare the confidence of the detected borders with the set confidence threshold, and retain the detected borders with a confidence greater than the confidence threshold. Step S3: Locate the coordinates of the detected borders not retained in Step S2 and determine whether there are element borders within the detected borders. If there are no element borders, the confidence threshold remains unchanged. If there are element borders, reduce the confidence threshold within the located area and recall the detected borders in this located area. Step S4: Obtain the coordinates of the detected borders retained in Step S2, the coordinates of the detected borders recalled in Step S3, and the coordinates of the element borders in Step S1, calculate the coordinate coincidence degree between the retained and recalled detected borders and the element borders, compare the coordinate coincidence degree with the set coordinate coincidence threshold. If it is less than the coincidence threshold, retain the original detected borders. If it is greater than the coincidence threshold, use the element borders to replace the corresponding detected borders to complete the correction of the detected borders.
2. The method for optimizing the alignment degree and improving the recall rate of object detection anchor boxes based on RPA according to claim 1, wherein: The element capture process in Step S1 is carried out according to the following steps: Step S1.1.1: Use RPA software to continuously monitor the mouse position and call the automated element interface. The automated element interface obtains the root of the UI elements of the application window, traverses the UI element tree, and the smallest element at the mouse position is the target element. Step S1.1.2: Obtain the upper left corner coordinates, width, and height in the target element information as the border of the element selected by the mouse, that is, the element border. Step S1.1.3: Move the mouse position evenly and repeat Step S1.1.1 and Step S1.1.
2. After traversing the application interface, obtain all element borders.
3. The method for optimizing the alignment degree and improving the recall rate of target detection anchor boxes based on RPA according to claim 2, wherein: In Step S1.1.3, the mouse is moved every 5 pixels to traverse the application interface.
4. The method for optimizing the alignment degree and improving the recall rate of target detection anchor boxes based on RPA according to claim 1, wherein: The data annotation process in Step S1 is carried out according to the following steps: Step S1.2.1: Import the screenshot and the element border data into the annotation tool. Step S2.2.2: Label the category according to the style of the element framed by the element border in the screenshot. Step S2.2.3: Export the element border data with the labeled category into the input format required by the target detection model for training.
5. The method for optimizing the alignment degree and improving the recall rate of target detection anchor boxes based on RPA according to claim 4, characterized in that: The annotation tool is Label Studio.
6. The method for optimizing the alignment degree and improving the recall rate of the target detection anchor box based on RPA according to claim 1, characterized in that: The target detection model in Step S1 is the YOLOv8 model; the YOLOv8 model includes an input layer, a feature extraction module, a feature fusion module, and an output module; the input layer receives the screenshot and processes it to the required size; the feature extraction module extracts the features in the screenshot; the feature fusion module fuses the features of different scales; the output module predicts the detected borders and the corresponding categories and confidences for the fused features.
7. The method for optimizing the alignment degree and improving the recall rate of target detection anchor boxes based on RPA according to claim 1, wherein: The step S4 obtains the coordinate coincidence degree through the intersection over union calculation. The specific process is to represent the upper left coordinate and the lower right coordinate of the element border as represent the upper left coordinate and the lower right coordinate of the detection border as The coincidence degree is calculated by the following formula: Area of Union=Area ele +Area pre -Area of Intersect; In the formula, is the abscissa of the upper left corner of the element border, is the ordinate of the upper left corner of the element border, is the abscissa of the lower right corner of the element border, is the ordinate of the lower right corner of the element border; is the abscissa of the upper left corner of the detection border, is the ordinate of the upper left corner of the detection border, is the abscissa of the lower right corner of the detection border, is the ordinate of the lower right corner of the detection border; Area ofIntersect is the area of the intersection region between the element border and the detection border; Area ele is the area of the element border; Area pre is the area of the detection border; Area of Union is the area of the union region between the element border and the detection border; IOU is the intersection over union.
8. The method for optimizing the alignment degree of target detection anchor boxes and improving the recall rate based on RPA according to claim 7, wherein: In Step S4, the way to replace the corresponding detected border with the element border is to assign the coordinates of the element border to the detected border.