Picture identification-based interface element searching and positioning method
By integrating the capabilities of automated user interface operation technology with large language models, identifying and numbering interface elements, and combining natural language operation instructions for comprehensive analysis, the RPA system's accurate positioning problem when processing graphic icons and dynamic interfaces is solved, and the accuracy and reliability of the system are significantly improved.
Patent Information
- Application Number
- CN202510072417.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-06
AI Technical Summary
When existing RPA systems based on natural language interactions process graphic icons and dynamic interfaces, it is difficult to accurately find the target operation elements and their coordinate positions, resulting in a high error rate during execution.
By integrating the capabilities of automated user interface operation technology with the large language model, interface elements are identified and unique numbers are given, and comprehensive analysis is carried out in combination with natural language operation instructions to accurately obtain coordinate information of target operation elements and perform corresponding operations.
It significantly improves the accuracy and reliability of the RPA system based on natural language interaction, improves the user experience, and can quickly and accurately find the target operation elements and their coordinate positions.
Smart Images

Figure CN119942553A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of Robotic Process Automation (RPA), and in particular to a method for searching and locating interface elements based on image recognition, which is used to quickly and accurately find target operation elements and their accurate coordinate positions. Background Art
[0002] In an RPA system based on natural language interaction, each operation needs to find the target operation element and accurately locate it. Currently, two methods, plain text mode and image recognition mode, are commonly used to achieve this goal. Plain text mode describes each interface element through an XML-like data structure, including the element type and coordinate position. However, this method has significant limitations when dealing with various icons without text descriptions. For example, many modern user interfaces contain a large number of graphic icons that convey information through graphic symbols, but plain text descriptions often cannot accurately express the meaning and function of these icons. This results in the large language model not being able to understand the interface content well, and thus being unable to accurately find the target operation element, resulting in a high error rate during execution. In addition, when dealing with dynamic interfaces or complex layouts, the plain text mode is also difficult to update in real time and accurately reflect interface changes, further affecting the reliability of the system and user experience. Image recognition mode requires the system to receive screenshots and analyze their content. Although this method has certain advantages in processing graphic icons, the current large language model's image recognition capabilities cannot determine the exact coordinate position of the target operation element. Specifically, when recognizing screenshots, the large language model can often only recognize the approximate location of the icon, but cannot accurately identify the specific coordinate points. As a result, when performing an operation, the system may not be able to accurately click or operate the target element, resulting in a high error rate during the execution process.
[0003] In order to solve the above problems, the present invention proposes a new method for searching and locating interface elements based on image recognition, which can quickly and accurately find the target operation elements and their coordinate positions, significantly improve the accuracy and reliability of the RPA system based on natural language interaction, and enhance the user experience. Summary of the invention
[0004] The present invention provides a method for locating interface elements based on image recognition. The method achieves accurate recognition and positioning of interface elements by integrating the advanced capabilities of automated user interface operation technology and large language models, significantly improving the accuracy and reliability of the system.
[0005] The specific steps include: Figure 1 ): Identify all operable elements, identify all operable elements on the interface, such as buttons, input boxes, links, etc.
[0006] Uniquely number each actionable element and add the number next to the element for easy identification in subsequent steps.
[0007] The screen image containing the number is captured and passed to the large language model, along with the natural language operation instructions entered by the user.
[0008] Return the operation instructions, use the large language model to comprehensively analyze the captured screen image and natural language operation instructions, and return operation instructions such as "Click A" (click element A).
[0009] Get the coordinate information of the target operation element, find the corresponding target operation element according to the returned target operation element number, read the UI attributes of the element, and then accurately obtain its coordinate position on the screen.
[0010] Move the cursor to perform corresponding operations on the target element, such as moving the cursor to a specified position, simulating a mouse click on the target element, entering text content in a specified input box, etc., thereby achieving efficient and accurate operation of the user interface. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 : Overall flow chart.
[0012] Figure 2 : Original interface picture.
[0013] Figure 3 : Give each operable element a unique number. DETAILED DESCRIPTION
[0014] The present invention provides an interface element positioning method based on image recognition. By integrating the advanced capabilities of automated user interface operation technology and large language models, it realizes the precise recognition and positioning of interface elements, and significantly improves the accuracy and reliability of the RPA system based on natural language interaction.
[0015] The specific implementation steps are as follows: Identify all operable elements, and use existing automated user interface operation technology to comprehensively scan and identify the operable elements on the interface to ensure that no elements that can be operated by users are missed, such as buttons, input boxes, links, etc.
[0016] The specific technologies are as follows: Web page: directly read and manipulate the DOM structure through JavaScript.
[0017] Windows desktop application: read and manipulate interface elements through Automation protocol.
[0018] Apple Mac applications: read and manipulate interface elements through the Apple Accessibility API protocol.
[0019] Uniquely number each actionable element. Uniquely number each identified actionable element in a circular manner from A to Z (such as: A, B, C...Z, AA, AB, AC...) to reduce the length, and add the number next to the element in a way that does not interfere with the normal visual experience, so that it can be accurately identified in subsequent steps.
[0020] Capture the screen image containing the number and pass it to the large language model. Capture the screen image containing the number. The image should clearly and completely display the interface layout and the number of each element.
[0021] At the same time, the natural language operation instructions input by the user are transmitted. The natural language operation instructions can be the desired operations described by the user in daily language, such as "click the login button" and "enter 'kidi' in the search box".
[0022] Return the operation instructions, use the powerful image recognition ability and deep understanding of natural language of the large language model to comprehensively analyze the captured screen image and natural language operation instructions, identify the target operation elements and operation methods for the next operation, and then find the corresponding number according to the target operation element, and return an operation instruction such as "Click_A" (click on element A).
[0023] "Click" indicates the operation mode, which can also be a mouse click, double-click, keyboard input, etc.; "A" is the number of the target operation element.
[0024] Obtain coordinate information, use existing automated user interface operation technology, find the corresponding target operation element according to the returned operation instructions, read the UI attributes of the element, and then accurately obtain its coordinate information on the screen, providing accurate positioning basis for subsequent operations.
[0025] The specific method of obtaining location information is as follows: Web page: Get the location attribute of the element through Javascript.
[0026] Windows desktop applications: Get the Bounding Rectangle property of an element through the UIAutomation protocol.
[0027] Apple Mac applications: Get the location property of an element through the Apple Accessibility API protocol and AppleScript.
[0028] Move the cursor to perform corresponding operations on the target element.
[0029] According to the operation instructions, the corresponding operation actions are performed, such as moving the cursor to the specified position, simulating the mouse to click the target element, entering text content in the specified input box, etc., so as to achieve efficient and accurate operation of the user interface. Example
[0030] In this embodiment, we will explain in detail how to apply the method of the present invention in a specific RPA system application scenario based on natural language interaction. Consider a typical user interface, such as a short video playback interface. Figure 2 shown.
[0031] On this interface, there are multiple icon buttons (such as like, favorite, comment, etc.). These buttons are displayed in the form of icons. We need to like this video through the RPA system based on natural language interaction.
[0032] The specific steps are: Identify all operable elements: In the short video playback interface, we first use the existing automated user interface operation technology to identify all operable elements, including text, links, like buttons (such as star icons), favorite buttons (such as heart icons), comment buttons (such as dialog box icons), etc.
[0033] Uniquely number each actionable element: Assign a unique number to each identified actionable element, for example, the Like button is numbered "BW", the Favorite button is numbered "CA", and the Comment button is numbered "BY". These numbers will be displayed in the upper left corner of the element, such as Figure 3 shown.
[0034] Capture the screen image containing the numbers and pass it to the large language model: Capture the screen image containing all the numbers and pass it to the large language model together with the user's natural language operation instructions (for example, "like the video").
[0035] Return operation instructions: After receiving the screen image containing the number, the large language model uses its image recognition ability and natural language understanding ability to identify the target operation elements and operation methods for the next operation, and returns an instruction similar to "Click_BW", where Click is a click operation and "BW" is the number of the "Like" icon.
[0036] Get coordinate information: According to the returned operation instruction, the system uses the existing automated user interface operation technology to obtain the coordinate information of the element numbered "BW". Here is a web page, and the coordinates can be obtained by obtaining the location attribute of the element through the Javascript language.
[0037] Move the cursor to perform the corresponding operation on the target element: Finally, the system moves the mouse to the specified coordinate position, accurately locates and clicks the star icon to complete the like operation. Innovation
[0038] Combine the image recognition capabilities of the large language model with the positioning and interface control capabilities of automated user interface operation technology to accurately locate and operate the target operating elements.
Claims
1. A method for searching and locating interface elements based on image recognition, characterized in that: include: Identify interface elements through automated UI manipulation techniques; Uniquely number each operable element and display it next to the element; capture a screen image containing the number and pass it to the large language model; use the large language model's image recognition and natural language understanding capabilities to identify the operation target; obtain the operation element and its coordinate information; and perform the corresponding operation.
2. The method according to claim 1, characterized in that The automated user interface manipulation technology is applicable to web pages, Windows desktop applications, and Apple Mac applications.
3. The method according to claim 1, characterized in that The operable elements include but are not limited to text, links, buttons, pictures, icons, video players, audio players, and form elements.
4. The method according to claim 1, characterized in that The unique number is numbered in a circular manner from A to Z (such as: A, B, C...Z, AA, AB, AC...), and is added next to the operable element in a style with low transparency, small font, short length, and remote but easy to identify position.
5. The method according to claim 1, characterized in that After the large language model recognizes the operation target, it returns an instruction similar to "Click_A" (where Click is the operation method and A is the number of the target operation element).
6. The method according to claim 1, characterized in that The coordinate information is obtained by reading attribute values in UI automation technology, specifically: web pages obtain it through the location attribute of the element, Windows desktop applications obtain it through the UIAutomation protocol and the BoundingRectangle attribute, and Apple Mac applications obtain it through the AppleAccessibility API protocol and the location attribute of AppleScript.
7. The method according to claim 1, characterized in that The operations include mouse movement, mouse click, mouse double click, holding down the mouse to drag, mouse scrolling up and down, keyboard input and option selection.
8. The method according to claim 1, characterized in that The method for identifying the interface elements includes reading and operating DOM elements through JavaScript, reading and operating interface elements through UIAutomation protocol, and reading and operating interface elements through AppleAccessibility API protocol and AppleScript scripting language.
9. The method according to claim 1, characterized in that The screenshot contains the user interface and element numbers and is passed to the large language model together with the natural language instructions entered by the user.
Citation Information
Cited By
Icon and text positioning method based on multi-modal large model
CN121561451A
Element positioning method and device, equipment and medium
CN121640476A
Element location methods, apparatus, equipment and media
CN121640476B