Graphical user interface positioning method and system based on dual-system thinking

By adopting an adaptive system switching mechanism and a dual-system thinking positioning method based on a multimodal large model, the problem of insufficient accuracy of existing GUI positioning methods in complex interfaces is solved, achieving a balance between high efficiency and high accuracy, especially when handling icon/widget positioning tasks.

CN120928977APending Publication Date: 2025-11-11ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510988663.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing GUI positioning methods lack accuracy when dealing with complex interfaces, cannot dynamically adjust processing strategies, and lack systematic analysis of interface layout and visual relationships, resulting in a tradeoff between efficiency and accuracy.

Method used

A graphical user interface localization method based on dual-system thinking is adopted. Through a multimodal large model adaptive system switching mechanism, a fast system or a slow system is dynamically selected for the localization task. The fast system directly predicts the coordinates, while the slow system performs interface summary analysis and visual focus analysis. The results of interface summary and visual focus are combined to make accurate coordinate prediction.

Benefits of technology

It achieves a balance between high efficiency and high accuracy in complex interfaces, with short average processing time and significantly improved accuracy, especially when handling icon/widget positioning tasks, improving accuracy by 1.4% to 17.7% compared to existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120928977A_ABST
    Figure CN120928977A_ABST
Patent Text Reader

Abstract

The invention discloses a graphical user interface positioning method and system based on dual-system thinking. The method comprises the steps that firstly, a GUI positioning task instruction which is input by a user and describes a target element needing to be positioned is obtained, meanwhile, a GUI screenshot needing to be positioned currently is obtained, and the GUI positioning task instruction and the GUI screenshot are input into a multi-modal large model trained through a GUI positioning task; and then a self-adaptive system switching mechanism is adopted, and a high-speed system or a low-speed system is dynamically determined to execute a positioning task according to task complexity. According to the method, the high efficiency and the high accuracy of GUI positioning are ensured by simulating a human dual-system cognition process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of human-computer interaction, specifically relating to a graphical user interface (GUI) positioning method and system based on dual-system thinking. Background Technology

[0002] With the rapid development of computer and internet technologies, graphical user interfaces (GUIs) have become the primary mode of human-computer interaction. To automate the operation of computer interfaces, GUI agent technology has emerged, enabling the automation of complex interface operation tasks. A core capability of GUI agents is GUI grounding: accurately locating and understanding elements within the interface based on natural language instructions. This capability directly impacts the agent's understanding of interface semantics and the effectiveness of its operations.

[0003] Traditional GUI location methods mainly rely on structured information, such as XML or DOM trees, to locate interface elements by parsing this structural information. Although this method can provide reliable structural information, it faces information redundancy and accessibility issues in practical applications. In recent years, with the development of Multimodal Large Language Models (MLLMs), screenshot-based GUI location methods have gradually emerged, such as SeeClick and ShowUI, which directly identify and locate target elements from screenshots. However, current MLLMs-based GUI location methods mainly adopt a direct prediction paradigm, attempting to locate target elements directly from screenshots without in-depth analysis. Although this method is efficient, it has two key limitations: (1) Limited interface understanding: These methods are not effective when dealing with complex interfaces containing multiple windows, nested menus, and hierarchical structures, lacking a systematic analysis of interface layout and element relationships; (2) Insufficient visual analysis: Accurate element recognition requires understanding various visual attributes (such as color, shape, and position) and their contextual relationships, while current methods attempt to directly predict target positions, similar to a human's rapid system, lacking the in-depth analytical thinking ability of humans.

[0004] This difference between direct localization and in-depth analysis reflects a fundamental characteristic of human cognitive task processing. Humans use two distinct cognitive systems when interpreting visual interfaces: a rapid, intuitive system for simple tasks and a deliberate, analytical system for complex scenarios. Current technologies cannot switch between these two modes of thinking as flexibly as humans, limiting their effectiveness in complex interface scenarios. Summary of the Invention

[0005] Existing GUI localization methods suffer from drawbacks such as insufficient accuracy due to rapid direct prediction, or inefficiency due to deep analysis and an inability to dynamically adjust processing strategies based on task complexity. Furthermore, existing methods have limited understanding of complex interface layouts and lack systematic interface analysis and visual relationship processing capabilities, resulting in low accuracy in complex GUI scenarios. Moreover, existing methods neglect the balance between rapid localization and in-depth analysis, failing to simultaneously achieve both efficiency and accuracy. To address these issues, this invention provides a graphical user interface localization method and system based on a dual-system approach.

[0006] The specific technical solution adopted in this invention is as follows:

[0007] In a first aspect, the present invention provides a graphical user interface positioning method based on dual-system thinking, comprising:

[0008] S1. Obtain the GUI positioning task instruction for the target element to be located as input by the user, and at the same time obtain the GUI screenshot that needs to be located. Input both into the multimodal large model trained by the GUI positioning task.

[0009] S2. An adaptive system switching mechanism is adopted to dynamically determine whether to use a fast or slow system to perform the positioning task based on the task complexity, wherein:

[0010] If the fast system is activated, the multimodal large model directly predicts the coordinates of the target element in the GUI screenshot based on the input GUI positioning task instructions;

[0011] If the slow system is activated, the multimodal large model first performs a task-oriented interface summary analysis on the GUI screenshot, capturing the interface summary that describes the overall layout structure and element relationships. Then, based on the interface summary, it determines the interface area in the GUI screenshot related to the target element and performs visual focus analysis on it, checking various visual attributes and contextual relationships. Finally, it combines the interface summary and visual focus analysis results to perform accurate coordinate prediction and output the coordinates of the target element in the GUI screenshot.

[0012] As a preferred embodiment of the first aspect mentioned above, if the format of the GUI screenshot does not conform to the input format standard of the multimodal large model, it needs to be converted into a standard format that can be processed by the multimodal large model.

[0013] As a preferred embodiment of the first aspect above, the graphical user interface positioning method operates on an electronic device with a display screen, and the GUI screenshot is obtained by taking a screenshot of the currently displayed GUI interface on the display screen while the user inputs the GUI positioning task command.

[0014] As a preferred embodiment of the first aspect, when the multimodal large model adopts an adaptive system switching mechanism to determine whether to use a fast system or a slow system, it first needs to predict the activation probability of the fast system and the activation probability of the slow system based on the input GUI positioning task command and GUI screenshot. Then, the activation probabilities of the two systems are adjusted by a preset scaling factor. Finally, the activation probabilities of the two systems after adjustment are compared, and the system with the higher activation probability performs the positioning task.

[0015] As a preferred embodiment of the first aspect, the method of performing task-oriented interface summary analysis on GUI screenshots by the multimodal large model is as follows: the multimodal large model is driven by task instruction prompts to first describe the overall layout structure of the interface in the GUI screenshot, then identify and describe the main functional areas and element hierarchical relationships in the GUI screenshot, and finally integrate the descriptive text output from the two steps into an interface summary.

[0016] As a preferred embodiment of the first aspect mentioned above, the method for performing visual focus analysis on GUI screenshots using the multimodal large model is as follows: the multimodal large model is driven by task instruction prompts to first preliminarily determine the target area where potential target elements are located in the GUI screenshot based on the interface summary, and then perform fine-grained visual analysis on the target area to check various visual attributes of the target element and the contextual relationship between the target element and surrounding interactive elements, and generate feature description text containing the target area and the results of fine-grained visual analysis, which is used as the visual focus analysis result to assist in determining the target element.

[0017] As a preferred embodiment of the first aspect mentioned above, the multimodal large model is pre-trained in stages using a synthesized training dataset. The method for synthesizing the training dataset is as follows: using an existing GUI localization model, coordinate prediction is directly performed on GUI screenshot samples according to the GUI localization task instructions. If the prediction is successful, fast system training data is constructed. Otherwise, an interface summary analysis step is added and prediction is performed again. If the prediction is successful, the first stage training data of the slow system is constructed. If the prediction still fails, a visual focus analysis step is added and prediction is performed again. If the prediction is successful, the complete training data of the slow system is constructed. Different data structure labels are introduced for the three types of data to guide the model to learn different inference chains.

[0018] Secondly, the present invention provides a graphical user interface positioning system based on dual-system thinking, comprising:

[0019] The data input module is used to obtain the GUI positioning task instruction that the user inputs to describe the target element to be located, and at the same time obtain the GUI screenshot that needs to be located. Both are then input into the multimodal large model trained by the GUI positioning task.

[0020] The element localization module employs an adaptive system switching mechanism to dynamically determine whether to use a fast or slow system to perform the localization task based on task complexity.

[0021] If the fast system is activated, the multimodal large model directly predicts the coordinates of the target element in the GUI screenshot based on the input GUI positioning task instructions;

[0022] If the slow system is activated, the multimodal large model first performs a task-oriented interface summary analysis on the GUI screenshot, capturing the interface summary that describes the overall layout structure and element relationships. Then, based on the interface summary, it determines the interface area in the GUI screenshot related to the target element and performs visual focus analysis on it, checking various visual attributes and contextual relationships. Finally, it combines the interface summary and visual focus analysis results to perform accurate coordinate prediction and output the coordinates of the target element in the GUI screenshot.

[0023] Thirdly, the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, can implement the graphical user interface positioning method based on dual-system thinking as described in any of the solutions of the first aspect above.

[0024] Fourthly, the present invention provides a computer electronic device, which includes a memory and a processor;

[0025] The memory is used to store computer programs;

[0026] The processor is configured to, when executing the computer program, implement the graphical user interface positioning method based on dual-system thinking as described in any of the solutions of the first aspect above.

[0027] Compared with the prior art, the present invention has the following advantages:

[0028] 1. This invention achieves a balance between high efficiency and high accuracy in GUI localization by simulating the human dual-system cognitive process. Experiments show that this method achieves an average accuracy of 77.4% on the ScreenSpot benchmark and 13.3% on the more challenging ScreenSpot-Pro, representing improvements of 1.4% and 17.7% respectively compared to existing methods.

[0029] 2. The adaptive system switching mechanism of the present invention can dynamically determine which processing strategy to use according to the complexity of the task, maintain high efficiency in simple tasks, and ensure high accuracy in complex tasks. The average processing time is only 2.6 seconds, which is better than 5.4 seconds for a system that relies entirely on a slow system.

[0030] 3. The progressive interface understanding method of the present invention (from overall overview to visual focus analysis) enables the system to better understand complex interface layouts and element relationships, and performs particularly well in handling icon / widget positioning tasks, with a significant improvement in accuracy.

[0031] 4. This invention achieves state-of-the-art performance using only 300K training data and a 2B parameter model, making it more efficient than other methods that require larger models and more data. Attached Figure Description

[0032] Figure 1 A flowchart illustrating the steps of a graphical user page location method based on dual-system thinking;

[0033] Figure 2 A schematic diagram illustrating the training method for a graphical user page localization model based on a dual-system architecture;

[0034] Figure 3 An example diagram illustrating the workflow of an adaptive system switching mechanism;

[0035] Figure 4 A schematic diagram of the module composition of a graphical user interface positioning system based on dual-system thinking;

[0036] Figure 5 This is a schematic diagram of the structure of a computer electronic device. Detailed Implementation

[0037] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. Technical features in various embodiments of the present invention can be combined accordingly without mutual conflict.

[0038] like Figure 1 As shown, in a preferred embodiment of the present invention, a graphical user interface positioning method based on dual-system thinking is provided, which includes the following steps S1 and S2. The specific implementation of the two steps will be described in detail below.

[0039] S1. Obtain the GUI positioning task instruction for the target element to be located, which is input by the user, and at the same time obtain the GUI screenshot that needs to be located. Input both into the multimodal large model trained by the GUI positioning task.

[0040] It should be noted that the aforementioned multimodal large model can be selected according to actual needs. In the embodiments of the present invention, the multimodal large model adopts the Qwen2-VL-2B-Instruct multimodal model, but it can also be replaced with other multimodal large language models, such as InternVL, simply by retraining accordingly.

[0041] Furthermore, the initial input of the multimodal large model consists of two parts: the first part is the GUI positioning task instruction, and the second part is the GUI screenshot. The GUI positioning task instruction is a statement describing the target element to be located, such as "click the search button" or "click the play music button," and the input format can be text input or voice input, among others. Additionally, the format of the GUI screenshot must conform to the input format standard of the multimodal large model. If the format of the GUI screenshot does not conform to the input format standard of the multimodal large model, it needs to be converted to a standard format that the multimodal large model can process. The GUI screenshot can be input offline or captured online. In embodiments of the present invention, if the graphical user interface positioning method runs on an electronic device with a display screen, the GUI screenshot is obtained by taking a screenshot of the currently displayed GUI interface on the display screen while the user inputs the GUI positioning task instruction, thereby potentially enabling the user to perform voice operations on the GUI interface currently displayed on the screen in real time.

[0042] S2. An adaptive system switching mechanism is adopted to dynamically determine whether to use a fast or slow system to perform the positioning task based on the task complexity, wherein:

[0043] If the fast system is activated, the multimodal large model directly predicts the coordinates of the target element in the GUI screenshot based on the input GUI positioning task instructions;

[0044] If the slow system is activated, the multimodal large model first performs a task-oriented interface summary analysis on the GUI screenshot, capturing the interface summary that describes the overall layout structure and element relationships. Then, based on the interface summary, it determines the interface area in the GUI screenshot related to the target element and performs visual focus analysis on it, checking various visual attributes and contextual relationships. Finally, it combines the interface summary and visual focus analysis results to perform accurate coordinate prediction and output the coordinates of the target element in the GUI screenshot.

[0045] In one embodiment of the present invention, such as Figure 2As shown, when a GUI localization task command (i) and the current screenshot (s) are input into the multimodal large model, the multimodal large model can use an adaptive system switching mechanism to decide whether to use the fast or slow system. First, based on the input GUI localization task command and the GUI screenshot, the activation probabilities of the fast and slow systems are predicted. Then, the activation probabilities of the two systems are adjusted using a preset scaling factor. Finally, the adjusted activation probabilities are compared, and the system with the higher activation probability performs the localization task. If the activation probability of the fast system is greater than that of the slow system, the fast system directly predicts the coordinates of the target element. If the activation probability of the slow system is greater than that of the fast system, the slow system predicts the coordinates of the target element through a two-step analysis method: first analyzing the interface overview, then performing visual focus analysis. For ease of implementation, the final output target element coordinates use normalized coordinates (x, y) from the GUI screenshot, where x ∈ [0, 1] and y ∈ [0, 1]. Combining the normalized coordinates (x, y) and the image width and height dimensions of the GUI screenshot, the pixel coordinates of the target element in the GUI screenshot can be further obtained.

[0046] It should be noted that in this invention, the scaling factor is denoted as α, and the fast system activation probability directly predicted by the multimodal large model based on the input GUI localization task instructions and GUI screenshots is denoted as p. fast The directly predicted activation probability of a slow system is denoted as p. slow The activation probability of the fast system after scaling factor adjustment is p' fast = (1-α)×p fast The activation probability of the slow system after scaling factor adjustment is p' slow =α×p slow .

[0047] Upon receiving the input GUI positioning task command and GUI screenshot, the multimodal large model of this invention dynamically adjusts its processing strategy according to the task complexity. For simple tasks, it uses a fast positioning system for efficient processing; for complex tasks, it activates a slow positioning system. This switching mechanism allows the invention to achieve a balance between efficiency and accuracy.

[0048] The slow system activation probability in this invention decomposes the localization process into several progressively advancing stages: interface overview analysis, visual focus analysis, and precise coordinate prediction. The first step, interface overview analysis, aims to capture the overall layout structure. The second step, visual focus analysis, aims to conduct in-depth analysis of relevant interface areas and their visual features, thus laying the foundation for the third step, precise coordinate prediction. This structured decomposition allows this invention to systematically understand the interface layout and visual relationships from a global to a local perspective, thereby accurately locating target elements in complex GUI interfaces. This process of gradually refining from global understanding to precise localization echoes the cognitive process humans use to solve complex tasks, enabling this invention to exhibit human-like cognitive flexibility.

[0049] In an embodiment of the present invention, the method of performing task-oriented interface summary analysis on GUI screenshots by a multimodal large model is as follows: the multimodal large model is driven by task instruction prompts to first describe the overall layout structure of the interface in the GUI screenshot, then identify and describe the main functional areas and element hierarchical relationships in the GUI screenshot, and finally integrate the description text output by the two steps into an interface summary.

[0050] It should be noted that the aforementioned multimodal large model's task-oriented interface summary analysis of GUI screenshots is guided by prompts. Specific prompts can be constructed and optimized through a prompt engineering process. These prompts require the multimodal large model to first describe the overall layout structure of the interface, then identify the main functional areas and element hierarchical relationships within the interface, thus forming a descriptive text of the interface summary. This interface summary, combined with the original input, can then be used for subsequent visual focus analysis.

[0051] As an example of this invention, a GUI screenshot shows a music playback interface of the Apple Music software. The corresponding GUI positioning task instruction is "turn off ads". Under the above two inputs, after the multimodal large model performs task-oriented interface summary analysis on the GUI screenshot, the generated interface summary text is: "The screen displays the Apple Music interface, with a prominent red promotional banner at the top center of the page. The interface looks like a mobile view, with a simple and concise layout, mainly featuring Apple Music ads."

[0052] In an embodiment of the present invention, the method for performing visual focus analysis on GUI screenshots using a multimodal large model is as follows: the multimodal large model is driven by task instruction prompts to first preliminarily determine the target area where potential target elements in the GUI screenshot are located based on the interface summary obtained in the previous step. Then, fine-grained visual analysis is performed on the target area to check various visual attributes of the target element (including the relative position, absolute position, shape, and color of the element). At the same time, the contextual relationship between the target element and the surrounding interactive elements is analyzed. Finally, feature description text containing the target area and the results of fine-grained visual analysis is generated to assist in determining the target element as a result of visual focus analysis.

[0053] Based on the above example, after performing visual focus analysis on the GUI screenshot, the multimodal large model generated the following visual focus analysis results: "The close / close button of the Apple Music ad is located in the upper right corner of the red banner. It appears as a small "X" or close icon in a light color (possibly white or light gray), which contrasts sharply with the red background of the banner. The button is roughly located about 20-30 pixels from the top and right edge of the banner."

[0054] In this invention, when the slow system is activated, precise coordinate prediction can only be performed after obtaining the interface summary and visual focus analysis results through the inference chain. The specific implementation method is as follows:

[0055] First, combine the global context provided by the interface overview with the local details provided by the visual focus analysis;

[0056] Secondly, the normalized coordinates (x, y) of the target element in the GUI screenshot are predicted by a multimodal large language model, where the values ​​of x and y are both in the range of [0, 1].

[0057] Finally, the multimodal large model outputs the final coordinate prediction results, completing the GUI element localization.

[0058] It should be noted that the multimodal large model used in steps S1 and S2 of this invention needs to be fine-tuned and trained using a localization task dataset before being used for inference. To effectively train the multimodal large model, this invention designs specialized data structure markers to guide the model's inference process. These markers are introduced as identifiers into different data samples, enabling the model to develop the ability to quickly locate data and perform slow analysis processing through supervised learning. In the embodiments of this invention, three types of data structure markers are introduced, namely:

[0059] Fast system tags: <|grounding_start|>, <|grounding_end|>;

[0060] Interface summary analysis markers: <|summary_start|>, <|summary_end|>;

[0061] Visual focus analysis markers: <|focus_start|>, <|focus_end|>.

[0062] The fast location marker encapsulates directly predicted coordinates, with an example format: `<|grounding_start|>element coordinates<|grounding_end|>`. The interface summary marker marks the interface summary analysis in the inference chain of the slow system, with an example format: `<|summary_start|>interface summary analysis text<|summary_end|>`. The visual focus analysis marker marks the visual focus analysis results in the inference chain of the slow system, with an example format: `<|focus_start|>visual focus analysis result text<|focus_end|>`. Through these specialized markers, the model develops different cognitive processes for fast and slow thinking paths. This structured design not only ensures consistent modeling of cognitive behavior but also provides clear guidance for different inference strategies.

[0063] In embodiments of the present invention, the aforementioned multimodal large model can be pre-trained in stages using a synthetic training dataset. For example... Figure 3 The diagram illustrates the model training process in an embodiment of the present invention, which employs a progressive data synthesis and two-stage training method. The method for progressively synthesizing the training dataset is as follows:

[0064] 1) Using existing GUI positioning models (such as ShowUI), directly predict the coordinates of GUI screenshot samples according to the GUI positioning task instructions. If the prediction is successful, the fast system training data is constructed. Otherwise, an interface summary analysis step is added and the prediction is made again. If the prediction is successful after adding the interface summary analysis, the first stage training data of the slow system is constructed. If the prediction still fails after adding the interface summary analysis, a visual focus analysis step is added and the prediction is made again. If the prediction is successful after adding the visual focus analysis, the complete training data of the slow system is constructed.

[0065] Therefore, based on the aforementioned three data structure notations, an exemplary fast system training data format is as follows:

[0066] <|grounding_start|>

[0067] (0.56, 0.77)

[0068] <|grounding_end|>

[0069] Additionally, an example of a complete training data format for a slow system is as follows:

[0070] <summary_start>

[0071] The screen displays an Apple Music interface with a navigation bar at the top and a playback control bar at the bottom. The main content area displays a red and white Apple Music promotional banner.

[0072] <summary_end>

[0073] <focus_start>

[0074] At the bottom of the playback control bar, there is a "Not Playing" text label on the left, followed by a small black triangle play button.

[0075] <focus_end>

[0076] <grounding_start>

[0077] (0.56, 0.77)

[0078] <grounding_end>

[0079] The first-stage training data for the slow system is similar to the complete training data for the slow system, the difference being the missing data.<focus_start> ,<focus_end> And the results of visual focus analysis between the two.

[0080] See also Figure 3 As shown, based on the three types of data structures mentioned above, differentiated labeled data samples can be used to perform supervised fine-tuning on open-source multimodal large models (such as Qwen2-VL-2B-Instruct). The special three-type label structure guides model training, and through supervised fine-tuning, the large model can learn different processing strategies and inference chains, further improving model performance. The trained model has the ability to switch between fast and slow systems.

[0081] In embodiments of the present invention, since the initial structure markers of the training data for the fast system and the training data for the slow system are different, the initial structure marker for the training data of the fast system is <|grounding_start|>, while the initial structure marker for the training data of the slow system is <|grounding_start|>.<summary_start> Therefore, in step S2 above, when the multimodal large model uses an adaptive system switching mechanism to decide whether to use a fast or slow system, the activation probabilities of the fast and slow systems can be represented by predicting the probabilities of two different initial structure markers. In other words, if the multimodal large model directly predicts the activation probability of the fast system based on the input GUI location task instructions and GUI screenshots, let's denot it as p. fast The directly predicted activation probability of a slow system is denoted as p. slow Then p fast =p(t=t) g |s,i), p slow =p(t=t) s |s,i), where t s and t g Let represent the probabilities of generating the <|summary_start|> and <|grounding_start|> markers, respectively, where s represents the GUI screenshot and i represents the GUI location task command. Therefore, when adjusting the system activation probability by introducing a scaling factor α, the adjusted slow system activation probability p′ can be calculated using the following formula. slow =α×p(t=t s |s,i), adjusted fast system activation probability p′ fast = (1-α)×p(t=t) g |s,i).

[0082] It should be noted that the method steps shown in S1 and S2 above can essentially be implemented in the form of computer programs or software functional modules.

[0083] Therefore, based on the same inventive concept, such as Figure 4 As shown, the present invention also provides a graphical user interface positioning system based on dual-system thinking, corresponding to the graphical user interface positioning method based on dual-system thinking provided in the above embodiments, which includes:

[0084] The data input module is used to obtain the GUI positioning task instruction that the user inputs to describe the target element to be located, and at the same time obtain the GUI screenshot that needs to be located. Both are then input into the multimodal large model trained by the GUI positioning task.

[0085] The element localization module employs an adaptive system switching mechanism to dynamically determine whether to use a fast or slow system to perform the localization task based on task complexity.

[0086] If the fast system is activated, the multimodal large model directly predicts the coordinates of the target element in the GUI screenshot based on the input GUI positioning task instructions;

[0087] If the slow system is activated, the multimodal large model first performs a task-oriented interface summary analysis on the GUI screenshot, capturing the interface summary that describes the overall layout structure and element relationships. Then, based on the interface summary, it determines the interface area in the GUI screenshot related to the target element and performs visual focus analysis on it, checking various visual attributes and contextual relationships. Finally, it combines the interface summary and visual focus analysis results to perform accurate coordinate prediction and output the coordinates of the target element in the GUI screenshot.

[0088] Furthermore, based on the same inventive concept, such as Figure 5 As shown, the present invention also provides a computer electronic device corresponding to the graphical user interface positioning method based on dual-system thinking provided in the above embodiments, which includes a memory and a processor;

[0089] The memory is used to store computer programs;

[0090] The processor is used to implement the graphical user interface positioning method based on dual-system thinking as described above when executing the computer program;

[0091] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0092] Therefore, based on the same inventive concept, the present invention provides a computer-readable storage medium corresponding to the graphical user interface positioning method based on dual-system thinking. The storage medium stores a computer program, which, when executed by a processor, can realize the graphical user interface positioning method based on dual-system thinking as described above.

[0093] Therefore, based on the same inventive concept, the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, can realize the graphical user interface positioning method based on dual-system thinking as described above.

[0094] Specifically, in the computer-readable storage medium of the above three embodiments, the stored computer program is executed by a processor, which can perform the aforementioned steps S1 to S2.

[0095] It is understood that the aforementioned storage media may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Furthermore, the storage media may also be various media capable of storing program code, such as USB flash drives, external hard drives, magnetic disks, or optical discs.

[0096] It is understood that the processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0097] It should also be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the system described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. In the embodiments provided in this application, the division of steps or modules in the system and method is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple modules or steps may be combined or integrated together, and a module or step may also be split.

[0098] The present invention will further demonstrate the detailed implementation process and technical effects of the graphical user interface positioning method based on dual-system thinking shown in steps S1 to S2 on a specific dataset through a specific embodiment, so as to facilitate understanding of the essence of the present invention.

[0099] Example

[0100] The steps in this embodiment are the same as the graphical user interface positioning method based on dual-system thinking shown in steps S1 to S2 above, and will not be repeated here. The main focus is on demonstrating the specific dataset, some specific parameter settings, and implementation results of this embodiment. For ease of description, the method shown in steps S1 to S2 will be referred to as the method of this invention below.

[0101] This embodiment performs a comprehensive evaluation on two publicly available GUI positioning benchmark datasets: ScreenSpot and ScreenSpot-Pro. ScreenSpot contains 1,272 samples covering mobile, desktop, and web platforms, primarily testing common interface scenarios and element types; ScreenSpot-Pro contains high-resolution interface samples from 23 professional applications, focusing on testing complex layouts in professional software environments.

[0102] This embodiment compares the method of the present invention with several existing methods, including closed-source multimodal models (such as GPT-4V and Gemini-1.5-pro) and open-source GUI localization models (such as ShowUI and CogAgent). Following the evaluation criteria of previous studies, a correct prediction is considered when the predicted coordinates (x, y) fall within the bounding box of the target element, and the average accuracy of all test samples is used as the evaluation metric.

[0103] In the ScreenSpot benchmark, the multimodal large model of this invention, using 2B parameters, achieved an average accuracy of 77.4%, surpassing ShowUI (75.1%) and InfiGUIAgent (76.3%), which also use 2B parameters. General multimodal models such as GPT-4V and Gemini-1.5-pro performed relatively poorly (16.7% and 53.2%, respectively). In particular, the method of this invention performs exceptionally well in icon / widget localization tasks on mobile platforms, achieving an accuracy of 78.2%.

[0104] In the more challenging ScreenSpot-Pro benchmark test, the method of this invention achieved an overall accuracy of 13.3%, significantly outperforming AriaUI (11.3%) and CogAgent (7.7%), despite having fewer parameters. The method also demonstrated a significant advantage in icon / widget recognition, achieving an accuracy of 3.9%, compared to ShowUI's 2.6%. The experimental results are shown in Tables 1 and 2, indicating that the localization method of this invention provides accurate element localization.

[0105] Table 1 shows the application of the method of the present invention on the ScreenSpot dataset.

[0106]

[0107] Table 2 shows the application of the method of the present invention on the ScreenSpot-Pro dataset.

[0108]

[0109] The embodiments described above are merely some preferred implementations of the present invention and are not intended to limit the invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the invention. Therefore, all technical solutions obtained through equivalent substitution or transformation fall within the protection scope of the present invention.

Claims

1. A graphical user interface positioning method based on dual-system thinking, characterized in that, include: S1. Obtain the GUI positioning task instruction for the target element to be located as input by the user, and at the same time obtain the GUI screenshot that needs to be located. Input both into the multimodal large model trained by the GUI positioning task. S2. An adaptive system switching mechanism is adopted to dynamically determine whether to use a fast or slow system to perform the positioning task based on the task complexity, wherein: If the fast system is activated, the multimodal large model directly predicts the coordinates of the target element in the GUI screenshot based on the input GUI positioning task instructions; If the slow system is activated, the multimodal large model first performs a task-oriented interface summary analysis on the GUI screenshot, capturing the interface summary that describes the overall layout structure and element relationships. Then, based on the interface summary, it determines the interface area in the GUI screenshot related to the target element and performs visual focus analysis on it, checking various visual attributes and contextual relationships. Finally, it combines the interface summary and visual focus analysis results to perform accurate coordinate prediction and output the coordinates of the target element in the GUI screenshot.

2. The graphical user interface positioning method based on dual-system thinking as described in claim 1, characterized in that, If the format of the GUI screenshot does not conform to the input format standard of the multimodal large model, it needs to be converted into a standard format that the multimodal large model can process.

3. The graphical user interface positioning method based on dual-system thinking as described in claim 1, characterized in that, The graphical user interface positioning method operates on an electronic device with a display screen. The GUI screenshot is obtained by taking a screenshot of the currently displayed GUI interface on the display screen while the user inputs the GUI positioning task command.

4. The graphical user interface positioning method based on dual-system thinking as described in claim 1, characterized in that, When the multimodal large model adopts an adaptive system switching mechanism to decide whether to use a fast system or a slow system, it first needs to predict the activation probability of the fast system and the activation probability of the slow system based on the input GUI positioning task command and GUI screenshot. Then, it adjusts the activation probabilities of the two systems through a preset scaling factor. Finally, it compares the activation probabilities of the two systems after adjustment, and the system with the higher activation probability executes the positioning task.

5. The graphical user interface positioning method based on dual-system thinking as described in claim 1, characterized in that, The method of performing task-oriented interface summary analysis on GUI screenshots by the multimodal large model is as follows: the multimodal large model is driven by task instruction prompts to first describe the overall layout structure of the interface in the GUI screenshot, then identify and describe the main functional areas and element hierarchical relationships in the GUI screenshot, and finally integrate the description text output by the two steps into an interface summary.

6. The graphical user interface positioning method based on dual-system thinking as described in claim 1, characterized in that, The method for performing visual focus analysis on GUI screenshots using the multimodal large model is as follows: The multimodal large model is driven by task instruction prompts to first preliminarily determine the target area where potential target elements are located in the GUI screenshot based on the interface summary. Then, fine-grained visual analysis is performed on the target area to check various visual attributes of the target element and the contextual relationship between the target element and surrounding interactive elements. Feature description text containing the target area and the results of fine-grained visual analysis is generated to assist in determining the target element as a result of visual focus analysis.

7. The graphical user interface positioning method based on dual-system thinking as described in claim 1, characterized in that, The multimodal large model is pre-trained in stages using a synthetic training dataset. The method for synthesizing the training dataset is as follows: using an existing GUI localization model, coordinate prediction is directly performed on GUI screenshot samples according to the GUI localization task instructions. If the prediction is successful, fast system training data is constructed. Otherwise, an interface summary analysis step is added and prediction is performed again. If the prediction is successful, the first stage training data of the slow system is constructed. If the prediction still fails, a visual focus analysis step is added and prediction is performed again. If the prediction is successful, the complete training data of the slow system is constructed. Different data structure labels are introduced for the three types of data to guide the model to learn different inference chains.

8. A graphical user interface positioning system based on dual-system thinking, characterized in that, include: The data input module is used to obtain the GUI positioning task instruction that the user inputs to describe the target element to be located, and at the same time obtain the GUI screenshot that needs to be located. Both are then input into the multimodal large model trained by the GUI positioning task. The element localization module employs an adaptive system switching mechanism to dynamically determine whether to use a fast or slow system to perform the localization task based on task complexity. If the fast system is activated, the multimodal large model directly predicts the coordinates of the target element in the GUI screenshot based on the input GUI positioning task instructions; If the slow system is activated, the multimodal large model first performs a task-oriented interface summary analysis on the GUI screenshot, capturing the interface summary that describes the overall layout structure and element relationships. Then, based on the interface summary, it determines the interface area in the GUI screenshot related to the target element and performs visual focus analysis on it, checking various visual attributes and contextual relationships. Finally, it combines the interface summary and visual focus analysis results to perform accurate coordinate prediction and output the coordinates of the target element in the GUI screenshot.

9. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it can realize the graphical user interface positioning method based on dual-system thinking as described in any one of claims 1 to 7.

10. A computer electronic device, characterized in that, Including memory and processor; The memory is used to store computer programs; The processor is configured to implement the graphical user interface positioning method based on dual-system thinking as described in any one of claims 1 to 7 when executing the computer program.