Automatic generation method of GUI (Graphical User Interface) view operation state machine
By conducting deep learning model training and state machine generation of GUI views, the problem of insufficient understanding of the deep structure and element details of GUI views in the existing technology is solved, accurate analysis and dynamic interaction recording of the GUI interface are realized, and the GUI interaction capabilities of the agent are improved.
Patent Information
- Application Number
- CN202510478439.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-25
AI Technical Summary
When processing GUI screenshots, it is difficult to accurately understand the deep structure and element details of the GUI view, especially the logical association and functional semantics between page elements, resulting in insufficient capabilities of the agent in GUI interaction.
The computer vision algorithm based on deep learning is adopted, the GUI images are retrained using the YOLOv8 object detection model, the page element relationship matrix is constructed, and the GUI interface state machine is generated based on user operation records. The visual tool is used to correct and identify errors, so as to realize accurate analysis of the GUI interface and dynamic interactive recording.
It significantly improves the versatility and accuracy of GUI parsing, can adapt to a variety of operating systems and application interfaces, supports backtracking and reproduction of operation processes, and provides reliable data support for GUI automation testing and human-computer intelligent interaction.
Smart Images

Figure CN120375152A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of graphical user interface processing for web pages and application programs, and particularly relates to a method for automatically generating a GUI view operation state machine. Background Art
[0002] With the rapid development of artificial intelligence technology, intelligent agents based on large language models have shown great potential in many fields, and intelligent agents for instruction prediction based on GUI screenshots have become a research hotspot. Such intelligent agents aim to understand the interaction requirements proposed by users based on UI screenshots and accurately predict subsequent operation steps, providing support for achieving more intelligent and convenient human-computer interaction. However, how to enable intelligent agents to efficiently and accurately understand the rich semantic information in GUI screenshots has become a key bottleneck in their development.
[0003] In the prior art, some are based on interface metadata to carry out GUI view parsing work. However, since it is difficult to obtain the metadata of most interfaces, researchers have proposed pixel-based UI parsing. Using pixel-level data such as interface screenshots for UI parsing eliminates the dependence on interface metadata. Pixel-based UI parsing has gradually been divided into a method of an end-to-end technical route that uses a deep learning model for automated processing from input to output and reduces the dependence on domain knowledge; and a technical route of a traditional method that relies on heuristics or classical computer vision techniques designed manually and requires rich domain knowledge.
[0004] Currently, when most intelligent agents process GUI screenshots, whether through end-to-end technical methods or through heuristic computer vision technical methods, they lack the ability to understand the deep structure and element details of general GUI views, and can only perform some shallow image feature recognition, making it difficult to truly grasp the logical relationships between page elements and the functional meanings they carry. In Schwerdtfeger R S.Making the gui talk.IBM[EB / OL], Outspoken is a screen reader that supports GUIs and can describe the text and graphical elements on the screen. The system maintains a database of graphical elements with oral descriptions and matches the elements on the screen with descriptions of similar items. However, the accuracy and granularity of recognition relying solely on this matching strategy largely depend on the labeled dataset, and the logical relationships between page elements cannot be obtained. In addition, the icon and picture elements of different application programs are often different, and the graphical element database of the system cannot contain the icons of all application programs, so the elements of most application program pages still cannot be parsed.
[0005] The prior art is mainly divided into two categories: one is the parsing method based on interface metadata, but its application is limited due to the difficulty of obtaining metadata; the other is the GUI parsing method based on pixels, including the end-to-end technical route using deep learning models and the traditional computer vision technical route relying on manually designed rules. However, the existing methods generally have the problem of insufficient understanding of the deep structure and element details of the GUI view, and it is difficult to accurately grasp the logical association and functional semantics between page elements. Summary of the Invention
[0006] In order to overcome the above-mentioned deficiencies of the prior art, the purpose of the present invention is to provide an automatic generation method for the operation state machine of the GUI view, providing structured, semantic input and complete scenario information for the intelligent agent based on the large language model; significantly improving the ability of the intelligent agent to process complex GUI interaction instructions and expanding its application scenarios.
[0007] In order to achieve the above purpose, the technical solution adopted by the present invention is:
[0008] An automatic generation method for the operation state machine of the GUI view, including the following steps;
[0009] Step 1: Retrain the YOLOv8 object detection model using the constructed GUI dataset to achieve accurate recognition of the type, content, and position information of interactive components in the GUI image, and learn the relative relationships between element components to construct a page element relationship matrix;
[0010] Step 2: Combine the recognition results of the YOLOv8 object detection model for the GUI image and the page element relationship matrix to generate the logical relationships between components, and complete the construction of the GUI interface structure based on the spatial position relationship and visual hierarchy features;
[0011] This interface structure is similar to the DOM tree structure and can clearly reflect the layout structure of the GUI interface and the hierarchical relationship between nodes;
[0012] Step 3: Based on the construction of the GUI interface structure, gradually construct a complete interface state machine by recording the interface jumps brought about by the user's operations on the dynamic elements of the interface, and map the operation steps to the state jump relationship between the nodes of the page structure tree;
[0013] Compared with the traditional method, the interface state machine model not only completely records the jump relationship between the interface structure trees, but also supports the backtracking and reproduction of the operation process, providing a standardized and structured description method for the analysis and evaluation of GUI interaction behaviors;
[0014] Step 4: After generating the GUI interface structure and recording the state machine model of page jumps, to ensure the accuracy of the identified components and structures, an artificial inspection mechanism is introduced. The state machine is visually presented, and manual correction of identified errors or missing elements is allowed to ensure a high level of accuracy in the parsing results of the GUI interface structure;
[0015] Compared with traditional methods, this state machine model can not only completely record the jump relationships between interface structures but also support the backtracking and reproduction of operation processes, providing a standardized and structured description method for the analysis and evaluation of GUI interaction behaviors.
[0016] In the above-mentioned Step 1, a computer vision algorithm based on deep learning is adopted to retrain the YOLOV8 object detection model using the constructed GUI dataset. With the marked supervised dataset, components are classified into multiple GUI component types including buttons, links, input boxes, dropdown menus, and checkboxes. The characteristics of the YOLOV8 object detection network are used to predict the classification categories and spatial position coordinates of elements, and at the same time, PaddleOCR is used to recognize the text content in the element components to meet the detection requirements of the main interactive elements of the GUI interface.
[0017] Component type recognition: Accurately recognize the interactive components of the GUI page, including buttons, links, input boxes, dropdown menus, and checkboxes. The input image X with size H*W is output through the YOLOv8 object detection model to obtain the prediction results on the S*S grid, including object confidence, coordinate representation of elements, and probability values of C component types;
[0018]
[0019] Component content extraction: Extract the text labels and typed content of the components;
[0020] Component position location: Accurately describe the position of the component in the interface through the bounding box coordinates (x, y, width, height);
[0021] Element relationship matrix construction: Use supervised learning with a multi-layer perceptron MLP(·) to construct a homogeneous element relationship matrix with n dimensions of the number of elements to represent the relationships between all elements in the page, and complete the subsequent construction of the page structure based on spatial position relationships and visual hierarchy features;
[0022] For the page element set E = {e1, e2,..., e n}, define the feature vector of each element as f i ∈R d , i = 1, 2,..., n, and obtain the relationship matrix R:
[0023] R = [r ij n*n , r ij = MLP([f i , f j ).
[0024] The specific steps of step 2 are as follows:
[0025] Automatically construct a page logic tree through supervised learning of a multi-layer perceptron. By extracting the spatial coordinates, visual attributes, and semantic features of GUI elements, construct feature vectors and train a supervised learning model to predict the logical relationships between elements and generate a relationship matrix;
[0026] Use topological sorting and spatial hierarchy checking to post-process the matrix, eliminate circular dependencies, and enforce inclusion relationships. Finally, construct a page logic tree that reflects the component hierarchy and layout structure, enabling clear acquisition of the web page layout structure and the detailed content contained in each node. This method integrates multi-modal features, adopts an end-to-end learning strategy, avoids manual rule design, supports cross-platform dynamic component parsing, and the generated tree structure can be further used in scenarios such as layout analysis and automated testing. The computational efficiency and generalization ability are optimized through block processing and platform feature input.
[0027] The interface state machine model in step 3 is constructed in the following way: intercept and parse the GUI interfaces of web pages and applications, record the user's operation process, including the type of user operation (such as click event, key-in event, scroll event), operation location, input content, and the new interface with interface changes brought about by the current operation;
[0028] Construct a state machine model based on the above data.
[0029] The specific steps of step 3 are as follows:
[0030] Construct a state machine model to record the user operation sequence. By recording the interaction operations on interactive dynamic elements, map each operation step to the state transition between page structure nodes to form a traceable operation process;
[0031] The specific implementation is as follows: First, through the operation capture module of the multi-platform event listening mechanism, real-time listen to and record user operation events, including the operation component ID and operation type, and construct GUI structure tree nodes and state machine states;
[0032] Construct a two-layer state model through the state definition module, dynamically associate the GUI structure tree node state with the global state machine state, and achieve an accurate mapping from operation events to state changes through predefined state transition rules;
[0033] The persistent storage adopts a hybrid storage strategy. The operation sequence is stored in a time-series database, and the state machine model is stored in XML format to describe the transfer logic;
[0034] The backtracking and replay module realizes reversible traversal of the operation process based on the management of the historical state stack, supports millisecond-level timing accuracy replay and a visual debugging interface, combines cross-platform consistency verification to ensure the accuracy of interactive behavior reproduction, provides a standardized and structured description method for the interaction between the user and the GUI, and realizes efficient processing and response to user instructions.
[0035] The specific steps of step 4 are as follows:
[0036] Use a visualization tool (such as Graphviz, PyQt, Tkinter, etc.) to graphically present the generated GUI interface structure;
[0037] First, abstract all components in the GUI interface structure to be analyzed into nodes, represent the relationship between nodes in a tree structure or hierarchical structure, and connect each node with an edge to represent the parent-child hierarchical relationship.
[0038] Use the Graphviz tool to achieve graphical presentation of the GUI structure, including the following steps: First, automatically extract the component information of the GUI interface structure to be analyzed to form structured data including component type, identifier, attributes, and parent-child hierarchical relationship;
[0039] Then convert the above data into a DOT format file supported by Graphviz;
[0040] Then, by configuring the graphical layout and visual parameters of Graphviz, use its built-in visualization engine to generate a tree structure diagram of GUI components; finally, conduct an intuitive comparison and analysis between the graphical component structure and the actual interface to facilitate quickly discovering and locating the differences between the GUI structure and the actual interface.
[0041] The method does not depend on a specific operating system or interface framework and can handle various GUI interfaces including web pages, desktop applications, and mobile applications.
[0042] A method for automatically generating the structure of a GUI view includes the following steps:
[0043] Detect interactive components in the image, extract component feature information and the layout relationship between components to construct a structured tree model;
[0044] Record the interface jumps caused by the user's operation of interactive components to generate a state machine model;
[0045] Visually present the state machine for manual verification and correction;
[0046] Output the complete state machine model.
[0047] A device for automatically generating the interface structure of a GUI view includes:
[0048] A memory for storing a computer program;
[0049] A processor for automatically generating the operation state machine of the GUI view when executing the computer program.
[0050] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can automatically generate a GUI view operation state machine according to the prediction method.
[0051] Advantages of the present invention:
[0052] 1. The present invention adopts a technical solution combining a vision-based GUI parsing method with a structured state machine, breaking through the dependence of traditional methods on specific platforms or interface frameworks, and significantly improving the generality and applicability of GUI parsing. First, the GUI view is parsed multi-dimensionally through computer vision technology to accurately identify the type, content, and location information of components; then, a structured tree model is constructed based on the spatial relationship and visual hierarchy features between components; then, the user operation trajectory is mapped into a state machine model to record the jump relationship between structural nodes; finally, a complete record and reproduction of GUI interaction behaviors are realized on the state machine model. This method significantly improves the generality and accuracy of GUI parsing and can adapt to multiple operating systems and application program interfaces.
[0053] 2. The present invention realizes the correlation analysis between interface elements and interaction behaviors by constructing a GUI operation state machine model, and does not limit the parsing to static interface elements. This method can not only accurately record the hierarchical relationship of interface components, but also completely capture the user operation process, providing reliable data support for GUI automated testing and human-machine intelligent interaction.
[0054] 3. The present invention combines specific GUI context information for interface parsing and interaction recording. Compared with traditional methods based on code analysis or fixed template matching, it has stronger environmental adaptability and higher parsing efficiency. By introducing visual feature analysis and dynamic state tracking, it can more accurately reflect the actual usage scenario of the GUI, providing a real and reliable data basis for human-machine interaction research and intelligent agent evaluation.
[0055] 4. The complete GUI operation state machine provided by the present invention provides a standardized benchmark for evaluating the execution results of instructions of intelligent agents based on large language models. Compared with existing evaluation methods, the present invention can provide more comprehensive and structured evaluation data, which helps to improve the human-machine interaction ability and task execution accuracy of intelligent agents. Description of the Drawings
[0056] Figure 1This is the workflow diagram of the present invention.
[0057] Figure 2 These are partial GUI pages of a personnel management system.
[0058] Figure 3 For attachment Figure 2 The corresponding page structure model.
[0059] Figure 4 These are partial operation state machine models of a personnel management system. Detailed implementation manners
[0060] The present invention will be further described in detail below with reference to the accompanying drawings.
[0061] The present invention aims to solve the technical bottleneck of insufficient understanding of the deep structure and element details of GUI views in the prior art, and discloses an automatic generation method for a GUI view operation state machine. Through innovative visual parsing technology, this method realizes the accurate parsing of GUI views, provides structured and semantic input support for agents based on large language models, thereby improving the understanding ability of agents for GUI screenshots and the accuracy of instruction prediction.
[0062] The following will be described in detail the specific implementation manners of the present invention with reference to Figure 1 shown below. The method includes:
[0063] Step 1: Retrain the relatively advanced YOLOv8 object detection model using the carefully constructed GUI dataset of the present invention to accurately identify the type, content, and location information of interactive components, and learn the relative relationships between elements to construct an element relationship matrix;
[0064] In this embodiment, the adopted YOLOv8 object detection model is pre-trained on a large-scale manually annotated GUI dataset to achieve accurate identification of interactive components in GUI screenshots.
[0065] The specific parsing process includes the following key steps:
[0066] Component type recognition: Accurately identify the interactive components of the GUI page, including buttons, links, input boxes, dropdown menus, checkboxes, etc. The input image X with size H*W is output through the YOLOv8 object detection model as the prediction results on the S*S grid, including the object confidence, the coordinate representation of the element, and the probability values of C component types;
[0067]
[0068]
[0069] Component content extraction: Extract the text labels and typed content of the components;Component position positioning: The position of the component in the interface is accurately described by the bounding box coordinates (x, y, width, height);
[0070] Element relationship matrix construction: Use the multi-layer perceptron MLP(·) for supervised learning to construct a homogeneous element relationship matrix with n dimensions of the number of elements to represent the relationships between all elements in the page, so as to complete the subsequent construction of the page structure based on the spatial position relationship and visual hierarchy features. For the page element set E = {e1, e2,..., e n}, define the feature vector of each element as f i ∈R d , i = 1, 2,..., n, and obtain the relationship matrix R:
[0071] R = [r ij n*n , r ij = MLP([f i , f j )
[0072] In this embodiment, the process of obtaining the GUI interface screenshot does not depend on a specific operating system or interface framework, and can handle various GUI interfaces including Web pages, desktop applications, and mobile applications.
[0073] The method supports multiple image formats (such as PNG, JPEG, etc.), and does not depend on the resolution or image size of the target device, and is applicable to the automatic generation of the general GUI interface structure.
[0074] In the specific implementation process, the device is set as a computer, and the GUI interface screenshot is automatically captured by real-time monitoring the device screen, and the keyboard and mouse events are synchronously monitored to completely record the operation behaviors triggered by the user, laying a data foundation for the subsequent generation of the GUI interface structure and the construction of the state machine model.
[0075] Step 2: Combine the detection results of the GUI image by the target detection network and the generated page element relationship matrix to generate the logical relationship between components, and complete the construction of the page structure based on the spatial position relationship and visual hierarchy features. This interface structure is similar to the DOM tree structure and can clearly reflect the layout structure of the GUI interface and the hierarchical relationship between nodes;
[0076] On the basis of identifying the components, combine the element relationship matrix predicted based on the spatial position relationship and visual hierarchy features, and generate a complete page structure. This structure is similar to the DOM tree structure and can clearly reflect the hierarchical layout of the GUI interface and the organizational relationship between each node.
[0077] In the specific implementation process, with respect to Figure 2 Taking the page of a certain personnel management system as an example, first, use the object detection network to generate the element component types, positions, and text information of the personnel management system page, use the multi-layer perceptron to predict the element relationship matrix, and construct a structure that highlights the visual hierarchy relationship with this information, such as Figure 3 shown, so as to more intuitively display the hierarchy and logical relationship of the interface elements.
[0078] Step 3: As the interface jumps caused by the user's operations on the dynamic elements of the interface, gradually record the complete interface state machine, and map the operation steps to the state transition relationship between the nodes of the page structure tree. Compared with the traditional method, this state machine model not only completely records the jump relationship between the interface structure trees, but also supports the backtracking and reproduction of the operation process, providing a standardized and structured description method for the analysis and evaluation of GUI interaction behaviors;
[0079] In this embodiment, the state machine model will record the operation process of some pages of a certain personnel management system as shown in Figure 4 the example. The state machine model mainly includes the following core components: State: A specific state of the GUI interface, such as "main interface", "login interface", or "form filling page", etc.; Event: The interaction operations of the user with the dynamic elements in the GUI view environment, such as click, input, scroll and other events; State transition: The process of transitioning from one state to another under the action of the user event. For example, after clicking the "Login" button, the interface jumps from the login page to the main page; State history: Stores the operation path of the user, which is used to support backtracking, replay, and analysis of user behaviors.
[0080] For the state machine model, in order to achieve efficient storage and backtracking, this example uses a directed graph to represent it. The nodes are different GUI interface states, and the edges represent the state transitions caused by user operations. The weights can be used to count the number of times the operations occur in the data. This method clearly represents the conversion relationship between the interfaces, can visually analyze the user interaction path, and supports backtracking and replay. It is also possible to store user operations based on time, similar to a log. This method is more suitable for analyzing interaction habits rather than recording the complete execution logic relationship of the web page.
[0081] By combining the state machine with the GUI structure, intelligent interaction analysis can be carried out: Detect component relationships (such as buttons, input boxes) to analyze user behaviors, predict possible operation paths (such as "next step" recommendations), analyze abnormal operations (such as accidentally clicking an invalid button), etc.
[0082] Step 4: After generating the GUI structure and recording the state machine model of the page jump, in order to ensure the accuracy of the identified components and structures, introduce an artificial inspection mechanism, visually present the interface state machine in a visual way, and allow manual correction of the identified errors or missing elements, so as to ensure the high accuracy of the GUI parsing result.
[0083] In this embodiment, for the convenience of manual inspection, a visualization tool is used to present the generated GUI interface state machine in a graphical manner.
[0084] The visualization tool can display the nodes (components) and edges (hierarchical relationships) in the structure in the form of a tree diagram or a hierarchical diagram, enabling the operator to intuitively compare the actual GUI interface with the structure. The operator checks the structure through the visualization tool and focuses on the following aspects:
[0085] Component recognition accuracy: Check whether there are errors in component type recognition (for example, a button is misrecognized as an input box);
[0086] Component recognition integrity: Check whether there are missing components (for example, some interactive controls are not detected);
[0087] Structural hierarchy correctness: Check whether the hierarchical relationships between components are correct (for example, whether the parent-child relationship or sibling relationship conforms to the actual interface layout).
[0088] To support manual correction, the present invention provides a visualization editing function based on front-end technologies such as D3.js, supports modifying the structure through operations such as clicking and dragging, and provides an interactive editing function. The operator can modify the structure in the following ways:
[0089] Add missing components: Manually add undetected components to the structure by adding new components and specify their types, contents, and location information;
[0090] Modify incorrect components: Correct the incorrect component type or content by clicking on the component. For example, correct a component misrecognized as an input box to a button;
[0091] Adjust the hierarchical relationship: Adjust the hierarchical relationship between components by dragging component nodes to ensure that the structure is consistent with the actual interface layout.
[0092] Therefore, the technical method of the present invention can be widely applied but is not limited to the following scenarios: Intelligent customer service assisting users in operating software: Users upload screenshots of software interfaces, and the intelligent agent analyzes the screenshots through the present invention and generates operation instructions to guide users to complete specific tasks; Automated software testing: Generating test scripts based on GUI screenshots to automatically simulate user operations and complete software function testing; Interface reconstruction and optimization: By parsing existing GUI interfaces, generating a complete state machine model of the interface to support interface reconstruction and optimization. Through deep learning models and state machine models, accurate parsing and dynamic interaction recording of GUI interfaces are achieved;
[0093] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, in any regard, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
[0094] In addition, it should be understood that although this specification is described in terms of embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A method for automatically generating a GUI view operation state machine, characterized in that, Including the following steps; Step 1: Retrain the YOLOv8 object detection model using the constructed GUI dataset to achieve accurate recognition of the types, contents, and location information of interactive components in GUI images, and learn the relative relationships between element components to construct a page element relationship matrix; Step 2: Combine the recognition results of the YOLOv8 object detection model for GUI images and the page element relationship matrix to generate the logical relationships between components, and complete the construction of the GUI interface structure based on spatial position relationships and visual hierarchy features; Step 3: Based on the construction of the GUI interface structure, gradually construct a complete interface state machine by recording the interface jumps brought about by the user's operations on dynamic elements of the interface, and map the operation steps to the state jump relationships between the nodes of the page structure tree; Step 4: After generating the GUI interface structure and recording the state machine model of page jumps, introduce an artificial inspection mechanism, visually present the state machine in a visual manner, and allow manual correction of misrecognized or missing elements to ensure the high accuracy of the GUI interface structure parsing results.
2. The automatic generation method of a GUI view operation state machine according to claim 1, characterized in that In the above Step 1, a computer vision algorithm based on deep learning is adopted to retrain the YOLOV8 object detection model using the constructed GUI dataset. Using the labeled supervised dataset, the components are classified into multiple GUI component types including buttons, links, input boxes, dropdown menus, and checkboxes. And utilize the characteristics of the YOLOV8 object detection network to predict the classification category and spatial position coordinates of elements. At the same time, use PaddleOCR to recognize the text content in the element components to meet the detection requirements of the main interactive elements in the GUI interface.
3. The automatic generation method of a GUI view operation state machine according to claim 2, characterized in that, Component type recognition: Accurately recognize the interactive components of the GUI page, including buttons, links, input boxes, dropdown menus, and checkboxes. Output the prediction results on the S*S grid for the input image X with size H*W through the YOLOv8 object detection model, including object confidence, the coordinate representation of elements, and the probability values of C component types; Component content extraction: Extract the text labels and typed contents of the components; Component position localization: Accurately describe the position of the component in the interface through the bounding box coordinates (x, y, width, height); Element relationship matrix construction: Use supervised learning with a multi-layer perceptron MLP(·) to construct a homogeneous element relationship matrix with n dimensions of the number of elements to represent the relationships between all elements in the page, and complete the subsequent construction of the page structure based on spatial position relationships and visual hierarchy features; For the page element set E = {e1, e2,..., e n}, the feature vector of each element is defined as f i ∈ R d , i = 1, 2,..., n, and the relationship matrix R is obtained: R = [r ij n*n , r ij = MLP([f i , f j )。 4. The automatic generation method of a GUI view operation state machine according to claim 1, characterized in that The specific content of the above Step 2 is as follows: Automatically construct a page logic tree through supervised learning of a multi-layer perceptron. Construct a feature vector and train a supervised learning model by extracting the spatial coordinates, visual attributes, and semantic features of GUI elements to predict the logical relationships between elements and generate a relationship matrix; Perform post-processing on the matrix using topological sorting and spatial hierarchy verification to eliminate circular dependencies and enforce inclusion constraints, and finally construct a page logic tree reflecting the component hierarchy and layout structure, clearly obtaining the layout structure of the web page and the detailed content included in each node.
5. The automatic generation method of a GUI view operation state machine according to claim 1, characterized in that, The interface state machine in step 3 is constructed as follows: intercept and parse the GUI interfaces of web pages and applications, record the user's operation process including the operation type, operation location, input content, and the new interface resulting from the current operation that brings about changes to the interface.
6. The automatic generation method of a GUI view operation state machine according to claim 1, characterized in that The specific content of step 3 is as follows: Construct a state machine model to record the user operation sequence. By recording the interaction operations on interactive dynamic elements, map each operation step to the state transition between page structure nodes to form a traceable operation process. The specific implementation is as follows: First, through the multi-platform event listening mechanism, the operation capture module listens to and records user operation events in real time, including the operation component ID and operation type, and constructs the GUI structure tree nodes and state machine states. Construct a two-layer state model through the state definition module, dynamically associate the GUI structure tree node states with the global state machine states, and achieve an accurate mapping from operation events to state changes through predefined state transition rules. For persistent storage, a hybrid storage strategy is adopted. The operation sequence is stored in a time-series database, and the state machine model is stored in XML format to describe the transfer logic. The backtracking and replay module realizes the reversible traversal of the operation process based on the historical state stack management, supports millisecond-level time-series accuracy replay and a visual debugging interface, combines cross-platform consistency verification to ensure the accuracy of interaction behavior reproduction, provides a standardized and structured description method for the interaction between the user and the GUI, and realizes the efficient processing and response to user instructions.
7. A method for automatically generating a GUI view operation state machine according to claim 1, characterized in that The specific content of step 4 is as follows: Use a visualization tool to present the generated GUI interface structure in a graphical manner. First, abstract all components in the GUI interface structure to be analyzed into nodes, represent the relationship between nodes in a tree structure or hierarchical structure, and connect each node with an edge to represent the parent-child hierarchical relationship. Use a visualization tool to achieve the graphical presentation of the GUI structure.
8. The automatic generation method of a GUI view operation state machine according to claim 7, characterized in that, It includes the following steps: The specific steps of the visualization Graphviz tool are as follows: First, automatically extract the component information of the GUI interface structure to be analyzed to form structured data including component type, identifier, attributes, and parent-child hierarchical relationship. Then convert the above data into a DOT format file supported by Graphviz. Then, by configuring the graphical layout and visual parameters of Graphviz, use its built-in visualization engine to generate a tree structure diagram of GUI components; finally, conduct an intuitive comparison and analysis between the graphical component structure and the actual interface to facilitate quickly discovering and locating the differences between the GUI structure and the actual interface.
9. A device for automatically generating an interface structure of a GUI view, characterized in that, It includes: A memory for storing computer programs. A processor for implementing the automatic generation method of the GUI view operation state machine according to any one of claims 1-7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it can automatically generate the GUI view operation state machine according to any one of claims 1-7 according to the prediction method.
Citation Information
Cited By
GUI (Graphical User Interface) element operation method, system and equipment and storage medium
CN121008724A
Intelligent right steward implementation method and device based on AI capability
CN121144621A
Web page automatic screenshot and anomaly recognition system, method and equipment based on vision-language multi-mode large model and storage medium
CN121170431A
Advertisement interference evaluation benchmark automatic construction method for mobile terminal graphical user interface agent
CN121597559A