UI automation testing method and related device

CN122817064APending Publication Date: 2026-09-25BEIJING HONGTENG INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510355026.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0002]传统的UI自动化测试方法依赖基于DOM的技术,如Selenium、Appium等,高度依赖UI结构、源码和平台,导致其存在诸多问题,增加了频繁维护脚本的负担、适用性狭窄且测试环境复杂

Benefits of technology

[0031]在本申请的一些实施例所提供的技术方案中,捕获UI界面截图后,通过UI界面识别模型对界面中的各UI元素进行识别,摆脱了对DOM结构的依赖,由于UI界面识别模型是神经网络模型,其测试环境也不复杂,适用性也更加广泛,只要给予多样性充足的样本进行训练,就可以摆脱UI结构、源码和平台的桎梏。在自学习机制下可以不必频繁维护脚本,在面对复杂应用和快速迭代的开发环境时,自学习机制也使其具有足够的灵活性和高效性。结合自动化的交互操作工具,可以实现高效、鲁棒的UI自动化测试。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817064A_ABST
    Figure CN122817064A_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a UI automation testing method and related equipment. The UI automation testing method comprises: capturing a UI interface screenshot; inputting the UI interface screenshot into a UI interface recognition model to obtain target element information, the target element information being element information of a target UI element, comprising target element coordinates and a target element type, the target UI element being a UI element in the UI interface screenshot; and performing interactive testing on each target UI element according to the target element coordinates and the target element type to obtain a test result. The technical solution of the embodiments of the present application, after capturing a UI interface screenshot, identifies each UI element in the interface through a UI interface recognition model, thereby breaking the dependence on a DOM structure. Since the UI interface recognition model is a neural network model, the test environment is not complex, and the applicability is more extensive. As long as a sufficient number of samples are given for training, the UI structure, source code and platform can be broken free from.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of computer and communication technology, and more specifically, to a UI automation testing method and related equipment. Background Technology

[0002] Traditional UI automation testing methods rely on DOM-based technologies such as Selenium and Appium, which are highly dependent on UI structure, source code, and platform. This leads to many problems, increases the burden of frequent script maintenance, narrow applicability, and complex testing environments.

[0003] Therefore, traditional UI automation testing methods cannot provide sufficient flexibility and efficiency when faced with complex applications and rapidly iterating development environments. Summary of the Invention

[0004] The embodiments of this application provide a UI automated testing method and related equipment, which can at least to some extent reduce the dependence on UI structure and platform, and improve the reliability and automation of testing.

[0005] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.

[0006] According to one aspect of the embodiments of this application, a UI automated testing method is provided, comprising: capturing a screenshot of a UI interface; inputting the screenshot of the UI interface into a UI interface recognition model to obtain target element information, wherein the target element information is element information of a target UI element, including target element coordinates and target element type, and the target UI element is a UI element in the screenshot of the UI interface; performing interaction tests on each target UI element according to the target element coordinates and target element type to obtain test results.

[0007] In some feasible embodiments of this application, the UI interface recognition model includes a scene recognition sub-model and multiple interface recognition sub-models. The step of inputting the UI interface screenshot into the UI interface recognition model to obtain target element information specifically includes: inputting the interface screenshot into the scene recognition sub-model to obtain an interface scene; calling the corresponding interface recognition sub-model based on the interface scene, wherein the interface recognition sub-model and the interface scene correspond one-to-one; and inputting the interface screenshot into the interface recognition sub-model to obtain target element information.

[0008] In some feasible embodiments of this application, the step of calling the corresponding interface recognition sub-model according to the interface scenario specifically includes: determining the corresponding interface recognition sub-model according to the interface scenario; and sending a calling instruction to the cache dictionary so as to load the interface recognition sub-model from the cache dictionary into memory.

[0009] In some feasible embodiments of this application, the UI automated testing method further includes: in response to the completion of use of the interface sub-model, unloading the interface sub-model from memory to the cache dictionary.

[0010] In some feasible embodiments of this application, the UI automated testing method further includes: acquiring a set of UI interface screenshot samples, training the UI interface recognition model, and obtaining a trained UI interface recognition model.

[0011] In some feasible embodiments of this application, the UI interface recognition model includes a scene recognition sub-model and multiple interface recognition sub-models. The step of obtaining a UI interface screenshot sample set and training the UI interface recognition model to obtain a trained UI interface recognition model specifically includes: obtaining a common object image recognition sample set and pre-training each of the interface recognition sub-models to obtain corresponding pre-trained interface recognition sub-models; obtaining a UI interface screenshot sample set, which contains multiple subsets of UI interface screenshot samples from different scenes, each subset containing multiple UI interface screenshot samples from the corresponding scene; inputting each UI interface screenshot sample from the subset into the pre-trained interface recognition sub-model for training in the corresponding scene to obtain a trained interface recognition sub-model; and inputting each UI interface screenshot sample into the scene recognition sub-model to train both the scene recognition sub-model and the interface recognition sub-models to obtain a trained UI interface recognition model.

[0012] In some feasible embodiments of this application, the recognition box label corresponding to each UI interface screenshot sample, and the step of inputting the UI interface screenshot samples one by one into the scene recognition sub-model, training the scene recognition sub-model and each of the interface recognition sub-models to obtain a trained UI interface recognition model, specifically includes: inputting the UI interface screenshot samples one by one into the scene recognition sub-model to obtain the corresponding scene result; calling the corresponding interface recognition sub-model according to the scene result; inputting the UI interface screenshot samples into the corresponding scene recognition sub-model to obtain target element information; updating the parameters of the scene recognition sub-model and each of the interface recognition sub-models according to the target element information and the recognition box label, until a predetermined termination condition is reached, ending the training, and obtaining a trained UI interface recognition model.

[0013] In some feasible embodiments of this application, the step of inputting the UI screenshot into the UI interface recognition model to obtain target element information specifically includes: inputting the UI screenshot into the UI interface recognition model to obtain the element position with corresponding confidence and target element type; and determining the target element coordinates based on the element position and corresponding confidence.

[0014] In some feasible embodiments of this application, determining the target element coordinates based on the element position and the corresponding confidence level specifically includes: determining the target element position based on the element position and the corresponding confidence level; and converting the target element position into target element coordinates based on the actual screen coordinates.

[0015] In some feasible embodiments of this application, the step of performing interaction tests on each of the target UI elements according to the target element coordinates and target element type to obtain test results specifically includes: performing interaction operations on each of the target UI elements according to the target element coordinates and target element type to obtain operation results; and obtaining the test results based on each operation result.

[0016] In some feasible embodiments of this application, the operation result includes a post-operation interface; obtaining the test result based on each of the operation results specifically includes: determining whether the operation was successful based on the post-operation interface; and determining the test result based on the result of whether the operation was successful based on the result of whether the operation was successful based on the post-operation interface.

[0017] According to one aspect of the embodiments of this application, a UI automated testing device is provided, the UI automated testing device comprising: an interface screenshot capture module for capturing UI interface screenshots; a UI interface recognition module for inputting the UI interface screenshots into a UI interface recognition model to obtain target element information, the target element information being element information of a target UI element, including target element coordinates and target element type, the target UI element being a UI element in the UI interface screenshot; and a test result generation module for performing interaction tests on each of the target UI elements according to the target element coordinates and target element type to obtain test results.

[0018] In some feasible embodiments of this application, the UI interface recognition model includes a scene recognition sub-model and multiple interface recognition sub-models. The UI interface recognition module specifically includes: an interface scene recognition sub-module, used to input the interface screenshot into the scene recognition sub-model to obtain an interface scene; a scene model calling sub-module, used to call the corresponding interface recognition sub-model according to the interface scene, wherein the interface recognition sub-model and the interface scene correspond one-to-one; and an element information recognition sub-module, used to input the interface screenshot into the interface recognition sub-model to obtain target element information.

[0019] In some feasible embodiments of this application, the scene model invocation submodule specifically includes: a submodel determination unit, used to determine the corresponding interface recognition submodel according to the interface scene; and a submodel invocation unit, used to send an invocation instruction to the cache dictionary so as to load the interface recognition submodel from the cache dictionary into memory.

[0020] In some feasible embodiments of this application, the UI automated testing device further includes: a sub-model unloading unit, used to unload the interface sub-model from memory to the cache dictionary in response to the completion of the use of the interface sub-model.

[0021] In some feasible embodiments of this application, the UI automated testing device further includes: a model training module, used to acquire a set of UI interface screenshot samples, train the UI interface recognition model, and obtain a trained UI interface recognition model.

[0022] In some feasible embodiments of this application, the UI interface recognition model includes a scene recognition sub-model and multiple interface recognition sub-models. The model training module specifically includes: a pre-training sub-module, used to acquire a common object image recognition sample set and pre-train each of the interface recognition sub-models to obtain a corresponding pre-trained interface recognition sub-model; a sample acquisition sub-module, used to acquire a UI interface screenshot sample set, which contains multiple subsets of UI interface screenshot samples under different scenes, and each subset of UI interface screenshot samples contains multiple UI interface screenshot samples under the corresponding scene; a separate training sub-module, used to input each UI interface screenshot sample in the subset of UI interface screenshot samples one by one into the pre-trained interface recognition sub-model under the corresponding scene for training to obtain a trained interface recognition sub-model; and a joint training sub-module, used to input each UI interface screenshot sample into the scene recognition sub-model one by one, and train the scene recognition sub-model and each of the interface recognition sub-models to obtain a trained UI interface recognition model.

[0023] In some feasible embodiments of this application, the recognition box label corresponding to each UI screenshot sample, the joint training submodule specifically includes: a sample input unit, used to input the UI screenshot samples one by one into the scene recognition sub-model to obtain the corresponding scene result; a model calling unit, used to call the corresponding interface recognition sub-model according to the scene result; an element recognition unit, used to input the UI screenshot samples into the corresponding scene recognition sub-model to obtain target element information; and a parameter update unit, used to update the parameters of the scene recognition sub-model and each of the interface recognition sub-models according to the target element information and the recognition box label, until a predetermined termination condition is reached, the training ends, and a trained UI interface recognition model is obtained.

[0024] In some feasible embodiments of this application, the UI interface recognition module specifically includes: a position determination submodule, used to input the UI interface screenshot into the UI interface recognition model to obtain the element position with the corresponding confidence level and target element type; and a coordinate transformation submodule, used to determine the target element coordinates based on the element position and the corresponding confidence level.

[0025] In some feasible embodiments of this application, determining the target element coordinates based on the element position and the corresponding confidence level specifically includes: a target element position unit, used to determine the target element position based on the element position and the corresponding confidence level; and a target element coordinate unit, used to convert the target element position into target element coordinates based on the actual screen coordinates.

[0026] In some feasible embodiments of this application, the test result generation module specifically includes: an interactive operation submodule, used to perform interactive operations on each of the target UI elements according to the target element coordinates and target element type, and obtain operation results; and a test result submodule, used to obtain the test results according to each of the operation results.

[0027] In some feasible embodiments of this application, the operation result includes a post-operation interface; the test result submodule specifically includes: an operation determination unit, used to determine whether the operation was successful based on the post-operation interface; and a test result unit, used to determine the test result based on the result of whether the operation was successful on each post-operation interface.

[0028] According to one aspect of the embodiments of this application, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processor, implements the UI automated testing method as described in the above embodiments.

[0029] According to one aspect of the embodiments of this application, an electronic device is provided, including: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the UI automation testing method as described in the above embodiments.

[0030] According to one aspect of the embodiments of this application, a computer program product is provided, including one or more computer programs that, when executed by one or more processors, implement the steps of the UI automated testing method as described in the above embodiments.

[0031] In some embodiments of this application, after capturing a screenshot of the UI interface, a UI interface recognition model is used to identify each UI element in the interface, thus eliminating the dependence on the DOM structure. Since the UI interface recognition model is a neural network model, its testing environment is not complex, and its applicability is wider. As long as it is trained with a sufficient variety of samples, it can break free from the constraints of UI structure, source code, and platform. The self-learning mechanism eliminates the need for frequent script maintenance. In the face of complex applications and rapidly iterating development environments, the self-learning mechanism also provides sufficient flexibility and efficiency. Combined with automated interactive operation tools, efficient and robust UI automated testing can be achieved.

[0032] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0033] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:

[0034] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of this application can be applied is shown.

[0035] Figure 2 The diagram shows a flowchart of a UI automation testing method provided in an embodiment of this application.

[0036] Figure 3 It shows that according to Figure 2 A flowchart illustrating a specific implementation of step S200 in the UI automation testing method shown in the corresponding embodiment.

[0037] Figure 4 A flowchart illustrating the training method of the UI interface recognition model provided in this application embodiment is shown.

[0038] Figure 5 It shows that according to Figure 2 A flowchart illustrating a specific implementation of step S300 in the UI automation testing method shown in the corresponding embodiment.

[0039] Figure 6 It shows that according to Figure 2 A flowchart illustrating a specific implementation of step S300 in the UI automation testing method shown in the corresponding embodiment.

[0040] Figure 7A schematic diagram of the structure of a UI automated testing device provided in an embodiment of this application is shown.

[0041] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation

[0042] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.

[0043] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.

[0044] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0045] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0046] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of this application can be applied is shown.

[0047] like Figure 1 As shown, the system architecture may include terminal devices (such as...) Figure 1 The device shown includes one or more of a smartphone 101, tablet 102, and portable computer 103 (which could also be a desktop computer, etc.), a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the terminal device and the server 105. The network 104 can include various connection types, such as wired communication links, wireless communication links, etc.

[0048] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, there can be any number of terminal devices, networks, and servers. For example, server 105 could be a server cluster composed of multiple servers.

[0049] Users can use terminal devices to interact with server 105 via network 104 to receive or send messages, etc. Server 105 can be a server that provides various services. For example, a user can upload a screenshot of a UI interface to server 105 using terminal device 103 (or terminal device 101 or 102). Server 105 can input the screenshot into a UI interface recognition model to obtain target element information. The target element information is the element information of the target UI element, including the target element coordinates and the target element type. The target UI element is the UI element in the screenshot. Based on the target element coordinates and target element type, interaction tests are performed on each target UI element to obtain test results.

[0050] It should be noted that the UI automated testing method provided in this application embodiment is generally executed by server 105, and correspondingly, the UI automated testing device is generally set in server 105. However, in other embodiments of this application, the terminal device may also have similar functions to the server, thereby executing the UI automated testing scheme provided in this application embodiment.

[0051] The implementation details of the technical solutions in the embodiments of this application are described in detail below:

[0052] Figure 2 A flowchart of a UI automation testing method according to an embodiment of this application is shown. This UI automation testing method can be executed by a server, which may be... Figure 1 The server shown. (Refer to...) Figure 2 As shown, this UI automation testing method includes at least the following:

[0053] S100, capturing a screenshot of the UI interface.

[0054] S200, input the UI interface screenshot into the UI interface recognition model to obtain target element information. The target element information is the element information of the target UI element, including the target element coordinates and the target element type. The target UI element is the UI element in the UI interface screenshot.

[0055] S300, based on the target element coordinates and target element type, perform interaction tests on each target UI element to obtain test results.

[0056] In this embodiment, after capturing a screenshot of the UI interface, the UI elements in the interface are identified through a UI interface recognition model, eliminating the dependence on the DOM structure. Since the UI interface recognition model is a neural network model, its testing environment is not complex, and its applicability is wider. As long as it is trained with a sufficient variety of samples, it can break free from the constraints of UI structure, source code, and platform. The self-learning mechanism eliminates the need for frequent script maintenance. In the face of complex applications and rapidly iterating development environments, the self-learning mechanism also provides sufficient flexibility and efficiency. Combined with automated interactive tools, efficient and robust automated UI testing can be achieved.

[0057] This embodiment utilizes a UI interface recognition model to automatically identify UI elements in screenshots and extract information from the target elements. Automated recognition reduces manual operation, lowers the probability of errors, and enhances the consistency and reliability of the testing process. Simultaneously, based on the precise location and recognition of UI elements, corresponding interactive operations are automatically executed for different types of UI elements, covering all UI element interactive operations to ensure no potential UI issues are missed, thus ensuring the comprehensiveness and accuracy of interaction testing. Furthermore, by automatically capturing screenshots, recognizing elements, executing interactions, and obtaining test results, the efficiency of UI automation testing is greatly improved, saving significant labor costs.

[0058] The S100 provides an efficient way to obtain interface images by capturing UI screenshots. UI screenshots contain information about all UI elements, especially static visual information such as buttons, text boxes, and icons, ensuring the comprehensiveness of subsequent testing, avoiding manual intervention, and obtaining accurate interface snapshots for later processing.

[0059] In S200, captured UI screenshots are input into the UI recognition model. This model, potentially based on deep learning or computer vision techniques (e.g., the YOLOv8 model), automatically identifies and extracts information about target elements, such as their coordinates and type. By learning from a large number of samples, the model can efficiently and accurately identify UI elements on the interface and extract their features. The target element coordinates refer to the position of the UI element within the interface. The target element type refers to the type of UI element, such as a button, text box, or checkbox.

[0060] This embodiment automatically identifies and extracts target element information, avoiding manual annotation and operation, improving the accuracy of identification, supporting element identification under complex interface layouts, adapting to different UI design styles, and improving the applicability of testing.

[0061] Specifically, in some embodiments, the specific implementation of step S200 can be found in [reference needed]. Figure 3 . Figure 3 It is based on Figure 2 According to the detailed description of step S200 in the UI automated testing method shown in the corresponding embodiment, the UI interface recognition model in the UI automated testing method includes a scene recognition sub-model and multiple interface recognition sub-models. Step S200 may include the following steps:

[0062] S210, input the screenshot of the interface into the scene recognition sub-model to obtain the interface scene.

[0063] S220, based on the interface scenario, call the corresponding interface recognition sub-model, the interface recognition sub-model and the interface scenario correspond one-to-one.

[0064] S230, input the screenshot of the interface into the interface recognition sub-model to obtain the target element information.

[0065] This embodiment inputs a screenshot of the UI interface into the scene recognition sub-model and calls the corresponding interface recognition sub-model according to the scene type. This method can flexibly adapt to different types of UI interfaces. Each recognition sub-model focuses on a specific scene, improving the recognition efficiency for complex UI interfaces.

[0066] By combining scene recognition and interface recognition sub-models, target elements in the UI can be identified efficiently and accurately, and their position and type information can be extracted. Through this hierarchical recognition approach, the testing method can automatically select the optimal recognition strategy in different UI scenarios, thereby solving the technical difficulties caused by the complexity of the interface in traditional methods.

[0067] In this embodiment, the combination of scene recognition and sub-models enables the method to adapt to various complex interfaces, automatically selecting recognition strategies to ensure accurate identification of each UI element. This avoids the misidentification and omission issues found in traditional methods, thus improving testing efficiency. Furthermore, through the collaborative work of multiple sub-models, the method can handle different UI designs and layouts, ensuring good adaptability to interfaces in various scenarios.

[0068] In S210, the screenshot is input into the scene recognition sub-model to identify the general scene of the interface, that is, to determine which category the interface belongs to (such as settings interface, login interface, homepage, etc.). Specifically, the scene recognition sub-model analyzes the layout, element combination and visual structure of the entire screenshot based on image features, thereby determining the scene category of the interface.

[0069] This step identifies different UI scenarios, allowing automated testing to select appropriate identification methods and interaction strategies for different types of interfaces. For example, the homepage might need to identify the navigation bar, while the settings interface focuses on input boxes and buttons, ensuring the model can adapt to different types of interface designs and layouts, avoiding the limitations of traditional methods on interface types.

[0070] In S220, once the scene recognition sub-model identifies the interface scene, it can call a specific interface recognition sub-model to further analyze the specific structure and elements of the interface based on the identified scene type. The interface recognition sub-model corresponds one-to-one with a specific interface scene, and each interface recognition sub-model is specifically designed for a certain type of UI interface, thus having higher recognition accuracy.

[0071] In this step, each interface recognition sub-model is trained and optimized based on specific interface scenarios to ensure accurate recognition of UI elements in each scenario. This on-demand approach improves the system's adaptability and accuracy. It supports multiple UI design styles and can perform personalized recognition for different types of interfaces (such as forms, dialog boxes, homepages, etc.), further enhancing automated testing capabilities.

[0072] Specifically, in some embodiments, the specific implementation of step S220 can be found in the following embodiments. This embodiment is based on... Figure 3 According to the detailed description of step S220 in the UI automation testing method shown in the corresponding embodiment, step S220 in the UI automation testing method may include the following steps:

[0073] Based on the interface scenario, determine the corresponding interface recognition sub-model.

[0074] Send a call instruction to the cache dictionary to load the interface recognition sub-model from the cache dictionary into memory.

[0075] In this embodiment, by introducing a cache dictionary, model reloading is avoided every time, reducing memory and computing resource consumption, optimizing resource management, and thus significantly improving model loading efficiency. Using a cache dictionary, the model can respond quickly while consuming fewer computing resources, reducing the time overhead of repeated loading and thereby improving the speed of UI automation testing. Combining scene recognition and a cache dictionary, appropriate sub-models can be called according to different UI interface scenarios, achieving high-precision and high-efficiency UI automation testing.

[0076] In one embodiment of this application, after the interface sub-model has been used, the UI automated testing method further includes:

[0077] In response to the completion of use of the interface sub-model, the interface sub-model is unloaded from memory and placed into the cache dictionary.

[0078] In this embodiment, once it is confirmed that the sub-model has finished using, it can be removed from memory and stored again in the previously mentioned cache dictionary. The cache dictionary, as a storage area for models, acts as a "temporary repository" for models. When the same sub-model needs to be used again, it can be quickly restored from the cache dictionary without reloading, avoiding the process of repeatedly loading models.

[0079] This embodiment frees up memory space by unloading sub-models when they are no longer needed, avoiding excessive memory consumption, optimizing memory management, and allowing memory to be used for other more demanding tasks. It also prevents excessive memory pressure caused by too many sub-models residing in memory. Simultaneously, by storing models in a cache dictionary, it ensures that they can be quickly restored from the cache when needed in the future, without having to load the models from scratch. This not only improves the speed of model invocation and loading but also reduces the demand on computing resources, improving the overall execution efficiency of UI automation testing. Furthermore, through the aforementioned intelligent unloading and caching mechanisms, the load is effectively controlled, avoiding performance bottlenecks caused by insufficient memory and improving the stability of testing tasks.

[0080] In S230, after the interface scene is determined, the interface screenshot is input into the corresponding interface recognition sub-model, which further extracts information about the UI elements. This information includes the position (coordinates), type (button, text box, checkbox, etc.), and other features of the UI elements.

[0081] Each interface recognition sub-model can identify and label the specific location of UI elements in the interface based on its own design, and classify the types of these elements.

[0082] This step ensures accurate identification of each interface element, avoiding potential misidentification and omissions in traditional methods. Through training with a deep learning model, the interface recognition sub-model can efficiently extract feature information for each element, ensuring the comprehensiveness and accuracy of the test.

[0083] Automated extraction of UI element information greatly reduces manual operation and intervention, improves testing efficiency, and reduces the possibility of human error.

[0084] It should be noted that the training methods for the aforementioned UI interface recognition model may include:

[0085] Obtain a sample set of UI interface screenshots, train the UI interface recognition model, and obtain a trained UI interface recognition model.

[0086] In this embodiment, by acquiring and training a diverse range of UI screenshot samples, the accuracy of the UI recognition model can be effectively improved, enabling it to adapt to different UI designs, device types, and screen resolutions. The training process allows the model to learn diverse interface elements and layout patterns, improving its adaptability to UI changes and ensuring testing accuracy across different versions and environments.

[0087] Through training, the model can accurately identify UI elements in different testing scenarios, improving the effectiveness of automated testing. A well-trained model can efficiently identify UI elements and execute corresponding test operations, saving significant manual intervention and increasing testing efficiency. Furthermore, a well-trained model can adapt to different interface and device environments, thereby improving the stability and reliability of the UI automated testing system.

[0088] In some embodiments, the specific training methods can be referred to Figure 4 In the embodiment shown, the UI interface recognition model includes a scene recognition sub-model and multiple interface recognition sub-models. The training method for the UI interface recognition model may specifically include:

[0089] S410, Obtain a common object image recognition sample set, and pre-train each of the interface recognition sub-models to obtain the corresponding pre-trained interface recognition sub-models.

[0090] S420, Obtain a UI screenshot sample set. The UI screenshot sample set contains multiple UI screenshot sample subsets under different scenarios, and each UI screenshot sample subset contains multiple UI screenshot samples under the corresponding scenario.

[0091] S430, each UI screenshot sample in the subset of UI screenshot samples is input into the pre-trained interface recognition sub-model in the corresponding scenario for training, and the trained interface recognition sub-model is obtained.

[0092] S440, input the UI interface screenshot samples one by one into the scene recognition sub-model, train the scene recognition sub-model and each of the interface recognition sub-models to obtain the trained UI interface recognition model.

[0093] In this embodiment, by combining a scene recognition sub-model and multiple interface recognition sub-models, the technical challenge of identifying different UI interfaces and scenes in UI automation testing is solved. Specifically, the scene recognition sub-model can identify the scene to which the entire UI interface belongs, while each interface recognition sub-model focuses on identifying interface elements within a specific scene. Through pre-training and training of these models, an accurate and efficient UI interface recognition model can be obtained, thereby greatly improving the accuracy and efficiency of automated testing.

[0094] In S410, a common object image sample set is acquired to provide initial learning data for each interface recognition sub-model. This sample set contains a large number of common UI elements, such as buttons, text boxes, and icons, aiming to give each interface recognition sub-model a basic understanding of these elements. The pre-training phase helps the model better recognize these common objects, thus laying the foundation for subsequent recognition of more complex UI screenshots.

[0095] In S420, each sample subset represents a different scene, such as the main interface, settings interface, menu interface, etc. Each scene contains multiple UI screenshots, covering different interface changes and states. These sample sets provide diverse training data for the scene recognition sub-model and the interface recognition sub-model, ensuring that the model can handle various different UI scenes and interfaces.

[0096] In S430, each UI screenshot sample is input into the pre-trained UI recognition sub-model for the corresponding scene. Through training, the UI recognition sub-model learns how to recognize the structure and content of UI elements in a specific scene, enabling each UI recognition sub-model to handle UI interfaces in that specific scene and allowing the model to more accurately identify and classify UI elements.

[0097] In S440, all UI screenshot samples are input into the scene recognition sub-model. The task of the scene recognition sub-model is to identify which scene the current UI belongs to (e.g., the main screen or settings screen). In this step, the scene recognition sub-model and each interface recognition sub-model are trained simultaneously so that, in actual testing, the model can simultaneously recognize the scene of the entire UI and specific UI elements.

[0098] After training, a complete UI interface recognition model is obtained, which can accurately recognize UI interface elements in different scenarios.

[0099] Specifically, in some embodiments, the specific implementation of step S440 can be found in the following embodiments. This embodiment is based on... Figure 4In the detailed description of step S440 of the UI automated testing method shown in the corresponding embodiment, the recognition box label corresponding to each UI interface screenshot sample in the UI automated testing method may include the following steps:

[0100] The UI screenshot samples are input one by one into the scene recognition sub-model to obtain the corresponding scene results.

[0101] Based on the scenario results, the corresponding interface recognition sub-model is invoked.

[0102] The UI screenshot sample is input into the corresponding scene recognition sub-model to obtain the target element information.

[0103] Based on the target element information and the recognition box label, the parameters of the scene recognition sub-model and each of the interface recognition sub-models are updated until the predetermined termination condition is met, and the training ends, resulting in a trained UI interface recognition model.

[0104] In this embodiment, by finely matching UI screenshot samples with corresponding labels, the model can more accurately adjust parameters during training, thereby improving recognition accuracy. Furthermore, using scene results to dynamically select the appropriate interface recognition sub-model avoids unnecessary computation and improves training efficiency. In the above training process, transfer learning can be employed, using yolov8m.pt as the pre-trained model, and the labeled dataset (1268 images in total, with a training set size of 634 and a validation set size of 634) is input into the model. When images are input into the YOLOv8m model, image preprocessing is first performed, including resizing the images to 640x640 and performing data augmentation (such as mosaic, scaling, color adjustment, etc.). Then, multi-scale features are extracted through the CSPDarknet53 backbone network and feature pyramid structure. Next, the detection head performs classification prediction, bounding box prediction, and confidence prediction, and calculates three key losses: bounding box loss (box_loss), classification loss (cls_loss), and distribution focus loss (dfl_loss). Finally, the final detection results are obtained through non-maximum suppression (NMS) and confidence thresholding, and the model performance is evaluated using metrics such as precision, recall, and mAP. Throughout the training process, the learning rate is continuously adjusted and the model parameters are updated by the optimizer, ultimately saving the weights of the model with the best performance.

[0105] We train separately for each UI automation scenario, with 100 training epochs (this can be adjusted based on the convergence of the validation set loss). The specific training parameters are as follows (for reference only), and need to be continuously optimized based on mAP50, precision, and recall.

[0106] Specifically, in other embodiments, the specific implementation of step S200 can be found in [reference needed]. Figure 5 . Figure 5 It is based on Figure 2 According to the detailed description of step S200 in the UI automation testing method shown in the corresponding embodiment, step S200 in the UI automation testing method may include the following steps:

[0107] S250, input the screenshot of the UI interface into the UI interface recognition model to obtain the element position with the corresponding confidence level and target element type.

[0108] S260, determine the coordinates of the target element based on the element position and the corresponding confidence level.

[0109] In this embodiment, by refining the UI recognition process, the accuracy of location and the quantification of confidence are improved, thereby providing a more accurate UI element recognition function. Simultaneously, by obtaining the target element's position coordinates, confidence level, and type, the reliability and stability of UI testing can be effectively improved.

[0110] This embodiment, by specifying the exact location of the element and its corresponding confidence level, can more accurately locate the target element on the UI interface, avoiding accidental operations. It also provides a quantified confidence value to help determine the reliability of the recognition results, thus providing decision support for subsequent operations. Simultaneously, more precise target element coordinates make the execution of automated scripts more accurate and efficient, reducing the occurrence of errors.

[0111] In S250, a screenshot of the UI is input into the UI recognition model. The model analyzes the input UI, first outputting the position of each element, i.e., its relative coordinates within the interface (e.g., top-left corner, bottom-right corner, etc.). Simultaneously, the model outputs the confidence score of the element based on image features, representing the model's confidence in recognizing that element (e.g., a 95% confidence score means the model considers the element to be correct with a 95% probability). Finally, the model determines the type of the target element (e.g., button, text box, checkbox, etc.). This means the model not only knows the exact location of the element but can also identify its specific function.

[0112] This step provides more accurate and reliable identification results for subsequent UI automation operations by outputting the element's location information, confidence level, and type, especially in complex or dynamic UI interfaces.

[0113] In S260, after obtaining the element's position and confidence level, the precise coordinates of the target element can be further determined based on this information. The coordinates are determined through refined calculations based on the element's position on the interface and the confidence level.

[0114] For example, the positions of elements with high confidence may be further confirmed, while elements with low confidence may be marked as unreliable, or further improved in subsequent training and optimization.

[0115] This step, through further refinement of element positions and quantification of confidence levels, ensures the accuracy of the final identified target element coordinates. By introducing confidence level judgments, erroneous positioning due to misidentification can be avoided, ensuring that every operation in automated testing is more precise.

[0116] Specifically, the user provides a screenshot of the UI interface, which is then input into a pre-trained UI recognition model. The model analyzes the image, extracts each UI element, and determines the position, type, and corresponding confidence level of each element based on existing training data. Through image processing and deep learning models, each element in the UI interface (such as buttons, text boxes, etc.) is identified, and its specific coordinates and type are provided. For example, the model might identify a button located in the upper left corner of the screen, classified as a "submit button," and assign a high confidence value (e.g., 0.95). Based on the element's position and confidence level, the precise coordinates of the target element are further confirmed. If the confidence level is low, additional processing may be performed or the element may be marked as unreliable; if the confidence level is high, the element's position is considered accurate, and further automated operations can be performed. Through this precise element location and confidence evaluation mechanism, UI automated testing can more efficiently identify and operate the correct UI elements, avoiding test failures or misoperations caused by incorrect identification.

[0117] Specifically, in some embodiments, the specific implementation of step S260 can be found in the following embodiments. This embodiment is based on... Figure 5 According to the detailed description of step S260 in the UI automation testing method shown in the corresponding embodiment, step S260 in the UI automation testing method may include the following steps:

[0118] The target element position is determined based on the element position and the corresponding confidence level.

[0119] Based on the actual screen coordinates, the position of the target element is converted into target element coordinates.

[0120] In this embodiment, by combining element position with confidence level and converting the position into actual screen coordinates, the problem of automated operation deviation caused by different screen resolutions and layouts is solved.

[0121] This embodiment, by considering confidence levels, can accurately determine the position of the target element based on the current interface state and convert this position into actual screen coordinates, ensuring the accuracy of the operation. It can also cope with device differences or changes in UI layout, thereby improving the robustness and stability of UI automation testing. Furthermore, by converting to screen coordinates, automated testing can be performed across devices with different resolutions and screen sizes, ensuring operational consistency.

[0122] Specifically, the system first obtains the relative position of UI elements (e.g., the coordinates of the element relative to the top left corner of the screen) and the confidence level of the element (whether it is operable, visible, etc.) based on the UI interface recognition model. For example, assuming a button is located in the top left corner of the interface with relative coordinates (50, 100) and a high confidence level, the system will determine that this element is the target element and process it.

[0123] Then, by combining the element's relative position and confidence level, the system determines whether it is the target element and processes it accordingly. If the element's confidence level is high, the system continues to process it; if the confidence level is low, the system may skip the element. For example, if the element's confidence level is high, the system determines the target element's position to be (50, 100).

[0124] Next, based on information such as the device's resolution and screen size, the relative position (e.g., 50, 100) is converted into absolute screen coordinates. Assuming the device's screen resolution is 1920x1080, after conversion, the element's absolute coordinates may become actual values ​​such as (1000, 200).

[0125] Once the screen coordinates of the target element are obtained, these coordinates can be used for subsequent automated operations, such as clicking the element or entering text. Through coordinate transformation, automated testing can ensure the accuracy of the operation and avoid deviations caused by different screen layouts or resolutions.

[0126] In the S300, UI element coordinates allow for precise element location, enabling automation tools to click, input, or select corresponding UI elements. Different types of UI elements require different interaction methods. For example, buttons require clicking, text boxes require input, and checkboxes require checking or dechecking. Automatically selecting the appropriate interaction method based on the element type can simulate real user behavior.

[0127] This step automatically executes interactive operations (such as clicking, text input, etc.) and tests them based on the target element coordinates and type extracted in the previous steps. By automating interactive testing, actual user operations can be simulated to evaluate whether the responsiveness and functionality of interface elements meet expectations.

[0128] Specifically, in some embodiments, the specific implementation of step S300 can be found in [reference needed]. Figure 6 . Figure 6 It is based on Figure 2 According to the detailed description of step S300 in the UI automation testing method shown in the corresponding embodiment, step S300 in the UI automation testing method may include the following steps:

[0129] S310, based on the target element coordinates and target element type, perform interactive operations on each of the target UI elements to obtain the operation results.

[0130] S320, based on the results of each operation, obtain the test results.

[0131] In this embodiment, by combining the target element coordinates with the target element type, the interaction operation method of each UI element is clarified, and test results are automatically generated based on the results of the interaction operation, thereby improving the automation level of the test and reducing manual intervention.

[0132] Furthermore, selecting the appropriate interaction method (such as clicking, text input, swiping, etc.) based on the element type ensures that each type of element is handled appropriately. Different types of UI elements may have different interaction requirements; by combining element class and coordinates, personalized operation strategies can be developed for each element to ensure the accuracy of the operation.

[0133] In S310, the system determines how to interact with the target element based on its type and coordinates. For example, if the target element is a button with coordinates (1000, 200), the system will simulate a mouse click at that location. If the target element is a text box, predefined text will be entered. After each interaction, the system obtains the result based on the UI's response.

[0134] The target element coordinates are the screen coordinates of the target element that were determined in the previous steps (i.e., the precise position of the element on the screen). This step will locate each target element based on these coordinates and perform the corresponding interactive operation.

[0135] The target element type has been determined in the preceding steps. Each UI element may belong to a different type, such as a button, text box, checkbox, dropdown menu, etc. Different types of elements require different operation methods. For example, a button requires a click operation; a text box requires a text input operation; a checkbox requires a check or decheck operation; and a dropdown menu requires a selection operation.

[0136] Operation results typically include whether an element responded to the operation (e.g., whether a button was successfully clicked, or whether text box content was successfully entered) and whether the UI interface changed as expected after the operation (e.g., whether a corresponding dialog box popped up after clicking a button, or whether the correct text was displayed after entering it into a text box). The criteria for judging the operation result may include changes in the state of UI elements (e.g., a button becoming disabled or a text box displaying new content) and changes in the display of the UI interface (e.g., pop-ups, tooltips, page transitions, etc.).

[0137] In S320, test reports or feedback information are generated based on the operation results.

[0138] Test results typically include whether the operation was completed successfully, whether the UI elements responded as expected, and whether UI errors or abnormal behavior occurred.

[0139] For example, suppose there is a button in the UI. Clicking the button successfully redirects the page to the expected new page. In this case, the operation is considered successful, and the test result will show that the operation passed. If the operation fails, such as an invalid button click or a text box failing to display the correct content, the test result will show failure or error, and may provide a detailed reason for the failure (such as the page not refreshing, the element not being activated, etc.).

[0140] The generated test results can be further used for regression testing and error logging.

[0141] Specifically, in some embodiments, the specific implementation of step S320 can be found in the following embodiments. This embodiment is based on... Figure 6 The detailed description of step S320 in the UI automation testing method shown in the corresponding embodiment, wherein the operation result includes the interface after the operation, and step S320 may include the following steps:

[0142] Based on the post-operation interface, determine whether the operation was successful.

[0143] The test results are determined based on whether the operation was successful on the interface after each operation.

[0144] In this embodiment, by introducing the post-operation interface as a standard for judging the success of the operation, the judgment criteria during the testing process are ensured to be more comprehensive and accurate. By monitoring the changes in the UI interface after the operation, it is possible to determine whether the operation had the expected effect on the UI, thus improving the accuracy of the test results. This embodiment not only relies on the operation result itself but also uses interface changes to determine the success of the operation, making the judgment more comprehensive and avoiding omissions or misjudgments. Combining the operation result and the post-operation interface changes, test reports can be generated more intelligently, ensuring the reliability and accuracy of the results.

[0145] After the operation is completed, the system enters the result verification and analysis phase to ensure the correctness of the operation and the achievement of the test objectives. In some embodiments, screenshots of the interface after the operation can be captured, and the categories identified in the images can be labeled and saved to determine whether the interface state meets expectations. Batch execution of test cases in multiple test scenarios can verify the system's performance under different environments and interface states. Analyzing the test results and recording the success rate and reasons for failure provides data support for system optimization. Automatically generating test reports, including test execution status (success / failure rate), time consumption, failure case analysis, etc., are presented intuitively in the form of charts and text.

[0146] The following describes an apparatus embodiment of this application, which can be used to execute the UI automated testing method in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the UI automated testing method described above.

[0147] Figure 7 A block diagram of a UI automation testing apparatus according to an embodiment of this application is shown.

[0148] Reference Figure 7 As shown, a UI automated testing device 700 according to an embodiment of this application includes: an interface screenshot capture module 710, a UI interface recognition module 720, and a test result generation module 730.

[0149] The interface screenshot capture module 710 is used to capture UI interface screenshots; the UI interface recognition module 720 is used to input the UI interface screenshots into the UI interface recognition model to obtain target element information, wherein the target element information is the element information of the target UI element, including the target element coordinates and the target element type, and the target UI element is the UI element in the UI interface screenshot; the test result generation module 730 is used to perform interaction tests on each of the target UI elements according to the target element coordinates and the target element type to obtain test results.

[0150] In some feasible embodiments of this application, the UI interface recognition model includes a scene recognition sub-model and multiple interface recognition sub-models. The UI interface recognition module specifically includes: an interface scene recognition sub-module, used to input the interface screenshot into the scene recognition sub-model to obtain an interface scene; a scene model calling sub-module, used to call the corresponding interface recognition sub-model according to the interface scene, wherein the interface recognition sub-model and the interface scene correspond one-to-one; and an element information recognition sub-module, used to input the interface screenshot into the interface recognition sub-model to obtain target element information.

[0151] In some feasible embodiments of this application, the scene model invocation submodule specifically includes: a submodel determination unit, used to determine the corresponding interface recognition submodel according to the interface scene; and a submodel invocation unit, used to send an invocation instruction to the cache dictionary so as to load the interface recognition submodel from the cache dictionary into memory.

[0152] In some feasible embodiments of this application, the UI automated testing device further includes: a sub-model unloading unit, used to unload the interface sub-model from memory to the cache dictionary in response to the completion of the use of the interface sub-model.

[0153] In some feasible embodiments of this application, the UI automated testing device further includes: a model training module, used to acquire a set of UI interface screenshot samples, train the UI interface recognition model, and obtain a trained UI interface recognition model.

[0154] In some feasible embodiments of this application, the UI interface recognition model includes a scene recognition sub-model and multiple interface recognition sub-models. The model training module specifically includes: a pre-training sub-module, used to acquire a common object image recognition sample set and pre-train each of the interface recognition sub-models to obtain a corresponding pre-trained interface recognition sub-model; a sample acquisition sub-module, used to acquire a UI interface screenshot sample set, which contains multiple subsets of UI interface screenshot samples under different scenes, and each subset of UI interface screenshot samples contains multiple UI interface screenshot samples under the corresponding scene; a separate training sub-module, used to input each UI interface screenshot sample in the subset of UI interface screenshot samples one by one into the pre-trained interface recognition sub-model under the corresponding scene for training to obtain a trained interface recognition sub-model; and a joint training sub-module, used to input each UI interface screenshot sample into the scene recognition sub-model one by one, and train the scene recognition sub-model and each of the interface recognition sub-models to obtain a trained UI interface recognition model.

[0155] In some feasible embodiments of this application, the recognition box label corresponding to each UI screenshot sample, the joint training submodule specifically includes: a sample input unit, used to input the UI screenshot samples one by one into the scene recognition sub-model to obtain the corresponding scene result; a model calling unit, used to call the corresponding interface recognition sub-model according to the scene result; an element recognition unit, used to input the UI screenshot samples into the corresponding scene recognition sub-model to obtain target element information; and a parameter update unit, used to update the parameters of the scene recognition sub-model and each of the interface recognition sub-models according to the target element information and the recognition box label, until a predetermined termination condition is reached, the training ends, and a trained UI interface recognition model is obtained.

[0156] In some feasible embodiments of this application, the UI interface recognition module specifically includes: a position determination submodule, used to input the UI interface screenshot into the UI interface recognition model to obtain the element position with the corresponding confidence level and target element type; and a coordinate transformation submodule, used to determine the target element coordinates based on the element position and the corresponding confidence level.

[0157] In some feasible embodiments of this application, determining the target element coordinates based on the element position and the corresponding confidence level specifically includes: a target element position unit, used to determine the target element position based on the element position and the corresponding confidence level; and a target element coordinate unit, used to convert the target element position into target element coordinates based on the actual screen coordinates.

[0158] In some feasible embodiments of this application, the test result generation module specifically includes: an interactive operation submodule, used to perform interactive operations on each of the target UI elements according to the target element coordinates and target element type, and obtain operation results; and a test result submodule, used to obtain the test results according to each of the operation results.

[0159] In some feasible embodiments of this application, the operation result includes a post-operation interface; the test result submodule specifically includes: an operation determination unit, used to determine whether the operation was successful based on the post-operation interface; and a test result unit, used to determine the test result based on the result of whether the operation was successful on each post-operation interface.

[0160] In this embodiment, after capturing a screenshot of the UI interface, the UI elements in the interface are identified through a UI interface recognition model, eliminating the dependence on the DOM structure. Since the UI interface recognition model is a neural network model, its testing environment is not complex, and its applicability is wider. As long as it is trained with a sufficient variety of samples, it can break free from the constraints of UI structure, source code, and platform. The self-learning mechanism eliminates the need for frequent script maintenance. In the face of complex applications and rapidly iterating development environments, the self-learning mechanism also provides sufficient flexibility and efficiency. Combined with automated interactive tools, efficient and robust automated UI testing can be achieved.

[0161] This embodiment utilizes a UI interface recognition model to automatically identify UI elements in screenshots and extract information from the target elements. Automated recognition reduces manual operation, lowers the probability of errors, and enhances the consistency and reliability of the testing process. Simultaneously, based on the precise location and recognition of UI elements, corresponding interactive operations are automatically executed for different types of UI elements, covering all UI element interactive operations to ensure no potential UI issues are missed, thus ensuring the comprehensiveness and accuracy of interaction testing. Furthermore, by automatically capturing screenshots, recognizing elements, executing interactions, and obtaining test results, the efficiency of UI automation testing is greatly improved, saving significant labor costs.

[0162] Figure 8 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown.

[0163] It should be noted that, Figure 8 The computer system of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0164] like Figure 8 As shown, the computer system includes a Central Processing Unit (CPU) 1801, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 1802 or programs loaded from storage portion 1808 into Random Access Memory (RAM) 1803, such as performing the methods described in the above embodiments. The RAM 1803 also stores various programs and data required for system operation. The CPU 1801, ROM 1802, and RAM 1803 are interconnected via a bus 1804. An Input / Output (I / O) interface 1805 is also connected to the bus 1804.

[0165] The following components are connected to I / O interface 1805: an input section 1806 including a keyboard, mouse, etc.; an output section 1807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1808 including a hard disk, etc.; and a communication section 1809 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 1809 performs communication processing via a network such as the Internet. A drive 1810 is also connected to I / O interface 1805 as needed. Removable media 1811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1810 as needed so that computer programs read from them can be installed into storage section 1808 as needed.

[0166] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1809, and / or installed from removable medium 1811. When the computer program is executed by central processing unit (CPU) 1801, it performs various functions defined in the system of this application.

[0167] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0168] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0169] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0170] In another aspect, this application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods described in the above embodiments.

[0171] This specification also provides a computer program product that stores at least one instruction, said at least one instruction being loaded and executed by the processor as described above. Figures 1-6 The method described in the illustrated embodiment can be found in the following document for a detailed execution process. Figures 1-6 The specific details of the illustrated embodiments will not be elaborated here.

[0172] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0173] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the method according to the embodiments of this application.

[0174] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

[0175] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A UI automation testing method, characterized in that, The UI automation testing methods include: Capture UI screenshots; The screenshot of the UI interface is input into the UI interface recognition model to obtain the target element information. The target element information is the element information of the target UI element, including the target element coordinates and the target element type. The target UI element is the UI element in the screenshot of the UI interface. Based on the target element coordinates and target element type, interaction tests are performed on each target UI element to obtain test results.

2. The UI automated testing method as described in claim 1, characterized in that, The UI interface recognition model includes a scene recognition sub-model and multiple interface recognition sub-models. The step of inputting the UI interface screenshot into the UI interface recognition model to obtain target element information specifically includes: The screenshot of the interface is input into the scene recognition sub-model to obtain the interface scene; Based on the interface scenario, the corresponding interface recognition sub-model is invoked, and the interface recognition sub-model corresponds one-to-one with the interface scenario. The screenshot of the interface is input into the interface recognition sub-model to obtain the target element information.

3. The UI automated testing method as described in claim 2, characterized in that, The step of calling the corresponding interface recognition sub-model according to the interface scenario specifically includes: Based on the interface scenario, determine the corresponding interface recognition sub-model; Send a call instruction to the cache dictionary to load the interface recognition sub-model from the cache dictionary into memory.

4. The UI automated testing method as described in claim 3, characterized in that, The UI automation testing method also includes: In response to the completion of use of the interface sub-model, the interface sub-model is unloaded from memory and placed into the cache dictionary.

5. The UI automated testing method as described in claim 1, characterized in that, The UI automation testing method also includes: Obtain a sample set of UI interface screenshots, train the UI interface recognition model, and obtain a trained UI interface recognition model.

6. The UI automated testing as described in claim 5, characterized in that, The UI interface recognition model includes a scene recognition sub-model and multiple interface recognition sub-models. The step of obtaining a UI interface screenshot sample set and training the UI interface recognition model to obtain a trained UI interface recognition model specifically includes: Obtain a common object image recognition sample set, and pre-train each of the interface recognition sub-models to obtain the corresponding pre-trained interface recognition sub-models. Obtain a UI screenshot sample set, which contains multiple UI screenshot sample subsets under different scenarios, and each UI screenshot sample subset contains multiple UI screenshot samples under the corresponding scenario. Each UI screenshot sample in the subset of UI screenshot samples is input into the pre-trained interface recognition sub-model in the corresponding scenario for training, and the trained interface recognition sub-model is obtained. The UI screenshot samples are input one by one into the scene recognition sub-model, and the scene recognition sub-model and each of the interface recognition sub-models are trained to obtain the trained UI interface recognition model.

7. A UI automated testing device, characterized in that, The UI automated testing device includes: The UI screenshot capture module is used to capture screenshots of the UI interface. The UI interface recognition module is used to input the UI interface screenshot into the UI interface recognition model to obtain target element information. The target element information is the element information of the target UI element, including the target element coordinates and the target element type. The target UI element is the UI element in the UI interface screenshot. The test result generation module is used to perform interaction tests on each of the target UI elements according to the target element coordinates and target element type, and obtain test results.

8. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the UI automated testing method as described in any one of claims 1 to 6.

9. An electronic device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the UI automation testing method as described in any one of claims 1 to 6.

10. A computer program product comprising one or more computer programs, characterized in that, When the one or more computer programs are executed by one or more processors, they implement the steps of the UI automated testing method according to any one of claims 1 to 6.