UI automated testing method, device, equipment, and readable storage medium
By using out-of-band image recognition models and humanoid models running in UI automation testing, UI images are identified and test operation paths are generated, resource occupation and compatibility problems caused by drivers in the prior art are solved, and efficient and reliable UI testing is achieved.
Patent Information
- Application Number
- CN202510320077.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-18
AI Technical Summary
The existing UI automation testing methods require the identification of UI elements through drivers, resulting in high resource usage, poor test performance and incompatible, especially in multi-operating system environments, which increases the testing cost and complexity.
The out-of-band image recognition model and human-like model are used to obtain target information by identifying UI images, and generate the optimal test operation path based on the target information and preset scripts to realize UI testing.
There is no need to install drivers on the device under test, which avoids resource usage and performance issues, improves the reuse efficiency of test scripts and the reliability of test results, and reduces development complexity and maintenance costs.
Smart Images

Figure CN119829468B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer application technology, and in particular to a UI automation testing method, device, equipment and readable storage medium. Background Art
[0002] At present, securities trading software usually needs to be deployed in a variety of operating system environments such as PC (such as MacOS, Windows and trusted systems, etc.) and mobile (such as Android, iOS, HarmonyOS, etc.); in this context, although the interface content of securities trading software on various platforms remains highly consistent, its UI (User Interface) automated testing faces challenges brought by differences in the underlying operating systems. Specifically, this requires that UI automated testing tools must have cross-operating system compatibility. At the same time, the writing of test case scripts also requires multi-version development for different systems, which significantly increases the development complexity of the testing tools and the workload of script writing; it is particularly worth noting that securities trading use cases are extremely diverse due to their data and are already huge in number. If traditional testing methods are used, the cost of UI automated testing will undoubtedly become a daunting challenge as operating systems become increasingly diversified.
[0003] After in-depth research, it was found that the current mainstream UI automation testing solutions usually require the installation of specific UI identification drivers on the device under test (i.e., the device that supports the operation of securities software) and the server that executes the script (such as installing the WebDriverAgent driver on iOS) to identify the UI of the application under test through the driver, and write test scripts based on business logic to simulate user operations, thereby achieving automated testing.
[0004] Although the above solution can realize the automated testing of UI, a single driver often only supports one operating system or terminal, and cannot support emerging systems (such as Xinchuang and Hongmeng). In addition, the incompatibility of drivers will also make the test scripts unable to be used across platforms, significantly increasing the cost of script writing and maintenance. In addition, some drivers will occupy too many resources during operation, and may even compete for resources with the device under test, which seriously affects the stability and efficiency of automated testing. It can be seen that implementing UI testing through drivers will not only consume the resources of the device under test and affect the performance of the test, but also cause compatibility problems. Summary of the invention
[0005] The present application provides a UI automated testing method, apparatus, device and readable storage medium, which can solve the technical problems existing in the prior art such as resource occupation, poor testing performance and incompatibility caused by performing UI testing through a driver.
[0006] In a first aspect, an embodiment of the present application provides a UI automation testing method, the UI automation testing method comprising:
[0007] Input the target UI image of the tested software into a preset image recognition model, so that the image recognition model can recognize the target UI image and obtain target information, wherein the target information includes UI element visual information, UI state information and environmental context information, wherein the UI element visual information includes UI element category, UI element position and UI element confidence, the UI state information includes page type, number of operable UI elements and UI layout information, and the environmental context information includes network status information, page change information and page refresh status;
[0008] The target information and the preset target script are input into a preset humanoid model, so that the humanoid model determines the optimal test operation path corresponding to the target user type based on the target information and the target script, and outputs the target operation instruction based on the preset keyboard and mouse mapping relationship and the optimal test operation path, so as to perform UI testing on the software under test through the target operation instruction, the target script includes operation tasks, user account information and expected operation results, and the target operation instruction includes the coordinates of the UI elements to be operated.
[0009] In combination with the first aspect, in one embodiment, the target user types include new users, old users, ordinary users, professional users, high-frequency trading users and low-frequency trading users, and the optimal test operation path includes operation sequence, click position, click frequency, operation stay time, operation end waiting time and operation error rate.
[0010] In combination with the first aspect, in one implementation, the humanoid model determines the optimal test operation path corresponding to the target user type based on the target information and the target script, including:
[0011] The humanoid model determines the target user type based on the user account information, and acquires a plurality of candidate test operation paths corresponding to the target user type according to the operation task;
[0012] The humanoid model selects an optimal test operation path from multiple candidate test operation paths through the UI element visual information, the UI state information and the environmental context information.
[0013] In combination with the first aspect, in one implementation, the humanoid model selects an optimal test operation path from multiple candidate test operation paths through the UI element visual information, the UI state information, and the environmental context information, including:
[0014] The humanoid model selects a target test operation path corresponding to the UI element visual information and the UI state information from a plurality of candidate test operation paths;
[0015] The humanoid model randomly adjusts one or more of the operation sequence, click position, click frequency, operation dwell time, operation end waiting time and operation error rate on the target test operation path through the environmental context information to generate the optimal test operation path.
[0016] In combination with the first aspect, in one implementation, after the step of performing a UI test on the software under test by using the target operation instruction, the method further includes:
[0017] Acquire an updated UI image, and transmit the updated UI image to an image recognition model, so that the image recognition model recognizes the updated UI image and obtains an actual operation result;
[0018] The actual operation result is input into the humanoid model, so that the humanoid model compares the actual operation result with the expected operation result to output the UI test result.
[0019] In combination with the first aspect, in one embodiment, before the step of inputting the target UI image of the tested software into a preset image recognition model, the step further includes:
[0020] Get the UI image to be tested of the software under test;
[0021] The UI image to be tested is subjected to data compression encoding processing to generate a target UI image.
[0022] In a second aspect, an embodiment of the present application provides a UI automation testing device, wherein the UI automation testing device comprises:
[0023] An image recognition model is used to recognize the target UI image of the software under test and obtain target information, wherein the target information includes UI element visual information, UI state information and environmental context information, wherein the UI element visual information includes UI element category, UI element position and UI element confidence, wherein the UI state information includes page type, number of operable UI elements and UI layout information, and wherein the environmental context information includes network status information, page change information and page refresh status;
[0024] A humanoid model is used to determine the optimal test operation path corresponding to the target user type based on the target information and a preset target script, and output target operation instructions based on a preset keyboard and mouse mapping relationship and the optimal test operation path, so as to perform UI testing on the software under test through the target operation instructions, wherein the target script includes operation tasks, user account information and expected operation results, and the target operation instructions include the coordinates of the UI elements to be operated.
[0025] In combination with the second aspect, in one implementation, the target user types include new users, old users, ordinary users, professional users, high-frequency trading users and low-frequency trading users, and the optimal test operation path includes operation sequence, click position, click frequency, operation stay time, operation end waiting time and operation error rate.
[0026] In conjunction with the second aspect, in one implementation, the humanoid model is specifically used for:
[0027] Determine a target user type based on the user account information, and acquire multiple candidate test operation paths corresponding to the target user type according to the operation task;
[0028] An optimal test operation path is screened out from a plurality of candidate test operation paths through the UI element visual information, the UI state information and the environmental context information.
[0029] In conjunction with the second aspect, in one implementation, the humanoid model is further used for:
[0030] Filtering out a target test operation path corresponding to the UI element visual information and the UI state information from a plurality of candidate test operation paths;
[0031] The environmental context information is used to randomly adjust one or more of the operation sequence, click position, click frequency, operation dwell time, operation end waiting time and operation error rate on the target test operation path to generate an optimal test operation path.
[0032] In combination with the second aspect, in one embodiment, the image recognition model is also used to obtain an updated UI image and identify the updated UI image to obtain an actual operation result; the actual operation result is input into the humanoid model so that the humanoid model can compare the actual operation result with the expected operation result to output a UI test result.
[0033] In conjunction with the second aspect, in one implementation, the device further includes a data processing module, which is used to:
[0034] Get the UI image to be tested of the software under test;
[0035] The UI image to be tested is subjected to data compression encoding processing to generate a target UI image.
[0036] In a third aspect, an embodiment of the present application provides a UI automation testing device, which includes a processor, a memory, and a UI automation testing program stored in the memory and executable by the processor, wherein when the UI automation testing program is executed by the processor, the steps of the aforementioned UI automation testing method are implemented.
[0037] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a UI automation test program is stored, wherein when the UI automation test program is executed by a processor, the steps of the aforementioned UI automation test method are implemented.
[0038] The beneficial effects brought by the technical solution provided in the embodiments of the present application include:
[0039] The target UI image of the tested software is identified by an out-of-band image recognition model to obtain target information including UI element category, UI element position, UI element confidence, page type, number of operable UI elements, UI layout information, network status information, page change information and page refresh status, and then the target information and the target script including operation tasks, user account information, operation targets and expected operation results are analyzed and decided by the out-of-band humanoid model to determine the optimal test operation path corresponding to the target user type, and the target operation instruction including the coordinates of the UI element to be operated is output based on the keyboard and mouse mapping relationship and the optimal test operation path, so as to realize the UI test of the tested software through the target operation instruction. It can be seen that the present application realizes UI testing in an out-of-band manner without installing a driver on the tested device, which not only solves the resource occupation, poor test performance and incompatibility problems caused by the driver, but also eliminates the need to find the UI element to be tested and its corresponding element information through the driver and modify the test script based on the element information, so as to reduce the script writing and maintenance costs and improve the reuse efficiency of the test script. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 This is a flowchart of an embodiment of the UI automation testing method of the present application;
[0041] Figure 2 This is a schematic diagram of the hardware structure of the UI automation testing equipment involved in the embodiment of the present application. DETAILED DESCRIPTION
[0042] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0043] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0044] In a first aspect, an embodiment of the present application provides a UI automation testing method.
[0045] In one embodiment, referring to Figure 1 , Figure 1 This is a flow chart of an embodiment of the UI automation testing method of this application. Figure 1 As shown, the UI automation testing method includes:
[0046] Step S10: Input the target UI image of the software under test into a preset image recognition model, so that the image recognition model can identify the target UI image and obtain target information, wherein the target information includes UI element visual information, UI state information and environmental context information, wherein the UI element visual information includes UI element category, UI element position and UI element confidence, the UI state information includes page type, number of operable UI elements and UI layout information, and the environmental context information includes network status information, page change information and page refresh status.
[0047] For example, in this embodiment, in order to solve the resource occupation, poor test performance and incompatible technical problems caused by UI testing through a driver, UI testing will be implemented through out-of-band operation, that is, an independent and out-of-band smart terminal including an image recognition model and a human-like model is constructed to implement UI testing through the image recognition model and the human-like model in the smart terminal.
[0048] It should be noted that the software under test refers to the software that needs to be tested for UI, including but not limited to securities software; the software under test will be installed on the device under test, which can be a PC or a handheld terminal, as long as it can support the operation of the software under test, and the operating system running on the device under test can also be determined according to actual needs, for example, the device under test can be a PC running a Windows operating system, a PC running a MacOS operating system, a PC running a UOS operating system, and a handheld terminal running an Android and IOS operating system. It should be understood that the main function of the device under test is to provide a basic platform for the software under test, and at the same time transmit the page image of the software under test to the smart terminal through universal interfaces such as HDMI (High Definition Multimedia Interface) or USB (Universal Serial Bus) and RFB (Remote Frame Buffer) protocol, and accept keyboard and mouse control operations from the smart terminal.
[0049] It should be understood that new message types can be added based on the RFB protocol to control the image transmission process. That is, during the execution of the UI use case, the newly added control messages are used to enable the smart terminal to obtain the latest complete page image of the software under test, stop the image data transmission, and wait until the image recognition model completes the image data analysis, and then restart the image data transmission and execute the operation instructions. This measure aims to reduce interference factors in the image data analysis process, thereby improving the overall performance of the smart terminal.
[0050] Among them, the device under test transmits the target UI image corresponding to the page of the software under test to the smart terminal running out-of-band, so that the smart terminal inputs the target UI image into the image recognition model, so as to identify the UI elements in the interface of the software under test through the image recognition model and feedback the page status, that is, the image recognition model performs feature extraction, target detection and result output on the target UI image to obtain target information including UI element visual information, UI status information and environmental context information.
[0051] Specifically, when detecting the target UI image, the image recognition model will output a set of feature information of the recognized UI elements (i.e., UI element visual information), which can help the human-like model understand the operable elements on the current interface; wherein, the UI element visual information includes but is not limited to the UI element category (Class), UI element position (Bounding Box Coordinates), and UI element confidence (Confidence Score); for example, assuming that the image recognition model detects that there is a "login button" and a "text box" on the target UI image of the software under test, it will output the following features: Element category: {Button, Text Box}, Element position: {Login button: x center = 0.5, y center = 0.8, width = 0.2, height = 0.05; Text box: x center = 0.5, y center = 0.6, width = 0.6, height = 0.1}, confidence: {login button: 0.95, text box: 0.90}; The above UI element visual information is used to tell the humanoid model that there is a high-confidence login button and text box on the current interface, and their positions and sizes have been determined, so that the humanoid model can perform operations based on this (such as clicking the login button or entering text).
[0052] UI state information is used to describe the characteristics of interface layout and element distribution, which can help the humanoid model determine the state of the current interface and make corresponding operation decisions; wherein, UI state information includes but is not limited to page type, number of operable UI elements and UI layout information. It can be understood that page type is used to characterize which page of the software under test the current interface belongs to (such as login page, transaction page, information page, etc.), number of operable UI elements refers to the number of operable elements on the interface, to help the humanoid model evaluate the space for operation selection, UI layout information refers to the distribution pattern of UI elements, such as whether there are specific layout rules (such as table arrangement, vertical / horizontal arrangement, etc.); for example, assuming that the image recognition model detects that the page currently displayed by the software under test is a "login page", it will output the following state information: page type is login page, number of operable UI elements: 2 (namely login button and user name input box), and layout information is vertical arrangement of elements, text box on top and button on bottom, so that the humanoid model can determine that this is a login page based on the UI state information, and understand the UI layout of the page, so as to perform operations in a specific order, such as entering the user name first and then clicking the login button.
[0053] Environmental context information is used to describe context information outside the current interface, such as the current network status, the system status of the application, the interface refresh status, etc., which can provide a more comprehensive basis for judgment for the humanoid model. Among them, environmental context information includes but is not limited to network status information, page change information and page refresh status. It can be understood that network status information includes information such as network delay, so as to characterize whether the current network status will affect the interface response speed through network status information, thereby helping the humanoid model to adjust the operation rhythm (such as extending the waiting time after clicking, etc.); the page refresh status is used to characterize whether the current page is in a refresh state or loading process, thereby preventing the model from performing operations when the page is not fully loaded; the page change information is used to characterize whether the interface jumps or layout changes during the operation of the software under test. For example, during the login page operation, the following environmental context information may be encountered: {Network delay: The current network condition is poor, which may cause the interface response time to be delayed by 1-2 seconds; Page refresh status: After entering the user name, the page will be refreshed briefly, and the humanoid model needs to wait for the refresh to complete before continuing to the next step; Page change information: After clicking the login button, the page jumps to the "Home" or "Error prompt" interface, and the humanoid model chooses subsequent operations based on the page changes}; It can be seen that environmental context information can help the humanoid model be more intelligent when performing operations to avoid premature or erroneous operations.
[0054] In this embodiment, in order to build a dedicated image recognition model for the securities system UI, the image recognition model is first trained and deeply learned, and various UI elements in the securities system are included in the image library; then, each UI element in the image library is accurately mapped, and a unique securities system UI recognition large model is formed on this basis; this model will play a key role in the subsequent test script command parsing and execution process to ensure accurate understanding and effective execution of the command. The following will explain the construction process of the image recognition model.
[0055] First, prepare the training data. Specifically, prepare various screenshots of the software interfaces of securities (such as login page, trading page, information page, etc.) that have been annotated, and ensure that the UI elements contained in each screenshot cover all key UI elements, such as buttons, text boxes, charts, etc.; use annotation tools (such as LabelImg tools) to annotate the UI elements in the screenshots to generate the annotation files required by the model, where each image will generate a corresponding TXT file, the specific format is as follows:
[0056] <object-class> <object-name><x_center><y_center> <width> <height>
[0057] <object-class>: The category of UI elements (such as buttons, text boxes, etc.)
[0058] <object-name>: The specific name of the UI element (such as "LoginButton", "SearchBar", etc.)
[0059] <x_center> and<y_center> : The coordinates of the center point of the UI element, normalized to the image width and height ratio
[0060] <width>and <height>: The width and height of the UI element, normalized to the image's aspect ratio.
[0061] The images are preprocessed by scaling, rotating, cropping, graying, and adjusting contrast to expand the data set, thereby improving the generalization ability of the model and enhancing the recognition effect. The preprocessed images are then adjusted to a fixed size such as 416×416 to meet the input requirements of the model, and normalized so that the pixel value range is between [0, 1] to accelerate training and improve stability. The data set is then divided into three parts: training set, validation set, and test set, and the proportions of the three are 70%, 15%, and 15%, respectively. Parameters are then set: for example, the initial learning rate is set to 0.001, and cosine annealing or step decay is used, and whenever the training set loss no longer decreases, the learning rate is reduced, for example, to 1 / 10 of the original; the weight decay is set to 0.0005 to prevent overfitting; the momentum is set to 0.9 to accelerate gradient descent.
[0062] Secondly, an image recognition model is constructed based on the YOLOv3 model (a single-stage target detection model based on convolutional neural network), that is, the image recognition model includes a Darknet-53 network and a feature pyramid network; wherein, the Darknet-53 network consists of a CBL module and a residual module, the CBL module consists of a convolutional layer, batch normalization and an activation function, and the convolution kernel size of the convolutional layer is 3×3 and 1×1, the 3×3 convolution kernel is used to extract spatial features, and the 1×1 convolution kernel is used to reduce the number of channels to reduce the amount of calculation and increase the expression ability of the network; batch normalization is applied after each convolution layer to speed up training and prevent gradient disappearance or explosion, and the Leaky ReLU function is used as the activation function so that negative input values have a small slope to avoid neuron "death"; residual connection is realized through the residual module, that is, residual connection is introduced to directly skip several convolutional layers and then add the input directly to the output to help the gradient flow to the deep layer, thereby enhancing deep feature learning.
[0063] Use the feature pyramid network for multi-scale detection and convolution layer operations, that is, make predictions from feature maps of different scales to achieve target detection; the feature pyramid network includes high-resolution layer, medium-resolution layer and low-resolution layer, the high-resolution layer (such as 52×52 feature map) is used to detect small targets (such as small icons, small buttons, etc.), the medium-resolution layer (such as 26×26 feature map) is used to detect medium-sized targets (such as large buttons, text boxes, etc.), and the low-resolution layer (such as 13×13 feature map) is used to detect large targets (such as menus, large area charts, etc.); it should be understood that the convolution layer corresponding to each scale will generate three predicted bounding boxes, each of which contains position, width, height, confidence and category probability for identifying UI elements, that is, the humanoid model can further operate the UI interface according to the bounding box and category output by the image recognition model.
[0064] The output of the image recognition model includes the following information for each grid cell: (1) bounding box coordinates, which are used to represent the coordinates (x, y) of the center point of the bounding box of the predicted object and the width w and height h of the bounding box; (2) confidence, which is used to indicate whether there is an object in the predicted bounding box, and the confidence score ranges from 0 to 1, where confidence Score = P(Object) × IoU pred,truth , P(Object) represents the probability of an object in the bounding box, IoU pred,truth It indicates the degree of overlap between the predicted bounding box and the true bounding box. It can be understood that the higher the confidence, the greater the possibility that the bounding box contains the object. (3) Class probability, that is, for each predicted bounding box, the probability value of the object belonging to each category is given. Assuming that the number of categories is C, then the length of the class probability list is C and the sum of all class probabilities is 1. Class probability ClassProbability = P (Class i / Object), where Class i represents the i-th type of target, P (Class i / Object) represents the probability that the i-th target belongs to the Object class; (4) The number of channels of the output feature map N channels , that is, N channels =B×(5 +C), B=3, that is, each grid unit predicts 3 bounding boxes, 5 means that each bounding box needs to predict 5 values (i.e. center point coordinates, width, height and confidence), C represents the number of categories. For example, if there are 10 categories of UI elements, the number of output channels is 3×(5+10)=45).
[0065] Finally, after the initial construction of the image recognition model is completed, the model is trained using the above dataset as the basic data so that the model can learn to detect relevant elements in any interface during subsequent use. Specifically, the image data with a batch size of 64 (if there is enough GPU memory, 64 images can be processed simultaneously to speed up the training process) is input into the image recognition model for training to output the predicted bounding box; for each predicted bounding box, the confidence threshold is first applied, that is, the bounding boxes below the confidence threshold are filtered out, and the overlapping bounding boxes with lower confidence are removed through non-maximum suppression (NMS) to retain the prediction results with higher confidence; after completing a training, the total loss will be calculated according to the loss function; among them, the total loss function is three losses (i.e., positioning loss Loss loc , confidence loss Loss conf , classification loss Loss class ) is the weighted sum of: Total Loss = λ coord ×Loss loc + Loss conf + Loss class ,λ coord Represents the weight of the positioning loss, which can be set to 5; then use the Adam optimizer and perform gradient updates on the model parameters based on the loss, and continue training for multiple Epochs according to the above principle (Epoch means that all batches of the entire dataset are input into the model once, that is, after each epoch, the model will perform a complete learning on the entire dataset) until the loss of the validation set converges or reaches the preset number of Epochs, and then the final image recognition model can be generated.
[0066] It should be noted that there will be multiple similar elements (such as stock code input box, stock price input box, etc.) on the same interface in the securities system software. Therefore, similar elements can be accurately identified and marked based on factors such as the location information of the elements and the characteristics of the surrounding elements, and different methods of deep learning can be adopted according to different types of interfaces, so as to form a fully functional image recognition model for securities system software interface recognition. For example, for interface elements that change in real time such as market conditions, such elements can be identified and marked by algorithms such as fuzzy matching during the image recognition process by extracting and training market characteristics; for active product interfaces that are updated and iterated quickly, learning can be carried out through factors such as requirement documents, product images, and product names, and the results can be added to the recognition model. In addition, the image recognition model can use frame difference methods to detect dynamic changes in UI elements in the interface, such as refreshing data, button status changes, etc. After receiving the operation instructions completed by the humanoid model, the page changes after the operation will be identified to provide real-time data for the humanoid model.
[0067] Step S20: Input the target information and the preset target script into the preset humanoid model, so that the humanoid model determines the optimal test operation path corresponding to the target user type based on the target information and the target script, and outputs the target operation instruction based on the preset keyboard and mouse mapping relationship and the optimal test operation path, so as to perform UI testing on the software under test through the target operation instruction, the target script includes operation tasks, user account information and expected operation results, and the target operation instruction includes the coordinates of the UI elements to be operated.
[0068] Exemplarily, in this embodiment, the target script refers to the use case data used to execute the test, which includes but is not limited to operation tasks, user account information and expected operation results; wherein, the operation task refers to the task to be completed, such as completing a login, viewing a certain page, completing a transaction, etc.; user account information includes but is not limited to user account and password; expected operation results refer to the operation results that are expected to occur.
[0069] The image recognition model converts the target information such as UI element type, UI element position, UI element confidence, page type, number of operable UI elements, UI layout information, network status information, page change information and page refresh status into structured data, and then inputs it into the human-like model in a format such as JSON. For example, for a stock code input box, its target information includes: {"name":"stock code","type":"text box","coordinates":[100, 100, 300, 50]}, where the format of the coordinates is [x-axis starting point, y-axis starting point, width, height]; at the same time, the target script is also input into the humanoid model, so that the humanoid model can determine the target user type (such as new user or old user or other type of user) based on the user account information in the target script, and the target user type is used for the humanoid model to decide to take different test operation paths; then, according to the operation task and target information, the optimal test operation path corresponding to the target user type is determined from the preset database, and the action corresponding to the optimal test operation path is executed to generate the target operation instruction according to the keyboard and mouse mapping relationship (used to convert the input of the keyboard and mouse into operations or instructions that can be recognized by other devices or software); then, according to the coordinates of the UI elements to be operated in the target operation instructions, each instruction is mapped to the actual operation of the software under test, that is, the interaction between the model and each UI element in the UI interface is realized through mouse clicks, screen sliding and other behaviors, for example, according to the coordinates of the UI elements to be operated, a corresponding position on the screen is clicked, text is input, etc., to realize the UI test of the software under test.
[0070] Since this embodiment uses coordinates to locate UI elements, there is no need to search for the UI elements to be tested and their corresponding element information through the driver program, and thus there is no need to frequently modify the test scripts according to different UI interface testing requirements, which can improve the reuse efficiency of the test scripts.
[0071] It can be seen that this embodiment implements UI testing in an out-of-band manner without installing a driver on the device under test. It not only solves the resource occupation, poor test performance and incompatibility problems caused by the driver, but also eliminates the need to use the driver to find the UI elements to be tested and their corresponding element information and modify the test script based on the element information, thereby reducing the cost of script writing and maintenance and improving the reuse efficiency of test scripts.
[0072] It should be understood that the humanoid model plays the role of simulating real user behavior in UI automation testing. It can predict user behavior by combining algorithms such as reinforcement learning, behavior modeling, and data-driven decision logic, thereby automating the operation of the UI interface. The following will explain the construction process of the humanoid model.
[0073] First, a humanoid model including a data interaction and processing module, a decision-making module, an operation module, a log recording and analysis module, and an exception handling module is constructed. The data interaction and processing module is used to process and analyze the input target script and the UI interface data passed in by the image recognition model to analyze the logic and basic process of the use case, and push UI update notifications to the image recognition model after each step of the operation is completed. The decision-making module is used to generate an operation process according to the logic and basic process of the use case, and then optimize and enrich the behavior sequence in the operation process, so as to gradually and finally adjust the operation path and further generate an operation path that is more in line with the real scene. The operation module is used to perform operations according to the type of operation (such as click, slide, etc.), position, delay and other information. The log recording and analysis module is used to record the operation to analyze the delay and feedback results of the API (Application Programming Interface). The exception handling module is used to monitor abnormal conditions during the operation process together with the image recognition model, such as unclickable buttons and page loading timeouts, and trigger corresponding response strategies.
[0074] Specifically, for the data interaction and processing module, it will receive the driving data of the use case to be executed in the current target script (for example: {use case name: normal Shenzhen Stock Exchange buy transaction example; counter data source: 1; buy type: limit buy; whether to verify: yes; expected result: the order has been placed, the order number is Xxx}), and generate the basic operation plan of the "normal Shenzhen Stock Exchange limit buy transaction example" use case according to the driving data (for example: the basic process is to find the account and stock code, price, etc. that meet the conditions → start the APP → enter the login page to log in → enter the normal buy interface → enter the stock code → enter the stock price → enter the number of stocks to buy → place an order → verify the result data).
[0075] For the decision-making module, it will optimize and enrich the basic operation plan generated in the data interaction and processing module: load the preset operation plan. If the preset operation plan specifies the test target and user type, the priority and preference of the operation will be determined according to the preset type; if the preset operation plan is not specified, generate test cases for corresponding scenarios for each test target and user type (such as novice users, ordinary users, active traders), and distribute all test cases to the corresponding devices for execution when there are sufficient devices under test. When there are insufficient devices, select one or several test cases as the result according to the preset rules and priorities.
[0076] The operation module will add human-like elements and plan paths on the basis of the basic operation plan. For example, for novice users, in the operation process such as login, randomness is introduced when determining the next operation, and different operation paths are selected or parameters (such as waiting time, click position, etc.) are slightly adjusted to simulate the operation frequency of novice users; for example, before entering the buy page, first enter the home page -> information page -> market page, and add sliding, clicking, returning and other operations to each page according to the characteristics of novice users; among them, on the market page, enter the buy page by searching the stock code, enter the stock code to use the code prompt function, and enter the purchase quantity and price through symbols such as + / -; after enriching the operation plan, the decision module receives the UI element information provided by the image recognition module (such as buttons, text boxes, labels, etc. of the current page), and determines the type of the current page (such as login page, home page, trading page, etc.) according to the UI elements recognized by the image recognition module; this step can be achieved by comparing specific identifying elements on the page (for example, there may be a navigation bar on the home page, and buy and sell buttons on the trading page).
[0077] The decision module will search for feasible operation paths through the decision tree under the current page status and operation plan. For example, on the "Login" page, the branches of the decision tree may be "Enter account number", "Enter password", "Slide", "Hide keyboard" and "Click to log in", etc. If there are multiple optional paths on the decision tree, the decision module will select the priority path according to the user type and page status. For example, if it is a novice user, the browsing information path is preferred on the homepage, while active traders prefer the trading path; after determining the path, the decision module will generate the corresponding operation instructions (such as "Enter user name", "Click the login button", etc.) and pass them to the operation module to perform the operation. In addition, the decision module can also record the loading time of key pages, API response time, page rendering time, etc., to provide data for analyzing and optimizing the performance of the software, and can also promptly capture and report exceptions that occur during the operation process, such as page loading errors, button unresponsiveness, etc.
[0078] It should be noted that if a page element exception makes it impossible to execute a certain operation path (for example, a network problem causes the login page to fail to load and the login operation cannot be performed), the exception handling module will be called to perform exception handling; if the exception handling rule is to report an error or take a screenshot, an error will be reported directly; if the exception handling rule is to repair it first, the exception type will be determined, an attempt will be made to repair it, and the use case will continue to be executed after the repair.
[0079] After the operation is completed, the decision module will obtain the updated state information of the page from the image recognition module to determine whether the actual operation result is consistent with the expected operation result (for example, after "clicking to log in", check whether it jumps to the homepage), that is, by comparing the expected state and the actual state to confirm whether the operation has caused the expected result. If the operation is successful, continue to the next operation; if the operation fails, the exception handling process is triggered. Among them, the decision module can use the Markov chain algorithm to dynamically adjust the operation path. For example, if a certain operation path fails multiple times, the decision module can reduce its execution priority or avoid the operation path in subsequent path selection.
[0080] It is understandable that in the training of humanoid models, the characteristics of input data are the basis for the model to make decisions. Therefore, this embodiment will use real user information, transaction data, embedded data, etc. for analysis and training to generate a set of models for securities users. Therefore, this embodiment will collect and annotate data, wherein the collected data set contains the operating habits and behavior patterns of different users (obtained through the collected embedded data and logs of online users, specifically including page dwell time, number of clicks, page access frequency, operation time interval, number of times a specific function is used, etc.), and contains the operation steps of various test scenarios (the basic operation steps of each test scenario have been preset and stored in the database, such as ordinary transaction buying includes steps such as starting the APP→login→order→verification, and the basic steps do not contain humanoid operation factors.
[0081] The collected data sets are classified to obtain UI element visual information, UI state information, user behavior characteristics and environmental context information; wherein, UI element visual information includes UI element category, UI element position and UI element confidence, UI state information includes page type, number of operable UI elements and UI layout information, and environmental context information includes network status information, page change information and page refresh status; user behavior characteristics describe the operation habits and behavior patterns of historical users on similar interfaces. These characteristics can help human-like models imitate the operation sequence, frequency, preferences, etc. of real users, including but not limited to operation sequence, dwell time, click frequency and operation error rate, etc. The operation sequence refers to the historical operation sequence performed by the user on a certain interface (for example, enter the user name first, then click the login button); dwell time refers to the user's dwell time on a certain UI element (for example, the user stays in the user name input box for 3 seconds), and click frequency refers to the frequency of a certain UI element being clicked, indicating the user's operation priority for this element; the operation error rate refers to the probability of errors made by the user in the past operation process (such as the probability of wrong clicks).
[0082] Suppose the system records a user's historical operation behavior on the login page: {Operation sequence: the user usually enters the user name first and then clicks the login button; Dwell time: the user stays in the user name input box for an average of 3 seconds; Click frequency: the login button is clicked most frequently with a frequency of 90%, and the user occasionally clicks other areas with a frequency of 10%; Operation error rate: when the user clicks other invalid areas, the error rate is 10%}. The above features can help the human-like model select an operation sequence that is more in line with the actual user behavior based on historical behavior, and simulate the user's operation rhythm through dwell time and click frequency.
[0083] It can be understood that using the above data as input to the humanoid model can help the humanoid model understand the current interface layout, executable operations, and possible user behaviors, so as to select the most appropriate operation.
[0084] Then, the humanoid model is trained through the decision tree classification algorithm so that it can classify different user types (such as novice users, ordinary users, active traders, etc.) based on user operation habits, that is, a decision tree is generated according to the characteristics of user operation behavior (such as click frequency, page dwell time, etc.) to classify users into different types; among them, the decision tree can judge the user type through a series of conditional branches, for example: if (page dwell time>300 seconds) and (button click frequency<2 times / minute): user type = novice user; elif (page dwell time<300 seconds) and (button click frequency>5 times / minute): user type = active trader; else: user type = ordinary user.
[0085] After determining the user type, the behavior data generator will generate the operation path for this type of user, that is, different operation paths will be generated for different types of users to ensure that the humanoid model covers different user habits; for example, the user's operation preferences and priorities are predefined according to the user type and test scenario, so that the behavior data generator generates a random sequence based on these priorities to simulate user behavior; and then the reinforcement learning algorithm is used to optimize the operation sequence to make it more in line with actual operation habits. Among them, the reinforcement learning algorithm gradually optimizes the operation path through repeated trials to reward successful operation sequences and punish failed operation sequences, thereby learning the best operation path.
[0086] This embodiment will preferably use Q-learning as a reinforcement learning algorithm for model training. The core of the training is to learn the optimal actions to be performed in different states by constantly interacting with the environment. For example, the humanoid model simulates the behavior of a user when entering the homepage of a securities software after logging in through Q-learning training. This scenario involves two different behavior patterns of new users and old users. The model needs to select different operation paths according to the user type (for example, new users browse information, while old users check the market). After reinforcement learning, the model can gradually learn the behavior patterns of these two users and optimize the operation strategy. Among them, the formula of Q-learning is: , Indicates in status s Perform the following operations The reward value, Indicates that in the next state All possible operations The maximum reward value, Indicates immediate reward, Represents the learning rate, which is used to control the fusion ratio of the new and old reward values. γ Represents a discount factor that balances the impact of current and future rewards.
[0087] It should be noted that after generating the operation sequence, the humanoid model needs to determine the specific operation logic in different test scenarios, that is, the operation flow chart and business logic. The training process is: 1) Use the business logic description in the data annotation phase to define the transition matrix in each test scenario; 2) Use the Markov chain to describe the transition probability between pages. Specifically, for different test scenarios, set the state transition matrix to simulate the user's flow between pages; where the state transition probability matrix is defined: , Indicates from the scene To the scene 3) During the operation, add logical judgment conditions to ensure that certain specific operations follow business rules, for example, blocking transactions when the balance is insufficient, to ensure compliance with business scenarios.
[0088] It should be understood that when the humanoid model performs each operation, it will record the operation results, notify the image recognition model to identify and feedback the latest page elements and status, so as to adjust the operation path and update the Q value through the reinforcement learning algorithm. Among them, the detailed information of each operation is recorded through the operation log, such as the page ID, operation type, timestamp, operation result, etc., and detect whether the operation is successful. If the expected effect is not achieved, it will be backtracked and the Q value will be updated.
[0089] The following describes the reinforcement learning process of Q-learning in combination with a specific scenario. For example, the scenario is: after a user successfully logs in, he enters the homepage of the securities APP; new users usually choose to browse information first; old users usually choose to check the market first. Among them, Q-learning will construct the current state representation through input data such as UI element visual information, UI state information, user behavior characteristics, and environmental context information s ,It can be seen that the state representation includes not only ,visual features, but also interface layout, user habits and ,environmental context.
[0090] It is understandable that for the above scenario, the status representation needs to include: {user type information: new user or old user; UI page status: UI elements on the homepage, such as "information button", "market button", etc.; historical operation information: whether the user has browsed information or viewed market conditions before}. For example, the status s 1: Home page status (user is a new user, home page UI elements: information button, market button); status s 2: Home page status (the user is an old user, home page UI elements: information button, market button).
[0091] A Q value table as shown in Table 1 is initialized to store the Q values of various actions performed in different states.
[0092] Table 1 Q value table
[0093]
[0094] For the action space, in the homepage state, the actions that the humanoid model can perform include: a 1: Click the information button to enter the information page; action a 2: Click the market button to enter the market page; Action a 3: Click other buttons (such as personal account, transaction page, etc.).
[0095] Starting from the initial state, the model ε -greedy strategy selects actions, and as training progresses, the model will gradually reduce exploration and choose more optimal actions. For example, in state s 1, the model can choose to perform the action first a 1: Click the information button; while in the status s 2, the model can choose to perform the action first a 2: Click the Quote button.
[0096] It should be understood that after selecting an action, the model executes the selected action and interacts with the software interface to observe the operation results and rewards obtained. Among them, the model should make the best choice based on the user type, so the reward function can be designed according to the following rules: (1) Positive reward: if the model's operation is consistent with the user's expected behavior (such as new users browsing information and old users checking the market), a positive reward is given; (2) Negative reward: if the model's operation is inconsistent with the user's expectations, or an irrelevant operation is selected (such as a new user directly entering the market page), a negative reward is given. For example, a new user in state s 1 Click the information button a 1, the reward is +10; old users in the state s 2. Click the Quote button a 2, the reward is +10; the new user is in state s 1 Click the Quote button a 2, the reward is -5; old users in the state s 2. Click the Information button a 1, the reward is -5. Assume that the new user is in state s 1 click on the information button a 1, and successfully entered the information page and received a +10 reward, the Q value table will be updated to: ; It can be seen that the Q value table Q ( s 1 ,a 1 ) will gradually increase to encourage the model to s 1. First click the information button a 1.
[0097] Repeating the above training, i.e. the model repeatedly performs actions, observes rewards, and updates Q values to gradually optimize the operation strategy, over time, the model will learn to distinguish between various users (such as new users and old users) and perform the best operation based on the user type. For example, after a new user logs in, the model will give priority to the information button, while after an old user logs in, the model will give priority to checking the market.
[0098] Through the above steps, the humanoid model can achieve accurate user behavior simulation, covering user type identification, operation path generation, test scenario selection and feedback optimization, and the model gradually adjusts the behavior path during the training process to make the automated test closer to the actual user operation. After the training of the humanoid model is completed, the trained operation path will be stored in the database for the humanoid model to select the test operation path in the database for UI testing.
[0099] It is understandable that after the image recognition model and the humanoid model are trained, the operation of the entire link will involve multiple links: the input of use case data, the generation of use cases, the execution of humanoid actions, and the feedback and evaluation of results.
[0100] Among them, for the input of use case data, the use case data is the starting point of the entire process, which defines the tasks that the humanoid model needs to complete in the securities software interface; for example, the tasks can be login operations, browsing information, checking market conditions, buying stocks, etc. Specifically, the task is to complete the login operation of the securities APP, the user type is an old user, and the expected interface state is the login page, which includes a user name input box, a password input box, and a login button; the operation goal is to enter content in the user name input box, enter the password in the password input box, and click the login button.
[0101] The image recognition model is responsible for identifying all operable elements from the current UI interface and passing these elements to the humanoid model. The purpose of this step is to convert the UI elements on the interface into a "state representation" that the humanoid model can understand and operate. The operating rules of the image recognition model are as follows: (1) Input: a screenshot of the current interface is used as input (such as a login page); (2) Output: the detected UI elements are used as output, which includes the category (button, text box, etc.), location (bounding box coordinates), confidence level and other information of each element; (3) Result transfer: the identified UI elements are transferred to the humanoid model for subsequent operations. Specifically, for example, the input is a screenshot of the login page of a securities software; the output includes: 1) User name input box: location x center = 0.5, y center = 0.6, width = 0.6, height = 0.1, confidence = 0.95; 2) Password input box: position x center = 0.5, y center = 0.5, width = 0.5, height = 0.1, confidence = 0.92; 3) Login button: location x center =0.5, y center = 0.7, width = 0.3, height = 0.1, confidence = 0.98; the above recognition information provides the state representation of the UI interface, and the subsequent humanoid model will perform corresponding operations based on these elements.
[0102] The humanoid model will make decisions and perform operations based on the UI elements provided by the image recognition model and the current task use case. That is, the humanoid model selects the optimal action in different states through reinforcement learning. Its goal is to complete the tasks in the use case, such as completing login, browsing information, etc. The operation logic of the humanoid model is as follows: (1) Input, which includes: 1) UI state information: UI element information transmitted by the image recognition model; 2) Use case data: including operation tasks, user types, and expected operation results.
[0103] The humanoid model combines UI elements and use case data to generate the current state representation, and then selects actions: 1) According to the decision logic of Q-learning, the humanoid model will select the optimal operation in the current state. The action selection of the model will be different for different user types. For example, for the new user login page, the actions include: Action 1-enter the user name (stay for a long time, such as 3 seconds), Action 2-enter the password, Action 3-click the login button, enter the homepage after successful login, Action 4-click the "Information" button to start browsing information; 2) Use ε -greedy strategy for action selection, that is, exploring new operation paths in some cases and choosing the known optimal path in most cases.
[0104] Finally, the humanoid model implements action execution: (1) The humanoid model maps the selected action to the actual operation, such as clicking a certain position on the screen, entering text, etc.; (2) The model interacts with the UI interface through mouse clicks, screen sliding, etc. Each time the humanoid model executes an action, the system will feedback the result, which includes whether the interface has the expected change, whether the operation is successful, whether the login is completed, etc.
[0105] In general, the data processing flow of the humanoid model includes but is not limited to: judging the user type and page status based on the use case data and image recognition results and building a decision tree to guide the operation of the humanoid model in different scenarios; defining the operation preferences and operation frequencies of each user type to form different operation flows, and then generating behavior sequences through predefined behavior data generators and based on the proportion of user types and operation distribution, and using reinforcement learning algorithms to optimize behavior sequences to gradually adjust the operation path to make the test more in line with real scenarios; defining the deviation range of the click area and the distribution of waiting time for different users, for example, novice users may randomly switch between browsing information and checking accounts, while ordinary users are more inclined to transaction behaviors, and mapping basic operations such as clicking, sliding, inputting, waiting, etc. to the software under test through keyboard and mouse mapping; recording detailed information of each operation (such as time, page, behavior, results, etc.) in the database through the automated logging function, so as to be monitored through performance monitoring plug-ins (such as Dynatrace, New Relic, etc.) captures performance indicators such as page loading time and API response time and integrates them into log records; after the humanoid model completes each operation and causes the page to refresh, it notifies the image recognition model to detect changes in UI elements and feedback the status of the current page to guide the humanoid model to complete the next operation.
[0106] Furthermore, in one embodiment, the target user types include new users, old users, ordinary users, professional users, high-frequency trading users and low-frequency trading users, and the optimal test operation path includes operation sequence, click position, click frequency, operation dwell time, operation end waiting time and operation error rate.
[0107] For example, it should be understood that the commonly used UI automation testing tools currently generally execute instructions continuously or add fixed delays, and operations such as text input are generally in the form of directly inputting strings through APIs, etc., and such operations are actually not in line with normal human operating habits during securities trading. In order to conduct comprehensive testing, the humanoid model in this embodiment can simulate different types of user scenarios to generate different test operation paths according to different user types. Among them, the target user types include but are not limited to new users, old users, ordinary users, professional users, high-frequency trading users and low-frequency trading users. The specific user type can also be determined according to actual needs and is not limited here.
[0108] For example, for a common buy order use case, the humanoid model can generate three test operation paths corresponding to different scenarios for novice users, ordinary users, and high-frequency trading users. For example, novice users generally have the following behavioral habits in the buy order use case: they may find stocks through recommendations on the homepage, consult and browse before placing an order, and have slow operation speed. Ordinary users have the following behavioral habits before placing an order: they may check account information, find stocks through searches, check stock consultation and market K-lines, and have relatively accurate operations. High-frequency trading users usually have the following behavioral habits: they generally view stocks from the watchlist and frequently place and cancel orders, quickly switch pages to view stock information, zoom in and out of K-lines, switch K-lines, and have short operation pauses.
[0109] It should be understood that the specific differences between different test operation paths mainly lie in the differences in operation sequence, click position, click frequency, operation dwell time, operation end waiting time and operation error rate. Among them, the operation sequence refers to the order of operations performed by the user on the page; the operation dwell time refers to the length of time the user stays on a certain UI element; the click frequency refers to the frequency with which a certain UI element is clicked, indicating the user's operation priority for this element; the operation end waiting time refers to the length of time required to wait after operating the UI element; the operation error rate refers to the probability of the user making an error in the past operation process. It should be noted that the specific parameters included in the test operation path can also be determined according to actual needs and are not limited here.
[0110] Furthermore, in one embodiment, the humanoid model determines the optimal test operation path corresponding to the target user type based on the target information and the target script, including:
[0111] The humanoid model determines the target user type based on the user account information, and acquires a plurality of candidate test operation paths corresponding to the target user type according to the operation task;
[0112] The humanoid model selects an optimal test operation path from multiple candidate test operation paths through the UI element visual information, the UI state information and the environmental context information.
[0113] Exemplarily, it should be understood that after the training is completed, the humanoid model has recorded the type of each user and the operation path corresponding to each type of user, and stored the mapping relationship between the user account and the user type and the mapping relationship between the user type and the operation path in the database. Therefore, in this embodiment, the humanoid model can determine the target user type corresponding to the user account from the database through the user account information, and filter out multiple operation paths corresponding to the operation task from all operation paths corresponding to the target user type, and use the multiple operation paths as candidate test operation paths; then perform path matching through UI element visual information, UI state information and environmental context information to filter out the optimal test operation path from multiple candidate test operation paths.
[0114] Furthermore, in one embodiment, the humanoid model selects the optimal test operation path from multiple candidate test operation paths through the UI element visual information, the UI state information and the environmental context information, including:
[0115] The humanoid model selects a target test operation path corresponding to the UI element visual information and the UI state information from a plurality of candidate test operation paths;
[0116] The humanoid model randomly adjusts one or more of the operation sequence, click position, click frequency, operation dwell time, operation end waiting time and operation error rate on the target test operation path through the environmental context information to generate the optimal test operation path.
[0117] Exemplarily, it should be understood that operating the UI through driving cannot completely simulate people's operating habits, so that the performance and authenticity of the securities software UI automation test are affected to a certain extent. In this embodiment, the alternative test operation path is first scene-matched through the UI element visual information and UI state information to screen out the target test operation path corresponding to the UI element visual information and UI state information; then the operation sequence, click position, click frequency, operation stay time, operation end waiting time and operation error rate on the target test operation path are randomly adjusted through environmental context information, user type, etc. to generate the optimal test operation path. For example, if there is a network delay, the operation stay time can be randomly extended in the target test operation path or actions such as click, long press, double click, slide and drag can be randomly added. Some operations such as false touch, over-slide, and secondary click can also be added to the target test operation path corresponding to elderly users and new users.
[0118] Furthermore, in one embodiment, after the step of performing UI testing on the software under test through the target operation instruction, the method further includes:
[0119] Acquire an updated UI image, and transmit the updated UI image to an image recognition model, so that the image recognition model recognizes the updated UI image and obtains an actual operation result;
[0120] The actual operation result is input into the humanoid model, so that the humanoid model compares the actual operation result with the expected operation result to output the UI test result.
[0121] Exemplarily, in this embodiment, after the humanoid model makes the software under test execute the target operation instruction, it will perform process correction and recovery through interactive collaboration with the image recognition model to meet the usage habits of real users in the software, thereby simulating user operations more realistically while ensuring the execution of the effective process of the use case. Specifically, the image recognition model first obtains the updated UI image and recognizes the updated UI image to obtain the actual operation result; then transmits the actual operation result to the humanoid model, so that the humanoid model can output the UI test result based on the comparison between the actual operation result and the expected operation result, so as to confirm whether the operation is successful or abnormal through the UI test result, and record the operation result; then the image recognition model is made to detect the new interface of the software under test again and transmit the detection result to the humanoid model, and the humanoid model makes a decision and executes the next operation according to the operation result fed back by the image recognition model, and continues to interact with the image recognition model until the test sequence is completed.
[0122] Furthermore, in one embodiment, before the step of inputting the target UI image of the software under test into a preset image recognition model, the method further includes:
[0123] Get the UI image to be tested of the software under test;
[0124] The UI image to be tested is subjected to data compression encoding processing to generate a target UI image.
[0125] Exemplarily, in this embodiment, the tester can monitor the interface and test results of the software under test by remotely accessing the smart terminal. This process usually includes the transmission of multiple data such as video data, image data, text data, and keyboard and mouse control signal lights. In order to save bandwidth and improve transmission efficiency, this embodiment will compress and encode the data according to the type of data transmitted. Specifically, before transmitting the UI image data to the image recognition model, the acquired UI image to be tested of the software under test is first subjected to data compression and encoding processing to generate a compressed target UI image; the compressed target UI image is then transmitted to the image recognition model for UI recognition, thereby saving bandwidth and improving transmission efficiency. It should be noted that the UI image to be tested can be a screenshot or can be generated by shooting with a camera device. The specific method is not limited here.
[0126] It should be understood that this embodiment supports CopyRect, RRE, Hextile, ZRLE, JPEG, MPEG and other encoding formats, and the application scenarios of these encoding formats are different. In the process of data transmission, a suitable encoding format can be selected according to the data format and the transmission scenario. Among them, the selection of compression encoding method can support two methods: one is through parameter control, that is, the compression encoding method is set through parameter configuration before executing the task, so that a fixed compression encoding format is used in the entire task execution process; and the other is to autonomously select the optimal encoding method based on factors such as use case scenarios and network types, such as using JPEG encoding format when transmitting screenshots, using RRE encoding format when transmitting real-time operation interface images of trading use cases, and using ZRLE encoding format when transmitting cross-sectional images of market-related use cases; in addition, the optimal encoding format can also be selected based on whether the remote monitoring host and the device under test are in the same LAN environment (i.e., intranet environment) or non-LAN environment (i.e., extranet environment).
[0127] In addition, since the remote monitoring host may be in an external network environment and access the smart terminal through VPN (Virtual Private Network) and other forms, it is easy to cause the intranet to be invaded if the data is not encrypted, which is not allowed in the securities industry; therefore, in order to achieve secure data transmission, the smart terminal in this embodiment will perform encryption processing when transmitting data. For example, when the remote monitoring host accesses the smart terminal in the intranet environment through VPN, the smart terminal will determine the network type based on the IP. If it is an external network access, the data will be encrypted with advanced encryption such as AES (Advanced Encryption Standard), using elements such as OA (Office Automation) as key information; and if it is an intranet access, a simple encryption algorithm will be used to improve data transmission efficiency. Similarly, the encryption method also supports two forms: parameter control and autonomous selection based on IP.
[0128] It should be noted that channel selection can be performed during data transmission, that is, the intelligent terminal can independently determine the environment type based on the visitor's IP information. If it is in a controlled environment, transmission efficiency is prioritized in channel selection; if it is in an uncontrolled environment, transmission security is prioritized, so that reasonable judgments and choices can be made in terms of data compression and encryption methods, thereby balancing the performance and security of the system.
[0129] It is understandable that the UI test of the securities system aims to ensure the reliability and stability of the system by simulating the actual transaction process of users. Although the environments in which customers use devices to run securities systems are diverse, the interface designs of these systems on different operating systems are usually highly consistent; therefore, this embodiment cleverly utilizes this feature, performs real-time sampling and analysis of images through an external edge intelligent terminal, and simulates manual operations through a humanoid model after deep learning, that is, adding artificial features to the test, making the test scenario more realistic and improving the reliability of the test results; finally, the keyboard and mouse operations of the intelligent terminal are mapped to the device under test, that is, the operation instructions are accurately converted into keyboard and mouse operations on the screen of the device under test, so as to achieve the UI automation test of securities system software running on various operating systems under the conditions of zero occupation of the resources of the device under test, no mutual influence, and more realistic humanoid operation.
[0130] In summary, the core of this embodiment is to deploy the test script in an independent, out-of-band smart terminal without installing any driver or execution script on the device under test, which effectively reduces the burden of the test tool on system resources. It not only avoids problems such as incompatibility, lack of support, and conflict with the system under test due to resource competition, but also enhances the overall performance and execution efficiency of the test system, and can collect more realistic performance data. It also greatly reduces the development complexity of UI automation testing tools, while significantly reducing the number of test scripts written and the cost investment, and improving the test efficiency, flexibility, and reuse efficiency of test scripts. In addition, the environment is judged by parameter control and autonomous identification, so as to more accurately select the transmission channel, encryption method, and compression encoding, so that the system can more intelligently improve performance, security, and stability. The smart terminal provided in this embodiment is highly versatile and can be deployed in various operating systems to conduct UI automation testing of securities systems.
[0131] In a second aspect, an embodiment of the present application also provides a UI automation testing device.
[0132] In one embodiment, the UI automation testing device includes:
[0133] An image recognition model is used to recognize the target UI image of the software under test and obtain target information, wherein the target information includes UI element visual information, UI state information and environmental context information, wherein the UI element visual information includes UI element category, UI element position and UI element confidence, wherein the UI state information includes page type, number of operable UI elements and UI layout information, and wherein the environmental context information includes network status information, page change information and page refresh status;
[0134] A humanoid model is used to determine the optimal test operation path corresponding to the target user type based on the target information and a preset target script, and output target operation instructions based on a preset keyboard and mouse mapping relationship and the optimal test operation path, so as to perform UI testing on the software under test through the target operation instructions, wherein the target script includes operation tasks, user account information and expected operation results, and the target operation instructions include the coordinates of the UI elements to be operated.
[0135] Furthermore, in one embodiment, the target user types include new users, old users, ordinary users, professional users, high-frequency trading users and low-frequency trading users, and the optimal test operation path includes operation sequence, click position, click frequency, operation dwell time, operation end waiting time and operation error rate.
[0136] Furthermore, in one embodiment, the humanoid model is specifically used for:
[0137] Determine a target user type based on the user account information, and acquire multiple candidate test operation paths corresponding to the target user type according to the operation task;
[0138] An optimal test operation path is screened out from a plurality of candidate test operation paths through the UI element visual information, the UI state information and the environmental context information.
[0139] Furthermore, in one embodiment, the humanoid model is also used for:
[0140] Filtering out a target test operation path corresponding to the UI element visual information and the UI state information from a plurality of candidate test operation paths;
[0141] The environmental context information is used to randomly adjust one or more of the operation sequence, click position, click frequency, operation dwell time, operation end waiting time and operation error rate on the target test operation path to generate an optimal test operation path.
[0142] Furthermore, in one embodiment, the image recognition model is also used to obtain an updated UI image and identify the updated UI image to obtain an actual operation result; the actual operation result is input into the humanoid model so that the humanoid model can compare the actual operation result with the expected operation result to output a UI test result.
[0143] Furthermore, in one embodiment, the device further includes a data processing module, which is used to:
[0144] Get the UI image to be tested of the software under test;
[0145] The UI image to be tested is subjected to data compression encoding processing to generate a target UI image.
[0146] Among them, the functional implementation of each part in the above-mentioned UI automation testing device corresponds to each step in the above-mentioned UI automation testing method embodiment, and its functions and implementation processes are no longer repeated here.
[0147] In a third aspect, an embodiment of the present application provides a UI automation testing device, which may be a device with data processing functions, such as a personal computer (PC), a laptop computer, or a server.
[0148] Reference Figure 2 , Figure 2 The hardware structure diagram of the UI automation test device involved in the embodiment of the present application is shown in FIG. In the embodiment of the present application, the UI automation test device may include a processor, a memory, a communication interface, and a communication bus.
[0149] The communication bus may be of any type and is used to interconnect the processor, the memory, and the communication interface.
[0150] Communication interfaces include input / output (I / O) interfaces, physical interfaces, and logical interfaces, which are used to interconnect devices within the UI automated test equipment, as well as interfaces used to interconnect UI automated test equipment with other devices (such as other computing devices or user devices). Physical interfaces can be Ethernet interfaces, optical fiber interfaces, ATM interfaces, etc.; user devices can be display screens (Display), keyboards (Keyboard), etc.
[0151] The memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical storage, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.
[0152] The processor may be a general-purpose processor, which may call the UI automation test program stored in the memory and execute the UI automation test method provided in the embodiment of the present application. For example, the general-purpose processor may be a central processing unit (CPU). The method executed when the UI automation test program is called may refer to the various embodiments of the UI automation test method of the present application, which will not be described in detail here.
[0153] Those skilled in the art will understand that Figure 2 The hardware structure shown in the figure does not constitute a limitation on the present application, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0154] In a fourth aspect, an embodiment of the present application also provides a computer-readable storage medium.
[0155] The readable storage medium of the present application stores a UI automation test program, wherein when the UI automation test program is executed by a processor, the steps of the UI automation test method described above are implemented.
[0156] Among them, the method implemented when the UI automation test program is executed can refer to the various embodiments of the UI automation test method of the present application, and will not be repeated here.
[0157] It should be noted that the serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0158] The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices. The terms "first", "second" and "third" are used to distinguish different objects, etc., and do not represent a sequence, nor do they limit "first", "second" and "third" to different types.
[0159] In the description of the embodiments of the present application, "exemplary", "for example" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary", "for example" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary", "for example" or "for example" is intended to present related concepts in a specific way.
[0160] In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; the "and / or" in the text is merely a description of the association relationship of associated objects, indicating that three relationships may exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, "multiple" refers to two or more than two.
[0161] In some processes described in the embodiments of the present application, multiple operations or steps that appear in a specific order are included, but it should be understood that these operations or steps may not be executed in the order in which they appear in the embodiments of the present application or in parallel, and the sequence number of the operation is only used to distinguish the different operations, and the sequence number itself does not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed in sequence or in parallel, and these operations or steps may be combined.
[0162] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, disk, CD) as described above, and includes a number of instructions for a terminal device to execute the methods described in each embodiment of the present application.
[0163] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.< / height> < / width> < / height> < / width> < / object-class>
Claims
1. A UI automation testing method, characterized in that: The UI automation testing method comprises: Input the target UI image of the tested software into a preset image recognition model, so that the image recognition model can recognize the target UI image and obtain target information, wherein the target information includes UI element visual information, UI state information and environmental context information, wherein the UI element visual information includes UI element category, UI element position and UI element confidence, the UI state information includes page type, number of operable UI elements and UI layout information, and the environmental context information includes network status information, page change information and page refresh status; The target information and the preset target script are input into the preset humanoid model, so that the humanoid model can determine the optimal test operation path corresponding to the target user type based on the target information and the target script, and output the target operation instruction based on the preset keyboard-mouse mapping relationship and the optimal test operation path, so as to perform UI testing on the software under test through the target operation instruction, wherein the target script includes the operation task, user account information and expected operation result, and the target operation instruction includes the coordinates of the UI element to be operated; The target user types include new users, old users, ordinary users, professional users, high-frequency trading users and low-frequency trading users, and the optimal test operation path includes operation sequence, click position, click frequency, operation stay time, operation end waiting time and operation error rate; The humanoid model determines the optimal test operation path corresponding to the target user type based on the target information and the target script, including: The humanoid model determines the target user type based on the user account information, and acquires a plurality of candidate test operation paths corresponding to the target user type according to the operation task; The humanoid model selects an optimal test operation path from multiple candidate test operation paths through the UI element visual information, the UI state information and the environmental context information.
2. The UI automation testing method according to claim 1, characterized in that: The humanoid model selects an optimal test operation path from a plurality of candidate test operation paths through the UI element visual information, the UI state information and the environmental context information, including: The humanoid model selects a target test operation path corresponding to the UI element visual information and the UI state information from a plurality of candidate test operation paths; The humanoid model randomly adjusts one or more of the operation sequence, click position, click frequency, operation dwell time, operation end waiting time and operation error rate on the target test operation path through the environmental context information to generate the optimal test operation path.
3. The UI automation testing method according to claim 1, characterized in that: After the step of performing UI testing on the software under test through the target operation instruction, the method further includes: Acquire an updated UI image, and transmit the updated UI image to an image recognition model, so that the image recognition model recognizes the updated UI image and obtains an actual operation result; The actual operation result is input into the humanoid model, so that the humanoid model compares the actual operation result with the expected operation result to output the UI test result.
4. The UI automation testing method according to claim 1, wherein: Before the step of inputting the target UI image of the tested software into the preset image recognition model, the method further includes: Get the UI image to be tested of the software under test; The UI image to be tested is subjected to data compression encoding processing to generate a target UI image.
5. A UI automated testing device, characterized in that: The UI automation testing device comprises: An image recognition model is used to recognize the target UI image of the software under test and obtain target information, wherein the target information includes UI element visual information, UI state information and environmental context information, wherein the UI element visual information includes UI element category, UI element position and UI element confidence, wherein the UI state information includes page type, number of operable UI elements and UI layout information, and wherein the environmental context information includes network status information, page change information and page refresh status; A humanoid model is used to determine the optimal test operation path corresponding to the target user type based on the target information and the preset target script, and output the target operation instruction based on the preset keyboard-mouse mapping relationship and the optimal test operation path, so as to perform UI testing on the tested software through the target operation instruction, wherein the target script includes the operation task, user account information and expected operation results, and the target operation instruction includes the coordinates of the UI element to be operated. The target user types include new users, old users, ordinary users, professional users, high-frequency trading users and low-frequency trading users, and the optimal test operation path includes operation sequence, click position, click frequency, operation stay time, operation end waiting time and operation error rate; The humanoid model is specifically used for: Determine a target user type based on the user account information, and acquire multiple candidate test operation paths corresponding to the target user type according to the operation task; An optimal test operation path is screened out from a plurality of candidate test operation paths through the UI element visual information, the UI state information and the environmental context information.
6. A UI automated testing device, characterized in that: The UI automation testing device includes a processor, a memory, and a UI automation testing program stored in the memory and executable by the processor, wherein when the UI automation testing program is executed by the processor, the steps of the UI automation testing method according to any one of claims 1 to 4 are implemented.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a UI automation test program, wherein when the UI automation test program is executed by a processor, the steps of the UI automation test method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Exception elimination method, device and equipment for automatic process and storage medium
CN110851299A
Operation simulation method, game test method and related devices
CN112742029A
Automatic testing method and device, storage medium and computer equipment
CN116401164A
Software testing method and device, storage medium and electronic equipment
CN117573544A
Automatic UI (User Interface) interactive exploration method based on multi-modal large model
CN118467032A