Model training method and device, test case generation method and device, medium and product

By building a functional area experience library and UI classification model, the inefficiency problem in cross-platform automated UI testing is solved, and efficient classification and test case generation of multi-platform UI information is realized.

CN120336168APending Publication Date: 2025-07-18CHINA UNITED NETWORK COMM GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510329238.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing technology has low efficiency problems in mid-to-cross-platform automated UI testing, mainly due to insufficient model generalization capabilities and complex testing processes.

Method used

The operation sequence is determined based on the operation video of the sample application, a functional area experience library is built, and the training samples are built using the functional area experience library, and UI classification model is trained to generate a UI classification model.

Benefits of technology

It improves the efficiency of automated testing, simplifies the test process, realizes effective classification and processing of multi-platform UI information, and generates effective test cases that meet multi-platform deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336168A_ABST
    Figure CN120336168A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model training method and device, a test case generation method and device, a medium and a product, and is applied to the technical field of automatic testing. The model training method comprises: according to each operation video of a sample application, determining an operation sequence of the operation video, the operation sequence comprising action information corresponding to each operation action in the operation video and a target video frame; according to the similarity between the different target video frames, a functional area experience library is constructed, the functional area experience library comprises functional area information of a plurality of functional areas, and the functional area information of each functional area comprises action information belonging to the functional area and an operation sequence corresponding to the action information; constructing a training sample according to the functional area experience library; and based on the training sample, performing model training to generate a UI classification model. The method is used for achieving the effect of improving the hall testing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of automated testing technologies, and in particular, to a model training method, a test case generation method, a device, a medium, and a product. Background Art

[0002] With the development of automated testing technologies, in order to increase the user base of an application, it is necessary to design and deploy the application for multiple deployment platforms on the market. After the application is deployed, in order to ensure the stability of the application, it is necessary to perform multi-platform testing on the application, that is, cross-platform testing. In cross-platform testing, the UI testing of the application becomes one of the important means to ensure the stable operation of the application.

[0003] In the prior art, the cross-platform UI testing method is mainly: processing the operation video of the target application through image processing and deep learning, performing multi-level abstraction processing on the organization and architecture of the target application to obtain the usage model of the target application. Using the constructed usage model to construct test cases, so as to achieve multi-platform UI testing of the target application.

[0004] Due to the limitations caused by different platforms in the cross-platform automated testing method in the prior art, there is a technical problem of low automated testing efficiency in the prior art. Summary of the Invention

[0005] The embodiments of this application provide a model training method, a test case generation method, a device, a medium, and a product, so as to achieve the technical effect of improving the automated testing efficiency.

[0006] In a first aspect, the embodiments of this application provide a model training method, including:

[0007] Determining an operation sequence of an operation video according to each operation video of a sample application, where the operation sequence includes action information corresponding to each operation action in the operation video and a target video frame;

[0008] Constructing a functional area experience library according to the similarity between different target video frames, where the functional area experience library includes functional area information of multiple functional areas, and the functional area information of each functional area includes action information belonging to the functional area and the operation sequence corresponding to the action information;

[0009] Constructing a training sample according to the functional area experience library;

[0010] Based on the training sample, performing model training to generate a user interface (UI) classification model.

[0011] In a possible implementation, the action information corresponding to each operation action includes the control type, text content, parent control type, operation action identifier, and screen transition identifier of the controlled control corresponding to the operation action. The screen transition identifier is used to indicate whether the controlled control will cause a screen transition when it is operated.

[0012] In a possible implementation, according to each operation video of the sample application, determine the operation sequence of the operation video, including:

[0013] According to each operation video of the sample application, determine the initial action information corresponding to each operation video;

[0014] If the operation action corresponding to the initial action information is a click action, and the touch point corresponding to the click action is within the keyboard area, then update the operation action corresponding to the initial action information from a click action to an input action, and update the text content corresponding to the initial action information to generate action information;

[0015] If the operation action corresponding to the initial action information is other actions, then determine the initial action information as action information;

[0016] If the operation action corresponding to the initial action information is a click action, and the touch point corresponding to the click action is not within the keyboard area, then determine the initial action information as action information;

[0017] Generate the operation sequence of the operation video according to the execution order of each operation action in the operation video.

[0018] In a possible implementation, updating the text content corresponding to the initial action information includes:

[0019] Obtain the first text content based on the continuous change of the touch point in the keyboard area;

[0020] According to the text recognition technology, recognize the second text content in the input display area;

[0021] Update the text content corresponding to the initial action information according to the first text content and the second text content.

[0022] In a possible implementation, construct training samples according to the function area experience library, including:

[0023] According to the function area experience library, determine the control type, text content, and parent control type of the controlled control included in each action information as training samples;

[0024] Determine the function area identifier corresponding to the action information, as well as the control type, parent control type, operation action identifier, and screen conversion identifier of the controlled control included in the action information, as the label information of the training sample corresponding to the action information.

[0025] In a possible implementation manner, according to the similarity between different target video frames, construct a function area experience library, including:

[0026] Perform grayscale processing on each target video frame to generate a grayscale target video frame;

[0027] Calculate the similarity between different grayscale target video frames;

[0028] Determine the action information corresponding to two grayscale target video frames with a similarity greater than the preset similarity as the action information of the same function area;

[0029] Construct a function area experience library according to the action information of each function area and the operation sequence corresponding to each action information.

[0030] In a second aspect, an embodiment of the present application provides a test case generation method, including:

[0031] Determine the information of the control to be detected in the image of the interface to be detected of the application to be tested, where the information of the control to be detected includes the control type, text content, and parent control type of the control to be detected in the image of the interface to be detected;

[0032] Input the information of the control to be detected into the user interface (UI) classification model to determine the target function area identifier corresponding to the control to be detected. The UI classification model is obtained by training the model through the model training method of any one of items 1-6 in the claims;

[0033] According to the target function area identifier, determine the target operation sequence corresponding to the information of the control to be detected from the target function area corresponding to the target function area identifier in the function area experience library;

[0034] Generate a prompt word based on the target operation sequence, the information of the control to be detected, and the application to be tested. The prompt word is used to guide the large language model to generate a test case for testing the control to be detected;

[0035] Input the prompt word into the large language model to obtain the test case output by the large language model.

[0036] In a possible implementation manner, determining the information of the control to be detected in the image of the interface to be detected of the application to be tested includes:

[0037] Process the image of the interface to be detected according to the target detection algorithm to determine the first control type of the first control;

[0038] Process the image of the interface to be detected according to the optical character recognition algorithm to determine the first text content of the first control;

[0039] Process the image of the interface to be detected according to the image segmentation technology to determine the second control type and the second text content of the second control;

[0040] Determine the control type and text content of the control to be detected according to the first control type, the first text content of the first control, the second control type and the second text content of the second control;

[0041] Determine the parent control type of the control to be detected;

[0042] Generate the information of the control to be detected according to the control type, the text content and the parent control type of the control to be detected.

[0043] In a third aspect, an embodiment of the present application provides a model training device, including:

[0044] The first processing module is used to determine the operation sequence of the operation video according to each operation video of the sample application. The operation sequence includes the action information corresponding to each operation action in the operation video and the target video frame;

[0045] The second processing module is used to construct a functional area experience library according to the similarity between different target video frames. The functional area experience library includes the functional area information of multiple functional areas. The functional area information of each functional area includes the action information belonging to the functional area and the operation sequence corresponding to the action information;

[0046] The third processing module is used to construct training samples according to the functional area experience library;

[0047] The fourth processing module is used to perform model training based on the training samples to generate a user interface UI classification model.

[0048] In a possible implementation manner, the action information corresponding to each operation action includes the control type, the text content, the parent control type, the operation action identifier and the screen conversion identifier of the control to be operated corresponding to the operation action. The screen conversion identifier is used to indicate whether the screen will be converted when the control to be operated is being operated.

[0049] In a possible implementation manner, the first processing module is further used to:

[0050] Determine the initial action information corresponding to each operation video according to each operation video of the sample application;

[0051] If the operation action corresponding to the initial action information is a click action, and the touch point corresponding to the click action is within the keyboard area, then update the operation action corresponding to the initial action information from a click action to an input action, and update the text content corresponding to the initial action information to generate action information;

[0052] If the operation action corresponding to the initial action information is other actions, then determine the initial action information as action information;

[0053] If the operation action corresponding to the initial action information is a click action, and the touch point corresponding to the click action is not within the keyboard area, then determine the initial action information as action information;

[0054] Generate an operation sequence of the operation video according to the execution order of each operation action in the operation video.

[0055] In a possible implementation manner, the first processing module is further configured to:

[0056] Obtain a first text content based on the continuous change of the touch point in the keyboard area;

[0057] Identify a second text content within the input display area according to text recognition technology;

[0058] Update the text content corresponding to the initial action information according to the first text content and the second text content.

[0059] In a possible implementation manner, the third processing module is further configured to:

[0060] According to the function area experience library, determine the control type, text content, and parent control type of the controlled control included in each action information as training samples;

[0061] Determine the function area identifier corresponding to the action information, and the control type, parent control type, operation action identifier, and screen conversion identifier of the controlled control included in the action information as the label information of the training sample corresponding to the action information.

[0062] In a possible implementation manner, the second processing module is further configured to:

[0063] Perform grayscale processing on each target video frame to generate a grayscale target video frame;

[0064] Calculate the similarity between different grayscale target video frames;

[0065] Determine the action information corresponding to two grayscale target video frames with a similarity greater than a preset similarity as the action information of the same function area;

[0066] Construct a function area experience library based on the action information of each function area and the operation sequence corresponding to each action information.

[0067] In a fourth aspect, an embodiment of the present application provides a test case generation device, including:

[0068] A first processing module, configured to determine the information of the control to be detected in the image of the interface to be tested of the application to be tested, where the information of the control to be detected includes the control type, text content, and parent control type of the control to be detected in the image of the interface to be tested;

[0069] A second processing module, configured to input the information of the control to be detected into a user interface (UI) classification model to determine the target function area identifier corresponding to the control to be detected, and the UI classification model is obtained by training the model through the model training method of any item in the first aspect;

[0070] A third processing module, configured to determine the target operation sequence corresponding to the information of the control to be detected from the target function area corresponding to the target function area identifier in the function area experience library according to the target function area identifier;

[0071] A fourth processing module, configured to generate a prompt word based on the target operation sequence, the information of the control to be detected, and the application to be tested, and the prompt word is used to guide a large language model to generate a test case for testing the control to be detected;

[0072] A fifth processing module, configured to input the prompt word into the large language model to obtain the test case output by the large language model.

[0073] In a possible implementation manner, the first processing module is further configured to:

[0074] Process the image of the interface to be tested according to a target detection algorithm to determine the first control type of the first control;

[0075] Process the image of the interface to be tested according to an optical character recognition algorithm to determine the first text content of the first control;

[0076] Process the image of the interface to be tested according to an image segmentation technology to determine the second control type and the second text content of the second control;

[0077] Determine the control type and text content of the control to be detected according to the first control type, the first text content of the first control, and the second control type and the second text content of the second control;

[0078] Determine the parent control type of the control to be detected;

[0079] Generate the information of the control to be detected according to the control type, text content, and parent control type of the control to be detected.

[0080] Fifth aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor;

[0081] The memory stores computer-executable instructions;

[0082] The processor executes the computer-executable instructions stored in the memory, so that the processor executes various possible implementation manners of the first aspect or the second aspect as above.

[0083] Sixth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement various possible implementation manners of the first aspect or the second aspect as above.

[0084] Seventh aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements various possible implementation manners of the first aspect or the second aspect as above.

[0085] The model training method, test case generation method, device, medium and product provided by the embodiments of the present application. In this model training method, based on the operation video of the sample application, the operation sequence of the operation video is determined; a functional area experience library is constructed based on the similarity of multiple different target video frames in the operation video; training samples are constructed by using the functional area experience library, and a UI classification model is obtained based on the training samples. In the model training method of the present application, the operation sequence in the operation video and the action information corresponding to each operation action in the operation sequence are extracted through the operation video of the sample application, and at the same time, the functional area is divided by using the similarity between the video frames corresponding to each action information, so as to construct a functional area experience library based on the operation actions; the training samples are constructed by using the multiple functional area information extracted from the functional area experience library, and the UI classification model obtained by training with the training samples can classify the functional areas according to the interface information of the application. Compared with the prior art, in the process of training the UI classification model in the present application, the functional area experience library is used to perform functional area abstraction processing on each operation video of the sample application, so as to obtain the operation sequences of the sample application on multiple platforms. Using the information in the functional area experience library for model training can effectively classify the UI information of multiple platforms, and the operation sequences in the functional area experience library record the operation sequences of each function or each action, which helps to improve the logical processing ability of the automated test and generate effective test cases that meet the premise of multi-platform deployment of the application, thus achieving the technical effect of improving the efficiency of the automated test. Description of the Drawings

[0086] The accompanying drawings here are incorporated into the description and form a part of this description, showing embodiments consistent with the present application, and are used together with the description to explain the principles of the present application.

[0087] Figure 1 Schematic flow of the model training method provided for the present application Figure 1 ;

[0088] Figure 2 Schematic flow of the model training method provided for the present application Figure 2 ;

[0089] Figure 3 Schematic flow chart of the test case generation method provided for the present application;

[0090] Figure 4 Schematic flow chart of the automated UI testing method for cross-platform applications provided for the embodiments of the present application;

[0091] Figure 5 Schematic structural diagram of the model training device provided for the present application;

[0092] Figure 6 Schematic structural diagram of the test case generation device provided for the present application;

[0093] Figure 7 Schematic structural diagram of the electronic device provided for the present application.

[0094] Through the above accompanying drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed implementation manners

[0095] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0096] First, the terms related to the present application are explained:

[0097] User Interface (UI): The user interface refers to the interface for users to interact with a computer system, software, or website. It includes graphical elements, text, buttons, menus, etc., aiming to improve the user experience and ensure that users can use and operate.

[0098] You Only Look Once (YOLO): It is an object detection algorithm that can quickly identify and locate multiple objects in videos or images. Different from traditional object detection methods, YOLO transforms object detection into a regression problem and processes images through a single neural network to achieve highly efficient detection.

[0099] Optical Character Recognition (OCR): Optical character recognition is a technology used to convert text in images into machine-readable text. This technology is widely applied in fields such as document digitization, automated data entry, and archive management.

[0100] Graphical User Interface (GUI): The graphical user interface is a way of interaction between users and devices. Through graphical elements such as windows, icons, and buttons, it displays information and receives input from users. The design of the GUI makes operations more intuitive, and compared with command-line interfaces, it is easier for users to understand and use.

[0101] Red Green Blue (RGB): RGB is a color representation model. By combining the three primary colors of red, green, and blue with different intensities, various colors can be generated. RGB is widely used in the fields of image processing, design, and display devices and is the basic standard for digital photography and video colors.

[0102] In the prior art, when solving cross-platform automated UI testing, mainly operation videos of multiple sample applications are collected, manual annotation is performed on the feature information extracted from the operation videos, and training data is generated. Based on the training data, the training of the test model is carried out.

[0103] However, in actual multi-platform automated UI testing, there are obvious limitations in the test results of the cross-platform migration ability of the test model. For example, when facing a new deployment platform, due to the low quantity of training data for this platform, the generalization ability of the model for this platform is low. At the same time, because the training data adopts the method of manual annotation, the overall model training efficiency is low. And when the prior art conducts automated test processing for multiple platforms or multiple applications and generates test cases in combination with the test model, it is necessary to design targeted test logics for each platform and each application, resulting in a complex overall test process. Therefore, there is a technical problem of low automated test efficiency in the prior art.

[0104] To address the above technical problems, the present application proposes the following technical concept: aiming at the technical problems of low automation testing efficiency caused by low model generalization ability and complex testing processes in the prior art. The present application proposes a model training method that improves the model training efficiency and generalization ability and simplifies the testing process. Specifically: based on the operation videos of sample applications, determine the operation sequences of the operation videos; construct a functional area experience library based on the similarity of multiple different target video frames in the operation videos; use the functional area experience library to construct training samples, and perform model training based on the training samples to obtain a UI classification model. Among them, the operation sequences extracted from the operation videos of sample applications are also stored in the functional area experience library, and the functional area information of each functional area includes the action information belonging to the functional area and the operation sequence corresponding to the action information. Compared with the prior art, in the process of training the UI classification model in the present application, the mapping relationship between each action information, functional area information, and operation sequence is determined, realizing the automatic annotation of training data; at the same time, the extracted operation sequences reflect the operation logics of multiple functions and actions on multiple platforms; thus, it can effectively classify the UI information of multiple platforms and improve the logical processing ability of automatic testing using the operation sequences; thereby achieving the technical effect of improving the efficiency of automatic testing.

[0105] The following uses specific embodiments to elaborate in detail on the technical solutions of the present application and how the technical solutions of the present application solve the above technical problems. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0106] Figure 1 Flow schematic of the model training method provided by the present application Figure 1 , such as Figure 1 shown, this method includes:

[0107] S101. According to each operation video of the sample application, determine the operation sequence of the operation video.

[0108] In this step, the operation sequence includes the action information corresponding to each operation action in the operation video and the target video frame.

[0109] Among them, the action information corresponding to each operation action includes the control type, text content, parent control type, operation action identifier, and screen conversion identifier of the controlled control corresponding to the operation action. The screen conversion identifier is used to indicate whether the controlled control will cause a screen conversion when being operated.

[0110] Exemplarily, the expression form of the action information can be: {control type of the controlled control: button; parent control type: form; text content: submit; operation action identifier: btn_submit; screen transition identifier: last}. Here, btn_submit refers to the identifier of the current operation action being the submit operation action, and the screen transition identifier last means that the current operation action is the last one in the screen transition sequence and can cause a screen transition.

[0111] In this step, the target video frames included in each operation sequence refer to the video frames corresponding to each operation action in the operation sequence.

[0112] Exemplarily, there is an operation video with 6000 video images. For each video image frame, the recognition of operation actions and the extraction of action information are performed. Finally, the target video frames corresponding to each operation action are determined. By using multiple loop video recognitions and analyses, combined with the extracted operation actions and action information, the operation sequences in this operation video are extracted according to the association relationships between the operation actions.

[0113] It should be noted that the specific implementation method for determining the operation sequence in this step will be further explained in detail in the embodiments shown below, and no redundant elaboration will be made here. Figure 2 shown in the embodiments shown below, and no redundant elaboration will be made here.

[0114] S102. Construct a functional area experience library according to the similarity between different target video frames.

[0115] In this step, the functional area experience library includes the functional area information of multiple functional areas. The functional area information of each functional area includes the action information belonging to the functional area and the operation sequence corresponding to the action information.

[0116] Among them, the functional area information of each functional area contains the action information of at least one operation action.

[0117] Exemplarily, there is a functional area referring to the user login function. The operation actions included in this functional area are respectively: Action One, Action Two, and Action Three. Among them, the action types of Action One and Action Three are click actions, and the action type of Action Two is an input action. Correspondingly, the functional area information contains the action information of these three operation actions: Action One, Action Two, and Action Three.

[0118] Optionally, a possible implementation manner for constructing the functional area experience library based on the similarity between different target video frames is:

[0119] S1021. Perform grayscale processing on each target video frame to generate the grayscale target video frame.

[0120] In this step, the method of grayscale processing for each video frame can be as follows:

[0121] Use the weighted average method to assign different weights to different colors in the target video frame to calculate the grayscale value. Or calculate the average value of the RGB values of each pixel in the target video frame and use it as the grayscale value.

[0122] S1022. Calculate the similarity between the target video frames after different grayscale conversions.

[0123] In this step, the method of calculating the similarity of the target video frames can be: calculate the error between each corresponding pixel in two target video frames and perform mean square processing on the error to obtain the similarity value of the two target video frames.

[0124] It should be noted that in this application, only examples of the method for calculating the similarity of the target video frames are given, and no specific restrictions are imposed on it.

[0125] S1023. Determine the action information corresponding to two grayscale target video frames with a similarity greater than the preset similarity as the action information of the same functional area.

[0126] In this step, the similarity of the target video frames is used to determine whether the current operation actions belong to the same functional area, so as to achieve the division of the functional area. Exemplarily, assume that the value of the preset similarity is 0.9. Then calculate the similarity between two target video frames after grayscale processing. If the calculated similarity is greater than 0.9, it indicates that the action information corresponding to these two target video frames belongs to the same functional area; otherwise, it indicates that the action information corresponding to these two target video frames does not belong to the same functional area.

[0127] S1024. Construct a functional area experience library according to the action information of each functional area and the operation sequence corresponding to each action information.

[0128] In this step, after determining the action information of each functional area, the target video frames saved in the operation sequence can be deleted. The finally obtained functional area experience library contains functional area information corresponding to multiple functional areas, and each functional area information contains at least one action information.

[0129] S103. Construct training samples according to the functional area experience library.

[0130] Optionally, a possible implementation method for constructing training samples according to the functional area experience library is:

[0131] S1031. According to the functional area experience library, determine the control type, text content, and parent control type of the controlled control included in each action information as training samples.

[0132] S1032. Determine the function area identifier corresponding to the action information, as well as the control type, parent control type, operation action identifier, and screen transition identifier of the controlled control included in the action information, as the label information of the training sample corresponding to the action information.

[0133] S104. Based on the training samples, perform model training to generate a UI classification model.

[0134] In this step, the input information of the UI classification model is the GUI information extracted from the interface image of the application. The GUI information includes the control information in the interface image and the text information related to the control information. The output information of the UI classification model is the classification result of the interface image of the application. The output information includes the function area identifier of the function area corresponding to the interface image, the operation action identifier of the operation action corresponding to the interface image, the screen transition identifier of the operation action corresponding to the interface image, and the sorted control type and parent control type corresponding to the operation action.

[0135] Among them, the control information refers to the control type existing in the interface image, and the text information refers to the text content existing in each control in the interface image or the text content related to the control.

[0136] The model training method provided in the embodiments of the present application determines the operation sequence of the operation video based on the operation video of the sample application, constructs a function area experience library based on the similarity of multiple different target video frames in the operation video, constructs training samples using the function area experience library, and performs model training based on the training samples to obtain a UI classification model. Compared with the prior art, in the process of training the UI classification model in the present application, the function area experience library construction method is used to perform function area abstraction processing on each operation video of the sample application, so as to obtain the operation sequences of the sample application on multiple platforms. Using the information in the function area experience library for model training can effectively classify the UI information of multiple platforms, and the operation sequences in the function area experience library record the operation sequences of each function or each action, which helps to improve the logical processing ability of automated testing and generate effective test cases that meet the premise of multi-platform deployment of the application, thus achieving the technical effect of improving the efficiency of automated testing.

[0137] Figure 2 Schematic flow of the model training method provided by the present application Figure 2 Based on the above Figure 1 On the basis of the embodiment shown, this embodiment further explains the determination of the operation sequence in step S101. As Figure 2 shown, the method includes:

[0138] S201. According to each operation video of the sample application, determine the initial action information corresponding to each operation video.

[0139] In this step, a possible implementation for determining the initial action information corresponding to the operation video is as follows:

[0140] S2011. Decompose each operation video to generate multiple video frames.

[0141] Among them, the method for generating video frames based on each operation video can be: extracting video frames from the operation video based on a preset time period, so as to obtain multiple video frames corresponding to each operation video.

[0142] S2012. Adjust the size of each video frame to a unified size.

[0143] S2013. Perform action detection on the adjusted video frames, detect the initial actions existing in the video frames, and generate corresponding initial action information.

[0144] Among them, the initial actions can be: click action, long - press action, and swipe action.

[0145] S202. If the operation action corresponding to the initial action information is a click action, and the touch point corresponding to the click action is within the keyboard area, then update the operation action corresponding to the initial action information from the click action to an input action, and update the text content corresponding to the initial action information to generate action information.

[0146] In this step, when it is detected that the operation action is a click action and the touch point corresponding to the click action is within the keyboard, it indicates that the current click action is related to keyboard input. Therefore, update the click action to an input action, and it is necessary to synchronously associate the control type and the parent control type corresponding to the click action to the updated input action.

[0147] Meanwhile, after updating the click action to an input action, it is necessary to update the text content corresponding to the input action to ensure the accuracy of the action information.

[0148] Optionally, a possible implementation for updating the text content corresponding to the initial action information is as follows:

[0149] S2021. Obtain the first text content based on the continuous change of the touch point in the keyboard area.

[0150] In this step, the continuous change of the touch point in the keyboard area refers to the information input on the keyboard, that is, the first text content, and this information is not equal to the text information actually displayed in the input control.

[0151] S2022. Identify the second text content in the input display area according to text recognition technology.

[0152] In this step, the second text content obtained by the text recognition technology from the input display area refers to the text information within the current display area, that is, the second text content.

[0153] S2023. Update the text content corresponding to the initial action information according to the first text content and the second text content.

[0154] In this step, updating the text content corresponding to the initial action information means using the first text content and the second text content to determine the effective text content that is actually input to the display area by the keyboard and can be effectively displayed.

[0155] Exemplarily, the first text content input in the keyboard area includes the actual effective text content and the invalid text content generated due to modification and deletion operations; the second text content recognized in the display area includes the default text content in the display area and the actual input effective text content. According to the overlapping part of the first text content and the second text content, the actual effective text content is determined, and thus the text content in the initial action information is updated based on the effective text content.

[0156] Optionally, in a possible implementation, if the display mode of the display area corresponding to the current input action is encrypted display, the second text content cannot be recognized from the display area by means of text recognition, then the second text content corresponding to the first text content can be obtained from the default configuration file according to the first text content.

[0157] Among them, the configuration file for obtaining the second text content can be a configuration file of encrypted information, and the encrypted information can be login information.

[0158] S203. If the operation action corresponding to the initial action information is other actions, determine the initial action information as the action information.

[0159] Exemplarily, there is a long - press action, and the action information corresponding to this action is: {type of the controlled control: list item; type of the parent control: list view; text content: delete; operation action identifier: acttion_long_press; screen transition identifier: nolast}, where acttion_long_press refers to the identifier of the long - press action, and nolast refers to that the current long - press action does not cause a screen transition.

[0160] S204. If the operation action corresponding to the initial action information is a click action, and the touch point corresponding to the click action is not within the keyboard area, determine the initial action information as the action information.

[0161] In this step, if the touch point corresponding to the click action is not within the keyboard area, it indicates that there is no input behavior for this click action, and the text content of the control being operated corresponding to the click action is the text content in the action information.

[0162] Exemplarily, there is a click action, and the action information corresponding to this action is: {control type of the control being operated: button; parent control type: view; text content: Confirm; operation action identifier: action_click; screen transition identifier: single}, where action_click refers to the operation action identifier for the current click operation, and single means that the current click action causes a screen transition, and there is only one such click action in the operation sequence that causes the screen transition to which this click action belongs.

[0163] S205. Generate an operation sequence for the operation video according to the execution order of each operation action in the operation video.

[0164] In this step, the operation sequence of the operation video includes action information corresponding to multiple operation actions and target video frames. Among them, the way to store the target video frames in the operation sequence is to store the address information of the target video frames for reading the target video frames from the storage control.

[0165] In this embodiment, by extracting action information from each operation video of the sample application, the operation sequence and action information in the operation video are obtained; at the same time, the position of the touch point is used to determine the input action, so as to newly extract the input action and enrich the obtained action information. Combining the action information and the operation sequence can reflect the execution order of each operation action in the operation video, which is beneficial to the generation of test cases in automated testing.

[0166] Figure 3 is a schematic flowchart of the test case generation method provided by this application. As Figure 3 shown, this method includes:

[0167] S301. Determine the information of the control to be detected in the image of the interface to be detected of the application to be tested.

[0168] In this step, the information of the control to be detected includes the control type, text content, and parent control type of the control to be detected in the image of the interface to be detected.

[0169] Optionally, a possible implementation manner for determining the information of the control to be detected in the image of the interface to be detected of the application to be tested is:

[0170] S3011. Process the image of the interface to be detected according to the target detection algorithm to determine the first control type of the first control.

[0171] In this step, the target detection algorithm used can be the open-source YOLO algorithm. By using the target detection algorithm to process the image of the interface to be detected, multiple UI elements in the image of the interface to be detected and the type of each UI element can be obtained. The multiple UI elements are determined as the first controls, and the type of the UI element is determined as the first control type.

[0172] S3012. Process the image of the interface to be detected according to the optical character recognition algorithm, and determine the first text content of the first control.

[0173] In this step, the optical character recognition algorithm can be the OCR algorithm. The way to obtain the first text content by using the OCR algorithm to process the image of the interface to be detected is as follows:

[0174] Extract text information from the image of the interface to be detected, determine the text information corresponding to the position of each first control based on the first control, and create a corresponding relationship between each first control and the text information, so as to determine the first text content of the first control.

[0175] S3013. Process the image of the interface to be detected according to the image segmentation technology, and determine the second control type and the second text content of the second control.

[0176] In this step, the image segmentation technology is used to detect the controls with rich colors and high content freedom in the image of the interface to be detected. The way to determine the second control type and the second text content of the second control by using the image segmentation technology is as follows:

[0177] Perform grayscale processing on the image of the interface to be detected to obtain a grayscale image after grayscale processing; process each pixel in the grayscale image by means of binaryzation to obtain a binary image; enhance the connectivity and robustness of the binary image by means of dilation processing, and remove impurities by filtering to obtain a dilated image; use the Hongfan algorithm to segment the dilated image to obtain the white areas in the image, and label each segmented white area as the second control, and use the segmented image to determine the second control type and the second text content corresponding to each second control.

[0178] Among them, when performing image processing by using the image segmentation technology, the visually overlapping segmentation image frames can be combined into a large segmentation image to restore the integrity of the segmentation image, avoid the problem of over-segmentation in the image segmentation process, and ensure the accuracy of the second control recognition.

[0179] S3014. Determine the control type and text content of the control to be detected according to the first control type, the first text content of the first control, and the second control type and the second text content of the second control.

[0180] In this step, based on the first control type and the first text content of the first control, as well as the second control type and the second text content of the second control, duplicate data is identified, and the part of the second control that overlaps with the first control is excluded, so as to obtain the control type and text content of the control to be detected.

[0181] Among them, in addition to excluding overlapping control information, inoperable control information can also be excluded. If there is a title bar control or a control when a dialog box is activated in the current first control or second control, it is determined as an inoperable control and excluded.

[0182] S3015. Determine the parent control type of the control to be detected.

[0183] In this step, the method for determining the parent control type can be: using the method defined by features to determine the parent control features of each control to be detected; combining the image of the interface to be detected to determine the overlapping relationship between each control; combining the overlapping relationship and the parent control features to determine the parent control type corresponding to each control to be detected.

[0184] S3016. Generate control information to be detected according to the control type, text content and parent control type of the control to be detected.

[0185] S302. Input the control information to be detected into the UI classification model to determine the target function area identifier corresponding to the control to be detected.

[0186] In this step, the UI classification model is obtained through Figure 1 and Figure 2 Any of the model training methods in the embodiments is trained to determine the function area identifier to which the control to be detected belongs.

[0187] Optionally, in a possible implementation manner, the UI classification model can generate a function area identifier, an operation action identifier, and a screen conversion identifier corresponding to each control to be detected based on the input control information to be detected. The function area identifier is used to determine the function area to which the control to be detected belongs, the operation action identifier is used to determine the operation action for operating the control to be detected, and the screen conversion identifier is used to determine whether the operation action will cause a screen conversion.

[0188] S303. According to the target function area identifier, determine the target operation sequence corresponding to the control information to be detected from the target function area corresponding to the target function area identifier in the function area experience library.

[0189] In this step, based on the ribbon identifier, the target ribbon can be determined from the ribbon experience library. The ribbon information of the target ribbon contains the action information belonging to this ribbon and the operation sequence corresponding to the action information. Based on the action information and operation sequence in the target ribbon, determine the target operation sequence corresponding to the control information to be detected, where the target operation sequence refers to the operation sequence related to the control to be detected.

[0190] Exemplarily, the target operation sequence contains three action information. The control type of the controlled control corresponding to the first action is the same as the control type of the control to be detected. The control type of the operating control corresponding to the second action is the same as the control type of the parent control of the control to be detected. The parent control type of the operating control corresponding to the third action is the same as the control type of the control to be detected. Then, based on the operation sequences of the three actions, the operation sequence related to the action to be detected can be determined.

[0191] S304. Generate a prompt word based on the target operation sequence, the control information to be detected, and the application to be tested.

[0192] In this step, the prompt word is used to guide the large language model to generate test cases for testing the control to be detected.

[0193] Among them, the target operation sequence contains the execution logic of the operation actions related to the control to be detected. The control information to be detected is the operation control included in the current image of the interface to be detected. The application to be tested indicates the execution object of the test case. Generate a prompt word based on the above information. The prompt word needs to include the execution object of the test case, the execution control of the test case, and the execution logic of the test case.

[0194] S305. Input the prompt word into the large language model to obtain the test cases output by the large language model.

[0195] Exemplarily, the form of the prompt word input into the large language model can be: {Execution control: Add button; Execution logic: Determine the quantity of selected products based on the quantity input box of the parent control of the add button, and add products based on the add button; Execution object: Application A}, and the test case: {Enter the product quantity 3 in the quantity input box and click the add button to add 3 products}.

[0196] A test case generation method provided by an embodiment of the present application identifies an image of an interface to be detected to obtain information about controls to be detected, uses a UI classification model to obtain classification information of the information about the controls to be detected, that is, a function area identifier; uses an operation sequence corresponding to the function area identifier in a function area experience library, combines the information about the controls to be detected and the application to be detected to generate a prompt, and uses a large language model to generate a test case corresponding to the prompt, thereby quickly generating a test case for automated testing while ensuring the logical integrity and accuracy of the test case, achieving the technical effect of improving the efficiency of automated testing.

[0197] Based on the above embodiment, an embodiment of the present application provides a method for automated UI testing for cross-platform applications. Figure 4 It is a schematic flowchart of the method for automated UI testing for cross-platform applications provided by an embodiment of the present application, as Figure 4 shown, this method includes:

[0198] Step 1: UI modeling. Specifically:

[0199] A1. Collect operation videos of a sample application on multiple platforms, and extract the operation sequences in the operation videos of the sample application.

[0200] A2. Use the target video frames in the operation sequence for function area division to create a function area experience library.

[0201] A3. Use the function area experience library to extract training samples, and for each piece of data in the training samples, perform sample annotation according to the function area to which it belongs to obtain the training samples after annotation.

[0202] A4. Obtain the picture of the interface to be detected of the application to be detected, and extract GUI information based on the picture of the interface to be detected.

[0203] A5. Perform model training based on the training samples to obtain a UI classification model. Use the UI classification model to perform UI analysis on the GUI information to obtain UI classification information.

[0204] Step 2: Scenario testing. Specifically:

[0205] B1. Extract sequence features corresponding to the UI classification information from the function area experience library based on the UI classification information.

[0206] B2. Determine the target operation sequence corresponding to the UI classification information based on the sequence features.

[0207] B3. Generate a prompt based on the target operation sequence, GUI information, and the application to be detected, input the prompt into a large language model to obtain a test case, and perform automated UI testing on the application to be detected based on the test case.

[0208] Figure 5 The structural schematic diagram of the model training device provided for this application is as follows Figure 5 As shown, the model training device provided in this embodiment includes:

[0209] The first processing module 501 is configured to determine an operation sequence of an operation video according to each operation video of a sample application. The operation sequence includes action information corresponding to each operation action in the operation video and a target video frame;

[0210] The second processing module 502 is configured to construct a functional area experience library according to the similarity between different target video frames. The functional area experience library includes functional area information of multiple functional areas. The functional area information of each functional area includes action information belonging to the functional area and an operation sequence corresponding to the action information;

[0211] The third processing module 503 is configured to construct training samples according to the functional area experience library;

[0212] The fourth processing module 504 is configured to perform model training based on the training samples to generate a user interface UI classification model.

[0213] In a possible implementation manner, the action information corresponding to each operation action includes the control type, text content, parent control type, operation action identifier, and screen conversion identifier of the controlled control corresponding to the operation action. The screen conversion identifier is used to indicate whether the controlled control will cause a screen conversion when being operated.

[0214] In a possible implementation manner, the first processing module 501 is further configured to:

[0215] Determine initial action information corresponding to each operation video according to each operation video of the sample application;

[0216] If the operation action corresponding to the initial action information is a click action, and the touch point corresponding to the click action is within the keyboard area, then update the operation action corresponding to the initial action information from a click action to an input action, and update the text content corresponding to the initial action information to generate action information;

[0217] If the operation action corresponding to the initial action information is other actions, then determine the initial action information as action information;

[0218] If the operation action corresponding to the initial action information is a click action, and the touch point corresponding to the click action is not within the keyboard area, then determine the initial action information as action information;

[0219] Generate an operation sequence of the operation video according to the execution order of each operation action in the operation video.

[0220] In a possible implementation, the first processing module 501 is further configured to:

[0221] Obtain first text content based on continuous changes of the touch point in the keyboard area;

[0222] Identify second text content within the input display area according to text recognition technology;

[0223] Update the text content corresponding to the initial action information according to the first text content and the second text content.

[0224] In a possible implementation, the third processing module 503 is further configured to:

[0225] Determine, according to the function area experience library, the control type, text content, and parent control type of the controlled control included in each action information as training samples;

[0226] Determine the function area identifier corresponding to the action information, and the control type, parent control type, operation action identifier, and screen conversion identifier of the controlled control included in the action information as the label information of the training sample corresponding to the action information.

[0227] In a possible implementation, the second processing module 502 is further configured to:

[0228] Perform grayscale processing on each target video frame to generate a grayscale target video frame;

[0229] Calculate the similarity between different grayscale target video frames;

[0230] Determine the action information corresponding to two grayscale target video frames with a similarity greater than a preset similarity as the action information of the same function area;

[0231] Construct a function area experience library according to the action information of each function area and the operation sequence corresponding to each action information.

[0232] The model training device provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.

[0233] Figure 6 It is a structural schematic diagram of the test case generation device provided in this application. As Figure 6 shown, the test case generation device provided in this embodiment includes:

[0234] A first processing module 601, configured to determine the controlled control information of the to-be-detected interface image of the to-be-tested application, where the controlled control information includes the control type, text content, and parent control type of the to-be-detected control in the to-be-detected interface image;

[0235] A second processing module 602, configured to input the information of the control to be detected into a user interface UI classification model, and determine a target function area identifier corresponding to the control to be detected. The UI classification model is obtained by training the model through the model training method of any item in the first aspect;

[0236] A third processing module 603, configured to determine a target operation sequence corresponding to the information of the control to be detected from the target function area corresponding to the target function area identifier in the function area experience library according to the target function area identifier;

[0237] A fourth processing module 604, configured to generate a prompt word based on the target operation sequence, the information of the control to be detected, and the application to be tested. The prompt word is used to guide a large language model to generate a test case for testing the control to be detected;

[0238] A fifth processing module 605, configured to input the prompt word into the large language model and obtain the test case output by the large language model.

[0239] In a possible implementation manner, the first processing module 601 is further configured to:

[0240] Process the image of the interface to be detected according to the target detection algorithm to determine the first control type of the first control;

[0241] Process the image of the interface to be detected according to the optical character recognition algorithm to determine the first text content of the first control;

[0242] Process the image of the interface to be detected according to the image segmentation technology to determine the second control type and the second text content of the second control;

[0243] Determine the control type and text content of the control to be detected according to the first control type, the first text content of the first control, the second control type of the second control, and the second text content;

[0244] Determine the type of the parent control of the control to be detected;

[0245] Generate the information of the control to be detected according to the control type, the text content, and the type of the parent control of the control to be detected.

[0246] The test case generation device provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.

[0247] Figure 7 It is a schematic structural diagram of an electronic device provided in this application. As Figure 7As shown in the figure, the electronic device provided in this embodiment includes: at least one processor 701 and a memory 702. Optionally, the device further includes a communication component 703. Among them, the processor 701, the memory 702, and the communication component 703 are connected through a bus 704.

[0248] In a specific implementation process, at least one processor 701 executes the computer-executable instructions stored in the memory 702, so that at least one processor 701 executes the above model training method or test case generation method.

[0249] For the specific implementation process of the processor 701, reference can be made to the above method embodiment. The implementation principle and technical effect are similar, and will not be elaborated here in this embodiment.

[0250] In the above embodiment, it should be understood that the processor may be a central processing unit (English: Central Processing Unit, abbreviated as: CPU), or other general-purpose processors, digital signal processors (English: Digital Signal Processor, abbreviated as: DSP), application specific integrated circuits (English: Application Specific Integrated Circuit, abbreviated as: ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor.

[0251] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.

[0252] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the bus in the attached drawings of this application is not limited to only one bus or one type of bus.

[0253] This application also provides a computer program product, including a computer program, which implements the above model training method or test case generation method when executed by a processor.

[0254] The present application further provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above model training method or test case generation method.

[0255] The above-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0256] An exemplary readable storage medium is coupled to the processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in a device.

[0257] The division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can be in electrical, mechanical, or other forms.

[0258] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0259] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0260] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., which are various media that can store program codes.

[0261] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When this program is executed, it executes the steps of including the above method embodiments; and the foregoing storage medium includes: ROMs, RAMs, magnetic disks, or optical discs, etc., which are various media that can store program codes.

[0262] Finally, it should be noted that: After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily think of other implementation manners of the present invention. The present invention is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include the common general knowledge or conventional technical means in the technical field not disclosed in the present invention. It is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A model training method, characterized in that, including: Determine the operation sequence of each operation video according to the sample application, where the operation sequence includes the action information corresponding to each operation action in the operation video and the target video frame; Construct a functional area experience library according to the similarity between different target video frames. The functional area experience library includes the functional area information of multiple functional areas. The functional area information of each functional area includes the action information belonging to the functional area and the operation sequence corresponding to the action information; Construct a training sample according to the functional area experience library; Based on the training sample, perform model training to generate a user interface (UI) classification model.

2. The method according to claim 1, wherein The action information corresponding to each operation action includes the control type, text content, parent control type, operation action identifier, and screen conversion identifier of the controlled control corresponding to the operation action. The screen conversion identifier is used to indicate whether the controlled control will cause a screen conversion when being operated.

3. The method according to claim 2, wherein The determining the operation sequence of the operation video according to each operation video of the sample application includes: Determine the initial action information corresponding to each operation video according to each operation video of the sample application; If the operation action corresponding to the initial action information is a click action and the touch point corresponding to the click action is within the keyboard area, update the operation action corresponding to the initial action information from the click action to an input action, and update the text content corresponding to the initial action information to generate the action information; If the operation action corresponding to the initial action information is other actions, determine the initial action information as the action information; If the operation action corresponding to the initial action information is a click action and the touch point corresponding to the click action is not within the keyboard area, determine the initial action information as the action information; Generate the operation sequence of the operation video according to the execution order of each operation action in the operation video.

4. The method according to claim 3, wherein The updating the text content corresponding to the initial action information includes: Obtain the first text content based on the continuous change of the touch point in the keyboard area; Identify the second text content in the input display area according to text recognition technology; Update the text content corresponding to the initial action information according to the first text content and the second text content.

5. The method according to any one of claims 2-4, characterized in that, The constructing a training sample according to the functional area experience library includes: According to the functional area experience library, determine the control type, text content, and parent control type of the controlled control included in each action information as the training sample; Determine the functional area identifier corresponding to the action information, and the control type, parent control type, operation action identifier, and screen conversion identifier of the controlled control included in the action information as the label information of the training sample corresponding to the action information.

6. The method according to any one of claims 1-4, characterized in that, The constructing a functional area experience library according to the similarity between different target video frames includes: Perform grayscale processing on each target video frame to generate a grayscale target video frame; Calculate the similarity between different grayscale target video frames; Determine the action information corresponding to the two grayscale target video frames with a similarity greater than the preset similarity as the action information of the same functional area; Construct the functional area experience library according to the action information of each functional area and the operation sequence corresponding to each action information.

7. A test case generation method, characterized in that Including: Determine the information of the control to be detected in the image of the interface to be detected of the application to be tested, where the information of the control to be detected includes the control type, text content, and parent control type of the control to be detected in the image of the interface to be detected; Input the information of the control to be detected into the user interface (UI) classification model to determine the target functional area identifier corresponding to the control to be detected, and the UI classification model is obtained by training the model through the model training method described in any one of claims 1-6; According to the target functional area identifier, determine the target operation sequence corresponding to the information of the control to be detected from the target functional area corresponding to the target functional area identifier in the functional area experience library; Generate a prompt word based on the target operation sequence, the information of the control to be detected, and the application to be tested, where the prompt word is used to guide the large language model to generate test cases for testing the control to be detected; Input the prompt word into the large language model to obtain the test cases output by the large language model.

8. The method according to claim 7, wherein The determination of the information of the control to be detected in the image of the interface to be detected of the application to be tested includes: Process the image of the interface to be detected according to the target detection algorithm to determine the first control type of the first control; Process the image of the interface to be detected according to the optical character recognition algorithm to determine the first text content of the first control; Process the image of the interface to be detected according to the image segmentation technology to determine the second control type and the second text content of the second control; Determine the control type and text content of the control to be detected according to the first control type and first text content of the first control, and the second control type and second text content of the second control; Determine the parent control type of the control to be detected; Generate the information of the control to be detected according to the control type, text content, and parent control type of the control to be detected.

9. A model training device, characterized in that, Including: The first processing module is used to determine the operation sequence of the operation video according to each operation video of the sample application, and the operation sequence includes the action information and the target video frame corresponding to each operation action in the operation video; The second processing module is used to construct a functional area experience library according to the similarity between different target video frames, and the functional area experience library includes the functional area information of multiple functional areas, and the functional area information of each functional area includes the action information belonging to the functional area and the operation sequence corresponding to the action information; The third processing module is used to construct training samples according to the functional area experience library; The fourth processing module is used to perform model training based on the training samples to generate a user interface (UI) classification model.

10. A test case generation device, characterized in that, Including: The first processing module is used to determine the information of the control to be detected in the image of the interface to be detected of the application to be tested, where the information of the control to be detected includes the control type, text content, and parent control type of the control to be detected in the image of the interface to be detected; A second processing module, configured to input the to-be-detected control information into a user interface (UI) classification model to determine a target function area identifier corresponding to the to-be-detected control, where the UI classification model is obtained by training a model through the model training method described in any one of claims 1-6; A third processing module, configured to determine a target operation sequence corresponding to the to-be-detected control information from a target function area corresponding to the target function area identifier in a function area experience library according to the target function area identifier; A fourth processing module, configured to generate a prompt word based on the target operation sequence, the to-be-detected control information, and the to-be-tested application, where the prompt word is used to guide a large language model to generate a test case for testing the to-be-detected control; A fifth processing module, configured to input the prompt word into the large language model to obtain the test case output by the large language model.

11. An electronic device, characterized in that, Comprising: A memory and a processor; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory, so that the processor executes the method described in any one of claims 1-6, or executes the method described in any one of claims 7-8.

12. A computer-readable storage medium, characterized in that, Computer execution instructions are stored in the computer-readable storage medium, and when the computer execution instructions are executed by a processor, they are used to implement the method described in any one of claims 1-6, or execute the method described in any one of claims 7-8.

13. A computer program product, characterized in that, Comprising a computer program, which when executed by a processor, implements the method described in any one of claims 1-6, or executes the method described in any one of claims 7-8.