Application test method and device, electronic equipment, product and storage medium

Through the test method based on the deep Q network, the deep Q network is trained to evaluate the stability of the application by utilizing the application's historical state space information and action space information, solving the problem that existing testing methods are difficult to achieve fine-grained stability testing, and achieving efficient and flexible testing coverage.

CN119938531APending Publication Date: 2025-05-06CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1

Patent Information

Application Number
CN202411999519.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing software testing methods such as traditional manual use cases and model-based testing methods have problems such as high labor costs, difficulty in covering all test scenarios, and large test granularity, and cannot effectively conduct fine-grained stability testing.

Method used

Using a test method based on the deep Q network, the deep Q network is trained to determine the stability report by obtaining the application's historical state space information and action space information, and to evaluate whether abnormal situations occur during the comprehensive traversal and testing of the application user interface.

Benefits of technology

This method can effectively evaluate the stability of the application and achieve fine-grained test coverage. It is more flexible than traditional methods, adapts to different scenarios, and improves testing efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938531A_ABST
    Figure CN119938531A_ABST
Patent Text Reader

Abstract

The invention discloses an application testing method and device, electronic equipment, a product and a storage medium. The method comprises the steps of obtaining first state space information of an application in a historical operation process and first action space information corresponding to the first state space information; determining a trained depth Q network based on the first state space information, the first action space information and a preset depth Q network; the preset deep Q network is a deep neural network based on a DQN algorithm; acquiring second state space information of the to-be-tested application in the current running process and second action space information corresponding to the second state space information; inputting the second state space information and the second action space information into the trained deep Q network, and outputting a test report corresponding to the to-be-tested application; the test report comprises a stability report about whether an abnormal condition occurs or not in the process of comprehensively traversing and testing the UI of the to-be-tested application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of software testing technology, and in particular to an application testing method, device, electronic equipment, product and storage medium. Background Art

[0002] In the prior art, software testing is usually performed using traditional manual use cases or model-based testing methods. However, the traditional manual use case testing method has high labor costs, and the written use cases are difficult to cover all possible test scenarios. The model-based testing method cannot test a group of pages, and its test granularity is large. Summary of the invention

[0003] To solve related technical problems, the embodiments of the present application provide a testing method, device, electronic device, product and storage medium for application.

[0004] The technical solution of the embodiment of the present application is implemented as follows:

[0005] The present application embodiment provides a method for testing an application, including:

[0006] Obtaining first state space information of the application during historical operation and first action space information corresponding to the first state space information;

[0007] Determining a trained deep Q network based on the first state space information, the first action space information and a preset deep Q network; the preset deep Q network is a deep neural network based on a deep Q network (Deep Q-Network, DQN) algorithm;

[0008] Acquire second state space information of the application to be tested during the current running process and second action space information corresponding to the second state space information;

[0009] The second state space information and the second action space information are input into the trained deep Q network, and a test report corresponding to the application to be tested is output; the test report includes a stability report on whether abnormal situations occur during a comprehensive traversal and testing of the user interface (UI) of the application to be tested.

[0010] In the above solution, the first state space information includes first page element information, first control layout information and first path information; the method further includes:

[0011] Determine executable operation information of the first page element information based on the first page element information, the first control layout information and the first path information;

[0012] The executable operation information of the first page element information is used as the first action space information.

[0013] In the above solution, determining a trained deep Q network based on the first state space information, the first action space information and a preset deep Q network includes:

[0014] Performing encoding processing on the first state space information to obtain encoded vector information;

[0015] The first state space information, the first action space information and the vector information are input into the preset deep Q network for training to obtain the trained deep Q network.

[0016] In the above solution, the first state space information includes first page element information and first path information; the encoding process of the first state space information to obtain the encoded vector information includes:

[0017] Determine a first vector based on the first page element information;

[0018] Encoding the first path information by one-hot encoding to obtain a second vector;

[0019] The first vector and the second vector are concatenated to obtain the encoded vector information.

[0020] In the above solution, the inputting the first state space information, the first action space information and the vector information into the preset deep Q network for training to obtain the trained deep Q network includes:

[0021] Determining reward parameters and action strategy parameters in the preset deep Q network; the reward parameters represent the degree of code coverage of the application during the stability test;

[0022] Determine an optimal action value function based on the reward parameter and the action strategy parameter;

[0023] Inputting the first state space information and the first action space information into the optimal action value function to perform state prediction to obtain predicted state information;

[0024] The preset deep Q network is trained by using the predicted state information, the first state space information, the first action space information and the vector information to obtain the trained deep Q network.

[0025] In the above scheme, the use of the predicted state information, the first state space information, the first action space information and the vector information to learn and train the preset deep Q network to obtain the trained deep Q network includes:

[0026] Optimizing the loss function in the preset deep Q network using the predicted state information, the first state space information, the first action space information and the vector information to obtain target parameters;

[0027] The trained deep Q network is determined based on the target parameter, the preset network training hyperparameter and the preset deep Q network.

[0028] In the above solution, the second state space information includes the current second page element information, the second control layout information and the second path information; the method further includes:

[0029] Determine executable operation information of the second page element information based on the second page element information, the second control layout information and the second path information;

[0030] The executable operation information of the second page element information is used as the second action space information.

[0031] In the above solution, the test report also includes an execution use case report; the step of inputting the second state space information and the second action space information into the trained deep Q network and outputting a test report corresponding to the application to be tested includes:

[0032] Inputting the second state space information and the second action space information into the trained deep Q network, and using the second path information to determine whether the application to be tested jumps out of a specified space or page when performing a user interface UI operation;

[0033] In a case where the application to be tested does not jump out of a specified space or page when executing a UI operation, the stability report and the execution use case report are determined.

[0034] In the above scheme, the method further comprises:

[0035] When the application to be tested jumps out of the specified space or page when executing UI operations, a preset debug bridge (Android Debug Bridge, ADB) tool is used to pull it back until the preset test time is reached to obtain the stability report and the execution use case report.

[0036] The present application also provides a testing device for application, including:

[0037] an obtaining unit, configured to obtain first state space information of the application during historical operation and first action space information corresponding to the first state space information;

[0038] A determination unit, configured to determine a trained deep Q network based on the first state space information, the first action space information, and a preset deep Q network; the preset deep Q network is a deep neural network based on a DQN algorithm;

[0039] an acquisition unit, configured to acquire second state space information of the application to be tested during the current running process and second action space information corresponding to the second state space information;

[0040] An output unit is used to input the second state space information and the second action space information into the trained deep Q network, and output a test report corresponding to the application to be tested; the test report includes a stability report on whether abnormal situations occur during the comprehensive traversal and testing of the user interface UI of the application to be tested.

[0041] The present application also provides an electronic device, including:

[0042] A memory for storing executable instructions;

[0043] The processor is used to implement any step of the above-mentioned method when executing the executable instructions stored in the memory.

[0044] The present application also provides a computer program product, including a computer program, which implements any step of the above method when executed by a processor.

[0045] The embodiment of the present application also provides a computer-readable storage medium storing executable instructions for implementing any step of the above-described method when executed by a processor.

[0046] The application testing method, device, electronic device, product and storage medium provided by the embodiments of the present application, wherein the method comprises: obtaining first state space information of the application in the historical operation process and first action space information corresponding to the first state space information; determining a trained deep Q network based on the first state space information, the first action space information and a preset deep Q network; the preset deep Q network is a deep neural network based on the DQN algorithm; obtaining second state space information of the application to be tested in the current operation process and second action space information corresponding to the second state space information; inputting the second state space information and the second action space information into the trained deep Q network, and outputting a test report corresponding to the application to be tested; the test report includes a stability report on whether abnormal situations occur during the process of comprehensively traversing and testing the user interface UI of the application to be tested. The solution of the embodiments of the present application is based on the first state space information of the application in the historical operation process. The trained deep Q network is determined by using time information, first action space information corresponding to the first state space information and a preset deep Q network; obtaining the second state space information of the application to be tested during the current operation and the second action space information corresponding to the second state space information; inputting the obtained second state space information of the application to be tested during the current operation and the second action space information corresponding to the second state space information into the trained deep Q network, and outputting a test report corresponding to the application to be tested; the test report includes a stability report on whether abnormal situations occur during the comprehensive traversal and testing of the UI of the application to be tested, that is, it integrates the test case historical information and source file information, fully explores the correlation between historical information and current source file information, uses neural networks to learn the relationship between multiple information, and evaluates the optimal actions of the application in different states to obtain the maximum test case coverage and complete the stability test. Compared with traditional methods, it is more flexible and can adapt to different scenarios, thereby realizing fine-grained stability testing. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 A schematic diagram of a test method flow for an application is provided for an embodiment of the present application;

[0048] Figure 2 A schematic diagram of another application test method flow is provided for the embodiment of the present application;

[0049] Figure 3 A schematic diagram of a deep Q network in a test method applied in an embodiment of the present application is provided;

[0050] Figure 4 This is a schematic diagram of the structure of a test device used in an embodiment of the present application;

[0051] Figure 5A schematic diagram of a hardware entity structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0052] The present application is further described in detail below in conjunction with the accompanying drawings and embodiments.

[0053] In the related art, Android application stability testing is a software testing method designed to evaluate and verify the stability and reliability of Android applications under various conditions. Stability testing mainly simulates user interaction operations through scripts, comprehensively traverses and tests the application's UI, and detects whether the application crashes, freezes, or becomes unresponsive during the test, thereby evaluating the stability of the application.

[0054] Traditional manual use case, random and model-based UI traversal testing methods are commonly used Android application stability testing methods. The traditional manual use case-based testing method is to manually write automated test scripts to cover the Android application code; the random-based testing method is to use the Monkey testing tool, adopt a randomization strategy, and randomly traverse and operate the application graphical user interface (GUI) to discover stability problems; the model-based testing method is to abstract the test activities into models, use different algorithms to automatically generate GUI test cases, and achieve code coverage and stability testing.

[0055] The disadvantages of the test method based on traditional manual use cases are high labor and time costs, cumbersome test case writing and difficulty in covering all possible test scenarios, resulting in hidden stability issues in the application not being discovered; the random test method is highly random and has poor results. During the test, the blank area on the page will be repeatedly operated, which will generate a large number of invalid inputs and greatly increase the time cost; the model-based test method has the problem that its model cannot fully and accurately reflect the UI operations of human users. The high complexity of the application will cause the state space to explode, and it is difficult to design a general model; and with the increase in the complexity of the application and the characteristics of vertical segmentation of the business, the existing methods cannot test on a set of pages specified in the application. The test process has no concept of boundaries, and the test granularity can only be the entire Android application. In short, the above existing methods are limited by the diversity of the test environment, the complexity of the application and the infinity of the state space, and it is difficult to cover the actual test scenarios and the fine-grained problems of the test object.

[0056] Based on this, the embodiment of the present application provides an application testing method, which is applied to electronic devices. The functions implemented by the method can be implemented by calling program codes by a processor in the electronic device. Of course, the program codes can be stored in a computer storage medium. It can be seen that the electronic device at least includes a processor and a storage medium. As an example, the electronic device can be a mobile phone, a computer, a terminal, an information transceiver device, a tablet device, a personal digital assistant, etc.

[0057] Figure 1 A schematic diagram of a test method flow for an application is provided for an embodiment of the present application; Figure 1 As shown, the method includes:

[0058] Step S101: obtaining first state space information of the application during historical operation and first action space information corresponding to the first state space information;

[0059] Step S102: determining a trained deep Q network based on the first state space information, the first action space information and a preset deep Q network; the preset deep Q network is a deep neural network based on a DQN algorithm;

[0060] Step S103: obtaining second state space information of the application to be tested during the current running process and second action space information corresponding to the second state space information;

[0061] Step S104: input the second state space information and the second action space information into the trained deep Q network, and output a test report corresponding to the application to be tested; the test report includes a stability report on whether abnormal situations occur during the comprehensive traversal and testing of the UI of the application to be tested.

[0062] It should be noted that the application can be determined according to actual conditions and is not limited here. As an example, the application can be an Android application.

[0063] The test method of the application can be determined according to the actual situation and is not limited here. As an example, the test method of the application can be an Android application stability test method, specifically, an Android application stability test method based on deep reinforcement learning.

[0064] In step 101, the first state space information can be determined according to actual conditions, which is not limited here. As an example, the first state space information includes first page element information, first control layout information and first path information.

[0065] The first action space information can be determined according to actual conditions and is not limited here. As an example, the first action space information can be understood as a set of actions to be performed. In practical applications, the set of actions to be performed can include clicking, long pressing, sliding, and inputting text.

[0066] The specific acquisition process of obtaining the first state space information of the application during the historical operation process and the first action space information corresponding to the first state space information can be determined according to actual conditions and is not limited here. As an example, the acquisition of the first state space information of the application during the historical operation process and the first action space information corresponding to the first state space information can be to collect the first state space information of the application during the historical operation process and the first action space information corresponding to the first state space information.

[0067] In step 102, the specific determination process in determining the trained deep Q network based on the first state space information, the first action space information and the preset deep Q network can be determined according to actual conditions and is not limited here. As an example, determining the trained deep Q network based on the first state space information, the first action space information and the preset deep Q network may include encoding the first state space information to obtain encoded vector information; inputting the first state space information, the first action space information and the vector information into the preset deep Q network for training to obtain the trained deep Q network.

[0068] In step 103, the application to be tested may be determined according to actual conditions, which is not limited here. As an example, the application to be tested may be an Android application to be tested.

[0069] The second state space information can be determined according to actual conditions and is not limited here. As an example, the second state space information can include current second page element information, second control layout information, and second path information.

[0070] The second action space information can be determined according to actual conditions and is not limited here. As an example, the second action space information can be understood as a set of actions to be performed. In practical applications, the set of actions to be performed can include clicking, long pressing, sliding, and inputting text.

[0071] In step 104, the test report includes a stability report on whether an abnormal situation occurs during the comprehensive traversal and testing of the user interface UI of the application to be tested; wherein the abnormal situation can be determined according to the actual situation and is not limited here. As an example, the abnormal situation may include whether the application crashes, freezes, or becomes unresponsive.

[0072] The specific process of inputting the second state space information and the second action space information into the trained deep Q network and outputting the test report corresponding to the application to be tested can be determined according to the actual situation and is not limited here. As an example, the test report also includes an execution use case report; the inputting the second state space information and the second action space information into the trained deep Q network and outputting the test report corresponding to the application to be tested may include inputting the second state space information and the second action space information into the trained deep Q network, and using the second path information to determine whether the application to be tested jumps out of the specified space or page when executing the user interface UI operation; when the application to be tested does not jump out of the specified space or page when executing the UI operation, determine the stability report and the execution use case report.

[0073] In practical applications, inputting the second state space information and the second action space information into the trained deep Q network and outputting the test report corresponding to the application to be tested can be understood as inputting the second state space information and the second action space information into the trained deep Q network for multiple traversal and learning processes until the test coverage converges, and outputting the test report corresponding to the application to be tested. That is, during the stability test, multiple traversal and learning processes are required until the test coverage converges and the test report corresponding to the application to be tested is output.

[0074] In an embodiment of the present application, a trained deep Q network is determined based on the first state space information of the application obtained during the historical operation process, the first action space information corresponding to the first state space information, and a preset deep Q network; the second state space information of the application to be tested during the current operation process and the second action space information corresponding to the second state space information are obtained; the second state space information of the application to be tested during the current operation process obtained and the second action space information corresponding to the second state space information are input into the trained deep Q network, and a test report corresponding to the application to be tested is output; the test report includes a stability report on whether abnormal situations occur during the comprehensive traversal and testing of the UI of the application to be tested, that is, it integrates the historical information of the test case and the source file information, fully explores the correlation between the historical information and the current source file information, uses the neural network to learn the relationship between multiple information, evaluates the optimal action of the application in different states, so as to obtain the maximum test case coverage, and completes the stability test. Compared with the traditional method, it is more flexible and can adapt to different scenarios, thereby realizing fine-grained stability testing.

[0075] In one embodiment, the first state space information includes first page element information, first control layout information and first path information; the method further includes:

[0076] Determine executable operation information of the first page element information based on the first page element information, the first control layout information and the first path information;

[0077] The executable operation information of the first page element information is used as the first action space information.

[0078] Among them, the first state space information, the first page element information, the first control layout information and the first path information can all be determined according to the actual situation, and are not limited here. As an example, the first state space information can be called the state space; the first page element information can be called the page element; the first control layout information can be called the control layout; the first path information can be called the Schema path. Among them, the state space can be denoted as S; each of the page elements can be denoted as e; the control layout can include controls such as text boxes, buttons, check boxes, drop-down lists, etc.; the control layout can be denoted as c; the Schema path can be denoted as s; the first state space information includes the first page element information, the first control layout information and the first path information; in actual applications, the application includes an agent, and when the test agent operates the Android application for t consecutive times, the state space S can be expressed as S={s1,s2,...,s t}.

[0079] The specific determination process of determining the executable operation information of the first page element information based on the first page element information, the first control layout information and the first path information can be determined according to actual conditions, which is not limited here. As an example, the executable operation information of the first page element information determined based on the first page element information, the first control layout information and the first path information can be the executable operation information of the first page element information inferred based on the first page element information, the first control layout information and the first path information. Among them, the executable operation information can be determined according to actual conditions, which is not limited here. As an example, the executable operation information may include operations such as clicking, long pressing, sliding and entering text.

[0080] The use of the executable operation information of the first page element information as the first action space information can be understood as the first action space information being the executable operation information of the first page element information; wherein, the first action space information can be called an action space; and the action space can be denoted as A.

[0081] In actual applications, the page elements, control layouts, and Schema paths of the Android application are collected during its operation, and the executable operations on the page elements are inferred, which are used as the state space and action space of the stability test agent. As an example, specifically, the state space is defined, and the Activity tree of the Android application is traversed. Each Activity tree is composed of multiple controls such as text boxes, buttons, check boxes, drop-down lists, etc. All page elements, control layouts, and Schema paths of the Android application are saved as the state space of the stability test process. The state space is a set of possible states of the agent during the test. Each page element is denoted as e, the control layout is denoted as c, and the Schema path is denoted as s. The test agent operates the Android application t times in a row, and the state space S can be expressed as S={s1,s2,...,s t}, where each state s t = [e,c,s,u] is defined by unique elements, control layout and Schema path. Define the action space. According to each state s in the test sequence t The saved UI elements can be clicked, long pressed, slid, and input text, and the action space A can be expressed as A = {a1, a2, ..., a t}. The action space is the set of actions that the agent can perform during testing.

[0082] In one embodiment, determining a trained deep Q network based on the first state space information, the first action space information, and a preset deep Q network includes:

[0083] Performing encoding processing on the first state space information to obtain encoded vector information;

[0084] The first state space information, the first action space information and the vector information are input into the preset deep Q network for training to obtain the trained deep Q network.

[0085] In this embodiment, the encoding process of the first state space information to obtain the encoded vector information can be determined according to actual conditions and is not limited here; as an example, the first state space information includes first page element information and first path information; the encoding process of the first state space information to obtain the encoded vector information may include determining a first vector based on the first page element information; encoding the first path information using a one-hot encoding method to obtain a second vector; and concatenating the first vector and the second vector to obtain the encoded vector information.

[0086] The inputting of the first state space information, the first action space information and the vector information into the preset deep Q network for training to obtain the specific training process in the trained deep Q network can be determined according to actual conditions and is not limited here; as an example, the inputting of the first state space information, the first action space information and the vector information into the preset deep Q network for training to obtain the trained deep Q network can include determining reward parameters and action strategy parameters in the preset deep Q network; the reward parameters characterize the degree of code coverage of the application during stability testing; determining the optimal action value function based on the reward parameters and action strategy parameters; inputting the first state space information and the first action space information into the optimal action value function for state prediction to obtain predicted state information; and using the predicted state information, the first state space information, the first action space information and the vector information to learn and train the preset deep Q network to obtain the trained deep Q network.

[0087] In one embodiment, the first state space information includes first page element information and first path information; and encoding the first state space information to obtain encoded vector information includes:

[0088] Determine a first vector based on the first page element information;

[0089] Encoding the first path information by one-hot encoding to obtain a second vector;

[0090] The first vector and the second vector are concatenated to obtain the encoded vector information.

[0091] It should be noted that the specific determination process of determining the first vector based on the first page element information can be determined according to actual conditions and is not limited here. As an example, the first page element information can be understood as the page element in the Android application testing process. Assuming that there are m elements, the first vector can be understood as an m-dimensional vector, which can be recorded as e=[e1,e2,…e m ]; The determination of the first vector based on the first page element information can be understood as writing the m elements in the Android application test process into an m-dimensional vector e=[e1,e2,…e m ].

[0092] The one-hot encoding method can be determined according to actual conditions and is not limited here. As an example, the one-hot encoding method can be one-host.

[0093] The first path information can be understood as a schema path with a length of n; the second vector can be understood as an n-dimensional vector, which can be recorded as u=[u1,u2…u n The first path information is encoded by one-hot encoding to obtain the second vector, which can be understood as encoding the schema path of length n into an n-dimensional vector represented by u=[u1, u2...u n ].

[0094] Concatenating the first vector and the second vector to obtain the encoded vector information can be understood as concatenating vectors e and u to obtain a vector s having a dimension of m+n.

[0095] In one embodiment, inputting the first state space information, the first action space information, and the vector information into the preset deep Q network for training to obtain the trained deep Q network includes:

[0096] Determining reward parameters and action strategy parameters in the preset deep Q network; the reward parameters represent the degree of code coverage of the application during the stability test;

[0097] Determine an optimal action value function based on the reward parameter and the action strategy parameter;

[0098] Inputting the first state space information and the first action space information into the optimal action value function to perform state prediction to obtain predicted state information;

[0099] The preset deep Q network is trained by using the predicted state information, the first state space information, the first action space information and the vector information to obtain the trained deep Q network.

[0100] In this embodiment, the reward parameters can be determined according to actual conditions and are not limited here. As an example, the reward parameters may include a new state reward, a feedback state reward, a discount reward, etc. Among them, the new state reward can be recorded as r t ; The reward of the feedback state can be recorded as X t ; The discount reward can be recorded as R t As an example, define the rewards and parameters. During execution, test the agent in s t Execution in state a t After the action, the new state s is obtained t+1 , reward r t It can refer to the following formula (1):

[0101]

[0102] When the code coverage rate increases positively during the Android application stability test, t The reward is 1. When the code coverage rate remains unchanged, the reward is r t When s is 0, the reward is -1 when the code coverage decreases negatively. t The action to be performed next in the state is recorded as a t+1 , the new state obtained is S t+1 . Test the agent on the Android application t The reward for executing UI actions and obtaining feedback at each step in a state can be recorded by referring to the following formula (2):

[0103] X t ={s t ,a t ,r t ,s t+1} (2);

[0104] Assuming that the decay coefficient of the agent performing the stability test is γ, the long-term cumulative discount reward can refer to the following formula (3):

[0105]

[0106] The action strategy parameter can be determined according to actual conditions and is not limited here. As an example, the action strategy parameter can be referred to as the action strategy, which can be recorded as π; the action strategy adopts the ε-greddy greedy strategy commonly used in Markov decision problems in reinforcement learning.

[0107] The specific determination process of determining the optimal action value function based on the reward parameter and the action strategy parameter can be determined according to the actual situation and is not limited here. Among them, the optimal action value function can be recorded as Q*(s,a).

[0108] For ease of understanding, an example is given here to design a DQN network. The DQN algorithm designed in this application uses a 4-layer fully connected network, including an input layer, a hidden layer, and an output layer, which is used to approximate the calculation of the Android application in state s t Next, perform action a t The Q value represents the optimal action value. The optimal Q value means that the code coverage executed during the Android application testing process is the largest, and the benefits obtained are better. The optimal action value function Q*(s,a) is defined as a function that follows the maximum total reward, and takes action a after observing the Android state s. The calculation process of Q*(s,a) can refer to the following formula (4):

[0109]

[0110] Where π is the corresponding action selection strategy, Q * (s, a) represents the optimal expected cumulative reward value of executing action a in state s. In this application, the action strategy adopts the ε-greddy greedy strategy commonly used in Markov decision problems in reinforcement learning.

[0111] The first state space information and the first action space information are input into the optimal action value function to perform state prediction, and obtain predicted state information; wherein the predicted state information can be determined according to actual conditions and is not limited here. As an example, the predicted state information may include a new state; the new state may be recorded as s t+1 In practical applications, an action a is selected based on the action strategy and the Q value predicted by the deep neural network. t , put a t Input into the environment to obtain a new state s t+1 and r.

[0112] The specific process of learning and training the preset deep Q network using the predicted state information, the first state space information, the first action space information, and the vector information to obtain the trained deep Q network can be determined according to actual conditions and is not limited here. As an example, learning and training the preset deep Q network using the predicted state information, the first state space information, the first action space information, and the vector information to obtain the trained deep Q network may include optimizing the loss function in the preset deep Q network using the predicted state information, the first state space information, the first action space information, and the vector information to obtain target parameters; determining the trained deep Q network based on the target parameters, preset network training hyperparameters, and the preset deep Q network.

[0113] In one embodiment, the using the predicted state information, the first state space information, the first action space information and the vector information to learn and train the preset deep Q network to obtain the trained deep Q network includes:

[0114] Optimizing the loss function in the preset deep Q network using the predicted state information, the first state space information, the first action space information and the vector information to obtain target parameters;

[0115] The trained deep Q network is determined based on the target parameter, the preset network training hyperparameter and the preset deep Q network.

[0116] In this embodiment, the specific optimization process of optimizing the loss function in the preset deep Q network using the predicted state information, the first state space information, the first action space information and the vector information to obtain the target parameters can be determined according to actual conditions and is not limited here.

[0117] As an example, the loss function in the DQN network training process can be set according to the following formula (5):

[0118]

[0119] Where i is the number of iterations, represents the expected value, ρ(s, a) is the probability distribution of sequences s and a, and θ i represents the target hyperparameter, y i To optimize the target value, refer to the following formula (6):

[0120]

[0121] Where r represents the reward obtained by the agent after performing action a, γ is the discount factor (between 0 and 1), It means that in the new state, according to the network parameters of the previous round (i-1 round), the Q value is calculated for all possible actions, and the maximum value is taken. Differentiating the loss function, the following gradient can be obtained by referring to the following formula (7):

[0122]

[0123] in Represents the loss function L i (θ i ) to find the gradient, θ i represents the network parameters of the network in round i training, It represents the mathematical expectation after sampling the state s under the state transition probability distribution ρ and sampling the action under the strategy ε. This expectation is used to measure the average direction of the policy gradient. Finally, the loss function is optimized by stochastic gradient descent.

[0124] The preset network training hyperparameter can be determined according to actual conditions and is not limited here. As an example, the preset network training hyperparameter can be understood as a set network training hyperparameter. The network training hyperparameter can be recorded as ω.

[0125] In practical applications, the DQN network is trained, the network parameters are initialized, and the gradient descent optimization algorithm is set. The network weights are continuously adjusted and optimized through back propagation to learn the optimal action execution strategy to maximize the expected value of the cumulative reward of the test agent performing the test. Specifically, the loss function in the DQN network training process is set according to the previous formula (5); where i is the number of iterations, ρ(s, a) is the probability distribution of the sequence s and a, and y i To optimize the target, refer to the previous formula (6); differentiate the loss function, and get the gradient by referring to the previous formula (7); finally, optimize the loss function through stochastic gradient descent. Set the network training hyperparameters. The network is trained for 100 epochs, and the batch size is set to 10 for each training. The loss function is calculated using mean square error. Set the learning rate of the DQN network training process to 0.001 and the attenuation coefficient γ to 0.9. Set the experience buffer. The process of the DQN network interacting with the Android application to calculate the Q value during the stability test is represented by f(s, a, ω), where s and a represent the state and action respectively, and ω represents the network hyperparameters. In the process of the DQN network interacting with the Android application, the state s of the next step of the program t+1 With the current status of the program tThere is a high correlation, and DQN may overfit and fail to converge. In order to break the data correlation and improve the neural network update efficiency and algorithm convergence effect, this application sets a 1GB experience buffer to store the four-tuple (s) consisting of state s, action a, reward r obtained by executing the work, and predicted Q value. t ,a t ,r t ,s t+1 ).

[0126] In one embodiment, the second state space information includes current second page element information, second control layout information and second path information; the method further includes:

[0127] Determine executable operation information of the second page element information based on the second page element information, the second control layout information and the second path information;

[0128] The executable operation information of the second page element information is used as the second action space information.

[0129] In this embodiment, the second state space information includes the current second page element information, second control layout information, and second path information. Among them, the second state space information, the second page element information, the second control layout information, and the second path information can all be determined according to actual conditions and are not limited here. As an example, the second state space information can be called a state space; the second page element information can be called a page element; the second control layout information can be called a control layout; and the second path information can be called a Schema path.

[0130] The specific determination process of determining the executable operation information of the second page element information based on the second page element information, the second control layout information, and the second path information can be determined according to actual conditions, which is not limited here. Among them, the executable operation information can be determined according to actual conditions, which is not limited here. As an example, the executable operation information can be understood as a set of executed actions. The executed action set may include clicking, long pressing, sliding, and entering text.

[0131] The executable operation information of the second page element information as the second action space information can be understood as the second action space information being the executable operation information of the second page element information; wherein the second action space information can be understood as a set of executed actions. In practical applications, the set of executed actions may include clicking, long pressing, sliding, and inputting text.

[0132] In one embodiment, the test report also includes an execution use case report; the inputting the second state space information and the second action space information into the trained deep Q network, and outputting a test report corresponding to the application to be tested, includes:

[0133] Inputting the second state space information and the second action space information into the trained deep Q network, and using the second path information to determine whether the application to be tested jumps out of a specified space or page when performing a user interface UI operation;

[0134] In a case where the application to be tested does not jump out of a specified space or page when executing a UI operation, the stability report and the execution use case report are determined.

[0135] In this embodiment, the second path information can be understood as the current schema path; the designated space can be determined according to the actual situation and is not limited here. As an example, the designated space can be understood as the designated schema space; the designated page can be determined according to the actual situation and is not limited here. As an example, the designated page can be a designated Schema page.

[0136] In one embodiment, the method further comprises:

[0137] When the application to be tested jumps out of the specified space or page when executing UI operations, a preset debug bridge ADB tool is used to perform pull-back processing until the preset test time is reached to obtain the stability report and the execution use case report.

[0138] In this embodiment, the preset test time can be determined according to actual conditions and is not limited here. As an example, the preset test time can be 3-5 days.

[0139] In actual applications, stability testing is performed. When the network converges, the network weights are saved to obtain the target DQN network f(s, a, ω′). When performing stability testing, each decision is based on the elements, controls, and schema paths of the current page of the Android application as the state s t , will s t Each action a in state k ={a1,a2,…,a k} and s t As input, we get Q for each action executed. tvalue, and then select the action with the highest Q value with probability 1-ε according to the ε-greddy greedy strategy, and select the random action with probability ε. Get the report. In the repeated test process, this application uses the directed graph model analysis method to walk the agent test towards walk = (s t ,a t ,s t+1 ) is constructed into a directed graph for analysis. When it is detected that the test agent repeats a certain test path infinitely, the agent test is interrupted to make it jump out of the current test direction. After each action is executed, it is also determined whether the current schema path is in the specified schema space. When the schema path meets the requirements, continue to observe and verify. If it does not meet the requirements, the application is pulled back through the ADB tool in the Andorid environment. Repeat the previous test steps until the test time reaches the predicted time, that is, the Android application stability test is completed, and finally the stability test report and execution case report are output.

[0140] For ease of understanding, an example is given here. The test method of the application can specifically be an Android application stability test method based on deep reinforcement learning. In order to solve the problems of state space explosion and fine-grained test objects, this application first collects page elements and schema path information on the Android application interface, adds page elements and schema paths to the state space, and fully mines the state information of Android applications. Then, a deep Q network is used to learn all the states, actions, and schema paths generated by the interaction between the program agent and the Android environment and approximate the optimal action decision for the calculation value. While meeting the maximum coverage of the program test, ensure that the Android environment is inside the specified application or on the page represented by the specified schema path.

[0141] This application first integrates the historical information of the test case and the source file information to fully explore the correlation between the historical information and the current source file information. Then, the neural network is used to learn the relationship between multiple information and adaptively adjust the weight of each information. The learned neural network can automatically sort the test case priority according to the input, which is more flexible than the traditional method and can adapt to different scenarios.

[0142] This application is based on deep reinforcement learning and proposes a method for testing the stability of Android applications based on deep reinforcement learning. This method regards the page elements and schema paths of Android applications as states, and records the executable operations of page elements as actions. The state is then encoded as a low-dimensional vector and input into the deep Q network to approximate the optimal Q value and train and update the network weights to maximize the stability test coverage. Through interaction and learning with Android applications, action execution decisions and value estimates are continuously optimized, and the optimal testing strategy is autonomously learned from experience to discover stability issues in Android applications. The following is combined with Figure 2 and Figure 3 The application principle of this application is described in detail. Figure 2 A schematic diagram of another application test method flow is provided for the embodiment of the present application; Figure 3 A schematic diagram of a deep Q network in a test method for application is provided in an embodiment of the present application. Specifically, the following steps are included:

[0143] Step 1, define the state space and action space: collect the page elements, control layout and Schema path during the running of the Android application, and infer the executable operations on the page elements as the state space and action space of the stability test agent respectively.

[0144] Step 2: Design a deep Q network DQN: Use it as a function approximation to calculate the expected value of the cumulative reward obtained by executing an action in a given state space and action space. The larger the expected value of the cumulative reward, the higher the code coverage of the test, indicating that the test is effective in exploring stability issues.

[0145] Step 3: Train the DQN network, initialize network parameters and set the gradient descent optimization algorithm, continuously adjust and optimize network weights through back propagation, and learn the optimal action execution strategy to maximize the expected value of the cumulative reward of the test agent executing the test.

[0146] Step 4: Execute the stability test action. Based on the response of the Android application, the stability test agent decides the next action to be performed. Repeat step 4 until the test is completed and output the stability test report.

[0147] The specific steps of step 1 above are as follows:

[0148] Step 1.1, define the state space, traverse the Activity tree of the Android application. Each Activity tree consists of multiple controls such as text boxes, buttons, check boxes, drop-down lists, etc. Save all page elements, control layouts and Schema paths of the Android application as the state space of the stability test process. The state space is the set of possible states of the agent during the test. Each page element is denoted as e, the space layout is denoted as c, and the Schema path is denoted as s. The test agent operates the Android application t times in a row, and the state space S can be expressed as S = {s1, s2, ..., s t}, where each state s = [e,u] is defined by a unique element, control layout, and Schema path.

[0149] Step 1.2, define the action space. According to each state s in the test sequence t The saved UI elements can be clicked, long pressed, slid, and input text, and the action space A can be expressed as A = {a1, a2, ..., a t}. The action space is the set of actions that the agent can perform during testing.

[0150] Step 1.3, define rewards and parameters. During execution, the test agent is tested in s t Execution in state a t After the action, the new state s is obtained t+1 , reward r t Please refer to the above formula (1).

[0151] When the code coverage rate increases positively during the Android application stability test, t The reward is 1. When the code coverage rate remains unchanged, the reward is r t When s is 0, the reward is -1 when the code coverage decreases negatively. t The action to be performed next in the state is recorded as a t+1 , the new state obtained is S t+1 . Test the agent on the Android application t The state and reward of executing UI actions at each step and obtaining feedback can refer to the above formula (2).

[0152] Assuming that the decay coefficient of the agent performing the stability test is γ, the definition of the long-term cumulative discounted reward can refer to the above formula (3).

[0153] The specific steps of step 2 above are as follows:

[0154] Step 2.1, define network input. The deep neural network designed in this application and the deep Q network (DQN) of Q-learning are used in the Android stability test framework. Figure 3 As shown in the figure, the page elements in the Android application testing process are assumed to have m elements, and all the elements are written as an m-dimensional vector e = [e1, e2, …e m For a schema path of length n, one-host encoding is used to encode the schema path into an n-dimensional vector represented as u = [u1, u2...u n ].

[0155] This application uses the page element e and the schema path m as the state s of the intelligent agent to perform stability testing, concatenates the vectors e and u, and obtains a vector s with a dimension of m+n.

[0156] Step 2.2, design the DQN network. The DQN algorithm designed in this application uses a 4-layer fully connected network, including an input layer, a hidden layer, and an output layer, which is used to approximate the calculation of the Android application in state s t Next, perform action a t The Q value represents the optimal action value. The optimal Q value means that the code coverage executed during the Android application testing process is the largest, and the benefits obtained are better. The optimal action value function Q*(s,a) is defined as a function that follows the maximum total reward and takes action a after observing the Android state s. The calculation process of Q*(s,a) can refer to the previous formula (4). Where π is the corresponding action selection strategy, Q * (s, a) represents the optimal expected cumulative reward value of executing action a in state s. In this application, the action strategy adopts the ε-greddy greedy strategy commonly used in Markov decision problems in reinforcement learning. According to the action strategy and the Q value predicted by the deep neural network, an action a is selected. t , put a t Input into the environment to obtain a new state s t+1 and r.

[0157] The specific steps of step 3 above are as follows:

[0158] Step 3.1, set the loss function during DQN network training by referring to the above formula (5).

[0159] Where i is the number of iterations, represents the expected value, ρ(s, a) is the probability distribution of sequences s and a, and θ i represents the target hyperparameter, y i To optimize the target value, refer to the above formula (6).

[0160] Where r represents the reward obtained by the agent after performing action a, γ is the discount factor (between 0 and 1), It means that in the new state, according to the network parameters of the previous round (i-1 round), the Q value is calculated for all possible actions, and the maximum value is taken. Differentiate the loss function and get the following gradient, which can refer to the above formula (7).

[0161] in Represents the loss function L i (θ i ) to find the gradient, θ i represents the network parameters of the network in round i training, It represents the mathematical expectation after sampling the state s under the state transition probability distribution ρ and sampling the action under the strategy ε. This expectation is used to measure the average direction of the policy gradient. Finally, the loss function is optimized by stochastic gradient descent.

[0162] Step 3.2, set the network training hyperparameters. The network is trained for 100 epochs, the batch size is set to 10 for each training, and the loss function is calculated using mean square error. The learning rate during DQN network training is set to 0.001, and the attenuation coefficient γ is set to 0.9.

[0163] Step 3.3, set the experience buffer. The process of DQN network interacting with Android application to calculate Q value during stability test is represented by f(s, a, ω), where s and a represent state and action respectively, and ω represents network hyperparameter. In the process of DQN network interacting with Android application, the state s of the next step of the program t+1 With the current status of the program t There is a high correlation, and DQN may overfit and fail to converge. In order to break the data correlation and improve the neural network update efficiency and algorithm convergence effect, this application sets a 1GB experience buffer to store the four-tuple (s) consisting of state s, action a, reward r obtained by executing the work, and predicted Q value. t ,a t ,r t ,s t+1 );

[0164] The specific steps of step 4 are as follows:

[0165] Step 4.1, perform stability test. When the network converges, save the network weights and get the target DQN network f(s, a, ω′). When performing stability test, each decision is based on the elements, controls and schema path of the current page of the Android application as the state st, and each action ak={a1, a2,…, ak} and st in the st state are used as input to get the Qt value of each action executed, and then select the action with the highest Q value with probability 1-ε according to the ε-greddy greedy strategy, and select a random action with probability ε.

[0166] Step 4.2, get the report. In the process of repeating the test in step 4.1, this application uses a directed graph model analysis method to construct the agent test walk = (st, at, st+1) into a directed graph for analysis. When it is detected that the test agent repeats a certain test path infinitely, the agent test is interrupted to make it jump out of the current test direction. After each action is executed, it is also determined whether the current schema path is in the specified schema space. When the schema path meets the requirements, continue to observe and verify. If it does not meet the requirements, the application will be pulled back through the ADB tool in the Andorid environment.

[0167] Step 4.3, repeat steps 4.1 and 4.2 until the test time reaches the predicted time, that is, the Android application stability test is completed, and finally the stability test report and execution case report are output.

[0168] This application proposes an Android application stability testing method based on deep reinforcement learning, which uses a deep Q network to learn and evaluate the optimal actions of Android applications in different states to maximize the test case coverage and complete the stability test.

[0169] A reinforcement learning method combining Android program page element information and schema path information is proposed. Two vectors are used to record the page element information and schema path information of the current Android application respectively, and the two vectors are fused through a fusion strategy. This vector combines program page information and schema information, so that the test agent can calculate the Q value in combination with schema path information when learning the optimal action decision, and ensure that the program will not jump out of the current Android application or the specified page, thereby realizing fine-grained stability testing.

[0170] A deep reinforcement learning Q network for Android application stability testing is designed. Android application page elements and schema paths are used as inputs of the deep Q network. Four fully connected layers and activation functions are used to approximate the Q values ​​corresponding to different actions in any state space of the Android application to solve the problem of infinite state space of complex Android applications. Then, the weight of each neuron is adjusted through the back-propagation algorithm to automatically learn the optimal decision in the process of Android program stability testing.

[0171] In order to improve the convergence speed of the deep Q network and the efficiency of the stability test execution, and avoid the program loop execution of repeated paths, the directed graph model is combined to dynamically monitor the test path in real time. When there is a loop in the directed graph model of the stability test path, the current path is immediately interrupted and backtracked until the loop is eliminated, and the next test path is executed.

[0172] The UI elements, page Schema path and executable actions of the Android application page are integrated to define the state space and action space of the Android application stability test problem, ensuring that the stability automation test program will not jump out of the set Schema space or page when performing UI operations. The prior art only uses UI elements and executable action information. If the Android application jumps out of a given page or even jumps out of the current application during the test, it will affect the subsequent test results. This application combines the Schema information of the current page to prevent the Android program from jumping out of the specified Schema space or page during the test.

[0173] Design and use a deep Q network to approximate the Q value of the UI action executed by the stability automation test program in a given state, and automatically adjust the weight parameters through the back-propagation gradient descent algorithm to guide the test agent on how to choose the optimal UI action in different states, so as to achieve the goal of maximizing the cumulative value of test case coverage. When existing model-based or random-based technologies are used for stability testing, it is often impossible to design a general model to deal with Android applications with high complexity and infinite state space. Through the training of a deep Q network, this application can adaptively interact and learn with Android applications, and has strong usability and robustness.

[0174] In order to solve the problem that the test agent repeatedly refreshes a certain page indefinitely during future tests, which affects the learning speed and effect of the deep Q network, this proposal can also improve the decision-making of executing actions. A directed graph model is constructed based on the test status and actions. When the test agent refreshes a certain number of times or reaches the upper limit of the page test time, it will jump out of the current state and return to normal testing or terminate the test.

[0175] This technology uses deep reinforcement learning technology in the field of artificial intelligence to implement an Android application stability testing method, which is in line with the current development trend of AI automated testing. Therefore, this technology has broad market application prospects.

[0176] In order to implement the method of the embodiment of the present application, the embodiment of the present application also provides an application test device 400, which is set on an electronic device, such as Figure 4 As shown, Figure 4 This is a schematic diagram of a test device structure used in an embodiment of the present application, including:

[0177] An obtaining unit 401 is used to obtain first state space information of the application during historical operation and first action space information corresponding to the first state space information;

[0178] A determination unit 402 is used to determine a trained deep Q network based on the first state space information, the first action space information and a preset deep Q network; the preset deep Q network is a deep neural network based on a DQN algorithm;

[0179] An acquisition unit 403 is used to acquire second state space information of the application to be tested during the current running process and second action space information corresponding to the second state space information;

[0180] The output unit 404 is used to input the second state space information and the second action space information into the trained deep Q network, and output a test report corresponding to the application to be tested; the test report includes a stability report on whether any abnormal situation occurs during the comprehensive traversal and testing of the user interface UI of the application to be tested.

[0181] Here, in one embodiment, the first state space information includes first page element information, first control layout information and first path information; the determination unit 402 is also used to determine the executable operation information of the first page element information based on the first page element information, the first control layout information and the first path information; and use the executable operation information of the first page element information as the first action space information.

[0182] Here, in one embodiment, the determination unit 402 is also used to encode the first state space information to obtain encoded vector information; input the first state space information, the first action space information and the vector information into the preset deep Q network for training to obtain the trained deep Q network.

[0183] Here, in one embodiment, the first state space information includes first page element information and first path information; the determination unit 402 is further used to determine a first vector based on the first page element information; encode the first path information using a one-hot encoding method to obtain a second vector; and concatenate the first vector and the second vector to obtain the encoded vector information.

[0184] Here, in one embodiment, the determination unit 402 is also used to determine the reward parameters and action strategy parameters in the preset deep Q network; the reward parameters characterize the degree of code coverage of the application during the stability test; the optimal action value function is determined based on the reward parameters and the action strategy parameters; the first state space information and the first action space information are input into the optimal action value function for state prediction to obtain predicted state information; the preset deep Q network is learned and trained using the predicted state information, the first state space information, the first action space information and the vector information to obtain the trained deep Q network.

[0185] Here, in one embodiment, the determination unit 402 is also used to optimize the loss function in the preset deep Q network using the predicted state information, the first state space information, the first action space information and the vector information to obtain target parameters; and determine the trained deep Q network based on the target parameters, preset network training hyperparameters and the preset deep Q network.

[0186] Here, in one embodiment, the second state space information includes the current second page element information, second control layout information and second path information; the determination unit 402 is also used to determine the executable operation information of the second page element information based on the second page element information, the second control layout information and the second path information; and use the executable operation information of the second page element information as the second action space information.

[0187] Here, in one embodiment, the test report also includes an execution use case report; the output unit 404 is also used to input the second state space information and the second action space information into the trained deep Q network, and use the second path information to determine whether the application to be tested jumps out of the specified space or page when executing the user interface UI operation; when the application to be tested does not jump out of the specified space or page when executing the UI operation, determine the stability report and the execution use case report.

[0188] Here, in one embodiment, the device 400 also includes a processing unit, which is used to use a preset debugging bridge ADB tool to perform pull-back processing when the application to be tested jumps out of the designated space or page when executing a UI operation, until a preset test time is reached, to obtain the stability report and the execution use case report.

[0189] It should be noted that: the application testing device provided in the above embodiment only uses the division of the above program modules as an example when testing the application. In actual applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device is divided into different program modules to complete all or part of the processing described above. In addition, the application testing device provided in the above embodiment and the application testing method embodiment belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0190] Based on the hardware implementation of the above-mentioned program module, an embodiment of the present application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, the steps in the application testing method provided in the above-mentioned embodiment are implemented.

[0191] An embodiment of the present application further provides a computer program product, including a computer program, which, when executed by a processor, implements the steps in the application testing method provided in the above embodiment.

[0192] Correspondingly, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the application testing method provided in the above embodiment are implemented.

[0193] It should be noted here that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.

[0194] It should be noted that Figure 5 is a schematic diagram of a hardware entity structure of an electronic device in an embodiment of the present application, such as Figure 5 As shown, the hardware entity of the electronic device 500 includes: a processor 501 and a memory 503 . Optionally, the electronic device 500 may also include a communication interface 502 .

[0195] It can be understood that the memory 503 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), and direct RAM bus random access memory (DRRAM, Direct Rambus Random Access Memory).The memory 503 described in the embodiments of the present application is intended to include but is not limited to these and any other suitable types of memories.

[0196] The method disclosed in the above embodiment of the present application can be applied to the processor 501, or implemented by the processor 501. The processor 501 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 501. The above processor 501 may be a general processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor 501 can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general processor can be a microprocessor or any conventional processor, etc. In combination with the steps of the method disclosed in the embodiment of the present application, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in the memory 503, and the processor 501 reads the information in the memory 503 and completes the steps of the above method in combination with its hardware.

[0197] In an exemplary embodiment, the device may be implemented by one or more application specific integrated circuits (ASIC), DSP, programmable logic device (PLD), complex programmable logic device (CPLD), field programmable gate array (FPGA), general processor, controller, microcontroller (MCU), microprocessor, or other electronic components to execute the aforementioned method.

[0198] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned sequence numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.

[0199] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0200] The methods disclosed in several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0201] The features disclosed in several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0202] The features disclosed in several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0203] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A testing method for application, characterized in that: include: Obtaining first state space information of the application during historical operation and first action space information corresponding to the first state space information; Determining a trained deep Q network based on the first state space information, the first action space information and a preset deep Q network; The preset deep Q network is a deep neural network based on the DQN algorithm; Acquire second state space information of the application to be tested during the current running process and second action space information corresponding to the second state space information; The second state space information and the second action space information are input into the trained deep Q network, and a test report corresponding to the application to be tested is output; the test report includes a stability report on whether abnormal situations occur during the comprehensive traversal and testing of the user interface UI of the application to be tested.

2. The method according to claim 1, characterized in that The first state space information includes first page element information, first control layout information and first path information; the method further includes: Determine executable operation information of the first page element information based on the first page element information, the first control layout information and the first path information; The executable operation information of the first page element information is used as the first action space information.

3. The method according to claim 1, characterized in that The determining a trained deep Q network based on the first state space information, the first action space information and a preset deep Q network includes: Performing encoding processing on the first state space information to obtain encoded vector information; The first state space information, the first action space information and the vector information are input into the preset deep Q network for training to obtain the trained deep Q network.

4. The method according to claim 3, characterized in that The first state space information includes first page element information and first path information; the encoding process of the first state space information to obtain the encoded vector information includes: Determine a first vector based on the first page element information; Encoding the first path information by one-hot encoding to obtain a second vector; The first vector and the second vector are concatenated to obtain the encoded vector information.

5. The method according to claim 3, characterized in that: The step of inputting the first state space information, the first action space information, and the vector information into the preset deep Q network for training to obtain the trained deep Q network includes: Determining reward parameters and action strategy parameters in the preset deep Q network; the reward parameters represent the degree of code coverage of the application during the stability test; Determine an optimal action value function based on the reward parameter and the action strategy parameter; Inputting the first state space information and the first action space information into the optimal action value function to perform state prediction to obtain predicted state information; The preset deep Q network is trained by using the predicted state information, the first state space information, the first action space information and the vector information to obtain the trained deep Q network.

6. The method according to claim 5, characterized in that The method of using the predicted state information, the first state space information, the first action space information, and the vector information to learn and train the preset deep Q network to obtain the trained deep Q network includes: Optimizing the loss function in the preset deep Q network using the predicted state information, the first state space information, the first action space information and the vector information to obtain target parameters; The trained deep Q network is determined based on the target parameter, the preset network training hyperparameter and the preset deep Q network.

7. The method according to claim 1, characterized in that The second state space information includes current second page element information, second control layout information and second path information; the method further includes: Determine executable operation information of the second page element information based on the second page element information, the second control layout information and the second path information; The executable operation information of the second page element information is used as the second action space information.

8. The method according to claim 7, characterized in that The test report also includes an execution use case report; the inputting the second state space information and the second action space information into the trained deep Q network, and outputting a test report corresponding to the application to be tested, including: Inputting the second state space information and the second action space information into the trained deep Q network, and using the second path information to determine whether the application to be tested jumps out of a specified space or page when performing a user interface UI operation; In a case where the application to be tested does not jump out of a specified space or page when executing a UI operation, the stability report and the execution use case report are determined.

9. The method according to claim 8, characterized in that The method further comprises: When the application to be tested jumps out of the specified space or page when executing UI operations, a preset debug bridge ADB tool is used to perform pull-back processing until the preset test time is reached to obtain the stability report and the execution use case report.

10. A testing device for application, characterized in that: include: an obtaining unit, configured to obtain first state space information of the application during historical operation and first action space information corresponding to the first state space information; A determining unit, configured to determine a trained deep Q network based on the first state space information, the first action space information, and a preset deep Q network; The preset deep Q network is a deep neural network based on the DQN algorithm; an acquisition unit, configured to acquire second state space information of the application to be tested during the current running process and second action space information corresponding to the second state space information; An output unit is used to input the second state space information and the second action space information into the trained deep Q network, and output a test report corresponding to the application to be tested; the test report includes a stability report on whether abnormal situations occur during the comprehensive traversal and testing of the user interface UI of the application to be tested.

11. An electronic device, characterized in that: include: A memory for storing executable instructions; A processor, configured to implement the test method for the application described in any one of claims 1 to 9 when executing the executable instructions stored in the memory.

12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the test method of the application according to any one of claims 1 to 9.

13. A computer-readable storage medium, characterized in that: Executable instructions are stored, and when executed by a processor, the test method for the application described in any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Application test method, electronic equipment and storage medium

    CN111694752A

  • Application testing method and system, equipment and computer readable storage medium

    CN113986769A

  • Mobile application cross-platform reinforcement learning traversal test technology based on depth image understanding

    CN114138653A

  • Intelligent software testing method based on DQN neural network

    CN116303007A

  • Deep reinforcement learning-based iOS application test method

    CN118585431A

Cited By

  • UI test script visual arrangement method and system based on reinforcement learning

    CN120670322A