Distributed Android application automatic testing method based on reinforcement learning
Through the design of multi-agent parallel testing framework and state similarity-driven reward function, the problem of inefficiency in Android application testing is solved, efficient state space exploration and fault detection are achieved, and testing efficiency and coverage are improved.
Patent Information
- Application Number
- CN202510347138.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-11
AI Technical Summary
The existing technology is inefficient in Android application testing, making it difficult to effectively explore the state space and discover faults, and the testing efficiency of a single device is limited.
The multi-agent parallel testing framework is adopted to achieve efficient exploration and test case generation through distributed architecture design, state similarity-driven reward functions and regular experience aggregation mechanism.
The testing efficiency and fault detection capabilities have been significantly improved, the code coverage rate has been increased by 16.5%-34.3%, the fault detection rate has reached 100%, and the test time has been shortened by 21.3 minutes.
Smart Images

Figure CN120295910A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of reinforcement learning, Android application automated testing, etc., and specifically relates to a distributed Android application automated testing method based on reinforcement learning. Background Art
[0002] With the popularization of mobile devices, many users integrate mobile applications into their daily lives. The average daily usage time of users on mobile applications exceeds 2 hours. This data highlights the importance of ensuring the correctness of applications. In practice, programmers usually improve the correctness of applications through testing. Although many serious faults have been discovered through testing, testing Android applications is still challenging. For example, if an application contains complex business functions, multiple interfaces, and executable events, then generating test cases requires enumerating a large number of possible combinations of events and transitions.
[0003] To generate high-quality Android application test cases, researchers have proposed various methods, which can be roughly divided into random testing, model-based testing, and system testing. Although these methods have successfully discovered many faults, they also have inherent limitations. Random testing generates pseudo-random events, which may generate invalid events and may not be able to trigger complex events. Model-based testing uses dynamic or static strategies to build a model of the Android application and uses this model to guide the generation of test cases. However, due to the complex states and behaviors that Android applications may have, it is very difficult to build a complete application model. If the model is incomplete, the quality of the generated test cases may be affected. System testing uses complex techniques such as symbolic execution to guide the generation of test cases. The scalability of these methods is poor, and their effectiveness in actual development is often questioned.
[0004] In recent years, researchers have begun to use reinforcement learning to test Android applications. The goal of Android testing is to explore more states and discover hidden faults in the application. To achieve this goal, reinforcement learning adjusts the exploration strategy by using the rewards obtained during the interaction between the agent and the real environment (i.e., the Android application). The search space for these faults is huge and requires a lot of time to explore. Therefore, test efficiency is crucial for achieving effective testing within a given test time budget. However, existing work only uses a single device for testing, and the test efficiency is very limited. Summary of the Invention
[0005] In view of the defects and deficiencies of the above-mentioned existing technologies, the present invention proposes a distributed Android application automated testing method based on reinforcement learning. Multiple agents execute the following steps in parallel: 1) In the independent testing stage, each agent generates a state representation according to the Android application interface structure, selects actions using a curiosity-driven reinforcement learning strategy, records state transitions and reward values, and maintains a global state set to track the states discovered by all agents; 2) In the experience aggregation stage, the Q-tables and action execution counts of each agent are aggregated regularly, and a global policy is generated according to the principle of "priority by access count"; 3) In the policy synchronization stage, the aggregated global policy is distributed to all agents as the initial policy for the next test round, and the above process is iterated until the test terminates. This method realizes efficient exploration and test case generation through a distributed architecture design, a reward function driven by state similarity, and a regular experience aggregation mechanism.
[0006] The technical solution specifically adopted by the present invention to solve its technical problems is as follows:
[0007] A distributed Android application automated testing method based on reinforcement learning:
[0008] It is executed in parallel by multiple agents through the following steps:
[0009] a) Independent testing stage: Each agent interacts with the application under test within a preset round duration and performs the following operations:
[0010] Generate a state representation according to the current interface structure;
[0011] Select and execute actions based on the reinforcement learning strategy, and record state transitions and reward values;
[0012] Maintain a global state set to track the states discovered by all agents;
[0013] b) Experience aggregation stage: At the end of the preset round duration, aggregate the learning experiences of all agents to generate a global policy. The learning experiences at least include Q-tables and action execution frequencies;
[0014] c) Policy synchronization stage: Distribute the global policy to all agents as the initial policy for the next time period;
[0015] Repeat steps a)-c) until the test terminates.
[0016] Furthermore, the aggregation satisfies: for the same state-action pair, traverse the access counts of all agents, and select the Q value corresponding to the agent with the highest access count as the Q value of the global policy.
[0017] Furthermore, the calculation of the reward function of the reinforcement learning strategy includes:
[0018] If the Activity in the current state does not exist in the global state set or triggers an app crash, give a first positive reward;
[0019] If the current state exists in the global set and the maximum similarity with the state of the same Activity does not exceed the preset threshold, give a dynamic reward inversely proportional to the similarity;
[0020] If the maximum similarity between the current state and the state of the same Activity exceeds the preset threshold, give a negative reward.
[0021] Furthermore, the calculation of the maximum similarity includes:
[0022] Extract the Activity name of the current state and the tree - level structure of the interface components;
[0023] Calculate the component path matching degree and the component attribute difference, and sum them with weights.
[0024] Furthermore, the state representation is generated in the following way:
[0025] Parse the tree - level structure of the current interface components, and extract the component path, class name, coordinates, and size attributes;
[0026] Incorporate system - level events that can trigger abnormal behaviors into the action set, where the events include at least two of screen rotation, simulated physical key operations, and resource occupancy events.
[0027] Furthermore, step a) also includes:
[0028] When the number of action steps executed in a single exploration reaches the preset maximum value, force - restart the application under test;
[0029] When an app crash is detected, record the action sequence that has been executed in the current test trajectory as a failure case.
[0030] Furthermore, the preset round duration is 10% of the total test duration, and a 5 - second waiting delay is inserted after each action execution.
[0031] Furthermore, the action selection of the reinforcement learning strategy includes:
[0032] With a probability of 1 - ε, select the action with the highest Q - value from the valid actions in the current state;
[0033] With a total probability of ε, divide the random exploration actions into two categories: user - interface interaction events and system - level events for execution, where:
[0034] The system - level events include at least one of screen rotation, simulated physical key operations, and resource occupancy events;
[0035] The execution probability allocation ratio between user interface interaction events and system-level events is a preset value.
[0036] Furthermore, at the end of the test, the code coverage data of all agents are merged to generate a global coverage report.
[0037] Furthermore, the calculation of the matching degree of the component paths includes:
[0038] If the paths of two components are exactly the same and the parent node class names are the same, it is determined as a complete match;
[0039] If the paths are the same but the parent node class names are different, it is determined as a partial match;
[0040] Components with different paths do not participate in the similarity calculation.
[0041] And, an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, when the processor executes the program, it implements the steps of the above-mentioned distributed Android application automated testing method based on reinforcement learning.
[0042] A non-transitory computer-readable storage medium, on which a computer program is stored, when the computer program is executed by a processor, it implements the steps of the above-mentioned distributed Android application automated testing method based on reinforcement learning.
[0043] Compared with the prior art, the present invention and its preferred solutions significantly improve the test effect through a multi-agent collaborative testing framework and the following technical improvements:
[0044] Test efficiency optimization: The distributed architecture combined with the regular experience aggregation mechanism enables the code coverage convergence speed to increase by 16.5% - 34.3%, and the average coverage rate is 1.1% higher than that of the single-agent method (NA-DAE) (57.7% vs. 56.6%), and the test time is shortened by 21.3 minutes (compared with 32.4 minutes of D-M).
[0045] Enhanced fault detection ability: The curiosity-driven reward function based on state similarity effectively guides the agent to explore new states, the fault detection rate reaches 100% (optimal for 10 / 10 applications), and the total number of detected faults increases to 16, significantly higher than random testing (9) and distributed random testing (12).
[0046] State space control: Through the state abstraction of the interface structure (extracting attributes such as Activity, component path, coordinates, etc.) and similarity threshold filtering, the preprocessing time only accounts for 2.8% of the total test time, avoiding the state explosion problem.
[0047] Experiments show that, among 10 open-source Android applications, both the code coverage rate and the fault detection rate of this method are better than existing benchmark algorithms (such as Monkey, Stoat, etc.), and the test efficiency and coverage achieve an optimal balance. Brief Description of the Drawings
[0048] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments:
[0049] Figure 1 It is a schematic overview diagram of the agent testing steps in the embodiment of the present invention;
[0050] Figure 2 It is a schematic overview diagram of the aggregation step in the embodiment of the present invention;
[0051] Figure 3 It is a diagram showing the change of code coverage rate during the operation of the application in the embodiment of the present invention. Specific Embodiments
[0052] To make the features and advantages of this patent more obvious and understandable, specific embodiments are given below for detailed description as follows:
[0053] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0054] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0055] In view of the defects and deficiencies of the existing technology, the embodiment of the present invention proposes a distributed Android test framework based on reinforcement learning, DistributedAndroidExplore (DAE). The aim is to use multiple agents to simultaneously perform reinforcement learning-based testing on application programs, perform testing using reinforcement learning, and construct a distributed test framework to adaptively explore Android programs and generate high-quality test cases.
[0056] The implementation process of this solution mainly includes two steps: the testing of the application under test by each agent and the aggregation. In the agent testing step, the DAE distributes and executes a set of test agents, which execute independently of each other and periodically aggregate their learning experiences. The period of independent agent testing is called an "episode". At the end of each episode, all agents send their Q-tables and the number of action executions to the controller. In the aggregation step, the aggregator receives the data of each agent from the controller, calculates a Q-table representing the cumulative learning experience based on this data, and sends back the aggregated Q-table for the next episode of testing.
[0057] Agent Testing
[0058] Figure 1 Figure 1 shows an overview of the agent testing process, which can be divided into two main parts: 1) The left part shows the preprocessing module, which is responsible for implementing state abstraction and action analysis based on the Android application interface structure. 2) The right part is the curiosity-driven reinforcement learning module. All agents jointly maintain a global set of visited states, incorporate the Android application state similarity calculation method into the design of the reward function, and use this to guide the optimization of the Q-learning exploration strategy.
[0059] Algorithm 1 details the testing steps of the agent. During the given test period, the agent continuously explores the application under test (line 1). To avoid getting stuck in a non-continuable state transition, a maximum number of steps for the action sequence is set (line 5). Once the upper limit is reached, the agent restarts the application to start a new round of testing. All agents jointly maintain a global set of visited states M to record the visited states (line 4), and use this to guide the reinforcement learning to explore unvisited states, that is, curious states. The preprocessing module abstracts the initial Android application state and action set from the application under test (line 3). Each round of testing starts from the initial state, selects and executes an action according to the policy, the application under test jumps to the next state, and the access count table is updated simultaneously (lines 6 - 9). If a failure state is detected, the action sequence that caused the failure is added to the failure case set (lines 10 - 13). If no failure is detected, continue to abstract the next Android application state and action set (line 14). After the application under test transitions to a new state, calculate the similarity between the current state and the global set of visited states according to the Android application state similarity calculation method, calculate the reward, and update the state set simultaneously (lines 15 - 20). Then, train the reinforcement learning strategy based on this state transition (line 22). When the episode of agent testing ends, the agent is ready to aggregate its learning experience with other agents (lines 24 - 25). Finally, when the Android application testing ends, the DAE uses the Jacoco tool to calculate the code coverage in this test (line 29).
[0060] Algorithm 1 Agent Test Input: The Android application under test AUT, the total test duration T total , the test duration of each round T round , the maximum number of steps L of the action sequence, the similarity threshold threshold
[0061] Output: The set F of test cases that discover Android application failures, the code coverage P
[0062] Initialization: The reinforcement learning policy π, the global set of visited states The set of test cases that discover Android application failures The test case trajectory The currently executed time step t = 0, the access count table N = 0
[0063] 1.
[0064] 2.reset(AUT)
[0065] 3.s t ,A t ←preprocessing(AUT,M)
[0066] 4.M←M∪s t
[0067] 5.for each t∈[1,L] do
[0068] 6.a t ←getAction(π,s t )
[0069] 7.traj←traj.append(s t ,a t )
[0070] 8.failed←execute(AUT,a t )
[0071] 9.N(s t ,a t )=N(s t ,a t )+1
[0072] 10.if failed then
[0073] 11.F←F∪traj
[0074] 12.break
[0075] 13. end if
[0076] 14. s t+1 , A t+1 ← preprocessing(AUT, M)
[0077] 15. for s in M do
[0078] 16. similarity ← max(sim(s t+1 , s))
[0079] 17. end for
[0080] 18. if similarity < threshold then
[0081] 19. M ← M ∪ s t+1
[0082] 20. end if
[0083] 21. r t ← R(s t , a t )
[0084] 22. According to (s t , a t , s t+1 , r t ) train the reinforcement learning policy π
[0085] 23. s t = s t+1
[0086] 24. if timeout(T round ) then
[0087] 25. Transmit the information to the controller and wait for feedback
[0088] 26. end if
[0089] 27. end for
[0090] 28. end while
[0091] 29. The test is over. Obtain the code coverage P according to the application under test AUT
[0092] Preprocessing
[0093] To perform policy learning through reinforcement learning, the state representation must be defined. Although directly using the screen image (i.e., screenshot) of an Android application as the state representation is an intuitive way, considering the dynamic characteristics of Android applications, this method may lead to an explosive growth of the state space. For example, different interactions of users may cause changes in the screen image, even if these changes are limited to the visual level and do not involve business logic. It can also be observed that pages with the same business logic usually have similar interface structures. Therefore, in this embodiment, the state of the Android application is abstracted according to its interface structure. Specifically, the automated framework (UIAutomator) is used to analyze the structure of the current application interface, extract information about Activities and various view components, and thus abstract the Android application state s t =(act, e1, e2, …, e n ), where act represents the Activity of the current Android application interface, and e i ={path i , classname i , x i , y i , width i , height i} represents the view components in the interface, which are arranged in a tree structure, where path i represents the location path of the element e i in the Android interface structure, classname i represents the class name of e i , and x i , y i , width i , height i represent the abscissa, ordinate, width, and height of e i in the interface, respectively.
[0094] Algorithm 2 details the working process of the preprocessing module. The preprocessing module takes the AUT as the input and outputs the current state s t and the action set A t . During the testing process, all agents jointly maintain a global set M of visited states to record the visited states. First, this module uses the automated framework (UIAutomator) to extract the application interface structure information and maps it into an application state s t(Lines 1-2). Secondly, the module traverses and analyzes the attributes of each view control (e.g., clickable, scrollable, etc.), and infers and generates a set of executable actions accordingly (Line 3). In addition, this embodiment also incorporates system-level events on the mobile device (such as screen rotation, volume control, etc.) into the action set to reveal more potential application failures (Line 4). Finally, to avoid the infinite growth of the state space, if the similarity between the current state s t and the visited state s is greater than the set threshold, then return state s; otherwise, the new state s t is returned (Lines 5-11).
[0095] Algorithm 2 Preprocessing
[0096] Input: The Android application AUT to be tested, the global set M of visited states
[0097] Output: The current Android application state s t , the set A of valid actions in the current application state t
[0098] 1. activity, pageInfo ← dumpAndAnalysis(AUT)
[0099] 2. Create the Android application state s using <activity, pageInfo> t
[0100] 3. A t ← analysisActionWith UIAutomator(AUT)
[0101] 4. Add the system events of the mobile device to the action set A t
[0102] 5. for s in M do
[0103] 6. similarity = sim(s t , s)
[0104] 7. if similarity > threshold then
[0105] 8. return s, A t
[0106] 9. end if
[0107] 10. end for
[0108] 11. return s t , At
[0109] Curiosity-driven Reinforcement Learning
[0110] In a typical reinforcement learning task, there is usually a clear and quantifiable optimization goal, which provides a clear basis for the design of the reward function. However, in the field of Android application testing, it is not easy to design an effective reward function. This is because the core goal of Android application testing is to explore as many different behaviors in the application as possible, which is relatively vague and difficult to directly quantify. To address this challenge, it can be noted that Android application pages with different functions often have different interface structures. Therefore, this embodiment proposes an effective reward function design aimed at encouraging the agent to explore new application states.
[0111] Previous work proposed an Android application state similarity method that can calculate the similarity sim(s1, s2) between two Android states. Based on this method, this embodiment designs a curiosity-driven reward function to encourage the discovery of new states by introducing an adaptive exploration strategy. Specifically, during testing, all agents jointly maintain a global set M of visited states to record all visited states. When reaching state s t at this time, if the Activity of state s t is different from the Activities of all visited states in M, or s t is a failure state, it can be considered that a completely new state has been discovered, and a positive reward of 100 is given. Otherwise, the current state s t is compared with all states in M that have the same Activity in terms of similarity, and a reward is given according to the maximum similarity value: the higher the similarity, the lower the reward; if the similarity exceeds a preset threshold, it is considered that this state has been visited and its exploration value is relatively low, and at this time a negative reward of -100 is given to avoid repeated exploration. The curiosity-driven reward function is shown in Formula 1:
[0112]
[0113] where s t represents the current state, M represents the global set of visited states, the g(s) function represents the Activity of state s, crash represents the application failure state, is the maximum similarity between state s and the visited states in set M.
[0114] In the present invention, DAE adopts a model-free reinforcement learning method, Q-learning, to optimize the curiosity-driven reward-based policy. Q-learning evaluates the value of state-action pairs by maintaining a Q-function, which assigns a Q-value to each pair of state and action, representing the cumulative reward expected to be obtained from the current state after taking a specific action. Since both the states and actions in Android application testing are discrete, this embodiment uses a Q-table to represent the Q-function. At each discrete time step t, DAE updates the Q-table (Equation 2) based on the state transition information (s t , a t , s t+1 ) feedback by the environment and the reward value r t calculated by the curiosity-driven reward function, thereby dynamically adjusting the exploration policy π of reinforcement learning.
[0115]
[0116] Q-learning follows the ∈-greedy policy according to the Q-function to select the next valid action a t+1 in state s t+1 , and the action selection strategy is shown in Equation 3:
[0117]
[0118] where the action with the maximum Q-value is selected with a probability of 1 - ∈, and the UI event or system event is randomly selected as the next action with a probability of respectively.
[0119] Thus, in the Android application automated testing under the guidance of reinforcement learning, when a new state is discovered, a higher curiosity reward will be given to this action, thereby motivating the agent to preferentially explore the unvisited states and increasing the probability of discovering new states. DAE further incorporates the Android application state similarity into the reward function design, such that the higher the similarity of a state to the visited states, the lower its exploration value is considered and the corresponding curiosity reward is also smaller. This can encourage the agent to focus on exploring those more unique and potentially valuable states. In addition, when two agents first visit the same state successively, the agent that arrives first will receive a positive reward of 100, and this state will be immediately added to the global visited state set; while the agent that arrives later will receive a negative reward of -100 due to repeatedly visiting the known state. Such a setting effectively drives each agent to disperse exploration, avoids redundant search, and thereby improves the exploration efficiency and coverage of the entire system.
[0120] Aggregation
[0121] Figure 2This is an overview of the aggregation step. At the end of each round, all agents send their Q-tables and access count tables to the controller. After all agents have transmitted their data to the controller, the aggregator receives the data of each agent from the controller, calculates a Q-table representing the cumulative learning experience based on this, and sends back the aggregated Q-table for testing in the next round. Algorithm 3 presents the detailed process of the aggregation step. The aggregator takes the Q-table and access count table of each agent as input and outputs the aggregated Q-table. The aggregator follows the principle of "the minority obeys the majority". Regarding the learning experience of multiple agents for the same action, it trusts more the agent with a higher number of executions of that action (lines 1 - 6).
[0122] Algorithm 3 Aggregation
[0123] Input: Q-table Q of each agent 1,2,…,n , access count table N of each agent 1,2,…,n
[0124] Output: Aggregated Q-table
[0125] 1. for (s t , a t ) in Q 1,2,…,n do
[0126] 2. if N i (s t , a t ) ≥ N 1,2,…,n (s t , a t ) then
[0127] 3. Q(s t , a t ) = Q i (s t , a t )
[0128] 4. end if
[0129] 5. end for
[0130] 6. return Q
[0131] Experimental Setup
[0132] The test experiment implemented DAE based on Python 3.8, UIAutomator2, and Xposed. To verify the effectiveness of the method of the present invention, a research benchmark consisting of 10 real-world Android applications was constructed, and on this basis, the method of the present invention was compared and analyzed with existing algorithms. Most of the applications were selected from Q-testing. Since some applications were outdated or could not be compiled, not all applications were selected. Some actively maintained applications were collected from GitHub to supplement and form a new research benchmark. These 10 real Android applications have different orders of magnitude of executable code lines (ELOC), which can further demonstrate the test performance of the method proposed by the present invention on Android applications with different levels of complexity. For the scalability of the research, the method of the present invention directly conducts end-to-end tests on these Android applications without specific adjustments for each application. Based on this benchmark, the method of the present invention was deeply compared with existing Android application automated testing methods in terms of three metrics: code coverage, number of faults exposed, and testing efficiency.
[0133] In addition, in the experimental evaluation, to evaluate the effectiveness of the method of the present invention in detecting Android application faults, system-level faults reported by the console were collected and analyzed. It should be noted that user-level faults do not necessarily cause system-level crashes, which depends on the robustness of the application and the input validation mechanism. In addition, different Android applications may have different definitions of these user-level faults, making it difficult to distinguish. Therefore, the present invention mainly focuses on system-level faults that cause application crashes or abnormal behaviors. All detected faults were manually reviewed and confirmed to ensure that the identified anomalies were indeed actual problems rather than false alarms.
[0134] To evaluate the effectiveness of DAE, the following comparison methods and their variants were selected as the benchmark comparison objects for the study: 1) Monkey (M) generates a pseudo-random stream of user events through random interactions with the screen coordinates. To ensure the fairness of the experiment, Monkey was further optimized so that its action space is the same as that of DAE; 2) Distributed-Monkey (D-M) extends Monkey into a distributed system, and multiple devices simultaneously perform random tests on the application under test; 3) Stoat (St) is a model-based method. It first uses a random finite state machine model to describe the behavior of the automated execution program, and then uses MCMC sampling to guide the mutation of the model and generate test cases; 4) Distributed-Stoat (D-St) extends Stoat into a distributed system, and multiple devices simultaneously perform Stoat-based tests on the application under test; 5) NoAggregation-DAE (NA-DAE) is a variant of DAE that does not perform the aggregation step during testing.
[0135] For all experiments, the same time budget (i.e., 60 minutes) was allocated to each method. To ensure that the page is fully loaded after complex actions, a 5-second waiting interval was added after each action. In the experiment, DAE was configured with 2 agents, and the duration of each "round" was 6 minutes. According to the settings of previous work, each of the two stages of the Stoat method was defaultly allocated 30 minutes. To ensure fairness, the number of devices for D-M and D-St was also set to 2. To reduce the influence of random factors from a statistical perspective, all experiments were repeated 5 times, and the average value was used as the final result. In terms of parameter settings, the discount factor coefficient λ in reinforcement learning was uniformly set to 0.96; for the action selection strategy, the ∈-greedy strategy was adopted, and ε was set to 0.2. To maintain consistency, the Jacoco tool was used to calculate the code coverage of all test methods. For the distributed system, the Jacoco merge command was used to merge the coverage data on multiple devices to obtain a comprehensive coverage value, which was used as the coverage of the entire distributed system.
[0136] Experimental Results and Analysis
[0137] Table 1 shows the average instruction coverage of M, D-M, St, D-St, NA-DAE, and DAE tested on 10 open-source applications, where the best results are shown in bold with a gray background. Android applications are sorted according to the ELOC extracted from the Jacoco coverage report. Overall, the proposed method achieved the best code coverage in 8 / 10 Android applications (ties are counted as first). Followed by NA-DAE, which achieved the best results on 5 / 10 Android applications. In terms of average coverage, DAE achieved an average instruction coverage of 57.7%, higher than M (41.1%), D-M (53.4%), St (46.8%), D-St (48.0%), and NA-DAE (56.6%).
[0138] Table 1 also shows the number of fault detections of M, D-M, St, D-St, NA-DAE, and DAE on 10 open-source applications, where the best results are highlighted in bold with a gray background. Overall, the method proposed in the present invention detected the most application faults on 10 / 10 Android applications (ties are counted as first). DAE detected 16 Android application faults, higher than M (9), D-M (12), St (10), D-St (10), and NA-DAE (15). In addition, the discovered faults have been reported to the developers and partial confirmations have been obtained.
[0139] Comparison of Test Results in Table 1
[0140]
[0141] To evaluate the test efficiency, the code coverage was collected every 3 minutes within a 60-minute test time, and a curve graph of the change in code coverage of each Android application during the test was plotted, as Figure 3As shown. This figure shows the curve of code coverage changes of a distributed system (including D-M, D-St, NA-DAE, and DAE) in 10 applications (the horizontal axis represents the test time, and the vertical axis represents the code coverage). Generally speaking, the DAE method has the highest test efficiency in Android application testing. Its code coverage growth curve converges on average within 21.3 minutes, while D-M, D-St, and NA-DAE converge within 32.4 minutes, 26.4 minutes, and 25.5 minutes respectively. The test efficiency of DAE is improved by 34.3%, 19.3%, and 16.5% compared with D-M, D-St, and NA-DAE respectively. These results show that by regularly iterating and aggregating the learning experiences accumulated by all agents, DAE makes full use of the learning ability of each agent, thus significantly improving the test efficiency and the growth rate of code coverage. This not only proves the effectiveness of DAE but also highlights its advantages in the field of automated testing.
[0142] Further evaluate the performance loss caused by the management and maintenance of agent distribution and synchronization in the aggregation step. During the 60-minute test time, DAE consumes an average of 1.67 minutes in the aggregation step, resulting in a 2.8% performance loss. It is considered that the performance loss caused by DAE is within an acceptable range.
[0143] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions. Specifically, it is used to load and execute one or more instructions in the computer storage medium to implement the above method.
[0144] It should be further noted that, based on the same inventive concept, the present invention also provides a computer storage medium, on which a computer program is stored, and when the computer program is run by a processor, it executes the above method. The storage medium can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, apparatus, or device.
[0145] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meanings understood by those with ordinary skills in the field to which the present invention belongs. The "first", "second", and similar terms used in the present invention do not indicate any order, quantity, or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right", etc. are only used to represent relative position relationships, and when the absolute position of the object being described changes, the relative position relationship may also change accordingly.
[0146] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.
[0147] This patent is not limited to the above-mentioned optimal implementation mode. Anyone inspired by this patent can come up with various other forms of distributed Android application automated testing methods based on reinforcement learning. All equivalent changes and modifications made within the scope of the patent application of this invention shall fall within the scope covered by this patent.
Claims
1. A distributed Android application automated testing method based on reinforcement learning, characterized in that: Multiple agents execute the following steps in parallel: a) Independent testing stage: Each agent interacts with the application under test within a preset round duration and performs the following operations: Generate a state representation based on the current interface structure; Select and execute an action based on the reinforcement learning strategy, and record the state transition and reward value; Maintain a global state set to track the states discovered by all agents; b) Experience aggregation stage: At the end of the preset round duration, aggregate the learning experiences of all agents to generate a global policy, and the learning experiences at least include the Q-table and the action execution frequency; c) Policy synchronization stage: Distribute the global policy to all agents as the initial policy for the next time period; Repeat steps a)-c) until the test terminates.
2. The distributed Android application automated testing method based on reinforcement learning according to claim 1, wherein: The aggregation satisfies: For the same state-action pair, traverse the access times of all agents, and select the Q value corresponding to the agent with the highest access times as the Q value of the global policy.
3. The distributed Android application automated testing method based on reinforcement learning according to claim 1, characterized in that: The calculation of the reward function of the reinforcement learning strategy includes: If the Activity of the current state does not exist in the global state set or triggers an application crash, give a first positive reward; If the current state exists in the global set and the maximum similarity with the state of the same Activity does not exceed the preset threshold, give a dynamic reward inversely proportional to the similarity; If the maximum similarity between the current state and the state of the same Activity exceeds the preset threshold, give a negative reward.
4. The distributed Android application automated testing method based on reinforcement learning according to claim 3, characterized in that: The calculation of the maximum similarity includes: Extract the Activity name and the tree-like hierarchical structure of the interface components of the current state; Calculate the component path matching degree and the component attribute difference, and perform a weighted sum of the two.
5. The distributed Android application automated testing method based on reinforcement learning according to claim 1, characterized in that: The state representation is generated in the following way: Parse the tree-like hierarchical structure of the current interface components, and extract the component path, class name, coordinates, and size attributes; Incorporate system-level events that can trigger abnormal behaviors into the action set, and the events include at least two of screen rotation, simulated physical key operations, and resource occupancy events.
6. The distributed Android application automated testing method based on reinforcement learning according to claim 1, characterized in that: Step a) further includes: When the number of action steps executed in a single exploration reaches the preset maximum value, forcefully restart the application under test; When an application crash is detected, record the executed action sequence in the current test trajectory as a failure case.
7. The distributed Android application automated testing method based on reinforcement learning according to claim 1, characterized in that: The preset round duration is 10% of the total test duration, and a 5-second waiting delay is inserted after each action execution.
8. The distributed Android application automated testing method based on reinforcement learning according to claim 1, characterized in that: The action selection of the reinforcement learning strategy includes: Select the action with the highest Q - value from the valid actions of the current state with probability 1 - ε; Execute the random exploration action as two types of events, namely user - interface interaction events and system - level events, with a total probability of ε. Among them: The system - level events include at least one of screen rotation, simulated physical key operation, and resource - occupancy event; The execution probability distribution ratio between user - interface interaction events and system - level events is a preset value.
9. The distributed Android application automated testing method based on reinforcement learning according to claim 1, characterized in that: At the end of the test, merge the code - coverage data of all agents to generate a global coverage report.
10. The distributed Android application automated testing method based on reinforcement learning according to claim 4, characterized in that: The calculation of the matching degree of the component path includes: If the paths of two components are exactly the same and the parent - node class names are the same, it is determined as a complete match; If the paths are the same but the parent - node class names are different, it is determined as a partial match; Components with different paths do not participate in the similarity calculation.
Citation Information
Cited By
Multi-robot full-coverage path planning method based on heuristic Q-Learning
CN121916925A
Distributed software self-healing maintenance method and system based on deep reinforcement learning
CN122111728A