Automated software testing

Through random exploration and machine learning models, the action space model is built, the action mode is identified, and the test machine is allocated for different types of tests, the problem of insufficient test coverage in undefined action space is solved, and efficient automated software testing is achieved.

CN119948497APending Publication Date: 2025-05-06MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380040094.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-20
Filing Date
2023-03-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing automated software testing technologies are difficult to effectively explore and understand undefined action space, resulting in insufficient test coverage and undetected potential vulnerabilities.

Method used

Through random exploration and machine learning models, models of action space are built, action patterns are identified that produce the desired results, and testing machines are allocated for random testing, playback testing and pioneering testing to fully explore action space.

Benefits of technology

It realizes efficient automation of software testing in undefined action space, improves test coverage and ability to discover vulnerabilities, and avoids test invalidity due to wrong action space or patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948497A_ABST
    Figure CN119948497A_ABST
Patent Text Reader

Abstract

The technology described herein provides an automated software testing platform that functions in an undefined action space. The techniques described herein begin from an undefined action space, but begin to understand the action space through random exploration. The actions taken during the test and the resultant status may both be communicated to a centralized test service. The techniques described herein also mine action telemetry data and status telemetry data to identify an action pattern that produces the sought result. Once multiple action modes are identified and at least a partial model of the action space is constructed, the test on the test machine can be divided into a random test mode, a replay test mode, and a pioneering test mode.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Automated software testing for complex environments, such as operating systems or applications running on them, should simulate a wide variety of ways that users interact with the software under test. Simulated use during testing allows vulnerabilities to be detected before they become usability or security issues after deployment. Simulated use attempts to test representative scenarios known to reveal vulnerabilities while providing enough variety to push the software under test and the operating system into a wide range of plausible states. Similarly, automated testing should test as many possible interaction scenarios as possible. Summary of the invention

[0002] The technology described herein provides an automated software testing platform that functions in an undefined action space. The action space is a collection of all actions that the testing platform may take on the tested software. In the undefined action space, the testing platform is unaware of the actions available to the testing platform outside of the current software state. A state is a description of the software and machine conditions at a certain point in time. For example, a software state includes visible interface objects. Interacting with the interface objects can produce a second state with different visible interface objects.

[0003] The techniques described herein start with an undefined action space and learn about it through random exploration. The random exploration is performed by multiple test machines running instances of the software under test. Both the actions taken and the resulting states are communicated to a centralized test service. The centralized test service can then use the action telemetry and state telemetry to build a model of the action space. The model of the action states is built by combining telemetry received from multiple test machines performing the test.

[0004] The technology described herein also mines action telemetry data and state telemetry data to identify action patterns that produce expected results. Action telemetry data describes input to a computer system, such as keyboard strokes, mouse clicks, or touch screen input. During use, users typically provide input. During testing, input is simulated. State telemetry data describes the state of the software and / or the machine on which the software is running. The expected result is defined by a specified state condition that reflects the achievement of the expected result. In one example, the expected state is an unhealthy state, such as a crash, hang, exception, assertion error, or the appearance of any error message. On the other hand, the expected state represents the completion of a scenario, such as the completion of a task within the software under test.

[0005] In various aspects, machine learning models are used to identify patterns within action telemetry data and state telemetry data that produce associated desired results. The pattern is identified from a large sequence of action and result state pairs. Actions and result states are described herein as events. When correctly identified, the pattern will include the actions required to produce the desired results, without irrelevant actions. Once identified, the pattern can be used for replay testing in the same or different versions of the software.

[0006] Once multiple action patterns are identified and at least part of the action space model is constructed, the tests on the test machine can be split into different patterns. The first part of the machine can be assigned to continue random testing, which serves an exploratory function. The second part of the machine can be assigned to replay testing, which seeks to replay the identified action pattern that produces the desired result. The third part of the machine can be assigned to pioneering testing. Pioneering testing performs random actions, but the random actions are directed to unexplored parts in the action space. Another part can be dedicated to running an action pattern generated by a calculated action pattern, which is derived from a heuristic method, a statistical method, or a machine learning algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The method and system disclosed herein will be described in detail below with reference to the accompanying drawings, in which:

[0008] Figure 1 is a block diagram of an example software testing environment in accordance with aspects of the techniques described herein;

[0009] Figure 2 is a diagram illustrating pattern identification in an example sequence of events in accordance with aspects of the techniques described herein;

[0010] Figure 3 is a diagram illustrating an example undefined action space in accordance with aspects of the techniques described herein;

[0011] Figure 4 is a block diagram illustrating an example pattern recognition system in accordance with aspects of the techniques described herein;

[0012] Figure 5 is a flow chart illustrating an example software testing method according to aspects of the technology described herein;

[0013] Figure 6 is a flow chart illustrating an example software testing method according to aspects of the technology described herein;

[0014] Figure 7 is a flow chart illustrating an example software testing method according to aspects of the technology described herein;

[0015] Figure 8 is a block diagram of an example distributed computing environment suitable for implementing aspects of the techniques described herein; and

[0016] Fig. 9 is a block diagram of an example computing environment suitable for implementing aspects of the techniques described herein. DETAILED DESCRIPTION

[0017] The technology described herein provides an automated software testing platform that works in an undefined action space. The action space is a collection of all actions that the test platform can take on the tested software. When testing user interface features, actions are interactions with user interface elements (e.g., buttons, menus). In the undefined action space, the test platform initially does not know the actions available to the test platform outside the current state. A state is a description of the software and machine conditions at a certain point in time. For example, a software state includes visible interface objects. In addition, in the undefined action space, the test platform also does not know the programming state changes (e.g., bold text, text input box to be typed, closing dialog box) that may result from taking available actions from the current state. This is in contrast to a variety of existing test systems, which require developers to provide a defined action space to facilitate testing.

[0018] Previous systems rely on heuristic-driven methods, such as generating random events or machine learning-based methods to navigate the user interface. Although these methods are adopted in testing, there are still limitations. For example, some machine learning-based models only utilize historically observed paths during testing, where in many cases, vulnerabilities or scenarios are found in previously unobserved paths. Existing machine learning models cannot effectively learn how to follow paths that are different from previously observed paths. The technology described in this article starts with an undefined action space, but begins to learn about the action space through random exploration. The random exploration is performed by multiple test machines running instances of the tested software. Each test machine includes a test agent that performs actions on the tested software. In the random exploration mode, the test agent receives information about the current state of the software. On the one hand, the state information is provided by accessibility functions embedded in the software and / or in the operating system running on the test machine. The state information of the current state of the software identifies the user interface elements that can be interacted with from the current state. The state information of the current state can also include the types of interactions that each component can receive. Then, the test agent randomly selects the user interface element to interact with, and if there may be multiple types of interactions, selects one type of interaction. The selected interaction type is then implemented on the selected user interface element to change the state of the software under test. Both the action taken and the resulting state are passed to the centralized test service. The action taken is described herein as part of the action telemetry data, and the resulting state is described as part of the state telemetry data. The centralized test service can then use the action telemetry data and the state telemetry data to build a model of the action space. The model of the action state is built by combining telemetry data received from multiple test machines performing the test.

[0019] The technology described herein also mines action telemetry data and state telemetry data to identify action patterns that produce expected results. Action telemetry data describes input to a computer system, such as keyboard strokes, mouse clicks, or touch screen input. During use, users typically provide input. During testing, input is simulated. State telemetry data describes the state of the software and / or the machine on which the software is running. The expected result is defined by a specified state condition that reflects the achievement of the expected result. In one example, the expected state is an unhealthy state, such as a crash, hang, exception, assertion error, or the appearance of any error message. In another example, the expected state represents the completion of a scenario, such as the completion of a task within the software under test.

[0020] The tasks may be taking a picture, entering text, bolding text, or any number of other possible tasks. Many of these tasks require a sequence of multiple actions. For example, a complex task may require opening a menu, selecting a menu item, providing input, and then selecting an input button. In contrast, simply opening a menu interface and then randomly closing the same user interface is an example of a scenario that cannot be completed. Likewise, the expected state is defined by the reward criteria, and the expected state is assigned a reward value when it is generated by the test agent.

[0021] Machine learning methods can be used to identify patterns within action telemetry and state telemetry that produce associated desired results. The pattern can be identified from a large sequence of action and result state pairs. Actions and result states are described as events in this article. When correctly identified, the pattern will include the actions required to produce the desired results, without irrelevant actions. Once identified, the pattern can be used for replay testing in the same or different versions of the software.

[0022] Replay testing is different from random testing. In replay testing, the test agent attempts to take actions associated with an identified pattern and determine what state will result. For example, when an action pattern is associated with the generation of an unhealthy state, replay seeks to reproduce this pattern to determine whether the vulnerability that produced the unhealthy state has been fixed. In other cases, the action pattern represents realistic actions that a real user might take. For example, a real user is unlikely to randomly open and close the same menu as in a random test. Instead, a real user might open a menu and attempt to perform actions that are performable through the menu. Attempting to complete realistic scenarios is generally considered a better use of testing resources because it provides coverage of scenarios that a user might attempt.

[0023] Once multiple action patterns are identified and at least a partial model of the action space is constructed, the test on the test machine can be divided into different modes. Once multiple action patterns are identified and at least a partial model of the action space is constructed, the test on the test machine can be divided into different modes. The first part of the machine can be allocated to continue random testing, which plays an exploratory role. The second part of the machine can be allocated to perform replay testing, which is intended to replay the identified action patterns that produce the desired results. The third part of the machine can be allocated to perform pioneering testing. Exploring testing performs random actions, but the random actions are aimed at the unexplored part of the action space. For example, available user interface elements that have not been interacted with by random testing before are selected as the starting point of pioneering testing. Exploring testing helps to ensure that all aspects of the tested software are covered. Additional parts can be dedicated to running action patterns generated by calculated action patterns, which are derived from heuristic methods, statistical methods, or machine learning algorithms.

[0024] In various aspects, replay testing is performed on a second version of the software in which a pattern of actions is identified. For example, when a vulnerability is detected during random testing of a first version of the software, the pattern that produced the vulnerability is identified. Presumably, the developer attempts to fix the vulnerability and releases a second version of the software. The replay test will then attempt to execute the pattern that revealed the vulnerability in the second version of the software.

[0025] The technology described in this article improves existing testing technology in several ways, including efficient use of computer resources. The goal is to find the most problems in the tested software using the least amount of resources. The technology described in this article takes the tested software and the reward criteria as the main inputs without defining the action space. In contrast, many existing technologies also require that the action space be defined as part of the input. Many existing technologies also receive various action patterns to guide the test. These inputs are usually provided for efficient use of test resources and to provide high efficiency. However, errors in the provided action space or action pattern will reduce the effectiveness of the current testing method because some areas of the software will be ignored (if part of the action space is missing), or if the action space includes parts that do not exist in the actual software, test errors will result. These input errors may occur as the various versions of the software progress without corresponding updates to the test input.

[0026] The techniques described herein do not need to use these same inputs to maintain high efficiency and effectiveness. Therefore, the techniques described herein also avoid ineffectiveness due to errors in the action space or action patterns that are typically provided as input. As described above, the techniques described herein learn action spaces and meaningful action patterns. As the system learns, more test resources can be used to test known spaces and known patterns, while fewer resources can be dedicated to exploration.

[0027] Automated testing environment

[0028] Now go to Figure 1 , according to one aspect of the technology described herein, an example software testing environment 100 is shown. In addition to other components not shown, the software testing environment 100 includes a test cloud 120 having a test machine (in Figure 1In the embodiment of the present invention, the present invention is abbreviated as TMA 122, test machine B 124 and test machine N 125 and test platform 130, all of which are connected through a computer network. It should be understood that this arrangement and other arrangements described herein are only set forth as examples. In addition to the arrangements and elements shown or as an alternative, other arrangements and elements (for example, machines, interfaces, functions, sequences and functional groupings) can also be used, and for the sake of clarity, some elements can be omitted completely. In addition, many elements described herein are functional entities, which can be implemented as discrete or distributed components or with other components, and can be implemented in any suitable combination and position. The various functions performed by one or more entities described herein can be performed by hardware, firmware, and / or software. For example, some functions can be performed by a processor that executes instructions stored in a memory.

[0029] Figure 1 Each component shown in FIG. 1 can be connected to any type of computing device (such as a Fig. 9 The components may be implemented by a computing device 900 described in the foregoing. These components may communicate with each other via a network, which may include one or more local area networks (LANs) and / or wide area networks (WANs). In an example implementation, the network includes the Internet and / or a cellular network, as well as a variety of possible public and / or private networks.

[0030] In addition, these components, functions performed by these components, or services performed by these components can be implemented on appropriate (multiple) abstraction layers (such as operating system layers, application layers, hardware layers, etc. of (multiple) computing systems). Alternatively or additionally, the functionality of these components and / or various aspects of the techniques described herein can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), etc. In addition, although the functionality is described herein for specific components shown in the software testing environment 100, it is contemplated that, in some aspects, the functionality of these components can be shared or distributed among other components.

[0031] The technology described herein includes a framework in which an agent interacts with multiple test machines (e.g., 30, 60, 100, 1000) simultaneously to collect test data. Each test machine 125 is pre-installed with an operating system and a software under test 126 (e.g., MICROSOFT WORD). A test agent 127 opens the software under test 126 (in Figure 1The test agent 127 observes the current state in the environment, takes action, and observes the next state.

[0032] The test cloud 120 includes a test machine A 122, a test machine B 124, and a test machine N 125. The N designation on the test machine N 125 is intended to indicate that any number of test machines may be used in the test cloud 120. Each test machine includes the software under test and a simulated computing environment, including an operating system. The test machine may be a virtual machine or a physical machine. The test director 144 may assign different test types to different machines. For example, a first group of machines may perform random walkthrough testing, while a second group of machines follows the action patterns predicted by the pattern detector 140 to complete the task. A third group of machines may perform pioneering exploration, which is directed to exploring areas of the action space that have not been previously explored and are therefore unknown.

[0033] The test platform 130 includes an action telemetry interface 132 , a state telemetry interface 134 , a reward component 136 , an event sequence builder 138 , a pattern detector 140 , an action space mapper 142 , and a test director 144 .

[0034] The action telemetry interface 132 receives action telemetry data 129 from test agents 127 running on multiple test machines. The action telemetry data 129 includes descriptions of actions taken by various test agents on the test machines (alternatively, simply described as "actions"). Actions include all possible interaction actions with the software interface. In other words, an action is any action (e.g., selection, hovering, entering text) that a user can perform with an interface element (e.g., a button, menu, text box). On the one hand, the action is determined by querying an accessibility layer (e.g., Microsoft User Interface (UI) Automation System). Applications such as screen readers use accessibility layers or functional frameworks for low vision users. The number of available actions for each state can be dynamic. Some software applications have very large action spaces. For example, some applications have 100,000 or more actions. During testing, the identified actions can be stored in a database. On the one hand, the identified actions are stored in a multi-model database service such as a key-value store.

[0035] The state telemetry interface 134 receives state telemetry data 128 from the software under test 126. The state telemetry data 128 includes new interface elements presented in response to actions and other changes made to the interface (e.g., content changes). The state telemetry data may also include system and software health information, such as whether the system crashes, hangs, etc. In some aspects, the state telemetry data takes the form of UI images generated by the action. Acquiring UI images requires a lot of resources, and collecting images for each test action may not be an effective approach. In various aspects, UI images are collected during action replay when the pattern being replayed is associated with a confidence score higher than a threshold score. Selectively collecting UI images saves resources required to capture and store UI images. In various aspects, the confidence score threshold can be 0.7, 0.8, and / or 0.9. The confidence score indicates the confidence of the pattern correctly identified by the pattern detector 140.

[0036] The reward component 136 evaluates the latest achieved state and assigns a reward. The reward is then associated with the state and the action that produced the state. The goal is to test the functionality that the user experiences in the application and operating system shell. The reward function can be formulated differently in different experiments. On the one hand, if the action taken by the agent matches the target action (for example, when the agent takes the action of clicking the bold button in the menu bar or clicking the font button) or achieves the target state, a positive reward is triggered. The actions and / or states and the associated rewards can be provided as training data.

[0037] The test system discerns when a desired state is achieved by comparing the new state to the reward criteria. In one example, the desired state represents a scenario completion, such as the completion of a task within the software under test. For example, a task may be taking a picture, entering text, bolding text, or any number of other possible tasks. As described above, many of these tasks require a sequence of multiple actions.

[0038] The event sequence builder 138 matches the action with the result state. The match is done by comparing the timestamp on the action with the timestamp on the state. The timestamp associated with the action and the timestamp associated with the result state may not match exactly. On the one hand, the result state is associated with the last action performed before the result state.

[0039] The pattern detector 140 identifies patterns that produce a desired state. An action pattern is the steps taken to perform a task, such as changing the font color to red. To change the font color to red, these steps may include opening a document, selecting text, opening a font menu, and selecting red from the available font colors. Note that different paths for performing the same task may be available. For example, in MICROSOFT WORD, the font color menu may be found on the task ribbon, automatically found by selecting text, or found through a drop-down menu. Therefore, there are at least three different font color change patterns available. During testing, it is desirable to find and test each pattern.

[0040] like Figure 4 As shown in , pattern detector 140 includes several different machine learning models that are trained to identify patterns that generate rewards. Different systems can be used to detect different types of patterns. For example, crash pattern detector 402 is trained to detect patterns that cause crashes, while hang pattern detector 404 is trained to detect patterns that cause hangs. State N pattern detector 406 is trained to detect state N, i.e., any other defined state. Each state can have its own detector, but a single detector can be trained to detect multiple states with similar patterns. Different machine learning models can be used to detect longer or shorter patterns. For example, natural language processing can be used to detect longer patterns, and machine learning models that evaluate action frequency and uniqueness can be used to detect shorter patterns. On the one hand, longer patterns require five or more actions to generate rewards. On the other hand, shorter patterns require four or fewer actions to generate rewards.

[0041] In an embodiment, a vectorization technique is used in a detection method for shorter patterns that indicates the importance of a specific action or pattern. For example, in a first detection method for shorter patterns, the technology described herein scores the importance of an action sequence to a reward by customizing a term frequency-inverse document frequency (TF-IDF) technique from the information retrieval domain. TF-IDF determines the frequency with which various actions and / or action patterns occur within an observed action. Actions that occur frequently but are not often associated with rewards tend to be excluded from the identified patterns, especially at the beginning of the pattern. Conversely, actions that are more frequently associated with rewards tend to be included. In addition, unique actions that occur near actions that generate rewards are more likely to be included in the identified patterns.

[0042] To learn longer action sequences, the techniques described herein can use word embedding models (e.g., Word2Vec, Global Vectors for Word Representation (GloVe), fastText) with continuous skip-grams to identify longer windows (e.g., up to 50 actions). Word embeddings are word representations in vector form that encode the meaning of words so that the meanings of words that are close in the vector space are expected to be similar. From there, the techniques described herein can look at the top actions for each reward and filter out the top five patterns that contain the most relevant actions. Longer action windows are required to perform some tasks. For example, a UI element may be inactive (e.g., grayed out) in a first user interface and therefore cannot be interacted with. In the example, activating the UI element may require navigating a settings menu and changing the settings, and then returning to the UI where the now active UI element is located.

[0043] Word embedding models can use neural network models to learn word associations from large amounts of text. Once trained, such models typically detect synonyms or suggest additional words for parts of sentences. The actions in action telemetry data are not words, nor are they "natural language." Despite this, natural language processing techniques are still suitable for identifying action patterns as described herein. Compared to traditional natural language evaluations, word embedding models can be configured to only evaluate actions that occur before the action that generates a reward. This configuration prevents the evaluation of tokens that occur after the action that generates the reward, because subsequent actions are not included in the pattern by definition. This contrasts with traditional natural language processing, in which words that appear after a particular word add context and meaning.

[0044] As a pre-processing step, actions or events (action / state combinations) are tokenized by an event tokenizer 408. The event tokenizer 408 assigns a token identifier to each action or event. The assigned token identifier is based on the properties of the action and / or the resulting state. Thus, similar actions receive similar (e.g., numerically close) token identifiers. For example, a first action of hovering the mouse over a first button should receive a similar token identifier as a second action of selecting the first button.

[0045] In this example, the training data includes labeled action patterns and action sequences that do not include meaningful patterns (e.g., negative examples). The labels represent the true values ​​of the training data (e.g., whether the training data is an action pattern that leads to a desired result or an action sequence that does not lead to a desired result). Action patterns and action sequences can be input as labeled identifiers with positive and negative labels.

[0046] In one or more embodiments, the machine learning model is trained by minimizing the loss (calculated by a loss function) between the training labels (e.g., action mode) and the actual predicted values ​​(e.g., non-action mode). Based on the loss determined by the loss function (e.g., mean squared error loss (MSEL), cross entropy loss, etc.), the training goal will be to reduce the prediction error over multiple batches or training periods so that, given the input, the neural network learns which features and weights indicate the correct inference.

[0047] In one or more embodiments, a neural network associated with the pattern detector 140 learns features from (multiple) training data inputs and applies weights to these features accordingly during training. A "weight" in the context of machine learning represents the importance or significance of a feature or feature value for a prediction. For example, each feature is associated with an integer or other real number, where the higher the real number, the more important the feature is for its prediction. In one or more embodiments, a weight in a neural network or other machine learning application represents the strength of a connection between nodes or neurons from one layer (input) to the next layer (output). A weight of 0 means that the input does not change the output, while a weight greater than 0 means that the output will change. The higher the input value or the closer the value is to 1, the more the output will change or increase. Similarly, there may also be negative weights. Negative weights will proportionally reduce the output value. For example, the more the input value increases, the more the output value will decrease. Negative weights will result in negative scores.

[0048] In one or more embodiments, after the neural network is trained, the (multiple) machine learning models (e.g., in a deployed state) receive one or more deployment inputs such as even sequences. In contrast to the training inputs, the deployment inputs are not labeled. When the machine learning model is deployed, it has typically been trained, tested, and packaged to process data it has never processed. In response, in one or more embodiments, the (multiple) deployment inputs are automatically converted into one or more labels, as described above. The machine learning model uses the labels to make predictions. In some embodiments, the (multiple) predictions are hard (e.g., class membership is a binary "yes" or "no") or soft (e.g., a probability or likelihood is attached to the label), which can be expressed as a confidence score.

[0049] Now back to Figure 1, the action space mapper 142 uses the action telemetry data and the state telemetry data 128 to understand the action space. The technology described herein starts with an undefined action space, but begins to learn about the action space through random exploration. When an action is taken during testing, both the action taken and the resulting state are passed to the centralized test service. The centralized test service can then begin to build a model of the action space using the action telemetry data and the state telemetry data. A model of the action state is built by combining the action and state telemetry data received from multiple test machines performing the test.

[0050] Test guide 144 assigns test tasks to various machines. In various aspects, tasks are assigned for a period of time (such as one hour), and then new test tasks are assigned. Once multiple action patterns are identified and at least a partial model of the action space is constructed, the test on the test machine can be divided into different modes. As previously mentioned, the test machine of the first part can be assigned to perform random testing, the second part is assigned to perform replay testing, and the third part is assigned to perform pioneering testing. Other parts can be devoted to running action scenarios generated by calculated action scenarios, which are derived from heuristic methods, statistical methods, or machine learning algorithms.

[0051] Various rules can be used to direct test resources to different test modes or different areas of the software. The test director 144 evaluates reward results from past runs and reduces the run time of branches that have no or few unique rewards in the past. This feature saves tester capacity. The test director 144 can evaluate rewards specific to the tested branch to focus efforts to reproduce rewards specific to that branch. On the one hand, as the number of rewards hit during random exploration decreases, the amount of resources allocated to random exploration decreases. Similarly, as the new action space discovered decreases, the amount of resources allocated to pioneering and / or random exploration decreases.

[0052] The technique described in this paper uses the learned system space to efficiently navigate the system while trying to achieve a state that satisfies the reward criteria. Telemetry data from the attempts is used to re-evaluate the model and retrain the technique described in this paper. The technique described in this paper starts with random exploration, and once it learns how to obtain rewards, it optimizes to focus on obtaining rewards.

[0053] Now go to Figure 2, according to aspects of the technology described herein, illustrates an example identification of patterns within an event sequence. As previously described, the test platform 130 receives action telemetry data and corresponding state telemetry data for multiple actions. The event sequence builder 138 associates specific actions with specific result states to form events. The result state is the state of the software under test directly after the action is taken and before the subsequent action is taken. The event sequence builder 138 can use timestamps associated with the action telemetry data and the state telemetry data to match specific actions with specific states to form events (or otherwise associate actions with states). The state is defined by a set of software and / or system properties and corresponding values.

[0054] Figure 2 Shown is a sequence of events 200. The sequence of events includes a first event 203, a second event 206, a third event 209, a fourth event 212, a fifth event 215, a sixth event 218, a seventh event 221, an eighth event 224, and a ninth event 227. These nine events may be only nine of hundreds, thousands, or more events recorded during testing.

[0055] The first event 203 includes a first action 201 and a first state 202. The first state 202 is a state generated by executing the first action 201. For example, if the first action 201 is to select a save icon, the first state 202 includes a "save interface" that is not displayed in the previous state. The second event 206 includes a second action 204 and a second state 205. The third event 209 includes a third action 207 and a third state 208. The third event 209 is also associated with the first reward 230. In various aspects, rewards are assigned to each state, and the rewards for the desired state are higher. In other aspects, rewards are assigned only when the state matches the desired state, such as a crash state, a suspended state, or other states indicating that the software memory is in a vulnerability or other problems. In other aspects, rewards are also assigned when the state matches the desired state, such as after completing a target task in the application (such as saving a file, taking a photo, or any other defined task that the tester may be particularly interested in).

[0056] The fourth event 212 includes the fourth action 210 and the fourth state 211. The fifth event 215 includes the fifth action 213 and the fifth state 214. The sixth event 218 includes the sixth action 216 and the sixth state 217. The seventh event 221 includes the seventh action 219 and the seventh state 220. The eighth event 224 includes the eighth action 222 and the eighth state 223. The eighth event 224 is associated with the second reward 232. The second reward 232 indicates that the eighth state 223 is the desired state. The ninth event 227 includes the ninth action 225 and the ninth state 226.

[0057] As previously described, the pattern detector 140 identifies a sequence of actions that produces a desired state. In this example, the first detected pattern 240 includes a first event 203, a second event 206, and a third event 209. The last event in the detected pattern is associated with a reward indicating that the desired state has been achieved. The challenge of detecting a sequence of actions that produces a desired state is to determine which action starts the sequence. The first 240 includes three events, but please note that the second detected pattern 250 only includes two events, and the fourth event 212, the fifth event 215, and the sixth event 218 are determined to be irrelevant to producing the eighth state 223. On the contrary, only the seventh action 219 and the eighth action 222 are required to produce the eighth state 223. In essence, these three excluded events are the results pursued by the test program, which are found to be tangents, and this does not produce the results sought. For example, the test program renamed the document during the fourth event 212, the fifth event 215, and the sixth event 218, which did not produce the desired state. However, the next two actions aimed at changing the document style produce a crash, which is the desired state to be detected during software testing. For the sake of illustration, assume that the rename and style change are independent and only the style change operation needs to be performed to generate the crash.

[0058] Now go to Figure 3 , according to various aspects of the technology described herein, an undefined action space is illustrated. The action space is a collection of actions that can be taken from different user interface states available in the software under test. In the defined action space, all available actions and the resulting states produced by taking the available actions are provided. In the undefined action space, the actions available to the test platform outside the current state are initially unknown to the test platform. In addition, in the undefined action space, the test platform also does not know the programming state changes (e.g., bold text, enter typed text into a box, close a dialog box) that should result from taking the available actions from the current state.

[0059] Action space 300A illustrates an undefined action space. Action space 300A includes a first state 302. First state 302 points to a user interface through which five different actions can be performed. These actions include a first action 301, a second action 304, a third action 307, a fourth action 310, and a fifth action 313. Note that the resulting state of taking any of these five actions is unknown.

[0060] Action space 300B illustrates what happens when a first action 301 is taken. In response to taking the first action 301, a second state 314 is generated. Three additional actions can be taken from the second state 314. The three additional actions include a sixth action 316, a seventh action 319, and an eighth action 322. As actions are taken, the techniques described herein can build an action space map. This is part of the learning process. The action space can then be used to run various scenarios during testing. For example, in a pioneering scenario, a client tester takes actions in an unknown portion of the action space with the goal of learning the new space.

[0061] See now Figure 5-Figure 7 Each block of the methods 500, 600, and 700 described herein includes a computing process that can be performed using any combination of hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in a memory. These methods can also be embodied as computer-usable instructions stored on a computer storage medium. The methods can be provided by a stand-alone application, a service, or a hosted service (standalone or in combination with another hosted service), to name a few. In addition, by Figure 1-Figure 4 Methods 500, 600, and 700 are described by way of example. However, these methods may additionally or alternatively be performed by any one system or any combination of systems, including the systems described herein.

[0062] Figure 5 A method 500 for testing software according to an aspect of the technology described herein is described. At step 510, the method 500 includes receiving action telemetry data that describes actions taken on a first version of the software. The action telemetry data is received from test agents running on multiple test machines. The action telemetry data includes descriptions of actions taken by various test agents on the test machines. The actions include all possible interaction actions with the software's (multiple) user interfaces. In other words, an action can be any action (e.g., selection, hovering, entering text) that a user can perform with a user interface element (e.g., a button, a menu, a text box). In one aspect, the action is determined by querying an accessibility layer (e.g., a Microsoft UI Automation system).

[0063] At step 520, method 500 includes receiving state telemetry data that describes the state of the first version of the software at a point in time during testing. The state telemetry data is received from the software on various test machines. The state telemetry data may include new user interface elements presented in response to the action and other changes made to the user interface (e.g., content changes). The state telemetry data may also include system and software health information, such as whether the system crashes, hangs, etc. In some aspects, the state telemetry data takes the form of UI images generated by the action. As described above, in various aspects, the UI images are collected during action replay when the replayed pattern is associated with a confidence score above a threshold.

[0064] At step 530, method 500 includes assigning a reward to a specific state within the state telemetry data that meets the reward criteria. The reward is associated with the state and the action that produced the state. The goal is to test the functionality as the user experiences the functionality in the application and the operating system shell. The reward function can be formulated in different ways in different experiments. In a first aspect, a positive reward is triggered if the action taken by the agent matches the target action (e.g., when the agent takes the action of "bold button", "click font") or achieves the target state.

[0065] At step 540, method 500 includes identifying an event by associating a corresponding action in the actions with a corresponding result state in the states, the corresponding result state being produced by the corresponding action. The action can be associated with the result state by timestamps on the action and result state.

[0066] At step 550 , method 500 includes generating a time series of the events.

[0067] At step 560, method 500 includes identifying an event pattern that produces a specific state within the time series of the event. The event pattern includes actions that produce the specific state, while excluding actions that are not required to produce the specific state. As previously described, various machine learning methods can be used to identify event patterns. These methods include TF-IDF and natural language processing, such as word embedding methods. Other pattern identification methods can be used to identify event time series. Suitable pattern identification methods include, but are not limited to, statistical pattern recognition, neural network pattern recognition, support vector machine pattern recognition, Bayesian pattern recognition, case-based pattern recognition, evolutionary computing pattern recognition, index vector pattern recognition, and random value optimization.

[0068] At step 570, method 500 includes storing the event pattern. The pattern can be stored for future replay testing and / or debugging work. For example, if the pattern is based on detecting an unhealthy condition, a developer can use the pattern to fix the cause of the unhealthy condition.

[0069] Figure 6 A method 600 of testing software according to an aspect of the technology described herein is described. At step 610, the method 600 includes receiving an event time sequence from testing a first version of the software. The events within the event time sequence include actions within the first version of the software and result states produced by the actions. The actions and states have been described previously. This information can be received from telemetry data of multiple test agents and the software instance under test.

[0070] At step 620, method 600 includes identifying event patterns within the event time series that produce specific states that meet specific reward criteria. As previously described, various machine learning methods can be used to identify event patterns. These methods include TF-IDF and natural language processing, such as word embedding methods.

[0071] At step 630, method 600 includes instructing the first plurality of test instances to test the second version of the software by reproducing the event pattern. This type of testing is described as replay testing. The goal of replay testing is to determine whether the vulnerability has been fixed. Replay testing performs tasks that might be performed by a real user.

[0072] Figure 7 A method 700 of testing software according to an aspect of the technology described herein is described. At step 710, the method 700 includes performing random walkthrough testing on a first version of the software within an undefined action space. Random walkthrough testing has been described previously. Generally speaking, random walkthrough testing randomly selects user interface elements to interact with and randomly selects the types of interactions to attempt.

[0073] At step 720, method 700 includes collecting test data from the random walkthrough test. The test data includes motion telemetry data and state telemetry data.

[0074] At step 730, method 700 includes identifying, from the test data, a pattern of actions that produces an unhealthy state in the first version of the software.The unhealthy state may be detected in the state data.

[0075] At step 740, method 700 includes replaying the action pattern during testing of a second version of the software to determine whether an unhealthy state is produced. The replay is used to determine whether the vulnerability has been fixed and to perform actions that may discover the vulnerability in a subsequent version.

[0076] Sample distributed computing environment

[0077] Reference now Figure 8 , Figure 8An example distributed computing environment 800 in which implementations of the present disclosure may be employed is illustrated. A data center may support a distributed computing environment 800 that includes a cloud computing platform 810, a rack 820, and nodes 830 (e.g., computing devices, processing units, or blades) in the rack 820. The system may be implemented with a cloud computing platform 810 that runs cloud services across different data centers and geographic regions. The cloud computing platform 810 may implement a fabric controller 840 component for provisioning and managing resource allocation, deployment, upgrades, and management of cloud services. Typically, the cloud computing platform 810 is used to store data or run service applications in a distributed manner. The cloud computing platform 810 in a data center may be configured to host and support the operation of endpoints for specific service applications. The cloud computing platform 810 may be a public cloud, a private cloud, or a dedicated cloud.

[0078] Node 830 can be provided with host 850 (for example, operating system or runtime environment), and this host runs the software stack defined on node 830. Node 830 can also be configured to perform special functionality (for example, computing node or storage node) in cloud computing platform 810. Node 830 is allocated to run one or more parts of the service application of tenant. Tenant can refer to the customer that utilizes the resources of cloud computing platform 810. The service application component of cloud computing platform 810 that supports specific tenant can be referred to as tenant infrastructure or tenant. The term "service application", "application" or "service" is used interchangeably here, and refers broadly to any software or software part that runs on top of data center or accesses storage and computing device locations in data center.

[0079] When node 830 supports more than one separate service application, node 830 can be divided into virtual machines (e.g., virtual machine 852 and virtual machine 854). Physical machines can also run separate service applications at the same time. Virtual machines or physical machines can be configured as personalized computing environments supported by resources 860 (e.g., hardware resources and software resources) in cloud computing platform 810. It is conceivable that resources can be configured for specific service applications. In addition, each service application can be divided into functional parts so that each functional part can run on a separate virtual machine. In cloud computing platform 810, multiple servers can be used to run service applications and perform data storage operations in a cluster. In particular, servers can perform data operations independently, but are disclosed as a single device called a cluster. Each server in a cluster can be implemented as a node.

[0080] The client device 880 may be linked to a service application in the cloud computing platform 810. The client device 880 may be any type of computing device, for example, it may correspond to a reference Fig. 9The computing device 900 described. The client device 880 can be configured to issue commands to the cloud computing platform 810. In an embodiment, the client device 880 can communicate with the service application through a virtual Internet Protocol (IP) and a load balancer or other means of directing communication requests to endpoints specified in the cloud computing platform 810. The components of the cloud computing platform 810 can communicate with each other through a network (not shown), which can include one or more local area networks (LANs) and / or wide area networks (WANs).

[0081] Sample computing environment

[0082] After briefly describing an overview of certain implementations of the present disclosure, an example operating environment in which embodiments of the present disclosure may be implemented is described below to provide a general background for various aspects of the present disclosure. Fig. 9 , an example operating environment for implementing embodiments of the present disclosure is shown and generally designated as computing device 900. Computing device 900 is only one example of a suitable computing environment and is not intended to suggest any limitation on the scope of use or functionality described herein. Neither should computing device 900 be interpreted as having any dependency or requirement relating to any one or combination of components illustrated.

[0083] The disclosed systems and methods may be described in the general context of computer code or machine-usable instructions, including computer-executable instructions (such as program modules) executed by a computer or other machine (such as a personal data assistant or other handheld device). In general, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs a specific task or implements a specific abstract data type. The disclosed systems and methods may be practiced in a wide variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The disclosed systems and methods may also be practiced in distributed computing environments, where tasks are performed by remote processing devices linked through a communications network.

[0084] refer to Fig. 9 , computing device 900 includes bus 910, which directly or indirectly couples the following devices: memory 912, one or more processors 914, one or more presentation components 916, input / output ports 918, input / output components 920, and an illustrative power supply 922. Bus 910 represents one or more buses (such as an address bus, a data bus, or a combination thereof). For conceptual clarity, Fig. 9 The various boxes of are represented by lines, and other arrangements of the components and / or component functionality are contemplated. For example, a presentation component such as a display device may be considered an I / O component. In addition, the processor has a memory. Thus, Fig. 9The figures illustrate only example computing devices that can be used in conjunction with one or more embodiments of the present disclosure. No distinction is made between categories such as "workstation," "server," "laptop," "handheld device," etc., as all of these are in the same Fig. 9 and reference to “computing device”.

[0085] The computing device 900 typically includes a variety of computer-readable media. Computer-readable media can be any available media that can be accessed by the computing device 900, including volatile and nonvolatile media, removable and non-removable media. For example, computer-readable media can include computer storage media and communication media.

[0086] Computer storage media includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer readable instructions, data structures, program modules, or other data. Computer storage media includes random access memory (RAM), read-only memory (ROM), electrically erasable programmable ROM (EEPROM), flash memory or other memory technology, compact disk ROM (CD-ROM), digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by the computing device 900. Computer storage media itself does not include signals.

[0087] Communication media typically embodies computer readable instructions, data structures, program modules, or other data in the form of a modulated data signal (such as a carrier wave or other transmission mechanism), and includes any information delivery media. The term "modulated data signal" means a signal whose one or more characteristics are set or changed to encode information in the signal. For example, communication media include wired media (such as a wired network or a direct wired connection) and wireless media (such as acoustic, RF, infrared and other wireless media). Any combination of the above is included within the scope of computer readable media.

[0088] Memory 912 includes computer storage media in the form of volatile and / or non-volatile memory. Memory can be removable, non-removable, or a combination thereof. Example hardware devices include solid-state memory, hard drives, optical drives, etc. Computing device 900 includes one or more processors that read data from various entities (such as memory 912 or I / O components 920). (Multiple) display components 916 display data indications to users or other devices. Example presentation components include display devices, speakers, printing components, vibration components, etc.

[0089] I / O ports 918 allow computing device 900 to be logically coupled to other devices including I / O components 920, some of which may be embedded. Illustrative I / O components include a microphone, joystick, game pad, satellite dish, scanner, printer, wireless devices, and the like.

[0090] End-to-end software-based systems can operate within system components to operate computer hardware to provide system functionality. At a low level, a hardware processor executes instructions selected from a given processor's machine language (also referred to as machine code or native) instruction set. The processor identifies native instructions and performs corresponding low-level functions associated with logic, control, and memory operations. Low-level software written in machine code can provide more complex functionality for more advanced software. As used herein, computer executable instructions include any software, including low-level software written in machine code, high-level software such as application software, and any combination thereof. In this regard, system components can manage resources and provide services for system functionality. Embodiments of the present disclosure contemplate any other changes and combinations thereof.

[0091] For example, the test environment may include an application interface (API) library that includes routines, data structures, object classes, and specifications for variables that can support the interaction between the hardware architecture of the device and the software framework of the test environment. These APIs include configuration specifications for the test environment so that different components therein can communicate with each other in the test environment, as described herein.

[0092] After determining the various components utilized herein, it should be understood that any number of components and arrangements can be used to achieve the functionality desired within the scope of this disclosure. For example, for conceptual clarity, the components in the embodiments shown in the accompanying drawings are represented by lines. Other arrangements of these components and other components can also be realized. For example, although some components are depicted as single components, many elements described herein can be implemented as discrete or distributed components or combined with other components, and are implemented in any suitable combination and location. Some elements can be omitted completely. In addition, the various functions performed by one or more entities described herein can be performed by hardware, firmware, and / or software, as described below. For example, various functions can be performed by a processor executing instructions stored in a memory. In this way, in addition to the arrangements and elements shown or as a substitute, other arrangements and elements (for example, machines, interfaces, functions, sequences, and functional groupings, etc.) can also be used.

[0093] The embodiments described in the following paragraphs may be combined with one or more of the specifically described alternatives. In particular, the claimed embodiments may contain alternative references to multiple other embodiments.

[0094] The subject matter of the disclosed embodiments is specifically described herein to meet statutory requirements. However, the description itself is not intended to limit the scope of this patent. On the contrary, the inventors have contemplated that the claimed subject matter may also be embodied in other ways in conjunction with other existing or future technologies to include different steps or combinations of steps similar to the steps described in this document. In addition, although the terms "step" and / or "box" may be used herein to imply different elements of the method employed, these terms should not be interpreted as implying any particular order between the various steps disclosed herein unless the order of the various steps is explicitly described.

[0095] For purposes of this disclosure, the term "including" has the same broad meaning as the term "comprising," and the term "accessing" includes "receiving," "referencing," or "retrieving." In addition, the term "communicating" has the same broad meaning as the term "receiving" or "transmitting," which is accomplished by a software or hardware-based bus, receiver, or transmitter using the communication media described herein. In addition, unless otherwise indicated to the contrary, words such as "a" and "an" include the plural and the singular. Thus, for example, the constraint of "feature" is satisfied when one or more features are present. In addition, the term "or" includes conjunctions, disjuncts, and both (thus, a or b includes a or b and a and b).

[0096] For the above detailed discussion, embodiments of the present disclosure are described with reference to a distributed computing environment; however, the distributed computing environment described herein is merely illustrative. Components may be configured to perform novel aspects of the embodiments, where the term "configured to" may refer to "programmed to" perform a specific task or implement a specific abstract data type using code. In addition, while embodiments of the present disclosure may generally relate to the test environment and schematics described herein, it should be understood that the technology may be extended to other implementation environments.

[0097] The embodiments of the present disclosure have been described with respect to specific embodiments, which are intended in all respects to be illustrative rather than restrictive. Alternative embodiments will become apparent to those skilled in the art to which the present disclosure pertains, without departing from its scope.

[0098] From the foregoing it will be seen that the system and method of the present disclosure are well adapted to attain all of the ends and objects set forth above, as well as the obvious advantages inherent in the structure.

[0099] It will be understood that certain features and subcombinations are of utility and may be employed without reference to other features or subcombinations. This is contemplated by the claims and is within the scope of the claims.

Claims

1. A method for automated software testing, comprising: receiving action telemetry data describing an action taken on the first version of the software; receiving state telemetry data describing a state of the first version of the software at a point in time during testing; Assigning a reward to a specific state within the state telemetry data that meets a reward criterion; identifying an event by associating a corresponding action among the actions with a corresponding result state among the states, the corresponding result state being produced by the corresponding action; generating a time series of said events; identifying a pattern of events within said temporal sequence of said events that produces said specific state; as well as The event pattern is stored.

2. The method of claim 1, wherein the motion telemetry data is collected during a random walkthrough of user interface elements generated by the first version of the software. 3 . The method of claim 2 , wherein the random walkthrough is performed in an undefined action space in which, when a first action is performed, a first result state produced by taking the first action is unknown. 4 . The method of claim 1 , wherein identifying the event patterns comprises using natural language processing that receives as input the events encoded as tokens.

5. The method of claim 4, wherein word embeddings are used in the natural language processing.

6. The method according to claim 1, wherein the method further comprises: The event model is run on a second version of the software. The method according to claim 1 , wherein the specific status is that the first version of the software crashes.

8. A computer system comprising: processor; as well as a memory configured to provide computer program instructions to the processor, the computer program instructions comprising a software testing platform, the software testing platform configured to: receiving a time sequence of events from testing a first version of software, wherein events within the time sequence of events include actions within the first version of the software and resultant states produced by the actions; identifying a pattern of events within the event time sequence that produces a specific state that satisfies a specific reward criterion; and The first plurality of test instances are directed to test the second version of the software by reproducing the event pattern.

9. The computer system of claim 8, wherein the actions and the resulting states in the event time sequence are collected by randomly walking an undefined action space within the first version of the software. 10 . The computer system of claim 8 , wherein the software testing platform is further configured to: instruct a second plurality of test instances to perform random walkthrough testing on the second version of the software.

11. The computer system of claim 8, wherein the software testing platform is further configured to: instruct a third plurality of test instances to perform an exploration test on the second version of the software, wherein the exploration test explores previously unexplored areas in an action space.

12. The computer system of claim 8, wherein the software testing platform is further configured to: construct an action space graph of the first version of the software using the event time sequence.

13. A computer system according to claim 8, wherein the software testing platform is further configured to: instruct the first plurality of test instances to generate an image of each user interface produced when the event pattern is reproduced when the event pattern is associated with a confidence score above a threshold, wherein the confidence score indicates the strength of the pattern identification.

14. The computer system of claim 8, wherein the action is a specific interaction with a user interface element.

15. The computer system of claim 8, wherein the specific state is a crash.