Software performance analysis method, device and equipment
By defining the execution process and calling the performance analysis tool, the software performance analysis is automated, the problems of manual dependence and low efficiency in the existing technology are solved, and the analysis efficiency and intelligence are improved.
Patent Information
- Application Number
- CN202311652202.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-04
- Publication Date
- 2025-06-06
AI Technical Summary
The existing software performance analysis methods rely on manual experience, have high professional capabilities, low efficiency and insufficient intelligence.
By defining the execution process and calling the corresponding performance analysis tools, we automatically analyze the execution order and conditions of the behavior, and realize automated analysis and detection of software performance.
Reduced manual participation, improved analysis efficiency and intelligence, reduced requirements for professional capabilities, and enabled automated testing and optimization suggestions.
Smart Images

Figure CN120104442A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a software performance analysis method, device and equipment. Background Art
[0002] Software performance analysis is an important part of the software development process. It is a technology that determines the software performance bottleneck by analyzing the information of the software running process, so as to determine whether the software needs to be optimized.
[0003] The commonly used software performance analysis method is that the personnel select the corresponding analysis tool from the performance analysis tools provided by the computer operating system according to one or more performances to be analyzed by the software, and manually start the collection of the index data required by these analysis tools. Then, these started analysis tools can detect and analyze the corresponding index data collected from the running process of the software to obtain the analysis results of the corresponding performance of the software.
[0004] However, this software performance analysis method relies on the experience of personnel, which not only requires high professional ability of personnel, but also has low performance analysis efficiency. Summary of the invention
[0005] The embodiments of the present application provide a software performance analysis method, device, electronic device, computer storage medium and computer program product, which can improve the intelligence of software performance analysis and reduce manual dependence.
[0006] In a first aspect, an embodiment of the present application provides a software performance analysis method, which may include calling an execution process for performing performance analysis on an application under test, wherein the execution process includes an execution order between multiple analysis behaviors and an execution condition of each analysis behavior; during the execution process, a performance analysis tool related to the corresponding analysis behavior is called according to the execution conditions achieved, and one analysis behavior is associated with at least one performance analysis tool; target data is processed by the called performance analysis tool to obtain a return result, wherein the target data includes performance indicator data exhibited during the operation of the application under test; after obtaining any return result, according to the execution condition satisfied by any return result, the performance analysis tool related to the next analysis behavior is continued to be scheduled in the execution order to process the target data until the execution process ends and the performance analysis result is obtained.
[0007] In this embodiment, the computing device can perform performance analysis on the local tested application or the tested application running on other devices or device clusters. The computing device can predefine the execution process through a decision tree, and then call the execution process to run when the tested application is subjected to performance analysis. Specifically, the computing device automatically calls the relevant performance analysis tool to obtain the target data generated during the operation of the tested application and perform statistical analysis according to the execution process (the corresponding execution condition is the start condition) to obtain the target data generated during the operation of the tested application and to obtain the return result, and then calls the performance analysis tool related to the next analysis behavior according to the execution condition satisfied by the analysis result. In this way, until the return result of the last analysis behavior is obtained, the final performance analysis result is formed, without the need for human participation, which reduces the professional ability requirements for personnel and has high efficiency of automated testing.
[0008] In some possible examples, calling an execution process for performance analysis of the application under test includes: obtaining startup parameters, the startup parameters including identification information of the application under test; and calling the execution process to start a performance analysis operation on the application under test according to the startup parameters.
[0009] In this example, the computing device can obtain the startup parameters input by the personnel through the interactive interface. The startup parameters can be used to turn on, turn off or switch specified functions during the software testing process, but are not limited to this. For example, they are used to specify the application under test and / or give the performance analysis tool the ability to track the threads (processes) of the application under test, thereby ensuring that the application under test runs according to the specified startup parameters. In this way, the performance analysis of the specified application under test can be performed using a general execution process, and the execution process can be customized to perform performance analysis on the test application. In addition, if some mature performance analysis tools used in the field need to be called in the execution process, but the original input and output formats of the tool are different from the formats defined in the execution process, these mature performance analysis tools can be reused to perform software performance testing by converting the input and output formats of the tool, without the need to redevelop new tools, which helps to reduce development costs.
[0010] In some possible examples, the returned result includes feature data, and the feature data is used to characterize the process running status or system status that forms the returned result. The target data is processed by the called performance analysis tool to obtain the returned result, including: obtaining the target data from the kernel by the called performance analysis tool; performing statistics on the target data according to the statistical dimension of the performance analysis tool to obtain statistical results; extracting feature data from the target data; and forming the returned result with the statistical results and feature data. Among them, the target data includes the performance indicator data required by the performance analysis tool, and the performance indicator data is generated during the operation of the application under test.
[0011] In this example, the performance analysis tool can be a performance analysis tool based on BPF technology, but is not limited thereto. The performance analysis tool can be non-invasively injected into the application under test based on the BPF framework, so that during the operation of the application under test, the performance analysis tool can collect the required performance indicator data (including kernel status information) from the kernel through the interface provided by the kernel or other interfaces for statistics, and extract the corresponding feature data together with the statistical results to form a return result, which is conducive to obtaining richer test data.
[0012] In some possible examples, the performance analysis results include return results and target optimization suggestions; according to the execution conditions satisfied by each return result, the next performance analysis tool related to the analysis behavior is continuously scheduled in the execution order to process the target data until the execution process ends and the performance analysis results are obtained, including:
[0013] According to the execution conditions satisfied by each returned result, continue to schedule the next analysis behavior-related performance analysis tool in the execution order to process the target data until the return result of the last analysis behavior-related performance analysis tool in the execution process is obtained; according to the return result of the last analysis behavior-related performance analysis tool, match the corresponding target optimization suggestion from the suggestion library, the suggestion library includes multiple optimization suggestions, and the optimization suggestions are used to optimize the performance of the application. Each optimization suggestion is associated with at least one return result, and the target optimization suggestion is at least one optimization suggestion in the suggestion library; the return result of the last analysis behavior-related performance analysis tool and the target optimization suggestion are combined to form a performance analysis result.
[0014] In this example, the performance analysis results obtained can also include optimization suggestions. In this way, when analyzing the root cause of the bottleneck point of the tested application, optimization suggestions for solving the root cause can also be given to improve the automation level of software testing. In some examples, when an optimization suggestion can be obtained based on a certain return result, the process of the corresponding branch can be terminated. On the contrary, if no effective suggestion can be obtained, the next analysis action will be defined in the execution process until the optimization suggestion is obtained.
[0015] In some possible examples, the target data includes stack information generated during the running of the application under test, and the stack information points to the code location of the application under test. As a specific example, the stack information can be data information in the function call stack, so as to locate the code address of the application under test. According to the execution condition satisfied by any return result, the next performance analysis tool related to the analysis behavior is continuously scheduled in the execution order to process the target data until the execution process ends and the performance analysis results are obtained, including:
[0016] According to the execution conditions satisfied by any return result, continue to schedule the next analysis behavior-related performance analysis tool in the execution order to process the target data until the return result of the last analysis behavior-related performance analysis tool in the execution flow is obtained; when the return result of the last analysis behavior-related performance analysis tool does not meet the test pass condition, determine the code location pointed to by the stack information from the target data processed by each performance analysis tool called in the execution flow; and form a performance analysis result with the return result and code location of the last analysis behavior-related performance analysis tool.
[0017] In this example, the obtained performance analysis results may also include the location of the code to be optimized. In this way, when analyzing the root cause of the bottleneck point of the tested application, the specific location of the code to be optimized can be efficiently located, making it easier for personnel to modify and optimize the tested application code.
[0018] In some possible examples, after obtaining the performance analysis result, the method further includes: displaying the performance analysis result and / or each returned result according to a predefined display rule.
[0019] In some possible examples, the execution process includes a time overhead ratio analysis behavior to analyze the running time ratio and idle waiting ratio during the running of the tested application; in the process of performing the execution process, according to the execution conditions achieved, the corresponding performance analysis tool related to the analysis behavior is called, including: according to the execution process, when the idle waiting ratio reaches a preset threshold, the idle cause analysis tool related to the idle cause analysis behavior is called to determine the waiting reason for the idle waiting of the tested application; and, when the running time ratio reaches a preset value, the hot function analysis tool related to the hot function analysis behavior is called to determine the resource occupancy during the running of the tested application.
[0020] In this example, by converting the software performance problem into a time problem, because the limitations of various resources will eventually affect the time, this example mainly analyzes the time cost of the application under test. If the idle waiting time is high, we can further analyze why we have to wait (idle reasons), whether it is reasonable, and whether there is room for improvement. We will analyze layer by layer until we find the cause. Similarly, if the normal operation time accounts for a high proportion, we can further analyze whether the CPU pipeline is fully utilized and whether there is room for optimization. Secondly, we can also analyze whether the compiled instructions are optimal, whether the CPU has been fully utilized, whether vectorization is performed, etc. In this way, by analyzing from top to bottom, layer by layer, it is conducive to quickly and efficiently analyzing the performance bottleneck of the application under test.
[0021] In the second aspect, an embodiment of the present application provides a software performance analysis device, which includes: a performance analysis model, which is used to define an execution process for performance analysis of an application under test, wherein the execution process includes an execution order between multiple analysis behaviors and an execution condition of each analysis behavior; an execution framework, which is used to load and run a performance analysis module, and in the process of executing the execution process, according to the execution conditions achieved, call a performance analysis tool related to the corresponding analysis behavior, and one analysis behavior is associated with at least one performance analysis tool; a performance analysis tool, which is used to process target data and obtain a return result, wherein the target data includes performance indicator data exhibited during the operation of the application under test; the execution framework is also used to, after obtaining any return result, continue to schedule the next analysis behavior-related performance analysis tool in accordance with the execution order to process the target data according to the execution conditions satisfied by any return result, until the execution process ends and the performance analysis result is obtained.
[0022] In some possible examples, the device also includes a visualization module, which is used to obtain startup parameters, and the startup parameters include identification information of the application under test; the execution framework is also used to call the execution process to start a performance analysis operation on the application under test according to the startup parameters.
[0023] In some possible examples, the returned result includes feature data, and the feature data is used to characterize the process operation status or system status that forms the returned result; the performance analysis tool is specifically used to: obtain target data from the kernel by calling the performance analysis tool; perform statistics on the target data according to the statistical dimensions of the performance analysis tool to obtain statistical results; extract feature data from the target data; and form a return result with the statistical results and the feature data.
[0024] In some possible examples, the performance analysis results include return results and target optimization suggestions; the execution framework is specifically used to: according to the execution conditions satisfied by any return result, continue to schedule the next analysis behavior-related performance analysis tool in the execution order to process the target data until the return result of the last analysis behavior-related performance analysis tool in the execution process is obtained; according to the return result of the last analysis behavior-related performance analysis tool, match the corresponding target optimization suggestion from the suggestion library, the suggestion library includes multiple optimization suggestions, the optimization suggestions are used to optimize the performance of the application, each optimization suggestion is associated with at least one return result, and the target optimization suggestion is at least one optimization suggestion in the suggestion library; the return result of the last analysis behavior-related performance analysis tool and the target optimization suggestion are combined to form a performance analysis result.
[0025] In some possible examples, the target data includes stack information generated during the running of the application under test, and the stack information points to the code location of the application under test; the execution framework is specifically used to: according to the execution conditions satisfied by any return result, continue to schedule the next analysis behavior-related performance analysis tool in the execution order to process the target data until the return result of the last analysis behavior-related performance analysis tool in the execution process is obtained; when the return result of the last analysis behavior-related performance analysis tool does not meet the test pass condition, determine the code location pointed to by the stack information from the target data processed by each performance analysis tool called in the execution process; and form a performance analysis result with the return result and code location of the last analysis behavior-related performance analysis tool.
[0026] In some possible examples, after obtaining the performance analysis result, the method further includes: the device further includes a visualization module, and the visualization module is used to: display the performance analysis result and / or each returned result according to a predefined display rule.
[0027] In some possible examples, the execution process includes a time overhead ratio analysis behavior to analyze the running time ratio and idle waiting ratio during the running of the tested application; the execution framework is specifically used to: according to the execution process, when the idle waiting ratio reaches a preset threshold, call the idle cause analysis tool related to the idle cause analysis behavior to determine the waiting reason for the idle waiting of the tested application; and, when the running time ratio reaches a preset value, call the hot function analysis tool related to the hot function analysis behavior to determine the resource occupancy during the running of the tested application.
[0028] In a third aspect, an embodiment of the present application provides an electronic device, comprising: at least one memory for storing programs; and at least one processor for executing programs stored in the memory; wherein, when the program stored in the memory is executed, the processor is used to execute the method described in the first aspect or any possible implementation of the first aspect.
[0029] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method described in the first aspect or any possible implementation of the first aspect.
[0030] In a fifth aspect, an embodiment of the present application provides a computer program product, characterized in that when the computer program product runs on a processor, the processor executes the method described in the first aspect or any possible implementation of the first aspect.
[0031] In the sixth aspect, an embodiment of the present application provides a chip, characterized in that it includes at least one processor and an interface; at least one processor obtains program instructions or data through the interface; at least one processor is used to execute program line instructions to implement the method described in the first aspect or any possible implementation method of the first aspect.
[0032] It can be understood that the beneficial effects of the second to sixth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a schematic diagram of possible application scenarios of the embodiments of the present application;
[0034] Figure 2 It is a structural diagram of a software performance analysis device provided in an embodiment of the present application.
[0035] Figure 3 It is a schematic diagram of the execution flow of the performance analysis model definition in a specific example of the present application;
[0036] Figure 4 A schematic diagram of a software performance analysis device performing application performance analysis in a specific example of the present application;
[0037] Figure 5 This is a schematic diagram of the structure of a toolset module in a specific example of this application;
[0038] Figure 6 It is a flowchart of a software performance analysis method provided in an embodiment of the present application;
[0039] Figure 7 is a flowchart of a software performance analysis method in a specific embodiment of the present application;
[0040] Figure 8 This is a schematic diagram of data information collected by a performance analysis tool in a specific example of this application;
[0041] Fig. 9 It is a schematic diagram of test data presented by a visualization module in a specific example of this application;
[0042] Fig.10 It is a schematic diagram of test data presented by a visualization module in a specific example of this application;
[0043] Fig.11 It is a schematic diagram of test data presented by a visualization module in a specific example of this application;
[0044] Fig.12 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application;
[0045] Fig.13 It is a schematic diagram of the structure of a chip provided in an embodiment of the present application. DETAILED DESCRIPTION
[0046] The term "and / or" in this article is a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The symbol " / " in this article indicates that the associated objects are in an or relationship, for example, A / B means A or B.
[0047] The terms "first" and "second" in the specification and claims herein are used to distinguish different objects rather than to describe a specific order of the objects. For example, a first response message and a second response message are used to distinguish different response messages rather than to describe a specific order of the response messages.
[0048] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0049] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more than two. For example, multiple processing units refer to two or more processing units, etc.; multiple elements refer to two or more elements, etc.
[0050] In order to obtain a highly accurate target detection model and reduce the labeling cost, a method is provided in an embodiment of the present application. The method is mainly as follows.
[0051] To facilitate understanding of the technical solutions of the embodiments of the present application, the terms and English abbreviations involved in this document are explained below.
[0052] vtune is a performance analysis tool provided by Intel.
[0053] Pref is a performance analysis tool provided by Linux (an operating system).
[0054] CPU (central processing unit) is the computing and control core of the computer system and the final execution unit for information processing and program running.
[0055] AIO (asynchronous input / output), non-blocking asynchronous input and output.
[0056] Lock, lock, ensures that when one thread is in the critical section of the code, another thread does not enter the critical section.
[0057] Poll: A thread or process in dormant state is awakened when the awakening condition is met or the waiting time is reached.
[0058] Schedule, scheduled scheduling
[0059] Page_fault, page fault exception.
[0060] BPF (Berkeley Packet Filter) performance analysis tool uses BPF technology to insert kernel executable code into kernel space to monitor and analyze the system. Similarly, eBPF (Extended Berkeley Packet Filter), literally translated into Chinese as extended Berkeley Packet Filter, is also BPF in essence. Similar to BPF performance analysis tool, eBPF performance analysis tool also allows the program to run without modifying the kernel source code or adding additional kernel modules.
[0061] In the related art, commonly used performance analysis tools for software performance analysis include Intel's vtune performance analyzer and the Pref tool provided by Linux.
[0062] The vtune performance analyzer can analyze the software and display various indicators of the software from multiple dimensions (such as CPU usage, memory usage, number of concurrency, etc.), and provide a performance snapshot function for overall display to guide users to perform the next step of analysis based on the overall situation. However, the vtune performance analyzer requires manual startup of data collection for multiple indicators and the startup of different feature analysis tools, and its intelligence level is relatively low. Moreover, the vtune performance analyzer requires personnel to operate the data collection for indicators and the startup of analysis tools based on their own rich experience. In addition, based solely on the analysis results output by the vtune performance analyzer, it is impossible to interpret the actual code that needs to be optimized in the software, and the suggestions provided are not specific enough and are not very operational.
[0063] When performing software analysis, the Perf tool can flexibly define software and hardware events and trigger the collection of current state context information. It can also perform relevant statistics through software and hardware counters. However, the data collection of different state information of this tool requires personnel to manually trigger it according to data analysis work, which requires high professional background of personnel and is inefficient.
[0064] Therefore, the above software performance analysis method has the problems of low intelligence and low efficiency.
[0065] In order to improve the intelligence level and efficiency in software performance analysis, the embodiment of the present application provides a software performance analysis method. The method mainly relies on a performance analysis model, which defines the execution process of various analysis behaviors. The corresponding performance analysis tool can be called based on the execution process to perform automatic data collection and detection analysis on the software to be tested, thereby obtaining analysis results. In this way, manual input can be reduced, the intelligence level of performance analysis operations can be improved, and analysis efficiency can be improved.
[0066] To facilitate understanding of the technical solution of the present application, an application scenario of an embodiment of the present application is first introduced below.
[0067] For example, Figure 1 A schematic diagram of an application scenario of an embodiment of the present application is shown. Figure 1 As shown in (1a) to (1c), one or more applications (a1, a2, ...) can be deployed on the server 10. The applications (a1, a2, ...) can be websites, forums, databases, etc., but are not limited thereto. The client 30 can access the server 10 through the network to obtain the services provided by these applications (a1, a2, ...).
[0068] In order to ensure service stability and user experience, before the application (a1, a2, ...) goes online, the operation and maintenance personnel can use the software performance analysis device 20 to perform performance analysis on the application (a1, a2, ...) on the server 10. The software performance analysis device 20 defines the execution process of multiple analysis behaviors. Based on the execution process, the software performance analysis device 20 can determine: what kind of analysis behavior needs to be executed at each stage, the indicator data that needs to be collected from the data generated by the application (a1, a2, ...) when each analysis behavior is executed, and the analysis tools that need to be called, and it can also determine what kind of analysis behavior needs to be executed next after the execution analysis behavior returns the result, etc., until the final performance analysis result is obtained. Based on this performance analysis result, the software performance analysis device 20 can determine the performance bottleneck point (the ultimate cause of time delay) of the application (a1, a2, ...), so as to optimize the specific code that causes the bottleneck point of the application (a1, a2, ...).
[0069] In one possible implementation, Figure 1 As shown in (1a), the software performance analysis device 20 can be deployed as an analysis tool and the applications (a1, a2, ...) in the same server 10, so that the applications (a1, a2, ...) running on the server 10 can be performance-tested through the locally deployed software performance analysis device 20.
[0070] In another possible implementation, Figure 1 As shown in (1b), the software performance analysis device 20 can be deployed on the proxy server 40 of the server 10. The proxy server 40 can collect relevant data of the applications (a1, a2, ...) running on the server 10 and transmit it to the software performance analysis device 20 to perform performance testing on these applications (a1, a2, ...).
[0071] In other possible implementations, such as Figure 1 As shown in (1c), the software performance analysis device 20 can be deployed on one or more servers 50. The software performance analysis device 20 virtualizes the resources (such as processors, memories) of multiple servers 50 and deploys them as remote services on the virtualized resources to provide remote software performance detection services for personnel (server 10 side).
[0072] In this embodiment, the software performance analysis device 20 may also have other suitable deployment modes, which will not be described in detail here.
[0073] Next, a software performance analysis device provided in an embodiment of the present application is introduced in conjunction with the accompanying drawings.
[0074] For example, Figure 2 FIG. 1 is a schematic diagram showing the structure of a software performance analysis device provided in an embodiment of the present application. It is understood that the software performance analysis device can be deployed in any device or device cluster with computing and processing capabilities, for example, Figure 1 The software performance analysis device 20 shown in FIG. Figure 2 As shown, the software performance analysis device 20 may include a performance analysis model 210, an execution framework 220, a tool set module 230, and a visualization module 240. Among them:
[0075] The performance analysis model 210 may be used to define an execution flow between multiple analysis behaviors, wherein the execution flow includes an execution order between multiple analysis behaviors and an execution condition of each analysis behavior;
[0076] The execution framework 220 may be used to load the performance analysis model 210 and, in the process of executing the flow, call the performance analysis tool related to the corresponding analysis behavior according to the execution conditions achieved, and one analysis behavior is associated with at least one performance analysis tool;
[0077] Performance analysis tools can be used to process target data and obtain return results. The target data includes performance indicator data displayed during the operation of the target application;
[0078] The execution framework 220 can also be used to, after obtaining any return result, continue to schedule the next analysis behavior-related performance analysis tool to process the target data in the execution order according to the execution conditions satisfied by any of the return results, until the execution process ends and the performance analysis result is obtained.
[0079] The visualization module 240 may be used to visualize the above-mentioned returned results and performance analysis results.
[0080] In this embodiment, the software performance analysis device 20 can call the corresponding performance analysis tool to automatically and orderly perform various analysis behaviors based on the execution process defined by the performance analysis model 210 during the operation of the target application, such as first performing time proportion overhead analysis, and then performing idle waiting cause analysis when the returned result indicates that the idle time proportion is high, etc., step by step until the final analysis result is obtained. As a result, manual input is reduced, the intelligence level of performance analysis operation is improved, and the versatility is strong.
[0081] The following is a detailed description of each module / unit of the software performance analysis device 20 in this embodiment.
[0082] In this embodiment, the performance analysis model 210 is a model created based on logic tree thinking. The performance analysis model 210 is used to define the execution flow between various performance analysis behaviors of the software through a decision tree.
[0083] Exemplarily, the performance analysis model 210 represents the corresponding analysis behavior through the nodes of the decision tree 211 (including the root node and each layer of child nodes), and the analysis behavior is the performance analysis behavior of the software, such as the analysis of the thread time overhead ratio, IO model analysis, IO delay analysis, thread idle reason analysis, system load analysis, etc. during the operation of the tested application. Each branch of each node (i.e., the branch from the current node to the next layer of nodes) is used to represent a return result of the corresponding analysis behavior, and each leaf node is used to store a corresponding analysis result, which may include existing problems and / or optimization suggestions.
[0084] For example, Figure 3 What is shown in FIG. is a schematic diagram of a decision tree 211. Figure 3As shown, in this example, the root node N0 of the decision tree 211 represents the time proportion overhead analysis behavior, which is used to count the proportion of idle waiting time and the proportion of normal operation time of the process in the tested application. After the analysis of the root node N0, if the idle time proportion is high, the idle reason analysis behavior represented by the sub-node N11 in the first level is executed. In this way, the result returned by the analysis behavior of the sub-node N11 can determine the reason for the tested application to wait, thereby judging whether the idle waiting state is reasonable, whether there is room for improvement, and so on. For example Figure 3 As shown in , the branches of subnode N11 respectively represent the results of high IO waiting ratio, high network waiting ratio, high synchronization time waiting ratio, high scheduling waiting ratio and high exchange partition waiting ratio. The second-level neutron nodes N21, N22, ... connected to these branches respectively represent IO model analysis behavior, network waiting analysis behavior, etc., so that the corresponding performance analysis behavior can be further performed according to the branch to which the return result of subnode N11 belongs. In this way, step by step, after executing the analysis behaviors represented by the third-level neutron nodes N31 and N32, the fourth-level neutron nodes N41, N42, N43, and the fifth-level subnodes N51 and N52, the output results represented by the leaf nodes E1 and E2 are obtained, that is, the problem to be solved is "high system load", or the output suggestion is "configure IO queue priority".
[0085] Similarly, if Figure 3 As shown in , after the analysis of the root node N0, if the normal operation time ratio is high, the hotspot function analysis behavior represented by the subnode N12 in the first level is executed. Therefore, it is also possible to further analyze whether the CPU flow meets expectations based on the returned results (the subsequent analysis behavior node of the subnode N12 is not in Figure 3 In the CPU pipeline, we can analyze whether the performance or efficiency shown in the CPU pipeline meets expectations, whether there is room for optimization, whether the compiled instructions are optimal, whether the CPU performance has reached its limit, and so on.
[0086] Understandably, Figure 3 Only a partial structure of the decision tree 211 is illustrated. The decision tree 211 may also include more or fewer nodes and branches. Each node and branch may be constructed according to actual application analysis requirements, which will not be described in detail here.
[0087] In this way, the performance analysis model 210 can perform performance analysis and measurement on the time proportion overhead, CPU occupancy, memory occupancy, network occupancy, number of concurrency, number of instructions, response time and other dimensional indicators shown during the operation of the application under test (i.e., the target application) based on the execution process of the decision tree 211, but is not limited to this. Considering that most performance problems of software are essentially time problems, various resource restrictions on software will eventually be reflected in time. Therefore, the decision tree 211 constructed by the performance analysis model 210 in this embodiment can analyze the time results shown by the application under test at different stages from top to bottom, layer by layer, for example Figure 3 As shown in the figure, the idle / running ratio, IO latency and other results shown in different stages of the application are analyzed to evaluate the performance of the tested application and obtain accurate and effective performance analysis results.
[0088] In this embodiment, various analysis behaviors are abstracted and solidified into model files to obtain the performance analysis model 210 to define these analysis behaviors at various stages of the software analysis process, thereby constructing an overall execution process that does not require human intervention, improving the intelligence of software performance analysis, and improving analysis efficiency.
[0089] In this embodiment, the performance analysis module 210 can be loaded and run by the execution framework 220. The execution framework 220 can be used as the execution entry of the performance analysis operation to execute the execution process defined by the performance analysis model 210, so as to call and execute the corresponding performance analysis tool according to the analysis behavior represented by each node of the execution process, and obtain the corresponding return result to execute the analysis behavior of the next node until all operations are completed.
[0090] Exemplary, reference Figure 4 As shown, the execution framework 220 can obtain the startup parameters in step 1. The startup parameters can be input by the user through the command line or the interface provided by the visualization module 240, so that the execution framework 220 determines the target application (such as the identification information of the target application (such as the application number), storage path or running path, etc., but not limited to this) according to the startup parameters, or gives the performance analysis tool corresponding capabilities. For example, for a performance analysis tool that does not support reading and writing process lists, the newly created threads or processes during the operation of the tested application are tracked by the performance analysis tool according to the startup parameters. Then, the execution framework 220 can be Figure 4 Step 2 shown in the figure loads and runs the performance analysis module 210, and step 3 calls the corresponding performance analysis tool (b1, b2, ...) from the tool set module 230 to perform analysis behavior operations.
[0091] In this example, the execution framework 220 can be considered as the executor of the performance analysis model 210, and each performance analysis tool (b1, b2, ...) called by the execution framework 220 is configured with a corresponding start condition, which is the execution condition of the analysis behavior defined in the performance analysis model 210. The execution framework 220 starts the performance analysis tool (b1, b2, ...) to perform the analysis task after the start condition is met, and notifies the performance analysis tool (b1, b2, ...) to end the collection of target data, feature data, etc. when the end condition is met. The data collected by the performance analysis tool (b1, b2, ...) can be stored in the database 250, and the execution framework 220 can read the data through step 5 and then pass it to the visualization module 240 for presentation through step 6.
[0092] In some examples, the analysis behavior represented by each node Node of the decision tree 211 can be associated with the corresponding performance analysis tool (b1, b2, ...) for the execution framework 220 to call according to the execution process. Among them, when each performance analysis tool (b1, b2, ...) is called and executed, it can automatically obtain the target data for performance analysis processing and obtain the return result. Among them, the target data is the indicator data required for the performance analysis tool to implement the corresponding performance analysis. As an example, the target data may include data generated during the operation of the application under test, such as stack information during process switching, and the target data may also include the return results of related performance analysis tools, such as the target data obtained by the sub-node N22 includes the return result of the upper-level sub-node N11, and so on. In addition, it can be understood that these performance analysis tools (b1, b2, ...) can all come from the tool set module 230.
[0093] The tool set module 230 and the performance analysis tool are introduced in detail below.
[0094] In this embodiment, if Figure 5 As shown, the tool set module 230 may include an application base library 231, which can be used to provide encapsulation of basic user-mode functions, such as map data reading and writing, database reading and writing, message interaction, maintenance encapsulation of process lists (pid information, etc.), and maintenance encapsulation of map data structures, so that the execution framework 220 (or bpf execution framework) and performance analysis tools (b1, b2, ...) can access these data. It can be understood that in some examples, functions such as maintenance encapsulation of map data structures can also rely on the bpf function library. It can be understood that map data is the storage form of target data collected by performance analysis tools (b1, b2, ...) in the database. In addition, test data such as target data, return results, feature data, etc. can be stored in a specified location in the database, and multiple copies can be saved. Multiple copies of test data can be presented from different dimensions.
[0095] The performance analysis tool (b1, b2, ...) is used to collect target data, analyze and process target data, and manage the returned results (e.g., output in a unified format and store in the database 250) through the interface provided by the kernel or other interfaces. The target data may include data generated during the running of the application under test, such as stack information when the thread of the application under test is switched, the number of surviving threads, the number of threads running off the CPU (offcpu), the last time the thread was switched out, the time of switching out, etc., but is not limited thereto.
[0096] Exemplarily, the performance analysis tools (b1, b2, ...) in the tool set module 230 may specifically include:
[0097] The time overhead ratio analysis tool is used to perform time overhead ratio analysis. The result returned is the thread idle time ratio and thread running time ratio.
[0098] The hot function analysis tool is used to perform hot function analysis to locate hot functions in the application under test when the return result of the thread time overhead ratio analysis tool indicates that the thread running time ratio is high (higher than the preset value or higher than the thread idle time ratio).
[0099] The idle cause analysis tool is used to determine the cause of the high idle ratio of threads when the idle time of threads is high (higher than the preset value or higher than the thread running time ratio). The return results of the tool include but are not limited to the cause types such as IO waiting, network waiting, synchronization event waiting, and exchange partition waiting. It can be understood that the call of the hot function analysis tool and the idle cause analysis tool are not mutually exclusive, and the execution can be started as long as the start conditions are met. Similarly, the cause types obtained by the idle cause analysis tool are not mutually exclusive. As long as the occurrence ratio of any cause type reaches the preset threshold, it can be used as a return result for the next analysis behavior operation. Therefore, after the idle cause analysis tool obtains the return result, multiple performance analysis tools may be activated. In addition, when the next behavior analysis is performed based on a certain cause type, the callable performance analysis tool is not unique. It can be selected from multiple performance analysis tools (b1, b2, ...), or multiple performance analysis tools (b1, b2, ...) can be analyzed at the same time. Different performance analysis tools (b1, b2, ...) can perform different judgment logics, which are not listed here one by one.
[0100] IO model analysis tool, used to classify IO types. This tool mainly classifies IO based on the size, continuity, randomness and other factors of each IO request. The output IO classification includes continuous IO, random IO, read IO, and write IO.
[0101] The IO delay analysis tool is used to measure the IO delay and determine whether the delay of an IO is too large in combination with the IO classification returned by the IO model analysis tool mentioned above.
[0102] The request (req) statistical analysis tool is used to analyze IO req requests when the IO delay is too large and determine whether IO preemption occurs.
[0103] The system load analysis tool is mainly used to determine whether there are other programs performing IO reading and writing in addition to the application being tested, and whether the ratio of IO reading and writing performed by other programs reaches the ratio threshold (that is, whether the system load is higher than the set value).
[0104] It can be understood that the toolset module 230 in this embodiment can also include more or fewer performance analysis tools. Moreover, in this example, each performance analysis tool (b1, b2, ...) can be relatively independent, and these performance analysis tools (b1, b2, ...) can adopt mature software performance analysis tools in the field, such as bpf tool (or ebpf tool) b3, perf tool, which are not listed here one by one. Of course, in this example, the performance analysis tools (b1, b2, ...) can also be developed by personnel according to the performance analysis model 210 and added to the characteristic analysis model 210 for association.
[0105] Exemplarily, a performance analysis tool can be called when performing multiple analysis behaviors. For example, a hotspot function analysis tool can be called when performing a hotspot function analysis behavior operation, and can also be called when performing a CPU idle rate analysis behavior operation. Therefore, in this example, each performance analysis tool (b1, b2, ...) can have an identity identifier and a call identifier, and the identity identifier is a unique identifier used to characterize the identity of the performance analysis tool (b1, b2, ...), and the call identifier is used to characterize the association between the performance analysis tool (b1, b2, ...) and the analysis behavior. In this way, when the same performance analysis tool is started multiple times in different analysis stages, it can be determined based on the identity identifier and the call identifier when executing which analysis behavior the performance analysis tool is called.
[0106] In this example, different performance analysis tools (b1, b2, ...) can exchange map data, database information, and messages through the application base library. If a performance analysis tool does not support reading and writing process lists, the execution framework 220 can grant this capability through startup parameters, so that new threads or processes are tracked by these performance analysis tools themselves during the running of the application under test.
[0107] Exemplarily, the characteristic analysis model 210 can define a unified input and output format for each performance analysis tool to facilitate the smooth execution of the execution process. The output (return result) of each performance analysis tool (b1, b2, ...) can include execution results, intermediate process data, etc., but is not limited thereto.
[0108] In addition, the performance analysis tool (b1, b2, ...) can also extract feature data based on the returned results for saving operations. The feature data is used to characterize the process operation status or system status that forms the returned results, and is the data obtained by the performance analysis tool (b1, b2, ...) through statistical analysis of the target data. For example, the feature data extracted by the idle cause analysis tool may include the number of occurrences of events classified by various reasons that cause the CPU to be idle, the total time of being out of CPU operation (offcpu_time), etc. In this example, the feature analysis model 210 can also define a unified storage format and presentation method (such as a table or histogram, etc.) for the feature data, so that the execution framework 220 obtains test data in a unified format and passes it to the visualization module 240 during the operation of the feature analysis model 210. The visualization module 240 does not need to perform content parsing based on this self-explanatory unified format, and only needs to present the data to the user according to the set presentation method, which is conducive to improving the intelligence of the test.
[0109] In some possible examples, the tool set module 230 also includes a suggestion set, which includes at least one optimization suggestion, and each optimization suggestion can be associated with a corresponding branch of the last child node in the execution process of the performance analysis model 210. In this way, after executing the last analysis behavior, the performance analysis model 210 can output corresponding optimization suggestions based on the return results of the performance analysis tool (b1, b2, ...) that implements the analysis behavior operation. For example, when the system load analysis tool determines that the system load is higher than the set value, it can be considered that the system load has a greater impact on the operation of the application under test, and the output analysis results may include optimization suggestions for solving the high system load.
[0110] In some possible examples, the execution framework 220 can also obtain the performance analysis results according to the execution logic defined by the performance analysis model 210, and then trace back the storage location of the corresponding program code from the stack information in the target data used by each performance analysis tool (b1, b2, ...), so as to provide the program code to the user as the code to be optimized, such as providing the storage path and code segment, etc. For example, if the thread time overhead ratio analysis behavior, idle reason analysis behavior, IO model analysis behavior and IO delay analysis behavior are performed according to the execution process, and finally it is determined that there is a problem of high system load, the execution framework 220 traces the stack information (such as the data information in the function call stack) included in the target data used when performing these analysis behaviors according to this result, determines the program code address of the application under test pointed to by these stack information, that is, finds the location of the code to be optimized. In this way, during the test process, it is not necessary to parse the application code, and the user code behavior can be deduced according to the performance analysis results and related feature data, kernel state data (included in the target data), etc., and the location of the code to be optimized can also be provided while providing optimization suggestions.
[0111] Exemplarily, the performance analysis tools (b1, b2, ...) in this embodiment can adopt bpf analysis tools, ebpf analysis tools, etc. For this purpose, the tool set module 230 can also include a bpf execution framework 233. The bpf execution framework 233 is the execution entry of the bpf program, and can maintain dependencies according to the definitions of performance analysis tools such as the bpf analysis tool and the ebpf analysis tool, and automatically start the dependent tools when starting the performance analysis tools that depend on other tools. In this way, based on the characteristics of the bpf execution framework 233, the purpose of kernel-mode code injection without modifying the kernel can be achieved. It can be understood that some performance analysis tools of the tool set module 230 can be implemented using other frameworks (such as the perf framework), or encapsulated using mature tools.
[0112] In this embodiment, the software performance analysis device 20 also includes a visualization module 240. The visualization module 240 provides a front-end visualization interface for the user. The return results and feature data obtained when performing the analysis behavior of each stage of the performance analysis module 210, as well as test data such as the code to be optimized, can all be passed to the visualization module 240 by the execution framework 220. Then, the visualization module 240 can present these test data on the visualization interface through a web page (web), command, etc. Among them, the presentation method of the test data on the visualization interface can include but is not limited to a table, a histogram, a timing diagram, a pie chart, etc. For example, the return results obtained by processing using a time proportion overhead analysis tool include the thread idle time and the running time proportion, which are then displayed in a table by the visualization module 240 according to the content of the thread, the thread idle time and the running time proportion.
[0113] In this example, the data display location, data selection rules, etc. of each presentation method can be pre-defined (can be customized by the user from the visualization interface). For example, if the presentation method is set to a pie chart, a column of data in a data table can be specified as a comparison object through the visualization interface. The data in the column where the comparison object is located can be a string. If a column of data is specified to calculate the comparison percentage, the column of data can be a percentage or a decimal, or a number. In this example, the execution framework 220 can provide a unified interface to process and parse the test data obtained from the performance analysis tool, and pass it to the visualization module 240 for normalization, so that it can be displayed according to the definition of the performance analysis model 210.
[0114] In this embodiment, the visualization interface based on the visualization module 240 can not only display the execution process and test data, but also provide an interface that supports manual triggering of a performance analysis tool, a page for modifying the performance analysis model (such as its execution process), etc., to facilitate users to perform relevant triggers or modifications according to their own needs.
[0115] Thus, in this embodiment, when the software performance analysis device 20 is used to perform performance testing on the application under test, the execution framework 220 can call the performance analysis tool based on the execution process defined by the performance analysis model 210 to automatically collect target data, automatically execute various analysis behaviors, and obtain test data of each test stage, and present it to the user. In this way, the manual input in the test process is greatly reduced, helping users to quickly find the root cause of the bottleneck of the application under test and present the external performance.
[0116] Next, a software performance analysis method provided by an embodiment of the present application is introduced. It can be understood that the method is proposed based on the content described above, and part or all of the content of the method can be referred to the description above.
[0117] For example, Figure 6 FIG. 1 is a flow chart of a software performance analysis method provided by an embodiment of the present application. It is understood that the method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities. Figure 6 As shown, the method may include:
[0118] S610: Obtain an execution process for analyzing the performance of a target application.
[0119] In this embodiment, before the computing device launches the application (i.e., the application under test), the root cause of the bottleneck of the application can be tested through performance analysis, thereby facilitating personnel to solve the problem in a targeted manner. For example, during the testing process of a web application, the user may encounter the problem of taking too long to log in to the application. The reason for the long user operation time (such as excessive system load) can be determined through performance analysis. Therefore, in this embodiment, a complete set of performance analysis execution processes is defined through a pre-created performance analysis model to solidify the execution conditions, end conditions, and execution order of various analysis behaviors. In this way, the automated testing of the application is completed by executing the execution process. It can be understood that the computing device in this embodiment can be Figure 1 The server 10, proxy server 40, or server 50 in the cluster shown in the figure shall be the same below.
[0120] For example, the execution process may be Figure 2 The performance analysis model 210 provided by the software performance analysis device 20 is defined by a decision tree 211. The root node and each child node of the decision tree 211 represent an analysis behavior respectively. Each branch in the decision tree 211 represents an execution condition respectively, and the leaf node represents an analysis result. In the process of running the execution flow, the analysis behavior of the next node can be executed according to the execution condition satisfied by the analysis result of the analysis behavior of any node until the execution flow is completed, that is, the test is completed. It can be understood that the application under test and the software performance analysis device 20 can be deployed on the same device or the same device cluster, but is not limited to this. The computing device can test local applications through the software performance analysis device 20, and can also test applications running on other devices.
[0121] As an example, the computing device 20 may obtain startup parameters input by the user through an interactive interface (which may be provided by the visualization module 240 in the software performance analysis apparatus 20), determine the application under test, and trigger operations of loading and running the corresponding execution process for the application under test.
[0122] In this example, the execution framework 220 of the performance analysis model 210 may be used to load and run the execution process to start a performance analysis operation on the target application.
[0123] S620: During the execution process, according to the execution conditions reached, a performance analysis tool related to the corresponding analysis behavior is called.
[0124] In this embodiment, the execution framework 220 determines the execution conditions reached in order according to the predefined execution process, and calls the performance analysis tool (b1, b2, ...) related to the analysis behavior corresponding to the execution condition. For example, after starting the execution process, the execution framework 220 calls the thread time overhead ratio analysis tool associated with the thread time overhead ratio analysis behavior represented by the root node N0 of the decision tree 211, or the execution framework 220 determines that the thread space time ratio of the tested application reaches a preset value, that is, it meets the execution condition for performing idle cause analysis, and then calls the idle cause analysis tool related to the idle cause analysis behavior.
[0125] Exemplarily, one analysis behavior is associated with at least one performance analysis tool, and a performance analysis tool associated with one analysis behavior may also be called depending on other performance analysis tools, and the performance analysis tools associated with different analysis behaviors may be the same or different. As an example, these performance analysis tools may be called from the tool set module 230 .
[0126] S630, processing the target data by calling the performance analysis tool to obtain a return result.
[0127] In this embodiment, the performance indicators required by each performance analysis tool (b1, b2, ...) are not the same. After the execution framework 220 calls the performance analysis tool (b1, b2, ...), the performance analysis tool (b1, b2, ...) can collect the target data generated during the operation of the tested application through the interface provided by the kernel according to the performance indicators required by itself. For example, when the idle cause analysis tool is running, it can obtain the current time when the process (including at least one thread) is scheduled to switch to the idle state, the number of surviving threads, the number of threads running off the CPU (offcpu), the instruction operation code (cpuid), the last time the process was switched out and the stack of the switch, the current switching time, the switch-out process, the switch-in process and the current stack of the switch-out process, the stack information and other target data during the operation of the tested application according to the performance indicators required by itself, but it is not limited to this.
[0128] The called performance analysis tool (b1, b2, ...) automatically obtains the required target data and performs statistical processing during operation. For example, the idle reason analysis tool counts various waiting types (i.e., idle reason types) caused by each thread, such as lock event waiting, IO waiting, etc., and counts the total time and total time proportion of each reason type away from CPU operation, etc., and returns the result to the execution box, 220.
[0129] In some examples, the performance analysis tool (b1, b2, . . . ) may also extract feature data corresponding to the returned result from the target data and transmit the same to the execution framework 220 .
[0130] S640, according to the execution condition satisfied by the returned result, call the next performance analysis tool related to the analysis behavior in the execution order to process the target data until the execution process ends to obtain the performance analysis result.
[0131] In this embodiment, the execution framework 220 calls the corresponding completed execution analysis behaviors according to the predefined logical order of the execution process, and obtains the analysis results of these analysis behaviors one by one (that is, the return results of the corresponding performance analysis tools). If the analysis results of any stage meet any execution condition, the next analysis behavior of the execution condition will continue to be executed until the execution process ends.
[0132] During the entire execution process, the execution framework 220 can parse the return results obtained at each stage of the software performance analysis, the feature data extracted by the performance analysis tool (b1, b2, ...), the optimization suggestions finally obtained, and the traced code location to be optimized and other test data, and pass them to the visualization module 240 for display as performance analysis results, so that users can understand the test results.
[0133] The software performance analysis method provided in the embodiment of the present application is described in detail below with reference to a specific example.
[0134] In this example, taking the analysis of the time results of each stage of the tested application as an example to perform performance testing, the performance analysis model 210 defines, for example Figure 3 The execution flow of the decision tree 211 shown in FIG. Figure 7 As shown, the software performance analysis method may specifically include:
[0135] S700, obtaining startup parameters.
[0136] In this step, the computing device obtains the startup parameters configured by the user through the interactive interface provided by the visualization module 240. The startup parameters may include the storage path (or running path) of the application under test, etc., and may also include configuration parameters for some performance analysis tools (b1, b2, ...) to give these performance analysis tools (b1, b2, ...) the ability to track the creation of threads (or processes), etc., but not limited to this.
[0137] S710: Calling an execution process for performing performance analysis on the application under test according to the startup parameters.
[0138] In this step, the computing device passes the startup parameters to the execution framework 220 , and the execution framework 220 loads and runs the performance analysis model 210 according to the startup parameters, and starts the execution process of performing performance analysis on the application under test.
[0139] In this example, the execution framework 220 serves as the execution entry of the performance analysis operation to execute the execution process defined by the performance analysis model 210. In addition, the execution framework 220 also maintains the process pid information, map structure data, etc. of the application under test according to the entry parameters.
[0140] S720, during the execution process, calling a performance analysis tool related to the corresponding analysis behavior according to the execution condition achieved;
[0141] S730, processing the target data by calling the performance analysis tool to obtain a return result.
[0142] In this example, the principles of steps S720 to S730 are similar to the description of S620 to S630 in the above embodiment. The execution framework 220 determines the execution conditions according to the execution process to execute corresponding analysis behaviors, and calls relevant performance analysis tools (b1, b2, ...) to perform target data analysis.
[0143] For example. During the application running time, there is a part of time that needs to consume the CPU, and there is a part of time that does not need to consume the CPU. The idle reason analysis behavior is mainly used to analyze the time and reasons that do not need to consume the CPU by counting the idle reasons of each thread switch of the tested application, so as to perform the next step according to the analysis results. Therefore, in this example, when the execution process proceeds to the stage of executing the idle reason analysis behavior, the execution framework 220 calls the associated idle reason analysis tool from the tool set module 230 to perform the analysis operation. At this time, combined with Figure 8 As shown, the idle cause analysis tool called to run can record the scheduled process and the corresponding call chain when the process scheduling of the application under test switches to the idle state through the interface provided by the kernel (such as through system call). The call chain contains the call information (target data) such as time, interface, level, result, etc. related to the process switching, so as to perform statistics based on these call information. The idle cause analysis tool can be a performance analysis tool based on ebpf, so that when performing target data (through Figure 8 In step 1, there is no need to modify the kernel during the acquisition and completion of the analysis behavior (real-time collection from kernel state), which improves the usability of code injection of performance analysis tools.
[0144] Specifically, when the idle reason analysis tool is running, the system call can be used to collect the current time when the process (including at least one thread) is scheduled to switch to the idle state, the number of surviving threads, the number of threads running off the CPU (offcpu), historical cpu distribution, instruction operation code (cpuid), the last time the process was switched out and the stack, the current switching time, the process being switched out, the process being switched in, the current stack of the process being switched out, the stack information, and other target data. Figure 8Then the idle reason analysis tool can be used to save the data in real time. Figure 8 Step 3 performs statistical analysis on these target data to determine the type of reason for waiting (category), number of times (count), total time off from CPU operation (off_time), and the proportion of each off_time (precent). By way of example and not limitation, the reason type may include IO waiting (such as non-blocking asynchronous input and output (AIO) waiting), synchronous event waiting (such as lock waiting), network waiting, scheduling waiting, swap partition waiting, etc., and may also include idleness caused by thread sleeping (sleep), waiting (wait), waking up (poll) or abnormal (unknown) state. The idle reason analysis tool counts the number of times, off_time and proportion of each reason type (the proportion of a single reason type = the off_time of a single reason type / the sum of the off_time of all reason types), determines the reason type with the highest idle time proportion among all idle reasons, and uses this to judge the state of each thread switch, clarify the reason for the switch, and form the return result of the idle reason analysis behavior. Together with the above-mentioned reason type (category), number (count), total time off CPU operation (off_time), and the proportion of each off_time (precent) and other characteristic data, it is used as test data through Figure 8 The steps 4 are summarized and saved to be passed to the execution framework 210 .
[0145] Next, the execution framework 220 may execute step S731, parse and process the test data received from the idle reason analysis tool, and then transmit the parsed test data to the visualization module 240 for display. The display may be in real time during the test process or after the test is completed, which is not limited in this example. The specific content of the test data presented by the visualization module 240 can be referred to in Fig. 9 shown.
[0146] Similarly, when the idle rate analysis tool is called for analysis, the tool can perform statistics based on the idle state of each record point (process or thread) to determine the percentage of the total process occupied by the idle process, thereby achieving the overall CPU idle situation evaluation at a small storage space cost. The return result of the idle rate analysis tool can be passed to other feature analysis tools as a data reference. The return result is passed by the execution framework to the visualization module 240 for presentation. The specific presentation content can be referred to Fig.10 .
[0147] S740, according to the execution condition satisfied by the returned result, calling the next performance analysis tool related to the analysis behavior in the execution order to process the target data until the execution process ends to obtain the performance analysis result.
[0148] The execution principle of this step is similar to the description of S640 in the above embodiment, and will not be repeated here.
[0149] S750: Display the performance analysis result.
[0150] In this step, the computing device can display each returned result, feature data, optimization suggestion, code to be optimized, etc. according to the rules defined by the performance analysis module 210 through the visualization module 240, so that the user can understand the specific content of each data, and improve the efficiency of visualization development and display. For example, the test data received by the visualization module 240 can display the classification statistics of the thread based on the thread (such as Fig.11 ), or display the online time (oncpu_time) and offline time (offcpu_time) of the tested application during the entire running period and the related ratios by thread (as shown in (11a)). Fig.11 ), or display statistics of each stack based on threads and directly point to the code to be optimized (such as Fig.11 as shown in (11c)).
[0151] In this way, the performance analysis process and execution framework for automated execution are defined through the mind tree method, and the performance analysis work is automatically assigned and executed without the need for personnel to receive checks, thereby reducing human participation and thus reducing the requirements for personnel's professional knowledge, and improving test results and problem location efficiency.
[0152] Based on the method in the above embodiment, Fig.12 As shown, an embodiment of the present application provides an electronic device 900. The electronic device 900 may include: at least one memory 910 storing a program; at least one processor 920 executing the program stored in the memory 910; wherein, when the program stored in the memory 910 is executed, the processor 920 is used to execute the method in the above embodiment. As a specific example, the software performance analysis device 20 described in the above embodiment may be stored in the memory 910, so that the processor 920 executes the steps in the software performance analysis method in the above embodiment.
[0153] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method in the above embodiment.
[0154] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product, characterized in that when the computer program product runs on a processor, the processor executes the method in the above embodiment.
[0155] Based on the method in the above embodiment, the present application embodiment also provides a chip. Fig.13 , Fig.13 This is a schematic diagram of the structure of a chip provided in an embodiment of the present application. Fig.13 As shown, chip 1000 includes one or more processors 1001 and an interface circuit 1002 .
[0156] Optionally, the chip 1000 may further include a bus 1003. Wherein:
[0157] Processor 1001 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in processor 1001 or an instruction in software form. The above processor 1001 can be a general-purpose processor, a digital communicator (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The disclosed methods and steps in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0158] The interface circuit 1002 can be used to send or receive data, instructions or information. The processor 1001 can use the data, instructions or other information received by the interface circuit 1002 to process, and can send the processing completion information through the interface circuit 1002.
[0159] Optionally, the chip 1000 further includes a memory, which may include a read-only memory and a random access memory, and provides operation instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory (NVRAM).
[0160] Optionally, the memory stores executable software modules or data structures, and the processor can perform corresponding operations by calling operation instructions stored in the memory (the operation instructions can be stored in the operating system).
[0161] Optionally, the interface circuit 1002 may be used to output the execution result of the processor 1001 .
[0162] It should be noted that the functions corresponding to the processor 1001 and the interface circuit 1002 can be implemented through hardware design, software design, or a combination of hardware and software, and there is no limitation here.
[0163] It should be understood that each step of the above method embodiment can be completed by a hardware-based logic circuit or a software-based instruction in a processor.
[0164] It is understandable that the size of the sequence number of each step in the above embodiment does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application. In addition, in some possible implementations, each step in the above embodiment can be selectively executed according to actual conditions, and can be partially executed or fully executed, which is not limited here.
[0165] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0166] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.
[0167] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions may be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state disk (SSD)), etc.
[0168] It should be understood that the various numerical numbers involved in the embodiments of the present application are only used for the convenience of description and are not used to limit the scope of the embodiments of the present application.
Claims
1. A software performance analysis method, It is characterized in that The method comprises: Invoking an execution process for performing performance analysis on the application under test, wherein the execution process includes an execution order between multiple analysis behaviors and an execution condition of each of the analysis behaviors; In the process of performing the execution process, according to the achieved execution conditions, a performance analysis tool related to the corresponding analysis behavior is called, and one of the analysis behaviors is associated with at least one of the performance analysis tools; Processing the target data by the called performance analysis tool to obtain a return result, wherein the target data includes performance indicator data displayed during the operation of the tested application; After obtaining any of the returned results, according to the execution conditions satisfied by any of the returned results, a performance analysis tool related to the corresponding analysis behavior is called in the execution order to process the target data to obtain a performance analysis result.
2. The method according to claim 1, It is characterized in that The execution process of calling the performance analysis of the application under test includes: Acquire startup parameters, where the startup parameters include identification information of the application under test; According to the startup parameters, the execution process is called to start a performance analysis operation on the application under test.
3. The method according to claim 1 or 2, It is characterized in that The returned result includes characteristic data, and the characteristic data is used to characterize the process operation status or system status that forms the returned result; The target data is processed by the performance analysis tool called to obtain a return result, including: Obtaining the target data from the kernel by calling the performance analysis tool; Perform statistics on the target data according to the statistical dimensions of the performance analysis tool to obtain statistical results; extracting the feature data from the target data; The statistical result and the characteristic data are combined to form the returned result.
4. The method according to any one of claims 1 to 3, It is characterized in that The performance analysis results include return results and target optimization suggestions; The step of calling a performance analysis tool related to a corresponding analysis behavior according to the execution condition satisfied by any of the returned results to process the target data in the execution order to obtain a performance analysis result includes: According to the execution condition satisfied by any of the returned results, continue to schedule the next analysis behavior-related performance analysis tool to process the target data in the execution order until the return result of the last analysis behavior-related performance analysis tool in the execution process is obtained; According to the return result of the performance analysis tool related to the last analysis behavior, matching a corresponding target optimization suggestion from a suggestion library, wherein the suggestion library includes multiple optimization suggestions, and the optimization suggestions are used to optimize the performance of the application, each of the optimization suggestions is associated with at least one of the return results, and the target optimization suggestion is at least one optimization suggestion in the suggestion library; The return result of the performance analysis tool related to the last analysis behavior and the target optimization suggestion are used to form the performance analysis result.
5. The method according to any one of claims 1 to 4, It is characterized in that The target data includes stack information generated during the running of the application under test, and the stack information points to the code location of the application under test; The step of calling a performance analysis tool related to a corresponding analysis behavior according to the execution condition satisfied by any of the returned results to process the target data in the execution order to obtain a performance analysis result includes: According to the execution condition satisfied by any of the returned results, continue to schedule the next analysis behavior-related performance analysis tool to process the target data in the execution order until the return result of the last analysis behavior-related performance analysis tool in the execution process is obtained; In the case that the return result of the performance analysis tool related to the last analysis behavior does not meet the test pass condition, determining the code position pointed to by the stack information from the target data processed by each performance analysis tool called in the execution process; The return result of the performance analysis tool related to the last analysis behavior and the code position are used to form the performance analysis result.
6. The method according to any one of claims 1 to 5, It is characterized in that After obtaining the performance analysis result, the method further includes: The performance analysis result and / or each of the returned results are displayed according to predefined display rules.
7. The method according to any one of claims 1 to 6, It is characterized in that The execution process includes time overhead ratio analysis behavior to analyze the running time ratio and idle waiting ratio during the running process of the tested application; In the process of executing the execution flow, according to the execution conditions reached, calling the performance analysis tool related to the corresponding analysis behavior includes: According to the execution process, when the idle waiting ratio reaches a preset threshold, an idle reason analysis tool related to the idle reason analysis behavior is called to determine the waiting reason for the idle waiting of the tested application; And, when the running time ratio reaches a preset value, a hot function analysis tool related to the hot function analysis behavior is called to determine the resource usage during the running of the tested application.
8. A software performance analysis device, It is characterized in that The device comprises: A performance analysis model is used to define an execution process for performing performance analysis on the application under test, wherein the execution process includes an execution order between multiple analysis behaviors and an execution condition of each of the analysis behaviors; An execution framework, used to load and run the performance analysis module, and in the process of performing the execution process, according to the execution conditions achieved, call the performance analysis tool related to the corresponding analysis behavior, and one of the analysis behaviors is associated with at least one of the performance analysis tools; The performance analysis tool is used to process target data and obtain return results, wherein the target data includes performance indicator data displayed during the operation of the tested application; The execution framework is also used to, after obtaining any of the return results, call the performance analysis tool related to the corresponding analysis behavior in accordance with the execution order according to the execution conditions satisfied by any of the return results to process the target data to obtain the performance analysis result.
9. The device according to claim 8, It is characterized in that The device further includes a visualization module, the visualization module being used to obtain startup parameters, the startup parameters including identification information of the application under test; The execution framework is also used to call the execution process according to the startup parameters to start the performance analysis operation on the application under test.
10. The device according to claim 8 or 9, It is characterized in that The returned result includes characteristic data, and the characteristic data is used to characterize the process operation status or system status that forms the returned result; The performance analysis tool is specifically used for: Obtaining the target data from the kernel by calling the performance analysis tool; Perform statistics on the target data according to the statistical dimensions of the performance analysis tool to obtain statistical results; extracting the feature data from the target data; The statistical result and the characteristic data are combined to form the returned result.
11. The device according to any one of claims 8 to 10, It is characterized in that The performance analysis results include return results and target optimization suggestions; The execution framework is specifically used for: According to the execution condition satisfied by any of the returned results, continue to schedule the next analysis behavior-related performance analysis tool to process the target data in the execution order until the return result of the last analysis behavior-related performance analysis tool in the execution process is obtained; According to the return result of the performance analysis tool related to the last analysis behavior, matching a corresponding target optimization suggestion from a suggestion library, wherein the suggestion library includes multiple optimization suggestions, and the optimization suggestions are used to optimize the performance of the application, each of the optimization suggestions is associated with at least one of the return results, and the target optimization suggestion is at least one optimization suggestion in the suggestion library; The return result of the performance analysis tool related to the last analysis behavior and the target optimization suggestion are used to form the performance analysis result.
12. The device according to any one of claims 8 to 11, It is characterized in that The target data includes stack information generated during the running of the application under test, and the stack information points to the code location of the application under test; The execution framework is specifically used for: According to the execution condition satisfied by any of the returned results, continue to schedule the next analysis behavior-related performance analysis tool to process the target data in the execution order until the return result of the last analysis behavior-related performance analysis tool in the execution process is obtained; In the case that the return result of the performance analysis tool related to the last analysis behavior does not meet the test pass condition, determining the code position pointed to by the stack information from the target data processed by each performance analysis tool called in the execution process; The return result of the performance analysis tool related to the last analysis behavior and the code position are used to form the performance analysis result.
13. The device according to any one of claims 8 to 12, It is characterized in that After obtaining the performance analysis result, the method further includes: the device further includes a visualization module, and the visualization module is used to: The performance analysis result and / or each of the returned results are displayed according to predefined display rules.
14. The device according to any one of claims 8 to 13, It is characterized in that The execution process includes time overhead ratio analysis behavior to analyze the running time ratio and idle waiting ratio during the running process of the tested application; The execution framework is specifically used for: According to the execution process, when the idle waiting ratio reaches a preset threshold, an idle reason analysis tool related to the idle reason analysis behavior is called to determine the waiting reason for the idle waiting of the tested application; And, when the running time ratio reaches a preset value, a hot function analysis tool related to the hot function analysis behavior is called to determine the resource usage during the running of the tested application.
15. An electronic device, It is characterized in that include: at least one memory for storing a program; at least one processor, configured to execute the program stored in the memory; Wherein, when the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1-7.
16. A computer-readable storage medium storing a computer program, wherein when the computer program is executed on a processor, the processor is caused to execute the method according to any one of claims 1 to 7.
17. A computer program product, It is characterized in that When the computer program product runs on a processor, the processor is caused to execute the method according to any one of claims 1 to 7.
18. A chip, It is characterized in that It comprises at least one processor and an interface; the at least one processor obtains program instructions or data through the interface; the at least one processor is used to execute program line instructions to implement any method as claimed in claims 1-7.