Method, device, readable storage medium, and program product for performance analysis of a microarchitecture

By acquiring and analyzing microarchitecture events and combining preset indicator evaluation models, the problem of insufficient accuracy of microarchitecture performance evaluation in the existing technology is solved, and high-precision function-level performance analysis is achieved.

CN119621493BActive Publication Date: 2025-06-20ALIBABA CLOUD COMPUTING CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510169799.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-20
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

The prior art has insufficient accuracy when evaluating the performance of cloud servers, which affects the accuracy of performance evaluation results.

Method used

By collecting microarchitecture events generated by the execution of various functions of the application, and determining the corresponding performance indicators of each function based on the preset indicator evaluation model, performing performance analysis to realize function-level microarchitecture performance analysis.

Benefits of technology

Improve the accuracy of performance analysis, so that the performance analysis results can be refined to the function level granularity, avoiding the complexity of using multiple tools to collect events and evaluate performance indicators separately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119621493B_ABST
    Figure CN119621493B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a method, device, readable storage medium, and program product for performance analysis of a microarchitecture, relating to the field of computer technology. The method includes: respectively collecting microarchitecture events generated by each function of the microarchitecture executing a predetermined application program; determining respective predetermined performance indicators corresponding to each function based on a preset metric evaluation model of the microarchitecture and the microarchitecture events corresponding to each function; performing performance analysis according to the respective predetermined performance indicators to obtain a performance analysis result of each function running on the microarchitecture. According to the performance analysis method of the embodiment of the present application, the accuracy of the performance analysis result can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and in particular, to a method, device, readable storage medium, and program product for performance analysis of a microarchitecture. Background Art

[0002] Cloud computing services are a service model that provides computing resources over the Internet. Cloud servers are used to provide efficient, scalable, and pay-per-use computing services built on cloud computing technology. Users expect to select high-performance cloud servers, and the performance of cloud servers depends to a large extent on their underlying hardware, especially the microarchitecture of the processor. Although the performance evaluation tools in the related art can evaluate the overall running performance of a program in a microarchitecture, there are deficiencies in accuracy, which affects the accuracy of the performance evaluation results of the microarchitecture. Summary of the Invention

[0003] Embodiments of this application provide a method, device, readable storage medium, and program product for performance analysis of a microarchitecture to alleviate or solve one or more technical problems existing in the prior art.

[0004] In a first aspect, embodiments of this application provide a method for performance analysis of a microarchitecture, including: respectively collecting microarchitecture events generated by each function of the microarchitecture executing a predetermined application program; determining, based on a preset metric evaluation model of the microarchitecture and the microarchitecture events corresponding to each function, each predetermined performance metric corresponding to each function; and performing performance analysis according to each predetermined performance metric to obtain a performance analysis result of each function running on the microarchitecture.

[0005] In a second aspect, embodiments of this application provide a method for performance analysis of a microarchitecture, including: obtaining microarchitecture events generated by each function of a first microarchitecture executing a predetermined application program, and processing the microarchitecture events corresponding to each function of the first microarchitecture by using the method in the first aspect to obtain a first result, where the first result is a performance analysis result of each function running on the first microarchitecture; obtaining microarchitecture events generated by each function of a second microarchitecture executing the predetermined application program, and processing the microarchitecture events corresponding to each function of the second microarchitecture by using the performance analysis method in the first aspect to obtain a second result, where the second result is a performance analysis result of each function running on the second microarchitecture; where the first microarchitecture and the second microarchitecture are different types of processor microarchitectures; and determining a performance comparison result between the first microarchitecture and the second microarchitecture according to the first result and the second result.

[0006] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program, the method according to any one of the embodiments of the present application is implemented.

[0007] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method according to any one of the embodiments of the present application is implemented.

[0008] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the method according to any one of the embodiments of the present application is implemented.

[0009] According to the performance analysis method of the microarchitecture in the embodiment of the present application, microarchitecture events generated by each function of the microarchitecture executing a predetermined application program can be collected, and then based on the preset index evaluation model of the microarchitecture and the microarchitecture events corresponding to each function, the respective predetermined performance indexes corresponding to each function are determined; thereby, performance analysis is performed according to each predetermined performance index, and a performance analysis result of each function running on the microarchitecture is obtained.

[0010] The method of the present application collects microarchitecture events generated by each function of the microarchitecture executing a predetermined application program, and determines the predetermined performance indexes corresponding to each function based on the preset index evaluation model, which avoids the complexity of using multiple tools to collect events and evaluate performance indexes respectively. Moreover, according to this method, performance analysis can be performed according to the respective predetermined performance indexes corresponding to each function, and a performance analysis result of each function running on the microarchitecture is obtained, realizing microarchitecture performance analysis at the function level, so that the performance analysis result can be refined to the granularity of the function level. Furthermore, the method of the present application can collect microarchitecture events corresponding to each function specifically executed based on the execution of the predetermined application program by the microarchitecture, rather than simply obtaining event-related information when a specific microarchitecture event occurs. Considering that each function in the application program execution process is often called multiple times, according to the method of the present application, the microarchitecture events of each function can be obtained more completely, and more complete function execution information can be obtained, which is beneficial to better evaluating the execution situation of the function and improving the accuracy of the performance analysis result.

[0011] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present application more obvious and understandable, the specific embodiments of the present application are hereinafter specifically exemplified. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the several views denote the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments in accordance with the present application and should not be considered as limiting the scope of the present application.

[0013] Figure 1 A flowchart of a method for performance analysis of a microarchitecture according to an embodiment of the present application is shown.

[0014] Figure 2 A detailed flowchart of a method for performance analysis of a microarchitecture according to an exemplary embodiment of the present application is shown.

[0015] Figure 3 A flowchart of a method for performance analysis of a microarchitecture according to another embodiment of the present application is shown.

[0016] Figure 4 A schematic structural diagram of a performance analysis device for a microarchitecture shown in an embodiment of the present application is shown.

[0017] Figure 5 A schematic structural diagram of a performance analysis device for a microarchitecture shown in another embodiment of the present application is shown.

[0018] Figure 6 A block diagram of an electronic device provided in an embodiment of the present application is shown. Detailed implementation manners

[0019] In the following, only some exemplary embodiments are simply described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the concept or scope of the present application. Therefore, the drawings and the description are considered to be exemplary in nature and not restrictive.

[0020] To facilitate understanding of the technical solutions of the embodiments of the present application, the related technologies of the embodiments of the present application are described below. The following related technologies can be arbitrarily combined with the technical solutions of the embodiments of the present application as optional solutions, and all of them fall within the protection scope of the embodiments of the present application.

[0021] In cloud computing services, cloud service providers can integrate resources through pooling technology for centralized resource management and dynamic allocation. Specifically, physical resources (such as servers, storage devices, network devices, etc.) are centralized through virtualization technology to form a shared resource pool, aiming to improve resource utilization and scalability and reduce the overhead of resource allocation and release. Moreover, after integrating resources through pooling technology, cloud service providers can sell resources of different specifications to users flexibly. Resources of different specifications refer to the computing resources with various configurations provided by cloud service providers according to parameters such as computing power, memory size, storage space, and network bandwidth. Flexible selling is a sales model that dynamically adjusts resource allocation and billing according to user needs. This model allows users to use and pay for the above-mentioned computing power, storage space, etc. on demand, thus optimizing costs and enhancing flexibility. Selling these resources of different specifications to users flexibly can meet the resource usage requirements in different scenarios. That is to say, users usually expect to select a cloud server that can run a certain application with higher performance. To achieve this goal, users often need to have the ability to comprehensively evaluate the execution efficiency of the program on different cloud servers, select a cloud server with great optimization potential, and perform corresponding customized optimization.

[0022] In actual scenarios, the performance of cloud servers depends to a large extent on their underlying hardware, especially the microarchitecture of the processor. Therefore, evaluating the execution efficiency of an application on different cloud servers also relies on the performance evaluation results of the microarchitecture. For most developers, the performance evaluation of the microarchitecture needs to be completed relying on existing performance tools. In related technologies, evaluated from the two dimensions of professionalism and generality, the higher the professionalism of a tool, the worse its generality performance usually is, resulting in the inability to be uniformly used across multiple platforms and making it difficult to unify the analysis between platforms. And the general-purpose tools that can be used for horizontal evaluation and comparison across multiple platforms can only cover some macroscopic indicators, and such indicators are difficult to play a specific guiding role in the performance evaluation and optimization of the microarchitecture. Therefore, developers usually need to rely on personal understanding to seek the help of multiple tools for microarchitecture performance evaluation, resulting in a complex and lengthy performance evaluation process, and the applicability of different performance evaluation tools is not clear. If developers use tools without understanding their applicable ranges, they may misuse the tools, thus obtaining inaccurate or irrelevant performance data, which is likely to lead to poor reliability of the evaluation results.

[0023] In the related art, a performance analysis tool (linuxperf) using performance counters can be used in an open-source system (Linux) for performance analysis of a micro-framework. A performance counter is a kernel-based subsystem that can provide a performance analysis framework, using a Performance Monitoring Unit (PMU), tracing tools, and counters in the kernel for performance analysis. Among them, the performance monitoring unit includes a set of event counters that are used to track and count underlying hardware events in the system, the tracing tools are used to capture microarchitecture events, and the counters are used to count microarchitecture events.

[0024] The performance analysis tool can read the hardware events detected by the performance monitoring unit in two ways. One way is to read the microarchitecture event count instantaneously at the start and end of the program, and record the difference in the number of microarchitecture events that occurred during this period. Another way is to read the microarchitecture event count at regular intervals, and the reading timing depends on the timing when the microarchitecture event count generates an overflow interrupt, which is issued after the microarchitecture event count reaches a certain fixed threshold.

[0025] It can be seen that this performance analysis tool can only determine the performance bottleneck at the program level, and cannot count the execution information refined to the function level, so it cannot determine the performance bottleneck at the function granularity, making it difficult to guide developers to optimize the program more deeply. Also, this performance analysis tool can only focus on a single performance evaluation index and cannot jointly process multiple performance evaluation indexes, making it difficult for developers to understand the specific execution efficiency of functions through the microarchitecture analysis method.

[0026] In the related art, a specific microarchitecture can provide multiple microarchitecture events that can be collected through the performance monitoring unit. However, it is very difficult to select truly useful events from these numerous events for analysis, which usually requires in-depth understanding of the microarchitecture design and the usage specifications of the performance monitoring unit to obtain useful information from the numerous originally collected microarchitecture event data.

[0027] It should be noted that the above application scenarios or application examples provided in the embodiments of the present application are for easy understanding, and the embodiments of the present application do not make specific limitations on the application of the technical solutions. In addition, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0028] The following uses specific embodiments to elaborate in detail on the technical solution of the present application and how the technical solution of the present application solves the foregoing technical problems. The several specific embodiments listed may be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The following will describe the embodiments of the present application in detail with reference to the drawings.

[0029] Figure 1 The flowchart of the performance analysis method for the microarchitecture of the embodiment of the present application is shown, as Figure 1 shown, the method may include the following steps S101 - S103.

[0030] S101, respectively collect the microarchitecture events generated by each function of the microarchitecture executing a predetermined application program.

[0031] S102, based on the preset metric evaluation model of the microarchitecture and the microarchitecture events corresponding to each function, determine the respective predetermined performance metrics corresponding to each function.

[0032] S103, perform performance analysis according to each predetermined performance metric to obtain the performance analysis results of each function running on the microarchitecture.

[0033] According to the performance analysis method for the microarchitecture of the embodiment of the present application, the microarchitecture events generated by each function of the microarchitecture executing a predetermined application program can be collected, and then based on the preset metric evaluation model of the microarchitecture and the microarchitecture events corresponding to each function, the respective predetermined performance metrics corresponding to each function can be determined; thereby, performance analysis is performed according to each predetermined performance metric to obtain the performance analysis results of each function running on the microarchitecture. Compared with the related art that requires the use of multiple performance evaluation tools, the method of the present application can directly collect the microarchitecture events generated by each function of the microarchitecture executing a predetermined application program, and determine the predetermined performance metrics corresponding to each function based on the preset metric evaluation model, which avoids the complexity of using multiple tools to collect events and evaluate performance metrics respectively. Moreover, according to this method, performance analysis can be performed according to the respective predetermined performance metrics corresponding to each function to obtain the performance analysis results of each function running on the microarchitecture, realizing function - level microarchitecture performance analysis, and making the performance analysis results refined to the function - level granularity. Also, the method of the present application can collect the microarchitecture events corresponding to each function specifically executed during the execution of a predetermined application program by the microarchitecture, rather than simply obtaining event - related information when specific microarchitecture events occur. Considering that each function in the application program execution process is often called multiple times, according to the method of the present application, the microarchitecture events of each function can be obtained more completely, and more complete function execution information can be obtained, which is beneficial to better evaluating the execution situation of the function and improving the accuracy of the performance analysis results.

[0034] In step S101, the microarchitecture events are hardware events defined and tracked by the performance detection unit. These events are used to detect and analyze various performance metrics and behaviors of the processor when executing each function included in a predetermined application. In some scenarios, the performance detection unit can also be referred to as a hardware performance counter.

[0035] Exemplarily, the microarchitecture events include, but are not limited to, at least one of the following hardware events: instruction-related events, cache-related events, clock cycle-related events, and floating-point operation-related events. Specifically, instruction-related events can include, for example, events such as the total number of executed instructions and the number of branch instructions. Cache-related events can include, for example, events such as the number of cache accesses, the number of cache hits and misses. Clock cycle-related events can include, for example, events such as the number of processor clock cycles and the number of instruction execution cycles. Floating-point operation-related events can include, for example, events such as the number of executed floating-point operation instructions and the throughput of floating-point operations.

[0036] From the specific examples of the above microarchitecture events, it can be seen that the microarchitecture events are dynamically generated and reflect the usage of hardware resources by the corresponding functions during execution. Therefore, the microarchitecture events can be collected during the execution of each function of the predetermined application.

[0037] In the above step S101, the microarchitecture events can include multiple collection methods. For example, the microarchitecture events corresponding to each function can be periodically collected according to a predetermined collection period. For example, the microarchitecture events corresponding to each function can be collected within a specific time period. For another example, these microarchitecture events can be continuously collected during the running of the predetermined application.

[0038] In some embodiments, step S101 can specifically include the following steps: identifying the microarchitecture that executes the predetermined application according to the running environment information of the predetermined application; collecting microarchitecture events according to a predetermined collection period to obtain at least one set of call stack information, and any set of call stack information contains at least one piece of collection information, and any piece of collection information is used to indicate: any function executed by the microarchitecture and the microarchitecture events generated by the executed function; wherein, the microarchitecture events collected corresponding to different microarchitectures are at least partially different.

[0039] Specifically, the running environment information refers to various conditions required for the operation of the predetermined application, such as including hardware environment information and software environment information. The microarchitecture is also called the processor architecture or microprocessor architecture, and is the specific implementation method and structural design inside the processor. The microarchitecture can be used to define the mechanism by which the processor executes instructions, processes data, and interacts with other system components.

[0040] Exemplarily, when identifying the microarchitecture, it mainly relies on the hardware environment information required for the operation of a predetermined application. This hardware environment information includes, for example, the model and specifications of the processor. The process of identifying the microarchitecture includes, for example, the following steps: obtaining the model and specifications of the processor through system commands or third-party tools, and based on the obtained model and specifications of the processor, its corresponding microarchitecture can be determined.

[0041] For example, in an open-source operating system (Linux system), the detailed information of the processor can be displayed through a command that displays processor-related information (lscpu command). Again, for example, the detailed information of the processor can be obtained by calling a third-party hardware detection tool. The above-mentioned detailed information of the processor includes at least information such as the model and specifications of the processor.

[0042] It should be understood that the above system commands and third-party tools are merely illustrative. When identifying the microarchitecture, the specific system commands or third-party tools used can be selected according to actual needs, and the embodiments of the present application do not make specific limitations. In the embodiments of the present application, a predetermined application (system or software) can be compatible and run on multiple different microarchitectures, with relatively high flexibility and compatibility. Different microarchitectures can be understood as different types of microarchitectures or different types of processor architectures. In some scenarios, the microarchitectures provided by different microarchitecture manufacturers are also referred to as microarchitectures of different platforms. These types of processors differ in terms of instruction sets, performance characteristics, power consumption, etc. Exemplarily, different server chip manufacturers can provide different microarchitectures. Or, the same server chip manufacturer can provide different microarchitectures.

[0043] In step S101, exemplarily, the predetermined collection period can be a predetermined time interval or a predetermined number of events. For example, a collection is performed when a certain time interval is reached, or a collection is performed after a certain number of events. The collection performed according to a predetermined time interval is simple to operate and easy to implement. The collection performed according to a predetermined number of events is conducive to more accurately capturing the occurrence frequency of specific events, so as to better understand the impact of the detected specific events on the microarchitecture performance.

[0044] In actual application scenarios, the implementation method of the predetermined collection period can be selected according to specific detection requirements, and the embodiments of the present application do not make specific limitations.

[0045] Exemplarily, a call stack is a data structure used to store function call information during the running of a program. In the embodiments of the present application, the call stack information is used to record the collection information of microarchitecture events. The collection information can not only record the functions executed by the microarchitecture, but also include the microarchitecture events generated by the executed functions.

[0046] In this embodiment, the microarchitecture can be automatically identified based on the running environment information of the application program, and corresponding performance analysis can be performed, so that the performance analysis method can be applied to different microarchitectures. Compared with the related art that requires the use of multiple performance evaluation tools, the method of the present application is conducive to providing corresponding performance analysis methods for different microarchitectures, improving the compatibility of the performance analysis method among different microarchitectures, and reducing the complexity of the performance evaluation method. Moreover, by collecting microarchitecture events once according to a predetermined collection period, a set of call stack information can be obtained. The call stack information of each group is saved as metadata, that is, the call stack information of each group can be saved in a structured format to a file or a database. The collection information in the call stack information of each group is used as metadata to achieve structured storage, which is convenient for the management and query of the call stack information of each group after the collection period ends, so as to meet the performance analysis requirements of the microarchitecture.

[0047] In some embodiments, step S102 above may specifically include: based on the call stack information of each group, counting the microarchitecture events corresponding to each function to obtain the occurrence times of the microarchitecture events corresponding to each function; based on a preset index evaluation model and the occurrence times of the microarchitecture events corresponding to each function, determining the respective predetermined performance indexes corresponding to each function.

[0048] Exemplarily, for any group of call stack information, obtain any piece of collection information in this group of call stack information, and according to the function indicated by this piece of collection information and the microarchitecture event generated by executing the corresponding function, accumulate the microarchitecture event count to the corresponding function. After performing the above-mentioned accumulation processing of the microarchitecture event count on the call stack information of each group, the occurrence times of the microarchitecture events corresponding to each function of the predetermined application program can be obtained.

[0049] Exemplarily, in the actual application scenario, due to the different structural designs and optimization goals of different microarchitectures, when performing performance analysis on different microarchitectures, the preset index evaluation model and the collected microarchitecture events are also different.

[0050] Among them, the preset index evaluation model is used to indicate: the corresponding relationship between each predetermined performance index and the occurrence times of the microarchitecture events corresponding to each function. The preset index evaluation model may include multiple sub-models. Any sub-model may correspond to a predetermined performance index. Specifically, any sub-model may be used to indicate: the association relationship between a predetermined performance index and the occurrence times of the corresponding microarchitecture events.

[0051] Exemplarily, for any sub-model, the association relationship can be expressed as a calculation expression or an algorithm model. The calculation expression is the index calculation formula, and the algorithm model can be a non-linear association relationship between a predetermined performance index and the occurrence times of corresponding micro-architecture events learned in advance through training data.

[0052] Specifically, for any predetermined performance index, when the association relationship indicated by its corresponding sub-model is an expression, the expression at least includes the following components: output, input, and weight. Among them, the output is the index value of the predetermined performance index. The input is the occurrence times of at least one micro-architecture event that needs to be collected corresponding to the predetermined performance index preset. The weight is the weight coefficient of each corresponding micro-architecture event, and each weight system is used to indicate the influence degree of the corresponding micro-architecture event on the predetermined performance index. For example, the larger the coefficient value, the higher the influence degree. The weight coefficient can be obtained through experience or experiments, for example.

[0053] In some scenarios, the expression can also include a constant term, which is used to represent the baseline performance when no micro-architecture event occurs. For example, when the constant term is zero, it means that when no micro-architecture event occurs, the index value of the predetermined performance index is zero. When the constant term is a non-zero value, it means that even if no micro-architecture event occurs, the micro-architecture has a certain basic performance.

[0054] Specifically, for any predetermined performance index, when the association relationship indicated by its corresponding sub-model is an algorithm model, the model can be trained in the following way. First, collect training data. Specifically, collect data samples containing the index value of the predetermined performance index and related micro-architecture events. These data can be generated by actually running an application program or a simulator. For the occurrence times of micro-architecture events and the index value of the predetermined performance index, they can be collected through at least one of the performance monitoring interfaces provided by the operating system, kernel programming, etc., and no specific limitation is made here. Secondly, select the model type. The model type can be, for example, a polynomial regression model, a neural network model, a support vector machine, etc., and can be specifically set according to actual needs, and no specific limitation is made in the embodiments of the present application. Then, perform model training. Specifically, use the training data to fit a non-linear model, and adjust the model parameters through an optimization algorithm (such as gradient descent) to minimize the prediction error. The prediction error is the difference between the model prediction value and the actual value. By adjusting the model parameters, the difference value is gradually reduced. When the difference value is less than or equal to a predetermined difference threshold, or the number of model training times reaches a preset number threshold, stop the model training to obtain the trained model.

[0055] In the embodiments of the present application, the sub-model for each predetermined performance metric can be a calculation expression or a pre-trained algorithm model. Through the calculation expression or the algorithm model, it can help quantify the impact of microarchitecture events on the performance metric, thereby providing a basis for performance evaluation.

[0056] Exemplarily, each of the predetermined performance metrics includes at least one of the following performance metrics: Frontend Bound, Backend Bound, Bad Speculation, and Retiring.

[0057] Specifically, a predetermined application can be understood as software composed of a series of instructions. Among them, the performance metric "Frontend Bound" is used to describe the performance limitation situation caused by insufficient performance of the processor front end (i.e., the instruction fetching and decoding stages) during the processing of the predetermined application.

[0058] The performance metric "Backend Bound" is used to describe the performance bottleneck caused by insufficient resources in the backend of the processor.

[0059] The performance metric "Bad Speculation" is used to describe the performance loss caused by incorrect prediction by the processor. For example, when the processor encounters a conditional branch, it will try to predict the direction of the branch. If the prediction is incorrect, the processor needs to discard the instructions on the incorrect branch and restart execution from the correct branch. This will cause a pause in the processing flow, thus easily resulting in a waste of the processor's execution resources.

[0060] The performance metric "Retiring" is used to describe that the instruction execution result of the processor is written into a register or memory, which means that it is determined that the instruction execution of the processor has entered the final completion stage. If most of the instructions of an application (the proportion of the corresponding number of instructions in the total number of instructions included in the application exceeds a predetermined ratio threshold) can successfully enter the retirement stage, it indicates that the execution efficiency of the application is relatively high.

[0061] In some embodiments, the above-mentioned predetermined performance metrics are based on the microarchitecture performance metrics in the Top-down Microarchitecture Analysis Method. This analysis method uses a top-down analysis framework, which can achieve selective analysis of the performance bottlenecks of the microframework. In this top-down analysis framework, the microarchitecture performance metrics are hierarchically divided. In each layer, the first layer includes the four performance metrics of front-end limitation, back-end limitation, misprediction, and commit described in the above embodiments. Each performance metric in the first layer can be further divided. Taking the performance metric "front-end limitation" as an example, this metric can be divided into two sub-performance metrics: Fetch Latency and Fetch Bandwidth.

[0062] In the embodiments of the present application, different types of microarchitectures can all adopt the top-down microarchitecture analysis method for performance analysis. The difference is that for the same performance metric, at least some of the sub-models in the preset metric evaluation model adopted and the microarchitecture events that each sub-model needs to collect are different. For example, the same performance metric can correspond to different calculation expressions in different microarchitectures, and the calculation expressions are supported by different microarchitecture events for calculation.

[0063] For example, any sub-model can correspond to a predetermined performance metric. Suppose a certain performance metric in the first layer can be expressed as Y1. Among the multiple sub-models included in the preset metric evaluation model, the sub-model corresponding to the performance metric Y1 can be expressed as: Y1 = K1*A1 + K2*A2 + …… + Kn*An, where n is an integer greater than or equal to 1.

[0064] In this expression, A1, A2, ……, An can represent the collected microarchitecture events, and at least some of the microarchitecture events collected by different microarchitectures are different. K1, K2, ……, Kn are the weight coefficients corresponding to the respective microarchitecture events in this sub-model. Among different microarchitectures, even if the same microarchitecture event is used in the sub-model of the same predetermined performance metric, the weight coefficients of the same microarchitecture event can be different.

[0065] In the embodiments of the present application, there are differences in the hardware design of different types of microarchitecture designs. The specific implementation of the hardware design includes but is not limited to at least one of the following: cache structure, memory structure, number of execution units, type of execution units, and network structure, etc.

[0066] Taking the cache structure as an example, different microarchitectures may adopt caches with different storage capacities, which will affect the definition of cache-related microarchitecture events. For example, the caches of microarchitectures can be divided into small cache designs and large cache designs. A small cache design means that the cache capacity is relatively small, such as the L1 cache. The storage capacity of the L1 cache ranges from 128 kilobytes (KB) to 2 megabytes (MB). A large cache design means that the cache capacity is relatively large. For example, the L2 cache and the L3 cache. The storage capacity of the L2 cache ranges from 256 kilobytes to 32 megabytes. The storage capacity of the L3 cache ranges from 1 megabyte to 128 megabytes.

[0067] If the cache of a microarchitecture adopts a small cache design, the cache miss rate may be higher, and the microarchitecture events that need to be collected include cache miss events, which can be used to evaluate the performance evaluation index of "front-end limited". If the cache of a microarchitecture adopts a large cache design, the cache miss events may be fewer, but the impact of cache access latency on performance is greater. The microarchitecture events that need to be collected include but are not limited to at least one of the following: cache hit rate time, cache line utilization events, which can be used to evaluate cache efficiency, and these events are usually related to the performance index of "back-end limited", and specifically can belong to the "memory limited" category in the performance index of "back-end limited".

[0068] Taking the memory structure as an example, the memory structure includes a unified memory architecture and a heterogeneous memory architecture. A unified memory architecture means that the memory resources are uniformly managed, and each processing unit shares the same memory. A heterogeneous memory architecture means that different types of processing units have their own memories.

[0069] For the unified memory architecture, since each processing unit shares the same memory, all memory access latencies are relatively low. However, since the bandwidth is shared by all processing units, bandwidth bottlenecks are likely to occur. Therefore, the microarchitecture events that need to be collected include: memory access latency, memory bandwidth utilization, and the number of memory method conflicts, etc. For the heterogeneous memory architecture, its characteristics are: the access latencies of different types of memories vary greatly and need to be optimized to reduce latency. And the bandwidths of different types of memories may be different and need to be reasonably allocated to improve efficiency. Therefore, the microarchitecture events that need to be collected include: the access latencies of different types of memories, the bandwidth utilization of different types of memories, and the number of memory page migrations. The number of memory page migrations is used to record the number of migrations of memory pages between different types of memories.

[0070] From the above examples, it can be seen that due to differences in hardware design, different types of microarchitectures result in at least partially different microarchitecture events required for evaluating predefined performance metrics. For example, the microarchitecture events required for different types of microarchitectures can be completely different (refer to the above description taking the cache structure as an example), or partially different (refer to the above description taking the memory structure as an example), and the embodiments of the present application do not make specific limitations.

[0071] In the embodiments of the present application, for different types of microarchitectures, the preset index evaluation model of the microarchitecture and the microarchitecture events to be collected can be obtained by looking up a table. However, for any type of microarchitecture, the preset index evaluation model and the microarchitecture events to be collected in the table are not subjectively set, but are based on objective data and scientific methods.

[0072] Exemplarily, first, the specific details of different types of microarchitecture designs can be obtained by querying the hardware design documents. Second, by running the same benchmark program on different microarchitectures, for any microarchitecture, the microarchitecture events and performance metric data collected during the running of the benchmark program. According to the experimental results, determine which microarchitecture events have the strongest correlation with specific performance metrics. During the experiment, the performance monitoring interfaces provided by the operating system, kernel programming, etc. can be used to collect microarchitecture events and performance metric data. The collection method can be specifically selected according to actual needs, and the embodiments of the present application do not make specific limitations. Third, use methods such as regression analysis and machine learning to fit the corresponding relationships (which can be calculation expressions or models) between different submodels and microarchitecture events based on the collected data to ensure the objectivity of the corresponding relationships. Then, use methods such as cross-validation to verify the accuracy and reliability of the model. Among them, the benchmark program is a dedicated program for testing hardware performance and can be customized according to actual needs. Cross-validation means dividing the data set into a training set and a test set. Train the model or fit the calculation expression on the training set and verify the accuracy of the model on the test set.

[0073] The embodiments of the present application can ensure the accuracy of the corresponding relationship between the performance metrics and microarchitecture events under different microarchitecture designs through specific implementation means such as hardware document analysis, experimental verification, and model fitting. This method not only relies on the characteristics of the hardware design but also is verified by experimental data, more accurately reflecting the corresponding relationship between each predetermined performance metric and microarchitecture events on different types of microarchitectures, and having high objectivity and adaptability.

[0074] It should be noted that the expressions of the submodels corresponding to the above performance metrics are only illustrative. In actual scenarios, the preset index evaluation model and its included submodels can have various expression forms, which can be linear models, non-linear models, or a mixed model of linear and non-linear. Specifically, the specific corresponding preset index evaluation models and the required microarchitecture events to be collected for different microarchitectures can be designed according to actual situations, and the embodiments of the present application do not make specific limitations.

[0075] In the top-down microarchitecture analysis method adopted by the microarchitecture of the embodiments of the present application, different preset index evaluation models can be designed in advance for different microarchitectures. The microarchitecture includes one or more predetermined performance indicators, and any one of the predetermined performance indicators can correspond to a sub-model in the preset index evaluation model. For the same predetermined performance indicator of different microarchitectures, the sub-model corresponding to the predetermined performance indicator can be different, and the microarchitecture events that need to be collected in the sub-model can be different. According to the method of the embodiments of the present application, the preset index evaluation models designed for different microarchitectures make the performance analysis more accurate and targeted, can fully consider the characteristics and differences of various microarchitectures, and avoid a one-size-fits-all analysis method. Moreover, between different microarchitectures, the sub-models corresponding to the same predetermined performance indicator and the microarchitecture events that need to be collected can be different, and this flexibility is beneficial to providing rich data support for the performance analysis of different microarchitectures.

[0076] Specifically, after determining the sub-model corresponding to any one of the predetermined performance indicators in each microarchitecture and the microarchitecture events that need to be collected according to the above embodiments, a correspondence set can be established, such as a data table. The data items in the data table include: microarchitecture identifier, each predetermined performance indicator and its corresponding sub-model, and the list of microarchitecture events that need to be collected. Among them, the microarchitecture identifier is used to indicate the microarchitecture, which can specifically be the microarchitecture name or microarchitecture type, and the embodiments of the present application do not make specific limitations. The list of microarchitecture events that need to be collected is a set of microarchitecture events, including all the microarchitecture events that need to be collected for the corresponding microarchitecture. After identifying the microarchitecture, the sub-model corresponding to each predetermined performance indicator of the microarchitecture and the microarchitecture events that need to be collected can be obtained by looking up the table.

[0077] In some embodiments, the step of counting the microarchitecture events corresponding to each function based on each group of call stack information to obtain the occurrence times of the microarchitecture events corresponding to each function can specifically include: generating a call relationship tree using the function call relationships of each function. The attribute information of any leaf node in the call relationship tree includes: the function name of any function and the occurrence times of the corresponding microarchitecture events; for the function and the microarchitecture event indicated by any piece of collection information, adding the occurrence times of the microarchitecture event to the target leaf node, and the target leaf node is the leaf node where the function is located; summarizing the occurrence times of the microarchitecture events accumulated in each leaf node to obtain the occurrence times of the microarchitecture events corresponding to each function.

[0078] Exemplarily, any piece of collection information included in the call stack information can be used to indicate: any function executed by the microarchitecture, the microarchitecture event generated by executing the function, and the function call relationship of the function. Among them, the function call relationship is used to indicate: the call and called relationship between any function and other functions.

[0079] Exemplarily, the call relationship tree is a tree structure used to represent the hierarchical relationship of function calls. The root node of the tree can be the entry function of a predetermined application, and other nodes represent functions that are called in sequence starting from the entry function. In the call relationship tree, each leaf node represents a function, and the edge represents the call relationship between functions, from the caller to the callee. The hierarchical structure of the tree is used to indicate the nested relationship of function calls, for example, a child node represents a function called by its parent node.

[0080] In this embodiment, for each piece of collected information contained in any group of call stack information, the function call relationship of each function is obtained from each piece of collected information, and a complete call relationship tree is established based on the function call relationship of each function, that is, a tree model can be used to model the function call relationship of the entire predetermined application program, and the call relationship tree helps to clearly display the call hierarchy and structure between functions. Accumulating the number of occurrences of micro-architecture events to the target leaf node is conducive to accurately counting the micro-architecture events generated by the execution of each function by the micro-architecture. After summarizing the number of occurrences of micro-architecture events of each leaf node, the number of occurrences of micro-architecture events corresponding to each function can be obtained, and the subsequent calculation of each predetermined performance indicator corresponding to each function provides accurate data support, realizing function-level micro-architecture performance evaluation. In addition, it is possible to combine multiple predetermined performance indicators of a function to dig out more function execution information, such as micro-architecture events during function execution, so as to perform multi-faceted performance evaluation of the micro-architecture.

[0081] In some embodiments, the above-mentioned step of determining each predetermined performance indicator corresponding to each function based on the preset indicator evaluation model and the number of occurrences of the micro-architecture events corresponding to each function may specifically include: obtaining each sub-model in the preset indicator evaluation model, any sub-model is used to indicate the correlation between any performance indicator and the number of occurrences of the corresponding micro-architecture event; for any function, based on the number of occurrences of the micro-architecture event corresponding to the function and the correlation indicated by each sub-model, calculate the predetermined performance indicators corresponding to the function.

[0082] Exemplarily, the attribute information of the leaf node corresponding to the function may also include: at least one predetermined performance indicator corresponding to the function, each predetermined performance indicator corresponding to a sub-model in the preset indicator evaluation model. After the number of occurrences of the micro-architecture event is accumulated to the leaf node of the corresponding function, the number of occurrences of the micro-architecture event corresponding to each function can be obtained.

[0083] In this embodiment, each sub-model can be used to calculate a predetermined performance metric. Through each sub-model, the correlation between each predetermined performance metric and the corresponding microarchitecture event can be accurately indicated. For any function, based on the number of occurrences of the corresponding microarchitecture event and the correlation indicated by the sub-model, the metric values of the respective predetermined performance metrics corresponding to the function can be accurately obtained, thereby realizing the quantitative evaluation of the function performance.

[0084] In some embodiments, before step S101, the method further includes: obtaining a preset correspondence set, where any correspondence in the correspondence set is used to indicate the correspondence information between the microarchitecture, the metric evaluation model, and at least one microarchitecture event; taking the identified microarchitecture as the target microarchitecture, and obtaining, from the correspondence set, the metric evaluation model corresponding to the target microarchitecture and the corresponding at least one microarchitecture event as the preset metric evaluation model and each microarchitecture event of the target microarchitecture.

[0085] Exemplarily, the correspondence set can be in the form of a data table such as a "microarchitecture metric set table", or in the form of a knowledge base. Managing the correspondence between different microarchitectures, metric evaluation models, and microarchitecture events through the preset correspondence set helps to quickly query and obtain the metric evaluation model required for evaluating the performance of the microarchitecture and the microarchitecture events to be collected according to the identified microarchitecture, thereby facilitating the provision of corresponding performance analysis methods for different microarchitectures and improving the compatibility of performance analysis methods among different microarchitectures.

[0086] In this embodiment, through the automatic identification of the microarchitecture, according to the identified microarchitecture, the preset corresponding microarchitecture events and the preset metric evaluation model of the microarchitecture can be obtained from the correspondence set. The preset metric evaluation model includes, for example, metric calculation formulas corresponding to the respective predetermined performance metrics. Through these formulas, the respective predetermined performance metrics of the microarchitecture are calculated, providing a data basis for subsequent microarchitecture performance analysis.

[0087] In some embodiments, in step S103, according to the respective predetermined performance metrics corresponding to the functions, various methods can be used for performance analysis to obtain the performance analysis results of the functions running on the microarchitecture.

[0088] For example, single - indicator anomaly detection can be performed. Specifically, for any function, when there is an anomaly in any predetermined performance indicator of the function, it can be determined that the function performance is abnormal. For example, multi - indicator comprehensive analysis can be performed. Specifically, for any function, when the indicator values of at least two predetermined performance indicators of the function are abnormal, it can be determined that the function performance is abnormal. For another example, analyze the predetermined performance indicators that are the key concerns of a specified function. When the indicator value of any predetermined performance indicator that is the key concern of any specified function is abnormal, it can be determined that the function performance is abnormal.

[0089] Among them, the specified functions include, but are not limited to: functions that are called multiple times (exceeding a predetermined call - count threshold) in a predetermined application or functions with high computational complexity (the number of instruction executions exceeds the count threshold, and / or the function execution time exceeds the execution - duration threshold, etc.). For functions that are called multiple times, the predetermined performance indicators that are key concerns include, for example, execution time and cache hit rate; for functions with high computational complexity, the predetermined performance indicators that are key concerns include, for example, the number of instruction executions and branch - prediction failure rate. Specifically, it can be set according to actual needs, and the embodiments of the present application do not make specific limitations.

[0090] After performing the above - mentioned performance analysis, a performance - analysis result can be generated based on the functions with performance anomalies and their corresponding function values. The performance - analysis result can include at least one of the following items: the function name of the function, the predetermined performance indicators with anomalies, and the corresponding indicator values.

[0091] In some embodiments, step S103 above can specifically include the following steps: for any function among all functions, when at least one performance indicator corresponding to the function meets the anomaly condition, it is determined that the operation of the function has performance anomalies. The anomaly condition is used to indicate that the corresponding performance indicator belongs to a preset value range in the case of performance anomalies; the function determined to belong to the case of performance anomalies is used as a problem function, and each performance indicator that meets the anomaly condition corresponding to the problem function is used as a problem performance indicator; a performance - analysis result corresponding to the function is generated based on the problem function and the problem performance indicators.

[0092] Specifically, the performance - analysis result can be generated in the form of display information or a document. The specific content of the performance - analysis result includes at least: the function name of the problem function and the specific value of the problem performance indicator. In some scenarios, the specific content of the performance - analysis result can also include: the normal range of the indicator value, which is used to indicate the predetermined normal value range or expected value range of the problem performance indicator. In some scenarios, the performance - analysis result can be visually displayed in the form of charts, tables, etc. The specific implementation form of the performance - analysis result can be determined according to the actual situation. The embodiments of the present application do not make specific limitations.

[0093] In this embodiment, if a performance metric of a certain function of a predetermined application is abnormal, the function with performance anomalies and the anomaly metrics during the running process can be determined. This is conducive to clearly identifying the problem function and its corresponding problem performance metrics, thereby helping developers quickly locate the root cause of the problem and perform performance optimization and debugging in a targeted manner.

[0094] In some embodiments, after generating the performance analysis result corresponding to the function according to the problem function and the problem performance metrics, the following steps are further included: obtaining any one of the problem performance metrics as the current metric, and obtaining the next-level metrics of the current metric, where the next-level metrics include each sub-performance metric obtained by pre-dividing the current metric; determining each sub-performance metric of the microarchitecture based on the new metric evaluation model and the new microarchitecture events, where the new metric evaluation model is used to indicate the association relationship between each sub-performance metric and the occurrence times of the corresponding new microarchitecture events; in the case where at least one sub-performance metric meets the abnormal condition, taking each sub-performance metric that meets the abnormal condition as the new problem performance metric, and taking any one of the new problem performance metrics as the new current metric, and starting to loop and execute from the step of obtaining the next-level metrics of the current metric until a stop processing instruction is received or the number of each sub-performance metric included in the next-level metrics of the new current metric is zero, and stopping the performance analysis processing of the microarchitecture.

[0095] Exemplarily, the new metric evaluation model and the corresponding at least one microarchitecture event can be queried from the corresponding relationship set described in the above embodiments. Assume that the performance metric of front-end limitation includes two sub-performance metrics, namely latency and bandwidth. If there is an anomaly in the performance metric of front-end limitation of a certain function, in order to further locate the problem, the calculation of sub-performance metrics can continue based on the new metric evaluation model and the new microarchitecture events corresponding to each function.

[0096] Exemplarily, assume that the current metric is already the last-level metric, or during the execution of the performance analysis method of the microarchitecture, in response to the received stop processing instruction, stop the performance analysis processing of the microarchitecture.

[0097] In this embodiment, by analyzing the microarchitecture performance metrics layer by layer, potential performance problems can be excavated clearly, ensuring the comprehensiveness and in-depthness of the analysis process. The method of performing performance analysis on layer-by-layer metrics makes the performance evaluation more targeted, can quickly locate the key performance bottlenecks, and provides a clear direction and basis for performance optimization based on the performance evaluation results.

[0098] According to the performance analysis method of the microarchitecture according to the embodiments of the present application, the microarchitecture can be automatically identified based on the running environment information of the application program, and corresponding performance analysis can be performed, so that the performance analysis method can be applicable to different microarchitectures. Compared with the related art that requires the use of multiple performance evaluation tools, the method of the present application is beneficial to providing corresponding performance analysis methods for different microarchitectures, improving the compatibility of the performance analysis method among different microarchitectures, and reducing the complexity of the performance evaluation method. Moreover, according to this method, performance analysis can be performed according to the respective predetermined performance indicators corresponding to each function, and the performance analysis results of each function running on the microarchitecture can be obtained, realizing microarchitecture performance analysis at the function level, making the performance analysis results refined to the function-level granularity, thereby being beneficial to improving the accuracy and reliability of the performance analysis results.

[0099] Figure 2 FIG. shows a detailed flowchart of the performance analysis method of the microarchitecture according to an exemplary embodiment of the present application. As Figure 2 shown, the method includes the following steps.

[0100] S201, Microarchitecture identification.

[0101] In this step, the microarchitecture of the predetermined application program can be identified based on the running environment information of the predetermined application program.

[0102] S202, Microarchitecture index look-up table.

[0103] In this step, based on the identified microarchitecture, the preset index evaluation model of the microarchitecture and the microarchitecture events to be collected can be obtained by looking up a table. Specifically, the preset index evaluation model that can be obtained by looking up the table can include multiple sub-models. The microarchitecture events to be collected are the microarchitecture event list, which contains all the microarchitecture events required to be collected for the corresponding microarchitecture. The data table looked up in this step is the data table established after determining the sub-model corresponding to any predetermined performance index in each microarchitecture and the required microarchitecture events in the above embodiments. Details are not described herein again.

[0104] Exemplarily, as Figure 2 shown, step S202 can specifically include the following steps.

[0105] S11, Obtain the microarchitecture index set table.

[0106] Exemplarily, the microarchitecture index set table is the corresponding relationship set in the above embodiments of the present application, which is used to indicate the corresponding relationship between different microarchitectures and the index evaluation model and microarchitecture events.

[0107] S12, Look up tables for multiple types of architectures.

[0108] In this step, using the identified microarchitecture, the corresponding metric evaluation model and the microarchitecture events to be collected are obtained from the microarchitecture metric set table. Different microarchitectures correspond to different metric evaluation models and at least partially different microarchitecture events. Exemplarily, as Figure 2 shown, the metric evaluation model can be a metric calculation formula, and different metric calculation formulas are supported and calculated by different microarchitecture events.

[0109] As Figure 2 schematically shown in the different types of microarchitectures "Architecture 1, Architecture 2, and Architecture 3" and the corresponding "metric calculation formulas" and "microarchitecture events" for each architecture, the microarchitecture metric set table stores the metric evaluation models and the microarchitecture events to be collected corresponding to different types of microarchitectures. Each sub-model included in the metric evaluation model can be, for example, the metric calculation formula corresponding to each sub-model. The microarchitecture events to be collected can be all the microarchitecture events required for the corresponding architecture.

[0110] S13, Formula and Event Output.

[0111] In this step, the metric evaluation model and the microarchitecture events required for the current microarchitecture obtained by looking up the table are obtained. The metric evaluation model can include: the metric calculation formulas corresponding to each sub-model.

[0112] S203, Drive Collection.

[0113] In this step, the corresponding microarchitecture events are collected by the driver program according to a predetermined collection period, and multiple sets of call stack information are obtained. The collected call stack information is saved as metadata. Each set of call stack information contains the following collection information: function name, function call relationship, and the occurrence times of the corresponding microarchitecture events.

[0114] S204, Metadata Processing.

[0115] In this step, the saved metadata is processed. Specifically, a complete call relationship tree can be established according to the call relationship between functions. The occurrence times of the corresponding microarchitecture events are accumulated to the leaf nodes of the corresponding functions to obtain the occurrence times of the microarchitecture events corresponding to each function, and subsequently, according to the metric calculation formulas obtained by looking up the table, the function-level microarchitecture metrics are calculated.

[0116] S205, Function-Level Microarchitecture Analysis Completed.

[0117] In this step, for any one of the functions, when at least one performance metric corresponding to the function meets the abnormal condition, it is determined that there is a performance anomaly in the operation of the function. The function determined to have a performance anomaly is regarded as a problem function, and each performance metric that meets the abnormal condition corresponding to the problem function is regarded as a problem performance metric. A performance analysis result corresponding to the function is generated based on the problem function and the problem performance metrics.

[0118] According to the method of the embodiment of the present application, the microarchitecture for executing a predetermined application program can be automatically identified, and the microarchitecture events to be collected and the preset metric evaluation models to be used can be queried and obtained through a preset corresponding relationship set, and the collected microarchitecture events are mapped to the function level, realizing a deep analysis of the microarchitecture performance at the function level. Moreover, the performance analysis result can be used for the optimization of the predetermined application program. According to this performance analysis method, it is beneficial to improve the accuracy and efficiency of problem positioning related to performance bottlenecks in the program optimization process, and it is beneficial for developers to obtain and insight into the differences hidden in the details of function-level code execution from a new perspective.

[0119] Figure 3 The flowchart showing the performance analysis method of the microarchitecture according to another embodiment of the present application is as follows Figure 3 As shown, the method includes the following steps.

[0120] Step S301: Obtain the microarchitecture events generated by each function of the first microarchitecture executing a predetermined application program, and use the performance analysis method in the above embodiment to process the microarchitecture events generated by each function corresponding to the first microarchitecture to obtain a first result, where the first result is the performance analysis result of each function running on the first microarchitecture.

[0121] Step S302: Obtain the microarchitecture events generated by each function of the second microarchitecture executing a predetermined application program, and use the performance analysis method in the above embodiment to process the microarchitecture events generated by each function corresponding to the second microarchitecture to obtain a second result, where the second result is the performance analysis result of each function running on the second microarchitecture; wherein, the first microarchitecture and the second microarchitecture are different types of processor microarchitectures.

[0122] Step S303: Determine the performance comparison result between the first microarchitecture and the second microarchitecture according to the first result and the second result.

[0123] Through the above steps S301 - S303, using the performance analysis method in the above embodiments, the performance of the first micro - architecture and the second micro - architecture are respectively analyzed to obtain the performance analysis results at the function level of both. By comparing the performance analysis results at the function level of both, a performance comparison result is obtained. According to this method, unified micro - architecture performance analysis can be carried out across different micro - architectures, supporting performance comparison between different micro - architectures. For different micro - architectures provided by different cloud servers, according to this method, performance analysis at the function level of different micro - architectures can be completed, helping developers quickly comprehensively evaluate the ability of a predetermined application to execute efficiently on different micro - architectures, select the cloud server belonging to the micro - architecture with better performance, discover the performance bottlenecks of specific functions of the application, and assist in precisely completing the performance optimization work.

[0124] In some embodiments, the above step S303 may specifically include: for any one of the preset predetermined performance metrics, obtaining the metric values of each function based on the performance metric from the first result as the first metric values, and obtaining the metric values of each function based on the performance metric from the second result as the second metric values; comparing the first metric values and the second metric values to obtain the performance comparison results of each function based on the performance metric.

[0125] In this embodiment, the metric values of the same performance metric of the same function can be obtained from the first result and the second result for comparison, so as to obtain the performance comparison result of the function based on the same performance metric in two different micro - architectures. For applications that need to pursue better performance in a diverse computing environment, according to the performance analysis method of the embodiments of the present application, it can provide strong support for selecting a micro - architecture with better performance in terms of efficiency and effectiveness. Through this method, it is beneficial to simplify the work processes of software debugging and tuning based on different micro - architectures, and it is also beneficial to explore important data support for performance analysis of different processor architectures, so as to efficiently identify and deeply analyze the performance bottlenecks at the internal function level of the application program on different micro - architectures provided by heterogeneous computing platforms, assist developers in finding the optimization direction of the application program more quickly and accurately, and promote the software engineering field to move towards a higher level.

[0126] It should be understood that for the specific details in the above steps of the embodiments of the present application, reference can be made to the specific details in the performance analysis method described in the above embodiments in combination with Figure 1 and Figure 2 and have the corresponding beneficial effects, which will not be elaborated here.

[0127] Figure 4 The structure diagram of the performance analysis device of the micro - architecture shown in the embodiments of the present application is shown. This device is used to execute the performance analysis method of the micro - architecture described in combination with Figure 1 and Figure 2 described, asFigure 4 As shown, the performance analysis device includes the following modules.

[0128] The acquisition module 410 is configured to acquire microarchitecture events generated by each function of the microarchitecture executing a predetermined application program respectively.

[0129] The determination module 420 is configured to determine respective predetermined performance indicators corresponding to each function based on a preset metric evaluation model of the microarchitecture and the microarchitecture events corresponding to each function.

[0130] The analysis module 430 is configured to perform performance analysis according to the respective predetermined performance indicators to obtain performance analysis results of each function running on the microarchitecture.

[0131] In some embodiments, the performance analysis device further includes: an identification module, configured to identify the microarchitecture executing the predetermined application program according to the running environment information of the predetermined application program; the acquisition module 410 is specifically configured to: acquire microarchitecture events according to a predetermined acquisition period to obtain at least one set of call stack information, and any one set of call stack information includes at least one piece of acquisition information, and any one piece of acquisition information is used to indicate: any function executed by the microarchitecture and the microarchitecture event generated by the executed function; wherein, the microarchitecture events acquired corresponding to different microarchitectures are at least partially different.

[0132] In some embodiments, the determination module 420 is specifically configured to: count the microarchitecture events corresponding to each function based on each set of call stack information to obtain the occurrence times of the microarchitecture events corresponding to each function; determine the respective predetermined performance indicators corresponding to each function based on the preset metric evaluation model and the occurrence times of the microarchitecture events corresponding to each function.

[0133] In some embodiments, any one piece of acquisition information is further used to indicate: the function call relationship of any function; when the determination module 420 is configured to count the microarchitecture events corresponding to each function based on each set of call stack information to obtain the occurrence times of the microarchitecture events corresponding to each function, it is specifically configured to: generate a call relationship tree by using the function call relationship of each function, and the attribute information of any leaf node in the call relationship tree includes: the function name of any function and the occurrence times of the corresponding microarchitecture event; for the function and the microarchitecture event indicated by any one piece of acquisition information, add the occurrence times of the microarchitecture event to the target leaf node, and the target leaf node is the leaf node where the function is located; summarize the occurrence times of the microarchitecture events accumulated in each leaf node to obtain the occurrence times of the microarchitecture events corresponding to each function.

[0134] In some embodiments, when the determining module 420 determines the respective predetermined performance metrics corresponding to the functions based on the occurrence times of the microarchitecture events corresponding to the preset metric evaluation model and each function, it is specifically configured to: obtain each sub-model in the preset metric evaluation model, where any sub-model is used to indicate the association relationship between any performance metric and the occurrence times of the corresponding microarchitecture event; for any function, calculate the respective predetermined performance metrics corresponding to the function based on the occurrence times of the microarchitecture events corresponding to the function and the association relationships indicated by each sub-model.

[0135] In some embodiments, the performance analysis apparatus further includes an obtaining module, configured to obtain a preset set of corresponding relationships before respectively collecting the microarchitecture events generated by each function of the microarchitecture executing a predetermined application program. Any corresponding relationship in the set of corresponding relationships is used to indicate: the corresponding relationship information between the microarchitecture and the metric evaluation model and at least one microarchitecture event; taking the identified microarchitecture as the target microarchitecture, and obtaining, from the set of corresponding relationships, the metric evaluation model corresponding to the target microarchitecture and the corresponding at least one microarchitecture event as the preset metric evaluation model and each microarchitecture event of the target microarchitecture.

[0136] In some embodiments, the analysis module 430 is specifically configured to: for any function among the functions, when at least one performance metric corresponding to the function satisfies an abnormal condition, determine that the operation of the function has a performance anomaly, where the abnormal condition is used to indicate that the corresponding performance metric belongs to a preset value range in the case of a performance anomaly; taking the function determined to belong to the performance anomaly as the problem function, and taking each performance metric corresponding to the problem function that satisfies the abnormal condition as the problem performance metric; generating a performance analysis result corresponding to the function according to the problem function and the problem performance metric.

[0137] In some embodiments, the analysis module 430 is further configured to: after generating a performance analysis result corresponding to the function according to the problem function and the problem performance metric, obtain any performance metric in the problem performance metric as the current metric, and obtain the next-level metric of the current metric, where the next-level metric includes each sub-performance metric obtained by previously dividing the current metric; determine each sub-performance metric of the microarchitecture based on a new metric evaluation model and new microarchitecture events, where the new metric evaluation model is used to indicate the association relationship between each sub-performance metric and the occurrence times of the corresponding new microarchitecture event; when at least one sub-performance metric satisfies the abnormal condition, taking each sub-performance metric that satisfies the abnormal condition as the new problem performance metric, and taking any performance metric in the new problem performance metric as the new current metric, and looping from the step of obtaining the next-level metric of the current metric until a stop processing instruction is received or the number of each sub-performance metric included in the next-level metric of the new current metric is zero, and stopping the performance analysis processing of the microarchitecture.

[0138] For the functions of each module in each device in the embodiments of the present application, reference may be made to the corresponding descriptions in the above methods, and they have corresponding beneficial effects, which will not be elaborated here.

[0139] Figure 5 The structural schematic diagram of the performance analysis device for the microarchitecture shown in another embodiment of the present application is shown. This device is used to execute the performance analysis method for the microarchitecture described above in combination with Figure 3 description. As Figure 5 shown, the performance analysis device includes the following modules.

[0140] The first processing module 510 is configured to obtain the microarchitecture events generated by each function of the first microarchitecture executing a predetermined application program, and use the performance analysis method in the above embodiments to process the microarchitecture events generated by each function corresponding to the first microarchitecture, so as to obtain a first result, where the first result is the performance analysis result of each function running on the first microarchitecture.

[0141] The second processing module 520 is configured to obtain the microarchitecture events generated by each function of the second microarchitecture executing a predetermined application program, and use the performance analysis method in the above embodiments to process the microarchitecture events generated by each function corresponding to the second microarchitecture, so as to obtain a second result, where the second result is the performance analysis result of each function running on the second microarchitecture; wherein, the first microarchitecture and the second microarchitecture are different types of processor microarchitectures.

[0142] The result comparison module 530 is configured to determine the performance comparison result between the first microarchitecture and the second microarchitecture according to the first result and the second result.

[0143] In some embodiments, the result comparison module 530 is specifically configured to: for any one of the preset predetermined performance indicators, obtain the indicator values of each function based on the performance indicator from the first result as the first indicator values, and obtain the indicator values of each function based on the performance indicator from the second result as the second indicator values; compare the first indicator values and the second indicator values to obtain the performance comparison result of each function based on the performance indicator.

[0144] For the functions of each module in each device in the embodiments of the present application, reference may be made to the corresponding descriptions in the above methods, and they have corresponding beneficial effects, which will not be elaborated here.

[0145] Figure 6 It is a block diagram of an electronic device for implementing the embodiments of the present application. As Figure 6As shown in the figure, the electronic device includes: a memory 601 and a processor 602. The memory 601 stores a computer program that can run on the processor 602. When the processor 602 executes the computer program, the methods in the above embodiments are implemented. The number of the memory 601 and the processor 602 can be one or more. Specifically, the electronic device may further include a communication interface 603 for communicating with external devices and performing data interaction and transmission.

[0146] Specifically, if the memory 601, the processor 602, and the communication interface 603 are implemented independently, the memory 601, the processor 602, and the communication interface 603 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 6 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0147] Optionally, specifically, if the memory 601, the processor 602, and the communication interface 603 are integrated on a single chip, the memory 601, the processor 602, and the communication interface 603 can communicate with each other through an internal interface.

[0148] The embodiment of the present application provides a computer-readable storage medium that stores a computer program. When the program is executed by a processor, the methods provided in the embodiments of the present application are implemented.

[0149] The embodiment of the present application provides a computer program product that includes a computer program. When the program is executed by a processor, the methods provided in the embodiments of the present application are implemented.

[0150] The embodiment of the present application further provides a chip that includes a processor for calling and running instructions stored in a memory, so that a communication device equipped with the chip executes the methods provided in the embodiments of the present application.

[0151] The embodiment of the present application further provides a chip that includes: an input interface, an output interface, a processor, and a memory. The input interface, the output interface, the processor, and the memory are connected through an internal connection path. The processor is used to execute the code in the memory. When the code is executed, the processor is used to execute the methods provided in the embodiments of the application.

[0152] It should be understood that the above-mentioned processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. It is worth noting that the processor can be a processor that supports the advanced RISC machines (ARM) architecture.

[0153] Furthermore, optionally, the above-mentioned memory can include a read-only memory and a random access memory. The memory can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).

[0154] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another.

[0155] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0156] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of the features. In the description of the present application, "a plurality of" means two or more, unless otherwise specifically defined.

[0157] Any process or method described in the flowchart or otherwise described herein can be understood to represent a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process. And the scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed.

[0158] The logic and / or steps described in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices.

[0159] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the method in the above embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. When this program is executed, it includes one or a combination of the steps of the method embodiment.

[0160] In addition, in each embodiment of the present application, each functional unit can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. If the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. This storage medium can be a read-only memory, a magnetic disk, an optical disk, or the like.

[0161] As described above, only the exemplary embodiments of the present application are provided, but the protection scope of the present application is not limited thereto. Any person skilled in the art in the technical field recorded in the present application can easily think of various changes or substitutions within the technical scope recorded in the present application, and these should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A performance analysis method for a microarchitecture, characterized in that: include: respectively collecting micro-architecture events generated by each function of the micro-architecture executing the predetermined application program; Determine the predetermined performance indicators corresponding to the functions based on the preset indicator evaluation model of the micro-architecture and the micro-architecture events corresponding to the functions; The calculation expression of each sub-model in the preset index evaluation model includes: the number of occurrences of each corresponding micro-architecture event, each weight coefficient and a constant term, any weight coefficient is used to characterize the influence of the corresponding micro-architecture event on the corresponding predetermined performance index, and the constant term is used to indicate the performance when no micro-architecture event occurs; each weight coefficient is obtained through experiments; Performing performance analysis according to the predetermined performance indicators to obtain performance analysis results of the functions running on the micro-architecture; Obtain any performance indicator among the problem performance indicators as the current indicator, and obtain the next-layer indicator of the current indicator, wherein the next-layer indicator includes each sub-performance indicator obtained by pre-dividing the current indicator; the problem performance indicator is: each performance indicator that satisfies the abnormal condition corresponding to the problem function, and the problem function is: a function in which at least one corresponding performance indicator satisfies the abnormal condition; the abnormal condition is used to indicate that the corresponding performance indicator belongs to a preset value range in the case of performance abnormality; Determine each sub-performance index of the micro-architecture based on a new index evaluation model and a new micro-architecture event, wherein the new index evaluation model is used to indicate: a correlation relationship between each sub-performance index and the number of occurrences of the corresponding new micro-architecture event; When at least one sub-performance indicator satisfies an abnormal condition, each sub-performance indicator that satisfies the abnormal condition is used as a new problem performance indicator, and any performance indicator among the new problem performance indicators is used as a new current indicator. The process is looped starting from the step of obtaining the next layer indicator of the current indicator until a stop processing instruction is received or the number of sub-performance indicators included in the next layer indicator of the new current indicator is zero, and the performance analysis processing of the microarchitecture is stopped.

2. The method according to claim 1, characterized in that The respectively collecting micro-architecture events generated by the micro-architecture executing each function of the predetermined application program includes: Identifying a micro-architecture for executing a predetermined application program based on the operating environment information of the predetermined application program; The microarchitecture events are collected according to a predetermined collection period to obtain at least one set of call stack information, any set of call stack information contains at least one piece of collection information, and any piece of collection information is used to indicate: any function executed by the microarchitecture and the microarchitecture events generated by the execution of the function; wherein the microarchitecture events collected corresponding to different microarchitectures are at least partially different.

3. The method according to claim 2, characterized in that The step of determining the predetermined performance indicators corresponding to the functions based on the preset indicator evaluation model of the micro-architecture and the micro-architecture events corresponding to the functions includes: Based on each group of call stack information, the micro-architecture events corresponding to each function are counted to obtain the number of occurrences of the micro-architecture events corresponding to each function; Based on the preset indicator evaluation model and the number of occurrences of the micro-architecture events corresponding to the functions, the predetermined performance indicators corresponding to the functions are determined.

4. The method according to claim 3, characterized in that Any of the collected information is also used to indicate: a function call relationship of any of the functions; based on each set of call stack information, the micro-architecture events corresponding to each of the functions are counted to obtain the number of occurrences of the micro-architecture events corresponding to each of the functions, including: Generate a call relationship tree using the function call relationship of each function, wherein the attribute information of any leaf node in the call relationship tree includes: the function name of any function and the number of occurrences of the corresponding micro-architecture event; For any function and micro-architecture event indicated by any piece of collected information, the number of occurrences of the micro-architecture event is accumulated to a target leaf node, where the target leaf node is the leaf node where the function is located; The number of occurrences of the micro-architecture events accumulated at each leaf node is summarized to obtain the number of occurrences of the micro-architecture events corresponding to each function.

5. The method according to claim 3, characterized in that: The determining of the predetermined performance indicators corresponding to the functions based on the preset indicator evaluation model and the number of occurrences of the micro-architecture events corresponding to the functions respectively includes: Obtaining each sub-model in the preset indicator evaluation model, wherein any sub-model is used to indicate the correlation between any performance indicator and the number of occurrences of a corresponding micro-architecture event; For any of the functions, based on the number of occurrences of the micro-architecture event corresponding to the function and the association relationships respectively indicated by the sub-models, the predetermined performance indicators corresponding to the function are calculated.

6. The method according to any one of claims 1 to 5, characterized in that Before respectively collecting micro-architecture events generated by the micro-architecture executing each function of the predetermined application, the method further includes: Obtaining a preset correspondence set, wherein any correspondence in the correspondence set is used to indicate: correspondence information between a micro-architecture and an indicator evaluation model and at least one micro-architecture event; The identified microarchitecture is taken as the target microarchitecture, and an indicator evaluation model and at least one corresponding microarchitecture event corresponding to the target microarchitecture are obtained from the corresponding relationship set as the preset indicator evaluation model and each microarchitecture event of the target microarchitecture.

7. The method according to claim 1, characterized in that The performing performance analysis according to the predetermined performance indicators to obtain performance analysis results of the functions running on the micro-architecture includes: For any function among the functions, when at least one performance indicator corresponding to the function satisfies an abnormal condition, determining that there is a performance abnormality in the operation of the function; The function determined to have performance anomalies is regarded as a problem function, and each performance index corresponding to the problem function and satisfying the anomaly condition is regarded as a problem performance index; A performance analysis result corresponding to the function is generated according to the problem function and the problem performance indicator.

8. A performance analysis method for a microarchitecture, characterized in that: The method comprises: Obtaining microarchitecture events generated by each function of a predetermined application program executed by a first microarchitecture, and processing the microarchitecture events generated by each function corresponding to the first microarchitecture using the method described in any one of claims 1 to 7 to obtain a first result, wherein the first result is a performance analysis result of each function running on the first microarchitecture; Obtaining microarchitecture events generated by each function of the predetermined application program executed by the second microarchitecture, and processing the microarchitecture events generated by each function corresponding to the second microarchitecture using the method described in any one of claims 1 to 7 to obtain a second result, wherein the second result is a performance analysis result of each function running on the second microarchitecture; wherein the first microarchitecture and the second microarchitecture are different types of processor microarchitectures; A performance comparison result between the first micro-architecture and the second micro-architecture is determined according to the first result and the second result.

9. The method according to claim 8, characterized in that Determining a performance comparison result between the first micro-architecture and the second micro-architecture according to the first result and the second result includes: For any performance indicator among the preset predetermined performance indicators, obtaining an indicator value of each function based on the performance indicator from the first result as a first indicator value, and obtaining an indicator value of each function based on the performance indicator from the second result as a second indicator value; The first indicator value and the second indicator value are compared to obtain performance comparison results of the functions based on the performance indicators.

10. An electronic device comprising a memory, a processor and a computer program stored in the memory, wherein the processor implements the method according to any one of claims 1 to 7 or any one of claims 8 to 9 when executing the computer program.

11. A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 or any one of claims 8 to 9 is implemented.

12. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 7 or any one of claims 8 to 9.

Citation Information

Patent Citations

  • Software performance testing method applied in embedded system

    CN101630285A

  • Performance index determination method and device, equipment and storage medium

    CN115248764A

  • Cross-architecture fine-grained operating system performance exception mining method and device

    CN117453502A