Performance analysis method, electronic device, readable medium and program product
By inserting a stub program at any point in any thread, combining it with performance events, and collecting PMU indicators, the problem of the inability to accurately identify fine-grained CPU indicators in existing technologies is solved, and precise CPU load analysis and power consumption control are achieved.
Patent Information
- Application Number
- CN202410269915.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-08
- Publication Date
- 2025-09-09
AI Technical Summary
Existing technologies cannot accurately identify the fine-grained PMU indicators of threads or specific functions executed on the CPU, and cannot achieve fine-grained performance analysis and power consumption control of hardware.
By inserting a stub program at any point in any thread, combining it with performance events pre-registered on each CPU core, and collecting PMU indicators, we can achieve fine-grained performance analysis of any function or thread.
It realizes accurate CPU load analysis between any points in any thread, and supports performance optimization and power consumption control of programs and hardware.
Smart Images

Figure CN120610862A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a performance analysis method, electronic equipment, readable medium, and program product. Background Art
[0002] A performance monitoring unit (PMU) refers to a counter within the system that is used to count the performance indicators of the underlying hardware, such as a counter that counts the performance indicators of the central processing unit (CPU) microarchitecture. The performance indicators recorded by this counter may include the number of CPU instructions, the number of CPU clock cycles, etc., referred to as PMU indicators. PMU indicators can accurately obtain the operating status of the CPU, which can help business parties such as application programs to analyze business codes based on PMU indicators, troubleshoot related problems, and optimize business logic. Obviously, if the business party wants to accurately and timely analyze the CPU operating status or program execution efficiency, it often needs to accurately and real-timely collect PMU indicators to accurately obtain some performance analysis results, such as CPU load.
[0003] However, current performance analysis solutions can only obtain the overall PMU count changes generated by the business side (such as applications, processes, etc.) running on the CPU, that is, they can only obtain the overall overhead of the program running on the CPU, and cannot accurately identify the fine-grained PMU indicators corresponding to the execution of a thread or a specific function on the CPU. For example, it is impossible to obtain the number of CPU instructions or CPU clock cycles consumed by the function executed between any position points of any thread. Summary of the Invention
[0004] The present application provides a performance analysis method, electronic device, readable medium and program product, which can collect multiple PMU indicators during the life cycle of any function or between any points in any thread, and based on the collected PMU indicators, more accurately analyze the function at any point in any thread or the CPU load between any points in any thread, etc., to achieve more fine-grained performance analysis of application functions and more fine-grained performance analysis of hardware such as processors, and can also further guide the optimization of application functions and help electronic devices control power consumption, etc.
[0005] In a first aspect, the present application provides a performance analysis method, applied to an electronic device, the method comprising: detecting a performance analysis request for a processor in the process of processing a target object, wherein the performance analysis request includes a performance indicator analysis interval, and the type of the target object includes a function or a thread; based on the performance analysis request, collecting performance indicators of the processor processing the target object in the performance indicator analysis interval; based on the collected performance indicators, analyzing the relevant performance of the processor executing the target object in the performance indicator analysis interval, and obtaining a performance analysis result of the processor processing the target object.
[0006] For example, the target object may include a target thread or a target function. The target function can be a function executed at any point in the target thread or corresponding thread, or a user-specified function in the target program for analysis, without limitation. Accordingly, the performance indicator analysis interval may correspond to the target function lifecycle or the execution time between any points in the target thread.
[0007] Based on this, the performance analysis method provided in the first aspect can, for example, collect performance indicators between any points in the target function life cycle / target thread, that is, collect performance indicators of the processor processing the target object, so as to perform a more fine-grained performance analysis of the execution efficiency or effect of the target function or target thread between any points. This can also further guide the optimization of application functions and help electronic devices control power consumption, etc. Among them, the above-mentioned performance indicators can, for example, include PMU indicators such as the number of CPU instructions or the number of CPU clock cycles, which are not limited here.
[0008] In a possible implementation of the first aspect above, the performance indicator analysis interval includes: a time period in which the processor processes a first function, where the first function is any function among multiple functions processed by the processor; or a first time period in which the processor processes a first thread, where the first thread is any thread among multiple threads processed by the processor, and the first time period is any time period in which the processor processes the first thread.
[0009] For example, the first function may be the target function described below, and the performance indicator analysis interval corresponds to the time period of the processor indicating the processing of the first function, which may be, for example, the life cycle of the target function.
[0010] For another example, the first thread may be the target thread described below. The performance indicator analysis interval corresponds to a first time period during which the processor processes the first thread, and may include any points within the target thread. These points may include the start and end points of any time period during which the processor processes the target thread.
[0011] Thus, the performance analysis method provided by this application can collect multiple PMU indicators during the life cycle of any function or between any points in any thread. Based on the collected PMU indicators, it can more accurately analyze the CPU load of any function at any point in any thread or between any points in any thread.
[0012] In a possible implementation of the first aspect above, the processing time period for the processor to process the first function includes: a time period corresponding to a start position of the processor processing the first function to an end position of the processor processing the first function.
[0013] For example, the starting position of the processor processing the first function may be a user-specified starting position during the processor processing the target thread, i.e., the starting position corresponding to the aforementioned arbitrary time period. The ending position of the processor processing the first function may be a user-specified ending position during the processor processing the target thread, i.e., the ending position corresponding to the aforementioned arbitrary time period.
[0014] In a possible implementation of the first aspect above, the performance indicators include first-category indicators collected corresponding to the instrumentation program running at at least one instrumentation point where the processor runs, wherein the at least one instrumentation point includes: a location where the instrumentation program is run, determined during the processor processing a target object based on a performance analysis request.
[0015] For example, the first category of indicators mentioned above may include the performance indicators of the switching start position and the start position point corresponding to the target object loaded on the last processor for processing based on perf event collection, which are recorded as the second PMU indicators in step 405 below. In the corresponding calculation formula, the second PMU indicators included in the first category of indicators mentioned above can be used with thread 切换开始 、thread 开始位置点 The first category of indicators mentioned above may also include a performance indicator based on the end point of the target object collected based on the perf event and loaded onto the last processor for processing, which is recorded as the third PMU indicator in step 406 below. In the corresponding calculation formula, the third PMU indicator included in the first category of indicators mentioned above can be expressed as thread 结束位置点 express.
[0016] In a possible implementation of the first aspect above, the instrumentation program includes a BPF program for implementing dynamic instrumentation during processing of a target object by a processor.
[0017] In a possible implementation of the first aspect above, the processor of the electronic device includes one or more CPU cores, and detecting a performance analysis request for the processor in the process of the processor processing a target object includes: detecting a CPU performance analysis request for the one or more CPU cores in the process of processing the target object.
[0018] That is, the performance analysis method provided in this application can be applied to electronic devices that use single-core or multi-core CPUs to perform processing tasks.
[0019] In a possible implementation of the first aspect above, the performance indicators also include a second type of indicators corresponding to performance events related to the target object on each CPU core, wherein the time when the performance event is registered on each CPU core is earlier than the time when the processor processes the target object.
[0020] For example, the second category of indicators mentioned above may include performance indicators of corresponding performance event records triggered by the system call that triggers thread switching, including performance indicators recorded at the beginning of switching and performance indicators recorded at the end of switching, which are recorded as the first PMU indicators in step 404 below. In the corresponding calculation formula, the first PMU indicator included in the second category of indicators mentioned above can be represented by switch 切换开始 、switch 切换结束 express.
[0021] That is, the performance analysis method provided in this application can combine the performance indicators recorded when the system call triggers the thread switch with the performance indicators triggered and obtained when the BPF instrumentation program at a specific insertion point, such as any location, is executed, to achieve fine-grained collection of multiple PMU indicators between any function life cycle or any location point of any thread.
[0022] In a possible implementation of the first aspect above, based on a performance analysis request, performance indicators of the processor processing the target object in the performance indicator analysis interval are collected, including: based on the performance analysis request, collecting at least one first-category indicator and at least one second-category indicator, and the sum of the number of items of the collected first-category indicators and second-category indicators is an even number.
[0023] It is understood that the first and second categories of indicators above correspond to performance indicators collected at the start and end locations, such as PMU indicators, including thread switching start and end points, and the start and end points corresponding to specific instrumentation points. Therefore, the number of PMU indicators collected for the first and second categories of indicators above can generally be an even number.
[0024] It can be understood that in some embodiments, if the value of the counter corresponding to the perf event obtained when the corresponding CPU switches to the target thread before the start position or the end position of the target thread is recorded as the same PMU indicator, then the sum of the number of items of the above-mentioned first category indicators and second category indicators may be an odd number.
[0025] In a possible implementation of the first aspect above, based on the collected performance indicators, the relevant performance of the processor in executing the target object in the performance indicator analysis interval is analyzed to obtain a performance analysis result of the processor processing the target object, including: performing a difference calculation based on the first type of indicators and the second type of indicators to obtain a first CPU load corresponding to the first CPU core executing the target object in the performance indicator analysis interval, and a second CPU load corresponding to the second CPU core executing the target object, wherein the multiple CPU cores include the first CPU core and the second CPU core; and taking the accumulated result of the first CPU load and the second CPU load as the performance analysis result of the processor processing the target object.
[0026] That is, on a multi-core CPU system, the performance analysis result may include the sum of the CPU loads generated by executing the target object on each CPU core.
[0027] In a possible implementation of the first aspect, the performance indicator includes a value of a PMU counter, and the PMU counter includes an instruction counter and / or a clock cycle counter.
[0028] The instruction counter can be used to record the number of CPU instructions, and the clock cycle counter can be used to record the CPU clock cycle.
[0029] In a second aspect, the present application provides an electronic device comprising: one or more processors; one or more memories; one or more memories storing one or more programs, which, when executed by one or more processors, enables the electronic device to execute the performance analysis method provided by the above-mentioned first aspect and various possible implementations of the first aspect.
[0030] In a third aspect, the present application provides a computer-readable medium having instructions stored thereon, which, when executed on a computer, causes the computer to execute the performance analysis method provided by the first aspect and various possible implementations of the first aspect.
[0031] In a fourth aspect, the present application provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the performance analysis method provided by the above-mentioned first aspect and various possible implementations of the first aspect.
[0032] The beneficial effects of the second to fourth aspects mentioned above can be found in the relevant descriptions of the first aspect and various possible implementations of the first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 The figure shows a schematic structural diagram of a CPU internal functional module provided in an embodiment of the present application.
[0034] Figure 2a The figure shows a schematic diagram of the principle of implementing performance analysis at the function level or any thread at any time on a single CPU core provided by an embodiment of the present application.
[0035] Figure 2b The figure shows a schematic diagram of the execution sequence of an instrumented function and an instrumented function on a CPU provided by an embodiment of the present application.
[0036] Figure 2c The figure shows a schematic diagram of the technical principle of implementing dynamic instrumentation provided by an embodiment of the present application.
[0037] Figure 3a The figure shows a schematic diagram of a process for registering performance events on each CPU core provided by an embodiment of the present application.
[0038] Figure 3b The figure shows another principle diagram of implementing performance analysis at the function level or at any thread at any time on multiple CPU cores provided by an embodiment of the present application.
[0039] Figure 4 The figure shows a schematic diagram of the implementation process of a performance analysis method provided in an embodiment of the present application.
[0040] Figure 5 Shown is a structural schematic diagram of a performance analysis tool / device provided in an embodiment of the present application.
[0041] Figure 6 Shown is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application.
[0042] Figure 7 The figure shows a schematic diagram of the software structure of an operating system of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0043] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described in detail below with reference to the accompanying drawings and specific implementation methods.
[0044] In order to facilitate those skilled in the art to understand the solutions in the embodiments of the present application, some concepts and terms involved in the embodiments of the present application are explained below.
[0045] (1) The number of CPU instructions, referred to as the instruction count, refers to the count data of the instruction counter in the CPU. The instruction counter in the CPU is essentially an accumulator register that indicates the number of instructions currently to be executed. When a program is executed, the initial value of the instruction counter is the address of the first instruction in the program. When the program is executed sequentially, the controller first retrieves an instruction from the memory according to the instruction address indicated by the instruction counter, then analyzes and executes the instruction, and at the same time adds 1 to the value of the instruction counter in the CPU, pointing to the address of the next instruction to be executed.
[0046] (2) A CPU clock cycle, also known as a machine cycle, is the time it takes a computer to complete a basic operation. Correspondingly, the number of CPU clock cycles refers to the number of clock cycles during CPU operation. The time it takes for a CPU to fetch an instruction and execute it is called an instruction cycle. The number of machine instructions executed by the CPU is referred to as the CPU instruction count.
[0047] It's understandable that each instruction executed by a CPU may take multiple clock cycles to complete. Generally speaking, the higher the CPU's clock frequency, the shorter the execution time per instruction, and the greater the number of instructions executed per unit time or within a certain period of time. Therefore, a CPU with a higher clock frequency typically executes more instructions in the same amount of time.
[0048] It's important to note that the number of CPU instructions and CPU clock cycles is also affected by other factors, such as the CPU architecture, instruction set, cache size, and memory bandwidth. Therefore, when analyzing CPU performance or thread / function performance based on the number of CPU instructions and / or CPU clock cycles, it's necessary to consider multiple factors.
[0049] It is understood that the electronic device in the embodiments of the present application may also be referred to as a terminal, user equipment (UE), mobile terminal (MT), etc. The electronic device may be a mobile phone, a tablet computer, a computer with wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical surgery, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, etc., without limitation herein.
[0050] Figure 1 According to an embodiment of the present application, a structural diagram of a CPU internal functional module is shown.
[0051] like Figure 1 As shown, the internal functional modules of the CPU include: a CPU core 101, an external cache 102, an internal memory 103, a general purpose unit 104, an accelerator 105 and an input / output control & interface unit 106.
[0052] The CPU is a multi-core system, consisting of n+1 CPU cores, where n is an integer greater than or equal to 0, such as CPUCore0 to CPUCoren. Within a single CPU core, a CPU core microarchitecture exists that can be divided according to logical function. This CPU microarchitecture is designed based on the execution process of the processor instruction set. The execution of a piece of software (i.e., a business entity, such as a program or process) often requires processing on one or more CPU cores.
[0053] However, existing technologies can only capture the overall power consumption of software running on the CPU. They cannot accurately identify fine-grained PMU metrics for threads or specific functions executed on the CPU. For example, it is impossible to obtain the number of CPU instructions or CPU clock cycles consumed by functions executed between any points in any thread. Similarly, it is impossible to implement fine-grained performance analysis and control of hardware (e.g., a single CPU core, multiple CPU cores, etc.), threads running software, or software functions.
[0054] For example, current performance analysis tools integrated into IDEs, such as Xcode Counter, can only collect performance metrics at the level of a complete thread, process, or even application. Therefore, this approach struggles to accurately analyze CPU load at the function level. However, because the performance of software (e.g., programs, processes) is closely related to the CPU load of functions along the entire call chain, this approach's accuracy in analyzing CPU load is also limited.
[0055] For example, other current tools that can collect performance metrics, such as the perf command on Linux, can only collect performance metrics related to the CPU load of the thread at the current snapshot (i.e., the moment of collection), and cannot collect performance metrics in real time to monitor the CPU load of the corresponding thread. Furthermore, this method cannot collect the CPU load corresponding to the function executed on the corresponding thread, that is, it cannot collect performance metrics related to function-level CPU load.
[0056] As an example of software, WeChat TM The program runs on mobile phones and other terminals, and WeChat cannot be obtained based on existing technology. TM It is also impossible to obtain the number of CPU instructions consumed by a thread of a function or task in any time period, nor the number of CPU instructions consumed by a function executed on a thread, and thus it is impossible to obtain the fine-grained overhead of the execution of the above thread or function on the corresponding CPU, so it is impossible to perform WeChat TM The program performs fine-grained power consumption monitoring, load analysis and control. More specifically, WeChat TM For example, some users use the Moments function more frequently, while others use the Video Account function more frequently. TMDifferent functions correspond to different threads and execute different functions, and the number of CPU instructions consumed when executing each thread and each function is also different. Existing technologies cannot obtain the number of CPU instructions consumed by threads corresponding to different functions in any time period, and therefore cannot perform a more fine-grained analysis of the performance experience felt by users when using the function based on the number of CPU instructions. For example, for functions that users feel have a poor response rate, it is impossible to specifically analyze the code segments on the threads that need to optimize the execution logic or the functions that need to be optimized. If the number of CPU instructions consumed by threads corresponding to each function in any time period can be obtained, then the above-mentioned more fine-grained analysis can be achieved.
[0057] Therefore, a technical solution is needed that can accurately, efficiently, and in real time collect PMU indicators at function granularity for performance analysis and power consumption control.
[0058] In order to solve the above problems, the present application provides a performance analysis method, which collects performance indicators (i.e., PMU indicators) at any point in any thread by executing a plug-in program of a specified plug-in point (hereinafter referred to as a hook point) according to the user's performance analysis request for CPU operation or function, thread, etc., or combines the performance events pre-registered on each CPU core to realize the collection of PMU indicators at any point in any thread. Among them, the arbitrary position point may include a specific position for collecting PMU indicators at the function level, and the above-mentioned collected PMU indicators may include PMU indicators when the function is executed to a specific position, etc. Furthermore, based on the above-mentioned collected PMU indicators, the overhead of each function or each thread on one or more CPU cores is calculated to achieve the fine-grained performance analysis goal at the function or thread level. In this way, the electronic device to which the performance analysis method provided by the present application is applied can more accurately analyze the functions executed between any position points of any thread or the CPU load during the execution time period between any position points of any thread, etc., so that it can be further used for performance analysis of programs or hardware and to help electronic devices perform power consumption control, etc.
[0059] Among them, any position point may include a hook point at any starting position and a hook point at any ending position below. The stub program executed by the above hook point can be a program that the user pre-inserts in the specified function or thread according to the performance analysis requirements. The program can implement stubs during the processing of the specified function or thread on the corresponding CPU core, and will not modify the original code executed by the specified function or thread. The above hook point refers to inserting a stub, that is, an stub program or stub code, at a specific location during the execution of the thread, so as to perform specific operations or checks at the specific location, such as collecting PMU indicators. Hook points are usually used in scenarios such as performance monitoring and debugging, which can help developers better understand the program running efficiency, CPU operation status, discover potential problems or generate related tuning or security control strategies.
[0060] The function executed at the specific location described above may be referred to as the target function, and any thread supporting the collection of PMU metrics at any location may be referred to as the target thread. That is, the target thread may be any thread among multiple threads running on the electronic device. In some embodiments, the execution time period between any location points of any thread described above may be referred to as the analysis time period or any time period, and may be arbitrarily specified based on a user's request for performance analysis of a CPU, function, thread, or the like.
[0061] The performance metrics collected above, namely PMU metrics, can include the number of CPU instructions or CPU clock cycles consumed during the execution of the target function or within any time period of the target thread. The following uses the number of CPU instructions (hereinafter referred to as instruction count) as an example to detail the specific implementation process of the performance analysis method provided in this application.
[0062] As an example, for an electronic device with a single CPU core system, such as the above Figure 1 In the illustrated multi-core system with one CPU core when n=0, the CPU load of the target function or the CPU load during the execution period between any points in the target thread can be calculated by collecting the instruction counts recorded by the instruction counter corresponding to any start and end points of the target thread processed by the processor. The CPU load during the execution period between any points in the target thread can be calculated by difference calculation. The time from the start to the end of the target function processing by the processor of the electronic device is the processor's processing period for the target function.
[0063] refer to Figure 2aTaking the target function lifecycle as an example, we can execute a performance metric collection instrumentation program at the beginning of the target function to collect PMU counter data, such as the instruction counter data, recorded as r0. Furthermore, we can execute a performance metric collection instrumentation program at the end of the target function to collect PMU counter data, such as the instruction counter data, recorded as r1. The difference between the collected instruction counts (r1 - r0) is the CPU load of the target function.
[0064] Similarly, taking the execution time period between any points of the target thread as an example, the PMU counter data of the first insertion point, such as the instruction counter data, can be collected by executing the instrumentation program for collecting performance indicators set at any starting position of the target thread, and recorded as r0. In addition, the PMU counter data of the second insertion point, such as the instruction counter data, can be collected by executing the instrumentation program for collecting performance indicators set at any ending position, and recorded as r1. If no thread switching occurs during this period, the difference in the number of instructions collected above (r1-r0) is the CPU load during the execution time period between any starting position and the ending position of the target thread.
[0065] refer to Figure 2b Continuing with the target function as an example, during the process of CPU executing code within the life cycle of the above-mentioned target function, the CPU can execute PMU code 00, target function code, and PMU code 01 in sequence. In some embodiments, the function that triggers the execution of "PMU code 00" and the function that triggers the execution of "PMU code 01" can both be called instrumented functions. "Target function code" is the function that is instrumented. Among them, "PMU code 00" can be the code of the instrumentation program for collecting performance indicators set at the beginning of the target function. When the code is executed, the PMU indicators of the first instrumentation point can be collected, such as the above-mentioned r0. "PMU code 01" can be the code of the instrumentation program for collecting performance indicators set at the end of the target function. When the code is executed, the PMU indicators of the second instrumentation point can be collected, such as the above-mentioned r1.
[0066] It is understood that in the process of collecting PMU indicators on a single CPU core of the above example, in order not to affect the program code running on the target thread or the program code to which the target function belongs, such as modifying and recompiling the program code, the present application can adopt the method of dynamic plugging to realize the purpose of executing the plugging program for collecting performance indicators at a specified position and / or specified time. That is, the plugging program for collecting performance indicators can be inserted into the corresponding program instruction stream by the method of dynamic plugging. Wherein, dynamic plugging refers to inserting a specific code block, i.e., a plugging program, during the operation of the program, so as to obtain a plugging technology of the execution status of the program in real time. The present application can realize dynamic plugging by dynamically inserting a Berkeley Packet Filter (BPF) program when the program is running. For example, when the program is running, the BPF loader can be used to dynamically insert the BPF program into the code or instruction stream loaded and executed by the thread running the program.
[0067] During the process of the processor processing the target thread, the target thread (also known as the user thread) can call the BPF system call to execute the above-mentioned BPF program for dynamic instrumentation. When the target thread calls the BPF system call, the BPF program can be compiled into bytecode and stored in a structure for storing the BPF bytecode. At the same time, the system will replace the first instruction of the instrumentation point set by the user with an interrupt instruction. Then, when the CPU executes the interrupt instruction, the corresponding BPF program will be executed. After the BPF program is executed, the next instruction in the target thread will be executed. Among them. The instrumentation point set by the above-mentioned user can be determined based on the instrumentation location description information input by the developer through the interface provided by the performance analysis tool running on the electronic device. Correspondingly, the above-mentioned "user" in the embodiment of the present application can refer to a developer, or an analyst.
[0068] refer to Figure 2c As shown in the figure, the program code instruction flow originally executed by the target thread is, for example, "cc 3d 81 5e 26 00ff 53 74 6e", where "cc 3d 81 5e 26 00 ff" is the instruction flow at the insertion point set by the user. In response to the user setting, when the target thread calls the BPF system call, the system can replace the first instruction of the user-set insertion point "cc 3d 81 5e 2600 ff" with an interrupt instruction, for example Figure 2c The instruction "83" shown in the figure is replaced. After replacement, the instruction flow including the interrupt instruction executed by the target program is "83 3d 81 5e 26 00 ff". Figure 2cAs shown in the "interrupt handling" section, when the CPU executes the instruction stream including the interrupt instruction, it will "correct the registers and stack" and execute the BPF program to collect the PMU indicators of the instrumentation point, such as the number of instructions recorded by the instruction counter. After the BPF program is executed, the next instruction in the target thread will be executed, for example Figure 2c Instruction "53" is shown.
[0069] In a multi-core system, in order to improve the computing efficiency, each CPU can execute multiple threads in parallel. Figure 1 For a multi-core system with two or more CPU cores corresponding to n ≥ 1 in the multi-core system shown, the CPU load on each core and the cumulative consumption on multiple CPU cores can be calculated to analyze and determine the CPU load during the execution period between any points of the target function or target thread.
[0070] refer to Figure 3a Each CPU can allocate execution time to different threads, and the execution time can be evenly distributed. For example, CPU0 can allocate execution time t1 to each thread, CPU1 can allocate execution time t2 to each thread, and CPU2 can allocate execution time t3 to each thread.
[0071] Before the target thread runs, the electronic device can register a PMU performance event on each core by calling the perf event interface, referred to as a performance event or perf event. The performance event can provide the value of the corresponding PMU counter and / or PMU accumulator to the corresponding executed instrumentation program as the monitored PMU indicator. Among them, the PMU accumulator can collect the accumulated value of the PMU counter data, refer to Figure 3a As shown in the "Count Accumulation" section, it is understood that for the performance analysis process of the target function, a perf event can be uniformly registered on each CPU before the target function is executed to collect PMU indicators at any starting position and any ending position of the target thread.
[0072] Among them, the above performance events refer to some electrical logic related to the internal functional modules of the processor, corresponding to the events generated when the processor performs logical operations, which can be counted and monitored by the PMU unit. Figure 1 In the illustrated CPU internal architecture, the PMU unit may include: a clock cycle counter, a performance event counter, interrupt and overflow registers, system control registers, a PMU register interface, and PMU event interfaces for other functional modules. The performance event counter may correspond to the aforementioned PMU accumulator and is used to count the number of performance events that occur within a preset time range.
[0073] It can be understood that in a multi-core system, if a thread switch occurs during the life cycle of the target function or the execution time period between any points in the target thread, it is necessary to calculate the start position, end position, and PMU counter data of the thread switch start and thread switch end respectively.
[0074] refer to Figure 3b When CPU0 switches to the target thread, calling the sched_switch system call triggers the execution of the instrumentation program, which obtains the value of the counter corresponding to the perf event, i.e., the PMU counter data, recorded as switch0. When the execution time allocated for the target thread on CPU0 ends, the target thread can continue to switch to CPU1. At this time, when the instrumentation program on CPU0 is executed, the value of the counter corresponding to the perf event can be obtained as the PMU counter data collected at any end position of the target thread, recorded as switch1. Correspondingly, the overhead of the target thread executing on CPU0 is (switch1-switch0).
[0075] Similarly, when CPU1 switches to the target thread, calling the sched_switch system call can trigger the execution of the instrumentation program to obtain the value of the counter corresponding to the perf event, that is, the PMU counter data, recorded as switch2. When the execution time allocated for the target thread on CPU1 ends, the instrumentation program executed by CPU1 can obtain the value of the counter corresponding to the perf event as the collected PMU counter data at any end position of the target thread, recorded as switch3. Correspondingly, the execution overhead of the target thread on CPU1 is (switch3-switch2). Afterwards, the target thread can continue to run in another execution time allocated for the target thread on CPU1. If the CPU1 executes to the designated instrumentation point (also recorded as a hook point) on the target thread before the execution ends, the PMU indicator can be collected by executing the BPF program of the designated instrumentation point.
[0076] Continue to refer Figure 3b Before executing to the designated insertion point (i.e., hook point), CPU1 switches to the target thread and calls the sched_switch system call. This triggers the execution of the instrumentation program to obtain the value of the counter corresponding to the perf event. This is used as the collected PMU counter data for the target thread switch start position, recorded as thread0. When CPU1 executes to the designated insertion point (i.e., hook point) in the target thread, it can execute the instrumentation program at that point (e.g., the start point) on the target thread, such as the BPF program. The corresponding collected PMU counter data can be recorded as thread1.
[0077] Correspondingly, the CPU overhead from the switch start position from CPU0 to the target thread execution to the hook point of the plug-in at any start position can be calculated by the following formula (1), denoted as r0, which is the cumulative CPU consumption before any start position (hook point).
[0078] r0=(switch1-switch0)+(switch3-switch2)+(thread1-thread0) (1)
[0079] In other embodiments, after the target thread finishes executing the execution time corresponding to the allocation on CPU1, it can also be switched to CPU2 for execution. Correspondingly, when CPU2 switches to the target thread for execution, calling the sched_switch system call can trigger the execution of the instrumentation program to obtain the value of the counter corresponding to the perf event, that is, the PMU counter data, recorded as switch4. When the execution time allocated to the target thread on CPU1 ends, the instrumentation program executed by CPU2 can obtain the value of the counter corresponding to the perf event as the PMU counter data of any end position of the target thread collected, recorded as switch5. Correspondingly, the overhead of the target thread executing on CPU2 is (switch5-switch4).
[0080] Afterwards, the target thread can continue to switch to CPU0 for execution. At this point, CPU0 calls the sched_switch system call to trigger the execution of the instrumentation program to obtain the value of the counter corresponding to the perf event, which is used as the collected PMU counter data for the target thread switch start position, recorded as thread2. When CPU0 executes to a specified instrumentation point (i.e., a hook point) in the target thread, it can execute the instrumentation program at that point (e.g., the end point) of the target thread, such as a BPF program, and the corresponding collected PMU counter data is recorded as thread3.
[0081] Correspondingly, the CPU overhead from the switch start position when CPU2 switches to the target thread execution to the hook point where the stub is inserted at any end position can be calculated by the following formula (2), denoted as r1, which is the cumulative CPU consumption before any end position (hook point).
[0082] r1=(switch1-switch0)+(switch3-switch2)+(switch5-switch4)+(thread3-thread2) (2)
[0083] Therefore, the CPU load during the execution period between any points (including any start position and any end position) of the target thread can be:
[0084] (r1-r0)=[(switch1-switch0)+(switch3-switch2)+(switch5-switch4)+(thread3-thread2)]-[(switch1-switch0)+(switch3-switch2)+(thread1-thread0)]
[0085] =(switch5-switch4)+(thread3-thread2)-(thread1-thread0)
[0086] Similarly, for the function executed between any points of any thread or the CPU load during the execution time period, the following formula (3) can be used to calculate:
[0087]
[0088] Among them, switch 切换结束 Indicates the value of the perfevent counter obtained at the starting position of the switch when the corresponding CPU switches to the target thread; switch 切换开始 Indicates the value of the counter corresponding to the perf event obtained at the end position of the switch from the corresponding CPU to the target thread. 开始位置点 Indicates the PMU counter data collected by the CPU core executing the instrumentation program at any starting position of the target thread or the hook point of the starting position; thread 结束位置点 Indicates the PMU counter data collected by the CPU core executing the instrumentation program at any end position of the target thread or the end position of the target thread; 切换开始 Indicates the value of the counter corresponding to the perf event obtained at the start position of the switch from the corresponding CPU to the target thread before the start position or end position of the target thread.
[0089] It is understood that the value of the counter corresponding to the perf event (i.e., the PMU indicator) obtained by the instrumentation program executed on the CPU may include data such as the number of instructions consumed by multiple threads executed on the corresponding CPU. Therefore, the overhead of the above-mentioned target thread on each CPU core can be filtered out from the PMU indicators of multiple threads stored based on the identification information of the target thread, such as the thread identity (ID), including the above-mentioned switch. 切换开始, such as switch0, switch2, switch4, etc., as well as the above switch 切换结束 , for example, switch1, switch3, switch5, etc.
[0090] It can be understood that if the life cycle of the target function is executed to the end position on the last CPU core but the execution time allocated by the CPU for the target function has not yet ended, the PMU indicator collection method and CPU load calculation method corresponding to the above formula (3) can also be used for calculation. In this case, the above thread 开始位置点 It can represent the PMU counter data collected when the BPF program with the target function life cycle is executed. In some embodiments, it can also be recorded as "function() 开始位置点 ". Correspondingly, the above thread 结束位置点 It can represent the PMU indicators collected when the BPF program with the target function lifecycle is executed. In some embodiments, it can also be recorded as "function() 结束位置点 ", no restrictions are imposed here.
[0091] It can be understood that for the target thread within the execution time between arbitrary position points, the arbitrary starting position can also be in the process of the target thread being processed on a certain CPU core. At this time, the performance indicators of the hook point of the arbitrary starting position can also be collected, and then combined with any ending position of the target thread executed on the CPU core, the CPU overhead of the target thread on the CPU can also be calculated.
[0092] Based on the process of collecting PMU indicators and calculating CPU load in the above example, the performance analysis method provided in this application can achieve more accurate analysis of the target function and the CPU load of the target thread in any time period, etc., which can be further used for performance analysis of programs or hardware and power consumption control of electronic devices.
[0093] It is understood that the performance analysis method provided by this application uses dynamic instrumentation to collect PMU counter values, etc., without intruding on the execution process of the original program code or original function executed on the target thread, without modifying the structure of the original code, and is independent of the logic before and after the execution of the original code or original function. At the same time, the performance analysis method provided by this application can achieve real-time function-level CPU load analysis by combining dynamic instrumentation technology with performance event-related technologies.
[0094] Figure 4 According to the embodiment of the present application, a schematic diagram of the implementation process of a performance analysis method is shown. Figure 4The execution subjects of each step in the process shown can be electronic devices operated by users to perform performance analysis, such as mobile phones, notebooks, tablet computers, etc., and performance analysis tools / devices can be run on the electronic devices to perform Figure 4 Each step in the process shown is to implement the performance analysis method provided by this application. It is understood that the electronic device can be a multi-core system, that is, a device with multiple CPU cores, and the electronic device can use a target thread to run a user-specified target program to analyze the target thread, and / or the target program can include a user-specified target function for analysis, without limitation herein.
[0095] In other embodiments, the electronic device that runs the performance analysis tool / device and the electronic device that runs the target thread or executes the target function may also be different devices.
[0096] It should also be stated that the steps in the methods and processes in the embodiments of the present application are numbered for ease of reference, rather than to limit the order of precedence. If there is a sequence between the steps, the written description shall prevail.
[0097] 401: Request for obtaining performance analysis of a processor in the process of processing a target object.
[0098] Exemplarily, the processor may include one or more CPU cores configured in the electronic device, and may also include registers / memory corresponding to each CPU core. The target object may include a function or a thread. For example, the target object may include the execution time between any points in a target function or a target thread. The target thread may be a function at a specific location in any thread running in the electronic device or performance analysis tool.
[0099] The performance analysis request may be a request, issued by a user through a performance analysis tool running on an electronic device, to analyze the overhead of a target thread running a target program within any time period, or may be a request to analyze the execution overhead corresponding to any function in the target program. The user-instructed analysis request may be initiated through input or a click operation on the interface provided by the performance analysis tool, without limitation.
[0100] 402: In response to the performance analysis request, register performance events for monitoring the target object on each processor.
[0101] For example, an electronic device using a multi-core system can respond to the above-mentioned performance analysis request and call the perf event interface on each CPU core to register a performance event, which is used to monitor the data changes of the PMU counter of the execution time between any position points of the target function or target thread and collect PMU indicators.
[0102] Refer to the above Figure 3a to Figure 3b As shown, when the execution conditions of the instrumentation program for obtaining the PMU indicators corresponding to the performance events are met, for example, when a thread switching system call is generated on the corresponding CPU core, the PMU indicators can be collected at this location based on the above perf event. It can be understood that the PMU indicators collected when a thread switch occurs on each CPU may include PMU counter data of other threads, such as instruction count data. In this case, the performance analysis tool running on the electronic device can filter out the PMU counter data of the target thread from the PMU indicators collected at the thread switching location on each CPU, or filter out the PMU counter data of the target function.
[0103] It can be understood that after completing the registration of perf events on each CPU core, the electronic device can continue to insert a BPF program that can realize dynamic instrumentation at the specified instrumentation point according to the instrumentation position indicated in the performance analysis request input by the user through the performance analysis tool, such as a BPF program that is inserted at any end position of the execution time between the end position of the target function or any point in the execution time of the target thread, so as to realize the collection of PMU indicators at the corresponding position.
[0104] 403: Process the target object in the execution time allocated to the target object by each processor.
[0105] Exemplarily, each CPU core of the electronic device can load and execute threads or functions in sequence according to the execution time allocated to each thread or function currently executed by the system, including each CPU core loading and executing the target function or target thread within the execution time allocated for the target function or target thread.
[0106] It can be understood that the target function or target thread can start from a CPU core loading the function or thread to start execution, but the end position of the execution time between any points of the target function or target thread may be within the time period when the last CPU core loads the function or thread to start execution, or it may be at any end position of the time period, and there is no restriction here.
[0107] 404: Detecting a system call that triggers thread switching, and collecting the first PMU indicator of the corresponding performance event record.
[0108] For example, during the execution of the target function or target thread, multiple first PMU indicators can be collected. Figure 3bDuring the execution of the target thread, the PMU indicator switch0 of any starting position of the target thread on CPU0 and the PMU indicator switch1 of any ending position of the target thread on CPU0 can be collected. The PMU indicator switch2 of any starting position of the target thread on CPU1 and the PMU indicator switch3 of any ending position of the target thread on CPU1 can also be collected. The PMU indicator switch4 of any starting position of the target thread on CPU2 and the PMU indicator switch5 of any ending position of the target thread on CPU2 can also be collected, and so on.
[0109] 405: Collect the performance index of any starting position of the target object processed on the last processor, and record it as the second PMU index.
[0110] For example, when the target function or the target thread at any position is loaded onto the last CPU core and starts to execute, the PMU indicator of the position can be collected based on the perf event, such as the above Figure 3b The thread0 or thread2 shown is recorded as the second PMU indicator for distinction.
[0111] It can be understood that in the process of processing the execution time between any points of the target function or target thread on the processor, a third PMU indicator on the CPU core can be collected, such as the above-mentioned number of instructions or number of clock cycles.
[0112] It can be understood that in some embodiments, for example, any starting position of the execution time between any position points of the target thread occurs during the process of a CPU core processing the target thread. At this time, it is impossible to collect PMU indicators based on perf events. The electronic device can also collect corresponding PMU indicators by running the plug-in program of the hook point.
[0113] 406: Collect the performance index of any end position of the target object executed on the last processor, and record it as the third PMU index.
[0114] Exemplarily, any end position of the target object executed on the last CPU core may include any end position of the execution time allocated to the target thread on the CPU, and may also include the execution position corresponding to the hook point of the plug-in in the target thread. Therefore, the performance indicators of any end position of the target object executed on the last CPU core may be obtained by executing the plug-in program based on the thread switching position to obtain the PMU indicators corresponding to the perf event, or may be directly collected when executing the plug-in program at the hook point, such as the BPF program. Among them, the PMU indicators collected by the BPF program executed at the hook point may include, for example, the above Figure 3b The thread 1 or thread 3 shown is recorded as the third PMU indicator for distinction.
[0115] It can be understood that during the performance analysis of the execution time between any points of the target function or target thread, a third PMU indicator can be collected.
[0116] It can be understood that the sum of the number of the first PMU indicator, the second PMU indicator, and the third PMU indicator can be an even number. For example, the processor of the electronic device can collect an odd number of first PMU indicators and one second PMU indicator during the process of processing the target object, or it can collect an odd number of first PMU indicators and one third PMU indicator, or it can collect an even number of first PMU indicators, one second PMU indicator, and one third PMU indicator, and so on. Among them, the second PMU indicator and the third PMU indicator can belong to the same category of indicators, for example, they are both PMU indicators collected based on the instrumentation program, recorded as first category indicators; the first PMU indicator can belong to another category of indicators, for example, it is a PMU indicator corresponding to the perf event obtained by executing the instrumentation program, recorded as a second category indicator.
[0117] 407 : Calculate the processor load during the process of processing the target object based on at least two of the collected first PMU indicator, the second PMU indicator, and the third PMU indicator.
[0118] Exemplarily, the electronic device can calculate the CPU load of the target function or target thread based on the above formula (3) based on the multiple first PMU indicators collected in the above step 404, the second PMU indicator collected in the above step 405, and the third PMU indicator collected in the above step 406, as the output function-level or any thread any time period performance analysis result.
[0119] Figure 5 According to an embodiment of the present application, a schematic diagram of the structure of a performance analysis tool is shown. In some embodiments, Figure 5 The performance analysis tool exemplified may also be referred to as a performance analysis device, so in Figure 5 The performance analysis tool / device 500 is taken as an example.
[0120] like Figure 5 As shown, the performance analysis tool / device 500 may include the following structures:
[0121] The performance event registration unit 510 is used to register a performance event for monitoring a target function or target thread, i.e., a perf event, on each CPU core according to a target object, so as to obtain data changes of a PMU counter of the execution time between any points of the target function or target thread and collect PMU indicators.
[0122] The counter configuration unit 520 is configured to configure the relevant parameters or values of the PMU counters and accumulators in conjunction with the perf events registered by the performance event registration unit 510. For example, the target thread's thread ID is configured as identification information in the instruction counter's related data to facilitate filtering data belonging to the target thread in the collected counter data.
[0123] It is understandable that in some embodiments, the functions performed by the above-mentioned performance event registration unit 510 and the counter configuration unit 520 can also be provided by the same functional unit or functional module, which is not limited here.
[0124] Loading unit 530 is configured to load a target function or target thread onto each CPU core for execution. Each CPU core can load and execute the corresponding function within the execution time allocated for each function, including the target function. Similarly, each CPU core can load and execute the corresponding thread within the execution time allocated for each thread, including the target thread. When the execution time on each CPU core polls the target function or target thread, the corresponding CPU can load the target function or target thread based on the functionality of loading unit 530.
[0125] The monitoring unit 540 is configured to monitor performance events on each CPU core. The performance event may be a perf event registered by the performance event registration unit 510 for the target function or target thread. The value of the counter corresponding to the performance event may be obtained by the instrumentation program executed when a thread switch occurs on each CPU core.
[0126] In an embodiment of the present application, the monitoring unit 540 is also used to monitor the CPU's execution event of the BPF program on the hook point on the target function or target thread. The execution event can trigger the execution of the corresponding hook point instrumentation program to obtain the PMU indicator corresponding to the perf event. As mentioned above, the above-mentioned hook point can be a designated instrumentation point set at the start position and / or end position of the target function, or set at any start position and / or any end position of the execution time between any position points of the target thread. The instrumentation point can collect PMU indicators based on the BPF program that implements dynamic instrumentation.
[0127] The calculation unit 550 is used to calculate the CPU load and other indicator analysis results of the execution time between any position points of the target function or target thread based on the PMU indicators corresponding to the performance events monitored by the monitoring unit 540, such as the first PMU indicator, and the PMU indicators corresponding to the perf event obtained according to the BPF program of the corresponding CPU execution hook point, such as the second PMU indicator and the third PMU indicator, in combination with the above formula (3) or the above formulas (1) and (2).
[0128] Output unit 560 is configured to output the calculation results of calculation unit 550 as performance analysis results. In some embodiments, the performance analysis results output by output unit 560 may be numerical values or graphs of analysis results such as CPU load displayed on an interface provided by a performance analysis tool. In other embodiments, the performance analysis results output by output unit 560 may also be converted into an optimization strategy for the target program running on the target function or target thread and displayed to the user. This is not limited to this and will not be further described herein.
[0129] It is understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the performance analysis tool / device 500. In other embodiments of the present application, the performance analysis tool / device 500 may include more or fewer structures than those illustrated. It is understood that the performance analysis tool / device 500 may be presented in the form of an independent software program or an embedded program, may be installed as a third-party application and run on an electronic device, or may be a system program of the electronic device to provide performance analysis functions or services to users.
[0130] Figure 6 According to an embodiment of the present application, a hardware structure diagram of an electronic device is shown. In the embodiment of the present application, the electronic device can be a device with one CPU core or multiple CPU cores, that is, a single-core system device or a multi-core system device.
[0131] like Figure 6As shown, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identity module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0132] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0133] The processor 110 may include one or more processing units, for example: the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units can be independent devices or integrated into one or more processors. In the embodiment of the present application, the processor 110 may include a single-core CPU or a multi-core CPU, which is not limited here.
[0134] The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of instruction fetching and execution.
[0135] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.
[0136] In the embodiment of the present application, the processor 110 can complete instruction fetching and instruction execution through the controller to implement the performance analysis method provided by the present application. It is understood that the processor 110 of the electronic device 100 can also control the operation of the performance analysis tool / device through the controller to implement the performance analysis method provided by the present application.
[0137] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM card interface, and / or a universal serial bus (USB) interface. Among them, the USB interface 130 is an interface that complies with USB standard specifications, and specifically can be a Mini USB interface, a MicroUSB interface, a USB Type-C interface, etc.
[0138] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.
[0139] The charging management module 140 is configured to receive charging input from a charger. The charger can be either a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also provide power to the electronic device via the power management module 141.
[0140] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, and provides power to the processor 110, the internal memory 121, the display 194, the camera 193, and the wireless communication module 160. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some other embodiments, the power management module 141 can also be set in the processor 110. In other embodiments, the power management module 141 and the charging management module 140 can also be set in the same device.
[0141] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.
[0142] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.
[0143] The mobile communication module 150 can provide wireless communication solutions including 2G / 3G / 4G / 5G applied on the electronic device 100.
[0144] The wireless communication module 160 can provide wireless communication solutions for application on the electronic device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication technology (NFC), infrared technology (IR), etc.
[0145] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, so that electronic device 100 can communicate with a network and other devices via wireless communication technologies. The above-mentioned wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The above-mentioned GNSS may include the global positioning system (GPS), the global navigation satellite system (GLONASS), the Beidou navigation satellite system (BDS), the quasi-zenith satellite system (QZSS) and / or the satellite based augmentation system (SBAS).
[0146] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.
[0147] The display screen 194 is used to display images, videos, etc. In the embodiment of the present application, the display screen 194 can be used to display performance analysis results or programs generated based on the performance analysis results, CPU performance tuning strategies or suggestions, etc.
[0148] Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini-LED, a Micro-LED, a Micro-OLED, a quantum dot light-emitting diode (QLED), or the like. In some embodiments, electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1.
[0149] The electronic device 100 can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor.
[0150] The ISP is used to process data fed back by the camera 193. In some embodiments, the ISP can be set in the camera 193.
[0151] The camera 193 is used to capture still images or videos. In some embodiments, the electronic device 100 may include 1 or N cameras 193, where N is a positive integer greater than 1.
[0152] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.
[0153] The internal memory 121 can be used to store computer executable program code, which includes instructions. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area can store data created during the use of the electronic device 100 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the electronic device 100 by running instructions stored in the internal memory 121 and / or instructions stored in a memory provided in the processor.
[0154] The electronic device 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.
[0155] The buttons 190 include a power button, a volume button, and the like. The buttons 190 may be mechanical buttons or touch buttons. The electronic device 100 may receive key inputs and generate key signal inputs related to user settings and function control of the electronic device 100.
[0156] Motor 191 can generate vibration prompts. Motor 191 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playback, etc.) can correspond to different vibration feedback effects. For touch operations acting on different areas of the display screen 194, motor 191 can also correspond to different vibration feedback effects. Different application scenarios (for example: time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.
[0157] The indicator 192 may be an indicator light, which may be used to indicate the charging status, power level changes, messages, missed calls, notifications, etc.
[0158] The SIM card interface 195 is used to connect a SIM card. The SIM card can be connected to and separated from the electronic device 100 by inserting it into or removing it from the SIM card interface 195. The electronic device 100 can support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 at the same time. The types of the above-mentioned multiple cards can be the same or different. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to implement functions such as calls and data communications. In some embodiments, the electronic device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.
[0159] Figure 7 According to an embodiment of the present application, a schematic diagram of the software structure of an operating system of an electronic device is shown.
[0160] It is understood that the operating system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a microservice architecture, or a cloud architecture. The embodiment of the present application takes the Android system of the layered architecture as an example to exemplify the system software structure of the electronic device 100.
[0161] The layered architecture divides the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, Android TM The system is divided into four layers, from top to bottom: application layer, application framework layer, Android TM Runtime (Android TM runtime) and system libraries, as well as the kernel layer.
[0162] like Figure 7 As shown, the application layer may include a series of application packages, including camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message and other applications.
[0163] The application framework layer provides an application programming interface (API) and programming framework for the applications in the application layer. The application framework layer includes some predefined functions.
[0164] The application framework layer may include a window manager, content provider, view system, telephony manager, resource manager, notification manager, etc.
[0165] The window manager is used to manage window programs. The window manager can obtain the display size, determine whether there is a status bar, lock the screen, take screenshots, etc.
[0166] Content providers are used to store and retrieve data and make it accessible to applications. This data can include videos, images, audio, calls made and received, browsing history and bookmarks, phone books, etc.
[0167] The view system includes visual controls, such as those for displaying text and images. The view system is used to build applications. A display interface can consist of one or more views. For example, a display interface containing a text notification icon might include a view for displaying text and a view for displaying images.
[0168] The phone manager is used to provide communication functions of the electronic device 100, such as management of call status (including answering, hanging up, etc.).
[0169] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.
[0170] The Notification Manager allows applications to display notifications in the status bar. These messages can be displayed briefly and then disappear automatically without user interaction. For example, the Notification Manager is used to notify users of completed downloads and message reminders. The Notification Manager can also display notifications in the top status bar of the system as icons or scrolling text, such as notifications from background applications, or as dialog windows on the screen. Examples include text messages in the status bar, beeps, vibrations on electronic devices, and flashing indicator lights.
[0171] Android TM Runtime includes core libraries and virtual machines. Android TM runtime is responsible for Android TM System scheduling and management.
[0172] The core library consists of two parts: one is the function that needs to be called by the Java language, and the other is the Android core library.
[0173] The application layer and application framework layer run in a virtual machine. The virtual machine executes Java files in the application layer and application framework layer as binary files. The virtual machine manages object lifecycles, stack management, thread management, security and exception management, and garbage collection.
[0174] The system library can include multiple functional modules, such as surface manager, media libraries, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.
[0175] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera driver, audio driver, and sensor driver.
[0176] The embodiments of the present application also provide a computer program product for implementing the performance analysis methods provided in the above embodiments.
[0177] The various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or a combination of these implementation methods. The embodiments of the present application can be implemented as computer program modules or module codes executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0178] A computer program module or module code can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.
[0179] Module code can be implemented with high-level modular language or object-oriented programming language to communicate with the processing system. When necessary, module code can also be implemented with assembly language or machine language. In fact, the mechanism described in this application is not limited to the scope of any specific programming language. In either case, the language can be a compiled language or an interpreted language.
[0180] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, instructions may be distributed over a network or through other computer-readable media. Therefore, a machine-readable medium may include any mechanism for storing or transmitting information in a machine (e.g., computer) readable form, including but not limited to a floppy disk, an optical disk, an optical disk, a magneto-optical disk, a read-only memory (ROM), a random access memory (RAM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic card or an optical card, a flash memory, or a tangible machine-readable memory for transmitting information (e.g., a carrier wave, an infrared signal, a digital signal, etc.) using the Internet in an electrical, optical, acoustic, or other form of propagation signal. Accordingly, machine-readable media includes any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a form readable by a machine (eg, a computer).
[0181] References in the specification to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one exemplary implementation or technique disclosed according to the embodiment of the present application. The appearances of the phrase "in one embodiment" in various places in the specification do not necessarily all refer to the same embodiment.
[0182] The disclosure of the embodiments of the present application also relates to an operating device for executing the text. The device can be constructed specifically for the required purpose or it can include a general-purpose computer that is selectively activated or reconfigured by a computer program stored in the computer. Such a computer program can be stored in a computer-readable medium, such as, but not limited to, any type of disk, including a floppy disk, an optical disk, a CD-ROM, a magneto-optical disk, a read-only memory (ROM), a random access memory (RAM), an EPROM, an EEPROM, a magnetic or optical card, an application-specific integrated circuit (ASIC) or any type of medium suitable for storing electronic instructions, and each can be coupled to a computer system bus. In addition, the computer mentioned in the specification can include a single processor or can be an architecture involving multiple processors for increased computing power.
[0183] In addition, the language used in this specification has been primarily selected for readability and instructional purposes and may not be selected to describe or limit the disclosed subject matter. Therefore, the present disclosure of embodiments is intended to illustrate, not to limit, the scope of the concepts discussed herein.
Claims
1. A performance analysis method, applied to electronic equipment, characterized in that: The method comprises: detecting a performance analysis request for a processor in a process of the processor processing a target object, wherein the performance analysis request includes a performance indicator analysis interval, and the type of the target object includes a function or a thread; Based on the performance analysis request, collecting the performance indicator of the processor processing the target object in the performance indicator analysis interval; Based on the collected performance indicators, relevant performance of the processor in executing the target object in the performance indicator analysis interval is analyzed to obtain a performance analysis result of the processor processing the target object.
2. The method according to claim 1, characterized in that The performance indicator analysis interval includes: a processing time period of a first function by a processor, wherein the first function is any function among a plurality of functions processed by the processor; or A first time period when a processor processes a first thread, wherein the first thread is any thread among multiple threads processed by the processor, and the first time period is any time period when the processor processes the first thread.
3. The method according to claim 2, characterized in that The processing time period during which the processor processes the first function includes: A time period corresponding to a start position of the processor processing the first function to an end position of the processor processing the first function.
4. The method according to claim 2, characterized in that The performance indicator includes a first type of indicator collected corresponding to the instrumentation program executed by the processor at at least one instrumentation point, wherein the at least one instrumentation point includes: A location for running an instrumentation program is determined during a process in which the processor processes a target object based on the performance analysis request.
5. The method according to claim 4, characterized in that The instrumentation program includes a BPF program for implementing dynamic instrumentation during the process of the processor processing a target object.
6. The method according to claim 4, characterized in that The processor of the electronic device includes one or more CPU cores, and The detecting a request for analyzing the performance of the processor in the process of the processor processing the target object includes: A CPU performance analysis request is detected for the one or more CPU cores processing a target object.
7. The method according to claim 6, characterized in that The performance indicators also include second-category indicators corresponding to performance events related to the target object on each CPU core, wherein: The time when the performance event is registered on each CPU core is earlier than the time when the processor processes the target object.
8. The method according to claim 7, characterized in that The collecting, based on the performance analysis request, the performance indicator of the processor processing the target object in the performance indicator analysis interval, includes: Based on the performance analysis request, at least one first-category indicator and at least one second-category indicator are collected, and The sum of the number of items of the first category of indicators and the second category of indicators collected is an even number.
9. The method according to claim 8, characterized in that The step of analyzing, based on the collected performance indicators, the performance of the processor in executing the target object in the performance indicator analysis interval to obtain a performance analysis result of the processor processing the target object includes: performing a difference calculation based on the first category of indicators and the second category of indicators, obtaining a first CPU load corresponding to the execution of the target object by the first CPU core and a second CPU load corresponding to the execution of the target object by the second CPU core within the performance indicator analysis interval, wherein the plurality of CPU cores include the first CPU core and the second CPU core; An accumulated result of the first CPU load and the second CPU load is used as a performance analysis result of the processor processing the target object.
10. The method according to any one of claims 1 to 9, characterized in that The performance indicator includes a value of a PMU counter, and the PMU counter includes an instruction counter and / or a clock cycle counter.
11. An electronic device, characterized in that: include: one or more processors; One or more memories; the one or more memories store one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device executes the performance analysis method according to any one of claims 1 to 10.
12. A computer-readable medium, characterized in that The readable medium stores instructions, which, when executed on a computer, cause the computer to execute the performance analysis method according to any one of claims 1 to 10.
13. A computer program product, characterized in that The method comprises a computer program / instruction, which implements the performance analysis method according to any one of claims 1 to 10 when executed by a processor.