Flame pattern generation method and device, storage medium and electronic equipment

By sampling the time information of the target thread on the processor and the time information that is called out/tuned into the processor, a flame diagram is generated, which solves the problem of single flame diagram information in the prior art, resulting in low accuracy, and achieves more accurate performance delay analysis.

CN120104440AActive Publication Date: 2025-06-06ALIBABA CLOUD COMPUTING CO LTD

Patent Information

Application Number
CN202510592235.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-06
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

In the prior art, the single information of the flame map results in low accuracy of the flame map, making it difficult to comprehensively evaluate performance bottlenecks in complex distributed systems or multi-task environments.

Method used

When the target thread is running, it samples its time information on the processor to obtain the first time proportion information; the time information of the target thread being called out and being called into the processor is sampled to obtain the second time proportion information; and then generates a target flame map based on these two time proportion information, which is used to analyze the performance delay information of the target thread.

Benefits of technology

By integrating the first time proportion information and the second time proportion information, the generated target flame map can more accurately reflect the performance delay characteristics of the target thread in the execution and waiting states, overcome the problem of singularity of traditional flame map information, and more accurately identify the causes of the delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104440A_ABST
    Figure CN120104440A_ABST
Patent Text Reader

Abstract

The invention discloses a flame pattern generation method and device, a storage medium and electronic equipment. Relates to the technical field of computers, and the method comprises the following steps: when a target thread runs, sampling time information when the target thread is executed on a processor to obtain first time proportion information corresponding to the target thread; sampling time information of calling the target thread out of the processor and calling the target thread into the processor to obtain second time proportion information corresponding to the target thread; according to the first time proportion information and the second time proportion information, a target flame map is generated, and the target flame map is used for analyzing performance delay information of the target thread. According to the invention, the technical problem of low accuracy of the flame pattern caused by single information contained in the flame pattern in the related art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to computer technology, and in particular, to a method and device for generating a flame graph, a storage medium, and an electronic device. Background Art

[0002] In the field of modern computer system performance analysis, flame graph is a widely used visualization method to reveal the time distribution characteristics of the system operation status. At present, existing flame graphs in the prior art often only include the time distribution characteristics of a single dimension. However, in complex distributed systems or multi-tasking environments, it is often difficult to fully evaluate performance bottlenecks by relying only on a single-dimensional analysis.

[0003] Currently, no effective solution has been proposed to address the technical problem that the flame graph in related technologies contains only a single piece of information, resulting in relatively low accuracy of the flame graph. Summary of the invention

[0004] The embodiments of the present application provide a method and device for generating a flame graph, a storage medium, and an electronic device, so as to at least solve the technical problem in the related art that the flame graph contains single information, resulting in relatively low accuracy of the flame graph.

[0005] According to one aspect of an embodiment of the present application, a method for generating a flame graph is provided, comprising: when a target thread is running, sampling time information of the target thread when it is executed on a processor to obtain first time proportion information corresponding to the target thread; sampling time information of the target thread being called out of the processor and called into the processor to obtain second time proportion information corresponding to the target thread; and generating a target flame graph based on the first time proportion information and the second time proportion information, wherein the target flame graph is used to analyze performance delay information of the target thread.

[0006] Furthermore, sampling the time information of the target thread when it is executed on the processor to obtain the first time proportion information corresponding to the target thread includes: sampling and processing the clock cycle events in the processor to obtain multiple first sampling point information; when sampling the clock cycle events, obtaining the first call stack information corresponding to the target thread at the sampling moment; based on the multiple first sampling point information and the first call stack information, obtaining the first time proportion information.

[0007] Furthermore, sampling the time information of the target thread being called out of the processor and called into the processor to obtain the second time proportion information corresponding to the target thread includes: sampling and processing the process switching event in the processor to obtain multiple second sampling point information; when sampling the process switching event, obtaining the second call stack information corresponding to the target thread at the sampling moment; and obtaining the second time proportion information based on the multiple second sampling point information and the second call stack information.

[0008] Further, generating a target flame graph according to the first time proportion information and the second time proportion information includes: generating a first flame graph corresponding to when the target thread is executed on the processor according to the first time proportion information; generating a second flame graph corresponding to when the target thread is called out of the processor and called into the processor according to the second time proportion information; generating the target flame graph according to the first flame graph and the second flame graph.

[0009] Further, generating, based on the first time proportion information, a first flame graph corresponding to when the target thread is executed on the processor includes: determining a first total duration based on a start count and an end count of a clock cycle event in the processor; calculating, based on the first time proportion information and the first total duration, to obtain first execution duration information of a function in first call stack information corresponding to a first sampling point in the multiple first sampling point information; and generating the first flame graph based on the first execution duration information and the first total duration.

[0010] Further, generating, based on the second time proportion information, a second flame graph corresponding to when the target thread is called out of the processor and called into the processor includes: determining a second total duration based on a start time and an end time of a process switching event in the processor; performing calculation based on the second time proportion information and the second total duration to obtain second execution duration information of a function in the second call stack information corresponding to a second sampling point in the plurality of second sampling point information; and generating the second flame graph based on the second execution duration information and the second total duration.

[0011] Further, generating the target flame graph according to the first flame graph and the second flame graph includes: determining a matching relationship between multiple first call stacks in the first flame graph and multiple second call stacks in the second flame graph, wherein the matching relationship includes: a first matching relationship and a second matching relationship, the first matching relationship indicating that the first call stack is the same as the second call stack, and the second matching relationship indicating that some function call relationships in the first call stack are the same as the second call stack; based on the matching relationship, merging the first flame graph and the second flame graph to obtain the target flame graph.

[0012] Further, based on the matching relationship, merging the first flame graph and the second flame graph to obtain the target flame graph includes: if there is a first target call stack in the multiple first call stacks and a second target call stack in the multiple second call stacks has a first matching relationship, merging the execution time of the first function in the second target call stack into the corresponding position of the first target call stack, and setting the execution time corresponding to the second function of the second target call stack at the top level of the first target call stack, the first function is a function other than the top level in the second target call stack, and the second function is a top level function in the second target call stack; if the multiple first call stacks have a first target call stack and a second target call stack in the multiple second call stacks have a first matching relationship, merging the execution time of the first function in the second target call stack into the corresponding position of the first target call stack, and setting the execution time corresponding to the second function of the second target call stack at the top level of the first target call stack, If a first target call stack exists in a call stack and a second target call stack in the multiple second call stacks has a second matching relationship, the execution time of the third function in the second target call stack is merged into the corresponding position of the first target call stack, and the execution time of the fourth function in the second target call stack is added to the first flame graph according to the position of the fourth function in the second target call stack, the third function is a function in the second target call stack that has the same call relationship as the function in the first target call stack, and the fourth function is a function in the second target call stack other than the third function; after the merging is completed, the first flame graph obtained by merging is proportionally adjusted to obtain the target flame graph.

[0013] According to another aspect of an embodiment of the present application, a flame graph generation method is also provided, including: receiving a flame graph generation request triggered by a client; in a cloud server, based on the flame graph generation request, when a target thread is running, sampling the time information of the target thread when it is executed on a processor to obtain first time proportion information corresponding to the target thread; sampling the time information of the target thread being called out of the processor and called into the processor to obtain second time proportion information corresponding to the target thread; generating a target flame graph based on the first time proportion information and the second time proportion information, wherein the target flame graph is used to analyze the performance delay information of the target thread; and returning the target flame graph to the client.

[0014] According to another aspect of an embodiment of the present application, a flame graph generation device is also provided, including: a first sampling unit, used to sample the time information of the target thread when it is executed on the processor when the target thread is running, and obtain first time proportion information corresponding to the target thread; a second sampling unit, used to sample the time information of the target thread being called out of the processor and called into the processor, and obtain second time proportion information corresponding to the target thread; a generation unit, used to generate a target flame graph based on the first time proportion information and the second time proportion information, wherein the target flame graph is used to analyze the performance delay information of the target thread.

[0015] Furthermore, the first sampling unit includes: a first sampling module, used to sample and process the clock cycle events in the processor to obtain multiple first sampling point information; a first acquisition module, used to obtain the first call stack information corresponding to the target thread at the sampling time when sampling the clock cycle events; a first determination module, used to obtain the first time proportion information based on the multiple first sampling point information and the first call stack information.

[0016] Furthermore, the second sampling unit includes: a second sampling module, used to sample and process the process switching event in the processor to obtain multiple second sampling point information; a second acquisition module, used to obtain the second call stack information corresponding to the target thread at the sampling time when sampling the process switching event; a second determination module, used to obtain the second time proportion information based on the multiple second sampling point information and the second call stack information.

[0017] Furthermore, the generation unit includes: a first generation module, used to generate a first flame graph corresponding to the target thread when it is executed on the processor according to the first time proportion information; a second generation module, used to generate a second flame graph corresponding to the target thread when it is called out of the processor and called into the processor according to the second time proportion information; a third generation module, used to generate the target flame graph according to the first flame graph and the second flame graph.

[0018] Further, the first generating module includes: a first determining submodule, used to determine a first total duration based on a start count and an end count of a clock cycle event in the processor; a first calculating submodule, used to calculate according to the first time proportion information and the first total duration to obtain the first execution duration information of the function in the first call stack information corresponding to the first sampling point in the multiple first sampling point information; and a first generating submodule, used to generate the first flame graph according to the first execution duration information and the first total duration.

[0019] Further, the second generating module includes: a second determining submodule, used to determine a second total duration based on the start time and the end time of the process switching event in the processor; a second calculating submodule, used to calculate according to the second time proportion information and the second total duration to obtain the second execution duration information of the function in the second call stack information corresponding to the second sampling point in the multiple second sampling point information; the second generating submodule is used to generate the second flame graph according to the second execution duration information and the second total duration.

[0020] Further, the third generation module includes: a third determination submodule, used to determine the matching relationship between multiple first call stacks in the first flame graph and multiple second call stacks in the second flame graph, wherein the matching relationship includes: a first matching relationship and a second matching relationship, the first matching relationship indicating that the first call stack is the same as the second call stack, and the second matching relationship indicating that there are some function call relationships in the first call stack that are the same as the second call stack; a merging submodule, used to merge the first flame graph and the second flame graph based on the matching relationship to obtain the target flame graph.

[0021] Furthermore, the merging submodule includes: a first merging submodule, which is used to merge the execution time of the first function in the second target call stack into the corresponding position of the first target call stack if there is a first target call stack in the multiple first call stacks and the matching relationship between the second target call stack in the multiple second call stacks is a first matching relationship, and set the execution time corresponding to the second function of the second target call stack at the top level of the first target call stack, the first function is a function other than the top level in the second target call stack, and the second function is a top level function in the second target call stack; a second merging submodule, which is used to merge the execution time of the first function in the second target call stack into the corresponding position of the first target call stack if there is a first target call stack in the multiple first call stacks and the matching relationship between the second target call stack in the multiple second call stacks is a first matching relationship If the matching relationship between the call stack and the second target call stack in the multiple second call stacks is a second matching relationship, the execution time of the third function in the second target call stack is merged into the corresponding position of the first target call stack, and the execution time of the fourth function is added to the first flame graph according to the position of the fourth function in the second target call stack, the third function is a function in the second target call stack that has the same function call relationship as the function in the first target call stack, and the fourth function is a function in the second target call stack other than the third function; an adjustment sub-module is used to proportionally adjust the merged first flame graph after the merging is completed to obtain the target flame graph.

[0022] According to another aspect of an embodiment of the present invention, there is further provided an electronic device, comprising: a memory storing an executable program; and a processor for running the program, wherein any one of the above flame graph generation methods is executed when the program is running.

[0023] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is provided, wherein the storage medium stores a program, wherein when the program is running, the device where the storage medium is located is controlled to execute any of the above-mentioned flame graph generation methods.

[0024] According to another aspect of an embodiment of the present invention, a computer program product is provided, including a computer program or an instruction, and when the computer program or the instruction is executed by a processor, the method for generating a flame graph of any one of the above items is implemented.

[0025] In an embodiment of the present application, the following steps are adopted: when the target thread is running, the time information of the target thread when it is executed on the processor is sampled to obtain the first time proportion information corresponding to the target thread; the time information of the target thread being called out of the processor and called into the processor is sampled to obtain the second time proportion information corresponding to the target thread; based on the first time proportion information and the second time proportion information, a target flame graph is generated, wherein the target flame graph is used to analyze the performance delay information of the target thread, thereby solving the technical problem in the related art that the information contained in the flame graph is single, resulting in relatively low accuracy of the flame graph.

[0026] In this solution, when the target thread is running, the time of the target thread executing on the processor is sampled, and the waiting time of the target thread being scheduled out of the processor and being scheduled back into the processor is sampled to obtain first time proportion information and second time proportion information. Finally, a target flame graph is generated based on the two time proportion information. The target flame graph obtained by integrating the first time proportion information and the second time proportion information can intuitively reflect the performance delay characteristics of the target thread in the execution and waiting states, overcome the problem of the singleness of traditional flame graph information, and can more accurately identify which function calls or system states are the cause of the delay, thereby achieving the effect of improving the accuracy of the flame graph. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0028] Figure 1 is a hardware structure block diagram of a computer terminal provided according to Embodiment 1 of the present application;

[0029] Figure 2is a flowchart of a method for generating a flame graph according to Embodiment 1 of the present application;

[0030] Figure 3 This is a schematic diagram of the flame diagram provided in Example 1 of the present application. Figure 1 ;

[0031] Figure 4 This is a schematic diagram of the flame diagram provided in Example 1 of the present application. Figure 2 ;

[0032] Figure 5 This is a schematic diagram of the flame diagram provided in Example 1 of the present application. Figure 3 ;

[0033] Figure 6 This is a schematic diagram of the flame diagram provided in Example 1 of the present application. Figure 4 ;

[0034] Figure 7 is a flowchart of a method for generating a flame graph according to Embodiment 2 of the present application;

[0035] Figure 8 is a schematic diagram of a flame graph generation device provided according to Embodiment 3 of the present application;

[0036] Fig. 9 This is a structural block diagram of an electronic device provided according to Embodiment 6 of the present application. DETAILED DESCRIPTION

[0037] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0038] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0039] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following explanations:

[0040] Thread: A thread is the basic unit of operating system scheduling and represents an independent execution path. Each thread has its own execution context (such as register status, stack space, etc.), and the operating system kernel allocates CPU time slices according to the scheduling policy.

[0041] Flame Graph: Flame Graph is a visualization tool based on stack tracing, which is used to show the call stack distribution when the program is running. Its feature is that the call stack is stacked and displayed according to the time ratio, and the width represents the time consumption ratio of the function or code path. Flame Graph is widely used in the field of performance analysis to help developers quickly locate hot functions or bottleneck codes.

[0042] Oncpu: refers to the time period during which a thread is actually executed on the CPU, that is, the state in which the thread occupies CPU resources for calculation or processing tasks. The length of oncpu time usually reflects the degree of demand for CPU resources by the thread. Excessive oncpu time may indicate the presence of computationally intensive tasks or CPU performance bottlenecks.

[0043] Offcpu: refers to the period of time when the thread is not executing on the CPU, usually because the thread is in a waiting state (such as I / O blocking, lock contention, sleep, etc.). The distribution of offcpu time can help identify non-CPU related performance issues in the system, such as disk latency, network latency, or resource contention.

[0044] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards in the relevant regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0045] Example 1

[0046] According to an embodiment of the present application, a method for generating a flame graph is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0047] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1FIG. 1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for generating a flame graph. Figure 1 As shown, the computer terminal (or mobile device) 10 may include a processor set 102 (the processor set 102 may include but is not limited to a processing device such as a microprocessor MCU (Microcontroller Unit) or a programmable logic device FPGA (Field-Programmable Gate Array), and the processor set 102 may include a processor set, Figure 1 102a, 102b, ..., 102n are used to illustrate), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.

[0048] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the computer terminal 10 (or mobile device). As involved in the embodiments of the present application, the data processing circuit acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0049] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage devices corresponding to the method for generating a flame graph in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, the method for generating the flame graph described above is realized. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0050] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (Radio Frequency, RF) module, which is used to communicate with the Internet wirelessly.

[0051] The display may be, for example, a touch screen liquid crystal display, which may enable a user to interact with a user interface of the computer terminal 10 (or mobile device).

[0052] Under the above operating environment, this application provides Figure 2 The flame graph shown is generated. Figure 2 : is a flow chart of a method for generating a flame graph according to Embodiment 1 of the present application. The method comprises:

[0053] Step S201, when a target thread is running, time information of the target thread when it is executed on a processor is sampled to obtain first time proportion information corresponding to the target thread.

[0054] Optionally, determine the target thread to be analyzed. The thread will call multiple functions during its life cycle, and these function calls will form a path called a call stack. When the target thread is running, sample the time information of the target thread when it is executed on the processor. The processor can be a central processing unit. It should be noted that the performance analysis tool can be used to sample the events of the hardware performance monitoring unit (PMU) to obtain the first time proportion information corresponding to the target thread. The first time proportion information is used to characterize the proportion of the time the target thread is executed on the central processing unit (CPU) relative to the total sampling time.

[0055] For example, under the set time period or event trigger conditions, whenever a PMU event occurs, the currently executing function will be sampled, and the call stack information, event type, timestamp and other related data will be recorded. During the entire sampling process, a large number of sampling points will be accumulated, and the sampling points contain the specific function being executed by the thread at that time and its context information. After the sampling is completed or the predetermined amount of data is reached, the sampled data is counted. The key indicator of the statistics is the number of sampling points of the function (or call stack), which directly reflects the frequency of function execution. The first time proportion information is obtained by calculating the total number of sampling points and the number of sampling points corresponding to the function.

[0056] Before sampling with a performance analysis tool, you need to configure the tool, specify the target thread to be sampled (by thread ID or process ID), and the sampling frequency or conditions. For example, you can set the sampling frequency of the performance analysis tool to once every millisecond.

[0057] Step S202: sampling the time information of the target thread being called out of the processor and called into the processor to obtain second time proportion information corresponding to the target thread.

[0058] Optionally, the time points when the target thread is scheduled out of the CPU and rescheduled back to the CPU can be sampled through a performance analysis tool to obtain the above-mentioned second time share information. The time difference between the time points when the target thread is scheduled out of the CPU and rescheduled back to the CPU can be used as the blocking time of the function executed when the target thread is scheduled out of the CPU. Therefore, the above-mentioned second time share information can be obtained based on the time points when the target thread is scheduled out of the CPU and rescheduled back to the CPU. The second time share information indicates the time distribution of the target thread in the offCPU (i.e., not executing on the processor) state, especially the blocking time percentage between when the thread is scheduled out of the CPU and when it is scheduled back to the CPU again due to waiting for resources, lock contention, I / O operations or other blocking events.

[0059] For example, the time points when the target thread is scheduled out of the CPU and rescheduled back to the CPU can be sampled by tracking the kernel function finish_task_switch or schedule. finish_task_switch or schedule is part of the operating system scheduling mechanism. Whenever a thread is transferred out of the CPU or a new thread is transferred into the CPU for execution, finish_task_switch or schedule will be called. With the help of performance analysis tools, the timestamps of these events and the status information at the time can be accurately recorded without affecting the operation of the system, including which threads are scheduled out and which are scheduled in, as well as their PID (process ID) and call stack information.

[0060] For example, use performance analysis tools to track the scheduling events of the target thread, such as tracking functions such as finish_task_switch or schedule. The call of these functions marks the change of the thread state from running to waiting or vice versa. Whenever the target thread is scheduled out of the CPU, record this time point (that is, the time when the blocking starts); similarly, when the thread is rescheduled back to the CPU, also record this time point (that is, the time when the blocking ends). This usually includes the thread ID, timestamp and current call stack information. By calculating the time difference between the above two time points, we can get the time when the target thread is in the off-CPU state, that is, the time between being scheduled out of the CPU and being scheduled back to the CPU again. This duration is the blocking time. According to the blocking time of the function and the total offCPU time in the entire sampling period, calculate the second time proportion information of each function or call stack.

[0061] For example, whenever the finish_task_switch event occurs, sampling is performed to determine the function that is scheduled out. During the entire sampling process, a large number of sampling points are accumulated, and the sampling points contain the scheduled function and its context information. After the sampling is completed (for example, the target thread is scheduled back to the CPU) or the predetermined amount of data is reached, the sampled data is counted. The key indicator of the statistics is the number of sampling points of the function (or call stack), which directly reflects the frequency of the function being scheduled out. The second time proportion information is obtained by calculating the total number of sampling points and the number of sampling points corresponding to the function.

[0062] Step S203: Generate a target flame graph according to the first time proportion information and the second time proportion information, wherein the target flame graph is used to analyze the performance delay information of the target thread.

[0063] Optionally, a target flame graph is generated based on the first time proportion information and the second time proportion information. For example, the oncpu and offcpu times of the functions involved in the target thread are calculated based on the first time proportion information and the second time proportion information, and then the above-mentioned target flame graph is generated based on the calculation results. It should be noted that the generated target flame graph distinguishes oncpu and offcpu times by different colors or styles, so that users can see at a glance which parts are delays caused by the execution of threads on the CPU and which parts are delays caused by the threads waiting or being scheduled out of the CPU. This intuitive visualization method greatly simplifies the understanding and analysis of thread performance delay information, helping users quickly locate bottlenecks that may cause slow system response.

[0064] In summary, when the target thread is running, the time of the target thread executing on the processor is sampled, and the waiting time of the target thread being scheduled out of the processor and scheduled into the processor again is sampled to obtain the first time proportion information and the second time proportion information. Finally, the target flame graph is generated according to the two time proportion information. The target flame graph obtained by integrating the first time proportion information and the second time proportion information can intuitively reflect the performance delay characteristics of the target thread in the execution and waiting states, overcome the problem of the singleness of the traditional flame graph information, and can more accurately identify which function calls or system states are the cause of the delay, thereby achieving the effect of improving the accuracy of the flame graph.

[0065] In order to improve the accuracy of obtaining the first time proportion information, in the flame graph generation method provided in the first embodiment of the present application, the time information of the target thread when executing on the processor is sampled to obtain the first time proportion information corresponding to the target thread, including: sampling and processing the clock cycle events in the processor to obtain multiple first sampling point information; when sampling the clock cycle events, obtaining the first call stack information corresponding to the target thread at the sampling time; based on the multiple first sampling point information and the first call stack information, obtaining the first time proportion information.

[0066] Optionally, the clock cycle events in the processor can be sampled and processed by a performance analysis tool to obtain the above-mentioned multiple first sampling point information. The first sampling point information can be a thread ID and a corresponding sampling point sequence number. The clock cycle event can be a cycle event (a cycle event is an event type supported by a hardware performance monitoring unit, which is used to detect the frequency or number of clock cycles of the processor executing instructions). The sampling of the cycle event can be regarded as uniform sampling during the CPU execution process. Through a large number of sampling samples, the CPU execution event distribution of the thread on different functions can be reflected.

[0067] When the performance analysis tool samples the clock cycle event, it also obtains the call stack information of the target thread at the sampling time, which is the first call stack information mentioned above. The call stack is a list of a series of functions or code snippets being executed. It can tell what the thread is doing at a certain moment and how it reached the current execution location. The first call stack information is crucial for analyzing the active path and execution efficiency of the thread.

[0068] Finally, the first time proportion information is obtained by analyzing the collected multiple first sampling point information and the first call stack information. For example, the sampling points of the same first call stack may be aggregated, and then the first time proportion information is obtained according to the ratio between the number of sampling points corresponding to the first call stack and the total number of sampling points. For example, the sampling points of the same call stack path may be aggregated, the total execution time of the function in the target thread on the CPU may be calculated, and then these times may be compared with the total CPU execution time of the thread, so as to obtain the execution time proportion in the target thread, that is, the first time proportion information mentioned above.

[0069] By sampling and acquiring call stack information, we can capture the actual execution status of threads on the CPU and quickly locate which ones consume the most CPU resources, thus improving the accuracy of performance analysis.

[0070] In order to improve the accuracy of obtaining the second time proportion information, in the flame graph generation method provided in the first embodiment of the present application, the time information of the target thread being called out of the processor and called into the processor is sampled to obtain the second time proportion information corresponding to the target thread, including: sampling and processing the process switching event in the processor to obtain multiple second sampling point information; when sampling the process switching event, obtaining the second call stack information corresponding to the target thread at the sampling time; based on the multiple second sampling point information and the second call stack information, obtaining the second time proportion information.

[0071] Optionally, the process switching event in the processor can be sampled by a performance analysis tool to obtain the above-mentioned multiple second sampling point information. The second sampling point information can be a thread ID and a corresponding sampling point sequence number. The process switching event can be finish_task_switch. Finish_task_switch is an important function in the thread scheduling process. The finish_task_switch event records the thread pid that is scheduled out of the cpu and the thread pid that is scheduled into the cpu, as well as the timestamp and call stack when the event occurs. Different from the cycle event, the finish_task_switch event is a full trace. Each time it is triggered, a trace record will be generated. In this way, the performance analysis tool will fully record the time when the thread is scheduled into and out of the cpu by the cpu. Therefore, multiple second sampling point information can be obtained by sampling finish_task_switch. For example, whenever the finish_task_switch event occurs, sampling will be performed to determine the function that is scheduled out.

[0072] At the same time, when sampling the process switching event, the second call stack information corresponding to the target thread at the sampling time is obtained, and finally, based on the above-sampled second sampling point information and the second call stack information, the time distribution of the target thread in the offcpu state is calculated to obtain the above-mentioned second time proportion information. For example, the sampling points of the same second call stack may be aggregated, and then the second time proportion information may be obtained based on the ratio between the number of sampling points corresponding to the second call stack and the total number of sampling points.

[0073] By fully tracking the timestamps and call stack information during the thread scheduling process, we can accurately calculate the time distribution of threads in the offcpu state, and then optimize resource management, reduce blocking waits, or adjust scheduling strategies to improve the continuity and efficiency of thread execution, thereby improving overall system performance.

[0074] In order to improve the accuracy of the target flame graph, in the flame graph generation method provided in the first embodiment of the present application, generating the target flame graph according to the first time proportion information and the second time proportion information includes: generating a first flame graph corresponding to when the target thread is executed on the processor according to the first time proportion information; generating a second flame graph corresponding to when the target thread is called out of the processor and called into the processor according to the second time proportion information; generating the target flame graph according to the first flame graph and the second flame graph.

[0075] Optionally, a first flame graph corresponding to the target thread when it is executed on the processor is generated according to the first time proportion information obtained from the CPU clock cycle event sampling. The first flame graph is to display the execution time distribution of the thread on the CPU, and the width represents the execution time proportion of the function, the height represents the call level, and the call stack is displayed from bottom to top, so as to clearly reveal the hot spots and bottlenecks of oncpu time.

[0076] Based on the second time proportion information obtained by sampling the thread's entry and exit events (such as finish_task_switch) during the CPU scheduling process, the second flame graph corresponding to the target thread being called out of the processor and called into the processor is generated. The second flame graph focuses on the time distribution of the thread not executing on the CPU, that is, the offcpu time. It analyzes the call stack information to show the time the thread spends waiting for resources, blocking or other non-computing activities, helping to identify the root cause of non-CPU related performance problems.

[0077] Finally, the information of the first flame graph and the second flame graph are fused to generate a target flame graph. For example, the offcpu time segment is matched with the oncpu call stack path to accumulate the offcpu time to the corresponding oncpu call stack to obtain the above target flame graph.

[0078] By generating the first flame graph and the second flame graph and integrating them into the target flame graph, a complete and accurate view of the thread performance delay distribution can be obtained, which helps to more effectively solve performance bottleneck problems in complex systems.

[0079] In order to improve the accuracy of the first flame graph, in the flame graph generation method provided in the first embodiment of the present application, generating the first flame graph corresponding to the target thread when it is executed on the processor according to the first time proportion information includes: determining the first total duration based on the start count and the end count of the clock cycle event in the processor; calculating according to the first time proportion information and the first total duration to obtain the first execution duration information of the function in the first call stack information corresponding to the first sampling point in the multiple first sampling point information; generating the first flame graph according to the first execution duration information and the first total duration.

[0080] Optionally, the first total duration is obtained by calculating the start count and the end count of all clock cycle events. For example, if one number is counted every 1s, the first total duration can be obtained according to the start count, the end count and the counting cycle. It should be noted that the counting starts when the target thread is executed on the processor (determining the above-mentioned start count), stops when the target thread is called out of the processor, and continues to count until the target thread is executed and the end count is obtained when it is called back. Then, the first execution duration information of the function in the first call stack information corresponding to the first sampling point is calculated according to the first time proportion information and the first total duration, and the first execution duration information is obtained by multiplying the first time proportion information and the first total duration. Then, the same first call stack is aggregated, and finally the first flame graph is generated according to the first execution duration information and the first total duration. The call stacks are stacked and displayed in the flame graph according to their execution duration, the width represents the time proportion, the height represents the call level, and they are arranged in the order of calls from bottom to top. Each layer of the first flame graph represents a function or call point, and its width intuitively reflects the proportion of the function in the total execution time of the target thread, allowing users to quickly identify functions that are computationally intensive or time-consuming.

[0081] The first flame graph accurately reflects the execution time distribution of the target thread on the CPU, which helps users more accurately identify which functions or code segments are computationally intensive and their degree of CPU resource consumption.

[0082] In order to improve the accuracy of the second flame graph, in the flame graph generation method provided in the first embodiment of the present application, generating the second flame graph corresponding to when the target thread is called out of the processor and called into the processor according to the second time proportion information includes: determining the second total duration based on the start time and the end time of the process switching event in the processor; calculating according to the second time proportion information and the second total duration to obtain the second execution duration information of the function in the second call stack information corresponding to the second sampling point in the multiple second sampling point information; generating the second flame graph according to the second execution duration information and the second total duration.

[0083] Optionally, based on sampling of events related to process scheduling in the processor (such as finish_task_switch, schedule), the precise timestamps of all these events are recorded. By analyzing the start time and end time of the event, the time interval when the target thread is called out and called back into the CPU can be calculated, that is, the second total duration.

[0084] According to the second time proportion information and the second total duration, the offcpu time corresponding to the call stack path in the second sampling point information is calculated, that is, the second execution duration information. The second execution duration information is obtained by multiplying the second time proportion information and the second total duration. By aggregating and accumulating the offcpu time segments of the same call stack path, it is possible to quantify the time spent by the function when the target thread is not executed on the CPU, as well as their proportion in the total offcpu time. Finally, based on the calculated second execution duration information and the second total duration, a second flame graph is generated. Unlike the first flame graph, the second flame graph focuses on showing the offcpu time distribution of the target thread. It also uses the form of a call stack and is stacked from bottom to top. The width of each layer represents the offcpu time proportion of a specific call stack path, while the height represents the depth of the call.

[0085] Through the second flame graph, you can clearly see which functions cause the thread to wait or block for a long time, improving the accuracy of subsequent performance analysis.

[0086] How to obtain the target flame graph is crucial. Therefore, in the flame graph generation method provided in Example 1 of the present application, generating the target flame graph based on the first flame graph and the second flame graph includes: determining the matching relationship between multiple first call stacks in the first flame graph and multiple second call stacks in the second flame graph, wherein the matching relationship includes: a first matching relationship and a second matching relationship, the first matching relationship indicates that the first call stack is the same as the second call stack, and the second matching relationship indicates that there are some function call relationships in the first call stack that are the same as the second call stack; based on the matching relationship, merging the first flame graph and the second flame graph to obtain the target flame graph.

[0087] Optionally, a matching relationship between multiple first call stacks in the first flame graph and multiple second call stacks in the second flame graph is determined. The call stack matching relationship is divided into two categories: the first matching relationship indicates that the first call stack is exactly the same as the second call stack, which means that the threads are in the same execution context during oncpu and offcpu time. The second matching relationship indicates that some function call relationships in the first call stack are the same as those in the second call stack, that is, the call stack paths partially overlap. After determining the matching relationship of the call stacks, the first flame graph and the second flame graph are merged according to the matching relationship to obtain the target flame graph.

[0088] In an optional embodiment, based on the matching relationship, the first flame graph and the second flame graph are merged to obtain the target flame graph, including: if there is a first target call stack in multiple first call stacks and a second target call stack in multiple second call stacks has a first matching relationship, then the execution time of the first function in the second target call stack is merged into the corresponding position of the first target call stack, and the execution time corresponding to the second function of the second target call stack is set at the top level of the first target call stack, the first function is a function other than the top level in the second target call stack, and the second function is the top level function in the second target call stack; if multiple If a matching relationship between a first target call stack and a second target call stack in multiple second call stacks is a second matching relationship in the first call stack, the execution time of the third function in the second target call stack is merged into the corresponding position of the first target call stack, and the execution time of the fourth function in the second target call stack is added to the first flame graph according to the position of the fourth function in the second target call stack, the third function is a function in the second target call stack that has the same function call relationship as the function in the first target call stack, and the fourth function is a function in the second target call stack other than the third function; after the merging is completed, the merged first flame graph is proportionally adjusted to obtain the target flame graph.

[0089] Optionally, after determining the matching relationship of the call stack, merging the first flame graph and the second flame graph based on the matching includes: when the matching relationship between the first target call stack and the second target call stack is the first matching relationship, it means that the two sets of call stacks are completely consistent. In this case, the execution time of the first function (i.e., the function other than the top-level function) in the second target call stack will be merged into the corresponding position of the first target call stack, which is essentially superimposing the offcpu time information on the basis of the oncpu time distribution. The execution time of the second function (the top-level function of the second target call stack) is directly set at the top level of the first target call stack, and is displayed in the form of an independent "spire" to ensure a clear distinction between the oncpu and offcpu states.

[0090] For example, Figure 3As shown, the first target call stack in the first flame graph includes function A-function B-function D, and the second target call stack in the second flame graph includes function A-function B-function D. The matching relationship between the first target call stack and the second target call stack is the first matching relationship, and the schematic diagram after the matching relationship between the first target call stack and the second target call stack is merged is as shown in Figure 3 As shown, the execution time of functions other than the top-level function is merged into the corresponding position of the first target call stack, and the execution time of the top-level function of the second target call stack is directly set at the top of the first target call stack, and displayed in the form of an independent "spire".

[0091] If there is a second matching relationship between the first target call stack and the second target call stack, it means that the call stacks of the two are partially the same, but not completely matched. For the second matching relationship, the execution time of the third function in the second target call stack (i.e., the function with the same call relationship as the function in the first target call stack) is merged into the corresponding position of the first target call stack, and according to the position of the fourth function in the second target call stack, its execution time is added to the first flame graph at an appropriate level. That is, only the offcpu time is accumulated to the longest common prefix of the oncpu time, and the incompletely matched part (i.e., the fourth function mentioned above) is presented as an independent rectangle in the flame graph to distinguish the difference between offcpu and oncpu time. It should be noted that if the offcpu fragment and the oncpu information of a certain call stack are completely matched, the functions at the top of the tower are not merged, and different colors can be set to distinguish the offcpu and oncpu times.

[0092] For example, Figure 4 As shown, the first target call stack in the first flame graph includes function A-function B-function D, and the second target call stack in the second flame graph includes function A-function B-function C. The matching relationship between the first target call stack and the second target call stack is the second matching relationship. The schematic diagram after the matching relationship between the first target call stack and the second target call stack is merged is as shown in Figure 4 As shown, function A-function B are merged, and function D and function C are set at the top level.

[0093] For example, Figure 5 As shown, the first target call stack in the first flame graph includes function A-function B-function D-function E, and the second target call stack in the second flame graph includes function A-function B-function C. The matching relationship between the first target call stack and the second target call stack is the second matching relationship, then function A-function B are merged, function C is set at the third layer, and function E is still at the top layer.

[0094] Since the original first flame graph and the second flame graph respectively show the hot spot distribution of oncpu time and offcpu time, but after the merge process, the total execution time of the function has changed because it includes both oncpu time and offcpu time. If the proportion is not adjusted, the width in the flame graph will not correctly reflect the proportion of the function in the combined oncpu time and offcpu time, which will cause distortion of the performance analysis results. Therefore, after completing the merge process, it is necessary to recalculate the width of the function in the target flame graph based on the new execution time information and total time to reflect its total proportion in oncpu and offcpu time, that is, the above-mentioned proportion adjustment of the merged first flame graph is used to obtain the target flame graph.

[0095] In an optional embodiment, the sampled thread pid is specified, and then events that can reflect oncpu and offcpu information are sampled simultaneously according to the thread pid. For oncpu, cycle events are sampled. For offcpu, any technology that supports full tracking of thread scheduling is acceptable, including but not limited to tracking of functions related to scheduling such as finish_task_switch and schedule. The sampled information includes timestamp, thread pid scheduled out of cpu, thread pid scheduled into cpu, and call stack.

[0096] For oncpu, aggregate the call stack and calculate the oncpu hotspot ratio of the functions on the call stack. Find the earliest and latest time of all cycle events, calculate the time difference to get the total time. Then for the function, calculate the oncpu time based on the hotspot ratio in the previous step.

[0097] For offcpu, extract the offcpu fragment from the offcpu event, calculate the time of the offcpu fragment, and aggregate the same call stack (the offcpu fragment time of the same call stack is accumulated). For the aggregated offcpu fragment, match the call stack in the oncpu information. Perhaps the offcpu fragment call stack cannot completely match the oncpu call stack, but the longest prefix should be matched as much as possible. The unmatched suffix will be presented as an independent "spire", and the time of the offcpu fragment is added to the matching prefix. If the offcpu fragment and a call stack of the oncpu information completely match, the functions at the top of the tower are not merged, so that they can be distinguished in the flame graph. Recalculate the proportion of each function to the whole, draw the flame graph, and the functions at the top of the flame graph tower either belong to oncpu or offcpu, and the two are distinguished by different tones. Then the final effect diagram is as follows Figure 6 As shown, the orange color at the top level represents oncpu, and the blue color represents offcpu.

[0098] The target flame graph not only shows the functions or call points that consume the most time, but also reveals their specific performance in the oncpu and offcpu states. This provides a clear direction for performance optimization. Users can optimize call stacks that account for a large proportion of offcpu time, such as reducing resource waiting time or adjusting oncpu functions to improve computing efficiency.

[0099] In the flame graph generation method provided in the first embodiment of the present application, by sampling the time information of the target thread when it is executed on the processor when the target thread is running, first time proportion information corresponding to the target thread is obtained; the time information of the target thread being called out of the processor and called into the processor is sampled to obtain second time proportion information corresponding to the target thread; based on the first time proportion information and the second time proportion information, a target flame graph is generated, wherein the target flame graph is used to analyze the performance delay information of the target thread, thereby solving the technical problem in the related art that the information contained in the flame graph is single, resulting in relatively low accuracy of the flame graph.

[0100] In this solution, when the target thread is running, the time of the target thread executing on the processor is sampled, and the waiting time of the target thread being scheduled out of the processor and being scheduled back into the processor is sampled to obtain first time proportion information and second time proportion information. Finally, a target flame graph is generated based on the two time proportion information. The target flame graph obtained by integrating the first time proportion information and the second time proportion information can intuitively reflect the performance delay characteristics of the target thread in the execution and waiting states, overcome the problem of the singleness of traditional flame graph information, and can more accurately identify which function calls or system states are the cause of the delay, thereby achieving the effect of improving the accuracy of the flame graph.

[0101] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0102] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, disk, CD), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.

[0103] Example 2

[0104] According to an embodiment of the present application, a method for generating a flame graph is also provided, such as Figure 7 As shown, the method for generating the flame graph includes:

[0105] Step S701, receiving a flame graph generation request triggered by a client;

[0106] Step S702: In the cloud server, based on the flame graph generation request, when the target thread is running, the time information of the target thread when it is executed on the processor is sampled to obtain first time proportion information corresponding to the target thread; the time information of the target thread being called out of the processor and called into the processor is sampled to obtain second time proportion information corresponding to the target thread; based on the first time proportion information and the second time proportion information, a target flame graph is generated, wherein the target flame graph is used to analyze the performance delay information of the target thread;

[0107] Step S703: Return the target flame graph to the client.

[0108] It should be noted that the method for generating the flame graph in the cloud server is the same as that in the first embodiment, and will not be described in detail here.

[0109] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0110] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, disk, CD), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.

[0111] Example 3

[0112] According to an embodiment of the present application, a flame graph generation device for implementing the flame graph generation method is also provided, such as Figure 8 As shown, the device includes: a first sampling unit 801, a second sampling unit 802 and a generating unit 803.

[0113] The first sampling unit 801 is used to sample the time information of the target thread when it is executed on the processor when the target thread is running, so as to obtain the first time proportion information corresponding to the target thread;

[0114] The second sampling unit 802 is used to sample the time information of the target thread being called out of the processor and called into the processor to obtain the second time proportion information corresponding to the target thread;

[0115] The generating unit 803 is used to generate a target flame graph according to the first time proportion information and the second time proportion information, wherein the target flame graph is used to analyze the performance delay information of the target thread.

[0116] Optionally, in the flame graph generation device provided in Example 3 of the present application, the first sampling unit includes: a first sampling module, used to sample and process clock cycle events in the processor to obtain multiple first sampling point information; a first acquisition module, used to obtain the first call stack information corresponding to the target thread at the sampling time when sampling the clock cycle events; a first determination module, used to obtain the first time proportion information based on the multiple first sampling point information and the first call stack information.

[0117] Optionally, in the flame graph generation device provided in Example 3 of the present application, the second sampling unit includes: a second sampling module, used to sample and process the process switching event in the processor to obtain multiple second sampling point information; a second acquisition module, used to obtain the second call stack information corresponding to the target thread at the sampling time when sampling the process switching event; a second determination module, used to obtain the second time proportion information based on the multiple second sampling point information and the second call stack information.

[0118] Optionally, in the flame graph generation device provided in Example 3 of the present application, the generation unit includes: a first generation module, used to generate a first flame graph corresponding to when the target thread is executed on the processor based on the first time proportion information; a second generation module, used to generate a second flame graph corresponding to when the target thread is called out of the processor and called into the processor based on the second time proportion information; a third generation module, used to generate a target flame graph based on the first flame graph and the second flame graph.

[0119] Optionally, in the flame graph generation device provided in Example 3 of the present application, the first generation module includes: a first determination submodule, used to determine the first total duration based on the start count and the end count of the clock cycle event in the processor; a first calculation submodule, used to calculate according to the first time proportion information and the first total duration to obtain the first execution duration information of the function in the first call stack information corresponding to the first sampling point in the multiple first sampling point information; the first generation submodule, used to generate the first flame graph according to the first execution duration information and the first total duration.

[0120] Optionally, in the flame graph generation device provided in Example 3 of the present application, the second generation module includes: a second determination submodule, used to determine the second total duration based on the start time and the end time of the process switching event in the processor; a second calculation submodule, used to calculate according to the second time proportion information and the second total duration, to obtain the second execution duration information of the function in the second call stack information corresponding to the second sampling point in the multiple second sampling point information; the second generation submodule is used to generate a second flame graph based on the second execution duration information and the second total duration.

[0121] Optionally, in the flame graph generation device provided in Example 3 of the present application, the third generation module includes: a third determination submodule, used to determine the matching relationship between multiple first call stacks in the first flame graph and multiple second call stacks in the second flame graph, wherein the matching relationship includes: a first matching relationship and a second matching relationship, the first matching relationship indicating that the first call stack is the same as the second call stack, and the second matching relationship indicating that there are some function call relationships in the first call stack that are the same as the second call stack; a merging submodule, used to merge the first flame graph and the second flame graph based on the matching relationship to obtain a target flame graph.

[0122] Optionally, in the flame graph generation device provided in Example 3 of the present application, the merging submodule includes: a first merging submodule, which is used to merge the execution time of the first function in the second target call stack into the corresponding position of the first target call stack if there is a first target call stack in multiple first call stacks and a second target call stack in multiple second call stacks has a first matching relationship, and set the execution time corresponding to the second function of the second target call stack at the top level of the first target call stack, the first function is a function other than the top level in the second target call stack, and the second function is the top level function in the second target call stack; a second merging submodule, which is used to merge the execution time of the first function in the second target call stack into the corresponding position of the first target call stack if there is a first target call stack in multiple first call stacks and a second target call stack in multiple second call stacks has a first matching relationship; and If there is a first target call stack in the first call stack and a second target call stack in multiple second call stacks, and the matching relationship is the second matching relationship, then the execution time of the third function in the second target call stack is merged into the corresponding position of the first target call stack, and the execution time of the fourth function in the second target call stack is added to the first flame graph according to the position of the fourth function in the second target call stack, the third function is a function in the second target call stack with the same function call relationship as the function in the first target call stack, and the fourth function is a function in the second target call stack except the third function; the adjustment sub-module is used to proportionally adjust the merged first flame graph after the merging is completed to obtain the target flame graph.

[0123] It should be noted that the first sampling unit 801, the second sampling unit 802 and the generating unit 803 described above correspond to steps S201 to S203 in the first embodiment, and the three units and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in the first embodiment. It should be noted that the above modules, as part of the device, can be run in the computer terminal 10 provided in the first embodiment.

[0124] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0125] Example 4

[0126] The embodiment of the present application may provide an electronic device, which may be any electronic device in an electronic device terminal group. Optionally, in this embodiment, the electronic device may also be replaced by a terminal device such as a mobile terminal.

[0127] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.

[0128] In this embodiment, the electronic device can execute the following program codes in the method for generating a flame graph: when the target thread is running, sampling the time information of the target thread when it is executed on the processor to obtain first time proportion information corresponding to the target thread; sampling the time information of the target thread being called out of the processor and called into the processor to obtain second time proportion information corresponding to the target thread; generating a target flame graph based on the first time proportion information and the second time proportion information, wherein the target flame graph is used to analyze the performance delay information of the target thread.

[0129] The electronic device can execute the following steps of the flame graph generation method: sampling the time information of the target thread when it is executed on the processor to obtain the first time proportion information corresponding to the target thread, including: sampling and processing the clock cycle events in the processor to obtain multiple first sampling point information; when sampling the clock cycle events, obtaining the first call stack information corresponding to the target thread at the sampling time; based on the multiple first sampling point information and the first call stack information, obtaining the first time proportion information.

[0130] The electronic device can execute the following steps of the flame graph generation method: sampling the time information of the target thread being called out of the processor and called into the processor to obtain the second time proportion information corresponding to the target thread, including: sampling and processing the process switching event in the processor to obtain multiple second sampling point information; when sampling the process switching event, obtaining the second call stack information corresponding to the target thread at the sampling time; based on the multiple second sampling point information and the second call stack information, obtaining the second time proportion information.

[0131] The electronic device can execute the following steps of the flame graph generation method: generating a target flame graph based on the first time proportion information and the second time proportion information includes: generating a first flame graph corresponding to when the target thread is executed on the processor based on the first time proportion information; generating a second flame graph corresponding to when the target thread is called out of the processor and called into the processor based on the second time proportion information; generating a target flame graph based on the first flame graph and the second flame graph.

[0132] The electronic device can execute the following steps of the flame graph generation method: generating a first flame graph corresponding to the target thread when it is executed on the processor according to the first time proportion information, including: determining a first total duration based on a start count and an end count of a clock cycle event in the processor; calculating according to the first time proportion information and the first total duration to obtain first execution duration information of a function in first call stack information corresponding to a first sampling point in multiple first sampling point information; generating a first flame graph according to the first execution duration information and the first total duration.

[0133] The electronic device can execute the following steps of the flame graph generation method: generating a second flame graph corresponding to when the target thread is called out of the processor and called into the processor according to the second time proportion information, including: determining the second total duration based on the start time and the end time of the process switching event in the processor; calculating according to the second time proportion information and the second total duration to obtain the second execution duration information of the function in the second call stack information corresponding to the second sampling point in the multiple second sampling point information; generating the second flame graph according to the second execution duration information and the second total duration.

[0134] The electronic device can execute the program code of the following steps in the method for generating a flame graph: generating a target flame graph based on the first flame graph and the second flame graph includes: determining a matching relationship between multiple first call stacks in the first flame graph and multiple second call stacks in the second flame graph, wherein the matching relationship includes: a first matching relationship and a second matching relationship, the first matching relationship indicating that the first call stack is the same as the second call stack, and the second matching relationship indicating that some function call relationships in the first call stack are the same as the second call stack; based on the matching relationship, merging the first flame graph and the second flame graph to obtain the target flame graph.

[0135] The electronic device can execute the program code of the following steps in the method for generating a flame graph: based on the matching relationship, the first flame graph and the second flame graph are merged to obtain the target flame graph, including: if there is a first target call stack in the multiple first call stacks and the second target call stack in the multiple second call stacks has a first matching relationship, then the execution time of the first function in the second target call stack is merged into the corresponding position of the first target call stack, and the execution time corresponding to the second function of the second target call stack is set at the top layer of the first target call stack, the first function is a function other than the top layer in the second target call stack, and the second function is the top layer function in the second target call stack; if there is a first target call stack in the multiple first call stacks and the second target call stack in the multiple second call stacks has a second matching relationship, then the execution time of the third function in the second target call stack is merged into the corresponding position of the first target call stack, and the execution time of the fourth function is added to the first flame graph according to the position of the fourth function in the second target call stack, the third function is a function in the second target call stack with the same function call relationship as that in the first target call stack; after the merging is completed, the first flame graph obtained by merging is proportionally adjusted to obtain the target flame graph.

[0136] Optionally, FIG. 1 is a structural block diagram of an electronic device according to an embodiment of the present application. Fig. 9 As shown, the electronic device 90 may include: one or more ( Fig. 9Only one is shown in the figure) processor 902, memory 904. The electronic device 90 may also include a storage controller, through which the memory 904 is controlled and managed; the electronic device 90 may also include a peripheral interface, through which the radio frequency module, audio module and display screen are connected.

[0137] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the method and device for generating the flame graph in the embodiment of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned method for generating the flame graph. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the electronic device 90 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0138] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: when the target thread is running, the time information of the target thread when it is executed on the processor is sampled to obtain the first time proportion information corresponding to the target thread; the time information of the target thread being called out of the processor and called into the processor is sampled to obtain the second time proportion information corresponding to the target thread; based on the first time proportion information and the second time proportion information, a target flame graph is generated, wherein the target flame graph is used to analyze the performance delay information of the target thread.

[0139] Optionally, the processor may also execute the following program code: sampling the time information of the target thread when it is executed on the processor to obtain the first time proportion information corresponding to the target thread, including: sampling and processing the clock cycle events in the processor to obtain multiple first sampling point information; when sampling the clock cycle events, obtaining the first call stack information corresponding to the target thread at the sampling moment; and obtaining the first time proportion information based on the multiple first sampling point information and the first call stack information.

[0140] Optionally, the processor may also execute program code of the following steps: sampling the time information of the target thread being called out of the processor and called into the processor to obtain second time proportion information corresponding to the target thread, including: sampling and processing the process switching event in the processor to obtain multiple second sampling point information; when sampling the process switching event, obtaining the second call stack information corresponding to the target thread at the sampling time; and obtaining the second time proportion information based on the multiple second sampling point information and the second call stack information.

[0141] Optionally, the processor may also execute program code of the following steps: generating a target flame graph based on the first time proportion information and the second time proportion information, including: generating a first flame graph corresponding to when the target thread is executed on the processor based on the first time proportion information; generating a second flame graph corresponding to when the target thread is called out of the processor and called into the processor based on the second time proportion information; generating a target flame graph based on the first flame graph and the second flame graph.

[0142] Optionally, the processor may also execute program code of the following steps: generating, based on the first time proportion information, a first flame graph corresponding to when the target thread is executed on the processor, including: determining a first total duration based on a start count and an end count of a clock cycle event in the processor; calculating, based on the first time proportion information and the first total duration, to obtain first execution duration information of a function in first call stack information corresponding to a first sampling point in multiple first sampling point information; generating a first flame graph based on the first execution duration information and the first total duration.

[0143] Optionally, the processor may also execute program code of the following steps: generating, based on the second time proportion information, a second flame graph corresponding to when the target thread is called out of the processor and called into the processor, including: determining a second total duration based on the start time and the end time of the process switching event in the processor; calculating, based on the second time proportion information and the second total duration, to obtain second execution duration information of a function in the second call stack information corresponding to a second sampling point in multiple second sampling point information; generating a second flame graph based on the second execution duration information and the second total duration.

[0144] Optionally, the processor may also execute program code of the following steps: generating a target flame graph based on the first flame graph and the second flame graph includes: determining a matching relationship between multiple first call stacks in the first flame graph and multiple second call stacks in the second flame graph, wherein the matching relationship includes: a first matching relationship and a second matching relationship, the first matching relationship indicating that the first call stack is the same as the second call stack, and the second matching relationship indicating that some function call relationships in the first call stack are the same as the second call stack; based on the matching relationship, merging the first flame graph and the second flame graph to obtain the target flame graph.

[0145] Optionally, the processor may also execute the following program code: based on the matching relationship, merging the first flame graph and the second flame graph to obtain a target flame graph including: if there is a first target call stack in multiple first call stacks and a second target call stack in multiple second call stacks has a first matching relationship, then merging the execution time of the first function in the second target call stack into the corresponding position of the first target call stack, and setting the execution time corresponding to the second function of the second target call stack at the top level of the first target call stack, the first function is the function other than the top level in the second target call stack, and the second function is the top level in the second target call stack. layer function; if there is a first target call stack in the multiple first call stacks and the second target call stack in the multiple second call stacks has a second matching relationship, then the execution time of the third function in the second target call stack is merged into the corresponding position of the first target call stack, and the execution time of the fourth function is added to the first flame graph according to the position of the fourth function in the second target call stack, the third function is a function in the second target call stack that has the same function call relationship as the function in the first target call stack, and the fourth function is a function in the second target call stack except the third function; after the merging is completed, the merged first flame graph is proportionally adjusted to obtain the target flame graph.

[0146] It can be understood by those skilled in the art that Fig. 9 The structure shown is for illustration only, and the electronic device 90 may also be a terminal device such as a smart phone, a tablet computer, a PDA, a mobile Internet device (MID), a PAD, etc. Fig. 9 The structure of the electronic device is not limited. Fig. 9 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Fig. 9 Different configurations are shown.

[0147] A person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing the hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0148] Example 5

[0149] The embodiment of the present application further provides a computer program product. Optionally, in this embodiment, the computer program product can be used to store the program code executed by the method for generating a flame graph provided in the first embodiment.

[0150] Optionally, in this embodiment, the computer program product may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.

[0151] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0152] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0153] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0154] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0155] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0156] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk and other media that can store program code.

[0157] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for generating a flame graph, characterized in that: include: When the target thread is running, sampling the time information of the target thread when it is executed on the processor to obtain first time proportion information corresponding to the target thread; Sampling time information of the target thread being called out of the processor and called into the processor to obtain second time proportion information corresponding to the target thread; A target flame graph is generated according to the first time proportion information and the second time proportion information, wherein the target flame graph is used to analyze the performance delay information of the target thread.

2. The method according to claim 1, characterized in that Sampling the time information of the target thread when it is executed on the processor to obtain the first time proportion information corresponding to the target thread includes: Sampling the clock cycle events in the processor to obtain a plurality of first sampling point information; When sampling the clock cycle event, obtaining first call stack information corresponding to the target thread at the sampling time; The first time proportion information is obtained based on the plurality of first sampling point information and the first call stack information.

3. The method according to claim 2, characterized in that Sampling the time information of the target thread being called out of the processor and called into the processor to obtain the second time proportion information corresponding to the target thread includes: Sampling the process switching event in the processor to obtain a plurality of second sampling point information; When sampling the process switching event, obtaining second call stack information corresponding to the target thread at the sampling time; The second time proportion information is obtained based on the plurality of second sampling point information and the second call stack information.

4. The method according to claim 3, characterized in that Generating a target flame graph according to the first time proportion information and the second time proportion information includes: Generate a first flame graph corresponding to the target thread when it is executed on the processor according to the first time proportion information; generating, according to the second time proportion information, a second flame graph corresponding to when the target thread is called out of the processor and called into the processor; The target flame graph is generated according to the first flame graph and the second flame graph.

5. The method according to claim 4, characterized in that Generating, according to the first time proportion information, a first flame graph corresponding to when the target thread is executed on the processor includes: Determining a first total duration based on a start count and an end count of a clock cycle event in the processor; Calculating according to the first time proportion information and the first total duration to obtain first execution duration information of a function in the first call stack information corresponding to a first sampling point in the plurality of first sampling point information; The first flame graph is generated according to the first execution duration information and the first total duration.

6. The method according to claim 4, characterized in that Generating, according to the second time proportion information, a second flame graph corresponding to when the target thread is called out of the processor and called into the processor includes: Determining a second total duration based on a start time and an end time of a process switching event in the processor; Calculating according to the second time proportion information and the second total duration to obtain second execution duration information of a function in the second call stack information corresponding to a second sampling point in the plurality of second sampling point information; The second flame graph is generated according to the second execution duration information and the second total duration.

7. The method according to claim 4, characterized in that Generating the target flame graph according to the first flame graph and the second flame graph includes: Determine a matching relationship between a plurality of first call stacks in the first flame graph and a plurality of second call stacks in the second flame graph, wherein the matching relationship includes: a first matching relationship and a second matching relationship, the first matching relationship indicating that the first call stack is the same as the second call stack, and the second matching relationship indicating that some function call relationships in the first call stack are the same as those in the second call stack; Based on the matching relationship, the first flame graph and the second flame graph are merged to obtain the target flame graph.

8. The method according to claim 7, characterized in that Based on the matching relationship, the first flame graph and the second flame graph are merged to obtain the target flame graph, including: If there is a first target call stack in the multiple first call stacks and a second target call stack in the multiple second call stacks has a first matching relationship, then the execution time of the first function in the second target call stack is merged into the corresponding position of the first target call stack, and the execution time corresponding to the second function of the second target call stack is set at the top layer of the first target call stack, the first function is a function other than the top layer in the second target call stack, and the second function is a top layer function in the second target call stack; If there is a first target call stack in the multiple first call stacks and the matching relationship between the second target call stack in the multiple second call stacks is a second matching relationship, then the execution time of the third function in the second target call stack is merged into the corresponding position of the first target call stack, and the execution time of the fourth function in the second target call stack is added to the first flame graph according to the position of the fourth function in the second target call stack, the third function is a function in the second target call stack with the same function call relationship as the function in the first target call stack, and the fourth function is a function in the second target call stack other than the third function; After the merging is completed, the first flame graph obtained by the merging is proportionally adjusted to obtain the target flame graph.

9. A method for generating a flame graph, characterized in that: include: Receive flame graph generation request triggered by client; In the cloud server, based on the flame graph generation request, when the target thread is running, time information of the target thread when it is executed on the processor is sampled to obtain first time proportion information corresponding to the target thread; Sampling time information of when the target thread is called out of the processor and called into the processor to obtain second time proportion information corresponding to the target thread; generating a target flame graph according to the first time proportion information and the second time proportion information, wherein the target flame graph is used to analyze performance delay information of the target thread; The target flame graph is returned to the client.

10. A flame graph generation device, characterized in that: include: A first sampling unit is used to sample time information of the target thread when it is executed on the processor when the target thread is running, so as to obtain first time proportion information corresponding to the target thread; A second sampling unit is used to sample the time information of the target thread being called out of the processor and called into the processor to obtain second time proportion information corresponding to the target thread; A generating unit is used to generate a target flame graph according to the first time proportion information and the second time proportion information, wherein the target flame graph is used to analyze the performance delay information of the target thread.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the storage medium is located is controlled to execute the flame graph generation method according to any one of claims 1 to 9.

12. An electronic device, characterized in that: include: A memory storing an executable program; A processor, configured to run the program, wherein the program, when running, executes the method for generating a flame graph according to any one of claims 1 to 9.

13. A computer program product, characterized in that The method comprises a computer program or an instruction, which, when executed by a processor, implements the method for generating a flame graph according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Data processing method, device and equipment

    CN110928750A

  • Fault determination method and device and terminal equipment

    CN115629904A

  • Software performance bottleneck processing method, electronic device and storage medium

    CN118779194A

  • Performance analysis method, electronic equipment, storage medium and program product

    CN119292885A

  • Distributed service performance analysis method and system and storage medium

    CN119621511A

Cited By

  • Instruction scheduling circuit, chip, data processing method and electronic equipment

    CN120909655A