Program performance analysis method, device and storage medium

By cross-analyzing and analyzing the program source database of target programs in high-performance computing systems, key performance indicators are automatically collected and analyzed, and problems of difficulty in data integration and limited user independent analysis capabilities in the existing technology are solved, and multi-dimensional program performance analysis and more accurate performance tuning decisions are achieved.

CN119645814BActive Publication Date: 2025-05-06国家超级计算天津中心
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510181149.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-05-06
Estimated Expiration
2045-02-19

AI Technical Summary

Technical Problem

When performing program performance analysis in high-performance computing systems in the prior art, there are problems such as inconsistent data storage and format, independent tool interface and limited user independent analysis capabilities, resulting in inaccurate performance tuning decisions.

Method used

By obtaining the program source database of the target program, cross-analysis and analysis are performed for each sub-database, key performance indicators are automatically collected and analyzed, and the analysis results are presented in a user-friendly manner, providing multi-dimensional program performance analysis conclusions.

Benefits of technology

It realizes automated acquisition and analysis of program performance on high-performance clusters, provides multi-dimensional analysis conclusions, and helps users make more accurate performance tuning decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119645814B_ABST
    Figure CN119645814B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of high-performance computing, and discloses a program performance analysis method, device, and storage medium, the method comprising: obtaining a program source database corresponding to a target program; wherein the target program is applied to a high-performance cluster; for each sub-database in the program source database, determining the performance analysis sub-result corresponding to the sub-database through cross-analysis and / or analysis one by one; determining the overview analysis results and the program type of the target program according to the program source database; determining the program performance analysis results corresponding to the target program according to the program type, the overview analysis results, and the performance analysis sub-results corresponding to each sub-database. Through the technical solution of the present invention, the key performance indicators of the target program in each performance analysis dimension on the high-performance cluster are automatically collected and analyzed, and the effect of visual display is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of high performance computing, and in particular to a program performance analysis method, device and storage medium. Background Art

[0002] With the rapid development of high performance computing (HPC) technology, the corresponding software and hardware technologies are also constantly improving. In this context, the performance analysis technology of various programs on supercomputer systems has also undergone rapid iteration and update.

[0003] At present, performance analysis mainly has the following problems: data storage and format are not unified, which makes data integration difficult; each performance analysis tool has an independent interface and only displays specific information, which limits the user's independent data analysis capabilities; and when users compare program performance under multiple runs or different configurations, they face the cumbersome process of manual comparison.

[0004] In view of this, the present invention is proposed. Summary of the invention

[0005] In order to solve the above technical problems, the present invention provides a program performance analysis method, device and storage medium, which realize the automatic collection and analysis of key performance indicators of the target program in the computing, reading and writing, and network communication dimensions on a high-performance cluster, display the analysis results in a user-friendly manner, and provide users with multi-dimensional program performance analysis conclusions, helping users make more accurate performance tuning decisions on high-performance clusters.

[0006] An embodiment of the present invention provides a program performance analysis method, the method comprising:

[0007] Obtaining a program source database corresponding to a target program; wherein the target program is applied to a high-performance cluster;

[0008] For each sub-database in the program source database, determining the performance analysis sub-result corresponding to the sub-database through cross analysis and / or analysis one by one;

[0009] Determining an overview analysis result and a program type of the target program according to the program source database;

[0010] The program performance analysis result corresponding to the target program is determined according to the program type, the overview analysis result, and the performance analysis sub-results corresponding to each sub-database.

[0011] An embodiment of the present invention provides an electronic device, the electronic device comprising:

[0012] Processor and memory;

[0013] The processor is used to execute the steps of the program performance analysis method described in any embodiment by calling the program or instruction stored in the memory.

[0014] An embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a program or instructions, wherein the program or instructions enable a computer to execute the steps of the program performance analysis method described in any embodiment.

[0015] The embodiments of the present invention have the following technical effects:

[0016] By obtaining the program source database corresponding to the target program applied to the high-performance cluster, each sub-database in the program source database is analyzed separately, and the performance analysis sub-results corresponding to the sub-database are determined through cross-analysis and / or one-by-one analysis, so as to perform separate analysis in each performance analysis dimension. Then, according to the program source database, the overview analysis results and the program type of the target program are determined, and according to the program type, the overview analysis results and the performance analysis sub-results corresponding to each sub-database, the program performance analysis results corresponding to the target program are determined. This realizes the automatic collection and analysis of the key performance indicators of the target program in each performance analysis dimension on the high-performance cluster, displays the analysis results in a user-friendly manner, and provides users with multi-dimensional program performance analysis conclusions, helping users make more accurate performance tuning decisions on high-performance clusters. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0018] Figure 1 is a flow chart of a program performance analysis method provided by an embodiment of the present invention;

[0019] Figure 2 It is a structural schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the scope of protection of the present invention.

[0021] The program performance analysis method provided by the embodiment of the present invention is mainly applicable to the case where a target program applied on a high-performance cluster is subjected to multi-dimensional performance analysis and program performance analysis results are obtained. The program performance analysis method provided by the embodiment of the present invention can be executed by an electronic device. Example

[0022] Figure 1 is a flow chart of a program performance analysis method provided by an embodiment of the present invention. Figure 1 , the program performance analysis method specifically includes:

[0023] S110: Obtain a program source database corresponding to the target program.

[0024] The target program is applied to the high-performance cluster, and the target program is the application program to be analyzed for performance. The program source database is a database composed of source data collected when the target program is running on the high-performance cluster and required for analyzing the performance of various dimensions.

[0025] Specifically, when the target program is applied in a high-performance cluster, basic data in different performance dimensions during the execution of the target program, namely, source data, are collected, and these source data are combined according to different performance dimensions to form a program source database.

[0026] Based on the above example, the program source database corresponding to the target program can be obtained in the following ways:

[0027] For each analysis dimension, according to the data collection module corresponding to the analysis dimension and the target program, an initial database corresponding to the analysis dimension is obtained;

[0028] According to the data type conversion rules corresponding to the analysis dimension, the data type of the initial database is converted to obtain the program source database corresponding to the analysis dimension.

[0029] The analysis dimensions are the dimensions for performance analysis of the target program, which may include the computing dimension, the reading and writing dimension, and the network communication dimension. The data acquisition module is a module for collecting relevant data for different analysis dimensions. The initial database is a database composed of data collected by the data acquisition module. The data type conversion rules are the conversion rules for converting the data types in the initial database corresponding to different analysis dimensions into the data types used for subsequent comprehensive performance analysis.

[0030] Specifically, for each analysis dimension, the data collection module corresponding to the analysis dimension is used to collect data when the target program is executed in the high-performance cluster, and the collected data can be used as the initial database corresponding to the analysis dimension. Then, according to the data type conversion rule corresponding to the analysis dimension, the data type of the initial database is converted, and after the conversion, the program source database corresponding to the analysis dimension can be obtained. In this way, the program source database corresponding to each analysis dimension can be obtained.

[0031] Exemplarily, the data acquisition module is used to sample various performance indicators of the target program and generate a corresponding initial database to provide data support for subsequent performance data analysis. The data acquisition module can support users to complete data acquisition in the following two ways: first, provide an executable program and its operating environment configuration, which is used to automatically use the tools in the data acquisition module to perform performance sampling and generate a corresponding initial database; second, directly input the address of each initial database that the user has sampled, and first determine whether the address of the initial database is valid, and use the address of the valid initial database as input for subsequent performance data analysis.

[0032] For example, for the computing dimension, you can use a module such as HPCToolkit (high-performance computing tool) to load the collected initial database into a target format, such as a Graphframe in a pandasdataframe (data analysis dataframe) through the data interface provided by the Hatchet tool. For the read-write dimension, you can use the pyDarshan (a tool for analyzing read-write performance) tool interface to convert the initial database to the target format. For the network communication dimension, you can specifically split and convert the initial database collected by mpiP (network communication performance sampling tool) into the target format.

[0033] S120 , for each sub-database in the program source database, determine the performance analysis sub-result corresponding to the sub-database through cross analysis and / or analysis one by one.

[0034] Among them, the sub-database is the partial database corresponding to different analysis dimensions in the program source database. The performance analysis sub-result is the result obtained by analyzing different analysis dimensions. Cross-analysis is the analysis of multiple runs of the target program, such as running in different operating environments. One-by-one analysis is the analysis of a single run of the target program.

[0035] Specifically, for each sub-database in the program source database, cross analysis and / or one-by-one analysis may be performed according to analysis requirements, and the obtained analysis result is the performance analysis sub-result corresponding to the sub-database.

[0036] Based on the above example, if the sub-database includes a computing source database, the performance analysis sub-result corresponding to the sub-database can be determined through cross-analysis in the following manner:

[0037] In a case where the computing source database includes two first databases to be processed in the operating environment, data cleaning is performed on each first database to be processed, and each first database to be processed is split into an application function database and a system function database according to the first field;

[0038] According to the second field, two application function databases are merged to obtain a first target database, two system function databases are merged to obtain a second target database, and two first to-be-processed databases are merged to obtain a third target database;

[0039] For each target database, the usage time difference and the usage time percentage difference corresponding to the central processor are determined according to the usage time of the central processor in the target database, and the first target data is determined from the target database as the performance analysis sub-result corresponding to the calculation source database according to the usage time difference and the usage time percentage difference.

[0040] Among them, the first database to be processed is a partial database of the computing source database in two operating environments. The first field is a field used to split the application function database and the system function database. The application function database is a partial database related to the function in the target program, and the system function database is a partial database related to the function other than the target program. The second field is used to connect the fields of databases corresponding to different operating environments. The first target database is the connection result of two application function databases, the second target database is the connection result of two system function databases, and the third target database is the connection result of two first databases to be processed. The target database includes the first target database, the second target database and the third target database. The central processing unit usage time is the data corresponding to the CPU (Central Processing Unit, central processing unit) usage time field. The usage time difference is the absolute value of the difference between the usage times of the two central processing units corresponding to each data in the target database, which can be understood as the difference in the CPU usage time of the same function in different operating environments. The usage time percentage difference is the absolute value of the difference in the proportion of the usage time of the two central processing units corresponding to each data in the target database in the total time, which can be understood as the difference in the proportion of the CPU usage time of the same function in different operating environments. The first target data is a preset number of data in the target database sorted from large to small according to the usage time difference, and a preset number of data in the target database sorted from large to small according to the usage time percentage difference.

[0041] Specifically, the calculation source database may contain the first database to be processed of the target program under two operating environments. Data cleaning can be performed on each first database to be processed, for example, useless fields can be removed, and data with a CPU usage time of 0 can be removed. Furthermore, each first database to be processed can be split into an application function database and a system function database according to the content of the first field. According to the second field, the two application function databases are merged to obtain the first target database, the two system function databases are merged to obtain the second target database, and the two first databases to be processed are merged to obtain the third target database. For each target database, the following processing is performed respectively: the central processing unit usage time under the two operating environments in the target database is searched, and a difference operation is performed to obtain the usage time difference corresponding to the central processing unit. For each operating environment, the central processing unit usage time corresponding to each piece of data is divided by the sum of the usage time of each central processing unit under the operating environment, so as to obtain the usage time percentage of each piece of data under the operating environment. Furthermore, the percentage of the usage time of the central processing unit under the two operating environments is subjected to a difference operation to obtain the usage time percentage difference. Sort by usage time difference from large to small, select a preset number of data (such as the first 20) from the target database, and sort by usage time percentage difference from large to small, select a preset number of data from the target database, and use these selected data as the first target data, that is, calculate a part of the performance analysis sub-result corresponding to the source database in the cross analysis.

[0042] Exemplarily, all data with CPU usage time of 0 in the two first databases to be processed are filtered, and then unnecessary data fields are removed. This operation is to reduce the amount of data and improve the speed of data analysis. Then, according to whether the first field in the data, such as "file" and "function", is null, the two first databases to be processed are split into an application function database and a system function database respectively. This operation is to more finely compare and analyze the characteristics of the program. The two application function databases that are split are connected according to the second field, such as "name" and "type", and "time_diff" and "time_diff_on_percentage" are calculated respectively, where "time_diff" is the difference in CPU usage time of the same function, and "time_diff_on_percentage" is to first calculate the percentage of each function in the total CPU running time, and then calculate the difference. Finally, the absolute values ​​of "time_diff" and "time_diff_on_percentage" are counted respectively. The first preset number (such as 20) of data and the statistical results are displayed. Similarly, the two split system function databases are also joined, and the absolute values ​​of "time_diff" and "time_diff_on_percentage" are counted for the preset number (such as 20) of data and the statistical results are displayed. Finally, the two first unsplit databases are joined to count the absolute values ​​of "time_diff" and "time_diff_on_percentage" for the preset number (such as 20) of data and the statistical results are displayed.

[0043] In addition, the performance analysis sub-results corresponding to the sub-databases can be determined by analyzing them one by one in the following manner:

[0044] For a first database to be processed in each operating environment in a computing source database, determining first original information based on a first preset script;

[0045] Determine a calling route according to the first database to be processed;

[0046] Determine the target route according to the CPU time percentage corresponding to each call route, and determine the call information to be displayed according to the target route and the preset depth;

[0047] According to the first preset chart requirement, the first to-be-processed database is processed to determine a first target chart;

[0048] Performing data cleaning on the first database to be processed, and splitting the first database to be processed into an application function database and a system function database according to the first field;

[0049] The first to-be-processed database, the application function database and the system function database are used as a central processing unit time-consuming analysis database;

[0050] For each CPU time consumption analysis database, determine the corresponding CPU usage time and load imbalance, and determine the second target data according to the CPU usage time and load imbalance respectively;

[0051] The first original information, the call information to be displayed, the target route, the first target chart, and the second target data are determined as the performance analysis sub-result corresponding to the calculation source database.

[0052] Among them, the first preset script is a pre-built script for collecting the overall information of the first database to be processed (the first original information). The first original information is the macro-level information in the first database to be processed, such as the job number, the number of nodes and computing cores used, the start time, the end time and the total running time of the program. The call route is the function call route involved in the first database to be processed, and the target route is the call route with the largest percentage of CPU usage time. The preset depth is the display depth set according to the display requirements. The call information to be displayed is the target route within the preset depth. The first preset chart requirement is the chart requirement when analyzing the calculation dimension, which may involve data and chart types. The first target chart is a chart generated according to the first preset chart requirement. The load imbalance is the variance of the corresponding central processing unit usage time on different processes when the function is executed. The second target data is a preset number of data in the central processing unit time consumption analysis database after being sorted from large to small according to the usage time, and a preset number of data after being sorted from large to small according to the load imbalance.

[0053] Specifically, for the first database to be processed in each operating environment in the computing source database, the following processing can be traversed and performed: the first database to be processed is read using a first preset script to generate first original information. Further, from the call relationship in the first database to be processed, the involved call routes are determined, and the CPU time ratio corresponding to each call route is solved, and the call route with the longest CPU time ratio is used as the target route. The part of the target route within the preset depth range is used as the call information to be displayed. According to the first preset chart requirement, the first database to be processed is processed to generate a first target chart. The first database to be processed is cleaned as in the above example, and the first database to be processed is split into an application function database and a system function database according to the first field. The first database to be processed, the application function database and the system function database are used as CPU time consumption analysis databases. For each CPU time consumption analysis database, the CPU usage time corresponding to each piece of data and the load imbalance of the CPU usage time are calculated respectively. The CPU usage time is sorted from large to small, and a preset number of data are selected. The load imbalance is sorted from large to small, and a preset number of data are selected, and the selected data are determined as target data. Furthermore, the first original information, the call information to be displayed, the target route, the first target chart, and the second target data are determined as a part of the performance analysis sub-result corresponding to the calculation source database in the analysis one by one.

[0054] Exemplarily, first, based on the first preset script, the first original information of the first database to be processed is generated, including the job number, the number of nodes and computing cores used, the start time, end time and total running time of the program, etc. Then, the hot function call graph of the program is generated, specifically, all the call routes inside the first database to be processed are traversed, and the call route (target route) with the longest CPU time percentage is selected for display. Due to performance reasons, only the call information within the preset depth (such as 20) from the program entry can be displayed, that is, the call information to be displayed. The first preset chart requirement can be a complete internal call graph, for example: when the cluster environment supports it and the amount of data in the database is relatively small (for example, the call information is less than 2000 lines), the corresponding instructions are used to generate the complete internal call graph of the first database to be processed, and it is displayed. The first preset chart requirement can be flame graph information, for example: when the cluster environment supports it, the flame graph information of the target program is generated using the response instruction, and it is displayed. According to the above example, the application function database and the system function database are determined. The CPU usage time and the load imbalance of the CPU usage time of the application function database, the system function database and the first database to be processed are analyzed, and a preset number (such as 20) of information (second target data) with the largest values ​​are respectively taken for analysis and display.

[0055] If the user does not enable the cross-analysis function, or the number of the first databases to be processed is less than 2, the cross-analysis function will be skipped and the logic of the analysis part will be directly performed. After completing the cross-analysis and / or the analysis part, the analysis of the CPU time of the target program will be summarized based on all the data information obtained from the analysis, and the analysis conclusion will be provided and displayed in the form of text description (the performance analysis sub-result corresponding to the calculation source database).

[0056] Based on the above example, if the sub-database includes a read-write source database, the performance analysis sub-result corresponding to the sub-database can be determined through cross-analysis in the following manner:

[0057] In the case where the read-write source database includes the second to-be-processed databases in two operating environments, for each read-write type, determining the number of first operations corresponding to the read-write type and the number of second operations corresponding to each preset file size interval from each second to-be-processed database;

[0058] An operation quantity difference is determined according to the first operation quantity and each second operation quantity corresponding to each operating environment, and each operation quantity difference is determined as a performance analysis sub-result corresponding to the read-write source database.

[0059] Among them, the second database to be processed is a partial database of the read-write source database in two operating environments. The read-write types include general storage read-write, network storage read-write, and program internal storage read-write. The first operation quantity is the number of operations corresponding to each read-write type. The preset file size interval is a file size range of multiple read-write operations set in advance. The second operation quantity is the number of operations of each read-write type in each preset file size interval. The operation quantity difference is the absolute value of the difference between the first operation quantity and the second operation quantity that have a corresponding relationship.

[0060] Specifically, the read-write source database may contain different analysis dimension databases of the second program source database to be processed under two operating environments of the target program. For each read-write type, the following processing can be performed respectively: from each second program source database to be processed, the total number of operations corresponding to the read-write type is counted as the first number of operations, and according to the file size corresponding to each operation, the second number of operations corresponding to each preset file size interval is counted. The first number of operations corresponding to the two operating environments is calculated by difference operation, and the corresponding second number of operations is calculated by difference operation, and each difference operation result is determined as the difference in the number of operations, and each difference in the number of operations is determined as a part of the performance analysis sub-result corresponding to the read-write source database in the cross analysis.

[0061] Exemplarily, if both of the two second databases to be processed include a POSIX (general storage read / write) module, the number of POSIX-related operations (first operation number) and the file size information of the operations in the two second databases to be processed are counted respectively, and the difference in the number of POSIX-related operations between the two second databases to be processed is calculated, and the number of file size information of the operations in each preset file size interval (second operation number) is counted, and the difference is counted, and the comparison result (performance analysis sub-result) is displayed. If both of the second databases to be processed include an MPI-IO (network storage read / write) module, the number of MPI-IO-related operations (first operation number) and the file size information of the operations in the two second databases to be processed are counted respectively, and the difference in the number of MPI-IO-related operations between the two second databases to be processed is calculated, and the number of file size information of the operations in each preset file size interval (second operation number) is counted, and the difference is counted, and the comparison result (performance analysis sub-result) is displayed. If both second databases to be processed contain STDIO (program internal storage read and write) modules, the number of STDIO related operations (first operation number) and the file size information of the operations in the two second databases to be processed are counted respectively, and the difference in the number of STDIO related operations between the two second databases to be processed is calculated, the number of file size information of the operations within each preset file size range is counted (second operation number), and the difference is counted, and the comparison results (performance analysis sub-results) are displayed.

[0062] In addition, the performance analysis sub-results corresponding to the sub-databases can be determined by analyzing them one by one in the following manner:

[0063] For each second database to be processed, determining second original information based on a second preset script;

[0064] For each read / write type, according to the read data and write data corresponding to the read / write type in the second to-be-processed database and in accordance with the second preset chart requirement, determine a second target chart corresponding to the read / write type;

[0065] Determine a read timing diagram according to the read data in the second database to be processed, and determine a write timing diagram according to the write data in the second database to be processed;

[0066] The second original information, the second target chart, the read timing chart, and the write timing chart are determined as performance analysis sub-results corresponding to the read and write source database.

[0067] Among them, the second preset script is a pre-built script for collecting the overall information of the second database to be processed (the second original information). The second original information is the macro-level information in the second database to be processed, such as the running job number, the user number submitting the job, the instructions for running the job, the number of nodes and computing cores used, the start time, end time and total running time of the job, etc. Read data and write data are the reading-related part and the writing-related part in the second database to be processed. The second preset chart requirement is the chart requirement when analyzing the read and write dimensions, which may involve data and chart types. The second target chart is a chart generated according to the second preset chart requirement. The read timing diagram is the drawing result of the read data in the timing, and the write timing diagram is the drawing result of the write data in the timing.

[0068] Specifically, for each second database to be processed, the following processing can be traversed and performed: Use the second preset script to read the second database to be processed and generate second original information. For each read-write type, distinguish the read data and write data corresponding to the read-write type in the second database to be processed. According to the second preset chart requirement, the second database to be processed, the read data and / or the write data are processed to generate a second target chart corresponding to the read-write type. Draw a read timing diagram through the read data in the second database to be processed, and draw a write timing diagram through the write data in the second database to be processed. The second original information, the second target chart, the read timing diagram, and the write timing diagram are determined as part of the performance analysis sub-results corresponding to the read-write source database in the one-by-one analysis.

[0069] Exemplarily, first, based on the second preset script, the second original information of the second database to be processed is generated, including the running job number, the user number submitting the job, the instruction for running the job, the number of nodes and computing cores used, the start time of the job, the end time and the total running time, etc. Then, the statistics of the number of operations of different read and write types of the second database to be processed are generated, including general storage read and write, network storage read and write, and program internal storage read and write, etc. If the second database to be processed contains general storage read and write operations, detailed statistics can be made to generate corresponding operand statistics tables and bar charts (second target charts), such as the distribution of operands of read operations in general storage read and write, the distribution of operands of write operations in general storage read and write, etc. At the same time, a table and a bar chart (second target chart) of the file size statistics operated by general storage read and write can also be generated, such as counting read data and write data according to each preset file size interval. For other read and write types such as network storage read and write and program internal storage read and write, a similar method as the above-mentioned general storage read and write can be performed, which will not be repeated here. It is also possible to generate a read timing diagram and a write timing diagram for the second database to be processed, with the program's running time as the vertical axis and the subscript of the core in which the program runs as the vertical axis, to display the corresponding read and write operations of each computing core in each time period, and to use different line types to distinguish different read and write types. For example, general storage read and write operations are represented by red line segments, network storage read and write operations are represented by yellow line segments, and program internal storage read and write operations are represented by blue line segments, etc.

[0070] If the user does not enable the cross-analysis function, or the number of the second databases to be processed is less than 2, the cross-analysis function will be skipped and the logic of the analysis part will be directly performed. After completing the cross-analysis and / or the analysis part, the analysis of the read and write operations of the target program will be summarized based on all the data information obtained from the analysis, and the analysis conclusion will be provided and displayed in the form of text description (the performance analysis sub-results corresponding to the read and write source database).

[0071] Based on the above example, if the sub-database includes a network communication source database, the performance analysis sub-result corresponding to the sub-database can be determined through cross-analysis in the following manner:

[0072] In the case where the network communication source database includes two third databases to be processed in the operating environment, determining a calculation index corresponding to the data synchronization function for each third database to be processed;

[0073] For each synchronization type, determine the number of operations and the average time consumption corresponding to the synchronization type from the third to-be-processed database;

[0074] According to the calculation index in each operating environment, the number of operations corresponding to each synchronization type in each operating environment, and the average time consumption, the data synchronization index difference is determined, and the data synchronization index difference is used as the performance analysis sub-result corresponding to the network communication source database.

[0075] Among them, the third database to be processed is a partial database of the network communication source database in two operating environments. The data synchronization function is a function for communication and synchronization between multiple processes. The calculation indicators may include the number of computing cores, the average running time of the target program, the average time of the target program MPI (Message Passing Interface, data synchronization) function, the average time of the target program excluding the MPI function, the percentage of the total running time of the MPI function in the target program, and the variance of the MPI function time of the target program on each core. The synchronization type is a variety of data synchronization operation types. The number of operations is the total number of operations of a synchronization type. The average time is the result of dividing the total operation time of a synchronization type by the number of operations.

[0076] Specifically, the network communication source database may contain a third database to be processed under two operating environments for the target program. For each third database to be processed, the calculation indicators corresponding to various data synchronization functions can be statistically determined. Moreover, in each third database to be processed, for each synchronization type, the number of operations and the average time consumption corresponding to the synchronization type can be statistically calculated from the third database to be processed. The difference between the calculation indicators in the two operating environments and the number of operations and the average time consumption corresponding to each synchronization type in each operating environment can be calculated respectively to obtain the data synchronization indicator difference. Furthermore, the data synchronization indicator difference can be used as a part of the performance analysis sub-result corresponding to the network communication source database in the cross analysis.

[0077] Exemplarily, the calculation indicators corresponding to each data synchronization function in the two third databases to be processed are counted separately, such as the number of computing cores used, the average running time of the target program, the average time of the MPI function of the target program, the average time excluding the MPI function in the target program, the percentage of the total running time of the MPI function in the target program, and the variance of the MPI time of the target program on each core, etc. And, for the two third databases to be processed, the difference of these calculation indicators is calculated. The number of operations of various synchronization types and the corresponding average time consumption in the two third databases to be processed are counted separately, and the difference in the number of operations of each synchronization type and the corresponding average time consumption in the two third databases to be processed is calculated. These differences are all part of the performance analysis sub-results corresponding to the network communication source database.

[0078] In addition, the performance analysis sub-results corresponding to the sub-databases can be determined by analyzing them one by one in the following manner:

[0079] For each of the third databases to be processed in the operating environment in the network communication source database, determining third original information based on a third preset script;

[0080] According to each synchronization type and the third to-be-processed database, and in accordance with the third preset chart requirement, a third target chart is determined;

[0081] Determine the first function data according to the number of calls corresponding to the data synchronization functions of each synchronization type in the third database to be processed, determine the second function data according to the function time corresponding to each data synchronization function in the third database to be processed, and determine the third function data according to the amount of function data corresponding to each data synchronization function in the third database to be processed;

[0082] A performance analysis sub-result corresponding to the network communication source database is determined according to the third original information, the third target chart, the first function data, the second function data and the third function data.

[0083] Among them, the third preset script is a pre-built script for collecting the overall information of the third database to be processed (the third original information). The third original information is the macro-level information in the third database to be processed, such as the instructions for running the program, the tool version used, the start time, end time and running time of the program, the type of timer used, the number of samplers and the process number of the sampler, etc. The third preset chart requirement is the chart requirement for analyzing the network communication dimension, which may involve data and chart types. The third target chart is a chart generated according to the third preset chart requirement. The first function data is the data of the preset number of types after the call times corresponding to the data synchronization functions of each synchronization type in the third database to be processed are sorted from large to small. The second function data is the data of the preset number of times after the function time corresponding to each data synchronization function in the third database to be processed is sorted from large to small. The third function data is the data of the preset number of times after the function data volume corresponding to each data synchronization function in the third database to be processed is sorted from large to small.

[0084] Specifically, for the third database to be processed under each operating environment in the network communication source database, the following processing can be traversed and performed: Use the third preset script to read the third database to be processed and generate the third original information. Split the third database to be processed according to each synchronization type, and select appropriate data in combination with the third preset chart requirements to generate a third target chart. Sort the number of calls corresponding to the data synchronization function of each synchronization type in the third database to be processed from large to small, select the data of the preset number of types, and determine it as the first function data. Sort the function time corresponding to each data synchronization function in the third database to be processed from large to small, select the data of the preset number of times, and determine it as the second function data. Sort the function data volume corresponding to each data synchronization function in the third database to be processed from large to small, select the data of the preset number of times, and determine it as the third function data. The third original information, the third target chart, the first function data, the second function data, and the third function data are used as part of the performance analysis sub-results corresponding to the network communication source database in the analysis one by one.

[0085] Exemplarily, first, based on the third preset script, the first original information of the third database to be processed is generated, including the instructions for running the program, the tool version used, the start time, end time and running time of the program, the type of timer used, the number of samplers and the process number of the sampler, etc. Then, a pie chart (third target chart) of the distribution of computing cores on computing nodes during the running of the target program can be generated. According to the time information of the data synchronization function, a pie chart (third target chart) of the proportion of the overall data synchronization function time in the total program time and a scatter plot of the time information of the data synchronization function on each computing core are generated (it can be understood that the higher the degree of discreteness, the more uneven the distribution of the data synchronization function and the greater the space for performance optimization). Furthermore, the number of calls of the data synchronization function in the third database to be processed can also be counted, and the number of calls can be sorted from large to small, and the statistical results of the data synchronization function with the preset number of types (such as 20) can be taken for display (first function data). The detailed information of the function calls of the preset number of times (such as 20 times) before the time consumption of a single data synchronization function in the third database to be processed can also be counted and displayed (second function data). It is also possible to count the function data volume of a single data synchronization function in the third database to be processed, sort it from large to small, and display the detailed information of the function calls of the first preset number of times (such as 20 times) (third function data). The data obtained from these analyses is part of the performance analysis sub-results corresponding to the network communication source database in the analysis one by one.

[0086] If the user does not enable the cross-analysis function, or the number of the third databases to be processed is less than 2, the cross-analysis function will be skipped and the logic of the analysis part will be directly performed. After completing the cross-analysis and / or analysis, the analysis of the target program running on the data synchronization function will be summarized based on all the data information obtained from the analysis, and the analysis conclusion will be provided and displayed in the form of text description (the performance analysis sub-result corresponding to the network communication source database).

[0087] S130: Determine the overview analysis result and the program type of the target program according to the program source database.

[0088] The overview analysis result is the result of analyzing the target program in all performance analysis dimensions, which can be understood as the key performance indicators during the operation of the target program. Program types include computing intensive, read-write intensive, and network communication intensive.

[0089] Specifically, the program source database is subjected to an overall performance analysis in each dimension, and the key performance indicators corresponding to each dimension are determined. These key performance indicators are used as an overview analysis result to provide the user with a macroscopic analysis result. Furthermore, the program type of the target program is determined based on the time-consuming information in the program source database. Specifically, the total computing time is obtained by using the time information in the computing source database, the total reading and writing time is obtained by using the time information in the reading and writing source database, and the total network communication time is obtained by using the time information in the network communication source database. The total computing time, the total reading and writing time, and the total network communication time are compared. If the total computing time is the largest, the program type of the target program is determined to be computing intensive. If the total reading and writing time is the largest, the program type of the target program is determined to be reading and writing intensive. If the total network communication time is the largest, the program type of the target program is determined to be network communication intensive.

[0090] Based on the above example, if the program source database includes a computing source database, a reading and writing source database, and a network communication source database, the overview analysis result can be determined according to the program source database in the following manner:

[0091] Based on the computational source database, determine the total number of instructions, the average number of instructions executed, the CPU utilization, and the function names with the highest time usage;

[0092] According to the read-write source database, determine the operation address corresponding to each piece of data, determine the overall read-write operation characteristics according to each operation address, and determine the read-write mode according to the overall read-write operation characteristics;

[0093] Determine the average bandwidth and peak bandwidth based on the read and write source database, and determine the throughput type based on the average bandwidth and peak bandwidth;

[0094] According to the network communication source database, determine the total synchronization communication time, the proportion of synchronization communication time, and the synchronization indicators corresponding to each synchronization type;

[0095] The total number of instructions, average number of executed instructions, CPU utilization, function name with the highest time consumption, read / write mode, throughput type, peak bandwidth, total synchronous communication time consumption, synchronous communication time consumption ratio, and synchronization indicators corresponding to each synchronization type are taken as the overview analysis results of the target program.

[0096] Among them, the total number of instructions is the total number of instructions executed during the execution of the target program. The average number of executed instructions is the ratio of the total number of instructions to the total execution time of the target program. The CPU utilization is the ratio of the time the CPU is in working state to the total execution time of the target program. The overall read and write operation characteristics include continuous operation addresses and operation address jumps. The read and write modes include sequential access and random access. The operation addresses corresponding to sequential access are continuous, and the operation addresses corresponding to random access are jumps. The throughput types include high throughput and low throughput. The total synchronous communication time is the total time used to perform data synchronization operations. The proportion of synchronous communication time is the ratio of the total synchronous communication time to the total execution time of the target program. The synchronization indicators are mainly the amount of data transmitted, the average delay information and the imbalance.

[0097] Specifically, by performing an overview analysis on the computing source database, the total number of instructions can be counted, and the ratio of the total number of instructions to the total execution time of the target program can be used as the average number of executed instructions. The ratio of the time the central processing unit is in a working state to the total execution time of the target program is used as the central processing unit utilization rate, and the time usage ratio of each function is calculated to obtain the function name with the highest time usage ratio. By performing an overview analysis on the read-write source database, the operation address corresponding to each data can be obtained, and the continuity of the operation address corresponding to each data can be determined to determine whether the overall read-write operation characteristics are continuous operation addresses or operation address jumps. Calculate the proportion of data with continuous operation addresses and the proportion of data with jumping operation addresses. If the data with a large proportion is continuous operation addresses, then the read-write mode of the target program is determined to be sequential read-write. If the data with a large proportion is operation address jumps, then the read-write mode of the target program is determined to be random read-write. Furthermore, multiple real-time bandwidths can be obtained from the read-write source database, and the ratio of the sum of each real-time bandwidth to the number is used as the average bandwidth, and the maximum value of each real-time bandwidth is used as the peak bandwidth. The average bandwidth and the peak bandwidth are calculated according to the preset weights to obtain a weighted value. If the weighted value is greater than the preset bandwidth threshold, the throughput type is determined to be high throughput, otherwise, the throughput type is determined to be low throughput. An overview analysis of the network communication source database can be performed to sum up the total time consumption of synchronous communication, and the ratio of the total time consumption of synchronous communication to the total execution time of the target program is used as the proportion of synchronous communication time consumption. Then, the synchronization indicators corresponding to each synchronization type are statistically obtained from the network communication source database, such as the amount of data transmitted, the average delay information and the imbalance. Finally, the total number of instructions, the average number of executed instructions, the CPU utilization rate, the function name with the highest time consumption, the read-write mode, the throughput type, the peak bandwidth, the total time consumption of synchronous communication, the proportion of synchronous communication time consumption, and the synchronization indicators corresponding to each synchronization type are used as the overview analysis results of the target program.

[0098] Exemplarily, the data statistics in the computing source database are used to generate the analysis conclusions of the target program on the CPU time, clarifying how many instructions (total number of instructions) the target program executes during operation, as well as the average number of instructions executed per second (average number of executed instructions), CPU utilization information (central processing unit utilization) and the function name with the highest time usage. Based on the data statistics in the read-write source database, the analysis conclusions of the target program on read and write operations are generated, clarifying the read and write mode (sequential access / random access) of the target program during operation, as well as the throughput type (high throughput / low throughput), and the peak bandwidth in reading and writing. Based on the data statistics in the network communication database, the analysis conclusions of the target program on network communication are generated, clarifying the total time consumed by MPI communication (total time consumed by synchronous communication) during the operation of the target program, as well as the percentage of the total operating time (percentage of synchronous communication time consumption), the main MPI type (synchronous type) and the amount of data transmitted, the average delay information and the degree of imbalance (synchronization index), etc.

[0099] S140: Determine the program performance analysis result corresponding to the target program according to the program type, the overview analysis result, and the performance analysis sub-results corresponding to each sub-database.

[0100] The program performance analysis results include program type, overview analysis results, and performance analysis sub-results corresponding to various performance analysis dimensions.

[0101] Specifically, the program type, overview analysis results, and performance analysis sub-results corresponding to each sub-database may be sorted according to a preset integration rule or arrangement order to obtain a comprehensive performance analysis display page as the program performance analysis result corresponding to the target program.

[0102] Optionally, the program performance analysis result corresponding to the target program may be determined according to the program type, the overview analysis result, and the performance analysis sub-results corresponding to each sub-database in the following manner: determining a program analysis overview according to the program type and the overview analysis result;

[0103] According to the program analysis overview and the performance analysis sub-results corresponding to each sub-database, the program performance analysis results corresponding to the target program are determined.

[0104] Among them, the program analysis overview can be understood as the homepage of the analysis, which is used to provide relevant performance analysis results of the target program at a macro level.

[0105] Specifically, the program type and the summary analysis results are used as the program analysis overview, and in the program analysis overview, jump links of the performance analysis sub-results corresponding to each sub-database can be set, so that when the user wants to view a certain performance analysis, the corresponding jump link can be triggered to switch the page to the corresponding performance analysis sub-result. The program analysis overview is used as the homepage of the program performance analysis results corresponding to the target program, and jump links are set to link the performance analysis sub-results corresponding to each sub-database, so as to construct the program performance analysis results for the convenience of user viewing.

[0106] Exemplarily, the program performance analysis results corresponding to the target program may be output to HTML (HyperTextMarkup Language) and a page of all program performance analysis results may be assembled.

[0107] Optionally, after determining the program performance analysis results corresponding to the target program, in order to improve operating efficiency, a data storage option is also provided. Users can choose to store the analyzed target format data in the converted Parquet format (a file format for column storage) locally. Parquet, as a column storage file format, supports advanced compression and encoding schemes, which can significantly reduce file size and improve read and write performance. Therefore, it is particularly suitable as a storage format for sampled data in high-performance computing applications. In this way, in subsequent analysis, the method provides an interface to directly read these compressed data, thereby speeding up data reading speed and data processing efficiency.

[0108] The present invention has the following technical effects: by obtaining a program source database corresponding to a target program applied to a high-performance cluster, analyzing each sub-database in the program source database separately, determining the performance analysis sub-results corresponding to the sub-database through cross-analysis and / or analysis one by one, so as to perform separate analysis in each performance analysis dimension, and then, according to the program source database, determining the overview analysis results and the program type of the target program, and determining the program performance analysis results corresponding to the target program according to the program type, the overview analysis results and the performance analysis sub-results corresponding to each sub-database, realizing the effect of automatically collecting and analyzing the key performance indicators of the target program in each performance analysis dimension on the high-performance cluster, displaying the analysis results in a user-friendly manner, and providing users with multi-dimensional program performance analysis conclusions, thereby helping users make more accurate performance tuning decisions on high-performance clusters. Example

[0109] Figure 2 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Figure 2 As shown, the electronic device 200 includes one or more processors 201 and a memory 202 .

[0110] The processor 201 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 200 to perform desired functions.

[0111] The memory 202 may include one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, a random access memory (RAM) and / or a cache memory (cache), etc. The non-volatile memory may include, for example, a read-only memory (ROM), a hard disk, a flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 201 may run the program instructions to implement the program performance analysis method of any embodiment of the present invention described above and / or other desired functions. Various contents such as initial external parameters, thresholds, etc. may also be stored in the computer-readable storage medium.

[0112] In one example, the electronic device 200 may further include: an input device 203 and an output device 204, which are interconnected via a bus system and / or other forms of connection mechanisms (not shown). The input device 203 may include, for example, a keyboard, a mouse, etc. The output device 204 may output various information to the outside, including early warning prompt information, braking force, etc. The output device 204 may include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto, etc.

[0113] Of course, to simplify, Figure 2 Only some of the components related to the present invention in the electronic device 200 are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, according to specific application conditions, the electronic device 200 may also include any other appropriate components. Example

[0114] In addition to the above methods and devices, an embodiment of the present invention may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the program performance analysis method provided by any embodiment of the present invention.

[0115] The computer program product may be written in any combination of one or more programming languages ​​to write program code for performing the operations of the embodiments of the present invention, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0116] In addition, an embodiment of the present invention may also be a computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are executed by a processor, the processor executes the steps of the program performance analysis method provided by any embodiment of the present invention.

[0117] The computer readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can include, for example, but is not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0118] It should be noted that the terms used in the present invention are only for describing specific embodiments, rather than limiting the scope of the present application. As shown in the present specification, unless the context clearly indicates an exception, the words "one", "a", "a kind of" and / or "the" do not specifically refer to the singular, but may also include the plural. The terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method or device. In the absence of further restrictions, the elements defined by the sentence "include one..." do not exclude the presence of other identical elements in the process, method or device including the elements.

[0119] It should also be noted that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. Unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", etc. should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be an indirect connection through an intermediate medium, or it can be a connection between the two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.

Claims

1. A program performance analysis method, characterized in that: include: Obtaining a program source database corresponding to a target program; wherein the target program is applied to a high-performance cluster; the program source database is a database consisting of source data required for analyzing performance in various dimensions collected when the target program is running on the high-performance cluster; For each sub-database in the program source database, determine the performance analysis sub-result corresponding to the sub-database through cross analysis and / or one-by-one analysis; wherein the sub-database is a partial database corresponding to different analysis dimensions in the program source database; the cross analysis is an analysis of multiple runs of the target program, and the one-by-one analysis is an analysis of a single run of the target program; Determining an overview analysis result and a program type of the target program according to the program source database; The program performance analysis result corresponding to the target program is determined according to the program type, the overview analysis result, and the performance analysis sub-results corresponding to each sub-database.

2. The method according to claim 1, characterized in that The sub-database includes a computing source database; Determining the performance analysis sub-results corresponding to the sub-database through cross analysis includes: In a case where the computing source database includes two first databases to be processed in the operating environment, data cleaning is performed on each first database to be processed, and each first database to be processed is split into an application function database and a system function database according to the first field; According to the second field, two application function databases are merged to obtain a first target database, two system function databases are merged to obtain a second target database, and two first to-be-processed databases are merged to obtain a third target database; For each target database, determine a usage time difference and a usage time percentage difference corresponding to the central processor according to the usage time of the central processor in the target database, and determine first target data from the target database as a performance analysis sub-result corresponding to the computing source database according to the usage time difference and the usage time percentage difference; The target database includes a first target database, a second target database and a third target database.

3. The method according to claim 1, characterized in that The sub-database includes a computing source database; By analyzing one by one, the performance analysis sub-results corresponding to the sub-database are determined, including: For the first database to be processed in each operating environment in the computing source database, determining first original information based on a first preset script; Determine a calling route according to the first to-be-processed database; Determine a target route according to the CPU time percentage corresponding to each call route, and determine the call information to be displayed according to the target route and the preset depth; According to the first preset chart requirement, the first to-be-processed database is processed to determine a first target chart; Performing data cleaning on the first database to be processed, and splitting the first database to be processed into an application function database and a system function database according to the first field; Using the first to-be-processed database, the application function database and the system function database as a central processor time-consuming analysis database; For each CPU time consumption analysis database, determine the corresponding CPU usage time and load imbalance, and determine the second target data according to the CPU usage time and the load imbalance respectively; The first original information, the to-be-displayed call information, the target route, the first target chart, and the second target data are determined as performance analysis sub-results corresponding to the computing source database.

4. The method according to claim 1, characterized in that: The sub-database includes a read-write source database; Determining the performance analysis sub-results corresponding to the sub-database through cross analysis includes: In the case where the read-write source database includes two second databases to be processed in the operating environment, for each read-write type, determining the number of first operations corresponding to the read-write type and the number of second operations corresponding to each preset file size interval from each second database to be processed; Determine a difference in the number of operations according to the first number of operations and each second number of operations corresponding to each operating environment, and determine each difference in the number of operations as a performance analysis sub-result corresponding to the read-write source database; Accordingly, the performance analysis sub-results corresponding to the sub-databases are determined by analyzing them one by one, including: For each second database to be processed, determining second original information based on a second preset script; For each read / write type, according to the read data and write data corresponding to the read / write type in the second to-be-processed database and in accordance with the second preset chart requirement, determine a second target chart corresponding to the read / write type; Determine a read timing diagram according to the read data in the second database to be processed, and determine a write timing diagram according to the write data in the second database to be processed; The second original information, the second target chart, the read timing chart, and the write timing chart are determined as performance analysis sub-results corresponding to the read-write source database.

5. The method according to claim 1, characterized in that The sub-database includes a network communication source database; Determining the performance analysis sub-results corresponding to the sub-database through cross analysis includes: In the case where the network communication source database includes two third databases to be processed under the operating environment, determining a calculation index corresponding to a data synchronization function for each third database to be processed; For each synchronization type, determining the number of operations and the average time consumption corresponding to the synchronization type from the third to-be-processed database; According to the calculation index in each operating environment, the number of operations corresponding to each synchronization type in each operating environment and the average time consumption, the data synchronization index difference is determined, and the data synchronization index difference is used as the performance analysis sub-result corresponding to the network communication source database.

6. The method according to claim 1, characterized in that The sub-database includes a network communication source database; By analyzing one by one, the performance analysis sub-results corresponding to the sub-database are determined, including: For the third database to be processed in each operating environment in the network communication source database, based on a third preset script, determining third original information; According to each synchronization type and the third database to be processed, and in accordance with the third preset chart requirement, a third target chart is determined; Determine the first function data according to the number of calls corresponding to the data synchronization functions of each synchronization type in the third database to be processed, determine the second function data according to the function time corresponding to each data synchronization function in the third database to be processed, and determine the third function data according to the amount of function data corresponding to each data synchronization function in the third database to be processed; A performance analysis sub-result corresponding to the network communication source database is determined according to the third original information, the third target chart, the first function data, the second function data, and the third function data.

7. The method according to claim 1, characterized in that The program source database includes a computing source database, a reading and writing source database, and a network communication source database; Determining the overview analysis result according to the program source database includes: Determine the total number of instructions, the average number of executed instructions, the CPU utilization rate, and the function name with the highest time usage according to the computing source database; Determine the operation address corresponding to each piece of data according to the read-write source database, determine the overall read-write operation characteristics according to each operation address, and determine the read-write mode according to the overall read-write operation characteristics; Determine an average bandwidth and a peak bandwidth according to the read-write source database, and determine a throughput type according to the average bandwidth and the peak bandwidth; Determine the total synchronization communication time, the synchronization communication time ratio and the synchronization index corresponding to each synchronization type according to the network communication source database; The total number of instructions, the average number of executed instructions, the CPU utilization, the function name with the highest time consumption, the read-write mode, the throughput type, the peak bandwidth, the total synchronous communication time consumption, the proportion of synchronous communication time consumption, and the synchronization indicators corresponding to each synchronization type are taken as the overview analysis results of the target program.

8. The method according to claim 1, characterized in that The step of obtaining a program source database corresponding to the target program includes: For each analysis dimension, according to the data collection module corresponding to the analysis dimension and the target program, an initial database corresponding to the analysis dimension is obtained; According to the data type conversion rule corresponding to the analysis dimension, the data type of the initial database is converted to obtain a program source database corresponding to the analysis dimension.

9. An electronic device, characterized in that: The electronic device comprises: Processor and memory; The processor is used to execute the steps of the program performance analysis method according to any one of claims 1 to 8 by calling the program or instruction stored in the memory.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program or an instruction, and the program or the instruction enables a computer to execute the steps of the program performance analysis method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method, system and equipment for monitoring performance index of application system and storage medium

    CN114020580A

  • Binary programmable method for application performance data collection

    US20090150874A1