Communication Trace Data Compression and Storage Method and System Based on Parallel Program Characteristics
By overloading the MPI function, the parallel program communication trace data is intercepted, hash value judgment and data structure recording are used, and the detailed stack information storage strategy is combined to solve the problem of redundancy of large-scale parallel program communication trace data, efficient data compression and storage is achieved, and performance analysis is supported with low overhead.
Patent Information
- Application Number
- CN202510480793.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The prior art has data redundancy problems in the processing of large-scale parallel program communication trace data, resulting in inefficient storage and degradation of computing performance, and existing tools may affect the accuracy of analysis during compression.
The communication trace data is intercepted by overloading the MPI function, the hash value is used to judge the first or repeated calls, and the records are taken using different data structures, and the data is compressed in combination with hash tables and linked lists. The detailed stack information is obtained by combining libunwind and elfutils, and the storage strategy is set for real-time storage.
It realizes a significant reduction in data scale while ensuring data integrity, supports low-overhead data analysis, improves analysis accuracy and storage efficiency, and can quickly obtain the latest performance data to optimize large-scale parallel programs.
Smart Images

Figure CN119988339B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication for large-scale parallel programs, and particularly to a method and system for compressing and storing communication trace data based on parallel program characteristics. Background Art
[0002] MPI (Message Passing Interface) plays a crucial role in the performance of high-performance computer systems. Through efficient task scheduling, high flexibility, and powerful parallel processing capabilities, MPI can significantly improve computing efficiency, shorten computing time, and ensure data consistency in complex computing tasks. Therefore, as high-performance computer systems become increasingly complex, it becomes particularly important to trace the execution, performance, and behavior of MPI communications in parallel programs, which will directly affect the efficiency and availability of the entire high-performance computer system. To efficiently locate and diagnose performance bottlenecks in parallel programs, existing performance analysis tools usually analyze based on trace data. These trace data can usually reflect information such as message passing, communication latency, packet size, and frequency between nodes and processes from the side. However, with the expansion of program and system scale, the amount of communication trace data generated is huge, and the problem of communication trace data redundancy becomes increasingly prominent, significantly affecting storage efficiency and computing performance. Since multiple processing units may generate and store similar or duplicate data simultaneously, resulting in waste of system resources and increased communication overhead, it is particularly important to effectively compress communication trace data. By implementing data compression technology, not only can the storage requirements and data transmission time be significantly reduced, but also the overall system operation efficiency and response speed can be improved. Facing the rapidly expanding field of parallel computing, adopting efficient compression algorithms and methods can effectively reduce data redundancy, optimize resource utilization, and thus promote the further development of high-performance computing.
[0003] To address this issue, mainstream performance analysis tools internationally, such as TAU, reduce the volume of generated trace data in various ways. For example, through selective instrumentation and restricting functions with short running times, the number of functions for which trace data needs to be collected is reduced. Or, trace data is merged after the fact and then converted into a more concise format such as oft2 to achieve compression of the trace data. HPCToolkit can sample a specified time period or a specified code range, thereby reducing the size of the collected trace data. Or, the sampling interval can be increased to reduce the amount of data collected by decreasing the frequency. Both of these methods will result in a decrease in information richness and affect the accuracy of the analysis. ITAC reduces the volume of trace data in two ways. One is by setting filtering options to collect only partial function trace data, and the other is by configuring the option VT_COMPRESS_RAW_DATA to compress the raw data (compress the raw data). However, since ITAC is commercial software, its source code and the generated trace files cannot be viewed.
[0004] Therefore, implementing a communication trace data compression and storage method based on the characteristics of parallel programs needs to face the following challenges: (1) To make the communication trace data of large-scale parallel programs accurate and sufficient, a means with high reliability is required to extract the trace data during the communication process of parallel programs; (2) As the scale expands, the communication trace data of parallel programs is complex and huge. For this characteristic of parallel programs, it is necessary to reduce the data scale without destroying the communication integrity and reduce the overhead required for data processing; (3) To avoid the situation of data loss caused by abnormal interruption of large-scale parallel programs, the communication trace data requires a good storage mechanism that can better observe whether the program execution state is abnormal while reducing the memory usage of the file system. Summary of the Invention
[0005] The technical problem to be solved by the present invention: In view of the above problems of the prior art, a communication trace data compression and storage method and system based on the characteristics of parallel programs are provided. The present invention aims to achieve the compression and storage of communication trace data, significantly reducing the data scale on the premise of ensuring the integrity and richness of the trace data, and providing low-overhead data support for users to analyze bottleneck problems related to the communication of large-scale parallel programs.
[0006] To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0007] A communication trace data compression and storage method based on the characteristics of parallel programs, comprising the following steps:
[0008] S1, for the target parallel program, intercept MPI calls by overloading MPI functions and collect communication trace data;
[0009] S2. For each piece of communication trace data, determine whether it is the communication trace data at the first call or the communication trace data generated by a repeated call. If it is the communication trace data at the first call, record it in the data header using the first data structure. If it is the communication trace data generated by a repeated call, record it in the data body using the second data structure, and the number of value fields in the second data structure is less than that in the first data structure.
[0010] S3. Store the communication trace data recorded in the data header and data body.
[0011] Optionally, in step S2, the processing of each piece of communication trace data includes:
[0012] S2.1 Save this piece of communication trace data to the temporary data key, and use the function name and stack address of the MPI function corresponding to this piece of communication trace data as inputs to generate a hash value tableindex for this MPI function.
[0013] S2.2 Find the index position index corresponding to the hash value tableindex in the communication trace data structure array table.
[0014] S2.3 Determine whether the index position index is empty. If the index position index is empty, return null. Otherwise, loop through the communication trace data structure array table until the communication trace data pointed to by the pointer ptr at the index position index is equal to the temporary data key, and return the pointer ptr at the index position index.
[0015] S2.4 If the result returned when determining whether the index position index is empty is null, determine that this piece of communication trace data is the communication trace data at the first call and record it in the data header using the first data structure. Otherwise, determine that this piece of communication trace data is the communication trace data generated by a repeated call and record it in the data body using the second data structure, and connect it to the pointer ptr at the index position index in a linked list manner.
[0016] Optionally, the communication trace data includes function number id, function name name, communication direction rank, start time star, execution time dur, computing time between communications computing, single - communication data size data, cumulative execution times count, cumulative execution time cumuative_time, and path number computing, where the computing time between communications computing is the difference between the start time star of the current MPI function and the end time of the previous MPI function.
[0017] Optionally, the first data structure includes all numerical fields in the communication trace data, and the second data structure includes the function number id, start time star, execution time dur, and communication computing time computing in the communication trace data.
[0018] Optionally, it further includes obtaining and storing the stack information of the MPI function:
[0019] S101, using the libunwind function to traverse the stack frames to obtain the stack address of the MPI function;
[0020] S102, using the elfutils function to parse the stack address to obtain and store the stack information. The stack information includes the source file name, function name, and line number, and the function name is the original function name obtained after restoring the function name modified by the compiler. When storing the stack information, first use an array to record all different stack addresses, and at the same time record the corresponding array subscript for each function according to the required stack address. At the end of the program, perform a unified parsing once. Directly search for the parsed stack information through the array subscript to ensure that a stack address will only be parsed once, and write the finally parsed stored stack information into the data header.
[0021] Optionally, when storing the communication trace data recorded in the data header and data body in step S3, it includes recording the call times of the MPI function and I / O operations and storing them through setting a storage policy: when the target parallel program starts, specify the MPI functions and I / O operations to be monitored by modifying the configuration file; create global counters for each monitored MPI function and I / O operation according to the specified parameters, and insert counter increment operations at the entrance of each monitored MPI function and I / O operation; during the execution of the target parallel program, detect the call of the MPI function and I / O operation through the counter increment operation inserted at the entrance of the MPI function and I / O operation, and record the number of times it is called by increasing the global counter. At the same time, record the current timestamp when the global counter is updated; regularly check the value of the global counter and the time interval between the values of the global counter; if the value of the global counter reaches the counter threshold or the time interval between the values of the global counter exceeds the preset expectation, trigger a storage operation to write the communication trace data recorded in the data header and data body into the file system, and clear the record to prepare to receive the communication trace data of the next cycle.
[0022] Optionally, when storing the communication trace data recorded in the data header and data body in step S3, it further includes periodically checking the memory size occupied by the communication trace data recorded in the data header and data body. If the memory size occupied by the communication trace data recorded in the data header and data body exceeds a preset maximum memory threshold, a storage operation is triggered to write the communication trace data recorded in the data header and data body into the file system, and the record is cleared to prepare to receive the communication trace data of the next cycle.
[0023] In addition, the present invention also provides a communication trace data compression and storage system based on parallel program characteristics, including a microprocessor and a memory connected to each other. The microprocessor is programmed or configured to execute the communication trace data compression and storage method based on parallel program characteristics.
[0024] In addition, the present invention also provides a computer-readable storage medium, in which a computer program or instruction is stored. The computer program or instruction is programmed or configured to execute the communication trace data compression and storage method based on parallel program characteristics through a processor.
[0025] In addition, the present invention also provides a computer program product, including a computer program or instruction. The computer program or instruction is programmed or configured to execute the communication trace data compression and storage method based on parallel program characteristics through a processor.
[0026] Compared with the prior art, the present invention can mainly achieve the following beneficial effects:
[0027] 1. The present invention obtains communication trace data based on the existing performance analysis tool mpiP, and through an efficient communication trace data compression and storage method, on the premise of ensuring the integrity and richness of the trace data, the data scale is significantly reduced compared with the previous situation, and it can provide low-overhead data support for users to analyze the bottleneck problems related to the communication of large-scale parallel programs.
[0028] 2. The communication trace data storage method of the present invention supports real-time storage, ensuring that analysts can quickly obtain the latest performance data and make responses and optimizations in a timely manner.
[0029] 3. By recording the detailed metadata of each communication event (such as timestamp, sending / receiving process, message type, etc.), the present invention can better analyze the communication mode and potential performance bottlenecks, providing strong support for optimizing large-scale parallel programs. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a schematic diagram of the basic process of the method of the embodiment of the present invention.
[0031] Figure 2Schematic diagram of the data collection process in the embodiment of the present invention, where the positions to be modified are within the dashed boxes.
[0032] Figure 3 Flowchart of the communication trace data processing in the embodiment of the present invention.
[0033] Figure 4 Interaction schematic diagram between the program and the data collector in the embodiment of the present invention.
[0034] Figure 5 Brief example of the data format compression of the communication trace data in the embodiment of the present invention, where the upper side of the arrow is the data format of the traditional sectional data; the lower side of the arrow is the data format of the communication trace data in the method of this embodiment.
[0035] Figure 6 Storage schematic diagram of the communication trace data in the embodiment of the present invention.
[0036] Figure 7 Schematic diagram of the first storage strategy in the embodiment of the present invention.
[0037] Figure 8 Schematic diagram of the second storage strategy in the embodiment of the present invention. Detailed implementation mode
[0038] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be further described in detail below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0039] As Figure 1 shown, the method for compressing and storing communication trace data based on the characteristics of parallel programs in this embodiment includes the following steps:
[0040] S1. Intercept MPI calls and collect communication trace data for the target parallel program by overloading MPI functions;
[0041] S2. Determine whether each communication trace data is the communication trace data at the first call or the communication trace data generated by repeated calls. If it is the communication trace data at the first call, record it into the data header using the first data structure. If it is the communication trace data generated by repeated calls, record it into the data body using the second data structure, and the numerical domain in the second data structure is less than that in the first data structure;
[0042] S3. Store the communication trace data recorded in the data header and the data body.
[0043] As an alternative implementation, specifically in this embodiment, communication trace data is obtained based on the existing performance analysis tool mpiP, and the data compression and storage method of this embodiment is applied to provide low-overhead data support for users to analyze bottleneck problems related to large-scale parallel program communication. As Figure 2 shown, the method of this embodiment needs to improve its data collection, data compression, and data analysis (including profiling analysis and trace analysis) for the existing performance analysis tool mpiP. Its overall working process is: data collection and data compression to obtain communication trace data and stack information, and then generate analysis reports and visualization reports through performance analysis.
[0044] As Figure 3 shown, in step S1 of this embodiment, it includes intercepting MPI calls to obtain the stack address and start time, calling the PMPI interface to complete communication, and obtaining communication trace data such as execution time and data size. This embodiment collects communication trace data by improving the data collection method of the performance analysis tool mpiP, and this method is called mpiPtrace. The performance analysis tool mpiP uses the PMPI mechanism to run pre-set data collection code before the execution of MPI functions, and finally uses the profiling mode to count data. The data collected in this way is accurate but not comprehensive enough. mpiPtrace adds a trace mode to this data collector, creates a structure for each call to record its data, and additionally collects computing information between communications that mpiP does not pay attention to. The pseudo-code is as follows:
[0045] MPI_ () { / / Overload the MPI function interface
[0046] mpiPif_MPI_ (); / / Data collection interface.
[0047] }
[0048] mpiPi_MPI_ () { ......
[0049] PMPI_ (); / / Call the PMPI interface to complete communication
[0050] ...... / / Collect data: data such as timestamp, message size, communication tag
[0051] mpiPi_updata(); / / Update and store data.
[0052] }
[0053] mpiP_updata() { ......
[0054] if (h_search(csp) == NULL) { / / csp is a pointer to a struct
[0055] csp = malloc(); / / Create a new struct to store data
[0056] }
[0057] mpiPi_cs_updata(csp); / / Update data in trace mode
[0058] }
[0059] Among them, refers to a custom name.
[0060] On the basis of retaining the dissection value information, mpiPtrace in this embodiment adds the acquisition of trace data. This requires certain modifications between data collection and storage to ensure that function information is not stored only in an accumulative manner. The trace data of mpiPtrace is collected through an improved data collector of mpiP. Figure 4This is an interaction diagram of the program and data collector in this embodiment. The performance data collected by the data collector includes three parts: communication performance data, computing time between communications, and stack information. mpiPtrace collects the communication performance data of MPI (information such as communication data size, communication direction, timestamp, etc.) through the PMPI interface. Compared with mpiP, it additionally collects communication direction, single communication execution time, and tag number information. mpiPtrace regards the difference between the end time of the current function and the start time of the next function as the computing time between two communications (collectively referred to as computing time). Using the computing time can quickly locate performance problems caused by non-communication part code. Although the computing time between communications can be calculated through communication trace data, the time overhead required for performance analysis will increase significantly. mpiPtrace sets two modes, namely the profiling mode and the tracing mode, for collecting information on the computing part between communications. The profiling mode statistics the cumulative execution time between two functions, and the tracing mode records each execution time. mpiPtrace combines libunwind and elfutils to obtain stack information, including source file name, function name, and line number. The libunwind function traverses the stack frames to obtain stack addresses, and elfutils resolves the stack addresses into function name and other information. Through the source file, line number, and function name, users can accurately find the specific location of the hot functions in the program code. At the same time, mpiPtrace decodes the function names of C++ (restore the function names modified by the compiler) to facilitate users to query data information due to the characteristic that C++ will modify function names. These three parts of data ensure the accuracy of subsequent performance analysis and provide important data support for subsequent performance analysis work.
[0061] To ensure that the data is complete enough to accurately locate the performance bottleneck of large-scale parallel programs. This embodiment records the necessary information for each call, such as function number, function name, communication direction, time information, computing time between communications, single communication data size, cumulative execution times, time and path number, etc. However, this direct output method will generate huge trace data files. Therefore, while ensuring the accuracy of performance analysis and reducing the trace data output file, this embodiment compresses each communication trace data in combination with the communication characteristics of the MPI program SPMD (Single Program Multiple Data). Specifically, in step S2 of this embodiment, the processing of each communication trace data includes:
[0062] S2.1, save the communication trace data to the temporary data key, and use the function name and stack address of the MPI function corresponding to the communication trace data as input to generate a hash value tableindex for the MPI function;
[0063] S2.2. Find the index position index corresponding to the hash value tableindex in the communication trace data structure array table;
[0064] S2.3. Determine whether the index position index is empty. If the index position index is empty, return null; otherwise, loop through the communication trace data structure array table until the communication trace data pointed to by the pointer ptr at the index position index is equal to the temporary data key, and return the pointer ptr at the index position index;
[0065] S2.4. If the result returned when determining whether the index position index is empty is null, it is determined that the communication trace data is the communication trace data at the first call, and it is recorded in the data header ("head") using the first data structure (callsite_stats_t); otherwise, it is determined that the communication trace data is the communication trace data generated by repeated calls, and it is recorded in the data body ("body") using the second data structure (mpiP_coll_data), and is connected to the pointer ptr (index→ptr) at the index position index in a linked list manner.
[0066] In this embodiment, when compressing each communication trace data in combination with the communication characteristics of the MPI program SPMD (Single Program Multiple Data), the "head" information is first output. On the basis of the dissection value output format, in order to quickly perform dissection value analysis, two types of information, the cumulative execution times and the cumulative execution time, are added between the single communication data size and the path number (corresponding one by one to id, name, etc. in Figure 5 . Then the "body" information is output. The "body" is the compressed data, including the time series information of each MPI function call, the calculation time, and the tag number (used for point-to-point communication matching). Therefore, the trace format output records the detailed information of each MPI function call additionally on the basis of the dissection value. Since some content is missing in the compressed data of the trace format, in order to facilitate subsequent performance analysis to reproduce the communication behavior of the program and ensure its accuracy, the trace output format outputs the performance data with the same communication behavior concentrated together, which is extremely convenient when filling (decompressing) the data for the "body". When decompressing, only the data in the "head" needs to be read to batch decompress the data in the "body".
[0067] After the trace function is implemented, how to effectively control the overhead caused by the implementation of the trace function becomes a challenge. The design, storage method and query method of the structure will affect the overhead caused by the trace function. In order to minimize this part of the overhead, this embodiment proposes a new communication trace data compression method. The "head" is combined with the "body" in the form of an array to store the "head" (callsite_stats_t structure). The "head" records the communication data at the first call, as well as the cumulative number of executions and the cumulative execution time. The communication data generated by repeated calls is recorded using the "body" (mpiP_coll_data) after compression, and is connected to the "head" in the form of a linked list. Since the hash value is used to locate the position of the "head" in the array, multiple "heads" may be stored in the same position. Compared with the initial use of data storage in a single structure callsite_stats_t, this method significantly compresses the amount of data. For a large amount of data in the trace mode, this embodiment uses a hash search method to query the storage location of the structure to reduce time complexity. In this embodiment, a combination of "head" and "body" is used. The structure callsite_stats_t that stores the profile information is regarded as the "head", and a linked list that stores the compressed information is added to it. The linked list consists of the structure mpiP_coll_data and is regarded as the "body". Only different data is recorded in the "body", such as timestamps, communication calculation time and tag information. When a hash conflict occurs, the "head" linked list is traversed, and its time complexity is O(n), where n is the number of elements in the linked list. When the value of n is extremely small, the upper limit depends only on the number of MPI function references in the source code. However, in order to ensure the orderliness of the output data (that is, the data of the same communication behavior is connected together for output), it is necessary to traverse each element in the linked list corresponding to the hash value to find the most suitable position for insertion. This time complexity is also O(n). Since the MPI function will loop multiple times in a large-scale parallel program, the value of n will increase significantly. In order to ensure the integrity of the information and eliminate redundant information at the same time, mpiPtrace improves the structure search and comparison method, through which mpiPtrace realizes trace data compression. mpiPtrace adds a series of parameters, such as the amount of data and the number of the sending and receiving processes, and connects the function data that are consistent after comparison through a linked list. When it is called for the first time, redundant data such as the function name, communication data, and stack information will be recorded in the "head". After the data information of subsequent repeated calls is compressed, only the time series information and tag number need to be recorded in the "body".
[0068] like Figure 5As shown, the communication trace data in this embodiment includes function number id, function name name, communication direction rank, start time star, execution time dur, computing time between communications computing, single - communication data size data, cumulative execution times count, cumulative execution time cumuative_time, and path number computing. Among them, the computing time between communications computing is the difference between the start time star of the current MPI function and the end time of the previous MPI function.
[0069] In this embodiment, the first data structure (callsite_stats_t) includes all numerical fields in the communication trace data. The second data structure (mpiP_coll_data) includes the function number id, start time star, execution time dur, and computing time between communications computing in the communication trace data. The communication trace data composed of the data header and data body obtained at a certain moment is as Figure 6 shown. Whenever a new communication trace data occurs, a first data structure (callsite_stats_t) will be generated and written into the data header. If there is communication trace data with the same communication behavior later, a second data structure (mpiP_coll_data) will be generated and written into the data body, and so on. Eventually, multiple first data structures will be formed in the data header, and multiple groups of second data structures (mpiP_coll_data) will be included in the data body.
[0070] After completing the collection and compression of the trace data, in order to accurately locate the program performance bottleneck, it is also necessary to obtain detailed stack information. Although mpiP supports stack parsing, it is found in actual use that the stack information parsed by mpiP is not comprehensive. To solve the above problems, this embodiment also includes obtaining and storing the stack information of MPI functions:
[0071] S101, use the libunwind function to traverse the stack frames to obtain the stack address of the MPI function;
[0072] S102, use elfutils functions to parse the stack address to obtain stack information and store it. The stack information includes the source file name, function name, and line number, and the function name is the original function name obtained after restoring the function name modified by the compiler. When storing the stack information, first use an array to record all different stack addresses, and at the same time record the corresponding array subscript for each function according to the required stack address. At the same time, after the program ends, perform a unified parsing, and directly find the parsed stack information through the array subscript to ensure that a stack address will only be parsed once, and write the finally parsed stored stack information into the data header. In this embodiment, the elfutils library is introduced into mpiPtrace to perform more detailed parsing of the stack information. There are a large number of duplicate parts in the stack information of functions, and the same stack address will be parsed multiple times, which leads to many duplicate parsing operations. To eliminate this part of the duplicate parsing operations, mpiPtrace first uses an array to write all different stack addresses, and at the same time records the corresponding array subscript for each function according to the required stack address. After the program ends, perform a unified parsing, and directly find the parsed stack information through the array subscript, so as to ensure that a stack address will only be parsed once. At the same time, the stack information will only be recorded in the "head". Use the NPB program with 256 processes to conduct experiments on it. The results show that the size of the mpiPtrace data file is reduced by up to 92.17% compared with Tau; in terms of time overhead, it is reduced by about 3% compared with Tau.
[0073] Such as Figure 7As shown, when storing the communication trace data recorded in the data header and data body in step S3, it includes recording the call times of MPI functions and I / O operations and storing them by setting a storage policy: when the target parallel program starts, specify the MPI functions and I / O operations to be monitored by modifying the configuration file; create global counters for each monitored MPI function and I / O operation according to the specified parameters, and insert counter increment operations at the entry of each monitored MPI function and I / O operation; during the execution of the target parallel program, detect the call of MPI functions and I / O operations through the counter increment operations inserted at the entry of MPI functions and I / O operations, record the number of times they are called by increasing the global counter, and record the current timestamp while the global counter is updated; regularly check the value of the global counter and the time interval between the values of the global counter; if the value of the global counter reaches the counter threshold or the time interval between the values of the global counter exceeds the preset expectation, trigger a storage operation to write the communication trace data recorded in the data header and data body to the file system, and clear the record to prepare to receive the communication trace data of the next cycle. By monitoring the call times of MPI functions and their time intervals, programmers can determine whether there are load imbalance problems in the program. For example, when some processes perform a large amount of calculations while other processes perform relatively few calculations, the slower processes may cause other processes to wait for data or synchronize. This unbalanced task allocation may cause some processes to be idle for a long time. In addition, with the call times of MPI functions, programmers can better observe whether the program execution status conforms to the expected design of the algorithm. This process is of great significance for optimizing the performance of the program. At the same time, we also record I / O operations and their time intervals to prevent data loss caused by abnormal interrupts in large-scale parallel programs. In this way, the integrity and reliability of important data can be ensured. This embodiment provides an effective storage mechanism in the above manner for recording the call times of MPI functions and I / O operations, and optimizing the management of communication trace data by setting a storage policy, thereby improving the efficiency of performance analysis. This embodiment also provides a user interface that enables users to query the current monitored data statistics in real time through this interface, including the call times, time intervals, memory status, etc. of each monitored MPI function and I / O operation. With the help of this interface, users can adjust the monitoring configuration and storage policy in a timely manner according to the program execution situation, which can better help users locate problems.
[0074] As Figure 8As shown, when storing the communication trace data recorded in the data header and data body in step S3, it also includes periodically checking the memory size occupied by the communication trace data recorded in the data header and data body. If the memory size occupied by the communication trace data recorded in the data header and data body exceeds the preset maximum memory threshold, a storage operation is triggered to write the communication trace data recorded in the data header and data body into the file system, and the record is cleared to prepare to receive the communication trace data of the next cycle. Through the above method, this embodiment provides another effective storage mechanism. By setting a maximum memory threshold, when the communication trace data accumulates to this threshold, a storage operation is automatically triggered to write the current communication trace data into the file system, and the communication trace data file is cleared to prepare to receive the communication trace data of the next cycle, so as to achieve automatic storage according to the file size.
[0075] In summary, this embodiment obtains communication trace data by overloading MPI functions. This embodiment proposes a communication trace data format and rearranges the format of the obtained communication trace data based on the characteristics of parallel programs. This embodiment uses the formatted communication trace data header as the key value and compresses the communication trace data using the hash method. This embodiment also proposes a new storage method, and stores the compressed communication trace data in the file system at an opportune moment according to the communication characteristics and trace data volume during runtime. The method of this embodiment significantly reduces the data scale compared with the previous one through efficient communication trace data compression and storage methods while ensuring the integrity and richness of the trace data. The method of this embodiment supports real-time storage, ensuring that analysts can quickly obtain the latest performance data and make timely responses and optimizations. By recording the detailed metadata of each communication event (such as timestamp, sending / receiving process, message type, etc.), the method of this embodiment can better analyze the communication pattern and potential performance bottlenecks, thus providing strong support for optimizing large-scale parallel programs.
[0076] In addition, this embodiment also provides a communication trace data compression and storage system based on the characteristics of parallel programs, including a microprocessor and a memory connected to each other. The microprocessor is programmed or configured to execute the communication trace data compression and storage method based on the characteristics of parallel programs.
[0077] In addition, this embodiment also provides a computer-readable storage medium, in which a computer program or instruction is stored. The computer program or instruction is programmed or configured to execute the communication trace data compression and storage method based on the characteristics of parallel programs through a processor.
[0078] In addition, this embodiment also provides a computer program product, including a computer program or instruction. The computer program or instruction is programmed or configured to execute the communication trace data compression and storage method based on the characteristics of parallel programs through a processor.
[0079] Those skilled in the art should understand that the technical solutions provided by the embodiments of the present invention can be in the form of methods, systems, or computer program products. Therefore, the present invention can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program codes. The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0080] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.
Claims
1. A communication trace data compression and storage method based on parallel program characteristics, characterized in that It includes the following steps: S1. Intercept MPI calls for the target parallel program by overloading MPI functions and collect communication trace data; S2. For each piece of communication trace data, determine whether it is the communication trace data at the first call or the communication trace data generated by repeated calls. If it is the communication trace data at the first call, record it into the data header using the first data structure. If it is the communication trace data generated by repeated calls, record it into the data body using the second data structure, and the number of value fields in the second data structure is less than that in the first data structure; S3. Store the communication trace data recorded in the data header and data body.
2. The communication trace data compression and storage method based on parallel program characteristics according to claim 1, wherein In step S2, the processing of each piece of communication trace data includes: S2.
1. Save this piece of communication trace data to the temporary data key, and use the function name and stack address of the MPI function corresponding to this piece of communication trace data as inputs to generate a hash value tableindex for this MPI function; S2.
2. Find the corresponding index position index in the communication trace data structure array table for the hash value tableindex; S2.
3. Determine whether the index position index is empty. If the index position index is empty, return null. Otherwise, loop through the communication trace data structure array table until the communication trace data pointed to by the pointer ptr at the index position index is equal to the temporary data key, and return the pointer ptr at the index position index; S2.
4. If the result returned when determining whether the index position index is empty is null, determine that this piece of communication trace data is the communication trace data at the first call and record it into the data header using the first data structure. Otherwise, determine that this piece of communication trace data is the communication trace data generated by repeated calls and record it into the data body using the second data structure, and connect it with the pointer ptr at the index position index in a linked list manner.
3. The communication trace data compression and storage method based on parallel program characteristics according to claim 2, characterized in that The communication trace data includes function number id, function name name, communication direction rank, start time star, execution time dur, computing time between communications computing, single - communication data size data, cumulative execution times count, cumulative execution time cumuative_time, and path number computing, where the computing time between communications computing is the difference between the start time star of the current MPI function and the end time of the previous MPI function.
4. The method for compressing and storing communication trace data based on parallel program characteristics according to claim 3, wherein The first data structure includes all value fields in the communication trace data, and the second data structure includes function number id, start time star, execution time dur, and computing time between communications computing in the communication trace data.
5. The communication trace data compression and storage method based on parallel program characteristics according to claim 1, characterized in that It also includes obtaining and storing the stack information of MPI functions: S101. Use the libunwind function to traverse the stack frames to obtain the stack address of the MPI function; S102. Use elfutils functions to parse the stack address to obtain stack information and store it. The stack information includes the source file name, function name, and line number, and the function name is the original function name obtained after restoring the function name modified by the compiler. When storing the stack information, first use an array to record all different stack addresses, and at the same time record the corresponding array subscript for each function according to the required stack address. At the end of the program, perform a unified parsing once. Directly search for the parsed stack information through the array subscript to ensure that a stack address is only parsed once, and write the finally parsed stored stack information into the data header.
6. The communication trace data compression and storage method based on parallel program features according to claim 1, wherein When storing the communication trace data recorded in the data header and data body in step S3, it includes recording the call times of MPI functions and I / O operations and storing them by setting a storage policy: when the target parallel program starts, specify the MPI functions and I / O operations to be monitored by modifying the configuration file; create global counters for each monitored MPI function and I / O operation according to the specified parameters, and insert counter increment operations at the entrance of each monitored MPI function and I / O operation. During the execution of the target parallel program, detect the call of MPI functions and I / O operations through the counter increment operations inserted at the entrance of MPI functions and I / O operations, record the number of times they are called by increasing the global counter, and record the current timestamp while updating the global counter; regularly check the value of the global counter and the time interval between the values of the global counter. If the value of the global counter reaches the counter threshold or the time interval between the values of the global counter exceeds the preset expectation, trigger a storage operation to write the communication trace data recorded in the data header and data body into the file system, and clear the record to prepare to receive the communication trace data of the next cycle.
7. The communication trace data compression and storage method based on parallel program characteristics according to claim 1, wherein When storing the communication trace data recorded in the data header and data body in step S3, it also includes regularly checking the memory size occupied by the communication trace data recorded in the data header and data body. If the memory size occupied by the communication trace data recorded in the data header and data body exceeds the preset maximum memory threshold, trigger a storage operation to write the communication trace data recorded in the data header and data body into the file system, and clear the record to prepare to receive the communication trace data of the next cycle.
8. A communication trace data compression and storage system based on parallel program characteristics, including a microprocessor and a memory connected to each other, characterized in that, The microprocessor is programmed or configured to execute the communication trace data compression and storage method based on the parallel program characteristics described in any one of claims 1 to 7.
9. A computer-readable storage medium storing a computer program or instructions, characterized in that, The computer program or instruction is programmed or configured to execute the communication trace data compression and storage method based on the parallel program characteristics described in any one of claims 1 to 7 through a processor.
10. A computer program product, comprising a computer program or instructions, characterized in that, The computer program or instruction is programmed or configured to execute the communication trace data compression and storage method based on the parallel program characteristics described in any one of claims 1 to 7 through a processor.
Citation Information
Patent Citations
Fast information interaction mechanism in multimedia system and network transmission method
CN107026887A
Function call stack analysis and backtracking method and device
CN115292201A