Communication trace data compression and storage method and system based on parallel program characteristics

By overloading MPI functions in parallel programs to collect trace data and using specific data structures and indexes for compression, the redundancy of communication trace data in large-scale parallel programs is solved, and the data scale is significantly reduced and analysis efficiency is improved.

CN119988339AActive Publication Date: 2025-05-13NAT UNIV OF DEFENSE TECH

Patent Information

Application Number
CN202510480793.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-13
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

In large-scale parallel programs, the amount of communication trace data is huge and the redundancy is obvious, which affects storage efficiency and computing performance. The information richness of the existing technology decreases when compressing data, affecting the accuracy of analysis.

Method used

The MPI call is intercepted by overloading the MPI function, the communication trace data is collected, and the first and second data structures are used to record the first and repeated call data respectively, and the hash table index and linked list connection are used to realize efficient compression and storage of data.

Benefits of technology

On the premise of ensuring the integrity and richness of trace data, it significantly reduces data scale, reduces data processing overhead, supports real-time storage, improves analysis efficiency, and promptly discovers performance bottlenecks of parallel programs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988339A_ABST
    Figure CN119988339A_ABST
Patent Text Reader

Abstract

The invention discloses a communication trace data compression and storage method and system based on parallel program characteristics. The method comprises the steps that MPI calling is intercepted through a reloading MPI function for a target parallel program, and communication trace data is collected; for each piece of communication trace data, if the data is the communication trace data during first calling, recording the data into a data head by adopting a first data structure body, and if the data is the communication trace data generated during repeated calling, recording the data into a data body by adopting a second data structure body, the numerical field in the second data structure body is less than the numerical field of the first data structure body; and storing the communication trace data recorded in the data head and the data body. The invention aims to realize communication trace data compression and storage, the data scale is obviously reduced on the premise of ensuring the trace data to be complete and rich in content, and low-overhead data support is provided for users to analyze bottleneck problems related to large-scale parallel program communication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of large-scale parallel program communication, and in particular to a communication trace data compression and storage method and system based on parallel program characteristics. Background Art

[0002] MPI (Message Passing Interface) plays a vital role in the performance of high-performance computer systems. Through efficient task scheduling, high flexibility and powerful parallel processing capabilities, MPI can significantly improve computing efficiency, shorten computing time, and ensure data consistency in complex computing tasks. Therefore, as high-performance computer systems become more and more complex, it is particularly important to track the execution, performance and behavior of MPI communication of parallel programs, which will directly affect the efficiency and availability of the entire high-performance computer system. In order to efficiently locate and diagnose the performance bottleneck of parallel programs, existing performance analysis tools are usually analyzed based on trace data. These trace data can usually reflect information such as message passing, communication delay, size and frequency of data packets between nodes and processes from the side. However, as the scale of programs and systems expands, the amount of communication trace data generated is huge, and the redundancy problem of communication trace data is becoming increasingly prominent, which significantly affects storage efficiency and computing performance. Since multiple processing units may generate and store similar or repeated data at the same time, resulting in waste of system resources and increased communication overhead, it is particularly important to effectively compress communication trace data. By implementing data compression technology, not only can storage requirements and data transmission time be significantly reduced, but the overall system computing efficiency and response speed can also be improved. In the face of the rapidly expanding field of parallel computing, adopting efficient compression algorithms and methods can effectively reduce data redundancy, optimize resource utilization, and promote the further development of high-performance computing.

[0003] In order to solve this problem, the current mainstream performance analysis tools in the world, such as TAU, reduce the amount of trace data generated in a variety of ways, such as selectively inserting stubs and limiting short-running functions to reduce the number of functions that need to collect trace data. Or merge the trace data afterwards and convert it into a more concise format such as oft2 to compress the trace data. HPCToolkit can sample a specified time period or a specified code range to reduce the size of the collected trace data. Or increase the sampling interval and reduce the amount of data collected by reducing the frequency. Both of the above methods will cause a decrease in information richness and affect the accuracy of the analysis. ITAC reduces the amount of trace data in two ways. One is to collect only part of the function trace data by setting the filtering option, and the other is to compress the raw data by configuring the option VT_COMPRESS_RAW_DATA (compressing the raw data). However, since ITAC is a commercial software, its source code and the generated trace files cannot be viewed.

[0004] Therefore, the implementation of communication trace data compression and storage methods based on the characteristics of parallel programs needs to face the following challenges: (1) In order to make the communication trace data of large-scale parallel programs accurate and sufficient, a highly reliable means is needed to extract the trace data during the parallel program communication process; (2) As the scale increases, the communication trace data of parallel programs becomes complex and large. In view of this characteristic of parallel programs, it is necessary to reduce the data size and the overhead required for data processing without destroying the integrity of communication; (3) In order to avoid data loss caused by abnormal interruptions of large-scale parallel programs, communication trace data requires a good storage mechanism that can better observe whether the program execution status is abnormal while reducing the file system memory usage. Summary of the invention

[0005] Technical problem to be solved by the present invention: In view of the above-mentioned problems in the prior art, a communication trace data compression and storage method and system based on the characteristics of parallel programs are provided. The present invention aims to realize communication trace data compression and storage, and significantly reduce the data size while ensuring that the trace data is complete and rich in content, so as to provide low-overhead data support for users to analyze bottleneck problems related to communication of large-scale parallel programs.

[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is: A communication trace data compression and storage method based on parallel program characteristics comprises the following steps: S1, intercepts MPI calls by overloading MPI functions for the target parallel program and collects communication trace data; S2, for each piece of communication trace data, determining whether it is the communication trace data of the first call or the communication trace data generated by repeated calls, if it is the communication trace data of the first call, recording it in the data header using a first data structure, if it is the communication trace data generated by repeated calls, recording it in the data body using a second data structure, and the value field in the second data structure is less than the value field in the first data structure; S3, stores the communication trace data recorded in the data header and data body.

[0007] Optionally, in step S2, the processing of each piece of communication trace data includes: S2.1, save the communication trace data to the temporary data key, take the function name and stack address of the MPI function corresponding to the communication trace data as input, and generate a hash value tableindex for the MPI function; S2.2, find the index position index corresponding to the hash value tableindex in the communication trace data structure array table; S2.3, determine whether the index position index is empty. If the index position index is empty, return empty; otherwise, loop through the communication trace data structure array table until the communication trace data pointed to by the pointer ptr of the index position index is found to be equal to the temporary data key, and return the pointer ptr of the index position index; S2.4, if the result returned when judging whether the index position index is empty is empty, it is determined that it is the communication trace data when the communication trace data is called for the first time, and the first data structure is used to record it in the data header; otherwise, it is determined that the communication trace data is the communication trace data generated by repeated calls, and the second data structure is used to record it in the data body, and it is connected to the pointer ptr of the index position index through a linked list.

[0008] Optionally, the communication trace data includes function number id, function name name, communication direction rank, start time star, execution time dur, communication computing time computing, single communication data size data, cumulative execution times count, cumulative execution time cumulative_time and path number computing, where the communication computing time computing is the difference between the start time star of the current MPI function and the end time of the previous MPI function.

[0009] Optionally, the first data structure includes all numerical fields in the communication trace data, and the second data structure includes the function number id, start time star, execution time dur, and communication calculation time computing in the communication trace data.

[0010] Optionally, it also includes obtaining and storing the stack information of the MPI function: S101, using libunwind function to traverse the stack frame to obtain the stack address of the MPI function; S102, using the elfutils function to parse the stack address to obtain stack information and store it, wherein the stack information includes the source file name, function name and line number, and the function name is the original function name obtained by restoring the function name modified by the compiler; and storing the stack information includes first using an array to record all different stack addresses, and recording the corresponding array index for each function according to the required stack address, and performing a unified analysis after the program ends, directly searching the stack information after analysis through the array index to ensure that a stack address is only analyzed once, and writing the storage stack information finally analyzed into the data header.

[0011] Optionally, when storing the communication trace data recorded in the data header and the data body in step S3, it includes recording the number of calls of the MPI function and the I / O operation and storing them by setting a storage strategy: when the target parallel program is started, the MPI function and I / O operation to be monitored are specified by modifying the configuration file; a global counter is created for each monitored MPI function and I / O operation according to the specified parameters, and a counter increment operation is inserted at the entry of each monitored MPI function and I / O operation; during the execution of the target parallel program, the calling MPI function and I / O operation is detected by the counter increment operation inserted at the entry of the MPI function and I / O operation, and the number of times it is called is recorded by increasing the global counter, and the current timestamp is recorded when the global counter is updated; the value of the global counter and the time interval between the values ​​of the global counter are regularly checked; if the value of the global counter reaches the counter threshold or the time interval between the values ​​of the global counter exceeds the preset expectation, the storage operation is triggered to write the communication trace data recorded in the data header and the data body to the file system, and the record is cleared to prepare for receiving the communication trace data of the next cycle.

[0012] Optionally, when storing the communication trace data recorded in the data header and the data body in step S3, it also includes periodically checking the memory size occupied by the communication trace data recorded in the data header and the data body. If the memory size occupied by the communication trace data recorded in the data header and the data body exceeds a preset maximum memory threshold, a storage operation is triggered to write the communication trace data recorded in the data header and the data body to the file system, and the record is cleared to prepare for receiving the communication trace data of the next cycle.

[0013] In addition, the present invention also provides a communication trace data compression and storage system based on parallel program features, including a microprocessor and a memory connected to each other, wherein the microprocessor is programmed or configured to execute the communication trace data compression and storage method based on parallel program features.

[0014] In addition, the present invention also provides a computer-readable storage medium, in which a computer program or instruction is stored. The computer program or instruction is programmed or configured to execute the communication trace data compression and storage method based on parallel program features through a processor.

[0015] In addition, the present invention also provides a computer program product, including a computer program or instructions, which are programmed or configured to execute the communication trace data compression and storage method based on parallel program features through a processor.

[0016] Compared with the prior art, the present invention can mainly achieve the following beneficial effects: 1. The present invention obtains communication trace data based on the existing performance analysis tool mpiP, and through an efficient communication trace data compression and storage method, the data scale is significantly reduced compared with before while ensuring that the trace data is complete and rich in content, which can provide low-overhead data support for users to analyze bottleneck problems related to communication in large-scale parallel programs.

[0017] 2. The communication trace data storage method of the present invention supports real-time storage, ensuring that analysts can quickly obtain the latest performance data and make timely responses and optimizations.

[0018] 3. By recording detailed metadata of each communication event (such as timestamp, sending / receiving process, message type, etc.), the present invention can better analyze communication patterns and potential performance bottlenecks, providing strong support for optimizing large-scale parallel programs. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 Schematic diagram of the basic flow of the method of the embodiment of the present invention.

[0020] Figure 2Schematic diagram of the data collection process in an embodiment of the present invention, wherein the dotted box indicates the position to be modified.

[0021] Figure 3 The figure is a flow chart of communication trace data processing in an embodiment of the present invention.

[0022] Figure 4 Schematic diagram of the interaction between the program and the data collector in an embodiment of the present invention.

[0023] Figure 5 This is a brief example of the data format compression of the communication trace data in the embodiment of the present invention, wherein the upper side of the arrow is the data format of the traditional profile data; the lower side of the arrow is the data format of the communication trace data in the method of this embodiment.

[0024] Figure 6 Schematic diagram of storage of communication trace data in an embodiment of the present invention.

[0025] Figure 7 This is a schematic diagram of the first storage strategy in an embodiment of the present invention.

[0026] Figure 8 Schematic diagram of the second storage strategy in an embodiment of the present invention. DETAILED DESCRIPTION

[0027] In order to enable those skilled in the art to better understand the technical solution of the present invention, the technical solution of the present invention will be further described in detail below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0028] like Figure 1 As shown, the communication trace data compression and storage method based on the parallel program feature of this embodiment includes the following steps: S1, intercepts MPI calls by overloading MPI functions for the target parallel program and collects communication trace data; S2, for each piece of communication trace data, determining whether it is the communication trace data of the first call or the communication trace data generated by repeated calls, if it is the communication trace data of the first call, recording it in the data header using a first data structure, if it is the communication trace data generated by repeated calls, recording it in the data body using a second data structure, and the value field in the second data structure is less than the value field in the first data structure; S3, stores the communication trace data recorded in the data header and data body.

[0029] As an optional implementation, this embodiment specifically obtains communication trace data based on the existing performance analysis tool mpiP, and applies the data compression and storage method of this embodiment to provide low-overhead data support for users to analyze bottleneck problems related to communication in large-scale parallel programs. Figure 2 As shown, the method of this embodiment aims at improving the data collection, data compression and data analysis (including profile analysis and trace analysis) of the existing performance analysis tool mpiP, and its overall working process is: data collection and data compression to obtain communication trace data and stack information, and then generating analysis reports and visualization reports through performance analysis.

[0030] like Figure 3 As shown, step S1 of this embodiment includes intercepting the MPI call to obtain the stack address and start time, calling the PMPI interface to complete the communication, and obtaining communication trace data such as execution time and data size. This embodiment collects communication trace data by improving the data collection method of the performance analysis tool mpiP, and this method is called mpiPtrace. The performance analysis tool mpiP uses the PMPI mechanism to run the preset data collection code before the MPI function is executed, and finally uses the profile mode to calculate the data. The data collected in this way is accurate but not comprehensive enough. mpiPtrace adds a trace mode to the data collector, creating a structure to record its data for each call. It also collects additional communication calculation information that mpiP does not pay attention to. The pseudo code is as follows: MPI_ () { / / Overload MPI function interface mpiPif_MPI_ (); / / Data collection interface.

[0031] } mpiPi_MPI_ () { ...... PMPI_ (); / / Call PMPI interface to complete communication ...... / / Collect data: timestamp, message size, communication tag, etc. mpiPi_updata(); / / Update and store data.

[0032] } mpiP_updata(){ ...... if(h_search(csp) == NULL){ / / csp is a structure pointer csp = malloc() / / Create a new structure to store data } mpiPi_cs_updata(csp); / / Update data in trace mode } in, Refers to the custom name.

[0033] On the basis of retaining the profile information, mpiPtrace in this embodiment adds the acquisition of trace data. This requires certain modifications between data collection and storage to ensure that function information is not only stored in an accumulated manner. The trace data of mpiPtrace is collected by improving the data collector of mpiP. Figure 4 The diagram of the interaction between the program and the data collector in this embodiment is shown in FIG. 1 . The performance data collected by the data collector includes three parts: communication performance data, communication calculation time and stack information. mpiPtrace collects the communication performance data (communication data size, communication direction, timestamp and other information) of MPI through the PMPI interface. Compared with mpiP, it additionally collects the communication direction, single communication execution time and tag number information. mpiPtrace regards the difference between the end time of the current function and the start time of the next function as the calculation time between two communications (collectively referred to as calculation time). The calculation time can be used to quickly locate performance problems caused by the non-communication part of the code. Although the communication calculation time can be calculated through the communication trace data, the time overhead required for performance analysis will increase significantly. mpiPtrace sets two modes, namely, profile and trace, for the collection of the calculation part information between communications. The profile mode counts the cumulative execution time between two functions, and the trace mode records the execution time of each time. mpiPtrace combines libunwind with elfutils to obtain stack information, including source file name, function name and line number. The libunwind function traverses the stack frame to obtain the stack address, and elfutils resolves the stack address into information such as function names. Through the source file, line number, and function name, users can accurately find the specific location of the hot function in the program code. At the same time, mpiPtrace decodes the C++ function name (restoring the function name modified by the compiler) based on the feature that C++ will modify the function name so that users can query data information. These three parts of data ensure the accuracy of subsequent performance analysis and provide important data support for subsequent performance analysis work.

[0034] In order to ensure that the data is complete enough to accurately locate the performance bottleneck of large-scale parallel programs. This embodiment records the necessary information of each call, such as function number, function name, communication direction, time information, communication calculation time, single communication data size, cumulative execution times, time and path number, etc. However, this direct output method will generate a huge trace data file. Therefore, in order to ensure the accuracy of performance analysis and reduce the trace data output file, this embodiment combines the communication characteristics of the MPI program SPMD (single program multiple data) to compress each communication trace data. Specifically, in step S2 of this embodiment, the processing of each communication trace data includes: S2.1, save the communication trace data to the temporary data key, take the function name and stack address of the MPI function corresponding to the communication trace data as input, and generate a hash value tableindex for the MPI function; S2.2, find the index position index corresponding to the hash value tableindex in the communication trace data structure array table; S2.3, determine whether the index position index is empty. If the index position index is empty, return empty; otherwise, loop through the communication trace data structure array table until the communication trace data pointed to by the pointer ptr of the index position index is found to be equal to the temporary data key, and return the pointer ptr of the index position index; S2.4, if the result returned when judging whether the index position index is empty is empty, it is determined that it is the communication trace data when the communication trace data is called for the first time, and the first data structure (callsite_stats_t) is used to record it in the data header ("head"); otherwise, it is determined that the communication trace data is the communication trace data generated by repeated calls, and the second data structure (mpiP_coll_data) is used to record it in the data body ("body"), and it is connected with the pointer ptr (index→ptr) of the index position index through a linked list.

[0035] In this embodiment, the communication characteristics of the MPI program SPMD (single program multiple data) are combined to compress each communication trace data. The "head" information is output first. The output format is based on the profile output format. In order to quickly perform profile analysis, two types of information, the cumulative execution count and the cumulative execution time, are added between the single communication data size and the path number (same as Figure 5 id, name, etc. in the file correspond one by one), and then output the "body" information. "Body" is compressed data, including the time series information, calculation time and tag number of each MPI function call (used for point-to-point communication matching). Therefore, the trace format output records the detailed information of each MPI function call on the basis of the profile value. Since the compressed data in the trace format lacks certain content, in order to facilitate the subsequent performance analysis and reproduce the communication behavior of the program and ensure its accuracy, the trace output format outputs the performance data with the same communication behavior together, which is very convenient when filling (decompressing) data for the "body". When decompressing, you only need to read the data in the "head" to batch decompress the data in the "body".

[0036] After the trace function is implemented, how to effectively control the overhead caused by the implementation of the trace function becomes a challenge. The design, storage method and query method of the structure will affect the overhead caused by the trace function. In order to minimize this part of the overhead, this embodiment proposes a new communication trace data compression method. The "head" is combined with the "body" in the form of an array to store the "head" (callsite_stats_t structure). The "head" records the communication data at the first call, as well as the cumulative number of executions and the cumulative execution time. The communication data generated by repeated calls is recorded using the "body" (mpiP_coll_data) after compression, and is connected to the "head" in the form of a linked list. Since the hash value is used to locate the position of the "head" in the array, multiple "heads" may be stored in the same position. Compared with the initial use of data storage in a single structure callsite_stats_t, this method significantly compresses the amount of data. For a large amount of data in the trace mode, this embodiment uses a hash search method to query the storage location of the structure to reduce time complexity. In this embodiment, a combination of "head" and "body" is used. The structure callsite_stats_t that stores the profile information is regarded as the "head", and a linked list that stores the compressed information is added to it. The linked list consists of the structure mpiP_coll_data and is regarded as the "body". Only different data is recorded in the "body", such as timestamps, communication calculation time and tag information. When a hash conflict occurs, the "head" linked list is traversed, and its time complexity is O(n), where n is the number of elements in the linked list. When the value of n is extremely small, the upper limit depends only on the number of MPI function references in the source code. However, in order to ensure the orderliness of the output data (that is, the data of the same communication behavior is connected together for output), it is necessary to traverse each element in the linked list corresponding to the hash value to find the most suitable position for insertion. This time complexity is also O(n). Since the MPI function will loop multiple times in a large-scale parallel program, the value of n will increase significantly. In order to ensure the integrity of the information and eliminate redundant information at the same time, mpiPtrace improves the structure search and comparison method, through which mpiPtrace realizes trace data compression. mpiPtrace adds a series of parameters, such as the amount of data and the number of the sending and receiving processes, and connects the function data that are consistent after comparison through a linked list. When it is called for the first time, redundant data such as the function name, communication data, and stack information will be recorded in the "head". After the data information of subsequent repeated calls is compressed, only the time series information and tag number need to be recorded in the "body".

[0037] like Figure 5As shown, the communication trace data in this embodiment includes function number id, function name name, communication direction rank, start time star, execution time dur, communication computing time computing, single communication data size data, cumulative execution times count, cumulative execution time cumulative_time and path number computing, wherein the communication computing time computing is the difference between the start time star of the current MPI function and the end time of the previous MPI function.

[0038] In this embodiment, the first data structure (callsite_stats_t) includes all the numerical fields in the communication trace data, and the second data structure (mpiP_coll_data) includes the function number id, the start time star, the execution time dur, and the communication calculation time computing in the communication trace data. The communication trace data composed of the data header and the data body obtained at a certain moment is as follows: Figure 6 As shown. Every time a new communication trace data occurs, a first data structure (callsite_stats_t) is generated and written into the data header. If there is communication trace data of the same communication behavior later, a second data structure (mpiP_coll_data) is generated and written into the data body. This is repeated and multiple first data structures are eventually formed in the data header, and multiple sets of second data structures (mpiP_coll_data) are included in the data body.

[0039] After completing the collection and compression of trace data, in order to accurately locate the program performance bottleneck, detailed stack information needs to be obtained. Although mpiP supports stack parsing, it is found in actual use that mpiP's analysis of stack information is not comprehensive. In order to solve the above problem, this embodiment also includes obtaining and storing the stack information of the MPI function: S101, using libunwind function to traverse the stack frame to obtain the stack address of the MPI function; S102, using the elfutils function to parse the stack address to obtain stack information and store it, the stack information includes the source file name, function name and line number, and the function name is the original function name obtained after restoring the function name modified by the compiler; and when storing the stack information, it includes first using an array to record all different stack addresses, and recording the corresponding array subscript for each function according to the required stack address, and performing a unified analysis after the program ends, directly searching the stack information after the analysis through the array subscript to ensure that a stack address will only be analyzed once, and writing the storage stack information finally analyzed into the data header. This embodiment introduces the elfutils library to mpiPtrace to perform a more detailed analysis of the stack information. There are a large number of repeated parts in the stack information of the function, and the same stack address will be analyzed multiple times, which results in many repeated analysis operations. In order to eliminate this part of the repeated analysis operation, mpiPtrace first uses an array to write all different stack addresses, and records the corresponding array subscript for each function according to the required stack address. After the program ends, a unified analysis is performed, and the stack information after the analysis is directly searched through the array subscript, so that a stack address will only be analyzed once. At the same time, the stack information will only be recorded in the "head". The experiment was conducted using the NPB program with 256 processes. The results showed that the mpiPtrace data file size was reduced by up to 92.17% compared with Tau, and the time overhead was reduced by about 3% compared with Tau.

[0040] like Figure 7As shown, when the communication trace data recorded in the data header and the data body is stored in step S3, it includes recording the number of calls of the MPI function and the I / O operation and storing it by setting a storage strategy: when the target parallel program is started, the MPI function and I / O operation to be monitored are specified by modifying the configuration file; a global counter is created for each monitored MPI function and I / O operation according to the specified parameters, and a counter increment operation is inserted at the entrance of each monitored MPI function and I / O operation; during the execution of the target parallel program, the MPI function and I / O operation are detected by the counter increment operation inserted at the entrance of the MPI function and I / O operation, and the number of times it is called is recorded by increasing the global counter, and the current timestamp is recorded while the global counter is updated; the value of the global counter and the time interval between the values ​​of the global counter are regularly checked; if the value of the global counter reaches the counter threshold or the time interval between the values ​​of the global counter exceeds the preset expectation, the storage operation is triggered to write the communication trace data recorded in the data header and the data body into the file system, and the record is cleared to prepare for receiving the communication trace data of the next cycle. By monitoring the number of calls of the MPI function and its time interval, the programmer can determine whether the program has a load imbalance problem. For example, when some processes perform a lot of calculations while other processes perform relatively few calculations, the slower processes may cause other processes to wait for data or synchronization. This unbalanced task distribution may cause some processes to be idle for a long time. In addition, with the help of the number of calls to the MPI function, programmers can better observe whether the program execution status meets the expected design of the algorithm. This process is of great significance for optimizing the performance of the program. At the same time, we also record I / O operations and their time intervals to prevent data loss caused by abnormal interruptions in large-scale parallel programs. In this way, the integrity and reliability of important data can be ensured. This embodiment provides an effective storage mechanism through the above method for recording the number of calls to MPI functions and I / O operations, and by setting storage strategies, optimizes the management of communication trace data, thereby improving the efficiency of performance analysis. This embodiment also provides a user interface that allows users to query the current monitored data statistics in real time through this interface, including the number of calls, time intervals, memory status, etc. of each monitored MPI function and I / O operation. With the help of this interface, users can adjust the monitoring configuration and storage strategy in time according to the execution of the program, which can better help users locate problems.

[0041] like Figure 8As shown, when the communication trace data recorded in the data header and the data body is stored in step S3, it also includes periodically checking the memory size occupied by the communication trace data recorded in the data header and the data body. If the memory size occupied by the communication trace data recorded in the data header and the data body exceeds the preset maximum memory threshold, the storage operation is triggered to write the communication trace data recorded in the data header and the data body into the file system, and the record is cleared to prepare for receiving the communication trace data of the next cycle. This embodiment provides another effective storage mechanism in the above manner. By setting a maximum memory threshold, when the communication trace data accumulates to the threshold, the storage operation is automatically triggered, the current communication trace data is written into the file system, and the communication trace data file is cleared to prepare for receiving the communication trace data of the next cycle, thereby realizing automatic storage according to the file size.

[0042] In summary, the present embodiment obtains communication trace data by overloading the MPI function. The present embodiment proposes a communication trace data format, and rearranges the format of the acquired communication trace data based on the characteristics of the parallel program. The present embodiment uses the formatted communication trace data header as the key value and uses the hash method to compress the communication trace data. The present embodiment also proposes a new storage method, which stores the compressed communication trace data in the file system at an appropriate time according to the communication characteristics and trace data volume at runtime. The method of the present embodiment uses an efficient communication trace data compression and storage method to significantly reduce the data size compared with the previous one while ensuring that the trace data is complete and rich in content. The method of the present embodiment supports real-time storage, ensuring that analysts can quickly obtain the latest performance data, respond and optimize in a timely manner. The method of the present embodiment can better analyze the communication mode and potential performance bottlenecks by recording detailed metadata of each communication event (such as timestamp, sending / receiving process, message type, etc.), thereby providing strong support for optimizing large-scale parallel programs.

[0043] In addition, this embodiment also provides a communication trace data compression and storage system based on parallel program features, including a microprocessor and a memory connected to each other, and the microprocessor is programmed or configured to execute the communication trace data compression and storage method based on parallel program features.

[0044] In addition, this embodiment also provides a computer-readable storage medium, in which a computer program or instruction is stored. The computer program or instruction is programmed or configured to execute the communication trace data compression and storage method based on parallel program features through a processor.

[0045] In addition, this embodiment also provides a computer program product, including a computer program or instructions, which are programmed or configured to execute the communication trace data compression and storage method based on parallel program features through a processor.

[0046] Those skilled in the art should understand that the technical solutions provided by the embodiments of the present invention may be in the form of methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes. The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, may be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that instructions executed by the processor of a computer or other programmable data processing device generate instructions for implementing the functions in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the functions specified in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the process in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0047] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

Claims

1. A communication trace data compression and storage method based on parallel program characteristics, characterized in that: The steps include: S1, intercepts MPI calls by overloading MPI functions for the target parallel program and collects communication trace data; S2, for each piece of communication trace data, determining whether it is the communication trace data of the first call or the communication trace data generated by repeated calls, if it is the communication trace data of the first call, recording it in the data header using a first data structure, if it is the communication trace data generated by repeated calls, recording it in the data body using a second data structure, and the value field in the second data structure is less than the value field in the first data structure; S3, stores the communication trace data recorded in the data header and data body.

2. The communication trace data compression and storage method based on parallel program characteristics according to claim 1 is characterized in that: In step S2, the processing of each communication trace data includes: S2.1, save the communication trace data to the temporary data key, take the function name and stack address of the MPI function corresponding to the communication trace data as input, and generate a hash value tableindex for the MPI function; S2.2, find the index position index corresponding to the hash value tableindex in the communication trace data structure array table; S2.3, determine whether the index position index is empty. If the index position index is empty, return empty; otherwise, loop through the communication trace data structure array table until the communication trace data pointed to by the pointer ptr of the index position index is found to be equal to the temporary data key, and return the pointer ptr of the index position index; S2.4, if the result returned when judging whether the index position index is empty is empty, it is determined that it is the communication trace data when the communication trace data is called for the first time, and the first data structure is used to record it in the data header; otherwise, it is determined that the communication trace data is the communication trace data generated by repeated calls, and the second data structure is used to record it in the data body, and it is connected to the pointer ptr of the index position index through a linked list.

3. The communication trace data compression and storage method based on parallel program characteristics according to claim 2 is characterized in that: The communication trace data includes function number id, function name name, communication direction rank, start time star, execution time dur, communication computing time computing, single communication data size data, cumulative execution times count, cumulative execution time cumulative_time and path number computing, where the communication computing time computing is the difference between the start time star of the current MPI function and the end time of the previous MPI function.

4. The communication trace data compression and storage method based on parallel program characteristics according to claim 3 is characterized in that: The first data structure includes all the numerical fields in the communication trace data, and the second data structure includes the function number id, the start time star, the execution time dur, and the computation time computing during communication in the communication trace data.

5. The communication trace data compression and storage method based on parallel program characteristics according to claim 1 is characterized in that: It also includes obtaining and storing the stack information of the MPI function: S101, using libunwind function to traverse the stack frame to obtain the stack address of the MPI function; S102, using the elfutils function to parse the stack address to obtain stack information and store it, wherein the stack information includes the source file name, function name and line number, and the function name is the original function name obtained by restoring the function name modified by the compiler; and storing the stack information includes first using an array to record all different stack addresses, and recording the corresponding array index for each function according to the required stack address, and performing a unified analysis after the program ends, directly searching the stack information after analysis through the array index to ensure that a stack address is only analyzed once, and writing the storage stack information finally analyzed into the data header.

6. The communication trace data compression and storage method based on parallel program characteristics according to claim 1 is characterized in that: When the communication trace data recorded in the data header and the data body are stored in step S3, the communication trace data is stored by recording the number of calls of the MPI function and the I / O operation and setting a storage strategy: when the target parallel program is started, the MPI function and the I / O operation to be monitored are specified by modifying the configuration file; a global counter is created for each monitored MPI function and I / O operation according to the specified parameters, and a counter increment operation is inserted at the entry of each monitored MPI function and I / O operation; During the execution of the target parallel program, the calling of MPI functions and I / O operations is detected by inserting counter increment operations at the entry of MPI functions and I / O operations, and the number of times they are called is recorded by increasing the global counter, and the current timestamp is recorded while the global counter is updated; the value of the global counter and the time interval between the values ​​of the global counter are checked regularly; If the value of the global counter reaches the counter threshold or the time interval between the values ​​of the global counter exceeds the preset expectation, the storage operation is triggered to write the communication trace data recorded in the data header and data body to the file system, and clear the record to prepare for receiving the communication trace data of the next cycle.

7. The communication trace data compression and storage method based on parallel program characteristics according to claim 1 is characterized in that: When storing the communication trace data recorded in the data header and the data body in step S3, it also includes periodically checking the memory size occupied by the communication trace data recorded in the data header and the data body. If the memory size occupied by the communication trace data recorded in the data header and the data body exceeds the preset maximum memory threshold, the storage operation is triggered to write the communication trace data recorded in the data header and the data body to the file system, and the record is cleared to prepare for receiving the communication trace data of the next cycle.

8. A communication trace data compression and storage system based on parallel program characteristics, comprising a microprocessor and a memory connected to each other, characterized in that: The microprocessor is programmed or configured to execute the communication trace data compression and storage method based on parallel program features as described in any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program or instruction stored therein, characterized in that: The computer program or instruction is programmed or configured to execute the communication trace data compression and storage method based on parallel program features as claimed in any one of claims 1 to 7 through a processor.

10. A computer program product comprising a computer program or instructions, characterized in that The computer program or instruction is programmed or configured to execute the communication trace data compression and storage method based on parallel program features as claimed in any one of claims 1 to 7 through a processor.

Citation Information

Patent Citations

  • Fast information interaction mechanism in multimedia system and network transmission method

    CN107026887A

  • Function call stack analysis and backtracking method and device

    CN115292201A

  • General analysis system and method for radar recording data

    CN116166845A

  • Monitoring stack memory usage to optimize programs

    US20230044935A1

Cited By

  • Iterative sensing MPI communication trace data compression method and system

    CN121455807A

  • Iterative perception-based mpi communication trace data compression method and system

    CN121455807B