Communication Performance Analysis Method and System Based on Multidimensional Features and Improved Clustering
Through the communication performance analysis method based on multi-dimensional features and improved clustering, the problem of lack of refined and targeted optimization in the existing MPI communication analysis methods is solved, and accurate communication pattern recognition and optimization of different HPC application fields is achieved.
Patent Information
- Application Number
- CN202510486669.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The existing MPI communication analysis methods lack refined communication mode analysis, making it difficult to identify the differences in communication modes in different HPC application fields, resulting in a lack of targeted optimization measures and failing to effectively identify MPI communication needs in different fields.
The communication performance analysis method based on multi-dimensional features and improved clustering is adopted. By refining communication feature classification and clustering analysis, the communication mode differences between applications in different fields are identified, the feature weights are dynamically adjusted using the improved clustering algorithm, the multi-dimensional communication feature matrix is constructed, and the adaptive similarity calculation method is combined to optimize the cluster analysis of communication modes.
It improves the accuracy of MPI communication performance analysis and clustering accuracy, can classify communication modes more effectively, provide accurate reference for subsequent optimization, provide customized optimization suggestions, and improve communication efficiency.
Smart Images

Figure CN120017527B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field related to digital data processing, and specifically, to a communication performance analysis method and system based on multi-dimensional features and improved clustering. Background Art
[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] High-Performance Computing (HPC) has been widely applied in many fields such as scientific computing, engineering simulation, weather prediction, and bioinformatics, and has become an important means to solve complex computing problems. In an HPC system, parallel computing is the core method to improve computing efficiency, and efficient data transmission and inter-process communication are the key factors for optimizing the performance of a parallel computing system.
[0004] In a High-Performance Computing (HPC) system, MPI (Message Passing Interface) is a widely used parallel communication standard that provides a complete set of message passing mechanisms for distributed computing, including point-to-point communication, collective communication, and synchronous and asynchronous communication modes to adapt to different parallel computing scenarios. In large-scale HPC applications, the MPI communication overhead usually occupies a significant proportion of the program running time. As the computing scale expands, the communication bottleneck may become the key factor restricting the overall performance improvement. Therefore, the analysis and optimization of MPI communication are crucial. The current MPI communication analysis mainly focuses on communication overhead evaluation, network topology optimization, and load balancing analysis to improve the computing efficiency and resource utilization rate of HPC applications.
[0005] Although the existing MPI communication analysis methods can reveal the overall communication overhead of a program, there are still many deficiencies. First, there is a lack of refined communication mode analysis, mainly focusing on the macroscopic statistical data of MPI communication, such as the total communication volume and communication delay. The single identified feature makes it difficult to accurately identify the features of various communication modes and their impacts on performance. For example, there are significant differences in communication modes, delays, and bandwidth usage between point-to-point communication and collective communication. Traditional analysis methods often cannot accurately identify the specific impacts of these differences on performance, resulting in lack of pertinence in optimization measures. Second, the communication mode differences in different HPC application fields cannot be effectively identified. Although fields such as weather simulation, bioinformatics, and computer graphics may adopt similar parallel computing methods, their MPI communication requirements are different. For example, weather simulation relies on large-scale broadcast communication, while bioinformatics analysis often involves high-frequency point-to-point communication. The existing methods fail to provide customized optimization suggestions for different application fields. Summary of the Invention
[0006] To solve the above problems, the present invention proposes a communication performance analysis method and system based on multi-dimensional features and improved clustering. By refining communication feature classification and clustering analysis, it deeply understands the communication behaviors of different function types, and analyzes the communication mode differences between applications in different fields and identifies abnormal message communication behaviors during application operation through an improved clustering method, so as to provide a scientific basis and effective support for subsequent communication performance optimization.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] One or more embodiments provide a communication performance analysis method based on multi-dimensional features and improved clustering, including the following steps:
[0009] Based on the semantic characteristics of MPI communication functions, divide the communication functions according to communication modes and performance impacts;
[0010] Obtain the communication record information of various MPI communication functions and extract basic communication features;
[0011] Calculate combined communication features based on the basic communication features;
[0012] Based on the basic communication features and combined communication features, construct a communication feature matrix, adopt an adaptive similarity calculation method based on feature variance, assign weights to different features according to the variances of each communication feature, and perform clustering analysis to obtain clustering results;
[0013] Based on the clustering results, identify the performance differences in the communication modes of each application program.
[0014] One or more embodiments provide a communication performance analysis system based on multi-dimensional features and improved clustering, including:
[0015] A division module configured to divide communication functions according to communication modes and performance impacts based on the semantic characteristics of MPI communication functions;
[0016] A feature extraction module configured to obtain the communication record information of various MPI communication functions and extract basic communication features;
[0017] A combined communication feature calculation module configured to calculate combined communication features based on the basic communication features;
[0018] A clustering module configured to construct a communication feature matrix based on the basic communication features and combined communication features, adopt an adaptive similarity calculation method based on feature variance, assign weights to different features according to the variances of each communication feature, and perform clustering analysis to obtain clustering results;
[0019] A result recognition module, configured to recognize performance differences in the communication patterns of each application program based on the clustering results.
[0020] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0021] In this embodiment, the improved clustering algorithm calculates the variance of each feature and assigns weights to different features according to the variance, dynamically adjusting the contributions of each feature in the similarity calculation to ensure that during the clustering process, features with larger variances (i.e., features that contribute more to the differences between application programs) occupy a more important position; by constructing a multi-dimensional communication feature matrix and combining an adaptive similarity calculation method based on variance, the clustering analysis of communication patterns is optimized, thereby improving the accuracy of MPI communication performance analysis. Traditional methods mainly rely on static statistical data and fail to deeply mine the detailed features of communication patterns, while the method of this embodiment can dynamically adjust feature weights, highlight important features, and improve the accuracy of clustering. At the same time, the improved clustering algorithm can more effectively classify communication patterns, and thus provide a more accurate reference for communication optimization.
[0022] The advantages of the present invention and the advantages of additional aspects will be described in detail in the following specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute a limitation to the present invention.
[0024] Figure 1 is a flowchart of the communication performance analysis method according to Embodiment 1 of the present invention;
[0025] Figure 2 is a schematic flowchart of the communication performance analysis method according to Embodiment 1 of the present invention;
[0026] Figure 3 is a block diagram of the structure of the communication performance analysis system according to Embodiment 2 of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] The present invention will be further described below in conjunction with the drawings and embodiments.
[0028] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0029] Note that the terms used here are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof. It should be noted that, without conflict, the various embodiments in the present invention and the features in the embodiments can be combined with each other. The embodiments will be described in detail below with reference to the drawings.
[0030] Technical term explanation:
[0031] MPI communication function: Refers to the communication operation functions provided by MPI (Message Passing Interface), such as MPI_Send, MPI_Recv, MPI_Bcast, etc., which are the basic components for data transmission in application programs.
[0032] Application program: Refers to a program that uses MPI for parallel computing, such as applications in meteorological simulation, bioinformatics analysis, computer graphics, etc. Different application programs will call different MPI communication functions, showing different communication modes and performance characteristics.
[0033] Each application program will call multiple MPI communication functions for data transmission. The usage of these communication functions, such as call frequency, message size, communication mode, etc., determines the communication characteristics of the application program. The analysis object of the present invention is the "application program", but the feature extraction is based on the "MPI communication function": The present invention collects the call records of MPI communication functions from the execution process of the application program, extracts the basic communication features, calculates the combined features based on the basic features of the MPI communication function, to more comprehensively describe the communication mode of the application program. Construct a communication feature matrix, improve the clustering method based on the variance of the features, input the features of the application program into the clustering analysis, so as to identify the communication modes of different application programs, and can provide guiding suggestions for optimizing MPI communication. The following will be analyzed with specific embodiments.
[0034] Embodiment 1
[0035] In the technical solutions disclosed in one or more embodiments, as Figures 1 to 2 shown, the communication performance analysis method based on multi-dimensional features and improved clustering includes the following steps:
[0036] Step 1: Based on the semantic characteristics of MPI communication functions, divide the communication functions according to the communication mode and performance impact;
[0037] Step 2: Obtain the communication record information of various MPI communication functions and extract the basic communication features;
[0038] Step 3: Calculate combined communication features based on basic communication features;
[0039] Step 4: Based on the basic communication features and the combined communication features, construct a communication feature matrix, adopt an adaptive similarity calculation method based on feature variance, assign weights to different features according to the variance of each communication feature, and perform clustering analysis to obtain a clustering result;
[0040] Step 5: Based on the clustering result, identify the performance differences in the communication patterns of each application program;
[0041] Based on the clustering result, the performance differences in the communication patterns can be identified. By analyzing the differences between different clustering categories, the advantages and disadvantages of different MPI communication patterns in terms of performance can be identified, such as calculating the communication overlap degree, communication bottleneck, message transmission delay, etc. According to the analysis results, guiding suggestions for optimizing MPI communication can be provided, such as adjusting the communication strategy, optimizing the load balance, etc.
[0042] In this embodiment, the improved clustering algorithm calculates the variance of each feature, assigns weights to different features according to the variance, and dynamically adjusts the contribution of each feature in the similarity calculation to ensure that during the clustering process, features with larger variances (i.e., features that contribute more to the differences between application programs) occupy a more important position; by constructing a multi-dimensional communication feature matrix and combining the adaptive similarity calculation method based on variance, the clustering analysis of the communication pattern is optimized, thereby improving the accuracy of MPI communication performance analysis. Traditional methods mainly rely on static statistical data and fail to deeply explore the detailed features of the communication pattern, while the method of this embodiment can dynamically adjust the feature weights, highlight important features, and improve the accuracy of clustering. At the same time, the improved clustering algorithm can more effectively classify the communication patterns, and thus provide more accurate references for communication optimization.
[0043] In Step 1, classify the MPI message communication functions: Different types of message communication functions exhibit different performance characteristics. Based on the semantic characteristics of the MPI communication functions, the communication functions are classified into four categories: point-to-point blocking communication, point-to-point non-blocking communication, one-to-many collective communication, and many-to-many collective communication according to the behavior pattern and performance impact;
[0044] Specifically, point-to-point communication includes:
[0045] Point-to-point blocking communication: When data is transmitted from the sender to the receiver, the sending operation will block until the data is completely received;
[0046] Point-to-point non-blocking communication: The sending operation does not wait for the data to be completely received, but returns immediately, allowing the sender to perform other operations;
[0047] Specifically, collective communication includes:
[0048] One-to-many collective communication: One process sends data to multiple processes, often used in broadcast or multicast operations;
[0049] Many-to-many collective communication: Data exchange occurs among multiple processes, commonly seen in allgather or alltoall operations.
[0050] As shown in Table 1 below, the following table lists the typical communication types for each category and their corresponding MPI function examples:
[0051] Table 1 Typical communication types of communication functions and their corresponding MPI function examples;
[0052]
[0053] In this embodiment, the MPI communication functions in the high-performance computing application program are classified into four categories according to the communication mode, where correspond to different communication types. Each type contains communications, and the corresponding communication functions are respectively. Each communication ( ) is a communication instance under the corresponding communication type;
[0054] In step 2, message communication record collection: For each communication type, the record data of MPI communication functions are collected through a performance analysis tool, including message size, call count, communication time, and the proportion of communication time, and the data is cleaned and processed;
[0055] To collect the communication data of high-performance computing applications, a supercomputing environment based on MPI is selected as the experimental platform. On this platform, multiple widely used high-performance computing (HPC) applications are run, covering multiple scientific fields such as physics, chemistry, oceanography, and biology, ensuring the diversity and comprehensiveness of the data. In this embodiment, the existing performance analysis tool is used to record the message communication records of each application program. During the communication process of the application program, n MPI communication functions are involved, which are respectively represented as ( ). For each MPI communication function (where ), the following key features are collected:
[0056] Message size: Represents the size of the data transmitted during the communication process, in bytes (Byte); The message size is related to each MPI communication function ;
[0057] Call count: Represents the MPI function The number of calls during communication, i.e., the number of times the function is called during one execution of the application.
[0058] The communication time of the MPI communication function: represents each MPI communication function 's execution time.
[0059] The proportion of communication time: represents each MPI communication function 's proportion of communication time in the overall application execution time.
[0060] Based on the collected message communication records, construct basic communication features. For each type of communication function ( ), there are the following four basic communication features. Summarize the features of all communication instances in the type to obtain the basic communication features of the corresponding communication type:
[0061] (1) Message size : The total message size transmitted by the th type of communication function, representing the total amount of communication data transmitted by this type of communication:
[0062] ;
[0063] Among them, represents the message size of the th communication instance, is the total number of communication instances under the th type of communication function.
[0064] (2) Number of calls : The total number of times the th communication function is called in the application:
[0065] ;
[0066] (3) Communication time : The total communication time of the th type of communication function, representing the sum of the execution times of all communication functions of this type:
[0067] ;
[0068] (4) Communication proportion : The proportion of the total communication time of the th type of communication function in the total execution time of the application:
[0069] ;
[0070] In step 3, combined communication features are calculated based on basic communication features. By combining multiple basic communication features, multiple combined communication features are obtained.
[0071] The combined communication features can further characterize the complex performance of communication functions, and can reveal the characteristics of communication behaviors from multiple perspectives such as communication efficiency, communication intensity, and communication load. In particular, potential problems that are difficult to discover by traditional analysis methods. These features can help developers deeply understand communication performance and provide a basis for targeted optimization.
[0072] Specifically, in this embodiment, for each function type Three combined communication features are designed, which are obtained by combining and calculating multiple basic communication features, including communication efficiency, communication intensity, and communication load, as follows:
[0073] (5) Communication efficiency: The product of the number of times each function type is called per unit time and the message size is used as the communication efficiency. The communication efficiency of the th function type
[0074] ;
[0075] The communication efficiency constructed in this embodiment is the core feature for evaluating network bandwidth utilization, and measures the amount of data transmitted per unit time. A higher communication efficiency indicates that the bandwidth is well utilized, while a lower efficiency may imply bandwidth waste or network bottlenecks.
[0076] (6) Communication intensity: The number of times each function type is called per unit time is used as the communication intensity. The communication intensity of the th function type
[0077] ;
[0078] The communication intensity constructed in this embodiment reflects the frequency of communication operations. A higher communication intensity may mean that the communication operations in the application are frequent and short, which may lead to frequent communication delays and performance losses.
[0079] (7) Communication load: The contribution of each communication function type to the total communication time per average call is used as the communication load. The communication load of the th communication function type
[0080] ;
[0081] The communication load constructed in this embodiment helps to identify the proportion of each communication function call in the total communication time. A higher communication load means that a single communication has a greater impact on the total communication time of the program and may be the root cause of performance bottlenecks.
[0082] In step 3, for the construction of the communication feature matrix, specifically, taking the basic features and combined features obtained for each communication function type as columns respectively, with each row corresponding to the features of a communication function type, and setting the rows of unused function types to zero, the communication feature matrix is obtained;
[0083] The feature matrix of each application program will contain four types of communication functions, and each type of communication function has four basic communication features and three combined communication features. The structure of the feature matrix is as follows:
[0084] ;
[0085] Among them, the 1st to 4th rows of the feature matrix respectively represent four types of communication functions, namely: point-to-point blocking, point-to-point non-blocking, one-to-many collective communication, and many-to-many collective communication. Each row contains four basic communication features and three combined communication features of this communication function type. For unused communication function types, all the feature values of this row will be filled with 0 values.
[0086] For example, if a certain application program only uses point-to-point blocking communication and point-to-point non-blocking communication, and does not use one-to-many collective communication and many-to-many collective communication, then the corresponding feature matrix is as follows:
[0087] ;
[0088] The clustering method based on multi-dimensional communication features proposed in this embodiment can accurately identify and classify the communication behaviors of MPI application programs. This method combines multiple communication dimensions (function type, feature design) for clustering analysis, has strong practical application value, and can especially help to identify application programs with similar communication patterns, further supporting the optimization of high-performance computing systems.
[0089] In step 4, an adaptive similarity calculation method based on feature variance is adopted. Weights are assigned to different features according to the variances of the features, and then clustering analysis is carried out to obtain the clustering result. The spectral clustering method based on adaptive similarity can be used to improve this clustering method through feature variance;
[0090] The spectral clustering method is an algorithm based on graph theory. It clusters data by constructing a similarity matrix and performing eigen decomposition, and can effectively capture the potential similarities of different applications. Compared with traditional methods, spectral clustering can better handle non-linear relationships and adapt to high-dimensional data, enabling accurate communication behavior analysis. Traditional spectral clustering methods usually use fixed similarity metrics, such as Euclidean distance or cosine similarity, which are not flexible enough for the feature differences of high-dimensional data. Based on this, this embodiment proposes an adaptive similarity calculation method based on feature variance to dynamically adjust the weights in similarity calculation.
[0091] Furthermore, an improved spectral clustering method based on feature variance is adopted. The method of assigning weights to different features according to the variance of each feature and performing clustering analysis on communication patterns to obtain clustering results includes the following steps:
[0092] Step 41, construction of the similarity matrix: An adaptive similarity calculation method based on feature variance is adopted to assign variance-weighted weights to different communication features, calculate the weighted Euclidean distance metric to measure the similarity between applications using the weights, and finally construct the similarity matrix through Gaussian kernel function transformation;
[0093] Among them, different communication features include basic communication features and combined communication features;
[0094] Similarity matrix is the core of spectral clustering, which represents the similarity between different applications. To better reflect the differences between applications, a weight coefficient is introduced for each feature. This coefficient is proportional to the variance of the feature. This method can dynamically adjust the similarity calculation according to the degree of change of each feature in the application, making the features with larger variances have a greater impact on the clustering results.
[0095] Based on feature variance, adaptively calculate the similarity between programs and construct the similarity matrix, including the following steps:
[0096] Step 41.1, calculate the variance of each MPI communication feature in the application, and use the variance to determine the importance of the feature, so as to adjust the contribution degree of each feature when calculating the similarity.
[0097] For each feature of each communication function , calculate its variance in all applications:
[0098] ;
[0099] Among them, represents the value of the th application on the th feature, is the The mean value of a feature is the total number of applications. , that is, 7 features in each function type are covered, and there are 4 function types in total.
[0100] Step 41.2: Standardize the calculated variance to obtain the weight of each communication feature:
[0101] After calculating the variance of each feature, standardize the variance through normalization to obtain the weight of each feature , as follows:
[0102] ;
[0103] Among them, the weight represents the feature 's importance in clustering.
[0104] Step 41.3: Based on the weight of each communication feature obtained, perform weighted calculation to obtain the Euclidean distance between applications:
[0105] In the traditional spectral clustering method, the Euclidean distance is usually used to measure the similarity between two data points. The Euclidean distance calculation formula is:
[0106] ;
[0107] Among them, and are the feature vectors of the p-th application and the q-th application respectively, is the dimension of the feature. This method assumes that all features contribute equally to the similarity.
[0108] In this embodiment, the previously calculated is used to dynamically adjust the contribution of each feature to the similarity calculation. Based on the weight assigned to each feature, calculate the weighted Euclidean distance, and the formula is as follows:
[0109] ;
[0110] Among them, represents the value of the p-th application on the k-th feature, then represents the value of the q-th application on the k-th feature; k is the index of the feature, indicating that there are a total of m features. The in the formula is the weight of the k-th feature, indicating the contribution size of each feature when calculating the weighted Euclidean distance.
[0111] The adaptive similarity measurement method of this embodiment dynamically adjusts the weights of features according to the variance of each feature, so that features with larger variances contribute more to the final similarity. Clustering algorithm for adaptive similarity based on feature variance: By introducing a weight factor and replacing the traditional Euclidean distance calculation, this method can dynamically adjust the similarity measurement according to the actual importance of features, thereby improving the clustering accuracy and precision. In a high-dimensional, multi-feature data environment, this method can better handle complex data sets and improve the effectiveness and stability of clustering results.
[0112] Step 41.4: Use the Gaussian kernel function to convert the weighted Euclidean distance into a similarity matrix to adjust the numerical range for suitability in clustering analysis;
[0113] Use the Gaussian kernel function to convert the weighted Euclidean distance into a similarity matrix :
[0114] ;
[0115] Among them, , representing the standard deviation of the Euclidean distances between all pairs of samples, is used to control the attenuation rate of similarity; the similarity matrix 𝑊 reflects the similarity between different applications.
[0116] In the above solution of this embodiment, the similarity calculation can dynamically adapt to the contribution degrees of different communication features, ensuring that features with higher discrimination for applications play a more important role in clustering analysis, thereby improving the accuracy of communication mode classification.
[0117] Step 42: Construct a degree matrix based on the constructed similarity matrix: Sum each row of the similarity matrix to obtain the total similarity between this application and other applications, and fill the calculated total similarity on the diagonal to obtain a diagonal matrix as the degree matrix;
[0118] The degree matrix 𝐷 represents the connection strength between each application and other applications. It is a diagonal matrix, where each diagonal element is the sum of the similarities between the p-th application and other applications:
[0119] ;
[0120] Among them, represents the weighted connection strength between the p-th application and the q-th application, and n is the total number of applications.
[0121] Step 43: Subtract the degree matrix D from the similarity matrix to obtain the Laplacian matrix;
[0122] The Laplacian matrix is the core matrix of spectral clustering and is used to reflect the graph structure of the data. In spectral clustering, the Laplacian matrix captures the structure of the data through the difference between the similarity matrix and the degree matrix :
[0123] ;
[0124] Step 44, Eigenvalue decomposition and clustering: Perform eigenvalue decomposition on the Laplacian matrix to obtain the eigenvectors corresponding to the first eigenvalues; Use the K-means clustering algorithm to cluster these eigenvectors and divide the applications into clusters;
[0125] In this embodiment, the eigenvectors after eigenvalue decomposition are used as low-dimensional representations, which can highlight the similarities between different applications. K-means clustering in the low-dimensional feature space can more accurately distinguish different communication patterns, thereby improving the accuracy of analysis.
[0126] The purpose of analyzing the clustering results in this embodiment is to identify and distinguish the differences in communication patterns among different applications and provide a basis for optimizing the MPI communication strategy. Through clustering analysis, the applications are grouped according to their communication characteristics. The applications in each group (cluster) exhibit similar communication patterns. Identifying these differences can determine the communication requirements of each application, including communication frequency, bandwidth requirements, latency requirements, etc. Through this process, not only can the differences between different applications be revealed, but also a targeted customization plan can be provided for the optimization strategy, thereby improving the overall communication efficiency and system performance.
[0127] The process of clustering analysis is essentially to classify applications according to their communication characteristics, and each cluster represents a class of applications with similar communication characteristics.
[0128] Furthermore, the association relationship between the clustering results and the application fields can be established in the following ways:
[0129] 1) Feature matching: Through clustering analysis, applications with similar communication characteristics are grouped into the same cluster, and then according to the characteristics of each application field, the application fields to which the applications included in each cluster belong are analyzed;
[0130] Specifically, in the field of weather simulation, applications usually exhibit the characteristics of high bandwidth requirements and large-scale broadcast communication; while in the field of molecular dynamics, applications may exhibit the characteristics of low communication load and frequent local communication. In this way, clustering analysis identifies the commonalities in communication patterns among applications in different fields.
[0131] 2) Combination of domain and communication mode: Within each cluster, based on the communication mode characteristics of the applications in the cluster and the actual requirements of the application domain, match the application domain with the applications within the cluster;
[0132] For example, in clusters with high bandwidth requirements, applications in the fields of engineering and materials may be involved, and the main requirement of these applications is the efficient utilization of bandwidth;
[0133] In this embodiment, through the above construction method of the association relationship, the communication characteristics of the applications within the cluster can be closely combined with the requirements of their respective domains, thereby providing customized optimization strategies for the applications in each domain.
[0134] Through the establishment of the above association relationship, the clustering results can not only reveal the communication characteristic differences between applications, but also combine these differences with the actual requirements of the application domain to provide domain-specific optimization suggestions. Through clustering analysis, the problem of differences in communication characteristics among different applications is solved. Originally, due to the differences in communication characteristics, optimizing the MPI communication strategy may lead to poor generalization effects. By performing refined clustering on applications, customized optimization suggestions can be provided according to the characteristics of the applications within the cluster and the domains in the application concentration within the cluster, thereby significantly improving communication efficiency and solving the technical problem that the existing methods fail to effectively identify and optimize the communication requirements of applications in different domains.
[0135] For a further technical solution, for the different clusters obtained from the clustering results, to distinguish the classification categories of the applications, identify different categories of applications, and summarize their communication mode characteristics, so as to provide a basis for optimizing the MPI communication strategy: divide the obtained clusters into specific application domains and define specific optimization directions, as follows:
[0136] (1) Clusters with high communication intensity and low latency: Applications usually rely on intensive computing and frequent communication, manifested as high computing loads and communication frequencies, mainly in the fields of computer science and mathematics. The clustering results reflect the requirements of applications for fast data exchange and low latency.
[0137] (2) Clusters with high bandwidth requirements: Applications process large-scale data sets and have high bandwidth requirements, mainly in the fields of engineering and materials. The optimization of applications should focus on bandwidth utilization and bottleneck reduction.
[0138] (3) Clusters with high bandwidth requirements and high communication loads: Applications usually process a large amount of weather data, have long communication cycles and large amounts of data, mainly in the fields of meteorology and oceanography. The optimization focus of applications should be placed on bandwidth management and optimization.
[0139] (4) Clusters with low communication load and low latency requirements: The application depends on frequent local communication and has low requirements for global communication, mainly in the field of molecular dynamics. The optimization of the application should focus on improving the efficiency of local communication and reducing the overhead of global communication.
[0140] The clustering results demonstrate extensive applicability in these practical applications, verify the obvious differences in communication characteristics among different applications, and indicate the rationality of the clustering analysis results.
[0141] According to the clustering analysis results and the communication characteristics of application programs in different fields, combined with the specific requirements of different fields and the application characteristics within the cluster, communication performance optimization strategies are proposed to optimize the communication efficiency specifically, including the following:
[0142] (51) Optimization for clusters with high bandwidth requirements, perform bandwidth - first scheduling, data compression, or / and parallel transmission;
[0143] By preferentially allocating bandwidth resources, compressing data to reduce the message size, and adopting multi - path transmission technology to meet high - bandwidth requirements
[0144] (52) Optimization strategies for clusters with low latency requirements, including using non - blocking functions to reduce waiting time, asynchronous communication to reduce latency, and network topology optimization to improve routing efficiency.
[0145] (53) Optimization strategies for clusters with high communication load, adjust task partitioning or adopt an efficient communication interface to reduce unnecessary MPI calls, and optimize the protocol to reduce repeated transmission and waiting.
[0146] (54) Optimization strategies for clusters with high computational load, including reasonably scheduling computational tasks to reduce the dependence between computation and communication, and optimizing caching to improve data locality and reduce cross - node communication.
[0147] Among them, the application characteristics within the cluster refer to, in the clustering analysis results, for all application programs within each cluster, summarizing and averaging the communication characteristics of each type of communication function (such as point - to - point blocking communication, point - to - point non - blocking communication, one - to - many collective communication, many - to - many collective communication), which reflects the overall communication behavior of all application programs within the cluster under different types of communication functions;
[0148] For example, for a clustering cluster, the application characteristics within the cluster include:
[0149] Message size: The average message size of all application programs under each type of communication function;
[0150] Call count: The average value of the call counts of all application programs under each type of communication function;
[0151] Communication time ratio: The average value of the communication time ratios of all application programs under each communication function type;
[0152] Communication efficiency, communication intensity, and communication load: The average values of the communication efficiency, communication intensity, and communication load of all application programs under each communication function type;
[0153] In this embodiment, by performing clustering analysis on the communication characteristics of application programs in different fields, the present invention can accurately identify the communication patterns of application programs in each field, such as communication efficiency, communication intensity, communication load, etc. It helps to identify the differences in communication characteristics of applications in different fields and the abnormal message communication behaviors during application operation. This data-driven analysis method enables the optimization scheme to more precisely adapt to different communication requirements. Compared with the traditional unified optimization method, the strategy of this embodiment has higher adaptability and pertinence.
[0154] In addition, by analyzing the typical characteristics within the clusters, customized optimization directions are provided for clusters with different characteristics. For example, for clusters with high bandwidth demand characteristics, the optimization strategy can focus on the reasonable allocation of bandwidth resources; while for clusters with low latency requirements, asynchronous communication should be optimized first and waiting time should be reduced. Through clustering analysis, a clear direction for optimization can be provided, and a favorable reference for the subsequent optimization of resource allocation can be obtained.
[0155] Embodiment 2
[0156] Based on Embodiment 1, a communication performance analysis system based on multi-dimensional features and improved clustering is provided in this embodiment, as Figure 3 shown, including:
[0157] A partitioning module, configured to partition communication functions according to communication modes and performance impacts based on the semantic characteristics of MPI communication functions;
[0158] A feature extraction module, configured to obtain the communication record information of various MPI communication functions and extract basic communication features;
[0159] A combined communication feature calculation module, configured to calculate combined communication features based on basic communication features;
[0160] A clustering module, configured to construct a communication feature matrix based on basic communication features and combined communication features, adopt an adaptive similarity calculation method based on feature variance, assign weights to different features according to the variances of each communication feature, and perform clustering analysis to obtain a clustering result;
[0161] A result identification module, configured to identify the performance differences in the communication patterns of each application program based on the clustering result.
[0162] Further, the combined communication features include communication efficiency, communication intensity, and communication load;
[0163] Take the product of the number of times each function type is called per unit time and the message size as the communication efficiency;
[0164] Take the number of times each function type is called per unit time as the communication intensity;
[0165] Take the contribution of each communication function type to the total communication time per average call as the communication load.
[0166] Further, an improved spectral clustering method based on feature variance is adopted. According to the variances of each feature, weights are assigned to different features, and a method for clustering analysis of communication patterns to obtain clustering results includes the following steps:
[0167] Adopt an adaptive similarity calculation method based on feature variance, calculate weights for different communication features, use the weights to calculate the weighted Euclidean distance to measure the similarity between applications, and construct a similarity matrix through Gaussian kernel function transformation;
[0168] Sum each row of the similarity matrix and fill it on the diagonal to obtain a diagonal matrix as the degree matrix;
[0169] Subtract the degree matrix from the similarity matrix to obtain the Laplacian matrix;
[0170] Perform eigenvalue decomposition on the Laplacian matrix to obtain the eigenvectors corresponding to the eigenvalues, and use the K-means clustering algorithm to cluster these eigenvectors to divide the applications into multiple clusters.
[0171] It should be noted here that each module in this embodiment corresponds one by one to each step in Embodiment 1, and its specific implementation process is the same, so it will not be repeated here.
[0172] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0173] Although the specific implementation manners of the present invention are described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications or deformations that can be made without creative labor on the basis of the technical solutions of the present invention are still within the protection scope of the present invention.
Claims
1. A communication performance analysis method based on multi-dimensional features and improved clustering, characterized in that, It includes the following steps: Based on the semantic characteristics of MPI communication functions, divide the communication functions according to communication modes and performance impacts; Obtain the communication record information of various MPI communication functions and extract basic communication features; The basic communication features include message size, call count, communication time, and communication time ratio; Calculate combined communication features based on the basic communication features; The combined communication features include communication efficiency, communication intensity, and communication load; Take the product of the call count per unit time and the message size of each function type as the communication efficiency; Take the call count per unit time of each function type as the communication intensity; Take the contribution of each call of each communication function type to the total communication time on average as the communication load; Based on the basic communication features and combined communication features, construct a communication feature matrix. Specifically: take the basic features and combined features obtained for each communication function type as columns respectively, with each row corresponding to the features of one communication function type, and set the rows of unused function types to zero to obtain the communication feature matrix; Adopt an adaptive similarity calculation method based on feature variance, assign weights to different features according to the variances of the communication features, and perform clustering analysis to obtain clustering results; Based on the clustering results, identify the performance differences in the communication modes of each application program.
2. The communication performance analysis method based on multi-dimensional features and improved clustering according to claim 1, wherein Divide the communication functions according to behavioral modes and performance impacts into: point-to-point blocking communication, point-to-point non-blocking communication, one-to-many collective communication, and many-to-many collective communication.
3. The communication performance analysis method based on multi-dimensional features and improved clustering according to claim 1, characterized in that: Take the basic features and combined features obtained for each communication function type as columns respectively, with each row corresponding to the features of one communication function type, and set the rows of unused function types to zero to obtain the communication feature matrix.
4. The communication performance analysis method based on multi-dimensional features and improved clustering according to claim 1, characterized in that Adopt a spectral clustering method based on feature variance, assign weights to different features according to the variances of the features, perform clustering analysis on the communication modes, and obtain clustering results, including the following steps: Adopt an adaptive similarity calculation method based on feature variance, calculate weights for different communication features, use the weights to calculate the weighted Euclidean distance to measure the similarity between application programs, and construct a similarity matrix through Gaussian kernel function transformation; Sum each row of the similarity matrix and fill it on the diagonal to obtain a diagonal matrix as the degree matrix; Subtract the degree matrix from the similarity matrix to obtain the Laplacian matrix; Perform eigenvalue decomposition on the Laplacian matrix to obtain the eigenvectors corresponding to the eigenvalues, and use the K-means clustering algorithm to cluster these eigenvectors to divide the application programs into multiple clusters.
5. The communication performance analysis method based on multi-dimensional features and improved clustering according to claim 4, characterized in that, Based on feature variance, adaptively calculate the similarity between programs and construct a similarity matrix, including the following steps: Calculate the variance of each MPI communication feature in the application program; Perform normalization processing on the calculated variances to obtain the weights of each communication feature; Based on the weights of each obtained communication feature, perform weighted calculation to obtain the Euclidean distance between application programs; Use the Gaussian kernel function to convert the weighted Euclidean distance into a similarity matrix.
6. A communication performance analysis system based on multi-dimensional features and improved clustering, characterized in that It includes: A division module configured to divide the communication functions according to communication modes and performance impacts based on the semantic characteristics of MPI communication functions; A feature extraction module, configured to obtain communication record information of various MPI communication functions and extract basic communication features; The basic communication features include message size, call count, communication time, and communication time ratio; A combined communication feature calculation module, configured to calculate combined communication features based on the basic communication features; the combined communication features include communication efficiency, communication intensity, and communication load; Taking the product of the call count per unit time and the message size of each function type as the communication efficiency; Taking the call count per unit time of each function type as the communication intensity; Taking the contribution of each call of each communication function type to the total communication time on average as the communication load; A clustering module, configured to construct a communication feature matrix based on the basic communication features and the combined communication features. Specifically: using the basic features and combined features obtained for each communication function type as columns respectively, with each row corresponding to the features of one communication function type, and setting the rows of unused function types to zero to obtain the communication feature matrix; adopting an adaptive similarity calculation method based on feature variance, assigning weights to different features according to the variances of the communication features, and performing clustering analysis to obtain the clustering result; A result identification module, configured to identify the performance differences in the communication patterns of each application program based on the clustering result.
7. The communication performance analysis system based on multi-dimensional features and improved clustering according to claim 6, wherein The combined communication features include communication efficiency, communication efficiency, and communication load; Taking the product of the call count per unit time and the message size of each function type as the communication efficiency; Taking the call count per unit time of each function type as the communication intensity; Taking the contribution of each call of each communication function type to the total communication time on average as the communication load.
8. The communication performance analysis system based on multi-dimensional features and improved clustering according to claim 6, wherein, A method for performing clustering analysis on communication patterns by using a spectral clustering method based on feature variance, assigning weights to different features according to the variances of the features, and obtaining the clustering result, includes the following steps: Adopting an adaptive similarity calculation method based on feature variance, calculating weights for different communication features, using the weights to calculate the weighted Euclidean distance to measure the similarity between application programs, and constructing a similarity matrix through Gaussian kernel function transformation; Summing each row of the similarity matrix and filling it on the diagonal to obtain a diagonal matrix as the degree matrix; Subtracting the degree matrix from the similarity matrix to obtain the Laplacian matrix; Performing eigenvalue decomposition on the Laplacian matrix to obtain the eigenvectors corresponding to the eigenvalues, and using the K-means clustering algorithm to cluster these eigenvectors to divide the application programs into multiple clusters.
Citation Information
Patent Citations
MPI (message passing interface) parallel program load problem three-dimensional visualized analysis method suitable for large-scale cluster
CN103019852A
Customer segmentation method and device based on cluster analysis
CN108734217A