HPC application program communication time prediction method and system based on machine learning

Through machine learning-based methods, MPI functions in HPC applications are classified and feature extracted, and communication time is predicted using the optimal machine learning model, which solves the problem that traditional methods are difficult to cope with complex environments, and achieves efficient and accurate communication performance prediction.

CN120104446APending Publication Date: 2025-06-06QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510172398.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-17
Filing Date
2025-02-17
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Traditional MPI communication performance prediction methods are difficult to cope with complex communication patterns and dynamically changing system environments, and lack detailed model measurement methods to explain it, resulting in application difficulties.

Method used

Using a machine learning-based method, we obtain the MPI functions in HPC applications, classify them according to the communication mode, and select the optimal machine learning fusion model (such as XGBoost and ANN models) for training, extract communication features to predict communication time.

Benefits of technology

It realizes accurate prediction of the time-consuming MPI communication in HPC applications, adapts to complex and changeable communication modes and system environments, reduces the requirements for researchers' domain knowledge, and improves the applicability and practicality of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104446A_ABST
    Figure CN120104446A_ABST
Patent Text Reader

Abstract

The invention provides an HPC application program communication time prediction method and system based on machine learning, and the method comprises the steps: obtaining MPI functions in an HPC application program, and classifying the MPI functions according to a communication mode; selecting an optimal machine learning fusion model according to the category of the MPI function; extracting communication features of the MPI function, and inputting the communication features to a trained optimal machine learning fusion model to obtain predicted communication time of the MPI function; and summing the predicted communication time of all MPI functions to obtain the predicted communication time of the HPC application program. According to the method, the performance model of the parallel application program is established by utilizing the comprehensive performance data, and the communication performance model is established by adopting a machine learning technology, so that the time consumption of MPI communication in the parallel application program can be accurately predicted. The method has good practicability and universality, can be universally suitable for high-performance computing application, and provides effective support for optimizing resource allocation and improving system efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of program performance prediction, and in particular to a method and system for predicting communication time of an HPC application program based on machine learning. Background Art

[0002] As the demand for computing power in scientific research and industrial applications continues to grow, high-performance computing (HPC) has gradually become a key technology to promote the progress of modern science and technology. Large-scale parallel computing can handle complex scientific computing, engineering simulation, and big data analysis tasks in a short time. The efficient operation of HPC applications depends on the efficiency of their underlying communication, among which the message passing interface (MPI) plays a vital role as a common parallel programming model in HPC systems. Modeling communication overhead in computer clusters is a critical and challenging issue. This study provides important insights for optimizing the deployment of parallel scientific applications in HPC. However, with the increasing complexity of network computing nodes, communication performance modeling faces new challenges. Therefore, how to accurately predict and optimize MPI communication performance has become a research hotspot in the HPC field.

[0003] Traditional MPI communication performance prediction methods usually rely on manual modeling and empirical analysis, which not only requires researchers to have rich domain knowledge, but also may have difficulty coping with the growth of data volume and the increase of system complexity when facing complex communication patterns and dynamically changing system environments. Although some studies have effectively demonstrated their models, most of the work focuses on theoretical analysis and lacks a detailed explanation of the model measurement methods. Users often need to infer the measurement methods of model parameters by themselves, which makes it difficult to apply these models in many cases. There are also challenges with replay-based models, which reconstruct application behavior from historical execution traces to predict performance. These methods require a lot of storage space to save traces, limiting their wider applicability. Summary of the invention

[0004] In order to solve the above problems, the present invention proposes a method and system for predicting the communication time of HPC applications based on machine learning, which uses comprehensive performance data to establish a performance model of parallel applications. By using machine learning technology to build a communication performance model, the time consumption of MPI communication in parallel applications can be accurately predicted. This method has good practicality and universality, can be widely applied to high-performance computing applications, and provides effective support for optimizing resource allocation and improving system efficiency.

[0005] In order to achieve the above object, the present invention adopts the following technical solution:

[0006] In a first aspect, the present invention provides a method for predicting communication time of an HPC application based on machine learning, comprising:

[0007] Obtain MPI functions in HPC applications and classify MPI functions based on communication patterns;

[0008] According to the category of the MPI function, an optimal machine learning fusion model is selected; the communication characteristics of the MPI function are extracted and input into the trained optimal machine learning fusion model to obtain the predicted communication time of the MPI function;

[0009] The predicted communication time of all MPI functions is summed to obtain the predicted communication time of the HPC application.

[0010] Preferably, the MPI functions include a message sending function, a message receiving function, a data reduction function, a broadcast function, a collection function, a global reduction function, a global collection function and a full interchange function.

[0011] Preferably, the MPI functions are classified according to communication modes, and the types include point-to-point blocking communication, point-to-point non-blocking communication, one-to-many collective communication and many-to-many collective communication.

[0012] Preferably, the communication characteristics include the number of processes, message size and communication frequency.

[0013] Preferably, selecting the optimal machine learning fusion model according to the category of the MPI function specifically includes:

[0014] The optimal machine learning fusion model includes an optimal XGBoost model and an optimal ANN model; the training process of the optimal XGBoost model and the optimal ANN model is:

[0015] The historical communication characteristics and historical communication time of each type of MPI function are obtained to construct the training set and validation set;

[0016] Set XGBoost model hyperparameters, perform parameter tuning based on grid search and cross-validation, use the training set to train the model, use the validation set to validate the various evaluation indicators of the model, and save the model as the optimal parameter model if it meets the standards, as the optimal XGBoost model suitable for different types of MPI functions;

[0017] Set ANN model hyperparameters, perform parameter tuning based on random search and early stopping strategies, use the training set to train the model, use the validation set to validate the various evaluation indicators of the model, and save the optimal parameter model if it meets the standards, as the optimal ANN model suitable for different types of MPI functions.

[0018] Preferably, the extracting the communication features of the MPI function and inputting them into a trained optimal machine learning fusion model to obtain the predicted communication time of the MPI function specifically includes:

[0019] The communication features are respectively input into the trained optimal XGBoost model and the optimal ANN model to obtain a first predicted communication time and a second predicted communication time;

[0020] The first predicted communication time and the second predicted communication time are weightedly fused to obtain the predicted communication time of the MPI function.

[0021] Preferably, the weights of the two models in the weighted fusion are obtained by the mean absolute error obtained through model training; the weights are determined in inverse proportion to the mean absolute error of the models to ensure that the model with a smaller mean absolute error occupies a larger weight in the fusion.

[0022] In a second aspect, the present invention provides a HPC application communication time prediction system based on machine learning, comprising:

[0023] A type classification module is used to obtain MPI functions in HPC applications and classify MPI functions according to communication modes;

[0024] A function communication time prediction module is used to select an optimal machine learning fusion model according to the category of the MPI function; extract the communication characteristics of the MPI function, input them into the trained optimal machine learning fusion model, and obtain the predicted communication time of the MPI function;

[0025] The program communication time prediction module is used to sum the predicted communication time of all MPI functions to obtain the predicted communication time of the HPC application.

[0026] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in a method for predicting communication time of an HPC application based on machine learning described in the first aspect.

[0027] In a fourth aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the method for predicting communication time of an HPC application based on machine learning described in the first aspect are implemented.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] In view of the problems that traditional manual modeling and empirical analysis methods have high requirements on researchers' domain knowledge and are difficult to cope with complex communication modes and dynamic system environments, the present invention adopts machine learning technology to automatically extract communication features from MPI functions, such as the number of processes, message size, and communication frequency, and uses machine learning algorithms to learn and recognize patterns of large amounts of data. This does not require researchers to have extremely rich manual modeling experience and deep domain knowledge, and can adapt to complex and changeable communication modes and system environments, and effectively deal with the challenges of growing data volume and increasing system complexity.

[0030] In view of the lack of detailed explanation of model measurement methods in traditional research, users need to infer the model parameter measurement method by themselves, which leads to application difficulties. This invention provides in detail the construction process of the optimal machine learning fusion model (including the optimal XGBoost model and the optimal ANN model). From obtaining the historical communication characteristics and historical communication time of each type of MPI function to construct training sets and validation sets, to parameter tuning based on grid search and cross-validation (XGBoost model) and random search and early stopping strategy (ANN model), and clearly giving the specific steps of model training, verification and determination of the optimal parameter model, users do not need to infer the model parameter measurement method by themselves, which greatly improves the applicability of the model.

[0031] Compared with the limitation that the replay-based model requires a large amount of storage space to save tracking traces, the present invention only extracts the communication characteristics of the MPI function for time prediction, without the need to store a large amount of historical execution trajectories, greatly reducing the demand for storage space, thereby improving the applicability and practicality of the model.

[0032] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their description are used to explain the present invention but do not constitute a limitation of the present invention.

[0034] Figure 1 A main flow chart of a method for predicting communication time of HPC applications based on machine learning provided by an embodiment of the present invention;

[0035] Figure 2 A detailed flow chart of a method for predicting communication time of HPC applications based on machine learning provided by an embodiment of the present invention;

[0036] Figure 3 This is a graph showing experimental data provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0037] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0038] The present invention collects data on MPI communication functions in high-performance computing applications through performance analysis tools, extracts features including the number of processes, message size, and communication frequency, and uses these features as input variables for the machine learning model. The execution time t of each MPI communication function is used as the prediction target of the model. By learning these features, the model can accurately predict the execution time of the communication function, thereby improving the ability to analyze the communication performance of high-performance computing applications. The collected data includes:

[0039] (1) MPI functions involved in the communication process (f);

[0040] (2) The number of processes (p), that is, the number of parallel processes involved in each communication operation;

[0041] (3) Message size (m), which indicates the amount of data transmitted during each MPI function communication process, usually in bytes;

[0042] (4) Communication frequency (k), that is, the number of calls to each MPI function;

[0043] (5) Communication time of MPI function (t), i.e., the execution time of each MPI function;

[0044] (6) Total communication time (T), that is, the total communication time consumed in the application.

[0045] The number of processes, message size, and communication frequency are extracted as key features to form the feature vector X of each MPI communication function. i =[p i ,m i ,k i ], where i = 1, 2, ..., n, and each feature vector corresponds to a communication function in the application. At the same time, for each MPI communication function, its actual execution time t is also recorded. i , which represents the time function f i The total time required to complete a communication operation. Execution time t i As the prediction target of the machine learning model, that is, through the model to the feature vector X i The learning of will output the predicted execution time of the communication function Therefore, the training goal of the model is to pass the feature X i =[p i ,m i ,k i ] to predict the execution time t of the communication functioni , to achieve accurate prediction of MPI communication time in high-performance computing applications.

[0046] Embodiment 1

[0047] like Figure 1 As shown, this embodiment discloses a method for predicting communication time of HPC applications based on machine learning, comprising the following steps:

[0048] S1: Obtain MPI functions in HPC applications and classify MPI functions according to communication modes;

[0049] S2: Selecting an optimal machine learning fusion model according to the category of the MPI function; extracting the communication features of the MPI function, inputting them into the trained optimal machine learning fusion model, and obtaining the predicted communication time of the MPI function;

[0050] S3: Sum the predicted communication time of all MPI functions to obtain the predicted communication time of the HPC application.

[0051] Next, combine Figure 2 , a method for predicting HPC application communication time based on machine learning disclosed in this embodiment is described in detail.

[0052] In S1, this embodiment processes MPI communication data in a high performance computing application.

[0053] First, according to the communication mode, the n MPI communication functions are divided into four categories: point-to-point blocking communication, point-to-point non-blocking communication, one-to-many collective communication, and many-to-many collective communication. For each category, assume that the category type j (i.e. j = 1, 2, 3, 4, corresponding to four categories respectively) MPI communication functions.

[0054] Specifically, the n MPI communication functions in high-performance computing applications are divided into four categories according to the communication mode:

[0055] (1) Point-to-point blocking communication: denoted as type 1 , including Communication functions;

[0056] (2) Point-to-point non-blocking communication: denoted as type 2 , including Communication functions;

[0057] (3) One-to-many collective communication: denoted as category type 3 , including Communication functions;

[0058] (4) Many-to-many collection communication: recorded as category type 4 , including A communication function.

[0059] Therefore, the total number of communication functions in the whole system is:

[0060]

[0061] Each communication function in each class has its own characteristics and execution time. j The i-th communication function f in i,j (i represents the function index, j represents the communication type), its feature vector can be expressed as X i,j =[p i,j ,m i,j ,k i,j ], where m i,j is the message size, k i,j is the communication frequency, p i,j is the number of communicating processes.

[0062] To better illustrate the classification of these communication functions and their roles in practical applications, the following Table 1 lists typical communication types and their corresponding MPI function examples for each category:

[0063] Table 1 Typical communication types and their corresponding MPI functions for each communication category

[0064]

[0065] Among them, MPI_Send is the message sending function, MPI_Recv is the message receiving function, MPI_Isend is the non-blocking message sending function, MPI_Irecv is the non-blocking message receiving function, MPI_Bcast is the broadcast function, MPI_Reduce is the reduction function, MPI_Gathe is the collection function, MPI_Allreduce is the global reduction function, MPI_Allgather is the global collection function, and MPI_Alltoall is the full exchange function.

[0066] The classification of these communication functions reflects the use of different communication modes in high-performance computing applications. By classifying the communication functions, a more accurate prediction model can be established for different types of communication functions, thus improving the accuracy of communication time prediction.

[0067] In S2, extract each communication function f i,j The eigenvector X i,j =[p i,j ,m i,j ,ki,j ], including the number of processes p i,j , message size m i,j and the communication frequency k i,j Execution time t i,j is the target variable of each type of communication function, through the feature vector X i,j The machine learning model will predict the execution time of the communication function.

[0068] In function-level modeling, this embodiment uses different machine learning models to model different types of MPI communication functions.

[0069] This embodiment uses XGBoost and ANN (artificial neural network) models for training, and performs parameter tuning and optimization on XGBoost and ANN models based on different types of MPI communication functions. Based on the characteristics of point-to-point blocking communication, point-to-point non-blocking communication, one-to-many collective communication, and many-to-many collective communication, model parameters suitable for each type of communication mode are selected, which improves the prediction accuracy of the model in different communication scenarios.

[0070] As a specific implementation method, the preprocessed data is divided into three parts: training set, validation set and test set to ensure that the training and evaluation of the model are representative. i,j , the input feature vector is X i,j =[p i,j ,m i,j ,k i,j ], the target variable is the execution time t of the communication function i,j , the model learns the feature vector X i,j To predict the execution time of the communication function This embodiment combines the characteristics of different communication types, constructs XGBoost and ANN models, and optimizes key hyperparameters through experiments to adapt to the needs of different communication modes.

[0071] (1) XGBoost model training and parameter tuning

[0072] The XGBoost model is a decision tree model based on gradient boosting, which is suitable for processing data with complex structures. When training the XGBoost model, the input feature is X i,j , the output target is t i,jXGBoost reduces prediction errors by combining multiple weak learners (decision trees) and optimizes performance differences for different communication types by adjusting hyperparameters such as learning rate (learning_rate), maximum tree depth (max_depth), child node weight (min_child_weight), and number of decision trees (n_estimators).

[0073] This example uses a grid search combined with cross-validation. Taking point-to-point blocking communication as an example, based on the characteristics of this communication type, the tuning process of the XGBoost model focuses on the selection of the following key hyperparameters:

[0074] learning_rate: The parameter is set to {0.01, 0.05, 0.1}.

[0075] max_depth: parameter is set to a depth range of {3, 5, 7, 9}.

[0076] min_child_weight: candidate values ​​are {1, 3, 5}.

[0077] n_estimators: candidate value is {100, 300, 500}, used to adjust the number of decision trees.

[0078] The above parameter combinations are traversed through grid search, and the performance of the model is evaluated through five-fold cross validation under each set of parameters, and finally the parameter combination with the best performance on the validation set is selected.

[0079] After a series of experimental tuning, this embodiment finally determined the optimal parameter configuration suitable for different communication types, as shown in Table 2 below:

[0080] Table 2 Optimal parameter configurations for different communication types in the XGBoost model

[0081]

[0082]

[0083] (2) ANN model training and parameter tuning

[0084] In ANN, the feature vector X i,j As input, it is transformed by nonlinear activation functions of multiple hidden layers and finally outputs the predicted communication time t i,jIn the optimization process of the ANN model, this embodiment focuses on adjusting the hidden layer structure, activation function, optimizer (solver) and other hyperparameters. Through experiments, it is found that for point-to-point blocking and point-to-point non-blocking communications, the relu activation function and a deeper neural network structure can better capture the characteristics of such communications, while when processing many-to-many set communications, the use of the tanh activation function and a shallower neural network structure can improve training efficiency and prediction accuracy.

[0085] This embodiment uses random search combined with early stopping strategy to optimize hyperparameters. Random search first sets a value range for key parameters, and randomly extracts different parameter combinations within this range for evaluation. Taking the point-to-point blocking communication function as an example, the key tuning parameters include:

[0086] hidden_layer_sizes: Set to [(50,30),(100,50),(100,75)].

[0087] activation: Select from {'relu','tanh'}.

[0088] solver: When selecting an optimization algorithm, choose from {'adam','sgd'}.

[0089] learning_rate_init: The initial learning rate is selected in {0.001, 0.01, 0.1}.

[0090] learning_rate: Set to {'constant','adaptive'}.

[0091] In random search, multiple models are trained for each parameter combination, and the performance on the validation set is monitored through the early stopping mechanism. Once the performance of the model on the validation set is found to have not improved in several rounds, the training is stopped to avoid overfitting. Finally, the parameter configuration with the best effect on the validation set is selected as the final parameters of the ANN model.

[0092] After a series of experimental tuning, this embodiment finally determined the optimal parameter configuration suitable for different communication types, as shown in Table 3 below:

[0093] Table 3 Optimal parameter configurations for different communication types in ANN models

[0094]

[0095] Through the above parameter tuning, the ANN model can optimize the prediction effect of execution time under different communication types, especially showing significant advantages in nonlinear and complex many-to-many communication scenarios.

[0096] To ensure the prediction accuracy of the model, a variety of evaluation indicators were used to evaluate the model performance, including the determination coefficient R 2 (used to measure the degree of model fit) and root mean square error (RMSE). Through the comprehensive evaluation of these indicators, it is ensured that the final selected model has high prediction accuracy and stability.

[0097] For each type of communication function, the corresponding prediction model is obtained Where type j Indicates the communication type, including "point-to-point blocking", "point-to-point non-blocking", "one-to-many collective communication" and "many-to-many collective communication".

[0098] Through experimental design and analysis, this paper implements systematic parameter tuning for different types of communication functions and gradually optimizes model performance. By adjusting relevant hyperparameters, we verify the optimal configuration scheme for each type of communication function, effectively improving the prediction accuracy and stability of the model in different communication scenarios.

[0099] On the basis of function-level modeling, in order to further improve the prediction accuracy, this embodiment introduces model fusion technology. Specifically, the prediction results of XGBoost and ANN models are fused using the weighted average method. The weight distribution in the fusion process is determined based on the performance of each model on the validation set, and the better the model, the higher the weight in the fusion. The fusion model can not only combine the advantages of different models, but also improve the robustness and generalization ability of the model, especially in the face of diverse communication modes and complex application environments, and can provide more accurate communication time prediction.

[0100] Specifically, in order to improve the prediction accuracy, the model fusion technology is used, and the final time prediction of each type of communication function is completed by the weighted fusion of the XGBoost and ANN models. For each type of communication function, the prediction value t of the XGBoost model is obtained. 1 and the predicted value t of the ANN model 2 , through the formula:

[0101] t 融合 =w 1 μt 1 +w 2 ·t 2

[0102] Among them, w 1 and w 2 is the fusion weight of the two models, weight w 1 and w 2The method of determining is as follows: through the validation set, the mean absolute error of each model is calculated, and the weight is determined inversely proportional to the model error to ensure that the model with smaller error has a larger weight in the fusion. The formula is:

[0103]

[0104] Among them, ε 1 and ε 2 They are the mean absolute errors of the XGBoost and ANN models on the validation set, respectively.

[0105] This embodiment classifies MPI functions, selects the optimal machine learning fusion model for different categories, and combines the weighted fusion strategy to integrate the advantages of different models, making the prediction results more accurate and reliable, and can better provide a basis for communication performance analysis and optimization of high-performance computing applications.

[0106] In S3, this embodiment builds an application-level communication time prediction model for MPI communication of high-performance computing applications. Different types of communication function models are integrated into a comprehensive application-level model to predict the communication time of the entire application. First, assume that the application contains multiple MPI functions. For each function, relevant features are extracted, including the number of processes, message size, and number of calls. Next, the communication time is predicted based on the extracted features using the different types of MPI communication function models established above. Then, the predicted communication time of all MPI functions is summarized to obtain the total communication time of the entire application.

[0107] Specifically, for the n MPI communication functions included in the actual HPC application, the feature vector X of each communication function is first extracted. i,j =[p i,j ,m i,j ,k i,j ], where p i,j Indicates the number of processes, m i,j Indicates the message size, k i,j These feature vectors fully reflect the communication mode and communication volume of each communication function. In order to accurately predict the communication time, this embodiment selects the corresponding prediction model for each communication function. to make predictions.

[0108] In the model selection process, first select the communication type of each communication function. j , (point-to-point blocking, point-to-point non-blocking, one-to-many collective communication, many-to-many collective communication) select the most suitable pre-trained prediction model. For each type of communication, there is a pre-trained model By inputting the feature vector Xi,j The predicted communication time of the communication function can be obtained Therefore, the predicted communication time of the i-th communication function can be expressed as:

[0109]

[0110] in Indicates the type of communication function with the i-th type j The corresponding prediction model. The total MPI communication time T of the application is obtained by adding the time consumption of each communication function. The specific formula is as follows:

[0111]

[0112] in, represents the number of the jth type of communication functions, For communication type j Model.

[0113] As a specific implementation method, this embodiment selects an MPI-based supercomputing environment as an experimental platform. On this platform, 50 widely used high-performance computing (HPC) applications are run, covering multiple scientific fields such as physics, chemistry, oceanography, and biology, ensuring the diversity and comprehensiveness of the data. The IPM performance analysis tool is used to collect the communication data of these applications under different numbers of processes (such as 2, 4, 8, 16, 32, 64, 128, 256, and 512). In the communication process of the application, n MPI communication functions are involved, which are represented as f 1 ,f 2 ,...,f n For each MPI communication function f i (where i=1, 2, ..., n), this embodiment collects the following key features:

[0114] (1) Number of processes (p i ): represents the MPI function f i The number of processes participating in the communication.

[0115] (2) Message size (m) i ): indicates the size of the data transmitted during the communication process, in bytes. The message size is related to each MPI communication function f i Therefore, for the i-th MPI function, the message size is m i .

[0116] (3) Communication frequency (k i ): represents the MPI function f iThe number of calls during the communication process, that is, the frequency with which the function is called in an application execution. For each MPI communication function f i , whose communication frequency is k i .

[0117] The relevant feature data of each type of communication function is organized as the input features of the model. The main features include: number of processes, message size, and communication frequency. Then, according to the different communication types, the corresponding feature data is input into the corresponding prediction model f type In (X), the execution time prediction values ​​of various communication functions are obtained.

[0118] The predicted times of all communication types are added together to obtain the overall MPI communication time prediction value of CPMD applications at different scales.

[0119] For each experimental configuration (i.e., under different process numbers), the overall communication time predicted by the model is compared with the communication time measured during the actual execution of CPMD. The relative error is calculated to evaluate the prediction accuracy of the model. The error calculation formula is:

[0120]

[0121] Here, Measured Time indicates measured communication time, and Predicted Time indicates predicted communication time.

[0122] The experimental results are as follows Figure 3 As shown in the figure, through experimental verification of CPMD application, the prediction performance of the model under different communication scales (2 to 512 processes) is demonstrated. In the figure, the blue line represents the actual measured communication time, the red line represents the communication time given by the prediction model, and the black line is the relative error percentage. It can be observed that when the number of processes increases, the predicted communication time is basically consistent with the actual measured communication time trend, indicating that the model effectively captures the characteristics of communication time changing with process scale. The experimental results show that the communication time prediction model based on the present invention has high practicality and accuracy, and can be used for communication time prediction of high-performance computing applications.

[0123] This embodiment significantly simplifies the process of MPI communication performance modeling by adopting a machine learning-based approach. Lightweight performance analysis tools are used to collect communication data, and performance prediction models are established in combination with machine learning techniques, which reduces the consumption of human resources in traditional performance modeling. In addition, the proposed classification modeling strategy can more accurately capture the characteristics of MPI communication functions. By classifying communication functions and building corresponding models, the MPI time consumption in high-performance computing applications can be accurately predicted.

[0124] The method provided by the present invention demonstrates good versatility and ease of use, and uses real high-performance computing applications as training data to provide a reliable reference for performance optimization. Overall, the present invention provides an efficient and reliable solution for MPI communication performance prediction in high-performance computing, laying a foundation for future research and application.

[0125] Embodiment 2

[0126] This embodiment provides a HPC application communication time prediction system based on machine learning, including:

[0127] A type classification module is used to obtain MPI functions in HPC applications and classify MPI functions according to communication modes;

[0128] A function communication time prediction module is used to select an optimal machine learning fusion model according to the category of the MPI function; extract the communication characteristics of the MPI function, input them into the trained optimal machine learning fusion model, and obtain the predicted communication time of the MPI function;

[0129] The program communication time prediction module is used to sum the predicted communication time of all MPI functions to obtain the predicted communication time of the HPC application.

[0130] Embodiment 3

[0131] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the steps in the method for predicting communication time of an HPC application based on machine learning as described in the first embodiment above are implemented.

[0132] Embodiment 4

[0133] This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps in the method for predicting communication time of an HPC application based on machine learning as described in the first embodiment above are implemented.

[0134] The steps or modules involved in the above embodiments 2 to 4 correspond to those in embodiment 1. For the specific implementation, please refer to the relevant description of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0135] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for predicting communication time of HPC applications based on machine learning, characterized in that: include: Obtain MPI functions in HPC applications and classify MPI functions based on communication patterns; According to the category of the MPI function, an optimal machine learning fusion model is selected; the communication characteristics of the MPI function are extracted and input into the trained optimal machine learning fusion model to obtain the predicted communication time of the MPI function; The predicted communication time of all MPI functions is summed to obtain the predicted communication time of the HPC application.

2. A method for predicting communication time of HPC applications based on machine learning as claimed in claim 1, characterized in that: The MPI functions include a message sending function, a message receiving function, a data reduction function, a broadcast function, a collection function, a global reduction function, a global collection function and a full exchange function.

3. The method for predicting communication time of HPC applications based on machine learning as claimed in claim 1, characterized in that: The MPI functions are classified according to the communication modes, and the types include point-to-point blocking communication, point-to-point non-blocking communication, one-to-many collective communication and many-to-many collective communication.

4. The method for predicting communication time of HPC applications based on machine learning according to claim 1, characterized in that: The communication characteristics include the number of processes, message size and communication frequency.

5. The method for predicting communication time of HPC applications based on machine learning according to claim 1, characterized in that: The selecting the optimal machine learning fusion model according to the category of the MPI function specifically includes: The optimal machine learning fusion model includes an optimal XGBoost model and an optimal ANN model; the training process of the optimal XGBoost model and the optimal ANN model is: The historical communication characteristics and historical communication time of each type of MPI function are obtained to construct the training set and validation set; Set XGBoost model hyperparameters, perform parameter tuning based on grid search and cross-validation, use the training set to train the model, use the validation set to validate the various evaluation indicators of the model, and save the model as the optimal parameter model if it meets the standards, as the optimal XGBoost model suitable for different types of MPI functions; Set ANN model hyperparameters, perform parameter tuning based on random search and early stopping strategies, use the training set to train the model, use the validation set to validate the various evaluation indicators of the model, and save the optimal parameter model if it meets the standards, as the optimal ANN model suitable for different types of MPI functions.

6. A method for predicting communication time of HPC applications based on machine learning as claimed in claim 5, characterized in that: The extracting of the communication features of the MPI function and inputting them into the trained optimal machine learning fusion model to obtain the predicted communication time of the MPI function specifically includes: The communication features are respectively input into the trained optimal XGBoost model and the optimal ANN model to obtain a first predicted communication time and a second predicted communication time; The first predicted communication time and the second predicted communication time are weightedly fused to obtain the predicted communication time of the MPI function.

7. A method for predicting communication time of HPC applications based on machine learning as claimed in claim 6, characterized in that: The weights of the two models in the weighted fusion are obtained by the mean absolute error obtained by model training; the weights are determined in inverse proportion to the mean absolute error of the models to ensure that the model with a smaller mean absolute error has a larger weight in the fusion.

8. A machine learning-based HPC application communication time prediction system, characterized in that: include: A type classification module is used to obtain MPI functions in HPC applications and classify MPI functions according to communication modes; A function communication time prediction module is used to select an optimal machine learning fusion model according to the category of the MPI function; extract the communication characteristics of the MPI function, input them into the trained optimal machine learning fusion model, and obtain the predicted communication time of the MPI function; The program communication time prediction module is used to sum the predicted communication time of all MPI functions to obtain the predicted communication time of the HPC application.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in a method for predicting communication time of an HPC application based on machine learning as described in any one of claims 1-7 are implemented.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the method for predicting communication time of HPC applications based on machine learning are implemented as described in any one of claims 1-7.

Citation Information

Cited By

  • Performance analysis method and device for source address verification, equipment and storage medium

    CN121644210A