Parallel Computing Application Runtime Optimization Method and System Based on Remote Intelligent Services
Through remote intelligent service automation, the hardware parameter configuration is solved, and the problem of resource utilization difficulties for scientific researchers in high-performance computing clusters is solved, efficient resource utilization and cost reduction are achieved, and the smooth progress of scientific research work is promoted.
Patent Information
- Application Number
- CN202510437036.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-09
AI Technical Summary
In the prior art, when scientific researchers use high-performance computing clusters, they lack detailed hardware configuration suggestions, which leads to high usage thresholds, difficulty in effectively utilizing resources, and lack of high-quality operation data sets, resulting in waste of resources and increased costs.
Through remote intelligent services, combined with local cluster characteristics and remote historical job database, the hardware parameter configuration of the job is automatically determined, parameter configuration suggestions are provided, and the user's requirement for scheduling of specified resources is reduced.
The parameter configuration of parallel computing applications has been optimized, resource utilization has been improved, user usage costs have been reduced, operational operation processes have been simplified, and scientific research has been promoted efficiently.
Smart Images

Figure CN119961004B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of supercomputing clusters, and particularly to a method and system for optimizing the runtime of parallel computing applications based on remote intelligent services. Background Art
[0002] With the rapid development of technology, the construction of high-performance computing platforms has advanced by leaps and bounds, bringing unprecedented impetus to complex computing and simulation tasks. This progress enables researchers to efficiently process and analyze large-scale datasets in the scientific research field within a reasonable time frame, significantly improving the efficiency and stability of computing task execution. Many scientific research software has been adapted to the parallel configuration of high-performance platforms, such as the first-principles calculation software VASP widely used in the field of materials science, and the classic molecular dynamics simulator LAMMPS, etc. Before these software run in parallel on rich computing resources, not only basic hardware resource parameters need to be set, such as the number of nodes allocated for running and the number of cores per node, etc., but software with more complex structures also has relatively high requirements for the parallel parameter configuration strongly related to subject area knowledge in the job. For example, for a job running on a cluster, by applying for corresponding allocation of a certain number of nodes and cores, the specific quantity and allocation of the two are often based on factors such as past experience in running related jobs and machine time overhead costs. During the actual final job execution process, the number of cores allocated on a node affects issues such as communication efficiency, computing speed, and storage security.
[0003] From the perspective of computer resource managers, maximizing the utilization of resources is the core goal. In this context, setting a set of optimized parallel parameters is crucial for improving job execution efficiency. Therefore, conducting in-depth exploration work on optimizing the parameters of scientific research software has important practical significance. However, currently deployed software on clusters generally lacks detailed hardware configuration suggestions, which brings troubles to researchers, especially computing users without a computer-related academic background. They often lack sufficient understanding of core allocation on high-performance computing clusters and are even more difficult to master the complex parallel parameter configuration and its optimization methods in various software. This undoubtedly raises the threshold for using scientific research tools to a certain extent and hinders the efficient progress of scientific research work. For researchers, while ensuring the correctness and running speed of jobs, how to save job running costs and effectively utilize resources within a reasonable budget has become an important goal they pursue. This is not only related to the cost control of scientific research projects but also to the sustainable development of scientific research work.
[0004] At present, there is a serious shortage of high-quality job running datasets for specific tasks, which has become a major bottleneck restricting the development of this field. The existing patent "A method for optimizing job running parameters applied to supercomputer cluster scheduling" provides an effective means to test and evaluate the advantages and disadvantages of various parameter configurations, and allows the running data of test jobs to be saved, thereby generating high-quality job data for a specific single cluster. However, analyzing data based only on a single cluster and building a complete dataset from scratch locally is not only time-consuming and laborious, but also storing the running data of common scientific computing software multiple times as a whole will inevitably lead to unnecessary data redundancy, further exacerbating resource waste. Summary of the Invention
[0005] The object of the present invention is to propose a method and system for optimizing the runtime of parallel computing applications based on remote intelligent services, providing better parameter configuration suggestions based on the analysis and processing of a large amount of running data, automatically determining the hardware parameter configuration of job applications, and eliminating the need for users to specify resource quantity scheduling, thereby reducing the threshold for using cluster software.
[0006] To achieve the above object, the present invention proposes a method for optimizing the runtime of parallel computing applications based on remote intelligent services, including the following steps:
[0007] Step S1: Parse the application job submitted by the user in the local cluster, obtain the application job and local cluster characteristics, and push them to the remote server;
[0008] Step S2: The remote server receives the processing applications of the application jobs and local cluster characteristics submitted by each local cluster, and allocates the corresponding parameter search sub-module through the parameter search module to predict the parameters of the application job, obtaining a list of parameter configuration values;
[0009] Step S3: The remote server sends the predicted list of parameter configuration values to the local cluster, and the list of parameter configuration values includes one or more groups of parameter configuration values;
[0010] Step S4: The local cluster filters out the most suitable parameter configuration through a custom method, feeds back the specific parameter configuration value to the user, and at the same time modifies the number of hardware resources applied for by the application job and the application job parameter settings, and submits them to the local cluster scheduling system to complete the actual application job running;
[0011] Step S5: The local cluster monitors the running of the application job and uploads the actual running status results to the remote server, and the results include the local actual running time, the category to which the application job belongs, the parameters related to the parallel running of the application job, and the characteristic values of the application job and hardware running environment parameters;
[0012] Step S6: Expand the dataset in the historical job database of the remote server. The remote server stores the received actual operation status result data into the corresponding application type dataset, and periodically updates the parameter search module based on the updated dataset.
[0013] Preferably, in step S1, the application job characteristics are parameters that affect the computational overhead of the application job to be run, including the amount of calculation data, the number of calculation tasks, and the number of communications; the local cluster characteristics are the characteristic expressions after processing the environmental parameter configurations of the cluster, and the environmental parameter configurations include the hardware performance parameters of the CPU, GPU, acceleration card, and the interconnection network between nodes or cards.
[0014] Preferably, in step S1, the application job submitted by the user further includes the job operation requirement parameters submitted by the user, and the requirement parameters include three modes: calculation speed priority, hardware efficiency priority, and speed-efficiency balance.
[0015] Preferably, in step S2, the parameter search module is divided into different types of parameter search sub-modules according to different application types and different Er values are configured. The parameter search sub-module is trained according to the job data of the corresponding application in the historical job database, and Er is the quality determination index.
[0016] Preferably, in step S3, the set of parameter configuration values includes the parallel parameters and their values of the sub-tasks in the application task, and the hardware resource parameters and their values of the application job.
[0017] Preferably, in step S4, the custom method is to select the most suitable hardware parameters for the current local cluster to run the target application job from the parameter configuration list returned by the remote server. The hardware parameters are the parameter configurations that maximize the utilization of idle resources under the condition of not exceeding the number of idle resources of the current local cluster. The specific steps are as follows:
[0018] Step S41: Select the candidate values of the parameter list that meet the application user's permissions and the current running conditions of the local cluster from the parameter configuration list;
[0019] Step S42: Sort the multiple sets of parameter configuration value lists that meet the conditions according to the requirement parameters submitted by the user, and select the optimal set of parameters.
[0020] Preferably, in step S6, the expanded dataset is collected from one or more groups of the optimal parameter configuration list issued by the remote server and applied to the local cluster job operation data.
[0021] The present invention also provides a parallel computing application runtime optimization system based on remote intelligent services, including a local cluster and a remote server;
[0022] The local cluster includes: a feature extraction module, a request access module, a resource configuration module, and a scheduling system module; among them, the feature extraction module: is used to extract feature description parameters related to the parallel computing tasks of the application jobs to be run; the request access module: is used to send the application jobs and local cluster features to the remote server and receive the list of parameter configuration values sent down by the remote server; the resource configuration module: is used to complete the screening and sorting of the parameter configuration values from the list of parameter configuration values according to the priorities set by the user and the cluster, and obtain the optimal parameter configuration to be run; the scheduling system module: is used to manage and schedule the actual operation of the application jobs.
[0023] The remote server includes: a parameter search module and a historical job database; among them, the parameter search module: generates parameter configuration suggestions for the application jobs to be run according to the received application jobs and local cluster features; the historical job database is used to store and expand job operation data: providing training data for the parameter search module.
[0024] Preferably, the parameter search module includes multiple parameter search sub-modules, which respectively process the parameter optimization problems corresponding to different types of applications and update the parameter search module regularly based on the updated data set.
[0025] Preferably, the optimization system supports the parameter optimization service of multiple clusters sharing the remote server.
[0026] Therefore, the present invention proposes a parallel computing application runtime optimization method and system based on remote intelligent services, and its beneficial effects are as follows:
[0027] (1) The parallel computing application runtime optimization method and system based on remote intelligent services proposed by the present invention optimize the parameters related to parallel efficiency by combining the local feature extraction module and the remote historical job database. The local feature extraction module extracts features based on specific jobs and the current cluster hardware configuration, and collects accurate local personalized configuration information. There are multiple remote parameter search modules, which respectively process the corresponding parameters of common software, and the remote high-performance servers provide a robust and easy-to-use parameter search model to predict suitable hardware parameters. The present invention provides users with a relatively comprehensive remote service and a parameter feature extraction module that can set corresponding preferences.
[0028] (2) The present invention improves the utilization rate of the operation results of jobs on the supercomputing platform and advocates the construction and expansion of a truly available job operation dataset for each application type. A complete job operation requires multiple steps such as the allocation of hardware resources, job scheduling system scheduling, and actual operation. Among them, situations such as queuing may further increase the operation cost. There are frequent and massive job application runs on the cluster platform, and these operation data cannot be fully utilized. At the same time, the conversion rate of the operation cost under each successful parameter configuration to the next operation experience is not high. Corresponding to a large amount of undervalued and idle data, there is a lack of high-quality job operation datasets for specific applications in reality. By requesting access to the remote parameter search model, not only can the parameters of the parameter search model be continuously updated, but also the corresponding dataset can be stably expanded. In the framework of the present invention, a historical job database established remotely collects the operation data of various applications and provides it for the model training and learning of the corresponding applications. Realizing basic optimized data sharing is beneficial to saving the cost of parallel parameter optimization for each dispersed cluster and can better collect operation data.
[0029] (3) The system proposed in the present invention can better exert the potential of the high-performance platform to provide user services. Users submit job characteristics to the remote server through the local cluster, and multiple clusters share the parameter optimization service of the remote server. The high-concurrency processing of parameter prediction requests by the high-performance platform can significantly improve the utilization rate of the remote server cluster.
[0030] (4) The present invention assists users on the supercomputing center cluster to complete the setting of the necessary number of hardware resources for job operation and the estimation of relevant parallel parameter numbers. That is, users do not need to contact the parameters strongly related to hardware parallelism on the cluster software at all, greatly simplifying the job operation process and reducing the learning cost of users using supercomputing software. Under the existing parameter search experience and methods, the parameter values specified by users are not necessarily accurate and reliable. The remote parameter search module in the present invention can, based on a large amount of data, mine the parameters related to job parallelism, the number of hardware resource allocations, and the hidden characteristics of the cluster hardware in the operation data, saving users a great deal of time in adjusting parameters and various operation expenses of jobs. Furthermore, it helps users automatically perform efficient core number setting under the condition of meeting the budget, enabling users to run jobs without being aware of the hardware, thus better promoting scientific research work based on the high-performance service platform.
[0031] The technical solution of the present invention will be further described in detail below through the accompanying drawings and embodiments. Description of the Drawings
[0032] Figure 1 It is a flowchart of the parallel computing application operation optimization method based on remote intelligent services of the present invention;
[0033] Figure 2It is the timing diagram of the parallel computing application runtime optimization method based on remote intelligent service in Embodiment 1;
[0034] Figure 3 It is the state diagram of the parallel computing application runtime optimization system based on remote intelligent service in the present invention;
[0035] Figure 4 It is the connection schematic diagram of the cluster and the remote server in the present invention. Detailed implementation manners
[0036] To make the technical solutions, advantages and objectives of the present invention clearer, the technical solutions of the embodiments of the present invention will be described clearly and completely below. The described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the described embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of this application.
[0037] Unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meanings understood by those with ordinary skills in the field to which the present invention belongs.
[0038] Embodiment 1
[0039] As Figure 1-2 shown, the present invention proposes a parallel computing application runtime optimization method based on remote intelligent service, including the following steps:
[0040] SA1. The local cluster obtains the application job submitted by the user.
[0041] SA2. The local cluster obtains the application category in the application job information, and estimates the characteristic input information of the job to be run according to the application category. The characteristic input information includes cluster characteristics and job characteristics. Among them, the cluster characteristics correspond to the input of the local hardware environment at the current moment, and the job characteristics correspond to the input parameters and their descriptions in the application job that are irrelevant to the parallel configuration such as the overall hardware resource quantity of the job and the resource parameters of specific subtasks; specifically, the job characteristics are parameters that affect the calculation overhead such as the amount of calculation data, the number of calculation tasks, and the number of communications. By checking the input parameters of the calculation program and the intermediate parameters obtained during the preprocessing process of the calculation program, verify whether they have an impact on the calculation overhead, and finally obtain a list of characteristic parameters and obtain their values.
[0042] SA3. The local cluster uploads the characteristic input information of the job to be run and applies for remote parameter optimization service, and obtains a list of parameter configuration values composed of multiple groups of parameter estimation information. Each group of information includes the overall hardware resource quantity of the job and the assignment of resource parameters of specific subtasks.
[0043] SA4. The remote server forwards the list of parameter configuration values of the job to be run to the local cluster.
[0044] SA5. The local cluster receives the list of parameter configuration values, obtains the optimal parameter combination through the screening and sorting method defined by the user and the cluster, and modifies the configuration of the job to be run.
[0045] SA6. The local cluster actually runs the job according to the optimized parameter configuration and feeds back the running data to the remote historical job database.
[0046] In step SA2, the local hardware environment refers to the unique hardware configuration of the local hardware, such as the number of nodes, the number of cores and their corresponding relationships, etc. The overall resource quantity of the job refers to the estimated hardware allocation value of the job to be run, such as the total number of CPU cores and GPUs applied for when the job runs. The parameter configuration of specific subtask resources refers to the parameter settings related to parallelization in jobs of specific different application types, which need to be determined based on the overall resource quantity of the aforementioned job.
[0047] Embodiment 2
[0048] As Figure 3 shown, the present invention also proposes a parallel computing application runtime optimization system based on remote intelligent services, including a local cluster and a remote server.
[0049] The local cluster includes: a feature extraction module, a request access module, a resource configuration module, and a scheduling system module. Among them, the feature extraction module: sub-module according to specific application categories, used to extract the feature representations of the job to be run and the local cluster hardware. The input of this module is divided into two parts, the job parameter description that extracts job features and has no direct association with job parallelization, and the necessary local hardware environment configuration that extracts cluster features. The output of this module is the feature input information that can be utilized by the remote parameter search module; the resource configuration module: used to complete the screening and sorting of parameter configuration values from the list of parameter configuration values according to the priorities set by the user and the cluster, and obtain the optimal parameter configuration to be run; the request access module: used to send job and cluster features to the remote server and receive the list of parameter configuration values sent by the remote server; the scheduling system module: used to manage and schedule the actual operation of the job.
[0050] The parameter feature extraction model can also be trained using the historical job data stored locally. This data includes the hardware operating environment, the original information of the job and its processing features, and the actual running resource assignment of the job. After the job runs, the parameter feature module is responsible for collecting the necessary job information including the actual running time of the job, and optionally adding it to the local historical job database for timely training of the local feature module.
[0051] The remote server includes: a parameter search module and a historical job database. Among them, the parameter search module: generates parameter configuration suggestions for the job to be run according to the received job and cluster characteristics, where the sorting in the parameter list refers to the Er value of the pros and cons judgment index; the historical job database is used to store and expand job operation data and provide training data for the parameter search module.
[0052] The parameter search module contains multiple subdivided parameter search sub-modules. Each sub-module is trained by the corresponding application job operation data in the historical job database and processes the parameter optimization problems corresponding to different types of applications. Since the data in the remote historical job database is less in the initial stage and lacks sufficient training, the parameter search module on the remote server side can obtain a list of estimated values of several parameter configurations for the job to be run through an empirical model in the relevant application field, and obtain the actual operation data of the corresponding configuration through the trial operation of the local cluster, while supplementing the data in the remote server side historical job database.
[0053] As Figure 4 shown, the parallel computing application runtime optimization method and system based on remote intelligent services provided by the present invention support multi-cluster sharing of remote cloud service parameter optimization.
[0054] Embodiment 3
[0055] In this embodiment, taking the commonly used job scheduling system SLURM job scheduling system and the commonly used scientific computing software VASP in the supercomputer platform as examples, a parallel computing application runtime optimization method and system based on remote intelligent services provided by the present invention are further explained. The specific operation steps are as follows:
[0056] Step 1: The user submits a VASP application job to the SLURM job management system through the submission system. The submission system obtains the job input information and transfers it to the feature extraction module. The job input information includes application type description, input file path, hardware resource application information, instruction sequence in the job, etc.
[0057] Step 2: The feature extraction module receives the job input information and assigns the job to the corresponding feature extraction sub-module according to the application type description therein. The feature extraction sub-module estimates whether different inputs and parameters in the previous processing process have an impact on the calculation overhead of the VASP job, and then obtains the parameters that have an impact, such as the number of energy bands, the number of atoms, the reduced number of K-point calculations, the energy cutoff value, the number of transition state calculation mirrors, the size of the Fourier transform array, etc. as job features. At the same time, this module collects the hardware parameters of the local cluster at the current moment and uses them as inputs to generate a similar local cluster feature expression. The request access module sends the feature input information containing job features and cluster features to the remote parameter search module.
[0058] Step 3: After receiving the feature input information, the remote parameter search module outputs a list of parameter estimated values including the overall job resource quantity and the parameters of specific job subtasks. The specific process is to allocate the job to the corresponding parameter search sub-module for parameter prediction according to the application type description therein. For the specific subtask of the VASP job, the parameters with relatively high correlation with the parallel running efficiency are NPAR, KPAR, and NCORE. These three parameters follow the constraint that their product is the total number of cores allocated for the job. Based on the parameter prediction sub-module to which the VASP job in the historical job database belongs, the hardware resource configuration values of the parameters related to parallelization of multiple groups of jobs and the overall job resource quantity are predicted.
[0059] Step 4: The local cluster receives the list of parameter configuration values returned by the remote parameter search module, and the resource configuration module performs two processing procedures: screening and sorting. The local cluster performs basic screening on the parameter configuration values in combination with the current idle resources of the cluster, and completes the high-order screening for further parameter value configuration in combination with the user's permissions and preferences, etc., to obtain the optimal parameter combination and modify the configuration of the job to be run.
[0060] Step 5: The local cluster actually runs the job according to the optimized parameter configuration and feeds back the running data to the remote historical job database.
[0061] In this embodiment, other parameters in the VASP job input file will affect the accuracy of the three main parameters related to parallelization in Step 2. For example, the relationship between the number of K points NKPTS and the number of parallel K-point calculations KPAR. The training feature extraction module can deeply explore the complex structure of job parameters and has higher efficiency and accuracy compared to empirical selection.
[0062] To complete the training of the parameter search sub-module in Step 3, a high-quality historical job running data set needs to be prepared. When the construction of the remote historical job database is insufficient, the parameter search module will use an empirical model to generate multiple parameter configuration recommendations for the local cluster to generate the running results of multiple test jobs, and add important running information such as the actual job running time of the remote cluster, the assignment of each parameter, and the setting of the hardware resource number to the remote database classified by application type for regularly updating the corresponding sub-module parameter search model. The sufficiency of the database data means that the number of running results actually fed back from the operation of each local cluster can train the remote parameter search sub-module until the error converges to a certain range, and this part is adjusted manually. When the construction of the remote historical job library is sufficient, the parameter search module will use one or more parameter search methods to generate a certain number of parameter configuration values, feedback the list to the user through the access module, and collect the running results. To sum up, the parameter prediction work can be set to different situations of parameter search, or a combination of empirical model and parameters.
[0063] In specific implementation, the high-order screening of the parameter configuration value list in step 4 can be based on two methods: user permissions and the job running requirements specified by the user. Screening the parameter configuration values according to user permissions means judging whether the maximum available hardware resource usage number allocated to the user by the cluster is consistent with the recommended hardware resource allocation in the current parameter configuration list; screening the most suitable parameter configuration according to the job running requirement parameters specified by the user means the parameter screening preference that needs to be clearly pointed out before the user applies for job running. This method combines automated screening and user demands by omitting the actual user screening configuration process. The methods include but are not limited to "computing speed first", "hardware efficiency first", and "speed-efficiency balance". Among them, "computing speed first" means that under the same conditions, it is inclined to allocate sufficient hardware resources to the job to be run to accelerate the computing process of the application job. The actual strategy is to select a set of search values with a larger number of computing cores in the available parameter configuration value list. "Hardware efficiency first" means that the cluster computing resources occupied to complete this task are as few as possible. The actual strategy is to select the configuration with a smaller number of computing cores in the available parameter list. "Speed-efficiency balance", as a compromise between the previous two screening methods, selects the parameter configuration with a medium number of computing cores to achieve the purpose of balancing the number of allocated hardware resources and accelerating computing.
[0064] For the local cluster configuration of a specific job scheduling system in this embodiment, such as the SLURM scheduling system allowing the submission of jobs to set the core number range, then the most suitable hardware resource configuration value finally selected can comprehensively consider several groups of configuration values with relatively high rankings.
[0065] For the situation where the receipt message of the local cluster applying for remote parameter search service and feedback job running information times out, the local request access control module automatically ends the parallel computing application running optimization process of the remote cloud service and promptly feedbacks the exception to the user.
[0066] It should be noted that the content not elaborated in detail in the present invention is all prior art and is well-known to those skilled in the art.
[0067] Therefore, the present invention provides a method and system for optimizing the running of parallel computing applications based on remote intelligent services, which can optimize the running parameters of parallel computing applications, reduce data redundancy, improve the utilization rate of the supercomputer platform, reduce the user's usage cost, and strongly promote scientific research work on high-performance service platforms.
[0068] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for optimizing the runtime of parallel computing applications based on remote intelligent services, characterized in that The steps are as follows: Step S1: Parse the application job submitted by the user in the local cluster, obtain the application job characteristics and local cluster characteristics, and push them to the remote server; Step S2: The remote server receives the processing applications of the application job characteristics and local cluster characteristics submitted by each local cluster, and allocates the corresponding parameter search sub-module through the parameter search module to perform parameter prediction on the application job, and obtains a list of parameter configuration values; Step S3: The remote server distributes the predicted list of parameter configuration values to the local cluster, and the list of parameter configuration values includes one or more groups of parameter configuration values; Step S4: The local cluster filters out the most suitable parameter configuration through a custom method, feeds back the specific parameter configuration value to the user, and at the same time modifies the number of hardware resources applied for by the application job and the application job parameter settings, and submits them to the local cluster scheduling system to complete the actual application job operation; Step S5: The local cluster monitors the operation of the application job and uploads the actual operation status results to the remote server. The results include the local actual operation time, the category to which the application job belongs, the parameters related to the parallel operation of the application job, and the characteristic values of the application job and the hardware operation environment parameters; Step S6: Expand the dataset in the remote server historical job database. The remote server stores the received actual operation status result data in the corresponding application type dataset, and periodically updates the parameter search module based on the updated dataset.
2. The parallel computing application runtime optimization method based on remote intelligent services according to claim 1, wherein In step S1, the application job characteristics are the parameters that affect the computing overhead of the application job to be run, including the amount of calculation data, the number of calculation tasks, and the number of communications; the local cluster characteristics are the characteristic expressions after processing the environmental parameter configuration of the cluster, and the environmental parameter configuration includes the hardware performance parameters of the CPU, GPU, acceleration card, and the interconnection network between nodes or cards.
3. The parallel computing application runtime optimization method based on remote intelligent services according to claim 1, characterized in that In step S1, the application job submitted by the user also includes the job operation requirement parameters submitted by the user, and the requirement parameters include three modes: calculation speed priority, hardware efficiency priority, and speed-efficiency balance.
4. The parallel computing application runtime optimization method based on remote intelligent services according to claim 1, characterized in that In step S2, the parameter search module is divided into different types of parameter search sub-modules according to different application types and different Er values are configured. The parameter search sub-module is trained according to the job data of the corresponding application in the historical job database, and the Er is the quality determination index.
5. The parallel computing application runtime optimization method based on remote intelligent services according to claim 1, characterized in that, In step S3, the group of parameter configuration values includes the parallel parameters and their values of the sub-tasks in the application task, and the hardware resource parameters and their values of the application job.
6. The parallel computing application runtime optimization method based on remote intelligent services according to claim 1, characterized in that In step S4, the custom method is to select the hardware parameters that are most suitable for the current local cluster to run the target application job from the parameter configuration list returned by the remote server. The hardware parameters are the parameter configurations that maximize the use of idle resources under the condition that the number of idle resources of the current local cluster is not exceeded. The specific steps are as follows: Step S41: Select the candidate values of the parameter list that meet the application user's permissions and the current running conditions of the local cluster from the parameter configuration list; Step S42: Sort the multiple groups of parameter configuration value lists that meet the conditions according to the requirement parameters submitted by the user, and select the optimal group of parameters.
7. The parallel computing application runtime optimization method based on remote intelligent services according to claim 1, wherein In step S6, the extended data set is collected from one or more groups of optimal parameter configuration lists sent by the remote server and applied to the local cluster job running data.
8. A parallel computing application runtime optimization system based on remote intelligent services, including a local cluster and a remote server, characterized in that the local cluster includes: a feature extraction module, a request access module, a resource configuration module, and a scheduling system module; among them, the feature extraction module: is used to extract the feature description parameters related to the parallel computing tasks of the application job to be run; the request access module: is used to send the application job and the local cluster features to the remote server and receive the parameter configuration value list sent by the remote server; the resource configuration module: is used to complete the screening and sorting of the parameter configuration values from the parameter configuration value list according to the priorities set by the user and the cluster, and obtain the optimal parameter configuration to be run; the scheduling system module: is used to manage and schedule the actual operation of the application job. The remote server includes: a parameter search module and a historical job database; among them, the parameter search module: generates parameter configuration suggestions for the application job to be run according to the received application job and local cluster features; the historical job database is used to store and expand job running data: providing training data for the parameter search module.
9. The parallel computing application runtime optimization system based on remote intelligent services according to claim 8, wherein, The parameter search module includes multiple parameter search sub-modules, which respectively process the parameter optimization problems corresponding to different types of applications and update the parameter search module regularly based on the updated data set.
10. The parallel computing application runtime optimization system based on remote intelligent services according to claim 8, characterized in that, The optimization system supports multiple clusters to share the parameter optimization service of the remote server.
Citation Information
Patent Citations
Job operation parameter optimization method applied to super-computing cluster scheduling
CN114048027A
Parameter configuration method and device in multi-computing-cluster environment and computer equipment
CN119597731A