Parallel computing application runtime optimization method and system based on remote intelligent service

By introducing runtime optimization methods and systems of parallel computing applications based on remote intelligent services on high-performance computing clusters, the hardware parameter configuration is automatically determined, which solves the problem of parallel parameter configuration when scientific research software is run on high-performance computing clusters, improves operational efficiency, reduces user usage costs, and promotes the efficiency and sustainability of scientific research work.

CN119961004AActive Publication Date: 2025-05-09UNIV OF SCI & TECH OF CHINA
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510437036.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-09
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

In the prior art, when scientific research software runs on high-performance computing clusters, there is a lack of detailed hardware configuration suggestions, which makes it difficult for scientific researchers to optimize parallel parameter configurations, improve operational efficiency, and lacks high-quality operational data sets, resulting in waste of resources and increased computing costs.

Method used

A parallel computing application runtime optimization method and system based on remote intelligent services is proposed. By analyzing the application jobs submitted by users and local cluster characteristics, the processing application is pushed to the remote server, and the hardware parameter configuration of the job application is automatically determined using the parameter search module, which lowers the threshold for users to use cluster software.

Benefits of technology

Through automated determination of hardware parameter configuration, users' operational complexity in high-performance computing clusters is reduced, job operation efficiency is improved, data redundancy and resource waste are reduced, user usage costs are reduced, and scientific research work is promoted efficient and sustainable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961004A_ABST
    Figure CN119961004A_ABST
Patent Text Reader

Abstract

The invention specifically discloses a remote intelligent service-based parallel computing application runtime optimization method and system, and relates to the technical field of super computing clusters. The method comprises the following steps: firstly, acquiring job and cluster characteristics and pushing the characteristics to a remote server; after the remote server receives the application, the parameter search module allocates sub-modules to carry out parameter prediction on the job and issues the parameter prediction result to the local cluster; the local cluster screens out the optimal parameter configuration through a resource configuration module self-defining method, feeds back the optimal parameter configuration to a user, modifies job related settings and submits the job related settings for operation at the same time; the local cluster monitors the operation and uploads the actual operation state result to a remote server; and finally, the remote server expands a historical job database, stores the received data to a corresponding application type data set, and updates the parameter search module regularly according to the updated data set. According to the method, the operation data is fully utilized to optimize parameter configuration, the use cost of a user is reduced, and high-efficiency promotion of scientific research work based on a high-performance service platform is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of supercomputing clusters, and in particular to a method and system for optimizing parallel computing applications at runtime based on remote intelligent services. Background Art

[0002] With the rapid development of science and technology, the construction of high-performance computing platforms is changing with each passing day, bringing unprecedented impetus to complex computing and simulation work. This progress enables researchers to efficiently process and analyze large-scale data sets in the field of scientific research within a reasonable time frame, significantly improving the efficiency and stability of computing tasks. Many scientific research software have been adapted to the parallel configuration of high-performance platforms, such as VASP, a first-principles calculation software widely used in the field of materials science, and LAMMPS, a classic molecular dynamics simulator. Before such software runs in parallel on abundant computing resources, it is not only necessary to set basic hardware resource parameters, such as the number of nodes and the number of node cores allocated for operation, but also software with more complex structures also puts forward higher requirements for the parallel parameter configuration that is strongly related to the knowledge of the subject area in the operation. For example, a job running on a cluster is allocated to a number of nodes and cores through application. The specific number and allocation of the two are often based on past experience in running related jobs, machine time overhead costs and other factors. In the actual final job running process, the number of cores allocated on a node affects communication efficiency, computing rate and storage security.

[0003] From the perspective of computer resource managers, maximizing resource utilization is the core goal. In this context, setting a set of optimized parallel parameters is crucial to improving job running efficiency. Therefore, conducting in-depth exploration of scientific research software parameter optimization is of great practical significance. However, the software currently deployed on clusters generally lacks detailed hardware configuration recommendations, which has caused trouble for researchers, especially computing users with non-computer-related backgrounds. They often lack sufficient understanding of the core number allocation on high-performance computing clusters, and it is even more difficult to master the complicated parallel parameter configuration and optimization methods in various software. This undoubtedly raises the threshold for the use of scientific research tools to a certain extent and hinders the efficient advancement of scientific research work. For researchers, how to save job running costs and achieve effective resource utilization within a reasonable budget while ensuring job correctness and running speed has become an important goal they pursue. This is not only related to the cost control of scientific research projects, but also to the sustainable development of scientific research work.

[0004] At present, there is a serious shortage of high-quality job running data sets for specific tasks, which has become a major bottleneck restricting the development of this field. The existing patent "A method for optimizing job running parameters for supercomputing cluster scheduling" provides an effective means to test and evaluate the advantages and disadvantages of multiple parameter configurations, and allows the running data of the test job to be saved, thereby generating high-quality job data for a specific single cluster. However, it is not only time-consuming and laborious to analyze only based on the data of a single cluster and build a complete data set from scratch locally, but also the overall storage of the running data of common scientific computing software multiple times will inevitably lead to unnecessary data redundancy, further exacerbating the waste of resources. Summary of the invention

[0005] The purpose of the present invention is to propose a runtime optimization method and system for parallel computing applications based on remote intelligent services, which provides better parameter configuration suggestions based on the analysis and processing of a large amount of operating data, and automatically determines the hardware parameter configuration of the job application without the need for the user to specify resource scheduling, thereby lowering the threshold for cluster software use.

[0006] To achieve the above object, the present invention proposes a runtime optimization method for parallel computing applications based on remote intelligent services, comprising the following steps: Step S1: parse the application job submitted by the user in the local cluster, obtain the application job and local cluster features, and push them to the remote server; Step S2: The remote server receives the application jobs and the processing applications of the local cluster features submitted by each local cluster, and allocates the corresponding parameter search submodule through the parameter search module to perform parameter prediction on the application jobs to obtain a parameter configuration value list; Step S3: The remote server sends the predicted parameter configuration value list to the local cluster, where the parameter configuration value list includes one or more groups of parameter configuration values. Step S4: The local cluster selects the most suitable parameter configuration through a custom method, feeds back the specific parameter configuration value to the user, modifies the number of hardware resources and application job parameter settings requested by the application job, and submits them to the local cluster scheduling system to complete the actual application job operation; Step S5: The local cluster monitors the running of the application job and uploads the actual running status result to the remote server, wherein the result includes the local actual running time, the category of the application job, the related parameters of the parallel running of the application job, and the characteristic values ​​of the application job and the hardware running environment parameters; Step S6: the data set in the historical operation database of the remote server is expanded. The remote server stores the received actual operation status result data into the corresponding application type data set, and periodically updates the parameter search module based on the updated data set.

[0007] Preferably, in step S1, the application job characteristics are parameters that affect the computing overhead of the application job to be run, including the amount of computing data, the number of computing tasks and the number of communications; the local cluster characteristics are characteristic expressions after processing the cluster's environmental parameter configuration, and the environmental parameter configuration includes hardware performance parameters of the CPU, GPU, accelerator card, and node or inter-card interconnection network.

[0008] Preferably, in step S1, the application job submitted by the user also includes job running requirement parameters submitted by the user, and the requirement parameters include three modes: computing speed priority, hardware efficiency priority and speed efficiency balance.

[0009] Preferably, in step S2, the parameter search module is divided into different types of parameter search sub-modules according to different application types and configured with different Er values. The parameter search sub-module is trained based on the job data of the corresponding application in the historical job database, and the Er is a good or bad judgment indicator.

[0010] Preferably, in step S3, the set of parameter configuration values ​​includes parallel parameters of subtasks in the application task and their values, and hardware resource parameters of the application job and their values.

[0011] Preferably, in step S4, the custom method is to select the hardware parameters that are most suitable for the current local cluster to run the target application job from the parameter configuration list returned by the remote server, and the hardware parameters are parameter configurations that maximize the use of idle resources without exceeding the number of idle resources in the current local cluster. The specific steps are as follows: Step S41: Select parameter list candidate values ​​that meet the application user's authority and the current operating conditions of the local cluster from the parameter configuration list; Step S42: sort the lists of multiple parameter configuration values ​​that meet the conditions according to the required parameters submitted by the user, and select the best set of parameters.

[0012] Preferably, in step S6, the expanded data set is collected from one or more groups of optimal parameter configuration lists issued by the remote server and applied to local cluster job running data.

[0013] The present invention also provides a parallel computing application runtime optimization system based on remote intelligent services, including a local cluster and a remote server; The local cluster includes: a feature extraction module, a request access module, a resource configuration module and a scheduling system module; wherein the feature extraction module is used to extract feature description parameters related to the parallel computing tasks of the application job to be run; the request access module is used to send the application job and the local cluster features to the remote server and receive the parameter configuration value list issued by the remote server; the resource configuration module is used to complete the screening and sorting of the parameter configuration values ​​from the parameter configuration value list according to the priority set by the user and the cluster, and obtain the optimal parameter configuration to be run; the scheduling system module is used to manage and schedule the actual operation of the application job; The remote server includes: a parameter search module and a historical job database; wherein the parameter search module generates parameter configuration suggestions for the application job to be run based on the received application job and local cluster characteristics; the historical job database is used to store and expand job running data: providing training data for the parameter search module.

[0014] Preferably, the parameter search module includes a plurality of parameter search submodules, which respectively process parameter optimization problems corresponding to different types of applications and regularly update the parameter search module based on the updated data set.

[0015] Preferably, the optimization system supports parameter optimization services for multiple clusters sharing a remote server.

[0016] Therefore, the present invention proposes a parallel computing application runtime optimization method and system based on remote intelligent services, and its beneficial effects are as follows: (1) The runtime optimization method and system for parallel computing applications based on remote intelligent services proposed in the present invention optimizes the parameters related to parallel efficiency by combining a local feature extraction module and a remote historical job database. The local feature extraction module extracts features based on specific jobs and the current cluster hardware configuration to collect local accurate personalized configuration information. The remote parameter search module includes multiple modules that process commonly used software corresponding parameters respectively. The remote high-performance server provides a robust and easy-to-use parameter search model to predict appropriate hardware parameters. The present invention provides users with a more comprehensive remote service and a parameter feature extraction module that can set corresponding preferences.

[0017] (2) The present invention improves the utilization rate of job running results on the supercomputing platform, and advocates the construction and expansion of job running data sets that are truly available under each application category. A complete job run requires multiple links such as hardware resource allocation, job scheduling system scheduling, and actual operation, which may involve queuing and waiting, which further increases the operating cost. There are frequent and massive job applications on the cluster platform, and these operating data cannot be fully utilized. At the same time, the operating cost under each successful parameter configuration has a low conversion rate for the next operating experience. Corresponding to the large amount of underestimated idle data is the lack of high-quality job running data sets for specific applications in reality. By requesting access to the remote parameter search model, not only can the parameters of the parameter search model be continuously updated, but also the corresponding data set can be stably expanded. In the framework of the present invention, the operating data of various applications are collected in the historical job database established remotely and provided to the model training and learning of the corresponding application. The realization of basic optimization data sharing is conducive to saving the cost of parallel parameter optimization of each distributed cluster and better collecting operating data.

[0018] (3) The system proposed in this invention can better leverage the potential of high-performance platforms to provide user services. Users submit job features to remote servers through local clusters, and multiple clusters share the parameter optimization services of remote servers. The high-performance platform can significantly improve the utilization rate of remote server clusters by processing parameter prediction applications with high concurrency.

[0019] (4) The present invention assists users on the supercomputing center cluster to complete the setting of the number of hardware resources necessary for job running and the estimation of the number of parallel parameters. That is, users do not need to touch the parameters on the cluster software that are strongly related to hardware parallelism, which greatly simplifies the job running process and reduces the learning cost of users using supercomputing software. Under the existing parameter search experience and methods, the parameter values ​​specified by the user are not necessarily accurate and reliable. The remote parameter search module in the present invention is able to mine the parameters related to job parallelism, the number of hardware resource allocations and the hidden features of cluster hardware in the running data based on a large amount of data, which greatly saves users the time to adjust parameters and various overheads of job running. Furthermore, it helps users to automatically set the number of cores efficiently while meeting the budget, and enables users to run jobs without hardware perception, thereby better promoting scientific research work based on high-performance service platforms.

[0020] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a flow chart of the runtime optimization method of parallel computing application based on remote intelligent service of the present invention; Figure 2It is a timing diagram of the runtime optimization method of parallel computing application based on remote intelligent service in Example 1; Figure 3 A state diagram of a runtime optimization system for a parallel computing application based on remote intelligent services in the present invention; Figure 4 It is a schematic diagram of the connection between the cluster and the remote server in the present invention. DETAILED DESCRIPTION

[0022] In order to make the technical solutions, advantages and purposes of the present invention clearer, the technical solutions of the embodiments of the present invention are clearly and completely described below. The described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the described embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the protection scope of this application.

[0023] Unless otherwise defined, technical or scientific terms used in the present invention shall have the common meanings understood by one having ordinary skills in the field to which the present invention belongs.

[0024] Embodiment 1 like Figure 1-2 As shown, the present invention proposes a parallel computing application runtime optimization method based on remote intelligent service, comprising the following steps: SA1. The local cluster obtains the application job submitted by the user.

[0025] SA2. The local cluster obtains the application category in the application job information, and estimates the feature input information with the running job based on the application category. The feature input information includes cluster features and job features, where the cluster feature corresponds to the local hardware environment at the current moment, and the job feature corresponds to the input parameters and their descriptions in the application job that are unrelated to the parallel configuration of the overall hardware resource quantity of the job, specific subtask resource parameters, etc.; specifically, the job characteristics are parameters that affect the computing overhead such as the amount of computing data, the number of computing tasks, and the number of communications. By checking the input parameters of the computing program and the intermediate parameters obtained during the pre-processing of the computing program, it is verified whether it has an impact on the computing overhead, and finally a list of feature parameters is obtained and its values ​​are obtained.

[0026] SA3. The local cluster uploads the feature input information of the job to be run and applies for remote parameter optimization service, and obtains multiple sets of parameter estimation information to form a parameter configuration value list. Each set of information contains the overall hardware resource quantity of the job and the resource parameter assignment of a specific subtask.

[0027] SA4. The remote server forwards the parameter configuration value list of the job to be run to the local cluster.

[0028] SA5. The local cluster receives the parameter configuration value list, obtains the optimal parameter combination based on the screening and sorting method defined by the user and the cluster, and modifies the configuration of the job to be run.

[0029] SA6. The local cluster actually runs the job according to the optimized parameter configuration and feeds the operation data back to the remote historical job database.

[0030] In step SA2, the local hardware environment refers to the hardware configuration unique to the local hardware, such as the number of nodes, the number of cores and their corresponding relationships. The overall resource quantity of the job refers to the estimated hardware allocation value of the job to be run, such as the total number of CPU cores and GPUs requested when the job is running. The specific subtask resource parameter configuration refers to the parameter settings related to parallelization in specific jobs of different application types, which must be determined based on the aforementioned overall resource quantity of the job.

[0031] Embodiment 2 like Figure 3 As shown, the present invention also proposes a parallel computing application runtime optimization system based on remote intelligent services, including a local cluster and a remote server.

[0032] The local cluster includes: feature extraction module, request access module, resource configuration module and scheduling system module, among which, feature extraction module: sub-modules are divided according to specific application categories, used to extract feature representations of jobs to be run and local cluster hardware. The input of this module is two parts: job parameter descriptions that extract job features and are not directly related to job parallelization, and local necessary hardware environment configurations that extract cluster features. The output of this module is feature input information that can be used by the remote parameter search module; resource configuration module: used to complete the screening and sorting of parameter configuration values ​​from the parameter configuration value list according to the priorities set by the user and the cluster, and obtain the optimal parameter configuration to be run; request access module: used to send job and cluster features to the remote server and receive the parameter configuration value list issued by the remote server; scheduling system module: used to manage and schedule the actual operation of jobs; The parameter feature extraction model can also be trained using the historical job data stored locally, which includes the hardware operating environment, the original information of the job and its processing characteristics, and the actual resource assignment of the job. After the job is completed, the parameter feature module is responsible for collecting the necessary job information including the actual running time of the job, and optionally adding it to the local historical job database for timely training of the local feature module.

[0033] The remote server includes: a parameter search module and a historical job database, wherein the parameter search module generates parameter configuration suggestions for the job to be run based on the received job and cluster characteristics, wherein the ranking in the parameter list refers to the Er value of the quality judgment index; the historical job database is used to store and expand job running data and provide training data for the parameter search module.

[0034] The parameter search module contains multiple sub-modules, each of which is trained with the corresponding application job running data in the historical job database, and handles the parameter optimization problems corresponding to different types of applications. Since the remote server-side parameter search module lacks sufficient training due to the lack of data in the remote historical job database in the early stage, it can derive a list of estimated parameter configurations for the jobs to be run through the empirical model of the relevant application field, obtain the actual running data of the corresponding configuration through the trial operation of the local cluster, and supplement the data of the remote server-side historical job database.

[0035] like Figure 4 As shown, the parallel computing application runtime optimization method and system based on remote intelligent service provided by the present invention support multi-cluster shared remote cloud service parameter optimization.

[0036] Embodiment 3 In this embodiment, the SLURM job scheduling system, a common job scheduling system for supercomputing platforms, and the commonly used scientific computing software VASP are taken as examples to further explain a parallel computing application runtime optimization method and system based on remote intelligent services provided by the present invention. The specific operation steps are as follows: Step 1: The user submits a VASP application job to the SLURM job management system through the submission system. The submission system obtains the job input information and passes it to the feature extraction module. The job input information includes application type description, input file path, hardware resource application information, instruction sequence in the job, etc.

[0037] Step 2: The feature extraction module receives the job input information and assigns the job to the corresponding feature extraction submodule according to the application type description. The feature extraction submodule estimates whether different inputs and parameters in the previous processing process have an impact on the VASP job computational overhead, and then obtains influential parameters, such as the number of energy bands, the number of atoms, the number of reduced points for K-point calculation, the energy cutoff value, the number of transition state calculation mirrors, the size of the Fourier transform array, etc. as job features. At the same time, the module collects various hardware parameters of the local cluster at the current moment and uses them as input to generate similar local cluster feature expressions. The access request module sends the feature input information containing job features and cluster features to the remote parameter search module.

[0038] Step 3: After receiving the feature input information, the remote parameter search module outputs a parameter estimation list containing the overall resource quantity of the job and the job-specific subtask parameters. The specific process is to assign the job to the corresponding parameter search submodule for parameter prediction based on the application type description. For the specific subtask of the VASP job, the parameters that are highly correlated with the parallel operation efficiency are NPAR, KPAR, and NCORE. The three parameters follow the constraint that the product is the total number of cores allocated to the job. The parameter prediction submodule of the VASP job based on the historical job database predicts the hardware resource configuration values ​​of the job-specific subtask parameters and the overall resource quantity of the job for multiple sets of parallelization-related parameters.

[0039] Step 4: The local cluster receives the parameter configuration value list returned by the remote parameter search module, and the resource configuration module performs two processing steps: screening and sorting. The local cluster performs basic screening of parameter configuration values ​​based on the current cluster idle resources, and completes high-level screening of further parameter value configuration based on user permissions and preferences, obtains the optimal parameter combination, and modifies the job configuration to be run.

[0040] Step 5: The local cluster actually runs the job according to the optimized parameter configuration and feeds the running data back to the remote historical job database.

[0041] In this embodiment, the presence of other parameters in the VASP job input file will affect the accuracy of the three main parameters related to the parallel operation in step 2, such as the relationship between the number of K points NKPTS and the number of K point parallel calculations KPAR. The training feature extraction module can deeply explore the complex structure of job parameters and has higher efficiency and accuracy than experience selection.

[0042] In order to complete the training of the parameter search submodule in step 3, it is necessary to prepare a high-quality historical job running data set. When the construction of the remote historical job database is insufficient, the parameter search module will use the empirical model to generate a variety of parameter configuration recommendations for the local cluster to generate the running results of multiple test jobs, and add the important parameter search running information such as the actual running time of the remote cluster, the assignment of various parameters, the number of hardware resources, etc. to the remote database classified by application type, which is used to regularly update the corresponding submodule parameter search model. The sufficient database data refers to the number of running result feedbacks actually generated from the running of each local cluster to train the remote parameter search submodule until the error converges to a certain range. This part is set and adjusted manually. When the construction of the remote historical job library is sufficient, the parameter search module will use one or more parameter search methods to generate a certain number of parameter configuration values, feedback the list to the user through the access module, and collect the running results. In summary, the parameter prediction work can be set as parameter search, or different situations where the empirical model and parameter are mixed.

[0043] In specific implementation, the high-level screening of the parameter configuration value list in step 4 can be based on two methods: user permissions and user-specified job running requirements. The parameter configuration values ​​are screened according to user permissions, that is, whether the maximum number of available hardware resources allocated to the user's cluster is consistent with the recommended hardware resource allocation in the current parameter configuration list; the most suitable parameter configuration is screened according to the job running requirement parameters specified by the user, that is, the parameter screening preference that needs to be clearly stated before the user applies for job running. This method omits the actual user screening configuration process and combines this operation with automatic screening and user demands. The methods include but are not limited to "computing speed priority", "hardware efficiency priority", and "speed efficiency balance". Among them, "computing speed priority" means that under the same circumstances, it is inclined to allocate sufficient hardware resources to the running job to accelerate the computing process of the application job. The actual strategy is to select a set of search values ​​with a larger number of computing cores in the available parameter configuration value list. "Hardware efficiency priority" means that the cluster computing resources occupied by the task are as small as possible. The actual strategy is to select a configuration with a smaller number of computing cores in the available parameter list. "Speed ​​efficiency balance", as a compromise between the previous two screening methods, selects a parameter configuration with a middle number of computing cores to achieve the balance between controlling the number of allocated hardware resources and accelerating computing.

[0044] For the local cluster configuration of the specific job scheduling system in this embodiment, such as the SLURM scheduling system that allows the submitted job to set the core number range, the optimal hardware resource configuration value can be finally selected by comprehensively considering the top ranked configuration values.

[0045] In the event of a timeout in receiving a receipt message for a local cluster's application for a remote parameter search service and feedback of job running information, the local request access control module automatically terminates the parallel computing application running optimization process of the remote cloud service and promptly feeds back the exception to the user.

[0046] It is worth noting that the contents not elaborated in detail in the present invention are all prior art and are well known to those skilled in the art.

[0047] Therefore, the present invention provides a runtime optimization method and system for parallel computing applications based on remote intelligent services, which can optimize the operating parameters of parallel computing applications, reduce data redundancy, improve the utilization rate of supercomputing platforms, reduce user usage costs, and effectively promote scientific research on high-performance service platforms.

[0048] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.

Claims

1. A runtime optimization method for parallel computing applications based on remote intelligent services, characterized in that: Here are the steps: Step S1: parse the application job submitted by the user in the local cluster, obtain the application job and local cluster features, and push them to the remote server; Step S2: The remote server receives the application jobs and the processing applications of the local cluster features submitted by each local cluster, and allocates the corresponding parameter search submodule through the parameter search module to perform parameter prediction on the application jobs to obtain a parameter configuration value list; Step S3: The remote server sends the predicted parameter configuration value list to the local cluster, where the parameter configuration value list includes one or more groups of parameter configuration values. Step S4: The local cluster selects the most suitable parameter configuration through a custom method, feeds back the specific parameter configuration value to the user, modifies the number of hardware resources and application job parameter settings requested by the application job, and submits them to the local cluster scheduling system to complete the actual application job operation; Step S5: The local cluster monitors the running of the application job and uploads the actual running status result to the remote server, wherein the result includes the local actual running time, the category of the application job, the related parameters of the parallel running of the application job, and the characteristic values ​​of the application job and the hardware running environment parameters; Step S6: the data set in the historical operation database of the remote server is expanded. The remote server stores the received actual operation status result data into the corresponding application type data set, and periodically updates the parameter search module based on the updated data set.

2. The method for optimizing parallel computing applications at runtime based on remote intelligent services according to claim 1, characterized in that: In step S1, the application job characteristics are parameters that affect the computing overhead of the application job to be run, including the amount of computing data, the number of computing tasks and the number of communications; the local cluster characteristics are characteristic expressions after the environmental parameter configuration of the cluster is processed, and the environmental parameter configuration includes the hardware performance parameters of the CPU, GPU, accelerator card and the node or card interconnection network.

3. The method for optimizing parallel computing applications at runtime based on remote intelligent services according to claim 1, characterized in that: In step S1, the application job submitted by the user also includes the job running requirement parameters submitted by the user, and the requirement parameters include three modes: computing speed priority, hardware efficiency priority and speed efficiency balance.

4. The method for optimizing parallel computing applications at runtime based on remote intelligent services according to claim 1, characterized in that: In step S2, the parameter search module is divided into different types of parameter search submodules according to different application types and configured with different Er values. The parameter search submodule is trained based on the job data of the corresponding application in the historical job database, and the Er is a good or bad judgment indicator.

5. The method for optimizing parallel computing applications at runtime based on remote intelligent services according to claim 1, characterized in that: In step S3, the set of parameter configuration values ​​includes parallel parameters of subtasks in the application task and their values, and hardware resource parameters of the application job and their values.

6. The method for optimizing parallel computing applications at runtime based on remote intelligent services according to claim 1, characterized in that: In step S4, the custom method is to select the hardware parameters that are most suitable for the current local cluster to run the target application job from the parameter configuration list returned by the remote server. The hardware parameters are parameter configurations that maximize the use of idle resources without exceeding the number of idle resources in the current local cluster. The specific steps are as follows: Step S41: Select parameter list candidate values ​​that meet the application user's authority and the current operating conditions of the local cluster from the parameter configuration list; Step S42: sort the lists of multiple parameter configuration values ​​that meet the conditions according to the required parameters submitted by the user, and select the best set of parameters.

7. The method for optimizing parallel computing applications at runtime based on remote intelligent services according to claim 1, characterized in that: In step S6, the expanded data set is collected from one or more groups of optimal parameter configuration lists issued by the remote server and applied to the local cluster job running data.

8. A parallel computing application runtime optimization system based on remote intelligent services, comprising a local cluster and a remote server, characterized in that: The local cluster includes: a feature extraction module, a request access module, a resource configuration module and a scheduling system module; wherein the feature extraction module is used to extract feature description parameters related to the parallel computing tasks of the application job to be run; the request access module is used to send the application job and the local cluster features to the remote server and receive the parameter configuration value list issued by the remote server; the resource configuration module is used to complete the screening and sorting of the parameter configuration values ​​from the parameter configuration value list according to the priority set by the user and the cluster, and obtain the optimal parameter configuration to be run; the scheduling system module is used to manage and schedule the actual operation of the application job; The remote server includes: a parameter search module and a historical job database; wherein the parameter search module generates parameter configuration suggestions for the application job to be run based on the received application job and local cluster characteristics; the historical job database is used to store and expand job running data: providing training data for the parameter search module.

9. The parallel computing application runtime optimization system based on remote intelligent service according to claim 8, characterized in that: The parameter search module includes a plurality of parameter search submodules, which respectively process parameter optimization problems corresponding to different types of applications and regularly update the parameter search module based on the updated data set.

10. The parallel computing application runtime optimization system based on remote intelligent service according to claim 8, characterized in that: The optimization system supports parameter optimization services of multiple clusters sharing remote servers.

Citation Information

Patent Citations

  • Method and device for uniform configuration of carrier-class clustered applications

    CN103516538A

  • Job operation parameter optimization method applied to super-computing cluster scheduling

    CN114048027A

  • Configuration method and device of cluster host, electronic equipment and storage medium

    CN116149853A

  • Remote co-simulation platform architecture based on super computing cloud and design method

    CN117390841A

  • Parameter configuration method and device in multi-computing-cluster environment and computer equipment

    CN119597731A