A Serverless Parallel Computing Method and System for MPI
Through dynamic address mapping and parallel function management technical means of managing runtime environment, IP addressing, parallel function calls and differentiated parallel collaborative execution of MPI parallel computing in Serverless environment is solved, and efficient MPI parallel computing support is achieved.
Patent Information
- Application Number
- CN202210837029.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-15
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-07-15
AI Technical Summary
The existing Serverless technology cannot effectively support MPI parallel computing, and there are IP addressing problems, parallel function calls problems, and differentiated parallel collaborative execution problems.
Through dynamic address mapping technology, the mapping relationship between function names and network addresses is established, and the parallel function management runtime environment, parallel function address access mechanism and parallel function scheduling mechanism are constructed to achieve support for the MPI parallel computing model.
In the Serverless environment, MPI parallel computing is implemented, solving the problems of IP addressing, parallel function calls and differentiated parallel collaborative execution, and improving the execution efficiency and resource utilization of the application.
Smart Images

Figure CN115357375B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of parallel computing and Serverless (serverless computing), and particularly to a Serverless parallel computing method and system for MPI. Background Art
[0002] With the wide application of multi-core processors and cloud computing systems, parallel computing has become an important means to effectively utilize resources. In a distributed architecture, data exchange and task coordination and cooperation in parallel computing can be achieved through message passing. MPI (Message Passing Interface) is currently the most common parallel programming method and the main programming environment for distributed parallel systems. Applications developed based on the MPI programming model cover most of the workloads of HPC (High Performance Computing), and these workloads need to run on a single machine or a distributed cluster with an MPI runtime environment. Parallel computing applications can be successfully deployed and run in traditional cloud infrastructures (including Infrastructure as a Service IaaS). Compared with running on a traditional cluster, by leveraging the elasticity of cloud resources, applications can benefit from the rich computing power resources in the cloud environment, thereby achieving performance improvement. However, although cloud computing frees users from physical infrastructure management, what is left to them is a large number of virtual resources that need to be managed, and currently the cloud does not solve all the challenges of distributed computing. For example, when the computing scale and data volume change, the fixed resource allocation scheme cannot adapt to the changing load requirements, resulting in over-allocation or under-allocation of resources. Another example is that in the cloud environment, users need to manually maintain cluster resource managers (such as Slurm, Torque, etc.), and the operation and maintenance of applications become very complex, especially for containerized workloads.
[0003] In recent years, the emergence of Serverless Computing can, to a certain extent, solve the problems existing in the operation of parallel computing on traditional cloud infrastructures. Serverless is a new cloud computing model aiming to build an architecture that requires no management of infrastructures such as servers at runtime. Among them, Function as a Service (FaaS) is an important means to implement the Serverless architecture currently. In the cloud environment, the traditional method for users is to use the IaaS of cloud providers to provide virtual machines (VMs) and use them in a way similar to an internal cluster; while in the Serverless environment, Serverless combines the full elasticity of cloud resources with a maximally simplified programming model: users only need to program stateless functions, and the cloud takes on the responsibility of fine-grained scheduling for the invocation of these functions. In terms of ease of use, Serverless enables users to focus on business logic without having to consider infrastructures such as servers, operating systems, or file systems, runtimes, or even container management, maximizing zero operation and maintenance and reducing the technical entry threshold for users. In terms of performance, in the face of parallel computing of different scales, Serverless can achieve efficient orchestration of a large number of parallel functions and extreme scaling, so as to improve the execution efficiency of applications, and in the pay-as-you-go mode, elastically allocate computing resources under different parallel degrees to achieve the goal of maximizing resource utilization by allocating resources on demand.
[0004] The problems and disadvantages of the prior art are as follows:
[0005] Currently, Serverless still cannot effectively support the development of MPI parallel computing. Due to many inherent characteristics of the parallel programming model for MPI, such as the inter-process communication (IPC) based on MPI needs to be achieved through IP addressing, the execution logic of MPI parallel processes is different according to their identities, the execution of multiple nodes requires the support of the MPI runtime, etc., while native Serverless functions are dynamically generated and stateless, without network addresses for addressing, there are no differences between function replicas, and they cannot be executed simultaneously, etc. These limitations lead to the fact that the existing Serverless technologies cannot well adapt to traditional MPI applications.
[0006] In traditional Serverless technologies, in the face of similar problems, data can only be indirectly communicated between functions by passing through slow and expensive storage. Each function call request can only call one function replica at a time, and it is impossible to call all relevant parallel functions at once. Moreover, the function replicas cannot achieve differential parallel collaborative execution required for data parallelism and task parallelism. Summary of the Invention
[0007] The object of the present invention is to provide a Serverless parallel computing method and system for MPI, which support Serverless MPI parallel computing to solve the problems of IP addressing, parallel function call, and differential parallel cooperative execution in the Serverless environment.
[0008] To achieve the above object, on the one hand, the present invention provides a Serverless parallel computing method for MPI, including:
[0009] A dynamic address mapping step of establishing a mapping relationship between a function name and a corresponding network address, and uniformly programming ordinary functions and parallel functions of the Serverless parallel computing platform;
[0010] A computing model construction step of constructing a parallel function management runtime environment, a parallel function address access mechanism, and a parallel function scheduling mechanism in the Serverless parallel computing platform, and supporting the parallel computing model through a function replica mechanism, and establishing a corresponding relationship between each function in the function replica set and each parallel computing process to facilitate the execution of the parallel computing process; and
[0011] A parallel computing implementation step of implementing an MPI parallel computing process by the Serverless parallel computing platform.
[0012] In an embodiment of the Serverless parallel computing method for MPI, in the dynamic address mapping step, it further includes:
[0013] Constructing an address mapping table to establish a mapping relationship between a function name and a corresponding network address; and
[0014] Constructing an address mapping manager to manage the content of the address mapping table.
[0015] In an embodiment of the Serverless parallel computing method for MPI, in the step of constructing the address mapping manager, it further includes:
[0016] A mapping relationship update step of, when the content of the function replica is updated, the address mapping manager updating the content of the address mapping table according to the updated content of the function replica;
[0017] A mapping relationship query step of, when the parallel function replica runtime environment queries the address mapping relationship of the function replica, the address mapping manager querying the content of the address mapping table and encapsulating it into a specific message format;
[0018] Mapping relationship synchronization step. When the address mapping table is managed in a distributed form, each distributed component synchronizes the content mapped by the address mapping table; and
[0019] Mapping relationship recovery step. When the content of the address mapping table is lost, the address mapping manager restores the content of the address mapping table by invoking the Serverless parallel computing platform.
[0020] In the described Serverless parallel computing method for MPI, in one embodiment, the computing model construction step further includes:
[0021] Encapsulate the management of parallel functions, support lifecycle management operations for parallel functions, and convert the corresponding operations into management operations for function replicas;
[0022] Manage the identification of the addresses of function replicas, establish a two-layer structure for function replicas. The first layer is the identity identifier based on the function name, and the second layer is the identity identifier based on the function address; and
[0023] Parallel function scheduling matches the monitored cluster resource information with the scheduling requirements of parallel functions, and maximizes the affinity scores of each parallel function under the same parallel task when the resource requirements of each parallel function are met.
[0024] In the described Serverless parallel computing method for MPI, in one embodiment, the Serverless parallel computing platform includes a client, a control node, computing nodes, and a function repository. The control node includes a gateway, a controller, a monitor, and a scaling manager. The computing nodes include an executor and a container; the parallel computing implementation step further includes:
[0025] MPI parallel program functional development step:
[0026] (1) Write MPI function processing code according to the Serverless template with the MPI runtime environment, and configure the function runtime environment; and
[0027] (2) Submit the function for automated construction, package the MPI environment together as a container image, and push it to the image repository address configured in the function runtime environment in step (1) to facilitate any node in the cluster to download the image.
[0028] The described Serverless parallel computing method for MPI. In one embodiment, the Serverless parallel computing platform includes a client, a control node, computing nodes, and a function repository. The control node includes a gateway, a controller, a monitor, and a scaling manager. The computing nodes include an executor and containers. The steps for implementing the parallel computing further include:
[0029] Parallel function deployment service steps:
[0030] (1) The client submits a deployment request.
[0031] (2) The deployment request enters the request queue of the control node gateway and queues up.
[0032] (3) The control node gateway takes out the first deployment request from the request queue and hands it over to the controller for scheduling.
[0033] (4) The controller parses the scheduling request and queries whether the same task has been deployed on the cluster. If not, it obtains the cluster resource information collected by the monitor, sets thresholds for resource change degree and parallel request count, and jumps to step (6). If so, it notifies the scaling manager to perform scaling and transfers to step (5).
[0034] (5) If there is a situation of insufficient resources or a decrease in the parallelism of the function to be deployed, the scaling manager issues an instruction to scale down the parallel function to the controller. If there is an increase in the parallelism of the function to be deployed, the scaling manager issues an instruction to scale up the parallel function to the controller, and transfers to step (6).
[0035] (6) The controller makes a match based on the cluster resource information and deployment request information collected by the monitor. If a function needs to be deployed, when each parallel function of the task is scheduled to a computing node with sufficient resources, it maximizes the affinity score of the parallel functions under the same task, allocates the function to this group of computing nodes and allocates corresponding exclusive resources. If a function needs to be shut down, it shuts down the containers of the corresponding function in the startup order.
[0036] (7) The controller writes the scheduled function information into the address mapping table.
[0037] (8) The controller notifies all the computing nodes to which the parallel function is scheduled to download the function image, and each computing node downloads it from the image repository; and
[0038] (9) The controller starts the executor on all the computing nodes to which the parallel function is scheduled, and the executor creates a parallel function container to run on this node.
[0039] The described Serverless parallel computing method for MPI. In one embodiment, the Serverless parallel computing platform includes a client, a control node, computing nodes, and a function repository. The control node includes a gateway, a controller, a monitor, and a scaling manager. The computing nodes include an executor and a container. The steps for implementing the parallel computing further include:
[0040] Parallel function triggering service steps:
[0041] (1) The client sends a call request to the gateway of the control node;
[0042] (2) The gateway forwards the request with parameters to the controller;
[0043] (3) The controller obtains the address from the address mapping table and forwards it to the executors of the task;
[0044] (4) The executor parses the forwarded request parameters, concatenates them to form an MPI command, and passes it to all the parallel functions under the task to call their execution, while monitoring the running parallel functions in real time; and
[0045] (5) When the parallel function finishes running, the executor receives a signal indicating normal or abnormal exit of the function, and recovers the execution status and results.
[0046] On the other hand, the present invention also provides a Serverless parallel computing system for MPI, including:
[0047] A dynamic address mapping module that establishes a mapping relationship between function names and corresponding network addresses, and uniformly compiles ordinary functions and parallel functions of the Serverless parallel computing platform;
[0048] A computing model construction module that constructs a runtime environment for managing parallel functions, an address access mechanism for parallel functions, and a scheduling mechanism for parallel functions in the Serverless parallel computing platform, and supports the parallel computing model through a function replica mechanism, establishing a one-to-one correspondence between each function in the function replica set and each parallel computing process to facilitate the execution of parallel computing programs; and
[0049] A Serverless parallel computing platform that implements MPI parallel computing programs.
[0050] In one embodiment of the described Serverless parallel computing system for MPI, the dynamic address mapping module includes:
[0051] An address mapping table that establishes a mapping relationship between function names and corresponding network addresses; and
[0052] The address mapping manager manages the content of the address mapping table.
[0053] In one embodiment of the described MPI-oriented Serverless parallel computing system, the address mapping manager is further configured to:
[0054] When the content of the function copy is updated, update the content of the address mapping table according to the updated content of the function copy;
[0055] When the parallel function copy queries the address mapping relationship of the function copy during the running environment, query the content of the address mapping table and encapsulate it into a specific message format;
[0056] When the address mapping table is managed in a distributed form, each distributed component synchronizes the content mapped by the address mapping table; and
[0057] When the content of the address mapping table is lost, call the Serverless parallel computing platform to recover the content of the address mapping table.
[0058] In one embodiment of the described MPI-oriented Serverless parallel computing system, the computing model construction module is further configured to:
[0059] Encapsulate the management of parallel functions, support the management operations of the life cycle of parallel functions, and convert the corresponding operations into the management operations of function copies;
[0060] Manage the identification of the addresses of function copies, establish a two-layer structure for function copies, the first layer is the identity identification based on the function name, and the second layer is the identity identification based on the function address; and
[0061] The parallel function scheduling matches the monitored cluster resource information with the scheduling requirements of parallel functions, and maximizes the affinity scores of each parallel function under the same parallel task when the resource requirements of each parallel function are met.
[0062] In one embodiment of the described MPI-oriented Serverless parallel computing system, the Serverless parallel computing platform includes a client, a control node, a computing node, and a function repository. The control node includes a gateway, a controller, a monitor, and a scaling manager. The computing node includes an executor and a container; wherein,
[0063] The client is used for the relevant management of MPI parallel computing tasks and the initiation of computing operations;
[0064] A control node, which is used to respond to the runtime environment requirements of MPI parallel computing tasks, and is responsible for receiving deployment and invocation requests in Serverless computing, request queue management, container scheduling, address mapping, resource monitoring, scaling management, starting a parallel executor, and forwarding invocation requests;
[0065] A computing node, which is responsible for starting, monitoring, executing, and destroying functions in Serverless computing;
[0066] A function repository, which stores the basic container environment required for writing functions and the built function container images.
[0067] In one embodiment of the described Serverless parallel computing system for MPI, this Serverless parallel computing platform is also used for MPI parallel program functional development services:
[0068] (1) Write MPI function processing code according to a Serverless template with an MPI runtime environment, and configure the function runtime environment; and
[0069] (2) Submit the function for automated building, package the MPI environment into a container image, and push it to the image repository address configured in the function runtime environment in step (1) to facilitate any node in the cluster to download this image.
[0070] In one embodiment of the described Serverless parallel computing system for MPI, this Serverless parallel computing platform is also used for parallel function deployment services:
[0071] (1) The client submits a deployment request;
[0072] (2) The deployment request enters the request queue of the control node gateway for queuing;
[0073] (3) The control node gateway takes out the first deployment request from the request queue and hands it over to the controller for scheduling;
[0074] (4) The controller parses the scheduling request and queries whether the same task has been deployed on the cluster. If not, it obtains the cluster resource information collected by the monitor, sets thresholds for resource change degree and parallel request number, and jumps to step (6); if so, it notifies the scaling manager to perform scaling and transfers to step (5);
[0075] (5) If there is a situation of insufficient resources or a decrease in the parallelism of the function to be deployed, the scaling manager sends an instruction to reduce the parallel function to the controller. If there is a situation of an increase in the parallelism of the function to be deployed, the scaling manager sends an instruction to expand the parallel function to the controller, and transfers to step (6);
[0076] (6) The controller matches the cluster resource information and deployment request information collected by the monitor. If a function needs to be deployed, when each parallel function of the task is scheduled to a computing node with sufficient resources, the affinity score of the parallel functions under the same task is maximized, and the function is allocated to this group of computing nodes and corresponding exclusive resources are allocated; if a function needs to be shut down, the containers of the corresponding function are shut down in the startup order.
[0077] (7) The controller writes the scheduled function information into the address mapping table.
[0078] (8) The controller notifies all computing nodes to which the parallel function is scheduled to download the function image, and each computing node downloads from the image repository; and
[0079] (9) The controller starts an executor on all computing nodes to which the parallel function is scheduled, and the executor creates a parallel function container to run on this node.
[0080] In the described MPI-oriented Serverless parallel computing system, in one embodiment, the Serverless parallel computing platform is also used for parallel function triggering services:
[0081] (1) The client sends a call request to the gateway of the control node.
[0082] (2) The gateway forwards the request with parameters to the controller.
[0083] (3) The controller obtains the address from the address mapping table and forwards it to each executor of the task.
[0084] (4) The executor parses the forwarded request parameters, concatenates them to form an MPI command, and passes it to all parallel functions under the task to call their execution, and at the same time monitors the running parallel functions in real time; and
[0085] (5) When the parallel function finishes running, the executor receives a signal that the function exits normally or abnormally, and recovers the execution status and results.
[0086] Adopting the above technical solutions, the beneficial technical effects of the present invention are as follows:
[0087] (1) Realize dynamic address mapping of parallel functions, and support parallel functions to achieve network communication-based parallel collaboration based on function addresses.
[0088] (2) Support the computing requirements of the MPI programming model through the parallel function computing model, and realize writing and running MPI parallel programs in the Serverless environment.
[0089] (3) Support the full - life - cycle management and services of MPI parallel programs in the Serverless environment, including the development, deployment, and invocation of Serverless - enabled MPI parallel programs, etc. Brief Description of the Drawings
[0090] Figure 1 It is a management framework diagram for dynamic function address mapping;
[0091] Figure 2 It is an interaction diagram for the parallel function runtime environment;
[0092] Figure 3 It is a schematic structural diagram of the Serverless parallel computing platform for MPI of the present invention;
[0093] Figure 4 It is a logical view and deployment example diagram of the MPI parallel computing task of the present invention;
[0094] Figure 5 It is a schematic diagram of the MPI parallel function development service of the present invention;
[0095] Figure 6 It is a schematic diagram of the process in the parallel function deployment stage of the present invention;
[0096] Figure 7 It is a schematic diagram of the process in the parallel function triggering stage of the present invention. Detailed Embodiment
[0097] To support Serverless - enabled MPI parallel computing and solve problems such as IP addressing, parallel function invocation, and differential parallel collaborative execution in the Serverless environment, the present invention proposes a Serverless parallel computing method and system for MPI. The main key technologies of the present invention are as follows:
[0098] 1. Serverless Function Dynamic Address Mapping Technology
[0099] The present invention proposes a Serverless function dynamic address mapping technology. This technology proposes a function dynamic address mapping table mechanism, which establishes the mapping relationship between the function name and its network address, uniformly compiles ordinary functions and parallel functions in the Serverless computing platform, supports the Serverless computing platform to dynamically obtain its network address through the function name, and the two - way mapping of obtaining the function name through the network address.
[0100] This technology consists of an address mapping manager and an address mapping table, and provides an API for accessing address mapping table information. The address mapping manager writes the scheduled function information of the computing platform into the address mapping table. This address mapping table supports the mapping and query of basic information and address information of functions. The basic information of a function may include, but is not limited to, function name, function copy information, etc. When a function exists in the form of a normal copy or a parallel function copy, the address information corresponding to its function name is an address set, including the address information of each copy, where the address information may include, but is not limited to, function IP address, service port, etc. The information in the address mapping table can be stored in various forms, such as cached in memory in the key-value form. The address mapping manager can also quickly restore the content in the address mapping table through the Serverless computing platform API to prevent the impact caused by memory loss. The APIs provided by this technology include operations such as querying address table information, obtaining function addresses based on function names, and obtaining function names based on function addresses. By calling the API query service of the function address, all the address information corresponding to the function can be obtained simultaneously, which can support the parallel collaboration of function parallel computing.
[0101] Adopting the Serverless function dynamic address mapping technology, the present invention can support obtaining function addresses according to function names, and support subsequent parallel collaboration based on the obtained parallel function addresses for parallel computing.
[0102] The present invention designs and implements a dynamic address mapping technology for functions, and uses it to improve the Serverless copy mechanism to support synchronous calls and differential executions of parallel functions, realizing Serverless computing for the MPI parallel programming model.
[0103] 2. Serverless Parallel Function Copy Calculation Model
[0104] The present invention proposes a Serverless parallel function copy calculation model. The Serverless parallel function copy calculation model proposes to build a parallel function management runtime environment, a parallel function address access mechanism, and a parallel function scheduling mechanism in the Serverless computing platform. It can support the parallel computing model through the function copy mechanism, establish a one-to-one correspondence between each function in the function copy set and each parallel computing process, so that each parallel computing process can use the function to implement operations such as deployment and expansion.
[0105] The Serverless parallel function copy calculation model builds a parallel function management runtime environment, which encapsulates the management of parallel functions, supports lifecycle management operations such as creating, deploying, and accessing addresses for parallel functions, and converts the corresponding operations into management operations of function copies.
[0106] The Serverless parallel function replica computing model manages the identification of the addresses of function replicas. Compared with the flat structure of the ordinary function replica model, the Serverless parallel function replica computing model establishes a two-layer structure in the function replica. The first layer is the identity identification based on the function name, and the second layer is the identity identification based on the function address. Among them, the mapping of the second-layer address can use the Serverless function dynamic address mapping technology to achieve address access.
[0107] The Serverless parallel function replica computing model provides optimized support for parallel function scheduling, and can achieve affinity scheduling of parallel computing tasks, etc. Parallel function scheduling matches the monitored cluster resource information with the scheduling requirements of parallel functions. On the premise of meeting the resource requirements of each parallel function, it maximizes the affinity scores of each parallel function under the same parallel task, and realizes the maximum aggregation in scheduling.
[0108] Adopting the Serverless parallel function replica computing model can support the computing requirements of parallel processes, and management operations can be realized through parallel function replicas, so that parallel programs can run in the Serverless environment and improve the parallel computing efficiency through scheduling technology.
[0109] 3. Serverless Parallel Computing System for MPI
[0110] The present invention proposes a Serverless parallel computing system for MPI. The Serverless parallel computing system for MPI applies the Serverless parallel function replica computing model to realize the support for MPI parallel computing programs. Among them, it includes the following modules:
[0111] Module 1, Client. It is used for the relevant management of MPI parallel computing tasks and the initiation of computing operations, including task startup, task viewing, task suspension, task resume, task stop, parallel function call, etc. The client supports three forms: Web interface, command-line tool, and RESTful API for corresponding operations and viewing operation results.
[0112] Module 2, Control Node. It responds to the runtime environment requirements of MPI parallel computing tasks and is responsible for receiving deployment and call requests in Serverless computing, request queue management, container scheduling, address mapping, resource monitoring, scaling management, starting parallel executors, and forwarding call requests. The control node mainly includes four components: gateway, controller, monitor, and scaling manager. Among them:
[0113] (1) The gateway is responsible for receiving client requests and forwarding requests to the controller according to the request type.
[0114] (2) The controller integrates the Serverless parallel function replica computing model and its runtime environment module, manages function scheduling, address mapping, and the start / stop and invocation of tasks of the parallel executor. Function scheduling matches the cluster resource information collected by monitoring with the MPI parallel function scheduling requests, and performs scheduling according to the system scheduling policy, such as parallel function affinity scheduling; address mapping maps functions to address information to achieve dynamic address mapping of Serverless functions; the parallel executor supports interaction with the MPI runtime environment, and its start / stop and invocation of tasks are that the controller starts or pauses the execution of the task on the scheduled node according to the scheduling result, and can forward the request to each parallel executor under the task after receiving the invocation request.
[0115] (3) The monitor is responsible for monitoring the resource information and request information of the entire cluster, and sets thresholds for the degree of resource change and the number of parallel requests. Once the cluster resource change exceeds the threshold, it notifies the scale-out / scale-in manager to perform scale-out / scale-in.
[0116] (4) The scale-out / scale-in manager is responsible for specifically executing the scale-out / scale-in of parallel functions. In case of insufficient resources or a decrease in the parallelism of the requested deployed functions, it sends an instruction to the controller to scale in the parallel functions. In case of an increase in the parallelism of the requested deployed functions, it sends an instruction to the controller to scale out the parallel functions. The controller is responsible for shutting down or starting the corresponding parallel function containers according to the scheduling policy.
[0117] Module 3, computing node. It is responsible for the start, monitoring, execution, and destruction of each parallel function in Serverless computing, including two parts: the executor and the container. The executor is responsible for receiving the start / stop and invocation requests from the controller, and managing the start / stop or invocation of functions belonging to the same task on this node. These functions will be started, stopped, or invoked synchronously. When invoked, it will also parse the forwarded request parameters and use MPI commands to pass them to the parallel function to call its execution, and at the same time monitor the running function containers in real time; the container is the runtime environment of the function, where the environment variables in the deployment request are configured, and it is the place responsible for receiving the function request parameters and executing the function.
[0118] Module 4, function repository. It is a remote storage facility that stores the basic container environment required for writing functions and the built function container images. After starting the container during the parallel function deployment phase, each container will download the corresponding function image from the function repository as the object to be started.
[0119] The system includes the following key services: MPI parallel program functional development, parallel function deployment, and parallel function triggering.
[0120] Specifically, the MPI parallel program functional development service includes the following steps:
[0121] (1) The user writes the MPI function processing code according to the Serverless template with the MPI runtime environment and configures the function runtime environment.
[0122] (2) Submitting the function for automated building will package the MPI environment together as a container image and push it to the image repository address configured in the function runtime environment in (1), facilitating any node in the cluster to download the image.
[0123] Specifically, the parallel function deployment service includes the following steps:
[0124] (1) The user submits a deployment request from the client and can choose any one of the three forms: Web interface, command-line tool, and RESTful API for request invocation.
[0125] (2) The deployment request enters the request queue of the control node gateway and is queued in FIFO (First In First Out).
[0126] (3) The control node gateway takes out the first deployment request from the request queue and hands it over to the controller for scheduling.
[0127] (4) The controller parses the scheduling request and queries whether the same task has been deployed on the cluster. If not, it obtains the cluster resource information collected by the monitor, sets the thresholds for resource change degree and parallel request count, and jumps to step (6); if so, it notifies the scaling manager to perform scaling and transfers to step (5).
[0128] (5) If there is a situation of insufficient resources or a decrease in the parallelism of the function to be deployed, the scaling manager sends an instruction to scale down the parallel function to the controller. If there is an increase in the parallelism of the function to be deployed, the scaling manager sends an instruction to scale up the parallel function to the controller and transfers to step (6).
[0129] (6) The controller matches according to the cluster resource information and deployment request information collected by the monitor. If a function needs to be deployed, on the premise that each parallel function of the task is scheduled to a computing node with sufficient resources, it maximizes the affinity score of the parallel functions under the same task, allocates the function to this group of computing nodes and allocates corresponding exclusive resources; if a function needs to be shut down, it shuts down the corresponding function containers in the startup order.
[0130] (7) The controller writes the scheduled function information into the address mapping table, which uniformly numbers the ordinary functions and parallel functions. The address mapping table caches the obtained function address information and the IP addresses of the computing nodes in the form of key-value in the memory. If memory loss occurs, all function address information is quickly restored from each node through the cluster API and the IP addresses of the computing nodes are remapped to restore the address mapping table;
[0131] (8) The controller notifies all the computing nodes to which the parallel function is scheduled to download the function image, and each computing node downloads it from the image repository;
[0132] (9) The controller starts the executor on all the computing nodes to which the parallel function is scheduled, and the executor creates the parallel function containers running on this computing node.
[0133] Specifically, the parallel function trigger service includes the following steps:
[0134] (1) The user sends a call request to the control node gateway through the client, and can choose any one of the three forms: Web interface, command line tool, and RESTful API for the request call;
[0135] (2) The gateway forwards the request with parameters to the controller;
[0136] (3) The controller obtains the addresses from the address mapping table and forwards them to the executors of this task;
[0137] (4) The executor parses the forwarded request parameters, splices them to form an MPI command, and passes it to all the parallel functions under this task to call their execution, and at the same time monitors the running parallel functions in real time;
[0138] (5) When the parallel function finishes running, the executor receives the signal of normal or abnormal exit of the function, and recovers the execution status and results.
[0139] By adopting the above steps, the present invention can support the full life cycle management of the Serverless MPI program.
[0140] Next, the technical solution of the present invention will be further described in conjunction with the drawings of the present invention. The described embodiments are part of the embodiments of the present invention, rather than all the embodiments.
[0141] For those skilled in the art, some well-known technologies may not be elaborated in detail.
[0142] The implementation process of the present invention includes building a Serverless function dynamic address mapping module, building a Serverless parallel function replica runtime environment, building a Serverless parallel computing system for MPI, building an MPI parallel function development service, building a parallel function deployment service, and building a parallel function trigger service.
[0143] 1. Build a Serverless function dynamic address mapping module
[0144] Building a Serverless function dynamic address mapping module includes building an address mapping manager and an address mapping table. The address mapping manager is responsible for managing the update and query of the address mapping table content, and responding to and serving the function dynamic address access requirements of other modules in the system, as Figure 1 shown. The main functions of the address manager include:
[0145] (1) Mapping relationship update. When the content of the function replica is updated, the address mapping manager updates the content of the address mapping table according to the updated content of the replica, such as adding, modifying, or deleting the corresponding function addresses in the function address set of the updated function replica.
[0146] (2) Mapping relationship query. When the parallel function replica runtime environment queries the address mapping relationship of the function replica, the address mapping manager queries the content of the address mapping table, encapsulates it into a specific message format, and returns it to the requesting module.
[0147] (3) Mapping relationship synchronization. When the function address mapping table is managed in a distributed form, each distributed component can synchronize the content of the address table.
[0148] (4) Mapping relationship recovery. When the content of the function address mapping table is lost due to reasons such as failures, the address mapping manager can restore the content of the address mapping table by calling the function replica management API of the computing platform.
[0149] 2. Build a Serverless parallel function replica runtime environment
[0150] The parallel function replica runtime environment responds to the parallel function replica management requirements of the Serverless parallel computing system for MPI, controls and manages the parallel function replica and the parallel executor, and accesses and updates the dynamic function address through the dynamic function address mapping manager, as Figure 2 shown. The parallel function runtime environment converts the management commands of the parallel function replica into relevant operations of ordinary function replicas, realizing the parallel management of ordinary function replicas.
[0151] 3. Build a Serverless parallel computing system for MPI
[0152] The Serverless parallel computing system for MPI can run on top of a physical machine or a virtual machine cluster. It can use the operating system, the system-level container runtime environment, and the container orchestrator as the basic support. The platform architecture is as follows: Figure 3 As shown below. The construction content of each module is as follows:
[0153] Module 1, Client. The client is built based on the Web interface, command line, and Restful API. Both the Web and the command line can be encapsulated based on the Restful API. The functions provided by the client include the management of parallel computing tasks and the initiation of computing operations, including task startup, task viewing, task suspension, task resume, task stop, parallel function call, etc.
[0154] Module 2, Control Node. In the control node, request queue management, container scheduling, address mapping, resource monitoring, scaling management, and parallel executor management are built, including four components: gateway, controller, monitor, and scaling manager. Among them, the gateway is responsible for receiving requests sent by the client and initiating a scheduling request or forwarding the request to the controller according to the request type. The controller is responsible for function scheduling, address mapping, and the start, stop, and call of the executor. The scheduling is to match the cluster resource information and scheduling request information collected by the monitor. On the premise that each parallel function of the task is scheduled to a computing node with sufficient resources, the affinity score of the parallel functions under the same task is maximized, and the function is assigned to this group of computing nodes and the corresponding exclusive resources are allocated; the address mapping is to map the function with the information after scheduling and cache the result of the dynamic scheduling of the function; the start, stop, and call of the executor are that the controller starts or pauses the executor of the task on the scheduled node according to the scheduling result, and can forward the request to each executor under the task after receiving the call request. The monitor is responsible for monitoring the resource information and request information of the entire cluster, and setting the thresholds of the resource change degree and the number of parallel requests. Once the cluster resource change exceeds the threshold, it notifies the scaling manager to perform scaling. The scaling manager is responsible for specifically executing the scaling of the parallel function. If there is a situation of insufficient resources or a decrease in the parallelism of the requested deployed function, it sends an instruction to the controller to scale down the parallel function. If there is a situation of an increase in the parallelism of the requested deployed function, it sends an instruction to the controller to scale up the parallel function. The controller is responsible for closing or starting the corresponding function containers according to the scheduling strategy.
[0155] Module 3, Computing Node. It is responsible for the startup, monitoring, execution, and destruction of functions in Serverless computing, including two parts: an executor and a container. The executor is responsible for receiving the startup, stop, and call requests from the controller, and starting, stopping, or calling the functions in all containers belonging to the same task in the same node below. These functions will be started, stopped, or called synchronously. When calling, it will also parse the forwarded request parameters and use the MPI command to pass them to the parallel function to call its execution, and at the same time, monitor the running function containers in real time; the container is the runtime environment of the function, where the environment variables in the deployment request are configured, and it is the place responsible for receiving the function request parameters and executing the function.
[0156] Module 4, Function Repository. It is a remote storage container that stores the basic container environment required for writing functions and the built function container images. After starting the container during the parallel function deployment phase, each container will download the corresponding function image from the function repository as the startup runtime environment.
[0157] The logical view and deployment example of the MPI parallel computing task are as Figure 4 . This example shows 3 tasks scheduled on 2 physical nodes. Among the 3 tasks, Task 1 has 2 parallel functions, Task 2 has 6 parallel functions, and Task 3 has 1 function, which is a single-function task. The scheduling result is that Task 1 and Task 2 are deployed on Computing Node 1, and Task 2 and Task 3 are deployed on Computing Node 2. Task 2 is a task deployed across two computing nodes, and 3 parallel functions are deployed on each of the two computing nodes.
[0158] The MPI parallel computing task has the following two constraints: all functions are executed simultaneously; after all functions are executed, the task is considered to end.
[0159] According to another aspect of the present invention, the present invention proposes a Serverless parallel computing method for MPI, which includes three stages in the implementation process: function writing and building, pushing the image, parallel function deployment, and parallel function triggering.
[0160] 4. Build the MPI parallel function development service
[0161] In an embodiment, as Figure 5 shown, the steps of function writing and building, pushing the image are as follows:
[0162] Step 501: The user writes the MPI function processing code according to the Serverless template with the MPI runtime environment and configures the function runtime environment;
[0163] Step 502: Submitting the function for automated building will package the MPI environment into a container image and push it to the image repository address configured in the function runtime environment in Step 501, facilitating the download of the image by any node in the cluster.
[0164] Further, the runtime environment configuration described in Step 501 includes the function name, the function processing code path, and the image repository address.
[0165] 5. Building a parallel function deployment service
[0166] In one embodiment, as Figure 6 shown, the steps of parallel function deployment are as follows:
[0167] Step 601: The user submits a deployment request from the client, and can choose any one of the three forms: Web interface, command-line tool, and RESTful API for request invocation;
[0168] Step 602: The deployment request enters the request queue of the control node gateway and is queued in FIFO (First In First Out);
[0169] Step 603: The control node gateway takes out the first deployment request from the request queue and hands it over to the controller for scheduling;
[0170] Step 604: The controller parses the scheduling request and queries whether the same task has been deployed on the cluster. If not, it obtains the cluster resource information collected by the monitor, sets the thresholds for resource change degree and the number of parallel requests, and jumps to Step 606; if so, it notifies the scaling manager to perform scaling and transfers to Step 605;
[0171] Step 605: In case of insufficient resources or a decrease in the parallelism of the function to be deployed, the scaling manager sends an instruction to scale down the parallel function to the controller. In case of an increase in the parallelism of the function to be deployed, the scaling manager sends an instruction to scale up the parallel function to the controller, and transfers to Step 606;
[0172] Step 606: The controller makes a match based on the cluster resource information and the deployment request information collected by the monitor. If a function needs to be deployed, on the premise that each parallel function of the task is scheduled to a computing node with sufficient resources, the affinity score of the parallel functions under the same task is maximized, and the function is assigned to this group of computing nodes and the corresponding exclusive resources are allocated; if a function needs to be shut down, the corresponding function containers are shut down in the startup order.
[0173] Step 607: The controller writes the scheduled function information into the address mapping table, which uniformly numbers the ordinary functions and parallel functions. The address mapping table caches the obtained function address information and the computing node IP addresses in the form of key-value in the memory. If memory loss occurs, all function address information is quickly restored from each node through the cluster API and the computing node IP addresses are remapped to restore the address mapping table;
[0174] Step 608: The controller notifies all the computing nodes to which the parallel function is scheduled to download the function image, and each node downloads it from the image repository;
[0175] Step 609: The controller starts an executor on all the computing nodes to which the parallel function is scheduled, and the executor creates a parallel function container to run on this node.
[0176] Further, the deployment request described in step 601 includes the gateway address of the specified running Serverless platform, the image repository address, the function name, the MPI parallelism, and the environment variables required by the application.
[0177] Further, the scheduling strategy described in step 606 includes the First Fit algorithm, the Best Fit algorithm, or a user-defined scheduling algorithm.
[0178] Further, the calculation formula for the parallel function affinity score under the same task described in step 606 is:
[0179]
[0180]
[0181]
[0182]
[0183]
[0184]
[0185]
[0186] Among them, the formula variables are as follows:
[0187] 1) Two minimization objectives, which are respectively to minimize the number of loaded nodes and the cost of task partitioning.
[0188] 2) i, i′: represent any two functions under the same task, and a task consists of n functions.
[0189] 3) j, j′: Represent any two nodes in the cluster, and the cluster consists of N nodes.
[0190] 4) E: Represents the communication relationship between functions. If {i, i′} ∈ E, it means that function i and function i′ have a communication relationship, otherwise they do not.
[0191] 5) A task consists of n parallel functions that require resources of The resource requirement of function i can be represented by a d-dimensional vector denoted as.
[0192] 6) The cluster has N nodes with a resource capacity of The capacity of node j can be represented by a d-dimensional vector denoted as.
[0193] 7) x ij : If function i is assigned to node j, then x ij = 1, otherwise x ij = 0.
[0194] 8) y j : If there is a task function already deployed on node j, then y j = 1, otherwise y j = 0.
[0195] 9) z ib : If function i is in subtask b, then z ib = 1, otherwise z ib = 0.
[0196] 10) z i′b : If function i′ is in subtask b, then z i′b = 1, otherwise z i′b = 0.
[0197] 11) e i,i′ : If functions i and i′ belonging to the same task are on the same node, then the variable value e i,i′ = 0, otherwise e i,i′ = 1.
[0198] In the formula, s.t. refers to the constraint conditions, which satisfy the two objectives (a) and (b) under the conditions (c) to (g). The objectives are to minimize the number of loaded nodes and the cost of task partitioning, which are represented by formula (a) and formula (b). Constraint formula (c) states that any function of a task should and must be deployed on one node, while formula (d) states that the sum of the function resource requirements on a node should not exceed the available capacity of the node, ensuring that resources are not over-allocated. Formulas (e) and (f) ensure the effective partitioning of a task, and formula (g) ensures that each function is exactly allocated to one partition (i.e., subtask). Finally, a set (x ij , y j ) can be obtained to match all functions of the task with the nodes in the cluster.
[0199] Furthermore, the function address information described in step 607 is the address customized by the address mapping table, which consists of the binary tuple <task ID, container ID>. The task ID is generated by the controller, and different tasks have different function IDs. For ordinary functions, since there is only one function under a task, the task ID can distinguish the function; for parallel functions, the task IDs of the functions belonging to the same task are the same, so it is necessary to rely on the binary tuple <function ID, container ID> to uniquely determine. Such an ID design can ensure quick traversal of parallel functions belonging to the same task.
[0200] Furthermore, considering the possibility that the task may run across nodes during the startup of the executor described in step 608, one executor will be started separately under all nodes where the task is scheduled. This executor is responsible for managing all parallel functions of the task on this node, including creation, passing parameters for execution, running monitoring, and result recovery, etc.
[0201] Furthermore, the creation of the parallel function container described in step 608 is to reduce the container cold start time when the function is triggered. After the container is created, it provides services externally in the role of an HTTP server.
[0202] 6. Build a parallel function trigger service
[0203] In an embodiment, as Figure 7 shown, the steps of triggering the parallel function are as follows:
[0204] Step 701: The user sends a call request to the control node gateway through the client, and can choose any one of the three forms: Web interface, command line tool, and RESTful API for the request call;
[0205] Step 702: The gateway forwards the request with parameters to the controller;
[0206] Step 703: The controller obtains the address from the address mapping table and forwards it to each executor of the task;
[0207] Step 704: The executor parses the forwarded request parameters, splices them to form an MPI command, and passes it to all parallel functions under the task to call their execution, while monitoring the running parallel functions in real time;
[0208] Step 705: When the parallel function finishes running, the executor receives a signal indicating normal or abnormal exit of the function, and recovers the execution status and result.
[0209] Further, the call request described in step 701 includes the gateway address of the specified running Serverless platform and call parameters;
[0210] Further, the splicing to form an MPI command described in step 704 includes splicing the IP addresses, port numbers, and container network card information of all parallel functions under the task, generating a command to start the MPI task, and starting the parallel functions based on this;
[0211] Further, after the parallel function described in step 705 finishes execution, if one of the parallel functions under the task exits abnormally, the executor determines that the task fails and returns an error status to the control node gateway; otherwise, it is successful and returns a success status to the control node gateway.
[0212] In view of many problems in implementing Serverless MPI parallel computing, the present invention proposes a Serverless parallel computing method and system for MPI. At the method level, it includes Serverless function dynamic address mapping technology and a parallel function replica computing model, which solve the problems of IP addressing, parallel function call, and parallel collaborative execution; at the system level, a Serverless parallel computing system for MPI is constructed, and the MPI programming model is transplanted into the Serverless environment, providing Serverless-enabled MPI runtime environment support. The method of the present invention can improve the existing Serverless platform to form a Serverless parallel computing system supporting MPI, enabling users to efficiently implement Serverless MPI parallel computing.
[0213] Of course, the present invention can also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and deformations according to the present invention, but these corresponding changes and deformations should all fall within the protection scope of the appended claims of the present invention.
Claims
1. A Serverless parallel computing method for MPI, characterized in that, Including: A dynamic address mapping step, establishing a mapping relationship between function names and corresponding network addresses, and uniformly programming ordinary functions and parallel functions of the Serverless parallel computing platform; A computing model construction step, constructing a parallel function management runtime environment, a parallel function address access mechanism, and a parallel function scheduling mechanism in the Serverless parallel computing platform, and supporting the parallel computing model through a function replica mechanism, establishing a correspondence between each function in the function replica set and each parallel computing process to facilitate the execution of the parallel computing process; And A parallel computing implementation step, implementing an MPI parallel computing process by the Serverless parallel computing platform; The computing model construction step further includes: Encapsulating the management of parallel functions, supporting lifecycle management operations on parallel functions, and converting corresponding operations into management operations of function replicas; Performing identity management on the addresses of function replicas, establishing a two-layer structure for function replicas, with the first layer being an identity identifier based on the function name and the second layer being an identity identifier based on the function address; and Parallel function scheduling matches the monitored cluster resource information with the scheduling requirements of parallel functions. When the resource requirements of each parallel function are met, the affinity scores of each parallel function under the same parallel task are maximized for maximum aggregation in scheduling; Among them, the mapping of the second-layer address uses the Serverless function dynamic address mapping technology to implement address access.
2. The Serverless parallel computing method for MPI according to claim 1, wherein In the dynamic address mapping step, it further includes: Constructing an address mapping table to establish a mapping relationship between function names and corresponding network addresses; and Constructing an address mapping manager to manage the content of the address mapping table.
3. The Serverless parallel computing method for MPI according to claim 2, wherein In the step of constructing the address mapping manager, it further includes: A mapping relationship update step. When the content of the function replica is updated, the address mapping manager updates the content of the address mapping table according to the updated content of the function replica; A mapping relationship query step. When the parallel function replica runtime environment queries the address mapping relationship of the function replica, the address mapping manager queries the content of the address mapping table and encapsulates it into a specific message format; A mapping relationship synchronization step. When the address mapping table is managed in a distributed form, each distributed component synchronizes the content mapped by the address mapping table; and A mapping relationship recovery step. When the content of the address mapping table is lost, the address mapping manager restores the content of the address mapping table by calling the Serverless parallel computing platform.
4. The Serverless parallel computing method for MPI according to claim 1, wherein The Serverless parallel computing platform includes a client, a control node, a computing node, and a function repository. The control node includes a gateway, a controller, a monitor, and a scaling manager. The computing node includes an executor and a container; The parallel computing implementation step further includes: An MPI parallel program functional development step: (1) Writing MPI function processing code according to a Serverless template with an MPI runtime environment and configuring the function runtime environment; and (2) Submit the function for automated building, package the MPI environment into a container image, and push it to the image repository address configured in the function runtime environment in step (1) to facilitate any node in the cluster to download the image.
5. The Serverless parallel computing method for MPI according to claim 1, characterized in that, The Serverless parallel computing platform includes a client, a control node, computing nodes, and a function repository. The control node includes a gateway, a controller, a monitor, and a scaling manager. The computing nodes include an executor and containers. The parallel computing implementation steps further include: Parallel function deployment service steps: (1) The client submits a deployment request. (2) The deployment request enters the request queue of the control node gateway for queuing. (3) The control node gateway takes out the first deployment request from the request queue and hands it over to the controller for scheduling. (4) The controller parses the scheduling request and queries whether the same task has been deployed on the cluster. If not, it obtains the cluster resource information collected by the monitor, sets the thresholds for resource change degree and parallel request number, and jumps to step (6); if so, it notifies the scaling manager to perform scaling and transfers to step (5). (5) If there is a situation of insufficient resources or a decrease in the parallelism of the function to be deployed, the scaling manager sends an instruction to shrink the parallel function to the controller. If there is an increase in the parallelism of the function to be deployed, the scaling manager sends an instruction to expand the parallel function to the controller and transfers to step (6). (6) The controller makes a match based on the cluster resource information and deployment request information collected by the monitor. If a function needs to be deployed, when each parallel function of the task is scheduled to a computing node with sufficient resources, maximize the affinity score of the parallel functions under the same task, allocate the function to the computing node and allocate corresponding exclusive resources; if a function needs to be shut down, shut down the containers of the corresponding function in the startup order. (7) The controller writes the scheduled function information into the address mapping table. (8) The controller notifies all the computing nodes to which the parallel function is scheduled to download the function image, and each computing node downloads it from the image repository; and (9) The controller starts the executor on all the computing nodes to which the parallel function is scheduled, and the executor creates a parallel function container to run on this node.
6. The Serverless parallel computing method for MPI according to claim 1, wherein The Serverless parallel computing platform includes a client, a control node, computing nodes, and a function repository. The control node includes a gateway, a controller, a monitor, and a scaling manager. The computing nodes include an executor and containers. The parallel computing implementation steps further include: Parallel function trigger service steps: (1) The client sends a call request to the gateway of the control node. (2) The gateway forwards the request with parameters to the controller. (3) The controller obtains the address from the address mapping table and forwards it to the executors of each task. (4) The executor parses the forwarded request parameters, concatenates them to form an MPI command, and passes it to all the parallel functions under this task to call their execution, and at the same time monitors the running parallel functions in real time; and (5) When the parallel function finishes running, the executor receives a signal that the function exits normally or abnormally, and recovers the execution status and results.
7. A Serverless parallel computing system for MPI, characterized in that Include: The dynamic address mapping module establishes the mapping relationship between function names and corresponding network addresses, and uniformly compiles ordinary functions and parallel functions of the Serverless parallel computing platform; The computing model construction module constructs a parallel function management runtime environment, a parallel function address access mechanism, and a parallel function scheduling mechanism in the Serverless parallel computing platform, and supports the parallel computing model through the function replica mechanism. A one-to-one correspondence is established between each function in the function replica set and each parallel computing process to facilitate the execution of parallel computing programs; and The Serverless parallel computing platform implements MPI parallel computing programs; The computing model construction module is further used for: Encapsulating the management of parallel functions, supporting lifecycle management operations on parallel functions, and converting corresponding operations into management operations of function replicas; Managing the identification of the addresses of function replicas, establishing a two-layer structure for function replicas. The first layer is the identity identification based on function names, and the second layer is the identity identification based on function addresses; and Parallel function scheduling matches the monitored cluster resource information with the scheduling requirements of parallel functions. When the resource requirements of each parallel function are met, the affinity scores of each parallel function under the same parallel task are maximized to achieve the greatest degree of aggregation in scheduling; Among them, the mapping of the second-layer address uses the Serverless function dynamic address mapping technology to achieve address access.
8. The Serverless parallel computing system for MPI according to claim 7, wherein The dynamic address mapping module includes: An address mapping table that establishes the mapping relationship between function names and corresponding network addresses; and An address mapping manager that manages the content of the address mapping table.
9. The Serverless parallel computing system for MPI according to claim 7, characterized in that, The address mapping manager is further used for: When the content of the function replica is updated, updating the content of the address mapping table according to the updated content of the function replica; When the parallel function replica runtime environment queries the address mapping relationship of the function replica, querying the content of the address mapping table and encapsulating it into a specific message format; When the address mapping table is managed in a distributed form, each distributed component synchronizes the content mapped by the address mapping table; and When the content of the address mapping table is lost, calling the Serverless parallel computing platform to restore the content of the address mapping table.
10. The Serverless parallel computing system for MPI according to claim 7, characterized in that, The Serverless parallel computing platform includes a client, a control node, computing nodes, and a function repository. The control node includes a gateway, a controller, a monitor, and a scaling manager. The computing nodes include executors and containers. Among them, The client is used for the relevant management of MPI parallel computing tasks and the initiation of computing operations; The control node is used to respond to the runtime environment requirements of MPI parallel computing tasks, and is responsible for receiving deployment and call requests in Serverless computing, request queue management, container scheduling, address mapping, resource monitoring, scaling management, starting parallel executors, and forwarding call requests; The computing nodes are responsible for starting, monitoring, executing, and destroying functions in Serverless computing; The function repository stores the basic container environment required for writing functions and the built function container images.
11. The Serverless parallel computing system for MPI according to claim 10, wherein This Serverless parallel computing platform is also used for the functional development service of MPI parallel programs: (1) Write MPI function processing code according to the Serverless template with the MPI runtime environment, and configure the function runtime environment; and (2) Submit the function for automated building, package the MPI environment into a container image, and push it to the image repository address configured in the function runtime environment in step (1) to facilitate any node in the cluster to download the image.
12. The Serverless parallel computing system for MPI according to claim 10, wherein This Serverless parallel computing platform is also used for parallel function deployment service: (1) The client submits a deployment request; (2) The deployment request enters the request queue of the control node gateway for queuing; (3) The control node gateway takes out the first deployment request from the request queue and hands it over to the controller for scheduling; (4) The controller parses the scheduling request and queries whether the same task has been deployed on the cluster. If not, it obtains the cluster resource information collected by the monitor, sets the thresholds for resource change degree and parallel request number, and jumps to step (6); if so, it notifies the scaling manager to perform scaling, and transfers to step (5); (5) If there is a situation of insufficient resources or a decrease in the parallelism of the function to be deployed, the scaling manager sends an instruction to scale down the parallel function to the controller. If there is a situation of an increase in the parallelism of the function to be deployed, the scaling manager sends an instruction to scale up the parallel function to the controller, and transfers to step (6); (6) The controller makes a match based on the cluster resource information and deployment request information collected by the monitor. If a function needs to be deployed, when each parallel function of the task is scheduled to a computing node with sufficient resources, maximize the affinity score of the parallel functions under the same task, allocate the function to the computing node and allocate the corresponding exclusive resources; if a function needs to be shut down, close the containers of the corresponding functions in the startup order; (7) The controller writes the scheduled function information into the address mapping table; (8) The controller notifies all the computing nodes to which the parallel function is scheduled to download the function image, and each computing node downloads it from the image repository; and (9) The controller starts an executor on all the computing nodes to which the parallel function is scheduled, and the executor creates a parallel function container to run on this node.
13. The Serverless parallel computing system for MPI according to claim 10, characterized in that, This Serverless parallel computing platform is also used for parallel function triggering service: (1) The client sends a call request to the gateway of the control node; (2) The gateway forwards the request with parameters to the controller; (3) The controller obtains the address from the address mapping table and forwards it to each executor of the task; (4) The executor parses the forwarded request parameters, concatenates them to form an MPI command, and passes it to all the parallel functions under this task to call their execution, and at the same time monitors the running parallel functions in real time; and (5) When the parallel function finishes running, the executor receives a signal that the function exits normally or abnormally, and recovers the execution status and results.
Citation Information
Patent Citations
Communication behavior information of device based on message passing interface extraction method and system thereof
CN101571814A
Communication optimizations for distributed machine learning
US20190205745A1