A method and device for providing supercomputing internet operator services
Through the supercomputing Internet operator service-oriented approach, the problem of insufficient flexibility in computing component abstraction and resource allocation in the supercomputing Internet has been solved, the elastic supply and efficient utilization of supercomputing resources have been achieved, and the flexible construction and operation of complex cross-domain collaborative computing tasks have been supported.
Patent Information
- Application Number
- CN202510522227.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-04-24
AI Technical Summary
Existing workflow execution engines lack the abstraction of computing components in the supercomputing Internet, and the user-defined task node execution process lacks flexibility and scalability, making it difficult to meet the construction requirements of complex cross-domain collaborative applications. Function-as-a-Service technology is limited by security management and resource allocation methods in supercomputing systems, making it difficult to directly apply to the supercomputing Internet.
A supercomputing Internet operator service-oriented method is introduced. Through the control-end core service unit and function execution management unit, computing unit functions are constructed, operator trigger rules are defined, and a dynamic resource application and release mechanism is adopted to realize the service-oriented construction and operation of operators, supporting the elastic supply and flexible deployment of supercomputing resources.
It achieves efficient utilization of supercomputing resources and flexible application development, supports the service-oriented construction of supercomputing Internet operators, optimizes operator loading, improves computing resource utilization and task execution efficiency, shields differences in heterogeneous computing resources, and realizes seamless integration and collaborative work.
Smart Images

Figure CN120469793B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of high-performance computing, and in particular to a method and device for servitizing supercomputing Internet operators. Background Art
[0002] Operators were originally a core concept in mathematics and physics. Their core concept can be summarized as rules or mappings for operating on certain objects. In computing, an operator is a rule or function used to perform a specific operation or transformation, processing input data to produce a desired output. Compared to functions, operators are more fundamental, universal, and standardized definitions of data manipulation. In traditional computing environments, operators are typically used for basic operations such as arithmetic, logical reasoning, and data processing, and are the building blocks for building complex computational processes and algorithms. Common high-performance computing operators include numerical linear algebra operators such as sparse matrix-vector multiplication (SpMV), fast Fourier transform (FFT), and large-scale dense matrix factorization (LU, Cholesky), as well as partial differential equation solvers such as the gradient operator and Laplace operator in finite difference methods. Artificial intelligence operators refer to functions that implement basic neural network operations in deep learning frameworks (such as TensorFlow and PyTorch). They are typically oriented towards tensor computations and emphasize automatic differentiation, hardware acceleration (GPU / TPU), and dynamic graph optimization.
[0003] With the advancement of computing technology, applications in scientific and engineering computing are becoming increasingly complex and expansive. Solving computational problems on a larger scale or with higher precision is driving the discovery of new substances or the study of new properties. This, in turn, places new demands on the scale and diversity of computing resources. The supercomputing internet has emerged as a response to this challenge. Through technological innovation, it aims to integrate multiple supercomputing centers and intelligent computing centers into a highly interconnected, resource-sharing, and collaborative computing environment. This will foster a healthy and sustainable ecosystem for high-performance computing software and hardware, fostering an open and shared computing power ecosystem.
[0004] The concept of operators has been further expanded in the context of the supercomputing internet. On the one hand, supercomputing internet operators should be common, basic computing interfaces for application domains, serving as the fundamental building blocks for building cross-domain collaborative computing applications. They should be able to adapt to the demands of high-performance computing and complex data processing, provide powerful and flexible computing capabilities, and enable seamless integration and collaboration between diverse computing resources and applications. On the other hand, supercomputing internet operators, through standardized interfaces and protocols, can effectively shield the differences between heterogeneous computing resources. By leveraging technologies such as adaptive adaptation, automatic operator measurement, and cross-domain collaborative scheduling, they can load optimized operators onto the corresponding computing resources during the computation process, improving resource utilization and the efficiency of computing task execution.
[0005] Workflow execution engines are commonly used tools for building applications across computing resources. They are software systems used to automate the management and execution of complex business processes or computing tasks. They define, orchestrate, and schedule multiple tasks or operations, ensuring that they execute in a predefined logical order while also handling dependencies, error recovery, and resource allocation. Existing workflow execution engines, such as the High-Performance Computing Application Unified Runtime Framework (BEE), and Mashup, a hybrid execution strategy for scientific computing workflows, are no longer limited to a single computing resource. Mashup can access a variety of computing resources, including container platforms, virtual machine platforms, and cloud computing platforms. They also provide an easy-to-use, unified user interface and execution environment, offering advantages in reducing execution time and costs. However, most workflow execution engines lack abstraction of computing components. Users only define the execution commands and inputs and outputs of each task node, and the execution process is controlled by the engine's built-in algorithms. This lacks scalability and flexibility in terms of computing resource acquisition and task execution, making it difficult to meet the requirements for building complex cross-domain collaborative applications on the supercomputing internet.
[0006] Function-as-a-Service (FaaS) is a new service model and technology following Infrastructure-as-a-Service, Platform-as-a-Service, and Software-as-a-Service. It offers advantages such as elastic resource scaling and automated management. Serverless computing is its corresponding resource management technology. In the Function-as-a-Service model, users only need to focus on the specific logic and business implementation of the function, without having to worry about the configuration and management of the underlying servers or containers. When a function is triggered, the platform automatically allocates the required computing resources and releases them after the task is completed. This "elasticity + automation" approach gives the function service platform a clear advantage in scenarios such as sudden loads, high-concurrency access, and short-lifecycle operations, and significantly lowers the threshold for application deployment and operation. Currently, Function-as-a-Service technology has successfully demonstrated its technical feasibility in scenarios such as big data processing, scientific workflows, and AI. However, due to limitations in the security management and resource allocation methods of supercomputing systems, it is difficult to directly apply to the construction and operation of applications on the supercomputing Internet. Summary of the Invention
[0007] In response to the defects of the existing technology, the present invention provides a method and device for supercomputing Internet operator service, which can effectively solve the above problems.
[0008] The technical solution adopted in the present invention is as follows:
[0009] The present invention provides a method for providing a supercomputing internet operator service, comprising the following steps:
[0010] Step S1, the supercomputing internet operator service device includes a control-end core service unit and a function execution management unit; wherein: the control-end core service unit includes an operator trigger, a function construction module, a function registration module, a function communication management module, a supercomputing resource application and release module, and a supercomputing resource dynamic perception module;
[0011] Step S2: deploying the control-end core service unit in the front-end server of the supercomputing cluster system; deploying the function execution management unit in the cluster shared file system of the supercomputing cluster system;
[0012] Step S3: construct one or more computing unit functions through the function construction module according to application requirements, and register them with the control end core service unit through the function registration module; store the function metadata of the registered computing unit functions in the front-end server of the supercomputing cluster system; and store the executable program of the registered computing unit functions in the cluster shared file system;
[0013] Step S4, defining operator trigger rules; the operator trigger rules include the computing unit functions that need to be called, the data dependencies between the computing unit functions, and the resource requirements of the computing unit functions;
[0014] Step S5: Send the operator trigger rule and the given input data to the operator trigger; the operator trigger executes the operator trigger rule to obtain output data; the specific execution method is:
[0015] Step S5.1: The operator trigger adds all the computing unit functions that need to be started to the function task queue in order based on the computing unit functions that need to be called and the data dependencies between the computing unit functions;
[0016] Step S5.2: The supercomputing resource application and release module maintains a computing node list; the computing node list stores the computing nodes applied for by the supercomputing resource application and release module, the performance of each computing node, and the status of each computing node; the status of the computing node includes an idle state and an occupied state;
[0017] The supercomputing resource application and release module evaluates whether the idle computing nodes in the current computing node list meet the execution requirements of the computing unit function to be executed this time in the function task queue; if not, the module dynamically applies for a computing node from the supercomputing cluster system at the granularity of a single computing node and adds it to the computing node list. At the same time, the module calls and starts the function execution management unit corresponding to the newly applied computing node from the cluster shared file system, records the correspondence between the newly applied computing node and the function execution management unit, and then executes step S5.3; if it meets the requirements, the module directly executes step S5.3;
[0018] Step S5.3, the operator trigger selects the computing node required by the computing unit function to be executed this time in the function task queue according to the strategy in the computing node list, and starts the corresponding function execution management unit through the computing node, and the function execution management unit executes the computing unit function to be executed this time in the function task queue on the computing node;
[0019] When the computing unit function executes the computing task, the computing unit function communicates with the external service interface through the function communication management module;
[0020] Step S5.4: After the computation unit function to be executed this time is completed, the operator trigger determines whether all computation unit functions in the function task queue have been completed; if so, step S5.5 is executed; otherwise, the next computation unit function to be executed is determined, and the process returns to step S5.3;
[0021] Step S5.5: The operator trigger obtains the result of executing the operator trigger rule this time and uses it as output data;
[0022] During the execution of steps S5.3 to S5.5, the supercomputing resource application and release module releases idle computing nodes based on the computing unit functions to be executed in the function task queue and the status of each computing node in the computing node list.
[0023] Preferably, the function construction module constructs multiple computing unit functions, and the specific construction method is:
[0024] Determine the input / output parameter definition and file information of the calculation unit function; use the calculation unit function code template to obtain the calculation unit function code;
[0025] Building a computing unit function container image based on the computing unit function code;
[0026] According to the computing unit function container image, function meta information is determined and registered; wherein the function meta information includes the loading and calling method of the function.
[0027] Preferably, the function execution management unit includes a function startup management module, a function input and output data management module, a function running environment module, a function heartbeat management module, and a resource allocation and function isolation module;
[0028] The function startup management module sends the amount of resources required to be occupied by one or more computing unit functions to be started to the resource allocation and function isolation module;
[0029] The resource allocation and function isolation module allocates resources to the computing unit functions that need to be started and implements resource isolation according to the internal resource occupancy of the computing node dynamically applied for;
[0030] The function startup management module determines the function runtime environment of the computing unit function to be started and sends it to the function runtime environment module, which loads the corresponding computing unit function from the cluster shared file system;
[0031] The function startup management module starts the computing unit function loaded by the function execution environment module in the computing node based on the resources allocated by the resource allocation and function isolation module;
[0032] During the execution of the computing unit function, the function heartbeat management module detects the running status of the computing unit function and reports it to the supercomputing resource dynamic perception module to detect whether the computing unit function is abnormally executed;
[0033] During the execution of the computing unit function, the function input and output data management module monitors the execution of the computing unit function and the function output data. When it is detected that the computing unit function has been completed, it triggers the execution of subsequent computing unit functions and sends the function output data to the subsequent computing unit functions.
[0034] Preferably, the supercomputing resource application and release module uses the following algorithm to dynamically apply for computing nodes from the supercomputing cluster system at the granularity of a single computing node:
[0035] The supercomputing resource application and release module evaluates whether the idle computing nodes in the current computing node list meet the execution requirements of the computing unit function to be executed this time; if not, it applies for one computing node from the supercomputing cluster system;
[0036] Then, the supercomputing resource application and release module calculates the number of computing unit functions that can be processed by the newly applied computing node according to the resource requirements of the computing unit functions to be executed;
[0037] Update the number of computing unit functions and the computing node list to be executed this time; evaluate whether the idle computing nodes in the current computing node list meet the execution requirements of the computing unit function to be executed this time; if not, apply for two computing nodes from the supercomputing cluster system;
[0038] Then, the supercomputing resource application and release module calculates the number of computing unit functions that can be processed by the newly applied computing node according to the resource requirements of the computing unit functions to be executed;
[0039] Update the number of computing unit functions and the computing node list to be executed this time; evaluate whether the idle computing nodes in the current computing node list meet the execution requirements of the computing unit function to be executed this time; if not, apply for 4 computing nodes from the supercomputing cluster system;
[0040] Similarly, a 2x progressive expansion strategy is used to apply for computing nodes.
[0041] Preferably, the supercomputing resource application and release module uses the following algorithm to release idle computing nodes according to the computing unit functions to be executed in the function task queue and the status of each computing node in the computing node list:
[0042] When the continuous idle time of the computing nodes in the computing node list exceeds a threshold, or exceeds the resource exponential moving average, a resource release operation is triggered.
[0043] Preferably, the resource exponential moving average is calculated as follows:
[0044] Step 1: Calculate the required number of computing nodes NodeNum based on the computing unit functions in the function task queue and the corresponding resource requirements;
[0045] Step ② Calculate the resource exponential moving average EMA at the current time t t :
[0046] When t=0, EMA t =NodeNum;
[0047] When t>0, EMA t =α*NodeNum+(1-α)*EMA t-1
[0048] Where: α is the smoothing factor constant; EMA t-1 is the moving average of resource exponential at time t-1;
[0049] Step 3: Determine whether the number of computing nodes in the computing node list is greater than EMA t Round up; if yes, release the computing node with the longest idle time in the computing node list;
[0050] Step ④ updates the computing node list; and returns to step ①.
[0051] The present invention also provides a supercomputing internet operator service-oriented device, which includes a control-end core service unit and a function execution management unit; wherein: the control-end core service unit includes an operator trigger, a function construction module, a function registration module, a function communication management module, a supercomputing resource application and release module, and a supercomputing resource dynamic perception module; the function execution management unit includes a function startup management module, a function input and output data management module, a function running environment module, a function heartbeat management module, and a resource allocation and function isolation module;
[0052] Deploy the control-end core service unit in the front-end server of the supercomputing cluster system; deploy the function execution management unit in the cluster shared file system of the supercomputing cluster system;
[0053] One or more computing unit functions are constructed through the function construction module according to application requirements, and registered with the control end core service unit through the function registration module; function meta-information of the registered computing unit functions is stored in the front-end server of the supercomputing cluster system; and the executable program of the registered computing unit functions is stored in the cluster shared file system;
[0054] Define operator trigger rules; the operator trigger rules include the computing unit functions that need to be called, the data dependencies between the computing unit functions, and the resource requirements of the computing unit functions;
[0055] The operator trigger rule and given input data are sent to the operator trigger; the operator trigger executes the operator trigger rule to obtain output data.
[0056] The method and device for providing supercomputing internet operator services provided by the present invention have the following advantages:
[0057] While complying with the current security management and resource allocation methods of supercomputing centers, the present invention introduces the technical idea of function-as-a-service into the supercomputing Internet, breaking the inherent resource allocation method of the original supercomputing cluster system that allocates a fixed number of resources according to the task calculation scale, separating the application and use of supercomputing resources, supporting the dynamic application and release of supercomputing resources, realizing the elastic supply of supercomputing resources, and supporting the service-oriented construction and operation of supercomputing Internet operators. The present invention decomposes complex cross-domain collaborative computing tasks into reusable operator units, and flexibly deploys and runs them in supercomputing resources in a service-oriented manner. At the same time, with the help of technologies such as adaptive adaptation, automatic operator measurement, and cross-domain collaborative scheduling, it optimizes operator loading, thereby achieving efficient resource utilization and flexible application development. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 Schematic diagram of the supercomputing Internet operator service device provided by the present invention;
[0059] Figure 2 A schematic diagram of the cross-domain collaborative application principle of the supercomputing Internet operator provided by the present invention;
[0060] Figure 3 Schematic diagram of function construction and function registration provided by the present invention;
[0061] Figure 4 This is an operation flow chart of the supercomputing Internet operator service device provided by the present invention. DETAILED DESCRIPTION
[0062] In order to make the technical problems, technical solutions and beneficial effects solved by the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0063] The technical problem to be solved by this invention is to introduce the technical concept of Function as a Service into the supercomputing Internet, while complying with the current security management and resource allocation methods of supercomputing centers, to achieve the elastic supply of supercomputing resources and support the service-oriented construction and operation of supercomputing Internet operators. This invention decomposes complex cross-domain collaborative computing tasks into reusable operator units, and flexibly deploys and runs them in supercomputing resources in a service-oriented manner. At the same time, it optimizes operator loading with the help of technologies such as adaptive adaptation, automatic operator measurement, and cross-domain collaborative scheduling, thereby achieving efficient resource utilization and flexible application development.
[0064] The present invention breaks the traditional supercomputing resource batch processing mode, separates the application and use of supercomputing resources, supports the dynamic application and release of supercomputing resources, runs the supercomputing Internet application operator unit in a service-oriented manner, starts and executes multiple computing unit subtasks decomposed by the supercomputing Internet operator in a function computing manner, and provides users with fast service response.
[0065] The present invention provides a supercomputing Internet operator service method, see Figures 1 to 4 , including the following steps:
[0066] Step S1, the supercomputing internet operator service device includes a control-end core service unit and a function execution management unit; wherein: the control-end core service unit includes an operator trigger, a function construction module, a function registration module, a function communication management module, a supercomputing resource application and release module, and a supercomputing resource dynamic perception module;
[0067] Step S2: deploying the control-end core service unit in the front-end server of the supercomputing cluster system; deploying the function execution management unit in the cluster shared file system of the supercomputing cluster system;
[0068] Step S3: construct one or more computing unit functions through the function construction module according to application requirements, and register them with the control end core service unit through the function registration module; store the function metadata of the registered computing unit functions in the front-end server of the supercomputing cluster system; and store the executable program of the registered computing unit functions in the cluster shared file system;
[0069] In this step, the function building module builds multiple computing unit functions, see Figure 3 , the specific construction method is:
[0070] Determine the input / output parameter definition and file information of the calculation unit function; use the calculation unit function code template to obtain the calculation unit function code;
[0071] Building a computing unit function container image based on the computing unit function code;
[0072] According to the computing unit function container image, function meta information is determined and registered; wherein the function meta information includes the loading and calling method of the function.
[0073] Step S4, defining operator trigger rules; the operator trigger rules include the computing unit functions that need to be called, the data dependencies between the computing unit functions, and the resource requirements of the computing unit functions;
[0074] Step S5, sending the operator trigger rule and the given input data to the operator trigger; the operator trigger executes the operator trigger rule to obtain output data; Figure 4 , the specific implementation method is:
[0075] Step S5.1: The operator trigger adds all the computing unit functions that need to be started to the function task queue in order based on the computing unit functions that need to be called and the data dependencies between the computing unit functions;
[0076] Step S5.2: The supercomputing resource application and release module maintains a computing node list; the computing node list stores the computing nodes applied for by the supercomputing resource application and release module, the performance of each computing node, and the status of each computing node; the status of the computing node includes an idle state and an occupied state;
[0077] The supercomputing resource application and release module evaluates whether the idle computing nodes in the current computing node list meet the execution requirements of the computing unit function to be executed this time in the function task queue; if not, the module dynamically applies for a computing node from the supercomputing cluster system at the granularity of a single computing node and adds it to the computing node list. At the same time, the module calls and starts the function execution management unit corresponding to the newly applied computing node from the cluster shared file system, records the correspondence between the newly applied computing node and the function execution management unit, and then executes step S5.3; if it meets the requirements, the module directly executes step S5.3;
[0078] Step S5.3, the operator trigger selects the computing node required by the computing unit function to be executed this time in the function task queue according to the strategy in the computing node list, and starts the corresponding function execution management unit through the computing node, and the function execution management unit executes the computing unit function to be executed this time in the function task queue on the computing node;
[0079] When the computing unit function executes the computing task, the computing unit function communicates with the external service interface through the function communication management module;
[0080] Step S5.4: After the computation unit function to be executed this time is completed, the operator trigger determines whether all computation unit functions in the function task queue have been completed; if so, step S5.5 is executed; otherwise, the next computation unit function to be executed is determined, and the process returns to step S5.3;
[0081] Step S5.5: The operator trigger obtains the result of executing the operator trigger rule this time and uses it as output data;
[0082] During the execution of steps S5.3 to S5.5, the supercomputing resource application and release module releases idle computing nodes based on the computing unit functions to be executed in the function task queue and the status of each computing node in the computing node list.
[0083] In the present invention, the function execution management unit includes a function startup management module, a function input and output data management module, a function running environment module, a function heartbeat management module, and a resource allocation and function isolation module;
[0084] The function startup management module sends the amount of resources required to be occupied by one or more computing unit functions to be started to the resource allocation and function isolation module;
[0085] The resource allocation and function isolation module allocates resources to the computing unit functions that need to be started and implements resource isolation according to the internal resource occupancy of the computing node dynamically applied for;
[0086] The function startup management module determines the function runtime environment of the computing unit function to be started and sends it to the function runtime environment module, which loads the corresponding computing unit function from the cluster shared file system;
[0087] The function startup management module starts the computing unit function loaded by the function execution environment module in the computing node based on the resources allocated by the resource allocation and function isolation module;
[0088] During the execution of the computing unit function, the function heartbeat management module detects the running status of the computing unit function and reports it to the supercomputing resource dynamic perception module to detect whether the computing unit function is abnormally executed;
[0089] During the execution of the computing unit function, the function input and output data management module monitors the execution of the computing unit function and the function output data. When it is detected that the computing unit function has been completed, it triggers the execution of subsequent computing unit functions and sends the function output data to the subsequent computing unit functions.
[0090] The supercomputing resource application and release module uses the following algorithm to dynamically apply for computing nodes from the supercomputing cluster system at the granularity of a single computing node:
[0091] The supercomputing resource application and release module evaluates whether the idle computing nodes in the current computing node list meet the execution requirements of the computing unit function to be executed this time; if not, it applies for one computing node from the supercomputing cluster system;
[0092] Then, the supercomputing resource application and release module calculates the number of computing unit functions that can be processed by the newly applied computing node according to the resource requirements of the computing unit functions to be executed;
[0093] Update the number of computing unit functions and the computing node list to be executed this time; evaluate whether the idle computing nodes in the current computing node list meet the execution requirements of the computing unit function to be executed this time; if not, apply for two computing nodes from the supercomputing cluster system;
[0094] Then, the supercomputing resource application and release module calculates the number of computing unit functions that can be processed by the newly applied computing node according to the resource requirements of the computing unit functions to be executed;
[0095] Update the number of computing unit functions and the computing node list to be executed this time; evaluate whether the idle computing nodes in the current computing node list meet the execution requirements of the computing unit function to be executed this time; if not, apply for 4 computing nodes from the supercomputing cluster system;
[0096] Similarly, a 2x progressive expansion strategy is used to apply for computing nodes.
[0097] The supercomputing resource application and release module uses the following algorithm to release idle computing nodes based on the computing unit functions to be executed in the function task queue and the status of each computing node in the computing node list:
[0098] When the continuous idle time of the computing nodes in the computing node list exceeds a threshold, or exceeds the resource exponential moving average, a resource release operation is triggered.
[0099] The moving average of the resource index is calculated as follows:
[0100] Step 1: Calculate the required number of computing nodes NodeNum based on the computing unit functions in the function task queue and the corresponding resource requirements;
[0101] Step ② Calculate the resource exponential moving average EMA at the current time t t :
[0102] When t=0, EMA t =NodeNum;
[0103] When t>0, EMA t =α*NodeNum+(1-α)*EMA t-1
[0104] Where: α is the smoothing factor constant; EMA t-1 is the moving average of resource exponential at time t-1;
[0105] Step 3: Determine whether the number of computing nodes in the computing node list is greater than EMA tRound up; if yes, release the computing node with the longest idle time in the computing node list;
[0106] Step ④ updates the computing node list; and returns to step ①.
[0107] The present invention also provides a supercomputing Internet operator service-oriented device, which includes a control-end core service unit and a function execution management unit; wherein: the control-end core service unit includes an operator trigger, a function construction module, a function registration module, a function communication management module, a supercomputing resource application and release module, and a supercomputing resource dynamic perception module; the function execution management unit includes a function startup management module, a function input and output data management module, a function running environment module, a function heartbeat management module, and a resource allocation and function isolation module.
[0108] Deploy the control-end core service unit in the front-end server of the supercomputing cluster system; deploy the function execution management unit in the cluster shared file system of the supercomputing cluster system;
[0109] One or more computing unit functions are constructed through the function construction module according to application requirements, and registered with the control end core service unit through the function registration module; function meta-information of the registered computing unit functions is stored in the front-end server of the supercomputing cluster system; and the executable program of the registered computing unit functions is stored in the cluster shared file system;
[0110] Define operator trigger rules; the operator trigger rules include the computing unit functions that need to be called, the data dependencies between the computing unit functions, and the resource requirements of the computing unit functions;
[0111] The operator trigger rule and given input data are sent to the operator trigger; the operator trigger executes the operator trigger rule to obtain output data.
[0112] A specific embodiment is described below:
[0113] The present invention discloses a supercomputing internet operator service-oriented device, comprising two modules: a control-side core service unit and a function execution management unit. The control-side core service unit is deployed and runs independently of the supercomputing cluster system's front-end server; the function execution management unit runs on the requested computing nodes, independently deployed and running on each computing node, with its lifecycle managed by the control-side core service unit.
[0114] (1) Control end core service unit:
[0115] The overall control module of the system is responsible for defining, registering, and starting functions and operators, as well as elastic management of supercomputing resources. Specifically, it includes:
[0116] Operator trigger rule definition: allows users to define operator trigger rules, including computing unit function call parameters, data dependencies, startup quantity, resource requirements, etc., register and store them in the device, and support runtime retrieval.
[0117] Function construction module: Provides users with a templated container image construction tool to assist users in encapsulating computing interfaces such as algorithm libraries or computational solver libraries into computing unit functions of this method.
[0118] Function registration module: registers the metadata of the calculation unit function in the system, including the loading and calling methods of the calculation unit function, to facilitate the startup of the calculation unit function when the calculation unit function is accessed later.
[0119] Function Communication Management Module: In a supercomputing cluster system, computing unit functions run on dynamically requested compute nodes. The Function Communication Management Module is responsible for managing message forwarding through the supercomputing cluster system's login nodes, enabling message communication between external service interfaces and computing unit functions running on compute nodes.
[0120] Supercomputing resource application and release module: Access and apply for computing nodes in the cluster through the resource management interface of the supercomputing cluster system. When the cluster load is high, the supercomputing resource application and release module needs to queue up and wait for resources. The computing node can only be obtained when the conditions are met. When all computing unit functions on a computing node applied for are executed and it is predicted that there will be no new computing unit function execution requests, the computing node is released to avoid excessive computing node costs.
[0121] Supercomputing resource dynamic perception module: responsible for monitoring and managing dynamically applied computing nodes, and can record and feedback the computing unit function load of the currently available computing nodes in real time.
[0122] (2) Function Execution Management Unit: Responsible for starting and managing the execution of one or more functions in the requested computing nodes, including:
[0123] Function startup management module: runs on one computing node in the device's available resources and only runs one function startup management module. It is the main program for computing node startup and is responsible for resource allocation, function startup and management within the current computing node.
[0124] Function input and output data management module: monitors the execution of calculation unit functions and data output. When it detects that the execution of a calculation unit function is completed, it triggers the execution of subsequent calculation unit functions and provides the calculation unit function output data to subsequent calculation unit functions.
[0125] Function runtime environment module: The software stack and process context required for computing unit function execution are created by loading a pre-built function image and are the specific function execution body.
[0126] Function heartbeat management module: Detects the running status of computing unit functions and reports it to the supercomputing resource dynamic perception module, allowing operator triggers to perceive task execution anomalies and take measures.
[0127] Resource allocation and function isolation module: Collaborative function startup management module, responsible for allocating reasonable resources and implementing isolation for the started function running environment based on the internal resource usage of the computing node.
[0128] A method for servitizing supercomputing Internet operators, the specific implementation of which includes three parts:
[0129] Part 1: Define, build, and register functions
[0130] Clarify the input / output parameter definitions and file information of each computing unit function, encapsulate the function code and its software technology stack based on lightweight virtualization container technology, and register and store function metadata in the device.
[0131] Generally speaking, computing unit functions are mostly in the form of executable binary code, which can be started and executed through the command line in a compatible operating environment, passing data through parameters and data files, and producing standard output, error output, and file output. At the same time, there are also computing unit functions provided in the form of shared libraries. Under the traditional usage method, an additional main program needs to be written and compiled into the same executable program before it can be run. In this method, its functions can be run through an additional service interface. For computing unit functions in the form of binary executable code, users can use container tools such as Docker to build the container image by themselves, and then directly execute step 3 to register the computing unit function in this device. For computing unit functions in the form of shared libraries, users can use steps 1 and 2 to assist in building the function container image, and then register the computing unit function in this device through step 3.
[0132] Step 1: Template Generation
[0133] The user executes the function definition command and passes in the required parameters through the command line, including the function name (multiple functions are supported) and the function writing language. The command verifies the user input and returns the function writing template file and the basic Dockerfile required to build the function image.
[0134] Step 2: Build the function image
[0135] The user writes function calls and other required code logic into the template file, modifies the basic Dockerfile according to the deployment dependencies, and then starts the code compilation and container image building through the command line and pushes them to the specified container repository.
[0136] Step 3: Function registration
[0137] The user provides a function mapping relationship definition (function name, function container image, function call address / command), registers the function to the device through a command, and stores and maintains information such as function ID, function name, function container image address, function type, creation time, parameter type list, etc. in the device.
[0138] The second part defines the operator trigger rules
[0139] The user defines the operator trigger rules, which include the calling parameters, data dependencies, startup quantity, resource requirements, etc. between one or more functions; the basic information of the rules is defined and recorded in the device through commands, which serves as the subsequent operator service execution script.
[0140] The third part starts the application and triggers the execution of operators and functions.
[0141] The service object of this method is cross-domain collaborative computing applications, supporting their optimized operation in the supercomputing Internet infrastructure. Therefore, the first step of execution is to decompose the application process. In this method, the cross-domain collaborative computing application can be decomposed into multiple supercomputing Internet operators. Each supercomputing Internet operator is the sum of execution tasks in the same computing cluster system, which can be further divided into one or more computing units, and is called, triggered and run as a function in the cluster system. The specific application decomposition and operator decomposition methods need to be designed one by one according to the specific application process and computing tools. The method of the present invention focuses on the specific steps of the service-oriented operation of the supercomputing Internet operator.
[0142] Specifically, when a cross-domain collaborative computing application is running, an external system calls or initiates the execution of a supercomputing internet operator trigger, providing input data. The operator trigger within the device allocates compute nodes to the compute unit functions defined in the operator trigger rules based on the user-defined operator trigger rules and the available compute node resources. It then initiates and manages function execution based on the dependencies defined in the operator trigger rules, establishing an input and output data pipeline, and returns the output data after all compute unit functions have completed execution.
[0143] Step 1: Trigger the operator trigger to execute
[0144] The operator trigger first needs to check whether there is a reusable function startup management module. If not, it requests a new computing node based on the supercomputing resource dynamic perception, supercomputing resource application and release module, and starts the corresponding function startup management module, and establishes communication with it through function communication management; then the operator trigger notifies the function startup management module to start the function according to the predefined operator trigger rules.
[0145] Step 2: Create a function execution environment through the function startup management module
[0146] The operator trigger sends an execution request and input data to the function startup management module engine. At this time, the function startup management module starts the function running environment according to the resource allocation and function isolation strategy within the computing node, pulls the function of the defined operator trigger rule to build an image startup container, allocates a specific computing node to it, and executes the function process.
[0147] Step 3: Execute the function in the function execution environment and manage input and output data
[0148] The function execution environment provides the basic software stack and execution context for function execution, exposes a function call interface, and executes functions upon receiving a request. It also creates an output data cache queue and a background data management thread to transmit the output data of the current function to the compute node where the subsequent function resides. If the subsequent function has not yet been allocated computing resources, the transmission begins after the compute node resources are ready. The function execution environment module collaborates with the function heartbeat management module, and is recognized by the supercomputing resource dynamic perception module, which updates data in real time.
[0149] Step 4: Trigger subsequent function execution
[0150] When all the output data and data information of the current function have completely reached the computing node where the subsequent function is located, the function running environment module triggers the function startup management module where the subsequent function is located to execute the function request; if the subsequent function is executed in a different cluster system, the operator trigger of the corresponding cluster system is triggered to execute.
[0151] Compared with the prior art, the present invention has the following beneficial effects:
[0152] 1) Provide function definition and registration mechanisms. Through standardized interfaces and protocols, support the definition of common basic computing interfaces for application fields, form reusable basic units for building cross-domain collaborative computing applications, and achieve seamless integration and collaboration between different computing node resources and applications.
[0153] 2) Through the definition of supercomputing Internet service operators, the differences in operating environment and execution mode among different heterogeneous computing resources in the supercomputing Internet can be effectively shielded, and the advantages of different computing node resources can be fully utilized to coordinate and efficiently run applications.
[0154] 3) Provide a flexible and scalable supercomputing resource supply model, separate the supercomputing resource application and usage processes, dynamically apply for or release supercomputing computing node resources based on task load, and reasonably and fully utilize the acquired supercomputing computing node resources through internal computing node resource scheduling.
[0155] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for providing supercomputing internet operator services, characterized in that: The following steps are involved: Step S1, the supercomputing internet operator service device includes a control-end core service unit and a function execution management unit; wherein: the control-end core service unit includes an operator trigger, a function construction module, a function registration module, a function communication management module, a supercomputing resource application and release module, and a supercomputing resource dynamic perception module; Step S2: deploying the control-end core service unit in the front-end server of the supercomputing cluster system; deploying the function execution management unit in the cluster shared file system of the supercomputing cluster system; Step S3: construct one or more computing unit functions through the function construction module according to application requirements, and register them with the control end core service unit through the function registration module; store the function metadata of the registered computing unit functions in the front-end server of the supercomputing cluster system; and store the executable program of the registered computing unit functions in the cluster shared file system; Step S4, defining operator trigger rules; the operator trigger rules include the computing unit functions that need to be called, the data dependencies between the computing unit functions, and the resource requirements of the computing unit functions; Step S5: Send the operator trigger rule and the given input data to the operator trigger; the operator trigger executes the operator trigger rule to obtain output data; the specific execution method is: Step S5.1: The operator trigger adds all the computing unit functions that need to be started to the function task queue in order based on the computing unit functions that need to be called and the data dependencies between the computing unit functions; Step S5.2: The supercomputing resource application and release module maintains a computing node list; the computing node list stores the computing nodes applied for by the supercomputing resource application and release module, the performance of each computing node, and the status of each computing node; the status of the computing node includes an idle state and an occupied state; The supercomputing resource application and release module evaluates whether the idle computing nodes in the current computing node list meet the execution requirements of the computing unit function to be executed this time in the function task queue; if not, the module dynamically applies for a computing node from the supercomputing cluster system at the granularity of a single computing node and adds it to the computing node list. At the same time, the module calls and starts the function execution management unit corresponding to the newly applied computing node from the cluster shared file system, records the correspondence between the newly applied computing node and the function execution management unit, and then executes step S5.3; if it meets the requirements, the module directly executes step S5.3; Step S5.3, the operator trigger selects the computing node required by the computing unit function to be executed this time in the function task queue according to the strategy in the computing node list, and starts the corresponding function execution management unit through the computing node, and the function execution management unit executes the computing unit function to be executed this time in the function task queue on the computing node; When the computing unit function executes the computing task, the computing unit function communicates with the external service interface through the function communication management module; Step S5.4: After the computation unit function to be executed this time is completed, the operator trigger determines whether all computation unit functions in the function task queue have been completed; if so, step S5.5 is executed; otherwise, the next computation unit function to be executed is determined, and the process returns to step S5.3; Step S5.5: The operator trigger obtains the result of executing the operator trigger rule this time and uses it as output data; During the execution of steps S5.3 to S5.5, the supercomputing resource application and release module releases idle computing nodes based on the computing unit functions to be executed in the function task queue and the status of each computing node in the computing node list.
2. A supercomputing Internet operator service method according to claim 1, characterized in that: The function construction module constructs multiple computing unit functions, and the specific construction method is as follows: Determine the input / output parameter definition and file information of the calculation unit function; use the calculation unit function code template to obtain the calculation unit function code; Building a computing unit function container image based on the computing unit function code; According to the computing unit function container image, function meta information is determined and registered; wherein the function meta information includes the loading and calling method of the function.
3. The method for providing supercomputing Internet operators as a service according to claim 1, characterized in that: The function execution management unit includes a function startup management module, a function input and output data management module, a function running environment module, a function heartbeat management module and a resource allocation and function isolation module; The function startup management module sends the amount of resources required to be occupied by one or more computing unit functions to be started to the resource allocation and function isolation module; The resource allocation and function isolation module allocates resources to the computing unit functions that need to be started and implements resource isolation according to the internal resource occupancy of the computing node dynamically applied for; The function startup management module determines the function runtime environment of the computing unit function to be started and sends it to the function runtime environment module, which loads the corresponding computing unit function from the cluster shared file system; The function startup management module starts the computing unit function loaded by the function execution environment module in the computing node based on the resources allocated by the resource allocation and function isolation module; During the execution of the computing unit function, the function heartbeat management module detects the running status of the computing unit function and reports it to the supercomputing resource dynamic perception module to detect whether the computing unit function is abnormally executed; During the execution of the computing unit function, the function input and output data management module monitors the execution of the computing unit function and the function output data. When it is detected that the computing unit function has been completed, it triggers the execution of subsequent computing unit functions and sends the function output data to the subsequent computing unit functions.
4. The method for providing supercomputing Internet operators as a service according to claim 1, characterized in that: The supercomputing resource application and release module uses the following algorithm to dynamically apply for computing nodes from the supercomputing cluster system at the granularity of a single computing node: The supercomputing resource application and release module evaluates whether the idle computing nodes in the current computing node list meet the execution requirements of the computing unit function to be executed this time; if not, it applies for one computing node from the supercomputing cluster system; Then, the supercomputing resource application and release module calculates the number of computing unit functions that can be processed by the newly applied computing node according to the resource requirements of the computing unit functions to be executed; Update the number of computing unit functions and the computing node list to be executed this time; evaluate whether the idle computing nodes in the current computing node list meet the execution requirements of the computing unit function to be executed this time; if not, apply for two computing nodes from the supercomputing cluster system; Then, the supercomputing resource application and release module calculates the number of computing unit functions that can be processed by the newly applied computing node according to the resource requirements of the computing unit functions to be executed; Update the number of computing unit functions and the computing node list to be executed this time; evaluate whether the idle computing nodes in the current computing node list meet the execution requirements of the computing unit function to be executed this time; if not, apply for 4 computing nodes from the supercomputing cluster system; Similarly, a 2x progressive expansion strategy is used to apply for computing nodes.
5. The method for providing supercomputing Internet operators as a service according to claim 1, characterized in that: The supercomputing resource application and release module uses the following algorithm to release idle computing nodes based on the computing unit functions to be executed in the function task queue and the status of each computing node in the computing node list: When the continuous idle time of the computing nodes in the computing node list exceeds a threshold, or exceeds the resource exponential moving average, a resource release operation is triggered.
6. A supercomputing Internet operator service method according to claim 5, characterized in that: The resource exponential moving average is calculated as: Step 1: Calculate the required number of computing nodes NodeNum based on the computing unit functions in the function task queue and the corresponding resource requirements; Step ② Calculate the resource exponential moving average EMA at the current time t t : When t=0, EMA t =NodeNum; When t>0,EMA t =α*NodeNum+(1-α)*EMA t-1 Where: α is the smoothing factor constant; EMA t-1 is the moving average of resource exponential at time t-1; Step 3: Determine whether the number of computing nodes in the computing node list is greater than EMA t Round up; if yes, release the computing node with the longest idle time in the computing node list; Step ④ updates the computing node list; and returns to step ①.
7. A supercomputing Internet operator service device, characterized in that: The supercomputing internet operator service-oriented device includes a control-end core service unit and a function execution management unit; wherein: the control-end core service unit includes an operator trigger, a function construction module, a function registration module, a function communication management module, a supercomputing resource application and release module, and a supercomputing resource dynamic perception module; the function execution management unit includes a function startup management module, a function input and output data management module, a function running environment module, a function heartbeat management module, and a resource allocation and function isolation module; Deploy the control-end core service unit in the front-end server of the supercomputing cluster system; deploy the function execution management unit in the cluster shared file system of the supercomputing cluster system; One or more computing unit functions are constructed through the function construction module according to application requirements, and registered with the control end core service unit through the function registration module; function meta-information of the registered computing unit functions is stored in the front-end server of the supercomputing cluster system; and the executable program of the registered computing unit functions is stored in the cluster shared file system; Define operator trigger rules; the operator trigger rules include the computing unit functions that need to be called, the data dependencies between the computing unit functions, and the resource requirements of the computing unit functions; The operator trigger rule and given input data are sent to the operator trigger; the operator trigger executes the operator trigger rule to obtain output data.
Citation Information
Patent Citations
Multidisciplinary data model calculation integrated collaborative scientific research cloud service system
CN115905723A
Multi-policy intelligent scheduling method and apparatus oriented to heterogeneous computing power
US20240111586A1