Resource Parallel Scheduling and Optimization Method and System for Large-Scale Difference Operators
By using multi-grained parallel computing and automatic parallelization middleware in a cluster environment, the resource scheduling of differential operators is optimized, and the efficient computing problem of large-scale differential operators in a cluster environment is solved, and the collaborative computing of CPU and GPU is realized, which improves computing efficiency and resource utilization.
Patent Information
- Application Number
- CN202211126572.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-16
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-09-16
AI Technical Summary
In a cluster environment, how to efficiently schedule and optimize the resource allocation of large-scale differential operators to improve the calculation rate and resource utilization rate.
The middleware that uses multi-grained parallel computing and automatic parallelization is adopted. By extracting the common parts of the difference operator and converting them into multi-threaded C code and CUDA code, combining distributed frameworks and message middleware, tasks are reasonably allocated to the CPU and GPU for calculations, and task allocation is optimized using sharding ideas and resource scheduling models.
It realizes efficient calculation of differential operators in a cluster environment, reduces calculation time, improves resource utilization, and supports parallel conversion of serial programs.
Smart Images

Figure CN115509743B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of parallel computing, middleware and resource scheduling in the computer discipline, and relates to a resource parallel scheduling and optimization method and system for large-scale difference operators, in particular to a resource parallel scheduling and optimization method and system for large-scale difference operators combining multi-granularity parallel computing, automatic parallelization and resource scheduling middleware in a cluster environment. Background Art
[0002] A difference operator is an operator. For any real function f(x), if we denote Δf(x) = f(x + 1) - f(x), then Δ is called the forward difference operator, simply referred to as the difference operator. Difference is one of the basic concepts in computational mathematics, referring to the change of a discrete function at discrete nodes. The difference operator is very useful in numerical integration, numerical differentiation and numerical solutions of differential equations.
[0003] Resource scheduling is a method of allocating work to the resources that perform the work. The work may be virtual computing elements, such as threads, processes or data streams, which are in turn scheduled onto hardware resources (such as processors, network links or expansion cards). How to reasonably allocate computing resources to improve the computing speed is the main content of this research.
[0004] Heterogeneous computing mainly refers to the computing method of a system composed of computing units with different types of instruction sets and architectures. Using computing units with different types of instruction sets and different architectures to form a hybrid system and performing calculations in a special way is called "heterogeneous computing".
[0005] Middleware is a type of software between application systems and system software. It uses the basic functions provided by the system software to connect various parts of the application system on the network or different applications, and can achieve the purpose of resource sharing and function sharing. Summary of the Invention
[0006] The purpose of the present invention is to provide a resource parallel scheduling and optimization method and system for large-scale difference operators combining multi-granularity parallel computing, automatic parallelization and resource scheduling middleware in a cluster environment to achieve optimal resource scheduling.
[0007] The technical solution adopted by the method of the present invention is: a resource parallel scheduling and optimization method for large-scale difference operators, including the following steps:
[0008] Step 1: Extract the difference operator, analyze the common parts therein, take their common expressions, extract the differences as parameters, and convert the obtained difference operator into multi-threaded C code and code under CUDA;
[0009] Step 2: Obtain the demand task, i.e., the number of times N that the difference operator needs to be repeatedly calculated; send the demand task to the cluster;
[0010] Step 3: Divide the total demand task, reasonably allocate the demand task to the CPUs and GPUs of the task execution units in the cluster for calculation, and store the calculation results.
[0011] The technical solution adopted by the system of the present invention is: a resource parallel scheduling and optimization system for large-scale difference operators, including the following modules:
[0012] Module 1 is used to extract the difference operator, analyze the common parts therein, take their common expressions, extract the differences as parameters, and convert the obtained difference operator into multi-threaded C code and code under CUDA;
[0013] Module 2 is used to obtain the demand task, i.e., the number of times N that the difference operator needs to be repeatedly calculated; send the demand task to the cluster;
[0014] Module 3 is used to divide the total demand task, reasonably allocate the demand task to the CPUs and GPUs of the task execution units in the cluster for calculation, and store the calculation results.
[0015] In Step 3, the total demand task is divided through a distributed framework, taking the total number of tasks N as the input and contacting through a message middleware; the message middleware distributes the total number of tasks N to each task execution unit respectively and monitors the running status of each task execution unit; when the operation of the task execution unit ends, the result is stored.
[0016] The resource scheduling and task allocation of the present invention use the idea of sharding.
[0017] The core idea of extracting the operator of the present invention is to integrate the operators with common properties or structures, and pass the different parts as changeable parameters, etc.
[0018] In the resource scheduling part of the present invention, the core technology is the middleware technology. The resource scheduling middleware and the message middleware used in the distributed framework.
[0019] The present invention tries to make the calculation times of the CPU and GPU the same so as to minimize the total calculation time.
[0020] Guided by the ideas of distributed computing, computer systems and structures, and sharding technology, the present invention uses a distributed framework to connect the task execution units to the host to form a cluster, and uses middleware in the cluster environment, having the functions of automatic parallelization translation and resource scheduling. For large-scale operators, analyze the operator structure, find the commonalities, extract the differences as parameters to pass, and translate its parallelized code.
[0021] The present invention can not only achieve efficient computing of operators on a cluster, but also help with the issue of converting other serial programs into parallel programs. After flexible changes, in addition to enabling the tasks to be processed to run on the cluster, if the task volume is small, the collaborative computing of CPU\GPU can also be achieved on a single host. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a flowchart of an embodiment of the present invention.
[0023] Figure 2 It is a schematic diagram of the working principle of the message middleware in an embodiment of the present invention.
[0024] Figure 3 It is a schematic diagram of the working principle of the distributed framework to achieve demand task partitioning in an embodiment of the present invention;
[0025] Figure 4 It is a schematic diagram of the working principle of the sharding idea in an embodiment of the present invention;
[0026] Figure 5 It is a resource scheduling model diagram of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] To facilitate the understanding and implementation of the present invention by those of ordinary skill in the art, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0028] The present invention is applicable to the case of calculating difference operators under a cluster. First, the multi-threaded C code of the operator and the code under CUDA are obtained. Then, the operators in the code are analyzed, commonalities of a series of operators are extracted, and the code is simplified. The extraction of operator commonalities is the main content of the present invention. After integrating the operator code, the target code is obtained, and then according to specific requirements, the target code and the number of calculations are resource-scheduled in the middleware, and the tasks are reasonably allocated to obtain the calculation result faster through the collaborative computing of CPU\GPU.
[0029] Please refer to Figure 1, A resource parallel scheduling and optimization method for large-scale difference operators provided by the present invention, the first part is the preparatory work. Extract the operators and convert the obtained operators into multi-threaded C code and code under CUDA. The second part is to summarize a series of operators, analyze the common parts among them, extract their common expressions, extract the differences as parameters, and then encapsulate them into functions. After the operator part is processed, obtain the requirement, that is, the number of times N that the operator needs to be repeatedly calculated. Send the task to the cluster system. Then divide the total task through the distributed framework and reasonably allocate it to each task execution unit. For example, the first task execution unit executes a quantity of n1, the second is n2, and so on. Where n1 + n2 +... + n i = N. Then the task execution unit obtains the number of tasks and starts the middleware to perform resource scheduling on it, that is, task allocation, and reasonably allocates the tasks to the CPU and GPU of the task unit. Perform calculations and store the calculation results.
[0030] In this embodiment, the middleware for resource scheduling is deployed on a single host of the cluster. The overall structure and working process of this middleware are as Figure 2 shown. Take the parallelized operator and the number of times to be calculated, that is, the number of tasks of a task execution unit as the input. Through the resource scheduling model, the tasks can be reasonably allocated to the CPU and GPU, achieving simultaneous calculation by both to reduce the running time.
[0031] The resource scheduling model of this embodiment is as shown in the appendix Figure 5 , First, extract the difference operators that need to be calculated and then enter the difference operator commonality analysis program to analyze and extract the commonalities of the difference operators. After the extraction is completed, perform a structured expression on the difference operators. Then generate the corresponding multi-threaded implementation code and CUDA implementation code through the code generation program. Shard the incoming data and assign metadata tags to mark the position of the data block in the source data for subsequent merging. For the incoming of new data, we first judge in the database whether the relevant parameters (the number of CPU blocks, the number of GPU blocks, and the number of threads created on the CPU mentioned above) configured for the Worker to be called exist. If the relevant parameters do not exist, send the task and let the Worker call the instrumentation program for fitting and obtain the relevant results (obtain these parameters (the same relevant parameters) by analyzing the information obtained by the instrumentation program after fitting). Then send these parameters together with the data to the task queue and wait for the slave machine to pull the task and run. The innovation of this model lies in using a mathematical model to pre-allocate the number of CPUs, GPUs, and threads in advance to achieve the purpose of using the CPU and GPU efficiently while the program is efficient.
[0032] To connect the task execution unit with the host to form a cluster, a distributed framework is required. The framework form provided by this embodiment is asFigure 3 As shown in the figure. The total number of tasks N is used as the input and is connected through a message middleware. This middleware distributes the total number of tasks N to each task execution unit respectively and monitors the running status of each task execution unit. After the operation of the task execution unit is completed, the result is stored in the backend.
[0033] Please refer to Figure 4 , in the resource scheduling and task allocation of the present invention, the idea of sharding is used.
[0034] The following further elaborates on the present invention by taking the operations of 12 basic difference operators derived from a partial differential equation in the used ocean model as an example.
[0035] This series of operators comes from the paper published by Xiaomeng Huang et al. in 2019 (Xiaomeng Huang, Xing Huang, Dong Wang, Qi Wu, Yi Li, Shixun Zhang, Yuwen Chen, Mingqing Wang, Yuan Gao, Qiang Tang, Yue Chen, Zheng Fang, Zhenya Song, and Guangwen Yang: OpenArray v1.0: a simple operator library for the decoupling of ocean modeling and parallel computing. Geosci. Model Dev., 12, 4729–4749, 2019). Table 1 shows the detailed content of the 12 basic difference operators.
[0036] Table 1
[0037] Serial number Operator name Discrete form 1 AXF [var(i,j,k)+var(i+1,j,k)] / 2 2 AXB [var(i,j,k)+var(i-1,j,k)] / 2 3 AYF [var(i,j,k)+var(i,j+1,k)] / 2 4 AYB [var(i,j,k)+var(i,j-1,k)] / 2 5 AZF [var(i,j,k)+var(i,j,k+1)] / 2 6 AZB [var(i,j,k)+var(i,j,k-1)] / 2 7 DXF [var(i+1,j,k)-var(i,j,k)] / dx(i,j) 8 DXB [var(i,j,k)-var(i-1,j,k)] / dx(i-1,j) 9 DYF [var(i,j+1,k)-var(i,j,k)] / dy(i,j) 10 DYB [var(i,j,k)-var(i,j-1,k)] / dy(i,j-1) 11 DZF [var(i,j,k+1)-var(i,j,k)] / dz(k) 12 DZB [var(i,j,k)-var(i,j,k-1)] / dz(k-1)
[0038] Using the dynamic code generation technology, the code corresponding to the operator can be obtained. For example,
[0039] (list_0[calc_id2(i,j,k,S0_0,S0_1)] + list_0[calc_id2(-1 + i,j,k,S0_0,S0_1)]) / 2 represents operator 2. For the convenience of subsequent description, in this embodiment, it is represented in another form, that is
[0040] 0.5 * (list_0[calc_id2(i,j,k,S0_0,S0_1)] + liSt_0[calc_id2(-1 + i,j,k,S0_0,S0_1)]);
[0041] (list_1[calc_id2(i,j,k,S1_0,S1_1)] + list_1[calc_id2(-1 + i,j,k,S1_0,S1_1)]) represents operator 8.
[0042] Through observation, the operator can be simplified into the following form: f * (A ± B). f is the coefficient, A is the left - hand array, and B is the right - hand array. ± is the operator between A and B, and + or - is selected according to the specific situation of different operators. In the formula, f is the coefficient, taking the value of 0.5 for the first six operators in Table 1; and 1.0 for the last six. Both A and B are in the form of three - dimensional arrays [i, j, k]. Due to different operators, the operations performed are also changed accordingly. For example, in AXB, no operation is performed on the left - hand side, while i - 1 is performed on the right - hand side. It is stipulated that when the array is [0, 0, 0], only the original data is passed in without performing operations. If operations are required, the corresponding positions of i, j, and k are changed. Here, 1 represents a +1 operation, and -1 represents a -1 operation. For example, when the right - hand side is [0, 1, 0], it means that the triple on the right becomes [i, j + 1, k]. The operation principle of the right - hand side and the left - hand side is the same. For the operator between the two terms, through analysis, it can be known that the first six items in Table 1 correspond to the operator +, and the last six items correspond to -. Therefore, the operator can be defined as a parameter op, where 1 corresponds to the + operation and -1 corresponds to the - operation. Pass f, A, B, and op as parameters in the program, and change the target operator to be calculated by changing the parameters.
[0043] Rewrite the integrated operator code into multi - threaded C code and CUDA code that can run on the cluster to obtain the required target code. After obtaining the code, the required number of operations N needs to be obtained. Pass N into the cluster, and reasonably allocate the tasks to different task execution units through the message middleware.
[0044] Start the resource scheduling middleware in the task execution unit. The middleware allocates the tasks to be processed by the task execution unit to the CPU and GPU in this unit. The number of threads on the CPU follows the task allocation scheme calculated by the resource scheduling model, which is restricted by conditions such as the number of CPU cores, the number of data blocks, and the block calculation acceleration rate based on GPU execution.
[0045] This embodiment provides a simplified mapping example, which specifically includes the following steps:
[0046] (1) Use the #program parallel tag to analyze the code blocks that can be parallelized.
[0047] (2) Create threads on the CPU according to the number of threads determined in the analysis phase, creating one more thread than the determined number of threads to control the operation of the GPU for the tasks.
[0048] (4) The CPU threads execute the allocated tasks in parallel.
[0049] (4.1) Under the control of the CPU threads, the CUDA kernels perform parallel computing on the tasks.
[0050] (4.2) After clarifying the tasks to be performed on the GPU, allocate GPU kernels and assign threads to the corresponding loops.
[0051] (4.3) Determine the dimensions and sizes of the grid and block, which affect the execution efficiency of CUDA.
[0052] (4.4) Perform coarse-grained inter-cluster scheduling through the resource scheduling model.
[0053] (5) Complete the computing tasks and merge the task data.
[0054] In the resource scheduling module, tasks are allocated according to the principle that the total time calculated when the CPU and GPU complete the tasks simultaneously is the shortest.
[0055] After the tasks are allocated, direct computing is performed. After the computing is completed, each task execution unit will store the results at the backend of the distributed framework, and the results and running time will be displayed on the interface.
[0056] It should be understood that the above description of the preferred embodiment is relatively detailed, and it should not be considered as a limitation to the protection scope of the present invention patent. Under the inspiration of the present invention, those of ordinary skill in the art can also make substitutions or deformations without departing from the protection scope defined by the claims of the present invention, and all fall within the protection scope of the present invention. The scope of protection requested by the present invention shall be subject to the appended claims.
Claims
1. A resource parallel scheduling and optimization method for large-scale difference operators, characterized in that, It includes the following steps: Step 1: Extract the difference operator, analyze the common parts therein, take their common expressions, extract the differences as parameters, and convert the obtained difference operator into multi-threaded C code and code under CUDA; Step 2: Obtain the required tasks, that is, the number of times N that the difference operator needs to be repeatedly calculated; send the required tasks to the cluster; Step 3: Divide the total required tasks, reasonably allocate the required tasks to the CPUs and GPUs of the task execution units in the cluster for calculation, and store the calculation results; Among them, the total required tasks are divided through a distributed framework, with the total number of tasks N as the input, and connected through a message middleware; the message middleware distributes the total number of tasks N to each task execution unit respectively, and monitors the running status of each task execution unit; when the operation of the task execution unit ends, the results are stored; The message middleware takes the parallelized difference operator and the number of times to be calculated, that is, the number of tasks of a task execution unit as the input, and through the resource scheduling model, allocates tasks according to the principle that the total time is the shortest when the CPU and GPU complete the tasks simultaneously, so as to realize the simultaneous calculation of the two; The resource scheduling model first extracts the difference operator to be calculated and then enters the commonality analysis of the difference operator, analyzes and extracts the commonality of the difference operator, and after the extraction is completed, makes a structured expression of the difference operator, and then generates multi-threaded implementation code and CUDA implementation code, slices the incoming data and assigns metadata tags to mark the position of the data block in the source data for subsequent merging; for the incoming of new data, first judge in the database whether the relevant parameters already configured exist for the Worker to be called, including the number of CPU blocks, the number of GPU blocks, and the number of threads created on the CPU; if the relevant parameters do not exist, send the task, let the Worker call the instrumentation program for fitting, and obtain these parameters from the obtained relevant results; then send them to the task queue together with the data according to these parameters and wait for the slave to pull the task to run.
2. A resource parallel scheduling and optimization system for large-scale differential operators, characterized in that It includes the following modules: Module 1 is used to extract the difference operator, analyze the common parts therein, take their common expressions, extract the differences as parameters, and convert the obtained difference operator into multi-threaded C code and code under CUDA; Module 2 is used to obtain the required tasks, that is, the number of times N that the difference operator needs to be repeatedly calculated; Send the required tasks to the cluster; Module 3 is used to divide the total required tasks, reasonably allocate the required tasks to the CPUs and GPUs of the task execution units in the cluster for calculation, and store the calculation results; Among them, the total required tasks are divided through a distributed framework, with the total number of tasks N as the input, and connected through a message middleware; the message middleware distributes the total number of tasks N to each task execution unit respectively, and monitors the running status of each task execution unit; when the operation of the task execution unit ends, the results are stored; The message middleware takes the parallelized difference operator and the number of times to be calculated, i.e., the number of tasks of a task execution unit, as inputs. Through the resource scheduling model, tasks are allocated according to the principle that the total time calculated is the shortest when the CPU and GPU complete tasks simultaneously, so as to achieve simultaneous calculation by both of them; For the resource scheduling model, first, the difference operators to be calculated are extracted and then enter the commonality analysis of the difference operators. The commonalities of the difference operators are analyzed and extracted. After the extraction is completed, the difference operators are structurally expressed. Then, multi-threaded implementation code and CUDA implementation code are generated. The incoming data is sharded and metadata tags are assigned to mark the position of the data block in the source data for subsequent merging; for the incoming of new data, first, it is judged in the database whether the relevant parameters already configured exist for the Worker to be called, including the number of CPU blocks, the number of GPU blocks, and the number of threads created on the CPU; if the relevant parameters do not exist, the task is sent, and the Worker is allowed to call the instrumentation program for fitting and obtain these parameters from the relevant results; then, according to these parameters, together with the data, they are sent to the task queue waiting for the slave to pull the task and run.
Citation Information
Patent Citations
GPU virtualization and resource scheduling method and device
CN111930522A
Methods, systems, and computer readable media for utilizing parallel adaptive rectangular decomposition (ARD) to perform acoustic simulations
US20160171131A1