Adaptive Load Balancing Method and Device for Ocean Model Operator
The adaptive load balancing method optimizes task distribution on heterogeneous clusters by using variance as a metric and CPU/GPU capability-based allocation, addressing inefficiencies and reducing execution times in ocean model calculations.
Patent Information
- Application Number
- CN202211207811.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-30
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-09-30
AI Technical Summary
In heterogeneous clusters, how to reasonably allocate ocean mode operator tasks to CPU and GPU to achieve load balancing and avoid idle resources and prolonged computing time.
Adaptive load balancing method is adopted to build a fine-grained model, use variance as a measurement indicator, and use the best task allocation algorithm to realize parallel computing of CPU and GPU, and optimize task allocation to achieve the minimum variance state.
It improves the computing efficiency of heterogeneous clusters, reduces task execution time, makes full use of the computing resources of CPU and GPU, and realizes load balancing.
Smart Images

Figure CN115525430B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of the computational complexity of ocean model operators, and in particular, to an adaptive load balancing method and device for ocean model operators. Background Art
[0002] The ocean model is a representative one of the Earth system models and is used to study the mechanism of ocean environmental evolution. By studying the characteristics of the ocean model operator tasks, it can be found that the operator tasks are large-scale compute-intensive tasks. In the process of computing the Earth system model, it is necessary to solve the numerical values of complex partial differential equations. The massive floating-point operations have high requirements for the performance of the processor; the process of numerically discretely solving the operator expression has high requirements for numerical accuracy. Usually, double-precision floating-point numbers are used to represent the operation results, and a large number of repeated dataset operations at high iteration times are required to obtain accurate values; the operator expression has strong decomposability. The basic operators can be encapsulated to obtain different composite operator expressions through permutation and combination, and each composite operator expression is decomposed into basic operators. Therefore, the operators have good parallelism and are suitable for parallel processing of operator tasks on a cluster. Using a high-performance cluster to process ocean model operator tasks can improve the computing speed.
[0003] The heterogeneous cluster is an effective way to solve large-scale complex computing tasks. In a heterogeneous cluster, the performance of each node is different, including computing power and memory size, etc., resulting in different computing times of tasks on different nodes. Inside the node, due to the performance differences between the CPU and the GPU, the two form a heterogeneous system. In order to make full use of the computing resources of the heterogeneous cluster and avoid the idle state of the CPU or the GPU, the CPU and the GPU of the node should both be used to process computing tasks. How to reasonably allocate tasks to the CPU and the GPU of each node to make the cluster reach the load balancing state and at the same time reduce the time for the cluster to complete tasks is a key point for improving the performance of the heterogeneous cluster. Therefore, developing an adaptive load balancing method and device for ocean model operators to effectively overcome the defects in the above-mentioned related technologies has become a technical problem urgently to be solved in the industry. Summary of the Invention
[0004] In view of the above problems existing in the prior art, the embodiments of the present invention provide an adaptive load balancing method and device for ocean model operators.
[0005] In a first aspect, the embodiments of the present invention provide an adaptive load balancing method for ocean model operators, including: constructing a fine-grained model to achieve fine-grained parallelism on a heterogeneous cluster; using variance as an index to measure the balance of the cluster load; and adopting an optimal task allocation algorithm to make the cluster adaptively reach the load balancing state.
[0006] Based on the content of the above method embodiments, the adaptive load balancing method for ocean model operators provided in the embodiments of the present invention, which realizes fine-grained parallelism on heterogeneous clusters, includes: creating a main thread on the CPU to run the serial part of the ocean model equation expression. When running into the parallelized operator task, the operator task is divided into two parts and processed on the CPU and GPU respectively. On the CPU, a predetermined number of threads are created, and the tasks assigned to the CPU are processed on each thread. On the GPU, the main thread is used to process the assigned tasks. After the CPU and GPU have processed their respective tasks, the multi-threads of the CPU end the tasks. At this time, the current operator task ends, and the next part of the equation is continued to run. This process is repeated until the entire ocean model expression is executed. In each node of the cluster, the CPU creates multi-threads, one of which is used to schedule the GPU, and the other threads execute the operator tasks. The GPU executes the operator tasks in parallel, and the operator tasks are run in parallel between all nodes.
[0007] Based on the content of the above method embodiments, the adaptive load balancing method for ocean model operators provided in the embodiments of the present invention, which uses variance as an index to measure the balance of the cluster load, includes: to maximize the mining of the computing resources of the cluster, the CPU multi-thread strategy is used, and the tasks assigned to the CPU are processed in parallel on the multi-threads to reduce the total task execution duration. When the CPU and GPU in the node complete the tasks and each node in the cluster completes the tasks, the cluster reaches the load balancing state. When the cluster is in the load balancing state, the CPU and GPU of each node will complete the tasks. Due to the different computing capabilities of the CPU and GPU, the number of tasks assigned is different. The computing power of the node depends on the computing capabilities of the CPU and GPU. The number of tasks assigned to each node is different. The stronger the computing power, the more tasks are assigned; conversely, the weaker the computing power, the fewer tasks are assigned. When the ratio of task assignment between the CPU and GPU and between nodes is reasonable, the variance of the running duration is the smallest, and all CPUs, GPUs, and nodes can complete the tasks. Then the task assignment algorithm follows the principle of minimum variance.
[0008] Based on the content of the above method embodiments, the adaptive load balancing method for ocean model operators provided in the embodiments of the present invention, which uses the optimal task assignment algorithm to make the cluster adaptively reach the load balancing state, includes: when the cluster reaches the load balancing state, assigning the number of tasks to each node, assigning the number of tasks to the CPU of each node, assigning the number of GPU tasks to each node, and assigning parameter values to each node.
[0009] Based on the content of the above method embodiments, the adaptive load balancing method for ocean model operators provided in the embodiments of the present invention further includes, after assigning parameter values to each node: when using the adaptive load balancing algorithm based on the minimum variance, preprocessing operator tasks on the CPUs and GPUs of each node in the cluster, calculating the running duration of a single operator task on the CPUs and GPUs of the node as a measure of the node's computing performance, obtaining the number of tasks assigned to the CPUs and GPUs within each node according to the number of tasks assigned to each node, and then obtaining the number of multi-threads created on each CPU. According to these parameters, the operator tasks run on the CPUs and GPUs of each node in the cluster according to the fine-grained parallel computing model, calculating the variance of the running duration of the cluster, and evaluating the performance of the adaptive load balancing algorithm based on the minimum variance.
[0010] In a second aspect, an embodiment of the present invention provides an adaptive load balancing device for ocean model operators, including: a first main module for constructing a fine-grained model to achieve fine-grained parallelism on a heterogeneous cluster; a second main module for using variance as an indicator to measure the balance of the cluster load; and a third main module for adopting an optimal task allocation algorithm to adaptively balance the load of the cluster to a balanced state.
[0011] In a third aspect, an embodiment of the present invention provides an electronic device, including:
[0012] at least one processor; and
[0013] at least one memory communicatively connected to the processor, where:
[0014] The memory stores program instructions executable by the processor, and the processor can execute the adaptive load balancing method for ocean model operators provided by any one of the various implementation manners in the first aspect by invoking the program instructions.
[0015] In a fourth aspect, an embodiment of the present invention provides a non-transitory computer-readable storage medium, and the non-transitory computer-readable storage medium stores computer instructions that cause a computer to execute the adaptive load balancing method for ocean model operators provided by any one of the various implementation manners in the first aspect.
[0016] The adaptive load balancing method and device for ocean model operators provided by the embodiments of the present invention are oriented to ocean model operators, run operator tasks on CPUs and GPUs, divide the operator tasks into two structures of CPU multi-threading and GPUs, and fully utilize the computing resources of CPUs and GPUs on heterogeneous clusters; use the variance of the running duration of tasks in the cluster to evaluate the load status of the cluster; calculate the number of tasks allocated on each node in the cluster and on the CPUs and GPUs within the nodes to minimize the variance of the running duration of the cluster. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 It is a flowchart of the adaptive load balancing method for ocean model operators provided by the embodiments of the present invention;
[0019] Figure 2 It is a schematic structural diagram of the adaptive load balancing device for ocean model operators provided by the embodiments of the present invention;
[0020] Figure 3 It is a schematic physical structure diagram of an electronic device provided by the embodiments of the present invention;
[0021] Figure 4 It is a schematic diagram of the intermediate calculation process principle of formula generation provided by the embodiments of the present invention;
[0022] Figure 5 It is a schematic diagram of the kernel fusion principle of the operation function provided by the embodiments of the present invention;
[0023] Figure 6 It is a schematic diagram of the minimum variance model structure provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. In addition, the technical features in each embodiment or individual embodiment provided by the present invention can be combined with each other arbitrarily to form a feasible technical solution. Such combination is not restricted by the order of steps and / or the pattern of structural composition, but must be based on the fact that those of ordinary skill in the art can implement it. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0025] An embodiment of the present invention provides an adaptive load balancing method for ocean model operators. Refer to Figure 1 , the method includes: constructing a fine-grained model to achieve fine-grained parallelism on a heterogeneous cluster; using variance as an indicator to measure the load balance of the cluster; and adopting an optimal task allocation algorithm to make the cluster adaptively reach a load-balanced state.
[0026] Based on the content of the above method embodiment, as an optional embodiment, in the adaptive load balancing method for ocean model operators provided in the embodiments of the present invention, the implementation of fine-grained parallelism on a heterogeneous cluster includes: creating a main thread on the CPU to run the serial part of the ocean model equation expression. When running to the parallelized operator task, the operator task is divided into two parts and processed on the CPU and GPU respectively. On the CPU, a predetermined number of threads are created, and the tasks assigned to the CPU are processed on each thread. On the GPU, the main thread is used to process the assigned tasks. After the CPU and GPU have processed their respective tasks, the multi-threads of the CPU end the tasks. At this time, the current operator task ends, and the next part of the equation is continued to run. This process is repeated until the entire ocean model expression is executed. In each node of the cluster, the CPU creates multi-threads, one of which is used to schedule the GPU, and the other threads execute the operator tasks. The GPU executes the operator tasks in parallel, and the operator tasks are executed in parallel between all nodes.
[0027] Based on the content of the above method embodiments, as an alternative embodiment, the adaptive load balancing method for ocean pattern operators provided in the embodiments of the present invention, which uses variance as an indicator to measure the balance of the cluster load, includes: to maximize the exploitation of the computing resources of the cluster, the CPU multi-threading strategy is used, and the tasks assigned to the CPU are processed in parallel on multiple threads to reduce the total task execution duration. When the CPU and GPU within a node complete the tasks, and each node within the cluster completes the tasks, the cluster reaches the load balancing state. When the cluster is in the load balancing state, the CPU and GPU of each node will complete the tasks. Due to the different computing capabilities of the CPU and GPU, the number of tasks assigned is different. The computing power of a node depends on the computing capabilities of the CPU and GPU, and the number of tasks assigned to each node is different. The stronger the computing power, the more tasks are assigned; conversely, the weaker the computing power, the fewer tasks are assigned. When the ratio of task allocation between the CPU and GPU and between nodes is reasonable, the variance of the running duration is the smallest, and all CPUs, GPUs, and nodes can complete the tasks, then the task allocation algorithm follows the principle of minimum variance.
[0028] Based on the content of the above method embodiments, as an alternative embodiment, the adaptive load balancing method for ocean pattern operators provided in the embodiments of the present invention, which uses the optimal task allocation algorithm to make the cluster adaptively reach the load balancing state, includes: when the cluster reaches the load balancing state, the number of tasks is assigned to each node, the number of tasks is assigned to the CPU of each node, the number of GPU tasks is assigned to each node, and parameter values are assigned to each node.
[0029] Based on the content of the above method embodiments, as an alternative embodiment, the adaptive load balancing method for ocean pattern operators provided in the embodiments of the present invention, after assigning parameter values to each node, further includes: when using the adaptive load balancing algorithm based on minimum variance, preprocessing operations are performed on the operator tasks on the CPU and GPU of each node in the cluster, and the running duration of a single operator task on the CPU and GPU of the node is calculated as a label for measuring the computing performance of the node. According to the number of tasks assigned to each node, the number of tasks assigned to the CPU and GPU within each node is obtained, and then the number of multi-threads created on each CPU is obtained. According to these parameters, the operator tasks run on the CPU and GPU of each node in the cluster according to the fine-grained parallel computing model, and the variance of the running duration of the cluster is calculated to evaluate the performance of the adaptive load balancing algorithm based on minimum variance.
[0030] The adaptive load balancing method for ocean model operators provided by the embodiments of the present invention is directed to ocean model operators, runs operator tasks on CPUs and GPUs, divides operator tasks into two structures: CPU multi-threading and GPUs, and fully utilizes the computing resources of CPUs and GPUs on heterogeneous clusters; uses the variance of the running duration of tasks in the cluster to evaluate the load status of the cluster; calculates the number of tasks allocated on each node in the cluster and on the CPUs and GPUs within the node, and minimizes the variance of the running duration of the cluster.
[0031] (1) Fine-grained parallel model
[0032] When ocean model operators run in a heterogeneous cluster, a fine-grained parallel strategy of CPU multi-threading and GPU parallel computing is adopted. Operator tasks are divided into CPU multi-threading computing tasks and GPU parallel computing tasks. Among them, the specific process of the fine-grained parallel strategy is as follows: First, create a main thread on the CPU to run the serial part of the ocean model equation expression. When running to the parallelized operator task, divide the operator task into two parts and process them on the CPU and GPU respectively. On the CPU, create a certain number of threads, and process the tasks allocated to the CPU on each thread. On the GPU, use the main thread to process the allocated tasks. After the CPU and GPU have processed the tasks respectively, the multi-threading of the CPU ends the task. At this time, the current operator task runs to completion, and continue to run the next part of the equation, repeating this process until the entire ocean model expression is executed. Within each node of the cluster, the CPU creates multi-threads, one of which is used to schedule the GPU, and the other threads execute operator tasks. The GPU executes operator tasks in parallel, and operator tasks run in parallel between all nodes.
[0033] In the ocean model equation, the input data is a three-dimensional matrix, which is represented by a three-dimensional array. Some mathematical operation processes are performed on the three-dimensional array. The common data operations are shown in Table 1.
[0034] Table 1
[0035]
[0036]
[0037] There are 12 types of difference operations for the partial differential equation of the ocean model. These 12 operations are encapsulated into basic operators, denoted as AXF, AXB, AYF, AYB, AZF, AZB, DXF, DXB, DYF, DYB, DZF, DZB, as shown in Table 2.
[0038] Table 2
[0039]
[0040]
[0041] In Table 2, according to whether the difference operation belongs to average difference or differential difference, which dimension of the matrix is selected, and whether it is forward or backward, there are 12 different difference operations, which are named with different letters. The first letter is D or A, indicating whether the operator belongs to a differential difference operator or an average difference operator. The second letter is X, Y, or Z, indicating the X direction, Y direction, or Z direction of the three-dimensional matrix. The third letter is F or B, indicating forward difference or backward difference. According to the selection of different letters in the three positions, the basic operator of the difference calculation is named with [A|D][X|Y|Z][F|B].
[0042] The following gives examples of these common data operations and 12 basic operators in ocean models.
[0043] In the calculation of the ocean model equation, a three-dimensional array must be created first, which is used as a three-dimensional matrix. The data type is Array, which includes a three-dimensional array for storing double-precision floating-point data, a pointer to the computational grid, a Message Passing Interface (MPI) communicator, and the size of the halo region. In the example, only the data of the three-dimensional array is shown. For a = seqs(m, n, k), the inputs m, n, and k are the sizes of each dimension in the three-dimensional matrix, and a three-dimensional array a of size m×n×k is returned, with each element incrementing sequentially starting from 0.
[0044] Create an array a of size 2×2×2, call display() to print the result, and the calculation result is expressed in the form of slicing the three-dimensional matrix. For the dimensions where k = 0 and k = 1, there are 8 elements in the i and j dimensions. It can be seen that starting from the first element 0, each subsequent element increments by 1 sequentially.
[0045] For AXB(a), it is the backward average difference of the three-dimensional array a in the X direction of the matrix, and the array a created in the first example is used as the input.
[0046] First, call seqs(2, 2, 2) to create an array a of size 2×2×2, call AXB(a) to perform the AXB operator operation on the array a, and then call display() to print the result. It can be seen that the array a is averaged backward in the X direction. For DYF(a), it is the forward differential difference of the three-dimensional array a in the Y direction of the matrix, and the array a created in the first example is used as the input.
[0047] First, call seqs(2, 2, 2) to create an array of size 2×2×2. Then call DYF(a) to perform the DYF operator operation on array a. Finally, call display() to print the result, and you can see that the differential operation is performed forward in the Y direction on array a. When the basic operators run on a heterogeneous cluster, a fine-grained parallel model is used. Each basic operator has a CPU serial version and a CUDA parallel version.
[0048] Under the fine-grained parallel strategy, the cluster distributes the computing tasks to all nodes. Each node divides the tasks into two parts. One part runs on the CPU multi-threads using the CPU serial code, and the other part runs on the GPU using the GPU parallel code. The two parts work together to complete the computing tasks.
[0049] For the ocean model composite operator expression, to describe the conversion process from the mathematical and physical equations of the ocean model to the operator expression, a partial differential equation describing the surface elevation of seawater is used as an example.
[0050] (where η is the height of the surface elevation of seawater, D is the depth of the seawater column, U is the zonal flow velocity of seawater, and V is the meridional flow velocity of seawater.)
[0051] When performing simulation calculations on a computer, the staggered Arakawa C grid scheme in the ocean model is used to discretize the formula. In the Arakawa C grid, D is calculated at the center, U is calculated on the left and right sides of D, V is calculated on the upper and lower sides of D, and a temporary variable tmpD is used instead. Perform a backward difference on the product of tmpD and U to obtain The discrete expression of:
[0052] tmpD(i + 1, j) = 0.5·(D(i + 1, j) + D(i, j))·U(i + 1, j) gives:
[0053]
[0054] where
[0055] dx(i, j) * = 0.5·(dx(i, j) + dx(i - 1, j))
[0056] After encapsulating with the basic operators, we get:
[0057] elf = elb - dt2·(DXF(AXB(D)·U) + DYF(AYB(D)·V))
[0058] Among them, elf is the result of solving the operator expression, the altitude at time t + 1, and elb is the altitude at time t - 1.
[0059] Simplify it to get:
[0060]
[0061]
[0062] Among them, δ is the difference operation, _ is the average operation, the superscripts x and y are the matrix directions, and the subscripts f and b are the forward or backward operations. For the forward difference operation in the x direction, For the backward difference operation in the y direction, For the forward average operation in the x direction, For the backward average operation in the y direction.
[0063] After encapsulating with basic operators, we get:
[0064] elf = elb - dt2·(DXF(AXB(D)·U) + DYF(AYB(D)·V))
[0065] Among them, elf is the result of solving the operator expression, the altitude at time t + 1, and elb is the altitude at time t - 1.
[0066] Save the operator expression as the formula for solving the surface altitude of seawater at a certain moment. When calculating the surface altitude of seawater at each moment, directly substitute the relevant values to obtain the result. During the calculation process, to obtain an accurate value, perform multiple loop iterations. Run the calculation tasks on a heterogeneous cluster. In the fine-grained parallel mode, put all tasks into each node of the cluster for processing. On each node, the tasks are divided into two parts. The operator expression is calculated using the basic operators of CPU serial and CUDA parallel. One part of the operator expression runs on the CPU multi-thread, and the other part runs on the GPU. Iteratively process the solution results of all tasks on the CPU main thread to obtain an accurate result.
[0067] For ocean model developers, according to the conversion relationship of basic operators, convert the discrete partial differential equations into operator expressions, generate operator program tasks on the computer, and finally execute the operator tasks on the heterogeneous cluster according to the fine-grained parallel strategy. During the conversion process, four steps should be followed.
[0068] 1. Convert the operator expression
[0069] For the partial differential equations for solving the ocean model state, they are converted into operator expressions according to the conversion rules shown in the previous subsection. For first-order differences, higher-order differences, and other more complex operations, they are obtained by arranging and combining basic operators. For example, based on the first-order difference it is used for the second-order difference, and the corresponding discrete expression is (var(i + 1, j, k)+var(i - 1, j, k)-2·var(i, j, k)) / dx 2 .
[0070] By converting all the ocean model equations into the form of operator expressions, without changing the semantics of the original equations, with a simple and direct correspondence relationship, it is not only convenient to remember but also easy to convert into program code, greatly simplifying the development and maintenance work of the ocean model.
[0071] 2. Construct the intermediate computation graph
[0072] In this process, the operator expression is converted into a directed acyclic graph, where each node in the graph stores the data of the array or the operation of the basic operator. Taking the formula elf = elb - dt2·(DXF(AXB(D)·U)+DYF(AYB(D)·V)) as an example, the conversion result is as Figure 4 shown. In Figure 4 It shows the process of converting the operator expression elf = elb - dt2·(DXF(AXB(D)·U) + DYF(AYB(D)·V)) into a computational graph. In the figure, the square boxes are the input and output variables. For example, elb is the altitude at time t - 1 as the input, D is the depth of the sea water column as the input, U is the zonal flow velocity of the sea water as the input, V is the meridional flow velocity of the sea water as the input, and elf is the altitude at time t + 1 as the final output. The circles in the figure are operators, all of which are operations on arrays. For example, "=" is the assignment function for arrays, "-" is the subtraction function for arrays, "*" is the multiplication function for arrays, "+" is the addition function for arrays, AXF and AYF are the average functions for arrays, and DXF and DYF are the differential functions for arrays. Through the arithmetic relationship between variables and operations, the entire operator expression is split and interpreted into the form of a computational graph, making the variable relationship and calculation process of the operator expression clearer and reducing the memory consumption of temporary arrays. For example, in the calculation process of DXF(AXB(D)·U), first calculate AXB(D) to obtain a temporary variable tmp1 of an array, then calculate tmp1·U to obtain a temporary variable tmp2, and finally calculate DXF(tmp2). During this process, two large array-type temporary variables are stored in memory, wasting memory space. Therefore, through the form of a computational graph, lazy evaluation is used to record the calculation process, and all function operations are overloaded to generate an intermediate computational graph, rather than obtaining the results of each function execution, and the variables are evaluated lazily when called by functions. Through this form, memory space is saved and the efficiency of operator operations is improved.
[0073] 3. Generate two-level parallel code
[0074] After analyzing the data dependencies through the computational graph, two-level parallel code at the underlying level needs to be generated. In the computational graph, the operation function of each circle is called a kernel. In the figure, all the kernels are fused into a large kernel function. When two kernel functions operate, the fused function will be called, and the shared data is stored in memory. This kernel function is called once to obtain the final output data, without temporary variables or intermediate variables, reducing the startup and scheduling overhead. The specific fusion process is as Figure 5 shown.
[0075] Analyze the computational graph to find the nodes that can be fused, store the results of the calculation in the subgraph, and when accessing a separate subgraph, assign the subgraph to the intermediate variables of the calculation process. By analyzing the subgraph and adopting just-in-time (JIT) compilation technology, the corresponding kernel function is generated, and multiple operator operations are fused into a large kernel function. Compared with executing each operator separately, using the form of a fused function reduces the memory bandwidth limitation and improves the computational performance.
[0076] The execution duration of each operator in the ocean model is relatively short. However, when running the operator expression task, due to the huge data volume of the three-dimensional matrix and the high number of iterations, the task running duration will be very long. Therefore, the duration overhead for generating the computational graph is very high, and it occupies a large amount of memory and performance during compilation and running of the kernel function. When calculating subgraphs, a fused kernel function is generated for each subgraph and placed in the function pool. If the same subgraph is used in subsequent calculations, the kernel function of this subgraph can be directly called in the memory pool, thus reducing the corresponding duration overhead. In addition, since the three-dimensional matrix is distributed in different computing units and involves multiple different operations on the data of a matrix, when the data of adjacent points in the matrix is used, the consistency of the matrix data facing multiple operator operations is checked before the fusion function. Regarding the partitioning of distributed matrix data, different strategies will bring different computational performances. Currently, the widely adopted method for partitioning data is the matrix block-based strategy, where the partial differential equation is solved in a structured grid, and different domain simulations in the grid process the allocated data blocks. During the actual operation process, the communication at the matrix boundary region is controlled. For example, in the average operator and differential operator, calculations are often performed on adjacent points of the matrix. Therefore, during the execution of the kernel function, a matrix boundary management mode is adopted to update and maintain the boundary information, and the communication problem of adjacent matrices is not concerned in the kernel function. In the boundary management mode, the update and maintenance of the boundary information are completed through asynchronous calls and communication. When fusing the kernel and the calling function, the process of asynchronous communication is hidden, and the internal details are transparent to developers, simplifying the research and development work of the ocean model.
[0077] In the final code generation stage, a lightweight code generation engine is used to analyze the computational graph and generate the corresponding code. The code file is dynamically compiled to obtain a shared library file, and the shared library file is dynamically linked to obtain the corresponding function for executing the expression tree. After passing in the parameters of the function, it is called and the result is returned. Each computational node of the computational graph is fused into the corresponding kernel function, and a unique hash value is calculated for each code file. When running to the same computational node each time, the code is not regenerated repeatedly. When running to the assignment operator corresponding to the final output result, the function codes of all computational nodes are executed at once.
[0078] The generated code will have different versions on different platforms. When the ocean model operator expression runs on the CPU, a code file with the.cpp suffix will be generated and compiled using gcc; when it runs on the GPU, a code file with the.cu suffix will be generated and compiled using nvcc. When the ocean model operator expression runs on a heterogeneous cluster, since both the CPU and GPU within each node participate in the running tasks, the generated code will include both.cpp and.cu code files, which will run serially on the CPU and in parallel on the GPU respectively. After creating multiple threads on the CPU, the.cpp code will be run within each thread. The parallel strategy for multiple threads is implemented on the CPU, combined with the GPU parallel strategy, to achieve a fine-grained two-level parallel code running strategy on the heterogeneous cluster.
[0079] 4. Map the fine-grained parallel template
[0080] To enable the ocean model operator to run in the environment of a heterogeneous cluster, the CPU and GPU computing resources of the heterogeneous cluster are fully utilized to transform the running mode of the ocean model operator. The fine-grained parallel strategy is used to make the operator adapt to the running mode of parallel computing on the CPU and GPU, and map the operator tasks to the fine-grained parallel model.
[0081] On the heterogeneous cluster, after the task division by the fine-grained parallel model, it shows the mode of cooperation between the CPU and GPU. The operator tasks are divided into CPU operator tasks and GPU operators, which run on the CPU and GPU respectively. "program CPU" is the program model on the CPU, "program GPU" is the program model on the GPU, "parallel" is the parallelized code block in the operator expression, and there are two different versions. Among them, "parallel cpp" is the operator code in the cpp version, and "parallelcuda" is the operator code in the cuda version. "t_count" is the number of threads created on the CPU. "ctasks" and "gtasks" are the tasks processed on the CPU and GPU respectively. "loop_func" is the core function of the CPU operator, and "__global__ void kernel_func" is the CUDA core code of the GPU operator.
[0082] First, within the main function of the CPU, calculate the appropriate number of multi-threads kthread created by the CPU (Chapter 4.3), assign it to t_count, create t_count threads using thread_create, call the loop_func function within each thread, run the parallel cpp loop code within the loop_func function, and process the operator tasks ctasks assigned to the CPU to implement the processing of operator tasks by CPU multi-threading. Then create another thread, call scheduling within the thread to schedule the operator tasks on the GPU. At this time, the code on the GPU starts to execute, call the __global__ void kenel_func function, run the parallel cuda loop code, and process the operator tasks gtask assigned to the GPU to implement the processing of operator tasks by the GPU. Finally, the operator tasks assigned to the CPU and the GPU are run and completed on the mapped two-level parallel template to implement the fine-grained parallel strategy on the heterogeneous cluster.
[0083] (2) Adaptive load balancing algorithm based on minimum variance
[0084] Due to the large performance differences between clusters and the large differences in the load of ocean model operator tasks, it is very difficult to uniformly and specifically describe whether the cluster is in a load-balanced state when processing tasks, or whether the degree of balance is high or low. For example, in a heterogeneous cluster, the performance of each node is different, resulting in a certain gap in the load-balanced state of each node. Therefore, when calculating the load balance of the entire cluster, a quantitative method is used to measure the load-balanced state of the cluster.
[0085] Variance is used to measure the deviation degree between a random variable and its mathematical expectation (mean). For a set of data, variance is the average of the squares of the differences between each data and the average. Let a set of data be x1, x2,..., x n , is the average, and the calculation formula for variance is:
[0086]
[0087]
[0088] The idea of minimum variance is used to measure the load balancing of the cluster. For a set of data on the running duration of cluster nodes, the running duration of the cluster depends on the node that finishes the task latest and has nothing to do with other nodes that finish the task earlier. If the variance is large, the differences in the running duration of each node are significant. At this time, some nodes are assigned too many tasks while some are assigned too few tasks, and the task allocation of the cluster is unreasonable. When the cluster has an unreasonable task allocation situation, nodes with relatively weak computing power are assigned more tasks, or nodes with relatively strong computing power are assigned fewer tasks. As a result, some nodes complete the assigned tasks in a shorter duration and then wait for other nodes to complete their tasks, during which period they are in an idle state, wasting computing resources. While other nodes are in a full-load state, with a large hardware pressure ratio, which easily leads to the aging and damage of the devices.
[0089] Therefore, when the variance of the running duration of the internal CPU and GPU of the node is the smallest, and the variance of the running duration of the cluster nodes is the smallest, it indicates that the task allocation of the cluster is the most reasonable and the cluster reaches the best load balancing state.
[0090] In a heterogeneous cluster, each node of the cluster is assigned computing tasks. Inside the node, the tasks are allocated to the CPU and GPU, and a part of the tasks are executed on the multi-threaded CPU while another part is executed on the GPU. The running duration of each node is the maximum value of the CPU and GPU, and the running duration of the entire cluster is the maximum value of the execution durations of all nodes.
[0091] For a single node, the running duration of the node depends on the one that finishes the task latest between the CPU and GPU and has nothing to do with the finishing duration of the other one. The variance of the two sets of data of the running durations of the CPU and GPU is the difference between the two values. If the difference between the two is too large, at this time, too many tasks are allocated to the CPU while too few are allocated to the GPU, or too many tasks are allocated to the GPU while too few are allocated to the CPU, indicating that the task allocation is unreasonable. The one that finishes the task earliest between the CPU and GPU waits for the other one to complete the task, resulting in an increase in the total running duration of the node.
[0092] Suppose there is a heterogeneous cluster C with n nodes in the cluster, which are N1, N2,..., N n , and each node has a CPU and a GPU. The cluster has a total of n CPUs and GPUs, which are CPU1, CPU2,..., CPU n , GPU1, GPU2,..., GPU n . Let T Ni be the running duration of the i-th node, T cpui be the running duration of the i-th CPU, and T gpui be the running duration of the i-th GPU, where 1 <= i <= n. Inside the i-th node, T cpui and T gpuiThe mean value of T cpui and T gpui The variance of
[0093] is the degree of deviation between the running times of the CPU and the GPU. When the variance is larger, the difference between T cpui and T gpui is larger, indicating that the task loads of the CPU and the GPU are more unbalanced, and the running time of the node is also longer. When the variance is smaller, T cpui and T gpui are closer, and the running time of the node is also shorter. If the variance reaches the minimum value of zero, T cpui and T gpui are the same, and the CPU and the GPU complete the task, indicating that the task loads of the CPU and the GPU reach the most balanced state at this time.
[0094] In cluster C, the mean value of T N1 , T N2 ,..., T Nn is T N1 , T N2 ,..., T Nn The variance of
[0095] is the degree of deviation of the running times of all nodes. When the variance is larger, the difference between T Ni is larger, indicating that the task loads of the cluster nodes are more unbalanced, and the running time of the cluster is also longer. When the variance is smaller, the running time of the cluster is shorter. When the variance reaches the minimum value of zero, the running times T Ni of each node are equal, and all nodes complete the task. At this time, the cluster reaches the best load balancing state.
[0096] The running time of the i-th node is the maximum of the running times of the i-th CPU and the i-th GPU: T Ni = max(T cpui , T gpui ), (1 <= i <= n). The running time of the cluster, which is the maximum of the running times of all nodes, is T c = max(T N1 , T N2 ,..., T Nn ). Under this condition, the minimum variance model of the heterogeneous cluster is as shown in Figure 6 as follows.
[0097] In Figure 6In order to maximize the utilization of the computing resources in the cluster, a CPU multi-threading strategy is adopted. The tasks assigned to the CPU are processed in parallel on multiple threads, reducing the total task execution time. When the CPU and GPU within a node complete their tasks, and each node in the cluster has completed its tasks, the cluster reaches a load balancing state. When the cluster is in the load balancing state, the CPU and GPU of each node will complete their tasks, and the running times of all CPUs and GPUs are equal. Due to the different computing capabilities of the CPU and GPU, the number of tasks assigned is different. The computing power of a node depends on the computing capabilities of the CPU and GPU, and the number of tasks assigned to each node is different. Generally speaking, the stronger the computing power, the more tasks are assigned; conversely, the weaker the computing power, the fewer tasks are assigned. When the ratio of task allocation between the CPU and GPU, and between nodes is reasonable, the variance of the running time is minimized, and all CPUs, GPUs, and nodes can complete their tasks. Therefore, the task allocation algorithm follows the principle of minimum variance.
[0098] (III) Optimal Task Allocation Algorithm
[0099] To achieve the optimal load balancing state in a heterogeneous cluster, according to the minimum variance model, the running times of each node are made equal, and the running times of the CPU and GPU within a node are made equal. Calculate the number of tasks assigned to each node, and the number of tasks assigned to the CPU and GPU within each node. The variables and their meanings used in the minimum variance load balancing algorithm are shown in Table 3.
[0100] Table 3
[0101]
[0102]
[0103]
[0104] Let s be the total number of tasks in the entire cluster, s i be the number of tasks assigned to the i-th node, s cpui be the number of tasks assigned to the CPU on the i-th node, s gpui be the number of tasks assigned to the GPU on the i-th node. t cpui be the running time of a single task on the single thread of the i-th CPU, t gpui the running time of a single task on the i-th GPU. k i be the number of threads of the i-th CPU, k i is less than the number of cores of the CPU. In operator operations, the overhead of multi-threading is much smaller than the CPU operation time, so it is ignored.
[0105] Within each node of the cluster, the sum of the number of tasks assigned to the CPU and the number of tasks assigned to the GPU is the number of tasks assigned to the current node, s cpui +s gpui =s i ; The speedup ratio of the i-th GPU is the ratio of the running time of a single task on the single-thread of the i-th CPU and the running time of a single task on the i-th GPU: The multi-thread speedup ratio of the i-th CPU is the ratio of the running time of a task on multi-threads and on a single thread, the number of multi-threads:
[0106] Since the CPU uses multi-threaded computing, the running time of the i-th CPU is the product of the number of tasks assigned to the CPU and the running time of a single task on the single-thread of the i-th CPU divided by the number of threads of the i CPUs: T cpui =s cpui t cpui / k i ; The running time of the i-th GPU is the product of the number of tasks assigned to the GPU and the running time of a single task on the i-th GPU: T gpui =s gpui t gpui ; The running time of the i-th node is the maximum of the running times of the CPU and the GPU: T Ni =max{T cpui ,T gpui}=max{s cpui t cpui / k i ,s gpui t gpui}.
[0107] According to the minimum variance model, when the running times of the CPU and the GPU are equal, the variance is the smallest, and the running time T Ni of the i-th node is the shortest, T cpui =T gpui s cpui t cpui / k i =s gpui t gpui , where t cpui 、t gpui and k i are known. From the above formula, the number of tasks assigned to the CPU is: The number of tasks assigned to the GPU are respectively:
[0108] At this time, the speedup ratio of the i-th node is the largest, which is the ratio of the running time of all tasks on the CPU and the shortest running time of the node, the sum of the CPU speedup ratio and the GPU speedup ratio:
[0109] When using the minimum variance-based adaptive load balancing algorithm, preprocessing operations are performed on operator tasks on the CPUs and GPUs of each node in the cluster, and the running durations of individual operator tasks on the CPUs and GPUs of the nodes are calculated as a measure of the computing performance of the nodes. According to the number of tasks assigned to each node, the number of tasks assigned to the CPUs and GPUs within each node is obtained, and then the number of multi-threads created on each CPU is obtained. Based on these parameters, the operator tasks run on the CPUs and GPUs of each node in the cluster according to the fine-grained parallel computing model. Finally, the variance of the running duration of the cluster is calculated to evaluate the performance of the minimum variance-based adaptive load balancing algorithm. The research of the present invention is based on ocean model operators and uses a heterogeneous cluster as the computing platform to establish a minimum variance-based adaptive load balancing algorithm.
[0110] The implementation basis of each embodiment of the present invention is achieved through programmed processing by a device with processor functions. Therefore, in engineering practice, the technical solutions and functions of each embodiment of the present invention can be encapsulated into various modules. Based on this actual situation, on the basis of the above embodiments, the embodiments of the present invention provide an adaptive load balancing device for ocean model operators, and this device is used to execute the adaptive load balancing method for ocean model operators in the above method embodiments. See Figure 2 , this device includes: a first main module for constructing a fine-grained model to achieve fine-grained parallelism on a heterogeneous cluster; a second main module for using variance as an index to measure the balance of the cluster load; a third main module for using the optimal task allocation algorithm to adaptively achieve the load balancing state of the cluster.
[0111] The adaptive load balancing device for ocean model operators provided by the embodiments of the present invention adopts Figure 2 several modules among them, is oriented to ocean model operators, runs operator tasks on CPUs and GPUs, divides the operator tasks into two structures of CPU multi-threads and GPUs, and fully utilizes the computing resources of CPUs and GPUs on the heterogeneous cluster; uses the variance of the running duration of tasks in the cluster to judge the load status of the cluster; calculates the number of tasks assigned to each node in the cluster and to the CPUs and GPUs within the node to minimize the variance of the running duration of the cluster.
[0112] It should be noted that the device in the device embodiment provided by the present invention can be used not only to implement the method in the above method embodiment, but also to implement the methods in other method embodiments provided by the present invention. The difference lies only in setting corresponding functional modules, and its principle is basically the same as that of the above device embodiment provided by the present invention. As long as those skilled in the art, on the basis of the above device embodiment, refer to the specific technical solutions in other method embodiments, obtain corresponding technical means by combining technical features, and the technical solutions composed of these technical means, and ensure the practicability of the technical solutions, they can improve the device in the above device embodiment, so as to obtain corresponding device type embodiments for implementing the methods in other method type embodiments. For example:
[0113] Based on the content of the above device embodiment, as an optional embodiment, the adaptive load balancing device for ocean mode operators provided in the embodiment of the present invention further includes: a first sub-module for implementing fine-grained parallelism on the heterogeneous cluster, including: creating a main thread on the CPU to run the serial part of the ocean mode equation expression. When running to the parallelized operator task, the operator task is divided into two parts and processed on the CPU and GPU respectively. On the CPU, a predetermined number of threads are created, and each thread processes the task assigned to the CPU. On the GPU, the main thread is used to process the assigned task. After the CPU and GPU each complete the task, the multi-thread of the CPU ends the task. At this time, the current operator task runs to completion, and the next part of the equation is continued to run. This process is repeated until the entire ocean mode expression is executed. Within each node of the cluster, the CPU creates multi-threads, where one thread is used to schedule the GPU and the other threads execute operator tasks. The GPU executes operator tasks in parallel, and operator tasks are run in parallel between all nodes.
[0114] Based on the content of the above device embodiments, as an alternative embodiment, the adaptive load balancing device for ocean mode operators provided in the embodiments of the present invention further includes: a second sub-module for implementing using variance as an index to measure the balance of cluster load, including: to maximize the excavation of the computing resources of the cluster, using the CPU multi-threading strategy, the tasks assigned to the CPU are processed in parallel on multiple threads to reduce the total task execution duration. When the CPU and GPU within a node complete the tasks, each node within the cluster completes the tasks, and the cluster reaches the load balance state. When the cluster is in the load balance state, the CPU and GPU of each node will complete the tasks. Due to the different computing capabilities of the CPU and GPU, the number of tasks assigned is different. The computing power of a node depends on the computing capabilities of the CPU and GPU. The number of tasks assigned to each node is different. The stronger the computing power, the more tasks are assigned; conversely, the weaker the computing power, the fewer tasks are assigned. When the ratio of task assignment between the CPU and GPU and between nodes is reasonable, the variance of the running duration is the smallest, and all CPUs, GPUs, and nodes can complete the tasks, then the task assignment algorithm follows the principle of minimum variance.
[0115] Based on the content of the above device embodiments, as an alternative embodiment, the adaptive load balancing device for ocean mode operators provided in the embodiments of the present invention further includes: a third sub-module for implementing using the optimal task assignment algorithm to make the cluster adaptively reach the load balance state, including: when the cluster reaches the load balance state, assigning the number of tasks to each node, assigning the number of tasks to the CPU of each node, assigning the number of GPU tasks to each node, and assigning parameter values to each node.
[0116] Based on the content of the above device embodiments, as an alternative embodiment, the adaptive load balancing device for ocean mode operators provided in the embodiments of the present invention further includes: a fourth sub-module. After assigning parameter values to each node, it further includes: when using the adaptive load balancing algorithm based on minimum variance, performing a preprocessing operation on the operator tasks on the CPU and GPU of each node in the cluster, calculating the running duration of a single operator task on the CPU and GPU of the node as a label for measuring the computing performance of the node. According to the number of tasks assigned to each node, obtaining the number of tasks assigned to the CPU and GPU within each node, and then obtaining the number of multi-threads created on each CPU. According to these parameters, the operator tasks run on the CPU and GPU of each node in the cluster according to the fine-grained parallel computing model, calculating the variance of the running duration of the cluster, and evaluating the performance of the adaptive load balancing algorithm based on minimum variance.
[0117] The method of the embodiments of the present invention is implemented relying on an electronic device. Therefore, it is necessary to introduce the relevant electronic device. For this purpose, an embodiment of the present invention provides an electronic device, as Figure 3 shown. The electronic device includes: at least one processor, a communications interface, at least one memory, and a communication bus. Among them, the at least one processor, the communications interface, and the at least one memory complete mutual communication through the communication bus. The at least one processor can call the logical instructions in the at least one memory to execute all or part of the steps of the methods provided in the foregoing method embodiments.
[0118] In addition, when the logical instructions in the foregoing at least one memory are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the method embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.
[0119] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0120] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0121] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to multiple embodiments of the present invention. Based on this understanding, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and sometimes they may be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0122] It should be noted that the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, the elements defined by the statement "comprising..." do not exclude the presence of additional identical elements in the process, method, article or device comprising the said elements.
[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An adaptive load balancing method for ocean model operators, characterized in that Including: Solving the partial differential equation of the ocean model state, and converting it into an operator expression according to the conversion rules; Based on the arithmetic relationship between variables and operations, splitting and interpreting the entire operator expression into the form of a computational graph, where the computational graph is a directed acyclic graph, and each node in the graph stores the data of an array or the operation of a basic operator; Analyzing the data dependence through the computational graph to generate low-level two-level parallel code, where the parallel code is the code running on the CPU and the code running on the GPU; Creating a main thread on the CPU to run the serial part of the ocean model equation expression. When running to the parallelized operator task, taking the above code as the operator task, dividing the operator task into two parts, and processing them on the CPU and GPU respectively. On the CPU, creating a predetermined number of threads, and processing the tasks assigned to the CPU on each thread. On the GPU, using the main thread to process the assigned tasks. After the CPU and GPU have processed their respective tasks, the multi-threading of the CPU ends the task. At this time, the current operator task runs to completion, and then continues to run the next part of the equation. Repeating this process until the entire ocean model expression is executed. Within each node of the cluster, the CPU creates multi-threads, where one thread is used to schedule the GPU, and the other threads execute the operator tasks. The GPU executes the operator tasks in parallel, and the operator tasks are run in parallel between all nodes; Calculating the running duration of a single operator task on the CPU and GPU of a node as a standard for measuring the computing performance of the node. According to the number of tasks assigned to each node, obtaining the number of tasks assigned to the CPU and GPU within each node, and then obtaining the number of multi-threads created on each CPU. Based on these parameters, calculating the variance of the running durations of the CPU and GPU of the cluster. When the variance is the smallest, it indicates that the task allocation of the cluster is the most reasonable, and the cluster reaches the best load balancing state.
2. The adaptive load balancing method for ocean-oriented mode operators according to claim 1, wherein The method further includes: To maximize the utilization of the computing resources of the cluster, using the CPU multi-threading strategy, processing the tasks assigned to the CPU in parallel on the multi-threads to reduce the total task execution duration. When the CPU and GPU within a node complete the tasks, and each node within the cluster completes the tasks, the cluster reaches the load balancing state. When the cluster is in the load balancing state, the CPU and GPU of each node will complete the tasks. Due to the different computing capabilities of the CPU and GPU, the number of tasks assigned is different. The computing power of a node depends on the computing capabilities of the CPU and GPU. The number of tasks assigned to each node is different. The stronger the computing power, the more tasks are assigned; conversely, the weaker the computing power, the fewer tasks are assigned. When the ratio of task allocation between the CPU and GPU and between nodes is reasonable, the variance of the running duration is the smallest, and all CPUs, GPUs, and nodes can complete the tasks, then the task allocation algorithm follows the principle of the smallest variance.
3. The adaptive load balancing method for the ocean-oriented mode operator according to claim 2, wherein The method further includes: When the cluster reaches the load balancing state, assigning the number of tasks to each node, assigning the number of tasks to the CPU of each node, assigning GPU tasks to each node, and assigning parameter values to each node.
4. An adaptive load balancing device for ocean model operators, characterized in that, The adaptive load balancing device for the ocean-oriented mode operator is used to execute the method described in any one of claims 1 to 3.
5. An electronic device, characterized in that, It includes: At least one processor, at least one memory, and a communication interface; wherein, The processor, the memory, and the communication interface communicate with each other; The memory stores program instructions executable by the processor, and the processor calls the program instructions to execute the method described in any one of claims 1 to 3.
6. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the method described in any one of claims 1 to 3.
Citation Information
Patent Citations
Sea flow induced magnetic field calculation method and system
CN113094915A
Automatic computing resource allocation method and system for three-level parallel middleware
CN114356550A