Computing power distribution method

By adopting technologies such as load balancing and GPU resource virtualization in universities, research institutions, and small and medium-sized enterprises, the resource integration problem of distributed and scattered devices has been solved, efficient computing power allocation and resource management have been achieved, and computing efficiency and resource utilization have been improved.

CN120803709APending Publication Date: 2025-10-17LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510900756.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

In universities, research institutions, and small and medium-sized enterprises, scattered high-performance computing equipment makes it difficult to form effective resource integration, resulting in low task execution efficiency and waste of resources. The existing computing power allocation method lacks a refined scheduling mechanism.

Method used

Through load balancing strategies, the target computing tasks are determined to be executed on multiple servers, divided into multiple operator tasks, and the execution order is determined based on resource requirements and dependencies. Combined with asynchronous API requests, GPU resource virtualization, short request packaging and multi-task GPU sharing, efficient computing power allocation is achieved.

Benefits of technology

It improves GPU resource utilization, reduces task execution delays, enhances system high availability and stability, provides on-demand elastic computing power services, and improves computing efficiency and resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803709A_ABST
    Figure CN120803709A_ABST
Patent Text Reader

Abstract

The invention discloses a computing power distribution method which is applied to a first server side, and the method comprises the steps that a second server side for executing a target computing task is determined at a plurality of server sides based on a load balancing strategy, and the first server side and the second server side are the same or different; the target calculation task is acquired from a client which is arranged in the same data sharing network with the server; dividing the target calculation task into a plurality of operator tasks; determining a first execution sequence of the plurality of operator tasks based on a resource demand and a dependency relationship of each operator task, the resource demand at least comprising a computing power demand and a storage demand, and the dependency relationship being used for representing the execution sequence of the operator tasks; executing the operator tasks based on the first execution sequence to obtain an execution result; and returning the execution result to the client.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The embodiment of the present application relates to the field of computer system resource management, and relates to but is not limited to a computing power allocation method. BACKGROUND

[0002] With the growing demand for artificial intelligence and big data processing, high-performance computing resources such as graphics processing units (GPUs) play an increasingly important role in various computing tasks. In particular, in environments such as universities, research institutions, and small and medium-sized enterprises, there are often multiple independently deployed high-performance devices, which are often scattered and have low utilization, making it difficult to form effective resource integration.

[0003] Current computing power allocation methods mostly lack fine scheduling mechanisms for task dependency relationships and resource requirements, resulting in low task execution efficiency, serious resource waste, and difficulty in meeting the efficient execution needs of complex computing tasks. SUMMARY

[0004] Therefore, the embodiment of the present application provides a computing power allocation method.

[0005] The technical scheme of the embodiment of the present application is as follows:

[0006] In a first aspect, the embodiment of the present application provides a computing power allocation method applied to a first server, the method comprising: determining, based on a load balancing strategy, a second server from a plurality of servers to execute a target computing task, wherein the first server and the second server are the same or different, and the target computing task is obtained from a client disposed in a same data sharing network as the server; dividing the target computing task into a plurality of operator tasks; determining a first execution order of the plurality of operator tasks based on resource requirements and dependency relationships of each operator task, wherein the resource requirements at least include computing power requirements and storage requirements, and the dependency relationships are used to represent the execution order of the operator tasks; executing the plurality of operator tasks based on the first execution order to obtain an execution result; and returning the execution result to the client.

[0007] In a second aspect, the embodiment of the present application provides a computing power allocation method applied to a client, the method comprising: obtaining a plurality of computing requests corresponding to a plurality of computing tasks; dividing the plurality of computing requests into a plurality of computing request groups based on the task properties of the computing tasks; dividing each computing request group into a plurality of packaged requests based on a packaged data volume threshold and a request format; and sending the plurality of packaged requests to a first server, wherein the packaged request includes a target computing task.

[0008] In a third aspect, the embodiment of the present application provides a computing power allocation device, comprising:

[0009] The first determining module is configured to determine, based on a load balancing strategy, a second server for executing a target computing task from a plurality of servers, wherein the first server and the second server are the same or different, and the target computing task is obtained from a client disposed in a same data sharing network as the servers.

[0010] The task division module is configured to divide the target computing task into a plurality of operator tasks.

[0011] The second determining module is configured to determine a first execution order of the plurality of operator tasks based on resource requirements and a dependency relationship of each operator task, wherein the resource requirements at least include computing power requirements and storage requirements, and the dependency relationship is used to represent the execution order of the operator tasks.

[0012] The task execution module is configured to execute the plurality of operator tasks based on the first execution order to obtain an execution result.

[0013] The result returning module is configured to return the execution result to the client.

[0014] In a fourth aspect, an example of the present application provides a computing power allocation apparatus, comprising:

[0015] The second obtaining module is configured to obtain a plurality of computing requests corresponding to a plurality of computing tasks.

[0016] The first division module is configured to divide the plurality of computing requests into a plurality of computing request groups based on task properties of the computing tasks.

[0017] The second division module is configured to divide each computing request group into a plurality of packaged requests based on a packaged data volume threshold and a request format.

[0018] The sending module is configured to send the plurality of packaged requests to a first server, wherein the packaged requests include target computing tasks.

[0019] In a fifth aspect, an example of the present application provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program capable of running on the processor, and the processor implements the above method when executing the program.

[0020] In a sixth aspect, an example of the present application provides a storage medium, which stores executable instructions for implementing the above method when executed by a processor.

[0021] In a seventh aspect, an example of the present application provides a computer program product, comprising a computer program or instructions, which, when executed by a processor, implements the steps in the above method. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1AA software architecture schematic diagram of the AI computing power sharing and resource management platform in a local area network provided by the embodiment of the present application is provided.

[0023] Figure 1B A schematic diagram of the logic control flow of the computing power allocation provided by the embodiment of the present application is provided.

[0024] Figure 1C A flow schematic diagram of the computing power allocation method provided by the embodiment of the present application is provided.

[0025] Figure 2 A flow schematic diagram of the operator task division provided by the embodiment of the present application is provided.

[0026] Figure 3 A flow schematic diagram of the operator task execution order determination provided by the embodiment of the present application is provided.

[0027] Figure 4 An implementation flow schematic diagram of the computing power allocation method provided by the embodiment of the present application is provided.

[0028] Figure 5A A component structure schematic diagram of the computing power allocation device provided by the embodiment of the present application is provided.

[0029] Figure 5B A component structure schematic diagram of the computing power allocation device provided by the embodiment of the present application is provided.

[0030] Figure 6 A hardware entity schematic diagram of the electronic device provided by the embodiment of the present application is provided. DETAILED DESCRIPTION

[0031] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the specific technical scheme of the embodiments of the present application will be further described in detail below with reference to the drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application, but not to limit the scope of the present application.

[0032] In the following description, "some embodiments" are described, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subset of all possible embodiments, and can be combined with each other without conflict.

[0033] In the following description, the terms "first\second\third" are only to distinguish similar objects, and do not represent a specific order of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence as allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description herein is for describing the embodiments of the application only and is not intended to be limiting of the application.

[0035] Figure 1A A software architecture diagram of an AI computing power sharing and resource management platform in a local area network provided by an embodiment of the application is shown in FIG. 1, which includes an application layer 11, a system architecture layer 12, an operating system layer 13, and a device layer 14, wherein, Figure 1A

[0036] An artificial intelligence (AI) computing power sharing and resource management platform is directed to an application using a compute unified device architecture (CUDA). An intermediate layer, i.e., the system architecture layer 12, is constructed, which is responsible for the scheduling and distribution of a desktop multi-GPU computing power group in the local area network. The system architecture layer 12 can reasonably schedule and distribute the computing power required by the application deployed on the client 121 to other high-performance computing power servers 122 in the local area network. The application using CUDA is deployed on the client 121, which can call the API function interface (cuBLAS to CUDA Runtime API) of CUDA. The server 122 is deployed on a high-performance GPU computing power desktop.

[0037] The system architecture layer 12 respectively shows the architecture of the client 121 and the server 122. As shown in FIG. 2, the dashed box in the client 121 is newly added in the embodiment of the application, which includes the application programming interface (API) function interface (cuBLAS to CUDA Runtime API) of CUDA, asynchronous API non-waiting call, and short request packaging. The dashed box in the server 122 is newly added in the embodiment of the application, which includes multi-GPU load balancing, GPU load sensing, and multi-task GPU sharing. Figure 1A

[0038] A schematic diagram of a computing power distribution logic control flow provided by an embodiment of the application is shown in FIG. 3, which includes: Figure 1B Figure 1B

[0039] The computing power demand side 15, the local area network GPU computing power management platform 16, and the computing power supply side 17, wherein,

[0040] The computing power demand side 15, i.e., Figure 1A ​​​The client 121 shown in FIG. 1. The power demander can provide the user interface for the user of the design class, the material science user and the biology user to use the power function or the model function.

[0041] The power supplier 17, that is Figure 1A The server 122 shown in FIG. 1, is configured to provide the power for the power demander 15.

[0042] The local area network GPU power management platform 16 can be configured in the local area network to meet the power distribution requirements of any one of Figure 1A The server 122 shown in FIG. 1, is configured to provide the power for the power demander 15.

[0043] The following describes Figure 1B The pseudo-CUDA library request interception, API request asynchronous response, GPU resource virtual mapping, short request packaging and remote calling, GPU load sensing, multi-task GPU sharing and multi-GPU load balancing described in FIG. 1.

[0044] I. Pseudo-CUDA library request interception

[0045] 1. The pseudo-CUDA request interception mechanism is as follows:

[0046] On the client side, all CUDA API calls (such as cudaMalloc, cudaMemcpy, cudaLaunchKernel, etc.) are usually invoked by directly linking the CUDA library. In order to forward these requests to the remote GPU for processing, the pseudo-CUDA library request interception can be used.

[0047] Pseudo-CUDA library implementation: create a pseudo-CUDA library on the client side to intercept all CUDA API function calls. When these functions are called, the pseudo-CUDA library can replace the native CUDA library (located on the client side, corresponding to Figure 1A GPU request interception) to forward the request initiated by the client to the scheduler on the remote server through the Remote Procedure Call (RPC) protocol instead of directly executing on the local GPU. This approach ensures the transparency and seamless operation of the client application. The RPC is a communication protocol that allows programs to call processes or functions located in different address spaces (usually different physical machines) on the network, just like calling local functions; the scheduler can be set on the server side, and the initiated request only performs the next step of distribution and scheduling in the scheduler, and the real operation is performed by the GPU set on the server side.

[0048] 2. The implementation of pseudo-CUDA library request interception is as follows:

[0049] Intercept and forward API calls using open source libraries such as cuda_hook, forward all CUDA calls to the scheduler server.

[0050] Intercept and forward flow: the client application calls the CUDA API (such as cudaMalloc); the pseudo-CUDA library intercepts this call and encapsulates it as an RPC request; the request is sent to the remote scheduler; the scheduler decides which GPU to execute this task and forwards the request to the target GPU set on the server for processing.

[0051] II. API request asynchronous response.

[0052] 1. The asynchronous API request mechanism is as follows:

[0053] In order to improve the response speed and throughput of the system, the asynchronous API request mechanism is used to avoid the delay and blocking caused by synchronous calls. Especially when multiple small tasks (such as memory allocation, Kernel launch, etc.) occur, asynchronous requests can significantly reduce the waiting time. Here, not all requests use the asynchronous API request mechanism, which depends on the specific task type, and in real scenarios, a mix of asynchronous and synchronous is mainly used. Asynchronous is suitable for compute-intensive tasks (such as matrix operations, deep learning inference, etc.), scheduling responses to remote resources, etc., while synchronous is suitable for IO-intensive operations (such as data transfer, etc.), some real-time tasks that require strong consistency of results.

[0054] Working mechanism: when requesting an API, the client does not wait for the response of the remote GPU, but sends the request as an asynchronous task. The client then continues to perform other tasks. When the GPU finishes processing and returns the result, the server notifies the client through the asynchronous callback mechanism, or obtains the result through polling. This can avoid the "idle waiting" time in GPU requests and reduce the total execution time of tasks. Asynchronous call is a kind of message or event mechanism, which solves the problem of synchronous blocking.

[0055] 2. The application scenarios of asynchronous request are as follows:

[0056] For memory allocation, data transfer, etc., instead of waiting for the remote GPU to complete the allocation, placeholders and / or delayed binding can be used to improve efficiency.

[0057] III. GPU resource virtual mapping.

[0058] In the implementation process, such as Figure 1BThe illustrated computing power request passes through the computing power demand side 15, and then is uniformly scheduled through the local area network GPU computing power management platform 16, and according to the scheduling result, the request is forwarded to the actual computing power supply side 17 for execution. In this process from the computing power demand side 15 to the local area network GPU computing power management platform 16, and then to the computing power supply side 17, the GPU resources of the server (the computing power supply side 17) need to be virtually mapped into a virtual address, and this address can be perceived by the client (the computing power demand side 15) and the local area network GPU computing power management platform 16, so that the request from the client can be forwarded to the actual computing power supply side 17 for execution.

[0059] 1. The virtual GPU resource mapping technology is as follows:

[0060] The GPU resource virtualization technology can virtualize remote physical GPU resources as logical resources of a local system. Each client request can access the remote GPU through a virtual address space without needing to care about the physical location of the GPU. Here, the scheduler can map the remote GPU as a virtual GPU address space that can be recognized by the local system (the client), that is, when the local system accesses these virtual addresses, the remote GPU can be recognized, which is similar to that the local system virtually mounts the GPU, but this GPU belongs to the remote server.

[0061] The virtualization principle is as follows: the resources (such as memory and computing units) of each GPU are managed by virtualization at the scheduler level. The client interacts with the remote GPU through a virtualized interface, which is similar to executing a task on a local GPU, that is, the client request is executed on the virtual resource mapped by the client. In this way, the remote GPU resources can be shared among multiple clients, thereby effectively optimizing the utilization of GPU resources.

[0062] 2. Resource isolation: the GPU tasks of each client are isolated through a virtual address space to avoid resource competition between tasks. Here, isolation means that different tasks are independently stored in a virtual address space.

[0063] Four, short request packaging and remote calling.

[0064] 1. Short request packaging.

[0065] Here, short request packaging is to reduce the communication overhead by packaging multiple small requests (such as multiple Kernel calls or inference requests) into a batch request to reduce the communication overhead of each request. In the implementation process, the packaged request can be on the CPU of the server where the scheduler is located, or on the client.

[0066] The scheduler or client combines multiple small requests into a request package after receiving them. The request package contains multiple small CUDA operations or inference tasks. The remote GPU returns a unified result to the client after executing the batch request.

[0067] The packaging strategy includes task grouping, request merging, defining a unified request format, and data packaging.

[0068] Task grouping: The scheduler or client can group tasks based on their properties (e.g., computation type, data dependency, etc.) after receiving multiple short requests. The short requests are divided into several logical groups, which usually contain multiple small CUDA operations, memory copies, or inference requests.

[0069] Request merging: Multiple small requests can be combined into a batch request, which can be a combination of multiple independent CUDA kernel calls, data transfers (such as cudaMemcpy), or inference requests into a "package". Define a unified request format: Batch requests use a standardized request format, including input data, task type, parameters, and computation order for all small tasks. For example, multiple CUDA kernel parameters can be packaged into a data structure and executed sequentially on the GPU.

[0070] Data packaging: In addition to merging tasks, related input data (such as tensors, matrices, etc.) is also packaged together. For multiple inference tasks, input data can be combined into a large tensor or matrix for batch processing, which can reduce the data transfer overhead for each request.

[0071] The request package contains multiple small CUDA operations or inference tasks. The remote GPU returns a unified result to the client after executing the batch request.

[0072] 2、Optimization effects are as follows:

[0073] Reducing latency: By batching requests, the latency of establishing a connection for each small request is reduced.

[0074] Improving performance: By reducing the number of communications, the network transmission load is reduced, and the system throughput is improved.

[0075] Five, GPU load awareness.

[0076] 1、Real-time GPU fine-grained resource monitoring is as follows:

[0077] GPU overall resource utilization: Obtain GPU overall utilization, memory usage, memory bandwidth, etc. through NVML (NVML is a monitoring tool for GPU hardware, which can view GPU temperature, power, fan, etc.) or DCGM (DCGM is more inclined to monitor and manage clusters, which can perform health checks, configure policies, and integrate with tools such as K8S). Real-time monitoring of GPU resource occupation.

[0078] Task layer monitoring: By tracking the demand of different tasks for GPU resources, the execution of the task can be judged. For example: memory bandwidth utilization: monitoring the memory read-write speed of the task, which can indirectly reflect the speed of task execution.

[0079] Kernel execution rate, i.e. the efficiency of the operating system kernel when processing tasks: estimate the execution speed of the task by the frequency of task sending kernel. If the rate of task sending kernel is high, it means that the task executes faster.

[0080] 2、GPU load sensing as follows:

[0081] Load-aware scheduling: The scheduler can intelligently select idle or lightly loaded GPUs for task allocation by obtaining GPU load information, avoiding overloading GPUs in high occupancy state for a long time.

[0082] Real-time dynamic adjustment: The scheduler can dynamically adjust the task allocation strategy according to real-time monitoring data to ensure reasonable utilization of each GPU resource and reduce computing bottlenecks.

[0083] Six, multi-task GPU sharing.

[0084] Here, multi-task GPU sharing and scheduling are described in the dimension of a single GPU.

[0085] 1、Task division and sharing: tasks can be divided into multiple operator layers or sub-task layers.

[0086] Request is the demand for an operation, usually initiated by a client and proposed to the server. A request can contain multiple tasks or a single task call. Task usually refers to a specific operation or computing unit that needs to be executed. For example, in GPU computing, a task may be the execution of a CUDA kernel, an image processing operation, a deep learning inference task, etc.

[0087] 2. Operator-level scheduling: The scheduler can dynamically decide the execution order of operators based on their resource requirements, optimizing resource sharing. For example, compute-intensive operators can be executed in parallel with memory-intensive operators to improve GPU throughput. Compute-intensive operators refer to operators that require a large amount of computational resources during execution. These operators typically involve complex mathematical operations, algorithmic processing, or data processing; memory-intensive operators refer to operators that require a large amount of memory resources during execution. These operators typically involve large-scale data storage, access, and transmission,

[0088] 3. Sub-task scheduling: Many AI tasks can be divided into multiple sub-tasks (such as DNN forward and backward propagation, etc.). By reasonably arranging the execution order of these sub-tasks, the overall throughput can be maximized while meeting delay constraints. For example, the scheduler can ensure that tasks can be cooperatively scheduled and executed by handling the dependencies between tasks. For example, the scheduler can improve GPU utilization by combining multiple sub-tasks into a batch for GPU computation through batch processing technology.

[0089] 4. Invasive compilation scheduling: Invasive modification of CUDA code for fine-grained control of task resource allocation. By binding the computational logic of tasks to GPU thread blocks, the execution time and resource occupancy of each operator on GPU cores can be directly controlled. This method can avoid the performance loss caused by frequent operator switching in traditional scheduling methods.

[0090] Invasive modification usually includes the following changes to CUDA code:

[0091] Operator task granularity control: In traditional CUDA scheduling, tasks are usually scheduled and executed in units of thread blocks. Invasive modification controls task granularity to accurately schedule the use of each thread block, thread, and even register resources.

[0092] Operator task and GPU core binding: In invasive compilation scheduling, operator tasks are refined to specific GPU cores for execution, rather than having the scheduler allocate tasks to any core by default. This approach helps reduce the frequency of task switching, as each task can maintain execution on a designated core, avoiding thread scheduling overhead.

[0093] The three ways to modify are as follows:

[0094] Code writing time: Developers need to explicitly define the granularity of tasks, memory access patterns, and dependencies between tasks when writing CUDA programs, so that reasonable resource allocation can be made in the subsequent scheduling stage.

[0095] Compile-time: The compiler can adjust thread blocks, memory layout, task scheduling, and other parameters using compile-time optimization strategies. This usually requires developers to provide additional parameters at compile time or use some specific compilation options (such as NVIDIA's nvcc compiler options).

[0096] Runtime: The runtime scheduler dynamically adjusts the execution strategy based on the state of the GPU, the dependencies of the tasks, the characteristics of the input data, and other factors.

[0097] 5. Fine-grained resource allocation: Based on GPU resource monitoring, the scheduler precisely controls the GPU resource occupancy of each operator to ensure resource isolation and predictability between different operators. Control strategies include:

[0098] (1) Resource allocation in GPU time and space dimensions: Based on the execution time and space requirements of operator tasks, dynamically allocate GPU time slices or space to ensure that the resource occupancy of each operator task does not exceed its requirements. Dynamic allocation refers to analyzing the execution time and space requirements of a task when it needs to be executed, and allocating GPU time slices or space in real time.

[0099] (2) Task scheduling and resource isolation: Through flow control or **Multi-Process Service (MPS)** technology, multiple operator tasks can be executed in parallel on the same GPU, while ensuring performance isolation between tasks.

[0100] Seven, multi-GPU load balancing.

[0101] Here, multi-task load balancing includes GPU task distribution and task dynamic migration.

[0102] 1. Multi-GPU load balancing: In a multi-GPU environment, the scheduler needs to monitor the load of each GPU in real time. Through load balancing strategies, tasks are distributed to GPUs with lower loads to avoid computational bottlenecks caused by some GPUs being overloaded.

[0103] 2. Task dynamic migration:

[0104] When detecting that a GPU is overloaded or has failed, the scheduler supports task migration. During task migration, the consistency of task execution and the consistency of data need to be ensured. When detecting that multiple tasks on a machine are competing for resources on the same GPU, the scheduler will perform task migration. Migration operations include:

[0105] GPU memory transfer: To ensure the consistency of task execution after migration, it is necessary to ensure the correct transfer of GPU memory data. The address of GPU memory may change, so an intermediate layer (located in the scheduler, in the task dynamic migration in Figure 1) is introduced to maintain the mapping relationship between the virtual address and the actual address of GPU memory.

[0106] Context synchronization: During migration, a new context (CUDA context is the execution environment or workspace of CUDA program on GPU device, similar to process or thread context in CPU programming) is first created on the target GPU and synchronized with the original context. After synchronization is completed, the task can be seamlessly migrated to the new GPU for execution.

[0107] In the embodiments of the present application, the following beneficial technical effects can be achieved:

[0108] 1. Efficient asynchronous call: Through the asynchronous API request mechanism, the blocking caused by synchronous call is avoided, and the throughput and response speed of the system are improved. In this way, the computing power of GPU can be fully utilized, and the computing efficiency in a multi-user environment can be improved.

[0109] 2. Reduce network delay: Through the short request packaging technology, multiple small requests are combined into one large request for sending, reducing communication overhead and network delay.

[0110] 3. Fine-grained resource monitoring: Real-time and fine-grained GPU monitoring is provided, which can comprehensively track the specific execution of each task (such as Kernel execution rate, GPU memory usage, etc.), which helps to optimize resource scheduling.

[0111] 4. Maximize resource utilization: Through GPU resource virtualization, multi-task sharing and load balancing mechanism, multiple clients can share the same GPU resource, thereby improving the utilization rate of GPU resources and avoiding the idle of resources.

[0112] 5. High availability: Support stateful migration of tasks when GPU fails to ensure that the computing task will not be interrupted, and improve the stability and high availability of the system.

[0113] 6. Elastic computing power service: Provide on-demand elastic computing power service, so that users can dynamically allocate GPU resources according to their own computing needs, save costs and improve resource utilization efficiency.

[0114] The computing power allocation method provided in the embodiments of the present application can be executed by a first server, wherein the first server can be any computing node deployed in a local area network, used to receive a task request of a client, and distribute the task to a suitable second server according to a load balancing strategy. The second server can be located on the same physical device as the first server, or can be located on different devices, depending on the actual resource distribution and load condition. As shown in Figure 1A , the first server and the second server can be the servers 122 set in the system architecture layer 12.

[0115] Figure 1C A flowchart of a computing power allocation method provided in the embodiments of the present application is shown in Figure 1C , which can be implemented through the following steps:

[0116] Step S110, determining a second server for executing a target computing task based on a load balancing strategy among a plurality of servers, wherein the first server and the second server are the same or different, and the target computing task is obtained from a client in the same data sharing network as the server;

[0117] Here, as shown in Figure 1B , the first server acts as a local area network GPU computing power scheduler, used to receive a target computing task from a computing power demander 15, and select the most suitable second server (computing power supplier 17) to execute the task through a load balancing strategy. The load balancing strategy here refers to a resource scheduling mechanism, used to reasonably allocate tasks among a plurality of servers, so as to avoid some servers being overloaded while others being idle, thereby improving the efficiency and resource utilization of the overall system. The strategy usually distributes tasks based on real-time load information of each server (such as GPU core utilization, video memory usage, memory bandwidth, etc.). For example, if the current GPU load of a certain server is low, it can be preferentially selected as the execution target of the task. In this way, the task can be allocated to the most suitable server, thereby improving the overall response speed and resource utilization of the system.

[0118] In the implementation process, the load balancing strategy can be implemented in various ways, such as round robin method, weighted round robin method, minimum connection number method, etc. Among them, the weighted round robin method can assign different weights to each server according to its performance configuration (such as GPU model, core number, video memory size, etc.), so as to more accurately reflect its processing capacity. In addition, real-time monitoring data (such as GPU temperature, power consumption, etc.) can also be combined for dynamic adjustment to adapt to the changing system state.

[0119] Step S120, dividing the target computing task into a plurality of operator tasks;

[0120] Here, the target computing task is usually a relatively complex computing process. In order to improve execution efficiency and parallelism, it needs to be split into multiple smaller execution units, namely operator tasks. Operator tasks are the smallest executable units that are further divided in the computing task. Each operator task has clear resource requirements (such as computing power requirements and storage requirements) and dependencies (such as predecessor / successor tasks). By analyzing these characteristics, the execution order of tasks can be optimized and the possibility of parallel execution can be increased. For example, a deep neural network training task can be split into multiple operator tasks, including data preprocessing, forward propagation, backpropagation, parameter update, etc. Each operator task can be executed independently or in parallel with other tasks, thereby fully utilizing the parallel computing capabilities of the GPU.

[0121] During implementation, task partitioning can be based on the structure of the tasks themselves, or program analysis tools can be used to automatically identify dependencies and resource requirements between tasks. For example, for a computational task consisting of multiple CUDAKernel calls, each kernel call can be treated as an operator task and scheduled based on their execution order and dependencies. Furthermore, task granularity control strategies can be used to dynamically adjust the size of operator tasks to balance task parallelism and scheduling overhead.

[0122] Step S130: determining a first execution order of the plurality of operator tasks based on the resource requirements and dependency relationships of each operator task, wherein the resource requirements include at least computing power requirements and storage requirements, and the dependency relationships are used to characterize the execution order of the operator tasks;

[0123] Here, each operator task has specific resource requirements and dependencies. Resource requirements include computing power requirements (such as the computing power of the GPU) and storage requirements (such as the amount of video memory used), which determine the task's occupancy of hardware resources. Dependency refers to the execution order between two or more operator tasks. If the execution of one operator task must depend on the result of another task, there is a dependency between the two tasks. In task scheduling, the execution order of tasks must be determined according to the dependency relationship to ensure the correctness of the calculation. For example, during the training process of a deep learning model, forward propagation must be completed before backpropagation; otherwise, parameter updates cannot be performed.

[0124] During implementation, task scheduling algorithms (such as topological sorting and dynamic priority scheduling) can be used to determine the execution order of operator tasks. For example, tasks without dependencies can be executed in parallel, while tasks with dependencies must be executed sequentially. Furthermore, resource demand information can be combined to prioritize tasks with lower resource requirements to improve resource utilization.

[0125] Step S140, performing the plurality of operator tasks based on the first execution order to obtain execution results.

[0126] In implementation, each operator task can be executed in sequence according to the determined first execution order, and the corresponding execution result can be generated after each task is completed. The execution result can be an intermediate calculation result or a final output result, depending on the nature of the task and the application scenario. For example, in deep learning model training, each operator task can generate parameter update values, and these parameter update values will be finally aggregated as the final parameters of the model.

[0127] In implementation, the execution process can be carried out in a synchronous or asynchronous manner. Synchronous execution is suitable for tasks with complex dependency relationships and strict order requirements, while asynchronous execution is suitable for tasks with high parallelism and less dependency between tasks. In addition, task pipeline technology can be used to overlap the execution processes of multiple tasks to improve overall execution efficiency. For example, while one operator task is being executed, the input data for the next task can be prepared at the same time, thereby reducing the waiting time between tasks.

[0128] Step S150, returning the execution results to the client.

[0129] In implementation, when all operator tasks are completed, the first server will collect the execution results of each operator task and integrate these results into the final output result, which is returned to the client. The client can perform subsequent processing based on these results, such as model inference, result display, etc. Among them, the result return can be realized in various ways, such as directly returning the final result, returning the intermediate result for further processing by the client, or notifying the client that the task has been completed through a callback function.

[0130] In the embodiments of the present application, through load balancing strategy, task division and scheduling, resource demand analysis and parallel execution mechanism, efficient computing power sharing and resource management between multiple servers are realized. This method can significantly improve the utilization rate of GPU resources, reduce task execution delay, and enhance the high availability and stability of the system.

[0131] In some embodiments, the above step S120 "dividing the target computing task into a plurality of operator tasks" can be realized by the following steps as shown in Figure 2

[0132] Step S210, dividing the target computing task into a plurality of stage tasks based on different models scheduled to execute the target computing task;

[0133] ​Here, the model refers to a structured framework or abstract representation used to describe and process the target computing task, usually used to guide the formulation of task partitioning and scheduling strategies. Different models may differ depending on the task type, resource requirements, or performance goals. For example, deep learning models, graph neural network models, traditional numerical computation models, etc. can be used, and each model may differ in the way tasks are partitioned and scheduling logic. By partitioning tasks based on different models, different computing scenarios and resource conditions can be more flexibly adapted, thereby improving the rationality of task partitioning and the overall execution efficiency of the system.

[0134] The stage task refers to the process of dividing the entire target computing task into a set of relatively independent but dependent sub-tasks according to the structure or scheduling strategy of the model during task partitioning. Each stage task usually corresponds to a processing level or computing stage in the model, such as data preprocessing stage, feature extraction stage, model inference stage, etc.

[0135] In the implementation process, resource scheduling can be realized in stages according to the granularity of sub-tasks: taking text-to-image as an example, the first stage is through LLM coding language prompts, and the second stage is through diffusion model to generate images. The granularity of the two-stage task is inconsistent. The scheduling granularity of LLM coding is Token, while the scheduling granularity of diffusion model is denoising step. Differential scheduling in stages is beneficial to load balancing.

[0136] By dividing the target computing task into multiple stage tasks, more fine-grained parallelization and resource optimization allocation can be achieved in subsequent operator task partitioning, thereby improving the flexibility and efficiency of task execution.

[0137] Step S220, based on the operation type set in the model, each stage task is divided into a plurality of operator tasks.

[0138] Here, the operation type refers to the specific calculation operation category defined in the model, such as matrix multiplication, convolution operation, normalization operation, activation function, etc. Different operation types have different computational complexity and resource requirements, and may require different hardware support or optimization strategies in actual execution. By identifying the operation types contained in each stage task in the model, each stage task can be further refined into multiple operator tasks, so that each operator task can be independently executed and efficiently scheduled. This helps to fully utilize heterogeneous computing resources such as GPUs, improving the efficiency of parallel computing and resource utilization.

[0139] An operator task refers to the smallest executable unit composed of a specific operation type, usually corresponding to a computing operation or a group of related operations in the model. By decomposing the stage tasks into multiple operator tasks according to the operation types, more fine-grained task scheduling and resource allocation can be achieved. For example, a stage task may contain multiple convolution operations and pooling operations, which can be scheduled and executed as independent operator tasks. This not only improves the flexibility of task division, but also helps to achieve better load balancing and resource utilization in a multi-GPU or multi-node environment.

[0140] In the implementation process, there is a close relationship between step S210 and step S220. Step S210 first divides the target computing task into multiple stage tasks according to the model, providing a structural basis for subsequent refinement; step S220 further decomposes each stage task into multiple operator tasks according to the operation types defined in the model, thereby achieving more fine-grained task scheduling and resource allocation. This hierarchical division method can effectively improve the parallel computing capability and resource utilization of the system.

[0141] In the embodiments of the present application, by dividing the target computing task into multiple stage tasks based on different models, and further dividing each stage task into multiple operator tasks based on the operation types in the model, more refined task division and resource scheduling can be achieved. In this way, the adaptability and execution efficiency of task division can be improved, thereby optimizing the utilization of computing resources and further improving the overall performance and task completion speed of the system.

[0142] In some embodiments, the above computing power allocation method further includes the following steps:

[0143] Step S160, determining a second execution order of the multiple stage tasks based on the scheduling order of different models for executing the target computing task;

[0144] Here, the scheduling order refers to the execution order formed by the system after sorting the models according to their priority, resource demand or task dependency relationship when executing the target computing task. The determination of the scheduling order can be based on various factors, such as the computational complexity of the model, the required GPU resources, the dependency relationship between tasks, etc. By determining the scheduling order, the system can optimize the execution process of the tasks, reduce resource conflicts and improve overall execution efficiency.

[0145] The second execution order refers to a more refined task execution order generated in combination with the model scheduling order, taking into account the influence of the model scheduling order.

[0146] Correspondingly, the above step S140 "executing the multiple operator tasks based on the first execution order" can be implemented through the following process:

[0147] perform the plurality of operator tasks based on the first execution order and the second execution order.

[0148] Here, the first execution order refers to the execution order of operator tasks in the time dimension, usually determined by factors such as dependency relationship, priority, etc. The second execution order is the task execution order generated in combination with the model scheduling order. By considering both the first execution order and the second execution order, the system can take into account both the time order of the model and the rationality of model resource allocation when executing multiple operator tasks, thereby avoiding resource conflicts and task blocking.

[0149] In the implementation process, the second execution order can be generated by analyzing the model scheduling order to provide more accurate scheduling basis for subsequent task execution; in the task execution phase, the first execution order and the second execution order are combined to ensure that the task meets both the time logic and the resource allocation requirements, thereby realizing efficient and orderly computation task scheduling.

[0150] In the embodiments of the present application, by comprehensively considering the first execution order and the second execution order, the system can realize more intelligent task scheduling and resource management in a complex computing environment, effectively improving the task execution efficiency and the overall performance of the system.

[0151] In some embodiments, the above step S130 of "determining the first execution order of the plurality of operator tasks based on the resource requirements and dependency relationships of each operator task" can be implemented by the following steps: Figure 3

[0152] Step S310, in the case that the first operator task and the second operator task have a dependency relationship, determining the execution order of the first operator task and the second operator task based on the dependency relationship;

[0153] Here, the dependency relationship refers to the logical constraint between two operator tasks that one must be executed before the other. For example, if the output of task A is the input of task B, then task B must be executed after task A is completed. This dependency relationship can be data dependency (such as computation result transmission), control dependency (such as conditional judgment), or resource dependency (such as sharing the same GPU resource). The scheduler determines the execution order by analyzing the dependency relationship between tasks, thereby avoiding task conflicts or data inconsistency problems.

[0154] In actual application, when the client submits multiple AI computing tasks, the scheduler first analyzes the dependency relationship between these tasks. For example, in deep learning model training, forward propagation must be completed before backward propagation, so there is a clear dependency relationship between the two tasks. The scheduler will place the forward propagation task first to ensure its priority execution to meet the data requirements of subsequent tasks. ​

[0155] In this way, the correctness and efficiency of the overall task execution can be improved, thereby optimizing the utilization of GPU resources and reducing task waiting time.

[0156] In step S320, the third computing power requirement and the third storage requirement of the third operator task and the fourth computing power requirement and the fourth storage requirement of the fourth operator task are determined in the case where there is no dependency relationship between the third operator task and the fourth operator task.

[0157] Here, the computing power requirement refers to the computing power required to execute an operator task, usually measured in floating-point operations per second (FLOPS). The storage requirement refers to the size of the video memory required during task execution, usually expressed in MB or GB. When there is no dependency relationship between two tasks, the scheduler evaluates their computing power and storage requirements respectively to determine whether they can be executed in parallel.

[0158] For example, in an image processing task, the third operator task may be a convolution operation that requires high computing power but low storage, and the fourth operator task may be a normalization operation that requires low computing power but high storage. At this time, the scheduler will obtain the computing power and storage requirements of the third operator task and the fourth operator task 4 respectively, and make the next decision accordingly.

[0159] In this way, the system can have accurate knowledge of the resource requirements of each task, so that appropriate execution strategies can be allocated to the tasks, thereby improving the flexibility and accuracy of resource scheduling.

[0160] In step S330, the third operator task and the fourth operator task are determined to be executed in parallel in the case where the third computing power requirement is greater than the fourth computing power requirement and the third storage requirement is less than the fourth storage requirement.

[0161] Here, parallel execution refers to the simultaneous running of two or more tasks within the same time period without waiting for each other. When there is no dependency relationship between two tasks, and one task has high computing power requirement but low storage requirement, and the other task has low computing power requirement but high storage requirement, the scheduler can decide to execute them in parallel. This strategy fully utilizes the heterogeneous computing power of the GPU, so that high-computing-power tasks and high-storage tasks can share GPU resources without resource contention.

[0162] For example, the third operator task has a computing power requirement of 5000 FLOPS and a storage requirement of 200 MB, and the fourth operator task has a computing power requirement of 3000 FLOPS and a storage requirement of 50 MB. At this time, the scheduler determines that the third operator task and the fourth operator task can be executed in parallel on the same GPU because their resource usage patterns are complementary and do not cause resource bottlenecks.

[0163] In this way, the utilization of the GPU can be effectively improved, thereby reducing the execution time of the tasks and reducing the resource idle rate, and thus the overall throughput and response speed of the system can be enhanced.

[0164] In the implementation process, first, the scheduler determines the execution order of the tasks by analyzing the dependency relationship between the tasks, which is the basis of the entire scheduling process. Then, for tasks without dependency relationship, the scheduler further evaluates their computing power and storage requirements to provide a basis for whether to allow parallel execution. Finally, according to the comparison result of the resource requirements, the scheduler decides whether to execute the tasks in parallel, thereby realizing the optimal configuration of resources. The three steps together constitute the mechanism of task scheduling, ensuring the efficiency of task execution and the rational use of system resources.

[0165] In the embodiments of the present application, the execution order is determined based on the dependency relationship, and whether to execute in parallel is decided according to the computing power and storage requirements. In this way, the execution order of the tasks and the allocation of resources can be reasonably arranged, thereby avoiding task conflicts and resource waste, and thus the execution efficiency and resource utilization of the system can be improved.

[0166] In some embodiments, the above step S330 "determining that the third operator task and the fourth operator task are executed in parallel in the case that the third computing power requirement is greater than the fourth computing power requirement and the third storage requirement is less than the fourth storage requirement" can be implemented by the following steps:

[0167] Step 331, determining that the target thread block of the image processor in the second server can simultaneously satisfy the third computing power requirement and the fourth computing power requirement, and the third computing power requirement is greater than the fourth computing power requirement and the third storage requirement is less than the fourth storage requirement;

[0168] Here, the target thread block refers to a group of thread resources in the image processor for executing computing tasks. Each thread block can run independently and has a certain computing power and memory access capability. The selection of the target thread block is based on whether it has sufficient computing power and storage capacity to support the execution of two operator tasks simultaneously. For example, in the NVIDIA GPU architecture, a thread block is usually composed of multiple threads that share the same set of registers and shared memory. By evaluating the computing power (such as FLOPS) and storage capacity (such as available memory) of the target thread block, the system can determine whether the thread block is suitable for parallel execution of two tasks.

[0169] In the implementation process, the system first evaluates whether the computing power of the target thread block can simultaneously meet the requirements of the two operator tasks. For example, if the computing power requirement of the first task (third operator task) is greater than that of the second task (fourth operator task), but its storage requirement is lower, then the target thread block only needs to have sufficient computing power. In this case, the system can preferentially select a thread block with high computing power to ensure that both tasks can be executed efficiently, while avoiding task blocking caused by storage bottlenecks.

[0170] Step 332, the third operator task and the fourth operator task are allocated to the target thread block for parallel execution.

[0171] Here, the two operator tasks are allocated to the same target thread block for parallel execution, that is, the two tasks will run concurrently on the same group of thread resources. This approach can significantly improve the utilization of the GPU, especially in multi-task scenarios. By reasonably arranging the execution order of the tasks and resource allocation, the system can achieve higher throughput and lower latency.

[0172] In the implementation process, for example, in the process of deep learning model training or inference, different operator tasks may have different requirements for computing power and storage. By pairing tasks with high computing power requirements but low storage requirements with tasks with low computing power requirements but high storage requirements, the system can improve overall performance without adding additional hardware resources. In addition, since the two tasks share the resources of the same thread block, the overhead caused by thread switching can also be reduced, further optimizing execution efficiency.

[0173] In the embodiments of the present application, by introducing the concept of target thread block, the task allocation is only performed when the target thread block meets the computing power and storage conditions. This ensures the safety and effectiveness of resource scheduling and avoids task failure or performance degradation due to insufficient resources. According to the computing power and storage requirements, the GPU resources can be more flexibly scheduled, thereby improving the utilization of the GPU and the efficiency of task execution, and further enabling more efficient AI computing power sharing and resource management.

[0174] In some embodiments, the above performing the third operator task and the fourth operator task further comprises the following steps:

[0175] Step 333, binding the third operator task and the fourth operator task to different computing units in the target thread block respectively;

[0176] Or,

[0177] Here, the third operator task and the fourth operator task are bound to different computing units in the target thread block respectively, that is, within the same thread block, two operator tasks are allocated to different computing units for running to achieve parallel execution. This binding method can avoid resource competition between operator tasks and improve the utilization of GPU. For example, a thread block contains multiple streaming multiprocessors (SMs), each SM can be regarded as a computing unit, and the system can allocate operator tasks to different SMs for running according to the types and resource requirements of the operator tasks. In this way, two tasks can be executed in parallel within the same thread block without interfering with each other, thereby improving the overall computing efficiency.

[0178] Step 334, allocating time slices and space of the image processor based on the execution time and space requirements corresponding to the third operator task and the fourth operator task respectively.

[0179] Here, the time slices and space of the image processor are allocated based on the execution time and space requirements corresponding to the third operator task and the fourth operator task respectively, that is, the scheduler can dynamically allocate time slices and space according to the resource requirements of the tasks. For example, if one task needs more video memory but has small computing amount, and another task is computationally intensive but has low video memory requirement, the scheduler can allocate more video memory space for the former and more time slices for the latter.

[0180] In some embodiments, a task can be executed on multiple time slices, and another task can be inserted to execute on other time slices.

[0181] In this way, it can ensure that both tasks can efficiently run within the required computing resources and time resources, avoiding resource waste or task blocking. In this way, the system can more flexibly adapt to the resource requirements of different tasks, improving the overall throughput and response speed of the GPU.

[0182] In the embodiments of the present application, by binding the third operator task and the fourth operator task to different computing units in the target thread block respectively, or dynamically allocating time slices and space of the image processor based on the execution time and space requirements, resource conflicts between tasks can be effectively avoided, thereby improving the parallel processing capability of the GPU, and further improving the overall performance and resource utilization of the system.

[0183] In some embodiments, the step of "determining a second service end for executing the target computing task based on a load balancing strategy among multiple service ends" in step S110 above can be implemented by the following steps:

[0184] Step 111, obtaining load information corresponding to each service end; wherein the load information includes at least one of the following: kernel utilization, GPU memory usage, memory bandwidth;

[0185] Here, the load information refers to a set of indicators for measuring the current working state and resource occupation of the GPU or computing device, which can reflect the real-time performance of the service end when executing the computing task. The collection and analysis of load information is the basis for efficient load balancing scheduling. Load information includes the following dimensions:

[0186] Kernel utilization (Kernel Utilization), which represents the proportion of time spent executing computing kernels (Kernel) on the GPU, reflecting whether the GPU's computing unit is in a high-load running state. The higher the kernel utilization, the more parallel computing tasks the GPU is executing, and it may face resource bottlenecks.

[0187] GPU memory usage (GPU Memory Usage), which represents the proportion of GPU memory currently in use, used to determine whether the GPU is at risk of running out of memory when processing large-scale data tasks. High GPU memory usage can cause task execution delays or failures.

[0188] Memory bandwidth (Memory Bandwidth), which represents the ability of the GPU to read and write memory in a unit of time, is an important indicator of the efficiency of GPU data transmission. The level of memory bandwidth directly affects the execution speed and throughput of tasks.

[0189] These indicators are usually collected in real time by underlying monitoring tools (such as NVIDIA's NVML, DCGM, etc.) and managed uniformly by the scheduler to ensure that tasks can be reasonably allocated to service ends with lower loads. By obtaining and analyzing these fine-grained load information, the system can avoid assigning new tasks to already overloaded GPUs, thereby improving the overall system's resource utilization and task response speed.

[0190] Step 112, determining the service end with the lowest load as the second service end based on the load information corresponding to each service end.

[0191] In the implementation process, for example, in an AI computing power sharing platform within a local area network, when multiple users submit computing tasks, the scheduler queries the load information of each GPU node in real time. If the kernel utilization rate of a certain GPU node is close to 100%, and the video memory usage rate is also more than 90%, it indicates that the GPU can no longer accept new tasks, at this time the scheduler allocates new tasks to other GPUs with lower load to execute, thereby achieving dynamic load balancing.

[0192] In the embodiments of the present application, by obtaining the load information of each service end, the available computing capacity of each service end can be accurately evaluated. In this way, the task can be avoided to be allocated to the GPU which is already full, so as to optimize the execution efficiency of the task, and further improve the resource utilization and task completion rate of the whole system.

[0193] In some embodiments, the method for allocating computing power further comprises the following steps:

[0194] Step S160, monitoring the running state of the second server end;

[0195] Here, the second server end refers to a GPU resource node deployed in a local area network, which provides computing resources for clients to call. The second server end usually includes one or more GPU devices and a scheduler and resource management module running on these devices. The server end not only receives computing task requests from clients, but also undertakes functions such as task scheduling, resource allocation, and load monitoring.

[0196] The "running state" refers to the current resource usage and health state of the second server end, including but not limited to CPU usage, memory occupancy, GPU video memory usage, GPU utilization, network connection state, whether the service process is running normally, etc. By monitoring these indicators, potential problems such as GPU overload, insufficient video memory, service abnormalities, etc. can be found in time, so that corresponding adjustments and optimizations can be made.

[0197] In the implementation process, the running state of the second server end can be monitored in multiple ways. For example, by using the management library (NVIDIA Management Library, NVML) interface provided by NVIDIA, detailed performance indicators of each GPU can be obtained in real time; through system-level monitoring tools (such as Prometheus, Grafana, etc.), comprehensive monitoring of the entire server can be realized; custom log analysis and alarm mechanisms can also be combined to control thresholds and detect abnormalities.

[0198] Step S170, in the case that the running state shows a fault, migrating the target computing task to a third service end for processing;

[0199] The third server and the second server are located in the same data sharing network, and the third server is different from the second server.

[0200] In the implementation process, when it is detected that a certain GPU is overloaded or fails, the scheduler supports task migration. During task migration, the consistency of task execution and the consistency of data need to be ensured.

[0201] Or,

[0202] Step S180, obtaining a virtual address corresponding to each of the servers, so that the first server can access the plurality of servers based on the virtual address;

[0203] Correspondingly, the above step S150 "returning the execution result to the client" can be realized by the following process:

[0204] Returning the execution result and the virtual address to the client.

[0205] Here, the GPU resource virtualization technology can virtualize remote physical GPU resources as logical resources of the local system. Each client request can access remote GPUs through a virtual address space without worrying about the physical location of the GPU. Here, the first server can map remote GPUs as virtual GPU address spaces that can be recognized by the local system (client), that is, when the local system accesses these virtual addresses, it can recognize the remote GPU, similar to the local system virtually mounting a GPU, but this GPU belongs to the remote server.

[0206] In the embodiment of the application, stateful migration of tasks is supported when the GPU fails, ensuring that the computing task will not be interrupted and improving the stability and high availability of the system. The first client interacts with the remote GPU through a virtualized interface, similar to executing a task on a local GPU, that is, the client requests are executed on the virtual resources mapped by the client. In this way, remote GPU resources can be shared among multiple clients, thereby effectively optimizing the utilization of GPU resources.

[0207] The application embodiment provides a computing power allocation method, which is applied to a client, as shown in the figure, which can be realized by the following steps: Figure 4

[0208] Step S410, obtaining a plurality of computing requests corresponding to a plurality of computing tasks;

[0209] Here, the client combines multiple small requests into a request package after receiving them. The request package contains multiple small CUDA operations or inference tasks. After the remote GPU executes the batch request, it returns a unified result to the client.​

[0210] Step S420, dividing the plurality of computing requests into a plurality of computing request groups based on task properties of the computing task;

[0211] Here, the request grouping: the scheduler or the client can group tasks according to the properties of the tasks (e.g., the computing type of the tasks, data dependency, etc.) after receiving a plurality of short requests, divide the short requests into several logical groups, and these groups usually contain a plurality of small CUDA operations, memory copying, or inference requests.

[0212] Step S430, dividing each of the computing request groups into a plurality of packed requests based on a packed data amount threshold and a request format;

[0213] Here, the packed data amount threshold can be determined based on the bandwidth of the transmitted data. A standardized request format can be used, which contains information such as input data, task type, parameters, and computing order of all small tasks.

[0214] In the implementation process, the plurality of small requests can be combined into a batch request, which can be a plurality of independent CUDA kernel calls, data transmission (such as cudaMemcpy) or inference requests combined into a "package" based on the packed data amount threshold and the request format. For example, the parameters of a plurality of CUDA kernels can be packed into a data structure and executed sequentially on the GPU side.

[0215] Step S440, sending a plurality of packed requests to a first server, wherein the target computing task is included in the packed request.

[0216] In the embodiments of the present application, the short request packing is to reduce the frequent communication overhead, and a plurality of small requests are packed into a batch request to reduce the communication overhead of each request.

[0217] In some embodiments, the above step S440 "sending a plurality of packed requests to a first server" can be implemented by the following process:

[0218] The plurality of packed requests are sent to the first server using an asynchronous request mechanism, and the asynchronous request mechanism can be implemented based on placeholders and / or delay binding.

[0219] Here, in order to improve the response speed and throughput of the system, the asynchronous API request mechanism is used to avoid the delay and blocking caused by synchronous calls. Especially when a plurality of small tasks (such as memory allocation, Kernel startup, etc.) occur, the asynchronous request can significantly reduce the waiting time.

[0220] The placeholder and the deferred binding in the asynchronous request mechanism are two independent implementation manners, and one of them or a combination of them can be selected according to specific scenes and requirements:

[0221] Independence: The placeholder and the deferred binding are two different implementation ideas, and they can exist independently. In some simple asynchronous request scenes, only the placeholder can be used to fill data; and in more complex scenes, the deferred binding can be more suitable for processing data.

[0222] Combination: Although they are in an "or" relationship, they can also be combined in actual application. For example, the placeholder can be used to reserve the position of data display on the page first, and then the data is filled into the placeholder through the deferred binding after the asynchronous request is completed, and other related processing operations are performed at the same time.

[0223] In the embodiment of the application, when sending a task request, the client does not wait for the response of the remote GPU, but wraps the request into an asynchronous task and sends it out. The client then continues to perform other tasks. When the GPU finishes processing and returns the result, the server notifies the client through an asynchronous callback mechanism, or obtains the result through a polling method. In this way, the "idle waiting" time in the GPU request can be avoided, and the total execution time of the task can be reduced.

[0224] In some embodiments, the first client can obtain a virtual address corresponding to each server in the same data sharing network as the client, so that the client can access the plurality of servers based on the virtual address. In this way, the remote GPU resources can be shared among multiple clients, thereby effectively optimizing the utilization of GPU resources.

[0225] Based on the foregoing embodiments, the application provides a computing power allocation device, which includes various modules, each module includes various sub-modules, each sub-module includes a unit, and can be implemented by a processor in an electronic device. Of course, it can also be implemented by a specific logic circuit. In the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA).

[0226] Figure 5A The composition structure diagram of the computing power allocation device provided in the embodiment of the application is shown as Figure 5A The device 500 includes:

[0227] The first determining module 501 is configured to determine, based on a load balancing strategy, a second server for executing a target computing task from a plurality of servers, wherein the first server and the second server are the same or different, and the target computing task is obtained from a client in a same data sharing network as the servers.

[0228] The task division module 502 is configured to divide the target computing task into a plurality of operator tasks.

[0229] The second determining module 503 is configured to determine a first execution order of the plurality of operator tasks based on resource requirements and dependency relationships of each operator task, wherein the resource requirements at least include computing power requirements and storage requirements, and the dependency relationships are used to represent the execution order of the operator tasks.

[0230] The task execution module 504 is configured to execute the plurality of operator tasks based on the first execution order to obtain an execution result.

[0231] The result returning module 505 is configured to return the execution result to the client.

[0232] In some embodiments, the task division module 502 includes a first division sub-module and a second division sub-module, wherein the first division sub-module is configured to divide the target computing task into a plurality of stage tasks based on different models for scheduling the target computing task; and the second division sub-module is configured to divide each stage task into a plurality of operator tasks based on an operation type set in the model.

[0233] In some embodiments, the computing power allocation apparatus further includes a third determining module configured to determine a second execution order of the plurality of stage tasks based on a scheduling order of different models for executing the target computing task; and correspondingly, the task execution module 504 is further configured to execute the plurality of operator tasks based on the first execution order and the second execution order.

[0234] In some embodiments, the second determining module 503 comprises a first determining submodule, a second determining submodule and a third determining submodule, wherein the first determining submodule is configured to determine the execution order of the first operator task and the second operator task based on the dependency relationship when the first operator task and the second operator task have a dependency relationship; the second determining submodule is configured to determine the third computing power requirement and the third storage requirement of the third operator task and the fourth computing power requirement and the fourth storage requirement of the fourth operator task when the third operator task and the fourth operator task do not have a dependency relationship; and the third determining submodule is configured to determine that the third operator task and the fourth operator task are executed in parallel when the third computing power requirement is greater than the fourth computing power requirement and the third storage requirement is less than the fourth storage requirement.

[0235] In some embodiments, the third determining submodule comprises a determining unit and an allocating unit, wherein the determining unit is configured to determine that a target thread block of an image processor in the second server can simultaneously meet the third computing power requirement and the fourth computing power requirement, and the third computing power requirement is greater than the fourth computing power requirement and the third storage requirement is less than the fourth storage requirement; and the allocating unit is configured to allocate the third operator task and the fourth operator task to the target thread block for parallel execution.

[0236] In some embodiments, the third determining submodule further comprises a binding unit or an allocating unit, wherein the binding unit is configured to bind the third operator task and the fourth operator task to different computing units in the target thread block, respectively; and the allocating unit is configured to allocate time slices and spaces of the image processor based on the execution time and space requirements corresponding to the third operator task and the fourth operator task, respectively.

[0237] In some embodiments, the first determining module 501 comprises an obtaining submodule and a fourth determining submodule, wherein the obtaining submodule is configured to obtain load information corresponding to each of the servers; and the load information comprises at least one of the following: kernel utilization rate, video memory usage rate and memory bandwidth; and the fourth determining submodule is configured to determine a server with the lowest load as the second server based on the load information corresponding to each of the servers.

[0238] In some embodiments, the computing power allocation apparatus further comprises a monitoring module, a fourth determination module, or a first acquisition module, wherein the monitoring module is configured to monitor a running state of the second server end; the fourth determination module is configured to determine, in a case where the running state shows a fault, to migrate the target computing task to a third server end for processing; the first acquisition module is configured to acquire a virtual address corresponding to each of the server ends, so that the first server end can access the plurality of server ends based on the virtual address; correspondingly, the result returning module 505 is further configured to return the execution result and the virtual address to the client end.

[0239] Figure 5B A schematic diagram of the composition structure of the computing power allocation apparatus provided by the embodiments of the present application is shown in FIG. 5, which comprises: Figure 5B

[0240] A second acquisition module 511 is configured to acquire a plurality of computing requests corresponding to a plurality of computing tasks.

[0241] A first division module 512 is configured to divide the plurality of computing requests into a plurality of computing request groups based on the task properties of the computing tasks.

[0242] A second division module 513 is configured to divide each of the computing request groups into a plurality of packaged requests based on a packaged data volume threshold and a request format.

[0243] A sending module 514 is configured to send the plurality of packaged requests to a first server end, wherein the packaged requests include target computing tasks.

[0244] In some embodiments, the sending module 514 is further configured to send the plurality of packaged requests to the first server end by using an asynchronous request mechanism, and the asynchronous request mechanism can be implemented based on a placeholder and / or a delay binding.

[0245] The above description of the apparatus embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the apparatus embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0246] ​It should be noted that, in the embodiments of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application can be embodied in the form of a software product in essence or the part that contributes to the related art, and the computer software product is stored in a storage medium, including a plurality of instructions for causing an electronic device (which can be a mobile phone, a tablet computer, a notebook computer, a desktop computer, etc.) to execute all or part of the method described in the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (Read Only Memory, ROM), a magnetic disk or an optical disk, and various media that can store program codes. Thus, the embodiments of the present application are not limited to any specific hardware and software combination.

[0247] Correspondingly, the embodiments of the present application provide a storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the computing power allocation method provided in the above embodiments.

[0248] Correspondingly, the embodiments of the present application provide an electronic device, Figure 6 A hardware entity schematic diagram of the electronic device provided in the embodiments of the present application is shown in FIG. 6, which includes a memory 601 and a processor 602. The memory 601 stores a computer program executable on the processor 602, and the processor 602 implements the steps of the computing power allocation method provided in the above embodiments when executing the program. Figure 6

[0249] The memory 601 is configured to store instructions and applications executable by the processor 602, and can also cache data (for example, image data, audio data, voice communication data and video communication data) to be processed by the processor 602 and each module in the electronic device 600. It can be realized by FLASH or Random Access Memory (RAM).

[0250] It should be noted that: the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects to the method embodiments. For technical details not disclosed in the storage medium and device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0251] ​It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.

[0252] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0253] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0254] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.

[0255] In addition, all functional units in the embodiments of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the above-mentioned integrated units can be implemented in the form of hardware or in the form of hardware plus software functional units.

[0256] Those skilled in the art can understand that all or part of the steps of the above-mentioned method embodiments can be completed by program instruction related hardware, and the foregoing program can be stored in a computer readable storage medium, and the program performs the steps of the above-mentioned method embodiments when executed; and the foregoing storage medium includes a mobile storage device, a read only memory (ROM), a magnetic disc or an optical disc, and various storage medium that can store program codes.

[0257] Alternatively, the integrated units in the above embodiments of the present application, if realized in the form of software function modules and sold or used as independent products, can also be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, and includes several instructions for making an electronic device (which can be a mobile phone, a tablet computer, a notebook computer, a desktop computer, etc.) execute all or part of the methods described in the embodiments of the present application. The foregoing storage medium includes a mobile storage device, a ROM, a magnetic disc or an optical disc, and various storage medium that can store program codes.

[0258] The methods disclosed in the several method embodiments provided by the present application can be combined arbitrarily without conflict to obtain new method embodiments.

[0259] The features disclosed in the several product embodiments provided by the present application can be combined arbitrarily without conflict to obtain new product embodiments.

[0260] The features disclosed in the several method or device embodiments provided by the present application can be combined arbitrarily without conflict to obtain new method embodiments or device embodiments.

[0261] The above is only the implementation manner of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed in the present application, and all of them should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A computing power allocation method, applied to a first server, comprising: Determining, based on a load balancing strategy, a second server end from a plurality of server ends to execute a target computing task, wherein the first server end and the second server end are the same or different, and the target computing task is obtained from a client end that is arranged in the same data sharing network as the server end; Dividing the target computing task into a plurality of operator tasks; Determining a first execution order of the plurality of operator tasks based on resource requirements and dependency relationships of each of the operator tasks, wherein the resource requirements include at least computing power requirements and storage requirements, and the dependency relationships are used to characterize the execution order of the operator tasks; Executing the plurality of operator tasks based on the first execution order to obtain an execution result; The execution result is returned to the client.

2. The method according to claim 1, wherein dividing the target computing task into a plurality of operator tasks comprises: Dividing the target computing task into multiple stage tasks based on different models for executing the target computing task scheduling; Each of the stage tasks is divided into a plurality of operator tasks based on the operation type set in the model.

3. The method of claim 2, further comprising: Determining a second execution order of the plurality of stage tasks based on a scheduling order of executing the target computing task for different models; Correspondingly, executing the multiple operator tasks based on the first execution order includes: The plurality of operator tasks are executed based on the first execution order and the second execution order.

4. The method according to any one of claims 1 to 3, wherein determining the first execution order of the plurality of operator tasks based on the resource requirements and dependencies of each operator task comprises: When it is determined that a dependency relationship exists between the first operator task and the second operator task, determining an execution order of the first operator task and the second operator task based on the dependency relationship; When it is determined that there is no dependency relationship between the third operator task and the fourth operator task, determine the third computing power requirement and the third storage requirement of the third operator task, and the fourth computing power requirement and the fourth storage requirement of the fourth operator task; When it is determined that the third computing power requirement is greater than the fourth computing power requirement and the third storage requirement is less than the fourth storage requirement, it is determined that the third operator task and the fourth operator task are executed in parallel.

5. The method of claim 4, wherein determining that the third computing power requirement is greater than the fourth computing power requirement and the third storage requirement is less than the fourth storage requirement, determining that the third operator task and the fourth operator task are executed in parallel comprises: Determining that a target thread block of the graphics processor in the second server can simultaneously meet the third computing power requirement and the fourth computing power requirement, the third computing power requirement is greater than the fourth computing power requirement, and the third storage requirement is less than the fourth storage requirement; Allocate the third operator task and the fourth operator task to the target thread block for parallel execution.

6. The method of claim 5, further comprising: Binding the third operator task and the fourth operator task to different computing units in the target thread block respectively; or, The time slice and space of the image processor are allocated based on the execution time and space requirements corresponding to the third operator task and the fourth operator task respectively.

7. The method according to any one of claims 1 to 3, wherein determining a second server to execute the target computing task from a plurality of server terminals based on a load balancing strategy comprises: Obtaining load information corresponding to each of the servers; wherein the load information includes at least one of the following: core utilization, video memory utilization, and memory bandwidth; Based on the load information corresponding to each of the servers, the server with the lowest load is determined as the second server.

8. The method according to any one of claims 1 to 3, further comprising: Monitoring the running status of the second server; When it is determined that the running status shows a fault, migrating the target computing task to a third server for processing; The third server and the second server are set in the same data sharing network, and the third server is different from the second server; or, Obtaining a virtual address corresponding to each of the servers, so that the first server can access the multiple servers based on the virtual address; Correspondingly, returning the execution result to the client includes: The execution result and the virtual address are returned to the client.

9. A computing power allocation method, applied to a client, comprising: Get multiple computing requests corresponding to multiple computing tasks; dividing the plurality of computing requests into a plurality of computing request groups based on the task properties of the computing tasks; dividing each computing request group into a plurality of packaged requests based on a packaged data amount threshold and a request format; Sending multiple packaging requests to the first server, wherein the packaging requests include target computing tasks.

10. A computing power distribution device, comprising: a first determining module, configured to determine, based on a load balancing strategy, a second server end from a plurality of server ends to execute a target computing task, wherein the first server end and the second server end are the same or different, and the target computing task is obtained from a client end that is provided in the same data sharing network as the server end; The task division module is used to divide the target computing task into multiple operator tasks; a second determining module, configured to determine a first execution order of the plurality of operator tasks based on resource requirements and dependency relationships of each operator task, wherein the resource requirements include at least computing power requirements and storage requirements, and the dependency relationships are used to characterize the execution order of the operator tasks; A task execution module, configured to execute multiple operator tasks based on the first execution order and obtain execution results; The result return module is used to return the execution result to the client.